Volume XXII, number 279Tuesday, October 6, 2026Latest message 44 minutes ago

The Git List

News and archive of git@vger.kernel.org, since April 2005

patchindex-pack: speed up promisor link recording

9 messages between Aug 2, 2026 and Aug 3, 2026, from Arijit Banerjee via GitGitGadget, brian m. carlson, Arijit Banerjee, Junio C Hamano, Collin Funk.

Plain Markdown or JSON for tools and agents. Diffs are folded; open one to read it.

Arijit Banerjee via GitGitGadgetAug 2, 2026, 21:33 UTC on lore
From: Arijit Banerjee <arijit@effectiveailabs.com>

When indexing a promisor pack, index-pack parses every reconstructed non-blob object into the shared object model to record its outgoing links. Since parse_object_buffer() runs under read_mutex, worker threads serialize while allocating persistent tree, commit, and tag structures that are only needed to enumerate those links.

Read the links directly from the reconstructed object buffers instead. Keep the strict and fsck paths unchanged, use worker-local typed oidmaps during normal promisor indexing, and merge them after the workers exit. Transfer entries during the merge so that it does not temporarily duplicate the complete link set.

The typed entries preserve checks previously performed as a side effect of object parsing. Reject malformed commit and tag headers, conflicting expected types, and targets whose actual type disagrees when the target is present in the pack. Preserve commit-graft handling and the existing policy of recording only subtree entries from trees.

With three runs per version on Debian 12, median end-to-end wall-clock time for a --filter=blob:none clone of linux.git decreased from 156 seconds to 133 seconds (15%). Trace2 attributed the change to the initial index-pack --promisor phase, whose median duration decreased from 121 seconds to 98 seconds (19%). System CPU time decreased by 46%.

Two paired spot checks against GitHub showed end-to-end reductions of 18% and 26%. These measurements include network and server variability and are therefore corroborating rather than controlled results. A third pair was not interpretable because the baseline request encountered a transport stall.

A full-clone control showed no material change, taking approximately 256 seconds with either version. This is expected because full clones do not exercise promisor-link recording.

t5302-pack-index.sh passed with both SHA-1 and SHA-256, while t0410-partial-clone.sh and t5616-partial-clone.sh also passed. New coverage checks malformed commit headers, conflicting link types, and mismatched tag target types.

Signed-off-by: Arijit Banerjee <arijit@effectiveailabs.com>
---
    index-pack: speed up promisor link recording
    
    AI assistance: OpenAI Codex was used to identify the bottleneck and
    assist with the implementation, testing, and benchmark analysis. I
    reviewed the resulting change and take responsibility for this
    submission.
Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-2191%2Farijit91%2Findex-pack-promisor-link-recording-v1
Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2191/arijit91/index-pack-promisor-link-recording-v1
Pull-Request: https://github.com/gitgitgadget/git/pull/2191
 builtin/index-pack.c  | 234 ++++++++++++++++++++++++++++++++++++++----
 t/t5302-pack-index.sh |  50 +++++++++
 2 files changed, 265 insertions(+), 19 deletions(-)
Show changes to 2 files +265 −19

builtin/index-pack.c, t/t5302-pack-index.sh

diff --git a/builtin/index-pack.c b/builtin/index-pack.c
index bc86925ad0..a973146757 100644
--- a/builtin/index-pack.c
+++ b/builtin/index-pack.c
@@ -23,7 +23,8 @@
 #include "odb.h"
 #include "odb/streaming.h"
 #include "oid-array.h"
-#include "oidset.h"
+#include "hash-lookup.h"
+#include "oidmap.h"
 #include "path.h"
 #include "replace-object.h"
 #include "tree-walk.h"
@@ -105,6 +106,12 @@ static size_t base_cache_limit;
 struct thread_local_data {
 	pthread_t thread;
 	int pack_fd;
+	struct oidmap outgoing_links;
+};
+
+struct outgoing_link {
+	struct oidmap_entry entry;
+	enum object_type type;
 };
 
 /* Remember to update object flag allocation in object.h */
@@ -155,11 +162,8 @@ static uint32_t input_crc32;
 static int input_fd, output_fd;
 static const char *curr_pack;
 
-/*
- * outgoing_links is guarded by read_mutex, and record_outgoing_links is
- * read-only in a thread.
- */
-static struct oidset outgoing_links = OIDSET_INIT;
+/* Worker-local maps are merged after all workers have exited. */
+static struct oidmap outgoing_links = OIDMAP_INIT;
 static int record_outgoing_links;
 
 static struct thread_local_data *thread_data;
@@ -196,6 +200,55 @@ static inline void unlock_mutex(pthread_mutex_t *mutex)
 		pthread_mutex_unlock(mutex);
 }
 
+static void record_outgoing_link_to(struct oidmap *map,
+				    const struct object_id *oid,
+				    enum object_type type)
+{
+	struct outgoing_link *link = oidmap_get(map, oid);
+
+	if (link) {
+		if (type != OBJ_ANY && link->type != OBJ_ANY &&
+		    type != link->type)
+			die(_("object %s is referred to as both a %s and a %s"),
+			    oid_to_hex(oid), type_name(link->type),
+			    type_name(type));
+		if (link->type == OBJ_ANY)
+			link->type = type;
+		return;
+	}
+
+	CALLOC_ARRAY(link, 1);
+	oidcpy(&link->entry.oid, oid);
+	link->type = type;
+	if (oidmap_put(map, link))
+		BUG("duplicate outgoing link");
+}
+
+static void merge_outgoing_links(struct oidmap *dest, struct oidmap *src)
+{
+	struct oidmap_iter iter;
+	struct outgoing_link *link;
+
+	while ((link = oidmap_iter_first(src, &iter))) {
+		struct outgoing_link *old = oidmap_get(dest, &link->entry.oid);
+
+		oidmap_remove(src, &link->entry.oid);
+		if (old) {
+			if (link->type != OBJ_ANY && old->type != OBJ_ANY &&
+			    link->type != old->type)
+				die(_("object %s is referred to as both a %s and a %s"),
+				    oid_to_hex(&link->entry.oid),
+				    type_name(old->type), type_name(link->type));
+			if (old->type == OBJ_ANY)
+				old->type = link->type;
+			free(link);
+		} else if (oidmap_put(dest, link)) {
+			BUG("duplicate outgoing link");
+		}
+	}
+	oidmap_clear(src, 0);
+}
+
 /*
  * Mutex and conditional variable can't be statically-initialized on Windows.
  */
@@ -211,6 +264,7 @@ static void init_thread(void)
 	CALLOC_ARRAY(thread_data, nr_threads);
 	for (i = 0; i < nr_threads; i++) {
 		thread_data[i].pack_fd = xopen(curr_pack, O_RDONLY);
+		oidmap_init(&thread_data[i].outgoing_links, 0);
 	}
 
 	threads_active = 1;
@@ -221,14 +275,17 @@ static void cleanup_thread(void)
 	int i;
 	if (!threads_active)
 		return;
-	threads_active = 0;
 	pthread_mutex_destroy(&read_mutex);
 	pthread_mutex_destroy(&counter_mutex);
 	pthread_mutex_destroy(&work_mutex);
 	if (show_stat)
 		pthread_mutex_destroy(&deepest_delta_mutex);
-	for (i = 0; i < nr_threads; i++)
+	for (i = 0; i < nr_threads; i++) {
+		merge_outgoing_links(&outgoing_links,
+				     &thread_data[i].outgoing_links);
 		close(thread_data[i].pack_fd);
+	}
+	threads_active = 0;
 	pthread_key_delete(key);
 	free(thread_data);
 }
@@ -818,9 +875,14 @@ static int check_collison(struct object_entry *entry)
 	return 0;
 }
 
-static void record_outgoing_link(const struct object_id *oid)
+static void record_outgoing_link(const struct object_id *oid,
+				 enum object_type type)
 {
-	oidset_insert(&outgoing_links, oid);
+	struct oidmap *map = &outgoing_links;
+
+	if (threads_active && !strict && !do_fsck_object)
+		map = &get_thread_data()->outgoing_links;
+	record_outgoing_link_to(map, oid, type);
 }
 
 static void maybe_record_name_entry(const struct name_entry *entry)
@@ -849,7 +911,97 @@ static void maybe_record_name_entry(const struct name_entry *entry)
 	 * pack, so it won't be GC-ed, the tradeoff seems worth it.
 	*/
 	if (S_ISDIR(entry->mode))
-		record_outgoing_link(&entry->oid);
+		record_outgoing_link(&entry->oid, OBJ_ANY);
+}
+
+static int parse_outgoing_link_oid(const char **buf, const char *tail,
+				   const char *header, struct object_id *oid)
+{
+	const char *end;
+	size_t header_len = strlen(header);
+
+	if (tail - *buf <= header_len + the_hash_algo->hexsz ||
+	    memcmp(*buf, header, header_len) ||
+	    parse_oid_hex_algop(*buf + header_len, oid, &end,
+				the_hash_algo) ||
+	    end >= tail || *end != '\n')
+		return -1;
+	*buf = end + 1;
+	return 0;
+}
+
+static void record_outgoing_links_from_data(const void *data,
+					    unsigned long size,
+					    enum object_type type,
+					    const struct object_id *oid)
+{
+	const char *buf = data;
+	const char *tail = buf + size;
+
+	if (type == OBJ_TREE) {
+		struct tree_desc desc;
+		struct name_entry entry;
+
+		if (init_tree_desc_gently(&desc, oid, data, size, 0))
+			return;
+		while (tree_entry_gently(&desc, &entry))
+			maybe_record_name_entry(&entry);
+	} else if (type == OBJ_COMMIT) {
+		struct object_id link;
+		struct commit_graft *graft;
+		int i;
+
+		if (threads_active &&
+		    !the_repository->parsed_objects->commit_graft_prepared)
+			BUG("commit grafts were not prepared before resolving deltas");
+		graft = lookup_commit_graft(the_repository, oid);
+
+		if (parse_outgoing_link_oid(&buf, tail, "tree ", &link))
+			die(_("invalid tree line in commit %s"),
+			    oid_to_hex(oid));
+		if (buf >= tail)
+			die(_("truncated commit %s after tree line"),
+			    oid_to_hex(oid));
+		record_outgoing_link(&link, OBJ_TREE);
+
+		while (tail - buf > 7 + the_hash_algo->hexsz &&
+		       starts_with(buf, "parent ")) {
+			if (parse_outgoing_link_oid(&buf, tail, "parent ",
+						    &link))
+				die(_("invalid parent line in commit %s"),
+				    oid_to_hex(oid));
+			if (buf >= tail)
+				die(_("truncated commit %s after parent line"),
+				    oid_to_hex(oid));
+			if (!graft ||
+			    (graft->nr_parent >= 0 && grafts_keep_true_parents))
+				record_outgoing_link(&link, OBJ_COMMIT);
+		}
+		if (graft)
+			for (i = 0; i < graft->nr_parent; i++)
+				record_outgoing_link(&graft->parent[i],
+						     OBJ_COMMIT);
+	} else if (type == OBJ_TAG) {
+		struct object_id link;
+		const char *line_end;
+		enum object_type target_type;
+
+		if (size < the_hash_algo->hexsz + 24 ||
+		    parse_outgoing_link_oid(&buf, tail, "object ", &link))
+			die(_("invalid object line in tag %s"), oid_to_hex(oid));
+		if (!skip_prefix(buf, "type ", &buf) ||
+		    !(line_end = memchr(buf, '\n', tail - buf)))
+			die(_("invalid type line in tag %s"), oid_to_hex(oid));
+		target_type = type_from_string_gently(buf, line_end - buf, 1);
+		if (target_type < 0)
+			die(_("invalid type line in tag %s"), oid_to_hex(oid));
+		buf = line_end + 1;
+		if (buf + 4 >= tail || !skip_prefix(buf, "tag ", &buf) ||
+		    !memchr(buf, '\n', tail - buf))
+			die(_("invalid tag name line in tag %s"),
+			    oid_to_hex(oid));
+		record_outgoing_link(&link, target_type);
+	}
 }
 
 static void do_record_outgoing_links(struct object *obj)
@@ -871,12 +1023,13 @@ static void do_record_outgoing_links(struct object *obj)
 		struct commit *commit = (struct commit *) obj;
 		struct commit_list *parents = commit->parents;
 
-		record_outgoing_link(get_commit_tree_oid(commit));
+		record_outgoing_link(get_commit_tree_oid(commit), OBJ_TREE);
 		for (; parents; parents = parents->next)
-			record_outgoing_link(&parents->item->object.oid);
+			record_outgoing_link(&parents->item->object.oid,
+					     OBJ_COMMIT);
 	} else if (obj->type == OBJ_TAG) {
 		struct tag *tag = (struct tag *) obj;
-		record_outgoing_link(get_tagged_oid(tag));
+		record_outgoing_link(get_tagged_oid(tag), tag->tagged->type);
 	}
 }
 
@@ -925,6 +1078,12 @@ static void sha1_object(const void *data, struct object_entry *obj_entry,
 		free(has_data);
 	}
 
+	if (record_outgoing_links && !strict && !do_fsck_object) {
+		if (type != OBJ_BLOB)
+			record_outgoing_links_from_data(data, size, type, oid);
+		goto out;
+	}
+
 	if (strict || do_fsck_object || record_outgoing_links) {
 		read_lock();
 		if (type == OBJ_BLOB) {
@@ -975,6 +1134,7 @@ static void sha1_object(const void *data, struct object_entry *obj_entry,
 		read_unlock();
 	}
 
+out:
 	free(new_data);
 }
 
@@ -1811,20 +1971,53 @@ static void show_pack_info(int stat_only)
 	free(chain_histogram);
 }
 
+static const struct object_id *idx_object_oid(size_t pos, const void *table)
+{
+	struct pack_idx_entry * const *entries = table;
+
+	return &entries[pos]->oid;
+}
+
+static void validate_outgoing_link_types(struct pack_idx_entry **sorted,
+					 int nr)
+{
+	struct oidmap_iter iter;
+	struct outgoing_link *link;
+
+	oidmap_iter_init(&outgoing_links, &iter);
+	while ((link = oidmap_iter_next(&iter))) {
+		int pos;
+		struct object_entry *actual;
+
+		if (link->type == OBJ_ANY)
+			continue;
+		pos = oid_pos(&link->entry.oid, sorted, nr, idx_object_oid);
+		if (pos < 0)
+			continue;
+		actual = container_of(sorted[pos], struct object_entry, idx);
+		if (actual->real_type != link->type)
+			die(_("object %s is a %s, but was referred to as a %s"),
+			    oid_to_hex(&link->entry.oid),
+			    type_name(actual->real_type),
+			    type_name(link->type));
+	}
+}
+
 static void repack_local_links(void)
 {
 	struct child_process cmd = CHILD_PROCESS_INIT;
 	FILE *out;
 	struct strbuf line = STRBUF_INIT;
-	struct oidset_iter iter;
-	struct object_id *oid;
+	struct oidmap_iter iter;
+	struct outgoing_link *link;
 	char *base_name = NULL;
 
-	if (!oidset_size(&outgoing_links))
+	if (!oidmap_get_size(&outgoing_links))
 		return;
 
-	oidset_iter_init(&outgoing_links, &iter);
-	while ((oid = oidset_iter_next(&iter))) {
+	oidmap_iter_init(&outgoing_links, &iter);
+	while ((link = oidmap_iter_next(&iter))) {
+		const struct object_id *oid = &link->entry.oid;
 		struct odb_source_info source_info;
 		struct object_info info = {
 			.source_infop = &source_info,
@@ -1919,6 +2112,7 @@ int cmd_index_pack(int argc,
 	fsck_options.walk = mark_link;
 
 	reset_pack_idx_option(&opts);
+	oidmap_init(&outgoing_links, 0);
 	opts.flags |= WRITE_REV;
 	repo_config(the_repository, git_index_pack_config, &opts);
 	if (prefix && chdir(prefix))
@@ -2102,6 +2296,7 @@ int cmd_index_pack(int argc,
 		idx_objects[i] = &objects[i].idx;
 	curr_index = write_idx_file(the_repository, index_name, idx_objects,
 				    nr_objects, &opts, pack_hash);
+	validate_outgoing_link_types(idx_objects, nr_objects);
 	if (rev_index)
 		curr_rev_index = write_rev_file(the_repository, rev_index_name,
 						idx_objects, nr_objects,
@@ -2146,6 +2341,7 @@ int cmd_index_pack(int argc,
 	free(curr_rev_index);
 
 	repack_local_links();
+	oidmap_clear(&outgoing_links, 1);
 
 	/*
 	 * Let the caller know this pack is not self contained
diff --git a/t/t5302-pack-index.sh b/t/t5302-pack-index.sh
index 735de1023e..2f99a49d9d 100755
--- a/t/t5302-pack-index.sh
+++ b/t/t5302-pack-index.sh
@@ -309,4 +309,54 @@ test_expect_success DEFAULT_HASH_ALGORITHM 'index-pack --fsck-objects outside of
 	)
 '
 
+test_expect_success 'index-pack --promisor rejects malformed commits' '
+	test_when_finished "rm -rf malformed-src malformed-dst malformed.pack bad-commit" &&
+	test_create_repo malformed-src &&
+	test_create_repo malformed-dst &&
+	printf "tree not-an-object\n\nmessage\n" >bad-commit &&
+	bad_oid=$(git -C malformed-src hash-object --literally -t commit \
+		-w --stdin <bad-commit) &&
+	printf "%s\n" "$bad_oid" |
+		git -C malformed-src pack-objects --stdout >malformed.pack &&
+	test_must_fail git -C malformed-dst index-pack --stdin --promisor \
+		<malformed.pack 2>err &&
+	test_grep "invalid tree line in commit $bad_oid" err
+'
+
+test_expect_success 'index-pack --promisor rejects conflicting link types' '
+	test_when_finished "rm -rf conflict-src conflict-dst conflict.pack bad-commit" &&
+	test_create_repo conflict-src &&
+	test_create_repo conflict-dst &&
+	tree_oid=$(git -C conflict-src mktree </dev/null) &&
+	{
+		printf "tree %s\n" "$tree_oid" &&
+		printf "parent %s\n\nmessage\n" "$tree_oid"
+	} >bad-commit &&
+	commit_oid=$(git -C conflict-src hash-object --literally -t commit \
+		-w --stdin <bad-commit) &&
+	printf "%s\n%s\n" "$tree_oid" "$commit_oid" |
+		git -C conflict-src pack-objects --stdout >conflict.pack &&
+	test_must_fail git -C conflict-dst index-pack --stdin --promisor \
+		<conflict.pack 2>err &&
+	test_grep "object $tree_oid is referred to as both a tree and a commit" err
+'
+
+test_expect_success 'index-pack --promisor verifies tag target types' '
+	test_when_finished "rm -rf tag-src tag-dst tag.pack bad-tag" &&
+	test_create_repo tag-src &&
+	test_create_repo tag-dst &&
+	tree_oid=$(git -C tag-src mktree </dev/null) &&
+	{
+		printf "object %s\n" "$tree_oid" &&
+		printf "type commit\ntag wrong-type\n\nmessage\n"
+	} >bad-tag &&
+	tag_oid=$(git -C tag-src hash-object --literally -t tag \
+		-w --stdin <bad-tag) &&
+	printf "%s\n%s\n" "$tree_oid" "$tag_oid" |
+		git -C tag-src pack-objects --stdout >tag.pack &&
+	test_must_fail git -C tag-dst index-pack --stdin --promisor \
+		<tag.pack 2>err &&
+	test_grep "object $tree_oid is a tree, but was referred to as a commit" err
+'
+
 test_done

base-commit: a97fcc37c2bc6340a8d7ce78dedf227aac4e9aa7
-- 
gitgitgadget
brian m. carlsonAug 2, 2026, 21:51 UTC in reply to Arijit Banerjee via GitGitGadget on lore

Re: [PATCH] index-pack: speed up promisor link recording

On 2026-08-02 at 21:33:15, Arijit Banerjee via GitGitGadget wrote:
Show 48 quoted lines
> From: Arijit Banerjee <arijit@effectiveailabs.com>
> 
> When indexing a promisor pack, index-pack parses every reconstructed
> non-blob object into the shared object model to record its outgoing links.
> Since parse_object_buffer() runs under read_mutex, worker threads serialize
> while allocating persistent tree, commit, and tag structures that are only
> needed to enumerate those links.
> 
> Read the links directly from the reconstructed object buffers instead. Keep
> the strict and fsck paths unchanged, use worker-local typed oidmaps during
> normal promisor indexing, and merge them after the workers exit. Transfer
> entries during the merge so that it does not temporarily duplicate the
> complete link set.
> 
> The typed entries preserve checks previously performed as a side effect of
> object parsing. Reject malformed commit and tag headers, conflicting
> expected types, and targets whose actual type disagrees when the target is
> present in the pack. Preserve commit-graft handling and the existing policy
> of recording only subtree entries from trees.
> 
> With three runs per version on Debian 12, median end-to-end wall-clock time
> for a --filter=blob:none clone of linux.git decreased from 156 seconds to
> 133 seconds (15%). Trace2 attributed the change to the initial index-pack
> --promisor phase, whose median duration decreased from 121 seconds to 98
> seconds (19%). System CPU time decreased by 46%.
> 
> Two paired spot checks against GitHub showed end-to-end reductions of 18%
> and 26%. These measurements include network and server variability and are
> therefore corroborating rather than controlled results. A third pair was not
> interpretable because the baseline request encountered a transport stall.
> 
> A full-clone control showed no material change, taking approximately 256
> seconds with either version. This is expected because full clones do not
> exercise promisor-link recording.
> 
> t5302-pack-index.sh passed with both SHA-1 and SHA-256, while
> t0410-partial-clone.sh and t5616-partial-clone.sh also passed. New coverage
> checks malformed commit headers, conflicting link types, and mismatched tag
> target types.
> 
> Signed-off-by: Arijit Banerjee <arijit@effectiveailabs.com>
> ---
>     index-pack: speed up promisor link recording
> 
>     AI assistance: OpenAI Codex was used to identify the bottleneck and
>     assist with the implementation, testing, and benchmark analysis. I
>     reviewed the resulting change and take responsibility for this
>     submission.

I don't think SubmittingPatches really allows more than trivial changes written by AI:

    The Developer's Certificate of Origin requires contributors to certify
    that they know the origin of their contributions to the project and
    that they have the right to submit it under the project's license.
    It's not yet clear that this can be legally satisfied when submitting
    significant amount of content that has been generated by AI tools.
    [...]
    To avoid these issues, we will reject anything that looks AI
    generated, that sounds overly formal or bloated, that looks like AI
    slop, that looks good on the surface but makes no sense, or that
    senders don’t understand or cannot explain.

This doesn't look like it's a trivial change, so I don't believe this patch can be accepted.

-- 
brian m. carlson (they/them)
Toronto, Ontario, CA
Arijit BanerjeeAug 2, 2026, 22:20 UTC in reply to brian m. carlson on lore

Re: [PATCH] index-pack: speed up promisor link recording

On Sun, Aug 2, 2026, brian m. carlson wrote:
> This doesn't look like it's a trivial change, so I don't believe this
> patch can be accepted.
Thanks, Brian. I am not trying to bypass the project's policy.

I do not claim to be an expert on this topic, but Codex appears to have found a material performance improvement of about 15% on end-to-end blobless clone times. Would it be appropriate to treat the current submission as an RFC for maintainers before deciding if the optimization is worth getting in? It seems worth trying to preserve the technical result.

Thanks, Arijit

On Sun, Aug 2, 2026 at 2:52 PM brian m. carlson <sandals@crustytoothpaste.net> wrote:

Show 72 quoted lines
>
> On 2026-08-02 at 21:33:15, Arijit Banerjee via GitGitGadget wrote:
> > From: Arijit Banerjee <arijit@effectiveailabs.com>
> >
> > When indexing a promisor pack, index-pack parses every reconstructed
> > non-blob object into the shared object model to record its outgoing links.
> > Since parse_object_buffer() runs under read_mutex, worker threads serialize
> > while allocating persistent tree, commit, and tag structures that are only
> > needed to enumerate those links.
> >
> > Read the links directly from the reconstructed object buffers instead. Keep
> > the strict and fsck paths unchanged, use worker-local typed oidmaps during
> > normal promisor indexing, and merge them after the workers exit. Transfer
> > entries during the merge so that it does not temporarily duplicate the
> > complete link set.
> >
> > The typed entries preserve checks previously performed as a side effect of
> > object parsing. Reject malformed commit and tag headers, conflicting
> > expected types, and targets whose actual type disagrees when the target is
> > present in the pack. Preserve commit-graft handling and the existing policy
> > of recording only subtree entries from trees.
> >
> > With three runs per version on Debian 12, median end-to-end wall-clock time
> > for a --filter=blob:none clone of linux.git decreased from 156 seconds to
> > 133 seconds (15%). Trace2 attributed the change to the initial index-pack
> > --promisor phase, whose median duration decreased from 121 seconds to 98
> > seconds (19%). System CPU time decreased by 46%.
> >
> > Two paired spot checks against GitHub showed end-to-end reductions of 18%
> > and 26%. These measurements include network and server variability and are
> > therefore corroborating rather than controlled results. A third pair was not
> > interpretable because the baseline request encountered a transport stall.
> >
> > A full-clone control showed no material change, taking approximately 256
> > seconds with either version. This is expected because full clones do not
> > exercise promisor-link recording.
> >
> > t5302-pack-index.sh passed with both SHA-1 and SHA-256, while
> > t0410-partial-clone.sh and t5616-partial-clone.sh also passed. New coverage
> > checks malformed commit headers, conflicting link types, and mismatched tag
> > target types.
> >
> > Signed-off-by: Arijit Banerjee <arijit@effectiveailabs.com>
> > ---
> >     index-pack: speed up promisor link recording
> >
> >     AI assistance: OpenAI Codex was used to identify the bottleneck and
> >     assist with the implementation, testing, and benchmark analysis. I
> >     reviewed the resulting change and take responsibility for this
> >     submission.
>
> I don't think SubmittingPatches really allows more than trivial changes
> written by AI:
>
>     The Developer's Certificate of Origin requires contributors to certify
>     that they know the origin of their contributions to the project and
>     that they have the right to submit it under the project's license.
>     It's not yet clear that this can be legally satisfied when submitting
>     significant amount of content that has been generated by AI tools.
>
>     [...]
>
>     To avoid these issues, we will reject anything that looks AI
>     generated, that sounds overly formal or bloated, that looks like AI
>     slop, that looks good on the surface but makes no sense, or that
>     senders don’t understand or cannot explain.
>
> This doesn't look like it's a trivial change, so I don't believe this
> patch can be accepted.
> --
> brian m. carlson (they/them)
> Toronto, Ontario, CA
brian m. carlsonAug 2, 2026, 22:32 UTC on lore

Re: [PATCH] index-pack: speed up promisor link recording

On 2026-08-02 at 22:12:16, Arijit Banerjee wrote:
Show 12 quoted lines
> Thanks, Brian. I am not trying to bypass the project's policy.
> 
> The investigation is in the same general spirit as the Git performance work
> being tracked here:
> https://openai-git-upstream.openai.chatgpt.site/
> 
> I do not claim to be an expert on this topic, but Codex appears to have found
> a material performance improvement of about 15% on end-to-end blobless clone
> times.
> 
> Would it be appropriate to treat the current submission as an RFC? It seems
> worth trying to preserve the technical result.

I don't think the project's policy prevents you from doing analysis and investigation with an LLM, although it does require you to verify the correctness of the results and be accountable for them. If, based on the analysis of the performance impact, you write some code without the use of an LLM that improves things, I think that would be allowed and probably welcome, assuming it is otherwise acceptable. Some contributors will be willing to review such a contribution and others will not, but it is not outside of the policy.

However, writing substantial code with an LLM doesn't appear to be allowed. The kinds of trivial changes that I think would be allowed to be generated would be things like fixing spelling errors or adding include guards to header files that lack them. Of course, these are also the kinds of things you could mostly fix with a small script, which is why they are generally considered so trivial as to be uncopyrightable.

So I think to have a patch accepted in this case, you would need to totally discard the existing patch and rewrite it by hand without recourse to the generated code.

I understand that the SubmittingPatches documentation is a bit long, but I do suggest giving it at least a glance so you know what to expect. I think reading this sort of contributing documentation is more important than ever since, in the era of LLMs, projects tend to have strong opinions on what is and is not acceptable, not only just in terms of LLM usage, but in how code and documentation are to be written and formatted.

-- 
brian m. carlson (they/them)
Toronto, Ontario, CA
Junio C HamanoAug 2, 2026, 22:46 UTC in reply to brian m. carlson on lore

Re: [PATCH] index-pack: speed up promisor link recording

"brian m. carlson" <sandals@crustytoothpaste.net> writes:
Show 25 quoted lines
>>     index-pack: speed up promisor link recording
>> 
>>     AI assistance: OpenAI Codex was used to identify the bottleneck and
>>     assist with the implementation, testing, and benchmark analysis. I
>>     reviewed the resulting change and take responsibility for this
>>     submission.
>
> I don't think SubmittingPatches really allows more than trivial changes
> written by AI:
>
>     The Developer's Certificate of Origin requires contributors to certify
>     that they know the origin of their contributions to the project and
>     that they have the right to submit it under the project's license.
>     It's not yet clear that this can be legally satisfied when submitting
>     significant amount of content that has been generated by AI tools.
>
>     [...]
>
>     To avoid these issues, we will reject anything that looks AI
>     generated, that sounds overly formal or bloated, that looks like AI
>     slop, that looks good on the surface but makes no sense, or that
>     senders don’t understand or cannot explain.
>
> This doesn't look like it's a trivial change, so I don't believe this
> patch can be accepted.
The project we borrowed DCO from has this to say on this topic:
 https://docs.kernel.org/process/coding-assistants.html
 * All contributions must comply with licensing reuqirements.
 * Only humans can attest DCO by Siging off their patches.
   The human submitter is responsible for reviewing all AI generated
   code, ensuring compliance with licensing requirements, certify
   DCO with their sign-off, and taking full responsibility for the
   contribution.

Now we are *not* kernel, but I think there is a general concensus in the community that, while we do not want to outright ban machine assisted contributions, we generally want to tread very carefully, especially in the DCO area.

It is very easy for anybody and their dog to say "I reviewed X" and it is very hard for others to assess how trustworthy such a statement is, so I am unsure how the kernel project is enforcing the "human submitter is responsible for these", and more importantly, even assuming that we would take a similar policy for ourselves, I am not sure what mechanism we can put in place to detect cases where these expectations are violated.

So...
Junio C HamanoAug 2, 2026, 22:52 UTC in reply to brian m. carlson on lore

Re: [PATCH] index-pack: speed up promisor link recording

"brian m. carlson" <sandals@crustytoothpaste.net> writes:
Show 7 quoted lines
> I understand that the SubmittingPatches documentation is a bit long, but
> I do suggest giving it at least a glance so you know what to expect.  I
> think reading this sort of contributing documentation is more important
> than ever since, in the era of LLMs, projects tend to have strong
> opinions on what is and is not acceptable, not only just in terms of LLM
> usage, but in how code and documentation are to be written and
> formatted.
Amen.

Since we seem to be drawn into the AI policy discussion, are there things that we should consider borrowing from policies battle-tested by other projects? I kind of like what LLVM has as "extractive contributions are rejected (whether it is AI or not AI)", as we are severely review-bandwidth limited these days.

Arijit BanerjeeAug 2, 2026, 23:19 UTC in reply to Junio C Hamano on lore

Re: [PATCH] index-pack: speed up promisor link recording

Two suggestions:
- Maybe a karma system?
- A special release train that is more indulgent towards AI generated
code? Brave users can try out features and they get baked into stable
releases after enough soak time

Definitely not trying to be extractive, this one appears to be a decent size perf improvement!

Still holding on to a SHA1-DC hardware acceleration change, that one would admittedly be much harder to review ;)

On Sun, Aug 2, 2026 at 3:52 PM Junio C Hamano <gitster@pobox.com> wrote:
Show 19 quoted lines
>
> "brian m. carlson" <sandals@crustytoothpaste.net> writes:
>
> > I understand that the SubmittingPatches documentation is a bit long, but
> > I do suggest giving it at least a glance so you know what to expect.  I
> > think reading this sort of contributing documentation is more important
> > than ever since, in the era of LLMs, projects tend to have strong
> > opinions on what is and is not acceptable, not only just in terms of LLM
> > usage, but in how code and documentation are to be written and
> > formatted.
>
> Amen.
>
> Since we seem to be drawn into the AI policy discussion, are there
> things that we should consider borrowing from policies battle-tested
> by other projects?  I kind of like what LLVM has as "extractive
> contributions are rejected (whether it is AI or not AI)", as we are
> severely review-bandwidth limited these days.
>
brian m. carlsonAug 3, 2026, 00:31 UTC on lore

Re: [PATCH] index-pack: speed up promisor link recording

[please avoid top-posting]
On 2026-08-02 at 22:54:27, Arijit Banerjee wrote:
Show 6 quoted lines
> Maybe software can ship experimental versions which is more indulging
> towards AI generated patches? Brave users get to try the features and they
> can baked into stable releases once there's enough soak time.
> 
> I have been holding onto my patch around hardware acceleration for SHA1-DC
> :)

The rationale, as Junio said, is based on the Developer's Certificate of Origin. That is a legal statement that a person has the legal right to contribute those changes under the license and if they make a false or misleading statement to that effect, they are responsible—legally and otherwise—for it.

Considering the extensive litigation over LLM output at the moment, I don't think anyone can clearly make that assertion. Most of the arguments I've heard are that it's fair use, which is a U.S. legal concept. That does not exist in Canada or the U.K., where there is fair dealing, which is much more restricted.

If Company X includes LLM-generated code in their proprietary product and it's found to be infringing in say, Germany, then they can simply not distribute their code in Germany. Git cannot do that: it's distributed in Linux distributions around the planet, even in countries subject to sanctions, such as Russia[0]. We must comply with the license and the law everywhere in every country or we risk liability for our contributors and distributors. I, for one, am not willing to be sued over this project and the project does not have the financial means to deal with extensive litigation.

There are also concerns about the quality of the code and whether submitters adequately understand the code well enough to have evaluated and reviewed it thoroughly. It's well known that when creating code becomes cheap, the burden shifts to review and review becomes extremely important. That has been seen in lots of places, but we are an open source project and we can't force contributors to do review like a company can. As Junio says, we already have trouble getting reviews through and we don't want to make the problem worse.

LLM-generated content also has a negative quality reputation (see the reaction to AI content in video games and books for an example) and while all software has bugs, I appreciate the reputation that Git has for quality and wish to retain that.

Those alone are reason enough for the policy, but there are other concerns about the environmental impact, the impact on electricity and hardware prices, the ethics of incorporating open source code without so much as a credit[1], and a lot more.

So I don't think it's likely we're going to accept nontrivial LLM-generated content in any capacity anytime soon and I don't think trying to argue this or persuade us to accept it is going to be productive or well received. Of course, anyone can distribute their own fork of Git with additional patches if they prefer, but we won't include them.

[0] Debian, which distributes Git, has mirrors in Russia and Belarus: https://www.debian.org/mirror/list [1] For instance, as a member of ACM, I have to follow § 1.5 (Respect the work required to produce new ideas, inventions, creative works, and computing artifacts) of the Code of Ethics: https://www.acm.org/code-of-ethics.

-- 
brian m. carlson (they/them)
Toronto, Ontario, CA
Collin FunkAug 3, 2026, 00:55 UTC in reply to brian m. carlson on lore

Re: [PATCH] index-pack: speed up promisor link recording

"brian m. carlson" <sandals@crustytoothpaste.net> writes:
Show 9 quoted lines
> If Company X includes LLM-generated code in their proprietary product
> and it's found to be infringing in say, Germany, then they can simply
> not distribute their code in Germany.  Git cannot do that: it's
> distributed in Linux distributions around the planet, even in countries
> subject to sanctions, such as Russia[0].  We must comply with the
> license and the law everywhere in every country or we risk liability for
> our contributors and distributors.  I, for one, am not willing to be
> sued over this project and the project does not have the financial means
> to deal with extensive litigation.

I generally agree. But some countries have more respected legal systems than others, to put it mildly. I certainly hope that being extradited to one of the sanctioned countries isn't too large of a concern.

Collin

Back to recent threads