{"thread":{"id":"53933","subject":"[PATCH 1/6] commit-graph: fix regression when computing bloom filter","startedAt":"2020-07-28T09:13:57Z","lastAt":"2021-02-01T18:27:06Z","messageCount":211,"participants":["Abhishek Kumar via GitGitGadget","Derrick Stolee","Taylor Blau","René Scharfe","Abhishek Kumar","Jakub Narębski","Junio C Hamano","Philip Oakley","SZEDER Gábor"],"isPatch":true,"patchVersion":1,"patchTotal":6},"messages":[{"id":"402183","messageId":"pull.676.git.1595927632.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":null,"subject":"[PATCH 0/6] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-07-28T09:13:45Z","receivedAt":"2020-07-28T09:13:57Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"This patch series implements the corrected commit date offsets as generation\nnumber v2, along with other pre-requisites.\n\nGit uses topological levels in the commit-graph file for commit-graph\ntraversal operations like git log --graph. Unfortunately, using topological\nlevels can result in a worse performance than without them when compared\nwith committer date as a heuristics. For example, git merge-base v4.8 v4.9 \non the Linux repository walks 635,579 commits using topological levels and\nwalks 167,468 using committer date.\n\nThus, the need for generation number v2 was born. New generation number\nneeded to provide good performance, increment updates, and backward\ncompatibility. Due to an unfortunate problem, we also needed a way to\ndistinguish between the old and new generation number without incrementing\ngraph version.\n\nVarious candidates were examined (https://github.com/derrickstolee/gen-test, \nhttps://github.com/abhishekkumar2718/git/pull/1). The proposed generation\nnumber v2, Corrected Commit Date with Mononotically Increasing Offsets \nperformed much worse than committer date (506,577 vs. 167,468 commits walked\nfor git merge-base v4.8 v4.9) and was dropped.\n\nUsing Generation Data chunk (GDAT) relieves the requirement of backward\ncompatibility as we would continue to store topological levels in Commit\nData (CDAT) chunk. Thus, Corrected Commit Date was chosen as generation\nnumber v2. The Corrected Commit Date is defined as:\n\nFor a commit C, let its corrected commit date be the maximum of the commit\ndate of C and the corrected commit dates of its parents. Then corrected\ncommit date offset is the difference between corrected commit date of C and\ncommit date of C.\n\nWe will introduce an additional commit-graph chunk, Generation Data chunk,\nand store corrected commit date offsets in GDAT chunk while storing\ntopological levels in CDAT chunk. The old versions of Git would ignore GDAT\nchunk, using topological levels from CDAT chunk. In contrast, new versions\nof Git would use corrected commit dates, falling back to topological level\nif the generation data chunk is absent in the commit-graph file.\n\nHere's what left for the PR (which I intend to take on with the second\nversion of pull request):\n\n 1. Add an option to skip writing generation data chunk (to test whether new\n    Git works without GDAT as intended).\n 2. Handle writing to commit-graph for mismatched version (that is, merging\n    all graphs into a new graph with a GDAT chunk).\n 3. Update technical documentation.\n\nI look forward to everyone's reviews!\n\nThanks\n\n * Abhishek\n\n\n----------------------------------------------------------------------------\n\nThe build fails for t9807-git-p4-submit.sh on osx-clang, which I feel is\nunrelated to my code changes. Still need to investigate further.\n\nAbhishek Kumar (6):\n  commit-graph: fix regression when computing bloom filter\n  revision: parse parent in indegree_walk_step()\n  commit-graph: consolidate fill_commit_graph_info\n  commit-graph: consolidate compare_commits_by_gen\n  commit-graph: implement generation data chunk\n  commit-graph: implement corrected commit date offset\n\n blame.c                       |   2 +-\n commit-graph.c                | 181 +++++++++++++++++++++-------------\n commit-graph.h                |   7 +-\n commit-reach.c                |  47 +++------\n commit-reach.h                |   2 +-\n commit.c                      |   9 +-\n commit.h                      |   3 +\n revision.c                    |  17 ++--\n t/helper/test-read-graph.c    |   2 +\n t/t4216-log-bloom.sh          |   4 +-\n t/t5000-tar-tree.sh           |   4 +-\n t/t5318-commit-graph.sh       |  21 ++--\n t/t5324-split-commit-graph.sh |  12 +--\n upload-pack.c                 |   2 +-\n 14 files changed, 178 insertions(+), 135 deletions(-)\n\n\nbase-commit: 47ae905ffb98cc4d4fd90083da6bc8dab55d9ecc\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-676%2Fabhishekkumar2718%2Fcorrected_commit_date-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-676/abhishekkumar2718/corrected_commit_date-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/676\n-- \ngitgitgadget\n"},{"id":"402180","messageId":"91e6e97a66aff88e0b860e34659dddc3396c7f28.1595927632.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.git.1595927632.gitgitgadget@gmail.com","subject":"[PATCH 1/6] commit-graph: fix regression when computing bloom filter","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-07-28T09:13:46Z","receivedAt":"2020-07-28T09:13:58Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith 3d112755 (commit-graph: examine commits by generation number), Git\nknew to sort by generation number before examining the diff when not\nusing pack order. c49c82aa (commit: move members graph_pos, generation\nto a slab, 2020-06-17) moved generation number into a slab and\nintroduced a helper which returns GENERATION_NUMBER_INFINITY when\nwriting the graph. Sorting is no longer useful and essentially reverts\nthe earlier commit.\n\nLet's fix this by accessing generation number directly through the slab.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 5 +++--\n 1 file changed, 3 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 1af68c297d..5d3c9bd23c 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -144,8 +144,9 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n+\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n+\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n+\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"402181","messageId":"d23f67dc80b85abe4eba9a9dfc39d50188e23bb7.1595927632.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.git.1595927632.gitgitgadget@gmail.com","subject":"[PATCH 2/6] revision: parse parent in indegree_walk_step()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-07-28T09:13:47Z","receivedAt":"2020-07-28T09:14:00Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn indegree_walk_step(), we add unvisited parents to the indegree queue.\nHowever, parents are not guaranteed to be parsed. As the indegree queue\nsorts by generation number, let's parse parents before inserting them to\nensure the correct priority order.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n revision.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/revision.c b/revision.c\nindex 6aa7f4f567..23287d26c3 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3343,6 +3343,9 @@ static void indegree_walk_step(struct rev_info *revs)\n \t\tstruct commit *parent = p->item;\n \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n \n+\t\tif (parse_commit_gently(parent, 1) < 0)\n+\t\t\treturn ;\n+\n \t\tif (*pi)\n \t\t\t(*pi)++;\n \t\telse\n-- \ngitgitgadget\n\n"},{"id":"402184","messageId":"701f5912369c0fcc07cf604c3129cb5017a125ce.1595927632.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.git.1595927632.gitgitgadget@gmail.com","subject":"[PATCH 3/6] commit-graph: consolidate fill_commit_graph_info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-07-28T09:13:48Z","receivedAt":"2020-07-28T09:14:01Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nBoth fill_commit_graph_info() and fill_commit_in_graph() parse\ninformation present in commit data chunk. Let's simplify the\nimplementation by calling fill_commit_graph_info() within\nfill_commit_in_graph().\n\nThe test 'generate tar with future mtime' creates a commit with commit\ntime of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\ngeneration number and has undefined behavior. The test used to pass as\nfill_commit_in_graph() did not read commit time from commit graph,\nreading commit date from odb instead.\n\nLet's fix that by setting commit time of (2 ^ 34 - 1) seconds.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c      | 31 ++++++++++++-------------------\n t/t5000-tar-tree.sh |  4 ++--\n 2 files changed, 14 insertions(+), 21 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 5d3c9bd23c..204eb454b2 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -735,15 +735,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n \tuint32_t lex_index;\n+\tuint64_t date_high, date_low;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n+\tif (pos >= g->num_commits + g->num_commits_in_base)\n+\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\n \tlex_index = pos - g->num_commits_in_base;\n \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n \n \tgraph_data = commit_graph_data_at(item);\n \tgraph_data->graph_pos = pos;\n+\n+\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n+\tdate_low = get_be32(commit_data + g->hash_len + 12);\n+\titem->date = (timestamp_t)((date_high << 32) | date_low);\n+\n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n@@ -758,38 +767,22 @@ static int fill_commit_in_graph(struct repository *r,\n {\n \tuint32_t edge_value;\n \tuint32_t *parent_data_ptr;\n-\tuint64_t date_low, date_high;\n \tstruct commit_list **pptr;\n-\tstruct commit_graph_data *graph_data;\n \tconst unsigned char *commit_data;\n \tuint32_t lex_index;\n \n+\tfill_commit_graph_info(item, g, pos);\n+\n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n-\tif (pos >= g->num_commits + g->num_commits_in_base)\n-\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n-\n-\t/*\n-\t * Store the \"full\" position, but then use the\n-\t * \"local\" position for the rest of the calculation.\n-\t */\n-\tgraph_data = commit_graph_data_at(item);\n-\tgraph_data->graph_pos = pos;\n \tlex_index = pos - g->num_commits_in_base;\n-\n-\tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n+\tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n \n \titem->object.parsed = 1;\n \n \tset_commit_tree(item, NULL);\n \n-\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n-\tdate_low = get_be32(commit_data + g->hash_len + 12);\n-\titem->date = (timestamp_t)((date_high << 32) | date_low);\n-\n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n-\n \tpptr = &item->parents;\n \n \tedge_value = get_be32(commit_data + g->hash_len);\ndiff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\nindex 37655a237c..1986354fc3 100755\n--- a/t/t5000-tar-tree.sh\n+++ b/t/t5000-tar-tree.sh\n@@ -406,7 +406,7 @@ test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n \trm -f .git/index &&\n \techo content >file &&\n \tgit add file &&\n-\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n+\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n \t\tgit commit -m \"tempori parendum\"\n '\n \n@@ -415,7 +415,7 @@ test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n '\n \n test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n-\techo 4147 >expect &&\n+\techo 2514 >expect &&\n \ttar_info future.tar | cut -d\" \" -f2 >actual &&\n \ttest_cmp expect actual\n '\n-- \ngitgitgadget\n\n"},{"id":"402182","messageId":"812fe75fc7252db0b7b6604f84b17dcf7324b922.1595927632.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.git.1595927632.gitgitgadget@gmail.com","subject":"[PATCH 4/6] commit-graph: consolidate compare_commits_by_gen","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-07-28T09:13:49Z","receivedAt":"2020-07-28T09:14:02Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nComparing commits by generation has been independently defined twice, in\ncommit-reach and commit. Let's simplify the implementation by moving\ncompare_commits_by_gen() to commit-graph.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 15 +++++++++++++++\n commit-graph.h |  2 ++\n commit-reach.c | 15 ---------------\n commit.c       |  9 +++------\n 4 files changed, 20 insertions(+), 21 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 204eb454b2..1c98f38d69 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -112,6 +112,21 @@ uint32_t commit_graph_generation(const struct commit *c)\n \treturn data->generation;\n }\n \n+int compare_commits_by_gen(const void *_a, const void *_b)\n+{\n+\tconst struct commit *a = _a, *b = _b;\n+\tconst uint32_t generation_a = commit_graph_generation(a);\n+\tconst uint32_t generation_b = commit_graph_generation(b);\n+\n+\t/* older commits first */\n+\tif (generation_a < generation_b)\n+\t\treturn -1;\n+\telse if (generation_a > generation_b)\n+\t\treturn 1;\n+\n+\treturn 0;\n+}\n+\n static struct commit_graph_data *commit_graph_data_at(const struct commit *c)\n {\n \tunsigned int i, nth_slab;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 28f89cdf3e..98cc5a3b9d 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -145,4 +145,6 @@ struct commit_graph_data {\n  */\n uint32_t commit_graph_generation(const struct commit *);\n uint32_t commit_graph_position(const struct commit *);\n+\n+int compare_commits_by_gen(const void *_a, const void *_b);\n #endif\ndiff --git a/commit-reach.c b/commit-reach.c\nindex efd5925cbb..c83cc291e7 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -561,21 +561,6 @@ int commit_contains(struct ref_filter *filter, struct commit *commit,\n \treturn repo_is_descendant_of(the_repository, commit, list);\n }\n \n-static int compare_commits_by_gen(const void *_a, const void *_b)\n-{\n-\tconst struct commit *a = *(const struct commit * const *)_a;\n-\tconst struct commit *b = *(const struct commit * const *)_b;\n-\n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n-\n-\tif (generation_a < generation_b)\n-\t\treturn -1;\n-\tif (generation_a > generation_b)\n-\t\treturn 1;\n-\treturn 0;\n-}\n-\n int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\ndiff --git a/commit.c b/commit.c\nindex 7128895c3a..bed63b41fb 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -731,14 +731,11 @@ int compare_commits_by_author_date(const void *a_, const void *b_,\n int compare_commits_by_gen_then_commit_date(const void *a_, const void *b_, void *unused)\n {\n \tconst struct commit *a = a_, *b = b_;\n-\tconst uint32_t generation_a = commit_graph_generation(a),\n-\t\t       generation_b = commit_graph_generation(b);\n+\tint ret_val = compare_commits_by_gen(a_, b_);\n \n \t/* newer commits first */\n-\tif (generation_a < generation_b)\n-\t\treturn 1;\n-\telse if (generation_a > generation_b)\n-\t\treturn -1;\n+\tif (ret_val)\n+\t\treturn -ret_val;\n \n \t/* use date as a heuristic when generations are equal */\n \tif (a->date < b->date)\n-- \ngitgitgadget\n\n"},{"id":"402186","messageId":"80ea7da3435396edcb19423ab602962d31585209.1595927632.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.git.1595927632.gitgitgadget@gmail.com","subject":"[PATCH 5/6] commit-graph: implement generation data chunk","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-07-28T09:13:50Z","receivedAt":"2020-07-28T09:14:04Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nOne of the essential pre-requisites before implementing generation\nnumber as to distinguish between generation numbers v1 and v2 while\nstill being compatible with old Git.\n\nWe are going to introduce a new chunk called Generation Data chunk (or\nGDAT). GDAT stores generation number v2 (and any subsequent versions),\nwhereas CDAT will still store topological level.\n\nOld Git does not understand GDAT chunk and would ignore it, reading\ntopological levels from CDAT. Newer versions of Git can parse GDAT and\ntake advantage of newer generation numbers, falling back to topological\nlevels when GDAT chunk is missing (as it would happen with a commit\ngraph written by old Git).\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                | 33 +++++++++++++++++++++++++++++----\n commit-graph.h                |  1 +\n t/helper/test-read-graph.c    |  2 ++\n t/t4216-log-bloom.sh          |  4 ++--\n t/t5318-commit-graph.sh       | 19 +++++++++++--------\n t/t5324-split-commit-graph.sh | 12 ++++++------\n 6 files changed, 51 insertions(+), 20 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 1c98f38d69..ab714f4a76 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -38,11 +38,12 @@ void git_test_write_commit_graph_or_die(void)\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n+#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n-#define MAX_NUM_CHUNKS 7\n+#define MAX_NUM_CHUNKS 8\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -389,6 +390,13 @@ struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size)\n \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n \t\t\tbreak;\n \n+\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n+\t\t\tif (graph->chunk_generation_data)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n+\t\t\tbreak;\n+\n \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n \t\t\tif (graph->chunk_extra_edges)\n \t\t\t\tchunk_repeated = 1;\n@@ -768,7 +776,10 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\tif (g->chunk_generation_data)\n+\t\tgraph_data->generation = get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n+\telse\n+\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n static inline void set_commit_tree(struct commit *c, struct tree *t)\n@@ -1100,6 +1111,17 @@ static void write_graph_chunk_data(struct hashfile *f, int hash_len,\n \t}\n }\n \n+static void write_graph_chunk_generation_data(struct hashfile *f,\n+\t\t\t\t\t      struct write_commit_graph_context *ctx)\n+{\n+\tstruct commit **list = ctx->commits.list;\n+\tint count;\n+\tfor (count = 0; count < ctx->commits.nr; count++, list++) {\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n+\t\thashwrite_be32(f, commit_graph_data_at(*list)->generation);\n+\t}\n+}\n+\n static void write_graph_chunk_extra_edges(struct hashfile *f,\n \t\t\t\t\t  struct write_commit_graph_context *ctx)\n {\n@@ -1605,7 +1627,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tuint64_t chunk_offsets[MAX_NUM_CHUNKS + 1];\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n \tstruct strbuf progress_title = STRBUF_INIT;\n-\tint num_chunks = 3;\n+\tint num_chunks = 4;\n \tstruct object_id file_hash;\n \tconst struct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \n@@ -1656,6 +1678,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tchunk_ids[0] = GRAPH_CHUNKID_OIDFANOUT;\n \tchunk_ids[1] = GRAPH_CHUNKID_OIDLOOKUP;\n \tchunk_ids[2] = GRAPH_CHUNKID_DATA;\n+\tchunk_ids[3] = GRAPH_CHUNKID_GENERATION_DATA;\n \tif (ctx->num_extra_edges) {\n \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_EXTRAEDGES;\n \t\tnum_chunks++;\n@@ -1677,8 +1700,9 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tchunk_offsets[1] = chunk_offsets[0] + GRAPH_FANOUT_SIZE;\n \tchunk_offsets[2] = chunk_offsets[1] + hashsz * ctx->commits.nr;\n \tchunk_offsets[3] = chunk_offsets[2] + (hashsz + 16) * ctx->commits.nr;\n+\tchunk_offsets[4] = chunk_offsets[3] + sizeof(uint32_t) * ctx->commits.nr;\n \n-\tnum_chunks = 3;\n+\tnum_chunks = 4;\n \tif (ctx->num_extra_edges) {\n \t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n \t\t\t\t\t\t4 * ctx->num_extra_edges;\n@@ -1728,6 +1752,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \twrite_graph_chunk_fanout(f, ctx);\n \twrite_graph_chunk_oids(f, hashsz, ctx);\n \twrite_graph_chunk_data(f, hashsz, ctx);\n+\twrite_graph_chunk_generation_data(f, ctx);\n \tif (ctx->num_extra_edges)\n \t\twrite_graph_chunk_extra_edges(f, ctx);\n \tif (ctx->changed_paths) {\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 98cc5a3b9d..e3d4ba96f4 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -67,6 +67,7 @@ struct commit_graph {\n \tconst uint32_t *chunk_oid_fanout;\n \tconst unsigned char *chunk_oid_lookup;\n \tconst unsigned char *chunk_commit_data;\n+\tconst unsigned char *chunk_generation_data;\n \tconst unsigned char *chunk_extra_edges;\n \tconst unsigned char *chunk_base_graphs;\n \tconst unsigned char *chunk_bloom_indexes;\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex 6d0c962438..1c2a5366c7 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -32,6 +32,8 @@ int cmd__read_graph(int argc, const char **argv)\n \t\tprintf(\" oid_lookup\");\n \tif (graph->chunk_commit_data)\n \t\tprintf(\" commit_metadata\");\n+\tif (graph->chunk_generation_data)\n+\t\tprintf(\" generation_data\");\n \tif (graph->chunk_extra_edges)\n \t\tprintf(\" extra_edges\");\n \tif (graph->chunk_bloom_indexes)\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex c855bcd3e7..780855e691 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -33,11 +33,11 @@ test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n \tgit commit-graph write --reachable --changed-paths\n '\n graph_read_expect () {\n-\tNUM_CHUNKS=5\n+\tNUM_CHUNKS=6\n \tcat >expect <<- EOF\n \theader: 43475048 1 1 $NUM_CHUNKS 0\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n \tEOF\n \ttest-tool read-graph >actual &&\n \ttest_cmp expect actual\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 26f332d6a3..3ec5248d70 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -71,16 +71,16 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n \n graph_read_expect() {\n \tOPTIONAL=\"\"\n-\tNUM_CHUNKS=3\n+\tNUM_CHUNKS=4\n \tif test ! -z $2\n \tthen\n \t\tOPTIONAL=\" $2\"\n-\t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n+\t\tNUM_CHUNKS=$((4 + $(echo \"$2\" | wc -w)))\n \tfi\n \tcat >expect <<- EOF\n \theader: 43475048 1 1 $NUM_CHUNKS 0\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata$OPTIONAL\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data$OPTIONAL\n \tEOF\n \ttest-tool read-graph >output &&\n \ttest_cmp expect output\n@@ -433,7 +433,7 @@ GRAPH_BYTE_HASH=5\n GRAPH_BYTE_CHUNK_COUNT=6\n GRAPH_CHUNK_LOOKUP_OFFSET=8\n GRAPH_CHUNK_LOOKUP_WIDTH=12\n-GRAPH_CHUNK_LOOKUP_ROWS=5\n+GRAPH_CHUNK_LOOKUP_ROWS=6\n GRAPH_BYTE_OID_FANOUT_ID=$GRAPH_CHUNK_LOOKUP_OFFSET\n GRAPH_BYTE_OID_LOOKUP_ID=$(($GRAPH_CHUNK_LOOKUP_OFFSET + \\\n \t\t\t    1 * $GRAPH_CHUNK_LOOKUP_WIDTH))\n@@ -451,11 +451,14 @@ GRAPH_BYTE_COMMIT_TREE=$GRAPH_COMMIT_DATA_OFFSET\n GRAPH_BYTE_COMMIT_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN))\n GRAPH_BYTE_COMMIT_EXTRA_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 4))\n GRAPH_BYTE_COMMIT_WRONG_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 3))\n-GRAPH_BYTE_COMMIT_GENERATION=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 11))\n GRAPH_BYTE_COMMIT_DATE=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 12))\n GRAPH_COMMIT_DATA_WIDTH=$(($HASH_LEN + 16))\n-GRAPH_OCTOPUS_DATA_OFFSET=$(($GRAPH_COMMIT_DATA_OFFSET + \\\n-\t\t\t     $GRAPH_COMMIT_DATA_WIDTH * $NUM_COMMITS))\n+GRAPH_GENERATION_DATA_OFFSET=$(($GRAPH_COMMIT_DATA_OFFSET + \\\n+\t\t\t\t$GRAPH_COMMIT_DATA_WIDTH * $NUM_COMMITS))\n+GRAPH_GENERATION_DATA_WIDTH=4\n+GRAPH_BYTE_COMMIT_GENERATION=$(($GRAPH_GENERATION_DATA_OFFSET + 3))\n+GRAPH_OCTOPUS_DATA_OFFSET=$(($GRAPH_GENERATION_DATA_OFFSET + \\\n+\t\t\t     $GRAPH_GENERATION_DATA_WIDTH * $NUM_COMMITS))\n GRAPH_BYTE_OCTOPUS=$(($GRAPH_OCTOPUS_DATA_OFFSET + 4))\n GRAPH_BYTE_FOOTER=$(($GRAPH_OCTOPUS_DATA_OFFSET + 4 * $NUM_OCTOPUS_EDGES))\n \n@@ -594,7 +597,7 @@ test_expect_success 'detect incorrect generation number' '\n '\n \n test_expect_success 'detect incorrect generation number' '\n-\tcorrupt_graph_and_verify $GRAPH_BYTE_COMMIT_GENERATION \"\\01\" \\\n+\tcorrupt_graph_and_verify $GRAPH_BYTE_COMMIT_GENERATION \"\\00\" \\\n \t\t\"non-zero generation number\"\n '\n \ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 269d0964a3..096a96ec41 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -14,11 +14,11 @@ test_expect_success 'setup repo' '\n \tgraphdir=\"$infodir/commit-graphs\" &&\n \ttest_oid_init &&\n \ttest_oid_cache <<-EOM\n-\tshallow sha1:1760\n-\tshallow sha256:2064\n+\tshallow sha1:2132\n+\tshallow sha256:2436\n \n-\tbase sha1:1376\n-\tbase sha256:1496\n+\tbase sha1:1408\n+\tbase sha256:1528\n \tEOM\n '\n \n@@ -29,9 +29,9 @@ graph_read_expect() {\n \t\tNUM_BASE=$2\n \tfi\n \tcat >expect <<- EOF\n-\theader: 43475048 1 1 3 $NUM_BASE\n+\theader: 43475048 1 1 4 $NUM_BASE\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n \tEOF\n \ttest-tool read-graph >output &&\n \ttest_cmp expect output\n-- \ngitgitgadget\n\n"},{"id":"402185","messageId":"647290d0368e385227614dd1822aa9083b0dba5e.1595927632.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.git.1595927632.gitgitgadget@gmail.com","subject":"[PATCH 6/6] commit-graph: implement corrected commit date offset","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-07-28T09:13:51Z","receivedAt":"2020-07-28T09:14:06Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith preparations done, let's implement corrected commit date offset.\nWe add a new commit-slab to store topological levels while writing\ncommit graph and upgrade number of struct commit_graph_data to 64-bits.\n\nWe have to touch many files, upgrading generation number from uint32_t\nto timestamp_t.\n\nWe drop 'detect incorrect generation number' from t5318-commit-graph.sh,\nwhich tests if verify can detect if a commit graph have\nGENERATION_NUMBER_ZERO for a commit, followed by a non-zero generation.\nWith corrected commit dates, GENERATION_NUMBER_ZERO is possible only if\none of dates is Unix epoch zero.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n blame.c                 |   2 +-\n commit-graph.c          | 109 ++++++++++++++++++++++------------------\n commit-graph.h          |   4 +-\n commit-reach.c          |  32 ++++++------\n commit-reach.h          |   2 +-\n commit.h                |   3 ++\n revision.c              |  14 +++---\n t/t5318-commit-graph.sh |   2 +-\n upload-pack.c           |   2 +-\n 9 files changed, 93 insertions(+), 77 deletions(-)\n\ndiff --git a/blame.c b/blame.c\nindex 82fa16d658..48aa632461 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -1272,7 +1272,7 @@ static int maybe_changed_path(struct repository *r,\n \tif (!bd)\n \t\treturn 1;\n \n-\tif (commit_graph_generation(origin->commit) == GENERATION_NUMBER_INFINITY)\n+\tif (commit_graph_generation(origin->commit) == GENERATION_NUMBER_V2_INFINITY)\n \t\treturn 1;\n \n \tfilter = get_bloom_filter(r, origin->commit, 0);\ndiff --git a/commit-graph.c b/commit-graph.c\nindex ab714f4a76..9647d9f0df 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -65,6 +65,8 @@ void git_test_write_commit_graph_or_die(void)\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n+define_commit_slab(topo_level_slab, uint32_t);\n+\n /* Keep track of the order in which commits are added to our list. */\n define_commit_slab(commit_pos, int);\n static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n@@ -100,15 +102,15 @@ uint32_t commit_graph_position(const struct commit *c)\n \treturn data ? data->graph_pos : COMMIT_NOT_FROM_GRAPH;\n }\n \n-uint32_t commit_graph_generation(const struct commit *c)\n+timestamp_t commit_graph_generation(const struct commit *c)\n {\n \tstruct commit_graph_data *data =\n \t\tcommit_graph_data_slab_peek(&commit_graph_data_slab, c);\n \n \tif (!data)\n-\t\treturn GENERATION_NUMBER_INFINITY;\n+\t\treturn GENERATION_NUMBER_V2_INFINITY;\n \telse if (data->graph_pos == COMMIT_NOT_FROM_GRAPH)\n-\t\treturn GENERATION_NUMBER_INFINITY;\n+\t\treturn GENERATION_NUMBER_V2_INFINITY;\n \n \treturn data->generation;\n }\n@@ -116,8 +118,8 @@ uint32_t commit_graph_generation(const struct commit *c)\n int compare_commits_by_gen(const void *_a, const void *_b)\n {\n \tconst struct commit *a = _a, *b = _b;\n-\tconst uint32_t generation_a = commit_graph_generation(a);\n-\tconst uint32_t generation_b = commit_graph_generation(b);\n+\tconst timestamp_t generation_a = commit_graph_generation(a);\n+\tconst timestamp_t generation_b = commit_graph_generation(b);\n \n \t/* older commits first */\n \tif (generation_a < generation_b)\n@@ -160,8 +162,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n-\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n+\ttimestamp_t generation_a = commit_graph_data_at(a)->generation;\n+\ttimestamp_t generation_b = commit_graph_data_at(b)->generation;\n \n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n@@ -169,11 +171,6 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \telse if (generation_a > generation_b)\n \t\treturn 1;\n \n-\t/* use date as a heuristic when generations are equal */\n-\tif (a->date < b->date)\n-\t\treturn -1;\n-\telse if (a->date > b->date)\n-\t\treturn 1;\n \treturn 0;\n }\n \n@@ -777,8 +774,13 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n \tif (g->chunk_generation_data)\n-\t\tgraph_data->generation = get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n+\t{\n+\t\t/* Read corrected commit date offset from GDAT */\n+\t\tgraph_data->generation = item->date +\n+\t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n+\t}\n \telse\n+\t\t/* Read topological level from CDAT */\n \t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n@@ -950,6 +952,7 @@ struct write_commit_graph_context {\n \tstruct progress *progress;\n \tint progress_done;\n \tuint64_t progress_cnt;\n+\tstruct topo_level_slab *topo_levels;\n \n \tchar *base_graph_name;\n \tint num_commit_graphs_before;\n@@ -1102,7 +1105,7 @@ static void write_graph_chunk_data(struct hashfile *f, int hash_len,\n \t\telse\n \t\t\tpackedDate[0] = 0;\n \n-\t\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n+\t\tpackedDate[0] |= htonl(*topo_level_slab_at(ctx->topo_levels, *list) << 2);\n \n \t\tpackedDate[1] = htonl((*list)->date);\n \t\thashwrite(f, packedDate, 8);\n@@ -1117,8 +1120,13 @@ static void write_graph_chunk_generation_data(struct hashfile *f,\n \tstruct commit **list = ctx->commits.list;\n \tint count;\n \tfor (count = 0; count < ctx->commits.nr; count++, list++) {\n+\t\ttimestamp_t offset = commit_graph_data_at(*list)->generation - (*list)->date;\n \t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n-\t\thashwrite_be32(f, commit_graph_data_at(*list)->generation);\n+\n+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX)\n+\t\t\toffset = GENERATION_NUMBER_V2_OFFSET_MAX;\n+\n+\t\thashwrite_be32(f, offset);\n \t}\n }\n \n@@ -1316,7 +1324,7 @@ static void close_reachable(struct write_commit_graph_context *ctx)\n \tstop_progress(&ctx->progress);\n }\n \n-static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n+static void compute_corrected_commit_date_offsets(struct write_commit_graph_context *ctx)\n {\n \tint i;\n \tstruct commit_list *list = NULL;\n@@ -1326,11 +1334,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n+\t\tuint32_t topo_level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n-\t\t    generation != GENERATION_NUMBER_ZERO)\n+\t\tif (topo_level != GENERATION_NUMBER_INFINITY &&\n+\t\t    topo_level != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1338,29 +1346,38 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit *current = list->item;\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n-\t\t\tuint32_t max_generation = 0;\n+\t\t\tuint32_t max_level = 0;\n+\t\t\ttimestamp_t max_corrected_commit_date = current->date;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n-\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n+\t\t\t\ttopo_level = *topo_level_slab_at(ctx->topo_levels, parent->item);\n \n-\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n-\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n+\t\t\t\tif (topo_level == GENERATION_NUMBER_INFINITY ||\n+\t\t\t\t    topo_level == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (generation > max_generation) {\n-\t\t\t\t\tmax_generation = generation;\n+\t\t\t\t} else {\n+\t\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(parent->item);\n+\n+\t\t\t\t\tif (topo_level > max_level)\n+\t\t\t\t\t\tmax_level = topo_level;\n+\n+\t\t\t\t\tif (data->generation > max_corrected_commit_date)\n+\t\t\t\t\t\tmax_corrected_commit_date = data->generation;\n \t\t\t\t}\n \t\t\t}\n \n \t\t\tif (all_parents_computed) {\n \t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n \n-\t\t\t\tdata->generation = max_generation + 1;\n-\t\t\t\tpop_commit(&list);\n+\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n+\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n+\n+\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n+\t\t\t\tdata->generation = max_corrected_commit_date + 1;\n \n-\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n-\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n+\t\t\t\tpop_commit(&list);\n \t\t\t}\n \t\t}\n \t}\n@@ -2085,6 +2102,7 @@ int write_commit_graph(struct object_directory *odb,\n \tuint32_t i, count_distinct = 0;\n \tint res = 0;\n \tint replace = 0;\n+\tstruct topo_level_slab topo_levels;\n \n \tif (!commit_graph_compatible(the_repository))\n \t\treturn 0;\n@@ -2099,6 +2117,9 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->changed_paths = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n \tctx->total_bloom_filter_data_size = 0;\n \n+\tinit_topo_level_slab(&topo_levels);\n+\tctx->topo_levels = &topo_levels;\n+\n \tif (ctx->split) {\n \t\tstruct commit_graph *g;\n \t\tprepare_commit_graph(ctx->r);\n@@ -2197,7 +2218,7 @@ int write_commit_graph(struct object_directory *odb,\n \t} else\n \t\tctx->num_commit_graphs_after = 1;\n \n-\tcompute_generation_numbers(ctx);\n+\tcompute_corrected_commit_date_offsets(ctx);\n \n \tif (ctx->changed_paths)\n \t\tcompute_bloom_filters(ctx);\n@@ -2325,8 +2346,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \tfor (i = 0; i < g->num_commits; i++) {\n \t\tstruct commit *graph_commit, *odb_commit;\n \t\tstruct commit_list *graph_parents, *odb_parents;\n-\t\tuint32_t max_generation = 0;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t max_parent_corrected_commit_date = 0;\n+\t\ttimestamp_t corrected_commit_date;\n \n \t\tdisplay_progress(progress, i + 1);\n \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n@@ -2365,9 +2386,9 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t\t\t\t     oid_to_hex(&graph_parents->item->object.oid),\n \t\t\t\t\t     oid_to_hex(&odb_parents->item->object.oid));\n \n-\t\t\tgeneration = commit_graph_generation(graph_parents->item);\n-\t\t\tif (generation > max_generation)\n-\t\t\t\tmax_generation = generation;\n+\t\t\tcorrected_commit_date = commit_graph_generation(graph_parents->item);\n+\t\t\tif (corrected_commit_date > max_parent_corrected_commit_date)\n+\t\t\t\tmax_parent_corrected_commit_date = corrected_commit_date;\n \n \t\t\tgraph_parents = graph_parents->next;\n \t\t\todb_parents = odb_parents->next;\n@@ -2389,20 +2410,12 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n \t\t\tcontinue;\n \n-\t\t/*\n-\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n-\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n-\t\t * extra logic in the following condition.\n-\t\t */\n-\t\tif (max_generation == GENERATION_NUMBER_MAX)\n-\t\t\tmax_generation--;\n-\n-\t\tgeneration = commit_graph_generation(graph_commit);\n-\t\tif (generation != max_generation + 1)\n-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n+\t\tcorrected_commit_date = commit_graph_generation(graph_commit);\n+\t\tif (corrected_commit_date < max_parent_corrected_commit_date + 1)\n+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n \t\t\t\t     oid_to_hex(&cur_oid),\n-\t\t\t\t     generation,\n-\t\t\t\t     max_generation + 1);\n+\t\t\t\t     corrected_commit_date,\n+\t\t\t\t     max_parent_corrected_commit_date + 1);\n \n \t\tif (graph_commit->date != odb_commit->date)\n \t\t\tgraph_report(_(\"commit date for commit %s in commit-graph is %\"PRItime\" != %\"PRItime),\ndiff --git a/commit-graph.h b/commit-graph.h\nindex e3d4ba96f4..20c5848587 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -138,13 +138,13 @@ void disable_commit_graph(struct repository *r);\n \n struct commit_graph_data {\n \tuint32_t graph_pos;\n-\tuint32_t generation;\n+\ttimestamp_t generation;\n };\n \n /*\n  * Commits should be parsed before accessing generation, graph positions.\n  */\n-uint32_t commit_graph_generation(const struct commit *);\n+timestamp_t commit_graph_generation(const struct commit *);\n uint32_t commit_graph_position(const struct commit *);\n \n int compare_commits_by_gen(const void *_a, const void *_b);\ndiff --git a/commit-reach.c b/commit-reach.c\nindex c83cc291e7..2ce9867ff3 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -32,12 +32,12 @@ static int queue_has_nonstale(struct prio_queue *queue)\n static struct commit_list *paint_down_to_common(struct repository *r,\n \t\t\t\t\t\tstruct commit *one, int n,\n \t\t\t\t\t\tstruct commit **twos,\n-\t\t\t\t\t\tint min_generation)\n+\t\t\t\t\t\ttimestamp_t min_generation)\n {\n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \tstruct commit_list *result = NULL;\n \tint i;\n-\tuint32_t last_gen = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t last_gen = GENERATION_NUMBER_V2_INFINITY;\n \n \tif (!min_generation)\n \t\tqueue.compare = compare_commits_by_commit_date;\n@@ -58,10 +58,10 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \t\tstruct commit *commit = prio_queue_get(&queue);\n \t\tstruct commit_list *parents;\n \t\tint flags;\n-\t\tuint32_t generation = commit_graph_generation(commit);\n+\t\ttimestamp_t generation = commit_graph_generation(commit);\n \n \t\tif (min_generation && generation > last_gen)\n-\t\t\tBUG(\"bad generation skip %8x > %8x at %s\",\n+\t\t\tBUG(\"bad generation skip %\"PRItime\" > %\"PRItime\" at %s\",\n \t\t\t    generation, last_gen,\n \t\t\t    oid_to_hex(&commit->object.oid));\n \t\tlast_gen = generation;\n@@ -177,12 +177,12 @@ static int remove_redundant(struct repository *r, struct commit **array, int cnt\n \t\trepo_parse_commit(r, array[i]);\n \tfor (i = 0; i < cnt; i++) {\n \t\tstruct commit_list *common;\n-\t\tuint32_t min_generation = commit_graph_generation(array[i]);\n+\t\ttimestamp_t min_generation = commit_graph_generation(array[i]);\n \n \t\tif (redundant[i])\n \t\t\tcontinue;\n \t\tfor (j = filled = 0; j < cnt; j++) {\n-\t\t\tuint32_t curr_generation;\n+\t\t\ttimestamp_t curr_generation;\n \t\t\tif (i == j || redundant[j])\n \t\t\t\tcontinue;\n \t\t\tfilled_index[filled] = j;\n@@ -321,7 +321,7 @@ int repo_in_merge_bases_many(struct repository *r, struct commit *commit,\n {\n \tstruct commit_list *bases;\n \tint ret = 0, i;\n-\tuint32_t generation, min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t generation, min_generation = GENERATION_NUMBER_V2_INFINITY;\n \n \tif (repo_parse_commit(r, commit))\n \t\treturn ret;\n@@ -470,7 +470,7 @@ static int in_commit_list(const struct commit_list *want, struct commit *c)\n static enum contains_result contains_test(struct commit *candidate,\n \t\t\t\t\t  const struct commit_list *want,\n \t\t\t\t\t  struct contains_cache *cache,\n-\t\t\t\t\t  uint32_t cutoff)\n+\t\t\t\t\t  timestamp_t cutoff)\n {\n \tenum contains_result *cached = contains_cache_at(cache, candidate);\n \n@@ -506,11 +506,11 @@ static enum contains_result contains_tag_algo(struct commit *candidate,\n {\n \tstruct contains_stack contains_stack = { 0, 0, NULL };\n \tenum contains_result result;\n-\tuint32_t cutoff = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t cutoff = GENERATION_NUMBER_V2_INFINITY;\n \tconst struct commit_list *p;\n \n \tfor (p = want; p; p = p->next) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = p->item;\n \t\tload_commit_graph_info(the_repository, c);\n \t\tgeneration = commit_graph_generation(c);\n@@ -565,7 +565,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation)\n+\t\t\t\t timestamp_t min_generation)\n {\n \tstruct commit **list = NULL;\n \tint i;\n@@ -666,13 +666,13 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \ttime_t min_commit_date = cutoff_by_min_date ? from->item->date : 0;\n \tstruct commit_list *from_iter = from, *to_iter = to;\n \tint result;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_V2_INFINITY;\n \n \twhile (from_iter) {\n \t\tadd_object_array(&from_iter->item->object, NULL, &from_objs);\n \n \t\tif (!parse_commit(from_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (from_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = from_iter->item->date;\n \n@@ -686,7 +686,7 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \n \twhile (to_iter) {\n \t\tif (!parse_commit(to_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (to_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = to_iter->item->date;\n \n@@ -726,13 +726,13 @@ struct commit_list *get_reachable_subset(struct commit **from, int nr_from,\n \tstruct commit_list *found_commits = NULL;\n \tstruct commit **to_last = to + nr_to;\n \tstruct commit **from_last = from + nr_from;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_V2_INFINITY;\n \tint num_to_find = 0;\n \n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \n \tfor (item = to; item < to_last; item++) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = *item;\n \n \t\tparse_commit(c);\ndiff --git a/commit-reach.h b/commit-reach.h\nindex b49ad71a31..148b56fea5 100644\n--- a/commit-reach.h\n+++ b/commit-reach.h\n@@ -87,7 +87,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation);\n+\t\t\t\t timestamp_t min_generation);\n int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \t\t       int commit_date_cutoff);\n \ndiff --git a/commit.h b/commit.h\nindex e901538909..dd17a81672 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -15,6 +15,9 @@\n #define GENERATION_NUMBER_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n \n+#define GENERATION_NUMBER_V2_INFINITY ((1ULL << 63) - 1)\n+#define GENERATION_NUMBER_V2_OFFSET_MAX 0xFFFFFFFF\n+\n struct commit_list {\n \tstruct commit *item;\n \tstruct commit_list *next;\ndiff --git a/revision.c b/revision.c\nindex 23287d26c3..b978e79601 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -725,7 +725,7 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \tif (!revs->repo->objects->commit_graph)\n \t\treturn -1;\n \n-\tif (commit_graph_generation(commit) == GENERATION_NUMBER_INFINITY)\n+\tif (commit_graph_generation(commit) == GENERATION_NUMBER_V2_INFINITY)\n \t\treturn -1;\n \n \tfilter = get_bloom_filter(revs->repo, commit, 0);\n@@ -3270,7 +3270,7 @@ define_commit_slab(indegree_slab, int);\n define_commit_slab(author_date_slab, timestamp_t);\n \n struct topo_walk_info {\n-\tuint32_t min_generation;\n+\ttimestamp_t min_generation;\n \tstruct prio_queue explore_queue;\n \tstruct prio_queue indegree_queue;\n \tstruct prio_queue topo_queue;\n@@ -3316,7 +3316,7 @@ static void explore_walk_step(struct rev_info *revs)\n }\n \n static void explore_to_depth(struct rev_info *revs,\n-\t\t\t     uint32_t gen_cutoff)\n+\t\t\t     timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3359,7 +3359,7 @@ static void indegree_walk_step(struct rev_info *revs)\n }\n \n static void compute_indegrees_to_depth(struct rev_info *revs,\n-\t\t\t\t       uint32_t gen_cutoff)\n+\t\t\t\t       timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3414,10 +3414,10 @@ static void init_topo_walk(struct rev_info *revs)\n \tinfo->explore_queue.compare = compare_commits_by_gen_then_commit_date;\n \tinfo->indegree_queue.compare = compare_commits_by_gen_then_commit_date;\n \n-\tinfo->min_generation = GENERATION_NUMBER_INFINITY;\n+\tinfo->min_generation = GENERATION_NUMBER_V2_INFINITY;\n \tfor (list = revs->commits; list; list = list->next) {\n \t\tstruct commit *c = list->item;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (parse_commit_gently(c, 1))\n \t\t\tcontinue;\n@@ -3478,7 +3478,7 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n \tfor (p = commit->parents; p; p = p->next) {\n \t\tstruct commit *parent = p->item;\n \t\tint *pi;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (parent->object.flags & UNINTERESTING)\n \t\t\tcontinue;\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 3ec5248d70..43801f07a5 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -596,7 +596,7 @@ test_expect_success 'detect incorrect generation number' '\n \t\t\"generation for commit\"\n '\n \n-test_expect_success 'detect incorrect generation number' '\n+test_expect_failure 'detect incorrect generation number' '\n \tcorrupt_graph_and_verify $GRAPH_BYTE_COMMIT_GENERATION \"\\00\" \\\n \t\t\"non-zero generation number\"\n '\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 951a2b23aa..db2332e687 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -489,7 +489,7 @@ static int got_oid(struct upload_pack_data *data,\n \n static int ok_to_give_up(struct upload_pack_data *data)\n {\n-\tuint32_t min_generation = GENERATION_NUMBER_ZERO;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_ZERO;\n \n \tif (!data->have_obj.nr)\n \t\treturn 0;\n-- \ngitgitgadget\n"},{"id":"402191","messageId":"b997b649-cfeb-4b55-9c83-1c0ee2a5677c@gmail.com","threadId":"53933","inReplyTo":"d23f67dc80b85abe4eba9a9dfc39d50188e23bb7.1595927632.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/6] revision: parse parent in indegree_walk_step()","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-07-28T13:00:51Z","receivedAt":"2020-07-28T13:00:56Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/28/2020 5:13 AM, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> In indegree_walk_step(), we add unvisited parents to the indegree queue.\n> However, parents are not guaranteed to be parsed. As the indegree queue\n> sorts by generation number, let's parse parents before inserting them to\n> ensure the correct priority order.\n\nYou mentioned this in your blog post. I'm sorry that such a small\nissue caused you pain. Perhaps you could summarize a little bit of\nhow that investigation led you to find this issue?\n\nQuestion: is this something that is only necessary when we change\nthe generation number, or is it something that is only _exposed_\nby the test suite when we change the generation number? It seems that\nit is likely to be an existing bug, but it might be hard to expose\nin a test case.\n\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  revision.c | 3 +++\n>  1 file changed, 3 insertions(+)\n> \n> diff --git a/revision.c b/revision.c\n> index 6aa7f4f567..23287d26c3 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -3343,6 +3343,9 @@ static void indegree_walk_step(struct rev_info *revs)\n>  \t\tstruct commit *parent = p->item;\n>  \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n>  \n> +\t\tif (parse_commit_gently(parent, 1) < 0)\n> +\t\t\treturn ;\n\nDrop the extra space.\n\nThanks,\n-Stolee\n"},{"id":"402192","messageId":"a9d50995-566d-cad2-ff67-8b8604b52eed@gmail.com","threadId":"53933","inReplyTo":"701f5912369c0fcc07cf604c3129cb5017a125ce.1595927632.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 3/6] commit-graph: consolidate fill_commit_graph_info","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-07-28T13:14:42Z","receivedAt":"2020-07-28T13:14:46Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/28/2020 5:13 AM, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> Both fill_commit_graph_info() and fill_commit_in_graph() parse\n> information present in commit data chunk. Let's simplify the\n> implementation by calling fill_commit_graph_info() within\n> fill_commit_in_graph().\n> \n> The test 'generate tar with future mtime' creates a commit with commit\n> time of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\n> generation number and has undefined behavior. The test used to pass as\n> fill_commit_in_graph() did not read commit time from commit graph,\n> reading commit date from odb instead.\n\nI was first confused as to why fill_commit_graph_info() did not\nload the timestamp, but the reason is that it is only used by\ntwo methods:\n\n1. fill_commit_in_graph(): this actually leaves the commit in a\n   \"parsed\" state, so the date must be correct. Thus, it parses\n   the date out of the commit-graph.\n\n2. load_commit_graph_info(): this only helps to guarantee we\n   know the graph_pos and generation number values.\n\nPerhaps add this extra context: you will _need_ the commit date\nfrom the commit-graph in order to populate the generation number\nv2 in fill_commit_graph_info().\n\n> Let's fix that by setting commit time of (2 ^ 34 - 1) seconds.\n\nThe timestamp limit placed in the commit-graph is more restrictive\nthan 64-bit timestamps, but as your test points out, the maximum\ntimestamp allowed takes place in the year 2514. That is far enough\naway for all real data.\n\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c      | 31 ++++++++++++-------------------\n>  t/t5000-tar-tree.sh |  4 ++--\n>  2 files changed, 14 insertions(+), 21 deletions(-)\n> \n> diff --git a/commit-graph.c b/commit-graph.c\n> index 5d3c9bd23c..204eb454b2 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -735,15 +735,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \tconst unsigned char *commit_data;\n>  \tstruct commit_graph_data *graph_data;\n>  \tuint32_t lex_index;\n> +\tuint64_t date_high, date_low;\n>  \n>  \twhile (pos < g->num_commits_in_base)\n>  \t\tg = g->base_graph;\n>  \n> +\tif (pos >= g->num_commits + g->num_commits_in_base)\n> +\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n> +\n>  \tlex_index = pos - g->num_commits_in_base;\n>  \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n>  \n>  \tgraph_data = commit_graph_data_at(item);\n>  \tgraph_data->graph_pos = pos;\n> +\n> +\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> +\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> +\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> +\n>  \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n>  }\n>  \n> @@ -758,38 +767,22 @@ static int fill_commit_in_graph(struct repository *r,\n>  {\n>  \tuint32_t edge_value;\n>  \tuint32_t *parent_data_ptr;\n> -\tuint64_t date_low, date_high;\n>  \tstruct commit_list **pptr;\n> -\tstruct commit_graph_data *graph_data;\n>  \tconst unsigned char *commit_data;\n>  \tuint32_t lex_index;\n>  \n> +\tfill_commit_graph_info(item, g, pos);\n> +\n>  \twhile (pos < g->num_commits_in_base)\n>  \t\tg = g->base_graph;\n\nThis 'while' loop happens in both implementations, so you could\nsave a miniscule amount of time by placing the call to\nfill_commit_graph_info() after the while loop.\n\n> -\tif (pos >= g->num_commits + g->num_commits_in_base)\n> -\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n\n> -\t/*\n> -\t * Store the \"full\" position, but then use the\n> -\t * \"local\" position for the rest of the calculation.\n> -\t */\n> -\tgraph_data = commit_graph_data_at(item);\n> -\tgraph_data->graph_pos = pos;\n>  \tlex_index = pos - g->num_commits_in_base;\n> -\n> -\tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n> +\tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n\nI was about to complain about this change, but GRAPH_DATA_WIDTH\nis a macro that does an equivalent thing (except the_hash_algo->rawsz\ninstead of g->hash_len).\n\n>  \n>  \titem->object.parsed = 1;\n>  \n>  \tset_commit_tree(item, NULL);\n>  \n> -\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> -\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> -\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> -\n> -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> -\n>  \tpptr = &item->parents;\n>  \n>  \tedge_value = get_be32(commit_data + g->hash_len);\n> diff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\n> index 37655a237c..1986354fc3 100755\n> --- a/t/t5000-tar-tree.sh\n> +++ b/t/t5000-tar-tree.sh\n> @@ -406,7 +406,7 @@ test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n>  \trm -f .git/index &&\n>  \techo content >file &&\n>  \tgit add file &&\n> -\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n> +\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n>  \t\tgit commit -m \"tempori parendum\"\n>  '\n>  \n> @@ -415,7 +415,7 @@ test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n>  '\n>  \n>  test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n> -\techo 4147 >expect &&\n> +\techo 2514 >expect &&\n>  \ttar_info future.tar | cut -d\" \" -f2 >actual &&\n>  \ttest_cmp expect actual\n>  '\n> \n\nThanks,\n-Stolee\n"},{"id":"402194","messageId":"20200728145458.GA87373@syl.lan","threadId":"53933","inReplyTo":"pull.676.git.1595927632.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/6] [GSoC] Implement Corrected Commit Date","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-07-28T14:54:58Z","receivedAt":"2020-07-28T14:55:03Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Hi Abhishek,\n\nOn Tue, Jul 28, 2020 at 09:13:45AM +0000, Abhishek Kumar via GitGitGadget wrote:\n> This patch series implements the corrected commit date offsets as generation\n> number v2, along with other pre-requisites.\n\nVery exciting. I have been eagerly following your blog and asking\nStolee about your progress, so I am excited to read these patches.\n\n> Git uses topological levels in the commit-graph file for commit-graph\n> traversal operations like git log --graph. Unfortunately, using topological\n> levels can result in a worse performance than without them when compared\n> with committer date as a heuristics. For example, git merge-base v4.8 v4.9\n> on the Linux repository walks 635,579 commits using topological levels and\n> walks 167,468 using committer date.\n>\n> Thus, the need for generation number v2 was born. New generation number\n> needed to provide good performance, increment updates, and backward\n> compatibility. Due to an unfortunate problem, we also needed a way to\n> distinguish between the old and new generation number without incrementing\n> graph version.\n>\n> Various candidates were examined (https://github.com/derrickstolee/gen-test,\n> https://github.com/abhishekkumar2718/git/pull/1). The proposed generation\n> number v2, Corrected Commit Date with Mononotically Increasing Offsets\n> performed much worse than committer date (506,577 vs. 167,468 commits walked\n> for git merge-base v4.8 v4.9) and was dropped.\n>\n> Using Generation Data chunk (GDAT) relieves the requirement of backward\n> compatibility as we would continue to store topological levels in Commit\n> Data (CDAT) chunk. Thus, Corrected Commit Date was chosen as generation\n> number v2. The Corrected Commit Date is defined as:\n>\n> For a commit C, let its corrected commit date be the maximum of the commit\n> date of C and the corrected commit dates of its parents. Then corrected\n> commit date offset is the difference between corrected commit date of C and\n> commit date of C.\n\nInterestingly, we use a very similar metric at GitHub to sort commits in\nvarious UI views which have lots of existing machinery that sorts\nan abstract collection by each element's \"date\". Since that sort is\nstable, and we want to respect the order that Git delivered, we take the\npairwise max of each successive pair of commits.\n\n> We will introduce an additional commit-graph chunk, Generation Data chunk,\n> and store corrected commit date offsets in GDAT chunk while storing\n> topological levels in CDAT chunk. The old versions of Git would ignore GDAT\n> chunk, using topological levels from CDAT chunk. In contrast, new versions\n> of Git would use corrected commit dates, falling back to topological level\n> if the generation data chunk is absent in the commit-graph file.\n\nI'm sure that I'll learn more when I get to this point, but I would like\nto hear more about why you want to store the offset rather than the\ncorrected commit date itself. It seems that the offset could be either\npositive or negative, so you'd only have the range of a signed integer\n(rather than storing 8 bytes of a time_t for the full breadth of\npossibilities).\n\nI know also that Peff is working on negative timestamp support, so I\nwould want to hear about what he thinks of this, too.\n\n> Here's what left for the PR (which I intend to take on with the second\n> version of pull request):\n>\n>  1. Add an option to skip writing generation data chunk (to test whether new\n>     Git works without GDAT as intended).\n\nThis will be good to gradually roll-out the new chunk. Another thought\nis to control whether or not the commit-graph machinery _reads_ this\nchunk if it's present. That can be useful for debugging too (eg., I have\na commit-graph with a GDAT chunk that is broken in some way, what\nhappens if I don't read that chunk?)\n\nMaybe something like `commitgraph.readsGenerationData`? Incidentally,\nI'm preparing a `commitgraph.readsChangedPaths` to control whether or\nnot we read the Bloom index and data chunks. I'll send that to the list\nshortly (it's in my fork somewhere if you want an earlier look), but\nthat may be a useful reference for you.\n\n>  2. Handle writing to commit-graph for mismatched version (that is, merging\n>     all graphs into a new graph with a GDAT chunk).\n>  3. Update technical documentation.\n>\n> I look forward to everyone's reviews!\n>\n> Thanks\n>\n>  * Abhishek\n>\n>\n> ----------------------------------------------------------------------------\n>\n> The build fails for t9807-git-p4-submit.sh on osx-clang, which I feel is\n> unrelated to my code changes. Still need to investigate further.\n>\n> Abhishek Kumar (6):\n>   commit-graph: fix regression when computing bloom filter\n>   revision: parse parent in indegree_walk_step()\n>   commit-graph: consolidate fill_commit_graph_info\n>   commit-graph: consolidate compare_commits_by_gen\n>   commit-graph: implement generation data chunk\n>   commit-graph: implement corrected commit date offset\n>\n>  blame.c                       |   2 +-\n>  commit-graph.c                | 181 +++++++++++++++++++++-------------\n>  commit-graph.h                |   7 +-\n>  commit-reach.c                |  47 +++------\n>  commit-reach.h                |   2 +-\n>  commit.c                      |   9 +-\n>  commit.h                      |   3 +\n>  revision.c                    |  17 ++--\n>  t/helper/test-read-graph.c    |   2 +\n>  t/t4216-log-bloom.sh          |   4 +-\n>  t/t5000-tar-tree.sh           |   4 +-\n>  t/t5318-commit-graph.sh       |  21 ++--\n>  t/t5324-split-commit-graph.sh |  12 +--\n>  upload-pack.c                 |   2 +-\n>  14 files changed, 178 insertions(+), 135 deletions(-)\n>\n>\n> base-commit: 47ae905ffb98cc4d4fd90083da6bc8dab55d9ecc\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-676%2Fabhishekkumar2718%2Fcorrected_commit_date-v1\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-676/abhishekkumar2718/corrected_commit_date-v1\n> Pull-Request: https://github.com/gitgitgadget/git/pull/676\n> --\n> gitgitgadget\n\nThanks,\nTaylor\n"},{"id":"402195","messageId":"d4a613c1-f3e8-3789-2548-8344c4b976e9@web.de","threadId":"53933","inReplyTo":"a9d50995-566d-cad2-ff67-8b8604b52eed@gmail.com","subject":"Re: [PATCH 3/6] commit-graph: consolidate fill_commit_graph_info","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2020-07-28T15:19:26Z","receivedAt":"2020-07-28T15:19:32Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"[Had to remove stolee@gmail.com because with it my mail provider\n rejected this email with the following error message:\n\n   Requested action not taken: mailbox unavailable\n   invalid DNS MX or A/AAAA resource record.]\n\nAm 28.07.20 um 15:14 schrieb Derrick Stolee:\n> On 7/28/2020 5:13 AM, Abhishek Kumar via GitGitGadget wrote:\n>> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>>\n>> Both fill_commit_graph_info() and fill_commit_in_graph() parse\n>> information present in commit data chunk. Let's simplify the\n>> implementation by calling fill_commit_graph_info() within\n>> fill_commit_in_graph().\n>>\n>> The test 'generate tar with future mtime' creates a commit with commit\n>> time of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\n>> generation number and has undefined behavior. The test used to pass as\n>> fill_commit_in_graph() did not read commit time from commit graph,\n>> reading commit date from odb instead.\n>\n> I was first confused as to why fill_commit_graph_info() did not\n> load the timestamp, but the reason is that it is only used by\n> two methods:\n>\n> 1. fill_commit_in_graph(): this actually leaves the commit in a\n>    \"parsed\" state, so the date must be correct. Thus, it parses\n>    the date out of the commit-graph.\n>\n> 2. load_commit_graph_info(): this only helps to guarantee we\n>    know the graph_pos and generation number values.\n>\n> Perhaps add this extra context: you will _need_ the commit date\n> from the commit-graph in order to populate the generation number\n> v2 in fill_commit_graph_info().\n>\n>> Let's fix that by setting commit time of (2 ^ 34 - 1) seconds.\n>\n> The timestamp limit placed in the commit-graph is more restrictive\n> than 64-bit timestamps, but as your test points out, the maximum\n> timestamp allowed takes place in the year 2514. That is far enough\n> away for all real data.\n\nWe all may feel like the end of the world is imminent, but do we really\nneed to set such an arbitrary limit?  OK, that limit was already set two\nyears ago, and I'm really late.  But still: It's sad to see anything\nelse than signed 64-bit timestamps to be used in fresh code (after Y2K).\nThe extra four bytes would fatten up the structures less than the\ntransition from SHA-1 to SHA-256 will, and no bit twiddling would be\nrequired.  *sigh*\n\nRené\n"},{"id":"402196","messageId":"20200728152844.GB87373@syl.lan","threadId":"53933","inReplyTo":"91e6e97a66aff88e0b860e34659dddc3396c7f28.1595927632.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/6] commit-graph: fix regression when computing bloom filter","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-07-28T15:28:44Z","receivedAt":"2020-07-28T15:28:50Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Jul 28, 2020 at 09:13:46AM +0000, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> With 3d112755 (commit-graph: examine commits by generation number), Git\n> knew to sort by generation number before examining the diff when not\n> using pack order. c49c82aa (commit: move members graph_pos, generation\n> to a slab, 2020-06-17) moved generation number into a slab and\n> introduced a helper which returns GENERATION_NUMBER_INFINITY when\n> writing the graph. Sorting is no longer useful and essentially reverts\n> the earlier commit.\n\nThis last sentence is slightly confusing. Do you think it would be more\nclear if you said elaborated a bit? Perhaps something like:\n\n  [...]\n\n  commit_gen_cmp is used when writing a commit-graph to sort commits in\n  generation order before computing Bloom filters. Since c49c82aa made\n  it so that 'commit_graph_generation()' returns\n  'GENERATION_NUMBER_INFINITY' during writing, we cannot call it within\n  this function. Instead, access the generation number directly through\n  the slab (i.e., by calling 'commit_graph_data_at(c)->generation') in\n  order to access it while writing.\n\nI think the above would be a good extra paragraph in the commit message\nprovided that you remove the sentence beginning with \"Sorting is no\nlonger useful...\"\n\n> Let's fix this by accessing generation number directly through the slab.\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 5 +++--\n>  1 file changed, 3 insertions(+), 2 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 1af68c297d..5d3c9bd23c 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -144,8 +144,9 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>  \tconst struct commit *a = *(const struct commit **)va;\n>  \tconst struct commit *b = *(const struct commit **)vb;\n>\n> -\tuint32_t generation_a = commit_graph_generation(a);\n> -\tuint32_t generation_b = commit_graph_generation(b);\n> +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n> +\n\nNit; this whitespace diff is extraneous, but it's not hurting anything\neither. Since it looks like you're rerolling anyway, it would be good to\njust get rid of it.\n\nOtherwise this fix makes sense to me.\n\n>  \t/* lower generation commits first */\n>  \tif (generation_a < generation_b)\n>  \t\treturn -1;\n> --\n> gitgitgadget\n\nThanks,\nTaylor\n"},{"id":"402197","messageId":"20200728153042.GC87373@syl.lan","threadId":"53933","inReplyTo":"b997b649-cfeb-4b55-9c83-1c0ee2a5677c@gmail.com","subject":"Re: [PATCH 2/6] revision: parse parent in indegree_walk_step()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-07-28T15:30:42Z","receivedAt":"2020-07-28T15:30:47Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Jul 28, 2020 at 09:00:51AM -0400, Derrick Stolee wrote:\n> On 7/28/2020 5:13 AM, Abhishek Kumar via GitGitGadget wrote:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > In indegree_walk_step(), we add unvisited parents to the indegree queue.\n> > However, parents are not guaranteed to be parsed. As the indegree queue\n> > sorts by generation number, let's parse parents before inserting them to\n> > ensure the correct priority order.\n>\n> You mentioned this in your blog post. I'm sorry that such a small\n> issue caused you pain. Perhaps you could summarize a little bit of\n> how that investigation led you to find this issue?\n\nIndeed ;-). I feel like forgetting to call 'parse_commit_gently()' is a\nrite of passage for this part of the code in some sense.\n\n> Question: is this something that is only necessary when we change\n> the generation number, or is it something that is only _exposed_\n> by the test suite when we change the generation number? It seems that\n> it is likely to be an existing bug, but it might be hard to expose\n> in a test case.\n\nI tend to agree that this bug probably existed before Abhishek's\nchanges, but that it's probably more trouble than it's worth to tickle\nwith a test case. So, I'd be fine with this fix as it is (provided that\nthe style nit is addressed below, too).\n\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  revision.c | 3 +++\n> >  1 file changed, 3 insertions(+)\n> >\n> > diff --git a/revision.c b/revision.c\n> > index 6aa7f4f567..23287d26c3 100644\n> > --- a/revision.c\n> > +++ b/revision.c\n> > @@ -3343,6 +3343,9 @@ static void indegree_walk_step(struct rev_info *revs)\n> >  \t\tstruct commit *parent = p->item;\n> >  \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n> >\n> > +\t\tif (parse_commit_gently(parent, 1) < 0)\n> > +\t\t\treturn ;\n>\n> Drop the extra space.\n>\n> Thanks,\n> -Stolee\n\nThanks,\nTaylor\n"},{"id":"402203","messageId":"e8646aaa-667f-b7d8-f8f2-efbaaeb8877d@gmail.com","threadId":"53933","inReplyTo":"647290d0368e385227614dd1822aa9083b0dba5e.1595927632.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 6/6] commit-graph: implement corrected commit date offset","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-07-28T15:55:12Z","receivedAt":"2020-07-28T15:55:17Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/28/2020 5:13 AM, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> With preparations done,...\n\nI feel like this commit could have been made smaller by doing the\nuint32_t -> timestamp_t conversion in a separate patch. That would\nmake it easier to focus on the changes to the generation number v2\nlogic.\n\n> let's implement corrected commit date offset.\n> We add a new commit-slab to store topological levels while writing\n\nIt's important to add: we store topological levels to ensure that older\nversions of Git will still have the performance benefits from generation\nnumber v1.\n\n> commit graph and upgrade number of struct commit_graph_data to 64-bits.\n\nDo you mean \"update the generation member in struct commit_graph_data\nto a 64-bit timestamp\"? The struct itself also has the 32-bit graph_pos\nmember.\n\n> We have to touch many files, upgrading generation number from uint32_t\n> to timestamp_t.\n\nYes, that's why I recommend doing that in a different step.\n\n> We drop 'detect incorrect generation number' from t5318-commit-graph.sh,\n> which tests if verify can detect if a commit graph have\n> GENERATION_NUMBER_ZERO for a commit, followed by a non-zero generation.\n> With corrected commit dates, GENERATION_NUMBER_ZERO is possible only if\n> one of dates is Unix epoch zero.\n\nWhat about the topological levels? Are we caring about verifying the data\nthat we start to ignore in this new version? I'm hesitant to drop this\nright now, but I'm open to it if we really don't see it as a valuable test.\n\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  blame.c                 |   2 +-\n>  commit-graph.c          | 109 ++++++++++++++++++++++------------------\n>  commit-graph.h          |   4 +-\n>  commit-reach.c          |  32 ++++++------\n>  commit-reach.h          |   2 +-\n>  commit.h                |   3 ++\n>  revision.c              |  14 +++---\n>  t/t5318-commit-graph.sh |   2 +-\n>  upload-pack.c           |   2 +-\n>  9 files changed, 93 insertions(+), 77 deletions(-)\n> \n> diff --git a/blame.c b/blame.c\n> index 82fa16d658..48aa632461 100644\n> --- a/blame.c\n> +++ b/blame.c\n> @@ -1272,7 +1272,7 @@ static int maybe_changed_path(struct repository *r,\n>  \tif (!bd)\n>  \t\treturn 1;\n>  \n> -\tif (commit_graph_generation(origin->commit) == GENERATION_NUMBER_INFINITY)\n> +\tif (commit_graph_generation(origin->commit) == GENERATION_NUMBER_V2_INFINITY)\n>  \t\treturn 1;\n\nI don't see value in changing the name of this macro. It\nis only used as the default value for a commit not in the\ncommit-graph. Changing its value to 0xFFFFFFFF works for\nboth versions when the type is updated to timestamp_t.\n\nThe actually-important change in this patch (not just the\ntype change) is here:\n\n> -static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> +static void compute_corrected_commit_date_offsets(struct write_commit_graph_context *ctx)\n>  {\n>  \tint i;\n>  \tstruct commit_list *list = NULL;\n> @@ -1326,11 +1334,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n>  \t\t\t\t\tctx->commits.nr);\n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n> -\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n> +\t\tuint32_t topo_level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n>  \n>  \t\tdisplay_progress(ctx->progress, i + 1);\n> -\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n> -\t\t    generation != GENERATION_NUMBER_ZERO)\n> +\t\tif (topo_level != GENERATION_NUMBER_INFINITY &&\n> +\t\t    topo_level != GENERATION_NUMBER_ZERO)\n>  \t\t\tcontinue;\n\nHere, our \"skip\" condition is that the topo_level has been computed.\nThis should be fine, as we are never reading that out of the commit-graph.\nWe will never be in a mode where topo_level is computed but corrected\ncommit-date is not.\n\n>  \t\tcommit_list_insert(ctx->commits.list[i], &list);\n> @@ -1338,29 +1346,38 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\tstruct commit *current = list->item;\n>  \t\t\tstruct commit_list *parent;\n>  \t\t\tint all_parents_computed = 1;\n> -\t\t\tuint32_t max_generation = 0;\n> +\t\t\tuint32_t max_level = 0;\n> +\t\t\ttimestamp_t max_corrected_commit_date = current->date;\n\nLater you assign data->generation to be \"max_corrected_commit_date + 1\",\nwhich made me think this should be \"current->date - 1\". Is that so? Or,\ndo we want most offsets to be one instead of zero? Is there value there?\n\n>  \n>  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n> -\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n> +\t\t\t\ttopo_level = *topo_level_slab_at(ctx->topo_levels, parent->item);\n>  \n> -\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n> -\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n> +\t\t\t\tif (topo_level == GENERATION_NUMBER_INFINITY ||\n> +\t\t\t\t    topo_level == GENERATION_NUMBER_ZERO) {\n>  \t\t\t\t\tall_parents_computed = 0;\n>  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n>  \t\t\t\t\tbreak;\n> -\t\t\t\t} else if (generation > max_generation) {\n> -\t\t\t\t\tmax_generation = generation;\n> +\t\t\t\t} else {\n> +\t\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(parent->item);\n> +\n> +\t\t\t\t\tif (topo_level > max_level)\n> +\t\t\t\t\t\tmax_level = topo_level;\n> +\n> +\t\t\t\t\tif (data->generation > max_corrected_commit_date)\n> +\t\t\t\t\t\tmax_corrected_commit_date = data->generation;\n>  \t\t\t\t}\n>  \t\t\t}\n>  \n>  \t\t\tif (all_parents_computed) {\n>  \t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n>  \n> -\t\t\t\tdata->generation = max_generation + 1;\n> -\t\t\t\tpop_commit(&list);\n> +\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n> +\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n> +\n> +\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n> +\t\t\t\tdata->generation = max_corrected_commit_date + 1;\n>  \n> -\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n> -\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n> +\t\t\t\tpop_commit(&list);\n>  \t\t\t}\n>  \t\t}\n>  \t}\n\nThis looks correct, and I've done a tiny bit of perf tests locally.\n\n> @@ -2085,6 +2102,7 @@ int write_commit_graph(struct object_directory *odb,\n>  \tuint32_t i, count_distinct = 0;\n>  \tint res = 0;\n>  \tint replace = 0;\n> +\tstruct topo_level_slab topo_levels;\n>  \n>  \tif (!commit_graph_compatible(the_repository))\n>  \t\treturn 0;\n> @@ -2099,6 +2117,9 @@ int write_commit_graph(struct object_directory *odb,\n>  \tctx->changed_paths = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n>  \tctx->total_bloom_filter_data_size = 0;\n>  \n> +\tinit_topo_level_slab(&topo_levels);\n> +\tctx->topo_levels = &topo_levels;\n> +\n>  \tif (ctx->split) {\n>  \t\tstruct commit_graph *g;\n>  \t\tprepare_commit_graph(ctx->r);\n> @@ -2197,7 +2218,7 @@ int write_commit_graph(struct object_directory *odb,\n>  \t} else\n>  \t\tctx->num_commit_graphs_after = 1;\n>  \n> -\tcompute_generation_numbers(ctx);\n> +\tcompute_corrected_commit_date_offsets(ctx);\n\nThis rename might not be necessary. You are computing both\nversions (v1 and v2) so the name change is actually less\naccurate than the old name.\n\nThanks,\n-Stolee\n\n"},{"id":"402204","messageId":"542e98f1-f793-5290-02ae-3e4706765b80@gmail.com","threadId":"53933","inReplyTo":"d4a613c1-f3e8-3789-2548-8344c4b976e9@web.de","subject":"Re: [PATCH 3/6] commit-graph: consolidate fill_commit_graph_info","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-07-28T15:58:28Z","receivedAt":"2020-07-28T15:58:32Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/28/2020 11:19 AM, René Scharfe wrote:\n> [Had to remove stolee@gmail.com because with it my mail provider\n>  rejected this email with the following error message:\n> \n>    Requested action not taken: mailbox unavailable\n>    invalid DNS MX or A/AAAA resource record.]\n> \n> Am 28.07.20 um 15:14 schrieb Derrick Stolee:\n>> On 7/28/2020 5:13 AM, Abhishek Kumar via GitGitGadget wrote:\n>>> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>>>\n>>> Both fill_commit_graph_info() and fill_commit_in_graph() parse\n>>> information present in commit data chunk. Let's simplify the\n>>> implementation by calling fill_commit_graph_info() within\n>>> fill_commit_in_graph().\n>>>\n>>> The test 'generate tar with future mtime' creates a commit with commit\n>>> time of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\n>>> generation number and has undefined behavior. The test used to pass as\n>>> fill_commit_in_graph() did not read commit time from commit graph,\n>>> reading commit date from odb instead.\n>>\n>> I was first confused as to why fill_commit_graph_info() did not\n>> load the timestamp, but the reason is that it is only used by\n>> two methods:\n>>\n>> 1. fill_commit_in_graph(): this actually leaves the commit in a\n>>    \"parsed\" state, so the date must be correct. Thus, it parses\n>>    the date out of the commit-graph.\n>>\n>> 2. load_commit_graph_info(): this only helps to guarantee we\n>>    know the graph_pos and generation number values.\n>>\n>> Perhaps add this extra context: you will _need_ the commit date\n>> from the commit-graph in order to populate the generation number\n>> v2 in fill_commit_graph_info().\n>>\n>>> Let's fix that by setting commit time of (2 ^ 34 - 1) seconds.\n>>\n>> The timestamp limit placed in the commit-graph is more restrictive\n>> than 64-bit timestamps, but as your test points out, the maximum\n>> timestamp allowed takes place in the year 2514. That is far enough\n>> away for all real data.\n> \n> We all may feel like the end of the world is imminent, but do we really\n> need to set such an arbitrary limit?  OK, that limit was already set two\n> years ago, and I'm really late.  But still: It's sad to see anything\n> else than signed 64-bit timestamps to be used in fresh code (after Y2K).\n> The extra four bytes would fatten up the structures less than the\n> transition from SHA-1 to SHA-256 will, and no bit twiddling would be\n> required.  *sigh*\n\nOne thing to consider after generation number v2 is out long enough\nis if we could drop the topo-levels and write zeroes for the topo-\nlevel portion. This was valid data in the first version of the\ncommit-graph, so it would still be valid. Then, we could allow\nfull 64-bit timestamps again.\n\nThis is something to think about again in a year, maybe.\n\nThanks,\n-Stolee\n\n"},{"id":"402205","messageId":"20200728160134.GD87373@syl.lan","threadId":"53933","inReplyTo":"a9d50995-566d-cad2-ff67-8b8604b52eed@gmail.com","subject":"Re: [PATCH 3/6] commit-graph: consolidate fill_commit_graph_info","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-07-28T16:01:34Z","receivedAt":"2020-07-28T16:01:40Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Jul 28, 2020 at 09:14:42AM -0400, Derrick Stolee wrote:\n> On 7/28/2020 5:13 AM, Abhishek Kumar via GitGitGadget wrote:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > Both fill_commit_graph_info() and fill_commit_in_graph() parse\n> > information present in commit data chunk. Let's simplify the\n> > implementation by calling fill_commit_graph_info() within\n> > fill_commit_in_graph().\n> >\n> > The test 'generate tar with future mtime' creates a commit with commit\n> > time of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\n> > generation number and has undefined behavior. The test used to pass as\n> > fill_commit_in_graph() did not read commit time from commit graph,\n> > reading commit date from odb instead.\n>\n> I was first confused as to why fill_commit_graph_info() did not\n> load the timestamp, but the reason is that it is only used by\n> two methods:\n>\n> 1. fill_commit_in_graph(): this actually leaves the commit in a\n>    \"parsed\" state, so the date must be correct. Thus, it parses\n>    the date out of the commit-graph.\n>\n> 2. load_commit_graph_info(): this only helps to guarantee we\n>    know the graph_pos and generation number values.\n>\n> Perhaps add this extra context: you will _need_ the commit date\n> from the commit-graph in order to populate the generation number\n> v2 in fill_commit_graph_info().\n>\n> > Let's fix that by setting commit time of (2 ^ 34 - 1) seconds.\n>\n> The timestamp limit placed in the commit-graph is more restrictive\n> than 64-bit timestamps, but as your test points out, the maximum\n> timestamp allowed takes place in the year 2514. That is far enough\n> away for all real data.\n>\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c      | 31 ++++++++++++-------------------\n> >  t/t5000-tar-tree.sh |  4 ++--\n> >  2 files changed, 14 insertions(+), 21 deletions(-)\n> >\n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index 5d3c9bd23c..204eb454b2 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -735,15 +735,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n> >  \tconst unsigned char *commit_data;\n> >  \tstruct commit_graph_data *graph_data;\n> >  \tuint32_t lex_index;\n> > +\tuint64_t date_high, date_low;\n> >\n> >  \twhile (pos < g->num_commits_in_base)\n> >  \t\tg = g->base_graph;\n> >\n> > +\tif (pos >= g->num_commits + g->num_commits_in_base)\n> > +\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n> > +\n> >  \tlex_index = pos - g->num_commits_in_base;\n> >  \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n> >\n> >  \tgraph_data = commit_graph_data_at(item);\n> >  \tgraph_data->graph_pos = pos;\n> > +\n> > +\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> > +\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> > +\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> > +\n> >  \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> >  }\n> >\n> > @@ -758,38 +767,22 @@ static int fill_commit_in_graph(struct repository *r,\n> >  {\n> >  \tuint32_t edge_value;\n> >  \tuint32_t *parent_data_ptr;\n> > -\tuint64_t date_low, date_high;\n> >  \tstruct commit_list **pptr;\n> > -\tstruct commit_graph_data *graph_data;\n> >  \tconst unsigned char *commit_data;\n> >  \tuint32_t lex_index;\n> >\n> > +\tfill_commit_graph_info(item, g, pos);\n> > +\n> >  \twhile (pos < g->num_commits_in_base)\n> >  \t\tg = g->base_graph;\n>\n> This 'while' loop happens in both implementations, so you could\n> save a miniscule amount of time by placing the call to\n> fill_commit_graph_info() after the while loop.\n>\n> > -\tif (pos >= g->num_commits + g->num_commits_in_base)\n> > -\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n>\n> > -\t/*\n> > -\t * Store the \"full\" position, but then use the\n> > -\t * \"local\" position for the rest of the calculation.\n> > -\t */\n> > -\tgraph_data = commit_graph_data_at(item);\n> > -\tgraph_data->graph_pos = pos;\n> >  \tlex_index = pos - g->num_commits_in_base;\n> > -\n> > -\tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n> > +\tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n>\n> I was about to complain about this change, but GRAPH_DATA_WIDTH\n> is a macro that does an equivalent thing (except the_hash_algo->rawsz\n> instead of g->hash_len).\n>\n> >\n> >  \titem->object.parsed = 1;\n> >\n> >  \tset_commit_tree(item, NULL);\n> >\n> > -\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> > -\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> > -\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> > -\n> > -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> > -\n> >  \tpptr = &item->parents;\n> >\n> >  \tedge_value = get_be32(commit_data + g->hash_len);\n> > diff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\n> > index 37655a237c..1986354fc3 100755\n> > --- a/t/t5000-tar-tree.sh\n> > +++ b/t/t5000-tar-tree.sh\n> > @@ -406,7 +406,7 @@ test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n> >  \trm -f .git/index &&\n> >  \techo content >file &&\n> >  \tgit add file &&\n> > -\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n> > +\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n> >  \t\tgit commit -m \"tempori parendum\"\n> >  '\n> >\n> > @@ -415,7 +415,7 @@ test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n> >  '\n> >\n> >  test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n> > -\techo 4147 >expect &&\n> > +\techo 2514 >expect &&\n> >  \ttar_info future.tar | cut -d\" \" -f2 >actual &&\n> >  \ttest_cmp expect actual\n> >  '\n> >\n>\n> Thanks,\n> -Stolee\n\nAgreed with Stolee's review, but otherwise this looks like a faithful\ntransformation.\n\nThanks,\nTaylor\n"},{"id":"402206","messageId":"20200728160326.GE87373@syl.lan","threadId":"53933","inReplyTo":"812fe75fc7252db0b7b6604f84b17dcf7324b922.1595927632.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 4/6] commit-graph: consolidate compare_commits_by_gen","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-07-28T16:03:26Z","receivedAt":"2020-07-28T16:03:31Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Jul 28, 2020 at 09:13:49AM +0000, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> Comparing commits by generation has been independently defined twice, in\n> commit-reach and commit. Let's simplify the implementation by moving\n> compare_commits_by_gen() to commit-graph.\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 15 +++++++++++++++\n>  commit-graph.h |  2 ++\n>  commit-reach.c | 15 ---------------\n>  commit.c       |  9 +++------\n>  4 files changed, 20 insertions(+), 21 deletions(-)\n\nAll looks good to me.\n\n  Reviewed-by: Taylor Blau <me@ttaylorr.com>\n\nThanks,\nTaylor\n"},{"id":"402207","messageId":"20200728161250.GF87373@syl.lan","threadId":"53933","inReplyTo":"80ea7da3435396edcb19423ab602962d31585209.1595927632.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 5/6] commit-graph: implement generation data chunk","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-07-28T16:12:50Z","receivedAt":"2020-07-28T16:12:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Jul 28, 2020 at 09:13:50AM +0000, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> One of the essential pre-requisites before implementing generation\n> number as to distinguish between generation numbers v1 and v2 while\n\ns/as/is\n\n> still being compatible with old Git.\n\nMaybe you could add a section here to talk about why this is needed\nspecifically? That is, you mention it's a prerequisite, but a reader in\na year or two may not remember why. Adding that information here would\nbe good.\n\n> We are going to introduce a new chunk called Generation Data chunk (or\n> GDAT). GDAT stores generation number v2 (and any subsequent versions),\n> whereas CDAT will still store topological level.\n>\n> Old Git does not understand GDAT chunk and would ignore it, reading\n> topological levels from CDAT. Newer versions of Git can parse GDAT and\n> take advantage of newer generation numbers, falling back to topological\n> levels when GDAT chunk is missing (as it would happen with a commit\n> graph written by old Git).\n\n...this is exactly the paragraph that I was looking for above. Could you\nswap the order of these last two paragraphs? I think that it would make\nthe patch message far clearer.\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c                | 33 +++++++++++++++++++++++++++++----\n>  commit-graph.h                |  1 +\n>  t/helper/test-read-graph.c    |  2 ++\n>  t/t4216-log-bloom.sh          |  4 ++--\n>  t/t5318-commit-graph.sh       | 19 +++++++++++--------\n>  t/t5324-split-commit-graph.sh | 12 ++++++------\n>  6 files changed, 51 insertions(+), 20 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 1c98f38d69..ab714f4a76 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -38,11 +38,12 @@ void git_test_write_commit_graph_or_die(void)\n>  #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n>  #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n>  #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n> +#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n>  #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n>  #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n>  #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n>  #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n> -#define MAX_NUM_CHUNKS 7\n> +#define MAX_NUM_CHUNKS 8\n\nUgh. I am simultaneously working on a new chunk myself (so a bad\nconflict resolution would look at both of us incrementing this number\nto the same value without generating a conflict.)\n\nI think the right thing to do here would be to define an enum over chunk\nnames, and then index an array by that enum (where the value at each\nindex is the chunk identifier). Then, the last value of that enum would\nbe a '__COUNT' which you could use to initialize the array (as well as\nwithin the commit-graph writing routines).\n\nAnyway, I think that it's probably not worth it in the meantime, but it\nis something that Junio should look out for when merging (if yours and\nmy topic happen to get merged around the same time, which they may not).\n\n>  #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n>\n> @@ -389,6 +390,13 @@ struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size)\n>  \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n>  \t\t\tbreak;\n>\n> +\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n> +\t\t\tif (graph->chunk_generation_data)\n> +\t\t\t\tchunk_repeated = 1;\n> +\t\t\telse\n> +\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n> +\t\t\tbreak;\n> +\n>  \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n>  \t\t\tif (graph->chunk_extra_edges)\n>  \t\t\t\tchunk_repeated = 1;\n> @@ -768,7 +776,10 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \tdate_low = get_be32(commit_data + g->hash_len + 12);\n>  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n>\n> -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> +\tif (g->chunk_generation_data)\n> +\t\tgraph_data->generation = get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n> +\telse\n> +\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n>  }\n>\n>  static inline void set_commit_tree(struct commit *c, struct tree *t)\n> @@ -1100,6 +1111,17 @@ static void write_graph_chunk_data(struct hashfile *f, int hash_len,\n>  \t}\n>  }\n>\n> +static void write_graph_chunk_generation_data(struct hashfile *f,\n> +\t\t\t\t\t      struct write_commit_graph_context *ctx)\n> +{\n> +\tstruct commit **list = ctx->commits.list;\n> +\tint count;\n> +\tfor (count = 0; count < ctx->commits.nr; count++, list++) {\n> +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n> +\t\thashwrite_be32(f, commit_graph_data_at(*list)->generation);\n> +\t}\n> +}\n> +\n\nThis pointer arithmetic is not necessary. Why not like:\n\n  int i;\n  for (i = 0; i < ctx->commits.nr; i++) {\n    struct commit *c = ctx->commits.list[i];\n    display_progress(ctx->progress, ++ctx->progress_cnt);\n    hashwrite_be32(f, commit_graph_data_at(c)->generation);\n  }\n\ninstead?\n\n>  static void write_graph_chunk_extra_edges(struct hashfile *f,\n>  \t\t\t\t\t  struct write_commit_graph_context *ctx)\n>  {\n> @@ -1605,7 +1627,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \tuint64_t chunk_offsets[MAX_NUM_CHUNKS + 1];\n>  \tconst unsigned hashsz = the_hash_algo->rawsz;\n>  \tstruct strbuf progress_title = STRBUF_INIT;\n> -\tint num_chunks = 3;\n> +\tint num_chunks = 4;\n>  \tstruct object_id file_hash;\n>  \tconst struct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n>\n> @@ -1656,6 +1678,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \tchunk_ids[0] = GRAPH_CHUNKID_OIDFANOUT;\n>  \tchunk_ids[1] = GRAPH_CHUNKID_OIDLOOKUP;\n>  \tchunk_ids[2] = GRAPH_CHUNKID_DATA;\n> +\tchunk_ids[3] = GRAPH_CHUNKID_GENERATION_DATA;\n>  \tif (ctx->num_extra_edges) {\n>  \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_EXTRAEDGES;\n>  \t\tnum_chunks++;\n> @@ -1677,8 +1700,9 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \tchunk_offsets[1] = chunk_offsets[0] + GRAPH_FANOUT_SIZE;\n>  \tchunk_offsets[2] = chunk_offsets[1] + hashsz * ctx->commits.nr;\n>  \tchunk_offsets[3] = chunk_offsets[2] + (hashsz + 16) * ctx->commits.nr;\n> +\tchunk_offsets[4] = chunk_offsets[3] + sizeof(uint32_t) * ctx->commits.nr;\n>\n> -\tnum_chunks = 3;\n> +\tnum_chunks = 4;\n>  \tif (ctx->num_extra_edges) {\n>  \t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n>  \t\t\t\t\t\t4 * ctx->num_extra_edges;\n> @@ -1728,6 +1752,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \twrite_graph_chunk_fanout(f, ctx);\n>  \twrite_graph_chunk_oids(f, hashsz, ctx);\n>  \twrite_graph_chunk_data(f, hashsz, ctx);\n> +\twrite_graph_chunk_generation_data(f, ctx);\n>  \tif (ctx->num_extra_edges)\n>  \t\twrite_graph_chunk_extra_edges(f, ctx);\n>  \tif (ctx->changed_paths) {\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 98cc5a3b9d..e3d4ba96f4 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -67,6 +67,7 @@ struct commit_graph {\n>  \tconst uint32_t *chunk_oid_fanout;\n>  \tconst unsigned char *chunk_oid_lookup;\n>  \tconst unsigned char *chunk_commit_data;\n> +\tconst unsigned char *chunk_generation_data;\n>  \tconst unsigned char *chunk_extra_edges;\n>  \tconst unsigned char *chunk_base_graphs;\n>  \tconst unsigned char *chunk_bloom_indexes;\n> diff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\n> index 6d0c962438..1c2a5366c7 100644\n> --- a/t/helper/test-read-graph.c\n> +++ b/t/helper/test-read-graph.c\n> @@ -32,6 +32,8 @@ int cmd__read_graph(int argc, const char **argv)\n>  \t\tprintf(\" oid_lookup\");\n>  \tif (graph->chunk_commit_data)\n>  \t\tprintf(\" commit_metadata\");\n> +\tif (graph->chunk_generation_data)\n> +\t\tprintf(\" generation_data\");\n>  \tif (graph->chunk_extra_edges)\n>  \t\tprintf(\" extra_edges\");\n>  \tif (graph->chunk_bloom_indexes)\n> diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n> index c855bcd3e7..780855e691 100755\n> --- a/t/t4216-log-bloom.sh\n> +++ b/t/t4216-log-bloom.sh\n> @@ -33,11 +33,11 @@ test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n>  \tgit commit-graph write --reachable --changed-paths\n>  '\n>  graph_read_expect () {\n> -\tNUM_CHUNKS=5\n> +\tNUM_CHUNKS=6\n>  \tcat >expect <<- EOF\n>  \theader: 43475048 1 1 $NUM_CHUNKS 0\n>  \tnum_commits: $1\n> -\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n> +\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n>  \tEOF\n>  \ttest-tool read-graph >actual &&\n>  \ttest_cmp expect actual\n> diff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\n> index 26f332d6a3..3ec5248d70 100755\n> --- a/t/t5318-commit-graph.sh\n> +++ b/t/t5318-commit-graph.sh\n> @@ -71,16 +71,16 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n>\n>  graph_read_expect() {\n>  \tOPTIONAL=\"\"\n> -\tNUM_CHUNKS=3\n> +\tNUM_CHUNKS=4\n>  \tif test ! -z $2\n>  \tthen\n>  \t\tOPTIONAL=\" $2\"\n> -\t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n> +\t\tNUM_CHUNKS=$((4 + $(echo \"$2\" | wc -w)))\n>  \tfi\n>  \tcat >expect <<- EOF\n>  \theader: 43475048 1 1 $NUM_CHUNKS 0\n>  \tnum_commits: $1\n> -\tchunks: oid_fanout oid_lookup commit_metadata$OPTIONAL\n> +\tchunks: oid_fanout oid_lookup commit_metadata generation_data$OPTIONAL\n>  \tEOF\n>  \ttest-tool read-graph >output &&\n>  \ttest_cmp expect output\n> @@ -433,7 +433,7 @@ GRAPH_BYTE_HASH=5\n>  GRAPH_BYTE_CHUNK_COUNT=6\n>  GRAPH_CHUNK_LOOKUP_OFFSET=8\n>  GRAPH_CHUNK_LOOKUP_WIDTH=12\n> -GRAPH_CHUNK_LOOKUP_ROWS=5\n> +GRAPH_CHUNK_LOOKUP_ROWS=6\n>  GRAPH_BYTE_OID_FANOUT_ID=$GRAPH_CHUNK_LOOKUP_OFFSET\n>  GRAPH_BYTE_OID_LOOKUP_ID=$(($GRAPH_CHUNK_LOOKUP_OFFSET + \\\n>  \t\t\t    1 * $GRAPH_CHUNK_LOOKUP_WIDTH))\n> @@ -451,11 +451,14 @@ GRAPH_BYTE_COMMIT_TREE=$GRAPH_COMMIT_DATA_OFFSET\n>  GRAPH_BYTE_COMMIT_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN))\n>  GRAPH_BYTE_COMMIT_EXTRA_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 4))\n>  GRAPH_BYTE_COMMIT_WRONG_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 3))\n> -GRAPH_BYTE_COMMIT_GENERATION=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 11))\n>  GRAPH_BYTE_COMMIT_DATE=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 12))\n>  GRAPH_COMMIT_DATA_WIDTH=$(($HASH_LEN + 16))\n> -GRAPH_OCTOPUS_DATA_OFFSET=$(($GRAPH_COMMIT_DATA_OFFSET + \\\n> -\t\t\t     $GRAPH_COMMIT_DATA_WIDTH * $NUM_COMMITS))\n> +GRAPH_GENERATION_DATA_OFFSET=$(($GRAPH_COMMIT_DATA_OFFSET + \\\n> +\t\t\t\t$GRAPH_COMMIT_DATA_WIDTH * $NUM_COMMITS))\n> +GRAPH_GENERATION_DATA_WIDTH=4\n> +GRAPH_BYTE_COMMIT_GENERATION=$(($GRAPH_GENERATION_DATA_OFFSET + 3))\n> +GRAPH_OCTOPUS_DATA_OFFSET=$(($GRAPH_GENERATION_DATA_OFFSET + \\\n> +\t\t\t     $GRAPH_GENERATION_DATA_WIDTH * $NUM_COMMITS))\n>  GRAPH_BYTE_OCTOPUS=$(($GRAPH_OCTOPUS_DATA_OFFSET + 4))\n>  GRAPH_BYTE_FOOTER=$(($GRAPH_OCTOPUS_DATA_OFFSET + 4 * $NUM_OCTOPUS_EDGES))\n>\n> @@ -594,7 +597,7 @@ test_expect_success 'detect incorrect generation number' '\n>  '\n>\n>  test_expect_success 'detect incorrect generation number' '\n> -\tcorrupt_graph_and_verify $GRAPH_BYTE_COMMIT_GENERATION \"\\01\" \\\n> +\tcorrupt_graph_and_verify $GRAPH_BYTE_COMMIT_GENERATION \"\\00\" \\\n>  \t\t\"non-zero generation number\"\n>  '\n>\n> diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> index 269d0964a3..096a96ec41 100755\n> --- a/t/t5324-split-commit-graph.sh\n> +++ b/t/t5324-split-commit-graph.sh\n> @@ -14,11 +14,11 @@ test_expect_success 'setup repo' '\n>  \tgraphdir=\"$infodir/commit-graphs\" &&\n>  \ttest_oid_init &&\n>  \ttest_oid_cache <<-EOM\n> -\tshallow sha1:1760\n> -\tshallow sha256:2064\n> +\tshallow sha1:2132\n> +\tshallow sha256:2436\n>\n> -\tbase sha1:1376\n> -\tbase sha256:1496\n> +\tbase sha1:1408\n> +\tbase sha256:1528\n>  \tEOM\n>  '\n>\n> @@ -29,9 +29,9 @@ graph_read_expect() {\n>  \t\tNUM_BASE=$2\n>  \tfi\n>  \tcat >expect <<- EOF\n> -\theader: 43475048 1 1 3 $NUM_BASE\n> +\theader: 43475048 1 1 4 $NUM_BASE\n>  \tnum_commits: $1\n> -\tchunks: oid_fanout oid_lookup commit_metadata\n> +\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n>  \tEOF\n>  \ttest-tool read-graph >output &&\n>  \ttest_cmp expect output\n> --\n> gitgitgadget\n>\n\nAll of this looks good to me.\n\nThanks,\nTaylor\n"},{"id":"402208","messageId":"20200728162315.GG87373@syl.lan","threadId":"53933","inReplyTo":"e8646aaa-667f-b7d8-f8f2-efbaaeb8877d@gmail.com","subject":"Re: [PATCH 6/6] commit-graph: implement corrected commit date offset","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-07-28T16:23:15Z","receivedAt":"2020-07-28T16:23:20Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Jul 28, 2020 at 11:55:12AM -0400, Derrick Stolee wrote:\n> On 7/28/2020 5:13 AM, Abhishek Kumar via GitGitGadget wrote:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > With preparations done,...\n>\n> I feel like this commit could have been made smaller by doing the\n> uint32_t -> timestamp_t conversion in a separate patch. That would\n> make it easier to focus on the changes to the generation number v2\n> logic.\n\nYep, agreed.\n\n> > let's implement corrected commit date offset.\n> > We add a new commit-slab to store topological levels while writing\n>\n> It's important to add: we store topological levels to ensure that older\n> versions of Git will still have the performance benefits from generation\n> number v1.\n>\n> > commit graph and upgrade number of struct commit_graph_data to 64-bits.\n>\n> Do you mean \"update the generation member in struct commit_graph_data\n> to a 64-bit timestamp\"? The struct itself also has the 32-bit graph_pos\n> member.\n>\n> > We have to touch many files, upgrading generation number from uint32_t\n> > to timestamp_t.\n>\n> Yes, that's why I recommend doing that in a different step.\n>\n> > We drop 'detect incorrect generation number' from t5318-commit-graph.sh,\n> > which tests if verify can detect if a commit graph have\n> > GENERATION_NUMBER_ZERO for a commit, followed by a non-zero generation.\n> > With corrected commit dates, GENERATION_NUMBER_ZERO is possible only if\n> > one of dates is Unix epoch zero.\n>\n> What about the topological levels? Are we caring about verifying the data\n> that we start to ignore in this new version? I'm hesitant to drop this\n> right now, but I'm open to it if we really don't see it as a valuable test.\n>\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  blame.c                 |   2 +-\n> >  commit-graph.c          | 109 ++++++++++++++++++++++------------------\n> >  commit-graph.h          |   4 +-\n> >  commit-reach.c          |  32 ++++++------\n> >  commit-reach.h          |   2 +-\n> >  commit.h                |   3 ++\n> >  revision.c              |  14 +++---\n> >  t/t5318-commit-graph.sh |   2 +-\n> >  upload-pack.c           |   2 +-\n> >  9 files changed, 93 insertions(+), 77 deletions(-)\n> >\n> > diff --git a/blame.c b/blame.c\n> > index 82fa16d658..48aa632461 100644\n> > --- a/blame.c\n> > +++ b/blame.c\n> > @@ -1272,7 +1272,7 @@ static int maybe_changed_path(struct repository *r,\n> >  \tif (!bd)\n> >  \t\treturn 1;\n> >\n> > -\tif (commit_graph_generation(origin->commit) == GENERATION_NUMBER_INFINITY)\n> > +\tif (commit_graph_generation(origin->commit) == GENERATION_NUMBER_V2_INFINITY)\n> >  \t\treturn 1;\n>\n> I don't see value in changing the name of this macro. It\n> is only used as the default value for a commit not in the\n> commit-graph. Changing its value to 0xFFFFFFFF works for\n> both versions when the type is updated to timestamp_t.\n>\n> The actually-important change in this patch (not just the\n> type change) is here:\n>\n> > -static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> > +static void compute_corrected_commit_date_offsets(struct write_commit_graph_context *ctx)\n> >  {\n> >  \tint i;\n> >  \tstruct commit_list *list = NULL;\n> > @@ -1326,11 +1334,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n> >  \t\t\t\t\tctx->commits.nr);\n> >  \tfor (i = 0; i < ctx->commits.nr; i++) {\n> > -\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n> > +\t\tuint32_t topo_level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n> >\n> >  \t\tdisplay_progress(ctx->progress, i + 1);\n> > -\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n> > -\t\t    generation != GENERATION_NUMBER_ZERO)\n> > +\t\tif (topo_level != GENERATION_NUMBER_INFINITY &&\n> > +\t\t    topo_level != GENERATION_NUMBER_ZERO)\n> >  \t\t\tcontinue;\n>\n> Here, our \"skip\" condition is that the topo_level has been computed.\n> This should be fine, as we are never reading that out of the commit-graph.\n> We will never be in a mode where topo_level is computed but corrected\n> commit-date is not.\n>\n> >  \t\tcommit_list_insert(ctx->commits.list[i], &list);\n> > @@ -1338,29 +1346,38 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\tstruct commit *current = list->item;\n> >  \t\t\tstruct commit_list *parent;\n> >  \t\t\tint all_parents_computed = 1;\n> > -\t\t\tuint32_t max_generation = 0;\n> > +\t\t\tuint32_t max_level = 0;\n> > +\t\t\ttimestamp_t max_corrected_commit_date = current->date;\n>\n> Later you assign data->generation to be \"max_corrected_commit_date + 1\",\n> which made me think this should be \"current->date - 1\". Is that so? Or,\n> do we want most offsets to be one instead of zero? Is there value there?\n>\n> >\n> >  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n> > -\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n> > +\t\t\t\ttopo_level = *topo_level_slab_at(ctx->topo_levels, parent->item);\n> >\n> > -\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n> > -\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n> > +\t\t\t\tif (topo_level == GENERATION_NUMBER_INFINITY ||\n> > +\t\t\t\t    topo_level == GENERATION_NUMBER_ZERO) {\n> >  \t\t\t\t\tall_parents_computed = 0;\n> >  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n> >  \t\t\t\t\tbreak;\n> > -\t\t\t\t} else if (generation > max_generation) {\n> > -\t\t\t\t\tmax_generation = generation;\n> > +\t\t\t\t} else {\n> > +\t\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(parent->item);\n> > +\n> > +\t\t\t\t\tif (topo_level > max_level)\n> > +\t\t\t\t\t\tmax_level = topo_level;\n> > +\n> > +\t\t\t\t\tif (data->generation > max_corrected_commit_date)\n> > +\t\t\t\t\t\tmax_corrected_commit_date = data->generation;\n> >  \t\t\t\t}\n> >  \t\t\t}\n> >\n> >  \t\t\tif (all_parents_computed) {\n> >  \t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n> >\n> > -\t\t\t\tdata->generation = max_generation + 1;\n> > -\t\t\t\tpop_commit(&list);\n> > +\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n> > +\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n> > +\n> > +\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n> > +\t\t\t\tdata->generation = max_corrected_commit_date + 1;\n> >\n> > -\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n> > -\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n> > +\t\t\t\tpop_commit(&list);\n> >  \t\t\t}\n> >  \t\t}\n> >  \t}\n>\n> This looks correct, and I've done a tiny bit of perf tests locally.\n>\n> > @@ -2085,6 +2102,7 @@ int write_commit_graph(struct object_directory *odb,\n> >  \tuint32_t i, count_distinct = 0;\n> >  \tint res = 0;\n> >  \tint replace = 0;\n> > +\tstruct topo_level_slab topo_levels;\n> >\n> >  \tif (!commit_graph_compatible(the_repository))\n> >  \t\treturn 0;\n> > @@ -2099,6 +2117,9 @@ int write_commit_graph(struct object_directory *odb,\n> >  \tctx->changed_paths = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n> >  \tctx->total_bloom_filter_data_size = 0;\n> >\n> > +\tinit_topo_level_slab(&topo_levels);\n> > +\tctx->topo_levels = &topo_levels;\n> > +\n> >  \tif (ctx->split) {\n> >  \t\tstruct commit_graph *g;\n> >  \t\tprepare_commit_graph(ctx->r);\n> > @@ -2197,7 +2218,7 @@ int write_commit_graph(struct object_directory *odb,\n> >  \t} else\n> >  \t\tctx->num_commit_graphs_after = 1;\n> >\n> > -\tcompute_generation_numbers(ctx);\n> > +\tcompute_corrected_commit_date_offsets(ctx);\n>\n> This rename might not be necessary. You are computing both\n> versions (v1 and v2) so the name change is actually less\n> accurate than the old name.\n>\n> Thanks,\n> -Stolee\n\nI don't have anything to add that Stolee hasn't already pointed out.\nThanks for your work on this series, and I'm looking forward to another\nreroll.\n\nThanks,\nTaylor\n"},{"id":"402210","messageId":"2ae8e6e0-9bf7-4e47-2a93-5e5092abe77a@gmail.com","threadId":"53933","inReplyTo":"pull.676.git.1595927632.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/6] [GSoC] Implement Corrected Commit Date","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-07-28T16:35:55Z","receivedAt":"2020-07-28T16:36:02Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/28/2020 5:13 AM, Abhishek Kumar via GitGitGadget wrote:\n> This patch series implements the corrected commit date offsets as generation\n> number v2, along with other pre-requisites.\n> \n> Git uses topological levels in the commit-graph file for commit-graph\n> traversal operations like git log --graph. Unfortunately, using topological\n> levels can result in a worse performance than without them when compared\n> with committer date as a heuristics. For example, git merge-base v4.8 v4.9 \n> on the Linux repository walks 635,579 commits using topological levels and\n> walks 167,468 using committer date.\n> \n> Thus, the need for generation number v2 was born. New generation number\n> needed to provide good performance, increment updates, and backward\n> compatibility. Due to an unfortunate problem, we also needed a way to\n> distinguish between the old and new generation number without incrementing\n> graph version.\n> \n> Various candidates were examined (https://github.com/derrickstolee/gen-test, \n> https://github.com/abhishekkumar2718/git/pull/1). The proposed generation\n> number v2, Corrected Commit Date with Mononotically Increasing Offsets \n> performed much worse than committer date (506,577 vs. 167,468 commits walked\n> for git merge-base v4.8 v4.9) and was dropped.\n> \n> Using Generation Data chunk (GDAT) relieves the requirement of backward\n> compatibility as we would continue to store topological levels in Commit\n> Data (CDAT) chunk. Thus, Corrected Commit Date was chosen as generation\n> number v2. The Corrected Commit Date is defined as:\n> \n> For a commit C, let its corrected commit date be the maximum of the commit\n> date of C and the corrected commit dates of its parents. Then corrected\n> commit date offset is the difference between corrected commit date of C and\n> commit date of C.\n> \n> We will introduce an additional commit-graph chunk, Generation Data chunk,\n> and store corrected commit date offsets in GDAT chunk while storing\n> topological levels in CDAT chunk. The old versions of Git would ignore GDAT\n> chunk, using topological levels from CDAT chunk. In contrast, new versions\n> of Git would use corrected commit dates, falling back to topological level\n> if the generation data chunk is absent in the commit-graph file.\n> \n> Here's what left for the PR (which I intend to take on with the second\n> version of pull request):\n> \n>  1. Add an option to skip writing generation data chunk (to test whether new\n>     Git works without GDAT as intended).\n\nThis would be a good idea, if only as a GIT_TEST_* environment variable.\nI think it important we have a test for the compatibility scenario where\nwe have an \"old\" commit-graph with the new code and test that reading and\nwriting still works properly.\n\n>  2. Handle writing to commit-graph for mismatched version (that is, merging\n>     all graphs into a new graph with a GDAT chunk).\n\nThis is an excellent thing to do. There are a few options when writing an\nincremental commit-graph when the base graphs do not have the GDAT chunk:\n\n   i. Do not write the GDAT chunk unless we are merging all levels\n      (based on the merging strategy).\n\n  ii. Merge all levels, then write the GDAT chunk.\n\n>  3. Update technical documentation.\n\nYes, I was going to ask for a patch that updates\nDocumentation/technical/commit-graph-format.txt.\n\nThis is an excellent v1. A lot of small things, but no\nreally big issues.\n\nThanks,\n-Stolee\n"},{"id":"402456","messageId":"20200730052429.GA50429@Abhishek-Arch","threadId":"53933","inReplyTo":"20200728152844.GB87373@syl.lan","subject":"Re: [PATCH 1/6] commit-graph: fix regression when computing bloom filter","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-07-30T05:24:29Z","receivedAt":"2020-07-30T05:26:36Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Jul 28, 2020 at 11:28:44AM -0400, Taylor Blau wrote:\n> On Tue, Jul 28, 2020 at 09:13:46AM +0000, Abhishek Kumar via GitGitGadget wrote:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > With 3d112755 (commit-graph: examine commits by generation number), Git\n> > knew to sort by generation number before examining the diff when not\n> > using pack order. c49c82aa (commit: move members graph_pos, generation\n> > to a slab, 2020-06-17) moved generation number into a slab and\n> > introduced a helper which returns GENERATION_NUMBER_INFINITY when\n> > writing the graph. Sorting is no longer useful and essentially reverts\n> > the earlier commit.\n> \n> This last sentence is slightly confusing. Do you think it would be more\n> clear if you said elaborated a bit? Perhaps something like:\n> \n>   [...]\n> \n>   commit_gen_cmp is used when writing a commit-graph to sort commits in\n>   generation order before computing Bloom filters. Since c49c82aa made\n>   it so that 'commit_graph_generation()' returns\n>   'GENERATION_NUMBER_INFINITY' during writing, we cannot call it within\n>   this function. Instead, access the generation number directly through\n>   the slab (i.e., by calling 'commit_graph_data_at(c)->generation') in\n>   order to access it while writing.\n> \n\nThanks! That is clearer. Will change.\n\n> I think the above would be a good extra paragraph in the commit message\n> provided that you remove the sentence beginning with \"Sorting is no\n> longer useful...\"\n> \n> > Let's fix this by accessing generation number directly through the slab.\n> >\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c | 5 +++--\n> >  1 file changed, 3 insertions(+), 2 deletions(-)\n> >\n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index 1af68c297d..5d3c9bd23c 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -144,8 +144,9 @@ static int commit_gen_cmp(const void *va, const void *vb)\n> >  \tconst struct commit *a = *(const struct commit **)va;\n> >  \tconst struct commit *b = *(const struct commit **)vb;\n> >\n> > -\tuint32_t generation_a = commit_graph_generation(a);\n> > -\tuint32_t generation_b = commit_graph_generation(b);\n> > +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> > +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n> > +\n> \n> Nit; this whitespace diff is extraneous, but it's not hurting anything\n> either. Since it looks like you're rerolling anyway, it would be good to\n> just get rid of it.\n> \n> Otherwise this fix makes sense to me.\n> \n> >  \t/* lower generation commits first */\n> >  \tif (generation_a < generation_b)\n> >  \t\treturn -1;\n> > --\n> > gitgitgadget\n> \n> Thanks,\n> Taylor\n"},{"id":"402457","messageId":"20200730060732.GB50429@Abhishek-Arch","threadId":"53933","inReplyTo":"a9d50995-566d-cad2-ff67-8b8604b52eed@gmail.com","subject":"Re: [PATCH 3/6] commit-graph: consolidate fill_commit_graph_info","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-07-30T06:07:32Z","receivedAt":"2020-07-30T06:09:35Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Jul 28, 2020 at 09:14:42AM -0400, Derrick Stolee wrote:\n> On 7/28/2020 5:13 AM, Abhishek Kumar via GitGitGadget wrote:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > \n> > Both fill_commit_graph_info() and fill_commit_in_graph() parse\n> > information present in commit data chunk. Let's simplify the\n> > implementation by calling fill_commit_graph_info() within\n> > fill_commit_in_graph().\n> > \n> > The test 'generate tar with future mtime' creates a commit with commit\n> > time of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\n> > generation number and has undefined behavior. The test used to pass as\n> > fill_commit_in_graph() did not read commit time from commit graph,\n> > reading commit date from odb instead.\n> \n> I was first confused as to why fill_commit_graph_info() did not\n> load the timestamp, but the reason is that it is only used by\n> two methods:\n> \n> 1. fill_commit_in_graph(): this actually leaves the commit in a\n>    \"parsed\" state, so the date must be correct. Thus, it parses\n>    the date out of the commit-graph.\n> \n> 2. load_commit_graph_info(): this only helps to guarantee we\n>    know the graph_pos and generation number values.\n> \n> Perhaps add this extra context: you will _need_ the commit date\n> from the commit-graph in order to populate the generation number\n> v2 in fill_commit_graph_info().\n\nThanks, that makes sense. I have revised the commit message to:\n\ncommit-graph: consolidate fill_commit_graph_info\n    \n    Both fill_commit_graph_info() and fill_commit_in_graph() parse\n    information present in commit data chunk. Let's simplify the\n    implementation by calling fill_commit_graph_info() within\n    fill_commit_in_graph().\n    \n    The test 'generate tar with future mtime' creates a commit with commit\n    time of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\n    generation number (within CDAT chunk) and has undefined behavior.\n    \n    The test used to pass as fill_commit_in_graph() guarantees the values of\n    graph position and generation number, and did not load timestamp.\n    However, with corrected commit date we will need load the timestamp as\n    well to populate the generation number.\n> \n> > Let's fix that by setting commit time of (2 ^ 34 - 1) seconds.\n> \n> The timestamp limit placed in the commit-graph is more restrictive\n> than 64-bit timestamps, but as your test points out, the maximum\n> timestamp allowed takes place in the year 2514. That is far enough\n> away for all real data.\n> \n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c      | 31 ++++++++++++-------------------\n> >  t/t5000-tar-tree.sh |  4 ++--\n> >  2 files changed, 14 insertions(+), 21 deletions(-)\n> > \n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index 5d3c9bd23c..204eb454b2 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -735,15 +735,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n> >  \tconst unsigned char *commit_data;\n> >  \tstruct commit_graph_data *graph_data;\n> >  \tuint32_t lex_index;\n> > +\tuint64_t date_high, date_low;\n> >  \n> >  \twhile (pos < g->num_commits_in_base)\n> >  \t\tg = g->base_graph;\n> >  \n> > +\tif (pos >= g->num_commits + g->num_commits_in_base)\n> > +\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n> > +\n> >  \tlex_index = pos - g->num_commits_in_base;\n> >  \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n> >  \n> >  \tgraph_data = commit_graph_data_at(item);\n> >  \tgraph_data->graph_pos = pos;\n> > +\n> > +\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> > +\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> > +\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> > +\n> >  \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> >  }\n> >  \n> > @@ -758,38 +767,22 @@ static int fill_commit_in_graph(struct repository *r,\n> >  {\n> >  \tuint32_t edge_value;\n> >  \tuint32_t *parent_data_ptr;\n> > -\tuint64_t date_low, date_high;\n> >  \tstruct commit_list **pptr;\n> > -\tstruct commit_graph_data *graph_data;\n> >  \tconst unsigned char *commit_data;\n> >  \tuint32_t lex_index;\n> >  \n> > +\tfill_commit_graph_info(item, g, pos);\n> > +\n> >  \twhile (pos < g->num_commits_in_base)\n> >  \t\tg = g->base_graph;\n> \n> This 'while' loop happens in both implementations, so you could\n> save a miniscule amount of time by placing the call to\n> fill_commit_graph_info() after the while loop.\n> \n> > -\tif (pos >= g->num_commits + g->num_commits_in_base)\n> > -\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n> \n> > -\t/*\n> > -\t * Store the \"full\" position, but then use the\n> > -\t * \"local\" position for the rest of the calculation.\n> > -\t */\n> > -\tgraph_data = commit_graph_data_at(item);\n> > -\tgraph_data->graph_pos = pos;\n> >  \tlex_index = pos - g->num_commits_in_base;\n> > -\n> > -\tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n> > +\tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n> \n> I was about to complain about this change, but GRAPH_DATA_WIDTH\n> is a macro that does an equivalent thing (except the_hash_algo->rawsz\n> instead of g->hash_len).\n> \n> >  \n> >  \titem->object.parsed = 1;\n> >  \n> >  \tset_commit_tree(item, NULL);\n> >  \n> > -\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> > -\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> > -\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> > -\n> > -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> > -\n> >  \tpptr = &item->parents;\n> >  \n> >  \tedge_value = get_be32(commit_data + g->hash_len);\n> > diff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\n> > index 37655a237c..1986354fc3 100755\n> > --- a/t/t5000-tar-tree.sh\n> > +++ b/t/t5000-tar-tree.sh\n> > @@ -406,7 +406,7 @@ test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n> >  \trm -f .git/index &&\n> >  \techo content >file &&\n> >  \tgit add file &&\n> > -\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n> > +\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n> >  \t\tgit commit -m \"tempori parendum\"\n> >  '\n> >  \n> > @@ -415,7 +415,7 @@ test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n> >  '\n> >  \n> >  test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n> > -\techo 4147 >expect &&\n> > +\techo 2514 >expect &&\n> >  \ttar_info future.tar | cut -d\" \" -f2 >actual &&\n> >  \ttest_cmp expect actual\n> >  '\n> > \n> \n> Thanks,\n> -Stolee\n"},{"id":"402458","messageId":"20200730065234.GA2395@Abhishek-Arch","threadId":"53933","inReplyTo":"20200728161250.GF87373@syl.lan","subject":"Re: [PATCH 5/6] commit-graph: implement generation data chunk","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-07-30T06:52:34Z","receivedAt":"2020-07-30T06:54:44Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Jul 28, 2020 at 12:12:50PM -0400, Taylor Blau wrote:\n> On Tue, Jul 28, 2020 at 09:13:50AM +0000, Abhishek Kumar via GitGitGadget wrote:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > One of the essential pre-requisites before implementing generation\n> > number as to distinguish between generation numbers v1 and v2 while\n> \n> s/as/is\n> \n> > still being compatible with old Git.\n> \n> Maybe you could add a section here to talk about why this is needed\n> specifically? That is, you mention it's a prerequisite, but a reader in\n> a year or two may not remember why. Adding that information here would\n> be good.\n> \n> > We are going to introduce a new chunk called Generation Data chunk (or\n> > GDAT). GDAT stores generation number v2 (and any subsequent versions),\n> > whereas CDAT will still store topological level.\n> >\n> > Old Git does not understand GDAT chunk and would ignore it, reading\n> > topological levels from CDAT. Newer versions of Git can parse GDAT and\n> > take advantage of newer generation numbers, falling back to topological\n> > levels when GDAT chunk is missing (as it would happen with a commit\n> > graph written by old Git).\n> \n> ...this is exactly the paragraph that I was looking for above. Could you\n> swap the order of these last two paragraphs? I think that it would make\n> the patch message far clearer.\n\nHere's revised commit message:\n\n  commit-graph: implement generation data chunk\n    \n  As discovered by Ævar, we cannot increment graph version to\n  distinguish between generation numbers v1 and v2 [1]. Thus, one of\n  pre-requistes before implementing generation number v2 was to\n  distinguish generation numbers in a backwards compatible manner\n  without increment graph version.\n  \n  We are going to introduce a new chunk called Generation Data chunk (or\n  GDAT). GDAT stores generation number v2 (and any subsequent versions),\n  whereas CDAT will still store topological level.\n  \n  Old Git does not understand GDAT chunk and would ignore it, reading\n  topological levels from CDAT. New Git can parse GDAT and take advantage\n  of newer generation numbers, falling back to topological levels when\n  GDAT chunk is missing (as it would happen with a commit graph written\n  by old Git).\n \n  [1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nFirst paragraph explains why we need this patch (cannot increment graph\nversion) second explains what this patch does (introduce a new chunk)\nand third proves why it works (Old Git ignores GDAT, New Git parses GDAT).\n\nCan we improve this commit message further? \n\n> >\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c                | 33 +++++++++++++++++++++++++++++----\n> >  commit-graph.h                |  1 +\n> >  t/helper/test-read-graph.c    |  2 ++\n> >  t/t4216-log-bloom.sh          |  4 ++--\n> >  t/t5318-commit-graph.sh       | 19 +++++++++++--------\n> >  t/t5324-split-commit-graph.sh | 12 ++++++------\n> >  6 files changed, 51 insertions(+), 20 deletions(-)\n> >\n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index 1c98f38d69..ab714f4a76 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -38,11 +38,12 @@ void git_test_write_commit_graph_or_die(void)\n> >  #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n> >  #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n> >  #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n> > +#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n> >  #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n> >  #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n> >  #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n> >  #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n> > -#define MAX_NUM_CHUNKS 7\n> > +#define MAX_NUM_CHUNKS 8\n> \n> Ugh. I am simultaneously working on a new chunk myself (so a bad\n> conflict resolution would look at both of us incrementing this number\n> to the same value without generating a conflict.)\n> \n> I think the right thing to do here would be to define an enum over chunk\n> names, and then index an array by that enum (where the value at each\n> index is the chunk identifier). Then, the last value of that enum would\n> be a '__COUNT' which you could use to initialize the array (as well as\n> within the commit-graph writing routines).\n> \n> Anyway, I think that it's probably not worth it in the meantime, but it\n> is something that Junio should look out for when merging (if yours and\n> my topic happen to get merged around the same time, which they may not).\n> \n> >  #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n> >\n> > @@ -389,6 +390,13 @@ struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size)\n> >  \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n> >  \t\t\tbreak;\n> >\n> > +\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n> > +\t\t\tif (graph->chunk_generation_data)\n> > +\t\t\t\tchunk_repeated = 1;\n> > +\t\t\telse\n> > +\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n> > +\t\t\tbreak;\n> > +\n> >  \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n> >  \t\t\tif (graph->chunk_extra_edges)\n> >  \t\t\t\tchunk_repeated = 1;\n> > @@ -768,7 +776,10 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n> >  \tdate_low = get_be32(commit_data + g->hash_len + 12);\n> >  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n> >\n> > -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> > +\tif (g->chunk_generation_data)\n> > +\t\tgraph_data->generation = get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n> > +\telse\n> > +\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> >  }\n> >\n> >  static inline void set_commit_tree(struct commit *c, struct tree *t)\n> > @@ -1100,6 +1111,17 @@ static void write_graph_chunk_data(struct hashfile *f, int hash_len,\n> >  \t}\n> >  }\n> >\n> > +static void write_graph_chunk_generation_data(struct hashfile *f,\n> > +\t\t\t\t\t      struct write_commit_graph_context *ctx)\n> > +{\n> > +\tstruct commit **list = ctx->commits.list;\n> > +\tint count;\n> > +\tfor (count = 0; count < ctx->commits.nr; count++, list++) {\n> > +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n> > +\t\thashwrite_be32(f, commit_graph_data_at(*list)->generation);\n> > +\t}\n> > +}\n> > +\n> \n> This pointer arithmetic is not necessary. Why not like:\n> \n>   int i;\n>   for (i = 0; i < ctx->commits.nr; i++) {\n>     struct commit *c = ctx->commits.list[i];\n>     display_progress(ctx->progress, ++ctx->progress_cnt);\n>     hashwrite_be32(f, commit_graph_data_at(c)->generation);\n>   }\n> \n> instead?\n> \n> >  static void write_graph_chunk_extra_edges(struct hashfile *f,\n> >  \t\t\t\t\t  struct write_commit_graph_context *ctx)\n> >  {\n> > @@ -1605,7 +1627,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n> >  \tuint64_t chunk_offsets[MAX_NUM_CHUNKS + 1];\n> >  \tconst unsigned hashsz = the_hash_algo->rawsz;\n> >  \tstruct strbuf progress_title = STRBUF_INIT;\n> > -\tint num_chunks = 3;\n> > +\tint num_chunks = 4;\n> >  \tstruct object_id file_hash;\n> >  \tconst struct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n> >\n> > @@ -1656,6 +1678,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n> >  \tchunk_ids[0] = GRAPH_CHUNKID_OIDFANOUT;\n> >  \tchunk_ids[1] = GRAPH_CHUNKID_OIDLOOKUP;\n> >  \tchunk_ids[2] = GRAPH_CHUNKID_DATA;\n> > +\tchunk_ids[3] = GRAPH_CHUNKID_GENERATION_DATA;\n> >  \tif (ctx->num_extra_edges) {\n> >  \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_EXTRAEDGES;\n> >  \t\tnum_chunks++;\n> > @@ -1677,8 +1700,9 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n> >  \tchunk_offsets[1] = chunk_offsets[0] + GRAPH_FANOUT_SIZE;\n> >  \tchunk_offsets[2] = chunk_offsets[1] + hashsz * ctx->commits.nr;\n> >  \tchunk_offsets[3] = chunk_offsets[2] + (hashsz + 16) * ctx->commits.nr;\n> > +\tchunk_offsets[4] = chunk_offsets[3] + sizeof(uint32_t) * ctx->commits.nr;\n> >\n> > -\tnum_chunks = 3;\n> > +\tnum_chunks = 4;\n> >  \tif (ctx->num_extra_edges) {\n> >  \t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n> >  \t\t\t\t\t\t4 * ctx->num_extra_edges;\n> > @@ -1728,6 +1752,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n> >  \twrite_graph_chunk_fanout(f, ctx);\n> >  \twrite_graph_chunk_oids(f, hashsz, ctx);\n> >  \twrite_graph_chunk_data(f, hashsz, ctx);\n> > +\twrite_graph_chunk_generation_data(f, ctx);\n> >  \tif (ctx->num_extra_edges)\n> >  \t\twrite_graph_chunk_extra_edges(f, ctx);\n> >  \tif (ctx->changed_paths) {\n> > diff --git a/commit-graph.h b/commit-graph.h\n> > index 98cc5a3b9d..e3d4ba96f4 100644\n> > --- a/commit-graph.h\n> > +++ b/commit-graph.h\n> > @@ -67,6 +67,7 @@ struct commit_graph {\n> >  \tconst uint32_t *chunk_oid_fanout;\n> >  \tconst unsigned char *chunk_oid_lookup;\n> >  \tconst unsigned char *chunk_commit_data;\n> > +\tconst unsigned char *chunk_generation_data;\n> >  \tconst unsigned char *chunk_extra_edges;\n> >  \tconst unsigned char *chunk_base_graphs;\n> >  \tconst unsigned char *chunk_bloom_indexes;\n> > diff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\n> > index 6d0c962438..1c2a5366c7 100644\n> > --- a/t/helper/test-read-graph.c\n> > +++ b/t/helper/test-read-graph.c\n> > @@ -32,6 +32,8 @@ int cmd__read_graph(int argc, const char **argv)\n> >  \t\tprintf(\" oid_lookup\");\n> >  \tif (graph->chunk_commit_data)\n> >  \t\tprintf(\" commit_metadata\");\n> > +\tif (graph->chunk_generation_data)\n> > +\t\tprintf(\" generation_data\");\n> >  \tif (graph->chunk_extra_edges)\n> >  \t\tprintf(\" extra_edges\");\n> >  \tif (graph->chunk_bloom_indexes)\n> > diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n> > index c855bcd3e7..780855e691 100755\n> > --- a/t/t4216-log-bloom.sh\n> > +++ b/t/t4216-log-bloom.sh\n> > @@ -33,11 +33,11 @@ test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n> >  \tgit commit-graph write --reachable --changed-paths\n> >  '\n> >  graph_read_expect () {\n> > -\tNUM_CHUNKS=5\n> > +\tNUM_CHUNKS=6\n> >  \tcat >expect <<- EOF\n> >  \theader: 43475048 1 1 $NUM_CHUNKS 0\n> >  \tnum_commits: $1\n> > -\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n> > +\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n> >  \tEOF\n> >  \ttest-tool read-graph >actual &&\n> >  \ttest_cmp expect actual\n> > diff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\n> > index 26f332d6a3..3ec5248d70 100755\n> > --- a/t/t5318-commit-graph.sh\n> > +++ b/t/t5318-commit-graph.sh\n> > @@ -71,16 +71,16 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n> >\n> >  graph_read_expect() {\n> >  \tOPTIONAL=\"\"\n> > -\tNUM_CHUNKS=3\n> > +\tNUM_CHUNKS=4\n> >  \tif test ! -z $2\n> >  \tthen\n> >  \t\tOPTIONAL=\" $2\"\n> > -\t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n> > +\t\tNUM_CHUNKS=$((4 + $(echo \"$2\" | wc -w)))\n> >  \tfi\n> >  \tcat >expect <<- EOF\n> >  \theader: 43475048 1 1 $NUM_CHUNKS 0\n> >  \tnum_commits: $1\n> > -\tchunks: oid_fanout oid_lookup commit_metadata$OPTIONAL\n> > +\tchunks: oid_fanout oid_lookup commit_metadata generation_data$OPTIONAL\n> >  \tEOF\n> >  \ttest-tool read-graph >output &&\n> >  \ttest_cmp expect output\n> > @@ -433,7 +433,7 @@ GRAPH_BYTE_HASH=5\n> >  GRAPH_BYTE_CHUNK_COUNT=6\n> >  GRAPH_CHUNK_LOOKUP_OFFSET=8\n> >  GRAPH_CHUNK_LOOKUP_WIDTH=12\n> > -GRAPH_CHUNK_LOOKUP_ROWS=5\n> > +GRAPH_CHUNK_LOOKUP_ROWS=6\n> >  GRAPH_BYTE_OID_FANOUT_ID=$GRAPH_CHUNK_LOOKUP_OFFSET\n> >  GRAPH_BYTE_OID_LOOKUP_ID=$(($GRAPH_CHUNK_LOOKUP_OFFSET + \\\n> >  \t\t\t    1 * $GRAPH_CHUNK_LOOKUP_WIDTH))\n> > @@ -451,11 +451,14 @@ GRAPH_BYTE_COMMIT_TREE=$GRAPH_COMMIT_DATA_OFFSET\n> >  GRAPH_BYTE_COMMIT_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN))\n> >  GRAPH_BYTE_COMMIT_EXTRA_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 4))\n> >  GRAPH_BYTE_COMMIT_WRONG_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 3))\n> > -GRAPH_BYTE_COMMIT_GENERATION=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 11))\n> >  GRAPH_BYTE_COMMIT_DATE=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 12))\n> >  GRAPH_COMMIT_DATA_WIDTH=$(($HASH_LEN + 16))\n> > -GRAPH_OCTOPUS_DATA_OFFSET=$(($GRAPH_COMMIT_DATA_OFFSET + \\\n> > -\t\t\t     $GRAPH_COMMIT_DATA_WIDTH * $NUM_COMMITS))\n> > +GRAPH_GENERATION_DATA_OFFSET=$(($GRAPH_COMMIT_DATA_OFFSET + \\\n> > +\t\t\t\t$GRAPH_COMMIT_DATA_WIDTH * $NUM_COMMITS))\n> > +GRAPH_GENERATION_DATA_WIDTH=4\n> > +GRAPH_BYTE_COMMIT_GENERATION=$(($GRAPH_GENERATION_DATA_OFFSET + 3))\n> > +GRAPH_OCTOPUS_DATA_OFFSET=$(($GRAPH_GENERATION_DATA_OFFSET + \\\n> > +\t\t\t     $GRAPH_GENERATION_DATA_WIDTH * $NUM_COMMITS))\n> >  GRAPH_BYTE_OCTOPUS=$(($GRAPH_OCTOPUS_DATA_OFFSET + 4))\n> >  GRAPH_BYTE_FOOTER=$(($GRAPH_OCTOPUS_DATA_OFFSET + 4 * $NUM_OCTOPUS_EDGES))\n> >\n> > @@ -594,7 +597,7 @@ test_expect_success 'detect incorrect generation number' '\n> >  '\n> >\n> >  test_expect_success 'detect incorrect generation number' '\n> > -\tcorrupt_graph_and_verify $GRAPH_BYTE_COMMIT_GENERATION \"\\01\" \\\n> > +\tcorrupt_graph_and_verify $GRAPH_BYTE_COMMIT_GENERATION \"\\00\" \\\n> >  \t\t\"non-zero generation number\"\n> >  '\n> >\n> > diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> > index 269d0964a3..096a96ec41 100755\n> > --- a/t/t5324-split-commit-graph.sh\n> > +++ b/t/t5324-split-commit-graph.sh\n> > @@ -14,11 +14,11 @@ test_expect_success 'setup repo' '\n> >  \tgraphdir=\"$infodir/commit-graphs\" &&\n> >  \ttest_oid_init &&\n> >  \ttest_oid_cache <<-EOM\n> > -\tshallow sha1:1760\n> > -\tshallow sha256:2064\n> > +\tshallow sha1:2132\n> > +\tshallow sha256:2436\n> >\n> > -\tbase sha1:1376\n> > -\tbase sha256:1496\n> > +\tbase sha1:1408\n> > +\tbase sha256:1528\n> >  \tEOM\n> >  '\n> >\n> > @@ -29,9 +29,9 @@ graph_read_expect() {\n> >  \t\tNUM_BASE=$2\n> >  \tfi\n> >  \tcat >expect <<- EOF\n> > -\theader: 43475048 1 1 3 $NUM_BASE\n> > +\theader: 43475048 1 1 4 $NUM_BASE\n> >  \tnum_commits: $1\n> > -\tchunks: oid_fanout oid_lookup commit_metadata\n> > +\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n> >  \tEOF\n> >  \ttest-tool read-graph >output &&\n> >  \ttest_cmp expect output\n> > --\n> > gitgitgadget\n> >\n> \n> All of this looks good to me.\n> \n> Thanks,\n> Taylor\n"},{"id":"402459","messageId":"20200730072714.GA964@Abhishek-Arch","threadId":"53933","inReplyTo":"e8646aaa-667f-b7d8-f8f2-efbaaeb8877d@gmail.com","subject":"Re: [PATCH 6/6] commit-graph: implement corrected commit date offset","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-07-30T07:27:14Z","receivedAt":"2020-07-30T07:29:23Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Jul 28, 2020 at 11:55:12AM -0400, Derrick Stolee wrote:\n> On 7/28/2020 5:13 AM, Abhishek Kumar via GitGitGadget wrote:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > \n> > With preparations done,...\n> \n> I feel like this commit could have been made smaller by doing the\n> uint32_t -> timestamp_t conversion in a separate patch. That would\n> make it easier to focus on the changes to the generation number v2\n> logic.\n> \n\nSure, would seperate into two patches.\n\n> > let's implement corrected commit date offset.\n> > We add a new commit-slab to store topological levels while writing\n> \n> It's important to add: we store topological levels to ensure that older\n> versions of Git will still have the performance benefits from generation\n> number v1.\n> \n\nWill do.\n\n> > commit graph and upgrade number of struct commit_graph_data to 64-bits.\n> \n> Do you mean \"update the generation member in struct commit_graph_data\n> to a 64-bit timestamp\"? The struct itself also has the 32-bit graph_pos\n> member.\n> \n\nYes, \"update the generation number\".\n\n> > We have to touch many files, upgrading generation number from uint32_t\n> > to timestamp_t.\n> \n> Yes, that's why I recommend doing that in a different step.\n> \n> > We drop 'detect incorrect generation number' from t5318-commit-graph.sh,\n> > which tests if verify can detect if a commit graph have\n> > GENERATION_NUMBER_ZERO for a commit, followed by a non-zero generation.\n> > With corrected commit dates, GENERATION_NUMBER_ZERO is possible only if\n> > one of dates is Unix epoch zero.\n> \n> What about the topological levels? Are we caring about verifying the data\n> that we start to ignore in this new version? I'm hesitant to drop this\n> right now, but I'm open to it if we really don't see it as a valuable test.\n> \n\nWe haven't tested the scenario \"New Git reads a commit graph without\nGDAT chunk\" yet. Verifying topological levels (along with many of the\nchanged offsets) would be a part of the scenario.\n\nNow that I think about it, those tests should have been included with\nthis patch.\n\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> [...]\n>\n> Later you assign data->generation to be \"max_corrected_commit_date + 1\",\n> which made me think this should be \"current->date - 1\". Is that so? Or,\n> do we want most offsets to be one instead of zero? Is there value there?\n> \n\nDoes it? \n\nI had hoped most of the offsets could have been zero, as we could take\nadvantage of the fact that commit-slab zero initializes values and avoid\na commit-slab access.\n\nRight, What I meant to do was:\n\n        /*\n         * max_parent_corrected_commit_date is initialized with zero and\n         * takes the maximum of\n         * (parent->item->date + commit_graph_data_at(parent->item)->generation)\n        */\n\n        if (max_parent_corrected_commit_date >= current->date)\n        {\n                struct commit_graph_data *data = commit_graph_data_at(current);\n                data->generation = max_parent_corrected_commit_date + 1;\n        }\n\nThanks for pointing this out!\n\n> [...]\n\n- Abhishek\n"},{"id":"402460","messageId":"20200730074701.GA10743@Abhishek-Arch","threadId":"53933","inReplyTo":"20200728145458.GA87373@syl.lan","subject":"Re: [PATCH 0/6] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-07-30T07:47:01Z","receivedAt":"2020-07-30T07:49:08Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Jul 28, 2020 at 10:54:58AM -0400, Taylor Blau wrote:\n> Hi Abhishek,\n> \n> On Tue, Jul 28, 2020 at 09:13:45AM +0000, Abhishek Kumar via GitGitGadget wrote:\n> > This patch series implements the corrected commit date offsets as generation\n> > number v2, along with other pre-requisites.\n> \n> Very exciting. I have been eagerly following your blog and asking\n> Stolee about your progress, so I am excited to read these patches.\n> \n\nI am so glad to hear that!\n\n> \n> [...]\n> \n> I'm sure that I'll learn more when I get to this point, but I would like\n> to hear more about why you want to store the offset rather than the\n> corrected commit date itself. It seems that the offset could be either\n> positive or negative, so you'd only have the range of a signed integer\n> (rather than storing 8 bytes of a time_t for the full breadth of\n> possibilities).\n> \n\nCorrected commit dates are at least as big as the committer date, so the\noffset (i.e. corrected date - committer date) would never be negative.\n\nWe store offsets instead of corrected commit dates because:\n- We save 4 bytes for each commit, which amounts to 7-8% of the size of\n  commit graph file (of course, dependent on the other chunks used).\n- We save some time while writing the commit-graph file too, around\n  ~200ms for the Linux repo.\n\nWhile the savings are modest, writing corrected dates does not offer any\nadvantage that we could think of, at the time.\n\n> I know also that Peff is working on negative timestamp support, so I\n> would want to hear about what he thinks of this, too.\n\nI have read up on Peff's work with negative timestamp support and it's\npretty exciting.\n\n> [...]\n\nThanks\n- Abhishek\n"},{"id":"402782","messageId":"85eeonutj4.fsf@gmail.com","threadId":"53933","inReplyTo":"91e6e97a66aff88e0b860e34659dddc3396c7f28.1595927632.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/6] commit-graph: fix regression when computing bloom filter","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-04T00:46:55Z","receivedAt":"2020-08-04T00:47:01Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> With 3d112755 (commit-graph: examine commits by generation number), Git\n> knew to sort by generation number before examining the diff when not\n> using pack order. c49c82aa (commit: move members graph_pos, generation\n> to a slab, 2020-06-17) moved generation number into a slab and\n> introduced a helper which returns GENERATION_NUMBER_INFINITY when\n> writing the graph. Sorting is no longer useful and essentially reverts\n> the earlier commit.\n>\n> Let's fix this by accessing generation number directly through the slab.\n\nIt looks like unfortunate and unforeseen consequence of putting together\ngraph position and generation number in the commit_graph_data struct.\nDuring writing of the commit-graph file generation number is computed,\nbut graph position is undefined (yet), and commit_graph_generation()\nuses graph_pos field to find if the data for commit is initialized;\nin this case wrongly.\n\nAnyway, when writing the commit graph we first compute generation\nnumber, then (if requested) the changed-paths Bloom filter.  Skipping\nthe unnecessary check is a good thing... assuming that commit_gen_cmp()\nis used only when writing the commit graph, and not when traversing it\n(because then some commits may not have generation number set, and maybe\neven do not have any data on the commit slab) - which is the case.\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 5 +++--\n>  1 file changed, 3 insertions(+), 2 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 1af68c297d..5d3c9bd23c 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n\nWe might want to add function comment either here or in the header that\nthis comparisonn function is to be used only for `git commit-graph\nwrite`, and not for graph traversal (even if similar funnction exists in\nother modules).\n\n> @@ -144,8 +144,9 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>  \tconst struct commit *a = *(const struct commit **)va;\n>  \tconst struct commit *b = *(const struct commit **)vb;\n>\n> -\tuint32_t generation_a = commit_graph_generation(a);\n> -\tuint32_t generation_b = commit_graph_generation(b);\n> +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n> +\n>  \t/* lower generation commits first */\n>  \tif (generation_a < generation_b)\n>  \t\treturn -1;\n\nBest,\n--\nJakub Narębski\n"},{"id":"402789","messageId":"20200804005658.GB75662@syl.lan","threadId":"53933","inReplyTo":"85eeonutj4.fsf@gmail.com","subject":"Re: [PATCH 1/6] commit-graph: fix regression when computing bloom filter","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-08-04T00:56:58Z","receivedAt":"2020-08-04T00:57:02Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Aug 04, 2020 at 02:46:55AM +0200, Jakub Narębski wrote:\n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > With 3d112755 (commit-graph: examine commits by generation number), Git\n> > knew to sort by generation number before examining the diff when not\n> > using pack order. c49c82aa (commit: move members graph_pos, generation\n> > to a slab, 2020-06-17) moved generation number into a slab and\n> > introduced a helper which returns GENERATION_NUMBER_INFINITY when\n> > writing the graph. Sorting is no longer useful and essentially reverts\n> > the earlier commit.\n> >\n> > Let's fix this by accessing generation number directly through the slab.\n>\n> It looks like unfortunate and unforeseen consequence of putting together\n> graph position and generation number in the commit_graph_data struct.\n> During writing of the commit-graph file generation number is computed,\n> but graph position is undefined (yet), and commit_graph_generation()\n> uses graph_pos field to find if the data for commit is initialized;\n> in this case wrongly.\n>\n> Anyway, when writing the commit graph we first compute generation\n> number, then (if requested) the changed-paths Bloom filter.  Skipping\n> the unnecessary check is a good thing... assuming that commit_gen_cmp()\n> is used only when writing the commit graph, and not when traversing it\n> (because then some commits may not have generation number set, and maybe\n> even do not have any data on the commit slab) - which is the case.\n>\n> >\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c | 5 +++--\n> >  1 file changed, 3 insertions(+), 2 deletions(-)\n> >\n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index 1af68c297d..5d3c9bd23c 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n>\n> We might want to add function comment either here or in the header that\n> this comparisonn function is to be used only for `git commit-graph\n> write`, and not for graph traversal (even if similar funnction exists in\n> other modules).\n\nI think that probably within the function is just fine, and that we can\navoid touching commit-graph.h here.\n\n>\n> > @@ -144,8 +144,9 @@ static int commit_gen_cmp(const void *va, const void *vb)\n> >  \tconst struct commit *a = *(const struct commit **)va;\n> >  \tconst struct commit *b = *(const struct commit **)vb;\n\nMaybe something like:\n\n  /*\n   * Access the generation number directly with\n   * 'commit_graph_data_at(...)->generation' instead of going through\n   * the slab as usual to avoid accessing a yet-uncomputed value.\n   */\n\nFolks that are curious for more can blame this commit and read there.\nI'd err on the side of being brief in the code comment and verbose in\nthe commit message than the other way around ;).\n\n> >\n> > -\tuint32_t generation_a = commit_graph_generation(a);\n> > -\tuint32_t generation_b = commit_graph_generation(b);\n> > +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> > +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n> > +\n> >  \t/* lower generation commits first */\n> >  \tif (generation_a < generation_b)\n> >  \t\treturn -1;\n>\n> Best,\n> --\n> Jakub Narębski\n\nThanks,\nTaylor\n"},{"id":"402803","messageId":"857duevo9m.fsf@gmail.com","threadId":"53933","inReplyTo":"85eeonutj4.fsf@gmail.com","subject":"Re: [PATCH 1/6] commit-graph: fix regression when computing bloom filter","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-04T07:55:17Z","receivedAt":"2020-08-04T07:55:22Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jakub Narębski <jnareb@gmail.com> writes:\n\n[...]\n>> @@ -144,8 +144,9 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>>  \tconst struct commit *a = *(const struct commit **)va;\n>>  \tconst struct commit *b = *(const struct commit **)vb;\n>>\n>> -\tuint32_t generation_a = commit_graph_generation(a);\n>> -\tuint32_t generation_b = commit_graph_generation(b);\n>> +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n>> +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n>> +\n>>  \t/* lower generation commits first */\n>>  \tif (generation_a < generation_b)\n>>  \t\treturn -1;\n\nNOTE: One more thing: we would want to check if corrected commit date\n(generation number v2) or topological level (generation number v1) is\nbetter for this purpose, that is gives better performance.\n\nThe commit 3d11275505 (commit-graph: examine commits by generation\nnumber) which introduced using commit_gen_cmp when writing commit graph\nwhen finding commits via `--reachable` flags describes the following\nperformance improvement:\n\n    On the Linux kernel repository, this change reduced the computation\n    time for 'git commit-graph write --reachable --changed-paths' from\n    3m00s to 1m37s.\n\nWe would probably want time for no sorting, for sorting by generation\nnumber v2, and for sorting by topological level (generation number v1)\nfor the same or similar case.\n\nBest,\n--\nJakub Narębski\n"},{"id":"402805","messageId":"85wo2eu3gc.fsf@gmail.com","threadId":"53933","inReplyTo":"20200804005658.GB75662@syl.lan","subject":"Re: [PATCH 1/6] commit-graph: fix regression when computing bloom filter","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-04T10:10:11Z","receivedAt":"2020-08-04T10:10:16Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n> On Tue, Aug 04, 2020 at 02:46:55AM +0200, Jakub Narębski wrote:\n>> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n[...]\n>>> diff --git a/commit-graph.c b/commit-graph.c\n>>> index 1af68c297d..5d3c9bd23c 100644\n>>> --- a/commit-graph.c\n>>> +++ b/commit-graph.c\n>>\n>> We might want to add function comment either here or in the header that\n>> this comparisonn function is to be used only for `git commit-graph\n>> write`, and not for graph traversal (even if similar funnction exists in\n>> other modules).\n>\n> I think that probably within the function is just fine, and that we can\n> avoid touching commit-graph.h here.\n>\n>>\n>>> @@ -144,8 +144,9 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>>>  \tconst struct commit *a = *(const struct commit **)va;\n>>>  \tconst struct commit *b = *(const struct commit **)vb;\n>\n> Maybe something like:\n>\n>   /*\n>    * Access the generation number directly with\n>    * 'commit_graph_data_at(...)->generation' instead of going through\n>    * the slab as usual to avoid accessing a yet-uncomputed value.\n>    */\n\nI think the last part of this comment should read:\n\n[...]\n     * 'commit_graph_data_at(...)->generation' instead of going through\n     * the commit_graph_generation() helper function to access just\n     * computed data [during `git commit-graph write --reachable --changed-paths`].\n     */\n\nOr something like that (the part in square brackets is optional; I am\nnot sure if adding it helps or not).\n\n>\n> Folks that are curious for more can blame this commit and read there.\n> I'd err on the side of being brief in the code comment and verbose in\n> the commit message than the other way around ;).\n\nI agree.\n\n>>>\n>>> -\tuint32_t generation_a = commit_graph_generation(a);\n>>> -\tuint32_t generation_b = commit_graph_generation(b);\n>>> +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n>>> +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n>>> +\n>>>  \t/* lower generation commits first */\n>>>  \tif (generation_a < generation_b)\n>>>  \t\treturn -1;\n\nBest,\n-- \nJakub Narębski\n"},{"id":"403007","messageId":"85pn84u1j9.fsf@gmail.com","threadId":"53933","inReplyTo":"d23f67dc80b85abe4eba9a9dfc39d50188e23bb7.1595927632.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/6] revision: parse parent in indegree_walk_step()","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-05T23:16:10Z","receivedAt":"2020-08-05T23:16:24Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> In indegree_walk_step(), we add unvisited parents to the indegree queue.\n> However, parents are not guaranteed to be parsed. As the indegree queue\n> sorts by generation number, let's parse parents before inserting them to\n> ensure the correct priority order.\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  revision.c | 3 +++\n>  1 file changed, 3 insertions(+)\n>\n> diff --git a/revision.c b/revision.c\n> index 6aa7f4f567..23287d26c3 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -3343,6 +3343,9 @@ static void indegree_walk_step(struct rev_info *revs)\n>  \t\tstruct commit *parent = p->item;\n>  \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n>\n> +\t\tif (parse_commit_gently(parent, 1) < 0)\n\nAll right, parse_commit_gently() avoids re-parsing objects, and makes\nuse of the commit-graph data.  If parents are not guaranteed to be\nparsed, this is a correct thing to do.\n\nThough I do wonder how this issue got missed by the test suite, just\nlike other reviewers...\n\n> +\t\t\treturn ;\n\nWhy this need to be 'return' and not 'continue'?\n\n> +\n>  \t\tif (*pi)\n>  \t\t\t(*pi)++;\n>  \t\telse\n\nBest,\n-- \nJakub Narębski\n"},{"id":"403217","messageId":"a962b9ae4b7a95a8263010ba76519f2bc0a73888.1596941624.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"[PATCH v2 01/10] commit-graph: fix regression when computing bloom filter","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-09T02:53:35Z","receivedAt":"2020-08-09T02:53:53Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\ncommit_gen_cmp is used when writing a commit-graph to sort commits in\ngeneration order before computing Bloom filters. Since c49c82aa (commit:\nmove members graph_pos, generation to a slab, 2020-06-17) made it so\nthat 'commit_graph_generation()' returns 'GENERATION_NUMBER_INFINITY'\nduring writing, we cannot call it within this function. Instead, access\nthe generation number directly through the slab (i.e., by calling\n'commit_graph_data_at(c)->generation') in order to access it while\nwriting.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex e51c91dd5b..ace7400a1a 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n+\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n+\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"403218","messageId":"32da955e318c143db9605a1b2b312598b3fc5231.1596941624.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"[PATCH v2 03/10] commit-graph: consolidate fill_commit_graph_info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-09T02:53:37Z","receivedAt":"2020-08-09T02:54:01Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nBoth fill_commit_graph_info() and fill_commit_in_graph() parse\ninformation present in commit data chunk. Let's simplify the\nimplementation by calling fill_commit_graph_info() within\nfill_commit_in_graph().\n\nThe test 'generate tar with future mtime' creates a commit with commit\ntime of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\ngeneration number (within CDAT chunk) and has undefined behavior.\n\nThe test used to pass as fill_commit_in_graph() guarantees the values of\ngraph position and generation number, and did not load timestamp.\nHowever, with corrected commit date we will need load the timestamp as\nwell to populate the generation number.\n\nLet's fix the test by setting a timestamp of (2 ^ 34 - 1) seconds.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c      | 29 +++++++++++------------------\n t/t5000-tar-tree.sh |  4 ++--\n 2 files changed, 13 insertions(+), 20 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex ace7400a1a..af8d9cc45e 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -725,15 +725,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n \tuint32_t lex_index;\n+\tuint64_t date_high, date_low;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n+\tif (pos >= g->num_commits + g->num_commits_in_base)\n+\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\n \tlex_index = pos - g->num_commits_in_base;\n \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n \n \tgraph_data = commit_graph_data_at(item);\n \tgraph_data->graph_pos = pos;\n+\n+\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n+\tdate_low = get_be32(commit_data + g->hash_len + 12);\n+\titem->date = (timestamp_t)((date_high << 32) | date_low);\n+\n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n@@ -748,38 +757,22 @@ static int fill_commit_in_graph(struct repository *r,\n {\n \tuint32_t edge_value;\n \tuint32_t *parent_data_ptr;\n-\tuint64_t date_low, date_high;\n \tstruct commit_list **pptr;\n-\tstruct commit_graph_data *graph_data;\n \tconst unsigned char *commit_data;\n \tuint32_t lex_index;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n-\tif (pos >= g->num_commits + g->num_commits_in_base)\n-\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\tfill_commit_graph_info(item, g, pos);\n \n-\t/*\n-\t * Store the \"full\" position, but then use the\n-\t * \"local\" position for the rest of the calculation.\n-\t */\n-\tgraph_data = commit_graph_data_at(item);\n-\tgraph_data->graph_pos = pos;\n \tlex_index = pos - g->num_commits_in_base;\n-\n-\tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n+\tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n \n \titem->object.parsed = 1;\n \n \tset_commit_tree(item, NULL);\n \n-\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n-\tdate_low = get_be32(commit_data + g->hash_len + 12);\n-\titem->date = (timestamp_t)((date_high << 32) | date_low);\n-\n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n-\n \tpptr = &item->parents;\n \n \tedge_value = get_be32(commit_data + g->hash_len);\ndiff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\nindex 37655a237c..1986354fc3 100755\n--- a/t/t5000-tar-tree.sh\n+++ b/t/t5000-tar-tree.sh\n@@ -406,7 +406,7 @@ test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n \trm -f .git/index &&\n \techo content >file &&\n \tgit add file &&\n-\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n+\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n \t\tgit commit -m \"tempori parendum\"\n '\n \n@@ -415,7 +415,7 @@ test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n '\n \n test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n-\techo 4147 >expect &&\n+\techo 2514 >expect &&\n \ttar_info future.tar | cut -d\" \" -f2 >actual &&\n \ttest_cmp expect actual\n '\n-- \ngitgitgadget\n\n"},{"id":"403219","messageId":"cf61239f9332a9cd91110a9b6319b1133ff3bd27.1596941624.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"[PATCH v2 02/10] revision: parse parent in indegree_walk_step()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-09T02:53:36Z","receivedAt":"2020-08-09T02:54:01Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn indegree_walk_step(), we add unvisited parents to the indegree queue.\nHowever, parents are not guaranteed to be parsed. As the indegree queue\nsorts by generation number, let's parse parents before inserting them to\nensure the correct priority order.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n revision.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/revision.c b/revision.c\nindex 6de29cdf7a..4ec82ed5ab 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3365,6 +3365,9 @@ static void indegree_walk_step(struct rev_info *revs)\n \t\tstruct commit *parent = p->item;\n \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n \n+\t\tif (parse_commit_gently(parent, 1) < 0)\n+\t\t\treturn;\n+\n \t\tif (*pi)\n \t\t\t(*pi)++;\n \t\telse\n-- \ngitgitgadget\n\n"},{"id":"403220","messageId":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.git.1595927632.gitgitgadget@gmail.com","subject":"[PATCH v2 00/10] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-09T02:53:34Z","receivedAt":"2020-08-09T02:54:01Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"This patch series implements the corrected commit date offsets as generation\nnumber v2, along with other pre-requisites.\n\nGit uses topological levels in the commit-graph file for commit-graph\ntraversal operations like git log --graph. Unfortunately, using topological\nlevels can result in a worse performance than without them when compared\nwith committer date as a heuristics. For example, git merge-base v4.8 v4.9 \non the Linux repository walks 635,579 commits using topological levels and\nwalks 167,468 using committer date.\n\nThus, the need for generation number v2 was born. New generation number\nneeded to provide good performance, increment updates, and backward\ncompatibility. Due to an unfortunate problem, we also needed a way to\ndistinguish between the old and new generation number without incrementing\ngraph version.\n\nVarious candidates were examined (https://github.com/derrickstolee/gen-test, \nhttps://github.com/abhishekkumar2718/git/pull/1). The proposed generation\nnumber v2, Corrected Commit Date with Mononotically Increasing Offsets \nperformed much worse than committer date (506,577 vs. 167,468 commits walked\nfor git merge-base v4.8 v4.9) and was dropped.\n\nUsing Generation Data chunk (GDAT) relieves the requirement of backward\ncompatibility as we would continue to store topological levels in Commit\nData (CDAT) chunk. Thus, Corrected Commit Date was chosen as generation\nnumber v2. The Corrected Commit Date is defined as:\n\nFor a commit C, let its corrected commit date be the maximum of the commit\ndate of C and the corrected commit dates of its parents. Then corrected\ncommit date offset is the difference between corrected commit date of C and\ncommit date of C.\n\nWe will introduce an additional commit-graph chunk, Generation Data chunk,\nand store corrected commit date offsets in GDAT chunk while storing\ntopological levels in CDAT chunk. The old versions of Git would ignore GDAT\nchunk, using topological levels from CDAT chunk. In contrast, new versions\nof Git would use corrected commit dates, falling back to topological level\nif the generation data chunk is absent in the commit-graph file.\n\nThanks to Dr. Stolee, Dr. Narębski, and Taylor for their reviews on the\nfirst version.\n\nI look forward to everyone's reviews!\n\nThanks\n\n * Abhishek\n\n\n----------------------------------------------------------------------------\n\nChanges in version 2:\n\n * Add tests for generation data chunk.\n * Add an option GIT_TEST_COMMIT_GRAPH_NO_GDAT to control whether to write\n   generation data chunk.\n * Compare commits with corrected commit dates if present in\n   paint_down_to_common().\n * Update technical documentation.\n * Handle mixed graph version commit chains.\n * Improve commit messages for\n * Revert unnecessary whitespace changes.\n * Split uint_32 -> timestamp_t change into a new commit.\n\nAbhishek Kumar (10):\n  commit-graph: fix regression when computing bloom filter\n  revision: parse parent in indegree_walk_step()\n  commit-graph: consolidate fill_commit_graph_info\n  commit-graph: consolidate compare_commits_by_gen\n  commit-graph: implement generation data chunk\n  commit-graph: return 64-bit generation number\n  commit-graph: implement corrected commit date\n  commit-graph: handle mixed generation commit chains\n  commit-reach: use corrected commit dates in paint_down_to_common()\n  doc: add corrected commit date info\n\n .../technical/commit-graph-format.txt         |  12 +-\n Documentation/technical/commit-graph.txt      |  45 ++--\n commit-graph.c                                | 203 ++++++++++++------\n commit-graph.h                                |  14 +-\n commit-reach.c                                |  49 ++---\n commit-reach.h                                |   2 +-\n commit.c                                      |   9 +-\n commit.h                                      |   4 +-\n revision.c                                    |  13 +-\n t/README                                      |   3 +\n t/helper/test-read-graph.c                    |   2 +\n t/t4216-log-bloom.sh                          |   4 +-\n t/t5000-tar-tree.sh                           |   4 +-\n t/t5318-commit-graph.sh                       |  27 +--\n t/t5324-split-commit-graph.sh                 |  78 ++++++-\n t/t6024-recursive-merge.sh                    |   4 +-\n t/t6600-test-reach.sh                         |  62 +++---\n upload-pack.c                                 |   2 +-\n 18 files changed, 354 insertions(+), 183 deletions(-)\n\n\nbase-commit: dc04167d378fb29d30e1647ff6ff51dd182bc9a3\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-676%2Fabhishekkumar2718%2Fcorrected_commit_date-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-676/abhishekkumar2718/corrected_commit_date-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/676\n\nRange-diff vs v1:\n\n  1:  91e6e97a66 !  1:  a962b9ae4b commit-graph: fix regression when computing bloom filter\n     @@ Metadata\n       ## Commit message ##\n          commit-graph: fix regression when computing bloom filter\n      \n     -    With 3d112755 (commit-graph: examine commits by generation number), Git\n     -    knew to sort by generation number before examining the diff when not\n     -    using pack order. c49c82aa (commit: move members graph_pos, generation\n     -    to a slab, 2020-06-17) moved generation number into a slab and\n     -    introduced a helper which returns GENERATION_NUMBER_INFINITY when\n     -    writing the graph. Sorting is no longer useful and essentially reverts\n     -    the earlier commit.\n     -\n     -    Let's fix this by accessing generation number directly through the slab.\n     +    commit_gen_cmp is used when writing a commit-graph to sort commits in\n     +    generation order before computing Bloom filters. Since c49c82aa (commit:\n     +    move members graph_pos, generation to a slab, 2020-06-17) made it so\n     +    that 'commit_graph_generation()' returns 'GENERATION_NUMBER_INFINITY'\n     +    during writing, we cannot call it within this function. Instead, access\n     +    the generation number directly through the slab (i.e., by calling\n     +    'commit_graph_data_at(c)->generation') in order to access it while\n     +    writing.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n     @@ commit-graph.c: static int commit_gen_cmp(const void *va, const void *vb)\n      -\tuint32_t generation_b = commit_graph_generation(b);\n      +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n      +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n     -+\n       \t/* lower generation commits first */\n       \tif (generation_a < generation_b)\n       \t\treturn -1;\n  2:  d23f67dc80 !  2:  cf61239f93 revision: parse parent in indegree_walk_step()\n     @@ revision.c: static void indegree_walk_step(struct rev_info *revs)\n       \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n       \n      +\t\tif (parse_commit_gently(parent, 1) < 0)\n     -+\t\t\treturn ;\n     ++\t\t\treturn;\n      +\n       \t\tif (*pi)\n       \t\t\t(*pi)++;\n  3:  701f591236 !  3:  32da955e31 commit-graph: consolidate fill_commit_graph_info\n     @@ Commit message\n      \n          The test 'generate tar with future mtime' creates a commit with commit\n          time of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\n     -    generation number and has undefined behavior. The test used to pass as\n     -    fill_commit_in_graph() did not read commit time from commit graph,\n     -    reading commit date from odb instead.\n     +    generation number (within CDAT chunk) and has undefined behavior.\n      \n     -    Let's fix that by setting commit time of (2 ^ 34 - 1) seconds.\n     +    The test used to pass as fill_commit_in_graph() guarantees the values of\n     +    graph position and generation number, and did not load timestamp.\n     +    However, with corrected commit date we will need load the timestamp as\n     +    well to populate the generation number.\n     +\n     +    Let's fix the test by setting a timestamp of (2 ^ 34 - 1) seconds.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n     @@ commit-graph.c: static int fill_commit_in_graph(struct repository *r,\n       \tconst unsigned char *commit_data;\n       \tuint32_t lex_index;\n       \n     -+\tfill_commit_graph_info(item, g, pos);\n     -+\n       \twhile (pos < g->num_commits_in_base)\n       \t\tg = g->base_graph;\n       \n      -\tif (pos >= g->num_commits + g->num_commits_in_base)\n      -\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n     --\n     ++\tfill_commit_graph_info(item, g, pos);\n     + \n      -\t/*\n      -\t * Store the \"full\" position, but then use the\n      -\t * \"local\" position for the rest of the calculation.\n  4:  812fe75fc7 !  4:  b254782858 commit-graph: consolidate compare_commits_by_gen\n     @@ Commit message\n          compare_commits_by_gen() to commit-graph.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n     +    Reviewed-by: Taylor Blau <me@ttaylorr.com>\n      \n       ## commit-graph.c ##\n      @@ commit-graph.c: uint32_t commit_graph_generation(const struct commit *c)\n  5:  80ea7da343 !  5:  cb797e20d7 commit-graph: implement generation data chunk\n     @@ Metadata\n       ## Commit message ##\n          commit-graph: implement generation data chunk\n      \n     -    One of the essential pre-requisites before implementing generation\n     -    number as to distinguish between generation numbers v1 and v2 while\n     -    still being compatible with old Git.\n     +    As discovered by Ævar, we cannot increment graph version to\n     +    distinguish between generation numbers v1 and v2 [1]. Thus, one of\n     +    pre-requistes before implementing generation number was to distinguish\n     +    between graph versions in a backwards compatible manner.\n      \n          We are going to introduce a new chunk called Generation Data chunk (or\n          GDAT). GDAT stores generation number v2 (and any subsequent versions),\n          whereas CDAT will still store topological level.\n      \n          Old Git does not understand GDAT chunk and would ignore it, reading\n     -    topological levels from CDAT. Newer versions of Git can parse GDAT and\n     -    take advantage of newer generation numbers, falling back to topological\n     -    levels when GDAT chunk is missing (as it would happen with a commit\n     -    graph written by old Git).\n     +    topological levels from CDAT. New Git can parse GDAT and take advantage\n     +    of newer generation numbers, falling back to topological levels when\n     +    GDAT chunk is missing (as it would happen with a commit graph written\n     +    by old Git).\n     +\n     +    We introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\n     +    which forces commit-graph file to be written without generation data\n     +    chunk to emulate a commit-graph file written by old Git.\n     +\n     +    [1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n     @@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct c\n       }\n       \n       static inline void set_commit_tree(struct commit *c, struct tree *t)\n     -@@ commit-graph.c: static void write_graph_chunk_data(struct hashfile *f, int hash_len,\n     - \t}\n     +@@ commit-graph.c: struct write_commit_graph_context {\n     + \t\t report_progress:1,\n     + \t\t split:1,\n     + \t\t changed_paths:1,\n     +-\t\t order_by_pack:1;\n     ++\t\t order_by_pack:1,\n     ++\t\t write_generation_data:1;\n     + \n     + \tconst struct split_commit_graph_opts *split_opts;\n     + \tsize_t total_bloom_filter_data_size;\n     +@@ commit-graph.c: static int write_graph_chunk_data(struct hashfile *f,\n     + \treturn 0;\n       }\n       \n     -+static void write_graph_chunk_generation_data(struct hashfile *f,\n     ++static int write_graph_chunk_generation_data(struct hashfile *f,\n      +\t\t\t\t\t      struct write_commit_graph_context *ctx)\n      +{\n     -+\tstruct commit **list = ctx->commits.list;\n     -+\tint count;\n     -+\tfor (count = 0; count < ctx->commits.nr; count++, list++) {\n     ++\tint i;\n     ++\tfor (i = 0; i < ctx->commits.nr; i++) {\n     ++\t\tstruct commit *c = ctx->commits.list[i];\n      +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n     -+\t\thashwrite_be32(f, commit_graph_data_at(*list)->generation);\n     ++\t\thashwrite_be32(f, commit_graph_data_at(c)->generation);\n      +\t}\n     ++\n     ++\treturn 0;\n      +}\n      +\n     - static void write_graph_chunk_extra_edges(struct hashfile *f,\n     - \t\t\t\t\t  struct write_commit_graph_context *ctx)\n     + static int write_graph_chunk_extra_edges(struct hashfile *f,\n     +-\t\t\t\t\t struct write_commit_graph_context *ctx)\n     ++\t\t\t\t\t  struct write_commit_graph_context *ctx)\n       {\n     + \tstruct commit **list = ctx->commits.list;\n     + \tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n      @@ commit-graph.c: static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n     - \tuint64_t chunk_offsets[MAX_NUM_CHUNKS + 1];\n     - \tconst unsigned hashsz = the_hash_algo->rawsz;\n     - \tstruct strbuf progress_title = STRBUF_INIT;\n     --\tint num_chunks = 3;\n     -+\tint num_chunks = 4;\n     - \tstruct object_id file_hash;\n     - \tconst struct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n     - \n     -@@ commit-graph.c: static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n     - \tchunk_ids[0] = GRAPH_CHUNKID_OIDFANOUT;\n     - \tchunk_ids[1] = GRAPH_CHUNKID_OIDLOOKUP;\n     - \tchunk_ids[2] = GRAPH_CHUNKID_DATA;\n     -+\tchunk_ids[3] = GRAPH_CHUNKID_GENERATION_DATA;\n     + \tchunks[2].id = GRAPH_CHUNKID_DATA;\n     + \tchunks[2].size = (hashsz + 16) * ctx->commits.nr;\n     + \tchunks[2].write_fn = write_graph_chunk_data;\n     ++\tif (ctx->write_generation_data) {\n     ++\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA;\n     ++\t\tchunks[num_chunks].size = sizeof(uint32_t) * ctx->commits.nr;\n     ++\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data;\n     ++\t\tnum_chunks++;\n     ++\t}\n       \tif (ctx->num_extra_edges) {\n     - \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_EXTRAEDGES;\n     - \t\tnum_chunks++;\n     -@@ commit-graph.c: static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n     - \tchunk_offsets[1] = chunk_offsets[0] + GRAPH_FANOUT_SIZE;\n     - \tchunk_offsets[2] = chunk_offsets[1] + hashsz * ctx->commits.nr;\n     - \tchunk_offsets[3] = chunk_offsets[2] + (hashsz + 16) * ctx->commits.nr;\n     -+\tchunk_offsets[4] = chunk_offsets[3] + sizeof(uint32_t) * ctx->commits.nr;\n     + \t\tchunks[num_chunks].id = GRAPH_CHUNKID_EXTRAEDGES;\n     + \t\tchunks[num_chunks].size = 4 * ctx->num_extra_edges;\n     +@@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n     + \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n     + \tctx->split_opts = split_opts;\n     + \tctx->total_bloom_filter_data_size = 0;\n     ++\tctx->write_generation_data = !git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0);\n       \n     --\tnum_chunks = 3;\n     -+\tnum_chunks = 4;\n     - \tif (ctx->num_extra_edges) {\n     - \t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n     - \t\t\t\t\t\t4 * ctx->num_extra_edges;\n     -@@ commit-graph.c: static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n     - \twrite_graph_chunk_fanout(f, ctx);\n     - \twrite_graph_chunk_oids(f, hashsz, ctx);\n     - \twrite_graph_chunk_data(f, hashsz, ctx);\n     -+\twrite_graph_chunk_generation_data(f, ctx);\n     - \tif (ctx->num_extra_edges)\n     - \t\twrite_graph_chunk_extra_edges(f, ctx);\n     - \tif (ctx->changed_paths) {\n     + \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n     + \t\tctx->changed_paths = 1;\n      \n       ## commit-graph.h ##\n     +@@\n     + #include \"oidset.h\"\n     + \n     + #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n     ++#define GIT_TEST_COMMIT_GRAPH_NO_GDAT \"GIT_TEST_COMMIT_GRAPH_NO_GDAT\"\n     + #define GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE \"GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE\"\n     + #define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n     + \n      @@ commit-graph.h: struct commit_graph {\n       \tconst uint32_t *chunk_oid_fanout;\n       \tconst unsigned char *chunk_oid_lookup;\n     @@ commit-graph.h: struct commit_graph {\n       \tconst unsigned char *chunk_base_graphs;\n       \tconst unsigned char *chunk_bloom_indexes;\n      \n     + ## t/README ##\n     +@@ t/README: GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n     + be written after every 'git commit' command, and overrides the\n     + 'core.commitGraph' setting to true.\n     + \n     ++GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n     ++commit-graph to be written without generation data chunk.\n     ++\n     + GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n     + commit-graph write to compute and write changed path Bloom filters for\n     + every 'git commit-graph write', as if the `--changed-paths` option was\n     +\n       ## t/helper/test-read-graph.c ##\n      @@ t/helper/test-read-graph.c: int cmd__read_graph(int argc, const char **argv)\n       \t\tprintf(\" oid_lookup\");\n     @@ t/t4216-log-bloom.sh: test_expect_success 'setup test - repo, commits, commit gr\n      \n       ## t/t5318-commit-graph.sh ##\n      @@ t/t5318-commit-graph.sh: graph_git_behavior 'no graph' full commits/3 commits/1\n     - \n       graph_read_expect() {\n       \tOPTIONAL=\"\"\n     --\tNUM_CHUNKS=3\n     -+\tNUM_CHUNKS=4\n     - \tif test ! -z $2\n     + \tNUM_CHUNKS=3\n     +-\tif test ! -z $2\n     ++\tif test ! -z \"$2\"\n       \tthen\n       \t\tOPTIONAL=\" $2\"\n     --\t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n     -+\t\tNUM_CHUNKS=$((4 + $(echo \"$2\" | wc -w)))\n     - \tfi\n     - \tcat >expect <<- EOF\n     - \theader: 43475048 1 1 $NUM_CHUNKS 0\n     - \tnum_commits: $1\n     --\tchunks: oid_fanout oid_lookup commit_metadata$OPTIONAL\n     -+\tchunks: oid_fanout oid_lookup commit_metadata generation_data$OPTIONAL\n     - \tEOF\n     - \ttest-tool read-graph >output &&\n     - \ttest_cmp expect output\n     -@@ t/t5318-commit-graph.sh: GRAPH_BYTE_HASH=5\n     - GRAPH_BYTE_CHUNK_COUNT=6\n     - GRAPH_CHUNK_LOOKUP_OFFSET=8\n     - GRAPH_CHUNK_LOOKUP_WIDTH=12\n     --GRAPH_CHUNK_LOOKUP_ROWS=5\n     -+GRAPH_CHUNK_LOOKUP_ROWS=6\n     - GRAPH_BYTE_OID_FANOUT_ID=$GRAPH_CHUNK_LOOKUP_OFFSET\n     - GRAPH_BYTE_OID_LOOKUP_ID=$(($GRAPH_CHUNK_LOOKUP_OFFSET + \\\n     - \t\t\t    1 * $GRAPH_CHUNK_LOOKUP_WIDTH))\n     -@@ t/t5318-commit-graph.sh: GRAPH_BYTE_COMMIT_TREE=$GRAPH_COMMIT_DATA_OFFSET\n     - GRAPH_BYTE_COMMIT_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN))\n     - GRAPH_BYTE_COMMIT_EXTRA_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 4))\n     - GRAPH_BYTE_COMMIT_WRONG_PARENT=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 3))\n     --GRAPH_BYTE_COMMIT_GENERATION=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 11))\n     - GRAPH_BYTE_COMMIT_DATE=$(($GRAPH_COMMIT_DATA_OFFSET + $HASH_LEN + 12))\n     - GRAPH_COMMIT_DATA_WIDTH=$(($HASH_LEN + 16))\n     --GRAPH_OCTOPUS_DATA_OFFSET=$(($GRAPH_COMMIT_DATA_OFFSET + \\\n     --\t\t\t     $GRAPH_COMMIT_DATA_WIDTH * $NUM_COMMITS))\n     -+GRAPH_GENERATION_DATA_OFFSET=$(($GRAPH_COMMIT_DATA_OFFSET + \\\n     -+\t\t\t\t$GRAPH_COMMIT_DATA_WIDTH * $NUM_COMMITS))\n     -+GRAPH_GENERATION_DATA_WIDTH=4\n     -+GRAPH_BYTE_COMMIT_GENERATION=$(($GRAPH_GENERATION_DATA_OFFSET + 3))\n     -+GRAPH_OCTOPUS_DATA_OFFSET=$(($GRAPH_GENERATION_DATA_OFFSET + \\\n     -+\t\t\t     $GRAPH_GENERATION_DATA_WIDTH * $NUM_COMMITS))\n     - GRAPH_BYTE_OCTOPUS=$(($GRAPH_OCTOPUS_DATA_OFFSET + 4))\n     - GRAPH_BYTE_FOOTER=$(($GRAPH_OCTOPUS_DATA_OFFSET + 4 * $NUM_OCTOPUS_EDGES))\n     - \n     -@@ t/t5318-commit-graph.sh: test_expect_success 'detect incorrect generation number' '\n     - '\n     - \n     - test_expect_success 'detect incorrect generation number' '\n     --\tcorrupt_graph_and_verify $GRAPH_BYTE_COMMIT_GENERATION \"\\01\" \\\n     -+\tcorrupt_graph_and_verify $GRAPH_BYTE_COMMIT_GENERATION \"\\00\" \\\n     - \t\t\"non-zero generation number\"\n     + \t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'exit with correct error on bad input to --stdin-commits' '\n     + \t# valid commit and tree OID\n     + \tgit rev-parse HEAD HEAD^{tree} >in &&\n     + \tgit commit-graph write --stdin-commits <in &&\n     +-\tgraph_read_expect 3\n     ++\tgraph_read_expect 3 generation_data\n     + '\n     + \n     + test_expect_success 'write graph' '\n     + \tcd \"$TRASH_DIRECTORY/full\" &&\n     + \tgit commit-graph write &&\n     + \ttest_path_is_file $objdir/info/commit-graph &&\n     +-\tgraph_read_expect \"3\"\n     ++\tgraph_read_expect \"3\" generation_data\n       '\n       \n     + test_expect_success POSIXPERM 'write graph has correct permissions' '\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'write graph with merges' '\n     + \tcd \"$TRASH_DIRECTORY/full\" &&\n     + \tgit commit-graph write &&\n     + \ttest_path_is_file $objdir/info/commit-graph &&\n     +-\tgraph_read_expect \"10\" \"extra_edges\"\n     ++\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n     + '\n     + \n     + graph_git_behavior 'merge 1 vs 2' full merge/1 merge/2\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'write graph with new commit' '\n     + \tcd \"$TRASH_DIRECTORY/full\" &&\n     + \tgit commit-graph write &&\n     + \ttest_path_is_file $objdir/info/commit-graph &&\n     +-\tgraph_read_expect \"11\" \"extra_edges\"\n     ++\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n     + '\n     + \n     + graph_git_behavior 'full graph, commit 8 vs merge 1' full commits/8 merge/1\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'write graph with nothing new' '\n     + \tcd \"$TRASH_DIRECTORY/full\" &&\n     + \tgit commit-graph write &&\n     + \ttest_path_is_file $objdir/info/commit-graph &&\n     +-\tgraph_read_expect \"11\" \"extra_edges\"\n     ++\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n     + '\n     + \n     + graph_git_behavior 'cleared graph, commit 8 vs merge 1' full commits/8 merge/1\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'build graph from latest pack with closure' '\n     + \tcd \"$TRASH_DIRECTORY/full\" &&\n     + \tcat new-idx | git commit-graph write --stdin-packs &&\n     + \ttest_path_is_file $objdir/info/commit-graph &&\n     +-\tgraph_read_expect \"9\" \"extra_edges\"\n     ++\tgraph_read_expect \"9\" \"generation_data extra_edges\"\n     + '\n     + \n     + graph_git_behavior 'graph from pack, commit 8 vs merge 1' full commits/8 merge/1\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'build graph from commits with closure' '\n     + \tgit rev-parse merge/1 >>commits-in &&\n     + \tcat commits-in | git commit-graph write --stdin-commits &&\n     + \ttest_path_is_file $objdir/info/commit-graph &&\n     +-\tgraph_read_expect \"6\"\n     ++\tgraph_read_expect \"6\" \"generation_data\"\n     + '\n     + \n     + graph_git_behavior 'graph from commits, commit 8 vs merge 1' full commits/8 merge/1\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'build graph from commits with append' '\n     + \tcd \"$TRASH_DIRECTORY/full\" &&\n     + \tgit rev-parse merge/3 | git commit-graph write --stdin-commits --append &&\n     + \ttest_path_is_file $objdir/info/commit-graph &&\n     +-\tgraph_read_expect \"10\" \"extra_edges\"\n     ++\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n     + '\n     + \n     + graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'build graph using --reachable' '\n     + \tcd \"$TRASH_DIRECTORY/full\" &&\n     + \tgit commit-graph write --reachable &&\n     + \ttest_path_is_file $objdir/info/commit-graph &&\n     +-\tgraph_read_expect \"11\" \"extra_edges\"\n     ++\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n     + '\n     + \n     + graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'write graph in bare repo' '\n     + \tcd \"$TRASH_DIRECTORY/bare\" &&\n     + \tgit commit-graph write &&\n     + \ttest_path_is_file $baredir/info/commit-graph &&\n     +-\tgraph_read_expect \"11\" \"extra_edges\"\n     ++\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n     + '\n     + \n     + graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'replace-objects invalidates commit-graph' '\n     + \n     + test_expect_success 'git commit-graph verify' '\n     + \tcd \"$TRASH_DIRECTORY/full\" &&\n     +-\tgit rev-parse commits/8 | git commit-graph write --stdin-commits &&\n     +-\tgit commit-graph verify >output\n     ++\tgit rev-parse commits/8 | GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --stdin-commits &&\n     ++\tgit commit-graph verify >output &&\n     ++\tgraph_read_expect 9 extra_edges\n     + '\n     + \n     + NUM_COMMITS=9\n      \n       ## t/t5324-split-commit-graph.sh ##\n      @@ t/t5324-split-commit-graph.sh: test_expect_success 'setup repo' '\n     @@ t/t5324-split-commit-graph.sh: graph_read_expect() {\n       \tEOF\n       \ttest-tool read-graph >output &&\n       \ttest_cmp expect output\n     +\n     + ## t/t6600-test-reach.sh ##\n     +@@ t/t6600-test-reach.sh: test_expect_success 'setup' '\n     + \tgit show-ref -s commit-5-5 | git commit-graph write --stdin-commits &&\n     + \tmv .git/objects/info/commit-graph commit-graph-half &&\n     + \tchmod u+w commit-graph-half &&\n     ++\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable &&\n     ++\tmv .git/objects/info/commit-graph commit-graph-no-gdat &&\n     ++\tchmod u+w commit-graph-no-gdat &&\n     + \tgit config core.commitGraph true\n     + '\n     + \n     +-run_three_modes () {\n     ++run_all_modes () {\n     + \ttest_when_finished rm -rf .git/objects/info/commit-graph &&\n     + \t\"$@\" <input >actual &&\n     + \ttest_cmp expect actual &&\n     +@@ t/t6600-test-reach.sh: run_three_modes () {\n     + \ttest_cmp expect actual &&\n     + \tcp commit-graph-half .git/objects/info/commit-graph &&\n     + \t\"$@\" <input >actual &&\n     ++\ttest_cmp expect actual &&\n     ++\tcp commit-graph-no-gdat .git/objects/info/commit-graph &&\n     ++\t\"$@\" <input >actual &&\n     + \ttest_cmp expect actual\n     + }\n     + \n     +-test_three_modes () {\n     +-\trun_three_modes test-tool reach \"$@\"\n     ++test_all_modes () {\n     ++\trun_all_modes test-tool reach \"$@\"\n     + }\n     + \n     + test_expect_success 'ref_newer:miss' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'ref_newer:miss' '\n     + \tB:commit-4-9\n     + \tEOF\n     + \techo \"ref_newer(A,B):0\" >expect &&\n     +-\ttest_three_modes ref_newer\n     ++\ttest_all_modes ref_newer\n     + '\n     + \n     + test_expect_success 'ref_newer:hit' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'ref_newer:hit' '\n     + \tB:commit-2-3\n     + \tEOF\n     + \techo \"ref_newer(A,B):1\" >expect &&\n     +-\ttest_three_modes ref_newer\n     ++\ttest_all_modes ref_newer\n     + '\n     + \n     + test_expect_success 'in_merge_bases:hit' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'in_merge_bases:hit' '\n     + \tB:commit-8-8\n     + \tEOF\n     + \techo \"in_merge_bases(A,B):1\" >expect &&\n     +-\ttest_three_modes in_merge_bases\n     ++\ttest_all_modes in_merge_bases\n     + '\n     + \n     + test_expect_success 'in_merge_bases:miss' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'in_merge_bases:miss' '\n     + \tB:commit-5-9\n     + \tEOF\n     + \techo \"in_merge_bases(A,B):0\" >expect &&\n     +-\ttest_three_modes in_merge_bases\n     ++\ttest_all_modes in_merge_bases\n     + '\n     + \n     + test_expect_success 'is_descendant_of:hit' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'is_descendant_of:hit' '\n     + \tX:commit-1-1\n     + \tEOF\n     + \techo \"is_descendant_of(A,X):1\" >expect &&\n     +-\ttest_three_modes is_descendant_of\n     ++\ttest_all_modes is_descendant_of\n     + '\n     + \n     + test_expect_success 'is_descendant_of:miss' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'is_descendant_of:miss' '\n     + \tX:commit-7-6\n     + \tEOF\n     + \techo \"is_descendant_of(A,X):0\" >expect &&\n     +-\ttest_three_modes is_descendant_of\n     ++\ttest_all_modes is_descendant_of\n     + '\n     + \n     + test_expect_success 'get_merge_bases_many' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'get_merge_bases_many' '\n     + \t\tgit rev-parse commit-5-6 \\\n     + \t\t\t      commit-4-7 | sort\n     + \t} >expect &&\n     +-\ttest_three_modes get_merge_bases_many\n     ++\ttest_all_modes get_merge_bases_many\n     + '\n     + \n     + test_expect_success 'reduce_heads' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'reduce_heads' '\n     + \t\t\t      commit-2-8 \\\n     + \t\t\t      commit-1-10 | sort\n     + \t} >expect &&\n     +-\ttest_three_modes reduce_heads\n     ++\ttest_all_modes reduce_heads\n     + '\n     + \n     + test_expect_success 'can_all_from_reach:hit' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'can_all_from_reach:hit' '\n     + \tY:commit-8-1\n     + \tEOF\n     + \techo \"can_all_from_reach(X,Y):1\" >expect &&\n     +-\ttest_three_modes can_all_from_reach\n     ++\ttest_all_modes can_all_from_reach\n     + '\n     + \n     + test_expect_success 'can_all_from_reach:miss' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'can_all_from_reach:miss' '\n     + \tY:commit-8-5\n     + \tEOF\n     + \techo \"can_all_from_reach(X,Y):0\" >expect &&\n     +-\ttest_three_modes can_all_from_reach\n     ++\ttest_all_modes can_all_from_reach\n     + '\n     + \n     + test_expect_success 'can_all_from_reach_with_flag: tags case' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'can_all_from_reach_with_flag: tags case' '\n     + \tY:commit-8-1\n     + \tEOF\n     + \techo \"can_all_from_reach_with_flag(X,_,_,0,0):1\" >expect &&\n     +-\ttest_three_modes can_all_from_reach_with_flag\n     ++\ttest_all_modes can_all_from_reach_with_flag\n     + '\n     + \n     + test_expect_success 'commit_contains:hit' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'commit_contains:hit' '\n     + \tX:commit-9-3\n     + \tEOF\n     + \techo \"commit_contains(_,A,X,_):1\" >expect &&\n     +-\ttest_three_modes commit_contains &&\n     +-\ttest_three_modes commit_contains --tag\n     ++\ttest_all_modes commit_contains &&\n     ++\ttest_all_modes commit_contains --tag\n     + '\n     + \n     + test_expect_success 'commit_contains:miss' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'commit_contains:miss' '\n     + \tX:commit-9-3\n     + \tEOF\n     + \techo \"commit_contains(_,A,X,_):0\" >expect &&\n     +-\ttest_three_modes commit_contains &&\n     +-\ttest_three_modes commit_contains --tag\n     ++\ttest_all_modes commit_contains &&\n     ++\ttest_all_modes commit_contains --tag\n     + '\n     + \n     + test_expect_success 'rev-list: basic topo-order' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: basic topo-order' '\n     + \t\tcommit-6-2 commit-5-2 commit-4-2 commit-3-2 commit-2-2 commit-1-2 \\\n     + \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n     + \t>expect &&\n     +-\trun_three_modes git rev-list --topo-order commit-6-6\n     ++\trun_all_modes git rev-list --topo-order commit-6-6\n     + '\n     + \n     + test_expect_success 'rev-list: first-parent topo-order' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: first-parent topo-order' '\n     + \t\tcommit-6-2 \\\n     + \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n     + \t>expect &&\n     +-\trun_three_modes git rev-list --first-parent --topo-order commit-6-6\n     ++\trun_all_modes git rev-list --first-parent --topo-order commit-6-6\n     + '\n     + \n     + test_expect_success 'rev-list: range topo-order' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: range topo-order' '\n     + \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n     + \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n     + \t>expect &&\n     +-\trun_three_modes git rev-list --topo-order commit-3-3..commit-6-6\n     ++\trun_all_modes git rev-list --topo-order commit-3-3..commit-6-6\n     + '\n     + \n     + test_expect_success 'rev-list: range topo-order' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: range topo-order' '\n     + \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n     + \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n     + \t>expect &&\n     +-\trun_three_modes git rev-list --topo-order commit-3-8..commit-6-6\n     ++\trun_all_modes git rev-list --topo-order commit-3-8..commit-6-6\n     + '\n     + \n     + test_expect_success 'rev-list: first-parent range topo-order' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: first-parent range topo-order' '\n     + \t\tcommit-6-2 \\\n     + \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n     + \t>expect &&\n     +-\trun_three_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n     ++\trun_all_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n     + '\n     + \n     + test_expect_success 'rev-list: ancestry-path topo-order' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: ancestry-path topo-order' '\n     + \t\tcommit-6-4 commit-5-4 commit-4-4 commit-3-4 \\\n     + \t\tcommit-6-3 commit-5-3 commit-4-3 \\\n     + \t>expect &&\n     +-\trun_three_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n     ++\trun_all_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n     + '\n     + \n     + test_expect_success 'rev-list: symmetric difference topo-order' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: symmetric difference topo-order' '\n     + \t\tcommit-3-8 commit-2-8 commit-1-8 \\\n     + \t\tcommit-3-7 commit-2-7 commit-1-7 \\\n     + \t>expect &&\n     +-\trun_three_modes git rev-list --topo-order commit-3-8...commit-6-6\n     ++\trun_all_modes git rev-list --topo-order commit-3-8...commit-6-6\n     + '\n     + \n     + test_expect_success 'get_reachable_subset:all' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'get_reachable_subset:all' '\n     + \t\t\t      commit-1-7 \\\n     + \t\t\t      commit-5-6 | sort\n     + \t) >expect &&\n     +-\ttest_three_modes get_reachable_subset\n     ++\ttest_all_modes get_reachable_subset\n     + '\n     + \n     + test_expect_success 'get_reachable_subset:some' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'get_reachable_subset:some' '\n     + \t\tgit rev-parse commit-3-3 \\\n     + \t\t\t      commit-1-7 | sort\n     + \t) >expect &&\n     +-\ttest_three_modes get_reachable_subset\n     ++\ttest_all_modes get_reachable_subset\n     + '\n     + \n     + test_expect_success 'get_reachable_subset:none' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'get_reachable_subset:none' '\n     + \tY:commit-2-8\n     + \tEOF\n     + \techo \"get_reachable_subset(X,Y)\" >expect &&\n     +-\ttest_three_modes get_reachable_subset\n     ++\ttest_all_modes get_reachable_subset\n     + '\n     + \n     + test_done\n  6:  647290d036 !  6:  1aa2a00a7a commit-graph: implement corrected commit date offset\n     @@ Metadata\n      Author: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n       ## Commit message ##\n     -    commit-graph: implement corrected commit date offset\n     +    commit-graph: return 64-bit generation number\n      \n     -    With preparations done, let's implement corrected commit date offset.\n     -    We add a new commit-slab to store topological levels while writing\n     -    commit graph and upgrade number of struct commit_graph_data to 64-bits.\n     -\n     -    We have to touch many files, upgrading generation number from uint32_t\n     -    to timestamp_t.\n     -\n     -    We drop 'detect incorrect generation number' from t5318-commit-graph.sh,\n     -    which tests if verify can detect if a commit graph have\n     -    GENERATION_NUMBER_ZERO for a commit, followed by a non-zero generation.\n     -    With corrected commit dates, GENERATION_NUMBER_ZERO is possible only if\n     -    one of dates is Unix epoch zero.\n     +    In a preparatory step, let's return timestamp_t values from\n     +    commit_graph_generation(), use timestamp_t for local variables and\n     +    define GENERATION_NUMBER_INFINITY as (2 ^ 63 - 1) instead.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n     - ## blame.c ##\n     -@@ blame.c: static int maybe_changed_path(struct repository *r,\n     - \tif (!bd)\n     - \t\treturn 1;\n     - \n     --\tif (commit_graph_generation(origin->commit) == GENERATION_NUMBER_INFINITY)\n     -+\tif (commit_graph_generation(origin->commit) == GENERATION_NUMBER_V2_INFINITY)\n     - \t\treturn 1;\n     - \n     - \tfilter = get_bloom_filter(r, origin->commit, 0);\n     -\n       ## commit-graph.c ##\n     -@@ commit-graph.c: void git_test_write_commit_graph_or_die(void)\n     - /* Remember to update object flag allocation in object.h */\n     - #define REACHABLE       (1u<<15)\n     - \n     -+define_commit_slab(topo_level_slab, uint32_t);\n     -+\n     - /* Keep track of the order in which commits are added to our list. */\n     - define_commit_slab(commit_pos, int);\n     - static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n      @@ commit-graph.c: uint32_t commit_graph_position(const struct commit *c)\n       \treturn data ? data->graph_pos : COMMIT_NOT_FROM_GRAPH;\n       }\n     @@ commit-graph.c: uint32_t commit_graph_position(const struct commit *c)\n       {\n       \tstruct commit_graph_data *data =\n       \t\tcommit_graph_data_slab_peek(&commit_graph_data_slab, c);\n     - \n     - \tif (!data)\n     --\t\treturn GENERATION_NUMBER_INFINITY;\n     -+\t\treturn GENERATION_NUMBER_V2_INFINITY;\n     - \telse if (data->graph_pos == COMMIT_NOT_FROM_GRAPH)\n     --\t\treturn GENERATION_NUMBER_INFINITY;\n     -+\t\treturn GENERATION_NUMBER_V2_INFINITY;\n     - \n     - \treturn data->generation;\n     - }\n      @@ commit-graph.c: uint32_t commit_graph_generation(const struct commit *c)\n       int compare_commits_by_gen(const void *_a, const void *_b)\n       {\n     @@ commit-graph.c: static int commit_gen_cmp(const void *va, const void *vb)\n       \n      -\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n      -\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n     -+\ttimestamp_t generation_a = commit_graph_data_at(a)->generation;\n     -+\ttimestamp_t generation_b = commit_graph_data_at(b)->generation;\n     - \n     ++\tconst timestamp_t generation_a = commit_graph_data_at(a)->generation;\n     ++\tconst timestamp_t generation_b = commit_graph_data_at(b)->generation;\n       \t/* lower generation commits first */\n       \tif (generation_a < generation_b)\n     -@@ commit-graph.c: static int commit_gen_cmp(const void *va, const void *vb)\n     - \telse if (generation_a > generation_b)\n     - \t\treturn 1;\n     - \n     --\t/* use date as a heuristic when generations are equal */\n     --\tif (a->date < b->date)\n     --\t\treturn -1;\n     --\telse if (a->date > b->date)\n     --\t\treturn 1;\n     - \treturn 0;\n     - }\n     - \n     -@@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n     - \titem->date = (timestamp_t)((date_high << 32) | date_low);\n     - \n     - \tif (g->chunk_generation_data)\n     --\t\tgraph_data->generation = get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     -+\t{\n     -+\t\t/* Read corrected commit date offset from GDAT */\n     -+\t\tgraph_data->generation = item->date +\n     -+\t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     -+\t}\n     - \telse\n     -+\t\t/* Read topological level from CDAT */\n     - \t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n     - }\n     - \n     -@@ commit-graph.c: struct write_commit_graph_context {\n     - \tstruct progress *progress;\n     - \tint progress_done;\n     - \tuint64_t progress_cnt;\n     -+\tstruct topo_level_slab *topo_levels;\n     - \n     - \tchar *base_graph_name;\n     - \tint num_commit_graphs_before;\n     -@@ commit-graph.c: static void write_graph_chunk_data(struct hashfile *f, int hash_len,\n     - \t\telse\n     - \t\t\tpackedDate[0] = 0;\n     - \n     --\t\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n     -+\t\tpackedDate[0] |= htonl(*topo_level_slab_at(ctx->topo_levels, *list) << 2);\n     - \n     - \t\tpackedDate[1] = htonl((*list)->date);\n     - \t\thashwrite(f, packedDate, 8);\n     -@@ commit-graph.c: static void write_graph_chunk_generation_data(struct hashfile *f,\n     - \tstruct commit **list = ctx->commits.list;\n     - \tint count;\n     - \tfor (count = 0; count < ctx->commits.nr; count++, list++) {\n     -+\t\ttimestamp_t offset = commit_graph_data_at(*list)->generation - (*list)->date;\n     - \t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n     --\t\thashwrite_be32(f, commit_graph_data_at(*list)->generation);\n     -+\n     -+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX)\n     -+\t\t\toffset = GENERATION_NUMBER_V2_OFFSET_MAX;\n     -+\n     -+\t\thashwrite_be32(f, offset);\n     - \t}\n     - }\n     - \n     -@@ commit-graph.c: static void close_reachable(struct write_commit_graph_context *ctx)\n     - \tstop_progress(&ctx->progress);\n     - }\n     - \n     --static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n     -+static void compute_corrected_commit_date_offsets(struct write_commit_graph_context *ctx)\n     - {\n     - \tint i;\n     - \tstruct commit_list *list = NULL;\n     + \t\treturn -1;\n      @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n     - \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n     - \t\t\t\t\tctx->commits.nr);\n     - \tfor (i = 0; i < ctx->commits.nr; i++) {\n     --\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n     -+\t\tuint32_t topo_level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n     + \t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n       \n       \t\tdisplay_progress(ctx->progress, i + 1);\n      -\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n     --\t\t    generation != GENERATION_NUMBER_ZERO)\n     -+\t\tif (topo_level != GENERATION_NUMBER_INFINITY &&\n     -+\t\t    topo_level != GENERATION_NUMBER_ZERO)\n     ++\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n     + \t\t    generation != GENERATION_NUMBER_ZERO)\n       \t\t\tcontinue;\n       \n     - \t\tcommit_list_insert(ctx->commits.list[i], &list);\n      @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n     - \t\t\tstruct commit *current = list->item;\n     - \t\t\tstruct commit_list *parent;\n     - \t\t\tint all_parents_computed = 1;\n     --\t\t\tuint32_t max_generation = 0;\n     -+\t\t\tuint32_t max_level = 0;\n     -+\t\t\ttimestamp_t max_corrected_commit_date = current->date;\n     - \n       \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n     --\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n     -+\t\t\t\ttopo_level = *topo_level_slab_at(ctx->topo_levels, parent->item);\n     + \t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n       \n      -\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n     --\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n     -+\t\t\t\tif (topo_level == GENERATION_NUMBER_INFINITY ||\n     -+\t\t\t\t    topo_level == GENERATION_NUMBER_ZERO) {\n     ++\t\t\t\tif (generation == GENERATION_NUMBER_V1_INFINITY ||\n     + \t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n       \t\t\t\t\tall_parents_computed = 0;\n       \t\t\t\t\tcommit_list_insert(parent->item, &list);\n     - \t\t\t\t\tbreak;\n     --\t\t\t\t} else if (generation > max_generation) {\n     --\t\t\t\t\tmax_generation = generation;\n     -+\t\t\t\t} else {\n     -+\t\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(parent->item);\n     -+\n     -+\t\t\t\t\tif (topo_level > max_level)\n     -+\t\t\t\t\t\tmax_level = topo_level;\n     -+\n     -+\t\t\t\t\tif (data->generation > max_corrected_commit_date)\n     -+\t\t\t\t\t\tmax_corrected_commit_date = data->generation;\n     - \t\t\t\t}\n     - \t\t\t}\n     - \n     - \t\t\tif (all_parents_computed) {\n     - \t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n     - \n     --\t\t\t\tdata->generation = max_generation + 1;\n     --\t\t\t\tpop_commit(&list);\n     -+\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n     -+\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n     -+\n     -+\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n     -+\t\t\t\tdata->generation = max_corrected_commit_date + 1;\n     - \n     --\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n     --\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n     -+\t\t\t\tpop_commit(&list);\n     - \t\t\t}\n     - \t\t}\n     - \t}\n     -@@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n     - \tuint32_t i, count_distinct = 0;\n     - \tint res = 0;\n     - \tint replace = 0;\n     -+\tstruct topo_level_slab topo_levels;\n     - \n     - \tif (!commit_graph_compatible(the_repository))\n     - \t\treturn 0;\n     -@@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n     - \tctx->changed_paths = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n     - \tctx->total_bloom_filter_data_size = 0;\n     - \n     -+\tinit_topo_level_slab(&topo_levels);\n     -+\tctx->topo_levels = &topo_levels;\n     -+\n     - \tif (ctx->split) {\n     - \t\tstruct commit_graph *g;\n     - \t\tprepare_commit_graph(ctx->r);\n     -@@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n     - \t} else\n     - \t\tctx->num_commit_graphs_after = 1;\n     - \n     --\tcompute_generation_numbers(ctx);\n     -+\tcompute_corrected_commit_date_offsets(ctx);\n     - \n     - \tif (ctx->changed_paths)\n     - \t\tcompute_bloom_filters(ctx);\n      @@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n       \tfor (i = 0; i < g->num_commits; i++) {\n       \t\tstruct commit *graph_commit, *odb_commit;\n       \t\tstruct commit_list *graph_parents, *odb_parents;\n      -\t\tuint32_t max_generation = 0;\n      -\t\tuint32_t generation;\n     -+\t\ttimestamp_t max_parent_corrected_commit_date = 0;\n     -+\t\ttimestamp_t corrected_commit_date;\n     ++\t\ttimestamp_t max_generation = 0;\n     ++\t\ttimestamp_t generation;\n       \n       \t\tdisplay_progress(progress, i + 1);\n       \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n     -@@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n     - \t\t\t\t\t     oid_to_hex(&graph_parents->item->object.oid),\n     - \t\t\t\t\t     oid_to_hex(&odb_parents->item->object.oid));\n     - \n     --\t\t\tgeneration = commit_graph_generation(graph_parents->item);\n     --\t\t\tif (generation > max_generation)\n     --\t\t\t\tmax_generation = generation;\n     -+\t\t\tcorrected_commit_date = commit_graph_generation(graph_parents->item);\n     -+\t\t\tif (corrected_commit_date > max_parent_corrected_commit_date)\n     -+\t\t\t\tmax_parent_corrected_commit_date = corrected_commit_date;\n     - \n     - \t\t\tgraph_parents = graph_parents->next;\n     - \t\t\todb_parents = odb_parents->next;\n     -@@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n     - \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n     - \t\t\tcontinue;\n     - \n     --\t\t/*\n     --\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n     --\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n     --\t\t * extra logic in the following condition.\n     --\t\t */\n     --\t\tif (max_generation == GENERATION_NUMBER_MAX)\n     --\t\t\tmax_generation--;\n     --\n     --\t\tgeneration = commit_graph_generation(graph_commit);\n     --\t\tif (generation != max_generation + 1)\n     --\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n     -+\t\tcorrected_commit_date = commit_graph_generation(graph_commit);\n     -+\t\tif (corrected_commit_date < max_parent_corrected_commit_date + 1)\n     -+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n     - \t\t\t\t     oid_to_hex(&cur_oid),\n     --\t\t\t\t     generation,\n     --\t\t\t\t     max_generation + 1);\n     -+\t\t\t\t     corrected_commit_date,\n     -+\t\t\t\t     max_parent_corrected_commit_date + 1);\n     - \n     - \t\tif (graph_commit->date != odb_commit->date)\n     - \t\t\tgraph_report(_(\"commit date for commit %s in commit-graph is %\"PRItime\" != %\"PRItime),\n      \n       ## commit-graph.h ##\n      @@ commit-graph.h: void disable_commit_graph(struct repository *r);\n     @@ commit-reach.c: static int queue_has_nonstale(struct prio_queue *queue)\n       \tstruct commit_list *result = NULL;\n       \tint i;\n      -\tuint32_t last_gen = GENERATION_NUMBER_INFINITY;\n     -+\ttimestamp_t last_gen = GENERATION_NUMBER_V2_INFINITY;\n     ++\ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n       \n       \tif (!min_generation)\n       \t\tqueue.compare = compare_commits_by_commit_date;\n     @@ commit-reach.c: int repo_in_merge_bases_many(struct repository *r, struct commit\n       \tstruct commit_list *bases;\n       \tint ret = 0, i;\n      -\tuint32_t generation, min_generation = GENERATION_NUMBER_INFINITY;\n     -+\ttimestamp_t generation, min_generation = GENERATION_NUMBER_V2_INFINITY;\n     ++\ttimestamp_t generation, min_generation = GENERATION_NUMBER_INFINITY;\n       \n       \tif (repo_parse_commit(r, commit))\n       \t\treturn ret;\n     @@ commit-reach.c: static enum contains_result contains_tag_algo(struct commit *can\n       \tstruct contains_stack contains_stack = { 0, 0, NULL };\n       \tenum contains_result result;\n      -\tuint32_t cutoff = GENERATION_NUMBER_INFINITY;\n     -+\ttimestamp_t cutoff = GENERATION_NUMBER_V2_INFINITY;\n     ++\ttimestamp_t cutoff = GENERATION_NUMBER_INFINITY;\n       \tconst struct commit_list *p;\n       \n       \tfor (p = want; p; p = p->next) {\n     @@ commit-reach.c: int can_all_from_reach(struct commit_list *from, struct commit_l\n       \tstruct commit_list *from_iter = from, *to_iter = to;\n       \tint result;\n      -\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n     -+\ttimestamp_t min_generation = GENERATION_NUMBER_V2_INFINITY;\n     ++\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n       \n       \twhile (from_iter) {\n       \t\tadd_object_array(&from_iter->item->object, NULL, &from_objs);\n     @@ commit-reach.c: struct commit_list *get_reachable_subset(struct commit **from, i\n       \tstruct commit **to_last = to + nr_to;\n       \tstruct commit **from_last = from + nr_from;\n      -\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n     -+\ttimestamp_t min_generation = GENERATION_NUMBER_V2_INFINITY;\n     ++\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n       \tint num_to_find = 0;\n       \n       \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n     @@ commit-reach.h: int can_all_from_reach_with_flag(struct object_array *from,\n      \n       ## commit.h ##\n      @@\n     + #include \"commit-slab.h\"\n     + \n     + #define COMMIT_NOT_FROM_GRAPH 0xFFFFFFFF\n     +-#define GENERATION_NUMBER_INFINITY 0xFFFFFFFF\n     ++#define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n     ++#define GENERATION_NUMBER_V1_INFINITY 0xFFFFFFFF\n       #define GENERATION_NUMBER_MAX 0x3FFFFFFF\n       #define GENERATION_NUMBER_ZERO 0\n       \n     -+#define GENERATION_NUMBER_V2_INFINITY ((1ULL << 63) - 1)\n     -+#define GENERATION_NUMBER_V2_OFFSET_MAX 0xFFFFFFFF\n     -+\n     - struct commit_list {\n     - \tstruct commit *item;\n     - \tstruct commit_list *next;\n      \n       ## revision.c ##\n     -@@ revision.c: static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n     - \tif (!revs->repo->objects->commit_graph)\n     - \t\treturn -1;\n     - \n     --\tif (commit_graph_generation(commit) == GENERATION_NUMBER_INFINITY)\n     -+\tif (commit_graph_generation(commit) == GENERATION_NUMBER_V2_INFINITY)\n     - \t\treturn -1;\n     - \n     - \tfilter = get_bloom_filter(revs->repo, commit, 0);\n      @@ revision.c: define_commit_slab(indegree_slab, int);\n       define_commit_slab(author_date_slab, timestamp_t);\n       \n     @@ revision.c: static void indegree_walk_step(struct rev_info *revs)\n       \tstruct topo_walk_info *info = revs->topo_walk_info;\n       \tstruct commit *c;\n      @@ revision.c: static void init_topo_walk(struct rev_info *revs)\n     - \tinfo->explore_queue.compare = compare_commits_by_gen_then_commit_date;\n     - \tinfo->indegree_queue.compare = compare_commits_by_gen_then_commit_date;\n     - \n     --\tinfo->min_generation = GENERATION_NUMBER_INFINITY;\n     -+\tinfo->min_generation = GENERATION_NUMBER_V2_INFINITY;\n     + \tinfo->min_generation = GENERATION_NUMBER_INFINITY;\n       \tfor (list = revs->commits; list; list = list->next) {\n       \t\tstruct commit *c = list->item;\n      -\t\tuint32_t generation;\n     @@ revision.c: static void expand_topo_walk(struct rev_info *revs, struct commit *c\n       \t\tif (parent->object.flags & UNINTERESTING)\n       \t\t\tcontinue;\n      \n     - ## t/t5318-commit-graph.sh ##\n     -@@ t/t5318-commit-graph.sh: test_expect_success 'detect incorrect generation number' '\n     - \t\t\"generation for commit\"\n     - '\n     - \n     --test_expect_success 'detect incorrect generation number' '\n     -+test_expect_failure 'detect incorrect generation number' '\n     - \tcorrupt_graph_and_verify $GRAPH_BYTE_COMMIT_GENERATION \"\\00\" \\\n     - \t\t\"non-zero generation number\"\n     - '\n     -\n       ## upload-pack.c ##\n      @@ upload-pack.c: static int got_oid(struct upload_pack_data *data,\n       \n  -:  ---------- >  7:  bfe1473201 commit-graph: implement corrected commit date\n  -:  ---------- >  8:  833779ad53 commit-graph: handle mixed generation commit chains\n  -:  ---------- >  9:  58a2d5da01 commit-reach: use corrected commit dates in paint_down_to_common()\n  -:  ---------- > 10:  4c34294602 doc: add corrected commit date info\n\n-- \ngitgitgadget\n"},{"id":"403221","messageId":"833779ad53eb4f57ae514f4e8964e397845f1ddd.1596941625.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"[PATCH v2 08/10] commit-graph: handle mixed generation commit chains","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-09T02:53:42Z","receivedAt":"2020-08-09T02:54:03Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nAs corrected commit dates and topological levels cannot be compared\ndirectly, we must handle commit graph chains with mixed generation\nnumber definitions.\n\nWhile reading a commit graph file, we disable generation numbers if the\nchain contains mixed generation numbers.\n\nWhile writing to commit graph chain, we write generation data chunk only\nif the previous tip of chain had a generation data chunk. Using\n`--split=replace` overwrites the existing chain and writes generation\ndata chunk regardless of previous tip.\n\nIn t5324-split-commit-graph, we set up a repo with twelve commits and\nwrite a base commit graph file with no generation data chunk. When add\nthree commits and write to chain again, Git does not write generation\ndata chunk even without setting GIT_TEST_COMMIT_GRAPH_NO_GDAT=1. Then,\nas we replace the existing chain, Git writes a commit graph file with\ngeneration data chunk.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                | 14 ++++++++\n t/t5324-split-commit-graph.sh | 66 +++++++++++++++++++++++++++++++++++\n 2 files changed, 80 insertions(+)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex d0f977852b..c6b6111adf 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -674,6 +674,14 @@ int generation_numbers_enabled(struct repository *r)\n \tif (!g->num_commits)\n \t\treturn 0;\n \n+\t/* We cannot compare topological levels and corrected commit dates */\n+\twhile (g->base_graph) {\n+\t\twarning(_(\"commit-graph-chain contains mixed generation versions\"));\n+\t\tif ((g->chunk_generation_data == NULL) ^ (g->base_graph->chunk_generation_data == NULL))\n+\t\t\treturn 0;\n+\t\tg = g->base_graph;\n+\t}\n+\n \tfirst_generation = get_be32(g->chunk_commit_data +\n \t\t\t\t    g->hash_len + 8) >> 2;\n \n@@ -2186,6 +2194,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \t\tg = ctx->r->objects->commit_graph;\n \n+\t\tif (g && !g->chunk_generation_data)\n+\t\t\tctx->write_generation_data = 0;\n+\n \t\twhile (g) {\n \t\t\tctx->num_commit_graphs_before++;\n \t\t\tg = g->base_graph;\n@@ -2204,6 +2215,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \t\tif (ctx->split_opts)\n \t\t\treplace = ctx->split_opts->flags & COMMIT_GRAPH_SPLIT_REPLACE;\n+\n+\t\tif (replace)\n+\t\t\tctx->write_generation_data = 1;\n \t}\n \n \tctx->approx_nr_objects = approximate_object_count();\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 6b25c3d9ce..1a9be5e656 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -425,4 +425,70 @@ done <<\\EOF\n 0600 -r--------\n EOF\n \n+test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n+\tmkdir mixed &&\n+\tgraphdir=\".git/objects/info/commit-graphs\" &&\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit init &&\n+\tgit config core.commitGraph true &&\n+\tgit config gc.writeCommitGraph false &&\n+\tfor i in $(test_seq 3)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit reset --hard commits/1 &&\n+\tfor i in $(test_seq 4 5)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit reset --hard commits/2 &&\n+\tfor i in $(test_seq 6 10)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit reset --hard commits/2 &&\n+\tgit merge commits/4 &&\n+\tgit branch merge/1 &&\n+\tgit reset --hard commits/4 &&\n+\tgit merge commits/6 &&\n+\tgit branch merge/2 &&\n+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split &&\n+\ttest-tool read-graph >output &&\n+\tcat >expect <<-EOF &&\n+\theader: 43475048 1 1 3 0\n+\tnum_commits: 12\n+\tchunks: oid_fanout oid_lookup commit_metadata\n+\tEOF\n+\ttest_cmp expect output\n+'\n+\n+test_expect_success 'does not write generation data chunk if not present on existing tip' '\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit reset --hard commits/3 &&\n+\tgit merge merge/1 &&\n+\tgit merge commits/5 &&\n+\tgit merge merge/2 &&\n+\tgit branch merge/3 &&\n+\tgit commit-graph write --reachable --split &&\n+\ttest-tool read-graph >output &&\n+\tcat >expect <<-EOF &&\n+\theader: 43475048 1 1 4 1\n+\tnum_commits: 3\n+\tchunks: oid_fanout oid_lookup commit_metadata\n+\tEOF\n+\ttest_cmp expect output\n+'\n+\n+test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit commit-graph write --reachable --split='replace' &&\n+\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\tverify_chain_files_exist $graphdir &&\n+\tgraph_read_expect 15\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"403222","messageId":"58a2d5da0105e6572305b07d4e39ef6be9ee0044.1596941625.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"[PATCH v2 09/10] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-09T02:53:43Z","receivedAt":"2020-08-09T02:54:03Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith corrected commit dates implemented, we no longer have to rely on\ncommit date as a heuristic in paint_down_to_common().\n\nt6024-recursive-merge setups a unique repository where all commits have\nthe same committer date without well-defined merge-base. As this has\nalready caused problems (as noted in 859fdc0 (commit-graph: define\nGIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph within the\ntest script.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c             | 14 ++++++++++++++\n commit-graph.h             |  6 ++++++\n commit-reach.c             |  2 +-\n t/t6024-recursive-merge.sh |  4 +++-\n 4 files changed, 24 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex c6b6111adf..eb78af3dad 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -688,6 +688,20 @@ int generation_numbers_enabled(struct repository *r)\n \treturn !!first_generation;\n }\n \n+int corrected_commit_dates_enabled(struct repository *r)\n+{\n+\tstruct commit_graph *g;\n+\tif (!prepare_commit_graph(r))\n+\t\treturn 0;\n+\n+\tg = r->objects->commit_graph;\n+\n+\tif (!g->num_commits)\n+\t\treturn 0;\n+\n+\treturn !!g->chunk_generation_data;\n+}\n+\n static void close_commit_graph_one(struct commit_graph *g)\n {\n \tif (!g)\ndiff --git a/commit-graph.h b/commit-graph.h\nindex f89614ecd5..d3a485faa6 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -89,6 +89,12 @@ struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size);\n  */\n int generation_numbers_enabled(struct repository *r);\n \n+/*\n+ * Return 1 if and only if the repository has a commit-graph\n+ * file and generation data chunk has been written for the file.\n+ */\n+int corrected_commit_dates_enabled(struct repository *r);\n+\n enum commit_graph_write_flags {\n \tCOMMIT_GRAPH_WRITE_APPEND     = (1 << 0),\n \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\ndiff --git a/commit-reach.c b/commit-reach.c\nindex 470bc80139..3a1b925274 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -39,7 +39,7 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \tint i;\n \ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n-\tif (!min_generation)\n+\tif (!min_generation && !corrected_commit_dates_enabled(r))\n \t\tqueue.compare = compare_commits_by_commit_date;\n \n \tone->object.flags |= PARENT1;\ndiff --git a/t/t6024-recursive-merge.sh b/t/t6024-recursive-merge.sh\nindex 332cfc53fd..d3def66e7d 100755\n--- a/t/t6024-recursive-merge.sh\n+++ b/t/t6024-recursive-merge.sh\n@@ -15,6 +15,8 @@ GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n export GIT_COMMITTER_DATE\n \n test_expect_success 'setup tests' '\n+\tGIT_TEST_COMMIT_GRAPH=0 &&\n+\texport GIT_TEST_COMMIT_GRAPH &&\n \techo 1 >a1 &&\n \tgit add a1 &&\n \tGIT_AUTHOR_DATE=\"2006-12-12 23:00:00\" git commit -m 1 a1 &&\n@@ -66,7 +68,7 @@ test_expect_success 'setup tests' '\n '\n \n test_expect_success 'combined merge conflicts' '\n-\ttest_must_fail env GIT_TEST_COMMIT_GRAPH=0 git merge -m final G\n+\ttest_must_fail git merge -m final G\n '\n \n test_expect_success 'result contains a conflict' '\n-- \ngitgitgadget\n\n"},{"id":"403223","messageId":"bfe14732014807ff19f943cdf51068f0d3043c30.1596941625.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"[PATCH v2 07/10] commit-graph: implement corrected commit date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-09T02:53:41Z","receivedAt":"2020-08-09T02:54:05Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith most of preparations done, let's implement corrected commit date\noffset. We add a new commit-slab to store topogical levels while\nwriting commit graph and upgrade the generation member in struct\ncommit_graph_data to a 64-bit timestamp. We store topological levels to\nensure that older versions of Git will still have the performance\nbenefits from generation number v2.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 89 ++++++++++++++++++++++++++++----------------------\n commit.h       |  1 +\n 2 files changed, 51 insertions(+), 39 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 42f3ec5460..d0f977852b 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -65,6 +65,8 @@ void git_test_write_commit_graph_or_die(void)\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n+define_commit_slab(topo_level_slab, uint32_t);\n+\n /* Keep track of the order in which commits are added to our list. */\n define_commit_slab(commit_pos, int);\n static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n@@ -168,11 +170,6 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \telse if (generation_a > generation_b)\n \t\treturn 1;\n \n-\t/* use date as a heuristic when generations are equal */\n-\tif (a->date < b->date)\n-\t\treturn -1;\n-\telse if (a->date > b->date)\n-\t\treturn 1;\n \treturn 0;\n }\n \n@@ -767,7 +764,10 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n \tif (g->chunk_generation_data)\n-\t\tgraph_data->generation = get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n+\t{\n+\t\tgraph_data->generation = item->date +\n+\t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n+\t}\n \telse\n \t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n@@ -948,6 +948,7 @@ struct write_commit_graph_context {\n \tstruct progress *progress;\n \tint progress_done;\n \tuint64_t progress_cnt;\n+\tstruct topo_level_slab *topo_levels;\n \n \tchar *base_graph_name;\n \tint num_commit_graphs_before;\n@@ -1106,7 +1107,7 @@ static int write_graph_chunk_data(struct hashfile *f,\n \t\telse\n \t\t\tpackedDate[0] = 0;\n \n-\t\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n+\t\tpackedDate[0] |= htonl(*topo_level_slab_at(ctx->topo_levels, *list) << 2);\n \n \t\tpackedDate[1] = htonl((*list)->date);\n \t\thashwrite(f, packedDate, 8);\n@@ -1123,8 +1124,13 @@ static int write_graph_chunk_generation_data(struct hashfile *f,\n \tint i;\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tstruct commit *c = ctx->commits.list[i];\n+\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n \t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n-\t\thashwrite_be32(f, commit_graph_data_at(c)->generation);\n+\n+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX)\n+\t\t\toffset = GENERATION_NUMBER_V2_OFFSET_MAX;\n+\n+\t\thashwrite_be32(f, offset);\n \t}\n \n \treturn 0;\n@@ -1360,11 +1366,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n+\t\tuint32_t topo_level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n-\t\t    generation != GENERATION_NUMBER_ZERO)\n+\t\tif (topo_level != GENERATION_NUMBER_V1_INFINITY &&\n+\t\t    topo_level != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1372,29 +1378,38 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit *current = list->item;\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n-\t\t\tuint32_t max_generation = 0;\n+\t\t\tuint32_t max_level = 0;\n+\t\t\ttimestamp_t max_corrected_commit_date = current->date - 1;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n-\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n+\t\t\t\ttopo_level = *topo_level_slab_at(ctx->topo_levels, parent->item);\n \n-\t\t\t\tif (generation == GENERATION_NUMBER_V1_INFINITY ||\n-\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n+\t\t\t\tif (topo_level == GENERATION_NUMBER_V1_INFINITY ||\n+\t\t\t\t    topo_level == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (generation > max_generation) {\n-\t\t\t\t\tmax_generation = generation;\n+\t\t\t\t} else {\n+\t\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(parent->item);\n+\n+\t\t\t\t\tif (topo_level > max_level)\n+\t\t\t\t\t\tmax_level = topo_level;\n+\n+\t\t\t\t\tif (data->generation > max_corrected_commit_date)\n+\t\t\t\t\t\tmax_corrected_commit_date = data->generation;\n \t\t\t\t}\n \t\t\t}\n \n \t\t\tif (all_parents_computed) {\n \t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n \n-\t\t\t\tdata->generation = max_generation + 1;\n-\t\t\t\tpop_commit(&list);\n+\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n+\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n+\n+\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n+\t\t\t\tdata->generation = max_corrected_commit_date + 1;\n \n-\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n-\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n+\t\t\t\tpop_commit(&list);\n \t\t\t}\n \t\t}\n \t}\n@@ -2132,6 +2147,7 @@ int write_commit_graph(struct object_directory *odb,\n \tuint32_t i, count_distinct = 0;\n \tint res = 0;\n \tint replace = 0;\n+\tstruct topo_level_slab topo_levels;\n \n \tif (!commit_graph_compatible(the_repository))\n \t\treturn 0;\n@@ -2146,6 +2162,9 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->total_bloom_filter_data_size = 0;\n \tctx->write_generation_data = !git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0);\n \n+\tinit_topo_level_slab(&topo_levels);\n+\tctx->topo_levels = &topo_levels;\n+\n \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n \t\tctx->changed_paths = 1;\n \tif (!(flags & COMMIT_GRAPH_NO_WRITE_BLOOM_FILTERS)) {\n@@ -2387,8 +2406,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \tfor (i = 0; i < g->num_commits; i++) {\n \t\tstruct commit *graph_commit, *odb_commit;\n \t\tstruct commit_list *graph_parents, *odb_parents;\n-\t\ttimestamp_t max_generation = 0;\n-\t\ttimestamp_t generation;\n+\t\ttimestamp_t max_parent_corrected_commit_date = 0;\n+\t\ttimestamp_t corrected_commit_date;\n \n \t\tdisplay_progress(progress, i + 1);\n \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n@@ -2427,9 +2446,9 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t\t\t\t     oid_to_hex(&graph_parents->item->object.oid),\n \t\t\t\t\t     oid_to_hex(&odb_parents->item->object.oid));\n \n-\t\t\tgeneration = commit_graph_generation(graph_parents->item);\n-\t\t\tif (generation > max_generation)\n-\t\t\t\tmax_generation = generation;\n+\t\t\tcorrected_commit_date = commit_graph_generation(graph_parents->item);\n+\t\t\tif (corrected_commit_date > max_parent_corrected_commit_date)\n+\t\t\t\tmax_parent_corrected_commit_date = corrected_commit_date;\n \n \t\t\tgraph_parents = graph_parents->next;\n \t\t\todb_parents = odb_parents->next;\n@@ -2451,20 +2470,12 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n \t\t\tcontinue;\n \n-\t\t/*\n-\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n-\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n-\t\t * extra logic in the following condition.\n-\t\t */\n-\t\tif (max_generation == GENERATION_NUMBER_MAX)\n-\t\t\tmax_generation--;\n-\n-\t\tgeneration = commit_graph_generation(graph_commit);\n-\t\tif (generation != max_generation + 1)\n-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n+\t\tcorrected_commit_date = commit_graph_generation(graph_commit);\n+\t\tif (corrected_commit_date < max_parent_corrected_commit_date + 1)\n+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n \t\t\t\t     oid_to_hex(&cur_oid),\n-\t\t\t\t     generation,\n-\t\t\t\t     max_generation + 1);\n+\t\t\t\t     corrected_commit_date,\n+\t\t\t\t     max_parent_corrected_commit_date + 1);\n \n \t\tif (graph_commit->date != odb_commit->date)\n \t\t\tgraph_report(_(\"commit date for commit %s in commit-graph is %\"PRItime\" != %\"PRItime),\ndiff --git a/commit.h b/commit.h\nindex bc0732a4fe..bb846e0025 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -15,6 +15,7 @@\n #define GENERATION_NUMBER_V1_INFINITY 0xFFFFFFFF\n #define GENERATION_NUMBER_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n+#define GENERATION_NUMBER_V2_OFFSET_MAX 0xFFFFFFFF\n \n struct commit_list {\n \tstruct commit *item;\n-- \ngitgitgadget\n\n"},{"id":"403224","messageId":"1aa2a00a7a2a0e4c884bf95261b5e308c3611fbc.1596941625.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"[PATCH v2 06/10] commit-graph: return 64-bit generation number","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-09T02:53:40Z","receivedAt":"2020-08-09T02:54:07Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a preparatory step, let's return timestamp_t values from\ncommit_graph_generation(), use timestamp_t for local variables and\ndefine GENERATION_NUMBER_INFINITY as (2 ^ 63 - 1) instead.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 18 +++++++++---------\n commit-graph.h |  4 ++--\n commit-reach.c | 32 ++++++++++++++++----------------\n commit-reach.h |  2 +-\n commit.h       |  3 ++-\n revision.c     | 10 +++++-----\n upload-pack.c  |  2 +-\n 7 files changed, 36 insertions(+), 35 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex d5da1e8028..42f3ec5460 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -100,7 +100,7 @@ uint32_t commit_graph_position(const struct commit *c)\n \treturn data ? data->graph_pos : COMMIT_NOT_FROM_GRAPH;\n }\n \n-uint32_t commit_graph_generation(const struct commit *c)\n+timestamp_t commit_graph_generation(const struct commit *c)\n {\n \tstruct commit_graph_data *data =\n \t\tcommit_graph_data_slab_peek(&commit_graph_data_slab, c);\n@@ -116,8 +116,8 @@ uint32_t commit_graph_generation(const struct commit *c)\n int compare_commits_by_gen(const void *_a, const void *_b)\n {\n \tconst struct commit *a = _a, *b = _b;\n-\tconst uint32_t generation_a = commit_graph_generation(a);\n-\tconst uint32_t generation_b = commit_graph_generation(b);\n+\tconst timestamp_t generation_a = commit_graph_generation(a);\n+\tconst timestamp_t generation_b = commit_graph_generation(b);\n \n \t/* older commits first */\n \tif (generation_a < generation_b)\n@@ -160,8 +160,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n-\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n+\tconst timestamp_t generation_a = commit_graph_data_at(a)->generation;\n+\tconst timestamp_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n@@ -1363,7 +1363,7 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n+\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n \t\t    generation != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n@@ -1377,7 +1377,7 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n \t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n \n-\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n+\t\t\t\tif (generation == GENERATION_NUMBER_V1_INFINITY ||\n \t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n@@ -2387,8 +2387,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \tfor (i = 0; i < g->num_commits; i++) {\n \t\tstruct commit *graph_commit, *odb_commit;\n \t\tstruct commit_list *graph_parents, *odb_parents;\n-\t\tuint32_t max_generation = 0;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t max_generation = 0;\n+\t\ttimestamp_t generation;\n \n \t\tdisplay_progress(progress, i + 1);\n \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex cc232e0678..f89614ecd5 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -140,13 +140,13 @@ void disable_commit_graph(struct repository *r);\n \n struct commit_graph_data {\n \tuint32_t graph_pos;\n-\tuint32_t generation;\n+\ttimestamp_t generation;\n };\n \n /*\n  * Commits should be parsed before accessing generation, graph positions.\n  */\n-uint32_t commit_graph_generation(const struct commit *);\n+timestamp_t commit_graph_generation(const struct commit *);\n uint32_t commit_graph_position(const struct commit *);\n \n int compare_commits_by_gen(const void *_a, const void *_b);\ndiff --git a/commit-reach.c b/commit-reach.c\nindex c83cc291e7..470bc80139 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -32,12 +32,12 @@ static int queue_has_nonstale(struct prio_queue *queue)\n static struct commit_list *paint_down_to_common(struct repository *r,\n \t\t\t\t\t\tstruct commit *one, int n,\n \t\t\t\t\t\tstruct commit **twos,\n-\t\t\t\t\t\tint min_generation)\n+\t\t\t\t\t\ttimestamp_t min_generation)\n {\n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \tstruct commit_list *result = NULL;\n \tint i;\n-\tuint32_t last_gen = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n \tif (!min_generation)\n \t\tqueue.compare = compare_commits_by_commit_date;\n@@ -58,10 +58,10 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \t\tstruct commit *commit = prio_queue_get(&queue);\n \t\tstruct commit_list *parents;\n \t\tint flags;\n-\t\tuint32_t generation = commit_graph_generation(commit);\n+\t\ttimestamp_t generation = commit_graph_generation(commit);\n \n \t\tif (min_generation && generation > last_gen)\n-\t\t\tBUG(\"bad generation skip %8x > %8x at %s\",\n+\t\t\tBUG(\"bad generation skip %\"PRItime\" > %\"PRItime\" at %s\",\n \t\t\t    generation, last_gen,\n \t\t\t    oid_to_hex(&commit->object.oid));\n \t\tlast_gen = generation;\n@@ -177,12 +177,12 @@ static int remove_redundant(struct repository *r, struct commit **array, int cnt\n \t\trepo_parse_commit(r, array[i]);\n \tfor (i = 0; i < cnt; i++) {\n \t\tstruct commit_list *common;\n-\t\tuint32_t min_generation = commit_graph_generation(array[i]);\n+\t\ttimestamp_t min_generation = commit_graph_generation(array[i]);\n \n \t\tif (redundant[i])\n \t\t\tcontinue;\n \t\tfor (j = filled = 0; j < cnt; j++) {\n-\t\t\tuint32_t curr_generation;\n+\t\t\ttimestamp_t curr_generation;\n \t\t\tif (i == j || redundant[j])\n \t\t\t\tcontinue;\n \t\t\tfilled_index[filled] = j;\n@@ -321,7 +321,7 @@ int repo_in_merge_bases_many(struct repository *r, struct commit *commit,\n {\n \tstruct commit_list *bases;\n \tint ret = 0, i;\n-\tuint32_t generation, min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t generation, min_generation = GENERATION_NUMBER_INFINITY;\n \n \tif (repo_parse_commit(r, commit))\n \t\treturn ret;\n@@ -470,7 +470,7 @@ static int in_commit_list(const struct commit_list *want, struct commit *c)\n static enum contains_result contains_test(struct commit *candidate,\n \t\t\t\t\t  const struct commit_list *want,\n \t\t\t\t\t  struct contains_cache *cache,\n-\t\t\t\t\t  uint32_t cutoff)\n+\t\t\t\t\t  timestamp_t cutoff)\n {\n \tenum contains_result *cached = contains_cache_at(cache, candidate);\n \n@@ -506,11 +506,11 @@ static enum contains_result contains_tag_algo(struct commit *candidate,\n {\n \tstruct contains_stack contains_stack = { 0, 0, NULL };\n \tenum contains_result result;\n-\tuint32_t cutoff = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t cutoff = GENERATION_NUMBER_INFINITY;\n \tconst struct commit_list *p;\n \n \tfor (p = want; p; p = p->next) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = p->item;\n \t\tload_commit_graph_info(the_repository, c);\n \t\tgeneration = commit_graph_generation(c);\n@@ -565,7 +565,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation)\n+\t\t\t\t timestamp_t min_generation)\n {\n \tstruct commit **list = NULL;\n \tint i;\n@@ -666,13 +666,13 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \ttime_t min_commit_date = cutoff_by_min_date ? from->item->date : 0;\n \tstruct commit_list *from_iter = from, *to_iter = to;\n \tint result;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \n \twhile (from_iter) {\n \t\tadd_object_array(&from_iter->item->object, NULL, &from_objs);\n \n \t\tif (!parse_commit(from_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (from_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = from_iter->item->date;\n \n@@ -686,7 +686,7 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \n \twhile (to_iter) {\n \t\tif (!parse_commit(to_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (to_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = to_iter->item->date;\n \n@@ -726,13 +726,13 @@ struct commit_list *get_reachable_subset(struct commit **from, int nr_from,\n \tstruct commit_list *found_commits = NULL;\n \tstruct commit **to_last = to + nr_to;\n \tstruct commit **from_last = from + nr_from;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \tint num_to_find = 0;\n \n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \n \tfor (item = to; item < to_last; item++) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = *item;\n \n \t\tparse_commit(c);\ndiff --git a/commit-reach.h b/commit-reach.h\nindex b49ad71a31..148b56fea5 100644\n--- a/commit-reach.h\n+++ b/commit-reach.h\n@@ -87,7 +87,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation);\n+\t\t\t\t timestamp_t min_generation);\n int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \t\t       int commit_date_cutoff);\n \ndiff --git a/commit.h b/commit.h\nindex e901538909..bc0732a4fe 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -11,7 +11,8 @@\n #include \"commit-slab.h\"\n \n #define COMMIT_NOT_FROM_GRAPH 0xFFFFFFFF\n-#define GENERATION_NUMBER_INFINITY 0xFFFFFFFF\n+#define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n+#define GENERATION_NUMBER_V1_INFINITY 0xFFFFFFFF\n #define GENERATION_NUMBER_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n \ndiff --git a/revision.c b/revision.c\nindex 4ec82ed5ab..bd7b39c806 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3292,7 +3292,7 @@ define_commit_slab(indegree_slab, int);\n define_commit_slab(author_date_slab, timestamp_t);\n \n struct topo_walk_info {\n-\tuint32_t min_generation;\n+\ttimestamp_t min_generation;\n \tstruct prio_queue explore_queue;\n \tstruct prio_queue indegree_queue;\n \tstruct prio_queue topo_queue;\n@@ -3338,7 +3338,7 @@ static void explore_walk_step(struct rev_info *revs)\n }\n \n static void explore_to_depth(struct rev_info *revs,\n-\t\t\t     uint32_t gen_cutoff)\n+\t\t\t     timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3381,7 +3381,7 @@ static void indegree_walk_step(struct rev_info *revs)\n }\n \n static void compute_indegrees_to_depth(struct rev_info *revs,\n-\t\t\t\t       uint32_t gen_cutoff)\n+\t\t\t\t       timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3439,7 +3439,7 @@ static void init_topo_walk(struct rev_info *revs)\n \tinfo->min_generation = GENERATION_NUMBER_INFINITY;\n \tfor (list = revs->commits; list; list = list->next) {\n \t\tstruct commit *c = list->item;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (parse_commit_gently(c, 1))\n \t\t\tcontinue;\n@@ -3500,7 +3500,7 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n \tfor (p = commit->parents; p; p = p->next) {\n \t\tstruct commit *parent = p->item;\n \t\tint *pi;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (parent->object.flags & UNINTERESTING)\n \t\t\tcontinue;\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 8673741070..18ee29db67 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -490,7 +490,7 @@ static int got_oid(struct upload_pack_data *data,\n \n static int ok_to_give_up(struct upload_pack_data *data)\n {\n-\tuint32_t min_generation = GENERATION_NUMBER_ZERO;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_ZERO;\n \n \tif (!data->have_obj.nr)\n \t\treturn 0;\n-- \ngitgitgadget\n\n"},{"id":"403225","messageId":"4c34294602b23f4427b024bbd38e4403a397fc50.1596941625.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"[PATCH v2 10/10] doc: add corrected commit date info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-09T02:53:44Z","receivedAt":"2020-08-09T02:54:09Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith generation data chunk and corrected commit dates implemented, let's\nupdate the technical documentation for commit-graph.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n .../technical/commit-graph-format.txt         | 12 ++---\n Documentation/technical/commit-graph.txt      | 45 ++++++++++++-------\n 2 files changed, 36 insertions(+), 21 deletions(-)\n\ndiff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\nindex 440541045d..71c43884ec 100644\n--- a/Documentation/technical/commit-graph-format.txt\n+++ b/Documentation/technical/commit-graph-format.txt\n@@ -4,11 +4,7 @@ Git commit graph format\n The Git commit graph stores a list of commit OIDs and some associated\n metadata, including:\n \n-- The generation number of the commit. Commits with no parents have\n-  generation number 1; commits with parents have generation number\n-  one more than the maximum generation number of its parents. We\n-  reserve zero as special, and can be used to mark a generation\n-  number invalid or as \"not computed\".\n+- The generation number of the commit.\n \n - The root tree OID.\n \n@@ -88,6 +84,12 @@ CHUNK DATA:\n       2 bits of the lowest byte, storing the 33rd and 34th bit of the\n       commit time.\n \n+  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes) [Optional]\n+    * This list of 4-byte values store corrected commit date offsets for the\n+      commits, arranged in the same order as commit data chunk.\n+    * This list can be later modified to store future generation number related\n+      data.\n+\n   Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n       This list of 4-byte values store the second through nth parents for\n       all octopus merges. The second parent value in the commit data stores\ndiff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\nindex 808fa30b99..f27145328c 100644\n--- a/Documentation/technical/commit-graph.txt\n+++ b/Documentation/technical/commit-graph.txt\n@@ -38,14 +38,27 @@ A consumer may load the following info for a commit from the graph:\n \n Values 1-4 satisfy the requirements of parse_commit_gently().\n \n-Define the \"generation number\" of a commit recursively as follows:\n+There are two definitions of generation number:\n+1. Corrected committer dates\n+2. Topological levels\n+\n+Define \"corrected committer date\" of a commit recursively as follows:\n+\n+  * A commit with no parents (a root commit) has corrected committer date\n+    equal to its committer date.\n+\n+  * A commit with at least one parent has corrected committer date equal to\n+    the maximum of its commiter date and one more than the largest corrected\n+    committer date among its parents.\n+\n+Define the \"topological level\" of a commit recursively as follows:\n \n  * A commit with no parents (a root commit) has generation number one.\n \n- * A commit with at least one parent has generation number one more than\n-   the largest generation number among its parents.\n+ * A commit with at least one parent has topological level one more than\n+   the largest topological level among its parents.\n \n-Equivalently, the generation number of a commit A is one more than the\n+Equivalently, the topological level of a commit A is one more than the\n length of a longest path from A to a root commit. The recursive definition\n is easier to use for computation and observing the following property:\n \n@@ -67,17 +80,12 @@ numbers, the general heuristic is the following:\n     If A and B are commits with commit time X and Y, respectively, and\n     X < Y, then A _probably_ cannot reach B.\n \n-This heuristic is currently used whenever the computation is allowed to\n-violate topological relationships due to clock skew (such as \"git log\"\n-with default order), but is not used when the topological order is\n-required (such as merge base calculations, \"git log --graph\").\n-\n In practice, we expect some commits to be created recently and not stored\n in the commit graph. We can treat these commits as having \"infinite\"\n generation number and walk until reaching commits with known generation\n number.\n \n-We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n+We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n in the commit-graph file. If a commit-graph file was written by a version\n of Git that did not compute generation numbers, then those commits will\n have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n@@ -93,12 +101,11 @@ fully-computed generation numbers. Using strict inequality may result in\n walking a few extra commits, but the simplicity in dealing with commits\n with generation number *_INFINITY or *_ZERO is valuable.\n \n-We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n-generation numbers are computed to be at least this value. We limit at\n-this value since it is the largest value that can be stored in the\n-commit-graph file using the 30 bits available to generation numbers. This\n-presents another case where a commit can have generation number equal to\n-that of a parent.\n+We use the macro GENERATION_NUMBER_MAX for commits whose generation numbers\n+are computed to be at least this value. We limit at this value since it is\n+the largest value that can be stored in the commit-graph file using the\n+available to generation numbers. This presents another case where a\n+commit can have generation number equal to that of a parent.\n \n Design Details\n --------------\n@@ -267,6 +274,12 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n number of commits) could be extracted into config settings for full\n flexibility.\n \n+We also merge commit-graph chains when we try to write a commit graph with\n+two different generation number definitions as they cannot be compared directly.\n+We overwrite the existing chain and create a commit-graph with the newer or more\n+efficient defintion. For example, overwriting topological levels commit graph\n+chain to create a corrected commit dates commit graph chain.\n+\n ## Deleting graph-{hash} files\n \n After a new tip file is written, some `graph-{hash}` files may no longer\n-- \ngitgitgadget\n"},{"id":"403226","messageId":"b2547828585be24f779b7b0672a2984e6b883f3c.1596941624.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"[PATCH v2 04/10] commit-graph: consolidate compare_commits_by_gen","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-09T02:53:38Z","receivedAt":"2020-08-09T02:54:10Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nComparing commits by generation has been independently defined twice, in\ncommit-reach and commit. Let's simplify the implementation by moving\ncompare_commits_by_gen() to commit-graph.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\nReviewed-by: Taylor Blau <me@ttaylorr.com>\n---\n commit-graph.c | 15 +++++++++++++++\n commit-graph.h |  2 ++\n commit-reach.c | 15 ---------------\n commit.c       |  9 +++------\n 4 files changed, 20 insertions(+), 21 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex af8d9cc45e..fb6e2bf18f 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -112,6 +112,21 @@ uint32_t commit_graph_generation(const struct commit *c)\n \treturn data->generation;\n }\n \n+int compare_commits_by_gen(const void *_a, const void *_b)\n+{\n+\tconst struct commit *a = _a, *b = _b;\n+\tconst uint32_t generation_a = commit_graph_generation(a);\n+\tconst uint32_t generation_b = commit_graph_generation(b);\n+\n+\t/* older commits first */\n+\tif (generation_a < generation_b)\n+\t\treturn -1;\n+\telse if (generation_a > generation_b)\n+\t\treturn 1;\n+\n+\treturn 0;\n+}\n+\n static struct commit_graph_data *commit_graph_data_at(const struct commit *c)\n {\n \tunsigned int i, nth_slab;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 09a97030dc..701e3d41aa 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -146,4 +146,6 @@ struct commit_graph_data {\n  */\n uint32_t commit_graph_generation(const struct commit *);\n uint32_t commit_graph_position(const struct commit *);\n+\n+int compare_commits_by_gen(const void *_a, const void *_b);\n #endif\ndiff --git a/commit-reach.c b/commit-reach.c\nindex efd5925cbb..c83cc291e7 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -561,21 +561,6 @@ int commit_contains(struct ref_filter *filter, struct commit *commit,\n \treturn repo_is_descendant_of(the_repository, commit, list);\n }\n \n-static int compare_commits_by_gen(const void *_a, const void *_b)\n-{\n-\tconst struct commit *a = *(const struct commit * const *)_a;\n-\tconst struct commit *b = *(const struct commit * const *)_b;\n-\n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n-\n-\tif (generation_a < generation_b)\n-\t\treturn -1;\n-\tif (generation_a > generation_b)\n-\t\treturn 1;\n-\treturn 0;\n-}\n-\n int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\ndiff --git a/commit.c b/commit.c\nindex 7128895c3a..bed63b41fb 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -731,14 +731,11 @@ int compare_commits_by_author_date(const void *a_, const void *b_,\n int compare_commits_by_gen_then_commit_date(const void *a_, const void *b_, void *unused)\n {\n \tconst struct commit *a = a_, *b = b_;\n-\tconst uint32_t generation_a = commit_graph_generation(a),\n-\t\t       generation_b = commit_graph_generation(b);\n+\tint ret_val = compare_commits_by_gen(a_, b_);\n \n \t/* newer commits first */\n-\tif (generation_a < generation_b)\n-\t\treturn 1;\n-\telse if (generation_a > generation_b)\n-\t\treturn -1;\n+\tif (ret_val)\n+\t\treturn -ret_val;\n \n \t/* use date as a heuristic when generations are equal */\n \tif (a->date < b->date)\n-- \ngitgitgadget\n\n"},{"id":"403227","messageId":"cb797e20d79e9dcd3e0b953e0db3ed1defb9aa7c.1596941625.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"[PATCH v2 05/10] commit-graph: implement generation data chunk","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-09T02:53:39Z","receivedAt":"2020-08-09T02:54:10Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nAs discovered by Ævar, we cannot increment graph version to\ndistinguish between generation numbers v1 and v2 [1]. Thus, one of\npre-requistes before implementing generation number was to distinguish\nbetween graph versions in a backwards compatible manner.\n\nWe are going to introduce a new chunk called Generation Data chunk (or\nGDAT). GDAT stores generation number v2 (and any subsequent versions),\nwhereas CDAT will still store topological level.\n\nOld Git does not understand GDAT chunk and would ignore it, reading\ntopological levels from CDAT. New Git can parse GDAT and take advantage\nof newer generation numbers, falling back to topological levels when\nGDAT chunk is missing (as it would happen with a commit graph written\nby old Git).\n\nWe introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\nwhich forces commit-graph file to be written without generation data\nchunk to emulate a commit-graph file written by old Git.\n\n[1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                | 40 +++++++++++++++++++---\n commit-graph.h                |  2 ++\n t/README                      |  3 ++\n t/helper/test-read-graph.c    |  2 ++\n t/t4216-log-bloom.sh          |  4 +--\n t/t5318-commit-graph.sh       | 27 +++++++--------\n t/t5324-split-commit-graph.sh | 12 +++----\n t/t6600-test-reach.sh         | 62 +++++++++++++++++++----------------\n 8 files changed, 99 insertions(+), 53 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex fb6e2bf18f..d5da1e8028 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -38,11 +38,12 @@ void git_test_write_commit_graph_or_die(void)\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n+#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n-#define MAX_NUM_CHUNKS 7\n+#define MAX_NUM_CHUNKS 8\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -392,6 +393,13 @@ struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size)\n \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n \t\t\tbreak;\n \n+\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n+\t\t\tif (graph->chunk_generation_data)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n+\t\t\tbreak;\n+\n \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n \t\t\tif (graph->chunk_extra_edges)\n \t\t\t\tchunk_repeated = 1;\n@@ -758,7 +766,10 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\tif (g->chunk_generation_data)\n+\t\tgraph_data->generation = get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n+\telse\n+\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n static inline void set_commit_tree(struct commit *c, struct tree *t)\n@@ -951,7 +962,8 @@ struct write_commit_graph_context {\n \t\t report_progress:1,\n \t\t split:1,\n \t\t changed_paths:1,\n-\t\t order_by_pack:1;\n+\t\t order_by_pack:1,\n+\t\t write_generation_data:1;\n \n \tconst struct split_commit_graph_opts *split_opts;\n \tsize_t total_bloom_filter_data_size;\n@@ -1105,8 +1117,21 @@ static int write_graph_chunk_data(struct hashfile *f,\n \treturn 0;\n }\n \n+static int write_graph_chunk_generation_data(struct hashfile *f,\n+\t\t\t\t\t      struct write_commit_graph_context *ctx)\n+{\n+\tint i;\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n+\t\thashwrite_be32(f, commit_graph_data_at(c)->generation);\n+\t}\n+\n+\treturn 0;\n+}\n+\n static int write_graph_chunk_extra_edges(struct hashfile *f,\n-\t\t\t\t\t struct write_commit_graph_context *ctx)\n+\t\t\t\t\t  struct write_commit_graph_context *ctx)\n {\n \tstruct commit **list = ctx->commits.list;\n \tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n@@ -1710,6 +1735,12 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tchunks[2].id = GRAPH_CHUNKID_DATA;\n \tchunks[2].size = (hashsz + 16) * ctx->commits.nr;\n \tchunks[2].write_fn = write_graph_chunk_data;\n+\tif (ctx->write_generation_data) {\n+\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA;\n+\t\tchunks[num_chunks].size = sizeof(uint32_t) * ctx->commits.nr;\n+\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_extra_edges) {\n \t\tchunks[num_chunks].id = GRAPH_CHUNKID_EXTRAEDGES;\n \t\tchunks[num_chunks].size = 4 * ctx->num_extra_edges;\n@@ -2113,6 +2144,7 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n \tctx->split_opts = split_opts;\n \tctx->total_bloom_filter_data_size = 0;\n+\tctx->write_generation_data = !git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0);\n \n \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n \t\tctx->changed_paths = 1;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 701e3d41aa..cc232e0678 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -6,6 +6,7 @@\n #include \"oidset.h\"\n \n #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n+#define GIT_TEST_COMMIT_GRAPH_NO_GDAT \"GIT_TEST_COMMIT_GRAPH_NO_GDAT\"\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE \"GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE\"\n #define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n \n@@ -67,6 +68,7 @@ struct commit_graph {\n \tconst uint32_t *chunk_oid_fanout;\n \tconst unsigned char *chunk_oid_lookup;\n \tconst unsigned char *chunk_commit_data;\n+\tconst unsigned char *chunk_generation_data;\n \tconst unsigned char *chunk_extra_edges;\n \tconst unsigned char *chunk_base_graphs;\n \tconst unsigned char *chunk_bloom_indexes;\ndiff --git a/t/README b/t/README\nindex 70ec61cf88..6647ef132e 100644\n--- a/t/README\n+++ b/t/README\n@@ -379,6 +379,9 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n+commit-graph to be written without generation data chunk.\n+\n GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n commit-graph write to compute and write changed path Bloom filters for\n every 'git commit-graph write', as if the `--changed-paths` option was\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex 6d0c962438..1c2a5366c7 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -32,6 +32,8 @@ int cmd__read_graph(int argc, const char **argv)\n \t\tprintf(\" oid_lookup\");\n \tif (graph->chunk_commit_data)\n \t\tprintf(\" commit_metadata\");\n+\tif (graph->chunk_generation_data)\n+\t\tprintf(\" generation_data\");\n \tif (graph->chunk_extra_edges)\n \t\tprintf(\" extra_edges\");\n \tif (graph->chunk_bloom_indexes)\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex c21cc160f3..55c94e9ebd 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -33,11 +33,11 @@ test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n \tgit commit-graph write --reachable --changed-paths\n '\n graph_read_expect () {\n-\tNUM_CHUNKS=5\n+\tNUM_CHUNKS=6\n \tcat >expect <<- EOF\n \theader: 43475048 1 1 $NUM_CHUNKS 0\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n \tEOF\n \ttest-tool read-graph >actual &&\n \ttest_cmp expect actual\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 2804b0dd45..fef05c33d7 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -72,7 +72,7 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n graph_read_expect() {\n \tOPTIONAL=\"\"\n \tNUM_CHUNKS=3\n-\tif test ! -z $2\n+\tif test ! -z \"$2\"\n \tthen\n \t\tOPTIONAL=\" $2\"\n \t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n@@ -99,14 +99,14 @@ test_expect_success 'exit with correct error on bad input to --stdin-commits' '\n \t# valid commit and tree OID\n \tgit rev-parse HEAD HEAD^{tree} >in &&\n \tgit commit-graph write --stdin-commits <in &&\n-\tgraph_read_expect 3\n+\tgraph_read_expect 3 generation_data\n '\n \n test_expect_success 'write graph' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"3\"\n+\tgraph_read_expect \"3\" generation_data\n '\n \n test_expect_success POSIXPERM 'write graph has correct permissions' '\n@@ -215,7 +215,7 @@ test_expect_success 'write graph with merges' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'merge 1 vs 2' full merge/1 merge/2\n@@ -250,7 +250,7 @@ test_expect_success 'write graph with new commit' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'full graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -260,7 +260,7 @@ test_expect_success 'write graph with nothing new' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'cleared graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -270,7 +270,7 @@ test_expect_success 'build graph from latest pack with closure' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tcat new-idx | git commit-graph write --stdin-packs &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"9\" \"extra_edges\"\n+\tgraph_read_expect \"9\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'graph from pack, commit 8 vs merge 1' full commits/8 merge/1\n@@ -283,7 +283,7 @@ test_expect_success 'build graph from commits with closure' '\n \tgit rev-parse merge/1 >>commits-in &&\n \tcat commits-in | git commit-graph write --stdin-commits &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"6\"\n+\tgraph_read_expect \"6\" \"generation_data\"\n '\n \n graph_git_behavior 'graph from commits, commit 8 vs merge 1' full commits/8 merge/1\n@@ -293,7 +293,7 @@ test_expect_success 'build graph from commits with append' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit rev-parse merge/3 | git commit-graph write --stdin-commits --append &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -303,7 +303,7 @@ test_expect_success 'build graph using --reachable' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write --reachable &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -324,7 +324,7 @@ test_expect_success 'write graph in bare repo' '\n \tcd \"$TRASH_DIRECTORY/bare\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $baredir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n@@ -421,8 +421,9 @@ test_expect_success 'replace-objects invalidates commit-graph' '\n \n test_expect_success 'git commit-graph verify' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n-\tgit rev-parse commits/8 | git commit-graph write --stdin-commits &&\n-\tgit commit-graph verify >output\n+\tgit rev-parse commits/8 | GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --stdin-commits &&\n+\tgit commit-graph verify >output &&\n+\tgraph_read_expect 9 extra_edges\n '\n \n NUM_COMMITS=9\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 9b850ea907..6b25c3d9ce 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -14,11 +14,11 @@ test_expect_success 'setup repo' '\n \tgraphdir=\"$infodir/commit-graphs\" &&\n \ttest_oid_init &&\n \ttest_oid_cache <<-EOM\n-\tshallow sha1:1760\n-\tshallow sha256:2064\n+\tshallow sha1:2132\n+\tshallow sha256:2436\n \n-\tbase sha1:1376\n-\tbase sha256:1496\n+\tbase sha1:1408\n+\tbase sha256:1528\n \tEOM\n '\n \n@@ -29,9 +29,9 @@ graph_read_expect() {\n \t\tNUM_BASE=$2\n \tfi\n \tcat >expect <<- EOF\n-\theader: 43475048 1 1 3 $NUM_BASE\n+\theader: 43475048 1 1 4 $NUM_BASE\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n \tEOF\n \ttest-tool read-graph >output &&\n \ttest_cmp expect output\ndiff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\nindex 475564bee7..d14b129f06 100755\n--- a/t/t6600-test-reach.sh\n+++ b/t/t6600-test-reach.sh\n@@ -55,10 +55,13 @@ test_expect_success 'setup' '\n \tgit show-ref -s commit-5-5 | git commit-graph write --stdin-commits &&\n \tmv .git/objects/info/commit-graph commit-graph-half &&\n \tchmod u+w commit-graph-half &&\n+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable &&\n+\tmv .git/objects/info/commit-graph commit-graph-no-gdat &&\n+\tchmod u+w commit-graph-no-gdat &&\n \tgit config core.commitGraph true\n '\n \n-run_three_modes () {\n+run_all_modes () {\n \ttest_when_finished rm -rf .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n \ttest_cmp expect actual &&\n@@ -67,11 +70,14 @@ run_three_modes () {\n \ttest_cmp expect actual &&\n \tcp commit-graph-half .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n+\ttest_cmp expect actual &&\n+\tcp commit-graph-no-gdat .git/objects/info/commit-graph &&\n+\t\"$@\" <input >actual &&\n \ttest_cmp expect actual\n }\n \n-test_three_modes () {\n-\trun_three_modes test-tool reach \"$@\"\n+test_all_modes () {\n+\trun_all_modes test-tool reach \"$@\"\n }\n \n test_expect_success 'ref_newer:miss' '\n@@ -80,7 +86,7 @@ test_expect_success 'ref_newer:miss' '\n \tB:commit-4-9\n \tEOF\n \techo \"ref_newer(A,B):0\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'ref_newer:hit' '\n@@ -89,7 +95,7 @@ test_expect_success 'ref_newer:hit' '\n \tB:commit-2-3\n \tEOF\n \techo \"ref_newer(A,B):1\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'in_merge_bases:hit' '\n@@ -98,7 +104,7 @@ test_expect_success 'in_merge_bases:hit' '\n \tB:commit-8-8\n \tEOF\n \techo \"in_merge_bases(A,B):1\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'in_merge_bases:miss' '\n@@ -107,7 +113,7 @@ test_expect_success 'in_merge_bases:miss' '\n \tB:commit-5-9\n \tEOF\n \techo \"in_merge_bases(A,B):0\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'is_descendant_of:hit' '\n@@ -118,7 +124,7 @@ test_expect_success 'is_descendant_of:hit' '\n \tX:commit-1-1\n \tEOF\n \techo \"is_descendant_of(A,X):1\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'is_descendant_of:miss' '\n@@ -129,7 +135,7 @@ test_expect_success 'is_descendant_of:miss' '\n \tX:commit-7-6\n \tEOF\n \techo \"is_descendant_of(A,X):0\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'get_merge_bases_many' '\n@@ -144,7 +150,7 @@ test_expect_success 'get_merge_bases_many' '\n \t\tgit rev-parse commit-5-6 \\\n \t\t\t      commit-4-7 | sort\n \t} >expect &&\n-\ttest_three_modes get_merge_bases_many\n+\ttest_all_modes get_merge_bases_many\n '\n \n test_expect_success 'reduce_heads' '\n@@ -166,7 +172,7 @@ test_expect_success 'reduce_heads' '\n \t\t\t      commit-2-8 \\\n \t\t\t      commit-1-10 | sort\n \t} >expect &&\n-\ttest_three_modes reduce_heads\n+\ttest_all_modes reduce_heads\n '\n \n test_expect_success 'can_all_from_reach:hit' '\n@@ -189,7 +195,7 @@ test_expect_success 'can_all_from_reach:hit' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach(X,Y):1\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach:miss' '\n@@ -211,7 +217,7 @@ test_expect_success 'can_all_from_reach:miss' '\n \tY:commit-8-5\n \tEOF\n \techo \"can_all_from_reach(X,Y):0\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach_with_flag: tags case' '\n@@ -234,7 +240,7 @@ test_expect_success 'can_all_from_reach_with_flag: tags case' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach_with_flag(X,_,_,0,0):1\" >expect &&\n-\ttest_three_modes can_all_from_reach_with_flag\n+\ttest_all_modes can_all_from_reach_with_flag\n '\n \n test_expect_success 'commit_contains:hit' '\n@@ -250,8 +256,8 @@ test_expect_success 'commit_contains:hit' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):1\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'commit_contains:miss' '\n@@ -267,8 +273,8 @@ test_expect_success 'commit_contains:miss' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):0\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'rev-list: basic topo-order' '\n@@ -280,7 +286,7 @@ test_expect_success 'rev-list: basic topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 commit-3-2 commit-2-2 commit-1-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent topo-order' '\n@@ -292,7 +298,7 @@ test_expect_success 'rev-list: first-parent topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -304,7 +310,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -316,7 +322,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent range topo-order' '\n@@ -328,7 +334,7 @@ test_expect_success 'rev-list: first-parent range topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: ancestry-path topo-order' '\n@@ -338,7 +344,7 @@ test_expect_success 'rev-list: ancestry-path topo-order' '\n \t\tcommit-6-4 commit-5-4 commit-4-4 commit-3-4 \\\n \t\tcommit-6-3 commit-5-3 commit-4-3 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: symmetric difference topo-order' '\n@@ -352,7 +358,7 @@ test_expect_success 'rev-list: symmetric difference topo-order' '\n \t\tcommit-3-8 commit-2-8 commit-1-8 \\\n \t\tcommit-3-7 commit-2-7 commit-1-7 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8...commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8...commit-6-6\n '\n \n test_expect_success 'get_reachable_subset:all' '\n@@ -372,7 +378,7 @@ test_expect_success 'get_reachable_subset:all' '\n \t\t\t      commit-1-7 \\\n \t\t\t      commit-5-6 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:some' '\n@@ -390,7 +396,7 @@ test_expect_success 'get_reachable_subset:some' '\n \t\tgit rev-parse commit-3-3 \\\n \t\t\t      commit-1-7 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:none' '\n@@ -404,7 +410,7 @@ test_expect_success 'get_reachable_subset:none' '\n \tY:commit-2-8\n \tEOF\n \techo \"get_reachable_subset(X,Y)\" >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_done\n-- \ngitgitgadget\n\n"},{"id":"403263","messageId":"a3910f82-ab2e-bf35-ac43-c30d77f3c96b@gmail.com","threadId":"53933","inReplyTo":"bfe14732014807ff19f943cdf51068f0d3043c30.1596941625.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 07/10] commit-graph: implement corrected commit date","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-10T14:23:32Z","receivedAt":"2020-08-10T14:23:36Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/8/2020 10:53 PM, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> With most of preparations done, let's implement corrected commit date\n> offset. We add a new commit-slab to store topogical levels while\n> writing commit graph and upgrade the generation member in struct\n> commit_graph_data to a 64-bit timestamp. We store topological levels to\n> ensure that older versions of Git will still have the performance\n> benefits from generation number v2.\n> \n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n\n> @@ -767,7 +764,10 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n>  \n>  \tif (g->chunk_generation_data)\n> -\t\tgraph_data->generation = get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n> +\t{\n> +\t\tgraph_data->generation = item->date +\n> +\t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n> +\t}\n\nYou don't need curly braces here, since this is only one line\nin the block. Even if you did, these braces are in the wrong\nlocation.\n\nThere is a subtle issue with this interpretation, and it\ninvolves the case where the following happens:\n\n1. A new version of Git writes a commit-graph using the\n   GDAT chunk.\n\n2. An older version of Git adds a new layer without the\n   GDAT chunk.\n\nAt that point, the tip commit-graph does not have GDAT,\nso the commits in that layer will get \"generation\" set\nwith the topological level, which is likely to be much\nlower than the corrected commit dates set in the\n\"generation\" field for commits in the lower layer.\n\nThe crux of the issue is that we are only considering\nthe current layer when interpreting the generation number\nvalue.\n\nThe patch below inserts a flag into fill_commit_graph_info()\ncorresponding to the \"global\" state of whether the top\ncommit-graph layer has a GDAT chunk. By your later protection\nto not write GDAT chunks on top of commit-graphs without\na GDAT chunk, this top commit-graph has all of the information\nwe need for this check.\n\nThanks,\n-Stolee\n\n--- >8 ---\n\nFrom 62189709fad3b051cedbd36193f5244fcce17e1f Mon Sep 17 00:00:00 2001\nFrom: Derrick Stolee <dstolee@microsoft.com>\nDate: Mon, 10 Aug 2020 10:06:47 -0400\nSubject: [PATCH] commit-graph: use generation v2 only if entire chain does\n\nSince there are released versions of Git that understand generation\nnumbers in the commit-graph's CDAT chunk but do not understand the GDAT\nchunk, the following scenario is possible:\n\n 1. \"New\" Git writes a commit-graph with the GDAT chunk.\n 2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n\nBecause of the current use of inspecting the current layer for a\ngeneration_data_chunk pointer, the commits in the lower layer will be\ninterpreted as having very large generation values (commit date plus\noffset) compared to the generation numbers in the top layer (topological\nlevel). This violates the expectation that the generation of a parent is\nstrictly smaller than the generation of a child.\n\nIt is difficult to expose this issue in a test. Since we _start_ with\nartificially low generation numbers, any commit walk that prioritizes\ngeneration numbers will walk all of the commits with high generation\nnumber before walking the commits with low generation number. In all the\ncases I tried, the commit-graph layers themselves \"protect\" any\nincorrect behavior since none of the commits in the lower layer can\nreach the commits in the upper layer.\n\nThis issue would manifest itself as a performance problem in this\ncase, especially with something like \"git log --graph\" since the low\ngeneration numbers would cause the in-degree queue to walk all of the\ncommits in the lower layer before allowing the topo-order queue to write\nanything to output (depending on the size of the upper layer).\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\n---\n commit-graph.c | 24 ++++++++++++++++++------\n 1 file changed, 18 insertions(+), 6 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex eb78af3dad..17623274d9 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -762,7 +762,9 @@ static struct commit_list **insert_parent_or_die(struct repository *r,\n \treturn &commit_list_insert(c, pptr)->next;\n }\n \n-static void fill_commit_graph_info(struct commit *item, struct commit_graph *g, uint32_t pos)\n+#define COMMIT_GRAPH_GENERATION_V2 (1 << 0)\n+\n+static void fill_commit_graph_info(struct commit *item, struct commit_graph *g, uint32_t pos, int flags)\n {\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n@@ -785,11 +787,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tif (g->chunk_generation_data)\n-\t{\n+\tif (g->chunk_generation_data && (flags & COMMIT_GRAPH_GENERATION_V2))\n \t\tgraph_data->generation = item->date +\n \t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n-\t}\n \telse\n \t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n@@ -799,6 +799,10 @@ static inline void set_commit_tree(struct commit *c, struct tree *t)\n \tc->maybe_tree = t;\n }\n \n+/*\n+ * In the case of a split commit-graph, this method expects the given\n+ * commit-graph 'g' to be the top layer.\n+ */\n static int fill_commit_in_graph(struct repository *r,\n \t\t\t\tstruct commit *item,\n \t\t\t\tstruct commit_graph *g, uint32_t pos)\n@@ -808,11 +812,15 @@ static int fill_commit_in_graph(struct repository *r,\n \tstruct commit_list **pptr;\n \tconst unsigned char *commit_data;\n \tuint32_t lex_index;\n+\tint flags = 0;\n+\n+\tif (g->chunk_generation_data)\n+\t\tflags |= COMMIT_GRAPH_GENERATION_V2;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n-\tfill_commit_graph_info(item, g, pos);\n+\tfill_commit_graph_info(item, g, pos, flags);\n \n \tlex_index = pos - g->num_commits_in_base;\n \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n@@ -904,10 +912,14 @@ int parse_commit_in_graph(struct repository *r, struct commit *item)\n void load_commit_graph_info(struct repository *r, struct commit *item)\n {\n \tuint32_t pos;\n+\tint flags = 0;\n+\n \tif (!prepare_commit_graph(r))\n \t\treturn;\n+\tif (r->objects->commit_graph->chunk_generation_data)\n+\t\tflags |= COMMIT_GRAPH_GENERATION_V2;\n \tif (find_commit_in_graph(item, r->objects->commit_graph, &pos))\n-\t\tfill_commit_graph_info(item, r->objects->commit_graph, pos);\n+\t\tfill_commit_graph_info(item, r->objects->commit_graph, pos, flags);\n }\n \n static struct tree *load_tree_for_commit(struct repository *r,\n-- \n2.28.0.38.gc6f546511c1\n\n"},{"id":"403275","messageId":"aee0ae56-3395-6848-d573-27a318d72755@gmail.com","threadId":"53933","inReplyTo":"cb797e20d79e9dcd3e0b953e0db3ed1defb9aa7c.1596941625.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 05/10] commit-graph: implement generation data chunk","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-10T16:28:10Z","receivedAt":"2020-08-10T16:28:14Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/8/2020 10:53 PM, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> As discovered by Ævar, we cannot increment graph version to\n> distinguish between generation numbers v1 and v2 [1]. Thus, one of\n> pre-requistes before implementing generation number was to distinguish\n> between graph versions in a backwards compatible manner.\n> \n> We are going to introduce a new chunk called Generation Data chunk (or\n> GDAT). GDAT stores generation number v2 (and any subsequent versions),\n> whereas CDAT will still store topological level.\n> \n> Old Git does not understand GDAT chunk and would ignore it, reading\n> topological levels from CDAT. New Git can parse GDAT and take advantage\n> of newer generation numbers, falling back to topological levels when\n> GDAT chunk is missing (as it would happen with a commit graph written\n> by old Git).\n\nThere is a philosophical problem with this patch, and I'm not sure\nabout the right way to fix it, or if there really is a problem at all.\nAt minimum, the commit message needs to be improved to make the issue\nclear:\n\nThis version of the chunk does not store corrected commit date offsets!\n\nThis commit add a chunk named \"GDAT\" and fills it with topological\nlevels. This is _different_ than the intended final format. For that\nreason, the commit-graph-format.txt document is not updated.\n\nThe reason I say this is a \"philosophical\" problem is that this patch\nintroduces a version of Git that has a different interpretation of the\nGDAT chunk than the version presented two patches later. While this\nversion would never be released, it still exists in history and could\npresent difficulty if someone were to bisect on an issue with the GDAT\nchunk (using external data, not data produced by the compiled binary\nat that version).\n\nThe justification for this commit the way you did it is clear: there\nis a lot of test fallout to just including a new chunk. The question\nis whether it is enough to justify this \"dummy\" implementation for\nnow?\n\nThe tricky bit is the series of three patches starting with this\none.\n\n1. The next patch \"commit-graph: return 64-bit generation number\" can\n   be reordered to be before this patch, no problem. I don't think\n   there will be any text conflicts _except_ inside the\n   write_graph_chunk_generation_data() method introduced here.\n\n2. The patch after that, \"commit-graph: implement corrected commit date\"\n   only has a small dependence: it writes to the GDAT chunk and parses\n   it out. If you remove the interaction with the GDAT chunk, then you\n   still have the computation as part of compute_generation_numbers()\n   that is valuable. You will need to be careful about the exit\n   condition, though, since you also introduce the topo_level chunk.\n\nPatches 5-7 could perhaps be reorganized as follows:\n\n  i. commit-graph: return 64-bit generation number, as-is.\n\n ii. Add a topo_level slab that is parsed from CDAT. Modify\n     compute_generation_numbers() to populate this value and modify\n     write_graph_chunk_data() to read this value. Simultaneously\n     populate the \"generation\" member with the same value.\n\niii. \"commit-graph: implement corrected commit date\" without any GDAT\n     chunk interaction. Make sure the algorithm in\n     compute_generation_numbers() walks commits if either topo_level or\n     generation are unset. There is a trick here: the generation value\n     _is_ set if the commit is parsed from the existing commit-graph!\n     Is this case covered by the existing logic to not write GDAT when\n     writing a split commit-graph file with a base that does not have\n     GDAT? Note that the non-split case does not load the commit-graph\n     for parsing, so the interesting case is \"--split-replace\". Worth\n     a test (after we write the GDAT chunk), which you have in \"commit-graph:\n     handle mixed generation commit chains\".\n\n iv. This patch, introducing the chunk and the read/write logic.\n\n  v. Add the remaining patches.\n\nAgain, this is a complicated patch-reorganization. The hope is that\nthe end result is something that is easy to review as well as something\nthat produces an as-sane-as-possible history for future bisecters.\n\nPerhaps other reviewers have similar feelings, or can say that I am\nbeing too picky.\n\n> We introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\n> which forces commit-graph file to be written without generation data\n> chunk to emulate a commit-graph file written by old Git.\n\nThank you for introducing this. It really makes it clear what the\nbenefit is when looking at the t6600-test-reach.sh changes. However,\nthe changes to that script are more \"here is an opportunity for extra\ncoverage\" as opposed to a necessary change immediately upon creating\nthe GDAT chunk. That could be separated out and justified on its own.\nRecall that the justification is that the new version of Git will\ncontinue to work with commit-graph files without a GDAT chunk.\n\n> +static int write_graph_chunk_generation_data(struct hashfile *f,\n> +\t\t\t\t\t      struct write_commit_graph_context *ctx)\n> +{\n> +\tint i;\n> +\tfor (i = 0; i < ctx->commits.nr; i++) {\n> +\t\tstruct commit *c = ctx->commits.list[i];\n> +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n> +\t\thashwrite_be32(f, commit_graph_data_at(c)->generation);\n\nHere is the \"incorrect\" data being written.\n\n> +\t}\n> +\n> +\treturn 0;\n> +}\n> +\n\n> --- a/t/t5318-commit-graph.sh\n> +++ b/t/t5318-commit-graph.sh\n> @@ -72,7 +72,7 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n>  graph_read_expect() {\n>  \tOPTIONAL=\"\"\n>  \tNUM_CHUNKS=3\n> -\tif test ! -z $2\n> +\tif test ! -z \"$2\"\n\nA subtle change, but important because we now have multiple \"extra\"\nchunks possible here. Good.\n\n>  graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n> @@ -421,8 +421,9 @@ test_expect_success 'replace-objects invalidates commit-graph' '\n>  \n>  test_expect_success 'git commit-graph verify' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n> -\tgit rev-parse commits/8 | git commit-graph write --stdin-commits &&\n> -\tgit commit-graph verify >output\n> +\tgit rev-parse commits/8 | GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --stdin-commits &&\n> +\tgit commit-graph verify >output &&\n> +\tgraph_read_expect 9 extra_edges\n>  '\n\nAnd it is this case as to why we don't just add \"generation_data\" to our\nlist of expected chunks.\n\n> @@ -29,9 +29,9 @@ graph_read_expect() {\n>  \t\tNUM_BASE=$2\n>  \tfi\n>  \tcat >expect <<- EOF\n> -\theader: 43475048 1 1 3 $NUM_BASE\n> +\theader: 43475048 1 1 4 $NUM_BASE\n>  \tnum_commits: $1\n> -\tchunks: oid_fanout oid_lookup commit_metadata\n> +\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n\nIn this script, you _do_ add it to the default chunk list, which\nsaves some extra work in the rest of the tests. Good.\n\n\nThanks,\n-Stolee\n"},{"id":"403277","messageId":"0d741fb2-e25a-be05-9f2b-81ba2b4ced3f@gmail.com","threadId":"53933","inReplyTo":"833779ad53eb4f57ae514f4e8964e397845f1ddd.1596941625.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 08/10] commit-graph: handle mixed generation commit chains","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-10T16:42:29Z","receivedAt":"2020-08-10T16:42:34Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/8/2020 10:53 PM, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> As corrected commit dates and topological levels cannot be compared\n> directly, we must handle commit graph chains with mixed generation\n> number definitions.\n> \n> While reading a commit graph file, we disable generation numbers if the\n> chain contains mixed generation numbers.\n> \n> While writing to commit graph chain, we write generation data chunk only\n> if the previous tip of chain had a generation data chunk. Using\n> `--split=replace` overwrites the existing chain and writes generation\n> data chunk regardless of previous tip.\n> \n> In t5324-split-commit-graph, we set up a repo with twelve commits and\n> write a base commit graph file with no generation data chunk. When add\n> three commits and write to chain again, Git does not write generation\n> data chunk even without setting GIT_TEST_COMMIT_GRAPH_NO_GDAT=1. Then,\n> as we replace the existing chain, Git writes a commit graph file with\n> generation data chunk.\n> \n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c                | 14 ++++++++\n>  t/t5324-split-commit-graph.sh | 66 +++++++++++++++++++++++++++++++++++\n>  2 files changed, 80 insertions(+)\n> \n> diff --git a/commit-graph.c b/commit-graph.c\n> index d0f977852b..c6b6111adf 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -674,6 +674,14 @@ int generation_numbers_enabled(struct repository *r)\n>  \tif (!g->num_commits)\n>  \t\treturn 0;\n>  \n> +\t/* We cannot compare topological levels and corrected commit dates */\n> +\twhile (g->base_graph) {\n> +\t\twarning(_(\"commit-graph-chain contains mixed generation versions\"));\n\nThis warning is premature. It will add a warning whenever we have\na split commit-graph, regardless of an incorrect chain.\n\n> +\t\tif ((g->chunk_generation_data == NULL) ^ (g->base_graph->chunk_generation_data == NULL))\n\nHm. A bit-wise XOR here? That seems unfortunate. I think that it\nis easier to focus on the \n\n\n\n> +\t\t\treturn 0;\n> +\t\tg = g->base_graph;\n> +\t}\n> +\n\nHm. So this scenario actually disables generation numbers completely\nin the event that anything in the chain disagrees. I think this is\nnot the right way to approach the situation, as it will significantly\npunish users in this state with slow performance.\n\nThe patch I sent [1] is probably better: it uses generation number\nv1 if the tip of the chain does not have a GDAT chunk.\n\n[1] https://lore.kernel.org/git/a3910f82-ab2e-bf35-ac43-c30d77f3c96b@gmail.com/\n\n>  \tfirst_generation = get_be32(g->chunk_commit_data +\n>  \t\t\t\t    g->hash_len + 8) >> 2;\n>  \n> @@ -2186,6 +2194,9 @@ int write_commit_graph(struct object_directory *odb,\n>  \n>  \t\tg = ctx->r->objects->commit_graph;\n>  \n> +\t\tif (g && !g->chunk_generation_data)\n> +\t\t\tctx->write_generation_data = 0;\n> +\n>  \t\twhile (g) {\n>  \t\t\tctx->num_commit_graphs_before++;\n>  \t\t\tg = g->base_graph;\n> @@ -2204,6 +2215,9 @@ int write_commit_graph(struct object_directory *odb,\n>  \n>  \t\tif (ctx->split_opts)\n>  \t\t\treplace = ctx->split_opts->flags & COMMIT_GRAPH_SPLIT_REPLACE;\n> +\n> +\t\tif (replace)\n> +\t\t\tctx->write_generation_data = 1;\n>  \t}\n\nPlease make a point to move the line that checks GIT_TEST_COMMIT_GRAPH_NO_GDAT\nfrom its current location to after this line. We want to make sure that the\nenvironment variable is checked _last_. The best location is likely the start\nof the implementation of compute_generation_numbers(), or immediately before\nthe call to the method.\n\n> +test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n> +\tmkdir mixed &&\n> +\tgraphdir=\".git/objects/info/commit-graphs\" &&\n> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> +\tgit init &&\n> +\tgit config core.commitGraph true &&\n> +\tgit config gc.writeCommitGraph false &&\n> +\tfor i in $(test_seq 3)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit reset --hard commits/1 &&\n> +\tfor i in $(test_seq 4 5)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit reset --hard commits/2 &&\n> +\tfor i in $(test_seq 6 10)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit reset --hard commits/2 &&\n> +\tgit merge commits/4 &&\n> +\tgit branch merge/1 &&\n> +\tgit reset --hard commits/4 &&\n> +\tgit merge commits/6 &&\n> +\tgit branch merge/2 &&\n> +\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split &&\n> +\ttest-tool read-graph >output &&\n> +\tcat >expect <<-EOF &&\n> +\theader: 43475048 1 1 3 0\n> +\tnum_commits: 12\n> +\tchunks: oid_fanout oid_lookup commit_metadata\n> +\tEOF\n> +\ttest_cmp expect output\n> +'\n> +\n> +test_expect_success 'does not write generation data chunk if not present on existing tip' '\n> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> +\tgit reset --hard commits/3 &&\n> +\tgit merge merge/1 &&\n> +\tgit merge commits/5 &&\n> +\tgit merge merge/2 &&\n> +\tgit branch merge/3 &&\n> +\tgit commit-graph write --reachable --split &&\n> +\ttest-tool read-graph >output &&\n> +\tcat >expect <<-EOF &&\n> +\theader: 43475048 1 1 4 1\n> +\tnum_commits: 3\n> +\tchunks: oid_fanout oid_lookup commit_metadata\n> +\tEOF\n> +\ttest_cmp expect output\n> +'\n> +\n> +test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> +\tgit commit-graph write --reachable --split='replace' &&\n> +\ttest_path_is_file $graphdir/commit-graph-chain &&\n> +\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n> +\tverify_chain_files_exist $graphdir &&\n> +\tgraph_read_expect 15\n> +'\n\nIt would be valuable to double-check here that the values in the GDAT chunk\nare correct. I'm concerned about the possibility that the 'generation'\nmember of struct commit_graph_data gets filled with topological level during\nparsing and then that is written as an offset into the CDAT chunk.\n\nPerhaps this is best left for a follow-up series that updates the 'verify'\nsubcommand to check the GDAT chunk.\n\nThanks,\n-Stolee\n\n"},{"id":"403278","messageId":"bf9ab386-39a0-e386-3828-4427fab3ee25@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 00/10] [GSoC] Implement Corrected Commit Date","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-10T16:47:26Z","receivedAt":"2020-08-10T16:47:31Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/8/2020 10:53 PM, Abhishek Kumar via GitGitGadget wrote:\n> This patch series implements the corrected commit date offsets as generation\n> number v2, along with other pre-requisites.\n> \n> Git uses topological levels in the commit-graph file for commit-graph\n> traversal operations like git log --graph. Unfortunately, using topological\n> levels can result in a worse performance than without them when compared\n> with committer date as a heuristics. For example, git merge-base v4.8 v4.9 \n> on the Linux repository walks 635,579 commits using topological levels and\n> walks 167,468 using committer date.\n> \n> Thus, the need for generation number v2 was born. New generation number\n> needed to provide good performance, increment updates, and backward\n> compatibility. Due to an unfortunate problem, we also needed a way to\n> distinguish between the old and new generation number without incrementing\n> graph version.\n> \n> Various candidates were examined (https://github.com/derrickstolee/gen-test, \n> https://github.com/abhishekkumar2718/git/pull/1). The proposed generation\n> number v2, Corrected Commit Date with Mononotically Increasing Offsets \n> performed much worse than committer date (506,577 vs. 167,468 commits walked\n> for git merge-base v4.8 v4.9) and was dropped.\n> \n> Using Generation Data chunk (GDAT) relieves the requirement of backward\n> compatibility as we would continue to store topological levels in Commit\n> Data (CDAT) chunk. Thus, Corrected Commit Date was chosen as generation\n> number v2. The Corrected Commit Date is defined as:\n> \n> For a commit C, let its corrected commit date be the maximum of the commit\n> date of C and the corrected commit dates of its parents. Then corrected\n> commit date offset is the difference between corrected commit date of C and\n> commit date of C.\n> \n> We will introduce an additional commit-graph chunk, Generation Data chunk,\n> and store corrected commit date offsets in GDAT chunk while storing\n> topological levels in CDAT chunk. The old versions of Git would ignore GDAT\n> chunk, using topological levels from CDAT chunk. In contrast, new versions\n> of Git would use corrected commit dates, falling back to topological level\n> if the generation data chunk is absent in the commit-graph file.\n> \n> Thanks to Dr. Stolee, Dr. Narębski, and Taylor for their reviews on the\n> first version.\n> \n> I look forward to everyone's reviews!\n> \n> Thanks\n> \n>  * Abhishek\n> \n> \n> ----------------------------------------------------------------------------\n> \n> Changes in version 2:\n> \n>  * Add tests for generation data chunk.\n>  * Add an option GIT_TEST_COMMIT_GRAPH_NO_GDAT to control whether to write\n>    generation data chunk.\n>  * Compare commits with corrected commit dates if present in\n>    paint_down_to_common().\n>  * Update technical documentation.\n>  * Handle mixed graph version commit chains.\n>  * Improve commit messages for\n>  * Revert unnecessary whitespace changes.\n>  * Split uint_32 -> timestamp_t change into a new commit.\n\nThis version looks to be in really good shape, for the most part. I feel\nthe need to point out that this is a very mature submission, and it covers\nmany of the bases we expect from a quality series, even at only the second\nversion.\n\nMy comments deal with some high-level \"patch series strategy\" considerations,\nmostly. There are a few concerns about correctness in the complicated cases\nof mixed commit-graphs with or without the GDAT chunk, and how we could\nperhaps test those cases. The best way would be to update the 'verify'\nsubcommand to check the GDAT chunk. That's a good feature to have, but would\nbe unnecessary bloat to this series. Depending on timing, we might want to\nhold this series in 'next' until such an implementation is submitted. (I\nexpect that Abhishek's GSoC internship would be over at that point, so I\nvolunteer to send those patches after this series stabilizes.)\n\nThank you for your careful work on such a complicated subject!\n\nThanks,\n-Stolee\n\n"},{"id":"403345","messageId":"20200811110316.GA3220@Abhishek-Arch","threadId":"53933","inReplyTo":"aee0ae56-3395-6848-d573-27a318d72755@gmail.com","subject":"Re: [PATCH v2 05/10] commit-graph: implement generation data chunk","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-08-11T11:03:16Z","receivedAt":"2020-08-11T11:05:32Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Mon, Aug 10, 2020 at 12:28:10PM -0400, Derrick Stolee wrote:\n> On 8/8/2020 10:53 PM, Abhishek Kumar via GitGitGadget wrote:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > \n> > As discovered by Ævar, we cannot increment graph version to\n> > distinguish between generation numbers v1 and v2 [1]. Thus, one of\n> > pre-requistes before implementing generation number was to distinguish\n> > between graph versions in a backwards compatible manner.\n> > \n> > We are going to introduce a new chunk called Generation Data chunk (or\n> > GDAT). GDAT stores generation number v2 (and any subsequent versions),\n> > whereas CDAT will still store topological level.\n> > \n> > Old Git does not understand GDAT chunk and would ignore it, reading\n> > topological levels from CDAT. New Git can parse GDAT and take advantage\n> > of newer generation numbers, falling back to topological levels when\n> > GDAT chunk is missing (as it would happen with a commit graph written\n> > by old Git).\n> \n> There is a philosophical problem with this patch, and I'm not sure\n> about the right way to fix it, or if there really is a problem at all.\n> At minimum, the commit message needs to be improved to make the issue\n> clear:\n> \n> This version of the chunk does not store corrected commit date offsets!\n> \n> This commit add a chunk named \"GDAT\" and fills it with topological\n> levels. This is _different_ than the intended final format. For that\n> reason, the commit-graph-format.txt document is not updated.\n> \n> The reason I say this is a \"philosophical\" problem is that this patch\n> introduces a version of Git that has a different interpretation of the\n> GDAT chunk than the version presented two patches later. While this\n> version would never be released, it still exists in history and could\n> present difficulty if someone were to bisect on an issue with the GDAT\n> chunk (using external data, not data produced by the compiled binary\n> at that version).\n> \n\nYes, that is correct. I did often wonder that our inference that \"commit\ngraph has a generation data chunk implies commit graph stores corrected\ncommit date offsets\" is not always true because of this \"dummy\"\nimplementation. \n\n> The justification for this commit the way you did it is clear: there\n> is a lot of test fallout to just including a new chunk. The question\n> is whether it is enough to justify this \"dummy\" implementation for\n> now?\n> \n> The tricky bit is the series of three patches starting with this\n> one.\n> \n> 1. The next patch \"commit-graph: return 64-bit generation number\" can\n>    be reordered to be before this patch, no problem. I don't think\n>    there will be any text conflicts _except_ inside the\n>    write_graph_chunk_generation_data() method introduced here.\n> \n> 2. The patch after that, \"commit-graph: implement corrected commit date\"\n>    only has a small dependence: it writes to the GDAT chunk and parses\n>    it out. If you remove the interaction with the GDAT chunk, then you\n>    still have the computation as part of compute_generation_numbers()\n>    that is valuable. You will need to be careful about the exit\n>    condition, though, since you also introduce the topo_level chunk.\n> \n> Patches 5-7 could perhaps be reorganized as follows:\n> \n>   i. commit-graph: return 64-bit generation number, as-is.\n> \n>  ii. Add a topo_level slab that is parsed from CDAT. Modify\n>      compute_generation_numbers() to populate this value and modify\n>      write_graph_chunk_data() to read this value. Simultaneously\n>      populate the \"generation\" member with the same value.\n> \n> iii. \"commit-graph: implement corrected commit date\" without any GDAT\n>      chunk interaction. Make sure the algorithm in\n>      compute_generation_numbers() walks commits if either topo_level or\n>      generation are unset. There is a trick here: the generation value\n>      _is_ set if the commit is parsed from the existing commit-graph!\n>      Is this case covered by the existing logic to not write GDAT when\n>      writing a split commit-graph file with a base that does not have\n>      GDAT? Note that the non-split case does not load the commit-graph\n>      for parsing, so the interesting case is \"--split-replace\". Worth\n>      a test (after we write the GDAT chunk), which you have in \"commit-graph:\n>      handle mixed generation commit chains\".\n> \n\nRight, so at the end of this patch we compute corrected commit dates but\ndon't write them to graph file.\n\nAlthough, writing ii. and iii. together in the same patch makes more\nsense to me. Would it be hard to follow for someone who has no context\nof this discussion?\n\n>  iv. This patch, introducing the chunk and the read/write logic.\n> \n>   v. Add the remaining patches.\n> \n> Again, this is a complicated patch-reorganization. The hope is that\n> the end result is something that is easy to review as well as something\n> that produces an as-sane-as-possible history for future bisecters.\n> \n> Perhaps other reviewers have similar feelings, or can say that I am\n> being too picky.\n> \n\nI can see how the reorganization helps us avoid a rather nasty\nsituation to be in. Should not be too hard to reorganize.\n\n> > We introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\n> > which forces commit-graph file to be written without generation data\n> > chunk to emulate a commit-graph file written by old Git.\n> \n> ...\n> \n> Thanks,\n> -Stolee\n\nThanks\n- Abhishek\n"},{"id":"403346","messageId":"20200811113621.GB3220@Abhishek-Arch","threadId":"53933","inReplyTo":"0d741fb2-e25a-be05-9f2b-81ba2b4ced3f@gmail.com","subject":"Re: [PATCH v2 08/10] commit-graph: handle mixed generation commit chains","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-08-11T11:36:21Z","receivedAt":"2020-08-11T11:38:33Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Mon, Aug 10, 2020 at 12:42:29PM -0400, Derrick Stolee wrote:\n> On 8/8/2020 10:53 PM, Abhishek Kumar via GitGitGadget wrote:\n> \n> ...\n>\n> Hm. So this scenario actually disables generation numbers completely\n> in the event that anything in the chain disagrees. I think this is\n> not the right way to approach the situation, as it will significantly\n> punish users in this state with slow performance.\n> \n> The patch I sent [1] is probably better: it uses generation number\n> v1 if the tip of the chain does not have a GDAT chunk.\n> \n> [1] https://lore.kernel.org/git/a3910f82-ab2e-bf35-ac43-c30d77f3c96b@gmail.com/\n> \n\nYes, the patch is an clear improvement over my (convoluted and incorrect)\nlogic. Will add.\n\n>\n> ...\n> \n> Please make a point to move the line that checks GIT_TEST_COMMIT_GRAPH_NO_GDAT\n> from its current location to after this line. We want to make sure that the\n> environment variable is checked _last_. The best location is likely the start\n> of the implementation of compute_generation_numbers(), or immediately before\n> the call to the method.\n> \n\nSure, will do.\n\n>\n> ...\n> \n> It would be valuable to double-check here that the values in the GDAT chunk\n> are correct. I'm concerned about the possibility that the 'generation'\n> member of struct commit_graph_data gets filled with topological level during\n> parsing and then that is written as an offset into the CDAT chunk.\n> \n> Perhaps this is best left for a follow-up series that updates the 'verify'\n> subcommand to check the GDAT chunk.\n\nIf I can understand it correctly, one of ways to update 'verify'\nsubcommand to check the GDAT chunk as well would to be make use of the\nflag variable introduced in your patch. We can isolate generation number\nrelated checks and run checks once with flag = 1 (checking corrected\ncommit dates) and once with flag = 0 (checking topological levels).\n\nThis has the unfortunate effect of filling all commits twice, but as we\ncannot change the commit_graph_data->generation any other way, I see no\nalternatives without changing how commit_graph_generation() works.\n\nWould it make more sense if we add the flag to struct commit_graph\ninstead of making it depend solely on g->chunk_generation_data and set\nit within parse_commit_graph()?\n\nWe would be able to control the behavior of fill_commit_graph_info() and\nwe will not need to check g->chunk_generation_data before filling every\ncommit.\n\n> \n> Thanks,\n> -Stolee\n> \n\nThanks\n- Abhishek\n"},{"id":"403351","messageId":"3c053281-f5af-1ac8-75ef-9eb8ce4f539d@gmail.com","threadId":"53933","inReplyTo":"20200811110316.GA3220@Abhishek-Arch","subject":"Re: [PATCH v2 05/10] commit-graph: implement generation data chunk","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-11T12:27:41Z","receivedAt":"2020-08-11T12:27:51Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/11/2020 7:03 AM, Abhishek Kumar wrote:\n> On Mon, Aug 10, 2020 at 12:28:10PM -0400, Derrick Stolee wrote:\n>> Patches 5-7 could perhaps be reorganized as follows:\n>>\n>>   i. commit-graph: return 64-bit generation number, as-is.\n>>\n>>  ii. Add a topo_level slab that is parsed from CDAT. Modify\n>>      compute_generation_numbers() to populate this value and modify\n>>      write_graph_chunk_data() to read this value. Simultaneously\n>>      populate the \"generation\" member with the same value.\n>>\n>> iii. \"commit-graph: implement corrected commit date\" without any GDAT\n>>      chunk interaction. Make sure the algorithm in\n>>      compute_generation_numbers() walks commits if either topo_level or\n>>      generation are unset. There is a trick here: the generation value\n>>      _is_ set if the commit is parsed from the existing commit-graph!\n>>      Is this case covered by the existing logic to not write GDAT when\n>>      writing a split commit-graph file with a base that does not have\n>>      GDAT? Note that the non-split case does not load the commit-graph\n>>      for parsing, so the interesting case is \"--split-replace\". Worth\n>>      a test (after we write the GDAT chunk), which you have in \"commit-graph:\n>>      handle mixed generation commit chains\".\n>>\n> \n> Right, so at the end of this patch we compute corrected commit dates but\n> don't write them to graph file.\n> \n> Although, writing ii. and iii. together in the same patch makes more\n> sense to me. Would it be hard to follow for someone who has no context\n> of this discussion?\n\nIt is always easier to combine two patches than to split one into two.\n\nWith that in mind, I recommend starting with a split version and then\nseeing how each patch looks. I think that these are \"independent enough\"\nideas that justify the separate patches.\n\n>>  iv. This patch, introducing the chunk and the read/write logic.\n>>\n>>   v. Add the remaining patches.\n>>\n>> Again, this is a complicated patch-reorganization. The hope is that\n>> the end result is something that is easy to review as well as something\n>> that produces an as-sane-as-possible history for future bisecters.\n>>\n>> Perhaps other reviewers have similar feelings, or can say that I am\n>> being too picky.\n>>\n> \n> I can see how the reorganization helps us avoid a rather nasty\n> situation to be in. Should not be too hard to reorganize.\n\nI hope not. Hopefully you get some more review on this version\nbefore jumping in on such a big reorg (in case someone else has\na different opinion).\n\nThanks,\n-Stolee\n"},{"id":"403352","messageId":"4043ffbc-84df-0cd6-5c75-af80383a56cf@gmail.com","threadId":"53933","inReplyTo":"20200811113621.GB3220@Abhishek-Arch","subject":"Re: [PATCH v2 08/10] commit-graph: handle mixed generation commit chains","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-11T12:43:59Z","receivedAt":"2020-08-11T12:44:04Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/11/2020 7:36 AM, Abhishek Kumar wrote:\n> On Mon, Aug 10, 2020 at 12:42:29PM -0400, Derrick Stolee wrote:\n>> On 8/8/2020 10:53 PM, Abhishek Kumar via GitGitGadget wrote:\n>>\n>> ...\n>>\n>> Hm. So this scenario actually disables generation numbers completely\n>> in the event that anything in the chain disagrees. I think this is\n>> not the right way to approach the situation, as it will significantly\n>> punish users in this state with slow performance.\n>>\n>> The patch I sent [1] is probably better: it uses generation number\n>> v1 if the tip of the chain does not have a GDAT chunk.\n>>\n>> [1] https://lore.kernel.org/git/a3910f82-ab2e-bf35-ac43-c30d77f3c96b@gmail.com/\n>>\n> \n> Yes, the patch is an clear improvement over my (convoluted and incorrect)\n> logic. Will add.\n> \n>>\n>> ...\n>>\n>> Please make a point to move the line that checks GIT_TEST_COMMIT_GRAPH_NO_GDAT\n>> from its current location to after this line. We want to make sure that the\n>> environment variable is checked _last_. The best location is likely the start\n>> of the implementation of compute_generation_numbers(), or immediately before\n>> the call to the method.\n>>\n> \n> Sure, will do.\n> \n>>\n>> ...\n>>\n>> It would be valuable to double-check here that the values in the GDAT chunk\n>> are correct. I'm concerned about the possibility that the 'generation'\n>> member of struct commit_graph_data gets filled with topological level during\n>> parsing and then that is written as an offset into the CDAT chunk.\n>>\n>> Perhaps this is best left for a follow-up series that updates the 'verify'\n>> subcommand to check the GDAT chunk.\n> \n> If I can understand it correctly, one of ways to update 'verify'\n> subcommand to check the GDAT chunk as well would to be make use of the\n> flag variable introduced in your patch. We can isolate generation number\n> related checks and run checks once with flag = 1 (checking corrected\n> commit dates) and once with flag = 0 (checking topological levels).\n> \n> This has the unfortunate effect of filling all commits twice, but as we\n> cannot change the commit_graph_data->generation any other way, I see no\n> alternatives without changing how commit_graph_generation() works.\n> \n> Would it make more sense if we add the flag to struct commit_graph\n> instead of making it depend solely on g->chunk_generation_data and set\n> it within parse_commit_graph()?\n> \n> We would be able to control the behavior of fill_commit_graph_info() and\n> we will not need to check g->chunk_generation_data before filling every\n> commit.\n\nI missed that you _already_ updated the logic in verify_commit_graph()\nbased on the generation. That logic should catch the problem, so it\nmight be enough to just add some \"git commit-graph verify\" commands into\nyour multi-level tests.\n\nSpecifically, the end result is this check:\n\n\tcorrected_commit_date = commit_graph_generation(graph_commit);\n\tif (corrected_commit_date < max_parent_corrected_commit_date + 1)\n\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n\t\t\t     oid_to_hex(&cur_oid),\n\t\t\t     corrected_commit_date,\n\t\t\t     max_parent_corrected_commit_date + 1);\n\nThis will catch the order violations I was proposing could happen. It\ndoesn't go the extra mile to ensure that the commit-graph stores the\nexact correct value or that the two bits of data are correct (both\ntopo-level and corrected commit date). That is fine for now, and we\ncan revisit if necessary.\n\nThe diff below makes some tweaks to your split-level test to show the\nlogic _was_ incorrect without my patch. Please incorporate the test\nchanges into your series. Note in particular that I added a base\nlayer that includes the GDAT chunk and _then_ adds a layer without\nthe GDAT chunk. That is an important case!\n\nThanks,\n-Stolee\n\n--- >8 ---\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 17623274d9..d891a8ba3a 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -674,14 +674,6 @@ int generation_numbers_enabled(struct repository *r)\n \tif (!g->num_commits)\n \t\treturn 0;\n \n-\t/* We cannot compare topological levels and corrected commit dates */\n-\twhile (g->base_graph) {\n-\t\twarning(_(\"commit-graph-chain contains mixed generation versions\"));\n-\t\tif ((g->chunk_generation_data == NULL) ^ (g->base_graph->chunk_generation_data == NULL))\n-\t\t\treturn 0;\n-\t\tg = g->base_graph;\n-\t}\n-\n \tfirst_generation = get_be32(g->chunk_commit_data +\n \t\t\t\t    g->hash_len + 8) >> 2;\n \n@@ -787,7 +779,7 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tif (g->chunk_generation_data && (flags & COMMIT_GRAPH_GENERATION_V2))\n+\tif (g->chunk_generation_data)\n \t\tgraph_data->generation = item->date +\n \t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n \telse\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 1a9be5e656..721515cc23 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -443,6 +443,7 @@ test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n \t\ttest_commit $i &&\n \t\tgit branch commits/$i || return 1\n \tdone &&\n+\tgit commit-graph write --reachable --split &&\n \tgit reset --hard commits/2 &&\n \tfor i in $(test_seq 6 10)\n \tdo\n@@ -455,14 +456,15 @@ test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n \tgit reset --hard commits/4 &&\n \tgit merge commits/6 &&\n \tgit branch merge/2 &&\n-\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split &&\n+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n \ttest-tool read-graph >output &&\n \tcat >expect <<-EOF &&\n-\theader: 43475048 1 1 3 0\n-\tnum_commits: 12\n+\theader: 43475048 1 1 4 1\n+\tnum_commits: 7\n \tchunks: oid_fanout oid_lookup commit_metadata\n \tEOF\n-\ttest_cmp expect output\n+\ttest_cmp expect output &&\n+\tgit commit-graph verify\n '\n \n test_expect_success 'does not write generation data chunk if not present on existing tip' '\n@@ -472,23 +474,25 @@ test_expect_success 'does not write generation data chunk if not present on exis\n \tgit merge commits/5 &&\n \tgit merge merge/2 &&\n \tgit branch merge/3 &&\n-\tgit commit-graph write --reachable --split &&\n+\tgit commit-graph write --reachable --split=no-merge &&\n \ttest-tool read-graph >output &&\n \tcat >expect <<-EOF &&\n \theader: 43475048 1 1 4 1\n \tnum_commits: 3\n \tchunks: oid_fanout oid_lookup commit_metadata\n \tEOF\n-\ttest_cmp expect output\n+\ttest_cmp expect output &&\n+\tgit commit-graph verify\n '\n \n test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n \tcd \"$TRASH_DIRECTORY/mixed\" &&\n-\tgit commit-graph write --reachable --split='replace' &&\n+\tgit commit-graph write --reachable --split=replace &&\n \ttest_path_is_file $graphdir/commit-graph-chain &&\n \ttest_line_count = 1 $graphdir/commit-graph-chain &&\n \tverify_chain_files_exist $graphdir &&\n-\tgraph_read_expect 15\n+\tgraph_read_expect 15 &&\n+\tgit commit-graph verify\n '\n \n test_done\n"},{"id":"403376","messageId":"20200811185839.GA34058@syl.lan","threadId":"53933","inReplyTo":"3c053281-f5af-1ac8-75ef-9eb8ce4f539d@gmail.com","subject":"Re: [PATCH v2 05/10] commit-graph: implement generation data chunk","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-08-11T18:58:39Z","receivedAt":"2020-08-11T18:58:44Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Aug 11, 2020 at 08:27:41AM -0400, Derrick Stolee wrote:\n> On 8/11/2020 7:03 AM, Abhishek Kumar wrote:\n> > On Mon, Aug 10, 2020 at 12:28:10PM -0400, Derrick Stolee wrote:\n> >> Patches 5-7 could perhaps be reorganized as follows:\n> >>\n> >>   i. commit-graph: return 64-bit generation number, as-is.\n> >>\n> >>  ii. Add a topo_level slab that is parsed from CDAT. Modify\n> >>      compute_generation_numbers() to populate this value and modify\n> >>      write_graph_chunk_data() to read this value. Simultaneously\n> >>      populate the \"generation\" member with the same value.\n> >>\n> >> iii. \"commit-graph: implement corrected commit date\" without any GDAT\n> >>      chunk interaction. Make sure the algorithm in\n> >>      compute_generation_numbers() walks commits if either topo_level or\n> >>      generation are unset. There is a trick here: the generation value\n> >>      _is_ set if the commit is parsed from the existing commit-graph!\n> >>      Is this case covered by the existing logic to not write GDAT when\n> >>      writing a split commit-graph file with a base that does not have\n> >>      GDAT? Note that the non-split case does not load the commit-graph\n> >>      for parsing, so the interesting case is \"--split-replace\". Worth\n> >>      a test (after we write the GDAT chunk), which you have in \"commit-graph:\n> >>      handle mixed generation commit chains\".\n> >>\n> >\n> > Right, so at the end of this patch we compute corrected commit dates but\n> > don't write them to graph file.\n> >\n> > Although, writing ii. and iii. together in the same patch makes more\n> > sense to me. Would it be hard to follow for someone who has no context\n> > of this discussion?\n>\n> It is always easier to combine two patches than to split one into two.\n>\n> With that in mind, I recommend starting with a split version and then\n> seeing how each patch looks. I think that these are \"independent enough\"\n> ideas that justify the separate patches.\n>\n> >>  iv. This patch, introducing the chunk and the read/write logic.\n> >>\n> >>   v. Add the remaining patches.\n> >>\n> >> Again, this is a complicated patch-reorganization. The hope is that\n> >> the end result is something that is easy to review as well as something\n> >> that produces an as-sane-as-possible history for future bisecters.\n> >>\n> >> Perhaps other reviewers have similar feelings, or can say that I am\n> >> being too picky.\n> >>\n> >\n> > I can see how the reorganization helps us avoid a rather nasty\n> > situation to be in. Should not be too hard to reorganize.\n>\n> I hope not. Hopefully you get some more review on this version\n> before jumping in on such a big reorg (in case someone else has\n> a different opinion).\n\nI think the direction makes sense. We should avoid having the dummy\nimplementation of the GDAT chunk in the interim if at all possible (and\nit seems like it is). What Stolee is proposing is what I'd suggest, too.\n\nPlease let us know if you need any help restructuring these patches.\nPlease make sure to give them a careful review since it is easy to move\na hunk into the wrong commit, or let a detail in the patch text become\nout-of-date. Looking over \"git log -p origin/master..HEAD\" and \"git\nrebase -x make -j8 DEVELOPER=1 test' origin/master\" never hurts ;).\n\n> Thanks,\n> -Stolee\n\nThanks,\nTaylor\n"},{"id":"403627","messageId":"20200814045957.GA1380@Abhishek-Arch","threadId":"53933","inReplyTo":"a3910f82-ab2e-bf35-ac43-c30d77f3c96b@gmail.com","subject":"Re: [PATCH v2 07/10] commit-graph: implement corrected commit date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-08-14T04:59:57Z","receivedAt":"2020-08-14T05:02:12Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Mon, Aug 10, 2020 at 10:23:32AM -0400, Derrick Stolee wrote:\n> ....\n> --- >8 ---\n> \n> From 62189709fad3b051cedbd36193f5244fcce17e1f Mon Sep 17 00:00:00 2001\n> From: Derrick Stolee <dstolee@microsoft.com>\n> Date: Mon, 10 Aug 2020 10:06:47 -0400\n> Subject: [PATCH] commit-graph: use generation v2 only if entire chain does\n> \n> Since there are released versions of Git that understand generation\n> numbers in the commit-graph's CDAT chunk but do not understand the GDAT\n> chunk, the following scenario is possible:\n> \n>  1. \"New\" Git writes a commit-graph with the GDAT chunk.\n>  2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n> \n> Because of the current use of inspecting the current layer for a\n> generation_data_chunk pointer, the commits in the lower layer will be\n> interpreted as having very large generation values (commit date plus\n> offset) compared to the generation numbers in the top layer (topological\n> level). This violates the expectation that the generation of a parent is\n> strictly smaller than the generation of a child.\n> \n> It is difficult to expose this issue in a test. Since we _start_ with\n> artificially low generation numbers, any commit walk that prioritizes\n> generation numbers will walk all of the commits with high generation\n> number before walking the commits with low generation number. In all the\n> cases I tried, the commit-graph layers themselves \"protect\" any\n> incorrect behavior since none of the commits in the lower layer can\n> reach the commits in the upper layer.\n> \n> This issue would manifest itself as a performance problem in this\n> case, especially with something like \"git log --graph\" since the low\n> generation numbers would cause the in-degree queue to walk all of the\n> commits in the lower layer before allowing the topo-order queue to write\n> anything to output (depending on the size of the upper layer).\n> \n> Signed-off-by: Derrick Stolee <dstolee@microsoft.com>\n> ---\n>  commit-graph.c | 24 ++++++++++++++++++------\n>  1 file changed, 18 insertions(+), 6 deletions(-)\n> \n> diff --git a/commit-graph.c b/commit-graph.c\n> index eb78af3dad..17623274d9 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -762,7 +762,9 @@ static struct commit_list **insert_parent_or_die(struct repository *r,\n>  \treturn &commit_list_insert(c, pptr)->next;\n>  }\n>  \n> -static void fill_commit_graph_info(struct commit *item, struct commit_graph *g, uint32_t pos)\n> +#define COMMIT_GRAPH_GENERATION_V2 (1 << 0)\n> +\n> +static void fill_commit_graph_info(struct commit *item, struct commit_graph *g, uint32_t pos, int flags)\n>  {\n>  \tconst unsigned char *commit_data;\n>  \tstruct commit_graph_data *graph_data;\n> @@ -785,11 +787,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \tdate_low = get_be32(commit_data + g->hash_len + 12);\n>  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n>  \n> -\tif (g->chunk_generation_data)\n> -\t{\n> +\tif (g->chunk_generation_data && (flags & COMMIT_GRAPH_GENERATION_V2))\n>  \t\tgraph_data->generation = item->date +\n>  \t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n> -\t}\n>  \telse\n>  \t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n>  }\n> @@ -799,6 +799,10 @@ static inline void set_commit_tree(struct commit *c, struct tree *t)\n>  \tc->maybe_tree = t;\n>  }\n>  \n> +/*\n> + * In the case of a split commit-graph, this method expects the given\n> + * commit-graph 'g' to be the top layer.\n> + */\n\nUnfortunately, this turns out to be an optimistic assumption. After\nadding in changes from this patch and extended tests for\nsplit-commit-graph [1], the tests fail with `git commit-graph verify` as\nGit tries to compare topological level with corrected commit dates.\n\nThe problem lies in assuming that `g` is always the top layer, without\nany way to assert if that's true. In case of `commit-graph verify`, we\nupdate `g = g->base_graph` and verify recursively.\n\nIf we can assume such behavior is only the part of verify subcommand, I\ncan update the tests to no longer verify at the end.\n\n[1]: https://lore.kernel.org/git/4043ffbc-84df-0cd6-5c75-af80383a56cf@gmail.com/\n\n>  static int fill_commit_in_graph(struct repository *r,\n>  \t\t\t\tstruct commit *item,\n>  \t\t\t\tstruct commit_graph *g, uint32_t pos)\n> @@ -808,11 +812,15 @@ static int fill_commit_in_graph(struct repository *r,\n>  \tstruct commit_list **pptr;\n>  \tconst unsigned char *commit_data;\n>  \tuint32_t lex_index;\n> +\tint flags = 0;\n> +\n> +\tif (g->chunk_generation_data)\n> +\t\tflags |= COMMIT_GRAPH_GENERATION_V2;\n>  \n>  \twhile (pos < g->num_commits_in_base)\n>  \t\tg = g->base_graph;\n>  \n> -\tfill_commit_graph_info(item, g, pos);\n> +\tfill_commit_graph_info(item, g, pos, flags);\n>  \n>  \tlex_index = pos - g->num_commits_in_base;\n>  \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n> @@ -904,10 +912,14 @@ int parse_commit_in_graph(struct repository *r, struct commit *item)\n>  void load_commit_graph_info(struct repository *r, struct commit *item)\n>  {\n>  \tuint32_t pos;\n> +\tint flags = 0;\n> +\n>  \tif (!prepare_commit_graph(r))\n>  \t\treturn;\n> +\tif (r->objects->commit_graph->chunk_generation_data)\n> +\t\tflags |= COMMIT_GRAPH_GENERATION_V2;\n>  \tif (find_commit_in_graph(item, r->objects->commit_graph, &pos))\n> -\t\tfill_commit_graph_info(item, r->objects->commit_graph, pos);\n> +\t\tfill_commit_graph_info(item, r->objects->commit_graph, pos, flags);\n>  }\n>  \n>  static struct tree *load_tree_for_commit(struct repository *r,\n> -- \n> 2.28.0.38.gc6f546511c1\n> \n\nI solved the issue by adding a new member to struct commit_graph\n`read_generation_data` to maintain the \"global\" state of the entire\ncommit-graph chain instead.\n\nThe relevant changes are in validate_mixed_generation_chain(),\nread_commit_graph_one() and fill_commit_graph_info().\n\nThanks\n- Abhishek\n\n--- >8 ---\nFrom aaf8c27bfec6e110a8bb12173c2dd612a8c6b8b9 Mon Sep 17 00:00:00 2001\nFrom: Abhishek Kumar <abhishekkumar8222@gmail.com>\nDate: Thu, 6 Aug 2020 19:08:52 +0530\nSubject: [PATCH] commit-graph: use generation v2 only if entire chain does\n\nSince there are released versions of Git that understand generation\nnumbers in the commit-graph's CDAT chunk but do not understand the GDAT\nchunk, the following scenario is possible:\n\n1. \"New\" Git writes a commit-graph with the GDAT chunk.\n2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n\nBecause of the current use of inspecting the current layer for a\nchunk_generation_data pointer, the commits in the lower layer will be\ninterpreted as having very large generation values (commit date plus\noffset) compared to the generation numbers in the top layer (topological\nlevel). This violates the expectation that the generation of a parent is\nstrictly smaller than the generation of a child.\n\nIt is difficult to expose this issue in a test. Since we _start_ with\nartificially low generation numbers, any commit walk that prioritizes\ngeneration numbers will walk all of the commits with high generation\nnumber before walking the commits with low generation number. In all the\ncases I tried, the commit-graph layers themselves \"protect\" any\nincorrect behavior since none of the commits in the lower layer can\nreach the commits in the upper layer.\n\nThis issue would manifest itself as a performance problem in this case,\nespecially with something like \"git log --graph\" since the low\ngeneration numbers would cause the in-degree queue to walk all of the\ncommits in the lower layer before allowing the topo-order queue to write\nanything to output (depending on the size of the upper layer).\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                | 36 ++++++++++++++++--\n commit-graph.h                |  1 +\n t/t5324-split-commit-graph.sh | 70 +++++++++++++++++++++++++++++++++++\n 3 files changed, 104 insertions(+), 3 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex d0f977852b..10309f870f 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -597,6 +597,27 @@ static struct commit_graph *load_commit_graph_chain(struct repository *r,\n \treturn graph_chain;\n }\n \n+static void validate_mixed_generation_chain(struct repository *r)\n+{\n+\tstruct commit_graph *g = r->objects->commit_graph;\n+\tint read_generation_data = 1;\n+\n+\twhile (g) {\n+\t\tif (!g->chunk_generation_data) {\n+\t\t\tread_generation_data = 0;\n+\t\t\tbreak;\n+\t\t}\n+\t\tg = g->base_graph;\n+\t}\n+\n+\tg = r->objects->commit_graph;\n+\n+\twhile (g) {\n+\t\tg->read_generation_data = read_generation_data;\n+\t\tg = g->base_graph;\n+\t}\n+}\n+\n struct commit_graph *read_commit_graph_one(struct repository *r,\n \t\t\t\t\t   struct object_directory *odb)\n {\n@@ -605,6 +626,8 @@ struct commit_graph *read_commit_graph_one(struct repository *r,\n \tif (!g)\n \t\tg = load_commit_graph_chain(r, odb);\n \n+\tvalidate_mixed_generation_chain(r);\n+\n \treturn g;\n }\n \n@@ -740,6 +763,8 @@ static struct commit_list **insert_parent_or_die(struct repository *r,\n \treturn &commit_list_insert(c, pptr)->next;\n }\n \n+#define COMMIT_GRAPH_GENERATION_V2 (1 << 0)\n+\n static void fill_commit_graph_info(struct commit *item, struct commit_graph *g, uint32_t pos)\n {\n \tconst unsigned char *commit_data;\n@@ -763,11 +788,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tif (g->chunk_generation_data)\n-\t{\n+\tif (g->chunk_generation_data && g->read_generation_data)\n \t\tgraph_data->generation = item->date +\n \t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n-\t}\n \telse\n \t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n@@ -884,6 +907,7 @@ void load_commit_graph_info(struct repository *r, struct commit *item)\n \tuint32_t pos;\n \tif (!prepare_commit_graph(r))\n \t\treturn;\n+\n \tif (find_commit_in_graph(item, r->objects->commit_graph, &pos))\n \t\tfill_commit_graph_info(item, r->objects->commit_graph, pos);\n }\n@@ -2186,6 +2210,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \t\tg = ctx->r->objects->commit_graph;\n \n+\t\tif (g && !g->chunk_generation_data)\n+\t\t\tctx->write_generation_data = 0;\n+\n \t\twhile (g) {\n \t\t\tctx->num_commit_graphs_before++;\n \t\t\tg = g->base_graph;\n@@ -2204,6 +2231,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \t\tif (ctx->split_opts)\n \t\t\treplace = ctx->split_opts->flags & COMMIT_GRAPH_SPLIT_REPLACE;\n+\n+\t\tif (replace)\n+\t\t\tctx->write_generation_data = 1;\n \t}\n \n \tctx->approx_nr_objects = approximate_object_count();\ndiff --git a/commit-graph.h b/commit-graph.h\nindex f89614ecd5..305c332b7e 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -63,6 +63,7 @@ struct commit_graph {\n \tstruct object_directory *odb;\n \n \tuint32_t num_commits_in_base;\n+\tuint32_t read_generation_data;\n \tstruct commit_graph *base_graph;\n \n \tconst uint32_t *chunk_oid_fanout;\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 531016f405..ac5e7783fb 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -424,4 +424,74 @@ done <<\\EOF\n 0600 -r--------\n EOF\n \n+test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n+\tmkdir mixed &&\n+\tgraphdir=\".git/objects/info/commit-graphs\" &&\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit init &&\n+\tgit config core.commitGraph true &&\n+\tgit config gc.writeCommitGraph false &&\n+\tfor i in $(test_seq 3)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit reset --hard commits/1 &&\n+\tfor i in $(test_seq 4 5)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit reset --hard commits/2 &&\n+\tfor i in $(test_seq 6 10)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit commit-graph write --reachable --split &&\n+\tgit reset --hard commits/2 &&\n+\tgit merge commits/4 &&\n+\tgit branch merge/1 &&\n+\tgit reset --hard commits/4 &&\n+\tgit merge commits/6 &&\n+\tgit branch merge/2 &&\n+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n+\ttest-tool read-graph >output &&\n+\tcat >expect <<-EOF &&\n+\theader: 43475048 1 1 4 1\n+\tnum_commits: 2\n+\tchunks: oid_fanout oid_lookup commit_metadata\n+\tEOF\n+\ttest_cmp expect output &&\n+\tgit commit-graph verify\n+'\n+\n+test_expect_success 'does not write generation data chunk if not present on existing tip' '\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit reset --hard commits/3 &&\n+\tgit merge merge/1 &&\n+\tgit merge commits/5 &&\n+\tgit merge merge/2 &&\n+\tgit branch merge/3 &&\n+\tgit commit-graph write --reachable --split=no-merge &&\n+\ttest-tool read-graph >output &&\n+\tcat >expect <<-EOF &&\n+\theader: 43475048 1 1 4 2\n+\tnum_commits: 3\n+\tchunks: oid_fanout oid_lookup commit_metadata\n+\tEOF\n+\ttest_cmp expect output &&\n+\tgit commit-graph verify\n+'\n+\n+test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit commit-graph write --reachable --split=replace &&\n+\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\tverify_chain_files_exist $graphdir &&\n+\tgraph_read_expect 15 &&\n+\tgit commit-graph verify\n+'\n+\n test_done\n-- \n2.28.0\n"},{"id":"403653","messageId":"2c7ee14e-1c80-9860-7bcc-633ac43910a6@gmail.com","threadId":"53933","inReplyTo":"20200814045957.GA1380@Abhishek-Arch","subject":"Re: [PATCH v2 07/10] commit-graph: implement corrected commit date","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-14T12:24:38Z","receivedAt":"2020-08-14T12:24:42Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/14/2020 12:59 AM, Abhishek Kumar wrote:\n> I solved the issue by adding a new member to struct commit_graph\n> `read_generation_data` to maintain the \"global\" state of the entire\n> commit-graph chain instead.\n> \n> The relevant changes are in validate_mixed_generation_chain(),\n> read_commit_graph_one() and fill_commit_graph_info().\n\nI think this is a good way to go. Adding that restriction about\nthe tip commit-graph was short-sighted of me and was likely to\nbreak in the future.\n\nI think your solution here to store extra state from the entire\nchain into each layer makes a lot of sense.\n\nThanks!\n-Stolee\n"},{"id":"403738","messageId":"c6b7ade7af92b6faf365a5609748f6d024ea0408.1597509583.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v3 01/11] commit-graph: fix regression when computing bloom filter","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:33Z","receivedAt":"2020-08-15T21:52:15Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\ncommit_gen_cmp is used when writing a commit-graph to sort commits in\ngeneration order before computing Bloom filters. Since c49c82aa (commit:\nmove members graph_pos, generation to a slab, 2020-06-17) made it so\nthat 'commit_graph_generation()' returns 'GENERATION_NUMBER_INFINITY'\nduring writing, we cannot call it within this function. Instead, access\nthe generation number directly through the slab (i.e., by calling\n'commit_graph_data_at(c)->generation') in order to access it while\nwriting.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex e51c91dd5b..ace7400a1a 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n+\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n+\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"403740","messageId":"6a0cde983d9ed20f043a4977313d714154602012.1597509583.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v3 04/11] commit-graph: consolidate compare_commits_by_gen","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:36Z","receivedAt":"2020-08-15T21:53:15Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nComparing commits by generation has been independently defined twice, in\ncommit-reach and commit. Let's simplify the implementation by moving\ncompare_commits_by_gen() to commit-graph.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\nReviewed-by: Taylor Blau <me@ttaylorr.com>\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 15 +++++++++++++++\n commit-graph.h |  2 ++\n commit-reach.c | 15 ---------------\n commit.c       |  9 +++------\n 4 files changed, 20 insertions(+), 21 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex af8d9cc45e..fb6e2bf18f 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -112,6 +112,21 @@ uint32_t commit_graph_generation(const struct commit *c)\n \treturn data->generation;\n }\n \n+int compare_commits_by_gen(const void *_a, const void *_b)\n+{\n+\tconst struct commit *a = _a, *b = _b;\n+\tconst uint32_t generation_a = commit_graph_generation(a);\n+\tconst uint32_t generation_b = commit_graph_generation(b);\n+\n+\t/* older commits first */\n+\tif (generation_a < generation_b)\n+\t\treturn -1;\n+\telse if (generation_a > generation_b)\n+\t\treturn 1;\n+\n+\treturn 0;\n+}\n+\n static struct commit_graph_data *commit_graph_data_at(const struct commit *c)\n {\n \tunsigned int i, nth_slab;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 09a97030dc..701e3d41aa 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -146,4 +146,6 @@ struct commit_graph_data {\n  */\n uint32_t commit_graph_generation(const struct commit *);\n uint32_t commit_graph_position(const struct commit *);\n+\n+int compare_commits_by_gen(const void *_a, const void *_b);\n #endif\ndiff --git a/commit-reach.c b/commit-reach.c\nindex efd5925cbb..c83cc291e7 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -561,21 +561,6 @@ int commit_contains(struct ref_filter *filter, struct commit *commit,\n \treturn repo_is_descendant_of(the_repository, commit, list);\n }\n \n-static int compare_commits_by_gen(const void *_a, const void *_b)\n-{\n-\tconst struct commit *a = *(const struct commit * const *)_a;\n-\tconst struct commit *b = *(const struct commit * const *)_b;\n-\n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n-\n-\tif (generation_a < generation_b)\n-\t\treturn -1;\n-\tif (generation_a > generation_b)\n-\t\treturn 1;\n-\treturn 0;\n-}\n-\n int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\ndiff --git a/commit.c b/commit.c\nindex 4ce8cb38d5..bd6d5e587f 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -731,14 +731,11 @@ int compare_commits_by_author_date(const void *a_, const void *b_,\n int compare_commits_by_gen_then_commit_date(const void *a_, const void *b_, void *unused)\n {\n \tconst struct commit *a = a_, *b = b_;\n-\tconst uint32_t generation_a = commit_graph_generation(a),\n-\t\t       generation_b = commit_graph_generation(b);\n+\tint ret_val = compare_commits_by_gen(a_, b_);\n \n \t/* newer commits first */\n-\tif (generation_a < generation_b)\n-\t\treturn 1;\n-\telse if (generation_a > generation_b)\n-\t\treturn -1;\n+\tif (ret_val)\n+\t\treturn -ret_val;\n \n \t/* use date as a heuristic when generations are equal */\n \tif (a->date < b->date)\n-- \ngitgitgadget\n\n"},{"id":"403746","messageId":"f6f91af30587ec24e2eee052c89a536cbff42c4f.1597509583.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v3 11/11] doc: add corrected commit date info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:43Z","receivedAt":"2020-08-15T21:55:42Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith generation data chunk and corrected commit dates implemented, let's\nupdate the technical documentation for commit-graph.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n .../technical/commit-graph-format.txt         | 12 ++---\n Documentation/technical/commit-graph.txt      | 45 ++++++++++++-------\n 2 files changed, 36 insertions(+), 21 deletions(-)\n\ndiff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\nindex 440541045d..71c43884ec 100644\n--- a/Documentation/technical/commit-graph-format.txt\n+++ b/Documentation/technical/commit-graph-format.txt\n@@ -4,11 +4,7 @@ Git commit graph format\n The Git commit graph stores a list of commit OIDs and some associated\n metadata, including:\n \n-- The generation number of the commit. Commits with no parents have\n-  generation number 1; commits with parents have generation number\n-  one more than the maximum generation number of its parents. We\n-  reserve zero as special, and can be used to mark a generation\n-  number invalid or as \"not computed\".\n+- The generation number of the commit.\n \n - The root tree OID.\n \n@@ -88,6 +84,12 @@ CHUNK DATA:\n       2 bits of the lowest byte, storing the 33rd and 34th bit of the\n       commit time.\n \n+  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes) [Optional]\n+    * This list of 4-byte values store corrected commit date offsets for the\n+      commits, arranged in the same order as commit data chunk.\n+    * This list can be later modified to store future generation number related\n+      data.\n+\n   Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n       This list of 4-byte values store the second through nth parents for\n       all octopus merges. The second parent value in the commit data stores\ndiff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\nindex 808fa30b99..f27145328c 100644\n--- a/Documentation/technical/commit-graph.txt\n+++ b/Documentation/technical/commit-graph.txt\n@@ -38,14 +38,27 @@ A consumer may load the following info for a commit from the graph:\n \n Values 1-4 satisfy the requirements of parse_commit_gently().\n \n-Define the \"generation number\" of a commit recursively as follows:\n+There are two definitions of generation number:\n+1. Corrected committer dates\n+2. Topological levels\n+\n+Define \"corrected committer date\" of a commit recursively as follows:\n+\n+  * A commit with no parents (a root commit) has corrected committer date\n+    equal to its committer date.\n+\n+  * A commit with at least one parent has corrected committer date equal to\n+    the maximum of its commiter date and one more than the largest corrected\n+    committer date among its parents.\n+\n+Define the \"topological level\" of a commit recursively as follows:\n \n  * A commit with no parents (a root commit) has generation number one.\n \n- * A commit with at least one parent has generation number one more than\n-   the largest generation number among its parents.\n+ * A commit with at least one parent has topological level one more than\n+   the largest topological level among its parents.\n \n-Equivalently, the generation number of a commit A is one more than the\n+Equivalently, the topological level of a commit A is one more than the\n length of a longest path from A to a root commit. The recursive definition\n is easier to use for computation and observing the following property:\n \n@@ -67,17 +80,12 @@ numbers, the general heuristic is the following:\n     If A and B are commits with commit time X and Y, respectively, and\n     X < Y, then A _probably_ cannot reach B.\n \n-This heuristic is currently used whenever the computation is allowed to\n-violate topological relationships due to clock skew (such as \"git log\"\n-with default order), but is not used when the topological order is\n-required (such as merge base calculations, \"git log --graph\").\n-\n In practice, we expect some commits to be created recently and not stored\n in the commit graph. We can treat these commits as having \"infinite\"\n generation number and walk until reaching commits with known generation\n number.\n \n-We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n+We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n in the commit-graph file. If a commit-graph file was written by a version\n of Git that did not compute generation numbers, then those commits will\n have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n@@ -93,12 +101,11 @@ fully-computed generation numbers. Using strict inequality may result in\n walking a few extra commits, but the simplicity in dealing with commits\n with generation number *_INFINITY or *_ZERO is valuable.\n \n-We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n-generation numbers are computed to be at least this value. We limit at\n-this value since it is the largest value that can be stored in the\n-commit-graph file using the 30 bits available to generation numbers. This\n-presents another case where a commit can have generation number equal to\n-that of a parent.\n+We use the macro GENERATION_NUMBER_MAX for commits whose generation numbers\n+are computed to be at least this value. We limit at this value since it is\n+the largest value that can be stored in the commit-graph file using the\n+available to generation numbers. This presents another case where a\n+commit can have generation number equal to that of a parent.\n \n Design Details\n --------------\n@@ -267,6 +274,12 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n number of commits) could be extracted into config settings for full\n flexibility.\n \n+We also merge commit-graph chains when we try to write a commit graph with\n+two different generation number definitions as they cannot be compared directly.\n+We overwrite the existing chain and create a commit-graph with the newer or more\n+efficient defintion. For example, overwriting topological levels commit graph\n+chain to create a corrected commit dates commit graph chain.\n+\n ## Deleting graph-{hash} files\n \n After a new tip file is written, some `graph-{hash}` files may no longer\n-- \ngitgitgadget\n"},{"id":"403748","messageId":"6be759a9542114e4de41422efa18491085e19682.1597509583.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v3 05/11] commit-graph: return 64-bit generation number","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:37Z","receivedAt":"2020-08-15T21:56:03Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a preparatory step, let's return timestamp_t values from\ncommit_graph_generation(), use timestamp_t for local variables and\ndefine GENERATION_NUMBER_INFINITY as (2 ^ 63 - 1) instead.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 18 +++++++++---------\n commit-graph.h |  4 ++--\n commit-reach.c | 32 ++++++++++++++++----------------\n commit-reach.h |  2 +-\n commit.h       |  3 ++-\n revision.c     | 10 +++++-----\n upload-pack.c  |  2 +-\n 7 files changed, 36 insertions(+), 35 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex fb6e2bf18f..7f9f858577 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -99,7 +99,7 @@ uint32_t commit_graph_position(const struct commit *c)\n \treturn data ? data->graph_pos : COMMIT_NOT_FROM_GRAPH;\n }\n \n-uint32_t commit_graph_generation(const struct commit *c)\n+timestamp_t commit_graph_generation(const struct commit *c)\n {\n \tstruct commit_graph_data *data =\n \t\tcommit_graph_data_slab_peek(&commit_graph_data_slab, c);\n@@ -115,8 +115,8 @@ uint32_t commit_graph_generation(const struct commit *c)\n int compare_commits_by_gen(const void *_a, const void *_b)\n {\n \tconst struct commit *a = _a, *b = _b;\n-\tconst uint32_t generation_a = commit_graph_generation(a);\n-\tconst uint32_t generation_b = commit_graph_generation(b);\n+\tconst timestamp_t generation_a = commit_graph_generation(a);\n+\tconst timestamp_t generation_b = commit_graph_generation(b);\n \n \t/* older commits first */\n \tif (generation_a < generation_b)\n@@ -159,8 +159,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n-\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n+\tconst timestamp_t generation_a = commit_graph_data_at(a)->generation;\n+\tconst timestamp_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n@@ -1338,7 +1338,7 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n+\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n \t\t    generation != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n@@ -1352,7 +1352,7 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n \t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n \n-\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n+\t\t\t\tif (generation == GENERATION_NUMBER_V1_INFINITY ||\n \t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n@@ -2355,8 +2355,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \tfor (i = 0; i < g->num_commits; i++) {\n \t\tstruct commit *graph_commit, *odb_commit;\n \t\tstruct commit_list *graph_parents, *odb_parents;\n-\t\tuint32_t max_generation = 0;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t max_generation = 0;\n+\t\ttimestamp_t generation;\n \n \t\tdisplay_progress(progress, i + 1);\n \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 701e3d41aa..430bc830bb 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -138,13 +138,13 @@ void disable_commit_graph(struct repository *r);\n \n struct commit_graph_data {\n \tuint32_t graph_pos;\n-\tuint32_t generation;\n+\ttimestamp_t generation;\n };\n \n /*\n  * Commits should be parsed before accessing generation, graph positions.\n  */\n-uint32_t commit_graph_generation(const struct commit *);\n+timestamp_t commit_graph_generation(const struct commit *);\n uint32_t commit_graph_position(const struct commit *);\n \n int compare_commits_by_gen(const void *_a, const void *_b);\ndiff --git a/commit-reach.c b/commit-reach.c\nindex c83cc291e7..470bc80139 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -32,12 +32,12 @@ static int queue_has_nonstale(struct prio_queue *queue)\n static struct commit_list *paint_down_to_common(struct repository *r,\n \t\t\t\t\t\tstruct commit *one, int n,\n \t\t\t\t\t\tstruct commit **twos,\n-\t\t\t\t\t\tint min_generation)\n+\t\t\t\t\t\ttimestamp_t min_generation)\n {\n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \tstruct commit_list *result = NULL;\n \tint i;\n-\tuint32_t last_gen = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n \tif (!min_generation)\n \t\tqueue.compare = compare_commits_by_commit_date;\n@@ -58,10 +58,10 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \t\tstruct commit *commit = prio_queue_get(&queue);\n \t\tstruct commit_list *parents;\n \t\tint flags;\n-\t\tuint32_t generation = commit_graph_generation(commit);\n+\t\ttimestamp_t generation = commit_graph_generation(commit);\n \n \t\tif (min_generation && generation > last_gen)\n-\t\t\tBUG(\"bad generation skip %8x > %8x at %s\",\n+\t\t\tBUG(\"bad generation skip %\"PRItime\" > %\"PRItime\" at %s\",\n \t\t\t    generation, last_gen,\n \t\t\t    oid_to_hex(&commit->object.oid));\n \t\tlast_gen = generation;\n@@ -177,12 +177,12 @@ static int remove_redundant(struct repository *r, struct commit **array, int cnt\n \t\trepo_parse_commit(r, array[i]);\n \tfor (i = 0; i < cnt; i++) {\n \t\tstruct commit_list *common;\n-\t\tuint32_t min_generation = commit_graph_generation(array[i]);\n+\t\ttimestamp_t min_generation = commit_graph_generation(array[i]);\n \n \t\tif (redundant[i])\n \t\t\tcontinue;\n \t\tfor (j = filled = 0; j < cnt; j++) {\n-\t\t\tuint32_t curr_generation;\n+\t\t\ttimestamp_t curr_generation;\n \t\t\tif (i == j || redundant[j])\n \t\t\t\tcontinue;\n \t\t\tfilled_index[filled] = j;\n@@ -321,7 +321,7 @@ int repo_in_merge_bases_many(struct repository *r, struct commit *commit,\n {\n \tstruct commit_list *bases;\n \tint ret = 0, i;\n-\tuint32_t generation, min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t generation, min_generation = GENERATION_NUMBER_INFINITY;\n \n \tif (repo_parse_commit(r, commit))\n \t\treturn ret;\n@@ -470,7 +470,7 @@ static int in_commit_list(const struct commit_list *want, struct commit *c)\n static enum contains_result contains_test(struct commit *candidate,\n \t\t\t\t\t  const struct commit_list *want,\n \t\t\t\t\t  struct contains_cache *cache,\n-\t\t\t\t\t  uint32_t cutoff)\n+\t\t\t\t\t  timestamp_t cutoff)\n {\n \tenum contains_result *cached = contains_cache_at(cache, candidate);\n \n@@ -506,11 +506,11 @@ static enum contains_result contains_tag_algo(struct commit *candidate,\n {\n \tstruct contains_stack contains_stack = { 0, 0, NULL };\n \tenum contains_result result;\n-\tuint32_t cutoff = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t cutoff = GENERATION_NUMBER_INFINITY;\n \tconst struct commit_list *p;\n \n \tfor (p = want; p; p = p->next) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = p->item;\n \t\tload_commit_graph_info(the_repository, c);\n \t\tgeneration = commit_graph_generation(c);\n@@ -565,7 +565,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation)\n+\t\t\t\t timestamp_t min_generation)\n {\n \tstruct commit **list = NULL;\n \tint i;\n@@ -666,13 +666,13 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \ttime_t min_commit_date = cutoff_by_min_date ? from->item->date : 0;\n \tstruct commit_list *from_iter = from, *to_iter = to;\n \tint result;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \n \twhile (from_iter) {\n \t\tadd_object_array(&from_iter->item->object, NULL, &from_objs);\n \n \t\tif (!parse_commit(from_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (from_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = from_iter->item->date;\n \n@@ -686,7 +686,7 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \n \twhile (to_iter) {\n \t\tif (!parse_commit(to_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (to_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = to_iter->item->date;\n \n@@ -726,13 +726,13 @@ struct commit_list *get_reachable_subset(struct commit **from, int nr_from,\n \tstruct commit_list *found_commits = NULL;\n \tstruct commit **to_last = to + nr_to;\n \tstruct commit **from_last = from + nr_from;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \tint num_to_find = 0;\n \n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \n \tfor (item = to; item < to_last; item++) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = *item;\n \n \t\tparse_commit(c);\ndiff --git a/commit-reach.h b/commit-reach.h\nindex b49ad71a31..148b56fea5 100644\n--- a/commit-reach.h\n+++ b/commit-reach.h\n@@ -87,7 +87,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation);\n+\t\t\t\t timestamp_t min_generation);\n int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \t\t       int commit_date_cutoff);\n \ndiff --git a/commit.h b/commit.h\nindex e901538909..bc0732a4fe 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -11,7 +11,8 @@\n #include \"commit-slab.h\"\n \n #define COMMIT_NOT_FROM_GRAPH 0xFFFFFFFF\n-#define GENERATION_NUMBER_INFINITY 0xFFFFFFFF\n+#define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n+#define GENERATION_NUMBER_V1_INFINITY 0xFFFFFFFF\n #define GENERATION_NUMBER_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n \ndiff --git a/revision.c b/revision.c\nindex ecf757c327..411852468b 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3290,7 +3290,7 @@ define_commit_slab(indegree_slab, int);\n define_commit_slab(author_date_slab, timestamp_t);\n \n struct topo_walk_info {\n-\tuint32_t min_generation;\n+\ttimestamp_t min_generation;\n \tstruct prio_queue explore_queue;\n \tstruct prio_queue indegree_queue;\n \tstruct prio_queue topo_queue;\n@@ -3336,7 +3336,7 @@ static void explore_walk_step(struct rev_info *revs)\n }\n \n static void explore_to_depth(struct rev_info *revs,\n-\t\t\t     uint32_t gen_cutoff)\n+\t\t\t     timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3379,7 +3379,7 @@ static void indegree_walk_step(struct rev_info *revs)\n }\n \n static void compute_indegrees_to_depth(struct rev_info *revs,\n-\t\t\t\t       uint32_t gen_cutoff)\n+\t\t\t\t       timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3437,7 +3437,7 @@ static void init_topo_walk(struct rev_info *revs)\n \tinfo->min_generation = GENERATION_NUMBER_INFINITY;\n \tfor (list = revs->commits; list; list = list->next) {\n \t\tstruct commit *c = list->item;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (parse_commit_gently(c, 1))\n \t\t\tcontinue;\n@@ -3498,7 +3498,7 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n \tfor (p = commit->parents; p; p = p->next) {\n \t\tstruct commit *parent = p->item;\n \t\tint *pi;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (parent->object.flags & UNINTERESTING)\n \t\t\tcontinue;\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 80ad9a38d8..bcb8b5dfda 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -497,7 +497,7 @@ static int got_oid(struct upload_pack_data *data,\n \n static int ok_to_give_up(struct upload_pack_data *data)\n {\n-\tuint32_t min_generation = GENERATION_NUMBER_ZERO;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_ZERO;\n \n \tif (!data->have_obj.nr)\n \t\treturn 0;\n-- \ngitgitgadget\n\n"},{"id":"403749","messageId":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v2.git.1596941624.gitgitgadget@gmail.com","subject":"[PATCH v3 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:32Z","receivedAt":"2020-08-15T21:56:34Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"This patch series implements the corrected commit date offsets as generation\nnumber v2, along with other pre-requisites.\n\nGit uses topological levels in the commit-graph file for commit-graph\ntraversal operations like git log --graph. Unfortunately, using topological\nlevels can result in a worse performance than without them when compared\nwith committer date as a heuristics. For example, git merge-base v4.8 v4.9 \non the Linux repository walks 635,579 commits using topological levels and\nwalks 167,468 using committer date.\n\nThus, the need for generation number v2 was born. New generation number\nneeded to provide good performance, increment updates, and backward\ncompatibility. Due to an unfortunate problem, we also needed a way to\ndistinguish between the old and new generation number without incrementing\ngraph version.\n\nVarious candidates were examined (https://github.com/derrickstolee/gen-test, \nhttps://github.com/abhishekkumar2718/git/pull/1). The proposed generation\nnumber v2, Corrected Commit Date with Mononotically Increasing Offsets \nperformed much worse than committer date (506,577 vs. 167,468 commits walked\nfor git merge-base v4.8 v4.9) and was dropped.\n\nUsing Generation Data chunk (GDAT) relieves the requirement of backward\ncompatibility as we would continue to store topological levels in Commit\nData (CDAT) chunk. Thus, Corrected Commit Date was chosen as generation\nnumber v2. The Corrected Commit Date is defined as:\n\nFor a commit C, let its corrected commit date be the maximum of the commit\ndate of C and the corrected commit dates of its parents. Then corrected\ncommit date offset is the difference between corrected commit date of C and\ncommit date of C.\n\nWe will introduce an additional commit-graph chunk, Generation Data chunk,\nand store corrected commit date offsets in GDAT chunk while storing\ntopological levels in CDAT chunk. The old versions of Git would ignore GDAT\nchunk, using topological levels from CDAT chunk. In contrast, new versions\nof Git would use corrected commit dates, falling back to topological level\nif the generation data chunk is absent in the commit-graph file.\n\nThanks to Dr. Stolee, Dr. Narębski, and Taylor for their reviews on the\nfirst version.\n\nI look forward to everyone's reviews!\n\nThanks\n\n * Abhishek\n\n\n----------------------------------------------------------------------------\n\nChanges in version 3:\n\n * Reordered patches as discussed in 1\n   [https://lore.kernel.org/git/aee0ae56-3395-6848-d573-27a318d72755@gmail.com/]\n * Split \"implement corrected commit date\" into two patches - one\n   introducing the topo level slab and other implementing corrected commit\n   dates.\n * Extended split-commit-graph tests to verify at the end of test.\n * Use topological levels as generation number if any of split commit-graph\n   files do not have generation data chunk.\n\nChanges in version 2:\n\n * Add tests for generation data chunk.\n * Add an option GIT_TEST_COMMIT_GRAPH_NO_GDAT to control whether to write\n   generation data chunk.\n * Compare commits with corrected commit dates if present in\n   paint_down_to_common().\n * Update technical documentation.\n * Handle mixed graph version commit chains.\n * Improve commit messages for\n * Revert unnecessary whitespace changes.\n * Split uint_32 -> timestamp_t change into a new commit.\n\nAbhishek Kumar (11):\n  commit-graph: fix regression when computing bloom filter\n  revision: parse parent in indegree_walk_step()\n  commit-graph: consolidate fill_commit_graph_info\n  commit-graph: consolidate compare_commits_by_gen\n  commit-graph: return 64-bit generation number\n  commit-graph: add a slab to store topological levels\n  commit-graph: implement corrected commit date\n  commit-graph: implement generation data chunk\n  commit-graph: use generation v2 only if entire chain does\n  commit-reach: use corrected commit dates in paint_down_to_common()\n  doc: add corrected commit date info\n\n .../technical/commit-graph-format.txt         |  12 +-\n Documentation/technical/commit-graph.txt      |  45 ++--\n commit-graph.c                                | 241 +++++++++++++-----\n commit-graph.h                                |  16 +-\n commit-reach.c                                |  49 ++--\n commit-reach.h                                |   2 +-\n commit.c                                      |   9 +-\n commit.h                                      |   4 +-\n revision.c                                    |  13 +-\n t/README                                      |   3 +\n t/helper/test-read-graph.c                    |   2 +\n t/t4216-log-bloom.sh                          |   4 +-\n t/t5000-tar-tree.sh                           |   4 +-\n t/t5318-commit-graph.sh                       |  27 +-\n t/t5324-split-commit-graph.sh                 |  82 +++++-\n t/t6024-recursive-merge.sh                    |   4 +-\n t/t6600-test-reach.sh                         |  62 +++--\n upload-pack.c                                 |   2 +-\n 18 files changed, 396 insertions(+), 185 deletions(-)\n\n\nbase-commit: 7814e8a05a59c0cf5fb186661d1551c75d1299b5\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-676%2Fabhishekkumar2718%2Fcorrected_commit_date-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-676/abhishekkumar2718/corrected_commit_date-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/676\n\nRange-diff vs v2:\n\n  1:  a962b9ae4b =  1:  c6b7ade7af commit-graph: fix regression when computing bloom filter\n  2:  cf61239f93 =  2:  e673867234 revision: parse parent in indegree_walk_step()\n  3:  32da955e31 =  3:  18d5864f81 commit-graph: consolidate fill_commit_graph_info\n  4:  b254782858 !  4:  6a0cde983d commit-graph: consolidate compare_commits_by_gen\n     @@ Commit message\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n          Reviewed-by: Taylor Blau <me@ttaylorr.com>\n     +    Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n       ## commit-graph.c ##\n      @@ commit-graph.c: uint32_t commit_graph_generation(const struct commit *c)\n  6:  1aa2a00a7a =  5:  6be759a954 commit-graph: return 64-bit generation number\n  7:  bfe1473201 !  6:  b347dbb01b commit-graph: implement corrected commit date\n     @@ Metadata\n      Author: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n       ## Commit message ##\n     -    commit-graph: implement corrected commit date\n     +    commit-graph: add a slab to store topological levels\n      \n     -    With most of preparations done, let's implement corrected commit date\n     -    offset. We add a new commit-slab to store topogical levels while\n     -    writing commit graph and upgrade the generation member in struct\n     -    commit_graph_data to a 64-bit timestamp. We store topological levels to\n     -    ensure that older versions of Git will still have the performance\n     -    benefits from generation number v2.\n     +    As we are writing topological levels to commit data chunk to ensure\n     +    backwards compatibility with \"Old\" Git and the member `generation` of\n     +    struct commit_graph_data will store corrected commit date in a later\n     +    commit, let's introduce a commit-slab to store topological levels while\n     +    writing commit-graph.\n     +\n     +    When Git creates a split commit-graph, it takes advantage of the\n     +    generation values that have been computed already and present in\n     +    existing commit-graph files.\n     +\n     +    So, let's add a pointer to struct commit_graph to the topological level\n     +    commit-slab and populate it with topological levels while writing a\n     +    split commit-graph.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n     @@ commit-graph.c: void git_test_write_commit_graph_or_die(void)\n       /* Keep track of the order in which commits are added to our list. */\n       define_commit_slab(commit_pos, int);\n       static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n     -@@ commit-graph.c: static int commit_gen_cmp(const void *va, const void *vb)\n     - \telse if (generation_a > generation_b)\n     - \t\treturn 1;\n     - \n     --\t/* use date as a heuristic when generations are equal */\n     --\tif (a->date < b->date)\n     --\t\treturn -1;\n     --\telse if (a->date > b->date)\n     --\t\treturn 1;\n     - \treturn 0;\n     - }\n     - \n      @@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n       \titem->date = (timestamp_t)((date_high << 32) | date_low);\n       \n     - \tif (g->chunk_generation_data)\n     --\t\tgraph_data->generation = get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     -+\t{\n     -+\t\tgraph_data->generation = item->date +\n     -+\t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     -+\t}\n     - \telse\n     - \t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n     + \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n     ++\n     ++\tif (g->topo_levels)\n     ++\t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n       }\n     + \n     + static inline void set_commit_tree(struct commit *c, struct tree *t)\n      @@ commit-graph.c: struct write_commit_graph_context {\n     - \tstruct progress *progress;\n     - \tint progress_done;\n     - \tuint64_t progress_cnt;\n     -+\tstruct topo_level_slab *topo_levels;\n     + \t\t changed_paths:1,\n     + \t\t order_by_pack:1;\n       \n     - \tchar *base_graph_name;\n     - \tint num_commit_graphs_before;\n     ++\tstruct topo_level_slab *topo_levels;\n     + \tconst struct split_commit_graph_opts *split_opts;\n     + \tsize_t total_bloom_filter_data_size;\n     + \tconst struct bloom_filter_settings *bloom_settings;\n      @@ commit-graph.c: static int write_graph_chunk_data(struct hashfile *f,\n       \t\telse\n       \t\t\tpackedDate[0] = 0;\n     @@ commit-graph.c: static int write_graph_chunk_data(struct hashfile *f,\n       \n       \t\tpackedDate[1] = htonl((*list)->date);\n       \t\thashwrite(f, packedDate, 8);\n     -@@ commit-graph.c: static int write_graph_chunk_generation_data(struct hashfile *f,\n     - \tint i;\n     - \tfor (i = 0; i < ctx->commits.nr; i++) {\n     - \t\tstruct commit *c = ctx->commits.list[i];\n     -+\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n     - \t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n     --\t\thashwrite_be32(f, commit_graph_data_at(c)->generation);\n     -+\n     -+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX)\n     -+\t\t\toffset = GENERATION_NUMBER_V2_OFFSET_MAX;\n     -+\n     -+\t\thashwrite_be32(f, offset);\n     - \t}\n     - \n     - \treturn 0;\n      @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n       \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n       \t\t\t\t\tctx->commits.nr);\n       \tfor (i = 0; i < ctx->commits.nr; i++) {\n      -\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n     -+\t\tuint32_t topo_level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n     ++\t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n       \n       \t\tdisplay_progress(ctx->progress, i + 1);\n      -\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n      -\t\t    generation != GENERATION_NUMBER_ZERO)\n     -+\t\tif (topo_level != GENERATION_NUMBER_V1_INFINITY &&\n     -+\t\t    topo_level != GENERATION_NUMBER_ZERO)\n     ++\t\tif (level != GENERATION_NUMBER_V1_INFINITY &&\n     ++\t\t    level != GENERATION_NUMBER_ZERO)\n       \t\t\tcontinue;\n       \n       \t\tcommit_list_insert(ctx->commits.list[i], &list);\n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n       \t\t\tint all_parents_computed = 1;\n      -\t\t\tuint32_t max_generation = 0;\n      +\t\t\tuint32_t max_level = 0;\n     -+\t\t\ttimestamp_t max_corrected_commit_date = current->date - 1;\n       \n       \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n      -\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n     -+\t\t\t\ttopo_level = *topo_level_slab_at(ctx->topo_levels, parent->item);\n     ++\t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n       \n      -\t\t\t\tif (generation == GENERATION_NUMBER_V1_INFINITY ||\n      -\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n     -+\t\t\t\tif (topo_level == GENERATION_NUMBER_V1_INFINITY ||\n     -+\t\t\t\t    topo_level == GENERATION_NUMBER_ZERO) {\n     ++\t\t\t\tif (level == GENERATION_NUMBER_V1_INFINITY ||\n     ++\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n       \t\t\t\t\tall_parents_computed = 0;\n       \t\t\t\t\tcommit_list_insert(parent->item, &list);\n       \t\t\t\t\tbreak;\n      -\t\t\t\t} else if (generation > max_generation) {\n      -\t\t\t\t\tmax_generation = generation;\n     -+\t\t\t\t} else {\n     -+\t\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(parent->item);\n     -+\n     -+\t\t\t\t\tif (topo_level > max_level)\n     -+\t\t\t\t\t\tmax_level = topo_level;\n     -+\n     -+\t\t\t\t\tif (data->generation > max_corrected_commit_date)\n     -+\t\t\t\t\t\tmax_corrected_commit_date = data->generation;\n     ++\t\t\t\t} else if (level > max_level) {\n     ++\t\t\t\t\tmax_level = level;\n       \t\t\t\t}\n       \t\t\t}\n       \n       \t\t\tif (all_parents_computed) {\n     - \t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n     - \n     +-\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n     +-\n      -\t\t\t\tdata->generation = max_generation + 1;\n     --\t\t\t\tpop_commit(&list);\n     -+\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n     -+\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n     -+\n     -+\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n     -+\t\t\t\tdata->generation = max_corrected_commit_date + 1;\n     + \t\t\t\tpop_commit(&list);\n       \n      -\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n      -\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n     -+\t\t\t\tpop_commit(&list);\n     ++\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n     ++\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n     ++\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n       \t\t\t}\n       \t\t}\n       \t}\n     @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \tif (!commit_graph_compatible(the_repository))\n       \t\treturn 0;\n      @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n     - \tctx->total_bloom_filter_data_size = 0;\n     - \tctx->write_generation_data = !git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0);\n     + \t\t}\n     + \t}\n       \n      +\tinit_topo_level_slab(&topo_levels);\n      +\tctx->topo_levels = &topo_levels;\n      +\n     - \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n     - \t\tctx->changed_paths = 1;\n     - \tif (!(flags & COMMIT_GRAPH_NO_WRITE_BLOOM_FILTERS)) {\n     -@@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n     - \tfor (i = 0; i < g->num_commits; i++) {\n     - \t\tstruct commit *graph_commit, *odb_commit;\n     - \t\tstruct commit_list *graph_parents, *odb_parents;\n     --\t\ttimestamp_t max_generation = 0;\n     --\t\ttimestamp_t generation;\n     -+\t\ttimestamp_t max_parent_corrected_commit_date = 0;\n     -+\t\ttimestamp_t corrected_commit_date;\n     - \n     - \t\tdisplay_progress(progress, i + 1);\n     - \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n     -@@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n     - \t\t\t\t\t     oid_to_hex(&graph_parents->item->object.oid),\n     - \t\t\t\t\t     oid_to_hex(&odb_parents->item->object.oid));\n     - \n     --\t\t\tgeneration = commit_graph_generation(graph_parents->item);\n     --\t\t\tif (generation > max_generation)\n     --\t\t\t\tmax_generation = generation;\n     -+\t\t\tcorrected_commit_date = commit_graph_generation(graph_parents->item);\n     -+\t\t\tif (corrected_commit_date > max_parent_corrected_commit_date)\n     -+\t\t\t\tmax_parent_corrected_commit_date = corrected_commit_date;\n     - \n     - \t\t\tgraph_parents = graph_parents->next;\n     - \t\t\todb_parents = odb_parents->next;\n     -@@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n     - \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n     - \t\t\tcontinue;\n     ++\tif (ctx->r->objects->commit_graph) {\n     ++\t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n     ++\n     ++\t\twhile (g) {\n     ++\t\t\tg->topo_levels = &topo_levels;\n     ++\t\t\tg = g->base_graph;\n     ++\t\t}\n     ++\t}\n     ++\n     + \tif (pack_indexes) {\n     + \t\tctx->order_by_pack = 1;\n     + \t\tif ((res = fill_oids_from_packs(ctx, pack_indexes)))\n     +\n     + ## commit-graph.h ##\n     +@@ commit-graph.h: struct commit_graph {\n     + \tconst unsigned char *chunk_bloom_indexes;\n     + \tconst unsigned char *chunk_bloom_data;\n     + \n     ++\tstruct topo_level_slab *topo_levels;\n     + \tstruct bloom_filter_settings *bloom_filter_settings;\n     + };\n       \n     --\t\t/*\n     --\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n     --\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n     --\t\t * extra logic in the following condition.\n     --\t\t */\n     --\t\tif (max_generation == GENERATION_NUMBER_MAX)\n     --\t\t\tmax_generation--;\n     --\n     --\t\tgeneration = commit_graph_generation(graph_commit);\n     --\t\tif (generation != max_generation + 1)\n     --\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n     -+\t\tcorrected_commit_date = commit_graph_generation(graph_commit);\n     -+\t\tif (corrected_commit_date < max_parent_corrected_commit_date + 1)\n     -+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n     - \t\t\t\t     oid_to_hex(&cur_oid),\n     --\t\t\t\t     generation,\n     --\t\t\t\t     max_generation + 1);\n     -+\t\t\t\t     corrected_commit_date,\n     -+\t\t\t\t     max_parent_corrected_commit_date + 1);\n     - \n     - \t\tif (graph_commit->date != odb_commit->date)\n     - \t\t\tgraph_report(_(\"commit date for commit %s in commit-graph is %\"PRItime\" != %\"PRItime),\n      \n       ## commit.h ##\n      @@\n  -:  ---------- >  7:  4074ace65b commit-graph: implement corrected commit date\n  5:  cb797e20d7 !  8:  4e746628ac commit-graph: implement generation data chunk\n     @@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct c\n       \n      -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n      +\tif (g->chunk_generation_data)\n     -+\t\tgraph_data->generation = get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     ++\t\tgraph_data->generation = item->date +\n     ++\t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n      +\telse\n      +\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n     - }\n       \n     - static inline void set_commit_tree(struct commit *c, struct tree *t)\n     + \tif (g->topo_levels)\n     + \t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n      @@ commit-graph.c: struct write_commit_graph_context {\n       \t\t report_progress:1,\n       \t\t split:1,\n     @@ commit-graph.c: struct write_commit_graph_context {\n      +\t\t order_by_pack:1,\n      +\t\t write_generation_data:1;\n       \n     + \tstruct topo_level_slab *topo_levels;\n       \tconst struct split_commit_graph_opts *split_opts;\n     - \tsize_t total_bloom_filter_data_size;\n      @@ commit-graph.c: static int write_graph_chunk_data(struct hashfile *f,\n       \treturn 0;\n       }\n     @@ commit-graph.c: static int write_graph_chunk_data(struct hashfile *f,\n      +\tint i;\n      +\tfor (i = 0; i < ctx->commits.nr; i++) {\n      +\t\tstruct commit *c = ctx->commits.list[i];\n     ++\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n      +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n     -+\t\thashwrite_be32(f, commit_graph_data_at(c)->generation);\n     ++\n     ++\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX)\n     ++\t\t\toffset = GENERATION_NUMBER_V2_OFFSET_MAX;\n     ++\t\thashwrite_be32(f, offset);\n      +\t}\n      +\n      +\treturn 0;\n     @@ commit-graph.c: static int write_commit_graph_file(struct write_commit_graph_con\n       \tchunks[2].id = GRAPH_CHUNKID_DATA;\n       \tchunks[2].size = (hashsz + 16) * ctx->commits.nr;\n       \tchunks[2].write_fn = write_graph_chunk_data;\n     ++\n     ++\tif (git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0))\n     ++\t\tctx->write_generation_data = 0;\n      +\tif (ctx->write_generation_data) {\n      +\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA;\n      +\t\tchunks[num_chunks].size = sizeof(uint32_t) * ctx->commits.nr;\n     @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n       \tctx->split_opts = split_opts;\n       \tctx->total_bloom_filter_data_size = 0;\n     -+\tctx->write_generation_data = !git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0);\n     ++\tctx->write_generation_data = 1;\n       \n       \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n       \t\tctx->changed_paths = 1;\n     @@ t/t5318-commit-graph.sh: test_expect_success 'replace-objects invalidates commit\n      \n       ## t/t5324-split-commit-graph.sh ##\n      @@ t/t5324-split-commit-graph.sh: test_expect_success 'setup repo' '\n     + \tinfodir=\".git/objects/info\" &&\n       \tgraphdir=\"$infodir/commit-graphs\" &&\n     - \ttest_oid_init &&\n       \ttest_oid_cache <<-EOM\n      -\tshallow sha1:1760\n      -\tshallow sha256:2064\n  8:  833779ad53 !  9:  5a147a9704 commit-graph: handle mixed generation commit chains\n     @@ Metadata\n      Author: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n       ## Commit message ##\n     -    commit-graph: handle mixed generation commit chains\n     +    commit-graph: use generation v2 only if entire chain does\n      \n     -    As corrected commit dates and topological levels cannot be compared\n     -    directly, we must handle commit graph chains with mixed generation\n     -    number definitions.\n     +    Since there are released versions of Git that understand generation\n     +    numbers in the commit-graph's CDAT chunk but do not understand the GDAT\n     +    chunk, the following scenario is possible:\n      \n     -    While reading a commit graph file, we disable generation numbers if the\n     -    chain contains mixed generation numbers.\n     +    1. \"New\" Git writes a commit-graph with the GDAT chunk.\n     +    2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n      \n     -    While writing to commit graph chain, we write generation data chunk only\n     -    if the previous tip of chain had a generation data chunk. Using\n     -    `--split=replace` overwrites the existing chain and writes generation\n     -    data chunk regardless of previous tip.\n     +    Because of the current use of inspecting the current layer for a\n     +    chunk_generation_data pointer, the commits in the lower layer will be\n     +    interpreted as having very large generation values (commit date plus\n     +    offset) compared to the generation numbers in the top layer (topological\n     +    level). This violates the expectation that the generation of a parent is\n     +    strictly smaller than the generation of a child.\n      \n     -    In t5324-split-commit-graph, we set up a repo with twelve commits and\n     -    write a base commit graph file with no generation data chunk. When add\n     -    three commits and write to chain again, Git does not write generation\n     -    data chunk even without setting GIT_TEST_COMMIT_GRAPH_NO_GDAT=1. Then,\n     -    as we replace the existing chain, Git writes a commit graph file with\n     -    generation data chunk.\n     +    It is difficult to expose this issue in a test. Since we _start_ with\n     +    artificially low generation numbers, any commit walk that prioritizes\n     +    generation numbers will walk all of the commits with high generation\n     +    number before walking the commits with low generation number. In all the\n     +    cases I tried, the commit-graph layers themselves \"protect\" any\n     +    incorrect behavior since none of the commits in the lower layer can\n     +    reach the commits in the upper layer.\n      \n     +    This issue would manifest itself as a performance problem in this case,\n     +    especially with something like \"git log --graph\" since the low\n     +    generation numbers would cause the in-degree queue to walk all of the\n     +    commits in the lower layer before allowing the topo-order queue to write\n     +    anything to output (depending on the size of the upper layer).\n     +\n     +    Signed-off-by: Derrick Stolee <dstolee@microsoft.com>\n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n       ## commit-graph.c ##\n     -@@ commit-graph.c: int generation_numbers_enabled(struct repository *r)\n     - \tif (!g->num_commits)\n     - \t\treturn 0;\n     +@@ commit-graph.c: static struct commit_graph *load_commit_graph_chain(struct repository *r,\n     + \treturn graph_chain;\n     + }\n       \n     -+\t/* We cannot compare topological levels and corrected commit dates */\n     -+\twhile (g->base_graph) {\n     -+\t\twarning(_(\"commit-graph-chain contains mixed generation versions\"));\n     -+\t\tif ((g->chunk_generation_data == NULL) ^ (g->base_graph->chunk_generation_data == NULL))\n     -+\t\t\treturn 0;\n     ++static void validate_mixed_generation_chain(struct repository *r)\n     ++{\n     ++\tstruct commit_graph *g = r->objects->commit_graph;\n     ++\tint read_generation_data = 1;\n     ++\n     ++\twhile (g) {\n     ++\t\tif (!g->chunk_generation_data) {\n     ++\t\t\tread_generation_data = 0;\n     ++\t\t\tbreak;\n     ++\t\t}\n      +\t\tg = g->base_graph;\n      +\t}\n      +\n     - \tfirst_generation = get_be32(g->chunk_commit_data +\n     - \t\t\t\t    g->hash_len + 8) >> 2;\n     ++\tg = r->objects->commit_graph;\n     ++\n     ++\twhile (g) {\n     ++\t\tg->read_generation_data = read_generation_data;\n     ++\t\tg = g->base_graph;\n     ++\t}\n     ++}\n     ++\n     + struct commit_graph *read_commit_graph_one(struct repository *r,\n     + \t\t\t\t\t   struct object_directory *odb)\n     + {\n     +@@ commit-graph.c: struct commit_graph *read_commit_graph_one(struct repository *r,\n     + \tif (!g)\n     + \t\tg = load_commit_graph_chain(r, odb);\n     + \n     ++\tvalidate_mixed_generation_chain(r);\n     ++\n     + \treturn g;\n     + }\n     + \n     +@@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n     + \tdate_low = get_be32(commit_data + g->hash_len + 12);\n     + \titem->date = (timestamp_t)((date_high << 32) | date_low);\n       \n     +-\tif (g->chunk_generation_data)\n     ++\tif (g->chunk_generation_data && g->read_generation_data)\n     + \t\tgraph_data->generation = item->date +\n     + \t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     + \telse\n     +@@ commit-graph.c: void load_commit_graph_info(struct repository *r, struct commit *item)\n     + \tuint32_t pos;\n     + \tif (!prepare_commit_graph(r))\n     + \t\treturn;\n     ++\n     + \tif (find_commit_in_graph(item, r->objects->commit_graph, &pos))\n     + \t\tfill_commit_graph_info(item, r->objects->commit_graph, pos);\n     + }\n      @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \n       \t\tg = ctx->r->objects->commit_graph;\n     @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \n       \tctx->approx_nr_objects = approximate_object_count();\n      \n     + ## commit-graph.h ##\n     +@@ commit-graph.h: struct commit_graph {\n     + \tstruct object_directory *odb;\n     + \n     + \tuint32_t num_commits_in_base;\n     ++\tuint32_t read_generation_data;\n     + \tstruct commit_graph *base_graph;\n     + \n     + \tconst uint32_t *chunk_oid_fanout;\n     +\n       ## t/t5324-split-commit-graph.sh ##\n      @@ t/t5324-split-commit-graph.sh: done <<\\EOF\n       0600 -r--------\n     @@ t/t5324-split-commit-graph.sh: done <<\\EOF\n      +\t\ttest_commit $i &&\n      +\t\tgit branch commits/$i || return 1\n      +\tdone &&\n     ++\tgit commit-graph write --reachable --split &&\n      +\tgit reset --hard commits/2 &&\n      +\tgit merge commits/4 &&\n      +\tgit branch merge/1 &&\n      +\tgit reset --hard commits/4 &&\n      +\tgit merge commits/6 &&\n      +\tgit branch merge/2 &&\n     -+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split &&\n     ++\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n      +\ttest-tool read-graph >output &&\n      +\tcat >expect <<-EOF &&\n     -+\theader: 43475048 1 1 3 0\n     -+\tnum_commits: 12\n     ++\theader: 43475048 1 1 4 1\n     ++\tnum_commits: 2\n      +\tchunks: oid_fanout oid_lookup commit_metadata\n      +\tEOF\n     -+\ttest_cmp expect output\n     ++\ttest_cmp expect output &&\n     ++\tgit commit-graph verify\n      +'\n      +\n      +test_expect_success 'does not write generation data chunk if not present on existing tip' '\n     @@ t/t5324-split-commit-graph.sh: done <<\\EOF\n      +\tgit merge commits/5 &&\n      +\tgit merge merge/2 &&\n      +\tgit branch merge/3 &&\n     -+\tgit commit-graph write --reachable --split &&\n     ++\tgit commit-graph write --reachable --split=no-merge &&\n      +\ttest-tool read-graph >output &&\n      +\tcat >expect <<-EOF &&\n     -+\theader: 43475048 1 1 4 1\n     ++\theader: 43475048 1 1 4 2\n      +\tnum_commits: 3\n      +\tchunks: oid_fanout oid_lookup commit_metadata\n      +\tEOF\n     -+\ttest_cmp expect output\n     ++\ttest_cmp expect output &&\n     ++\tgit commit-graph verify\n      +'\n      +\n      +test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n      +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n     -+\tgit commit-graph write --reachable --split='replace' &&\n     ++\tgit commit-graph write --reachable --split=replace &&\n      +\ttest_path_is_file $graphdir/commit-graph-chain &&\n      +\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n      +\tverify_chain_files_exist $graphdir &&\n     -+\tgraph_read_expect 15\n     ++\tgraph_read_expect 15 &&\n     ++\tgit commit-graph verify\n      +'\n      +\n       test_done\n  9:  58a2d5da01 = 10:  439adc1718 commit-reach: use corrected commit dates in paint_down_to_common()\n 10:  4c34294602 = 11:  f6f91af305 doc: add corrected commit date info\n\n-- \ngitgitgadget\n"},{"id":"403754","messageId":"439adc1718d6cc37f18c1eaeafd605f5c2961733.1597509583.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v3 10/11] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:42Z","receivedAt":"2020-08-15T21:57:56Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith corrected commit dates implemented, we no longer have to rely on\ncommit date as a heuristic in paint_down_to_common().\n\nt6024-recursive-merge setups a unique repository where all commits have\nthe same committer date without well-defined merge-base. As this has\nalready caused problems (as noted in 859fdc0 (commit-graph: define\nGIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph within the\ntest script.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c             | 14 ++++++++++++++\n commit-graph.h             |  6 ++++++\n commit-reach.c             |  2 +-\n t/t6024-recursive-merge.sh |  4 +++-\n 4 files changed, 24 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex c1292f8e08..6411068411 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -703,6 +703,20 @@ int generation_numbers_enabled(struct repository *r)\n \treturn !!first_generation;\n }\n \n+int corrected_commit_dates_enabled(struct repository *r)\n+{\n+\tstruct commit_graph *g;\n+\tif (!prepare_commit_graph(r))\n+\t\treturn 0;\n+\n+\tg = r->objects->commit_graph;\n+\n+\tif (!g->num_commits)\n+\t\treturn 0;\n+\n+\treturn !!g->chunk_generation_data;\n+}\n+\n static void close_commit_graph_one(struct commit_graph *g)\n {\n \tif (!g)\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 3cf89d895d..e22ec1e626 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -91,6 +91,12 @@ struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size);\n  */\n int generation_numbers_enabled(struct repository *r);\n \n+/*\n+ * Return 1 if and only if the repository has a commit-graph\n+ * file and generation data chunk has been written for the file.\n+ */\n+int corrected_commit_dates_enabled(struct repository *r);\n+\n enum commit_graph_write_flags {\n \tCOMMIT_GRAPH_WRITE_APPEND     = (1 << 0),\n \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\ndiff --git a/commit-reach.c b/commit-reach.c\nindex 470bc80139..3a1b925274 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -39,7 +39,7 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \tint i;\n \ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n-\tif (!min_generation)\n+\tif (!min_generation && !corrected_commit_dates_enabled(r))\n \t\tqueue.compare = compare_commits_by_commit_date;\n \n \tone->object.flags |= PARENT1;\ndiff --git a/t/t6024-recursive-merge.sh b/t/t6024-recursive-merge.sh\nindex 332cfc53fd..d3def66e7d 100755\n--- a/t/t6024-recursive-merge.sh\n+++ b/t/t6024-recursive-merge.sh\n@@ -15,6 +15,8 @@ GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n export GIT_COMMITTER_DATE\n \n test_expect_success 'setup tests' '\n+\tGIT_TEST_COMMIT_GRAPH=0 &&\n+\texport GIT_TEST_COMMIT_GRAPH &&\n \techo 1 >a1 &&\n \tgit add a1 &&\n \tGIT_AUTHOR_DATE=\"2006-12-12 23:00:00\" git commit -m 1 a1 &&\n@@ -66,7 +68,7 @@ test_expect_success 'setup tests' '\n '\n \n test_expect_success 'combined merge conflicts' '\n-\ttest_must_fail env GIT_TEST_COMMIT_GRAPH=0 git merge -m final G\n+\ttest_must_fail git merge -m final G\n '\n \n test_expect_success 'result contains a conflict' '\n-- \ngitgitgadget\n\n"},{"id":"403759","messageId":"5a147a9704f0f8d8644c92ea38583e966378b931.1597509583.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v3 09/11] commit-graph: use generation v2 only if entire chain does","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:41Z","receivedAt":"2020-08-15T21:59:11Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nSince there are released versions of Git that understand generation\nnumbers in the commit-graph's CDAT chunk but do not understand the GDAT\nchunk, the following scenario is possible:\n\n1. \"New\" Git writes a commit-graph with the GDAT chunk.\n2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n\nBecause of the current use of inspecting the current layer for a\nchunk_generation_data pointer, the commits in the lower layer will be\ninterpreted as having very large generation values (commit date plus\noffset) compared to the generation numbers in the top layer (topological\nlevel). This violates the expectation that the generation of a parent is\nstrictly smaller than the generation of a child.\n\nIt is difficult to expose this issue in a test. Since we _start_ with\nartificially low generation numbers, any commit walk that prioritizes\ngeneration numbers will walk all of the commits with high generation\nnumber before walking the commits with low generation number. In all the\ncases I tried, the commit-graph layers themselves \"protect\" any\nincorrect behavior since none of the commits in the lower layer can\nreach the commits in the upper layer.\n\nThis issue would manifest itself as a performance problem in this case,\nespecially with something like \"git log --graph\" since the low\ngeneration numbers would cause the in-degree queue to walk all of the\ncommits in the lower layer before allowing the topo-order queue to write\nanything to output (depending on the size of the upper layer).\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                | 32 +++++++++++++++-\n commit-graph.h                |  1 +\n t/t5324-split-commit-graph.sh | 70 +++++++++++++++++++++++++++++++++++\n 3 files changed, 102 insertions(+), 1 deletion(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex b7a72b40db..c1292f8e08 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -597,6 +597,27 @@ static struct commit_graph *load_commit_graph_chain(struct repository *r,\n \treturn graph_chain;\n }\n \n+static void validate_mixed_generation_chain(struct repository *r)\n+{\n+\tstruct commit_graph *g = r->objects->commit_graph;\n+\tint read_generation_data = 1;\n+\n+\twhile (g) {\n+\t\tif (!g->chunk_generation_data) {\n+\t\t\tread_generation_data = 0;\n+\t\t\tbreak;\n+\t\t}\n+\t\tg = g->base_graph;\n+\t}\n+\n+\tg = r->objects->commit_graph;\n+\n+\twhile (g) {\n+\t\tg->read_generation_data = read_generation_data;\n+\t\tg = g->base_graph;\n+\t}\n+}\n+\n struct commit_graph *read_commit_graph_one(struct repository *r,\n \t\t\t\t\t   struct object_directory *odb)\n {\n@@ -605,6 +626,8 @@ struct commit_graph *read_commit_graph_one(struct repository *r,\n \tif (!g)\n \t\tg = load_commit_graph_chain(r, odb);\n \n+\tvalidate_mixed_generation_chain(r);\n+\n \treturn g;\n }\n \n@@ -763,7 +786,7 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tif (g->chunk_generation_data)\n+\tif (g->chunk_generation_data && g->read_generation_data)\n \t\tgraph_data->generation = item->date +\n \t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n \telse\n@@ -885,6 +908,7 @@ void load_commit_graph_info(struct repository *r, struct commit *item)\n \tuint32_t pos;\n \tif (!prepare_commit_graph(r))\n \t\treturn;\n+\n \tif (find_commit_in_graph(item, r->objects->commit_graph, &pos))\n \t\tfill_commit_graph_info(item, r->objects->commit_graph, pos);\n }\n@@ -2192,6 +2216,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \t\tg = ctx->r->objects->commit_graph;\n \n+\t\tif (g && !g->chunk_generation_data)\n+\t\t\tctx->write_generation_data = 0;\n+\n \t\twhile (g) {\n \t\t\tctx->num_commit_graphs_before++;\n \t\t\tg = g->base_graph;\n@@ -2210,6 +2237,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \t\tif (ctx->split_opts)\n \t\t\treplace = ctx->split_opts->flags & COMMIT_GRAPH_SPLIT_REPLACE;\n+\n+\t\tif (replace)\n+\t\t\tctx->write_generation_data = 1;\n \t}\n \n \tctx->approx_nr_objects = approximate_object_count();\ndiff --git a/commit-graph.h b/commit-graph.h\nindex f78c892fc0..3cf89d895d 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -63,6 +63,7 @@ struct commit_graph {\n \tstruct object_directory *odb;\n \n \tuint32_t num_commits_in_base;\n+\tuint32_t read_generation_data;\n \tstruct commit_graph *base_graph;\n \n \tconst uint32_t *chunk_oid_fanout;\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 531016f405..ac5e7783fb 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -424,4 +424,74 @@ done <<\\EOF\n 0600 -r--------\n EOF\n \n+test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n+\tmkdir mixed &&\n+\tgraphdir=\".git/objects/info/commit-graphs\" &&\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit init &&\n+\tgit config core.commitGraph true &&\n+\tgit config gc.writeCommitGraph false &&\n+\tfor i in $(test_seq 3)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit reset --hard commits/1 &&\n+\tfor i in $(test_seq 4 5)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit reset --hard commits/2 &&\n+\tfor i in $(test_seq 6 10)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit commit-graph write --reachable --split &&\n+\tgit reset --hard commits/2 &&\n+\tgit merge commits/4 &&\n+\tgit branch merge/1 &&\n+\tgit reset --hard commits/4 &&\n+\tgit merge commits/6 &&\n+\tgit branch merge/2 &&\n+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n+\ttest-tool read-graph >output &&\n+\tcat >expect <<-EOF &&\n+\theader: 43475048 1 1 4 1\n+\tnum_commits: 2\n+\tchunks: oid_fanout oid_lookup commit_metadata\n+\tEOF\n+\ttest_cmp expect output &&\n+\tgit commit-graph verify\n+'\n+\n+test_expect_success 'does not write generation data chunk if not present on existing tip' '\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit reset --hard commits/3 &&\n+\tgit merge merge/1 &&\n+\tgit merge commits/5 &&\n+\tgit merge merge/2 &&\n+\tgit branch merge/3 &&\n+\tgit commit-graph write --reachable --split=no-merge &&\n+\ttest-tool read-graph >output &&\n+\tcat >expect <<-EOF &&\n+\theader: 43475048 1 1 4 2\n+\tnum_commits: 3\n+\tchunks: oid_fanout oid_lookup commit_metadata\n+\tEOF\n+\ttest_cmp expect output &&\n+\tgit commit-graph verify\n+'\n+\n+test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit commit-graph write --reachable --split=replace &&\n+\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\tverify_chain_files_exist $graphdir &&\n+\tgraph_read_expect 15 &&\n+\tgit commit-graph verify\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"403761","messageId":"4e746628acdb49af5e8eb788864156f54724d4fa.1597509583.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v3 08/11] commit-graph: implement generation data chunk","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:40Z","receivedAt":"2020-08-15T22:00:21Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nAs discovered by Ævar, we cannot increment graph version to\ndistinguish between generation numbers v1 and v2 [1]. Thus, one of\npre-requistes before implementing generation number was to distinguish\nbetween graph versions in a backwards compatible manner.\n\nWe are going to introduce a new chunk called Generation Data chunk (or\nGDAT). GDAT stores generation number v2 (and any subsequent versions),\nwhereas CDAT will still store topological level.\n\nOld Git does not understand GDAT chunk and would ignore it, reading\ntopological levels from CDAT. New Git can parse GDAT and take advantage\nof newer generation numbers, falling back to topological levels when\nGDAT chunk is missing (as it would happen with a commit graph written\nby old Git).\n\nWe introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\nwhich forces commit-graph file to be written without generation data\nchunk to emulate a commit-graph file written by old Git.\n\n[1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                | 48 ++++++++++++++++++++++++---\n commit-graph.h                |  2 ++\n t/README                      |  3 ++\n t/helper/test-read-graph.c    |  2 ++\n t/t4216-log-bloom.sh          |  4 +--\n t/t5318-commit-graph.sh       | 27 +++++++--------\n t/t5324-split-commit-graph.sh | 12 +++----\n t/t6600-test-reach.sh         | 62 +++++++++++++++++++----------------\n 8 files changed, 107 insertions(+), 53 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex fd69534dd5..b7a72b40db 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -38,11 +38,12 @@ void git_test_write_commit_graph_or_die(void)\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n+#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n-#define MAX_NUM_CHUNKS 7\n+#define MAX_NUM_CHUNKS 8\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -389,6 +390,13 @@ struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size)\n \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n \t\t\tbreak;\n \n+\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n+\t\t\tif (graph->chunk_generation_data)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n+\t\t\tbreak;\n+\n \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n \t\t\tif (graph->chunk_extra_edges)\n \t\t\t\tchunk_repeated = 1;\n@@ -755,7 +763,11 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\tif (g->chunk_generation_data)\n+\t\tgraph_data->generation = item->date +\n+\t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n+\telse\n+\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n \n \tif (g->topo_levels)\n \t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n@@ -951,7 +963,8 @@ struct write_commit_graph_context {\n \t\t report_progress:1,\n \t\t split:1,\n \t\t changed_paths:1,\n-\t\t order_by_pack:1;\n+\t\t order_by_pack:1,\n+\t\t write_generation_data:1;\n \n \tstruct topo_level_slab *topo_levels;\n \tconst struct split_commit_graph_opts *split_opts;\n@@ -1106,8 +1119,25 @@ static int write_graph_chunk_data(struct hashfile *f,\n \treturn 0;\n }\n \n+static int write_graph_chunk_generation_data(struct hashfile *f,\n+\t\t\t\t\t      struct write_commit_graph_context *ctx)\n+{\n+\tint i;\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n+\n+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX)\n+\t\t\toffset = GENERATION_NUMBER_V2_OFFSET_MAX;\n+\t\thashwrite_be32(f, offset);\n+\t}\n+\n+\treturn 0;\n+}\n+\n static int write_graph_chunk_extra_edges(struct hashfile *f,\n-\t\t\t\t\t struct write_commit_graph_context *ctx)\n+\t\t\t\t\t  struct write_commit_graph_context *ctx)\n {\n \tstruct commit **list = ctx->commits.list;\n \tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n@@ -1726,6 +1756,15 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tchunks[2].id = GRAPH_CHUNKID_DATA;\n \tchunks[2].size = (hashsz + 16) * ctx->commits.nr;\n \tchunks[2].write_fn = write_graph_chunk_data;\n+\n+\tif (git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0))\n+\t\tctx->write_generation_data = 0;\n+\tif (ctx->write_generation_data) {\n+\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA;\n+\t\tchunks[num_chunks].size = sizeof(uint32_t) * ctx->commits.nr;\n+\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_extra_edges) {\n \t\tchunks[num_chunks].id = GRAPH_CHUNKID_EXTRAEDGES;\n \t\tchunks[num_chunks].size = 4 * ctx->num_extra_edges;\n@@ -2130,6 +2169,7 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n \tctx->split_opts = split_opts;\n \tctx->total_bloom_filter_data_size = 0;\n+\tctx->write_generation_data = 1;\n \n \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n \t\tctx->changed_paths = 1;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 1152a9642e..f78c892fc0 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -6,6 +6,7 @@\n #include \"oidset.h\"\n \n #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n+#define GIT_TEST_COMMIT_GRAPH_NO_GDAT \"GIT_TEST_COMMIT_GRAPH_NO_GDAT\"\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE \"GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE\"\n #define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n \n@@ -67,6 +68,7 @@ struct commit_graph {\n \tconst uint32_t *chunk_oid_fanout;\n \tconst unsigned char *chunk_oid_lookup;\n \tconst unsigned char *chunk_commit_data;\n+\tconst unsigned char *chunk_generation_data;\n \tconst unsigned char *chunk_extra_edges;\n \tconst unsigned char *chunk_base_graphs;\n \tconst unsigned char *chunk_bloom_indexes;\ndiff --git a/t/README b/t/README\nindex 70ec61cf88..6647ef132e 100644\n--- a/t/README\n+++ b/t/README\n@@ -379,6 +379,9 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n+commit-graph to be written without generation data chunk.\n+\n GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n commit-graph write to compute and write changed path Bloom filters for\n every 'git commit-graph write', as if the `--changed-paths` option was\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex 6d0c962438..1c2a5366c7 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -32,6 +32,8 @@ int cmd__read_graph(int argc, const char **argv)\n \t\tprintf(\" oid_lookup\");\n \tif (graph->chunk_commit_data)\n \t\tprintf(\" commit_metadata\");\n+\tif (graph->chunk_generation_data)\n+\t\tprintf(\" generation_data\");\n \tif (graph->chunk_extra_edges)\n \t\tprintf(\" extra_edges\");\n \tif (graph->chunk_bloom_indexes)\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex c21cc160f3..55c94e9ebd 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -33,11 +33,11 @@ test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n \tgit commit-graph write --reachable --changed-paths\n '\n graph_read_expect () {\n-\tNUM_CHUNKS=5\n+\tNUM_CHUNKS=6\n \tcat >expect <<- EOF\n \theader: 43475048 1 1 $NUM_CHUNKS 0\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n \tEOF\n \ttest-tool read-graph >actual &&\n \ttest_cmp expect actual\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 044cf8a3de..b41b2160c6 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -71,7 +71,7 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n graph_read_expect() {\n \tOPTIONAL=\"\"\n \tNUM_CHUNKS=3\n-\tif test ! -z $2\n+\tif test ! -z \"$2\"\n \tthen\n \t\tOPTIONAL=\" $2\"\n \t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n@@ -98,14 +98,14 @@ test_expect_success 'exit with correct error on bad input to --stdin-commits' '\n \t# valid commit and tree OID\n \tgit rev-parse HEAD HEAD^{tree} >in &&\n \tgit commit-graph write --stdin-commits <in &&\n-\tgraph_read_expect 3\n+\tgraph_read_expect 3 generation_data\n '\n \n test_expect_success 'write graph' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"3\"\n+\tgraph_read_expect \"3\" generation_data\n '\n \n test_expect_success POSIXPERM 'write graph has correct permissions' '\n@@ -214,7 +214,7 @@ test_expect_success 'write graph with merges' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'merge 1 vs 2' full merge/1 merge/2\n@@ -249,7 +249,7 @@ test_expect_success 'write graph with new commit' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'full graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -259,7 +259,7 @@ test_expect_success 'write graph with nothing new' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'cleared graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -269,7 +269,7 @@ test_expect_success 'build graph from latest pack with closure' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tcat new-idx | git commit-graph write --stdin-packs &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"9\" \"extra_edges\"\n+\tgraph_read_expect \"9\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'graph from pack, commit 8 vs merge 1' full commits/8 merge/1\n@@ -282,7 +282,7 @@ test_expect_success 'build graph from commits with closure' '\n \tgit rev-parse merge/1 >>commits-in &&\n \tcat commits-in | git commit-graph write --stdin-commits &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"6\"\n+\tgraph_read_expect \"6\" \"generation_data\"\n '\n \n graph_git_behavior 'graph from commits, commit 8 vs merge 1' full commits/8 merge/1\n@@ -292,7 +292,7 @@ test_expect_success 'build graph from commits with append' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit rev-parse merge/3 | git commit-graph write --stdin-commits --append &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -302,7 +302,7 @@ test_expect_success 'build graph using --reachable' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write --reachable &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -323,7 +323,7 @@ test_expect_success 'write graph in bare repo' '\n \tcd \"$TRASH_DIRECTORY/bare\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $baredir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n@@ -420,8 +420,9 @@ test_expect_success 'replace-objects invalidates commit-graph' '\n \n test_expect_success 'git commit-graph verify' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n-\tgit rev-parse commits/8 | git commit-graph write --stdin-commits &&\n-\tgit commit-graph verify >output\n+\tgit rev-parse commits/8 | GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --stdin-commits &&\n+\tgit commit-graph verify >output &&\n+\tgraph_read_expect 9 extra_edges\n '\n \n NUM_COMMITS=9\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex ea28d522b8..531016f405 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -13,11 +13,11 @@ test_expect_success 'setup repo' '\n \tinfodir=\".git/objects/info\" &&\n \tgraphdir=\"$infodir/commit-graphs\" &&\n \ttest_oid_cache <<-EOM\n-\tshallow sha1:1760\n-\tshallow sha256:2064\n+\tshallow sha1:2132\n+\tshallow sha256:2436\n \n-\tbase sha1:1376\n-\tbase sha256:1496\n+\tbase sha1:1408\n+\tbase sha256:1528\n \tEOM\n '\n \n@@ -28,9 +28,9 @@ graph_read_expect() {\n \t\tNUM_BASE=$2\n \tfi\n \tcat >expect <<- EOF\n-\theader: 43475048 1 1 3 $NUM_BASE\n+\theader: 43475048 1 1 4 $NUM_BASE\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n \tEOF\n \ttest-tool read-graph >output &&\n \ttest_cmp expect output\ndiff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\nindex 475564bee7..d14b129f06 100755\n--- a/t/t6600-test-reach.sh\n+++ b/t/t6600-test-reach.sh\n@@ -55,10 +55,13 @@ test_expect_success 'setup' '\n \tgit show-ref -s commit-5-5 | git commit-graph write --stdin-commits &&\n \tmv .git/objects/info/commit-graph commit-graph-half &&\n \tchmod u+w commit-graph-half &&\n+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable &&\n+\tmv .git/objects/info/commit-graph commit-graph-no-gdat &&\n+\tchmod u+w commit-graph-no-gdat &&\n \tgit config core.commitGraph true\n '\n \n-run_three_modes () {\n+run_all_modes () {\n \ttest_when_finished rm -rf .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n \ttest_cmp expect actual &&\n@@ -67,11 +70,14 @@ run_three_modes () {\n \ttest_cmp expect actual &&\n \tcp commit-graph-half .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n+\ttest_cmp expect actual &&\n+\tcp commit-graph-no-gdat .git/objects/info/commit-graph &&\n+\t\"$@\" <input >actual &&\n \ttest_cmp expect actual\n }\n \n-test_three_modes () {\n-\trun_three_modes test-tool reach \"$@\"\n+test_all_modes () {\n+\trun_all_modes test-tool reach \"$@\"\n }\n \n test_expect_success 'ref_newer:miss' '\n@@ -80,7 +86,7 @@ test_expect_success 'ref_newer:miss' '\n \tB:commit-4-9\n \tEOF\n \techo \"ref_newer(A,B):0\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'ref_newer:hit' '\n@@ -89,7 +95,7 @@ test_expect_success 'ref_newer:hit' '\n \tB:commit-2-3\n \tEOF\n \techo \"ref_newer(A,B):1\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'in_merge_bases:hit' '\n@@ -98,7 +104,7 @@ test_expect_success 'in_merge_bases:hit' '\n \tB:commit-8-8\n \tEOF\n \techo \"in_merge_bases(A,B):1\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'in_merge_bases:miss' '\n@@ -107,7 +113,7 @@ test_expect_success 'in_merge_bases:miss' '\n \tB:commit-5-9\n \tEOF\n \techo \"in_merge_bases(A,B):0\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'is_descendant_of:hit' '\n@@ -118,7 +124,7 @@ test_expect_success 'is_descendant_of:hit' '\n \tX:commit-1-1\n \tEOF\n \techo \"is_descendant_of(A,X):1\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'is_descendant_of:miss' '\n@@ -129,7 +135,7 @@ test_expect_success 'is_descendant_of:miss' '\n \tX:commit-7-6\n \tEOF\n \techo \"is_descendant_of(A,X):0\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'get_merge_bases_many' '\n@@ -144,7 +150,7 @@ test_expect_success 'get_merge_bases_many' '\n \t\tgit rev-parse commit-5-6 \\\n \t\t\t      commit-4-7 | sort\n \t} >expect &&\n-\ttest_three_modes get_merge_bases_many\n+\ttest_all_modes get_merge_bases_many\n '\n \n test_expect_success 'reduce_heads' '\n@@ -166,7 +172,7 @@ test_expect_success 'reduce_heads' '\n \t\t\t      commit-2-8 \\\n \t\t\t      commit-1-10 | sort\n \t} >expect &&\n-\ttest_three_modes reduce_heads\n+\ttest_all_modes reduce_heads\n '\n \n test_expect_success 'can_all_from_reach:hit' '\n@@ -189,7 +195,7 @@ test_expect_success 'can_all_from_reach:hit' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach(X,Y):1\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach:miss' '\n@@ -211,7 +217,7 @@ test_expect_success 'can_all_from_reach:miss' '\n \tY:commit-8-5\n \tEOF\n \techo \"can_all_from_reach(X,Y):0\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach_with_flag: tags case' '\n@@ -234,7 +240,7 @@ test_expect_success 'can_all_from_reach_with_flag: tags case' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach_with_flag(X,_,_,0,0):1\" >expect &&\n-\ttest_three_modes can_all_from_reach_with_flag\n+\ttest_all_modes can_all_from_reach_with_flag\n '\n \n test_expect_success 'commit_contains:hit' '\n@@ -250,8 +256,8 @@ test_expect_success 'commit_contains:hit' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):1\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'commit_contains:miss' '\n@@ -267,8 +273,8 @@ test_expect_success 'commit_contains:miss' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):0\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'rev-list: basic topo-order' '\n@@ -280,7 +286,7 @@ test_expect_success 'rev-list: basic topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 commit-3-2 commit-2-2 commit-1-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent topo-order' '\n@@ -292,7 +298,7 @@ test_expect_success 'rev-list: first-parent topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -304,7 +310,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -316,7 +322,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent range topo-order' '\n@@ -328,7 +334,7 @@ test_expect_success 'rev-list: first-parent range topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: ancestry-path topo-order' '\n@@ -338,7 +344,7 @@ test_expect_success 'rev-list: ancestry-path topo-order' '\n \t\tcommit-6-4 commit-5-4 commit-4-4 commit-3-4 \\\n \t\tcommit-6-3 commit-5-3 commit-4-3 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: symmetric difference topo-order' '\n@@ -352,7 +358,7 @@ test_expect_success 'rev-list: symmetric difference topo-order' '\n \t\tcommit-3-8 commit-2-8 commit-1-8 \\\n \t\tcommit-3-7 commit-2-7 commit-1-7 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8...commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8...commit-6-6\n '\n \n test_expect_success 'get_reachable_subset:all' '\n@@ -372,7 +378,7 @@ test_expect_success 'get_reachable_subset:all' '\n \t\t\t      commit-1-7 \\\n \t\t\t      commit-5-6 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:some' '\n@@ -390,7 +396,7 @@ test_expect_success 'get_reachable_subset:some' '\n \t\tgit rev-parse commit-3-3 \\\n \t\t\t      commit-1-7 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:none' '\n@@ -404,7 +410,7 @@ test_expect_success 'get_reachable_subset:none' '\n \tY:commit-2-8\n \tEOF\n \techo \"get_reachable_subset(X,Y)\" >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_done\n-- \ngitgitgadget\n\n"},{"id":"403762","messageId":"b347dbb01b9254ab8d79fbbd0f7c2b637efde62e.1597509583.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v3 06/11] commit-graph: add a slab to store topological levels","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:38Z","receivedAt":"2020-08-15T22:01:21Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nAs we are writing topological levels to commit data chunk to ensure\nbackwards compatibility with \"Old\" Git and the member `generation` of\nstruct commit_graph_data will store corrected commit date in a later\ncommit, let's introduce a commit-slab to store topological levels while\nwriting commit-graph.\n\nWhen Git creates a split commit-graph, it takes advantage of the\ngeneration values that have been computed already and present in\nexisting commit-graph files.\n\nSo, let's add a pointer to struct commit_graph to the topological level\ncommit-slab and populate it with topological levels while writing a\nsplit commit-graph.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 47 ++++++++++++++++++++++++++++++++---------------\n commit-graph.h |  1 +\n commit.h       |  1 +\n 3 files changed, 34 insertions(+), 15 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 7f9f858577..a2f15b2825 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -64,6 +64,8 @@ void git_test_write_commit_graph_or_die(void)\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n+define_commit_slab(topo_level_slab, uint32_t);\n+\n /* Keep track of the order in which commits are added to our list. */\n define_commit_slab(commit_pos, int);\n static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n@@ -759,6 +761,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\n+\tif (g->topo_levels)\n+\t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n static inline void set_commit_tree(struct commit *c, struct tree *t)\n@@ -953,6 +958,7 @@ struct write_commit_graph_context {\n \t\t changed_paths:1,\n \t\t order_by_pack:1;\n \n+\tstruct topo_level_slab *topo_levels;\n \tconst struct split_commit_graph_opts *split_opts;\n \tsize_t total_bloom_filter_data_size;\n \tconst struct bloom_filter_settings *bloom_settings;\n@@ -1094,7 +1100,7 @@ static int write_graph_chunk_data(struct hashfile *f,\n \t\telse\n \t\t\tpackedDate[0] = 0;\n \n-\t\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n+\t\tpackedDate[0] |= htonl(*topo_level_slab_at(ctx->topo_levels, *list) << 2);\n \n \t\tpackedDate[1] = htonl((*list)->date);\n \t\thashwrite(f, packedDate, 8);\n@@ -1335,11 +1341,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n+\t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n-\t\t    generation != GENERATION_NUMBER_ZERO)\n+\t\tif (level != GENERATION_NUMBER_V1_INFINITY &&\n+\t\t    level != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1347,29 +1353,27 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit *current = list->item;\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n-\t\t\tuint32_t max_generation = 0;\n+\t\t\tuint32_t max_level = 0;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n-\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n+\t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n \n-\t\t\t\tif (generation == GENERATION_NUMBER_V1_INFINITY ||\n-\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n+\t\t\t\tif (level == GENERATION_NUMBER_V1_INFINITY ||\n+\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (generation > max_generation) {\n-\t\t\t\t\tmax_generation = generation;\n+\t\t\t\t} else if (level > max_level) {\n+\t\t\t\t\tmax_level = level;\n \t\t\t\t}\n \t\t\t}\n \n \t\t\tif (all_parents_computed) {\n-\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n-\n-\t\t\t\tdata->generation = max_generation + 1;\n \t\t\t\tpop_commit(&list);\n \n-\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n-\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n+\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n+\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n+\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n \t\t\t}\n \t\t}\n \t}\n@@ -2101,6 +2105,7 @@ int write_commit_graph(struct object_directory *odb,\n \tuint32_t i, count_distinct = 0;\n \tint res = 0;\n \tint replace = 0;\n+\tstruct topo_level_slab topo_levels;\n \n \tif (!commit_graph_compatible(the_repository))\n \t\treturn 0;\n@@ -2179,6 +2184,18 @@ int write_commit_graph(struct object_directory *odb,\n \t\t}\n \t}\n \n+\tinit_topo_level_slab(&topo_levels);\n+\tctx->topo_levels = &topo_levels;\n+\n+\tif (ctx->r->objects->commit_graph) {\n+\t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n+\n+\t\twhile (g) {\n+\t\t\tg->topo_levels = &topo_levels;\n+\t\t\tg = g->base_graph;\n+\t\t}\n+\t}\n+\n \tif (pack_indexes) {\n \t\tctx->order_by_pack = 1;\n \t\tif ((res = fill_oids_from_packs(ctx, pack_indexes)))\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 430bc830bb..1152a9642e 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -72,6 +72,7 @@ struct commit_graph {\n \tconst unsigned char *chunk_bloom_indexes;\n \tconst unsigned char *chunk_bloom_data;\n \n+\tstruct topo_level_slab *topo_levels;\n \tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \ndiff --git a/commit.h b/commit.h\nindex bc0732a4fe..bb846e0025 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -15,6 +15,7 @@\n #define GENERATION_NUMBER_V1_INFINITY 0xFFFFFFFF\n #define GENERATION_NUMBER_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n+#define GENERATION_NUMBER_V2_OFFSET_MAX 0xFFFFFFFF\n \n struct commit_list {\n \tstruct commit *item;\n-- \ngitgitgadget\n\n"},{"id":"403772","messageId":"18d5864f81e89585cc94cd12eca166a9d8b929a5.1597509583.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v3 03/11] commit-graph: consolidate fill_commit_graph_info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:35Z","receivedAt":"2020-08-15T22:05:09Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nBoth fill_commit_graph_info() and fill_commit_in_graph() parse\ninformation present in commit data chunk. Let's simplify the\nimplementation by calling fill_commit_graph_info() within\nfill_commit_in_graph().\n\nThe test 'generate tar with future mtime' creates a commit with commit\ntime of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\ngeneration number (within CDAT chunk) and has undefined behavior.\n\nThe test used to pass as fill_commit_in_graph() guarantees the values of\ngraph position and generation number, and did not load timestamp.\nHowever, with corrected commit date we will need load the timestamp as\nwell to populate the generation number.\n\nLet's fix the test by setting a timestamp of (2 ^ 34 - 1) seconds.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c      | 29 +++++++++++------------------\n t/t5000-tar-tree.sh |  4 ++--\n 2 files changed, 13 insertions(+), 20 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex ace7400a1a..af8d9cc45e 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -725,15 +725,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n \tuint32_t lex_index;\n+\tuint64_t date_high, date_low;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n+\tif (pos >= g->num_commits + g->num_commits_in_base)\n+\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\n \tlex_index = pos - g->num_commits_in_base;\n \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n \n \tgraph_data = commit_graph_data_at(item);\n \tgraph_data->graph_pos = pos;\n+\n+\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n+\tdate_low = get_be32(commit_data + g->hash_len + 12);\n+\titem->date = (timestamp_t)((date_high << 32) | date_low);\n+\n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n@@ -748,38 +757,22 @@ static int fill_commit_in_graph(struct repository *r,\n {\n \tuint32_t edge_value;\n \tuint32_t *parent_data_ptr;\n-\tuint64_t date_low, date_high;\n \tstruct commit_list **pptr;\n-\tstruct commit_graph_data *graph_data;\n \tconst unsigned char *commit_data;\n \tuint32_t lex_index;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n-\tif (pos >= g->num_commits + g->num_commits_in_base)\n-\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\tfill_commit_graph_info(item, g, pos);\n \n-\t/*\n-\t * Store the \"full\" position, but then use the\n-\t * \"local\" position for the rest of the calculation.\n-\t */\n-\tgraph_data = commit_graph_data_at(item);\n-\tgraph_data->graph_pos = pos;\n \tlex_index = pos - g->num_commits_in_base;\n-\n-\tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n+\tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n \n \titem->object.parsed = 1;\n \n \tset_commit_tree(item, NULL);\n \n-\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n-\tdate_low = get_be32(commit_data + g->hash_len + 12);\n-\titem->date = (timestamp_t)((date_high << 32) | date_low);\n-\n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n-\n \tpptr = &item->parents;\n \n \tedge_value = get_be32(commit_data + g->hash_len);\ndiff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\nindex 37655a237c..1986354fc3 100755\n--- a/t/t5000-tar-tree.sh\n+++ b/t/t5000-tar-tree.sh\n@@ -406,7 +406,7 @@ test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n \trm -f .git/index &&\n \techo content >file &&\n \tgit add file &&\n-\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n+\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n \t\tgit commit -m \"tempori parendum\"\n '\n \n@@ -415,7 +415,7 @@ test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n '\n \n test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n-\techo 4147 >expect &&\n+\techo 2514 >expect &&\n \ttar_info future.tar | cut -d\" \" -f2 >actual &&\n \ttest_cmp expect actual\n '\n-- \ngitgitgadget\n\n"},{"id":"403774","messageId":"4074ace65be3094d35dd0aaedb89eb5a0ec98cee.1597509583.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v3 07/11] commit-graph: implement corrected commit date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:39Z","receivedAt":"2020-08-15T22:05:17Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith most of preparations done, let's implement corrected commit date.\n\nThe corrected commit date for a commit is defined as:\n\n* A commit with no parents (a root commit) has corrected commit date\n  equal to its committer date.\n* A commit with at least one parent has corrected commit date equal to\n  the maximum of its commit date and one more than the largest corrected\n  commit date among its parents.\n\nTo minimize the space required to store corrected commit date, Git\nstores corrected commit date offsets into the commit-graph file. The\ncorrected commit date offset for a commit is defined as the difference\nbetween its corrected commit date and actual commit date.\n\nWhile Git does not write out offsets at this stage, Git stores the\ncorrected commit dates in member generation of struct commit_graph_data.\nIt will begin writing commit date offsets with the introduction of\ngeneration data chunk.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 58 +++++++++++++++++++++++++++-----------------------\n 1 file changed, 31 insertions(+), 27 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex a2f15b2825..fd69534dd5 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -169,11 +169,6 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \telse if (generation_a > generation_b)\n \t\treturn 1;\n \n-\t/* use date as a heuristic when generations are equal */\n-\tif (a->date < b->date)\n-\t\treturn -1;\n-\telse if (a->date > b->date)\n-\t\treturn 1;\n \treturn 0;\n }\n \n@@ -1342,10 +1337,14 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n+\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n \t\tif (level != GENERATION_NUMBER_V1_INFINITY &&\n-\t\t    level != GENERATION_NUMBER_ZERO)\n+\t\t    level != GENERATION_NUMBER_ZERO &&\n+\t\t    corrected_commit_date != GENERATION_NUMBER_INFINITY &&\n+\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO\n+\t\t    )\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1354,17 +1353,26 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n \t\t\tuint32_t max_level = 0;\n+\t\t\ttimestamp_t max_corrected_commit_date = 0;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n+\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n \n \t\t\t\tif (level == GENERATION_NUMBER_V1_INFINITY ||\n-\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n+\t\t\t\t    level == GENERATION_NUMBER_ZERO ||\n+\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_INFINITY ||\n+\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO\n+\t\t\t\t    ) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (level > max_level) {\n-\t\t\t\t\tmax_level = level;\n+\t\t\t\t} else {\n+\t\t\t\t\tif (level > max_level)\n+\t\t\t\t\t\tmax_level = level;\n+\n+\t\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n+\t\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n \t\t\t\t}\n \t\t\t}\n \n@@ -1374,6 +1382,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n \t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n+\n+\t\t\t\tif (current->date > max_corrected_commit_date)\n+\t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n+\t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n \t\t\t}\n \t\t}\n \t}\n@@ -2372,8 +2384,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \tfor (i = 0; i < g->num_commits; i++) {\n \t\tstruct commit *graph_commit, *odb_commit;\n \t\tstruct commit_list *graph_parents, *odb_parents;\n-\t\ttimestamp_t max_generation = 0;\n-\t\ttimestamp_t generation;\n+\t\ttimestamp_t max_corrected_commit_date = 0;\n+\t\ttimestamp_t corrected_commit_date;\n \n \t\tdisplay_progress(progress, i + 1);\n \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n@@ -2412,9 +2424,9 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t\t\t\t     oid_to_hex(&graph_parents->item->object.oid),\n \t\t\t\t\t     oid_to_hex(&odb_parents->item->object.oid));\n \n-\t\t\tgeneration = commit_graph_generation(graph_parents->item);\n-\t\t\tif (generation > max_generation)\n-\t\t\t\tmax_generation = generation;\n+\t\t\tcorrected_commit_date = commit_graph_generation(graph_parents->item);\n+\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n+\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n \n \t\t\tgraph_parents = graph_parents->next;\n \t\t\todb_parents = odb_parents->next;\n@@ -2436,20 +2448,12 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n \t\t\tcontinue;\n \n-\t\t/*\n-\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n-\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n-\t\t * extra logic in the following condition.\n-\t\t */\n-\t\tif (max_generation == GENERATION_NUMBER_MAX)\n-\t\t\tmax_generation--;\n-\n-\t\tgeneration = commit_graph_generation(graph_commit);\n-\t\tif (generation != max_generation + 1)\n-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n+\t\tcorrected_commit_date = commit_graph_generation(graph_commit);\n+\t\tif (corrected_commit_date < max_corrected_commit_date + 1)\n+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n \t\t\t\t     oid_to_hex(&cur_oid),\n-\t\t\t\t     generation,\n-\t\t\t\t     max_generation + 1);\n+\t\t\t\t     corrected_commit_date,\n+\t\t\t\t     max_corrected_commit_date + 1);\n \n \t\tif (graph_commit->date != odb_commit->date)\n \t\t\tgraph_report(_(\"commit date for commit %s in commit-graph is %\"PRItime\" != %\"PRItime),\n-- \ngitgitgadget\n\n"},{"id":"403775","messageId":"e6738672349254c6405f7dde48f612b82af9299f.1597509583.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v3 02/11] revision: parse parent in indegree_walk_step()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-08-15T16:39:34Z","receivedAt":"2020-08-15T22:05:21Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn indegree_walk_step(), we add unvisited parents to the indegree queue.\nHowever, parents are not guaranteed to be parsed. As the indegree queue\nsorts by generation number, let's parse parents before inserting them to\nensure the correct priority order.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n revision.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/revision.c b/revision.c\nindex 3dcf689341..ecf757c327 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3363,6 +3363,9 @@ static void indegree_walk_step(struct rev_info *revs)\n \t\tstruct commit *parent = p->item;\n \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n \n+\t\tif (parse_commit_gently(parent, 1) < 0)\n+\t\t\treturn;\n+\n \t\tif (*pi)\n \t\t\t(*pi)++;\n \t\telse\n-- \ngitgitgadget\n\n"},{"id":"403799","messageId":"85zh6uxh7l.fsf@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-17T00:13:18Z","receivedAt":"2020-08-17T00:13:28Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> This patch series implements the corrected commit date offsets as generation\n> number v2, along with other pre-requisites.\n\nI'm not sure if this level of detail is required in the cover letter for\nthe series, but generation number v2 is corrected commit date; corrected\ncommit date offsets is how we store this value in the commit-graph file.\n\n>\n> Git uses topological levels in the commit-graph file for commit-graph\n> traversal operations like git log --graph. Unfortunately, using topological\n> levels can result in a worse performance than without them when compared\n> with committer date as a heuristics. For example, git merge-base v4.8 v4.9\n> on the Linux repository walks 635,579 commits using topological levels and\n> walks 167,468 using committer date.\n\nI would say \"committer date heuristics\" instead of just \"committer\ndate\", to be more exact.\n\nIs this data generated using https://github.com/derrickstolee/gen-test\nscripts?\n\n>\n> Thus, the need for generation number v2 was born. New generation number\n> needed to provide good performance, increment updates, and backward\n> compatibility. Due to an unfortunate problem, we also needed a way to\n> distinguish between the old and new generation number without incrementing\n> graph version.\n\nIt would be nice to have reference email (or other place with details)\nfor \"unfortunate problem\".\n\n>\n> Various candidates were examined (https://github.com/derrickstolee/gen-test,\n> https://github.com/abhishekkumar2718/git/pull/1). The proposed generation\n> number v2, Corrected Commit Date with Mononotically Increasing Offsets\n> performed much worse than committer date (506,577 vs. 167,468 commits walked\n> for git merge-base v4.8 v4.9) and was dropped.\n>\n> Using Generation Data chunk (GDAT) relieves the requirement of backward\n> compatibility as we would continue to store topological levels in Commit\n> Data (CDAT) chunk. Thus, Corrected Commit Date was chosen as generation\n> number v2.\n\nThis is a bit of simplification, but good enough for a cover letter.\n\nTo be more exact, from various candidates the Corrected Commit Date was\nchosen.  Then it turned out that old Git crashes on changed commit-graph\nformat version value, so if the generation number v2 was to replace v1\nit needed to be backward-compatibile: hence the idea of Corrected Commit\nDate with Monotonically Increasing Offsets.  But with GDAT chunk to\nstore generation number v2 (and for the time being leaving generation\nnumber v1, i.e. Topological Levels, in CDAT), we are no longer\nconstrained by the requirement of backward-compatibility to make old Git\nwork with commit-graph file created by new Git.  So we could go back to\nCorrected Commit Date, and as you wrote above the backward-compatibile\nvariant performs worse.\n\n> The Corrected Commit Date is defined as:\n>\n> For a commit C, let its corrected commit date be the maximum of the commit\n> date of C and the corrected commit dates of its parents.\n\nActually it needs to be \"corrected commit dates of its parents plus 1\"\nto fulfill the reachability condition for a generation number for a\ncommit:\n\n      A can reach B   =>  gen(A) < gen(B)\n\nOf course it can be computed in simpler way, because\n\n  max_P (gen(P) + 1)  ==  max_P (gen(P)) + 1\n\n\n>                                                           Then corrected\n> commit date offset is the difference between corrected commit date of C and\n> commit date of C.\n\nAll right.\n\n>\n> We will introduce an additional commit-graph chunk, Generation Data chunk,\n> and store corrected commit date offsets in GDAT chunk while storing\n> topological levels in CDAT chunk. The old versions of Git would ignore GDAT\n> chunk, using topological levels from CDAT chunk. In contrast, new versions\n> of Git would use corrected commit dates, falling back to topological level\n> if the generation data chunk is absent in the commit-graph file.\n\nAll right.\n\nHowever I think the cover letter should also describe what should happen\nin a mixed version environment (for example new Git on command line,\ncopy of old Git used by GUI client), and in particular what should\nhappen in a mixed-chain case - both for reading and for writing the\ncommit-graph file.\n\nFor *writing*: because old Git would create commit-graph layers without\nthe GDAT chunk, to simplify the behavior and make easy to reason about\ncommit-graph data (the situation should be not that common, and\ntransient -- it should get more rare as the time goes), we want the\nfollowing behavior from new Git:\n\n- If top layer contains the GDAT chunk, or we are rewriting commit-graph\n  file (--split=replace), or we are merging layers and there are no\n  layers without GDAT chunk below set of layers that are merged, then\n\n     write commit-graph file or commit-graph layer with GDAT chunk,\n\n  otherwise\n\n     write commit-graph layer without GDAT chunk.\n\n  This means that there are commit-graph layers without GDAT chunk if\n  and only if the top layer is also without GDAT chunk.\n\n\nFor *reading* we want to use generation number v2 (corrected commit\ndate) if possible, and fall back to generation number v1 (topological\nlevels).\n\n- If the top layer contains the GDAT chunk (or maybe even if the topmost\n  layer that involves all commits in question, not necessarily the top\n  layer in the full commit-graph chain), then use generation number v2\n\n  - commit_graph_data->generation stores corrected commit date,\n    computed as sum of committer date (from CDAT) and offset (from GDAT)\n\n  - A can reach B   =>  gen(A) < gen(B)\n\n  - there is no need for committer date heuristics, and no need for\n    limiting use of generation number to where there is a cutoff (to not\n    hamper performance).\n\n- If there are layers without GDAT chunks, which thanks to the write\n  behavior means simply top layer without GDAT chunk, we need to turn\n  off use of generation numbers or fall back to using topological levels\n\n  - commit_graph_data->generation stores topological levels,\n    taken from CDAT chunk (30-bits)\n\n  - A can reach B   =>  gen(A) < gen(B)\n\n  - we probably want to keep tie-breaking of sorting by generation\n    number via committer date, and limit use of generation number as\n    opposed to using committer date heuristics (with slop) to not make\n    performance worse.\n\n>\n> Thanks to Dr. Stolee, Dr. Narębski, and Taylor for their reviews on the\n> first version.\n>\n> I look forward to everyone's reviews!\n>\n> Thanks\n>\n>  * Abhishek\n>\n>\n> ----------------------------------------------------------------------------\n>\n> Changes in version 3:\n>\n>  * Reordered patches as discussed in 1\n>    [https://lore.kernel.org/git/aee0ae56-3395-6848-d573-27a318d72755@gmail.com/]\n\nIf I remember it correctly this was done to always store in GDAT chunk\ncorrected commit date offsets, isn't it?\n\n>  * Split \"implement corrected commit date\" into two patches - one\n>    introducing the topo level slab and other implementing corrected commit\n>    dates.\n\nAll right.\n\nI think it might be good idea to split off the change to tar file tests\n(as a preparatory patch), to make reviews and bisecting easier.\n\n>  * Extended split-commit-graph tests to verify at the end of test.\n\nDo we also test for proper merging of split commit-graph layers, not\nonly adding a new layer and a full rewrite (--split=replace)?\n\n>  * Use topological levels as generation number if any of split commit-graph\n>    files do not have generation data chunk.\n\nThat is good for performance.\n\n>\n> Changes in version 2:\n>\n>  * Add tests for generation data chunk.\n\nGood.\n\n>  * Add an option GIT_TEST_COMMIT_GRAPH_NO_GDAT to control whether to write\n>    generation data chunk.\n\nGood, that is needed for testing mixed-version behavior.\n\n>  * Compare commits with corrected commit dates if present in\n>    paint_down_to_common().\n\nAll right, but see the caveat.\n\n>  * Update technical documentation.\n\nAlways a good thing.\n\n>  * Handle mixed graph version commit chains.\n\nWhere by \"version\" you mean generation number version - the commit-graph\nversion number unfortunately needs to stay the same...\n\n>  * Improve commit messages for\n                                ^^^^^^\nSomething missing in this point, the sentence ends abruptly.\n\n>  * Revert unnecessary whitespace changes.\n\nThanks.\n\n>  * Split uint_32 -> timestamp_t change into a new commit.\n\nIt is usually better to keep the commits small.  Good.\n\n\nGood work!\n\n>\n> Abhishek Kumar (11):\n>   commit-graph: fix regression when computing bloom filter\n>   revision: parse parent in indegree_walk_step()\n>   commit-graph: consolidate fill_commit_graph_info\n>   commit-graph: consolidate compare_commits_by_gen\n>   commit-graph: return 64-bit generation number\n>   commit-graph: add a slab to store topological levels\n>   commit-graph: implement corrected commit date\n>   commit-graph: implement generation data chunk\n>   commit-graph: use generation v2 only if entire chain does\n>   commit-reach: use corrected commit dates in paint_down_to_common()\n>   doc: add corrected commit date info\n>\n>  .../technical/commit-graph-format.txt         |  12 +-\n>  Documentation/technical/commit-graph.txt      |  45 ++--\n>  commit-graph.c                                | 241 +++++++++++++-----\n>  commit-graph.h                                |  16 +-\n>  commit-reach.c                                |  49 ++--\n>  commit-reach.h                                |   2 +-\n>  commit.c                                      |   9 +-\n>  commit.h                                      |   4 +-\n>  revision.c                                    |  13 +-\n>  t/README                                      |   3 +\n>  t/helper/test-read-graph.c                    |   2 +\n>  t/t4216-log-bloom.sh                          |   4 +-\n>  t/t5000-tar-tree.sh                           |   4 +-\n>  t/t5318-commit-graph.sh                       |  27 +-\n>  t/t5324-split-commit-graph.sh                 |  82 +++++-\n>  t/t6024-recursive-merge.sh                    |   4 +-\n>  t/t6600-test-reach.sh                         |  62 +++--\n>  upload-pack.c                                 |   2 +-\n>  18 files changed, 396 insertions(+), 185 deletions(-)\n>\n>\n> base-commit: 7814e8a05a59c0cf5fb186661d1551c75d1299b5\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-676%2Fabhishekkumar2718%2Fcorrected_commit_date-v3\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-676/abhishekkumar2718/corrected_commit_date-v3\n> Pull-Request: https://github.com/gitgitgadget/git/pull/676\n[...]\n\nBest,\n--\nJakub Narębski\n"},{"id":"403800","messageId":"20200817013206.GA57201@syl.lan","threadId":"53933","inReplyTo":"CANQwDwdKp7oKy9BeKdvKhwPUiq0R5MS8TCw-eWGCYCoMGv=G-g@mail.gmail.com","subject":"Re: Fwd: [PATCH v3 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-08-17T01:32:06Z","receivedAt":"2020-08-17T01:32:17Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Mon, Aug 17, 2020 at 02:16:08AM +0200, Jakub Narębski wrote:\n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> > We will introduce an additional commit-graph chunk, Generation Data chunk,\n> > and store corrected commit date offsets in GDAT chunk while storing\n> > topological levels in CDAT chunk. The old versions of Git would ignore GDAT\n> > chunk, using topological levels from CDAT chunk. In contrast, new versions\n> > of Git would use corrected commit dates, falling back to topological level\n> > if the generation data chunk is absent in the commit-graph file.\n>\n> All right.\n>\n> However I think the cover letter should also describe what should happen\n> in a mixed version environment (for example new Git on command line,\n> copy of old Git used by GUI client), and in particular what should\n> happen in a mixed-chain case - both for reading and for writing the\n> commit-graph file.\n>\n> For *writing*: because old Git would create commit-graph layers without\n> the GDAT chunk, to simplify the behavior and make easy to reason about\n> commit-graph data (the situation should be not that common, and\n> transient -- it should get more rare as the time goes), we want the\n> following behavior from new Git:\n>\n> - If top layer contains the GDAT chunk, or we are rewriting commit-graph\n>   file (--split=replace), or we are merging layers and there are no\n>   layers without GDAT chunk below set of layers that are merged, then\n>\n>      write commit-graph file or commit-graph layer with GDAT chunk,\n>\n>   otherwise\n>\n>      write commit-graph layer without GDAT chunk.\n>\n>   This means that there are commit-graph layers without GDAT chunk if\n>   and only if the top layer is also without GDAT chunk.\n\nThis seems very sane to me, and I'd be glad to see it spelled out in\nmore specific detail. I was wondering this myself, and had to double\ncheck with Stolee off-list that my interpretation of Abhishek's code was\ncorrect.\n\nBut yes, only writing GDAT chunks when all layers in the chain have GDAT\nchunks makes sense, since we can't interoperate between corrected dates\nand topological levels. Since we can't fill in the GDAT data of layers\ngenerated in pre-GDAT versions of Git without invalidating the GDAT\nlayers on-disk, there's no point to speculatively computing both chunks.\n\nMerging rules are obviously correct, which is good. For what it's worth,\nthe '--split=replace' case is what we'll really care about at GitHub,\nsince it's unlikely we'd drop all existing commit-graph chains and\nrebuild them from scratch. More likely is that we'll let the new GDAT\nchunks trickle in over time when we run 'git commit-graph write' with\n'--split=replace', which happens \"every so often\".\n\n> For *reading* we want to use generation number v2 (corrected commit\n> date) if possible, and fall back to generation number v1 (topological\n> levels).\n>\n> - If the top layer contains the GDAT chunk (or maybe even if the topmost\n>   layer that involves all commits in question, not necessarily the top\n>   layer in the full commit-graph chain), then use generation number v2\n\nI don't follow this. If we have a multi-layer chain, either all or none\nof the layers have a GDAT chunk. So, \"if the top layer contains the GDAT\nchunk\" makes sense, since it implies that all layers have the GDAT\nchunk. I don't see how \"even if the topmost layer that involves all\ncommits in question\" would be possible, since (if I'm understanding your\ndescription correctly), we can't have *some* of the layers having a GDAT\nchunk with others only having a CDAT chunk.\n\nI'm a little confused here.\n\n>   - commit_graph_data->generation stores corrected commit date,\n>     computed as sum of committer date (from CDAT) and offset (from GDAT)\n>\n>   - A can reach B   =>  gen(A) < gen(B)\n>\n>   - there is no need for committer date heuristics, and no need for\n>     limiting use of generation number to where there is a cutoff (to not\n>     hamper performance).\n>\n> - If there are layers without GDAT chunks, which thanks to the write\n>   behavior means simply top layer without GDAT chunk, we need to turn\n>   off use of generation numbers or fall back to using topological levels\n\nGood, I'm glad that this can be a quick check (that we can cache for\nfuture reads, but I'm not even sure the caching would be necessary\nwithout measuring).\n>\n>   - commit_graph_data->generation stores topological levels,\n>     taken from CDAT chunk (30-bits)\n>\n>   - A can reach B   =>  gen(A) < gen(B)\n>\n>   - we probably want to keep tie-breaking of sorting by generation\n>     number via committer date, and limit use of generation number as\n>     opposed to using committer date heuristics (with slop) to not make\n>     performance worse.\n\nAll makes very good sense, except for the one point I raised above.\n\n> >\n> > Thanks to Dr. Stolee, Dr. Narębski, and Taylor for their reviews on the\n> > first version.\n\nThanks, Abhishek for your great work on this. I was feeling bad that I\nwasn't more involved in the early discussions about the transition plan,\nbut what you, Stolee, and Jakub came up with all seems like what I would\nhave suggested, anyway ;-).\n\n> Jakub Narebski\n\nThanks,\nTaylor\n"},{"id":"403801","messageId":"CANQwDwebQXS8pghXYCBMHvQhCLr9PKFZYsWO4hLAqV=JctV8dA@mail.gmail.com","threadId":"53933","inReplyTo":"20200817013206.GA57201@syl.lan","subject":"Re: Fwd: [PATCH v3 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-17T07:56:46Z","receivedAt":"2020-08-17T07:57:26Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Mon, 17 Aug 2020 at 03:32, Taylor Blau <me@ttaylorr.com> wrote:\n> On Mon, Aug 17, 2020 at 02:16:08AM +0200, Jakub Narębski wrote:\n> > \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> > >\n> > > We will introduce an additional commit-graph chunk, Generation Data chunk,\n> > > and store corrected commit date offsets in GDAT chunk while storing\n> > > topological levels in CDAT chunk. The old versions of Git would ignore GDAT\n> > > chunk, using topological levels from CDAT chunk. In contrast, new versions\n> > > of Git would use corrected commit dates, falling back to topological level\n> > > if the generation data chunk is absent in the commit-graph file.\n> >\n> > All right.\n> >\n> > However I think the cover letter should also describe what should happen\n> > in a mixed version environment (for example new Git on command line,\n> > copy of old Git used by GUI client), and in particular what should\n> > happen in a mixed-chain case - both for reading and for writing the\n> > commit-graph file.\n> >\n> > For *writing*: because old Git would create commit-graph layers without\n> > the GDAT chunk, to simplify the behavior and make easy to reason about\n> > commit-graph data (the situation should be not that common, and\n> > transient -- it should get more rare as the time goes), we want the\n> > following behavior from new Git:\n> >\n> > - If top layer contains the GDAT chunk, or we are rewriting commit-graph\n> >   file (--split=replace), or we are merging layers and there are no\n> >   layers without GDAT chunk below set of layers that are merged, then\n> >\n> >      write commit-graph file or commit-graph layer with GDAT chunk,\n> >\n> >   otherwise\n> >\n> >      write commit-graph layer without GDAT chunk.\n> >\n> >   This means that there are commit-graph layers without GDAT chunk if\n> >   and only if the top layer is also without GDAT chunk.\n>\n> This seems very sane to me, and I'd be glad to see it spelled out in\n> more specific detail. I was wondering this myself, and had to double\n> check with Stolee off-list that my interpretation of Abhishek's code was\n> correct.\n>\n> But yes, only writing GDAT chunks when all layers in the chain have GDAT\n> chunks makes sense, since we can't interoperate between corrected dates\n> and topological levels. Since we can't fill in the GDAT data of layers\n> generated in pre-GDAT versions of Git without invalidating the GDAT\n> layers on-disk, there's no point to speculatively computing both chunks.\n>\n> Merging rules are obviously correct, which is good. For what it's worth,\n> the '--split=replace' case is what we'll really care about at GitHub,\n> since it's unlikely we'd drop all existing commit-graph chains and\n> rebuild them from scratch. More likely is that we'll let the new GDAT\n> chunks trickle in over time when we run 'git commit-graph write' with\n> '--split=replace', which happens \"every so often\".\n\nTo be more detailed, without '--split=replace' we would want the following\nlayer merging behavior:\n\n   [layer with GDAT][with GDAT][without GDAT][without GDAT][without GDAT]\n\nIn the split commit-graph chain above, merging two topmost layers\nshould create a layer without GDAT; merging three topmost layers\n(and any other layers, e.g. two middle ones) should create a layer\n with GDAT.\n\n> > For *reading* we want to use generation number v2 (corrected commit\n> > date) if possible, and fall back to generation number v1 (topological\n> > levels).\n> >\n> > - If the top layer contains the GDAT chunk (or maybe even if the topmost\n> >   layer that involves all commits in question, not necessarily the top\n> >   layer in the full commit-graph chain), then use generation number v2\n>\n> I don't follow this. If we have a multi-layer chain, either all or none\n> of the layers have a GDAT chunk. So, \"if the top layer contains the GDAT\n> chunk\" makes sense, since it implies that all layers have the GDAT\n> chunk. I don't see how \"even if the topmost layer that involves all\n> commits in question\" would be possible, since (if I'm understanding your\n> description correctly), we can't have *some* of the layers having a GDAT\n> chunk with others only having a CDAT chunk.\n>\n> I'm a little confused here.\n\nThis is only speculative, and most probably totally unnecessary\ncomplication (either that, or something that we would get for free).\nAssume that the command in question operates only on\nhistorical data; for example `git log --topo-order HEAD~1000`.\nIf all commits (or, what's equivalent, most recent commits\ni.e. HEAD~1000) have their data in split commit-graph layers\nwith GDAT, we can theoretically use generation number v2,\neven if there are some newer commits that have their data\nin layers without GDAT (and some even newer ones outside\ncommit-graph files).\n\nI hope that this explains my (possibly harebrained) idea.\n\n> >   - commit_graph_data->generation stores corrected commit date,\n> >     computed as sum of committer date (from CDAT) and offset (from GDAT)\n> >\n> >   - A can reach B   =>  gen(A) < gen(B)\n> >\n> >   - there is no need for committer date heuristics, and no need for\n> >     limiting use of generation number to where there is a cutoff (to not\n> >     hamper performance).\n> >\n> > - If there are layers without GDAT chunks, which thanks to the write\n> >   behavior means simply top layer without GDAT chunk, we need to turn\n> >   off use of generation numbers or fall back to using topological levels\n>\n> Good, I'm glad that this can be a quick check (that we can cache for\n> future reads, but I'm not even sure the caching would be necessary\n> without measuring).\n\nThere is a question where to store the information that we cannot\nuse generation number v2 (that 'generation' contains topological\nlevels and not corrected commit date):\n- create new global variable\n- store it in `struct split_commit_graph_opts`\n- set `chunk_generation_data` to NULL for all graphs\n  in chain (it is in `struct commit_graph`)?\n\n> >\n> >   - commit_graph_data->generation stores topological levels,\n> >     taken from CDAT chunk (30-bits)\n> >\n> >   - A can reach B   =>  gen(A) < gen(B)\n> >\n> >   - we probably want to keep tie-breaking of sorting by generation\n> >     number via committer date, and limit use of generation number as\n> >     opposed to using committer date heuristics (with slop) to not make\n> >     performance worse.\n>\n> All makes very good sense, except for the one point I raised above.\n>\n> > >\n> > > Thanks to Dr. Stolee, Dr. Narębski, and Taylor for their reviews on the\n> > > first version.\n>\n> Thanks, Abhishek for your great work on this. I was feeling bad that I\n> wasn't more involved in the early discussions about the transition plan,\n> but what you, Stolee, and Jakub came up with all seems like what I would\n> have suggested, anyway ;-).\n\nThank you for your work on improving this feature.\n\nBest,\n-- \nJakub Narebski\n"},{"id":"403816","messageId":"79ab9d8c-9767-430a-9744-01fa81bdf9bc@gmail.com","threadId":"53933","inReplyTo":"6a0cde983d9ed20f043a4977313d714154602012.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 04/11] commit-graph: consolidate compare_commits_by_gen","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-17T13:22:46Z","receivedAt":"2020-08-17T13:22:56Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/15/2020 12:39 PM, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> Comparing commits by generation has been independently defined twice, in\n> commit-reach and commit. Let's simplify the implementation by moving\n> compare_commits_by_gen() to commit-graph.\n> \n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> Reviewed-by: Taylor Blau <me@ttaylorr.com>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWhoops! Be sure to add the \"Reviewed-by: above the \"Signed-off-by\" line(s)\nso re-signing off doesn't add a duplicate like this.\n\nThanks,\n-Stolee\n"},{"id":"403901","messageId":"85r1s4ykgm.fsf@gmail.com","threadId":"53933","inReplyTo":"c6b7ade7af92b6faf365a5609748f6d024ea0408.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 01/11] commit-graph: fix regression when computing bloom filter","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-17T22:30:01Z","receivedAt":"2020-08-17T22:30:06Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Subject: [PATCH v3 01/11] commit-graph: fix regression when computing bloom filter\n\ns/bloom filter/Bloom filters/\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> commit_gen_cmp is used when writing a commit-graph to sort commits in\n> generation order before computing Bloom filters. Since c49c82aa (commit:\n> move members graph_pos, generation to a slab, 2020-06-17) made it so\n> that 'commit_graph_generation()' returns 'GENERATION_NUMBER_INFINITY'\n> during writing, we cannot call it within this function. Instead, access\n> the generation number directly through the slab (i.e., by calling\n> 'commit_graph_data_at(c)->generation') in order to access it while\n> writing.\n\nTwo things that might not be obvious from the commit message:\n\n- Is commit_gen_cmp in commit-graph.c used by anything but writing\n  Bloom filters for changed paths?\n\n- That the generation number is computed during `commit-graph write`\n  before computing Bloom filters.\n\nAlso, after this series 'generation' would be generation number v2, that\nis corrected commit date, and not v1, that is topological levels.  We\nshould check, just in case, that it does not lead to significant\nperformance regression for `git commit-graph write --reachable <...>`\ncase (the one that uses commit_gen_cmp sort).\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 4 ++--\n>  1 file changed, 2 insertions(+), 2 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index e51c91dd5b..ace7400a1a 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>  \tconst struct commit *a = *(const struct commit **)va;\n>  \tconst struct commit *b = *(const struct commit **)vb;\n>\n> -\tuint32_t generation_a = commit_graph_generation(a);\n> -\tuint32_t generation_b = commit_graph_generation(b);\n> +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n\nNice and easy.\n\n>  \t/* lower generation commits first */\n>  \tif (generation_a < generation_b)\n>  \t\treturn -1;\n\nBest,\n-- \nJakub Narębski\n"},{"id":"403922","messageId":"20200818061220.GA28571@Abhishek-Arch","threadId":"53933","inReplyTo":"85zh6uxh7l.fsf@gmail.com","subject":"Re: [PATCH v3 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-08-18T06:12:20Z","receivedAt":"2020-08-18T06:14:39Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Mon, Aug 17, 2020 at 02:13:18AM +0200, Jakub Narębski wrote:\n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > This patch series implements the corrected commit date offsets as generation\n> > number v2, along with other pre-requisites.\n> \n> I'm not sure if this level of detail is required in the cover letter for\n> the series, but generation number v2 is corrected commit date; corrected\n> commit date offsets is how we store this value in the commit-graph file.\n> \n> >\n> > Git uses topological levels in the commit-graph file for commit-graph\n> > traversal operations like git log --graph. Unfortunately, using topological\n> > levels can result in a worse performance than without them when compared\n> > with committer date as a heuristics. For example, git merge-base v4.8 v4.9\n> > on the Linux repository walks 635,579 commits using topological levels and\n> > walks 167,468 using committer date.\n> \n> I would say \"committer date heuristics\" instead of just \"committer\n> date\", to be more exact.\n> \n> Is this data generated using https://github.com/derrickstolee/gen-test\n> scripts?\n> \n\nYes, it is.\n\n> >\n> > Thus, the need for generation number v2 was born. New generation number\n> > needed to provide good performance, increment updates, and backward\n> > compatibility. Due to an unfortunate problem, we also needed a way to\n> > distinguish between the old and new generation number without incrementing\n> > graph version.\n> \n> It would be nice to have reference email (or other place with details)\n> for \"unfortunate problem\".\n> \n\nWill add.\n\n> >\n> > Various candidates were examined (https://github.com/derrickstolee/gen-test,\n> > https://github.com/abhishekkumar2718/git/pull/1). The proposed generation\n> > number v2, Corrected Commit Date with Mononotically Increasing Offsets\n> > performed much worse than committer date (506,577 vs. 167,468 commits walked\n> > for git merge-base v4.8 v4.9) and was dropped.\n> >\n> > Using Generation Data chunk (GDAT) relieves the requirement of backward\n> > compatibility as we would continue to store topological levels in Commit\n> > Data (CDAT) chunk. Thus, Corrected Commit Date was chosen as generation\n> > number v2.\n> \n> This is a bit of simplification, but good enough for a cover letter.\n> \n> To be more exact, from various candidates the Corrected Commit Date was\n> chosen.  Then it turned out that old Git crashes on changed commit-graph\n> format version value, so if the generation number v2 was to replace v1\n> it needed to be backward-compatibile: hence the idea of Corrected Commit\n> Date with Monotonically Increasing Offsets.  But with GDAT chunk to\n> store generation number v2 (and for the time being leaving generation\n> number v1, i.e. Topological Levels, in CDAT), we are no longer\n> constrained by the requirement of backward-compatibility to make old Git\n> work with commit-graph file created by new Git.  So we could go back to\n> Corrected Commit Date, and as you wrote above the backward-compatibile\n> variant performs worse.\n> \n> > The Corrected Commit Date is defined as:\n> >\n> > For a commit C, let its corrected commit date be the maximum of the commit\n> > date of C and the corrected commit dates of its parents.\n> \n> Actually it needs to be \"corrected commit dates of its parents plus 1\"\n> to fulfill the reachability condition for a generation number for a\n> commit:\n> \n>       A can reach B   =>  gen(A) < gen(B)\n> \n> Of course it can be computed in simpler way, because\n> \n>   max_P (gen(P) + 1)  ==  max_P (gen(P)) + 1\n> \n> \n> >                                                           Then corrected\n> > commit date offset is the difference between corrected commit date of C and\n> > commit date of C.\n> \n> All right.\n> \n> >\n> > We will introduce an additional commit-graph chunk, Generation Data chunk,\n> > and store corrected commit date offsets in GDAT chunk while storing\n> > topological levels in CDAT chunk. The old versions of Git would ignore GDAT\n> > chunk, using topological levels from CDAT chunk. In contrast, new versions\n> > of Git would use corrected commit dates, falling back to topological level\n> > if the generation data chunk is absent in the commit-graph file.\n> \n> All right.\n> \n> However I think the cover letter should also describe what should happen\n> in a mixed version environment (for example new Git on command line,\n> copy of old Git used by GUI client), and in particular what should\n> happen in a mixed-chain case - both for reading and for writing the\n> commit-graph file.\n> \n\nYes, definitely. Will add\n\n> For *writing*: because old Git would create commit-graph layers without\n> the GDAT chunk, to simplify the behavior and make easy to reason about\n> commit-graph data (the situation should be not that common, and\n> transient -- it should get more rare as the time goes), we want the\n> following behavior from new Git:\n> \n> - If top layer contains the GDAT chunk, or we are rewriting commit-graph\n>   file (--split=replace), or we are merging layers and there are no\n>   layers without GDAT chunk below set of layers that are merged, then\n> \n>      write commit-graph file or commit-graph layer with GDAT chunk,\n> \n>   otherwise\n> \n>      write commit-graph layer without GDAT chunk.\n> \n>   This means that there are commit-graph layers without GDAT chunk if\n>   and only if the top layer is also without GDAT chunk.\n> \n> \n> For *reading* we want to use generation number v2 (corrected commit\n> date) if possible, and fall back to generation number v1 (topological\n> levels).\n> \n> - If the top layer contains the GDAT chunk (or maybe even if the topmost\n>   layer that involves all commits in question, not necessarily the top\n>   layer in the full commit-graph chain), then use generation number v2\n> \n\nThe current implementation checks the entire chain for GDAT, rather than\njust the topmost layer as we cannot assert that `g` would be the topmost\nlayer of the chain.\n\nSee the discussion here: https://lore.kernel.org/git/20200814045957.GA1380@Abhishek-Arch/\n\nIt's one of drawbacks of having a single member 64-bit `generation`\ninstead of two 32-bit members `level` and `odate`.\n\n>\n>\n>   - commit_graph_data->generation stores corrected commit date,\n>     computed as sum of committer date (from CDAT) and offset (from GDAT)\n> \n>   - A can reach B   =>  gen(A) < gen(B)\n> \n>   - there is no need for committer date heuristics, and no need for\n>     limiting use of generation number to where there is a cutoff (to not\n>     hamper performance).\n> \n> - If there are layers without GDAT chunks, which thanks to the write\n>   behavior means simply top layer without GDAT chunk, we need to turn\n>   off use of generation numbers or fall back to using topological levels\n> \n>   - commit_graph_data->generation stores topological levels,\n>     taken from CDAT chunk (30-bits)\n> \n>   - A can reach B   =>  gen(A) < gen(B)\n> \n>   - we probably want to keep tie-breaking of sorting by generation\n>     number via committer date, and limit use of generation number as\n>     opposed to using committer date heuristics (with slop) to not make\n>     performance worse.\n> \n> >\n> > Thanks to Dr. Stolee, Dr. Narębski, and Taylor for their reviews on the\n> > first version.\n> >\n> > I look forward to everyone's reviews!\n> >\n> > Thanks\n> >\n> >  * Abhishek\n> >\n> >\n> > ----------------------------------------------------------------------------\n> >\n> > Changes in version 3:\n> >\n> >  * Reordered patches as discussed in 1\n> >    [https://lore.kernel.org/git/aee0ae56-3395-6848-d573-27a318d72755@gmail.com/]\n> \n> If I remember it correctly this was done to always store in GDAT chunk\n> corrected commit date offsets, isn't it?\n> \n\nYes.\n\n> >  * Split \"implement corrected commit date\" into two patches - one\n> >    introducing the topo level slab and other implementing corrected commit\n> >    dates.\n> \n> All right.\n> \n> I think it might be good idea to split off the change to tar file tests\n> (as a preparatory patch), to make reviews and bisecting easier.\n> \n> >  * Extended split-commit-graph tests to verify at the end of test.\n> \n> Do we also test for proper merging of split commit-graph layers, not\n> only adding a new layer and a full rewrite (--split=replace)?\n> \n\nWe do not, will add a test at end. Thanks for pointing this out.\n\n> >  * Use topological levels as generation number if any of split commit-graph\n> >    files do not have generation data chunk.\n> \n> That is good for performance.\n> \n> >\n> > Changes in version 2:\n> >\n> >  * Add tests for generation data chunk.\n> \n> Good.\n> \n> >  * Add an option GIT_TEST_COMMIT_GRAPH_NO_GDAT to control whether to write\n> >    generation data chunk.\n> \n> Good, that is needed for testing mixed-version behavior.\n> \n> >  * Compare commits with corrected commit dates if present in\n> >    paint_down_to_common().\n> \n> All right, but see the caveat.\n> \n> >  * Update technical documentation.\n> \n> Always a good thing.\n> \n> >  * Handle mixed graph version commit chains.\n> \n> Where by \"version\" you mean generation number version - the commit-graph\n> version number unfortunately needs to stay the same...\n> \n\nYes, clarified.\n\n> >  * Improve commit messages for\n>                                 ^^^^^^\n> Something missing in this point, the sentence ends abruptly.\n\nI didn't finish the sentence. Meant to say:\n\n- Improve commit messages for \"commit-graph: fix regression when computing bloom filter\", \"commit-graph: consolidate fill_commit_graph_info\",\n> \n> >  * Revert unnecessary whitespace changes.\n> \n> Thanks.\n> \n> >  * Split uint_32 -> timestamp_t change into a new commit.\n> \n> It is usually better to keep the commits small.  Good.\n> \n> \n> Good work!\n> \n> ...\n> --\n> Jakub Narębski\n"},{"id":"403933","messageId":"85imdgxck6.fsf@gmail.com","threadId":"53933","inReplyTo":"e6738672349254c6405f7dde48f612b82af9299f.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 02/11] revision: parse parent in indegree_walk_step()","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-18T14:18:17Z","receivedAt":"2020-08-18T14:18:24Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> In indegree_walk_step(), we add unvisited parents to the indegree queue.\n> However, parents are not guaranteed to be parsed. As the indegree queue\n> sorts by generation number, let's parse parents before inserting them to\n> ensure the correct priority order.\n\nAll right, we need to have commit parsed to have correct value for its\ngeneration number.\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  revision.c | 3 +++\n>  1 file changed, 3 insertions(+)\n>\n> diff --git a/revision.c b/revision.c\n> index 3dcf689341..ecf757c327 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -3363,6 +3363,9 @@ static void indegree_walk_step(struct rev_info *revs)\n>  \t\tstruct commit *parent = p->item;\n>  \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n>\n> +\t\tif (parse_commit_gently(parent, 1) < 0)\n> +\t\t\treturn;\n> +\n\nAll right, this is exactly what is done in this function for commit 'c'\ntaken from indegree_queue, whose parents we process here:\n\n\tif (parse_commit_gently(c, 1) < 0)\n\t\treturn;\n\n>  \t\tif (*pi)\n>  \t\t\t(*pi)++;\n>  \t\telse\n\nLooks good to me.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404050","messageId":"857dtuo71v.fsf@gmail.com","threadId":"53933","inReplyTo":"18d5864f81e89585cc94cd12eca166a9d8b929a5.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 03/11] commit-graph: consolidate fill_commit_graph_info","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-19T17:54:20Z","receivedAt":"2020-08-19T17:54:38Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> Both fill_commit_graph_info() and fill_commit_in_graph() parse\n> information present in commit data chunk. Let's simplify the\n> implementation by calling fill_commit_graph_info() within\n> fill_commit_in_graph().\n>\n> The test 'generate tar with future mtime' creates a commit with commit\n> time of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\n> generation number (within CDAT chunk) and has undefined behavior.\n>\n> The test used to pass\n\nCould you please tell us how does this test starts to fail without the\nchange to the test described there?  What is the error message, etc.?\n\n>                       as fill_commit_in_graph() guarantees the values of\n                           ^^^^^^^^^^^^^^^^^^^^^^\n\ns/fill_commit_in_graph()/fill_commit_graph_info()/\n\nIt is fill_commit_graph_info() that changes its behavior in this patch.\n\n> graph position and generation number, and did not load timestamp.\n> However, with corrected commit date we will need load the timestamp as\n> well to populate the generation number.\n>\n> Let's fix the test by setting a timestamp of (2 ^ 34 - 1) seconds.\n\nI think this commit should be split into two commits:\n- fix to the 'generate tar with future mtime' test\n- simplify implementation of fill_commit_in_graph()\n\nThe test 'generate tar with future mtime' in t/t5000-tar-tree.sh creates\na commit with commit time of (2 ^ 36 - 1) seconds since EPOCH\n(68719476737). However, the commit-graph file format version 1 provides\nonly 34-bits for storing committer date (32 + 2 bits), not 64-bits.\nTherefore maximum far in the future commit time can only be at most\n(2 ^ 34 - 1) seconds since EPOCH, as Stolee said in commet for v1\nof this series.\n\nThis \"limitation\" is not a problem in practice, because the maximum\ntimestamp allowed takes place in the year 2514. I hope at that time\nthere would be no Git version in use that still crashes on changing the\nversion field in the commit-graph format -- then we can simply get rid\nof storing topological levels (generation number v1) in those 30 bits of\nCDAT chunk and use full 64 bits for committer date.\n\nGit does not perform any bounds checking for committer date value in\nwrite_graph_chunk_data():\n\n\tuint32_t packedDate[2];\n\n\t/* ... */\n\n\tif (sizeof((*list)->date) > 4)\n\t\tpackedDate[0] = htonl(((*list)->date >> 32) & 0x3);\n\telse\n\t\tpackedDate[0] = 0;\n\n\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n\n\tpackedDate[1] = htonl((*list)->date);\n\thashwrite(f, packedDate, 8);\n\nThis means that the date is trimmed to 34 bits on save discarding most\nsignificant bits, assuming that unsigned overflow simply discards most\nsignificant bits truncating the signed (?) value.\n\nIn this case running the test with GIT_TEST_COMMIT_GRAPH=1 would lead to\nerrors, as the committer date read from the commit graph would be\nincorrect, and therefore generation number v2 would also be incorrect.\n\n\nI don't quite understand however how second part of this patch in its\ncurrent iteration, namely simplifing the implementation of\nfill_commit_in_graph() makes this bug / error shows...\n\nDo I understand it correctly that before this change the committer date\nwould always be parsed out of the commit object, instead of reading it\nfrom the commit-graph file?  However the only user of static\nfill_commit_in_graph() is the parse_commit_in_graph(), which in turn is\nused by parse_commit_gently(); but fill_commit_in_graph() read commit\ndate from commit-graph before this change... color me confused.\n\nAh, after the change fill_commit_graph_info() changes its behavior, not\nfill_commit_in_graph() as said in the commit message. Before this commit\nit used to only load graph position and generation number, and did not\nload the timestamp. The function fill_commit_graph_info() is used in\nturn by public-facing load_commit_graph_info():\n\n  /*\n   * It is possible that we loaded commit contents from the commit buffer,\n   * but we also want to ensure the commit-graph content is correctly\n   * checked and filled. Fill the graph_pos and generation members of\n   * the given commit.\n   */\n  void load_commit_graph_info(struct repository *r, struct commit *item);\n\nThis function is used in turn by get_bloom_filter(), contains_tag_algo()\nand parse_commit_buffer(), change in any of which behavior can lead to\nfailing 'generate tar with future mtime' test.\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c      | 29 +++++++++++------------------\n>  t/t5000-tar-tree.sh |  4 ++--\n>  2 files changed, 13 insertions(+), 20 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index ace7400a1a..af8d9cc45e 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -725,15 +725,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \tconst unsigned char *commit_data;\n>  \tstruct commit_graph_data *graph_data;\n>  \tuint32_t lex_index;\n> +\tuint64_t date_high, date_low;\n>\n>  \twhile (pos < g->num_commits_in_base)\n>  \t\tg = g->base_graph;\n>\n> +\tif (pos >= g->num_commits + g->num_commits_in_base)\n> +\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n> +\n\nAll right, if we want to use fill_commit_graph_info() function to load\nthe graph data (graph position and generation number) in the\nfill_commit_in_graph() we need to perform this check.\n\nI'd think that this check should be here from the beginning, just in\ncase.\n\n\nSidenote: I wonder if it would be good idea to print more information in\nthe above error message, for example:\n\n\tdie(_(\"invalid commit position %ld. commit-graph '%s' is likely corrupt\"),\n        pos, g->filename);\n\nBut this is unrelated thing, tangential to this change, and it might not\nadd anything useful.\n\n>  \tlex_index = pos - g->num_commits_in_base;\n>  \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n>\n>  \tgraph_data = commit_graph_data_at(item);\n>  \tgraph_data->graph_pos = pos;\n> +\n> +\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> +\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> +\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> +\n\nI think this change, moving loading of commit date from commit-graph out\nof fill_commit_in_graph() and into fill_commit_graph_info(), is in my\nopinion a bit inadequatly described in the commit message. As I\nunderstand it this change prepares fill_commit_graph_info() for\ngeneration number v2, that is Corrected Commit Date, where loading\ncommit date from CDAT together with loading offset from GDAT would be\nnecessary to correctly set the 'generation' field of 'struct\ncommit_graph_data' (on the commit_graph_data_slab).\n\nI'm not sure if it would be worth it splitting this refactoring change\n(Move Statements into Function) into a separate patch -- it would split\nthis commit into three, changing 11 part series into 13 part series.\n\n\nNote that we might want to update the description of\nload_commit_graph_info() in commit-graph.h to include that it\nincidentally loads commit date from the commit-graph.  Butthis might be\nnot worth it -- it is a side effect, not the major goal of this\nfunction.\n\n>  \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n>  }\n>\n> @@ -748,38 +757,22 @@ static int fill_commit_in_graph(struct repository *r,\n>  {\n>  \tuint32_t edge_value;\n>  \tuint32_t *parent_data_ptr;\n> -\tuint64_t date_low, date_high;\n>  \tstruct commit_list **pptr;\n> -\tstruct commit_graph_data *graph_data;\n>  \tconst unsigned char *commit_data;\n>  \tuint32_t lex_index;\n>\n>  \twhile (pos < g->num_commits_in_base)\n>  \t\tg = g->base_graph;\n>\n> -\tif (pos >= g->num_commits + g->num_commits_in_base)\n> -\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n> +\tfill_commit_graph_info(item, g, pos);\n\nAll right, the check got moved into fill_commit_graph_info().\n\n>\n> -\t/*\n> -\t * Store the \"full\" position, but then use the\n> -\t * \"local\" position for the rest of the calculation.\n> -\t */\n> -\tgraph_data = commit_graph_data_at(item);\n> -\tgraph_data->graph_pos = pos;\n\nAll right, 'graph_pos' field in the graph data (on commit slab) got\nfilled by just called load_commit_graph_info().\n\n>  \tlex_index = pos - g->num_commits_in_base;\n> -\n> -\tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n> +\tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n\nAll right, unrelated cleanup in the neighbourhood.\n\n>\n>  \titem->object.parsed = 1;\n>\n>  \tset_commit_tree(item, NULL);\n>\n> -\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> -\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> -\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> -\n\nAll right, this code got moved down the call chain into just called\nload_commit_graph_info().\n\n> -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> -\n\n\nAll right, 'generation' field in the graph data (on commit slab) got\nfilled by load_commit_graph_info() called at the beginning of the\nfunction.\n\n\n\n>  \tpptr = &item->parents;\n>\n>  \tedge_value = get_be32(commit_data + g->hash_len);\n> diff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\n> index 37655a237c..1986354fc3 100755\n> --- a/t/t5000-tar-tree.sh\n> +++ b/t/t5000-tar-tree.sh\n> @@ -406,7 +406,7 @@ test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n>  \trm -f .git/index &&\n>  \techo content >file &&\n>  \tgit add file &&\n> -\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n> +\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n>  \t\tgit commit -m \"tempori parendum\"\n>  '\n>\n> @@ -415,7 +415,7 @@ test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n>  '\n>\n>  test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n> -\techo 4147 >expect &&\n> +\techo 2514 >expect &&\n>  \ttar_info future.tar | cut -d\" \" -f2 >actual &&\n>  \ttest_cmp expect actual\n>  '\n\nLooks good to me.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404119","messageId":"20200821041124.GA39355@Abhishek-Arch","threadId":"53933","inReplyTo":"857dtuo71v.fsf@gmail.com","subject":"Re: [PATCH v3 03/11] commit-graph: consolidate fill_commit_graph_info","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-08-21T04:11:24Z","receivedAt":"2020-08-21T04:13:47Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Wed, Aug 19, 2020 at 07:54:20PM +0200, Jakub Narębski wrote:\n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > Both fill_commit_graph_info() and fill_commit_in_graph() parse\n> > information present in commit data chunk. Let's simplify the\n> > implementation by calling fill_commit_graph_info() within\n> > fill_commit_in_graph().\n> >\n> > The test 'generate tar with future mtime' creates a commit with commit\n> > time of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\n> > generation number (within CDAT chunk) and has undefined behavior.\n> >\n> > The test used to pass\n> \n> Could you please tell us how does this test starts to fail without the\n> change to the test described there?  What is the error message, etc.?\n> \n\nHere's what I revised the commit message to:\n\ncommit-graph: consolidate fill_commit_graph_info\n\nBoth fill_commit_graph_info() and fill_commit_in_graph() parse\ninformation present in commit data chunk. Let's simplify the\nimplementation by calling fill_commit_graph_info() within\nfill_commit_in_graph().\n\nThe test 'generate tar with future mtime' creates a commit with commit\ntime of (2 ^ 36 + 1) seconds since EPOCH. The CDAT chunk provides\n34-bits for storing commiter date, thus committer time overflows into\ngeneration number (within CDAT chunk) and has undefined behavior.\n\nThe test used to pass as fill_commit_graph_info() would not set struct\nmember `date` of struct commit and loads committer date from the object\ndatabase, generating a tar file with the expected mtime.\n\nHowever, with corrected commit date, we will load the committer date\nfrom CDAT chunk (truncated to lower 34-bits) to populate the generation\nnumber. Thus, fill_commit_graph_info() sets date and generates tar file\nwith the truncated mtime and the test fails.\n\nLet's fix the test by setting a timestamp of (2 ^ 34 - 1) seconds, which\nwill not be truncated.\n\n> >                       as fill_commit_in_graph() guarantees the values of\n>                            ^^^^^^^^^^^^^^^^^^^^^^\n> \n> s/fill_commit_in_graph()/fill_commit_graph_info()/\n> \n> ...\n> \n> Ah, after the change fill_commit_graph_info() changes its behavior, not\n> fill_commit_in_graph() as said in the commit message. Before this commit\n> it used to only load graph position and generation number, and did not\n> load the timestamp. The function fill_commit_graph_info() is used in\n> turn by public-facing load_commit_graph_info():\n> \n\nThat's exactly it. I should have elaborated better in the commit\nmessage. Thanks for the through investigation.\n\n>   /*\n>    * It is possible that we loaded commit contents from the commit buffer,\n>    * but we also want to ensure the commit-graph content is correctly\n>    * checked and filled. Fill the graph_pos and generation members of\n>    * the given commit.\n>    */\n>   void load_commit_graph_info(struct repository *r, struct commit *item);\n> \n> ...\n> \n> Looks good to me.\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"404129","messageId":"85sgcgmf80.fsf@gmail.com","threadId":"53933","inReplyTo":"6a0cde983d9ed20f043a4977313d714154602012.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 04/11] commit-graph: consolidate compare_commits_by_gen","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-21T11:05:19Z","receivedAt":"2020-08-21T11:05:24Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> Comparing commits by generation has been independently defined twice, in\n> commit-reach and commit. Let's simplify the implementation by moving\n> compare_commits_by_gen() to commit-graph.\n\nAll right, seems reasonable.\n\nThough it might be not obvious that the second repetition of code\ncomparing commits by generation is part of commit.c's\ncompare_commits_by_gen_then_commit_date().\n\nIs't it micro-pessimization though, or can the compiler inline function\nacross different files?  On the other hand it reduces code duplication...\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> Reviewed-by: Taylor Blau <me@ttaylorr.com>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 15 +++++++++++++++\n>  commit-graph.h |  2 ++\n>  commit-reach.c | 15 ---------------\n>  commit.c       |  9 +++------\n>  4 files changed, 20 insertions(+), 21 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index af8d9cc45e..fb6e2bf18f 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -112,6 +112,21 @@ uint32_t commit_graph_generation(const struct commit *c)\n>  \treturn data->generation;\n>  }\n>\n> +int compare_commits_by_gen(const void *_a, const void *_b)\n> +{\n> +\tconst struct commit *a = _a, *b = _b;\n> +\tconst uint32_t generation_a = commit_graph_generation(a);\n> +\tconst uint32_t generation_b = commit_graph_generation(b);\n\nAll right, this function used protected access to generation number of a\ncommit, that is it correctly handles the case where commit '_a' and/or\n'_b' are new enough to be not present in the commit graph.\n\nThat is why we cannot simply use commit_gen_cmp(), that is the function\nused for sorting during `git commit-graph write --reachable --changed-paths`,\nbecause after 1st patch it access the slab directly.\n\n> +\n> +\t/* older commits first */\n\nNice!  Thanks for adding this comment.\n\nThough it might be good idea to add this comment also to the header\nfile, commit-graph.h, because the fact that compare_commits_by_gen()\nand compare_commits_by_gen_then_commit_date() sort in different\norder is not something that we can see from their names.  Well,\nthey have slightly different sigatures...\n\n> +\tif (generation_a < generation_b)\n> +\t\treturn -1;\n> +\telse if (generation_a > generation_b)\n> +\t\treturn 1;\n> +\n> +\treturn 0;\n> +}\n> +\n>  static struct commit_graph_data *commit_graph_data_at(const struct commit *c)\n>  {\n>  \tunsigned int i, nth_slab;\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 09a97030dc..701e3d41aa 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -146,4 +146,6 @@ struct commit_graph_data {\n>   */\n>  uint32_t commit_graph_generation(const struct commit *);\n>  uint32_t commit_graph_position(const struct commit *);\n> +\n> +int compare_commits_by_gen(const void *_a, const void *_b);\n\nAll right.\n\n>  #endif\n> diff --git a/commit-reach.c b/commit-reach.c\n> index efd5925cbb..c83cc291e7 100644\n> --- a/commit-reach.c\n> +++ b/commit-reach.c\n> @@ -561,21 +561,6 @@ int commit_contains(struct ref_filter *filter, struct commit *commit,\n>  \treturn repo_is_descendant_of(the_repository, commit, list);\n>  }\n>\n> -static int compare_commits_by_gen(const void *_a, const void *_b)\n> -{\n> -\tconst struct commit *a = *(const struct commit * const *)_a;\n> -\tconst struct commit *b = *(const struct commit * const *)_b;\n> -\n> -\tuint32_t generation_a = commit_graph_generation(a);\n> -\tuint32_t generation_b = commit_graph_generation(b);\n> -\n> -\tif (generation_a < generation_b)\n> -\t\treturn -1;\n> -\tif (generation_a > generation_b)\n> -\t\treturn 1;\n> -\treturn 0;\n> -}\n\nAll right, commit-reach.c includes commit-graph.h, so now it simply uses\ncompare_commits_by_gen() that was copied to commit-graph.c.\n\n> -\n>  int can_all_from_reach_with_flag(struct object_array *from,\n>  \t\t\t\t unsigned int with_flag,\n>  \t\t\t\t unsigned int assign_flag,\n> diff --git a/commit.c b/commit.c\n> index 4ce8cb38d5..bd6d5e587f 100644\n> --- a/commit.c\n> +++ b/commit.c\n> @@ -731,14 +731,11 @@ int compare_commits_by_author_date(const void *a_, const void *b_,\n>  int compare_commits_by_gen_then_commit_date(const void *a_, const void *b_, void *unused)\n>  {\n>  \tconst struct commit *a = a_, *b = b_;\n> -\tconst uint32_t generation_a = commit_graph_generation(a),\n> -\t\t       generation_b = commit_graph_generation(b);\n> +\tint ret_val = compare_commits_by_gen(a_, b_);\n>\n>  \t/* newer commits first */\n\nMaybe this comment should be put in the header file, near this functionn\ndeclaration?\n\n> -\tif (generation_a < generation_b)\n> -\t\treturn 1;\n> -\telse if (generation_a > generation_b)\n> -\t\treturn -1;\n> +\tif (ret_val)\n> +\t\treturn -ret_val;\n\nAll right, this handles reversed sorting order of compare_commits_by_gen().\n\n>\n>  \t/* use date as a heuristic when generations are equal */\n>  \tif (a->date < b->date)\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404137","messageId":"85d03km98l.fsf@gmail.com","threadId":"53933","inReplyTo":"6be759a9542114e4de41422efa18491085e19682.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 05/11] commit-graph: return 64-bit generation number","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-21T13:14:34Z","receivedAt":"2020-08-21T13:14:42Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> In a preparatory step, let's return timestamp_t values from\n> commit_graph_generation(), use timestamp_t for local variables\n\nAll right, this is all good.\n\n> and define GENERATION_NUMBER_INFINITY as (2 ^ 63 - 1) instead.\n\nThis needs more detailed examination.  There are two similar constants,\nGENERATION_NUMBER_INFINITY and GENERATION_NUMBER_MAX.  The former is\nused for newest commits outside the commit-graph, while the latter is\nmaximum number that commits in the commit-graph can have (because of the\nstorage limitations).  We therefore need GENERATION_NUMBER_INFINITY\nto be larger than GENERATION_NUMBER_MAX, and it is (and was).\n\nThe GENERATION_NUMBER_INFINITY is because of the above requirement\ntraditionally taken as maximum value that can be represented in the data\ntype used to store commit's generation number _in memory_, but it can be\nless.  For timestamp_t the maximum value that can be represented\nis (2 ^ 63 - 1).\n\nAll right then.\n\n>\n\nThe commit message says nothing about the new symbolic constant\nGENERATION_NUMBER_V1_INFINITY, though.\n\nI'm not sure it is even needed (see comments below).\n\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 18 +++++++++---------\n>  commit-graph.h |  4 ++--\n>  commit-reach.c | 32 ++++++++++++++++----------------\n>  commit-reach.h |  2 +-\n>  commit.h       |  3 ++-\n>  revision.c     | 10 +++++-----\n>  upload-pack.c  |  2 +-\n>  7 files changed, 36 insertions(+), 35 deletions(-)\n\nI hope that changing the type returned by commit_graph_generation() and\nstored in 'generation' field of `struct commit_graph_data` would mean\nthat the compiler or at least the linter would catch all the places that\nneed updating the type.\n\nJust in case, I have performed a simple code search and it agrees with\nthe above list (one search result missing, in commit.c, was handled by\nprevious patch).\n\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index fb6e2bf18f..7f9f858577 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -99,7 +99,7 @@ uint32_t commit_graph_position(const struct commit *c)\n>  \treturn data ? data->graph_pos : COMMIT_NOT_FROM_GRAPH;\n>  }\n>\n> -uint32_t commit_graph_generation(const struct commit *c)\n> +timestamp_t commit_graph_generation(const struct commit *c)\n>  {\n>  \tstruct commit_graph_data *data =\n>  \t\tcommit_graph_data_slab_peek(&commit_graph_data_slab, c);\n> @@ -115,8 +115,8 @@ uint32_t commit_graph_generation(const struct commit *c)\n>  int compare_commits_by_gen(const void *_a, const void *_b)\n>  {\n>  \tconst struct commit *a = _a, *b = _b;\n> -\tconst uint32_t generation_a = commit_graph_generation(a);\n> -\tconst uint32_t generation_b = commit_graph_generation(b);\n> +\tconst timestamp_t generation_a = commit_graph_generation(a);\n> +\tconst timestamp_t generation_b = commit_graph_generation(b);\n>\n>  \t/* older commits first */\n>  \tif (generation_a < generation_b)\n> @@ -159,8 +159,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>  \tconst struct commit *a = *(const struct commit **)va;\n>  \tconst struct commit *b = *(const struct commit **)vb;\n>\n> -\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> -\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n> +\tconst timestamp_t generation_a = commit_graph_data_at(a)->generation;\n> +\tconst timestamp_t generation_b = commit_graph_data_at(b)->generation;\n>  \t/* lower generation commits first */\n>  \tif (generation_a < generation_b)\n>  \t\treturn -1;\n\nAll right.\n\n> @@ -1338,7 +1338,7 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n\nShouldn't this be\n\n-  \t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n+  \t\ttimestamp_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n\n\n>\n>  \t\tdisplay_progress(ctx->progress, i + 1);\n> -\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n> +\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n\nThen there would be no need for this change, isn't it?\n\n>  \t\t    generation != GENERATION_NUMBER_ZERO)\n>  \t\t\tcontinue;\n>\n> @@ -1352,7 +1352,7 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n>  \t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n>\n> -\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n> +\t\t\t\tif (generation == GENERATION_NUMBER_V1_INFINITY ||\n\nAnd this one either.\n\n>  \t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n>  \t\t\t\t\tall_parents_computed = 0;\n>  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n> @@ -2355,8 +2355,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n>  \tfor (i = 0; i < g->num_commits; i++) {\n>  \t\tstruct commit *graph_commit, *odb_commit;\n>  \t\tstruct commit_list *graph_parents, *odb_parents;\n> -\t\tuint32_t max_generation = 0;\n> -\t\tuint32_t generation;\n> +\t\ttimestamp_t max_generation = 0;\n> +\t\ttimestamp_t generation;\n\nAll right.\n\n>\n>  \t\tdisplay_progress(progress, i + 1);\n>  \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 701e3d41aa..430bc830bb 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -138,13 +138,13 @@ void disable_commit_graph(struct repository *r);\n>\n>  struct commit_graph_data {\n>  \tuint32_t graph_pos;\n> -\tuint32_t generation;\n> +\ttimestamp_t generation;\n>  };\n\nAll right; this is the main part of this change.\n\n>\n>  /*\n>   * Commits should be parsed before accessing generation, graph positions.\n>   */\n> -uint32_t commit_graph_generation(const struct commit *);\n> +timestamp_t commit_graph_generation(const struct commit *);\n>  uint32_t commit_graph_position(const struct commit *);\n\nAs is this one.\n\n>\n>  int compare_commits_by_gen(const void *_a, const void *_b);\n> diff --git a/commit-reach.c b/commit-reach.c\n> index c83cc291e7..470bc80139 100644\n> --- a/commit-reach.c\n> +++ b/commit-reach.c\n> @@ -32,12 +32,12 @@ static int queue_has_nonstale(struct prio_queue *queue)\n>  static struct commit_list *paint_down_to_common(struct repository *r,\n>  \t\t\t\t\t\tstruct commit *one, int n,\n>  \t\t\t\t\t\tstruct commit **twos,\n> -\t\t\t\t\t\tint min_generation)\n> +\t\t\t\t\t\ttimestamp_t min_generation)\n>  {\n>  \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n>  \tstruct commit_list *result = NULL;\n>  \tint i;\n> -\tuint32_t last_gen = GENERATION_NUMBER_INFINITY;\n> +\ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n>\n>  \tif (!min_generation)\n>  \t\tqueue.compare = compare_commits_by_commit_date;\n\nAll right.\n\n> @@ -58,10 +58,10 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n>  \t\tstruct commit *commit = prio_queue_get(&queue);\n>  \t\tstruct commit_list *parents;\n>  \t\tint flags;\n> -\t\tuint32_t generation = commit_graph_generation(commit);\n> +\t\ttimestamp_t generation = commit_graph_generation(commit);\n>\n>  \t\tif (min_generation && generation > last_gen)\n> -\t\t\tBUG(\"bad generation skip %8x > %8x at %s\",\n> +\t\t\tBUG(\"bad generation skip %\"PRItime\" > %\"PRItime\" at %s\",\n>  \t\t\t    generation, last_gen,\n>  \t\t\t    oid_to_hex(&commit->object.oid));\n>  \t\tlast_gen = generation;\n\nAll right.\n\n> @@ -177,12 +177,12 @@ static int remove_redundant(struct repository *r, struct commit **array, int cnt\n>  \t\trepo_parse_commit(r, array[i]);\n>  \tfor (i = 0; i < cnt; i++) {\n>  \t\tstruct commit_list *common;\n> -\t\tuint32_t min_generation = commit_graph_generation(array[i]);\n> +\t\ttimestamp_t min_generation = commit_graph_generation(array[i]);\n>\n>  \t\tif (redundant[i])\n>  \t\t\tcontinue;\n>  \t\tfor (j = filled = 0; j < cnt; j++) {\n> -\t\t\tuint32_t curr_generation;\n> +\t\t\ttimestamp_t curr_generation;\n>  \t\t\tif (i == j || redundant[j])\n>  \t\t\t\tcontinue;\n>  \t\t\tfilled_index[filled] = j;\n\nAll right.\n\n> @@ -321,7 +321,7 @@ int repo_in_merge_bases_many(struct repository *r, struct commit *commit,\n>  {\n>  \tstruct commit_list *bases;\n>  \tint ret = 0, i;\n> -\tuint32_t generation, min_generation = GENERATION_NUMBER_INFINITY;\n> +\ttimestamp_t generation, min_generation = GENERATION_NUMBER_INFINITY;\n>\n>  \tif (repo_parse_commit(r, commit))\n>  \t\treturn ret;\n\nAll right,\n\n> @@ -470,7 +470,7 @@ static int in_commit_list(const struct commit_list *want, struct commit *c)\n>  static enum contains_result contains_test(struct commit *candidate,\n>  \t\t\t\t\t  const struct commit_list *want,\n>  \t\t\t\t\t  struct contains_cache *cache,\n> -\t\t\t\t\t  uint32_t cutoff)\n> +\t\t\t\t\t  timestamp_t cutoff)\n\nAll right.\n\n(Sidenote: this one I have missed in my simple search.)\n\n>  {\n>  \tenum contains_result *cached = contains_cache_at(cache, candidate);\n>\n> @@ -506,11 +506,11 @@ static enum contains_result contains_tag_algo(struct commit *candidate,\n>  {\n>  \tstruct contains_stack contains_stack = { 0, 0, NULL };\n>  \tenum contains_result result;\n> -\tuint32_t cutoff = GENERATION_NUMBER_INFINITY;\n> +\ttimestamp_t cutoff = GENERATION_NUMBER_INFINITY;\n>  \tconst struct commit_list *p;\n>\n>  \tfor (p = want; p; p = p->next) {\n> -\t\tuint32_t generation;\n> +\t\ttimestamp_t generation;\n>  \t\tstruct commit *c = p->item;\n>  \t\tload_commit_graph_info(the_repository, c);\n>  \t\tgeneration = commit_graph_generation(c);\n\nAll right.\n\n> @@ -565,7 +565,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n>  \t\t\t\t unsigned int with_flag,\n>  \t\t\t\t unsigned int assign_flag,\n>  \t\t\t\t time_t min_commit_date,\n> -\t\t\t\t uint32_t min_generation)\n> +\t\t\t\t timestamp_t min_generation)\n>  {\n>  \tstruct commit **list = NULL;\n>  \tint i;\n> @@ -666,13 +666,13 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n>  \ttime_t min_commit_date = cutoff_by_min_date ? from->item->date : 0;\n>  \tstruct commit_list *from_iter = from, *to_iter = to;\n>  \tint result;\n> -\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n> +\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n>\n>  \twhile (from_iter) {\n>  \t\tadd_object_array(&from_iter->item->object, NULL, &from_objs);\n>\n>  \t\tif (!parse_commit(from_iter->item)) {\n> -\t\t\tuint32_t generation;\n> +\t\t\ttimestamp_t generation;\n>  \t\t\tif (from_iter->item->date < min_commit_date)\n>  \t\t\t\tmin_commit_date = from_iter->item->date;\n>\n> @@ -686,7 +686,7 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n>\n>  \twhile (to_iter) {\n>  \t\tif (!parse_commit(to_iter->item)) {\n> -\t\t\tuint32_t generation;\n> +\t\t\ttimestamp_t generation;\n>  \t\t\tif (to_iter->item->date < min_commit_date)\n>  \t\t\t\tmin_commit_date = to_iter->item->date;\n>\n\nAll right.\n\n> @@ -726,13 +726,13 @@ struct commit_list *get_reachable_subset(struct commit **from, int nr_from,\n>  \tstruct commit_list *found_commits = NULL;\n>  \tstruct commit **to_last = to + nr_to;\n>  \tstruct commit **from_last = from + nr_from;\n> -\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n> +\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n>  \tint num_to_find = 0;\n>\n>  \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n>\n>  \tfor (item = to; item < to_last; item++) {\n> -\t\tuint32_t generation;\n> +\t\ttimestamp_t generation;\n>  \t\tstruct commit *c = *item;\n>\n>  \t\tparse_commit(c);\n\nAll right.\n\n> diff --git a/commit-reach.h b/commit-reach.h\n> index b49ad71a31..148b56fea5 100644\n> --- a/commit-reach.h\n> +++ b/commit-reach.h\n> @@ -87,7 +87,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n>  \t\t\t\t unsigned int with_flag,\n>  \t\t\t\t unsigned int assign_flag,\n>  \t\t\t\t time_t min_commit_date,\n> -\t\t\t\t uint32_t min_generation);\n> +\t\t\t\t timestamp_t min_generation);\n>  int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n>  \t\t       int commit_date_cutoff);\n>\n\nAll right.\n\n> diff --git a/commit.h b/commit.h\n> index e901538909..bc0732a4fe 100644\n> --- a/commit.h\n> +++ b/commit.h\n> @@ -11,7 +11,8 @@\n>  #include \"commit-slab.h\"\n>\n>  #define COMMIT_NOT_FROM_GRAPH 0xFFFFFFFF\n> -#define GENERATION_NUMBER_INFINITY 0xFFFFFFFF\n> +#define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n> +#define GENERATION_NUMBER_V1_INFINITY 0xFFFFFFFF\n>  #define GENERATION_NUMBER_MAX 0x3FFFFFFF\n>  #define GENERATION_NUMBER_ZERO 0\n>\n\nWhy do we even need GENERATION_NUMBER_V1_INFINITY?  It is about marking\nout-of-graph commits, and it is about in-memory storage.\n\nWe would need separate GENERATION_NUMBER_V1_MAX and GENERATION_NUMBER_V2_MAX\nbecause of different _on-disk_ storage, or in other words file format\nlimitations.  But that is for the future commit.\n\n> diff --git a/revision.c b/revision.c\n> index ecf757c327..411852468b 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -3290,7 +3290,7 @@ define_commit_slab(indegree_slab, int);\n>  define_commit_slab(author_date_slab, timestamp_t);\n>\n>  struct topo_walk_info {\n> -\tuint32_t min_generation;\n> +\ttimestamp_t min_generation;\n>  \tstruct prio_queue explore_queue;\n>  \tstruct prio_queue indegree_queue;\n>  \tstruct prio_queue topo_queue;\n\nAll right.\n\n> @@ -3336,7 +3336,7 @@ static void explore_walk_step(struct rev_info *revs)\n>  }\n>\n>  static void explore_to_depth(struct rev_info *revs,\n> -\t\t\t     uint32_t gen_cutoff)\n> +\t\t\t     timestamp_t gen_cutoff)\n>  {\n>  \tstruct topo_walk_info *info = revs->topo_walk_info;\n>  \tstruct commit *c;\n\nAll right.\n\n> @@ -3379,7 +3379,7 @@ static void indegree_walk_step(struct rev_info *revs)\n>  }\n>\n>  static void compute_indegrees_to_depth(struct rev_info *revs,\n> -\t\t\t\t       uint32_t gen_cutoff)\n> +\t\t\t\t       timestamp_t gen_cutoff)\n>  {\n>  \tstruct topo_walk_info *info = revs->topo_walk_info;\n>  \tstruct commit *c;\n> @@ -3437,7 +3437,7 @@ static void init_topo_walk(struct rev_info *revs)\n>  \tinfo->min_generation = GENERATION_NUMBER_INFINITY;\n>  \tfor (list = revs->commits; list; list = list->next) {\n>  \t\tstruct commit *c = list->item;\n> -\t\tuint32_t generation;\n> +\t\ttimestamp_t generation;\n>\n>  \t\tif (parse_commit_gently(c, 1))\n>  \t\t\tcontinue;\n\nAll right.\n\n> @@ -3498,7 +3498,7 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n>  \tfor (p = commit->parents; p; p = p->next) {\n>  \t\tstruct commit *parent = p->item;\n>  \t\tint *pi;\n> -\t\tuint32_t generation;\n> +\t\ttimestamp_t generation;\n>\n>  \t\tif (parent->object.flags & UNINTERESTING)\n>  \t\t\tcontinue;\n\nAll right.\n\n> diff --git a/upload-pack.c b/upload-pack.c\n> index 80ad9a38d8..bcb8b5dfda 100644\n> --- a/upload-pack.c\n> +++ b/upload-pack.c\n> @@ -497,7 +497,7 @@ static int got_oid(struct upload_pack_data *data,\n>\n>  static int ok_to_give_up(struct upload_pack_data *data)\n>  {\n> -\tuint32_t min_generation = GENERATION_NUMBER_ZERO;\n> +\ttimestamp_t min_generation = GENERATION_NUMBER_ZERO;\n>\n>  \tif (!data->have_obj.nr)\n>  \t\treturn 0;\n\nAll right.\n\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404180","messageId":"85d03jlu05.fsf@gmail.com","threadId":"53933","inReplyTo":"b347dbb01b9254ab8d79fbbd0f7c2b637efde62e.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 06/11] commit-graph: add a slab to store topological levels","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-21T18:43:38Z","receivedAt":"2020-08-21T18:43:45Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello,\n\n\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> As we are writing topological levels to commit data chunk to ensure\n> backwards compatibility with \"Old\" Git and the member `generation` of\n> struct commit_graph_data will store corrected commit date in a later\n> commit, let's introduce a commit-slab to store topological levels while\n> writing commit-graph.\n\nIn my opinion the above it would be easier to follow if rephrased in the\nfollowing way:\n\n  In a later commit we will introduce corrected commit date as the\n  generation number v2.  This value will be stored in the new separate\n  GDAT chunk.  However to ensure backwards compatibility with \"Old\" Git\n  we need to continue to write generation number v1, which is\n  topological level, to the commit data chunk (CDAT).  This means that\n  we need to compute both versions of generation numbers when writing\n  the commit-graph file.  Let's therefore introduce a commit-slab\n  to store topological levels; corrected commit date will be stored\n  in the member `generation` of struct commit_graph_data.\n\nWhat do you think?\n\n\nBy the way, do I understand it correctly that in backward-compatibility\nmode (that is, in mixed-version environment where at least some\ncommit-graph files were written by \"Old\" Git and are lacking GDAT chunk\nand generation number v2 data) the `generation` member of commit graph\ndata chunk will be populated and will store generation number v1, that\nis topological level? And that the commit-slab for topological levels is\nonly there for writing and re-writing?\n\n>\n> When Git creates a split commit-graph, it takes advantage of the\n> generation values that have been computed already and present in\n> existing commit-graph files.\n>\n> So, let's add a pointer to struct commit_graph to the topological level\n> commit-slab and populate it with topological levels while writing a\n> split commit-graph.\n\nAll right, looks sensible.\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 47 ++++++++++++++++++++++++++++++++---------------\n>  commit-graph.h |  1 +\n>  commit.h       |  1 +\n>  3 files changed, 34 insertions(+), 15 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 7f9f858577..a2f15b2825 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -64,6 +64,8 @@ void git_test_write_commit_graph_or_die(void)\n>  /* Remember to update object flag allocation in object.h */\n>  #define REACHABLE       (1u<<15)\n>\n> +define_commit_slab(topo_level_slab, uint32_t);\n> +\n\nAll right.\n\nAlso, here we might need GENERATION_NUMBER_V1_INFINITY, but I don't\nthink it would be necessary.\n\n>  /* Keep track of the order in which commits are added to our list. */\n>  define_commit_slab(commit_pos, int);\n>  static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n> @@ -759,6 +761,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n>\n>  \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> +\n> +\tif (g->topo_levels)\n> +\t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n>  }\n\nAll right, here we store topological levels on commit-slab to avoid\nrecomputing them.\n\nDo I understand it correctly that the `topo_levels` member of the `struct\ncommit_graph` would be non-null only when we are updating the\ncommit-graph?\n\n>\n>  static inline void set_commit_tree(struct commit *c, struct tree *t)\n> @@ -953,6 +958,7 @@ struct write_commit_graph_context {\n>  \t\t changed_paths:1,\n>  \t\t order_by_pack:1;\n>\n> +\tstruct topo_level_slab *topo_levels;\n>  \tconst struct split_commit_graph_opts *split_opts;\n>  \tsize_t total_bloom_filter_data_size;\n>  \tconst struct bloom_filter_settings *bloom_settings;\n\nWhy do we need `topo_levels` member *both* in `struct commit_graph` and\nin `struct write_commit_graph_context`?\n\n[After examining the change further I have realized why both are needed,\n and written about the reasoning later in this email.]\n\n\nNote that the commit message talks only about `struct commit_graph`...\n\n> @@ -1094,7 +1100,7 @@ static int write_graph_chunk_data(struct hashfile *f,\n>  \t\telse\n>  \t\t\tpackedDate[0] = 0;\n>\n> -\t\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n> +\t\tpackedDate[0] |= htonl(*topo_level_slab_at(ctx->topo_levels, *list) << 2);\n\nAll right, here we prepare for writing to the CDAT chunk using data that\nis now stored on newly introduced topo_levels slab (either computed, or\ntaken from commit-graph file being rewritten).\n\nAssuming that ctx->topo_levels is not-null, and that the values are\nproperly calculated before this -- and we did compute topological levels\nbefore writing the commit-graph.\n\n>\n>  \t\tpackedDate[1] = htonl((*list)->date);\n>  \t\thashwrite(f, packedDate, 8);\n> @@ -1335,11 +1341,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n>  \t\t\t\t\tctx->commits.nr);\n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n> -\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n> +\t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n\nAll right, so that is why this 'generation' variable was not converted\nto timestamp_t type.\n\n>\n>  \t\tdisplay_progress(ctx->progress, i + 1);\n> -\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n> -\t\t    generation != GENERATION_NUMBER_ZERO)\n> +\t\tif (level != GENERATION_NUMBER_V1_INFINITY &&\n> +\t\t    level != GENERATION_NUMBER_ZERO)\n>  \t\t\tcontinue;\n\nHere we use GENERATION_NUMBER*_INFINITY to check if the commit is\noutside commit-graph files, and therefore we would need its topological\nlevel computed.\n\nHowever, I don't understand how it works.  We have had created the\ncommit_graph_data_at() and use it instead of commit_graph_data_slab_at()\nto provide default values for `struct commit_graph`... but only for\n`graph_pos` member.  It is commit_graph_generation() that returns\nGENERATION_NUMBER_INFINITY for commits not in graph.\n\nBut neither commit_graph_data_at()->generation nor topo_level_slab_at()\nhandles this special case, so I don't see how 'generation' variable can\n*ever* be GENERATION_NUMBER_INFINITY, and 'level' variable can ever be\nGENERATION_NUMBER_V1_INFINITY for commits not in graph.\n\nDoes it work *accidentally*, because the default value for uninitialized\ndata on commit-slab is 0, which matches GENERATION_NUMBER_ZERO?  It\ncertainly looks like it does.  And GENERATION_NUMBER_ZERO is an artifact\nof commit-graph feature development history, namely the short time where\nGit didn't use any generation numbers and stored 0 in the place set for\nit in the commit-graph format...  On the other hand this is not the case\nfor corrected commit date (generation number v2), as it could\n\"legitimately\" be 0 if some root commit (without any parents) had\ncommitterdate of epoch 0, i.e. 1 January 1970 00:00:00 UTC, perhaps\ncaused by malformed but valid commit object.\n\nUgh...\n\n>\n>  \t\tcommit_list_insert(ctx->commits.list[i], &list);\n> @@ -1347,29 +1353,27 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\tstruct commit *current = list->item;\n>  \t\t\tstruct commit_list *parent;\n>  \t\t\tint all_parents_computed = 1;\n> -\t\t\tuint32_t max_generation = 0;\n> +\t\t\tuint32_t max_level = 0;\n>\n>  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n> -\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n> +\t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n>\n> -\t\t\t\tif (generation == GENERATION_NUMBER_V1_INFINITY ||\n> -\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n> +\t\t\t\tif (level == GENERATION_NUMBER_V1_INFINITY ||\n> +\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n>  \t\t\t\t\tall_parents_computed = 0;\n>  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n>  \t\t\t\t\tbreak;\n> -\t\t\t\t} else if (generation > max_generation) {\n> -\t\t\t\t\tmax_generation = generation;\n> +\t\t\t\t} else if (level > max_level) {\n> +\t\t\t\t\tmax_level = level;\n>  \t\t\t\t}\n>  \t\t\t}\n\nThis is the same case as for previous chunk; see the comment above.\n\nThis code checks if parents have generation number / topological level\ncomputed, and tracks maximum value of it among all parents.\n\n>\n>  \t\t\tif (all_parents_computed) {\n> -\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n> -\n> -\t\t\t\tdata->generation = max_generation + 1;\n>  \t\t\t\tpop_commit(&list);\n>\n> -\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n> -\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n> +\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n> +\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n> +\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n\nOK, this is safer way of handling GENERATION_NUMBER*_MAX, especially if\nthis value can be maximum value that can be safely stored in a given\ndata type.  Previously GENERATION_NUMBER_MAX was smaller than maximum\nvalue that can be safely stored in uint32_t, so generation+1 had no\nchance to overflow.  This is no longer the case; the reorganization done\nhere leads to more defensive code (safer).\n\nAll good.  However I think that we should clamp the value of topological\nlevel to the maximum value that can be safely stored *on disk*, in the\n30 bits of the CDAT chunk reserved for generation number v1.  Otherwise\nthe code to write topological level would get more complicated.\n\nIn my opinion the symbolic constant used here should be named\nGENERATION_NUMBER_V1_MAX, and its value should be at most (2 ^ 30 - 1);\nit should be the current value of GENERATION_NUMBER_MAX, that is\n0x3FFFFFFF.\n\n>  \t\t\t}\n>  \t\t}\n>  \t}\n> @@ -2101,6 +2105,7 @@ int write_commit_graph(struct object_directory *odb,\n>  \tuint32_t i, count_distinct = 0;\n>  \tint res = 0;\n>  \tint replace = 0;\n> +\tstruct topo_level_slab topo_levels;\n>\n\nAll right, we will be using topo_level slab for writing the\ncommit-graph, and only for this purpose, so it is good to put it here.\n\n>  \tif (!commit_graph_compatible(the_repository))\n>  \t\treturn 0;\n> @@ -2179,6 +2184,18 @@ int write_commit_graph(struct object_directory *odb,\n>  \t\t}\n>  \t}\n>\n> +\tinit_topo_level_slab(&topo_levels);\n> +\tctx->topo_levels = &topo_levels;\n> +\n> +\tif (ctx->r->objects->commit_graph) {\n> +\t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n> +\n> +\t\twhile (g) {\n> +\t\t\tg->topo_levels = &topo_levels;\n> +\t\t\tg = g->base_graph;\n> +\t\t}\n> +\t}\n\nAll right, now I see why we need `topo_levels` member both in the\n`struct write_commit_graph_context` and in `struct commit_graph`.\nThe former is for functions that write the commit-graph, the latter for\nfill_commit_graph_info() functions that is deep in the callstack, but it\nneeds to know whether to load topological level to commit-slab, or maybe\nput it as generation number (and in the future -- discard it, if not\nneeded).\n\n\nSidenote: this fragment of code, that fills with a given value some\nmember of the `struct commit_graph` throughout the split commit-graph\nchain, will be repeated as similar code in patches later in series.\nHowever without resorting to preprocessor macros I have no idea how to\ngeneralize it to avoid code duplication (well, almost).\n\n> +\n>  \tif (pack_indexes) {\n>  \t\tctx->order_by_pack = 1;\n>  \t\tif ((res = fill_oids_from_packs(ctx, pack_indexes)))\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 430bc830bb..1152a9642e 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -72,6 +72,7 @@ struct commit_graph {\n>  \tconst unsigned char *chunk_bloom_indexes;\n>  \tconst unsigned char *chunk_bloom_data;\n>\n> +\tstruct topo_level_slab *topo_levels;\n>  \tstruct bloom_filter_settings *bloom_filter_settings;\n>  };\n\nAll right: `struct commit_graph` is public, `struct\nwrite_commit_graph_context` is not.\n\n>\n> diff --git a/commit.h b/commit.h\n> index bc0732a4fe..bb846e0025 100644\n> --- a/commit.h\n> +++ b/commit.h\n> @@ -15,6 +15,7 @@\n>  #define GENERATION_NUMBER_V1_INFINITY 0xFFFFFFFF\n>  #define GENERATION_NUMBER_MAX 0x3FFFFFFF\n\nThe name GENERATION_NUMBER_MAX for 0x3FFFFFFF should be instead\nGENERATION_NUMBER_V1_MAX, but that may be done in a later commit.\n\n>  #define GENERATION_NUMBER_ZERO 0\n> +#define GENERATION_NUMBER_V2_OFFSET_MAX 0xFFFFFFFF\n\nThis value is never used, so why it is defined in this commit.\n\n>\n>  struct commit_list {\n>  \tstruct commit *item;\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404245","messageId":"85wo1rk0iy.fsf@gmail.com","threadId":"53933","inReplyTo":"4074ace65be3094d35dd0aaedb89eb5a0ec98cee.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 07/11] commit-graph: implement corrected commit date","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-22T00:05:41Z","receivedAt":"2020-08-22T00:05:50Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> With most of preparations done, let's implement corrected commit date.\n>\n> The corrected commit date for a commit is defined as:\n>\n> * A commit with no parents (a root commit) has corrected commit date\n>   equal to its committer date.\n> * A commit with at least one parent has corrected commit date equal to\n>   the maximum of its commit date and one more than the largest corrected\n>   commit date among its parents.\n\nGood.\n\n>\n> To minimize the space required to store corrected commit date, Git\n> stores corrected commit date offsets into the commit-graph file. The\n> corrected commit date offset for a commit is defined as the difference\n> between its corrected commit date and actual commit date.\n\nPerhaps we should add more details about data type sizes in question.\n\nStoring corrected commit date requires sizeof(timestamp_t) bytes, which\nin most cases is 64 bits (uintmax_t).  However corrected commit date\noffsets can be safely stored^* using only 32 bits.  This halves the size\nof GDAT chunk, reducing per-commit storage from 2*H + 16 + 8 bytes to\n2*H + 16 + 4 bytes, which is reduction of around 6%, not including\nheader, fanout table (OIDF) and extra edges list (EDGE).\n\nWhich might mean that the extra complication is not worth it, and we\nshould store corrected commit date directly instead.\n\n*) unless for example one of commits is malformed but valid,\n   and has committerdate of 0 Unix time, 1 January 1970.\n\n>\n> While Git does not write out offsets at this stage, Git stores the\n> corrected commit dates in member generation of struct commit_graph_data.\n> It will begin writing commit date offsets with the introduction of\n> generation data chunk.\n\nOK, so the agenda for introducing geeration number v2 is as follows:\n- compute generation numbers v2, i.e. corrected commit date\n- store corrected commit date [offsets] in new GDAT chunk,\n  unless backward-compatibility concerns require us to not to\n- load [and compute] corrected commit date from commit-graph\n  storing it as 'generation' field of `struct commit_graph_data`,\n  unless backward-compatibility concerns require us to store\n  topological levels (generation number v1) in there instead\n\nBecause the reachability condition for corrected commit date and for\ntopological level is exactly the same, we don't need to do anything to\ntake advantage of generation number v2.\n\nThough we can use generation number v2 in more cases, where we turned\noff use of generation numbers because v1 gave worse performance than\ndate heuristics.\n\nDid I got this right?\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 58 +++++++++++++++++++++++++++-----------------------\n>  1 file changed, 31 insertions(+), 27 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index a2f15b2825..fd69534dd5 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -169,11 +169,6 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>  \telse if (generation_a > generation_b)\n>  \t\treturn 1;\n>\n> -\t/* use date as a heuristic when generations are equal */\n> -\tif (a->date < b->date)\n> -\t\treturn -1;\n> -\telse if (a->date > b->date)\n> -\t\treturn 1;\n\nAt first I was wondering why this tie-breaking is beig removed; wouldn't\nbe needed for backward-compatibility?  But then I remembered that this\ncomparison function is used _only_ for sorting commits when writing\nBloom filters, for `git commit-graph write --reachable --changed-paths ...`\n\nAssuming that when writing the commit graph we always compute geeration\nnumber v2 and 'generation' field stores corrected commit date, we don't\nneed to use date as a heuristic when generations are equal, and it would\nnot help in tie-breaking anyway.\n\nAll right.\n\n>  \treturn 0;\n>  }\n>\n> @@ -1342,10 +1337,14 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\t\t\tctx->commits.nr);\n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>  \t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n> +\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n\nAll right, so the pattern is to add 'corrected_commit_date' stuff after\n'topological_level' stuff.\n\n>\n>  \t\tdisplay_progress(ctx->progress, i + 1);\n>  \t\tif (level != GENERATION_NUMBER_V1_INFINITY &&\n> -\t\t    level != GENERATION_NUMBER_ZERO)\n> +\t\t    level != GENERATION_NUMBER_ZERO &&\n> +\t\t    corrected_commit_date != GENERATION_NUMBER_INFINITY &&\n> +\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO\n> +\t\t    )\n>  \t\t\tcontinue;\n>\n>  \t\tcommit_list_insert(ctx->commits.list[i], &list);\n> @@ -1354,17 +1353,26 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\tstruct commit_list *parent;\n>  \t\t\tint all_parents_computed = 1;\n>  \t\t\tuint32_t max_level = 0;\n> +\t\t\ttimestamp_t max_corrected_commit_date = 0;\n>\n>  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n>  \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n> +\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n>\n>  \t\t\t\tif (level == GENERATION_NUMBER_V1_INFINITY ||\n> -\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n> +\t\t\t\t    level == GENERATION_NUMBER_ZERO ||\n> +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_INFINITY ||\n> +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO\n> +\t\t\t\t    ) {\n>  \t\t\t\t\tall_parents_computed = 0;\n>  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n>  \t\t\t\t\tbreak;\n> -\t\t\t\t} else if (level > max_level) {\n> -\t\t\t\t\tmax_level = level;\n> +\t\t\t\t} else {\n> +\t\t\t\t\tif (level > max_level)\n> +\t\t\t\t\t\tmax_level = level;\n> +\n> +\t\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n> +\t\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n>  \t\t\t\t}\n>  \t\t\t}\n>\n> @@ -1374,6 +1382,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n>  \t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n>  \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n> +\n> +\t\t\t\tif (current->date > max_corrected_commit_date)\n> +\t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n> +\t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n>  \t\t\t}\n>  \t\t}\n>  \t}\n\nAll right.  Looks good to me.\n\n> @@ -2372,8 +2384,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n>  \tfor (i = 0; i < g->num_commits; i++) {\n>  \t\tstruct commit *graph_commit, *odb_commit;\n>  \t\tstruct commit_list *graph_parents, *odb_parents;\n> -\t\ttimestamp_t max_generation = 0;\n> -\t\ttimestamp_t generation;\n> +\t\ttimestamp_t max_corrected_commit_date = 0;\n> +\t\ttimestamp_t corrected_commit_date;\n\nThis is simple, and perhaps unnecessary, rename of variables.\nShouldn't we however verify *both* topological level, and\n(if exists) corrected commit date?\n\n>\n>  \t\tdisplay_progress(progress, i + 1);\n>  \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n> @@ -2412,9 +2424,9 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n>  \t\t\t\t\t     oid_to_hex(&graph_parents->item->object.oid),\n>  \t\t\t\t\t     oid_to_hex(&odb_parents->item->object.oid));\n>\n> -\t\t\tgeneration = commit_graph_generation(graph_parents->item);\n> -\t\t\tif (generation > max_generation)\n> -\t\t\t\tmax_generation = generation;\n> +\t\t\tcorrected_commit_date = commit_graph_generation(graph_parents->item);\n> +\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n> +\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n\nActually, commit_graph_generation(<commit>) can return either corrected\ncommit date, or topological level, the latter in backward-compatibility\ncase (if at least one commit-graph file is lacking GDAT chunk, because\n[some of] it was created by the \"Old\" Git).\n\n>\n>  \t\t\tgraph_parents = graph_parents->next;\n>  \t\t\todb_parents = odb_parents->next;\n> @@ -2436,20 +2448,12 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n>  \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n>  \t\t\tcontinue;\n>\n> -\t\t/*\n> -\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n> -\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n> -\t\t * extra logic in the following condition.\n> -\t\t */\n> -\t\tif (max_generation == GENERATION_NUMBER_MAX)\n> -\t\t\tmax_generation--;\n\nAll right, this was needed for checking the correctness of topological\nlevels (generation number v1) because we were checking not that it\nfullfills the reachability condition, but more strict one: namely that\ntopological level of commit is equal to maximum of topological levels of\nits parents plus one.\n\nThe comment about checking both generation number v1 and v2 still\napplies.\n\n> -\n> -\t\tgeneration = commit_graph_generation(graph_commit);\n> -\t\tif (generation != max_generation + 1)\n> -\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n> +\t\tcorrected_commit_date = commit_graph_generation(graph_commit);\n> +\t\tif (corrected_commit_date < max_corrected_commit_date + 1)\n> +\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n>  \t\t\t\t     oid_to_hex(&cur_oid),\n> -\t\t\t\t     generation,\n> -\t\t\t\t     max_generation + 1);\n> +\t\t\t\t     corrected_commit_date,\n> +\t\t\t\t     max_corrected_commit_date + 1);\n\nAll right, we check less strict condition for corrected commit date.\n\n>\n>  \t\tif (graph_commit->date != odb_commit->date)\n>  \t\t\tgraph_report(_(\"commit date for commit %s in commit-graph is %\"PRItime\" != %\"PRItime),\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404253","messageId":"85tuwuj08g.fsf@gmail.com","threadId":"53933","inReplyTo":"4e746628acdb49af5e8eb788864156f54724d4fa.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 08/11] commit-graph: implement generation data chunk","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-22T13:09:35Z","receivedAt":"2020-08-22T13:09:45Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"All right, it looks like this patch implements first part of step 2.)\nand step 3.) in the following plan of adding support for geeration\nnumber v2:\n\n1. compute generation numbers v2, i.e. corrected commit date\n2. store corrected commit date [offsets] in new GDAT chunk,\n   unless backward-compatibility concerns require us to not to\n3. load [and compute] corrected commit date from commit-graph\n   storing it as 'generation' field of `struct commit_graph_data`,\n   unless backward-compatibility concerns require us to store\n   topological levels (generation number v1) in there instead\n4. use generation number v2 in more places, where we had to turn\n   off using v1 for performance reasons\n\n\n\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> As discovered by Ævar, we cannot increment graph version to\n> distinguish between generation numbers v1 and v2 [1]. Thus, one of\n> pre-requistes before implementing generation number was to distinguish\n> between graph versions in a backwards compatible manner.\n\nFortunately, we have fixed this issue, and Git does no longer die when\nit encounters commit-graph format version that it does not understand.\n\n>\n> We are going to introduce a new chunk called Generation Data chunk (or\n> GDAT). GDAT stores generation number v2 (and any subsequent versions),\n> whereas CDAT will still store topological level.\n\nShould we say anything about storing 64 bit corrected commit date\n(geeration number v2) as 32 bit corrected commit date offset?\n\n>\n> Old Git does not understand GDAT chunk and would ignore it, reading\n> topological levels from CDAT. New Git can parse GDAT and take advantage\n> of newer generation numbers, falling back to topological levels when\n> GDAT chunk is missing (as it would happen with a commit graph written\n> by old Git).\n\nNote that the fact that we do not have special code for handling\nmixed-version layers in split commit-graph is not [that] dangerous, as\nwe don't read this new data yet.  Splitting it to patch 09/11 (next\npatch) makes this patch simpler.\n\n>\n> We introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\n> which forces commit-graph file to be written without generation data\n> chunk to emulate a commit-graph file written by old Git.\n\nAll right.\n\n>\n> [1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c                | 48 ++++++++++++++++++++++++---\n>  commit-graph.h                |  2 ++\n>  t/README                      |  3 ++\n>  t/helper/test-read-graph.c    |  2 ++\n>  t/t4216-log-bloom.sh          |  4 +--\n>  t/t5318-commit-graph.sh       | 27 +++++++--------\n>  t/t5324-split-commit-graph.sh | 12 +++----\n>  t/t6600-test-reach.sh         | 62 +++++++++++++++++++----------------\n>  8 files changed, 107 insertions(+), 53 deletions(-)\n\nIt might be a good idea to add documentation of this chunk (and only\nabout this chunk) to Documentation/technical/commit-graph-format.txt\n\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index fd69534dd5..b7a72b40db 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -38,11 +38,12 @@ void git_test_write_commit_graph_or_die(void)\n>  #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n>  #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n>  #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n> +#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n>  #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n>  #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n>  #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n>  #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n> -#define MAX_NUM_CHUNKS 7\n> +#define MAX_NUM_CHUNKS 8\n\nAll right, define new chunk and increase the maximum number of chunks\ncommit-graph file can have.\n\n>\n>  #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n>\n> @@ -389,6 +390,13 @@ struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size)\n>  \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n>  \t\t\tbreak;\n>\n> +\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n> +\t\t\tif (graph->chunk_generation_data)\n> +\t\t\t\tchunk_repeated = 1;\n> +\t\t\telse\n> +\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n> +\t\t\tbreak;\n> +\n\nAll right.  The size of GDAT chunk is defined by the number of commits,\nso nothink more is needed.\n\n>  \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n>  \t\t\tif (graph->chunk_extra_edges)\n>  \t\t\t\tchunk_repeated = 1;\n> @@ -755,7 +763,11 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \tdate_low = get_be32(commit_data + g->hash_len + 12);\n>  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n>\n> -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> +\tif (g->chunk_generation_data)\n> +\t\tgraph_data->generation = item->date +\n> +\t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n\nWARNING: The above does not properly handle clamped data.\n\nIf the offset is maximum value that can be stored in 32 bit field, that\nis GENERATION_NUMBER_V2_OFFSET_MAX for more that one commit, we wouldn't\nknow which commit has greater generation number v2 because of this clamping.\nThe 'generation' data needs to be set to GENERATION_NUMBER_V2_MAX (which\nin turn needs to be smaller than GENERATION_NUMBER_INFINITY).  This is\nnot done here!\n\nAll the above complication would not be an issue if we stored 64 bit\ncorrected commit date directly, instead of storing 32 bit corrected\ncommit date offsets. Storing offsets saves at most 4/(2*H + 16 + 4) = 7%\nof commit-graph file size (OIDL + CDAT + GDAT with offsets), when\n160-bit/20-byte SHA-1 hash is used.\n\n> +\telse\n> +\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n\nShould we convert GENERATION_NUMBER_V1_MAX into GENERATION_NUMBER_MAX\nhere, for easier handling of clamped values of generation numbers later\non in backward-compatibile way, without special-casing for v1 and v2?\n\n\nHere we load and perhaps compute generation number, using v2 if possible\n(from GDAT + commit date), with fallback to v1 (from CDAT).\n\nAlmost all right... but should this reading be a part of this patch, or\nsplit off into separate patch?\n\n>\n>  \tif (g->topo_levels)\n>  \t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n> @@ -951,7 +963,8 @@ struct write_commit_graph_context {\n>  \t\t report_progress:1,\n>  \t\t split:1,\n>  \t\t changed_paths:1,\n> -\t\t order_by_pack:1;\n> +\t\t order_by_pack:1,\n> +\t\t write_generation_data:1;\n\nAll right, here we store the iformation if we should write the GDAT\nchunk, taking into account among other things state of the\nGIT_TEST_COMMIT_GRAPH_NO_GDAT enviroment variable.\n\n>\n>  \tstruct topo_level_slab *topo_levels;\n>  \tconst struct split_commit_graph_opts *split_opts;\n> @@ -1106,8 +1119,25 @@ static int write_graph_chunk_data(struct hashfile *f,\n>  \treturn 0;\n>  }\n>\n> +static int write_graph_chunk_generation_data(struct hashfile *f,\n> +\t\t\t\t\t      struct write_commit_graph_context *ctx)\n> +{\n> +\tint i;\n> +\tfor (i = 0; i < ctx->commits.nr; i++) {\n\nSide note: it is a bit funny that some of write_graph_chunk_*()\nfunctions use `for` loop, and some `while` loop to process commits.\n\n> +\t\tstruct commit *c = ctx->commits.list[i];\n> +\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n> +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n> +\n> +\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX)\n> +\t\t\toffset = GENERATION_NUMBER_V2_OFFSET_MAX;\n\nThis GENERATION_NUMBER_V2_OFFSET_MAX symbolic constant (or equivalet)\nshould be defined in this patch and not in previous one, in my opinion.\n\n> +\t\thashwrite_be32(f, offset);\n> +\t}\n> +\n> +\treturn 0;\n> +}\n> +\n>  static int write_graph_chunk_extra_edges(struct hashfile *f,\n> -\t\t\t\t\t struct write_commit_graph_context *ctx)\n> +\t\t\t\t\t  struct write_commit_graph_context *ctx)\n\nThis change in whitespace is, I think, incorrect.\n\n>  {\n>  \tstruct commit **list = ctx->commits.list;\n>  \tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n> @@ -1726,6 +1756,15 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \tchunks[2].id = GRAPH_CHUNKID_DATA;\n>  \tchunks[2].size = (hashsz + 16) * ctx->commits.nr;\n>  \tchunks[2].write_fn = write_graph_chunk_data;\n> +\n> +\tif (git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0))\n> +\t\tctx->write_generation_data = 0;\n\nAll right, here we handle newly introduced GIT_TEST_COMMIT_GRAPH_NO_GDAT\nenvironment variable.\n\n> +\tif (ctx->write_generation_data) {\n> +\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA;\n> +\t\tchunks[num_chunks].size = sizeof(uint32_t) * ctx->commits.nr;\n> +\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data;\n> +\t\tnum_chunks++;\n> +\t}\n\nHmmm... so the GDAT chunk does not have a header, and is not versioned.\nDoes this mean that to move to generation number v3, or add some\nadditional reachability labeling we would need to either rename the\nchunk or add new one?\n\n>  \tif (ctx->num_extra_edges) {\n>  \t\tchunks[num_chunks].id = GRAPH_CHUNKID_EXTRAEDGES;\n>  \t\tchunks[num_chunks].size = 4 * ctx->num_extra_edges;\n> @@ -2130,6 +2169,7 @@ int write_commit_graph(struct object_directory *odb,\n>  \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n>  \tctx->split_opts = split_opts;\n>  \tctx->total_bloom_filter_data_size = 0;\n> +\tctx->write_generation_data = 1;\n\nBut by default we do write the GDAT chunk.\n\n>\n>  \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n>  \t\tctx->changed_paths = 1;\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 1152a9642e..f78c892fc0 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -6,6 +6,7 @@\n>  #include \"oidset.h\"\n>\n>  #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n> +#define GIT_TEST_COMMIT_GRAPH_NO_GDAT \"GIT_TEST_COMMIT_GRAPH_NO_GDAT\"\n>  #define GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE \"GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE\"\n>  #define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n>\n\nAll right.\n\n(Though I wonder about the ordering -- but better not to start \"bike\nshed\" discussion).\n\n> @@ -67,6 +68,7 @@ struct commit_graph {\n>  \tconst uint32_t *chunk_oid_fanout;\n>  \tconst unsigned char *chunk_oid_lookup;\n>  \tconst unsigned char *chunk_commit_data;\n> +\tconst unsigned char *chunk_generation_data;\n>  \tconst unsigned char *chunk_extra_edges;\n>  \tconst unsigned char *chunk_base_graphs;\n>  \tconst unsigned char *chunk_bloom_indexes;\n\nAll right, we need to store position of new GDAT chunk.\n\n> diff --git a/t/README b/t/README\n> index 70ec61cf88..6647ef132e 100644\n> --- a/t/README\n> +++ b/t/README\n> @@ -379,6 +379,9 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n>  be written after every 'git commit' command, and overrides the\n>  'core.commitGraph' setting to true.\n>\n> +GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n> +commit-graph to be written without generation data chunk.\n> +\n\nAll right.\n\nThis description could have been more detailed, but I think it is good\nenough.\n\n>  GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n>  commit-graph write to compute and write changed path Bloom filters for\n>  every 'git commit-graph write', as if the `--changed-paths` option was\n> diff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\n> index 6d0c962438..1c2a5366c7 100644\n> --- a/t/helper/test-read-graph.c\n> +++ b/t/helper/test-read-graph.c\n> @@ -32,6 +32,8 @@ int cmd__read_graph(int argc, const char **argv)\n>  \t\tprintf(\" oid_lookup\");\n>  \tif (graph->chunk_commit_data)\n>  \t\tprintf(\" commit_metadata\");\n> +\tif (graph->chunk_generation_data)\n> +\t\tprintf(\" generation_data\");\n>  \tif (graph->chunk_extra_edges)\n>  \t\tprintf(\" extra_edges\");\n>  \tif (graph->chunk_bloom_indexes)\n\nAll right, we examine if GDAT chunk is present.\n\nMany commit-graph tests would probably need to be updated; at least\nthose that make use if `git test-tool read-graph`.\n\n> diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n> index c21cc160f3..55c94e9ebd 100755\n> --- a/t/t4216-log-bloom.sh\n> +++ b/t/t4216-log-bloom.sh\n> @@ -33,11 +33,11 @@ test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n>  \tgit commit-graph write --reachable --changed-paths\n>  '\n>  graph_read_expect () {\n> -\tNUM_CHUNKS=5\n> +\tNUM_CHUNKS=6\n>  \tcat >expect <<- EOF\n>  \theader: 43475048 1 1 $NUM_CHUNKS 0\n>  \tnum_commits: $1\n> -\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n> +\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n>  \tEOF\n>  \ttest-tool read-graph >actual &&\n>  \ttest_cmp expect actual\n\nAll right.\n\n> diff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\n> index 044cf8a3de..b41b2160c6 100755\n> --- a/t/t5318-commit-graph.sh\n> +++ b/t/t5318-commit-graph.sh\n> @@ -71,7 +71,7 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n>  graph_read_expect() {\n>  \tOPTIONAL=\"\"\n>  \tNUM_CHUNKS=3\n> -\tif test ! -z $2\n> +\tif test ! -z \"$2\"\n>  \tthen\n>  \t\tOPTIONAL=\" $2\"\n>  \t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n\nAll right, a fix for issue that is relevant only after this change, now\nthat there can be more than one extra chunk.\n\nSide note: how I wish that helper function in tests were documented...\n\n> @@ -98,14 +98,14 @@ test_expect_success 'exit with correct error on bad input to --stdin-commits' '\n>  \t# valid commit and tree OID\n>  \tgit rev-parse HEAD HEAD^{tree} >in &&\n>  \tgit commit-graph write --stdin-commits <in &&\n> -\tgraph_read_expect 3\n> +\tgraph_read_expect 3 generation_data\n\nAll right, we need to treat generation_data as extra, because it can be\nnot there (it's existence is conditional).\n\n>  '\n>\n>  test_expect_success 'write graph' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit commit-graph write &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"3\"\n> +\tgraph_read_expect \"3\" generation_data\n\nSide note: I wonder why here we have\n\n  \tgraph_read_expect \"3\" generation_data\n\nbut one test earlier we have\n\n  \tgraph_read_expect 3 generation_data\n\nwithout quotes.\n\n>  '\n>\n>  test_expect_success POSIXPERM 'write graph has correct permissions' '\n> @@ -214,7 +214,7 @@ test_expect_success 'write graph with merges' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit commit-graph write &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"10\" \"extra_edges\"\n> +\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n\nAll right.  It is why we needed to fix graph_read_expect().\n\n>  '\n>\n>  graph_git_behavior 'merge 1 vs 2' full merge/1 merge/2\n> @@ -249,7 +249,7 @@ test_expect_success 'write graph with new commit' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit commit-graph write &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"11\" \"extra_edges\"\n> +\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n>  '\n>\n>  graph_git_behavior 'full graph, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -259,7 +259,7 @@ test_expect_success 'write graph with nothing new' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit commit-graph write &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"11\" \"extra_edges\"\n> +\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n>  '\n>\n>  graph_git_behavior 'cleared graph, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -269,7 +269,7 @@ test_expect_success 'build graph from latest pack with closure' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tcat new-idx | git commit-graph write --stdin-packs &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"9\" \"extra_edges\"\n> +\tgraph_read_expect \"9\" \"generation_data extra_edges\"\n>  '\n>\n>  graph_git_behavior 'graph from pack, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -282,7 +282,7 @@ test_expect_success 'build graph from commits with closure' '\n>  \tgit rev-parse merge/1 >>commits-in &&\n>  \tcat commits-in | git commit-graph write --stdin-commits &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"6\"\n> +\tgraph_read_expect \"6\" \"generation_data\"\n>  '\n>\n>  graph_git_behavior 'graph from commits, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -292,7 +292,7 @@ test_expect_success 'build graph from commits with append' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit rev-parse merge/3 | git commit-graph write --stdin-commits --append &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"10\" \"extra_edges\"\n> +\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n>  '\n>\n>  graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -302,7 +302,7 @@ test_expect_success 'build graph using --reachable' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit commit-graph write --reachable &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"11\" \"extra_edges\"\n> +\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n>  '\n>\n>  graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -323,7 +323,7 @@ test_expect_success 'write graph in bare repo' '\n>  \tcd \"$TRASH_DIRECTORY/bare\" &&\n>  \tgit commit-graph write &&\n>  \ttest_path_is_file $baredir/info/commit-graph &&\n> -\tgraph_read_expect \"11\" \"extra_edges\"\n> +\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n>  '\n\nAll right, those were the straightforward changes.\n\n>\n>  graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n> @@ -420,8 +420,9 @@ test_expect_success 'replace-objects invalidates commit-graph' '\n>\n>  test_expect_success 'git commit-graph verify' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n> -\tgit rev-parse commits/8 | git commit-graph write --stdin-commits &&\n> -\tgit commit-graph verify >output\n> +\tgit rev-parse commits/8 | GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --stdin-commits &&\n> +\tgit commit-graph verify >output &&\n> +\tgraph_read_expect 9 extra_edges\n>  '\n\nWhat this change is about?  Is it about the fact that we have not added\nsupport for checking correctness of GDAT chunk to `git commit-graph\nverify`?  But in previous commit we did modify verify_commit_graph()...\n\n>\n>  NUM_COMMITS=9\n> diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> index ea28d522b8..531016f405 100755\n> --- a/t/t5324-split-commit-graph.sh\n> +++ b/t/t5324-split-commit-graph.sh\n> @@ -13,11 +13,11 @@ test_expect_success 'setup repo' '\n>  \tinfodir=\".git/objects/info\" &&\n>  \tgraphdir=\"$infodir/commit-graphs\" &&\n>  \ttest_oid_cache <<-EOM\n> -\tshallow sha1:1760\n> -\tshallow sha256:2064\n> +\tshallow sha1:2132\n> +\tshallow sha256:2436\n>\n> -\tbase sha1:1376\n> -\tbase sha256:1496\n> +\tbase sha1:1408\n> +\tbase sha256:1528\n>  \tEOM\n>  '\n\nI guess this change is because with GDAT chunk present the positions of\nrelevant bits that we want to corrupt change (I guess because we have\nextra chunk in chunk lookup section).\n\nSomeone else would have to verify if this change is in fact correct, if\nit was not done already.\n\n>\n> @@ -28,9 +28,9 @@ graph_read_expect() {\n>  \t\tNUM_BASE=$2\n>  \tfi\n>  \tcat >expect <<- EOF\n> -\theader: 43475048 1 1 3 $NUM_BASE\n> +\theader: 43475048 1 1 4 $NUM_BASE\n>  \tnum_commits: $1\n> -\tchunks: oid_fanout oid_lookup commit_metadata\n> +\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n\nAll right, we now have 4 chunks not 3 (old ones + generation_data).\n\n>  \tEOF\n>  \ttest-tool read-graph >output &&\n>  \ttest_cmp expect output\n> diff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\n> index 475564bee7..d14b129f06 100755\n> --- a/t/t6600-test-reach.sh\n> +++ b/t/t6600-test-reach.sh\n> @@ -55,10 +55,13 @@ test_expect_success 'setup' '\n>  \tgit show-ref -s commit-5-5 | git commit-graph write --stdin-commits &&\n>  \tmv .git/objects/info/commit-graph commit-graph-half &&\n>  \tchmod u+w commit-graph-half &&\n> +\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable &&\n> +\tmv .git/objects/info/commit-graph commit-graph-no-gdat &&\n> +\tchmod u+w commit-graph-no-gdat &&\n>  \tgit config core.commitGraph true\n>  '\n\nAll right, we add setup for testing no GDAT case (as if the commit-graph\nfile was written by \"Old\" Git).\n\n>\n> -run_three_modes () {\n> +run_all_modes () {\n\nAll right, we compare more modes, among others the no-GDAT case.\n\nNice futureproofing!\n\n>  \ttest_when_finished rm -rf .git/objects/info/commit-graph &&\n>  \t\"$@\" <input >actual &&\n>  \ttest_cmp expect actual &&\n> @@ -67,11 +70,14 @@ run_three_modes () {\n>  \ttest_cmp expect actual &&\n>  \tcp commit-graph-half .git/objects/info/commit-graph &&\n>  \t\"$@\" <input >actual &&\n> +\ttest_cmp expect actual &&\n> +\tcp commit-graph-no-gdat .git/objects/info/commit-graph &&\n> +\t\"$@\" <input >actual &&\n>  \ttest_cmp expect actual\n>  }\n\nOK, now we test yet another variant of commit-graph file: one without\nthe GDAT chunk (testing for backward compatibility with \"Old\" Git).\n\n>\n> -test_three_modes () {\n> -\trun_three_modes test-tool reach \"$@\"\n> +test_all_modes () {\n> +\trun_all_modes test-tool reach \"$@\"\n>  }\n\nAll right.\n\n>\n>  test_expect_success 'ref_newer:miss' '\n> @@ -80,7 +86,7 @@ test_expect_success 'ref_newer:miss' '\n>  \tB:commit-4-9\n>  \tEOF\n>  \techo \"ref_newer(A,B):0\" >expect &&\n> -\ttest_three_modes ref_newer\n> +\ttest_all_modes ref_newer\n>  '\n>\n>  test_expect_success 'ref_newer:hit' '\n> @@ -89,7 +95,7 @@ test_expect_success 'ref_newer:hit' '\n>  \tB:commit-2-3\n>  \tEOF\n>  \techo \"ref_newer(A,B):1\" >expect &&\n> -\ttest_three_modes ref_newer\n> +\ttest_all_modes ref_newer\n>  '\n>\n>  test_expect_success 'in_merge_bases:hit' '\n> @@ -98,7 +104,7 @@ test_expect_success 'in_merge_bases:hit' '\n>  \tB:commit-8-8\n>  \tEOF\n>  \techo \"in_merge_bases(A,B):1\" >expect &&\n> -\ttest_three_modes in_merge_bases\n> +\ttest_all_modes in_merge_bases\n>  '\n>\n>  test_expect_success 'in_merge_bases:miss' '\n> @@ -107,7 +113,7 @@ test_expect_success 'in_merge_bases:miss' '\n>  \tB:commit-5-9\n>  \tEOF\n>  \techo \"in_merge_bases(A,B):0\" >expect &&\n> -\ttest_three_modes in_merge_bases\n> +\ttest_all_modes in_merge_bases\n>  '\n>\n>  test_expect_success 'is_descendant_of:hit' '\n> @@ -118,7 +124,7 @@ test_expect_success 'is_descendant_of:hit' '\n>  \tX:commit-1-1\n>  \tEOF\n>  \techo \"is_descendant_of(A,X):1\" >expect &&\n> -\ttest_three_modes is_descendant_of\n> +\ttest_all_modes is_descendant_of\n>  '\n>\n>  test_expect_success 'is_descendant_of:miss' '\n> @@ -129,7 +135,7 @@ test_expect_success 'is_descendant_of:miss' '\n>  \tX:commit-7-6\n>  \tEOF\n>  \techo \"is_descendant_of(A,X):0\" >expect &&\n> -\ttest_three_modes is_descendant_of\n> +\ttest_all_modes is_descendant_of\n>  '\n>\n>  test_expect_success 'get_merge_bases_many' '\n> @@ -144,7 +150,7 @@ test_expect_success 'get_merge_bases_many' '\n>  \t\tgit rev-parse commit-5-6 \\\n>  \t\t\t      commit-4-7 | sort\n>  \t} >expect &&\n> -\ttest_three_modes get_merge_bases_many\n> +\ttest_all_modes get_merge_bases_many\n>  '\n>\n>  test_expect_success 'reduce_heads' '\n> @@ -166,7 +172,7 @@ test_expect_success 'reduce_heads' '\n>  \t\t\t      commit-2-8 \\\n>  \t\t\t      commit-1-10 | sort\n>  \t} >expect &&\n> -\ttest_three_modes reduce_heads\n> +\ttest_all_modes reduce_heads\n>  '\n>\n>  test_expect_success 'can_all_from_reach:hit' '\n> @@ -189,7 +195,7 @@ test_expect_success 'can_all_from_reach:hit' '\n>  \tY:commit-8-1\n>  \tEOF\n>  \techo \"can_all_from_reach(X,Y):1\" >expect &&\n> -\ttest_three_modes can_all_from_reach\n> +\ttest_all_modes can_all_from_reach\n>  '\n>\n>  test_expect_success 'can_all_from_reach:miss' '\n> @@ -211,7 +217,7 @@ test_expect_success 'can_all_from_reach:miss' '\n>  \tY:commit-8-5\n>  \tEOF\n>  \techo \"can_all_from_reach(X,Y):0\" >expect &&\n> -\ttest_three_modes can_all_from_reach\n> +\ttest_all_modes can_all_from_reach\n>  '\n>\n>  test_expect_success 'can_all_from_reach_with_flag: tags case' '\n> @@ -234,7 +240,7 @@ test_expect_success 'can_all_from_reach_with_flag: tags case' '\n>  \tY:commit-8-1\n>  \tEOF\n>  \techo \"can_all_from_reach_with_flag(X,_,_,0,0):1\" >expect &&\n> -\ttest_three_modes can_all_from_reach_with_flag\n> +\ttest_all_modes can_all_from_reach_with_flag\n>  '\n>\n>  test_expect_success 'commit_contains:hit' '\n> @@ -250,8 +256,8 @@ test_expect_success 'commit_contains:hit' '\n>  \tX:commit-9-3\n>  \tEOF\n>  \techo \"commit_contains(_,A,X,_):1\" >expect &&\n> -\ttest_three_modes commit_contains &&\n> -\ttest_three_modes commit_contains --tag\n> +\ttest_all_modes commit_contains &&\n> +\ttest_all_modes commit_contains --tag\n>  '\n>\n>  test_expect_success 'commit_contains:miss' '\n> @@ -267,8 +273,8 @@ test_expect_success 'commit_contains:miss' '\n>  \tX:commit-9-3\n>  \tEOF\n>  \techo \"commit_contains(_,A,X,_):0\" >expect &&\n> -\ttest_three_modes commit_contains &&\n> -\ttest_three_modes commit_contains --tag\n> +\ttest_all_modes commit_contains &&\n> +\ttest_all_modes commit_contains --tag\n>  '\n>\n>  test_expect_success 'rev-list: basic topo-order' '\n> @@ -280,7 +286,7 @@ test_expect_success 'rev-list: basic topo-order' '\n>  \t\tcommit-6-2 commit-5-2 commit-4-2 commit-3-2 commit-2-2 commit-1-2 \\\n>  \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --topo-order commit-6-6\n> +\trun_all_modes git rev-list --topo-order commit-6-6\n>  '\n>\n>  test_expect_success 'rev-list: first-parent topo-order' '\n> @@ -292,7 +298,7 @@ test_expect_success 'rev-list: first-parent topo-order' '\n>  \t\tcommit-6-2 \\\n>  \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --first-parent --topo-order commit-6-6\n> +\trun_all_modes git rev-list --first-parent --topo-order commit-6-6\n>  '\n>\n>  test_expect_success 'rev-list: range topo-order' '\n> @@ -304,7 +310,7 @@ test_expect_success 'rev-list: range topo-order' '\n>  \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n>  \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --topo-order commit-3-3..commit-6-6\n> +\trun_all_modes git rev-list --topo-order commit-3-3..commit-6-6\n>  '\n>\n>  test_expect_success 'rev-list: range topo-order' '\n> @@ -316,7 +322,7 @@ test_expect_success 'rev-list: range topo-order' '\n>  \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n>  \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --topo-order commit-3-8..commit-6-6\n> +\trun_all_modes git rev-list --topo-order commit-3-8..commit-6-6\n>  '\n>\n>  test_expect_success 'rev-list: first-parent range topo-order' '\n> @@ -328,7 +334,7 @@ test_expect_success 'rev-list: first-parent range topo-order' '\n>  \t\tcommit-6-2 \\\n>  \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n> +\trun_all_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n>  '\n>\n>  test_expect_success 'rev-list: ancestry-path topo-order' '\n> @@ -338,7 +344,7 @@ test_expect_success 'rev-list: ancestry-path topo-order' '\n>  \t\tcommit-6-4 commit-5-4 commit-4-4 commit-3-4 \\\n>  \t\tcommit-6-3 commit-5-3 commit-4-3 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n> +\trun_all_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n>  '\n>\n>  test_expect_success 'rev-list: symmetric difference topo-order' '\n> @@ -352,7 +358,7 @@ test_expect_success 'rev-list: symmetric difference topo-order' '\n>  \t\tcommit-3-8 commit-2-8 commit-1-8 \\\n>  \t\tcommit-3-7 commit-2-7 commit-1-7 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --topo-order commit-3-8...commit-6-6\n> +\trun_all_modes git rev-list --topo-order commit-3-8...commit-6-6\n>  '\n>\n>  test_expect_success 'get_reachable_subset:all' '\n> @@ -372,7 +378,7 @@ test_expect_success 'get_reachable_subset:all' '\n>  \t\t\t      commit-1-7 \\\n>  \t\t\t      commit-5-6 | sort\n>  \t) >expect &&\n> -\ttest_three_modes get_reachable_subset\n> +\ttest_all_modes get_reachable_subset\n>  '\n>\n>  test_expect_success 'get_reachable_subset:some' '\n> @@ -390,7 +396,7 @@ test_expect_success 'get_reachable_subset:some' '\n>  \t\tgit rev-parse commit-3-3 \\\n>  \t\t\t      commit-1-7 | sort\n>  \t) >expect &&\n> -\ttest_three_modes get_reachable_subset\n> +\ttest_all_modes get_reachable_subset\n>  '\n>\n>  test_expect_success 'get_reachable_subset:none' '\n> @@ -404,7 +410,7 @@ test_expect_success 'get_reachable_subset:none' '\n>  \tY:commit-2-8\n>  \tEOF\n>  \techo \"get_reachable_subset(X,Y)\" >expect &&\n> -\ttest_three_modes get_reachable_subset\n> +\ttest_all_modes get_reachable_subset\n>  '\n\nAll right, s/test_three_modes/test_all_modes/... which admitedly could\nhave been separate pure refactoring patch, but it is not necessary.\n\n>\n>  test_done\n\nBest,\n--\nJakub Narębski\n"},{"id":"404256","messageId":"85pn7ihabl.fsf@gmail.com","threadId":"53933","inReplyTo":"5a147a9704f0f8d8644c92ea38583e966378b931.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 09/11] commit-graph: use generation v2 only if entire chain does","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-22T17:14:38Z","receivedAt":"2020-08-22T17:15:47Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hi Abhishek,\n\n\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> Since there are released versions of Git that understand generation\n> numbers in the commit-graph's CDAT chunk but do not understand the GDAT\n> chunk, the following scenario is possible:\n>\n> 1. \"New\" Git writes a commit-graph with the GDAT chunk.\n> 2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n>\n> Because of the current use of inspecting the current layer for a\n> chunk_generation_data pointer, the commits in the lower layer will be\n> interpreted as having very large generation values (commit date plus\n> offset) compared to the generation numbers in the top layer (topological\n> level). This violates the expectation that the generation of a parent is\n> strictly smaller than the generation of a child.\n\nAll right, this explains changes to the *reading* side.  If there is\nsplit commit-graph layer without GDAT chunk (or to be more exact,\ncurrent code checks if there is any layer without GDAT chunk), then we\nfall back to using topological levels, as if all layers / all graphs\nwere without GDAT.  This allows to avoid the issue explained above,\nwhere for some commits 'generation' holds corrected commit date, and for\nsome it holds topological levels, breaking the reachability condition\nguarantee.\n\n\nHowever the commit message do not say anything about the *writing* side.\n\nWe have decided to not write the GDAT chunk when writing the new layer\nin split commit-graph, and top layer doesn't itself have GDAT chunks.\nThat makes for easier reasoning and safer handling: in mixed-version\nenvironment the only possible arrangement is for the lower layers\n(possibly zero) have GDAT chunk, and higher layers (possibly zero) not\nhave GDAT chunks.\n\nRewriting layers follows similar approach: if the topmost layer below\nset of layers being rewritten (in the split commit-graph chain) exists,\nand it does not contain GDAT chunk, then the result of rewrite should\nnot have GDAT chunks either.\n\n\nTo be more detailed, without '--split=replace' we would want the following\nlayer merging behavior:\n\n   [layer with GDAT][with GDAT][without GDAT][without GDAT][without GDAT]\n           1              2           3             4            5\n\nIn the split commit-graph chain above, merging two topmost layers\n(layers 4 and 5) should create a layer without GDAT; merging three\ntopmost layers (and any other layers, e.g. two middle ones, i.e. 3 and\n4) should create a new layer with GDAT.\n\n   [layer with GDAT][with GDAT][without GDAT][-------without GDAT-------]\n           1              2           3               merged\n\n   [layer with GDAT][with GDAT][-------------with GDAT------------------]\n           1              2                    merged\n\nI hope those ASCII-art pictures help understanding it\n\n>\n> It is difficult to expose this issue in a test. Since we _start_ with\n> artificially low generation numbers, any commit walk that prioritizes\n> generation numbers will walk all of the commits with high generation\n> number before walking the commits with low generation number. In all the\n> cases I tried, the commit-graph layers themselves \"protect\" any\n> incorrect behavior since none of the commits in the lower layer can\n> reach the commits in the upper layer.\n>\n> This issue would manifest itself as a performance problem in this case,\n> especially with something like \"git log --graph\" since the low\n> generation numbers would cause the in-degree queue to walk all of the\n> commits in the lower layer before allowing the topo-order queue to write\n> anything to output (depending on the size of the upper layer).\n\nWouldn't breaking the reachability condition promise make some Git\ncommands to return *incorrect* results if they short-circuit, stop\nwalking if generation number shows that A cannot reach B?\n\nI am talking here about commands that return boolean, or select subset\nfrom given set of revisions:\n- git merge-base --is-ancestor <B> <A>\n- git branch branch-A <A> && git branch --contains <B>\n- git branch branch-B <B> && git branch --merged <A>\n\nGit assumes that generation numbers fulfill the following condition:\n\n  if A can reach B, then gen(A) > gen(B)\n\nNotably this includes commits not in commit-graph, and clamped values.\n\nHowever, in the following case\n\n* if commit A is from higher layer without GDAT\n  and uses topological levels for 'generation', e.g. 115 (in a small repo)\n* and commit B is from lower layer with GDAT\n  and uses corrected commit date as 'generation', for example 1598112896,\n\nit may happen that A (later commit) can reach B (earlier commit), but\ngen(B) > gen(A).  The reachability condition promise for generation\nnumbers is broken.\n\n>\n> Signed-off-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n\nI have reordered files in the patch itself to make it easier to review\nthe proposed changes.\n\n>  commit-graph.h                |  1 +\n>  commit-graph.c                | 32 +++++++++++++++-\n>  t/t5324-split-commit-graph.sh | 70 +++++++++++++++++++++++++++++++++++\n>  3 files changed, 102 insertions(+), 1 deletion(-)\n>\n> diff --git a/commit-graph.h b/commit-graph.h\n> index f78c892fc0..3cf89d895d 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -63,6 +63,7 @@ struct commit_graph {\n>  \tstruct object_directory *odb;\n>\n>  \tuint32_t num_commits_in_base;\n> +\tuint32_t read_generation_data;\n>  \tstruct commit_graph *base_graph;\n>\n\nFirst, why `read_generation_data` is of uint32_t type, when it stores\n(as far as I understand it), a \"boolean\" value of either 0 or 1?\n\nSecond, couldn't we simply set chunk_generation_data to NULL?  Or would\nthat interfere with the case of rewriting, where we want to use existing\nGDAT data when writing new commit-graph with GDAT chunk?\n\n>  \tconst uint32_t *chunk_oid_fanout;\n> diff --git a/commit-graph.c b/commit-graph.c\n> index b7a72b40db..c1292f8e08 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -597,6 +597,27 @@ static struct commit_graph *load_commit_graph_chain(struct repository *r,\n>  \treturn graph_chain;\n>  }\n>\n> +static void validate_mixed_generation_chain(struct repository *r)\n> +{\n> +\tstruct commit_graph *g = r->objects->commit_graph;\n> +\tint read_generation_data = 1;\n> +\n> +\twhile (g) {\n> +\t\tif (!g->chunk_generation_data) {\n> +\t\t\tread_generation_data = 0;\n> +\t\t\tbreak;\n> +\t\t}\n> +\t\tg = g->base_graph;\n> +\t}\n\nThis loop checks whole split commit-graph chain for existence of layers\nwithout GDAT chunk.\n\nOn one hand it is more than needed _if_ we assume that the fact that\nonly topmost layers can be without GDAT holds true. On the other hand it\nis safer (an example of defensive coding), and as the length of chain is\nlimited it should be not much of a performance penalty.\n\n> +\n> +\tg = r->objects->commit_graph;\n> +\n> +\twhile (g) {\n> +\t\tg->read_generation_data = read_generation_data;\n> +\t\tg = g->base_graph;\n> +\t}\n\nAll right... though one of earlier commits introduced similar loop, but\nit set chunk_generation_data to NULL instead.  Or did I remember it wrong?\n\n> +}\n> +\n>  struct commit_graph *read_commit_graph_one(struct repository *r,\n>  \t\t\t\t\t   struct object_directory *odb)\n>  {\n> @@ -605,6 +626,8 @@ struct commit_graph *read_commit_graph_one(struct repository *r,\n>  \tif (!g)\n>  \t\tg = load_commit_graph_chain(r, odb);\n>\n> +\tvalidate_mixed_generation_chain(r);\n> +\n\nAll right, when reading the commit-graph, check if we are in forced\nbackward-compatibile mode, and we need to use topological levels for\ngeneration numbers.\n\n>  \treturn g;\n>  }\n>\n> @@ -763,7 +786,7 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \tdate_low = get_be32(commit_data + g->hash_len + 12);\n>  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n>\n> -\tif (g->chunk_generation_data)\n> +\tif (g->chunk_generation_data && g->read_generation_data)\n\nAll right, when deciding whether to use corrected commit date\n(generation number v2 from GDAT chunk), or fall back to using\ntopological levels (generation number v1 from CDAT chunk), we need to\ntake into accout other layers, to not mix v1 and v2.\n\nWe have earlier checked whether we can use generation number v2, now we\nuse the result of this check, propagated down the commit-graph chain.\n\n>  \t\tgraph_data->generation = item->date +\n>  \t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n>  \telse\n> @@ -885,6 +908,7 @@ void load_commit_graph_info(struct repository *r, struct commit *item)\n>  \tuint32_t pos;\n>  \tif (!prepare_commit_graph(r))\n>  \t\treturn;\n> +\n>  \tif (find_commit_in_graph(item, r->objects->commit_graph, &pos))\n>  \t\tfill_commit_graph_info(item, r->objects->commit_graph, pos);\n>  }\n\nThis is unrelated whitespace fix, a \"while at it\" in neighbourhood of\nchanges.  All right then.\n\n> @@ -2192,6 +2216,9 @@ int write_commit_graph(struct object_directory *odb,\n>\n>  \t\tg = ctx->r->objects->commit_graph;\n>\n> +\t\tif (g && !g->chunk_generation_data)\n> +\t\t\tctx->write_generation_data = 0;\n> +\n>  \t\twhile (g) {\n>  \t\t\tctx->num_commit_graphs_before++;\n>  \t\t\tg = g->base_graph;\n> @@ -2210,6 +2237,9 @@ int write_commit_graph(struct object_directory *odb,\n>\n>  \t\tif (ctx->split_opts)\n>  \t\t\treplace = ctx->split_opts->flags & COMMIT_GRAPH_SPLIT_REPLACE;\n> +\n> +\t\tif (replace)\n> +\t\t\tctx->write_generation_data = 1;\n>  \t}\n\n\nThe previous commit introduced `write_generation_data` member in\n`struct write_commit_graph_context`, then used to handle support for\nGIT_TEST_COMMIT_GRAPH_NO_GDAT environment variable.\n\nThose two hunks of changes above are both inside\n\n   if (ctx->split) {\n      ...\n   }\n\nHere we examine the topmost layer of split commit-graph chain, and if it\ndoes not contain GDAT chunk, then we do not store the GDAT chunk, unless\n`git commit-graph write` is ru with `--split=replace` option.\n\nHowever this is overly strict condition.  If we merge layer without GDAT\nwith layer with GDAT below, then we surely can write GDAT; the condition\nfor GDAT-less layers would be still fulfilled (met).  However we can\nconsider it 'good enough' for now, and relax this condition in later\ncommits.\n\n\nNote that it is the first time in this patch were we make use of\nassumption that if there are layers without GDAT then topmost layer is\nwithout GDAT.\n\n>\n>  \tctx->approx_nr_objects = approximate_object_count();\n> diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> index 531016f405..ac5e7783fb 100755\n> --- a/t/t5324-split-commit-graph.sh\n> +++ b/t/t5324-split-commit-graph.sh\n> @@ -424,4 +424,74 @@ done <<\\EOF\n>  0600 -r--------\n>  EOF\n>\n\nIt would be nice to have an ASCII-art graph of commits, but earlier\ntests do not have it either...\n\n> +test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n> +\tmkdir mixed &&\n> +\tgraphdir=\".git/objects/info/commit-graphs\" &&\n> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> +\tgit init &&\n> +\tgit config core.commitGraph true &&\n> +\tgit config gc.writeCommitGraph false &&\n> +\tfor i in $(test_seq 3)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit reset --hard commits/1 &&\n> +\tfor i in $(test_seq 4 5)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit reset --hard commits/2 &&\n> +\tfor i in $(test_seq 6 10)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit commit-graph write --reachable --split &&\n> +\tgit reset --hard commits/2 &&\n> +\tgit merge commits/4 &&\n> +\tgit branch merge/1 &&\n> +\tgit reset --hard commits/4 &&\n> +\tgit merge commits/6 &&\n> +\tgit branch merge/2 &&\n> +\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n> +\ttest-tool read-graph >output &&\n> +\tcat >expect <<-EOF &&\n> +\theader: 43475048 1 1 4 1\n> +\tnum_commits: 2\n> +\tchunks: oid_fanout oid_lookup commit_metadata\n> +\tEOF\n> +\ttest_cmp expect output &&\n> +\tgit commit-graph verify\n> +'\n\nLooks all right to me.\n\n> +\n> +test_expect_success 'does not write generation data chunk if not present on existing tip' '\n> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> +\tgit reset --hard commits/3 &&\n> +\tgit merge merge/1 &&\n> +\tgit merge commits/5 &&\n> +\tgit merge merge/2 &&\n> +\tgit branch merge/3 &&\n> +\tgit commit-graph write --reachable --split=no-merge &&\n> +\ttest-tool read-graph >output &&\n> +\tcat >expect <<-EOF &&\n> +\theader: 43475048 1 1 4 2\n> +\tnum_commits: 3\n> +\tchunks: oid_fanout oid_lookup commit_metadata\n> +\tEOF\n> +\ttest_cmp expect output &&\n> +\tgit commit-graph verify\n> +'\n> +\n> +test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> +\tgit commit-graph write --reachable --split=replace &&\n> +\ttest_path_is_file $graphdir/commit-graph-chain &&\n> +\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n> +\tverify_chain_files_exist $graphdir &&\n> +\tgraph_read_expect 15 &&\n> +\tgit commit-graph verify\n> +'\n\nIt would be nice to have an example with merging layers (whether we\nwould handle it in strict or relaxed way).\n\n> +\n>  test_done\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404258","messageId":"85imdah50e.fsf@gmail.com","threadId":"53933","inReplyTo":"439adc1718d6cc37f18c1eaeafd605f5c2961733.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 10/11] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-22T19:09:21Z","receivedAt":"2020-08-22T19:09:29Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello Abhishek,\n\n\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> With corrected commit dates implemented, we no longer have to rely on\n> commit date as a heuristic in paint_down_to_common().\n\nAll right, but it would be nice to have some benchmark data: what were\nperformance when using topological levels, what was performance when\nusing commit date heuristics (before this patch), what is performace now\nwhen using corrected commit date.\n\n>\n> t6024-recursive-merge setups a unique repository where all commits have\n> the same committer date without well-defined merge-base. As this has\n> already caused problems (as noted in 859fdc0 (commit-graph: define\n> GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph within the\n> test script.\n\nOK?\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c             | 14 ++++++++++++++\n>  commit-graph.h             |  6 ++++++\n>  commit-reach.c             |  2 +-\n>  t/t6024-recursive-merge.sh |  4 +++-\n>  4 files changed, 24 insertions(+), 2 deletions(-)\n>\n\nI have reorderd files for easier review.\n\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 3cf89d895d..e22ec1e626 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -91,6 +91,12 @@ struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size);\n>   */\n>  int generation_numbers_enabled(struct repository *r);\n>\n> +/*\n> + * Return 1 if and only if the repository has a commit-graph\n> + * file and generation data chunk has been written for the file.\n> + */\n> +int corrected_commit_dates_enabled(struct repository *r);\n> +\n>  enum commit_graph_write_flags {\n>  \tCOMMIT_GRAPH_WRITE_APPEND     = (1 << 0),\n>  \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\n\nAll right.\n\n> diff --git a/commit-graph.c b/commit-graph.c\n> index c1292f8e08..6411068411 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -703,6 +703,20 @@ int generation_numbers_enabled(struct repository *r)\n>  \treturn !!first_generation;\n>  }\n>\n> +int corrected_commit_dates_enabled(struct repository *r)\n> +{\n> +\tstruct commit_graph *g;\n> +\tif (!prepare_commit_graph(r))\n> +\t\treturn 0;\n> +\n> +\tg = r->objects->commit_graph;\n> +\n> +\tif (!g->num_commits)\n> +\t\treturn 0;\n> +\n> +\treturn !!g->chunk_generation_data;\n> +}\n\nThe previous commit introduced validate_mixed_generation_chain(), which\nwalked whole split commit-graph chain, and set `read_generation_data`\nfield in `struct commit_graph` for all layers in the chain.\n\nThis function examines only the top layer, so it follows the assumption\nthat Git would behave in such way that oly topmost layers in the chai\ncan be GDAT-less.\n\nWhy the difference?  Couldn't validate_mixed_generation_chain() simply\ncall corrected_commit_dates_enabled()?\n\n> +\n>  static void close_commit_graph_one(struct commit_graph *g)\n>  {\n>  \tif (!g)\n> diff --git a/commit-reach.c b/commit-reach.c\n> index 470bc80139..3a1b925274 100644\n> --- a/commit-reach.c\n> +++ b/commit-reach.c\n> @@ -39,7 +39,7 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n>  \tint i;\n>  \ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n>\n> -\tif (!min_generation)\n\nThis check was added in 091f4cf (commit: don't use generation numbers if\nnot needed, 2018-08-30) by Derrick Stolee, and its commit message\nincludes benchmark results for running 'git merge-base v4.8 v4.9' in\nLinux kernel repository:\n\n      v2.18.0: 0.122s    167,468 walked\n  v2.19.0-rc1: 0.547s    635,579 walked\n         HEAD: 0.127s\n\n> +\tif (!min_generation && !corrected_commit_dates_enabled(r))\n>  \t\tqueue.compare = compare_commits_by_commit_date;\n\nIt would be nice to have similar benchmark for this change... unless of\ncourse there is no change in performance, but I think then it needs to\nbe stated explicitly.  I think.\n\n>\n>  \tone->object.flags |= PARENT1;\n> diff --git a/t/t6024-recursive-merge.sh b/t/t6024-recursive-merge.sh\n> index 332cfc53fd..d3def66e7d 100755\n> --- a/t/t6024-recursive-merge.sh\n> +++ b/t/t6024-recursive-merge.sh\n> @@ -15,6 +15,8 @@ GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n>  export GIT_COMMITTER_DATE\n>\n>  test_expect_success 'setup tests' '\n> +\tGIT_TEST_COMMIT_GRAPH=0 &&\n> +\texport GIT_TEST_COMMIT_GRAPH &&\n>  \techo 1 >a1 &&\n>  \tgit add a1 &&\n>  \tGIT_AUTHOR_DATE=\"2006-12-12 23:00:00\" git commit -m 1 a1 &&\n> @@ -66,7 +68,7 @@ test_expect_success 'setup tests' '\n>  '\n>\n>  test_expect_success 'combined merge conflicts' '\n> -\ttest_must_fail env GIT_TEST_COMMIT_GRAPH=0 git merge -m final G\n> +\ttest_must_fail git merge -m final G\n>  '\n>\n>  test_expect_success 'result contains a conflict' '\n\nOK, so instead of disabling commit-graph for this test, now we disable\nit for the whole script.\n\nMaybe this change should be in a separate patch?\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404263","messageId":"85y2m6fhkm.fsf@gmail.com","threadId":"53933","inReplyTo":"f6f91af30587ec24e2eee052c89a536cbff42c4f.1597509583.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 11/11] doc: add corrected commit date info","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-22T22:20:57Z","receivedAt":"2020-08-22T22:21:08Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello,\n\n\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> With generation data chunk and corrected commit dates implemented, let's\n> update the technical documentation for commit-graph.\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nAll right.\n\n> ---\n>  .../technical/commit-graph-format.txt         | 12 ++---\n>  Documentation/technical/commit-graph.txt      | 45 ++++++++++++-------\n>  2 files changed, 36 insertions(+), 21 deletions(-)\n>\n> diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n> index 440541045d..71c43884ec 100644\n> --- a/Documentation/technical/commit-graph-format.txt\n> +++ b/Documentation/technical/commit-graph-format.txt\n> @@ -4,11 +4,7 @@ Git commit graph format\n>  The Git commit graph stores a list of commit OIDs and some associated\n>  metadata, including:\n>\n> -- The generation number of the commit. Commits with no parents have\n> -  generation number 1; commits with parents have generation number\n> -  one more than the maximum generation number of its parents. We\n> -  reserve zero as special, and can be used to mark a generation\n> -  number invalid or as \"not computed\".\n> +- The generation number of the commit.\n\nAll right, that was duplicated information.  Now that we need to talk\nabout two of them, it would not make sense to duplicate that.\n\n>\n>  - The root tree OID.\n>\n> @@ -88,6 +84,12 @@ CHUNK DATA:\n\nShouldn't we also replace 'generation number' occurences in description\nof the Commit Data (CDAT) chunk with either 'topological level' or\n'generation number v1'?\n\n>        2 bits of the lowest byte, storing the 33rd and 34th bit of the\n>        commit time.\n>\n> +  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes) [Optional]\n\nIt is not exactly 'optional', as it implies that we need to turn it on\n(or that we can turn it off).  It is more 'conditional', as it can be\nnot present due to outside influences (mixed-version environment).\n\n> +    * This list of 4-byte values store corrected commit date offsets for the\n> +      commits, arranged in the same order as commit data chunk.\n\nI have just realized purely theoretical, but possible, problem with\nstoring non-monotinic generation number related values like corrected\ncommit date offset in constrained space.  There are problems with\nclamping them.\n\nSay that somewhere in the ancestry chain there is a commit A with commit\ndate far in the future by mistake, for example 2120-08-22; it is\nimportant for that date to be not able to be represented using uint32_t.\nSay that a later descendant commit B is malformed, and has committer\ndate of 0, that is 1970-01-01. This means that the corrected commit date\nfor B must be larger than 2120-08-22 - which for this commit means that\ncorrected commit date offset do not fit in 32 bits, and must be clamped\n(replaced) with GENERATION_NUMBER_V2_OFFSET_MAX.\n\nSay that we have commit C that is child of B, and it has correct commit\ndate.  Because of mistake in commit A, it has corrected commit date of\nmore than 2120-08-22 (corrected commit date degenerated into topological\nlevel plus constant).\n\nNow C can reach B, and B can reach A.  However, if we recover corrected\ncommit date of B out of its date=0 and offset=GENERATION_NUMBER_V2_OFFSET_MAX\nwe get a number that is smaller than correct corrected commit date.  We\nwill have\n\n   gen(A) > date(B) + offset(B) < gen(C)\n\nWhich breaks reachability condition guarantee.\n\nIf instead we use GENERATION_NUMBER_V2_MAX for commits with clamped\ncorrected commit date, that is offset=GENERATION_NUMBER_V2_OFFSET_MAX,\nwe would get\n\n  gen(A) < GENERATION_NUMBER_V2_MAX > gen(C)\n\nAnd again reachability condition is broken.\n\nThis is a very contrived but possible example.  This shouldn't happen,\nbut ufortunately it can happen.\n\n\nThe question is how to deal with this issue.  Ignore it as unlikely?\nSwitch to storing corrected commit date, which is monotonic, so if there\nis commit with GENERATION_NUMBER_V2_MAX, then subsequent descendant\ncommits will also have GENERATION_NUMBER_V2_MAX -- and pay with up to 7%\nlarger commit-graph file?\n\n> +    * This list can be later modified to store future generation number related\n> +      data.\n\nHow can it be later modified?  There is no header, no version number.\nHow would we add another generation number data?\n\n> +\n>    Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n>        This list of 4-byte values store the second through nth parents for\n>        all octopus merges. The second parent value in the commit data stores\n> diff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\n> index 808fa30b99..f27145328c 100644\n> --- a/Documentation/technical/commit-graph.txt\n> +++ b/Documentation/technical/commit-graph.txt\n> @@ -38,14 +38,27 @@ A consumer may load the following info for a commit from the graph:\n>\n>  Values 1-4 satisfy the requirements of parse_commit_gently().\n>\n> -Define the \"generation number\" of a commit recursively as follows:\n> +There are two definitions of generation number:\n> +1. Corrected committer dates\n> +2. Topological levels\n\nShould we add versioning info, that is:\n\n  +1. Corrected committer dates  (generation number v2)\n  +2. Topological levels  (generation number v1)\n\n> +\n> +Define \"corrected committer date\" of a commit recursively as follows:\n> +\n> +  * A commit with no parents (a root commit) has corrected committer date\n> +    equal to its committer date.\n> +\n> +  * A commit with at least one parent has corrected committer date equal to\n> +    the maximum of its commiter date and one more than the largest corrected\n> +    committer date among its parents.\n> +\n> +Define the \"topological level\" of a commit recursively as follows:\n>\n>   * A commit with no parents (a root commit) has generation number one.\n\nShouldn't this be\n\n    * A commit with no parents (a root commit) has topological level of one.\n\n>\n> - * A commit with at least one parent has generation number one more than\n> -   the largest generation number among its parents.\n> + * A commit with at least one parent has topological level one more than\n> +   the largest topological level among its parents.\n>\n> -Equivalently, the generation number of a commit A is one more than the\n> +Equivalently, the topological level of a commit A is one more than the\n>  length of a longest path from A to a root commit. The recursive definition\n>  is easier to use for computation and observing the following property:\n\nWe should probably explicitly state that the property state applies to\nboth versions of generation number, not only to topological level.\n\n>\n> @@ -67,17 +80,12 @@ numbers, the general heuristic is the following:\n>      If A and B are commits with commit time X and Y, respectively, and\n>      X < Y, then A _probably_ cannot reach B.\n>\n> -This heuristic is currently used whenever the computation is allowed to\n> -violate topological relationships due to clock skew (such as \"git log\"\n> -with default order), but is not used when the topological order is\n> -required (such as merge base calculations, \"git log --graph\").\n> -\n\nTo be overly pedantic, this heuristic is still used, but now in much\nmore rare case.  In addition to what is stated above, at least one layer\nin the split commit-graph chain must have been generated by \"Old\" Git,\nfor the date heuristic to be used.\n\nBut that might be unnecessary level of detail.\n\n>  In practice, we expect some commits to be created recently and not stored\n>  in the commit graph. We can treat these commits as having \"infinite\"\n>  generation number and walk until reaching commits with known generation\n>  number.\n>\n> -We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n> +We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n\nAll right.\n\n>  in the commit-graph file. If a commit-graph file was written by a version\n>  of Git that did not compute generation numbers, then those commits will\n>  have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n> @@ -93,12 +101,11 @@ fully-computed generation numbers. Using strict inequality may result in\n>  walking a few extra commits, but the simplicity in dealing with commits\n>  with generation number *_INFINITY or *_ZERO is valuable.\n>\n> -We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n> -generation numbers are computed to be at least this value. We limit at\n> -this value since it is the largest value that can be stored in the\n> -commit-graph file using the 30 bits available to generation numbers. This\n> -presents another case where a commit can have generation number equal to\n> -that of a parent.\n> +We use the macro GENERATION_NUMBER_MAX for commits whose generation numbers\n> +are computed to be at least this value. We limit at this value since it is\n> +the largest value that can be stored in the commit-graph file using the\n> +available to generation numbers. This presents another case where a\n> +commit can have generation number equal to that of a parent.\n\nAll right, though it could have been done without re-wrapping, so that\nonly first line would be marked as changed.\n\nAs I wrote, there is theoretical problem with this for offsets.\n\n>\n>  Design Details\n>  --------------\n> @@ -267,6 +274,12 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n>  number of commits) could be extracted into config settings for full\n>  flexibility.\n>\n> +We also merge commit-graph chains when we try to write a commit graph with\n> +two different generation number definitions as they cannot be compared directly.\n> +We overwrite the existing chain and create a commit-graph with the newer or more\n> +efficient defintion. For example, overwriting topological levels commit graph\n> +chain to create a corrected commit dates commit graph chain.\n> +\n\nThis is more complicated than that.\n\nI think we should explicitly state that Git ensures that in split\ncommit-graph chain, if there are layers without the GDAT chunk (that\nforce Git to use topological levels for generation numbers), then they\nare top layers.  So if there is commit-graph file created by \"Old\" Git,\nthen when addig new layer it would also be GDAT-less.\n\nNow how to write this...\n\n>  ## Deleting graph-{hash} files\n>\n>  After a new tip file is written, some `graph-{hash}` files may no longer\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404296","messageId":"85blj1e619.fsf@gmail.com","threadId":"53933","inReplyTo":"85zh6uxh7l.fsf@gmail.com","subject":"Re: [PATCH v3 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-23T15:27:46Z","receivedAt":"2020-08-23T15:30:31Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello,\n\nHere is a summary of my comments and thoughts after carefully reviewing\nall patches in the series.\n\nJakub Narębski <jnareb@gmail.com> writes:\n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n[...]\n>> The Corrected Commit Date is defined as:\n>>\n>> For a commit C, let its corrected commit date be the maximum of the commit\n>> date of C and the corrected commit dates of its parents.\n>\n> Actually it needs to be \"corrected commit dates of its parents plus 1\"\n> to fulfill the reachability condition for a generation number for a\n> commit:\n>\n>       A can reach B   =>  gen(A) < gen(B)\n>\n> Of course it can be computed in simpler way, because\n>\n>   max_P (gen(P) + 1)  ==  max_P (gen(P)) + 1\n\nI see that it is defined correctly in the documentation, which is more\nimportant than cover letter (which would not be stored in the repository\nfor memory, even in the commit message for the merge).\n\n[...]\n> However I think the cover letter should also describe what should happen\n> in a mixed version environment (for example new Git on command line,\n> copy of old Git used by GUI client), and in particular what should\n> happen in a mixed-chain case - both for reading and for writing the\n> commit-graph file.\n>\n> For *writing*: because old Git would create commit-graph layers without\n> the GDAT chunk, to simplify the behavior and make easy to reason about\n> commit-graph data (the situation should be not that common, and\n> transient -- it should get more rare as the time goes), we want the\n> following behavior from new Git:\n>\n> - If top layer contains the GDAT chunk, or we are rewriting commit-graph\n>   file (--split=replace), or we are merging layers and there are no\n>   layers without GDAT chunk below set of layers that are merged, then\n>\n>      write commit-graph file or commit-graph layer with GDAT chunk,\n\nActually we can simplify handling of merging layers, while still\nretaining the property that in mixed-version chain all GDAT-full layers\nare before / below GDAT-less layers.\n\nNamely, if merging layers, and at least one layer being merged doesn't\nhave GDAT chunk, then the result of merge wouldn't have either.\n\nWe can always switch to slightly more complicated behavior proposed\nabove in quoted part later, perhaps as a followup commit.\n\n>\n>   otherwise\n>\n>      write commit-graph layer without GDAT chunk.\n>\n>   This means that there are commit-graph layers without GDAT chunk if\n>   and only if the top layer is also without GDAT chunk.\n\nThis might be not necessary if Git would always check the whole chain of\nsplit commit-graph layers for presence of GDAT-less layers.\n\nBut I still think it is a good idea to avoid having GDAT-less \"holes\".\n\n>\n> For *reading* we want to use generation number v2 (corrected commit\n> date) if possible, and fall back to generation number v1 (topological\n> levels).\n>\n> - If the top layer contains the GDAT chunk (or maybe even if the topmost\n>   layer that involves all commits in question, not necessarily the top\n>   layer in the full commit-graph chain), then use generation number v2\n>\n>   - commit_graph_data->generation stores corrected commit date,\n>     computed as sum of committer date (from CDAT) and offset (from GDAT)\n\nOr stored directly in GDAT, at the cost of increasing the file size by\nat most 7% (if I have done my math correctly).\n\nSee also the issue with clamping offsets (GENERATION_NUMBER_V2_OFFSET_MAX).\n\n>\n>   - A can reach B   =>  gen(A) < gen(B)\n>\n>   - there is no need for committer date heuristics, and no need for\n>     limiting use of generation number to where there is a cutoff (to not\n>     hamper performance).\n>\n> - If there are layers without GDAT chunks, which thanks to the write\n>   behavior means simply top layer without GDAT chunk, we need to turn\n>   off use of generation numbers or fall back to using topological levels\n>\n>   - commit_graph_data->generation stores topological levels,\n>     taken from CDAT chunk (30-bits)\n>\n>   - A can reach B   =>  gen(A) < gen(B)\n>\n>   - we probably want to keep tie-breaking of sorting by generation\n>     number via committer date, and limit use of generation number as\n>     opposed to using committer date heuristics (with slop) to not make\n>     performance worse.\n\nAnd this is being done in this patch series.  Good!\n\nThe thing I was worrying about turned out to be non-issue, as the\ncomparison function in question is used only when writing, and in this\ncase we have corrected commit date computer - though perhaps not being\nwritten (as far as I understand it, but I might be mistaken).\n\n[...]\n>>\n>> Abhishek Kumar (11):\n>>   commit-graph: fix regression when computing bloom filter\n\nNo problems, maybe just expanding a commit message and/or adding a comment.\n\n>>   revision: parse parent in indegree_walk_step()\n\nLooks good to me.\n\n>>   commit-graph: consolidate fill_commit_graph_info\n\nI think this commit could be split into three:\n- fix to the 'generate tar with future mtime' test\n  as it is a hidden bug (when using commit-graph)\n- simplify implementation of fill_commit_in_graph()\n  by using fill_commit_graph_info()\n- move loading date into fill_commit_graph_info()\n  that uncovers the issue with 'generate tar with future mtime'\n\nOn the other hand because they are inter-related, those changes might be\nkept in a single commit.\n\nIn commit message greater care needs to be taken with\nfill_commit_graph_info() and fill_commit_in_graph(), when to use one and\nwhen to use the other. For example it is fill_commit_graph_info() that\nchanges its behavior, and it is fill_commit_in_graph() that is getting\nsimplified.\n\n>>   commit-graph: consolidate compare_commits_by_gen\n\nLooks good to me, though it might be good idea to add comments about the\nsorting order (inside comment) to appropriate header files.\n\n>>   commit-graph: return 64-bit generation number\n\nThis conversion misses one place, though it would be changed to use\ntopological levels slab in next patch.\n\nThis patch also unnecessary introduces GENERATION_NUMBER_V1_INFINITY.\nThere is no need for it: `generation` field can always simply use\nGENERATION_NUMBER_INFINITY for commits not in commit-graph.\n\n>>   commit-graph: add a slab to store topological levels\n\nThis is the patch that needs GENERATION_NUMBER_V1_INFINITY (or\nTOPOLOGICAL_LEVEL_INFINITY), not the previous patch.\n\nDetailed analysis uncovered hidden bug in the code of\ncompute_generation_numbers() that works only because of historical\nreasons (that topological levels in Git start from 1, not from 0). The\nproblem is that the 'level' / 'generation' variable for commits not in\ngraph, and therefore ones that needs to have its generation number\ncomputed, is 0 (default value on commit slab) and not\nGENERATION_NUMBER*_INFINITY as it should.\n\nThis issue is present since moving commit graph info data to\ncommit-slab.  We can simply document it and ignore (it works, if by\naccident), or try to fix it.\n\nWe need to handle GENERATION_NUMBER*_MAX clamping carefully\nin the future patches.\n\nI think also that this patch needs a bit more detailed commit message.\n\n>>   commit-graph: implement corrected commit date\n\nLooks good, though verify_commit_graph() now verifies *a* generation\nnumber used, be it v1 (topological level) or v2 (corrected commit date),\nso the variable rename is unnecessary.  We verify that they fulfill the\nreachability condition promise, that is gen(parent) <= gen(commit),\n(the '=' is to include GENERATION_NUMBER*_MAX case), as it is what is\nused to speed up graph traversal.\n\nWe probably want to verify both topological level values in CDAT, and if\nthey exist also corrected commit date values in GDAT.  But that might be\nleft for the future commit.\n\n>>   commit-graph: implement generation data chunk\n\nTo save up to 6-7% on commit-graph file size we store 32-bits corrected\ncommit date offsets, instead of storing 64-bits corrected commit date.\n\nHowever, as far as I understand it, using non-monotonic values for\non-disk storage with limited field size, like 32-bits corrected\ncommit date offset, leads to problems with GENERATION_NUMBER*_MAX.\nNamely, as I have written in detail in my reply to patch 11/11 in this\nseries, there is no way to fulfill the reachability condition promise if\nwe have to store offset which true value do not fit in 32-bits reserved\nfor it.\n\nThis is extremly unlikely to happen in practice, but we need to be able\nto handle it somehow.  We can store 64-bit corrected commit date, which\nhas graph-monotonic values, and the problem goes away in theory and in\npractice (we would never have values that do not fit).  We can keep\nstoring 32-bit offsets, and simply do not use GDAT chunk if there is\noffset value that do not fit.\n\nAll this of course, provided that I am not wrong about this issue...\n\n>>   commit-graph: use generation v2 only if entire chain does\n\nThe commit message do not say anything about the *writing* side.\n\nHowever, if we want to keep the requirement that GDAT-less layers in the\nsplit commit-graph chain must all be above any GDAT-containing layers,\nwe need to consider how we want layer merging to behave.  We could\neither opt for using GDAT whenever possible, or similify code and skip\nusing GDAT if we are unsure.\n\nThe first approach would mean that if the topmost layer below set of\nlayers being rewritten (in the split commit-graph chain) exists, and it\ndoes not contain GDAT chunk, then and only then the result of rewrite\nshould not have GDAT chunk either.\n\nThe second approach is, I think, simpler: if any of layers that is being\nrewritten is GDAT-less (we need to check only the top layer, though),\nand we are not running full rewrite ('--split=replace'), then the result\nof rewrite should not have GDAT chunk either.  We can switch to the\nfirst algorithm in later commit.\n\nWhether we choose one or the other, we need test that doing layer\nmerging do not break GDAT-inness requirement stated above.\n\n\nAlso, we can probably test that we are not using v1 and v2 at the same\ntime with tests involving --is-ancestor, or --contains / --merged.\n\n>>   commit-reach: use corrected commit dates in paint_down_to_common()\n\nI think this commit could be split into two:\n- disable commit graph for entire t6024-recursive-merge test\n- use corrected commit dates in paint_down_to_common()\n\nOn the other hand because they are inter-related, those changes might be\nbetter kept in a single commit.\n\nIt would be nice to have some benchmark data, or at least stating that\nperformance does not change (within the error bounds), using for example\n'git merge-base v4.8 v4.9' in Linux kernel repository.\n\n>>   doc: add corrected commit date info\n\nLooks good, there are a few places where 'generation number' (referring\nto the v1 version) should have been replaced with 'topological level'.\n\nI am also unsure how the GDAT chunk can be \"later modified to store\nfuture generation number related data.\".  I'd like to have an example,\nor for this statement to be removed (if it turns out to not be true, not\nwithout introducing yet another chunk type).\n\n>>  18 files changed, 396 insertions(+), 185 deletions(-)\n\nThanks for all your work.\n\nBest regards,\n--\nJakub Narębski\n"},{"id":"404304","messageId":"20200824024909.GA38636@Abhishek-Arch","threadId":"53933","inReplyTo":"85blj1e619.fsf@gmail.com","subject":"Re: [PATCH v3 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-08-24T02:49:09Z","receivedAt":"2020-08-24T02:51:31Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"Hello,\n\nOn Sun, Aug 23, 2020 at 05:27:46PM +0200, Jakub Narębski wrote:\n> Hello,\n> \n> Here is a summary of my comments and thoughts after carefully reviewing\n> all patches in the series.\n> \n\nThanks for the detailed review. It must have taken you a lot of time and\nfocus to go through the patches, so I have really appreciate the effort.\n\nI have been going over the comments and they are reasonable and very\nmuch needed. I will respond to the mails in the specific threads (to\nkeep the discussion locally scoped and thus manageable), with doubts\nand comments that I have as I try to follow through the suggestions.\n\n> Best regards,\n> --\n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"404390","messageId":"20200825050448.GA21012@Abhishek-Arch","threadId":"53933","inReplyTo":"85d03km98l.fsf@gmail.com","subject":"Re: [PATCH v3 05/11] commit-graph: return 64-bit generation number","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-08-25T05:04:48Z","receivedAt":"2020-08-25T05:07:07Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Fri, Aug 21, 2020 at 03:14:34PM +0200, Jakub Narębski wrote:\n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > In a preparatory step, let's return timestamp_t values from\n> > commit_graph_generation(), use timestamp_t for local variables\n> \n> All right, this is all good.\n> \n> > and define GENERATION_NUMBER_INFINITY as (2 ^ 63 - 1) instead.\n> \n> This needs more detailed examination.  There are two similar constants,\n> GENERATION_NUMBER_INFINITY and GENERATION_NUMBER_MAX.  The former is\n> used for newest commits outside the commit-graph, while the latter is\n> maximum number that commits in the commit-graph can have (because of the\n> storage limitations).  We therefore need GENERATION_NUMBER_INFINITY\n> to be larger than GENERATION_NUMBER_MAX, and it is (and was).\n> \n> The GENERATION_NUMBER_INFINITY is because of the above requirement\n> traditionally taken as maximum value that can be represented in the data\n> type used to store commit's generation number _in memory_, but it can be\n> less.  For timestamp_t the maximum value that can be represented\n> is (2 ^ 63 - 1).\n> \n> All right then.\n> \n> >\n\nRelated to this, by the end of this series we are using\nGENERATION_NUMBER_MAX in just one place - compute_generation_numbers()\nto make sure the topological levels fit within 30 bits.\n\nWould it be more appropriate to rename GENERATION_NUMBER_MAX to\nGENERATION_NUMBER_V1_MAX (along the lines of\nGENERATION_NUMBER_V2_OFFSET_MAX)  to correctly describe that is a\nlimit on topological levels, rather than generation number value?\n\n> \n> The commit message says nothing about the new symbolic constant\n> GENERATION_NUMBER_V1_INFINITY, though.\n> \n> I'm not sure it is even needed (see comments below).\n\nYes, you are correct. I tried it out with your suggestions and it wasn't\nreally needed.\n\nThanks for catching this!\n\n> ...\n\nThanks\n- Abhishek\n"},{"id":"404391","messageId":"20200825061418.GA629699@Abhishek-Arch","threadId":"53933","inReplyTo":"85d03jlu05.fsf@gmail.com","subject":"Re: [PATCH v3 06/11] commit-graph: add a slab to store topological levels","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-08-25T06:14:18Z","receivedAt":"2020-08-25T06:16:39Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Fri, Aug 21, 2020 at 08:43:38PM +0200, Jakub Narębski wrote:\n> Hello,\n> \n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > As we are writing topological levels to commit data chunk to ensure\n> > backwards compatibility with \"Old\" Git and the member `generation` of\n> > struct commit_graph_data will store corrected commit date in a later\n> > commit, let's introduce a commit-slab to store topological levels while\n> > writing commit-graph.\n> \n> In my opinion the above it would be easier to follow if rephrased in the\n> following way:\n> \n>   In a later commit we will introduce corrected commit date as the\n>   generation number v2.  This value will be stored in the new separate\n>   GDAT chunk.  However to ensure backwards compatibility with \"Old\" Git\n>   we need to continue to write generation number v1, which is\n>   topological level, to the commit data chunk (CDAT).  This means that\n>   we need to compute both versions of generation numbers when writing\n>   the commit-graph file.  Let's therefore introduce a commit-slab\n>   to store topological levels; corrected commit date will be stored\n>   in the member `generation` of struct commit_graph_data.\n> \n> What do you think?\n> \n\nYes, that's better.\n\n> \n> By the way, do I understand it correctly that in backward-compatibility\n> mode (that is, in mixed-version environment where at least some\n> commit-graph files were written by \"Old\" Git and are lacking GDAT chunk\n> and generation number v2 data) the `generation` member of commit graph\n> data chunk will be populated and will store generation number v1, that\n> is topological level? And that the commit-slab for topological levels is\n> only there for writing and re-writing?\n> \n\nNo, the topo_levels commit-slab would be always populated when we write\na commit data chunk. The topo_level slab is a workaround for the fact\nthat we are trying to write two independent values (corrected commit\ndate offset, topological levels) but have one struct member to store them in\n(data->generation).\n\nIf we are in a mixed-version environment, we could avoid initializing\nthe slab and fill the topological levels into data->generation instead,\nbut that's not how it is implemented right now.\n\n> >\n> > When Git creates a split commit-graph, it takes advantage of the\n> > generation values that have been computed already and present in\n> > existing commit-graph files.\n> >\n> > So, let's add a pointer to struct commit_graph to the topological level\n> > commit-slab and populate it with topological levels while writing a\n> > split commit-graph.\n> \n> All right, looks sensible.\n\nI have extend the last paragraph to include write_commit_graph_context\nas well as:\n\n  So, let's add a pointer to struct commit_graph as well as struct\n  write_commit_graph_context to the topological level commit-slab and\n  populate it with topological levels while writing a commit-graph file.\n\n> \n> >\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c | 47 ++++++++++++++++++++++++++++++++---------------\n> >  commit-graph.h |  1 +\n> >  commit.h       |  1 +\n> >  3 files changed, 34 insertions(+), 15 deletions(-)\n> >\n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index 7f9f858577..a2f15b2825 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -64,6 +64,8 @@ void git_test_write_commit_graph_or_die(void)\n> >  /* Remember to update object flag allocation in object.h */\n> >  #define REACHABLE       (1u<<15)\n> >\n> > +define_commit_slab(topo_level_slab, uint32_t);\n> > +\n> \n> All right.\n> \n> Also, here we might need GENERATION_NUMBER_V1_INFINITY, but I don't\n> think it would be necessary.\n> \n> >  /* Keep track of the order in which commits are added to our list. */\n> >  define_commit_slab(commit_pos, int);\n> >  static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n> > @@ -759,6 +761,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n> >  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n> >\n> >  \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> > +\n> > +\tif (g->topo_levels)\n> > +\t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n> >  }\n> \n> All right, here we store topological levels on commit-slab to avoid\n> recomputing them.\n> \n> Do I understand it correctly that the `topo_levels` member of the `struct\n> commit_graph` would be non-null only when we are updating the\n> commit-graph?\n> \n\nYes, that's correct.\n\n> >\n> >  static inline void set_commit_tree(struct commit *c, struct tree *t)\n> > @@ -953,6 +958,7 @@ struct write_commit_graph_context {\n> >  \t\t changed_paths:1,\n> >  \t\t order_by_pack:1;\n> >\n> > +\tstruct topo_level_slab *topo_levels;\n> >  \tconst struct split_commit_graph_opts *split_opts;\n> >  \tsize_t total_bloom_filter_data_size;\n> >  \tconst struct bloom_filter_settings *bloom_settings;\n> \n> Why do we need `topo_levels` member *both* in `struct commit_graph` and\n> in `struct write_commit_graph_context`?\n> \n> [After examining the change further I have realized why both are needed,\n>  and written about the reasoning later in this email.]\n> \n> \n> Note that the commit message talks only about `struct commit_graph`...\n> \n> > @@ -1094,7 +1100,7 @@ static int write_graph_chunk_data(struct hashfile *f,\n> >  \t\telse\n> >  \t\t\tpackedDate[0] = 0;\n> >\n> > -\t\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n> > +\t\tpackedDate[0] |= htonl(*topo_level_slab_at(ctx->topo_levels, *list) << 2);\n> \n> All right, here we prepare for writing to the CDAT chunk using data that\n> is now stored on newly introduced topo_levels slab (either computed, or\n> taken from commit-graph file being rewritten).\n> \n> Assuming that ctx->topo_levels is not-null, and that the values are\n> properly calculated before this -- and we did compute topological levels\n> before writing the commit-graph.\n> \n> >\n> >  \t\tpackedDate[1] = htonl((*list)->date);\n> >  \t\thashwrite(f, packedDate, 8);\n> > @@ -1335,11 +1341,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n> >  \t\t\t\t\tctx->commits.nr);\n> >  \tfor (i = 0; i < ctx->commits.nr; i++) {\n> > -\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n> > +\t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n> \n> All right, so that is why this 'generation' variable was not converted\n> to timestamp_t type.\n> \n> >\n> >  \t\tdisplay_progress(ctx->progress, i + 1);\n> > -\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n> > -\t\t    generation != GENERATION_NUMBER_ZERO)\n> > +\t\tif (level != GENERATION_NUMBER_V1_INFINITY &&\n> > +\t\t    level != GENERATION_NUMBER_ZERO)\n> >  \t\t\tcontinue;\n> \n> Here we use GENERATION_NUMBER*_INFINITY to check if the commit is\n> outside commit-graph files, and therefore we would need its topological\n> level computed.\n> \n> However, I don't understand how it works.  We have had created the\n> commit_graph_data_at() and use it instead of commit_graph_data_slab_at()\n> to provide default values for `struct commit_graph`... but only for\n> `graph_pos` member.  It is commit_graph_generation() that returns\n> GENERATION_NUMBER_INFINITY for commits not in graph.\n> \n> But neither commit_graph_data_at()->generation nor topo_level_slab_at()\n> handles this special case, so I don't see how 'generation' variable can\n> *ever* be GENERATION_NUMBER_INFINITY, and 'level' variable can ever be\n> GENERATION_NUMBER_V1_INFINITY for commits not in graph.\n> \n> Does it work *accidentally*, because the default value for uninitialized\n> data on commit-slab is 0, which matches GENERATION_NUMBER_ZERO?  It\n> certainly looks like it does.  And GENERATION_NUMBER_ZERO is an artifact\n> of commit-graph feature development history, namely the short time where\n> Git didn't use any generation numbers and stored 0 in the place set for\n> it in the commit-graph format...  On the other hand this is not the case\n> for corrected commit date (generation number v2), as it could\n> \"legitimately\" be 0 if some root commit (without any parents) had\n> committerdate of epoch 0, i.e. 1 January 1970 00:00:00 UTC, perhaps\n> caused by malformed but valid commit object.\n> \n> Ugh...\n\nIt works accidentally.\n\nOur decision to avoid the cost of initializing both\ncommit_graph_data->generation and commit_graph_data->graph_pos has\nled to some unwieldy choices - the complexity of helper functions,\nbypassing helper functions when writing a commit-graph file [1].\n\nI want to re-visit how commit_graph_data slab is defined in a future series.\n\n[1]: https://lore.kernel.org/git/be28ab7b-0ae4-2cc5-7f2b-92075de3723a@gmail.com/\n\n> \n> >\n> >  \t\tcommit_list_insert(ctx->commits.list[i], &list);\n> > @@ -1347,29 +1353,27 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\tstruct commit *current = list->item;\n> >  \t\t\tstruct commit_list *parent;\n> >  \t\t\tint all_parents_computed = 1;\n> > -\t\t\tuint32_t max_generation = 0;\n> > +\t\t\tuint32_t max_level = 0;\n> >\n> >  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n> > -\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n> > +\t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n> >\n> > -\t\t\t\tif (generation == GENERATION_NUMBER_V1_INFINITY ||\n> > -\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n> > +\t\t\t\tif (level == GENERATION_NUMBER_V1_INFINITY ||\n> > +\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n> >  \t\t\t\t\tall_parents_computed = 0;\n> >  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n> >  \t\t\t\t\tbreak;\n> > -\t\t\t\t} else if (generation > max_generation) {\n> > -\t\t\t\t\tmax_generation = generation;\n> > +\t\t\t\t} else if (level > max_level) {\n> > +\t\t\t\t\tmax_level = level;\n> >  \t\t\t\t}\n> >  \t\t\t}\n> \n> This is the same case as for previous chunk; see the comment above.\n> \n> This code checks if parents have generation number / topological level\n> computed, and tracks maximum value of it among all parents.\n> \n> >\n> >  \t\t\tif (all_parents_computed) {\n> > -\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n> > -\n> > -\t\t\t\tdata->generation = max_generation + 1;\n> >  \t\t\t\tpop_commit(&list);\n> >\n> > -\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n> > -\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n> > +\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n> > +\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n> > +\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n> \n> OK, this is safer way of handling GENERATION_NUMBER*_MAX, especially if\n> this value can be maximum value that can be safely stored in a given\n> data type.  Previously GENERATION_NUMBER_MAX was smaller than maximum\n> value that can be safely stored in uint32_t, so generation+1 had no\n> chance to overflow.  This is no longer the case; the reorganization done\n> here leads to more defensive code (safer).\n> \n> All good.  However I think that we should clamp the value of topological\n> level to the maximum value that can be safely stored *on disk*, in the\n> 30 bits of the CDAT chunk reserved for generation number v1.  Otherwise\n> the code to write topological level would get more complicated.\n> \n> In my opinion the symbolic constant used here should be named\n> GENERATION_NUMBER_V1_MAX, and its value should be at most (2 ^ 30 - 1);\n> it should be the current value of GENERATION_NUMBER_MAX, that is\n> 0x3FFFFFFF.\n> \n> >  \t\t\t}\n> >  \t\t}\n> >  \t}\n> > @@ -2101,6 +2105,7 @@ int write_commit_graph(struct object_directory *odb,\n> >  \tuint32_t i, count_distinct = 0;\n> >  \tint res = 0;\n> >  \tint replace = 0;\n> > +\tstruct topo_level_slab topo_levels;\n> >\n> \n> All right, we will be using topo_level slab for writing the\n> commit-graph, and only for this purpose, so it is good to put it here.\n> \n> >  \tif (!commit_graph_compatible(the_repository))\n> >  \t\treturn 0;\n> > @@ -2179,6 +2184,18 @@ int write_commit_graph(struct object_directory *odb,\n> >  \t\t}\n> >  \t}\n> >\n> > +\tinit_topo_level_slab(&topo_levels);\n> > +\tctx->topo_levels = &topo_levels;\n> > +\n> > +\tif (ctx->r->objects->commit_graph) {\n> > +\t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n> > +\n> > +\t\twhile (g) {\n> > +\t\t\tg->topo_levels = &topo_levels;\n> > +\t\t\tg = g->base_graph;\n> > +\t\t}\n> > +\t}\n> \n> All right, now I see why we need `topo_levels` member both in the\n> `struct write_commit_graph_context` and in `struct commit_graph`.\n> The former is for functions that write the commit-graph, the latter for\n> fill_commit_graph_info() functions that is deep in the callstack, but it\n> needs to know whether to load topological level to commit-slab, or maybe\n> put it as generation number (and in the future -- discard it, if not\n> needed).\n> \n> \n> Sidenote: this fragment of code, that fills with a given value some\n> member of the `struct commit_graph` throughout the split commit-graph\n> chain, will be repeated as similar code in patches later in series.\n> However without resorting to preprocessor macros I have no idea how to\n> generalize it to avoid code duplication (well, almost).\n> \n\nThe pattern is: iterate over the commit-graph chain and assign a member\n(here, topo_level and in the other patch, read_generation_data) a value\n(the address of topo_level slab, 1/0 depending on whether it is a mixed\ngeneration chain).\n\nWe could generalize this in a future series but I don't think it is\nworthwhile.\n\n> > +\n> >  \tif (pack_indexes) {\n> >  \t\tctx->order_by_pack = 1;\n> >  \t\tif ((res = fill_oids_from_packs(ctx, pack_indexes)))\n> > diff --git a/commit-graph.h b/commit-graph.h\n> > index 430bc830bb..1152a9642e 100644\n> > --- a/commit-graph.h\n> > +++ b/commit-graph.h\n> > @@ -72,6 +72,7 @@ struct commit_graph {\n> >  \tconst unsigned char *chunk_bloom_indexes;\n> >  \tconst unsigned char *chunk_bloom_data;\n> >\n> > +\tstruct topo_level_slab *topo_levels;\n> >  \tstruct bloom_filter_settings *bloom_filter_settings;\n> >  };\n> \n> All right: `struct commit_graph` is public, `struct\n> write_commit_graph_context` is not.\n> \n> >\n> > diff --git a/commit.h b/commit.h\n> > index bc0732a4fe..bb846e0025 100644\n> > --- a/commit.h\n> > +++ b/commit.h\n> > @@ -15,6 +15,7 @@\n> >  #define GENERATION_NUMBER_V1_INFINITY 0xFFFFFFFF\n> >  #define GENERATION_NUMBER_MAX 0x3FFFFFFF\n> \n> The name GENERATION_NUMBER_MAX for 0x3FFFFFFF should be instead\n> GENERATION_NUMBER_V1_MAX, but that may be done in a later commit.\n> \n> >  #define GENERATION_NUMBER_ZERO 0\n> > +#define GENERATION_NUMBER_V2_OFFSET_MAX 0xFFFFFFFF\n> \n> This value is never used, so why it is defined in this commit.\n> \n\nMoved down to the patch actually uses it.\n\n> >\n> >  struct commit_list {\n> >  \tstruct commit *item;\n> \n> Best,\n> -- \n> Jakub Narębski\n"},{"id":"404392","messageId":"20200825064954.GA645690@Abhishek-Arch","threadId":"53933","inReplyTo":"85wo1rk0iy.fsf@gmail.com","subject":"Re: [PATCH v3 07/11] commit-graph: implement corrected commit date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-08-25T06:49:54Z","receivedAt":"2020-08-25T06:52:17Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Sat, Aug 22, 2020 at 02:05:41AM +0200, Jakub Narębski wrote:\n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > With most of preparations done, let's implement corrected commit date.\n> >\n> > The corrected commit date for a commit is defined as:\n> >\n> > * A commit with no parents (a root commit) has corrected commit date\n> >   equal to its committer date.\n> > * A commit with at least one parent has corrected commit date equal to\n> >   the maximum of its commit date and one more than the largest corrected\n> >   commit date among its parents.\n> \n> Good.\n> \n> >\n> > To minimize the space required to store corrected commit date, Git\n> > stores corrected commit date offsets into the commit-graph file. The\n> > corrected commit date offset for a commit is defined as the difference\n> > between its corrected commit date and actual commit date.\n> \n> Perhaps we should add more details about data type sizes in question.\n\nWill add.\n\n> \n> Storing corrected commit date requires sizeof(timestamp_t) bytes, which\n> in most cases is 64 bits (uintmax_t).  However corrected commit date\n> offsets can be safely stored^* using only 32 bits.  This halves the size\n> of GDAT chunk, reducing per-commit storage from 2*H + 16 + 8 bytes to\n> 2*H + 16 + 4 bytes, which is reduction of around 6%, not including\n> header, fanout table (OIDF) and extra edges list (EDGE).\n> \n> Which might mean that the extra complication is not worth it, and we\n> should store corrected commit date directly instead.\n> \n> *) unless for example one of commits is malformed but valid,\n>    and has committerdate of 0 Unix time, 1 January 1970.\n> \n> >\n> > While Git does not write out offsets at this stage, Git stores the\n> > corrected commit dates in member generation of struct commit_graph_data.\n> > It will begin writing commit date offsets with the introduction of\n> > generation data chunk.\n> \n> OK, so the agenda for introducing geeration number v2 is as follows:\n> - compute generation numbers v2, i.e. corrected commit date\n> - store corrected commit date [offsets] in new GDAT chunk,\n>   unless backward-compatibility concerns require us to not to\n> - load [and compute] corrected commit date from commit-graph\n>   storing it as 'generation' field of `struct commit_graph_data`,\n>   unless backward-compatibility concerns require us to store\n>   topological levels (generation number v1) in there instead\n> \n\nThe last point is not correct. We always store topological levels into\nthe topo_levels slab introduced and always store corrected commit date\ninto data->generation, regardless of backward compatibility concerns.\n\nWe could avoid initializing topo_slab if we are not writing generation\ndata chunk (and thus don't need corrected commit dates) but that\nwouldn't have an impact on run time while writing commit-graph because\ncomputing corrected commit dates is cheap as the main cost is in walking\nthe graph and writing the file.\n\n> Because the reachability condition for corrected commit date and for\n> topological level is exactly the same, we don't need to do anything to\n> take advantage of generation number v2.\n> \n> Though we can use generation number v2 in more cases, where we turned\n> off use of generation numbers because v1 gave worse performance than\n> date heuristics.\n> \n> Did I got this right?\n> \n> >\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c | 58 +++++++++++++++++++++++++++-----------------------\n> >  1 file changed, 31 insertions(+), 27 deletions(-)\n> >\n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index a2f15b2825..fd69534dd5 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -169,11 +169,6 @@ static int commit_gen_cmp(const void *va, const void *vb)\n> >  \telse if (generation_a > generation_b)\n> >  \t\treturn 1;\n> >\n> > -\t/* use date as a heuristic when generations are equal */\n> > -\tif (a->date < b->date)\n> > -\t\treturn -1;\n> > -\telse if (a->date > b->date)\n> > -\t\treturn 1;\n> \n> At first I was wondering why this tie-breaking is beig removed; wouldn't\n> be needed for backward-compatibility?  But then I remembered that this\n> comparison function is used _only_ for sorting commits when writing\n> Bloom filters, for `git commit-graph write --reachable --changed-paths ...`\n> \n> Assuming that when writing the commit graph we always compute geeration\n> number v2 and 'generation' field stores corrected commit date, we don't\n> need to use date as a heuristic when generations are equal, and it would\n> not help in tie-breaking anyway.\n> \n> All right.\n> \n> >  \treturn 0;\n> >  }\n> >\n> > @@ -1342,10 +1337,14 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\t\t\tctx->commits.nr);\n> >  \tfor (i = 0; i < ctx->commits.nr; i++) {\n> >  \t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n> > +\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n> \n> All right, so the pattern is to add 'corrected_commit_date' stuff after\n> 'topological_level' stuff.\n> \n> >\n> >  \t\tdisplay_progress(ctx->progress, i + 1);\n> >  \t\tif (level != GENERATION_NUMBER_V1_INFINITY &&\n> > -\t\t    level != GENERATION_NUMBER_ZERO)\n> > +\t\t    level != GENERATION_NUMBER_ZERO &&\n> > +\t\t    corrected_commit_date != GENERATION_NUMBER_INFINITY &&\n> > +\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO\n> > +\t\t    )\n> >  \t\t\tcontinue;\n> >\n> >  \t\tcommit_list_insert(ctx->commits.list[i], &list);\n> > @@ -1354,17 +1353,26 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\tstruct commit_list *parent;\n> >  \t\t\tint all_parents_computed = 1;\n> >  \t\t\tuint32_t max_level = 0;\n> > +\t\t\ttimestamp_t max_corrected_commit_date = 0;\n> >\n> >  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n> >  \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n> > +\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n> >\n> >  \t\t\t\tif (level == GENERATION_NUMBER_V1_INFINITY ||\n> > -\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n> > +\t\t\t\t    level == GENERATION_NUMBER_ZERO ||\n> > +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_INFINITY ||\n> > +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO\n> > +\t\t\t\t    ) {\n> >  \t\t\t\t\tall_parents_computed = 0;\n> >  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n> >  \t\t\t\t\tbreak;\n> > -\t\t\t\t} else if (level > max_level) {\n> > -\t\t\t\t\tmax_level = level;\n> > +\t\t\t\t} else {\n> > +\t\t\t\t\tif (level > max_level)\n> > +\t\t\t\t\t\tmax_level = level;\n> > +\n> > +\t\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n> > +\t\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n> >  \t\t\t\t}\n> >  \t\t\t}\n> >\n> > @@ -1374,6 +1382,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n> >  \t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n> >  \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n> > +\n> > +\t\t\t\tif (current->date > max_corrected_commit_date)\n> > +\t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n> > +\t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n> >  \t\t\t}\n> >  \t\t}\n> >  \t}\n> \n> All right.  Looks good to me.\n> \n> > @@ -2372,8 +2384,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n> >  \tfor (i = 0; i < g->num_commits; i++) {\n> >  \t\tstruct commit *graph_commit, *odb_commit;\n> >  \t\tstruct commit_list *graph_parents, *odb_parents;\n> > -\t\ttimestamp_t max_generation = 0;\n> > -\t\ttimestamp_t generation;\n> > +\t\ttimestamp_t max_corrected_commit_date = 0;\n> > +\t\ttimestamp_t corrected_commit_date;\n> \n> This is simple, and perhaps unnecessary, rename of variables.\n> Shouldn't we however verify *both* topological level, and\n> (if exists) corrected commit date?\n\nThe problem with verifying both topological level and corrected commit\ndates is that we would have to re-fill commit_graph_data slab with commit\ndata chunk as we cannot modify data->generation otherwise, essentially\nrepeating the whole verification process.\n\nWhile it's okay for now, I might take this up in a future series [1].\n\n[1]: https://lore.kernel.org/git/4043ffbc-84df-0cd6-5c75-af80383a56cf@gmail.com/\n\n> \n> >\n> >  \t\tdisplay_progress(progress, i + 1);\n> >  \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n> > @@ -2412,9 +2424,9 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n> >  \t\t\t\t\t     oid_to_hex(&graph_parents->item->object.oid),\n> >  \t\t\t\t\t     oid_to_hex(&odb_parents->item->object.oid));\n> >\n> > -\t\t\tgeneration = commit_graph_generation(graph_parents->item);\n> > -\t\t\tif (generation > max_generation)\n> > -\t\t\t\tmax_generation = generation;\n> > +\t\t\tcorrected_commit_date = commit_graph_generation(graph_parents->item);\n> > +\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n> > +\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n> \n> Actually, commit_graph_generation(<commit>) can return either corrected\n> commit date, or topological level, the latter in backward-compatibility\n> case (if at least one commit-graph file is lacking GDAT chunk, because\n> [some of] it was created by the \"Old\" Git).\n> \n> >\n> >  \t\t\tgraph_parents = graph_parents->next;\n> >  \t\t\todb_parents = odb_parents->next;\n> > @@ -2436,20 +2448,12 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n> >  \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n> >  \t\t\tcontinue;\n> >\n> > -\t\t/*\n> > -\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n> > -\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n> > -\t\t * extra logic in the following condition.\n> > -\t\t */\n> > -\t\tif (max_generation == GENERATION_NUMBER_MAX)\n> > -\t\t\tmax_generation--;\n> \n> All right, this was needed for checking the correctness of topological\n> levels (generation number v1) because we were checking not that it\n> fullfills the reachability condition, but more strict one: namely that\n> topological level of commit is equal to maximum of topological levels of\n> its parents plus one.\n> \n> The comment about checking both generation number v1 and v2 still\n> applies.\n> \n> > -\n> > -\t\tgeneration = commit_graph_generation(graph_commit);\n> > -\t\tif (generation != max_generation + 1)\n> > -\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n> > +\t\tcorrected_commit_date = commit_graph_generation(graph_commit);\n> > +\t\tif (corrected_commit_date < max_corrected_commit_date + 1)\n> > +\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n> >  \t\t\t\t     oid_to_hex(&cur_oid),\n> > -\t\t\t\t     generation,\n> > -\t\t\t\t     max_generation + 1);\n> > +\t\t\t\t     corrected_commit_date,\n> > +\t\t\t\t     max_corrected_commit_date + 1);\n> \n> All right, we check less strict condition for corrected commit date.\n> \n> >\n> >  \t\tif (graph_commit->date != odb_commit->date)\n> >  \t\t\tgraph_report(_(\"commit date for commit %s in commit-graph is %\"PRItime\" != %\"PRItime),\n> \n> Best,\n> -- \n> Jakub Narębski\n"},{"id":"404393","messageId":"855z97dvsp.fsf@gmail.com","threadId":"53933","inReplyTo":"20200825061418.GA629699@Abhishek-Arch","subject":"Re: [PATCH v3 06/11] commit-graph: add a slab to store topological levels","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-25T07:33:26Z","receivedAt":"2020-08-25T07:33:32Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello,\n\nAbhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Fri, Aug 21, 2020 at 08:43:38PM +0200, Jakub Narębski wrote:\n>> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>\n>>> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n[...]\n>>> @@ -1335,11 +1341,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>>>  \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n>>>  \t\t\t\t\tctx->commits.nr);\n>>>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>>> -\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n>>> +\t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n>>\n>> All right, so that is why this 'generation' variable was not converted\n>> to timestamp_t type.\n>>\n>>>\n>>>  \t\tdisplay_progress(ctx->progress, i + 1);\n>>> -\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n>>> -\t\t    generation != GENERATION_NUMBER_ZERO)\n>>> +\t\tif (level != GENERATION_NUMBER_V1_INFINITY &&\n>>> +\t\t    level != GENERATION_NUMBER_ZERO)\n>>>  \t\t\tcontinue;\n>>\n>> Here we use GENERATION_NUMBER*_INFINITY to check if the commit is\n>> outside commit-graph files, and therefore we would need its topological\n>> level computed.\n>>\n>> However, I don't understand how it works.  We have had created the\n>> commit_graph_data_at() and use it instead of commit_graph_data_slab_at()\n>> to provide default values for `struct commit_graph`... but only for\n>> `graph_pos` member.  It is commit_graph_generation() that returns\n>> GENERATION_NUMBER_INFINITY for commits not in graph.\n>>\n>> But neither commit_graph_data_at()->generation nor topo_level_slab_at()\n>> handles this special case, so I don't see how 'generation' variable can\n>> *ever* be GENERATION_NUMBER_INFINITY, and 'level' variable can ever be\n>> GENERATION_NUMBER_V1_INFINITY for commits not in graph.\n>>\n>> Does it work *accidentally*, because the default value for uninitialized\n>> data on commit-slab is 0, which matches GENERATION_NUMBER_ZERO?  It\n>> certainly looks like it does.  And GENERATION_NUMBER_ZERO is an artifact\n>> of commit-graph feature development history, namely the short time where\n>> Git didn't use any generation numbers and stored 0 in the place set for\n>> it in the commit-graph format...  On the other hand this is not the case\n>> for corrected commit date (generation number v2), as it could\n>> \"legitimately\" be 0 if some root commit (without any parents) had\n>> committerdate of epoch 0, i.e. 1 January 1970 00:00:00 UTC, perhaps\n>> caused by malformed but valid commit object.\n>>\n>> Ugh...\n>\n> It works accidentally.\n>\n> Our decision to avoid the cost of initializing both\n> commit_graph_data->generation and commit_graph_data->graph_pos has\n> led to some unwieldy choices - the complexity of helper functions,\n> bypassing helper functions when writing a commit-graph file [1].\n>\n> I want to re-visit how commit_graph_data slab is defined in a future series.\n>\n> [1]: https://lore.kernel.org/git/be28ab7b-0ae4-2cc5-7f2b-92075de3723a@gmail.com/\n\nAll right, we might want to make use of the fact that the value of 0 for\ntopological level here always mean that its value for a commit needs to\nbe computed, that 0 is not a valid value for topological levels.\n- if the value 0 came from commit-graph file, it means that it came\n  from Git version that used commit-graph but didn't compute generation\n  numbers; the value is GENERATION_NUMBER_ZERO\n- the value 0 might came from the fact that commit is not in graph,\n  and that commit-slab zero-initializes the values stored; let's\n  call this value GENERATION_NUMBER_UNINITIALIZED\n\nIf we ensure that corrected commit date can never be zero (which is\nextremely unlikely, as one of root commits would have to be malformed or\nwritten on badly misconfigured computer, with value of 0 for committer\ntimestamp), then this \"happy accident\" can keep working.\n\n  As a special case, commit date with timestamp of zero (01.01.1970 00:00:00Z)\n  has corrected commit date of one, to be able to distinguish\n  uninitialized values.\n\nOr something like that.\n\nActually, it is not even necessary, as corrected commit date of 0 just\nmeans that this single value (well, for every root commit with commit\ndate of 0) would be unnecessary recomputed in compute_generation_numbers().\n\nAnyway, we would want to document this fact in the commit message.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404396","messageId":"CANQwDwdsV0mSos7M_d7UP1CjT1rCyA_GfaYarMKUZaFdDZ0WRg@mail.gmail.com","threadId":"53933","inReplyTo":"855z97dvsp.fsf@gmail.com","subject":"Re: [PATCH v3 06/11] commit-graph: add a slab to store topological levels","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-25T07:56:44Z","receivedAt":"2020-08-25T07:57:23Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, 25 Aug 2020 at 09:33, Jakub Narębski <jnareb@gmail.com> wrote:\n> Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> > On Fri, Aug 21, 2020 at 08:43:38PM +0200, Jakub Narębski wrote:\n> >> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> >>\n> >>> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> [...]\n> >>> @@ -1335,11 +1341,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >>>                                     _(\"Computing commit graph generation numbers\"),\n> >>>                                     ctx->commits.nr);\n> >>>     for (i = 0; i < ctx->commits.nr; i++) {\n> >>> -           uint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n> >>> +           uint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n> >>\n> >> All right, so that is why this 'generation' variable was not converted\n> >> to timestamp_t type.\n> >>\n> >>>\n> >>>             display_progress(ctx->progress, i + 1);\n> >>> -           if (generation != GENERATION_NUMBER_V1_INFINITY &&\n> >>> -               generation != GENERATION_NUMBER_ZERO)\n> >>> +           if (level != GENERATION_NUMBER_V1_INFINITY &&\n> >>> +               level != GENERATION_NUMBER_ZERO)\n> >>>                     continue;\n> >>\n> >> Here we use GENERATION_NUMBER*_INFINITY to check if the commit is\n> >> outside commit-graph files, and therefore we would need its topological\n> >> level computed.\n> >>\n> >> However, I don't understand how it works.  We have had created the\n> >> commit_graph_data_at() and use it instead of commit_graph_data_slab_at()\n> >> to provide default values for `struct commit_graph`... but only for\n> >> `graph_pos` member.  It is commit_graph_generation() that returns\n> >> GENERATION_NUMBER_INFINITY for commits not in graph.\n> >>\n> >> But neither commit_graph_data_at()->generation nor topo_level_slab_at()\n> >> handles this special case, so I don't see how 'generation' variable can\n> >> *ever* be GENERATION_NUMBER_INFINITY, and 'level' variable can ever be\n> >> GENERATION_NUMBER_V1_INFINITY for commits not in graph.\n> >>\n> >> Does it work *accidentally*, because the default value for uninitialized\n> >> data on commit-slab is 0, which matches GENERATION_NUMBER_ZERO?  It\n> >> certainly looks like it does.  And GENERATION_NUMBER_ZERO is an artifact\n> >> of commit-graph feature development history, namely the short time where\n> >> Git didn't use any generation numbers and stored 0 in the place set for\n> >> it in the commit-graph format...  On the other hand this is not the case\n> >> for corrected commit date (generation number v2), as it could\n> >> \"legitimately\" be 0 if some root commit (without any parents) had\n> >> committerdate of epoch 0, i.e. 1 January 1970 00:00:00 UTC, perhaps\n> >> caused by malformed but valid commit object.\n> >>\n> >> Ugh...\n> >\n> > It works accidentally.\n> >\n> > Our decision to avoid the cost of initializing both\n> > commit_graph_data->generation and commit_graph_data->graph_pos has\n> > led to some unwieldy choices - the complexity of helper functions,\n> > bypassing helper functions when writing a commit-graph file [1].\n> >\n> > I want to re-visit how commit_graph_data slab is defined in a future series.\n> >\n> > [1]: https://lore.kernel.org/git/be28ab7b-0ae4-2cc5-7f2b-92075de3723a@gmail.com/\n>\n> All right, we might want to make use of the fact that the value of 0 for\n> topological level here always mean that its value for a commit needs to\n> be computed, that 0 is not a valid value for topological levels.\n> - if the value 0 came from commit-graph file, it means that it came\n>   from Git version that used commit-graph but didn't compute generation\n>   numbers; the value is GENERATION_NUMBER_ZERO\n> - the value 0 might came from the fact that commit is not in graph,\n>   and that commit-slab zero-initializes the values stored; let's\n>   call this value GENERATION_NUMBER_UNINITIALIZED\n>\n> If we ensure that corrected commit date can never be zero (which is\n> extremely unlikely, as one of root commits would have to be malformed or\n> written on badly misconfigured computer, with value of 0 for committer\n> timestamp), then this \"happy accident\" can keep working.\n>\n>   As a special case, commit date with timestamp of zero (01.01.1970 00:00:00Z)\n>   has corrected commit date of one, to be able to distinguish\n>   uninitialized values.\n>\n> Or something like that.\n>\n> Actually, it is not even necessary, as corrected commit date of 0 just\n> means that this single value (well, for every root commit with commit\n> date of 0) would be unnecessary recomputed in compute_generation_numbers().\n>\n> Anyway, we would want to document this fact in the commit message.\n\nAlternatively, instead of comparing 'level' (and later in series also\n'corrected_commit_date') against GENERATION_NUMBER_INFINITY,\nwe could load at no extra cost `graph_pos` value and compare it\nagainst COMMIT_NOT_FROM_GRAPH.\n\nBut with this solution we could never get rid of graph_pos, if we\nthink it is unnecessary. If we split commit_graph_data into separate\nslabs (as it was in early versions of respective patch series), we\nwould have to pay additional cost.\n\nBut it is an alternative.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404397","messageId":"85wo1nca3u.fsf@gmail.com","threadId":"53933","inReplyTo":"20200825064954.GA645690@Abhishek-Arch","subject":"Re: [PATCH v3 07/11] commit-graph: implement corrected commit date","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-25T10:07:17Z","receivedAt":"2020-08-25T10:07:23Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello,\n\nAbhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Sat, Aug 22, 2020 at 02:05:41AM +0200, Jakub Narębski wrote:\n>> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>\n>>> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n[...]\n>>> To minimize the space required to store corrected commit date, Git\n>>> stores corrected commit date offsets into the commit-graph file. The\n>>> corrected commit date offset for a commit is defined as the difference\n>>> between its corrected commit date and actual commit date.\n>>\n>> Perhaps we should add more details about data type sizes in question.\n>\n> Will add.\n\nNote however that we need to solve the problem of storing values which\nare not monotonic wrt. parent relation (partial order) in limited disk\nspace, that is GENERATION_NUMBER_V2_OFFSET_MAX vs GENERATION_NUMBER_MAX;\nsee comments in 11/11 and 00/11.\n\n>>\n>> Storing corrected commit date requires sizeof(timestamp_t) bytes, which\n>> in most cases is 64 bits (uintmax_t).  However corrected commit date\n>> offsets can be safely stored^* using only 32 bits.  This halves the size\n>> of GDAT chunk, reducing per-commit storage from 2*H + 16 + 8 bytes to\n>> 2*H + 16 + 4 bytes, which is reduction of around 6%, not including\n>> header, fanout table (OIDF) and extra edges list (EDGE).\n>>\n>> Which might mean that the extra complication is not worth it, and we\n>> should store corrected commit date directly instead.\n>>\n>> *) unless for example one of commits is malformed but valid,\n>>    and has committerdate of 0 Unix time, 1 January 1970.\n\nSee above.\n\n>>> While Git does not write out offsets at this stage, Git stores the\n>>> corrected commit dates in member generation of struct commit_graph_data.\n>>> It will begin writing commit date offsets with the introduction of\n>>> generation data chunk.\n>>\n>> OK, so the agenda for introducing geeration number v2 is as follows:\n>> - compute generation numbers v2, i.e. corrected commit date\n>> - store corrected commit date [offsets] in new GDAT chunk,\n>>   unless backward-compatibility concerns require us to not to\n>> - load [and compute] corrected commit date from commit-graph\n>>   storing it as 'generation' field of `struct commit_graph_data`,\n>>   unless backward-compatibility concerns require us to store\n>>   topological levels (generation number v1) in there instead\n>>\n>\n> The last point is not correct. We always store topological levels into\n> the topo_levels slab introduced and always store corrected commit date\n> into data->generation, regardless of backward compatibility concerns.\n\nI think I was not clear enough (in trying to be brief).  I meant here\nloading available generation numbers for use in graph traversal,\ndone in later patches in this series.\n\nIn _next_ commit we store topological levels in `generation` field:\n\n  @@ -755,7 +763,11 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n   \tdate_low = get_be32(commit_data + g->hash_len + 12);\n   \titem->date = (timestamp_t)((date_high << 32) | date_low);\n\n  -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n  +\tif (g->chunk_generation_data)\n  +\t\tgraph_data->generation = item->date +\n  +\t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n  +\telse\n  +\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n\n\nWe use topo_level slab only when writing the commit-graph file.\n\n> We could avoid initializing topo_slab if we are not writing generation\n> data chunk (and thus don't need corrected commit dates) but that\n> wouldn't have an impact on run time while writing commit-graph because\n> computing corrected commit dates is cheap as the main cost is in walking\n> the graph and writing the file.\n\nRight.\n\nThough you need to add the cost of allocation and managing extra\ncommit slab, I think that amortized cost is negligible.\n\nBut what would be better is showing benchmark data: does writing the\ncommit graph without GDAT take not insigificant more time than without\nthis patch?\n\n[...]\n>>> @@ -2372,8 +2384,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n>>>  \tfor (i = 0; i < g->num_commits; i++) {\n>>>  \t\tstruct commit *graph_commit, *odb_commit;\n>>>  \t\tstruct commit_list *graph_parents, *odb_parents;\n>>> -\t\ttimestamp_t max_generation = 0;\n>>> -\t\ttimestamp_t generation;\n>>> +\t\ttimestamp_t max_corrected_commit_date = 0;\n>>> +\t\ttimestamp_t corrected_commit_date;\n>>\n>> This is simple, and perhaps unnecessary, rename of variables.\n>> Shouldn't we however verify *both* topological level, and\n>> (if exists) corrected commit date?\n>\n> The problem with verifying both topological level and corrected commit\n> dates is that we would have to re-fill commit_graph_data slab with commit\n> data chunk as we cannot modify data->generation otherwise, essentially\n> repeating the whole verification process.\n>\n> While it's okay for now, I might take this up in a future series [1].\n>\n> [1]: https://lore.kernel.org/git/4043ffbc-84df-0cd6-5c75-af80383a56cf@gmail.com/\n\nAll right, I believe you that verifying both topological level and\ncorrected commit date would be more difficult.\n\nThat doesn't change the conclusion that this variable should remain to\nbe named `generation`, as when verifying GDAT-less commit-graph files it\nwould check topological levels (it uses commit_graph_generation(), which\nin turn uses `generation` field in commit graph info, which as I have\nshow above in later patch could be v1 or v2 generation number).\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404399","messageId":"85mu2jc75c.fsf@gmail.com","threadId":"53933","inReplyTo":"20200821041124.GA39355@Abhishek-Arch","subject":"Re: [PATCH v3 03/11] commit-graph: consolidate fill_commit_graph_info","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-25T11:11:11Z","receivedAt":"2020-08-25T11:11:18Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello,\n\nAbhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Wed, Aug 19, 2020 at 07:54:20PM +0200, Jakub Narębski wrote:\n>> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>\n>>> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>>>\n>>> Both fill_commit_graph_info() and fill_commit_in_graph() parse\n>>> information present in commit data chunk. Let's simplify the\n>>> implementation by calling fill_commit_graph_info() within\n>>> fill_commit_in_graph().\n>>>\n>>> The test 'generate tar with future mtime' creates a commit with commit\n>>> time of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\n>>> generation number (within CDAT chunk) and has undefined behavior.\n>>>\n>>> The test used to pass\n>>\n>> Could you please tell us how does this test starts to fail without the\n>> change to the test described there?  What is the error message, etc.?\n>>\n>\n> Here's what I revised the commit message to:\n>\n> commit-graph: consolidate fill_commit_graph_info\n>\n> Both fill_commit_graph_info() and fill_commit_in_graph() parse\n> information present in commit data chunk. Let's simplify the\n> implementation by calling fill_commit_graph_info() within\n> fill_commit_in_graph().\n\nAll right.\n\nWe might want to add here the information that we also move loading the\ncommit date from the commit-graph file from fill_commit_in_graph() down\nthe [new] call chain into fill_commit_graph_info().  The commit date\nwould be needed in fill_commit_graph_info() in the next commit to\ncompute corrected commit date out of corrected commit date offset, and\nstore it as generation number.\n\n\nNOTE that this means that if we switch to storing 64-bit corrected\ncommit date directly in the commit-graph file, instead of storing 32-bit\noffsets, neither this Move Statement Into Function Out of Caller\nrefactoring nor change to the 'generate tar with future mtime' test\nwould be necessary.\n\n>\n> The test 'generate tar with future mtime' creates a commit with commit\n> time of (2 ^ 36 + 1) seconds since EPOCH. The CDAT chunk provides\n> 34-bits for storing commiter date, thus committer time overflows into\n> generation number (within CDAT chunk) and has undefined behavior.\n>\n> The test used to pass as fill_commit_graph_info() would not set struct\n> member `date` of struct commit and loads committer date from the object\n> database, generating a tar file with the expected mtime.\n\nI guess that in the case of generating a tar file we would read the\ncommit out of 'object database', and then only add commit-graph specific\ninfo with fill_commit_graph_info().  Possibly because we need more\ninformation that commit-graph provides for a commit.\n\n>\n> However, with corrected commit date, we will load the committer date\n> from CDAT chunk (truncated to lower 34-bits) to populate the generation\n> number. Thus, fill_commit_graph_info() sets date and generates tar file\n> with the truncated mtime and the test fails.\n>\n> Let's fix the test by setting a timestamp of (2 ^ 34 - 1) seconds, which\n> will not be truncated.\n\nNow I got interested why the value of (2 ^ 36 + 1) seconds since EPOCH\nwas used.\n\nThe commit that introduced the 'generate tar with future mtime' test,\nnamely e51217e15 (t5000: test tar files that overflow ustar headers,\n30-06-2016), says:\n\n\tThe ustar format only has room for 11 (or 12, depending on\n\tsome implementations) octal digits for the size and mtime of\n\teach file. For values larger than this, we have to add pax\n\textended headers to specify the real data, and git does not\n\tyet know how to do so.\n\n\tBefore fixing that, let's start off with some test\n\tinfrastructure [...]\n\nThe value of 2 ^ 36 equals 2 ^ 3*12 = (2 ^ 3) ^ 12 = 8 ^ 12.\nSo we need the value of (2 ^ 36 + 1) for this test do do its job.\nPossibly the value of 8 ^ 11 + 1 = 2 ^ 33 + 1 would be enough\n(if we skip testing \"some implementations\").\n\nSo I think to make this test more clear (for inquisitive minds) we\nshould set a timestamp of (2 ^ 33 + 1), not (2 ^ 34 - 1) seconds\nsince EPOCH.  Maybe even add a variant of this test that uses the\norigial value of (2 ^ 36 + 1) seconds since EPOCH, but turns off\nuse of serialized commit-graph.\n\nI'm sorry for not checking this earlier.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404408","messageId":"85ft8adilr.fsf@gmail.com","threadId":"53933","inReplyTo":"20200825050448.GA21012@Abhishek-Arch","subject":"Re: [PATCH v3 05/11] commit-graph: return 64-bit generation number","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-25T12:18:24Z","receivedAt":"2020-08-25T12:18:53Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Fri, Aug 21, 2020 at 03:14:34PM +0200, Jakub Narębski wrote:\n>> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>\n>>> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>>>\n>>> In a preparatory step, let's return timestamp_t values from\n>>> commit_graph_generation(), use timestamp_t for local variables\n>>\n>> All right, this is all good.\n>>\n>>> and define GENERATION_NUMBER_INFINITY as (2 ^ 63 - 1) instead.\n>>\n>> This needs more detailed examination.  There are two similar constants,\n>> GENERATION_NUMBER_INFINITY and GENERATION_NUMBER_MAX.  The former is\n>> used for newest commits outside the commit-graph, while the latter is\n>> maximum number that commits in the commit-graph can have (because of the\n>> storage limitations).  We therefore need GENERATION_NUMBER_INFINITY\n>> to be larger than GENERATION_NUMBER_MAX, and it is (and was).\n>>\n>> The GENERATION_NUMBER_INFINITY is because of the above requirement\n>> traditionally taken as maximum value that can be represented in the data\n>> type used to store commit's generation number _in memory_, but it can be\n>> less.  For timestamp_t the maximum value that can be represented\n>> is (2 ^ 63 - 1).\n>>\n>> All right then.\n>\n> Related to this, by the end of this series we are using\n> GENERATION_NUMBER_MAX in just one place - compute_generation_numbers()\n> to make sure the topological levels fit within 30 bits.\n>\n> Would it be more appropriate to rename GENERATION_NUMBER_MAX to\n> GENERATION_NUMBER_V1_MAX (along the lines of\n> GENERATION_NUMBER_V2_OFFSET_MAX)  to correctly describe that is a\n> limit on topological levels, rather than generation number value?\n\nYes, I think that at the end of this patch series we should be using\nGENERATION_NUMBER_V1_MAX and GENERATION_NUMBER_V2_OFFSET_MAX to describe\nstorage limits, and GENERATION_NUMBER_INFINITY (the latter as generation\nnumber value for commits not in graph).\n\nWe need to ensure that both GENERATION_NUMBER_V1_MAX and\nGENERATION_NUMBER_V2_OFFSET_MAX are smaller than\nGENERATION_NUMBER_INFINITY.\n\n\nHowever, as I wrote, handling GENERATION_NUMBER_V2_OFFSET_MAX is\ndifficult.  As far as I can see, we can choose one of the *three*\nsolutions (the third one is _new_):\n\na. store 64-bit corrected commit date in the GDAT chunk\n   all possible values are able to be stored, no need for\n   GENERATION_NUMBER_V2_MAX,\n\nb. store 32-bit corrected commit date offset in the GDAT chunk,\n   if its value is larger than GENERATION_NUMBER_V2_OFFSET_MAX,\n   do not write GDAT chunk at all (like for backward compatibility\n   with mixed-version chains of split commit-graph layers),\n\nc. store 32-bit corrected commit date offset in the GDAT chunk,\n   using some kind of overflow handling scheme; for example if\n   the most significant bit of 32-bit value is 1, then the\n   rest 31-bits are position in GDOV chunk, which uses 64-bit\n   to store those corrected commit date offsets that do not\n   fit in 32 bits.\n\nThis type of schema is used in other places in Git code, if I remember\nit correctly.\n\n>> The commit message says nothing about the new symbolic constant\n>> GENERATION_NUMBER_V1_INFINITY, though.\n>>\n>> I'm not sure it is even needed (see comments below).\n>\n> Yes, you are correct. I tried it out with your suggestions and it wasn't\n> really needed.\n>\n> Thanks for catching this!\n\nMistakes can happen when changig how the series is split into commits.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404502","messageId":"20200826071519.GA6805@Abhishek-Arch","threadId":"53933","inReplyTo":"85pn7ihabl.fsf@gmail.com","subject":"Re: [PATCH v3 09/11] commit-graph: use generation v2 only if entire chain does","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-08-26T07:15:19Z","receivedAt":"2020-08-26T07:17:39Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Sat, Aug 22, 2020 at 07:14:38PM +0200, Jakub Narębski wrote:\n> Hi Abhishek,\n> \n> ... \n> \n> However the commit message do not say anything about the *writing* side.\n> \n\nRevised the commit message to include the following at the end:\n\nWhen writing the new layer in split commit-graph, we write a GDAT chunk\nonly if the topmost layer has a GDAT chunk. This guarantees that if a\nlayer has GDAT chunk, all lower layers must have a GDAT chunk as well.\n\nRewriting layers follows similar approach: if the topmost layer below\nset of layers being rewritten (in the split commit-graph chain) exists,\nand it does not contain GDAT chunk, then the result of rewrite does not\nhave GDAT chunks either.\n\n> \n> ...\n> \n> To be more detailed, without '--split=replace' we would want the following\n> layer merging behavior:\n> \n>    [layer with GDAT][with GDAT][without GDAT][without GDAT][without GDAT]\n>            1              2           3             4            5\n> \n> In the split commit-graph chain above, merging two topmost layers\n> (layers 4 and 5) should create a layer without GDAT; merging three\n> topmost layers (and any other layers, e.g. two middle ones, i.e. 3 and\n> 4) should create a new layer with GDAT.\n> \n>    [layer with GDAT][with GDAT][without GDAT][-------without GDAT-------]\n>            1              2           3               merged\n> \n>    [layer with GDAT][with GDAT][-------------with GDAT------------------]\n>            1              2                    merged\n> \n> I hope those ASCII-art pictures help understanding it\n> \n\nThanks! There were helpful.\n\nWhile we work as expected in the first scenario i.e merging 4 and 5, we\nwould *still* write a layer without GDAT in the second scenario.\n\nI have tweaked split_graph_merge_strategy() to fix this:\n\n----------------------------------------------\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 6d54d9a286..246fad030d 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1973,6 +1973,9 @@ static void split_graph_merge_strategy(struct write_commit_graph_context *ctx)\n \t\t}\n \t}\n \n+\tif (!ctx->write_generation_data && g->chunk_generation_data)\n+\t\tctx->write_generation_data = 1;\n+\n \tif (flags != COMMIT_GRAPH_SPLIT_REPLACE)\n \t\tctx->new_base_graph = g;\n \telse if (ctx->num_commit_graphs_after != 1)\n\n----------------------------------------------------\n\nThat is, if we were not writing generation data (because of mixed\ngeneration concerns) but the new topmost layer has a generation data\nchunk, we have merged all layers without GDAT chunk and can now write a\nGDAT chunk safely.\n\n> >\n> > It is difficult to expose this issue in a test. Since we _start_ with\n> > artificially low generation numbers, any commit walk that prioritizes\n> > generation numbers will walk all of the commits with high generation\n> > number before walking the commits with low generation number. In all the\n> > cases I tried, the commit-graph layers themselves \"protect\" any\n> > incorrect behavior since none of the commits in the lower layer can\n> > reach the commits in the upper layer.\n> >\n> > This issue would manifest itself as a performance problem in this case,\n> > especially with something like \"git log --graph\" since the low\n> > generation numbers would cause the in-degree queue to walk all of the\n> > commits in the lower layer before allowing the topo-order queue to write\n> > anything to output (depending on the size of the upper layer).\n> \n> Wouldn't breaking the reachability condition promise make some Git\n> commands to return *incorrect* results if they short-circuit, stop\n> walking if generation number shows that A cannot reach B?\n> \n> I am talking here about commands that return boolean, or select subset\n> from given set of revisions:\n> - git merge-base --is-ancestor <B> <A>\n> - git branch branch-A <A> && git branch --contains <B>\n> - git branch branch-B <B> && git branch --merged <A>\n> \n> Git assumes that generation numbers fulfill the following condition:\n> \n>   if A can reach B, then gen(A) > gen(B)\n> \n> Notably this includes commits not in commit-graph, and clamped values.\n> \n> However, in the following case\n> \n> * if commit A is from higher layer without GDAT\n>   and uses topological levels for 'generation', e.g. 115 (in a small repo)\n> * and commit B is from lower layer with GDAT\n>   and uses corrected commit date as 'generation', for example 1598112896,\n> \n> it may happen that A (later commit) can reach B (earlier commit), but\n> gen(B) > gen(A).  The reachability condition promise for generation\n> numbers is broken.\n> \n> >\n> > Signed-off-by: Derrick Stolee <dstolee@microsoft.com>\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> \n> I have reordered files in the patch itself to make it easier to review\n> the proposed changes.\n> \n> >  commit-graph.h                |  1 +\n> >  commit-graph.c                | 32 +++++++++++++++-\n> >  t/t5324-split-commit-graph.sh | 70 +++++++++++++++++++++++++++++++++++\n> >  3 files changed, 102 insertions(+), 1 deletion(-)\n> >\n> > diff --git a/commit-graph.h b/commit-graph.h\n> > index f78c892fc0..3cf89d895d 100644\n> > --- a/commit-graph.h\n> > +++ b/commit-graph.h\n> > @@ -63,6 +63,7 @@ struct commit_graph {\n> >  \tstruct object_directory *odb;\n> >\n> >  \tuint32_t num_commits_in_base;\n> > +\tuint32_t read_generation_data;\n> >  \tstruct commit_graph *base_graph;\n> >\n> \n> First, why `read_generation_data` is of uint32_t type, when it stores\n> (as far as I understand it), a \"boolean\" value of either 0 or 1?\n> \n\nYes, using unsigned int instead of uint32_t (although in most of cases\nit would be same).  If commit_graph had other flags as well, we could\nhave used a bit field.\n\n> Second, couldn't we simply set chunk_generation_data to NULL?  Or would\n> that interfere with the case of rewriting, where we want to use existing\n> GDAT data when writing new commit-graph with GDAT chunk?\n\nIt interferes with rewriting the split commit-graph, as you might have\nguessed from the above code snippet.\n\n> \n> ...\n>\n> > diff --git a/commit-graph.c b/commit-graph.c\n> \n> >  \t\tgraph_data->generation = item->date +\n> >  \t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n> >  \telse\n> > @@ -885,6 +908,7 @@ void load_commit_graph_info(struct repository *r, struct commit *item)\n> >  \tuint32_t pos;\n> >  \tif (!prepare_commit_graph(r))\n> >  \t\treturn;\n> > +\n> >  \tif (find_commit_in_graph(item, r->objects->commit_graph, &pos))\n> >  \t\tfill_commit_graph_info(item, r->objects->commit_graph, pos);\n> >  }\n> \n> This is unrelated whitespace fix, a \"while at it\" in neighbourhood of\n> changes.  All right then.\n> \n\nReverted this change, as it's unimportant.\n\n> > @@ -2192,6 +2216,9 @@ int write_commit_graph(struct object_directory *odb,\n>\n> ...\n> \n> It would be nice to have an example with merging layers (whether we\n> would handle it in strict or relaxed way).\n> \n\nSure, will add.\n\n> > +\n> >  test_done\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"404506","messageId":"857dtld75f.fsf@gmail.com","threadId":"53933","inReplyTo":"20200826071519.GA6805@Abhishek-Arch","subject":"Re: [PATCH v3 09/11] commit-graph: use generation v2 only if entire chain does","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-26T10:38:04Z","receivedAt":"2020-08-26T10:38:13Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hi Abhishek,\n\nAbhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Sat, Aug 22, 2020 at 07:14:38PM +0200, Jakub Narębski wrote:\n>> Hi Abhishek,\n>>\n>> ...\n>>\n>> However the commit message do not say anything about the *writing* side.\n>>\n>\n> Revised the commit message to include the following at the end:\n>\n> When writing the new layer in split commit-graph, we write a GDAT chunk\n> only if the topmost layer has a GDAT chunk. This guarantees that if a\n> layer has GDAT chunk, all lower layers must have a GDAT chunk as well.\n>\n\nAll right.\n\n> Rewriting layers follows similar approach: if the topmost layer below\n> set of layers being rewritten (in the split commit-graph chain) exists,\n> and it does not contain GDAT chunk, then the result of rewrite does not\n> have GDAT chunks either.\n\nAll right.\n\nI see that you went with proposed more complex (but better) solution...\n\n>>\n>> ...\n>>\n>> To be more detailed, without '--split=replace' we would want the following\n>> layer merging behavior:\n>>\n>>    [layer with GDAT][with GDAT][without GDAT][without GDAT][without GDAT]\n>>            1              2           3             4            5\n>>\n>> In the split commit-graph chain above, merging two topmost layers\n>> (layers 4 and 5) should create a layer without GDAT; merging three\n>> topmost layers (and any other layers, e.g. two middle ones, i.e. 3 and\n>> 4) should create a new layer with GDAT.\n\nA simpler solution would be to create a new merged layer without GDAT if\nany of the layers being merged do not have GDAT.\n\nIn this solution merging 3+4+5, 3+4, and even 2+3 would result with\nlayer without GDAT, and only merging 1+2 would result in layer with GDAT.\n\n>>\n>>    [layer with GDAT][with GDAT][without GDAT][-------without GDAT-------]\n>>            1              2           3               merged\n>>\n>>    [layer with GDAT][with GDAT][-------------with GDAT------------------]\n>>            1              2                    merged\n>>\n>> I hope those ASCII-art pictures help understanding it\n>>\n>\n> Thanks! There were helpful.\n>\n> While we work as expected in the first scenario i.e merging 4 and 5, we\n> would *still* write a layer without GDAT in the second scenario.\n>\n> I have tweaked split_graph_merge_strategy() to fix this:\n>\n> ----------------------------------------------\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 6d54d9a286..246fad030d 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -1973,6 +1973,9 @@ static void split_graph_merge_strategy(struct write_commit_graph_context *ctx)\n>  \t\t}\n>  \t}\n>\n> +\tif (!ctx->write_generation_data && g->chunk_generation_data)\n> +\t\tctx->write_generation_data = 1;\n> +\n>  \tif (flags != COMMIT_GRAPH_SPLIT_REPLACE)\n>  \t\tctx->new_base_graph = g;\n>  \telse if (ctx->num_commit_graphs_after != 1)\n\n...which turned out to be not that complicated.  Nice work!\n\nThough this needs tests that if fulfills the stated condition (because I\nam not sure if it is entirely correct: we are not checking the layer\nbelow current one, isn't it?... ah, you explain it below).\n\nOne possible solution would be to grep `test-tool read-graph` output for\n\"^chunks: \", then pass it through `uniq` (without `sort`!), check that\nthe number of lines is less or equal 2, and if there are two lines then\ncheck that we get the following contents:\n\n  chunks: oid_fanout oid_lookup commit_metadata generation_data\n  chunks: oid_fanout oid_lookup commit_metadata\n\n(assuming that information about layers is added in top-down order).\n\nThis test must be run with GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0, which\nI think is the default.\n\n> ----------------------------------------------------\n>\n> That is, if we were not writing generation data (because of mixed\n> generation concerns) but the new topmost layer has a generation data\n> chunk, we have merged all layers without GDAT chunk and can now write a\n> GDAT chunk safely.\n\nAll right.\n\n[...]\n>>> diff --git a/commit-graph.h b/commit-graph.h\n>>> index f78c892fc0..3cf89d895d 100644\n>>> --- a/commit-graph.h\n>>> +++ b/commit-graph.h\n>>> @@ -63,6 +63,7 @@ struct commit_graph {\n>>>  \tstruct object_directory *odb;\n>>>\n>>>  \tuint32_t num_commits_in_base;\n>>> +\tuint32_t read_generation_data;\n>>>  \tstruct commit_graph *base_graph;\n>>>\n>>\n>> First, why `read_generation_data` is of uint32_t type, when it stores\n>> (as far as I understand it), a \"boolean\" value of either 0 or 1?\n>\n> Yes, using unsigned int instead of uint32_t (although in most of cases\n> it would be same).  If commit_graph had other flags as well, we could\n> have used a bit field.\n\nOK.\n\n>> Second, couldn't we simply set chunk_generation_data to NULL?  Or would\n>> that interfere with the case of rewriting, where we want to use existing\n>> GDAT data when writing new commit-graph with GDAT chunk?\n>\n> It interferes with rewriting the split commit-graph, as you might have\n> guessed from the above code snippet.\n\nAll right.\n\n[...]\n>>> @@ -885,6 +908,7 @@ void load_commit_graph_info(struct repository *r, struct commit *item)\n>>>  \tuint32_t pos;\n>>>  \tif (!prepare_commit_graph(r))\n>>>  \t\treturn;\n>>> +\n>>>  \tif (find_commit_in_graph(item, r->objects->commit_graph, &pos))\n>>>  \t\tfill_commit_graph_info(item, r->objects->commit_graph, pos);\n>>>  }\n>>\n>> This is unrelated whitespace fix, a \"while at it\" in neighbourhood of\n>> changes.  All right then.\n>>\n>\n> Reverted this change, as it's unimportant.\n\nActually I am not against fixing the whitespace in the neighbourhood of\nchanges, so you can keep it or revert it (discard).\n\n>>> @@ -2192,6 +2216,9 @@ int write_commit_graph(struct object_directory *odb,\n>>\n>> ...\n>>\n>> It would be nice to have an example with merging layers (whether we\n>> would handle it in strict or relaxed way).\n>>\n>\n> Sure, will add.\n\nThanks.\n\n\nBest,\n--\nJakub Narębski\n"},{"id":"404586","messageId":"20200827063951.GA16268@Abhishek-Arch","threadId":"53933","inReplyTo":"85y2m6fhkm.fsf@gmail.com","subject":"Re: [PATCH v3 11/11] doc: add corrected commit date info","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-08-27T06:39:51Z","receivedAt":"2020-08-27T06:42:19Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Sun, Aug 23, 2020 at 12:20:57AM +0200, Jakub Narębski wrote:\n> Hello,\n> \n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > With generation data chunk and corrected commit dates implemented, let's\n> > update the technical documentation for commit-graph.\n> >\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> All right.\n> \n> > ---\n> >  .../technical/commit-graph-format.txt         | 12 ++---\n> >  Documentation/technical/commit-graph.txt      | 45 ++++++++++++-------\n> >  2 files changed, 36 insertions(+), 21 deletions(-)\n> >\n> > diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n> > index 440541045d..71c43884ec 100644\n> > --- a/Documentation/technical/commit-graph-format.txt\n> > +++ b/Documentation/technical/commit-graph-format.txt\n> > @@ -4,11 +4,7 @@ Git commit graph format\n> >  The Git commit graph stores a list of commit OIDs and some associated\n> >  metadata, including:\n> >\n> > -- The generation number of the commit. Commits with no parents have\n> > -  generation number 1; commits with parents have generation number\n> > -  one more than the maximum generation number of its parents. We\n> > -  reserve zero as special, and can be used to mark a generation\n> > -  number invalid or as \"not computed\".\n> > +- The generation number of the commit.\n> \n> All right, that was duplicated information.  Now that we need to talk\n> about two of them, it would not make sense to duplicate that.\n> \n> >\n> >  - The root tree OID.\n> >\n> > @@ -88,6 +84,12 @@ CHUNK DATA:\n> \n> Shouldn't we also replace 'generation number' occurences in description\n> of the Commit Data (CDAT) chunk with either 'topological level' or\n> 'generation number v1'?\n\nYes, we should.\n\n> \n> >        2 bits of the lowest byte, storing the 33rd and 34th bit of the\n> >        commit time.\n> >\n> > +  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes) [Optional]\n> \n> It is not exactly 'optional', as it implies that we need to turn it on\n> (or that we can turn it off).  It is more 'conditional', as it can be\n> not present due to outside influences (mixed-version environment).\n> \n> > +    * This list of 4-byte values store corrected commit date offsets for the\n> > +      commits, arranged in the same order as commit data chunk.\n> \n> I have just realized purely theoretical, but possible, problem with\n> storing non-monotinic generation number related values like corrected\n> commit date offset in constrained space.  There are problems with\n> clamping them.\n> \n> Say that somewhere in the ancestry chain there is a commit A with commit\n> date far in the future by mistake, for example 2120-08-22; it is\n> important for that date to be not able to be represented using uint32_t.\n> Say that a later descendant commit B is malformed, and has committer\n> date of 0, that is 1970-01-01. This means that the corrected commit date\n> for B must be larger than 2120-08-22 - which for this commit means that\n> corrected commit date offset do not fit in 32 bits, and must be clamped\n> (replaced) with GENERATION_NUMBER_V2_OFFSET_MAX.\n> \n> Say that we have commit C that is child of B, and it has correct commit\n> date.  Because of mistake in commit A, it has corrected commit date of\n> more than 2120-08-22 (corrected commit date degenerated into topological\n> level plus constant).\n> \n> Now C can reach B, and B can reach A.  However, if we recover corrected\n> commit date of B out of its date=0 and offset=GENERATION_NUMBER_V2_OFFSET_MAX\n> we get a number that is smaller than correct corrected commit date.  We\n> will have\n> \n>    gen(A) > date(B) + offset(B) < gen(C)\n> \n> Which breaks reachability condition guarantee.\n> \n> If instead we use GENERATION_NUMBER_V2_MAX for commits with clamped\n> corrected commit date, that is offset=GENERATION_NUMBER_V2_OFFSET_MAX,\n> we would get\n> \n>   gen(A) < GENERATION_NUMBER_V2_MAX > gen(C)\n> \n> And again reachability condition is broken.\n> \n> This is a very contrived but possible example.  This shouldn't happen,\n> but ufortunately it can happen.\n> \n\nYes, that's very unfortunate. \n\nHere's a much simpler example:\n\nA commit P has an reasonable commit date (i.e. after release of Git to\npresent) D and has a child commit C with committer date 0. Now, the \ncorrected commiter date of C would D + 1 and the offset would be same too,\nas the committer date is zero. This overflows as reasonable dates are of\nthe order 2 ^ 34.\n\n> \n> The question is how to deal with this issue.  Ignore it as unlikely?\n> Switch to storing corrected commit date, which is monotonic, so if there\n> is commit with GENERATION_NUMBER_V2_MAX, then subsequent descendant\n> commits will also have GENERATION_NUMBER_V2_MAX -- and pay with up to 7%\n> larger commit-graph file?\n> \n\nTo be honest, I would prefer storing corrected committer dates over\nstoring offsets.\n\nWhile it is 7% of the size of commit-graph file, it is also *only* around\n~3.5 MB for a repository of the size of linux kernel (and IIRC\ncorrectly, the Windows repo has ~2M commits, it amounts to ~8 MB).\n\nMinimizing space and memory requirements are a top priority, but\nshouldn't making sure our program is correct and efficient to be a\ngreater priority?\n\nI would love to hear your and Dr. Stolee's opinions on this.\n\n> > +    * This list can be later modified to store future generation number related\n> > +      data.\n> \n> How can it be later modified?  There is no header, no version number.\n> How would we add another generation number data?\n> \n\nWe could modify the graph version in future. Here's how I think it would\nwork:\n\nGraph Version 1, No GDAT -> Topological level\nGraph Version 2, GDAT    -> Corrected committer dates\nGraph Version 3, GDAT    -> Generation number v3\n\nand so on.\n\nOf course, we do not have to update generation number definition for\neach graph version.\n\nHowever, my statement could still be wrong for things that we do not\nforesee (similar to how we missed the hard die on different graph version),\nso I am removing the statement.\n\n> > +\n> >    Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n> >        This list of 4-byte values store the second through nth parents for\n> >        all octopus merges. The second parent value in the commit data stores\n> > diff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\n> > index 808fa30b99..f27145328c 100644\n> > --- a/Documentation/technical/commit-graph.txt\n> > +++ b/Documentation/technical/commit-graph.txt\n> > @@ -38,14 +38,27 @@ A consumer may load the following info for a commit from the graph:\n> >\n> >  Values 1-4 satisfy the requirements of parse_commit_gently().\n> >\n> > -Define the \"generation number\" of a commit recursively as follows:\n> > +There are two definitions of generation number:\n> > +1. Corrected committer dates\n> > +2. Topological levels\n> \n> Should we add versioning info, that is:\n> \n>   +1. Corrected committer dates  (generation number v2)\n>   +2. Topological levels  (generation number v1)\n> \n\nYes, added.\n\n> > +\n> > +Define \"corrected committer date\" of a commit recursively as follows:\n> > +\n> > +  * A commit with no parents (a root commit) has corrected committer date\n> > +    equal to its committer date.\n> > +\n> > +  * A commit with at least one parent has corrected committer date equal to\n> > +    the maximum of its commiter date and one more than the largest corrected\n> > +    committer date among its parents.\n> > +\n> > +Define the \"topological level\" of a commit recursively as follows:\n> >\n> >   * A commit with no parents (a root commit) has generation number one.\n> \n> Shouldn't this be\n> \n>     * A commit with no parents (a root commit) has topological level of one.\n> \n\nThanks, fixed!\n\n> >\n> > - * A commit with at least one parent has generation number one more than\n> > -   the largest generation number among its parents.\n> > + * A commit with at least one parent has topological level one more than\n> > +   the largest topological level among its parents.\n> >\n> > -Equivalently, the generation number of a commit A is one more than the\n> > +Equivalently, the topological level of a commit A is one more than the\n> >  length of a longest path from A to a root commit. The recursive definition\n> >  is easier to use for computation and observing the following property:\n> \n> We should probably explicitly state that the property state applies to\n> both versions of generation number, not only to topological level.\n> \n> >\n> > @@ -67,17 +80,12 @@ numbers, the general heuristic is the following:\n> >      If A and B are commits with commit time X and Y, respectively, and\n> >      X < Y, then A _probably_ cannot reach B.\n> >\n> > -This heuristic is currently used whenever the computation is allowed to\n> > -violate topological relationships due to clock skew (such as \"git log\"\n> > -with default order), but is not used when the topological order is\n> > -required (such as merge base calculations, \"git log --graph\").\n> > -\n> \n> To be overly pedantic, this heuristic is still used, but now in much\n> more rare case.  In addition to what is stated above, at least one layer\n> in the split commit-graph chain must have been generated by \"Old\" Git,\n> for the date heuristic to be used.\n> \n> But that might be unnecessary level of detail.\n> \n> >  In practice, we expect some commits to be created recently and not stored\n> >  in the commit graph. We can treat these commits as having \"infinite\"\n> >  generation number and walk until reaching commits with known generation\n> >  number.\n> >\n> > -We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n> > +We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n> \n> All right.\n> \n> >  in the commit-graph file. If a commit-graph file was written by a version\n> >  of Git that did not compute generation numbers, then those commits will\n> >  have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n> > @@ -93,12 +101,11 @@ fully-computed generation numbers. Using strict inequality may result in\n> >  walking a few extra commits, but the simplicity in dealing with commits\n> >  with generation number *_INFINITY or *_ZERO is valuable.\n> >\n> > -We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n> > -generation numbers are computed to be at least this value. We limit at\n> > -this value since it is the largest value that can be stored in the\n> > -commit-graph file using the 30 bits available to generation numbers. This\n> > -presents another case where a commit can have generation number equal to\n> > -that of a parent.\n> > +We use the macro GENERATION_NUMBER_MAX for commits whose generation numbers\n> > +are computed to be at least this value. We limit at this value since it is\n> > +the largest value that can be stored in the commit-graph file using the\n> > +available to generation numbers. This presents another case where a\n> > +commit can have generation number equal to that of a parent.\n> \n> All right, though it could have been done without re-wrapping, so that\n> only first line would be marked as changed.\n> \n> As I wrote, there is theoretical problem with this for offsets.\n> \n> >\n> >  Design Details\n> >  --------------\n> > @@ -267,6 +274,12 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n> >  number of commits) could be extracted into config settings for full\n> >  flexibility.\n> >\n> > +We also merge commit-graph chains when we try to write a commit graph with\n> > +two different generation number definitions as they cannot be compared directly.\n> > +We overwrite the existing chain and create a commit-graph with the newer or more\n> > +efficient defintion. For example, overwriting topological levels commit graph\n> > +chain to create a corrected commit dates commit graph chain.\n> > +\n> \n> This is more complicated than that.\n> \n> I think we should explicitly state that Git ensures that in split\n> commit-graph chain, if there are layers without the GDAT chunk (that\n> force Git to use topological levels for generation numbers), then they\n> are top layers.  So if there is commit-graph file created by \"Old\" Git,\n> then when addig new layer it would also be GDAT-less.\n> \n> Now how to write this...\n\nThinking about this, I feel creating a new section called \"Handling\nMixed Generation Number Chains\" made more sense:\n\n  ## Handling Mixed Generation Number Chains\n\n  With the introduction of generation number v2 and generation data chunk,\n  the following scenario is possible:\n\n  1. \"New\" Git writes a commit-graph with a GDAT chunk.\n  2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n\n  The commits in the lower layer will be interpreted as having very large\n  generation values (commit date plus offset) compared to the generation\n  numbers in the top layer (toplogical level). This violates the\n  expectation that the generation of a parent is strictly smaller than the\n  generation of a child. In such cases, we revert to using topological\n  levels for all layers to maintain backwards compatability.\n\n  When writing a new layer in split commit-graph, we write a GDAT chunk\n  only if the topmost layer has a GDAT chunk. This guarantees that if a\n  lyer has GDAT chunk, all lower layers must have a GDAT chunk as well.\n\n  Rewriting layers follows similar approach: if the topmost layer below\n  set of layers being rewriteen (in the split commit-graph chain) exists,\n  and it does not contain GDAT chunk, then the result of rewrite does not\n  have GDAT chunks either.\n\n> \n> >  ## Deleting graph-{hash} files\n> >\n> >  After a new tip file is written, some `graph-{hash}` files may no longer\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"404601","messageId":"85o8mwb6nq.fsf@gmail.com","threadId":"53933","inReplyTo":"20200827063951.GA16268@Abhishek-Arch","subject":"Re: [PATCH v3 11/11] doc: add corrected commit date info","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-08-27T12:43:53Z","receivedAt":"2020-08-27T12:44:51Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello,\n\nAbhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Sun, Aug 23, 2020 at 12:20:57AM +0200, Jakub Narębski wrote:\n>> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n[...]\n\n>>> +    * This list of 4-byte values store corrected commit date offsets for the\n>>> +      commits, arranged in the same order as commit data chunk.\n>>\n>> I have just realized purely theoretical, but possible, problem with\n>> storing non-monotinic generation number related values like corrected\n>> commit date offset in constrained space.  There are problems with\n>> clamping them.\n>>\n>> Say that somewhere in the ancestry chain there is a commit A with commit\n>> date far in the future by mistake, for example 2120-08-22; it is\n>> important for that date to be not able to be represented using uint32_t.\n>> Say that a later descendant commit B is malformed, and has committer\n>> date of 0, that is 1970-01-01. This means that the corrected commit date\n>> for B must be larger than 2120-08-22 - which for this commit means that\n>> corrected commit date offset do not fit in 32 bits, and must be clamped\n>> (replaced) with GENERATION_NUMBER_V2_OFFSET_MAX.\n>>\n>> Say that we have commit C that is child of B, and it has correct commit\n>> date.  Because of mistake in commit A, it has corrected commit date of\n>> more than 2120-08-22 (corrected commit date degenerated into topological\n>> level plus constant).\n>>\n>> Now C can reach B, and B can reach A.  However, if we recover corrected\n>> commit date of B out of its date=0 and offset=GENERATION_NUMBER_V2_OFFSET_MAX\n>> we get a number that is smaller than correct corrected commit date.  We\n>> will have\n>>\n>>    gen(A) > date(B) + offset(B) < gen(C)\n>>\n>> Which breaks reachability condition guarantee.\n>>\n>> If instead we use GENERATION_NUMBER_V2_MAX for commits with clamped\n>> corrected commit date, that is offset=GENERATION_NUMBER_V2_OFFSET_MAX,\n>> we would get\n>>\n>>   gen(A) < GENERATION_NUMBER_V2_MAX > gen(C)\n>>\n>> And again reachability condition is broken.\n>>\n>> This is a very contrived but possible example.  This shouldn't happen,\n>> but ufortunately it can happen.\n>>\n>\n> Yes, that's very unfortunate.\n>\n> Here's a much simpler example:\n>\n> A commit P has an reasonable commit date (i.e. after release of Git to\n> present) D and has a child commit C with committer date 0. Now, the\n> corrected commiter date of C would D + 1 and the offset would be same too,\n> as the committer date is zero. This overflows as reasonable dates are of\n> the order 2 ^ 34.\n\nNo, we need the value of date D that doesn't fit in 2^32 _unsigned_ value,\nso it needs to be even more in the future than Y2k38 (2038-01-19 03:14:07),\nwhich is related to storing date as a _signed_ 32-bit integer\n\nThe current-ish Unix epoch time is 1598524281 - let's use it for value\nof D.  Then the offset for commit C would be 1598524282.  The current\nproposal uses 32 bits to store commit date offsets (as unsigned value).\nThe maximum value of offset that we can store is therefore 2^32 - 1,\nwhich is 4294967295.\n\n   corrected commit date offset(C) = 1,598,524,282\n   GENERATION_NUMBER_V2_MAX        = 4,294,967,295\n\nAs you can see there is no overflow in the simplified example.\n\n>>\n>> The question is how to deal with this issue.  Ignore it as unlikely?\n>> Switch to storing corrected commit date, which is monotonic, so if there\n>> is commit with GENERATION_NUMBER_V2_MAX, then subsequent descendant\n>> commits will also have GENERATION_NUMBER_V2_MAX -- and pay with up to 7%\n>> larger commit-graph file?\n>>\n>\n> To be honest, I would prefer storing corrected committer dates over\n> storing offsets.\n>\n> While it is 7% of the size of commit-graph file, it is also *only* around\n> ~3.5 MB for a repository of the size of linux kernel (and IIRC\n> correctly, the Windows repo has ~2M commits, it amounts to ~8 MB).\n\nIt is up to 7% of per-commit data, and it doesn't take into account EDGE\nchunk (for octopus merges), and it doesn't also take into account the\nsize of changed-paths Bloom filters data take in the commit-graph.\n\n> Minimizing space and memory requirements are a top priority, but\n> shouldn't making sure our program is correct and efficient to be a\n> greater priority?\n\nOn the other hand the case where we would encounter offsets that do not\nfit in uint32_t is extremply unlikely in sane repositories.\n\nI can think of three solutions:\n\n1. use 64-bit corrected commit dates\n   - advantages:\n     * simplest code,\n     * no need for overflow handling, as we can store all possible values\n       of timestamp_t\n   - disadvantages:\n     * commit-graph size increased by up to 7%\n\n2. use 32-bit corrected commit date offsets,\n   but simply do not store GDAT chunk if there is offset that would not\n   fit in 32-bit wide field\n   - advantages:\n     * commit-graph is smaller\n     * relatively simple overflow handling\n   - disadvantages:\n     * performance penalty (generation number v1 vs v2) for abnormal\n       repositories (with overflow not fitting in uint32_t)\n     * tests would be needed to exercise the overflow code\n\n3. use 32-bit for corrected commit date offset,\n   with oveflow handling, for example using most significant bit\n   to denote that other bits store position into offset overflow\n   with 64-bits for those offsets that do not fit in 31-bits\n   - advantages:\n     * commit-graph is smaller, increasing for abnormal repos\n   - disadvantages:\n     * most complex code of all proposed solutions\n     * smaller overflow limit of 2^31 - 1\n     * tests would be needed to exercise the overflow code\n\nI think because the situation where we encounter overflow in 32-bit\ncorrected commit date offset is rare, we should go with either 1 or 2\nsolution.\n\n> I would love to hear your and Dr. Stolee's opinions on this.\n\nI have CC-ed Junio C Hamano to ask for his opinion.\n\n>>> +    * This list can be later modified to store future generation number related\n>>> +      data.\n>>\n>> How can it be later modified?  There is no header, no version number.\n>> How would we add another generation number data?\n>>\n>\n> We could modify the graph version in future. Here's how I think it would\n> work:\n>\n> Graph Version 1, No GDAT -> Topological level\n> Graph Version 2, GDAT    -> Corrected committer dates\n> Graph Version 3, GDAT    -> Generation number v3\n>\n> and so on.\n>\n> Of course, we do not have to update generation number definition for\n> each graph version.\n\nSo it was about generic mechanism, not something specific to the GDAT chunk.\n\n> However, my statement could still be wrong for things that we do not\n> foresee (similar to how we missed the hard die on different graph version),\n> so I am removing the statement.\n\nGood.\n\n[...]\n>>> +We also merge commit-graph chains when we try to write a commit graph with\n>>> +two different generation number definitions as they cannot be compared directly.\n>>> +We overwrite the existing chain and create a commit-graph with the newer or more\n>>> +efficient defintion. For example, overwriting topological levels commit graph\n>>> +chain to create a corrected commit dates commit graph chain.\n>>> +\n>>\n>> This is more complicated than that.\n>>\n>> I think we should explicitly state that Git ensures that in split\n>> commit-graph chain, if there are layers without the GDAT chunk (that\n>> force Git to use topological levels for generation numbers), then they\n>> are top layers.  So if there is commit-graph file created by \"Old\" Git,\n>> then when addig new layer it would also be GDAT-less.\n>>\n>> Now how to write this...\n>\n> Thinking about this, I feel creating a new section called \"Handling\n> Mixed Generation Number Chains\" made more sense:\n>\n>   ## Handling Mixed Generation Number Chains\n>\n>   With the introduction of generation number v2 and generation data chunk,\n>   the following scenario is possible:\n>\n>   1. \"New\" Git writes a commit-graph with a GDAT chunk.\n>   2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n>\n>   The commits in the lower layer will be interpreted as having very large\n>   generation values (commit date plus offset) compared to the generation\n>   numbers in the top layer (toplogical level). This violates the\n>   expectation that the generation of a parent is strictly smaller than the\n>   generation of a child. In such cases, we revert to using topological\n>   levels for all layers to maintain backwards compatability.\n>\n>   When writing a new layer in split commit-graph, we write a GDAT chunk\n>   only if the topmost layer has a GDAT chunk. This guarantees that if a\n>   lyer has GDAT chunk, all lower layers must have a GDAT chunk as well.\n>\n>   Rewriting layers follows similar approach: if the topmost layer below\n>   set of layers being rewriteen (in the split commit-graph chain) exists,\n>   and it does not contain GDAT chunk, then the result of rewrite does not\n>   have GDAT chunks either.\n\nGood idea, and nice writeup.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404605","messageId":"e7bbce30-93a6-b7e2-844b-5f2af4dbddf3@gmail.com","threadId":"53933","inReplyTo":"20200827063951.GA16268@Abhishek-Arch","subject":"Re: [PATCH v3 11/11] doc: add corrected commit date info","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-27T13:15:56Z","receivedAt":"2020-08-27T14:53:54Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/27/2020 2:39 AM, Abhishek Kumar wrote:\n> Thinking about this, I feel creating a new section called \"Handling\n> Mixed Generation Number Chains\" made more sense:\n> \n>   ## Handling Mixed Generation Number Chains\n> \n>   With the introduction of generation number v2 and generation data chunk,\n>   the following scenario is possible:\n> \n>   1. \"New\" Git writes a commit-graph with a GDAT chunk.\n>   2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n\nI like the idea of this section, and this setup is good.\n\n>   The commits in the lower layer will be interpreted as having very large\n>   generation values (commit date plus offset) compared to the generation\n>   numbers in the top layer (toplogical level). This violates the\n>   expectation that the generation of a parent is strictly smaller than the\n>   generation of a child. In such cases, we revert to using topological\n>   levels for all layers to maintain backwards compatability.\n\ns/toplogical/topological\n\nBut also, we don't want to phrase this as \"in this case, we do the wrong\nthing\" but instead\n\n  A naive approach of using the newest available generation number from\n  each layer would lead to violated expectations: the lower layer would\n  use corrected commit dates which are much larger than the topological\n  levels of the higher layer. For this reason, Git inspects each layer\n  to see if any layer is missing corrected commit dates. In such a case,\n  Git only uses topological levels.\n\n>   When writing a new layer in split commit-graph, we write a GDAT chunk\n>   only if the topmost layer has a GDAT chunk. This guarantees that if a\n>   lyer has GDAT chunk, all lower layers must have a GDAT chunk as well.\n\ns/lyer/layer\n\nPerhaps leaving this at a higher level than referencing \"GDAT chunk\" is\nadvisable. Perhaps use \"we write corrected commit dates\" or \"all lower\nlayers must store corrected commit dates as well\", for example.\n\n>   Rewriting layers follows similar approach: if the topmost layer below\n>   set of layers being rewriteen (in the split commit-graph chain) exists,\n>   and it does not contain GDAT chunk, then the result of rewrite does not\n>   have GDAT chunks either.\n\nThis could use more positive language to make it clear that sometimes\nwe _do_ want to write corrected commit dates when merging layers:\n\n  When merging layers, we do not consider whether the merged layers had\n  corrected commit dates. Instead, the new layer will have corrected\n  commit dates if and only if all existing layers below the new layer\n  have corrected commit dates.\n\nThanks,\n-Stolee\n"},{"id":"404840","messageId":"20200901100828.GA10388@Abhishek-Arch","threadId":"53933","inReplyTo":"85imdah50e.fsf@gmail.com","subject":"Re: [PATCH v3 10/11] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-09-01T10:08:28Z","receivedAt":"2020-09-01T10:11:40Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Sat, Aug 22, 2020 at 09:09:21PM +0200, Jakub Narębski wrote:\n> Hello Abhishek,\n> \n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > With corrected commit dates implemented, we no longer have to rely on\n> > commit date as a heuristic in paint_down_to_common().\n> \n> All right, but it would be nice to have some benchmark data: what were\n> performance when using topological levels, what was performance when\n> using commit date heuristics (before this patch), what is performace now\n> when using corrected commit date.\n> \n> >\n> > t6024-recursive-merge setups a unique repository where all commits have\n> > the same committer date without well-defined merge-base. As this has\n> > already caused problems (as noted in 859fdc0 (commit-graph: define\n> > GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph within the\n> > test script.\n> \n> OK?\n\nIn hindsight, that is a terrible explanation. Here's what I have revised\nthis to:\n\n  With corrected commit dates implemented, we no longer have to rely on\n  commit date as a heuristic in paint_down_to_common().\n\n  While using corrected commit dates Git walks nearly the same number of\n  commits as commit date, the process is slower as for each comparision we\n  have to access the commit-slab (for corrected committer date) instead of\n  accessing struct member (for committer date).\n\n  For example, the command `git merge-base v4.8 v4.9` on the linux\n  repository walks 167468 commits, taking 0.135s for committer date and\n  167496 commits, taking 0.157s for corrected committer date respectively.\n\n  t6404-recursive-merge setups a unique repository where all commits have\n  the same committer date without well-defined merge-base. As this has\n  already caused problems (as noted in 859fdc0 (commit-graph: define\n  GIT_TEST_COMMIT_GRAPH, 2018-08-29)).\n\n  While running tests with GIT_TEST_COMMIT_GRAPH unset, we use committer\n  date as a heuristic in paint_down_to_common(). 6404.1 'combined merge\n  conflicts' merges commits in the order:\n  - Merge C with B to form a intermediate commit.\n  - Merge the intermediate commit with A.\n\n  With GIT_TEST_COMMIT_GRAPH=1, we write a commit-graph and subsequently\n  use the corrected committer date, which changes the order in which\n  commits are merged:\n  - Merge A with B to form a intermediate commit.\n  - Merge the intermediate commit with C.\n\n  While resulting repositories are equivalent, 6404.4 'virtual trees were\n  processed' fails with GIT_TEST_COMMIT_GRAPH=1 as we are selecting\n  different merge-bases and thus have different object ids for the\n  intermediate commits.\n\n  As this has already causes problems (as noted in 859fdc0 (commit-graph:\n  define GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph\n  within t6404-recursive-merge.\n> \n> >\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c             | 14 ++++++++++++++\n> >  commit-graph.h             |  6 ++++++\n> >  commit-reach.c             |  2 +-\n> >  t/t6024-recursive-merge.sh |  4 +++-\n> >  4 files changed, 24 insertions(+), 2 deletions(-)\n> >\n> \n> I have reorderd files for easier review.\n> \n> > diff --git a/commit-graph.h b/commit-graph.h\n> > index 3cf89d895d..e22ec1e626 100644\n> > --- a/commit-graph.h\n> > +++ b/commit-graph.h\n> > @@ -91,6 +91,12 @@ struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size);\n> >   */\n> >  int generation_numbers_enabled(struct repository *r);\n> >\n> > +/*\n> > + * Return 1 if and only if the repository has a commit-graph\n> > + * file and generation data chunk has been written for the file.\n> > + */\n> > +int corrected_commit_dates_enabled(struct repository *r);\n> > +\n> >  enum commit_graph_write_flags {\n> >  \tCOMMIT_GRAPH_WRITE_APPEND     = (1 << 0),\n> >  \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\n> \n> All right.\n> \n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index c1292f8e08..6411068411 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -703,6 +703,20 @@ int generation_numbers_enabled(struct repository *r)\n> >  \treturn !!first_generation;\n> >  }\n> >\n> > +int corrected_commit_dates_enabled(struct repository *r)\n> > +{\n> > +\tstruct commit_graph *g;\n> > +\tif (!prepare_commit_graph(r))\n> > +\t\treturn 0;\n> > +\n> > +\tg = r->objects->commit_graph;\n> > +\n> > +\tif (!g->num_commits)\n> > +\t\treturn 0;\n> > +\n> > +\treturn !!g->chunk_generation_data;\n> > +}\n> \n> The previous commit introduced validate_mixed_generation_chain(), which\n> walked whole split commit-graph chain, and set `read_generation_data`\n> field in `struct commit_graph` for all layers in the chain.\n> \n> This function examines only the top layer, so it follows the assumption\n> that Git would behave in such way that oly topmost layers in the chai\n> can be GDAT-less.\n> \n> Why the difference?  Couldn't validate_mixed_generation_chain() simply\n> call corrected_commit_dates_enabled()?\n\nThe previous commit didn't need to walk the whole split commit-graph\nchain. Because of how we are handling writing in a mixed generation data\nchunk, if a layer has generation data chunk, all layers below it have a\ngeneration data chunk as well.\n\nSo, there are two cases at hand:\n\n- Topmost layer has generation data chunk, so we know all layers below\n  it has generation data chunk and we can read values from it.\n- Topmost layer does not have generation data chunk, so we know we can't\n  read from generation data chunk.\n\nJust checking the topmost layer suffices - modified the previous commit.\n\nThen, this function is more or less the same as\n`g->read_generation_data` that is, if we are reading from generation\ndata chunk, we are using corrected commit dates.\n\n> \n> > +\n> >  static void close_commit_graph_one(struct commit_graph *g)\n> >  {\n> >  \tif (!g)\n> > diff --git a/commit-reach.c b/commit-reach.c\n> > index 470bc80139..3a1b925274 100644\n> > --- a/commit-reach.c\n> > +++ b/commit-reach.c\n> > @@ -39,7 +39,7 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n> >  \tint i;\n> >  \ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n> >\n> > -\tif (!min_generation)\n> \n> This check was added in 091f4cf (commit: don't use generation numbers if\n> not needed, 2018-08-30) by Derrick Stolee, and its commit message\n> includes benchmark results for running 'git merge-base v4.8 v4.9' in\n> Linux kernel repository:\n> \n>       v2.18.0: 0.122s    167,468 walked\n>   v2.19.0-rc1: 0.547s    635,579 walked\n>          HEAD: 0.127s\n> \n> > +\tif (!min_generation && !corrected_commit_dates_enabled(r))\n> >  \t\tqueue.compare = compare_commits_by_commit_date;\n> \n> It would be nice to have similar benchmark for this change... unless of\n> course there is no change in performance, but I think then it needs to\n> be stated explicitly.  I think.\n> \n\nMentioned in the commit message - we walk (nearly) the same number of\ncommits but take somewhat longer.\n\n> >\n> >  \tone->object.flags |= PARENT1;\n> > diff --git a/t/t6024-recursive-merge.sh b/t/t6024-recursive-merge.sh\n> > index 332cfc53fd..d3def66e7d 100755\n> > --- a/t/t6024-recursive-merge.sh\n> > +++ b/t/t6024-recursive-merge.sh\n> > @@ -15,6 +15,8 @@ GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n> >  export GIT_COMMITTER_DATE\n> >\n> >  test_expect_success 'setup tests' '\n> > +\tGIT_TEST_COMMIT_GRAPH=0 &&\n> > +\texport GIT_TEST_COMMIT_GRAPH &&\n> >  \techo 1 >a1 &&\n> >  \tgit add a1 &&\n> >  \tGIT_AUTHOR_DATE=\"2006-12-12 23:00:00\" git commit -m 1 a1 &&\n> > @@ -66,7 +68,7 @@ test_expect_success 'setup tests' '\n> >  '\n> >\n> >  test_expect_success 'combined merge conflicts' '\n> > -\ttest_must_fail env GIT_TEST_COMMIT_GRAPH=0 git merge -m final G\n> > +\ttest_must_fail git merge -m final G\n> >  '\n> >\n> >  test_expect_success 'result contains a conflict' '\n> \n> OK, so instead of disabling commit-graph for this test, now we disable\n> it for the whole script.\n> \n> Maybe this change should be in a separate patch?\n\nWith the explanation in commit message, it's clear to see how using\ncorrected commit dates leads to an (incorrectly) failing test. Does it\nstill make sense to seperate them?\n\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"404841","messageId":"20200901102624.GB10388@Abhishek-Arch","threadId":"53933","inReplyTo":"CANQwDwdsV0mSos7M_d7UP1CjT1rCyA_GfaYarMKUZaFdDZ0WRg@mail.gmail.com","subject":"Re: [PATCH v3 06/11] commit-graph: add a slab to store topological levels","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-09-01T10:26:24Z","receivedAt":"2020-09-01T10:28:48Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Aug 25, 2020 at 09:56:44AM +0200, Jakub Narębski wrote:\n> On Tue, 25 Aug 2020 at 09:33, Jakub Narębski <jnareb@gmail.com> wrote:\n>\n> ...\n>\n> >\n> > All right, we might want to make use of the fact that the value of 0 for\n> > topological level here always mean that its value for a commit needs to\n> > be computed, that 0 is not a valid value for topological levels.\n> > - if the value 0 came from commit-graph file, it means that it came\n> >   from Git version that used commit-graph but didn't compute generation\n> >   numbers; the value is GENERATION_NUMBER_ZERO\n> > - the value 0 might came from the fact that commit is not in graph,\n> >   and that commit-slab zero-initializes the values stored; let's\n> >   call this value GENERATION_NUMBER_UNINITIALIZED\n> >\n> > If we ensure that corrected commit date can never be zero (which is\n> > extremely unlikely, as one of root commits would have to be malformed or\n> > written on badly misconfigured computer, with value of 0 for committer\n> > timestamp), then this \"happy accident\" can keep working.\n> >\n> >   As a special case, commit date with timestamp of zero (01.01.1970 00:00:00Z)\n> >   has corrected commit date of one, to be able to distinguish\n> >   uninitialized values.\n> >\n> > Or something like that.\n> >\n> > Actually, it is not even necessary, as corrected commit date of 0 just\n> > means that this single value (well, for every root commit with commit\n> > date of 0) would be unnecessary recomputed in compute_generation_numbers().\n> >\n> > Anyway, we would want to document this fact in the commit message.\n> \n> Alternatively, instead of comparing 'level' (and later in series also\n> 'corrected_commit_date') against GENERATION_NUMBER_INFINITY,\n> we could load at no extra cost `graph_pos` value and compare it\n> against COMMIT_NOT_FROM_GRAPH.\n> \n> But with this solution we could never get rid of graph_pos, if we\n> think it is unnecessary. If we split commit_graph_data into separate\n> slabs (as it was in early versions of respective patch series), we\n> would have to pay additional cost.\n> \n> But it is an alternative.\n> \n> Best,\n> -- \n> Jakub Narębski\n\nI think updating a commit date with timestampt of zero to use corrected\ncommit date of one would leave us more options down the line.\n\nChanging this is easy enough.\n\nFor a root commit with timestamp zero, current->date would be zero and \nmax_corrected_commit_date would be zero as well. So we can set \ncorrected commit date as `max_corrected_commit_date + 1`, instead of the\nearlier `(current->date - 1) + 1`.\n\n----\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 7ed0a33ad6..e3c5e30405 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1389,7 +1389,7 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n \n-\t\t\t\tif (current->date > max_corrected_commit_date)\n+\t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n \t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n \t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n \t\t\t}\n"},{"id":"404854","messageId":"20200901110141.GC10388@Abhishek-Arch","threadId":"53933","inReplyTo":"85wo1nca3u.fsf@gmail.com","subject":"Re: [PATCH v3 07/11] commit-graph: implement corrected commit date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-09-01T11:01:41Z","receivedAt":"2020-09-01T11:13:47Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Aug 25, 2020 at 12:07:17PM +0200, Jakub Narębski wrote:\n> Hello,\n> \n> ...\n> \n> I think I was not clear enough (in trying to be brief).  I meant here\n> loading available generation numbers for use in graph traversal,\n> done in later patches in this series.\n> \n> In _next_ commit we store topological levels in `generation` field:\n> \n>   @@ -755,7 +763,11 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>    \tdate_low = get_be32(commit_data + g->hash_len + 12);\n>    \titem->date = (timestamp_t)((date_high << 32) | date_low);\n> \n>   -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n>   +\tif (g->chunk_generation_data)\n>   +\t\tgraph_data->generation = item->date +\n>   +\t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n>   +\telse\n>   +\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> \n> \n> We use topo_level slab only when writing the commit-graph file.\n> \n\nRight, I thought the agenda outlined points in the process of writing\ncommit-graph file.\n\n>\n> > We could avoid initializing topo_slab if we are not writing generation\n> > data chunk (and thus don't need corrected commit dates) but that\n> > wouldn't have an impact on run time while writing commit-graph because\n> > computing corrected commit dates is cheap as the main cost is in walking\n> > the graph and writing the file.\n> \n> Right.\n> \n> Though you need to add the cost of allocation and managing extra\n> commit slab, I think that amortized cost is negligible.\n> \n> But what would be better is showing benchmark data: does writing the\n> commit graph without GDAT take not insigificant more time than without\n> this patch?\n\nRight, we could compare time taken by master and series until (but not\nincluding this patcth) to write a commit-graph file. Will add.\n\n> \n> [...]\n> >>> @@ -2372,8 +2384,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n> >>>  \tfor (i = 0; i < g->num_commits; i++) {\n> >>>  \t\tstruct commit *graph_commit, *odb_commit;\n> >>>  \t\tstruct commit_list *graph_parents, *odb_parents;\n> >>> -\t\ttimestamp_t max_generation = 0;\n> >>> -\t\ttimestamp_t generation;\n> >>> +\t\ttimestamp_t max_corrected_commit_date = 0;\n> >>> +\t\ttimestamp_t corrected_commit_date;\n> >>\n> >> This is simple, and perhaps unnecessary, rename of variables.\n> >> Shouldn't we however verify *both* topological level, and\n> >> (if exists) corrected commit date?\n> >\n> > The problem with verifying both topological level and corrected commit\n> > dates is that we would have to re-fill commit_graph_data slab with commit\n> > data chunk as we cannot modify data->generation otherwise, essentially\n> > repeating the whole verification process.\n> >\n> > While it's okay for now, I might take this up in a future series [1].\n> >\n> > [1]: https://lore.kernel.org/git/4043ffbc-84df-0cd6-5c75-af80383a56cf@gmail.com/\n> \n> All right, I believe you that verifying both topological level and\n> corrected commit date would be more difficult.\n> \n> That doesn't change the conclusion that this variable should remain to\n> be named `generation`, as when verifying GDAT-less commit-graph files it\n> would check topological levels (it uses commit_graph_generation(), which\n> in turn uses `generation` field in commit graph info, which as I have\n> show above in later patch could be v1 or v2 generation number).\n> \n\nRight, I completely misunderstood you initially. Reverted the variable\nname changes.\n\n> Best,\n> -- \n> Jakub Narębski\n"},{"id":"404855","messageId":"20200901113524.GD10388@Abhishek-Arch","threadId":"53933","inReplyTo":"85mu2jc75c.fsf@gmail.com","subject":"Re: [PATCH v3 03/11] commit-graph: consolidate fill_commit_graph_info","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-09-01T11:35:24Z","receivedAt":"2020-09-01T11:42:30Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Aug 25, 2020 at 01:11:11PM +0200, Jakub Narębski wrote:\n> Hello,\n> \n> ...\n> \n> All right.\n> \n> We might want to add here the information that we also move loading the\n> commit date from the commit-graph file from fill_commit_in_graph() down\n> the [new] call chain into fill_commit_graph_info().  The commit date\n> would be needed in fill_commit_graph_info() in the next commit to\n> compute corrected commit date out of corrected commit date offset, and\n> store it as generation number.\n> \n> \n> NOTE that this means that if we switch to storing 64-bit corrected\n> commit date directly in the commit-graph file, instead of storing 32-bit\n> offsets, neither this Move Statement Into Function Out of Caller\n> refactoring nor change to the 'generate tar with future mtime' test\n> would be necessary.\n> \n> >\n> > The test 'generate tar with future mtime' creates a commit with commit\n> > time of (2 ^ 36 + 1) seconds since EPOCH. The CDAT chunk provides\n> > 34-bits for storing commiter date, thus committer time overflows into\n> > generation number (within CDAT chunk) and has undefined behavior.\n> >\n> > The test used to pass as fill_commit_graph_info() would not set struct\n> > member `date` of struct commit and loads committer date from the object\n> > database, generating a tar file with the expected mtime.\n> \n> I guess that in the case of generating a tar file we would read the\n> commit out of 'object database', and then only add commit-graph specific\n> info with fill_commit_graph_info().  Possibly because we need more\n> information that commit-graph provides for a commit.\n> \n> >\n> > However, with corrected commit date, we will load the committer date\n> > from CDAT chunk (truncated to lower 34-bits) to populate the generation\n> > number. Thus, fill_commit_graph_info() sets date and generates tar file\n> > with the truncated mtime and the test fails.\n> >\n> > Let's fix the test by setting a timestamp of (2 ^ 34 - 1) seconds, which\n> > will not be truncated.\n> \n> Now I got interested why the value of (2 ^ 36 + 1) seconds since EPOCH\n> was used.\n> \n> The commit that introduced the 'generate tar with future mtime' test,\n> namely e51217e15 (t5000: test tar files that overflow ustar headers,\n> 30-06-2016), says:\n> \n> \tThe ustar format only has room for 11 (or 12, depending on\n> \tsome implementations) octal digits for the size and mtime of\n> \teach file. For values larger than this, we have to add pax\n> \textended headers to specify the real data, and git does not\n> \tyet know how to do so.\n> \n> \tBefore fixing that, let's start off with some test\n> \tinfrastructure [...]\n> \n> The value of 2 ^ 36 equals 2 ^ 3*12 = (2 ^ 3) ^ 12 = 8 ^ 12.\n> So we need the value of (2 ^ 36 + 1) for this test do do its job.\n> Possibly the value of 8 ^ 11 + 1 = 2 ^ 33 + 1 would be enough\n> (if we skip testing \"some implementations\").\n> \n> So I think to make this test more clear (for inquisitive minds) we\n> should set a timestamp of (2 ^ 33 + 1), not (2 ^ 34 - 1) seconds\n> since EPOCH.  Maybe even add a variant of this test that uses the\n> origial value of (2 ^ 36 + 1) seconds since EPOCH, but turns off\n> use of serialized commit-graph.\n\nThat's pretty interesting! I didn't look into this either, will modify\nthe existing test and add a new test for it.\n\nThanks for investigating this further.\n\n> \n> I'm sorry for not checking this earlier.\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"404857","messageId":"20200901120653.GA59580@Abhishek-Arch","threadId":"53933","inReplyTo":"85ft8adilr.fsf@gmail.com","subject":"Re: [PATCH v3 05/11] commit-graph: return 64-bit generation number","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-09-01T12:06:53Z","receivedAt":"2020-09-01T12:46:26Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Aug 25, 2020 at 02:18:24PM +0200, Jakub Narębski wrote:\n> Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> \n> ...\n> \n> However, as I wrote, handling GENERATION_NUMBER_V2_OFFSET_MAX is\n> difficult.  As far as I can see, we can choose one of the *three*\n> solutions (the third one is _new_):\n> \n> a. store 64-bit corrected commit date in the GDAT chunk\n>    all possible values are able to be stored, no need for\n>    GENERATION_NUMBER_V2_MAX,\n> \n> b. store 32-bit corrected commit date offset in the GDAT chunk,\n>    if its value is larger than GENERATION_NUMBER_V2_OFFSET_MAX,\n>    do not write GDAT chunk at all (like for backward compatibility\n>    with mixed-version chains of split commit-graph layers),\n> \n> c. store 32-bit corrected commit date offset in the GDAT chunk,\n>    using some kind of overflow handling scheme; for example if\n>    the most significant bit of 32-bit value is 1, then the\n>    rest 31-bits are position in GDOV chunk, which uses 64-bit\n>    to store those corrected commit date offsets that do not\n>    fit in 32 bits.\n> \n\nAlright, so the third solution leverages the fact that in practice,\nvery few offsets would overflow the 32-bit limit. Using 64-bits for all\noffsets would be wasteful, we can trade off a miniscule amount of\ncomputation to save large amounts of disk space.\n\n>\n> This type of schema is used in other places in Git code, if I remember\n> it correctly.\n> \n\nYes, it's a similar idea to the extra edge list chunk, where the most\nsignificant bit of second parent indicates whether they are more than\ntwo parents.\n\nIt's definitely feasible, albeit a little complex.\n\nWhat's the overall consensus on the third solution?\n\n>\n> >> The commit message says nothing about the new symbolic constant\n> >> GENERATION_NUMBER_V1_INFINITY, though.\n> >>\n> >> I'm not sure it is even needed (see comments below).\n> >\n> > Yes, you are correct. I tried it out with your suggestions and it wasn't\n> > really needed.\n> >\n> > Thanks for catching this!\n> \n> Mistakes can happen when changig how the series is split into commits.\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"404859","messageId":"20200901130152.GA4186@Abhishek-Arch","threadId":"53933","inReplyTo":"e7bbce30-93a6-b7e2-844b-5f2af4dbddf3@gmail.com","subject":"Re: [PATCH v3 11/11] doc: add corrected commit date info","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-09-01T13:01:52Z","receivedAt":"2020-09-01T13:07:40Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Thu, Aug 27, 2020 at 09:15:56AM -0400, Derrick Stolee wrote:\n> On 8/27/2020 2:39 AM, Abhishek Kumar wrote:\n> > Thinking about this, I feel creating a new section called \"Handling\n> > Mixed Generation Number Chains\" made more sense:\n> > \n> >   ## Handling Mixed Generation Number Chains\n> > \n> >   With the introduction of generation number v2 and generation data chunk,\n> >   the following scenario is possible:\n> > \n> >   1. \"New\" Git writes a commit-graph with a GDAT chunk.\n> >   2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n> \n> I like the idea of this section, and this setup is good.\n> \n> >   The commits in the lower layer will be interpreted as having very large\n> >   generation values (commit date plus offset) compared to the generation\n> >   numbers in the top layer (toplogical level). This violates the\n> >   expectation that the generation of a parent is strictly smaller than the\n> >   generation of a child. In such cases, we revert to using topological\n> >   levels for all layers to maintain backwards compatability.\n> \n> s/toplogical/topological\n> \n> But also, we don't want to phrase this as \"in this case, we do the wrong\n> thing\" but instead\n> \n>   A naive approach of using the newest available generation number from\n>   each layer would lead to violated expectations: the lower layer would\n>   use corrected commit dates which are much larger than the topological\n>   levels of the higher layer. For this reason, Git inspects each layer\n>   to see if any layer is missing corrected commit dates. In such a case,\n>   Git only uses topological levels.\n> \n> >   When writing a new layer in split commit-graph, we write a GDAT chunk\n> >   only if the topmost layer has a GDAT chunk. This guarantees that if a\n> >   lyer has GDAT chunk, all lower layers must have a GDAT chunk as well.\n> \n> s/lyer/layer\n> \n> Perhaps leaving this at a higher level than referencing \"GDAT chunk\" is\n> advisable. Perhaps use \"we write corrected commit dates\" or \"all lower\n> layers must store corrected commit dates as well\", for example.\n> \n> >   Rewriting layers follows similar approach: if the topmost layer below\n> >   set of layers being rewriteen (in the split commit-graph chain) exists,\n> >   and it does not contain GDAT chunk, then the result of rewrite does not\n> >   have GDAT chunks either.\n> \n> This could use more positive language to make it clear that sometimes\n> we _do_ want to write corrected commit dates when merging layers:\n> \n>   When merging layers, we do not consider whether the merged layers had\n>   corrected commit dates. Instead, the new layer will have corrected\n>   commit dates if and only if all existing layers below the new layer\n>   have corrected commit dates.\n\nThanks, that is a great suggestion! Using positive language is more\nstraightforward and easier to understand.\n\n> \n> Thanks,\n> -Stolee\n"},{"id":"404961","messageId":"85imcvb4ag.fsf@gmail.com","threadId":"53933","inReplyTo":"20200901102624.GB10388@Abhishek-Arch","subject":"Re: [PATCH v3 06/11] commit-graph: add a slab to store topological levels","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-09-03T09:25:27Z","receivedAt":"2020-09-03T09:25:36Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Tue, Aug 25, 2020 at 09:56:44AM +0200, Jakub Narębski wrote:\n>> On Tue, 25 Aug 2020 at 09:33, Jakub Narębski <jnareb@gmail.com> wrote:\n>>\n>> ...\n>>\n>>>\n>>> All right, we might want to make use of the fact that the value of 0 for\n>>> topological level here always mean that its value for a commit needs to\n>>> be computed, that 0 is not a valid value for topological levels.\n>>> - if the value 0 came from commit-graph file, it means that it came\n>>>   from Git version that used commit-graph but didn't compute generation\n>>>   numbers; the value is GENERATION_NUMBER_ZERO\n>>> - the value 0 might came from the fact that commit is not in graph,\n>>>   and that commit-slab zero-initializes the values stored; let's\n>>>   call this value GENERATION_NUMBER_UNINITIALIZED\n>>>\n>>> If we ensure that corrected commit date can never be zero (which is\n>>> extremely unlikely, as one of root commits would have to be malformed or\n>>> written on badly misconfigured computer, with value of 0 for committer\n>>> timestamp), then this \"happy accident\" can keep working.\n>>>\n>>>   As a special case, commit date with timestamp of zero (01.01.1970 00:00:00Z)\n>>>   has corrected commit date of one, to be able to distinguish\n>>>   uninitialized values.\n>>>\n>>> Or something like that.\n>>>\n>>> Actually, it is not even necessary, as corrected commit date of 0 just\n>>> means that this single value (well, for every root commit with commit\n>>> date of 0) would be unnecessary recomputed in compute_generation_numbers().\n>>>\n>>> Anyway, we would want to document this fact in the commit message.\n>> \n>> Alternatively, instead of comparing 'level' (and later in series also\n>> 'corrected_commit_date') against GENERATION_NUMBER_INFINITY,\n>> we could load at no extra cost `graph_pos` value and compare it\n>> against COMMIT_NOT_FROM_GRAPH.\n>> \n>> But with this solution we could never get rid of graph_pos, if we\n>> think it is unnecessary. If we split commit_graph_data into separate\n>> slabs (as it was in early versions of respective patch series), we\n>> would have to pay additional cost.\n>> \n>> But it is an alternative.\n>\n> I think updating a commit date with timestampt of zero to use corrected\n> commit date of one would leave us more options down the line.\n>\n> Changing this is easy enough.\n>\n> For a root commit with timestamp zero, current->date would be zero and \n> max_corrected_commit_date would be zero as well. So we can set \n> corrected commit date as `max_corrected_commit_date + 1`, instead of the\n> earlier `(current->date - 1) + 1`.\n>\n> ----\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 7ed0a33ad6..e3c5e30405 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -1389,7 +1389,7 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n>  \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n>  \n> -\t\t\t\tif (current->date > max_corrected_commit_date)\n> +\t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n>  \t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n>  \t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n>  \t\t\t}\n\nIt turned out to be much easier than I have expected: a one-line change,\nadding simply a new condition.  Good work!\n\nPerhaps it would be better to write it as current->date == GENERATION_NUMBER_UNINITIALIZED\n(or *_ZERO, or *_NO_DATA,...), but current version is quite idiomatic\nand easy to read.\n\nWith this change we should, of course, also change the commit-graph\nformat docs.\n\n\nOn the other hand it is a bit unnecessary.  If `generation` is zero,\nusing it would still work, and it would just mean that it would be\nunnecessarily recomputed - but corrected commit date equal zero is\npossible only for root commits.\n\nBut the above solution is more consistent, using 0 to mark not\ninitialized values...  it is cleaner, at the cost of one more corner\ncase, single line change, and possibly an insignificant amount of\nperformance penalty due to adding unlikely true branch to the\nconditional.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404969","messageId":"851rjjasdo.fsf@gmail.com","threadId":"53933","inReplyTo":"20200901120653.GA59580@Abhishek-Arch","subject":"Re: [PATCH v3 05/11] commit-graph: return 64-bit generation number","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-09-03T13:42:43Z","receivedAt":"2020-09-03T14:58:55Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hi,\n\nAbhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Tue, Aug 25, 2020 at 02:18:24PM +0200, Jakub Narębski wrote:\n>> Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n>> \n>> ...\n>> \n>> However, as I wrote, handling GENERATION_NUMBER_V2_OFFSET_MAX is\n>> difficult.  As far as I can see, we can choose one of the *three*\n>> solutions (the third one is _new_):\n>> \n>> a. store 64-bit corrected commit date in the GDAT chunk\n>>    all possible values are able to be stored, no need for\n>>    GENERATION_NUMBER_V2_MAX,\n>> \n>> b. store 32-bit corrected commit date offset in the GDAT chunk,\n>>    if its value is larger than GENERATION_NUMBER_V2_OFFSET_MAX,\n>>    do not write GDAT chunk at all (like for backward compatibility\n>>    with mixed-version chains of split commit-graph layers),\n>> \n>> c. store 32-bit corrected commit date offset in the GDAT chunk,\n>>    using some kind of overflow handling scheme; for example if\n>>    the most significant bit of 32-bit value is 1, then the\n>>    rest 31-bits are position in GDOV chunk, which uses 64-bit\n>>    to store those corrected commit date offsets that do not\n>>    fit in 32 bits.\n\nNote that I have posted more detailed analysis of advantages and\ndisadvantages of each of the above solutions in response to 11/11\nhttps://public-inbox.org/git/85o8mwb6nq.fsf@gmail.com/\n\nI can think of yet another solution, a variant of approach 'c' with\ndifferent overflow handling scheme:\n\nc'.  Store 32-bit corrected commit date offset in the GDAT chunk,\n     using the following overflow handling scheme: if the value\n     is 0xFFFFFFFF (all bits set to 1, the maximum possible value\n     for uint32_t), then the corrected commit date or corrected\n     commit date offset can be found in GDOV chunk (Generation\n     Data OVerflow handling).\n\n     The GDOV chunk is composed of:\n     - H bytes of commit OID, or 4 bytes (32 bits) of commit pos\n     - 8 bytes (64 bits) of corrected commit date or its offset\n     \n     Commits in GDOV chunk are sorted; as we expect for the number\n     of commits that require GDOV to be zero or a very small number\n     there is no need for GDO Fanout chunk.\n     \n   - advantages:\n     * commit-graph is smaller, increasing for abnormal repos\n     * overflow limit reduced only by 1 (a single value)\n   - disadvantages:\n     * most complex code of all proposed solutions\n       even more complicated than for solution 'c',\n       different from EDGE chunk handling\n     * tests would be needed to exercise the overflow code\n\nOr we can split overflow handling into two chunks: GDOI (Generation Data\nOverflow Index) and GDOV, where GDOI would be composed of H bytes of\ncommit OID or 4 bytes of commit graph position (sorted), and GDOV would\nbe composed oly of 8 bytes (64 bits) of corrected commit date data.\n\nThis c'') variant has the same advantages and disadvantages as c'), with\nnegligibly slightly larger disk size and possibly slightly better\nperformance because of better data locality.\n\n>\n> Alright, so the third solution leverages the fact that in practice,\n> very few offsets would overflow the 32-bit limit. Using 64-bits for all\n> offsets would be wasteful, we can trade off a miniscule amount of\n> computation to save large amounts of disk space.\n\nOn the other hand we can say that we can trade negligible increase of\ncommit-graph disk space size (less than 7% in worst case: no octopus\nmerges, no changed-path Bloom filter data, using SHA-1 for object ids,\nlarge repository so header size + OIFD is negligible) for simpler code\nwith no need for overflow handling at all (and a minuscule amount of\nless computations).\n\n>>\n>> This type of schema is used in other places in Git code, if I remember\n>> it correctly.\n>> \n>\n> Yes, it's a similar idea to the extra edge list chunk, where the most\n> significant bit of second parent indicates whether they are more than\n> two parents.\n\nYes and no.  Yes, the solution 'c' uses exactly the same mechanism as\nthe pointer from Commid Data chunk into Extra Edges List chunk:\n\n      [...]  If there are more than two parents, the second value\n      has its most-significant bit on and the other bits store an array\n      position into the Extra Edge List chunk.\n\nOn the other hand we need to have some kind of overflow handling for the\nlist of parents, as the number of parents is not limited in Git (there\nis no technical upper limit on the number of parents a commits can\nhave), as opposed to for example Mercurial.  This is not the case for\nstoring corrected commit date (or corrected commit date offset), as 64\nbits is all we would ever need.\n\n> It's definitely feasible, albeit a little complex.\n>\n> What's the overall consensus on the third solution?\n\nStill waiting for others to weight in.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"404977","messageId":"85pn72ad4w.fsf@gmail.com","threadId":"53933","inReplyTo":"20200901100828.GA10388@Abhishek-Arch","subject":"Re: [PATCH v3 10/11] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-09-03T19:11:59Z","receivedAt":"2020-09-03T19:12:08Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello,\n\nAbhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Sat, Aug 22, 2020 at 09:09:21PM +0200, Jakub Narębski wrote:\n>> \n>> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>> \n>>> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>>>\n>>> With corrected commit dates implemented, we no longer have to rely on\n>>> commit date as a heuristic in paint_down_to_common().\n>> \n>> All right, but it would be nice to have some benchmark data: what were\n>> performance when using topological levels, what was performance when\n>> using commit date heuristics (before this patch), what is performace now\n>> when using corrected commit date.\n\nAll right, the new proposed commit message has this benchmark data.\nThanks.\n\n>>> t6024-recursive-merge setups a unique repository where all commits have\n>>> the same committer date without well-defined merge-base. As this has\n>>> already caused problems (as noted in 859fdc0 (commit-graph: define\n>>> GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph within the\n>>> test script.\n>> \n>> OK?\n>\n> In hindsight, that is a terrible explanation. Here's what I have revised\n> this to:\n>\n>   With corrected commit dates implemented, we no longer have to rely on\n>   commit date as a heuristic in paint_down_to_common().\n>\n>   While using corrected commit dates Git walks nearly the same number of\n>   commits as commit date, the process is slower as for each comparision we\n>   have to access the commit-slab (for corrected committer date) instead of\n>   accessing struct member (for committer date).\n>\n>   For example, the command `git merge-base v4.8 v4.9` on the linux\n>   repository walks 167468 commits, taking 0.135s for committer date and\n>   167496 commits, taking 0.157s for corrected committer date respectively.\n>\n>   t6404-recursive-merge setups a unique repository where all commits have\n>   the same committer date without well-defined merge-base. As this has\n>   already caused problems (as noted in 859fdc0 (commit-graph: define\n>   GIT_TEST_COMMIT_GRAPH, 2018-08-29)).\n>\n>   While running tests with GIT_TEST_COMMIT_GRAPH unset, we use committer\n>   date as a heuristic in paint_down_to_common(). 6404.1 'combined merge\n>   conflicts' merges commits in the order:\n>   - Merge C with B to form a intermediate commit.\n>   - Merge the intermediate commit with A.\n>\n>   With GIT_TEST_COMMIT_GRAPH=1, we write a commit-graph and subsequently\n>   use the corrected committer date, which changes the order in which\n>   commits are merged:\n>   - Merge A with B to form a intermediate commit.\n>   - Merge the intermediate commit with C.\n>\n>   While resulting repositories are equivalent, 6404.4 'virtual trees were\n>   processed' fails with GIT_TEST_COMMIT_GRAPH=1 as we are selecting\n>   different merge-bases and thus have different object ids for the\n>   intermediate commits.\n>\n>   As this has already causes problems (as noted in 859fdc0 (commit-graph:\n>   define GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph\n>   within t6404-recursive-merge.\n\nMuch better.  Thanks a lot.\n\n[...]\n>>> diff --git a/commit-graph.c b/commit-graph.c\n>>> index c1292f8e08..6411068411 100644\n>>> --- a/commit-graph.c\n>>> +++ b/commit-graph.c\n>>> @@ -703,6 +703,20 @@ int generation_numbers_enabled(struct repository *r)\n>>>  \treturn !!first_generation;\n>>>  }\n>>>\n>>> +int corrected_commit_dates_enabled(struct repository *r)\n>>> +{\n>>> +\tstruct commit_graph *g;\n>>> +\tif (!prepare_commit_graph(r))\n>>> +\t\treturn 0;\n>>> +\n>>> +\tg = r->objects->commit_graph;\n>>> +\n>>> +\tif (!g->num_commits)\n>>> +\t\treturn 0;\n>>> +\n>>> +\treturn !!g->chunk_generation_data;\n>>> +}\n>> \n>> The previous commit introduced validate_mixed_generation_chain(), which\n>> walked whole split commit-graph chain, and set `read_generation_data`\n>> field in `struct commit_graph` for all layers in the chain.\n>> \n>> This function examines only the top layer, so it follows the assumption\n>> that Git would behave in such way that oly topmost layers in the chai\n>> can be GDAT-less.\n>> \n>> Why the difference?  Couldn't validate_mixed_generation_chain() simply\n>> call corrected_commit_dates_enabled()?\n>\n> The previous commit didn't need to walk the whole split commit-graph\n> chain.\n\nErrr... but `validate_mixed_generation_chain()` introduced in previous\ncommit in this patch series *does* walk all the layers of the whole\nsplit commit-graph chain.\n\n\tstatic void validate_mixed_generation_chain(struct repository *r)\n\t{\n\t\tstruct commit_graph *g = r->objects->commit_graph;\n\t\tint read_generation_data = 1;\n\t\n\t\twhile (g) {\n\t\t\tif (!g->chunk_generation_data) {\n\t\t\t\tread_generation_data = 0;\n\t\t\t\tbreak;\n\t\t\t}\n\t\t\tg = g->base_graph;\n\t\t}\n\t\n\t\tg = r->objects->commit_graph;\n\t\n\t\twhile (g) {\n\t\t\tg->read_generation_data = read_generation_data;\n\t\t\tg = g->base_graph;\n\t\t}\n\t}\n\nMoreover it \"marks up\" the whole chain, actually walking it twice.\n\nYou wrote somewhere else (possibly after I wrote this post) that this\nwas needed to handle `git commit-graph validate`, if I remember it\ncorrectly.\n\nIf it is true, then we need both approaches: the less expensive one\n(relying on our assumptions) and the more expensive one.  But we need to\nbetter explain both: why we need more expensive one, why we can use the\nless expensive onne (how we ensure that the requirements are fulfilled).\n\n>       Because of how we are handling writing in a mixed generation data\n> chunk, if a layer has generation data chunk, all layers below it have a\n> generation data chunk as well.\n>\n> So, there are two cases at hand:\n>\n> - Topmost layer has generation data chunk, so we know all layers below\n>   it has generation data chunk and we can read values from it.\n> - Topmost layer does not have generation data chunk, so we know we can't\n>   read from generation data chunk.\n>\n> Just checking the topmost layer suffices - modified the previous commit.\n>\n> Then, this function is more or less the same as\n> `g->read_generation_data` that is, if we are reading from generation\n> data chunk, we are using corrected commit dates.\n\nAll right.  That explains how corrected_commit_dates_enabled() works,\nbut not why we need also validate_mixed_generation_chain() that sets\ng->read_generation_data for every layer in the chain.\n\n>> \n>>> +\n>>>  static void close_commit_graph_one(struct commit_graph *g)\n>>>  {\n>>>  \tif (!g)\n>>> diff --git a/commit-reach.c b/commit-reach.c\n>>> index 470bc80139..3a1b925274 100644\n>>> --- a/commit-reach.c\n>>> +++ b/commit-reach.c\n>>> @@ -39,7 +39,7 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n>>>  \tint i;\n>>>  \ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n>>>\n>>> -\tif (!min_generation)\n>> \n>> This check was added in 091f4cf (commit: don't use generation numbers if\n>> not needed, 2018-08-30) by Derrick Stolee, and its commit message\n>> includes benchmark results for running 'git merge-base v4.8 v4.9' in\n>> Linux kernel repository:\n>> \n>>       v2.18.0: 0.122s    167,468 walked\n>>   v2.19.0-rc1: 0.547s    635,579 walked\n>>          HEAD: 0.127s\n>> \n>>> +\tif (!min_generation && !corrected_commit_dates_enabled(r))\n>>>  \t\tqueue.compare = compare_commits_by_commit_date;\n>> \n>> It would be nice to have similar benchmark for this change... unless of\n>> course there is no change in performance, but I think then it needs to\n>> be stated explicitly.  I think.\n>> \n>\n> Mentioned in the commit message - we walk (nearly) the same number of\n> commits but take somewhat longer.\n\nAll right, the new proposed commit message has it.\n\nSidenote: this is outside of the scope of this patch series, but perhaps\nwe should think about bringing the `generation` field from the\ncommit-slab back as a member of the `struct commit`; this would need\nprofiling and benchmarking of the typical workload to get amortized\nperformance across many git commands.\n\n>>>\n>>>  \tone->object.flags |= PARENT1;\n>>> diff --git a/t/t6024-recursive-merge.sh b/t/t6024-recursive-merge.sh\n>>> index 332cfc53fd..d3def66e7d 100755\n>>> --- a/t/t6024-recursive-merge.sh\n>>> +++ b/t/t6024-recursive-merge.sh\n\nNote: this might be now t/t6404-recursive-merge.sh -- 6404 ot 6024.\n\n>>> @@ -15,6 +15,8 @@ GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n>>>  export GIT_COMMITTER_DATE\n>>>\n>>>  test_expect_success 'setup tests' '\n>>> +\tGIT_TEST_COMMIT_GRAPH=0 &&\n>>> +\texport GIT_TEST_COMMIT_GRAPH &&\n>>>  \techo 1 >a1 &&\n>>>  \tgit add a1 &&\n>>>  \tGIT_AUTHOR_DATE=\"2006-12-12 23:00:00\" git commit -m 1 a1 &&\n>>> @@ -66,7 +68,7 @@ test_expect_success 'setup tests' '\n>>>  '\n>>>\n>>>  test_expect_success 'combined merge conflicts' '\n>>> -\ttest_must_fail env GIT_TEST_COMMIT_GRAPH=0 git merge -m final G\n>>> +\ttest_must_fail git merge -m final G\n>>>  '\n>>>\n>>>  test_expect_success 'result contains a conflict' '\n>> \n>> OK, so instead of disabling commit-graph for this test, now we disable\n>> it for the whole script.\n>> \n>> Maybe this change should be in a separate patch?\n>\n> With the explanation in commit message, it's clear to see how using\n> corrected commit dates leads to an (incorrectly) failing test. Does it\n> still make sense to seperate them?\n\nNo, I think that the new commit message explains why those changes are\ntogether.\n\nOn the other hand it might be a good idea to add a TODO comment to this\ntest to mark it as fragile (fixing it is certainly out of scope of this\npatch series, but better have something to remind us about the issue).\nPerhaps:\n\n  # TODO: fragile test, relies on specific resolving of ambiguity\n\nOr something like that.  The original commit that added\nGIT_TEST_COMMIT_GRAPH=0 (for a single test) explained:\n\n  There is one test in t6024-recursive-merge.sh that relies on the\n  merge-base algorithm picking one of two ambiguous merge-bases, and\n  the commit-graph feature changes which merge-base is picked.\n\nI'm not sure of we could salvage some of this test as it is now adding\n`env GIT_TEST_COMMIT_GRAPH=0` in more individual tests instead of\nturning it off for the whole test script.  But that is something that we\ncan do later.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"405065","messageId":"20200905172127.GA1382@Abhishek-Arch","threadId":"53933","inReplyTo":"851rjjasdo.fsf@gmail.com","subject":"Re: [PATCH v3 05/11] commit-graph: return 64-bit generation number","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-09-05T17:21:27Z","receivedAt":"2020-09-05T17:23:59Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Thu, Sep 03, 2020 at 03:42:43PM +0200, Jakub Narębski wrote:\n> Hi,\n> \n> Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> > On Tue, Aug 25, 2020 at 02:18:24PM +0200, Jakub Narębski wrote:\n> >> Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> >> \n> >> ...\n> >> \n> >> However, as I wrote, handling GENERATION_NUMBER_V2_OFFSET_MAX is\n> >> difficult.  As far as I can see, we can choose one of the *three*\n> >> solutions (the third one is _new_):\n> >> \n> >> a. store 64-bit corrected commit date in the GDAT chunk\n> >>    all possible values are able to be stored, no need for\n> >>    GENERATION_NUMBER_V2_MAX,\n> >> \n> >> b. store 32-bit corrected commit date offset in the GDAT chunk,\n> >>    if its value is larger than GENERATION_NUMBER_V2_OFFSET_MAX,\n> >>    do not write GDAT chunk at all (like for backward compatibility\n> >>    with mixed-version chains of split commit-graph layers),\n> >> \n> >> c. store 32-bit corrected commit date offset in the GDAT chunk,\n> >>    using some kind of overflow handling scheme; for example if\n> >>    the most significant bit of 32-bit value is 1, then the\n> >>    rest 31-bits are position in GDOV chunk, which uses 64-bit\n> >>    to store those corrected commit date offsets that do not\n> >>    fit in 32 bits.\n> \n> Note that I have posted more detailed analysis of advantages and\n> disadvantages of each of the above solutions in response to 11/11\n> https://public-inbox.org/git/85o8mwb6nq.fsf@gmail.com/\n> \n> I can think of yet another solution, a variant of approach 'c' with\n> different overflow handling scheme:\n> \n> c'.  Store 32-bit corrected commit date offset in the GDAT chunk,\n>      using the following overflow handling scheme: if the value\n>      is 0xFFFFFFFF (all bits set to 1, the maximum possible value\n>      for uint32_t), then the corrected commit date or corrected\n>      commit date offset can be found in GDOV chunk (Generation\n>      Data OVerflow handling).\n> \n>      The GDOV chunk is composed of:\n>      - H bytes of commit OID, or 4 bytes (32 bits) of commit pos\n>      - 8 bytes (64 bits) of corrected commit date or its offset\n>      \n>      Commits in GDOV chunk are sorted; as we expect for the number\n>      of commits that require GDOV to be zero or a very small number\n>      there is no need for GDO Fanout chunk.\n>      \n>    - advantages:\n>      * commit-graph is smaller, increasing for abnormal repos\n>      * overflow limit reduced only by 1 (a single value)\n>    - disadvantages:\n>      * most complex code of all proposed solutions\n>        even more complicated than for solution 'c',\n>        different from EDGE chunk handling\n>      * tests would be needed to exercise the overflow code\n> \n> Or we can split overflow handling into two chunks: GDOI (Generation Data\n> Overflow Index) and GDOV, where GDOI would be composed of H bytes of\n> commit OID or 4 bytes of commit graph position (sorted), and GDOV would\n> be composed oly of 8 bytes (64 bits) of corrected commit date data.\n> \n> This c'') variant has the same advantages and disadvantages as c'), with\n> negligibly slightly larger disk size and possibly slightly better\n> performance because of better data locality.\n> \n\nThe primary benefit of c') over c) seems to the range of valid offsets -\nc') can range from [0, 0xFFFFFFFF) whereas offsets for c) can range\nbetwen [0, 0x7FFFFFF].\n\nIn other words, we should prefer c') over c) only if the offsets are\nusually in the range [0x7FFFFFFF + 1, 0xFFFFFFFF)\n\nCommits were overflowing corrected committer date offsets are rare, and\noffsets in that particular range are doubly rare. To be wrong within\nthe range would be have an offset of 68 to 136 years, so that's possible\nonly if the corrupted timestamp is in future (it's been 50.68 years since\nUnix epoch 0 so far).\n\nThikning back to the linux repository, the largest offset was around of\nthe order of 2 ^ 25 (offset of 1.06 years) and I would assume that holds\n\nOverall, I don't think the added complexity (compared to c) approach)\nmakes up for by greater versatility.\n\n[1]: https://lore.kernel.org/git/20200703082842.GA28027@Abhishek-Arch/\n\nThanks\n- Abhishek\n"},{"id":"405502","messageId":"85v9gh1yaz.fsf@gmail.com","threadId":"53933","inReplyTo":"20200905172127.GA1382@Abhishek-Arch","subject":"Re: [PATCH v3 05/11] commit-graph: return 64-bit generation number","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-09-13T15:39:00Z","receivedAt":"2020-09-13T15:39:09Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello,\n\nAbhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Thu, Sep 03, 2020 at 03:42:43PM +0200, Jakub Narębski wrote:\n>> Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n>>> On Tue, Aug 25, 2020 at 02:18:24PM +0200, Jakub Narębski wrote:\n>>>> Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n>>>>\n>>>> ...\n>>>>\n>>>> However, as I wrote, handling GENERATION_NUMBER_V2_OFFSET_MAX is\n>>>> difficult.  As far as I can see, we can choose one of the *three*\n>>>> solutions (the third one is _new_):\n>>>>\n>>>> a. store 64-bit corrected commit date in the GDAT chunk\n>>>>    all possible values are able to be stored, no need for\n>>>>    GENERATION_NUMBER_V2_MAX,\n>>>>\n>>>> b. store 32-bit corrected commit date offset in the GDAT chunk,\n>>>>    if its value is larger than GENERATION_NUMBER_V2_OFFSET_MAX,\n>>>>    do not write GDAT chunk at all (like for backward compatibility\n>>>>    with mixed-version chains of split commit-graph layers),\n>>>>\n>>>> c. store 32-bit corrected commit date offset in the GDAT chunk,\n>>>>    using some kind of overflow handling scheme; for example if\n>>>>    the most significant bit of 32-bit value is 1, then the\n>>>>    rest 31-bits are position in GDOV chunk, which uses 64-bit\n>>>>    to store those corrected commit date offsets that do not\n>>>>    fit in 32 bits.\n>>\n>> Note that I have posted more detailed analysis of advantages and\n>> disadvantages of each of the above solutions in response to 11/11\n>> https://public-inbox.org/git/85o8mwb6nq.fsf@gmail.com/\n>>\n>> I can think of yet another solution, a variant of approach 'c' with\n>> different overflow handling scheme:\n>>\n>> c'.  Store 32-bit corrected commit date offset in the GDAT chunk,\n>>      using the following overflow handling scheme: if the value\n>>      is 0xFFFFFFFF (all bits set to 1, the maximum possible value\n>>      for uint32_t), then the corrected commit date or corrected\n>>      commit date offset can be found in GDOV chunk (Generation\n>>      Data OVerflow handling).\n>>\n>>      The GDOV chunk is composed of:\n>>      - H bytes of commit OID, or 4 bytes (32 bits) of commit pos\n>>      - 8 bytes (64 bits) of corrected commit date or its offset\n>>\n>>      Commits in GDOV chunk are sorted; as we expect for the number\n>>      of commits that require GDOV to be zero or a very small number\n>>      there is no need for GDO Fanout chunk.\n>>\n>>    - advantages:\n>>      * commit-graph is smaller, increasing for abnormal repos\n>>      * overflow limit reduced only by 1 (a single value)\n>>    - disadvantages:\n>>      * most complex code of all proposed solutions\n>>        even more complicated than for solution 'c',\n>>        different from EDGE chunk handling\n>>      * tests would be needed to exercise the overflow code\n>>\n>> Or we can split overflow handling into two chunks: GDOI (Generation Data\n>> Overflow Index) and GDOV, where GDOI would be composed of H bytes of\n>> commit OID or 4 bytes of commit graph position (sorted), and GDOV would\n>> be composed oly of 8 bytes (64 bits) of corrected commit date data.\n>>\n>> This c'') variant has the same advantages and disadvantages as c'), with\n>> negligibly slightly larger disk size and possibly slightly better\n>> performance because of better data locality.\n>>\n>\n> The primary benefit of c') over c) seems to the range of valid offsets -\n> c') can range from [0, 0xFFFFFFFF) whereas offsets for c) can range\n> betwen [0, 0x7FFFFFF].\n>\n> In other words, we should prefer c') over c) only if the offsets are\n> usually in the range [0x7FFFFFFF + 1, 0xFFFFFFFF)\n>\n> Commits were overflowing corrected committer date offsets are rare, and\n> offsets in that particular range are doubly rare. To be wrong within\n> the range would be have an offset of 68 to 136 years, so that's possible\n> only if the corrupted timestamp is in future (it's been 50.68 years since\n> Unix epoch 0 so far).\n\nRight, the c) variant has the same limitation as if corrected commit\ndate offsets were stored as signed 32-bit integer (int32_t), so to have\noverflow we would have date post Y2k38 followed by date of Unix epoch 0.\nVery unlikely.\n\n>\n> Thinking back to the linux repository, the largest offset was around of\n> the order of 2 ^ 25 (offset of 1.06 years) and I would assume that holds\n>\n> Overall, I don't think the added complexity (compared to c) approach)\n> makes up for by greater versatility.\n>\n> [1]: https://lore.kernel.org/git/20200703082842.GA28027@Abhishek-Arch/\n\nAll right.\n\nWith variant c) we have additional advantage in that we can pattern the\ncode on the code for EDGE chunk handling, as you said.\n\nI wanted to warn about the need for sanity checking, like ensuring that\nwe have GDOV chunk and that it is large enough -- but it turns out that\nwe skip this bounds checking for extra edges / EDGE chunk:\n\n\tif (!(edge_value & GRAPH_EXTRA_EDGES_NEEDED)) {\n\t\tpptr = insert_parent_or_die(r, g, edge_value, pptr);\n\t\treturn 1;\n\t}\n\n\tparent_data_ptr = (uint32_t*)(g->chunk_extra_edges +\n\t\t\t  4 * (uint64_t)(edge_value & GRAPH_EDGE_LAST_MASK));\n\tdo {\n\t\tedge_value = get_be32(parent_data_ptr);\n\t\tpptr = insert_parent_or_die(r, g,\n\t\t\t\t\t    edge_value & GRAPH_EDGE_LAST_MASK,\n\t\t\t\t\t    pptr);\n\t\tparent_data_ptr++;\n\t} while (!(edge_value & GRAPH_LAST_EDGE));\n\nBest,\n-- \nJakub Narębski\n"},{"id":"406601","messageId":"857dsdvai2.fsf@gmail.com","threadId":"53933","inReplyTo":"85v9gh1yaz.fsf@gmail.com","subject":"Re: [PATCH v3 05/11] commit-graph: return 64-bit generation number","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-09-28T21:48:05Z","receivedAt":"2020-09-29T00:13:51Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello,\n\nJakub Narębski <jnareb@gmail.com> writes:\n> Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n>> On Thu, Sep 03, 2020 at 03:42:43PM +0200, Jakub Narębski wrote:\n>>> Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n>>>> On Tue, Aug 25, 2020 at 02:18:24PM +0200, Jakub Narębski wrote:\n>>>>> Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n>>>>>\n>>>>> ...\n>>>>>\n>>>>> However, as I wrote, handling GENERATION_NUMBER_V2_OFFSET_MAX is\n>>>>> difficult.  As far as I can see, we can choose one of the *three*\n>>>>> solutions (the third one is _new_):\n>>>>>\n>>>>> a. store 64-bit corrected commit date in the GDAT chunk\n>>>>>    all possible values are able to be stored, no need for\n>>>>>    GENERATION_NUMBER_V2_MAX,\n>>>>>\n>>>>> b. store 32-bit corrected commit date offset in the GDAT chunk,\n>>>>>    if its value is larger than GENERATION_NUMBER_V2_OFFSET_MAX,\n>>>>>    do not write GDAT chunk at all (like for backward compatibility\n>>>>>    with mixed-version chains of split commit-graph layers),\n>>>>>\n>>>>> c. store 32-bit corrected commit date offset in the GDAT chunk,\n>>>>>    using some kind of overflow handling scheme; for example if\n>>>>>    the most significant bit of 32-bit value is 1, then the\n>>>>>    rest 31-bits are position in GDOV chunk, which uses 64-bit\n>>>>>    to store those corrected commit date offsets that do not\n>>>>>    fit in 32 bits.\n>>>\n>>> Note that I have posted more detailed analysis of advantages and\n>>> disadvantages of each of the above solutions in response to 11/11\n>>> https://public-inbox.org/git/85o8mwb6nq.fsf@gmail.com/\n>>>\n>>> I can think of yet another solution, a variant of approach 'c'\n[...]\n>>\n>> The primary benefit of c') over c) seems to the range of valid offsets -\n>> c') can range from [0, 0xFFFFFFFF) whereas offsets for c) can range\n>> betwen [0, 0x7FFFFFF].\n>>\n>> In other words, we should prefer c') over c) only if the offsets are\n>> usually in the range [0x7FFFFFFF + 1, 0xFFFFFFFF)\n>>\n>> Commits were overflowing corrected committer date offsets are rare, and\n>> offsets in that particular range are doubly rare. To be wrong within\n>> the range would be have an offset of 68 to 136 years, so that's possible\n>> only if the corrupted timestamp is in future (it's been 50.68 years since\n>> Unix epoch 0 so far).\n>\n> Right, the c) variant has the same limitation as if corrected commit\n> date offsets were stored as signed 32-bit integer (int32_t), so to have\n> overflow we would have date post Y2k38 followed by date of Unix epoch 0.\n> Very unlikely.\n>\n>> Thinking back to the linux repository, the largest offset was around of\n>> the order of 2 ^ 25 (offset of 1.06 years) and I would assume that holds\n>>\n>> Overall, I don't think the added complexity (compared to c) approach)\n>> makes up for by greater versatility.\n>>\n>> [1]: https://lore.kernel.org/git/20200703082842.GA28027@Abhishek-Arch/\n>\n> All right.\n>\n> With variant c) we have additional advantage in that we can pattern the\n> code on the code for EDGE chunk handling, as you said.\n>\n> I wanted to warn about the need for sanity checking, like ensuring that\n> we have GDOV chunk and that it is large enough -- but it turns out that\n> we skip this bounds checking for extra edges / EDGE chunk:\n>\n> \tif (!(edge_value & GRAPH_EXTRA_EDGES_NEEDED)) {\n> \t\tpptr = insert_parent_or_die(r, g, edge_value, pptr);\n> \t\treturn 1;\n> \t}\n>\n> \tparent_data_ptr = (uint32_t*)(g->chunk_extra_edges +\n> \t\t\t  4 * (uint64_t)(edge_value & GRAPH_EDGE_LAST_MASK));\n> \tdo {\n> \t\tedge_value = get_be32(parent_data_ptr);\n> \t\tpptr = insert_parent_or_die(r, g,\n> \t\t\t\t\t    edge_value & GRAPH_EDGE_LAST_MASK,\n> \t\t\t\t\t    pptr);\n> \t\tparent_data_ptr++;\n> \t} while (!(edge_value & GRAPH_LAST_EDGE));\n\nBoth Taylor Blau and Junio C Hamano agree that it is better to store\ncorrected commit date offsets and have overload handling to halve (to\nreduce by 50%) the size of the GDAT chunk.  I have not heard from\nDerrick Stolee.\n\nIt looks then that it is the way to go; as I said that you have\nconvinced me that variant 'c' (EDGE-like) is the best solution for\noverflow handling.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"406892","messageId":"20201005052500.GA7276@Abhishek-Arch","threadId":"53933","inReplyTo":"857dsdvai2.fsf@gmail.com","subject":"Re: [PATCH v3 05/11] commit-graph: return 64-bit generation number","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-10-05T05:25:00Z","receivedAt":"2020-10-05T05:27:45Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Mon, Sep 28, 2020 at 11:48:05PM +0200, Jakub Narębski wrote:\n> Hello,\n> \n> \n> Both Taylor Blau and Junio C Hamano agree that it is better to store\n> corrected commit date offsets and have overload handling to halve (to\n> reduce by 50%) the size of the GDAT chunk.  I have not heard from\n> Derrick Stolee.\n> \n> It looks then that it is the way to go; as I said that you have\n> convinced me that variant 'c' (EDGE-like) is the best solution for\n> overflow handling.\n> \n\nGreat!\n\nSo I have implemented the variant 'c' and I am unsure whether my tests\nare exhaustive enough. Can you preview the commit \"commit-graph:\nimplement generation data chunk\" [1] on the pull request?.\n\nApart from that, I am ready to publish the v4 to the mailing list.\n\n[1]: https://github.com/gitgitgadget/git/pull/676/commits/390973da1d744cbb8a08a3b99c991f6d04ae9baf\n\n> Best,\n> -- \n> Jakub Narębski\n"},{"id":"407053","messageId":"4470d916428a28bb8277dfc4c3da84e08110e88e.1602079786.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"[PATCH v4 02/10] revision: parse parent in indegree_walk_step()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-10-07T14:09:37Z","receivedAt":"2020-10-07T14:09:53Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn indegree_walk_step(), we add unvisited parents to the indegree queue.\nHowever, parents are not guaranteed to be parsed. As the indegree queue\nsorts by generation number, let's parse parents before inserting them to\nensure the correct priority order.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n revision.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/revision.c b/revision.c\nindex aa62212040..c97abcdde1 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3381,6 +3381,9 @@ static void indegree_walk_step(struct rev_info *revs)\n \t\tstruct commit *parent = p->item;\n \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n \n+\t\tif (repo_parse_commit_gently(revs->repo, parent, 1) < 0)\n+\t\t\treturn;\n+\n \t\tif (*pi)\n \t\t\t(*pi)++;\n \t\telse\n-- \ngitgitgadget\n\n"},{"id":"407054","messageId":"fae81b534b14c8227454ff94e385fb16faee0e99.1602079785.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"[PATCH v4 01/10] commit-graph: fix regression when computing Bloom filters","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-10-07T14:09:36Z","receivedAt":"2020-10-07T14:09:53Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\ncommit_gen_cmp is used when writing a commit-graph to sort commits in\ngeneration order before computing Bloom filters. Since c49c82aa (commit:\nmove members graph_pos, generation to a slab, 2020-06-17) made it so\nthat 'commit_graph_generation()' returns 'GENERATION_NUMBER_INFINITY'\nduring writing, we cannot call it within this function. Instead, access\nthe generation number directly through the slab (i.e., by calling\n'commit_graph_data_at(c)->generation') in order to access it while\nwriting.\n\nWhile measuring performance with `git commit-graph write --reachable\n--changed-paths` on the linux repository led to around 1m40s for both\nHEAD and master (and could be due to fault in my measurements), it is\nstill the \"right\" thing to do.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex cb042bdba8..94503e584b 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n+\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n+\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"407056","messageId":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v3.git.1597509583.gitgitgadget@gmail.com","subject":"[PATCH v4 00/10] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-10-07T14:09:35Z","receivedAt":"2020-10-07T14:09:54Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"This patch series implements the corrected commit date offsets as generation\nnumber v2, along with other pre-requisites.\n\nGit uses topological levels in the commit-graph file for commit-graph\ntraversal operations like git log --graph. Unfortunately, using topological\nlevels can result in a worse performance than without them when compared\nwith committer date as a heuristics. For example, git merge-base v4.8 v4.9 \non the Linux repository walks 635,579 commits using topological levels and\nwalks 167,468 using committer date.\n\nThus, the need for generation number v2 was born. New generation number\nneeded to provide good performance, increment updates, and backward\ncompatibility. Due to an unfortunate problem 1\n[https://public-inbox.org/git/87a7gdspo4.fsf@evledraar.gmail.com/], we also\nneeded a way to distinguish between the old and new generation number\nwithout incrementing graph version.\n\nVarious candidates were examined (https://github.com/derrickstolee/gen-test, \nhttps://github.com/abhishekkumar2718/git/pull/1). The proposed generation\nnumber v2, Corrected Commit Date with Mononotically Increasing Offsets \nperformed much worse than committer date (506,577 vs. 167,468 commits walked\nfor git merge-base v4.8 v4.9) and was dropped.\n\nUsing Generation Data chunk (GDAT) relieves the requirement of backward\ncompatibility as we would continue to store topological levels in Commit\nData (CDAT) chunk. Thus, Corrected Commit Date was chosen as generation\nnumber v2. The Corrected Commit Date is defined as:\n\nFor a commit C, let its corrected commit date be the maximum of the commit\ndate of C and the corrected commit dates of its parents plus 1. Then \ncorrected commit date offset is the difference between corrected commit date\nof C and commit date of C. As a special case, a root commit with timestamp\nzero has corrected commit date of 1 to be able distinguish it from\nGENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit date).\n\nWe will introduce an additional commit-graph chunk, Generation Data chunk,\nand store corrected commit date offsets in GDAT chunk while storing\ntopological levels in CDAT chunk. The old versions of Git would ignore GDAT\nchunk, using topological levels from CDAT chunk. In contrast, new versions\nof Git would use corrected commit dates, falling back to topological level\nif the generation data chunk is absent in the commit-graph file.\n\nWhile storing corrected commit date offsets saves us 4 bytes per commit (as\ncompared with storing corrected commit dates directly), it's possible for\nthe offset to overflow the space allocated. To handle such cases, we\nintroduce a new chunk, Generation Data Overflow (GDOV) that stores the\ncorrected commit date. For overflowing offsets, we set MSB and store the\nposition into the GDOV chunk, in a mechanism similar to the Extra Edges list\nchunk.\n\nFor mixed generation number environment (for example new Git on the command\nline, old Git used by GUI client), we can encounter a mixed-chain\ncommit-graph (a commit-graph chain where some of split commit-graph files\nhave GDAT chunk and others do not). As backward compatibility is one of the\ngoals, we can define the following behavior:\n\nWhile reading a mixed-chain commit-graph version, we fall back on\ntopological levels as corrected commit dates and topological levels cannot\nbe compared directly.\n\nWhile writing on top of a split commit-graph, we check if the tip of the\nchain has a GDAT chunk. If it does, we append to the chain, writing GDAT\nchunk. Thus, we guarantee if the topmost split commit-graph file has a GDAT\nchunk, rest of the chain does too.\n\nIf the topmost split commit-graph file does not have a GDAT chunk (meaning\nit has been appended by the old Git), we write without GDAT chunk. We do\nwrite a GDAT chunk when the existing chain does not have GDAT chunk - when\nwe are writing to the commit-graph chain with the 'replace' strategy.\n\nThanks to Dr. Stolee, Dr. Narębski, and Taylor for their reviews.\n\nI look forward to everyone's reviews!\n\nThanks\n\n * Abhishek\n\n\n----------------------------------------------------------------------------\n\nChanges in version 4:\n\n * Added GDOV to handle overflows in generation data.\n * Added a test for writing tip graph for a generation number v2 graph chain\n   in t5324-split-commit-graph.sh\n * Added a section on how mixed generation number chains are handled in \n   Documentation/technical/commit-graph-format.txt\n * Reverted unimportant whitespace, style changes in commit-graph.c\n * Added header comments about the order of comparision for\n   compare_commits_by_gen_then_commit_date in commit.h,\n   compare_commits_by_gen in commit-graph.h\n * Elaborated on why t6404 fails with corrected commit date and must be run\n   with GIT_TEST_COMMIT_GRAPH=1in the commit \"commit-reach: use corrected\n   commit dates in paint_down_to_common()\"\n * Elaborated on write behavior for mixed generation number chains in the\n   commit \"commit-graph: use generation v2 only if entire chain does\"\n * Added notes about adding the topo_level slab to struct\n   write_commit_graph_context as well as struct commit_graph.\n * Clarified commit message for \"commit-graph: consolidate\n   fill_commit_graph_info\"\n * Removed the claim \"GDAT can store future generation numbers\" because it\n   hasn't been tested yet.\n\nChanges in version 3:\n\n * Reordered patches as discussed in 2\n   [https://lore.kernel.org/git/aee0ae56-3395-6848-d573-27a318d72755@gmail.com/]\n   .\n * Split \"implement corrected commit date\" into two patches - one\n   introducing the topo level slab and other implementing corrected commit\n   dates.\n * Extended split-commit-graph tests to verify at the end of test.\n * Use topological levels as generation number if any of split commit-graph\n   files do not have generation data chunk.\n\nChanges in version 2:\n\n * Add tests for generation data chunk.\n * Add an option GIT_TEST_COMMIT_GRAPH_NO_GDAT to control whether to write\n   generation data chunk.\n * Compare commits with corrected commit dates if present in\n   paint_down_to_common().\n * Update technical documentation.\n * Handle mixed generation commit chains.\n * Improve commit messages for \"commit-graph: fix regression when computing\n   bloom filter\", \"commit-graph: consolidate fill_commit_graph_info\",\n * Revert unnecessary whitespace changes.\n * Split uint_32 -> timestamp_t change into a new commit.\n\nAbhishek Kumar (10):\n  commit-graph: fix regression when computing Bloom filters\n  revision: parse parent in indegree_walk_step()\n  commit-graph: consolidate fill_commit_graph_info\n  commit-graph: return 64-bit generation number\n  commit-graph: add a slab to store topological levels\n  commit-graph: implement corrected commit date\n  commit-graph: implement generation data chunk\n  commit-graph: use generation v2 only if entire chain does\n  commit-reach: use corrected commit dates in paint_down_to_common()\n  doc: add corrected commit date info\n\n .../technical/commit-graph-format.txt         |  21 +-\n Documentation/technical/commit-graph.txt      |  62 ++++-\n commit-graph.c                                | 256 ++++++++++++++----\n commit-graph.h                                |  17 +-\n commit-reach.c                                |  38 +--\n commit-reach.h                                |   2 +-\n commit.c                                      |   4 +-\n commit.h                                      |   5 +-\n revision.c                                    |  13 +-\n t/README                                      |   3 +\n t/helper/test-read-graph.c                    |   4 +\n t/t4216-log-bloom.sh                          |   4 +-\n t/t5000-tar-tree.sh                           |  20 +-\n t/t5318-commit-graph.sh                       |  70 ++++-\n t/t5324-split-commit-graph.sh                 |  98 ++++++-\n t/t6404-recursive-merge.sh                    |   5 +-\n t/t6600-test-reach.sh                         |  68 ++---\n upload-pack.c                                 |   2 +-\n 18 files changed, 534 insertions(+), 158 deletions(-)\n\n\nbase-commit: d98273ba77e1ab9ec755576bc86c716a97bf59d7\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-676%2Fabhishekkumar2718%2Fcorrected_commit_date-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-676/abhishekkumar2718/corrected_commit_date-v4\nPull-Request: https://github.com/gitgitgadget/git/pull/676\n\nRange-diff vs v3:\n\n  1:  c6b7ade7af !  1:  fae81b534b commit-graph: fix regression when computing bloom filter\n     @@ Metadata\n      Author: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n       ## Commit message ##\n     -    commit-graph: fix regression when computing bloom filter\n     +    commit-graph: fix regression when computing Bloom filters\n      \n          commit_gen_cmp is used when writing a commit-graph to sort commits in\n          generation order before computing Bloom filters. Since c49c82aa (commit:\n     @@ Commit message\n          'commit_graph_data_at(c)->generation') in order to access it while\n          writing.\n      \n     +    While measuring performance with `git commit-graph write --reachable\n     +    --changed-paths` on the linux repository led to around 1m40s for both\n     +    HEAD and master (and could be due to fault in my measurements), it is\n     +    still the \"right\" thing to do.\n     +\n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n       ## commit-graph.c ##\n  2:  e673867234 !  2:  4470d91642 revision: parse parent in indegree_walk_step()\n     @@ revision.c: static void indegree_walk_step(struct rev_info *revs)\n       \t\tstruct commit *parent = p->item;\n       \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n       \n     -+\t\tif (parse_commit_gently(parent, 1) < 0)\n     ++\t\tif (repo_parse_commit_gently(revs->repo, parent, 1) < 0)\n      +\t\t\treturn;\n      +\n       \t\tif (*pi)\n  3:  18d5864f81 !  3:  18bb3318a1 commit-graph: consolidate fill_commit_graph_info\n     @@ Commit message\n          implementation by calling fill_commit_graph_info() within\n          fill_commit_in_graph().\n      \n     -    The test 'generate tar with future mtime' creates a commit with commit\n     -    time of (2 ^ 36 + 1) seconds since EPOCH. The commit time overflows into\n     -    generation number (within CDAT chunk) and has undefined behavior.\n     +    fill_commit_graph_info() used to not load committer data from commit data\n     +    chunk. However, with the corrected committer date, we have to load\n     +    committer date to calculate generation number value.\n      \n     -    The test used to pass as fill_commit_in_graph() guarantees the values of\n     -    graph position and generation number, and did not load timestamp.\n     -    However, with corrected commit date we will need load the timestamp as\n     -    well to populate the generation number.\n     +    e51217e15 (t5000: test tar files that overflow ustar headers,\n     +    30-06-2016) introduced a test 'generate tar with future mtime' that\n     +    creates a commit with committer date of (2 ^ 36 + 1) seconds since\n     +    EPOCH. The CDAT chunk provides 34-bits for storing committer date, thus\n     +    committer time overflows into generation number (within CDAT chunk) and\n     +    has undefined behavior.\n      \n     -    Let's fix the test by setting a timestamp of (2 ^ 34 - 1) seconds.\n     +    The test used to pass as fill_commit_graph_info() would not set struct\n     +    member `date` of struct commit and loads committer date from the object\n     +    database, generating a tar file with the expected mtime.\n     +\n     +    However, with corrected commit date, we will load the committer date\n     +    from CDAT chunk (truncated to lower 34-bits to populate the generation\n     +    number. Thus, Git sets date and generates tar file with the truncated\n     +    mtime.\n     +\n     +    The ustar format (the header format used by most modern tar programs)\n     +    only has room for 11 (or 12, depending om some implementations) octal\n     +    digits for the size and mtime of each files.\n     +\n     +    Thus, setting a timestamp of 2 ^ 33 + 1 would overflow the 11-octal\n     +    digit implementations while still fitting into commit data chunk.\n     +\n     +    Since we want to test 12-octal digit implementations of ustar as well,\n     +    let's modify the existing test to no longer use commit-graph file.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n     @@ commit-graph.c: static int fill_commit_in_graph(struct repository *r,\n      -\tgraph_data->graph_pos = pos;\n       \tlex_index = pos - g->num_commits_in_base;\n      -\n     --\tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n     -+\tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n     + \tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n       \n       \titem->object.parsed = 1;\n       \n     @@ commit-graph.c: static int fill_commit_in_graph(struct repository *r,\n       \tedge_value = get_be32(commit_data + g->hash_len);\n      \n       ## t/t5000-tar-tree.sh ##\n     -@@ t/t5000-tar-tree.sh: test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n     +@@ t/t5000-tar-tree.sh: test_expect_success TAR_HUGE,LONG_IS_64BIT 'system tar can read our huge size' '\n     + \ttest_cmp expect actual\n     + '\n     + \n     ++test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n     ++\trm -f .git/index &&\n     ++\techo foo >file &&\n     ++\tgit add file &&\n     ++\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n     ++\t\tgit commit -m \"tempori parendum\"\n     ++'\n     ++\n     ++test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n     ++\tgit archive HEAD >future.tar\n     ++'\n     ++\n     ++test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n     ++\techo 2514 >expect &&\n     ++\ttar_info future.tar | cut -d\" \" -f2 >actual &&\n     ++\ttest_cmp expect actual\n     ++'\n     ++\n     + test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n       \trm -f .git/index &&\n       \techo content >file &&\n       \tgit add file &&\n      -\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n     -+\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n     ++\tGIT_TEST_COMMIT_GRAPH=0 GIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n       \t\tgit commit -m \"tempori parendum\"\n       '\n       \n     -@@ t/t5000-tar-tree.sh: test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n     - '\n     - \n     - test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n     --\techo 4147 >expect &&\n     -+\techo 2514 >expect &&\n     - \ttar_info future.tar | cut -d\" \" -f2 >actual &&\n     - \ttest_cmp expect actual\n     - '\n  4:  6a0cde983d <  -:  ---------- commit-graph: consolidate compare_commits_by_gen\n  5:  6be759a954 !  4:  011b0aa497 commit-graph: return 64-bit generation number\n     @@ Commit message\n          commit_graph_generation(), use timestamp_t for local variables and\n          define GENERATION_NUMBER_INFINITY as (2 ^ 63 - 1) instead.\n      \n     +    We rename GENERATION_NUMBER_MAX to GENERATION_NUMBER_V1_MAX to\n     +    represent the largest topological level we can store in the commit data\n     +    chunk.\n     +\n     +    With corrected commit dates implemented, we will have two such *_MAX\n     +    variables to denote the largest offset and largest topological level\n     +    that can be stored.\n     +\n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n       ## commit-graph.c ##\n     @@ commit-graph.c: uint32_t commit_graph_position(const struct commit *c)\n       {\n       \tstruct commit_graph_data *data =\n       \t\tcommit_graph_data_slab_peek(&commit_graph_data_slab, c);\n     -@@ commit-graph.c: uint32_t commit_graph_generation(const struct commit *c)\n     - int compare_commits_by_gen(const void *_a, const void *_b)\n     - {\n     - \tconst struct commit *a = _a, *b = _b;\n     --\tconst uint32_t generation_a = commit_graph_generation(a);\n     --\tconst uint32_t generation_b = commit_graph_generation(b);\n     -+\tconst timestamp_t generation_a = commit_graph_generation(a);\n     -+\tconst timestamp_t generation_b = commit_graph_generation(b);\n     - \n     - \t/* older commits first */\n     - \tif (generation_a < generation_b)\n      @@ commit-graph.c: static int commit_gen_cmp(const void *va, const void *vb)\n       \tconst struct commit *a = *(const struct commit **)va;\n       \tconst struct commit *b = *(const struct commit **)vb;\n     @@ commit-graph.c: static int commit_gen_cmp(const void *va, const void *vb)\n       \tif (generation_a < generation_b)\n       \t\treturn -1;\n      @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n     - \t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n     + \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n     + \t\t\t\t\tctx->commits.nr);\n     + \tfor (i = 0; i < ctx->commits.nr; i++) {\n     +-\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n     ++\t\ttimestamp_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n       \n       \t\tdisplay_progress(ctx->progress, i + 1);\n     --\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n     -+\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n     - \t\t    generation != GENERATION_NUMBER_ZERO)\n     - \t\t\tcontinue;\n     - \n     + \t\tif (generation != GENERATION_NUMBER_INFINITY &&\n      @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n     - \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n     - \t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n     - \n     --\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n     -+\t\t\t\tif (generation == GENERATION_NUMBER_V1_INFINITY ||\n     - \t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n     - \t\t\t\t\tall_parents_computed = 0;\n     - \t\t\t\t\tcommit_list_insert(parent->item, &list);\n     + \t\t\t\tdata->generation = max_generation + 1;\n     + \t\t\t\tpop_commit(&list);\n     + \n     +-\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n     +-\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n     ++\t\t\t\tif (data->generation > GENERATION_NUMBER_V1_MAX)\n     ++\t\t\t\t\tdata->generation = GENERATION_NUMBER_V1_MAX;\n     + \t\t\t}\n     + \t\t}\n     + \t}\n      @@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n       \tfor (i = 0; i < g->num_commits; i++) {\n       \t\tstruct commit *graph_commit, *odb_commit;\n     @@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_grap\n       \n       \t\tdisplay_progress(progress, i + 1);\n       \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n     +@@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n     + \t\t\tcontinue;\n     + \n     + \t\t/*\n     +-\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n     +-\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n     ++\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n     ++\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n     + \t\t * extra logic in the following condition.\n     + \t\t */\n     +-\t\tif (max_generation == GENERATION_NUMBER_MAX)\n     ++\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n     + \t\t\tmax_generation--;\n     + \n     + \t\tgeneration = commit_graph_generation(graph_commit);\n      \n       ## commit-graph.h ##\n      @@ commit-graph.h: void disable_commit_graph(struct repository *r);\n     @@ commit-graph.h: void disable_commit_graph(struct repository *r);\n      -uint32_t commit_graph_generation(const struct commit *);\n      +timestamp_t commit_graph_generation(const struct commit *);\n       uint32_t commit_graph_position(const struct commit *);\n     - \n     - int compare_commits_by_gen(const void *_a, const void *_b);\n     + #endif\n      \n       ## commit-reach.c ##\n      @@ commit-reach.c: static int queue_has_nonstale(struct prio_queue *queue)\n     @@ commit-reach.c: int repo_in_merge_bases_many(struct repository *r, struct commit\n       {\n       \tstruct commit_list *bases;\n       \tint ret = 0, i;\n     --\tuint32_t generation, min_generation = GENERATION_NUMBER_INFINITY;\n     -+\ttimestamp_t generation, min_generation = GENERATION_NUMBER_INFINITY;\n     +-\tuint32_t generation, max_generation = GENERATION_NUMBER_ZERO;\n     ++\ttimestamp_t generation, max_generation = GENERATION_NUMBER_INFINITY;\n       \n       \tif (repo_parse_commit(r, commit))\n       \t\treturn ret;\n     @@ commit-reach.c: static enum contains_result contains_tag_algo(struct commit *can\n       \t\tstruct commit *c = p->item;\n       \t\tload_commit_graph_info(the_repository, c);\n       \t\tgeneration = commit_graph_generation(c);\n     +@@ commit-reach.c: static int compare_commits_by_gen(const void *_a, const void *_b)\n     + \tconst struct commit *a = *(const struct commit * const *)_a;\n     + \tconst struct commit *b = *(const struct commit * const *)_b;\n     + \n     +-\tuint32_t generation_a = commit_graph_generation(a);\n     +-\tuint32_t generation_b = commit_graph_generation(b);\n     ++\ttimestamp_t generation_a = commit_graph_generation(a);\n     ++\ttimestamp_t generation_b = commit_graph_generation(b);\n     + \n     + \tif (generation_a < generation_b)\n     + \t\treturn -1;\n      @@ commit-reach.c: int can_all_from_reach_with_flag(struct object_array *from,\n       \t\t\t\t unsigned int with_flag,\n       \t\t\t\t unsigned int assign_flag,\n     @@ commit-reach.h: int can_all_from_reach_with_flag(struct object_array *from,\n       \t\t       int commit_date_cutoff);\n       \n      \n     + ## commit.c ##\n     +@@ commit.c: int compare_commits_by_author_date(const void *a_, const void *b_,\n     + int compare_commits_by_gen_then_commit_date(const void *a_, const void *b_, void *unused)\n     + {\n     + \tconst struct commit *a = a_, *b = b_;\n     +-\tconst uint32_t generation_a = commit_graph_generation(a),\n     +-\t\t       generation_b = commit_graph_generation(b);\n     ++\tconst timestamp_t generation_a = commit_graph_generation(a),\n     ++\t\t\t  generation_b = commit_graph_generation(b);\n     + \n     + \t/* newer commits first */\n     + \tif (generation_a < generation_b)\n     +\n       ## commit.h ##\n      @@\n       #include \"commit-slab.h\"\n       \n       #define COMMIT_NOT_FROM_GRAPH 0xFFFFFFFF\n      -#define GENERATION_NUMBER_INFINITY 0xFFFFFFFF\n     +-#define GENERATION_NUMBER_MAX 0x3FFFFFFF\n      +#define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n     -+#define GENERATION_NUMBER_V1_INFINITY 0xFFFFFFFF\n     - #define GENERATION_NUMBER_MAX 0x3FFFFFFF\n     ++#define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n       #define GENERATION_NUMBER_ZERO 0\n       \n     + struct commit_list {\n      \n       ## revision.c ##\n      @@ revision.c: define_commit_slab(indegree_slab, int);\n     @@ revision.c: static void init_topo_walk(struct rev_info *revs)\n      -\t\tuint32_t generation;\n      +\t\ttimestamp_t generation;\n       \n     - \t\tif (parse_commit_gently(c, 1))\n     + \t\tif (repo_parse_commit_gently(revs->repo, c, 1))\n       \t\t\tcontinue;\n      @@ revision.c: static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n       \tfor (p = commit->parents; p; p = p->next) {\n  6:  b347dbb01b !  5:  e067f653ad commit-graph: add a slab to store topological levels\n     @@ Metadata\n       ## Commit message ##\n          commit-graph: add a slab to store topological levels\n      \n     -    As we are writing topological levels to commit data chunk to ensure\n     -    backwards compatibility with \"Old\" Git and the member `generation` of\n     -    struct commit_graph_data will store corrected commit date in a later\n     -    commit, let's introduce a commit-slab to store topological levels while\n     -    writing commit-graph.\n     +    In a later commit we will introduce corrected commit date as the\n     +    generation number v2. This value will be stored in the new seperate\n     +    Generation Data chunk. However, to ensure backwards compatibility with\n     +    \"Old\" Git we need to continue to write generation number v1, which is\n     +    topological level, to the commit data chunk. This means that we need to\n     +    compute both versions of generation numbers when writing the\n     +    commit-graph file. Therefore, let's introduce a commit-slab to store\n     +    topological levels; corrected commit date will be stored in the member\n     +    `generation` of struct commit_graph_data.\n      \n          When Git creates a split commit-graph, it takes advantage of the\n          generation values that have been computed already and present in\n          existing commit-graph files.\n      \n     -    So, let's add a pointer to struct commit_graph to the topological level\n     -    commit-slab and populate it with topological levels while writing a\n     -    split commit-graph.\n     +    So, let's add a pointer to struct commit_graph as well as struct\n     +    write_commit_graph_context to the topological level commit-slab\n     +    and populate it with topological levels while writing a commit-graph\n     +    file.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n     @@ commit-graph.c: struct write_commit_graph_context {\n       \t\t order_by_pack:1;\n       \n      +\tstruct topo_level_slab *topo_levels;\n     - \tconst struct split_commit_graph_opts *split_opts;\n     + \tconst struct commit_graph_opts *opts;\n       \tsize_t total_bloom_filter_data_size;\n       \tconst struct bloom_filter_settings *bloom_settings;\n      @@ commit-graph.c: static int write_graph_chunk_data(struct hashfile *f,\n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n       \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n       \t\t\t\t\tctx->commits.nr);\n       \tfor (i = 0; i < ctx->commits.nr; i++) {\n     --\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n     -+\t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n     +-\t\ttimestamp_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n     ++\t\ttimestamp_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n       \n       \t\tdisplay_progress(ctx->progress, i + 1);\n     --\t\tif (generation != GENERATION_NUMBER_V1_INFINITY &&\n     +-\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n      -\t\t    generation != GENERATION_NUMBER_ZERO)\n     -+\t\tif (level != GENERATION_NUMBER_V1_INFINITY &&\n     ++\t\tif (level != GENERATION_NUMBER_INFINITY &&\n      +\t\t    level != GENERATION_NUMBER_ZERO)\n       \t\t\tcontinue;\n       \n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n      -\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n      +\t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n       \n     --\t\t\t\tif (generation == GENERATION_NUMBER_V1_INFINITY ||\n     +-\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n      -\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n     -+\t\t\t\tif (level == GENERATION_NUMBER_V1_INFINITY ||\n     ++\t\t\t\tif (level == GENERATION_NUMBER_INFINITY ||\n      +\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n       \t\t\t\t\tall_parents_computed = 0;\n       \t\t\t\t\tcommit_list_insert(parent->item, &list);\n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n      -\t\t\t\tdata->generation = max_generation + 1;\n       \t\t\t\tpop_commit(&list);\n       \n     --\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n     --\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n     -+\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n     -+\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n     +-\t\t\t\tif (data->generation > GENERATION_NUMBER_V1_MAX)\n     +-\t\t\t\t\tdata->generation = GENERATION_NUMBER_V1_MAX;\n     ++\t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n     ++\t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n      +\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n       \t\t\t}\n       \t\t}\n       \t}\n      @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n     - \tuint32_t i, count_distinct = 0;\n       \tint res = 0;\n       \tint replace = 0;\n     + \tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n      +\tstruct topo_level_slab topo_levels;\n       \n       \tif (!commit_graph_compatible(the_repository))\n       \t\treturn 0;\n      @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n     - \t\t}\n     - \t}\n     + \t\t\t\t\t\t\t bloom_settings.max_changed_paths);\n     + \tctx->bloom_settings = &bloom_settings;\n       \n      +\tinit_topo_level_slab(&topo_levels);\n      +\tctx->topo_levels = &topo_levels;\n     @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n      +\t\t}\n      +\t}\n      +\n     - \tif (pack_indexes) {\n     - \t\tctx->order_by_pack = 1;\n     - \t\tif ((res = fill_oids_from_packs(ctx, pack_indexes)))\n     + \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n     + \t\tctx->changed_paths = 1;\n     + \tif (!(flags & COMMIT_GRAPH_NO_WRITE_BLOOM_FILTERS)) {\n      \n       ## commit-graph.h ##\n      @@ commit-graph.h: struct commit_graph {\n     @@ commit-graph.h: struct commit_graph {\n       \tstruct bloom_filter_settings *bloom_filter_settings;\n       };\n       \n     -\n     - ## commit.h ##\n     -@@\n     - #define GENERATION_NUMBER_V1_INFINITY 0xFFFFFFFF\n     - #define GENERATION_NUMBER_MAX 0x3FFFFFFF\n     - #define GENERATION_NUMBER_ZERO 0\n     -+#define GENERATION_NUMBER_V2_OFFSET_MAX 0xFFFFFFFF\n     - \n     - struct commit_list {\n     - \tstruct commit *item;\n  7:  4074ace65b !  6:  694ef1ec08 commit-graph: implement corrected commit date\n     @@ Commit message\n            the maximum of its commit date and one more than the largest corrected\n            commit date among its parents.\n      \n     +    As a special case, a root commit with timestamp of zero (01.01.1970\n     +    00:00:00Z) has corrected commit date of one, to be able to distinguish\n     +    from GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit\n     +    date).\n     +\n          To minimize the space required to store corrected commit date, Git\n          stores corrected commit date offsets into the commit-graph file. The\n          corrected commit date offset for a commit is defined as the difference\n          between its corrected commit date and actual commit date.\n      \n     +    Storing corrected commit date requires sizeof(timestamp_t) bytes, which\n     +    in most cases is 64 bits (uintmax_t). However, corrected commit date\n     +    offsets can be safely stored using only 32-bits. This halves the size\n     +    of GDAT chunk, which is a reduction of around 6% in the size of\n     +    commit-graph file.\n     +\n     +    However, using offsets be problematic if one of commits is malformed but\n     +    valid and has committerdate of 0 Unix time, as the offset would be the\n     +    same as corrected commit date and thus require 64-bits to be stored\n     +    properly.\n     +\n          While Git does not write out offsets at this stage, Git stores the\n          corrected commit dates in member generation of struct commit_graph_data.\n          It will begin writing commit date offsets with the introduction of\n     @@ commit-graph.c: static int commit_gen_cmp(const void *va, const void *vb)\n      @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n       \t\t\t\t\tctx->commits.nr);\n       \tfor (i = 0; i < ctx->commits.nr; i++) {\n     - \t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n     + \t\ttimestamp_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n      +\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n       \n       \t\tdisplay_progress(ctx->progress, i + 1);\n     - \t\tif (level != GENERATION_NUMBER_V1_INFINITY &&\n     + \t\tif (level != GENERATION_NUMBER_INFINITY &&\n      -\t\t    level != GENERATION_NUMBER_ZERO)\n      +\t\t    level != GENERATION_NUMBER_ZERO &&\n      +\t\t    corrected_commit_date != GENERATION_NUMBER_INFINITY &&\n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n       \n       \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n       \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n     +-\n      +\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n     - \n     - \t\t\t\tif (level == GENERATION_NUMBER_V1_INFINITY ||\n     + \t\t\t\tif (level == GENERATION_NUMBER_INFINITY ||\n      -\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n      +\t\t\t\t    level == GENERATION_NUMBER_ZERO ||\n      +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_INFINITY ||\n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n       \t\t\t}\n       \n      @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n     - \t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n     - \t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n     + \t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n     + \t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n       \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n      +\n     -+\t\t\t\tif (current->date > max_corrected_commit_date)\n     ++\t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n      +\t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n      +\t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n       \t\t\t}\n       \t\t}\n       \t}\n     -@@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n     - \tfor (i = 0; i < g->num_commits; i++) {\n     - \t\tstruct commit *graph_commit, *odb_commit;\n     - \t\tstruct commit_list *graph_parents, *odb_parents;\n     --\t\ttimestamp_t max_generation = 0;\n     --\t\ttimestamp_t generation;\n     -+\t\ttimestamp_t max_corrected_commit_date = 0;\n     -+\t\ttimestamp_t corrected_commit_date;\n     - \n     - \t\tdisplay_progress(progress, i + 1);\n     - \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n     -@@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n     - \t\t\t\t\t     oid_to_hex(&graph_parents->item->object.oid),\n     - \t\t\t\t\t     oid_to_hex(&odb_parents->item->object.oid));\n     - \n     --\t\t\tgeneration = commit_graph_generation(graph_parents->item);\n     --\t\t\tif (generation > max_generation)\n     --\t\t\t\tmax_generation = generation;\n     -+\t\t\tcorrected_commit_date = commit_graph_generation(graph_parents->item);\n     -+\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n     -+\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n     - \n     - \t\t\tgraph_parents = graph_parents->next;\n     - \t\t\todb_parents = odb_parents->next;\n      @@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n       \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n       \t\t\tcontinue;\n       \n      -\t\t/*\n     --\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n     --\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n     +-\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n     +-\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n      -\t\t * extra logic in the following condition.\n      -\t\t */\n     --\t\tif (max_generation == GENERATION_NUMBER_MAX)\n     +-\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n      -\t\t\tmax_generation--;\n      -\n     --\t\tgeneration = commit_graph_generation(graph_commit);\n     + \t\tgeneration = commit_graph_generation(graph_commit);\n      -\t\tif (generation != max_generation + 1)\n      -\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n     -+\t\tcorrected_commit_date = commit_graph_generation(graph_commit);\n     -+\t\tif (corrected_commit_date < max_corrected_commit_date + 1)\n     ++\t\tif (generation < max_generation + 1)\n      +\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n       \t\t\t\t     oid_to_hex(&cur_oid),\n     --\t\t\t\t     generation,\n     --\t\t\t\t     max_generation + 1);\n     -+\t\t\t\t     corrected_commit_date,\n     -+\t\t\t\t     max_corrected_commit_date + 1);\n     - \n     - \t\tif (graph_commit->date != odb_commit->date)\n     - \t\t\tgraph_report(_(\"commit date for commit %s in commit-graph is %\"PRItime\" != %\"PRItime),\n     + \t\t\t\t     generation,\n     + \t\t\t\t     max_generation + 1);\n  8:  4e746628ac !  7:  b903efe2ea commit-graph: implement generation data chunk\n     @@ Commit message\n          between graph versions in a backwards compatible manner.\n      \n          We are going to introduce a new chunk called Generation Data chunk (or\n     -    GDAT). GDAT stores generation number v2 (and any subsequent versions),\n     -    whereas CDAT will still store topological level.\n     +    GDAT). GDAT stores corrected committer date offsets whereas CDAT will\n     +    still store topological level.\n      \n          Old Git does not understand GDAT chunk and would ignore it, reading\n          topological levels from CDAT. New Git can parse GDAT and take advantage\n     @@ Commit message\n          which forces commit-graph file to be written without generation data\n          chunk to emulate a commit-graph file written by old Git.\n      \n     +    While storing corrected commit date offset instead of the corrected\n     +    commit date saves us 4 bytes per commit, it's possible for the offsets\n     +    to overflow the 4-bytes allocated. As such overflows are exceedingly\n     +    rare, we use the following overflow management scheme:\n     +\n     +    We introduce a new commit-graph chunk, GENERATION_DATA_OVERFLOW ('GDOV')\n     +    to store corrected commit dates for commits with offsets greater than\n     +    GENERATION_NUMBER_V2_OFFSET_MAX.\n     +\n     +    If the offset is greater than GENERATION_NUMBER_V2_OFFSET_MAX, we set\n     +    the MSB of the offset and the other bits store the position of corrected\n     +    commit date in GDOV chunk, similar to how Extra Edge List is maintained.\n     +\n     +    We test the overflow-related code with the following repo history:\n     +\n     +               F - N - U\n     +              /         \\\n     +    U - N - U            N\n     +             \\          /\n     +              N - F - N\n     +\n     +    Where the commits denoted by U have committer date of zero seconds\n     +    since Unix epoch, the commits denoted by N have committer date of\n     +    1112354055 (default committer date for the test suite) seconds since\n     +    Unix epoch and the commits denoted by F have committer date of\n     +    (2 ^ 31 - 2) seconds since Unix epoch.\n     +\n     +    The largest offset observed is 2 ^ 31, just large enough to overflow.\n     +\n          [1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n     @@ commit-graph.c: void git_test_write_commit_graph_or_die(void)\n       #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n       #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n      +#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n     ++#define GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW 0x47444f56 /* \"GDOV\" */\n       #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n       #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n       #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n       #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n      -#define MAX_NUM_CHUNKS 7\n     -+#define MAX_NUM_CHUNKS 8\n     ++#define MAX_NUM_CHUNKS 9\n       \n       #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n       \n     -@@ commit-graph.c: struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size)\n     +@@ commit-graph.c: void git_test_write_commit_graph_or_die(void)\n     + #define GRAPH_MIN_SIZE (GRAPH_HEADER_SIZE + 4 * GRAPH_CHUNKLOOKUP_WIDTH \\\n     + \t\t\t+ GRAPH_FANOUT_SIZE + the_hash_algo->rawsz)\n     + \n     ++#define CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW (1ULL << 31)\n     ++\n     + /* Remember to update object flag allocation in object.h */\n     + #define REACHABLE       (1u<<15)\n     + \n     +@@ commit-graph.c: struct commit_graph *parse_commit_graph(struct repository *r,\n       \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n       \t\t\tbreak;\n       \n     @@ commit-graph.c: struct commit_graph *parse_commit_graph(void *graph_map, size_t\n      +\t\t\telse\n      +\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n      +\t\t\tbreak;\n     ++\n     ++\t\tcase GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW:\n     ++\t\t\tif (graph->chunk_generation_data_overflow)\n     ++\t\t\t\tchunk_repeated = 1;\n     ++\t\t\telse\n     ++\t\t\t\tgraph->chunk_generation_data_overflow = data + chunk_offset;\n     ++\t\t\tbreak;\n      +\n       \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n       \t\t\tif (graph->chunk_extra_edges)\n       \t\t\t\tchunk_repeated = 1;\n     +@@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n     + {\n     + \tconst unsigned char *commit_data;\n     + \tstruct commit_graph_data *graph_data;\n     +-\tuint32_t lex_index;\n     +-\tuint64_t date_high, date_low;\n     ++\tuint32_t lex_index, offset_pos;\n     ++\tuint64_t date_high, date_low, offset;\n     + \n     + \twhile (pos < g->num_commits_in_base)\n     + \t\tg = g->base_graph;\n      @@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n       \tdate_low = get_be32(commit_data + g->hash_len + 12);\n       \titem->date = (timestamp_t)((date_high << 32) | date_low);\n       \n      -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n     -+\tif (g->chunk_generation_data)\n     -+\t\tgraph_data->generation = item->date +\n     -+\t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     -+\telse\n     ++\tif (g->chunk_generation_data) {\n     ++\t\toffset = (timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     ++\n     ++\t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n     ++\t\t\toffset_pos = offset ^ CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW;\n     ++\t\t\tgraph_data->generation = get_be64(g->chunk_generation_data_overflow + 8 * offset_pos);\n     ++\t\t} else\n     ++\t\t\tgraph_data->generation = item->date + offset;\n     ++\t} else\n      +\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n       \n       \tif (g->topo_levels)\n       \t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n     +@@ commit-graph.c: struct write_commit_graph_context {\n     + \tstruct packed_oid_list oids;\n     + \tstruct packed_commit_list commits;\n     + \tint num_extra_edges;\n     ++\tint num_generation_data_overflows;\n     + \tunsigned long approx_nr_objects;\n     + \tstruct progress *progress;\n     + \tint progress_done;\n      @@ commit-graph.c: struct write_commit_graph_context {\n       \t\t report_progress:1,\n       \t\t split:1,\n     @@ commit-graph.c: struct write_commit_graph_context {\n      +\t\t write_generation_data:1;\n       \n       \tstruct topo_level_slab *topo_levels;\n     - \tconst struct split_commit_graph_opts *split_opts;\n     + \tconst struct commit_graph_opts *opts;\n      @@ commit-graph.c: static int write_graph_chunk_data(struct hashfile *f,\n       \treturn 0;\n       }\n     @@ commit-graph.c: static int write_graph_chunk_data(struct hashfile *f,\n      +static int write_graph_chunk_generation_data(struct hashfile *f,\n      +\t\t\t\t\t      struct write_commit_graph_context *ctx)\n      +{\n     -+\tint i;\n     ++\tint i, num_generation_data_overflows = 0;\n      +\tfor (i = 0; i < ctx->commits.nr; i++) {\n      +\t\tstruct commit *c = ctx->commits.list[i];\n      +\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n      +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n      +\n     -+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX)\n     -+\t\t\toffset = GENERATION_NUMBER_V2_OFFSET_MAX;\n     ++\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n     ++\t\t\toffset = CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW | num_generation_data_overflows;\n     ++\t\t\tnum_generation_data_overflows++;\n     ++\t\t}\n     ++\n      +\t\thashwrite_be32(f, offset);\n      +\t}\n      +\n      +\treturn 0;\n      +}\n     ++\n     ++static int write_graph_chunk_generation_data_overflow(struct hashfile *f,\n     ++\t\t\t\t\t\t       struct write_commit_graph_context *ctx)\n     ++{\n     ++\tint i;\n     ++\tfor (i = 0; i < ctx->commits.nr; i++) {\n     ++\t\tstruct commit *c = ctx->commits.list[i];\n     ++\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n     ++\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n     ++\n     ++\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n     ++\t\t\thashwrite_be32(f, offset >> 32);\n     ++\t\t\thashwrite_be32(f, (uint32_t) offset);\n     ++\t\t}\n     ++\t}\n     ++\n     ++\treturn 0;\n     ++}\n      +\n       static int write_graph_chunk_extra_edges(struct hashfile *f,\n     --\t\t\t\t\t struct write_commit_graph_context *ctx)\n     -+\t\t\t\t\t  struct write_commit_graph_context *ctx)\n     + \t\t\t\t\t struct write_commit_graph_context *ctx)\n       {\n     - \tstruct commit **list = ctx->commits.list;\n     - \tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n     +@@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n     + \n     + \t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n     + \t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n     ++\n     + \t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n     ++\n     ++\t\t\t\tif (commit_graph_data_at(current)->generation - current->date > GENERATION_NUMBER_V2_OFFSET_MAX)\n     ++\t\t\t\t\tctx->num_generation_data_overflows++;\n     + \t\t\t}\n     + \t\t}\n     + \t}\n      @@ commit-graph.c: static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n       \tchunks[2].id = GRAPH_CHUNKID_DATA;\n       \tchunks[2].size = (hashsz + 16) * ctx->commits.nr;\n     @@ commit-graph.c: static int write_commit_graph_file(struct write_commit_graph_con\n      +\t\tchunks[num_chunks].size = sizeof(uint32_t) * ctx->commits.nr;\n      +\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data;\n      +\t\tnum_chunks++;\n     ++\t}\n     ++\tif (ctx->num_generation_data_overflows) {\n     ++\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW;\n     ++\t\tchunks[num_chunks].size = sizeof(timestamp_t) * ctx->num_generation_data_overflows;\n     ++\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data_overflow;\n     ++\t\tnum_chunks++;\n      +\t}\n       \tif (ctx->num_extra_edges) {\n       \t\tchunks[num_chunks].id = GRAPH_CHUNKID_EXTRAEDGES;\n       \t\tchunks[num_chunks].size = 4 * ctx->num_extra_edges;\n      @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n     - \tctx->split_opts = split_opts;\n     + \tctx->opts = opts;\n       \tctx->total_bloom_filter_data_size = 0;\n      +\tctx->write_generation_data = 1;\n     ++\tctx->num_generation_data_overflows = 0;\n       \n     - \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n     - \t\tctx->changed_paths = 1;\n     + \tbloom_settings.bits_per_entry = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\",\n     + \t\t\t\t\t\t      bloom_settings.bits_per_entry);\n      \n       ## commit-graph.h ##\n      @@\n     @@ commit-graph.h: struct commit_graph {\n       \tconst unsigned char *chunk_oid_lookup;\n       \tconst unsigned char *chunk_commit_data;\n      +\tconst unsigned char *chunk_generation_data;\n     ++\tconst unsigned char *chunk_generation_data_overflow;\n       \tconst unsigned char *chunk_extra_edges;\n       \tconst unsigned char *chunk_base_graphs;\n       \tconst unsigned char *chunk_bloom_indexes;\n      \n     + ## commit.h ##\n     +@@\n     + #define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n     + #define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n     + #define GENERATION_NUMBER_ZERO 0\n     ++#define GENERATION_NUMBER_V2_OFFSET_MAX ((1ULL << 31) - 1)\n     + \n     + struct commit_list {\n     + \tstruct commit *item;\n     +\n       ## t/README ##\n      @@ t/README: GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n       be written after every 'git commit' command, and overrides the\n     @@ t/helper/test-read-graph.c: int cmd__read_graph(int argc, const char **argv)\n       \t\tprintf(\" commit_metadata\");\n      +\tif (graph->chunk_generation_data)\n      +\t\tprintf(\" generation_data\");\n     ++\tif (graph->chunk_generation_data_overflow)\n     ++\t\tprintf(\" generation_data_overflow\");\n       \tif (graph->chunk_extra_edges)\n       \t\tprintf(\" extra_edges\");\n       \tif (graph->chunk_bloom_indexes)\n      \n       ## t/t4216-log-bloom.sh ##\n      @@ t/t4216-log-bloom.sh: test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n     - \tgit commit-graph write --reachable --changed-paths\n       '\n     + \n       graph_read_expect () {\n      -\tNUM_CHUNKS=5\n      +\tNUM_CHUNKS=6\n       \tcat >expect <<- EOF\n     - \theader: 43475048 1 1 $NUM_CHUNKS 0\n     + \theader: 43475048 1 $(test_oid oid_version) $NUM_CHUNKS 0\n       \tnum_commits: $1\n      -\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n      +\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n     @@ t/t5318-commit-graph.sh: test_expect_success 'write graph in bare repo' '\n       '\n       \n       graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n     -@@ t/t5318-commit-graph.sh: test_expect_success 'replace-objects invalidates commit-graph' '\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'warn on improper hash version' '\n       \n       test_expect_success 'git commit-graph verify' '\n       \tcd \"$TRASH_DIRECTORY/full\" &&\n     @@ t/t5318-commit-graph.sh: test_expect_success 'replace-objects invalidates commit\n       '\n       \n       NUM_COMMITS=9\n     +@@ t/t5318-commit-graph.sh: test_expect_success 'corrupt commit-graph write (missing tree)' '\n     + \t)\n     + '\n     + \n     ++test_commit_with_date() {\n     ++  file=\"$1.t\" &&\n     ++  echo \"$1\" >\"$file\" &&\n     ++  git add \"$file\" &&\n     ++  GIT_COMMITTER_DATE=\"$2\" GIT_AUTHOR_DATE=\"$2\" git commit -m \"$1\"\n     ++  git tag \"$1\"\n     ++}\n     ++\n     ++test_expect_success 'overflow corrected commit date offset' '\n     ++\tobjdir=\".git/objects\" &&\n     ++\tUNIX_EPOCH_ZERO=\"1970-01-01 00:00 +0000\" &&\n     ++\tFUTURE_DATE=\"@2147483646 +0000\" &&\n     ++\ttest_oid_cache <<-EOF &&\n     ++\toid_version sha1:1\n     ++\toid_version sha256:2\n     ++\tEOF\n     ++\tcd \"$TRASH_DIRECTORY\" &&\n     ++\tmkdir repo &&\n     ++\tcd repo &&\n     ++\tgit init &&\n     ++\ttest_commit_with_date 1 \"$UNIX_EPOCH_ZERO\" &&\n     ++\ttest_commit 2 &&\n     ++\ttest_commit_with_date 3 \"$UNIX_EPOCH_ZERO\" &&\n     ++\tgit commit-graph write --reachable &&\n     ++\tgraph_read_expect 3 generation_data &&\n     ++\ttest_commit_with_date 4 \"$FUTURE_DATE\" &&\n     ++\ttest_commit 5 &&\n     ++\ttest_commit_with_date 6 \"$UNIX_EPOCH_ZERO\" &&\n     ++\tgit branch left &&\n     ++\tgit reset --hard 3 &&\n     ++\ttest_commit 7 &&\n     ++\ttest_commit_with_date 8 \"$FUTURE_DATE\" &&\n     ++\ttest_commit 9 &&\n     ++\tgit branch right &&\n     ++\tgit reset --hard 3 &&\n     ++\tgit merge left right &&\n     ++\tgit commit-graph write --reachable &&\n     ++\tgraph_read_expect 10 \"generation_data generation_data_overflow\" &&\n     ++\tgit commit-graph verify\n     ++'\n     ++\n     ++graph_git_behavior 'overflow corrected commit date offset' repo left right\n     ++\n     + test_done\n      \n       ## t/t5324-split-commit-graph.sh ##\n      @@ t/t5324-split-commit-graph.sh: test_expect_success 'setup repo' '\n     @@ t/t5324-split-commit-graph.sh: test_expect_success 'setup repo' '\n      -\tbase sha256:1496\n      +\tbase sha1:1408\n      +\tbase sha256:1528\n     - \tEOM\n     - '\n       \n     + \toid_version sha1:1\n     + \toid_version sha256:2\n      @@ t/t5324-split-commit-graph.sh: graph_read_expect() {\n       \t\tNUM_BASE=$2\n       \tfi\n       \tcat >expect <<- EOF\n     --\theader: 43475048 1 1 3 $NUM_BASE\n     -+\theader: 43475048 1 1 4 $NUM_BASE\n     +-\theader: 43475048 1 $(test_oid oid_version) 3 $NUM_BASE\n     ++\theader: 43475048 1 $(test_oid oid_version) 4 $NUM_BASE\n       \tnum_commits: $1\n      -\tchunks: oid_fanout oid_lookup commit_metadata\n      +\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n     @@ t/t6600-test-reach.sh: test_expect_success 'in_merge_bases:miss' '\n      +\ttest_all_modes in_merge_bases\n       '\n       \n     + test_expect_success 'in_merge_bases_many:hit' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'in_merge_bases_many:hit' '\n     + \tX:commit-5-7\n     + \tEOF\n     + \techo \"in_merge_bases_many(A,X):1\" >expect &&\n     +-\ttest_three_modes in_merge_bases_many\n     ++\ttest_all_modes in_merge_bases_many\n     + '\n     + \n     + test_expect_success 'in_merge_bases_many:miss' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'in_merge_bases_many:miss' '\n     + \tX:commit-8-6\n     + \tEOF\n     + \techo \"in_merge_bases_many(A,X):0\" >expect &&\n     +-\ttest_three_modes in_merge_bases_many\n     ++\ttest_all_modes in_merge_bases_many\n     + '\n     + \n     + test_expect_success 'in_merge_bases_many:miss-heuristic' '\n     +@@ t/t6600-test-reach.sh: test_expect_success 'in_merge_bases_many:miss-heuristic' '\n     + \tX:commit-6-6\n     + \tEOF\n     + \techo \"in_merge_bases_many(A,X):0\" >expect &&\n     +-\ttest_three_modes in_merge_bases_many\n     ++\ttest_all_modes in_merge_bases_many\n     + '\n     + \n       test_expect_success 'is_descendant_of:hit' '\n      @@ t/t6600-test-reach.sh: test_expect_success 'is_descendant_of:hit' '\n       \tX:commit-1-1\n  9:  5a147a9704 !  8:  8ec119edc6 commit-graph: use generation v2 only if entire chain does\n     @@ Commit message\n          commits in the lower layer before allowing the topo-order queue to write\n          anything to output (depending on the size of the upper layer).\n      \n     +    When writing the new layer in split commit-graph, we write a GDAT chunk\n     +    only if the topmost layer has a GDAT chunk. This guarantees that if a\n     +    layer has GDAT chunk, all lower layers must have a GDAT chunk as well.\n     +\n     +    Rewriting layers follows similar approach: if the topmost layer below\n     +    the set of layers being rewritten (in the split commit-graph chain)\n     +    exists, and it does not contain GDAT chunk, then the result of rewrite\n     +    does not have GDAT chunks either.\n     +\n          Signed-off-by: Derrick Stolee <dstolee@microsoft.com>\n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n     @@ commit-graph.c: static struct commit_graph *load_commit_graph_chain(struct repos\n       \treturn graph_chain;\n       }\n       \n     -+static void validate_mixed_generation_chain(struct repository *r)\n     ++static void validate_mixed_generation_chain(struct commit_graph *g)\n      +{\n     -+\tstruct commit_graph *g = r->objects->commit_graph;\n     -+\tint read_generation_data = 1;\n     ++\tint read_generation_data;\n      +\n     -+\twhile (g) {\n     -+\t\tif (!g->chunk_generation_data) {\n     -+\t\t\tread_generation_data = 0;\n     -+\t\t\tbreak;\n     -+\t\t}\n     -+\t\tg = g->base_graph;\n     -+\t}\n     ++\tif (!g)\n     ++\t\treturn;\n      +\n     -+\tg = r->objects->commit_graph;\n     ++\tread_generation_data = !!g->chunk_generation_data;\n      +\n      +\twhile (g) {\n      +\t\tg->read_generation_data = read_generation_data;\n     @@ commit-graph.c: struct commit_graph *read_commit_graph_one(struct repository *r,\n       \tif (!g)\n       \t\tg = load_commit_graph_chain(r, odb);\n       \n     -+\tvalidate_mixed_generation_chain(r);\n     ++\tvalidate_mixed_generation_chain(g);\n      +\n       \treturn g;\n       }\n     @@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct c\n       \tdate_low = get_be32(commit_data + g->hash_len + 12);\n       \titem->date = (timestamp_t)((date_high << 32) | date_low);\n       \n     --\tif (g->chunk_generation_data)\n     -+\tif (g->chunk_generation_data && g->read_generation_data)\n     - \t\tgraph_data->generation = item->date +\n     - \t\t\t(timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     - \telse\n     -@@ commit-graph.c: void load_commit_graph_info(struct repository *r, struct commit *item)\n     - \tuint32_t pos;\n     - \tif (!prepare_commit_graph(r))\n     - \t\treturn;\n     +-\tif (g->chunk_generation_data) {\n     ++\tif (g->chunk_generation_data && g->read_generation_data) {\n     + \t\toffset = (timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     + \n     + \t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n     +@@ commit-graph.c: static void split_graph_merge_strategy(struct write_commit_graph_context *ctx)\n     + \t\t}\n     + \t}\n     + \n     ++\tif (!ctx->write_generation_data && g->chunk_generation_data)\n     ++\t\tctx->write_generation_data = 1;\n      +\n     - \tif (find_commit_in_graph(item, r->objects->commit_graph, &pos))\n     - \t\tfill_commit_graph_info(item, r->objects->commit_graph, pos);\n     - }\n     + \tif (flags != COMMIT_GRAPH_SPLIT_REPLACE)\n     + \t\tctx->new_base_graph = g;\n     + \telse if (ctx->num_commit_graphs_after != 1)\n     +@@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n     + \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n     + \n     + \t\twhile (g) {\n     ++\t\t\tg->read_generation_data = 1;\n     + \t\t\tg->topo_levels = &topo_levels;\n     + \t\t\tg = g->base_graph;\n     + \t\t}\n      @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \n       \t\tg = ctx->r->objects->commit_graph;\n     @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \t\t\tg = g->base_graph;\n      @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \n     - \t\tif (ctx->split_opts)\n     - \t\t\treplace = ctx->split_opts->flags & COMMIT_GRAPH_SPLIT_REPLACE;\n     + \t\tif (ctx->opts)\n     + \t\t\treplace = ctx->opts->split_flags & COMMIT_GRAPH_SPLIT_REPLACE;\n      +\n      +\t\tif (replace)\n      +\t\t\tctx->write_generation_data = 1;\n     @@ commit-graph.h: struct commit_graph {\n       \tstruct object_directory *odb;\n       \n       \tuint32_t num_commits_in_base;\n     -+\tuint32_t read_generation_data;\n     ++\tunsigned int read_generation_data;\n       \tstruct commit_graph *base_graph;\n       \n       \tconst uint32_t *chunk_oid_fanout;\n      \n       ## t/t5324-split-commit-graph.sh ##\n     -@@ t/t5324-split-commit-graph.sh: done <<\\EOF\n     - 0600 -r--------\n     - EOF\n     +@@ t/t5324-split-commit-graph.sh: test_expect_success '--split=replace with partial Bloom data' '\n     + \tverify_chain_files_exist $graphdir\n     + '\n       \n      +test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n      +\tmkdir mixed &&\n      +\tgraphdir=\".git/objects/info/commit-graphs\" &&\n     ++\ttest_oid_cache <<-EOM &&\n     ++\toid_version sha1:1\n     ++\toid_version sha256:2\n     ++\tEOM\n      +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n      +\tgit init &&\n      +\tgit config core.commitGraph true &&\n     @@ t/t5324-split-commit-graph.sh: done <<\\EOF\n      +\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n      +\ttest-tool read-graph >output &&\n      +\tcat >expect <<-EOF &&\n     -+\theader: 43475048 1 1 4 1\n     ++\theader: 43475048 1 $(test_oid oid_version) 4 1\n      +\tnum_commits: 2\n      +\tchunks: oid_fanout oid_lookup commit_metadata\n      +\tEOF\n     @@ t/t5324-split-commit-graph.sh: done <<\\EOF\n      +\tgit commit-graph write --reachable --split=no-merge &&\n      +\ttest-tool read-graph >output &&\n      +\tcat >expect <<-EOF &&\n     -+\theader: 43475048 1 1 4 2\n     ++\theader: 43475048 1 $(test_oid oid_version) 4 2\n      +\tnum_commits: 3\n      +\tchunks: oid_fanout oid_lookup commit_metadata\n      +\tEOF\n     @@ t/t5324-split-commit-graph.sh: done <<\\EOF\n      +\tgraph_read_expect 15 &&\n      +\tgit commit-graph verify\n      +'\n     ++\n     ++test_expect_success 'add one commit, write a tip graph' '\n     ++\tcd \"$TRASH_DIRECTORY/mixed\" &&\n     ++\ttest_commit 11 &&\n     ++\tgit branch commits/11 &&\n     ++\tgit commit-graph write --reachable --split &&\n     ++\ttest_path_is_missing $infodir/commit-graph &&\n     ++\ttest_path_is_file $graphdir/commit-graph-chain &&\n     ++\tls $graphdir/graph-*.graph >graph-files &&\n     ++\ttest_line_count = 2 graph-files &&\n     ++\tverify_chain_files_exist $graphdir\n     ++'\n      +\n       test_done\n 10:  439adc1718 !  9:  bb9b02af32 commit-reach: use corrected commit dates in paint_down_to_common()\n     @@ Commit message\n          With corrected commit dates implemented, we no longer have to rely on\n          commit date as a heuristic in paint_down_to_common().\n      \n     -    t6024-recursive-merge setups a unique repository where all commits have\n     -    the same committer date without well-defined merge-base. As this has\n     -    already caused problems (as noted in 859fdc0 (commit-graph: define\n     -    GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph within the\n     -    test script.\n     +    While using corrected commit dates Git walks nearly the same number of\n     +    commits as commit date, the process is slower as for each comparision we\n     +    have to access a commit-slab (for corrected committer date) instead of\n     +    accessing struct member (for committer date).\n     +\n     +    For example, the command `git merge-base v4.8 v4.9` on the linux\n     +    repository walks 167468 commits, taking 0.135s for committer date and\n     +    167496 commits, taking 0.157s for corrected committer date respectively.\n     +\n     +    t6404-recursive-merge setups a unique repository where all commits have\n     +    the same committer date without well-defined merge-base.\n     +\n     +    While running tests with GIT_TEST_COMMIT_GRAPH unset, we use committer\n     +    date as a heuristic in paint_down_to_common(). 6404.1 'combined merge\n     +    conflicts' merges commits in the order:\n     +    - Merge C with B to form a intermediate commit.\n     +    - Merge the intermediate commit with A.\n     +\n     +    With GIT_TEST_COMMIT_GRAPH=1, we write a commit-graph and subsequently\n     +    use the corrected committer date, which changes the order in which\n     +    commits are merged:\n     +    - Merge A with B to form a intermediate commit.\n     +    - Merge the intermediate commit with C.\n     +\n     +    While resulting repositories are equivalent, 6404.4 'virtual trees were\n     +    processed' fails with GIT_TEST_COMMIT_GRAPH=1 as we are selecting\n     +    different merge-bases and thus have different object ids for the\n     +    intermediate commits.\n     +\n     +    As this has already causes problems (as noted in 859fdc0 (commit-graph:\n     +    define GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph\n     +    within t6404-recursive-merge.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n     @@ commit-graph.c: int generation_numbers_enabled(struct repository *r)\n      +\tif (!g->num_commits)\n      +\t\treturn 0;\n      +\n     -+\treturn !!g->chunk_generation_data;\n     ++\treturn g->read_generation_data;\n      +}\n      +\n     - static void close_commit_graph_one(struct commit_graph *g)\n     + struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r)\n       {\n     - \tif (!g)\n     + \tstruct commit_graph *g = r->objects->commit_graph;\n      \n       ## commit-graph.h ##\n     -@@ commit-graph.h: struct commit_graph *parse_commit_graph(void *graph_map, size_t graph_size);\n     +@@ commit-graph.h: struct commit_graph *read_commit_graph_one(struct repository *r,\n     + struct commit_graph *parse_commit_graph(struct repository *r,\n     + \t\t\t\t\tvoid *graph_map, size_t graph_size);\n     + \n     ++struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n     ++\n     + /*\n     +  * Return 1 if and only if the repository has a commit-graph\n     +  * file and generation numbers are computed in that file.\n        */\n       int generation_numbers_enabled(struct repository *r);\n       \n     +-struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n      +/*\n      + * Return 1 if and only if the repository has a commit-graph\n      + * file and generation data chunk has been written for the file.\n      + */\n      +int corrected_commit_dates_enabled(struct repository *r);\n     -+\n     + \n       enum commit_graph_write_flags {\n       \tCOMMIT_GRAPH_WRITE_APPEND     = (1 << 0),\n     - \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\n      \n       ## commit-reach.c ##\n      @@ commit-reach.c: static struct commit_list *paint_down_to_common(struct repository *r,\n     @@ commit-reach.c: static struct commit_list *paint_down_to_common(struct repositor\n       \n       \tone->object.flags |= PARENT1;\n      \n     - ## t/t6024-recursive-merge.sh ##\n     -@@ t/t6024-recursive-merge.sh: GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n     + ## t/t6404-recursive-merge.sh ##\n     +@@ t/t6404-recursive-merge.sh: GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n       export GIT_COMMITTER_DATE\n       \n       test_expect_success 'setup tests' '\n     @@ t/t6024-recursive-merge.sh: GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n       \techo 1 >a1 &&\n       \tgit add a1 &&\n       \tGIT_AUTHOR_DATE=\"2006-12-12 23:00:00\" git commit -m 1 a1 &&\n     -@@ t/t6024-recursive-merge.sh: test_expect_success 'setup tests' '\n     +@@ t/t6404-recursive-merge.sh: test_expect_success 'setup tests' '\n       '\n       \n       test_expect_success 'combined merge conflicts' '\n     @@ t/t6024-recursive-merge.sh: test_expect_success 'setup tests' '\n       '\n       \n       test_expect_success 'result contains a conflict' '\n     +@@ t/t6404-recursive-merge.sh: test_expect_success 'result contains a conflict' '\n     + '\n     + \n     + test_expect_success 'virtual trees were processed' '\n     ++\t# TODO: fragile test, relies on ambigious merge-base resolution\n     + \tgit ls-files --stage >out &&\n     + \n     + \tcat >expect <<-EOF &&\n 11:  f6f91af305 ! 10:  9ada43967d doc: add corrected commit date info\n     @@ Documentation/technical/commit-graph-format.txt: Git commit graph format\n       - The root tree OID.\n       \n      @@ Documentation/technical/commit-graph-format.txt: CHUNK DATA:\n     +       position. If there are more than two parents, the second value\n     +       has its most-significant bit on and the other bits store an array\n     +       position into the Extra Edge List chunk.\n     +-    * The next 8 bytes store the generation number of the commit and\n     ++    * The next 8 bytes store the topological level (generation number v1)\n     ++      of the commit and\n     +       the commit time in seconds since EPOCH. The generation number\n     +       uses the higher 30 bits of the first 4 bytes, while the commit\n     +       time uses the 32 bits of the second 4 bytes, along with the lowest\n             2 bits of the lowest byte, storing the 33rd and 34th bit of the\n             commit time.\n       \n     -+  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes) [Optional]\n     ++  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes)\n      +    * This list of 4-byte values store corrected commit date offsets for the\n      +      commits, arranged in the same order as commit data chunk.\n     -+    * This list can be later modified to store future generation number related\n     -+      data.\n     ++    * If the corrected commit date offset cannot be stored within 31 bits,\n     ++      the value has its most-significant bit on and the other bits store\n     ++      the position of corrected commit date into the Generation Data Overflow\n     ++      chunk.\n     ++\n     ++  Generation Data Overflow (ID: {'G', 'D', 'O', 'V' }) [Optional]\n     ++    * This list of 8-byte values stores the corrected commit dates for commits\n     ++      with corrected commit date offsets that cannot be stored within 31 bits.\n      +\n         Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n             This list of 4-byte values store the second through nth parents for\n     @@ Documentation/technical/commit-graph.txt: A consumer may load the following info\n       \n      -Define the \"generation number\" of a commit recursively as follows:\n      +There are two definitions of generation number:\n     -+1. Corrected committer dates\n     -+2. Topological levels\n     -+\n     ++1. Corrected committer dates (generation number v2)\n     ++2. Topological levels (generation nummber v1)\n     + \n     +- * A commit with no parents (a root commit) has generation number one.\n      +Define \"corrected committer date\" of a commit recursively as follows:\n     -+\n     + \n     +- * A commit with at least one parent has generation number one more than\n     +-   the largest generation number among its parents.\n      +  * A commit with no parents (a root commit) has corrected committer date\n      +    equal to its committer date.\n     -+\n     + \n     +-Equivalently, the generation number of a commit A is one more than the\n      +  * A commit with at least one parent has corrected committer date equal to\n      +    the maximum of its commiter date and one more than the largest corrected\n      +    committer date among its parents.\n      +\n     ++  * As a special case, a root commit with timestamp zero has corrected commit\n     ++    date of 1, to be able to distinguish it from GENERATION_NUMBER_ZERO\n     ++    (that is, an uncomputed corrected commit date).\n     ++\n      +Define the \"topological level\" of a commit recursively as follows:\n     - \n     -  * A commit with no parents (a root commit) has generation number one.\n     - \n     -- * A commit with at least one parent has generation number one more than\n     --   the largest generation number among its parents.\n     ++\n     ++ * A commit with no parents (a root commit) has topological level of one.\n     ++\n      + * A commit with at least one parent has topological level one more than\n      +   the largest topological level among its parents.\n     - \n     --Equivalently, the generation number of a commit A is one more than the\n     ++\n      +Equivalently, the topological level of a commit A is one more than the\n       length of a longest path from A to a root commit. The recursive definition\n       is easier to use for computation and observing the following property:\n       \n     +@@ Documentation/technical/commit-graph.txt: is easier to use for computation and observing the following property:\n     +     generation numbers, then we always expand the boundary commit with highest\n     +     generation number and can easily detect the stopping condition.\n     + \n     ++The properties applies to both versions of generation number, that is both\n     ++corrected committer dates and topological levels.\n     ++\n     + This property can be used to significantly reduce the time it takes to\n     + walk commits and determine topological relationships. Without generation\n     + numbers, the general heuristic is the following:\n      @@ Documentation/technical/commit-graph.txt: numbers, the general heuristic is the following:\n           If A and B are commits with commit time X and Y, respectively, and\n           X < Y, then A _probably_ cannot reach B.\n       \n      -This heuristic is currently used whenever the computation is allowed to\n     --violate topological relationships due to clock skew (such as \"git log\"\n     --with default order), but is not used when the topological order is\n     --required (such as merge base calculations, \"git log --graph\").\n     --\n     - In practice, we expect some commits to be created recently and not stored\n     - in the commit graph. We can treat these commits as having \"infinite\"\n     ++In absence of corrected commit dates (for example, old versions of Git or\n     ++mixed generation graph chains),\n     ++this heuristic is currently used whenever the computation is allowed to\n     + violate topological relationships due to clock skew (such as \"git log\"\n     + with default order), but is not used when the topological order is\n     + required (such as merge base calculations, \"git log --graph\").\n     +@@ Documentation/technical/commit-graph.txt: in the commit graph. We can treat these commits as having \"infinite\"\n       generation number and walk until reaching commits with known generation\n       number.\n       \n     @@ Documentation/technical/commit-graph.txt: fully-computed generation numbers. Usi\n       with generation number *_INFINITY or *_ZERO is valuable.\n       \n      -We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n     --generation numbers are computed to be at least this value. We limit at\n     --this value since it is the largest value that can be stored in the\n     --commit-graph file using the 30 bits available to generation numbers. This\n     --presents another case where a commit can have generation number equal to\n     --that of a parent.\n     -+We use the macro GENERATION_NUMBER_MAX for commits whose generation numbers\n     -+are computed to be at least this value. We limit at this value since it is\n     -+the largest value that can be stored in the commit-graph file using the\n     -+available to generation numbers. This presents another case where a\n     -+commit can have generation number equal to that of a parent.\n     - \n     - Design Details\n     - --------------\n     ++We use the macro GENERATION_NUMBER_MAX for commits whose\n     + generation numbers are computed to be at least this value. We limit at\n     + this value since it is the largest value that can be stored in the\n     + commit-graph file using the 30 bits available to generation numbers. This\n      @@ Documentation/technical/commit-graph.txt: The merge strategy values (2 for the size multiple, 64,000 for the maximum\n       number of commits) could be extracted into config settings for full\n       flexibility.\n       \n     -+We also merge commit-graph chains when we try to write a commit graph with\n     -+two different generation number definitions as they cannot be compared directly.\n     -+We overwrite the existing chain and create a commit-graph with the newer or more\n     -+efficient defintion. For example, overwriting topological levels commit graph\n     -+chain to create a corrected commit dates commit graph chain.\n     ++## Handling Mixed Generation Number Chains\n     ++\n     ++With the introduction of generation number v2 and generation data chunk, the\n     ++following scenario is possible:\n     ++\n     ++1. \"New\" Git writes a commit-graph with the corrected commit dates.\n     ++2. \"Old\" Git writes a split commit-graph on top without corrected commit dates.\n     ++\n     ++A naive approach of using the newest available generation number from\n     ++each layer would lead to violated expectations: the lower layer would\n     ++use corrected commit dates which are much larger than the topological\n     ++levels of the higher layer. For this reason, Git inspects each layer to\n     ++see if any layer is missing corrected commit dates. In such a case, Git\n     ++only uses topological level\n     ++\n     ++When writing a new layer in split commit-graph, we write corrected commit\n     ++dates if the topmost layer has corrected commit dates written. This\n     ++guarantees that if a layer has corrected commit dates, all lower layers\n     ++must have corrected commit dates as well.\n     ++\n     ++When merging layers, we do not consider whether the merged layers had corrected\n     ++commit dates. Instead, the new layer will have corrected commit dates if and\n     ++only if all existing layers below the new layer have corrected commit dates.\n      +\n       ## Deleting graph-{hash} files\n       \n\n-- \ngitgitgadget\n"},{"id":"407057","messageId":"18bb3318a12c859c21c8e95285d551c48d31b54b.1602079786.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"[PATCH v4 03/10] commit-graph: consolidate fill_commit_graph_info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-10-07T14:09:38Z","receivedAt":"2020-10-07T14:09:56Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nBoth fill_commit_graph_info() and fill_commit_in_graph() parse\ninformation present in commit data chunk. Let's simplify the\nimplementation by calling fill_commit_graph_info() within\nfill_commit_in_graph().\n\nfill_commit_graph_info() used to not load committer data from commit data\nchunk. However, with the corrected committer date, we have to load\ncommitter date to calculate generation number value.\n\ne51217e15 (t5000: test tar files that overflow ustar headers,\n30-06-2016) introduced a test 'generate tar with future mtime' that\ncreates a commit with committer date of (2 ^ 36 + 1) seconds since\nEPOCH. The CDAT chunk provides 34-bits for storing committer date, thus\ncommitter time overflows into generation number (within CDAT chunk) and\nhas undefined behavior.\n\nThe test used to pass as fill_commit_graph_info() would not set struct\nmember `date` of struct commit and loads committer date from the object\ndatabase, generating a tar file with the expected mtime.\n\nHowever, with corrected commit date, we will load the committer date\nfrom CDAT chunk (truncated to lower 34-bits to populate the generation\nnumber. Thus, Git sets date and generates tar file with the truncated\nmtime.\n\nThe ustar format (the header format used by most modern tar programs)\nonly has room for 11 (or 12, depending om some implementations) octal\ndigits for the size and mtime of each files.\n\nThus, setting a timestamp of 2 ^ 33 + 1 would overflow the 11-octal\ndigit implementations while still fitting into commit data chunk.\n\nSince we want to test 12-octal digit implementations of ustar as well,\nlet's modify the existing test to no longer use commit-graph file.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c      | 27 ++++++++++-----------------\n t/t5000-tar-tree.sh | 20 +++++++++++++++++++-\n 2 files changed, 29 insertions(+), 18 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 94503e584b..e8362e144e 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -749,15 +749,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n \tuint32_t lex_index;\n+\tuint64_t date_high, date_low;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n+\tif (pos >= g->num_commits + g->num_commits_in_base)\n+\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\n \tlex_index = pos - g->num_commits_in_base;\n \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n \n \tgraph_data = commit_graph_data_at(item);\n \tgraph_data->graph_pos = pos;\n+\n+\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n+\tdate_low = get_be32(commit_data + g->hash_len + 12);\n+\titem->date = (timestamp_t)((date_high << 32) | date_low);\n+\n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n@@ -772,38 +781,22 @@ static int fill_commit_in_graph(struct repository *r,\n {\n \tuint32_t edge_value;\n \tuint32_t *parent_data_ptr;\n-\tuint64_t date_low, date_high;\n \tstruct commit_list **pptr;\n-\tstruct commit_graph_data *graph_data;\n \tconst unsigned char *commit_data;\n \tuint32_t lex_index;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n-\tif (pos >= g->num_commits + g->num_commits_in_base)\n-\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\tfill_commit_graph_info(item, g, pos);\n \n-\t/*\n-\t * Store the \"full\" position, but then use the\n-\t * \"local\" position for the rest of the calculation.\n-\t */\n-\tgraph_data = commit_graph_data_at(item);\n-\tgraph_data->graph_pos = pos;\n \tlex_index = pos - g->num_commits_in_base;\n-\n \tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n \n \titem->object.parsed = 1;\n \n \tset_commit_tree(item, NULL);\n \n-\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n-\tdate_low = get_be32(commit_data + g->hash_len + 12);\n-\titem->date = (timestamp_t)((date_high << 32) | date_low);\n-\n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n-\n \tpptr = &item->parents;\n \n \tedge_value = get_be32(commit_data + g->hash_len);\ndiff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\nindex 3ebb0d3b65..8f41cdc509 100755\n--- a/t/t5000-tar-tree.sh\n+++ b/t/t5000-tar-tree.sh\n@@ -431,11 +431,29 @@ test_expect_success TAR_HUGE,LONG_IS_64BIT 'system tar can read our huge size' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n+\trm -f .git/index &&\n+\techo foo >file &&\n+\tgit add file &&\n+\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n+\t\tgit commit -m \"tempori parendum\"\n+'\n+\n+test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n+\tgit archive HEAD >future.tar\n+'\n+\n+test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n+\techo 2514 >expect &&\n+\ttar_info future.tar | cut -d\" \" -f2 >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n \trm -f .git/index &&\n \techo content >file &&\n \tgit add file &&\n-\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n+\tGIT_TEST_COMMIT_GRAPH=0 GIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n \t\tgit commit -m \"tempori parendum\"\n '\n \n-- \ngitgitgadget\n\n"},{"id":"407055","messageId":"e067f653ad5d474eee5f40c13bb02fde26ebdb9b.1602079786.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"[PATCH v4 05/10] commit-graph: add a slab to store topological levels","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-10-07T14:09:40Z","receivedAt":"2020-10-07T14:09:57Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a later commit we will introduce corrected commit date as the\ngeneration number v2. This value will be stored in the new seperate\nGeneration Data chunk. However, to ensure backwards compatibility with\n\"Old\" Git we need to continue to write generation number v1, which is\ntopological level, to the commit data chunk. This means that we need to\ncompute both versions of generation numbers when writing the\ncommit-graph file. Therefore, let's introduce a commit-slab to store\ntopological levels; corrected commit date will be stored in the member\n`generation` of struct commit_graph_data.\n\nWhen Git creates a split commit-graph, it takes advantage of the\ngeneration values that have been computed already and present in\nexisting commit-graph files.\n\nSo, let's add a pointer to struct commit_graph as well as struct\nwrite_commit_graph_context to the topological level commit-slab\nand populate it with topological levels while writing a commit-graph\nfile.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 47 ++++++++++++++++++++++++++++++++---------------\n commit-graph.h |  1 +\n 2 files changed, 33 insertions(+), 15 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex bfc532de6f..cedd311024 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -64,6 +64,8 @@ void git_test_write_commit_graph_or_die(void)\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n+define_commit_slab(topo_level_slab, uint32_t);\n+\n /* Keep track of the order in which commits are added to our list. */\n define_commit_slab(commit_pos, int);\n static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n@@ -768,6 +770,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\n+\tif (g->topo_levels)\n+\t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n static inline void set_commit_tree(struct commit *c, struct tree *t)\n@@ -962,6 +967,7 @@ struct write_commit_graph_context {\n \t\t changed_paths:1,\n \t\t order_by_pack:1;\n \n+\tstruct topo_level_slab *topo_levels;\n \tconst struct commit_graph_opts *opts;\n \tsize_t total_bloom_filter_data_size;\n \tconst struct bloom_filter_settings *bloom_settings;\n@@ -1108,7 +1114,7 @@ static int write_graph_chunk_data(struct hashfile *f,\n \t\telse\n \t\t\tpackedDate[0] = 0;\n \n-\t\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n+\t\tpackedDate[0] |= htonl(*topo_level_slab_at(ctx->topo_levels, *list) << 2);\n \n \t\tpackedDate[1] = htonl((*list)->date);\n \t\thashwrite(f, packedDate, 8);\n@@ -1350,11 +1356,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\ttimestamp_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n+\t\ttimestamp_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n-\t\t    generation != GENERATION_NUMBER_ZERO)\n+\t\tif (level != GENERATION_NUMBER_INFINITY &&\n+\t\t    level != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1362,29 +1368,27 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit *current = list->item;\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n-\t\t\tuint32_t max_generation = 0;\n+\t\t\tuint32_t max_level = 0;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n-\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n+\t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n \n-\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n-\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n+\t\t\t\tif (level == GENERATION_NUMBER_INFINITY ||\n+\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (generation > max_generation) {\n-\t\t\t\t\tmax_generation = generation;\n+\t\t\t\t} else if (level > max_level) {\n+\t\t\t\t\tmax_level = level;\n \t\t\t\t}\n \t\t\t}\n \n \t\t\tif (all_parents_computed) {\n-\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n-\n-\t\t\t\tdata->generation = max_generation + 1;\n \t\t\t\tpop_commit(&list);\n \n-\t\t\t\tif (data->generation > GENERATION_NUMBER_V1_MAX)\n-\t\t\t\t\tdata->generation = GENERATION_NUMBER_V1_MAX;\n+\t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n+\t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n+\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n \t\t\t}\n \t\t}\n \t}\n@@ -2142,6 +2146,7 @@ int write_commit_graph(struct object_directory *odb,\n \tint res = 0;\n \tint replace = 0;\n \tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n+\tstruct topo_level_slab topo_levels;\n \n \tif (!commit_graph_compatible(the_repository))\n \t\treturn 0;\n@@ -2163,6 +2168,18 @@ int write_commit_graph(struct object_directory *odb,\n \t\t\t\t\t\t\t bloom_settings.max_changed_paths);\n \tctx->bloom_settings = &bloom_settings;\n \n+\tinit_topo_level_slab(&topo_levels);\n+\tctx->topo_levels = &topo_levels;\n+\n+\tif (ctx->r->objects->commit_graph) {\n+\t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n+\n+\t\twhile (g) {\n+\t\t\tg->topo_levels = &topo_levels;\n+\t\t\tg = g->base_graph;\n+\t\t}\n+\t}\n+\n \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n \t\tctx->changed_paths = 1;\n \tif (!(flags & COMMIT_GRAPH_NO_WRITE_BLOOM_FILTERS)) {\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 8be247fa35..2e9aa7824e 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -73,6 +73,7 @@ struct commit_graph {\n \tconst unsigned char *chunk_bloom_indexes;\n \tconst unsigned char *chunk_bloom_data;\n \n+\tstruct topo_level_slab *topo_levels;\n \tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \n-- \ngitgitgadget\n\n"},{"id":"407058","messageId":"011b0aa497d1352bf54ac6a9e2e22ed92d409e64.1602079786.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"[PATCH v4 04/10] commit-graph: return 64-bit generation number","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-10-07T14:09:39Z","receivedAt":"2020-10-07T14:10:04Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a preparatory step, let's return timestamp_t values from\ncommit_graph_generation(), use timestamp_t for local variables and\ndefine GENERATION_NUMBER_INFINITY as (2 ^ 63 - 1) instead.\n\nWe rename GENERATION_NUMBER_MAX to GENERATION_NUMBER_V1_MAX to\nrepresent the largest topological level we can store in the commit data\nchunk.\n\nWith corrected commit dates implemented, we will have two such *_MAX\nvariables to denote the largest offset and largest topological level\nthat can be stored.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 22 +++++++++++-----------\n commit-graph.h |  4 ++--\n commit-reach.c | 36 ++++++++++++++++++------------------\n commit-reach.h |  2 +-\n commit.c       |  4 ++--\n commit.h       |  4 ++--\n revision.c     | 10 +++++-----\n upload-pack.c  |  2 +-\n 8 files changed, 42 insertions(+), 42 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex e8362e144e..bfc532de6f 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -99,7 +99,7 @@ uint32_t commit_graph_position(const struct commit *c)\n \treturn data ? data->graph_pos : COMMIT_NOT_FROM_GRAPH;\n }\n \n-uint32_t commit_graph_generation(const struct commit *c)\n+timestamp_t commit_graph_generation(const struct commit *c)\n {\n \tstruct commit_graph_data *data =\n \t\tcommit_graph_data_slab_peek(&commit_graph_data_slab, c);\n@@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n-\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n+\tconst timestamp_t generation_a = commit_graph_data_at(a)->generation;\n+\tconst timestamp_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n@@ -1350,7 +1350,7 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n+\t\ttimestamp_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n \t\tif (generation != GENERATION_NUMBER_INFINITY &&\n@@ -1383,8 +1383,8 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\tdata->generation = max_generation + 1;\n \t\t\t\tpop_commit(&list);\n \n-\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n-\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n+\t\t\t\tif (data->generation > GENERATION_NUMBER_V1_MAX)\n+\t\t\t\t\tdata->generation = GENERATION_NUMBER_V1_MAX;\n \t\t\t}\n \t\t}\n \t}\n@@ -2404,8 +2404,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \tfor (i = 0; i < g->num_commits; i++) {\n \t\tstruct commit *graph_commit, *odb_commit;\n \t\tstruct commit_list *graph_parents, *odb_parents;\n-\t\tuint32_t max_generation = 0;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t max_generation = 0;\n+\t\ttimestamp_t generation;\n \n \t\tdisplay_progress(progress, i + 1);\n \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n@@ -2469,11 +2469,11 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t\tcontinue;\n \n \t\t/*\n-\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n-\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n+\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n+\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n \t\t * extra logic in the following condition.\n \t\t */\n-\t\tif (max_generation == GENERATION_NUMBER_MAX)\n+\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n \t\t\tmax_generation--;\n \n \t\tgeneration = commit_graph_generation(graph_commit);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex f8e92500c6..8be247fa35 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -144,12 +144,12 @@ void disable_commit_graph(struct repository *r);\n \n struct commit_graph_data {\n \tuint32_t graph_pos;\n-\tuint32_t generation;\n+\ttimestamp_t generation;\n };\n \n /*\n  * Commits should be parsed before accessing generation, graph positions.\n  */\n-uint32_t commit_graph_generation(const struct commit *);\n+timestamp_t commit_graph_generation(const struct commit *);\n uint32_t commit_graph_position(const struct commit *);\n #endif\ndiff --git a/commit-reach.c b/commit-reach.c\nindex 50175b159e..20b48b872b 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -32,12 +32,12 @@ static int queue_has_nonstale(struct prio_queue *queue)\n static struct commit_list *paint_down_to_common(struct repository *r,\n \t\t\t\t\t\tstruct commit *one, int n,\n \t\t\t\t\t\tstruct commit **twos,\n-\t\t\t\t\t\tint min_generation)\n+\t\t\t\t\t\ttimestamp_t min_generation)\n {\n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \tstruct commit_list *result = NULL;\n \tint i;\n-\tuint32_t last_gen = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n \tif (!min_generation)\n \t\tqueue.compare = compare_commits_by_commit_date;\n@@ -58,10 +58,10 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \t\tstruct commit *commit = prio_queue_get(&queue);\n \t\tstruct commit_list *parents;\n \t\tint flags;\n-\t\tuint32_t generation = commit_graph_generation(commit);\n+\t\ttimestamp_t generation = commit_graph_generation(commit);\n \n \t\tif (min_generation && generation > last_gen)\n-\t\t\tBUG(\"bad generation skip %8x > %8x at %s\",\n+\t\t\tBUG(\"bad generation skip %\"PRItime\" > %\"PRItime\" at %s\",\n \t\t\t    generation, last_gen,\n \t\t\t    oid_to_hex(&commit->object.oid));\n \t\tlast_gen = generation;\n@@ -177,12 +177,12 @@ static int remove_redundant(struct repository *r, struct commit **array, int cnt\n \t\trepo_parse_commit(r, array[i]);\n \tfor (i = 0; i < cnt; i++) {\n \t\tstruct commit_list *common;\n-\t\tuint32_t min_generation = commit_graph_generation(array[i]);\n+\t\ttimestamp_t min_generation = commit_graph_generation(array[i]);\n \n \t\tif (redundant[i])\n \t\t\tcontinue;\n \t\tfor (j = filled = 0; j < cnt; j++) {\n-\t\t\tuint32_t curr_generation;\n+\t\t\ttimestamp_t curr_generation;\n \t\t\tif (i == j || redundant[j])\n \t\t\t\tcontinue;\n \t\t\tfilled_index[filled] = j;\n@@ -321,7 +321,7 @@ int repo_in_merge_bases_many(struct repository *r, struct commit *commit,\n {\n \tstruct commit_list *bases;\n \tint ret = 0, i;\n-\tuint32_t generation, max_generation = GENERATION_NUMBER_ZERO;\n+\ttimestamp_t generation, max_generation = GENERATION_NUMBER_INFINITY;\n \n \tif (repo_parse_commit(r, commit))\n \t\treturn ret;\n@@ -470,7 +470,7 @@ static int in_commit_list(const struct commit_list *want, struct commit *c)\n static enum contains_result contains_test(struct commit *candidate,\n \t\t\t\t\t  const struct commit_list *want,\n \t\t\t\t\t  struct contains_cache *cache,\n-\t\t\t\t\t  uint32_t cutoff)\n+\t\t\t\t\t  timestamp_t cutoff)\n {\n \tenum contains_result *cached = contains_cache_at(cache, candidate);\n \n@@ -506,11 +506,11 @@ static enum contains_result contains_tag_algo(struct commit *candidate,\n {\n \tstruct contains_stack contains_stack = { 0, 0, NULL };\n \tenum contains_result result;\n-\tuint32_t cutoff = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t cutoff = GENERATION_NUMBER_INFINITY;\n \tconst struct commit_list *p;\n \n \tfor (p = want; p; p = p->next) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = p->item;\n \t\tload_commit_graph_info(the_repository, c);\n \t\tgeneration = commit_graph_generation(c);\n@@ -566,8 +566,8 @@ static int compare_commits_by_gen(const void *_a, const void *_b)\n \tconst struct commit *a = *(const struct commit * const *)_a;\n \tconst struct commit *b = *(const struct commit * const *)_b;\n \n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n+\ttimestamp_t generation_a = commit_graph_generation(a);\n+\ttimestamp_t generation_b = commit_graph_generation(b);\n \n \tif (generation_a < generation_b)\n \t\treturn -1;\n@@ -580,7 +580,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation)\n+\t\t\t\t timestamp_t min_generation)\n {\n \tstruct commit **list = NULL;\n \tint i;\n@@ -681,13 +681,13 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \ttime_t min_commit_date = cutoff_by_min_date ? from->item->date : 0;\n \tstruct commit_list *from_iter = from, *to_iter = to;\n \tint result;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \n \twhile (from_iter) {\n \t\tadd_object_array(&from_iter->item->object, NULL, &from_objs);\n \n \t\tif (!parse_commit(from_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (from_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = from_iter->item->date;\n \n@@ -701,7 +701,7 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \n \twhile (to_iter) {\n \t\tif (!parse_commit(to_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (to_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = to_iter->item->date;\n \n@@ -741,13 +741,13 @@ struct commit_list *get_reachable_subset(struct commit **from, int nr_from,\n \tstruct commit_list *found_commits = NULL;\n \tstruct commit **to_last = to + nr_to;\n \tstruct commit **from_last = from + nr_from;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \tint num_to_find = 0;\n \n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \n \tfor (item = to; item < to_last; item++) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = *item;\n \n \t\tparse_commit(c);\ndiff --git a/commit-reach.h b/commit-reach.h\nindex b49ad71a31..148b56fea5 100644\n--- a/commit-reach.h\n+++ b/commit-reach.h\n@@ -87,7 +87,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation);\n+\t\t\t\t timestamp_t min_generation);\n int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \t\t       int commit_date_cutoff);\n \ndiff --git a/commit.c b/commit.c\nindex f53429c0ac..3b488381d5 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -731,8 +731,8 @@ int compare_commits_by_author_date(const void *a_, const void *b_,\n int compare_commits_by_gen_then_commit_date(const void *a_, const void *b_, void *unused)\n {\n \tconst struct commit *a = a_, *b = b_;\n-\tconst uint32_t generation_a = commit_graph_generation(a),\n-\t\t       generation_b = commit_graph_generation(b);\n+\tconst timestamp_t generation_a = commit_graph_generation(a),\n+\t\t\t  generation_b = commit_graph_generation(b);\n \n \t/* newer commits first */\n \tif (generation_a < generation_b)\ndiff --git a/commit.h b/commit.h\nindex 5467786c7b..33c66b2177 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -11,8 +11,8 @@\n #include \"commit-slab.h\"\n \n #define COMMIT_NOT_FROM_GRAPH 0xFFFFFFFF\n-#define GENERATION_NUMBER_INFINITY 0xFFFFFFFF\n-#define GENERATION_NUMBER_MAX 0x3FFFFFFF\n+#define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n+#define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n \n struct commit_list {\ndiff --git a/revision.c b/revision.c\nindex c97abcdde1..2861f1c45c 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3308,7 +3308,7 @@ define_commit_slab(indegree_slab, int);\n define_commit_slab(author_date_slab, timestamp_t);\n \n struct topo_walk_info {\n-\tuint32_t min_generation;\n+\ttimestamp_t min_generation;\n \tstruct prio_queue explore_queue;\n \tstruct prio_queue indegree_queue;\n \tstruct prio_queue topo_queue;\n@@ -3354,7 +3354,7 @@ static void explore_walk_step(struct rev_info *revs)\n }\n \n static void explore_to_depth(struct rev_info *revs,\n-\t\t\t     uint32_t gen_cutoff)\n+\t\t\t     timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3397,7 +3397,7 @@ static void indegree_walk_step(struct rev_info *revs)\n }\n \n static void compute_indegrees_to_depth(struct rev_info *revs,\n-\t\t\t\t       uint32_t gen_cutoff)\n+\t\t\t\t       timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3455,7 +3455,7 @@ static void init_topo_walk(struct rev_info *revs)\n \tinfo->min_generation = GENERATION_NUMBER_INFINITY;\n \tfor (list = revs->commits; list; list = list->next) {\n \t\tstruct commit *c = list->item;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (repo_parse_commit_gently(revs->repo, c, 1))\n \t\t\tcontinue;\n@@ -3516,7 +3516,7 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n \tfor (p = commit->parents; p; p = p->next) {\n \t\tstruct commit *parent = p->item;\n \t\tint *pi;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (parent->object.flags & UNINTERESTING)\n \t\t\tcontinue;\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 3b858eb457..fdb82885b6 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -497,7 +497,7 @@ static int got_oid(struct upload_pack_data *data,\n \n static int ok_to_give_up(struct upload_pack_data *data)\n {\n-\tuint32_t min_generation = GENERATION_NUMBER_ZERO;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_ZERO;\n \n \tif (!data->have_obj.nr)\n \t\treturn 0;\n-- \ngitgitgadget\n\n"},{"id":"407059","messageId":"9ada43967d29a3ec717b6a8db0de5b09e6d916b1.1602079786.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"[PATCH v4 10/10] doc: add corrected commit date info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-10-07T14:09:45Z","receivedAt":"2020-10-07T14:10:06Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith generation data chunk and corrected commit dates implemented, let's\nupdate the technical documentation for commit-graph.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n .../technical/commit-graph-format.txt         | 21 +++++--\n Documentation/technical/commit-graph.txt      | 62 ++++++++++++++++---\n 2 files changed, 69 insertions(+), 14 deletions(-)\n\ndiff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\nindex b3b58880b9..08d9026ad4 100644\n--- a/Documentation/technical/commit-graph-format.txt\n+++ b/Documentation/technical/commit-graph-format.txt\n@@ -4,11 +4,7 @@ Git commit graph format\n The Git commit graph stores a list of commit OIDs and some associated\n metadata, including:\n \n-- The generation number of the commit. Commits with no parents have\n-  generation number 1; commits with parents have generation number\n-  one more than the maximum generation number of its parents. We\n-  reserve zero as special, and can be used to mark a generation\n-  number invalid or as \"not computed\".\n+- The generation number of the commit.\n \n - The root tree OID.\n \n@@ -86,13 +82,26 @@ CHUNK DATA:\n       position. If there are more than two parents, the second value\n       has its most-significant bit on and the other bits store an array\n       position into the Extra Edge List chunk.\n-    * The next 8 bytes store the generation number of the commit and\n+    * The next 8 bytes store the topological level (generation number v1)\n+      of the commit and\n       the commit time in seconds since EPOCH. The generation number\n       uses the higher 30 bits of the first 4 bytes, while the commit\n       time uses the 32 bits of the second 4 bytes, along with the lowest\n       2 bits of the lowest byte, storing the 33rd and 34th bit of the\n       commit time.\n \n+  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes)\n+    * This list of 4-byte values store corrected commit date offsets for the\n+      commits, arranged in the same order as commit data chunk.\n+    * If the corrected commit date offset cannot be stored within 31 bits,\n+      the value has its most-significant bit on and the other bits store\n+      the position of corrected commit date into the Generation Data Overflow\n+      chunk.\n+\n+  Generation Data Overflow (ID: {'G', 'D', 'O', 'V' }) [Optional]\n+    * This list of 8-byte values stores the corrected commit dates for commits\n+      with corrected commit date offsets that cannot be stored within 31 bits.\n+\n   Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n       This list of 4-byte values store the second through nth parents for\n       all octopus merges. The second parent value in the commit data stores\ndiff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\nindex f14a7659aa..75f71c4c7b 100644\n--- a/Documentation/technical/commit-graph.txt\n+++ b/Documentation/technical/commit-graph.txt\n@@ -38,14 +38,31 @@ A consumer may load the following info for a commit from the graph:\n \n Values 1-4 satisfy the requirements of parse_commit_gently().\n \n-Define the \"generation number\" of a commit recursively as follows:\n+There are two definitions of generation number:\n+1. Corrected committer dates (generation number v2)\n+2. Topological levels (generation nummber v1)\n \n- * A commit with no parents (a root commit) has generation number one.\n+Define \"corrected committer date\" of a commit recursively as follows:\n \n- * A commit with at least one parent has generation number one more than\n-   the largest generation number among its parents.\n+  * A commit with no parents (a root commit) has corrected committer date\n+    equal to its committer date.\n \n-Equivalently, the generation number of a commit A is one more than the\n+  * A commit with at least one parent has corrected committer date equal to\n+    the maximum of its commiter date and one more than the largest corrected\n+    committer date among its parents.\n+\n+  * As a special case, a root commit with timestamp zero has corrected commit\n+    date of 1, to be able to distinguish it from GENERATION_NUMBER_ZERO\n+    (that is, an uncomputed corrected commit date).\n+\n+Define the \"topological level\" of a commit recursively as follows:\n+\n+ * A commit with no parents (a root commit) has topological level of one.\n+\n+ * A commit with at least one parent has topological level one more than\n+   the largest topological level among its parents.\n+\n+Equivalently, the topological level of a commit A is one more than the\n length of a longest path from A to a root commit. The recursive definition\n is easier to use for computation and observing the following property:\n \n@@ -60,6 +77,9 @@ is easier to use for computation and observing the following property:\n     generation numbers, then we always expand the boundary commit with highest\n     generation number and can easily detect the stopping condition.\n \n+The properties applies to both versions of generation number, that is both\n+corrected committer dates and topological levels.\n+\n This property can be used to significantly reduce the time it takes to\n walk commits and determine topological relationships. Without generation\n numbers, the general heuristic is the following:\n@@ -67,7 +87,9 @@ numbers, the general heuristic is the following:\n     If A and B are commits with commit time X and Y, respectively, and\n     X < Y, then A _probably_ cannot reach B.\n \n-This heuristic is currently used whenever the computation is allowed to\n+In absence of corrected commit dates (for example, old versions of Git or\n+mixed generation graph chains),\n+this heuristic is currently used whenever the computation is allowed to\n violate topological relationships due to clock skew (such as \"git log\"\n with default order), but is not used when the topological order is\n required (such as merge base calculations, \"git log --graph\").\n@@ -77,7 +99,7 @@ in the commit graph. We can treat these commits as having \"infinite\"\n generation number and walk until reaching commits with known generation\n number.\n \n-We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n+We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n in the commit-graph file. If a commit-graph file was written by a version\n of Git that did not compute generation numbers, then those commits will\n have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n@@ -93,7 +115,7 @@ fully-computed generation numbers. Using strict inequality may result in\n walking a few extra commits, but the simplicity in dealing with commits\n with generation number *_INFINITY or *_ZERO is valuable.\n \n-We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n+We use the macro GENERATION_NUMBER_MAX for commits whose\n generation numbers are computed to be at least this value. We limit at\n this value since it is the largest value that can be stored in the\n commit-graph file using the 30 bits available to generation numbers. This\n@@ -267,6 +289,30 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n number of commits) could be extracted into config settings for full\n flexibility.\n \n+## Handling Mixed Generation Number Chains\n+\n+With the introduction of generation number v2 and generation data chunk, the\n+following scenario is possible:\n+\n+1. \"New\" Git writes a commit-graph with the corrected commit dates.\n+2. \"Old\" Git writes a split commit-graph on top without corrected commit dates.\n+\n+A naive approach of using the newest available generation number from\n+each layer would lead to violated expectations: the lower layer would\n+use corrected commit dates which are much larger than the topological\n+levels of the higher layer. For this reason, Git inspects each layer to\n+see if any layer is missing corrected commit dates. In such a case, Git\n+only uses topological level\n+\n+When writing a new layer in split commit-graph, we write corrected commit\n+dates if the topmost layer has corrected commit dates written. This\n+guarantees that if a layer has corrected commit dates, all lower layers\n+must have corrected commit dates as well.\n+\n+When merging layers, we do not consider whether the merged layers had corrected\n+commit dates. Instead, the new layer will have corrected commit dates if and\n+only if all existing layers below the new layer have corrected commit dates.\n+\n ## Deleting graph-{hash} files\n \n After a new tip file is written, some `graph-{hash}` files may no longer\n-- \ngitgitgadget\n"},{"id":"407061","messageId":"bb9b02af32d028fc0c26d372aa490e260c74e74d.1602079786.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"[PATCH v4 09/10] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-10-07T14:09:44Z","receivedAt":"2020-10-07T14:10:07Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith corrected commit dates implemented, we no longer have to rely on\ncommit date as a heuristic in paint_down_to_common().\n\nWhile using corrected commit dates Git walks nearly the same number of\ncommits as commit date, the process is slower as for each comparision we\nhave to access a commit-slab (for corrected committer date) instead of\naccessing struct member (for committer date).\n\nFor example, the command `git merge-base v4.8 v4.9` on the linux\nrepository walks 167468 commits, taking 0.135s for committer date and\n167496 commits, taking 0.157s for corrected committer date respectively.\n\nt6404-recursive-merge setups a unique repository where all commits have\nthe same committer date without well-defined merge-base.\n\nWhile running tests with GIT_TEST_COMMIT_GRAPH unset, we use committer\ndate as a heuristic in paint_down_to_common(). 6404.1 'combined merge\nconflicts' merges commits in the order:\n- Merge C with B to form a intermediate commit.\n- Merge the intermediate commit with A.\n\nWith GIT_TEST_COMMIT_GRAPH=1, we write a commit-graph and subsequently\nuse the corrected committer date, which changes the order in which\ncommits are merged:\n- Merge A with B to form a intermediate commit.\n- Merge the intermediate commit with C.\n\nWhile resulting repositories are equivalent, 6404.4 'virtual trees were\nprocessed' fails with GIT_TEST_COMMIT_GRAPH=1 as we are selecting\ndifferent merge-bases and thus have different object ids for the\nintermediate commits.\n\nAs this has already causes problems (as noted in 859fdc0 (commit-graph:\ndefine GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph\nwithin t6404-recursive-merge.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c             | 14 ++++++++++++++\n commit-graph.h             |  8 +++++++-\n commit-reach.c             |  2 +-\n t/t6404-recursive-merge.sh |  5 ++++-\n 4 files changed, 26 insertions(+), 3 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 5d15a1399b..3de1933ede 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -705,6 +705,20 @@ int generation_numbers_enabled(struct repository *r)\n \treturn !!first_generation;\n }\n \n+int corrected_commit_dates_enabled(struct repository *r)\n+{\n+\tstruct commit_graph *g;\n+\tif (!prepare_commit_graph(r))\n+\t\treturn 0;\n+\n+\tg = r->objects->commit_graph;\n+\n+\tif (!g->num_commits)\n+\t\treturn 0;\n+\n+\treturn g->read_generation_data;\n+}\n+\n struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r)\n {\n \tstruct commit_graph *g = r->objects->commit_graph;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex ad52130883..d2c048dc64 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -89,13 +89,19 @@ struct commit_graph *read_commit_graph_one(struct repository *r,\n struct commit_graph *parse_commit_graph(struct repository *r,\n \t\t\t\t\tvoid *graph_map, size_t graph_size);\n \n+struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n+\n /*\n  * Return 1 if and only if the repository has a commit-graph\n  * file and generation numbers are computed in that file.\n  */\n int generation_numbers_enabled(struct repository *r);\n \n-struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n+/*\n+ * Return 1 if and only if the repository has a commit-graph\n+ * file and generation data chunk has been written for the file.\n+ */\n+int corrected_commit_dates_enabled(struct repository *r);\n \n enum commit_graph_write_flags {\n \tCOMMIT_GRAPH_WRITE_APPEND     = (1 << 0),\ndiff --git a/commit-reach.c b/commit-reach.c\nindex 20b48b872b..46f5a9e638 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -39,7 +39,7 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \tint i;\n \ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n-\tif (!min_generation)\n+\tif (!min_generation && !corrected_commit_dates_enabled(r))\n \t\tqueue.compare = compare_commits_by_commit_date;\n \n \tone->object.flags |= PARENT1;\ndiff --git a/t/t6404-recursive-merge.sh b/t/t6404-recursive-merge.sh\nindex 332cfc53fd..7055771b62 100755\n--- a/t/t6404-recursive-merge.sh\n+++ b/t/t6404-recursive-merge.sh\n@@ -15,6 +15,8 @@ GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n export GIT_COMMITTER_DATE\n \n test_expect_success 'setup tests' '\n+\tGIT_TEST_COMMIT_GRAPH=0 &&\n+\texport GIT_TEST_COMMIT_GRAPH &&\n \techo 1 >a1 &&\n \tgit add a1 &&\n \tGIT_AUTHOR_DATE=\"2006-12-12 23:00:00\" git commit -m 1 a1 &&\n@@ -66,7 +68,7 @@ test_expect_success 'setup tests' '\n '\n \n test_expect_success 'combined merge conflicts' '\n-\ttest_must_fail env GIT_TEST_COMMIT_GRAPH=0 git merge -m final G\n+\ttest_must_fail git merge -m final G\n '\n \n test_expect_success 'result contains a conflict' '\n@@ -82,6 +84,7 @@ test_expect_success 'result contains a conflict' '\n '\n \n test_expect_success 'virtual trees were processed' '\n+\t# TODO: fragile test, relies on ambigious merge-base resolution\n \tgit ls-files --stage >out &&\n \n \tcat >expect <<-EOF &&\n-- \ngitgitgadget\n\n"},{"id":"407063","messageId":"b903efe2ea11bc0b7e1ef8f239ed34f72caa4f03.1602079786.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"[PATCH v4 07/10] commit-graph: implement generation data chunk","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-10-07T14:09:42Z","receivedAt":"2020-10-07T14:10:07Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nAs discovered by Ævar, we cannot increment graph version to\ndistinguish between generation numbers v1 and v2 [1]. Thus, one of\npre-requistes before implementing generation number was to distinguish\nbetween graph versions in a backwards compatible manner.\n\nWe are going to introduce a new chunk called Generation Data chunk (or\nGDAT). GDAT stores corrected committer date offsets whereas CDAT will\nstill store topological level.\n\nOld Git does not understand GDAT chunk and would ignore it, reading\ntopological levels from CDAT. New Git can parse GDAT and take advantage\nof newer generation numbers, falling back to topological levels when\nGDAT chunk is missing (as it would happen with a commit graph written\nby old Git).\n\nWe introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\nwhich forces commit-graph file to be written without generation data\nchunk to emulate a commit-graph file written by old Git.\n\nWhile storing corrected commit date offset instead of the corrected\ncommit date saves us 4 bytes per commit, it's possible for the offsets\nto overflow the 4-bytes allocated. As such overflows are exceedingly\nrare, we use the following overflow management scheme:\n\nWe introduce a new commit-graph chunk, GENERATION_DATA_OVERFLOW ('GDOV')\nto store corrected commit dates for commits with offsets greater than\nGENERATION_NUMBER_V2_OFFSET_MAX.\n\nIf the offset is greater than GENERATION_NUMBER_V2_OFFSET_MAX, we set\nthe MSB of the offset and the other bits store the position of corrected\ncommit date in GDOV chunk, similar to how Extra Edge List is maintained.\n\nWe test the overflow-related code with the following repo history:\n\n           F - N - U\n          /         \\\nU - N - U            N\n         \\          /\n\t  N - F - N\n\nWhere the commits denoted by U have committer date of zero seconds\nsince Unix epoch, the commits denoted by N have committer date of\n1112354055 (default committer date for the test suite) seconds since\nUnix epoch and the commits denoted by F have committer date of\n(2 ^ 31 - 2) seconds since Unix epoch.\n\nThe largest offset observed is 2 ^ 31, just large enough to overflow.\n\n[1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                | 98 +++++++++++++++++++++++++++++++++--\n commit-graph.h                |  3 ++\n commit.h                      |  1 +\n t/README                      |  3 ++\n t/helper/test-read-graph.c    |  4 ++\n t/t4216-log-bloom.sh          |  4 +-\n t/t5318-commit-graph.sh       | 70 ++++++++++++++++++++-----\n t/t5324-split-commit-graph.sh | 12 ++---\n t/t6600-test-reach.sh         | 68 +++++++++++++-----------\n 9 files changed, 206 insertions(+), 57 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 03948adfce..71d0b243db 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -38,11 +38,13 @@ void git_test_write_commit_graph_or_die(void)\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n+#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n+#define GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW 0x47444f56 /* \"GDOV\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n-#define MAX_NUM_CHUNKS 7\n+#define MAX_NUM_CHUNKS 9\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -61,6 +63,8 @@ void git_test_write_commit_graph_or_die(void)\n #define GRAPH_MIN_SIZE (GRAPH_HEADER_SIZE + 4 * GRAPH_CHUNKLOOKUP_WIDTH \\\n \t\t\t+ GRAPH_FANOUT_SIZE + the_hash_algo->rawsz)\n \n+#define CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW (1ULL << 31)\n+\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n@@ -385,6 +389,20 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n \t\t\tbreak;\n \n+\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n+\t\t\tif (graph->chunk_generation_data)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW:\n+\t\t\tif (graph->chunk_generation_data_overflow)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_generation_data_overflow = data + chunk_offset;\n+\t\t\tbreak;\n+\n \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n \t\t\tif (graph->chunk_extra_edges)\n \t\t\t\tchunk_repeated = 1;\n@@ -745,8 +763,8 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n {\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n-\tuint32_t lex_index;\n-\tuint64_t date_high, date_low;\n+\tuint32_t lex_index, offset_pos;\n+\tuint64_t date_high, date_low, offset;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n@@ -764,7 +782,16 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\tif (g->chunk_generation_data) {\n+\t\toffset = (timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n+\n+\t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n+\t\t\toffset_pos = offset ^ CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW;\n+\t\t\tgraph_data->generation = get_be64(g->chunk_generation_data_overflow + 8 * offset_pos);\n+\t\t} else\n+\t\t\tgraph_data->generation = item->date + offset;\n+\t} else\n+\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n \n \tif (g->topo_levels)\n \t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n@@ -942,6 +969,7 @@ struct write_commit_graph_context {\n \tstruct packed_oid_list oids;\n \tstruct packed_commit_list commits;\n \tint num_extra_edges;\n+\tint num_generation_data_overflows;\n \tunsigned long approx_nr_objects;\n \tstruct progress *progress;\n \tint progress_done;\n@@ -960,7 +988,8 @@ struct write_commit_graph_context {\n \t\t report_progress:1,\n \t\t split:1,\n \t\t changed_paths:1,\n-\t\t order_by_pack:1;\n+\t\t order_by_pack:1,\n+\t\t write_generation_data:1;\n \n \tstruct topo_level_slab *topo_levels;\n \tconst struct commit_graph_opts *opts;\n@@ -1120,6 +1149,44 @@ static int write_graph_chunk_data(struct hashfile *f,\n \treturn 0;\n }\n \n+static int write_graph_chunk_generation_data(struct hashfile *f,\n+\t\t\t\t\t      struct write_commit_graph_context *ctx)\n+{\n+\tint i, num_generation_data_overflows = 0;\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n+\n+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n+\t\t\toffset = CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW | num_generation_data_overflows;\n+\t\t\tnum_generation_data_overflows++;\n+\t\t}\n+\n+\t\thashwrite_be32(f, offset);\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int write_graph_chunk_generation_data_overflow(struct hashfile *f,\n+\t\t\t\t\t\t       struct write_commit_graph_context *ctx)\n+{\n+\tint i;\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n+\n+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n+\t\t\thashwrite_be32(f, offset >> 32);\n+\t\t\thashwrite_be32(f, (uint32_t) offset);\n+\t\t}\n+\t}\n+\n+\treturn 0;\n+}\n+\n static int write_graph_chunk_extra_edges(struct hashfile *f,\n \t\t\t\t\t struct write_commit_graph_context *ctx)\n {\n@@ -1399,7 +1466,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \n \t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n \t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n+\n \t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n+\n+\t\t\t\tif (commit_graph_data_at(current)->generation - current->date > GENERATION_NUMBER_V2_OFFSET_MAX)\n+\t\t\t\t\tctx->num_generation_data_overflows++;\n \t\t\t}\n \t\t}\n \t}\n@@ -1765,6 +1836,21 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tchunks[2].id = GRAPH_CHUNKID_DATA;\n \tchunks[2].size = (hashsz + 16) * ctx->commits.nr;\n \tchunks[2].write_fn = write_graph_chunk_data;\n+\n+\tif (git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0))\n+\t\tctx->write_generation_data = 0;\n+\tif (ctx->write_generation_data) {\n+\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA;\n+\t\tchunks[num_chunks].size = sizeof(uint32_t) * ctx->commits.nr;\n+\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data;\n+\t\tnum_chunks++;\n+\t}\n+\tif (ctx->num_generation_data_overflows) {\n+\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW;\n+\t\tchunks[num_chunks].size = sizeof(timestamp_t) * ctx->num_generation_data_overflows;\n+\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data_overflow;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_extra_edges) {\n \t\tchunks[num_chunks].id = GRAPH_CHUNKID_EXTRAEDGES;\n \t\tchunks[num_chunks].size = 4 * ctx->num_extra_edges;\n@@ -2170,6 +2256,8 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n \tctx->opts = opts;\n \tctx->total_bloom_filter_data_size = 0;\n+\tctx->write_generation_data = 1;\n+\tctx->num_generation_data_overflows = 0;\n \n \tbloom_settings.bits_per_entry = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\",\n \t\t\t\t\t\t      bloom_settings.bits_per_entry);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 2e9aa7824e..19a02001fd 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -6,6 +6,7 @@\n #include \"oidset.h\"\n \n #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n+#define GIT_TEST_COMMIT_GRAPH_NO_GDAT \"GIT_TEST_COMMIT_GRAPH_NO_GDAT\"\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE \"GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE\"\n #define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n \n@@ -68,6 +69,8 @@ struct commit_graph {\n \tconst uint32_t *chunk_oid_fanout;\n \tconst unsigned char *chunk_oid_lookup;\n \tconst unsigned char *chunk_commit_data;\n+\tconst unsigned char *chunk_generation_data;\n+\tconst unsigned char *chunk_generation_data_overflow;\n \tconst unsigned char *chunk_extra_edges;\n \tconst unsigned char *chunk_base_graphs;\n \tconst unsigned char *chunk_bloom_indexes;\ndiff --git a/commit.h b/commit.h\nindex 33c66b2177..251d877fcf 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -14,6 +14,7 @@\n #define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n #define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n+#define GENERATION_NUMBER_V2_OFFSET_MAX ((1ULL << 31) - 1)\n \n struct commit_list {\n \tstruct commit *item;\ndiff --git a/t/README b/t/README\nindex 2adaf7c2d2..975c054bc9 100644\n--- a/t/README\n+++ b/t/README\n@@ -379,6 +379,9 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n+commit-graph to be written without generation data chunk.\n+\n GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n commit-graph write to compute and write changed path Bloom filters for\n every 'git commit-graph write', as if the `--changed-paths` option was\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex 5f585a1725..75927b2c81 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -33,6 +33,10 @@ int cmd__read_graph(int argc, const char **argv)\n \t\tprintf(\" oid_lookup\");\n \tif (graph->chunk_commit_data)\n \t\tprintf(\" commit_metadata\");\n+\tif (graph->chunk_generation_data)\n+\t\tprintf(\" generation_data\");\n+\tif (graph->chunk_generation_data_overflow)\n+\t\tprintf(\" generation_data_overflow\");\n \tif (graph->chunk_extra_edges)\n \t\tprintf(\" extra_edges\");\n \tif (graph->chunk_bloom_indexes)\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex d11040ce41..dbde016188 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -40,11 +40,11 @@ test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n '\n \n graph_read_expect () {\n-\tNUM_CHUNKS=5\n+\tNUM_CHUNKS=6\n \tcat >expect <<- EOF\n \theader: 43475048 1 $(test_oid oid_version) $NUM_CHUNKS 0\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n \tEOF\n \ttest-tool read-graph >actual &&\n \ttest_cmp expect actual\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 2ed0c1544d..0328e98564 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -76,7 +76,7 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n graph_read_expect() {\n \tOPTIONAL=\"\"\n \tNUM_CHUNKS=3\n-\tif test ! -z $2\n+\tif test ! -z \"$2\"\n \tthen\n \t\tOPTIONAL=\" $2\"\n \t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n@@ -103,14 +103,14 @@ test_expect_success 'exit with correct error on bad input to --stdin-commits' '\n \t# valid commit and tree OID\n \tgit rev-parse HEAD HEAD^{tree} >in &&\n \tgit commit-graph write --stdin-commits <in &&\n-\tgraph_read_expect 3\n+\tgraph_read_expect 3 generation_data\n '\n \n test_expect_success 'write graph' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"3\"\n+\tgraph_read_expect \"3\" generation_data\n '\n \n test_expect_success POSIXPERM 'write graph has correct permissions' '\n@@ -219,7 +219,7 @@ test_expect_success 'write graph with merges' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'merge 1 vs 2' full merge/1 merge/2\n@@ -254,7 +254,7 @@ test_expect_success 'write graph with new commit' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'full graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -264,7 +264,7 @@ test_expect_success 'write graph with nothing new' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'cleared graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -274,7 +274,7 @@ test_expect_success 'build graph from latest pack with closure' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tcat new-idx | git commit-graph write --stdin-packs &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"9\" \"extra_edges\"\n+\tgraph_read_expect \"9\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'graph from pack, commit 8 vs merge 1' full commits/8 merge/1\n@@ -287,7 +287,7 @@ test_expect_success 'build graph from commits with closure' '\n \tgit rev-parse merge/1 >>commits-in &&\n \tcat commits-in | git commit-graph write --stdin-commits &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"6\"\n+\tgraph_read_expect \"6\" \"generation_data\"\n '\n \n graph_git_behavior 'graph from commits, commit 8 vs merge 1' full commits/8 merge/1\n@@ -297,7 +297,7 @@ test_expect_success 'build graph from commits with append' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit rev-parse merge/3 | git commit-graph write --stdin-commits --append &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -307,7 +307,7 @@ test_expect_success 'build graph using --reachable' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write --reachable &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -328,7 +328,7 @@ test_expect_success 'write graph in bare repo' '\n \tcd \"$TRASH_DIRECTORY/bare\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $baredir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n@@ -454,8 +454,9 @@ test_expect_success 'warn on improper hash version' '\n \n test_expect_success 'git commit-graph verify' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n-\tgit rev-parse commits/8 | git commit-graph write --stdin-commits &&\n-\tgit commit-graph verify >output\n+\tgit rev-parse commits/8 | GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --stdin-commits &&\n+\tgit commit-graph verify >output &&\n+\tgraph_read_expect 9 extra_edges\n '\n \n NUM_COMMITS=9\n@@ -741,4 +742,47 @@ test_expect_success 'corrupt commit-graph write (missing tree)' '\n \t)\n '\n \n+test_commit_with_date() {\n+  file=\"$1.t\" &&\n+  echo \"$1\" >\"$file\" &&\n+  git add \"$file\" &&\n+  GIT_COMMITTER_DATE=\"$2\" GIT_AUTHOR_DATE=\"$2\" git commit -m \"$1\"\n+  git tag \"$1\"\n+}\n+\n+test_expect_success 'overflow corrected commit date offset' '\n+\tobjdir=\".git/objects\" &&\n+\tUNIX_EPOCH_ZERO=\"1970-01-01 00:00 +0000\" &&\n+\tFUTURE_DATE=\"@2147483646 +0000\" &&\n+\ttest_oid_cache <<-EOF &&\n+\toid_version sha1:1\n+\toid_version sha256:2\n+\tEOF\n+\tcd \"$TRASH_DIRECTORY\" &&\n+\tmkdir repo &&\n+\tcd repo &&\n+\tgit init &&\n+\ttest_commit_with_date 1 \"$UNIX_EPOCH_ZERO\" &&\n+\ttest_commit 2 &&\n+\ttest_commit_with_date 3 \"$UNIX_EPOCH_ZERO\" &&\n+\tgit commit-graph write --reachable &&\n+\tgraph_read_expect 3 generation_data &&\n+\ttest_commit_with_date 4 \"$FUTURE_DATE\" &&\n+\ttest_commit 5 &&\n+\ttest_commit_with_date 6 \"$UNIX_EPOCH_ZERO\" &&\n+\tgit branch left &&\n+\tgit reset --hard 3 &&\n+\ttest_commit 7 &&\n+\ttest_commit_with_date 8 \"$FUTURE_DATE\" &&\n+\ttest_commit 9 &&\n+\tgit branch right &&\n+\tgit reset --hard 3 &&\n+\tgit merge left right &&\n+\tgit commit-graph write --reachable &&\n+\tgraph_read_expect 10 \"generation_data generation_data_overflow\" &&\n+\tgit commit-graph verify\n+'\n+\n+graph_git_behavior 'overflow corrected commit date offset' repo left right\n+\n test_done\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex c334ee9155..651df89ab2 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -13,11 +13,11 @@ test_expect_success 'setup repo' '\n \tinfodir=\".git/objects/info\" &&\n \tgraphdir=\"$infodir/commit-graphs\" &&\n \ttest_oid_cache <<-EOM\n-\tshallow sha1:1760\n-\tshallow sha256:2064\n+\tshallow sha1:2132\n+\tshallow sha256:2436\n \n-\tbase sha1:1376\n-\tbase sha256:1496\n+\tbase sha1:1408\n+\tbase sha256:1528\n \n \toid_version sha1:1\n \toid_version sha256:2\n@@ -31,9 +31,9 @@ graph_read_expect() {\n \t\tNUM_BASE=$2\n \tfi\n \tcat >expect <<- EOF\n-\theader: 43475048 1 $(test_oid oid_version) 3 $NUM_BASE\n+\theader: 43475048 1 $(test_oid oid_version) 4 $NUM_BASE\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n \tEOF\n \ttest-tool read-graph >output &&\n \ttest_cmp expect output\ndiff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\nindex f807276337..e2d33a8a4c 100755\n--- a/t/t6600-test-reach.sh\n+++ b/t/t6600-test-reach.sh\n@@ -55,10 +55,13 @@ test_expect_success 'setup' '\n \tgit show-ref -s commit-5-5 | git commit-graph write --stdin-commits &&\n \tmv .git/objects/info/commit-graph commit-graph-half &&\n \tchmod u+w commit-graph-half &&\n+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable &&\n+\tmv .git/objects/info/commit-graph commit-graph-no-gdat &&\n+\tchmod u+w commit-graph-no-gdat &&\n \tgit config core.commitGraph true\n '\n \n-run_three_modes () {\n+run_all_modes () {\n \ttest_when_finished rm -rf .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n \ttest_cmp expect actual &&\n@@ -67,11 +70,14 @@ run_three_modes () {\n \ttest_cmp expect actual &&\n \tcp commit-graph-half .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n+\ttest_cmp expect actual &&\n+\tcp commit-graph-no-gdat .git/objects/info/commit-graph &&\n+\t\"$@\" <input >actual &&\n \ttest_cmp expect actual\n }\n \n-test_three_modes () {\n-\trun_three_modes test-tool reach \"$@\"\n+test_all_modes () {\n+\trun_all_modes test-tool reach \"$@\"\n }\n \n test_expect_success 'ref_newer:miss' '\n@@ -80,7 +86,7 @@ test_expect_success 'ref_newer:miss' '\n \tB:commit-4-9\n \tEOF\n \techo \"ref_newer(A,B):0\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'ref_newer:hit' '\n@@ -89,7 +95,7 @@ test_expect_success 'ref_newer:hit' '\n \tB:commit-2-3\n \tEOF\n \techo \"ref_newer(A,B):1\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'in_merge_bases:hit' '\n@@ -98,7 +104,7 @@ test_expect_success 'in_merge_bases:hit' '\n \tB:commit-8-8\n \tEOF\n \techo \"in_merge_bases(A,B):1\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'in_merge_bases:miss' '\n@@ -107,7 +113,7 @@ test_expect_success 'in_merge_bases:miss' '\n \tB:commit-5-9\n \tEOF\n \techo \"in_merge_bases(A,B):0\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'in_merge_bases_many:hit' '\n@@ -117,7 +123,7 @@ test_expect_success 'in_merge_bases_many:hit' '\n \tX:commit-5-7\n \tEOF\n \techo \"in_merge_bases_many(A,X):1\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'in_merge_bases_many:miss' '\n@@ -127,7 +133,7 @@ test_expect_success 'in_merge_bases_many:miss' '\n \tX:commit-8-6\n \tEOF\n \techo \"in_merge_bases_many(A,X):0\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'in_merge_bases_many:miss-heuristic' '\n@@ -137,7 +143,7 @@ test_expect_success 'in_merge_bases_many:miss-heuristic' '\n \tX:commit-6-6\n \tEOF\n \techo \"in_merge_bases_many(A,X):0\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'is_descendant_of:hit' '\n@@ -148,7 +154,7 @@ test_expect_success 'is_descendant_of:hit' '\n \tX:commit-1-1\n \tEOF\n \techo \"is_descendant_of(A,X):1\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'is_descendant_of:miss' '\n@@ -159,7 +165,7 @@ test_expect_success 'is_descendant_of:miss' '\n \tX:commit-7-6\n \tEOF\n \techo \"is_descendant_of(A,X):0\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'get_merge_bases_many' '\n@@ -174,7 +180,7 @@ test_expect_success 'get_merge_bases_many' '\n \t\tgit rev-parse commit-5-6 \\\n \t\t\t      commit-4-7 | sort\n \t} >expect &&\n-\ttest_three_modes get_merge_bases_many\n+\ttest_all_modes get_merge_bases_many\n '\n \n test_expect_success 'reduce_heads' '\n@@ -196,7 +202,7 @@ test_expect_success 'reduce_heads' '\n \t\t\t      commit-2-8 \\\n \t\t\t      commit-1-10 | sort\n \t} >expect &&\n-\ttest_three_modes reduce_heads\n+\ttest_all_modes reduce_heads\n '\n \n test_expect_success 'can_all_from_reach:hit' '\n@@ -219,7 +225,7 @@ test_expect_success 'can_all_from_reach:hit' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach(X,Y):1\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach:miss' '\n@@ -241,7 +247,7 @@ test_expect_success 'can_all_from_reach:miss' '\n \tY:commit-8-5\n \tEOF\n \techo \"can_all_from_reach(X,Y):0\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach_with_flag: tags case' '\n@@ -264,7 +270,7 @@ test_expect_success 'can_all_from_reach_with_flag: tags case' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach_with_flag(X,_,_,0,0):1\" >expect &&\n-\ttest_three_modes can_all_from_reach_with_flag\n+\ttest_all_modes can_all_from_reach_with_flag\n '\n \n test_expect_success 'commit_contains:hit' '\n@@ -280,8 +286,8 @@ test_expect_success 'commit_contains:hit' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):1\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'commit_contains:miss' '\n@@ -297,8 +303,8 @@ test_expect_success 'commit_contains:miss' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):0\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'rev-list: basic topo-order' '\n@@ -310,7 +316,7 @@ test_expect_success 'rev-list: basic topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 commit-3-2 commit-2-2 commit-1-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent topo-order' '\n@@ -322,7 +328,7 @@ test_expect_success 'rev-list: first-parent topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -334,7 +340,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -346,7 +352,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent range topo-order' '\n@@ -358,7 +364,7 @@ test_expect_success 'rev-list: first-parent range topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: ancestry-path topo-order' '\n@@ -368,7 +374,7 @@ test_expect_success 'rev-list: ancestry-path topo-order' '\n \t\tcommit-6-4 commit-5-4 commit-4-4 commit-3-4 \\\n \t\tcommit-6-3 commit-5-3 commit-4-3 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: symmetric difference topo-order' '\n@@ -382,7 +388,7 @@ test_expect_success 'rev-list: symmetric difference topo-order' '\n \t\tcommit-3-8 commit-2-8 commit-1-8 \\\n \t\tcommit-3-7 commit-2-7 commit-1-7 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8...commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8...commit-6-6\n '\n \n test_expect_success 'get_reachable_subset:all' '\n@@ -402,7 +408,7 @@ test_expect_success 'get_reachable_subset:all' '\n \t\t\t      commit-1-7 \\\n \t\t\t      commit-5-6 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:some' '\n@@ -420,7 +426,7 @@ test_expect_success 'get_reachable_subset:some' '\n \t\tgit rev-parse commit-3-3 \\\n \t\t\t      commit-1-7 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:none' '\n@@ -434,7 +440,7 @@ test_expect_success 'get_reachable_subset:none' '\n \tY:commit-2-8\n \tEOF\n \techo \"get_reachable_subset(X,Y)\" >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_done\n-- \ngitgitgadget\n\n"},{"id":"407060","messageId":"8ec119edc66814ad4d63908c79437a7f9dd3c08c.1602079786.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"[PATCH v4 08/10] commit-graph: use generation v2 only if entire chain does","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-10-07T14:09:43Z","receivedAt":"2020-10-07T14:10:08Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nSince there are released versions of Git that understand generation\nnumbers in the commit-graph's CDAT chunk but do not understand the GDAT\nchunk, the following scenario is possible:\n\n1. \"New\" Git writes a commit-graph with the GDAT chunk.\n2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n\nBecause of the current use of inspecting the current layer for a\nchunk_generation_data pointer, the commits in the lower layer will be\ninterpreted as having very large generation values (commit date plus\noffset) compared to the generation numbers in the top layer (topological\nlevel). This violates the expectation that the generation of a parent is\nstrictly smaller than the generation of a child.\n\nIt is difficult to expose this issue in a test. Since we _start_ with\nartificially low generation numbers, any commit walk that prioritizes\ngeneration numbers will walk all of the commits with high generation\nnumber before walking the commits with low generation number. In all the\ncases I tried, the commit-graph layers themselves \"protect\" any\nincorrect behavior since none of the commits in the lower layer can\nreach the commits in the upper layer.\n\nThis issue would manifest itself as a performance problem in this case,\nespecially with something like \"git log --graph\" since the low\ngeneration numbers would cause the in-degree queue to walk all of the\ncommits in the lower layer before allowing the topo-order queue to write\nanything to output (depending on the size of the upper layer).\n\nWhen writing the new layer in split commit-graph, we write a GDAT chunk\nonly if the topmost layer has a GDAT chunk. This guarantees that if a\nlayer has GDAT chunk, all lower layers must have a GDAT chunk as well.\n\nRewriting layers follows similar approach: if the topmost layer below\nthe set of layers being rewritten (in the split commit-graph chain)\nexists, and it does not contain GDAT chunk, then the result of rewrite\ndoes not have GDAT chunks either.\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                | 29 +++++++++++-\n commit-graph.h                |  1 +\n t/t5324-split-commit-graph.sh | 86 +++++++++++++++++++++++++++++++++++\n 3 files changed, 115 insertions(+), 1 deletion(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 71d0b243db..5d15a1399b 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -605,6 +605,21 @@ static struct commit_graph *load_commit_graph_chain(struct repository *r,\n \treturn graph_chain;\n }\n \n+static void validate_mixed_generation_chain(struct commit_graph *g)\n+{\n+\tint read_generation_data;\n+\n+\tif (!g)\n+\t\treturn;\n+\n+\tread_generation_data = !!g->chunk_generation_data;\n+\n+\twhile (g) {\n+\t\tg->read_generation_data = read_generation_data;\n+\t\tg = g->base_graph;\n+\t}\n+}\n+\n struct commit_graph *read_commit_graph_one(struct repository *r,\n \t\t\t\t\t   struct object_directory *odb)\n {\n@@ -613,6 +628,8 @@ struct commit_graph *read_commit_graph_one(struct repository *r,\n \tif (!g)\n \t\tg = load_commit_graph_chain(r, odb);\n \n+\tvalidate_mixed_generation_chain(g);\n+\n \treturn g;\n }\n \n@@ -782,7 +799,7 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tif (g->chunk_generation_data) {\n+\tif (g->chunk_generation_data && g->read_generation_data) {\n \t\toffset = (timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n \n \t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n@@ -2030,6 +2047,9 @@ static void split_graph_merge_strategy(struct write_commit_graph_context *ctx)\n \t\t}\n \t}\n \n+\tif (!ctx->write_generation_data && g->chunk_generation_data)\n+\t\tctx->write_generation_data = 1;\n+\n \tif (flags != COMMIT_GRAPH_SPLIT_REPLACE)\n \t\tctx->new_base_graph = g;\n \telse if (ctx->num_commit_graphs_after != 1)\n@@ -2274,6 +2294,7 @@ int write_commit_graph(struct object_directory *odb,\n \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n \n \t\twhile (g) {\n+\t\t\tg->read_generation_data = 1;\n \t\t\tg->topo_levels = &topo_levels;\n \t\t\tg = g->base_graph;\n \t\t}\n@@ -2300,6 +2321,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \t\tg = ctx->r->objects->commit_graph;\n \n+\t\tif (g && !g->chunk_generation_data)\n+\t\t\tctx->write_generation_data = 0;\n+\n \t\twhile (g) {\n \t\t\tctx->num_commit_graphs_before++;\n \t\t\tg = g->base_graph;\n@@ -2318,6 +2342,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \t\tif (ctx->opts)\n \t\t\treplace = ctx->opts->split_flags & COMMIT_GRAPH_SPLIT_REPLACE;\n+\n+\t\tif (replace)\n+\t\t\tctx->write_generation_data = 1;\n \t}\n \n \tctx->approx_nr_objects = approximate_object_count();\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 19a02001fd..ad52130883 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -64,6 +64,7 @@ struct commit_graph {\n \tstruct object_directory *odb;\n \n \tuint32_t num_commits_in_base;\n+\tunsigned int read_generation_data;\n \tstruct commit_graph *base_graph;\n \n \tconst uint32_t *chunk_oid_fanout;\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 651df89ab2..d0949a9eb8 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -440,4 +440,90 @@ test_expect_success '--split=replace with partial Bloom data' '\n \tverify_chain_files_exist $graphdir\n '\n \n+test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n+\tmkdir mixed &&\n+\tgraphdir=\".git/objects/info/commit-graphs\" &&\n+\ttest_oid_cache <<-EOM &&\n+\toid_version sha1:1\n+\toid_version sha256:2\n+\tEOM\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit init &&\n+\tgit config core.commitGraph true &&\n+\tgit config gc.writeCommitGraph false &&\n+\tfor i in $(test_seq 3)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit reset --hard commits/1 &&\n+\tfor i in $(test_seq 4 5)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit reset --hard commits/2 &&\n+\tfor i in $(test_seq 6 10)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit commit-graph write --reachable --split &&\n+\tgit reset --hard commits/2 &&\n+\tgit merge commits/4 &&\n+\tgit branch merge/1 &&\n+\tgit reset --hard commits/4 &&\n+\tgit merge commits/6 &&\n+\tgit branch merge/2 &&\n+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n+\ttest-tool read-graph >output &&\n+\tcat >expect <<-EOF &&\n+\theader: 43475048 1 $(test_oid oid_version) 4 1\n+\tnum_commits: 2\n+\tchunks: oid_fanout oid_lookup commit_metadata\n+\tEOF\n+\ttest_cmp expect output &&\n+\tgit commit-graph verify\n+'\n+\n+test_expect_success 'does not write generation data chunk if not present on existing tip' '\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit reset --hard commits/3 &&\n+\tgit merge merge/1 &&\n+\tgit merge commits/5 &&\n+\tgit merge merge/2 &&\n+\tgit branch merge/3 &&\n+\tgit commit-graph write --reachable --split=no-merge &&\n+\ttest-tool read-graph >output &&\n+\tcat >expect <<-EOF &&\n+\theader: 43475048 1 $(test_oid oid_version) 4 2\n+\tnum_commits: 3\n+\tchunks: oid_fanout oid_lookup commit_metadata\n+\tEOF\n+\ttest_cmp expect output &&\n+\tgit commit-graph verify\n+'\n+\n+test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\tgit commit-graph write --reachable --split=replace &&\n+\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\tverify_chain_files_exist $graphdir &&\n+\tgraph_read_expect 15 &&\n+\tgit commit-graph verify\n+'\n+\n+test_expect_success 'add one commit, write a tip graph' '\n+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n+\ttest_commit 11 &&\n+\tgit branch commits/11 &&\n+\tgit commit-graph write --reachable --split &&\n+\ttest_path_is_missing $infodir/commit-graph &&\n+\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\tls $graphdir/graph-*.graph >graph-files &&\n+\ttest_line_count = 2 graph-files &&\n+\tverify_chain_files_exist $graphdir\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"407062","messageId":"694ef1ec08d9dc96a74a2631b2710ad206397dbc.1602079786.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"[PATCH v4 06/10] commit-graph: implement corrected commit date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-10-07T14:09:41Z","receivedAt":"2020-10-07T14:10:09Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith most of preparations done, let's implement corrected commit date.\n\nThe corrected commit date for a commit is defined as:\n\n* A commit with no parents (a root commit) has corrected commit date\n  equal to its committer date.\n* A commit with at least one parent has corrected commit date equal to\n  the maximum of its commit date and one more than the largest corrected\n  commit date among its parents.\n\nAs a special case, a root commit with timestamp of zero (01.01.1970\n00:00:00Z) has corrected commit date of one, to be able to distinguish\nfrom GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit\ndate).\n\nTo minimize the space required to store corrected commit date, Git\nstores corrected commit date offsets into the commit-graph file. The\ncorrected commit date offset for a commit is defined as the difference\nbetween its corrected commit date and actual commit date.\n\nStoring corrected commit date requires sizeof(timestamp_t) bytes, which\nin most cases is 64 bits (uintmax_t). However, corrected commit date\noffsets can be safely stored using only 32-bits. This halves the size\nof GDAT chunk, which is a reduction of around 6% in the size of\ncommit-graph file.\n\nHowever, using offsets be problematic if one of commits is malformed but\nvalid and has committerdate of 0 Unix time, as the offset would be the\nsame as corrected commit date and thus require 64-bits to be stored\nproperly.\n\nWhile Git does not write out offsets at this stage, Git stores the\ncorrected commit dates in member generation of struct commit_graph_data.\nIt will begin writing commit date offsets with the introduction of\ngeneration data chunk.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 43 +++++++++++++++++++++++--------------------\n 1 file changed, 23 insertions(+), 20 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex cedd311024..03948adfce 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -154,11 +154,6 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \telse if (generation_a > generation_b)\n \t\treturn 1;\n \n-\t/* use date as a heuristic when generations are equal */\n-\tif (a->date < b->date)\n-\t\treturn -1;\n-\telse if (a->date > b->date)\n-\t\treturn 1;\n \treturn 0;\n }\n \n@@ -1357,10 +1352,14 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\ttimestamp_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n+\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n \t\tif (level != GENERATION_NUMBER_INFINITY &&\n-\t\t    level != GENERATION_NUMBER_ZERO)\n+\t\t    level != GENERATION_NUMBER_ZERO &&\n+\t\t    corrected_commit_date != GENERATION_NUMBER_INFINITY &&\n+\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO\n+\t\t    )\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1369,17 +1368,25 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n \t\t\tuint32_t max_level = 0;\n+\t\t\ttimestamp_t max_corrected_commit_date = 0;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n-\n+\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n \t\t\t\tif (level == GENERATION_NUMBER_INFINITY ||\n-\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n+\t\t\t\t    level == GENERATION_NUMBER_ZERO ||\n+\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_INFINITY ||\n+\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO\n+\t\t\t\t    ) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (level > max_level) {\n-\t\t\t\t\tmax_level = level;\n+\t\t\t\t} else {\n+\t\t\t\t\tif (level > max_level)\n+\t\t\t\t\t\tmax_level = level;\n+\n+\t\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n+\t\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n \t\t\t\t}\n \t\t\t}\n \n@@ -1389,6 +1396,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n \t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n+\n+\t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n+\t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n+\t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n \t\t\t}\n \t\t}\n \t}\n@@ -2485,17 +2496,9 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n \t\t\tcontinue;\n \n-\t\t/*\n-\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n-\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n-\t\t * extra logic in the following condition.\n-\t\t */\n-\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n-\t\t\tmax_generation--;\n-\n \t\tgeneration = commit_graph_generation(graph_commit);\n-\t\tif (generation != max_generation + 1)\n-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n+\t\tif (generation < max_generation + 1)\n+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n \t\t\t\t     oid_to_hex(&cur_oid),\n \t\t\t\t     generation,\n \t\t\t\t     max_generation + 1);\n-- \ngitgitgadget\n\n"},{"id":"408302","messageId":"85wnzf43kv.fsf@gmail.com","threadId":"53933","inReplyTo":"fae81b534b14c8227454ff94e385fb16faee0e99.1602079785.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 01/10] commit-graph: fix regression when computing Bloom filters","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-10-24T23:16:48Z","receivedAt":"2020-10-24T23:16:59Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> commit_gen_cmp is used when writing a commit-graph to sort commits in\n> generation order before computing Bloom filters. Since c49c82aa (commit:\n> move members graph_pos, generation to a slab, 2020-06-17) made it so\n> that 'commit_graph_generation()' returns 'GENERATION_NUMBER_INFINITY'\n> during writing, we cannot call it within this function. Instead, access\n> the generation number directly through the slab (i.e., by calling\n> 'commit_graph_data_at(c)->generation') in order to access it while\n> writing.\n\nThis description is all right, but I think it can be made more clear:\n\n  When running `git commit-graph write --reachable --changed-paths` to\n  compute Bloom filters for changed paths, commits are first sorted by\n  generation number using 'commit_gen_cmp()'.  Commits with similar\n  generation are more likely to have many trees in common, making the\n  diff faster, see 3d112755.\n\n  However, since c49c82aa (commit: move members graph_pos, generation to\n  a slab, 2020-06-17) made it so that 'commit_graph_generation()'\n  returns 'GENERATION_NUMBER_INFINITY' during writing, we cannot call it\n  within this function.  Instead, access the generation number directly\n  through the slab (i.e., by calling 'commit_graph_data_at(c)->generation')\n  in order to access it while writing.\n\nOr something like that.\n\nWe should also add an explanation why avoiding getter is safe here,\nperhaps adding the following line to the second paragraph:\n\n  It is safe to do because 'commit_gen_cmp()' from commit-graph.c is\n  static and used only when writing Bloom filters, and because writing\n  changed-paths filters is done after computing generation numbers (if\n  necessary).\n\nOr something like that.\n\n>\n> While measuring performance with `git commit-graph write --reachable\n> --changed-paths` on the linux repository led to around 1m40s for both\n> HEAD and master (and could be due to fault in my measurements), it is\n> still the \"right\" thing to do.\n\nI had to read the above paragraph several times to understand it,\npossibly because I have expected here to be a fix for a performance\nregression.  The commit message for 3d112755 (commit-graph: examine\ncommits by generation number) describes reduction of computation time\nfrom 3m00s to 1m37s.  So I would expect performance with HEAD (i.e.\nbefore those changes) to be around 3m, not the same before and after\nchanges being around 1m40s.\n\nCan anyone recheck this before-and-after benchmark, please?\n\nAnyway, it might be more clear to write it as the following:\n\n  On the Linux kernel repository, this patch didn't reduce the\n  computation time for 'git commit-graph write --reachable\n  --changed-paths', which is around 1m40s both before and after this\n  change.  This could be a fault in my measurements; it is still the\n  \"right\" thing to do.\n\nOr something like that.\n\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n\nAnyway, it is nice and clear change.\n\n>  commit-graph.c | 4 ++--\n>  1 file changed, 2 insertions(+), 2 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index cb042bdba8..94503e584b 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>  \tconst struct commit *a = *(const struct commit **)va;\n>  \tconst struct commit *b = *(const struct commit **)vb;\n>  \n> -\tuint32_t generation_a = commit_graph_generation(a);\n> -\tuint32_t generation_b = commit_graph_generation(b);\n> +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n>  \t/* lower generation commits first */\n>  \tif (generation_a < generation_b)\n>  \t\treturn -1;\n\nBest,\n-- \nJakub Narębski\n"},{"id":"408303","messageId":"85o8kr42g0.fsf@gmail.com","threadId":"53933","inReplyTo":"4470d916428a28bb8277dfc4c3da84e08110e88e.1602079786.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 02/10] revision: parse parent in indegree_walk_step()","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-10-24T23:41:19Z","receivedAt":"2020-10-24T23:44:32Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> In indegree_walk_step(), we add unvisited parents to the indegree queue.\n> However, parents are not guaranteed to be parsed. As the indegree queue\n> sorts by generation number, let's parse parents before inserting them to\n> ensure the correct priority order.\n\nAll right, we need to ensure the parent commit is parsed to know its\ngeneration number, to insert in into priority queue in a correct order.\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nLooks good.\n\n> ---\n>  revision.c | 3 +++\n>  1 file changed, 3 insertions(+)\n>\n> diff --git a/revision.c b/revision.c\n> index aa62212040..c97abcdde1 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -3381,6 +3381,9 @@ static void indegree_walk_step(struct rev_info *revs)\n>  \t\tstruct commit *parent = p->item;\n>  \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n>  \n> +\t\tif (repo_parse_commit_gently(revs->repo, parent, 1) < 0)\n> +\t\t\treturn;\n> +\n>  \t\tif (*pi)\n>  \t\t\t(*pi)++;\n>  \t\telse\n\nBest,\n-- \nJakub Narębski\n"},{"id":"408344","messageId":"85blgq4lxh.fsf@gmail.com","threadId":"53933","inReplyTo":"18bb3318a12c859c21c8e95285d551c48d31b54b.1602079786.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 03/10] commit-graph: consolidate fill_commit_graph_info","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-10-25T10:52:42Z","receivedAt":"2020-10-25T10:52:48Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hi Abhishek,\n\nIn short: everything is all right, except for the now duplicated test\nnames in t5000 after this commit.\n\n\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> Both fill_commit_graph_info() and fill_commit_in_graph() parse\n> information present in commit data chunk. Let's simplify the\n> implementation by calling fill_commit_graph_info() within\n> fill_commit_in_graph().\n>\n> fill_commit_graph_info() used to not load committer data from commit data\n> chunk. However, with the corrected committer date, we have to load\n> committer date to calculate generation number value.\n\nNice writeup, however the last sentence would in my opinion read better\nin the future tense: we don't use generation number v2 yet.  For\nexample:\n\n  However, with upcoming switch to using corrected committer date as\n  generation number v2, we will have to load committer date to compute\n  generation number value anyway.\n\nOr something like that - notice the minor addition and changes.\n\nThe following is slightly unrelated change, but we agreed that it would\nbe better to not separate them; the need for change to the t5000 test is\ncaused by the change described above.\n\n>\n> e51217e15 (t5000: test tar files that overflow ustar headers,\n> 30-06-2016) introduced a test 'generate tar with future mtime' that\n> creates a commit with committer date of (2 ^ 36 + 1) seconds since\n> EPOCH. The CDAT chunk provides 34-bits for storing committer date, thus\n> committer time overflows into generation number (within CDAT chunk) and\n> has undefined behavior.\n>\n> The test used to pass as fill_commit_graph_info() would not set struct\n> member `date` of struct commit and loads committer date from the object\n> database, generating a tar file with the expected mtime.\n\nI think it should be s/loads/load/, as in \"would load\", but I am not a\nnative English speaker.\n\n>\n> However, with corrected commit date, we will load the committer date\n> from CDAT chunk (truncated to lower 34-bits to populate the generation\n> number. Thus, Git sets date and generates tar file with the truncated\n> mtime.\n>\n> The ustar format (the header format used by most modern tar programs)\n> only has room for 11 (or 12, depending om some implementations) octal\n> digits for the size and mtime of each files.\n>\n> Thus, setting a timestamp of 2 ^ 33 + 1 would overflow the 11-octal\n> digit implementations while still fitting into commit data chunk.\n>\n> Since we want to test 12-octal digit implementations of ustar as well,\n> let's modify the existing test to no longer use commit-graph file.\n\nThe description above is for me does not make it entirely clear that we\nadd new test for handling possible 11-octal digit overflow nearly\nidentical to the existing one, and turn off use of commit-graph file for\ntest that checks handling 12-octal digit overflow.\n\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c      | 27 ++++++++++-----------------\n>  t/t5000-tar-tree.sh | 20 +++++++++++++++++++-\n>  2 files changed, 29 insertions(+), 18 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 94503e584b..e8362e144e 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -749,15 +749,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \tconst unsigned char *commit_data;\n>  \tstruct commit_graph_data *graph_data;\n>  \tuint32_t lex_index;\n> +\tuint64_t date_high, date_low;\n>  \n>  \twhile (pos < g->num_commits_in_base)\n>  \t\tg = g->base_graph;\n>  \n> +\tif (pos >= g->num_commits + g->num_commits_in_base)\n> +\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n> +\n>  \tlex_index = pos - g->num_commits_in_base;\n>  \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n>  \n>  \tgraph_data = commit_graph_data_at(item);\n>  \tgraph_data->graph_pos = pos;\n> +\n> +\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> +\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> +\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> +\n>  \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n>  }\n>  \n> @@ -772,38 +781,22 @@ static int fill_commit_in_graph(struct repository *r,\n>  {\n>  \tuint32_t edge_value;\n>  \tuint32_t *parent_data_ptr;\n> -\tuint64_t date_low, date_high;\n>  \tstruct commit_list **pptr;\n> -\tstruct commit_graph_data *graph_data;\n>  \tconst unsigned char *commit_data;\n>  \tuint32_t lex_index;\n>  \n>  \twhile (pos < g->num_commits_in_base)\n>  \t\tg = g->base_graph;\n>  \n> -\tif (pos >= g->num_commits + g->num_commits_in_base)\n> -\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n> +\tfill_commit_graph_info(item, g, pos);\n>  \n> -\t/*\n> -\t * Store the \"full\" position, but then use the\n> -\t * \"local\" position for the rest of the calculation.\n> -\t */\n> -\tgraph_data = commit_graph_data_at(item);\n> -\tgraph_data->graph_pos = pos;\n>  \tlex_index = pos - g->num_commits_in_base;\n> -\n>  \tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n>  \n>  \titem->object.parsed = 1;\n>  \n>  \tset_commit_tree(item, NULL);\n>  \n> -\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> -\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> -\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> -\n> -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> -\n>  \tpptr = &item->parents;\n>  \n>  \tedge_value = get_be32(commit_data + g->hash_len);\n\nAll right, looks good for me.\n\nHere second change begins.\n\n> diff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\n> index 3ebb0d3b65..8f41cdc509 100755\n> --- a/t/t5000-tar-tree.sh\n> +++ b/t/t5000-tar-tree.sh\n> @@ -431,11 +431,29 @@ test_expect_success TAR_HUGE,LONG_IS_64BIT 'system tar can read our huge size' '\n>  \ttest_cmp expect actual\n>  '\n>  \n> +test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n> +\trm -f .git/index &&\n> +\techo foo >file &&\n> +\tgit add file &&\n> +\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n> +\t\tgit commit -m \"tempori parendum\"\n> +'\n> +\n> +test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n> +\tgit archive HEAD >future.tar\n> +'\n> +\n> +test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n> +\techo 2514 >expect &&\n> +\ttar_info future.tar | cut -d\" \" -f2 >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n\nEverything is all right, except we now have duplicated test names.\n\nPerhaps in the three following tests we should use 'far-far-future\ncommit' and 'far future mtime' in place of current 'far-future commit'\nand 'future mtime' for tests checking handling 12-digital ditgits\noverflow, or add description how far the future is, for example\n'far-future commit (2^11 + 1)', etc.\n\n>  test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n>  \trm -f .git/index &&\n>  \techo content >file &&\n>  \tgit add file &&\n> -\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n> +\tGIT_TEST_COMMIT_GRAPH=0 GIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n>  \t\tgit commit -m \"tempori parendum\"\n>  '\n\nBest,\n-- \nJakub Narębski\n"},{"id":"408347","messageId":"85imay2z84.fsf@gmail.com","threadId":"53933","inReplyTo":"011b0aa497d1352bf54ac6a9e2e22ed92d409e64.1602079786.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 04/10] commit-graph: return 64-bit generation number","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-10-25T13:48:27Z","receivedAt":"2020-10-25T13:48:33Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hi Abhishek,\n\n\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> In a preparatory step, let's return timestamp_t values from\n> commit_graph_generation(), use timestamp_t for local variables and\n> define GENERATION_NUMBER_INFINITY as (2 ^ 63 - 1) instead.\n\nI think it would be easier to understand if it was explicitely said what\nthis preparatory step prepares for, e.g.:\n\n  In a preparatory step for introducing corrected commit dates as\n  generation number, let's return timestamp_t values from...\n\nOr even\n\n  generation number, let's change the return type of\n  commit_graph_generation() to timestamp_t, and use ...\n\nOtherwise it looks good.\n\n>\n> We rename GENERATION_NUMBER_MAX to GENERATION_NUMBER_V1_MAX to\n> represent the largest topological level we can store in the commit data\n> chunk.\n>\n> With corrected commit dates implemented, we will have two such *_MAX\n> variables to denote the largest offset and largest topological level\n> that can be stored.\n\nAll right, nice explanation.\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nNote that there are two changes that are not mentioned in the commit\nmessage, namely adding 'const'-ness to generation_a/b local variables in\ncommit_gen_cmp() from commit-graph.c, and switching from\nGENERATION_NUMBER_ZERO to GENERATION_NUMBER_INFINITY as the default\n(initial) value for 'max_generation' in repo_in_merge_bases_many().\n\nWhile the former is a simple \"while-at-it\" change that shouldn't affect\ncorrectness, the latter needs an explanation (or fixing if it is wrong).\n\n> ---\n>  commit-graph.c | 22 +++++++++++-----------\n>  commit-graph.h |  4 ++--\n>  commit-reach.c | 36 ++++++++++++++++++------------------\n>  commit-reach.h |  2 +-\n>  commit.c       |  4 ++--\n>  commit.h       |  4 ++--\n>  revision.c     | 10 +++++-----\n>  upload-pack.c  |  2 +-\n>  8 files changed, 42 insertions(+), 42 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index e8362e144e..bfc532de6f 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -99,7 +99,7 @@ uint32_t commit_graph_position(const struct commit *c)\n>  \treturn data ? data->graph_pos : COMMIT_NOT_FROM_GRAPH;\n>  }\n>  \n> -uint32_t commit_graph_generation(const struct commit *c)\n> +timestamp_t commit_graph_generation(const struct commit *c)\n\nAll right.\n\n>  {\n>  \tstruct commit_graph_data *data =\n>  \t\tcommit_graph_data_slab_peek(&commit_graph_data_slab, c);\n> @@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>  \tconst struct commit *a = *(const struct commit **)va;\n>  \tconst struct commit *b = *(const struct commit **)vb;\n>  \n> -\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> -\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n> +\tconst timestamp_t generation_a = commit_graph_data_at(a)->generation;\n> +\tconst timestamp_t generation_b = commit_graph_data_at(b)->generation;\n\nAll right... but this also adds 'const' qualifier.  I understand that\nyou don't want to create separate commit for this \"while at it\"\nchange...\n\n>  \t/* lower generation commits first */\n>  \tif (generation_a < generation_b)\n>  \t\treturn -1;\n> @@ -1350,7 +1350,7 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n>  \t\t\t\t\tctx->commits.nr);\n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n> -\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n> +\t\ttimestamp_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n\nAll right.\n\n>  \n>  \t\tdisplay_progress(ctx->progress, i + 1);\n>  \t\tif (generation != GENERATION_NUMBER_INFINITY &&\n> @@ -1383,8 +1383,8 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\t\tdata->generation = max_generation + 1;\n>  \t\t\t\tpop_commit(&list);\n>  \n> -\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n> -\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n> +\t\t\t\tif (data->generation > GENERATION_NUMBER_V1_MAX)\n> +\t\t\t\t\tdata->generation = GENERATION_NUMBER_V1_MAX;\n\nAll right, this is the other mentioned change.\n\n>  \t\t\t}\n>  \t\t}\n>  \t}\n> @@ -2404,8 +2404,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n>  \tfor (i = 0; i < g->num_commits; i++) {\n>  \t\tstruct commit *graph_commit, *odb_commit;\n>  \t\tstruct commit_list *graph_parents, *odb_parents;\n> -\t\tuint32_t max_generation = 0;\n> -\t\tuint32_t generation;\n> +\t\ttimestamp_t max_generation = 0;\n> +\t\ttimestamp_t generation;\n\nAll right.\n\n>  \n>  \t\tdisplay_progress(progress, i + 1);\n>  \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n> @@ -2469,11 +2469,11 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n>  \t\t\tcontinue;\n>  \n>  \t\t/*\n> -\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n> -\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n> +\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n> +\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n>  \t\t * extra logic in the following condition.\n>  \t\t */\n> -\t\tif (max_generation == GENERATION_NUMBER_MAX)\n> +\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n>  \t\t\tmax_generation--;\n\nAll right.  Nice fixing a comment too.\n\n>  \n>  \t\tgeneration = commit_graph_generation(graph_commit);\n> diff --git a/commit-graph.h b/commit-graph.h\n> index f8e92500c6..8be247fa35 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -144,12 +144,12 @@ void disable_commit_graph(struct repository *r);\n>  \n>  struct commit_graph_data {\n>  \tuint32_t graph_pos;\n> -\tuint32_t generation;\n> +\ttimestamp_t generation;\n>  };\n\nAll right.\n\n>  \n>  /*\n>   * Commits should be parsed before accessing generation, graph positions.\n>   */\n> -uint32_t commit_graph_generation(const struct commit *);\n> +timestamp_t commit_graph_generation(const struct commit *);\n>  uint32_t commit_graph_position(const struct commit *);\n>  #endif\n\nAll right.\n\n> diff --git a/commit-reach.c b/commit-reach.c\n> index 50175b159e..20b48b872b 100644\n> --- a/commit-reach.c\n> +++ b/commit-reach.c\n> @@ -32,12 +32,12 @@ static int queue_has_nonstale(struct prio_queue *queue)\n>  static struct commit_list *paint_down_to_common(struct repository *r,\n>  \t\t\t\t\t\tstruct commit *one, int n,\n>  \t\t\t\t\t\tstruct commit **twos,\n> -\t\t\t\t\t\tint min_generation)\n> +\t\t\t\t\t\ttimestamp_t min_generation)\n>  {\n>  \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n>  \tstruct commit_list *result = NULL;\n>  \tint i;\n> -\tuint32_t last_gen = GENERATION_NUMBER_INFINITY;\n> +\ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n\nAll right.\n\n>  \n>  \tif (!min_generation)\n>  \t\tqueue.compare = compare_commits_by_commit_date;\n> @@ -58,10 +58,10 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n>  \t\tstruct commit *commit = prio_queue_get(&queue);\n>  \t\tstruct commit_list *parents;\n>  \t\tint flags;\n> -\t\tuint32_t generation = commit_graph_generation(commit);\n> +\t\ttimestamp_t generation = commit_graph_generation(commit);\n\nAll right.\n\n>  \n>  \t\tif (min_generation && generation > last_gen)\n> -\t\t\tBUG(\"bad generation skip %8x > %8x at %s\",\n> +\t\t\tBUG(\"bad generation skip %\"PRItime\" > %\"PRItime\" at %s\",\n\nAll right; nice of you noticing this issue.\n\n>  \t\t\t    generation, last_gen,\n>  \t\t\t    oid_to_hex(&commit->object.oid));\n>  \t\tlast_gen = generation;\n> @@ -177,12 +177,12 @@ static int remove_redundant(struct repository *r, struct commit **array, int cnt\n>  \t\trepo_parse_commit(r, array[i]);\n>  \tfor (i = 0; i < cnt; i++) {\n>  \t\tstruct commit_list *common;\n> -\t\tuint32_t min_generation = commit_graph_generation(array[i]);\n> +\t\ttimestamp_t min_generation = commit_graph_generation(array[i]);\n>  \n>  \t\tif (redundant[i])\n>  \t\t\tcontinue;\n>  \t\tfor (j = filled = 0; j < cnt; j++) {\n> -\t\t\tuint32_t curr_generation;\n> +\t\t\ttimestamp_t curr_generation;\n>  \t\t\tif (i == j || redundant[j])\n>  \t\t\t\tcontinue;\n>  \t\t\tfilled_index[filled] = j;\n\nAll right.\n\n> @@ -321,7 +321,7 @@ int repo_in_merge_bases_many(struct repository *r, struct commit *commit,\n>  {\n>  \tstruct commit_list *bases;\n>  \tint ret = 0, i;\n> -\tuint32_t generation, max_generation = GENERATION_NUMBER_ZERO;\n> +\ttimestamp_t generation, max_generation = GENERATION_NUMBER_INFINITY;\n\nThe change of type from uint32_t to timestamp_t is expected, but the\nchange from GENERATION_NUMBER_ZERO to GENERATION_NUMBER_INFINITY is not.\n\nThis might be caused by the fact that repo_in_merge_bases_many()\nswitched from using min_generation and GENERATION_NUMBER_INFINITY to\nusing max_generation and GENERATION_NUMBER_ZERO. Or the reverse: I see\none version on https://github.com/git/git, and other version in 'master'\npulled from https://github.com/git-for-windows/git\n\nCertainly max_generation should be paired with GENERATION_NUMBER_ZERO,\nand min_generation with GENERATION_NUMBER_INFINITY.\n\n>  \n>  \tif (repo_parse_commit(r, commit))\n>  \t\treturn ret;\n> @@ -470,7 +470,7 @@ static int in_commit_list(const struct commit_list *want, struct commit *c)\n>  static enum contains_result contains_test(struct commit *candidate,\n>  \t\t\t\t\t  const struct commit_list *want,\n>  \t\t\t\t\t  struct contains_cache *cache,\n> -\t\t\t\t\t  uint32_t cutoff)\n> +\t\t\t\t\t  timestamp_t cutoff)\n\nAll right.\n\nSidenote: this parameter should probably be named gen_cutoff, for\nconsistency and better readability (but that was the existing state),\nbut this would also mean more changes.\n\n\n>  {\n>  \tenum contains_result *cached = contains_cache_at(cache, candidate);\n>  \n> @@ -506,11 +506,11 @@ static enum contains_result contains_tag_algo(struct commit *candidate,\n>  {\n>  \tstruct contains_stack contains_stack = { 0, 0, NULL };\n>  \tenum contains_result result;\n> -\tuint32_t cutoff = GENERATION_NUMBER_INFINITY;\n> +\ttimestamp_t cutoff = GENERATION_NUMBER_INFINITY;\n\nSidenote: this variable should probably be named gen_cutoff, for\nconsistency and better readability (but that was the existing state).\nHowever changing it would pollute this commit with unrelated changes;\nit is not that big of an isseu that it *requires* fixing.\n\n>  \tconst struct commit_list *p;\n>  \n>  \tfor (p = want; p; p = p->next) {\n> -\t\tuint32_t generation;\n> +\t\ttimestamp_t generation;\n>  \t\tstruct commit *c = p->item;\n>  \t\tload_commit_graph_info(the_repository, c);\n>  \t\tgeneration = commit_graph_generation(c);\n\nAll right.\n\n> @@ -566,8 +566,8 @@ static int compare_commits_by_gen(const void *_a, const void *_b)\n>  \tconst struct commit *a = *(const struct commit * const *)_a;\n>  \tconst struct commit *b = *(const struct commit * const *)_b;\n>  \n> -\tuint32_t generation_a = commit_graph_generation(a);\n> -\tuint32_t generation_b = commit_graph_generation(b);\n> +\ttimestamp_t generation_a = commit_graph_generation(a);\n> +\ttimestamp_t generation_b = commit_graph_generation(b);\n\nAll right.\n\n>  \n>  \tif (generation_a < generation_b)\n>  \t\treturn -1;\n> @@ -580,7 +580,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n>  \t\t\t\t unsigned int with_flag,\n>  \t\t\t\t unsigned int assign_flag,\n>  \t\t\t\t time_t min_commit_date,\n> -\t\t\t\t uint32_t min_generation)\n> +\t\t\t\t timestamp_t min_generation)\n>  {\n>  \tstruct commit **list = NULL;\n>  \tint i;\n\nAll right.\n\n> @@ -681,13 +681,13 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n>  \ttime_t min_commit_date = cutoff_by_min_date ? from->item->date : 0;\n>  \tstruct commit_list *from_iter = from, *to_iter = to;\n>  \tint result;\n> -\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n> +\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n>  \n>  \twhile (from_iter) {\n>  \t\tadd_object_array(&from_iter->item->object, NULL, &from_objs);\n>  \n>  \t\tif (!parse_commit(from_iter->item)) {\n> -\t\t\tuint32_t generation;\n> +\t\t\ttimestamp_t generation;\n>  \t\t\tif (from_iter->item->date < min_commit_date)\n>  \t\t\t\tmin_commit_date = from_iter->item->date;\n>\n\nAll right.\n\n> @@ -701,7 +701,7 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n>  \n>  \twhile (to_iter) {\n>  \t\tif (!parse_commit(to_iter->item)) {\n> -\t\t\tuint32_t generation;\n> +\t\t\ttimestamp_t generation;\n>  \t\t\tif (to_iter->item->date < min_commit_date)\n>  \t\t\t\tmin_commit_date = to_iter->item->date;\n>\n\nAll right.\n\n> @@ -741,13 +741,13 @@ struct commit_list *get_reachable_subset(struct commit **from, int nr_from,\n>  \tstruct commit_list *found_commits = NULL;\n>  \tstruct commit **to_last = to + nr_to;\n>  \tstruct commit **from_last = from + nr_from;\n> -\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n> +\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n>  \tint num_to_find = 0;\n>  \n>  \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n>  \n>  \tfor (item = to; item < to_last; item++) {\n> -\t\tuint32_t generation;\n> +\t\ttimestamp_t generation;\n>  \t\tstruct commit *c = *item;\n>  \n>  \t\tparse_commit(c);\n\nAll right.\n\n> diff --git a/commit-reach.h b/commit-reach.h\n> index b49ad71a31..148b56fea5 100644\n> --- a/commit-reach.h\n> +++ b/commit-reach.h\n> @@ -87,7 +87,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n>  \t\t\t\t unsigned int with_flag,\n>  \t\t\t\t unsigned int assign_flag,\n>  \t\t\t\t time_t min_commit_date,\n> -\t\t\t\t uint32_t min_generation);\n> +\t\t\t\t timestamp_t min_generation);\n>  int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n>  \t\t       int commit_date_cutoff);\n>\n\nAll right.\n\n> diff --git a/commit.c b/commit.c\n> index f53429c0ac..3b488381d5 100644\n> --- a/commit.c\n> +++ b/commit.c\n> @@ -731,8 +731,8 @@ int compare_commits_by_author_date(const void *a_, const void *b_,\n>  int compare_commits_by_gen_then_commit_date(const void *a_, const void *b_, void *unused)\n>  {\n>  \tconst struct commit *a = a_, *b = b_;\n> -\tconst uint32_t generation_a = commit_graph_generation(a),\n> -\t\t       generation_b = commit_graph_generation(b);\n> +\tconst timestamp_t generation_a = commit_graph_generation(a),\n> +\t\t\t  generation_b = commit_graph_generation(b);\n>\n\nAll right (assuming that the indent after change looks all right; but\neven if it doesn't t would be a very minor issue).\n\n>  \t/* newer commits first */\n>  \tif (generation_a < generation_b)\n> diff --git a/commit.h b/commit.h\n> index 5467786c7b..33c66b2177 100644\n> --- a/commit.h\n> +++ b/commit.h\n> @@ -11,8 +11,8 @@\n>  #include \"commit-slab.h\"\n>  \n>  #define COMMIT_NOT_FROM_GRAPH 0xFFFFFFFF\n> -#define GENERATION_NUMBER_INFINITY 0xFFFFFFFF\n> -#define GENERATION_NUMBER_MAX 0x3FFFFFFF\n> +#define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n> +#define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n>  #define GENERATION_NUMBER_ZERO 0\n>\n\nAll right, we redefine GENERATION_NUMBER_INFINITY and rename\nGENERATION_NUMBER_MAX.\n\n>  struct commit_list {\n> diff --git a/revision.c b/revision.c\n> index c97abcdde1..2861f1c45c 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -3308,7 +3308,7 @@ define_commit_slab(indegree_slab, int);\n>  define_commit_slab(author_date_slab, timestamp_t);\n>  \n>  struct topo_walk_info {\n> -\tuint32_t min_generation;\n> +\ttimestamp_t min_generation;\n>  \tstruct prio_queue explore_queue;\n>  \tstruct prio_queue indegree_queue;\n>  \tstruct prio_queue topo_queue;\n\nAll right.\n\n> @@ -3354,7 +3354,7 @@ static void explore_walk_step(struct rev_info *revs)\n>  }\n>  \n>  static void explore_to_depth(struct rev_info *revs,\n> -\t\t\t     uint32_t gen_cutoff)\n> +\t\t\t     timestamp_t gen_cutoff)\n>  {\n>  \tstruct topo_walk_info *info = revs->topo_walk_info;\n>  \tstruct commit *c;\n\nAll right.\n\n> @@ -3397,7 +3397,7 @@ static void indegree_walk_step(struct rev_info *revs)\n>  }\n>  \n>  static void compute_indegrees_to_depth(struct rev_info *revs,\n> -\t\t\t\t       uint32_t gen_cutoff)\n> +\t\t\t\t       timestamp_t gen_cutoff)\n>  {\n>  \tstruct topo_walk_info *info = revs->topo_walk_info;\n>  \tstruct commit *c;\n\nAll right.\n\n> @@ -3455,7 +3455,7 @@ static void init_topo_walk(struct rev_info *revs)\n>  \tinfo->min_generation = GENERATION_NUMBER_INFINITY;\n>  \tfor (list = revs->commits; list; list = list->next) {\n>  \t\tstruct commit *c = list->item;\n> -\t\tuint32_t generation;\n> +\t\ttimestamp_t generation;\n>  \n>  \t\tif (repo_parse_commit_gently(revs->repo, c, 1))\n>  \t\t\tcontinue;\n\nAll right.\n\n> @@ -3516,7 +3516,7 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n>  \tfor (p = commit->parents; p; p = p->next) {\n>  \t\tstruct commit *parent = p->item;\n>  \t\tint *pi;\n> -\t\tuint32_t generation;\n> +\t\ttimestamp_t generation;\n>  \n>  \t\tif (parent->object.flags & UNINTERESTING)\n>  \t\t\tcontinue;\n\nAll right.\n\n> diff --git a/upload-pack.c b/upload-pack.c\n> index 3b858eb457..fdb82885b6 100644\n> --- a/upload-pack.c\n> +++ b/upload-pack.c\n> @@ -497,7 +497,7 @@ static int got_oid(struct upload_pack_data *data,\n>  \n>  static int ok_to_give_up(struct upload_pack_data *data)\n>  {\n> -\tuint32_t min_generation = GENERATION_NUMBER_ZERO;\n> +\ttimestamp_t min_generation = GENERATION_NUMBER_ZERO;\n>  \n>  \tif (!data->have_obj.nr)\n>  \t\treturn 0;\n\nAll right.\n\nThe only thing to check is if you have changed the type in all the\nplaces that need it. My cursory examination shows that those are all\nplaces than need fixing.\n\nNote that the 'generation' variable in git-name-rev, git-fsck and in\ngit-show-branch (snd sha1-name.c) means something different.\n\nAlso, 'first_generation' variable in generation_numbers_enabled() (part\nof commit-graph.c) examines and will examine generation number v1 i.e.\ntopological levels, and do not need type change... though it may require\nname change in some time in the future; the generation number\ncomputation path also does not require change type, though variables\nwould be renamed in the future commit.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"408356","messageId":"X5Xm5nlFK39o0rkJ@nand.local","threadId":"53933","inReplyTo":"85wnzf43kv.fsf@gmail.com","subject":"Re: [PATCH v4 01/10] commit-graph: fix regression when computing Bloom filters","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-10-25T20:58:14Z","receivedAt":"2020-10-25T21:01:53Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Sun, Oct 25, 2020 at 01:16:48AM +0200, Jakub Narębski wrote:\n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > While measuring performance with `git commit-graph write --reachable\n> > --changed-paths` on the linux repository led to around 1m40s for both\n> > HEAD and master (and could be due to fault in my measurements), it is\n> > still the \"right\" thing to do.\n>\n> I had to read the above paragraph several times to understand it,\n> possibly because I have expected here to be a fix for a performance\n> regression.  The commit message for 3d112755 (commit-graph: examine\n> commits by generation number) describes reduction of computation time\n> from 3m00s to 1m37s.  So I would expect performance with HEAD (i.e.\n> before those changes) to be around 3m, not the same before and after\n> changes being around 1m40s.\n>\n> Can anyone recheck this before-and-after benchmark, please?\n\nMy hunch is that our heuristic to fall back to the commits 'date'\nvalue is saving us here. commit_gen_cmp() first compares the generation\nnumbers, breaking ties by 'date' as a heuristic. But since all\ngeneration number queries return GENERATION_NUMBER_INFINITY during\nwriting, we're relying on our heuristic entirely.\n\nI haven't looked much further than that, other than to see that I could\nget about a ~4sec speed-up with this patch as compared to v2.29.1 in the\ncomputing Bloom filters region on the kernel.\n\n> Anyway, it might be more clear to write it as the following:\n>\n>   On the Linux kernel repository, this patch didn't reduce the\n>   computation time for 'git commit-graph write --reachable\n>   --changed-paths', which is around 1m40s both before and after this\n>   change.  This could be a fault in my measurements; it is still the\n>   \"right\" thing to do.\n>\n> Or something like that.\n\nAssuming that we are in fact being saved by the \"date\" heuristic, I'd\nprobably write the following commit message instead:\n\n  Before computing Bloom filters, the commit-graph machinery uses\n  commit_gen_cmp to sort commits by generation order for improved diff\n  performance. 3d11275505 (commit-graph: examine commits by generation\n  number, 2020-03-30) claims that this sort can reduce the time spent to\n  compute Bloom filters by nearly half.\n\n  But since c49c82aa4c (commit: move members graph_pos, generation to a\n  slab, 2020-06-17), this optimization is broken, since asking for\n  'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\n  while writing.\n\n  Not all hope is lost, though: 'commit_graph_generation()' falls\n  back to comparing commits by their date when they have equal generation\n  number, and so since c49c82aa4c is purely a date comparison function.\n  This heuristic is good enough that we don't seem to loose appreciable\n  performance while computing Bloom filters. [Benchmark that we loose\n  about ~4sec before/after c49c82aa4c9...]\n\n  So, avoid the uesless 'commit_graph_generation()' while writing by\n  instead accessing the slab directly. This returns the newly-computed\n  generation numbers, and allows us to avoid the heuristic by directly\n  comparing generation numbers.\n\nThanks,\nTaylor\n"},{"id":"408358","messageId":"85tuui0x3m.fsf@gmail.com","threadId":"53933","inReplyTo":"e067f653ad5d474eee5f40c13bb02fde26ebdb9b.1602079786.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 05/10] commit-graph: add a slab to store topological levels","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-10-25T22:17:17Z","receivedAt":"2020-10-25T22:17:23Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> In a later commit we will introduce corrected commit date as the\n> generation number v2. This value will be stored in the new seperate\n> Generation Data chunk. However, to ensure backwards compatibility with\n> \"Old\" Git we need to continue to write generation number v1, which is\n> topological level, to the commit data chunk. This means that we need to\n> compute both versions of generation numbers when writing the\n> commit-graph file. Therefore, let's introduce a commit-slab to store\n> topological levels; corrected commit date will be stored in the member\n> `generation` of struct commit_graph_data.\n>\n> When Git creates a split commit-graph, it takes advantage of the\n> generation values that have been computed already and present in\n> existing commit-graph files.\n>\n> So, let's add a pointer to struct commit_graph as well as struct\n> write_commit_graph_context to the topological level commit-slab\n> and populate it with topological levels while writing a commit-graph\n> file.\n\nI think you meant here \"add a pointer in `struct commit_graph` as well\nas in `struct write_commit_graph_context`...\".\n\nPerhaps we should add the information that it is done that way to be\nable to allocate topo_level_slab only when needed, in the\nwrite_commit_graph(), and adding new member to those struct is required\nto pass it through the call chain (modifying `struct commit_graph` is\nneeded for fill_commit_graph_info()).  But that might be too much detail\nto put in the commit message.\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 47 ++++++++++++++++++++++++++++++++---------------\n>  commit-graph.h |  1 +\n>  2 files changed, 33 insertions(+), 15 deletions(-)\n>\n\nLet me reorder those files for easier review.\n\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 8be247fa35..2e9aa7824e 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -73,6 +73,7 @@ struct commit_graph {\n>  \tconst unsigned char *chunk_bloom_indexes;\n>  \tconst unsigned char *chunk_bloom_data;\n>  \n> +\tstruct topo_level_slab *topo_levels;\n>  \tstruct bloom_filter_settings *bloom_filter_settings;\n>  };\n\nAll right, here we add new member to `struct commit_graph` type.\n\n> diff --git a/commit-graph.c b/commit-graph.c\n> index bfc532de6f..cedd311024 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -962,6 +967,7 @@ struct write_commit_graph_context {\n>  \t\t changed_paths:1,\n>  \t\t order_by_pack:1;\n>  \n> +\tstruct topo_level_slab *topo_levels;\n>  \tconst struct commit_graph_opts *opts;\n>  \tsize_t total_bloom_filter_data_size;\n>  \tconst struct bloom_filter_settings *bloom_settings;\n\nAll right, here we add new member to `struct write_commit_graph_context`\ntype, which is local to commit-graph.c.\n\n> @@ -64,6 +64,8 @@ void git_test_write_commit_graph_or_die(void)\n>  /* Remember to update object flag allocation in object.h */\n>  #define REACHABLE       (1u<<15)\n>  \n> +define_commit_slab(topo_level_slab, uint32_t);\n> +\n\nAll right, here we define new slab for storing topological levels; this\njust defines new type. Note that we do not define any setters and\ngetters to handle non-zero initialization, like we have for\ncommit_graph_data_slab.\n\n>  /* Keep track of the order in which commits are added to our list. */\n>  define_commit_slab(commit_pos, int);\n>  static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n> @@ -768,6 +770,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n>  \n>  \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> +\n> +\tif (g->topo_levels)\n> +\t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n\nI guess using get_be32() is repeated in this newly added part of code\nbecause previous part would be changed to read in generation number v2,\nif available, and we won't be then able to use\n\n\t\t*topo_level_slab_at(g->topo_levels, item) = graph_data->generation;\n\nAll right, that's smart.\n\n\nI guess that in fill_commit_graph_info() we don't know if we are reading\ncommit-graph, when topo levels slab is not present, or whether we are\nextending and writing the commit-graph file, when we need to fill it\nwith current commit-graph data.\n\nThe fact that fill_commit_graph_info() takes 'struct commit_graph' also\nexplains why we need to add pointer to a topo_levels slab to both\nstructs.\n\n>  }\n>  \n>  static inline void set_commit_tree(struct commit *c, struct tree *t)\n[...]\n> @@ -2142,6 +2146,7 @@ int write_commit_graph(struct object_directory *odb,\n>  \tint res = 0;\n>  \tint replace = 0;\n>  \tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n> +\tstruct topo_level_slab topo_levels;\n>  \n>  \tif (!commit_graph_compatible(the_repository))\n>  \t\treturn 0;\n> @@ -2163,6 +2168,18 @@ int write_commit_graph(struct object_directory *odb,\n>  \t\t\t\t\t\t\t bloom_settings.max_changed_paths);\n>  \tctx->bloom_settings = &bloom_settings;\n>  \n> +\tinit_topo_level_slab(&topo_levels);\n> +\tctx->topo_levels = &topo_levels;\n> +\n> +\tif (ctx->r->objects->commit_graph) {\n> +\t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n> +\n> +\t\twhile (g) {\n> +\t\t\tg->topo_levels = &topo_levels;\n> +\t\t\tg = g->base_graph;\n> +\t\t}\n> +\t}\n> +\n>  \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n>  \t\tctx->changed_paths = 1;\n>  \tif (!(flags & COMMIT_GRAPH_NO_WRITE_BLOOM_FILTERS)) {\n\nAll right, we need topo_level_slab only for writing the commit-graph, so\nwe allocate it with init_*_slab() in write_commit_graph(), and set\npointers to it in `struct write_commit_graph_context *ctx` and in\n`struct commit_graph` for each layer in the commit graph.  This is\nneeded to pass it down the call-chain.\n\nLooks good to me.\n\n> @@ -1108,7 +1114,7 @@ static int write_graph_chunk_data(struct hashfile *f,\n>  \t\telse\n>  \t\t\tpackedDate[0] = 0;\n>  \n> -\t\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n> +\t\tpackedDate[0] |= htonl(*topo_level_slab_at(ctx->topo_levels, *list) << 2);\n>\n\nAll right, write_graph_chunk_data() is called from write_commit_graph(),\nso we know that cxt->topo_levels is not NULL.\n\n>  \t\tpackedDate[1] = htonl((*list)->date);\n>  \t\thashwrite(f, packedDate, 8);\n> @@ -1350,11 +1356,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n>  \t\t\t\t\tctx->commits.nr);\n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n> -\t\ttimestamp_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n> +\t\ttimestamp_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n>\n\nAll right, we know that compute_generation_numbers() is called by the\nwrite_commit_graph(), so we know that cxt->topo_levels is not NULL.\n\nAlso, we rename 'generation' to 'level' in preparation for the time when\nwe would be computing *both* topological level (for backward\ncompatibility) and corrected committer date (to be used as generation\nnumber v2).  All right.\n\n>  \t\tdisplay_progress(ctx->progress, i + 1);\n> -\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n> -\t\t    generation != GENERATION_NUMBER_ZERO)\n> +\t\tif (level != GENERATION_NUMBER_INFINITY &&\n> +\t\t    level != GENERATION_NUMBER_ZERO)\n>  \t\t\tcontinue;\n\nSame here, the results of renaming of 'generation' local variable to\n'level'.\n\n>  \n>  \t\tcommit_list_insert(ctx->commits.list[i], &list);\n> @@ -1362,29 +1368,27 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\tstruct commit *current = list->item;\n>  \t\t\tstruct commit_list *parent;\n>  \t\t\tint all_parents_computed = 1;\n> -\t\t\tuint32_t max_generation = 0;\n> +\t\t\tuint32_t max_level = 0;\n\nSimilarly, we rename 'max_generation' to 'max_level'.\n\n>  \n>  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n> -\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n> +\t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n>  \n> -\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n> -\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n> +\t\t\t\tif (level == GENERATION_NUMBER_INFINITY ||\n> +\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n>  \t\t\t\t\tall_parents_computed = 0;\n>  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n>  \t\t\t\t\tbreak;\n> -\t\t\t\t} else if (generation > max_generation) {\n> -\t\t\t\t\tmax_generation = generation;\n> +\t\t\t\t} else if (level > max_level) {\n> +\t\t\t\t\tmax_level = level;\n>  \t\t\t\t}\n>  \t\t\t}\n\nContinuation of those renames.\n\n>  \n>  \t\t\tif (all_parents_computed) {\n> -\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n> -\n> -\t\t\t\tdata->generation = max_generation + 1;\n>  \t\t\t\tpop_commit(&list);\n>  \n> -\t\t\t\tif (data->generation > GENERATION_NUMBER_V1_MAX)\n> -\t\t\t\t\tdata->generation = GENERATION_NUMBER_V1_MAX;\n> +\t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n> +\t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n> +\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n\nThis is a bit safer way to handle possible overflow: instead of\n\n  final = max_found + 1;            /* set to maximum plus 1 */\n  if (final > MAX_POSSIBLE_VALUE)   /* handle overflow */\n      final = MAX_POSSIBLE_VALUE;\n\nwhere we can have problems if MAX_POSSIBLE_VALUE overflows, we use the\nfollowing pattern:\n\n  if (max_found > MAX_POSSIBLE_VALUE - 1)  /* handle overflow */\n      max_found > MAX_POSSIBLE_VALUE - 1;\n  final = max_found + 1;                   /* set to maximum plus 1 */\n\nIt is just a bit obscured by renaming variable and switch to using\ncommit slab.\n\nIt is not that important for topological level, where\nGENERATION_NUMBER_V1_MAX is smaller than maximum possible value, but it\nwould be important for generation number v2.\n\n>  \t\t\t}\n>  \t\t}\n>  \t}\n\nBest,\n--\nJakub Narębski\n"},{"id":"408484","messageId":"20201027063306.GA15674@Abhishek-Arch","threadId":"53933","inReplyTo":"85blgq4lxh.fsf@gmail.com","subject":"Re: [PATCH v4 03/10] commit-graph: consolidate fill_commit_graph_info","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-10-27T06:33:24Z","receivedAt":"2020-10-27T06:36:18Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"Hello Dr. Narębski,\n\nOn Sun, Oct 25, 2020 at 11:52:42AM +0100, Jakub Narębski wrote:\n> Hi Abhishek,\n> \n> In short: everything is all right, except for the now duplicated test\n> names in t5000 after this commit.\n> \n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > Both fill_commit_graph_info() and fill_commit_in_graph() parse\n> > information present in commit data chunk. Let's simplify the\n> > implementation by calling fill_commit_graph_info() within\n> > fill_commit_in_graph().\n> >\n> > fill_commit_graph_info() used to not load committer data from commit data\n> > chunk. However, with the corrected committer date, we have to load\n> > committer date to calculate generation number value.\n> \n> Nice writeup, however the last sentence would in my opinion read better\n> in the future tense: we don't use generation number v2 yet.  For\n> example:\n> \n>   However, with upcoming switch to using corrected committer date as\n>   generation number v2, we will have to load committer date to compute\n>   generation number value anyway.\n> \n> Or something like that - notice the minor addition and changes.\n> \n\nThanks for the change, it looks better!\n\n> The following is slightly unrelated change, but we agreed that it would\n> be better to not separate them; the need for change to the t5000 test is\n> caused by the change described above.\n\n> \n> >\n> > e51217e15 (t5000: test tar files that overflow ustar headers,\n> > 30-06-2016) introduced a test 'generate tar with future mtime' that\n> > creates a commit with committer date of (2 ^ 36 + 1) seconds since\n> > EPOCH. The CDAT chunk provides 34-bits for storing committer date, thus\n> > committer time overflows into generation number (within CDAT chunk) and\n> > has undefined behavior.\n> >\n> > The test used to pass as fill_commit_graph_info() would not set struct\n> > member `date` of struct commit and loads committer date from the object\n> > database, generating a tar file with the expected mtime.\n> \n> I think it should be s/loads/load/, as in \"would load\", but I am not a\n> native English speaker.\n> \n\nThat's correct - since I have used \"would not set\" in the first half of\nsentence, the later half should follow suit too.\n\n> >\n> > However, with corrected commit date, we will load the committer date\n> > from CDAT chunk (truncated to lower 34-bits to populate the generation\n> > number. Thus, Git sets date and generates tar file with the truncated\n> > mtime.\n> >\n> > The ustar format (the header format used by most modern tar programs)\n> > only has room for 11 (or 12, depending om some implementations) octal\n> > digits for the size and mtime of each files.\n> >\n> > Thus, setting a timestamp of 2 ^ 33 + 1 would overflow the 11-octal\n> > digit implementations while still fitting into commit data chunk.\n> >\n> > Since we want to test 12-octal digit implementations of ustar as well,\n> > let's modify the existing test to no longer use commit-graph file.\n> \n> The description above is for me does not make it entirely clear that we\n> add new test for handling possible 11-octal digit overflow nearly\n> identical to the existing one, and turn off use of commit-graph file for\n> test that checks handling 12-octal digit overflow.\n> \n\nRevised the last paragraphs to:\n\n  The ustar format (the header format used by most modern tar programs)\n  only has room for 11 (or 12, depending on some implementations) octal\n  digits for the size and mtime of each file.\n\n  To test the 11-octal digit implementation, we create a future commit\n  with committer date of 2^34 - 1, which overflows 11-octal digits\n  without overflowing 34-bits of the Commit Data chunk.\n\n  To test the 12-octal digit implementation, the smallest committer date\n  possible is 2^36, which overflows the Commit Data chunk and thus\n  commit-graph must be disabled for the test.\n\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c      | 27 ++++++++++-----------------\n> >  t/t5000-tar-tree.sh | 20 +++++++++++++++++++-\n> >  2 files changed, 29 insertions(+), 18 deletions(-)\n> >\n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index 94503e584b..e8362e144e 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -749,15 +749,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n> >  \tconst unsigned char *commit_data;\n> >  \tstruct commit_graph_data *graph_data;\n> >  \tuint32_t lex_index;\n> > +\tuint64_t date_high, date_low;\n> >  \n> >  \twhile (pos < g->num_commits_in_base)\n> >  \t\tg = g->base_graph;\n> >  \n> > +\tif (pos >= g->num_commits + g->num_commits_in_base)\n> > +\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n> > +\n> >  \tlex_index = pos - g->num_commits_in_base;\n> >  \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n> >  \n> >  \tgraph_data = commit_graph_data_at(item);\n> >  \tgraph_data->graph_pos = pos;\n> > +\n> > +\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> > +\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> > +\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> > +\n> >  \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> >  }\n> >  \n> > @@ -772,38 +781,22 @@ static int fill_commit_in_graph(struct repository *r,\n> >  {\n> >  \tuint32_t edge_value;\n> >  \tuint32_t *parent_data_ptr;\n> > -\tuint64_t date_low, date_high;\n> >  \tstruct commit_list **pptr;\n> > -\tstruct commit_graph_data *graph_data;\n> >  \tconst unsigned char *commit_data;\n> >  \tuint32_t lex_index;\n> >  \n> >  \twhile (pos < g->num_commits_in_base)\n> >  \t\tg = g->base_graph;\n> >  \n> > -\tif (pos >= g->num_commits + g->num_commits_in_base)\n> > -\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n> > +\tfill_commit_graph_info(item, g, pos);\n> >  \n> > -\t/*\n> > -\t * Store the \"full\" position, but then use the\n> > -\t * \"local\" position for the rest of the calculation.\n> > -\t */\n> > -\tgraph_data = commit_graph_data_at(item);\n> > -\tgraph_data->graph_pos = pos;\n> >  \tlex_index = pos - g->num_commits_in_base;\n> > -\n> >  \tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n> >  \n> >  \titem->object.parsed = 1;\n> >  \n> >  \tset_commit_tree(item, NULL);\n> >  \n> > -\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n> > -\tdate_low = get_be32(commit_data + g->hash_len + 12);\n> > -\titem->date = (timestamp_t)((date_high << 32) | date_low);\n> > -\n> > -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> > -\n> >  \tpptr = &item->parents;\n> >  \n> >  \tedge_value = get_be32(commit_data + g->hash_len);\n> \n> All right, looks good for me.\n> \n> Here second change begins.\n> \n> > diff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\n> > index 3ebb0d3b65..8f41cdc509 100755\n> > --- a/t/t5000-tar-tree.sh\n> > +++ b/t/t5000-tar-tree.sh\n> > @@ -431,11 +431,29 @@ test_expect_success TAR_HUGE,LONG_IS_64BIT 'system tar can read our huge size' '\n> >  \ttest_cmp expect actual\n> >  '\n> >  \n> > +test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n> > +\trm -f .git/index &&\n> > +\techo foo >file &&\n> > +\tgit add file &&\n> > +\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n> > +\t\tgit commit -m \"tempori parendum\"\n> > +'\n> > +\n> > +test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n> > +\tgit archive HEAD >future.tar\n> > +'\n> > +\n> > +test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n> > +\techo 2514 >expect &&\n> > +\ttar_info future.tar | cut -d\" \" -f2 >actual &&\n> > +\ttest_cmp expect actual\n> > +'\n> > +\n> \n> Everything is all right, except we now have duplicated test names.\n> \n> Perhaps in the three following tests we should use 'far-far-future\n> commit' and 'far future mtime' in place of current 'far-future commit'\n> and 'future mtime' for tests checking handling 12-digital ditgits\n> overflow, or add description how far the future is, for example\n> 'far-future commit (2^11 + 1)', etc.\n> \n\nChanged, thanks for pointing this out.\n\n> >  test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n> >  \trm -f .git/index &&\n> >  \techo content >file &&\n> >  \tgit add file &&\n> > -\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n> > +\tGIT_TEST_COMMIT_GRAPH=0 GIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n> >  \t\tgit commit -m \"tempori parendum\"\n> >  '\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"408519","messageId":"85r1pjzejg.fsf@gmail.com","threadId":"53933","inReplyTo":"694ef1ec08d9dc96a74a2631b2710ad206397dbc.1602079786.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 06/10] commit-graph: implement corrected commit date","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-10-27T18:53:23Z","receivedAt":"2020-10-27T18:53:32Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> With most of preparations done, let's implement corrected commit date.\n>\n> The corrected commit date for a commit is defined as:\n>\n> * A commit with no parents (a root commit) has corrected commit date\n>   equal to its committer date.\n> * A commit with at least one parent has corrected commit date equal to\n>   the maximum of its commit date and one more than the largest corrected\n>   commit date among its parents.\n\nAll right.  We might want to say that it fulfills the same reachability\ncriteria as topological level, but perhaps this level of detail is not\nnecessary here.\n\n> As a special case, a root commit with timestamp of zero (01.01.1970\n> 00:00:00Z) has corrected commit date of one, to be able to distinguish\n> from GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit\n> date).\n\nI'm not sure if this special case is really necessary, but it makes for\ncleaner reasoning.\n\n> To minimize the space required to store corrected commit date, Git\n> stores corrected commit date offsets into the commit-graph file. The\n> corrected commit date offset for a commit is defined as the difference\n> between its corrected commit date and actual commit date.\n>\n> Storing corrected commit date requires sizeof(timestamp_t) bytes, which\n> in most cases is 64 bits (uintmax_t). However, corrected commit date\n> offsets can be safely stored using only 32-bits. This halves the size\n> of GDAT chunk, which is a reduction of around 6% in the size of\n> commit-graph file.\n>\n> However, using offsets be problematic if one of commits is malformed but\n> valid and has committerdate of 0 Unix time, as the offset would be the\n> same as corrected commit date and thus require 64-bits to be stored\n> properly.\n>\n> While Git does not write out offsets at this stage, Git stores the\n> corrected commit dates in member generation of struct commit_graph_data.\n> It will begin writing commit date offsets with the introduction of\n> generation data chunk.\n\nAll right.\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nSomewhere in the commit message we should also describe that this commit\nchanges how commit-graph is verified: from checking that the generation\nnumber agrees with _topological level definition_, that is that for a\ngiven commit it is 1 more than maximum of its parents (with the caveat\nthat we need to handle GENERATION_NUMBER_V1_MAX values correctly), to\nchecking that slightly weaker condition fulfilled by both topological\nlevels (generation number v1) and by corrected commit date (generation\nnumber v2) that for a given commit its generation number is 1 more than\nmaximum of its parents or larger.\n\nBut, as far as I understand it, current code does not handle correctly\nGENERATION_NUMBER_V1_MAX case (if we use generation number v1).\n\nOn the other hand we could have simpy use functional check, that\ngeneration number used (which can be v1 or v2, or any similar other)\nfulfills the reachability condition for each edge, which can be\nsimplified to checking that generation(parents) <= generation(commit).\nIf the reachability condition is true for each edge, then it is true for\neach path, and for each commit.\n\n> ---\n>  commit-graph.c | 43 +++++++++++++++++++++++--------------------\n>  1 file changed, 23 insertions(+), 20 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index cedd311024..03948adfce 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -154,11 +154,6 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>  \telse if (generation_a > generation_b)\n>  \t\treturn 1;\n>  \n> -\t/* use date as a heuristic when generations are equal */\n> -\tif (a->date < b->date)\n> -\t\treturn -1;\n> -\telse if (a->date > b->date)\n> -\t\treturn 1;\n\nWhy this change?  It is not described in the commit message.\n\nNote that while this tie-breaking fallback doesn't make much sense for\ncorrected committer date generation number v2, this tie-breaking helps\nif we have to use topological levels (generation number v2).\n\n>  \treturn 0;\n>  }\n>  \n> @@ -1357,10 +1352,14 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\t\t\tctx->commits.nr);\n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>  \t\ttimestamp_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n\nSidenote: I haven't noticed it earlier, but here 'uint32_t' might be\nenough; no need for 'timestamp_t' for 'level' variable.\n\n> +\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n>\n\nAll right, we compute both generation numbers: topological levels and\ncorrected commit date.\n\nI guess we use 'corrected_commit_date' instead of simply 'generation' to\nmake it asier to remember which is which.\n\n>  \t\tdisplay_progress(ctx->progress, i + 1);\n>  \t\tif (level != GENERATION_NUMBER_INFINITY &&\n> -\t\t    level != GENERATION_NUMBER_ZERO)\n> +\t\t    level != GENERATION_NUMBER_ZERO &&\n> +\t\t    corrected_commit_date != GENERATION_NUMBER_INFINITY &&\n> +\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO\n\nStraightforward addition.\n\n> +\t\t    )\n\nWhy this closing parenthesis is now in separated line?\n\n>  \t\t\tcontinue;\n>  \n>  \t\tcommit_list_insert(ctx->commits.list[i], &list);\n> @@ -1369,17 +1368,25 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\tstruct commit_list *parent;\n>  \t\t\tint all_parents_computed = 1;\n>  \t\t\tuint32_t max_level = 0;\n> +\t\t\ttimestamp_t max_corrected_commit_date = 0;\n\nAll right, straightforward addition.\n\n>  \n>  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n>  \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n> -\n\nWhy we have removed this empty line?\n\n> +\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n\nAll right.\n\n>  \t\t\t\tif (level == GENERATION_NUMBER_INFINITY ||\n> -\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n> +\t\t\t\t    level == GENERATION_NUMBER_ZERO ||\n> +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_INFINITY ||\n> +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO\n> +\t\t\t\t    ) {\n\nAll right, same as above.\n\n>  \t\t\t\t\tall_parents_computed = 0;\n>  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n>  \t\t\t\t\tbreak;\n> -\t\t\t\t} else if (level > max_level) {\n> -\t\t\t\t\tmax_level = level;\n> +\t\t\t\t} else {\n> +\t\t\t\t\tif (level > max_level)\n> +\t\t\t\t\t\tmax_level = level;\n> +\n> +\t\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n> +\t\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n>  \t\t\t\t}\n\nAll right, reasonable and straightforward.\n\n>  \t\t\t}\n>  \n> @@ -1389,6 +1396,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n>  \t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n>  \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n> +\n> +\t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n> +\t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n> +\t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n\nAll right.\n\nHere we use the same trick as in previous commit (and as above) to avoid\nany possible overflow, to minimize number of conditionals.  The fact\nthat max_corrected_commit_date might store incorrect value doesn't\nmatter, as it is reset at beginning of this loop.\n\n>  \t\t\t}\n>  \t\t}\n>  \t}\n> @@ -2485,17 +2496,9 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n>  \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n>  \t\t\tcontinue;\n>  \n> -\t\t/*\n> -\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n> -\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n> -\t\t * extra logic in the following condition.\n> -\t\t */\n> -\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n> -\t\t\tmax_generation--;\n> -\n\nPerhaps in the future we should check that both topological levels, and\nalso corrected committer date (if it exists) for correctness according\nto their definition.  Then the above removed part would be restored (but\nwith s/max_generation/max_level/).\n\n>  \t\tgeneration = commit_graph_generation(graph_commit);\n> -\t\tif (generation != max_generation + 1)\n> -\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n> +\t\tif (generation < max_generation + 1)\n> +\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n\nAll right, so we relaxed the check so that it will be fulfilled by\ngeneration number v2 (and also by generation number v1, as it implies\nthe more strict check for v1).\n\nWhat would happen however if generation holds topological levels, and it\nis GENERATION_NUMBER_V1_MAX for at least one parent, which means it is\nGENERATION_NUMBER_V1_MAX for a commit?  As you can check, the condition\nwould be true: GENERATION_NUMBER_V1_MAX < GENERATION_NUMBER_V1_MAX + 1,\nso the `git commit-graph verify` would incorrectly say that there is\na problem with generation number, while there isn't one (false positive\ndetection of error).\n\nSidenote: I think we don't have to worry about having to introduce\nGENERATION_NUMBER_V2_MAX, as the in-memory size (of reconstructed from\ndisck representation) corrected commiter date is the same as of commiter\ndate itself, plus some, and I don't see us coming close to 64-bit limit\nof timestamp_t for commit dates.\n\n>  \t\t\t\t     oid_to_hex(&cur_oid),\n>  \t\t\t\t     generation,\n>  \t\t\t\t     max_generation + 1);\n\nBest,\n-- \nJakub Narębski\n"},{"id":"408738","messageId":"854kmbx4pi.fsf@gmail.com","threadId":"53933","inReplyTo":"b903efe2ea11bc0b7e1ef8f239ed34f72caa4f03.1602079786.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 07/10] commit-graph: implement generation data chunk","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-10-30T12:45:29Z","receivedAt":"2020-10-30T12:56:07Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Tl;dr summary: the code writing GDOV chunk could be made more performant\n(I think), but that could be left for the future commit.\n\n\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> As discovered by Ævar, we cannot increment graph version to\n> distinguish between generation numbers v1 and v2 [1]. Thus, one of\n> pre-requistes before implementing generation number was to distinguish\n> between graph versions in a backwards compatible manner.\n\nMinor nitpick: I think you meant \"implementing generation number v2\",\nto be more precise.\n\n>\n> We are going to introduce a new chunk called Generation Data chunk (or\n\nVery minor nitpick: perhaps s/Generation Data/Generation DATa/, to provide\nmnemonics for chunk name.\n\n> GDAT). GDAT stores corrected committer date offsets whereas CDAT will\n> still store topological level.\n\nMinor nitpick: I think the second sentence should use consistent\ngrammatical tense (but I am not a native English speaker); also\ns/level/levels/:\n\n    GDAT will store corrected committer date offsets, whereas CDAT will\n    still store topological levels.\n\nBut it is perfectly understandable as it is.\n\n>\n> Old Git does not understand GDAT chunk and would ignore it, reading\n> topological levels from CDAT. New Git can parse GDAT and take advantage\n> of newer generation numbers, falling back to topological levels when\n> GDAT chunk is missing (as it would happen with a commit graph written\n> by old Git).\n\nMinor nitpick: I think we use commit-graph with dash when writing about\nthe commit-graph file, like below.\n\n>\n> We introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\n> which forces commit-graph file to be written without generation data\n> chunk to emulate a commit-graph file written by old Git.\n\nAll right.\n\n>\n> While storing corrected commit date offset instead of the corrected\n> commit date saves us 4 bytes per commit, it's possible for the offsets\n> to overflow the 4-bytes allocated. As such overflows are exceedingly\n> rare, we use the following overflow management scheme:\n\nPerhaps it would be good idea to write the idea in full from start, as\nthe commit message is intended to be read stadalone and not in the\ncontext of the patch series.  On the other hand it might be too much\ndetail in already [necessarily] lengthty commit message.\n\nPerhaps something like the following proposal would read better.\n\n  To minimize the space required to store corrected commit date, Git\n  stores corrected commit date offsets into the commit-graph file,\n  instead of corrected commit dates themselves. This saves us 4 bytes\n  per commit, decreasing the GDAT chunk size by half, but it's possible\n  for the offset to overflow the 4-bytes allocated for storage. As such\n  overflows are and should be exceedingly rare, we use the following\n  overflow management scheme:\n\n\nNOTE: this overflow handling is a *new* code (or new-ish code, as it is\ninspired and similar to EDGE chunk data handling), so it needs more\ncareful review.\n\n>\n> We introduce a new commit-graph chunk, GENERATION_DATA_OVERFLOW ('GDOV')\n\nMinor issue: why GENERATION_DATA_OVERFLOW and not Generation Data\nOVerflow, like for the GDAT chunk?\n\n> to store corrected commit dates for commits with offsets greater than\n> GENERATION_NUMBER_V2_OFFSET_MAX.\n>\n> If the offset is greater than GENERATION_NUMBER_V2_OFFSET_MAX, we set\n> the MSB of the offset and the other bits store the position of corrected\n> commit date in GDOV chunk, similar to how Extra Edge List is maintained.\n>\n> We test the overflow-related code with the following repo history:\n>\n>            F - N - U\n>           /         \\\n> U - N - U            N\n>          \\          /\n>            N - F - N\n\nDo we need such complex history? I guess we need to test the handling of\nmerge commits too.\n\n>\n> Where the commits denoted by U have committer date of zero seconds\n> since Unix epoch, the commits denoted by N have committer date of\n> 1112354055 (default committer date for the test suite) seconds since\n> Unix epoch and the commits denoted by F have committer date of\n> (2 ^ 31 - 2) seconds since Unix epoch.\n>\n> The largest offset observed is 2 ^ 31, just large enough to overflow.\n>\n> [1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c                | 98 +++++++++++++++++++++++++++++++++--\n>  commit-graph.h                |  3 ++\n>  commit.h                      |  1 +\n>  t/README                      |  3 ++\n>  t/helper/test-read-graph.c    |  4 ++\n>  t/t4216-log-bloom.sh          |  4 +-\n>  t/t5318-commit-graph.sh       | 70 ++++++++++++++++++++-----\n>  t/t5324-split-commit-graph.sh | 12 ++---\n>  t/t6600-test-reach.sh         | 68 +++++++++++++-----------\n>  9 files changed, 206 insertions(+), 57 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 03948adfce..71d0b243db 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -38,11 +38,13 @@ void git_test_write_commit_graph_or_die(void)\n>  #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n>  #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n>  #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n> +#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n> +#define GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW 0x47444f56 /* \"GDOV\" */\n>  #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n>  #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n>  #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n>  #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n> -#define MAX_NUM_CHUNKS 7\n> +#define MAX_NUM_CHUNKS 9\n\nAll right.\n\n>  \n>  #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n>  \n> @@ -61,6 +63,8 @@ void git_test_write_commit_graph_or_die(void)\n>  #define GRAPH_MIN_SIZE (GRAPH_HEADER_SIZE + 4 * GRAPH_CHUNKLOOKUP_WIDTH \\\n>  \t\t\t+ GRAPH_FANOUT_SIZE + the_hash_algo->rawsz)\n>  \n> +#define CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW (1ULL << 31)\n> +\n\nAll right, though the naming convention is different from the one used\nfor EDGE chunk: GRAPH_EXTRA_EDGES_NEEDED and GRAPH_EDGE_LAST_MASK.\n\n>  /* Remember to update object flag allocation in object.h */\n>  #define REACHABLE       (1u<<15)\n>  \n> @@ -385,6 +389,20 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n>  \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n>  \t\t\tbreak;\n>  \n> +\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n> +\t\t\tif (graph->chunk_generation_data)\n> +\t\t\t\tchunk_repeated = 1;\n> +\t\t\telse\n> +\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n> +\t\t\tbreak;\n> +\n> +\t\tcase GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW:\n> +\t\t\tif (graph->chunk_generation_data_overflow)\n> +\t\t\t\tchunk_repeated = 1;\n> +\t\t\telse\n> +\t\t\t\tgraph->chunk_generation_data_overflow = data + chunk_offset;\n> +\t\t\tbreak;\n> +\n\nNecessary but unavoidable boilerplate for adding new chunks to the\ncommit-graph file format.  All right.\n\n>  \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n>  \t\t\tif (graph->chunk_extra_edges)\n>  \t\t\t\tchunk_repeated = 1;\n> @@ -745,8 +763,8 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  {\n>  \tconst unsigned char *commit_data;\n>  \tstruct commit_graph_data *graph_data;\n> -\tuint32_t lex_index;\n> -\tuint64_t date_high, date_low;\n> +\tuint32_t lex_index, offset_pos;\n> +\tuint64_t date_high, date_low, offset;\n\nAll right, we are adding two new variables: `offset` to read data stored\nin GDAT chunk, and `offset_pos` to help read data from GDOV chunk if\nnecessary i.e. to handle overflow in corrected commit data offset\nstorage.\n\n>  \n>  \twhile (pos < g->num_commits_in_base)\n>  \t\tg = g->base_graph;\n> @@ -764,7 +782,16 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \tdate_low = get_be32(commit_data + g->hash_len + 12);\n>  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n>  \n> -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> +\tif (g->chunk_generation_data) {\n> +\t\toffset = (timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n\nStyle: why space after the `(timestamp_t)` cast operator?\n\nThough CodingGuidelines do not say anything on this topic... perhaps the\nspace after cast operator makes it more readable?\n\n> +\n> +\t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n\nAll right, so the CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW is equivalent of\nGRAPH_EXTRA_EDGES_NEEDED.\n\n> +\t\t\toffset_pos = offset ^ CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW;\n\nHmmm... instead of using bitwise and on an equivalent to the\nGRAPH_EDGE_LAST_MASK, we utilize the fact that we know that the MSB bit\nis set, so we can clear it with bitwise xor.  Clever trick.\n\n> +\t\t\tgraph_data->generation = get_be64(g->chunk_generation_data_overflow + 8 * offset_pos);\n> +\t\t} else\n> +\t\t\tgraph_data->generation = item->date + offset;\n\nAll right, this handles the case when we have generation number v2, with\nor without overflow.\n\n> +\t} else\n> +\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n\nAll right, this handles the case where we have only generation number\nv1, like for commit-graph file written by old Git.\n\n>  \n>  \tif (g->topo_levels)\n>  \t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n> @@ -942,6 +969,7 @@ struct write_commit_graph_context {\n>  \tstruct packed_oid_list oids;\n>  \tstruct packed_commit_list commits;\n>  \tint num_extra_edges;\n> +\tint num_generation_data_overflows;\n>  \tunsigned long approx_nr_objects;\n>  \tstruct progress *progress;\n>  \tint progress_done;\n> @@ -960,7 +988,8 @@ struct write_commit_graph_context {\n>  \t\t report_progress:1,\n>  \t\t split:1,\n>  \t\t changed_paths:1,\n> -\t\t order_by_pack:1;\n> +\t\t order_by_pack:1,\n> +\t\t write_generation_data:1;\n>  \n>  \tstruct topo_level_slab *topo_levels;\n>  \tconst struct commit_graph_opts *opts;\n\nAll right, this adds necessary fields to `struct write_commit_graph_context`.\n\n> @@ -1120,6 +1149,44 @@ static int write_graph_chunk_data(struct hashfile *f,\n>  \treturn 0;\n>  }\n>  \n> +static int write_graph_chunk_generation_data(struct hashfile *f,\n> +\t\t\t\t\t      struct write_commit_graph_context *ctx)\n> +{\n> +\tint i, num_generation_data_overflows = 0;\n\nMinor nitpick: in my opinion there should be empty line here, between\nthe variables declaration and the code... however not all\nwrite_graph_chunk_*() functions have it.\n\n> +\tfor (i = 0; i < ctx->commits.nr; i++) {\n> +\t\tstruct commit *c = ctx->commits.list[i];\n> +\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n> +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n\nAll right.\n\n> +\n> +\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n> +\t\t\toffset = CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW | num_generation_data_overflows;\n> +\t\t\tnum_generation_data_overflows++;\n> +\t\t}\n\nHmmm... shouldn't we store these commits that need overflow handling\n(with corrected commit date offset greater than GENERATION_NUMBER_V2_OFFSET_MAX)\nin a list or a queue, to remember them for writing GDOV chunk?\n\nWe could store oids, or we could store commits themselves, for example:\n\n\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n\t\t\toffset = CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW | num_generation_data_overflows;\n\t\t\tnum_generation_data_overflows++;\n\n\t\t\tALLOC_GROW(ctx->gdov_commits.list, ctx->gdov_commits.nr + 1, ctx->gdov_commits.alloc);\n\t\t\tctx->commits.list[ctx->gdov_commits.nr] = c;\n            ctx->gdov_commits.nr++;\n\t\t}\n\nThough in the above proposal we could get rid of `num_generation_data_overflows`, \nas it should be the same as `ctx->gdov_commits.nr`.\n\nI have called the extra commit list member of write_commit_graph_context\n`gdov_commits`, but perhaps a better name would be `commits_gen_v2_overflow`, \nor similar more descriptive name.\n\n> +\n> +\t\thashwrite_be32(f, offset);\n> +\t}\n> +\n> +\treturn 0;\n> +}\n\nAll right.\n\n> +\n> +static int write_graph_chunk_generation_data_overflow(struct hashfile *f,\n> +\t\t\t\t\t\t       struct write_commit_graph_context *ctx)\n> +{\n> +\tint i;\n> +\tfor (i = 0; i < ctx->commits.nr; i++) {\n\nHere we loop over *all* commits again, instead of looping over those\nvery rare commits that need overflow handling for their corrected commit\ndate data.\n\nThough this possible performance issue^* could be fixed in the future commit.\n\n*) It needs to be actually benchmarked which version is faster.\n\nWith the change proposed above (and required changes to the `struct\nwrite_commit_graph_context`) it could look like this:\n\n\tfor (i = 0; i < ctx->gcov_commits.nr; i++) {\n\n\n> +\t\tstruct commit *c = ctx->commits.list[i];\n> +\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n> +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n> +\n> +\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n> +\t\t\thashwrite_be32(f, offset >> 32);\n> +\t\t\thashwrite_be32(f, (uint32_t) offset);\n> +\t\t}\n> +\t}\n\nThe above would be as simple as the following:\n\n\t\tstruct commit *c = ctx->gcov_commits.list[i];\n\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n\n\t\thashwrite_be64(f, offset);\n\nAssumming that there would be hashwrite_be64(), it would be the\nfollowing otherwise:\n\n\t\thashwrite_be32(f, offset >> 32);\n\t\thashwrite_be32(f, (uint32_t)offset);\n\n> +\n> +\treturn 0;\n> +}\n> +\n>  static int write_graph_chunk_extra_edges(struct hashfile *f,\n>  \t\t\t\t\t struct write_commit_graph_context *ctx)\n>  {\n> @@ -1399,7 +1466,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \n>  \t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n>  \t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n> +\n\nThis is a bit unrelated change, adding this empty line.\n\n>  \t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n> +\n> +\t\t\t\tif (commit_graph_data_at(current)->generation - current->date > GENERATION_NUMBER_V2_OFFSET_MAX)\n> +\t\t\t\t\tctx->num_generation_data_overflows++;\n\nAll right, we need to track number of commits that need overflow\nhandling for generation number v2 to know what size GDOV chunk would\nneed to be.\n\n>  \t\t\t}\n>  \t\t}\n>  \t}\n> @@ -1765,6 +1836,21 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \tchunks[2].id = GRAPH_CHUNKID_DATA;\n>  \tchunks[2].size = (hashsz + 16) * ctx->commits.nr;\n>  \tchunks[2].write_fn = write_graph_chunk_data;\n> +\n> +\tif (git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0))\n> +\t\tctx->write_generation_data = 0;\n\nAll right, here we handle GIT_TEST_COMMIT_GRAPH_NO_GDAT.\n\n> +\tif (ctx->write_generation_data) {\n> +\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA;\n> +\t\tchunks[num_chunks].size = sizeof(uint32_t) * ctx->commits.nr;\n> +\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data;\n> +\t\tnum_chunks++;\n> +\t}\n\nAll right, the GDAT chunk consist of <number of commits> entries.\n\n> +\tif (ctx->num_generation_data_overflows) {\n> +\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW;\n> +\t\tchunks[num_chunks].size = sizeof(timestamp_t) * ctx->num_generation_data_overflows;\n> +\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data_overflow;\n> +\t\tnum_chunks++;\n> +\t}\n\nAll right, that's what num_generation_data_overflows was for.\n\n>  \tif (ctx->num_extra_edges) {\n>  \t\tchunks[num_chunks].id = GRAPH_CHUNKID_EXTRAEDGES;\n>  \t\tchunks[num_chunks].size = 4 * ctx->num_extra_edges;\n> @@ -2170,6 +2256,8 @@ int write_commit_graph(struct object_directory *odb,\n>  \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n>  \tctx->opts = opts;\n>  \tctx->total_bloom_filter_data_size = 0;\n> +\tctx->write_generation_data = 1;\n> +\tctx->num_generation_data_overflows = 0;\n>  \n>  \tbloom_settings.bits_per_entry = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\",\n>  \t\t\t\t\t\t      bloom_settings.bits_per_entry);\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 2e9aa7824e..19a02001fd 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -6,6 +6,7 @@\n>  #include \"oidset.h\"\n>  \n>  #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n> +#define GIT_TEST_COMMIT_GRAPH_NO_GDAT \"GIT_TEST_COMMIT_GRAPH_NO_GDAT\"\n>  #define GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE \"GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE\"\n>  #define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n>  \n> @@ -68,6 +69,8 @@ struct commit_graph {\n>  \tconst uint32_t *chunk_oid_fanout;\n>  \tconst unsigned char *chunk_oid_lookup;\n>  \tconst unsigned char *chunk_commit_data;\n> +\tconst unsigned char *chunk_generation_data;\n> +\tconst unsigned char *chunk_generation_data_overflow;\n\nAll right, two new chunks: GDAT and GDOV.\n\n>  \tconst unsigned char *chunk_extra_edges;\n>  \tconst unsigned char *chunk_base_graphs;\n>  \tconst unsigned char *chunk_bloom_indexes;\n> diff --git a/commit.h b/commit.h\n> index 33c66b2177..251d877fcf 100644\n> --- a/commit.h\n> +++ b/commit.h\n> @@ -14,6 +14,7 @@\n>  #define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n>  #define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n>  #define GENERATION_NUMBER_ZERO 0\n> +#define GENERATION_NUMBER_V2_OFFSET_MAX ((1ULL << 31) - 1)\n\nShould we use this form, or hexadecimal constant?\n\n   #define GENERATION_NUMBER_V2_OFFSET_MAX 0x7FFFFFFF\n\nBut I think the current definition is more explicit: all bits set to one\nexcept for the most significant digit.  All right.\n\n>  \n>  struct commit_list {\n>  \tstruct commit *item;\n> diff --git a/t/README b/t/README\n> index 2adaf7c2d2..975c054bc9 100644\n> --- a/t/README\n> +++ b/t/README\n> @@ -379,6 +379,9 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n>  be written after every 'git commit' command, and overrides the\n>  'core.commitGraph' setting to true.\n>  \n> +GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n> +commit-graph to be written without generation data chunk.\n> +\n\nAll right. Nice have it documented.\n\n>  GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n>  commit-graph write to compute and write changed path Bloom filters for\n>  every 'git commit-graph write', as if the `--changed-paths` option was\n> diff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\n> index 5f585a1725..75927b2c81 100644\n> --- a/t/helper/test-read-graph.c\n> +++ b/t/helper/test-read-graph.c\n> @@ -33,6 +33,10 @@ int cmd__read_graph(int argc, const char **argv)\n>  \t\tprintf(\" oid_lookup\");\n>  \tif (graph->chunk_commit_data)\n>  \t\tprintf(\" commit_metadata\");\n> +\tif (graph->chunk_generation_data)\n> +\t\tprintf(\" generation_data\");\n> +\tif (graph->chunk_generation_data_overflow)\n> +\t\tprintf(\" generation_data_overflow\");\n>  \tif (graph->chunk_extra_edges)\n>  \t\tprintf(\" extra_edges\");\n>  \tif (graph->chunk_bloom_indexes)\n\nAll right, updating `test-tool read-graph` with new chunks.\n\n> diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n> index d11040ce41..dbde016188 100755\n> --- a/t/t4216-log-bloom.sh\n> +++ b/t/t4216-log-bloom.sh\n> @@ -40,11 +40,11 @@ test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n>  '\n>  \n>  graph_read_expect () {\n> -\tNUM_CHUNKS=5\n> +\tNUM_CHUNKS=6\n>  \tcat >expect <<- EOF\n\nSidenote: I have just noticed this, and as I see it is not something you\nwrote, but usually we write it with no space after the dash and before\n'EOF':\n\n   \tcat >expect <<-EOF\n\n>  \theader: 43475048 1 $(test_oid oid_version) $NUM_CHUNKS 0\n>  \tnum_commits: $1\n> -\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n> +\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n>  \tEOF\n>  \ttest-tool read-graph >actual &&\n>  \ttest_cmp expect actual\n\nAll right, updating expect value for `test-tool read-graph` in the usual\ncase, with generation number chunk.\n\n> diff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\n> index 2ed0c1544d..0328e98564 100755\n> --- a/t/t5318-commit-graph.sh\n> +++ b/t/t5318-commit-graph.sh\n> @@ -76,7 +76,7 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n>  graph_read_expect() {\n>  \tOPTIONAL=\"\"\n>  \tNUM_CHUNKS=3\n> -\tif test ! -z $2\n> +\tif test ! -z \"$2\"\n\nAll right, that is straighforward fix, which is now needed.\n\n>  \tthen\n>  \t\tOPTIONAL=\" $2\"\n>  \t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n> @@ -103,14 +103,14 @@ test_expect_success 'exit with correct error on bad input to --stdin-commits' '\n>  \t# valid commit and tree OID\n>  \tgit rev-parse HEAD HEAD^{tree} >in &&\n>  \tgit commit-graph write --stdin-commits <in &&\n> -\tgraph_read_expect 3\n> +\tgraph_read_expect 3 generation_data\n>  '\n>  \n>  test_expect_success 'write graph' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit commit-graph write &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"3\"\n> +\tgraph_read_expect \"3\" generation_data\n>  '\n>  \n>  test_expect_success POSIXPERM 'write graph has correct permissions' '\n> @@ -219,7 +219,7 @@ test_expect_success 'write graph with merges' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit commit-graph write &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"10\" \"extra_edges\"\n> +\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n>  '\n>  \n>  graph_git_behavior 'merge 1 vs 2' full merge/1 merge/2\n> @@ -254,7 +254,7 @@ test_expect_success 'write graph with new commit' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit commit-graph write &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"11\" \"extra_edges\"\n> +\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n>  '\n>  \n>  graph_git_behavior 'full graph, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -264,7 +264,7 @@ test_expect_success 'write graph with nothing new' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit commit-graph write &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"11\" \"extra_edges\"\n> +\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n>  '\n>  \n>  graph_git_behavior 'cleared graph, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -274,7 +274,7 @@ test_expect_success 'build graph from latest pack with closure' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tcat new-idx | git commit-graph write --stdin-packs &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"9\" \"extra_edges\"\n> +\tgraph_read_expect \"9\" \"generation_data extra_edges\"\n>  '\n>  \n>  graph_git_behavior 'graph from pack, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -287,7 +287,7 @@ test_expect_success 'build graph from commits with closure' '\n>  \tgit rev-parse merge/1 >>commits-in &&\n>  \tcat commits-in | git commit-graph write --stdin-commits &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"6\"\n> +\tgraph_read_expect \"6\" \"generation_data\"\n>  '\n>  \n>  graph_git_behavior 'graph from commits, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -297,7 +297,7 @@ test_expect_success 'build graph from commits with append' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit rev-parse merge/3 | git commit-graph write --stdin-commits --append &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"10\" \"extra_edges\"\n> +\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n>  '\n>  \n>  graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -307,7 +307,7 @@ test_expect_success 'build graph using --reachable' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n>  \tgit commit-graph write --reachable &&\n>  \ttest_path_is_file $objdir/info/commit-graph &&\n> -\tgraph_read_expect \"11\" \"extra_edges\"\n> +\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n>  '\n>  \n>  graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n> @@ -328,7 +328,7 @@ test_expect_success 'write graph in bare repo' '\n>  \tcd \"$TRASH_DIRECTORY/bare\" &&\n>  \tgit commit-graph write &&\n>  \ttest_path_is_file $baredir/info/commit-graph &&\n> -\tgraph_read_expect \"11\" \"extra_edges\"\n> +\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n>  '\n\nAll those just add \"generation_data\" (aka GDAT) to expected chunks. All\nright.\n\n>  \n>  graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n> @@ -454,8 +454,9 @@ test_expect_success 'warn on improper hash version' '\n>  \n>  test_expect_success 'git commit-graph verify' '\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n> -\tgit rev-parse commits/8 | git commit-graph write --stdin-commits &&\n> -\tgit commit-graph verify >output\n> +\tgit rev-parse commits/8 | GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --stdin-commits &&\n> +\tgit commit-graph verify >output &&\n\nAll right, this simply adds GIT_TEST_COMMIT_GRAPH_NO_GDAT=1.  I assume\nthis is needed because this test is also setup for the following commits\n_without_ even saying that in the test name (bad practice, in my\nopinion), and the comment above this test says the following:\n\n  # the verify tests below expect the commit-graph to contain\n  # exactly the commits reachable from the commits/8 branch.\n  # If the file changes the set of commits in the list, then the\n  # offsets into the binary file will result in different edits\n  # and the tests will likely break.\n\nSo the following tests are fragile (though perhaps unavoidably fragile),\nand without this change they would not work, I assume.\n\n> +\tgraph_read_expect 9 extra_edges\n\nI guess that this is here to check that GIT_TEST_COMMIT_GRAPH_NO_GDAT=1\nwork as intended, and that the following \"verify\" tests wouldn't break.\nI understand its necessity, even if I don't quite like having a test\nthat checks multiple things.  This is a minor issue, though.\n\nAll right.\n\n\nWe might want to have a separate test that checks that we get\ncommit-graph with and without GDAT chunk depending on whether we use\nGIT_TEST_COMMIT_GRAPH_NO_GDAT=1.  On the other hand, this environment\nvariable is there purely for tests, so the question is should we test\nthe test infrastructure?\n\n>  '\n>  \n>  NUM_COMMITS=9\n> @@ -741,4 +742,47 @@ test_expect_success 'corrupt commit-graph write (missing tree)' '\n>  \t)\n>  '\n>  \n> +test_commit_with_date() {\n> +  file=\"$1.t\" &&\n> +  echo \"$1\" >\"$file\" &&\n> +  git add \"$file\" &&\n> +  GIT_COMMITTER_DATE=\"$2\" GIT_AUTHOR_DATE=\"$2\" git commit -m \"$1\"\n> +  git tag \"$1\"\n> +}\n\nHere we add a helper function.  All right.\n\nI wonder though if it wouldn't be a better idea to add `--date <date>`\noption to the test_commit() function in test-lib-functions.sh (which\noption would set GIT_COMMITTER_DATE and GIT_AUTHOR_DATE, and also\nset notick=yes).\n\nFor example:\n\ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex f1ae935fee..a1f9a2b09b 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -202,6 +202,12 @@ test_commit () {\n \t\t--signoff)\n \t\t\tsignoff=\"$1\"\n \t\t\t;;\n+        --date)\n+            notick=yes\n+            GIT_COMMITTER_DATE=\"$2\"\n+            GIT_AUTHOR_DATE=\"$2\"\n+            shift\n+            ;;\n \t\t-C)\n \t\t\tindir=\"$2\"\n \t\t\tshift\n\n\n> +\n\nIt would be nice to have there comment describing the shape of the\nrevision history we generate here, that currenly is present only in the\ncommmit message.\n\n# We test the overflow-related code with the following repo history:\n#\n#               4:F - 5:N - 6:U\n#              /               \\\n# 1:U - 2:N - 3:U               M:N\n#              \\               /\n#               7:N - 8:F - 9:N\n#\n# Here the commits denoted by U have committer date of zero seconds\n# since Unix epoch, the commits denoted by N have committer date\n# starting from 1112354055 seconds since Unix epoch (default committer\n# date for the test suite), and the commits denoted by F have committer\n# date of (2 ^ 31 - 2) seconds since Unix epoch.\n#\n# The largest offset observed is 2 ^ 31, just large enough to overflow.\n#\n\n> +test_expect_success 'overflow corrected commit date offset' '\n> +\tobjdir=\".git/objects\" &&\n> +\tUNIX_EPOCH_ZERO=\"1970-01-01 00:00 +0000\" &&\n> +\tFUTURE_DATE=\"@2147483646 +0000\" &&\n\nIt is a bit funny to see UNIX_EPOCH_ZERO spelled one way, and\nFUTURE_DATE other way.\n\nWouldn't be more readable to use UNIX_EPOCH_ZERO=\"@0 +0000\"?\n\n> +\ttest_oid_cache <<-EOF &&\n> +\toid_version sha1:1\n> +\toid_version sha256:2\n> +\tEOF\n> +\tcd \"$TRASH_DIRECTORY\" &&\n> +\tmkdir repo &&\n> +\tcd repo &&\n> +\tgit init &&\n> +\ttest_commit_with_date 1 \"$UNIX_EPOCH_ZERO\" &&\n> +\ttest_commit 2 &&\n> +\ttest_commit_with_date 3 \"$UNIX_EPOCH_ZERO\" &&\n> +\tgit commit-graph write --reachable &&\n> +\tgraph_read_expect 3 generation_data &&\n> +\ttest_commit_with_date 4 \"$FUTURE_DATE\" &&\n> +\ttest_commit 5 &&\n> +\ttest_commit_with_date 6 \"$UNIX_EPOCH_ZERO\" &&\n> +\tgit branch left &&\n> +\tgit reset --hard 3 &&\n> +\ttest_commit 7 &&\n> +\ttest_commit_with_date 8 \"$FUTURE_DATE\" &&\n> +\ttest_commit 9 &&\n> +\tgit branch right &&\n> +\tgit reset --hard 3 &&\n> +\tgit merge left right &&\n\nWe have test_merge() function in test-lib-functions.sh, perhaps we\nshould use it here.\n\n> +\tgit commit-graph write --reachable &&\n> +\tgraph_read_expect 10 \"generation_data generation_data_overflow\" &&\n\nAll right, we write the commit-graph and check that it has both GDAT and\nGDOV chunks present.\n\n> +\tgit commit-graph verify\n\nAll right, we checks that created commit graph with GDAT and GDOV passes\n'git commit-graph verify` checks.\n\n> +'\n> +\n> +graph_git_behavior 'overflow corrected commit date offset' repo left right\n\nAll right, here we compare the Git behavior with the commit-graph to the\nbehavior without it... however I think that those two tests really\nshould have distinct (different) test names. Currently they both use\n'overflow corrected commit date offset'.\n\n> +\n>  test_done\n> diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> index c334ee9155..651df89ab2 100755\n> --- a/t/t5324-split-commit-graph.sh\n> +++ b/t/t5324-split-commit-graph.sh\n> @@ -13,11 +13,11 @@ test_expect_success 'setup repo' '\n>  \tinfodir=\".git/objects/info\" &&\n>  \tgraphdir=\"$infodir/commit-graphs\" &&\n>  \ttest_oid_cache <<-EOM\n> -\tshallow sha1:1760\n> -\tshallow sha256:2064\n> +\tshallow sha1:2132\n> +\tshallow sha256:2436\n>  \n> -\tbase sha1:1376\n> -\tbase sha256:1496\n> +\tbase sha1:1408\n> +\tbase sha256:1528\n>  \n>  \toid_version sha1:1\n>  \toid_version sha256:2\n> @@ -31,9 +31,9 @@ graph_read_expect() {\n>  \t\tNUM_BASE=$2\n>  \tfi\n>  \tcat >expect <<- EOF\n> -\theader: 43475048 1 $(test_oid oid_version) 3 $NUM_BASE\n> +\theader: 43475048 1 $(test_oid oid_version) 4 $NUM_BASE\n>  \tnum_commits: $1\n> -\tchunks: oid_fanout oid_lookup commit_metadata\n> +\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n>  \tEOF\n>  \ttest-tool read-graph >output &&\n>  \ttest_cmp expect output\n\nAll right, we now expect the commit graph to include the GDAT chunk...\nthough shouldn't be there old expected value for no GDAT, for future\ntests?  But perhaps this is not necessary.\n\nNote that I have not checked the details, but it looks OK to me.\n\n> diff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\n> index f807276337..e2d33a8a4c 100755\n> --- a/t/t6600-test-reach.sh\n> +++ b/t/t6600-test-reach.sh\n> @@ -55,10 +55,13 @@ test_expect_success 'setup' '\n>  \tgit show-ref -s commit-5-5 | git commit-graph write --stdin-commits &&\n>  \tmv .git/objects/info/commit-graph commit-graph-half &&\n>  \tchmod u+w commit-graph-half &&\n> +\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable &&\n> +\tmv .git/objects/info/commit-graph commit-graph-no-gdat &&\n> +\tchmod u+w commit-graph-no-gdat &&\n\nAll right, this prepares for testing one more mode.  The run_all_modes()\nfunction would test the following cases:\n - no commit-graph\n - commit-graph for all commits, with GDAT\n - commit-graph with half of commits, with GDAT\n - commit-graph for all commits, without GDAT\n\n>  \tgit config core.commitGraph true\n>  '\n>  \n> -run_three_modes () {\n> +run_all_modes () {\n>  \ttest_when_finished rm -rf .git/objects/info/commit-graph &&\n>  \t\"$@\" <input >actual &&\n>  \ttest_cmp expect actual &&\n> @@ -67,11 +70,14 @@ run_three_modes () {\n>  \ttest_cmp expect actual &&\n>  \tcp commit-graph-half .git/objects/info/commit-graph &&\n>  \t\"$@\" <input >actual &&\n> +\ttest_cmp expect actual &&\n> +\tcp commit-graph-no-gdat .git/objects/info/commit-graph &&\n> +\t\"$@\" <input >actual &&\n>  \ttest_cmp expect actual\n>  }\n>  \n> -test_three_modes () {\n> -\trun_three_modes test-tool reach \"$@\"\n> +test_all_modes () {\n> +\trun_all_modes test-tool reach \"$@\"\n>  }\n\nAll right.\n\nThough to reduce \"noise\" in this patch, the rename of run_three_modes()\nto run_all_modes() and test_three_modes() to test_all_modes() could have\nbeen done in a separate preparatory patch.  It would be pure refactoring\npatch, without introducing any new functionality.\n\n>  \n>  test_expect_success 'ref_newer:miss' '\n> @@ -80,7 +86,7 @@ test_expect_success 'ref_newer:miss' '\n>  \tB:commit-4-9\n>  \tEOF\n>  \techo \"ref_newer(A,B):0\" >expect &&\n> -\ttest_three_modes ref_newer\n> +\ttest_all_modes ref_newer\n>  '\n>  \n>  test_expect_success 'ref_newer:hit' '\n> @@ -89,7 +95,7 @@ test_expect_success 'ref_newer:hit' '\n>  \tB:commit-2-3\n>  \tEOF\n>  \techo \"ref_newer(A,B):1\" >expect &&\n> -\ttest_three_modes ref_newer\n> +\ttest_all_modes ref_newer\n>  '\n>  \n>  test_expect_success 'in_merge_bases:hit' '\n> @@ -98,7 +104,7 @@ test_expect_success 'in_merge_bases:hit' '\n>  \tB:commit-8-8\n>  \tEOF\n>  \techo \"in_merge_bases(A,B):1\" >expect &&\n> -\ttest_three_modes in_merge_bases\n> +\ttest_all_modes in_merge_bases\n>  '\n>  \n>  test_expect_success 'in_merge_bases:miss' '\n> @@ -107,7 +113,7 @@ test_expect_success 'in_merge_bases:miss' '\n>  \tB:commit-5-9\n>  \tEOF\n>  \techo \"in_merge_bases(A,B):0\" >expect &&\n> -\ttest_three_modes in_merge_bases\n> +\ttest_all_modes in_merge_bases\n>  '\n>  \n>  test_expect_success 'in_merge_bases_many:hit' '\n> @@ -117,7 +123,7 @@ test_expect_success 'in_merge_bases_many:hit' '\n>  \tX:commit-5-7\n>  \tEOF\n>  \techo \"in_merge_bases_many(A,X):1\" >expect &&\n> -\ttest_three_modes in_merge_bases_many\n> +\ttest_all_modes in_merge_bases_many\n>  '\n>  \n>  test_expect_success 'in_merge_bases_many:miss' '\n> @@ -127,7 +133,7 @@ test_expect_success 'in_merge_bases_many:miss' '\n>  \tX:commit-8-6\n>  \tEOF\n>  \techo \"in_merge_bases_many(A,X):0\" >expect &&\n> -\ttest_three_modes in_merge_bases_many\n> +\ttest_all_modes in_merge_bases_many\n>  '\n>  \n>  test_expect_success 'in_merge_bases_many:miss-heuristic' '\n> @@ -137,7 +143,7 @@ test_expect_success 'in_merge_bases_many:miss-heuristic' '\n>  \tX:commit-6-6\n>  \tEOF\n>  \techo \"in_merge_bases_many(A,X):0\" >expect &&\n> -\ttest_three_modes in_merge_bases_many\n> +\ttest_all_modes in_merge_bases_many\n>  '\n>  \n>  test_expect_success 'is_descendant_of:hit' '\n> @@ -148,7 +154,7 @@ test_expect_success 'is_descendant_of:hit' '\n>  \tX:commit-1-1\n>  \tEOF\n>  \techo \"is_descendant_of(A,X):1\" >expect &&\n> -\ttest_three_modes is_descendant_of\n> +\ttest_all_modes is_descendant_of\n>  '\n>  \n>  test_expect_success 'is_descendant_of:miss' '\n> @@ -159,7 +165,7 @@ test_expect_success 'is_descendant_of:miss' '\n>  \tX:commit-7-6\n>  \tEOF\n>  \techo \"is_descendant_of(A,X):0\" >expect &&\n> -\ttest_three_modes is_descendant_of\n> +\ttest_all_modes is_descendant_of\n>  '\n>  \n>  test_expect_success 'get_merge_bases_many' '\n> @@ -174,7 +180,7 @@ test_expect_success 'get_merge_bases_many' '\n>  \t\tgit rev-parse commit-5-6 \\\n>  \t\t\t      commit-4-7 | sort\n>  \t} >expect &&\n> -\ttest_three_modes get_merge_bases_many\n> +\ttest_all_modes get_merge_bases_many\n>  '\n>  \n>  test_expect_success 'reduce_heads' '\n> @@ -196,7 +202,7 @@ test_expect_success 'reduce_heads' '\n>  \t\t\t      commit-2-8 \\\n>  \t\t\t      commit-1-10 | sort\n>  \t} >expect &&\n> -\ttest_three_modes reduce_heads\n> +\ttest_all_modes reduce_heads\n>  '\n>  \n>  test_expect_success 'can_all_from_reach:hit' '\n> @@ -219,7 +225,7 @@ test_expect_success 'can_all_from_reach:hit' '\n>  \tY:commit-8-1\n>  \tEOF\n>  \techo \"can_all_from_reach(X,Y):1\" >expect &&\n> -\ttest_three_modes can_all_from_reach\n> +\ttest_all_modes can_all_from_reach\n>  '\n>  \n>  test_expect_success 'can_all_from_reach:miss' '\n> @@ -241,7 +247,7 @@ test_expect_success 'can_all_from_reach:miss' '\n>  \tY:commit-8-5\n>  \tEOF\n>  \techo \"can_all_from_reach(X,Y):0\" >expect &&\n> -\ttest_three_modes can_all_from_reach\n> +\ttest_all_modes can_all_from_reach\n>  '\n>  \n>  test_expect_success 'can_all_from_reach_with_flag: tags case' '\n> @@ -264,7 +270,7 @@ test_expect_success 'can_all_from_reach_with_flag: tags case' '\n>  \tY:commit-8-1\n>  \tEOF\n>  \techo \"can_all_from_reach_with_flag(X,_,_,0,0):1\" >expect &&\n> -\ttest_three_modes can_all_from_reach_with_flag\n> +\ttest_all_modes can_all_from_reach_with_flag\n>  '\n>  \n>  test_expect_success 'commit_contains:hit' '\n> @@ -280,8 +286,8 @@ test_expect_success 'commit_contains:hit' '\n>  \tX:commit-9-3\n>  \tEOF\n>  \techo \"commit_contains(_,A,X,_):1\" >expect &&\n> -\ttest_three_modes commit_contains &&\n> -\ttest_three_modes commit_contains --tag\n> +\ttest_all_modes commit_contains &&\n> +\ttest_all_modes commit_contains --tag\n>  '\n>  \n>  test_expect_success 'commit_contains:miss' '\n> @@ -297,8 +303,8 @@ test_expect_success 'commit_contains:miss' '\n>  \tX:commit-9-3\n>  \tEOF\n>  \techo \"commit_contains(_,A,X,_):0\" >expect &&\n> -\ttest_three_modes commit_contains &&\n> -\ttest_three_modes commit_contains --tag\n> +\ttest_all_modes commit_contains &&\n> +\ttest_all_modes commit_contains --tag\n>  '\n>  \n>  test_expect_success 'rev-list: basic topo-order' '\n> @@ -310,7 +316,7 @@ test_expect_success 'rev-list: basic topo-order' '\n>  \t\tcommit-6-2 commit-5-2 commit-4-2 commit-3-2 commit-2-2 commit-1-2 \\\n>  \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --topo-order commit-6-6\n> +\trun_all_modes git rev-list --topo-order commit-6-6\n>  '\n>  \n>  test_expect_success 'rev-list: first-parent topo-order' '\n> @@ -322,7 +328,7 @@ test_expect_success 'rev-list: first-parent topo-order' '\n>  \t\tcommit-6-2 \\\n>  \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --first-parent --topo-order commit-6-6\n> +\trun_all_modes git rev-list --first-parent --topo-order commit-6-6\n>  '\n>  \n>  test_expect_success 'rev-list: range topo-order' '\n> @@ -334,7 +340,7 @@ test_expect_success 'rev-list: range topo-order' '\n>  \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n>  \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --topo-order commit-3-3..commit-6-6\n> +\trun_all_modes git rev-list --topo-order commit-3-3..commit-6-6\n>  '\n>  \n>  test_expect_success 'rev-list: range topo-order' '\n> @@ -346,7 +352,7 @@ test_expect_success 'rev-list: range topo-order' '\n>  \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n>  \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --topo-order commit-3-8..commit-6-6\n> +\trun_all_modes git rev-list --topo-order commit-3-8..commit-6-6\n>  '\n>  \n>  test_expect_success 'rev-list: first-parent range topo-order' '\n> @@ -358,7 +364,7 @@ test_expect_success 'rev-list: first-parent range topo-order' '\n>  \t\tcommit-6-2 \\\n>  \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n> +\trun_all_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n>  '\n>  \n>  test_expect_success 'rev-list: ancestry-path topo-order' '\n> @@ -368,7 +374,7 @@ test_expect_success 'rev-list: ancestry-path topo-order' '\n>  \t\tcommit-6-4 commit-5-4 commit-4-4 commit-3-4 \\\n>  \t\tcommit-6-3 commit-5-3 commit-4-3 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n> +\trun_all_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n>  '\n>  \n>  test_expect_success 'rev-list: symmetric difference topo-order' '\n> @@ -382,7 +388,7 @@ test_expect_success 'rev-list: symmetric difference topo-order' '\n>  \t\tcommit-3-8 commit-2-8 commit-1-8 \\\n>  \t\tcommit-3-7 commit-2-7 commit-1-7 \\\n>  \t>expect &&\n> -\trun_three_modes git rev-list --topo-order commit-3-8...commit-6-6\n> +\trun_all_modes git rev-list --topo-order commit-3-8...commit-6-6\n>  '\n>  \n>  test_expect_success 'get_reachable_subset:all' '\n> @@ -402,7 +408,7 @@ test_expect_success 'get_reachable_subset:all' '\n>  \t\t\t      commit-1-7 \\\n>  \t\t\t      commit-5-6 | sort\n>  \t) >expect &&\n> -\ttest_three_modes get_reachable_subset\n> +\ttest_all_modes get_reachable_subset\n>  '\n>  \n>  test_expect_success 'get_reachable_subset:some' '\n> @@ -420,7 +426,7 @@ test_expect_success 'get_reachable_subset:some' '\n>  \t\tgit rev-parse commit-3-3 \\\n>  \t\t\t      commit-1-7 | sort\n>  \t) >expect &&\n> -\ttest_three_modes get_reachable_subset\n> +\ttest_all_modes get_reachable_subset\n>  '\n>  \n>  test_expect_success 'get_reachable_subset:none' '\n> @@ -434,7 +440,7 @@ test_expect_success 'get_reachable_subset:none' '\n>  \tY:commit-2-8\n>  \tEOF\n>  \techo \"get_reachable_subset(X,Y)\" >expect &&\n> -\ttest_three_modes get_reachable_subset\n> +\ttest_all_modes get_reachable_subset\n\nAll those are pure renames of test_three_modes() to test_all_modes(),\nwhich now does tests for one more mode -- without GDAT.\n\nAll right.\n\n>  '\n>  \n>  test_done\n\nBest,\n-- \nJakub Narębski\n"},{"id":"408805","messageId":"85zh41sxow.fsf@gmail.com","threadId":"53933","inReplyTo":"8ec119edc66814ad4d63908c79437a7f9dd3c08c.1602079786.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 08/10] commit-graph: use generation v2 only if entire chain does","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-11-01T00:55:11Z","receivedAt":"2020-11-01T00:55:16Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hi Abhishek,\n\n\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> Since there are released versions of Git that understand generation\n> numbers in the commit-graph's CDAT chunk but do not understand the GDAT\n> chunk, the following scenario is possible:\n>\n> 1. \"New\" Git writes a commit-graph with the GDAT chunk.\n> 2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n\nAll right.\n\n>\n> Because of the current use of inspecting the current layer for a\n> chunk_generation_data pointer, the commits in the lower layer will be\n> interpreted as having very large generation values (commit date plus\n> offset) compared to the generation numbers in the top layer (topological\n> level). This violates the expectation that the generation of a parent is\n> strictly smaller than the generation of a child.\n\nI think this paragraphs tries too much to be concise, with the result it\nis less clear than it could be.  Perhaps it would be better to separate\n\"what-if\" from the current behavior.\n\n  If each layer of split commit-graph is treated independently, as it\n  were the case before this commit, with Git inspecting only the current\n  layer for chunk_generation_data pointer, commits in the lower layer\n  (one with GDAT) would have corrected commit date as their generation\n  number, while commits in the upper layer would have topological levels\n  as their generation.  Corrected commit dates have usually much larger\n  values than topological levels.  This means that if we take two\n  commits, one from the upper layer, and one reachable from it in the\n  lower layer, then the expectation that the generation of a parent is\n  smaller than the generation of a child would be violated.\n\n>\n> It is difficult to expose this issue in a test. Since we _start_ with\n> artificially low generation numbers, any commit walk that prioritizes\n> generation numbers will walk all of the commits with high generation\n> number before walking the commits with low generation number. In all the\n> cases I tried, the commit-graph layers themselves \"protect\" any\n> incorrect behavior since none of the commits in the lower layer can\n> reach the commits in the upper layer.\n\nI don't quite understand the issue here. Unless none of the following\nquery commands short-circuit and all walk the commit graph regardless of\nwhat generation numbers tell them, they should give different results\nwith and without the commit graph, if we take two commits one from lower\nlayer of split commit graph with GDAT, and one commit from the higher\nlayer without GDAT, one lower reachable from the other higher.\n\nWe have the following query commands that we can check:\n  $ git merge-base --is-ancestor <lower> <higher>\n  $ git merge-base --independent <lower> <higher>\n  \n  $ git tag --contains <tag-to-lower>\n  $ git tag --merged <tag-to-higher>\n  $ git branch --contains <branch-to-lower>\n  $ git branch --merged <branch-to-higher>\n\nThe second set of queries require for those commits to be tagged, or\nhave branch pointing at them, respectively.\n\nAlso, shouldn't `git commit-graph verify` fail with split commit graph\nwhere the top layer is created with GIT_TEST_COMMIT_GRAPH_NO_GDAT=1?\n\n\nLet's assume that we have the following history, with newer commits\nshown on top like in `git log --graph --oneline --all`:\n\n          topological     corrected         generation\n          level           commit date       number^*\n\n      d    3                                3\n      |\n   c  |    3                                3\n   |  |                                                 without GDAT\n ..|..|.....[layer.boundary]........................................\n   |  |                                                    with GDAT\n   |  b    2              1112912113        1112912113\n   |  |\n   a  |    2              1112912053        1112912053\n   | /\n   |/\n   r       1              1112911993        1112911993\n\n*) each layer inspected individually.\n\nWith such history, we can for example reach 'a' from 'c', thus\n`git merge-base --is-ancestor a b` should return true value, but\nwithout this commit gen(a) > gen(c), instead of gen(a) <= gen(c);\nI use here weaker reachability condition, but the one that works\nalso for commits outside the commit-graph (and those for which\ngeneration numbers overflows).\n\n>\n> This issue would manifest itself as a performance problem in this case,\n> especially with something like \"git log --graph\" since the low\n> generation numbers would cause the in-degree queue to walk all of the\n> commits in the lower layer before allowing the topo-order queue to write\n> anything to output (depending on the size of the upper layer).\n\nAll right, that's good explanation.\n\n>\n> When writing the new layer in split commit-graph, we write a GDAT chunk\n> only if the topmost layer has a GDAT chunk. This guarantees that if a\n> layer has GDAT chunk, all lower layers must have a GDAT chunk as well.\n>\n> Rewriting layers follows similar approach: if the topmost layer below\n> the set of layers being rewritten (in the split commit-graph chain)\n> exists, and it does not contain GDAT chunk, then the result of rewrite\n> does not have GDAT chunks either.\n\nAll right, very good explanation; the only minor suggestion would be to\nadd some 'intro' to the first of those two paragraphs, for example:\n\n  Therefore, when writing the new layer in split commit-graph...\n\nThough I am not sure if it is necessary.\n\n>\n> Signed-off-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c                | 29 +++++++++++-\n>  commit-graph.h                |  1 +\n>  t/t5324-split-commit-graph.sh | 86 +++++++++++++++++++++++++++++++++++\n>  3 files changed, 115 insertions(+), 1 deletion(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 71d0b243db..5d15a1399b 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -605,6 +605,21 @@ static struct commit_graph *load_commit_graph_chain(struct repository *r,\n>  \treturn graph_chain;\n>  }\n>  \n> +static void validate_mixed_generation_chain(struct commit_graph *g)\n> +{\n> +\tint read_generation_data;\n> +\n> +\tif (!g)\n> +\t\treturn;\n> +\n> +\tread_generation_data = !!g->chunk_generation_data;\n> +\n> +\twhile (g) {\n> +\t\tg->read_generation_data = read_generation_data;\n> +\t\tg = g->base_graph;\n> +\t}\n> +}\n\nAll right, this function checks assumedly topmost layer if it is\nGDAT-less, and if it is propagates this status down the layers of split\ncommit graph.  This is needed because if we have mixed-generation commit\ngraph, then for each and every layer we need to use topological levels\nas generation number.\n\nThe only minor issue is the name of this function (it does not hint that\nit propagates the GDAT status downwards), but I don't have a better\nidea, unfortunately.  And it does reflect what this function is used for.\n\n> +\n>  struct commit_graph *read_commit_graph_one(struct repository *r,\n>  \t\t\t\t\t   struct object_directory *odb)\n>  {\n> @@ -613,6 +628,8 @@ struct commit_graph *read_commit_graph_one(struct repository *r,\n>  \tif (!g)\n>  \t\tg = load_commit_graph_chain(r, odb);\n>  \n> +\tvalidate_mixed_generation_chain(g);\n> +\n\nAll right, this looks like a good place to put this new check: just\nafter reading commit-graph chain.\n\n>  \treturn g;\n>  }\n>  \n> @@ -782,7 +799,7 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \tdate_low = get_be32(commit_data + g->hash_len + 12);\n>  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n>  \n> -\tif (g->chunk_generation_data) {\n> +\tif (g->chunk_generation_data && g->read_generation_data) {\n>  \t\toffset = (timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n\nAll right, instead of simply checking if the current layer has\ngeneration data chunk, we need to also check if the whole graph allows\nfor it (if there are no mixed-generation layers).\n\nThe g->read_generation_data should be filled correctly, because\nfill_commit_graph_info() is always preceded by read_commit_graph(), if I\nunderstand it correctly.\n\n>  \n>  \t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n> @@ -2030,6 +2047,9 @@ static void split_graph_merge_strategy(struct write_commit_graph_context *ctx)\n>  \t\t}\n>  \t}\n>  \n> +\tif (!ctx->write_generation_data && g->chunk_generation_data)\n> +\t\tctx->write_generation_data = 1;\n> +\n\nThis needs more careful examination, and looking at larger context of\nthose lines.\n\nAt this point, unless `--split=replace` option is used, 'g' points to\nthe bottom layer out of all topmost layers being merged. We know that if\nthere are GDAT-less layers then these must be top layers, so this means\nthat we can write GDAT chunk in the result of the merge -- because we\nwould be replacing all possible GDAT-less layers (and maybe some with\nGDAT) with a single layer with the GDAT chunk.\n\nThe ctx->write_generation_data is set to true unless environment\nvariable GIT_TEST_COMMIT_GRAPH_NO_GDAT is true, and that in\nwrite_commit_graph() it would be set to false if topmost layer doesn't\nhave GDAT chunk, and to true if `--split=replace` option is used; see\nbelow.\n\nLooks good to me.\n\n\nNOTE that this means that GIT_TEST_COMMIT_GRAPH_NO_GDAT prevents from\nwriting GDAT chunk with generation data v2 unless we are merging layers,\nor replacing all of them with a single layer: then it is _ignored_.\n\nShould we clarify this fact in the description of GIT_TEST_COMMIT_GRAPH_NO_GDAT\nin t/README?  Currently it reads:\n\n  GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n  commit-graph to be written without generation data chunk.\n\n>  \tif (flags != COMMIT_GRAPH_SPLIT_REPLACE)\n>  \t\tctx->new_base_graph = g;\n>  \telse if (ctx->num_commit_graphs_after != 1)\n> @@ -2274,6 +2294,7 @@ int write_commit_graph(struct object_directory *odb,\n>  \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n>  \n>  \t\twhile (g) {\n> +\t\t\tg->read_generation_data = 1;\n>  \t\t\tg->topo_levels = &topo_levels;\n>  \t\t\tg = g->base_graph;\n>  \t\t}\n\nAll right, when writing the commit graph we want to make use of existing\ngeneration data chunks.  This is safe, because when computing generation\nnumbers for writing we have separate place to store topoogical levels\n(`topo_levels`) so they would not be mixed with corrected commit dates:\ngeneration number v1 and v2 are kept separate.\n\n> @@ -2300,6 +2321,9 @@ int write_commit_graph(struct object_directory *odb,\n>  \n>  \t\tg = ctx->r->objects->commit_graph;\n>  \n> +\t\tif (g && !g->chunk_generation_data)\n> +\t\t\tctx->write_generation_data = 0;\n> +\n\nAll right, if current (topmost) layed does not have GDAT, then when\ncreating a new layer do not create GDAT layer either (merging layers and\nrewriting the commit-graph is handled separately).\n\n>  \t\twhile (g) {\n>  \t\t\tctx->num_commit_graphs_before++;\n>  \t\t\tg = g->base_graph;\n> @@ -2318,6 +2342,9 @@ int write_commit_graph(struct object_directory *odb,\n>  \n>  \t\tif (ctx->opts)\n>  \t\t\treplace = ctx->opts->split_flags & COMMIT_GRAPH_SPLIT_REPLACE;\n> +\n> +\t\tif (replace)\n> +\t\t\tctx->write_generation_data = 1;\n\nAll right, when replacing all layers (`git commit-graph write --split=replace`),\nthen we can safely write the GDAT chunk.\n\nNote however that here we don't take into account the value of the\nenvironment variable GIT_TEST_COMMIT_GRAPH_NO_GDAT.  Which maybe is what\nwe want...\n\n>  \t}\n>  \n>  \tctx->approx_nr_objects = approximate_object_count();\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 19a02001fd..ad52130883 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -64,6 +64,7 @@ struct commit_graph {\n>  \tstruct object_directory *odb;\n>  \n>  \tuint32_t num_commits_in_base;\n> +\tunsigned int read_generation_data;\n>  \tstruct commit_graph *base_graph;\n\nAll right, this new field is here to propagate to each layer the\ninformation whether we can read from the generation number v2 data\nchunk.\n\nThough I am not sure whether this field should be added here, and\nwhether it should be `unsigned int` (we don't have to be that careful\nabout saving space for this type).\n\n>  \n>  \tconst uint32_t *chunk_oid_fanout;\n> diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> index 651df89ab2..d0949a9eb8 100755\n> --- a/t/t5324-split-commit-graph.sh\n> +++ b/t/t5324-split-commit-graph.sh\n> @@ -440,4 +440,90 @@ test_expect_success '--split=replace with partial Bloom data' '\n>  \tverify_chain_files_exist $graphdir\n>  '\n>  \n> +test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n> +\tmkdir mixed &&\n\nThis should probably go just before cd-ing into just created\nsubdirectory.\n\n> +\tgraphdir=\".git/objects/info/commit-graphs\" &&\n> +\ttest_oid_cache <<-EOM &&\n> +\toid_version sha1:1\n> +\toid_version sha256:2\n> +\tEOM\n\nMinor nitpick: Why use \"EOM\", which is used only twice in Git the test\nsuite, and not the conventional \"EOF\" (used at least 4000 times)?\n\n> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n\nThe t/README says:\n\n   - Don't chdir around in tests.  It is not sufficient to chdir to\n     somewhere and then chdir back to the original location later in\n     the test, as any intermediate step can fail and abort the test,\n     causing the next test to start in an unexpected directory.  Do so\n     inside a subshell if necessary.\n\nThough I am not sure if it should apply also to this situation.\n\n> +\tgit init &&\n> +\tgit config core.commitGraph true &&\n> +\tgit config gc.writeCommitGraph false &&\n\nAll right.\n\n> +\tfor i in $(test_seq 3)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit reset --hard commits/1 &&\n> +\tfor i in $(test_seq 4 5)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit reset --hard commits/2 &&\n> +\tfor i in $(test_seq 6 10)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit commit-graph write --reachable --split &&\n\nIs there a reason why we do not check just written commit-graph file\nwith `test-tool read-graph >output-layer-1`?\n\n> +\tgit reset --hard commits/2 &&\n> +\tgit merge commits/4 &&\n\nShouldn't we use `test_merge` instead of `git merge`; I am not sure when\nto use one or the other?\n\n> +\tgit branch merge/1 &&\n> +\tgit reset --hard commits/4 &&\n> +\tgit merge commits/6 &&\n> +\tgit branch merge/2 &&\n\nIt would be nice to have ASCII-art of the history (of the graph of\nrevisions) created here for subsequent tests:\n\n                                        \n           /- 6 <-- 7 <-- 8 <-- 9 <-- 10*\n          /    \\-\\\n         /        \\\n  1 <-- 2 <-- 3*   \\--\\\n  |      \\             \\ \n  |       \\-----\\       \\\n   \\             \\       \\\n    \\-- 4*<------ M/1     M/2\n        |\\               /  \n        | \\-- 5*        /\n        \\              /\n         \\------------/\n\n  * - 1st layer  \n\nThough I am not sure if what I have created is readable; I think a\nbetter way to draw this graph is possible, for example:\n\n               /- 3*\n              /\n             /\n  1 <------ 2 <---- 6 <-- 7 <-- 8 <-- 9 <-- 10*\n   \\         \\       \\\n    \\         \\       \\\n     \\         \\       \\\n      \\- 4* <-- M/1     \\    \n         |\\              \\\n         | \\------------- M/2\n         \\\n          \\---- 5*\n\nEdit: as I see the history gets even more complicated, so perhaps\nASCII-art diagram of the history with layers marked would be too\ncomplicated, and wouldn't bring much.\n\nWhy do we need such shape of the history in the repository?\n\n> +\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n> +\ttest-tool read-graph >output &&\n> +\tcat >expect <<-EOF &&\n> +\theader: 43475048 1 $(test_oid oid_version) 4 1\n> +\tnum_commits: 2\n> +\tchunks: oid_fanout oid_lookup commit_metadata\n> +\tEOF\n> +\ttest_cmp expect output &&\n\nAll right, we check that we have 2 commits, and that there is no GDAT\nchunk.\n\n> +\tgit commit-graph verify\n\nAll right, we verify commit-graph as a whole (both layers).\n\n> +'\n> +\n> +test_expect_success 'does not write generation data chunk if not present on existing tip' '\n\nHmmm... I wonder if we can come up with a better name for this test;\nfor example should it be \"does not write\" or \"do not write\"?\n\n> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> +\tgit reset --hard commits/3 &&\n> +\tgit merge merge/1 &&\n> +\tgit merge commits/5 &&\n> +\tgit merge merge/2 &&\n> +\tgit branch merge/3 &&\n\nThe commit graph gets complicated, so it would not be easy to visualize\nit with ASCII-art diagram without any crossed lines.  Maybe `git log\n--graph --oneline --all` would help:\n\n*   (merge/3) Merge branch 'merge/2'\n|\\\n| *   (merge/2) Merge branch 'commits/6'\n| |\\\n* | \\   Merge branch 'commits/5'\n|\\ \\ \\\n| * | | (commits/5) 5\n| |/ /\n* | |   Merge branch 'merge/1'\n|\\ \\ \\\n| * | | (merge/1) Merge branch 'commits/4'\n| |\\| |\n| | * | (commits/4) 4\n* | | | (commits/3) 3\n|/ / /\n| | | * (commits/10) 10\n| | | * (commits/9) 9\n| | | * (commits/8) 8\n| | | * (commits/7) 7\n| | |/\n| | * (commits/6) 6\n| |/\n|/|\n* | (commits/2) 2\n|/\n* (commits/1) 1\n\n\n> +\tgit commit-graph write --reachable --split=no-merge &&\n> +\ttest-tool read-graph >output &&\n> +\tcat >expect <<-EOF &&\n> +\theader: 43475048 1 $(test_oid oid_version) 4 2\n> +\tnum_commits: 3\n> +\tchunks: oid_fanout oid_lookup commit_metadata\n> +\tEOF\n> +\ttest_cmp expect output &&\n> +\tgit commit-graph verify\n\nAll right, so here we check that we have layer without GDAT at the top,\nand we request not to merge layers thus new layer will be created, then\nthe new layer also does not have GDAT chunk (and has 3 commits).\n\nMinor nitpick: shouldn't those test be indented?\n\n> +'\n> +\n> +test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> +\tgit commit-graph write --reachable --split=replace &&\n> +\ttest_path_is_file $graphdir/commit-graph-chain &&\n> +\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n> +\tverify_chain_files_exist $graphdir &&\n\nAll right, this checks that we have split commit-graph chain that\nconsist of a single layer, and that the commit-graph file for this\nsingle layer exists.\n\n> +\tgraph_read_expect 15 &&\n\nShouldn't we use `test-tool read-graph` to check whether generation_data\nchunk is present... ah, sorry, I have realized that after previous\npatches `graph_read_expect 15` implicitly checks the latter, because in\nits' use of `test-tool read-graph` it does expect generation_data chunk.\n\nSo we use `test-tool read-graph` manually to check that generation_data\nchunk is absent, and we use graph_read_expect to check that it is\npresent (and in both cases that the number of commits matches).  I\nwonder if it would be possible to simplify that...\n\n> +\tgit commit-graph verify\n\nAll right.\n\n> +'\n> +\n> +test_expect_success 'add one commit, write a tip graph' '\n> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> +\ttest_commit 11 &&\n> +\tgit branch commits/11 &&\n> +\tgit commit-graph write --reachable --split &&\n> +\ttest_path_is_missing $infodir/commit-graph &&\n> +\ttest_path_is_file $graphdir/commit-graph-chain &&\n> +\tls $graphdir/graph-*.graph >graph-files &&\n> +\ttest_line_count = 2 graph-files &&\n> +\tverify_chain_files_exist $graphdir\n> +'\n\nWhat it is meant to test?  That adding single-commit to a 15 commit\ncommit-graph file in split mode does not result in layers merging, and\nactually adds a new layer: we check that we have exactly two layers and\nthat they are all OK.\n\nWe don't check here that the newly created top layer commit-graph does\nhave GDAT chunk, as it should be if the top layer (in this case the only\nlayer) has GDAT chunk.\n\n> +\n>  test_done\n\nOne test we are missing is testing that merging layers is done\ncorrectly, namely that if we are merging layers in split commit-graph\nfile, and the layer below the ones we are merging lacks GDAT chunk, then\nthe result of the merge should also be without GDAT chunk.  This would\nrequire at least two GDAT-less layers in a setup.\n\nI'm not sure how difficult writing such test should be.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"408983","messageId":"20201103053629.GA13228@Abhishek-Arch","threadId":"53933","inReplyTo":"X5Xm5nlFK39o0rkJ@nand.local","subject":"Re: [PATCH v4 01/10] commit-graph: fix regression when computing Bloom filters","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-11-03T05:36:29Z","receivedAt":"2020-11-03T05:39:22Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Sun, Oct 25, 2020 at 04:58:14PM -0400, Taylor Blau wrote:\n> On Sun, Oct 25, 2020 at 01:16:48AM +0200, Jakub Narębski wrote:\n> > \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> >\n> > > While measuring performance with `git commit-graph write --reachable\n> > > --changed-paths` on the linux repository led to around 1m40s for both\n> > > HEAD and master (and could be due to fault in my measurements), it is\n> > > still the \"right\" thing to do.\n> >\n> > I had to read the above paragraph several times to understand it,\n> > possibly because I have expected here to be a fix for a performance\n> > regression.  The commit message for 3d112755 (commit-graph: examine\n> > commits by generation number) describes reduction of computation time\n> > from 3m00s to 1m37s.  So I would expect performance with HEAD (i.e.\n> > before those changes) to be around 3m, not the same before and after\n> > changes being around 1m40s.\n> >\n> > Can anyone recheck this before-and-after benchmark, please?\n> \n> My hunch is that our heuristic to fall back to the commits 'date'\n> value is saving us here. commit_gen_cmp() first compares the generation\n> numbers, breaking ties by 'date' as a heuristic. But since all\n> generation number queries return GENERATION_NUMBER_INFINITY during\n> writing, we're relying on our heuristic entirely.\n> \n> I haven't looked much further than that, other than to see that I could\n> get about a ~4sec speed-up with this patch as compared to v2.29.1 in the\n> computing Bloom filters region on the kernel.\n> \n\nThanks for benchmarking it. I wasn't sure if I am testing it correctly\nor the patch made no difference.\n\n> > Anyway, it might be more clear to write it as the following:\n> >\n> >   On the Linux kernel repository, this patch didn't reduce the\n> >   computation time for 'git commit-graph write --reachable\n> >   --changed-paths', which is around 1m40s both before and after this\n> >   change.  This could be a fault in my measurements; it is still the\n> >   \"right\" thing to do.\n> >\n> > Or something like that.\n> \n> Assuming that we are in fact being saved by the \"date\" heuristic, I'd\n> probably write the following commit message instead:\n> \n>   Before computing Bloom filters, the commit-graph machinery uses\n>   commit_gen_cmp to sort commits by generation order for improved diff\n>   performance. 3d11275505 (commit-graph: examine commits by generation\n>   number, 2020-03-30) claims that this sort can reduce the time spent to\n>   compute Bloom filters by nearly half.\n> \n>   But since c49c82aa4c (commit: move members graph_pos, generation to a\n>   slab, 2020-06-17), this optimization is broken, since asking for\n>   'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\n>   while writing.\n> \n>   Not all hope is lost, though: 'commit_graph_generation()' falls\n>   back to comparing commits by their date when they have equal generation\n>   number, and so since c49c82aa4c is purely a date comparison function.\n>   This heuristic is good enough that we don't seem to loose appreciable\n>   performance while computing Bloom filters. [Benchmark that we loose\n>   about ~4sec before/after c49c82aa4c9...]\n> \n>   So, avoid the uesless 'commit_graph_generation()' while writing by\n>   instead accessing the slab directly. This returns the newly-computed\n>   generation numbers, and allows us to avoid the heuristic by directly\n>   comparing generation numbers.\n> \n\nThat's a lot better, will change.\n\n> Thanks,\n> Taylor\n"},{"id":"408984","messageId":"20201103064058.GB13228@Abhishek-Arch","threadId":"53933","inReplyTo":"85imay2z84.fsf@gmail.com","subject":"Re: [PATCH v4 04/10] commit-graph: return 64-bit generation number","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-11-03T06:40:58Z","receivedAt":"2020-11-03T06:43:46Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Sun, Oct 25, 2020 at 02:48:27PM +0100, Jakub Narębski wrote:\n> Hi Abhishek,\n> \n> Note that there are two changes that are not mentioned in the commit\n> message, namely adding 'const'-ness to generation_a/b local variables in\n> commit_gen_cmp() from commit-graph.c, and switching from\n> GENERATION_NUMBER_ZERO to GENERATION_NUMBER_INFINITY as the default\n> (initial) value for 'max_generation' in repo_in_merge_bases_many().\n> \n> While the former is a simple \"while-at-it\" change that shouldn't affect\n> correctness, the latter needs an explanation (or fixing if it is wrong).\n> \n\nThe change from GENERATION_NUMBER_ZERO to GENERATION_NUMBER_INFINITY was\nincorrect. While fixing merge conflicts on rebasing to master again, I\ndidn't notice that repo_in_merge_bases_many() switched from using\nmin_generation and GENERATION_NUMBER_ZERO to max_generation and\nGENERATION_NUMBER_ZERO.\n\nThanks for noticing!\n\n> ...\n> \n> Best,\n> -- \n> Jakub Narębski\n"},{"id":"408991","messageId":"20201103114432.GA3577@Abhishek-Arch","threadId":"53933","inReplyTo":"85r1pjzejg.fsf@gmail.com","subject":"Re: [PATCH v4 06/10] commit-graph: implement corrected commit date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-11-03T11:44:32Z","receivedAt":"2020-11-03T11:47:26Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Oct 27, 2020 at 07:53:23PM +0100, Jakub Narębski wrote:\n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ...\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> Somewhere in the commit message we should also describe that this commit\n> changes how commit-graph is verified: from checking that the generation\n> number agrees with _topological level definition_, that is that for a\n> given commit it is 1 more than maximum of its parents (with the caveat\n> that we need to handle GENERATION_NUMBER_V1_MAX values correctly), to\n> checking that slightly weaker condition fulfilled by both topological\n> levels (generation number v1) and by corrected commit date (generation\n> number v2) that for a given commit its generation number is 1 more than\n> maximum of its parents or larger.\n\nSure, that makes sense. Will add.\n\n> \n> But, as far as I understand it, current code does not handle correctly\n> GENERATION_NUMBER_V1_MAX case (if we use generation number v1).\n> \n> On the other hand we could have simpy use functional check, that\n> generation number used (which can be v1 or v2, or any similar other)\n> fulfills the reachability condition for each edge, which can be\n> simplified to checking that generation(parents) <= generation(commit).\n> If the reachability condition is true for each edge, then it is true for\n> each path, and for each commit.\n> \n> > ---\n> >  commit-graph.c | 43 +++++++++++++++++++++++--------------------\n> >  1 file changed, 23 insertions(+), 20 deletions(-)\n> >\n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index cedd311024..03948adfce 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -154,11 +154,6 @@ static int commit_gen_cmp(const void *va, const void *vb)\n> >  \telse if (generation_a > generation_b)\n> >  \t\treturn 1;\n> >  \n> > -\t/* use date as a heuristic when generations are equal */\n> > -\tif (a->date < b->date)\n> > -\t\treturn -1;\n> > -\telse if (a->date > b->date)\n> > -\t\treturn 1;\n> \n> Why this change?  It is not described in the commit message.\n> \n> Note that while this tie-breaking fallback doesn't make much sense for\n> corrected committer date generation number v2, this tie-breaking helps\n> if we have to use topological levels (generation number v2).\n> \n\nRight, I should have mentioned this change (and it's not something that\nmakes a difference either way).\n\nWe call commit_gen_cmp() only when we are sorting commits by generation\nto speed up computation of Bloom filters i.e. while writing a commit\ngraph (either split commit-graph or a simple commit-graph).\n\nSince we are always computing and storing corrected commit date when we\nare writing (whether we write a GDAT chunk or not), using date as\nheuristic is longer required.\n\n> >  \treturn 0;\n> >  }\n> >  \n> > @@ -1357,10 +1352,14 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\t\t\tctx->commits.nr);\n> >  \tfor (i = 0; i < ctx->commits.nr; i++) {\n> >  \t\ttimestamp_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n> \n> Sidenote: I haven't noticed it earlier, but here 'uint32_t' might be\n> enough; no need for 'timestamp_t' for 'level' variable.\n> \n> > +\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n> >\n\nWe need the 'timestamp_t' as we are comparing level with the now 64-bits\nGENERATION_NUMBER_INFINITY. I thought uint32_t would be promoted to\ntimestamp_t. I have a hunch that since we are explicitly using a fixed\nwidth data type, compiler is unwilling to type coerce into broader data\ntypes.\n\nAdvice on this appreciated.\n\n> \n> All right, we compute both generation numbers: topological levels and\n> corrected commit date.\n> \n> I guess we use 'corrected_commit_date' instead of simply 'generation' to\n> make it asier to remember which is which.\n> \n> >  \t\tdisplay_progress(ctx->progress, i + 1);\n> >  \t\tif (level != GENERATION_NUMBER_INFINITY &&\n> > -\t\t    level != GENERATION_NUMBER_ZERO)\n> > +\t\t    level != GENERATION_NUMBER_ZERO &&\n> > +\t\t    corrected_commit_date != GENERATION_NUMBER_INFINITY &&\n> > +\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO\n> \n> Straightforward addition.\n> \n> > +\t\t    )\n> \n> Why this closing parenthesis is now in separated line?\n> \n> >  \t\t\tcontinue;\n> >  \n> >  \t\tcommit_list_insert(ctx->commits.list[i], &list);\n> > @@ -1369,17 +1368,25 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\tstruct commit_list *parent;\n> >  \t\t\tint all_parents_computed = 1;\n> >  \t\t\tuint32_t max_level = 0;\n> > +\t\t\ttimestamp_t max_corrected_commit_date = 0;\n> \n> All right, straightforward addition.\n> \n> >  \n> >  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n> >  \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n> > -\n> \n> Why we have removed this empty line?\n> \n> > +\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n> \n> All right.\n> \n> >  \t\t\t\tif (level == GENERATION_NUMBER_INFINITY ||\n> > -\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n> > +\t\t\t\t    level == GENERATION_NUMBER_ZERO ||\n> > +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_INFINITY ||\n> > +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO\n> > +\t\t\t\t    ) {\n> \n> All right, same as above.\n> \n> >  \t\t\t\t\tall_parents_computed = 0;\n> >  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n> >  \t\t\t\t\tbreak;\n> > -\t\t\t\t} else if (level > max_level) {\n> > -\t\t\t\t\tmax_level = level;\n> > +\t\t\t\t} else {\n> > +\t\t\t\t\tif (level > max_level)\n> > +\t\t\t\t\t\tmax_level = level;\n> > +\n> > +\t\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n> > +\t\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n> >  \t\t\t\t}\n> \n> All right, reasonable and straightforward.\n> \n> >  \t\t\t}\n> >  \n> > @@ -1389,6 +1396,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n> >  \t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n> >  \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n> > +\n> > +\t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n> > +\t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n> > +\t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n> \n> All right.\n> \n> Here we use the same trick as in previous commit (and as above) to avoid\n> any possible overflow, to minimize number of conditionals.  The fact\n> that max_corrected_commit_date might store incorrect value doesn't\n> matter, as it is reset at beginning of this loop.\n> \n> >  \t\t\t}\n> >  \t\t}\n> >  \t}\n> > @@ -2485,17 +2496,9 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n> >  \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n> >  \t\t\tcontinue;\n> >  \n> > -\t\t/*\n> > -\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n> > -\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n> > -\t\t * extra logic in the following condition.\n> > -\t\t */\n> > -\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n> > -\t\t\tmax_generation--;\n> > -\n> \n> Perhaps in the future we should check that both topological levels, and\n> also corrected committer date (if it exists) for correctness according\n> to their definition.  Then the above removed part would be restored (but\n> with s/max_generation/max_level/).\n> \n> >  \t\tgeneration = commit_graph_generation(graph_commit);\n> > -\t\tif (generation != max_generation + 1)\n> > -\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n> > +\t\tif (generation < max_generation + 1)\n> > +\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n> \n> All right, so we relaxed the check so that it will be fulfilled by\n> generation number v2 (and also by generation number v1, as it implies\n> the more strict check for v1).\n> \n> What would happen however if generation holds topological levels, and it\n> is GENERATION_NUMBER_V1_MAX for at least one parent, which means it is\n> GENERATION_NUMBER_V1_MAX for a commit?  As you can check, the condition\n> would be true: GENERATION_NUMBER_V1_MAX < GENERATION_NUMBER_V1_MAX + 1,\n> so the `git commit-graph verify` would incorrectly say that there is\n> a problem with generation number, while there isn't one (false positive\n> detection of error).\n\nAlright, so the above block still makes sense if we are working with\ntopological levels but not with corrected commit dates. Instead of\nremoving it, I will modify the condition to check that one of our parents\nhas GENERATION_NUMBER_V1_MAX and the graph uses topological levels.\n\nSuprised that no test breaks by this change.\n\nI have also moved changes in the verify function to the next patch, as\nwe cannot write or read corrected commit dates yet - so little sense in\nmodifying verify.\n\n> \n> Sidenote: I think we don't have to worry about having to introduce\n> GENERATION_NUMBER_V2_MAX, as the in-memory size (of reconstructed from\n> disck representation) corrected commiter date is the same as of commiter\n> date itself, plus some, and I don't see us coming close to 64-bit limit\n> of timestamp_t for commit dates.\n> \n> >  \t\t\t\t     oid_to_hex(&cur_oid),\n> >  \t\t\t\t     generation,\n> >  \t\t\t\t     max_generation + 1);\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"409025","messageId":"85y2jiqq3c.fsf@gmail.com","threadId":"53933","inReplyTo":"bb9b02af32d028fc0c26d372aa490e260c74e74d.1602079786.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 09/10] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-11-03T17:59:03Z","receivedAt":"2020-11-03T17:59:43Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> With corrected commit dates implemented, we no longer have to rely on\n> commit date as a heuristic in paint_down_to_common().\n>\n> While using corrected commit dates Git walks nearly the same number of\n> commits as commit date, the process is slower as for each comparision we\n> have to access a commit-slab (for corrected committer date) instead of\n> accessing struct member (for committer date).\n\nSomething for the future: I wonder if it would be worth it to bring back\ngeneration number from the commit-slab into `struct commit`.\n\n>\n> For example, the command `git merge-base v4.8 v4.9` on the linux\n> repository walks 167468 commits, taking 0.135s for committer date and\n> 167496 commits, taking 0.157s for corrected committer date respectively.\n\nI think it would be good idea to explicitly refer to the commit that\nchanged paint_down_to_common() to *not* use generation numbers v1\n(topological levels) in the cases such as this, namely 091f4cf3 (commit:\ndon't use generation numbers if not needed).  In this commit we have the\nfollowing:\n\n  This change makes a concrete difference depending on the topology\n  of the commit graph. For instance, computing the merge-base between\n  consecutive versions of the Linux kernel has no effect for versions\n  after v4.9, but 'git merge-base v4.8 v4.9' presents a performance\n  regression:\n\n      v2.18.0: 0.122s\n  v2.19.0-rc1: 0.547s\n         HEAD: 0.127s\n\n  To determine that this was simply an ordering issue, I inserted\n  a counter within the while loop of paint_down_to_common() and\n  found that the loop runs 167,468 times in v2.18.0 and 635,579\n  times in v2.19.0-rc1.\n\nThe times you report (0.135s and 0.157s) are close to 0.122s / 0.127s\nreported in 091f4cf3 - that is most probably because of the differences\nin the system performance (hardware, operating system, load, etc.).\nNumbers of commits walked for the committed date heuristics, that is\n167,468 agrees with your results; 167,496 (+28) for corrected commit\ndate (generation number v2) is significantly smaller (-468,083) than\n635,579 reported for topological levels (generation number v1).\n\nI suspect that there are cases (with date skew) where corrected commit\ndate gives better performance than committer date heuristics, and I am\nquite sure that generation number v2 can give better performance in case\nwhere paint_down_to_common() uses generation numbers.\n\n.................................................................\n\nHere begins separate second change, which is not put into separate\ncommit because it is fairly tightly connected to the change described\nabove.  It would be good idea, in my opinion, to add a sentence that\nexplicitely marks this switch, for example:\n\n  This change accidentally broke fragile t6404-recursive-merge test.\n  t6404-recursive-merge setups a unique repository...\n\nMaybe with s/accidentaly/incidentally/.\n\nOr add some other way of connection those two parts of the commit\nmessages.\n\n> t6404-recursive-merge setups a unique repository where all commits have\n> the same committer date without well-defined merge-base.\n>\n> While running tests with GIT_TEST_COMMIT_GRAPH unset, we use committer\n> date as a heuristic in paint_down_to_common(). 6404.1 'combined merge\n> conflicts' merges commits in the order:\n> - Merge C with B to form a intermediate commit.\n> - Merge the intermediate commit with A.\n>\n> With GIT_TEST_COMMIT_GRAPH=1, we write a commit-graph and subsequently\n> use the corrected committer date, which changes the order in which\n> commits are merged:\n> - Merge A with B to form a intermediate commit.\n> - Merge the intermediate commit with C.\n>\n> While resulting repositories are equivalent, 6404.4 'virtual trees were\n> processed' fails with GIT_TEST_COMMIT_GRAPH=1 as we are selecting\n> different merge-bases and thus have different object ids for the\n> intermediate commits.\n>\n> As this has already causes problems (as noted in 859fdc0 (commit-graph:\n> define GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph\n> within t6404-recursive-merge.\n\nVery nice explanation.\n\nPerhaps in the future we could make this test less fragile.\n\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c             | 14 ++++++++++++++\n>  commit-graph.h             |  8 +++++++-\n>  commit-reach.c             |  2 +-\n>  t/t6404-recursive-merge.sh |  5 ++++-\n>  4 files changed, 26 insertions(+), 3 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 5d15a1399b..3de1933ede 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -705,6 +705,20 @@ int generation_numbers_enabled(struct repository *r)\n>  \treturn !!first_generation;\n>  }\n>  \n> +int corrected_commit_dates_enabled(struct repository *r)\n> +{\n> +\tstruct commit_graph *g;\n> +\tif (!prepare_commit_graph(r))\n> +\t\treturn 0;\n> +\n> +\tg = r->objects->commit_graph;\n> +\n> +\tif (!g->num_commits)\n> +\t\treturn 0;\n> +\n> +\treturn g->read_generation_data;\n> +}\n\nVery nice abstraction.\n\nMinor issue: I wonder if it would be better to use _available() or\n\"_present()\" rather than _enabled() suffix.\n\n> +\n>  struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r)\n>  {\n>  \tstruct commit_graph *g = r->objects->commit_graph;\n> diff --git a/commit-graph.h b/commit-graph.h\n> index ad52130883..d2c048dc64 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -89,13 +89,19 @@ struct commit_graph *read_commit_graph_one(struct repository *r,\n>  struct commit_graph *parse_commit_graph(struct repository *r,\n>  \t\t\t\t\tvoid *graph_map, size_t graph_size);\n>  \n> +struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n> +\n>  /*\n>   * Return 1 if and only if the repository has a commit-graph\n>   * file and generation numbers are computed in that file.\n>   */\n>  int generation_numbers_enabled(struct repository *r);\n>  \n> -struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n\nThis moving get_bloom_filter_settings() before generation_numbers_enabled() \nlooks like accidental change.  If not, why it is here?\n\n> +/*\n> + * Return 1 if and only if the repository has a commit-graph\n> + * file and generation data chunk has been written for the file.\n> + */\n> +int corrected_commit_dates_enabled(struct repository *r);\n>\n\nAll right, nice to have documentation for the public function.\n\n>  enum commit_graph_write_flags {\n>  \tCOMMIT_GRAPH_WRITE_APPEND     = (1 << 0),\n> diff --git a/commit-reach.c b/commit-reach.c\n> index 20b48b872b..46f5a9e638 100644\n> --- a/commit-reach.c\n> +++ b/commit-reach.c\n> @@ -39,7 +39,7 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n>  \tint i;\n>  \ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n>  \n> -\tif (!min_generation)\n> +\tif (!min_generation && !corrected_commit_dates_enabled(r))\n>  \t\tqueue.compare = compare_commits_by_commit_date;\n>  \n>  \tone->object.flags |= PARENT1;\n\nAll right, this is the meat of the first change.\n\n> diff --git a/t/t6404-recursive-merge.sh b/t/t6404-recursive-merge.sh\n> index 332cfc53fd..7055771b62 100755\n> --- a/t/t6404-recursive-merge.sh\n> +++ b/t/t6404-recursive-merge.sh\n> @@ -15,6 +15,8 @@ GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n>  export GIT_COMMITTER_DATE\n>  \n>  test_expect_success 'setup tests' '\n> +\tGIT_TEST_COMMIT_GRAPH=0 &&\n> +\texport GIT_TEST_COMMIT_GRAPH &&\n>  \techo 1 >a1 &&\n>  \tgit add a1 &&\n>  \tGIT_AUTHOR_DATE=\"2006-12-12 23:00:00\" git commit -m 1 a1 &&\n\nAll right, we turn off running this test with commit-graph for the whole\nscript, not only for a single test.  As this is a setup, it would be run\neven if we are skipping some tests.\n\n> @@ -66,7 +68,7 @@ test_expect_success 'setup tests' '\n>  '\n>  \n>  test_expect_success 'combined merge conflicts' '\n> -\ttest_must_fail env GIT_TEST_COMMIT_GRAPH=0 git merge -m final G\n> +\ttest_must_fail git merge -m final G\n>  '\n\nAll right, it is no longer necessary to run this specific test with\nGIT_TEST_COMMIT_GRAPH=0 as now the whole script is run with this\nsetting.\n\n>  \n>  test_expect_success 'result contains a conflict' '\n> @@ -82,6 +84,7 @@ test_expect_success 'result contains a conflict' '\n>  '\n>  \n>  test_expect_success 'virtual trees were processed' '\n> +\t# TODO: fragile test, relies on ambigious merge-base resolution\n>  \tgit ls-files --stage >out &&\n>  \n>  \tcat >expect <<-EOF &&\n\nGood call!  Nice adding TODO comment for the future.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"409026","messageId":"xmqq7dr2s3oy.fsf@gitster.c.googlers.com","threadId":"53933","inReplyTo":"85y2jiqq3c.fsf@gmail.com","subject":"Re: [PATCH v4 09/10] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-11-03T18:19:57Z","receivedAt":"2020-11-03T18:20:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"jnareb@gmail.com (Jakub Narębski) writes:\n\n> I suspect that there are cases (with date skew) where corrected commit\n> date gives better performance than committer date heuristics, and I am\n> quite sure that generation number v2 can give better performance in case\n> where paint_down_to_common() uses generation numbers.\n\nThanks for a well reasoned review.\n\n>\n> .................................................................\n>\n> Here begins separate second change, which is not put into separate\n> commit because it is fairly tightly connected to the change described\n> above.  It would be good idea, in my opinion, to add a sentence that\n> explicitely marks this switch, for example:\n>\n>   This change accidentally broke fragile t6404-recursive-merge test.\n>   t6404-recursive-merge setups a unique repository...\n>\n> Maybe with s/accidentaly/incidentally/.\n\nAlso \"setup\" is not a verb.  \"... sets up a unique repository\".\n\n> Or add some other way of connection those two parts of the commit\n> messages.\n> ...\n>> As this has already causes problems (as noted in 859fdc0 (commit-graph:\n>> define GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph\n>> within t6404-recursive-merge.\n>\n> Very nice explanation.\n>\n> Perhaps in the future we could make this test less fragile.\n\nIf \"separate second change\" is distracting, would it be an option to\nfix the test before this step, perhaps?\n\nThanks.\n"},{"id":"409060","messageId":"85tuu5q4uy.fsf@gmail.com","threadId":"53933","inReplyTo":"9ada43967d29a3ec717b6a8db0de5b09e6d916b1.1602079786.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 10/10] doc: add corrected commit date info","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-11-04T01:37:41Z","receivedAt":"2020-11-04T01:37:47Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>\n> With generation data chunk and corrected commit dates implemented, let's\n> update the technical documentation for commit-graph.\n>\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nNice.\n\n> ---\n>  .../technical/commit-graph-format.txt         | 21 +++++--\n>  Documentation/technical/commit-graph.txt      | 62 ++++++++++++++++---\n>  2 files changed, 69 insertions(+), 14 deletions(-)\n>\n> diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n> index b3b58880b9..08d9026ad4 100644\n> --- a/Documentation/technical/commit-graph-format.txt\n> +++ b/Documentation/technical/commit-graph-format.txt\n> @@ -4,11 +4,7 @@ Git commit graph format\n>  The Git commit graph stores a list of commit OIDs and some associated\n>  metadata, including:\n>  \n> -- The generation number of the commit. Commits with no parents have\n> -  generation number 1; commits with parents have generation number\n> -  one more than the maximum generation number of its parents. We\n> -  reserve zero as special, and can be used to mark a generation\n> -  number invalid or as \"not computed\".\n> +- The generation number of the commit.\n\nAll right, because we could store both generation number v1 and\ngeneration number v2 in the commit-graph file, and we need to describe\nboth, the description is now consolidated and in only one place.\n\n>  \n>  - The root tree OID.\n>  \n> @@ -86,13 +82,26 @@ CHUNK DATA:\n>        position. If there are more than two parents, the second value\n>        has its most-significant bit on and the other bits store an array\n>        position into the Extra Edge List chunk.\n> -    * The next 8 bytes store the generation number of the commit and\n> +    * The next 8 bytes store the topological level (generation number v1)\n> +      of the commit and\n\nAll right, this is updated information about CDAT chunk.\n\n>        the commit time in seconds since EPOCH. The generation number\n>        uses the higher 30 bits of the first 4 bytes, while the commit\n>        time uses the 32 bits of the second 4 bytes, along with the lowest\n>        2 bits of the lowest byte, storing the 33rd and 34th bit of the\n>        commit time.\n>  \n> +  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes)\n\nShould we mark this chunk as \"[Optional]\"?  Its absence is not an error.\n\n> +    * This list of 4-byte values store corrected commit date offsets for the\n> +      commits, arranged in the same order as commit data chunk.\n> +    * If the corrected commit date offset cannot be stored within 31 bits,\n> +      the value has its most-significant bit on and the other bits store\n> +      the position of corrected commit date into the Generation Data Overflow\n> +      chunk.\n\nAll right.\n\n> +\n> +  Generation Data Overflow (ID: {'G', 'D', 'O', 'V' }) [Optional]\n> +    * This list of 8-byte values stores the corrected commit dates for commits\n> +      with corrected commit date offsets that cannot be stored within 31 bits.\n\nA question: do we store 8-byte / 64-bit corrected commit date *directly*,\nor do we store corrected commit date *offset* as 8-byte / 64-bit value?\n\nPerhaps we should add the information that [like the EDGE chunk] it is\npresent only when necessary, and that it is present only when GDAT chunk\nis present (it might be obvious, but it could be better to state\nthis explicitly).\n\n> +\n\nAll right, this is the information about two new chunks (with the\nmentioned above caveat about the clarity of the description of\noverflow-handling chunk).\n\n>    Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n>        This list of 4-byte values store the second through nth parents for\n>        all octopus merges. The second parent value in the commit data stores\n> diff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\n> index f14a7659aa..75f71c4c7b 100644\n> --- a/Documentation/technical/commit-graph.txt\n> +++ b/Documentation/technical/commit-graph.txt\n> @@ -38,14 +38,31 @@ A consumer may load the following info for a commit from the graph:\n>  \n>  Values 1-4 satisfy the requirements of parse_commit_gently().\n>  \n> -Define the \"generation number\" of a commit recursively as follows:\n> +There are two definitions of generation number:\n> +1. Corrected committer dates (generation number v2)\n> +2. Topological levels (generation nummber v1)\n\nAll right.\n\n>  \n> - * A commit with no parents (a root commit) has generation number one.\n> +Define \"corrected committer date\" of a commit recursively as follows:\n>  \n> - * A commit with at least one parent has generation number one more than\n> -   the largest generation number among its parents.\n> +  * A commit with no parents (a root commit) has corrected committer date\n> +    equal to its committer date.\n\nMinor nitpick: the above point has been accidentally indented one space\nmore than necessary, and than is indented in other places.  Or maybe\nthat fixes / unifies the formatting... I am not sure.\n\n>  \n> -Equivalently, the generation number of a commit A is one more than the\n> +  * A commit with at least one parent has corrected committer date equal to\n> +    the maximum of its commiter date and one more than the largest corrected\n> +    committer date among its parents.\n> +\n> +  * As a special case, a root commit with timestamp zero has corrected commit\n> +    date of 1, to be able to distinguish it from GENERATION_NUMBER_ZERO\n> +    (that is, an uncomputed corrected commit date).\n\nAll right.  Looks good.\n\n> +\n> +Define the \"topological level\" of a commit recursively as follows:\n> +\n> + * A commit with no parents (a root commit) has topological level of one.\n> +\n> + * A commit with at least one parent has topological level one more than\n> +   the largest topological level among its parents.\n> +\n\nAll right, this just repeats what was written before, or in other words\nmove existing contents lower/later, just with 'generation number'\nreplaced by 'topological level' (though it might be not obvious from the\npatch because of the latter change).\n\n> +Equivalently, the topological level of a commit A is one more than the\n>  length of a longest path from A to a root commit. The recursive definition\n>  is easier to use for computation and observing the following property:\n>  \n> @@ -60,6 +77,9 @@ is easier to use for computation and observing the following property:\n>      generation numbers, then we always expand the boundary commit with highest\n>      generation number and can easily detect the stopping condition.\n>  \n> +The properties applies to both versions of generation number, that is both\n> +corrected committer dates and topological levels.\n> +\n\nI think it should be \"This property\" or \"The property\", not \"The\nproperties\"; it is a single property, a single condition.\n\nWe can alternatively say \"This condition is fulfilled by both versions...\",\nor \"This condition is true for both versions...\".\n\n>  This property can be used to significantly reduce the time it takes to\n>  walk commits and determine topological relationships. Without generation\n>  numbers, the general heuristic is the following:\n> @@ -67,7 +87,9 @@ numbers, the general heuristic is the following:\n>      If A and B are commits with commit time X and Y, respectively, and\n>      X < Y, then A _probably_ cannot reach B.\n>  \n> -This heuristic is currently used whenever the computation is allowed to\n> +In absence of corrected commit dates (for example, old versions of Git or\n> +mixed generation graph chains),\n> +this heuristic is currently used whenever the computation is allowed to\n>  violate topological relationships due to clock skew (such as \"git log\"\n>  with default order), but is not used when the topological order is\n>  required (such as merge base calculations, \"git log --graph\").\n\nAll right, this explains when commit date heuristics is used (which is\nless often than before).\n\n> @@ -77,7 +99,7 @@ in the commit graph. We can treat these commits as having \"infinite\"\n>  generation number and walk until reaching commits with known generation\n>  number.\n>  \n> -We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n> +We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n\nAll right, 64-bit GENERATION_NUMBER_INFINITY = 0xFFFFFFFFFFFFFFFF is a\nbit unwieldy...\n\n>  in the commit-graph file. If a commit-graph file was written by a version\n>  of Git that did not compute generation numbers, then those commits will\n>  have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n> @@ -93,7 +115,7 @@ fully-computed generation numbers. Using strict inequality may result in\n>  walking a few extra commits, but the simplicity in dealing with commits\n>  with generation number *_INFINITY or *_ZERO is valuable.\n>  \n> -We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n> +We use the macro GENERATION_NUMBER_MAX for commits whose\n\nThis should be\n\n  +We use the macro GENERATION_NUMBER_V1_MAX = 0x3FFFFFFF to for commits whose\n  +topological levels (generation number v1) are computed to be at least this value. We limit at\n   this value since it is the largest value that can be stored in the\n  +commit-graph file using the 30 bits available to topological levels. This\n\nWe need to use \"topological levels\" or \"generation numbers v1\" thorough\nthe rest of this section.\n\n>  generation numbers are computed to be at least this value. We limit at\n>  this value since it is the largest value that can be stored in the\n>  commit-graph file using the 30 bits available to generation numbers. This\n> @@ -267,6 +289,30 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n>  number of commits) could be extracted into config settings for full\n>  flexibility.\n>\n\nAll right, I agree that we don't need to write about overflow handling\nfor storing corrected committer dates (generation number v2) as offsets;\nthis is something format-specific, and this documentation is more about\nusing commit-graph data.  What is present in commit-graph-format.txt\nshould be enough information.\n\nSidenote: I wonder if other Git implementations such as JGit, Dulwich,\nGitoxide (gix), go-git have support for the commit-graph file...\n\n> +## Handling Mixed Generation Number Chains\n> +\n> +With the introduction of generation number v2 and generation data chunk, the\n> +following scenario is possible:\n> +\n> +1. \"New\" Git writes a commit-graph with the corrected commit dates.\n> +2. \"Old\" Git writes a split commit-graph on top without corrected commit dates.\n> +\n> +A naive approach of using the newest available generation number from\n> +each layer would lead to violated expectations: the lower layer would\n> +use corrected commit dates which are much larger than the topological\n> +levels of the higher layer. For this reason, Git inspects each layer to\n> +see if any layer is missing corrected commit dates. In such a case, Git\n> +only uses topological level\n\nThis should end in full stop:\n\n  +only uses topological levels.\n\nOr maybe we should expand the last sentence a bit:\n\n  +only uses topological levels for generation numbers.\n\nSidenote: it is a good explanation, even if Git can make use of the\nproperty described below that only topmost layers might be missing\ncorrected commit graph by the construction (so it needs to check only\nthe top layer).\n\n> +\n> +When writing a new layer in split commit-graph, we write corrected commit\n> +dates if the topmost layer has corrected commit dates written. This\n> +guarantees that if a layer has corrected commit dates, all lower layers\n> +must have corrected commit dates as well.\n> +\n> +When merging layers, we do not consider whether the merged layers had corrected\n> +commit dates. Instead, the new layer will have corrected commit dates if and\n> +only if all existing layers below the new layer have corrected commit dates.\n> +\n\nPerhaps we should explicitly say that when rewriting split commit-graph\nas a single file (`--split=replace`) then the newly created single layer\nwould store corrected commit dates.\n\n>  ## Deleting graph-{hash} files\n>  \n>  After a new tip file is written, some `graph-{hash}` files may no longer\n\nBest,\n-- \nJakub Narębski\n"},{"id":"409096","messageId":"85pn4tnk8u.fsf@gmail.com","threadId":"53933","inReplyTo":"20201103114432.GA3577@Abhishek-Arch","subject":"Re: [PATCH v4 06/10] commit-graph: implement corrected commit date","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-11-04T16:45:53Z","receivedAt":"2020-11-04T16:45:57Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello Abhishek,\n\nAbhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Tue, Oct 27, 2020 at 07:53:23PM +0100, Jakub Narębski wrote:\n>> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>> \n>>> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>>> ...\n>>> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>> \n>> Somewhere in the commit message we should also describe that this commit\n>> changes how commit-graph is verified: from checking that the generation\n>> number agrees with _topological level definition_, that is that for a\n>> given commit it is 1 more than maximum of its parents (with the caveat\n>> that we need to handle GENERATION_NUMBER_V1_MAX values correctly), to\n>> checking that slightly weaker condition fulfilled by both topological\n>> levels (generation number v1) and by corrected commit date (generation\n>> number v2) that for a given commit its generation number is 1 more than\n>> maximum of its parents or larger.\n>\n> Sure, that makes sense. Will add.\n\nActually this description should match whatever we decide about\nmechanism for verifying correctness of generation numbers (see below).\nBecause we have to choose one.\n\n>> \n>> But, as far as I understand it, current code does not handle correctly\n>> GENERATION_NUMBER_V1_MAX case (if we use generation number v1).\n>> \n>> On the other hand we could have simpy use functional check, that\n>> generation number used (which can be v1 or v2, or any similar other)\n>> fulfills the reachability condition for each edge, which can be\n>> simplified to checking that generation(parents) <= generation(commit).\n>> If the reachability condition is true for each edge, then it is true for\n>> each path, and for each commit.\n\nSee below.\n\n>>> ---\n>>>  commit-graph.c | 43 +++++++++++++++++++++++--------------------\n>>>  1 file changed, 23 insertions(+), 20 deletions(-)\n>>>\n>>> diff --git a/commit-graph.c b/commit-graph.c\n>>> index cedd311024..03948adfce 100644\n>>> --- a/commit-graph.c\n>>> +++ b/commit-graph.c\n>>> @@ -154,11 +154,6 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>>>  \telse if (generation_a > generation_b)\n>>>  \t\treturn 1;\n>>>  \n>>> -\t/* use date as a heuristic when generations are equal */\n>>> -\tif (a->date < b->date)\n>>> -\t\treturn -1;\n>>> -\telse if (a->date > b->date)\n>>> -\t\treturn 1;\n>> \n>> Why this change?  It is not described in the commit message.\n>> \n>> Note that while this tie-breaking fallback doesn't make much sense for\n>> corrected committer date generation number v2, this tie-breaking helps\n>> if we have to use topological levels (generation number v2).\n>> \n>\n> Right, I should have mentioned this change (and it's not something that\n> makes a difference either way).\n>\n> We call commit_gen_cmp() only when we are sorting commits by generation\n> to speed up computation of Bloom filters i.e. while writing a commit\n> graph (either split commit-graph or a simple commit-graph).\n>\n> Since we are always computing and storing corrected commit date when we\n> are writing (whether we write a GDAT chunk or not), using date as\n> heuristic is longer required.\n\nThanks.  This description really should be added to the commit message,\nbecause (yet again?) I was confused by this change.\n\nSidenote: it is not obvious at least to me that this function is used\nonly for sorting commits to speed up computation of Bloom filters while\nwriting the commit-graph (`git commit-graph write --changed-paths [other\noptions]`).\n\n>>>  \treturn 0;\n>>>  }\n>>>  \n>>> @@ -1357,10 +1352,14 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>>>  \t\t\t\t\tctx->commits.nr);\n>>>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>>>  \t\ttimestamp_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n>> \n>> Sidenote: I haven't noticed it earlier, but here 'uint32_t' might be\n>> enough; no need for 'timestamp_t' for 'level' variable.\n>> \n>>> +\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n>>>\n>\n> We need the 'timestamp_t' as we are comparing level with the now 64-bits\n> GENERATION_NUMBER_INFINITY. I thought uint32_t would be promoted to\n> timestamp_t. I have a hunch that since we are explicitly using a fixed\n> width data type, compiler is unwilling to type coerce into broader data\n> types.\n>\n> Advice on this appreciated.\n\nAll right, so the wider type is used because of comparison with\nwide-uint GENERATION_NUMBER_INFINITY.  I stand corrected.\n\n[...]\n>>> @@ -2485,17 +2496,9 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n>>>  \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n>>>  \t\t\tcontinue;\n>>>  \n>>> -\t\t/*\n>>> -\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n>>> -\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n>>> -\t\t * extra logic in the following condition.\n>>> -\t\t */\n>>> -\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n>>> -\t\t\tmax_generation--;\n>>> -\n>> \n>> Perhaps in the future we should check that both topological levels, and\n>> also corrected committer date (if it exists) for correctness according\n>> to their definition.  Then the above removed part would be restored (but\n>> with s/max_generation/max_level/).\n>> \n>>>  \t\tgeneration = commit_graph_generation(graph_commit);\n>>> -\t\tif (generation != max_generation + 1)\n>>> -\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n>>> +\t\tif (generation < max_generation + 1)\n>>> +\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n>> \n>> All right, so we relaxed the check so that it will be fulfilled by\n>> generation number v2 (and also by generation number v1, as it implies\n>> the more strict check for v1).\n>> \n>> What would happen however if generation holds topological levels, and it\n>> is GENERATION_NUMBER_V1_MAX for at least one parent, which means it is\n>> GENERATION_NUMBER_V1_MAX for a commit?  As you can check, the condition\n>> would be true: GENERATION_NUMBER_V1_MAX < GENERATION_NUMBER_V1_MAX + 1,\n>> so the `git commit-graph verify` would incorrectly say that there is\n>> a problem with generation number, while there isn't one (false positive\n>> detection of error).\n>\n> Alright, so the above block still makes sense if we are working with\n> topological levels but not with corrected commit dates. Instead of\n> removing it, I will modify the condition to check that one of our parents\n> has GENERATION_NUMBER_V1_MAX and the graph uses topological levels.\n\nThat is one of the 3 possible solutions I can think of.\n\n\nI. First solution is to switch from checking that generation number\nmatches its definition to checking that the [weaker] reachability\ncondition for the generation number is true, that is:\n\n \tif (generation < max_generation)\n \t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n\nThe [weaker] reachability condition for generation numbers states that\n\n   A reachable from B    =>    gen(A) <= gen(B)\n\nThis condition is true even if one or more generation numbers is\nGENERATION_NUMBER_ZERO (uninitialized or written by old git version),\nGENERATION_NUMBER_V1_MAX (we hit storage limitations, can happen only\nfor generation number v1), or GENERATION_NUMBER_INFINITY (for commits\noutside of the serialized commit-graph, doesn't matter and cannot happen\nduring verification of the commit-graph data by definition).\n\nThis means that if P* is the parent of C with the maximal generation\nnumber, and gen(C) < gen(P*) is true (while gen(P*) <= gen(C) should be\ntrue), then there is a problem with generation number.\n\nThis is why I thought you were going for, and what I have proposed.\n\nAdvantages:\n- we are testing what actually matters for speeding up reachability\n  queries, namely that the reachability property holds true\n- the test works for generation number v1, generation number v2,\n  and any possible future use-compatibile generation number\n  (not that I think we would need any)\n- least complicated solution\n\nDisadvantages:\n- weaker test that we have had for generation number v1 (topological\n  levels), and weaker that possible test for generation number v2\n  that we could have (see below)\n\n\nII. Verify corrected committed date (generation number v2) if available,\nand verify topological levels (generation number v1) otherwise, checking\nthat it matches the definition of it -- using version-specific checks.\n\nThis would probably mean adding a conditional around the code verifying\nthat given generation number is correct, possibly:\n\n  if (g->read_generation_data) {\n  \t/* verify corrected commit date */\n  } else {\n  \t/* current code for verifying topological levels */\n  }\n\nII.a. For topological levels (generation number v1) we would continue\nchecking that it matches the definition, that is that the following\ncondition holds:\n\n  gen(C) = max_{P: P ∈ parents(C)} gen(P) + 1\n\nThis includes code for handling the case where `max_generation`, holding \nmax_{P: P ∈ parents(C)} gen(P), is GENERATION_NUMBER_V1_MAX.\n\nII.b. For corrected commiter dates (generation number v2) we can use the\ncode proposed by this revision of this commit, namely we check if the\nfollowing condition holds:\n\n  gen(P) + 1 <= gen(C)   for each P \\in parents(C)\n\nor, in other words:\n\n  max_{P: P ∈ parents(C)} { gen(P) } + 1  <=  gen(C)\n\nWhich could be checked using the following code (i.e. current state\nafter this revision of this patch):\n\n\tif (generation < max_generation + 1)\n\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n\nThis is what I think you are proposing now.\n\nAdditionally, theoretically we could also check that the following\ncondition holds for corrected commiter date:\n\n   committer_date(C) <= gen_v2(C)\n\nbut this is automatically fufilled because we use non-negative offsets\nto store corrected committed date info.\n\nAlternatively we can check for compliance with the definition of the\ncorrected committer date:\n\n  if (max_generation + 1 <= graph_commit->date) {\n  \t/* commit date does not need correction */\n  \tif (generation != graph_commit->date)\n    \tgraph_report(_(\"commit-graph corrected commit date for commit %s \"\n  \t\t               \"is %\"PRItime\" != %\"PRItime\" commit date\"),\n                     ...);\n  } else {\n  \tif (generation != max_generation + 1)\n  \t\tgraph_report(_(\"commit-graph generation v2 for commit %s is %\"PRItime\" != %\"PRItime),\n                     ...);\n  }\n\nThough I think it might be overkill.\n\nAdvantages:\n- more strict tests, checking generation numbers (v2 if present, v1\n  otherwise) against their definition\n- if there is no GDAT chunk, verify works just like it did before\n\nDisadvantages:\n- more complicated code\n- possibly measurable performance degradation due to extra conditional\n\n\nIII. Like II., but if there is generation numbers chunk (GDAT chunk), we\nverify *both* topological levels (v1) and corrected commit date (v2)\nagainst their definition.  If GDAT chunk is not present, it reduces to\ncurrent code (before this patch series).\n\nAdvantages:\n- if there is no GDAT chunk, verify works just like it did before\n- most strict tests, verifying all the data: both generation number v1\n  and generation number v2 -- if possible\n\nDisadvantages:\n- most complex code; we need to somehow extract topological levels\n  if the GDAT chunk is present (they are not on graph data slab in this\n  case); I have not even started to think how it could be done\n- slower verification\n\n> Suprised that no test breaks by this change.\n\nI don't whink we have any test that created commit graph with\ntopological levels greater than GENERATION_NUMBER_V1_MAX; this would be\nexpensive and have to be of course protected by GIT_TEST_LONG aka\nEXPENSIVE prerequisite.\n\n  # GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 is here to force verification of topological levels\n  test_expect_success EXPENSIVE 'verify handles topological levels > GENERATION_NUMBER_V1_MAX' '\n  \trm -rf long_chain &&\n  \tgit init long_chain &&\n  \ttest_commit_bulk -C long_chain 1073741824 &&\n    (\n  \t\tcd long_chain &&\n  \t\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write &&\n  \t\tgit commit-graph verify\n    )\n  '\n\nThis however lies slightly outside the scope of this patch series,\nthough if you could add this test (in a separate patch), after testing\nit, it would be very nice.\n\n>\n> I have also moved changes in the verify function to the next patch, as\n> we cannot write or read corrected commit dates yet - so little sense in\n> modifying verify.\n\nI think putting changes to the verify function in a separate patch, be\nit before or after this one (depending on the choice of the algorithm\nfor verification, see above) would be a good idea.\n\n>> \n>> Sidenote: I think we don't have to worry about having to introduce\n>> GENERATION_NUMBER_V2_MAX, as the in-memory size (of reconstructed from\n>> disck representation) corrected commiter date is the same as of commiter\n>> date itself, plus some, and I don't see us coming close to 64-bit limit\n>> of timestamp_t for commit dates.\n>> \n>>>  \t\t\t\t     oid_to_hex(&cur_oid),\n>>>  \t\t\t\t     generation,\n>>>  \t\t\t\t     max_generation + 1);\n\nBest,\n-- \nJakub Narębski\n"},{"id":"409177","messageId":"85tuu4lmlu.fsf@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 00/10] [GSoC] Implement Corrected Commit Date","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-11-04T23:37:49Z","receivedAt":"2020-11-04T23:37:54Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hi Abhishek,\n\n\"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> This patch series implements the corrected commit date offsets as generation\n> number v2, along with other pre-requisites.\n\nThanks a lot for continued working on this patch series.\n\n>\n> Git uses topological levels in the commit-graph file for commit-graph\n> traversal operations like git log --graph. Unfortunately, using topological\n> levels can result in a worse performance than without them when compared\n> with committer date as a heuristics. For example, git merge-base v4.8 v4.9 \n> on the Linux repository walks 635,579 commits using topological levels and\n> walks 167,468 using committer date.\n\nVery minor nitpick: it would make it easier to read if the commands\nthemself would be put inside single quotes or backticks, e.g. `git log\n--graph` and `git merge-base v4.8 v4.9`.\n\nI wonder if it is worth mentioning (probably not) that this performance\nhit was the reason why since 091f4cf3 `git merge-base` uses committer\ndate heuristics unless there is a cutoff and using topological levels\n(generation date v1) is expected to give better performance.\n\n>\n> Thus, the need for generation number v2 was born. New generation number\n> needed to provide good performance, increment updates, and backward\n> compatibility. Due to an unfortunate problem 1\n\nMinor issue: this looks a bit strange; is there an error in formatting\nthis part?\n\n> [https://public-inbox.org/git/87a7gdspo4.fsf@evledraar.gmail.com/], we also\n> needed a way to distinguish between the old and new generation number\n> without incrementing graph version.\n>\n> Various candidates were examined (https://github.com/derrickstolee/gen-test, \n> https://github.com/abhishekkumar2718/git/pull/1). The proposed generation\n> number v2, Corrected Commit Date with Mononotically Increasing Offsets \n> performed much worse than committer date (506,577 vs. 167,468 commits walked\n> for git merge-base v4.8 v4.9) and was dropped.\n>\n> Using Generation Data chunk (GDAT) relieves the requirement of backward\n> compatibility as we would continue to store topological levels in Commit\n> Data (CDAT) chunk.\n\nNice writeup about the history of generation number v2, much appreciated.\n\n>                    Thus, Corrected Commit Date was chosen as generation\n> number v2. The Corrected Commit Date is defined as:\n\nMinor nitpick: it would be probably better to use \"is defined as\nfollows.\" instead of \"is defined as:\".\n\n>\n> For a commit C, let its corrected commit date be the maximum of the commit\n> date of C and the corrected commit dates of its parents plus 1. Then \n> corrected commit date offset is the difference between corrected commit date\n> of C and commit date of C. As a special case, a root commit with timestamp\n> zero has corrected commit date of 1 to be able distinguish it from\n> GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit date).\n\nVery minor nitpick: s/with timestamp/with *the* timestamp/, and\ns/to be able distinguish/to be able *to* distinguish/ (without the '*'\nused to mark the additions).\n\n>\n> We will introduce an additional commit-graph chunk, Generation Data chunk,\n\nOr \"Generation DATa chunk\", if we want to emphasize where its name came\nfrom, or even \"Generation DATa (GDAT) chunk\". But it is fine as it is\nnow, though it would be good idea to write \"Generation Data (GDAT)\nchunk\" to explicitly state its name / shortcut.\n\n> and store corrected commit date offsets in GDAT chunk while storing\n> topological levels in CDAT chunk. The old versions of Git would ignore GDAT\n> chunk, using topological levels from CDAT chunk. In contrast, new versions\n> of Git would use corrected commit dates, falling back to topological level\n> if the generation data chunk is absent in the commit-graph file.\n\nNice writeup of handling the backward compatibility.\n\n>\n> While storing corrected commit date offsets saves us 4 bytes per commit (as\n> compared with storing corrected commit dates directly), it's possible for\n> the offset to overflow the space allocated. To handle such cases, we\n> introduce a new chunk, Generation Data Overflow (GDOV) that stores the\n> corrected commit date. For overflowing offsets, we set MSB and store the\n> position into the GDOV chunk, in a mechanism similar to the Extra Edges list\n> chunk.\n\nVery minor suggestion: perhaps it would be better to use \"it's however\npossible\".\n\nVery minor suggestion: \"it's possible for the offset to overflow\" could\nbe simplified to just \"the offset can overflow\"... though the simplified\nversion loses a bit of hint that the overflow should be very rare in\nreal repositories.\n\nBut it is just fine as it is now; I am not a native English speaker to\njudge which version is better.\n\n>\n> For mixed generation number environment (for example new Git on the command\n> line, old Git used by GUI client), we can encounter a mixed-chain\n> commit-graph (a commit-graph chain where some of split commit-graph files\n> have GDAT chunk and others do not). As backward compatibility is one of the\n> goals, we can define the following behavior:\n>\n> While reading a mixed-chain commit-graph version, we fall back on\n> topological levels as corrected commit dates and topological levels cannot\n> be compared directly.\n>\n> While writing on top of a split commit-graph, we check if the tip of the\n> chain has a GDAT chunk. If it does, we append to the chain, writing GDAT\n> chunk. Thus, we guarantee if the topmost split commit-graph file has a GDAT\n> chunk, rest of the chain does too.\n>\n> If the topmost split commit-graph file does not have a GDAT chunk (meaning\n> it has been appended by the old Git), we write without GDAT chunk. We do\n> write a GDAT chunk when the existing chain does not have GDAT chunk - when\n> we are writing to the commit-graph chain with the 'replace' strategy.\n\nI think the last paragraph can be simplified (or added to) by explicitly\nstating the goal:\n\n  When adding new layer to the split commit-graph file, and when merging\n  some or all layers (replacing them in the latter case), the new layer\n  will have GDAT chunk if and only if in the final result there would be\n  no layer without GDAT chunk just below it.\n\n>\n> Thanks to Dr. Stolee, Dr. Narębski, and Taylor for their reviews.\n\nYou are welcome.\n\n>\n> I look forward to everyone's reviews!\n>\n> Thanks\n>\n>  * Abhishek\n>\n>\n> ----------------------------------------------------------------------------\n>\n> Changes in version 4:\n>\n>  * Added GDOV to handle overflows in generation data.\n>  * Added a test for writing tip graph for a generation number v2 graph chain\n>    in t5324-split-commit-graph.sh\n>  * Added a section on how mixed generation number chains are handled in \n>    Documentation/technical/commit-graph-format.txt\n>  * Reverted unimportant whitespace style changes in commit-graph.c\n>  * Added header comments about the order of comparision for\n>    compare_commits_by_gen_then_commit_date in commit.h,\n>    compare_commits_by_gen in commit-graph.h\n>  * Elaborated on why t6404 fails with corrected commit date and must be run\n>    with GIT_TEST_COMMIT_GRAPH=1 in the commit \"commit-reach: use corrected\n>    commit dates in paint_down_to_common()\"\n>  * Elaborated on write behavior for mixed generation number chains in the\n>    commit \"commit-graph: use generation v2 only if entire chain does\"\n>  * Added notes about adding the topo_level slab to struct\n>    write_commit_graph_context as well as struct commit_graph.\n>  * Clarified commit message for \"commit-graph: consolidate\n>    fill_commit_graph_info\"\n>  * Removed the claim \"GDAT can store future generation numbers\" because it\n>    hasn't been tested yet.\n>\n> Changes in version 3:\n>\n>  * Reordered patches as discussed in 2\n>    [https://lore.kernel.org/git/aee0ae56-3395-6848-d573-27a318d72755@gmail.com/]\n>    .\n>  * Split \"implement corrected commit date\" into two patches - one\n>    introducing the topo level slab and other implementing corrected commit\n>    dates.\n>  * Extended split-commit-graph tests to verify at the end of test.\n>  * Use topological levels as generation number if any of split commit-graph\n>    files do not have generation data chunk.\n>\n> Changes in version 2:\n>\n>  * Add tests for generation data chunk.\n>  * Add an option GIT_TEST_COMMIT_GRAPH_NO_GDAT to control whether to write\n>    generation data chunk.\n>  * Compare commits with corrected commit dates if present in\n>    paint_down_to_common().\n>  * Update technical documentation.\n>  * Handle mixed generation commit chains.\n>  * Improve commit messages for \"commit-graph: fix regression when computing\n>    bloom filter\", \"commit-graph: consolidate fill_commit_graph_info\",\n>  * Revert unnecessary whitespace changes.\n>  * Split uint_32 -> timestamp_t change into a new commit.\n\nAfter careful review of those 10 patches it looks like the series is\nclose to being ready, requiring only small changes to progress.\n\n> Abhishek Kumar (10):\n>   commit-graph: fix regression when computing Bloom filters\n\n    All good, beside possible improvement to the commit message.\n    Thanks to Taylor Blau for discovering possible reason for strange\n    no change in performance.\n\n>   revision: parse parent in indegree_walk_step()\n\n    Looks good.\n\n>   commit-graph: consolidate fill_commit_graph_info\n\n    Needs to fix now duplicated test names (minor change).\n    Proposed possible improvement to the commit message.\n\n>   commit-graph: return 64-bit generation number\n\n    Needs fixing due to mismerge: there should be no switch from\n    using GENERATION_NUMBER_ZERO to using GENERATION_NUMBER_INFINITY.\n    Possible minor improvement to the commit message.\n\n>   commit-graph: add a slab to store topological levels\n\n    Possible minor improvement to the commit message.\n    \n    There is also not very important issue, but something that would be\n    nice to explain, namely that checks for GENERATION_NUMBER_INFINITY \n    can never be true, as topo_level_slab_at() returns 0 for commits\n    outside the commit-graph, not GENERATION_NUMBER_INFINITY.  It works\n    but it is not obvious why.\n\n>   commit-graph: implement corrected commit date\n\n    The change to commit-graph verification needs fixing, and we need to\n    decide how verifying generation numbers should work.  Perhaps a test\n    for handling topological level of GENERATION_NUMBER_V1_MAX could be\n    added (though this might be left for ater).\n\n    The changes to `git commit-graph verify` code could be put into\n    separate patch, either before or after this one.\n\n>   commit-graph: implement generation data chunk\n\n    Proposed possible improvement to the commit message.\n    The commit message does not explain why given shape of history is\n    needed to test handling corrected commit date offset overflow.\n\n    Proposed minor corrections to the coding style.\n\n    Instead of looping again through all commits when handling overflow\n    in corrected commit date offsets, while there should be at most a\n    few commits needing it, why not save those commits on list and loop\n    only through those commits?  Though this _possible_ performance\n    improvement could be left to the followup...\n\n    test_commit_with_date() could be instead implemented via adding\n    `--date <date>` option to test_commit() in test-lib-functions.sh.\n\n    Also, to reduce \"noise\" in this patch, the rename of\n    run_three_modes() to run_all_modes() and test_three_modes() to\n    test_all_modes() could have been done in a separate preparatory\n    patch. It would be pure refactoring patch, without introducing any\n    new functionality.  But it is not something that is necessary.\n\n>   commit-graph: use generation v2 only if entire chain does\n\n    Proposed possible improvement to the commit message.\n    Proposed minor corrections to the coding style (also in tests).\n\n    There is a question whether merging layers or replacing them should\n    honor GIT_TEST_COMMIT_GRAPH_NO_GDAT.\n\n    Tests possibly could be made more strict, and check more things\n    explicitly. One test we are missing is testing that merging layers\n    is done correctly, namely that if we are merging layers in split\n    commit-graph file, and the layer below the ones we are merging lacks\n    GDAT chunk, then the result of the merge should also be without GDAT\n    chunk -- but that might be left for later.\n\n>   commit-reach: use corrected commit dates in paint_down_to_common()\n\n    This patch consist of two slightly interleaved changes, which\n    possibly could be separated: change to paint_down_to_common() and\n    change to t6404-recursive-merge test.\n\n    In the commit message for the paint_down_to_common() we should\n    explicitly mention 091f4cf3, which this one partially reverts.\n\n    Possible accidental change, question about function naming.\n\n>   doc: add corrected commit date info\n\n    Needs further improvements to the documentation, like adding\n    \"[Optional]\" to chunk description, and leftover switching from\n    \"generation numbers\" to \"topological levels\" in one place.\n\n>\n>  .../technical/commit-graph-format.txt         |  21 +-\n>  Documentation/technical/commit-graph.txt      |  62 ++++-\n>  commit-graph.c                                | 256 ++++++++++++++----\n>  commit-graph.h                                |  17 +-\n>  commit-reach.c                                |  38 +--\n>  commit-reach.h                                |   2 +-\n>  commit.c                                      |   4 +-\n>  commit.h                                      |   5 +-\n>  revision.c                                    |  13 +-\n>  t/README                                      |   3 +\n>  t/helper/test-read-graph.c                    |   4 +\n>  t/t4216-log-bloom.sh                          |   4 +-\n>  t/t5000-tar-tree.sh                           |  20 +-\n>  t/t5318-commit-graph.sh                       |  70 ++++-\n>  t/t5324-split-commit-graph.sh                 |  98 ++++++-\n>  t/t6404-recursive-merge.sh                    |   5 +-\n>  t/t6600-test-reach.sh                         |  68 ++---\n>  upload-pack.c                                 |   2 +-\n>  18 files changed, 534 insertions(+), 158 deletions(-)\n>\n>\n> base-commit: d98273ba77e1ab9ec755576bc86c716a97bf59d7\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-676%2Fabhishekkumar2718%2Fcorrected_commit_date-v4\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-676/abhishekkumar2718/corrected_commit_date-v4\n> Pull-Request: https://github.com/gitgitgadget/git/pull/676\n[...]\n\nBest,\n-- \nJakub Narębski\n"},{"id":"409201","messageId":"efa3488a-3983-3435-e5e4-2eb71e76a33a@iee.email","threadId":"53933","inReplyTo":"85pn4tnk8u.fsf@gmail.com","subject":"Re: [PATCH v4 06/10] commit-graph: implement corrected commit date","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2020-11-05T14:05:06Z","receivedAt":"2020-11-05T14:05:36Z","isPatch":true,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Hi Abhishek,\n\nOn 04/11/2020 16:45, Jakub Narębski wrote:\n> Hello Abhishek,\n>\n> Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n>> On Tue, Oct 27, 2020 at 07:53:23PM +0100, Jakub Narębski wrote:\n>>> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>>\n>>>> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>>>> ...\n>>>> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>>> Somewhere in the commit message we should also describe that this commit\n>>> changes how commit-graph is verified: from checking that the generation\n>>> number agrees with _topological level definition_, that is that for a\n>>> given commit it is 1 more than maximum of its parents (with the caveat\n>>> that we need to handle GENERATION_NUMBER_V1_MAX values correctly), to\n>>> checking that slightly weaker condition fulfilled by both topological\n>>> levels (generation number v1) and by corrected commit date (generation\n>>> number v2) that for a given commit its generation number is 1 more than\n>>> maximum of its parents or larger.\n>> Sure, that makes sense. Will add.\n> Actually this description should match whatever we decide about\n> mechanism for verifying correctness of generation numbers (see below).\n> Because we have to choose one.\n\nThis may be not part of the the main project, but could you consider, if\ntime permits, also adding some entries into the Git Glossary (`git help\nglossary`) for the various terms we are using here and elsewhere, e.g.\n'topological levels', 'generation number', 'corrected commit date' (and\nits fancy technical name for the use of date heuristics e.g. the\n'chronological ordering';).\n\nThe glossary can provide a reference, once the issues are resolved. The\nHistory Simplification and Commit Ordering section of git-log maybe a\nuseful guide to some of the terms that would link to the glossary.\n--\nPhilip\n\n>>> But, as far as I understand it, current code does not handle correctly\n>>> GENERATION_NUMBER_V1_MAX case (if we use generation number v1).\n>>>\n>>> On the other hand we could have simpy use functional check, that\n>>> generation number used (which can be v1 or v2, or any similar other)\n>>> fulfills the reachability condition for each edge, which can be\n>>> simplified to checking that generation(parents) <= generation(commit).\n>>> If the reachability condition is true for each edge, then it is true for\n>>> each path, and for each commit.\n> See below.\n>\n>>>> ---\n>>>>  commit-graph.c | 43 +++++++++++++++++++++++--------------------\n>>>>  1 file changed, 23 insertions(+), 20 deletions(-)\n>>>>\n>>>> diff --git a/commit-graph.c b/commit-graph.c\n>>>> index cedd311024..03948adfce 100644\n>>>> --- a/commit-graph.c\n>>>> +++ b/commit-graph.c\n>>>> @@ -154,11 +154,6 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>>>>  \telse if (generation_a > generation_b)\n>>>>  \t\treturn 1;\n>>>>  \n>>>> -\t/* use date as a heuristic when generations are equal */\n>>>> -\tif (a->date < b->date)\n>>>> -\t\treturn -1;\n>>>> -\telse if (a->date > b->date)\n>>>> -\t\treturn 1;\n>>> Why this change?  It is not described in the commit message.\n>>>\n>>> Note that while this tie-breaking fallback doesn't make much sense for\n>>> corrected committer date generation number v2, this tie-breaking helps\n>>> if we have to use topological levels (generation number v2).\n>>>\n>> Right, I should have mentioned this change (and it's not something that\n>> makes a difference either way).\n>>\n>> We call commit_gen_cmp() only when we are sorting commits by generation\n>> to speed up computation of Bloom filters i.e. while writing a commit\n>> graph (either split commit-graph or a simple commit-graph).\n>>\n>> Since we are always computing and storing corrected commit date when we\n>> are writing (whether we write a GDAT chunk or not), using date as\n>> heuristic is longer required.\n> Thanks.  This description really should be added to the commit message,\n> because (yet again?) I was confused by this change.\n>\n> Sidenote: it is not obvious at least to me that this function is used\n> only for sorting commits to speed up computation of Bloom filters while\n> writing the commit-graph (`git commit-graph write --changed-paths [other\n> options]`).\n>\n>>>>  \treturn 0;\n>>>>  }\n>>>>  \n>>>> @@ -1357,10 +1352,14 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>>>>  \t\t\t\t\tctx->commits.nr);\n>>>>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>>>>  \t\ttimestamp_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n>>> Sidenote: I haven't noticed it earlier, but here 'uint32_t' might be\n>>> enough; no need for 'timestamp_t' for 'level' variable.\n>>>\n>>>> +\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n>>>>\n>> We need the 'timestamp_t' as we are comparing level with the now 64-bits\n>> GENERATION_NUMBER_INFINITY. I thought uint32_t would be promoted to\n>> timestamp_t. I have a hunch that since we are explicitly using a fixed\n>> width data type, compiler is unwilling to type coerce into broader data\n>> types.\n>>\n>> Advice on this appreciated.\n> All right, so the wider type is used because of comparison with\n> wide-uint GENERATION_NUMBER_INFINITY.  I stand corrected.\n>\n> [...]\n>>>> @@ -2485,17 +2496,9 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n>>>>  \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n>>>>  \t\t\tcontinue;\n>>>>  \n>>>> -\t\t/*\n>>>> -\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n>>>> -\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n>>>> -\t\t * extra logic in the following condition.\n>>>> -\t\t */\n>>>> -\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n>>>> -\t\t\tmax_generation--;\n>>>> -\n>>> Perhaps in the future we should check that both topological levels, and\n>>> also corrected committer date (if it exists) for correctness according\n>>> to their definition.  Then the above removed part would be restored (but\n>>> with s/max_generation/max_level/).\n>>>\n>>>>  \t\tgeneration = commit_graph_generation(graph_commit);\n>>>> -\t\tif (generation != max_generation + 1)\n>>>> -\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n>>>> +\t\tif (generation < max_generation + 1)\n>>>> +\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n>>> All right, so we relaxed the check so that it will be fulfilled by\n>>> generation number v2 (and also by generation number v1, as it implies\n>>> the more strict check for v1).\n>>>\n>>> What would happen however if generation holds topological levels, and it\n>>> is GENERATION_NUMBER_V1_MAX for at least one parent, which means it is\n>>> GENERATION_NUMBER_V1_MAX for a commit?  As you can check, the condition\n>>> would be true: GENERATION_NUMBER_V1_MAX < GENERATION_NUMBER_V1_MAX + 1,\n>>> so the `git commit-graph verify` would incorrectly say that there is\n>>> a problem with generation number, while there isn't one (false positive\n>>> detection of error).\n>> Alright, so the above block still makes sense if we are working with\n>> topological levels but not with corrected commit dates. Instead of\n>> removing it, I will modify the condition to check that one of our parents\n>> has GENERATION_NUMBER_V1_MAX and the graph uses topological levels.\n> That is one of the 3 possible solutions I can think of.\n>\n>\n> I. First solution is to switch from checking that generation number\n> matches its definition to checking that the [weaker] reachability\n> condition for the generation number is true, that is:\n>\n>  \tif (generation < max_generation)\n>  \t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n>\n> The [weaker] reachability condition for generation numbers states that\n>\n>    A reachable from B    =>    gen(A) <= gen(B)\n>\n> This condition is true even if one or more generation numbers is\n> GENERATION_NUMBER_ZERO (uninitialized or written by old git version),\n> GENERATION_NUMBER_V1_MAX (we hit storage limitations, can happen only\n> for generation number v1), or GENERATION_NUMBER_INFINITY (for commits\n> outside of the serialized commit-graph, doesn't matter and cannot happen\n> during verification of the commit-graph data by definition).\n>\n> This means that if P* is the parent of C with the maximal generation\n> number, and gen(C) < gen(P*) is true (while gen(P*) <= gen(C) should be\n> true), then there is a problem with generation number.\n>\n> This is why I thought you were going for, and what I have proposed.\n>\n> Advantages:\n> - we are testing what actually matters for speeding up reachability\n>   queries, namely that the reachability property holds true\n> - the test works for generation number v1, generation number v2,\n>   and any possible future use-compatibile generation number\n>   (not that I think we would need any)\n> - least complicated solution\n>\n> Disadvantages:\n> - weaker test that we have had for generation number v1 (topological\n>   levels), and weaker that possible test for generation number v2\n>   that we could have (see below)\n>\n>\n> II. Verify corrected committed date (generation number v2) if available,\n> and verify topological levels (generation number v1) otherwise, checking\n> that it matches the definition of it -- using version-specific checks.\n>\n> This would probably mean adding a conditional around the code verifying\n> that given generation number is correct, possibly:\n>\n>   if (g->read_generation_data) {\n>   \t/* verify corrected commit date */\n>   } else {\n>   \t/* current code for verifying topological levels */\n>   }\n>\n> II.a. For topological levels (generation number v1) we would continue\n> checking that it matches the definition, that is that the following\n> condition holds:\n>\n>   gen(C) = max_{P: P ∈ parents(C)} gen(P) + 1\n>\n> This includes code for handling the case where `max_generation`, holding \n> max_{P: P ∈ parents(C)} gen(P), is GENERATION_NUMBER_V1_MAX.\n>\n> II.b. For corrected commiter dates (generation number v2) we can use the\n> code proposed by this revision of this commit, namely we check if the\n> following condition holds:\n>\n>   gen(P) + 1 <= gen(C)   for each P \\in parents(C)\n>\n> or, in other words:\n>\n>   max_{P: P ∈ parents(C)} { gen(P) } + 1  <=  gen(C)\n>\n> Which could be checked using the following code (i.e. current state\n> after this revision of this patch):\n>\n> \tif (generation < max_generation + 1)\n> \t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n>\n> This is what I think you are proposing now.\n>\n> Additionally, theoretically we could also check that the following\n> condition holds for corrected commiter date:\n>\n>    committer_date(C) <= gen_v2(C)\n>\n> but this is automatically fufilled because we use non-negative offsets\n> to store corrected committed date info.\n>\n> Alternatively we can check for compliance with the definition of the\n> corrected committer date:\n>\n>   if (max_generation + 1 <= graph_commit->date) {\n>   \t/* commit date does not need correction */\n>   \tif (generation != graph_commit->date)\n>     \tgraph_report(_(\"commit-graph corrected commit date for commit %s \"\n>   \t\t               \"is %\"PRItime\" != %\"PRItime\" commit date\"),\n>                      ...);\n>   } else {\n>   \tif (generation != max_generation + 1)\n>   \t\tgraph_report(_(\"commit-graph generation v2 for commit %s is %\"PRItime\" != %\"PRItime),\n>                      ...);\n>   }\n>\n> Though I think it might be overkill.\n>\n> Advantages:\n> - more strict tests, checking generation numbers (v2 if present, v1\n>   otherwise) against their definition\n> - if there is no GDAT chunk, verify works just like it did before\n>\n> Disadvantages:\n> - more complicated code\n> - possibly measurable performance degradation due to extra conditional\n>\n>\n> III. Like II., but if there is generation numbers chunk (GDAT chunk), we\n> verify *both* topological levels (v1) and corrected commit date (v2)\n> against their definition.  If GDAT chunk is not present, it reduces to\n> current code (before this patch series).\n>\n> Advantages:\n> - if there is no GDAT chunk, verify works just like it did before\n> - most strict tests, verifying all the data: both generation number v1\n>   and generation number v2 -- if possible\n>\n> Disadvantages:\n> - most complex code; we need to somehow extract topological levels\n>   if the GDAT chunk is present (they are not on graph data slab in this\n>   case); I have not even started to think how it could be done\n> - slower verification\n>\n>> Suprised that no test breaks by this change.\n> I don't whink we have any test that created commit graph with\n> topological levels greater than GENERATION_NUMBER_V1_MAX; this would be\n> expensive and have to be of course protected by GIT_TEST_LONG aka\n> EXPENSIVE prerequisite.\n>\n>   # GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 is here to force verification of topological levels\n>   test_expect_success EXPENSIVE 'verify handles topological levels > GENERATION_NUMBER_V1_MAX' '\n>   \trm -rf long_chain &&\n>   \tgit init long_chain &&\n>   \ttest_commit_bulk -C long_chain 1073741824 &&\n>     (\n>   \t\tcd long_chain &&\n>   \t\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write &&\n>   \t\tgit commit-graph verify\n>     )\n>   '\n>\n> This however lies slightly outside the scope of this patch series,\n> though if you could add this test (in a separate patch), after testing\n> it, it would be very nice.\n>\n>> I have also moved changes in the verify function to the next patch, as\n>> we cannot write or read corrected commit dates yet - so little sense in\n>> modifying verify.\n> I think putting changes to the verify function in a separate patch, be\n> it before or after this one (depending on the choice of the algorithm\n> for verification, see above) would be a good idea.\n>\n>>> Sidenote: I think we don't have to worry about having to introduce\n>>> GENERATION_NUMBER_V2_MAX, as the in-memory size (of reconstructed from\n>>> disck representation) corrected commiter date is the same as of commiter\n>>> date itself, plus some, and I don't see us coming close to 64-bit limit\n>>> of timestamp_t for commit dates.\n>>>\n>>>>  \t\t\t\t     oid_to_hex(&cur_oid),\n>>>>  \t\t\t\t     generation,\n>>>>  \t\t\t\t     max_generation + 1);\n> Best,\n\n"},{"id":"409205","messageId":"xmqqtuu3k6jf.fsf@gitster.c.googlers.com","threadId":"53933","inReplyTo":"efa3488a-3983-3435-e5e4-2eb71e76a33a@iee.email","subject":"Re: [PATCH v4 06/10] commit-graph: implement corrected commit date","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-11-05T18:22:28Z","receivedAt":"2020-11-05T18:22:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Philip Oakley <philipoakley@iee.email> writes:\n\n> This may be not part of the the main project, but could you consider, if\n> time permits, also adding some entries into the Git Glossary (`git help\n> glossary`) for the various terms we are using here and elsewhere, e.g.\n> 'topological levels', 'generation number', 'corrected commit date' (and\n> its fancy technical name for the use of date heuristics e.g. the\n> 'chronological ordering';).\n>\n> The glossary can provide a reference, once the issues are resolved. The\n> History Simplification and Commit Ordering section of git-log maybe a\n> useful guide to some of the terms that would link to the glossary.\n\nAh, I first thought that Documentation/rev-list-options.txt (which\nis the relevant part of \"git log\" documentation you mention here)\nalready have references to deep technical terms explained in the\nglossary and you are suggesting Abhishek to mimic the arrangement by\nadding new and agreed-upon terms to the glossary and referring to\nthem from the commit-graph documentation updated by this series.\n\nBut sadly that is not the case.  What you are saying is that you\nnoticed that rev-list-options.txt needs a similar \"the terms we use\nto explain these two sections should be defined and explained in the\nglossary (if they are not) and new references to glossary should be\nadded there\" update.\n\nIn any case, that is a very good suggestion.  I agree that updating\n\"git log\" doc may be outside the scope of Abhishek's theme, but it\nwould be very good to have such an update by anybody ;-)\n\nThanks\n"},{"id":"409271","messageId":"20201106112513.GA1553@Abhishek-Arch","threadId":"53933","inReplyTo":"854kmbx4pi.fsf@gmail.com","subject":"Re: [PATCH v4 07/10] commit-graph: implement generation data chunk","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-11-06T11:25:13Z","receivedAt":"2020-11-06T11:28:12Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Fri, Oct 30, 2020 at 01:45:29PM +0100, Jakub Narębski wrote:\n> \n> ...\n>\n> >\n> > While storing corrected commit date offset instead of the corrected\n> > commit date saves us 4 bytes per commit, it's possible for the offsets\n> > to overflow the 4-bytes allocated. As such overflows are exceedingly\n> > rare, we use the following overflow management scheme:\n> \n> Perhaps it would be good idea to write the idea in full from start, as\n> the commit message is intended to be read stadalone and not in the\n> context of the patch series.  On the other hand it might be too much\n> detail in already [necessarily] lengthty commit message.\n> \n> Perhaps something like the following proposal would read better.\n> \n>   To minimize the space required to store corrected commit date, Git\n>   stores corrected commit date offsets into the commit-graph file,\n>   instead of corrected commit dates themselves. This saves us 4 bytes\n>   per commit, decreasing the GDAT chunk size by half, but it's possible\n>   for the offset to overflow the 4-bytes allocated for storage. As such\n>   overflows are and should be exceedingly rare, we use the following\n>   overflow management scheme:\n>\n\nThanks, that's better.\n\n> \n> ...\n>\n> > We test the overflow-related code with the following repo history:\n> >\n> >            F - N - U\n> >           /         \\\n> > U - N - U            N\n> >          \\          /\n> >            N - F - N\n> \n> Do we need such complex history? I guess we need to test the handling of\n> merge commits too.\n> \n\nI wanted to test three cases - a root epoch zero commit, a commit that's\nfar enough in past to overflow the offset and a commit that's far enough\nin the future to overflow the offset.\n\n> >\n> > Where the commits denoted by U have committer date of zero seconds\n> > since Unix epoch, the commits denoted by N have committer date of\n> > 1112354055 (default committer date for the test suite) seconds since\n> > Unix epoch and the commits denoted by F have committer date of\n> > (2 ^ 31 - 2) seconds since Unix epoch.\n> >\n> > The largest offset observed is 2 ^ 31, just large enough to overflow.\n> >\n> > [1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n> >\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c                | 98 +++++++++++++++++++++++++++++++++--\n> >  commit-graph.h                |  3 ++\n> >  commit.h                      |  1 +\n> >  t/README                      |  3 ++\n> >  t/helper/test-read-graph.c    |  4 ++\n> >  t/t4216-log-bloom.sh          |  4 +-\n> >  t/t5318-commit-graph.sh       | 70 ++++++++++++++++++++-----\n> >  t/t5324-split-commit-graph.sh | 12 ++---\n> >  t/t6600-test-reach.sh         | 68 +++++++++++++-----------\n> >  9 files changed, 206 insertions(+), 57 deletions(-)\n> >\n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index 03948adfce..71d0b243db 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -38,11 +38,13 @@ void git_test_write_commit_graph_or_die(void)\n> >  #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n> >  #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n> >  #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n> > +#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n> > +#define GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW 0x47444f56 /* \"GDOV\" */\n> >  #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n> >  #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n> >  #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n> >  #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n> > -#define MAX_NUM_CHUNKS 7\n> > +#define MAX_NUM_CHUNKS 9\n> \n> All right.\n> \n> >  \n> >  #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n> >  \n> > @@ -61,6 +63,8 @@ void git_test_write_commit_graph_or_die(void)\n> >  #define GRAPH_MIN_SIZE (GRAPH_HEADER_SIZE + 4 * GRAPH_CHUNKLOOKUP_WIDTH \\\n> >  \t\t\t+ GRAPH_FANOUT_SIZE + the_hash_algo->rawsz)\n> >  \n> > +#define CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW (1ULL << 31)\n> > +\n> \n> All right, though the naming convention is different from the one used\n> for EDGE chunk: GRAPH_EXTRA_EDGES_NEEDED and GRAPH_EDGE_LAST_MASK.\n> \n> >  /* Remember to update object flag allocation in object.h */\n> >  #define REACHABLE       (1u<<15)\n> >  \n> > @@ -385,6 +389,20 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n> >  \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n> >  \t\t\tbreak;\n> >  \n> > +\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n> > +\t\t\tif (graph->chunk_generation_data)\n> > +\t\t\t\tchunk_repeated = 1;\n> > +\t\t\telse\n> > +\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n> > +\t\t\tbreak;\n> > +\n> > +\t\tcase GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW:\n> > +\t\t\tif (graph->chunk_generation_data_overflow)\n> > +\t\t\t\tchunk_repeated = 1;\n> > +\t\t\telse\n> > +\t\t\t\tgraph->chunk_generation_data_overflow = data + chunk_offset;\n> > +\t\t\tbreak;\n> > +\n> \n> Necessary but unavoidable boilerplate for adding new chunks to the\n> commit-graph file format.  All right.\n> \n> >  \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n> >  \t\t\tif (graph->chunk_extra_edges)\n> >  \t\t\t\tchunk_repeated = 1;\n> > @@ -745,8 +763,8 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n> >  {\n> >  \tconst unsigned char *commit_data;\n> >  \tstruct commit_graph_data *graph_data;\n> > -\tuint32_t lex_index;\n> > -\tuint64_t date_high, date_low;\n> > +\tuint32_t lex_index, offset_pos;\n> > +\tuint64_t date_high, date_low, offset;\n> \n> All right, we are adding two new variables: `offset` to read data stored\n> in GDAT chunk, and `offset_pos` to help read data from GDOV chunk if\n> necessary i.e. to handle overflow in corrected commit data offset\n> storage.\n> \n> >  \n> >  \twhile (pos < g->num_commits_in_base)\n> >  \t\tg = g->base_graph;\n> > @@ -764,7 +782,16 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n> >  \tdate_low = get_be32(commit_data + g->hash_len + 12);\n> >  \titem->date = (timestamp_t)((date_high << 32) | date_low);\n> >  \n> > -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> > +\tif (g->chunk_generation_data) {\n> > +\t\toffset = (timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n> \n> Style: why space after the `(timestamp_t)` cast operator?\n> \n> Though CodingGuidelines do not say anything on this topic... perhaps the\n> space after cast operator makes it more readable?\n> \n> > +\n> > +\t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n> \n> All right, so the CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW is equivalent of\n> GRAPH_EXTRA_EDGES_NEEDED.\n> \n> > +\t\t\toffset_pos = offset ^ CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW;\n> \n> Hmmm... instead of using bitwise and on an equivalent to the\n> GRAPH_EDGE_LAST_MASK, we utilize the fact that we know that the MSB bit\n> is set, so we can clear it with bitwise xor.  Clever trick.\n>\n> \n> > +\t\t\tgraph_data->generation = get_be64(g->chunk_generation_data_overflow + 8 * offset_pos);\n> > +\t\t} else\n> > +\t\t\tgraph_data->generation = item->date + offset;\n> \n> All right, this handles the case when we have generation number v2, with\n> or without overflow.\n> \n> > +\t} else\n> > +\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n> \n> All right, this handles the case where we have only generation number\n> v1, like for commit-graph file written by old Git.\n> \n> >  \n> >  \tif (g->topo_levels)\n> >  \t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n> > @@ -942,6 +969,7 @@ struct write_commit_graph_context {\n> >  \tstruct packed_oid_list oids;\n> >  \tstruct packed_commit_list commits;\n> >  \tint num_extra_edges;\n> > +\tint num_generation_data_overflows;\n> >  \tunsigned long approx_nr_objects;\n> >  \tstruct progress *progress;\n> >  \tint progress_done;\n> > @@ -960,7 +988,8 @@ struct write_commit_graph_context {\n> >  \t\t report_progress:1,\n> >  \t\t split:1,\n> >  \t\t changed_paths:1,\n> > -\t\t order_by_pack:1;\n> > +\t\t order_by_pack:1,\n> > +\t\t write_generation_data:1;\n> >  \n> >  \tstruct topo_level_slab *topo_levels;\n> >  \tconst struct commit_graph_opts *opts;\n> \n> All right, this adds necessary fields to `struct write_commit_graph_context`.\n> \n> > @@ -1120,6 +1149,44 @@ static int write_graph_chunk_data(struct hashfile *f,\n> >  \treturn 0;\n> >  }\n> >  \n> > +static int write_graph_chunk_generation_data(struct hashfile *f,\n> > +\t\t\t\t\t      struct write_commit_graph_context *ctx)\n> > +{\n> > +\tint i, num_generation_data_overflows = 0;\n> \n> Minor nitpick: in my opinion there should be empty line here, between\n> the variables declaration and the code... however not all\n> write_graph_chunk_*() functions have it.\n> \n> > +\tfor (i = 0; i < ctx->commits.nr; i++) {\n> > +\t\tstruct commit *c = ctx->commits.list[i];\n> > +\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n> > +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n> \n> All right.\n> \n> > +\n> > +\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n> > +\t\t\toffset = CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW | num_generation_data_overflows;\n> > +\t\t\tnum_generation_data_overflows++;\n> > +\t\t}\n> \n> Hmmm... shouldn't we store these commits that need overflow handling\n> (with corrected commit date offset greater than GENERATION_NUMBER_V2_OFFSET_MAX)\n> in a list or a queue, to remember them for writing GDOV chunk?\n> \n\nWe could, although write_graph_chunk_extra_edges() (just like this function)\nprefers to iterate over all commits again. Both octopus merges and\noverflowing corrected commit dates are exceedingly rare, might be\nworthwhile to trade some memory to avoid looping again.\n\n> We could store oids, or we could store commits themselves, for example:\n> \n> \t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n> \t\t\toffset = CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW | num_generation_data_overflows;\n> \t\t\tnum_generation_data_overflows++;\n> \n> \t\t\tALLOC_GROW(ctx->gdov_commits.list, ctx->gdov_commits.nr + 1, ctx->gdov_commits.alloc);\n> \t\t\tctx->commits.list[ctx->gdov_commits.nr] = c;\n>             ctx->gdov_commits.nr++;\n> \t\t}\n> \n> Though in the above proposal we could get rid of `num_generation_data_overflows`, \n> as it should be the same as `ctx->gdov_commits.nr`.\n> \n> I have called the extra commit list member of write_commit_graph_context\n> `gdov_commits`, but perhaps a better name would be `commits_gen_v2_overflow`, \n> or similar more descriptive name.\n> \n> > +\n> > +\t\thashwrite_be32(f, offset);\n> > +\t}\n> > +\n> > +\treturn 0;\n> > +}\n> \n> All right.\n> \n> > +\n> > +static int write_graph_chunk_generation_data_overflow(struct hashfile *f,\n> > +\t\t\t\t\t\t       struct write_commit_graph_context *ctx)\n> > +{\n> > +\tint i;\n> > +\tfor (i = 0; i < ctx->commits.nr; i++) {\n> \n> Here we loop over *all* commits again, instead of looping over those\n> very rare commits that need overflow handling for their corrected commit\n> date data.\n> \n> Though this possible performance issue^* could be fixed in the future commit.\n> \n> *) It needs to be actually benchmarked which version is faster.\n> \n> ...\n> \n> >  \n> >  graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n> > @@ -454,8 +454,9 @@ test_expect_success 'warn on improper hash version' '\n> >  \n> >  test_expect_success 'git commit-graph verify' '\n> >  \tcd \"$TRASH_DIRECTORY/full\" &&\n> > -\tgit rev-parse commits/8 | git commit-graph write --stdin-commits &&\n> > -\tgit commit-graph verify >output\n> > +\tgit rev-parse commits/8 | GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --stdin-commits &&\n> > +\tgit commit-graph verify >output &&\n> \n> All right, this simply adds GIT_TEST_COMMIT_GRAPH_NO_GDAT=1.  I assume\n> this is needed because this test is also setup for the following commits\n> _without_ even saying that in the test name (bad practice, in my\n> opinion), and the comment above this test says the following:\n> \n>   # the verify tests below expect the commit-graph to contain\n>   # exactly the commits reachable from the commits/8 branch.\n>   # If the file changes the set of commits in the list, then the\n>   # offsets into the binary file will result in different edits\n>   # and the tests will likely break.\n> \n> So the following tests are fragile (though perhaps unavoidably fragile),\n> and without this change they would not work, I assume.\n> \n> > +\tgraph_read_expect 9 extra_edges\n> \n> I guess that this is here to check that GIT_TEST_COMMIT_GRAPH_NO_GDAT=1\n> work as intended, and that the following \"verify\" tests wouldn't break.\n> I understand its necessity, even if I don't quite like having a test\n> that checks multiple things.  This is a minor issue, though.\n> \n> All right.\n> \n> \n> We might want to have a separate test that checks that we get\n> commit-graph with and without GDAT chunk depending on whether we use\n> GIT_TEST_COMMIT_GRAPH_NO_GDAT=1.  On the other hand, this environment\n> variable is there purely for tests, so the question is should we test\n> the test infrastructure?\n> \n\n\n> >  '\n> >  \n> >  NUM_COMMITS=9\n> > @@ -741,4 +742,47 @@ test_expect_success 'corrupt commit-graph write (missing tree)' '\n> >  \t)\n> >  '\n> >  \n> > +test_commit_with_date() {\n> > +  file=\"$1.t\" &&\n> > +  echo \"$1\" >\"$file\" &&\n> > +  git add \"$file\" &&\n> > +  GIT_COMMITTER_DATE=\"$2\" GIT_AUTHOR_DATE=\"$2\" git commit -m \"$1\"\n> > +  git tag \"$1\"\n> > +}\n> \n> Here we add a helper function.  All right.\n> \n> I wonder though if it wouldn't be a better idea to add `--date <date>`\n> option to the test_commit() function in test-lib-functions.sh (which\n> option would set GIT_COMMITTER_DATE and GIT_AUTHOR_DATE, and also\n> set notick=yes).\n> \n\nYes, that's a better idea - I didn't know how to change test_commit()\nwell enough to tinker with what's working.\n\n> For example:\n> \n> diff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\n> index f1ae935fee..a1f9a2b09b 100644\n> --- a/t/test-lib-functions.sh\n> +++ b/t/test-lib-functions.sh\n> @@ -202,6 +202,12 @@ test_commit () {\n>  \t\t--signoff)\n>  \t\t\tsignoff=\"$1\"\n>  \t\t\t;;\n> +        --date)\n> +            notick=yes\n> +            GIT_COMMITTER_DATE=\"$2\"\n> +            GIT_AUTHOR_DATE=\"$2\"\n> +            shift\n> +            ;;\n>  \t\t-C)\n>  \t\t\tindir=\"$2\"\n>  \t\t\tshift\n> \n> \n> > +\n> \n> It would be nice to have there comment describing the shape of the\n> revision history we generate here, that currenly is present only in the\n> commmit message.\n> \n> # We test the overflow-related code with the following repo history:\n> #\n> #               4:F - 5:N - 6:U\n> #              /               \\\n> # 1:U - 2:N - 3:U               M:N\n> #              \\               /\n> #               7:N - 8:F - 9:N\n> #\n> # Here the commits denoted by U have committer date of zero seconds\n> # since Unix epoch, the commits denoted by N have committer date\n> # starting from 1112354055 seconds since Unix epoch (default committer\n> # date for the test suite), and the commits denoted by F have committer\n> # date of (2 ^ 31 - 2) seconds since Unix epoch.\n> #\n> # The largest offset observed is 2 ^ 31, just large enough to overflow.\n> #\n\nYes, it would. Added.\n> \n> > +test_expect_success 'overflow corrected commit date offset' '\n> > +\tobjdir=\".git/objects\" &&\n> > +\tUNIX_EPOCH_ZERO=\"1970-01-01 00:00 +0000\" &&\n> > +\tFUTURE_DATE=\"@2147483646 +0000\" &&\n> \n> It is a bit funny to see UNIX_EPOCH_ZERO spelled one way, and\n> FUTURE_DATE other way.\n> \n> Wouldn't be more readable to use UNIX_EPOCH_ZERO=\"@0 +0000\"?\n\nIt would, for some reason - I couldn't figure out the valid format for\nthis. Changed.\n\n> \n> > +\ttest_oid_cache <<-EOF &&\n> > +\toid_version sha1:1\n> > +\toid_version sha256:2\n> > +\tEOF\n> > +\tcd \"$TRASH_DIRECTORY\" &&\n> > +\tmkdir repo &&\n> > +\tcd repo &&\n> > +\tgit init &&\n> > +\ttest_commit_with_date 1 \"$UNIX_EPOCH_ZERO\" &&\n> > +\ttest_commit 2 &&\n> > +\ttest_commit_with_date 3 \"$UNIX_EPOCH_ZERO\" &&\n> > +\tgit commit-graph write --reachable &&\n> > +\tgraph_read_expect 3 generation_data &&\n> > +\ttest_commit_with_date 4 \"$FUTURE_DATE\" &&\n> > +\ttest_commit 5 &&\n> > +\ttest_commit_with_date 6 \"$UNIX_EPOCH_ZERO\" &&\n> > +\tgit branch left &&\n> > +\tgit reset --hard 3 &&\n> > +\ttest_commit 7 &&\n> > +\ttest_commit_with_date 8 \"$FUTURE_DATE\" &&\n> > +\ttest_commit 9 &&\n> > +\tgit branch right &&\n> > +\tgit reset --hard 3 &&\n> > +\tgit merge left right &&\n> \n> We have test_merge() function in test-lib-functions.sh, perhaps we\n> should use it here.\n> \n> > +\tgit commit-graph write --reachable &&\n> > +\tgraph_read_expect 10 \"generation_data generation_data_overflow\" &&\n> \n> All right, we write the commit-graph and check that it has both GDAT and\n> GDOV chunks present.\n> \n> > +\tgit commit-graph verify\n> \n> All right, we checks that created commit graph with GDAT and GDOV passes\n> 'git commit-graph verify` checks.\n> \n> > +'\n> > +\n> > +graph_git_behavior 'overflow corrected commit date offset' repo left right\n> \n> All right, here we compare the Git behavior with the commit-graph to the\n> behavior without it... however I think that those two tests really\n> should have distinct (different) test names. Currently they both use\n> 'overflow corrected commit date offset'.\n> \n\nFollowing the earlier tests, the first test could be \"set up and verify\nrepo with generation data overflow chunk\" and the git behavior test can\nbe \"generation data overflow chunk repo\"\n\n> > +\n> >  test_done\n> > diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> > index c334ee9155..651df89ab2 100755\n> > --- a/t/t5324-split-commit-graph.sh\n> > +++ b/t/t5324-split-commit-graph.sh\n> > @@ -13,11 +13,11 @@ test_expect_success 'setup repo' '\n> >  \tinfodir=\".git/objects/info\" &&\n> >  \tgraphdir=\"$infodir/commit-graphs\" &&\n> >  \ttest_oid_cache <<-EOM\n> > -\tshallow sha1:1760\n> > -\tshallow sha256:2064\n> > +\tshallow sha1:2132\n> > +\tshallow sha256:2436\n> >  \n> > -\tbase sha1:1376\n> > -\tbase sha256:1496\n> > +\tbase sha1:1408\n> > +\tbase sha256:1528\n> >  \n> >  \toid_version sha1:1\n> >  \toid_version sha256:2\n> > @@ -31,9 +31,9 @@ graph_read_expect() {\n> >  \t\tNUM_BASE=$2\n> >  \tfi\n> >  \tcat >expect <<- EOF\n> > -\theader: 43475048 1 $(test_oid oid_version) 3 $NUM_BASE\n> > +\theader: 43475048 1 $(test_oid oid_version) 4 $NUM_BASE\n> >  \tnum_commits: $1\n> > -\tchunks: oid_fanout oid_lookup commit_metadata\n> > +\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n> >  \tEOF\n> >  \ttest-tool read-graph >output &&\n> >  \ttest_cmp expect output\n> \n> All right, we now expect the commit graph to include the GDAT chunk...\n> though shouldn't be there old expected value for no GDAT, for future\n> tests?  But perhaps this is not necessary.\n> \n> Note that I have not checked the details, but it looks OK to me.\n> \n> > diff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\n> > index f807276337..e2d33a8a4c 100755\n> > --- a/t/t6600-test-reach.sh\n> > +++ b/t/t6600-test-reach.sh\n> > @@ -55,10 +55,13 @@ test_expect_success 'setup' '\n> >  \tgit show-ref -s commit-5-5 | git commit-graph write --stdin-commits &&\n> >  \tmv .git/objects/info/commit-graph commit-graph-half &&\n> >  \tchmod u+w commit-graph-half &&\n> > +\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable &&\n> > +\tmv .git/objects/info/commit-graph commit-graph-no-gdat &&\n> > +\tchmod u+w commit-graph-no-gdat &&\n> \n> All right, this prepares for testing one more mode.  The run_all_modes()\n> function would test the following cases:\n>  - no commit-graph\n>  - commit-graph for all commits, with GDAT\n>  - commit-graph with half of commits, with GDAT\n>  - commit-graph for all commits, without GDAT\n> \n> >  \tgit config core.commitGraph true\n> >  '\n> >  \n> > -run_three_modes () {\n> > +run_all_modes () {\n> >  \ttest_when_finished rm -rf .git/objects/info/commit-graph &&\n> >  \t\"$@\" <input >actual &&\n> >  \ttest_cmp expect actual &&\n> > @@ -67,11 +70,14 @@ run_three_modes () {\n> >  \ttest_cmp expect actual &&\n> >  \tcp commit-graph-half .git/objects/info/commit-graph &&\n> >  \t\"$@\" <input >actual &&\n> > +\ttest_cmp expect actual &&\n> > +\tcp commit-graph-no-gdat .git/objects/info/commit-graph &&\n> > +\t\"$@\" <input >actual &&\n> >  \ttest_cmp expect actual\n> >  }\n> >  \n> > -test_three_modes () {\n> > -\trun_three_modes test-tool reach \"$@\"\n> > +test_all_modes () {\n> > +\trun_all_modes test-tool reach \"$@\"\n> >  }\n> \n> All right.\n> \n> Though to reduce \"noise\" in this patch, the rename of run_three_modes()\n> to run_all_modes() and test_three_modes() to test_all_modes() could have\n> been done in a separate preparatory patch.  It would be pure refactoring\n> patch, without introducing any new functionality.\n> \n\nSure, that makes sense to me - this is patch is over 200 lines long\nalready.\n\n> ...\n\nThanks\n- Abhishek\n"},{"id":"409281","messageId":"858sbel67n.fsf@gmail.com","threadId":"53933","inReplyTo":"20201106112513.GA1553@Abhishek-Arch","subject":"Re: [PATCH v4 07/10] commit-graph: implement generation data chunk","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-11-06T17:56:28Z","receivedAt":"2020-11-06T17:56:34Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"In short: I think that because current implementation of writing GDOV\nchunk follows an example of writing EDGE chunk, it should be left as it\nis now (simple), and posible performance improvements be postponed to\nsome future commit.\n\nAbhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Fri, Oct 30, 2020 at 01:45:29PM +0100, Jakub Narębski wrote:\n[...]\n>>> We test the overflow-related code with the following repo history:\n>>>\n>>>            F - N - U\n>>>           /         \\\n>>> U - N - U            N\n>>>          \\          /\n>>>            N - F - N\n>> \n>> Do we need such complex history? I guess we need to test the handling of\n>> merge commits too.\n>> \n>\n> I wanted to test three cases - a root epoch zero commit, a commit that's\n> far enough in past to overflow the offset and a commit that's far enough\n> in the future to overflow the offset.\n\nAll right, if I understand this correctly this would be U as root, U-F\npair of commits and N-F pair of commits, respectively.  Did I get it\nright?\n\nAnyway, it might be a good idea to put this explanation in the commit\nmessage.\n\n>>>\n>>> Where the commits denoted by U have committer date of zero seconds\n>>> since Unix epoch, the commits denoted by N have committer date of\n>>> 1112354055 (default committer date for the test suite) seconds since\n>>> Unix epoch and the commits denoted by F have committer date of\n>>> (2 ^ 31 - 2) seconds since Unix epoch.\n>>>\n>>> The largest offset observed is 2 ^ 31, just large enough to overflow.\n\n[...]\n>>> @@ -1120,6 +1149,44 @@ static int write_graph_chunk_data(struct hashfile *f,\n>>>  \treturn 0;\n>>>  }\n>>>  \n>>> +static int write_graph_chunk_generation_data(struct hashfile *f,\n>>> +\t\t\t\t\t      struct write_commit_graph_context *ctx)\n>>> +{\n>>> +\tint i, num_generation_data_overflows = 0;\n>> \n>> Minor nitpick: in my opinion there should be empty line here, between\n>> the variables declaration and the code... however not all\n>> write_graph_chunk_*() functions have it.\n>> \n>>> +\tfor (i = 0; i < ctx->commits.nr; i++) {\n>>> +\t\tstruct commit *c = ctx->commits.list[i];\n>>> +\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n>>> +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n>> \n>> All right.\n>> \n>>> +\n>>> +\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n>>> +\t\t\toffset = CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW | num_generation_data_overflows;\n>>> +\t\t\tnum_generation_data_overflows++;\n>>> +\t\t}\n>> \n>> Hmmm... shouldn't we store these commits that need overflow handling\n>> (with corrected commit date offset greater than GENERATION_NUMBER_V2_OFFSET_MAX)\n>> in a list or a queue, to remember them for writing GDOV chunk?\n>> \n>\n> We could, although write_graph_chunk_extra_edges() (just like this function)\n> prefers to iterate over all commits again. Both octopus merges and\n> overflowing corrected commit dates are exceedingly rare, might be\n> worthwhile to trade some memory to avoid looping again.\n\nI'm sorry, I have not looked what write_graph_chunk_extra_edges() does,\nor rather how it does what it does -- it is a good idea to pattern your\nsolution in similar existing code.\n\nFor me this is an even stronger hint that we should strive for\nsimplicity first, and leave possible performance improvements for the\nfuture commit.  Especially that you perform the most significant\noptimization for this overflow handling: ensuring that we do not perform\nany work if there are no commits with generation data overflow.\n\nMaybe, maybe we should add that information about similarity between\nwrite_graph_chunk_generation_data_overflow() and write_graph_chunk_extra_edges() \nin the commit message.  I am unsure...\n\n>> We could store oids, or we could store commits themselves, for example:\n>> \n>> \t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n>> \t\t\toffset = CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW | num_generation_data_overflows;\n>> \t\t\tnum_generation_data_overflows++;\n>> \n>> \t\t\tALLOC_GROW(ctx->gdov_commits.list, ctx->gdov_commits.nr + 1, ctx->gdov_commits.alloc);\n>> \t\t\tctx->commits.list[ctx->gdov_commits.nr] = c;\n>>             ctx->gdov_commits.nr++;\n>> \t\t}\n>> \n>> Though in the above proposal we could get rid of `num_generation_data_overflows`, \n>> as it should be the same as `ctx->gdov_commits.nr`.\n>> \n>> I have called the extra commit list member of write_commit_graph_context\n>> `gdov_commits`, but perhaps a better name would be `commits_gen_v2_overflow`, \n>> or similar more descriptive name.\n\n[...]\n>>> @@ -741,4 +742,47 @@ test_expect_success 'corrupt commit-graph write (missing tree)' '\n>>>  \t)\n>>>  '\n>>>  \n>>> +test_commit_with_date() {\n>>> +  file=\"$1.t\" &&\n>>> +  echo \"$1\" >\"$file\" &&\n>>> +  git add \"$file\" &&\n>>> +  GIT_COMMITTER_DATE=\"$2\" GIT_AUTHOR_DATE=\"$2\" git commit -m \"$1\"\n>>> +  git tag \"$1\"\n>>> +}\n>> \n>> Here we add a helper function.  All right.\n>> \n>> I wonder though if it wouldn't be a better idea to add `--date <date>`\n>> option to the test_commit() function in test-lib-functions.sh (which\n>> option would set GIT_COMMITTER_DATE and GIT_AUTHOR_DATE, and also\n>> set notick=yes).\n>> \n>\n> Yes, that's a better idea - I didn't know how to change test_commit()\n> well enough to tinker with what's working.\n>\n>> For example:\n>> \n>> diff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\n>> index f1ae935fee..a1f9a2b09b 100644\n>> --- a/t/test-lib-functions.sh\n>> +++ b/t/test-lib-functions.sh\n>> @@ -202,6 +202,12 @@ test_commit () {\n>>  \t\t--signoff)\n>>  \t\t\tsignoff=\"$1\"\n>>  \t\t\t;;\n>> +        --date)\n>> +            notick=yes\n>> +            GIT_COMMITTER_DATE=\"$2\"\n>> +            GIT_AUTHOR_DATE=\"$2\"\n>> +            shift\n>> +            ;;\n>>  \t\t-C)\n>>  \t\t\tindir=\"$2\"\n>>  \t\t\tshift\n\nNote however that I have while I have followed example of other options\n(namely '-C <directory>'), I have not actually tested this proposed\nimplementation in tests; I have just tested that it looks like it works\nOK.\n\n[...]\n>>> +test_expect_success 'overflow corrected commit date offset' '\n>>> +\tobjdir=\".git/objects\" &&\n>>> +\tUNIX_EPOCH_ZERO=\"1970-01-01 00:00 +0000\" &&\n>>> +\tFUTURE_DATE=\"@2147483646 +0000\" &&\n>> \n>> It is a bit funny to see UNIX_EPOCH_ZERO spelled one way, and\n>> FUTURE_DATE other way.\n>> \n>> Wouldn't be more readable to use UNIX_EPOCH_ZERO=\"@0 +0000\"?\n>\n> It would, for some reason - I couldn't figure out the valid format for\n> this. Changed.\n\nWell, if \"@2147483646 +0000\" works (i.e. \"@<Unix epoch/timestamp> <offset>\"),\nwhy the same for timestamp 0, i.e. \"@0 +0000\", wouldn't work?\n\n[...]\n>>> +graph_git_behavior 'overflow corrected commit date offset' repo left right\n>> \n>> All right, here we compare the Git behavior with the commit-graph to the\n>> behavior without it... however I think that those two tests really\n>> should have distinct (different) test names. Currently they both use\n>> 'overflow corrected commit date offset'.\n>> \n>\n> Following the earlier tests, the first test could be \"set up and verify\n> repo with generation data overflow chunk\" and the git behavior test can\n> be \"generation data overflow chunk repo\"\n\nFirst is OK, the second could possibly be improved but is all right.\n\n[...]\n>> Though to reduce \"noise\" in this patch, the rename of run_three_modes()\n>> to run_all_modes() and test_three_modes() to test_all_modes() could have\n>> been done in a separate preparatory patch.  It would be pure refactoring\n>> patch, without introducing any new functionality.\n>> \n>\n> Sure, that makes sense to me - this is patch is over 200 lines long\n> already.\n\nThanks in advance.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"409284","messageId":"85zh3ujq9c.fsf_-_@gmail.com","threadId":"53933","inReplyTo":"xmqqtuu3k6jf.fsf@gitster.c.googlers.com","subject":"Extending and updating gitglossary (was: Re: [PATCH v4 06/10] commit-graph: implement corrected commit date)","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-11-06T18:26:23Z","receivedAt":"2020-11-06T18:26:28Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n> Philip Oakley <philipoakley@iee.email> writes:\n>\n>> This may be not part of the the main project, but could you consider, if\n>> time permits, also adding some entries into the Git Glossary (`git help\n>> glossary`) for the various terms we are using here and elsewhere, e.g.\n>> 'topological levels', 'generation number', 'corrected commit date' (and\n>> its fancy technical name for the use of date heuristics e.g. the\n>> 'chronological ordering';).\n>>\n>> The glossary can provide a reference, once the issues are resolved. The\n>> History Simplification and Commit Ordering section of git-log maybe a\n>> useful guide to some of the terms that would link to the glossary.\n>\n> Ah, I first thought that Documentation/rev-list-options.txt (which\n> is the relevant part of \"git log\" documentation you mention here)\n> already have references to deep technical terms explained in the\n> glossary and you are suggesting Abhishek to mimic the arrangement by\n> adding new and agreed-upon terms to the glossary and referring to\n> them from the commit-graph documentation updated by this series.\n>\n> But sadly that is not the case.  What you are saying is that you\n> noticed that rev-list-options.txt needs a similar \"the terms we use\n> to explain these two sections should be defined and explained in the\n> glossary (if they are not) and new references to glossary should be\n> added there\" update.\n>\n> In any case, that is a very good suggestion.  I agree that updating\n> \"git log\" doc may be outside the scope of Abhishek's theme, but it\n> would be very good to have such an update by anybody ;-)\n\nThe only possible problem I see with this suggestion is that some of\nthose terms (like 'topological levels' and 'corrected commit date') are\ntechnical terms that should be not of concern for Git user, only for\ndevelopers working on Git.  (However one could encounter the term\n\"generation number\" in `git commit-graph verify` output.)\n\nI don't think adding technical terms that the user won't encounter in\nthe documentation or among messages that Git outputs would be not a good\nidea.  It could confuse users, rather than help them.\n\nConversely, perhaps we should add Documentation/technical/glossary.txt\nto help developers.\n\n\nP.S. By the way, when looking at Documentation/glossary-content.txt, I\nhave noticed few obsolescent entries, like \"Git archive\", few that have\ndescription that soon could be or is obsolete and would need updating,\nlike \"master\" (when default branch switch to \"main\"), or \"object\nidentifier\" and \"SHA-1\" (when Git switches away from SHA-1 as hash\nfunction).\n\nBest,\n--\nJakub Narębski\n"},{"id":"409290","messageId":"xmqqa6vufffk.fsf@gitster.c.googlers.com","threadId":"53933","inReplyTo":"85zh3ujq9c.fsf_-_@gmail.com","subject":"Re: Extending and updating gitglossary","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-11-06T19:33:51Z","receivedAt":"2020-11-06T19:34:01Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narębski <jnareb@gmail.com> writes:\n\n> I don't think adding technical terms that the user won't encounter in\n> the documentation or among messages that Git outputs would be not a good\n> idea.  It could confuse users, rather than help them.\n>\n> Conversely, perhaps we should add Documentation/technical/glossary.txt\n> to help developers.\n\nThanks for a thoughtful suggestion to help the target audience.  I\nagree 100% with the above two paragraphs.\n"},{"id":"409335","messageId":"8d43335d-a0b4-511e-f132-057343234503@iee.email","threadId":"53933","inReplyTo":"85zh3ujq9c.fsf_-_@gmail.com","subject":"Re: Extending and updating gitglossary (was: Re: [PATCH v4 06/10] commit-graph: implement corrected commit date)","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2020-11-08T17:23:28Z","receivedAt":"2020-11-08T17:23:36Z","isPatch":true,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Hi Jakub,\n\nOn 06/11/2020 18:26, Jakub Narębski wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n>> Philip Oakley <philipoakley@iee.email> writes:\n>>\n>>> This may be not part of the the main project, but could you consider, if\n>>> time permits, also adding some entries into the Git Glossary (`git help\n>>> glossary`) for the various terms we are using here and elsewhere, e.g.\n>>> 'topological levels', 'generation number', 'corrected commit date' (and\n>>> its fancy technical name for the use of date heuristics e.g. the\n>>> 'chronological ordering';).\n>>>\n>>> The glossary can provide a reference, once the issues are resolved. The\n>>> History Simplification and Commit Ordering section of git-log maybe a\n>>> useful guide to some of the terms that would link to the glossary.\n>> Ah, I first thought that Documentation/rev-list-options.txt (which\n>> is the relevant part of \"git log\" documentation you mention here)\n>> already have references to deep technical terms explained in the\n>> glossary and you are suggesting Abhishek to mimic the arrangement by\n>> adding new and agreed-upon terms to the glossary and referring to\n>> them from the commit-graph documentation updated by this series.\n>>\n>> But sadly that is not the case.  What you are saying is that you\n>> noticed that rev-list-options.txt needs a similar \"the terms we use\n>> to explain these two sections should be defined and explained in the\n>> glossary (if they are not) and new references to glossary should be\n>> added there\" update.\n>>\n>> In any case, that is a very good suggestion.  I agree that updating\n>> \"git log\" doc may be outside the scope of Abhishek's theme, but it\n>> would be very good to have such an update by anybody ;-)\n> The only possible problem I see with this suggestion is that some of\n> those terms (like 'topological levels' and 'corrected commit date') are\n> technical terms that should be not of concern for Git user, only for\n> developers working on Git.  (However one could encounter the term\n> \"generation number\" in `git commit-graph verify` output.)\nHowever we do mention \"topolog*\"  in a number of the manual pages, and\nrather less, as yet, in the technical pages.\n\n\"Lexicographic\" and \"chronological\" are in the same group of fancy\ntechnical words ;-)\n\n>\n> I don't think adding technical terms that the user won't encounter in\n> the documentation or among messages that Git outputs would be not a good\n> idea.  It could confuse users, rather than help them.\n>\n> Conversely, perhaps we should add Documentation/technical/glossary.txt\n> to help developers.\n\nI would agree that the Glossary probably ought to be split into the\nprimary, secondary and background terms so that the core concepts are\nseparated from the academic/developer style terms.\n\nGit does rip up most of what folks think about version \"control\",\nusually based on the imperfect replication of physical artefacts.\n>\n> P.S. By the way, when looking at Documentation/glossary-content.txt, I\n> have noticed few obsolescent entries, like \"Git archive\", few that have\n> description that soon could be or is obsolete and would need updating,\n> like \"master\" (when default branch switch to \"main\"), or \"object\n> identifier\" and \"SHA-1\" (when Git switches away from SHA-1 as hash\n> function).\nThe obsolescent items can be updated. I'm expecting that the 'main' and\n'SHA-' changes will eventually be picked up as part of the respective\npatch series, hopefully as part of the global replacements.\n\n--\nPhilip\n"},{"id":"409453","messageId":"85imaej8nh.fsf@gmail.com","threadId":"53933","inReplyTo":"8d43335d-a0b4-511e-f132-057343234503@iee.email","subject":"Re: Extending and updating gitglossary","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-11-10T01:35:46Z","receivedAt":"2020-11-10T01:35:54Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello Philip,\n\nPhilip Oakley <philipoakley@iee.email> writes:\n> On 06/11/2020 18:26, Jakub Narębski wrote:\n>> Junio C Hamano <gitster@pobox.com> writes:\n>>> Philip Oakley <philipoakley@iee.email> writes:\n>>>\n>>>> This may be not part of the the main project, but could you consider, if\n>>>> time permits, also adding some entries into the Git Glossary (`git help\n>>>> glossary`) for the various terms we are using here and elsewhere, e.g.\n>>>> 'topological levels', 'generation number', 'corrected commit date' (and\n>>>> its fancy technical name for the use of date heuristics e.g. the\n>>>> 'chronological ordering';).\n>>>>\n>>>> The glossary can provide a reference, once the issues are resolved. The\n>>>> History Simplification and Commit Ordering section of git-log maybe a\n>>>> useful guide to some of the terms that would link to the glossary.\n>>>\n>>> Ah, I first thought that Documentation/rev-list-options.txt (which\n>>> is the relevant part of \"git log\" documentation you mention here)\n>>> already have references to deep technical terms explained in the\n>>> glossary and you are suggesting Abhishek to mimic the arrangement by\n>>> adding new and agreed-upon terms to the glossary and referring to\n>>> them from the commit-graph documentation updated by this series.\n>>>\n>>> But sadly that is not the case.  What you are saying is that you\n>>> noticed that rev-list-options.txt needs a similar \"the terms we use\n>>> to explain these two sections should be defined and explained in the\n>>> glossary (if they are not) and new references to glossary should be\n>>> added there\" update.\n\nWhat terms you feel need glossary entry?\n\n>>> In any case, that is a very good suggestion.  I agree that updating\n>>> \"git log\" doc may be outside the scope of Abhishek's theme, but it\n>>> would be very good to have such an update by anybody ;-)\n>>\n>> The only possible problem I see with this suggestion is that some of\n>> those terms (like 'topological levels' and 'corrected commit date') are\n>> technical terms that should be not of concern for Git user, only for\n>> developers working on Git.  (However one could encounter the term\n>> \"generation number\" in `git commit-graph verify` output.)\n\nTo be more precise, I think that user-facing glossary should include\nonly terms that appear in user-facing documentation and in output\nmessages of Git commands (with the possible exception of maybe output\nmessages of some low-level plumbing).\n\nI think that the developer-facing glossary should include terms that\nappear in technical documentation, and in commit messages in Git\nhistory.\n\n> However we do mention \"topolog*\"  in a number of the manual pages, and\n> rather less, as yet, in the technical pages.\n>\n> \"Lexicographic\" and \"chronological\" are in the same group of fancy\n> technical words ;-)\n\nI think that 'topological level' would appear only in technical\ndocumentation; if it would be the case then there is no reason to add it\nto user-facing glossary (to gitglossary manpage).\n\n'Topological order' or 'topological sort', 'lexicographical order' and\n'chronological order' are not Git-specific terms, and there are no\nGit-specific ambiguities.  I am therefore a bit unsure about adding them\nto *Git* glossary.\n\n- In computer science, a _topological sort_ or _topological_ ordering of\n  a directed graph is a linear ordering of its vertices such that for\n  every directed edge uv from vertex u to vertex v, u comes before v in\n  the ordering.\n\n  For Git it means that top to bottom, commits always appear before\n  their parents. With `--graph` or `--topo-order` Git also avoids\n  showing commits on multiple lines of history intermixed.\n\n- In mathematics, the _lexicographic_ or _lexicographical order_ (also\n  known as lexical order, dictionary order, etc.) is a generalization of\n  the alphabetical order.\n\n  For Git it is simply alphabetical order.\n\n- _Chronological order_ is the arrangement of things following one after\n  another in time; or in other words date order.\n\n  Note that `git log --date-order` commits also always appear before\n  their parents, but otherwise commits are shown in the commit timestamp\n  order (committer date order)\n\n>>\n>> I don't think adding technical terms that the user won't encounter in\n>> the documentation or among messages that Git outputs would be not a good\n>> idea.  It could confuse users, rather than help them.\n>>\n>> Conversely, perhaps we should add Documentation/technical/glossary.txt\n>> to help developers.\n>\n> I would agree that the Glossary probably ought to be split into the\n> primary, secondary and background terms so that the core concepts are\n> separated from the academic/developer style terms.\n\nI don't thing we need three separate layers; in my opinion separating\nterms that user of Git might encounter from terms that somebody working\non developing Git may encounter would be enough.\n\nThe technical glossary / dictionary could also help onboarding...\n\n>\n> Git does rip up most of what folks think about version \"control\",\n> usually based on the imperfect replication of physical artefacts.\n\nI don't quite understand what you wanted to say there.  Could you\nexplain in more detail, please?\n\n>> P.S. By the way, when looking at Documentation/glossary-content.txt, I\n>> have noticed few obsolescent entries, like \"Git archive\", few that have\n>> description that soon could be or is obsolete and would need updating,\n>> like \"master\" (when default branch switch to \"main\"), or \"object\n>> identifier\" and \"SHA-1\" (when Git switches away from SHA-1 as hash\n>> function).\n>\n> The obsolescent items can be updated. I'm expecting that the 'main' and\n> 'SHA-' changes will eventually be picked up as part of the respective\n> patch series, hopefully as part of the global replacements.\n\nHere I meant that \"Git archive\" entry is not important anymore, as I\nthink there are no active users of GNU arch version control system (no\n\"arch people\"); arch's last release was in 2006, and its replacement,\nBazaar (or 'bzr') doesn't use this term. So I think it can be safely\nremoved in 2020, after 14 years after last release of arch.\n\nIn most cases \"SHA-1\" in the descriptions of terms in glossary should be\nreplaced by \"object identifier\" (to be more generic).  This can be\nsafely done before switch to NewHash is ready and announced.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"409486","messageId":"8b6d49e3-fdb8-ce63-650e-c937a2da2c7a@iee.email","threadId":"53933","inReplyTo":"85imaej8nh.fsf@gmail.com","subject":"Re: Extending and updating gitglossary","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2020-11-10T14:04:23Z","receivedAt":"2020-11-10T14:04:35Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Hi Jakub,\n\nOn 10/11/2020 01:35, Jakub Narębski wrote:\n> Hello Philip,\n>\n> Philip Oakley <philipoakley@iee.email> writes:\n>> On 06/11/2020 18:26, Jakub Narębski wrote:\n>>> Junio C Hamano <gitster@pobox.com> writes:\n>>>> Philip Oakley <philipoakley@iee.email> writes:\n>>>>\n>>>>> This may be not part of the the main project, but could you consider, if\n>>>>> time permits, also adding some entries into the Git Glossary (`git help\n>>>>> glossary`) for the various terms we are using here and elsewhere, e.g.\n>>>>> 'topological levels', 'generation number', 'corrected commit date' (and\n>>>>> its fancy technical name for the use of date heuristics e.g. the\n>>>>> 'chronological ordering';).\n>>>>>\n>>>>> The glossary can provide a reference, once the issues are resolved. The\n>>>>> History Simplification and Commit Ordering section of git-log maybe a\n>>>>> useful guide to some of the terms that would link to the glossary.\n>>>> Ah, I first thought that Documentation/rev-list-options.txt (which\n>>>> is the relevant part of \"git log\" documentation you mention here)\n>>>> already have references to deep technical terms explained in the\n>>>> glossary and you are suggesting Abhishek to mimic the arrangement by\n>>>> adding new and agreed-upon terms to the glossary and referring to\n>>>> them from the commit-graph documentation updated by this series.\n>>>>\n>>>> But sadly that is not the case.  What you are saying is that you\n>>>> noticed that rev-list-options.txt needs a similar \"the terms we use\n>>>> to explain these two sections should be defined and explained in the\n>>>> glossary (if they are not) and new references to glossary should be\n>>>> added there\" update.\n> What terms you feel need glossary entry?\nWhile it was Junio that made the comment, I'd agree that we should be\nusing the glossary to explain, in a general sense, the terms that are\nused is a specialist sense. As the user community expands, their natural\nunderstanding of some of the terms diminishes.\n>\n>>>> In any case, that is a very good suggestion.  I agree that updating\n>>>> \"git log\" doc may be outside the scope of Abhishek's theme, but it\n>>>> would be very good to have such an update by anybody ;-)\n>>> The only possible problem I see with this suggestion is that some of\n>>> those terms (like 'topological levels' and 'corrected commit date') are\n>>> technical terms that should be not of concern for Git user, only for\n>>> developers working on Git.  (However one could encounter the term\n>>> \"generation number\" in `git commit-graph verify` output.)\n> To be more precise, I think that user-facing glossary should include\n> only terms that appear in user-facing documentation and in output\n> messages of Git commands (with the possible exception of maybe output\n> messages of some low-level plumbing).\nAnd where implied, the underlying concepts when they aren't obvious, or\nlack general terms (e.g. the 'staging area' discussions)\n>\n> I think that the developer-facing glossary should include terms that\n> appear in technical documentation, and in commit messages in Git\n> history.\n>\n>> However we do mention \"topolog*\"  in a number of the manual pages, and\n>> rather less, as yet, in the technical pages.\n>>\n>> \"Lexicographic\" and \"chronological\" are in the same group of fancy\n>> technical words ;-)\n> I think that 'topological level' would appear only in technical\n> documentation; if it would be the case then there is no reason to add it\n> to user-facing glossary (to gitglossary manpage).\n>\n> 'Topological order' or 'topological sort', 'lexicographical order' and\n> 'chronological order' are not Git-specific terms, and there are no\n> Git-specific ambiguities.  I am therefore a bit unsure about adding them\n> to *Git* glossary.\n\nIt is that they aren't terms used in normal speech, so many folks do not\ncomprehend the implied precision that the docs assume, nor the problems\nthey may hide.\n>\n> - In computer science, a _topological sort_ or _topological_ ordering of\n>   a directed graph is a linear ordering of its vertices such that for\n>   every directed edge uv from vertex u to vertex v, u comes before v in\n>   the ordering.\nDoes this imply that those who aren't computer scientists shouldn't be\nusing Git?\n>\n>   For Git it means that top to bottom, commits always appear before\n>   their parents. With `--graph` or `--topo-order` Git also avoids\n>   showing commits on multiple lines of history intermixed.\n>\n> - In mathematics, the _lexicographic_ or _lexicographical order_ (also\n>   known as lexical order, dictionary order, etc.) is a generalization of\n>   the alphabetical order.\n>\n>   For Git it is simply alphabetical order. \nASCII order, Case sensitivity, Special characters, etc.\n>\n> - _Chronological order_ is the arrangement of things following one after\n>   another in time; or in other words date order.\nGiven that most  résumés (the thing most folk see that asks for date\norder) is latest first, does this clarify which way chronological is? (I\nsee this regularly in my other volunteer work).\n>\n>   Note that `git log --date-order` commits also always appear before\n>   their parents, but otherwise commits are shown in the commit timestamp\n>   order (committer date order)\n\n>\n>>> I don't think adding technical terms that the user won't encounter in\n>>> the documentation or among messages that Git outputs would be not a good\n>>> idea.  It could confuse users, rather than help them.\n>>>\n>>> Conversely, perhaps we should add Documentation/technical/glossary.txt\n>>> to help developers.\n>> I would agree that the Glossary probably ought to be split into the\n>> primary, secondary and background terms so that the core concepts are\n>> separated from the academic/developer style terms.\n> I don't thing we need three separate layers; in my opinion separating\n> terms that user of Git might encounter from terms that somebody working\n> on developing Git may encounter would be enough.\n>\n> The technical glossary / dictionary could also help onboarding...\n>\n>> Git does rip up most of what folks think about version \"control\",\n>> usually based on the imperfect replication of physical artefacts.\n> I don't quite understand what you wanted to say there.  Could you\n> explain in more detail, please?\nBackground, I see Git & Version Control from an engineers view point,\nrather than developers view.\n\nIn the \"real\" world there are no perfect copies, we serialise key items\nso that we can track their degradation, and replace them when required.\nWe attempt to \"Control\" what is happening. Our documentation and\nmonitoring systems have layers of control to ensure only suitably\nqualified persons may access and inspect critical items, can record and\naccess previous status reports, etc. There is only one \"Mona Lisa\", with\ncritical access controls, even though there are 'copies'\nhttps://en.wikipedia.org/wiki/Mona_Lisa#Early_versions_and_copies.\nAlmost all of our terminology for configuration control comes from the\n'real' world, i.e. pre-modern computing.\n\nGit turns all that on its head. We can make perfect duplicates (they're\nnot copies, not replicas..). The Object name is immutable. It's either\nright or wrong (exempt the SHAttered sha-1 breakage; were moving to\nsha-256). Git does *not* provide any access control. It supports the\n'software freedoms' by distributing the control to the user. The\nrepository is a version storage system, and the OIDs allow easy\nauthentication between folks that they are looking at the same object,\nand all its implied descendants.\n\nGit has ripped up classical 'real' world version control. In many areas\nwe need new or alternative terms, and documents that explain them to\nscreen writers(*) and the many other non CS-major users of Git (and some\nengineers;-)\n\n(*) there's a diff pattern for them, IIRC, or at least one was proposed.\n>\n>>> P.S. By the way, when looking at Documentation/glossary-content.txt, I\n>>> have noticed few obsolescent entries, like \"Git archive\", few that have\n>>> description that soon could be or is obsolete and would need updating,\n>>> like \"master\" (when default branch switch to \"main\"), or \"object\n>>> identifier\" and \"SHA-1\" (when Git switches away from SHA-1 as hash\n>>> function).\n>> The obsolescent items can be updated. I'm expecting that the 'main' and\n>> 'SHA-' changes will eventually be picked up as part of the respective\n>> patch series, hopefully as part of the global replacements.\n> Here I meant that \"Git archive\" entry is not important anymore, as I\n> think there are no active users of GNU arch version control system (no\n> \"arch people\"); arch's last release was in 2006, and its replacement,\n> Bazaar (or 'bzr') doesn't use this term. So I think it can be safely\n> removed in 2020, after 14 years after last release of arch.\n>\n> In most cases \"SHA-1\" in the descriptions of terms in glossary should be\n> replaced by \"object identifier\" (to be more generic).  This can be\n> safely done before switch to NewHash is ready and announced.\n>\n> Best,\n--\nPhilip\n"},{"id":"409578","messageId":"85v9ecixct.fsf@gmail.com","threadId":"53933","inReplyTo":"8b6d49e3-fdb8-ce63-650e-c937a2da2c7a@iee.email","subject":"Re: Extending and updating gitglossary","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-11-10T23:52:02Z","receivedAt":"2020-11-10T23:52:33Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hello Philip,\n\nPhilip Oakley <philipoakley@iee.email> writes:\n> On 10/11/2020 01:35, Jakub Narębski wrote:\n>> Philip Oakley <philipoakley@iee.email> writes:\n>>> On 06/11/2020 18:26, Jakub Narębski wrote:\n>>>> Junio C Hamano <gitster@pobox.com> writes:\n>>>>> Philip Oakley <philipoakley@iee.email> writes:\n>>>>>\n>>>>>> This may be not part of the the main project, but could you consider, if\n>>>>>> time permits, also adding some entries into the Git Glossary (`git help\n>>>>>> glossary`) for the various terms we are using here and elsewhere, e.g.\n>>>>>> 'topological levels', 'generation number', 'corrected commit date' (and\n>>>>>> its fancy technical name for the use of date heuristics e.g. the\n>>>>>> 'chronological ordering';).\n>>>>>>\n>>>>>> The glossary can provide a reference, once the issues are resolved. The\n>>>>>> History Simplification and Commit Ordering section of git-log maybe a\n>>>>>> useful guide to some of the terms that would link to the glossary.\n[...]\n>> What terms you feel need glossary entry?\n>\n> While it was Junio that made the comment, I'd agree that we should be\n> using the glossary to explain, in a general sense, the terms that are\n> used is a specialist sense. As the user community expands, their natural\n> understanding of some of the terms diminishes.\n\nI was hoping for a list of terms from the abovementioned sections of\ngit-log manpage you feel need entry in gitglosary(7).\n\n[...]\n>> To be more precise, I think that user-facing glossary should include\n>> only terms that appear in user-facing documentation and in output\n>> messages of Git commands (with the possible exception of maybe output\n>> messages of some low-level plumbing).\n>\n> And where implied, the underlying concepts when they aren't obvious, or\n> lack general terms (e.g. the 'staging area' discussions)\n\nTrue, 'staging area' should IMVHO be in glossary (replacing or in\naddition to older less specific term 'index', previous name for 'staging\narea' term).\n\n>> I think that the developer-facing glossary should include terms that\n>> appear in technical documentation, and in commit messages in Git\n>> history.\n\nSuch as 'topological levels', 'commit slab' / 'on the slab', etc.\n\n>>> However we do mention \"topolog*\"  in a number of the manual pages, and\n>>> rather less, as yet, in the technical pages.\n>>>\n>>> \"Lexicographic\" and \"chronological\" are in the same group of fancy\n>>> technical words ;-)\n>>\n>> I think that 'topological level' would appear only in technical\n>> documentation; if it would be the case then there is no reason to add it\n>> to user-facing glossary (to gitglossary manpage).\n>>\n>> 'Topological order' or 'topological sort', 'lexicographical order' and\n>> 'chronological order' are not Git-specific terms, and there are no\n>> Git-specific ambiguities.  I am therefore a bit unsure about adding them\n>> to *Git* glossary.\n>\n> It is that they aren't terms used in normal speech, so many folks do not\n> comprehend the implied precision that the docs assume, nor the problems\n> they may hide.\n\nRight.\n\n>> - In computer science, a _topological sort_ or _topological_ ordering of\n>>   a directed graph is a linear ordering of its vertices such that for\n>>   every directed edge uv from vertex u to vertex v, u comes before v in\n>>   the ordering.\n>\n> Does this imply that those who aren't computer scientists shouldn't be\n> using Git?\n\nI think that in most cases where we refer to topological order in the\ndocumentation we describe it there.  It might be good idea to add it to\nthe glossary, especially because Git uses it often in a very specific\nsense.\n\nOn the other hand, should we define 'topology' or 'graph' as well? Or\n'glossary' ;-) ? Those don't have any special meaning in Git, and can be\nas well found in the dictionary or Wikipedia.\n\n>>   For Git it means that top to bottom, commits always appear before\n>>   their parents. With `--graph` or `--topo-order` Git also avoids\n>>   showing commits on multiple lines of history intermixed.\n>>\n>> - In mathematics, the _lexicographic_ or _lexicographical order_ (also\n>>   known as lexical order, dictionary order, etc.) is a generalization of\n>>   the alphabetical order.\n>>\n>>   For Git it is simply alphabetical order. \n>\n> ASCII order, Case sensitivity, Special characters, etc.\n\nActually I don't know. Let me check: the only place this term appears in\nthe documentation is in git-tag(1) manpage and related documentation.\nIt simplly uses strcmp(), or strcasecmp() when using `--ignore-case`\noption; so by default case sensitive.\n\nIt looks like it does not take locale-specific rules.\n\n>> - _Chronological order_ is the arrangement of things following one after\n>>   another in time; or in other words date order.\n>\n> Given that most résumés (the thing most folk see that asks for date\n> order) is latest first, does this clarify which way chronological is? (I\n> see this regularly in my other volunteer work).\n\nRight, it might be not obvious at first glance that Git outputs most\nrecent commits first, that is newest commits are on top. Though if you\nthink about it in more detail, it is the only ordering that makes sense,\nespecially for projects with a long history; first, it is newest commits\nthat are most interesting, and second Git always walks the history from\nchild to parent.\n\n>>   Note that `git log --date-order` commits also always appear before\n>>   their parents, but otherwise commits are shown in the commit timestamp\n>>   order (committer date order)\n\n[...]\n>>> Git does rip up most of what folks think about version \"control\",\n>>> usually based on the imperfect replication of physical artefacts.\n>>\n>> I don't quite understand what you wanted to say there.  Could you\n>> explain in more detail, please?\n>\n> Background, I see Git & Version Control from an engineers view point,\n> rather than developers view.\n>\n> In the \"real\" world there are no perfect copies, we serialise key items\n> so that we can track their degradation, and replace them when required.\n> We attempt to \"Control\" what is happening. Our documentation and\n> monitoring systems have layers of control to ensure only suitably\n> qualified persons may access and inspect critical items, can record and\n> access previous status reports, etc. There is only one \"Mona Lisa\", with\n> critical access controls, even though there are 'copies'\n> https://en.wikipedia.org/wiki/Mona_Lisa#Early_versions_and_copies.\n> Almost all of our terminology for configuration control comes from the\n> 'real' world, i.e. pre-modern computing.\n>\n> Git turns all that on its head. We can make perfect duplicates (they're\n> not copies, not replicas..). The Object name is immutable. It's either\n> right or wrong (exempt the SHAttered sha-1 breakage; were moving to\n> sha-256). Git does *not* provide any access control. It supports the\n> 'software freedoms' by distributing the control to the user. The\n> repository is a version storage system, and the OIDs allow easy\n> authentication between folks that they are looking at the same object,\n> and all its implied descendants.\n>\n> Git has ripped up classical 'real' world version control. In many areas\n> we need new or alternative terms, and documents that explain them to\n> screen writers(*) and the many other non CS-major users of Git (and some\n> engineers;-)\n>\n> (*) there's a diff pattern for them, IIRC, or at least one was proposed.\n\nRight, though for me the concept of 'version control' was by default\nalways about the digital, usually the source code.\n\nThere are different editions of books, changes to non-digital technical\ndrawings and plans (AFAIK often in the form of physical foil overlays as\nsubsequent layers, if done well; overdrawing on the same layer if not),\namendment and changes to laws, etc.\n\n\nAnyway, the question is what level of knowledge can we assume from the\naverage Git user -- this would affect the spread of terms that should be\nconsidered for the Git glossary.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"409751","messageId":"20201112100133.GA2691@Abhishek-Arch","threadId":"53933","inReplyTo":"85zh41sxow.fsf@gmail.com","subject":"Re: [PATCH v4 08/10] commit-graph: use generation v2 only if entire chain does","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-11-12T10:01:33Z","receivedAt":"2020-11-12T10:04:30Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Sun, Nov 01, 2020 at 01:55:11AM +0100, Jakub Narębski wrote:\n> Hi Abhishek,\n> \n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > Since there are released versions of Git that understand generation\n> > numbers in the commit-graph's CDAT chunk but do not understand the GDAT\n> > chunk, the following scenario is possible:\n> >\n> > 1. \"New\" Git writes a commit-graph with the GDAT chunk.\n> > 2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n> \n> All right.\n> \n> >\n> > Because of the current use of inspecting the current layer for a\n> > chunk_generation_data pointer, the commits in the lower layer will be\n> > interpreted as having very large generation values (commit date plus\n> > offset) compared to the generation numbers in the top layer (topological\n> > level). This violates the expectation that the generation of a parent is\n> > strictly smaller than the generation of a child.\n> \n> I think this paragraphs tries too much to be concise, with the result it\n> is less clear than it could be.  Perhaps it would be better to separate\n> \"what-if\" from the current behavior.\n> \n>   If each layer of split commit-graph is treated independently, as it\n>   were the case before this commit, with Git inspecting only the current\n>   layer for chunk_generation_data pointer, commits in the lower layer\n>   (one with GDAT) would have corrected commit date as their generation\n>   number, while commits in the upper layer would have topological levels\n>   as their generation.  Corrected commit dates have usually much larger\n>   values than topological levels.  This means that if we take two\n>   commits, one from the upper layer, and one reachable from it in the\n>   lower layer, then the expectation that the generation of a parent is\n>   smaller than the generation of a child would be violated.\n> \n\nThanks, that's better.\n\n> >\n> > It is difficult to expose this issue in a test. Since we _start_ with\n> > artificially low generation numbers, any commit walk that prioritizes\n> > generation numbers will walk all of the commits with high generation\n> > number before walking the commits with low generation number. In all the\n> > cases I tried, the commit-graph layers themselves \"protect\" any\n> > incorrect behavior since none of the commits in the lower layer can\n> > reach the commits in the upper layer.\n> \n> I don't quite understand the issue here. Unless none of the following\n> query commands short-circuit and all walk the commit graph regardless of\n> what generation numbers tell them, they should give different results\n> with and without the commit graph, if we take two commits one from lower\n> layer of split commit graph with GDAT, and one commit from the higher\n> layer without GDAT, one lower reachable from the other higher.\n> \n> We have the following query commands that we can check:\n>   $ git merge-base --is-ancestor <lower> <higher>\n>   $ git merge-base --independent <lower> <higher>\n>   \n>   $ git tag --contains <tag-to-lower>\n>   $ git tag --merged <tag-to-higher>\n>   $ git branch --contains <branch-to-lower>\n>   $ git branch --merged <branch-to-higher>\n> \n> The second set of queries require for those commits to be tagged, or\n> have branch pointing at them, respectively.\n> \n> Also, shouldn't `git commit-graph verify` fail with split commit graph\n> where the top layer is created with GIT_TEST_COMMIT_GRAPH_NO_GDAT=1?\n> \n> Let's assume that we have the following history, with newer commits\n> shown on top like in `git log --graph --oneline --all`:\n> \n>           topological     corrected         generation\n>           level           commit date       number^*\n> \n>       d    3                                3\n>       |\n>    c  |    3                                3\n>    |  |                                                 without GDAT\n>  ..|..|.....[layer.boundary]........................................\n>    |  |                                                    with GDAT\n>    |  b    2              1112912113        1112912113\n>    |  |\n>    a  |    2              1112912053        1112912053\n>    | /\n>    |/\n>    r       1              1112911993        1112911993\n> \n> *) each layer inspected individually.\n> \n> With such history, we can for example reach 'a' from 'c', thus\n> `git merge-base --is-ancestor a b` should return true value, but\n> without this commit gen(a) > gen(c), instead of gen(a) <= gen(c);\n> I use here weaker reachability condition, but the one that works\n> also for commits outside the commit-graph (and those for which\n> generation numbers overflows).\n> \n\nThe original explanation was given by Dr. Stolee and he might not have\nthought exhaustively about the issue.\n\nIn any case, your explanation and the history make sense to me. I will\ntry to add test and report back to the mailing list if something goes\nwrong.\n\nThank you for clarifying in such detail.\n\n> >\n> > This issue would manifest itself as a performance problem in this case,\n> > especially with something like \"git log --graph\" since the low\n> > generation numbers would cause the in-degree queue to walk all of the\n> > commits in the lower layer before allowing the topo-order queue to write\n> > anything to output (depending on the size of the upper layer).\n> \n> All right, that's good explanation.\n> \n> ...\n>\n> > @@ -2030,6 +2047,9 @@ static void split_graph_merge_strategy(struct write_commit_graph_context *ctx)\n> >  \t\t}\n> >  \t}\n> >  \n> > +\tif (!ctx->write_generation_data && g->chunk_generation_data)\n> > +\t\tctx->write_generation_data = 1;\n> > +\n> \n> This needs more careful examination, and looking at larger context of\n> those lines.\n> \n> At this point, unless `--split=replace` option is used, 'g' points to\n> the bottom layer out of all topmost layers being merged. We know that if\n> there are GDAT-less layers then these must be top layers, so this means\n> that we can write GDAT chunk in the result of the merge -- because we\n> would be replacing all possible GDAT-less layers (and maybe some with\n> GDAT) with a single layer with the GDAT chunk.\n> \n> The ctx->write_generation_data is set to true unless environment\n> variable GIT_TEST_COMMIT_GRAPH_NO_GDAT is true, and that in\n> write_commit_graph() it would be set to false if topmost layer doesn't\n> have GDAT chunk, and to true if `--split=replace` option is used; see\n> below.\n> \n> Looks good to me.\n> \n> \n> NOTE that this means that GIT_TEST_COMMIT_GRAPH_NO_GDAT prevents from\n> writing GDAT chunk with generation data v2 unless we are merging layers,\n> or replacing all of them with a single layer: then it is _ignored_.\n> \n> Should we clarify this fact in the description of GIT_TEST_COMMIT_GRAPH_NO_GDAT\n> in t/README?  Currently it reads:\n> \n>   GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n>   commit-graph to be written without generation data chunk.\n\nI think it's better to *not* write generation data chunk if\nGIT_TEST_COMMIT_GRAPH_NO_GDAT is set even though all GDAT-less layers\nare merged, that is:\n\n  if (!ctx->write_generation_data &&\n      g->chunk_generation_data &&\n     !git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0))\n    ctx->write_generation_data = 1;\n\nWith this change, we would have a method to force-write commit-graph\nwithout generation data chunk regardless of the shape of split\ncommit-graph files.\n\n> \n> ...\n>\n> > diff --git a/commit-graph.h b/commit-graph.h\n> > index 19a02001fd..ad52130883 100644\n> > --- a/commit-graph.h\n> > +++ b/commit-graph.h\n> > @@ -64,6 +64,7 @@ struct commit_graph {\n> >  \tstruct object_directory *odb;\n> >  \n> >  \tuint32_t num_commits_in_base;\n> > +\tunsigned int read_generation_data;\n> >  \tstruct commit_graph *base_graph;\n> \n> All right, this new field is here to propagate to each layer the\n> information whether we can read from the generation number v2 data\n> chunk.\n> \n> Though I am not sure whether this field should be added here, and\n> whether it should be `unsigned int` (we don't have to be that careful\n> about saving space for this type).\n> \n\nI cannot think of a more appropriate struct than `struct commit_graph`. \nAny particular suggestions?\n\n> > \n> >  \tconst uint32_t *chunk_oid_fanout;\n> > diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> > index 651df89ab2..d0949a9eb8 100755\n> > --- a/t/t5324-split-commit-graph.sh\n> > +++ b/t/t5324-split-commit-graph.sh\n> > @@ -440,4 +440,90 @@ test_expect_success '--split=replace with partial Bloom data' '\n> >  \tverify_chain_files_exist $graphdir\n> >  '\n> >  \n> > +test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n> > +\tmkdir mixed &&\n> \n> This should probably go just before cd-ing into just created\n> subdirectory.\n> \n> > +\tgraphdir=\".git/objects/info/commit-graphs\" &&\n> > +\ttest_oid_cache <<-EOM &&\n> > +\toid_version sha1:1\n> > +\toid_version sha256:2\n> > +\tEOM\n> \n> Minor nitpick: Why use \"EOM\", which is used only twice in Git the test\n> suite, and not the conventional \"EOF\" (used at least 4000 times)?\n\nRight, both instances of \"EOM\" are actually my own. I looked up some\ntest script for oid cache that did use EOM when I first wrote the tests\nbut it's changed now. Will replace.\n> \n> > +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> \n> The t/README says:\n> \n>    - Don't chdir around in tests.  It is not sufficient to chdir to\n>      somewhere and then chdir back to the original location later in\n>      the test, as any intermediate step can fail and abort the test,\n>      causing the next test to start in an unexpected directory.  Do so\n>      inside a subshell if necessary.\n> \n> Though I am not sure if it should apply also to this situation.\n\nWhile I cannot avoid changing directory, using a subshell would be best\nto avoid causing the later tests to start in unexpected directories.\n\n> \n> > +\tgit init &&\n> > +\tgit config core.commitGraph true &&\n> > +\tgit config gc.writeCommitGraph false &&\n> \n> All right.\n> \n> > +\tfor i in $(test_seq 3)\n> > +\tdo\n> > +\t\ttest_commit $i &&\n> > +\t\tgit branch commits/$i || return 1\n> > +\tdone &&\n> > +\tgit reset --hard commits/1 &&\n> > +\tfor i in $(test_seq 4 5)\n> > +\tdo\n> > +\t\ttest_commit $i &&\n> > +\t\tgit branch commits/$i || return 1\n> > +\tdone &&\n> > +\tgit reset --hard commits/2 &&\n> > +\tfor i in $(test_seq 6 10)\n> > +\tdo\n> > +\t\ttest_commit $i &&\n> > +\t\tgit branch commits/$i || return 1\n> > +\tdone &&\n> > +\tgit commit-graph write --reachable --split &&\n> \n> Is there a reason why we do not check just written commit-graph file\n> with `test-tool read-graph >output-layer-1`?\n\nWe could check the written commit-graph file at this point but it's same\nas existing tests as above.\n\n> \n> > +\tgit reset --hard commits/2 &&\n> > +\tgit merge commits/4 &&\n> \n> Shouldn't we use `test_merge` instead of `git merge`; I am not sure when\n> to use one or the other?\n\n`test_merge` is used in 26 places whereas `git merge` is used in over a\nthousand places. `test_merge` is just not widely adopted and this lack\nof adoption prevents further use.\n\n> \n> > +\tgit branch merge/1 &&\n> > +\tgit reset --hard commits/4 &&\n> > +\tgit merge commits/6 &&\n> > +\tgit branch merge/2 &&\n> \n> It would be nice to have ASCII-art of the history (of the graph of\n> revisions) created here for subsequent tests:\n> \n>                                         \n>            /- 6 <-- 7 <-- 8 <-- 9 <-- 10*\n>           /    \\-\\\n>          /        \\\n>   1 <-- 2 <-- 3*   \\--\\\n>   |      \\             \\ \n>   |       \\-----\\       \\\n>    \\             \\       \\\n>     \\-- 4*<------ M/1     M/2\n>         |\\               /  \n>         | \\-- 5*        /\n>         \\              /\n>          \\------------/\n> \n>   * - 1st layer  \n> \n> Though I am not sure if what I have created is readable; I think a\n> better way to draw this graph is possible, for example:\n> \n>                /- 3*\n>               /\n>              /\n>   1 <------ 2 <---- 6 <-- 7 <-- 8 <-- 9 <-- 10*\n>    \\         \\       \\\n>     \\         \\       \\\n>      \\         \\       \\\n>       \\- 4* <-- M/1     \\    \n>          |\\              \\\n>          | \\------------- M/2\n>          \\\n>           \\---- 5*\n> \n> Edit: as I see the history gets even more complicated, so perhaps\n> ASCII-art diagram of the history with layers marked would be too\n> complicated, and wouldn't bring much.\n> \n> Why do we need such shape of the history in the repository?\n\nWe don't need such a complicated shape. Any commit-graph file with 2-3\nlayers regardless of how commits are related should suffice. Will\nsimplify.\n\n> \n> > +\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n> > +\ttest-tool read-graph >output &&\n> > +\tcat >expect <<-EOF &&\n> > +\theader: 43475048 1 $(test_oid oid_version) 4 1\n> > +\tnum_commits: 2\n> > +\tchunks: oid_fanout oid_lookup commit_metadata\n> > +\tEOF\n> > +\ttest_cmp expect output &&\n> \n> All right, we check that we have 2 commits, and that there is no GDAT\n> chunk.\n> \n> > +\tgit commit-graph verify\n> \n> All right, we verify commit-graph as a whole (both layers).\n> \n> > +'\n> > +\n> > +test_expect_success 'does not write generation data chunk if not present on existing tip' '\n> \n> Hmmm... I wonder if we can come up with a better name for this test;\n> for example should it be \"does not write\" or \"do not write\"?\n\nThat's better.\n\n> \n> > +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> > +\tgit reset --hard commits/3 &&\n> > +\tgit merge merge/1 &&\n> > +\tgit merge commits/5 &&\n> > +\tgit merge merge/2 &&\n> > +\tgit branch merge/3 &&\n> \n> The commit graph gets complicated, so it would not be easy to visualize\n> it with ASCII-art diagram without any crossed lines.  Maybe `git log\n> --graph --oneline --all` would help:\n> \n> *   (merge/3) Merge branch 'merge/2'\n> |\\\n> | *   (merge/2) Merge branch 'commits/6'\n> | |\\\n> * | \\   Merge branch 'commits/5'\n> |\\ \\ \\\n> | * | | (commits/5) 5\n> | |/ /\n> * | |   Merge branch 'merge/1'\n> |\\ \\ \\\n> | * | | (merge/1) Merge branch 'commits/4'\n> | |\\| |\n> | | * | (commits/4) 4\n> * | | | (commits/3) 3\n> |/ / /\n> | | | * (commits/10) 10\n> | | | * (commits/9) 9\n> | | | * (commits/8) 8\n> | | | * (commits/7) 7\n> | | |/\n> | | * (commits/6) 6\n> | |/\n> |/|\n> * | (commits/2) 2\n> |/\n> * (commits/1) 1\n> \n> \n> > +\tgit commit-graph write --reachable --split=no-merge &&\n> > +\ttest-tool read-graph >output &&\n> > +\tcat >expect <<-EOF &&\n> > +\theader: 43475048 1 $(test_oid oid_version) 4 2\n> > +\tnum_commits: 3\n> > +\tchunks: oid_fanout oid_lookup commit_metadata\n> > +\tEOF\n> > +\ttest_cmp expect output &&\n> > +\tgit commit-graph verify\n> \n> All right, so here we check that we have layer without GDAT at the top,\n> and we request not to merge layers thus new layer will be created, then\n> the new layer also does not have GDAT chunk (and has 3 commits).\n> \n> Minor nitpick: shouldn't those test be indented?\n> \n\nThe tests look indented to me and `git diff HEAD^ --check` gives nothing.\n\nDid you mean the lines enclosed by EOF delimiter?\n\n> > +'\n> > +\n> > +test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n> > +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> > +\tgit commit-graph write --reachable --split=replace &&\n> > +\ttest_path_is_file $graphdir/commit-graph-chain &&\n> > +\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n> > +\tverify_chain_files_exist $graphdir &&\n> \n> All right, this checks that we have split commit-graph chain that\n> consist of a single layer, and that the commit-graph file for this\n> single layer exists.\n> \n> > +\tgraph_read_expect 15 &&\n> \n> Shouldn't we use `test-tool read-graph` to check whether generation_data\n> chunk is present... ah, sorry, I have realized that after previous\n> patches `graph_read_expect 15` implicitly checks the latter, because in\n> its' use of `test-tool read-graph` it does expect generation_data chunk.\n> \n> So we use `test-tool read-graph` manually to check that generation_data\n> chunk is absent, and we use graph_read_expect to check that it is\n> present (and in both cases that the number of commits matches).  I\n> wonder if it would be possible to simplify that...\n>\n\nThe problem here is graph_read_expect() as defined in\nt5324-split-commit-graph takes two parameters - number of commits and\nnumber of base graphs. If the number of base graphs is not passed to\nthe function call, it's assumed to be zero. Using a default parameter\nis tricky - I can fix it by manually adding a zero to each of \ngraph_read_expect() in an additional preparatory patch.\n\nAny other suggestions are welcome too.\n\n> \n> > +\tgit commit-graph verify\n> \n> All right.\n> \n> > +'\n> > +\n> > +test_expect_success 'add one commit, write a tip graph' '\n> > +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n> > +\ttest_commit 11 &&\n> > +\tgit branch commits/11 &&\n> > +\tgit commit-graph write --reachable --split &&\n> > +\ttest_path_is_missing $infodir/commit-graph &&\n> > +\ttest_path_is_file $graphdir/commit-graph-chain &&\n> > +\tls $graphdir/graph-*.graph >graph-files &&\n> > +\ttest_line_count = 2 graph-files &&\n> > +\tverify_chain_files_exist $graphdir\n> > +'\n> \n> What it is meant to test?  That adding single-commit to a 15 commit\n> commit-graph file in split mode does not result in layers merging, and\n> actually adds a new layer: we check that we have exactly two layers and\n> that they are all OK.\n\nThis test is meant to check writing to a split graph in \"normal\"\nconditions (i.e. all existing layers have generation data chunk). The\nabove tests are special cases as they involve merging layers with mixed \ngeneration number versions.\n\n> \n> We don't check here that the newly created top layer commit-graph does\n> have GDAT chunk, as it should be if the top layer (in this case the only\n> layer) has GDAT chunk.\n> > +\n> >  test_done\n> \n> One test we are missing is testing that merging layers is done\n> correctly, namely that if we are merging layers in split commit-graph\n> file, and the layer below the ones we are merging lacks GDAT chunk, then\n> the result of the merge should also be without GDAT chunk.  This would\n> require at least two GDAT-less layers in a setup.\n> \n> I'm not sure how difficult writing such test should be.\n\nIt wouldn't be too hard. \n\nAfter the last test, I can write some more commits and write split \ncommit-graph file without GDAT chunk. Then write some more commits \nand merge layers using `git commit-graph write --max-commits=<nr>`.\n\nThanks for pointing this out!\n\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"409864","messageId":"85ft5dinm6.fsf@gmail.com","threadId":"53933","inReplyTo":"20201112100133.GA2691@Abhishek-Arch","subject":"Re: [PATCH v4 08/10] commit-graph: use generation v2 only if entire chain does","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-11-13T09:59:13Z","receivedAt":"2020-11-13T09:59:22Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Abhishek Kumar <abhishekkumar8222@gmail.com> writes:\n> On Sun, Nov 01, 2020 at 01:55:11AM +0100, Jakub Narębski wrote:\n>> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n[...]\n>>> It is difficult to expose this issue in a test. Since we _start_ with\n>>> artificially low generation numbers, any commit walk that prioritizes\n>>> generation numbers will walk all of the commits with high generation\n>>> number before walking the commits with low generation number. In all the\n>>> cases I tried, the commit-graph layers themselves \"protect\" any\n>>> incorrect behavior since none of the commits in the lower layer can\n>>> reach the commits in the upper layer.\n>> \n>> I don't quite understand the issue here. Unless none of the following\n>> query commands short-circuit and all walk the commit graph regardless of\n>> what generation numbers tell them, they should give different results\n>> with and without the commit graph, if we take two commits one from lower\n>> layer of split commit graph with GDAT, and one commit from the higher\n>> layer without GDAT, one lower reachable from the other higher.\n>> \n>> We have the following query commands that we can check:\n>>   $ git merge-base --is-ancestor <lower> <higher>\n>>   $ git merge-base --independent <lower> <higher>\n>>   \n>>   $ git tag --contains <tag-to-lower>\n>>   $ git tag --merged <tag-to-higher>\n>>   $ git branch --contains <branch-to-lower>\n>>   $ git branch --merged <branch-to-higher>\n>> \n>> The second set of queries require for those commits to be tagged, or\n>> have branch pointing at them, respectively.\n>> \n>> Also, shouldn't `git commit-graph verify` fail with split commit graph\n>> where the top layer is created with GIT_TEST_COMMIT_GRAPH_NO_GDAT=1?\n>> \n>> Let's assume that we have the following history, with newer commits\n>> shown on top like in `git log --graph --oneline --all`:\n>> \n>>           topological     corrected         generation\n>>           level           commit date       number^*\n>> \n>>       d    3                                3\n>>       |\n>>    c  |    3                                3\n>>    |  |                                                 without GDAT\n>>  ..|..|.....[layer.boundary]........................................\n>>    |  |                                                    with GDAT\n>>    |  b    2              1112912113        1112912113\n>>    |  |\n>>    a  |    2              1112912053        1112912053\n>>    | /\n>>    |/\n>>    r       1              1112911993        1112911993\n>> \n>> *) each layer inspected individually.\n>> \n>> With such history, we can for example reach 'a' from 'c', thus\n>> `git merge-base --is-ancestor a b` should return true value, but\n>> without this commit gen(a) > gen(c), instead of gen(a) <= gen(c);\n>> I use here weaker reachability condition, but the one that works\n>> also for commits outside the commit-graph (and those for which\n>> generation numbers overflows).\n>> \n>\n> The original explanation was given by Dr. Stolee and he might not have\n> thought exhaustively about the issue.\n>\n> In any case, your explanation and the history make sense to me. I will\n> try to add test and report back to the mailing list if something goes\n> wrong.\n>\n> Thank you for clarifying in such detail.\n\nI don't think you need to add any new test.  It should be enough to check\nthat the first test introduced in this patch, namely 'setup repo for\nmixed generation commit-graph-chain', fails without the change in this\npatch -- as I think it does.  This is because `git commit-graph verify`\nshould fail with mixed-version split commit-graph with GDAT-less layer\non top without this change.\n\nReporting this (possibly as from one sentence to one paragraph in the\ncommit message) would be enough, in my opinion.\n\n[...]\n>> NOTE that this means that GIT_TEST_COMMIT_GRAPH_NO_GDAT prevents from\n>> writing GDAT chunk with generation data v2 unless we are merging layers,\n>> or replacing all of them with a single layer: then it is _ignored_.\n>> \n>> Should we clarify this fact in the description of GIT_TEST_COMMIT_GRAPH_NO_GDAT\n>> in t/README?  Currently it reads:\n>> \n>>   GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n>>   commit-graph to be written without generation data chunk.\n>\n> I think it's better to *not* write generation data chunk if\n> GIT_TEST_COMMIT_GRAPH_NO_GDAT is set even though all GDAT-less layers\n> are merged, that is:\n>\n>   if (!ctx->write_generation_data &&\n>       g->chunk_generation_data &&\n>      !git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0))\n>     ctx->write_generation_data = 1;\n>\n> With this change, we would have a method to force-write commit-graph\n> without generation data chunk regardless of the shape of split\n> commit-graph files.\n\nWhile it would be more consistent to always behave like the old Git with\nGIT_TEST_COMMIT_GRAPH_NO_GDAT=1, it is in my opinion not necessary.\n\nThe only thing we need to test the mixed-version commit-graph chain is\nthe ability to add new layer on top without GDAT.  It does not matter if\nthis layer is created from new commits or a result of partial or full\nmerge of layers.\n\nSo the alternative to extending what GIT_TEST_COMMIT_GRAPH_NO_GDAT does\nthat you propose here would be simply improving the description of t in\nt/README, e.g.\n\n     GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n     commit-graph, or new layer in split commit-graph chain, to be written\n     without generation data chunk.  It does not affect merging of layers.\n\nFor me either solution is fine.\n\n[...]\n>>> diff --git a/commit-graph.h b/commit-graph.h\n>>> index 19a02001fd..ad52130883 100644\n>>> --- a/commit-graph.h\n>>> +++ b/commit-graph.h\n>>> @@ -64,6 +64,7 @@ struct commit_graph {\n>>>  \tstruct object_directory *odb;\n>>>  \n>>>  \tuint32_t num_commits_in_base;\n>>> +\tunsigned int read_generation_data;\n>>>  \tstruct commit_graph *base_graph;\n>> \n>> All right, this new field is here to propagate to each layer the\n>> information whether we can read from the generation number v2 data\n>> chunk.\n>> \n>> Though I am not sure whether this field should be added here, and\n>> whether it should be `unsigned int` (we don't have to be that careful\n>> about saving space for this type).\n>\n> I cannot think of a more appropriate struct than `struct commit_graph`. \n> Any particular suggestions?\n\nAfter thinking about it a bit more, I think it is fine to have it here\nin `struct commit_graph`, it is better than using a global variable\n(which would make code non-reentrant; not that we use multiple threads\nfor reading multiple layers of the commit graph, but we might want to in\nthe future).\n\n[...]\n>>> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n>> \n>> The t/README says:\n>> \n>>    - Don't chdir around in tests.  It is not sufficient to chdir to\n>>      somewhere and then chdir back to the original location later in\n>>      the test, as any intermediate step can fail and abort the test,\n>>      causing the next test to start in an unexpected directory.  Do so\n>>      inside a subshell if necessary.\n>> \n>> Though I am not sure if it should apply also to this situation.\n>\n> While I cannot avoid changing directory, using a subshell would be best\n> to avoid causing the later tests to start in unexpected directories.\n\nThis would allow for easier skipping of tests, and failed tests would\nnot propagate the error (because of subsequent tests after a failed one\nstarting in unexpected directory).\n\n[...]\n>>> +\tfor i in $(test_seq 3)\n>>> +\tdo\n>>> +\t\ttest_commit $i &&\n>>> +\t\tgit branch commits/$i || return 1\n>>> +\tdone &&\n>>> +\tgit reset --hard commits/1 &&\n>>> +\tfor i in $(test_seq 4 5)\n>>> +\tdo\n>>> +\t\ttest_commit $i &&\n>>> +\t\tgit branch commits/$i || return 1\n>>> +\tdone &&\n>>> +\tgit reset --hard commits/2 &&\n>>> +\tfor i in $(test_seq 6 10)\n>>> +\tdo\n>>> +\t\ttest_commit $i &&\n>>> +\t\tgit branch commits/$i || return 1\n>>> +\tdone &&\n>>> +\tgit commit-graph write --reachable --split &&\n>> \n>> Is there a reason why we do not check just written commit-graph file\n>> with `test-tool read-graph >output-layer-1`?\n>\n> We could check the written commit-graph file at this point but it's same\n> as existing tests as above.\n\nAll right, thanks for an explanation.\n\n>> \n>>> +\tgit reset --hard commits/2 &&\n>>> +\tgit merge commits/4 &&\n>> \n>> Shouldn't we use `test_merge` instead of `git merge`; I am not sure when\n>> to use one or the other?\n>\n> `test_merge` is used in 26 places whereas `git merge` is used in over a\n> thousand places. `test_merge` is just not widely adopted and this lack\n> of adoption prevents further use.\n\nAll right then.\n\n>>> +\tgit branch merge/1 &&\n>>> +\tgit reset --hard commits/4 &&\n>>> +\tgit merge commits/6 &&\n>>> +\tgit branch merge/2 &&\n>> \n>> It would be nice to have ASCII-art of the history (of the graph of\n>> revisions) created here for subsequent tests:\n>> \n>>                                         \n>>            /- 6 <-- 7 <-- 8 <-- 9 <-- 10*\n>>           /    \\-\\\n>>          /        \\\n>>   1 <-- 2 <-- 3*   \\--\\\n>>   |      \\             \\ \n>>   |       \\-----\\       \\\n>>    \\             \\       \\\n>>     \\-- 4*<------ M/1     M/2\n>>         |\\               /  \n>>         | \\-- 5*        /\n>>         \\              /\n>>          \\------------/\n>> \n>>   * - 1st layer  \n>> \n>> Though I am not sure if what I have created is readable; I think a\n>> better way to draw this graph is possible, for example:\n>> \n>>                /- 3*\n>>               /\n>>              /\n>>   1 <------ 2 <---- 6 <-- 7 <-- 8 <-- 9 <-- 10*\n>>    \\         \\       \\\n>>     \\         \\       \\\n>>      \\         \\       \\\n>>       \\- 4* <-- M/1     \\    \n>>          |\\              \\\n>>          | \\------------- M/2\n>>          \\\n>>           \\---- 5*\n>> \n>> Edit: as I see the history gets even more complicated, so perhaps\n>> ASCII-art diagram of the history with layers marked would be too\n>> complicated, and wouldn't bring much.\n>> \n>> Why do we need such shape of the history in the repository?\n>\n> We don't need such a complicated shape. Any commit-graph file with 2-3\n> layers regardless of how commits are related should suffice. Will\n> simplify.\n\nIf you are unsire if we need this shape of history to properly test all\ncorner cases of the algorithm, or whether simple history would be\nenough, you can simply compare code coverage.  Git Makefile ha the\n'coverage' target (which requires 'gcov' tool).\n\nNOTE: if it is possible to run 'make coverage' for you, it can be used\nto check if there are any parts of the new code that are not tested.\n\n[...]\n>>> +\tgit commit-graph write --reachable --split=no-merge &&\n>>> +\ttest-tool read-graph >output &&\n>>> +\tcat >expect <<-EOF &&\n>>> +\theader: 43475048 1 $(test_oid oid_version) 4 2\n>>> +\tnum_commits: 3\n>>> +\tchunks: oid_fanout oid_lookup commit_metadata\n>>> +\tEOF\n>>> +\ttest_cmp expect output &&\n>>> +\tgit commit-graph verify\n>> \n>> All right, so here we check that we have layer without GDAT at the top,\n>> and we request not to merge layers thus new layer will be created, then\n>> the new layer also does not have GDAT chunk (and has 3 commits).\n>> \n>> Minor nitpick: shouldn't those test be indented?\n>> \n>\n> The tests look indented to me and `git diff HEAD^ --check` gives nothing.\n>\n> Did you mean the lines enclosed by EOF delimiter?\n\nI'm sorry, that was my mistake -- tabs are used for indent, and the\ntabstop (in my newsreader) when being quoted made it look like it was\nnot indented.\n\n[...]\n>>> +test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n>>> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n>>> +\tgit commit-graph write --reachable --split=replace &&\n>>> +\ttest_path_is_file $graphdir/commit-graph-chain &&\n>>> +\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n>>> +\tverify_chain_files_exist $graphdir &&\n>> \n>> All right, this checks that we have split commit-graph chain that\n>> consist of a single layer, and that the commit-graph file for this\n>> single layer exists.\n>> \n>>> +\tgraph_read_expect 15 &&\n>> \n>> Shouldn't we use `test-tool read-graph` to check whether generation_data\n>> chunk is present... ah, sorry, I have realized that after previous\n>> patches `graph_read_expect 15` implicitly checks the latter, because in\n>> its' use of `test-tool read-graph` it does expect generation_data chunk.\n>> \n>> So we use `test-tool read-graph` manually to check that generation_data\n>> chunk is absent, and we use graph_read_expect to check that it is\n>> present (and in both cases that the number of commits matches).  I\n>> wonder if it would be possible to simplify that...\n\nWhat I wanted to say that it might be better to have a second variant of\ngraph_read_expect() for GDAT-less layers -- but this might be\nunnecessary complication.\n\n> The problem here is graph_read_expect() as defined in\n> t5324-split-commit-graph takes two parameters - number of commits and\n> number of base graphs. If the number of base graphs is not passed to\n> the function call, it's assumed to be zero. Using a default parameter\n> is tricky - I can fix it by manually adding a zero to each of \n> graph_read_expect() in an additional preparatory patch.\n\nAll right, thanks for an explanation.  I should have examined\ngraph_read_expect() in more detail.\n\n> Any other suggestions are welcome too.\n\n[...]\n>>> +test_expect_success 'add one commit, write a tip graph' '\n>>> +\tcd \"$TRASH_DIRECTORY/mixed\" &&\n>>> +\ttest_commit 11 &&\n>>> +\tgit branch commits/11 &&\n>>> +\tgit commit-graph write --reachable --split &&\n>>> +\ttest_path_is_missing $infodir/commit-graph &&\n>>> +\ttest_path_is_file $graphdir/commit-graph-chain &&\n>>> +\tls $graphdir/graph-*.graph >graph-files &&\n>>> +\ttest_line_count = 2 graph-files &&\n>>> +\tverify_chain_files_exist $graphdir\n>>> +'\n>> \n>> What it is meant to test?  That adding single-commit to a 15 commit\n>> commit-graph file in split mode does not result in layers merging, and\n>> actually adds a new layer: we check that we have exactly two layers and\n>> that they are all OK.\n>\n> This test is meant to check writing to a split graph in \"normal\"\n> conditions (i.e. all existing layers have generation data chunk). The\n> above tests are special cases as they involve merging layers with mixed \n> generation number versions.\n\nAll right.\n\n>> \n>> We don't check here that the newly created top layer commit-graph does\n>> have GDAT chunk, as it should be if the top layer (in this case the only\n>> layer) has GDAT chunk.\n>>> +\n>>>  test_done\n>> \n>> One test we are missing is testing that merging layers is done\n>> correctly, namely that if we are merging layers in split commit-graph\n>> file, and the layer below the ones we are merging lacks GDAT chunk, then\n>> the result of the merge should also be without GDAT chunk.  This would\n>> require at least two GDAT-less layers in a setup.\n>> \n>> I'm not sure how difficult writing such test should be.\n>\n> It wouldn't be too hard. \n>\n> After the last test, I can write some more commits and write split \n> commit-graph file without GDAT chunk. Then write some more commits \n> and merge layers using `git commit-graph write --max-commits=<nr>`.\n>\n> Thanks for pointing this out!\n\nGood.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"410411","messageId":"X7ebaubi/FhLMtVO@Abhishek-Arch","threadId":"53933","inReplyTo":"85y2jiqq3c.fsf@gmail.com","subject":"Re: [PATCH v4 09/10] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-11-20T10:33:14Z","receivedAt":"2020-11-20T10:32:55Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Nov 03, 2020 at 06:59:03PM +0100, Jakub Narębski wrote:\n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > With corrected commit dates implemented, we no longer have to rely on\n> > commit date as a heuristic in paint_down_to_common().\n> >\n> > While using corrected commit dates Git walks nearly the same number of\n> > commits as commit date, the process is slower as for each comparision we\n> > have to access a commit-slab (for corrected committer date) instead of\n> > accessing struct member (for committer date).\n> \n> Something for the future: I wonder if it would be worth it to bring back\n> generation number from the commit-slab into `struct commit`.\n> \n> >\n> > For example, the command `git merge-base v4.8 v4.9` on the linux\n> > repository walks 167468 commits, taking 0.135s for committer date and\n> > 167496 commits, taking 0.157s for corrected committer date respectively.\n> \n> I think it would be good idea to explicitly refer to the commit that\n> changed paint_down_to_common() to *not* use generation numbers v1\n> (topological levels) in the cases such as this, namely 091f4cf3 (commit:\n> don't use generation numbers if not needed).  In this commit we have the\n> following:\n> ...\n>\n\nI have re-arranged the first half of commit message: \n\n  091f4cf3 (commit: don't use generation numbers if not needed,\n  2018-08-30) changed paint_down_to_common() to use commit dates instead\n  of generation numbers v1 (topological levels) as the performance\n  regressed on certain topologies. With generation number v2 (corrected\n  commit dates) implemented, we no longer have to rely on commit dates and\n  can use generation numbers.\n  \n  For example, the command `git merge-base v4.8 v4.9` on the Linux\n  repository walks 167468 commits, taking 0.135s for committer date and\n  167496 commits, taking 0.157s for corrected committer date respectively.\n  \n  While using corrected commit dates Git walks nearly the same number of\n  commits as commit date, the process is slower as for each comparision we\n  have to access a commit-slab (for corrected committer date) instead of\n  accessing struct member (for committer date).\n\n> \n> The times you report (0.135s and 0.157s) are close to 0.122s / 0.127s\n> reported in 091f4cf3 - that is most probably because of the differences\n> in the system performance (hardware, operating system, load, etc.).\n> Numbers of commits walked for the committed date heuristics, that is\n> 167,468 agrees with your results; 167,496 (+28) for corrected commit\n> date (generation number v2) is significantly smaller (-468,083) than\n> 635,579 reported for topological levels (generation number v1).\n> \n> I suspect that there are cases (with date skew) where corrected commit\n> date gives better performance than committer date heuristics, and I am\n> quite sure that generation number v2 can give better performance in case\n> where paint_down_to_common() uses generation numbers.\n> \n> .................................................................\n> \n> Here begins separate second change, which is not put into separate\n> commit because it is fairly tightly connected to the change described\n> above.  It would be good idea, in my opinion, to add a sentence that\n> explicitely marks this switch, for example:\n> \n>   This change accidentally broke fragile t6404-recursive-merge test.\n>   t6404-recursive-merge setups a unique repository...\n> \n> Maybe with s/accidentaly/incidentally/.\n> \n\nThanks, will add.\n\n> Or add some other way of connection those two parts of the commit\n> messages.\n> ...\n> >  \n> > +int corrected_commit_dates_enabled(struct repository *r)\n> > +{\n> > +\tstruct commit_graph *g;\n> > +\tif (!prepare_commit_graph(r))\n> > +\t\treturn 0;\n> > +\n> > +\tg = r->objects->commit_graph;\n> > +\n> > +\tif (!g->num_commits)\n> > +\t\treturn 0;\n> > +\n> > +\treturn g->read_generation_data;\n> > +}\n> \n> Very nice abstraction.\n> \n> Minor issue: I wonder if it would be better to use _available() or\n> \"_present()\" rather than _enabled() suffix.\n> \n\nWe could, but that breaks conformity with `generation_numbers_enabled()`.\n\nI see both functions to be similar in nature, to answer whether the\ncommit-graph has X? X could be topological levels or corrected commit\ndates.\n\n> > +\n> >  struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r)\n> >  {\n> >  \tstruct commit_graph *g = r->objects->commit_graph;\n> > diff --git a/commit-graph.h b/commit-graph.h\n> > index ad52130883..d2c048dc64 100644\n> > --- a/commit-graph.h\n> > +++ b/commit-graph.h\n> > @@ -89,13 +89,19 @@ struct commit_graph *read_commit_graph_one(struct repository *r,\n> >  struct commit_graph *parse_commit_graph(struct repository *r,\n> >  \t\t\t\t\tvoid *graph_map, size_t graph_size);\n> >  \n> > +struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n> > +\n> >  /*\n> >   * Return 1 if and only if the repository has a commit-graph\n> >   * file and generation numbers are computed in that file.\n> >   */\n> >  int generation_numbers_enabled(struct repository *r);\n> >  \n> > -struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n> \n> This moving get_bloom_filter_settings() before generation_numbers_enabled() \n> looks like accidental change.  If not, why it is here?\n\nRight, that's an accidental change. I wanted to group\ngeneration_numbers_enabled() and corrected_commit_dates_enabled()\ntogether.\n\n> \n> ...\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"410456","messageId":"X7iz9VhjHG5sbe9p@Abhishek-Arch","threadId":"53933","inReplyTo":"85tuu5q4uy.fsf@gmail.com","subject":"Re: [PATCH v4 10/10] doc: add corrected commit date info","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-11-21T06:30:13Z","receivedAt":"2020-11-21T06:33:29Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Wed, Nov 04, 2020 at 02:37:41AM +0100, Jakub Narębski wrote:\n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> >\n> > With generation data chunk and corrected commit dates implemented, let's\n> > update the technical documentation for commit-graph.\n> >\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> Nice.\n> \n> > ---\n> >  .../technical/commit-graph-format.txt         | 21 +++++--\n> >  Documentation/technical/commit-graph.txt      | 62 ++++++++++++++++---\n> >  2 files changed, 69 insertions(+), 14 deletions(-)\n> >\n> > diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n> > index b3b58880b9..08d9026ad4 100644\n> > --- a/Documentation/technical/commit-graph-format.txt\n> > +++ b/Documentation/technical/commit-graph-format.txt\n> > @@ -4,11 +4,7 @@ Git commit graph format\n> >  The Git commit graph stores a list of commit OIDs and some associated\n> >  metadata, including:\n> >  \n> > -- The generation number of the commit. Commits with no parents have\n> > -  generation number 1; commits with parents have generation number\n> > -  one more than the maximum generation number of its parents. We\n> > -  reserve zero as special, and can be used to mark a generation\n> > -  number invalid or as \"not computed\".\n> > +- The generation number of the commit.\n> \n> All right, because we could store both generation number v1 and\n> generation number v2 in the commit-graph file, and we need to describe\n> both, the description is now consolidated and in only one place.\n> \n> >  \n> >  - The root tree OID.\n> >  \n> > @@ -86,13 +82,26 @@ CHUNK DATA:\n> >        position. If there are more than two parents, the second value\n> >        has its most-significant bit on and the other bits store an array\n> >        position into the Extra Edge List chunk.\n> > -    * The next 8 bytes store the generation number of the commit and\n> > +    * The next 8 bytes store the topological level (generation number v1)\n> > +      of the commit and\n> \n> All right, this is updated information about CDAT chunk.\n> \n> >        the commit time in seconds since EPOCH. The generation number\n> >        uses the higher 30 bits of the first 4 bytes, while the commit\n> >        time uses the 32 bits of the second 4 bytes, along with the lowest\n> >        2 bits of the lowest byte, storing the 33rd and 34th bit of the\n> >        commit time.\n> >  \n> > +  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes)\n> \n> Should we mark this chunk as \"[Optional]\"?  Its absence is not an error.\n\nI think we should mark it as \"optional\", although optional might not\nhave been the best choice word. \n\nOptional (for me) implies that it is configurable and decided by the end-user\ndirectly.  However, it is *conditional* - on the existing commit graph file(s)\n(if any) and the version of Git.\n\n> > +    * This list of 4-byte values store corrected commit date offsets for the\n> > +      commits, arranged in the same order as commit data chunk.\n> > +    * If the corrected commit date offset cannot be stored within 31 bits,\n> > +      the value has its most-significant bit on and the other bits store\n> > +      the position of corrected commit date into the Generation Data Overflow\n> > +      chunk.\n> \n> All right.\n> \n> > +\n> > +  Generation Data Overflow (ID: {'G', 'D', 'O', 'V' }) [Optional]\n> > +    * This list of 8-byte values stores the corrected commit dates for commits\n> > +      with corrected commit date offsets that cannot be stored within 31 bits.\n> \n> A question: do we store 8-byte / 64-bit corrected commit date *directly*,\n> or do we store corrected commit date *offset* as 8-byte / 64-bit value?\n> \n\nWe store the dates directly rather 8-byte offsets. Will clarify.\n\n> Perhaps we should add the information that [like the EDGE chunk] it is\n> present only when necessary, and that it is present only when GDAT chunk\n> is present (it might be obvious, but it could be better to state\n> this explicitly).\n> \n\nIt's always better to be explicit. Thanks for the detailed review.\n\n> > +\n> \n> All right, this is the information about two new chunks (with the\n> mentioned above caveat about the clarity of the description of\n> overflow-handling chunk).\n> \n> >    Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n> >        This list of 4-byte values store the second through nth parents for\n> >        all octopus merges. The second parent value in the commit data stores\n> > diff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\n> > index f14a7659aa..75f71c4c7b 100644\n> > --- a/Documentation/technical/commit-graph.txt\n> > +++ b/Documentation/technical/commit-graph.txt\n> > @@ -38,14 +38,31 @@ A consumer may load the following info for a commit from the graph:\n> >  \n> >  Values 1-4 satisfy the requirements of parse_commit_gently().\n> >  \n> > -Define the \"generation number\" of a commit recursively as follows:\n> > +There are two definitions of generation number:\n> > +1. Corrected committer dates (generation number v2)\n> > +2. Topological levels (generation nummber v1)\n> \n> All right.\n> \n> >  \n> > - * A commit with no parents (a root commit) has generation number one.\n> > +Define \"corrected committer date\" of a commit recursively as follows:\n> >  \n> > - * A commit with at least one parent has generation number one more than\n> > -   the largest generation number among its parents.\n> > +  * A commit with no parents (a root commit) has corrected committer date\n> > +    equal to its committer date.\n> \n> Minor nitpick: the above point has been accidentally indented one space\n> more than necessary, and than is indented in other places.  Or maybe\n> that fixes / unifies the formatting... I am not sure.\n> \n\nThat's a force of habit - I like to write markdown with greater\nindentation. Should have been indented with one space instead of two.\n\n> >  \n> > -Equivalently, the generation number of a commit A is one more than the\n> > +  * A commit with at least one parent has corrected committer date equal to\n> > +    the maximum of its commiter date and one more than the largest corrected\n> > +    committer date among its parents.\n> > +\n> > +  * As a special case, a root commit with timestamp zero has corrected commit\n> > +    date of 1, to be able to distinguish it from GENERATION_NUMBER_ZERO\n> > +    (that is, an uncomputed corrected commit date).\n> \n> All right.  Looks good.\n> \n> > +\n> > +Define the \"topological level\" of a commit recursively as follows:\n> > +\n> > + * A commit with no parents (a root commit) has topological level of one.\n> > +\n> > + * A commit with at least one parent has topological level one more than\n> > +   the largest topological level among its parents.\n> > +\n> \n> All right, this just repeats what was written before, or in other words\n> move existing contents lower/later, just with 'generation number'\n> replaced by 'topological level' (though it might be not obvious from the\n> patch because of the latter change).\n> \n> > +Equivalently, the topological level of a commit A is one more than the\n> >  length of a longest path from A to a root commit. The recursive definition\n> >  is easier to use for computation and observing the following property:\n> >  \n> > @@ -60,6 +77,9 @@ is easier to use for computation and observing the following property:\n> >      generation numbers, then we always expand the boundary commit with highest\n> >      generation number and can easily detect the stopping condition.\n> >  \n> > +The properties applies to both versions of generation number, that is both\n> > +corrected committer dates and topological levels.\n> > +\n> \n> I think it should be \"This property\" or \"The property\", not \"The\n> properties\"; it is a single property, a single condition.\n> \n> We can alternatively say \"This condition is fulfilled by both versions...\",\n> or \"This condition is true for both versions...\".\n> \n> >  This property can be used to significantly reduce the time it takes to\n> >  walk commits and determine topological relationships. Without generation\n> >  numbers, the general heuristic is the following:\n> > @@ -67,7 +87,9 @@ numbers, the general heuristic is the following:\n> >      If A and B are commits with commit time X and Y, respectively, and\n> >      X < Y, then A _probably_ cannot reach B.\n> >  \n> > -This heuristic is currently used whenever the computation is allowed to\n> > +In absence of corrected commit dates (for example, old versions of Git or\n> > +mixed generation graph chains),\n> > +this heuristic is currently used whenever the computation is allowed to\n> >  violate topological relationships due to clock skew (such as \"git log\"\n> >  with default order), but is not used when the topological order is\n> >  required (such as merge base calculations, \"git log --graph\").\n> \n> All right, this explains when commit date heuristics is used (which is\n> less often than before).\n> \n> > @@ -77,7 +99,7 @@ in the commit graph. We can treat these commits as having \"infinite\"\n> >  generation number and walk until reaching commits with known generation\n> >  number.\n> >  \n> > -We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n> > +We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n> \n> All right, 64-bit GENERATION_NUMBER_INFINITY = 0xFFFFFFFFFFFFFFFF is a\n> bit unwieldy...\n> \n> >  in the commit-graph file. If a commit-graph file was written by a version\n> >  of Git that did not compute generation numbers, then those commits will\n> >  have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n> > @@ -93,7 +115,7 @@ fully-computed generation numbers. Using strict inequality may result in\n> >  walking a few extra commits, but the simplicity in dealing with commits\n> >  with generation number *_INFINITY or *_ZERO is valuable.\n> >  \n> > -We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n> > +We use the macro GENERATION_NUMBER_MAX for commits whose\n> \n> This should be\n> \n>   +We use the macro GENERATION_NUMBER_V1_MAX = 0x3FFFFFFF to for commits whose\n>   +topological levels (generation number v1) are computed to be at least this value. We limit at\n>    this value since it is the largest value that can be stored in the\n>   +commit-graph file using the 30 bits available to topological levels. This\n> \n> We need to use \"topological levels\" or \"generation numbers v1\" thorough\n> the rest of this section.\n> \n> >  generation numbers are computed to be at least this value. We limit at\n> >  this value since it is the largest value that can be stored in the\n> >  commit-graph file using the 30 bits available to generation numbers. This\n> > @@ -267,6 +289,30 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n> >  number of commits) could be extracted into config settings for full\n> >  flexibility.\n> >\n> \n> All right, I agree that we don't need to write about overflow handling\n> for storing corrected committer dates (generation number v2) as offsets;\n> this is something format-specific, and this documentation is more about\n> using commit-graph data.  What is present in commit-graph-format.txt\n> should be enough information.\n> \n> Sidenote: I wonder if other Git implementations such as JGit, Dulwich,\n> Gitoxide (gix), go-git have support for the commit-graph file...\n> \n> > +## Handling Mixed Generation Number Chains\n> > +\n> > +With the introduction of generation number v2 and generation data chunk, the\n> > +following scenario is possible:\n> > +\n> > +1. \"New\" Git writes a commit-graph with the corrected commit dates.\n> > +2. \"Old\" Git writes a split commit-graph on top without corrected commit dates.\n> > +\n> > +A naive approach of using the newest available generation number from\n> > +each layer would lead to violated expectations: the lower layer would\n> > +use corrected commit dates which are much larger than the topological\n> > +levels of the higher layer. For this reason, Git inspects each layer to\n> > +see if any layer is missing corrected commit dates. In such a case, Git\n> > +only uses topological level\n> \n> This should end in full stop:\n> \n>   +only uses topological levels.\n> \n> Or maybe we should expand the last sentence a bit:\n> \n>   +only uses topological levels for generation numbers.\n> \n> Sidenote: it is a good explanation, even if Git can make use of the\n> property described below that only topmost layers might be missing\n> corrected commit graph by the construction (so it needs to check only\n> the top layer).\n> \n> > +\n> > +When writing a new layer in split commit-graph, we write corrected commit\n> > +dates if the topmost layer has corrected commit dates written. This\n> > +guarantees that if a layer has corrected commit dates, all lower layers\n> > +must have corrected commit dates as well.\n> > +\n> > +When merging layers, we do not consider whether the merged layers had corrected\n> > +commit dates. Instead, the new layer will have corrected commit dates if and\n> > +only if all existing layers below the new layer have corrected commit dates.\n> > +\n> \n> Perhaps we should explicitly say that when rewriting split commit-graph\n> as a single file (`--split=replace`) then the newly created single layer\n> would store corrected commit dates.\n> \n\nRewriting split commit-graph as a single file is a case where there are\nno \"existing layers below the new layer\". We should clarify that if the\nnew layer is the only layer, it will always have corrected commit dates\nwhen written by compatible versions of Git. \n\nI have appended a paragraph at the end:\n\n  While writing or merging layers, if the new layer is the only layer,\n  it will have corrected commit dates when written by compatible\n  versions of Git. Thus, rewriting split commit-graph as a singel file\n  (`--split=replace`) creates a single layer with corrected commit\n  dates.\n\n> >  ## Deleting graph-{hash} files\n> >  \n> >  After a new tip file is written, some `graph-{hash}` files may no longer\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"410484","messageId":"X7n3rxnGthonwElD@Abhishek-Arch","threadId":"53933","inReplyTo":"85tuu4lmlu.fsf@gmail.com","subject":"Re: [PATCH v4 00/10] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2020-11-22T05:31:40Z","receivedAt":"2020-11-22T05:33:14Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Thu, Nov 05, 2020 at 12:37:49AM +0100, Jakub Narębski wrote:\n> Hi Abhishek,\n> \n> \"Abhishek Kumar via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n> > This patch series implements the corrected commit date offsets as generation\n> > number v2, along with other pre-requisites.\n> \n> Thanks a lot for continued working on this patch series.\n\nThank you so much for the careful review of the series.\n\n> \n> >\n> > Git uses topological levels in the commit-graph file for commit-graph\n> > traversal operations like git log --graph. Unfortunately, using topological\n> > levels can result in a worse performance than without them when compared\n> > with committer date as a heuristics. For example, git merge-base v4.8 v4.9 \n> > on the Linux repository walks 635,579 commits using topological levels and\n> > walks 167,468 using committer date.\n> \n> Very minor nitpick: it would make it easier to read if the commands\n> themself would be put inside single quotes or backticks, e.g. `git log\n> --graph` and `git merge-base v4.8 v4.9`.\n\nThat's unexpected - I wrote the commands within single quotes in the pull\nrequest. Since backticks are rendered as \"code-tags\" on Github, let me \ntry single quotes.\n\n> \n> I wonder if it is worth mentioning (probably not) that this performance\n> hit was the reason why since 091f4cf3 `git merge-base` uses committer\n> date heuristics unless there is a cutoff and using topological levels\n> (generation date v1) is expected to give better performance.\n> \n\nI think that's useful context for someone wondering whether we continue\nto take the performance hit with topological levels or have abandoned\ntopological levels or chosen some another alternative altogether.\n\n> >\n> > Thus, the need for generation number v2 was born. New generation number\n> > needed to provide good performance, increment updates, and backward\n> > compatibility. Due to an unfortunate problem 1\n> \n> Minor issue: this looks a bit strange; is there an error in formatting\n> this part?\n\nYes. The plaintext in pull request description reads as follows:\n\n  Thus, the need for generation number v2 was born. New generation number\n  needed to provide good performance, increment updates, and backward\n  compatibility. Due to an unfortunate problem [1], we also needed a way\n  to distinguish between the old and new generation number without\n  incrementing graph version.\n\n  [1]: https://public-inbox.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nI have been reviewing other pull request descriptions to match their\nstyle (and hope the cover letter renders correctly) and Dr. Stolee in\nhis patch series to \"add --literal value\" option to configuration has\nwritten:\n  \n  As reported [1], 'git maintenance unregister' fails when a repository\n  is located in a directory with regex glob characters.\n\n  [1] https://lore.kernel.org/git/2c2db228-069a-947d-8446-89f4d3f6181a@gmail.com/T/#mb96fa4187a0d6aeda097cd95804a8aafc0273022\n\n(Note the lack of colon after [1])\n\n> \n> > [https://public-inbox.org/git/87a7gdspo4.fsf@evledraar.gmail.com/], we also\n> > needed a way to distinguish between the old and new generation number\n> > without incrementing graph version.\n> >\n> > Various candidates were examined (https://github.com/derrickstolee/gen-test, \n> > https://github.com/abhishekkumar2718/git/pull/1). The proposed generation\n> > number v2, Corrected Commit Date with Mononotically Increasing Offsets \n> > performed much worse than committer date (506,577 vs. 167,468 commits walked\n> > for git merge-base v4.8 v4.9) and was dropped.\n> >\n> > Using Generation Data chunk (GDAT) relieves the requirement of backward\n> > compatibility as we would continue to store topological levels in Commit\n> > Data (CDAT) chunk.\n> \n> Nice writeup about the history of generation number v2, much appreciated.\n> \n> >                    Thus, Corrected Commit Date was chosen as generation\n> > number v2. The Corrected Commit Date is defined as:\n> \n> Minor nitpick: it would be probably better to use \"is defined as\n> follows.\" instead of \"is defined as:\".\n> \n> >\n> > For a commit C, let its corrected commit date be the maximum of the commit\n> > date of C and the corrected commit dates of its parents plus 1. Then \n> > corrected commit date offset is the difference between corrected commit date\n> > of C and commit date of C. As a special case, a root commit with timestamp\n> > zero has corrected commit date of 1 to be able distinguish it from\n> > GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit date).\n> \n> Very minor nitpick: s/with timestamp/with *the* timestamp/, and\n> s/to be able distinguish/to be able *to* distinguish/ (without the '*'\n> used to mark the additions).\n> \n> >\n> > We will introduce an additional commit-graph chunk, Generation Data chunk,\n> \n> Or \"Generation DATa chunk\", if we want to emphasize where its name came\n> from, or even \"Generation DATa (GDAT) chunk\". But it is fine as it is\n> now, though it would be good idea to write \"Generation Data (GDAT)\n> chunk\" to explicitly state its name / shortcut.\n> \n> > and store corrected commit date offsets in GDAT chunk while storing\n> > topological levels in CDAT chunk. The old versions of Git would ignore GDAT\n> > chunk, using topological levels from CDAT chunk. In contrast, new versions\n> > of Git would use corrected commit dates, falling back to topological level\n> > if the generation data chunk is absent in the commit-graph file.\n> \n> Nice writeup of handling the backward compatibility.\n> \n> >\n> > While storing corrected commit date offsets saves us 4 bytes per commit (as\n> > compared with storing corrected commit dates directly), it's possible for\n> > the offset to overflow the space allocated. To handle such cases, we\n> > introduce a new chunk, Generation Data Overflow (GDOV) that stores the\n> > corrected commit date. For overflowing offsets, we set MSB and store the\n> > position into the GDOV chunk, in a mechanism similar to the Extra Edges list\n> > chunk.\n> \n> Very minor suggestion: perhaps it would be better to use \"it's however\n> possible\".\n> \n> Very minor suggestion: \"it's possible for the offset to overflow\" could\n> be simplified to just \"the offset can overflow\"... though the simplified\n> version loses a bit of hint that the overflow should be very rare in\n> real repositories.\n> \n> But it is just fine as it is now; I am not a native English speaker to\n> judge which version is better.\n> \n\nI think it is better to indicate the rareness of overflows.\n\n> >\n> > For mixed generation number environment (for example new Git on the command\n> > line, old Git used by GUI client), we can encounter a mixed-chain\n> > commit-graph (a commit-graph chain where some of split commit-graph files\n> > have GDAT chunk and others do not). As backward compatibility is one of the\n> > goals, we can define the following behavior:\n> >\n> > While reading a mixed-chain commit-graph version, we fall back on\n> > topological levels as corrected commit dates and topological levels cannot\n> > be compared directly.\n> >\n> > While writing on top of a split commit-graph, we check if the tip of the\n> > chain has a GDAT chunk. If it does, we append to the chain, writing GDAT\n> > chunk. Thus, we guarantee if the topmost split commit-graph file has a GDAT\n> > chunk, rest of the chain does too.\n> >\n> > If the topmost split commit-graph file does not have a GDAT chunk (meaning\n> > it has been appended by the old Git), we write without GDAT chunk. We do\n> > write a GDAT chunk when the existing chain does not have GDAT chunk - when\n> > we are writing to the commit-graph chain with the 'replace' strategy.\n> \n> I think the last paragraph can be simplified (or added to) by explicitly\n> stating the goal:\n> \n>   When adding new layer to the split commit-graph file, and when merging\n>   some or all layers (replacing them in the latter case), the new layer\n>   will have GDAT chunk if and only if in the final result there would be\n>   no layer without GDAT chunk just below it.\n> \n\nThanks, that is much clearer to understand.\n\n> ...\n> \n> After careful review of those 10 patches it looks like the series is\n> close to being ready, requiring only small changes to progress.\n> \n\nThank you for writing this handy reference for changes.\n\n> > Abhishek Kumar (10):\n> >   commit-graph: fix regression when computing Bloom filters\n> \n>     All good, beside possible improvement to the commit message.\n>     Thanks to Taylor Blau for discovering possible reason for strange\n>     no change in performance.\n> \n> >   revision: parse parent in indegree_walk_step()\n> \n>     Looks good.\n> \n> >   commit-graph: consolidate fill_commit_graph_info\n> \n>     Needs to fix now duplicated test names (minor change).\n>     Proposed possible improvement to the commit message.\n> \n> >   commit-graph: return 64-bit generation number\n> \n>     Needs fixing due to mismerge: there should be no switch from\n>     using GENERATION_NUMBER_ZERO to using GENERATION_NUMBER_INFINITY.\n>     Possible minor improvement to the commit message.\n> \n> >   commit-graph: add a slab to store topological levels\n> \n>     Possible minor improvement to the commit message.\n>     \n>     There is also not very important issue, but something that would be\n>     nice to explain, namely that checks for GENERATION_NUMBER_INFINITY \n>     can never be true, as topo_level_slab_at() returns 0 for commits\n>     outside the commit-graph, not GENERATION_NUMBER_INFINITY.  It works\n>     but it is not obvious why.\n> \n> >   commit-graph: implement corrected commit date\n> \n>     The change to commit-graph verification needs fixing, and we need to\n>     decide how verifying generation numbers should work.  Perhaps a test\n>     for handling topological level of GENERATION_NUMBER_V1_MAX could be\n>     added (though this might be left for ater).\n> \n>     The changes to `git commit-graph verify` code could be put into\n>     separate patch, either before or after this one.\n> \n> >   commit-graph: implement generation data chunk\n> \n>     Proposed possible improvement to the commit message.\n>     The commit message does not explain why given shape of history is\n>     needed to test handling corrected commit date offset overflow.\n> \n>     Proposed minor corrections to the coding style.\n> \n>     Instead of looping again through all commits when handling overflow\n>     in corrected commit date offsets, while there should be at most a\n>     few commits needing it, why not save those commits on list and loop\n>     only through those commits?  Though this _possible_ performance\n>     improvement could be left to the followup...\n\nSince the improvement can be applied to both \n`write_graph_chunk_generation_data_overflow()` and \n`write_graph_chunk_extra_edges()`, I am planning to cover this in a\nfollowup.\n\n> \n>     test_commit_with_date() could be instead implemented via adding\n>     `--date <date>` option to test_commit() in test-lib-functions.sh.\n> \n>     Also, to reduce \"noise\" in this patch, the rename of\n>     run_three_modes() to run_all_modes() and test_three_modes() to\n>     test_all_modes() could have been done in a separate preparatory\n>     patch. It would be pure refactoring patch, without introducing any\n>     new functionality.  But it is not something that is necessary.\n> \n> >   commit-graph: use generation v2 only if entire chain does\n> \n>     Proposed possible improvement to the commit message.\n>     Proposed minor corrections to the coding style (also in tests).\n> \n>     There is a question whether merging layers or replacing them should\n>     honor GIT_TEST_COMMIT_GRAPH_NO_GDAT.\n> \n>     Tests possibly could be made more strict, and check more things\n>     explicitly. One test we are missing is testing that merging layers\n>     is done correctly, namely that if we are merging layers in split\n>     commit-graph file, and the layer below the ones we are merging lacks\n>     GDAT chunk, then the result of the merge should also be without GDAT\n>     chunk -- but that might be left for later.\n> \n> >   commit-reach: use corrected commit dates in paint_down_to_common()\n> \n>     This patch consist of two slightly interleaved changes, which\n>     possibly could be separated: change to paint_down_to_common() and\n>     change to t6404-recursive-merge test.\n> \n>     In the commit message for the paint_down_to_common() we should\n>     explicitly mention 091f4cf3, which this one partially reverts.\n> \n>     Possible accidental change, question about function naming.\n> \n> >   doc: add corrected commit date info\n> \n>     Needs further improvements to the documentation, like adding\n>     \"[Optional]\" to chunk description, and leftover switching from\n>     \"generation numbers\" to \"topological levels\" in one place.\n> \n> ...\n> \n> Best,\n> -- \n> Jakub Narębski\n\nThanks\n- Abhishek\n"},{"id":"413052","messageId":"c4e817abf7dbcd6c99da404507ea940305c521b6.1609154168.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v5 01/11] commit-graph: fix regression when computing Bloom filters","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:15:58Z","receivedAt":"2020-12-28T11:17:04Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nBefore computing Bloom fitlers, the commit-graph machinery uses\ncommit_gen_cmp to sort commits by generation order for improved diff\nperformance. 3d11275505 (commit-graph: examine commits by generation\nnumber, 2020-03-30) claims that this sort can reduce the time spent to\ncompute Bloom filters by nearly half.\n\nBut since c49c82aa4c (commit: move members graph_pos, generation to a\nslab, 2020-06-17), this optimization is broken, since asking for a\n'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\nwhile writing.\n\nNot all hope is lost, though: 'commit_graph_generation()' falls back to\ncomparing commits by their date when they have equal generation number,\nand so since c49c82aa4c is purely a date comparision function. This\nheuristic is good enough that we don't seem to loose appreciable\nperformance while computing Bloom filters. Applying this patch (compared\nwith v2.29.1) speeds up computing Bloom filters by around ~4\nseconds.\n\nSo, avoid the useless 'commit_graph_generation()' while writing by\ninstead accessing the slab directly. This returns the newly-computed\ngeneration numbers, and allows us to avoid the heuristic by directly\ncomparing generation numbers.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 06f8dc1d896..caf823295f4 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n+\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n+\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"413053","messageId":"7645e0bcef08c8fd726148b3545c0ca3adeefdec.1609154168.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v5 02/11] revision: parse parent in indegree_walk_step()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:15:59Z","receivedAt":"2020-12-28T11:17:05Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn indegree_walk_step(), we add unvisited parents to the indegree queue.\nHowever, parents are not guaranteed to be parsed. As the indegree queue\nsorts by generation number, let's parse parents before inserting them to\nensure the correct priority order.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n revision.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/revision.c b/revision.c\nindex 9dff845bed6..de8e45f462f 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3373,6 +3373,9 @@ static void indegree_walk_step(struct rev_info *revs)\n \t\tstruct commit *parent = p->item;\n \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n \n+\t\tif (repo_parse_commit_gently(revs->repo, parent, 1) < 0)\n+\t\t\treturn;\n+\n \t\tif (*pi)\n \t\t\t(*pi)++;\n \t\telse\n-- \ngitgitgadget\n\n"},{"id":"413054","messageId":"ca646912b2b3c9f629c23460a40bda7f0e6d2b2f.1609154168.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v5 03/11] commit-graph: consolidate fill_commit_graph_info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:16:00Z","receivedAt":"2020-12-28T11:17:05Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nBoth fill_commit_graph_info() and fill_commit_in_graph() parse\ninformation present in commit data chunk. Let's simplify the\nimplementation by calling fill_commit_graph_info() within\nfill_commit_in_graph().\n\nfill_commit_graph_info() used to not load committer data from commit data\nchunk. However, with the upcoming switch to using corrected committer\ndate as generation number v2, we will have to load committer date to\ncompute generation number value anyway.\n\ne51217e15 (t5000: test tar files that overflow ustar headers,\n30-06-2016) introduced a test 'generate tar with future mtime' that\ncreates a commit with committer date of (2^36 + 1) seconds since\nEPOCH. The CDAT chunk provides 34-bits for storing committer date, thus\ncommitter time overflows into generation number (within CDAT chunk) and\nhas undefined behavior.\n\nThe test used to pass as fill_commit_graph_info() would not set struct\nmember `date` of struct commit and load committer date from the object\ndatabase, generating a tar file with the expected mtime.\n\nHowever, with corrected commit date, we will load the committer date\nfrom CDAT chunk (truncated to lower 34-bits to populate the generation\nnumber. Thus, Git sets date and generates tar file with the truncated\nmtime.\n\nThe ustar format (the header format used by most modern tar programs)\nonly has room for 11 (or 12, depending on some implementations) octal\ndigits for the size and mtime of each file.\n\nAs the CDAT chunk is overflow by 12-octal digits but not 11-octal\ndigits, we split the existing tests to test both implementations\nseparately and add a new explicit test for 11-digit implementation.\n\nTo test the 11-octal digit implementation, we create a future commit\nwith committer date of 2^34 - 1, which overflows 11-octal digits without\noverflowing 34-bits of the Commit Date chunks.\n\nTo test the 12-octal digit implementation, the smallest committer date\npossible is 2^36 + 1, which overflows the CDAT chunk and thus\ncommit-graph must be disabled for the test.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c      | 27 ++++++++++-----------------\n t/t5000-tar-tree.sh | 24 +++++++++++++++++++++---\n 2 files changed, 31 insertions(+), 20 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex caf823295f4..d5b33b4f7ac 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -749,15 +749,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n \tuint32_t lex_index;\n+\tuint64_t date_high, date_low;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n+\tif (pos >= g->num_commits + g->num_commits_in_base)\n+\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\n \tlex_index = pos - g->num_commits_in_base;\n \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n \n \tgraph_data = commit_graph_data_at(item);\n \tgraph_data->graph_pos = pos;\n+\n+\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n+\tdate_low = get_be32(commit_data + g->hash_len + 12);\n+\titem->date = (timestamp_t)((date_high << 32) | date_low);\n+\n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n@@ -772,38 +781,22 @@ static int fill_commit_in_graph(struct repository *r,\n {\n \tuint32_t edge_value;\n \tuint32_t *parent_data_ptr;\n-\tuint64_t date_low, date_high;\n \tstruct commit_list **pptr;\n-\tstruct commit_graph_data *graph_data;\n \tconst unsigned char *commit_data;\n \tuint32_t lex_index;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n-\tif (pos >= g->num_commits + g->num_commits_in_base)\n-\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\tfill_commit_graph_info(item, g, pos);\n \n-\t/*\n-\t * Store the \"full\" position, but then use the\n-\t * \"local\" position for the rest of the calculation.\n-\t */\n-\tgraph_data = commit_graph_data_at(item);\n-\tgraph_data->graph_pos = pos;\n \tlex_index = pos - g->num_commits_in_base;\n-\n \tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n \n \titem->object.parsed = 1;\n \n \tset_commit_tree(item, NULL);\n \n-\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n-\tdate_low = get_be32(commit_data + g->hash_len + 12);\n-\titem->date = (timestamp_t)((date_high << 32) | date_low);\n-\n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n-\n \tpptr = &item->parents;\n \n \tedge_value = get_be32(commit_data + g->hash_len);\ndiff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\nindex 3ebb0d3b652..7204799a0b5 100755\n--- a/t/t5000-tar-tree.sh\n+++ b/t/t5000-tar-tree.sh\n@@ -431,15 +431,33 @@ test_expect_success TAR_HUGE,LONG_IS_64BIT 'system tar can read our huge size' '\n \ttest_cmp expect actual\n '\n \n-test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n+test_expect_success TIME_IS_64BIT 'set up repository with far-future (2^34 - 1) commit' '\n+\trm -f .git/index &&\n+\techo foo >file &&\n+\tgit add file &&\n+\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n+\t\tgit commit -m \"tempori parendum\"\n+'\n+\n+test_expect_success TIME_IS_64BIT 'generate tar with far-future mtime' '\n+\tgit archive HEAD >future.tar\n+'\n+\n+test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n+\techo 2514 >expect &&\n+\ttar_info future.tar | cut -d\" \" -f2 >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success TIME_IS_64BIT 'set up repository with far-far-future (2^36 + 1) commit' '\n \trm -f .git/index &&\n \techo content >file &&\n \tgit add file &&\n-\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n+\tGIT_TEST_COMMIT_GRAPH=0 GIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n \t\tgit commit -m \"tempori parendum\"\n '\n \n-test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n+test_expect_success TIME_IS_64BIT 'generate tar with far-far-future mtime' '\n \tgit archive HEAD >future.tar\n '\n \n-- \ngitgitgadget\n\n"},{"id":"413055","messageId":"591935075f1dce264b24c91715c81ce1ff15fd47.1609154168.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v5 04/11] t6600-test-reach: generalize *_three_modes","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:16:01Z","receivedAt":"2020-12-28T11:17:05Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a preparatory step to implement generation number v2, we add tests to\nensure Git can read and parse commit-graph files without Generation Data\nchunk. These files represent commit-graph files written by Old Git and\nare neccesary for backward compatability.\n\nWe extend run_three_modes() and test_three_modes() to *_all_modes() with\nthe fourth mode being \"commit-graph without generation data chunk\".\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n t/t6600-test-reach.sh | 62 +++++++++++++++++++++----------------------\n 1 file changed, 31 insertions(+), 31 deletions(-)\n\ndiff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\nindex f807276337d..af10f0dc090 100755\n--- a/t/t6600-test-reach.sh\n+++ b/t/t6600-test-reach.sh\n@@ -58,7 +58,7 @@ test_expect_success 'setup' '\n \tgit config core.commitGraph true\n '\n \n-run_three_modes () {\n+run_all_modes () {\n \ttest_when_finished rm -rf .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n \ttest_cmp expect actual &&\n@@ -70,8 +70,8 @@ run_three_modes () {\n \ttest_cmp expect actual\n }\n \n-test_three_modes () {\n-\trun_three_modes test-tool reach \"$@\"\n+test_all_modes () {\n+\trun_all_modes test-tool reach \"$@\"\n }\n \n test_expect_success 'ref_newer:miss' '\n@@ -80,7 +80,7 @@ test_expect_success 'ref_newer:miss' '\n \tB:commit-4-9\n \tEOF\n \techo \"ref_newer(A,B):0\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'ref_newer:hit' '\n@@ -89,7 +89,7 @@ test_expect_success 'ref_newer:hit' '\n \tB:commit-2-3\n \tEOF\n \techo \"ref_newer(A,B):1\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'in_merge_bases:hit' '\n@@ -98,7 +98,7 @@ test_expect_success 'in_merge_bases:hit' '\n \tB:commit-8-8\n \tEOF\n \techo \"in_merge_bases(A,B):1\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'in_merge_bases:miss' '\n@@ -107,7 +107,7 @@ test_expect_success 'in_merge_bases:miss' '\n \tB:commit-5-9\n \tEOF\n \techo \"in_merge_bases(A,B):0\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'in_merge_bases_many:hit' '\n@@ -117,7 +117,7 @@ test_expect_success 'in_merge_bases_many:hit' '\n \tX:commit-5-7\n \tEOF\n \techo \"in_merge_bases_many(A,X):1\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'in_merge_bases_many:miss' '\n@@ -127,7 +127,7 @@ test_expect_success 'in_merge_bases_many:miss' '\n \tX:commit-8-6\n \tEOF\n \techo \"in_merge_bases_many(A,X):0\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'in_merge_bases_many:miss-heuristic' '\n@@ -137,7 +137,7 @@ test_expect_success 'in_merge_bases_many:miss-heuristic' '\n \tX:commit-6-6\n \tEOF\n \techo \"in_merge_bases_many(A,X):0\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'is_descendant_of:hit' '\n@@ -148,7 +148,7 @@ test_expect_success 'is_descendant_of:hit' '\n \tX:commit-1-1\n \tEOF\n \techo \"is_descendant_of(A,X):1\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'is_descendant_of:miss' '\n@@ -159,7 +159,7 @@ test_expect_success 'is_descendant_of:miss' '\n \tX:commit-7-6\n \tEOF\n \techo \"is_descendant_of(A,X):0\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'get_merge_bases_many' '\n@@ -174,7 +174,7 @@ test_expect_success 'get_merge_bases_many' '\n \t\tgit rev-parse commit-5-6 \\\n \t\t\t      commit-4-7 | sort\n \t} >expect &&\n-\ttest_three_modes get_merge_bases_many\n+\ttest_all_modes get_merge_bases_many\n '\n \n test_expect_success 'reduce_heads' '\n@@ -196,7 +196,7 @@ test_expect_success 'reduce_heads' '\n \t\t\t      commit-2-8 \\\n \t\t\t      commit-1-10 | sort\n \t} >expect &&\n-\ttest_three_modes reduce_heads\n+\ttest_all_modes reduce_heads\n '\n \n test_expect_success 'can_all_from_reach:hit' '\n@@ -219,7 +219,7 @@ test_expect_success 'can_all_from_reach:hit' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach(X,Y):1\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach:miss' '\n@@ -241,7 +241,7 @@ test_expect_success 'can_all_from_reach:miss' '\n \tY:commit-8-5\n \tEOF\n \techo \"can_all_from_reach(X,Y):0\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach_with_flag: tags case' '\n@@ -264,7 +264,7 @@ test_expect_success 'can_all_from_reach_with_flag: tags case' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach_with_flag(X,_,_,0,0):1\" >expect &&\n-\ttest_three_modes can_all_from_reach_with_flag\n+\ttest_all_modes can_all_from_reach_with_flag\n '\n \n test_expect_success 'commit_contains:hit' '\n@@ -280,8 +280,8 @@ test_expect_success 'commit_contains:hit' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):1\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'commit_contains:miss' '\n@@ -297,8 +297,8 @@ test_expect_success 'commit_contains:miss' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):0\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'rev-list: basic topo-order' '\n@@ -310,7 +310,7 @@ test_expect_success 'rev-list: basic topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 commit-3-2 commit-2-2 commit-1-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent topo-order' '\n@@ -322,7 +322,7 @@ test_expect_success 'rev-list: first-parent topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -334,7 +334,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -346,7 +346,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent range topo-order' '\n@@ -358,7 +358,7 @@ test_expect_success 'rev-list: first-parent range topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: ancestry-path topo-order' '\n@@ -368,7 +368,7 @@ test_expect_success 'rev-list: ancestry-path topo-order' '\n \t\tcommit-6-4 commit-5-4 commit-4-4 commit-3-4 \\\n \t\tcommit-6-3 commit-5-3 commit-4-3 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: symmetric difference topo-order' '\n@@ -382,7 +382,7 @@ test_expect_success 'rev-list: symmetric difference topo-order' '\n \t\tcommit-3-8 commit-2-8 commit-1-8 \\\n \t\tcommit-3-7 commit-2-7 commit-1-7 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8...commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8...commit-6-6\n '\n \n test_expect_success 'get_reachable_subset:all' '\n@@ -402,7 +402,7 @@ test_expect_success 'get_reachable_subset:all' '\n \t\t\t      commit-1-7 \\\n \t\t\t      commit-5-6 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:some' '\n@@ -420,7 +420,7 @@ test_expect_success 'get_reachable_subset:some' '\n \t\tgit rev-parse commit-3-3 \\\n \t\t\t      commit-1-7 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:none' '\n@@ -434,7 +434,7 @@ test_expect_success 'get_reachable_subset:none' '\n \tY:commit-2-8\n \tEOF\n \techo \"get_reachable_subset(X,Y)\" >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_done\n-- \ngitgitgadget\n\n"},{"id":"413056","messageId":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v4.git.1602079785.gitgitgadget@gmail.com","subject":"[PATCH v5 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:15:57Z","receivedAt":"2020-12-28T11:17:05Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"This patch series implements the corrected commit date offsets as generation\nnumber v2, along with other pre-requisites.\n\nGit uses topological levels in the commit-graph file for commit-graph\ntraversal operations like 'git log --graph'. Unfortunately, using\ntopological levels can result in a worse performance than without them when\ncompared with committer date as a heuristics. For example, 'git merge-base\nv4.8 v4.9' on the Linux repository walks 635,579 commits using topological\nlevels and walks 167,468 using committer date. Since 091f4cf3 (commit: don't\nuse generation numbers if not needed, 2018-08-30), 'git merge-base' uses\ncommitter date heuristic unless there is a cutoff because of the performance\nhit.\n\nThus, the need for generation number v2 was born. New generation number\nneeded to provide good performance, increment updates, and backward\ncompatibility. Due to an unfortunate problem [1], we also needed a way to\ndistinguish between the old and new generation number without incrementing\ngraph version.\n\n[1] https://public-inbox.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nVarious candidates were examined (https://github.com/derrickstolee/gen-test,\nhttps://github.com/abhishekkumar2718/git/pull/1). The proposed generation\nnumber v2, Corrected Commit Date with Mononotically Increasing Offsets\nperformed much worse than committer date (506,577 vs. 167,468 commits walked\nfor 'git merge-base v4.8 v4.9') and was dropped.\n\nUsing Generation Data chunk (GDAT) relieves the requirement of backward\ncompatibility as we would continue to store topological levels in Commit\nData (CDAT) chunk. Thus, Corrected Commit Date was chosen as generation\nnumber v2. The Corrected Commit Date is defined as follows:\n\nFor a commit C, let its corrected commit date be the maximum of the commit\ndate of C and the corrected commit dates of its parents plus 1. Then\ncorrected commit date offset is the difference between corrected commit date\nof C and commit date of C. As a special case, a root commit with the\ntimestamp zero has corrected commit date of 1 to be able to distinguish it\nfrom GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit date).\n\nWe will introduce an additional commit-graph chunk, Generation DATa (GDAT)\nchunk, and store corrected commit date offsets in GDAT chunk while storing\ntopological levels in CDAT chunk. The old versions of Git would ignore GDAT\nchunk, using topological levels from CDAT chunk. In contrast, new versions\nof Git would use corrected commit dates, falling back to topological level\nif the generation data chunk is absent in the commit-graph file.\n\nWhile storing corrected commit date offsets saves us 4 bytes per commit (as\ncompared with storing corrected commit dates directly), it's however\npossible for the offset to overflow the space allocated. To handle such\ncases, we introduce a new chunk, Generation Data Overflow (GDOV) that stores\nthe corrected commit date. For overflowing offsets, we set MSB and store the\nposition into the GDOV chunk, in a mechanism similar to the Extra Edges list\nchunk.\n\nFor mixed generation number environment (for example new Git on the command\nline, old Git used by GUI client), we can encounter a mixed-chain\ncommit-graph (a commit-graph chain where some of split commit-graph files\nhave GDAT chunk and others do not). As backward compatibility is one of the\ngoals, we can define the following behavior:\n\nWhile reading a mixed-chain commit-graph version, we fall back on\ntopological levels as corrected commit dates and topological levels cannot\nbe compared directly.\n\nWhen adding new layer to the split commit-graph file, and when merging some\nor all layers (replacing them in the latter case), the new layer will have\nGDAT chunk if and only if in the final result there would be no layer\nwithout GDAT chunk just below it.\n\nThanks to Dr. Stolee, Dr. Narębski, and Taylor for their reviews.\n\nI look forward to everyone's reviews!\n\nThanks\n\n * Abhishek\n\n----------------------------------------------------------------------------\n\nImprovements left for a future series:\n\n * Save commits with generation data overflow and extra edge commits instead\n   of looping over all commits. cf. 858sbel67n.fsf@gmail.com\n * Verify both topological levels and corrected commit dates when present.\n   cf. 85pn4tnk8u.fsf@gmail.com\n\nChanges in version 5:\n\n * Explained a possible reason for no change in performance for\n   \"commit-graph: fix regression when computing bloom-filters\"\n * Clarified about the addition of a new test for 11-digit octal\n   implementations of ustar.\n * Fixed duplicate test names in \"commit-graph: consolidate\n   fill_commit_graph_info\".\n * Swapped the order \"commit-graph: return 64-bit generation number\",\n   \"commit-graph: add a slab to store topological levels\" to minimize lines\n   changed.\n * Fixed the mismerge in \"commit-graph: return 64-bit generation number\"\n * Clarified the preparatory steps are for the larger goal of implementing\n   generation number v2 in \"commit-graph: return 64-bit generation number\".\n * Moved the rename of \"run_three_modes()\" to \"run_all_modes()\" into a new\n   patch \"t6600-test-reach: generalize *_three_modes\".\n * Explained and removed the checks for GENERATION_NUMBER_INFINITY that can\n   never be true in \"commit-graph: add a slab to store topological levels\".\n * Fixed incorrect logic for verifying commit-graph in \"commit-graph:\n   implement corrected commit date\".\n * Added minor improvements to commit message of \"commit-graph: implement\n   generation data chunk\".\n * Added '--date ' option to test_commit() in 'test-lib-functions.sh' in\n   \"commit-graph: implement generation data chunk\".\n * Improved coding style (also in tests) for \"commit-graph: use generation\n   v2 only if entire chain does\".\n * Simplified test repository structure in \"commit-graph: use generation v2\n   only if entire chain does\" as only the number of commits in a split\n   commit-graph layer are relevant.\n * Added a new test in \"commit-graph: use generation v2 only if entire chain\n   does\" to check if the layers are merged correctly.\n * Explicitly mentioned commit \"091f4cf3\" in the commit-message of\n   \"commit-graph: use corrected commit dates in paint_down_to_common()\".\n * Minor corrections to documentation in \"doc: add corrected commit date\n   info\".\n * Minor corrections to coding style.\n\nChanges in version 4:\n\n * Added GDOV to handle overflows in generation data.\n * Added a test for writing tip graph for a generation number v2 graph chain\n   in t5324-split-commit-graph.sh\n * Added a section on how mixed generation number chains are handled in\n   Documentation/technical/commit-graph-format.txt\n * Reverted unimportant whitespace, style changes in commit-graph.c\n * Added header comments about the order of comparision for\n   compare_commits_by_gen_then_commit_date in commit.h,\n   compare_commits_by_gen in commit-graph.h\n * Elaborated on why t6404 fails with corrected commit date and must be run\n   with GIT_TEST_COMMIT_GRAPH=1in the commit \"commit-reach: use corrected\n   commit dates in paint_down_to_common()\"\n * Elaborated on write behavior for mixed generation number chains in the\n   commit \"commit-graph: use generation v2 only if entire chain does\"\n * Added notes about adding the topo_level slab to struct\n   write_commit_graph_context as well as struct commit_graph.\n * Clarified commit message for \"commit-graph: consolidate\n   fill_commit_graph_info\"\n * Removed the claim \"GDAT can store future generation numbers\" because it\n   hasn't been tested yet.\n\nChanges in version 3:\n\n * Reordered patches as discussed in 2\n   [https://lore.kernel.org/git/aee0ae56-3395-6848-d573-27a318d72755@gmail.com/].\n * Split \"implement corrected commit date\" into two patches - one\n   introducing the topo level slab and other implementing corrected commit\n   dates.\n * Extended split-commit-graph tests to verify at the end of test.\n * Use topological levels as generation number if any of split commit-graph\n   files do not have generation data chunk.\n\nChanges in version 2:\n\n * Add tests for generation data chunk.\n * Add an option GIT_TEST_COMMIT_GRAPH_NO_GDAT to control whether to write\n   generation data chunk.\n * Compare commits with corrected commit dates if present in\n   paint_down_to_common().\n * Update technical documentation.\n * Handle mixed generation commit chains.\n * Improve commit messages for \"commit-graph: fix regression when computing\n   bloom filter\", \"commit-graph: consolidate fill_commit_graph_info\",\n * Revert unnecessary whitespace changes.\n * Split uint_32 -> timestamp_t change into a new commit.\n\nAbhishek Kumar (11):\n  commit-graph: fix regression when computing Bloom filters\n  revision: parse parent in indegree_walk_step()\n  commit-graph: consolidate fill_commit_graph_info\n  t6600-test-reach: generalize *_three_modes\n  commit-graph: add a slab to store topological levels\n  commit-graph: return 64-bit generation number\n  commit-graph: implement corrected commit date\n  commit-graph: implement generation data chunk\n  commit-graph: use generation v2 only if entire chain does\n  commit-reach: use corrected commit dates in paint_down_to_common()\n  doc: add corrected commit date info\n\n .../technical/commit-graph-format.txt         |  28 +-\n Documentation/technical/commit-graph.txt      |  77 +++++-\n commit-graph.c                                | 243 ++++++++++++++----\n commit-graph.h                                |  15 +-\n commit-reach.c                                |  38 +--\n commit-reach.h                                |   2 +-\n commit.c                                      |   4 +-\n commit.h                                      |   5 +-\n revision.c                                    |  13 +-\n t/README                                      |   3 +\n t/helper/test-read-graph.c                    |   4 +\n t/t4216-log-bloom.sh                          |   4 +-\n t/t5000-tar-tree.sh                           |  24 +-\n t/t5318-commit-graph.sh                       |  79 +++++-\n t/t5324-split-commit-graph.sh                 | 193 +++++++++++++-\n t/t6404-recursive-merge.sh                    |   5 +-\n t/t6600-test-reach.sh                         |  68 ++---\n t/test-lib-functions.sh                       |   6 +\n upload-pack.c                                 |   2 +-\n 19 files changed, 659 insertions(+), 154 deletions(-)\n\n\nbase-commit: 4a0de43f4923993377dbbc42cfc0a1054b6c5ccf\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-676%2Fabhishekkumar2718%2Fcorrected_commit_date-v5\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-676/abhishekkumar2718/corrected_commit_date-v5\nPull-Request: https://github.com/gitgitgadget/git/pull/676\n\nRange-diff vs v4:\n\n  1:  fae81b534b1 !  1:  c4e817abf7d commit-graph: fix regression when computing Bloom filters\n     @@ Metadata\n       ## Commit message ##\n          commit-graph: fix regression when computing Bloom filters\n      \n     -    commit_gen_cmp is used when writing a commit-graph to sort commits in\n     -    generation order before computing Bloom filters. Since c49c82aa (commit:\n     -    move members graph_pos, generation to a slab, 2020-06-17) made it so\n     -    that 'commit_graph_generation()' returns 'GENERATION_NUMBER_INFINITY'\n     -    during writing, we cannot call it within this function. Instead, access\n     -    the generation number directly through the slab (i.e., by calling\n     -    'commit_graph_data_at(c)->generation') in order to access it while\n     -    writing.\n     +    Before computing Bloom fitlers, the commit-graph machinery uses\n     +    commit_gen_cmp to sort commits by generation order for improved diff\n     +    performance. 3d11275505 (commit-graph: examine commits by generation\n     +    number, 2020-03-30) claims that this sort can reduce the time spent to\n     +    compute Bloom filters by nearly half.\n      \n     -    While measuring performance with `git commit-graph write --reachable\n     -    --changed-paths` on the linux repository led to around 1m40s for both\n     -    HEAD and master (and could be due to fault in my measurements), it is\n     -    still the \"right\" thing to do.\n     +    But since c49c82aa4c (commit: move members graph_pos, generation to a\n     +    slab, 2020-06-17), this optimization is broken, since asking for a\n     +    'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\n     +    while writing.\n     +\n     +    Not all hope is lost, though: 'commit_graph_generation()' falls back to\n     +    comparing commits by their date when they have equal generation number,\n     +    and so since c49c82aa4c is purely a date comparision function. This\n     +    heuristic is good enough that we don't seem to loose appreciable\n     +    performance while computing Bloom filters. Applying this patch (compared\n     +    with v2.29.1) speeds up computing Bloom filters by around ~4\n     +    seconds.\n     +\n     +    So, avoid the useless 'commit_graph_generation()' while writing by\n     +    instead accessing the slab directly. This returns the newly-computed\n     +    generation numbers, and allows us to avoid the heuristic by directly\n     +    comparing generation numbers.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n  2:  4470d916428 =  2:  7645e0bcef0 revision: parse parent in indegree_walk_step()\n  3:  18bb3318a12 !  3:  ca646912b2b commit-graph: consolidate fill_commit_graph_info\n     @@ Commit message\n          fill_commit_in_graph().\n      \n          fill_commit_graph_info() used to not load committer data from commit data\n     -    chunk. However, with the corrected committer date, we have to load\n     -    committer date to calculate generation number value.\n     +    chunk. However, with the upcoming switch to using corrected committer\n     +    date as generation number v2, we will have to load committer date to\n     +    compute generation number value anyway.\n      \n          e51217e15 (t5000: test tar files that overflow ustar headers,\n          30-06-2016) introduced a test 'generate tar with future mtime' that\n     -    creates a commit with committer date of (2 ^ 36 + 1) seconds since\n     +    creates a commit with committer date of (2^36 + 1) seconds since\n          EPOCH. The CDAT chunk provides 34-bits for storing committer date, thus\n          committer time overflows into generation number (within CDAT chunk) and\n          has undefined behavior.\n      \n          The test used to pass as fill_commit_graph_info() would not set struct\n     -    member `date` of struct commit and loads committer date from the object\n     +    member `date` of struct commit and load committer date from the object\n          database, generating a tar file with the expected mtime.\n      \n          However, with corrected commit date, we will load the committer date\n     @@ Commit message\n          mtime.\n      \n          The ustar format (the header format used by most modern tar programs)\n     -    only has room for 11 (or 12, depending om some implementations) octal\n     -    digits for the size and mtime of each files.\n     +    only has room for 11 (or 12, depending on some implementations) octal\n     +    digits for the size and mtime of each file.\n      \n     -    Thus, setting a timestamp of 2 ^ 33 + 1 would overflow the 11-octal\n     -    digit implementations while still fitting into commit data chunk.\n     +    As the CDAT chunk is overflow by 12-octal digits but not 11-octal\n     +    digits, we split the existing tests to test both implementations\n     +    separately and add a new explicit test for 11-digit implementation.\n      \n     -    Since we want to test 12-octal digit implementations of ustar as well,\n     -    let's modify the existing test to no longer use commit-graph file.\n     +    To test the 11-octal digit implementation, we create a future commit\n     +    with committer date of 2^34 - 1, which overflows 11-octal digits without\n     +    overflowing 34-bits of the Commit Date chunks.\n     +\n     +    To test the 12-octal digit implementation, the smallest committer date\n     +    possible is 2^36 + 1, which overflows the CDAT chunk and thus\n     +    commit-graph must be disabled for the test.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n     @@ t/t5000-tar-tree.sh: test_expect_success TAR_HUGE,LONG_IS_64BIT 'system tar can\n       \ttest_cmp expect actual\n       '\n       \n     -+test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n     +-test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n     ++test_expect_success TIME_IS_64BIT 'set up repository with far-future (2^34 - 1) commit' '\n      +\trm -f .git/index &&\n      +\techo foo >file &&\n      +\tgit add file &&\n     @@ t/t5000-tar-tree.sh: test_expect_success TAR_HUGE,LONG_IS_64BIT 'system tar can\n      +\t\tgit commit -m \"tempori parendum\"\n      +'\n      +\n     -+test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n     ++test_expect_success TIME_IS_64BIT 'generate tar with far-future mtime' '\n      +\tgit archive HEAD >future.tar\n      +'\n      +\n     @@ t/t5000-tar-tree.sh: test_expect_success TAR_HUGE,LONG_IS_64BIT 'system tar can\n      +\ttest_cmp expect actual\n      +'\n      +\n     - test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n     ++test_expect_success TIME_IS_64BIT 'set up repository with far-far-future (2^36 + 1) commit' '\n       \trm -f .git/index &&\n       \techo content >file &&\n       \tgit add file &&\n     @@ t/t5000-tar-tree.sh: test_expect_success TAR_HUGE,LONG_IS_64BIT 'system tar can\n       \t\tgit commit -m \"tempori parendum\"\n       '\n       \n     +-test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n     ++test_expect_success TIME_IS_64BIT 'generate tar with far-far-future mtime' '\n     + \tgit archive HEAD >future.tar\n     + '\n     + \n  -:  ----------- >  4:  591935075f1 t6600-test-reach: generalize *_three_modes\n  5:  e067f653ad5 !  5:  baae7006764 commit-graph: add a slab to store topological levels\n     @@ Commit message\n          commit-graph: add a slab to store topological levels\n      \n          In a later commit we will introduce corrected commit date as the\n     -    generation number v2. This value will be stored in the new seperate\n     -    Generation Data chunk. However, to ensure backwards compatibility with\n     -    \"Old\" Git we need to continue to write generation number v1, which is\n     -    topological level, to the commit data chunk. This means that we need to\n     -    compute both versions of generation numbers when writing the\n     -    commit-graph file. Therefore, let's introduce a commit-slab to store\n     +    generation number v2. Corrected commit dates will be stored in the new\n     +    seperate Generation Data chunk. However, to ensure backwards\n     +    compatibility with \"Old\" Git we need to continue to write generation\n     +    number v1 (topological levels) to the commit data chunk. Thus, we need\n     +    to compute and store both versions of generation numbers to write the\n     +    commit-graph file.\n     +\n     +    Therefore, let's introduce a commit-slab `topo_level_slab` to store\n          topological levels; corrected commit date will be stored in the member\n          `generation` of struct commit_graph_data.\n      \n     -    When Git creates a split commit-graph, it takes advantage of the\n     -    generation values that have been computed already and present in\n     -    existing commit-graph files.\n     +    The macros `GENERATION_NUMBER_INFINITY` and `GENERATION_NUMBER_ZERO`\n     +    mark commits not in the commit-graph file and commits written by a\n     +    version of Git that did not compute generation numbers respectively.\n     +    Generation numbers are computed identically for both kinds of commits.\n     +\n     +    A \"slab-miss\" should return `GENERATION_NUMBER_INFINITY` as the commit\n     +    is not in the commit-graph file. However, since the slab is\n     +    zero-initialized, it returns 0 (or rather `GENERATION_NUMBER_ZERO`).\n     +    Thus, we no longer need to check if the topological level of a commit is\n     +    `GENERATION_NUMBER_INFINITY`.\n      \n     -    So, let's add a pointer to struct commit_graph as well as struct\n     -    write_commit_graph_context to the topological level commit-slab\n     -    and populate it with topological levels while writing a commit-graph\n     -    file.\n     +    We will add a pointer to the slab in `struct write_commit_graph_context`\n     +    and `struct commit_graph` to populate the slab in\n     +    `fill_commit_graph_info` if the commit has a pre-computed topological\n     +    level as in case of split commit-graphs.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n       \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n       \t\t\t\t\tctx->commits.nr);\n       \tfor (i = 0; i < ctx->commits.nr; i++) {\n     --\t\ttimestamp_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n     -+\t\ttimestamp_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n     +-\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n     ++\t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n       \n       \t\tdisplay_progress(ctx->progress, i + 1);\n      -\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n      -\t\t    generation != GENERATION_NUMBER_ZERO)\n     -+\t\tif (level != GENERATION_NUMBER_INFINITY &&\n     -+\t\t    level != GENERATION_NUMBER_ZERO)\n     ++\t\tif (level != GENERATION_NUMBER_ZERO)\n       \t\t\tcontinue;\n       \n       \t\tcommit_list_insert(ctx->commits.list[i], &list);\n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n       \n      -\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n      -\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n     -+\t\t\t\tif (level == GENERATION_NUMBER_INFINITY ||\n     -+\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n     ++\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n       \t\t\t\t\tall_parents_computed = 0;\n       \t\t\t\t\tcommit_list_insert(parent->item, &list);\n       \t\t\t\t\tbreak;\n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n      -\t\t\t\tdata->generation = max_generation + 1;\n       \t\t\t\tpop_commit(&list);\n       \n     --\t\t\t\tif (data->generation > GENERATION_NUMBER_V1_MAX)\n     --\t\t\t\t\tdata->generation = GENERATION_NUMBER_V1_MAX;\n     -+\t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n     -+\t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n     +-\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n     +-\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n     ++\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n     ++\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n      +\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n       \t\t\t}\n       \t\t}\n     @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n      +\tstruct topo_level_slab topo_levels;\n       \n     - \tif (!commit_graph_compatible(the_repository))\n     - \t\treturn 0;\n     + \tprepare_repo_settings(the_repository);\n     + \tif (!the_repository->settings.core_commit_graph) {\n      @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \t\t\t\t\t\t\t bloom_settings.max_changed_paths);\n       \tctx->bloom_settings = &bloom_settings;\n  4:  011b0aa497d !  6:  26bd6f49100 commit-graph: return 64-bit generation number\n     @@ Metadata\n       ## Commit message ##\n          commit-graph: return 64-bit generation number\n      \n     -    In a preparatory step, let's return timestamp_t values from\n     -    commit_graph_generation(), use timestamp_t for local variables and\n     -    define GENERATION_NUMBER_INFINITY as (2 ^ 63 - 1) instead.\n     +    In a preparatory step for introducing corrected commit dates, let's\n     +    return timestamp_t values from commit_graph_generation(), use\n     +    timestamp_t for local variables and define GENERATION_NUMBER_INFINITY\n     +    as (2 ^ 63 - 1) instead.\n      \n          We rename GENERATION_NUMBER_MAX to GENERATION_NUMBER_V1_MAX to\n          represent the largest topological level we can store in the commit data\n     @@ commit-graph.c: static int commit_gen_cmp(const void *va, const void *vb)\n       \tif (generation_a < generation_b)\n       \t\treturn -1;\n      @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n     - \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n     - \t\t\t\t\tctx->commits.nr);\n     - \tfor (i = 0; i < ctx->commits.nr; i++) {\n     --\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n     -+\t\ttimestamp_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n     - \n     - \t\tdisplay_progress(ctx->progress, i + 1);\n     - \t\tif (generation != GENERATION_NUMBER_INFINITY &&\n     -@@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n     - \t\t\t\tdata->generation = max_generation + 1;\n     + \t\t\tif (all_parents_computed) {\n       \t\t\t\tpop_commit(&list);\n       \n     --\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n     --\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n     -+\t\t\t\tif (data->generation > GENERATION_NUMBER_V1_MAX)\n     -+\t\t\t\t\tdata->generation = GENERATION_NUMBER_V1_MAX;\n     +-\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n     +-\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n     ++\t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n     ++\t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n     + \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n       \t\t\t}\n       \t\t}\n     - \t}\n      @@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n       \tfor (i = 0; i < g->num_commits; i++) {\n       \t\tstruct commit *graph_commit, *odb_commit;\n     @@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_grap\n       \t\t\tmax_generation--;\n       \n       \t\tgeneration = commit_graph_generation(graph_commit);\n     + \t\tif (generation != max_generation + 1)\n     +-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n     ++\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" != %\"PRItime),\n     + \t\t\t\t     oid_to_hex(&cur_oid),\n     + \t\t\t\t     generation,\n     + \t\t\t\t     max_generation + 1);\n      \n       ## commit-graph.h ##\n      @@ commit-graph.h: void disable_commit_graph(struct repository *r);\n     @@ commit-reach.c: int repo_in_merge_bases_many(struct repository *r, struct commit\n       \tstruct commit_list *bases;\n       \tint ret = 0, i;\n      -\tuint32_t generation, max_generation = GENERATION_NUMBER_ZERO;\n     -+\ttimestamp_t generation, max_generation = GENERATION_NUMBER_INFINITY;\n     ++\ttimestamp_t generation, max_generation = GENERATION_NUMBER_ZERO;\n       \n       \tif (repo_parse_commit(r, commit))\n       \t\treturn ret;\n  6:  694ef1ec08d !  7:  859c39eff52 commit-graph: implement corrected commit date\n     @@ Commit message\n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n       ## commit-graph.c ##\n     -@@ commit-graph.c: static int commit_gen_cmp(const void *va, const void *vb)\n     - \telse if (generation_a > generation_b)\n     - \t\treturn 1;\n     - \n     --\t/* use date as a heuristic when generations are equal */\n     --\tif (a->date < b->date)\n     --\t\treturn -1;\n     --\telse if (a->date > b->date)\n     --\t\treturn 1;\n     - \treturn 0;\n     - }\n     - \n      @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n       \t\t\t\t\tctx->commits.nr);\n       \tfor (i = 0; i < ctx->commits.nr; i++) {\n     - \t\ttimestamp_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n     + \t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n      +\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n       \n       \t\tdisplay_progress(ctx->progress, i + 1);\n     - \t\tif (level != GENERATION_NUMBER_INFINITY &&\n     --\t\t    level != GENERATION_NUMBER_ZERO)\n     -+\t\t    level != GENERATION_NUMBER_ZERO &&\n     -+\t\t    corrected_commit_date != GENERATION_NUMBER_INFINITY &&\n     -+\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO\n     -+\t\t    )\n     +-\t\tif (level != GENERATION_NUMBER_ZERO)\n     ++\t\tif (level != GENERATION_NUMBER_ZERO &&\n     ++\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO)\n       \t\t\tcontinue;\n       \n       \t\tcommit_list_insert(ctx->commits.list[i], &list);\n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n       \n       \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n       \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n     --\n      +\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n     - \t\t\t\tif (level == GENERATION_NUMBER_INFINITY ||\n     --\t\t\t\t    level == GENERATION_NUMBER_ZERO) {\n     -+\t\t\t\t    level == GENERATION_NUMBER_ZERO ||\n     -+\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_INFINITY ||\n     -+\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO\n     -+\t\t\t\t    ) {\n     + \n     +-\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n     ++\t\t\t\tif (level == GENERATION_NUMBER_ZERO ||\n     ++\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO) {\n       \t\t\t\t\tall_parents_computed = 0;\n       \t\t\t\t\tcommit_list_insert(parent->item, &list);\n       \t\t\t\t\tbreak;\n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n       \t\t\t}\n       \t\t}\n       \t}\n     -@@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n     - \t\tif (generation_zero == GENERATION_ZERO_EXISTS)\n     - \t\t\tcontinue;\n     - \n     --\t\t/*\n     --\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n     --\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n     --\t\t * extra logic in the following condition.\n     --\t\t */\n     --\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n     --\t\t\tmax_generation--;\n     --\n     - \t\tgeneration = commit_graph_generation(graph_commit);\n     --\t\tif (generation != max_generation + 1)\n     --\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n     -+\t\tif (generation < max_generation + 1)\n     -+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n     - \t\t\t\t     oid_to_hex(&cur_oid),\n     - \t\t\t\t     generation,\n     - \t\t\t\t     max_generation + 1);\n  7:  b903efe2ea1 !  8:  8403c4d0257 commit-graph: implement generation data chunk\n     @@ Commit message\n      \n          As discovered by Ævar, we cannot increment graph version to\n          distinguish between generation numbers v1 and v2 [1]. Thus, one of\n     -    pre-requistes before implementing generation number was to distinguish\n     -    between graph versions in a backwards compatible manner.\n     +    pre-requistes before implementing generation number v2 was to\n     +    distinguish between graph versions in a backwards compatible manner.\n      \n     -    We are going to introduce a new chunk called Generation Data chunk (or\n     -    GDAT). GDAT stores corrected committer date offsets whereas CDAT will\n     -    still store topological level.\n     +    We are going to introduce a new chunk called Generation DATa chunk (or\n     +    GDAT). GDAT will store corrected committer date offsets whereas CDAT\n     +    will still store topological level.\n      \n          Old Git does not understand GDAT chunk and would ignore it, reading\n          topological levels from CDAT. New Git can parse GDAT and take advantage\n          of newer generation numbers, falling back to topological levels when\n     -    GDAT chunk is missing (as it would happen with a commit graph written\n     +    GDAT chunk is missing (as it would happen with a commit-graph written\n          by old Git).\n      \n          We introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\n          which forces commit-graph file to be written without generation data\n          chunk to emulate a commit-graph file written by old Git.\n      \n     -    While storing corrected commit date offset instead of the corrected\n     -    commit date saves us 4 bytes per commit, it's possible for the offsets\n     -    to overflow the 4-bytes allocated. As such overflows are exceedingly\n     -    rare, we use the following overflow management scheme:\n     +    To minimize the space required to store corrrected commit date, Git\n     +    stores corrected commit date offsets into the commit-graph file, instea\n     +    of corrected commit dates. This saves us 4 bytes per commit, decreasing\n     +    the GDAT chunk size by half, but it's possible for the offset to\n     +    overflow the 4-bytes allocated for storage. As such overflows are and\n     +    should be exceedingly rare, we use the following overflow management\n     +    scheme:\n      \n     -    We introduce a new commit-graph chunk, GENERATION_DATA_OVERFLOW ('GDOV')\n     +    We introduce a new commit-graph chunk, Generation Data OVerflow ('GDOV')\n          to store corrected commit dates for commits with offsets greater than\n          GENERATION_NUMBER_V2_OFFSET_MAX.\n      \n     @@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct c\n       \n      -\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n      +\tif (g->chunk_generation_data) {\n     -+\t\toffset = (timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     ++\t\toffset = (timestamp_t)get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n      +\n      +\t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n      +\t\t\toffset_pos = offset ^ CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW;\n     @@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct c\n       \tif (g->topo_levels)\n       \t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n      @@ commit-graph.c: struct write_commit_graph_context {\n     - \tstruct packed_oid_list oids;\n     + \tstruct oid_array oids;\n       \tstruct packed_commit_list commits;\n       \tint num_extra_edges;\n      +\tint num_generation_data_overflows;\n     @@ commit-graph.c: static int write_graph_chunk_data(struct hashfile *f,\n      +\t\t\t\t\t      struct write_commit_graph_context *ctx)\n      +{\n      +\tint i, num_generation_data_overflows = 0;\n     ++\n      +\tfor (i = 0; i < ctx->commits.nr; i++) {\n      +\t\tstruct commit *c = ctx->commits.list[i];\n      +\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n     @@ commit-graph.c: static int write_graph_chunk_data(struct hashfile *f,\n       \t\t\t\t\t struct write_commit_graph_context *ctx)\n       {\n      @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n     - \n       \t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n       \t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n     -+\n       \t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n      +\n      +\t\t\t\tif (commit_graph_data_at(current)->generation - current->date > GENERATION_NUMBER_V2_OFFSET_MAX)\n     @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \n       \tbloom_settings.bits_per_entry = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\",\n       \t\t\t\t\t\t      bloom_settings.bits_per_entry);\n     +@@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n     + \t\t\tcontinue;\n     + \n     + \t\t/*\n     +-\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n     +-\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n     +-\t\t * extra logic in the following condition.\n     ++\t\t * If we are using topological level and one of our parents has\n     ++\t\t * generation GENERATION_NUMBER_V1_MAX, then our generation is\n     ++\t\t * also GENERATION_NUMBER_V1_MAX. Decrement to avoid extra logic\n     ++\t\t * in the following condition.\n     + \t\t */\n     +-\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n     ++\t\tif (!g->chunk_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n     + \t\t\tmax_generation--;\n     + \n     + \t\tgeneration = commit_graph_generation(graph_commit);\n     +-\t\tif (generation != max_generation + 1)\n     +-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" != %\"PRItime),\n     ++\t\tif (generation < max_generation + 1)\n     ++\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n     + \t\t\t\t     oid_to_hex(&cur_oid),\n     + \t\t\t\t     generation,\n     + \t\t\t\t     max_generation + 1);\n      \n       ## commit-graph.h ##\n      @@\n     @@ t/t5318-commit-graph.sh: test_expect_success 'corrupt commit-graph write (missin\n       \t)\n       '\n       \n     -+test_commit_with_date() {\n     -+  file=\"$1.t\" &&\n     -+  echo \"$1\" >\"$file\" &&\n     -+  git add \"$file\" &&\n     -+  GIT_COMMITTER_DATE=\"$2\" GIT_AUTHOR_DATE=\"$2\" git commit -m \"$1\"\n     -+  git tag \"$1\"\n     -+}\n     ++# We test the overflow-related code with the following repo history:\n     ++#\n     ++#               4:F - 5:N - 6:U\n     ++#              /                \\\n     ++# 1:U - 2:N - 3:U                M:N\n     ++#              \\                /\n     ++#               7:N - 8:F - 9:N\n     ++#\n     ++# Here the commits denoted by U have committer date of zero seconds\n     ++# since Unix epoch, the commits denoted by N have committer date\n     ++# starting from 1112354055 seconds since Unix epoch (default committer\n     ++# date for the test suite), and the commits denoted by F have committer\n     ++# date of (2 ^ 31 - 2) seconds since Unix epoch.\n     ++#\n     ++# The largest offset observed is 2 ^ 31, just large enough to overflow.\n     ++#\n      +\n     -+test_expect_success 'overflow corrected commit date offset' '\n     ++test_expect_success 'set up and verify repo with generation data overflow chunk' '\n      +\tobjdir=\".git/objects\" &&\n     -+\tUNIX_EPOCH_ZERO=\"1970-01-01 00:00 +0000\" &&\n     ++\tUNIX_EPOCH_ZERO=\"@0 +0000\" &&\n      +\tFUTURE_DATE=\"@2147483646 +0000\" &&\n      +\ttest_oid_cache <<-EOF &&\n      +\toid_version sha1:1\n     @@ t/t5318-commit-graph.sh: test_expect_success 'corrupt commit-graph write (missin\n      +\tmkdir repo &&\n      +\tcd repo &&\n      +\tgit init &&\n     -+\ttest_commit_with_date 1 \"$UNIX_EPOCH_ZERO\" &&\n     ++\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 1 &&\n      +\ttest_commit 2 &&\n     -+\ttest_commit_with_date 3 \"$UNIX_EPOCH_ZERO\" &&\n     ++\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 3 &&\n      +\tgit commit-graph write --reachable &&\n      +\tgraph_read_expect 3 generation_data &&\n     -+\ttest_commit_with_date 4 \"$FUTURE_DATE\" &&\n     ++\ttest_commit --date \"$FUTURE_DATE\" 4 &&\n      +\ttest_commit 5 &&\n     -+\ttest_commit_with_date 6 \"$UNIX_EPOCH_ZERO\" &&\n     ++\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 6 &&\n      +\tgit branch left &&\n      +\tgit reset --hard 3 &&\n      +\ttest_commit 7 &&\n     -+\ttest_commit_with_date 8 \"$FUTURE_DATE\" &&\n     ++\ttest_commit --date \"$FUTURE_DATE\" 8 &&\n      +\ttest_commit 9 &&\n      +\tgit branch right &&\n      +\tgit reset --hard 3 &&\n     -+\tgit merge left right &&\n     ++\ttest_merge M left right &&\n      +\tgit commit-graph write --reachable &&\n      +\tgraph_read_expect 10 \"generation_data generation_data_overflow\" &&\n      +\tgit commit-graph verify\n      +'\n      +\n     -+graph_git_behavior 'overflow corrected commit date offset' repo left right\n     ++graph_git_behavior 'generation data overflow chunk repo' repo left right\n      +\n       test_done\n      \n     @@ t/t6600-test-reach.sh: test_expect_success 'setup' '\n       \tgit config core.commitGraph true\n       '\n       \n     --run_three_modes () {\n     -+run_all_modes () {\n     - \ttest_when_finished rm -rf .git/objects/info/commit-graph &&\n     - \t\"$@\" <input >actual &&\n     - \ttest_cmp expect actual &&\n     -@@ t/t6600-test-reach.sh: run_three_modes () {\n     +@@ t/t6600-test-reach.sh: run_all_modes () {\n       \ttest_cmp expect actual &&\n       \tcp commit-graph-half .git/objects/info/commit-graph &&\n       \t\"$@\" <input >actual &&\n     @@ t/t6600-test-reach.sh: run_three_modes () {\n       \ttest_cmp expect actual\n       }\n       \n     --test_three_modes () {\n     --\trun_three_modes test-tool reach \"$@\"\n     -+test_all_modes () {\n     -+\trun_all_modes test-tool reach \"$@\"\n     - }\n     - \n     - test_expect_success 'ref_newer:miss' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'ref_newer:miss' '\n     - \tB:commit-4-9\n     - \tEOF\n     - \techo \"ref_newer(A,B):0\" >expect &&\n     --\ttest_three_modes ref_newer\n     -+\ttest_all_modes ref_newer\n     - '\n     - \n     - test_expect_success 'ref_newer:hit' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'ref_newer:hit' '\n     - \tB:commit-2-3\n     - \tEOF\n     - \techo \"ref_newer(A,B):1\" >expect &&\n     --\ttest_three_modes ref_newer\n     -+\ttest_all_modes ref_newer\n     - '\n     - \n     - test_expect_success 'in_merge_bases:hit' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'in_merge_bases:hit' '\n     - \tB:commit-8-8\n     - \tEOF\n     - \techo \"in_merge_bases(A,B):1\" >expect &&\n     --\ttest_three_modes in_merge_bases\n     -+\ttest_all_modes in_merge_bases\n     - '\n     - \n     - test_expect_success 'in_merge_bases:miss' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'in_merge_bases:miss' '\n     - \tB:commit-5-9\n     - \tEOF\n     - \techo \"in_merge_bases(A,B):0\" >expect &&\n     --\ttest_three_modes in_merge_bases\n     -+\ttest_all_modes in_merge_bases\n     - '\n     - \n     - test_expect_success 'in_merge_bases_many:hit' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'in_merge_bases_many:hit' '\n     - \tX:commit-5-7\n     - \tEOF\n     - \techo \"in_merge_bases_many(A,X):1\" >expect &&\n     --\ttest_three_modes in_merge_bases_many\n     -+\ttest_all_modes in_merge_bases_many\n     - '\n     - \n     - test_expect_success 'in_merge_bases_many:miss' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'in_merge_bases_many:miss' '\n     - \tX:commit-8-6\n     - \tEOF\n     - \techo \"in_merge_bases_many(A,X):0\" >expect &&\n     --\ttest_three_modes in_merge_bases_many\n     -+\ttest_all_modes in_merge_bases_many\n     - '\n     - \n     - test_expect_success 'in_merge_bases_many:miss-heuristic' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'in_merge_bases_many:miss-heuristic' '\n     - \tX:commit-6-6\n     - \tEOF\n     - \techo \"in_merge_bases_many(A,X):0\" >expect &&\n     --\ttest_three_modes in_merge_bases_many\n     -+\ttest_all_modes in_merge_bases_many\n     - '\n     - \n     - test_expect_success 'is_descendant_of:hit' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'is_descendant_of:hit' '\n     - \tX:commit-1-1\n     - \tEOF\n     - \techo \"is_descendant_of(A,X):1\" >expect &&\n     --\ttest_three_modes is_descendant_of\n     -+\ttest_all_modes is_descendant_of\n     - '\n     - \n     - test_expect_success 'is_descendant_of:miss' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'is_descendant_of:miss' '\n     - \tX:commit-7-6\n     - \tEOF\n     - \techo \"is_descendant_of(A,X):0\" >expect &&\n     --\ttest_three_modes is_descendant_of\n     -+\ttest_all_modes is_descendant_of\n     - '\n     - \n     - test_expect_success 'get_merge_bases_many' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'get_merge_bases_many' '\n     - \t\tgit rev-parse commit-5-6 \\\n     - \t\t\t      commit-4-7 | sort\n     - \t} >expect &&\n     --\ttest_three_modes get_merge_bases_many\n     -+\ttest_all_modes get_merge_bases_many\n     - '\n     - \n     - test_expect_success 'reduce_heads' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'reduce_heads' '\n     - \t\t\t      commit-2-8 \\\n     - \t\t\t      commit-1-10 | sort\n     - \t} >expect &&\n     --\ttest_three_modes reduce_heads\n     -+\ttest_all_modes reduce_heads\n     - '\n     - \n     - test_expect_success 'can_all_from_reach:hit' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'can_all_from_reach:hit' '\n     - \tY:commit-8-1\n     - \tEOF\n     - \techo \"can_all_from_reach(X,Y):1\" >expect &&\n     --\ttest_three_modes can_all_from_reach\n     -+\ttest_all_modes can_all_from_reach\n     - '\n     - \n     - test_expect_success 'can_all_from_reach:miss' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'can_all_from_reach:miss' '\n     - \tY:commit-8-5\n     - \tEOF\n     - \techo \"can_all_from_reach(X,Y):0\" >expect &&\n     --\ttest_three_modes can_all_from_reach\n     -+\ttest_all_modes can_all_from_reach\n     - '\n     - \n     - test_expect_success 'can_all_from_reach_with_flag: tags case' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'can_all_from_reach_with_flag: tags case' '\n     - \tY:commit-8-1\n     - \tEOF\n     - \techo \"can_all_from_reach_with_flag(X,_,_,0,0):1\" >expect &&\n     --\ttest_three_modes can_all_from_reach_with_flag\n     -+\ttest_all_modes can_all_from_reach_with_flag\n     - '\n     - \n     - test_expect_success 'commit_contains:hit' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'commit_contains:hit' '\n     - \tX:commit-9-3\n     - \tEOF\n     - \techo \"commit_contains(_,A,X,_):1\" >expect &&\n     --\ttest_three_modes commit_contains &&\n     --\ttest_three_modes commit_contains --tag\n     -+\ttest_all_modes commit_contains &&\n     -+\ttest_all_modes commit_contains --tag\n     - '\n     - \n     - test_expect_success 'commit_contains:miss' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'commit_contains:miss' '\n     - \tX:commit-9-3\n     - \tEOF\n     - \techo \"commit_contains(_,A,X,_):0\" >expect &&\n     --\ttest_three_modes commit_contains &&\n     --\ttest_three_modes commit_contains --tag\n     -+\ttest_all_modes commit_contains &&\n     -+\ttest_all_modes commit_contains --tag\n     - '\n     - \n     - test_expect_success 'rev-list: basic topo-order' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: basic topo-order' '\n     - \t\tcommit-6-2 commit-5-2 commit-4-2 commit-3-2 commit-2-2 commit-1-2 \\\n     - \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n     - \t>expect &&\n     --\trun_three_modes git rev-list --topo-order commit-6-6\n     -+\trun_all_modes git rev-list --topo-order commit-6-6\n     - '\n     - \n     - test_expect_success 'rev-list: first-parent topo-order' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: first-parent topo-order' '\n     - \t\tcommit-6-2 \\\n     - \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n     - \t>expect &&\n     --\trun_three_modes git rev-list --first-parent --topo-order commit-6-6\n     -+\trun_all_modes git rev-list --first-parent --topo-order commit-6-6\n     - '\n     - \n     - test_expect_success 'rev-list: range topo-order' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: range topo-order' '\n     - \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n     - \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n     - \t>expect &&\n     --\trun_three_modes git rev-list --topo-order commit-3-3..commit-6-6\n     -+\trun_all_modes git rev-list --topo-order commit-3-3..commit-6-6\n     - '\n     - \n     - test_expect_success 'rev-list: range topo-order' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: range topo-order' '\n     - \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n     - \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n     - \t>expect &&\n     --\trun_three_modes git rev-list --topo-order commit-3-8..commit-6-6\n     -+\trun_all_modes git rev-list --topo-order commit-3-8..commit-6-6\n     - '\n     - \n     - test_expect_success 'rev-list: first-parent range topo-order' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: first-parent range topo-order' '\n     - \t\tcommit-6-2 \\\n     - \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n     - \t>expect &&\n     --\trun_three_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n     -+\trun_all_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n     - '\n     - \n     - test_expect_success 'rev-list: ancestry-path topo-order' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: ancestry-path topo-order' '\n     - \t\tcommit-6-4 commit-5-4 commit-4-4 commit-3-4 \\\n     - \t\tcommit-6-3 commit-5-3 commit-4-3 \\\n     - \t>expect &&\n     --\trun_three_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n     -+\trun_all_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n     - '\n     - \n     - test_expect_success 'rev-list: symmetric difference topo-order' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'rev-list: symmetric difference topo-order' '\n     - \t\tcommit-3-8 commit-2-8 commit-1-8 \\\n     - \t\tcommit-3-7 commit-2-7 commit-1-7 \\\n     - \t>expect &&\n     --\trun_three_modes git rev-list --topo-order commit-3-8...commit-6-6\n     -+\trun_all_modes git rev-list --topo-order commit-3-8...commit-6-6\n     - '\n     - \n     - test_expect_success 'get_reachable_subset:all' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'get_reachable_subset:all' '\n     - \t\t\t      commit-1-7 \\\n     - \t\t\t      commit-5-6 | sort\n     - \t) >expect &&\n     --\ttest_three_modes get_reachable_subset\n     -+\ttest_all_modes get_reachable_subset\n     - '\n     - \n     - test_expect_success 'get_reachable_subset:some' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'get_reachable_subset:some' '\n     - \t\tgit rev-parse commit-3-3 \\\n     - \t\t\t      commit-1-7 | sort\n     - \t) >expect &&\n     --\ttest_three_modes get_reachable_subset\n     -+\ttest_all_modes get_reachable_subset\n     - '\n     - \n     - test_expect_success 'get_reachable_subset:none' '\n     -@@ t/t6600-test-reach.sh: test_expect_success 'get_reachable_subset:none' '\n     - \tY:commit-2-8\n     - \tEOF\n     - \techo \"get_reachable_subset(X,Y)\" >expect &&\n     --\ttest_three_modes get_reachable_subset\n     -+\ttest_all_modes get_reachable_subset\n     - '\n     - \n     - test_done\n     +\n     + ## t/test-lib-functions.sh ##\n     +@@ t/test-lib-functions.sh: test_commit () {\n     + \t\t--signoff)\n     + \t\t\tsignoff=\"$1\"\n     + \t\t\t;;\n     ++\t\t--date)\n     ++\t\t\tnotick=yes\n     ++\t\t\tGIT_COMMITTER_DATE=\"$2\"\n     ++\t\t\tGIT_AUTHOR_DATE=\"$2\"\n     ++\t\t\tshift\n     ++\t\t\t;;\n     + \t\t-C)\n     + \t\t\tindir=\"$2\"\n     + \t\t\tshift\n  8:  8ec119edc66 !  9:  a3a70a1edd0 commit-graph: use generation v2 only if entire chain does\n     @@ Commit message\n          1. \"New\" Git writes a commit-graph with the GDAT chunk.\n          2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n      \n     -    Because of the current use of inspecting the current layer for a\n     -    chunk_generation_data pointer, the commits in the lower layer will be\n     -    interpreted as having very large generation values (commit date plus\n     -    offset) compared to the generation numbers in the top layer (topological\n     -    level). This violates the expectation that the generation of a parent is\n     -    strictly smaller than the generation of a child.\n     +    If each layer of split commit-graph is treated independently, as it was\n     +    the case before this commit, with Git inspecting only the current layer\n     +    for chunk_generation_data pointer, commits in the lower layer (one with\n     +    GDAT) whould have corrected commit date as their generation number,\n     +    while commits in the upper layer would have topological levels as their\n     +    generation. Corrected commit dates usually have much larger values than\n     +    topological levels. This means that if we take two commits, one from the\n     +    upper layer, and one reachable from it in the lower layer, then the\n     +    expectation that the generation of a parent is smaller than the\n     +    generation of a child would be violated.\n      \n          It is difficult to expose this issue in a test. Since we _start_ with\n          artificially low generation numbers, any commit walk that prioritizes\n     @@ Commit message\n          commits in the lower layer before allowing the topo-order queue to write\n          anything to output (depending on the size of the upper layer).\n      \n     -    When writing the new layer in split commit-graph, we write a GDAT chunk\n     -    only if the topmost layer has a GDAT chunk. This guarantees that if a\n     -    layer has GDAT chunk, all lower layers must have a GDAT chunk as well.\n     +    Therefore, When writing the new layer in split commit-graph, we write a\n     +    GDAT chunk only if the topmost layer has a GDAT chunk. This guarantees\n     +    that if a layer has GDAT chunk, all lower layers must have a GDAT chunk\n     +    as well.\n      \n          Rewriting layers follows similar approach: if the topmost layer below\n          the set of layers being rewritten (in the split commit-graph chain)\n     @@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct c\n       \titem->date = (timestamp_t)((date_high << 32) | date_low);\n       \n      -\tif (g->chunk_generation_data) {\n     -+\tif (g->chunk_generation_data && g->read_generation_data) {\n     - \t\toffset = (timestamp_t) get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n     ++\tif (g->read_generation_data) {\n     + \t\toffset = (timestamp_t)get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n       \n       \t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n      @@ commit-graph.c: static void split_graph_merge_strategy(struct write_commit_graph_context *ctx)\n     - \t\t}\n     - \t}\n     + \t\tif (i < ctx->num_commit_graphs_after)\n     + \t\t\tctx->commit_graph_hash_after[i] = xstrdup(oid_to_hex(&g->oid));\n       \n     -+\tif (!ctx->write_generation_data && g->chunk_generation_data)\n     -+\t\tctx->write_generation_data = 1;\n     ++\t\t/*\n     ++\t\t * If the topmost remaining layer has generation data chunk, the\n     ++\t\t * resultant layer also has generation data chunk.\n     ++\t\t */\n     ++\t\tif (i == ctx->num_commit_graphs_after - 2)\n     ++\t\t\tctx->write_generation_data = !!g->chunk_generation_data;\n      +\n     - \tif (flags != COMMIT_GRAPH_SPLIT_REPLACE)\n     - \t\tctx->new_base_graph = g;\n     - \telse if (ctx->num_commit_graphs_after != 1)\n     + \t\ti--;\n     + \t\tg = g->base_graph;\n     + \t}\n      @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n       \n     @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n       \t\t\tg->topo_levels = &topo_levels;\n       \t\t\tg = g->base_graph;\n       \t\t}\n     -@@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n     - \n     - \t\tg = ctx->r->objects->commit_graph;\n     +@@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n     + \t\t * also GENERATION_NUMBER_V1_MAX. Decrement to avoid extra logic\n     + \t\t * in the following condition.\n     + \t\t */\n     +-\t\tif (!g->chunk_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n     ++\t\tif (!g->read_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n     + \t\t\tmax_generation--;\n       \n     -+\t\tif (g && !g->chunk_generation_data)\n     -+\t\t\tctx->write_generation_data = 0;\n     -+\n     - \t\twhile (g) {\n     - \t\t\tctx->num_commit_graphs_before++;\n     - \t\t\tg = g->base_graph;\n     -@@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n     - \n     - \t\tif (ctx->opts)\n     - \t\t\treplace = ctx->opts->split_flags & COMMIT_GRAPH_SPLIT_REPLACE;\n     -+\n     -+\t\tif (replace)\n     -+\t\t\tctx->write_generation_data = 1;\n     - \t}\n     - \n     - \tctx->approx_nr_objects = approximate_object_count();\n     + \t\tgeneration = commit_graph_generation(graph_commit);\n      \n       ## commit-graph.h ##\n      @@ commit-graph.h: struct commit_graph {\n     @@ commit-graph.h: struct commit_graph {\n       \tconst uint32_t *chunk_oid_fanout;\n      \n       ## t/t5324-split-commit-graph.sh ##\n     -@@ t/t5324-split-commit-graph.sh: test_expect_success '--split=replace with partial Bloom data' '\n     - \tverify_chain_files_exist $graphdir\n     +@@ t/t5324-split-commit-graph.sh: test_expect_success 'prevent regression for duplicate commits across layers' '\n     + \tgit -C dup commit-graph verify\n       '\n       \n     ++NUM_FIRST_LAYER_COMMITS=64\n     ++NUM_SECOND_LAYER_COMMITS=16\n     ++NUM_THIRD_LAYER_COMMITS=7\n     ++NUM_FOURTH_LAYER_COMMITS=8\n     ++NUM_FIFTH_LAYER_COMMITS=16\n     ++SECOND_LAYER_SEQUENCE_START=$(($NUM_FIRST_LAYER_COMMITS + 1))\n     ++SECOND_LAYER_SEQUENCE_END=$(($SECOND_LAYER_SEQUENCE_START + $NUM_SECOND_LAYER_COMMITS - 1))\n     ++THIRD_LAYER_SEQUENCE_START=$(($SECOND_LAYER_SEQUENCE_END + 1))\n     ++THIRD_LAYER_SEQUENCE_END=$(($THIRD_LAYER_SEQUENCE_START + $NUM_THIRD_LAYER_COMMITS - 1))\n     ++FOURTH_LAYER_SEQUENCE_START=$(($THIRD_LAYER_SEQUENCE_END + 1))\n     ++FOURTH_LAYER_SEQUENCE_END=$(($FOURTH_LAYER_SEQUENCE_START + $NUM_FOURTH_LAYER_COMMITS - 1))\n     ++FIFTH_LAYER_SEQUENCE_START=$(($FOURTH_LAYER_SEQUENCE_END + 1))\n     ++FIFTH_LAYER_SEQUENCE_END=$(($FIFTH_LAYER_SEQUENCE_START + $NUM_FIFTH_LAYER_COMMITS - 1))\n     ++\n     ++# Current split graph chain:\n     ++#\n     ++#     16 commits (No GDAT)\n     ++# ------------------------\n     ++#     64 commits (GDAT)\n     ++#\n      +test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n     -+\tmkdir mixed &&\n      +\tgraphdir=\".git/objects/info/commit-graphs\" &&\n     -+\ttest_oid_cache <<-EOM &&\n     ++\ttest_oid_cache <<-EOF &&\n      +\toid_version sha1:1\n      +\toid_version sha256:2\n     -+\tEOM\n     -+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n     -+\tgit init &&\n     -+\tgit config core.commitGraph true &&\n     -+\tgit config gc.writeCommitGraph false &&\n     -+\tfor i in $(test_seq 3)\n     -+\tdo\n     -+\t\ttest_commit $i &&\n     -+\t\tgit branch commits/$i || return 1\n     -+\tdone &&\n     -+\tgit reset --hard commits/1 &&\n     -+\tfor i in $(test_seq 4 5)\n     -+\tdo\n     -+\t\ttest_commit $i &&\n     -+\t\tgit branch commits/$i || return 1\n     -+\tdone &&\n     -+\tgit reset --hard commits/2 &&\n     -+\tfor i in $(test_seq 6 10)\n     -+\tdo\n     -+\t\ttest_commit $i &&\n     -+\t\tgit branch commits/$i || return 1\n     -+\tdone &&\n     -+\tgit commit-graph write --reachable --split &&\n     -+\tgit reset --hard commits/2 &&\n     -+\tgit merge commits/4 &&\n     -+\tgit branch merge/1 &&\n     -+\tgit reset --hard commits/4 &&\n     -+\tgit merge commits/6 &&\n     -+\tgit branch merge/2 &&\n     -+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n     -+\ttest-tool read-graph >output &&\n     -+\tcat >expect <<-EOF &&\n     -+\theader: 43475048 1 $(test_oid oid_version) 4 1\n     -+\tnum_commits: 2\n     -+\tchunks: oid_fanout oid_lookup commit_metadata\n      +\tEOF\n     -+\ttest_cmp expect output &&\n     -+\tgit commit-graph verify\n     ++\tgit init mixed &&\n     ++\t(\n     ++\t\tcd mixed &&\n     ++\t\tgit config core.commitGraph true &&\n     ++\t\tgit config gc.writeCommitGraph false &&\n     ++\t\tfor i in $(test_seq $NUM_FIRST_LAYER_COMMITS)\n     ++\t\tdo\n     ++\t\t\ttest_commit $i &&\n     ++\t\t\tgit branch commits/$i || return 1\n     ++\t\tdone &&\n     ++\t\tgit commit-graph write --reachable --split &&\n     ++\t\tgraph_read_expect $NUM_FIRST_LAYER_COMMITS &&\n     ++\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n     ++\t\tfor i in $(test_seq $SECOND_LAYER_SEQUENCE_START $SECOND_LAYER_SEQUENCE_END)\n     ++\t\tdo\n     ++\t\t\ttest_commit $i &&\n     ++\t\t\tgit branch commits/$i || return 1\n     ++\t\tdone &&\n     ++\t\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n     ++\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n     ++\t\ttest-tool read-graph >output &&\n     ++\t\tcat >expect <<-EOF &&\n     ++\t\theader: 43475048 1 $(test_oid oid_version) 4 1\n     ++\t\tnum_commits: $NUM_SECOND_LAYER_COMMITS\n     ++\t\tchunks: oid_fanout oid_lookup commit_metadata\n     ++\t\tEOF\n     ++\t\ttest_cmp expect output &&\n     ++\t\tgit commit-graph verify &&\n     ++\t\tcat $graphdir/commit-graph-chain\n     ++\t)\n      +'\n      +\n     -+test_expect_success 'does not write generation data chunk if not present on existing tip' '\n     -+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n     -+\tgit reset --hard commits/3 &&\n     -+\tgit merge merge/1 &&\n     -+\tgit merge commits/5 &&\n     -+\tgit merge merge/2 &&\n     -+\tgit branch merge/3 &&\n     -+\tgit commit-graph write --reachable --split=no-merge &&\n     -+\ttest-tool read-graph >output &&\n     -+\tcat >expect <<-EOF &&\n     -+\theader: 43475048 1 $(test_oid oid_version) 4 2\n     -+\tnum_commits: 3\n     -+\tchunks: oid_fanout oid_lookup commit_metadata\n     -+\tEOF\n     -+\ttest_cmp expect output &&\n     -+\tgit commit-graph verify\n     ++# The new layer will be added without generation data chunk as it was not\n     ++# present on the layer underneath it.\n     ++#\n     ++#      7 commits (No GDAT)\n     ++# ------------------------\n     ++#     16 commits (No GDAT)\n     ++# ------------------------\n     ++#     64 commits (GDAT)\n     ++#\n     ++test_expect_success 'do not write generation data chunk if not present on existing tip' '\n     ++\tgit clone mixed mixed-no-gdat &&\n     ++\t(\n     ++\t\tcd mixed-no-gdat &&\n     ++\t\tfor i in $(test_seq $THIRD_LAYER_SEQUENCE_START $THIRD_LAYER_SEQUENCE_END)\n     ++\t\tdo\n     ++\t\t\ttest_commit $i &&\n     ++\t\t\tgit branch commits/$i || return 1\n     ++\t\tdone &&\n     ++\t\tgit commit-graph write --reachable --split=no-merge &&\n     ++\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n     ++\t\ttest-tool read-graph >output &&\n     ++\t\tcat >expect <<-EOF &&\n     ++\t\theader: 43475048 1 $(test_oid oid_version) 4 2\n     ++\t\tnum_commits: $NUM_THIRD_LAYER_COMMITS\n     ++\t\tchunks: oid_fanout oid_lookup commit_metadata\n     ++\t\tEOF\n     ++\t\ttest_cmp expect output &&\n     ++\t\tgit commit-graph verify\n     ++\t)\n     ++'\n     ++\n     ++# Number of commits in each layer of the split-commit graph before merge:\n     ++#\n     ++#      8 commits (No GDAT)\n     ++# ------------------------\n     ++#      7 commits (No GDAT)\n     ++# ------------------------\n     ++#     16 commits (No GDAT)\n     ++# ------------------------\n     ++#     64 commits (GDAT)\n     ++#\n     ++# The top two layers are merged and do not have generation data chunk as layer below them does\n     ++# not have generation data chunk.\n     ++#\n     ++#     15 commits (No GDAT)\n     ++# ------------------------\n     ++#     16 commits (No GDAT)\n     ++# ------------------------\n     ++#     64 commits (GDAT)\n     ++#\n     ++test_expect_success 'do not write generation data chunk if the topmost remaining layer does not have generation data chunk' '\n     ++\tgit clone mixed-no-gdat mixed-merge-no-gdat &&\n     ++\t(\n     ++\t\tcd mixed-merge-no-gdat &&\n     ++\t\tfor i in $(test_seq $FOURTH_LAYER_SEQUENCE_START $FOURTH_LAYER_SEQUENCE_END)\n     ++\t\tdo\n     ++\t\t\ttest_commit $i &&\n     ++\t\t\tgit branch commits/$i || return 1\n     ++\t\tdone &&\n     ++\t\tgit commit-graph write --reachable --split --size-multiple 1 &&\n     ++\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n     ++\t\ttest-tool read-graph >output &&\n     ++\t\tcat >expect <<-EOF &&\n     ++\t\theader: 43475048 1 $(test_oid oid_version) 4 2\n     ++\t\tnum_commits: $(($NUM_THIRD_LAYER_COMMITS + $NUM_FOURTH_LAYER_COMMITS))\n     ++\t\tchunks: oid_fanout oid_lookup commit_metadata\n     ++\t\tEOF\n     ++\t\ttest_cmp expect output &&\n     ++\t\tgit commit-graph verify\n     ++\t)\n      +'\n      +\n     -+test_expect_success 'writes generation data chunk when commit-graph chain is replaced' '\n     -+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n     -+\tgit commit-graph write --reachable --split=replace &&\n     -+\ttest_path_is_file $graphdir/commit-graph-chain &&\n     -+\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n     -+\tverify_chain_files_exist $graphdir &&\n     -+\tgraph_read_expect 15 &&\n     -+\tgit commit-graph verify\n     ++# Number of commits in each layer of the split-commit graph before merge:\n     ++#\n     ++#     16 commits (No GDAT)\n     ++# ------------------------\n     ++#     15 commits (No GDAT)\n     ++# ------------------------\n     ++#     16 commits (No GDAT)\n     ++# ------------------------\n     ++#     64 commits (GDAT)\n     ++#\n     ++# The top three layers are merged and has generation data chunk as the topmost remaining layer\n     ++# has generation data chunk.\n     ++#\n     ++#     47 commits (GDAT)\n     ++# ------------------------\n     ++#     64 commits (GDAT)\n     ++#\n     ++test_expect_success 'write generation data chunk if topmost remaining layer has generation data chunk' '\n     ++\tgit clone mixed-merge-no-gdat mixed-merge-gdat &&\n     ++\t(\n     ++\t\tcd mixed-merge-gdat &&\n     ++\t\tfor i in $(test_seq $FIFTH_LAYER_SEQUENCE_START $FIFTH_LAYER_SEQUENCE_END)\n     ++\t\tdo\n     ++\t\t\ttest_commit $i &&\n     ++\t\t\tgit branch commits/$i || return 1\n     ++\t\tdone &&\n     ++\t\tgit commit-graph write --reachable --split --size-multiple 1 &&\n     ++\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n     ++\t\ttest-tool read-graph >output &&\n     ++\t\tcat >expect <<-EOF &&\n     ++\t\theader: 43475048 1 $(test_oid oid_version) 5 1\n     ++\t\tnum_commits: $(($NUM_SECOND_LAYER_COMMITS + $NUM_THIRD_LAYER_COMMITS + $NUM_FOURTH_LAYER_COMMITS + $NUM_FIFTH_LAYER_COMMITS))\n     ++\t\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n     ++\t\tEOF\n     ++\t\ttest_cmp expect output\n     ++\t)\n      +'\n      +\n     -+test_expect_success 'add one commit, write a tip graph' '\n     -+\tcd \"$TRASH_DIRECTORY/mixed\" &&\n     -+\ttest_commit 11 &&\n     -+\tgit branch commits/11 &&\n     -+\tgit commit-graph write --reachable --split &&\n     -+\ttest_path_is_missing $infodir/commit-graph &&\n     -+\ttest_path_is_file $graphdir/commit-graph-chain &&\n     -+\tls $graphdir/graph-*.graph >graph-files &&\n     -+\ttest_line_count = 2 graph-files &&\n     -+\tverify_chain_files_exist $graphdir\n     ++test_expect_success 'write generation data chunk when commit-graph chain is replaced' '\n     ++\tgit clone mixed mixed-replace &&\n     ++\t(\n     ++\t\tcd mixed-replace &&\n     ++\t\tgit commit-graph write --reachable --split=replace &&\n     ++\t\ttest_path_is_file $graphdir/commit-graph-chain &&\n     ++\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n     ++\t\tverify_chain_files_exist $graphdir &&\n     ++\t\tgraph_read_expect $(($NUM_FIRST_LAYER_COMMITS + $NUM_SECOND_LAYER_COMMITS)) &&\n     ++\t\tgit commit-graph verify\n     ++\t)\n      +'\n      +\n       test_done\n  9:  bb9b02af32d ! 10:  093101f908b commit-reach: use corrected commit dates in paint_down_to_common()\n     @@ Metadata\n       ## Commit message ##\n          commit-reach: use corrected commit dates in paint_down_to_common()\n      \n     -    With corrected commit dates implemented, we no longer have to rely on\n     -    commit date as a heuristic in paint_down_to_common().\n     +    091f4cf (commit: don't use generation numbers if not needed,\n     +    2018-08-30) changed paint_down_to_common() to use commit dates instead\n     +    of generation numbers v1 (topological levels) as the performance\n     +    regressed on certain topologies. With generation number v2 (corrected\n     +    commit dates) implemented, we no longer have to rely on commit dates and\n     +    can use generation numbers.\n      \n     -    While using corrected commit dates Git walks nearly the same number of\n     +    For example, the command `git merge-base v4.8 v4.9` on the Linux\n     +    repository walks 167468 commits, taking 0.135s for committer date and\n     +    167496 commits, taking 0.157s for corrected committer date respectively.\n     +\n     +    While using corrected commit dates, Git walks nearly the same number of\n          commits as commit date, the process is slower as for each comparision we\n          have to access a commit-slab (for corrected committer date) instead of\n          accessing struct member (for committer date).\n      \n     -    For example, the command `git merge-base v4.8 v4.9` on the linux\n     -    repository walks 167468 commits, taking 0.135s for committer date and\n     -    167496 commits, taking 0.157s for corrected committer date respectively.\n     -\n     -    t6404-recursive-merge setups a unique repository where all commits have\n     -    the same committer date without well-defined merge-base.\n     +    This change incidentally broke the fragile t6404-recursive-merge test.\n     +    t6404-recursive-merge sets up a unique repository where all commits have\n     +    the same committer date without a well-defined merge-base.\n      \n          While running tests with GIT_TEST_COMMIT_GRAPH unset, we use committer\n          date as a heuristic in paint_down_to_common(). 6404.1 'combined merge\n          conflicts' merges commits in the order:\n     -    - Merge C with B to form a intermediate commit.\n     +    - Merge C with B to form an intermediate commit.\n          - Merge the intermediate commit with A.\n      \n          With GIT_TEST_COMMIT_GRAPH=1, we write a commit-graph and subsequently\n          use the corrected committer date, which changes the order in which\n          commits are merged:\n     -    - Merge A with B to form a intermediate commit.\n     +    - Merge A with B to form an intermediate commit.\n          - Merge the intermediate commit with C.\n      \n          While resulting repositories are equivalent, 6404.4 'virtual trees were\n     @@ commit-graph.c: int generation_numbers_enabled(struct repository *r)\n       \tstruct commit_graph *g = r->objects->commit_graph;\n      \n       ## commit-graph.h ##\n     -@@ commit-graph.h: struct commit_graph *read_commit_graph_one(struct repository *r,\n     - struct commit_graph *parse_commit_graph(struct repository *r,\n     - \t\t\t\t\tvoid *graph_map, size_t graph_size);\n     - \n     -+struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n     -+\n     - /*\n     -  * Return 1 if and only if the repository has a commit-graph\n     -  * file and generation numbers are computed in that file.\n     +@@ commit-graph.h: struct commit_graph *parse_commit_graph(struct repository *r,\n        */\n       int generation_numbers_enabled(struct repository *r);\n       \n     --struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n      +/*\n      + * Return 1 if and only if the repository has a commit-graph\n      + * file and generation data chunk has been written for the file.\n      + */\n      +int corrected_commit_dates_enabled(struct repository *r);\n     ++\n     + struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n       \n       enum commit_graph_write_flags {\n     - \tCOMMIT_GRAPH_WRITE_APPEND     = (1 << 0),\n      \n       ## commit-reach.c ##\n      @@ commit-reach.c: static struct commit_list *paint_down_to_common(struct repository *r,\n 10:  9ada43967d2 ! 11:  20299e57457 doc: add corrected commit date info\n     @@ Documentation/technical/commit-graph-format.txt: CHUNK DATA:\n             2 bits of the lowest byte, storing the 33rd and 34th bit of the\n             commit time.\n       \n     -+  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes)\n     ++  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes) [Optional]\n      +    * This list of 4-byte values store corrected commit date offsets for the\n      +      commits, arranged in the same order as commit data chunk.\n      +    * If the corrected commit date offset cannot be stored within 31 bits,\n      +      the value has its most-significant bit on and the other bits store\n      +      the position of corrected commit date into the Generation Data Overflow\n      +      chunk.\n     ++    * Generation Data chunk is present only when commit-graph file is written\n     ++      by compatible versions of Git and in case of split commit-graph chains,\n     ++      the topmost layer also has Generation Data chunk.\n      +\n      +  Generation Data Overflow (ID: {'G', 'D', 'O', 'V' }) [Optional]\n     -+    * This list of 8-byte values stores the corrected commit dates for commits\n     -+      with corrected commit date offsets that cannot be stored within 31 bits.\n     ++    * This list of 8-byte values stores the corrected commit date offsets\n     ++      for commits with corrected commit date offsets that cannot be\n     ++      stored within 31 bits.\n     ++    * Generation Data Overflow chunk is present only when Generation Data\n     ++      chunk is present and atleast one corrected commit date offset cannot\n     ++      be stored within 31 bits.\n      +\n         Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n             This list of 4-byte values store the second through nth parents for\n     @@ Documentation/technical/commit-graph.txt: A consumer may load the following info\n       \n      - * A commit with at least one parent has generation number one more than\n      -   the largest generation number among its parents.\n     -+  * A commit with no parents (a root commit) has corrected committer date\n     ++ * A commit with no parents (a root commit) has corrected committer date\n      +    equal to its committer date.\n       \n      -Equivalently, the generation number of a commit A is one more than the\n     -+  * A commit with at least one parent has corrected committer date equal to\n     ++ * A commit with at least one parent has corrected committer date equal to\n      +    the maximum of its commiter date and one more than the largest corrected\n      +    committer date among its parents.\n      +\n     -+  * As a special case, a root commit with timestamp zero has corrected commit\n     ++ * As a special case, a root commit with timestamp zero has corrected commit\n      +    date of 1, to be able to distinguish it from GENERATION_NUMBER_ZERO\n      +    (that is, an uncomputed corrected commit date).\n      +\n     @@ Documentation/technical/commit-graph.txt: is easier to use for computation and o\n           generation numbers, then we always expand the boundary commit with highest\n           generation number and can easily detect the stopping condition.\n       \n     -+The properties applies to both versions of generation number, that is both\n     ++The property applies to both versions of generation number, that is both\n      +corrected committer dates and topological levels.\n      +\n       This property can be used to significantly reduce the time it takes to\n     @@ Documentation/technical/commit-graph.txt: fully-computed generation numbers. Usi\n       with generation number *_INFINITY or *_ZERO is valuable.\n       \n      -We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n     -+We use the macro GENERATION_NUMBER_MAX for commits whose\n     - generation numbers are computed to be at least this value. We limit at\n     - this value since it is the largest value that can be stored in the\n     - commit-graph file using the 30 bits available to generation numbers. This\n     +-generation numbers are computed to be at least this value. We limit at\n     +-this value since it is the largest value that can be stored in the\n     +-commit-graph file using the 30 bits available to generation numbers. This\n     +-presents another case where a commit can have generation number equal to\n     +-that of a parent.\n     ++We use the macro GENERATION_NUMBER_V1_MAX = 0x3FFFFFFF for commits whose\n     ++topological levels (generation number v1) are computed to be at least\n     ++this value. We limit at this value since it is the largest value that\n     ++can be stored in the commit-graph file using the 30 bits available\n     ++to topological levels. This presents another case where a commit can\n     ++have generation number equal to that of a parent.\n     + \n     + Design Details\n     + --------------\n      @@ Documentation/technical/commit-graph.txt: The merge strategy values (2 for the size multiple, 64,000 for the maximum\n       number of commits) could be extracted into config settings for full\n       flexibility.\n     @@ Documentation/technical/commit-graph.txt: The merge strategy values (2 for the s\n      +A naive approach of using the newest available generation number from\n      +each layer would lead to violated expectations: the lower layer would\n      +use corrected commit dates which are much larger than the topological\n     -+levels of the higher layer. For this reason, Git inspects each layer to\n     -+see if any layer is missing corrected commit dates. In such a case, Git\n     -+only uses topological level\n     ++levels of the higher layer. For this reason, Git inspects the topmost\n     ++layer to see if the layer is missing corrected commit dates. In such a case\n     ++Git only uses topological level for generation numbers.\n      +\n      +When writing a new layer in split commit-graph, we write corrected commit\n      +dates if the topmost layer has corrected commit dates written. This\n     @@ Documentation/technical/commit-graph.txt: The merge strategy values (2 for the s\n      +must have corrected commit dates as well.\n      +\n      +When merging layers, we do not consider whether the merged layers had corrected\n     -+commit dates. Instead, the new layer will have corrected commit dates if and\n     -+only if all existing layers below the new layer have corrected commit dates.\n     ++commit dates. Instead, the new layer will have corrected commit dates if the\n     ++layer below the new layer has corrected commit dates.\n     ++\n     ++While writing or merging layers, if the new layer is the only layer, it will\n     ++have corrected commit dates when written by compatible versions of Git. Thus,\n     ++rewriting split commit-graph as a single file (`--split=replace`) creates a\n     ++single layer with corrected commit dates.\n      +\n       ## Deleting graph-{hash} files\n       \n\n-- \ngitgitgadget\n"},{"id":"413057","messageId":"baae700676405fba5153ec9d8c8f32760bda0a70.1609154168.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v5 05/11] commit-graph: add a slab to store topological levels","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:16:02Z","receivedAt":"2020-12-28T11:17:51Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a later commit we will introduce corrected commit date as the\ngeneration number v2. Corrected commit dates will be stored in the new\nseperate Generation Data chunk. However, to ensure backwards\ncompatibility with \"Old\" Git we need to continue to write generation\nnumber v1 (topological levels) to the commit data chunk. Thus, we need\nto compute and store both versions of generation numbers to write the\ncommit-graph file.\n\nTherefore, let's introduce a commit-slab `topo_level_slab` to store\ntopological levels; corrected commit date will be stored in the member\n`generation` of struct commit_graph_data.\n\nThe macros `GENERATION_NUMBER_INFINITY` and `GENERATION_NUMBER_ZERO`\nmark commits not in the commit-graph file and commits written by a\nversion of Git that did not compute generation numbers respectively.\nGeneration numbers are computed identically for both kinds of commits.\n\nA \"slab-miss\" should return `GENERATION_NUMBER_INFINITY` as the commit\nis not in the commit-graph file. However, since the slab is\nzero-initialized, it returns 0 (or rather `GENERATION_NUMBER_ZERO`).\nThus, we no longer need to check if the topological level of a commit is\n`GENERATION_NUMBER_INFINITY`.\n\nWe will add a pointer to the slab in `struct write_commit_graph_context`\nand `struct commit_graph` to populate the slab in\n`fill_commit_graph_info` if the commit has a pre-computed topological\nlevel as in case of split commit-graphs.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 45 ++++++++++++++++++++++++++++++---------------\n commit-graph.h |  1 +\n 2 files changed, 31 insertions(+), 15 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex d5b33b4f7ac..c98e8910fe2 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -64,6 +64,8 @@ void git_test_write_commit_graph_or_die(void)\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n+define_commit_slab(topo_level_slab, uint32_t);\n+\n /* Keep track of the order in which commits are added to our list. */\n define_commit_slab(commit_pos, int);\n static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n@@ -768,6 +770,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\n+\tif (g->topo_levels)\n+\t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n static inline void set_commit_tree(struct commit *c, struct tree *t)\n@@ -956,6 +961,7 @@ struct write_commit_graph_context {\n \t\t changed_paths:1,\n \t\t order_by_pack:1;\n \n+\tstruct topo_level_slab *topo_levels;\n \tconst struct commit_graph_opts *opts;\n \tsize_t total_bloom_filter_data_size;\n \tconst struct bloom_filter_settings *bloom_settings;\n@@ -1102,7 +1108,7 @@ static int write_graph_chunk_data(struct hashfile *f,\n \t\telse\n \t\t\tpackedDate[0] = 0;\n \n-\t\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n+\t\tpackedDate[0] |= htonl(*topo_level_slab_at(ctx->topo_levels, *list) << 2);\n \n \t\tpackedDate[1] = htonl((*list)->date);\n \t\thashwrite(f, packedDate, 8);\n@@ -1332,11 +1338,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n+\t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n-\t\t    generation != GENERATION_NUMBER_ZERO)\n+\t\tif (level != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1344,29 +1349,26 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit *current = list->item;\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n-\t\t\tuint32_t max_generation = 0;\n+\t\t\tuint32_t max_level = 0;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n-\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n+\t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n \n-\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n-\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n+\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (generation > max_generation) {\n-\t\t\t\t\tmax_generation = generation;\n+\t\t\t\t} else if (level > max_level) {\n+\t\t\t\t\tmax_level = level;\n \t\t\t\t}\n \t\t\t}\n \n \t\t\tif (all_parents_computed) {\n-\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n-\n-\t\t\t\tdata->generation = max_generation + 1;\n \t\t\t\tpop_commit(&list);\n \n-\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n-\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n+\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n+\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n+\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n \t\t\t}\n \t\t}\n \t}\n@@ -2102,6 +2104,7 @@ int write_commit_graph(struct object_directory *odb,\n \tint res = 0;\n \tint replace = 0;\n \tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n+\tstruct topo_level_slab topo_levels;\n \n \tprepare_repo_settings(the_repository);\n \tif (!the_repository->settings.core_commit_graph) {\n@@ -2128,6 +2131,18 @@ int write_commit_graph(struct object_directory *odb,\n \t\t\t\t\t\t\t bloom_settings.max_changed_paths);\n \tctx->bloom_settings = &bloom_settings;\n \n+\tinit_topo_level_slab(&topo_levels);\n+\tctx->topo_levels = &topo_levels;\n+\n+\tif (ctx->r->objects->commit_graph) {\n+\t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n+\n+\t\twhile (g) {\n+\t\t\tg->topo_levels = &topo_levels;\n+\t\t\tg = g->base_graph;\n+\t\t}\n+\t}\n+\n \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n \t\tctx->changed_paths = 1;\n \tif (!(flags & COMMIT_GRAPH_NO_WRITE_BLOOM_FILTERS)) {\ndiff --git a/commit-graph.h b/commit-graph.h\nindex f8e92500c6e..00f00745b79 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -73,6 +73,7 @@ struct commit_graph {\n \tconst unsigned char *chunk_bloom_indexes;\n \tconst unsigned char *chunk_bloom_data;\n \n+\tstruct topo_level_slab *topo_levels;\n \tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \n-- \ngitgitgadget\n\n"},{"id":"413058","messageId":"26bd6f4910059a0900a9ce48b2d6b668da6e34e7.1609154168.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v5 06/11] commit-graph: return 64-bit generation number","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:16:03Z","receivedAt":"2020-12-28T11:17:51Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a preparatory step for introducing corrected commit dates, let's\nreturn timestamp_t values from commit_graph_generation(), use\ntimestamp_t for local variables and define GENERATION_NUMBER_INFINITY\nas (2 ^ 63 - 1) instead.\n\nWe rename GENERATION_NUMBER_MAX to GENERATION_NUMBER_V1_MAX to\nrepresent the largest topological level we can store in the commit data\nchunk.\n\nWith corrected commit dates implemented, we will have two such *_MAX\nvariables to denote the largest offset and largest topological level\nthat can be stored.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 22 +++++++++++-----------\n commit-graph.h |  4 ++--\n commit-reach.c | 36 ++++++++++++++++++------------------\n commit-reach.h |  2 +-\n commit.c       |  4 ++--\n commit.h       |  4 ++--\n revision.c     | 10 +++++-----\n upload-pack.c  |  2 +-\n 8 files changed, 42 insertions(+), 42 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex c98e8910fe2..1b2a015f92f 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -101,7 +101,7 @@ uint32_t commit_graph_position(const struct commit *c)\n \treturn data ? data->graph_pos : COMMIT_NOT_FROM_GRAPH;\n }\n \n-uint32_t commit_graph_generation(const struct commit *c)\n+timestamp_t commit_graph_generation(const struct commit *c)\n {\n \tstruct commit_graph_data *data =\n \t\tcommit_graph_data_slab_peek(&commit_graph_data_slab, c);\n@@ -146,8 +146,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n-\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n+\tconst timestamp_t generation_a = commit_graph_data_at(a)->generation;\n+\tconst timestamp_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n@@ -1366,8 +1366,8 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tif (all_parents_computed) {\n \t\t\t\tpop_commit(&list);\n \n-\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n-\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n+\t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n+\t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n \t\t\t}\n \t\t}\n@@ -2363,8 +2363,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \tfor (i = 0; i < g->num_commits; i++) {\n \t\tstruct commit *graph_commit, *odb_commit;\n \t\tstruct commit_list *graph_parents, *odb_parents;\n-\t\tuint32_t max_generation = 0;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t max_generation = 0;\n+\t\ttimestamp_t generation;\n \n \t\tdisplay_progress(progress, i + 1);\n \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n@@ -2428,16 +2428,16 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t\tcontinue;\n \n \t\t/*\n-\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n-\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n+\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n+\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n \t\t * extra logic in the following condition.\n \t\t */\n-\t\tif (max_generation == GENERATION_NUMBER_MAX)\n+\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n \t\t\tmax_generation--;\n \n \t\tgeneration = commit_graph_generation(graph_commit);\n \t\tif (generation != max_generation + 1)\n-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" != %\"PRItime),\n \t\t\t\t     oid_to_hex(&cur_oid),\n \t\t\t\t     generation,\n \t\t\t\t     max_generation + 1);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 00f00745b79..2e9aa7824ee 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -145,12 +145,12 @@ void disable_commit_graph(struct repository *r);\n \n struct commit_graph_data {\n \tuint32_t graph_pos;\n-\tuint32_t generation;\n+\ttimestamp_t generation;\n };\n \n /*\n  * Commits should be parsed before accessing generation, graph positions.\n  */\n-uint32_t commit_graph_generation(const struct commit *);\n+timestamp_t commit_graph_generation(const struct commit *);\n uint32_t commit_graph_position(const struct commit *);\n #endif\ndiff --git a/commit-reach.c b/commit-reach.c\nindex 50175b159e7..9b24b0378d5 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -32,12 +32,12 @@ static int queue_has_nonstale(struct prio_queue *queue)\n static struct commit_list *paint_down_to_common(struct repository *r,\n \t\t\t\t\t\tstruct commit *one, int n,\n \t\t\t\t\t\tstruct commit **twos,\n-\t\t\t\t\t\tint min_generation)\n+\t\t\t\t\t\ttimestamp_t min_generation)\n {\n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \tstruct commit_list *result = NULL;\n \tint i;\n-\tuint32_t last_gen = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n \tif (!min_generation)\n \t\tqueue.compare = compare_commits_by_commit_date;\n@@ -58,10 +58,10 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \t\tstruct commit *commit = prio_queue_get(&queue);\n \t\tstruct commit_list *parents;\n \t\tint flags;\n-\t\tuint32_t generation = commit_graph_generation(commit);\n+\t\ttimestamp_t generation = commit_graph_generation(commit);\n \n \t\tif (min_generation && generation > last_gen)\n-\t\t\tBUG(\"bad generation skip %8x > %8x at %s\",\n+\t\t\tBUG(\"bad generation skip %\"PRItime\" > %\"PRItime\" at %s\",\n \t\t\t    generation, last_gen,\n \t\t\t    oid_to_hex(&commit->object.oid));\n \t\tlast_gen = generation;\n@@ -177,12 +177,12 @@ static int remove_redundant(struct repository *r, struct commit **array, int cnt\n \t\trepo_parse_commit(r, array[i]);\n \tfor (i = 0; i < cnt; i++) {\n \t\tstruct commit_list *common;\n-\t\tuint32_t min_generation = commit_graph_generation(array[i]);\n+\t\ttimestamp_t min_generation = commit_graph_generation(array[i]);\n \n \t\tif (redundant[i])\n \t\t\tcontinue;\n \t\tfor (j = filled = 0; j < cnt; j++) {\n-\t\t\tuint32_t curr_generation;\n+\t\t\ttimestamp_t curr_generation;\n \t\t\tif (i == j || redundant[j])\n \t\t\t\tcontinue;\n \t\t\tfilled_index[filled] = j;\n@@ -321,7 +321,7 @@ int repo_in_merge_bases_many(struct repository *r, struct commit *commit,\n {\n \tstruct commit_list *bases;\n \tint ret = 0, i;\n-\tuint32_t generation, max_generation = GENERATION_NUMBER_ZERO;\n+\ttimestamp_t generation, max_generation = GENERATION_NUMBER_ZERO;\n \n \tif (repo_parse_commit(r, commit))\n \t\treturn ret;\n@@ -470,7 +470,7 @@ static int in_commit_list(const struct commit_list *want, struct commit *c)\n static enum contains_result contains_test(struct commit *candidate,\n \t\t\t\t\t  const struct commit_list *want,\n \t\t\t\t\t  struct contains_cache *cache,\n-\t\t\t\t\t  uint32_t cutoff)\n+\t\t\t\t\t  timestamp_t cutoff)\n {\n \tenum contains_result *cached = contains_cache_at(cache, candidate);\n \n@@ -506,11 +506,11 @@ static enum contains_result contains_tag_algo(struct commit *candidate,\n {\n \tstruct contains_stack contains_stack = { 0, 0, NULL };\n \tenum contains_result result;\n-\tuint32_t cutoff = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t cutoff = GENERATION_NUMBER_INFINITY;\n \tconst struct commit_list *p;\n \n \tfor (p = want; p; p = p->next) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = p->item;\n \t\tload_commit_graph_info(the_repository, c);\n \t\tgeneration = commit_graph_generation(c);\n@@ -566,8 +566,8 @@ static int compare_commits_by_gen(const void *_a, const void *_b)\n \tconst struct commit *a = *(const struct commit * const *)_a;\n \tconst struct commit *b = *(const struct commit * const *)_b;\n \n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n+\ttimestamp_t generation_a = commit_graph_generation(a);\n+\ttimestamp_t generation_b = commit_graph_generation(b);\n \n \tif (generation_a < generation_b)\n \t\treturn -1;\n@@ -580,7 +580,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation)\n+\t\t\t\t timestamp_t min_generation)\n {\n \tstruct commit **list = NULL;\n \tint i;\n@@ -681,13 +681,13 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \ttime_t min_commit_date = cutoff_by_min_date ? from->item->date : 0;\n \tstruct commit_list *from_iter = from, *to_iter = to;\n \tint result;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \n \twhile (from_iter) {\n \t\tadd_object_array(&from_iter->item->object, NULL, &from_objs);\n \n \t\tif (!parse_commit(from_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (from_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = from_iter->item->date;\n \n@@ -701,7 +701,7 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \n \twhile (to_iter) {\n \t\tif (!parse_commit(to_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (to_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = to_iter->item->date;\n \n@@ -741,13 +741,13 @@ struct commit_list *get_reachable_subset(struct commit **from, int nr_from,\n \tstruct commit_list *found_commits = NULL;\n \tstruct commit **to_last = to + nr_to;\n \tstruct commit **from_last = from + nr_from;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \tint num_to_find = 0;\n \n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \n \tfor (item = to; item < to_last; item++) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = *item;\n \n \t\tparse_commit(c);\ndiff --git a/commit-reach.h b/commit-reach.h\nindex b49ad71a317..148b56fea50 100644\n--- a/commit-reach.h\n+++ b/commit-reach.h\n@@ -87,7 +87,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation);\n+\t\t\t\t timestamp_t min_generation);\n int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \t\t       int commit_date_cutoff);\n \ndiff --git a/commit.c b/commit.c\nindex fe1fa3dc41f..17abf92a2d2 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -731,8 +731,8 @@ int compare_commits_by_author_date(const void *a_, const void *b_,\n int compare_commits_by_gen_then_commit_date(const void *a_, const void *b_, void *unused)\n {\n \tconst struct commit *a = a_, *b = b_;\n-\tconst uint32_t generation_a = commit_graph_generation(a),\n-\t\t       generation_b = commit_graph_generation(b);\n+\tconst timestamp_t generation_a = commit_graph_generation(a),\n+\t\t\t  generation_b = commit_graph_generation(b);\n \n \t/* newer commits first */\n \tif (generation_a < generation_b)\ndiff --git a/commit.h b/commit.h\nindex 5467786c7be..33c66b2177c 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -11,8 +11,8 @@\n #include \"commit-slab.h\"\n \n #define COMMIT_NOT_FROM_GRAPH 0xFFFFFFFF\n-#define GENERATION_NUMBER_INFINITY 0xFFFFFFFF\n-#define GENERATION_NUMBER_MAX 0x3FFFFFFF\n+#define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n+#define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n \n struct commit_list {\ndiff --git a/revision.c b/revision.c\nindex de8e45f462f..d55c2e4d566 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3300,7 +3300,7 @@ define_commit_slab(indegree_slab, int);\n define_commit_slab(author_date_slab, timestamp_t);\n \n struct topo_walk_info {\n-\tuint32_t min_generation;\n+\ttimestamp_t min_generation;\n \tstruct prio_queue explore_queue;\n \tstruct prio_queue indegree_queue;\n \tstruct prio_queue topo_queue;\n@@ -3346,7 +3346,7 @@ static void explore_walk_step(struct rev_info *revs)\n }\n \n static void explore_to_depth(struct rev_info *revs,\n-\t\t\t     uint32_t gen_cutoff)\n+\t\t\t     timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3389,7 +3389,7 @@ static void indegree_walk_step(struct rev_info *revs)\n }\n \n static void compute_indegrees_to_depth(struct rev_info *revs,\n-\t\t\t\t       uint32_t gen_cutoff)\n+\t\t\t\t       timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3447,7 +3447,7 @@ static void init_topo_walk(struct rev_info *revs)\n \tinfo->min_generation = GENERATION_NUMBER_INFINITY;\n \tfor (list = revs->commits; list; list = list->next) {\n \t\tstruct commit *c = list->item;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (repo_parse_commit_gently(revs->repo, c, 1))\n \t\t\tcontinue;\n@@ -3508,7 +3508,7 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n \tfor (p = commit->parents; p; p = p->next) {\n \t\tstruct commit *parent = p->item;\n \t\tint *pi;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (parent->object.flags & UNINTERESTING)\n \t\t\tcontinue;\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 3b66bf92ba8..b87607e0dd4 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -500,7 +500,7 @@ static int got_oid(struct upload_pack_data *data,\n \n static int ok_to_give_up(struct upload_pack_data *data)\n {\n-\tuint32_t min_generation = GENERATION_NUMBER_ZERO;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_ZERO;\n \n \tif (!data->have_obj.nr)\n \t\treturn 0;\n-- \ngitgitgadget\n\n"},{"id":"413059","messageId":"093101f908b166099af41d99250cf8e79d921740.1609154169.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v5 10/11] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:16:07Z","receivedAt":"2020-12-28T11:17:51Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\n091f4cf (commit: don't use generation numbers if not needed,\n2018-08-30) changed paint_down_to_common() to use commit dates instead\nof generation numbers v1 (topological levels) as the performance\nregressed on certain topologies. With generation number v2 (corrected\ncommit dates) implemented, we no longer have to rely on commit dates and\ncan use generation numbers.\n\nFor example, the command `git merge-base v4.8 v4.9` on the Linux\nrepository walks 167468 commits, taking 0.135s for committer date and\n167496 commits, taking 0.157s for corrected committer date respectively.\n\nWhile using corrected commit dates, Git walks nearly the same number of\ncommits as commit date, the process is slower as for each comparision we\nhave to access a commit-slab (for corrected committer date) instead of\naccessing struct member (for committer date).\n\nThis change incidentally broke the fragile t6404-recursive-merge test.\nt6404-recursive-merge sets up a unique repository where all commits have\nthe same committer date without a well-defined merge-base.\n\nWhile running tests with GIT_TEST_COMMIT_GRAPH unset, we use committer\ndate as a heuristic in paint_down_to_common(). 6404.1 'combined merge\nconflicts' merges commits in the order:\n- Merge C with B to form an intermediate commit.\n- Merge the intermediate commit with A.\n\nWith GIT_TEST_COMMIT_GRAPH=1, we write a commit-graph and subsequently\nuse the corrected committer date, which changes the order in which\ncommits are merged:\n- Merge A with B to form an intermediate commit.\n- Merge the intermediate commit with C.\n\nWhile resulting repositories are equivalent, 6404.4 'virtual trees were\nprocessed' fails with GIT_TEST_COMMIT_GRAPH=1 as we are selecting\ndifferent merge-bases and thus have different object ids for the\nintermediate commits.\n\nAs this has already causes problems (as noted in 859fdc0 (commit-graph:\ndefine GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph\nwithin t6404-recursive-merge.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c             | 14 ++++++++++++++\n commit-graph.h             |  6 ++++++\n commit-reach.c             |  2 +-\n t/t6404-recursive-merge.sh |  5 ++++-\n 4 files changed, 25 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 41a65d98738..c8d7ed13302 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -710,6 +710,20 @@ int generation_numbers_enabled(struct repository *r)\n \treturn !!first_generation;\n }\n \n+int corrected_commit_dates_enabled(struct repository *r)\n+{\n+\tstruct commit_graph *g;\n+\tif (!prepare_commit_graph(r))\n+\t\treturn 0;\n+\n+\tg = r->objects->commit_graph;\n+\n+\tif (!g->num_commits)\n+\t\treturn 0;\n+\n+\treturn g->read_generation_data;\n+}\n+\n struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r)\n {\n \tstruct commit_graph *g = r->objects->commit_graph;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex ad52130883b..97f3497c279 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -95,6 +95,12 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n  */\n int generation_numbers_enabled(struct repository *r);\n \n+/*\n+ * Return 1 if and only if the repository has a commit-graph\n+ * file and generation data chunk has been written for the file.\n+ */\n+int corrected_commit_dates_enabled(struct repository *r);\n+\n struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n \n enum commit_graph_write_flags {\ndiff --git a/commit-reach.c b/commit-reach.c\nindex 9b24b0378d5..e38771ca5a1 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -39,7 +39,7 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \tint i;\n \ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n-\tif (!min_generation)\n+\tif (!min_generation && !corrected_commit_dates_enabled(r))\n \t\tqueue.compare = compare_commits_by_commit_date;\n \n \tone->object.flags |= PARENT1;\ndiff --git a/t/t6404-recursive-merge.sh b/t/t6404-recursive-merge.sh\nindex b1c3d4dda49..86f74ae5847 100755\n--- a/t/t6404-recursive-merge.sh\n+++ b/t/t6404-recursive-merge.sh\n@@ -15,6 +15,8 @@ GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n export GIT_COMMITTER_DATE\n \n test_expect_success 'setup tests' '\n+\tGIT_TEST_COMMIT_GRAPH=0 &&\n+\texport GIT_TEST_COMMIT_GRAPH &&\n \techo 1 >a1 &&\n \tgit add a1 &&\n \tGIT_AUTHOR_DATE=\"2006-12-12 23:00:00\" git commit -m 1 a1 &&\n@@ -66,7 +68,7 @@ test_expect_success 'setup tests' '\n '\n \n test_expect_success 'combined merge conflicts' '\n-\ttest_must_fail env GIT_TEST_COMMIT_GRAPH=0 git merge -m final G\n+\ttest_must_fail git merge -m final G\n '\n \n test_expect_success 'result contains a conflict' '\n@@ -82,6 +84,7 @@ test_expect_success 'result contains a conflict' '\n '\n \n test_expect_success 'virtual trees were processed' '\n+\t# TODO: fragile test, relies on ambigious merge-base resolution\n \tgit ls-files --stage >out &&\n \n \tcat >expect <<-EOF &&\n-- \ngitgitgadget\n\n"},{"id":"413060","messageId":"859c39eff52e32ad322969d024184971acec82e7.1609154168.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v5 07/11] commit-graph: implement corrected commit date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:16:04Z","receivedAt":"2020-12-28T11:17:51Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith most of preparations done, let's implement corrected commit date.\n\nThe corrected commit date for a commit is defined as:\n\n* A commit with no parents (a root commit) has corrected commit date\n  equal to its committer date.\n* A commit with at least one parent has corrected commit date equal to\n  the maximum of its commit date and one more than the largest corrected\n  commit date among its parents.\n\nAs a special case, a root commit with timestamp of zero (01.01.1970\n00:00:00Z) has corrected commit date of one, to be able to distinguish\nfrom GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit\ndate).\n\nTo minimize the space required to store corrected commit date, Git\nstores corrected commit date offsets into the commit-graph file. The\ncorrected commit date offset for a commit is defined as the difference\nbetween its corrected commit date and actual commit date.\n\nStoring corrected commit date requires sizeof(timestamp_t) bytes, which\nin most cases is 64 bits (uintmax_t). However, corrected commit date\noffsets can be safely stored using only 32-bits. This halves the size\nof GDAT chunk, which is a reduction of around 6% in the size of\ncommit-graph file.\n\nHowever, using offsets be problematic if one of commits is malformed but\nvalid and has committerdate of 0 Unix time, as the offset would be the\nsame as corrected commit date and thus require 64-bits to be stored\nproperly.\n\nWhile Git does not write out offsets at this stage, Git stores the\ncorrected commit dates in member generation of struct commit_graph_data.\nIt will begin writing commit date offsets with the introduction of\ngeneration data chunk.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 21 +++++++++++++++++----\n 1 file changed, 17 insertions(+), 4 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 1b2a015f92f..bfc3aae5f93 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1339,9 +1339,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n+\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (level != GENERATION_NUMBER_ZERO)\n+\t\tif (level != GENERATION_NUMBER_ZERO &&\n+\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1350,16 +1352,23 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n \t\t\tuint32_t max_level = 0;\n+\t\t\ttimestamp_t max_corrected_commit_date = 0;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n+\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n \n-\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n+\t\t\t\tif (level == GENERATION_NUMBER_ZERO ||\n+\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (level > max_level) {\n-\t\t\t\t\tmax_level = level;\n+\t\t\t\t} else {\n+\t\t\t\t\tif (level > max_level)\n+\t\t\t\t\t\tmax_level = level;\n+\n+\t\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n+\t\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n \t\t\t\t}\n \t\t\t}\n \n@@ -1369,6 +1378,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n \t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n+\n+\t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n+\t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n+\t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n \t\t\t}\n \t\t}\n \t}\n-- \ngitgitgadget\n\n"},{"id":"413061","messageId":"20299e574574690ba11961d493aad378b804e5b5.1609154169.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v5 11/11] doc: add corrected commit date info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:16:08Z","receivedAt":"2020-12-28T11:17:51Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith generation data chunk and corrected commit dates implemented, let's\nupdate the technical documentation for commit-graph.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n .../technical/commit-graph-format.txt         | 28 +++++--\n Documentation/technical/commit-graph.txt      | 77 +++++++++++++++----\n 2 files changed, 86 insertions(+), 19 deletions(-)\n\ndiff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\nindex b3b58880b92..b6658eff188 100644\n--- a/Documentation/technical/commit-graph-format.txt\n+++ b/Documentation/technical/commit-graph-format.txt\n@@ -4,11 +4,7 @@ Git commit graph format\n The Git commit graph stores a list of commit OIDs and some associated\n metadata, including:\n \n-- The generation number of the commit. Commits with no parents have\n-  generation number 1; commits with parents have generation number\n-  one more than the maximum generation number of its parents. We\n-  reserve zero as special, and can be used to mark a generation\n-  number invalid or as \"not computed\".\n+- The generation number of the commit.\n \n - The root tree OID.\n \n@@ -86,13 +82,33 @@ CHUNK DATA:\n       position. If there are more than two parents, the second value\n       has its most-significant bit on and the other bits store an array\n       position into the Extra Edge List chunk.\n-    * The next 8 bytes store the generation number of the commit and\n+    * The next 8 bytes store the topological level (generation number v1)\n+      of the commit and\n       the commit time in seconds since EPOCH. The generation number\n       uses the higher 30 bits of the first 4 bytes, while the commit\n       time uses the 32 bits of the second 4 bytes, along with the lowest\n       2 bits of the lowest byte, storing the 33rd and 34th bit of the\n       commit time.\n \n+  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes) [Optional]\n+    * This list of 4-byte values store corrected commit date offsets for the\n+      commits, arranged in the same order as commit data chunk.\n+    * If the corrected commit date offset cannot be stored within 31 bits,\n+      the value has its most-significant bit on and the other bits store\n+      the position of corrected commit date into the Generation Data Overflow\n+      chunk.\n+    * Generation Data chunk is present only when commit-graph file is written\n+      by compatible versions of Git and in case of split commit-graph chains,\n+      the topmost layer also has Generation Data chunk.\n+\n+  Generation Data Overflow (ID: {'G', 'D', 'O', 'V' }) [Optional]\n+    * This list of 8-byte values stores the corrected commit date offsets\n+      for commits with corrected commit date offsets that cannot be\n+      stored within 31 bits.\n+    * Generation Data Overflow chunk is present only when Generation Data\n+      chunk is present and atleast one corrected commit date offset cannot\n+      be stored within 31 bits.\n+\n   Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n       This list of 4-byte values store the second through nth parents for\n       all octopus merges. The second parent value in the commit data stores\ndiff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\nindex f14a7659aa8..f05e7bda1a9 100644\n--- a/Documentation/technical/commit-graph.txt\n+++ b/Documentation/technical/commit-graph.txt\n@@ -38,14 +38,31 @@ A consumer may load the following info for a commit from the graph:\n \n Values 1-4 satisfy the requirements of parse_commit_gently().\n \n-Define the \"generation number\" of a commit recursively as follows:\n+There are two definitions of generation number:\n+1. Corrected committer dates (generation number v2)\n+2. Topological levels (generation nummber v1)\n \n- * A commit with no parents (a root commit) has generation number one.\n+Define \"corrected committer date\" of a commit recursively as follows:\n \n- * A commit with at least one parent has generation number one more than\n-   the largest generation number among its parents.\n+ * A commit with no parents (a root commit) has corrected committer date\n+    equal to its committer date.\n \n-Equivalently, the generation number of a commit A is one more than the\n+ * A commit with at least one parent has corrected committer date equal to\n+    the maximum of its commiter date and one more than the largest corrected\n+    committer date among its parents.\n+\n+ * As a special case, a root commit with timestamp zero has corrected commit\n+    date of 1, to be able to distinguish it from GENERATION_NUMBER_ZERO\n+    (that is, an uncomputed corrected commit date).\n+\n+Define the \"topological level\" of a commit recursively as follows:\n+\n+ * A commit with no parents (a root commit) has topological level of one.\n+\n+ * A commit with at least one parent has topological level one more than\n+   the largest topological level among its parents.\n+\n+Equivalently, the topological level of a commit A is one more than the\n length of a longest path from A to a root commit. The recursive definition\n is easier to use for computation and observing the following property:\n \n@@ -60,6 +77,9 @@ is easier to use for computation and observing the following property:\n     generation numbers, then we always expand the boundary commit with highest\n     generation number and can easily detect the stopping condition.\n \n+The property applies to both versions of generation number, that is both\n+corrected committer dates and topological levels.\n+\n This property can be used to significantly reduce the time it takes to\n walk commits and determine topological relationships. Without generation\n numbers, the general heuristic is the following:\n@@ -67,7 +87,9 @@ numbers, the general heuristic is the following:\n     If A and B are commits with commit time X and Y, respectively, and\n     X < Y, then A _probably_ cannot reach B.\n \n-This heuristic is currently used whenever the computation is allowed to\n+In absence of corrected commit dates (for example, old versions of Git or\n+mixed generation graph chains),\n+this heuristic is currently used whenever the computation is allowed to\n violate topological relationships due to clock skew (such as \"git log\"\n with default order), but is not used when the topological order is\n required (such as merge base calculations, \"git log --graph\").\n@@ -77,7 +99,7 @@ in the commit graph. We can treat these commits as having \"infinite\"\n generation number and walk until reaching commits with known generation\n number.\n \n-We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n+We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n in the commit-graph file. If a commit-graph file was written by a version\n of Git that did not compute generation numbers, then those commits will\n have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n@@ -93,12 +115,12 @@ fully-computed generation numbers. Using strict inequality may result in\n walking a few extra commits, but the simplicity in dealing with commits\n with generation number *_INFINITY or *_ZERO is valuable.\n \n-We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n-generation numbers are computed to be at least this value. We limit at\n-this value since it is the largest value that can be stored in the\n-commit-graph file using the 30 bits available to generation numbers. This\n-presents another case where a commit can have generation number equal to\n-that of a parent.\n+We use the macro GENERATION_NUMBER_V1_MAX = 0x3FFFFFFF for commits whose\n+topological levels (generation number v1) are computed to be at least\n+this value. We limit at this value since it is the largest value that\n+can be stored in the commit-graph file using the 30 bits available\n+to topological levels. This presents another case where a commit can\n+have generation number equal to that of a parent.\n \n Design Details\n --------------\n@@ -267,6 +289,35 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n number of commits) could be extracted into config settings for full\n flexibility.\n \n+## Handling Mixed Generation Number Chains\n+\n+With the introduction of generation number v2 and generation data chunk, the\n+following scenario is possible:\n+\n+1. \"New\" Git writes a commit-graph with the corrected commit dates.\n+2. \"Old\" Git writes a split commit-graph on top without corrected commit dates.\n+\n+A naive approach of using the newest available generation number from\n+each layer would lead to violated expectations: the lower layer would\n+use corrected commit dates which are much larger than the topological\n+levels of the higher layer. For this reason, Git inspects the topmost\n+layer to see if the layer is missing corrected commit dates. In such a case\n+Git only uses topological level for generation numbers.\n+\n+When writing a new layer in split commit-graph, we write corrected commit\n+dates if the topmost layer has corrected commit dates written. This\n+guarantees that if a layer has corrected commit dates, all lower layers\n+must have corrected commit dates as well.\n+\n+When merging layers, we do not consider whether the merged layers had corrected\n+commit dates. Instead, the new layer will have corrected commit dates if the\n+layer below the new layer has corrected commit dates.\n+\n+While writing or merging layers, if the new layer is the only layer, it will\n+have corrected commit dates when written by compatible versions of Git. Thus,\n+rewriting split commit-graph as a single file (`--split=replace`) creates a\n+single layer with corrected commit dates.\n+\n ## Deleting graph-{hash} files\n \n After a new tip file is written, some `graph-{hash}` files may no longer\n-- \ngitgitgadget\n"},{"id":"413063","messageId":"8403c4d025727bbc4b69ca12c42dd1db7826159b.1609154169.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v5 08/11] commit-graph: implement generation data chunk","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:16:05Z","receivedAt":"2020-12-28T11:17:51Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nAs discovered by Ævar, we cannot increment graph version to\ndistinguish between generation numbers v1 and v2 [1]. Thus, one of\npre-requistes before implementing generation number v2 was to\ndistinguish between graph versions in a backwards compatible manner.\n\nWe are going to introduce a new chunk called Generation DATa chunk (or\nGDAT). GDAT will store corrected committer date offsets whereas CDAT\nwill still store topological level.\n\nOld Git does not understand GDAT chunk and would ignore it, reading\ntopological levels from CDAT. New Git can parse GDAT and take advantage\nof newer generation numbers, falling back to topological levels when\nGDAT chunk is missing (as it would happen with a commit-graph written\nby old Git).\n\nWe introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\nwhich forces commit-graph file to be written without generation data\nchunk to emulate a commit-graph file written by old Git.\n\nTo minimize the space required to store corrrected commit date, Git\nstores corrected commit date offsets into the commit-graph file, instea\nof corrected commit dates. This saves us 4 bytes per commit, decreasing\nthe GDAT chunk size by half, but it's possible for the offset to\noverflow the 4-bytes allocated for storage. As such overflows are and\nshould be exceedingly rare, we use the following overflow management\nscheme:\n\nWe introduce a new commit-graph chunk, Generation Data OVerflow ('GDOV')\nto store corrected commit dates for commits with offsets greater than\nGENERATION_NUMBER_V2_OFFSET_MAX.\n\nIf the offset is greater than GENERATION_NUMBER_V2_OFFSET_MAX, we set\nthe MSB of the offset and the other bits store the position of corrected\ncommit date in GDOV chunk, similar to how Extra Edge List is maintained.\n\nWe test the overflow-related code with the following repo history:\n\n           F - N - U\n          /         \\\nU - N - U            N\n         \\          /\n\t  N - F - N\n\nWhere the commits denoted by U have committer date of zero seconds\nsince Unix epoch, the commits denoted by N have committer date of\n1112354055 (default committer date for the test suite) seconds since\nUnix epoch and the commits denoted by F have committer date of\n(2 ^ 31 - 2) seconds since Unix epoch.\n\nThe largest offset observed is 2 ^ 31, just large enough to overflow.\n\n[1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                | 111 ++++++++++++++++++++++++++++++----\n commit-graph.h                |   3 +\n commit.h                      |   1 +\n t/README                      |   3 +\n t/helper/test-read-graph.c    |   4 ++\n t/t4216-log-bloom.sh          |   4 +-\n t/t5318-commit-graph.sh       |  79 ++++++++++++++++++++----\n t/t5324-split-commit-graph.sh |  12 ++--\n t/t6600-test-reach.sh         |   6 ++\n t/test-lib-functions.sh       |   6 ++\n 10 files changed, 197 insertions(+), 32 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex bfc3aae5f93..629b2f17fbc 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -38,11 +38,13 @@ void git_test_write_commit_graph_or_die(void)\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n+#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n+#define GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW 0x47444f56 /* \"GDOV\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n-#define MAX_NUM_CHUNKS 7\n+#define MAX_NUM_CHUNKS 9\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -61,6 +63,8 @@ void git_test_write_commit_graph_or_die(void)\n #define GRAPH_MIN_SIZE (GRAPH_HEADER_SIZE + 4 * GRAPH_CHUNKLOOKUP_WIDTH \\\n \t\t\t+ GRAPH_FANOUT_SIZE + the_hash_algo->rawsz)\n \n+#define CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW (1ULL << 31)\n+\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n@@ -390,6 +394,20 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n \t\t\tbreak;\n \n+\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n+\t\t\tif (graph->chunk_generation_data)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW:\n+\t\t\tif (graph->chunk_generation_data_overflow)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_generation_data_overflow = data + chunk_offset;\n+\t\t\tbreak;\n+\n \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n \t\t\tif (graph->chunk_extra_edges)\n \t\t\t\tchunk_repeated = 1;\n@@ -750,8 +768,8 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n {\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n-\tuint32_t lex_index;\n-\tuint64_t date_high, date_low;\n+\tuint32_t lex_index, offset_pos;\n+\tuint64_t date_high, date_low, offset;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n@@ -769,7 +787,16 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\tif (g->chunk_generation_data) {\n+\t\toffset = (timestamp_t)get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n+\n+\t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n+\t\t\toffset_pos = offset ^ CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW;\n+\t\t\tgraph_data->generation = get_be64(g->chunk_generation_data_overflow + 8 * offset_pos);\n+\t\t} else\n+\t\t\tgraph_data->generation = item->date + offset;\n+\t} else\n+\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n \n \tif (g->topo_levels)\n \t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n@@ -941,6 +968,7 @@ struct write_commit_graph_context {\n \tstruct oid_array oids;\n \tstruct packed_commit_list commits;\n \tint num_extra_edges;\n+\tint num_generation_data_overflows;\n \tunsigned long approx_nr_objects;\n \tstruct progress *progress;\n \tint progress_done;\n@@ -959,7 +987,8 @@ struct write_commit_graph_context {\n \t\t report_progress:1,\n \t\t split:1,\n \t\t changed_paths:1,\n-\t\t order_by_pack:1;\n+\t\t order_by_pack:1,\n+\t\t write_generation_data:1;\n \n \tstruct topo_level_slab *topo_levels;\n \tconst struct commit_graph_opts *opts;\n@@ -1119,6 +1148,45 @@ static int write_graph_chunk_data(struct hashfile *f,\n \treturn 0;\n }\n \n+static int write_graph_chunk_generation_data(struct hashfile *f,\n+\t\t\t\t\t      struct write_commit_graph_context *ctx)\n+{\n+\tint i, num_generation_data_overflows = 0;\n+\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n+\n+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n+\t\t\toffset = CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW | num_generation_data_overflows;\n+\t\t\tnum_generation_data_overflows++;\n+\t\t}\n+\n+\t\thashwrite_be32(f, offset);\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int write_graph_chunk_generation_data_overflow(struct hashfile *f,\n+\t\t\t\t\t\t       struct write_commit_graph_context *ctx)\n+{\n+\tint i;\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n+\n+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n+\t\t\thashwrite_be32(f, offset >> 32);\n+\t\t\thashwrite_be32(f, (uint32_t) offset);\n+\t\t}\n+\t}\n+\n+\treturn 0;\n+}\n+\n static int write_graph_chunk_extra_edges(struct hashfile *f,\n \t\t\t\t\t struct write_commit_graph_context *ctx)\n {\n@@ -1382,6 +1450,9 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n \t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n \t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n+\n+\t\t\t\tif (commit_graph_data_at(current)->generation - current->date > GENERATION_NUMBER_V2_OFFSET_MAX)\n+\t\t\t\t\tctx->num_generation_data_overflows++;\n \t\t\t}\n \t\t}\n \t}\n@@ -1715,6 +1786,21 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tchunks[2].id = GRAPH_CHUNKID_DATA;\n \tchunks[2].size = (hashsz + 16) * ctx->commits.nr;\n \tchunks[2].write_fn = write_graph_chunk_data;\n+\n+\tif (git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0))\n+\t\tctx->write_generation_data = 0;\n+\tif (ctx->write_generation_data) {\n+\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA;\n+\t\tchunks[num_chunks].size = sizeof(uint32_t) * ctx->commits.nr;\n+\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data;\n+\t\tnum_chunks++;\n+\t}\n+\tif (ctx->num_generation_data_overflows) {\n+\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW;\n+\t\tchunks[num_chunks].size = sizeof(timestamp_t) * ctx->num_generation_data_overflows;\n+\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data_overflow;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_extra_edges) {\n \t\tchunks[num_chunks].id = GRAPH_CHUNKID_EXTRAEDGES;\n \t\tchunks[num_chunks].size = 4 * ctx->num_extra_edges;\n@@ -2135,6 +2221,8 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n \tctx->opts = opts;\n \tctx->total_bloom_filter_data_size = 0;\n+\tctx->write_generation_data = 1;\n+\tctx->num_generation_data_overflows = 0;\n \n \tbloom_settings.bits_per_entry = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\",\n \t\t\t\t\t\t      bloom_settings.bits_per_entry);\n@@ -2441,16 +2529,17 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t\tcontinue;\n \n \t\t/*\n-\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n-\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n-\t\t * extra logic in the following condition.\n+\t\t * If we are using topological level and one of our parents has\n+\t\t * generation GENERATION_NUMBER_V1_MAX, then our generation is\n+\t\t * also GENERATION_NUMBER_V1_MAX. Decrement to avoid extra logic\n+\t\t * in the following condition.\n \t\t */\n-\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n+\t\tif (!g->chunk_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n \t\t\tmax_generation--;\n \n \t\tgeneration = commit_graph_generation(graph_commit);\n-\t\tif (generation != max_generation + 1)\n-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" != %\"PRItime),\n+\t\tif (generation < max_generation + 1)\n+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n \t\t\t\t     oid_to_hex(&cur_oid),\n \t\t\t\t     generation,\n \t\t\t\t     max_generation + 1);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 2e9aa7824ee..19a02001fde 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -6,6 +6,7 @@\n #include \"oidset.h\"\n \n #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n+#define GIT_TEST_COMMIT_GRAPH_NO_GDAT \"GIT_TEST_COMMIT_GRAPH_NO_GDAT\"\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE \"GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE\"\n #define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n \n@@ -68,6 +69,8 @@ struct commit_graph {\n \tconst uint32_t *chunk_oid_fanout;\n \tconst unsigned char *chunk_oid_lookup;\n \tconst unsigned char *chunk_commit_data;\n+\tconst unsigned char *chunk_generation_data;\n+\tconst unsigned char *chunk_generation_data_overflow;\n \tconst unsigned char *chunk_extra_edges;\n \tconst unsigned char *chunk_base_graphs;\n \tconst unsigned char *chunk_bloom_indexes;\ndiff --git a/commit.h b/commit.h\nindex 33c66b2177c..251d877fcf6 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -14,6 +14,7 @@\n #define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n #define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n+#define GENERATION_NUMBER_V2_OFFSET_MAX ((1ULL << 31) - 1)\n \n struct commit_list {\n \tstruct commit *item;\ndiff --git a/t/README b/t/README\nindex c730a707705..8a121487279 100644\n--- a/t/README\n+++ b/t/README\n@@ -393,6 +393,9 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n+commit-graph to be written without generation data chunk.\n+\n GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n commit-graph write to compute and write changed path Bloom filters for\n every 'git commit-graph write', as if the `--changed-paths` option was\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex 5f585a17256..75927b2c81d 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -33,6 +33,10 @@ int cmd__read_graph(int argc, const char **argv)\n \t\tprintf(\" oid_lookup\");\n \tif (graph->chunk_commit_data)\n \t\tprintf(\" commit_metadata\");\n+\tif (graph->chunk_generation_data)\n+\t\tprintf(\" generation_data\");\n+\tif (graph->chunk_generation_data_overflow)\n+\t\tprintf(\" generation_data_overflow\");\n \tif (graph->chunk_extra_edges)\n \t\tprintf(\" extra_edges\");\n \tif (graph->chunk_bloom_indexes)\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex d11040ce41c..dbde0161882 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -40,11 +40,11 @@ test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n '\n \n graph_read_expect () {\n-\tNUM_CHUNKS=5\n+\tNUM_CHUNKS=6\n \tcat >expect <<- EOF\n \theader: 43475048 1 $(test_oid oid_version) $NUM_CHUNKS 0\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n \tEOF\n \ttest-tool read-graph >actual &&\n \ttest_cmp expect actual\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 2ed0c1544da..fa27df579a5 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -76,7 +76,7 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n graph_read_expect() {\n \tOPTIONAL=\"\"\n \tNUM_CHUNKS=3\n-\tif test ! -z $2\n+\tif test ! -z \"$2\"\n \tthen\n \t\tOPTIONAL=\" $2\"\n \t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n@@ -103,14 +103,14 @@ test_expect_success 'exit with correct error on bad input to --stdin-commits' '\n \t# valid commit and tree OID\n \tgit rev-parse HEAD HEAD^{tree} >in &&\n \tgit commit-graph write --stdin-commits <in &&\n-\tgraph_read_expect 3\n+\tgraph_read_expect 3 generation_data\n '\n \n test_expect_success 'write graph' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"3\"\n+\tgraph_read_expect \"3\" generation_data\n '\n \n test_expect_success POSIXPERM 'write graph has correct permissions' '\n@@ -219,7 +219,7 @@ test_expect_success 'write graph with merges' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'merge 1 vs 2' full merge/1 merge/2\n@@ -254,7 +254,7 @@ test_expect_success 'write graph with new commit' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'full graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -264,7 +264,7 @@ test_expect_success 'write graph with nothing new' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'cleared graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -274,7 +274,7 @@ test_expect_success 'build graph from latest pack with closure' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tcat new-idx | git commit-graph write --stdin-packs &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"9\" \"extra_edges\"\n+\tgraph_read_expect \"9\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'graph from pack, commit 8 vs merge 1' full commits/8 merge/1\n@@ -287,7 +287,7 @@ test_expect_success 'build graph from commits with closure' '\n \tgit rev-parse merge/1 >>commits-in &&\n \tcat commits-in | git commit-graph write --stdin-commits &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"6\"\n+\tgraph_read_expect \"6\" \"generation_data\"\n '\n \n graph_git_behavior 'graph from commits, commit 8 vs merge 1' full commits/8 merge/1\n@@ -297,7 +297,7 @@ test_expect_success 'build graph from commits with append' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit rev-parse merge/3 | git commit-graph write --stdin-commits --append &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -307,7 +307,7 @@ test_expect_success 'build graph using --reachable' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write --reachable &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -328,7 +328,7 @@ test_expect_success 'write graph in bare repo' '\n \tcd \"$TRASH_DIRECTORY/bare\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $baredir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n@@ -454,8 +454,9 @@ test_expect_success 'warn on improper hash version' '\n \n test_expect_success 'git commit-graph verify' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n-\tgit rev-parse commits/8 | git commit-graph write --stdin-commits &&\n-\tgit commit-graph verify >output\n+\tgit rev-parse commits/8 | GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --stdin-commits &&\n+\tgit commit-graph verify >output &&\n+\tgraph_read_expect 9 extra_edges\n '\n \n NUM_COMMITS=9\n@@ -741,4 +742,56 @@ test_expect_success 'corrupt commit-graph write (missing tree)' '\n \t)\n '\n \n+# We test the overflow-related code with the following repo history:\n+#\n+#               4:F - 5:N - 6:U\n+#              /                \\\n+# 1:U - 2:N - 3:U                M:N\n+#              \\                /\n+#               7:N - 8:F - 9:N\n+#\n+# Here the commits denoted by U have committer date of zero seconds\n+# since Unix epoch, the commits denoted by N have committer date\n+# starting from 1112354055 seconds since Unix epoch (default committer\n+# date for the test suite), and the commits denoted by F have committer\n+# date of (2 ^ 31 - 2) seconds since Unix epoch.\n+#\n+# The largest offset observed is 2 ^ 31, just large enough to overflow.\n+#\n+\n+test_expect_success 'set up and verify repo with generation data overflow chunk' '\n+\tobjdir=\".git/objects\" &&\n+\tUNIX_EPOCH_ZERO=\"@0 +0000\" &&\n+\tFUTURE_DATE=\"@2147483646 +0000\" &&\n+\ttest_oid_cache <<-EOF &&\n+\toid_version sha1:1\n+\toid_version sha256:2\n+\tEOF\n+\tcd \"$TRASH_DIRECTORY\" &&\n+\tmkdir repo &&\n+\tcd repo &&\n+\tgit init &&\n+\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 1 &&\n+\ttest_commit 2 &&\n+\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 3 &&\n+\tgit commit-graph write --reachable &&\n+\tgraph_read_expect 3 generation_data &&\n+\ttest_commit --date \"$FUTURE_DATE\" 4 &&\n+\ttest_commit 5 &&\n+\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 6 &&\n+\tgit branch left &&\n+\tgit reset --hard 3 &&\n+\ttest_commit 7 &&\n+\ttest_commit --date \"$FUTURE_DATE\" 8 &&\n+\ttest_commit 9 &&\n+\tgit branch right &&\n+\tgit reset --hard 3 &&\n+\ttest_merge M left right &&\n+\tgit commit-graph write --reachable &&\n+\tgraph_read_expect 10 \"generation_data generation_data_overflow\" &&\n+\tgit commit-graph verify\n+'\n+\n+graph_git_behavior 'generation data overflow chunk repo' repo left right\n+\n test_done\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 4d3842b83b9..587757b62d9 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -13,11 +13,11 @@ test_expect_success 'setup repo' '\n \tinfodir=\".git/objects/info\" &&\n \tgraphdir=\"$infodir/commit-graphs\" &&\n \ttest_oid_cache <<-EOM\n-\tshallow sha1:1760\n-\tshallow sha256:2064\n+\tshallow sha1:2132\n+\tshallow sha256:2436\n \n-\tbase sha1:1376\n-\tbase sha256:1496\n+\tbase sha1:1408\n+\tbase sha256:1528\n \n \toid_version sha1:1\n \toid_version sha256:2\n@@ -31,9 +31,9 @@ graph_read_expect() {\n \t\tNUM_BASE=$2\n \tfi\n \tcat >expect <<- EOF\n-\theader: 43475048 1 $(test_oid oid_version) 3 $NUM_BASE\n+\theader: 43475048 1 $(test_oid oid_version) 4 $NUM_BASE\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n \tEOF\n \ttest-tool read-graph >output &&\n \ttest_cmp expect output\ndiff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\nindex af10f0dc090..e2d33a8a4c4 100755\n--- a/t/t6600-test-reach.sh\n+++ b/t/t6600-test-reach.sh\n@@ -55,6 +55,9 @@ test_expect_success 'setup' '\n \tgit show-ref -s commit-5-5 | git commit-graph write --stdin-commits &&\n \tmv .git/objects/info/commit-graph commit-graph-half &&\n \tchmod u+w commit-graph-half &&\n+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable &&\n+\tmv .git/objects/info/commit-graph commit-graph-no-gdat &&\n+\tchmod u+w commit-graph-no-gdat &&\n \tgit config core.commitGraph true\n '\n \n@@ -67,6 +70,9 @@ run_all_modes () {\n \ttest_cmp expect actual &&\n \tcp commit-graph-half .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n+\ttest_cmp expect actual &&\n+\tcp commit-graph-no-gdat .git/objects/info/commit-graph &&\n+\t\"$@\" <input >actual &&\n \ttest_cmp expect actual\n }\n \ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex 999982fe4a9..3ad712c3acc 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -202,6 +202,12 @@ test_commit () {\n \t\t--signoff)\n \t\t\tsignoff=\"$1\"\n \t\t\t;;\n+\t\t--date)\n+\t\t\tnotick=yes\n+\t\t\tGIT_COMMITTER_DATE=\"$2\"\n+\t\t\tGIT_AUTHOR_DATE=\"$2\"\n+\t\t\tshift\n+\t\t\t;;\n \t\t-C)\n \t\t\tindir=\"$2\"\n \t\t\tshift\n-- \ngitgitgadget\n\n"},{"id":"413062","messageId":"a3a70a1edd0949ff3088fae625afa68fc61975df.1609154169.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v5 09/11] commit-graph: use generation v2 only if entire chain does","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-12-28T11:16:06Z","receivedAt":"2020-12-28T11:17:52Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nSince there are released versions of Git that understand generation\nnumbers in the commit-graph's CDAT chunk but do not understand the GDAT\nchunk, the following scenario is possible:\n\n1. \"New\" Git writes a commit-graph with the GDAT chunk.\n2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n\nIf each layer of split commit-graph is treated independently, as it was\nthe case before this commit, with Git inspecting only the current layer\nfor chunk_generation_data pointer, commits in the lower layer (one with\nGDAT) whould have corrected commit date as their generation number,\nwhile commits in the upper layer would have topological levels as their\ngeneration. Corrected commit dates usually have much larger values than\ntopological levels. This means that if we take two commits, one from the\nupper layer, and one reachable from it in the lower layer, then the\nexpectation that the generation of a parent is smaller than the\ngeneration of a child would be violated.\n\nIt is difficult to expose this issue in a test. Since we _start_ with\nartificially low generation numbers, any commit walk that prioritizes\ngeneration numbers will walk all of the commits with high generation\nnumber before walking the commits with low generation number. In all the\ncases I tried, the commit-graph layers themselves \"protect\" any\nincorrect behavior since none of the commits in the lower layer can\nreach the commits in the upper layer.\n\nThis issue would manifest itself as a performance problem in this case,\nespecially with something like \"git log --graph\" since the low\ngeneration numbers would cause the in-degree queue to walk all of the\ncommits in the lower layer before allowing the topo-order queue to write\nanything to output (depending on the size of the upper layer).\n\nTherefore, When writing the new layer in split commit-graph, we write a\nGDAT chunk only if the topmost layer has a GDAT chunk. This guarantees\nthat if a layer has GDAT chunk, all lower layers must have a GDAT chunk\nas well.\n\nRewriting layers follows similar approach: if the topmost layer below\nthe set of layers being rewritten (in the split commit-graph chain)\nexists, and it does not contain GDAT chunk, then the result of rewrite\ndoes not have GDAT chunks either.\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                |  29 +++++-\n commit-graph.h                |   1 +\n t/t5324-split-commit-graph.sh | 181 ++++++++++++++++++++++++++++++++++\n 3 files changed, 209 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 629b2f17fbc..41a65d98738 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -610,6 +610,21 @@ static struct commit_graph *load_commit_graph_chain(struct repository *r,\n \treturn graph_chain;\n }\n \n+static void validate_mixed_generation_chain(struct commit_graph *g)\n+{\n+\tint read_generation_data;\n+\n+\tif (!g)\n+\t\treturn;\n+\n+\tread_generation_data = !!g->chunk_generation_data;\n+\n+\twhile (g) {\n+\t\tg->read_generation_data = read_generation_data;\n+\t\tg = g->base_graph;\n+\t}\n+}\n+\n struct commit_graph *read_commit_graph_one(struct repository *r,\n \t\t\t\t\t   struct object_directory *odb)\n {\n@@ -618,6 +633,8 @@ struct commit_graph *read_commit_graph_one(struct repository *r,\n \tif (!g)\n \t\tg = load_commit_graph_chain(r, odb);\n \n+\tvalidate_mixed_generation_chain(g);\n+\n \treturn g;\n }\n \n@@ -787,7 +804,7 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tif (g->chunk_generation_data) {\n+\tif (g->read_generation_data) {\n \t\toffset = (timestamp_t)get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n \n \t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n@@ -2012,6 +2029,13 @@ static void split_graph_merge_strategy(struct write_commit_graph_context *ctx)\n \t\tif (i < ctx->num_commit_graphs_after)\n \t\t\tctx->commit_graph_hash_after[i] = xstrdup(oid_to_hex(&g->oid));\n \n+\t\t/*\n+\t\t * If the topmost remaining layer has generation data chunk, the\n+\t\t * resultant layer also has generation data chunk.\n+\t\t */\n+\t\tif (i == ctx->num_commit_graphs_after - 2)\n+\t\t\tctx->write_generation_data = !!g->chunk_generation_data;\n+\n \t\ti--;\n \t\tg = g->base_graph;\n \t}\n@@ -2239,6 +2263,7 @@ int write_commit_graph(struct object_directory *odb,\n \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n \n \t\twhile (g) {\n+\t\t\tg->read_generation_data = 1;\n \t\t\tg->topo_levels = &topo_levels;\n \t\t\tg = g->base_graph;\n \t\t}\n@@ -2534,7 +2559,7 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t * also GENERATION_NUMBER_V1_MAX. Decrement to avoid extra logic\n \t\t * in the following condition.\n \t\t */\n-\t\tif (!g->chunk_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n+\t\tif (!g->read_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n \t\t\tmax_generation--;\n \n \t\tgeneration = commit_graph_generation(graph_commit);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 19a02001fde..ad52130883b 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -64,6 +64,7 @@ struct commit_graph {\n \tstruct object_directory *odb;\n \n \tuint32_t num_commits_in_base;\n+\tunsigned int read_generation_data;\n \tstruct commit_graph *base_graph;\n \n \tconst uint32_t *chunk_oid_fanout;\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 587757b62d9..8e90f3423b8 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -453,4 +453,185 @@ test_expect_success 'prevent regression for duplicate commits across layers' '\n \tgit -C dup commit-graph verify\n '\n \n+NUM_FIRST_LAYER_COMMITS=64\n+NUM_SECOND_LAYER_COMMITS=16\n+NUM_THIRD_LAYER_COMMITS=7\n+NUM_FOURTH_LAYER_COMMITS=8\n+NUM_FIFTH_LAYER_COMMITS=16\n+SECOND_LAYER_SEQUENCE_START=$(($NUM_FIRST_LAYER_COMMITS + 1))\n+SECOND_LAYER_SEQUENCE_END=$(($SECOND_LAYER_SEQUENCE_START + $NUM_SECOND_LAYER_COMMITS - 1))\n+THIRD_LAYER_SEQUENCE_START=$(($SECOND_LAYER_SEQUENCE_END + 1))\n+THIRD_LAYER_SEQUENCE_END=$(($THIRD_LAYER_SEQUENCE_START + $NUM_THIRD_LAYER_COMMITS - 1))\n+FOURTH_LAYER_SEQUENCE_START=$(($THIRD_LAYER_SEQUENCE_END + 1))\n+FOURTH_LAYER_SEQUENCE_END=$(($FOURTH_LAYER_SEQUENCE_START + $NUM_FOURTH_LAYER_COMMITS - 1))\n+FIFTH_LAYER_SEQUENCE_START=$(($FOURTH_LAYER_SEQUENCE_END + 1))\n+FIFTH_LAYER_SEQUENCE_END=$(($FIFTH_LAYER_SEQUENCE_START + $NUM_FIFTH_LAYER_COMMITS - 1))\n+\n+# Current split graph chain:\n+#\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n+\tgraphdir=\".git/objects/info/commit-graphs\" &&\n+\ttest_oid_cache <<-EOF &&\n+\toid_version sha1:1\n+\toid_version sha256:2\n+\tEOF\n+\tgit init mixed &&\n+\t(\n+\t\tcd mixed &&\n+\t\tgit config core.commitGraph true &&\n+\t\tgit config gc.writeCommitGraph false &&\n+\t\tfor i in $(test_seq $NUM_FIRST_LAYER_COMMITS)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split &&\n+\t\tgraph_read_expect $NUM_FIRST_LAYER_COMMITS &&\n+\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\t\tfor i in $(test_seq $SECOND_LAYER_SEQUENCE_START $SECOND_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n+\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 4 1\n+\t\tnum_commits: $NUM_SECOND_LAYER_COMMITS\n+\t\tchunks: oid_fanout oid_lookup commit_metadata\n+\t\tEOF\n+\t\ttest_cmp expect output &&\n+\t\tgit commit-graph verify &&\n+\t\tcat $graphdir/commit-graph-chain\n+\t)\n+'\n+\n+# The new layer will be added without generation data chunk as it was not\n+# present on the layer underneath it.\n+#\n+#      7 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'do not write generation data chunk if not present on existing tip' '\n+\tgit clone mixed mixed-no-gdat &&\n+\t(\n+\t\tcd mixed-no-gdat &&\n+\t\tfor i in $(test_seq $THIRD_LAYER_SEQUENCE_START $THIRD_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split=no-merge &&\n+\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 4 2\n+\t\tnum_commits: $NUM_THIRD_LAYER_COMMITS\n+\t\tchunks: oid_fanout oid_lookup commit_metadata\n+\t\tEOF\n+\t\ttest_cmp expect output &&\n+\t\tgit commit-graph verify\n+\t)\n+'\n+\n+# Number of commits in each layer of the split-commit graph before merge:\n+#\n+#      8 commits (No GDAT)\n+# ------------------------\n+#      7 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+# The top two layers are merged and do not have generation data chunk as layer below them does\n+# not have generation data chunk.\n+#\n+#     15 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'do not write generation data chunk if the topmost remaining layer does not have generation data chunk' '\n+\tgit clone mixed-no-gdat mixed-merge-no-gdat &&\n+\t(\n+\t\tcd mixed-merge-no-gdat &&\n+\t\tfor i in $(test_seq $FOURTH_LAYER_SEQUENCE_START $FOURTH_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split --size-multiple 1 &&\n+\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 4 2\n+\t\tnum_commits: $(($NUM_THIRD_LAYER_COMMITS + $NUM_FOURTH_LAYER_COMMITS))\n+\t\tchunks: oid_fanout oid_lookup commit_metadata\n+\t\tEOF\n+\t\ttest_cmp expect output &&\n+\t\tgit commit-graph verify\n+\t)\n+'\n+\n+# Number of commits in each layer of the split-commit graph before merge:\n+#\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     15 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+# The top three layers are merged and has generation data chunk as the topmost remaining layer\n+# has generation data chunk.\n+#\n+#     47 commits (GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'write generation data chunk if topmost remaining layer has generation data chunk' '\n+\tgit clone mixed-merge-no-gdat mixed-merge-gdat &&\n+\t(\n+\t\tcd mixed-merge-gdat &&\n+\t\tfor i in $(test_seq $FIFTH_LAYER_SEQUENCE_START $FIFTH_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split --size-multiple 1 &&\n+\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 5 1\n+\t\tnum_commits: $(($NUM_SECOND_LAYER_COMMITS + $NUM_THIRD_LAYER_COMMITS + $NUM_FOURTH_LAYER_COMMITS + $NUM_FIFTH_LAYER_COMMITS))\n+\t\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n+\t\tEOF\n+\t\ttest_cmp expect output\n+\t)\n+'\n+\n+test_expect_success 'write generation data chunk when commit-graph chain is replaced' '\n+\tgit clone mixed mixed-replace &&\n+\t(\n+\t\tcd mixed-replace &&\n+\t\tgit commit-graph write --reachable --split=replace &&\n+\t\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\t\tverify_chain_files_exist $graphdir &&\n+\t\tgraph_read_expect $(($NUM_FIRST_LAYER_COMMITS + $NUM_SECOND_LAYER_COMMITS)) &&\n+\t\tgit commit-graph verify\n+\t)\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"413158","messageId":"1694688e-0253-9d67-1982-8ce483183162@gmail.com","threadId":"53933","inReplyTo":"c4e817abf7dbcd6c99da404507ea940305c521b6.1609154168.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 01/11] commit-graph: fix regression when computing Bloom filters","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-12-30T01:35:56Z","receivedAt":"2020-12-30T01:36:55Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/28/2020 6:15 AM, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> Before computing Bloom fitlers, the commit-graph machinery uses\n\ns/fitlers/filters/\n\n> commit_gen_cmp to sort commits by generation order for improved diff\n> performance. 3d11275505 (commit-graph: examine commits by generation\n> number, 2020-03-30) claims that this sort can reduce the time spent to\n> compute Bloom filters by nearly half.\n> \n> But since c49c82aa4c (commit: move members graph_pos, generation to a\n> slab, 2020-06-17), this optimization is broken, since asking for a\n> 'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\n> while writing.\n> \n> Not all hope is lost, though: 'commit_graph_generation()' falls back to\n> comparing commits by their date when they have equal generation number,\n> and so since c49c82aa4c is purely a date comparision function. This\n\ns/comparision/comparison/\n\n> heuristic is good enough that we don't seem to loose appreciable\n> performance while computing Bloom filters. Applying this patch (compared\n> with v2.29.1) speeds up computing Bloom filters by around ~4\n> seconds.\n\nUsing \"~4 seconds\" here is odd since there is no baseline. Which\nrepository did you use?\n\nPrevious discussion used relative terms. Something like \"speeds up by\na factor of 1.25\" or something might be interesting.\n\n> So, avoid the useless 'commit_graph_generation()' while writing by\n> instead accessing the slab directly. This returns the newly-computed\n> generation numbers, and allows us to avoid the heuristic by directly\n> comparing generation numbers.\n\nThis introduces some timing restrictions to the ability for this\ncomparison function. It would be dangerous if someone extracted\nthe method for another purpose. A comment above these lines could\nwarn future developers from making that mistake, but they would\nprobably use the comparison functions in commit.c instead.\n\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 4 ++--\n>  1 file changed, 2 insertions(+), 2 deletions(-)\n> \n> diff --git a/commit-graph.c b/commit-graph.c\n> index 06f8dc1d896..caf823295f4 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>  \tconst struct commit *a = *(const struct commit **)va;\n>  \tconst struct commit *b = *(const struct commit **)vb;\n>  \n> -\tuint32_t generation_a = commit_graph_generation(a);\n> -\tuint32_t generation_b = commit_graph_generation(b);\n> +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n>  \t/* lower generation commits first */\n>  \tif (generation_a < generation_b)\n>  \t\treturn -1;\n> \n\n"},{"id":"413159","messageId":"7a0eaa06-131f-ce2d-a335-b624d64ec7e4@gmail.com","threadId":"53933","inReplyTo":"859c39eff52e32ad322969d024184971acec82e7.1609154168.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 07/11] commit-graph: implement corrected commit date","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-12-30T01:53:11Z","receivedAt":"2020-12-30T01:54:29Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/28/2020 6:16 AM, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> With most of preparations done, let's implement corrected commit date.\n> \n> The corrected commit date for a commit is defined as:\n> \n> * A commit with no parents (a root commit) has corrected commit date\n>   equal to its committer date.\n> * A commit with at least one parent has corrected commit date equal to\n>   the maximum of its commit date and one more than the largest corrected\n>   commit date among its parents.\n> \n> As a special case, a root commit with timestamp of zero (01.01.1970\n> 00:00:00Z) has corrected commit date of one, to be able to distinguish\n> from GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit\n> date).\n> \n> To minimize the space required to store corrected commit date, Git\n> stores corrected commit date offsets into the commit-graph file. The\n> corrected commit date offset for a commit is defined as the difference\n> between its corrected commit date and actual commit date.\n> \n> Storing corrected commit date requires sizeof(timestamp_t) bytes, which\n> in most cases is 64 bits (uintmax_t). However, corrected commit date\n> offsets can be safely stored using only 32-bits. This halves the size\n> of GDAT chunk, which is a reduction of around 6% in the size of\n> commit-graph file.\n> \n> However, using offsets be problematic if one of commits is malformed but\n\nHowever, using 32-bit offsets is problematic if a commit is malformed...\n\n> valid and has committerdate of 0 Unix time, as the offset would be the\n\ns/committerdate/committer date/\n\n> same as corrected commit date and thus require 64-bits to be stored\n> properly.\n> \n> While Git does not write out offsets at this stage, Git stores the\n> corrected commit dates in member generation of struct commit_graph_data.\n> It will begin writing commit date offsets with the introduction of\n> generation data chunk.\n> \n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 21 +++++++++++++++++----\n>  1 file changed, 17 insertions(+), 4 deletions(-)\n> \n> diff --git a/commit-graph.c b/commit-graph.c\n> index 1b2a015f92f..bfc3aae5f93 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -1339,9 +1339,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\t\t\tctx->commits.nr);\n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>  \t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n> +\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n>  \n>  \t\tdisplay_progress(ctx->progress, i + 1);\n> -\t\tif (level != GENERATION_NUMBER_ZERO)\n> +\t\tif (level != GENERATION_NUMBER_ZERO &&\n> +\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO)\n>  \t\t\tcontinue;\n>  \n>  \t\tcommit_list_insert(ctx->commits.list[i], &list);\n> @@ -1350,16 +1352,23 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \t\t\tstruct commit_list *parent;\n>  \t\t\tint all_parents_computed = 1;\n>  \t\t\tuint32_t max_level = 0;\n> +\t\t\ttimestamp_t max_corrected_commit_date = 0;\n>  \n>  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n>  \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n> +\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n>  \n> -\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n> +\t\t\t\tif (level == GENERATION_NUMBER_ZERO ||\n> +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO) {\n>  \t\t\t\t\tall_parents_computed = 0;\n>  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n>  \t\t\t\t\tbreak;\n> -\t\t\t\t} else if (level > max_level) {\n> -\t\t\t\t\tmax_level = level;\n> +\t\t\t\t} else {\n> +\t\t\t\t\tif (level > max_level)\n> +\t\t\t\t\t\tmax_level = level;\n> +\n> +\t\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n> +\t\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n\nnit: the \"break\" in the first case makes it so this large else block\nis unnecessary. \n\n-\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n+\t\t\t\tif (level == GENERATION_NUMBER_ZERO ||\n+\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (level > max_level) {\n-\t\t\t\t\tmax_level = level;\n+\t\t\t\t\n+\t\t\t\tif (level > max_level)\n+\t\t\t\t\tmax_level = level;\n+\n+\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n+\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n-\t\t\t\t}\n \t\t\t}\n\nThanks,\n-Stolee\n\n"},{"id":"413160","messageId":"2e89c6e1-e8e8-0d51-5670-038b4e296d93@gmail.com","threadId":"53933","inReplyTo":"a3a70a1edd0949ff3088fae625afa68fc61975df.1609154169.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 09/11] commit-graph: use generation v2 only if entire chain does","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-12-30T03:23:54Z","receivedAt":"2020-12-30T03:24:53Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/28/2020 6:16 AM, Abhishek Kumar via GitGitGadget wrote:\n> From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\n...\n\n> +static void validate_mixed_generation_chain(struct commit_graph *g)\n> +{\n> +\tint read_generation_data;\n> +\n> +\tif (!g)\n> +\t\treturn;\n> +\n> +\tread_generation_data = !!g->chunk_generation_data;\n> +\n> +\twhile (g) {\n> +\t\tg->read_generation_data = read_generation_data;\n> +\t\tg = g->base_graph;\n> +\t}\n> +}\n> +\n\nThis method exists to say \"use generation v2 if the top layer has it\"\nand that helps with the future layer checks.\n\n> @@ -2239,6 +2263,7 @@ int write_commit_graph(struct object_directory *odb,\n>  \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n>  \n>  \t\twhile (g) {\n> +\t\t\tg->read_generation_data = 1;\n>  \t\t\tg->topo_levels = &topo_levels;\n>  \t\t\tg = g->base_graph;\n>  \t\t}\n\nHowever, here you just turn them on automatically.\n\nI think the diff you want is here:\n\n \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n \n+ \t\tvalidate_mixed_generation_chain(g);\n+ \n \t\twhile (g) {\n \t\t\tg->topo_levels = &topo_levels;\n \t\t\tg = g->base_graph;\n \t\t}\n\nBut maybe you have a good reason for what you already have.\n\nI paid attention to this because I hit a problem in my local testing.\nAfter trying to reproduce it, I think the root cause is that I had a\ncommit-graph that was written by an older version of your series, so\nit caused an unexpected pairing of an \"offset required\" bit but no\noffset chunk.\n\nPerhaps this diff is required in the proper place to avoid the\nsegfault I hit, in the case of a malformed commit-graph file:\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex c8d7ed1330..d264c90868 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -822,6 +822,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \t\toffset = (timestamp_t)get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n \n \t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n+\t\t\tif (!g->chunk_generation_data_overflow)\n+\t\t\t\tdie(_(\"commit-graph requires overflow generation data but has none\"));\n+\n \t\t\toffset_pos = offset ^ CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW;\n \t\t\tgraph_data->generation = get_be64(g->chunk_generation_data_overflow + 8 * offset_pos);\n \t\t} else\n\nYour tests in this patch seem very thorough, covering all the cases\nI could think to create this strange situation. I even tried creating\ncases where the overflow would be necessary. The following test actually\nfails on the \"graph_read_expect 6\" due to the extra chunk, not the 'write'\nprocess I was trying to trick into failure.\n\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 8e90f3423b..cfef8e52b9 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -453,6 +453,20 @@ test_expect_success 'prevent regression for duplicate commits across layers' '\n        git -C dup commit-graph verify\n '\n \n+test_expect_success 'upgrade to generation data succeeds when there was none' '\n+\t(\n+\t\tcd dup &&\n+\t\trm -rf .git/objects/info/commit-graph* &&\n+\t\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph \\\n+\t\t\twrite --reachable &&\n+\t\tGIT_COMMITTER_DATE=\"1980-01-01 00:00\" git commit --allow-empty -m one &&\n+\t\tGIT_COMMITTER_DATE=\"2090-01-01 00:00\" git commit --allow-empty -m two &&\n+\t\tGIT_COMMITTER_DATE=\"2000-01-01 00:00\" git commit --allow-empty -m three &&\n+\t\tgit commit-graph write --reachable &&\n+\t\tgraph_read_expect 6\n+\t)\n+'\n+\n NUM_FIRST_LAYER_COMMITS=64\n NUM_SECOND_LAYER_COMMITS=16\n NUM_THIRD_LAYER_COMMITS=7\n\nThanks,\n-Stolee\n"},{"id":"413163","messageId":"1adabda6-b80b-d543-f6c0-570dadbe589b@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-12-30T04:35:56Z","receivedAt":"2020-12-30T04:36:39Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/28/2020 6:15 AM, Abhishek Kumar via GitGitGadget wrote:\n> This patch series implements the corrected commit date offsets as generation\n> number v2, along with other pre-requisites.\n\nAbhishek,\n\nThank you for this version. I appreciate your hard work on this topic,\nespecially after GSoC ended and you returned to being a full-time student.\n\nMy hope was that I could completely approve this series and only provide\nforward-fixes from here on out, as necessary. I think there are a few minor\ntypos that you might want to address, but I was also able to understand your\nintention.\n\nI did make a particular case about a SEGFAULT I hit that I have been unable\nto replicate. I saw it both in my copy of torvalds/linux and of\nchromium/chromium. I have the file for chromium/chromium that is in a bad\nstate where a GDAT value includes the bit saying it should be in the long\noffsets chunk, but that chunk doesn't exist. Further, that chunk doesn't\nexist in a from-scratch write.\n\nI'm now taking backups of my existing commit-graph files before any later\ntest, but it doesn't repro for my Git repository or any other repo I try on\npurpose.\n\nHowever, I did some performance testing to double-check your numbers. I sent\na patch [1] that helps with some of the hard numbers.\n\n[1] https://lore.kernel.org/git/pull.828.git.1609302714183.gitgitgadget@gmail.com/\n\nThe big question is whether the overhead from using a slab to store the\ngeneration values is worth it. I still think it is, for these reasons:\n\n1. Generation number v2 is measurably better than v1 in most user cases.\n\n2. Generation number v2 is slower than using committer date due to the\n   overhead, but _guarantees correctness_.\n\nI like to use \"git log --graph -<N>\" to compare against topological levels\n(v1), for various levels of <N>. When <N> is small, we hope to minimize\nthe amount we need to walk using the extra commit-date information as an\nassistance. Repos like git/git and torvalds/linux use the philosophy of\n\"base your changes on oldest applicable commit\" enough that v1 struggles\nsometimes.\n\ngit/git: N=1000\n\n\tBenchmark #1: baseline\n\tTime (mean ± σ):     100.3 ms ±   4.2 ms    [User: 89.0 ms, System: 11.3 ms]\n\tRange (min … max):    94.5 ms … 105.1 ms    28 runs\n\t\n\tBenchmark #2: test\n\tTime (mean ± σ):      35.8 ms ±   3.1 ms    [User: 29.6 ms, System: 6.2 ms]\n\tRange (min … max):    29.8 ms …  40.6 ms    81 runs\n\t\n\tSummary\n\t'test' ran\n\t2.80 ± 0.27 times faster than 'baseline'\n\nThis is a dramatic improvement! Using my topo-walk stats commit, I see that\nv1 walks 58,805 commits as part of the in-degree walk while v2 only walks\n4,335 commits!\n\ntorvalds/linux: N=1000 (starting at v5.10)\n\n\tBenchmark #1: baseline\n\tTime (mean ± σ):      90.8 ms ±   3.7 ms    [User: 75.2 ms, System: 15.6 ms]\n\tRange (min … max):    85.2 ms …  96.2 ms    31 runs\n\t\n\tBenchmark #2: test\n\tTime (mean ± σ):      49.2 ms ±   3.5 ms    [User: 36.9 ms, System: 12.3 ms]\n\tRange (min … max):    42.9 ms …  54.0 ms    61 runs\n\t\n\tSummary\n\t'test' ran\n\t1.85 ± 0.15 times faster than 'baseline'\n\nSimilarly, v1 walked 38,161 commits compared to 4,340 by v2.\n\nIf I increase N to something like 10,000, then usually these values get\nwashed out due to the width of the parallel topics.\n\nThe place we were still using commit-date as a heuristic was paint_down_to_common\nwhich caused a regression the first time we used v1, at least for certain cases.\n\nSpecifically, computing the merge-base in torvalds/linux between v4.8 and v4.9\nhit a strangeness about a pair of recent commits both based on a very old commit,\nbut the generation numbers forced walking farther than necessary. This doesn't\nhappen with v2, but we see the overhead cost of the slabs:\n\n\tBenchmark #1: baseline\n\tTime (mean ± σ):     112.9 ms ±   2.8 ms    [User: 96.5 ms, System: 16.3 ms]\n\tRange (min … max):   107.7 ms … 118.0 ms    26 runs\n\t\n\tBenchmark #2: test\n\tTime (mean ± σ):     147.1 ms ±   5.2 ms    [User: 132.7 ms, System: 14.3 ms]\n\tRange (min … max):   141.4 ms … 162.2 ms    18 runs\n\t\n\tSummary\n\t'baseline' ran\n\t1.30 ± 0.06 times faster than 'test'\n\nThe overhead still exists for a more recent pair of versions (v5.0 and v5.1):\n\n\tBenchmark #1: baseline\n\tTime (mean ± σ):      25.1 ms ±   3.2 ms    [User: 18.6 ms, System: 6.5 ms]\n\tRange (min … max):    19.0 ms …  32.8 ms    99 runs\n\t\n\tBenchmark #2: test\n\tTime (mean ± σ):      33.3 ms ±   3.3 ms    [User: 26.5 ms, System: 6.9 ms]\n\tRange (min … max):    27.0 ms …  38.4 ms    105 runs\n\t\n\tSummary\n\t'baseline' ran\n\t1.33 ± 0.22 times faster than 'test'\n\nI still think this overhead is worth it. In case not everyone agrees, it _might_\nbe worth a command-line option to skip the GDAT chunk. That also prevents an\nability to eventually wean entirely of generation number v1 and allow the commit\ndate to take the full 64-bit column (instead of only 34 bits, saving 30 for\ntopo-levels).\n\nAgain, such a modification should not be considered required for this series.\n\n> ----------------------------------------------------------------------------\n> \n> Improvements left for a future series:\n> \n>  * Save commits with generation data overflow and extra edge commits instead\n>    of looping over all commits. cf. 858sbel67n.fsf@gmail.com\n>  * Verify both topological levels and corrected commit dates when present.\n>    cf. 85pn4tnk8u.fsf@gmail.com\n\nThese seem like reasonable things to delay for a later series\nor for #leftoverbits\n\nThanks,\n-Stolee\n\n"},{"id":"413440","messageId":"20210105094535.GN8396@szeder.dev","threadId":"53933","inReplyTo":"c4e817abf7dbcd6c99da404507ea940305c521b6.1609154168.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 01/11] commit-graph: fix regression when computing Bloom filters","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2021-01-05T09:45:35Z","receivedAt":"2021-01-05T09:46:36Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Dec 28, 2020 at 11:15:58AM +0000, Abhishek Kumar via GitGitGadget wrote:\n> Before computing Bloom fitlers, the commit-graph machinery uses\n> commit_gen_cmp to sort commits by generation order for improved diff\n> performance. 3d11275505 (commit-graph: examine commits by generation\n> number, 2020-03-30) claims that this sort can reduce the time spent to\n> compute Bloom filters by nearly half.\n\nThat's true, though there are repositories where it has basically no\neffect.  Alas we can't directly test it, because in 3d11275505 there\nis no '--changed-paths' option yet... one has to revert 3d11275505 on\ntop of d38e07b8c4 (commit-graph: add --changed-paths option to write\nsubcommand, 2020-04-06) to make any runtime comparisons ('git\ncommit-graph write --reachable --changed-paths', best of five):\n\n                   Sorting by\n               pack    | generation\n             position  |\n    -------------------+------------\n    gcc      114.821s  |    38.963s \n    git        8.896s  |     5.620s\n    linux    209.984s  |   104.900s\n    webkit    35.193s  |    35.482s\n\nNote the almost 3x speedup in the gcc repository, and the basically\nnegligible slowdown in the webkit repo.\n\n> But since c49c82aa4c (commit: move members graph_pos, generation to a\n> slab, 2020-06-17), this optimization is broken, since asking for a\n> 'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\n> while writing.\n\nI wouldn't say that c49c82aa4c broke this optimisation, because:\n\ndid not break that optimization.  Though, sadly, it's not\nmentioned in 3d11275505's commit message, when commit_gen_cmp()\ncompares two commits with identical generation numbers, then it\ndoesn't leave them unsorted, but falls back to use their committer\ndate as a tie-braker.  This means that after c49c82aa4c the commits\nare sorted by committer date, which appears to be so good a heuristic\nfor Bloom filter computation that there is barely any slowdown\ncompared to sorting by generation numbers:\n\n> Not all hope is lost, though: 'commit_graph_generation()' falls back to\n\nYou mean commit_gen_cmp() here.\n\n> comparing commits by their date when they have equal generation number,\n> and so since c49c82aa4c is purely a date comparision function. This\n> heuristic is good enough that we don't seem to loose appreciable\n> performance while computing Bloom filters.\n\nIndeed, c49c82aa4c barely caused any runtime difference in the\nrepositories I usually use to test modified path Bloom filter\nperformance:\n\n                 c49c82aa4c^  c49c82aa4c\n  ---------------------------------------------\n  android-base     43.057s     43.091s   0.07%\n  cmssw            21.781s     21.856s   0.34%\n  cpython           9.626s      9.724s   1.01%\n  elasticsearch    18.049s     18.224s   0.96%\n  gcc              40.312s     40.255s  -0.14%\n  gecko-dev       104.515s    104.740s   0.21%\n  git               5.559s      5.570s   0.19%\n  glibc             4.455s      4.468s   0.29%\n  go                4.009s      4.016s   0.17%\n  homebrew-cask    30.759s     30.523s  -0.76%\n  homebrew-core    57.122s     56.553s  -0.99%\n  jdk              18.297s     18.364s   0.36%\n  linux           104.499s    105.302s   0.76%\n  llvm-project     34.074s     34.446s   1.09%\n  rails             6.472s      6.486s   0.21%\n  rust             14.943s     14.947s   0.02%\n  tensorflow       13.362s     13.477s   0.86%\n  webkit           34.583s     34.601s   0.05%\n\n> Applying this patch (compared\n> with v2.29.1) speeds up computing Bloom filters by around ~4\n> seconds.\n\nWithout a baseline and knowing which repo, this \"~4 seconds\" is\nmeaningless.\n\nHere are my results comparing this fix to v2.30.0, best of five:\n\n                              v2.30.0 +\n                   v2.30.0    this fix\n  ---------------------------------------------\n  android-base     42.786s     42.933s   0.34%\n  cmssw            20.229s     20.160s  -0.34%\n  cpython           9.616s      9.647s   0.32%\n  elasticsearch    16.859s     16.936s   0.45%\n  gcc              38.909s     36.889s  -5.19%\n  gecko-dev        99.417s     98.558s  -0.86%\n  git               5.620s      5.509s  -1.97%\n  glibc             4.307s      4.301s  -0.13%\n  go                3.971s      3.938s  -0.83%\n  homebrew-cask    31.262s     30.283s  -3.13%\n  homebrew-core    57.842s     55.663s  -3.76%\n  jdk              12.557s     12.251s  -2.43%\n  linux            94.335s     94.760s   0.45%\n  llvm-project     34.432s     33.988s  -1.28%\n  rails             6.481s      6.454s  -0.41%\n  rust             14.772s     14.601s  -1.15%\n  tensorflow       11.759s     11.711s  -0.40%\n  webkit           33.917s     33.759s  -0.46%\n\n> So, avoid the useless 'commit_graph_generation()' while writing by\n> instead accessing the slab directly. This returns the newly-computed\n> generation numbers, and allows us to avoid the heuristic by directly\n> comparing generation numbers.\n> \n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  commit-graph.c | 4 ++--\n>  1 file changed, 2 insertions(+), 2 deletions(-)\n> \n> diff --git a/commit-graph.c b/commit-graph.c\n> index 06f8dc1d896..caf823295f4 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n>  \tconst struct commit *a = *(const struct commit **)va;\n>  \tconst struct commit *b = *(const struct commit **)vb;\n>  \n> -\tuint32_t generation_a = commit_graph_generation(a);\n> -\tuint32_t generation_b = commit_graph_generation(b);\n> +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n>  \t/* lower generation commits first */\n>  \tif (generation_a < generation_b)\n>  \t\treturn -1;\n> -- \n> gitgitgadget\n> \n"},{"id":"413441","messageId":"20210105094740.GO8396@szeder.dev","threadId":"53933","inReplyTo":"20210105094535.GN8396@szeder.dev","subject":"Re: [PATCH v5 01/11] commit-graph: fix regression when computing Bloom filters","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2021-01-05T09:47:40Z","receivedAt":"2021-01-05T09:49:12Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Tue, Jan 05, 2021 at 10:45:35AM +0100, SZEDER Gábor wrote:\n> > But since c49c82aa4c (commit: move members graph_pos, generation to a\n> > slab, 2020-06-17), this optimization is broken, since asking for a\n> > 'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\n> > while writing.\n> \n> I wouldn't say that c49c82aa4c broke this optimisation, because:\n> \n> did not break that optimization.  Though, sadly, it's not\n> mentioned in 3d11275505's commit message, when commit_gen_cmp()\n> compares two commits with identical generation numbers, then it\n> doesn't leave them unsorted, but falls back to use their committer\n> date as a tie-braker.  This means that after c49c82aa4c the commits\n> are sorted by committer date, which appears to be so good a heuristic\n> for Bloom filter computation that there is barely any slowdown\n> compared to sorting by generation numbers:\n\nGaah, scratch this paragraph; I first misunderstood what you wrote in\nthe paragraph below, but then forgot to remove it.\n\n> > Not all hope is lost, though: 'commit_graph_generation()' falls back to\n> \n> You mean commit_gen_cmp() here.\n> \n> > comparing commits by their date when they have equal generation number,\n> > and so since c49c82aa4c is purely a date comparision function. This\n> > heuristic is good enough that we don't seem to loose appreciable\n> > performance while computing Bloom filters.\n"},{"id":"413781","messageId":"X/fxlVc8UK7FQRpP@Abhishek-Arch","threadId":"53933","inReplyTo":"1694688e-0253-9d67-1982-8ce483183162@gmail.com","subject":"Re: [PATCH v5 01/11] commit-graph: fix regression when computing Bloom filters","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2021-01-08T05:45:57Z","receivedAt":"2021-01-08T05:46:35Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Dec 29, 2020 at 08:35:56PM -0500, Derrick Stolee wrote:\n> On 12/28/2020 6:15 AM, Abhishek Kumar via GitGitGadget wrote:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > \n> > Before computing Bloom fitlers, the commit-graph machinery uses\n> \n> s/fitlers/filters/\n> \n> > commit_gen_cmp to sort commits by generation order for improved diff\n> > performance. 3d11275505 (commit-graph: examine commits by generation\n> > number, 2020-03-30) claims that this sort can reduce the time spent to\n> > compute Bloom filters by nearly half.\n> > \n> > But since c49c82aa4c (commit: move members graph_pos, generation to a\n> > slab, 2020-06-17), this optimization is broken, since asking for a\n> > 'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\n> > while writing.\n> > \n> > Not all hope is lost, though: 'commit_graph_generation()' falls back to\n> > comparing commits by their date when they have equal generation number,\n> > and so since c49c82aa4c is purely a date comparision function. This\n> \n> s/comparision/comparison/\n> \n> > heuristic is good enough that we don't seem to loose appreciable\n> > performance while computing Bloom filters. Applying this patch (compared\n> > with v2.29.1) speeds up computing Bloom filters by around ~4\n> > seconds.\n> \n> Using \"~4 seconds\" here is odd since there is no baseline. Which\n> repository did you use?\n> \n\nI used the linux repository, will mention that.\n\n> Previous discussion used relative terms. Something like \"speeds up by\n> a factor of 1.25\" or something might be interesting.\n> \n\nAs SZEDER Gábor found, the improvements are rather minor - ranging from\n0.40% to 5.19% [1]. I want to make sure this is the correct way to word\nin the commit message:\n\nApplying this patch (compared with v2.30.0) speeds up computing Bloom\nfilters by factors ranging from 0.40% to 5.19% on various\nrepositories. \n\nhttps://lore.kernel.org/git/20210105094535.GN8396@szeder.dev/\n\n> > So, avoid the useless 'commit_graph_generation()' while writing by\n> > instead accessing the slab directly. This returns the newly-computed\n> > generation numbers, and allows us to avoid the heuristic by directly\n> > comparing generation numbers.\n> \n> This introduces some timing restrictions to the ability for this\n> comparison function. It would be dangerous if someone extracted\n> the method for another purpose. A comment above these lines could\n> warn future developers from making that mistake, but they would\n> probably use the comparison functions in commit.c instead.\n> \n\nSure, will add a comment above.\n\n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c | 4 ++--\n> >  1 file changed, 2 insertions(+), 2 deletions(-)\n> > \n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index 06f8dc1d896..caf823295f4 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n> >  \tconst struct commit *a = *(const struct commit **)va;\n> >  \tconst struct commit *b = *(const struct commit **)vb;\n> >  \n> > -\tuint32_t generation_a = commit_graph_generation(a);\n> > -\tuint32_t generation_b = commit_graph_generation(b);\n> > +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> > +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n> >  \t/* lower generation commits first */\n> >  \tif (generation_a < generation_b)\n> >  \t\treturn -1;\n> > \n> \n"},{"id":"413782","messageId":"X/fy6vWfCCVuApTE@Abhishek-Arch","threadId":"53933","inReplyTo":"20210105094535.GN8396@szeder.dev","subject":"Re: [PATCH v5 01/11] commit-graph: fix regression when computing Bloom filters","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2021-01-08T05:51:38Z","receivedAt":"2021-01-08T05:52:11Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Jan 05, 2021 at 10:45:35AM +0100, SZEDER Gábor wrote:\n> On Mon, Dec 28, 2020 at 11:15:58AM +0000, Abhishek Kumar via GitGitGadget wrote:\n> > Before computing Bloom fitlers, the commit-graph machinery uses\n> > commit_gen_cmp to sort commits by generation order for improved diff\n> > performance. 3d11275505 (commit-graph: examine commits by generation\n> > number, 2020-03-30) claims that this sort can reduce the time spent to\n> > compute Bloom filters by nearly half.\n> \n> That's true, though there are repositories where it has basically no\n> effect.  Alas we can't directly test it, because in 3d11275505 there\n> is no '--changed-paths' option yet... one has to revert 3d11275505 on\n> top of d38e07b8c4 (commit-graph: add --changed-paths option to write\n> subcommand, 2020-04-06) to make any runtime comparisons ('git\n> commit-graph write --reachable --changed-paths', best of five):\n> \n>                    Sorting by\n>                pack    | generation\n>              position  |\n>     -------------------+------------\n>     gcc      114.821s  |    38.963s \n>     git        8.896s  |     5.620s\n>     linux    209.984s  |   104.900s\n>     webkit    35.193s  |    35.482s\n> \n> Note the almost 3x speedup in the gcc repository, and the basically\n> negligible slowdown in the webkit repo.\n> \n> > But since c49c82aa4c (commit: move members graph_pos, generation to a\n> > slab, 2020-06-17), this optimization is broken, since asking for a\n> > 'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\n> > while writing.\n> \n> I wouldn't say that c49c82aa4c broke this optimisation, because:\n> \n> did not break that optimization.  Though, sadly, it's not\n> mentioned in 3d11275505's commit message, when commit_gen_cmp()\n> compares two commits with identical generation numbers, then it\n> doesn't leave them unsorted, but falls back to use their committer\n> date as a tie-braker.  This means that after c49c82aa4c the commits\n> are sorted by committer date, which appears to be so good a heuristic\n> for Bloom filter computation that there is barely any slowdown\n> compared to sorting by generation numbers:\n> \n> > Not all hope is lost, though: 'commit_graph_generation()' falls back to\n> \n> You mean commit_gen_cmp() here.\n> \n\nYes, fixed.\n\n> > comparing commits by their date when they have equal generation number,\n> > and so since c49c82aa4c is purely a date comparision function. This\n> > heuristic is good enough that we don't seem to loose appreciable\n> > performance while computing Bloom filters.\n> \n> Indeed, c49c82aa4c barely caused any runtime difference in the\n> repositories I usually use to test modified path Bloom filter\n> performance:\n> \n>                  c49c82aa4c^  c49c82aa4c\n>   ---------------------------------------------\n>   android-base     43.057s     43.091s   0.07%\n>   cmssw            21.781s     21.856s   0.34%\n>   cpython           9.626s      9.724s   1.01%\n>   elasticsearch    18.049s     18.224s   0.96%\n>   gcc              40.312s     40.255s  -0.14%\n>   gecko-dev       104.515s    104.740s   0.21%\n>   git               5.559s      5.570s   0.19%\n>   glibc             4.455s      4.468s   0.29%\n>   go                4.009s      4.016s   0.17%\n>   homebrew-cask    30.759s     30.523s  -0.76%\n>   homebrew-core    57.122s     56.553s  -0.99%\n>   jdk              18.297s     18.364s   0.36%\n>   linux           104.499s    105.302s   0.76%\n>   llvm-project     34.074s     34.446s   1.09%\n>   rails             6.472s      6.486s   0.21%\n>   rust             14.943s     14.947s   0.02%\n>   tensorflow       13.362s     13.477s   0.86%\n>   webkit           34.583s     34.601s   0.05%\n> \n> > Applying this patch (compared\n> > with v2.29.1) speeds up computing Bloom filters by around ~4\n> > seconds.\n> \n> Without a baseline and knowing which repo, this \"~4 seconds\" is\n> meaningless.\n> \n> Here are my results comparing this fix to v2.30.0, best of five:\n> \n>                               v2.30.0 +\n>                    v2.30.0    this fix\n>   ---------------------------------------------\n>   android-base     42.786s     42.933s   0.34%\n>   cmssw            20.229s     20.160s  -0.34%\n>   cpython           9.616s      9.647s   0.32%\n>   elasticsearch    16.859s     16.936s   0.45%\n>   gcc              38.909s     36.889s  -5.19%\n>   gecko-dev        99.417s     98.558s  -0.86%\n>   git               5.620s      5.509s  -1.97%\n>   glibc             4.307s      4.301s  -0.13%\n>   go                3.971s      3.938s  -0.83%\n>   homebrew-cask    31.262s     30.283s  -3.13%\n>   homebrew-core    57.842s     55.663s  -3.76%\n>   jdk              12.557s     12.251s  -2.43%\n>   linux            94.335s     94.760s   0.45%\n>   llvm-project     34.432s     33.988s  -1.28%\n>   rails             6.481s      6.454s  -0.41%\n>   rust             14.772s     14.601s  -1.15%\n>   tensorflow       11.759s     11.711s  -0.40%\n>   webkit           33.917s     33.759s  -0.46%\n>\n\nThank you for the detailed performance benchmarking.\n\n> \n> > So, avoid the useless 'commit_graph_generation()' while writing by\n> > instead accessing the slab directly. This returns the newly-computed\n> > generation numbers, and allows us to avoid the heuristic by directly\n> > comparing generation numbers.\n> > \n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c | 4 ++--\n> >  1 file changed, 2 insertions(+), 2 deletions(-)\n> > \n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index 06f8dc1d896..caf823295f4 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -144,8 +144,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n> >  \tconst struct commit *a = *(const struct commit **)va;\n> >  \tconst struct commit *b = *(const struct commit **)vb;\n> >  \n> > -\tuint32_t generation_a = commit_graph_generation(a);\n> > -\tuint32_t generation_b = commit_graph_generation(b);\n> > +\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n> > +\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n> >  \t/* lower generation commits first */\n> >  \tif (generation_a < generation_b)\n> >  \t\treturn -1;\n> > -- \n> > gitgitgadget\n> > \n\nThanks\n- Abhishek\n"},{"id":"413962","messageId":"X/rxL+ofMI4LNSYw@Abhishek-Arch","threadId":"53933","inReplyTo":"7a0eaa06-131f-ce2d-a335-b624d64ec7e4@gmail.com","subject":"Re: [PATCH v5 07/11] commit-graph: implement corrected commit date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2021-01-10T12:21:03Z","receivedAt":"2021-01-10T12:21:27Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Dec 29, 2020 at 08:53:11PM -0500, Derrick Stolee wrote:\n> On 12/28/2020 6:16 AM, Abhishek Kumar via GitGitGadget wrote:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > \n> > With most of preparations done, let's implement corrected commit date.\n> > \n> > The corrected commit date for a commit is defined as:\n> > \n> > * A commit with no parents (a root commit) has corrected commit date\n> >   equal to its committer date.\n> > * A commit with at least one parent has corrected commit date equal to\n> >   the maximum of its commit date and one more than the largest corrected\n> >   commit date among its parents.\n> > \n> > As a special case, a root commit with timestamp of zero (01.01.1970\n> > 00:00:00Z) has corrected commit date of one, to be able to distinguish\n> > from GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit\n> > date).\n> > \n> > To minimize the space required to store corrected commit date, Git\n> > stores corrected commit date offsets into the commit-graph file. The\n> > corrected commit date offset for a commit is defined as the difference\n> > between its corrected commit date and actual commit date.\n> > \n> > Storing corrected commit date requires sizeof(timestamp_t) bytes, which\n> > in most cases is 64 bits (uintmax_t). However, corrected commit date\n> > offsets can be safely stored using only 32-bits. This halves the size\n> > of GDAT chunk, which is a reduction of around 6% in the size of\n> > commit-graph file.\n> > \n> > However, using offsets be problematic if one of commits is malformed but\n> \n> However, using 32-bit offsets is problematic if a commit is malformed...\n> \n> > valid and has committerdate of 0 Unix time, as the offset would be the\n> \n> s/committerdate/committer date/\n> \n> > same as corrected commit date and thus require 64-bits to be stored\n> > properly.\n> > \n> > While Git does not write out offsets at this stage, Git stores the\n> > corrected commit dates in member generation of struct commit_graph_data.\n> > It will begin writing commit date offsets with the introduction of\n> > generation data chunk.\n> > \n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  commit-graph.c | 21 +++++++++++++++++----\n> >  1 file changed, 17 insertions(+), 4 deletions(-)\n> > \n> > diff --git a/commit-graph.c b/commit-graph.c\n> > index 1b2a015f92f..bfc3aae5f93 100644\n> > --- a/commit-graph.c\n> > +++ b/commit-graph.c\n> > @@ -1339,9 +1339,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\t\t\tctx->commits.nr);\n> >  \tfor (i = 0; i < ctx->commits.nr; i++) {\n> >  \t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n> > +\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n> >  \n> >  \t\tdisplay_progress(ctx->progress, i + 1);\n> > -\t\tif (level != GENERATION_NUMBER_ZERO)\n> > +\t\tif (level != GENERATION_NUMBER_ZERO &&\n> > +\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO)\n> >  \t\t\tcontinue;\n> >  \n> >  \t\tcommit_list_insert(ctx->commits.list[i], &list);\n> > @@ -1350,16 +1352,23 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n> >  \t\t\tstruct commit_list *parent;\n> >  \t\t\tint all_parents_computed = 1;\n> >  \t\t\tuint32_t max_level = 0;\n> > +\t\t\ttimestamp_t max_corrected_commit_date = 0;\n> >  \n> >  \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n> >  \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n> > +\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n> >  \n> > -\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n> > +\t\t\t\tif (level == GENERATION_NUMBER_ZERO ||\n> > +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO) {\n> >  \t\t\t\t\tall_parents_computed = 0;\n> >  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n> >  \t\t\t\t\tbreak;\n> > -\t\t\t\t} else if (level > max_level) {\n> > -\t\t\t\t\tmax_level = level;\n> > +\t\t\t\t} else {\n> > +\t\t\t\t\tif (level > max_level)\n> > +\t\t\t\t\t\tmax_level = level;\n> > +\n> > +\t\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n> > +\t\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n> \n> nit: the \"break\" in the first case makes it so this large else block\n> is unnecessary. \n\nThanks, removed.\n\n> \n> -\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n> +\t\t\t\tif (level == GENERATION_NUMBER_ZERO ||\n> +\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO) {\n>  \t\t\t\t\tall_parents_computed = 0;\n>  \t\t\t\t\tcommit_list_insert(parent->item, &list);\n>  \t\t\t\t\tbreak;\n> -\t\t\t\t} else if (level > max_level) {\n> -\t\t\t\t\tmax_level = level;\n> +\t\t\t\t\n> +\t\t\t\tif (level > max_level)\n> +\t\t\t\t\tmax_level = level;\n> +\n> +\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n> +\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n> -\t\t\t\t}\n>  \t\t\t}\n> \n> Thanks,\n> -Stolee\n> \n\nThanks\n- Abhishek\n"},{"id":"413966","messageId":"X/r9i0HJFEGxuyW/@Abhishek-Arch","threadId":"53933","inReplyTo":"2e89c6e1-e8e8-0d51-5670-038b4e296d93@gmail.com","subject":"Re: [PATCH v5 09/11] commit-graph: use generation v2 only if entire chain does","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2021-01-10T13:13:47Z","receivedAt":"2021-01-10T13:14:23Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Dec 29, 2020 at 10:23:54PM -0500, Derrick Stolee wrote:\n> On 12/28/2020 6:16 AM, Abhishek Kumar via GitGitGadget wrote:\n> > From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> \n> ...\n> \n> > +static void validate_mixed_generation_chain(struct commit_graph *g)\n> > +{\n> > +\tint read_generation_data;\n> > +\n> > +\tif (!g)\n> > +\t\treturn;\n> > +\n> > +\tread_generation_data = !!g->chunk_generation_data;\n> > +\n> > +\twhile (g) {\n> > +\t\tg->read_generation_data = read_generation_data;\n> > +\t\tg = g->base_graph;\n> > +\t}\n> > +}\n> > +\n> \n> This method exists to say \"use generation v2 if the top layer has it\"\n> and that helps with the future layer checks.\n> \n> > @@ -2239,6 +2263,7 @@ int write_commit_graph(struct object_directory *odb,\n> >  \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n> >  \n> >  \t\twhile (g) {\n> > +\t\t\tg->read_generation_data = 1;\n> >  \t\t\tg->topo_levels = &topo_levels;\n> >  \t\t\tg = g->base_graph;\n> >  \t\t}\n> \n> However, here you just turn them on automatically.\n> \n> I think the diff you want is here:\n> \n>  \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n>  \n> + \t\tvalidate_mixed_generation_chain(g);\n> + \n>  \t\twhile (g) {\n>  \t\t\tg->topo_levels = &topo_levels;\n>  \t\t\tg = g->base_graph;\n>  \t\t}\n> \n> But maybe you have a good reason for what you already have.\n> \n\nThanks, that was an oversight.\n\nMy (incorrect) reasoning at the time was:\n\nSince we are computing both topological levels and corrected commit\ndates, we can read corrected commit dates from layers with a GDAT chunk\nhidden below non-GDAT layer.\n\nBut we end up storing both corrected commit date offsets (for a layers with\nGDAT chunk) and topological level (for layers without GDAT chunk) in the\nsame slab with no way to distinguish between the two!\n\n> I paid attention to this because I hit a problem in my local testing.\n> After trying to reproduce it, I think the root cause is that I had a\n> commit-graph that was written by an older version of your series, so\n> it caused an unexpected pairing of an \"offset required\" bit but no\n> offset chunk.\n> \n> Perhaps this diff is required in the proper place to avoid the\n> segfault I hit, in the case of a malformed commit-graph file:\n> \n> diff --git a/commit-graph.c b/commit-graph.c\n> index c8d7ed1330..d264c90868 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -822,6 +822,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n>  \t\toffset = (timestamp_t)get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n>  \n>  \t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n> +\t\t\tif (!g->chunk_generation_data_overflow)\n> +\t\t\t\tdie(_(\"commit-graph requires overflow generation data but has none\"));\n> +\n>  \t\t\toffset_pos = offset ^ CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW;\n>  \t\t\tgraph_data->generation = get_be64(g->chunk_generation_data_overflow + 8 * offset_pos);\n>  \t\t} else\n> \n> Your tests in this patch seem very thorough, covering all the cases\n> I could think to create this strange situation. I even tried creating\n> cases where the overflow would be necessary. The following test actually\n> fails on the \"graph_read_expect 6\" due to the extra chunk, not the 'write'\n> process I was trying to trick into failure.\n> \n> diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> index 8e90f3423b..cfef8e52b9 100755\n> --- a/t/t5324-split-commit-graph.sh\n> +++ b/t/t5324-split-commit-graph.sh\n> @@ -453,6 +453,20 @@ test_expect_success 'prevent regression for duplicate commits across layers' '\n>         git -C dup commit-graph verify\n>  '\n>  \n> +test_expect_success 'upgrade to generation data succeeds when there was none' '\n> +\t(\n> +\t\tcd dup &&\n> +\t\trm -rf .git/objects/info/commit-graph* &&\n> +\t\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph \\\n> +\t\t\twrite --reachable &&\n> +\t\tGIT_COMMITTER_DATE=\"1980-01-01 00:00\" git commit --allow-empty -m one &&\n> +\t\tGIT_COMMITTER_DATE=\"2090-01-01 00:00\" git commit --allow-empty -m two &&\n> +\t\tGIT_COMMITTER_DATE=\"2000-01-01 00:00\" git commit --allow-empty -m three &&\n> +\t\tgit commit-graph write --reachable &&\n> +\t\tgraph_read_expect 6\n> +\t)\n> +'\n\nI am not sure what this test adds over the existing generation data\noverflow related tests added in t5318-commit-graph.sh\n\n> +\n>  NUM_FIRST_LAYER_COMMITS=64\n>  NUM_SECOND_LAYER_COMMITS=16\n>  NUM_THIRD_LAYER_COMMITS=7\n> \n> Thanks,\n> -Stolee\n\nThanks\n- Abhishek\n"},{"id":"413968","messageId":"X/sJ7gd7g+nnvqNt@Abhishek-Arch","threadId":"53933","inReplyTo":"1adabda6-b80b-d543-f6c0-570dadbe589b@gmail.com","subject":"Re: [PATCH v5 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2021-01-10T14:06:38Z","receivedAt":"2021-01-10T14:07:15Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Tue, Dec 29, 2020 at 11:35:56PM -0500, Derrick Stolee wrote:\n> On 12/28/2020 6:15 AM, Abhishek Kumar via GitGitGadget wrote:\n> > This patch series implements the corrected commit date offsets as generation\n> > number v2, along with other pre-requisites.\n> \n> Abhishek,\n> \n> Thank you for this version. I appreciate your hard work on this topic,\n> especially after GSoC ended and you returned to being a full-time student.\n> \n> My hope was that I could completely approve this series and only provide\n> forward-fixes from here on out, as necessary. I think there are a few minor\n> typos that you might want to address, but I was also able to understand your\n> intention.\n> \n> I did make a particular case about a SEGFAULT I hit that I have been unable\n> to replicate. I saw it both in my copy of torvalds/linux and of\n> chromium/chromium. I have the file for chromium/chromium that is in a bad\n> state where a GDAT value includes the bit saying it should be in the long\n> offsets chunk, but that chunk doesn't exist. Further, that chunk doesn't\n> exist in a from-scratch write.\n\nI hope validating mixed generation chain while writing as well was\nenough to fix the SEGFAULT.\n\n>\n> I'm now taking backups of my existing commit-graph files before any later\n> test, but it doesn't repro for my Git repository or any other repo I try on\n> purpose.\n> \n> However, I did some performance testing to double-check your numbers. I sent\n> a patch [1] that helps with some of the hard numbers.\n> \n> [1] https://lore.kernel.org/git/pull.828.git.1609302714183.gitgitgadget@gmail.com/\n> \n> The big question is whether the overhead from using a slab to store the\n> generation values is worth it. I still think it is, for these reasons:\n> \n> 1. Generation number v2 is measurably better than v1 in most user cases.\n> \n> 2. Generation number v2 is slower than using committer date due to the\n>    overhead, but _guarantees correctness_.\n> \n> I like to use \"git log --graph -<N>\" to compare against topological levels\n> (v1), for various levels of <N>. When <N> is small, we hope to minimize\n> the amount we need to walk using the extra commit-date information as an\n> assistance. Repos like git/git and torvalds/linux use the philosophy of\n> \"base your changes on oldest applicable commit\" enough that v1 struggles\n> sometimes.\n> \n> git/git: N=1000\n> \n> \tBenchmark #1: baseline\n> \tTime (mean ± σ):     100.3 ms ±   4.2 ms    [User: 89.0 ms, System: 11.3 ms]\n> \tRange (min … max):    94.5 ms … 105.1 ms    28 runs\n> \t\n> \tBenchmark #2: test\n> \tTime (mean ± σ):      35.8 ms ±   3.1 ms    [User: 29.6 ms, System: 6.2 ms]\n> \tRange (min … max):    29.8 ms …  40.6 ms    81 runs\n> \t\n> \tSummary\n> \t'test' ran\n> \t2.80 ± 0.27 times faster than 'baseline'\n> \n> This is a dramatic improvement! Using my topo-walk stats commit, I see that\n> v1 walks 58,805 commits as part of the in-degree walk while v2 only walks\n> 4,335 commits!\n> \n> torvalds/linux: N=1000 (starting at v5.10)\n> \n> \tBenchmark #1: baseline\n> \tTime (mean ± σ):      90.8 ms ±   3.7 ms    [User: 75.2 ms, System: 15.6 ms]\n> \tRange (min … max):    85.2 ms …  96.2 ms    31 runs\n> \t\n> \tBenchmark #2: test\n> \tTime (mean ± σ):      49.2 ms ±   3.5 ms    [User: 36.9 ms, System: 12.3 ms]\n> \tRange (min … max):    42.9 ms …  54.0 ms    61 runs\n> \t\n> \tSummary\n> \t'test' ran\n> \t1.85 ± 0.15 times faster than 'baseline'\n> \n> Similarly, v1 walked 38,161 commits compared to 4,340 by v2.\n> \n> If I increase N to something like 10,000, then usually these values get\n> washed out due to the width of the parallel topics.\n\nThat's not too bad, as large N would be needed rather infrequently.\n\n> \n> The place we were still using commit-date as a heuristic was paint_down_to_common\n> which caused a regression the first time we used v1, at least for certain cases.\n> \n> Specifically, computing the merge-base in torvalds/linux between v4.8 and v4.9\n> hit a strangeness about a pair of recent commits both based on a very old commit,\n> but the generation numbers forced walking farther than necessary. This doesn't\n> happen with v2, but we see the overhead cost of the slabs:\n> \n> \tBenchmark #1: baseline\n> \tTime (mean ± σ):     112.9 ms ±   2.8 ms    [User: 96.5 ms, System: 16.3 ms]\n> \tRange (min … max):   107.7 ms … 118.0 ms    26 runs\n> \t\n> \tBenchmark #2: test\n> \tTime (mean ± σ):     147.1 ms ±   5.2 ms    [User: 132.7 ms, System: 14.3 ms]\n> \tRange (min … max):   141.4 ms … 162.2 ms    18 runs\n> \t\n> \tSummary\n> \t'baseline' ran\n> \t1.30 ± 0.06 times faster than 'test'\n> \n> The overhead still exists for a more recent pair of versions (v5.0 and v5.1):\n> \n> \tBenchmark #1: baseline\n> \tTime (mean ± σ):      25.1 ms ±   3.2 ms    [User: 18.6 ms, System: 6.5 ms]\n> \tRange (min … max):    19.0 ms …  32.8 ms    99 runs\n> \t\n> \tBenchmark #2: test\n> \tTime (mean ± σ):      33.3 ms ±   3.3 ms    [User: 26.5 ms, System: 6.9 ms]\n> \tRange (min … max):    27.0 ms …  38.4 ms    105 runs\n> \t\n> \tSummary\n> \t'baseline' ran\n> \t1.33 ± 0.22 times faster than 'test'\n> \n> I still think this overhead is worth it. In case not everyone agrees, it _might_\n> be worth a command-line option to skip the GDAT chunk. That also prevents an\n> ability to eventually wean entirely of generation number v1 and allow the commit\n> date to take the full 64-bit column (instead of only 34 bits, saving 30 for\n> topo-levels).\n\nThank you for the detailed benchmarking and discussion. \n\nI don't think there is any disagreement on utility of corrected commit\ndates so far. \n\nWe will run out of 34-bits for the commit date by the year 2514, so I\nam not exactly worried about weaning of generation number v1 anytime\nsoon.\n\n> \n> Again, such a modification should not be considered required for this series.\n> \n> > ----------------------------------------------------------------------------\n> > \n> > Improvements left for a future series:\n> > \n> >  * Save commits with generation data overflow and extra edge commits instead\n> >    of looping over all commits. cf. 858sbel67n.fsf@gmail.com\n> >  * Verify both topological levels and corrected commit dates when present.\n> >    cf. 85pn4tnk8u.fsf@gmail.com\n> \n> These seem like reasonable things to delay for a later series\n> or for #leftoverbits\n> \n> Thanks,\n> -Stolee\n> \n\nThanks\n- Abhishek\n"},{"id":"414042","messageId":"003c1892-cbfe-7437-f8ce-fbae58f0cb83@gmail.com","threadId":"53933","inReplyTo":"X/r9i0HJFEGxuyW/@Abhishek-Arch","subject":"Re: [PATCH v5 09/11] commit-graph: use generation v2 only if entire chain does","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-01-11T12:43:31Z","receivedAt":"2021-01-11T12:44:14Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 1/10/2021 8:13 AM, Abhishek Kumar wrote:\n> On Tue, Dec 29, 2020 at 10:23:54PM -0500, Derrick Stolee wrote:\n>> Your tests in this patch seem very thorough, covering all the cases\n>> I could think to create this strange situation. I even tried creating\n>> cases where the overflow would be necessary. The following test actually\n>> fails on the \"graph_read_expect 6\" due to the extra chunk, not the 'write'\n>> process I was trying to trick into failure.\n>>\n>> diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n>> index 8e90f3423b..cfef8e52b9 100755\n>> --- a/t/t5324-split-commit-graph.sh\n>> +++ b/t/t5324-split-commit-graph.sh\n>> @@ -453,6 +453,20 @@ test_expect_success 'prevent regression for duplicate commits across layers' '\n>>         git -C dup commit-graph verify\n>>  '\n>>  \n>> +test_expect_success 'upgrade to generation data succeeds when there was none' '\n>> +\t(\n>> +\t\tcd dup &&\n>> +\t\trm -rf .git/objects/info/commit-graph* &&\n>> +\t\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph \\\n>> +\t\t\twrite --reachable &&\n>> +\t\tGIT_COMMITTER_DATE=\"1980-01-01 00:00\" git commit --allow-empty -m one &&\n>> +\t\tGIT_COMMITTER_DATE=\"2090-01-01 00:00\" git commit --allow-empty -m two &&\n>> +\t\tGIT_COMMITTER_DATE=\"2000-01-01 00:00\" git commit --allow-empty -m three &&\n>> +\t\tgit commit-graph write --reachable &&\n>> +\t\tgraph_read_expect 6\n>> +\t)\n>> +'\n> \n> I am not sure what this test adds over the existing generation data\n> overflow related tests added in t5318-commit-graph.sh\n\nGood point.\n\n-Stolee\n"},{"id":"414513","messageId":"05dcb8628186d8a71b84cc3cd2ce4877abee039a.1610820679.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v6 02/11] revision: parse parent in indegree_walk_step()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:09Z","receivedAt":"2021-01-16T18:12:37Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn indegree_walk_step(), we add unvisited parents to the indegree queue.\nHowever, parents are not guaranteed to be parsed. As the indegree queue\nsorts by generation number, let's parse parents before inserting them to\nensure the correct priority order.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n revision.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/revision.c b/revision.c\nindex 1bb590ece78..be2d828a4cc 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3397,6 +3397,9 @@ static void indegree_walk_step(struct rev_info *revs)\n \t\tstruct commit *parent = p->item;\n \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n \n+\t\tif (repo_parse_commit_gently(revs->repo, parent, 1) < 0)\n+\t\t\treturn;\n+\n \t\tif (*pi)\n \t\t\t(*pi)++;\n \t\telse\n-- \ngitgitgadget\n\n"},{"id":"414514","messageId":"4d8eb415578e67ab91ebbeefa158425da226ebbd.1610820679.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v6 01/11] commit-graph: fix regression when computing Bloom filters","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:08Z","receivedAt":"2021-01-16T18:12:37Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nBefore computing Bloom filters, the commit-graph machinery uses\ncommit_gen_cmp to sort commits by generation order for improved diff\nperformance. 3d11275505 (commit-graph: examine commits by generation\nnumber, 2020-03-30) claims that this sort can reduce the time spent to\ncompute Bloom filters by nearly half.\n\nBut since c49c82aa4c (commit: move members graph_pos, generation to a\nslab, 2020-06-17), this optimization is broken, since asking for a\n'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\nwhile writing.\n\nNot all hope is lost, though: 'commit_gen_cmp()' falls back to\ncomparing commits by their date when they have equal generation number,\nand so since c49c82aa4c is purely a date comparison function. This\nheuristic is good enough that we don't seem to loose appreciable\nperformance while computing Bloom filters.\n\nApplying this patch (compared with v2.30.0) speeds up computing Bloom\nfilters by factors ranging from 0.40% to 5.19% on various repositories [1].\n\nSo, avoid the useless 'commit_graph_generation()' while writing by\ninstead accessing the slab directly. This returns the newly-computed\ngeneration numbers, and allows us to avoid the heuristic by directly\ncomparing generation numbers.\n\n[1]: https://lore.kernel.org/git/20210105094535.GN8396@szeder.dev/\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 8 ++++++--\n 1 file changed, 6 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex e9124d4a412..0267886e76c 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -139,13 +139,17 @@ static struct commit_graph_data *commit_graph_data_at(const struct commit *c)\n \treturn data;\n }\n \n+/* \n+ * Should be used only while writing commit-graph as it compares\n+ * generation value of commits by directly accessing commit-slab.\n+ */\n static int commit_gen_cmp(const void *va, const void *vb)\n {\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n+\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n+\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"414515","messageId":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v5.git.1609154168.gitgitgadget@gmail.com","subject":"[PATCH v6 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:07Z","receivedAt":"2021-01-16T18:12:37Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"This patch series implements the corrected commit date offsets as generation\nnumber v2, along with other pre-requisites.\n\nGit uses topological levels in the commit-graph file for commit-graph\ntraversal operations like 'git log --graph'. Unfortunately, using\ntopological levels can result in a worse performance than without them when\ncompared with committer date as a heuristics. For example, 'git merge-base\nv4.8 v4.9' on the Linux repository walks 635,579 commits using topological\nlevels and walks 167,468 using committer date. Since 091f4cf3 (commit: don't\nuse generation numbers if not needed, 2018-08-30), 'git merge-base' uses\ncommitter date heuristic unless there is a cutoff because of the performance\nhit.\n\nThus, the need for generation number v2 was born. New generation number\nneeded to provide good performance, increment updates, and backward\ncompatibility. Due to an unfortunate problem [1], we also needed a way to\ndistinguish between the old and new generation number without incrementing\ngraph version.\n\n[1] https://public-inbox.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nVarious candidates were examined (https://github.com/derrickstolee/gen-test,\nhttps://github.com/abhishekkumar2718/git/pull/1). The proposed generation\nnumber v2, Corrected Commit Date with Mononotically Increasing Offsets\nperformed much worse than committer date (506,577 vs. 167,468 commits walked\nfor 'git merge-base v4.8 v4.9') and was dropped.\n\nUsing Generation Data chunk (GDAT) relieves the requirement of backward\ncompatibility as we would continue to store topological levels in Commit\nData (CDAT) chunk. Thus, Corrected Commit Date was chosen as generation\nnumber v2. The Corrected Commit Date is defined as follows:\n\nFor a commit C, let its corrected commit date be the maximum of the commit\ndate of C and the corrected commit dates of its parents plus 1. Then\ncorrected commit date offset is the difference between corrected commit date\nof C and commit date of C. As a special case, a root commit with the\ntimestamp zero has corrected commit date of 1 to be able to distinguish it\nfrom GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit date).\n\nWe will introduce an additional commit-graph chunk, Generation DATa (GDAT)\nchunk, and store corrected commit date offsets in GDAT chunk while storing\ntopological levels in CDAT chunk. The old versions of Git would ignore GDAT\nchunk, using topological levels from CDAT chunk. In contrast, new versions\nof Git would use corrected commit dates, falling back to topological level\nif the generation data chunk is absent in the commit-graph file.\n\nWhile storing corrected commit date offsets saves us 4 bytes per commit (as\ncompared with storing corrected commit dates directly), it's however\npossible for the offset to overflow the space allocated. To handle such\ncases, we introduce a new chunk, Generation Data Overflow (GDOV) that stores\nthe corrected commit date. For overflowing offsets, we set MSB and store the\nposition into the GDOV chunk, in a mechanism similar to the Extra Edges list\nchunk.\n\nFor mixed generation number environment (for example new Git on the command\nline, old Git used by GUI client), we can encounter a mixed-chain\ncommit-graph (a commit-graph chain where some of split commit-graph files\nhave GDAT chunk and others do not). As backward compatibility is one of the\ngoals, we can define the following behavior:\n\nWhile reading a mixed-chain commit-graph version, we fall back on\ntopological levels as corrected commit dates and topological levels cannot\nbe compared directly.\n\nWhen adding new layer to the split commit-graph file, and when merging some\nor all layers (replacing them in the latter case), the new layer will have\nGDAT chunk if and only if in the final result there would be no layer\nwithout GDAT chunk just below it.\n\nThanks to Dr. Stolee, Dr. Narębski, and Taylor for their reviews.\n\nI look forward to everyone's reviews!\n\nThanks\n\n * Abhishek\n\n----------------------------------------------------------------------------\n\nImprovements left for a future series:\n\n * Save commits with generation data overflow and extra edge commits instead\n   of looping over all commits. cf. 858sbel67n.fsf@gmail.com\n * Verify both topological levels and corrected commit dates when present.\n   cf. 85pn4tnk8u.fsf@gmail.com\n\nChanges in version 6:\n\n * Fixed typos in commit message for \"commit-graph: implement corrected\n   commit date\".\n * Removed an unnecessary else-block in \"commit-graph: implement corrected\n   commit date\".\n * Validate mixed generation chain correctly while writing in \"commit-graph:\n   use generation v2 only if the entire chain does\".\n * Die if the GDAT chunk indicates data has overflown but there are is no\n   generation data overflow chunk.\n\nChanges in version 5:\n\n * Explained a possible reason for no change in performance for\n   \"commit-graph: fix regression when computing bloom-filters\"\n * Clarified about the addition of a new test for 11-digit octal\n   implementations of ustar.\n * Fixed duplicate test names in \"commit-graph: consolidate\n   fill_commit_graph_info\".\n * Swapped the order \"commit-graph: return 64-bit generation number\",\n   \"commit-graph: add a slab to store topological levels\" to minimize lines\n   changed.\n * Fixed the mismerge in \"commit-graph: return 64-bit generation number\"\n * Clarified the preparatory steps are for the larger goal of implementing\n   generation number v2 in \"commit-graph: return 64-bit generation number\".\n * Moved the rename of \"run_three_modes()\" to \"run_all_modes()\" into a new\n   patch \"t6600-test-reach: generalize *_three_modes\".\n * Explained and removed the checks for GENERATION_NUMBER_INFINITY that can\n   never be true in \"commit-graph: add a slab to store topological levels\".\n * Fixed incorrect logic for verifying commit-graph in \"commit-graph:\n   implement corrected commit date\".\n * Added minor improvements to commit message of \"commit-graph: implement\n   generation data chunk\".\n * Added '--date ' option to test_commit() in 'test-lib-functions.sh' in\n   \"commit-graph: implement generation data chunk\".\n * Improved coding style (also in tests) for \"commit-graph: use generation\n   v2 only if entire chain does\".\n * Simplified test repository structure in \"commit-graph: use generation v2\n   only if entire chain does\" as only the number of commits in a split\n   commit-graph layer are relevant.\n * Added a new test in \"commit-graph: use generation v2 only if entire chain\n   does\" to check if the layers are merged correctly.\n * Explicitly mentioned commit \"091f4cf3\" in the commit-message of\n   \"commit-graph: use corrected commit dates in paint_down_to_common()\".\n * Minor corrections to documentation in \"doc: add corrected commit date\n   info\".\n * Minor corrections to coding style.\n\nChanges in version 4:\n\n * Added GDOV to handle overflows in generation data.\n * Added a test for writing tip graph for a generation number v2 graph chain\n   in t5324-split-commit-graph.sh\n * Added a section on how mixed generation number chains are handled in\n   Documentation/technical/commit-graph-format.txt\n * Reverted unimportant whitespace, style changes in commit-graph.c\n * Added header comments about the order of comparision for\n   compare_commits_by_gen_then_commit_date in commit.h,\n   compare_commits_by_gen in commit-graph.h\n * Elaborated on why t6404 fails with corrected commit date and must be run\n   with GIT_TEST_COMMIT_GRAPH=1in the commit \"commit-reach: use corrected\n   commit dates in paint_down_to_common()\"\n * Elaborated on write behavior for mixed generation number chains in the\n   commit \"commit-graph: use generation v2 only if entire chain does\"\n * Added notes about adding the topo_level slab to struct\n   write_commit_graph_context as well as struct commit_graph.\n * Clarified commit message for \"commit-graph: consolidate\n   fill_commit_graph_info\"\n * Removed the claim \"GDAT can store future generation numbers\" because it\n   hasn't been tested yet.\n\nChanges in version 3:\n\n * Reordered patches as discussed in 2\n   [https://lore.kernel.org/git/aee0ae56-3395-6848-d573-27a318d72755@gmail.com/].\n * Split \"implement corrected commit date\" into two patches - one\n   introducing the topo level slab and other implementing corrected commit\n   dates.\n * Extended split-commit-graph tests to verify at the end of test.\n * Use topological levels as generation number if any of split commit-graph\n   files do not have generation data chunk.\n\nChanges in version 2:\n\n * Add tests for generation data chunk.\n * Add an option GIT_TEST_COMMIT_GRAPH_NO_GDAT to control whether to write\n   generation data chunk.\n * Compare commits with corrected commit dates if present in\n   paint_down_to_common().\n * Update technical documentation.\n * Handle mixed generation commit chains.\n * Improve commit messages for \"commit-graph: fix regression when computing\n   bloom filter\", \"commit-graph: consolidate fill_commit_graph_info\",\n * Revert unnecessary whitespace changes.\n * Split uint_32 -> timestamp_t change into a new commit.\n\nAbhishek Kumar (11):\n  commit-graph: fix regression when computing Bloom filters\n  revision: parse parent in indegree_walk_step()\n  commit-graph: consolidate fill_commit_graph_info\n  t6600-test-reach: generalize *_three_modes\n  commit-graph: add a slab to store topological levels\n  commit-graph: return 64-bit generation number\n  commit-graph: implement corrected commit date\n  commit-graph: implement generation data chunk\n  commit-graph: use generation v2 only if entire chain does\n  commit-reach: use corrected commit dates in paint_down_to_common()\n  doc: add corrected commit date info\n\n .../technical/commit-graph-format.txt         |  28 +-\n Documentation/technical/commit-graph.txt      |  77 +++++-\n commit-graph.c                                | 251 ++++++++++++++----\n commit-graph.h                                |  15 +-\n commit-reach.c                                |  38 +--\n commit-reach.h                                |   2 +-\n commit.c                                      |   4 +-\n commit.h                                      |   5 +-\n revision.c                                    |  13 +-\n t/README                                      |   3 +\n t/helper/test-read-graph.c                    |   4 +\n t/t4216-log-bloom.sh                          |   4 +-\n t/t5000-tar-tree.sh                           |  24 +-\n t/t5318-commit-graph.sh                       |  79 +++++-\n t/t5324-split-commit-graph.sh                 | 193 +++++++++++++-\n t/t6404-recursive-merge.sh                    |   5 +-\n t/t6600-test-reach.sh                         |  68 ++---\n t/test-lib-functions.sh                       |   6 +\n upload-pack.c                                 |   2 +-\n 19 files changed, 667 insertions(+), 154 deletions(-)\n\n\nbase-commit: 4151fdb1c76c1a190ac9241b67223efd19f3e478\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-676%2Fabhishekkumar2718%2Fcorrected_commit_date-v6\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-676/abhishekkumar2718/corrected_commit_date-v6\nPull-Request: https://github.com/gitgitgadget/git/pull/676\n\nRange-diff vs v5:\n\n  1:  c4e817abf7d !  1:  4d8eb415578 commit-graph: fix regression when computing Bloom filters\n     @@ Metadata\n       ## Commit message ##\n          commit-graph: fix regression when computing Bloom filters\n      \n     -    Before computing Bloom fitlers, the commit-graph machinery uses\n     +    Before computing Bloom filters, the commit-graph machinery uses\n          commit_gen_cmp to sort commits by generation order for improved diff\n          performance. 3d11275505 (commit-graph: examine commits by generation\n          number, 2020-03-30) claims that this sort can reduce the time spent to\n     @@ Commit message\n          'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\n          while writing.\n      \n     -    Not all hope is lost, though: 'commit_graph_generation()' falls back to\n     +    Not all hope is lost, though: 'commit_gen_cmp()' falls back to\n          comparing commits by their date when they have equal generation number,\n     -    and so since c49c82aa4c is purely a date comparision function. This\n     +    and so since c49c82aa4c is purely a date comparison function. This\n          heuristic is good enough that we don't seem to loose appreciable\n     -    performance while computing Bloom filters. Applying this patch (compared\n     -    with v2.29.1) speeds up computing Bloom filters by around ~4\n     -    seconds.\n     +    performance while computing Bloom filters.\n     +\n     +    Applying this patch (compared with v2.30.0) speeds up computing Bloom\n     +    filters by factors ranging from 0.40% to 5.19% on various repositories [1].\n      \n          So, avoid the useless 'commit_graph_generation()' while writing by\n          instead accessing the slab directly. This returns the newly-computed\n          generation numbers, and allows us to avoid the heuristic by directly\n          comparing generation numbers.\n      \n     +    [1]: https://lore.kernel.org/git/20210105094535.GN8396@szeder.dev/\n     +\n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n       ## commit-graph.c ##\n     -@@ commit-graph.c: static int commit_gen_cmp(const void *va, const void *vb)\n     +@@ commit-graph.c: static struct commit_graph_data *commit_graph_data_at(const struct commit *c)\n     + \treturn data;\n     + }\n     + \n     ++/* \n     ++ * Should be used only while writing commit-graph as it compares\n     ++ * generation value of commits by directly accessing commit-slab.\n     ++ */\n     + static int commit_gen_cmp(const void *va, const void *vb)\n     + {\n       \tconst struct commit *a = *(const struct commit **)va;\n       \tconst struct commit *b = *(const struct commit **)vb;\n       \n  2:  7645e0bcef0 =  2:  05dcb862818 revision: parse parent in indegree_walk_step()\n  3:  ca646912b2b =  3:  dcb9891d819 commit-graph: consolidate fill_commit_graph_info\n  4:  591935075f1 =  4:  4fbdee7ac90 t6600-test-reach: generalize *_three_modes\n  5:  baae7006764 =  5:  fbd8feb5d8c commit-graph: add a slab to store topological levels\n  6:  26bd6f49100 =  6:  855ff662a44 commit-graph: return 64-bit generation number\n  7:  859c39eff52 !  7:  8fbe7486405 commit-graph: implement corrected commit date\n     @@ Commit message\n          of GDAT chunk, which is a reduction of around 6% in the size of\n          commit-graph file.\n      \n     -    However, using offsets be problematic if one of commits is malformed but\n     -    valid and has committerdate of 0 Unix time, as the offset would be the\n     -    same as corrected commit date and thus require 64-bits to be stored\n     -    properly.\n     +    However, using offsets be problematic if a commit is malformed but valid\n     +    and has committer date of 0 Unix time, as the offset would be the same\n     +    as corrected commit date and thus require 64-bits to be stored properly.\n      \n          While Git does not write out offsets at this stage, Git stores the\n          corrected commit dates in member generation of struct commit_graph_data.\n     @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph\n       \t\t\t\t\tbreak;\n      -\t\t\t\t} else if (level > max_level) {\n      -\t\t\t\t\tmax_level = level;\n     -+\t\t\t\t} else {\n     -+\t\t\t\t\tif (level > max_level)\n     -+\t\t\t\t\t\tmax_level = level;\n     -+\n     -+\t\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n     -+\t\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n       \t\t\t\t}\n     ++\n     ++\t\t\t\tif (level > max_level)\n     ++\t\t\t\t\tmax_level = level;\n     ++\n     ++\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n     ++\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n       \t\t\t}\n       \n     + \t\t\tif (all_parents_computed) {\n      @@ commit-graph.c: static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n       \t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n       \t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n  8:  8403c4d0257 !  8:  6d0696ae216 commit-graph: implement generation data chunk\n     @@ commit-graph.c: static void fill_commit_graph_info(struct commit *item, struct c\n      +\t\toffset = (timestamp_t)get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n      +\n      +\t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n     ++\t\t\tif (!g->chunk_generation_data_overflow)\n     ++\t\t\t\tdie(_(\"commit-graph requires overflow generation data but has none\"));\n     ++\n      +\t\t\toffset_pos = offset ^ CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW;\n      +\t\t\tgraph_data->generation = get_be64(g->chunk_generation_data_overflow + 8 * offset_pos);\n      +\t\t} else\n  9:  a3a70a1edd0 !  9:  fba0d7f3dfe commit-graph: use generation v2 only if entire chain does\n     @@ commit-graph.c: static void split_graph_merge_strategy(struct write_commit_graph\n       \t\tg = g->base_graph;\n       \t}\n      @@ commit-graph.c: int write_commit_graph(struct object_directory *odb,\n     - \t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n     + \t} else\n     + \t\tctx->num_commit_graphs_after = 1;\n       \n     - \t\twhile (g) {\n     -+\t\t\tg->read_generation_data = 1;\n     - \t\t\tg->topo_levels = &topo_levels;\n     - \t\t\tg = g->base_graph;\n     - \t\t}\n     ++\tvalidate_mixed_generation_chain(ctx->r->objects->commit_graph);\n     ++\n     + \tcompute_generation_numbers(ctx);\n     + \n     + \tif (ctx->changed_paths)\n      @@ commit-graph.c: int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n       \t\t * also GENERATION_NUMBER_V1_MAX. Decrement to avoid extra logic\n       \t\t * in the following condition.\n 10:  093101f908b = 10:  ba1f2c5555f commit-reach: use corrected commit dates in paint_down_to_common()\n 11:  20299e57457 = 11:  e571f03d8bd doc: add corrected commit date info\n\n-- \ngitgitgadget\n"},{"id":"414517","messageId":"dcb9891d819d9d848d26c93de8b5ed80f912535f.1610820679.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v6 03/11] commit-graph: consolidate fill_commit_graph_info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:10Z","receivedAt":"2021-01-16T18:12:37Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nBoth fill_commit_graph_info() and fill_commit_in_graph() parse\ninformation present in commit data chunk. Let's simplify the\nimplementation by calling fill_commit_graph_info() within\nfill_commit_in_graph().\n\nfill_commit_graph_info() used to not load committer data from commit data\nchunk. However, with the upcoming switch to using corrected committer\ndate as generation number v2, we will have to load committer date to\ncompute generation number value anyway.\n\ne51217e15 (t5000: test tar files that overflow ustar headers,\n30-06-2016) introduced a test 'generate tar with future mtime' that\ncreates a commit with committer date of (2^36 + 1) seconds since\nEPOCH. The CDAT chunk provides 34-bits for storing committer date, thus\ncommitter time overflows into generation number (within CDAT chunk) and\nhas undefined behavior.\n\nThe test used to pass as fill_commit_graph_info() would not set struct\nmember `date` of struct commit and load committer date from the object\ndatabase, generating a tar file with the expected mtime.\n\nHowever, with corrected commit date, we will load the committer date\nfrom CDAT chunk (truncated to lower 34-bits to populate the generation\nnumber. Thus, Git sets date and generates tar file with the truncated\nmtime.\n\nThe ustar format (the header format used by most modern tar programs)\nonly has room for 11 (or 12, depending on some implementations) octal\ndigits for the size and mtime of each file.\n\nAs the CDAT chunk is overflow by 12-octal digits but not 11-octal\ndigits, we split the existing tests to test both implementations\nseparately and add a new explicit test for 11-digit implementation.\n\nTo test the 11-octal digit implementation, we create a future commit\nwith committer date of 2^34 - 1, which overflows 11-octal digits without\noverflowing 34-bits of the Commit Date chunks.\n\nTo test the 12-octal digit implementation, the smallest committer date\npossible is 2^36 + 1, which overflows the CDAT chunk and thus\ncommit-graph must be disabled for the test.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c      | 27 ++++++++++-----------------\n t/t5000-tar-tree.sh | 24 +++++++++++++++++++++---\n 2 files changed, 31 insertions(+), 20 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 0267886e76c..3d59b8b905d 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -753,15 +753,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n \tuint32_t lex_index;\n+\tuint64_t date_high, date_low;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n+\tif (pos >= g->num_commits + g->num_commits_in_base)\n+\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\n \tlex_index = pos - g->num_commits_in_base;\n \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n \n \tgraph_data = commit_graph_data_at(item);\n \tgraph_data->graph_pos = pos;\n+\n+\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n+\tdate_low = get_be32(commit_data + g->hash_len + 12);\n+\titem->date = (timestamp_t)((date_high << 32) | date_low);\n+\n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n@@ -776,38 +785,22 @@ static int fill_commit_in_graph(struct repository *r,\n {\n \tuint32_t edge_value;\n \tuint32_t *parent_data_ptr;\n-\tuint64_t date_low, date_high;\n \tstruct commit_list **pptr;\n-\tstruct commit_graph_data *graph_data;\n \tconst unsigned char *commit_data;\n \tuint32_t lex_index;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n-\tif (pos >= g->num_commits + g->num_commits_in_base)\n-\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\tfill_commit_graph_info(item, g, pos);\n \n-\t/*\n-\t * Store the \"full\" position, but then use the\n-\t * \"local\" position for the rest of the calculation.\n-\t */\n-\tgraph_data = commit_graph_data_at(item);\n-\tgraph_data->graph_pos = pos;\n \tlex_index = pos - g->num_commits_in_base;\n-\n \tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n \n \titem->object.parsed = 1;\n \n \tset_commit_tree(item, NULL);\n \n-\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n-\tdate_low = get_be32(commit_data + g->hash_len + 12);\n-\titem->date = (timestamp_t)((date_high << 32) | date_low);\n-\n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n-\n \tpptr = &item->parents;\n \n \tedge_value = get_be32(commit_data + g->hash_len);\ndiff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\nindex 3ebb0d3b652..7204799a0b5 100755\n--- a/t/t5000-tar-tree.sh\n+++ b/t/t5000-tar-tree.sh\n@@ -431,15 +431,33 @@ test_expect_success TAR_HUGE,LONG_IS_64BIT 'system tar can read our huge size' '\n \ttest_cmp expect actual\n '\n \n-test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n+test_expect_success TIME_IS_64BIT 'set up repository with far-future (2^34 - 1) commit' '\n+\trm -f .git/index &&\n+\techo foo >file &&\n+\tgit add file &&\n+\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n+\t\tgit commit -m \"tempori parendum\"\n+'\n+\n+test_expect_success TIME_IS_64BIT 'generate tar with far-future mtime' '\n+\tgit archive HEAD >future.tar\n+'\n+\n+test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n+\techo 2514 >expect &&\n+\ttar_info future.tar | cut -d\" \" -f2 >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success TIME_IS_64BIT 'set up repository with far-far-future (2^36 + 1) commit' '\n \trm -f .git/index &&\n \techo content >file &&\n \tgit add file &&\n-\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n+\tGIT_TEST_COMMIT_GRAPH=0 GIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n \t\tgit commit -m \"tempori parendum\"\n '\n \n-test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n+test_expect_success TIME_IS_64BIT 'generate tar with far-far-future mtime' '\n \tgit archive HEAD >future.tar\n '\n \n-- \ngitgitgadget\n\n"},{"id":"414516","messageId":"4fbdee7ac9070345ad20d226a260b55767e375d1.1610820679.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v6 04/11] t6600-test-reach: generalize *_three_modes","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:11Z","receivedAt":"2021-01-16T18:12:38Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a preparatory step to implement generation number v2, we add tests to\nensure Git can read and parse commit-graph files without Generation Data\nchunk. These files represent commit-graph files written by Old Git and\nare neccesary for backward compatability.\n\nWe extend run_three_modes() and test_three_modes() to *_all_modes() with\nthe fourth mode being \"commit-graph without generation data chunk\".\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n t/t6600-test-reach.sh | 62 +++++++++++++++++++++----------------------\n 1 file changed, 31 insertions(+), 31 deletions(-)\n\ndiff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\nindex f807276337d..af10f0dc090 100755\n--- a/t/t6600-test-reach.sh\n+++ b/t/t6600-test-reach.sh\n@@ -58,7 +58,7 @@ test_expect_success 'setup' '\n \tgit config core.commitGraph true\n '\n \n-run_three_modes () {\n+run_all_modes () {\n \ttest_when_finished rm -rf .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n \ttest_cmp expect actual &&\n@@ -70,8 +70,8 @@ run_three_modes () {\n \ttest_cmp expect actual\n }\n \n-test_three_modes () {\n-\trun_three_modes test-tool reach \"$@\"\n+test_all_modes () {\n+\trun_all_modes test-tool reach \"$@\"\n }\n \n test_expect_success 'ref_newer:miss' '\n@@ -80,7 +80,7 @@ test_expect_success 'ref_newer:miss' '\n \tB:commit-4-9\n \tEOF\n \techo \"ref_newer(A,B):0\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'ref_newer:hit' '\n@@ -89,7 +89,7 @@ test_expect_success 'ref_newer:hit' '\n \tB:commit-2-3\n \tEOF\n \techo \"ref_newer(A,B):1\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'in_merge_bases:hit' '\n@@ -98,7 +98,7 @@ test_expect_success 'in_merge_bases:hit' '\n \tB:commit-8-8\n \tEOF\n \techo \"in_merge_bases(A,B):1\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'in_merge_bases:miss' '\n@@ -107,7 +107,7 @@ test_expect_success 'in_merge_bases:miss' '\n \tB:commit-5-9\n \tEOF\n \techo \"in_merge_bases(A,B):0\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'in_merge_bases_many:hit' '\n@@ -117,7 +117,7 @@ test_expect_success 'in_merge_bases_many:hit' '\n \tX:commit-5-7\n \tEOF\n \techo \"in_merge_bases_many(A,X):1\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'in_merge_bases_many:miss' '\n@@ -127,7 +127,7 @@ test_expect_success 'in_merge_bases_many:miss' '\n \tX:commit-8-6\n \tEOF\n \techo \"in_merge_bases_many(A,X):0\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'in_merge_bases_many:miss-heuristic' '\n@@ -137,7 +137,7 @@ test_expect_success 'in_merge_bases_many:miss-heuristic' '\n \tX:commit-6-6\n \tEOF\n \techo \"in_merge_bases_many(A,X):0\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'is_descendant_of:hit' '\n@@ -148,7 +148,7 @@ test_expect_success 'is_descendant_of:hit' '\n \tX:commit-1-1\n \tEOF\n \techo \"is_descendant_of(A,X):1\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'is_descendant_of:miss' '\n@@ -159,7 +159,7 @@ test_expect_success 'is_descendant_of:miss' '\n \tX:commit-7-6\n \tEOF\n \techo \"is_descendant_of(A,X):0\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'get_merge_bases_many' '\n@@ -174,7 +174,7 @@ test_expect_success 'get_merge_bases_many' '\n \t\tgit rev-parse commit-5-6 \\\n \t\t\t      commit-4-7 | sort\n \t} >expect &&\n-\ttest_three_modes get_merge_bases_many\n+\ttest_all_modes get_merge_bases_many\n '\n \n test_expect_success 'reduce_heads' '\n@@ -196,7 +196,7 @@ test_expect_success 'reduce_heads' '\n \t\t\t      commit-2-8 \\\n \t\t\t      commit-1-10 | sort\n \t} >expect &&\n-\ttest_three_modes reduce_heads\n+\ttest_all_modes reduce_heads\n '\n \n test_expect_success 'can_all_from_reach:hit' '\n@@ -219,7 +219,7 @@ test_expect_success 'can_all_from_reach:hit' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach(X,Y):1\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach:miss' '\n@@ -241,7 +241,7 @@ test_expect_success 'can_all_from_reach:miss' '\n \tY:commit-8-5\n \tEOF\n \techo \"can_all_from_reach(X,Y):0\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach_with_flag: tags case' '\n@@ -264,7 +264,7 @@ test_expect_success 'can_all_from_reach_with_flag: tags case' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach_with_flag(X,_,_,0,0):1\" >expect &&\n-\ttest_three_modes can_all_from_reach_with_flag\n+\ttest_all_modes can_all_from_reach_with_flag\n '\n \n test_expect_success 'commit_contains:hit' '\n@@ -280,8 +280,8 @@ test_expect_success 'commit_contains:hit' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):1\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'commit_contains:miss' '\n@@ -297,8 +297,8 @@ test_expect_success 'commit_contains:miss' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):0\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'rev-list: basic topo-order' '\n@@ -310,7 +310,7 @@ test_expect_success 'rev-list: basic topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 commit-3-2 commit-2-2 commit-1-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent topo-order' '\n@@ -322,7 +322,7 @@ test_expect_success 'rev-list: first-parent topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -334,7 +334,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -346,7 +346,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent range topo-order' '\n@@ -358,7 +358,7 @@ test_expect_success 'rev-list: first-parent range topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: ancestry-path topo-order' '\n@@ -368,7 +368,7 @@ test_expect_success 'rev-list: ancestry-path topo-order' '\n \t\tcommit-6-4 commit-5-4 commit-4-4 commit-3-4 \\\n \t\tcommit-6-3 commit-5-3 commit-4-3 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: symmetric difference topo-order' '\n@@ -382,7 +382,7 @@ test_expect_success 'rev-list: symmetric difference topo-order' '\n \t\tcommit-3-8 commit-2-8 commit-1-8 \\\n \t\tcommit-3-7 commit-2-7 commit-1-7 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8...commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8...commit-6-6\n '\n \n test_expect_success 'get_reachable_subset:all' '\n@@ -402,7 +402,7 @@ test_expect_success 'get_reachable_subset:all' '\n \t\t\t      commit-1-7 \\\n \t\t\t      commit-5-6 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:some' '\n@@ -420,7 +420,7 @@ test_expect_success 'get_reachable_subset:some' '\n \t\tgit rev-parse commit-3-3 \\\n \t\t\t      commit-1-7 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:none' '\n@@ -434,7 +434,7 @@ test_expect_success 'get_reachable_subset:none' '\n \tY:commit-2-8\n \tEOF\n \techo \"get_reachable_subset(X,Y)\" >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_done\n-- \ngitgitgadget\n\n"},{"id":"414518","messageId":"8fbe74864059ddde1c4107c0986b99518183de40.1610820679.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v6 07/11] commit-graph: implement corrected commit date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:14Z","receivedAt":"2021-01-16T18:13:03Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith most of preparations done, let's implement corrected commit date.\n\nThe corrected commit date for a commit is defined as:\n\n* A commit with no parents (a root commit) has corrected commit date\n  equal to its committer date.\n* A commit with at least one parent has corrected commit date equal to\n  the maximum of its commit date and one more than the largest corrected\n  commit date among its parents.\n\nAs a special case, a root commit with timestamp of zero (01.01.1970\n00:00:00Z) has corrected commit date of one, to be able to distinguish\nfrom GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit\ndate).\n\nTo minimize the space required to store corrected commit date, Git\nstores corrected commit date offsets into the commit-graph file. The\ncorrected commit date offset for a commit is defined as the difference\nbetween its corrected commit date and actual commit date.\n\nStoring corrected commit date requires sizeof(timestamp_t) bytes, which\nin most cases is 64 bits (uintmax_t). However, corrected commit date\noffsets can be safely stored using only 32-bits. This halves the size\nof GDAT chunk, which is a reduction of around 6% in the size of\ncommit-graph file.\n\nHowever, using offsets be problematic if a commit is malformed but valid\nand has committer date of 0 Unix time, as the offset would be the same\nas corrected commit date and thus require 64-bits to be stored properly.\n\nWhile Git does not write out offsets at this stage, Git stores the\ncorrected commit dates in member generation of struct commit_graph_data.\nIt will begin writing commit date offsets with the introduction of\ngeneration data chunk.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 21 +++++++++++++++++----\n 1 file changed, 17 insertions(+), 4 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 6d42e30cd9a..a899f429093 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1343,9 +1343,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n+\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (level != GENERATION_NUMBER_ZERO)\n+\t\tif (level != GENERATION_NUMBER_ZERO &&\n+\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1354,17 +1356,24 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n \t\t\tuint32_t max_level = 0;\n+\t\t\ttimestamp_t max_corrected_commit_date = 0;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n+\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n \n-\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n+\t\t\t\tif (level == GENERATION_NUMBER_ZERO ||\n+\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (level > max_level) {\n-\t\t\t\t\tmax_level = level;\n \t\t\t\t}\n+\n+\t\t\t\tif (level > max_level)\n+\t\t\t\t\tmax_level = level;\n+\n+\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n+\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n \t\t\t}\n \n \t\t\tif (all_parents_computed) {\n@@ -1373,6 +1382,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n \t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n+\n+\t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n+\t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n+\t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n \t\t\t}\n \t\t}\n \t}\n-- \ngitgitgadget\n\n"},{"id":"414519","messageId":"855ff662a445154d71cdecde5b4ee3fec510c700.1610820679.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v6 06/11] commit-graph: return 64-bit generation number","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:13Z","receivedAt":"2021-01-16T18:13:03Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a preparatory step for introducing corrected commit dates, let's\nreturn timestamp_t values from commit_graph_generation(), use\ntimestamp_t for local variables and define GENERATION_NUMBER_INFINITY\nas (2 ^ 63 - 1) instead.\n\nWe rename GENERATION_NUMBER_MAX to GENERATION_NUMBER_V1_MAX to\nrepresent the largest topological level we can store in the commit data\nchunk.\n\nWith corrected commit dates implemented, we will have two such *_MAX\nvariables to denote the largest offset and largest topological level\nthat can be stored.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 22 +++++++++++-----------\n commit-graph.h |  4 ++--\n commit-reach.c | 36 ++++++++++++++++++------------------\n commit-reach.h |  2 +-\n commit.c       |  4 ++--\n commit.h       |  4 ++--\n revision.c     | 10 +++++-----\n upload-pack.c  |  2 +-\n 8 files changed, 42 insertions(+), 42 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 3b69c3cc329..6d42e30cd9a 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -101,7 +101,7 @@ uint32_t commit_graph_position(const struct commit *c)\n \treturn data ? data->graph_pos : COMMIT_NOT_FROM_GRAPH;\n }\n \n-uint32_t commit_graph_generation(const struct commit *c)\n+timestamp_t commit_graph_generation(const struct commit *c)\n {\n \tstruct commit_graph_data *data =\n \t\tcommit_graph_data_slab_peek(&commit_graph_data_slab, c);\n@@ -150,8 +150,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n-\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n+\tconst timestamp_t generation_a = commit_graph_data_at(a)->generation;\n+\tconst timestamp_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n@@ -1370,8 +1370,8 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tif (all_parents_computed) {\n \t\t\t\tpop_commit(&list);\n \n-\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n-\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n+\t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n+\t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n \t\t\t}\n \t\t}\n@@ -2367,8 +2367,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \tfor (i = 0; i < g->num_commits; i++) {\n \t\tstruct commit *graph_commit, *odb_commit;\n \t\tstruct commit_list *graph_parents, *odb_parents;\n-\t\tuint32_t max_generation = 0;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t max_generation = 0;\n+\t\ttimestamp_t generation;\n \n \t\tdisplay_progress(progress, i + 1);\n \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n@@ -2432,16 +2432,16 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t\tcontinue;\n \n \t\t/*\n-\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n-\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n+\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n+\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n \t\t * extra logic in the following condition.\n \t\t */\n-\t\tif (max_generation == GENERATION_NUMBER_MAX)\n+\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n \t\t\tmax_generation--;\n \n \t\tgeneration = commit_graph_generation(graph_commit);\n \t\tif (generation != max_generation + 1)\n-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" != %\"PRItime),\n \t\t\t\t     oid_to_hex(&cur_oid),\n \t\t\t\t     generation,\n \t\t\t\t     max_generation + 1);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 00f00745b79..2e9aa7824ee 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -145,12 +145,12 @@ void disable_commit_graph(struct repository *r);\n \n struct commit_graph_data {\n \tuint32_t graph_pos;\n-\tuint32_t generation;\n+\ttimestamp_t generation;\n };\n \n /*\n  * Commits should be parsed before accessing generation, graph positions.\n  */\n-uint32_t commit_graph_generation(const struct commit *);\n+timestamp_t commit_graph_generation(const struct commit *);\n uint32_t commit_graph_position(const struct commit *);\n #endif\ndiff --git a/commit-reach.c b/commit-reach.c\nindex 50175b159e7..9b24b0378d5 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -32,12 +32,12 @@ static int queue_has_nonstale(struct prio_queue *queue)\n static struct commit_list *paint_down_to_common(struct repository *r,\n \t\t\t\t\t\tstruct commit *one, int n,\n \t\t\t\t\t\tstruct commit **twos,\n-\t\t\t\t\t\tint min_generation)\n+\t\t\t\t\t\ttimestamp_t min_generation)\n {\n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \tstruct commit_list *result = NULL;\n \tint i;\n-\tuint32_t last_gen = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n \tif (!min_generation)\n \t\tqueue.compare = compare_commits_by_commit_date;\n@@ -58,10 +58,10 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \t\tstruct commit *commit = prio_queue_get(&queue);\n \t\tstruct commit_list *parents;\n \t\tint flags;\n-\t\tuint32_t generation = commit_graph_generation(commit);\n+\t\ttimestamp_t generation = commit_graph_generation(commit);\n \n \t\tif (min_generation && generation > last_gen)\n-\t\t\tBUG(\"bad generation skip %8x > %8x at %s\",\n+\t\t\tBUG(\"bad generation skip %\"PRItime\" > %\"PRItime\" at %s\",\n \t\t\t    generation, last_gen,\n \t\t\t    oid_to_hex(&commit->object.oid));\n \t\tlast_gen = generation;\n@@ -177,12 +177,12 @@ static int remove_redundant(struct repository *r, struct commit **array, int cnt\n \t\trepo_parse_commit(r, array[i]);\n \tfor (i = 0; i < cnt; i++) {\n \t\tstruct commit_list *common;\n-\t\tuint32_t min_generation = commit_graph_generation(array[i]);\n+\t\ttimestamp_t min_generation = commit_graph_generation(array[i]);\n \n \t\tif (redundant[i])\n \t\t\tcontinue;\n \t\tfor (j = filled = 0; j < cnt; j++) {\n-\t\t\tuint32_t curr_generation;\n+\t\t\ttimestamp_t curr_generation;\n \t\t\tif (i == j || redundant[j])\n \t\t\t\tcontinue;\n \t\t\tfilled_index[filled] = j;\n@@ -321,7 +321,7 @@ int repo_in_merge_bases_many(struct repository *r, struct commit *commit,\n {\n \tstruct commit_list *bases;\n \tint ret = 0, i;\n-\tuint32_t generation, max_generation = GENERATION_NUMBER_ZERO;\n+\ttimestamp_t generation, max_generation = GENERATION_NUMBER_ZERO;\n \n \tif (repo_parse_commit(r, commit))\n \t\treturn ret;\n@@ -470,7 +470,7 @@ static int in_commit_list(const struct commit_list *want, struct commit *c)\n static enum contains_result contains_test(struct commit *candidate,\n \t\t\t\t\t  const struct commit_list *want,\n \t\t\t\t\t  struct contains_cache *cache,\n-\t\t\t\t\t  uint32_t cutoff)\n+\t\t\t\t\t  timestamp_t cutoff)\n {\n \tenum contains_result *cached = contains_cache_at(cache, candidate);\n \n@@ -506,11 +506,11 @@ static enum contains_result contains_tag_algo(struct commit *candidate,\n {\n \tstruct contains_stack contains_stack = { 0, 0, NULL };\n \tenum contains_result result;\n-\tuint32_t cutoff = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t cutoff = GENERATION_NUMBER_INFINITY;\n \tconst struct commit_list *p;\n \n \tfor (p = want; p; p = p->next) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = p->item;\n \t\tload_commit_graph_info(the_repository, c);\n \t\tgeneration = commit_graph_generation(c);\n@@ -566,8 +566,8 @@ static int compare_commits_by_gen(const void *_a, const void *_b)\n \tconst struct commit *a = *(const struct commit * const *)_a;\n \tconst struct commit *b = *(const struct commit * const *)_b;\n \n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n+\ttimestamp_t generation_a = commit_graph_generation(a);\n+\ttimestamp_t generation_b = commit_graph_generation(b);\n \n \tif (generation_a < generation_b)\n \t\treturn -1;\n@@ -580,7 +580,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation)\n+\t\t\t\t timestamp_t min_generation)\n {\n \tstruct commit **list = NULL;\n \tint i;\n@@ -681,13 +681,13 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \ttime_t min_commit_date = cutoff_by_min_date ? from->item->date : 0;\n \tstruct commit_list *from_iter = from, *to_iter = to;\n \tint result;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \n \twhile (from_iter) {\n \t\tadd_object_array(&from_iter->item->object, NULL, &from_objs);\n \n \t\tif (!parse_commit(from_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (from_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = from_iter->item->date;\n \n@@ -701,7 +701,7 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \n \twhile (to_iter) {\n \t\tif (!parse_commit(to_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (to_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = to_iter->item->date;\n \n@@ -741,13 +741,13 @@ struct commit_list *get_reachable_subset(struct commit **from, int nr_from,\n \tstruct commit_list *found_commits = NULL;\n \tstruct commit **to_last = to + nr_to;\n \tstruct commit **from_last = from + nr_from;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \tint num_to_find = 0;\n \n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \n \tfor (item = to; item < to_last; item++) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = *item;\n \n \t\tparse_commit(c);\ndiff --git a/commit-reach.h b/commit-reach.h\nindex b49ad71a317..148b56fea50 100644\n--- a/commit-reach.h\n+++ b/commit-reach.h\n@@ -87,7 +87,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation);\n+\t\t\t\t timestamp_t min_generation);\n int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \t\t       int commit_date_cutoff);\n \ndiff --git a/commit.c b/commit.c\nindex bab8d5ab07c..4c717329ee0 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -753,8 +753,8 @@ int compare_commits_by_author_date(const void *a_, const void *b_,\n int compare_commits_by_gen_then_commit_date(const void *a_, const void *b_, void *unused)\n {\n \tconst struct commit *a = a_, *b = b_;\n-\tconst uint32_t generation_a = commit_graph_generation(a),\n-\t\t       generation_b = commit_graph_generation(b);\n+\tconst timestamp_t generation_a = commit_graph_generation(a),\n+\t\t\t  generation_b = commit_graph_generation(b);\n \n \t/* newer commits first */\n \tif (generation_a < generation_b)\ndiff --git a/commit.h b/commit.h\nindex f4e7b0158e2..742d96c41e8 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -11,8 +11,8 @@\n #include \"commit-slab.h\"\n \n #define COMMIT_NOT_FROM_GRAPH 0xFFFFFFFF\n-#define GENERATION_NUMBER_INFINITY 0xFFFFFFFF\n-#define GENERATION_NUMBER_MAX 0x3FFFFFFF\n+#define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n+#define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n \n struct commit_list {\ndiff --git a/revision.c b/revision.c\nindex be2d828a4cc..31fd3219e65 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3300,7 +3300,7 @@ define_commit_slab(indegree_slab, int);\n define_commit_slab(author_date_slab, timestamp_t);\n \n struct topo_walk_info {\n-\tuint32_t min_generation;\n+\ttimestamp_t min_generation;\n \tstruct prio_queue explore_queue;\n \tstruct prio_queue indegree_queue;\n \tstruct prio_queue topo_queue;\n@@ -3368,7 +3368,7 @@ static void explore_walk_step(struct rev_info *revs)\n }\n \n static void explore_to_depth(struct rev_info *revs,\n-\t\t\t     uint32_t gen_cutoff)\n+\t\t\t     timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3413,7 +3413,7 @@ static void indegree_walk_step(struct rev_info *revs)\n }\n \n static void compute_indegrees_to_depth(struct rev_info *revs,\n-\t\t\t\t       uint32_t gen_cutoff)\n+\t\t\t\t       timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3471,7 +3471,7 @@ static void init_topo_walk(struct rev_info *revs)\n \tinfo->min_generation = GENERATION_NUMBER_INFINITY;\n \tfor (list = revs->commits; list; list = list->next) {\n \t\tstruct commit *c = list->item;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (repo_parse_commit_gently(revs->repo, c, 1))\n \t\t\tcontinue;\n@@ -3539,7 +3539,7 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n \tfor (p = commit->parents; p; p = p->next) {\n \t\tstruct commit *parent = p->item;\n \t\tint *pi;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (parent->object.flags & UNINTERESTING)\n \t\t\tcontinue;\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 3b66bf92ba8..b87607e0dd4 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -500,7 +500,7 @@ static int got_oid(struct upload_pack_data *data,\n \n static int ok_to_give_up(struct upload_pack_data *data)\n {\n-\tuint32_t min_generation = GENERATION_NUMBER_ZERO;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_ZERO;\n \n \tif (!data->have_obj.nr)\n \t\treturn 0;\n-- \ngitgitgadget\n\n"},{"id":"414520","messageId":"e571f03d8bd0b2def8e16df68f6cc53ffcf02082.1610820679.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v6 11/11] doc: add corrected commit date info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:18Z","receivedAt":"2021-01-16T18:13:03Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith generation data chunk and corrected commit dates implemented, let's\nupdate the technical documentation for commit-graph.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n .../technical/commit-graph-format.txt         | 28 +++++--\n Documentation/technical/commit-graph.txt      | 77 +++++++++++++++----\n 2 files changed, 86 insertions(+), 19 deletions(-)\n\ndiff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\nindex b3b58880b92..b6658eff188 100644\n--- a/Documentation/technical/commit-graph-format.txt\n+++ b/Documentation/technical/commit-graph-format.txt\n@@ -4,11 +4,7 @@ Git commit graph format\n The Git commit graph stores a list of commit OIDs and some associated\n metadata, including:\n \n-- The generation number of the commit. Commits with no parents have\n-  generation number 1; commits with parents have generation number\n-  one more than the maximum generation number of its parents. We\n-  reserve zero as special, and can be used to mark a generation\n-  number invalid or as \"not computed\".\n+- The generation number of the commit.\n \n - The root tree OID.\n \n@@ -86,13 +82,33 @@ CHUNK DATA:\n       position. If there are more than two parents, the second value\n       has its most-significant bit on and the other bits store an array\n       position into the Extra Edge List chunk.\n-    * The next 8 bytes store the generation number of the commit and\n+    * The next 8 bytes store the topological level (generation number v1)\n+      of the commit and\n       the commit time in seconds since EPOCH. The generation number\n       uses the higher 30 bits of the first 4 bytes, while the commit\n       time uses the 32 bits of the second 4 bytes, along with the lowest\n       2 bits of the lowest byte, storing the 33rd and 34th bit of the\n       commit time.\n \n+  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes) [Optional]\n+    * This list of 4-byte values store corrected commit date offsets for the\n+      commits, arranged in the same order as commit data chunk.\n+    * If the corrected commit date offset cannot be stored within 31 bits,\n+      the value has its most-significant bit on and the other bits store\n+      the position of corrected commit date into the Generation Data Overflow\n+      chunk.\n+    * Generation Data chunk is present only when commit-graph file is written\n+      by compatible versions of Git and in case of split commit-graph chains,\n+      the topmost layer also has Generation Data chunk.\n+\n+  Generation Data Overflow (ID: {'G', 'D', 'O', 'V' }) [Optional]\n+    * This list of 8-byte values stores the corrected commit date offsets\n+      for commits with corrected commit date offsets that cannot be\n+      stored within 31 bits.\n+    * Generation Data Overflow chunk is present only when Generation Data\n+      chunk is present and atleast one corrected commit date offset cannot\n+      be stored within 31 bits.\n+\n   Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n       This list of 4-byte values store the second through nth parents for\n       all octopus merges. The second parent value in the commit data stores\ndiff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\nindex f14a7659aa8..f05e7bda1a9 100644\n--- a/Documentation/technical/commit-graph.txt\n+++ b/Documentation/technical/commit-graph.txt\n@@ -38,14 +38,31 @@ A consumer may load the following info for a commit from the graph:\n \n Values 1-4 satisfy the requirements of parse_commit_gently().\n \n-Define the \"generation number\" of a commit recursively as follows:\n+There are two definitions of generation number:\n+1. Corrected committer dates (generation number v2)\n+2. Topological levels (generation nummber v1)\n \n- * A commit with no parents (a root commit) has generation number one.\n+Define \"corrected committer date\" of a commit recursively as follows:\n \n- * A commit with at least one parent has generation number one more than\n-   the largest generation number among its parents.\n+ * A commit with no parents (a root commit) has corrected committer date\n+    equal to its committer date.\n \n-Equivalently, the generation number of a commit A is one more than the\n+ * A commit with at least one parent has corrected committer date equal to\n+    the maximum of its commiter date and one more than the largest corrected\n+    committer date among its parents.\n+\n+ * As a special case, a root commit with timestamp zero has corrected commit\n+    date of 1, to be able to distinguish it from GENERATION_NUMBER_ZERO\n+    (that is, an uncomputed corrected commit date).\n+\n+Define the \"topological level\" of a commit recursively as follows:\n+\n+ * A commit with no parents (a root commit) has topological level of one.\n+\n+ * A commit with at least one parent has topological level one more than\n+   the largest topological level among its parents.\n+\n+Equivalently, the topological level of a commit A is one more than the\n length of a longest path from A to a root commit. The recursive definition\n is easier to use for computation and observing the following property:\n \n@@ -60,6 +77,9 @@ is easier to use for computation and observing the following property:\n     generation numbers, then we always expand the boundary commit with highest\n     generation number and can easily detect the stopping condition.\n \n+The property applies to both versions of generation number, that is both\n+corrected committer dates and topological levels.\n+\n This property can be used to significantly reduce the time it takes to\n walk commits and determine topological relationships. Without generation\n numbers, the general heuristic is the following:\n@@ -67,7 +87,9 @@ numbers, the general heuristic is the following:\n     If A and B are commits with commit time X and Y, respectively, and\n     X < Y, then A _probably_ cannot reach B.\n \n-This heuristic is currently used whenever the computation is allowed to\n+In absence of corrected commit dates (for example, old versions of Git or\n+mixed generation graph chains),\n+this heuristic is currently used whenever the computation is allowed to\n violate topological relationships due to clock skew (such as \"git log\"\n with default order), but is not used when the topological order is\n required (such as merge base calculations, \"git log --graph\").\n@@ -77,7 +99,7 @@ in the commit graph. We can treat these commits as having \"infinite\"\n generation number and walk until reaching commits with known generation\n number.\n \n-We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n+We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n in the commit-graph file. If a commit-graph file was written by a version\n of Git that did not compute generation numbers, then those commits will\n have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n@@ -93,12 +115,12 @@ fully-computed generation numbers. Using strict inequality may result in\n walking a few extra commits, but the simplicity in dealing with commits\n with generation number *_INFINITY or *_ZERO is valuable.\n \n-We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n-generation numbers are computed to be at least this value. We limit at\n-this value since it is the largest value that can be stored in the\n-commit-graph file using the 30 bits available to generation numbers. This\n-presents another case where a commit can have generation number equal to\n-that of a parent.\n+We use the macro GENERATION_NUMBER_V1_MAX = 0x3FFFFFFF for commits whose\n+topological levels (generation number v1) are computed to be at least\n+this value. We limit at this value since it is the largest value that\n+can be stored in the commit-graph file using the 30 bits available\n+to topological levels. This presents another case where a commit can\n+have generation number equal to that of a parent.\n \n Design Details\n --------------\n@@ -267,6 +289,35 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n number of commits) could be extracted into config settings for full\n flexibility.\n \n+## Handling Mixed Generation Number Chains\n+\n+With the introduction of generation number v2 and generation data chunk, the\n+following scenario is possible:\n+\n+1. \"New\" Git writes a commit-graph with the corrected commit dates.\n+2. \"Old\" Git writes a split commit-graph on top without corrected commit dates.\n+\n+A naive approach of using the newest available generation number from\n+each layer would lead to violated expectations: the lower layer would\n+use corrected commit dates which are much larger than the topological\n+levels of the higher layer. For this reason, Git inspects the topmost\n+layer to see if the layer is missing corrected commit dates. In such a case\n+Git only uses topological level for generation numbers.\n+\n+When writing a new layer in split commit-graph, we write corrected commit\n+dates if the topmost layer has corrected commit dates written. This\n+guarantees that if a layer has corrected commit dates, all lower layers\n+must have corrected commit dates as well.\n+\n+When merging layers, we do not consider whether the merged layers had corrected\n+commit dates. Instead, the new layer will have corrected commit dates if the\n+layer below the new layer has corrected commit dates.\n+\n+While writing or merging layers, if the new layer is the only layer, it will\n+have corrected commit dates when written by compatible versions of Git. Thus,\n+rewriting split commit-graph as a single file (`--split=replace`) creates a\n+single layer with corrected commit dates.\n+\n ## Deleting graph-{hash} files\n \n After a new tip file is written, some `graph-{hash}` files may no longer\n-- \ngitgitgadget\n"},{"id":"414521","messageId":"fbd8feb5d8c0a52a9451c82db930baa4fa5f840f.1610820679.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v6 05/11] commit-graph: add a slab to store topological levels","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:12Z","receivedAt":"2021-01-16T18:13:03Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a later commit we will introduce corrected commit date as the\ngeneration number v2. Corrected commit dates will be stored in the new\nseperate Generation Data chunk. However, to ensure backwards\ncompatibility with \"Old\" Git we need to continue to write generation\nnumber v1 (topological levels) to the commit data chunk. Thus, we need\nto compute and store both versions of generation numbers to write the\ncommit-graph file.\n\nTherefore, let's introduce a commit-slab `topo_level_slab` to store\ntopological levels; corrected commit date will be stored in the member\n`generation` of struct commit_graph_data.\n\nThe macros `GENERATION_NUMBER_INFINITY` and `GENERATION_NUMBER_ZERO`\nmark commits not in the commit-graph file and commits written by a\nversion of Git that did not compute generation numbers respectively.\nGeneration numbers are computed identically for both kinds of commits.\n\nA \"slab-miss\" should return `GENERATION_NUMBER_INFINITY` as the commit\nis not in the commit-graph file. However, since the slab is\nzero-initialized, it returns 0 (or rather `GENERATION_NUMBER_ZERO`).\nThus, we no longer need to check if the topological level of a commit is\n`GENERATION_NUMBER_INFINITY`.\n\nWe will add a pointer to the slab in `struct write_commit_graph_context`\nand `struct commit_graph` to populate the slab in\n`fill_commit_graph_info` if the commit has a pre-computed topological\nlevel as in case of split commit-graphs.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 45 ++++++++++++++++++++++++++++++---------------\n commit-graph.h |  1 +\n 2 files changed, 31 insertions(+), 15 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 3d59b8b905d..3b69c3cc329 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -64,6 +64,8 @@ void git_test_write_commit_graph_or_die(void)\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n+define_commit_slab(topo_level_slab, uint32_t);\n+\n /* Keep track of the order in which commits are added to our list. */\n define_commit_slab(commit_pos, int);\n static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n@@ -772,6 +774,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\n+\tif (g->topo_levels)\n+\t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n static inline void set_commit_tree(struct commit *c, struct tree *t)\n@@ -960,6 +965,7 @@ struct write_commit_graph_context {\n \t\t changed_paths:1,\n \t\t order_by_pack:1;\n \n+\tstruct topo_level_slab *topo_levels;\n \tconst struct commit_graph_opts *opts;\n \tsize_t total_bloom_filter_data_size;\n \tconst struct bloom_filter_settings *bloom_settings;\n@@ -1106,7 +1112,7 @@ static int write_graph_chunk_data(struct hashfile *f,\n \t\telse\n \t\t\tpackedDate[0] = 0;\n \n-\t\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n+\t\tpackedDate[0] |= htonl(*topo_level_slab_at(ctx->topo_levels, *list) << 2);\n \n \t\tpackedDate[1] = htonl((*list)->date);\n \t\thashwrite(f, packedDate, 8);\n@@ -1336,11 +1342,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n+\t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n-\t\t    generation != GENERATION_NUMBER_ZERO)\n+\t\tif (level != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1348,29 +1353,26 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit *current = list->item;\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n-\t\t\tuint32_t max_generation = 0;\n+\t\t\tuint32_t max_level = 0;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n-\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n+\t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n \n-\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n-\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n+\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (generation > max_generation) {\n-\t\t\t\t\tmax_generation = generation;\n+\t\t\t\t} else if (level > max_level) {\n+\t\t\t\t\tmax_level = level;\n \t\t\t\t}\n \t\t\t}\n \n \t\t\tif (all_parents_computed) {\n-\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n-\n-\t\t\t\tdata->generation = max_generation + 1;\n \t\t\t\tpop_commit(&list);\n \n-\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n-\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n+\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n+\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n+\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n \t\t\t}\n \t\t}\n \t}\n@@ -2106,6 +2108,7 @@ int write_commit_graph(struct object_directory *odb,\n \tint res = 0;\n \tint replace = 0;\n \tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n+\tstruct topo_level_slab topo_levels;\n \n \tprepare_repo_settings(the_repository);\n \tif (!the_repository->settings.core_commit_graph) {\n@@ -2132,6 +2135,18 @@ int write_commit_graph(struct object_directory *odb,\n \t\t\t\t\t\t\t bloom_settings.max_changed_paths);\n \tctx->bloom_settings = &bloom_settings;\n \n+\tinit_topo_level_slab(&topo_levels);\n+\tctx->topo_levels = &topo_levels;\n+\n+\tif (ctx->r->objects->commit_graph) {\n+\t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n+\n+\t\twhile (g) {\n+\t\t\tg->topo_levels = &topo_levels;\n+\t\t\tg = g->base_graph;\n+\t\t}\n+\t}\n+\n \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n \t\tctx->changed_paths = 1;\n \tif (!(flags & COMMIT_GRAPH_NO_WRITE_BLOOM_FILTERS)) {\ndiff --git a/commit-graph.h b/commit-graph.h\nindex f8e92500c6e..00f00745b79 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -73,6 +73,7 @@ struct commit_graph {\n \tconst unsigned char *chunk_bloom_indexes;\n \tconst unsigned char *chunk_bloom_data;\n \n+\tstruct topo_level_slab *topo_levels;\n \tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \n-- \ngitgitgadget\n\n"},{"id":"414522","messageId":"6d0696ae216156e820a8d19ebbae00039a0f3509.1610820679.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v6 08/11] commit-graph: implement generation data chunk","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:15Z","receivedAt":"2021-01-16T18:13:03Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nAs discovered by Ævar, we cannot increment graph version to\ndistinguish between generation numbers v1 and v2 [1]. Thus, one of\npre-requistes before implementing generation number v2 was to\ndistinguish between graph versions in a backwards compatible manner.\n\nWe are going to introduce a new chunk called Generation DATa chunk (or\nGDAT). GDAT will store corrected committer date offsets whereas CDAT\nwill still store topological level.\n\nOld Git does not understand GDAT chunk and would ignore it, reading\ntopological levels from CDAT. New Git can parse GDAT and take advantage\nof newer generation numbers, falling back to topological levels when\nGDAT chunk is missing (as it would happen with a commit-graph written\nby old Git).\n\nWe introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\nwhich forces commit-graph file to be written without generation data\nchunk to emulate a commit-graph file written by old Git.\n\nTo minimize the space required to store corrrected commit date, Git\nstores corrected commit date offsets into the commit-graph file, instea\nof corrected commit dates. This saves us 4 bytes per commit, decreasing\nthe GDAT chunk size by half, but it's possible for the offset to\noverflow the 4-bytes allocated for storage. As such overflows are and\nshould be exceedingly rare, we use the following overflow management\nscheme:\n\nWe introduce a new commit-graph chunk, Generation Data OVerflow ('GDOV')\nto store corrected commit dates for commits with offsets greater than\nGENERATION_NUMBER_V2_OFFSET_MAX.\n\nIf the offset is greater than GENERATION_NUMBER_V2_OFFSET_MAX, we set\nthe MSB of the offset and the other bits store the position of corrected\ncommit date in GDOV chunk, similar to how Extra Edge List is maintained.\n\nWe test the overflow-related code with the following repo history:\n\n           F - N - U\n          /         \\\nU - N - U            N\n         \\          /\n\t  N - F - N\n\nWhere the commits denoted by U have committer date of zero seconds\nsince Unix epoch, the commits denoted by N have committer date of\n1112354055 (default committer date for the test suite) seconds since\nUnix epoch and the commits denoted by F have committer date of\n(2 ^ 31 - 2) seconds since Unix epoch.\n\nThe largest offset observed is 2 ^ 31, just large enough to overflow.\n\n[1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                | 114 ++++++++++++++++++++++++++++++----\n commit-graph.h                |   3 +\n commit.h                      |   1 +\n t/README                      |   3 +\n t/helper/test-read-graph.c    |   4 ++\n t/t4216-log-bloom.sh          |   4 +-\n t/t5318-commit-graph.sh       |  79 +++++++++++++++++++----\n t/t5324-split-commit-graph.sh |  12 ++--\n t/t6600-test-reach.sh         |   6 ++\n t/test-lib-functions.sh       |   6 ++\n 10 files changed, 200 insertions(+), 32 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex a899f429093..7365958d9d3 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -38,11 +38,13 @@ void git_test_write_commit_graph_or_die(void)\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n+#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n+#define GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW 0x47444f56 /* \"GDOV\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n-#define MAX_NUM_CHUNKS 7\n+#define MAX_NUM_CHUNKS 9\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -61,6 +63,8 @@ void git_test_write_commit_graph_or_die(void)\n #define GRAPH_MIN_SIZE (GRAPH_HEADER_SIZE + 4 * GRAPH_CHUNKLOOKUP_WIDTH \\\n \t\t\t+ GRAPH_FANOUT_SIZE + the_hash_algo->rawsz)\n \n+#define CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW (1ULL << 31)\n+\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n@@ -394,6 +398,20 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n \t\t\tbreak;\n \n+\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n+\t\t\tif (graph->chunk_generation_data)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW:\n+\t\t\tif (graph->chunk_generation_data_overflow)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_generation_data_overflow = data + chunk_offset;\n+\t\t\tbreak;\n+\n \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n \t\t\tif (graph->chunk_extra_edges)\n \t\t\t\tchunk_repeated = 1;\n@@ -754,8 +772,8 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n {\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n-\tuint32_t lex_index;\n-\tuint64_t date_high, date_low;\n+\tuint32_t lex_index, offset_pos;\n+\tuint64_t date_high, date_low, offset;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n@@ -773,7 +791,19 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\tif (g->chunk_generation_data) {\n+\t\toffset = (timestamp_t)get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n+\n+\t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n+\t\t\tif (!g->chunk_generation_data_overflow)\n+\t\t\t\tdie(_(\"commit-graph requires overflow generation data but has none\"));\n+\n+\t\t\toffset_pos = offset ^ CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW;\n+\t\t\tgraph_data->generation = get_be64(g->chunk_generation_data_overflow + 8 * offset_pos);\n+\t\t} else\n+\t\t\tgraph_data->generation = item->date + offset;\n+\t} else\n+\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n \n \tif (g->topo_levels)\n \t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n@@ -945,6 +975,7 @@ struct write_commit_graph_context {\n \tstruct oid_array oids;\n \tstruct packed_commit_list commits;\n \tint num_extra_edges;\n+\tint num_generation_data_overflows;\n \tunsigned long approx_nr_objects;\n \tstruct progress *progress;\n \tint progress_done;\n@@ -963,7 +994,8 @@ struct write_commit_graph_context {\n \t\t report_progress:1,\n \t\t split:1,\n \t\t changed_paths:1,\n-\t\t order_by_pack:1;\n+\t\t order_by_pack:1,\n+\t\t write_generation_data:1;\n \n \tstruct topo_level_slab *topo_levels;\n \tconst struct commit_graph_opts *opts;\n@@ -1123,6 +1155,45 @@ static int write_graph_chunk_data(struct hashfile *f,\n \treturn 0;\n }\n \n+static int write_graph_chunk_generation_data(struct hashfile *f,\n+\t\t\t\t\t      struct write_commit_graph_context *ctx)\n+{\n+\tint i, num_generation_data_overflows = 0;\n+\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n+\n+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n+\t\t\toffset = CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW | num_generation_data_overflows;\n+\t\t\tnum_generation_data_overflows++;\n+\t\t}\n+\n+\t\thashwrite_be32(f, offset);\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int write_graph_chunk_generation_data_overflow(struct hashfile *f,\n+\t\t\t\t\t\t       struct write_commit_graph_context *ctx)\n+{\n+\tint i;\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n+\n+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n+\t\t\thashwrite_be32(f, offset >> 32);\n+\t\t\thashwrite_be32(f, (uint32_t) offset);\n+\t\t}\n+\t}\n+\n+\treturn 0;\n+}\n+\n static int write_graph_chunk_extra_edges(struct hashfile *f,\n \t\t\t\t\t struct write_commit_graph_context *ctx)\n {\n@@ -1386,6 +1457,9 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n \t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n \t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n+\n+\t\t\t\tif (commit_graph_data_at(current)->generation - current->date > GENERATION_NUMBER_V2_OFFSET_MAX)\n+\t\t\t\t\tctx->num_generation_data_overflows++;\n \t\t\t}\n \t\t}\n \t}\n@@ -1719,6 +1793,21 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tchunks[2].id = GRAPH_CHUNKID_DATA;\n \tchunks[2].size = (hashsz + 16) * ctx->commits.nr;\n \tchunks[2].write_fn = write_graph_chunk_data;\n+\n+\tif (git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0))\n+\t\tctx->write_generation_data = 0;\n+\tif (ctx->write_generation_data) {\n+\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA;\n+\t\tchunks[num_chunks].size = sizeof(uint32_t) * ctx->commits.nr;\n+\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data;\n+\t\tnum_chunks++;\n+\t}\n+\tif (ctx->num_generation_data_overflows) {\n+\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW;\n+\t\tchunks[num_chunks].size = sizeof(timestamp_t) * ctx->num_generation_data_overflows;\n+\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data_overflow;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_extra_edges) {\n \t\tchunks[num_chunks].id = GRAPH_CHUNKID_EXTRAEDGES;\n \t\tchunks[num_chunks].size = 4 * ctx->num_extra_edges;\n@@ -2139,6 +2228,8 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n \tctx->opts = opts;\n \tctx->total_bloom_filter_data_size = 0;\n+\tctx->write_generation_data = 1;\n+\tctx->num_generation_data_overflows = 0;\n \n \tbloom_settings.bits_per_entry = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\",\n \t\t\t\t\t\t      bloom_settings.bits_per_entry);\n@@ -2445,16 +2536,17 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t\tcontinue;\n \n \t\t/*\n-\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n-\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n-\t\t * extra logic in the following condition.\n+\t\t * If we are using topological level and one of our parents has\n+\t\t * generation GENERATION_NUMBER_V1_MAX, then our generation is\n+\t\t * also GENERATION_NUMBER_V1_MAX. Decrement to avoid extra logic\n+\t\t * in the following condition.\n \t\t */\n-\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n+\t\tif (!g->chunk_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n \t\t\tmax_generation--;\n \n \t\tgeneration = commit_graph_generation(graph_commit);\n-\t\tif (generation != max_generation + 1)\n-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" != %\"PRItime),\n+\t\tif (generation < max_generation + 1)\n+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n \t\t\t\t     oid_to_hex(&cur_oid),\n \t\t\t\t     generation,\n \t\t\t\t     max_generation + 1);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 2e9aa7824ee..19a02001fde 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -6,6 +6,7 @@\n #include \"oidset.h\"\n \n #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n+#define GIT_TEST_COMMIT_GRAPH_NO_GDAT \"GIT_TEST_COMMIT_GRAPH_NO_GDAT\"\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE \"GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE\"\n #define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n \n@@ -68,6 +69,8 @@ struct commit_graph {\n \tconst uint32_t *chunk_oid_fanout;\n \tconst unsigned char *chunk_oid_lookup;\n \tconst unsigned char *chunk_commit_data;\n+\tconst unsigned char *chunk_generation_data;\n+\tconst unsigned char *chunk_generation_data_overflow;\n \tconst unsigned char *chunk_extra_edges;\n \tconst unsigned char *chunk_base_graphs;\n \tconst unsigned char *chunk_bloom_indexes;\ndiff --git a/commit.h b/commit.h\nindex 742d96c41e8..eff94f3f7c2 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -14,6 +14,7 @@\n #define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n #define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n+#define GENERATION_NUMBER_V2_OFFSET_MAX ((1ULL << 31) - 1)\n \n struct commit_list {\n \tstruct commit *item;\ndiff --git a/t/README b/t/README\nindex c730a707705..8a121487279 100644\n--- a/t/README\n+++ b/t/README\n@@ -393,6 +393,9 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n+commit-graph to be written without generation data chunk.\n+\n GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n commit-graph write to compute and write changed path Bloom filters for\n every 'git commit-graph write', as if the `--changed-paths` option was\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex 5f585a17256..75927b2c81d 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -33,6 +33,10 @@ int cmd__read_graph(int argc, const char **argv)\n \t\tprintf(\" oid_lookup\");\n \tif (graph->chunk_commit_data)\n \t\tprintf(\" commit_metadata\");\n+\tif (graph->chunk_generation_data)\n+\t\tprintf(\" generation_data\");\n+\tif (graph->chunk_generation_data_overflow)\n+\t\tprintf(\" generation_data_overflow\");\n \tif (graph->chunk_extra_edges)\n \t\tprintf(\" extra_edges\");\n \tif (graph->chunk_bloom_indexes)\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex d11040ce41c..dbde0161882 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -40,11 +40,11 @@ test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n '\n \n graph_read_expect () {\n-\tNUM_CHUNKS=5\n+\tNUM_CHUNKS=6\n \tcat >expect <<- EOF\n \theader: 43475048 1 $(test_oid oid_version) $NUM_CHUNKS 0\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n \tEOF\n \ttest-tool read-graph >actual &&\n \ttest_cmp expect actual\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 2ed0c1544da..fa27df579a5 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -76,7 +76,7 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n graph_read_expect() {\n \tOPTIONAL=\"\"\n \tNUM_CHUNKS=3\n-\tif test ! -z $2\n+\tif test ! -z \"$2\"\n \tthen\n \t\tOPTIONAL=\" $2\"\n \t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n@@ -103,14 +103,14 @@ test_expect_success 'exit with correct error on bad input to --stdin-commits' '\n \t# valid commit and tree OID\n \tgit rev-parse HEAD HEAD^{tree} >in &&\n \tgit commit-graph write --stdin-commits <in &&\n-\tgraph_read_expect 3\n+\tgraph_read_expect 3 generation_data\n '\n \n test_expect_success 'write graph' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"3\"\n+\tgraph_read_expect \"3\" generation_data\n '\n \n test_expect_success POSIXPERM 'write graph has correct permissions' '\n@@ -219,7 +219,7 @@ test_expect_success 'write graph with merges' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'merge 1 vs 2' full merge/1 merge/2\n@@ -254,7 +254,7 @@ test_expect_success 'write graph with new commit' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'full graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -264,7 +264,7 @@ test_expect_success 'write graph with nothing new' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'cleared graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -274,7 +274,7 @@ test_expect_success 'build graph from latest pack with closure' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tcat new-idx | git commit-graph write --stdin-packs &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"9\" \"extra_edges\"\n+\tgraph_read_expect \"9\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'graph from pack, commit 8 vs merge 1' full commits/8 merge/1\n@@ -287,7 +287,7 @@ test_expect_success 'build graph from commits with closure' '\n \tgit rev-parse merge/1 >>commits-in &&\n \tcat commits-in | git commit-graph write --stdin-commits &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"6\"\n+\tgraph_read_expect \"6\" \"generation_data\"\n '\n \n graph_git_behavior 'graph from commits, commit 8 vs merge 1' full commits/8 merge/1\n@@ -297,7 +297,7 @@ test_expect_success 'build graph from commits with append' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit rev-parse merge/3 | git commit-graph write --stdin-commits --append &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -307,7 +307,7 @@ test_expect_success 'build graph using --reachable' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write --reachable &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -328,7 +328,7 @@ test_expect_success 'write graph in bare repo' '\n \tcd \"$TRASH_DIRECTORY/bare\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $baredir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n@@ -454,8 +454,9 @@ test_expect_success 'warn on improper hash version' '\n \n test_expect_success 'git commit-graph verify' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n-\tgit rev-parse commits/8 | git commit-graph write --stdin-commits &&\n-\tgit commit-graph verify >output\n+\tgit rev-parse commits/8 | GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --stdin-commits &&\n+\tgit commit-graph verify >output &&\n+\tgraph_read_expect 9 extra_edges\n '\n \n NUM_COMMITS=9\n@@ -741,4 +742,56 @@ test_expect_success 'corrupt commit-graph write (missing tree)' '\n \t)\n '\n \n+# We test the overflow-related code with the following repo history:\n+#\n+#               4:F - 5:N - 6:U\n+#              /                \\\n+# 1:U - 2:N - 3:U                M:N\n+#              \\                /\n+#               7:N - 8:F - 9:N\n+#\n+# Here the commits denoted by U have committer date of zero seconds\n+# since Unix epoch, the commits denoted by N have committer date\n+# starting from 1112354055 seconds since Unix epoch (default committer\n+# date for the test suite), and the commits denoted by F have committer\n+# date of (2 ^ 31 - 2) seconds since Unix epoch.\n+#\n+# The largest offset observed is 2 ^ 31, just large enough to overflow.\n+#\n+\n+test_expect_success 'set up and verify repo with generation data overflow chunk' '\n+\tobjdir=\".git/objects\" &&\n+\tUNIX_EPOCH_ZERO=\"@0 +0000\" &&\n+\tFUTURE_DATE=\"@2147483646 +0000\" &&\n+\ttest_oid_cache <<-EOF &&\n+\toid_version sha1:1\n+\toid_version sha256:2\n+\tEOF\n+\tcd \"$TRASH_DIRECTORY\" &&\n+\tmkdir repo &&\n+\tcd repo &&\n+\tgit init &&\n+\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 1 &&\n+\ttest_commit 2 &&\n+\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 3 &&\n+\tgit commit-graph write --reachable &&\n+\tgraph_read_expect 3 generation_data &&\n+\ttest_commit --date \"$FUTURE_DATE\" 4 &&\n+\ttest_commit 5 &&\n+\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 6 &&\n+\tgit branch left &&\n+\tgit reset --hard 3 &&\n+\ttest_commit 7 &&\n+\ttest_commit --date \"$FUTURE_DATE\" 8 &&\n+\ttest_commit 9 &&\n+\tgit branch right &&\n+\tgit reset --hard 3 &&\n+\ttest_merge M left right &&\n+\tgit commit-graph write --reachable &&\n+\tgraph_read_expect 10 \"generation_data generation_data_overflow\" &&\n+\tgit commit-graph verify\n+'\n+\n+graph_git_behavior 'generation data overflow chunk repo' repo left right\n+\n test_done\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 4d3842b83b9..587757b62d9 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -13,11 +13,11 @@ test_expect_success 'setup repo' '\n \tinfodir=\".git/objects/info\" &&\n \tgraphdir=\"$infodir/commit-graphs\" &&\n \ttest_oid_cache <<-EOM\n-\tshallow sha1:1760\n-\tshallow sha256:2064\n+\tshallow sha1:2132\n+\tshallow sha256:2436\n \n-\tbase sha1:1376\n-\tbase sha256:1496\n+\tbase sha1:1408\n+\tbase sha256:1528\n \n \toid_version sha1:1\n \toid_version sha256:2\n@@ -31,9 +31,9 @@ graph_read_expect() {\n \t\tNUM_BASE=$2\n \tfi\n \tcat >expect <<- EOF\n-\theader: 43475048 1 $(test_oid oid_version) 3 $NUM_BASE\n+\theader: 43475048 1 $(test_oid oid_version) 4 $NUM_BASE\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n \tEOF\n \ttest-tool read-graph >output &&\n \ttest_cmp expect output\ndiff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\nindex af10f0dc090..e2d33a8a4c4 100755\n--- a/t/t6600-test-reach.sh\n+++ b/t/t6600-test-reach.sh\n@@ -55,6 +55,9 @@ test_expect_success 'setup' '\n \tgit show-ref -s commit-5-5 | git commit-graph write --stdin-commits &&\n \tmv .git/objects/info/commit-graph commit-graph-half &&\n \tchmod u+w commit-graph-half &&\n+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable &&\n+\tmv .git/objects/info/commit-graph commit-graph-no-gdat &&\n+\tchmod u+w commit-graph-no-gdat &&\n \tgit config core.commitGraph true\n '\n \n@@ -67,6 +70,9 @@ run_all_modes () {\n \ttest_cmp expect actual &&\n \tcp commit-graph-half .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n+\ttest_cmp expect actual &&\n+\tcp commit-graph-no-gdat .git/objects/info/commit-graph &&\n+\t\"$@\" <input >actual &&\n \ttest_cmp expect actual\n }\n \ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex 999982fe4a9..3ad712c3acc 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -202,6 +202,12 @@ test_commit () {\n \t\t--signoff)\n \t\t\tsignoff=\"$1\"\n \t\t\t;;\n+\t\t--date)\n+\t\t\tnotick=yes\n+\t\t\tGIT_COMMITTER_DATE=\"$2\"\n+\t\t\tGIT_AUTHOR_DATE=\"$2\"\n+\t\t\tshift\n+\t\t\t;;\n \t\t-C)\n \t\t\tindir=\"$2\"\n \t\t\tshift\n-- \ngitgitgadget\n\n"},{"id":"414523","messageId":"ba1f2c5555f152b41b27835ed3980155d76eff95.1610820679.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v6 10/11] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:17Z","receivedAt":"2021-01-16T18:13:03Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\n091f4cf (commit: don't use generation numbers if not needed,\n2018-08-30) changed paint_down_to_common() to use commit dates instead\nof generation numbers v1 (topological levels) as the performance\nregressed on certain topologies. With generation number v2 (corrected\ncommit dates) implemented, we no longer have to rely on commit dates and\ncan use generation numbers.\n\nFor example, the command `git merge-base v4.8 v4.9` on the Linux\nrepository walks 167468 commits, taking 0.135s for committer date and\n167496 commits, taking 0.157s for corrected committer date respectively.\n\nWhile using corrected commit dates, Git walks nearly the same number of\ncommits as commit date, the process is slower as for each comparision we\nhave to access a commit-slab (for corrected committer date) instead of\naccessing struct member (for committer date).\n\nThis change incidentally broke the fragile t6404-recursive-merge test.\nt6404-recursive-merge sets up a unique repository where all commits have\nthe same committer date without a well-defined merge-base.\n\nWhile running tests with GIT_TEST_COMMIT_GRAPH unset, we use committer\ndate as a heuristic in paint_down_to_common(). 6404.1 'combined merge\nconflicts' merges commits in the order:\n- Merge C with B to form an intermediate commit.\n- Merge the intermediate commit with A.\n\nWith GIT_TEST_COMMIT_GRAPH=1, we write a commit-graph and subsequently\nuse the corrected committer date, which changes the order in which\ncommits are merged:\n- Merge A with B to form an intermediate commit.\n- Merge the intermediate commit with C.\n\nWhile resulting repositories are equivalent, 6404.4 'virtual trees were\nprocessed' fails with GIT_TEST_COMMIT_GRAPH=1 as we are selecting\ndifferent merge-bases and thus have different object ids for the\nintermediate commits.\n\nAs this has already causes problems (as noted in 859fdc0 (commit-graph:\ndefine GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph\nwithin t6404-recursive-merge.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c             | 14 ++++++++++++++\n commit-graph.h             |  6 ++++++\n commit-reach.c             |  2 +-\n t/t6404-recursive-merge.sh |  5 ++++-\n 4 files changed, 25 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex d32492f3724..d3d14601d4d 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -714,6 +714,20 @@ int generation_numbers_enabled(struct repository *r)\n \treturn !!first_generation;\n }\n \n+int corrected_commit_dates_enabled(struct repository *r)\n+{\n+\tstruct commit_graph *g;\n+\tif (!prepare_commit_graph(r))\n+\t\treturn 0;\n+\n+\tg = r->objects->commit_graph;\n+\n+\tif (!g->num_commits)\n+\t\treturn 0;\n+\n+\treturn g->read_generation_data;\n+}\n+\n struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r)\n {\n \tstruct commit_graph *g = r->objects->commit_graph;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex ad52130883b..97f3497c279 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -95,6 +95,12 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n  */\n int generation_numbers_enabled(struct repository *r);\n \n+/*\n+ * Return 1 if and only if the repository has a commit-graph\n+ * file and generation data chunk has been written for the file.\n+ */\n+int corrected_commit_dates_enabled(struct repository *r);\n+\n struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n \n enum commit_graph_write_flags {\ndiff --git a/commit-reach.c b/commit-reach.c\nindex 9b24b0378d5..e38771ca5a1 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -39,7 +39,7 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \tint i;\n \ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n-\tif (!min_generation)\n+\tif (!min_generation && !corrected_commit_dates_enabled(r))\n \t\tqueue.compare = compare_commits_by_commit_date;\n \n \tone->object.flags |= PARENT1;\ndiff --git a/t/t6404-recursive-merge.sh b/t/t6404-recursive-merge.sh\nindex b1c3d4dda49..86f74ae5847 100755\n--- a/t/t6404-recursive-merge.sh\n+++ b/t/t6404-recursive-merge.sh\n@@ -15,6 +15,8 @@ GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n export GIT_COMMITTER_DATE\n \n test_expect_success 'setup tests' '\n+\tGIT_TEST_COMMIT_GRAPH=0 &&\n+\texport GIT_TEST_COMMIT_GRAPH &&\n \techo 1 >a1 &&\n \tgit add a1 &&\n \tGIT_AUTHOR_DATE=\"2006-12-12 23:00:00\" git commit -m 1 a1 &&\n@@ -66,7 +68,7 @@ test_expect_success 'setup tests' '\n '\n \n test_expect_success 'combined merge conflicts' '\n-\ttest_must_fail env GIT_TEST_COMMIT_GRAPH=0 git merge -m final G\n+\ttest_must_fail git merge -m final G\n '\n \n test_expect_success 'result contains a conflict' '\n@@ -82,6 +84,7 @@ test_expect_success 'result contains a conflict' '\n '\n \n test_expect_success 'virtual trees were processed' '\n+\t# TODO: fragile test, relies on ambigious merge-base resolution\n \tgit ls-files --stage >out &&\n \n \tcat >expect <<-EOF &&\n-- \ngitgitgadget\n\n"},{"id":"414524","messageId":"fba0d7f3dfe14ea43a1e3307e52d43b9fcf6c683.1610820679.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v6 09/11] commit-graph: use generation v2 only if entire chain does","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-01-16T18:11:16Z","receivedAt":"2021-01-16T18:13:03Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nSince there are released versions of Git that understand generation\nnumbers in the commit-graph's CDAT chunk but do not understand the GDAT\nchunk, the following scenario is possible:\n\n1. \"New\" Git writes a commit-graph with the GDAT chunk.\n2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n\nIf each layer of split commit-graph is treated independently, as it was\nthe case before this commit, with Git inspecting only the current layer\nfor chunk_generation_data pointer, commits in the lower layer (one with\nGDAT) whould have corrected commit date as their generation number,\nwhile commits in the upper layer would have topological levels as their\ngeneration. Corrected commit dates usually have much larger values than\ntopological levels. This means that if we take two commits, one from the\nupper layer, and one reachable from it in the lower layer, then the\nexpectation that the generation of a parent is smaller than the\ngeneration of a child would be violated.\n\nIt is difficult to expose this issue in a test. Since we _start_ with\nartificially low generation numbers, any commit walk that prioritizes\ngeneration numbers will walk all of the commits with high generation\nnumber before walking the commits with low generation number. In all the\ncases I tried, the commit-graph layers themselves \"protect\" any\nincorrect behavior since none of the commits in the lower layer can\nreach the commits in the upper layer.\n\nThis issue would manifest itself as a performance problem in this case,\nespecially with something like \"git log --graph\" since the low\ngeneration numbers would cause the in-degree queue to walk all of the\ncommits in the lower layer before allowing the topo-order queue to write\nanything to output (depending on the size of the upper layer).\n\nTherefore, When writing the new layer in split commit-graph, we write a\nGDAT chunk only if the topmost layer has a GDAT chunk. This guarantees\nthat if a layer has GDAT chunk, all lower layers must have a GDAT chunk\nas well.\n\nRewriting layers follows similar approach: if the topmost layer below\nthe set of layers being rewritten (in the split commit-graph chain)\nexists, and it does not contain GDAT chunk, then the result of rewrite\ndoes not have GDAT chunks either.\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                |  30 +++++-\n commit-graph.h                |   1 +\n t/t5324-split-commit-graph.sh | 181 ++++++++++++++++++++++++++++++++++\n 3 files changed, 210 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 7365958d9d3..d32492f3724 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -614,6 +614,21 @@ static struct commit_graph *load_commit_graph_chain(struct repository *r,\n \treturn graph_chain;\n }\n \n+static void validate_mixed_generation_chain(struct commit_graph *g)\n+{\n+\tint read_generation_data;\n+\n+\tif (!g)\n+\t\treturn;\n+\n+\tread_generation_data = !!g->chunk_generation_data;\n+\n+\twhile (g) {\n+\t\tg->read_generation_data = read_generation_data;\n+\t\tg = g->base_graph;\n+\t}\n+}\n+\n struct commit_graph *read_commit_graph_one(struct repository *r,\n \t\t\t\t\t   struct object_directory *odb)\n {\n@@ -622,6 +637,8 @@ struct commit_graph *read_commit_graph_one(struct repository *r,\n \tif (!g)\n \t\tg = load_commit_graph_chain(r, odb);\n \n+\tvalidate_mixed_generation_chain(g);\n+\n \treturn g;\n }\n \n@@ -791,7 +808,7 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tif (g->chunk_generation_data) {\n+\tif (g->read_generation_data) {\n \t\toffset = (timestamp_t)get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n \n \t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n@@ -2019,6 +2036,13 @@ static void split_graph_merge_strategy(struct write_commit_graph_context *ctx)\n \t\tif (i < ctx->num_commit_graphs_after)\n \t\t\tctx->commit_graph_hash_after[i] = xstrdup(oid_to_hex(&g->oid));\n \n+\t\t/*\n+\t\t * If the topmost remaining layer has generation data chunk, the\n+\t\t * resultant layer also has generation data chunk.\n+\t\t */\n+\t\tif (i == ctx->num_commit_graphs_after - 2)\n+\t\t\tctx->write_generation_data = !!g->chunk_generation_data;\n+\n \t\ti--;\n \t\tg = g->base_graph;\n \t}\n@@ -2343,6 +2367,8 @@ int write_commit_graph(struct object_directory *odb,\n \t} else\n \t\tctx->num_commit_graphs_after = 1;\n \n+\tvalidate_mixed_generation_chain(ctx->r->objects->commit_graph);\n+\n \tcompute_generation_numbers(ctx);\n \n \tif (ctx->changed_paths)\n@@ -2541,7 +2567,7 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t * also GENERATION_NUMBER_V1_MAX. Decrement to avoid extra logic\n \t\t * in the following condition.\n \t\t */\n-\t\tif (!g->chunk_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n+\t\tif (!g->read_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n \t\t\tmax_generation--;\n \n \t\tgeneration = commit_graph_generation(graph_commit);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 19a02001fde..ad52130883b 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -64,6 +64,7 @@ struct commit_graph {\n \tstruct object_directory *odb;\n \n \tuint32_t num_commits_in_base;\n+\tunsigned int read_generation_data;\n \tstruct commit_graph *base_graph;\n \n \tconst uint32_t *chunk_oid_fanout;\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 587757b62d9..8e90f3423b8 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -453,4 +453,185 @@ test_expect_success 'prevent regression for duplicate commits across layers' '\n \tgit -C dup commit-graph verify\n '\n \n+NUM_FIRST_LAYER_COMMITS=64\n+NUM_SECOND_LAYER_COMMITS=16\n+NUM_THIRD_LAYER_COMMITS=7\n+NUM_FOURTH_LAYER_COMMITS=8\n+NUM_FIFTH_LAYER_COMMITS=16\n+SECOND_LAYER_SEQUENCE_START=$(($NUM_FIRST_LAYER_COMMITS + 1))\n+SECOND_LAYER_SEQUENCE_END=$(($SECOND_LAYER_SEQUENCE_START + $NUM_SECOND_LAYER_COMMITS - 1))\n+THIRD_LAYER_SEQUENCE_START=$(($SECOND_LAYER_SEQUENCE_END + 1))\n+THIRD_LAYER_SEQUENCE_END=$(($THIRD_LAYER_SEQUENCE_START + $NUM_THIRD_LAYER_COMMITS - 1))\n+FOURTH_LAYER_SEQUENCE_START=$(($THIRD_LAYER_SEQUENCE_END + 1))\n+FOURTH_LAYER_SEQUENCE_END=$(($FOURTH_LAYER_SEQUENCE_START + $NUM_FOURTH_LAYER_COMMITS - 1))\n+FIFTH_LAYER_SEQUENCE_START=$(($FOURTH_LAYER_SEQUENCE_END + 1))\n+FIFTH_LAYER_SEQUENCE_END=$(($FIFTH_LAYER_SEQUENCE_START + $NUM_FIFTH_LAYER_COMMITS - 1))\n+\n+# Current split graph chain:\n+#\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n+\tgraphdir=\".git/objects/info/commit-graphs\" &&\n+\ttest_oid_cache <<-EOF &&\n+\toid_version sha1:1\n+\toid_version sha256:2\n+\tEOF\n+\tgit init mixed &&\n+\t(\n+\t\tcd mixed &&\n+\t\tgit config core.commitGraph true &&\n+\t\tgit config gc.writeCommitGraph false &&\n+\t\tfor i in $(test_seq $NUM_FIRST_LAYER_COMMITS)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split &&\n+\t\tgraph_read_expect $NUM_FIRST_LAYER_COMMITS &&\n+\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\t\tfor i in $(test_seq $SECOND_LAYER_SEQUENCE_START $SECOND_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n+\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 4 1\n+\t\tnum_commits: $NUM_SECOND_LAYER_COMMITS\n+\t\tchunks: oid_fanout oid_lookup commit_metadata\n+\t\tEOF\n+\t\ttest_cmp expect output &&\n+\t\tgit commit-graph verify &&\n+\t\tcat $graphdir/commit-graph-chain\n+\t)\n+'\n+\n+# The new layer will be added without generation data chunk as it was not\n+# present on the layer underneath it.\n+#\n+#      7 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'do not write generation data chunk if not present on existing tip' '\n+\tgit clone mixed mixed-no-gdat &&\n+\t(\n+\t\tcd mixed-no-gdat &&\n+\t\tfor i in $(test_seq $THIRD_LAYER_SEQUENCE_START $THIRD_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split=no-merge &&\n+\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 4 2\n+\t\tnum_commits: $NUM_THIRD_LAYER_COMMITS\n+\t\tchunks: oid_fanout oid_lookup commit_metadata\n+\t\tEOF\n+\t\ttest_cmp expect output &&\n+\t\tgit commit-graph verify\n+\t)\n+'\n+\n+# Number of commits in each layer of the split-commit graph before merge:\n+#\n+#      8 commits (No GDAT)\n+# ------------------------\n+#      7 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+# The top two layers are merged and do not have generation data chunk as layer below them does\n+# not have generation data chunk.\n+#\n+#     15 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'do not write generation data chunk if the topmost remaining layer does not have generation data chunk' '\n+\tgit clone mixed-no-gdat mixed-merge-no-gdat &&\n+\t(\n+\t\tcd mixed-merge-no-gdat &&\n+\t\tfor i in $(test_seq $FOURTH_LAYER_SEQUENCE_START $FOURTH_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split --size-multiple 1 &&\n+\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 4 2\n+\t\tnum_commits: $(($NUM_THIRD_LAYER_COMMITS + $NUM_FOURTH_LAYER_COMMITS))\n+\t\tchunks: oid_fanout oid_lookup commit_metadata\n+\t\tEOF\n+\t\ttest_cmp expect output &&\n+\t\tgit commit-graph verify\n+\t)\n+'\n+\n+# Number of commits in each layer of the split-commit graph before merge:\n+#\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     15 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+# The top three layers are merged and has generation data chunk as the topmost remaining layer\n+# has generation data chunk.\n+#\n+#     47 commits (GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'write generation data chunk if topmost remaining layer has generation data chunk' '\n+\tgit clone mixed-merge-no-gdat mixed-merge-gdat &&\n+\t(\n+\t\tcd mixed-merge-gdat &&\n+\t\tfor i in $(test_seq $FIFTH_LAYER_SEQUENCE_START $FIFTH_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split --size-multiple 1 &&\n+\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 5 1\n+\t\tnum_commits: $(($NUM_SECOND_LAYER_COMMITS + $NUM_THIRD_LAYER_COMMITS + $NUM_FOURTH_LAYER_COMMITS + $NUM_FIFTH_LAYER_COMMITS))\n+\t\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n+\t\tEOF\n+\t\ttest_cmp expect output\n+\t)\n+'\n+\n+test_expect_success 'write generation data chunk when commit-graph chain is replaced' '\n+\tgit clone mixed mixed-replace &&\n+\t(\n+\t\tcd mixed-replace &&\n+\t\tgit commit-graph write --reachable --split=replace &&\n+\t\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\t\tverify_chain_files_exist $graphdir &&\n+\t\tgraph_read_expect $(($NUM_FIRST_LAYER_COMMITS + $NUM_SECOND_LAYER_COMMITS)) &&\n+\t\tgit commit-graph verify\n+\t)\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"414624","messageId":"2437ba7c-f9d9-34bd-5e08-eff96cadcf91@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-01-18T21:04:14Z","receivedAt":"2021-01-18T21:05:16Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 1/16/2021 1:11 PM, Abhishek Kumar via GitGitGadget wrote:\n> This patch series implements the corrected commit date offsets as generation\n> number v2, along with other pre-requisites.\n...\n> Changes in version 6:\n> \n>  * Fixed typos in commit message for \"commit-graph: implement corrected\n>    commit date\".\n>  * Removed an unnecessary else-block in \"commit-graph: implement corrected\n>    commit date\".\n>  * Validate mixed generation chain correctly while writing in \"commit-graph:\n>    use generation v2 only if the entire chain does\".\n>  * Die if the GDAT chunk indicates data has overflown but there are is no\n>    generation data overflow chunk.\n\nI checked the range-diff and looked once more through the patch\nseries. This version is good to go by my standards.\n\nReviewed-by: Derrick Stolee <dstolee@microsoft.com>\n\nThanks, Abhishek!\n\n"},{"id":"414633","messageId":"YAYFCbVvEL+GbQOl@nand.local","threadId":"53933","inReplyTo":"2437ba7c-f9d9-34bd-5e08-eff96cadcf91@gmail.com","subject":"Re: [PATCH v6 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-18T22:00:41Z","receivedAt":"2021-01-18T22:01:38Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Mon, Jan 18, 2021 at 04:04:14PM -0500, Derrick Stolee wrote:\n> I checked the range-diff and looked once more through the patch\n> series. This version is good to go by my standards.\n>\n> Reviewed-by: Derrick Stolee <dstolee@microsoft.com>\n\nI re-read this series now that it seems to have stabilized, and I agree\nwith Stolee that it LGTM.\n\n  Reviewed-by: Taylor Blau <me@ttaylorr.com>\n\n> Thanks, Abhishek!\n\nIncredible work!\n\nThanks,\nTaylor\n"},{"id":"414646","messageId":"xmqqft2xiyj0.fsf@gitster.c.googlers.com","threadId":"53933","inReplyTo":"2437ba7c-f9d9-34bd-5e08-eff96cadcf91@gmail.com","subject":"Re: [PATCH v6 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-01-19T00:02:59Z","receivedAt":"2021-01-19T00:04:02Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n> On 1/16/2021 1:11 PM, Abhishek Kumar via GitGitGadget wrote:\n>> This patch series implements the corrected commit date offsets as generation\n>> number v2, along with other pre-requisites.\n> ...\n>> Changes in version 6:\n>> \n>>  * Fixed typos in commit message for \"commit-graph: implement corrected\n>>    commit date\".\n>>  * Removed an unnecessary else-block in \"commit-graph: implement corrected\n>>    commit date\".\n>>  * Validate mixed generation chain correctly while writing in \"commit-graph:\n>>    use generation v2 only if the entire chain does\".\n>>  * Die if the GDAT chunk indicates data has overflown but there are is no\n>>    generation data overflow chunk.\n>\n> I checked the range-diff and looked once more through the patch\n> series. This version is good to go by my standards.\n>\n> Reviewed-by: Derrick Stolee <dstolee@microsoft.com>\n\nThanks, both.  I'll give it a (hopefully) final read-over after\nreplacing what we have kept in 'seen'.\n"},{"id":"415045","messageId":"YAwRbsfdErIlTP6v@Abhishek-Arch","threadId":"53933","inReplyTo":"2437ba7c-f9d9-34bd-5e08-eff96cadcf91@gmail.com","subject":"Re: [PATCH v6 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2021-01-23T12:07:10Z","receivedAt":"2021-01-23T12:07:38Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Mon, Jan 18, 2021 at 04:04:14PM -0500, Derrick Stolee wrote:\n> On 1/16/2021 1:11 PM, Abhishek Kumar via GitGitGadget wrote:\n> > This patch series implements the corrected commit date offsets as generation\n> > number v2, along with other pre-requisites.\n> ...\n> > Changes in version 6:\n> > \n> >  * Fixed typos in commit message for \"commit-graph: implement corrected\n> >    commit date\".\n> >  * Removed an unnecessary else-block in \"commit-graph: implement corrected\n> >    commit date\".\n> >  * Validate mixed generation chain correctly while writing in \"commit-graph:\n> >    use generation v2 only if the entire chain does\".\n> >  * Die if the GDAT chunk indicates data has overflown but there are is no\n> >    generation data overflow chunk.\n> \n> I checked the range-diff and looked once more through the patch\n> series. This version is good to go by my standards.\n> \n> Reviewed-by: Derrick Stolee <dstolee@microsoft.com>\n> \n> Thanks, Abhishek!\n> \n\nThanks a lot for the review and continued guidance through out the\npatch series!\n\n- Abhishek\n"},{"id":"415046","messageId":"YAwScMcw4sP1ZAJb@Abhishek-Arch","threadId":"53933","inReplyTo":"YAYFCbVvEL+GbQOl@nand.local","subject":"Re: [PATCH v6 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2021-01-23T12:11:28Z","receivedAt":"2021-01-23T12:12:06Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Mon, Jan 18, 2021 at 05:00:41PM -0500, Taylor Blau wrote:\n> On Mon, Jan 18, 2021 at 04:04:14PM -0500, Derrick Stolee wrote:\n> > I checked the range-diff and looked once more through the patch\n> > series. This version is good to go by my standards.\n> >\n> > Reviewed-by: Derrick Stolee <dstolee@microsoft.com>\n> \n> I re-read this series now that it seems to have stabilized, and I agree\n> with Stolee that it LGTM.\n> \n>   Reviewed-by: Taylor Blau <me@ttaylorr.com>\n> \n> > Thanks, Abhishek!\n> \n> Incredible work!\n\nThanks a lot for the reviews and help in identifying the reason behind\n(relatively) minor performance increase when we switched from useless\n'commit_graph_generation()' calls to direct slab calls.\n\n> \n> Thanks,\n> Taylor\n\nThanks\n- Abhishek\n"},{"id":"415369","messageId":"20210127000454.GA1440011@szeder.dev","threadId":"53933","inReplyTo":"e571f03d8bd0b2def8e16df68f6cc53ffcf02082.1610820679.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 11/11] doc: add corrected commit date info","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2021-01-27T00:04:54Z","receivedAt":"2021-01-27T03:59:13Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Sat, Jan 16, 2021 at 06:11:18PM +0000, Abhishek Kumar via GitGitGadget wrote:\n> With generation data chunk and corrected commit dates implemented, let's\n> update the technical documentation for commit-graph.\n\nThis patch should come much earlier in this series, before patch 07/11\n(commit-graph: implement corrected commit date), or perhaps even\nearlier.  That way if someone were to investigate an issue in this\nseries and checks out one of its commits, then the specification and\nthe will be right there under 'Documentation/technical/'.\n\nFurthermore, a patch introducing a new chunk format is the right place\nto justify the introduction of said new chunk.  What problems does a\nchunk of corrected commit dates solve?  Why does it solve them?  Why\ndo we need corrected commit dates instead of simple commit dates?\nWhat alternatives were considered [1]?  Any other design considerations\nworth mentioning for the benefit of future readers?\n\nNone of the patches' log messages properly explain these, and while\nmuch of these is indeed explained in the cover letter, the cover\nletter will not be part of the history.  Requiring to look up mailing\nlist archives for the justification puts unnecessary burden on other\ndevelopers who might get interested in this feature in the future.\n\nYou might want to take\nhttps://public-inbox.org/git/20200529085038.26008-16-szeder.dev@gmail.com/\nas an inspiration.\n\n\n[1] Please remember the following snippet from SubmittingPatches:\n    \"Try to make sure your explanation can be understood without\n    external resources. Instead of giving a URL to a mailing list\n    archive, summarize the relevant points of the discussion.\"\n\n> Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> ---\n>  .../technical/commit-graph-format.txt         | 28 +++++--\n>  Documentation/technical/commit-graph.txt      | 77 +++++++++++++++----\n>  2 files changed, 86 insertions(+), 19 deletions(-)\n> \n> diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n> index b3b58880b92..b6658eff188 100644\n> --- a/Documentation/technical/commit-graph-format.txt\n> +++ b/Documentation/technical/commit-graph-format.txt\n> @@ -4,11 +4,7 @@ Git commit graph format\n>  The Git commit graph stores a list of commit OIDs and some associated\n>  metadata, including:\n>  \n> -- The generation number of the commit. Commits with no parents have\n> -  generation number 1; commits with parents have generation number\n> -  one more than the maximum generation number of its parents. We\n> -  reserve zero as special, and can be used to mark a generation\n> -  number invalid or as \"not computed\".\n> +- The generation number of the commit.\n>  \n>  - The root tree OID.\n>  \n> @@ -86,13 +82,33 @@ CHUNK DATA:\n>        position. If there are more than two parents, the second value\n>        has its most-significant bit on and the other bits store an array\n>        position into the Extra Edge List chunk.\n> -    * The next 8 bytes store the generation number of the commit and\n> +    * The next 8 bytes store the topological level (generation number v1)\n> +      of the commit and\n>        the commit time in seconds since EPOCH. The generation number\n>        uses the higher 30 bits of the first 4 bytes, while the commit\n>        time uses the 32 bits of the second 4 bytes, along with the lowest\n>        2 bits of the lowest byte, storing the 33rd and 34th bit of the\n>        commit time.\n>  \n> +  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes) [Optional]\n> +    * This list of 4-byte values store corrected commit date offsets for the\n> +      commits, arranged in the same order as commit data chunk.\n> +    * If the corrected commit date offset cannot be stored within 31 bits,\n> +      the value has its most-significant bit on and the other bits store\n> +      the position of corrected commit date into the Generation Data Overflow\n> +      chunk.\n> +    * Generation Data chunk is present only when commit-graph file is written\n> +      by compatible versions of Git and in case of split commit-graph chains,\n> +      the topmost layer also has Generation Data chunk.\n> +\n> +  Generation Data Overflow (ID: {'G', 'D', 'O', 'V' }) [Optional]\n> +    * This list of 8-byte values stores the corrected commit date offsets\n> +      for commits with corrected commit date offsets that cannot be\n> +      stored within 31 bits.\n> +    * Generation Data Overflow chunk is present only when Generation Data\n> +      chunk is present and atleast one corrected commit date offset cannot\n> +      be stored within 31 bits.\n> +\n>    Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n>        This list of 4-byte values store the second through nth parents for\n>        all octopus merges. The second parent value in the commit data stores\n> diff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\n> index f14a7659aa8..f05e7bda1a9 100644\n> --- a/Documentation/technical/commit-graph.txt\n> +++ b/Documentation/technical/commit-graph.txt\n> @@ -38,14 +38,31 @@ A consumer may load the following info for a commit from the graph:\n>  \n>  Values 1-4 satisfy the requirements of parse_commit_gently().\n>  \n> -Define the \"generation number\" of a commit recursively as follows:\n> +There are two definitions of generation number:\n> +1. Corrected committer dates (generation number v2)\n> +2. Topological levels (generation nummber v1)\n>  \n> - * A commit with no parents (a root commit) has generation number one.\n> +Define \"corrected committer date\" of a commit recursively as follows:\n>  \n> - * A commit with at least one parent has generation number one more than\n> -   the largest generation number among its parents.\n> + * A commit with no parents (a root commit) has corrected committer date\n> +    equal to its committer date.\n>  \n> -Equivalently, the generation number of a commit A is one more than the\n> + * A commit with at least one parent has corrected committer date equal to\n> +    the maximum of its commiter date and one more than the largest corrected\n> +    committer date among its parents.\n> +\n> + * As a special case, a root commit with timestamp zero has corrected commit\n> +    date of 1, to be able to distinguish it from GENERATION_NUMBER_ZERO\n> +    (that is, an uncomputed corrected commit date).\n> +\n> +Define the \"topological level\" of a commit recursively as follows:\n> +\n> + * A commit with no parents (a root commit) has topological level of one.\n> +\n> + * A commit with at least one parent has topological level one more than\n> +   the largest topological level among its parents.\n> +\n> +Equivalently, the topological level of a commit A is one more than the\n>  length of a longest path from A to a root commit. The recursive definition\n>  is easier to use for computation and observing the following property:\n>  \n> @@ -60,6 +77,9 @@ is easier to use for computation and observing the following property:\n>      generation numbers, then we always expand the boundary commit with highest\n>      generation number and can easily detect the stopping condition.\n>  \n> +The property applies to both versions of generation number, that is both\n> +corrected committer dates and topological levels.\n> +\n>  This property can be used to significantly reduce the time it takes to\n>  walk commits and determine topological relationships. Without generation\n>  numbers, the general heuristic is the following:\n> @@ -67,7 +87,9 @@ numbers, the general heuristic is the following:\n>      If A and B are commits with commit time X and Y, respectively, and\n>      X < Y, then A _probably_ cannot reach B.\n>  \n> -This heuristic is currently used whenever the computation is allowed to\n> +In absence of corrected commit dates (for example, old versions of Git or\n> +mixed generation graph chains),\n> +this heuristic is currently used whenever the computation is allowed to\n>  violate topological relationships due to clock skew (such as \"git log\"\n>  with default order), but is not used when the topological order is\n>  required (such as merge base calculations, \"git log --graph\").\n> @@ -77,7 +99,7 @@ in the commit graph. We can treat these commits as having \"infinite\"\n>  generation number and walk until reaching commits with known generation\n>  number.\n>  \n> -We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n> +We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n>  in the commit-graph file. If a commit-graph file was written by a version\n>  of Git that did not compute generation numbers, then those commits will\n>  have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n> @@ -93,12 +115,12 @@ fully-computed generation numbers. Using strict inequality may result in\n>  walking a few extra commits, but the simplicity in dealing with commits\n>  with generation number *_INFINITY or *_ZERO is valuable.\n>  \n> -We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n> -generation numbers are computed to be at least this value. We limit at\n> -this value since it is the largest value that can be stored in the\n> -commit-graph file using the 30 bits available to generation numbers. This\n> -presents another case where a commit can have generation number equal to\n> -that of a parent.\n> +We use the macro GENERATION_NUMBER_V1_MAX = 0x3FFFFFFF for commits whose\n> +topological levels (generation number v1) are computed to be at least\n> +this value. We limit at this value since it is the largest value that\n> +can be stored in the commit-graph file using the 30 bits available\n> +to topological levels. This presents another case where a commit can\n> +have generation number equal to that of a parent.\n>  \n>  Design Details\n>  --------------\n> @@ -267,6 +289,35 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n>  number of commits) could be extracted into config settings for full\n>  flexibility.\n>  \n> +## Handling Mixed Generation Number Chains\n> +\n> +With the introduction of generation number v2 and generation data chunk, the\n> +following scenario is possible:\n> +\n> +1. \"New\" Git writes a commit-graph with the corrected commit dates.\n> +2. \"Old\" Git writes a split commit-graph on top without corrected commit dates.\n> +\n> +A naive approach of using the newest available generation number from\n> +each layer would lead to violated expectations: the lower layer would\n> +use corrected commit dates which are much larger than the topological\n> +levels of the higher layer. For this reason, Git inspects the topmost\n> +layer to see if the layer is missing corrected commit dates. In such a case\n> +Git only uses topological level for generation numbers.\n> +\n> +When writing a new layer in split commit-graph, we write corrected commit\n> +dates if the topmost layer has corrected commit dates written. This\n> +guarantees that if a layer has corrected commit dates, all lower layers\n> +must have corrected commit dates as well.\n> +\n> +When merging layers, we do not consider whether the merged layers had corrected\n> +commit dates. Instead, the new layer will have corrected commit dates if the\n> +layer below the new layer has corrected commit dates.\n> +\n> +While writing or merging layers, if the new layer is the only layer, it will\n> +have corrected commit dates when written by compatible versions of Git. Thus,\n> +rewriting split commit-graph as a single file (`--split=replace`) creates a\n> +single layer with corrected commit dates.\n> +\n>  ## Deleting graph-{hash} files\n>  \n>  After a new tip file is written, some `graph-{hash}` files may no longer\n> -- \n> gitgitgadget\n"},{"id":"415631","messageId":"YBTuoTsrnbzLtX0j@Abhishek-Arch","threadId":"53933","inReplyTo":"20210127000454.GA1440011@szeder.dev","subject":"Re: [PATCH v6 11/11] doc: add corrected commit date info","fromName":"Abhishek Kumar","fromEmail":"abhishekkumar8222@gmail.com","sentAt":"2021-01-30T05:29:05Z","receivedAt":"2021-01-30T05:33:29Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"On Wed, Jan 27, 2021 at 01:04:54AM +0100, SZEDER Gábor wrote:\n> On Sat, Jan 16, 2021 at 06:11:18PM +0000, Abhishek Kumar via GitGitGadget wrote:\n> > With generation data chunk and corrected commit dates implemented, let's\n> > update the technical documentation for commit-graph.\n> \n> This patch should come much earlier in this series, before patch 07/11\n> (commit-graph: implement corrected commit date), or perhaps even\n> earlier.  That way if someone were to investigate an issue in this\n> series and checks out one of its commits, then the specification and\n> the will be right there under 'Documentation/technical/'.\n> \n> Furthermore, a patch introducing a new chunk format is the right place\n> to justify the introduction of said new chunk.  What problems does a\n> chunk of corrected commit dates solve?  Why does it solve them?  Why\n> do we need corrected commit dates instead of simple commit dates?\n> What alternatives were considered [1]?  Any other design considerations\n> worth mentioning for the benefit of future readers?\n> \n> None of the patches' log messages properly explain these, and while\n> much of these is indeed explained in the cover letter, the cover\n> letter will not be part of the history.  Requiring to look up mailing\n> list archives for the justification puts unnecessary burden on other\n> developers who might get interested in this feature in the future.\n> \n> You might want to take\n> https://public-inbox.org/git/20200529085038.26008-16-szeder.dev@gmail.com/\n> as an inspiration.\n> \n\nAlright, the suggestion makes a lot of sense and the patch introducing\ndocumentation is the perfect place to justify the introduction of new\nchunk format.\n\n> \n> [1] Please remember the following snippet from SubmittingPatches:\n>     \"Try to make sure your explanation can be understood without\n>     external resources. Instead of giving a URL to a mailing list\n>     archive, summarize the relevant points of the discussion.\"\n> \n> > Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n> > ---\n> >  .../technical/commit-graph-format.txt         | 28 +++++--\n> >  Documentation/technical/commit-graph.txt      | 77 +++++++++++++++----\n> >  2 files changed, 86 insertions(+), 19 deletions(-)\n> > \n> > diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n> > index b3b58880b92..b6658eff188 100644\n> > --- a/Documentation/technical/commit-graph-format.txt\n> > +++ b/Documentation/technical/commit-graph-format.txt\n> > @@ -4,11 +4,7 @@ Git commit graph format\n> >  The Git commit graph stores a list of commit OIDs and some associated\n> >  metadata, including:\n> >  \n> > -- The generation number of the commit. Commits with no parents have\n> > -  generation number 1; commits with parents have generation number\n> > -  one more than the maximum generation number of its parents. We\n> > -  reserve zero as special, and can be used to mark a generation\n> > -  number invalid or as \"not computed\".\n> > +- The generation number of the commit.\n> >  \n> >  - The root tree OID.\n> >  \n> > @@ -86,13 +82,33 @@ CHUNK DATA:\n> >        position. If there are more than two parents, the second value\n> >        has its most-significant bit on and the other bits store an array\n> >        position into the Extra Edge List chunk.\n> > -    * The next 8 bytes store the generation number of the commit and\n> > +    * The next 8 bytes store the topological level (generation number v1)\n> > +      of the commit and\n> >        the commit time in seconds since EPOCH. The generation number\n> >        uses the higher 30 bits of the first 4 bytes, while the commit\n> >        time uses the 32 bits of the second 4 bytes, along with the lowest\n> >        2 bits of the lowest byte, storing the 33rd and 34th bit of the\n> >        commit time.\n> >  \n> > +  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes) [Optional]\n> > +    * This list of 4-byte values store corrected commit date offsets for the\n> > +      commits, arranged in the same order as commit data chunk.\n> > +    * If the corrected commit date offset cannot be stored within 31 bits,\n> > +      the value has its most-significant bit on and the other bits store\n> > +      the position of corrected commit date into the Generation Data Overflow\n> > +      chunk.\n> > +    * Generation Data chunk is present only when commit-graph file is written\n> > +      by compatible versions of Git and in case of split commit-graph chains,\n> > +      the topmost layer also has Generation Data chunk.\n> > +\n> > +  Generation Data Overflow (ID: {'G', 'D', 'O', 'V' }) [Optional]\n> > +    * This list of 8-byte values stores the corrected commit date offsets\n> > +      for commits with corrected commit date offsets that cannot be\n> > +      stored within 31 bits.\n> > +    * Generation Data Overflow chunk is present only when Generation Data\n> > +      chunk is present and atleast one corrected commit date offset cannot\n> > +      be stored within 31 bits.\n> > +\n> >    Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n> >        This list of 4-byte values store the second through nth parents for\n> >        all octopus merges. The second parent value in the commit data stores\n> > diff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\n> > index f14a7659aa8..f05e7bda1a9 100644\n> > --- a/Documentation/technical/commit-graph.txt\n> > +++ b/Documentation/technical/commit-graph.txt\n> > @@ -38,14 +38,31 @@ A consumer may load the following info for a commit from the graph:\n> >  \n> >  Values 1-4 satisfy the requirements of parse_commit_gently().\n> >  \n> > -Define the \"generation number\" of a commit recursively as follows:\n> > +There are two definitions of generation number:\n> > +1. Corrected committer dates (generation number v2)\n> > +2. Topological levels (generation nummber v1)\n> >  \n> > - * A commit with no parents (a root commit) has generation number one.\n> > +Define \"corrected committer date\" of a commit recursively as follows:\n> >  \n> > - * A commit with at least one parent has generation number one more than\n> > -   the largest generation number among its parents.\n> > + * A commit with no parents (a root commit) has corrected committer date\n> > +    equal to its committer date.\n> >  \n> > -Equivalently, the generation number of a commit A is one more than the\n> > + * A commit with at least one parent has corrected committer date equal to\n> > +    the maximum of its commiter date and one more than the largest corrected\n> > +    committer date among its parents.\n> > +\n> > + * As a special case, a root commit with timestamp zero has corrected commit\n> > +    date of 1, to be able to distinguish it from GENERATION_NUMBER_ZERO\n> > +    (that is, an uncomputed corrected commit date).\n> > +\n> > +Define the \"topological level\" of a commit recursively as follows:\n> > +\n> > + * A commit with no parents (a root commit) has topological level of one.\n> > +\n> > + * A commit with at least one parent has topological level one more than\n> > +   the largest topological level among its parents.\n> > +\n> > +Equivalently, the topological level of a commit A is one more than the\n> >  length of a longest path from A to a root commit. The recursive definition\n> >  is easier to use for computation and observing the following property:\n> >  \n> > @@ -60,6 +77,9 @@ is easier to use for computation and observing the following property:\n> >      generation numbers, then we always expand the boundary commit with highest\n> >      generation number and can easily detect the stopping condition.\n> >  \n> > +The property applies to both versions of generation number, that is both\n> > +corrected committer dates and topological levels.\n> > +\n> >  This property can be used to significantly reduce the time it takes to\n> >  walk commits and determine topological relationships. Without generation\n> >  numbers, the general heuristic is the following:\n> > @@ -67,7 +87,9 @@ numbers, the general heuristic is the following:\n> >      If A and B are commits with commit time X and Y, respectively, and\n> >      X < Y, then A _probably_ cannot reach B.\n> >  \n> > -This heuristic is currently used whenever the computation is allowed to\n> > +In absence of corrected commit dates (for example, old versions of Git or\n> > +mixed generation graph chains),\n> > +this heuristic is currently used whenever the computation is allowed to\n> >  violate topological relationships due to clock skew (such as \"git log\"\n> >  with default order), but is not used when the topological order is\n> >  required (such as merge base calculations, \"git log --graph\").\n> > @@ -77,7 +99,7 @@ in the commit graph. We can treat these commits as having \"infinite\"\n> >  generation number and walk until reaching commits with known generation\n> >  number.\n> >  \n> > -We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n> > +We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n> >  in the commit-graph file. If a commit-graph file was written by a version\n> >  of Git that did not compute generation numbers, then those commits will\n> >  have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n> > @@ -93,12 +115,12 @@ fully-computed generation numbers. Using strict inequality may result in\n> >  walking a few extra commits, but the simplicity in dealing with commits\n> >  with generation number *_INFINITY or *_ZERO is valuable.\n> >  \n> > -We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n> > -generation numbers are computed to be at least this value. We limit at\n> > -this value since it is the largest value that can be stored in the\n> > -commit-graph file using the 30 bits available to generation numbers. This\n> > -presents another case where a commit can have generation number equal to\n> > -that of a parent.\n> > +We use the macro GENERATION_NUMBER_V1_MAX = 0x3FFFFFFF for commits whose\n> > +topological levels (generation number v1) are computed to be at least\n> > +this value. We limit at this value since it is the largest value that\n> > +can be stored in the commit-graph file using the 30 bits available\n> > +to topological levels. This presents another case where a commit can\n> > +have generation number equal to that of a parent.\n> >  \n> >  Design Details\n> >  --------------\n> > @@ -267,6 +289,35 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n> >  number of commits) could be extracted into config settings for full\n> >  flexibility.\n> >  \n> > +## Handling Mixed Generation Number Chains\n> > +\n> > +With the introduction of generation number v2 and generation data chunk, the\n> > +following scenario is possible:\n> > +\n> > +1. \"New\" Git writes a commit-graph with the corrected commit dates.\n> > +2. \"Old\" Git writes a split commit-graph on top without corrected commit dates.\n> > +\n> > +A naive approach of using the newest available generation number from\n> > +each layer would lead to violated expectations: the lower layer would\n> > +use corrected commit dates which are much larger than the topological\n> > +levels of the higher layer. For this reason, Git inspects the topmost\n> > +layer to see if the layer is missing corrected commit dates. In such a case\n> > +Git only uses topological level for generation numbers.\n> > +\n> > +When writing a new layer in split commit-graph, we write corrected commit\n> > +dates if the topmost layer has corrected commit dates written. This\n> > +guarantees that if a layer has corrected commit dates, all lower layers\n> > +must have corrected commit dates as well.\n> > +\n> > +When merging layers, we do not consider whether the merged layers had corrected\n> > +commit dates. Instead, the new layer will have corrected commit dates if the\n> > +layer below the new layer has corrected commit dates.\n> > +\n> > +While writing or merging layers, if the new layer is the only layer, it will\n> > +have corrected commit dates when written by compatible versions of Git. Thus,\n> > +rewriting split commit-graph as a single file (`--split=replace`) creates a\n> > +single layer with corrected commit dates.\n> > +\n> >  ## Deleting graph-{hash} files\n> >  \n> >  After a new tip file is written, some `graph-{hash}` files may no longer\n> > -- \n> > gitgitgadget\n\nThanks\n- Abhishek\n"},{"id":"415674","messageId":"YBYLwpKdUfxCNwaz@nand.local","threadId":"53933","inReplyTo":"YBTuoTsrnbzLtX0j@Abhishek-Arch","subject":"Re: [PATCH v6 11/11] doc: add corrected commit date info","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-01-31T01:45:38Z","receivedAt":"2021-01-31T01:46:26Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Sat, Jan 30, 2021 at 10:59:05AM +0530, Abhishek Kumar wrote:\n> > You might want to take\n> > https://public-inbox.org/git/20200529085038.26008-16-szeder.dev@gmail.com/\n> > as an inspiration.\n> >\n> Alright, the suggestion makes a lot of sense and the patch introducing\n> documentation is the perfect place to justify the introduction of new\n> chunk format.\n\nI don't have any strong feelings about Gábor's suggestion itself, but\nnote that there isn't any work for you to do in this series, since the\npatches are on track to be merged to master.\n\nThanks,\nTaylor\n"},{"id":"415711","messageId":"9ac331b63ee609f5380649d3b395f420e57e56f8.1612162726.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"[PATCH v7 01/11] commit-graph: fix regression when computing Bloom filters","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:35Z","receivedAt":"2021-02-01T06:59:35Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nBefore computing Bloom filters, the commit-graph machinery uses\ncommit_gen_cmp to sort commits by generation order for improved diff\nperformance. 3d11275505 (commit-graph: examine commits by generation\nnumber, 2020-03-30) claims that this sort can reduce the time spent to\ncompute Bloom filters by nearly half.\n\nBut since c49c82aa4c (commit: move members graph_pos, generation to a\nslab, 2020-06-17), this optimization is broken, since asking for a\n'commit_graph_generation()' directly returns GENERATION_NUMBER_INFINITY\nwhile writing.\n\nNot all hope is lost, though: 'commit_gen_cmp()' falls back to\ncomparing commits by their date when they have equal generation number,\nand so since c49c82aa4c is purely a date comparison function. This\nheuristic is good enough that we don't seem to loose appreciable\nperformance while computing Bloom filters.\n\nApplying this patch (compared with v2.30.0) speeds up computing Bloom\nfilters by factors ranging from 0.40% to 5.19% on various repositories [1].\n\nSo, avoid the useless 'commit_graph_generation()' while writing by\ninstead accessing the slab directly. This returns the newly-computed\ngeneration numbers, and allows us to avoid the heuristic by directly\ncomparing generation numbers.\n\n[1]: https://lore.kernel.org/git/20210105094535.GN8396@szeder.dev/\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 8 ++++++--\n 1 file changed, 6 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex f3486ec18f1..78de312ccec 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -139,13 +139,17 @@ static struct commit_graph_data *commit_graph_data_at(const struct commit *c)\n \treturn data;\n }\n \n+/* \n+ * Should be used only while writing commit-graph as it compares\n+ * generation value of commits by directly accessing commit-slab.\n+ */\n static int commit_gen_cmp(const void *va, const void *vb)\n {\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n+\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n+\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"415712","messageId":"90ca0a1fd697b91180a01d82eeaed54511c9719f.1612162726.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"[PATCH v7 02/11] revision: parse parent in indegree_walk_step()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:36Z","receivedAt":"2021-02-01T06:59:38Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn indegree_walk_step(), we add unvisited parents to the indegree queue.\nHowever, parents are not guaranteed to be parsed. As the indegree queue\nsorts by generation number, let's parse parents before inserting them to\nensure the correct priority order.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n revision.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/revision.c b/revision.c\nindex 0b5c7231401..5474001331a 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3399,6 +3399,9 @@ static void indegree_walk_step(struct rev_info *revs)\n \t\tstruct commit *parent = p->item;\n \t\tint *pi = indegree_slab_at(&info->indegree, parent);\n \n+\t\tif (repo_parse_commit_gently(revs->repo, parent, 1) < 0)\n+\t\t\treturn;\n+\n \t\tif (*pi)\n \t\t\t(*pi)++;\n \t\telse\n-- \ngitgitgadget\n\n"},{"id":"415713","messageId":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v6.git.1610820679.gitgitgadget@gmail.com","subject":"[PATCH v7 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:34Z","receivedAt":"2021-02-01T06:59:38Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"This patch series implements the corrected commit date offsets as generation\nnumber v2, along with other pre-requisites.\n\nGit uses topological levels in the commit-graph file for commit-graph\ntraversal operations like 'git log --graph'. Unfortunately, topological\nlevels can perform worse than committer date when parents of a commit differ\ngreatly in generation numbers [1]. For example, 'git merge-base v4.8 v4.9'\non the Linux repository walks 635,579 commits using topological levels and\nwalks 167,468 using committer date. Since 091f4cf3 (commit: don't use\ngeneration numbers if not needed, 2018-08-30), 'git merge-base' uses\ncommitter date heuristic unless there is a cutoff because of the performance\nhit.\n\n[1]\nhttps://lore.kernel.org/git/efa3720fb40638e5d61c6130b55e3348d8e4339e.1535633886.git.gitgitgadget@gmail.com/\n\nThus, the need for generation number v2 was born. As Git used to die when\ngraph version understood by it and in the commit-graph file are different\n[2], we needed a way to distinguish between the old and new generation\nnumber without incrementing the graph version.\n\n[2] https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nThe following candidates were proposed\n(https://github.com/derrickstolee/gen-test,\nhttps://github.com/abhishekkumar2718/git/pull/1):\n\n * (Epoch, Date) Pairs.\n * Maximum Generation Numbers.\n * Corrected Commit Date.\n * FELINE Index.\n * Corrected Commit Date with Monotonically Increasing Offsets.\n\nBased on performance, local computability, and immutability (along with the\nintroduction of an additional commit-graph chunk which relieved the\nrequirement of backwards-compatibility) Corrected Commit Date was chosen as\ngeneration number v2 and is defined as follows:\n\nFor a commit C, let its corrected commit date be the maximum of the commit\ndate of C and the corrected commit dates of its parents plus 1. Then\ncorrected commit date offset is the difference between corrected commit date\nof C and commit date of C. As a special case, a root commit with the\ntimestamp zero has corrected commit date of 1 to distinguish it from\nGENERATION_NUMBER_ZERO (that is, an uncomputed generation number).\n\nWhile it was proposed initially to store corrected commit date offsets\nwithin Commit Data Chunk, storing the offsets in a new chunk did not affect\nthe performance measurably. The new chunk is \"Generation DATa (GDAT) chunk\"\nand it stores corrected commit date offsets while CDAT chunk stores\ntopological level. The old versions of Git would ignore GDAT chunk, using\ntopological levels from CDAT chunk. In contrast, new versions of Git would\nuse corrected commit dates, falling back to topological level if the\ngeneration data chunk is absent in the commit-graph file.\n\nWhile storing corrected commit date offsets saves us 4 bytes per commit (as\ncompared with storing corrected commit dates directly), it's however\npossible for the offset to overflow the space allocated. To handle such\ncases, we introduce a new chunk, Generation Data Overflow (GDOV) that stores\nthe corrected commit date. For overflowing offsets, we set MSB and store the\nposition into the GDOV chunk, in a mechanism similar to the Extra Edges list\nchunk.\n\nFor mixed generation number environment (for example new Git on the command\nline, old Git used by GUI client), we can encounter a mixed-chain\ncommit-graph (a commit-graph chain where some of split commit-graph files\nhave GDAT chunk and others do not). As backward compatibility is one of the\ngoals, we can define the following behavior:\n\nWhile reading a mixed-chain commit-graph version, we fall back on\ntopological levels as corrected commit dates and topological levels cannot\nbe compared directly.\n\nWhen adding new layer to the split commit-graph file, and when merging some\nor all layers (replacing them in the latter case), the new layer will have\nGDAT chunk if and only if in the final result there would be no layer\nwithout GDAT chunk just below it.\n\nThanks to Dr. Stolee, Dr. Narębski, Taylor Blau and SZEDER Gábor for their\nreviews.\n\nI look forward to everyone's reviews!\n\nThanks\n\n * Abhishek\n\n----------------------------------------------------------------------------\n\nImprovements left for a future series:\n\n * Save commits with generation data overflow and extra edge commits instead\n   of looping over all commits. cf. 858sbel67n.fsf@gmail.com\n * Verify both topological levels and corrected commit dates when present.\n   cf. 85pn4tnk8u.fsf@gmail.com\n\nChanges in version 7:\n\n * Moved the documentation patch ahead of \"commit-graph: implement corrected\n   commit date\" and elaborated on the introduction of generation number v2.\n\nChanges in version 6:\n\n * Fixed typos in commit message for \"commit-graph: implement corrected\n   commit date\".\n * Removed an unnecessary else-block in \"commit-graph: implement corrected\n   commit date\".\n * Validate mixed generation chain correctly while writing in \"commit-graph:\n   use generation v2 only if the entire chain does\".\n * Die if the GDAT chunk indicates data has overflown but there are is no\n   generation data overflow chunk.\n\nChanges in version 5:\n\n * Explained a possible reason for no change in performance for\n   \"commit-graph: fix regression when computing bloom-filters\"\n * Clarified about the addition of a new test for 11-digit octal\n   implementations of ustar.\n * Fixed duplicate test names in \"commit-graph: consolidate\n   fill_commit_graph_info\".\n * Swapped the order \"commit-graph: return 64-bit generation number\",\n   \"commit-graph: add a slab to store topological levels\" to minimize lines\n   changed.\n * Fixed the mismerge in \"commit-graph: return 64-bit generation number\"\n * Clarified the preparatory steps are for the larger goal of implementing\n   generation number v2 in \"commit-graph: return 64-bit generation number\".\n * Moved the rename of \"run_three_modes()\" to \"run_all_modes()\" into a new\n   patch \"t6600-test-reach: generalize *_three_modes\".\n * Explained and removed the checks for GENERATION_NUMBER_INFINITY that can\n   never be true in \"commit-graph: add a slab to store topological levels\".\n * Fixed incorrect logic for verifying commit-graph in \"commit-graph:\n   implement corrected commit date\".\n * Added minor improvements to commit message of \"commit-graph: implement\n   generation data chunk\".\n * Added '--date ' option to test_commit() in 'test-lib-functions.sh' in\n   \"commit-graph: implement generation data chunk\".\n * Improved coding style (also in tests) for \"commit-graph: use generation\n   v2 only if entire chain does\".\n * Simplified test repository structure in \"commit-graph: use generation v2\n   only if entire chain does\" as only the number of commits in a split\n   commit-graph layer are relevant.\n * Added a new test in \"commit-graph: use generation v2 only if entire chain\n   does\" to check if the layers are merged correctly.\n * Explicitly mentioned commit \"091f4cf3\" in the commit-message of\n   \"commit-graph: use corrected commit dates in paint_down_to_common()\".\n * Minor corrections to documentation in \"doc: add corrected commit date\n   info\".\n * Minor corrections to coding style.\n\nChanges in version 4:\n\n * Added GDOV to handle overflows in generation data.\n * Added a test for writing tip graph for a generation number v2 graph chain\n   in t5324-split-commit-graph.sh\n * Added a section on how mixed generation number chains are handled in\n   Documentation/technical/commit-graph-format.txt\n * Reverted unimportant whitespace, style changes in commit-graph.c\n * Added header comments about the order of comparision for\n   compare_commits_by_gen_then_commit_date in commit.h,\n   compare_commits_by_gen in commit-graph.h\n * Elaborated on why t6404 fails with corrected commit date and must be run\n   with GIT_TEST_COMMIT_GRAPH=1in the commit \"commit-reach: use corrected\n   commit dates in paint_down_to_common()\"\n * Elaborated on write behavior for mixed generation number chains in the\n   commit \"commit-graph: use generation v2 only if entire chain does\"\n * Added notes about adding the topo_level slab to struct\n   write_commit_graph_context as well as struct commit_graph.\n * Clarified commit message for \"commit-graph: consolidate\n   fill_commit_graph_info\"\n * Removed the claim \"GDAT can store future generation numbers\" because it\n   hasn't been tested yet.\n\nChanges in version 3:\n\n * Reordered patches to implement corrected commit date before generation\n   data chunk [3].\n * Split \"implement corrected commit date\" into two patches - one\n   introducing the topo level slab and other implementing corrected commit\n   dates.\n * Extended split-commit-graph tests to verify at the end of test.\n * Use topological levels as generation number if any of split commit-graph\n   files do not have generation data chunk.\n\n[3]\nhttps://lore.kernel.org/git/aee0ae56-3395-6848-d573-27a318d72755@gmail.com/\n\nChanges in version 2:\n\n * Add tests for generation data chunk.\n * Add an option GIT_TEST_COMMIT_GRAPH_NO_GDAT to control whether to write\n   generation data chunk.\n * Compare commits with corrected commit dates if present in\n   paint_down_to_common().\n * Update technical documentation.\n * Handle mixed generation commit chains.\n * Improve commit messages for \"commit-graph: fix regression when computing\n   bloom filter\", \"commit-graph: consolidate fill_commit_graph_info\",\n * Revert unnecessary whitespace changes.\n * Split uint_32 -> timestamp_t change into a new commit.\n\nAbhishek Kumar (11):\n  commit-graph: fix regression when computing Bloom filters\n  revision: parse parent in indegree_walk_step()\n  commit-graph: consolidate fill_commit_graph_info\n  t6600-test-reach: generalize *_three_modes\n  commit-graph: add a slab to store topological levels\n  commit-graph: return 64-bit generation number\n  commit-graph: document generation number v2\n  commit-graph: implement corrected commit date\n  commit-graph: implement generation data chunk\n  commit-graph: use generation v2 only if entire chain does\n  commit-reach: use corrected commit dates in paint_down_to_common()\n\n .../technical/commit-graph-format.txt         |  28 +-\n Documentation/technical/commit-graph.txt      |  77 +++++-\n commit-graph.c                                | 251 ++++++++++++++----\n commit-graph.h                                |  15 +-\n commit-reach.c                                |  38 +--\n commit-reach.h                                |   2 +-\n commit.c                                      |   4 +-\n commit.h                                      |   5 +-\n revision.c                                    |  13 +-\n t/README                                      |   3 +\n t/helper/test-read-graph.c                    |   4 +\n t/t4216-log-bloom.sh                          |   4 +-\n t/t5000-tar-tree.sh                           |  24 +-\n t/t5318-commit-graph.sh                       |  79 +++++-\n t/t5324-split-commit-graph.sh                 | 193 +++++++++++++-\n t/t6404-recursive-merge.sh                    |   5 +-\n t/t6600-test-reach.sh                         |  68 ++---\n t/test-lib-functions.sh                       |   6 +\n upload-pack.c                                 |   2 +-\n 19 files changed, 667 insertions(+), 154 deletions(-)\n\n\nbase-commit: e6362826a0409539642a5738db61827e5978e2e4\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-676%2Fabhishekkumar2718%2Fcorrected_commit_date-v7\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-676/abhishekkumar2718/corrected_commit_date-v7\nPull-Request: https://github.com/gitgitgadget/git/pull/676\n\nRange-diff vs v6:\n\n  1:  4d8eb415578 =  1:  9ac331b63ee commit-graph: fix regression when computing Bloom filters\n  2:  05dcb862818 =  2:  90ca0a1fd69 revision: parse parent in indegree_walk_step()\n  3:  dcb9891d819 =  3:  b3040696d43 commit-graph: consolidate fill_commit_graph_info\n  4:  4fbdee7ac90 =  4:  085085a4330 t6600-test-reach: generalize *_three_modes\n  5:  fbd8feb5d8c =  5:  3b1aae4106a commit-graph: add a slab to store topological levels\n  6:  855ff662a44 =  6:  ea32cba16ef commit-graph: return 64-bit generation number\n 11:  e571f03d8bd !  7:  8647b5d2e38 doc: add corrected commit date info\n     @@ Metadata\n      Author: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n       ## Commit message ##\n     -    doc: add corrected commit date info\n     +    commit-graph: document generation number v2\n      \n     -    With generation data chunk and corrected commit dates implemented, let's\n     -    update the technical documentation for commit-graph.\n     +    Git uses topological levels in the commit-graph file for commit-graph\n     +    traversal operations like 'git log --graph'. Unfortunately, topological\n     +    levels can perform worse than committer date when parents of a commit\n     +    differ greatly in generation numbers [1]. For example, 'git merge-base\n     +    v4.8 v4.9' on the Linux repository walks 635,579 commits using\n     +    topological levels and walks 167,468 using committer date. Since\n     +    091f4cf3 (commit: don't use generation numbers if not needed,\n     +    2018-08-30), 'git merge-base' uses committer date heuristic unless there\n     +    is a cutoff because of the performance hit.\n     +\n     +    [1] https://lore.kernel.org/git/efa3720fb40638e5d61c6130b55e3348d8e4339e.1535633886.git.gitgitgadget@gmail.com/\n     +\n     +    Thus, the need for generation number v2 was born. As Git used to die\n     +    when graph version understood by it and in the commit-graph file are\n     +    different [2], we needed a way to distinguish between the old and new\n     +    generation number without incrementing the graph version.\n     +\n     +    [2] https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n     +\n     +    The following candidates were proposed (https://github.com/derrickstolee/gen-test,\n     +    https://github.com/abhishekkumar2718/git/pull/1):\n     +    - (Epoch, Date) Pairs.\n     +    - Maximum Generation Numbers.\n     +    - Corrected Commit Date.\n     +    - FELINE Index.\n     +    - Corrected Commit Date with Monotonically Increasing Offsets.\n     +\n     +    Based on performance, local computability, and immutability (along with\n     +    the introduction of an additional commit-graph chunk which relieved the\n     +    requirement of backwards-compatibility) Corrected Commit Date was chosen\n     +    as generation number v2 and is defined as follows:\n     +\n     +    For a commit C, let its corrected commit date  be the maximum of the\n     +    commit date of C and the corrected commit dates of its parents plus 1.\n     +    Then corrected commit date offset is the difference between corrected\n     +    commit date of C and commit date of C. As a special case, a root commit\n     +    with the timestamp zero has corrected commit date of 1 to distinguish it\n     +    from GENERATION_NUMBER_ZERO (that is, an uncomputed generation number).\n     +\n     +    While it was proposed initially to store corrected commit date offsets\n     +    within Commit Data Chunk, storing the offsets in a new chunk did not\n     +    affect the performance measurably. The new chunk is \"Generation DATa\n     +    (GDAT) chunk\" and it stores corrected commit date offsets while CDAT\n     +    chunk stores topological level. The old versions of Git would ignore\n     +    GDAT chunk, using topological levels from CDAT chunk. In contrast, new\n     +    versions of Git would use corrected commit dates, falling back to\n     +    topological level if the generation data chunk is absent in the\n     +    commit-graph file.\n     +\n     +    While storing corrected commit date offsets saves us 4 bytes per commit\n     +    (as compared with storing corrected commit dates directly), it's however\n     +    possible for the offset to overflow the space allocated. To handle such\n     +    cases, we introduce a new chunk, _Generation Data Overflow_ (GDOV) that\n     +    stores the corrected commit date. For overflowing offsets, we set MSB\n     +    and store the position into the GDOV chunk, in a mechanism similar to\n     +    the Extra Edges list chunk.\n     +\n     +    For mixed generation number environment (for example new Git on the\n     +    command line, old Git used by GUI client), we can encounter a\n     +    mixed-chain commit-graph (a commit-graph chain where some of split\n     +    commit-graph files have GDAT chunk and others do not). As backward\n     +    compatibility is one of the goals, we can define the following behavior:\n     +\n     +    While reading a mixed-chain commit-graph version, we fall back on\n     +    topological levels as corrected commit dates and topological levels\n     +    cannot be compared directly.\n     +\n     +    When adding new layer to the split commit-graph file, and when merging\n     +    some or all layers (replacing them in the latter case), the new layer\n     +    will have GDAT chunk if and only if in the final result there would be\n     +    no layer without GDAT chunk just below it.\n      \n          Signed-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n      \n  7:  8fbe7486405 =  8:  ec598f1d500 commit-graph: implement corrected commit date\n  8:  6d0696ae216 =  9:  71d81518857 commit-graph: implement generation data chunk\n  9:  fba0d7f3dfe = 10:  07a88f1aae6 commit-graph: use generation v2 only if entire chain does\n 10:  ba1f2c5555f = 11:  523e2d4a902 commit-reach: use corrected commit dates in paint_down_to_common()\n\n-- \ngitgitgadget\n"},{"id":"415714","messageId":"b3040696d43f1abb6d7a50590b4e181cd2eb74aa.1612162726.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"[PATCH v7 03/11] commit-graph: consolidate fill_commit_graph_info","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:37Z","receivedAt":"2021-02-01T06:59:41Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nBoth fill_commit_graph_info() and fill_commit_in_graph() parse\ninformation present in commit data chunk. Let's simplify the\nimplementation by calling fill_commit_graph_info() within\nfill_commit_in_graph().\n\nfill_commit_graph_info() used to not load committer data from commit data\nchunk. However, with the upcoming switch to using corrected committer\ndate as generation number v2, we will have to load committer date to\ncompute generation number value anyway.\n\ne51217e15 (t5000: test tar files that overflow ustar headers,\n30-06-2016) introduced a test 'generate tar with future mtime' that\ncreates a commit with committer date of (2^36 + 1) seconds since\nEPOCH. The CDAT chunk provides 34-bits for storing committer date, thus\ncommitter time overflows into generation number (within CDAT chunk) and\nhas undefined behavior.\n\nThe test used to pass as fill_commit_graph_info() would not set struct\nmember `date` of struct commit and load committer date from the object\ndatabase, generating a tar file with the expected mtime.\n\nHowever, with corrected commit date, we will load the committer date\nfrom CDAT chunk (truncated to lower 34-bits to populate the generation\nnumber. Thus, Git sets date and generates tar file with the truncated\nmtime.\n\nThe ustar format (the header format used by most modern tar programs)\nonly has room for 11 (or 12, depending on some implementations) octal\ndigits for the size and mtime of each file.\n\nAs the CDAT chunk is overflow by 12-octal digits but not 11-octal\ndigits, we split the existing tests to test both implementations\nseparately and add a new explicit test for 11-digit implementation.\n\nTo test the 11-octal digit implementation, we create a future commit\nwith committer date of 2^34 - 1, which overflows 11-octal digits without\noverflowing 34-bits of the Commit Date chunks.\n\nTo test the 12-octal digit implementation, the smallest committer date\npossible is 2^36 + 1, which overflows the CDAT chunk and thus\ncommit-graph must be disabled for the test.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c      | 27 ++++++++++-----------------\n t/t5000-tar-tree.sh | 24 +++++++++++++++++++++---\n 2 files changed, 31 insertions(+), 20 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 78de312ccec..955418bd6e5 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -753,15 +753,24 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n \tuint32_t lex_index;\n+\tuint64_t date_high, date_low;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n+\tif (pos >= g->num_commits + g->num_commits_in_base)\n+\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\n \tlex_index = pos - g->num_commits_in_base;\n \tcommit_data = g->chunk_commit_data + GRAPH_DATA_WIDTH * lex_index;\n \n \tgraph_data = commit_graph_data_at(item);\n \tgraph_data->graph_pos = pos;\n+\n+\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n+\tdate_low = get_be32(commit_data + g->hash_len + 12);\n+\titem->date = (timestamp_t)((date_high << 32) | date_low);\n+\n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n@@ -776,38 +785,22 @@ static int fill_commit_in_graph(struct repository *r,\n {\n \tuint32_t edge_value;\n \tuint32_t *parent_data_ptr;\n-\tuint64_t date_low, date_high;\n \tstruct commit_list **pptr;\n-\tstruct commit_graph_data *graph_data;\n \tconst unsigned char *commit_data;\n \tuint32_t lex_index;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n \n-\tif (pos >= g->num_commits + g->num_commits_in_base)\n-\t\tdie(_(\"invalid commit position. commit-graph is likely corrupt\"));\n+\tfill_commit_graph_info(item, g, pos);\n \n-\t/*\n-\t * Store the \"full\" position, but then use the\n-\t * \"local\" position for the rest of the calculation.\n-\t */\n-\tgraph_data = commit_graph_data_at(item);\n-\tgraph_data->graph_pos = pos;\n \tlex_index = pos - g->num_commits_in_base;\n-\n \tcommit_data = g->chunk_commit_data + (g->hash_len + 16) * lex_index;\n \n \titem->object.parsed = 1;\n \n \tset_commit_tree(item, NULL);\n \n-\tdate_high = get_be32(commit_data + g->hash_len + 8) & 0x3;\n-\tdate_low = get_be32(commit_data + g->hash_len + 12);\n-\titem->date = (timestamp_t)((date_high << 32) | date_low);\n-\n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n-\n \tpptr = &item->parents;\n \n \tedge_value = get_be32(commit_data + g->hash_len);\ndiff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\nindex 3ebb0d3b652..7204799a0b5 100755\n--- a/t/t5000-tar-tree.sh\n+++ b/t/t5000-tar-tree.sh\n@@ -431,15 +431,33 @@ test_expect_success TAR_HUGE,LONG_IS_64BIT 'system tar can read our huge size' '\n \ttest_cmp expect actual\n '\n \n-test_expect_success TIME_IS_64BIT 'set up repository with far-future commit' '\n+test_expect_success TIME_IS_64BIT 'set up repository with far-future (2^34 - 1) commit' '\n+\trm -f .git/index &&\n+\techo foo >file &&\n+\tgit add file &&\n+\tGIT_COMMITTER_DATE=\"@17179869183 +0000\" \\\n+\t\tgit commit -m \"tempori parendum\"\n+'\n+\n+test_expect_success TIME_IS_64BIT 'generate tar with far-future mtime' '\n+\tgit archive HEAD >future.tar\n+'\n+\n+test_expect_success TAR_HUGE,TIME_IS_64BIT,TIME_T_IS_64BIT 'system tar can read our future mtime' '\n+\techo 2514 >expect &&\n+\ttar_info future.tar | cut -d\" \" -f2 >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success TIME_IS_64BIT 'set up repository with far-far-future (2^36 + 1) commit' '\n \trm -f .git/index &&\n \techo content >file &&\n \tgit add file &&\n-\tGIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n+\tGIT_TEST_COMMIT_GRAPH=0 GIT_COMMITTER_DATE=\"@68719476737 +0000\" \\\n \t\tgit commit -m \"tempori parendum\"\n '\n \n-test_expect_success TIME_IS_64BIT 'generate tar with future mtime' '\n+test_expect_success TIME_IS_64BIT 'generate tar with far-far-future mtime' '\n \tgit archive HEAD >future.tar\n '\n \n-- \ngitgitgadget\n\n"},{"id":"415715","messageId":"085085a433072076ffa45d149cdf4e0b6b55d918.1612162726.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"[PATCH v7 04/11] t6600-test-reach: generalize *_three_modes","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:38Z","receivedAt":"2021-02-01T06:59:47Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a preparatory step to implement generation number v2, we add tests to\nensure Git can read and parse commit-graph files without Generation Data\nchunk. These files represent commit-graph files written by Old Git and\nare neccesary for backward compatability.\n\nWe extend run_three_modes() and test_three_modes() to *_all_modes() with\nthe fourth mode being \"commit-graph without generation data chunk\".\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n t/t6600-test-reach.sh | 62 +++++++++++++++++++++----------------------\n 1 file changed, 31 insertions(+), 31 deletions(-)\n\ndiff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\nindex f807276337d..af10f0dc090 100755\n--- a/t/t6600-test-reach.sh\n+++ b/t/t6600-test-reach.sh\n@@ -58,7 +58,7 @@ test_expect_success 'setup' '\n \tgit config core.commitGraph true\n '\n \n-run_three_modes () {\n+run_all_modes () {\n \ttest_when_finished rm -rf .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n \ttest_cmp expect actual &&\n@@ -70,8 +70,8 @@ run_three_modes () {\n \ttest_cmp expect actual\n }\n \n-test_three_modes () {\n-\trun_three_modes test-tool reach \"$@\"\n+test_all_modes () {\n+\trun_all_modes test-tool reach \"$@\"\n }\n \n test_expect_success 'ref_newer:miss' '\n@@ -80,7 +80,7 @@ test_expect_success 'ref_newer:miss' '\n \tB:commit-4-9\n \tEOF\n \techo \"ref_newer(A,B):0\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'ref_newer:hit' '\n@@ -89,7 +89,7 @@ test_expect_success 'ref_newer:hit' '\n \tB:commit-2-3\n \tEOF\n \techo \"ref_newer(A,B):1\" >expect &&\n-\ttest_three_modes ref_newer\n+\ttest_all_modes ref_newer\n '\n \n test_expect_success 'in_merge_bases:hit' '\n@@ -98,7 +98,7 @@ test_expect_success 'in_merge_bases:hit' '\n \tB:commit-8-8\n \tEOF\n \techo \"in_merge_bases(A,B):1\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'in_merge_bases:miss' '\n@@ -107,7 +107,7 @@ test_expect_success 'in_merge_bases:miss' '\n \tB:commit-5-9\n \tEOF\n \techo \"in_merge_bases(A,B):0\" >expect &&\n-\ttest_three_modes in_merge_bases\n+\ttest_all_modes in_merge_bases\n '\n \n test_expect_success 'in_merge_bases_many:hit' '\n@@ -117,7 +117,7 @@ test_expect_success 'in_merge_bases_many:hit' '\n \tX:commit-5-7\n \tEOF\n \techo \"in_merge_bases_many(A,X):1\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'in_merge_bases_many:miss' '\n@@ -127,7 +127,7 @@ test_expect_success 'in_merge_bases_many:miss' '\n \tX:commit-8-6\n \tEOF\n \techo \"in_merge_bases_many(A,X):0\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'in_merge_bases_many:miss-heuristic' '\n@@ -137,7 +137,7 @@ test_expect_success 'in_merge_bases_many:miss-heuristic' '\n \tX:commit-6-6\n \tEOF\n \techo \"in_merge_bases_many(A,X):0\" >expect &&\n-\ttest_three_modes in_merge_bases_many\n+\ttest_all_modes in_merge_bases_many\n '\n \n test_expect_success 'is_descendant_of:hit' '\n@@ -148,7 +148,7 @@ test_expect_success 'is_descendant_of:hit' '\n \tX:commit-1-1\n \tEOF\n \techo \"is_descendant_of(A,X):1\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'is_descendant_of:miss' '\n@@ -159,7 +159,7 @@ test_expect_success 'is_descendant_of:miss' '\n \tX:commit-7-6\n \tEOF\n \techo \"is_descendant_of(A,X):0\" >expect &&\n-\ttest_three_modes is_descendant_of\n+\ttest_all_modes is_descendant_of\n '\n \n test_expect_success 'get_merge_bases_many' '\n@@ -174,7 +174,7 @@ test_expect_success 'get_merge_bases_many' '\n \t\tgit rev-parse commit-5-6 \\\n \t\t\t      commit-4-7 | sort\n \t} >expect &&\n-\ttest_three_modes get_merge_bases_many\n+\ttest_all_modes get_merge_bases_many\n '\n \n test_expect_success 'reduce_heads' '\n@@ -196,7 +196,7 @@ test_expect_success 'reduce_heads' '\n \t\t\t      commit-2-8 \\\n \t\t\t      commit-1-10 | sort\n \t} >expect &&\n-\ttest_three_modes reduce_heads\n+\ttest_all_modes reduce_heads\n '\n \n test_expect_success 'can_all_from_reach:hit' '\n@@ -219,7 +219,7 @@ test_expect_success 'can_all_from_reach:hit' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach(X,Y):1\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach:miss' '\n@@ -241,7 +241,7 @@ test_expect_success 'can_all_from_reach:miss' '\n \tY:commit-8-5\n \tEOF\n \techo \"can_all_from_reach(X,Y):0\" >expect &&\n-\ttest_three_modes can_all_from_reach\n+\ttest_all_modes can_all_from_reach\n '\n \n test_expect_success 'can_all_from_reach_with_flag: tags case' '\n@@ -264,7 +264,7 @@ test_expect_success 'can_all_from_reach_with_flag: tags case' '\n \tY:commit-8-1\n \tEOF\n \techo \"can_all_from_reach_with_flag(X,_,_,0,0):1\" >expect &&\n-\ttest_three_modes can_all_from_reach_with_flag\n+\ttest_all_modes can_all_from_reach_with_flag\n '\n \n test_expect_success 'commit_contains:hit' '\n@@ -280,8 +280,8 @@ test_expect_success 'commit_contains:hit' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):1\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'commit_contains:miss' '\n@@ -297,8 +297,8 @@ test_expect_success 'commit_contains:miss' '\n \tX:commit-9-3\n \tEOF\n \techo \"commit_contains(_,A,X,_):0\" >expect &&\n-\ttest_three_modes commit_contains &&\n-\ttest_three_modes commit_contains --tag\n+\ttest_all_modes commit_contains &&\n+\ttest_all_modes commit_contains --tag\n '\n \n test_expect_success 'rev-list: basic topo-order' '\n@@ -310,7 +310,7 @@ test_expect_success 'rev-list: basic topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 commit-3-2 commit-2-2 commit-1-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent topo-order' '\n@@ -322,7 +322,7 @@ test_expect_success 'rev-list: first-parent topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 commit-3-1 commit-2-1 commit-1-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -334,7 +334,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: range topo-order' '\n@@ -346,7 +346,7 @@ test_expect_success 'rev-list: range topo-order' '\n \t\tcommit-6-2 commit-5-2 commit-4-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: first-parent range topo-order' '\n@@ -358,7 +358,7 @@ test_expect_success 'rev-list: first-parent range topo-order' '\n \t\tcommit-6-2 \\\n \t\tcommit-6-1 commit-5-1 commit-4-1 \\\n \t>expect &&\n-\trun_three_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n+\trun_all_modes git rev-list --first-parent --topo-order commit-3-8..commit-6-6\n '\n \n test_expect_success 'rev-list: ancestry-path topo-order' '\n@@ -368,7 +368,7 @@ test_expect_success 'rev-list: ancestry-path topo-order' '\n \t\tcommit-6-4 commit-5-4 commit-4-4 commit-3-4 \\\n \t\tcommit-6-3 commit-5-3 commit-4-3 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n+\trun_all_modes git rev-list --topo-order --ancestry-path commit-3-3..commit-6-6\n '\n \n test_expect_success 'rev-list: symmetric difference topo-order' '\n@@ -382,7 +382,7 @@ test_expect_success 'rev-list: symmetric difference topo-order' '\n \t\tcommit-3-8 commit-2-8 commit-1-8 \\\n \t\tcommit-3-7 commit-2-7 commit-1-7 \\\n \t>expect &&\n-\trun_three_modes git rev-list --topo-order commit-3-8...commit-6-6\n+\trun_all_modes git rev-list --topo-order commit-3-8...commit-6-6\n '\n \n test_expect_success 'get_reachable_subset:all' '\n@@ -402,7 +402,7 @@ test_expect_success 'get_reachable_subset:all' '\n \t\t\t      commit-1-7 \\\n \t\t\t      commit-5-6 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:some' '\n@@ -420,7 +420,7 @@ test_expect_success 'get_reachable_subset:some' '\n \t\tgit rev-parse commit-3-3 \\\n \t\t\t      commit-1-7 | sort\n \t) >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_expect_success 'get_reachable_subset:none' '\n@@ -434,7 +434,7 @@ test_expect_success 'get_reachable_subset:none' '\n \tY:commit-2-8\n \tEOF\n \techo \"get_reachable_subset(X,Y)\" >expect &&\n-\ttest_three_modes get_reachable_subset\n+\ttest_all_modes get_reachable_subset\n '\n \n test_done\n-- \ngitgitgadget\n\n"},{"id":"415716","messageId":"3b1aae4106a409cf4652e07295683f92d7bf1fb2.1612162726.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"[PATCH v7 05/11] commit-graph: add a slab to store topological levels","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:39Z","receivedAt":"2021-02-01T07:00:15Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a later commit we will introduce corrected commit date as the\ngeneration number v2. Corrected commit dates will be stored in the new\nseperate Generation Data chunk. However, to ensure backwards\ncompatibility with \"Old\" Git we need to continue to write generation\nnumber v1 (topological levels) to the commit data chunk. Thus, we need\nto compute and store both versions of generation numbers to write the\ncommit-graph file.\n\nTherefore, let's introduce a commit-slab `topo_level_slab` to store\ntopological levels; corrected commit date will be stored in the member\n`generation` of struct commit_graph_data.\n\nThe macros `GENERATION_NUMBER_INFINITY` and `GENERATION_NUMBER_ZERO`\nmark commits not in the commit-graph file and commits written by a\nversion of Git that did not compute generation numbers respectively.\nGeneration numbers are computed identically for both kinds of commits.\n\nA \"slab-miss\" should return `GENERATION_NUMBER_INFINITY` as the commit\nis not in the commit-graph file. However, since the slab is\nzero-initialized, it returns 0 (or rather `GENERATION_NUMBER_ZERO`).\nThus, we no longer need to check if the topological level of a commit is\n`GENERATION_NUMBER_INFINITY`.\n\nWe will add a pointer to the slab in `struct write_commit_graph_context`\nand `struct commit_graph` to populate the slab in\n`fill_commit_graph_info` if the commit has a pre-computed topological\nlevel as in case of split commit-graphs.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 45 ++++++++++++++++++++++++++++++---------------\n commit-graph.h |  1 +\n 2 files changed, 31 insertions(+), 15 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 955418bd6e5..2f344cce151 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -64,6 +64,8 @@ void git_test_write_commit_graph_or_die(void)\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n+define_commit_slab(topo_level_slab, uint32_t);\n+\n /* Keep track of the order in which commits are added to our list. */\n define_commit_slab(commit_pos, int);\n static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n@@ -772,6 +774,9 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n \tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\n+\tif (g->topo_levels)\n+\t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n }\n \n static inline void set_commit_tree(struct commit *c, struct tree *t)\n@@ -960,6 +965,7 @@ struct write_commit_graph_context {\n \t\t changed_paths:1,\n \t\t order_by_pack:1;\n \n+\tstruct topo_level_slab *topo_levels;\n \tconst struct commit_graph_opts *opts;\n \tsize_t total_bloom_filter_data_size;\n \tconst struct bloom_filter_settings *bloom_settings;\n@@ -1106,7 +1112,7 @@ static int write_graph_chunk_data(struct hashfile *f,\n \t\telse\n \t\t\tpackedDate[0] = 0;\n \n-\t\tpackedDate[0] |= htonl(commit_graph_data_at(*list)->generation << 2);\n+\t\tpackedDate[0] |= htonl(*topo_level_slab_at(ctx->topo_levels, *list) << 2);\n \n \t\tpackedDate[1] = htonl((*list)->date);\n \t\thashwrite(f, packedDate, 8);\n@@ -1336,11 +1342,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\t_(\"Computing commit graph generation numbers\"),\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\tuint32_t generation = commit_graph_data_at(ctx->commits.list[i])->generation;\n+\t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (generation != GENERATION_NUMBER_INFINITY &&\n-\t\t    generation != GENERATION_NUMBER_ZERO)\n+\t\tif (level != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1348,29 +1353,26 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit *current = list->item;\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n-\t\t\tuint32_t max_generation = 0;\n+\t\t\tuint32_t max_level = 0;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n-\t\t\t\tgeneration = commit_graph_data_at(parent->item)->generation;\n+\t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n \n-\t\t\t\tif (generation == GENERATION_NUMBER_INFINITY ||\n-\t\t\t\t    generation == GENERATION_NUMBER_ZERO) {\n+\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (generation > max_generation) {\n-\t\t\t\t\tmax_generation = generation;\n+\t\t\t\t} else if (level > max_level) {\n+\t\t\t\t\tmax_level = level;\n \t\t\t\t}\n \t\t\t}\n \n \t\t\tif (all_parents_computed) {\n-\t\t\t\tstruct commit_graph_data *data = commit_graph_data_at(current);\n-\n-\t\t\t\tdata->generation = max_generation + 1;\n \t\t\t\tpop_commit(&list);\n \n-\t\t\t\tif (data->generation > GENERATION_NUMBER_MAX)\n-\t\t\t\t\tdata->generation = GENERATION_NUMBER_MAX;\n+\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n+\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n+\t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n \t\t\t}\n \t\t}\n \t}\n@@ -2106,6 +2108,7 @@ int write_commit_graph(struct object_directory *odb,\n \tint res = 0;\n \tint replace = 0;\n \tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n+\tstruct topo_level_slab topo_levels;\n \n \tprepare_repo_settings(the_repository);\n \tif (!the_repository->settings.core_commit_graph) {\n@@ -2132,6 +2135,18 @@ int write_commit_graph(struct object_directory *odb,\n \t\t\t\t\t\t\t bloom_settings.max_changed_paths);\n \tctx->bloom_settings = &bloom_settings;\n \n+\tinit_topo_level_slab(&topo_levels);\n+\tctx->topo_levels = &topo_levels;\n+\n+\tif (ctx->r->objects->commit_graph) {\n+\t\tstruct commit_graph *g = ctx->r->objects->commit_graph;\n+\n+\t\twhile (g) {\n+\t\t\tg->topo_levels = &topo_levels;\n+\t\t\tg = g->base_graph;\n+\t\t}\n+\t}\n+\n \tif (flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS)\n \t\tctx->changed_paths = 1;\n \tif (!(flags & COMMIT_GRAPH_NO_WRITE_BLOOM_FILTERS)) {\ndiff --git a/commit-graph.h b/commit-graph.h\nindex f8e92500c6e..00f00745b79 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -73,6 +73,7 @@ struct commit_graph {\n \tconst unsigned char *chunk_bloom_indexes;\n \tconst unsigned char *chunk_bloom_data;\n \n+\tstruct topo_level_slab *topo_levels;\n \tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \n-- \ngitgitgadget\n\n"},{"id":"415717","messageId":"ea32cba16ef8533117d583aea46cd5a254ef4e24.1612162726.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"[PATCH v7 06/11] commit-graph: return 64-bit generation number","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:40Z","receivedAt":"2021-02-01T07:00:28Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nIn a preparatory step for introducing corrected commit dates, let's\nreturn timestamp_t values from commit_graph_generation(), use\ntimestamp_t for local variables and define GENERATION_NUMBER_INFINITY\nas (2 ^ 63 - 1) instead.\n\nWe rename GENERATION_NUMBER_MAX to GENERATION_NUMBER_V1_MAX to\nrepresent the largest topological level we can store in the commit data\nchunk.\n\nWith corrected commit dates implemented, we will have two such *_MAX\nvariables to denote the largest offset and largest topological level\nthat can be stored.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 22 +++++++++++-----------\n commit-graph.h |  4 ++--\n commit-reach.c | 36 ++++++++++++++++++------------------\n commit-reach.h |  2 +-\n commit.c       |  4 ++--\n commit.h       |  4 ++--\n revision.c     | 10 +++++-----\n upload-pack.c  |  2 +-\n 8 files changed, 42 insertions(+), 42 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 2f344cce151..8f17815021d 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -101,7 +101,7 @@ uint32_t commit_graph_position(const struct commit *c)\n \treturn data ? data->graph_pos : COMMIT_NOT_FROM_GRAPH;\n }\n \n-uint32_t commit_graph_generation(const struct commit *c)\n+timestamp_t commit_graph_generation(const struct commit *c)\n {\n \tstruct commit_graph_data *data =\n \t\tcommit_graph_data_slab_peek(&commit_graph_data_slab, c);\n@@ -150,8 +150,8 @@ static int commit_gen_cmp(const void *va, const void *vb)\n \tconst struct commit *a = *(const struct commit **)va;\n \tconst struct commit *b = *(const struct commit **)vb;\n \n-\tuint32_t generation_a = commit_graph_data_at(a)->generation;\n-\tuint32_t generation_b = commit_graph_data_at(b)->generation;\n+\tconst timestamp_t generation_a = commit_graph_data_at(a)->generation;\n+\tconst timestamp_t generation_b = commit_graph_data_at(b)->generation;\n \t/* lower generation commits first */\n \tif (generation_a < generation_b)\n \t\treturn -1;\n@@ -1370,8 +1370,8 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tif (all_parents_computed) {\n \t\t\t\tpop_commit(&list);\n \n-\t\t\t\tif (max_level > GENERATION_NUMBER_MAX - 1)\n-\t\t\t\t\tmax_level = GENERATION_NUMBER_MAX - 1;\n+\t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n+\t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n \t\t\t}\n \t\t}\n@@ -2367,8 +2367,8 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \tfor (i = 0; i < g->num_commits; i++) {\n \t\tstruct commit *graph_commit, *odb_commit;\n \t\tstruct commit_list *graph_parents, *odb_parents;\n-\t\tuint32_t max_generation = 0;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t max_generation = 0;\n+\t\ttimestamp_t generation;\n \n \t\tdisplay_progress(progress, i + 1);\n \t\thashcpy(cur_oid.hash, g->chunk_oid_lookup + g->hash_len * i);\n@@ -2432,16 +2432,16 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t\tcontinue;\n \n \t\t/*\n-\t\t * If one of our parents has generation GENERATION_NUMBER_MAX, then\n-\t\t * our generation is also GENERATION_NUMBER_MAX. Decrement to avoid\n+\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n+\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n \t\t * extra logic in the following condition.\n \t\t */\n-\t\tif (max_generation == GENERATION_NUMBER_MAX)\n+\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n \t\t\tmax_generation--;\n \n \t\tgeneration = commit_graph_generation(graph_commit);\n \t\tif (generation != max_generation + 1)\n-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %u != %u\"),\n+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" != %\"PRItime),\n \t\t\t\t     oid_to_hex(&cur_oid),\n \t\t\t\t     generation,\n \t\t\t\t     max_generation + 1);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 00f00745b79..2e9aa7824ee 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -145,12 +145,12 @@ void disable_commit_graph(struct repository *r);\n \n struct commit_graph_data {\n \tuint32_t graph_pos;\n-\tuint32_t generation;\n+\ttimestamp_t generation;\n };\n \n /*\n  * Commits should be parsed before accessing generation, graph positions.\n  */\n-uint32_t commit_graph_generation(const struct commit *);\n+timestamp_t commit_graph_generation(const struct commit *);\n uint32_t commit_graph_position(const struct commit *);\n #endif\ndiff --git a/commit-reach.c b/commit-reach.c\nindex 50175b159e7..9b24b0378d5 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -32,12 +32,12 @@ static int queue_has_nonstale(struct prio_queue *queue)\n static struct commit_list *paint_down_to_common(struct repository *r,\n \t\t\t\t\t\tstruct commit *one, int n,\n \t\t\t\t\t\tstruct commit **twos,\n-\t\t\t\t\t\tint min_generation)\n+\t\t\t\t\t\ttimestamp_t min_generation)\n {\n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \tstruct commit_list *result = NULL;\n \tint i;\n-\tuint32_t last_gen = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n \tif (!min_generation)\n \t\tqueue.compare = compare_commits_by_commit_date;\n@@ -58,10 +58,10 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \t\tstruct commit *commit = prio_queue_get(&queue);\n \t\tstruct commit_list *parents;\n \t\tint flags;\n-\t\tuint32_t generation = commit_graph_generation(commit);\n+\t\ttimestamp_t generation = commit_graph_generation(commit);\n \n \t\tif (min_generation && generation > last_gen)\n-\t\t\tBUG(\"bad generation skip %8x > %8x at %s\",\n+\t\t\tBUG(\"bad generation skip %\"PRItime\" > %\"PRItime\" at %s\",\n \t\t\t    generation, last_gen,\n \t\t\t    oid_to_hex(&commit->object.oid));\n \t\tlast_gen = generation;\n@@ -177,12 +177,12 @@ static int remove_redundant(struct repository *r, struct commit **array, int cnt\n \t\trepo_parse_commit(r, array[i]);\n \tfor (i = 0; i < cnt; i++) {\n \t\tstruct commit_list *common;\n-\t\tuint32_t min_generation = commit_graph_generation(array[i]);\n+\t\ttimestamp_t min_generation = commit_graph_generation(array[i]);\n \n \t\tif (redundant[i])\n \t\t\tcontinue;\n \t\tfor (j = filled = 0; j < cnt; j++) {\n-\t\t\tuint32_t curr_generation;\n+\t\t\ttimestamp_t curr_generation;\n \t\t\tif (i == j || redundant[j])\n \t\t\t\tcontinue;\n \t\t\tfilled_index[filled] = j;\n@@ -321,7 +321,7 @@ int repo_in_merge_bases_many(struct repository *r, struct commit *commit,\n {\n \tstruct commit_list *bases;\n \tint ret = 0, i;\n-\tuint32_t generation, max_generation = GENERATION_NUMBER_ZERO;\n+\ttimestamp_t generation, max_generation = GENERATION_NUMBER_ZERO;\n \n \tif (repo_parse_commit(r, commit))\n \t\treturn ret;\n@@ -470,7 +470,7 @@ static int in_commit_list(const struct commit_list *want, struct commit *c)\n static enum contains_result contains_test(struct commit *candidate,\n \t\t\t\t\t  const struct commit_list *want,\n \t\t\t\t\t  struct contains_cache *cache,\n-\t\t\t\t\t  uint32_t cutoff)\n+\t\t\t\t\t  timestamp_t cutoff)\n {\n \tenum contains_result *cached = contains_cache_at(cache, candidate);\n \n@@ -506,11 +506,11 @@ static enum contains_result contains_tag_algo(struct commit *candidate,\n {\n \tstruct contains_stack contains_stack = { 0, 0, NULL };\n \tenum contains_result result;\n-\tuint32_t cutoff = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t cutoff = GENERATION_NUMBER_INFINITY;\n \tconst struct commit_list *p;\n \n \tfor (p = want; p; p = p->next) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = p->item;\n \t\tload_commit_graph_info(the_repository, c);\n \t\tgeneration = commit_graph_generation(c);\n@@ -566,8 +566,8 @@ static int compare_commits_by_gen(const void *_a, const void *_b)\n \tconst struct commit *a = *(const struct commit * const *)_a;\n \tconst struct commit *b = *(const struct commit * const *)_b;\n \n-\tuint32_t generation_a = commit_graph_generation(a);\n-\tuint32_t generation_b = commit_graph_generation(b);\n+\ttimestamp_t generation_a = commit_graph_generation(a);\n+\ttimestamp_t generation_b = commit_graph_generation(b);\n \n \tif (generation_a < generation_b)\n \t\treturn -1;\n@@ -580,7 +580,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation)\n+\t\t\t\t timestamp_t min_generation)\n {\n \tstruct commit **list = NULL;\n \tint i;\n@@ -681,13 +681,13 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \ttime_t min_commit_date = cutoff_by_min_date ? from->item->date : 0;\n \tstruct commit_list *from_iter = from, *to_iter = to;\n \tint result;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \n \twhile (from_iter) {\n \t\tadd_object_array(&from_iter->item->object, NULL, &from_objs);\n \n \t\tif (!parse_commit(from_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (from_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = from_iter->item->date;\n \n@@ -701,7 +701,7 @@ int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \n \twhile (to_iter) {\n \t\tif (!parse_commit(to_iter->item)) {\n-\t\t\tuint32_t generation;\n+\t\t\ttimestamp_t generation;\n \t\t\tif (to_iter->item->date < min_commit_date)\n \t\t\t\tmin_commit_date = to_iter->item->date;\n \n@@ -741,13 +741,13 @@ struct commit_list *get_reachable_subset(struct commit **from, int nr_from,\n \tstruct commit_list *found_commits = NULL;\n \tstruct commit **to_last = to + nr_to;\n \tstruct commit **from_last = from + nr_from;\n-\tuint32_t min_generation = GENERATION_NUMBER_INFINITY;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_INFINITY;\n \tint num_to_find = 0;\n \n \tstruct prio_queue queue = { compare_commits_by_gen_then_commit_date };\n \n \tfor (item = to; item < to_last; item++) {\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \t\tstruct commit *c = *item;\n \n \t\tparse_commit(c);\ndiff --git a/commit-reach.h b/commit-reach.h\nindex b49ad71a317..148b56fea50 100644\n--- a/commit-reach.h\n+++ b/commit-reach.h\n@@ -87,7 +87,7 @@ int can_all_from_reach_with_flag(struct object_array *from,\n \t\t\t\t unsigned int with_flag,\n \t\t\t\t unsigned int assign_flag,\n \t\t\t\t time_t min_commit_date,\n-\t\t\t\t uint32_t min_generation);\n+\t\t\t\t timestamp_t min_generation);\n int can_all_from_reach(struct commit_list *from, struct commit_list *to,\n \t\t       int commit_date_cutoff);\n \ndiff --git a/commit.c b/commit.c\nindex bab8d5ab07c..4c717329ee0 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -753,8 +753,8 @@ int compare_commits_by_author_date(const void *a_, const void *b_,\n int compare_commits_by_gen_then_commit_date(const void *a_, const void *b_, void *unused)\n {\n \tconst struct commit *a = a_, *b = b_;\n-\tconst uint32_t generation_a = commit_graph_generation(a),\n-\t\t       generation_b = commit_graph_generation(b);\n+\tconst timestamp_t generation_a = commit_graph_generation(a),\n+\t\t\t  generation_b = commit_graph_generation(b);\n \n \t/* newer commits first */\n \tif (generation_a < generation_b)\ndiff --git a/commit.h b/commit.h\nindex f4e7b0158e2..742d96c41e8 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -11,8 +11,8 @@\n #include \"commit-slab.h\"\n \n #define COMMIT_NOT_FROM_GRAPH 0xFFFFFFFF\n-#define GENERATION_NUMBER_INFINITY 0xFFFFFFFF\n-#define GENERATION_NUMBER_MAX 0x3FFFFFFF\n+#define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n+#define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n \n struct commit_list {\ndiff --git a/revision.c b/revision.c\nindex 5474001331a..a54d2bd28df 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -3302,7 +3302,7 @@ define_commit_slab(indegree_slab, int);\n define_commit_slab(author_date_slab, timestamp_t);\n \n struct topo_walk_info {\n-\tuint32_t min_generation;\n+\ttimestamp_t min_generation;\n \tstruct prio_queue explore_queue;\n \tstruct prio_queue indegree_queue;\n \tstruct prio_queue topo_queue;\n@@ -3370,7 +3370,7 @@ static void explore_walk_step(struct rev_info *revs)\n }\n \n static void explore_to_depth(struct rev_info *revs,\n-\t\t\t     uint32_t gen_cutoff)\n+\t\t\t     timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3415,7 +3415,7 @@ static void indegree_walk_step(struct rev_info *revs)\n }\n \n static void compute_indegrees_to_depth(struct rev_info *revs,\n-\t\t\t\t       uint32_t gen_cutoff)\n+\t\t\t\t       timestamp_t gen_cutoff)\n {\n \tstruct topo_walk_info *info = revs->topo_walk_info;\n \tstruct commit *c;\n@@ -3473,7 +3473,7 @@ static void init_topo_walk(struct rev_info *revs)\n \tinfo->min_generation = GENERATION_NUMBER_INFINITY;\n \tfor (list = revs->commits; list; list = list->next) {\n \t\tstruct commit *c = list->item;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (repo_parse_commit_gently(revs->repo, c, 1))\n \t\t\tcontinue;\n@@ -3541,7 +3541,7 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n \tfor (p = commit->parents; p; p = p->next) {\n \t\tstruct commit *parent = p->item;\n \t\tint *pi;\n-\t\tuint32_t generation;\n+\t\ttimestamp_t generation;\n \n \t\tif (parent->object.flags & UNINTERESTING)\n \t\t\tcontinue;\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 3b66bf92ba8..b87607e0dd4 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -500,7 +500,7 @@ static int got_oid(struct upload_pack_data *data,\n \n static int ok_to_give_up(struct upload_pack_data *data)\n {\n-\tuint32_t min_generation = GENERATION_NUMBER_ZERO;\n+\ttimestamp_t min_generation = GENERATION_NUMBER_ZERO;\n \n \tif (!data->have_obj.nr)\n \t\treturn 0;\n-- \ngitgitgadget\n\n"},{"id":"415718","messageId":"ec598f1d500b542953e8786f67f35115c2b29fec.1612162726.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"[PATCH v7 08/11] commit-graph: implement corrected commit date","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:42Z","receivedAt":"2021-02-01T07:00:40Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nWith most of preparations done, let's implement corrected commit date.\n\nThe corrected commit date for a commit is defined as:\n\n* A commit with no parents (a root commit) has corrected commit date\n  equal to its committer date.\n* A commit with at least one parent has corrected commit date equal to\n  the maximum of its commit date and one more than the largest corrected\n  commit date among its parents.\n\nAs a special case, a root commit with timestamp of zero (01.01.1970\n00:00:00Z) has corrected commit date of one, to be able to distinguish\nfrom GENERATION_NUMBER_ZERO (that is, an uncomputed corrected commit\ndate).\n\nTo minimize the space required to store corrected commit date, Git\nstores corrected commit date offsets into the commit-graph file. The\ncorrected commit date offset for a commit is defined as the difference\nbetween its corrected commit date and actual commit date.\n\nStoring corrected commit date requires sizeof(timestamp_t) bytes, which\nin most cases is 64 bits (uintmax_t). However, corrected commit date\noffsets can be safely stored using only 32-bits. This halves the size\nof GDAT chunk, which is a reduction of around 6% in the size of\ncommit-graph file.\n\nHowever, using offsets be problematic if a commit is malformed but valid\nand has committer date of 0 Unix time, as the offset would be the same\nas corrected commit date and thus require 64-bits to be stored properly.\n\nWhile Git does not write out offsets at this stage, Git stores the\ncorrected commit dates in member generation of struct commit_graph_data.\nIt will begin writing commit date offsets with the introduction of\ngeneration data chunk.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c | 21 +++++++++++++++++----\n 1 file changed, 17 insertions(+), 4 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 8f17815021d..d1e6ced8647 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1343,9 +1343,11 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\t\tctx->commits.nr);\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tuint32_t level = *topo_level_slab_at(ctx->topo_levels, ctx->commits.list[i]);\n+\t\ttimestamp_t corrected_commit_date = commit_graph_data_at(ctx->commits.list[i])->generation;\n \n \t\tdisplay_progress(ctx->progress, i + 1);\n-\t\tif (level != GENERATION_NUMBER_ZERO)\n+\t\tif (level != GENERATION_NUMBER_ZERO &&\n+\t\t    corrected_commit_date != GENERATION_NUMBER_ZERO)\n \t\t\tcontinue;\n \n \t\tcommit_list_insert(ctx->commits.list[i], &list);\n@@ -1354,17 +1356,24 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\tstruct commit_list *parent;\n \t\t\tint all_parents_computed = 1;\n \t\t\tuint32_t max_level = 0;\n+\t\t\ttimestamp_t max_corrected_commit_date = 0;\n \n \t\t\tfor (parent = current->parents; parent; parent = parent->next) {\n \t\t\t\tlevel = *topo_level_slab_at(ctx->topo_levels, parent->item);\n+\t\t\t\tcorrected_commit_date = commit_graph_data_at(parent->item)->generation;\n \n-\t\t\t\tif (level == GENERATION_NUMBER_ZERO) {\n+\t\t\t\tif (level == GENERATION_NUMBER_ZERO ||\n+\t\t\t\t    corrected_commit_date == GENERATION_NUMBER_ZERO) {\n \t\t\t\t\tall_parents_computed = 0;\n \t\t\t\t\tcommit_list_insert(parent->item, &list);\n \t\t\t\t\tbreak;\n-\t\t\t\t} else if (level > max_level) {\n-\t\t\t\t\tmax_level = level;\n \t\t\t\t}\n+\n+\t\t\t\tif (level > max_level)\n+\t\t\t\t\tmax_level = level;\n+\n+\t\t\t\tif (corrected_commit_date > max_corrected_commit_date)\n+\t\t\t\t\tmax_corrected_commit_date = corrected_commit_date;\n \t\t\t}\n \n \t\t\tif (all_parents_computed) {\n@@ -1373,6 +1382,10 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\tif (max_level > GENERATION_NUMBER_V1_MAX - 1)\n \t\t\t\t\tmax_level = GENERATION_NUMBER_V1_MAX - 1;\n \t\t\t\t*topo_level_slab_at(ctx->topo_levels, current) = max_level + 1;\n+\n+\t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n+\t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n+\t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n \t\t\t}\n \t\t}\n \t}\n-- \ngitgitgadget\n\n"},{"id":"415719","messageId":"07a88f1aae6f7f7812ab7a5937eac73131c2139e.1612162726.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"[PATCH v7 10/11] commit-graph: use generation v2 only if entire chain does","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:44Z","receivedAt":"2021-02-01T07:00:42Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nSince there are released versions of Git that understand generation\nnumbers in the commit-graph's CDAT chunk but do not understand the GDAT\nchunk, the following scenario is possible:\n\n1. \"New\" Git writes a commit-graph with the GDAT chunk.\n2. \"Old\" Git writes a split commit-graph on top without a GDAT chunk.\n\nIf each layer of split commit-graph is treated independently, as it was\nthe case before this commit, with Git inspecting only the current layer\nfor chunk_generation_data pointer, commits in the lower layer (one with\nGDAT) whould have corrected commit date as their generation number,\nwhile commits in the upper layer would have topological levels as their\ngeneration. Corrected commit dates usually have much larger values than\ntopological levels. This means that if we take two commits, one from the\nupper layer, and one reachable from it in the lower layer, then the\nexpectation that the generation of a parent is smaller than the\ngeneration of a child would be violated.\n\nIt is difficult to expose this issue in a test. Since we _start_ with\nartificially low generation numbers, any commit walk that prioritizes\ngeneration numbers will walk all of the commits with high generation\nnumber before walking the commits with low generation number. In all the\ncases I tried, the commit-graph layers themselves \"protect\" any\nincorrect behavior since none of the commits in the lower layer can\nreach the commits in the upper layer.\n\nThis issue would manifest itself as a performance problem in this case,\nespecially with something like \"git log --graph\" since the low\ngeneration numbers would cause the in-degree queue to walk all of the\ncommits in the lower layer before allowing the topo-order queue to write\nanything to output (depending on the size of the upper layer).\n\nTherefore, When writing the new layer in split commit-graph, we write a\nGDAT chunk only if the topmost layer has a GDAT chunk. This guarantees\nthat if a layer has GDAT chunk, all lower layers must have a GDAT chunk\nas well.\n\nRewriting layers follows similar approach: if the topmost layer below\nthe set of layers being rewritten (in the split commit-graph chain)\nexists, and it does not contain GDAT chunk, then the result of rewrite\ndoes not have GDAT chunks either.\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                |  30 +++++-\n commit-graph.h                |   1 +\n t/t5324-split-commit-graph.sh | 181 ++++++++++++++++++++++++++++++++++\n 3 files changed, 210 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex d2afcc83283..77fef5a240e 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -614,6 +614,21 @@ static struct commit_graph *load_commit_graph_chain(struct repository *r,\n \treturn graph_chain;\n }\n \n+static void validate_mixed_generation_chain(struct commit_graph *g)\n+{\n+\tint read_generation_data;\n+\n+\tif (!g)\n+\t\treturn;\n+\n+\tread_generation_data = !!g->chunk_generation_data;\n+\n+\twhile (g) {\n+\t\tg->read_generation_data = read_generation_data;\n+\t\tg = g->base_graph;\n+\t}\n+}\n+\n struct commit_graph *read_commit_graph_one(struct repository *r,\n \t\t\t\t\t   struct object_directory *odb)\n {\n@@ -622,6 +637,8 @@ struct commit_graph *read_commit_graph_one(struct repository *r,\n \tif (!g)\n \t\tg = load_commit_graph_chain(r, odb);\n \n+\tvalidate_mixed_generation_chain(g);\n+\n \treturn g;\n }\n \n@@ -791,7 +808,7 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tif (g->chunk_generation_data) {\n+\tif (g->read_generation_data) {\n \t\toffset = (timestamp_t)get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n \n \t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n@@ -2019,6 +2036,13 @@ static void split_graph_merge_strategy(struct write_commit_graph_context *ctx)\n \t\tif (i < ctx->num_commit_graphs_after)\n \t\t\tctx->commit_graph_hash_after[i] = xstrdup(oid_to_hex(&g->oid));\n \n+\t\t/*\n+\t\t * If the topmost remaining layer has generation data chunk, the\n+\t\t * resultant layer also has generation data chunk.\n+\t\t */\n+\t\tif (i == ctx->num_commit_graphs_after - 2)\n+\t\t\tctx->write_generation_data = !!g->chunk_generation_data;\n+\n \t\ti--;\n \t\tg = g->base_graph;\n \t}\n@@ -2343,6 +2367,8 @@ int write_commit_graph(struct object_directory *odb,\n \t} else\n \t\tctx->num_commit_graphs_after = 1;\n \n+\tvalidate_mixed_generation_chain(ctx->r->objects->commit_graph);\n+\n \tcompute_generation_numbers(ctx);\n \n \tif (ctx->changed_paths)\n@@ -2541,7 +2567,7 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t * also GENERATION_NUMBER_V1_MAX. Decrement to avoid extra logic\n \t\t * in the following condition.\n \t\t */\n-\t\tif (!g->chunk_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n+\t\tif (!g->read_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n \t\t\tmax_generation--;\n \n \t\tgeneration = commit_graph_generation(graph_commit);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 19a02001fde..ad52130883b 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -64,6 +64,7 @@ struct commit_graph {\n \tstruct object_directory *odb;\n \n \tuint32_t num_commits_in_base;\n+\tunsigned int read_generation_data;\n \tstruct commit_graph *base_graph;\n \n \tconst uint32_t *chunk_oid_fanout;\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 587757b62d9..8e90f3423b8 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -453,4 +453,185 @@ test_expect_success 'prevent regression for duplicate commits across layers' '\n \tgit -C dup commit-graph verify\n '\n \n+NUM_FIRST_LAYER_COMMITS=64\n+NUM_SECOND_LAYER_COMMITS=16\n+NUM_THIRD_LAYER_COMMITS=7\n+NUM_FOURTH_LAYER_COMMITS=8\n+NUM_FIFTH_LAYER_COMMITS=16\n+SECOND_LAYER_SEQUENCE_START=$(($NUM_FIRST_LAYER_COMMITS + 1))\n+SECOND_LAYER_SEQUENCE_END=$(($SECOND_LAYER_SEQUENCE_START + $NUM_SECOND_LAYER_COMMITS - 1))\n+THIRD_LAYER_SEQUENCE_START=$(($SECOND_LAYER_SEQUENCE_END + 1))\n+THIRD_LAYER_SEQUENCE_END=$(($THIRD_LAYER_SEQUENCE_START + $NUM_THIRD_LAYER_COMMITS - 1))\n+FOURTH_LAYER_SEQUENCE_START=$(($THIRD_LAYER_SEQUENCE_END + 1))\n+FOURTH_LAYER_SEQUENCE_END=$(($FOURTH_LAYER_SEQUENCE_START + $NUM_FOURTH_LAYER_COMMITS - 1))\n+FIFTH_LAYER_SEQUENCE_START=$(($FOURTH_LAYER_SEQUENCE_END + 1))\n+FIFTH_LAYER_SEQUENCE_END=$(($FIFTH_LAYER_SEQUENCE_START + $NUM_FIFTH_LAYER_COMMITS - 1))\n+\n+# Current split graph chain:\n+#\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'setup repo for mixed generation commit-graph-chain' '\n+\tgraphdir=\".git/objects/info/commit-graphs\" &&\n+\ttest_oid_cache <<-EOF &&\n+\toid_version sha1:1\n+\toid_version sha256:2\n+\tEOF\n+\tgit init mixed &&\n+\t(\n+\t\tcd mixed &&\n+\t\tgit config core.commitGraph true &&\n+\t\tgit config gc.writeCommitGraph false &&\n+\t\tfor i in $(test_seq $NUM_FIRST_LAYER_COMMITS)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split &&\n+\t\tgraph_read_expect $NUM_FIRST_LAYER_COMMITS &&\n+\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\t\tfor i in $(test_seq $SECOND_LAYER_SEQUENCE_START $SECOND_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable --split=no-merge &&\n+\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 4 1\n+\t\tnum_commits: $NUM_SECOND_LAYER_COMMITS\n+\t\tchunks: oid_fanout oid_lookup commit_metadata\n+\t\tEOF\n+\t\ttest_cmp expect output &&\n+\t\tgit commit-graph verify &&\n+\t\tcat $graphdir/commit-graph-chain\n+\t)\n+'\n+\n+# The new layer will be added without generation data chunk as it was not\n+# present on the layer underneath it.\n+#\n+#      7 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'do not write generation data chunk if not present on existing tip' '\n+\tgit clone mixed mixed-no-gdat &&\n+\t(\n+\t\tcd mixed-no-gdat &&\n+\t\tfor i in $(test_seq $THIRD_LAYER_SEQUENCE_START $THIRD_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split=no-merge &&\n+\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 4 2\n+\t\tnum_commits: $NUM_THIRD_LAYER_COMMITS\n+\t\tchunks: oid_fanout oid_lookup commit_metadata\n+\t\tEOF\n+\t\ttest_cmp expect output &&\n+\t\tgit commit-graph verify\n+\t)\n+'\n+\n+# Number of commits in each layer of the split-commit graph before merge:\n+#\n+#      8 commits (No GDAT)\n+# ------------------------\n+#      7 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+# The top two layers are merged and do not have generation data chunk as layer below them does\n+# not have generation data chunk.\n+#\n+#     15 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'do not write generation data chunk if the topmost remaining layer does not have generation data chunk' '\n+\tgit clone mixed-no-gdat mixed-merge-no-gdat &&\n+\t(\n+\t\tcd mixed-merge-no-gdat &&\n+\t\tfor i in $(test_seq $FOURTH_LAYER_SEQUENCE_START $FOURTH_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split --size-multiple 1 &&\n+\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 4 2\n+\t\tnum_commits: $(($NUM_THIRD_LAYER_COMMITS + $NUM_FOURTH_LAYER_COMMITS))\n+\t\tchunks: oid_fanout oid_lookup commit_metadata\n+\t\tEOF\n+\t\ttest_cmp expect output &&\n+\t\tgit commit-graph verify\n+\t)\n+'\n+\n+# Number of commits in each layer of the split-commit graph before merge:\n+#\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     15 commits (No GDAT)\n+# ------------------------\n+#     16 commits (No GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+# The top three layers are merged and has generation data chunk as the topmost remaining layer\n+# has generation data chunk.\n+#\n+#     47 commits (GDAT)\n+# ------------------------\n+#     64 commits (GDAT)\n+#\n+test_expect_success 'write generation data chunk if topmost remaining layer has generation data chunk' '\n+\tgit clone mixed-merge-no-gdat mixed-merge-gdat &&\n+\t(\n+\t\tcd mixed-merge-gdat &&\n+\t\tfor i in $(test_seq $FIFTH_LAYER_SEQUENCE_START $FIFTH_LAYER_SEQUENCE_END)\n+\t\tdo\n+\t\t\ttest_commit $i &&\n+\t\t\tgit branch commits/$i || return 1\n+\t\tdone &&\n+\t\tgit commit-graph write --reachable --split --size-multiple 1 &&\n+\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\t\ttest-tool read-graph >output &&\n+\t\tcat >expect <<-EOF &&\n+\t\theader: 43475048 1 $(test_oid oid_version) 5 1\n+\t\tnum_commits: $(($NUM_SECOND_LAYER_COMMITS + $NUM_THIRD_LAYER_COMMITS + $NUM_FOURTH_LAYER_COMMITS + $NUM_FIFTH_LAYER_COMMITS))\n+\t\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n+\t\tEOF\n+\t\ttest_cmp expect output\n+\t)\n+'\n+\n+test_expect_success 'write generation data chunk when commit-graph chain is replaced' '\n+\tgit clone mixed mixed-replace &&\n+\t(\n+\t\tcd mixed-replace &&\n+\t\tgit commit-graph write --reachable --split=replace &&\n+\t\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\t\tverify_chain_files_exist $graphdir &&\n+\t\tgraph_read_expect $(($NUM_FIRST_LAYER_COMMITS + $NUM_SECOND_LAYER_COMMITS)) &&\n+\t\tgit commit-graph verify\n+\t)\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"415720","messageId":"71d815188571c3ea421244918b73f762cd7963e3.1612162726.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"[PATCH v7 09/11] commit-graph: implement generation data chunk","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:43Z","receivedAt":"2021-02-01T07:00:44Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nAs discovered by Ævar, we cannot increment graph version to\ndistinguish between generation numbers v1 and v2 [1]. Thus, one of\npre-requistes before implementing generation number v2 was to\ndistinguish between graph versions in a backwards compatible manner.\n\nWe are going to introduce a new chunk called Generation DATa chunk (or\nGDAT). GDAT will store corrected committer date offsets whereas CDAT\nwill still store topological level.\n\nOld Git does not understand GDAT chunk and would ignore it, reading\ntopological levels from CDAT. New Git can parse GDAT and take advantage\nof newer generation numbers, falling back to topological levels when\nGDAT chunk is missing (as it would happen with a commit-graph written\nby old Git).\n\nWe introduce a test environment variable 'GIT_TEST_COMMIT_GRAPH_NO_GDAT'\nwhich forces commit-graph file to be written without generation data\nchunk to emulate a commit-graph file written by old Git.\n\nTo minimize the space required to store corrrected commit date, Git\nstores corrected commit date offsets into the commit-graph file, instea\nof corrected commit dates. This saves us 4 bytes per commit, decreasing\nthe GDAT chunk size by half, but it's possible for the offset to\noverflow the 4-bytes allocated for storage. As such overflows are and\nshould be exceedingly rare, we use the following overflow management\nscheme:\n\nWe introduce a new commit-graph chunk, Generation Data OVerflow ('GDOV')\nto store corrected commit dates for commits with offsets greater than\nGENERATION_NUMBER_V2_OFFSET_MAX.\n\nIf the offset is greater than GENERATION_NUMBER_V2_OFFSET_MAX, we set\nthe MSB of the offset and the other bits store the position of corrected\ncommit date in GDOV chunk, similar to how Extra Edge List is maintained.\n\nWe test the overflow-related code with the following repo history:\n\n           F - N - U\n          /         \\\nU - N - U            N\n         \\          /\n\t  N - F - N\n\nWhere the commits denoted by U have committer date of zero seconds\nsince Unix epoch, the commits denoted by N have committer date of\n1112354055 (default committer date for the test suite) seconds since\nUnix epoch and the commits denoted by F have committer date of\n(2 ^ 31 - 2) seconds since Unix epoch.\n\nThe largest offset observed is 2 ^ 31, just large enough to overflow.\n\n[1]: https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c                | 114 ++++++++++++++++++++++++++++++----\n commit-graph.h                |   3 +\n commit.h                      |   1 +\n t/README                      |   3 +\n t/helper/test-read-graph.c    |   4 ++\n t/t4216-log-bloom.sh          |   4 +-\n t/t5318-commit-graph.sh       |  79 +++++++++++++++++++----\n t/t5324-split-commit-graph.sh |  12 ++--\n t/t6600-test-reach.sh         |   6 ++\n t/test-lib-functions.sh       |   6 ++\n 10 files changed, 200 insertions(+), 32 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex d1e6ced8647..d2afcc83283 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -38,11 +38,13 @@ void git_test_write_commit_graph_or_die(void)\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n+#define GRAPH_CHUNKID_GENERATION_DATA 0x47444154 /* \"GDAT\" */\n+#define GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW 0x47444f56 /* \"GDOV\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n #define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n-#define MAX_NUM_CHUNKS 7\n+#define MAX_NUM_CHUNKS 9\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -61,6 +63,8 @@ void git_test_write_commit_graph_or_die(void)\n #define GRAPH_MIN_SIZE (GRAPH_HEADER_SIZE + 4 * GRAPH_CHUNKLOOKUP_WIDTH \\\n \t\t\t+ GRAPH_FANOUT_SIZE + the_hash_algo->rawsz)\n \n+#define CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW (1ULL << 31)\n+\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n@@ -394,6 +398,20 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n \t\t\t\tgraph->chunk_commit_data = data + chunk_offset;\n \t\t\tbreak;\n \n+\t\tcase GRAPH_CHUNKID_GENERATION_DATA:\n+\t\t\tif (graph->chunk_generation_data)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_generation_data = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW:\n+\t\t\tif (graph->chunk_generation_data_overflow)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_generation_data_overflow = data + chunk_offset;\n+\t\t\tbreak;\n+\n \t\tcase GRAPH_CHUNKID_EXTRAEDGES:\n \t\t\tif (graph->chunk_extra_edges)\n \t\t\t\tchunk_repeated = 1;\n@@ -754,8 +772,8 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n {\n \tconst unsigned char *commit_data;\n \tstruct commit_graph_data *graph_data;\n-\tuint32_t lex_index;\n-\tuint64_t date_high, date_low;\n+\tuint32_t lex_index, offset_pos;\n+\tuint64_t date_high, date_low, offset;\n \n \twhile (pos < g->num_commits_in_base)\n \t\tg = g->base_graph;\n@@ -773,7 +791,19 @@ static void fill_commit_graph_info(struct commit *item, struct commit_graph *g,\n \tdate_low = get_be32(commit_data + g->hash_len + 12);\n \titem->date = (timestamp_t)((date_high << 32) | date_low);\n \n-\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n+\tif (g->chunk_generation_data) {\n+\t\toffset = (timestamp_t)get_be32(g->chunk_generation_data + sizeof(uint32_t) * lex_index);\n+\n+\t\tif (offset & CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW) {\n+\t\t\tif (!g->chunk_generation_data_overflow)\n+\t\t\t\tdie(_(\"commit-graph requires overflow generation data but has none\"));\n+\n+\t\t\toffset_pos = offset ^ CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW;\n+\t\t\tgraph_data->generation = get_be64(g->chunk_generation_data_overflow + 8 * offset_pos);\n+\t\t} else\n+\t\t\tgraph_data->generation = item->date + offset;\n+\t} else\n+\t\tgraph_data->generation = get_be32(commit_data + g->hash_len + 8) >> 2;\n \n \tif (g->topo_levels)\n \t\t*topo_level_slab_at(g->topo_levels, item) = get_be32(commit_data + g->hash_len + 8) >> 2;\n@@ -945,6 +975,7 @@ struct write_commit_graph_context {\n \tstruct oid_array oids;\n \tstruct packed_commit_list commits;\n \tint num_extra_edges;\n+\tint num_generation_data_overflows;\n \tunsigned long approx_nr_objects;\n \tstruct progress *progress;\n \tint progress_done;\n@@ -963,7 +994,8 @@ struct write_commit_graph_context {\n \t\t report_progress:1,\n \t\t split:1,\n \t\t changed_paths:1,\n-\t\t order_by_pack:1;\n+\t\t order_by_pack:1,\n+\t\t write_generation_data:1;\n \n \tstruct topo_level_slab *topo_levels;\n \tconst struct commit_graph_opts *opts;\n@@ -1123,6 +1155,45 @@ static int write_graph_chunk_data(struct hashfile *f,\n \treturn 0;\n }\n \n+static int write_graph_chunk_generation_data(struct hashfile *f,\n+\t\t\t\t\t      struct write_commit_graph_context *ctx)\n+{\n+\tint i, num_generation_data_overflows = 0;\n+\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n+\n+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n+\t\t\toffset = CORRECTED_COMMIT_DATE_OFFSET_OVERFLOW | num_generation_data_overflows;\n+\t\t\tnum_generation_data_overflows++;\n+\t\t}\n+\n+\t\thashwrite_be32(f, offset);\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int write_graph_chunk_generation_data_overflow(struct hashfile *f,\n+\t\t\t\t\t\t       struct write_commit_graph_context *ctx)\n+{\n+\tint i;\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\ttimestamp_t offset = commit_graph_data_at(c)->generation - c->date;\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n+\n+\t\tif (offset > GENERATION_NUMBER_V2_OFFSET_MAX) {\n+\t\t\thashwrite_be32(f, offset >> 32);\n+\t\t\thashwrite_be32(f, (uint32_t) offset);\n+\t\t}\n+\t}\n+\n+\treturn 0;\n+}\n+\n static int write_graph_chunk_extra_edges(struct hashfile *f,\n \t\t\t\t\t struct write_commit_graph_context *ctx)\n {\n@@ -1386,6 +1457,9 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \t\t\t\tif (current->date && current->date > max_corrected_commit_date)\n \t\t\t\t\tmax_corrected_commit_date = current->date - 1;\n \t\t\t\tcommit_graph_data_at(current)->generation = max_corrected_commit_date + 1;\n+\n+\t\t\t\tif (commit_graph_data_at(current)->generation - current->date > GENERATION_NUMBER_V2_OFFSET_MAX)\n+\t\t\t\t\tctx->num_generation_data_overflows++;\n \t\t\t}\n \t\t}\n \t}\n@@ -1719,6 +1793,21 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tchunks[2].id = GRAPH_CHUNKID_DATA;\n \tchunks[2].size = (hashsz + 16) * ctx->commits.nr;\n \tchunks[2].write_fn = write_graph_chunk_data;\n+\n+\tif (git_env_bool(GIT_TEST_COMMIT_GRAPH_NO_GDAT, 0))\n+\t\tctx->write_generation_data = 0;\n+\tif (ctx->write_generation_data) {\n+\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA;\n+\t\tchunks[num_chunks].size = sizeof(uint32_t) * ctx->commits.nr;\n+\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data;\n+\t\tnum_chunks++;\n+\t}\n+\tif (ctx->num_generation_data_overflows) {\n+\t\tchunks[num_chunks].id = GRAPH_CHUNKID_GENERATION_DATA_OVERFLOW;\n+\t\tchunks[num_chunks].size = sizeof(timestamp_t) * ctx->num_generation_data_overflows;\n+\t\tchunks[num_chunks].write_fn = write_graph_chunk_generation_data_overflow;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_extra_edges) {\n \t\tchunks[num_chunks].id = GRAPH_CHUNKID_EXTRAEDGES;\n \t\tchunks[num_chunks].size = 4 * ctx->num_extra_edges;\n@@ -2139,6 +2228,8 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n \tctx->opts = opts;\n \tctx->total_bloom_filter_data_size = 0;\n+\tctx->write_generation_data = 1;\n+\tctx->num_generation_data_overflows = 0;\n \n \tbloom_settings.bits_per_entry = git_env_ulong(\"GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY\",\n \t\t\t\t\t\t      bloom_settings.bits_per_entry);\n@@ -2445,16 +2536,17 @@ int verify_commit_graph(struct repository *r, struct commit_graph *g, int flags)\n \t\t\tcontinue;\n \n \t\t/*\n-\t\t * If one of our parents has generation GENERATION_NUMBER_V1_MAX, then\n-\t\t * our generation is also GENERATION_NUMBER_V1_MAX. Decrement to avoid\n-\t\t * extra logic in the following condition.\n+\t\t * If we are using topological level and one of our parents has\n+\t\t * generation GENERATION_NUMBER_V1_MAX, then our generation is\n+\t\t * also GENERATION_NUMBER_V1_MAX. Decrement to avoid extra logic\n+\t\t * in the following condition.\n \t\t */\n-\t\tif (max_generation == GENERATION_NUMBER_V1_MAX)\n+\t\tif (!g->chunk_generation_data && max_generation == GENERATION_NUMBER_V1_MAX)\n \t\t\tmax_generation--;\n \n \t\tgeneration = commit_graph_generation(graph_commit);\n-\t\tif (generation != max_generation + 1)\n-\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" != %\"PRItime),\n+\t\tif (generation < max_generation + 1)\n+\t\t\tgraph_report(_(\"commit-graph generation for commit %s is %\"PRItime\" < %\"PRItime),\n \t\t\t\t     oid_to_hex(&cur_oid),\n \t\t\t\t     generation,\n \t\t\t\t     max_generation + 1);\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 2e9aa7824ee..19a02001fde 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -6,6 +6,7 @@\n #include \"oidset.h\"\n \n #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n+#define GIT_TEST_COMMIT_GRAPH_NO_GDAT \"GIT_TEST_COMMIT_GRAPH_NO_GDAT\"\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE \"GIT_TEST_COMMIT_GRAPH_DIE_ON_PARSE\"\n #define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n \n@@ -68,6 +69,8 @@ struct commit_graph {\n \tconst uint32_t *chunk_oid_fanout;\n \tconst unsigned char *chunk_oid_lookup;\n \tconst unsigned char *chunk_commit_data;\n+\tconst unsigned char *chunk_generation_data;\n+\tconst unsigned char *chunk_generation_data_overflow;\n \tconst unsigned char *chunk_extra_edges;\n \tconst unsigned char *chunk_base_graphs;\n \tconst unsigned char *chunk_bloom_indexes;\ndiff --git a/commit.h b/commit.h\nindex 742d96c41e8..eff94f3f7c2 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -14,6 +14,7 @@\n #define GENERATION_NUMBER_INFINITY ((1ULL << 63) - 1)\n #define GENERATION_NUMBER_V1_MAX 0x3FFFFFFF\n #define GENERATION_NUMBER_ZERO 0\n+#define GENERATION_NUMBER_V2_OFFSET_MAX ((1ULL << 31) - 1)\n \n struct commit_list {\n \tstruct commit *item;\ndiff --git a/t/README b/t/README\nindex c730a707705..8a121487279 100644\n--- a/t/README\n+++ b/t/README\n@@ -393,6 +393,9 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_COMMIT_GRAPH_NO_GDAT=<boolean>, when true, forces the\n+commit-graph to be written without generation data chunk.\n+\n GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n commit-graph write to compute and write changed path Bloom filters for\n every 'git commit-graph write', as if the `--changed-paths` option was\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex 5f585a17256..75927b2c81d 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -33,6 +33,10 @@ int cmd__read_graph(int argc, const char **argv)\n \t\tprintf(\" oid_lookup\");\n \tif (graph->chunk_commit_data)\n \t\tprintf(\" commit_metadata\");\n+\tif (graph->chunk_generation_data)\n+\t\tprintf(\" generation_data\");\n+\tif (graph->chunk_generation_data_overflow)\n+\t\tprintf(\" generation_data_overflow\");\n \tif (graph->chunk_extra_edges)\n \t\tprintf(\" extra_edges\");\n \tif (graph->chunk_bloom_indexes)\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 0f16c4b9d52..50f206db550 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -43,11 +43,11 @@ test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n '\n \n graph_read_expect () {\n-\tNUM_CHUNKS=5\n+\tNUM_CHUNKS=6\n \tcat >expect <<- EOF\n \theader: 43475048 1 $(test_oid oid_version) $NUM_CHUNKS 0\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data bloom_indexes bloom_data\n \tEOF\n \ttest-tool read-graph >actual &&\n \ttest_cmp expect actual\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 2ed0c1544da..fa27df579a5 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -76,7 +76,7 @@ graph_git_behavior 'no graph' full commits/3 commits/1\n graph_read_expect() {\n \tOPTIONAL=\"\"\n \tNUM_CHUNKS=3\n-\tif test ! -z $2\n+\tif test ! -z \"$2\"\n \tthen\n \t\tOPTIONAL=\" $2\"\n \t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n@@ -103,14 +103,14 @@ test_expect_success 'exit with correct error on bad input to --stdin-commits' '\n \t# valid commit and tree OID\n \tgit rev-parse HEAD HEAD^{tree} >in &&\n \tgit commit-graph write --stdin-commits <in &&\n-\tgraph_read_expect 3\n+\tgraph_read_expect 3 generation_data\n '\n \n test_expect_success 'write graph' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"3\"\n+\tgraph_read_expect \"3\" generation_data\n '\n \n test_expect_success POSIXPERM 'write graph has correct permissions' '\n@@ -219,7 +219,7 @@ test_expect_success 'write graph with merges' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'merge 1 vs 2' full merge/1 merge/2\n@@ -254,7 +254,7 @@ test_expect_success 'write graph with new commit' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'full graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -264,7 +264,7 @@ test_expect_success 'write graph with nothing new' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'cleared graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -274,7 +274,7 @@ test_expect_success 'build graph from latest pack with closure' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tcat new-idx | git commit-graph write --stdin-packs &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"9\" \"extra_edges\"\n+\tgraph_read_expect \"9\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'graph from pack, commit 8 vs merge 1' full commits/8 merge/1\n@@ -287,7 +287,7 @@ test_expect_success 'build graph from commits with closure' '\n \tgit rev-parse merge/1 >>commits-in &&\n \tcat commits-in | git commit-graph write --stdin-commits &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"6\"\n+\tgraph_read_expect \"6\" \"generation_data\"\n '\n \n graph_git_behavior 'graph from commits, commit 8 vs merge 1' full commits/8 merge/1\n@@ -297,7 +297,7 @@ test_expect_success 'build graph from commits with append' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit rev-parse merge/3 | git commit-graph write --stdin-commits --append &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"10\" \"extra_edges\"\n+\tgraph_read_expect \"10\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -307,7 +307,7 @@ test_expect_success 'build graph using --reachable' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n \tgit commit-graph write --reachable &&\n \ttest_path_is_file $objdir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'append graph, commit 8 vs merge 1' full commits/8 merge/1\n@@ -328,7 +328,7 @@ test_expect_success 'write graph in bare repo' '\n \tcd \"$TRASH_DIRECTORY/bare\" &&\n \tgit commit-graph write &&\n \ttest_path_is_file $baredir/info/commit-graph &&\n-\tgraph_read_expect \"11\" \"extra_edges\"\n+\tgraph_read_expect \"11\" \"generation_data extra_edges\"\n '\n \n graph_git_behavior 'bare repo with graph, commit 8 vs merge 1' bare commits/8 merge/1\n@@ -454,8 +454,9 @@ test_expect_success 'warn on improper hash version' '\n \n test_expect_success 'git commit-graph verify' '\n \tcd \"$TRASH_DIRECTORY/full\" &&\n-\tgit rev-parse commits/8 | git commit-graph write --stdin-commits &&\n-\tgit commit-graph verify >output\n+\tgit rev-parse commits/8 | GIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --stdin-commits &&\n+\tgit commit-graph verify >output &&\n+\tgraph_read_expect 9 extra_edges\n '\n \n NUM_COMMITS=9\n@@ -741,4 +742,56 @@ test_expect_success 'corrupt commit-graph write (missing tree)' '\n \t)\n '\n \n+# We test the overflow-related code with the following repo history:\n+#\n+#               4:F - 5:N - 6:U\n+#              /                \\\n+# 1:U - 2:N - 3:U                M:N\n+#              \\                /\n+#               7:N - 8:F - 9:N\n+#\n+# Here the commits denoted by U have committer date of zero seconds\n+# since Unix epoch, the commits denoted by N have committer date\n+# starting from 1112354055 seconds since Unix epoch (default committer\n+# date for the test suite), and the commits denoted by F have committer\n+# date of (2 ^ 31 - 2) seconds since Unix epoch.\n+#\n+# The largest offset observed is 2 ^ 31, just large enough to overflow.\n+#\n+\n+test_expect_success 'set up and verify repo with generation data overflow chunk' '\n+\tobjdir=\".git/objects\" &&\n+\tUNIX_EPOCH_ZERO=\"@0 +0000\" &&\n+\tFUTURE_DATE=\"@2147483646 +0000\" &&\n+\ttest_oid_cache <<-EOF &&\n+\toid_version sha1:1\n+\toid_version sha256:2\n+\tEOF\n+\tcd \"$TRASH_DIRECTORY\" &&\n+\tmkdir repo &&\n+\tcd repo &&\n+\tgit init &&\n+\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 1 &&\n+\ttest_commit 2 &&\n+\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 3 &&\n+\tgit commit-graph write --reachable &&\n+\tgraph_read_expect 3 generation_data &&\n+\ttest_commit --date \"$FUTURE_DATE\" 4 &&\n+\ttest_commit 5 &&\n+\ttest_commit --date \"$UNIX_EPOCH_ZERO\" 6 &&\n+\tgit branch left &&\n+\tgit reset --hard 3 &&\n+\ttest_commit 7 &&\n+\ttest_commit --date \"$FUTURE_DATE\" 8 &&\n+\ttest_commit 9 &&\n+\tgit branch right &&\n+\tgit reset --hard 3 &&\n+\ttest_merge M left right &&\n+\tgit commit-graph write --reachable &&\n+\tgraph_read_expect 10 \"generation_data generation_data_overflow\" &&\n+\tgit commit-graph verify\n+'\n+\n+graph_git_behavior 'generation data overflow chunk repo' repo left right\n+\n test_done\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 4d3842b83b9..587757b62d9 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -13,11 +13,11 @@ test_expect_success 'setup repo' '\n \tinfodir=\".git/objects/info\" &&\n \tgraphdir=\"$infodir/commit-graphs\" &&\n \ttest_oid_cache <<-EOM\n-\tshallow sha1:1760\n-\tshallow sha256:2064\n+\tshallow sha1:2132\n+\tshallow sha256:2436\n \n-\tbase sha1:1376\n-\tbase sha256:1496\n+\tbase sha1:1408\n+\tbase sha256:1528\n \n \toid_version sha1:1\n \toid_version sha256:2\n@@ -31,9 +31,9 @@ graph_read_expect() {\n \t\tNUM_BASE=$2\n \tfi\n \tcat >expect <<- EOF\n-\theader: 43475048 1 $(test_oid oid_version) 3 $NUM_BASE\n+\theader: 43475048 1 $(test_oid oid_version) 4 $NUM_BASE\n \tnum_commits: $1\n-\tchunks: oid_fanout oid_lookup commit_metadata\n+\tchunks: oid_fanout oid_lookup commit_metadata generation_data\n \tEOF\n \ttest-tool read-graph >output &&\n \ttest_cmp expect output\ndiff --git a/t/t6600-test-reach.sh b/t/t6600-test-reach.sh\nindex af10f0dc090..e2d33a8a4c4 100755\n--- a/t/t6600-test-reach.sh\n+++ b/t/t6600-test-reach.sh\n@@ -55,6 +55,9 @@ test_expect_success 'setup' '\n \tgit show-ref -s commit-5-5 | git commit-graph write --stdin-commits &&\n \tmv .git/objects/info/commit-graph commit-graph-half &&\n \tchmod u+w commit-graph-half &&\n+\tGIT_TEST_COMMIT_GRAPH_NO_GDAT=1 git commit-graph write --reachable &&\n+\tmv .git/objects/info/commit-graph commit-graph-no-gdat &&\n+\tchmod u+w commit-graph-no-gdat &&\n \tgit config core.commitGraph true\n '\n \n@@ -67,6 +70,9 @@ run_all_modes () {\n \ttest_cmp expect actual &&\n \tcp commit-graph-half .git/objects/info/commit-graph &&\n \t\"$@\" <input >actual &&\n+\ttest_cmp expect actual &&\n+\tcp commit-graph-no-gdat .git/objects/info/commit-graph &&\n+\t\"$@\" <input >actual &&\n \ttest_cmp expect actual\n }\n \ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex 6bca0023168..df5bba07295 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -218,6 +218,12 @@ test_commit () {\n \t\t--signoff)\n \t\t\tsignoff=\"$1\"\n \t\t\t;;\n+\t\t--date)\n+\t\t\tnotick=yes\n+\t\t\tGIT_COMMITTER_DATE=\"$2\"\n+\t\t\tGIT_AUTHOR_DATE=\"$2\"\n+\t\t\tshift\n+\t\t\t;;\n \t\t-C)\n \t\t\tindir=\"$2\"\n \t\t\tshift\n-- \ngitgitgadget\n\n"},{"id":"415721","messageId":"8647b5d2e38d7c00b4b189b5cba27cc4f61b778c.1612162726.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"[PATCH v7 07/11] commit-graph: document generation number v2","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:41Z","receivedAt":"2021-02-01T07:00:48Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\nGit uses topological levels in the commit-graph file for commit-graph\ntraversal operations like 'git log --graph'. Unfortunately, topological\nlevels can perform worse than committer date when parents of a commit\ndiffer greatly in generation numbers [1]. For example, 'git merge-base\nv4.8 v4.9' on the Linux repository walks 635,579 commits using\ntopological levels and walks 167,468 using committer date. Since\n091f4cf3 (commit: don't use generation numbers if not needed,\n2018-08-30), 'git merge-base' uses committer date heuristic unless there\nis a cutoff because of the performance hit.\n\n[1] https://lore.kernel.org/git/efa3720fb40638e5d61c6130b55e3348d8e4339e.1535633886.git.gitgitgadget@gmail.com/\n\nThus, the need for generation number v2 was born. As Git used to die\nwhen graph version understood by it and in the commit-graph file are\ndifferent [2], we needed a way to distinguish between the old and new\ngeneration number without incrementing the graph version.\n\n[2] https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n\nThe following candidates were proposed (https://github.com/derrickstolee/gen-test,\nhttps://github.com/abhishekkumar2718/git/pull/1):\n- (Epoch, Date) Pairs.\n- Maximum Generation Numbers.\n- Corrected Commit Date.\n- FELINE Index.\n- Corrected Commit Date with Monotonically Increasing Offsets.\n\nBased on performance, local computability, and immutability (along with\nthe introduction of an additional commit-graph chunk which relieved the\nrequirement of backwards-compatibility) Corrected Commit Date was chosen\nas generation number v2 and is defined as follows:\n\nFor a commit C, let its corrected commit date  be the maximum of the\ncommit date of C and the corrected commit dates of its parents plus 1.\nThen corrected commit date offset is the difference between corrected\ncommit date of C and commit date of C. As a special case, a root commit\nwith the timestamp zero has corrected commit date of 1 to distinguish it\nfrom GENERATION_NUMBER_ZERO (that is, an uncomputed generation number).\n\nWhile it was proposed initially to store corrected commit date offsets\nwithin Commit Data Chunk, storing the offsets in a new chunk did not\naffect the performance measurably. The new chunk is \"Generation DATa\n(GDAT) chunk\" and it stores corrected commit date offsets while CDAT\nchunk stores topological level. The old versions of Git would ignore\nGDAT chunk, using topological levels from CDAT chunk. In contrast, new\nversions of Git would use corrected commit dates, falling back to\ntopological level if the generation data chunk is absent in the\ncommit-graph file.\n\nWhile storing corrected commit date offsets saves us 4 bytes per commit\n(as compared with storing corrected commit dates directly), it's however\npossible for the offset to overflow the space allocated. To handle such\ncases, we introduce a new chunk, _Generation Data Overflow_ (GDOV) that\nstores the corrected commit date. For overflowing offsets, we set MSB\nand store the position into the GDOV chunk, in a mechanism similar to\nthe Extra Edges list chunk.\n\nFor mixed generation number environment (for example new Git on the\ncommand line, old Git used by GUI client), we can encounter a\nmixed-chain commit-graph (a commit-graph chain where some of split\ncommit-graph files have GDAT chunk and others do not). As backward\ncompatibility is one of the goals, we can define the following behavior:\n\nWhile reading a mixed-chain commit-graph version, we fall back on\ntopological levels as corrected commit dates and topological levels\ncannot be compared directly.\n\nWhen adding new layer to the split commit-graph file, and when merging\nsome or all layers (replacing them in the latter case), the new layer\nwill have GDAT chunk if and only if in the final result there would be\nno layer without GDAT chunk just below it.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n .../technical/commit-graph-format.txt         | 28 +++++--\n Documentation/technical/commit-graph.txt      | 77 +++++++++++++++----\n 2 files changed, 86 insertions(+), 19 deletions(-)\n\ndiff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\nindex b3b58880b92..b6658eff188 100644\n--- a/Documentation/technical/commit-graph-format.txt\n+++ b/Documentation/technical/commit-graph-format.txt\n@@ -4,11 +4,7 @@ Git commit graph format\n The Git commit graph stores a list of commit OIDs and some associated\n metadata, including:\n \n-- The generation number of the commit. Commits with no parents have\n-  generation number 1; commits with parents have generation number\n-  one more than the maximum generation number of its parents. We\n-  reserve zero as special, and can be used to mark a generation\n-  number invalid or as \"not computed\".\n+- The generation number of the commit.\n \n - The root tree OID.\n \n@@ -86,13 +82,33 @@ CHUNK DATA:\n       position. If there are more than two parents, the second value\n       has its most-significant bit on and the other bits store an array\n       position into the Extra Edge List chunk.\n-    * The next 8 bytes store the generation number of the commit and\n+    * The next 8 bytes store the topological level (generation number v1)\n+      of the commit and\n       the commit time in seconds since EPOCH. The generation number\n       uses the higher 30 bits of the first 4 bytes, while the commit\n       time uses the 32 bits of the second 4 bytes, along with the lowest\n       2 bits of the lowest byte, storing the 33rd and 34th bit of the\n       commit time.\n \n+  Generation Data (ID: {'G', 'D', 'A', 'T' }) (N * 4 bytes) [Optional]\n+    * This list of 4-byte values store corrected commit date offsets for the\n+      commits, arranged in the same order as commit data chunk.\n+    * If the corrected commit date offset cannot be stored within 31 bits,\n+      the value has its most-significant bit on and the other bits store\n+      the position of corrected commit date into the Generation Data Overflow\n+      chunk.\n+    * Generation Data chunk is present only when commit-graph file is written\n+      by compatible versions of Git and in case of split commit-graph chains,\n+      the topmost layer also has Generation Data chunk.\n+\n+  Generation Data Overflow (ID: {'G', 'D', 'O', 'V' }) [Optional]\n+    * This list of 8-byte values stores the corrected commit date offsets\n+      for commits with corrected commit date offsets that cannot be\n+      stored within 31 bits.\n+    * Generation Data Overflow chunk is present only when Generation Data\n+      chunk is present and atleast one corrected commit date offset cannot\n+      be stored within 31 bits.\n+\n   Extra Edge List (ID: {'E', 'D', 'G', 'E'}) [Optional]\n       This list of 4-byte values store the second through nth parents for\n       all octopus merges. The second parent value in the commit data stores\ndiff --git a/Documentation/technical/commit-graph.txt b/Documentation/technical/commit-graph.txt\nindex f14a7659aa8..f05e7bda1a9 100644\n--- a/Documentation/technical/commit-graph.txt\n+++ b/Documentation/technical/commit-graph.txt\n@@ -38,14 +38,31 @@ A consumer may load the following info for a commit from the graph:\n \n Values 1-4 satisfy the requirements of parse_commit_gently().\n \n-Define the \"generation number\" of a commit recursively as follows:\n+There are two definitions of generation number:\n+1. Corrected committer dates (generation number v2)\n+2. Topological levels (generation nummber v1)\n \n- * A commit with no parents (a root commit) has generation number one.\n+Define \"corrected committer date\" of a commit recursively as follows:\n \n- * A commit with at least one parent has generation number one more than\n-   the largest generation number among its parents.\n+ * A commit with no parents (a root commit) has corrected committer date\n+    equal to its committer date.\n \n-Equivalently, the generation number of a commit A is one more than the\n+ * A commit with at least one parent has corrected committer date equal to\n+    the maximum of its commiter date and one more than the largest corrected\n+    committer date among its parents.\n+\n+ * As a special case, a root commit with timestamp zero has corrected commit\n+    date of 1, to be able to distinguish it from GENERATION_NUMBER_ZERO\n+    (that is, an uncomputed corrected commit date).\n+\n+Define the \"topological level\" of a commit recursively as follows:\n+\n+ * A commit with no parents (a root commit) has topological level of one.\n+\n+ * A commit with at least one parent has topological level one more than\n+   the largest topological level among its parents.\n+\n+Equivalently, the topological level of a commit A is one more than the\n length of a longest path from A to a root commit. The recursive definition\n is easier to use for computation and observing the following property:\n \n@@ -60,6 +77,9 @@ is easier to use for computation and observing the following property:\n     generation numbers, then we always expand the boundary commit with highest\n     generation number and can easily detect the stopping condition.\n \n+The property applies to both versions of generation number, that is both\n+corrected committer dates and topological levels.\n+\n This property can be used to significantly reduce the time it takes to\n walk commits and determine topological relationships. Without generation\n numbers, the general heuristic is the following:\n@@ -67,7 +87,9 @@ numbers, the general heuristic is the following:\n     If A and B are commits with commit time X and Y, respectively, and\n     X < Y, then A _probably_ cannot reach B.\n \n-This heuristic is currently used whenever the computation is allowed to\n+In absence of corrected commit dates (for example, old versions of Git or\n+mixed generation graph chains),\n+this heuristic is currently used whenever the computation is allowed to\n violate topological relationships due to clock skew (such as \"git log\"\n with default order), but is not used when the topological order is\n required (such as merge base calculations, \"git log --graph\").\n@@ -77,7 +99,7 @@ in the commit graph. We can treat these commits as having \"infinite\"\n generation number and walk until reaching commits with known generation\n number.\n \n-We use the macro GENERATION_NUMBER_INFINITY = 0xFFFFFFFF to mark commits not\n+We use the macro GENERATION_NUMBER_INFINITY to mark commits not\n in the commit-graph file. If a commit-graph file was written by a version\n of Git that did not compute generation numbers, then those commits will\n have generation number represented by the macro GENERATION_NUMBER_ZERO = 0.\n@@ -93,12 +115,12 @@ fully-computed generation numbers. Using strict inequality may result in\n walking a few extra commits, but the simplicity in dealing with commits\n with generation number *_INFINITY or *_ZERO is valuable.\n \n-We use the macro GENERATION_NUMBER_MAX = 0x3FFFFFFF to for commits whose\n-generation numbers are computed to be at least this value. We limit at\n-this value since it is the largest value that can be stored in the\n-commit-graph file using the 30 bits available to generation numbers. This\n-presents another case where a commit can have generation number equal to\n-that of a parent.\n+We use the macro GENERATION_NUMBER_V1_MAX = 0x3FFFFFFF for commits whose\n+topological levels (generation number v1) are computed to be at least\n+this value. We limit at this value since it is the largest value that\n+can be stored in the commit-graph file using the 30 bits available\n+to topological levels. This presents another case where a commit can\n+have generation number equal to that of a parent.\n \n Design Details\n --------------\n@@ -267,6 +289,35 @@ The merge strategy values (2 for the size multiple, 64,000 for the maximum\n number of commits) could be extracted into config settings for full\n flexibility.\n \n+## Handling Mixed Generation Number Chains\n+\n+With the introduction of generation number v2 and generation data chunk, the\n+following scenario is possible:\n+\n+1. \"New\" Git writes a commit-graph with the corrected commit dates.\n+2. \"Old\" Git writes a split commit-graph on top without corrected commit dates.\n+\n+A naive approach of using the newest available generation number from\n+each layer would lead to violated expectations: the lower layer would\n+use corrected commit dates which are much larger than the topological\n+levels of the higher layer. For this reason, Git inspects the topmost\n+layer to see if the layer is missing corrected commit dates. In such a case\n+Git only uses topological level for generation numbers.\n+\n+When writing a new layer in split commit-graph, we write corrected commit\n+dates if the topmost layer has corrected commit dates written. This\n+guarantees that if a layer has corrected commit dates, all lower layers\n+must have corrected commit dates as well.\n+\n+When merging layers, we do not consider whether the merged layers had corrected\n+commit dates. Instead, the new layer will have corrected commit dates if the\n+layer below the new layer has corrected commit dates.\n+\n+While writing or merging layers, if the new layer is the only layer, it will\n+have corrected commit dates when written by compatible versions of Git. Thus,\n+rewriting split commit-graph as a single file (`--split=replace`) creates a\n+single layer with corrected commit dates.\n+\n ## Deleting graph-{hash} files\n \n After a new tip file is written, some `graph-{hash}` files may no longer\n-- \ngitgitgadget\n\n"},{"id":"415722","messageId":"523e2d4a902b22fb7e0234bf9c207f7b071b7187.1612162726.git.gitgitgadget@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"[PATCH v7 11/11] commit-reach: use corrected commit dates in paint_down_to_common()","fromName":"Abhishek Kumar via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-02-01T06:58:45Z","receivedAt":"2021-02-01T07:00:53Z","isPatch":true,"sender":{"key":"abhishekkumar8222@gmail.com","avatar":"https://avatars.githubusercontent.com/u/31231064?v=4"},"body":"From: Abhishek Kumar <abhishekkumar8222@gmail.com>\n\n091f4cf (commit: don't use generation numbers if not needed,\n2018-08-30) changed paint_down_to_common() to use commit dates instead\nof generation numbers v1 (topological levels) as the performance\nregressed on certain topologies. With generation number v2 (corrected\ncommit dates) implemented, we no longer have to rely on commit dates and\ncan use generation numbers.\n\nFor example, the command `git merge-base v4.8 v4.9` on the Linux\nrepository walks 167468 commits, taking 0.135s for committer date and\n167496 commits, taking 0.157s for corrected committer date respectively.\n\nWhile using corrected commit dates, Git walks nearly the same number of\ncommits as commit date, the process is slower as for each comparision we\nhave to access a commit-slab (for corrected committer date) instead of\naccessing struct member (for committer date).\n\nThis change incidentally broke the fragile t6404-recursive-merge test.\nt6404-recursive-merge sets up a unique repository where all commits have\nthe same committer date without a well-defined merge-base.\n\nWhile running tests with GIT_TEST_COMMIT_GRAPH unset, we use committer\ndate as a heuristic in paint_down_to_common(). 6404.1 'combined merge\nconflicts' merges commits in the order:\n- Merge C with B to form an intermediate commit.\n- Merge the intermediate commit with A.\n\nWith GIT_TEST_COMMIT_GRAPH=1, we write a commit-graph and subsequently\nuse the corrected committer date, which changes the order in which\ncommits are merged:\n- Merge A with B to form an intermediate commit.\n- Merge the intermediate commit with C.\n\nWhile resulting repositories are equivalent, 6404.4 'virtual trees were\nprocessed' fails with GIT_TEST_COMMIT_GRAPH=1 as we are selecting\ndifferent merge-bases and thus have different object ids for the\nintermediate commits.\n\nAs this has already causes problems (as noted in 859fdc0 (commit-graph:\ndefine GIT_TEST_COMMIT_GRAPH, 2018-08-29)), we disable commit graph\nwithin t6404-recursive-merge.\n\nSigned-off-by: Abhishek Kumar <abhishekkumar8222@gmail.com>\n---\n commit-graph.c             | 14 ++++++++++++++\n commit-graph.h             |  6 ++++++\n commit-reach.c             |  2 +-\n t/t6404-recursive-merge.sh |  5 ++++-\n 4 files changed, 25 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 77fef5a240e..bf735fac4ea 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -714,6 +714,20 @@ int generation_numbers_enabled(struct repository *r)\n \treturn !!first_generation;\n }\n \n+int corrected_commit_dates_enabled(struct repository *r)\n+{\n+\tstruct commit_graph *g;\n+\tif (!prepare_commit_graph(r))\n+\t\treturn 0;\n+\n+\tg = r->objects->commit_graph;\n+\n+\tif (!g->num_commits)\n+\t\treturn 0;\n+\n+\treturn g->read_generation_data;\n+}\n+\n struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r)\n {\n \tstruct commit_graph *g = r->objects->commit_graph;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex ad52130883b..97f3497c279 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -95,6 +95,12 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n  */\n int generation_numbers_enabled(struct repository *r);\n \n+/*\n+ * Return 1 if and only if the repository has a commit-graph\n+ * file and generation data chunk has been written for the file.\n+ */\n+int corrected_commit_dates_enabled(struct repository *r);\n+\n struct bloom_filter_settings *get_bloom_filter_settings(struct repository *r);\n \n enum commit_graph_write_flags {\ndiff --git a/commit-reach.c b/commit-reach.c\nindex 9b24b0378d5..e38771ca5a1 100644\n--- a/commit-reach.c\n+++ b/commit-reach.c\n@@ -39,7 +39,7 @@ static struct commit_list *paint_down_to_common(struct repository *r,\n \tint i;\n \ttimestamp_t last_gen = GENERATION_NUMBER_INFINITY;\n \n-\tif (!min_generation)\n+\tif (!min_generation && !corrected_commit_dates_enabled(r))\n \t\tqueue.compare = compare_commits_by_commit_date;\n \n \tone->object.flags |= PARENT1;\ndiff --git a/t/t6404-recursive-merge.sh b/t/t6404-recursive-merge.sh\nindex c7ab7048f58..eaf48e941e2 100755\n--- a/t/t6404-recursive-merge.sh\n+++ b/t/t6404-recursive-merge.sh\n@@ -18,6 +18,8 @@ GIT_COMMITTER_DATE=\"2006-12-12 23:28:00 +0100\"\n export GIT_COMMITTER_DATE\n \n test_expect_success 'setup tests' '\n+\tGIT_TEST_COMMIT_GRAPH=0 &&\n+\texport GIT_TEST_COMMIT_GRAPH &&\n \techo 1 >a1 &&\n \tgit add a1 &&\n \tGIT_AUTHOR_DATE=\"2006-12-12 23:00:00\" git commit -m 1 a1 &&\n@@ -69,7 +71,7 @@ test_expect_success 'setup tests' '\n '\n \n test_expect_success 'combined merge conflicts' '\n-\ttest_must_fail env GIT_TEST_COMMIT_GRAPH=0 git merge -m final G\n+\ttest_must_fail git merge -m final G\n '\n \n test_expect_success 'result contains a conflict' '\n@@ -85,6 +87,7 @@ test_expect_success 'result contains a conflict' '\n '\n \n test_expect_success 'virtual trees were processed' '\n+\t# TODO: fragile test, relies on ambigious merge-base resolution\n \tgit ls-files --stage >out &&\n \n \tcat >expect <<-EOF &&\n-- \ngitgitgadget\n"},{"id":"415749","messageId":"3c48f860-e743-afbe-63e8-99804036a965@gmail.com","threadId":"53933","inReplyTo":"pull.676.v7.git.1612162726.gitgitgadget@gmail.com","subject":"Re: [PATCH v7 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-02-01T13:14:34Z","receivedAt":"2021-02-01T13:15:34Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 2/1/2021 1:58 AM, Abhishek Kumar via GitGitGadget wrote:\n\n> Changes in version 7:\n> \n>  * Moved the documentation patch ahead of \"commit-graph: implement corrected\n>    commit date\" and elaborated on the introduction of generation number v2.\n\nThe only change in this version is this commit message:\n\n>  11:  e571f03d8bd !  7:  8647b5d2e38 doc: add corrected commit date info\n>      @@ Metadata\n>       Author: Abhishek Kumar <abhishekkumar8222@gmail.com>\n>       \n>        ## Commit message ##\n>      -    doc: add corrected commit date info\n>      +    commit-graph: document generation number v2\n>       \n>      -    With generation data chunk and corrected commit dates implemented, let's\n>      -    update the technical documentation for commit-graph.\n>      +    Git uses topological levels in the commit-graph file for commit-graph\n>      +    traversal operations like 'git log --graph'. Unfortunately, topological\n>      +    levels can perform worse than committer date when parents of a commit\n>      +    differ greatly in generation numbers [1]. For example, 'git merge-base\n>      +    v4.8 v4.9' on the Linux repository walks 635,579 commits using\n>      +    topological levels and walks 167,468 using committer date. Since\n>      +    091f4cf3 (commit: don't use generation numbers if not needed,\n>      +    2018-08-30), 'git merge-base' uses committer date heuristic unless there\n>      +    is a cutoff because of the performance hit.\n>      +\n>      +    [1] https://lore.kernel.org/git/efa3720fb40638e5d61c6130b55e3348d8e4339e.1535633886.git.gitgitgadget@gmail.com/\n>      +\n>      +    Thus, the need for generation number v2 was born. As Git used to die\n>      +    when graph version understood by it and in the commit-graph file are\n>      +    different [2], we needed a way to distinguish between the old and new\n>      +    generation number without incrementing the graph version.\n>      +\n>      +    [2] https://lore.kernel.org/git/87a7gdspo4.fsf@evledraar.gmail.com/\n>      +\n>      +    The following candidates were proposed (https://github.com/derrickstolee/gen-test,\n>      +    https://github.com/abhishekkumar2718/git/pull/1):\n>      +    - (Epoch, Date) Pairs.\n>      +    - Maximum Generation Numbers.\n>      +    - Corrected Commit Date.\n>      +    - FELINE Index.\n>      +    - Corrected Commit Date with Monotonically Increasing Offsets.\n>      +\n>      +    Based on performance, local computability, and immutability (along with\n>      +    the introduction of an additional commit-graph chunk which relieved the\n>      +    requirement of backwards-compatibility) Corrected Commit Date was chosen\n>      +    as generation number v2 and is defined as follows:\n>      +\n>      +    For a commit C, let its corrected commit date  be the maximum of the\n>      +    commit date of C and the corrected commit dates of its parents plus 1.\n>      +    Then corrected commit date offset is the difference between corrected\n>      +    commit date of C and commit date of C. As a special case, a root commit\n>      +    with the timestamp zero has corrected commit date of 1 to distinguish it\n>      +    from GENERATION_NUMBER_ZERO (that is, an uncomputed generation number).\n>      +\n>      +    While it was proposed initially to store corrected commit date offsets\n>      +    within Commit Data Chunk, storing the offsets in a new chunk did not\n>      +    affect the performance measurably. The new chunk is \"Generation DATa\n>      +    (GDAT) chunk\" and it stores corrected commit date offsets while CDAT\n>      +    chunk stores topological level. The old versions of Git would ignore\n>      +    GDAT chunk, using topological levels from CDAT chunk. In contrast, new\n>      +    versions of Git would use corrected commit dates, falling back to\n>      +    topological level if the generation data chunk is absent in the\n>      +    commit-graph file.\n>      +\n>      +    While storing corrected commit date offsets saves us 4 bytes per commit\n>      +    (as compared with storing corrected commit dates directly), it's however\n>      +    possible for the offset to overflow the space allocated. To handle such\n>      +    cases, we introduce a new chunk, _Generation Data Overflow_ (GDOV) that\n>      +    stores the corrected commit date. For overflowing offsets, we set MSB\n>      +    and store the position into the GDOV chunk, in a mechanism similar to\n>      +    the Extra Edges list chunk.\n>      +\n>      +    For mixed generation number environment (for example new Git on the\n>      +    command line, old Git used by GUI client), we can encounter a\n>      +    mixed-chain commit-graph (a commit-graph chain where some of split\n>      +    commit-graph files have GDAT chunk and others do not). As backward\n>      +    compatibility is one of the goals, we can define the following behavior:\n>      +\n>      +    While reading a mixed-chain commit-graph version, we fall back on\n>      +    topological levels as corrected commit dates and topological levels\n>      +    cannot be compared directly.\n>      +\n>      +    When adding new layer to the split commit-graph file, and when merging\n>      +    some or all layers (replacing them in the latter case), the new layer\n>      +    will have GDAT chunk if and only if in the final result there would be\n>      +    no layer without GDAT chunk just below it.\n\nWhile that is a quality message, v6 has landed in 'next' and I've begun\nworking off of that version. As Taylor attempted to say [1], this topic\nshould be considered final and updates should be follow-ups on top.\n\n[1] https://lore.kernel.org/git/YBYLwpKdUfxCNwaz@nand.local/\n\n(Of course, if Junio says differently, then listen to him.)\n\nThanks,\n-Stolee\n\n"},{"id":"415784","messageId":"xmqq7dnrmypw.fsf@gitster.c.googlers.com","threadId":"53933","inReplyTo":"3c48f860-e743-afbe-63e8-99804036a965@gmail.com","subject":"Re: [PATCH v7 00/11] [GSoC] Implement Corrected Commit Date","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-02-01T18:26:03Z","receivedAt":"2021-02-01T18:27:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n>>      +    When adding new layer to the split commit-graph file, and when merging\n>>      +    some or all layers (replacing them in the latter case), the new layer\n>>      +    will have GDAT chunk if and only if in the final result there would be\n>>      +    no layer without GDAT chunk just below it.\n>\n> While that is a quality message, v6 has landed in 'next' and I've begun\n> working off of that version. As Taylor attempted to say [1], this topic\n> should be considered final and updates should be follow-ups on top.\n>\n> [1] https://lore.kernel.org/git/YBYLwpKdUfxCNwaz@nand.local/\n\nSounds sensible, modulo s/final/solid enough/ ;-)\n\nI would imagine that the \"quality message\" has something of value to\nkeep to help future developers, and if that is the case, a follow-up\npatch to add to the Documentation/technical/ would be appropriate.\n\nThanks all, for a quality series.\n"}]}