{"thread":{"id":"58758","subject":"[PATCH 01/30] hashfile: allow skipping the hash function","startedAt":"2022-11-07T18:36:12Z","lastAt":"2022-12-02T18:28:54Z","messageCount":56,"participants":["Derrick Stolee via GitGitGadget","Derrick Stolee","Elijah Newren","Ævar Arnfjörð Bjarmason","Junio C Hamano","Taylor Blau","Han-Wen Nienhuys","Phillip Wood","Sean Allred"],"isPatch":true,"patchVersion":1,"patchTotal":30},"messages":[{"id":"466683","messageId":"71c76d4ccbe577f82e820fb08fe93e5177177804.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 01/30] hashfile: allow skipping the hash function","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:35Z","receivedAt":"2022-11-07T18:36:12Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe hashfile API is useful for generating files that include a trailing\nhash of the file's contents up to that point. Using such a hash is\nhelpful for verifying the file for corruption-at-rest, such as a faulty\ndrive causing flipped bits.\n\nSince the commit-graph and multi-pack-index files both use this trailing\nhash, the chunk-format API uses a 'struct hashfile' to handle the I/O to\nthe file. This was very convenient to allow using the hashfile methods\nduring these operations.\n\nHowever, hashing the file contents during write comes at a performance\npenalty. It's slower to hash the bytes on their way to the disk than\nwithout that step. If we wish to use the chunk-format API to upgrade\nother file types, then this hashing is a performance penalty that might\nnot be worth the benefit of a trailing hash.\n\nFor example, if we create a chunk-format version of the packed-refs\nfile, then the file format could shrink by using raw object IDs instead\nof hexadecimal representations in ASCII. That reduction in size is not\nenough to counteract the performance penalty of hashing the file\ncontents. In cases such as deleting a reference that appears in the\npacked-refs file, that write-time performance is critical. This is in\ncontrast to the commit-graph and multi-pack-index files which are mainly\nupdated in non-critical paths such as background maintenance.\n\nOne way to allow future chunked formats to not suffer this penalty would\nbe to create an abstraction layer around the 'struct hashfile' using a\nvtable of function pointers. This would allow placing a different\nrepresentation in place of the hashfile. This option would be cumbersome\nfor a few reasons. First, the hashfile's buffered writes are already\nhighly optimized and would need to be duplicated in another code path.\nThe second is that the chunk-format API calls the chunk_write_fn\npointers using a hashfile. If we change that to an abstraction layer,\nthen those that _do_ use the hashfile API would need to change all of\ntheir instances of hashwrite(), hashwrite_be32(), and others to use the\nnew abstraction layer.\n\nInstead, this change opts for a simpler change. Introduce a new\n'skip_hash' option to 'struct hashfile'. When set, the update_fn and\nfinal_fn members of the_hash_algo are skipped. When finalizing the\nhashfile, the trailing hash is replaced with the null hash.\n\nThis use of a trailing null hash would be desireable in either case,\nsince we do not want to special case a file format to have a different\nlength depending on whether it was hashed or not. When the final bytes\nof a file are all zero, we can infer that it was written without\nhashing, and thus that verification is not available as a check for file\nconsistency. This also means that we could easily toggle hashing for any\nfile format we desire. For the commit-graph and multi-pack-index file,\nit may be possible to allow the null hash without incrementing the file\nformat version, since it technically fits the structure of the file\nformat. The only issue is that older versions would trigger a failure\nduring 'git fsck'. For these file formats, we may want to delay such a\nchange until it is justified.\n\nHowever, the index file is written in critical paths. It is also\nfrequently updated, so corruption at rest is less likely to be an issue\nthan in those other file formats. This could be a good candidate to\ncreate an option that skips the hashing operation.\n\nA version of this patch has existed in the microsoft/git fork since\n2017 [1] (the linked commit was rebased in 2018, but the original dates\nback to January 2017). Here, the change to make the index use this fast\npath is delayed until a later change.\n\n[1] https://github.com/microsoft/git/commit/21fed2d91410f45d85279467f21d717a2db45201\n\nCo-authored-by: Kevin Willford <kewillf@microsoft.com>\nSigned-off-by: Kevin Willford <kewillf@microsoft.com>\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n csum-file.c | 14 +++++++++++---\n csum-file.h |  7 +++++++\n 2 files changed, 18 insertions(+), 3 deletions(-)\n\ndiff --git a/csum-file.c b/csum-file.c\nindex 59ef3398ca2..3243473c3d7 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -45,7 +45,8 @@ void hashflush(struct hashfile *f)\n \tunsigned offset = f->offset;\n \n \tif (offset) {\n-\t\tthe_hash_algo->update_fn(&f->ctx, f->buffer, offset);\n+\t\tif (!f->skip_hash)\n+\t\t\tthe_hash_algo->update_fn(&f->ctx, f->buffer, offset);\n \t\tflush(f, f->buffer, offset);\n \t\tf->offset = 0;\n \t}\n@@ -64,7 +65,12 @@ int finalize_hashfile(struct hashfile *f, unsigned char *result,\n \tint fd;\n \n \thashflush(f);\n-\tthe_hash_algo->final_fn(f->buffer, &f->ctx);\n+\n+\tif (f->skip_hash)\n+\t\tmemset(f->buffer, 0, the_hash_algo->rawsz);\n+\telse\n+\t\tthe_hash_algo->final_fn(f->buffer, &f->ctx);\n+\n \tif (result)\n \t\thashcpy(result, f->buffer);\n \tif (flags & CSUM_HASH_IN_STREAM)\n@@ -108,7 +114,8 @@ void hashwrite(struct hashfile *f, const void *buf, unsigned int count)\n \t\t\t * the hashfile's buffer. In this block,\n \t\t\t * f->offset is necessarily zero.\n \t\t\t */\n-\t\t\tthe_hash_algo->update_fn(&f->ctx, buf, nr);\n+\t\t\tif (!f->skip_hash)\n+\t\t\t\tthe_hash_algo->update_fn(&f->ctx, buf, nr);\n \t\t\tflush(f, buf, nr);\n \t\t} else {\n \t\t\t/*\n@@ -153,6 +160,7 @@ static struct hashfile *hashfd_internal(int fd, const char *name,\n \tf->tp = tp;\n \tf->name = name;\n \tf->do_crc = 0;\n+\tf->skip_hash = 0;\n \tthe_hash_algo->init_fn(&f->ctx);\n \n \tf->buffer_len = buffer_len;\ndiff --git a/csum-file.h b/csum-file.h\nindex 0d29f528fbc..29468067f81 100644\n--- a/csum-file.h\n+++ b/csum-file.h\n@@ -20,6 +20,13 @@ struct hashfile {\n \tsize_t buffer_len;\n \tunsigned char *buffer;\n \tunsigned char *check_buffer;\n+\n+\t/**\n+\t * If set to 1, skip_hash indicates that we should\n+\t * not actually compute the hash for this hashfile and\n+\t * instead only use it as a buffered write.\n+\t */\n+\tunsigned int skip_hash;\n };\n \n /* Checkpoint */\n-- \ngitgitgadget\n\n"},{"id":"466684","messageId":"pull.1408.git.1667846164.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":null,"subject":"[PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:34Z","receivedAt":"2022-11-07T18:36:13Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"\nIntroduction\n============\n\nI became interested in our packed-ref format based on the asymmetry between\nref updates and ref deletions: if we delete a packed ref, then the\npacked-refs file needs to be rewritten. Compared to writing a loose ref,\nthis is an O(N) cost instead of O(1).\n\nIn this way, I set out with some goals:\n\n * (Primary) Make packed ref deletions be nearly as fast as loose ref\n   updates.\n * (Secondary) Allow using a packed ref format for all refs, dropping loose\n   refs and creating a clear way to snapshot all refs at a given point in\n   time.\n\nI also had one major non-goal to keep things focused:\n\n * (Non-goal) Update the reflog format.\n\nAfter carefully considering several options, it seemed that there are two\nsolutions that can solve this effectively:\n\n 1. Wait for reftable to be integrated into Git.\n 2. Update the packed-refs backend to have a stacked version.\n\nThe reftable work seems currently dormant. The format is pretty complicated\nand I have a difficult time seeing a way forward for it to be fully\nintegrated into Git. Personally, I'd prefer a more incremental approach with\nformats that are built for a basic filesystem. During the process, we can\ncreate APIs within Git that can benefit other file formats within Git.\n\nFurther, there is a simpler model that satisfies my primary goal without the\ncomplication required for the secondary goal. Suppose we create a stacked\npacked-refs file but only have two layers: the first (base) layer is created\nwhen git pack-refs collapses the full stack and adds the loose ref updates\nto the packed-refs file; the second (top) layer contains only ref deletions\n(allowing null OIDs to indicate a deleted ref). Then, ref deletions would\nonly need to rewrite that top layer, making ref deletions take O(deletions)\ntime instead of O(all refs) time. With a reasonable schedule to squash the\npacked-refs stack, this would be a dramatic improvement. (A prototype\nimplementation showed that updating a layer of 1,000 deletions takes only\ntwice the time as writing a single loose ref.)\n\nIf we want to satisfy the secondary goal of passing all ref updates through\nthe packed storage, then more complicated layering would be necessary. The\npoint of bringing this up is that we have incremental goals along the way to\nthat final state that give us good stopping points to test the benefits of\neach step.\n\nStacking the packed-refs format introduces several interesting strategy\npoints that are complicated to resolve. Before we can do that, we first need\nto establish a way to modify the ref format of a Git repository. Hence, we\nneed a new extension for the ref formats.\n\nTo simplify the first update to the ref formats, it seemed better to add a\nnew file format version to the existing packed-refs file format. This format\nhas the exact lock/write/rename mechanics of the current packed-refs format,\nbut uses a file format that structures the information in a more compact\nway. It uses the chunk-format API, with some tweaks. This format update is\nuseful to the final goal of a stacked packed-refs API, since each layer will\nhave faster reads and writes. The main reason to do this first is that it is\nmuch simpler to understand the value-add (smaller files means faster\nperformance).\n\n\nRFC Organization\n================\n\nThis RFC is quite long, but the length seemed necessary to actually provide\nand end-to-end implementation that demonstrates the packed-refs v2 format\nalong with test coverage (via the new GIT_TEST_PACKED_REFS_VERSION\nvariable).\n\nFor convenience, I've broken each section of the full RFC into parts, which\nresembles how I intend to submit the pieces for full review. These parts are\navailable as pull requests in my fork, but here is a breakdown:\n\n\nPart I: Optionally hash the index\n=================================\n\n[1] https://github.com/derrickstolee/git/pull/23 Packed-refs v2 Part I:\nOptionally hash the index (Patches 1-2)\n\nThe chunk-format API uses the hashfile API as a buffered write, but also all\nexisting formats that use the chunk-format API also have a trailing hash as\npart of the format. Since the packed-refs file has a critical path involving\nits write speed (deleting a packed ref), it seemed important to allow\napples-to-apples comparison between the v1 and v2 format by skipping the\nhashing. This is later toggled by a config option.\n\nIn this part, the focus is on allowing the hashfile API to ignore updating\nthe hash during the buffered writes. We've been using this in microsoft/git\nto optionally speed up index writes, which patch 2 introduces here. The file\nformat instead writes a null OID which would look like a corrupt file to an\nolder 'git fsck'. Before submitting a full version, I would update 'git\nfsck' to ignore a null OID in all of our file formats that include a\ntrailing hash. Since the index is more short-lived than other formats (such\nas pack-files) this trailing hash is less useful. The write time is also\ncritical as the performance tests demonstrate.\n\n\nPart II: Create extensions.refFormat\n====================================\n\n[2] https://github.com/derrickstolee/git/pull/24 Packed-refs v2 Part II:\ncreate extensions.refFormat (Patches 3-7)\n\nThis part is a critical concept that has yet to be defined in the Git\ncodebase. We have no way to incrementally modify the ref format. Since refs\nare so critical, we cannot add an optionally-understood layer on top (like\nwe did with the multi-pack-index and commit-graph files). The reftable draft\n[6] proposes the same extension name (extensions.refFormat) but focuses\ninstead on only a single value. This means that the reftable must be defined\nat git init or git clone time and cannot be upgraded from the files backend.\n\nIn this RFC, I propose a different model that allows for more customization\nand incremental updates. The extensions.refFormat config key is multi-valued\nand defaults to the list of files and packed. In the context of this RFC,\nthe intention is to be able to add packed-v2 so the list of all three values\nwould allow Git to write and read either file format version (v1 or v2). In\nthe larger scheme, the extension could allow restricting to only loose refs\n(just files) or only packed-refs (just packed) or even later when reftable\nis complete, files and reftable could mean that loose refs are the primary\nref storage, but the reftable format serves as a drop-in replacement for the\npacked-refs file. Not all combinations need to be understood by Git, but\nhaving them available as an option could be useful for flexibility,\nespecially when trying to upgrade existing repositories to new formats.\n\nIn the future, beyond the scope of this RFC, it would be good to add a\nstacked value that allows a stack of files in packed-refs format (whose\nversion is specified by the packed or packed-v2 values) so we can further\nspeed up writes to the packed layer. Depending on how well that works, we\ncould focus on speeding up ref deletions or sending all ref writes straight\nto the packed-refs layer. With the option to keep the loose refs storage, we\nhave flexibility to explore that space incrementally when we have time to\nget to it.\n\n\nPart III: Allow a trailing table-of-contents in the chunk-format API\n====================================================================\n\n[3] https://github.com/derrickstolee/git/pull/25 Packed-refs v2 Part III:\ntrailing table of contents in chunk-format (Patches 8-17)\n\nIn order to optimize the write speed of the packed-refs v2 file format, we\nwant to write immediately to the file as we stream existing refs from the\ncurrent refs. The current chunk-format API requires computing the chunk\nlengths in advance, which can slow down the write and take more memory than\nnecessary. Using a trailing table of contents solves this problem, and was\nrecommended earlier [7]. We just didn't have enough evidence to justify the\nwork to update the existing chunk formats. Here, we update the API in\nadvance of using in the packed-refs v2 format.\n\nWe could consider updating the commit-graph and multi-pack-index formats to\nuse trailing table of contents, but it requires a version bump. That might\nbe worth it in the case of the commit-graph where computing the size of the\nchanged-path Bloom filters chunk requires a lot of memory at the moment.\nAfter this chunk-format API update is reviewed and merged, we can pursue\nthose directions more closely. We would want to investigate the formats more\ncarefully to see if we want to update the chunks themselves as well as some\nheader information.\n\n\nPart IV: Abstract some parts of the v1 file format\n==================================================\n\n[4] https://github.com/derrickstolee/git/pull/26 Packed-refs v2 Part IV:\nabstract some parts of the v1 file format (Patches 18-21)\n\nThese patches move the part of the refs/packed-backend.c file that deal with\nthe specifics of the packed-refs v1 file format into a new file:\nrefs/packed-format-v1.c. This also creates an abstraction layer that will\nallow inserting the v2 format more easily.\n\nOne thing that doesn't exist currently is a documentation file describing\nthe packed-refs file format. I would add that file in this part before\nsubmitting it for full review. (I also haven't written the file format doc\nfor the packed-refs v2 format, either.)\n\n\nPart V: Implement the v2 file format\n====================================\n\n[5] https://github.com/derrickstolee/git/pull/27 Packed-refs v2 Part V: the\nv2 file format (Patches 22-35)\n\nThis is the real meat of the work. Perhaps there are ways to split it\nfurther, but for now this is what I have ready. The very last patch does a\ncomplete performance comparison for a repo with many refs.\n\nThe format is not yet documented, but is broken up into these pieces:\n\n 1. The refs data chunk stores the same data as the packed-refs file, but\n    each ref is broken down as follows: the ref name (with trailing zero),\n    the OID for the ref in its raw bytes, and (if necessary) the peeled OID\n    for the ref in its raw bytes. The refs are sorted lexicographically.\n\n 2. The ref offsets chunk is a single column of 64-bit offsets into the refs\n    chunk indicating where each ref starts. The most-significant bit of that\n    value indicates whether or not there is a peeled OID.\n\n 3. The prefix data chunk lists a set of ref prefixes (currently writes only\n    allow depth-2 prefixes, such as refs/heads/ and refs/tags/). When\n    present, these prefixes are written in this chunk and not in the refs\n    data chunk. The prefixes are sorted lexicographically.\n\n 4. The prefix offset chunk has two 32-bit integer columns. The first column\n    stores the offset within the prefix data chunk to the start of the\n    prefix string. The second column points to the row position for the\n    first ref that has name greater than this prefix (the 0th prefix is\n    assumed to start at row 0, so we can interpret the prefix range from\n    row[i-1] and row[i]).\n\nBetween using raw OIDs and storing the depth-2 prefixes only once, this\nformat compresses the file to ~60% of its v1 size. (The format allows not\nwriting the prefix chunks, and the prefix chunks are implemented after the\nbasics of the ref chunks are complete.)\n\nThe write times are reduced in a similar fraction to the size difference.\nReads are sped up somewhat, and we have the potential to do a ref count by\nprefix much faster by doing a binary search for the start and end of the\nprefix and then subtracting the row positions instead of scanning the file\nbetween to count refs.\n\n\nRelationship to Reftable\n========================\n\nI mentioned earlier that I had considered using reftable as a way to achieve\nthe stated goals. With the current state of that work, I'm not confident\nthat it is the right approach here.\n\nMy main worry is that the reftable is more complicated than we need for a\ntypical Git repository that is based on a typical filesystem. This makes\ntesting the format very critical, and we seem to not be near reaching that\napproach. The v2 format here is very similar to existing Git file formats\nsince it uses the chunk-format API. This means that the amount of code\ncustom to just the v2 format is quite small.\n\nAs mentioned, the current extension plan [6] only allows reftable or files\nand does not allow for a mix of both. This RFC introduces the possibility\nthat both could co-exist. Using that multi-valued approach means that I'm\nable to test the v2 packed-refs file format almost as well as the v1 file\nformat within this RFC. (More tests need to be added that are specific to\nthis format, but I'm waiting for confirmation that this is an acceptable\ndirection.) At the very least, this multi-valued approach could be used as a\nway to allow using the reftable format as a drop-in replacement for the\npacked-refs file, as well as upgrading an existing repo to use reftable.\nThat might even help the integration process to allow the reftable format to\nbe tested at least by some subset of tests instead of waiting for a full\ntest suite update.\n\nI'm interested to hear from people more involved in the reftable work to see\nthe status of that project and how it matches or differs from my\nperspective.\n\nThe one thing I can say is that if the reftable work had not already begun,\nthen this is RFC is how I would have approached a new ref format.\n\nI look forward to your feedback!\n\nThanks,\n\n * Stolee\n\n[6]\nhttps://github.com/git/git/pull/1215/files#diff-a30f88b458b1f01e7a67e72576584b5b77ddb0362e40da6f7bf4a9ddf79db7b8R41-R48\nThe draft version of extensions.refFormat for reftable.\n\n[7] https://lore.kernel.org/git/4696bd93-9406-0abd-25ec-a739665a24d5@web.de/\nRe: [PATCH 00/15] Refactor chunk-format into an API (where René recommends a\ntrailing table of contents)\n\nDerrick Stolee (30):\n  hashfile: allow skipping the hash function\n  read-cache: add index.computeHash config option\n  extensions: add refFormat extension\n  config: fix multi-level bulleted list\n  repository: wire ref extensions to ref backends\n  refs: allow loose files without packed-refs\n  chunk-format: number of chunks is optional\n  chunk-format: document trailing table of contents\n  chunk-format: store chunk offset during write\n  chunk-format: allow trailing table of contents\n  chunk-format: parse trailing table of contents\n  refs: extract packfile format to new file\n  packed-backend: extract add_write_error()\n  packed-backend: extract iterator/updates merge\n  packed-backend: create abstraction for writing refs\n  config: add config values for packed-refs v2\n  packed-backend: create shell of v2 writes\n  packed-refs: write file format version 2\n  packed-refs: read file format v2\n  packed-refs: read optional prefix chunks\n  packed-refs: write prefix chunks\n  packed-backend: create GIT_TEST_PACKED_REFS_VERSION\n  t1409: test with packed-refs v2\n  t5312: allow packed-refs v2 format\n  t5502: add PACKED_REFS_V1 prerequisite\n  t3210: require packed-refs v1 for some tests\n  t*: skip packed-refs v2 over http tests\n  ci: run GIT_TEST_PACKED_REFS_VERSION=2 in some builds\n  p1401: create performance test for ref operations\n  refs: skip hashing when writing packed-refs v2\n\n Documentation/config.txt            |   2 +\n Documentation/config/extensions.txt |  76 ++-\n Documentation/config/index.txt      |   8 +\n Documentation/config/refs.txt       |  13 +\n Documentation/gitformat-chunk.txt   |  26 +-\n Makefile                            |   2 +\n cache.h                             |   2 +\n chunk-format.c                      | 109 +++-\n chunk-format.h                      |  18 +-\n ci/run-build-and-tests.sh           |   1 +\n commit-graph.c                      |   2 +-\n csum-file.c                         |  14 +-\n csum-file.h                         |   7 +\n midx.c                              |   2 +-\n read-cache.c                        |  22 +-\n refs.c                              |  24 +-\n refs/files-backend.c                |   8 +-\n refs/packed-backend.c               | 880 +++++++---------------------\n refs/packed-backend.h               | 281 +++++++++\n refs/packed-format-v1.c             | 456 ++++++++++++++\n refs/packed-format-v2.c             | 624 ++++++++++++++++++++\n refs/refs-internal.h                |   9 +\n repository.c                        |   2 +\n repository.h                        |   7 +\n setup.c                             |  26 +\n t/perf/p1401-ref-operations.sh      |  52 ++\n t/t1409-avoid-packing-refs.sh       |  22 +-\n t/t1600-index.sh                    |   8 +\n t/t3210-pack-refs.sh                |   8 +-\n t/t3212-ref-formats.sh              | 100 ++++\n t/t5502-quickfetch.sh               |   2 +-\n t/t5539-fetch-http-shallow.sh       |   7 +\n t/t5541-http-push-smart.sh          |   7 +\n t/t5542-push-http-shallow.sh        |   7 +\n t/t5551-http-fetch-smart.sh         |   7 +\n t/t5558-clone-bundle-uri.sh         |   7 +\n t/test-lib.sh                       |   4 +\n 37 files changed, 2157 insertions(+), 695 deletions(-)\n create mode 100644 Documentation/config/refs.txt\n create mode 100644 refs/packed-format-v1.c\n create mode 100644 refs/packed-format-v2.c\n create mode 100755 t/perf/p1401-ref-operations.sh\n create mode 100755 t/t3212-ref-formats.sh\n\n\nbase-commit: c03801e19cb8ab36e9c0d17ff3d5e0c3b0f24193\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1408%2Fderrickstolee%2Frefs%2Frfc-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1408/derrickstolee/refs/rfc-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/1408\n-- \ngitgitgadget\n"},{"id":"466685","messageId":"030d76f52af654470026b0c4b1dfba2b6c996885.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 02/30] read-cache: add index.computeHash config option","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:36Z","receivedAt":"2022-11-07T18:36:17Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe previous change allowed skipping the hashing portion of the\nhashwrite API, using it instead as a buffered write API. Disabling the\nhashwrite can be particularly helpful when the write operation is in a\ncritical path.\n\nOne such critical path is the writing of the index. This operation is so\ncritical that the sparse index was created specifically to reduce the\nsize of the index to make these writes (and reads) faster.\n\nFollowing a similar approach to one used in the microsoft/git fork [1],\nadd a new config option that allows disabling this hashing during the\nindex write. The cost is that we can no longer validate the contents for\ncorruption-at-rest using the trailing hash.\n\n[1] https://github.com/microsoft/git/commit/21fed2d91410f45d85279467f21d717a2db45201\n\nWhile older Git versions will not recognize the null hash as a special\ncase, the file format itself is still being met in terms of its\nstructure. Using this null hash will still allow Git operations to\nfunction across older versions.\n\nThe one exception is 'git fsck' which checks the hash of the index file.\nHere, we disable this check if the trailing hash is all zeroes. We add a\nwarning to the config option that this may cause undesirable behavior\nwith older Git versions.\n\nAs a quick comparison, I tested 'git update-index --force-write' with\nand without index.computHash=false on a copy of the Linux kernel\nrepository.\n\nBenchmark 1: with hash\n  Time (mean ± σ):      46.3 ms ±  13.8 ms    [User: 34.3 ms, System: 11.9 ms]\n  Range (min … max):    34.3 ms …  79.1 ms    82 runs\n\nBenchmark 2: without hash\n  Time (mean ± σ):      26.0 ms ±   7.9 ms    [User: 11.8 ms, System: 14.2 ms]\n  Range (min … max):    16.3 ms …  42.0 ms    69 runs\n\nSummary\n  'without hash' ran\n    1.78 ± 0.76 times faster than 'with hash'\n\nThese performance benefits are substantial enough to allow users the\nability to opt-in to this feature, even with the potential confusion\nwith older 'git fsck' versions.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n Documentation/config/index.txt |  8 ++++++++\n read-cache.c                   | 22 +++++++++++++++++++++-\n t/t1600-index.sh               |  8 ++++++++\n 3 files changed, 37 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/config/index.txt b/Documentation/config/index.txt\nindex 75f3a2d1054..709ba72f622 100644\n--- a/Documentation/config/index.txt\n+++ b/Documentation/config/index.txt\n@@ -30,3 +30,11 @@ index.version::\n \tSpecify the version with which new index files should be\n \tinitialized.  This does not affect existing repositories.\n \tIf `feature.manyFiles` is enabled, then the default is 4.\n+\n+index.computeHash::\n+\tWhen enabled, compute the hash of the index file as it is written\n+\tand store the hash at the end of the content. This is enabled by\n+\tdefault.\n++\n+If you disable `index.computHash`, then older Git clients may report that\n+your index is corrupt during `git fsck`.\ndiff --git a/read-cache.c b/read-cache.c\nindex 32024029274..f24d96de4d3 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1817,6 +1817,8 @@ static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n \tgit_hash_ctx c;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n \tint hdr_version;\n+\tint all_zeroes = 1;\n+\tunsigned char *start, *end;\n \n \tif (hdr->hdr_signature != htonl(CACHE_SIGNATURE))\n \t\treturn error(_(\"bad signature 0x%08x\"), hdr->hdr_signature);\n@@ -1827,10 +1829,23 @@ static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n \tif (!verify_index_checksum)\n \t\treturn 0;\n \n+\tend = (unsigned char *)hdr + size;\n+\tstart = end - the_hash_algo->rawsz;\n+\twhile (start < end) {\n+\t\tif (*start != 0) {\n+\t\t\tall_zeroes = 0;\n+\t\t\tbreak;\n+\t\t}\n+\t\tstart++;\n+\t}\n+\n+\tif (all_zeroes)\n+\t\treturn 0;\n+\n \tthe_hash_algo->init_fn(&c);\n \tthe_hash_algo->update_fn(&c, hdr, size - the_hash_algo->rawsz);\n \tthe_hash_algo->final_fn(hash, &c);\n-\tif (!hasheq(hash, (unsigned char *)hdr + size - the_hash_algo->rawsz))\n+\tif (!hasheq(hash, end - the_hash_algo->rawsz))\n \t\treturn error(_(\"bad index file sha1 signature\"));\n \treturn 0;\n }\n@@ -2917,9 +2932,14 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tint ieot_entries = 1;\n \tstruct index_entry_offset_table *ieot = NULL;\n \tint nr, nr_threads;\n+\tint compute_hash;\n \n \tf = hashfd(tempfile->fd, tempfile->filename.buf);\n \n+\tif (!git_config_get_maybe_bool(\"index.computehash\", &compute_hash) &&\n+\t    !compute_hash)\n+\t\tf->skip_hash = 1;\n+\n \tfor (i = removed = extended = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n \t\t\tremoved++;\ndiff --git a/t/t1600-index.sh b/t/t1600-index.sh\nindex 010989f90e6..24ab90ca047 100755\n--- a/t/t1600-index.sh\n+++ b/t/t1600-index.sh\n@@ -103,4 +103,12 @@ test_expect_success 'index version config precedence' '\n \ttest_index_version 0 true 2 2\n '\n \n+test_expect_success 'index.computeHash config option' '\n+\t(\n+\t\trm -f .git/index &&\n+\t\tgit -c index.computeHash=false add a &&\n+\t\tgit fsck\n+\t)\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"466686","messageId":"4013f992d15aab69346bf6f8eafe38511b923595.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 03/30] extensions: add refFormat extension","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:37Z","receivedAt":"2022-11-07T18:36:30Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nGit's reference storage is critical to its function. Creating new\nstorage formats for references requires adding an extension. This\nprevents third-party tools that do not understand that format from\noperating incorrectly on the repository. This makes updating ref formats\nmore difficult than other optional indexes, such as the commit-graph or\nmulti-pack-index.\n\nHowever, there are a number of potential ref storage enhancements that\nare underway or could be created. Git needs an established mechanism for\ncoordinating between these different options.\n\nThe first obvious format update is the reftable format as documented in\nDocumentation/technical/reftable.txt. This format has much of its\nimplementation already in Git, but its connection as a ref backend is\nnot complete. This change is similar to some changes within one of the\npatches intended for the reftable effort [1].\n\n[1] https://lore.kernel.org/git/pull.1215.git.git.1644351400761.gitgitgadget@gmail.com/\n\nHowever, this change makes a distinct strategy change from the one\nrecommended by reftable. Here, the extensions.refFormat extension is\nprovided as a multi-valued list. In the reftable RFC, the extension has\na single value, \"files\" or \"reftable\" and explicitly states that this\nshould not change after 'git init' or 'git clone'.\n\nThe single-valued approach has some major drawbacks, including the idea\nthat the \"files\" backend cannot coexist with the \"reftable\" backend at\nthe same time. In this way, it would not be possible to create a\nrepository that can write loose references and combine them into a\nreftable in the background. With the multi-valued approach, we could\nintegrate reftable as a drop-in replacement for the packed-refs file and\nallow that to be a faster way to do the integration since the test suite\nwould only need updates when the test is explicitly testing packed-refs.\n\nWhen upgrading a repository from the \"files\" backend to the \"reftable\"\nbackend, it can help to have a transition period where both are present,\nthen finally removing the \"files\" backend after all loose refs are\ncollected into the reftable.\n\nBut the reftable is not the only approach available.\n\nOne obvious improvement could be a new file format version for the\npacked-refs file. Its current plaintext-based format is inefficient due\nto storing object IDs as hexadecimal representations instead of in\ntheir raw format. This extra cost will get worse with SHA-256. In\naddition, binary searches need to guess a position and scan to find\nnewlines for a refname entry. A structured binary format could allow for\nmore compact representation and faster access. Adding such a format\ncould be seen as \"files-v2\", but it is really \"packed-v2\".\n\nThe reftable approach has a concept of a \"stack\" of reftable files. This\nidea would also work for a stack of packed-refs files (in v1 or v2\nformat). It would be helpful to describe that the refs could be stored\nin a stack of packed-ref files independently of whether that is in file\nformat v1 or v2.\n\nEven in these two options, it might be helpful to indicate whether or\nnot loose ref files are present. That is one reason to not make them\nappear as \"files-v2\" or \"files-v3\" options in a single-valued extension.\nEven as \"packed-v2\" or \"packed-v3\" options, this approach would require\nthird-party tools to understand the \"v2\" version if they want to support\nthe \"v3\" options. Instead, by splitting the format from the layout, we\ncan allow third-party tools to integrate only with the most-desired\nformat options.\n\nFor these reasons, this change is defining the extensions.refFormat\nextension as well as how the two existing values interact. By default,\nGit will assume \"files\" and \"packed\" in the list. If any other value\nis provided, then the extension is marked as unrecognized.\n\nAdd tests that check the behavior of extensions.refFormat, both in that\nit requires core.repositoryFormatVersion=1, and Git will refuse to work\nwith an unknown value of the extension.\n\nThere is a gap in the current implementation, though. What happens if\nexactly one of \"files\" or \"packed\" is provided? The presence of only one\nwould imply that the other is not available. A later change can\ncommunicate the list contents to the repository struct and then the\nreference backend could ignore one of these two layers.\n\nSpecifically, having only \"files\" would mean that Git should not read or\nwrite the packed-refs file and instead only read and write loose ref\nfiles. By contrast, having only \"packed\" would mean that Git should not\nread or write loose ref files and instead always update the packed-refs\nfile on every ref update.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n Documentation/config/extensions.txt | 41 +++++++++++++++++++++++++++++\n setup.c                             |  5 ++++\n t/t3212-ref-formats.sh              | 27 +++++++++++++++++++\n 3 files changed, 73 insertions(+)\n create mode 100755 t/t3212-ref-formats.sh\n\ndiff --git a/Documentation/config/extensions.txt b/Documentation/config/extensions.txt\nindex bccaec7a963..ce8185adf53 100644\n--- a/Documentation/config/extensions.txt\n+++ b/Documentation/config/extensions.txt\n@@ -7,6 +7,47 @@ Note that this setting should only be set by linkgit:git-init[1] or\n linkgit:git-clone[1].  Trying to change it after initialization will not\n work and will produce hard-to-diagnose issues.\n \n+extensions.refFormat::\n+\tSpecify the reference storage mechanisms used by the repoitory as a\n+\tmulti-valued list. The acceptable values are `files` and `packed`.\n+\tIf not specified, the list of `files` and `packed` is assumed. It\n+\tis an error to specify this key unless `core.repositoryFormatVersion`\n+\tis 1.\n++\n+As new ref formats are added, Git commands may modify this list before and\n+after upgrading the on-disk reference storage files. The specific values\n+indicate the existence of different layers:\n++\n+--\n+`files`;;\n+\tWhen present, references may be stored as \"loose\" reference files\n+\tin the `$GIT_DIR/refs/` directory. The name of the reference\n+\tcorresponds to the filename after `$GIT_DIR` and the file contains\n+\tan object ID as a hexadecimal string. If a loose reference file\n+\texists, then its value takes precedence over all other formats.\n+\n+`packed`;;\n+\tWhen present, references may be stored as a group in a\n+\t`packed-refs` file in its version 1 format. When grouped with\n+\t`\"files\"` or provided on its own, this file is located at\n+\t`$GIT_DIR/packed-refs`. This file contains a list of distinct\n+\treference names, paired with their object IDs. When combined with\n+\t`files`, the `packed` format will only be used to group multiple\n+\tloose object files upon request via the `git pack-refs` command or\n+\tvia the `pack-refs` maintenance task.\n+--\n++\n+The following combinations are supported by this version of Git:\n++\n+--\n+`files` and `packed`;;\n+\tThis set of values indicates that references are stored both as\n+\tloose reference files and in the `packed-refs` file in its v1\n+\tformat. Loose references are preferred, and the `packed-refs` file\n+\tis updated only when deleting a reference that is stored in the\n+\t`packed-refs` file or during a `git pack-refs` command.\n+--\n+\n extensions.worktreeConfig::\n \tIf enabled, then worktrees will load config settings from the\n \t`$GIT_DIR/config.worktree` file in addition to the\ndiff --git a/setup.c b/setup.c\nindex cefd5f63c46..f5eb50c969a 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -577,6 +577,11 @@ static enum extension_result handle_extension(const char *var,\n \t\t\t\t     \"extensions.objectformat\", value);\n \t\tdata->hash_algo = format;\n \t\treturn EXTENSION_OK;\n+\t} else if (!strcmp(ext, \"refformat\")) {\n+\t\tif (strcmp(value, \"files\") && strcmp(value, \"packed\"))\n+\t\t\treturn error(_(\"invalid value for '%s': '%s'\"),\n+\t\t\t\t     \"extensions.refFormat\", value);\n+\t\treturn EXTENSION_OK;\n \t}\n \treturn EXTENSION_UNKNOWN;\n }\ndiff --git a/t/t3212-ref-formats.sh b/t/t3212-ref-formats.sh\nnew file mode 100755\nindex 00000000000..bc554e7c701\n--- /dev/null\n+++ b/t/t3212-ref-formats.sh\n@@ -0,0 +1,27 @@\n+#!/bin/sh\n+\n+test_description='test across ref formats'\n+\n+. ./test-lib.sh\n+\n+test_expect_success 'extensions.refFormat requires core.repositoryFormatVersion=1' '\n+\ttest_when_finished rm -rf broken &&\n+\n+\t# Force sha1 to ensure GIT_TEST_DEFAULT_HASH does\n+\t# not imply a value of core.repositoryFormatVersion.\n+\tgit init --object-format=sha1 broken &&\n+\tgit -C broken config extensions.refFormat files &&\n+\ttest_must_fail git -C broken status 2>err &&\n+\tgrep \"repo version is 0, but v1-only extension found\" err\n+'\n+\n+test_expect_success 'invalid extensions.refFormat' '\n+\ttest_when_finished rm -rf broken &&\n+\tgit init broken &&\n+\tgit -C broken config core.repositoryFormatVersion 1 &&\n+\tgit -C broken config extensions.refFormat bogus &&\n+\ttest_must_fail git -C broken status 2>err &&\n+\tgrep \"invalid value for '\\''extensions.refFormat'\\'': '\\''bogus'\\''\" err\n+'\n+\n+test_done\n-- \ngitgitgadget\n\n"},{"id":"466687","messageId":"531bf1b6db0f5bbaf1508de5ea33f2e6d114f820.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 06/30] refs: allow loose files without packed-refs","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:40Z","receivedAt":"2022-11-07T18:36:30Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe extensions.refFormat extension is a multi-valued config that\nspecifies which ref formats are available to the current repository. By\ndefault, Git assumes the list of \"files\" and \"packed\", unless there is\nat least one of these extensions specified.\n\nWith the current values, it is possible for a user to specify only\n\"files\" or only \"packed\". The only-\"packed\" option was already ruled as\ninvalid since Git's current code has too many places that require a\nloose reference. This could change in the future.\n\nHowever, we can now allow the user to specify extensions.refFormat=files\nalone, making it impossible to create a packed-refs file (or to read one\nthat might exist).\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n Documentation/config/extensions.txt |  5 +++++\n refs/files-backend.c                |  6 ++++++\n refs/packed-backend.c               |  3 +++\n refs/refs-internal.h                |  5 +++++\n t/t3212-ref-formats.sh              | 20 ++++++++++++++++++++\n 5 files changed, 39 insertions(+)\n\ndiff --git a/Documentation/config/extensions.txt b/Documentation/config/extensions.txt\nindex 18ed1c58126..18071c336d0 100644\n--- a/Documentation/config/extensions.txt\n+++ b/Documentation/config/extensions.txt\n@@ -46,6 +46,11 @@ The following combinations are supported by this version of Git:\n \tformat. Loose references are preferred, and the `packed-refs` file\n \tis updated only when deleting a reference that is stored in the\n \t`packed-refs` file or during a `git pack-refs` command.\n+\n+`files`;;\n+\tWhen only this value is present, Git will ignore the `packed-refs`\n+\tfile and refuse to write one during `git pack-refs`. All references\n+\twill be read from and written to loose reference files.\n --\n \n extensions.worktreeConfig::\ndiff --git a/refs/files-backend.c b/refs/files-backend.c\nindex db6c8e434c6..4a18aed6204 100644\n--- a/refs/files-backend.c\n+++ b/refs/files-backend.c\n@@ -1198,6 +1198,12 @@ static int files_pack_refs(struct ref_store *ref_store, unsigned int flags)\n \tstruct strbuf err = STRBUF_INIT;\n \tstruct ref_transaction *transaction;\n \n+\tif (!packed_refs_enabled(refs->store_flags)) {\n+\t\twarning(_(\"refusing to create '%s' file because '%s' is not set\"),\n+\t\t\t\"packed-refs\", \"extensions.refFormat=packed\");\n+\t\treturn -1;\n+\t}\n+\n \ttransaction = ref_store_transaction_begin(refs->packed_ref_store, &err);\n \tif (!transaction)\n \t\treturn -1;\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex c1c71d183ea..a4371b711b9 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -478,6 +478,9 @@ static int load_contents(struct snapshot *snapshot)\n \tsize_t size;\n \tssize_t bytes_read;\n \n+\tif (!packed_refs_enabled(snapshot->refs->store_flags))\n+\t\treturn 0;\n+\n \tfd = open(snapshot->refs->path, O_RDONLY);\n \tif (fd < 0) {\n \t\tif (errno == ENOENT) {\ndiff --git a/refs/refs-internal.h b/refs/refs-internal.h\nindex 41520c945e4..a1900848a87 100644\n--- a/refs/refs-internal.h\n+++ b/refs/refs-internal.h\n@@ -524,6 +524,11 @@ struct ref_store;\n #define REF_STORE_FORMAT_FILES\t\t(1 << 8) /* can use loose ref files */\n #define REF_STORE_FORMAT_PACKED\t\t(1 << 9) /* can use packed-refs file */\n \n+static inline int packed_refs_enabled(int flags)\n+{\n+\treturn flags & REF_STORE_FORMAT_PACKED;\n+}\n+\n /*\n  * Initialize the ref_store for the specified gitdir. These functions\n  * should call base_ref_store_init() to initialize the shared part of\ndiff --git a/t/t3212-ref-formats.sh b/t/t3212-ref-formats.sh\nindex 8c4e70196a0..67aa65c116f 100755\n--- a/t/t3212-ref-formats.sh\n+++ b/t/t3212-ref-formats.sh\n@@ -36,4 +36,24 @@ test_expect_success 'extensions.refFormat=packed only' '\n \t)\n '\n \n+test_expect_success 'extensions.refFormat=files only' '\n+\ttest_commit T &&\n+\tgit pack-refs --all &&\n+\tgit init only-loose &&\n+\t(\n+\t\tcd only-loose &&\n+\t\tgit config core.repositoryFormatVersion 1 &&\n+\t\tgit config extensions.refFormat files &&\n+\t\ttest_commit A &&\n+\t\ttest_commit B &&\n+\t\ttest_must_fail git pack-refs 2>err &&\n+\t\tgrep \"refusing to create\" err &&\n+\t\ttest_path_is_missing .git/packed-refs &&\n+\n+\t\t# Refuse to parse a packed-refs file.\n+\t\tcp ../.git/packed-refs .git/packed-refs &&\n+\t\ttest_must_fail git rev-parse refs/tags/T\n+\t)\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"466688","messageId":"4fcbfed2c7c78c804c7eeeed5b7080b9fd812bb7.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 07/30] chunk-format: number of chunks is optional","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:41Z","receivedAt":"2022-11-07T18:36:33Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nEven though the commit-graph and multi-pack-index file formats specify a\nnumber of chunks in their header information, this is optional. The\ntable of contents terminates with a null chunk ID, which can be used\ninstead. The extra value is helpful for some checks, but is ultimately\nnot necessary for the format.\n\nThis will be important in some future formats.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n Documentation/gitformat-chunk.txt | 5 +++--\n 1 file changed, 3 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/gitformat-chunk.txt b/Documentation/gitformat-chunk.txt\nindex 57202ede273..c01f5567c4f 100644\n--- a/Documentation/gitformat-chunk.txt\n+++ b/Documentation/gitformat-chunk.txt\n@@ -24,8 +24,9 @@ how they use the chunks to describe structured data.\n \n A chunk-based file format begins with some header information custom to\n that format. That header should include enough information to identify\n-the file type, format version, and number of chunks in the file. From this\n-information, that file can determine the start of the chunk-based region.\n+the file type, format version, and (optionally) the number of chunks in\n+the file. From this information, that file can determine the start of the\n+chunk-based region.\n \n The chunk-based region starts with a table of contents describing where\n each chunk starts and ends. This consists of (C+1) rows of 12 bytes each,\n-- \ngitgitgadget\n\n"},{"id":"466689","messageId":"0cf654925f8d16a439871499a02125d75140ee36.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 04/30] config: fix multi-level bulleted list","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:38Z","receivedAt":"2022-11-07T18:36:35Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe documentation for 'extensions.worktreeConfig' includes a bulletted\nlist describing certain config values that need to be moved into the\nworktree config instead of the repository config file. However, since we\nare already in a bulletted list, the documentation tools do not know\nwhen that inner list is complete. Paragraphs intended to not be part of\nthat inner list are rendered as part of the last bullet.\n\nModify the format to match a similar doubly-nested list from the\n'column.ui' config documentation. Reword the descriptions slightly to\nmake the config keys appear as their own heading in the inner list.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n Documentation/config/extensions.txt | 13 +++++++++----\n 1 file changed, 9 insertions(+), 4 deletions(-)\n\ndiff --git a/Documentation/config/extensions.txt b/Documentation/config/extensions.txt\nindex ce8185adf53..18ed1c58126 100644\n--- a/Documentation/config/extensions.txt\n+++ b/Documentation/config/extensions.txt\n@@ -62,10 +62,15 @@ When enabling `extensions.worktreeConfig`, you must be careful to move\n certain values from the common config file to the main working tree's\n `config.worktree` file, if present:\n +\n-* `core.worktree` must be moved from `$GIT_COMMON_DIR/config` to\n-  `$GIT_COMMON_DIR/config.worktree`.\n-* If `core.bare` is true, then it must be moved from `$GIT_COMMON_DIR/config`\n-  to `$GIT_COMMON_DIR/config.worktree`.\n+--\n+`core.worktree`;;\n+\tThis config value must be moved from `$GIT_COMMON_DIR/config` to\n+\t`$GIT_COMMON_DIR/config.worktree`.\n+\n+`core.bare`;;\n+\tIf true, then this value must be moved from\n+\t`$GIT_COMMON_DIR/config` to `$GIT_COMMON_DIR/config.worktree`.\n+--\n +\n It may also be beneficial to adjust the locations of `core.sparseCheckout`\n and `core.sparseCheckoutCone` depending on your desire for customizable\n-- \ngitgitgadget\n\n"},{"id":"466690","messageId":"3121334256d8ab9afb2922a389ec22f9faaa08cf.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 05/30] repository: wire ref extensions to ref backends","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:39Z","receivedAt":"2022-11-07T18:36:37Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe previous change introduced the extensions.refFormat config option.\nIt is a multi-valued config option that currently understands \"files\"\nand \"packed\", with both values assumed by default. If any value is\nprovided explicitly, this default is ignored and the provided settings\nare used instead.\n\nThe multi-valued nature of this extension presents a way to allow a user\nto specify that they never want a packed-refs file (only use \"files\") or\nthat they never want loose reference files (only use \"packed\"). However,\nthat functionality is not currently connected.\n\nBefore actually modifying the files backend to understand these\nextension settings, do the basic wiring that connects the\nextensions.refFormat parsing to the creation of the ref backend. A\nfuture change will actually change the ref backend initialization based\non these settings, but this communication of the extension is\nsufficiently complicated to be worth an isolated change.\n\nFor now, also forbid the setting of only \"packed\". This is done by\nredirecting the choice of backend to the packed backend when that\nselection is made. A later change will make the \"files\"-only extension\nvalue ignore the packed backend.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n cache.h                |  2 ++\n refs.c                 | 22 ++++++++++++++++++++--\n refs/files-backend.c   |  2 +-\n refs/refs-internal.h   |  3 +++\n repository.c           |  2 ++\n repository.h           |  6 ++++++\n setup.c                | 18 +++++++++++++++++-\n t/t3212-ref-formats.sh | 12 ++++++++++++\n 8 files changed, 63 insertions(+), 4 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 26ed03bd6de..13e9c251ac3 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1155,6 +1155,8 @@ struct repository_format {\n \tint hash_algo;\n \tint sparse_index;\n \tchar *work_tree;\n+\tint ref_format_count;\n+\tenum ref_format_flags ref_format;\n \tstruct string_list unknown_extensions;\n \tstruct string_list v1_only_extensions;\n };\ndiff --git a/refs.c b/refs.c\nindex 1491ae937eb..21441ddb162 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -1982,6 +1982,15 @@ static struct ref_store *lookup_ref_store_map(struct hashmap *map,\n \treturn entry ? entry->refs : NULL;\n }\n \n+static int add_ref_format_flags(enum ref_format_flags flags, int caps) {\n+\tif (flags & REF_FORMAT_FILES)\n+\t\tcaps |= REF_STORE_FORMAT_FILES;\n+\tif (flags & REF_FORMAT_PACKED)\n+\t\tcaps |= REF_STORE_FORMAT_PACKED;\n+\n+\treturn caps;\n+}\n+\n /*\n  * Create, record, and return a ref_store instance for the specified\n  * gitdir.\n@@ -1991,9 +2000,17 @@ static struct ref_store *ref_store_init(struct repository *repo,\n \t\t\t\t\tunsigned int flags)\n {\n \tconst char *be_name = \"files\";\n-\tstruct ref_storage_be *be = find_ref_storage_backend(be_name);\n+\tstruct ref_storage_be *be;\n \tstruct ref_store *refs;\n \n+\tflags = add_ref_format_flags(repo->ref_format, flags);\n+\n+\tif (!(flags & REF_STORE_FORMAT_FILES) &&\n+\t    (flags & REF_STORE_FORMAT_PACKED))\n+\t\tbe_name = \"packed\";\n+\n+\tbe = find_ref_storage_backend(be_name);\n+\n \tif (!be)\n \t\tBUG(\"reference backend %s is unknown\", be_name);\n \n@@ -2009,7 +2026,8 @@ struct ref_store *get_main_ref_store(struct repository *r)\n \tif (!r->gitdir)\n \t\tBUG(\"attempting to get main_ref_store outside of repository\");\n \n-\tr->refs_private = ref_store_init(r, r->gitdir, REF_STORE_ALL_CAPS);\n+\tr->refs_private = ref_store_init(r, r->gitdir,\n+\t\t\t\t\t REF_STORE_ALL_CAPS);\n \tr->refs_private = maybe_debug_wrap_ref_store(r->gitdir, r->refs_private);\n \treturn r->refs_private;\n }\ndiff --git a/refs/files-backend.c b/refs/files-backend.c\nindex b89954355de..db6c8e434c6 100644\n--- a/refs/files-backend.c\n+++ b/refs/files-backend.c\n@@ -3274,7 +3274,7 @@ static int files_init_db(struct ref_store *ref_store, struct strbuf *err UNUSED)\n }\n \n struct ref_storage_be refs_be_files = {\n-\t.next = NULL,\n+\t.next = &refs_be_packed,\n \t.name = \"files\",\n \t.init = files_ref_store_create,\n \t.init_db = files_init_db,\ndiff --git a/refs/refs-internal.h b/refs/refs-internal.h\nindex 69f93b0e2ac..41520c945e4 100644\n--- a/refs/refs-internal.h\n+++ b/refs/refs-internal.h\n@@ -521,6 +521,9 @@ struct ref_store;\n \t\t\t\t REF_STORE_ODB | \\\n \t\t\t\t REF_STORE_MAIN)\n \n+#define REF_STORE_FORMAT_FILES\t\t(1 << 8) /* can use loose ref files */\n+#define REF_STORE_FORMAT_PACKED\t\t(1 << 9) /* can use packed-refs file */\n+\n /*\n  * Initialize the ref_store for the specified gitdir. These functions\n  * should call base_ref_store_init() to initialize the shared part of\ndiff --git a/repository.c b/repository.c\nindex 5d166b692c8..96533fc76be 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -182,6 +182,8 @@ int repo_init(struct repository *repo,\n \trepo->repository_format_partial_clone = format.partial_clone;\n \tformat.partial_clone = NULL;\n \n+\trepo->ref_format = format.ref_format;\n+\n \tif (worktree)\n \t\trepo_set_worktree(repo, worktree);\n \ndiff --git a/repository.h b/repository.h\nindex 24316ac944e..5cfde4282c5 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -61,6 +61,11 @@ struct repo_path_cache {\n \tchar *shallow;\n };\n \n+enum ref_format_flags {\n+\tREF_FORMAT_FILES = (1 << 0),\n+\tREF_FORMAT_PACKED = (1 << 1),\n+};\n+\n struct repository {\n \t/* Environment */\n \t/*\n@@ -95,6 +100,7 @@ struct repository {\n \t * the ref object.\n \t */\n \tstruct ref_store *refs_private;\n+\tenum ref_format_flags ref_format;\n \n \t/*\n \t * Contains path to often used file names.\ndiff --git a/setup.c b/setup.c\nindex f5eb50c969a..a5e63479558 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -578,9 +578,14 @@ static enum extension_result handle_extension(const char *var,\n \t\tdata->hash_algo = format;\n \t\treturn EXTENSION_OK;\n \t} else if (!strcmp(ext, \"refformat\")) {\n-\t\tif (strcmp(value, \"files\") && strcmp(value, \"packed\"))\n+\t\tif (!strcmp(value, \"files\"))\n+\t\t\tdata->ref_format |= REF_FORMAT_FILES;\n+\t\telse if (!strcmp(value, \"packed\"))\n+\t\t\tdata->ref_format |= REF_FORMAT_PACKED;\n+\t\telse\n \t\t\treturn error(_(\"invalid value for '%s': '%s'\"),\n \t\t\t\t     \"extensions.refFormat\", value);\n+\t\tdata->ref_format_count++;\n \t\treturn EXTENSION_OK;\n \t}\n \treturn EXTENSION_UNKNOWN;\n@@ -723,6 +728,11 @@ int read_repository_format(struct repository_format *format, const char *path)\n \tgit_config_from_file(check_repo_format, path, format);\n \tif (format->version == -1)\n \t\tclear_repository_format(format);\n+\n+\t/* Set default ref_format if no extensions.refFormat exists. */\n+\tif (!format->ref_format_count)\n+\t\tformat->ref_format = REF_FORMAT_FILES | REF_FORMAT_PACKED;\n+\n \treturn format->version;\n }\n \n@@ -1425,6 +1435,9 @@ int discover_git_directory(struct strbuf *commondir,\n \t\tcandidate.partial_clone;\n \tcandidate.partial_clone = NULL;\n \n+\t/* take ownership of candidate.ref_format */\n+\tthe_repository->ref_format = candidate.ref_format;\n+\n \tclear_repository_format(&candidate);\n \treturn 0;\n }\n@@ -1561,6 +1574,8 @@ const char *setup_git_directory_gently(int *nongit_ok)\n \t\t\tthe_repository->repository_format_partial_clone =\n \t\t\t\trepo_fmt.partial_clone;\n \t\t\trepo_fmt.partial_clone = NULL;\n+\n+\t\t\tthe_repository->ref_format = repo_fmt.ref_format;\n \t\t}\n \t}\n \t/*\n@@ -1650,6 +1665,7 @@ void check_repository_format(struct repository_format *fmt)\n \trepo_set_hash_algo(the_repository, fmt->hash_algo);\n \tthe_repository->repository_format_partial_clone =\n \t\txstrdup_or_null(fmt->partial_clone);\n+\tthe_repository->ref_format = fmt->ref_format;\n \tclear_repository_format(&repo_fmt);\n }\n \ndiff --git a/t/t3212-ref-formats.sh b/t/t3212-ref-formats.sh\nindex bc554e7c701..8c4e70196a0 100755\n--- a/t/t3212-ref-formats.sh\n+++ b/t/t3212-ref-formats.sh\n@@ -24,4 +24,16 @@ test_expect_success 'invalid extensions.refFormat' '\n \tgrep \"invalid value for '\\''extensions.refFormat'\\'': '\\''bogus'\\''\" err\n '\n \n+test_expect_success 'extensions.refFormat=packed only' '\n+\tgit init only-packed &&\n+\t(\n+\t\tcd only-packed &&\n+\t\tgit config core.repositoryFormatVersion 1 &&\n+\t\tgit config extensions.refFormat packed &&\n+\t\ttest_commit A &&\n+\t\ttest_path_exists .git/packed-refs &&\n+\t\ttest_path_is_missing .git/refs/tags/A\n+\t)\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"466691","messageId":"a7bf8cbec45859e79ac71dea06be391f75a0a524.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 08/30] chunk-format: document trailing table of contents","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:42Z","receivedAt":"2022-11-07T18:36:39Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nIt will be helpful to allow a trailing table of contents when writing\nsome file types with the chunk-format API. The main reason is that it\nallows dynamically computing the chunk sizes while writing the file.\nThis can use fewer resources than precomputing all chunk sizes in\nadvance.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n Documentation/gitformat-chunk.txt | 21 ++++++++++++++++++++-\n 1 file changed, 20 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/gitformat-chunk.txt b/Documentation/gitformat-chunk.txt\nindex c01f5567c4f..ee3718c4306 100644\n--- a/Documentation/gitformat-chunk.txt\n+++ b/Documentation/gitformat-chunk.txt\n@@ -52,8 +52,27 @@ The final entry in the table of contents must be four zero bytes. This\n confirms that the table of contents is ending and provides the offset for\n the end of the chunk-based data.\n \n+The default chunk format assumes the table of contents appears at the\n+beginning of the file (after the header information) and the chunks are\n+ordered by increasing offset. Alternatively, the chunk format allows a\n+table of contents that is placed at the end of the file (before the\n+trailing hash) and the offsets are in descending order. In this trailing\n+table of contents case, the data in order looks instead like the following\n+table:\n+\n+  | Chunk ID (4 bytes) | Chunk Offset (8 bytes) |\n+  |--------------------|------------------------|\n+  | 0x0000             | OFFSET[C+1]            |\n+  | ID[C]              | OFFSET[C]              |\n+  | ...                | ...                    |\n+  | ID[0]              | OFFSET[0]              |\n+\n+The concrete file format that uses the chunk format will mention that it\n+uses a trailing table of contents if it uses it. By default, the table of\n+contents is in ascending order before all chunk data.\n+\n Note: The chunk-based format expects that the file contains _at least_ a\n-trailing hash after `OFFSET[C+1]`.\n+trailing hash after either `OFFSET[C+1]` or the trailing table of contents.\n \n Functions for working with chunk-based file formats are declared in\n `chunk-format.h`. Using these methods provide extra checks that assist\n-- \ngitgitgadget\n\n"},{"id":"466692","messageId":"ff176b52306345fbb2ad96193b890839d7959015.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 09/30] chunk-format: store chunk offset during write","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:43Z","receivedAt":"2022-11-07T18:36:41Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nAs a preparatory step to allowing trailing table of contents, store the\noffsets of each chunk as we write them. This replaces an existing use of\na local variable, but the stored value will be used in the next change.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n chunk-format.c | 7 ++++---\n 1 file changed, 4 insertions(+), 3 deletions(-)\n\ndiff --git a/chunk-format.c b/chunk-format.c\nindex 0275b74a895..f1b2c8a8b36 100644\n--- a/chunk-format.c\n+++ b/chunk-format.c\n@@ -13,6 +13,7 @@ struct chunk_info {\n \tchunk_write_fn write_fn;\n \n \tconst void *start;\n+\toff_t offset;\n };\n \n struct chunkfile {\n@@ -78,16 +79,16 @@ int write_chunkfile(struct chunkfile *cf, void *data)\n \thashwrite_be64(cf->f, cur_offset);\n \n \tfor (i = 0; i < cf->chunks_nr; i++) {\n-\t\toff_t start_offset = hashfile_total(cf->f);\n+\t\tcf->chunks[i].offset = hashfile_total(cf->f);\n \t\tresult = cf->chunks[i].write_fn(cf->f, data);\n \n \t\tif (result)\n \t\t\tgoto cleanup;\n \n-\t\tif (hashfile_total(cf->f) - start_offset != cf->chunks[i].size)\n+\t\tif (hashfile_total(cf->f) - cf->chunks[i].offset != cf->chunks[i].size)\n \t\t\tBUG(\"expected to write %\"PRId64\" bytes to chunk %\"PRIx32\", but wrote %\"PRId64\" instead\",\n \t\t\t    cf->chunks[i].size, cf->chunks[i].id,\n-\t\t\t    hashfile_total(cf->f) - start_offset);\n+\t\t\t    hashfile_total(cf->f) - cf->chunks[i].offset);\n \t}\n \n cleanup:\n-- \ngitgitgadget\n\n"},{"id":"466693","messageId":"ebc719f92dd99bb6f5ae92104e87a05e520664d2.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 11/30] chunk-format: parse trailing table of contents","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:45Z","receivedAt":"2022-11-07T18:36:43Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe new read_trailing_table_of_contents() mimics\nread_table_of_contents() except that it reads the table of contents in\nreverse from the end of the given hashfile. The file is given as a\nmemory-mapped section of memory and a size. Automatically calculate the\nstart of the trailing hash and read the table of contents in revers from\nthat position.\n\nThe errors come along from those in read_table_of_contents(). The one\nexception is that the chunk_offset cannot be checked as going into the\ntable of contents since we do not have that length automatically. That\nmay have some surprising results for some narrow forms of corruption.\nHowever, we do still limit the size to the size of the file plus the\npart of the table of contents read so far. At minimum, the given sizes\ncan be used to limit parsing within the file itself.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n chunk-format.c | 53 ++++++++++++++++++++++++++++++++++++++++++++++++++\n chunk-format.h |  9 +++++++++\n 2 files changed, 62 insertions(+)\n\ndiff --git a/chunk-format.c b/chunk-format.c\nindex 3f5cc9b5ddf..e836a121c5c 100644\n--- a/chunk-format.c\n+++ b/chunk-format.c\n@@ -173,6 +173,59 @@ int read_table_of_contents(struct chunkfile *cf,\n \treturn 0;\n }\n \n+int read_trailing_table_of_contents(struct chunkfile *cf,\n+\t\t\t\t    const unsigned char *mfile,\n+\t\t\t\t    size_t mfile_size)\n+{\n+\tint i;\n+\tuint32_t chunk_id;\n+\tconst unsigned char *table_of_contents = mfile + mfile_size - the_hash_algo->rawsz;\n+\n+\twhile (1) {\n+\t\tuint64_t chunk_offset;\n+\n+\t\ttable_of_contents -= CHUNK_TOC_ENTRY_SIZE;\n+\n+\t\tchunk_id = get_be32(table_of_contents);\n+\t\tchunk_offset = get_be64(table_of_contents + 4);\n+\n+\t\t/* Calculate the previous chunk size, if it exists. */\n+\t\tif (cf->chunks_nr) {\n+\t\t\toff_t previous_offset = cf->chunks[cf->chunks_nr - 1].offset;\n+\n+\t\t\tif (chunk_offset < previous_offset ||\n+\t\t\t    chunk_offset > table_of_contents - mfile) {\n+\t\t\t\terror(_(\"improper chunk offset(s) %\"PRIx64\" and %\"PRIx64\"\"),\n+\t\t\t\tprevious_offset, chunk_offset);\n+\t\t\t\treturn -1;\n+\t\t\t}\n+\n+\t\t\tcf->chunks[cf->chunks_nr - 1].size = chunk_offset - previous_offset;\n+\t\t}\n+\n+\t\t/* Stop at the null chunk. We only need it for the last size. */\n+\t\tif (!chunk_id)\n+\t\t\tbreak;\n+\n+\t\tfor (i = 0; i < cf->chunks_nr; i++) {\n+\t\t\tif (cf->chunks[i].id == chunk_id) {\n+\t\t\t\terror(_(\"duplicate chunk ID %\"PRIx32\" found\"),\n+\t\t\t\t\tchunk_id);\n+\t\t\t\treturn -1;\n+\t\t\t}\n+\t\t}\n+\n+\t\tALLOC_GROW(cf->chunks, cf->chunks_nr + 1, cf->chunks_alloc);\n+\n+\t\tcf->chunks[cf->chunks_nr].id = chunk_id;\n+\t\tcf->chunks[cf->chunks_nr].start = mfile + chunk_offset;\n+\t\tcf->chunks[cf->chunks_nr].offset = chunk_offset;\n+\t\tcf->chunks_nr++;\n+\t}\n+\n+\treturn 0;\n+}\n+\n static int pair_chunk_fn(const unsigned char *chunk_start,\n \t\t\t size_t chunk_size,\n \t\t\t void *data)\ndiff --git a/chunk-format.h b/chunk-format.h\nindex 39e8967e950..acb8dfbce80 100644\n--- a/chunk-format.h\n+++ b/chunk-format.h\n@@ -46,6 +46,15 @@ int read_table_of_contents(struct chunkfile *cf,\n \t\t\t   uint64_t toc_offset,\n \t\t\t   int toc_length);\n \n+/**\n+ * Read the given chunkfile, but read the table of contents from the\n+ * end of the given mfile. The file is expected to be a hashfile with\n+ * the_hash_file->rawsz bytes at the end storing the hash.\n+ */\n+int read_trailing_table_of_contents(struct chunkfile *cf,\n+\t\t\t\t    const unsigned char *mfile,\n+\t\t\t\t    size_t mfile_size);\n+\n #define CHUNK_NOT_FOUND (-2)\n \n /*\n-- \ngitgitgadget\n\n"},{"id":"466694","messageId":"78e585cf4df2bb82a2569cee226a6b97d0ea7629.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 10/30] chunk-format: allow trailing table of contents","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:44Z","receivedAt":"2022-11-07T18:36:45Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe existing chunk formats use the table of contents at the beginning of\nthe file. This is intended as a way to speed up the initial loading of\nthe file, but comes at a cost during writes. Each example needs to fully\ncompute how big each chunk will be in advance, which usually requires\nstoring the full file contents in memory.\n\nFuture file formats may want to use the chunk format API in cases where\nthe writing stage is critical to performance, so we may want to stream\nupdates from an existing file and then only write the table of contents\nat the end.\n\nAdd a new 'flags' parameter to write_chunkfile() that allows this\nbehavior. When this is specified, the defensive programming that checks\nthat the chunks are written with the precomputed sizes is disabled.\nThen, the table of contents is written in reverse order at the end of\nthe hashfile, so a parser can read the chunk list starting from the end\nof the file (minus the hash).\n\nThe parsing of these table of contents will come in a later change.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n chunk-format.c | 53 +++++++++++++++++++++++++++++++++++---------------\n chunk-format.h |  9 ++++++++-\n commit-graph.c |  2 +-\n midx.c         |  2 +-\n 4 files changed, 47 insertions(+), 19 deletions(-)\n\ndiff --git a/chunk-format.c b/chunk-format.c\nindex f1b2c8a8b36..3f5cc9b5ddf 100644\n--- a/chunk-format.c\n+++ b/chunk-format.c\n@@ -57,26 +57,31 @@ void add_chunk(struct chunkfile *cf,\n \tcf->chunks_nr++;\n }\n \n-int write_chunkfile(struct chunkfile *cf, void *data)\n+int write_chunkfile(struct chunkfile *cf,\n+\t\t    enum chunkfile_flags flags,\n+\t\t    void *data)\n {\n \tint i, result = 0;\n-\tuint64_t cur_offset = hashfile_total(cf->f);\n \n \ttrace2_region_enter(\"chunkfile\", \"write\", the_repository);\n \n-\t/* Add the table of contents to the current offset */\n-\tcur_offset += (cf->chunks_nr + 1) * CHUNK_TOC_ENTRY_SIZE;\n+\tif (!(flags & CHUNKFILE_TRAILING_TOC)) {\n+\t\tuint64_t cur_offset = hashfile_total(cf->f);\n \n-\tfor (i = 0; i < cf->chunks_nr; i++) {\n-\t\thashwrite_be32(cf->f, cf->chunks[i].id);\n-\t\thashwrite_be64(cf->f, cur_offset);\n+\t\t/* Add the table of contents to the current offset */\n+\t\tcur_offset += (cf->chunks_nr + 1) * CHUNK_TOC_ENTRY_SIZE;\n \n-\t\tcur_offset += cf->chunks[i].size;\n-\t}\n+\t\tfor (i = 0; i < cf->chunks_nr; i++) {\n+\t\t\thashwrite_be32(cf->f, cf->chunks[i].id);\n+\t\t\thashwrite_be64(cf->f, cur_offset);\n \n-\t/* Trailing entry marks the end of the chunks */\n-\thashwrite_be32(cf->f, 0);\n-\thashwrite_be64(cf->f, cur_offset);\n+\t\t\tcur_offset += cf->chunks[i].size;\n+\t\t}\n+\n+\t\t/* Trailing entry marks the end of the chunks */\n+\t\thashwrite_be32(cf->f, 0);\n+\t\thashwrite_be64(cf->f, cur_offset);\n+\t}\n \n \tfor (i = 0; i < cf->chunks_nr; i++) {\n \t\tcf->chunks[i].offset = hashfile_total(cf->f);\n@@ -85,10 +90,26 @@ int write_chunkfile(struct chunkfile *cf, void *data)\n \t\tif (result)\n \t\t\tgoto cleanup;\n \n-\t\tif (hashfile_total(cf->f) - cf->chunks[i].offset != cf->chunks[i].size)\n-\t\t\tBUG(\"expected to write %\"PRId64\" bytes to chunk %\"PRIx32\", but wrote %\"PRId64\" instead\",\n-\t\t\t    cf->chunks[i].size, cf->chunks[i].id,\n-\t\t\t    hashfile_total(cf->f) - cf->chunks[i].offset);\n+\t\tif (!(flags & CHUNKFILE_TRAILING_TOC)) {\n+\t\t\tif (hashfile_total(cf->f) - cf->chunks[i].offset != cf->chunks[i].size)\n+\t\t\t\tBUG(\"expected to write %\"PRId64\" bytes to chunk %\"PRIx32\", but wrote %\"PRId64\" instead\",\n+\t\t\t\t    cf->chunks[i].size, cf->chunks[i].id,\n+\t\t\t\t    hashfile_total(cf->f) - cf->chunks[i].offset);\n+\t\t}\n+\n+\t\tcf->chunks[i].size = hashfile_total(cf->f) - cf->chunks[i].offset;\n+\t}\n+\n+\tif (flags & CHUNKFILE_TRAILING_TOC) {\n+\t\tsize_t last_chunk_tail = hashfile_total(cf->f);\n+\t\t/* First entry marks the end of the chunks */\n+\t\thashwrite_be32(cf->f, 0);\n+\t\thashwrite_be64(cf->f, last_chunk_tail);\n+\n+\t\tfor (i = cf->chunks_nr - 1; i >= 0; i--) {\n+\t\t\thashwrite_be32(cf->f, cf->chunks[i].id);\n+\t\t\thashwrite_be64(cf->f, cf->chunks[i].offset);\n+\t\t}\n \t}\n \n cleanup:\ndiff --git a/chunk-format.h b/chunk-format.h\nindex 7885aa08487..39e8967e950 100644\n--- a/chunk-format.h\n+++ b/chunk-format.h\n@@ -31,7 +31,14 @@ void add_chunk(struct chunkfile *cf,\n \t       uint32_t id,\n \t       size_t size,\n \t       chunk_write_fn fn);\n-int write_chunkfile(struct chunkfile *cf, void *data);\n+\n+enum chunkfile_flags {\n+\tCHUNKFILE_TRAILING_TOC = (1 << 0),\n+};\n+\n+int write_chunkfile(struct chunkfile *cf,\n+\t\t    enum chunkfile_flags flags,\n+\t\t    void *data);\n \n int read_table_of_contents(struct chunkfile *cf,\n \t\t\t   const unsigned char *mfile,\ndiff --git a/commit-graph.c b/commit-graph.c\nindex a7d87559328..c927b81250d 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1932,7 +1932,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t\t\tget_num_chunks(cf) * ctx->commits.nr);\n \t}\n \n-\twrite_chunkfile(cf, ctx);\n+\twrite_chunkfile(cf, 0, ctx);\n \n \tstop_progress(&ctx->progress);\n \tstrbuf_release(&progress_title);\ndiff --git a/midx.c b/midx.c\nindex 7cfad04a240..03d947a5d33 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -1510,7 +1510,7 @@ static int write_midx_internal(const char *object_dir,\n \t}\n \n \twrite_midx_header(f, get_num_chunks(cf), ctx.nr - dropped_packs);\n-\twrite_chunkfile(cf, &ctx);\n+\twrite_chunkfile(cf, 0, &ctx);\n \n \tfinalize_hashfile(f, midx_hash, FSYNC_COMPONENT_PACK_METADATA,\n \t\t\t  CSUM_FSYNC | CSUM_HASH_IN_STREAM);\n-- \ngitgitgadget\n\n"},{"id":"466695","messageId":"f141d8561ab850df36da0a4efb80315e34e7261a.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 13/30] packed-backend: extract add_write_error()","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:47Z","receivedAt":"2022-11-07T18:37:09Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe write_with_updates() method uses a write_error label to jump to code\nthat adds an error message before exiting with an error. This appears\nboth when the packed-refs file header is written, but also when a ref\nline is written to the packed-refs file.\n\nA future change will abstract the loop that writes the refs out of\nwrite_with_updates(), making the goto an inconvenient pattern. For now,\nremove the distinction between \"goto write_error\" and \"goto error\" by\nadding the message in-line using the new static method\nadd_write_error(). This is functionally equivalent, but will make the\nnext step easier.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n refs/packed-backend.c | 28 ++++++++++++++++++----------\n 1 file changed, 18 insertions(+), 10 deletions(-)\n\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex afaf6f53233..ef8060f2e08 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -529,6 +529,12 @@ static int packed_init_db(struct ref_store *ref_store UNUSED,\n \treturn 0;\n }\n \n+static void add_write_error(struct packed_ref_store *refs, struct strbuf *err)\n+{\n+\tstrbuf_addf(err, \"error writing to %s: %s\",\n+\t\t    get_tempfile_path(refs->tempfile), strerror(errno));\n+}\n+\n /*\n  * Write the packed refs from the current snapshot to the packed-refs\n  * tempfile, incorporating any changes from `updates`. `updates` must\n@@ -577,8 +583,10 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t\tgoto error;\n \t}\n \n-\tif (write_packed_file_header_v1(out) < 0)\n-\t\tgoto write_error;\n+\tif (write_packed_file_header_v1(out) < 0) {\n+\t\tadd_write_error(refs, err);\n+\t\tgoto error;\n+\t}\n \n \t/*\n \t * We iterate in parallel through the current list of refs and\n@@ -673,8 +681,10 @@ static int write_with_updates(struct packed_ref_store *refs,\n \n \t\t\tif (write_packed_entry_v1(out, iter->refname,\n \t\t\t\t\t\t  iter->oid,\n-\t\t\t\t\t\t  peel_error ? NULL : &peeled))\n-\t\t\t\tgoto write_error;\n+\t\t\t\t\t\t  peel_error ? NULL : &peeled)) {\n+\t\t\t\tadd_write_error(refs, err);\n+\t\t\t\tgoto error;\n+\t\t\t}\n \n \t\t\tif ((ok = ref_iterator_advance(iter)) != ITER_OK)\n \t\t\t\titer = NULL;\n@@ -694,8 +704,10 @@ static int write_with_updates(struct packed_ref_store *refs,\n \n \t\t\tif (write_packed_entry_v1(out, update->refname,\n \t\t\t\t\t\t  &update->new_oid,\n-\t\t\t\t\t\t  peel_error ? NULL : &peeled))\n-\t\t\t\tgoto write_error;\n+\t\t\t\t\t\t  peel_error ? NULL : &peeled)) {\n+\t\t\t\tadd_write_error(refs, err);\n+\t\t\t\tgoto error;\n+\t\t\t}\n \n \t\t\ti++;\n \t\t}\n@@ -719,10 +731,6 @@ static int write_with_updates(struct packed_ref_store *refs,\n \n \treturn 0;\n \n-write_error:\n-\tstrbuf_addf(err, \"error writing to %s: %s\",\n-\t\t    get_tempfile_path(refs->tempfile), strerror(errno));\n-\n error:\n \tif (iter)\n \t\tref_iterator_abort(iter);\n-- \ngitgitgadget\n\n"},{"id":"466696","messageId":"a171a84da65b777d9da4d0c1c841901712aca026.1667846164.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 12/30] refs: extract packfile format to new file","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:46Z","receivedAt":"2022-11-07T18:37:10Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nIn preparation for adding a new packed-refs file format, extract all\ncode from refs/packed-backend.c that involves knowledge of the plaintext\nfile format. This includes any parsing logic that cares about the\nheader, plaintext lines of the form \"<oid> <ref>\" or \"^<peeled>\", and\nthe error messages when there is an issue in the file. This also\nincludes the writing logic that writes the header or the individual\nreferences.\n\nFuture changes will perform more refactoring to abstract away more of\nthe writing process to be more generic, but this is enough of a chunk of\ncode movement.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n Makefile                |   1 +\n refs/packed-backend.c   | 595 ++--------------------------------------\n refs/packed-backend.h   | 195 +++++++++++++\n refs/packed-format-v1.c | 453 ++++++++++++++++++++++++++++++\n 4 files changed, 667 insertions(+), 577 deletions(-)\n create mode 100644 refs/packed-format-v1.c\n\ndiff --git a/Makefile b/Makefile\nindex 4927379184c..3dc887941d4 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1057,6 +1057,7 @@ LIB_OBJS += refs/debug.o\n LIB_OBJS += refs/files-backend.o\n LIB_OBJS += refs/iterator.o\n LIB_OBJS += refs/packed-backend.o\n+LIB_OBJS += refs/packed-format-v1.o\n LIB_OBJS += refs/ref-cache.o\n LIB_OBJS += refspec.o\n LIB_OBJS += remote.o\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex a4371b711b9..afaf6f53233 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -36,121 +36,6 @@ static enum mmap_strategy mmap_strategy = MMAP_TEMPORARY;\n static enum mmap_strategy mmap_strategy = MMAP_OK;\n #endif\n \n-struct packed_ref_store;\n-\n-/*\n- * A `snapshot` represents one snapshot of a `packed-refs` file.\n- *\n- * Normally, this will be a mmapped view of the contents of the\n- * `packed-refs` file at the time the snapshot was created. However,\n- * if the `packed-refs` file was not sorted, this might point at heap\n- * memory holding the contents of the `packed-refs` file with its\n- * records sorted by refname.\n- *\n- * `snapshot` instances are reference counted (via\n- * `acquire_snapshot()` and `release_snapshot()`). This is to prevent\n- * an instance from disappearing while an iterator is still iterating\n- * over it. Instances are garbage collected when their `referrers`\n- * count goes to zero.\n- *\n- * The most recent `snapshot`, if available, is referenced by the\n- * `packed_ref_store`. Its freshness is checked whenever\n- * `get_snapshot()` is called; if the existing snapshot is obsolete, a\n- * new snapshot is taken.\n- */\n-struct snapshot {\n-\t/*\n-\t * A back-pointer to the packed_ref_store with which this\n-\t * snapshot is associated:\n-\t */\n-\tstruct packed_ref_store *refs;\n-\n-\t/* Is the `packed-refs` file currently mmapped? */\n-\tint mmapped;\n-\n-\t/*\n-\t * The contents of the `packed-refs` file:\n-\t *\n-\t * - buf -- a pointer to the start of the memory\n-\t * - start -- a pointer to the first byte of actual references\n-\t *   (i.e., after the header line, if one is present)\n-\t * - eof -- a pointer just past the end of the reference\n-\t *   contents\n-\t *\n-\t * If the `packed-refs` file was already sorted, `buf` points\n-\t * at the mmapped contents of the file. If not, it points at\n-\t * heap-allocated memory containing the contents, sorted. If\n-\t * there were no contents (e.g., because the file didn't\n-\t * exist), `buf`, `start`, and `eof` are all NULL.\n-\t */\n-\tchar *buf, *start, *eof;\n-\n-\t/*\n-\t * What is the peeled state of the `packed-refs` file that\n-\t * this snapshot represents? (This is usually determined from\n-\t * the file's header.)\n-\t */\n-\tenum { PEELED_NONE, PEELED_TAGS, PEELED_FULLY } peeled;\n-\n-\t/*\n-\t * Count of references to this instance, including the pointer\n-\t * from `packed_ref_store::snapshot`, if any. The instance\n-\t * will not be freed as long as the reference count is\n-\t * nonzero.\n-\t */\n-\tunsigned int referrers;\n-\n-\t/*\n-\t * The metadata of the `packed-refs` file from which this\n-\t * snapshot was created, used to tell if the file has been\n-\t * replaced since we read it.\n-\t */\n-\tstruct stat_validity validity;\n-};\n-\n-/*\n- * A `ref_store` representing references stored in a `packed-refs`\n- * file. It implements the `ref_store` interface, though it has some\n- * limitations:\n- *\n- * - It cannot store symbolic references.\n- *\n- * - It cannot store reflogs.\n- *\n- * - It does not support reference renaming (though it could).\n- *\n- * On the other hand, it can be locked outside of a reference\n- * transaction. In that case, it remains locked even after the\n- * transaction is done and the new `packed-refs` file is activated.\n- */\n-struct packed_ref_store {\n-\tstruct ref_store base;\n-\n-\tunsigned int store_flags;\n-\n-\t/* The path of the \"packed-refs\" file: */\n-\tchar *path;\n-\n-\t/*\n-\t * A snapshot of the values read from the `packed-refs` file,\n-\t * if it might still be current; otherwise, NULL.\n-\t */\n-\tstruct snapshot *snapshot;\n-\n-\t/*\n-\t * Lock used for the \"packed-refs\" file. Note that this (and\n-\t * thus the enclosing `packed_ref_store`) must not be freed.\n-\t */\n-\tstruct lock_file lock;\n-\n-\t/*\n-\t * Temporary file used when rewriting new contents to the\n-\t * \"packed-refs\" file. Note that this (and thus the enclosing\n-\t * `packed_ref_store`) must not be freed.\n-\t */\n-\tstruct tempfile *tempfile;\n-};\n-\n /*\n  * Increment the reference count of `*snapshot`.\n  */\n@@ -164,7 +49,7 @@ static void acquire_snapshot(struct snapshot *snapshot)\n  * memory and close the file, or free the memory. Then set the buffer\n  * pointers to NULL.\n  */\n-static void clear_snapshot_buffer(struct snapshot *snapshot)\n+void clear_snapshot_buffer(struct snapshot *snapshot)\n {\n \tif (snapshot->mmapped) {\n \t\tif (munmap(snapshot->buf, snapshot->eof - snapshot->buf))\n@@ -245,224 +130,6 @@ static void clear_snapshot(struct packed_ref_store *refs)\n \t}\n }\n \n-static NORETURN void die_unterminated_line(const char *path,\n-\t\t\t\t\t   const char *p, size_t len)\n-{\n-\tif (len < 80)\n-\t\tdie(\"unterminated line in %s: %.*s\", path, (int)len, p);\n-\telse\n-\t\tdie(\"unterminated line in %s: %.75s...\", path, p);\n-}\n-\n-static NORETURN void die_invalid_line(const char *path,\n-\t\t\t\t      const char *p, size_t len)\n-{\n-\tconst char *eol = memchr(p, '\\n', len);\n-\n-\tif (!eol)\n-\t\tdie_unterminated_line(path, p, len);\n-\telse if (eol - p < 80)\n-\t\tdie(\"unexpected line in %s: %.*s\", path, (int)(eol - p), p);\n-\telse\n-\t\tdie(\"unexpected line in %s: %.75s...\", path, p);\n-\n-}\n-\n-struct snapshot_record {\n-\tconst char *start;\n-\tsize_t len;\n-};\n-\n-static int cmp_packed_ref_records(const void *v1, const void *v2)\n-{\n-\tconst struct snapshot_record *e1 = v1, *e2 = v2;\n-\tconst char *r1 = e1->start + the_hash_algo->hexsz + 1;\n-\tconst char *r2 = e2->start + the_hash_algo->hexsz + 1;\n-\n-\twhile (1) {\n-\t\tif (*r1 == '\\n')\n-\t\t\treturn *r2 == '\\n' ? 0 : -1;\n-\t\tif (*r1 != *r2) {\n-\t\t\tif (*r2 == '\\n')\n-\t\t\t\treturn 1;\n-\t\t\telse\n-\t\t\t\treturn (unsigned char)*r1 < (unsigned char)*r2 ? -1 : +1;\n-\t\t}\n-\t\tr1++;\n-\t\tr2++;\n-\t}\n-}\n-\n-/*\n- * Compare a snapshot record at `rec` to the specified NUL-terminated\n- * refname.\n- */\n-static int cmp_record_to_refname(const char *rec, const char *refname)\n-{\n-\tconst char *r1 = rec + the_hash_algo->hexsz + 1;\n-\tconst char *r2 = refname;\n-\n-\twhile (1) {\n-\t\tif (*r1 == '\\n')\n-\t\t\treturn *r2 ? -1 : 0;\n-\t\tif (!*r2)\n-\t\t\treturn 1;\n-\t\tif (*r1 != *r2)\n-\t\t\treturn (unsigned char)*r1 < (unsigned char)*r2 ? -1 : +1;\n-\t\tr1++;\n-\t\tr2++;\n-\t}\n-}\n-\n-/*\n- * `snapshot->buf` is not known to be sorted. Check whether it is, and\n- * if not, sort it into new memory and munmap/free the old storage.\n- */\n-static void sort_snapshot(struct snapshot *snapshot)\n-{\n-\tstruct snapshot_record *records = NULL;\n-\tsize_t alloc = 0, nr = 0;\n-\tint sorted = 1;\n-\tconst char *pos, *eof, *eol;\n-\tsize_t len, i;\n-\tchar *new_buffer, *dst;\n-\n-\tpos = snapshot->start;\n-\teof = snapshot->eof;\n-\n-\tif (pos == eof)\n-\t\treturn;\n-\n-\tlen = eof - pos;\n-\n-\t/*\n-\t * Initialize records based on a crude estimate of the number\n-\t * of references in the file (we'll grow it below if needed):\n-\t */\n-\tALLOC_GROW(records, len / 80 + 20, alloc);\n-\n-\twhile (pos < eof) {\n-\t\teol = memchr(pos, '\\n', eof - pos);\n-\t\tif (!eol)\n-\t\t\t/* The safety check should prevent this. */\n-\t\t\tBUG(\"unterminated line found in packed-refs\");\n-\t\tif (eol - pos < the_hash_algo->hexsz + 2)\n-\t\t\tdie_invalid_line(snapshot->refs->path,\n-\t\t\t\t\t pos, eof - pos);\n-\t\teol++;\n-\t\tif (eol < eof && *eol == '^') {\n-\t\t\t/*\n-\t\t\t * Keep any peeled line together with its\n-\t\t\t * reference:\n-\t\t\t */\n-\t\t\tconst char *peeled_start = eol;\n-\n-\t\t\teol = memchr(peeled_start, '\\n', eof - peeled_start);\n-\t\t\tif (!eol)\n-\t\t\t\t/* The safety check should prevent this. */\n-\t\t\t\tBUG(\"unterminated peeled line found in packed-refs\");\n-\t\t\teol++;\n-\t\t}\n-\n-\t\tALLOC_GROW(records, nr + 1, alloc);\n-\t\trecords[nr].start = pos;\n-\t\trecords[nr].len = eol - pos;\n-\t\tnr++;\n-\n-\t\tif (sorted &&\n-\t\t    nr > 1 &&\n-\t\t    cmp_packed_ref_records(&records[nr - 2],\n-\t\t\t\t\t   &records[nr - 1]) >= 0)\n-\t\t\tsorted = 0;\n-\n-\t\tpos = eol;\n-\t}\n-\n-\tif (sorted)\n-\t\tgoto cleanup;\n-\n-\t/* We need to sort the memory. First we sort the records array: */\n-\tQSORT(records, nr, cmp_packed_ref_records);\n-\n-\t/*\n-\t * Allocate a new chunk of memory, and copy the old memory to\n-\t * the new in the order indicated by `records` (not bothering\n-\t * with the header line):\n-\t */\n-\tnew_buffer = xmalloc(len);\n-\tfor (dst = new_buffer, i = 0; i < nr; i++) {\n-\t\tmemcpy(dst, records[i].start, records[i].len);\n-\t\tdst += records[i].len;\n-\t}\n-\n-\t/*\n-\t * Now munmap the old buffer and use the sorted buffer in its\n-\t * place:\n-\t */\n-\tclear_snapshot_buffer(snapshot);\n-\tsnapshot->buf = snapshot->start = new_buffer;\n-\tsnapshot->eof = new_buffer + len;\n-\n-cleanup:\n-\tfree(records);\n-}\n-\n-/*\n- * Return a pointer to the start of the record that contains the\n- * character `*p` (which must be within the buffer). If no other\n- * record start is found, return `buf`.\n- */\n-static const char *find_start_of_record(const char *buf, const char *p)\n-{\n-\twhile (p > buf && (p[-1] != '\\n' || p[0] == '^'))\n-\t\tp--;\n-\treturn p;\n-}\n-\n-/*\n- * Return a pointer to the start of the record following the record\n- * that contains `*p`. If none is found before `end`, return `end`.\n- */\n-static const char *find_end_of_record(const char *p, const char *end)\n-{\n-\twhile (++p < end && (p[-1] != '\\n' || p[0] == '^'))\n-\t\t;\n-\treturn p;\n-}\n-\n-/*\n- * We want to be able to compare mmapped reference records quickly,\n- * without totally parsing them. We can do so because the records are\n- * LF-terminated, and the refname should start exactly (GIT_SHA1_HEXSZ\n- * + 1) bytes past the beginning of the record.\n- *\n- * But what if the `packed-refs` file contains garbage? We're willing\n- * to tolerate not detecting the problem, as long as we don't produce\n- * totally garbled output (we can't afford to check the integrity of\n- * the whole file during every Git invocation). But we do want to be\n- * sure that we never read past the end of the buffer in memory and\n- * perform an illegal memory access.\n- *\n- * Guarantee that minimum level of safety by verifying that the last\n- * record in the file is LF-terminated, and that it has at least\n- * (GIT_SHA1_HEXSZ + 1) characters before the LF. Die if either of\n- * these checks fails.\n- */\n-static void verify_buffer_safe(struct snapshot *snapshot)\n-{\n-\tconst char *start = snapshot->start;\n-\tconst char *eof = snapshot->eof;\n-\tconst char *last_line;\n-\n-\tif (start == eof)\n-\t\treturn;\n-\n-\tlast_line = find_start_of_record(start, eof - 1);\n-\tif (*(eof - 1) != '\\n' || eof - last_line < the_hash_algo->hexsz + 2)\n-\t\tdie_invalid_line(snapshot->refs->path,\n-\t\t\t\t last_line, eof - last_line);\n-}\n-\n #define SMALL_FILE_SIZE (32*1024)\n \n /*\n@@ -524,67 +191,6 @@ static int load_contents(struct snapshot *snapshot)\n \treturn 1;\n }\n \n-/*\n- * Find the place in `snapshot->buf` where the start of the record for\n- * `refname` starts. If `mustexist` is true and the reference doesn't\n- * exist, then return NULL. If `mustexist` is false and the reference\n- * doesn't exist, then return the point where that reference would be\n- * inserted, or `snapshot->eof` (which might be NULL) if it would be\n- * inserted at the end of the file. In the latter mode, `refname`\n- * doesn't have to be a proper reference name; for example, one could\n- * search for \"refs/replace/\" to find the start of any replace\n- * references.\n- *\n- * The record is sought using a binary search, so `snapshot->buf` must\n- * be sorted.\n- */\n-static const char *find_reference_location(struct snapshot *snapshot,\n-\t\t\t\t\t   const char *refname, int mustexist)\n-{\n-\t/*\n-\t * This is not *quite* a garden-variety binary search, because\n-\t * the data we're searching is made up of records, and we\n-\t * always need to find the beginning of a record to do a\n-\t * comparison. A \"record\" here is one line for the reference\n-\t * itself and zero or one peel lines that start with '^'. Our\n-\t * loop invariant is described in the next two comments.\n-\t */\n-\n-\t/*\n-\t * A pointer to the character at the start of a record whose\n-\t * preceding records all have reference names that come\n-\t * *before* `refname`.\n-\t */\n-\tconst char *lo = snapshot->start;\n-\n-\t/*\n-\t * A pointer to a the first character of a record whose\n-\t * reference name comes *after* `refname`.\n-\t */\n-\tconst char *hi = snapshot->eof;\n-\n-\twhile (lo != hi) {\n-\t\tconst char *mid, *rec;\n-\t\tint cmp;\n-\n-\t\tmid = lo + (hi - lo) / 2;\n-\t\trec = find_start_of_record(lo, mid);\n-\t\tcmp = cmp_record_to_refname(rec, refname);\n-\t\tif (cmp < 0) {\n-\t\t\tlo = find_end_of_record(mid, hi);\n-\t\t} else if (cmp > 0) {\n-\t\t\thi = rec;\n-\t\t} else {\n-\t\t\treturn rec;\n-\t\t}\n-\t}\n-\n-\tif (mustexist)\n-\t\treturn NULL;\n-\telse\n-\t\treturn lo;\n-}\n-\n /*\n  * Create a newly-allocated `snapshot` of the `packed-refs` file in\n  * its current state and return it. The return value will already have\n@@ -630,54 +236,22 @@ static struct snapshot *create_snapshot(struct packed_ref_store *refs)\n \tif (!load_contents(snapshot))\n \t\treturn snapshot;\n \n-\t/* If the file has a header line, process it: */\n-\tif (snapshot->buf < snapshot->eof && *snapshot->buf == '#') {\n-\t\tchar *tmp, *p, *eol;\n-\t\tstruct string_list traits = STRING_LIST_INIT_NODUP;\n-\n-\t\teol = memchr(snapshot->buf, '\\n',\n-\t\t\t     snapshot->eof - snapshot->buf);\n-\t\tif (!eol)\n-\t\t\tdie_unterminated_line(refs->path,\n-\t\t\t\t\t      snapshot->buf,\n-\t\t\t\t\t      snapshot->eof - snapshot->buf);\n-\n-\t\ttmp = xmemdupz(snapshot->buf, eol - snapshot->buf);\n-\n-\t\tif (!skip_prefix(tmp, \"# pack-refs with:\", (const char **)&p))\n-\t\t\tdie_invalid_line(refs->path,\n-\t\t\t\t\t snapshot->buf,\n-\t\t\t\t\t snapshot->eof - snapshot->buf);\n-\n-\t\tstring_list_split_in_place(&traits, p, ' ', -1);\n-\n-\t\tif (unsorted_string_list_has_string(&traits, \"fully-peeled\"))\n-\t\t\tsnapshot->peeled = PEELED_FULLY;\n-\t\telse if (unsorted_string_list_has_string(&traits, \"peeled\"))\n-\t\t\tsnapshot->peeled = PEELED_TAGS;\n-\n-\t\tsorted = unsorted_string_list_has_string(&traits, \"sorted\");\n-\n-\t\t/* perhaps other traits later as well */\n-\n-\t\t/* The \"+ 1\" is for the LF character. */\n-\t\tsnapshot->start = eol + 1;\n-\n-\t\tstring_list_clear(&traits, 0);\n-\t\tfree(tmp);\n+\tif (parse_packed_format_v1_header(refs, snapshot, &sorted)) {\n+\t\tclear_snapshot(refs);\n+\t\treturn NULL;\n \t}\n \n-\tverify_buffer_safe(snapshot);\n+\tverify_buffer_safe_v1(snapshot);\n \n \tif (!sorted) {\n-\t\tsort_snapshot(snapshot);\n+\t\tsort_snapshot_v1(snapshot);\n \n \t\t/*\n \t\t * Reordering the records might have moved a short one\n \t\t * to the end of the buffer, so verify the buffer's\n \t\t * safety again:\n \t\t */\n-\t\tverify_buffer_safe(snapshot);\n+\t\tverify_buffer_safe_v1(snapshot);\n \t}\n \n \tif (mmap_strategy != MMAP_OK && snapshot->mmapped) {\n@@ -735,55 +309,11 @@ static int packed_read_raw_ref(struct ref_store *ref_store, const char *refname,\n \tstruct packed_ref_store *refs =\n \t\tpacked_downcast(ref_store, REF_STORE_READ, \"read_raw_ref\");\n \tstruct snapshot *snapshot = get_snapshot(refs);\n-\tconst char *rec;\n-\n-\t*type = 0;\n \n-\trec = find_reference_location(snapshot, refname, 1);\n-\n-\tif (!rec) {\n-\t\t/* refname is not a packed reference. */\n-\t\t*failure_errno = ENOENT;\n-\t\treturn -1;\n-\t}\n-\n-\tif (get_oid_hex(rec, oid))\n-\t\tdie_invalid_line(refs->path, rec, snapshot->eof - rec);\n-\n-\t*type = REF_ISPACKED;\n-\treturn 0;\n+\treturn packed_read_raw_ref_v1(refs, snapshot, refname,\n+\t\t\t\t      oid, type, failure_errno);\n }\n \n-/*\n- * This value is set in `base.flags` if the peeled value of the\n- * current reference is known. In that case, `peeled` contains the\n- * correct peeled value for the reference, which might be `null_oid`\n- * if the reference is not a tag or if it is broken.\n- */\n-#define REF_KNOWS_PEELED 0x40\n-\n-/*\n- * An iterator over a snapshot of a `packed-refs` file.\n- */\n-struct packed_ref_iterator {\n-\tstruct ref_iterator base;\n-\n-\tstruct snapshot *snapshot;\n-\n-\t/* The current position in the snapshot's buffer: */\n-\tconst char *pos;\n-\n-\t/* The end of the part of the buffer that will be iterated over: */\n-\tconst char *eof;\n-\n-\t/* Scratch space for current values: */\n-\tstruct object_id oid, peeled;\n-\tstruct strbuf refname_buf;\n-\n-\tstruct repository *repo;\n-\tunsigned int flags;\n-};\n-\n /*\n  * Move the iterator to the next record in the snapshot, without\n  * respect for whether the record is actually required by the current\n@@ -793,68 +323,7 @@ struct packed_ref_iterator {\n  */\n static int next_record(struct packed_ref_iterator *iter)\n {\n-\tconst char *p = iter->pos, *eol;\n-\n-\tstrbuf_reset(&iter->refname_buf);\n-\n-\tif (iter->pos == iter->eof)\n-\t\treturn ITER_DONE;\n-\n-\titer->base.flags = REF_ISPACKED;\n-\n-\tif (iter->eof - p < the_hash_algo->hexsz + 2 ||\n-\t    parse_oid_hex(p, &iter->oid, &p) ||\n-\t    !isspace(*p++))\n-\t\tdie_invalid_line(iter->snapshot->refs->path,\n-\t\t\t\t iter->pos, iter->eof - iter->pos);\n-\n-\teol = memchr(p, '\\n', iter->eof - p);\n-\tif (!eol)\n-\t\tdie_unterminated_line(iter->snapshot->refs->path,\n-\t\t\t\t      iter->pos, iter->eof - iter->pos);\n-\n-\tstrbuf_add(&iter->refname_buf, p, eol - p);\n-\titer->base.refname = iter->refname_buf.buf;\n-\n-\tif (check_refname_format(iter->base.refname, REFNAME_ALLOW_ONELEVEL)) {\n-\t\tif (!refname_is_safe(iter->base.refname))\n-\t\t\tdie(\"packed refname is dangerous: %s\",\n-\t\t\t    iter->base.refname);\n-\t\toidclr(&iter->oid);\n-\t\titer->base.flags |= REF_BAD_NAME | REF_ISBROKEN;\n-\t}\n-\tif (iter->snapshot->peeled == PEELED_FULLY ||\n-\t    (iter->snapshot->peeled == PEELED_TAGS &&\n-\t     starts_with(iter->base.refname, \"refs/tags/\")))\n-\t\titer->base.flags |= REF_KNOWS_PEELED;\n-\n-\titer->pos = eol + 1;\n-\n-\tif (iter->pos < iter->eof && *iter->pos == '^') {\n-\t\tp = iter->pos + 1;\n-\t\tif (iter->eof - p < the_hash_algo->hexsz + 1 ||\n-\t\t    parse_oid_hex(p, &iter->peeled, &p) ||\n-\t\t    *p++ != '\\n')\n-\t\t\tdie_invalid_line(iter->snapshot->refs->path,\n-\t\t\t\t\t iter->pos, iter->eof - iter->pos);\n-\t\titer->pos = p;\n-\n-\t\t/*\n-\t\t * Regardless of what the file header said, we\n-\t\t * definitely know the value of *this* reference. But\n-\t\t * we suppress it if the reference is broken:\n-\t\t */\n-\t\tif ((iter->base.flags & REF_ISBROKEN)) {\n-\t\t\toidclr(&iter->peeled);\n-\t\t\titer->base.flags &= ~REF_KNOWS_PEELED;\n-\t\t} else {\n-\t\t\titer->base.flags |= REF_KNOWS_PEELED;\n-\t\t}\n-\t} else {\n-\t\toidclr(&iter->peeled);\n-\t}\n-\n-\treturn ITER_OK;\n+\treturn next_record_v1(iter);\n }\n \n static int packed_ref_iterator_advance(struct ref_iterator *ref_iterator)\n@@ -942,7 +411,7 @@ static struct ref_iterator *packed_ref_iterator_begin(\n \tsnapshot = get_snapshot(refs);\n \n \tif (prefix && *prefix)\n-\t\tstart = find_reference_location(snapshot, prefix, 0);\n+\t\tstart = find_reference_location_v1(snapshot, prefix, 0);\n \telse\n \t\tstart = snapshot->start;\n \n@@ -972,23 +441,6 @@ static struct ref_iterator *packed_ref_iterator_begin(\n \treturn ref_iterator;\n }\n \n-/*\n- * Write an entry to the packed-refs file for the specified refname.\n- * If peeled is non-NULL, write it as the entry's peeled value. On\n- * error, return a nonzero value and leave errno set at the value left\n- * by the failing call to `fprintf()`.\n- */\n-static int write_packed_entry(FILE *fh, const char *refname,\n-\t\t\t      const struct object_id *oid,\n-\t\t\t      const struct object_id *peeled)\n-{\n-\tif (fprintf(fh, \"%s %s\\n\", oid_to_hex(oid), refname) < 0 ||\n-\t    (peeled && fprintf(fh, \"^%s\\n\", oid_to_hex(peeled)) < 0))\n-\t\treturn -1;\n-\n-\treturn 0;\n-}\n-\n int packed_refs_lock(struct ref_store *ref_store, int flags, struct strbuf *err)\n {\n \tstruct packed_ref_store *refs =\n@@ -1070,17 +522,6 @@ int packed_refs_is_locked(struct ref_store *ref_store)\n \treturn is_lock_file_locked(&refs->lock);\n }\n \n-/*\n- * The packed-refs header line that we write out. Perhaps other traits\n- * will be added later.\n- *\n- * Note that earlier versions of Git used to parse these traits by\n- * looking for \" trait \" in the line. For this reason, the space after\n- * the colon and the trailing space are required.\n- */\n-static const char PACKED_REFS_HEADER[] =\n-\t\"# pack-refs with: peeled fully-peeled sorted \\n\";\n-\n static int packed_init_db(struct ref_store *ref_store UNUSED,\n \t\t\t  struct strbuf *err UNUSED)\n {\n@@ -1136,7 +577,7 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t\tgoto error;\n \t}\n \n-\tif (fprintf(out, \"%s\", PACKED_REFS_HEADER) < 0)\n+\tif (write_packed_file_header_v1(out) < 0)\n \t\tgoto write_error;\n \n \t/*\n@@ -1230,9 +671,9 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t\t\tstruct object_id peeled;\n \t\t\tint peel_error = ref_iterator_peel(iter, &peeled);\n \n-\t\t\tif (write_packed_entry(out, iter->refname,\n-\t\t\t\t\t       iter->oid,\n-\t\t\t\t\t       peel_error ? NULL : &peeled))\n+\t\t\tif (write_packed_entry_v1(out, iter->refname,\n+\t\t\t\t\t\t  iter->oid,\n+\t\t\t\t\t\t  peel_error ? NULL : &peeled))\n \t\t\t\tgoto write_error;\n \n \t\t\tif ((ok = ref_iterator_advance(iter)) != ITER_OK)\n@@ -1251,9 +692,9 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t\t\tint peel_error = peel_object(&update->new_oid,\n \t\t\t\t\t\t     &peeled);\n \n-\t\t\tif (write_packed_entry(out, update->refname,\n-\t\t\t\t\t       &update->new_oid,\n-\t\t\t\t\t       peel_error ? NULL : &peeled))\n+\t\t\tif (write_packed_entry_v1(out, update->refname,\n+\t\t\t\t\t\t  &update->new_oid,\n+\t\t\t\t\t\t  peel_error ? NULL : &peeled))\n \t\t\t\tgoto write_error;\n \n \t\t\ti++;\ndiff --git a/refs/packed-backend.h b/refs/packed-backend.h\nindex 9dd8a344c34..143ed6d4f6c 100644\n--- a/refs/packed-backend.h\n+++ b/refs/packed-backend.h\n@@ -1,6 +1,10 @@\n #ifndef REFS_PACKED_BACKEND_H\n #define REFS_PACKED_BACKEND_H\n \n+#include \"../cache.h\"\n+#include \"refs-internal.h\"\n+#include \"../lockfile.h\"\n+\n struct repository;\n struct ref_transaction;\n \n@@ -36,4 +40,195 @@ int packed_refs_is_locked(struct ref_store *ref_store);\n int is_packed_transaction_needed(struct ref_store *ref_store,\n \t\t\t\t struct ref_transaction *transaction);\n \n+struct packed_ref_store;\n+\n+/*\n+ * A `snapshot` represents one snapshot of a `packed-refs` file.\n+ *\n+ * Normally, this will be a mmapped view of the contents of the\n+ * `packed-refs` file at the time the snapshot was created. However,\n+ * if the `packed-refs` file was not sorted, this might point at heap\n+ * memory holding the contents of the `packed-refs` file with its\n+ * records sorted by refname.\n+ *\n+ * `snapshot` instances are reference counted (via\n+ * `acquire_snapshot()` and `release_snapshot()`). This is to prevent\n+ * an instance from disappearing while an iterator is still iterating\n+ * over it. Instances are garbage collected when their `referrers`\n+ * count goes to zero.\n+ *\n+ * The most recent `snapshot`, if available, is referenced by the\n+ * `packed_ref_store`. Its freshness is checked whenever\n+ * `get_snapshot()` is called; if the existing snapshot is obsolete, a\n+ * new snapshot is taken.\n+ */\n+struct snapshot {\n+\t/*\n+\t * A back-pointer to the packed_ref_store with which this\n+\t * snapshot is associated:\n+\t */\n+\tstruct packed_ref_store *refs;\n+\n+\t/* Is the `packed-refs` file currently mmapped? */\n+\tint mmapped;\n+\n+\t/*\n+\t * The contents of the `packed-refs` file:\n+\t *\n+\t * - buf -- a pointer to the start of the memory\n+\t * - start -- a pointer to the first byte of actual references\n+\t *   (i.e., after the header line, if one is present)\n+\t * - eof -- a pointer just past the end of the reference\n+\t *   contents\n+\t *\n+\t * If the `packed-refs` file was already sorted, `buf` points\n+\t * at the mmapped contents of the file. If not, it points at\n+\t * heap-allocated memory containing the contents, sorted. If\n+\t * there were no contents (e.g., because the file didn't\n+\t * exist), `buf`, `start`, and `eof` are all NULL.\n+\t */\n+\tchar *buf, *start, *eof;\n+\n+\t/*\n+\t * What is the peeled state of the `packed-refs` file that\n+\t * this snapshot represents? (This is usually determined from\n+\t * the file's header.)\n+\t */\n+\tenum { PEELED_NONE, PEELED_TAGS, PEELED_FULLY } peeled;\n+\n+\t/*\n+\t * Count of references to this instance, including the pointer\n+\t * from `packed_ref_store::snapshot`, if any. The instance\n+\t * will not be freed as long as the reference count is\n+\t * nonzero.\n+\t */\n+\tunsigned int referrers;\n+\n+\t/*\n+\t * The metadata of the `packed-refs` file from which this\n+\t * snapshot was created, used to tell if the file has been\n+\t * replaced since we read it.\n+\t */\n+\tstruct stat_validity validity;\n+};\n+\n+/*\n+ * If the buffer in `snapshot` is active, then either munmap the\n+ * memory and close the file, or free the memory. Then set the buffer\n+ * pointers to NULL.\n+ */\n+void clear_snapshot_buffer(struct snapshot *snapshot);\n+\n+/*\n+ * A `ref_store` representing references stored in a `packed-refs`\n+ * file. It implements the `ref_store` interface, though it has some\n+ * limitations:\n+ *\n+ * - It cannot store symbolic references.\n+ *\n+ * - It cannot store reflogs.\n+ *\n+ * - It does not support reference renaming (though it could).\n+ *\n+ * On the other hand, it can be locked outside of a reference\n+ * transaction. In that case, it remains locked even after the\n+ * transaction is done and the new `packed-refs` file is activated.\n+ */\n+struct packed_ref_store {\n+\tstruct ref_store base;\n+\n+\tunsigned int store_flags;\n+\n+\t/* The path of the \"packed-refs\" file: */\n+\tchar *path;\n+\n+\t/*\n+\t * A snapshot of the values read from the `packed-refs` file,\n+\t * if it might still be current; otherwise, NULL.\n+\t */\n+\tstruct snapshot *snapshot;\n+\n+\t/*\n+\t * Lock used for the \"packed-refs\" file. Note that this (and\n+\t * thus the enclosing `packed_ref_store`) must not be freed.\n+\t */\n+\tstruct lock_file lock;\n+\n+\t/*\n+\t * Temporary file used when rewriting new contents to the\n+\t * \"packed-refs\" file. Note that this (and thus the enclosing\n+\t * `packed_ref_store`) must not be freed.\n+\t */\n+\tstruct tempfile *tempfile;\n+};\n+\n+/*\n+ * This value is set in `base.flags` if the peeled value of the\n+ * current reference is known. In that case, `peeled` contains the\n+ * correct peeled value for the reference, which might be `null_oid`\n+ * if the reference is not a tag or if it is broken.\n+ */\n+#define REF_KNOWS_PEELED 0x40\n+\n+/*\n+ * An iterator over a snapshot of a `packed-refs` file.\n+ */\n+struct packed_ref_iterator {\n+\tstruct ref_iterator base;\n+\n+\tstruct snapshot *snapshot;\n+\n+\t/* The current position in the snapshot's buffer: */\n+\tconst char *pos;\n+\n+\t/* The end of the part of the buffer that will be iterated over: */\n+\tconst char *eof;\n+\n+\t/* Scratch space for current values: */\n+\tstruct object_id oid, peeled;\n+\tstruct strbuf refname_buf;\n+\n+\tstruct repository *repo;\n+\tunsigned int flags;\n+};\n+\n+/**\n+ * Parse the buffer at the given snapshot to verify that it is a\n+ * packed-refs file in version 1 format. Update the snapshot->peeled\n+ * value according to the header information. Update the given\n+ * 'sorted' value with whether or not the packed-refs file is sorted.\n+ */\n+int parse_packed_format_v1_header(struct packed_ref_store *refs,\n+\t\t\t\t  struct snapshot *snapshot,\n+\t\t\t\t  int *sorted);\n+\n+/*\n+ * Find the place in `snapshot->buf` where the start of the record for\n+ * `refname` starts. If `mustexist` is true and the reference doesn't\n+ * exist, then return NULL. If `mustexist` is false and the reference\n+ * doesn't exist, then return the point where that reference would be\n+ * inserted, or `snapshot->eof` (which might be NULL) if it would be\n+ * inserted at the end of the file. In the latter mode, `refname`\n+ * doesn't have to be a proper reference name; for example, one could\n+ * search for \"refs/replace/\" to find the start of any replace\n+ * references.\n+ *\n+ * The record is sought using a binary search, so `snapshot->buf` must\n+ * be sorted.\n+ */\n+const char *find_reference_location_v1(struct snapshot *snapshot,\n+\t\t\t\t       const char *refname, int mustexist);\n+\n+int packed_read_raw_ref_v1(struct packed_ref_store *refs, struct snapshot *snapshot,\n+\t\t\t   const char *refname, struct object_id *oid,\n+\t\t\t   unsigned int *type, int *failure_errno);\n+\n+void verify_buffer_safe_v1(struct snapshot *snapshot);\n+void sort_snapshot_v1(struct snapshot *snapshot);\n+int write_packed_file_header_v1(FILE *out);\n+int next_record_v1(struct packed_ref_iterator *iter);\n+int write_packed_entry_v1(FILE *fh, const char *refname,\n+\t\t\t  const struct object_id *oid,\n+\t\t\t  const struct object_id *peeled);\n+\n #endif /* REFS_PACKED_BACKEND_H */\ndiff --git a/refs/packed-format-v1.c b/refs/packed-format-v1.c\nnew file mode 100644\nindex 00000000000..ef9e6618c89\n--- /dev/null\n+++ b/refs/packed-format-v1.c\n@@ -0,0 +1,453 @@\n+#include \"../cache.h\"\n+#include \"../config.h\"\n+#include \"../refs.h\"\n+#include \"refs-internal.h\"\n+#include \"packed-backend.h\"\n+#include \"../iterator.h\"\n+#include \"../lockfile.h\"\n+#include \"../chdir-notify.h\"\n+\n+static NORETURN void die_unterminated_line(const char *path,\n+\t\t\t\t\t   const char *p, size_t len)\n+{\n+\tif (len < 80)\n+\t\tdie(\"unterminated line in %s: %.*s\", path, (int)len, p);\n+\telse\n+\t\tdie(\"unterminated line in %s: %.75s...\", path, p);\n+}\n+\n+static NORETURN void die_invalid_line(const char *path,\n+\t\t\t\t      const char *p, size_t len)\n+{\n+\tconst char *eol = memchr(p, '\\n', len);\n+\n+\tif (!eol)\n+\t\tdie_unterminated_line(path, p, len);\n+\telse if (eol - p < 80)\n+\t\tdie(\"unexpected line in %s: %.*s\", path, (int)(eol - p), p);\n+\telse\n+\t\tdie(\"unexpected line in %s: %.75s...\", path, p);\n+}\n+\n+struct snapshot_record {\n+\tconst char *start;\n+\tsize_t len;\n+};\n+\n+static int cmp_packed_ref_records(const void *v1, const void *v2)\n+{\n+\tconst struct snapshot_record *e1 = v1, *e2 = v2;\n+\tconst char *r1 = e1->start + the_hash_algo->hexsz + 1;\n+\tconst char *r2 = e2->start + the_hash_algo->hexsz + 1;\n+\n+\twhile (1) {\n+\t\tif (*r1 == '\\n')\n+\t\t\treturn *r2 == '\\n' ? 0 : -1;\n+\t\tif (*r1 != *r2) {\n+\t\t\tif (*r2 == '\\n')\n+\t\t\t\treturn 1;\n+\t\t\telse\n+\t\t\t\treturn (unsigned char)*r1 < (unsigned char)*r2 ? -1 : +1;\n+\t\t}\n+\t\tr1++;\n+\t\tr2++;\n+\t}\n+}\n+\n+/*\n+ * Compare a snapshot record at `rec` to the specified NUL-terminated\n+ * refname.\n+ */\n+static int cmp_record_to_refname(const char *rec, const char *refname)\n+{\n+\tconst char *r1 = rec + the_hash_algo->hexsz + 1;\n+\tconst char *r2 = refname;\n+\n+\twhile (1) {\n+\t\tif (*r1 == '\\n')\n+\t\t\treturn *r2 ? -1 : 0;\n+\t\tif (!*r2)\n+\t\t\treturn 1;\n+\t\tif (*r1 != *r2)\n+\t\t\treturn (unsigned char)*r1 < (unsigned char)*r2 ? -1 : +1;\n+\t\tr1++;\n+\t\tr2++;\n+\t}\n+}\n+\n+/*\n+ * `snapshot->buf` is not known to be sorted. Check whether it is, and\n+ * if not, sort it into new memory and munmap/free the old storage.\n+ */\n+void sort_snapshot_v1(struct snapshot *snapshot)\n+{\n+\tstruct snapshot_record *records = NULL;\n+\tsize_t alloc = 0, nr = 0;\n+\tint sorted = 1;\n+\tconst char *pos, *eof, *eol;\n+\tsize_t len, i;\n+\tchar *new_buffer, *dst;\n+\n+\tpos = snapshot->start;\n+\teof = snapshot->eof;\n+\n+\tif (pos == eof)\n+\t\treturn;\n+\n+\tlen = eof - pos;\n+\n+\t/*\n+\t * Initialize records based on a crude estimate of the number\n+\t * of references in the file (we'll grow it below if needed):\n+\t */\n+\tALLOC_GROW(records, len / 80 + 20, alloc);\n+\n+\twhile (pos < eof) {\n+\t\teol = memchr(pos, '\\n', eof - pos);\n+\t\tif (!eol)\n+\t\t\t/* The safety check should prevent this. */\n+\t\t\tBUG(\"unterminated line found in packed-refs\");\n+\t\tif (eol - pos < the_hash_algo->hexsz + 2)\n+\t\t\tdie_invalid_line(snapshot->refs->path,\n+\t\t\t\t\t pos, eof - pos);\n+\t\teol++;\n+\t\tif (eol < eof && *eol == '^') {\n+\t\t\t/*\n+\t\t\t * Keep any peeled line together with its\n+\t\t\t * reference:\n+\t\t\t */\n+\t\t\tconst char *peeled_start = eol;\n+\n+\t\t\teol = memchr(peeled_start, '\\n', eof - peeled_start);\n+\t\t\tif (!eol)\n+\t\t\t\t/* The safety check should prevent this. */\n+\t\t\t\tBUG(\"unterminated peeled line found in packed-refs\");\n+\t\t\teol++;\n+\t\t}\n+\n+\t\tALLOC_GROW(records, nr + 1, alloc);\n+\t\trecords[nr].start = pos;\n+\t\trecords[nr].len = eol - pos;\n+\t\tnr++;\n+\n+\t\tif (sorted &&\n+\t\t    nr > 1 &&\n+\t\t    cmp_packed_ref_records(&records[nr - 2],\n+\t\t\t\t\t   &records[nr - 1]) >= 0)\n+\t\t\tsorted = 0;\n+\n+\t\tpos = eol;\n+\t}\n+\n+\tif (sorted)\n+\t\tgoto cleanup;\n+\n+\t/* We need to sort the memory. First we sort the records array: */\n+\tQSORT(records, nr, cmp_packed_ref_records);\n+\n+\t/*\n+\t * Allocate a new chunk of memory, and copy the old memory to\n+\t * the new in the order indicated by `records` (not bothering\n+\t * with the header line):\n+\t */\n+\tnew_buffer = xmalloc(len);\n+\tfor (dst = new_buffer, i = 0; i < nr; i++) {\n+\t\tmemcpy(dst, records[i].start, records[i].len);\n+\t\tdst += records[i].len;\n+\t}\n+\n+\t/*\n+\t * Now munmap the old buffer and use the sorted buffer in its\n+\t * place:\n+\t */\n+\tclear_snapshot_buffer(snapshot);\n+\tsnapshot->buf = snapshot->start = new_buffer;\n+\tsnapshot->eof = new_buffer + len;\n+\n+cleanup:\n+\tfree(records);\n+}\n+\n+/*\n+ * Return a pointer to the start of the record that contains the\n+ * character `*p` (which must be within the buffer). If no other\n+ * record start is found, return `buf`.\n+ */\n+static const char *find_start_of_record(const char *buf, const char *p)\n+{\n+\twhile (p > buf && (p[-1] != '\\n' || p[0] == '^'))\n+\t\tp--;\n+\treturn p;\n+}\n+\n+/*\n+ * Return a pointer to the start of the record following the record\n+ * that contains `*p`. If none is found before `end`, return `end`.\n+ */\n+static const char *find_end_of_record(const char *p, const char *end)\n+{\n+\twhile (++p < end && (p[-1] != '\\n' || p[0] == '^'))\n+\t\t;\n+\treturn p;\n+}\n+\n+/*\n+ * We want to be able to compare mmapped reference records quickly,\n+ * without totally parsing them. We can do so because the records are\n+ * LF-terminated, and the refname should start exactly (GIT_SHA1_HEXSZ\n+ * + 1) bytes past the beginning of the record.\n+ *\n+ * But what if the `packed-refs` file contains garbage? We're willing\n+ * to tolerate not detecting the problem, as long as we don't produce\n+ * totally garbled output (we can't afford to check the integrity of\n+ * the whole file during every Git invocation). But we do want to be\n+ * sure that we never read past the end of the buffer in memory and\n+ * perform an illegal memory access.\n+ *\n+ * Guarantee that minimum level of safety by verifying that the last\n+ * record in the file is LF-terminated, and that it has at least\n+ * (GIT_SHA1_HEXSZ + 1) characters before the LF. Die if either of\n+ * these checks fails.\n+ */\n+void verify_buffer_safe_v1(struct snapshot *snapshot)\n+{\n+\tconst char *start = snapshot->start;\n+\tconst char *eof = snapshot->eof;\n+\tconst char *last_line;\n+\n+\tif (start == eof)\n+\t\treturn;\n+\n+\tlast_line = find_start_of_record(start, eof - 1);\n+\tif (*(eof - 1) != '\\n' || eof - last_line < the_hash_algo->hexsz + 2)\n+\t\tdie_invalid_line(snapshot->refs->path,\n+\t\t\t\t last_line, eof - last_line);\n+}\n+\n+/*\n+ * Find the place in `snapshot->buf` where the start of the record for\n+ * `refname` starts. If `mustexist` is true and the reference doesn't\n+ * exist, then return NULL. If `mustexist` is false and the reference\n+ * doesn't exist, then return the point where that reference would be\n+ * inserted, or `snapshot->eof` (which might be NULL) if it would be\n+ * inserted at the end of the file. In the latter mode, `refname`\n+ * doesn't have to be a proper reference name; for example, one could\n+ * search for \"refs/replace/\" to find the start of any replace\n+ * references.\n+ *\n+ * The record is sought using a binary search, so `snapshot->buf` must\n+ * be sorted.\n+ */\n+const char *find_reference_location_v1(struct snapshot *snapshot,\n+\t\t\t\t       const char *refname, int mustexist)\n+{\n+\t/*\n+\t * This is not *quite* a garden-variety binary search, because\n+\t * the data we're searching is made up of records, and we\n+\t * always need to find the beginning of a record to do a\n+\t * comparison. A \"record\" here is one line for the reference\n+\t * itself and zero or one peel lines that start with '^'. Our\n+\t * loop invariant is described in the next two comments.\n+\t */\n+\n+\t/*\n+\t * A pointer to the character at the start of a record whose\n+\t * preceding records all have reference names that come\n+\t * *before* `refname`.\n+\t */\n+\tconst char *lo = snapshot->start;\n+\n+\t/*\n+\t * A pointer to a the first character of a record whose\n+\t * reference name comes *after* `refname`.\n+\t */\n+\tconst char *hi = snapshot->eof;\n+\n+\twhile (lo != hi) {\n+\t\tconst char *mid, *rec;\n+\t\tint cmp;\n+\n+\t\tmid = lo + (hi - lo) / 2;\n+\t\trec = find_start_of_record(lo, mid);\n+\t\tcmp = cmp_record_to_refname(rec, refname);\n+\t\tif (cmp < 0) {\n+\t\t\tlo = find_end_of_record(mid, hi);\n+\t\t} else if (cmp > 0) {\n+\t\t\thi = rec;\n+\t\t} else {\n+\t\t\treturn rec;\n+\t\t}\n+\t}\n+\n+\tif (mustexist)\n+\t\treturn NULL;\n+\telse\n+\t\treturn lo;\n+}\n+\n+int parse_packed_format_v1_header(struct packed_ref_store *refs,\n+\t\t\t\t  struct snapshot *snapshot,\n+\t\t\t\t  int *sorted)\n+{\n+\t*sorted = 0;\n+\t/* If the file has a header line, process it: */\n+\tif (snapshot->buf < snapshot->eof && *snapshot->buf == '#') {\n+\t\tchar *tmp, *p, *eol;\n+\t\tstruct string_list traits = STRING_LIST_INIT_NODUP;\n+\n+\t\teol = memchr(snapshot->buf, '\\n',\n+\t\t\t     snapshot->eof - snapshot->buf);\n+\t\tif (!eol)\n+\t\t\tdie_unterminated_line(refs->path,\n+\t\t\t\t\t      snapshot->buf,\n+\t\t\t\t\t      snapshot->eof - snapshot->buf);\n+\n+\t\ttmp = xmemdupz(snapshot->buf, eol - snapshot->buf);\n+\n+\t\tif (!skip_prefix(tmp, \"# pack-refs with:\", (const char **)&p))\n+\t\t\tdie_invalid_line(refs->path,\n+\t\t\t\t\t snapshot->buf,\n+\t\t\t\t\t snapshot->eof - snapshot->buf);\n+\n+\t\tstring_list_split_in_place(&traits, p, ' ', -1);\n+\n+\t\tif (unsorted_string_list_has_string(&traits, \"fully-peeled\"))\n+\t\t\tsnapshot->peeled = PEELED_FULLY;\n+\t\telse if (unsorted_string_list_has_string(&traits, \"peeled\"))\n+\t\t\tsnapshot->peeled = PEELED_TAGS;\n+\n+\t\t*sorted = unsorted_string_list_has_string(&traits, \"sorted\");\n+\n+\t\t/* perhaps other traits later as well */\n+\n+\t\t/* The \"+ 1\" is for the LF character. */\n+\t\tsnapshot->start = eol + 1;\n+\n+\t\tstring_list_clear(&traits, 0);\n+\t\tfree(tmp);\n+\t}\n+\n+\treturn 0;\n+}\n+\n+int packed_read_raw_ref_v1(struct packed_ref_store *refs, struct snapshot *snapshot,\n+\t\t\t   const char *refname, struct object_id *oid,\n+\t\t\t   unsigned int *type, int *failure_errno)\n+{\n+\tconst char *rec;\n+\n+\t*type = 0;\n+\n+\trec = find_reference_location_v1(snapshot, refname, 1);\n+\n+\tif (!rec) {\n+\t\t/* refname is not a packed reference. */\n+\t\t*failure_errno = ENOENT;\n+\t\treturn -1;\n+\t}\n+\n+\tif (get_oid_hex(rec, oid))\n+\t\tdie_invalid_line(refs->path, rec, snapshot->eof - rec);\n+\n+\t*type = REF_ISPACKED;\n+\treturn 0;\n+}\n+\n+int next_record_v1(struct packed_ref_iterator *iter)\n+{\n+\tconst char *p = iter->pos, *eol;\n+\n+\tstrbuf_reset(&iter->refname_buf);\n+\n+\tif (iter->pos == iter->eof)\n+\t\treturn ITER_DONE;\n+\n+\titer->base.flags = REF_ISPACKED;\n+\n+\tif (iter->eof - p < the_hash_algo->hexsz + 2 ||\n+\t    parse_oid_hex(p, &iter->oid, &p) ||\n+\t    !isspace(*p++))\n+\t\tdie_invalid_line(iter->snapshot->refs->path,\n+\t\t\t\t iter->pos, iter->eof - iter->pos);\n+\n+\teol = memchr(p, '\\n', iter->eof - p);\n+\tif (!eol)\n+\t\tdie_unterminated_line(iter->snapshot->refs->path,\n+\t\t\t\t      iter->pos, iter->eof - iter->pos);\n+\n+\tstrbuf_add(&iter->refname_buf, p, eol - p);\n+\titer->base.refname = iter->refname_buf.buf;\n+\n+\tif (check_refname_format(iter->base.refname, REFNAME_ALLOW_ONELEVEL)) {\n+\t\tif (!refname_is_safe(iter->base.refname))\n+\t\t\tdie(\"packed refname is dangerous: %s\",\n+\t\t\t    iter->base.refname);\n+\t\toidclr(&iter->oid);\n+\t\titer->base.flags |= REF_BAD_NAME | REF_ISBROKEN;\n+\t}\n+\tif (iter->snapshot->peeled == PEELED_FULLY ||\n+\t    (iter->snapshot->peeled == PEELED_TAGS &&\n+\t     starts_with(iter->base.refname, \"refs/tags/\")))\n+\t\titer->base.flags |= REF_KNOWS_PEELED;\n+\n+\titer->pos = eol + 1;\n+\n+\tif (iter->pos < iter->eof && *iter->pos == '^') {\n+\t\tp = iter->pos + 1;\n+\t\tif (iter->eof - p < the_hash_algo->hexsz + 1 ||\n+\t\t    parse_oid_hex(p, &iter->peeled, &p) ||\n+\t\t    *p++ != '\\n')\n+\t\t\tdie_invalid_line(iter->snapshot->refs->path,\n+\t\t\t\t\t iter->pos, iter->eof - iter->pos);\n+\t\titer->pos = p;\n+\n+\t\t/*\n+\t\t * Regardless of what the file header said, we\n+\t\t * definitely know the value of *this* reference. But\n+\t\t * we suppress it if the reference is broken:\n+\t\t */\n+\t\tif ((iter->base.flags & REF_ISBROKEN)) {\n+\t\t\toidclr(&iter->peeled);\n+\t\t\titer->base.flags &= ~REF_KNOWS_PEELED;\n+\t\t} else {\n+\t\t\titer->base.flags |= REF_KNOWS_PEELED;\n+\t\t}\n+\t} else {\n+\t\toidclr(&iter->peeled);\n+\t}\n+\n+\treturn ITER_OK;\n+}\n+\n+/*\n+ * The packed-refs header line that we write out. Perhaps other traits\n+ * will be added later.\n+ *\n+ * Note that earlier versions of Git used to parse these traits by\n+ * looking for \" trait \" in the line. For this reason, the space after\n+ * the colon and the trailing space are required.\n+ */\n+static const char PACKED_REFS_HEADER[] =\n+\t\"# pack-refs with: peeled fully-peeled sorted \\n\";\n+\n+int write_packed_file_header_v1(FILE *out)\n+{\n+\treturn fprintf(out, \"%s\", PACKED_REFS_HEADER);\n+}\n+\n+/*\n+ * Write an entry to the packed-refs file for the specified refname.\n+ * If peeled is non-NULL, write it as the entry's peeled value. On\n+ * error, return a nonzero value and leave errno set at the value left\n+ * by the failing call to `fprintf()`.\n+ */\n+int write_packed_entry_v1(FILE *fh, const char *refname,\n+\t\t\t  const struct object_id *oid,\n+\t\t\t  const struct object_id *peeled)\n+{\n+\tif (fprintf(fh, \"%s %s\\n\", oid_to_hex(oid), refname) < 0 ||\n+\t    (peeled && fprintf(fh, \"^%s\\n\", oid_to_hex(peeled)) < 0))\n+\t\treturn -1;\n+\n+\treturn 0;\n+}\n-- \ngitgitgadget\n\n"},{"id":"466697","messageId":"cca445cd7d8ae0ab891fcbe1e2087b3552ac3576.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 14/30] packed-backend: extract iterator/updates merge","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:48Z","receivedAt":"2022-11-07T18:37:12Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nTBD\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n refs/packed-backend.c | 117 +++++++++++++++++++++++-------------------\n 1 file changed, 64 insertions(+), 53 deletions(-)\n\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex ef8060f2e08..0dff78f02c8 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -535,58 +535,13 @@ static void add_write_error(struct packed_ref_store *refs, struct strbuf *err)\n \t\t    get_tempfile_path(refs->tempfile), strerror(errno));\n }\n \n-/*\n- * Write the packed refs from the current snapshot to the packed-refs\n- * tempfile, incorporating any changes from `updates`. `updates` must\n- * be a sorted string list whose keys are the refnames and whose util\n- * values are `struct ref_update *`. On error, rollback the tempfile,\n- * write an error message to `err`, and return a nonzero value.\n- *\n- * The packfile must be locked before calling this function and will\n- * remain locked when it is done.\n- */\n-static int write_with_updates(struct packed_ref_store *refs,\n-\t\t\t      struct string_list *updates,\n-\t\t\t      struct strbuf *err)\n+static int merge_iterator_and_updates(struct packed_ref_store *refs,\n+\t\t\t\t      struct string_list *updates,\n+\t\t\t\t      struct strbuf *err,\n+\t\t\t\t      FILE *out)\n {\n \tstruct ref_iterator *iter = NULL;\n-\tsize_t i;\n-\tint ok;\n-\tFILE *out;\n-\tstruct strbuf sb = STRBUF_INIT;\n-\tchar *packed_refs_path;\n-\n-\tif (!is_lock_file_locked(&refs->lock))\n-\t\tBUG(\"write_with_updates() called while unlocked\");\n-\n-\t/*\n-\t * If packed-refs is a symlink, we want to overwrite the\n-\t * symlinked-to file, not the symlink itself. Also, put the\n-\t * staging file next to it:\n-\t */\n-\tpacked_refs_path = get_locked_file_path(&refs->lock);\n-\tstrbuf_addf(&sb, \"%s.new\", packed_refs_path);\n-\tfree(packed_refs_path);\n-\trefs->tempfile = create_tempfile(sb.buf);\n-\tif (!refs->tempfile) {\n-\t\tstrbuf_addf(err, \"unable to create file %s: %s\",\n-\t\t\t    sb.buf, strerror(errno));\n-\t\tstrbuf_release(&sb);\n-\t\treturn -1;\n-\t}\n-\tstrbuf_release(&sb);\n-\n-\tout = fdopen_tempfile(refs->tempfile, \"w\");\n-\tif (!out) {\n-\t\tstrbuf_addf(err, \"unable to fdopen packed-refs tempfile: %s\",\n-\t\t\t    strerror(errno));\n-\t\tgoto error;\n-\t}\n-\n-\tif (write_packed_file_header_v1(out) < 0) {\n-\t\tadd_write_error(refs, err);\n-\t\tgoto error;\n-\t}\n+\tint ok, i;\n \n \t/*\n \t * We iterate in parallel through the current list of refs and\n@@ -713,6 +668,65 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t\t}\n \t}\n \n+error:\n+\tif (iter)\n+\t\tref_iterator_abort(iter);\n+\treturn ok;\n+}\n+\n+/*\n+ * Write the packed refs from the current snapshot to the packed-refs\n+ * tempfile, incorporating any changes from `updates`. `updates` must\n+ * be a sorted string list whose keys are the refnames and whose util\n+ * values are `struct ref_update *`. On error, rollback the tempfile,\n+ * write an error message to `err`, and return a nonzero value.\n+ *\n+ * The packfile must be locked before calling this function and will\n+ * remain locked when it is done.\n+ */\n+static int write_with_updates(struct packed_ref_store *refs,\n+\t\t\t      struct string_list *updates,\n+\t\t\t      struct strbuf *err)\n+{\n+\tint ok;\n+\tFILE *out;\n+\tstruct strbuf sb = STRBUF_INIT;\n+\tchar *packed_refs_path;\n+\n+\tif (!is_lock_file_locked(&refs->lock))\n+\t\tBUG(\"write_with_updates() called while unlocked\");\n+\n+\t/*\n+\t * If packed-refs is a symlink, we want to overwrite the\n+\t * symlinked-to file, not the symlink itself. Also, put the\n+\t * staging file next to it:\n+\t */\n+\tpacked_refs_path = get_locked_file_path(&refs->lock);\n+\tstrbuf_addf(&sb, \"%s.new\", packed_refs_path);\n+\tfree(packed_refs_path);\n+\trefs->tempfile = create_tempfile(sb.buf);\n+\tif (!refs->tempfile) {\n+\t\tstrbuf_addf(err, \"unable to create file %s: %s\",\n+\t\t\t    sb.buf, strerror(errno));\n+\t\tstrbuf_release(&sb);\n+\t\treturn -1;\n+\t}\n+\tstrbuf_release(&sb);\n+\n+\tout = fdopen_tempfile(refs->tempfile, \"w\");\n+\tif (!out) {\n+\t\tstrbuf_addf(err, \"unable to fdopen packed-refs tempfile: %s\",\n+\t\t\t    strerror(errno));\n+\t\tgoto error;\n+\t}\n+\n+\tif (write_packed_file_header_v1(out) < 0) {\n+\t\tadd_write_error(refs, err);\n+\t\tgoto error;\n+\t}\n+\n+\tok = merge_iterator_and_updates(refs, updates, err, out);\n+\n \tif (ok != ITER_DONE) {\n \t\tstrbuf_addstr(err, \"unable to write packed-refs file: \"\n \t\t\t      \"error iterating over old contents\");\n@@ -732,9 +746,6 @@ static int write_with_updates(struct packed_ref_store *refs,\n \treturn 0;\n \n error:\n-\tif (iter)\n-\t\tref_iterator_abort(iter);\n-\n \tdelete_tempfile(&refs->tempfile);\n \treturn -1;\n }\n-- \ngitgitgadget\n\n"},{"id":"466698","messageId":"7c1f6a1ad609ecd33ceda5655cd8fc02137f3e5e.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 16/30] config: add config values for packed-refs v2","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:50Z","receivedAt":"2022-11-07T18:37:15Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nWhen updating the file format version for something as critical as ref\nstorage, the file format version must come with an extension change. The\nextensions.refFormat config value is a multi-valued config value that\ndefaults to the pair \"files\" and \"packed\".\n\nAdd \"packed-v2\" as a possible value to extensions.refFormat. This\nvalue specifies that the packed-refs file may exist in the version 2\nformat. (If the \"packed\" value does not exist, then the packed-refs file\nmust exist in version 2, not version 1.)\n\nIn order to select version 2 for writing, the user will have two\noptions. First, the user could remove \"packed\" and add \"packed-v2\" to\nthe extensions.refFormat list. This would imply that version 2 is the\nonly format available. However, this also means that version 1 files\nwould be ignored at read time, so this does not allow users to upgrade\nrepositories with existing packed-refs files.\n\nAdd a new refs.packedRefsVersion config option which allows specifying\nwhich version to use during writes. Thus, when both \"packed\" and\n\"packed-v2\" are in the extensions.refFormat list, the user can upgrade\nfrom version 1 to version 2, or downgrade from 2 to 1.\n\nCurrently, the implementation does not use refs.packedRefsVersion, as\nthat is delayed until we have the code to write that file format\nversion. However, we can add the necessary enum values and flag\nconstants to communicate the presence of \"packed-v2\" in the\nextensions.refFormat list.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n Documentation/config.txt            |  2 ++\n Documentation/config/extensions.txt | 27 ++++++++++++++++++++++-----\n Documentation/config/refs.txt       | 13 +++++++++++++\n refs.c                              |  4 +++-\n refs/packed-backend.c               | 17 ++++++++++++++++-\n refs/refs-internal.h                |  5 +++--\n repository.h                        |  1 +\n setup.c                             |  2 ++\n t/t3212-ref-formats.sh              | 19 +++++++++++++++++++\n 9 files changed, 81 insertions(+), 9 deletions(-)\n create mode 100644 Documentation/config/refs.txt\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex 0e93aef8626..e480f99c3e1 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -493,6 +493,8 @@ include::config/rebase.txt[]\n \n include::config/receive.txt[]\n \n+include::config/refs.txt[]\n+\n include::config/remote.txt[]\n \n include::config/remotes.txt[]\ndiff --git a/Documentation/config/extensions.txt b/Documentation/config/extensions.txt\nindex 18071c336d0..05abb821e07 100644\n--- a/Documentation/config/extensions.txt\n+++ b/Documentation/config/extensions.txt\n@@ -35,17 +35,34 @@ indicate the existence of different layers:\n \t`files`, the `packed` format will only be used to group multiple\n \tloose object files upon request via the `git pack-refs` command or\n \tvia the `pack-refs` maintenance task.\n+\n+`packed-v2`;;\n+\tWhen present, references may be stored as a group in a\n+\t`packed-refs` file in its version 2 format. This file is in the\n+\tsame position and interacts with loose refs the same as when the\n+\t`packed` value exists. Both `packed` and `packed-v2` must exist to\n+\tupgrade an existing `packed-refs` file from version 1 to version 2\n+\tor to downgrade from version 2 to version 1. When both are\n+\tpresent, the `refs.packedRefsVersion` config value indicates which\n+\tfile format version is used during writes, but both versions are\n+\tunderstood when reading the file.\n --\n +\n The following combinations are supported by this version of Git:\n +\n --\n-`files` and `packed`;;\n+`files` and (`packed` and/or `packed-v2`);;\n \tThis set of values indicates that references are stored both as\n-\tloose reference files and in the `packed-refs` file in its v1\n-\tformat. Loose references are preferred, and the `packed-refs` file\n-\tis updated only when deleting a reference that is stored in the\n-\t`packed-refs` file or during a `git pack-refs` command.\n+\tloose reference files and in the `packed-refs` file. Loose\n+\treferences are preferred, and the `packed-refs` file is updated\n+\tonly when deleting a reference that is stored in the `packed-refs`\n+\tfile or during a `git pack-refs` command.\n++\n+The presence of `packed` and `packed-v2` specifies whether the `packed-refs`\n+file is allowed to be in its v1 or v2 formats, respectively. When only one\n+is present, Git will refuse to read the `packed-refs` file that do not\n+match the expected format. When both are present, the `refs.packedRefsVersion`\n+config option indicates which file format is used during writes.\n \n `files`;;\n \tWhen only this value is present, Git will ignore the `packed-refs`\ndiff --git a/Documentation/config/refs.txt b/Documentation/config/refs.txt\nnew file mode 100644\nindex 00000000000..b2fdb2923f7\n--- /dev/null\n+++ b/Documentation/config/refs.txt\n@@ -0,0 +1,13 @@\n+refs.packedRefsVersion::\n+\tSpecifies the file format version to use when writing a `packed-refs`\n+\tfile. Defaults to `1`.\n++\n+The only other value currently allowed is `2`, which uses a structured file\n+format to result in a smaller `packed-refs` file. In order to write this\n+file format version, the repository must also have the `packed-v2` extension\n+enabled. The most typical setup will include the\n+`core.repositoryFormatVersion=1` config value and the `extensions.refFormat`\n+key will have three values: `files`, `packed`, and `packed-v2`.\n++\n+If `extensions.refFormat` has the value `packed-v2` and not `packed`, then\n+`refs.packedRefsVersion` defaults to `2`.\ndiff --git a/refs.c b/refs.c\nindex 21441ddb162..bf53d1445f2 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -1987,6 +1987,8 @@ static int add_ref_format_flags(enum ref_format_flags flags, int caps) {\n \t\tcaps |= REF_STORE_FORMAT_FILES;\n \tif (flags & REF_FORMAT_PACKED)\n \t\tcaps |= REF_STORE_FORMAT_PACKED;\n+\tif (flags & REF_FORMAT_PACKED_V2)\n+\t\tcaps |= REF_STORE_FORMAT_PACKED_V2;\n \n \treturn caps;\n }\n@@ -2006,7 +2008,7 @@ static struct ref_store *ref_store_init(struct repository *repo,\n \tflags = add_ref_format_flags(repo->ref_format, flags);\n \n \tif (!(flags & REF_STORE_FORMAT_FILES) &&\n-\t    (flags & REF_STORE_FORMAT_PACKED))\n+\t    packed_refs_enabled(flags))\n \t\tbe_name = \"packed\";\n \n \tbe = find_ref_storage_backend(be_name);\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex 7ed9475812c..655aab939be 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -236,7 +236,13 @@ static struct snapshot *create_snapshot(struct packed_ref_store *refs)\n \tif (!load_contents(snapshot))\n \t\treturn snapshot;\n \n-\tif (parse_packed_format_v1_header(refs, snapshot, &sorted)) {\n+\t/*\n+\t * If this is a v1 file format, but we don't have v1 enabled,\n+\t * then ignore it the same way we would as if we didn't\n+\t * understand it.\n+\t */\n+\tif (parse_packed_format_v1_header(refs, snapshot, &sorted) ||\n+\t    !(refs->store_flags & REF_STORE_FORMAT_PACKED)) {\n \t\tclear_snapshot(refs);\n \t\treturn NULL;\n \t}\n@@ -310,6 +316,12 @@ static int packed_read_raw_ref(struct ref_store *ref_store, const char *refname,\n \t\tpacked_downcast(ref_store, REF_STORE_READ, \"read_raw_ref\");\n \tstruct snapshot *snapshot = get_snapshot(refs);\n \n+\tif (!snapshot) {\n+\t\t/* refname is not a packed reference. */\n+\t\t*failure_errno = ENOENT;\n+\t\treturn -1;\n+\t}\n+\n \treturn packed_read_raw_ref_v1(refs, snapshot, refname,\n \t\t\t\t      oid, type, failure_errno);\n }\n@@ -410,6 +422,9 @@ static struct ref_iterator *packed_ref_iterator_begin(\n \t */\n \tsnapshot = get_snapshot(refs);\n \n+\tif (!snapshot)\n+\t\treturn empty_ref_iterator_begin();\n+\n \tif (prefix && *prefix)\n \t\tstart = find_reference_location_v1(snapshot, prefix, 0);\n \telse\ndiff --git a/refs/refs-internal.h b/refs/refs-internal.h\nindex a1900848a87..39b93fce97c 100644\n--- a/refs/refs-internal.h\n+++ b/refs/refs-internal.h\n@@ -522,11 +522,12 @@ struct ref_store;\n \t\t\t\t REF_STORE_MAIN)\n \n #define REF_STORE_FORMAT_FILES\t\t(1 << 8) /* can use loose ref files */\n-#define REF_STORE_FORMAT_PACKED\t\t(1 << 9) /* can use packed-refs file */\n+#define REF_STORE_FORMAT_PACKED\t\t(1 << 9) /* can use v1 packed-refs file */\n+#define REF_STORE_FORMAT_PACKED_V2\t(1 << 10) /* can use v2 packed-refs file */\n \n static inline int packed_refs_enabled(int flags)\n {\n-\treturn flags & REF_STORE_FORMAT_PACKED;\n+\treturn flags & (REF_STORE_FORMAT_PACKED | REF_STORE_FORMAT_PACKED_V2);\n }\n \n /*\ndiff --git a/repository.h b/repository.h\nindex 5cfde4282c5..ee3a90efc72 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -64,6 +64,7 @@ struct repo_path_cache {\n enum ref_format_flags {\n \tREF_FORMAT_FILES = (1 << 0),\n \tREF_FORMAT_PACKED = (1 << 1),\n+\tREF_FORMAT_PACKED_V2 = (1 << 2),\n };\n \n struct repository {\ndiff --git a/setup.c b/setup.c\nindex a5e63479558..72bfa289ade 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -582,6 +582,8 @@ static enum extension_result handle_extension(const char *var,\n \t\t\tdata->ref_format |= REF_FORMAT_FILES;\n \t\telse if (!strcmp(value, \"packed\"))\n \t\t\tdata->ref_format |= REF_FORMAT_PACKED;\n+\t\telse if (!strcmp(value, \"packed-v2\"))\n+\t\t\tdata->ref_format |= REF_FORMAT_PACKED_V2;\n \t\telse\n \t\t\treturn error(_(\"invalid value for '%s': '%s'\"),\n \t\t\t\t     \"extensions.refFormat\", value);\ndiff --git a/t/t3212-ref-formats.sh b/t/t3212-ref-formats.sh\nindex 67aa65c116f..cd1b399bbb8 100755\n--- a/t/t3212-ref-formats.sh\n+++ b/t/t3212-ref-formats.sh\n@@ -56,4 +56,23 @@ test_expect_success 'extensions.refFormat=files only' '\n \t)\n '\n \n+test_expect_success 'extensions.refFormat=files,packed-v2' '\n+\ttest_commit Q &&\n+\tgit pack-refs --all &&\n+\tgit init no-packed-v1 &&\n+\t(\n+\t\tcd no-packed-v1 &&\n+\t\tgit config core.repositoryFormatVersion 1 &&\n+\t\tgit config extensions.refFormat files &&\n+\t\tgit config --add extensions.refFormat packed-v2 &&\n+\t\ttest_commit A &&\n+\t\ttest_commit B &&\n+\n+\t\t# Refuse to parse a v1 packed-refs file.\n+\t\tcp ../.git/packed-refs .git/packed-refs &&\n+\t\ttest_must_fail git rev-parse refs/tags/Q &&\n+\t\trm -f .git/packed-refs\n+\t)\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"466699","messageId":"a3819f665977194a8f44061d4a52f86d44cf6e0a.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 15/30] packed-backend: create abstraction for writing refs","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:49Z","receivedAt":"2022-11-07T18:37:17Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe packed-refs file is a plaintext file format that starts with a\nheader line, then each ref is given as one or two lines (two if there is\na peeled value). These lines are written as part of a sequence of\nupdates which are merged with the existing ref iterator in\nmerge_iterator_and_updates(). That method is currently tied directly to\nwrite_packed_entry_v1().\n\nWhen creating a new version of the packed-file format, it would be\nvaluable to use this merging logic in an identical way. Create a new\nfunction pointer type, write_ref_fn, and use that type in\nmerge_iterator_and_updates().\n\nNotably, the function pointer type no longer depends on a FILE pointer,\nbut instead takes an arbitrary \"void *write_data\" parameter. This\nflexibility will be critical in the future, since the planned v2 format\nwill use the chunk-format API and need a more complicated structure than\nthe output FILE.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n refs/packed-backend.c   | 26 +++++++++++++++-----------\n refs/packed-backend.h   | 16 ++++++++++++++--\n refs/packed-format-v1.c |  7 +++++--\n 3 files changed, 34 insertions(+), 15 deletions(-)\n\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex 0dff78f02c8..7ed9475812c 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -535,10 +535,11 @@ static void add_write_error(struct packed_ref_store *refs, struct strbuf *err)\n \t\t    get_tempfile_path(refs->tempfile), strerror(errno));\n }\n \n-static int merge_iterator_and_updates(struct packed_ref_store *refs,\n-\t\t\t\t      struct string_list *updates,\n-\t\t\t\t      struct strbuf *err,\n-\t\t\t\t      FILE *out)\n+int merge_iterator_and_updates(struct packed_ref_store *refs,\n+\t\t\t       struct string_list *updates,\n+\t\t\t       struct strbuf *err,\n+\t\t\t       write_ref_fn write_fn,\n+\t\t\t       void *write_data)\n {\n \tstruct ref_iterator *iter = NULL;\n \tint ok, i;\n@@ -634,9 +635,10 @@ static int merge_iterator_and_updates(struct packed_ref_store *refs,\n \t\t\tstruct object_id peeled;\n \t\t\tint peel_error = ref_iterator_peel(iter, &peeled);\n \n-\t\t\tif (write_packed_entry_v1(out, iter->refname,\n-\t\t\t\t\t\t  iter->oid,\n-\t\t\t\t\t\t  peel_error ? NULL : &peeled)) {\n+\t\t\tif (write_fn(iter->refname,\n+\t\t\t\t     iter->oid,\n+\t\t\t\t     peel_error ? NULL : &peeled,\n+\t\t\t\t     write_data)) {\n \t\t\t\tadd_write_error(refs, err);\n \t\t\t\tgoto error;\n \t\t\t}\n@@ -657,9 +659,10 @@ static int merge_iterator_and_updates(struct packed_ref_store *refs,\n \t\t\tint peel_error = peel_object(&update->new_oid,\n \t\t\t\t\t\t     &peeled);\n \n-\t\t\tif (write_packed_entry_v1(out, update->refname,\n-\t\t\t\t\t\t  &update->new_oid,\n-\t\t\t\t\t\t  peel_error ? NULL : &peeled)) {\n+\t\t\tif (write_fn(update->refname,\n+\t\t\t\t     &update->new_oid,\n+\t\t\t\t     peel_error ? NULL : &peeled,\n+\t\t\t\t     write_data)) {\n \t\t\t\tadd_write_error(refs, err);\n \t\t\t\tgoto error;\n \t\t\t}\n@@ -725,7 +728,8 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t\tgoto error;\n \t}\n \n-\tok = merge_iterator_and_updates(refs, updates, err, out);\n+\tok = merge_iterator_and_updates(refs, updates, err,\n+\t\t\t\t\twrite_packed_entry_v1, out);\n \n \tif (ok != ITER_DONE) {\n \t\tstrbuf_addstr(err, \"unable to write packed-refs file: \"\ndiff --git a/refs/packed-backend.h b/refs/packed-backend.h\nindex 143ed6d4f6c..b6908bb002c 100644\n--- a/refs/packed-backend.h\n+++ b/refs/packed-backend.h\n@@ -192,6 +192,17 @@ struct packed_ref_iterator {\n \tunsigned int flags;\n };\n \n+typedef int (*write_ref_fn)(const char *refname,\n+\t\t\t    const struct object_id *oid,\n+\t\t\t    const struct object_id *peeled,\n+\t\t\t    void *write_data);\n+\n+int merge_iterator_and_updates(struct packed_ref_store *refs,\n+\t\t\t       struct string_list *updates,\n+\t\t\t       struct strbuf *err,\n+\t\t\t       write_ref_fn write_fn,\n+\t\t\t       void *write_data);\n+\n /**\n  * Parse the buffer at the given snapshot to verify that it is a\n  * packed-refs file in version 1 format. Update the snapshot->peeled\n@@ -227,8 +238,9 @@ void verify_buffer_safe_v1(struct snapshot *snapshot);\n void sort_snapshot_v1(struct snapshot *snapshot);\n int write_packed_file_header_v1(FILE *out);\n int next_record_v1(struct packed_ref_iterator *iter);\n-int write_packed_entry_v1(FILE *fh, const char *refname,\n+int write_packed_entry_v1(const char *refname,\n \t\t\t  const struct object_id *oid,\n-\t\t\t  const struct object_id *peeled);\n+\t\t\t  const struct object_id *peeled,\n+\t\t\t  void *write_data);\n \n #endif /* REFS_PACKED_BACKEND_H */\ndiff --git a/refs/packed-format-v1.c b/refs/packed-format-v1.c\nindex ef9e6618c89..2d071567c02 100644\n--- a/refs/packed-format-v1.c\n+++ b/refs/packed-format-v1.c\n@@ -441,10 +441,13 @@ int write_packed_file_header_v1(FILE *out)\n  * error, return a nonzero value and leave errno set at the value left\n  * by the failing call to `fprintf()`.\n  */\n-int write_packed_entry_v1(FILE *fh, const char *refname,\n+int write_packed_entry_v1(const char *refname,\n \t\t\t  const struct object_id *oid,\n-\t\t\t  const struct object_id *peeled)\n+\t\t\t  const struct object_id *peeled,\n+\t\t\t  void *write_data)\n {\n+\tFILE *fh = write_data;\n+\n \tif (fprintf(fh, \"%s %s\\n\", oid_to_hex(oid), refname) < 0 ||\n \t    (peeled && fprintf(fh, \"^%s\\n\", oid_to_hex(peeled)) < 0))\n \t\treturn -1;\n-- \ngitgitgadget\n\n"},{"id":"466700","messageId":"eb6152f1d39310b2e0527bbe199e25b941744753.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 17/30] packed-backend: create shell of v2 writes","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:51Z","receivedAt":"2022-11-07T18:37:21Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n Makefile                |  1 +\n refs/packed-backend.c   | 75 +++++++++++++++++++++++++++++++++++------\n refs/packed-backend.h   |  7 ++++\n refs/packed-format-v2.c | 38 +++++++++++++++++++++\n 4 files changed, 110 insertions(+), 11 deletions(-)\n create mode 100644 refs/packed-format-v2.c\n\ndiff --git a/Makefile b/Makefile\nindex 3dc887941d4..16cd245e0ad 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1058,6 +1058,7 @@ LIB_OBJS += refs/files-backend.o\n LIB_OBJS += refs/iterator.o\n LIB_OBJS += refs/packed-backend.o\n LIB_OBJS += refs/packed-format-v1.o\n+LIB_OBJS += refs/packed-format-v2.o\n LIB_OBJS += refs/ref-cache.o\n LIB_OBJS += refspec.o\n LIB_OBJS += remote.o\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex 655aab939be..09f7b74584f 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -692,6 +692,45 @@ error:\n \treturn ok;\n }\n \n+static int write_with_updates_v1(struct packed_ref_store *refs,\n+\t\t\t\t struct string_list *updates,\n+\t\t\t\t struct strbuf *err)\n+{\n+\tFILE *out;\n+\n+\tout = fdopen_tempfile(refs->tempfile, \"w\");\n+\tif (!out) {\n+\t\tstrbuf_addf(err, \"unable to fdopen packed-refs tempfile: %s\",\n+\t\t\t    strerror(errno));\n+\t\tgoto error;\n+\t}\n+\n+\tif (write_packed_file_header_v1(out) < 0) {\n+\t\tadd_write_error(refs, err);\n+\t\tgoto error;\n+\t}\n+\n+\treturn merge_iterator_and_updates(refs, updates, err,\n+\t\t\t\t\t  write_packed_entry_v1, out);\n+\n+error:\n+\treturn -1;\n+}\n+\n+static int write_with_updates_v2(struct packed_ref_store *refs,\n+\t\t\t\t struct string_list *updates,\n+\t\t\t\t struct strbuf *err)\n+{\n+\tstruct write_packed_refs_v2_context *ctx = create_v2_context(refs, updates, err);\n+\tint ok = -1;\n+\n+\tif ((ok = write_packed_refs_v2(ctx)) < 0)\n+\t\tadd_write_error(refs, err);\n+\n+\tfree_v2_context(ctx);\n+\treturn ok;\n+}\n+\n /*\n  * Write the packed refs from the current snapshot to the packed-refs\n  * tempfile, incorporating any changes from `updates`. `updates` must\n@@ -707,9 +746,9 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t\t\t      struct strbuf *err)\n {\n \tint ok;\n-\tFILE *out;\n \tstruct strbuf sb = STRBUF_INIT;\n \tchar *packed_refs_path;\n+\tint version;\n \n \tif (!is_lock_file_locked(&refs->lock))\n \t\tBUG(\"write_with_updates() called while unlocked\");\n@@ -731,21 +770,35 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t}\n \tstrbuf_release(&sb);\n \n-\tout = fdopen_tempfile(refs->tempfile, \"w\");\n-\tif (!out) {\n-\t\tstrbuf_addf(err, \"unable to fdopen packed-refs tempfile: %s\",\n-\t\t\t    strerror(errno));\n-\t\tgoto error;\n+\tif (git_config_get_int(\"refs.packedrefsversion\", &version)) {\n+\t\t/*\n+\t\t * Set the default depending on the current extension\n+\t\t * list. Default to version 1 if available, but allow a\n+\t\t * default of 2 if only \"packed-v2\" exists.\n+\t\t */\n+\t\tif (refs->store_flags & REF_STORE_FORMAT_PACKED)\n+\t\t\tversion = 1;\n+\t\telse if (refs->store_flags & REF_STORE_FORMAT_PACKED_V2)\n+\t\t\tversion = 2;\n+\t\telse\n+\t\t\tBUG(\"writing a packed-refs file without an extension\");\n \t}\n \n-\tif (write_packed_file_header_v1(out) < 0) {\n-\t\tadd_write_error(refs, err);\n+\tswitch (version) {\n+\tcase 1:\n+\t\tok = write_with_updates_v1(refs, updates, err);\n+\t\tbreak;\n+\n+\tcase 2:\n+\t\tok = write_with_updates_v2(refs, updates, err);\n+\t\tbreak;\n+\n+\tdefault:\n+\t\tstrbuf_addf(err, \"unknown packed-refs version: %d\",\n+\t\t\t    version);\n \t\tgoto error;\n \t}\n \n-\tok = merge_iterator_and_updates(refs, updates, err,\n-\t\t\t\t\twrite_packed_entry_v1, out);\n-\n \tif (ok != ITER_DONE) {\n \t\tstrbuf_addstr(err, \"unable to write packed-refs file: \"\n \t\t\t      \"error iterating over old contents\");\ndiff --git a/refs/packed-backend.h b/refs/packed-backend.h\nindex b6908bb002c..e76f26bfc46 100644\n--- a/refs/packed-backend.h\n+++ b/refs/packed-backend.h\n@@ -243,4 +243,11 @@ int write_packed_entry_v1(const char *refname,\n \t\t\t  const struct object_id *peeled,\n \t\t\t  void *write_data);\n \n+struct write_packed_refs_v2_context;\n+struct write_packed_refs_v2_context *create_v2_context(struct packed_ref_store *refs,\n+\t\t\t\t\t\t       struct string_list *updates,\n+\t\t\t\t\t\t       struct strbuf *err);\n+int write_packed_refs_v2(struct write_packed_refs_v2_context *ctx);\n+void free_v2_context(struct write_packed_refs_v2_context *ctx);\n+\n #endif /* REFS_PACKED_BACKEND_H */\ndiff --git a/refs/packed-format-v2.c b/refs/packed-format-v2.c\nnew file mode 100644\nindex 00000000000..ecf3cc93694\n--- /dev/null\n+++ b/refs/packed-format-v2.c\n@@ -0,0 +1,38 @@\n+#include \"../cache.h\"\n+#include \"../config.h\"\n+#include \"../refs.h\"\n+#include \"refs-internal.h\"\n+#include \"packed-backend.h\"\n+#include \"../iterator.h\"\n+#include \"../lockfile.h\"\n+#include \"../chdir-notify.h\"\n+\n+struct write_packed_refs_v2_context {\n+\tstruct packed_ref_store *refs;\n+\tstruct string_list *updates;\n+\tstruct strbuf *err;\n+};\n+\n+struct write_packed_refs_v2_context *create_v2_context(struct packed_ref_store *refs,\n+\t\t\t\t\t\t       struct string_list *updates,\n+\t\t\t\t\t\t       struct strbuf *err)\n+{\n+\tstruct write_packed_refs_v2_context *ctx;\n+\tCALLOC_ARRAY(ctx, 1);\n+\n+\tctx->refs = refs;\n+\tctx->updates = updates;\n+\tctx->err = err;\n+\n+\treturn ctx;\n+}\n+\n+int write_packed_refs_v2(struct write_packed_refs_v2_context *ctx)\n+{\n+\treturn 0;\n+}\n+\n+void free_v2_context(struct write_packed_refs_v2_context *ctx)\n+{\n+\tfree(ctx);\n+}\n-- \ngitgitgadget\n\n"},{"id":"466701","messageId":"740c2f6e6d1e628a84dc4e1927fef70b5d8d624c.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 18/30] packed-refs: write file format version 2","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:52Z","receivedAt":"2022-11-07T18:37:23Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nTODO: add writing tests.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n refs/packed-backend.c   |   3 +-\n refs/packed-format-v2.c | 108 ++++++++++++++++++++++++++++++++++++++++\n t/t3212-ref-formats.sh  |   6 ++-\n 3 files changed, 115 insertions(+), 2 deletions(-)\n\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex 09f7b74584f..3429e63620a 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -790,7 +790,8 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t\tbreak;\n \n \tcase 2:\n-\t\tok = write_with_updates_v2(refs, updates, err);\n+\t\t/* Convert the normal error codes to ITER_DONE. */\n+\t\tok = write_with_updates_v2(refs, updates, err) ? -2 : ITER_DONE;\n \t\tbreak;\n \n \tdefault:\ndiff --git a/refs/packed-format-v2.c b/refs/packed-format-v2.c\nindex ecf3cc93694..044cc9f629a 100644\n--- a/refs/packed-format-v2.c\n+++ b/refs/packed-format-v2.c\n@@ -6,11 +6,30 @@\n #include \"../iterator.h\"\n #include \"../lockfile.h\"\n #include \"../chdir-notify.h\"\n+#include \"../chunk-format.h\"\n+#include \"../csum-file.h\"\n+\n+#define OFFSET_IS_PEELED (((uint64_t)1) << 63)\n+\n+#define PACKED_REFS_SIGNATURE          0x50524546 /* \"PREF\" */\n+#define CHREFS_CHUNKID_OFFSETS         0x524F4646 /* \"ROFF\" */\n+#define CHREFS_CHUNKID_REFS            0x52454653 /* \"REFS\" */\n \n struct write_packed_refs_v2_context {\n \tstruct packed_ref_store *refs;\n \tstruct string_list *updates;\n \tstruct strbuf *err;\n+\n+\tstruct hashfile *f;\n+\tstruct chunkfile *cf;\n+\n+\t/*\n+\t * As we stream the ref names to the refs chunk, store these\n+\t * values in-memory. These arrays are populated one for every ref.\n+\t */\n+\tuint64_t *offsets;\n+\tsize_t nr;\n+\tsize_t offsets_alloc;\n };\n \n struct write_packed_refs_v2_context *create_v2_context(struct packed_ref_store *refs,\n@@ -24,15 +43,104 @@ struct write_packed_refs_v2_context *create_v2_context(struct packed_ref_store *\n \tctx->updates = updates;\n \tctx->err = err;\n \n+\tif (!fdopen_tempfile(refs->tempfile, \"w\")) {\n+\t\tstrbuf_addf(err, \"unable to fdopen packed-refs tempfile: %s\",\n+\t\t\t    strerror(errno));\n+\t\treturn ctx;\n+\t}\n+\n+\tctx->f = hashfd(refs->tempfile->fd, refs->tempfile->filename.buf);\n+\tctx->cf = init_chunkfile(ctx->f);\n+\n \treturn ctx;\n }\n \n+static int write_packed_entry_v2(const char *refname,\n+\t\t\t\t const struct object_id *oid,\n+\t\t\t\t const struct object_id *peeled,\n+\t\t\t\t void *write_data)\n+{\n+\tstruct write_packed_refs_v2_context *ctx = write_data;\n+\tsize_t reflen = strlen(refname) + 1;\n+\tsize_t i = ctx->nr;\n+\n+\tALLOC_GROW(ctx->offsets, i + 1, ctx->offsets_alloc);\n+\n+\t/* Write entire ref, including null terminator. */\n+\thashwrite(ctx->f, refname, reflen);\n+\thashwrite(ctx->f, oid->hash, the_hash_algo->rawsz);\n+\tif (peeled)\n+\t\thashwrite(ctx->f, peeled->hash, the_hash_algo->rawsz);\n+\n+\tif (i)\n+\t\tctx->offsets[i] = (ctx->offsets[i - 1] & (~OFFSET_IS_PEELED));\n+\telse\n+\t\tctx->offsets[i] = 0;\n+\tctx->offsets[i] += reflen + the_hash_algo->rawsz;\n+\n+\tif (peeled) {\n+\t\tctx->offsets[i] += the_hash_algo->rawsz;\n+\t\tctx->offsets[i] |= OFFSET_IS_PEELED;\n+\t}\n+\n+\tctx->nr++;\n+\treturn 0;\n+}\n+\n+static int write_refs_chunk_refs(struct hashfile *f,\n+\t\t\t\t void *data)\n+{\n+\tstruct write_packed_refs_v2_context *ctx = data;\n+\tint ok;\n+\n+\ttrace2_region_enter(\"refs\", \"refs-chunk\", the_repository);\n+\tok = merge_iterator_and_updates(ctx->refs, ctx->updates, ctx->err,\n+\t\t\t\t\twrite_packed_entry_v2, ctx);\n+\ttrace2_region_leave(\"refs\", \"refs-chunk\", the_repository);\n+\n+\treturn ok != ITER_DONE;\n+}\n+\n+static int write_refs_chunk_offsets(struct hashfile *f,\n+\t\t\t\t    void *data)\n+{\n+\tstruct write_packed_refs_v2_context *ctx = data;\n+\tsize_t i;\n+\n+\ttrace2_region_enter(\"refs\", \"offsets\", the_repository);\n+\tfor (i = 0; i < ctx->nr; i++)\n+\t\thashwrite_be64(f, ctx->offsets[i]);\n+\n+\ttrace2_region_leave(\"refs\", \"offsets\", the_repository);\n+\treturn 0;\n+}\n+\n int write_packed_refs_v2(struct write_packed_refs_v2_context *ctx)\n {\n+\tunsigned char file_hash[GIT_MAX_RAWSZ];\n+\n+\tadd_chunk(ctx->cf, CHREFS_CHUNKID_REFS, 0, write_refs_chunk_refs);\n+\tadd_chunk(ctx->cf, CHREFS_CHUNKID_OFFSETS, 0, write_refs_chunk_offsets);\n+\n+\thashwrite_be32(ctx->f, PACKED_REFS_SIGNATURE);\n+\thashwrite_be32(ctx->f, 2);\n+\thashwrite_be32(ctx->f, the_hash_algo->format_id);\n+\n+\tif (write_chunkfile(ctx->cf, CHUNKFILE_TRAILING_TOC, ctx))\n+\t\tgoto failure;\n+\n+\tfinalize_hashfile(ctx->f, file_hash, FSYNC_COMPONENT_REFERENCE,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_FSYNC);\n+\n \treturn 0;\n+\n+failure:\n+\treturn -1;\n }\n \n void free_v2_context(struct write_packed_refs_v2_context *ctx)\n {\n+\tif (ctx->cf)\n+\t\tfree_chunkfile(ctx->cf);\n \tfree(ctx);\n }\ndiff --git a/t/t3212-ref-formats.sh b/t/t3212-ref-formats.sh\nindex cd1b399bbb8..03c713ac4f6 100755\n--- a/t/t3212-ref-formats.sh\n+++ b/t/t3212-ref-formats.sh\n@@ -71,7 +71,11 @@ test_expect_success 'extensions.refFormat=files,packed-v2' '\n \t\t# Refuse to parse a v1 packed-refs file.\n \t\tcp ../.git/packed-refs .git/packed-refs &&\n \t\ttest_must_fail git rev-parse refs/tags/Q &&\n-\t\trm -f .git/packed-refs\n+\t\trm -f .git/packed-refs &&\n+\n+\t\t# Create a v2 packed-refs file\n+\t\tgit pack-refs --all &&\n+\t\ttest_path_exists .git/packed-refs\n \t)\n '\n \n-- \ngitgitgadget\n\n"},{"id":"466702","messageId":"701c5ad22e7787cfd27628ff613b7849e24fc675.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 19/30] packed-refs: read file format v2","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:53Z","receivedAt":"2022-11-07T18:37:41Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n refs/packed-backend.c   | 129 ++++++++++++++++---------\n refs/packed-backend.h   |  72 ++++++++++++--\n refs/packed-format-v2.c | 209 ++++++++++++++++++++++++++++++++++++++++\n t/t3212-ref-formats.sh  |  17 +++-\n 4 files changed, 372 insertions(+), 55 deletions(-)\n\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex 3429e63620a..549cce1f84a 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -66,7 +66,7 @@ void clear_snapshot_buffer(struct snapshot *snapshot)\n  * Decrease the reference count of `*snapshot`. If it goes to zero,\n  * free `*snapshot` and return true; otherwise return false.\n  */\n-static int release_snapshot(struct snapshot *snapshot)\n+int release_snapshot(struct snapshot *snapshot)\n {\n \tif (!--snapshot->referrers) {\n \t\tstat_validity_clear(&snapshot->validity);\n@@ -142,7 +142,6 @@ static int load_contents(struct snapshot *snapshot)\n {\n \tint fd;\n \tstruct stat st;\n-\tsize_t size;\n \tssize_t bytes_read;\n \n \tif (!packed_refs_enabled(snapshot->refs->store_flags))\n@@ -168,25 +167,25 @@ static int load_contents(struct snapshot *snapshot)\n \n \tif (fstat(fd, &st) < 0)\n \t\tdie_errno(\"couldn't stat %s\", snapshot->refs->path);\n-\tsize = xsize_t(st.st_size);\n+\tsnapshot->buflen = xsize_t(st.st_size);\n \n-\tif (!size) {\n+\tif (!snapshot->buflen) {\n \t\tclose(fd);\n \t\treturn 0;\n-\t} else if (mmap_strategy == MMAP_NONE || size <= SMALL_FILE_SIZE) {\n-\t\tsnapshot->buf = xmalloc(size);\n-\t\tbytes_read = read_in_full(fd, snapshot->buf, size);\n-\t\tif (bytes_read < 0 || bytes_read != size)\n+\t} else if (mmap_strategy == MMAP_NONE || snapshot->buflen <= SMALL_FILE_SIZE) {\n+\t\tsnapshot->buf = xmalloc(snapshot->buflen);\n+\t\tbytes_read = read_in_full(fd, snapshot->buf, snapshot->buflen);\n+\t\tif (bytes_read < 0 || bytes_read != snapshot->buflen)\n \t\t\tdie_errno(\"couldn't read %s\", snapshot->refs->path);\n \t\tsnapshot->mmapped = 0;\n \t} else {\n-\t\tsnapshot->buf = xmmap(NULL, size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\t\tsnapshot->buf = xmmap(NULL, snapshot->buflen, PROT_READ, MAP_PRIVATE, fd, 0);\n \t\tsnapshot->mmapped = 1;\n \t}\n \tclose(fd);\n \n \tsnapshot->start = snapshot->buf;\n-\tsnapshot->eof = snapshot->buf + size;\n+\tsnapshot->eof = snapshot->buf + snapshot->buflen;\n \n \treturn 1;\n }\n@@ -232,46 +231,52 @@ static struct snapshot *create_snapshot(struct packed_ref_store *refs)\n \tsnapshot->refs = refs;\n \tacquire_snapshot(snapshot);\n \tsnapshot->peeled = PEELED_NONE;\n+\tsnapshot->version = 1;\n \n \tif (!load_contents(snapshot))\n \t\treturn snapshot;\n \n-\t/*\n-\t * If this is a v1 file format, but we don't have v1 enabled,\n-\t * then ignore it the same way we would as if we didn't\n-\t * understand it.\n-\t */\n-\tif (parse_packed_format_v1_header(refs, snapshot, &sorted) ||\n-\t    !(refs->store_flags & REF_STORE_FORMAT_PACKED)) {\n-\t\tclear_snapshot(refs);\n-\t\treturn NULL;\n-\t}\n+\tif ((refs->store_flags & REF_STORE_FORMAT_PACKED) &&\n+\t    !detect_packed_format_v2_header(refs, snapshot)) {\n+\t\tparse_packed_format_v1_header(refs, snapshot, &sorted);\n+\t\tsnapshot->version = 1;\n+\t\tverify_buffer_safe_v1(snapshot);\n \n-\tverify_buffer_safe_v1(snapshot);\n+\t\tif (!sorted) {\n+\t\t\tsort_snapshot_v1(snapshot);\n \n-\tif (!sorted) {\n-\t\tsort_snapshot_v1(snapshot);\n+\t\t\t/*\n+\t\t\t* Reordering the records might have moved a short one\n+\t\t\t* to the end of the buffer, so verify the buffer's\n+\t\t\t* safety again:\n+\t\t\t*/\n+\t\t\tverify_buffer_safe_v1(snapshot);\n+\t\t}\n \n-\t\t/*\n-\t\t * Reordering the records might have moved a short one\n-\t\t * to the end of the buffer, so verify the buffer's\n-\t\t * safety again:\n-\t\t */\n-\t\tverify_buffer_safe_v1(snapshot);\n+\t\tif (mmap_strategy != MMAP_OK && snapshot->mmapped) {\n+\t\t\t/*\n+\t\t\t* We don't want to leave the file mmapped, so we are\n+\t\t\t* forced to make a copy now:\n+\t\t\t*/\n+\t\t\tchar *buf_copy = xmalloc(snapshot->buflen);\n+\n+\t\t\tmemcpy(buf_copy, snapshot->start, snapshot->buflen);\n+\t\t\tclear_snapshot_buffer(snapshot);\n+\t\t\tsnapshot->buf = snapshot->start = buf_copy;\n+\t\t\tsnapshot->eof = buf_copy + snapshot->buflen;\n+\t\t}\n+\n+\t\treturn snapshot;\n \t}\n \n-\tif (mmap_strategy != MMAP_OK && snapshot->mmapped) {\n+\tif (refs->store_flags & REF_STORE_FORMAT_PACKED_V2) {\n \t\t/*\n-\t\t * We don't want to leave the file mmapped, so we are\n-\t\t * forced to make a copy now:\n+\t\t * Assume we are in v2 format mode, now.\n+\t\t *\n+\t\t * fill_snapshot_v2() will die() if parsing fails.\n \t\t */\n-\t\tsize_t size = snapshot->eof - snapshot->start;\n-\t\tchar *buf_copy = xmalloc(size);\n-\n-\t\tmemcpy(buf_copy, snapshot->start, size);\n-\t\tclear_snapshot_buffer(snapshot);\n-\t\tsnapshot->buf = snapshot->start = buf_copy;\n-\t\tsnapshot->eof = buf_copy + size;\n+\t\tfill_snapshot_v2(snapshot);\n+\t\tsnapshot->version = 2;\n \t}\n \n \treturn snapshot;\n@@ -322,8 +327,18 @@ static int packed_read_raw_ref(struct ref_store *ref_store, const char *refname,\n \t\treturn -1;\n \t}\n \n-\treturn packed_read_raw_ref_v1(refs, snapshot, refname,\n-\t\t\t\t      oid, type, failure_errno);\n+\tswitch (snapshot->version) {\n+\tcase 1:\n+\t\treturn packed_read_raw_ref_v1(refs, snapshot, refname,\n+\t\t\t\t\t      oid, type, failure_errno);\n+\n+\tcase 2:\n+\t\treturn packed_read_raw_ref_v2(refs, snapshot, refname,\n+\t\t\t\t\t      oid, type, failure_errno);\n+\n+\tdefault:\n+\t\treturn -1;\n+\t}\n }\n \n /*\n@@ -335,7 +350,16 @@ static int packed_read_raw_ref(struct ref_store *ref_store, const char *refname,\n  */\n static int next_record(struct packed_ref_iterator *iter)\n {\n-\treturn next_record_v1(iter);\n+\tswitch (iter->version) {\n+\tcase 1:\n+\t\treturn next_record_v1(iter);\n+\n+\tcase 2:\n+\t\treturn next_record_v2(iter);\n+\n+\tdefault:\n+\t\treturn -1;\n+\t}\n }\n \n static int packed_ref_iterator_advance(struct ref_iterator *ref_iterator)\n@@ -410,6 +434,7 @@ static struct ref_iterator *packed_ref_iterator_begin(\n \tstruct packed_ref_iterator *iter;\n \tstruct ref_iterator *ref_iterator;\n \tunsigned int required_flags = REF_STORE_READ;\n+\tsize_t v2_row = 0;\n \n \tif (!(flags & DO_FOR_EACH_INCLUDE_BROKEN))\n \t\trequired_flags |= REF_STORE_ODB;\n@@ -422,13 +447,21 @@ static struct ref_iterator *packed_ref_iterator_begin(\n \t */\n \tsnapshot = get_snapshot(refs);\n \n-\tif (!snapshot)\n+\tif (!snapshot || snapshot->version < 0 || snapshot->version > 2)\n \t\treturn empty_ref_iterator_begin();\n \n-\tif (prefix && *prefix)\n-\t\tstart = find_reference_location_v1(snapshot, prefix, 0);\n-\telse\n-\t\tstart = snapshot->start;\n+\tif (prefix && *prefix) {\n+\t\tif (snapshot->version == 1)\n+\t\t\tstart = find_reference_location_v1(snapshot, prefix, 0);\n+\t\telse\n+\t\t\tstart = find_reference_location_v2(snapshot, prefix, 0,\n+\t\t\t\t\t\t\t   &v2_row);\n+\t} else {\n+\t\tif (snapshot->version == 1)\n+\t\t\tstart = snapshot->start;\n+\t\telse\n+\t\t\tstart = snapshot->refs_chunk;\n+\t}\n \n \tif (start == snapshot->eof)\n \t\treturn empty_ref_iterator_begin();\n@@ -439,6 +472,8 @@ static struct ref_iterator *packed_ref_iterator_begin(\n \n \titer->snapshot = snapshot;\n \tacquire_snapshot(snapshot);\n+\titer->version = snapshot->version;\n+\titer->row = v2_row;\n \n \titer->pos = start;\n \titer->eof = snapshot->eof;\ndiff --git a/refs/packed-backend.h b/refs/packed-backend.h\nindex e76f26bfc46..3a8649857f1 100644\n--- a/refs/packed-backend.h\n+++ b/refs/packed-backend.h\n@@ -72,6 +72,9 @@ struct snapshot {\n \t/* Is the `packed-refs` file currently mmapped? */\n \tint mmapped;\n \n+\t/* which file format version is this file? */\n+\tint version;\n+\n \t/*\n \t * The contents of the `packed-refs` file:\n \t *\n@@ -96,6 +99,14 @@ struct snapshot {\n \t */\n \tenum { PEELED_NONE, PEELED_TAGS, PEELED_FULLY } peeled;\n \n+\t/*************************\n+\t * packed-refs v2 values *\n+\t *************************/\n+\tsize_t nr;\n+\tsize_t buflen;\n+\tconst unsigned char *offset_chunk;\n+\tconst char *refs_chunk;\n+\n \t/*\n \t * Count of references to this instance, including the pointer\n \t * from `packed_ref_store::snapshot`, if any. The instance\n@@ -112,6 +123,8 @@ struct snapshot {\n \tstruct stat_validity validity;\n };\n \n+int release_snapshot(struct snapshot *snapshot);\n+\n /*\n  * If the buffer in `snapshot` is active, then either munmap the\n  * memory and close the file, or free the memory. Then set the buffer\n@@ -175,21 +188,30 @@ struct packed_ref_store {\n  */\n struct packed_ref_iterator {\n \tstruct ref_iterator base;\n-\n \tstruct snapshot *snapshot;\n+\tstruct repository *repo;\n+\tunsigned int flags;\n+\tint version;\n+\n+\t/* Scratch space for current values: */\n+\tstruct object_id oid, peeled;\n+\tstruct strbuf refname_buf;\n \n \t/* The current position in the snapshot's buffer: */\n \tconst char *pos;\n \n+\t/***********************************\n+\t * packed-refs v1 iterator values. *\n+\t ***********************************/\n+\n \t/* The end of the part of the buffer that will be iterated over: */\n \tconst char *eof;\n \n-\t/* Scratch space for current values: */\n-\tstruct object_id oid, peeled;\n-\tstruct strbuf refname_buf;\n-\n-\tstruct repository *repo;\n-\tunsigned int flags;\n+\t/***********************************\n+\t * packed-refs v2 iterator values. *\n+\t ***********************************/\n+\tsize_t nr;\n+\tsize_t row;\n };\n \n typedef int (*write_ref_fn)(const char *refname,\n@@ -243,6 +265,42 @@ int write_packed_entry_v1(const char *refname,\n \t\t\t  const struct object_id *peeled,\n \t\t\t  void *write_data);\n \n+/**\n+ * Parse the buffer at the given snapshot to verify that it is a\n+ * packed-refs file in version 1 format. Update the snapshot->peeled\n+ * value according to the header information. Update the given\n+ * 'sorted' value with whether or not the packed-refs file is sorted.\n+ */\n+int parse_packed_format_v1_header(struct packed_ref_store *refs,\n+\t\t\t\t  struct snapshot *snapshot,\n+\t\t\t\t  int *sorted);\n+\n+int detect_packed_format_v2_header(struct packed_ref_store *refs,\n+\t\t\t\t   struct snapshot *snapshot);\n+/*\n+ * Find the place in `snapshot->buf` where the start of the record for\n+ * `refname` starts. If `mustexist` is true and the reference doesn't\n+ * exist, then return NULL. If `mustexist` is false and the reference\n+ * doesn't exist, then return the point where that reference would be\n+ * inserted, or `snapshot->eof` (which might be NULL) if it would be\n+ * inserted at the end of the file. In the latter mode, `refname`\n+ * doesn't have to be a proper reference name; for example, one could\n+ * search for \"refs/replace/\" to find the start of any replace\n+ * references.\n+ *\n+ * The record is sought using a binary search, so `snapshot->buf` must\n+ * be sorted.\n+ */\n+const char *find_reference_location_v2(struct snapshot *snapshot,\n+\t\t\t\t       const char *refname, int mustexist,\n+\t\t\t\t       size_t *pos);\n+\n+int packed_read_raw_ref_v2(struct packed_ref_store *refs, struct snapshot *snapshot,\n+\t\t\t   const char *refname, struct object_id *oid,\n+\t\t\t   unsigned int *type, int *failure_errno);\n+int next_record_v2(struct packed_ref_iterator *iter);\n+void fill_snapshot_v2(struct snapshot *snapshot);\n+\n struct write_packed_refs_v2_context;\n struct write_packed_refs_v2_context *create_v2_context(struct packed_ref_store *refs,\n \t\t\t\t\t\t       struct string_list *updates,\ndiff --git a/refs/packed-format-v2.c b/refs/packed-format-v2.c\nindex 044cc9f629a..d75df9545ec 100644\n--- a/refs/packed-format-v2.c\n+++ b/refs/packed-format-v2.c\n@@ -15,6 +15,215 @@\n #define CHREFS_CHUNKID_OFFSETS         0x524F4646 /* \"ROFF\" */\n #define CHREFS_CHUNKID_REFS            0x52454653 /* \"REFS\" */\n \n+int detect_packed_format_v2_header(struct packed_ref_store *refs,\n+\t\t\t\t   struct snapshot *snapshot)\n+{\n+\t/*\n+\t * packed-refs v1 might not have a header, so check instead\n+\t * that the v2 signature is not present.\n+\t */\n+\treturn get_be32(snapshot->buf) == PACKED_REFS_SIGNATURE;\n+}\n+\n+static const char *get_nth_ref(struct snapshot *snapshot,\n+\t\t\t       size_t n)\n+{\n+\tuint64_t offset;\n+\n+\tif (n >= snapshot->nr)\n+\t\tBUG(\"asking for position %\"PRIu64\" outside of bounds (%\"PRIu64\")\",\n+\t\t    (uint64_t)n, (uint64_t)snapshot->nr);\n+\n+\tif (n)\n+\t\toffset = get_be64(snapshot->offset_chunk + (n-1) * sizeof(uint64_t))\n+\t\t\t\t  & ~OFFSET_IS_PEELED;\n+\telse\n+\t\toffset = 0;\n+\n+\treturn snapshot->refs_chunk + offset;\n+}\n+\n+/*\n+ * Find the place in `snapshot->buf` where the start of the record for\n+ * `refname` starts. If `mustexist` is true and the reference doesn't\n+ * exist, then return NULL. If `mustexist` is false and the reference\n+ * doesn't exist, then return the point where that reference would be\n+ * inserted, or `snapshot->eof` (which might be NULL) if it would be\n+ * inserted at the end of the file. In the latter mode, `refname`\n+ * doesn't have to be a proper reference name; for example, one could\n+ * search for \"refs/replace/\" to find the start of any replace\n+ * references.\n+ *\n+ * The record is sought using a binary search, so `snapshot->buf` must\n+ * be sorted.\n+ */\n+const char *find_reference_location_v2(struct snapshot *snapshot,\n+\t\t\t\t       const char *refname, int mustexist,\n+\t\t\t\t       size_t *pos)\n+{\n+\tsize_t lo = 0, hi = snapshot->nr;\n+\n+\twhile (lo != hi) {\n+\t\tconst char *rec;\n+\t\tint cmp;\n+\t\tsize_t mid = lo + (hi - lo) / 2;\n+\n+\t\trec = get_nth_ref(snapshot, mid);\n+\t\tcmp = strcmp(rec, refname);\n+\t\tif (cmp < 0) {\n+\t\t\tlo = mid + 1;\n+\t\t} else if (cmp > 0) {\n+\t\t\thi = mid;\n+\t\t} else {\n+\t\t\tif (pos)\n+\t\t\t\t*pos = mid;\n+\t\t\treturn rec;\n+\t\t}\n+\t}\n+\n+\tif (mustexist) {\n+\t\treturn NULL;\n+\t} else {\n+\t\tconst char *ret;\n+\t\t/*\n+\t\t * We are likely doing a prefix match, so use the current\n+\t\t * 'lo' position as the indicator.\n+\t\t */\n+\t\tif (pos)\n+\t\t\t*pos = lo;\n+\t\tif (lo >= snapshot->nr)\n+\t\t\treturn NULL;\n+\n+\t\tret = get_nth_ref(snapshot, lo);\n+\t\treturn ret;\n+\t}\n+}\n+\n+int packed_read_raw_ref_v2(struct packed_ref_store *refs, struct snapshot *snapshot,\n+\t\t\t   const char *refname, struct object_id *oid,\n+\t\t\t   unsigned int *type, int *failure_errno)\n+{\n+\tconst char *rec;\n+\n+\t*type = 0;\n+\n+\trec = find_reference_location_v2(snapshot, refname, 1, NULL);\n+\n+\tif (!rec) {\n+\t\t/* refname is not a packed reference. */\n+\t\t*failure_errno = ENOENT;\n+\t\treturn -1;\n+\t}\n+\n+\thashcpy(oid->hash, (const unsigned char *)rec + strlen(rec) + 1);\n+\toid->algo = hash_algo_by_ptr(the_hash_algo);\n+\n+\t*type = REF_ISPACKED;\n+\treturn 0;\n+}\n+\n+static int packed_refs_read_offsets(const unsigned char *chunk_start,\n+\t\t\t\t     size_t chunk_size, void *data)\n+{\n+\tstruct snapshot *snapshot = data;\n+\n+\tsnapshot->offset_chunk = chunk_start;\n+\tsnapshot->nr = chunk_size / sizeof(uint64_t);\n+\treturn 0;\n+}\n+\n+void fill_snapshot_v2(struct snapshot *snapshot)\n+{\n+\tuint32_t file_signature, file_version, hash_version;\n+\tstruct chunkfile *cf;\n+\n+\tfile_signature = get_be32(snapshot->buf);\n+\tif (file_signature != PACKED_REFS_SIGNATURE)\n+\t\tdie(_(\"%s file signature %X does not match signature %X\"),\n+\t\t    \"packed-ref\", file_signature, PACKED_REFS_SIGNATURE);\n+\n+\tfile_version = get_be32(snapshot->buf + sizeof(uint32_t));\n+\tif (file_version != 2)\n+\t\tdie(_(\"format version %u does not match expected file version %u\"),\n+\t\t    file_version, 2);\n+\n+\thash_version = get_be32(snapshot->buf + 2 * sizeof(uint32_t));\n+\tif (hash_version != the_hash_algo->format_id)\n+\t\tdie(_(\"hash version %X does not match expected hash version %X\"),\n+\t\t    hash_version, the_hash_algo->format_id);\n+\n+\tcf = init_chunkfile(NULL);\n+\n+\tif (read_trailing_table_of_contents(cf, (const unsigned char *)snapshot->buf, snapshot->buflen)) {\n+\t\trelease_snapshot(snapshot);\n+\t\tsnapshot = NULL;\n+\t\tgoto cleanup;\n+\t}\n+\n+\tread_chunk(cf, CHREFS_CHUNKID_OFFSETS, packed_refs_read_offsets, snapshot);\n+\tpair_chunk(cf, CHREFS_CHUNKID_REFS, (const unsigned char**)&snapshot->refs_chunk);\n+\n+\t/* TODO: add error checks for invalid chunk combinations. */\n+\n+cleanup:\n+\tfree_chunkfile(cf);\n+}\n+\n+/*\n+ * Move the iterator to the next record in the snapshot, without\n+ * respect for whether the record is actually required by the current\n+ * iteration. Adjust the fields in `iter` and return `ITER_OK` or\n+ * `ITER_DONE`. This function does not free the iterator in the case\n+ * of `ITER_DONE`.\n+ */\n+int next_record_v2(struct packed_ref_iterator *iter)\n+{\n+\tuint64_t offset;\n+\tconst char *pos = iter->pos;\n+\tstrbuf_reset(&iter->refname_buf);\n+\n+\tif (iter->row == iter->snapshot->nr)\n+\t\treturn ITER_DONE;\n+\n+\titer->base.flags = REF_ISPACKED;\n+\n+\tstrbuf_addstr(&iter->refname_buf, pos);\n+\titer->base.refname = iter->refname_buf.buf;\n+\tpos += strlen(pos) + 1;\n+\n+\thashcpy(iter->oid.hash, (const unsigned char *)pos);\n+\titer->oid.algo = hash_algo_by_ptr(the_hash_algo);\n+\tpos += the_hash_algo->rawsz;\n+\n+\tif (check_refname_format(iter->base.refname, REFNAME_ALLOW_ONELEVEL)) {\n+\t\tif (!refname_is_safe(iter->base.refname))\n+\t\t\tdie(\"packed refname is dangerous: %s\",\n+\t\t\t    iter->base.refname);\n+\t\toidclr(&iter->oid);\n+\t\titer->base.flags |= REF_BAD_NAME | REF_ISBROKEN;\n+\t}\n+\n+\t/* We always know the peeled value! */\n+\titer->base.flags |= REF_KNOWS_PEELED;\n+\n+\toffset = get_be64(iter->snapshot->offset_chunk + sizeof(uint64_t) * iter->row);\n+\tif (offset & OFFSET_IS_PEELED) {\n+\t\thashcpy(iter->peeled.hash, (const unsigned char *)pos);\n+\t\titer->peeled.algo = hash_algo_by_ptr(the_hash_algo);\n+\t} else {\n+\t\toidclr(&iter->peeled);\n+\t}\n+\n+\t/* TODO: somehow all tags are getting OFFSET_IS_PEELED even though\n+\t * some are not annotated tags.\n+\t */\n+\titer->pos = iter->snapshot->refs_chunk + (offset & (~OFFSET_IS_PEELED));\n+\n+\titer->row++;\n+\n+\treturn ITER_OK;\n+}\n+\n struct write_packed_refs_v2_context {\n \tstruct packed_ref_store *refs;\n \tstruct string_list *updates;\ndiff --git a/t/t3212-ref-formats.sh b/t/t3212-ref-formats.sh\nindex 03c713ac4f6..571ba518ef1 100755\n--- a/t/t3212-ref-formats.sh\n+++ b/t/t3212-ref-formats.sh\n@@ -73,9 +73,24 @@ test_expect_success 'extensions.refFormat=files,packed-v2' '\n \t\ttest_must_fail git rev-parse refs/tags/Q &&\n \t\trm -f .git/packed-refs &&\n \n+\t\tgit for-each-ref --format=\"%(refname) %(objectname)\" >expect-all &&\n+\t\tgit for-each-ref --format=\"%(refname) %(objectname)\" \\\n+\t\t\trefs/tags/* >expect-tags &&\n+\n \t\t# Create a v2 packed-refs file\n \t\tgit pack-refs --all &&\n-\t\ttest_path_exists .git/packed-refs\n+\t\ttest_path_exists .git/packed-refs &&\n+\t\tfor t in A B\n+\t\tdo\n+\t\t\ttest_path_is_missing .git/refs/tags/$t &&\n+\t\t\tgit rev-parse refs/tags/$t || return 1\n+\t\tdone &&\n+\n+\t\tgit for-each-ref --format=\"%(refname) %(objectname)\" >actual-all &&\n+\t\ttest_cmp expect-all actual-all &&\n+\t\tgit for-each-ref --format=\"%(refname) %(objectname)\" \\\n+\t\t\trefs/tags/* >actual-tags &&\n+\t\ttest_cmp expect-tags actual-tags\n \t)\n '\n \n-- \ngitgitgadget\n\n"},{"id":"466703","messageId":"9b3bd93e51e5ed4358c76263e96c4b4e218987b7.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 20/30] packed-refs: read optional prefix chunks","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:54Z","receivedAt":"2022-11-07T18:37:43Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n refs/packed-backend.c   |   2 +\n refs/packed-backend.h   |   9 +++\n refs/packed-format-v2.c | 159 ++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 170 insertions(+)\n\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex 549cce1f84a..ae904de9014 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -475,6 +475,8 @@ static struct ref_iterator *packed_ref_iterator_begin(\n \titer->version = snapshot->version;\n \titer->row = v2_row;\n \n+\tinit_iterator_prefix_info(prefix, iter);\n+\n \titer->pos = start;\n \titer->eof = snapshot->eof;\n \tstrbuf_init(&iter->refname_buf, 0);\ndiff --git a/refs/packed-backend.h b/refs/packed-backend.h\nindex 3a8649857f1..1936bb5c76c 100644\n--- a/refs/packed-backend.h\n+++ b/refs/packed-backend.h\n@@ -103,9 +103,12 @@ struct snapshot {\n \t * packed-refs v2 values *\n \t *************************/\n \tsize_t nr;\n+\tsize_t prefixes_nr;\n \tsize_t buflen;\n \tconst unsigned char *offset_chunk;\n \tconst char *refs_chunk;\n+\tconst unsigned char *prefix_offsets_chunk;\n+\tconst char *prefix_chunk;\n \n \t/*\n \t * Count of references to this instance, including the pointer\n@@ -212,6 +215,9 @@ struct packed_ref_iterator {\n \t ***********************************/\n \tsize_t nr;\n \tsize_t row;\n+\tsize_t prefix_row_end;\n+\tsize_t prefix_i;\n+\tconst char *cur_prefix;\n };\n \n typedef int (*write_ref_fn)(const char *refname,\n@@ -308,4 +314,7 @@ struct write_packed_refs_v2_context *create_v2_context(struct packed_ref_store *\n int write_packed_refs_v2(struct write_packed_refs_v2_context *ctx);\n void free_v2_context(struct write_packed_refs_v2_context *ctx);\n \n+void init_iterator_prefix_info(const char *prefix,\n+\t\t\t       struct packed_ref_iterator *iter);\n+\n #endif /* REFS_PACKED_BACKEND_H */\ndiff --git a/refs/packed-format-v2.c b/refs/packed-format-v2.c\nindex d75df9545ec..0ab277f7ad4 100644\n--- a/refs/packed-format-v2.c\n+++ b/refs/packed-format-v2.c\n@@ -14,6 +14,79 @@\n #define PACKED_REFS_SIGNATURE          0x50524546 /* \"PREF\" */\n #define CHREFS_CHUNKID_OFFSETS         0x524F4646 /* \"ROFF\" */\n #define CHREFS_CHUNKID_REFS            0x52454653 /* \"REFS\" */\n+#define CHREFS_CHUNKID_PREFIX_DATA     0x50465844 /* \"PFXD\" */\n+#define CHREFS_CHUNKID_PREFIX_OFFSETS  0x5046584F /* \"PFXO\" */\n+\n+static const char *get_nth_prefix(struct snapshot *snapshot,\n+\t\t\t\t  size_t n, size_t *len)\n+{\n+\tuint64_t offset, next_offset;\n+\n+\tif (n >= snapshot->prefixes_nr)\n+\t\tBUG(\"asking for prefix %\"PRIu64\" outside of bounds (%\"PRIu64\")\",\n+\t\t    (uint64_t)n, (uint64_t)snapshot->prefixes_nr);\n+\n+\tif (n)\n+\t\toffset = get_be32(snapshot->prefix_offsets_chunk +\n+\t\t\t\t  2 * sizeof(uint32_t) * (n - 1));\n+\telse\n+\t\toffset = 0;\n+\n+\tif (len) {\n+\t\tnext_offset = get_be32(snapshot->prefix_offsets_chunk +\n+\t\t\t\t       2 * sizeof(uint32_t) * n);\n+\n+\t\t/* Prefix includes null terminator. */\n+\t\t*len = next_offset - offset - 1;\n+\t}\n+\n+\treturn snapshot->prefix_chunk + offset;\n+}\n+\n+/*\n+ * Find the place in `snapshot->buf` where the start of the record for\n+ * `refname` starts. If `mustexist` is true and the reference doesn't\n+ * exist, then return NULL. If `mustexist` is false and the reference\n+ * doesn't exist, then return the point where that reference would be\n+ * inserted, or `snapshot->eof` (which might be NULL) if it would be\n+ * inserted at the end of the file. In the latter mode, `refname`\n+ * doesn't have to be a proper reference name; for example, one could\n+ * search for \"refs/replace/\" to find the start of any replace\n+ * references.\n+ *\n+ * The record is sought using a binary search, so `snapshot->buf` must\n+ * be sorted.\n+ */\n+static const char *find_prefix_location(struct snapshot *snapshot,\n+\t\t\t\t\tconst char *refname, size_t *pos)\n+{\n+\tsize_t lo = 0, hi = snapshot->prefixes_nr;\n+\n+\twhile (lo != hi) {\n+\t\tconst char *rec;\n+\t\tint cmp;\n+\t\tsize_t len;\n+\t\tsize_t mid = lo + (hi - lo) / 2;\n+\n+\t\trec = get_nth_prefix(snapshot, mid, &len);\n+\t\tcmp = strncmp(rec, refname, len);\n+\t\tif (cmp < 0) {\n+\t\t\tlo = mid + 1;\n+\t\t} else if (cmp > 0) {\n+\t\t\thi = mid;\n+\t\t} else {\n+\t\t\t/* we have a prefix match! */\n+\t\t\t*pos = mid;\n+\t\t\treturn rec;\n+\t\t}\n+\t}\n+\n+\t*pos = lo;\n+\tif (lo < snapshot->prefixes_nr)\n+\t\treturn get_nth_prefix(snapshot, lo, NULL);\n+\telse\n+\t\treturn NULL;\n+}\n \n int detect_packed_format_v2_header(struct packed_ref_store *refs,\n \t\t\t\t   struct snapshot *snapshot)\n@@ -63,6 +136,46 @@ const char *find_reference_location_v2(struct snapshot *snapshot,\n {\n \tsize_t lo = 0, hi = snapshot->nr;\n \n+\tif (snapshot->prefix_chunk) {\n+\t\tsize_t prefix_row;\n+\t\tconst char *prefix;\n+\t\tint found = 1;\n+\n+\t\tprefix = find_prefix_location(snapshot, refname, &prefix_row);\n+\n+\t\tif (!prefix || !starts_with(refname, prefix)) {\n+\t\t\tif (mustexist)\n+\t\t\t\treturn NULL;\n+\t\t\tfound = 0;\n+\t\t}\n+\n+\t\t/* The second 4-byte column of the prefix offsets */\n+\t\tif (prefix_row) {\n+\t\t\t/* if prefix_row == 0, then lo = 0, which is already true. */\n+\t\t\tlo = get_be32(snapshot->prefix_offsets_chunk +\n+\t\t\t\t2 * sizeof(uint32_t) * (prefix_row - 1) + sizeof(uint32_t));\n+\t\t}\n+\n+\t\tif (!found) {\n+\t\t\tconst char *ret;\n+\t\t\t/* Terminate early with this lo position as the insertion point. */\n+\t\t\tif (pos)\n+\t\t\t\t*pos = lo;\n+\n+\t\t\tif (lo >= snapshot->nr)\n+\t\t\t\treturn NULL;\n+\n+\t\t\tret = get_nth_ref(snapshot, lo);\n+\t\t\treturn ret;\n+\t\t}\n+\n+\t\thi = get_be32(snapshot->prefix_offsets_chunk +\n+\t\t\t      2 * sizeof(uint32_t) * prefix_row + sizeof(uint32_t));\n+\n+\t\tif (prefix)\n+\t\t\trefname += strlen(prefix);\n+\t}\n+\n \twhile (lo != hi) {\n \t\tconst char *rec;\n \t\tint cmp;\n@@ -132,6 +245,16 @@ static int packed_refs_read_offsets(const unsigned char *chunk_start,\n \treturn 0;\n }\n \n+static int packed_refs_read_prefix_offsets(const unsigned char *chunk_start,\n+\t\t\t\t\t    size_t chunk_size, void *data)\n+{\n+\tstruct snapshot *snapshot = data;\n+\n+\tsnapshot->prefix_offsets_chunk = chunk_start;\n+\tsnapshot->prefixes_nr = chunk_size / sizeof(uint64_t);\n+\treturn 0;\n+}\n+\n void fill_snapshot_v2(struct snapshot *snapshot)\n {\n \tuint32_t file_signature, file_version, hash_version;\n@@ -163,6 +286,9 @@ void fill_snapshot_v2(struct snapshot *snapshot)\n \tread_chunk(cf, CHREFS_CHUNKID_OFFSETS, packed_refs_read_offsets, snapshot);\n \tpair_chunk(cf, CHREFS_CHUNKID_REFS, (const unsigned char**)&snapshot->refs_chunk);\n \n+\tread_chunk(cf, CHREFS_CHUNKID_PREFIX_OFFSETS, packed_refs_read_prefix_offsets, snapshot);\n+\tpair_chunk(cf, CHREFS_CHUNKID_PREFIX_DATA, (const unsigned char**)&snapshot->prefix_chunk);\n+\n \t/* TODO: add error checks for invalid chunk combinations. */\n \n cleanup:\n@@ -187,6 +313,8 @@ int next_record_v2(struct packed_ref_iterator *iter)\n \n \titer->base.flags = REF_ISPACKED;\n \n+\tif (iter->cur_prefix)\n+\t\tstrbuf_addstr(&iter->refname_buf, iter->cur_prefix);\n \tstrbuf_addstr(&iter->refname_buf, pos);\n \titer->base.refname = iter->refname_buf.buf;\n \tpos += strlen(pos) + 1;\n@@ -221,9 +349,40 @@ int next_record_v2(struct packed_ref_iterator *iter)\n \n \titer->row++;\n \n+\tif (iter->row == iter->prefix_row_end && iter->snapshot->prefix_chunk) {\n+\t\tsize_t prefix_pos = get_be32(iter->snapshot->prefix_offsets_chunk +\n+\t\t\t\t\t     2 * sizeof(uint32_t) * iter->prefix_i);\n+\t\titer->cur_prefix = iter->snapshot->prefix_chunk + prefix_pos;\n+\t\titer->prefix_i++;\n+\t\titer->prefix_row_end = get_be32(iter->snapshot->prefix_offsets_chunk +\n+\t\t\t\t\t\t2 * sizeof(uint32_t) * iter->prefix_i + sizeof(uint32_t));\n+\t}\n+\n \treturn ITER_OK;\n }\n \n+void init_iterator_prefix_info(const char *prefix,\n+\t\t\t       struct packed_ref_iterator *iter)\n+{\n+\tstruct snapshot *snapshot = iter->snapshot;\n+\n+\tif (snapshot->version != 2 || !snapshot->prefix_chunk) {\n+\t\titer->prefix_row_end = snapshot->nr;\n+\t\treturn;\n+\t}\n+\n+\tif (prefix)\n+\t\titer->cur_prefix = find_prefix_location(snapshot, prefix, &iter->prefix_i);\n+\telse {\n+\t\titer->cur_prefix = snapshot->prefix_chunk;\n+\t\titer->prefix_i = 0;\n+\t}\n+\n+\titer->prefix_row_end = get_be32(snapshot->prefix_offsets_chunk +\n+\t\t\t\t\t2 * sizeof(uint32_t) * iter->prefix_i +\n+\t\t\t\t\tsizeof(uint32_t));\n+}\n+\n struct write_packed_refs_v2_context {\n \tstruct packed_ref_store *refs;\n \tstruct string_list *updates;\n-- \ngitgitgadget\n\n"},{"id":"466704","messageId":"36f9aa02ebfb967799036c4a0a648ab332c2612b.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 21/30] packed-refs: write prefix chunks","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:55Z","receivedAt":"2022-11-07T18:37:44Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nTests already cover that we will start reading these prefixes.\n\nTODO: discuss time and space savings over typical approach.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n refs/packed-format-v2.c | 103 ++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 103 insertions(+)\n\ndiff --git a/refs/packed-format-v2.c b/refs/packed-format-v2.c\nindex 0ab277f7ad4..2cd45a5987a 100644\n--- a/refs/packed-format-v2.c\n+++ b/refs/packed-format-v2.c\n@@ -398,6 +398,18 @@ struct write_packed_refs_v2_context {\n \tuint64_t *offsets;\n \tsize_t nr;\n \tsize_t offsets_alloc;\n+\n+\tint write_prefixes;\n+\tconst char *cur_prefix;\n+\tsize_t cur_prefix_len;\n+\n+\tchar **prefixes;\n+\tuint32_t *prefix_offsets;\n+\tuint32_t *prefix_rows;\n+\tsize_t prefix_nr;\n+\tsize_t prefixes_alloc;\n+\tsize_t prefix_offsets_alloc;\n+\tsize_t prefix_rows_alloc;\n };\n \n struct write_packed_refs_v2_context *create_v2_context(struct packed_ref_store *refs,\n@@ -434,6 +446,56 @@ static int write_packed_entry_v2(const char *refname,\n \n \tALLOC_GROW(ctx->offsets, i + 1, ctx->offsets_alloc);\n \n+\tif (ctx->write_prefixes) {\n+\t\tif (ctx->cur_prefix && starts_with(refname, ctx->cur_prefix)) {\n+\t\t\t/* skip ahead! */\n+\t\t\trefname += ctx->cur_prefix_len;\n+\t\t\treflen -= ctx->cur_prefix_len;\n+\t\t} else {\n+\t\t\tsize_t len;\n+\t\t\tconst char *slash, *slashslash = NULL;\n+\t\t\tif (ctx->prefix_nr) {\n+\t\t\t\t/* close out the old prefix. */\n+\t\t\t\tctx->prefix_rows[ctx->prefix_nr - 1] = ctx->nr;\n+\t\t\t}\n+\n+\t\t\t/* Find the new prefix. */\n+\t\t\tslash = strchr(refname, '/');\n+\t\t\tif (slash)\n+\t\t\t\tslashslash = strchr(slash + 1, '/');\n+\t\t\t/* If there are two slashes, use that. */\n+\t\t\tslash = slashslash ? slashslash : slash;\n+\t\t\t/*\n+\t\t\t * If there is at least one slash, use that,\n+\t\t\t * and include the slash in the string.\n+\t\t\t * Otherwise, use the end of the ref.\n+\t\t\t */\n+\t\t\tslash = slash ? slash + 1 : refname + strlen(refname);\n+\n+\t\t\tlen = slash - refname;\n+\t\t\tALLOC_GROW(ctx->prefixes, ctx->prefix_nr + 1, ctx->prefixes_alloc);\n+\t\t\tALLOC_GROW(ctx->prefix_offsets, ctx->prefix_nr + 1, ctx->prefix_offsets_alloc);\n+\t\t\tALLOC_GROW(ctx->prefix_rows, ctx->prefix_nr + 1, ctx->prefix_rows_alloc);\n+\n+\t\t\tif (ctx->prefix_nr)\n+\t\t\t\tctx->prefix_offsets[ctx->prefix_nr] = ctx->prefix_offsets[ctx->prefix_nr - 1] + len + 1;\n+\t\t\telse\n+\t\t\t\tctx->prefix_offsets[ctx->prefix_nr] = len + 1;\n+\n+\t\t\tctx->prefixes[ctx->prefix_nr] = xstrndup(refname, len);\n+\t\t\tctx->cur_prefix = ctx->prefixes[ctx->prefix_nr];\n+\t\t\tctx->prefix_nr++;\n+\n+\t\t\trefname += len;\n+\t\t\treflen -= len;\n+\t\t\tctx->cur_prefix_len = len;\n+\t\t}\n+\n+\t\t/* Update the last row continually. */\n+\t\tctx->prefix_rows[ctx->prefix_nr - 1] = i + 1;\n+\t}\n+\n+\n \t/* Write entire ref, including null terminator. */\n \thashwrite(ctx->f, refname, reflen);\n \thashwrite(ctx->f, oid->hash, the_hash_algo->rawsz);\n@@ -483,13 +545,54 @@ static int write_refs_chunk_offsets(struct hashfile *f,\n \treturn 0;\n }\n \n+static int write_refs_chunk_prefix_data(struct hashfile *f,\n+\t\t\t\t\tvoid *data)\n+{\n+\tstruct write_packed_refs_v2_context *ctx = data;\n+\tsize_t i;\n+\n+\ttrace2_region_enter(\"refs\", \"prefix-data\", the_repository);\n+\tfor (i = 0; i < ctx->prefix_nr; i++) {\n+\t\tsize_t len = strlen(ctx->prefixes[i]) + 1;\n+\t\thashwrite(f, ctx->prefixes[i], len);\n+\n+\t\t/* TODO: assert the prefix lengths match the stored offsets? */\n+\t}\n+\n+\ttrace2_region_leave(\"refs\", \"prefix-data\", the_repository);\n+\treturn 0;\n+}\n+\n+static int write_refs_chunk_prefix_offsets(struct hashfile *f,\n+\t\t\t\t    void *data)\n+{\n+\tstruct write_packed_refs_v2_context *ctx = data;\n+\tsize_t i;\n+\n+\ttrace2_region_enter(\"refs\", \"prefix-offsets\", the_repository);\n+\tfor (i = 0; i < ctx->prefix_nr; i++) {\n+\t\thashwrite_be32(f, ctx->prefix_offsets[i]);\n+\t\thashwrite_be32(f, ctx->prefix_rows[i]);\n+\t}\n+\n+\ttrace2_region_leave(\"refs\", \"prefix-offsets\", the_repository);\n+\treturn 0;\n+}\n+\n int write_packed_refs_v2(struct write_packed_refs_v2_context *ctx)\n {\n \tunsigned char file_hash[GIT_MAX_RAWSZ];\n \n+\tctx->write_prefixes = git_env_bool(\"GIT_TEST_WRITE_PACKED_REFS_PREFIXES\", 1);\n+\n \tadd_chunk(ctx->cf, CHREFS_CHUNKID_REFS, 0, write_refs_chunk_refs);\n \tadd_chunk(ctx->cf, CHREFS_CHUNKID_OFFSETS, 0, write_refs_chunk_offsets);\n \n+\tif (ctx->write_prefixes) {\n+\t\tadd_chunk(ctx->cf, CHREFS_CHUNKID_PREFIX_DATA, 0, write_refs_chunk_prefix_data);\n+\t\tadd_chunk(ctx->cf, CHREFS_CHUNKID_PREFIX_OFFSETS, 0, write_refs_chunk_prefix_offsets);\n+\t}\n+\n \thashwrite_be32(ctx->f, PACKED_REFS_SIGNATURE);\n \thashwrite_be32(ctx->f, 2);\n \thashwrite_be32(ctx->f, the_hash_algo->format_id);\n-- \ngitgitgadget\n\n"},{"id":"466705","messageId":"6fe60ef2f53e680b047581eefc629144048b2224.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 22/30] packed-backend: create GIT_TEST_PACKED_REFS_VERSION","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:56Z","receivedAt":"2022-11-07T18:37:47Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nWhen set, this will create a default value for the packed-refs file\nversion on writes. When set to \"2\", it will automatically add the\n\"packed-v2\" value to extensions.refFormat.\n\nNot all tests pass with GIT_TEST_PACKED_REFS_VERSION=2 because they care\nspecifically about the content of the packed-refs file. These tests will\nbe updated in following changes.\n\nTo start, though, disable the GIT_TEST_PACKED_REFS_VERSION environment\nvariable in t3212-ref-formats.sh, since that script already tests both\nversions, including upgrade scenarios.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n refs/packed-backend.c  | 3 ++-\n setup.c                | 5 ++++-\n t/t3212-ref-formats.sh | 3 +++\n 3 files changed, 9 insertions(+), 2 deletions(-)\n\ndiff --git a/refs/packed-backend.c b/refs/packed-backend.c\nindex ae904de9014..e84f669c42e 100644\n--- a/refs/packed-backend.c\n+++ b/refs/packed-backend.c\n@@ -807,7 +807,8 @@ static int write_with_updates(struct packed_ref_store *refs,\n \t}\n \tstrbuf_release(&sb);\n \n-\tif (git_config_get_int(\"refs.packedrefsversion\", &version)) {\n+\tif (!(version = git_env_ulong(\"GIT_TEST_PACKED_REFS_VERSION\", 0)) &&\n+\t    git_config_get_int(\"refs.packedrefsversion\", &version)) {\n \t\t/*\n \t\t * Set the default depending on the current extension\n \t\t * list. Default to version 1 if available, but allow a\ndiff --git a/setup.c b/setup.c\nindex 72bfa289ade..a4525732fe9 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -732,8 +732,11 @@ int read_repository_format(struct repository_format *format, const char *path)\n \t\tclear_repository_format(format);\n \n \t/* Set default ref_format if no extensions.refFormat exists. */\n-\tif (!format->ref_format_count)\n+\tif (!format->ref_format_count) {\n \t\tformat->ref_format = REF_FORMAT_FILES | REF_FORMAT_PACKED;\n+\t\tif (git_env_ulong(\"GIT_TEST_PACKED_REFS_VERSION\", 0) == 2)\n+\t\t\tformat->ref_format |= REF_FORMAT_PACKED_V2;\n+\t}\n \n \treturn format->version;\n }\ndiff --git a/t/t3212-ref-formats.sh b/t/t3212-ref-formats.sh\nindex 571ba518ef1..5583f16db41 100755\n--- a/t/t3212-ref-formats.sh\n+++ b/t/t3212-ref-formats.sh\n@@ -2,6 +2,9 @@\n \n test_description='test across ref formats'\n \n+GIT_TEST_PACKED_REFS_VERSION=0\n+export GIT_TEST_PACKED_REFS_VERSION\n+\n . ./test-lib.sh\n \n test_expect_success 'extensions.refFormat requires core.repositoryFormatVersion=1' '\n-- \ngitgitgadget\n\n"},{"id":"466706","messageId":"191ad7fdef6880738c25c307bbc3c1d66b6378b5.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 24/30] t5312: allow packed-refs v2 format","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:58Z","receivedAt":"2022-11-07T18:37:48Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nOne test in t5312 uses 'grep' to detect that a ref is written in the\npacked-refs file instead of a loose object. This does not work when the\npacked-refs file is in v2 format, such as when\nGIT_TEST_PACKED_REFS_VERSION=2.\n\nSince the test already checks that the loose ref is missing, it suffices\nto check that 'git rev-parse' succeeds.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n t/t3210-pack-refs.sh | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/t/t3210-pack-refs.sh b/t/t3210-pack-refs.sh\nindex 577f32dc71f..fe6c97d9087 100755\n--- a/t/t3210-pack-refs.sh\n+++ b/t/t3210-pack-refs.sh\n@@ -159,7 +159,7 @@ test_expect_success 'delete ref while another dangling packed ref' '\n test_expect_success 'pack ref directly below refs/' '\n \tgit update-ref refs/top HEAD &&\n \tgit pack-refs --all --prune &&\n-\tgrep refs/top .git/packed-refs &&\n+\tgit rev-parse refs/top &&\n \ttest_path_is_missing .git/refs/top\n '\n \n-- \ngitgitgadget\n\n"},{"id":"466707","messageId":"188a55ddcb876aab4e5476234da5412ace053b7b.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 23/30] t1409: test with packed-refs v2","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:57Z","receivedAt":"2022-11-07T18:37:50Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nt1409-avoid-packing-refs.sh seeks to test that the packed-refs file is\nnot modified unnecessarily. One way it does this is by creating a\npacked-refs file, then munging its contents and verifying that the\nmunged data remains after other commands.\n\nFor packed-refs v1, it suffices to add a line that is similar to a\ncomment. For packed-refs v2, we cannot even add to the file without\nmessing up the trailing table of contents of its chunked format.\nHowever, we can manipulate the last bytes that are within the trailing\nhash and use 'tail -c 4' to read them.\n\nThis makes t1409 pass with GIT_TEST_PACKED_REFS_VERSION=2.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n t/t1409-avoid-packing-refs.sh | 22 +++++++++++++++++++---\n 1 file changed, 19 insertions(+), 3 deletions(-)\n\ndiff --git a/t/t1409-avoid-packing-refs.sh b/t/t1409-avoid-packing-refs.sh\nindex be12fb63506..dc8d58432c8 100755\n--- a/t/t1409-avoid-packing-refs.sh\n+++ b/t/t1409-avoid-packing-refs.sh\n@@ -8,13 +8,29 @@ test_description='avoid rewriting packed-refs unnecessarily'\n # shouldn't upset readers, and it should be omitted if the file is\n # ever rewritten.\n mark_packed_refs () {\n-\tsed -e \"s/^\\(#.*\\)/\\1 t1409 /\" .git/packed-refs >.git/packed-refs.new &&\n-\tmv .git/packed-refs.new .git/packed-refs\n+\tif test \"$GIT_TEST_PACKED_REFS_VERSION\" = \"2\"\n+\tthen\n+\t\tsize=$(wc -c < .git/packed-refs) &&\n+\t\tpos=$(expr $size - 4) &&\n+\t\tprintf \"FAKE\" | dd of=\".git/packed-refs\" bs=1 seek=\"$pos\" conv=notrunc\n+\telse\n+\t\tsed -e \"s/^\\(#.*\\)/\\1 t1409 /\" .git/packed-refs >.git/packed-refs.new &&\n+\t\tmv .git/packed-refs.new .git/packed-refs\n+\tfi\n }\n \n # Verify that the packed-refs file is still marked.\n check_packed_refs_marked () {\n-\tgrep -q '^#.* t1409 ' .git/packed-refs\n+\tif test \"$GIT_TEST_PACKED_REFS_VERSION\" = \"2\"\n+\tthen\n+\t\tsize=$(wc -c < .git/packed-refs) &&\n+\t\tpos=$(expr $size - 4) &&\n+\t\ttail -c 4 .git/packed-refs >actual &&\n+\t\tprintf \"FAKE\" >expect &&\n+\t\ttest_cmp expect actual\n+\telse\n+\t\tgrep -q '^#.* t1409 ' .git/packed-refs\n+\tfi\n }\n \n test_expect_success 'setup' '\n-- \ngitgitgadget\n\n"},{"id":"466708","messageId":"e6ed6e7607c4ce81876f281152c675fc14f0498c.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 26/30] t3210: require packed-refs v1 for some tests","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:36:00Z","receivedAt":"2022-11-07T18:37:53Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThree tests in t3210-pack-refs.sh corrupt a packed-refs file to test\nthat Git properly discovers and handles those failures. These tests\nassume that the file is in the v1 format, so add the PACKED_REFS_V1\nprereq to skip these tests when GIT_TEST_PACKED_REFS_VERSION=2.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n t/t3210-pack-refs.sh | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/t/t3210-pack-refs.sh b/t/t3210-pack-refs.sh\nindex fe6c97d9087..76251dfe05a 100755\n--- a/t/t3210-pack-refs.sh\n+++ b/t/t3210-pack-refs.sh\n@@ -197,7 +197,7 @@ test_expect_success 'notice d/f conflict with existing ref' '\n \ttest_must_fail git branch foo/bar/baz/lots/of/extra/components\n '\n \n-test_expect_success 'reject packed-refs with unterminated line' '\n+test_expect_success PACKED_REFS_V1 'reject packed-refs with unterminated line' '\n \tcp .git/packed-refs .git/packed-refs.bak &&\n \ttest_when_finished \"mv .git/packed-refs.bak .git/packed-refs\" &&\n \tprintf \"%s\" \"$HEAD refs/zzzzz\" >>.git/packed-refs &&\n@@ -206,7 +206,7 @@ test_expect_success 'reject packed-refs with unterminated line' '\n \ttest_cmp expected_err err\n '\n \n-test_expect_success 'reject packed-refs containing junk' '\n+test_expect_success PACKED_REFS_V1 'reject packed-refs containing junk' '\n \tcp .git/packed-refs .git/packed-refs.bak &&\n \ttest_when_finished \"mv .git/packed-refs.bak .git/packed-refs\" &&\n \tprintf \"%s\\n\" \"bogus content\" >>.git/packed-refs &&\n@@ -215,7 +215,7 @@ test_expect_success 'reject packed-refs containing junk' '\n \ttest_cmp expected_err err\n '\n \n-test_expect_success 'reject packed-refs with a short SHA-1' '\n+test_expect_success PACKED_REFS_V1 'reject packed-refs with a short SHA-1' '\n \tcp .git/packed-refs .git/packed-refs.bak &&\n \ttest_when_finished \"mv .git/packed-refs.bak .git/packed-refs\" &&\n \tprintf \"%.7s %s\\n\" $HEAD refs/zzzzz >>.git/packed-refs &&\n-- \ngitgitgadget\n\n"},{"id":"466709","messageId":"f8b9f355f0fd11857928de90e79eb8fc284fb009.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 25/30] t5502: add PACKED_REFS_V1 prerequisite","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:35:59Z","receivedAt":"2022-11-07T18:37:55Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe last test in t5502-quickfetch.sh exploits the packed-refs v1 file\nformat by appending 1000 lines to the packed-refs file. If the\npacked-refs file is in the v2 format, this corrupts the file as\nunreadable.\n\nInstead of making the test slower, let's ignore it when\nGIT_TEST_PACKED_REFS_VERSION=2. The test is really about 'git fetch',\nnot the packed-refs format. Create a prerequisite in case we want to use\nthis technique again in the future.\n\nAn alternative would be to write those 1000 refs using a different\nmechanism, but let's opt for the simpler case for now.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n t/t5502-quickfetch.sh | 2 +-\n t/test-lib.sh         | 4 ++++\n 2 files changed, 5 insertions(+), 1 deletion(-)\n\ndiff --git a/t/t5502-quickfetch.sh b/t/t5502-quickfetch.sh\nindex b160f8b7fb7..0c4aadebae6 100755\n--- a/t/t5502-quickfetch.sh\n+++ b/t/t5502-quickfetch.sh\n@@ -122,7 +122,7 @@ test_expect_success 'quickfetch should not copy from alternate' '\n \n '\n \n-test_expect_success 'quickfetch should handle ~1000 refs (on Windows)' '\n+test_expect_success PACKED_REFS_V1 'quickfetch should handle ~1000 refs (on Windows)' '\n \n \tgit gc &&\n \thead=$(git rev-parse HEAD) &&\ndiff --git a/t/test-lib.sh b/t/test-lib.sh\nindex 6db377f68b8..a244cd75c06 100644\n--- a/t/test-lib.sh\n+++ b/t/test-lib.sh\n@@ -1954,3 +1954,7 @@ test_lazy_prereq FSMONITOR_DAEMON '\n \tgit version --build-options >output &&\n \tgrep \"feature: fsmonitor--daemon\" output\n '\n+\n+test_lazy_prereq PACKED_REFS_V1 '\n+\ttest \"$GIT_TEST_PACKED_REFS_VERSION\" -ne \"2\"\n+'\n\\ No newline at end of file\n-- \ngitgitgadget\n\n"},{"id":"466710","messageId":"5aa0d4080291dda854fc1ea7655037822b53111a.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 27/30] t*: skip packed-refs v2 over http tests","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:36:01Z","receivedAt":"2022-11-07T18:38:02Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe GIT_TEST_PACKED_REFS_VERSION=2 environment variable helps us test\nthe packed-refs file format in its v2 version. This variable makes the\nGit process act as if the extensions.refFormat config key has\n\"packed-v2\" in its list. This means that if the environment variable is\nremoved, the repository is in a bad state. This is sufficient for most\ntest cases.\n\nHowever, tests that fetch over HTTP appear to lose this environment\nvariable when executed through the HTTP server. Since the repositories\nare created via Git commands in the tests, the packed-refs files end up\nin the v2 format, but the server processes do not understand this and\nstart serving empty payloads since they do not recognize any refs.\n\nThe preferred long-term solution would be to ensure that the GIT_TEST_*\nenvironment variable persists into the HTTP server. However, these tests\nare not exercising any particularly tricky parts of the packed-refs file\nformat. It may not be worth the effort to pass the environment variable\nand instead we can unset the environment variable (with a comment\nexplaining why) in these tests.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n t/t5539-fetch-http-shallow.sh | 7 +++++++\n t/t5541-http-push-smart.sh    | 7 +++++++\n t/t5542-push-http-shallow.sh  | 7 +++++++\n t/t5551-http-fetch-smart.sh   | 7 +++++++\n t/t5558-clone-bundle-uri.sh   | 7 +++++++\n 5 files changed, 35 insertions(+)\n\ndiff --git a/t/t5539-fetch-http-shallow.sh b/t/t5539-fetch-http-shallow.sh\nindex 3ea75d34ca0..5e3b4304367 100755\n--- a/t/t5539-fetch-http-shallow.sh\n+++ b/t/t5539-fetch-http-shallow.sh\n@@ -5,6 +5,13 @@ test_description='fetch/clone from a shallow clone over http'\n GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n+# If GIT_TEST_PACKED_REFS_VERSION=2, then the packed-refs file will\n+# be written in v2 format without extensions.refFormat=packed-v2. This\n+# causes issues for the HTTP server which does not carry over the\n+# environment variable to the server process.\n+GIT_TEST_PACKED_REFS_VERSION=0\n+export GIT_TEST_PACKED_REFS_VERSION\n+\n . ./test-lib.sh\n . \"$TEST_DIRECTORY\"/lib-httpd.sh\n start_httpd\ndiff --git a/t/t5541-http-push-smart.sh b/t/t5541-http-push-smart.sh\nindex fbad2d5ff5e..495437dd3c7 100755\n--- a/t/t5541-http-push-smart.sh\n+++ b/t/t5541-http-push-smart.sh\n@@ -7,6 +7,13 @@ test_description='test smart pushing over http via http-backend'\n GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n+# If GIT_TEST_PACKED_REFS_VERSION=2, then the packed-refs file will\n+# be written in v2 format without extensions.refFormat=packed-v2. This\n+# causes issues for the HTTP server which does not carry over the\n+# environment variable to the server process.\n+GIT_TEST_PACKED_REFS_VERSION=0\n+export GIT_TEST_PACKED_REFS_VERSION\n+\n . ./test-lib.sh\n \n ROOT_PATH=\"$PWD\"\ndiff --git a/t/t5542-push-http-shallow.sh b/t/t5542-push-http-shallow.sh\nindex c2cc83182f9..c47b18b9faa 100755\n--- a/t/t5542-push-http-shallow.sh\n+++ b/t/t5542-push-http-shallow.sh\n@@ -5,6 +5,13 @@ test_description='push from/to a shallow clone over http'\n GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n+# If GIT_TEST_PACKED_REFS_VERSION=2, then the packed-refs file will\n+# be written in v2 format without extensions.refFormat=packed-v2. This\n+# causes issues for the HTTP server which does not carry over the\n+# environment variable to the server process.\n+GIT_TEST_PACKED_REFS_VERSION=0\n+export GIT_TEST_PACKED_REFS_VERSION\n+\n . ./test-lib.sh\n . \"$TEST_DIRECTORY\"/lib-httpd.sh\n start_httpd\ndiff --git a/t/t5551-http-fetch-smart.sh b/t/t5551-http-fetch-smart.sh\nindex 6a38294a476..61f2e90eabe 100755\n--- a/t/t5551-http-fetch-smart.sh\n+++ b/t/t5551-http-fetch-smart.sh\n@@ -4,6 +4,13 @@ test_description='test smart fetching over http via http-backend'\n GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n+# If GIT_TEST_PACKED_REFS_VERSION=2, then the packed-refs file will\n+# be written in v2 format without extensions.refFormat=packed-v2. This\n+# causes issues for the HTTP server which does not carry over the\n+# environment variable to the server process.\n+GIT_TEST_PACKED_REFS_VERSION=0\n+export GIT_TEST_PACKED_REFS_VERSION\n+\n . ./test-lib.sh\n . \"$TEST_DIRECTORY\"/lib-httpd.sh\n start_httpd\ndiff --git a/t/t5558-clone-bundle-uri.sh b/t/t5558-clone-bundle-uri.sh\nindex 9155f31fa2c..3e35322155e 100755\n--- a/t/t5558-clone-bundle-uri.sh\n+++ b/t/t5558-clone-bundle-uri.sh\n@@ -2,6 +2,13 @@\n \n test_description='test fetching bundles with --bundle-uri'\n \n+# If GIT_TEST_PACKED_REFS_VERSION=2, then the packed-refs file will\n+# be written in v2 format without extensions.refFormat=packed-v2. This\n+# causes issues for the HTTP server which does not carry over the\n+# environment variable to the server process.\n+GIT_TEST_PACKED_REFS_VERSION=0\n+export GIT_TEST_PACKED_REFS_VERSION\n+\n . ./test-lib.sh\n \n test_expect_success 'fail to clone from non-existent file' '\n-- \ngitgitgadget\n\n"},{"id":"466711","messageId":"9d261a55403f8c9d207cfb363689ba9964a57c57.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 28/30] ci: run GIT_TEST_PACKED_REFS_VERSION=2 in some builds","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:36:02Z","receivedAt":"2022-11-07T18:38:03Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe linux-TEST-vars CI build helps us check that certain opt-in features\nare still exercised in at least one environment. The new\nGIT_TEST_PACKED_REFS_VERSION environment variable now passes the test\nsuite when set to \"2\", so add this to that list of variables.\n\nThis provides nearly the same coverage of the v2 format as we had in the\nv1 format.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n ci/run-build-and-tests.sh | 1 +\n 1 file changed, 1 insertion(+)\n\ndiff --git a/ci/run-build-and-tests.sh b/ci/run-build-and-tests.sh\nindex 8ebff425967..e93574ca262 100755\n--- a/ci/run-build-and-tests.sh\n+++ b/ci/run-build-and-tests.sh\n@@ -30,6 +30,7 @@ linux-TEST-vars)\n \texport GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=master\n \texport GIT_TEST_WRITE_REV_INDEX=1\n \texport GIT_TEST_CHECKOUT_WORKERS=2\n+\texport GIT_TEST_PACKED_REFS_VERSION=2\n \t;;\n linux-clang)\n \texport GIT_TEST_DEFAULT_HASH=sha1\n-- \ngitgitgadget\n\n"},{"id":"466712","messageId":"a4a69d8ee91d23465f945488ad42fa38818c2651.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 29/30] p1401: create performance test for ref operations","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:36:03Z","receivedAt":"2022-11-07T18:38:08Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nTBD\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n t/perf/p1401-ref-operations.sh | 47 ++++++++++++++++++++++++++++++++++\n 1 file changed, 47 insertions(+)\n create mode 100755 t/perf/p1401-ref-operations.sh\n\ndiff --git a/t/perf/p1401-ref-operations.sh b/t/perf/p1401-ref-operations.sh\nnew file mode 100755\nindex 00000000000..1c372ba0ee8\n--- /dev/null\n+++ b/t/perf/p1401-ref-operations.sh\n@@ -0,0 +1,47 @@\n+#!/bin/sh\n+\n+test_description=\"Tests performance of ref operations\"\n+\n+. ./perf-lib.sh\n+\n+test_perf_large_repo\n+\n+test_perf 'git pack-refs (v1)' '\n+\tgit commit --allow-empty -m \"change one ref\" &&\n+\tgit pack-refs --all\n+'\n+\n+test_perf 'git for-each-ref (v1)' '\n+\tgit for-each-ref --format=\"%(refname)\" >/dev/null\n+'\n+\n+test_perf 'git for-each-ref prefix (v1)' '\n+\tgit for-each-ref --format=\"%(refname)\" refs/tags/ >/dev/null\n+'\n+\n+test_expect_success 'configure packed-refs v2' '\n+\tgit config core.repositoryFormatVersion 1 &&\n+\tgit config --add extensions.refFormat files &&\n+\tgit config --add extensions.refFormat packed &&\n+\tgit config --add extensions.refFormat packed-v2 &&\n+\tgit config refs.packedRefsVersion 2 &&\n+\tgit commit --allow-empty -m \"change one ref\" &&\n+\tgit pack-refs --all &&\n+\ttest_copy_bytes 16 .git/packed-refs | xxd >actual &&\n+\tgrep PREF actual\n+'\n+\n+test_perf 'git pack-refs (v2)' '\n+\tgit commit --allow-empty -m \"change one ref\" &&\n+\tgit pack-refs --all\n+'\n+\n+test_perf 'git for-each-ref (v2)' '\n+\tgit for-each-ref --format=\"%(refname)\" >/dev/null\n+'\n+\n+test_perf 'git for-each-ref prefix (v2)' '\n+\tgit for-each-ref --format=\"%(refname)\" refs/tags/ >/dev/null\n+'\n+\n+test_done\n-- \ngitgitgadget\n\n"},{"id":"466713","messageId":"37fb4e73ca711f642351d10e1db51c330a1544f1.1667846165.git.gitgitgadget@gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"[PATCH 30/30] refs: skip hashing when writing packed-refs v2","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-11-07T18:36:04Z","receivedAt":"2022-11-07T18:38:09Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <derrickstolee@github.com>\n\nThe 'skip_hash' option in 'struct hashfile' indicates that we want to\nuse the hashfile API as a buffered writer, and not use the hash function\nto create a trailing hash. We still write a trailing null hash to\nindicate that we do not have a checksum at the end. This feature is\nenabled for index writes using the 'index.computeHash' config key.\n\nCreate a similar (currently hidden) option for the packed-refs v2 file\nformat: refs.hashPackedRefs. This defaults to false because performance\nis compared to the packed-refs v1 file format which does have a checksum\nanywhere.\n\nThis change results in improvements to p1401 when using a repository\nwith a 42 MB packed-refs file (600,000+ refs).\n\nTest                        HEAD~1            HEAD\n--------------------------------------------------------------------\n1401.1: git pack-refs (v1)  0.38(0.31+0.52)   0.37(0.28+0.52) -2.6%\n1401.5: git pack-refs (v2)  0.39(0.33+0.52)   0.30(0.28+0.46) -23.1%\n\nNote that these tests update a ref and then repack the packed-refs file.\nThe following benchmarks are from a hyperfine experiment that only ran\nthe 'git pack-refs --all' command for the two formats, but also compared\nthe effect when refs.hashPackedRefs=true.\n\nBenchmark 1: v1\n  Time (mean ± σ):     163.5 ms ±  18.1 ms    [User: 117.8 ms, System: 38.1 ms]\n  Range (min … max):   131.3 ms … 190.4 ms    50 runs\n\nBenchmark 2: v2-no-hash\n  Time (mean ± σ):      95.8 ms ±  15.1 ms    [User: 72.5 ms, System: 23.0 ms]\n  Range (min … max):    82.9 ms … 131.2 ms    50 runs\n\nBenchmark 3: v2-hashing\n  Time (mean ± σ):     100.8 ms ±  16.4 ms    [User: 77.2 ms, System: 23.1 ms]\n  Range (min … max):    83.0 ms … 131.1 ms    50 runs\n\nSummary\n  'v2-no-hash' ran\n    1.05 ± 0.24 times faster than 'v2-hashing'\n    1.71 ± 0.33 times faster than 'v1'\n\nIn this case of repeatedly rewriting the same refs seems to demonstrate\na smaller improvement than the p1401 test. However, the overall\nreduction from v1 matches the expected reduction in file size. In my\ntests, the 42 MB packed-refs (v1) file was compacted to 28 MB in the v2\nformat.\n\nSigned-off-by: Derrick Stolee <derrickstolee@github.com>\n---\n refs/packed-format-v2.c        | 7 +++++++\n t/perf/p1401-ref-operations.sh | 5 +++++\n 2 files changed, 12 insertions(+)\n\ndiff --git a/refs/packed-format-v2.c b/refs/packed-format-v2.c\nindex 2cd45a5987a..ada34bf9bf0 100644\n--- a/refs/packed-format-v2.c\n+++ b/refs/packed-format-v2.c\n@@ -417,6 +417,7 @@ struct write_packed_refs_v2_context *create_v2_context(struct packed_ref_store *\n \t\t\t\t\t\t       struct strbuf *err)\n {\n \tstruct write_packed_refs_v2_context *ctx;\n+\tint do_skip_hash;\n \tCALLOC_ARRAY(ctx, 1);\n \n \tctx->refs = refs;\n@@ -430,6 +431,12 @@ struct write_packed_refs_v2_context *create_v2_context(struct packed_ref_store *\n \t}\n \n \tctx->f = hashfd(refs->tempfile->fd, refs->tempfile->filename.buf);\n+\n+\t/* Default to true, so skip_hash if not set. */\n+\tif (git_config_get_maybe_bool(\"refs.hashpackedrefs\", &do_skip_hash) ||\n+\t    do_skip_hash)\n+\t\tctx->f->skip_hash = 1;\n+\n \tctx->cf = init_chunkfile(ctx->f);\n \n \treturn ctx;\ndiff --git a/t/perf/p1401-ref-operations.sh b/t/perf/p1401-ref-operations.sh\nindex 1c372ba0ee8..0b88a2f531a 100755\n--- a/t/perf/p1401-ref-operations.sh\n+++ b/t/perf/p1401-ref-operations.sh\n@@ -36,6 +36,11 @@ test_perf 'git pack-refs (v2)' '\n \tgit pack-refs --all\n '\n \n+test_perf 'git pack-refs (v2;hashing)' '\n+\tgit commit --allow-empty -m \"change one ref\" &&\n+\tgit -c refs.hashPackedRefs=true pack-refs --all\n+'\n+\n test_perf 'git for-each-ref (v2)' '\n \tgit for-each-ref --format=\"%(refname)\" >/dev/null\n '\n-- \ngitgitgadget\n"},{"id":"466961","messageId":"5ca6f871-3b12-a6e6-2996-9f1a79ee0977@github.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-11-09T15:15:37Z","receivedAt":"2022-11-09T15:15:44Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/7/2022 1:35 PM, Derrick Stolee via GitGitGadget wrote:\n> This RFC is quite long, but the length seemed necessary to actually provide\n> and end-to-end implementation that demonstrates the packed-refs v2 format\n> along with test coverage (via the new GIT_TEST_PACKED_REFS_VERSION\n> variable).\n\nApologies. I had intended to CC a long list of people, but I messed up\nthe \"cc:\" lines on my GGG PR. Appropriate people are CC'd. Please add\nanyone who I might have missed.\n\nThanks,\n-Stolee\n"},{"id":"467176","messageId":"CABPp-BEZK2KJHY+=Ta3VUzNjJKY=evPiAtp5UQFTVLMD0qreVQ@mail.gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-11-11T23:28:18Z","receivedAt":"2022-11-11T23:28:38Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Mon, Nov 7, 2022 at 11:01 AM Derrick Stolee via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> Introduction\n> ============\n>\n> I became interested in our packed-ref format based on the asymmetry between\n> ref updates and ref deletions: if we delete a packed ref, then the\n> packed-refs file needs to be rewritten. Compared to writing a loose ref,\n> this is an O(N) cost instead of O(1).\n>\n> In this way, I set out with some goals:\n>\n>  * (Primary) Make packed ref deletions be nearly as fast as loose ref\n>    updates.\n\nPerformance is always nice.  :-)\n\n>  * (Secondary) Allow using a packed ref format for all refs, dropping loose\n>    refs and creating a clear way to snapshot all refs at a given point in\n>    time.\n\nIs this secondary goal the actual goal you have, or just the\nimplementation by which you get the real underlying goal?\n\nTo me, it appears that such a capability would solve both (a) D/F\nconflict problems (i.e. the ability to simultaneously have a\nrefs/heads/feature and refs/heads/feature/shiny ref), and (b) case\nsensitivity issues in refnames (i.e. inability of some users to work\nwith both a refs/heads/feature and a refs/heads/FeAtUrE, due to\nconstraints of their filesystem and the loose storage mechanism).  Are\neither of those the goal you are trying to achieve (I think both would\nbe really nice, more so than the performance goal you have), or is\nthere another?\n\n> I also had one major non-goal to keep things focused:\n>\n>  * (Non-goal) Update the reflog format.\n>\n> After carefully considering several options, it seemed that there are two\n> solutions that can solve this effectively:\n>\n>  1. Wait for reftable to be integrated into Git.\n>  2. Update the packed-refs backend to have a stacked version.\n>\n> The reftable work seems currently dormant. The format is pretty complicated\n> and I have a difficult time seeing a way forward for it to be fully\n> integrated into Git. Personally, I'd prefer a more incremental approach with\n> formats that are built for a basic filesystem. During the process, we can\n> create APIs within Git that can benefit other file formats within Git.\n>\n> Further, there is a simpler model that satisfies my primary goal without the\n> complication required for the secondary goal. Suppose we create a stacked\n> packed-refs file but only have two layers: the first (base) layer is created\n> when git pack-refs collapses the full stack and adds the loose ref updates\n> to the packed-refs file; the second (top) layer contains only ref deletions\n> (allowing null OIDs to indicate a deleted ref). Then, ref deletions would\n> only need to rewrite that top layer, making ref deletions take O(deletions)\n> time instead of O(all refs) time. With a reasonable schedule to squash the\n> packed-refs stack, this would be a dramatic improvement. (A prototype\n> implementation showed that updating a layer of 1,000 deletions takes only\n> twice the time as writing a single loose ref.)\n\nMakes sense.  If a ref is re-introduced after deletion, then do you\nremove it from the deletion layer and then write the single loose ref?\n\n> If we want to satisfy the secondary goal of passing all ref updates through\n> the packed storage, then more complicated layering would be necessary. The\n> point of bringing this up is that we have incremental goals along the way to\n> that final state that give us good stopping points to test the benefits of\n> each step.\n\nI like the incremental plan.  Your primary goal perhaps benefits\nhosting providers the most, while the second appears to me to be an\ninteresting usability improvement (some of my users might argue it's\neven a bugfix) that would affect users with far fewer refs as well.\nSo, lots of benefits and we get some along the way to the final plan.\n\n> Stacking the packed-refs format introduces several interesting strategy\n> points that are complicated to resolve. Before we can do that, we first need\n> to establish a way to modify the ref format of a Git repository. Hence, we\n> need a new extension for the ref formats.\n>\n> To simplify the first update to the ref formats, it seemed better to add a\n> new file format version to the existing packed-refs file format. This format\n> has the exact lock/write/rename mechanics of the current packed-refs format,\n> but uses a file format that structures the information in a more compact\n> way. It uses the chunk-format API, with some tweaks. This format update is\n> useful to the final goal of a stacked packed-refs API, since each layer will\n> have faster reads and writes. The main reason to do this first is that it is\n> much simpler to understand the value-add (smaller files means faster\n> performance).\n>\n>\n> RFC Organization\n> ================\n>\n> This RFC is quite long, but the length seemed necessary to actually provide\n> and end-to-end implementation that demonstrates the packed-refs v2 format\n> along with test coverage (via the new GIT_TEST_PACKED_REFS_VERSION\n> variable).\n>\n> For convenience, I've broken each section of the full RFC into parts, which\n> resembles how I intend to submit the pieces for full review. These parts are\n> available as pull requests in my fork, but here is a breakdown:\n\n\n> Part I: Optionally hash the index\n> =================================\n>\n> [1] https://github.com/derrickstolee/git/pull/23 Packed-refs v2 Part I:\n> Optionally hash the index (Patches 1-2)\n>\n> The chunk-format API uses the hashfile API as a buffered write, but also all\n> existing formats that use the chunk-format API also have a trailing hash as\n> part of the format. Since the packed-refs file has a critical path involving\n> its write speed (deleting a packed ref), it seemed important to allow\n> apples-to-apples comparison between the v1 and v2 format by skipping the\n> hashing. This is later toggled by a config option.\n>\n> In this part, the focus is on allowing the hashfile API to ignore updating\n> the hash during the buffered writes. We've been using this in microsoft/git\n> to optionally speed up index writes, which patch 2 introduces here. The file\n> format instead writes a null OID which would look like a corrupt file to an\n> older 'git fsck'. Before submitting a full version, I would update 'git\n> fsck' to ignore a null OID in all of our file formats that include a\n> trailing hash. Since the index is more short-lived than other formats (such\n> as pack-files) this trailing hash is less useful. The write time is also\n> critical as the performance tests demonstrate.\n\nThis feels like a diversion from your goals.  Should it really be up-front?\n\nReading through the patches, the first patch does appear to be\nnecessary but isn't well motivated from the cover letter.  The second\npatch seems orthogonal to your series, though it is really nice to see\nindex writes dropping almost to half the time.\n\n> Part II: Create extensions.refFormat\n> ====================================\n>\n> [2] https://github.com/derrickstolee/git/pull/24 Packed-refs v2 Part II:\n> create extensions.refFormat (Patches 3-7)\n>\n> This part is a critical concept that has yet to be defined in the Git\n> codebase. We have no way to incrementally modify the ref format. Since refs\n> are so critical, we cannot add an optionally-understood layer on top (like\n> we did with the multi-pack-index and commit-graph files). The reftable draft\n> [6] proposes the same extension name (extensions.refFormat) but focuses\n> instead on only a single value. This means that the reftable must be defined\n> at git init or git clone time and cannot be upgraded from the files backend.\n>\n> In this RFC, I propose a different model that allows for more customization\n> and incremental updates. The extensions.refFormat config key is multi-valued\n> and defaults to the list of files and packed.\n\nThis last sentence doesn't parse that well for me.  Perhaps \"...and\ndefaults to a combination of 'files' and 'packed', meaning supporting\nboth loose refs and packed refs \"?\n\n> In the context of this RFC,\n> the intention is to be able to add packed-v2 so the list of all three values\n> would allow Git to write and read either file format version (v1 or v2). In\n> the larger scheme, the extension could allow restricting to only loose refs\n> (just files) or only packed-refs (just packed) or even later when reftable\n> is complete, files and reftable could mean that loose refs are the primary\n> ref storage, but the reftable format serves as a drop-in replacement for the\n> packed-refs file. Not all combinations need to be understood by Git, but\n> having them available as an option could be useful for flexibility,\n> especially when trying to upgrade existing repositories to new formats.\n>\n> In the future, beyond the scope of this RFC, it would be good to add a\n> stacked value that allows a stack of files in packed-refs format (whose\n> version is specified by the packed or packed-v2 values) so we can further\n> speed up writes to the packed layer. Depending on how well that works, we\n> could focus on speeding up ref deletions or sending all ref writes straight\n> to the packed-refs layer. With the option to keep the loose refs storage, we\n> have flexibility to explore that space incrementally when we have time to\n> get to it.\n>\n>\n> Part III: Allow a trailing table-of-contents in the chunk-format API\n> ====================================================================\n>\n> [3] https://github.com/derrickstolee/git/pull/25 Packed-refs v2 Part III:\n> trailing table of contents in chunk-format (Patches 8-17)\n>\n> In order to optimize the write speed of the packed-refs v2 file format, we\n> want to write immediately to the file as we stream existing refs from the\n> current refs. The current chunk-format API requires computing the chunk\n> lengths in advance, which can slow down the write and take more memory than\n> necessary. Using a trailing table of contents solves this problem, and was\n> recommended earlier [7]. We just didn't have enough evidence to justify the\n> work to update the existing chunk formats. Here, we update the API in\n> advance of using in the packed-refs v2 format.\n>\n> We could consider updating the commit-graph and multi-pack-index formats to\n> use trailing table of contents, but it requires a version bump. That might\n> be worth it in the case of the commit-graph where computing the size of the\n> changed-path Bloom filters chunk requires a lot of memory at the moment.\n> After this chunk-format API update is reviewed and merged, we can pursue\n> those directions more closely. We would want to investigate the formats more\n> carefully to see if we want to update the chunks themselves as well as some\n> header information.\n\nI like how you point out additional benefits the series could provide,\nbut leave them out.  Perhaps do the same with the optional index\nhashing in patch 2?\n\n> Part IV: Abstract some parts of the v1 file format\n> ==================================================\n>\n> [4] https://github.com/derrickstolee/git/pull/26 Packed-refs v2 Part IV:\n> abstract some parts of the v1 file format (Patches 18-21)\n>\n> These patches move the part of the refs/packed-backend.c file that deal with\n> the specifics of the packed-refs v1 file format into a new file:\n> refs/packed-format-v1.c. This also creates an abstraction layer that will\n> allow inserting the v2 format more easily.\n>\n> One thing that doesn't exist currently is a documentation file describing\n> the packed-refs file format. I would add that file in this part before\n> submitting it for full review. (I also haven't written the file format doc\n> for the packed-refs v2 format, either.)\n\nSounds like another win-win opportunity to get someone from the\ncommunity to contribute.  :-)   ..Or maybe that's not the best\nstrategy, since recent empirical evidence suggests that trick doesn't\nwork.  Oh well, it was worth a shot.\n\n> Part V: Implement the v2 file format\n> ====================================\n>\n> [5] https://github.com/derrickstolee/git/pull/27 Packed-refs v2 Part V: the\n> v2 file format (Patches 22-35)\n>\n> This is the real meat of the work. Perhaps there are ways to split it\n> further, but for now this is what I have ready. The very last patch does a\n> complete performance comparison for a repo with many refs.\n>\n> The format is not yet documented, but is broken up into these pieces:\n>\n>  1. The refs data chunk stores the same data as the packed-refs file, but\n>     each ref is broken down as follows: the ref name (with trailing zero),\n>     the OID for the ref in its raw bytes, and (if necessary) the peeled OID\n>     for the ref in its raw bytes. The refs are sorted lexicographically.\n>\n>  2. The ref offsets chunk is a single column of 64-bit offsets into the refs\n>     chunk indicating where each ref starts. The most-significant bit of that\n>     value indicates whether or not there is a peeled OID.\n>\n>  3. The prefix data chunk lists a set of ref prefixes (currently writes only\n>     allow depth-2 prefixes, such as refs/heads/ and refs/tags/). When\n>     present, these prefixes are written in this chunk and not in the refs\n>     data chunk. The prefixes are sorted lexicographically.\n>\n>  4. The prefix offset chunk has two 32-bit integer columns. The first column\n>     stores the offset within the prefix data chunk to the start of the\n>     prefix string. The second column points to the row position for the\n>     first ref that has name greater than this prefix (the 0th prefix is\n>     assumed to start at row 0, so we can interpret the prefix range from\n>     row[i-1] and row[i]).\n>\n> Between using raw OIDs and storing the depth-2 prefixes only once, this\n> format compresses the file to ~60% of its v1 size. (The format allows not\n> writing the prefix chunks, and the prefix chunks are implemented after the\n> basics of the ref chunks are complete.)\n>\n> The write times are reduced in a similar fraction to the size difference.\n> Reads are sped up somewhat, and we have the potential to do a ref count by\n> prefix much faster by doing a binary search for the start and end of the\n> prefix and then subtracting the row positions instead of scanning the file\n> between to count refs.\n>\n>\n> Relationship to Reftable\n> ========================\n>\n> I mentioned earlier that I had considered using reftable as a way to achieve\n> the stated goals. With the current state of that work, I'm not confident\n> that it is the right approach here.\n>\n> My main worry is that the reftable is more complicated than we need for a\n> typical Git repository that is based on a typical filesystem. This makes\n> testing the format very critical, and we seem to not be near reaching that\n> approach. The v2 format here is very similar to existing Git file formats\n> since it uses the chunk-format API. This means that the amount of code\n> custom to just the v2 format is quite small.\n>\n> As mentioned, the current extension plan [6] only allows reftable or files\n> and does not allow for a mix of both. This RFC introduces the possibility\n> that both could co-exist. Using that multi-valued approach means that I'm\n> able to test the v2 packed-refs file format almost as well as the v1 file\n> format within this RFC. (More tests need to be added that are specific to\n> this format, but I'm waiting for confirmation that this is an acceptable\n> direction.) At the very least, this multi-valued approach could be used as a\n> way to allow using the reftable format as a drop-in replacement for the\n> packed-refs file, as well as upgrading an existing repo to use reftable.\n> That might even help the integration process to allow the reftable format to\n> be tested at least by some subset of tests instead of waiting for a full\n> test suite update.\n\nThanks for providing this background; I also like how it potentially\nmakes it easier to adopt reftable in the future.\n\n> I'm interested to hear from people more involved in the reftable work to see\n> the status of that project and how it matches or differs from my\n> perspective.\n\nThat wouldn't be me, but I appreciate the comparisons to help me\norient where things are.\n"},{"id":"467177","messageId":"CABPp-BEvF3XF+udzTkEtgrtXqYuYEeYi0R65EY5gCespwZgOeg@mail.gmail.com","threadId":"58758","inReplyTo":"030d76f52af654470026b0c4b1dfba2b6c996885.1667846164.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 02/30] read-cache: add index.computeHash config option","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-11-11T23:31:01Z","receivedAt":"2022-11-11T23:31:18Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Mon, Nov 7, 2022 at 10:48 AM Derrick Stolee via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Derrick Stolee <derrickstolee@github.com>\n>\n> The previous change allowed skipping the hashing portion of the\n> hashwrite API, using it instead as a buffered write API. Disabling the\n> hashwrite can be particularly helpful when the write operation is in a\n> critical path.\n>\n> One such critical path is the writing of the index. This operation is so\n> critical that the sparse index was created specifically to reduce the\n> size of the index to make these writes (and reads) faster.\n>\n> Following a similar approach to one used in the microsoft/git fork [1],\n> add a new config option that allows disabling this hashing during the\n> index write. The cost is that we can no longer validate the contents for\n> corruption-at-rest using the trailing hash.\n>\n> [1] https://github.com/microsoft/git/commit/21fed2d91410f45d85279467f21d717a2db45201\n>\n> While older Git versions will not recognize the null hash as a special\n> case, the file format itself is still being met in terms of its\n> structure. Using this null hash will still allow Git operations to\n> function across older versions.\n>\n> The one exception is 'git fsck' which checks the hash of the index file.\n> Here, we disable this check if the trailing hash is all zeroes. We add a\n> warning to the config option that this may cause undesirable behavior\n> with older Git versions.\n>\n> As a quick comparison, I tested 'git update-index --force-write' with\n> and without index.computHash=false on a copy of the Linux kernel\n> repository.\n>\n> Benchmark 1: with hash\n>   Time (mean ± σ):      46.3 ms ±  13.8 ms    [User: 34.3 ms, System: 11.9 ms]\n>   Range (min … max):    34.3 ms …  79.1 ms    82 runs\n>\n> Benchmark 2: without hash\n>   Time (mean ± σ):      26.0 ms ±   7.9 ms    [User: 11.8 ms, System: 14.2 ms]\n>   Range (min … max):    16.3 ms …  42.0 ms    69 runs\n>\n> Summary\n>   'without hash' ran\n>     1.78 ± 0.76 times faster than 'with hash'\n>\n> These performance benefits are substantial enough to allow users the\n> ability to opt-in to this feature, even with the potential confusion\n> with older 'git fsck' versions.\n\nThis is impressive and interesting...but an improvement unrelated to\nthis series other than the fact that it builds on some of it.  Perhaps\npull this patch out?\n\nAlso, would it make sense to integrate index.computeHash with feature.manyFiles?\n\n>\n> Signed-off-by: Derrick Stolee <derrickstolee@github.com>\n> ---\n>  Documentation/config/index.txt |  8 ++++++++\n>  read-cache.c                   | 22 +++++++++++++++++++++-\n>  t/t1600-index.sh               |  8 ++++++++\n>  3 files changed, 37 insertions(+), 1 deletion(-)\n>\n> diff --git a/Documentation/config/index.txt b/Documentation/config/index.txt\n> index 75f3a2d1054..709ba72f622 100644\n> --- a/Documentation/config/index.txt\n> +++ b/Documentation/config/index.txt\n> @@ -30,3 +30,11 @@ index.version::\n>         Specify the version with which new index files should be\n>         initialized.  This does not affect existing repositories.\n>         If `feature.manyFiles` is enabled, then the default is 4.\n> +\n> +index.computeHash::\n> +       When enabled, compute the hash of the index file as it is written\n> +       and store the hash at the end of the content. This is enabled by\n> +       default.\n> ++\n> +If you disable `index.computHash`, then older Git clients may report that\n> +your index is corrupt during `git fsck`.\n> diff --git a/read-cache.c b/read-cache.c\n> index 32024029274..f24d96de4d3 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1817,6 +1817,8 @@ static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n>         git_hash_ctx c;\n>         unsigned char hash[GIT_MAX_RAWSZ];\n>         int hdr_version;\n> +       int all_zeroes = 1;\n> +       unsigned char *start, *end;\n>\n>         if (hdr->hdr_signature != htonl(CACHE_SIGNATURE))\n>                 return error(_(\"bad signature 0x%08x\"), hdr->hdr_signature);\n> @@ -1827,10 +1829,23 @@ static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n>         if (!verify_index_checksum)\n>                 return 0;\n>\n> +       end = (unsigned char *)hdr + size;\n> +       start = end - the_hash_algo->rawsz;\n> +       while (start < end) {\n> +               if (*start != 0) {\n> +                       all_zeroes = 0;\n> +                       break;\n> +               }\n> +               start++;\n> +       }\n> +\n> +       if (all_zeroes)\n> +               return 0;\n> +\n>         the_hash_algo->init_fn(&c);\n>         the_hash_algo->update_fn(&c, hdr, size - the_hash_algo->rawsz);\n>         the_hash_algo->final_fn(hash, &c);\n> -       if (!hasheq(hash, (unsigned char *)hdr + size - the_hash_algo->rawsz))\n> +       if (!hasheq(hash, end - the_hash_algo->rawsz))\n>                 return error(_(\"bad index file sha1 signature\"));\n>         return 0;\n>  }\n> @@ -2917,9 +2932,14 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>         int ieot_entries = 1;\n>         struct index_entry_offset_table *ieot = NULL;\n>         int nr, nr_threads;\n> +       int compute_hash;\n>\n>         f = hashfd(tempfile->fd, tempfile->filename.buf);\n>\n> +       if (!git_config_get_maybe_bool(\"index.computehash\", &compute_hash) &&\n> +           !compute_hash)\n> +               f->skip_hash = 1;\n> +\n>         for (i = removed = extended = 0; i < entries; i++) {\n>                 if (cache[i]->ce_flags & CE_REMOVE)\n>                         removed++;\n> diff --git a/t/t1600-index.sh b/t/t1600-index.sh\n> index 010989f90e6..24ab90ca047 100755\n> --- a/t/t1600-index.sh\n> +++ b/t/t1600-index.sh\n> @@ -103,4 +103,12 @@ test_expect_success 'index version config precedence' '\n>         test_index_version 0 true 2 2\n>  '\n>\n> +test_expect_success 'index.computeHash config option' '\n> +       (\n> +               rm -f .git/index &&\n> +               git -c index.computeHash=false add a &&\n> +               git fsck\n> +       )\n> +'\n> +\n>  test_done\n> --\n> gitgitgadget\n\nPretty simple change, though.  Very nice.  :-)\n"},{"id":"467178","messageId":"CABPp-BGqkSmExRN=bV2014sM5n0msiSOXBG-q-REc8Of4CM4wg@mail.gmail.com","threadId":"58758","inReplyTo":"4013f992d15aab69346bf6f8eafe38511b923595.1667846164.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 03/30] extensions: add refFormat extension","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-11-11T23:39:33Z","receivedAt":"2022-11-11T23:39:51Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Mon, Nov 7, 2022 at 10:48 AM Derrick Stolee via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n[...]\n> One obvious improvement could be a new file format version for the\n> packed-refs file. Its current plaintext-based format is inefficient due\n> to storing object IDs as hexadecimal representations instead of in\n> their raw format. This extra cost will get worse with SHA-256.\n\n> In addition, binary searches need to guess a position and scan to find\n> newlines for a refname entry. A structured binary format could allow for\n> more compact representation and faster access.\n\nThis doesn't parse very well at all.  The scanning is due to refname\nentries being of variable length, and changing hexadecimal\nrepresentation of object IDs to binary values isn't going to help\nthat.\n\nI _think_, after re-scanning your RFC cover letter that you had other\nideas to allow a binary search in order to read a single ref's value,\nand that the juxtaposing of these sentences together leads to an\nunfortunate assumption that one change is related to the both goals,\nbut something extra here to clarify would help.\n\n> diff --git a/Documentation/config/extensions.txt b/Documentation/config/extensions.txt\n> index bccaec7a963..ce8185adf53 100644\n> --- a/Documentation/config/extensions.txt\n> +++ b/Documentation/config/extensions.txt\n> @@ -7,6 +7,47 @@ Note that this setting should only be set by linkgit:git-init[1] or\n>  linkgit:git-clone[1].  Trying to change it after initialization will not\n>  work and will produce hard-to-diagnose issues.\n>\n> +extensions.refFormat::\n> +       Specify the reference storage mechanisms used by the repoitory as a\n> +       multi-valued list. The acceptable values are `files` and `packed`.\n\n> +       If not specified, the list of `files` and `packed` is assumed.\n\nThis sentence doesn't parse for me.\n\n> +       It\n> +       is an error to specify this key unless `core.repositoryFormatVersion`\n> +       is 1.\n\n...is at least 1?  Or are we trying to be incompatible with potential\nfuture core.repositoryFormatVersion values?\n"},{"id":"467229","messageId":"0e156172-0670-2832-78cb-c7dfe2599192@github.com","threadId":"58758","inReplyTo":"CABPp-BEZK2KJHY+=Ta3VUzNjJKY=evPiAtp5UQFTVLMD0qreVQ@mail.gmail.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-11-14T00:07:21Z","receivedAt":"2022-11-14T00:07:28Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/11/22 6:28 PM, Elijah Newren wrote:\n> On Mon, Nov 7, 2022 at 11:01 AM Derrick Stolee via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n>>\n>> Introduction\n>> ============\n>>\n>> I became interested in our packed-ref format based on the asymmetry between\n>> ref updates and ref deletions: if we delete a packed ref, then the\n>> packed-refs file needs to be rewritten. Compared to writing a loose ref,\n>> this is an O(N) cost instead of O(1).\n>>\n>> In this way, I set out with some goals:\n>>\n>>  * (Primary) Make packed ref deletions be nearly as fast as loose ref\n>>    updates.\n> \n> Performance is always nice.  :-)\n> \n>>  * (Secondary) Allow using a packed ref format for all refs, dropping loose\n>>    refs and creating a clear way to snapshot all refs at a given point in\n>>    time.\n> \n> Is this secondary goal the actual goal you have, or just the\n> implementation by which you get the real underlying goal?\n\nTo me, the primary goal takes precedence. It turns out that the best\nway to solve for that goal happens to also make it possible to store\nall refs in a packed form, because we can update the packed form\nmuch faster than our current setup. There are alternatives that I\nconsidered (and prototyped) that were more specific to the deletions\ncase, but they were not actually as fast as the stacked method. Those\nalternatives also would never help reach the secondary goal, but I\nprobably would have considered them anyway if they were faster, if\nonly for their simplicity.\n\n> To me, it appears that such a capability would solve both (a) D/F\n> conflict problems (i.e. the ability to simultaneously have a\n> refs/heads/feature and refs/heads/feature/shiny ref), and (b) case\n> sensitivity issues in refnames (i.e. inability of some users to work\n> with both a refs/heads/feature and a refs/heads/FeAtUrE, due to\n> constraints of their filesystem and the loose storage mechanism).  Are\n> either of those the goal you are trying to achieve (I think both would\n> be really nice, more so than the performance goal you have), or is\n> there another?\n\nFor a Git host provider, these D/F conflict and case-sensitivity\nsituations probably would need to stay as restrictions on the\nserver side for quite some time because we don't want users on\nolder Git clients to be unable to fetch a repository just because\nwe updated our ref storage to allow for such possibilities.\n\nThe biggest benefit on the server side is actually for consistency\nchecks. Using a stacked packed-refs (especially with a tip file\nthat describes all of the layers) allows an atomic way to take a\nsnapshot of the refs and run a checksum operation on their values.\nWith loose refs, concurrent updates can modify the checksum during\nits computation. This is a super niche reason for this, but it's\nnice that the performance-only focus also ends up with a design\nthat satisfies this goal.\n\n...\n\n>> Further, there is a simpler model that satisfies my primary goal without the\n>> complication required for the secondary goal. Suppose we create a stacked\n>> packed-refs file but only have two layers: the first (base) layer is created\n>> when git pack-refs collapses the full stack and adds the loose ref updates\n>> to the packed-refs file; the second (top) layer contains only ref deletions\n>> (allowing null OIDs to indicate a deleted ref). Then, ref deletions would\n>> only need to rewrite that top layer, making ref deletions take O(deletions)\n>> time instead of O(all refs) time. With a reasonable schedule to squash the\n>> packed-refs stack, this would be a dramatic improvement. (A prototype\n>> implementation showed that updating a layer of 1,000 deletions takes only\n>> twice the time as writing a single loose ref.)\n> \n> Makes sense.  If a ref is re-introduced after deletion, then do you\n> remove it from the deletion layer and then write the single loose ref?\n\nLoose refs always take precedence over the packed layer, so if the loose\nref exists we ignore its status in the packed layer. That allows us to not\nupdate the packed layer unless it is a ref deletion or ref maintenance.\n\n>> If we want to satisfy the secondary goal of passing all ref updates through\n>> the packed storage, then more complicated layering would be necessary. The\n>> point of bringing this up is that we have incremental goals along the way to\n>> that final state that give us good stopping points to test the benefits of\n>> each step.\n> \n> I like the incremental plan.  Your primary goal perhaps benefits\n> hosting providers the most, while the second appears to me to be an\n> interesting usability improvement (some of my users might argue it's\n> even a bugfix) that would affect users with far fewer refs as well.\n> So, lots of benefits and we get some along the way to the final plan.\n\nAs I mentioned earlier in this reply, we have a ways to go before we\ncan realize the usability issue, but we can get started somewhere and\nsee how things progress.\n\n>> In this part, the focus is on allowing the hashfile API to ignore updating\n>> the hash during the buffered writes. We've been using this in microsoft/git\n>> to optionally speed up index writes, which patch 2 introduces here. The file\n>> format instead writes a null OID which would look like a corrupt file to an\n>> older 'git fsck'. Before submitting a full version, I would update 'git\n>> fsck' to ignore a null OID in all of our file formats that include a\n>> trailing hash. Since the index is more short-lived than other formats (such\n>> as pack-files) this trailing hash is less useful. The write time is also\n>> critical as the performance tests demonstrate.\n> \n> This feels like a diversion from your goals.  Should it really be up-front?\n> \n> Reading through the patches, the first patch does appear to be\n> necessary but isn't well motivated from the cover letter.  The second\n> patch seems orthogonal to your series, though it is really nice to see\n> index writes dropping almost to half the time.\n\nThis one I put up front only because it is a good candidate for\nsubmitting soon, in parallel to any discussion about the rest of the\nRFC.\n\nThe last part uses the chunk-format API for the packed-refs v2 format,\nbut the write speed is critical and hence we need this ability to skip\nthe hashing. Patch 1 mentions the application in the refs space as a\npotential future, but the immediate benefit to index updates can help\nlots of users right now.\n\nEven if the community said \"this packed-refs v2 is a bad idea\" I would\nstill want to submit these two patches. That's the main reason they are\nup front.\n\n(Also: one thing I didn't mention in this cover letter or in the later\npatches is that we could enable the hashing on the packed-refs v2 via\na config value, eventually. That would allow users who really care\nabout validating their file hashes an option to do so. While this\nwould make packed-refs v2 writes slower than v1 writes, the v1 format\ndoes not have a checksum available, so that might be a valuable option\nto those users.)\n\n>> Part II: Create extensions.refFormat\n>> ====================================\n>>\n>> [2] https://github.com/derrickstolee/git/pull/24 Packed-refs v2 Part II:\n>> create extensions.refFormat (Patches 3-7)\n>>\n>> This part is a critical concept that has yet to be defined in the Git\n>> codebase. We have no way to incrementally modify the ref format. Since refs\n>> are so critical, we cannot add an optionally-understood layer on top (like\n>> we did with the multi-pack-index and commit-graph files). The reftable draft\n>> [6] proposes the same extension name (extensions.refFormat) but focuses\n>> instead on only a single value. This means that the reftable must be defined\n>> at git init or git clone time and cannot be upgraded from the files backend.\n>>\n>> In this RFC, I propose a different model that allows for more customization\n>> and incremental updates. The extensions.refFormat config key is multi-valued\n>> and defaults to the list of files and packed.\n> \n> This last sentence doesn't parse that well for me.  Perhaps \"...and\n> defaults to a combination of 'files' and 'packed', meaning supporting\n> both loose refs and packed refs \"?\n\nSounds good to me. Thanks.\n\n>> Part III: Allow a trailing table-of-contents in the chunk-format API\n>> ====================================================================\n>>\n>> [3] https://github.com/derrickstolee/git/pull/25 Packed-refs v2 Part III:\n>> trailing table of contents in chunk-format (Patches 8-17)\n>>\n>> In order to optimize the write speed of the packed-refs v2 file format, we\n>> want to write immediately to the file as we stream existing refs from the\n>> current refs. The current chunk-format API requires computing the chunk\n>> lengths in advance, which can slow down the write and take more memory than\n>> necessary. Using a trailing table of contents solves this problem, and was\n>> recommended earlier [7]. We just didn't have enough evidence to justify the\n>> work to update the existing chunk formats. Here, we update the API in\n>> advance of using in the packed-refs v2 format.\n>>\n>> We could consider updating the commit-graph and multi-pack-index formats to\n>> use trailing table of contents, but it requires a version bump. That might\n>> be worth it in the case of the commit-graph where computing the size of the\n>> changed-path Bloom filters chunk requires a lot of memory at the moment.\n>> After this chunk-format API update is reviewed and merged, we can pursue\n>> those directions more closely. We would want to investigate the formats more\n>> carefully to see if we want to update the chunks themselves as well as some\n>> header information.\n> \n> I like how you point out additional benefits the series could provide,\n> but leave them out.  Perhaps do the same with the optional index\n> hashing in patch 2?\n\nThe index hashing is just _so easy_ that I couldn't bring myself to\nleave it out. I didn't include the skip_hash option for these other\nformats since the write times are not as critical for these files as\nwrite time is for the index.\n\nUpdating these formats to a v2 that uses a trailing format (and\nlikely other small deviations based on what we've learned since they\nwere first created) would be an interesting direction to pursue with\ncare. Absolutely not something to do while blocking the refs work.\n\n>> Part IV: Abstract some parts of the v1 file format\n>> ==================================================\n>>\n>> [4] https://github.com/derrickstolee/git/pull/26 Packed-refs v2 Part IV:\n>> abstract some parts of the v1 file format (Patches 18-21)\n>>\n>> These patches move the part of the refs/packed-backend.c file that deal with\n>> the specifics of the packed-refs v1 file format into a new file:\n>> refs/packed-format-v1.c. This also creates an abstraction layer that will\n>> allow inserting the v2 format more easily.\n>>\n>> One thing that doesn't exist currently is a documentation file describing\n>> the packed-refs file format. I would add that file in this part before\n>> submitting it for full review. (I also haven't written the file format doc\n>> for the packed-refs v2 format, either.)\n> \n> Sounds like another win-win opportunity to get someone from the\n> community to contribute.  :-)   ..Or maybe that's not the best\n> strategy, since recent empirical evidence suggests that trick doesn't\n> work.  Oh well, it was worth a shot.\n\nHey, if someone beats me to it, I won't complain. They should expect\nme to CC them on reviews for the packed-refs v2 format ;)\n\nThanks,\n-Stolee\n"},{"id":"467254","messageId":"ca39e1d3-b08c-6ed3-8ab5-238efbd1f2dc@github.com","threadId":"58758","inReplyTo":"CABPp-BEvF3XF+udzTkEtgrtXqYuYEeYi0R65EY5gCespwZgOeg@mail.gmail.com","subject":"Re: [PATCH 02/30] read-cache: add index.computeHash config option","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-11-14T16:30:41Z","receivedAt":"2022-11-14T16:35:02Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/11/2022 6:31 PM, Elijah Newren wrote:\n> On Mon, Nov 7, 2022 at 10:48 AM Derrick Stolee via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n>>\n>> From: Derrick Stolee <derrickstolee@github.com>\n>>\n>> The previous change allowed skipping the hashing portion of the\n>> hashwrite API, using it instead as a buffered write API. Disabling the\n>> hashwrite can be particularly helpful when the write operation is in a\n>> critical path.\n>>\n>> One such critical path is the writing of the index. This operation is so\n>> critical that the sparse index was created specifically to reduce the\n>> size of the index to make these writes (and reads) faster.\n>>\n>> Following a similar approach to one used in the microsoft/git fork [1],\n>> add a new config option that allows disabling this hashing during the\n>> index write. The cost is that we can no longer validate the contents for\n>> corruption-at-rest using the trailing hash.\n>>\n>> [1] https://github.com/microsoft/git/commit/21fed2d91410f45d85279467f21d717a2db45201\n>>\n>> While older Git versions will not recognize the null hash as a special\n>> case, the file format itself is still being met in terms of its\n>> structure. Using this null hash will still allow Git operations to\n>> function across older versions.\n>>\n>> The one exception is 'git fsck' which checks the hash of the index file.\n>> Here, we disable this check if the trailing hash is all zeroes. We add a\n>> warning to the config option that this may cause undesirable behavior\n>> with older Git versions.\n>>\n>> As a quick comparison, I tested 'git update-index --force-write' with\n>> and without index.computHash=false on a copy of the Linux kernel\n>> repository.\n>>\n>> Benchmark 1: with hash\n>>   Time (mean ± σ):      46.3 ms ±  13.8 ms    [User: 34.3 ms, System: 11.9 ms]\n>>   Range (min … max):    34.3 ms …  79.1 ms    82 runs\n>>\n>> Benchmark 2: without hash\n>>   Time (mean ± σ):      26.0 ms ±   7.9 ms    [User: 11.8 ms, System: 14.2 ms]\n>>   Range (min … max):    16.3 ms …  42.0 ms    69 runs\n>>\n>> Summary\n>>   'without hash' ran\n>>     1.78 ± 0.76 times faster than 'with hash'\n>>\n>> These performance benefits are substantial enough to allow users the\n>> ability to opt-in to this feature, even with the potential confusion\n>> with older 'git fsck' versions.\n> \n> This is impressive and interesting...but an improvement unrelated to\n> this series other than the fact that it builds on some of it.  Perhaps\n> pull this patch out?\n\nWhile patch 1 is required for the packed-refs work, this one is an easy\nway to take advantage of it. I'll submit these two patches soon on their\nown as the rest of the RFC is discussed.\n\n> Also, would it make sense to integrate index.computeHash with feature.manyFiles?\n\nIt would make sense to include in feature.manyFiles and Scalar's recommended\nconfig. I expect that it would be good to have the config available in a Git\nrelease before updating those configs to include it. Perhaps that is too\nconservative, though.\n\nThanks,\n-Stolee\n"},{"id":"467294","messageId":"CABPp-BFNvUQx7exLgqDvzhgn1s=xSFKbJWdr8qfxLTXEFDQQig@mail.gmail.com","threadId":"58758","inReplyTo":"0e156172-0670-2832-78cb-c7dfe2599192@github.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-11-15T02:47:00Z","receivedAt":"2022-11-15T02:48:06Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sun, Nov 13, 2022 at 4:07 PM Derrick Stolee <derrickstolee@github.com> wrote:\n>\n> On 11/11/22 6:28 PM, Elijah Newren wrote:\n> > On Mon, Nov 7, 2022 at 11:01 AM Derrick Stolee via GitGitGadget\n> > <gitgitgadget@gmail.com> wrote:\n> >>\n> >> Introduction\n> >> ============\n> >>\n> >> I became interested in our packed-ref format based on the asymmetry between\n> >> ref updates and ref deletions: if we delete a packed ref, then the\n> >> packed-refs file needs to be rewritten. Compared to writing a loose ref,\n> >> this is an O(N) cost instead of O(1).\n> >>\n> >> In this way, I set out with some goals:\n> >>\n> >>  * (Primary) Make packed ref deletions be nearly as fast as loose ref\n> >>    updates.\n> >\n> > Performance is always nice.  :-)\n> >\n> >>  * (Secondary) Allow using a packed ref format for all refs, dropping loose\n> >>    refs and creating a clear way to snapshot all refs at a given point in\n> >>    time.\n> >\n> > Is this secondary goal the actual goal you have, or just the\n> > implementation by which you get the real underlying goal?\n>\n> To me, the primary goal takes precedence. It turns out that the best\n> way to solve for that goal happens to also make it possible to store\n> all refs in a packed form, because we can update the packed form\n> much faster than our current setup. There are alternatives that I\n> considered (and prototyped) that were more specific to the deletions\n> case, but they were not actually as fast as the stacked method. Those\n> alternatives also would never help reach the secondary goal, but I\n> probably would have considered them anyway if they were faster, if\n> only for their simplicity.\n\nThat's orthogonal to my question, though.  For your primary goal, you\nstated it in a form where it was obvious what benefit it would provide\nto end users.  Your secondary goal, as stated, didn't list any benefit\nto end users that I could see (update: reading the rest of your\nresponse it appears I just didn't understand it), so I was trying to\nguess at why your secondary goal might be a goal, i.e. what the real\nsecondary goal was.\n\n> > To me, it appears that such a capability would solve both (a) D/F\n> > conflict problems (i.e. the ability to simultaneously have a\n> > refs/heads/feature and refs/heads/feature/shiny ref), and (b) case\n> > sensitivity issues in refnames (i.e. inability of some users to work\n> > with both a refs/heads/feature and a refs/heads/FeAtUrE, due to\n> > constraints of their filesystem and the loose storage mechanism).  Are\n> > either of those the goal you are trying to achieve (I think both would\n> > be really nice, more so than the performance goal you have), or is\n> > there another?\n>\n> For a Git host provider, these D/F conflict and case-sensitivity\n> situations probably would need to stay as restrictions on the\n> server side for quite some time because we don't want users on\n> older Git clients to be unable to fetch a repository just because\n> we updated our ref storage to allow for such possibilities.\n\nOkay, but even if not used on the server side, this capability could\nstill be used on the client side and provide a big benefit to end\nusers.\n\nBut I think there's a minor issue with what you stated; as far as I\ncan tell, there is no case-sensitivity restriction on the server side\nfor GitHub currently, and users do currently have problems cloning and\nusing repositories with branches that differ in case only.  See e.g.\nhttps://github.com/newren/git-filter-repo/issues/48 and the multiple\nduplicates which reference that issue.  We've also had issues at\n$DAYJOB, though for GHE we added some hooks to deny creating branches\nthat differ only in case from another branch to avoid the problem.\n\nAlso, D/F restrictions on the server do not stop users from having D/F\nproblems when fetching.  If users forget to use `--prune`, then when a\nrefs/heads/foo has already been fetched is deleted and replaced by a\nrefs/heads/foo/bar, then the user gets errors.  This issue actually\ncaused a bit of a fire-drill for us just recently.\n\nSo both kinds of problems already exist, for users with any git client\nversion (although the former only for users with unfortunate file\nsystems).  And both problems cause pain.  Both issues are caused by\nloose refs, so limiting git storage to packed refs would fix both\nissues.\n\n> The biggest benefit on the server side is actually for consistency\n> checks. Using a stacked packed-refs (especially with a tip file\n> that describes all of the layers) allows an atomic way to take a\n> snapshot of the refs and run a checksum operation on their values.\n> With loose refs, concurrent updates can modify the checksum during\n> its computation. This is a super niche reason for this, but it's\n> nice that the performance-only focus also ends up with a design\n> that satisfies this goal.\n\nAh...so this is the reason for your secondary goal?  Re-reading it\nlooks like you did state this, I just missed it without the longer\nexplanation.\n\nAnyway, it might be worth calling out in your cover letter that there\nare (at least) three benefits to this secondary goal of yours -- the\none you list here, plus the two I list above.\n"},{"id":"467367","messageId":"d14744cf-5ac3-e1c2-0239-4c33b701d5bf@github.com","threadId":"58758","inReplyTo":"CABPp-BGqkSmExRN=bV2014sM5n0msiSOXBG-q-REc8Of4CM4wg@mail.gmail.com","subject":"Re: [PATCH 03/30] extensions: add refFormat extension","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-11-16T14:37:03Z","receivedAt":"2022-11-16T14:37:09Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/11/22 6:39 PM, Elijah Newren wrote:\n> On Mon, Nov 7, 2022 at 10:48 AM Derrick Stolee via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n> [...]\n>> One obvious improvement could be a new file format version for the\n>> packed-refs file. Its current plaintext-based format is inefficient due\n>> to storing object IDs as hexadecimal representations instead of in\n>> their raw format. This extra cost will get worse with SHA-256.\n> \n>> In addition, binary searches need to guess a position and scan to find\n>> newlines for a refname entry. A structured binary format could allow for\n>> more compact representation and faster access.\n> \n> This doesn't parse very well at all.  The scanning is due to refname\n> entries being of variable length, and changing hexadecimal\n> representation of object IDs to binary values isn't going to help\n> that.\n> \n> I _think_, after re-scanning your RFC cover letter that you had other\n> ideas to allow a binary search in order to read a single ref's value,\n> and that the juxtaposing of these sentences together leads to an\n> unfortunate assumption that one change is related to the both goals,\n> but something extra here to clarify would help.\n\nThe v2 format has a structured list of offsets that can be used to\nnavigate directly to the ith ref in the file. Thus, we can use a\nmore precise form of binary search. Since we have these values, we\ndo not need to scan for newlines or spaces for the end of the ref\nstrings. This allows us to use the raw OIDs since we are not using\nspecial characters as string boundaries.\n\nI will work to clarify when I submit this for review.\n\n>> diff --git a/Documentation/config/extensions.txt b/Documentation/config/extensions.txt\n>> index bccaec7a963..ce8185adf53 100644\n>> --- a/Documentation/config/extensions.txt\n>> +++ b/Documentation/config/extensions.txt\n>> @@ -7,6 +7,47 @@ Note that this setting should only be set by linkgit:git-init[1] or\n>>  linkgit:git-clone[1].  Trying to change it after initialization will not\n>>  work and will produce hard-to-diagnose issues.\n>>\n>> +extensions.refFormat::\n>> +       Specify the reference storage mechanisms used by the repoitory as a\n>> +       multi-valued list. The acceptable values are `files` and `packed`.\n> \n>> +       If not specified, the list of `files` and `packed` is assumed.\n> \n> This sentence doesn't parse for me.\n> \n>> +       It\n>> +       is an error to specify this key unless `core.repositoryFormatVersion`\n>> +       is 1.\n> \n> ...is at least 1?  Or are we trying to be incompatible with potential\n> future core.repositoryFormatVersion values?\n\nSpecifying exactly 1 is consistent across our extensions documentation.\nThe intention of the extensions system is that we should never need a\nvalue 2 here. If we do, then we should consider all extensions to be\nredesigned from scratch. Perhaps we'd have different defaults, or older\noptions not possible anymore.\n\nThanks,\n-Stolee\n"},{"id":"467369","messageId":"01063560-8f57-4e40-5707-f8d8ecfe6cca@github.com","threadId":"58758","inReplyTo":"CABPp-BFNvUQx7exLgqDvzhgn1s=xSFKbJWdr8qfxLTXEFDQQig@mail.gmail.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-11-16T14:45:13Z","receivedAt":"2022-11-16T14:45:36Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/14/22 9:47 PM, Elijah Newren wrote:\n> On Sun, Nov 13, 2022 at 4:07 PM Derrick Stolee <derrickstolee@github.com> wrote:\n>>\n>> On 11/11/22 6:28 PM, Elijah Newren wrote:\n>>> On Mon, Nov 7, 2022 at 11:01 AM Derrick Stolee via GitGitGadget\n>>> <gitgitgadget@gmail.com> wrote:\n>>>>\n>>>> Introduction\n>>>> ============\n>>>>\n>>>> I became interested in our packed-ref format based on the asymmetry between\n>>>> ref updates and ref deletions: if we delete a packed ref, then the\n>>>> packed-refs file needs to be rewritten. Compared to writing a loose ref,\n>>>> this is an O(N) cost instead of O(1).\n>>>>\n>>>> In this way, I set out with some goals:\n>>>>\n>>>>  * (Primary) Make packed ref deletions be nearly as fast as loose ref\n>>>>    updates.\n>>>\n>>> Performance is always nice.  :-)\n>>>\n>>>>  * (Secondary) Allow using a packed ref format for all refs, dropping loose\n>>>>    refs and creating a clear way to snapshot all refs at a given point in\n>>>>    time.\n>>>\n>>> Is this secondary goal the actual goal you have, or just the\n>>> implementation by which you get the real underlying goal?\n>>\n>> To me, the primary goal takes precedence. It turns out that the best\n>> way to solve for that goal happens to also make it possible to store\n>> all refs in a packed form, because we can update the packed form\n>> much faster than our current setup. There are alternatives that I\n>> considered (and prototyped) that were more specific to the deletions\n>> case, but they were not actually as fast as the stacked method. Those\n>> alternatives also would never help reach the secondary goal, but I\n>> probably would have considered them anyway if they were faster, if\n>> only for their simplicity.\n> \n> That's orthogonal to my question, though.  For your primary goal, you\n> stated it in a form where it was obvious what benefit it would provide\n> to end users.  Your secondary goal, as stated, didn't list any benefit\n> to end users that I could see (update: reading the rest of your\n> response it appears I just didn't understand it), so I was trying to\n> guess at why your secondary goal might be a goal, i.e. what the real\n> secondary goal was.\n\nThe reason is in the goal \"creating a clear way to snapshot all refs\nat a given point in time\". This is a server-side benefit with no\nvisible benefit to users, immediately.\n\nThe D/F conflicts and case-sensitive parts that could fall from that\nare not included in my goals. Part of that is because we would need a\nnew reflog format to complete that part. Let's take things one step\nat a time and handle reflogs after we have ref update performance\nhandled.\n\n>>> To me, it appears that such a capability would solve both (a) D/F\n>>> conflict problems (i.e. the ability to simultaneously have a\n>>> refs/heads/feature and refs/heads/feature/shiny ref), and (b) case\n>>> sensitivity issues in refnames (i.e. inability of some users to work\n>>> with both a refs/heads/feature and a refs/heads/FeAtUrE, due to\n>>> constraints of their filesystem and the loose storage mechanism).  Are\n>>> either of those the goal you are trying to achieve (I think both would\n>>> be really nice, more so than the performance goal you have), or is\n>>> there another?\n>>\n>> For a Git host provider, these D/F conflict and case-sensitivity\n>> situations probably would need to stay as restrictions on the\n>> server side for quite some time because we don't want users on\n>> older Git clients to be unable to fetch a repository just because\n>> we updated our ref storage to allow for such possibilities.\n> \n> Okay, but even if not used on the server side, this capability could\n> still be used on the client side and provide a big benefit to end\n> users.\n> \n> But I think there's a minor issue with what you stated; as far as I\n> can tell, there is no case-sensitivity restriction on the server side\n> for GitHub currently, and users do currently have problems cloning and\n> using repositories with branches that differ in case only.  See e.g.\n> https://github.com/newren/git-filter-repo/issues/48 and the multiple\n> duplicates which reference that issue.  We've also had issues at\n> $DAYJOB, though for GHE we added some hooks to deny creating branches\n> that differ only in case from another branch to avoid the problem.\n\nYes, you're right here. We could do better in rejecting case-sensitive\nmatches upon request.\n\n> Also, D/F restrictions on the server do not stop users from having D/F\n> problems when fetching.  If users forget to use `--prune`, then when a\n> refs/heads/foo has already been fetched is deleted and replaced by a\n> refs/heads/foo/bar, then the user gets errors.  This issue actually\n> caused a bit of a fire-drill for us just recently.\n\nAnd similar to the case-sensitive situation, I'm not sure if we have\nchecks to avoid D/F conflicts if they happen across the loose/packed\nboundary. We might just be using the filesystem as a constraint. I'll\nneed to dig in more here.\n\n(This is all the more reason why this space is already complicated and\nwill take some time to unwind.)\n\n> So both kinds of problems already exist, for users with any git client\n> version (although the former only for users with unfortunate file\n> systems).  And both problems cause pain.  Both issues are caused by\n> loose refs, so limiting git storage to packed refs would fix both\n> issues.\n> \n>> The biggest benefit on the server side is actually for consistency\n>> checks. Using a stacked packed-refs (especially with a tip file\n>> that describes all of the layers) allows an atomic way to take a\n>> snapshot of the refs and run a checksum operation on their values.\n>> With loose refs, concurrent updates can modify the checksum during\n>> its computation. This is a super niche reason for this, but it's\n>> nice that the performance-only focus also ends up with a design\n>> that satisfies this goal.\n> \n> Ah...so this is the reason for your secondary goal?  Re-reading it\n> looks like you did state this, I just missed it without the longer\n> explanation.\n> \n> Anyway, it might be worth calling out in your cover letter that there\n> are (at least) three benefits to this secondary goal of yours -- the\n> one you list here, plus the two I list above.\n\nI suppose I assumed that the D/F and case conflicts were a \"known\"\nbenefit and a huge motivation of the reftable work. Instead of trying\nto solve all of the ref problems at once, I wanted to focus on the\nsubset that I knew could be solved with a simpler solution, leaving\nthe full solution to later steps. It would help to be explicit about\nhow this direction helps solve this problem while also being clear\nabout how it does not solve it completely.\n\nThanks,\n-Stolee\n"},{"id":"467414","messageId":"CABPp-BFsvZeC34=VKN9ir+KM0tx4rt0eiGuyKzrD=OAi9sABNw@mail.gmail.com","threadId":"58758","inReplyTo":"01063560-8f57-4e40-5707-f8d8ecfe6cca@github.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-11-17T04:28:00Z","receivedAt":"2022-11-17T04:29:02Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Wed, Nov 16, 2022 at 6:45 AM Derrick Stolee <derrickstolee@github.com> wrote:\n>\n> On 11/14/22 9:47 PM, Elijah Newren wrote:\n> > On Sun, Nov 13, 2022 at 4:07 PM Derrick Stolee <derrickstolee@github.com> wrote:\n> >>\n> >> On 11/11/22 6:28 PM, Elijah Newren wrote:\n> >>> On Mon, Nov 7, 2022 at 11:01 AM Derrick Stolee via GitGitGadget\n> >>> <gitgitgadget@gmail.com> wrote:\n[...]\n> >>>>  * (Secondary) Allow using a packed ref format for all refs, dropping loose\n> >>>>    refs and creating a clear way to snapshot all refs at a given point in\n> >>>>    time.\n[...]\n>\n> The reason is in the goal \"creating a clear way to snapshot all refs\n> at a given point in time\". This is a server-side benefit with no\n> visible benefit to users, immediately.\n\nYes, sorry, I just missed it.  I didn't understand it and wrongly\nassumed it was continuing to talk about the implementation details\nrather than the benefit details.  My bad.\n\nThanks for patiently correcting me.\n\n> The D/F conflicts and case-sensitive parts that could fall from that\n> are not included in my goals. Part of that is because we would need a\n> new reflog format to complete that part. Let's take things one step\n> at a time and handle reflogs after we have ref update performance\n> handled.\n\nAh, right, I can see how reflog would affect both of those problems\nnow that you highlight it, but it hadn't occurred to me before.\n\n> >> The biggest benefit on the server side is actually for consistency\n> >> checks. Using a stacked packed-refs (especially with a tip file\n> >> that describes all of the layers) allows an atomic way to take a\n> >> snapshot of the refs and run a checksum operation on their values.\n> >> With loose refs, concurrent updates can modify the checksum during\n> >> its computation. This is a super niche reason for this, but it's\n> >> nice that the performance-only focus also ends up with a design\n> >> that satisfies this goal.\n> >\n> > Ah...so this is the reason for your secondary goal?  Re-reading it\n> > looks like you did state this, I just missed it without the longer\n> > explanation.\n> >\n> > Anyway, it might be worth calling out in your cover letter that there\n> > are (at least) three benefits to this secondary goal of yours -- the\n> > one you list here, plus the two I list above.\n>\n> I suppose I assumed that the D/F and case conflicts were a \"known\"\n> benefit and a huge motivation of the reftable work.\n\nYes, and I thought you had just found a simpler solution to those\nproblems that might not provide all the benefits of reftable (e.g.\nperformance with huge numbers of refs) but did solve those particular\nproblems.  I've only looked at reftable from the surface from a\ndistance, and I was unaware previously that reflog also affected these\ntwo problems (though it seems obvious in hindsight).  And I do\nremember you calling out that you weren't changing the reflog format\nin your cover letter, but I didn't understand the ramifications of\nthat statement at the time.\n\n> Instead of trying\n> to solve all of the ref problems at once, I wanted to focus on the\n> subset that I knew could be solved with a simpler solution, leaving\n> the full solution to later steps. It would help to be explicit about\n> how this direction helps solve this problem while also being clear\n> about how it does not solve it completely.\n\nIt certainly would have helped me.  :-)\n\nThanks for explaining all these details.\n"},{"id":"467450","messageId":"221117.8635ahik7e.gmgdl@evledraar.gmail.com","threadId":"58758","inReplyTo":"030d76f52af654470026b0c4b1dfba2b6c996885.1667846164.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 02/30] read-cache: add index.computeHash config option","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-11-17T16:13:45Z","receivedAt":"2022-11-17T17:07:08Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Nov 07 2022, Derrick Stolee via GitGitGadget wrote:\n\n> Summary\n>   'without hash' ran\n>     1.78 ± 0.76 times faster than 'with hash'\n>\n> These performance benefits are substantial enough to allow users the\n> ability to opt-in to this feature, even with the potential confusion\n> with older 'git fsck' versions.\n\nThe 0.76 part of that is probably just fs caches etc. screwing things\nup. I tried it on a ramdisk with CFLAGS=-O3:\n\t\n\t$ hyperfine -L v false,true './git -c index.computeHash={v} -C /dev/shm/linux update-index --force-write' -w 1 -r 10\n\tBenchmark 1: ./git -c index.computeHash=false -C /dev/shm/linux update-index --force-write\n\t  Time (mean ± σ):      13.3 ms ±   0.3 ms    [User: 7.1 ms, System: 6.1 ms]\n\t  Range (min … max):    12.7 ms …  13.6 ms    10 runs\n\t \n\tBenchmark 2: ./git -c index.computeHash=true -C /dev/shm/linux update-index --force-write\n\t  Time (mean ± σ):      34.8 ms ±   0.4 ms    [User: 28.9 ms, System: 5.8 ms]\n\t  Range (min … max):    34.2 ms …  35.1 ms    10 runs\n\t \n\tSummary\n\t  './git -c index.computeHash=false -C /dev/shm/linux update-index --force-write' ran\n\t    2.62 ± 0.07 times faster than './git -c index.computeHash=true -C /dev/shm/linux update-index --force-write'\n\nI also see that if I compile with OPENSSL_SHA1=Y, then:\n\t\n\t$ hyperfine -L v false,true './git -c index.computeHash={v} -C /dev/shm/linux update-index --force-write' \n\tBenchmark 1: ./git -c index.computeHash=false -C /dev/shm/linux update-index --force-write\n\t  Time (mean ± σ):      14.0 ms ±   1.3 ms    [User: 7.7 ms, System: 6.2 ms]\n\t  Range (min … max):    13.1 ms …  21.7 ms    206 runs\n\t \n\t  Warning: Statistical outliers were detected. Consider re-running this benchmark on a quiet PC without any interferences from other programs. It might help to use the '--warmup' or '--prepare' \n\toptions.\n\t \n\tBenchmark 2: ./git -c index.computeHash=true -C /dev/shm/linux update-index --force-write\n\t  Time (mean ± σ):      21.0 ms ±   1.0 ms    [User: 15.0 ms, System: 6.0 ms]\n\t  Range (min … max):    20.1 ms …  28.4 ms    138 runs\n\t \n\t  Warning: Statistical outliers were detected. Consider re-running this benchmark on a quiet PC without any interferences from other programs. It might help to use the '--warmup' or '--prepare' \n\toptions.\n\t \n\tSummary\n\t  './git -c index.computeHash=false -C /dev/shm/linux update-index --force-write' ran\n\t    1.50 ± 0.15 times faster than './git -c index.computeHash=true -C /dev/shm/linux update-index --force-write'\n\nWhich, FWIW is something worth considering. I.e. when we introduced\nsha1dc we did so with the \"big hammer\" of the existing hashing API,\nwhich is all or nothing, and we pick the hash when we compile git.\n\nBut that left a lot of things slower for no good reason, e.g. when we do\nthis hashing of the trailers. So if we could just compile with two\nimplementations, and give users the choice of \"use the faster hash when\nyou're not communicating with other git repos\" we could make things\nfaster in some cases, without the potential format interop issues.\n\n> From: Derrick Stolee <derrickstolee@github.com>\n> [...]\n> +index.computeHash::\n> +\tWhen enabled, compute the hash of the index file as it is written\n> +\tand store the hash at the end of the content. This is enabled by\n> +\tdefault.\n> ++\n\nIf we have a boolean option it makes sense to make its name reflect the\nopt-in nature. So \"index.skipHash\". Then just say \"If enabled\", and skip\nthe \"this is enabled by default, and then later this code:\n\n> +\tint compute_hash;\n> [...]\n> +\tif (!git_config_get_maybe_bool(\"index.computehash\", &compute_hash) &&\n> +\t    !compute_hash)\n> +\t\tf->skip_hash = 1;\n\nCan just become:\n\n\tgit_config_get_maybe_bool(\"index.skipHash\", &f->skip_hash);\n\nI.e. git_config_get_maybe_bool() leaves the passed-in dest value alone\nif it doesn't have it in the config, and you only use this\n\"compute_hash\" as an inverted version of \"skip_hash\".\n\n> +If you disable `index.computHash`, then older Git clients may report that\n> +your index is corrupt during `git fsck`.\n> diff --git a/read-cache.c b/read-cache.c\n> index 32024029274..f24d96de4d3 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1817,6 +1817,8 @@ static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n>  \tgit_hash_ctx c;\n>  \tunsigned char hash[GIT_MAX_RAWSZ];\n>  \tint hdr_version;\n> +\tint all_zeroes = 1;\n> +\tunsigned char *start, *end;\n>  \n>  \tif (hdr->hdr_signature != htonl(CACHE_SIGNATURE))\n>  \t\treturn error(_(\"bad signature 0x%08x\"), hdr->hdr_signature);\n> @@ -1827,10 +1829,23 @@ static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n>  \tif (!verify_index_checksum)\n>  \t\treturn 0;\n>  \n> +\tend = (unsigned char *)hdr + size;\n> +\tstart = end - the_hash_algo->rawsz;\n> +\twhile (start < end) {\n> +\t\tif (*start != 0) {\n> +\t\t\tall_zeroes = 0;\n> +\t\t\tbreak;\n> +\t\t}\n> +\t\tstart++;\n> +\t}\n\nDidn't you just re-invent oidread()? :)\n\nJust to narrate my way through this. Before we called verify_hdr() we\ndid:\n\n        hdr = (const struct cache_header *)mmap;\n        if (verify_hdr(hdr, mmap_size) < 0)\n\nSo, we mmap()'d the index on disk, and whe \"hdr\" is the struct version\nof this data, we then cast that back to an \"unsigned char *\" here,\nbecause we're interested in just the raw bytes.\n\nThen we \"jump to the end\" here, and start iterating over the rawsz at\nthe end, because we're just reading if we have a null_oid().\n\nThen, right after that verify_hdr() call, the veriy next thing we'll do is:\n\n\toidread(&istate->oid, (const unsigned char *)hdr + mmap_size - the_hash_algo->rawsz);\n\nSo, maybe I'm missing some subtlety still, and some of this is existing\nbaggage in the pre-image (we used to have the sha1 in the struct, a\n*long* time ago).\n\nBut isn't this equivalent?:\n\t\n\tdiff --git a/read-cache.c b/read-cache.c\n\tindex f24d96de4d3..39b5b8419f5 100644\n\t--- a/read-cache.c\n\t+++ b/read-cache.c\n\t@@ -1812,13 +1812,14 @@ int verify_index_checksum;\n\t /* Allow fsck to force verification of the cache entry order. */\n\t int verify_ce_order;\n\t \n\t-static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n\t+static int verify_hdr(const char *const mmap, const size_t size,\n\t+\t\t      const struct cache_header **hdrp, struct object_id *oid)\n\t {\n\t+\tconst struct cache_header *hdr = (const struct cache_header *)mmap;\n\t \tgit_hash_ctx c;\n\t \tunsigned char hash[GIT_MAX_RAWSZ];\n\t \tint hdr_version;\n\t-\tint all_zeroes = 1;\n\t-\tunsigned char *start, *end;\n\t+\tconst unsigned char *end = (unsigned char *)mmap + size;\n\t \n\t \tif (hdr->hdr_signature != htonl(CACHE_SIGNATURE))\n\t \t\treturn error(_(\"bad signature 0x%08x\"), hdr->hdr_signature);\n\t@@ -1826,20 +1827,12 @@ static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n\t \tif (hdr_version < INDEX_FORMAT_LB || INDEX_FORMAT_UB < hdr_version)\n\t \t\treturn error(_(\"bad index version %d\"), hdr_version);\n\t \n\t+\t*hdrp = hdr;\n\t+\toidread(oid, end - the_hash_algo->rawsz);\n\t+\n\t \tif (!verify_index_checksum)\n\t \t\treturn 0;\n\t-\n\t-\tend = (unsigned char *)hdr + size;\n\t-\tstart = end - the_hash_algo->rawsz;\n\t-\twhile (start < end) {\n\t-\t\tif (*start != 0) {\n\t-\t\t\tall_zeroes = 0;\n\t-\t\t\tbreak;\n\t-\t\t}\n\t-\t\tstart++;\n\t-\t}\n\t-\n\t-\tif (all_zeroes)\n\t+\tif (is_null_oid(oid))\n\t \t\treturn 0;\n\t \n\t \tthe_hash_algo->init_fn(&c);\n\t@@ -2358,11 +2351,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n\t \t\t\tmmap_os_err());\n\t \tclose(fd);\n\t \n\t-\thdr = (const struct cache_header *)mmap;\n\t-\tif (verify_hdr(hdr, mmap_size) < 0)\n\t+\tif (verify_hdr(mmap, mmap_size, &hdr, &istate->oid) < 0)\n\t \t\tgoto unmap;\n\t-\n\t-\toidread(&istate->oid, (const unsigned char *)hdr + mmap_size - the_hash_algo->rawsz);\n\t \tistate->version = ntohl(hdr->hdr_version);\n\t \tistate->cache_nr = ntohl(hdr->hdr_entries);\n\t \tistate->cache_alloc = alloc_nr(istate->cache_nr);\n\nI.e. we just make the verify function be in charge of populating our\n\"oid\", which we can do that early, as we'd error out later in the\nfunction if it doesn't match.\n\nWe could avoid the \"hdrp\" there, but if we're doing the cast it's\nprobably good for readability to just do it once.\n\n> +test_expect_success 'index.computeHash config option' '\n> +\t(\n> +\t\trm -f .git/index &&\n> +\t\tgit -c index.computeHash=false add a &&\n> +\t\tgit fsck\n> +\t)\n> +'\n\nYou can skip the subshell here, but for a non-RFC let's leave the test\nin a nice state for the next test someone adds, so maybe:\n\n\ttest_when_finished \"rm -rf repo\" &&\n\tgit clone . repo &&\n\t[...]\n\nLastly, on this again:\n\n> These performance benefits are substantial enough to allow users the\n> ability to opt-in to this feature, even with the potential confusion\n> with older 'git fsck' versions.\n\nIsn't an unstated major caveat here that it's not \"an older verison\",\nbut if you on *your version* set the config to \"true\" your index doesn't\nhave a hash, so it's persisted until you wipe the index?\n\n"},{"id":"467544","messageId":"xmqqiljbkfg9.fsf@gitster.g","threadId":"58758","inReplyTo":"0e156172-0670-2832-78cb-c7dfe2599192@github.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-11-18T23:31:18Z","receivedAt":"2022-11-19T00:06:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <derrickstolee@github.com> writes:\n\n> On 11/11/22 6:28 PM, Elijah Newren wrote:\n>> On Mon, Nov 7, 2022 at 11:01 AM Derrick Stolee via GitGitGadget\n>> <gitgitgadget@gmail.com> wrote:\n>>>\n>>> Introduction\n>>> ============\n>>>\n>>> I became interested in our packed-ref format based on the asymmetry between\n>>> ref updates and ref deletions: if we delete a packed ref, then the\n>>> packed-refs file needs to be rewritten. Compared to writing a loose ref,\n>>> this is an O(N) cost instead of O(1).\n>>>\n>>> In this way, I set out with some goals:\n>>>\n>>>  * (Primary) Make packed ref deletions be nearly as fast as loose ref\n>>>    updates.\n>> \n>> Performance is always nice.  :-)\n>> \n>>>  * (Secondary) Allow using a packed ref format for all refs, dropping loose\n>>>    refs and creating a clear way to snapshot all refs at a given point in\n>>>    time.\n>> \n>> Is this secondary goal the actual goal you have, or just the\n>> implementation by which you get the real underlying goal?\n>\n> To me, the primary goal takes precedence. It turns out that the best\n> way to solve for that goal happens to also make it possible to store\n> all refs in a packed form, because we can update the packed form\n> much faster than our current setup. There are alternatives that I\n> considered (and prototyped) that were more specific to the deletions\n> case, but they were not actually as fast as the stacked method. Those\n> alternatives also would never help reach the secondary goal, but I\n> probably would have considered them anyway if they were faster, if\n> only for their simplicity.\n\nI have been and am still offline and haven't examined this proposal\nin detail, but would it be a better longer-term approach to improve\nreftable backend, instead of piling more effort on loose+packed\nfilesystem based backend?\n"},{"id":"467546","messageId":"CABPp-BGqXbO9SyF_V_fPEOcZ2uQEWzr0V+KrdcHmfWOq3upniQ@mail.gmail.com","threadId":"58758","inReplyTo":"xmqqiljbkfg9.fsf@gitster.g","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-11-19T00:41:25Z","receivedAt":"2022-11-19T01:38:03Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Fri, Nov 18, 2022 at 3:31 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Derrick Stolee <derrickstolee@github.com> writes:\n>\n> > On 11/11/22 6:28 PM, Elijah Newren wrote:\n> >> On Mon, Nov 7, 2022 at 11:01 AM Derrick Stolee via GitGitGadget\n> >> <gitgitgadget@gmail.com> wrote:\n\n> I have been and am still offline and haven't examined this proposal\n> in detail, but would it be a better longer-term approach to improve\n> reftable backend, instead of piling more effort on loose+packed\n> filesystem based backend?\n\nWell, Stolee explicitly brought this up multiple times in his cover\nletter with various arguments about why he thinks this approach is a\nbetter way to move us on the path towards improved ref handling, and\ndoesn't see it as excluding the reftable option but just opening us up\nto more incremental (and incrementally testable) improvements.  This\nquestion came up early and often in the cover letter; he even ends\nwith a \"Relationship to reftable\" section.\n\nBut he is clearly open to feedback about whether others agree or\ndisagree with his thesis.\n\n(I haven't looked much at reftable, so I can't opine on that question,\nbut Stolee's approach did seem eminently easier to review.  I did have\nsome questions about his proposal(s) because I didn't quite understand\nthem, in part due to being unfamiliar with the area.)\n"},{"id":"467556","messageId":"Y3hG5VL24K2p9G+S@nand.local","threadId":"58758","inReplyTo":"CABPp-BGqXbO9SyF_V_fPEOcZ2uQEWzr0V+KrdcHmfWOq3upniQ@mail.gmail.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-11-19T03:00:53Z","receivedAt":"2022-11-19T03:01:59Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Nov 18, 2022 at 04:41:25PM -0800, Elijah Newren wrote:\n> (I haven't looked much at reftable, so I can't opine on that question,\n> but Stolee's approach did seem eminently easier to review.  I did have\n> some questions about his proposal(s) because I didn't quite understand\n> them, in part due to being unfamiliar with the area.)\n\nFor what it's worth, I'm in the same boat as you are.\n\nThat being said, I do find it somewhat sad that we have reftable bits in\nJunio's tree that don't appear to be progressing all that much. So as\nmuch as I would like to see us have fewer reference backends, I'd rather\nsee whatever the \"next-gen\" backend be have good support and momentum,\neven if that means carrying more code.\n\nThanks,\nTaylor\n"},{"id":"468122","messageId":"CAFQ2z_MZd150kQNTcxaDRVvALpZcCUbRj_81pt-VBY8DRaoRNw@mail.gmail.com","threadId":"58758","inReplyTo":"pull.1408.git.1667846164.gitgitgadget@gmail.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2022-11-28T18:56:55Z","receivedAt":"2022-11-28T18:57:11Z","isPatch":true,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Mon, Nov 7, 2022 at 7:36 PM Derrick Stolee via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n>\n> Introduction\n> ============\n>\n> I became interested in our packed-ref format based on the asymmetry between\n> ref updates and ref deletions: if we delete a packed ref, then the\n> packed-refs file needs to be rewritten. Compared to writing a loose ref,\n> this is an O(N) cost instead of O(1).\n>\n> In this way, I set out with some goals:\n>\n>  * (Primary) Make packed ref deletions be nearly as fast as loose ref\n>    updates.\n>  * (Secondary) Allow using a packed ref format for all refs, dropping loose\n>    refs and creating a clear way to snapshot all refs at a given point in\n>    time.\n>\n> I also had one major non-goal to keep things focused:\n>\n>  * (Non-goal) Update the reflog format.\n>\n> After carefully considering several options, it seemed that there are two\n> solutions that can solve this effectively:\n>\n>  1. Wait for reftable to be integrated into Git.\n>  2. Update the packed-refs backend to have a stacked version.\n>\n> The reftable work seems currently dormant. The format is pretty complicated\n> and I have a difficult time seeing a way forward for it to be fully\n> integrated into Git.\n\nThe format is somewhat complicated, and I think it would have been\npossible to design a block-oriented sorted-table approach that is\nsimpler, but the JGit implementation has set it in stone. But, to put\nthis in perspective, the amount of work for getting the format to\nread/write correctly has been completely dwarfed by the effort needed\nto make the refs API in git represent a true abstraction boundary.\nAlso, if you're introducing a new format, one might as well try to\noptimize it a bit.\n\nHere are some of the hard problems that I encountered\n\n* Worktrees and the main repository have a separate view of the ref\nnamespace. This is not explicit in the ref backend API, and there is a\ntechnical limitation that the packed-refs file cannot be in a\nworktree. This means that worktrees will always continue to use\nloose-ref storage if you only extend the packed-refs backend.\n\n* Symrefs are refs too, but for some reason the packed-refs file\ndoesn't support them. Does packed-refs v2 support symrefs too?  If you\nwant to snapshot the state of refs, do you want to snapshot the value\nof HEAD too?\n\n* By not changing reflogs, you are making things simpler. (if a\ntransaction updates the branch that HEAD points to, the reflog for\nHEAD has to be updated too. Because reftable updates the reflog\ntransactionally, this was some extra work)\nThen again, I feel the current way that reflogs work are a bit messy,\nbecause directory/file conflicts force reflogs to be deleted at times\nthat don't make sense from a user-perspective.\n\n* There are a lot of commands that store SHA1s in files under .git/,\nand access them as if they are a ref (for example: rebase-apply/ ,\nCHERRY_PICK_HEAD etc.).\n\n> In this RFC, I propose a different model that allows for more customization\n> and incremental updates. The extensions.refFormat config key is multi-valued\n> and defaults to the list of files and packed. In the context of this RFC,\n> the intention is to be able to add packed-v2 so the list of all three values\n> would allow Git to write and read either file format version (v1 or v2). In\n> the larger scheme, the extension could allow restricting to only loose refs\n> (just files) or only packed-refs (just packed) or even later when reftable\n> is complete, files and reftable could mean that loose refs are the primary\n> ref storage, but the reftable format serves as a drop-in replacement for the\n> packed-refs file. Not all combinations need to be understood by Git, but\n\nI'm not sure how feasible this is. reftable also holds reflog data. A\nsetting {files,reftable} would either not work, or necessitate hairy\nmerging of data to get the reflogs working correctly.\n\n> In order to optimize the write speed of the packed-refs v2 file format, we\n> want to write immediately to the file as we stream existing refs from the\n> current refs. The current chunk-format API requires computing the chunk\n> lengths in advance, which can slow down the write and take more memory than\n\nyes, this sounds sensible. reftable has the secondary indexes trailing the data.\n\n> Between using raw OIDs and storing the depth-2 prefixes only once, this\n> format compresses the file to ~60% of its v1 size. (The format allows not\n> writing the prefix chunks, and the prefix chunks are implemented after the\n> basics of the ref chunks are complete.)\n>\n> The write times are reduced in a similar fraction to the size difference.\n> Reads are sped up somewhat, and we have the potential to do a ref count by\n\nDo you mean 'enumerate refs' ? Why would you want to count refs by prefix?\n\n> I mentioned earlier that I had considered using reftable as a way to achieve\n> the stated goals. With the current state of that work, I'm not confident\n> that it is the right approach here.\n>\n> My main worry is that the reftable is more complicated than we need for a\n> typical Git repository that is based on a typical filesystem. This makes\n> testing the format very critical, and we seem to not be near reaching that\n> approach.\n\nI think the base code of reading and writing the reftable format is\nexercised quite exhaustively tested in unit tests. You say 'seem', but\ndo you have anything concrete to say?\n\n> As mentioned, the current extension plan [6] only allows reftable or files\n> and does not allow for a mix of both. This RFC introduces the possibility\n> that both could co-exist. Using that multi-valued approach means that I'm\n> able to test the v2 packed-refs file format almost as well as the v1 file\n> format within this RFC. (More tests need to be added that are specific to\n> this format, but I'm waiting for confirmation that this is an acceptable\n> direction.) At the very least, this multi-valued approach could be used as a\n> way to allow using the reftable format as a drop-in replacement for the\n> packed-refs file, as well as upgrading an existing repo to use reftable.\n\nThe multi-value approach creates more combinations of code of how\ndifferent pieces of code can interact, so I think it actually makes it\nmore error-prone.\nAlso,\n\n> That might even help the integration process to allow the reftable format to\n> be tested at least by some subset of tests instead of waiting for a full\n> test suite update.\n\nI don't understand this comment. In the current state,\nhttps://github.com/git/git/pull/1215 already passes 922 of the 968\ntest files if you set GIT_TEST_REFTABLE=1.\n\nSee https://github.com/git/git/pull/1215#issuecomment-1329579459 for\ndetails. As you can see, for most test files, it's just a few\nindividual test cases that fail.\n\n> I'm interested to hear from people more involved in the reftable work to see\n> the status of that project and how it matches or differs from my\n> perspective.\n\nOverall, I found that the loose/packed ref code hard to understand and\nfull of arbitrary limitations (dir/file conflicts, deleting reflogs\nwhen branches are deleted, locking across loose/packed refs etc.).\nThe way reftable stacks are setup (with both reflog and ref data\nincluding symrefs in the same file) make it much easier to verify that\nit behaves transactionally.\n\nFor deleting refs quickly, it seems that you only need to support\n$ZEROID in packed-refs and then implement a ref database as a stack of\npacked-ref files? If you're going for minimal effort and minimal\ndisruption wouldn't that be the place to start?\n\nYou're concerned about the reftable file format (and maybe rightly\nso), but if you're changing the file format anyway and you're not\npicking reftable, why not create a block-based, indexed format that\ncan support storing reflog entries at some point in the future too,\nrather than build on (the limitations) of packed-refs? Or is\npacked-refs v2 backward compatible with v1 (could an old git client\nread v2 files? I think not, right?).\n\nThe reftable project has gotten into a slump because my work\nresponsibilities have increased over the last 1.5 year squeezing down\nhow much time I have for 'fun' projects. I chatted with John Cai, who\nwas trying to staff this project out of Gitlab resources. I don't know\nwhere that stands, though.\n\n> The one thing I can say is that if the reftable work had not already begun,\n> then this is RFC is how I would have approached a new ref format.\n>\n> I look forward to your feedback!\n\nHope this helps.\n\n\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\n\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\n\nRegistergericht und -nummer: Hamburg, HRB 86891\n\nSitz der Gesellschaft: Hamburg\n\nGeschäftsführer: Paul Manicle, Liana Sebastian\n"},{"id":"468252","messageId":"f1c45bd5-692e-85db-90c3-c516003f47e5@github.com","threadId":"58758","inReplyTo":"CAFQ2z_MZd150kQNTcxaDRVvALpZcCUbRj_81pt-VBY8DRaoRNw@mail.gmail.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-11-30T15:16:52Z","receivedAt":"2022-11-30T15:17:04Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/28/2022 1:56 PM, Han-Wen Nienhuys wrote:\n\nHan-Wen,\n\nThanks for taking the time to reply. I was specifically hoping for your\nperspective on the ideas here.\n\n> On Mon, Nov 7, 2022 at 7:36 PM Derrick Stolee via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n>> After carefully considering several options, it seemed that there are two\n>> solutions that can solve this effectively:\n>>\n>>  1. Wait for reftable to be integrated into Git.\n>>  2. Update the packed-refs backend to have a stacked version.\n>>\n>> The reftable work seems currently dormant. The format is pretty complicated\n>> and I have a difficult time seeing a way forward for it to be fully\n>> integrated into Git.\n>\n> The format is somewhat complicated, and I think it would have been\n> possible to design a block-oriented sorted-table approach that is\n> simpler, but the JGit implementation has set it in stone.\n\nI agree that if we pursue reftable, that we should use the format as\nagreed upon and implemented in JGit. I do want to say that while I admire\nJGit's dedication to being compatible with repositories created by Git, I\ndon't think the reverse is a goal of the Git project.\n\n> But, to put\n> this in perspective, the amount of work for getting the format to\n> read/write correctly has been completely dwarfed by the effort needed\n> to make the refs API in git represent a true abstraction boundary.\n> Also, if you're introducing a new format, one might as well try to\n> optimize it a bit.\n\nThat's another reason why I was able to make an incremental improvement so\nquickly in this RFC: I worked within the existing API, reducing the\noverall impact of the change. It's easier to evaluate the performance\ndifference of packed-refs v2 versus packed-refs v1 because the change is\nisolated.\n\nThat work to make the Git refs API work with the reftable library is\nfurther ahead than I though (in your draft PR) but it is also completely\nmissing from the current Git tree, so that work still needs to be arranged\ninto a reviewable series before it is available to us. That does seem like\na substantial amount of work, but I might have been overestimating how\nmuch work it will be compared to these changes I am advocating for.\n\n> Here are some of the hard problems that I encountered\n\nThanks for including these.\n\n> * Worktrees and the main repository have a separate view of the ref\n> namespace. This is not explicit in the ref backend API, and there is a\n> technical limitation that the packed-refs file cannot be in a\n> worktree. This means that worktrees will always continue to use\n> loose-ref storage if you only extend the packed-refs backend.\n\nIf I'm understanding it correctly [1], only the special refs (like HEAD or\nREBASE_HEAD) are worktree-specific, and all refs under \"refs/*\" are\nrepository-scoped. I don't actually think of those special refs as \"loose\"\nrefs and thus they should still work under the \"only packed-refs\" value\nfor extensions.refFormat. I should definitely cover this in the\ndocumentation, though. Also, [1] probably needs updating because it calls\nHEAD a pseudo ref even though it explicitly is not [2].\n\n[1] https://git-scm.com/docs/git-worktree#_refs\n[2] https://git-scm.com/docs/gitglossary#Documentation/gitglossary.txt-aiddefpseudorefapseudoref\n\n> * Symrefs are refs too, but for some reason the packed-refs file\n> doesn't support them. Does packed-refs v2 support symrefs too?  If you\n> want to snapshot the state of refs, do you want to snapshot the value\n> of HEAD too?\n\nI forgot that loose refs under .git/refs/ can be symrefs. This definitely\nis a limitation that I should mention. Again, pseudorefs like HEAD are not\nincluded and are stored separately, but symrefs within refs/* are not\navailable in packed-refs (v1 or v2). That should be explicitly called out\nin the extensions.refFormat docs.\n\nI imagine that such symrefs are uncommon, and users can make their own\nevaluation of whether that use is worth keeping loose refs or not. We can\nstill have the {files, packed[-v2]} extension value while having a\nwriting strategy that writes as much as possible into the packed layer.\n\n> * By not changing reflogs, you are making things simpler. (if a\n> transaction updates the branch that HEAD points to, the reflog for\n> HEAD has to be updated too. Because reftable updates the reflog\n> transactionally, this was some extra work)\n> Then again, I feel the current way that reflogs work are a bit messy,\n> because directory/file conflicts force reflogs to be deleted at times\n> that don't make sense from a user-perspective.\n\nI agree that reflogs are messy. I also think that reflogs have different\nneeds than the ref storage, so separating their needs is valuable.\n\n> * There are a lot of commands that store SHA1s in files under .git/,\n> and access them as if they are a ref (for example: rebase-apply/ ,\n> CHERRY_PICK_HEAD etc.).\n\nYes, I think these pseudorefs are stored differently from usual refs, and\nhence the {packed[-v2]} extension value would still work, but I'll confirm\nthis with more testing.\n\n>> In this RFC, I propose a different model that allows for more customization\n>> and incremental updates. The extensions.refFormat config key is multi-valued\n>> and defaults to the list of files and packed. In the context of this RFC,\n>> the intention is to be able to add packed-v2 so the list of all three values\n>> would allow Git to write and read either file format version (v1 or v2). In\n>> the larger scheme, the extension could allow restricting to only loose refs\n>> (just files) or only packed-refs (just packed) or even later when reftable\n>> is complete, files and reftable could mean that loose refs are the primary\n>> ref storage, but the reftable format serves as a drop-in replacement for the\n>> packed-refs file. Not all combinations need to be understood by Git, but\n>\n> I'm not sure how feasible this is. reftable also holds reflog data. A\n> setting {files,reftable} would either not work, or necessitate hairy\n> merging of data to get the reflogs working correctly.\n\nIn this setup, would it be possible to continue using the \"loose reflog\"\nformat while using reftable as the packed layer? I personally think this\ncombination of formats to be critical to upgrading existing repositories\nto reftable.\n\n(Note: there is a strategy that doesn't need this approach, but it's a bit\ncomplicated. It would involve rotating all replicas to new repositories\nthat are configured to use reftable upon creation, getting the refs from\nother replicas via fetches. In my opinion, this is prohibitively\nexpensive.)\n\n>> Between using raw OIDs and storing the depth-2 prefixes only once, this\n>> format compresses the file to ~60% of its v1 size. (The format allows not\n>> writing the prefix chunks, and the prefix chunks are implemented after the\n>> basics of the ref chunks are complete.)\n>>\n>> The write times are reduced in a similar fraction to the size difference.\n>> Reads are sped up somewhat, and we have the potential to do a ref count by\n>\n> Do you mean 'enumerate refs' ? Why would you want to count refs by prefix?\n\nGenerally, I mean these kind of operations:\n\n* 'git for-each-ref' enumerates all refs within a prefix.\n\n* Serving the ref advertisement enumerates all refs.\n\n* There was a GitHub feature that counted refs and tags, but wanted to\n  ignore internal ref prefixes (outside of refs/heads/* or refs/tags/*).\n  It turns out that we didn't actually need the full count but an\n  existence indicator, but it would be helpful to quickly identify how\n  many branches or tags are in a repository at a glance. Packed-refs v1\n  requires scanning the whole file while packed-refs v2 does a fixed\n  number of binary searches followed by a subtraction of row indexes.\n\n>> I mentioned earlier that I had considered using reftable as a way to achieve\n>> the stated goals. With the current state of that work, I'm not confident\n>> that it is the right approach here.\n>>\n>> My main worry is that the reftable is more complicated than we need for a\n>> typical Git repository that is based on a typical filesystem. This makes\n>> testing the format very critical, and we seem to not be near reaching that\n>> approach.\n>\n> I think the base code of reading and writing the reftable format is\n> exercised quite exhaustively tested in unit tests. You say 'seem', but\n> do you have anything concrete to say?\n\nOur test suite is focused on integration tests at the command level. While\nunit tests are helpful, I'm not sure if all of the corner cases would be\ncovered by tests that check Git commands only.\n\n>> As mentioned, the current extension plan [6] only allows reftable or files\n>> and does not allow for a mix of both. This RFC introduces the possibility\n>> that both could co-exist. Using that multi-valued approach means that I'm\n>> able to test the v2 packed-refs file format almost as well as the v1 file\n>> format within this RFC. (More tests need to be added that are specific to\n>> this format, but I'm waiting for confirmation that this is an acceptable\n>> direction.) At the very least, this multi-valued approach could be used as a\n>> way to allow using the reftable format as a drop-in replacement for the\n>> packed-refs file, as well as upgrading an existing repo to use reftable.\n>\n> The multi-value approach creates more combinations of code of how\n> different pieces of code can interact, so I think it actually makes it\n> more error-prone.\n\nAs multiple values are added, it will be important to indicate which\nvalues are not compatible with each other. However, the plan for the\npacked-refs improvements do add values that are orthogonal to each other.\nIt does make testing all combinations more difficult.\n\nOf course, if reftable is truly incompatible with loose refs, then Git can\nsay that {reftable} is the only set of values that can use reftable, and\nmake {files, reftable} an incompatible set (which could be understood by\na later version of Git, if those barriers are overcome). However, if we do\nnot specify the extension as multi-valued from the start, then we cannot\nlater add this multi-valued option without changing the extension name.\n\n>> That might even help the integration process to allow the reftable format to\n>> be tested at least by some subset of tests instead of waiting for a full\n>> test suite update.\n>\n> I don't understand this comment. In the current state,\n> https://github.com/git/git/pull/1215 already passes 922 of the 968\n> test files if you set GIT_TEST_REFTABLE=1.\n>\n> See https://github.com/git/git/pull/1215#issuecomment-1329579459 for\n> details. As you can see, for most test files, it's just a few\n> individual test cases that fail.\n\nMy point is that to get those remaining tests passing requires a\nsignificant update to the test suite. I imagined that the complexity of\nthat update was the blocker to completing the reftable work.\n\nIt seems that my estimation of that complexity was overly high compared to\nwhat you appear to be describing.\n\n>> I'm interested to hear from people more involved in the reftable work to see\n>> the status of that project and how it matches or differs from my\n>> perspective.\n>\n> Overall, I found that the loose/packed ref code hard to understand and\n> full of arbitrary limitations (dir/file conflicts, deleting reflogs\n> when branches are deleted, locking across loose/packed refs etc.).\n> The way reftable stacks are setup (with both reflog and ref data\n> including symrefs in the same file) make it much easier to verify that\n> it behaves transactionally.\n\nI believe you that starting with a new data model makes many of these\nthings easier to reason with.\n\n> For deleting refs quickly, it seems that you only need to support\n> $ZEROID in packed-refs and then implement a ref database as a stack of\n> packed-ref files? If you're going for minimal effort and minimal\n> disruption wouldn't that be the place to start?\n\nI disagree that jumping straight to stacked packed-refs is minimal effort\nor minimal disruption.\n\nCreating the stack approach does require changing the semantics of the\npacked-refs format to include $ZEROID, which will modify some meanings in\nthe iteration code. The use of a stack, as well as how layers are combined\nduring a ref write or also during maintenance, adds complications to the\nlocking semantics that are decently complicated.\n\nBy contrast, the v2 format is isolated to the on-disk format. None of the\nwriting or reading semantics are changed in terms of which files to look\nat or write in which order. Instead, it's relatively simple to see from\nthe format exactly how it reduces the file size but otherwise has exactly\nthe same read/write behavior. In fact, since the refs and OIDs are all\nlocated in the same chunk in a similar order to the v1 file, we can even\ndeduce that page cache semantics will only improve in the new format.\n\nThe reason to start with this step is that the benefits and risks are\nclearly understood, which can motivate us to establish the mechanism for\nchanging the ref format by defining the extension.\n\n> You're concerned about the reftable file format (and maybe rightly\n> so), but if you're changing the file format anyway and you're not\n> picking reftable, why not create a block-based, indexed format that\n> can support storing reflog entries at some point in the future too,\n> rather than build on (the limitations) of packed-refs?\n\nMy personal feeling is that storing ref tips and storing the history of a\nref are sufficiently different problems that should have their own data\nstructures. Even if they could be combined by a common format, I don't\nthink it is safe to transition every part of every ref operation to a new\nformat all at once.\n\nLooking at reftable from the perspective of a hosting provider, I'm very\nhesitant to recommend transitioning to it because of how it is an \"all or\nnothing\" switch. It does not fit with my expectations for safe deployment\npractices.\n\nYes, packed-refs have some limitations, but those limitations are known\nand we are working within them right now. I'd rather make a change to\nwrite smaller versions of the file with the same semantics as a first\nstep.\n\n> Or is\n> packed-refs v2 backward compatible with v1 (could an old git client\n> read v2 files? I think not, right?).\n\nNo, it is not backward compatible. That's why the extension is needed.\n\n> The reftable project has gotten into a slump because my work\n> responsibilities have increased over the last 1.5 year squeezing down\n> how much time I have for 'fun' projects. I chatted with John Cai, who\n> was trying to staff this project out of Gitlab resources. I don't know\n> where that stands, though.\n\nI'll have my EM reach out to John to see where that stands to see how we\ncan coordinate in this space.\n\n>> The one thing I can say is that if the reftable work had not already begun,\n>> then this is RFC is how I would have approached a new ref format.\n>>\n>> I look forward to your feedback!\n>\n> Hope this helps.\n\nIt does help clarify where the reftable project currently stands as well\nas some key limitations of the packed-refs format. You've given me a lot\nto think about so I'll do some poking around in your branch (and do some\nperformance tests) to see what I can make of it.\n\nLet me attempt to summarize my understanding, now that you've added\nclarity:\n\n* The reftable work needs its refs backend implemented, but your draft PR\n  has a prototype of this and some basic test suite integration. There are\n  54 test files that have one or more failing tests, and likely these just\n  need to be adjusted to not care about loose references.\n\n* The reftable is currently fundamentally different enough that it could\n  not be used as a replacement for the packed-refs file underneath loose\n  refs (primarily due to its integration with the reflog). Doing so would\n  require significant work on top of your prototype.\n\n* This further indicates that moving to reftable is an \"all or nothing\"\n  transition, and even requires starting a repository from scratch with\n  reftable enabled. This is a bit of a blocker for a hosting provider to\n  transition to the format, and will likely be difficult for clients to\n  adopt the feature.\n\n* The plan established by this RFC does _not_ block reftable progress, but\n  generally we prefer not having competing formats in Git, so it would be\n  better to have only one, unless there is enough of a justification to\n  have different formats for different use cases.\n\nI'm going to take the following actions on my end to better understand the\nsituation:\n\n1. I'll take your draft PR branch and do some performance evaluations on\n   the speed of ref updates compared to loose refs and my prototype of a\n   two-stack packed-ref where the second layer of the stack is only for\n   deleted refs.\n\n2. I'll consult with my peers to determine how expensive it would be to\n   roll out reftable via a complete replacement of our hosted\n   repositories. I'll also try to discover ways to roll out the feature to\n   subsets of the fleet to create a safe deployment strategy.\n\n3. My EM and I will reach out to John Cai to learn about plans to push\n   reftable over the finish line.\n\n4. I will split out the \"skip_hash\" part of this RFC into its own series,\n   after adding the necessary details to fsck to understand a null\n   trailing hash.\n\nPlease let me know if I'm missing anything I should be investigating here.\n\nThanks,\n-Stolee\n"},{"id":"468253","messageId":"c104b0e6-1287-b4ab-dd4a-b4b6eef9996c@github.com","threadId":"58758","inReplyTo":"xmqqiljbkfg9.fsf@gitster.g","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-11-30T15:31:09Z","receivedAt":"2022-11-30T15:31:16Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/18/2022 6:31 PM, Junio C Hamano wrote:\n\n> I have been and am still offline and haven't examined this proposal\n> in detail, but would it be a better longer-term approach to improve\n> reftable backend, instead of piling more effort on loose+packed\n> filesystem based backend?\n\nIf reftable was complete and stable, then I would have carefully examined\nit to check that it solves the problems at hand. I interpreted the lack of\nprogress in the area to be due to significant work required or hard\nproblems blocking its completion. That appears to be a wrong assumption,\nso we are exploring what that will take to get it complete.\n\nI am still wary of it requiring a clean slate and not having any way to\nupgrade from an existing repository to one with reftable. I'm going to\nreevaluate this to see how expensive it would be to upgrade to reftable\nand how we can deploy that change safely.\n\nThese upgrade concerns may require us to eventually consider a world where\nwe can upgrade a repository by replacing the packed-refs file with a\nreftable file, then later removing the ability to read or write loose\nrefs. To do so might benefit from the multi-valued extensions.refFormat\nthat is proposed in this RFC, even if packed-v2 does not become a\nrecognized value.\n\nMy personal opinion is that if reftable was not already implemented in\nJGit and was not already partially contributed to Git, then we would not\nchoose that format or that \"all or nothing\" upgrade path. Instead,\nincremental improvements on the existing ref formats are easier to\nunderstand and test in parts.\n\nBut my opinion is not the most important one. I'll defer to the\ncommunity in this. I thought it worthwhile to present an alternative.\n\nThanks,\n-Stolee\n\n"},{"id":"468254","messageId":"0e3ab2bb-68ed-8e77-ac0b-afeb5ca88def@dunelm.org.uk","threadId":"58758","inReplyTo":"f1c45bd5-692e-85db-90c3-c516003f47e5@github.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2022-11-30T15:38:29Z","receivedAt":"2022-11-30T15:38:37Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Stolee\n\nOn 30/11/2022 15:16, Derrick Stolee wrote:\n> On 11/28/2022 1:56 PM, Han-Wen Nienhuys wrote:\n>> * Worktrees and the main repository have a separate view of the ref\n>> namespace. This is not explicit in the ref backend API, and there is a\n>> technical limitation that the packed-refs file cannot be in a\n>> worktree. This means that worktrees will always continue to use\n>> loose-ref storage if you only extend the packed-refs backend.\n> \n> If I'm understanding it correctly [1], only the special refs (like HEAD or\n> REBASE_HEAD) are worktree-specific, and all refs under \"refs/*\" are\n> repository-scoped. I don't actually think of those special refs as \"loose\"\n> refs and thus they should still work under the \"only packed-refs\" value\n> for extensions.refFormat. I should definitely cover this in the\n> documentation, though. Also, [1] probably needs updating because it calls\n> HEAD a pseudo ref even though it explicitly is not [2].\n>\n > [1] https://git-scm.com/docs/git-worktree#_refs\n > [2] \nhttps://git-scm.com/docs/gitglossary#Documentation/gitglossary.txt-aiddefpseudorefapseudoref\n\nUnfortunately I think it is a little messier than that (see \nrefs.c:is_per_worktree_ref()). refs/bisect/*, refs/rewritten/* and \nrefs/worktree/* are all worktree specific.\n\nBest Wishes\n\nPhillip\n"},{"id":"468257","messageId":"Y4eG1SDxBqjwUhwM@nand.local","threadId":"58758","inReplyTo":"f1c45bd5-692e-85db-90c3-c516003f47e5@github.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-11-30T16:37:41Z","receivedAt":"2022-11-30T16:37:59Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Nov 30, 2022 at 10:16:52AM -0500, Derrick Stolee wrote:\n> > Do you mean 'enumerate refs' ? Why would you want to count refs by prefix?\n>\n> * There was a GitHub feature that counted refs and tags, but wanted to\n>   ignore internal ref prefixes (outside of refs/heads/* or refs/tags/*).\n>   It turns out that we didn't actually need the full count but an\n>   existence indicator, but it would be helpful to quickly identify how\n>   many branches or tags are in a repository at a glance. Packed-refs v1\n>   requires scanning the whole file while packed-refs v2 does a fixed\n>   number of binary searches followed by a subtraction of row indexes.\n\nTrue. On the surface, it seemed odd to use a function which returns\nsomething like:\n\n    { \"refs/heads/*\" => NNNN, \"refs/tags/*\" => MMMM }\n\nonly to check whether or not NNNN and MMMM are zero or non-zero.\n\nBut there's a little more to the story. That emptiness check does occur\nat the beginning of many page loads. But when it responds \"non-empty\",\nwe then care about how many branches and tags there actually are.\n\nSo calling count_refs() (the name of the internal RPC that powers all of\nthis) was an optimization written under the assumption that we actually\nare going to ask about the exact number of branches/tags very shortly\nafter querying for emptiness.\n\nIt turns out that empirically it's faster to do something like:\n\n    $ git show-ref --[heads|tags] | head -n 1\n\nto check if there are any branches and tags at all[^1], and then a\nfollow up 'git show-ref --heads | wc -l' to check how many there are.\n\nBut it would be nice to do both operations quickly without having\nactually scan all of the entries in each prefix.\n\nThanks,\nTaylor\n\n[^1]: Some may remember my series in\n  https://lore.kernel.org/git/cover.1654552560.git.me@ttaylorr.com/\n  which replaced '| head -n 1' with a '--count=1' option. This matches\n  what GitHub runs in production where piping one command to another\n  from Ruby is unfortunately quite complicated.\n"},{"id":"468268","messageId":"CAFQ2z_MLwUoaSTG04LJYHgJH-QYJEuZ9bQcTsV8mXwxBbz7Egg@mail.gmail.com","threadId":"58758","inReplyTo":"f1c45bd5-692e-85db-90c3-c516003f47e5@github.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2022-11-30T18:30:06Z","receivedAt":"2022-11-30T18:30:31Z","isPatch":true,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Wed, Nov 30, 2022 at 4:16 PM Derrick Stolee <derrickstolee@github.com> wrote:\n> > * Symrefs are refs too, but for some reason the packed-refs file\n> > doesn't support them. Does packed-refs v2 support symrefs too?  If you\n> > want to snapshot the state of refs, do you want to snapshot the value\n> > of HEAD too?\n>\n> I forgot that loose refs under .git/refs/ can be symrefs. This definitely\n> is a limitation that I should mention. Again, pseudorefs like HEAD are not\n> included and are stored separately, but symrefs within refs/* are not\n> available in packed-refs (v1 or v2). That should be explicitly called out\n> in the extensions.refFormat docs.\n>\n> I imagine that such symrefs are uncommon, and users can make their own\n> evaluation of whether that use is worth keeping loose refs or not. We can\n> still have the {files, packed[-v2]} extension value while having a\n> writing strategy that writes as much as possible into the packed layer.\n\nTo be honest, I don't understand why symrefs are such a generic\nconcept; I've only ever seen them used for HEAD.\n\n> > * By not changing reflogs, you are making things simpler. (if a\n> > transaction updates the branch that HEAD points to, the reflog for\n> > HEAD has to be updated too. Because reftable updates the reflog\n> > transactionally, this was some extra work)\n> > Then again, I feel the current way that reflogs work are a bit messy,\n> > because directory/file conflicts force reflogs to be deleted at times\n> > that don't make sense from a user-perspective.\n>\n> I agree that reflogs are messy. I also think that reflogs have different\n> needs than the ref storage, so separating their needs is valuable.\n\nIf the reflog records the history of the ref database, then ideally,\nan update of a ref should be transactional across the ref database and\nthe reflog. I think you can never make this work unless you tie the\nstorage of both together.\n\nI can't judge how many hosting providers really care about this. At\ngoogle, we really care, but we keep the ref database and the refllog\nin a global Spanner database. Reftable is only used for per-datacenter\nserving. (I discovered some bugs in the JGit reflog code when I ported\nit to local filesystem repos, because it was never exercised at\nGoogle)\n\n> > * There are a lot of commands that store SHA1s in files under .git/,\n> > and access them as if they are a ref (for example: rebase-apply/ ,\n> > CHERRY_PICK_HEAD etc.).\n>\n> Yes, I think these pseudorefs are stored differently from usual refs, and\n> hence the {packed[-v2]} extension value would still work, but I'll confirm\n> this with more testing.\n\nThey will work as long as you keep support for loose refs, because\nthere is no distinction between \"a entry in the ref database\" and \"any\nfile randomly written into .git/ \".\n\n> >> In this RFC, I propose a different model that allows for more customization\n> >> and incremental updates. The extensions.refFormat config key is multi-valued\n> >> and defaults to the list of files and packed. In the context of this RFC,\n> >> the intention is to be able to add packed-v2 so the list of all three values\n> >> would allow Git to write and read either file format version (v1 or v2). In\n> >> the larger scheme, the extension could allow restricting to only loose refs\n> >> (just files) or only packed-refs (just packed) or even later when reftable\n> >> is complete, files and reftable could mean that loose refs are the primary\n> >> ref storage, but the reftable format serves as a drop-in replacement for the\n> >> packed-refs file. Not all combinations need to be understood by Git, but\n> >\n> > I'm not sure how feasible this is. reftable also holds reflog data. A\n> > setting {files,reftable} would either not work, or necessitate hairy\n> > merging of data to get the reflogs working correctly.\n>\n> In this setup, would it be possible to continue using the \"loose reflog\"\n> format while using reftable as the packed layer? I personally think this\n> combination of formats to be critical to upgrading existing repositories\n> to reftable.\n\nI suppose so? If you only store refs and tags (and don't handle\nreflogs, symrefs or use the inverse object mapping) then the reftable\nfile format is just a highly souped-up version of packed-refs.\n\n> (Note: there is a strategy that doesn't need this approach, but it's a bit\n> complicated. It would involve rotating all replicas to new repositories\n> that are configured to use reftable upon creation, getting the refs from\n> other replicas via fetches. In my opinion, this is prohibitively\n> expensive.)\n\nI'm not sure I understand the problem. Any deletion of a ref (that is\nin packed-refs) today already requires rewriting the entire\npacked-refs file (\"all or nothing\" operation). Whether you write a\npacked-refs or reftable is roughly equally expensive.\n\nAre you looking for a way to upgrade a repo, while concurrent git\nprocess may write updates into the repository during the update? That\nmay be hard to pull off, because you probably need to rename more than\none file atomically. If you accept momentarily failed writes, you\ncould do\n\n* rename refs/ to refs.old/ (loose ref writes will fail now)\n* collect loose refs under refs.old/ , put into packed-refs\n* populate the reftable/ dir\n* set refFormat extension.\n* rename refs.old/ to refs/ with a refs/heads a file (as described in\nthe reftable spec.)\n\nSee also https://gerrit.googlesource.com/jgit/+/ca166a0c62af2ea87fdedf2728ac19cb59a12601/org.eclipse.jgit/src/org/eclipse/jgit/internal/storage/file/FileRepository.java#734\n\n> >> I mentioned earlier that I had considered using reftable as a way to achieve\n> >> the stated goals. With the current state of that work, I'm not confident\n> >> that it is the right approach here.\n> >>\n> >> My main worry is that the reftable is more complicated than we need for a\n> >> typical Git repository that is based on a typical filesystem. This makes\n> >> testing the format very critical, and we seem to not be near reaching that\n> >> approach.\n> >\n> > I think the base code of reading and writing the reftable format is\n> > exercised quite exhaustively tested in unit tests. You say 'seem', but\n> > do you have anything concrete to say?\n>\n> Our test suite is focused on integration tests at the command level. While\n> unit tests are helpful, I'm not sure if all of the corner cases would be\n> covered by tests that check Git commands only.\n\nIt's actually easier to test all of the nooks of the format through\nunittests, because you can tweak parameters (eg. blocksize) that\naren't normally available in the command-line\n\n> >> That might even help the integration process to allow the reftable format to\n> >> be tested at least by some subset of tests instead of waiting for a full\n> >> test suite update.\n> >\n> > I don't understand this comment. In the current state,\n> > https://github.com/git/git/pull/1215 already passes 922 of the 968\n> > test files if you set GIT_TEST_REFTABLE=1.\n> >\n> > See https://github.com/git/git/pull/1215#issuecomment-1329579459 for\n> > details. As you can see, for most test files, it's just a few\n> > individual test cases that fail.\n>\n> My point is that to get those remaining tests passing requires a\n> significant update to the test suite. I imagined that the complexity of\n> that update was the blocker to completing the reftable work.\n>\n> It seems that my estimation of that complexity was overly high compared to\n> what you appear to be describing.\n\nTo be honest, i'm not quite sure how significant the work is: for\nthings like worktrees, it wasn't that obvious to me how things should\nwork in the first place. That makes it hard to make estimates. I\nthought there might be a month of full-time work left, but these days\nI can barely make a couple of hours of time per week to work on\nreftable  if at all.\n\n> > For deleting refs quickly, it seems that you only need to support\n> > $ZEROID in packed-refs and then implement a ref database as a stack of\n> > packed-ref files? If you're going for minimal effort and minimal\n> > disruption wouldn't that be the place to start?\n>\n> I disagree that jumping straight to stacked packed-refs is minimal effort\n> or minimal disruption.\n>\n> Creating the stack approach does require changing the semantics of the\n> packed-refs format to include $ZEROID, which will modify some meanings in\n> the iteration code. The use of a stack, as well as how layers are combined\n> during a ref write or also during maintenance, adds complications to the\n> locking semantics that are decently complicated.\n>\n> By contrast, the v2 format is isolated to the on-disk format. None of the\n> writing or reading semantics are changed in terms of which files to look\n> at or write in which order. Instead, it's relatively simple to see from\n> the format exactly how it reduces the file size but otherwise has exactly\n> the same read/write behavior. In fact, since the refs and OIDs are all\n> located in the same chunk in a similar order to the v1 file, we can even\n> deduce that page cache semantics will only improve in the new format.\n>\n> The reason to start with this step is that the benefits and risks are\n> clearly understood, which can motivate us to establish the mechanism for\n> changing the ref format by defining the extension.\n\nI believe that the v2 format is a safe change with performance\nimprovements, but it's a backward incompatible format change with only\nmodest payoff. I also don't understand how it will help you do a stack\nof tables,\nwhich you need for your primary goal (ie. transactions/deletions\nwriting only the delta, rather than rewriting the whole file?).\n\n> > You're concerned about the reftable file format (and maybe rightly\n> > so), but if you're changing the file format anyway and you're not\n> > picking reftable, why not create a block-based, indexed format that\n> > can support storing reflog entries at some point in the future too,\n> > rather than build on (the limitations) of packed-refs?\n>\n> My personal feeling is that storing ref tips and storing the history of a\n> ref are sufficiently different problems that should have their own data\n> structures. Even if they could be combined by a common format, I don't\n> think it is safe to transition every part of every ref operation to a new\n> format all at once.\n>\n> Looking at reftable from the perspective of a hosting provider, I'm very\n> hesitant to recommend transitioning to it because of how it is an \"all or\n> nothing\" switch. It does not fit with my expectations for safe deployment\n> practices.\n\nYou'd have to consult with your SRE team, how to do this best, but\nhere's my $.02. If you are a hosting provider, I assume you have 3 or\n5 copies of each repo in diffrent datacenters for\nredundancy/availability. You could have one of the datacenters use the\nnew format for while, and see if there are any errors or discrepancies\n(both in terms of data consistency and latency metrics)\n\n> * The reftable work needs its refs backend implemented, but your draft PR\n>   has a prototype of this and some basic test suite integration. There are\n>   54 test files that have one or more failing tests, and likely these just\n>   need to be adjusted to not care about loose references.\n>\n> * The reftable is currently fundamentally different enough that it could\n>   not be used as a replacement for the packed-refs file underneath loose\n>   refs (primarily due to its integration with the reflog). Doing so would\n>   require significant work on top of your prototype.\n\nIt could, but I don't see the point.\n\n> I'm going to take the following actions on my end to better understand the\n> situation:\n>\n> 1. I'll take your draft PR branch and do some performance evaluations on\n>    the speed of ref updates compared to loose refs and my prototype of a\n>    two-stack packed-ref where the second layer of the stack is only for\n>    deleted refs.\n\n(tangent) - wouldn't that design perform poorly once the number of\ndeletions gets large? You'd basically have to rewrite the\ndeleted-packed-refs file all the time.\n\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\n\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\n\nRegistergericht und -nummer: Hamburg, HRB 86891\n\nSitz der Gesellschaft: Hamburg\n\nGeschäftsführer: Paul Manicle, Liana Sebastian\n"},{"id":"468271","messageId":"87cz94xozi.fsf@gmail.com","threadId":"58758","inReplyTo":"CAFQ2z_MLwUoaSTG04LJYHgJH-QYJEuZ9bQcTsV8mXwxBbz7Egg@mail.gmail.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Sean Allred","fromEmail":"allred.sean@gmail.com","sentAt":"2022-11-30T18:37:21Z","receivedAt":"2022-11-30T18:43:34Z","isPatch":true,"sender":{"key":"allred.sean@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2082195?v=4"},"body":"\nHan-Wen Nienhuys <hanwen@google.com> writes:\n> To be honest, I don't understand why symrefs are such a generic\n> concept; I've only ever seen them used for HEAD.\n\nI've been only lurking in this thread (and loosely following along,\neven!) but I do want to call out that I have recently considered perhaps\nabusing symrefs to point to normal feature branches. In our workflow, we\nhave documentation records identified by a numeric ID -- the code\nchanges corresponding to that documentation (testing instructions, etc.)\nuse formulaic branch names like `feature/123456`.\n\nIt is sometimes beneficial for two or more of these documentation\nrecords to perform their work on the same code branch. There are myriad\nreasons for this, some better than others, but I want to avoid getting\nmired in whether or not this is a good idea. It does happen and is\nsometimes even the best way to do it.\n\nIn these scenarios, I've considered having `feature/2` be a symref to\n`feature/1` so that both features can always 'know' what to call their\nbranch for operations like checkout. I've done this on a smaller scale\nin the past to great effect.\n\nNothing is set in stone here for us, but I did want to call this out as\na potential real-world use case.\n\n--\nSean Allred\n"},{"id":"468292","messageId":"xmqqedtkuk6m.fsf@gitster.g","threadId":"58758","inReplyTo":"f1c45bd5-692e-85db-90c3-c516003f47e5@github.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-11-30T22:55:13Z","receivedAt":"2022-11-30T22:55:17Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <derrickstolee@github.com> writes:\n\n> I do want to say that while I admire\n> JGit's dedication to being compatible with repositories created by Git, I\n> don't think the reverse is a goal of the Git project.\n\nThe world works better if cross-pollination happens both ways,\nthough.\n\n>> * Symrefs are refs too, but for some reason the packed-refs file\n>> doesn't support them. Does packed-refs v2 support symrefs too?  If you\n>> want to snapshot the state of refs, do you want to snapshot the value\n>> of HEAD too?\n>\n> I forgot that loose refs under .git/refs/ can be symrefs. This definitely\n> is a limitation that I should mention. Again, pseudorefs like HEAD are not\n> included and are stored separately, but symrefs within refs/* are not\n> available in packed-refs (v1 or v2). That should be explicitly called out\n> in the extensions.refFormat docs.\n\nI expect that, in a typical individual-contributor repository, there\nare at least two symbolic refs, e.g.\n\n    .git/HEAD\n    .git/refs/remotes/origin/HEAD\n\nHaving to fall back on the loose ref hierarchy only to be able to\nstore the latter is a bit of shame---as long as you are revamping\nthe format, the design should allow us to eventually migrate all\nrefs to the new format without having to do the \"check there, and if\nthere isn't then check this other place\", which is what the current\nloose + packed combination do, I would think.\n"},{"id":"468351","messageId":"f5370fec-d517-eaa9-8e16-82fa20ac8532@github.com","threadId":"58758","inReplyTo":"CAFQ2z_MLwUoaSTG04LJYHgJH-QYJEuZ9bQcTsV8mXwxBbz7Egg@mail.gmail.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-12-01T20:18:57Z","receivedAt":"2022-12-01T20:19:07Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/30/2022 1:30 PM, Han-Wen Nienhuys wrote:\n> On Wed, Nov 30, 2022 at 4:16 PM Derrick Stolee <derrickstolee@github.com> wrote:\n>> (Note: there is a strategy that doesn't need this approach, but it's a bit\n>> complicated. It would involve rotating all replicas to new repositories\n>> that are configured to use reftable upon creation, getting the refs from\n>> other replicas via fetches. In my opinion, this is prohibitively\n>> expensive.)\n> \n> I'm not sure I understand the problem. Any deletion of a ref (that is\n> in packed-refs) today already requires rewriting the entire\n> packed-refs file (\"all or nothing\" operation). Whether you write a\n> packed-refs or reftable is roughly equally expensive.\n> \n> Are you looking for a way to upgrade a repo, while concurrent git\n> process may write updates into the repository during the update? That\n> may be hard to pull off, because you probably need to rename more than\n> one file atomically. If you accept momentarily failed writes, you\n> could do\n> \n> * rename refs/ to refs.old/ (loose ref writes will fail now)\n> * collect loose refs under refs.old/ , put into packed-refs\n> * populate the reftable/ dir\n> * set refFormat extension.\n> * rename refs.old/ to refs/ with a refs/heads a file (as described in\n> the reftable spec.)\n>\n> See also https://gerrit.googlesource.com/jgit/+/ca166a0c62af2ea87fdedf2728ac19cb59a12601/org.eclipse.jgit/src/org/eclipse/jgit/internal/storage/file/FileRepository.java#734\n\nYes, I would ideally like for the repository to \"upgrade\" its ref\nstorage mechanism during routine maintenance in a non-blocking way\nwhile other writes and reads continue as normal.\n\nAfter discussing it a bit internally, we _could_ avoid the \"rotate\nthe replicas\" solution if there was a \"git upgrade-ref-format\"\ncommand that could switch from one to another, but it would still\ninvolve pulling that replica out of the rotation and then having\nit catch up to the other replicas after that is complete. If I'm\nreading your draft correctly, that is not currently available in\nyour work, but we could add it after the fact.\n\nRequiring pulling replicas out of rotation is still a bit heavy-\nhanded for my liking, but it's much less expensive than moving\nall of the Git data.\n\n>> The reason to start with this step is that the benefits and risks are\n>> clearly understood, which can motivate us to establish the mechanism for\n>> changing the ref format by defining the extension.\n> \n> I believe that the v2 format is a safe change with performance\n> improvements, but it's a backward incompatible format change with only\n> modest payoff. I also don't understand how it will help you do a stack\n> of tables,\n> which you need for your primary goal (ie. transactions/deletions\n> writing only the delta, rather than rewriting the whole file?).\n\nThe v2 format doesn't help me on its own, but it has other benefits\nin terms of size and speed, as well as the \"ref count\" functionality.\n\nThe important thing is that the definition of extensions.refFormat\nthat I'm proposing in this RFC establishes a way to make incremental\nprogress on the ref format, allowing the stacked format to come in\nlater with less friction.\n \n>> * The reftable is currently fundamentally different enough that it could\n>>   not be used as a replacement for the packed-refs file underneath loose\n>>   refs (primarily due to its integration with the reflog). Doing so would\n>>   require significant work on top of your prototype.\n> \n> It could, but I don't see the point.\n\nMy point is that we can upgrade repositories by replacing packed-refs\nwith reftable during routine maintenance instead of the heavier\napproaches discussed earlier.\n\n* Step 1: replace packed-refs with reftable.\n* Step 2: stop writing loose refs, only update reftable (but still read loose refs).\n* Step 3: collapse all loose refs into reftable, stop reading or writing loose refs.\n \n>> I'm going to take the following actions on my end to better understand the\n>> situation:\n>>\n>> 1. I'll take your draft PR branch and do some performance evaluations on\n>>    the speed of ref updates compared to loose refs and my prototype of a\n>>    two-stack packed-ref where the second layer of the stack is only for\n>>    deleted refs.\n> \n> (tangent) - wouldn't that design perform poorly once the number of\n> deletions gets large? You'd basically have to rewrite the\n> deleted-packed-refs file all the time.\n \nWe have regular maintenance that is triggered by pushes that rewrites\nthe packed-refs file frequently, anyway. The maintenance currently is\nblocked on the amount of time spent repacking object data, so a large\nnumber of ref updates can come in during this process. (That maintenance\nstep would collapse the deleted-refs layer into the base layer.)\n\nI've tested a simple version of this stack that shows that rewriting the\nfile with 1,000 deletions is still within 2x the cost of updating a loose\nref, so it solves the immediate problem using a much simpler stack model,\nat least in the most-common case where ref deletions are less frequent\nthan other updates. Even if the size outgrew the 2x cost limit, the\ndeleted file is still going to be much smaller than the base packed-refs\nfile, which is currently rewritten for every deletion, so it is still an\nimprovement.\n\nThe more complicated stack model would be required to funnel all ref\nupdates into that structure and away from loose refs.\n\nThanks,\n-Stolee\n"},{"id":"468445","messageId":"CAFQ2z_NKpgsEsrDdkdp=HDajrzpUDjiUcUdR8TMkYpXZBU0k+g@mail.gmail.com","threadId":"58758","inReplyTo":"f5370fec-d517-eaa9-8e16-82fa20ac8532@github.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2022-12-02T16:46:51Z","receivedAt":"2022-12-02T16:47:08Z","isPatch":true,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Thu, Dec 1, 2022 at 9:19 PM Derrick Stolee <derrickstolee@github.com> wrote:\n> >> The reason to start with this step is that the benefits and risks are\n> >> clearly understood, which can motivate us to establish the mechanism for\n> >> changing the ref format by defining the extension.\n> >\n> > I believe that the v2 format is a safe change with performance\n> > improvements, but it's a backward incompatible format change with only\n> > modest payoff. I also don't understand how it will help you do a stack\n> > of tables,\n> > which you need for your primary goal (ie. transactions/deletions\n> > writing only the delta, rather than rewriting the whole file?).\n>\n> The v2 format doesn't help me on its own, but it has other benefits\n> in terms of size and speed, as well as the \"ref count\" functionality.\n>\n> The important thing is that the definition of extensions.refFormat\n> that I'm proposing in this RFC establishes a way to make incremental\n> progress on the ref format, allowing the stacked format to come in\n> later with less friction.\n\nI guess you want to move the read/write stack under the loose storage\n(packed backend), and introduce (read loose/packed + write packed\nonly) mode that is transitional?\n\nBefore you embark on this incremental route, I think it would be best\nto think through carefully how an online upgrade would work in detail\n(I think it's currently not specified?) If ultimately it's not\nfeasible to do incrementally, then the added complexity of the\nincremental approach will be for naught.\n\nThe incremental mode would only be of interest to hosting providers.\nIt will only be used transitionally. It is inherently going to be\ncomplex, because it has to consider both storage modes at the same\ntime, and because it is transitional, it will get less real life\ntesting. At the same time, the ref database is comparatively small, so\nthe availability blip that converting the storage offline will impair\nis going to be small. So, the incremental approach is rather expensive\nfor a comparatively small benefit.\n\nI also thought a bit about how you could make the transition seamless,\nbut I can't see a good way: you have to coordinate between tables.list\n(the list of reftables active or whatever file signals the presence of\na stack) and files under refs/heads/. I don't know how to do\ntransactions across multiple files without cooperative locking.\n\nIf you assume you can use filesystem locks, then you could do\nsomething simpler: if a git repository is marked 'transitional', git\nprocesses take an FS read lock on .git/ .  The process that converts\nthe storage can take an exclusive (write) lock on .git/, so it knows\nnobody will interfere. I think this only works if the repo is on local\ndisk rather than NFS, though.\n\n> * Step 1: replace packed-refs with reftable.\n> * Step 2: stop writing loose refs, only update reftable (but still read loose refs).\n\nDoes that work? A long running process might not notice the switch in\nstep 2, so it could still write a ref as loose, while another process\nracing might write a different value to the same ref through reftable.\n\nPS. I'll be away from work until Jan 9th.\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\nRegistergericht und -nummer: Hamburg, HRB 86891\nSitz der Gesellschaft: Hamburg\nGeschäftsführer: Paul Manicle, Liana Sebastian\n"},{"id":"468456","messageId":"221202.86fsdxfyny.gmgdl@evledraar.gmail.com","threadId":"58758","inReplyTo":"CAFQ2z_NKpgsEsrDdkdp=HDajrzpUDjiUcUdR8TMkYpXZBU0k+g@mail.gmail.com","subject":"Re: [PATCH 00/30] [RFC] extensions.refFormat and packed-refs v2 file format","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-12-02T18:24:34Z","receivedAt":"2022-12-02T18:28:54Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Dec 02 2022, Han-Wen Nienhuys wrote:\n\n> On Thu, Dec 1, 2022 at 9:19 PM Derrick Stolee <derrickstolee@github.com> wrote:\n>> >> The reason to start with this step is that the benefits and risks are\n>> >> clearly understood, which can motivate us to establish the mechanism for\n>> >> changing the ref format by defining the extension.\n>> >\n>> > I believe that the v2 format is a safe change with performance\n>> > improvements, but it's a backward incompatible format change with only\n>> > modest payoff. I also don't understand how it will help you do a stack\n>> > of tables,\n>> > which you need for your primary goal (ie. transactions/deletions\n>> > writing only the delta, rather than rewriting the whole file?).\n>>\n>> The v2 format doesn't help me on its own, but it has other benefits\n>> in terms of size and speed, as well as the \"ref count\" functionality.\n>>\n>> The important thing is that the definition of extensions.refFormat\n>> that I'm proposing in this RFC establishes a way to make incremental\n>> progress on the ref format, allowing the stacked format to come in\n>> later with less friction.\n>\n> I guess you want to move the read/write stack under the loose storage\n> (packed backend), and introduce (read loose/packed + write packed\n> only) mode that is transitional?\n>\n> Before you embark on this incremental route, I think it would be best\n> to think through carefully how an online upgrade would work in detail\n> (I think it's currently not specified?) If ultimately it's not\n> feasible to do incrementally, then the added complexity of the\n> incremental approach will be for naught.\n>\n> The incremental mode would only be of interest to hosting providers.\n> It will only be used transitionally. It is inherently going to be\n> complex, because it has to consider both storage modes at the same\n> time, and because it is transitional, it will get less real life\n> testing. At the same time, the ref database is comparatively small, so\n> the availability blip that converting the storage offline will impair\n> is going to be small. So, the incremental approach is rather expensive\n> for a comparatively small benefit.\n>\n> I also thought a bit about how you could make the transition seamless,\n> but I can't see a good way: you have to coordinate between tables.list\n> (the list of reftables active or whatever file signals the presence of\n> a stack) and files under refs/heads/. I don't know how to do\n> transactions across multiple files without cooperative locking.\n\nA multi-backend transaction would be hard to do at the best of times,\nbut we'd also presumably run into the issue that not all ref operations\ncurrently use the transaction mechanism (e.g. branch copying/moving). So\nif one or the other fails there all bets are off as far as getting back\nto a consistent state.\n\nPerhaps a more doable & interesting approach would be to have a \"slave\"\nbackend that would follow along, i.e. we'd replay all operations from\n\"master\" to \"slave\" (as with DB replication, just within a single\nrepository).\n\nWe might get out of sync, but as the \"master\" is always the source of\ntruth presumably we could run some one-off re-exporting of the refspace\nget back up-to-date, and hopefully not get out of sync again.\n\nThen once we're ready, we could flip the switch indicating what becomes\nthe canonical backend.\n\nFor reftable the FS layout under .git/* is incompatible, so we'd also\nneed to support writing to some alternate directory to make such a thing\nwork...\n"}]}