{"thread":{"id":"56996","subject":"[PATCH 00/17] cruft packs","startedAt":"2021-11-29T22:25:55Z","lastAt":"2023-06-01T13:01:52Z","messageCount":201,"participants":["Taylor Blau","Derrick Stolee","brian m. carlson","Junio C Hamano","Elijah Newren","Ævar Arnfjörð Bjarmason","Jonathan Nieder","rsbecker@nexbridge.com","René Scharfe","Andreas Schwab"],"isPatch":true,"patchVersion":1,"patchTotal":17},"messages":[{"id":"442515","messageId":"cover.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":null,"subject":"[PATCH 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:02Z","receivedAt":"2021-11-29T22:25:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This series implements \"cruft packs\", a pack which stores accumulated\nunreachable objects, along with a new \".mtimes\" file which tracks each\nobject's last known modification time.\n\nThis idea was discussed recently-ish in [1], but the most thorough\ndiscussion I could find is in [2]. The approach settled on in this\nseries is laid out in detail by the first patch.\n\nFor the uninitiated, cruft packs enable repositories to safely run\n`git repack -Ad` by storing unreachable objects which have not yet\n\"aged out\" in a separate pack. This prevents repositories from storing\na potentially large number of these such objects as loose.\n\nThis series is structured as follows:\n\n  - The first patch describes the technical details of cruft packs.\n  - The next five patches implement reading and writing the new\n    `.mtimes` format.\n  - The next six patches implement `git pack-objects --cruft`. The\n    first five implement this mode when no grace period is specified,\n    and the six patch adds support for the grace period.\n  - The next five patches integrate cruft packs with `git repack`,\n    including the new-ish `--geometric` mode.\n  - The final patch handles object freshening for objects stored in a\n    cruft pack.\n\nThanks in advance for your review.\n\n[1]: https://lore.kernel.org/git/20170610080626.sjujpmgkli4muh7h@sigill.intra.peff.net/\n[2]: https://lore.kernel.org/git/E1SdhJ9-0006B1-6p@tytso-glaptop.cam.corp.google.com/\n\nTaylor Blau (17):\n  Documentation/technical: add cruft-packs.txt\n  pack-mtimes: support reading .mtimes files\n  pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'\n  chunk-format.h: extract oid_version()\n  pack-mtimes: support writing pack .mtimes files\n  t/helper: add 'pack-mtimes' test-tool\n  builtin/pack-objects.c: return from create_object_entry()\n  builtin/pack-objects.c: --cruft without expiration\n  reachable: add options to add_unseen_recent_objects_to_traversal\n  reachable: report precise timestamps from objects in cruft packs\n  builtin/pack-objects.c: --cruft with expiration\n  builtin/repack.c: support generating a cruft pack\n  builtin/repack.c: allow configuring cruft pack generation\n  builtin/repack.c: use named flags for existing_packs\n  builtin/repack.c: add cruft packs to MIDX during geometric repack\n  builtin/gc.c: conditionally avoid pruning objects via loose\n  sha1-file.c: don't freshen cruft packs\n\n Documentation/Makefile                  |   1 +\n Documentation/config/gc.txt             |  21 +-\n Documentation/config/repack.txt         |   9 +\n Documentation/git-gc.txt                |   5 +\n Documentation/git-pack-objects.txt      |  23 +\n Documentation/git-repack.txt            |  11 +\n Documentation/technical/cruft-packs.txt |  95 ++++\n Documentation/technical/pack-format.txt |  22 +\n Makefile                                |   2 +\n builtin/gc.c                            |  10 +-\n builtin/pack-objects.c                  | 306 ++++++++++-\n builtin/repack.c                        | 189 ++++++-\n bulk-checkin.c                          |   2 +-\n chunk-format.c                          |  12 +\n chunk-format.h                          |   3 +\n commit-graph.c                          |  18 +-\n midx.c                                  |  18 +-\n object-file.c                           |   4 +-\n object-store.h                          |   7 +-\n pack-mtimes.c                           | 139 +++++\n pack-mtimes.h                           |  16 +\n pack-objects.c                          |   6 +\n pack-objects.h                          |  20 +\n pack-write.c                            |  90 +++-\n pack.h                                  |   4 +\n packfile.c                              |  18 +-\n packfile.h                              |   1 +\n reachable.c                             |  58 +-\n reachable.h                             |   9 +-\n t/helper/test-pack-mtimes.c             |  53 ++\n t/helper/test-tool.c                    |   1 +\n t/helper/test-tool.h                    |   1 +\n t/t5327-pack-objects-cruft.sh           | 685 ++++++++++++++++++++++++\n 33 files changed, 1757 insertions(+), 102 deletions(-)\n create mode 100644 Documentation/technical/cruft-packs.txt\n create mode 100644 pack-mtimes.c\n create mode 100644 pack-mtimes.h\n create mode 100644 t/helper/test-pack-mtimes.c\n create mode 100755 t/t5327-pack-objects-cruft.sh\n\n-- \n2.34.1.25.gb3157a20e6\n"},{"id":"442516","messageId":"a9f7c738e0ffbc5cdedc26768a0623446c98d239.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:05Z","receivedAt":"2021-11-29T22:25:58Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Create a technical document to explain cruft packs. It contains a brief\noverview of the problem, some background, details on the implementation,\nand a couple of alternative approaches not considered here.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/Makefile                  |  1 +\n Documentation/technical/cruft-packs.txt | 95 +++++++++++++++++++++++++\n 2 files changed, 96 insertions(+)\n create mode 100644 Documentation/technical/cruft-packs.txt\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex ed656db2ae..0b01c9408e 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -91,6 +91,7 @@ TECH_DOCS += MyFirstContribution\n TECH_DOCS += MyFirstObjectWalk\n TECH_DOCS += SubmittingPatches\n TECH_DOCS += technical/bundle-format\n+TECH_DOCS += technical/cruft-packs\n TECH_DOCS += technical/hash-function-transition\n TECH_DOCS += technical/http-protocol\n TECH_DOCS += technical/index-format\ndiff --git a/Documentation/technical/cruft-packs.txt b/Documentation/technical/cruft-packs.txt\nnew file mode 100644\nindex 0000000000..bb54cce1b1\n--- /dev/null\n+++ b/Documentation/technical/cruft-packs.txt\n@@ -0,0 +1,95 @@\n+= Cruft packs\n+\n+Cruft packs offer an alternative to Git's traditional mechanism of removing\n+unreachable objects. This document provides an overview of Git's pruning\n+mechanism, and how cruft packs can be used instead to accomplish the same.\n+\n+== Background\n+\n+To remove unreachable objects from your repository, Git offers `git repack -Ad`\n+(see linkgit:git-repack[1]). Quoting from the documentation:\n+\n+[quote]\n+[...] unreachable objects in a previous pack become loose, unpacked objects,\n+instead of being left in the old pack. [...] loose unreachable objects will be\n+pruned according to normal expiry rules with the next 'git gc' invocation.\n+\n+Unreachable objects aren't removed immediately, since doing so could race with\n+an incoming push which may reference an object which is about to be deleted.\n+Instead, those unreachable objects are stored as loose object and stay that way\n+until they are older than the expiration window, at which point they are removed\n+by linkgit:git-prune[1].\n+\n+Git must store these unreachable objects loose in order to keep track of their\n+per-object mtimes. If these unreachable objects were written into one big pack,\n+then either freshening that pack (because an object contained within it was\n+re-written) or creating a new pack of unreachable objects would cause the pack's\n+mtime to get updated, and the objects within it would never leave the expiration\n+window. Instead, objects are stored loose in order to keep track of the\n+individual object mtimes and avoid a situation where all cruft objects are\n+freshened at once.\n+\n+This can lead to undesirable situations when a repository contains many\n+unreachable objects which have not yet left the grace period. Having large\n+directories in the shards of `.git/objects` can lead to decreased performance in\n+the repository. But given enough unreachable objects, this can lead to inode\n+starvation and degrade the performance of the whole system. Since we\n+can never pack those objects, these repositories often take up a large amount of\n+disk space, since we can only zlib compress them, but not store them in delta\n+chains.\n+\n+== Cruft packs\n+\n+Cruft packs are designed to eliminate the need for storing unreachable objects\n+in a loose state by including the per-object mtimes in a separate file alongside\n+a single pack containing all loose objects.\n+\n+A cruft pack is written by `git repack --cruft` when generating a new pack.\n+linkgit:git-pack-objects[1]'s `--cruft` option. Note that `git repack --cruft`\n+is a classic all-into-one repack, meaning that everything in the resulting pack is\n+reachable, and everything else is unreachable. Once written, the `--cruft`\n+option instructs `git repack` to generate another pack containing only objects\n+not packed in the previous step (which equates to packing all unreachable\n+objects together). This progresses as follows:\n+\n+  1. Enumerate every object, marking any object which is (a) not contained in a\n+     kept-pack, and (b) whose mtime is within the grace period as a traversal\n+     tip.\n+\n+  2. Perform a reachability traversal based on the tips gathered in the previous\n+     step, adding every object along the way to the pack.\n+\n+  3. Write the pack out, along with a `.mtimes` file that records the per-object\n+     timestamps.\n+\n+This mode is invoked internally by linkgit:git-repack[1] when instructed to\n+write a cruft pack. Crucially, the set of in-core kept packs is exactly the set\n+of packs which will not be deleted by the repack; in other words, they contain\n+all of the repository's reachable objects.\n+\n+When a repository already has a cruft pack, `git repack --cruft` typically only\n+adds objects to it. An exception to this is when `git repack` is given the\n+`--cruft-expiration` option, which allows the generated cruft pack to omit\n+expired objects instead of waiting for linkgit:git-gc[1] to expire those objects\n+later on.\n+\n+It is linkgit:git-gc[1] that is typically responsible for removing expired\n+unreachable objects.\n+\n+== Alternatives\n+\n+Notable alternatives to this design include:\n+\n+  - The location of the per-object mtime data, and\n+  - Whether cruft packs should be incremental or not.\n+\n+On the location of mtime data, a new auxiliary file tied to the pack was chosen\n+to avoid complicating the `.idx` format. If the `.idx` format were ever to gain\n+support for optional chunks of data, it may make sense to consolidate the\n+`.mtimes` format into the `.idx` itself.\n+\n+Incremental cruft packs (i.e., where each time a repository is repacked a new\n+cruft pack is generated containing only the unreachable objects introduced since\n+the last time a cruft pack was written) are significantly more complicated to\n+construct, and so aren't pursued here. The obvious drawback to the current\n+implementation is that the entire cruft pack must be re-written from scratch.\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442517","messageId":"7d4ae7bd3e28e2ec904abb37b6f26505e37531c5.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:08Z","receivedAt":"2021-11-29T22:26:00Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"To store the individual mtimes of objects in a cruft pack, introduce a\nnew `.mtimes` format that can optionally accompany a single pack in the\nrepository.\n\nThe format is defined in Documentation/technical/pack-format.txt, and\nstores a 4-byte network order timestamp for each object in name (index)\norder.\n\nThis patch prepares for cruft packs by defining the `.mtimes` format,\nand introducing a basic API that callers can use to read out individual\nmtimes.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/technical/pack-format.txt |  22 ++++\n Makefile                                |   1 +\n builtin/repack.c                        |   1 +\n object-store.h                          |   5 +-\n pack-mtimes.c                           | 139 ++++++++++++++++++++++++\n pack-mtimes.h                           |  16 +++\n packfile.c                              |  18 ++-\n packfile.h                              |   1 +\n 8 files changed, 200 insertions(+), 3 deletions(-)\n create mode 100644 pack-mtimes.c\n create mode 100644 pack-mtimes.h\n\ndiff --git a/Documentation/technical/pack-format.txt b/Documentation/technical/pack-format.txt\nindex 8d2f42f29e..61d8d960e7 100644\n--- a/Documentation/technical/pack-format.txt\n+++ b/Documentation/technical/pack-format.txt\n@@ -294,6 +294,28 @@ Pack file entry: <+\n \n All 4-byte numbers are in network order.\n \n+== pack-*.mtimes files have the format:\n+\n+  - A 4-byte magic number '0x4d544d45' ('MTME').\n+\n+  - A 4-byte version identifier (= 1).\n+\n+  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n+\n+  - A table of mtimes (one per packed object, num_objects in total, each\n+    a 4-byte unsigned integer in network order), in the same order as\n+    objects appear in the index file (e.g., the first entry in the mtime\n+    table corresponds to the object with the lowest lexically-sorted\n+    oid). The mtimes count standard epoch seconds.\n+\n+  - A trailer, containing a:\n+\n+    checksum of the corresponding packfile, and\n+\n+    a checksum of all of the above.\n+\n+All 4-byte numbers are in network order.\n+\n == multi-pack-index (MIDX) files have the following format:\n \n The multi-pack-index files refer to multiple pack-files and loose objects.\ndiff --git a/Makefile b/Makefile\nindex 12be39ac49..efd5e00717 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -949,6 +949,7 @@ LIB_OBJS += oidtree.o\n LIB_OBJS += pack-bitmap-write.o\n LIB_OBJS += pack-bitmap.o\n LIB_OBJS += pack-check.o\n+LIB_OBJS += pack-mtimes.o\n LIB_OBJS += pack-objects.o\n LIB_OBJS += pack-revindex.o\n LIB_OBJS += pack-write.o\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 0b2d1e5d82..acbb7b8c3b 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -212,6 +212,7 @@ static struct {\n } exts[] = {\n \t{\".pack\"},\n \t{\".rev\", 1},\n+\t{\".mtimes\", 1},\n \t{\".bitmap\", 1},\n \t{\".promisor\", 1},\n \t{\".idx\"},\ndiff --git a/object-store.h b/object-store.h\nindex 952efb6a4b..d87481f101 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -89,12 +89,15 @@ struct packed_git {\n \t\t freshened:1,\n \t\t do_not_close:1,\n \t\t pack_promisor:1,\n-\t\t multi_pack_index:1;\n+\t\t multi_pack_index:1,\n+\t\t is_cruft:1;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n \tstruct revindex_entry *revindex;\n \tconst uint32_t *revindex_data;\n \tconst uint32_t *revindex_map;\n \tsize_t revindex_size;\n+\tconst uint32_t *mtimes_map;\n+\tsize_t mtimes_size;\n \t/* something like \".git/objects/pack/xxxxx.pack\" */\n \tchar pack_name[FLEX_ARRAY]; /* more */\n };\ndiff --git a/pack-mtimes.c b/pack-mtimes.c\nnew file mode 100644\nindex 0000000000..4c7c00fa67\n--- /dev/null\n+++ b/pack-mtimes.c\n@@ -0,0 +1,139 @@\n+#include \"pack-mtimes.h\"\n+#include \"object-store.h\"\n+#include \"packfile.h\"\n+\n+static char *pack_mtimes_filename(struct packed_git *p)\n+{\n+\tsize_t len;\n+\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n+\t\tBUG(\"pack_name does not end in .pack\");\n+\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n+\treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n+}\n+\n+int pack_has_mtimes(struct packed_git *p)\n+{\n+\tstruct stat st;\n+\tchar *fname = pack_mtimes_filename(p);\n+\n+\tif (stat(fname, &st) < 0) {\n+\t\tif (errno == ENOENT)\n+\t\t\treturn 0;\n+\t\tdie_errno(_(\"could not stat %s\"), fname);\n+\t}\n+\n+\tfree(fname);\n+\treturn 1;\n+}\n+\n+#define MTIMES_HEADER_SIZE (12)\n+#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n+\n+struct mtimes_header {\n+\tuint32_t signature;\n+\tuint32_t version;\n+\tuint32_t hash_id;\n+};\n+\n+static int load_pack_mtimes_file(char *mtimes_file,\n+\t\t\t\t uint32_t num_objects,\n+\t\t\t\t const uint32_t **data_p, size_t *len_p)\n+{\n+\tint fd, ret = 0;\n+\tstruct stat st;\n+\tvoid *data = NULL;\n+\tsize_t mtimes_size;\n+\tuint32_t *hdr;\n+\n+\tfd = git_open(mtimes_file);\n+\n+\tif (fd < 0) {\n+\t\tret = -1;\n+\t\tgoto cleanup;\n+\t}\n+\tif (fstat(fd, &st)) {\n+\t\tret = error_errno(_(\"failed to read %s\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tmtimes_size = xsize_t(st.st_size);\n+\n+\tif (mtimes_size < MTIMES_MIN_SIZE) {\n+\t\tret = error(_(\"mtimes file %s is too small\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t), num_objects)) {\n+\t\tret = error(_(\"mtimes file %s is corrupt\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tdata = hdr = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\n+\tif (ntohl(*hdr) != MTIMES_SIGNATURE) {\n+\t\tret = error(_(\"mtimes file %s has unknown signature\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (ntohl(*++hdr) != 1) {\n+\t\tret = error(_(\"mtimes file %s has unsupported version %\"PRIu32),\n+\t\t\t    mtimes_file, ntohl(*hdr));\n+\t\tgoto cleanup;\n+\t}\n+\thdr++;\n+\tif (!(ntohl(*hdr) == 1 || ntohl(*hdr) == 2)) {\n+\t\tret = error(_(\"mtimes file %s has unsupported hash id %\"PRIu32),\n+\t\t\t    mtimes_file, ntohl(*hdr));\n+\t\tgoto cleanup;\n+\t}\n+\n+cleanup:\n+\tif (ret) {\n+\t\tif (data)\n+\t\t\tmunmap(data, mtimes_size);\n+\t} else {\n+\t\t*len_p = mtimes_size;\n+\t\t*data_p = (const uint32_t *)data;\n+\t}\n+\n+\tclose(fd);\n+\treturn ret;\n+}\n+\n+int load_pack_mtimes(struct packed_git *p)\n+{\n+\tchar *mtimes_name = NULL;\n+\tint ret = 0;\n+\n+\tif (!p->is_cruft)\n+\t\treturn ret; /* not a cruft pack */\n+\tif (p->mtimes_map)\n+\t\treturn ret; /* already loaded */\n+\n+\tret = open_pack_index(p);\n+\tif (ret < 0)\n+\t\tgoto cleanup;\n+\n+\tmtimes_name = pack_mtimes_filename(p);\n+\tret = load_pack_mtimes_file(mtimes_name,\n+\t\t\t\t    p->num_objects,\n+\t\t\t\t    &p->mtimes_map,\n+\t\t\t\t    &p->mtimes_size);\n+\tif (ret)\n+\t\tgoto cleanup;\n+\n+cleanup:\n+\tfree(mtimes_name);\n+\treturn ret;\n+}\n+\n+uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos)\n+{\n+\tif (!p->mtimes_map)\n+\t\tBUG(\"pack .mtimes file not loaded for %s\", p->pack_name);\n+\tif (p->num_objects <= pos)\n+\t\tBUG(\"pack .mtimes out-of-bounds (%\"PRIu32\" vs %\"PRIu32\")\",\n+\t\t    pos, p->num_objects);\n+\n+\treturn get_be32(p->mtimes_map + pos + 3);\n+}\ndiff --git a/pack-mtimes.h b/pack-mtimes.h\nnew file mode 100644\nindex 0000000000..ac4247bb5e\n--- /dev/null\n+++ b/pack-mtimes.h\n@@ -0,0 +1,16 @@\n+#ifndef PACK_MTIMES_H\n+#define PACK_MTIMES_H\n+\n+#include \"git-compat-util.h\"\n+\n+#define MTIMES_SIGNATURE 0x4d544d45 /* \"MTME\" */\n+#define MTIMES_VERSION 1\n+\n+struct packed_git;\n+\n+int pack_has_mtimes(struct packed_git *p);\n+int load_pack_mtimes(struct packed_git *p);\n+\n+uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos);\n+\n+#endif\ndiff --git a/packfile.c b/packfile.c\nindex 89402cfc69..ae79ac644e 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -333,12 +333,21 @@ void close_pack_revindex(struct packed_git *p) {\n \tp->revindex_data = NULL;\n }\n \n+void close_pack_mtimes(struct packed_git *p) {\n+\tif (!p->mtimes_map)\n+\t\treturn;\n+\n+\tmunmap((void *)p->mtimes_map, p->mtimes_size);\n+\tp->mtimes_map = NULL;\n+}\n+\n void close_pack(struct packed_git *p)\n {\n \tclose_pack_windows(p);\n \tclose_pack_fd(p);\n \tclose_pack_index(p);\n \tclose_pack_revindex(p);\n+\tclose_pack_mtimes(p);\n \toidset_clear(&p->bad_objects);\n }\n \n@@ -362,7 +371,7 @@ void close_object_store(struct raw_object_store *o)\n \n void unlink_pack_path(const char *pack_name, int force_delete)\n {\n-\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\"};\n+\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\", \".mtimes\"};\n \tint i;\n \tstruct strbuf buf = STRBUF_INIT;\n \tsize_t plen;\n@@ -717,6 +726,10 @@ struct packed_git *add_packed_git(const char *path, size_t path_len, int local)\n \tif (!access(p->pack_name, F_OK))\n \t\tp->pack_promisor = 1;\n \n+\txsnprintf(p->pack_name + path_len, alloc - path_len, \".mtimes\");\n+\tif (!access(p->pack_name, F_OK))\n+\t\tp->is_cruft = 1;\n+\n \txsnprintf(p->pack_name + path_len, alloc - path_len, \".pack\");\n \tif (stat(p->pack_name, &st) || !S_ISREG(st.st_mode)) {\n \t\tfree(p);\n@@ -868,7 +881,8 @@ static void prepare_pack(const char *full_name, size_t full_name_len,\n \t    ends_with(file_name, \".pack\") ||\n \t    ends_with(file_name, \".bitmap\") ||\n \t    ends_with(file_name, \".keep\") ||\n-\t    ends_with(file_name, \".promisor\"))\n+\t    ends_with(file_name, \".promisor\") ||\n+\t    ends_with(file_name, \".mtimes\"))\n \t\tstring_list_append(data->garbage, full_name);\n \telse\n \t\treport_garbage(PACKDIR_FILE_GARBAGE, full_name);\ndiff --git a/packfile.h b/packfile.h\nindex 186146779d..32201d8af7 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -91,6 +91,7 @@ uint32_t get_pack_fanout(struct packed_git *p, uint32_t value);\n unsigned char *use_pack(struct packed_git *, struct pack_window **, off_t, unsigned long *);\n void close_pack_windows(struct packed_git *);\n void close_pack_revindex(struct packed_git *);\n+void close_pack_mtimes(struct packed_git *p);\n void close_pack(struct packed_git *);\n void close_object_store(struct raw_object_store *o);\n void unuse_pack(struct pack_window **);\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442518","messageId":"ea245b7216067093fdd3a5b2e3a9390f634c8af0.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 04/17] chunk-format.h: extract oid_version()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:13Z","receivedAt":"2021-11-29T22:26:01Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"There are three definitions of an identical function which converts\n`the_hash_algo` into either 1 (for SHA-1) or 2 (for SHA-256). There is a\ncopy of this function for writing both the commit-graph and\nmulti-pack-index file, and another inline definition used to write the\n.rev header.\n\nConsolidate these into a single definition in chunk-format.h. It's not\nclear that this is the best header to define this function in, but it\nshould do for now.\n\n(Worth noting, the .rev caller expects a 4-byte unsigned, but the other\ntwo callers work with a single unsigned byte. The consolidated version\nuses the latter type, and lets the compiler widen it when required).\n\nAnother caller will be added in a subsequent patch.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n chunk-format.c | 12 ++++++++++++\n chunk-format.h |  3 +++\n commit-graph.c | 18 +++---------------\n midx.c         | 18 +++---------------\n pack-write.c   | 15 ++-------------\n 5 files changed, 23 insertions(+), 43 deletions(-)\n\ndiff --git a/chunk-format.c b/chunk-format.c\nindex 1c3dca62e2..0275b74a89 100644\n--- a/chunk-format.c\n+++ b/chunk-format.c\n@@ -181,3 +181,15 @@ int read_chunk(struct chunkfile *cf,\n \n \treturn CHUNK_NOT_FOUND;\n }\n+\n+uint8_t oid_version(const struct git_hash_algo *algop)\n+{\n+\tswitch (hash_algo_by_ptr(algop)) {\n+\tcase GIT_HASH_SHA1:\n+\t\treturn 1;\n+\tcase GIT_HASH_SHA256:\n+\t\treturn 2;\n+\tdefault:\n+\t\tdie(_(\"invalid hash version\"));\n+\t}\n+}\ndiff --git a/chunk-format.h b/chunk-format.h\nindex 9ccbe00377..7885aa0848 100644\n--- a/chunk-format.h\n+++ b/chunk-format.h\n@@ -2,6 +2,7 @@\n #define CHUNK_FORMAT_H\n \n #include \"git-compat-util.h\"\n+#include \"hash.h\"\n \n struct hashfile;\n struct chunkfile;\n@@ -65,4 +66,6 @@ int read_chunk(struct chunkfile *cf,\n \t       chunk_read_fn fn,\n \t       void *data);\n \n+uint8_t oid_version(const struct git_hash_algo *algop);\n+\n #endif\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 2706683acf..1f08152a35 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -193,18 +193,6 @@ char *get_commit_graph_chain_filename(struct object_directory *odb)\n \treturn xstrfmt(\"%s/info/commit-graphs/commit-graph-chain\", odb->path);\n }\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n static struct commit_graph *alloc_commit_graph(void)\n {\n \tstruct commit_graph *g = xcalloc(1, sizeof(*g));\n@@ -365,9 +353,9 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n \t}\n \n \thash_version = *(unsigned char*)(data + 5);\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != oid_version(the_hash_algo)) {\n \t\terror(_(\"commit-graph hash version %X does not match version %X\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, oid_version(the_hash_algo));\n \t\treturn NULL;\n \t}\n \n@@ -1908,7 +1896,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \thashwrite_be32(f, GRAPH_SIGNATURE);\n \n \thashwrite_u8(f, GRAPH_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, oid_version(the_hash_algo));\n \thashwrite_u8(f, get_num_chunks(cf));\n \thashwrite_u8(f, ctx->num_commit_graphs_after - 1);\n \ndiff --git a/midx.c b/midx.c\nindex 8433086ac1..756ae6a206 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -40,18 +40,6 @@\n \n #define PACK_EXPIRED UINT_MAX\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n const unsigned char *get_midx_checksum(struct multi_pack_index *m)\n {\n \treturn m->data + m->data_len - the_hash_algo->rawsz;\n@@ -131,9 +119,9 @@ struct multi_pack_index *load_multi_pack_index(const char *object_dir, int local\n \t\t      m->version);\n \n \thash_version = m->data[MIDX_BYTE_HASH_VERSION];\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != oid_version(the_hash_algo)) {\n \t\terror(_(\"multi-pack-index hash version %u does not match version %u\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, oid_version(the_hash_algo));\n \t\tgoto cleanup_fail;\n \t}\n \tm->hash_len = the_hash_algo->rawsz;\n@@ -413,7 +401,7 @@ static size_t write_midx_header(struct hashfile *f,\n {\n \thashwrite_be32(f, MIDX_SIGNATURE);\n \thashwrite_u8(f, MIDX_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, oid_version(the_hash_algo));\n \thashwrite_u8(f, num_chunks);\n \thashwrite_u8(f, 0); /* unused */\n \thashwrite_be32(f, num_packs);\ndiff --git a/pack-write.c b/pack-write.c\nindex d594e3008e..ff305b404c 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -2,6 +2,7 @@\n #include \"pack.h\"\n #include \"csum-file.h\"\n #include \"remote.h\"\n+#include \"chunk-format.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -181,21 +182,9 @@ static int pack_order_cmp(const void *va, const void *vb, void *ctx)\n \n static void write_rev_header(struct hashfile *f)\n {\n-\tuint32_t oid_version;\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\toid_version = 1;\n-\t\tbreak;\n-\tcase GIT_HASH_SHA256:\n-\t\toid_version = 2;\n-\t\tbreak;\n-\tdefault:\n-\t\tdie(\"write_rev_header: unknown hash version\");\n-\t}\n-\n \thashwrite_be32(f, RIDX_SIGNATURE);\n \thashwrite_be32(f, RIDX_VERSION);\n-\thashwrite_be32(f, oid_version);\n+\thashwrite_be32(f, oid_version(the_hash_algo));\n }\n \n static void write_rev_index_positions(struct hashfile *f,\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442519","messageId":"7f4612e859dd9923014c4e8a28bf5caea84d971a.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 03/17] pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:10Z","receivedAt":"2021-11-29T22:26:03Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This structure will be used to communicate the per-object mtimes when\nwriting a cruft pack. Here, we need the full packing_data structure\nbecause the mtime information is stored in an array there, not on the\nindividual object_entry's themselves (to avoid paying the overhead in\nstructure width for operations which do not generate a cruft pack).\n\nWe haven't passed this information down before because one of the two\ncallers (in bulk-checkin.c) does not have a packing_data structure at\nall. In that case (where no cruft pack will be generated), NULL is\npassed instead.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 3 ++-\n bulk-checkin.c         | 2 +-\n pack-write.c           | 1 +\n pack.h                 | 3 +++\n 4 files changed, 7 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 1a3dd445f8..bf45ffbc57 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1254,7 +1254,8 @@ static void write_pack_file(void)\n \n \t\t\tstage_tmp_packfiles(&tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n-\t\t\t\t\t    &pack_idx_opts, hash, &idx_tmp_name);\n+\t\t\t\t\t    &to_pack, &pack_idx_opts, hash,\n+\t\t\t\t\t    &idx_tmp_name);\n \n \t\t\tif (write_bitmap_index) {\n \t\t\t\tsize_t tmpname_len = tmpname.len;\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac80..99f7596c4e 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -33,7 +33,7 @@ static void finish_tmp_packfile(struct strbuf *basename,\n \tchar *idx_tmp_name = NULL;\n \n \tstage_tmp_packfiles(basename, pack_tmp_name, written_list, nr_written,\n-\t\t\t    pack_idx_opts, hash, &idx_tmp_name);\n+\t\t\t    NULL, pack_idx_opts, hash, &idx_tmp_name);\n \trename_tmp_packfile_idx(basename, &idx_tmp_name);\n \n \tfree(idx_tmp_name);\ndiff --git a/pack-write.c b/pack-write.c\nindex a5846f3a34..d594e3008e 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -483,6 +483,7 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t const char *pack_tmp_name,\n \t\t\t struct pack_idx_entry **written_list,\n \t\t\t uint32_t nr_written,\n+\t\t\t struct packing_data *to_pack,\n \t\t\t struct pack_idx_option *pack_idx_opts,\n \t\t\t unsigned char hash[],\n \t\t\t char **idx_tmp_name)\ndiff --git a/pack.h b/pack.h\nindex b22bfc4a18..fd27cfdfd7 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -109,11 +109,14 @@ int encode_in_pack_object_header(unsigned char *hdr, int hdr_len,\n #define PH_ERROR_PROTOCOL\t(-3)\n int read_pack_header(int fd, struct pack_header *);\n \n+struct packing_data;\n+\n struct hashfile *create_tmp_packfile(char **pack_tmp_name);\n void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t const char *pack_tmp_name,\n \t\t\t struct pack_idx_entry **written_list,\n \t\t\t uint32_t nr_written,\n+\t\t\t struct packing_data *to_pack,\n \t\t\t struct pack_idx_option *pack_idx_opts,\n \t\t\t unsigned char hash[],\n \t\t\t char **idx_tmp_name);\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442520","messageId":"deece9eb70e9750bb8350946679b521e59139fe2.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 05/17] pack-mtimes: support writing pack .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:15Z","receivedAt":"2021-11-29T22:26:04Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Now that the `.mtimes` format is defined, supplement the pack-write API\nto be able to conditionally write an `.mtimes` file along with a pack by\nsetting an additional flag and passing an oidmap that contains the\ntimestamps corresponding to each object in the pack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n pack-objects.c |  6 ++++\n pack-objects.h | 20 ++++++++++++++\n pack-write.c   | 74 ++++++++++++++++++++++++++++++++++++++++++++++++++\n pack.h         |  1 +\n 4 files changed, 101 insertions(+)\n\ndiff --git a/pack-objects.c b/pack-objects.c\nindex fe2a4eace9..272e8d4517 100644\n--- a/pack-objects.c\n+++ b/pack-objects.c\n@@ -170,6 +170,9 @@ struct object_entry *packlist_alloc(struct packing_data *pdata,\n \n \t\tif (pdata->layer)\n \t\t\tREALLOC_ARRAY(pdata->layer, pdata->nr_alloc);\n+\n+\t\tif (pdata->cruft_mtime)\n+\t\t\tREALLOC_ARRAY(pdata->cruft_mtime, pdata->nr_alloc);\n \t}\n \n \tnew_entry = pdata->objects + pdata->nr_objects++;\n@@ -198,6 +201,9 @@ struct object_entry *packlist_alloc(struct packing_data *pdata,\n \tif (pdata->layer)\n \t\tpdata->layer[pdata->nr_objects - 1] = 0;\n \n+\tif (pdata->cruft_mtime)\n+\t\tpdata->cruft_mtime[pdata->nr_objects - 1] = 0;\n+\n \treturn new_entry;\n }\n \ndiff --git a/pack-objects.h b/pack-objects.h\nindex dca2351ef9..f17119de26 100644\n--- a/pack-objects.h\n+++ b/pack-objects.h\n@@ -168,6 +168,9 @@ struct packing_data {\n \t/* delta islands */\n \tunsigned int *tree_depth;\n \tunsigned char *layer;\n+\n+\t/* cruft packs */\n+\tuint32_t *cruft_mtime;\n };\n \n void prepare_packing_data(struct repository *r, struct packing_data *pdata);\n@@ -289,4 +292,21 @@ static inline void oe_set_layer(struct packing_data *pack,\n \tpack->layer[e - pack->objects] = layer;\n }\n \n+static inline uint32_t oe_cruft_mtime(struct packing_data *pack,\n+\t\t\t\t      struct object_entry *e)\n+{\n+\tif (!pack->cruft_mtime)\n+\t\treturn 0;\n+\treturn pack->cruft_mtime[e - pack->objects];\n+}\n+\n+static inline void oe_set_cruft_mtime(struct packing_data *pack,\n+\t\t\t\t      struct object_entry *e,\n+\t\t\t\t      uint32_t mtime)\n+{\n+\tif (!pack->cruft_mtime)\n+\t\tCALLOC_ARRAY(pack->cruft_mtime, pack->nr_alloc);\n+\tpack->cruft_mtime[e - pack->objects] = mtime;\n+}\n+\n #endif\ndiff --git a/pack-write.c b/pack-write.c\nindex ff305b404c..8c3efda2c3 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -3,6 +3,10 @@\n #include \"csum-file.h\"\n #include \"remote.h\"\n #include \"chunk-format.h\"\n+#include \"pack-mtimes.h\"\n+#include \"oidmap.h\"\n+#include \"chunk-format.h\"\n+#include \"pack-objects.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -276,6 +280,65 @@ const char *write_rev_file_order(const char *rev_name,\n \treturn rev_name;\n }\n \n+static void write_mtimes_header(struct hashfile *f)\n+{\n+\thashwrite_be32(f, MTIMES_SIGNATURE);\n+\thashwrite_be32(f, MTIMES_VERSION);\n+\thashwrite_be32(f, oid_version(the_hash_algo));\n+}\n+\n+static void write_mtimes_objects(struct hashfile *f,\n+\t\t\t\t struct packing_data *to_pack,\n+\t\t\t\t struct pack_idx_entry **objects,\n+\t\t\t\t uint32_t nr_objects)\n+{\n+\tuint32_t i;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\tstruct object_entry *e = (struct object_entry*)objects[i];\n+\t\thashwrite_be32(f, oe_cruft_mtime(to_pack, e));\n+\t}\n+}\n+\n+static void write_mtimes_trailer(struct hashfile *f, const unsigned char *hash)\n+{\n+\thashwrite(f, hash, the_hash_algo->rawsz);\n+}\n+\n+static const char *write_mtimes_file(const char *mtimes_name,\n+\t\t\t\t     struct packing_data *to_pack,\n+\t\t\t\t     struct pack_idx_entry **objects,\n+\t\t\t\t     uint32_t nr_objects,\n+\t\t\t\t     const unsigned char *hash)\n+{\n+\tstruct hashfile *f;\n+\tint fd;\n+\n+\tif (!to_pack)\n+\t\tBUG(\"cannot call write_mtimes_file with NULL packing_data\");\n+\n+\tif (!mtimes_name) {\n+\t\tstruct strbuf tmp_file = STRBUF_INIT;\n+\t\tfd = odb_mkstemp(&tmp_file, \"pack/tmp_mtimes_XXXXXX\");\n+\t\tmtimes_name = strbuf_detach(&tmp_file, NULL);\n+\t} else {\n+\t\tunlink(mtimes_name);\n+\t\tfd = xopen(mtimes_name, O_CREAT|O_EXCL|O_WRONLY, 0600);\n+\t}\n+\tf = hashfd(fd, mtimes_name);\n+\n+\twrite_mtimes_header(f);\n+\twrite_mtimes_objects(f, to_pack, objects, nr_objects);\n+\twrite_mtimes_trailer(f, hash);\n+\n+\tif (mtimes_name && adjust_shared_perm(mtimes_name) < 0)\n+\t\tdie(_(\"failed to make %s readable\"), mtimes_name);\n+\n+\tfinalize_hashfile(f, NULL,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE | CSUM_FSYNC);\n+\n+\treturn mtimes_name;\n+}\n+\n off_t write_pack_header(struct hashfile *f, uint32_t nr_entries)\n {\n \tstruct pack_header hdr;\n@@ -478,6 +541,7 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t char **idx_tmp_name)\n {\n \tconst char *rev_tmp_name = NULL;\n+\tconst char *mtimes_tmp_name = NULL;\n \n \tif (adjust_shared_perm(pack_tmp_name))\n \t\tdie_errno(\"unable to make temporary pack file readable\");\n@@ -490,9 +554,19 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \trev_tmp_name = write_rev_file(NULL, written_list, nr_written, hash,\n \t\t\t\t      pack_idx_opts->flags);\n \n+\tif (pack_idx_opts->flags & WRITE_MTIMES) {\n+\t\tmtimes_tmp_name = write_mtimes_file(NULL, to_pack, written_list,\n+\t\t\t\t\t\t    nr_written,\n+\t\t\t\t\t\t    hash);\n+\t\tif (adjust_shared_perm(mtimes_tmp_name))\n+\t\t\tdie_errno(\"unable to make temporary mtimes file readable\");\n+\t}\n+\n \trename_tmp_packfile(name_buffer, pack_tmp_name, \"pack\");\n \tif (rev_tmp_name)\n \t\trename_tmp_packfile(name_buffer, rev_tmp_name, \"rev\");\n+\tif (mtimes_tmp_name)\n+\t\trename_tmp_packfile(name_buffer, mtimes_tmp_name, \"mtimes\");\n }\n \n void write_promisor_file(const char *promisor_name, struct ref **sought, int nr_sought)\ndiff --git a/pack.h b/pack.h\nindex fd27cfdfd7..01d385903a 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -44,6 +44,7 @@ struct pack_idx_option {\n #define WRITE_IDX_STRICT 02\n #define WRITE_REV 04\n #define WRITE_REV_VERIFY 010\n+#define WRITE_MTIMES 020\n \n \tuint32_t version;\n \tuint32_t off32_limit;\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442521","messageId":"02f7fce788c1ecf9b2804329cae60984b0478853.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 09/17] reachable: add options to add_unseen_recent_objects_to_traversal","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:25Z","receivedAt":"2021-11-29T22:26:11Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This function behaves very similarly to what we will need in\npack-objects in order to implement cruft packs with expiration. But it\nis lacking a couple of things. Namely, it needs:\n\n  - a mechanism to communicate the timestamps of individual recent\n    objects to some external caller\n\n  - and, in the case of packed objects, our future caller will also want\n    to know the originating pack, as well as the offset within that pack\n    at which the object can be found\n\n  - finally, it needs a way to skip over packs which are marked as kept\n    in-core.\n\nTo address the first two, add a callback interface in this patch which\nreports the time of each recent object, as well as a (packed_git,\noff_t) pair for packed objects.\n\nLikewise, add a new option to the packed object iterators to skip over\npacks which are marked as kept in core. This option will become\nimplicitly tested in a future patch.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c |  2 +-\n reachable.c            | 51 +++++++++++++++++++++++++++++++++++-------\n reachable.h            |  9 +++++++-\n 3 files changed, 52 insertions(+), 10 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex b12e79e4b1..2c592d369a 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -3951,7 +3951,7 @@ static void get_object_list(int ac, const char **av)\n \tif (unpack_unreachable_expiration) {\n \t\trevs.ignore_missing_links = 1;\n \t\tif (add_unseen_recent_objects_to_traversal(&revs,\n-\t\t\t\tunpack_unreachable_expiration))\n+\t\t\t\tunpack_unreachable_expiration, NULL, 0))\n \t\t\tdie(_(\"unable to add recent objects\"));\n \t\tif (prepare_revision_walk(&revs))\n \t\t\tdie(_(\"revision walk setup failed\"));\ndiff --git a/reachable.c b/reachable.c\nindex 84e3d0d75e..0eb9909f47 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -60,9 +60,13 @@ static void mark_commit(struct commit *c, void *data)\n struct recent_data {\n \tstruct rev_info *revs;\n \ttimestamp_t timestamp;\n+\treport_recent_object_fn *cb;\n+\tint ignore_in_core_kept_packs;\n };\n \n static void add_recent_object(const struct object_id *oid,\n+\t\t\t      struct packed_git *pack,\n+\t\t\t      off_t offset,\n \t\t\t      timestamp_t mtime,\n \t\t\t      struct recent_data *data)\n {\n@@ -103,13 +107,29 @@ static void add_recent_object(const struct object_id *oid,\n \t\tdie(\"unable to lookup %s\", oid_to_hex(oid));\n \n \tadd_pending_object(data->revs, obj, \"\");\n+\tif (data->cb)\n+\t\tdata->cb(obj, pack, offset, mtime);\n+}\n+\n+static int want_recent_object(struct recent_data *data,\n+\t\t\t      const struct object_id *oid)\n+{\n+\tif (data->ignore_in_core_kept_packs &&\n+\t    has_object_kept_pack(oid, IN_CORE_KEEP_PACKS))\n+\t\treturn 0;\n+\treturn 1;\n }\n \n static int add_recent_loose(const struct object_id *oid,\n \t\t\t    const char *path, void *data)\n {\n \tstruct stat st;\n-\tstruct object *obj = lookup_object(the_repository, oid);\n+\tstruct object *obj;\n+\n+\tif (!want_recent_object(data, oid))\n+\t\treturn 0;\n+\n+\tobj = lookup_object(the_repository, oid);\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n@@ -126,7 +146,7 @@ static int add_recent_loose(const struct object_id *oid,\n \t\treturn error_errno(\"unable to stat %s\", oid_to_hex(oid));\n \t}\n \n-\tadd_recent_object(oid, st.st_mtime, data);\n+\tadd_recent_object(oid, NULL, 0, st.st_mtime, data);\n \treturn 0;\n }\n \n@@ -134,29 +154,43 @@ static int add_recent_packed(const struct object_id *oid,\n \t\t\t     struct packed_git *p, uint32_t pos,\n \t\t\t     void *data)\n {\n-\tstruct object *obj = lookup_object(the_repository, oid);\n+\tstruct object *obj;\n+\n+\tif (!want_recent_object(data, oid))\n+\t\treturn 0;\n+\n+\tobj = lookup_object(the_repository, oid);\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n-\tadd_recent_object(oid, p->mtime, data);\n+\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), p->mtime, data);\n \treturn 0;\n }\n \n int add_unseen_recent_objects_to_traversal(struct rev_info *revs,\n-\t\t\t\t\t   timestamp_t timestamp)\n+\t\t\t\t\t   timestamp_t timestamp,\n+\t\t\t\t\t   report_recent_object_fn *cb,\n+\t\t\t\t\t   int ignore_in_core_kept_packs)\n {\n \tstruct recent_data data;\n+\tenum for_each_object_flags flags;\n \tint r;\n \n \tdata.revs = revs;\n \tdata.timestamp = timestamp;\n+\tdata.cb = cb;\n+\tdata.ignore_in_core_kept_packs = ignore_in_core_kept_packs;\n \n \tr = for_each_loose_object(add_recent_loose, &data,\n \t\t\t\t  FOR_EACH_OBJECT_LOCAL_ONLY);\n \tif (r)\n \t\treturn r;\n-\treturn for_each_packed_object(add_recent_packed, &data,\n-\t\t\t\t      FOR_EACH_OBJECT_LOCAL_ONLY);\n+\n+\tflags = FOR_EACH_OBJECT_LOCAL_ONLY | FOR_EACH_OBJECT_PACK_ORDER;\n+\tif (ignore_in_core_kept_packs)\n+\t\tflags |= FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS;\n+\n+\treturn for_each_packed_object(add_recent_packed, &data, flags);\n }\n \n static int mark_object_seen(const struct object_id *oid,\n@@ -217,7 +251,8 @@ void mark_reachable_objects(struct rev_info *revs, int mark_reflog,\n \n \tif (mark_recent) {\n \t\trevs->ignore_missing_links = 1;\n-\t\tif (add_unseen_recent_objects_to_traversal(revs, mark_recent))\n+\t\tif (add_unseen_recent_objects_to_traversal(revs, mark_recent,\n+\t\t\t\t\t\t\t   NULL, 0))\n \t\t\tdie(\"unable to mark recent objects\");\n \t\tif (prepare_revision_walk(revs))\n \t\t\tdie(\"revision walk setup failed\");\ndiff --git a/reachable.h b/reachable.h\nindex 5df932ad8f..b776761baa 100644\n--- a/reachable.h\n+++ b/reachable.h\n@@ -1,11 +1,18 @@\n #ifndef REACHEABLE_H\n #define REACHEABLE_H\n \n+#include \"object.h\"\n+\n struct progress;\n struct rev_info;\n \n+typedef void report_recent_object_fn(const struct object *, struct packed_git *,\n+\t\t\t\t     off_t, time_t);\n+\n int add_unseen_recent_objects_to_traversal(struct rev_info *revs,\n-\t\t\t\t\t   timestamp_t timestamp);\n+\t\t\t\t\t   timestamp_t timestamp,\n+\t\t\t\t\t   report_recent_object_fn cb,\n+\t\t\t\t\t   int ignore_in_core_kept_packs);\n void mark_reachable_objects(struct rev_info *revs, int mark_reflog,\n \t\t\t    timestamp_t mark_recent, struct progress *);\n \n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442522","messageId":"0d2dfaa0620f6e29133aec6e3b87176fb2336ab2.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 13/17] builtin/repack.c: allow configuring cruft pack generation","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:35Z","receivedAt":"2021-11-29T22:26:15Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In servers which set the pack.window configuration to a large value, we\ncan wind up spending quite a lot of time finding new bases when breaking\ndelta chains between reachable and unreachable objects while generating\na cruft pack.\n\nIntroduce a handful of `repack.cruft*` configuration variables to\ncontrol the parameters used by pack-objects when generating a cruft\npack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/repack.txt |  9 ++++\n builtin/repack.c                | 50 ++++++++++++++------\n t/t5327-pack-objects-cruft.sh   | 83 +++++++++++++++++++++++++++++++++\n 3 files changed, 128 insertions(+), 14 deletions(-)\n\ndiff --git a/Documentation/config/repack.txt b/Documentation/config/repack.txt\nindex 9c413e177e..fd18d1fb89 100644\n--- a/Documentation/config/repack.txt\n+++ b/Documentation/config/repack.txt\n@@ -25,3 +25,12 @@ repack.writeBitmaps::\n \tspace and extra time spent on the initial repack.  This has\n \tno effect if multiple packfiles are created.\n \tDefaults to true on bare repos, false otherwise.\n+\n+repack.cruftWindow::\n+repack.cruftWindowMemory::\n+repack.cruftDepth::\n+repack.cruftThreads::\n+\tParameters used by linkgit:git-pack-objects[1] when generating\n+\ta cruft pack and the respective parameters are not given over\n+\tthe command line. See similarly named `pack.*` configuration\n+\tvariables for defaults and meaning.\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 68b4bdf06f..cefa906344 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -40,9 +40,21 @@ static const char incremental_bitmap_conflict_error[] = N_(\n \"--no-write-bitmap-index or disable the pack.writebitmaps configuration.\"\n );\n \n+struct pack_objects_args {\n+\tconst char *window;\n+\tconst char *window_memory;\n+\tconst char *depth;\n+\tconst char *threads;\n+\tconst char *max_pack_size;\n+\tint no_reuse_delta;\n+\tint no_reuse_object;\n+\tint quiet;\n+\tint local;\n+};\n \n static int repack_config(const char *var, const char *value, void *cb)\n {\n+\tstruct pack_objects_args *cruft_po_args = cb;\n \tif (!strcmp(var, \"repack.usedeltabaseoffset\")) {\n \t\tdelta_base_offset = git_config_bool(var, value);\n \t\treturn 0;\n@@ -61,6 +73,15 @@ static int repack_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"repack.cruftwindow\"))\n+\t\treturn git_config_string(&cruft_po_args->window, var, value);\n+\tif (!strcmp(var, \"repack.cruftwindowmemory\"))\n+\t\treturn git_config_string(&cruft_po_args->window_memory, var, value);\n+\tif (!strcmp(var, \"repack.cruftdepth\"))\n+\t\treturn git_config_string(&cruft_po_args->depth, var, value);\n+\tif (!strcmp(var, \"repack.cruftthreads\"))\n+\t\treturn git_config_string(&cruft_po_args->threads, var, value);\n+\n \treturn git_default_config(var, value, cb);\n }\n \n@@ -153,18 +174,6 @@ static void remove_redundant_pack(const char *dir_name, const char *base_name)\n \tstrbuf_release(&buf);\n }\n \n-struct pack_objects_args {\n-\tconst char *window;\n-\tconst char *window_memory;\n-\tconst char *depth;\n-\tconst char *threads;\n-\tconst char *max_pack_size;\n-\tint no_reuse_delta;\n-\tint no_reuse_object;\n-\tint quiet;\n-\tint local;\n-};\n-\n static void prepare_pack_objects(struct child_process *cmd,\n \t\t\t\t const struct pack_objects_args *args)\n {\n@@ -687,6 +696,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tint no_update_server_info = 0;\n \tstruct pack_objects_args po_args = {NULL};\n+\tstruct pack_objects_args cruft_po_args = {NULL};\n \tint geometric_factor = 0;\n \tint write_midx = 0;\n \n@@ -741,7 +751,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tOPT_END()\n \t};\n \n-\tgit_config(repack_config, NULL);\n+\tgit_config(repack_config, &cruft_po_args);\n \n \targc = parse_options(argc, argv, prefix, builtin_repack_options,\n \t\t\t\tgit_repack_usage, 0);\n@@ -920,7 +930,19 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tif (*pack_prefix == '/')\n \t\t\tpack_prefix++;\n \n-\t\tret = write_cruft_pack(&po_args, pack_prefix, &names,\n+\t\tif (!cruft_po_args.window)\n+\t\t\tcruft_po_args.window = po_args.window;\n+\t\tif (!cruft_po_args.window_memory)\n+\t\t\tcruft_po_args.window_memory = po_args.window_memory;\n+\t\tif (!cruft_po_args.depth)\n+\t\t\tcruft_po_args.depth = po_args.depth;\n+\t\tif (!cruft_po_args.threads)\n+\t\t\tcruft_po_args.threads = po_args.threads;\n+\n+\t\tcruft_po_args.local = po_args.local;\n+\t\tcruft_po_args.quiet = po_args.quiet;\n+\n+\t\tret = write_cruft_pack(&cruft_po_args, pack_prefix, &names,\n \t\t\t\t       &existing_nonkept_packs,\n \t\t\t\t       &existing_kept_packs);\n \t\tif (ret)\ndiff --git a/t/t5327-pack-objects-cruft.sh b/t/t5327-pack-objects-cruft.sh\nindex ed1a113ab6..750e9d6d6f 100755\n--- a/t/t5327-pack-objects-cruft.sh\n+++ b/t/t5327-pack-objects-cruft.sh\n@@ -511,4 +511,87 @@ test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n \t)\n '\n \n+test_expect_success 'cruft repack respects repack.cruftWindow' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tGIT_TRACE2_EVENT=$(pwd)/event.trace \\\n+\t\tgit -c pack.window=1 -c repack.cruftWindow=2 repack \\\n+\t\t       --cruft --window=3 &&\n+\n+\t\tgrep \"pack-objects.*--window=2.*--cruft\" event.trace\n+\t)\n+'\n+\n+test_expect_success 'cruft repack respects --window by default' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tGIT_TRACE2_EVENT=$(pwd)/event.trace \\\n+\t\tgit -c pack.window=2 repack --cruft --window=3 &&\n+\n+\t\tgrep \"pack-objects.*--window=3.*--cruft\" event.trace\n+\t)\n+'\n+\n+test_expect_success 'cruft repack respects --quiet' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\t\tGIT_PROGRESS_DELAY=0 git repack --cruft --quiet 2>err &&\n+\t\ttest_must_be_empty err\n+\t)\n+'\n+\n+test_expect_success 'cruft --local drops unreachable objects' '\n+\tgit init alternate &&\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr alternate repo\" &&\n+\n+\ttest_commit -C alternate base &&\n+\t# Pack all objects in alterate so that the cruft repack in \"repo\" sees\n+\t# the object it dropped due to `--local` as packed. Otherwise this\n+\t# object would not appear packed anywhere (since it is not packed in\n+\t# alternate and likewise not part of the cruft pack in the other repo\n+\t# because of `--local`).\n+\tgit -C alternate repack -ad &&\n+\n+\t(\n+\t\tcd repo &&\n+\n+\t\tobject=\"$(git -C ../alternate rev-parse HEAD:base.t)\" &&\n+\t\tgit -C ../alternate cat-file -p $object >contents &&\n+\n+\t\t# Write some reachable objects and two unreachable ones: one\n+\t\t# that the alternate has and another that is unique.\n+\t\ttest_commit other &&\n+\t\tgit hash-object -w -t blob contents &&\n+\t\tcruft=\"$(echo cruft | git hash-object -w -t blob --stdin)\" &&\n+\n+\t\t( cd ../alternate/.git/objects && pwd ) \\\n+\t\t       >.git/objects/info/alternates &&\n+\n+\t\ttest_path_is_file $objdir/$(test_oid_to_path $cruft) &&\n+\t\ttest_path_is_file $objdir/$(test_oid_to_path $object) &&\n+\n+\t\tgit repack -d --cruft --local &&\n+\n+\t\ttest-tool pack-mtimes \"$(basename $(ls $packdir/pack-*.mtimes))\" \\\n+\t\t       >objects &&\n+\t\t! grep $object objects &&\n+\t\tgrep $cruft objects\n+\t)\n+'\n+\n test_done\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442523","messageId":"e0a7b3b310c69350d8e2c0561e0991bb7045a66d.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 06/17] t/helper: add 'pack-mtimes' test-tool","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:17Z","receivedAt":"2021-11-29T22:26:16Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In the next patch, we will implement and test support for writing a\ncruft pack via a special mode of `git pack-objects`. To make sure that\nobjects are written with the correct timestamps, and a new test-tool\nthat can dump the object names and corresponding timestamps from a given\n`.mtimes` file.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Makefile                    |  1 +\n t/helper/test-pack-mtimes.c | 53 +++++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c        |  1 +\n t/helper/test-tool.h        |  1 +\n 4 files changed, 56 insertions(+)\n create mode 100644 t/helper/test-pack-mtimes.c\n\ndiff --git a/Makefile b/Makefile\nindex efd5e00717..a7382cbfc1 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -721,6 +721,7 @@ TEST_BUILTINS_OBJS += test-oid-array.o\n TEST_BUILTINS_OBJS += test-oidmap.o\n TEST_BUILTINS_OBJS += test-oidtree.o\n TEST_BUILTINS_OBJS += test-online-cpus.o\n+TEST_BUILTINS_OBJS += test-pack-mtimes.o\n TEST_BUILTINS_OBJS += test-parse-options.o\n TEST_BUILTINS_OBJS += test-parse-pathspec-file.o\n TEST_BUILTINS_OBJS += test-partial-clone.o\ndiff --git a/t/helper/test-pack-mtimes.c b/t/helper/test-pack-mtimes.c\nnew file mode 100644\nindex 0000000000..b143f62520\n--- /dev/null\n+++ b/t/helper/test-pack-mtimes.c\n@@ -0,0 +1,53 @@\n+#include \"git-compat-util.h\"\n+#include \"test-tool.h\"\n+#include \"strbuf.h\"\n+#include \"object-store.h\"\n+#include \"packfile.h\"\n+#include \"pack-mtimes.h\"\n+\n+static int dump_mtimes(struct packed_git *p)\n+{\n+\tuint32_t i;\n+\tif (load_pack_mtimes(p) < 0)\n+\t\tdie(\"could not load pack .mtimes\");\n+\n+\tfor (i = 0; i < p->num_objects; i++) {\n+\t\tstruct object_id oid;\n+\t\tif (nth_packed_object_id(&oid, p, i) < 0)\n+\t\t\tdie(\"could not load object id at position %\"PRIu32, i);\n+\n+\t\tprintf(\"%s %\"PRIu32\"\\n\",\n+\t\t       oid_to_hex(&oid), nth_packed_mtime(p, i));\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static const char *pack_mtimes_usage = \"\\n\"\n+\"  test-tool pack-mtimes <pack-name.mtimes>\";\n+\n+int cmd__pack_mtimes(int argc, const char **argv)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct packed_git *p;\n+\n+\tsetup_git_directory();\n+\n+\tif (argc != 2)\n+\t\tusage(pack_mtimes_usage);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tstrbuf_addstr(&buf, basename(p->pack_name));\n+\t\tstrbuf_strip_suffix(&buf, \".pack\");\n+\t\tstrbuf_addstr(&buf, \".mtimes\");\n+\n+\t\tif (!strcmp(buf.buf, argv[1]))\n+\t\t\tbreak;\n+\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstrbuf_release(&buf);\n+\n+\treturn p ? dump_mtimes(p) : 1;\n+}\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex 3ce5585e53..1bb1c4b562 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -46,6 +46,7 @@ static struct test_cmd cmds[] = {\n \t{ \"oidmap\", cmd__oidmap },\n \t{ \"oidtree\", cmd__oidtree },\n \t{ \"online-cpus\", cmd__online_cpus },\n+\t{ \"pack-mtimes\", cmd__pack_mtimes },\n \t{ \"parse-options\", cmd__parse_options },\n \t{ \"parse-pathspec-file\", cmd__parse_pathspec_file },\n \t{ \"partial-clone\", cmd__partial_clone },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex 9f0f522850..07a2d3f94e 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -35,6 +35,7 @@ int cmd__mktemp(int argc, const char **argv);\n int cmd__oidmap(int argc, const char **argv);\n int cmd__oidtree(int argc, const char **argv);\n int cmd__online_cpus(int argc, const char **argv);\n+int cmd__pack_mtimes(int argc, const char **argv);\n int cmd__parse_options(int argc, const char **argv);\n int cmd__parse_pathspec_file(int argc, const char** argv);\n int cmd__partial_clone(int argc, const char **argv);\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442524","messageId":"37fda94785f1e689c7a7c32e69c6ff16fee7da4f.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 11/17] builtin/pack-objects.c: --cruft with expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:30Z","receivedAt":"2021-11-29T22:26:17Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a previous patch, pack-objects learned how to generate a cruft pack\nso long as no objects are dropped.\n\nThis patch teaches pack-objects to handle the case where a non-never\n`--cruft-expiration` value is passed. This case is slightly more\ncomplicated than before, because we want pack-objects to save\nunreachable objects which would have been pruned when there is another\nrecent (i.e., non-prunable) unreachable object which reaches the other.\nWe'll call these objects \"unreachable but reachable-from-recent\".\n\nHere is how pack-objects handles `--cruft-expiration`:\n\n  - Instead of adding all objects outside of the kept pack(s) into the\n    packing list, only handle the ones whose mtime is within the grace\n    period.\n\n  - Construct a reachability traversal whose tips are the\n    unreachable-but-recent objects.\n\n  - Then, walk along that traversal, stopping if we reach an object in\n    the kept pack. At each step along the traversal, we add the object\n    we are visiting to the packing list.\n\nIn the majority of these cases, any object we visit in this traversal\nwill already be in our packing list. But we will sometimes encounter\nreachable-from-recent cruft objects, which we want to retain even if\nthey aged out of the grace period.\n\nThe most subtle point of this process is that we actually don't need to\nbother to update the rescued object's mtime. Even though we will write\nan .mtimes file with a value that is older than the expiration window,\nit will continue to survive cruft repacks so long as any objects which\nreach it haven't aged out.\n\nThat is, a future repack will also exclude that object from the initial\npacking list, only to discover it later on when doing the reachability\ntraversal.\n\nFinally, stopping early once an object is found in a kept pack is safe\nto do because the kept packs ordinarily represent which packs will\nsurvive after repacking. Assuming that it _isn't_ safe to halt a\ntraversal early would mean that there is some ancestor object which is\nmissing, which implies repository corruption (i.e., the complete set of\nreachable objects isn't present).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c        |  84 +++++++++++++++++++-\n t/t5327-pack-objects-cruft.sh | 143 ++++++++++++++++++++++++++++++++++\n 2 files changed, 226 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 2c592d369a..a38fa34479 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -3439,6 +3439,44 @@ static int add_cruft_object_entry(const struct object_id *oid, enum object_type\n \treturn 1;\n }\n \n+static void show_cruft_object(struct object *obj, const char *name, void *data)\n+{\n+\t/*\n+\t * if we did not record it earlier, it's at least as old as our\n+\t * expiration value. Rather than find it exactly, just use that\n+\t * value.  This may bump it forward from its real mtime, but it\n+\t * will still be \"too old\" next time we run with the same\n+\t * expiration.\n+\t *\n+\t * if obj does appear in the packing list, this call is a noop (or may\n+\t * set the namehash).\n+\t */\n+\tadd_cruft_object_entry(&obj->oid, obj->type, NULL, 0, name, cruft_expiration);\n+}\n+\n+static void show_cruft_commit(struct commit *commit, void *data)\n+{\n+\tshow_cruft_object((struct object*)commit, NULL, data);\n+}\n+\n+static int cruft_include_check_obj(struct object *obj, void *data)\n+{\n+\treturn !has_object_kept_pack(&obj->oid, IN_CORE_KEEP_PACKS);\n+}\n+\n+static int cruft_include_check(struct commit *commit, void *data)\n+{\n+\treturn cruft_include_check_obj((struct object*)commit, data);\n+}\n+\n+static void set_cruft_mtime(const struct object *object,\n+\t\t\t    struct packed_git *pack,\n+\t\t\t    off_t offset, time_t mtime)\n+{\n+\tadd_cruft_object_entry(&object->oid, object->type, pack, offset, NULL,\n+\t\t\t       mtime);\n+}\n+\n static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n {\n \tstruct string_list_item *item = NULL;\n@@ -3464,6 +3502,50 @@ static void enumerate_cruft_objects(void)\n \tstop_progress(&progress_state);\n }\n \n+static void enumerate_and_traverse_cruft_objects(struct string_list *fresh_packs)\n+{\n+\tstruct packed_git *p;\n+\tstruct rev_info revs;\n+\tint ret;\n+\n+\trepo_init_revisions(the_repository, &revs, NULL);\n+\n+\trevs.tag_objects = 1;\n+\trevs.tree_objects = 1;\n+\trevs.blob_objects = 1;\n+\n+\trevs.include_check = cruft_include_check;\n+\trevs.include_check_obj = cruft_include_check_obj;\n+\n+\trevs.ignore_missing_links = 1;\n+\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Enumerating cruft objects\"), 0);\n+\tret = add_unseen_recent_objects_to_traversal(&revs, cruft_expiration,\n+\t\t\t\t\t\t     set_cruft_mtime, 1);\n+\tstop_progress(&progress_state);\n+\n+\tif (ret)\n+\t\tdie(_(\"unable to add cruft objects\"));\n+\n+\t/*\n+\t * Re-mark only the fresh packs as kept so that objects in\n+\t * unknown packs do not halt the reachability traversal early.\n+\t */\n+\tfor (p = get_all_packs(the_repository); p; p = p->next)\n+\t\tp->pack_keep_in_core = 0;\n+\tmark_pack_kept_in_core(fresh_packs, 1);\n+\n+\tif (prepare_revision_walk(&revs))\n+\t\tdie(_(\"revision walk setup failed\"));\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Traversing cruft objects\"), 0);\n+\tnr_seen = 0;\n+\ttraverse_commit_list(&revs, show_cruft_commit, show_cruft_object, NULL);\n+\n+\tstop_progress(&progress_state);\n+}\n+\n static void read_cruft_objects(void)\n {\n \tstruct strbuf buf = STRBUF_INIT;\n@@ -3515,7 +3597,7 @@ static void read_cruft_objects(void)\n \tmark_pack_kept_in_core(&discard_packs, 0);\n \n \tif (cruft_expiration)\n-\t\tdie(\"--cruft-expiration not yet implemented\");\n+\t\tenumerate_and_traverse_cruft_objects(&fresh_packs);\n \telse\n \t\tenumerate_cruft_objects();\n \ndiff --git a/t/t5327-pack-objects-cruft.sh b/t/t5327-pack-objects-cruft.sh\nindex 543a80e9bf..31d4a561fe 100755\n--- a/t/t5327-pack-objects-cruft.sh\n+++ b/t/t5327-pack-objects-cruft.sh\n@@ -214,5 +214,148 @@ basic_cruft_pack_tests () {\n }\n \n basic_cruft_pack_tests never\n+basic_cruft_pack_tests 2.weeks.ago\n+\n+test_expect_success 'cruft tags rescue tagged objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit tagged &&\n+\t\tgit tag -a annotated -m tag &&\n+\n+\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\twhile read oid\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $oid)\"\n+\t\tdone <objects &&\n+\n+\t\ttest-tool chmtime -500 \\\n+\t\t\t\"$objdir/$(test_oid_to_path $(git rev-parse annotated))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\t\tcut -f1 -d\" \" <actual.raw | sort >actual &&\n+\n+\t\t(\n+\t\t\tcat objects &&\n+\t\t\tgit rev-parse annotated\n+\t\t) >expect.raw &&\n+\t\tsort <expect.raw >expect &&\n+\n+\t\ttest_cmp expect actual &&\n+\t\tcat actual\n+\t)\n+'\n+\n+test_expect_success 'cruft commits rescue parents, trees' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit old &&\n+\t\ttest_commit new &&\n+\n+\t\tgit rev-list --objects --no-object-names packed..new >objects &&\n+\t\twhile read object\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $object)\"\n+\t\tdone <objects &&\n+\t\ttest-tool chmtime +500 \"$objdir/$(test_oid_to_path \\\n+\t\t\t$(git rev-parse HEAD))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\tcut -d\" \" -f1 <actual.raw | sort >actual &&\n+\t\tsort <objects >expect &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'cruft trees rescue sub-trees, blobs' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\tmkdir -p dir/sub &&\n+\t\techo foo >foo &&\n+\t\techo bar >dir/bar &&\n+\t\techo baz >dir/sub/baz &&\n+\n+\t\ttest_tick &&\n+\t\tgit add . &&\n+\t\tgit commit -m \"pruned\" &&\n+\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD^{tree}))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:foo))\" &&\n+\t\ttest-tool chmtime  -500 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/bar))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/sub))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/sub/baz))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\t\tcut -f1 -d\" \" <actual.raw | sort >actual &&\n+\n+\t\tgit rev-parse HEAD:dir HEAD:dir/bar HEAD:dir/sub HEAD:dir/sub/baz >expect.raw &&\n+\t\tsort <expect.raw >expect &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'expired objects are pruned' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit pruned &&\n+\n+\t\tgit rev-list --objects --no-object-names packed..pruned >objects &&\n+\t\twhile read object\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $object)\"\n+\t\tdone <objects &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual &&\n+\t\ttest_must_be_empty actual\n+\t)\n+'\n \n test_done\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442525","messageId":"52e9ac571070154e4669b7f5f68685ccdb9e5337.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 10/17] reachable: report precise timestamps from objects in cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:27Z","receivedAt":"2021-11-29T22:26:19Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When generating a cruft pack, the caller within pack-objects will want\nto know the precise timestamps of cruft objects (i.e., their\ncorresponding values in the .mtimes table) rather than the mtime of the\ncruft pack itself.\n\nTeach add_recent_packed() to lookup each object's precise mtime from the\n.mtimes file if one exists (indicated by the is_cruft bit on the\npacked_git structure).\n\nA couple of small things worth noting here:\n\n  - load_pack_mtimes() needs to be called before asking for\n    nth_packed_mtime(), and that call is done lazily here. That function\n    exits early if the .mtimes file has already been opened and parsed,\n    so only the first call is slow.\n\n  - Checking the is_cruft bit can be done without any extra work on the\n    caller's behalf, since it is set up for us automatically as a\n    side-effect of calling add_packed_git() (just like the 'pack_keep'\n    and 'pack_promisor' bits).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n reachable.c | 9 ++++++++-\n 1 file changed, 8 insertions(+), 1 deletion(-)\n\ndiff --git a/reachable.c b/reachable.c\nindex 0eb9909f47..9ec8e6bd5b 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -13,6 +13,7 @@\n #include \"worktree.h\"\n #include \"object-store.h\"\n #include \"pack-bitmap.h\"\n+#include \"pack-mtimes.h\"\n \n struct connectivity_progress {\n \tstruct progress *progress;\n@@ -155,6 +156,7 @@ static int add_recent_packed(const struct object_id *oid,\n \t\t\t     void *data)\n {\n \tstruct object *obj;\n+\ttimestamp_t mtime = p->mtime;\n \n \tif (!want_recent_object(data, oid))\n \t\treturn 0;\n@@ -163,7 +165,12 @@ static int add_recent_packed(const struct object_id *oid,\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n-\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), p->mtime, data);\n+\tif (p->is_cruft) {\n+\t\tif (load_pack_mtimes(p) < 0)\n+\t\t\tdie(_(\"could not load cruft pack .mtimes\"));\n+\t\tmtime = nth_packed_mtime(p, pos);\n+\t}\n+\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), mtime, data);\n \treturn 0;\n }\n \n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442526","messageId":"5710933127b01125ebcfe232868abbe87fce0d87.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 07/17] builtin/pack-objects.c: return from create_object_entry()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:20Z","receivedAt":"2021-11-29T22:26:22Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"A new caller in the next commit will want to immediately modify the\nobject_entry structure created by create_object_entry(). Instead of\nforcing that caller to wastefully look-up the entry we just created,\nreturn it from create_object_entry() instead.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 16 +++++++++-------\n 1 file changed, 9 insertions(+), 7 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex bf45ffbc57..3fb10529ba 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1508,13 +1508,13 @@ static int want_object_in_pack(const struct object_id *oid,\n \treturn 1;\n }\n \n-static void create_object_entry(const struct object_id *oid,\n-\t\t\t\tenum object_type type,\n-\t\t\t\tuint32_t hash,\n-\t\t\t\tint exclude,\n-\t\t\t\tint no_try_delta,\n-\t\t\t\tstruct packed_git *found_pack,\n-\t\t\t\toff_t found_offset)\n+static struct object_entry *create_object_entry(const struct object_id *oid,\n+\t\t\t\t\t\tenum object_type type,\n+\t\t\t\t\t\tuint32_t hash,\n+\t\t\t\t\t\tint exclude,\n+\t\t\t\t\t\tint no_try_delta,\n+\t\t\t\t\t\tstruct packed_git *found_pack,\n+\t\t\t\t\t\toff_t found_offset)\n {\n \tstruct object_entry *entry;\n \n@@ -1531,6 +1531,8 @@ static void create_object_entry(const struct object_id *oid,\n \t}\n \n \tentry->no_try_delta = no_try_delta;\n+\n+\treturn entry;\n }\n \n static const char no_closure_warning[] = N_(\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442527","messageId":"66165917a4660f63ce60b820d178d52a51304d20.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:22Z","receivedAt":"2021-11-29T22:26:23Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Teach `pack-objects` how to generate a cruft pack when no objects are\ndropped (i.e., `--cruft-expiration=never`). Later patches will teach\n`pack-objects` how to generate a cruft pack that prunes objects.\n\nWhen generating a cruft pack which does not prune objects, we want to\ncollect all unreachable objects into a single pack (noting and updating\ntheir mtimes as we accumulate them). Ordinary use will pass the result\nof a `git repack -A` as a kept pack, so when this patch says \"kept\npack\", readers should think \"reachable objects\".\n\nGenerating a non-expiring cruft packs works as follows:\n\n  - Callers provide a list of every pack they know about, and indicate\n    which packs are about to be removed.\n\n  - All packs which are going to be removed (we'll call these the\n    redundant ones) are marked as kept in-core, as well as any packs\n    that `pack-objects` found but the caller did not specify.\n\n    These packs are presumed to have entered the repository between\n    the caller collecting packs and invoking `pack-objects`. Since we\n    do not want to include objects in these packs (because we don't know\n    which of their objects are or aren't reachable), these are also\n    marked as kept in-core.\n\n  - Then, we enumerate all objects in the repository, and add them to\n    our packing list if they do not appear in an in-core kept pack.\n\nThis results in a new cruft pack which contains all known objects that\naren't included in the kept packs. When the kept pack is the result of\n`git repack -A`, the resulting pack contains all unreachable objects.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-pack-objects.txt |  23 +++\n builtin/pack-objects.c             | 203 ++++++++++++++++++++++++++-\n object-file.c                      |   2 +-\n object-store.h                     |   2 +\n t/t5327-pack-objects-cruft.sh      | 218 +++++++++++++++++++++++++++++\n 5 files changed, 442 insertions(+), 6 deletions(-)\n create mode 100755 t/t5327-pack-objects-cruft.sh\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex dbfd1f9017..573c18afcd 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -13,6 +13,7 @@ SYNOPSIS\n \t[--no-reuse-delta] [--delta-base-offset] [--non-empty]\n \t[--local] [--incremental] [--window=<n>] [--depth=<n>]\n \t[--revs [--unpacked | --all]] [--keep-pack=<pack-name>]\n+\t[--cruft] [--cruft-expiration=<time>]\n \t[--stdout [--filter=<filter-spec>] | base-name]\n \t[--shallow] [--keep-true-parents] [--[no-]sparse] < object-list\n \n@@ -95,6 +96,28 @@ base-name::\n Incompatible with `--revs`, or options that imply `--revs` (such as\n `--all`), with the exception of `--unpacked`, which is compatible.\n \n+--cruft::\n+\tPacks unreachable objects into a separate \"cruft\" pack, denoted\n+\tby the existence of a `.mtimes` file. Pack names provided over\n+\tstdin indicate which packs will remain after a `git repack`.\n+\tPack names prefixed with a `-` indicate those which will be\n+\tremoved. The contents of the cruft pack are all objects not\n+\tcontained in the surviving packs specified by `--keep-pack`)\n+\twhich have not exceeded the grace period (see\n+\t`--cruft-expiration` below), or which have exceeded the grace\n+\tperiod, but are reachable from an other object which hasn't.\n++\n+Incompatible with `--unpack-unreachable`, `--keep-unreachable`,\n+`--pack-loose-unreachable`, `--stdin-packs`, as well as any other\n+options which imply `--revs`. Also incompatible with `--max-pack-size`;\n+when this option is set, the maximum pack size is not inferred from\n+`pack.packSizeLimit`.\n+\n+--cruft-expiration=<approxidate>::\n+\tIf specified, objects are eliminated from the cruft pack if they\n+\thave an mtime older than `<approxidate>`. If unspecified (and\n+\tgiven `--cruft`), then no objects are eliminated.\n+\n --window=<n>::\n --depth=<n>::\n \tThese two options affect how the objects contained in\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 3fb10529ba..b12e79e4b1 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -36,6 +36,7 @@\n #include \"trace2.h\"\n #include \"shallow.h\"\n #include \"promisor-remote.h\"\n+#include \"pack-mtimes.h\"\n \n /*\n  * Objects we are going to pack are collected in the `to_pack` structure.\n@@ -194,6 +195,8 @@ static int reuse_delta = 1, reuse_object = 1;\n static int keep_unreachable, unpack_unreachable, include_tag;\n static timestamp_t unpack_unreachable_expiration;\n static int pack_loose_unreachable;\n+static int cruft;\n+static timestamp_t cruft_expiration;\n static int local;\n static int have_non_local_packs;\n static int incremental;\n@@ -1252,6 +1255,9 @@ static void write_pack_file(void)\n \t\t\t\t\t&to_pack, written_list, nr_written);\n \t\t\t}\n \n+\t\t\tif (cruft)\n+\t\t\t\tpack_idx_opts.flags |= WRITE_MTIMES;\n+\n \t\t\tstage_tmp_packfiles(&tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n \t\t\t\t\t    &to_pack, &pack_idx_opts, hash,\n@@ -3389,6 +3395,135 @@ static void read_packs_list_from_stdin(void)\n \tstring_list_clear(&exclude_packs, 0);\n }\n \n+static int add_cruft_object_entry(const struct object_id *oid, enum object_type type,\n+\t\t\t\t  struct packed_git *pack, off_t offset,\n+\t\t\t\t  const char *name, uint32_t mtime)\n+{\n+\tstruct object_entry *entry;\n+\n+\tdisplay_progress(progress_state, ++nr_seen);\n+\n+\tentry = packlist_find(&to_pack, oid);\n+\tif (entry) {\n+\t\tif (name) {\n+\t\t\tentry->hash = pack_name_hash(name);\n+\t\t\tentry->no_try_delta = name && no_try_delta(name);\n+\t\t}\n+\t} else {\n+\t\tif (!want_object_in_pack(oid, 0, &pack, &offset))\n+\t\t\treturn 0;\n+\t\tif (!pack && type == OBJ_BLOB && !has_loose_object(oid)) {\n+\t\t\t/*\n+\t\t\t * If a traversed tree has a missing blob then we want\n+\t\t\t * to avoid adding that missing object to our pack.\n+\t\t\t *\n+\t\t\t * This only applies to missing blobs, not trees,\n+\t\t\t * because the traversal needs to parse sub-trees but\n+\t\t\t * not blobs.\n+\t\t\t *\n+\t\t\t * Note we only perform this check when we couldn't\n+\t\t\t * already find the object in a pack, so we're really\n+\t\t\t * limited to \"ensure non-tip blobs which don't exist in\n+\t\t\t * packs do exist via loose objects\". Confused?\n+\t\t\t */\n+\t\t\treturn 0;\n+\t\t}\n+\n+\t\tentry = create_object_entry(oid, type, pack_name_hash(name),\n+\t\t\t\t\t    0, name && no_try_delta(name),\n+\t\t\t\t\t    pack, offset);\n+\t}\n+\n+\tif (mtime > oe_cruft_mtime(&to_pack, entry))\n+\t\toe_set_cruft_mtime(&to_pack, entry, mtime);\n+\treturn 1;\n+}\n+\n+static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n+{\n+\tstruct string_list_item *item = NULL;\n+\tfor_each_string_list_item(item, packs) {\n+\t\tstruct packed_git *p = item->util;\n+\t\tif (!p)\n+\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n+\t\tp->pack_keep_in_core = keep;\n+\t}\n+}\n+\n+static void add_unreachable_loose_objects(void);\n+static void add_objects_in_unpacked_packs(void);\n+\n+static void enumerate_cruft_objects(void)\n+{\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Enumerating cruft objects\"), 0);\n+\n+\tadd_objects_in_unpacked_packs();\n+\tadd_unreachable_loose_objects();\n+\n+\tstop_progress(&progress_state);\n+}\n+\n+static void read_cruft_objects(void)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct string_list discard_packs = STRING_LIST_INIT_DUP;\n+\tstruct string_list fresh_packs = STRING_LIST_INIT_DUP;\n+\tstruct packed_git *p;\n+\n+\tignore_packed_keep_in_core = 1;\n+\n+\twhile (strbuf_getline(&buf, stdin) != EOF) {\n+\t\tif (!buf.len)\n+\t\t\tcontinue;\n+\n+\t\tif (*buf.buf == '-')\n+\t\t\tstring_list_append(&discard_packs, buf.buf + 1);\n+\t\telse\n+\t\t\tstring_list_append(&fresh_packs, buf.buf);\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstring_list_sort(&discard_packs);\n+\tstring_list_sort(&fresh_packs);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tconst char *pack_name = pack_basename(p);\n+\t\tstruct string_list_item *item;\n+\n+\t\titem = string_list_lookup(&fresh_packs, pack_name);\n+\t\tif (!item)\n+\t\t\titem = string_list_lookup(&discard_packs, pack_name);\n+\n+\t\tif (item) {\n+\t\t\titem->util = p;\n+\t\t} else {\n+\t\t\t/*\n+\t\t\t * This pack wasn't mentioned in either the \"fresh\" or\n+\t\t\t * \"discard\" list, so the caller didn't know about it.\n+\t\t\t *\n+\t\t\t * Mark it as kept so that its objects are ignored by\n+\t\t\t * add_unseen_recent_objects_to_traversal(). We'll\n+\t\t\t * unmark it before starting the traversal so it doesn't\n+\t\t\t * halt the traversal early.\n+\t\t\t */\n+\t\t\tp->pack_keep_in_core = 1;\n+\t\t}\n+\t}\n+\n+\tmark_pack_kept_in_core(&fresh_packs, 1);\n+\tmark_pack_kept_in_core(&discard_packs, 0);\n+\n+\tif (cruft_expiration)\n+\t\tdie(\"--cruft-expiration not yet implemented\");\n+\telse\n+\t\tenumerate_cruft_objects();\n+\n+\tstrbuf_release(&buf);\n+\tstring_list_clear(&discard_packs, 0);\n+\tstring_list_clear(&fresh_packs, 0);\n+}\n+\n static void read_object_list_from_stdin(void)\n {\n \tchar line[GIT_MAX_HEXSZ + 1 + PATH_MAX + 2];\n@@ -3521,7 +3656,24 @@ static int add_object_in_unpacked_pack(const struct object_id *oid,\n \t\t\t\t       uint32_t pos,\n \t\t\t\t       void *_data)\n {\n-\tadd_object_entry(oid, OBJ_NONE, \"\", 0);\n+\tif (cruft) {\n+\t\toff_t offset;\n+\t\ttime_t mtime;\n+\n+\t\tif (pack->is_cruft) {\n+\t\t\tif (load_pack_mtimes(pack) < 0)\n+\t\t\t\tdie(_(\"could not load cruft pack .mtimes\"));\n+\t\t\tmtime = nth_packed_mtime(pack, pos);\n+\t\t} else {\n+\t\t\tmtime = pack->mtime;\n+\t\t}\n+\t\toffset = nth_packed_object_offset(pack, pos);\n+\n+\t\tadd_cruft_object_entry(oid, OBJ_NONE, pack, offset,\n+\t\t\t\t       NULL, mtime);\n+\t} else {\n+\t\tadd_object_entry(oid, OBJ_NONE, \"\", 0);\n+\t}\n \treturn 0;\n }\n \n@@ -3545,7 +3697,19 @@ static int add_loose_object(const struct object_id *oid, const char *path,\n \t\treturn 0;\n \t}\n \n-\tadd_object_entry(oid, type, \"\", 0);\n+\tif (cruft) {\n+\t\tstruct stat st;\n+\t\tif (stat(path, &st) < 0) {\n+\t\t\tif (errno == ENOENT)\n+\t\t\t\treturn 0;\n+\t\t\treturn error_errno(\"unable to stat %s\", oid_to_hex(oid));\n+\t\t}\n+\n+\t\tadd_cruft_object_entry(oid, type, NULL, 0, NULL,\n+\t\t\t\t       st.st_mtime);\n+\t} else {\n+\t\tadd_object_entry(oid, type, \"\", 0);\n+\t}\n \treturn 0;\n }\n \n@@ -3864,6 +4028,20 @@ static int option_parse_unpack_unreachable(const struct option *opt,\n \treturn 0;\n }\n \n+static int option_parse_cruft_expiration(const struct option *opt,\n+\t\t\t\t\t const char *arg, int unset)\n+{\n+\tif (unset) {\n+\t\tcruft = 0;\n+\t\tcruft_expiration = 0;\n+\t} else {\n+\t\tcruft = 1;\n+\t\tif (arg)\n+\t\t\tcruft_expiration = approxidate(arg);\n+\t}\n+\treturn 0;\n+}\n+\n int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n {\n \tint use_internal_rev_list = 0;\n@@ -3936,6 +4114,10 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tOPT_CALLBACK_F(0, \"unpack-unreachable\", NULL, N_(\"time\"),\n \t\t  N_(\"unpack unreachable objects newer than <time>\"),\n \t\t  PARSE_OPT_OPTARG, option_parse_unpack_unreachable),\n+\t\tOPT_BOOL(0, \"cruft\", &cruft, N_(\"create a cruft pack\")),\n+\t\tOPT_CALLBACK_F(0, \"cruft-expiration\", NULL, N_(\"time\"),\n+\t\t  N_(\"expire cruft objects older than <time>\"),\n+\t\t  PARSE_OPT_OPTARG, option_parse_cruft_expiration),\n \t\tOPT_BOOL(0, \"sparse\", &sparse,\n \t\t\t N_(\"use the sparse reachability algorithm\")),\n \t\tOPT_BOOL(0, \"thin\", &thin,\n@@ -4060,7 +4242,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \n \tif (!HAVE_THREADS && delta_search_threads != 1)\n \t\twarning(_(\"no threads support, ignoring --threads\"));\n-\tif (!pack_to_stdout && !pack_size_limit)\n+\tif (!pack_to_stdout && !pack_size_limit && !cruft)\n \t\tpack_size_limit = pack_size_limit_cfg;\n \tif (pack_to_stdout && pack_size_limit)\n \t\tdie(_(\"--max-pack-size cannot be used to build a pack for transfer\"));\n@@ -4087,6 +4269,15 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (stdin_packs && use_internal_rev_list)\n \t\tdie(_(\"cannot use internal rev list with --stdin-packs\"));\n \n+\tif (cruft) {\n+\t\tif (use_internal_rev_list)\n+\t\t\tdie(_(\"cannot use internal rev list with --cruft\"));\n+\t\tif (stdin_packs)\n+\t\t\tdie(_(\"cannot use --stdin-packs with --cruft\"));\n+\t\tif (pack_size_limit)\n+\t\t\tdie(_(\"cannot use --max-pack-size with --cruft\"));\n+\t}\n+\n \t/*\n \t * \"soft\" reasons not to use bitmaps - for on-disk repack by default we want\n \t *\n@@ -4143,7 +4334,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t    the_repository);\n \tprepare_packing_data(the_repository, &to_pack);\n \n-\tif (progress)\n+\tif (progress && !cruft)\n \t\tprogress_state = start_progress(_(\"Enumerating objects\"), 0);\n \tif (stdin_packs) {\n \t\t/* avoids adding objects in excluded packs */\n@@ -4151,7 +4342,9 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tread_packs_list_from_stdin();\n \t\tif (rev_list_unpacked)\n \t\t\tadd_unreachable_loose_objects();\n-\t} else if (!use_internal_rev_list)\n+\t} else if (cruft)\n+\t\tread_cruft_objects();\n+\telse if (!use_internal_rev_list)\n \t\tread_object_list_from_stdin();\n \telse {\n \t\tget_object_list(rp.nr, rp.v);\ndiff --git a/object-file.c b/object-file.c\nindex c3d866a287..7ddb38b64a 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -956,7 +956,7 @@ int has_loose_object_nonlocal(const struct object_id *oid)\n \treturn check_and_freshen_nonlocal(oid, 0);\n }\n \n-static int has_loose_object(const struct object_id *oid)\n+int has_loose_object(const struct object_id *oid)\n {\n \treturn check_and_freshen(oid, 0);\n }\ndiff --git a/object-store.h b/object-store.h\nindex d87481f101..a79c1c91ab 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -308,6 +308,8 @@ int repo_has_object_file_with_flags(struct repository *r,\n  */\n int has_loose_object_nonlocal(const struct object_id *);\n \n+int has_loose_object(const struct object_id *);\n+\n void assert_oid_type(const struct object_id *oid, enum object_type expect);\n \n /*\ndiff --git a/t/t5327-pack-objects-cruft.sh b/t/t5327-pack-objects-cruft.sh\nnew file mode 100755\nindex 0000000000..543a80e9bf\n--- /dev/null\n+++ b/t/t5327-pack-objects-cruft.sh\n@@ -0,0 +1,218 @@\n+#!/bin/sh\n+\n+test_description='cruft pack related pack-objects tests'\n+. ./test-lib.sh\n+\n+objdir=.git/objects\n+packdir=$objdir/pack\n+\n+basic_cruft_pack_tests () {\n+\texpire=\"$1\"\n+\n+\ttest_expect_success \"unreachable loose objects are packed (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit base &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit loose &&\n+\n+\t\t\ttest-tool chmtime +2000 \"$objdir/$(test_oid_to_path \\\n+\t\t\t\t$(git rev-parse loose:loose.t))\" &&\n+\t\t\ttest-tool chmtime +1000 \"$objdir/$(test_oid_to_path \\\n+\t\t\t\t$(git rev-parse loose^{tree}))\" &&\n+\n+\t\t\t(\n+\t\t\t\tgit rev-list --objects --no-object-names base..loose |\n+\t\t\t\twhile read oid\n+\t\t\t\tdo\n+\t\t\t\t\tpath=\"$objdir/$(test_oid_to_path \"$oid\")\" &&\n+\t\t\t\t\tprintf \"%s %d\\n\" \"$oid\" \"$(test-tool chmtime --get \"$path\")\"\n+\t\t\t\tdone |\n+\t\t\t\tsort -k1\n+\t\t\t) >expect &&\n+\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t\t--cruft-expiration=\"$expire\" $packdir/pack)\" &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"unreachable packed objects are packed (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit packed &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit other &&\n+\n+\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\t\tother=\"$(git pack-objects --delta-base-offset \\\n+\t\t\t\t$packdir/pack <objects)\" &&\n+\t\t\tgit prune-packed &&\n+\n+\t\t\ttest-tool chmtime --get -100 \"$packdir/pack-$other.pack\" >expect &&\n+\n+\t\t\tcruft=\"$(git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-pack-$other.pack\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\t\tcut -d\" \" -f2 <actual.raw | sort -u >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"unreachable cruft objects are repacked (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit packed &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit other &&\n+\n+\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\n+\t\t\tcruft_a=\"$(echo $keep | git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack)\" &&\n+\t\t\tgit prune-packed &&\n+\t\t\tcruft_b=\"$(git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-pack-$cruft_a.pack\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft_a.mtimes\" >expect.raw &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft_b.mtimes\" >actual.raw &&\n+\n+\t\t\tsort <expect.raw >expect &&\n+\t\t\tsort <actual.raw >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"multiple cruft packs (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\tgit repack -Ad &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\n+\t\t\ttest_commit cruft &&\n+\t\t\tloose=\"$objdir/$(test_oid_to_path $(git rev-parse cruft))\" &&\n+\n+\t\t\t# generate three copies of the cruft object in different\n+\t\t\t# cruft packs, each with a unique mtime:\n+\t\t\t#   - one expired (1000 seconds ago)\n+\t\t\t#   - two non-expired (one 1000 seconds in the future,\n+\t\t\t#     one 1500 seconds in the future)\n+\t\t\ttest-tool chmtime =-1000 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-A <<-EOF &&\n+\t\t\t$keep\n+\t\t\tEOF\n+\t\t\ttest-tool chmtime =+1000 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-B <<-EOF &&\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\tEOF\n+\t\t\ttest-tool chmtime =+1500 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-C <<-EOF &&\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-B-*.pack))\n+\t\t\tEOF\n+\n+\t\t\t# ensure the resulting cruft pack takes the most recent\n+\t\t\t# mtime among all copies\n+\t\t\tcruft=\"$(git pack-objects --cruft \\\n+\t\t\t\t--cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-B-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-C-*.pack))\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\n+\t\t\ttest-tool pack-mtimes \"$(basename $(ls $packdir/pack-C-*.mtimes))\" >expect.raw &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\t\tsort expect.raw >expect &&\n+\t\t\tsort actual.raw >actual &&\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"cruft packs tolerate missing trees (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\ttest_commit cruft &&\n+\n+\t\t\ttree=\"$(git rev-parse cruft^{tree})\" &&\n+\n+\t\t\tgit reset --hard reachable &&\n+\t\t\tgit tag -d cruft &&\n+\t\t\trm -fr .git/logs &&\n+\n+\t\t\t# remove the unreachable tree, but leave the commit\n+\t\t\t# which has it as its root tree in-tact\n+\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$tree\")\" &&\n+\n+\t\t\tgit repack -Ad &&\n+\t\t\tbasename $(ls $packdir/pack-*.pack) >in &&\n+\t\t\tgit pack-objects --cruft --cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <in\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"cruft packs tolerate missing blobs (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\ttest_commit cruft &&\n+\n+\t\t\tblob=\"$(git rev-parse cruft:cruft.t)\" &&\n+\n+\t\t\tgit reset --hard reachable &&\n+\t\t\tgit tag -d cruft &&\n+\t\t\trm -fr .git/logs &&\n+\n+\t\t\t# remove the unreachable blob, but leave the commit (and\n+\t\t\t# the root tree of that commit) in-tact\n+\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$blob\")\" &&\n+\n+\t\t\tgit repack -Ad &&\n+\t\t\tbasename $(ls $packdir/pack-*.pack) >in &&\n+\t\t\tgit pack-objects --cruft --cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <in\n+\t\t)\n+\t'\n+}\n+\n+basic_cruft_pack_tests never\n+\n+test_done\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442528","messageId":"a05675ab834ac5e8bc3ab72847b0621a563e0e1b.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 12/17] builtin/repack.c: support generating a cruft pack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:33Z","receivedAt":"2021-11-29T22:26:27Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Expose a way to split the contents of a repository into a main and cruft\npack when doing an all-into-one repack with `git repack --cruft -d`, and\na complementary configuration variable.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-repack.txt            |  11 ++\n Documentation/technical/cruft-packs.txt |   2 +-\n builtin/repack.c                        | 112 ++++++++++++++++-\n t/t5327-pack-objects-cruft.sh           | 153 ++++++++++++++++++++++++\n 4 files changed, 272 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/git-repack.txt b/Documentation/git-repack.txt\nindex 7183fb498f..4f8f4b5a1f 100644\n--- a/Documentation/git-repack.txt\n+++ b/Documentation/git-repack.txt\n@@ -63,6 +63,17 @@ to the new separate pack will be written.\n \tAlso run  'git prune-packed' to remove redundant\n \tloose object files.\n \n+--cruft::\n+\tSame as `-a`, unless `-d` is used. Then any unreachable objects\n+\tare packed into a separate cruft pack. Unreachable objects can\n+\tbe pruned using the normal expiry rules with the next `git gc`\n+\tinvocation (see linkgit:git-gc[1]). Incompatible with `-k`.\n+\n+--cruft-expiration=<approxidate>::\n+\tExpire unreachable objects older than `<approxidate>`\n+\timmediately instead of waiting for the next `git gc` invocation.\n+\tOnly useful with `--cruft -d`.\n+\n -l::\n \tPass the `--local` option to 'git pack-objects'. See\n \tlinkgit:git-pack-objects[1].\ndiff --git a/Documentation/technical/cruft-packs.txt b/Documentation/technical/cruft-packs.txt\nindex bb54cce1b1..b7daad2e3e 100644\n--- a/Documentation/technical/cruft-packs.txt\n+++ b/Documentation/technical/cruft-packs.txt\n@@ -16,7 +16,7 @@ pruned according to normal expiry rules with the next 'git gc' invocation.\n \n Unreachable objects aren't removed immediately, since doing so could race with\n an incoming push which may reference an object which is about to be deleted.\n-Instead, those unreachable objects are stored as loose object and stay that way\n+Instead, those unreachable objects are stored as loose objects and stay that way\n until they are older than the expiration window, at which point they are removed\n by linkgit:git-prune[1].\n \ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex acbb7b8c3b..68b4bdf06f 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -18,11 +18,17 @@\n #include \"pack-bitmap.h\"\n #include \"refs.h\"\n \n+#define ALL_INTO_ONE 1\n+#define LOOSEN_UNREACHABLE 2\n+#define PACK_CRUFT 4\n+\n+static int pack_everything;\n static int delta_base_offset = 1;\n static int pack_kept_objects = -1;\n static int write_bitmaps = -1;\n static int use_delta_islands;\n static char *packdir, *packtmp_name, *packtmp;\n+static char *cruft_expiration;\n \n static const char *const git_repack_usage[] = {\n \tN_(\"git repack [<options>]\"),\n@@ -54,6 +60,7 @@ static int repack_config(const char *var, const char *value, void *cb)\n \t\tuse_delta_islands = git_config_bool(var, value);\n \t\treturn 0;\n \t}\n+\n \treturn git_default_config(var, value, cb);\n }\n \n@@ -298,9 +305,6 @@ static void repack_promisor_objects(const struct pack_objects_args *args,\n \t\tdie(_(\"could not finish pack-objects to repack promisor objects\"));\n }\n \n-#define ALL_INTO_ONE 1\n-#define LOOSEN_UNREACHABLE 2\n-\n struct pack_geometry {\n \tstruct packed_git **pack;\n \tuint32_t pack_nr, pack_alloc;\n@@ -337,6 +341,8 @@ static void init_pack_geometry(struct pack_geometry **geometry_p)\n \tfor (p = get_all_packs(the_repository); p; p = p->next) {\n \t\tif (!pack_kept_objects && p->pack_keep)\n \t\t\tcontinue;\n+\t\tif (p->is_cruft)\n+\t\t\tcontinue;\n \n \t\tALLOC_GROW(geometry->pack,\n \t\t\t   geometry->pack_nr + 1,\n@@ -598,6 +604,67 @@ static int write_midx_included_packs(struct string_list *include,\n \treturn finish_command(&cmd);\n }\n \n+static int write_cruft_pack(const struct pack_objects_args *args,\n+\t\t\t    const char *pack_prefix,\n+\t\t\t    struct string_list *names,\n+\t\t\t    struct string_list *existing_packs,\n+\t\t\t    struct string_list *existing_kept_packs)\n+{\n+\tstruct child_process cmd = CHILD_PROCESS_INIT;\n+\tstruct strbuf line = STRBUF_INIT;\n+\tstruct string_list_item *item;\n+\tFILE *in, *out;\n+\tint ret;\n+\n+\tprepare_pack_objects(&cmd, args);\n+\n+\tstrvec_push(&cmd.args, \"--cruft\");\n+\tif (cruft_expiration)\n+\t\tstrvec_pushf(&cmd.args, \"--cruft-expiration=%s\",\n+\t\t\t     cruft_expiration);\n+\n+\tstrvec_push(&cmd.args, \"--honor-pack-keep\");\n+\tstrvec_push(&cmd.args, \"--non-empty\");\n+\tstrvec_push(&cmd.args, \"--max-pack-size=0\");\n+\n+\tcmd.in = -1;\n+\n+\tret = start_command(&cmd);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\t/*\n+\t * names has a confusing double use: it both provides the list\n+\t * of just-written new packs, and accepts the name of the cruft\n+\t * pack we are writing.\n+\t *\n+\t * By the time it is read here, it contains only the pack(s)\n+\t * that were just written, which is exactly the set of packs we\n+\t * want to consider kept.\n+\t */\n+\tin = xfdopen(cmd.in, \"w\");\n+\tfor_each_string_list_item(item, names)\n+\t\tfprintf(in, \"%s-%s.pack\\n\", pack_prefix, item->string);\n+\tfor_each_string_list_item(item, existing_packs)\n+\t\tfprintf(in, \"-%s.pack\\n\", item->string);\n+\tfor_each_string_list_item(item, existing_kept_packs)\n+\t\tfprintf(in, \"%s.pack\\n\", item->string);\n+\tfclose(in);\n+\n+\tout = xfdopen(cmd.out, \"r\");\n+\twhile (strbuf_getline_lf(&line, out) != EOF) {\n+\t\tif (line.len != the_hash_algo->hexsz)\n+\t\t\tdie(_(\"repack: Expecting full hex object ID lines only \"\n+\t\t\t      \"from pack-objects.\"));\n+\t\tstring_list_append(names, line.buf);\n+\t}\n+\tfclose(out);\n+\n+\tstrbuf_release(&line);\n+\n+\treturn finish_command(&cmd);\n+}\n+\n int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n \tstruct child_process cmd = CHILD_PROCESS_INIT;\n@@ -614,7 +681,6 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tint show_progress = isatty(2);\n \n \t/* variables to be filled by option parsing */\n-\tint pack_everything = 0;\n \tint delete_redundant = 0;\n \tconst char *unpack_unreachable = NULL;\n \tint keep_unreachable = 0;\n@@ -630,6 +696,11 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tOPT_BIT('A', NULL, &pack_everything,\n \t\t\t\tN_(\"same as -a, and turn unreachable objects loose\"),\n \t\t\t\t   LOOSEN_UNREACHABLE | ALL_INTO_ONE),\n+\t\tOPT_BIT(0, \"cruft\", &pack_everything,\n+\t\t\t\tN_(\"same as -a, pack unreachable cruft objects separately\"),\n+\t\t\t\t   PACK_CRUFT | ALL_INTO_ONE),\n+\t\tOPT_STRING(0, \"cruft-expiration\", &cruft_expiration, N_(\"approxidate\"),\n+\t\t\t\tN_(\"with -C, expire objects older than this\")),\n \t\tOPT_BOOL('d', NULL, &delete_redundant,\n \t\t\t\tN_(\"remove redundant packs, and run git-prune-packed\")),\n \t\tOPT_BOOL('f', NULL, &po_args.no_reuse_delta,\n@@ -681,6 +752,14 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (keep_unreachable &&\n \t    (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE)))\n \t\tdie(_(\"--keep-unreachable and -A are incompatible\"));\n+\tif (pack_everything & PACK_CRUFT && delete_redundant) {\n+\t\tif (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE))\n+\t\t\tdie(_(\"--cruft and -A are incompatible\"));\n+\t\tif (keep_unreachable)\n+\t\t\tdie(_(\"--cruft and -k are incompatible\"));\n+\t\tif (!(pack_everything & ALL_INTO_ONE))\n+\t\t\tdie(_(\"--cruft must be combined with all-into-one\"));\n+\t}\n \n \tif (write_bitmaps < 0) {\n \t\tif (!write_midx &&\n@@ -763,7 +842,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (pack_everything & ALL_INTO_ONE) {\n \t\trepack_promisor_objects(&po_args, &names);\n \n-\t\tif (existing_nonkept_packs.nr && delete_redundant) {\n+\t\tif (existing_nonkept_packs.nr && delete_redundant &&\n+\t\t    !(pack_everything & PACK_CRUFT)) {\n \t\t\tfor_each_string_list_item(item, &names) {\n \t\t\t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s-%s.pack\",\n \t\t\t\t\t     packtmp_name, item->string);\n@@ -798,6 +878,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\treturn ret;\n \n \tif (geometry) {\n+\t\tstruct packed_git *p;\n \t\tFILE *in = xfdopen(cmd.in, \"w\");\n \t\t/*\n \t\t * The resulting pack should contain all objects in packs that\n@@ -808,6 +889,12 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\tfprintf(in, \"%s\\n\", pack_basename(geometry->pack[i]));\n \t\tfor (i = geometry->split; i < geometry->pack_nr; i++)\n \t\t\tfprintf(in, \"^%s\\n\", pack_basename(geometry->pack[i]));\n+\n+\t\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\t\tif (!p->is_cruft)\n+\t\t\t\tcontinue;\n+\t\t\tfprintf(in, \"^%s\\n\", pack_basename(p));\n+\t\t}\n \t\tfclose(in);\n \t}\n \n@@ -825,6 +912,21 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (!names.nr && !po_args.quiet)\n \t\tprintf_ln(_(\"Nothing new to pack.\"));\n \n+\tif (pack_everything & PACK_CRUFT) {\n+\t\tconst char *pack_prefix;\n+\t\tif (!skip_prefix(packtmp, packdir, &pack_prefix))\n+\t\t\tdie(_(\"pack prefix %s does not begin with objdir %s\"),\n+\t\t\t    packtmp, packdir);\n+\t\tif (*pack_prefix == '/')\n+\t\t\tpack_prefix++;\n+\n+\t\tret = write_cruft_pack(&po_args, pack_prefix, &names,\n+\t\t\t\t       &existing_nonkept_packs,\n+\t\t\t\t       &existing_kept_packs);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t}\n+\n \tfor_each_string_list_item(item, &names) {\n \t\titem->util = (void *)(uintptr_t)populate_pack_exts(item->string);\n \t}\ndiff --git a/t/t5327-pack-objects-cruft.sh b/t/t5327-pack-objects-cruft.sh\nindex 31d4a561fe..ed1a113ab6 100755\n--- a/t/t5327-pack-objects-cruft.sh\n+++ b/t/t5327-pack-objects-cruft.sh\n@@ -358,4 +358,157 @@ test_expect_success 'expired objects are pruned' '\n \t)\n '\n \n+test_expect_success 'repack --cruft generates a cruft pack' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit branch -M main &&\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit unreachable &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d unreachable &&\n+\t\t# objects are not cruft if they are contained in the reflogs\n+\t\trm -fr .git/logs &&\n+\n+\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\t\tsort <reachable.raw >reachable &&\n+\t\tcomm -13 reachable objects >unreachable &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tcruft=$(basename $(ls $packdir/pack-*.mtimes) .mtimes) &&\n+\t\tpack=$(basename $(ls $packdir/pack-*.pack | grep -v $cruft) .pack) &&\n+\n+\t\tgit show-index <$packdir/$pack.idx >actual.raw &&\n+\t\tcut -f2 -d\" \" actual.raw | sort >actual &&\n+\t\ttest_cmp reachable actual &&\n+\n+\t\tgit show-index <$packdir/$cruft.idx >actual.raw &&\n+\t\tcut -f2 -d\" \" actual.raw | sort >actual &&\n+\t\ttest_cmp unreachable actual\n+\t)\n+'\n+\n+test_expect_success 'loose objects mtimes upsert others' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit cruft &&\n+\t\t# incremental repack, leaving existing objects loose (so\n+\t\t# they can be \"freshened\")\n+\t\tgit repack &&\n+\n+\t\ttip=\"$(git rev-parse cruft)\" &&\n+\t\tpath=\"$objdir/$(test_oid_to_path \"$(git rev-parse cruft)\")\" &&\n+\t\ttest-tool chmtime --get +1000 \"$path\" >expect &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d cruft &&\n+\t\trm -fr .git/logs &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tmtimes=\"$(basename $(ls $packdir/pack-*.mtimes))\" &&\n+\t\ttest-tool pack-mtimes \"$mtimes\" >actual.raw &&\n+\t\tgrep \"$tip\" actual.raw | cut -d\" \" -f2 >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'cruft packs are not included in geometric repack' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit cruft &&\n+\t\tgit repack -d &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d cruft &&\n+\t\trm -fr .git/logs &&\n+\n+\t\tgit repack --cruft &&\n+\n+\t\tfind $packdir -type f | sort >before &&\n+\t\tgit repack --geometric=2 -d &&\n+\t\tfind $packdir -type f | sort >after &&\n+\n+\t\ttest_cmp before after\n+\t)\n+'\n+test_expect_success 'cruft repack with no reachable objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\t\tgit repack -ad &&\n+\n+\t\tbase=\"$(git rev-parse base)\" &&\n+\n+\t\tgit for-each-ref --format=\"delete %(refname)\" >in &&\n+\t\tgit update-ref --stdin <in &&\n+\t\trm -fr .git/logs &&\n+\t\trm -fr .git/index &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tgit cat-file -t $base\n+\t)\n+'\n+\n+test_expect_success 'cruft repack ignores --max-pack-size' '\n+\tgit init max-pack-size &&\n+\t(\n+\t\tcd max-pack-size &&\n+\t\ttest_commit base &&\n+\t\t# two cruft objects which exceed the maximum pack size\n+\t\ttest-tool genrandom foo 1048576 | git hash-object --stdin -w &&\n+\t\ttest-tool genrandom bar 1048576 | git hash-object --stdin -w &&\n+\t\tgit repack --cruft --max-pack-size=1M &&\n+\t\tfind $packdir -name \"*.mtimes\" >cruft &&\n+\t\ttest_line_count = 1 cruft &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(cat cruft)\")\" >objects &&\n+\t\ttest_line_count = 2 objects\n+\t)\n+'\n+\n+test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n+\t(\n+\t\tcd max-pack-size &&\n+\t\t# repack everything back together to remove the existing cruft\n+\t\t# pack (but to keep its objects)\n+\t\tgit repack -adk &&\n+\t\tgit -c pack.packSizeLimit=1M repack --cruft &&\n+\t\t# ensure the same post condition is met when --max-pack-size\n+\t\t# would otherwise be inferred from the configuration\n+\t\tfind $packdir -name \"*.mtimes\" >cruft &&\n+\t\ttest_line_count = 1 cruft &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(cat cruft)\")\" >objects &&\n+\t\ttest_line_count = 2 objects\n+\t)\n+'\n+\n test_done\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442529","messageId":"b2937ceda7ad1180f46b478f3d7c5a990e00c9cd.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 15/17] builtin/repack.c: add cruft packs to MIDX during geometric repack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:40Z","receivedAt":"2021-11-29T22:26:29Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When using cruft packs, the following race can occur when a geometric\nrepack that writes a MIDX bitmap takes place afterwords:\n\n  - First, create an unreachable object and do an all-into-one cruft\n    repack which stores that object in the repository's cruft pack.\n  - Then make that object reachable.\n  - Finally, do a geometric repack and write a MIDX bitmap.\n\nAssuming that we are sufficiently unlucky as to select a commit from the\nMIDX which reaches that object for bitmapping, then the `git\nmulti-pack-index` process will complain that that object is missing.\n\nThe reason is because we don't include cruft packs in the MIDX when\ndoing a geometric repack. Since the \"make that object reachable\" doesn't\nnecessarily mean that we'll create a new copy of that object in one of\nthe packs that will get rolled up as part of a geometric repack, it's\npossible that the MIDX won't see any copies of that now-reachable\nobject.\n\nOf course, it's desirable to avoid including cruft packs in the MIDX\nbecause it causes the MIDX to store a bunch of objects which are likely\nto get thrown away. But excluding that pack does open us up to the above\nrace.\n\nThis patch demonstrates the bug, and resolves it by including cruft\npacks in the MIDX even when doing a geometric repack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/repack.c              | 19 +++++++++++++++++--\n t/t5327-pack-objects-cruft.sh | 26 ++++++++++++++++++++++++++\n 2 files changed, 43 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex cd4d789d27..5a201063e7 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -23,6 +23,7 @@\n #define PACK_CRUFT 4\n \n #define DELETE_PACK 1\n+#define CRUFT_PACK 2\n \n static int pack_everything;\n static int delta_base_offset = 1;\n@@ -158,8 +159,11 @@ static void collect_pack_filenames(struct string_list *fname_nonkept_list,\n \t\tif ((extra_keep->nr > 0 && i < extra_keep->nr) ||\n \t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname))))\n \t\t\tstring_list_append_nodup(fname_kept_list, fname);\n-\t\telse\n-\t\t\tstring_list_append_nodup(fname_nonkept_list, fname);\n+\t\telse {\n+\t\t\tstruct string_list_item *item = string_list_append_nodup(fname_nonkept_list, fname);\n+\t\t\tif (file_exists(mkpath(\"%s/%s.mtimes\", packdir, fname)))\n+\t\t\t\titem->util = (void*)(uintptr_t)CRUFT_PACK;\n+\t\t}\n \t}\n \tclosedir(dir);\n }\n@@ -559,6 +563,17 @@ static void midx_included_packs(struct string_list *include,\n \n \t\t\tstring_list_insert(include, strbuf_detach(&buf, NULL));\n \t\t}\n+\n+\t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n+\t\t\tif (!((uintptr_t)item->util & CRUFT_PACK)) {\n+\t\t\t\t/*\n+\t\t\t\t * no need to check DELETE_PACK, since we're not\n+\t\t\t\t * doing an ALL_INTO_ONE repack\n+\t\t\t\t */\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tstring_list_insert(include, xstrfmt(\"%s.idx\", item->string));\n+\t\t}\n \t} else {\n \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n \t\t\tif ((uintptr_t)item->util & DELETE_PACK)\ndiff --git a/t/t5327-pack-objects-cruft.sh b/t/t5327-pack-objects-cruft.sh\nindex 750e9d6d6f..857f9e8855 100755\n--- a/t/t5327-pack-objects-cruft.sh\n+++ b/t/t5327-pack-objects-cruft.sh\n@@ -594,4 +594,30 @@ test_expect_success 'cruft --local drops unreachable objects' '\n \t)\n '\n \n+test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\ttest_commit cruft &&\n+\t\tunreachable=\"$(git rev-parse cruft)\" &&\n+\n+\t\tgit reset --hard $unreachable^ &&\n+\t\tgit tag -d cruft &&\n+\t\trm -fr .git/logs &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\t# resurrect the unreachable object via a new commit. the\n+\t\t# new commit will get selected for a bitmap, but be\n+\t\t# missing one of its parents from the selected packs.\n+\t\tgit reset --hard $unreachable &&\n+\t\ttest_commit resurrect &&\n+\n+\t\tgit repack --write-midx --write-bitmap-index --geometric=2 -d\n+\t)\n+'\n+\n test_done\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442530","messageId":"394de0199f9543208438b18a2c26b92c99f36d11.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 16/17] builtin/gc.c: conditionally avoid pruning objects via loose","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:43Z","receivedAt":"2021-11-29T22:26:31Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Expose the new `git repack --cruft` mode from `git gc` via a new opt-in\nflag. When invoked like `git gc --cruft`, `git gc` will avoid exploding\nunreachable objects as loose ones, and instead create a cruft pack and\n`.mtimes` file.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/gc.txt   | 21 +++++++++++++-------\n Documentation/git-gc.txt      |  5 +++++\n builtin/gc.c                  | 10 +++++++++-\n t/t5327-pack-objects-cruft.sh | 37 +++++++++++++++++++++++++++++++++++\n 4 files changed, 65 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/config/gc.txt b/Documentation/config/gc.txt\nindex c834e07991..38fea076a2 100644\n--- a/Documentation/config/gc.txt\n+++ b/Documentation/config/gc.txt\n@@ -81,14 +81,21 @@ gc.packRefs::\n \tto enable it within all non-bare repos or it can be set to a\n \tboolean value.  The default is `true`.\n \n+gc.cruftPacks::\n+\tStore unreachable objects in a cruft pack (see\n+\tlinkgit:git-repack[1]) instead of as loose objects. The default\n+\tis `false`.\n+\n gc.pruneExpire::\n-\tWhen 'git gc' is run, it will call 'prune --expire 2.weeks.ago'.\n-\tOverride the grace period with this config variable.  The value\n-\t\"now\" may be used to disable this grace period and always prune\n-\tunreachable objects immediately, or \"never\" may be used to\n-\tsuppress pruning.  This feature helps prevent corruption when\n-\t'git gc' runs concurrently with another process writing to the\n-\trepository; see the \"NOTES\" section of linkgit:git-gc[1].\n+\tWhen 'git gc' is run, it will call 'prune --expire 2.weeks.ago'\n+\t(and 'repack --cruft --cruft-expiration 2.weeks.ago' if using\n+\tcruft packs via `gc.cruftPacks` or `--cruft`).  Override the\n+\tgrace period with this config variable.  The value \"now\" may be\n+\tused to disable this grace period and always prune unreachable\n+\tobjects immediately, or \"never\" may be used to suppress pruning.\n+\tThis feature helps prevent corruption when 'git gc' runs\n+\tconcurrently with another process writing to the repository; see\n+\tthe \"NOTES\" section of linkgit:git-gc[1].\n \n gc.worktreePruneExpire::\n \tWhen 'git gc' is run, it calls\ndiff --git a/Documentation/git-gc.txt b/Documentation/git-gc.txt\nindex 853967dea0..ba4e67700e 100644\n--- a/Documentation/git-gc.txt\n+++ b/Documentation/git-gc.txt\n@@ -54,6 +54,11 @@ other housekeeping tasks (e.g. rerere, working trees, reflog...) will\n be performed as well.\n \n \n+--cruft::\n+\tWhen expiring unreachable objects, pack them separately into a\n+\tcruft pack instead of storing the loose objects as loose\n+\tobjects.\n+\n --prune=<date>::\n \tPrune loose objects older than date (default is 2 weeks ago,\n \toverridable by the config variable `gc.pruneExpire`).\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex bcef6a4c8d..c16cef0285 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -42,6 +42,7 @@ static const char * const builtin_gc_usage[] = {\n \n static int pack_refs = 1;\n static int prune_reflogs = 1;\n+static int cruft_packs = 0;\n static int aggressive_depth = 50;\n static int aggressive_window = 250;\n static int gc_auto_threshold = 6700;\n@@ -152,6 +153,7 @@ static void gc_config(void)\n \tgit_config_get_int(\"gc.auto\", &gc_auto_threshold);\n \tgit_config_get_int(\"gc.autopacklimit\", &gc_auto_pack_limit);\n \tgit_config_get_bool(\"gc.autodetach\", &detach_auto);\n+\tgit_config_get_bool(\"gc.cruftpacks\", &cruft_packs);\n \tgit_config_get_expiry(\"gc.pruneexpire\", &prune_expire);\n \tgit_config_get_expiry(\"gc.worktreepruneexpire\", &prune_worktrees_expire);\n \tgit_config_get_expiry(\"gc.logexpiry\", &gc_log_expire);\n@@ -331,7 +333,11 @@ static void add_repack_all_option(struct string_list *keep_pack)\n {\n \tif (prune_expire && !strcmp(prune_expire, \"now\"))\n \t\tstrvec_push(&repack, \"-a\");\n-\telse {\n+\telse if (cruft_packs) {\n+\t\tstrvec_push(&repack, \"--cruft\");\n+\t\tif (prune_expire)\n+\t\t\tstrvec_pushf(&repack, \"--cruft-expiration=%s\", prune_expire);\n+\t} else {\n \t\tstrvec_push(&repack, \"-A\");\n \t\tif (prune_expire)\n \t\t\tstrvec_pushf(&repack, \"--unpack-unreachable=%s\", prune_expire);\n@@ -550,6 +556,7 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t{ OPTION_STRING, 0, \"prune\", &prune_expire, N_(\"date\"),\n \t\t\tN_(\"prune unreferenced objects\"),\n \t\t\tPARSE_OPT_OPTARG, NULL, (intptr_t)prune_expire },\n+\t\tOPT_BOOL(0, \"cruft\", &cruft_packs, N_(\"pack unreferenced objects separately\")),\n \t\tOPT_BOOL(0, \"aggressive\", &aggressive, N_(\"be more thorough (increased runtime)\")),\n \t\tOPT_BOOL_F(0, \"auto\", &auto_gc, N_(\"enable auto-gc mode\"),\n \t\t\t   PARSE_OPT_NOCOMPLETE),\n@@ -668,6 +675,7 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t\tdie(FAILED_RUN, repack.v[0]);\n \n \t\tif (prune_expire) {\n+\t\t\t/* run `git prune` even if using cruft packs */\n \t\t\tstrvec_push(&prune, prune_expire);\n \t\t\tif (quiet)\n \t\t\t\tstrvec_push(&prune, \"--no-progress\");\ndiff --git a/t/t5327-pack-objects-cruft.sh b/t/t5327-pack-objects-cruft.sh\nindex 857f9e8855..4cd0f0cf57 100755\n--- a/t/t5327-pack-objects-cruft.sh\n+++ b/t/t5327-pack-objects-cruft.sh\n@@ -429,6 +429,43 @@ test_expect_success 'loose objects mtimes upsert others' '\n \t)\n '\n \n+test_expect_success 'expiring cruft objects with git gc' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit branch -M main &&\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit unreachable &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d unreachable &&\n+\t\t# objects are not cruft if they are contained in the reflogs\n+\t\trm -fr .git/logs &&\n+\n+\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\t\tsort <reachable.raw >reachable &&\n+\t\tcomm -13 reachable objects >unreachable &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tmtimes=$(ls .git/objects/pack/pack-*.mtimes) &&\n+\t\ttest_path_is_file $mtimes &&\n+\n+\t\tgit gc --cruft --prune=now &&\n+\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\n+\t\tcomm -23 unreachable objects >removed &&\n+\t\ttest_cmp unreachable removed &&\n+\t\ttest_path_is_missing $mtimes\n+\t)\n+'\n+\n test_expect_success 'cruft packs are not included in geometric repack' '\n \tgit init repo &&\n \ttest_when_finished \"rm -fr repo\" &&\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442531","messageId":"99aace8e16e4a2b36153e8aca7a6ab518065ff54.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 17/17] sha1-file.c: don't freshen cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:46Z","receivedAt":"2021-11-29T22:26:35Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"We don't bother to freshen objects stored in a cruft pack individually\nby updating the `.mtimes` file. This is because we can't portably `mmap`\nand write into the middle of a file (i.e., to update the mtime of just\none object). Instead, we would have to rewrite the entire `.mtimes` file\nwhich may incur some wasted effort especially if there a lot of cruft\nobjects and they are freshened infrequently.\n\nInstead, force the freshening code to avoid an optimizing write by\nwriting out the object loose and letting it pick up a current mtime.\n\nThis works because we prefer the mtime of the loose copy of an object\nwhen both a loose and packed one exist (whether or not the packed copy\ncomes from a cruft pack or not).\n\nThis could certainly do with a test and/or be included earlier in this\nseries/PR, but I want to wait until after I have a chance to clean up\nthe overly-repetitive nature of the cruft pack tests in general.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n object-file.c                 |  2 ++\n t/t5327-pack-objects-cruft.sh | 25 +++++++++++++++++++++++++\n 2 files changed, 27 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex 7ddb38b64a..dddc1bdd2c 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1946,6 +1946,8 @@ static int freshen_packed_object(const struct object_id *oid)\n \tstruct pack_entry e;\n \tif (!find_pack_entry(the_repository, oid, &e))\n \t\treturn 0;\n+\tif (e.p->is_cruft)\n+\t\treturn 0;\n \tif (e.p->freshened)\n \t\treturn 1;\n \tif (!freshen_file(e.p->pack_name))\ndiff --git a/t/t5327-pack-objects-cruft.sh b/t/t5327-pack-objects-cruft.sh\nindex 4cd0f0cf57..ff87701bbf 100755\n--- a/t/t5327-pack-objects-cruft.sh\n+++ b/t/t5327-pack-objects-cruft.sh\n@@ -657,4 +657,29 @@ test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n \t)\n '\n \n+test_expect_success 'cruft objects are freshend via loose' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\techo \"cruft\" >contents &&\n+\t\tblob=\"$(git hash-object -w -t blob contents)\" &&\n+\t\tloose=\"$objdir/$(test_oid_to_path $blob)\" &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\ttest_path_is_missing \"$loose\" &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(ls $packdir/pack-*.mtimes)\")\" >cruft &&\n+\t\tgrep \"$blob\" cruft &&\n+\n+\t\t# write the same object again\n+\t\tgit hash-object -w -t blob contents &&\n+\n+\t\ttest_path_is_file \"$loose\"\n+\t)\n+'\n+\n test_done\n-- \n2.34.1.25.gb3157a20e6\n"},{"id":"442532","messageId":"fd50c396577dedee00a513a9ba5a684dd78e5274.1638224692.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH 14/17] builtin/repack.c: use named flags for existing_packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-11-29T22:25:38Z","receivedAt":"2021-11-29T22:27:04Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"We use the `util` pointer for items in the `existing_packs` string list\nto indicate which packs are going to be deleted. Since that has so far\nbeen the only use of that `util` pointer, we just set it to 0 or 1.\n\nBut we're going to add an additional state to this field in the next\npatch, so prepare for that by adding a #define for the first bit so we\ncan more expressively inspect the flags state.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/repack.c | 9 ++++++---\n 1 file changed, 6 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex cefa906344..cd4d789d27 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -22,6 +22,8 @@\n #define LOOSEN_UNREACHABLE 2\n #define PACK_CRUFT 4\n \n+#define DELETE_PACK 1\n+\n static int pack_everything;\n static int delta_base_offset = 1;\n static int pack_kept_objects = -1;\n@@ -559,7 +561,7 @@ static void midx_included_packs(struct string_list *include,\n \t\t}\n \t} else {\n \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n-\t\t\tif (item->util)\n+\t\t\tif ((uintptr_t)item->util & DELETE_PACK)\n \t\t\t\tcontinue;\n \t\t\tstring_list_insert(include, xstrfmt(\"%s.idx\", item->string));\n \t\t}\n@@ -1002,7 +1004,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t * was given) and that we will actually delete this pack\n \t\t\t * (if `-d` was given).\n \t\t\t */\n-\t\t\titem->util = (void*)(intptr_t)!string_list_has_string(&names, sha1);\n+\t\t\tif (!string_list_has_string(&names, sha1))\n+\t\t\t\titem->util = (void*)(uintptr_t)((size_t)item->util | DELETE_PACK);\n \t\t}\n \t}\n \n@@ -1026,7 +1029,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (delete_redundant) {\n \t\tint opts = 0;\n \t\tfor_each_string_list_item(item, &existing_nonkept_packs) {\n-\t\t\tif (!item->util)\n+\t\t\tif (!((uintptr_t)item->util & DELETE_PACK))\n \t\t\t\tcontinue;\n \t\t\tremove_redundant_pack(packdir, item->string);\n \t\t}\n-- \n2.34.1.25.gb3157a20e6\n\n"},{"id":"442866","messageId":"df8990d5-98da-a05c-31cf-d3f5ce33f498@gmail.com","threadId":"56996","inReplyTo":"a9f7c738e0ffbc5cdedc26768a0623446c98d239.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-02T14:33:51Z","receivedAt":"2021-12-02T14:33:56Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/29/2021 5:25 PM, Taylor Blau wrote:\n> +Notable alternatives to this design include:\n> +\n> +  - The location of the per-object mtime data, and\n> +  - Whether cruft packs should be incremental or not.\n\nIt was not obvious from this sentence that \"incremental\" meant that\nwe could store a number of cruft packs and use the mtime of each pack\nas the time for all contained objects.\n\n> +On the location of mtime data, a new auxiliary file tied to the pack was chosen\n> +to avoid complicating the `.idx` format. If the `.idx` format were ever to gain\n> +support for optional chunks of data, it may make sense to consolidate the\n> +`.mtimes` format into the `.idx` itself.\n> +\n> +Incremental cruft packs (i.e., where each time a repository is repacked a new\n> +cruft pack is generated containing only the unreachable objects introduced since\n> +the last time a cruft pack was written) are significantly more complicated to\n> +construct, and so aren't pursued here. The obvious drawback to the current\n> +implementation is that the entire cruft pack must be re-written from scratch.\n\nBut you seem to be pointing that direction here. The difference being\nthat you don't discuss how a list of cruft packs could avoid the .mtimes\nfile.\n\nI think what is hidden underneath \"significantly more complicated to\nconstruct\" are situations such as \"this object was in an old cruft\npack, but then became reachable, but now is unreachable again\". I'll\ntry to remember to come back to this after seeing the situations you\ncover in your tests.\n\nThanks,\n-Stolee\n"},{"id":"442873","messageId":"ef10c824-e2d9-f113-f010-6a1ac307427a@gmail.com","threadId":"56996","inReplyTo":"7d4ae7bd3e28e2ec904abb37b6f26505e37531c5.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 02/17] pack-mtimes: support reading .mtimes files","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-02T15:06:07Z","receivedAt":"2021-12-02T15:06:37Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/29/2021 5:25 PM, Taylor Blau wrote:\n\n> +== pack-*.mtimes files have the format:\n> +\n> +  - A 4-byte magic number '0x4d544d45' ('MTME').\n> +\n> +  - A 4-byte version identifier (= 1).\n> +\n> +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n\nI vaguely remember complaints about using a 1-byte identifier in\nthe commit-graph and multi-pack-index formats because the \"standard\"\nway to refer to these hash functions was a magic number that had a\nmeaning in ASCII that helped human readers a bit. I cannot find an\nexample of such 4-byte identifiers, but perhaps brian (CC'd) could\nremind us.\n\nYou are using a 4-byte identifier, but using the same values as\nthose 1-byte identifiers.\n\n> +  - A table of mtimes (one per packed object, num_objects in total, each\n> +    a 4-byte unsigned integer in network order), in the same order as\n> +    objects appear in the index file (e.g., the first entry in the mtime\n> +    table corresponds to the object with the lowest lexically-sorted\n> +    oid). The mtimes count standard epoch seconds.\n\nThis paragraph seemed awkward. Here is a rephrasing that might be\nless awkward:\n\n - A table of 4-byte unsigned integers in network order. The ith value\n   is the modified time (mtime) of the ith object of the corresponding\n   pack in lexicographic order. The mtime represents standard epoch\n   seconds.\n\nStoring these mtimes in 32-bits means we will hit the 2038 problem.\nThe commit-graph stores commit times with an extra two bits to extend\nthe lifetime by another hundred years or so.\n\nCould we extend the lifetime of cruft packs by decreasing the granularity\nhere? Should 'mtime' store a number of _minutes_ instead of seconds? That\nshould be enough granularity for these purposes.\n\n> +  - A trailer, containing a:\n> +\n> +    checksum of the corresponding packfile, and\n> +\n> +    a checksum of all of the above.\n\nCould you specify the checksum as having length according to the\nspecified hash function?\n\n> +All 4-byte numbers are in network order.\n> +\n\nMaybe this could be at the start of the format, since the file\nversion and hash function are both 4-byte numbers here and we\ncould remove the mention of network order from the mtime values.\n\n> +static char *pack_mtimes_filename(struct packed_git *p)\n> +{\n> +\tsize_t len;\n> +\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n> +\t\tBUG(\"pack_name does not end in .pack\");\n> +\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n> +\treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n> +}\n\nI see your NEEDSWORK here and you are probably referring to this:\n\nstatic char *pack_revindex_filename(struct packed_git *p)\n{\n\tsize_t len;\n\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n\t\tBUG(\"pack_name does not end in .pack\");\n\treturn xstrfmt(\"%.*s.rev\", (int)len, p->pack_name);\n}\n\nand the implementation is identical except for the new trailer\n(which exist in the exts[] array in builtin/repack.c, but could\nalso be pulled out into a header somewhere.\n\nI'm happy to delay any cleanup of these code clones until later,\nif at all, because doing it right might mean moving more code\nthan we like. Such refactorings aren't worth it most of the time.\n\n> +static int load_pack_mtimes_file(char *mtimes_file,\n> +\t\t\t\t uint32_t num_objects,\n> +\t\t\t\t const uint32_t **data_p, size_t *len_p)\n> +{\n\n> +\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t), num_objects)) {\n> +\t\tret = error(_(\"mtimes file %s is corrupt\"), mtimes_file);\n\nThis message could be more informative: \"mtimes file %s has the wrong size\"?\n\n> +\tdata = hdr = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n> +\n> +\tif (ntohl(*hdr) != MTIMES_SIGNATURE) {\n> +\t\tret = error(_(\"mtimes file %s has unknown signature\"), mtimes_file);\n> +\t\tgoto cleanup;\n> +\t}\n\nInteresting that you defined 'struct mtimes_header' before this\nmethod, but don't use it here (in favor of moving a uint32_t\npointer). Perhaps you are avoiding pointing the struct at the\nmemory map, but you could also do this:\n\n\tstruct mtimes_header header;\n\n\theader.signature = ntohl(hdr[0]);\n\theader.version = ntohl(hdr[1]);\n\theader.hash_id = ntohl(hdr[2]);\n\nAnd then operate on the struct for your validation.\n\nAt the very least, 'struct mtimes_header' is defined but not\nused in this patch. If you decide to not use it this way, then\nmaybe delay its definition.\n\n> +\n> +\tif (ntohl(*++hdr) != 1) {\n> +\t\tret = error(_(\"mtimes file %s has unsupported version %\"PRIu32),\n> +\t\t\t    mtimes_file, ntohl(*hdr));\n\nUnlike the commit-graph, if we don't understand the version we\ncannot simply ignore the data. error() is appropriate here.\n\n> +int load_pack_mtimes(struct packed_git *p)\n> +{\n> +\tchar *mtimes_name = NULL;\n> +\tint ret = 0;\n> +\n> +\tif (!p->is_cruft)\n> +\t\treturn ret; /* not a cruft pack */\n\nInteresting that this indicator is essentially \"we have an mtimes\nfile for this pack\", but it makes sense to include that check next\nto the .keep and .promisor checks.\n\n> +uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos)\n> +{\n> +\tif (!p->mtimes_map)\n> +\t\tBUG(\"pack .mtimes file not loaded for %s\", p->pack_name);\n> +\tif (p->num_objects <= pos)\n> +\t\tBUG(\"pack .mtimes out-of-bounds (%\"PRIu32\" vs %\"PRIu32\")\",\n> +\t\t    pos, p->num_objects);\n> +\n> +\treturn get_be32(p->mtimes_map + pos + 3);\n> +}\n\nA nice safe access method. Good.\n\n> -\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\"};\n> +\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\", \".mtimes\"};\n\n(Speaking of that refactoring earlier, here is a second definition of\nexts[] that would be valuable to unify.)\n\nThe hunks I did not comment on look good. Nice standard file format\nstuff.\n\nThanks,\n-Stolee\n"},{"id":"442874","messageId":"2d5456f6-5a4d-1600-83f3-2b6d3e1b270a@gmail.com","threadId":"56996","inReplyTo":"ea245b7216067093fdd3a5b2e3a9390f634c8af0.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 04/17] chunk-format.h: extract oid_version()","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-02T15:22:05Z","receivedAt":"2021-12-02T15:22:09Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/29/2021 5:25 PM, Taylor Blau wrote:\n> There are three definitions of an identical function which converts\n> `the_hash_algo` into either 1 (for SHA-1) or 2 (for SHA-256). There is a\n> copy of this function for writing both the commit-graph and\n> multi-pack-index file, and another inline definition used to write the\n> .rev header.\n> \n> Consolidate these into a single definition in chunk-format.h. It's not\n> clear that this is the best header to define this function in, but it\n> should do for now.\n\nThanks for consolidating these!\n \n> (Worth noting, the .rev caller expects a 4-byte unsigned, but the other\n> two callers work with a single unsigned byte. The consolidated version\n> uses the latter type, and lets the compiler widen it when required).\n> \n> Another caller will be added in a subsequent patch.\n\n>  chunk-format.c | 12 ++++++++++++\n>  chunk-format.h |  3 +++\n>  commit-graph.c | 18 +++---------------\n>  midx.c         | 18 +++---------------\n>  pack-write.c   | 15 ++-------------\n\nI notice that you don't use this in load_pack_mtimes_file(),\nin pack-mtimes.c but you could at this point.\n\nThe code you do touch looks good.\n\nThanks,\n-Stolee\n"},{"id":"442876","messageId":"2a26da15-8d4f-92b9-d727-debf8b969899@gmail.com","threadId":"56996","inReplyTo":"deece9eb70e9750bb8350946679b521e59139fe2.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 05/17] pack-mtimes: support writing pack .mtimes files","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-02T15:36:16Z","receivedAt":"2021-12-02T15:36:20Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/29/2021 5:25 PM, Taylor Blau wrote:> @@ -168,6 +168,9 @@ struct packing_data {\n>  \t/* delta islands */\n>  \tunsigned int *tree_depth;\n>  \tunsigned char *layer;\n> +\n> +\t/* cruft packs */\n> +\tuint32_t *cruft_mtime;\n\nThis comment is a bit terse. Perhaps...\n\n\t/* Used when writing cruft packs. */\n\n> +static inline uint32_t oe_cruft_mtime(struct packing_data *pack,\n> +\t\t\t\t      struct object_entry *e)\n> +{\n> +\tif (!pack->cruft_mtime)\n> +\t\treturn 0;\n> +\treturn pack->cruft_mtime[e - pack->objects];\n> +}\n\nWhen writing a pack, it appears that the cruft_mtime array\nmaps to objects in pack-order, not idx-order, correct? That\nmight be worth mentioning in the struct definition because\nit differs from the .mtimes file.\n\n> +static void write_mtimes_objects(struct hashfile *f,\n> +\t\t\t\t struct packing_data *to_pack,\n> +\t\t\t\t struct pack_idx_entry **objects,\n> +\t\t\t\t uint32_t nr_objects)\n> +{\n> +\tuint32_t i;\n> +\tfor (i = 0; i < nr_objects; i++) {\n> +\t\tstruct object_entry *e = (struct object_entry*)objects[i];\n> +\t\thashwrite_be32(f, oe_cruft_mtime(to_pack, e));\n> +\t}\n\nThe name \"objects\" here confused me at first, thinking it\ncorresponded to the objects member of 'struct packing_data', but\nthat is being handled by the fact that 'objects' is actually a\nlex-sorted list of pack_idx_entry pointers (and they happen to\nalso point to 'struct object_entry' values because the 'struct\npack_idx_entry' is the first member.\n\nSo this is (very densely) handling the translation from pack-order\nto lex-order through the double pointer 'objects'. I'm not sure if\nthere is a way to make it more clear or if every reader will need\nto do the same mental gymnastics I had to do.\n\n> +}\n> +\n> +static void write_mtimes_trailer(struct hashfile *f, const unsigned char *hash)\n> +{\n> +\thashwrite(f, hash, the_hash_algo->rawsz);\n> +}\n> +\n> +static const char *write_mtimes_file(const char *mtimes_name,\n> +\t\t\t\t     struct packing_data *to_pack,\n> +\t\t\t\t     struct pack_idx_entry **objects,\n> +\t\t\t\t     uint32_t nr_objects,\n> +\t\t\t\t     const unsigned char *hash)\n> +{\n> +\tstruct hashfile *f;\n> +\tint fd;\n> +\n> +\tif (!to_pack)\n> +\t\tBUG(\"cannot call write_mtimes_file with NULL packing_data\");\n> +\n> +\tif (!mtimes_name) {\n> +\t\tstruct strbuf tmp_file = STRBUF_INIT;\n> +\t\tfd = odb_mkstemp(&tmp_file, \"pack/tmp_mtimes_XXXXXX\");\n> +\t\tmtimes_name = strbuf_detach(&tmp_file, NULL);\n> +\t} else {\n> +\t\tunlink(mtimes_name);\n> +\t\tfd = xopen(mtimes_name, O_CREAT|O_EXCL|O_WRONLY, 0600);\n> +\t}\n> +\tf = hashfd(fd, mtimes_name);\n> +\n> +\twrite_mtimes_header(f);\n> +\twrite_mtimes_objects(f, to_pack, objects, nr_objects);\n> +\twrite_mtimes_trailer(f, hash);\n> +\n> +\tif (mtimes_name && adjust_shared_perm(mtimes_name) < 0)\n> +\t\tdie(_(\"failed to make %s readable\"), mtimes_name);\n\nWhat could cause 'mtimes_name' to be NULL here? It seems that it would\nbe initialized in the \"if (!mtimes_name)\" block above.\n\n> +\n> +\tfinalize_hashfile(f, NULL,\n> +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE | CSUM_FSYNC);\n> +\n> +\treturn mtimes_name;\n\nNote that you return the name here...\n\n> +\tif (pack_idx_opts->flags & WRITE_MTIMES) {\n> +\t\tmtimes_tmp_name = write_mtimes_file(NULL, to_pack, written_list,\n> +\t\t\t\t\t\t    nr_written,\n> +\t\t\t\t\t\t    hash);\n> +\t\tif (adjust_shared_perm(mtimes_tmp_name))\n> +\t\t\tdie_errno(\"unable to make temporary mtimes file readable\");\n\n...and then adjust the perms again. I think that this adjustment is\nredundant, because it already happened within the write_mtimes_file()\nmethod.\n\n> +\t}\n> +\n>  \trename_tmp_packfile(name_buffer, pack_tmp_name, \"pack\");\n>  \tif (rev_tmp_name)\n>  \t\trename_tmp_packfile(name_buffer, rev_tmp_name, \"rev\");\n> +\tif (mtimes_tmp_name)\n> +\t\trename_tmp_packfile(name_buffer, mtimes_tmp_name, \"mtimes\");\n\nAnd then it is finally renamed here, if it had a temporary name to\nstart.\n\nThanks,\n-Stolee\n"},{"id":"442917","messageId":"YalJgGJFoGGgCazx@camp.crustytoothpaste.net","threadId":"56996","inReplyTo":"ef10c824-e2d9-f113-f010-6a1ac307427a@gmail.com","subject":"Re: [PATCH 02/17] pack-mtimes: support reading .mtimes files","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2021-12-02T22:32:32Z","receivedAt":"2021-12-02T22:32:37Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2021-12-02 at 15:06:07, Derrick Stolee wrote:\n> On 11/29/2021 5:25 PM, Taylor Blau wrote:\n> \n> > +== pack-*.mtimes files have the format:\n> > +\n> > +  - A 4-byte magic number '0x4d544d45' ('MTME').\n> > +\n> > +  - A 4-byte version identifier (= 1).\n> > +\n> > +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n> \n> I vaguely remember complaints about using a 1-byte identifier in\n> the commit-graph and multi-pack-index formats because the \"standard\"\n> way to refer to these hash functions was a magic number that had a\n> meaning in ASCII that helped human readers a bit. I cannot find an\n> example of such 4-byte identifiers, but perhaps brian (CC'd) could\n> remind us.\n> \n> You are using a 4-byte identifier, but using the same values as\n> those 1-byte identifiers.\n\nThe preferred value is the_hash_algo->format_id.  For SHA-1, that's\n\"sha1\", big-endian (0x73686131) and for SHA-256 it's \"s256\", big-endian\n(0x73323536).\n\nThere's also hash_algo_by_id to turn the format ID into an index into\nthe hash_algos array, but you need to check for GIT_HASH_UNKNOWN (0)\nfirst.\n\nThese will be used in index v3, which I haven't sent out patches for\nyet.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"443020","messageId":"xmqq5ys5sbzc.fsf@gitster.g","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 00/17] cruft packs","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-12-03T19:51:51Z","receivedAt":"2021-12-03T19:52:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> This series implements \"cruft packs\", a pack which stores accumulated\n> unreachable objects, along with a new \".mtimes\" file which tracks each\n> object's last known modification time.\n\nLet me rephrase the above to test my understanding, since I need to\nwrite a summary for the  \"What's cooking\" report.\n\n Instead of leaving unreachable objects in loose form when packing,\n or ejecting them into loose form when repacking, gather them in a\n packfile with an auxiliary file that records the last-use time of\n these objects.\n\nThat way, we do not have to waste so many inodes for loose objects\nthat is not likely to be used, which feels like a win.\n\n>   - The final patch handles object freshening for objects stored in a\n>     cruft pack.\n\nI am not going to read it today, but I think this is the most\ninteresting part of the series.  Instead of using mtime of an\nindividual loose object file, we'd need to record the time of\nlast use for each object in a pack.\n\nStepping back a bit, I do not see how we can get away without doing\nthe same .mtimes file for non-cruft packs.  An object that is in a\nnon-cruft pack may be referenced immediately after the repack that\ncreated the pack, but the ref that was referencing the object may\nhave gone away and now the pack is a month old.  If we were to\nrepack the object, we do not know when was the last time the object\nwas reachable from any of the refs and index entries (collectively\nknown as anchor points).  Of course, recording all mtimes for all\npacked objects all the time would involve quite a lot of overhead.\nI am guessing (I will not spend time today to figure it out myself)\nthat .mtimes update at runtime will happen in-place (i.e. via\nseek(2)+write(2), or pwrite()), and I wonder what the safety concern\nwould be (which is the primary reason why we tend not to do in-place\nupdates but recreate-and-rename updates).\n\nThanks for working on such an interesting topic.\n"},{"id":"443025","messageId":"Yap5INmX2ACfjoda@nand.local","threadId":"56996","inReplyTo":"xmqq5ys5sbzc.fsf@gitster.g","subject":"Re: [PATCH 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-12-03T20:08:00Z","receivedAt":"2021-12-03T20:08:03Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Dec 03, 2021 at 11:51:51AM -0800, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> > This series implements \"cruft packs\", a pack which stores accumulated\n> > unreachable objects, along with a new \".mtimes\" file which tracks each\n> > object's last known modification time.\n>\n> Let me rephrase the above to test my understanding, since I need to\n> write a summary for the  \"What's cooking\" report.\n>\n>  Instead of leaving unreachable objects in loose form when packing,\n>  or ejecting them into loose form when repacking, gather them in a\n>  packfile with an auxiliary file that records the last-use time of\n>  these objects.\n\nExactly. Thanks for such a concise and accurate description of the\ntopic.\n\n> That way, we do not have to waste so many inodes for loose objects\n> that is not likely to be used, which feels like a win.\n\nYes. This had historically been a problem for GitHub. We don't\nautomatically prune unreachable objects during repacking, but sometimes\ncustomers will ask us to do it on their behalf (if, for example, they\naccidentally pushed sensitive information to us, and then force-pushed\nover it).\n\nBut occasionally we'd get bitten by exploding many years of loose\nobjects (because we used to freshen packfiles too aggressively when\nmoving them around).\n\nWe've been running this series in production for the past few months,\nand it's been a huge relief on the folks who typically run these pruning\nGCs.\n\n> >   - The final patch handles object freshening for objects stored in a\n> >     cruft pack.\n>\n> I am not going to read it today, but I think this is the most\n> interesting part of the series.  Instead of using mtime of an\n> individual loose object file, we'd need to record the time of\n> last use for each object in a pack.\n>\n> Stepping back a bit, I do not see how we can get away without doing\n> the same .mtimes file for non-cruft packs.  An object that is in a\n> non-cruft pack may be referenced immediately after the repack that\n> created the pack, but the ref that was referencing the object may\n> have gone away and now the pack is a month old.  If we were to\n> repack the object, we do not know when was the last time the object\n> was reachable from any of the refs and index entries (collectively\n> known as anchor points).\n\nIn that situation, we would use the mtime of the pack which contains\nthat object itself as a proxy (or the mtime of a loose copy of the\nobject, if it is more recent).\n\nThat isn't perfect, as you note, since if the pack isn't otherwise\nfreshened, we'd consider that object to be a month old, even if the\nreference pointing at it was deleted a mere second ago.\n\nI can't recall if Peff and I talked about this off-list, but I have a\nvague sense we probably did (and I forgot the details).\n\n> Of course, recording all mtimes for all\n> packed objects all the time would involve quite a lot of overhead.\n> I am guessing (I will not spend time today to figure it out myself)\n> that .mtimes update at runtime will happen in-place (i.e. via\n> seek(2)+write(2), or pwrite()), and I wonder what the safety concern\n> would be (which is the primary reason why we tend not to do in-place\n> updates but recreate-and-rename updates).\n\nYeah, this series avoids doing an in-place update, and similarly avoids\nrecreating the entire .mtimes file before moving into place. Instead,\nfreshening an object stored in a cruft pack takes place by rewriting a\ncopy of the object loose, since we consider an object's mtime to be the\nmost recent of (a) what's in the .mtimes file, (b) the mtime of the\ncontaining pack, and (c) the mtime of a loose copy (if one exists).\n\nIt can be wasteful, but in practice \"resurrecting\" an object in a cruft\npack is pretty rare, so on balance it ends up costing less work to do.\n\n> Thanks for working on such an interesting topic.\n\nI'm glad to have piqued your interest.\n\nTaylor\n"},{"id":"443030","messageId":"YaqCZ7BPwuMGmkZY@nand.local","threadId":"56996","inReplyTo":"Yap5INmX2ACfjoda@nand.local","subject":"Re: [PATCH 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-12-03T20:47:35Z","receivedAt":"2021-12-03T20:47:38Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Dec 03, 2021 at 03:08:00PM -0500, Taylor Blau wrote:\n> On Fri, Dec 03, 2021 at 11:51:51AM -0800, Junio C Hamano wrote:\n> > Stepping back a bit, I do not see how we can get away without doing\n> > the same .mtimes file for non-cruft packs.  An object that is in a\n> > non-cruft pack may be referenced immediately after the repack that\n> > created the pack, but the ref that was referencing the object may\n> > have gone away and now the pack is a month old.  If we were to\n> > repack the object, we do not know when was the last time the object\n> > was reachable from any of the refs and index entries (collectively\n> > known as anchor points).\n>\n> In that situation, we would use the mtime of the pack which contains\n> that object itself as a proxy (or the mtime of a loose copy of the\n> object, if it is more recent).\n>\n> That isn't perfect, as you note, since if the pack isn't otherwise\n> freshened, we'd consider that object to be a month old, even if the\n> reference pointing at it was deleted a mere second ago.\n>\n> I can't recall if Peff and I talked about this off-list, but I have a\n> vague sense we probably did (and I forgot the details).\n\nMaybe I can rephrase the problem as being orthogonal to what we're\naddressing here. Modification time can be a useful-ish proxy for \"last\nreferenced time\", but they are ultimately different.\n\nForgetting cruft packs for a moment, our behavior today in that\nsituation would be to prune the object if our grace period did not cover\nthe time in which the pack was last modified. So if the pack was a month\nold, the grace period was two weeks, but the reference pointing at some\nobject in that pack was deleted only a second before starting a pruning\nGC, we'd prune that object before this series (just as we would do the\nsame thing with this series).\n\nAside from pruning, what happens to the value recorded in the .mtimes\nfile is more interesting. For the case you're talking about, we'll err\non the side of newer mtimes (either the original timestamp is recorded,\nor some future time when the containing pack was rewritten). But the\nmore interesting case is when an object becomes re-referenced. Since the\nref-update doesn't cause the object to be rewritten, we wouldn't change\nthe timestamp.\n\nAnyway, both of these are still independent from cruft packs, so we're\nnot changing the status quo there, I don't think.\n\nThanks,\nTaylor\n"},{"id":"443040","messageId":"YaqR64y0ddWOhHjP@nand.local","threadId":"56996","inReplyTo":"df8990d5-98da-a05c-31cf-d3f5ce33f498@gmail.com","subject":"Re: [PATCH 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-12-03T21:53:47Z","receivedAt":"2021-12-03T21:53:51Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, Dec 02, 2021 at 09:33:51AM -0500, Derrick Stolee wrote:\n> On 11/29/2021 5:25 PM, Taylor Blau wrote:\n> > +Notable alternatives to this design include:\n> > +\n> > +  - The location of the per-object mtime data, and\n> > +  - Whether cruft packs should be incremental or not.\n>\n> It was not obvious from this sentence that \"incremental\" meant that\n> we could store a number of cruft packs and use the mtime of each pack\n> as the time for all contained objects.\n\nYes, I think I meant \"incremental\" in the sense of \"incremental commit-\ngraphs\". But it's clearer to say \"storing unreachable objects in\nmultiple cruft packs\" (and then giving an example later on). Thanks!\n\n> I think what is hidden underneath \"significantly more complicated to\n> construct\" are situations such as \"this object was in an old cruft\n> pack, but then became reachable, but now is unreachable again\". I'll\n> try to remember to come back to this after seeing the situations you\n> cover in your tests.\n\nYeah, I'm being deliberately vague here, since the aim of this paragraph\nis to illustrate \"this is much more complicated than what we implement\nhere, and the trade-offs are...\"\n\nThanks,\nTaylor\n"},{"id":"443044","messageId":"YaqZA02FsCFA9qBi@nand.local","threadId":"56996","inReplyTo":"ef10c824-e2d9-f113-f010-6a1ac307427a@gmail.com","subject":"Re: [PATCH 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-12-03T22:24:03Z","receivedAt":"2021-12-03T22:24:07Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, Dec 02, 2021 at 10:06:07AM -0500, Derrick Stolee wrote:\n> On 11/29/2021 5:25 PM, Taylor Blau wrote:\n>\n> > +== pack-*.mtimes files have the format:\n> > +\n> > +  - A 4-byte magic number '0x4d544d45' ('MTME').\n> > +\n> > +  - A 4-byte version identifier (= 1).\n> > +\n> > +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n>\n> I vaguely remember complaints about using a 1-byte identifier in\n> the commit-graph and multi-pack-index formats because the \"standard\"\n> way to refer to these hash functions was a magic number that had a\n> meaning in ASCII that helped human readers a bit. I cannot find an\n> example of such 4-byte identifiers, but perhaps brian (CC'd) could\n> remind us.\n>\n> You are using a 4-byte identifier, but using the same values as\n> those 1-byte identifiers.\n\nYeah, I'm definitely borrowing from the commit-graph and multi-pack\nindex formats here. Though I believe we did the same thing for .rev\nfiles, too (and checking with Documentation/technical/pack-format.txt\nconfirms as much).\n\nI don't have a strong feeling about using the 4-byte identifier or not.\nBut making this field four bytes wide is very much intentional, since it\nmakes sure that all of our reads are aligned, which should yield much\nbetter cache performance (assuming the page size is also a multiple of\nfour).\n\nI don't, but if others feel strongly we could write the magic\nidentifiers brian points out downthread here instead. (It would be\nmildly inconvenient for GitHub, which has many hundreds of thousands of\nthese files laying around everywhere with '1' as the identifier. But\nsince the magic identifiers don't collide with the values proposed here,\nGitHub's fork could easily be taught to accept both on the reading side,\nbut only write out the special identifier).\n\n> > +  - A table of mtimes (one per packed object, num_objects in total, each\n> > +    a 4-byte unsigned integer in network order), in the same order as\n> > +    objects appear in the index file (e.g., the first entry in the mtime\n> > +    table corresponds to the object with the lowest lexically-sorted\n> > +    oid). The mtimes count standard epoch seconds.\n>\n> This paragraph seemed awkward. Here is a rephrasing that might be\n> less awkward:\n>\n>  - A table of 4-byte unsigned integers in network order. The ith value\n>    is the modified time (mtime) of the ith object of the corresponding\n>    pack in lexicographic order. The mtime represents standard epoch\n>    seconds.\n\nThanks, this is clearer. I went with a blend of the two:\n\n    - A table of 4-byte unsigned integers in network order. The ith\n      value is the modification time (mtime) of the ith object in the\n      corresponding pack by lexicographic (index) order. The mtimes\n      count standard epoch seconds.\n\n> Storing these mtimes in 32-bits means we will hit the 2038 problem.\n> The commit-graph stores commit times with an extra two bits to extend\n> the lifetime by another hundred years or so.\n>\n> Could we extend the lifetime of cruft packs by decreasing the granularity\n> here? Should 'mtime' store a number of _minutes_ instead of seconds? That\n> should be enough granularity for these purposes.\n\nPerhaps, though it does add some complexity to the code that deals with\nthis format at the expense of some future-proofing. I'm open to it,\nthough.\n\n>\n> > +  - A trailer, containing a:\n> > +\n> > +    checksum of the corresponding packfile, and\n> > +\n> > +    a checksum of all of the above.\n>\n> Could you specify the checksum as having length according to the\n> specified hash function?\n\nGreat suggestion, thanks.\n\n> > +All 4-byte numbers are in network order.\n> > +\n>\n> Maybe this could be at the start of the format, since the file\n> version and hash function are both 4-byte numbers here and we\n> could remove the mention of network order from the mtime values.\n\nThis is copy-and-pasted from the .rev section above, where I think I\nadded the \"All 4-byte numbers are in network order\" bit at the end in\nresponse to a suggestion opposite yours ;).\n\nHere I would probably rather stay consistent with the surrounding\nsections.\n\n> > +static char *pack_mtimes_filename(struct packed_git *p)\n> > +{\n> > +\tsize_t len;\n> > +\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n> > +\t\tBUG(\"pack_name does not end in .pack\");\n> > +\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n> > +\treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n> > +}\n>\n> I see your NEEDSWORK here and you are probably referring to this:\n>\n> static char *pack_revindex_filename(struct packed_git *p)\n> {\n> \tsize_t len;\n> \tif (!strip_suffix(p->pack_name, \".pack\", &len))\n> \t\tBUG(\"pack_name does not end in .pack\");\n> \treturn xstrfmt(\"%.*s.rev\", (int)len, p->pack_name);\n> }\n>\n> and the implementation is identical except for the new trailer\n> (which exist in the exts[] array in builtin/repack.c, but could\n> also be pulled out into a header somewhere.\n>\n> I'm happy to delay any cleanup of these code clones until later,\n> if at all, because doing it right might mean moving more code\n> than we like. Such refactorings aren't worth it most of the time.\n\nYeah, I think your thoughts matched my own when writing this. Which is\nto say, I felt it prudent to call out that there is an opportunity to\nDRY these two up, but I'm not convinced that such a clean up would be\nworthwhile.\n\n> > +static int load_pack_mtimes_file(char *mtimes_file,\n> > +\t\t\t\t uint32_t num_objects,\n> > +\t\t\t\t const uint32_t **data_p, size_t *len_p)\n> > +{\n>\n> > +\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t), num_objects)) {\n> > +\t\tret = error(_(\"mtimes file %s is corrupt\"), mtimes_file);\n>\n> This message could be more informative: \"mtimes file %s has the wrong size\"?\n\nCopy-and-pasting here again from the corresponding code for the .rev\nfile, which is why I didn't opt to change the message here. Probably\nmany of these checks could be extracted out and shared between the two\npaths, but I don't think we should attempt it here.\n\n> > +\tdata = hdr = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n> > +\n> > +\tif (ntohl(*hdr) != MTIMES_SIGNATURE) {\n> > +\t\tret = error(_(\"mtimes file %s has unknown signature\"), mtimes_file);\n> > +\t\tgoto cleanup;\n> > +\t}\n>\n> Interesting that you defined 'struct mtimes_header' before this\n> method, but don't use it here (in favor of moving a uint32_t\n> pointer). Perhaps you are avoiding pointing the struct at the\n> memory map, but you could also do this:\n>\n> \tstruct mtimes_header header;\n>\n> \theader.signature = ntohl(hdr[0]);\n> \theader.version = ntohl(hdr[1]);\n> \theader.hash_id = ntohl(hdr[2]);\n>\n> And then operate on the struct for your validation.\n>\n> At the very least, 'struct mtimes_header' is defined but not\n> used in this patch. If you decide to not use it this way, then\n> maybe delay its definition.\n\nYeah, not reading directly out of the struct is intentional, since the\ncompiler is free to insert padding between these members, which would\nbreak any subsequent reads out of the struct.\n\nBut I like your idea to assign the fields manually, thanks!\n\n> > +int load_pack_mtimes(struct packed_git *p)\n> > +{\n> > +\tchar *mtimes_name = NULL;\n> > +\tint ret = 0;\n> > +\n> > +\tif (!p->is_cruft)\n> > +\t\treturn ret; /* not a cruft pack */\n>\n> Interesting that this indicator is essentially \"we have an mtimes\n> file for this pack\", but it makes sense to include that check next\n> to the .keep and .promisor checks.\n\nI think I had originally called it \"mtimes\" but changed it to \"cruft\",\nsince it makes sense as a prefix similar to the others (that is, \"keep\npack\", \"promisor pack\", and \"cruft pack\", not \"mtimes pack\").\n\n> The hunks I did not comment on look good. Nice standard file format\n> stuff.\n\nThanks for your review!\n\nThanks,\nTaylor\n"},{"id":"443047","messageId":"Yaqc6FYo1pLhsNSB@nand.local","threadId":"56996","inReplyTo":"2d5456f6-5a4d-1600-83f3-2b6d3e1b270a@gmail.com","subject":"Re: [PATCH 04/17] chunk-format.h: extract oid_version()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-12-03T22:40:40Z","receivedAt":"2021-12-03T22:40:43Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, Dec 02, 2021 at 10:22:05AM -0500, Derrick Stolee wrote:\n> I notice that you don't use this in load_pack_mtimes_file(),\n> in pack-mtimes.c but you could at this point.\n\nHmm, I'm confused. Te extracted function converts a pointer to a struct\ngit_hash_algo into a uint32, but here we just care about reading the\nfour byte value we wrote.\n\nThanks,\nTaylor\n"},{"id":"443048","messageId":"YaqiYGM48p5F9lS1@nand.local","threadId":"56996","inReplyTo":"2a26da15-8d4f-92b9-d727-debf8b969899@gmail.com","subject":"Re: [PATCH 05/17] pack-mtimes: support writing pack .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-12-03T23:04:00Z","receivedAt":"2021-12-03T23:04:03Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, Dec 02, 2021 at 10:36:16AM -0500, Derrick Stolee wrote:\n> On 11/29/2021 5:25 PM, Taylor Blau wrote:> @@ -168,6 +168,9 @@ struct packing_data {\n> >  \t/* delta islands */\n> >  \tunsigned int *tree_depth;\n> >  \tunsigned char *layer;\n> > +\n> > +\t/* cruft packs */\n> > +\tuint32_t *cruft_mtime;\n>\n> This comment is a bit terse. Perhaps...\n>\n> \t/* Used when writing cruft packs. */\n\nSure; here I was imitating the terseness of the \"delta islands\" comment\na few lines above. But I don't mind changing it here.\n\n> > +static inline uint32_t oe_cruft_mtime(struct packing_data *pack,\n> > +\t\t\t\t      struct object_entry *e)\n> > +{\n> > +\tif (!pack->cruft_mtime)\n> > +\t\treturn 0;\n> > +\treturn pack->cruft_mtime[e - pack->objects];\n> > +}\n>\n> When writing a pack, it appears that the cruft_mtime array\n> maps to objects in pack-order, not idx-order, correct? That\n> might be worth mentioning in the struct definition because\n> it differs from the .mtimes file.\n\nGreat observation and suggestion, thank you! The comment that I\nultimately settled on is:\n\n  /*\n   * Used when writing cruft packs.\n   *\n   * Object mtimes  are stored in pack order when writing, but\n   * written out in lexicographic (index) order.\n   */\n   uint32_t *cruft_mtime;\n\n> > +static void write_mtimes_objects(struct hashfile *f,\n> > +\t\t\t\t struct packing_data *to_pack,\n> > +\t\t\t\t struct pack_idx_entry **objects,\n> > +\t\t\t\t uint32_t nr_objects)\n> > +{\n> > +\tuint32_t i;\n> > +\tfor (i = 0; i < nr_objects; i++) {\n> > +\t\tstruct object_entry *e = (struct object_entry*)objects[i];\n> > +\t\thashwrite_be32(f, oe_cruft_mtime(to_pack, e));\n> > +\t}\n>\n> The name \"objects\" here confused me at first, thinking it\n> corresponded to the objects member of 'struct packing_data', but\n> that is being handled by the fact that 'objects' is actually a\n> lex-sorted list of pack_idx_entry pointers (and they happen to\n> also point to 'struct object_entry' values because the 'struct\n> pack_idx_entry' is the first member.\n>\n> So this is (very densely) handling the translation from pack-order\n> to lex-order through the double pointer 'objects'. I'm not sure if\n> there is a way to make it more clear or if every reader will need\n> to do the same mental gymnastics I had to do.\n\nExactly, and sorry that I didn't point this out more clearly. It's been\nlong enough since I wrote this code that I can sympathize with the\nmental gymnastics required ;).\n\n> > +}\n> > +\n> > +static void write_mtimes_trailer(struct hashfile *f, const unsigned char *hash)\n> > +{\n> > +\thashwrite(f, hash, the_hash_algo->rawsz);\n> > +}\n> > +\n> > +static const char *write_mtimes_file(const char *mtimes_name,\n> > +\t\t\t\t     struct packing_data *to_pack,\n> > +\t\t\t\t     struct pack_idx_entry **objects,\n> > +\t\t\t\t     uint32_t nr_objects,\n> > +\t\t\t\t     const unsigned char *hash)\n> > +{\n> > +\tstruct hashfile *f;\n> > +\tint fd;\n> > +\n> > +\tif (!to_pack)\n> > +\t\tBUG(\"cannot call write_mtimes_file with NULL packing_data\");\n> > +\n> > +\tif (!mtimes_name) {\n> > +\t\tstruct strbuf tmp_file = STRBUF_INIT;\n> > +\t\tfd = odb_mkstemp(&tmp_file, \"pack/tmp_mtimes_XXXXXX\");\n> > +\t\tmtimes_name = strbuf_detach(&tmp_file, NULL);\n> > +\t} else {\n> > +\t\tunlink(mtimes_name);\n> > +\t\tfd = xopen(mtimes_name, O_CREAT|O_EXCL|O_WRONLY, 0600);\n> > +\t}\n> > +\tf = hashfd(fd, mtimes_name);\n> > +\n> > +\twrite_mtimes_header(f);\n> > +\twrite_mtimes_objects(f, to_pack, objects, nr_objects);\n> > +\twrite_mtimes_trailer(f, hash);\n> > +\n> > +\tif (mtimes_name && adjust_shared_perm(mtimes_name) < 0)\n> > +\t\tdie(_(\"failed to make %s readable\"), mtimes_name);\n>\n> What could cause 'mtimes_name' to be NULL here? It seems that it would\n> be initialized in the \"if (!mtimes_name)\" block above.\n\nYou're right, it's impossible for it to be NULL here. I'll remove the\nredundant side of the &&-expression here.\n\n> > +\n> > +\tfinalize_hashfile(f, NULL,\n> > +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE | CSUM_FSYNC);\n> > +\n> > +\treturn mtimes_name;\n>\n> Note that you return the name here...\n>\n> > +\tif (pack_idx_opts->flags & WRITE_MTIMES) {\n> > +\t\tmtimes_tmp_name = write_mtimes_file(NULL, to_pack, written_list,\n> > +\t\t\t\t\t\t    nr_written,\n> > +\t\t\t\t\t\t    hash);\n> > +\t\tif (adjust_shared_perm(mtimes_tmp_name))\n> > +\t\t\tdie_errno(\"unable to make temporary mtimes file readable\");\n>\n> ...and then adjust the perms again. I think that this adjustment is\n> redundant, because it already happened within the write_mtimes_file()\n> method.\n\nYep, thanks. I'll clean it up here to just call adjust_shared_perm()\nwitin write_mtimes_file().\n\nThanks,\nTaylor\n"},{"id":"443092","messageId":"CABPp-BGW3t1LnUGNSLSzYtQfAJuAQKxpHJMJOdh5T2pUyaDWAw@mail.gmail.com","threadId":"56996","inReplyTo":"a9f7c738e0ffbc5cdedc26768a0623446c98d239.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2021-12-04T22:20:23Z","receivedAt":"2021-12-04T22:20:37Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Mon, Nov 29, 2021 at 7:29 PM Taylor Blau <me@ttaylorr.com> wrote:\n>\n> Create a technical document to explain cruft packs. It contains a brief\n> overview of the problem, some background, details on the implementation,\n> and a couple of alternative approaches not considered here.\n>\n> Signed-off-by: Taylor Blau <me@ttaylorr.com>\n> ---\n>  Documentation/Makefile                  |  1 +\n>  Documentation/technical/cruft-packs.txt | 95 +++++++++++++++++++++++++\n>  2 files changed, 96 insertions(+)\n>  create mode 100644 Documentation/technical/cruft-packs.txt\n>\n> diff --git a/Documentation/Makefile b/Documentation/Makefile\n> index ed656db2ae..0b01c9408e 100644\n> --- a/Documentation/Makefile\n> +++ b/Documentation/Makefile\n> @@ -91,6 +91,7 @@ TECH_DOCS += MyFirstContribution\n>  TECH_DOCS += MyFirstObjectWalk\n>  TECH_DOCS += SubmittingPatches\n>  TECH_DOCS += technical/bundle-format\n> +TECH_DOCS += technical/cruft-packs\n>  TECH_DOCS += technical/hash-function-transition\n>  TECH_DOCS += technical/http-protocol\n>  TECH_DOCS += technical/index-format\n> diff --git a/Documentation/technical/cruft-packs.txt b/Documentation/technical/cruft-packs.txt\n> new file mode 100644\n> index 0000000000..bb54cce1b1\n> --- /dev/null\n> +++ b/Documentation/technical/cruft-packs.txt\n> @@ -0,0 +1,95 @@\n> += Cruft packs\n> +\n> +Cruft packs offer an alternative to Git's traditional mechanism of removing\n> +unreachable objects. This document provides an overview of Git's pruning\n> +mechanism, and how cruft packs can be used instead to accomplish the same.\n> +\n> +== Background\n> +\n> +To remove unreachable objects from your repository, Git offers `git repack -Ad`\n> +(see linkgit:git-repack[1]). Quoting from the documentation:\n> +\n> +[quote]\n> +[...] unreachable objects in a previous pack become loose, unpacked objects,\n> +instead of being left in the old pack. [...] loose unreachable objects will be\n> +pruned according to normal expiry rules with the next 'git gc' invocation.\n> +\n> +Unreachable objects aren't removed immediately, since doing so could race with\n> +an incoming push which may reference an object which is about to be deleted.\n> +Instead, those unreachable objects are stored as loose object and stay that way\n> +until they are older than the expiration window, at which point they are removed\n> +by linkgit:git-prune[1].\n> +\n> +Git must store these unreachable objects loose in order to keep track of their\n> +per-object mtimes. If these unreachable objects were written into one big pack,\n> +then either freshening that pack (because an object contained within it was\n> +re-written) or creating a new pack of unreachable objects would cause the pack's\n> +mtime to get updated, and the objects within it would never leave the expiration\n> +window. Instead, objects are stored loose in order to keep track of the\n> +individual object mtimes and avoid a situation where all cruft objects are\n> +freshened at once.\n> +\n> +This can lead to undesirable situations when a repository contains many\n> +unreachable objects which have not yet left the grace period. Having large\n> +directories in the shards of `.git/objects` can lead to decreased performance in\n> +the repository. But given enough unreachable objects, this can lead to inode\n> +starvation and degrade the performance of the whole system. Since we\n> +can never pack those objects, these repositories often take up a large amount of\n> +disk space, since we can only zlib compress them, but not store them in delta\n> +chains.\n> +\n> +== Cruft packs\n> +\n> +Cruft packs are designed to eliminate the need for storing unreachable objects\n> +in a loose state by including the per-object mtimes in a separate file alongside\n> +a single pack containing all loose objects.\n\nI had the same question as Stolee here: why not use the cruft-pack's\nmtime for all the objects in it?  Much later below, you make it clear\nthat a repository will generally only have one cruft pack which kind\nof answers the question, but the repeated mention of \"cruft packs\"\nthroughout the document subtly made me make the opposite assumption.\nIt might be nice to address the almost-always-only-one-cruft-pack\nearlier on, which may also help answer the question about why you need\nto store individual mtimes in an additional file.\n\n> +A cruft pack is written by `git repack --cruft` when generating a new pack.\n> +linkgit:git-pack-objects[1]'s `--cruft` option. Note that `git repack --cruft`\n> +is a classic all-into-one repack, meaning that everything in the resulting pack is\n> +reachable, and everything else is unreachable. Once written, the `--cruft`\n> +option instructs `git repack` to generate another pack containing only objects\n> +not packed in the previous step (which equates to packing all unreachable\n> +objects together). This progresses as follows:\n> +\n> +  1. Enumerate every object, marking any object which is (a) not contained in a\n> +     kept-pack, and (b) whose mtime is within the grace period as a traversal\n> +     tip.\n> +\n> +  2. Perform a reachability traversal based on the tips gathered in the previous\n> +     step, adding every object along the way to the pack.\n> +\n> +  3. Write the pack out, along with a `.mtimes` file that records the per-object\n> +     timestamps.\n> +\n> +This mode is invoked internally by linkgit:git-repack[1] when instructed to\n> +write a cruft pack. Crucially, the set of in-core kept packs is exactly the set\n> +of packs which will not be deleted by the repack; in other words, they contain\n> +all of the repository's reachable objects.\n> +\n> +When a repository already has a cruft pack, `git repack --cruft` typically only\n> +adds objects to it. An exception to this is when `git repack` is given the\n> +`--cruft-expiration` option, which allows the generated cruft pack to omit\n> +expired objects instead of waiting for linkgit:git-gc[1] to expire those objects\n> +later on.\n> +\n> +It is linkgit:git-gc[1] that is typically responsible for removing expired\n> +unreachable objects.\n> +\n> +== Alternatives\n> +\n> +Notable alternatives to this design include:\n> +\n> +  - The location of the per-object mtime data, and\n> +  - Whether cruft packs should be incremental or not.\n> +\n> +On the location of mtime data, a new auxiliary file tied to the pack was chosen\n> +to avoid complicating the `.idx` format. If the `.idx` format were ever to gain\n> +support for optional chunks of data, it may make sense to consolidate the\n> +`.mtimes` format into the `.idx` itself.\n> +\n> +Incremental cruft packs (i.e., where each time a repository is repacked a new\n> +cruft pack is generated containing only the unreachable objects introduced since\n> +the last time a cruft pack was written) are significantly more complicated to\n> +construct, and so aren't pursued here. The obvious drawback to the current\n> +implementation is that the entire cruft pack must be re-written from scratch.\n> --\n> 2.34.1.25.gb3157a20e6\n>\n"},{"id":"443095","messageId":"Yav6pDmSSDzm7lZO@nand.local","threadId":"56996","inReplyTo":"CABPp-BGW3t1LnUGNSLSzYtQfAJuAQKxpHJMJOdh5T2pUyaDWAw@mail.gmail.com","subject":"Re: [PATCH 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2021-12-04T23:32:52Z","receivedAt":"2021-12-04T23:32:56Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Sat, Dec 04, 2021 at 02:20:23PM -0800, Elijah Newren wrote:\n> > +== Cruft packs\n> > +\n> > +Cruft packs are designed to eliminate the need for storing unreachable objects\n> > +in a loose state by including the per-object mtimes in a separate file alongside\n> > +a single pack containing all loose objects.\n>\n> I had the same question as Stolee here: why not use the cruft-pack's\n> mtime for all the objects in it?  Much later below, you make it clear\n> that a repository will generally only have one cruft pack which kind\n> of answers the question, but the repeated mention of \"cruft packs\"\n> throughout the document subtly made me make the opposite assumption.\n> It might be nice to address the almost-always-only-one-cruft-pack\n> earlier on, which may also help answer the question about why you need\n> to store individual mtimes in an additional file.\n\nResponding to your suggestions out of order ;-). Throughout the\ndocument, I wrote \"cruft packs\" in the sense of \"the feature this series\nimplements\", not \"multiple cruft packs\".\n\nBut my wording is unintentionally vague, especially because this\ndocument does talk about why this series stores unreachable objects in a\nsingle cruft pack. I updated my copy to make clear the difference\nbetween the two, which should hopefully avoid any confusion here in the\nfuture.\n\nAs far as why not use the cruft pack's timestamp as the mtime for all of\nthe unreachable objects contained within it, there are a few reasons:\n\nIt makes freshening objects more complicated. Not because we couldn't\nfreshen individual objects (we would likely do so in the same way this\nseries does, by rewriting it loose and using the loose copy's mtime\ninstead), but because it makes it complicated to repack a repository\nwith many cruft packs. If I have a handful of cruft packs, and freshen a\nhandful of objects within them, I now need to update many cruft packs,\nor pay the price of storing their objects twice (if I instead don't\nrewrite them and keep the loose copies around).\n\nIt also makes it impossible to share deltas between cruft objects that\ndon't have the same timestamp, unless the cruft packs are stored thin\n(in which case it becomes much more complicated to figure out which\ncruft packs can be safely pruned without storing information about which\nother packs a thin pack has deltas against).\n\nI'm sure there were others, but these are the ones that I could recall\noff the top of my head. This all felt like a little too much detail for\nthe \"alternative designs\" section, but if you think some or all of this\nwould be interesting to memorialize not just on the mailing list, let me\nknow.\n\nThanks,\nTaylor\n"},{"id":"443118","messageId":"xmqqk0gikcf8.fsf@gitster.g","threadId":"56996","inReplyTo":"a05675ab834ac5e8bc3ab72847b0621a563e0e1b.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 12/17] builtin/repack.c: support generating a cruft pack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-12-05T20:46:19Z","receivedAt":"2021-12-05T20:46:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Various thoughts on just this part, as the hunk got my attention\nwhile merging with other topics in 'seen'.\n\n> +\tif (pack_everything & PACK_CRUFT && delete_redundant) {\n> +\t\tif (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE))\n> +\t\t\tdie(_(\"--cruft and -A are incompatible\"));\n> +\t\tif (keep_unreachable)\n> +\t\t\tdie(_(\"--cruft and -k are incompatible\"));\n> +\t\tif (!(pack_everything & ALL_INTO_ONE))\n> +\t\t\tdie(_(\"--cruft must be combined with all-into-one\"));\n> +\t}\n\nThe \"reuse similar messages for i18n\" topic will encourage us to\nturn this part into:\n\n\tif (pack_everything & PACK_CRUFT && delete_redundant) {\n\t\tif (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE))\n\t\t\tdie(_(\"%s and %s are mutually exclusive\"),\n\t\t\t    \"--cruft\", \"-A\");\n\t\tif (keep_unreachable)\n\t\t\tdie(_(\"%s and %s are mutually exclusive\"),\n\t\t\t    \"--cruft\", \"-k\");\n\t\tif (!(pack_everything & ALL_INTO_ONE))\n\t\t\tdie(_(\"--cruft must be combined with all-into-one\"));\n\t}\n\nThe conditionals are a bit unpleasant to read and maintain, but I\nguess we cannot help it?\n\nSaying ALL_INTO_ONE is a bit unfriendly to the end user, who would\nprobably not know that it is the name the code gave to the bit that\nis turned on when given an option externally known under a different\nname (is that \"-a\"?).\n\nIf \"--cruft\" must be used with \"all into one\", I wonder if it makes\nsense to make it imply that?  Not in the sense that OPT_BIT()\ninitially flips the ALL_INTO_ONE bit on upon seeing \"--cruft\", but\nafter parse_options() returns, we check PACK_CRUFT and if it is on\nturn ALL_INTO_ONE also on (so even if '-a' gains '--all-into-one'\noption, the user won't break us by giving \"--no-all-into-one\" after\nthey gave us \"--cruft\")?  I didn't think about this part thoroughly\nenough, though.\n\nThanks.\n\n\n\n\n\n\n"},{"id":"443179","messageId":"518f6dde-e849-f9f1-ef9c-274cd9179ce0@gmail.com","threadId":"56996","inReplyTo":"Yaqc6FYo1pLhsNSB@nand.local","subject":"Re: [PATCH 04/17] chunk-format.h: extract oid_version()","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-06T17:33:30Z","receivedAt":"2021-12-06T17:33:35Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/3/21 5:40 PM, Taylor Blau wrote:\n> On Thu, Dec 02, 2021 at 10:22:05AM -0500, Derrick Stolee wrote:\n>> I notice that you don't use this in load_pack_mtimes_file(),\n>> in pack-mtimes.c but you could at this point.\n> \n> Hmm, I'm confused. Te extracted function converts a pointer to a struct\n> git_hash_algo into a uint32, but here we just care about reading the\n> four byte value we wrote.\n\nAh. I got mixed up here. Sorry.\n\n-Stolee\n"},{"id":"443197","messageId":"15f1bbc6-7ae6-0ed3-872a-51feebd1296c@gmail.com","threadId":"56996","inReplyTo":"e0a7b3b310c69350d8e2c0561e0991bb7045a66d.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 06/17] t/helper: add 'pack-mtimes' test-tool","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-06T21:16:04Z","receivedAt":"2021-12-06T21:16:26Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/29/2021 5:25 PM, Taylor Blau wrote:\n> +static int dump_mtimes(struct packed_git *p)\n\nnit: you return an int here so you can use it as an error code...\n\n> +{\n> +\tuint32_t i;\n> +\tif (load_pack_mtimes(p) < 0)\n> +\t\tdie(\"could not load pack .mtimes\");\n> +\n> +\tfor (i = 0; i < p->num_objects; i++) {\n> +\t\tstruct object_id oid;\n> +\t\tif (nth_packed_object_id(&oid, p, i) < 0)\n> +\t\t\tdie(\"could not load object id at position %\"PRIu32, i);\n> +\n> +\t\tprintf(\"%s %\"PRIu32\"\\n\",\n> +\t\t       oid_to_hex(&oid), nth_packed_mtime(p, i));\n> +\t}\n> +\n> +\treturn 0;\n\nBut always return 0 unless you die().\n\n> +\treturn p ? dump_mtimes(p) : 1;\n\nIt makes this line concise, I suppose.\n\nPerhaps just use \"return dump_mtimes(p)\" and have dump_mtimes()\nreturn 1 if the given pack is NULL?\n\nThanks,\n-Stolee\n"},{"id":"443200","messageId":"b3a30e27-7821-1fcb-bacc-07a6d2b3df76@gmail.com","threadId":"56996","inReplyTo":"66165917a4660f63ce60b820d178d52a51304d20.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-06T21:44:31Z","receivedAt":"2021-12-06T21:44:41Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/29/2021 5:25 PM, Taylor Blau wrote:\n> Generating a non-expiring cruft packs works as follows:\n\nI had trouble parsing the documentation changes below, so I came back\nto this commit message to see if that helps.\n \n>   - Callers provide a list of every pack they know about, and indicate\n>     which packs are about to be removed.\n\nThis corresponds to the list over stdin.\n \n>   - All packs which are going to be removed (we'll call these the\n>     redundant ones) are marked as kept in-core, as well as any packs\n>     that `pack-objects` found but the caller did not specify.\n\nOk, so as an implementation detail we mark these as keep packs.\n\n>     These packs are presumed to have entered the repository between\n>     the caller collecting packs and invoking `pack-objects`. Since we\n>     do not want to include objects in these packs (because we don't know\n>     which of their objects are or aren't reachable), these are also\n>     marked as kept in-core.\n\nHere, \"are presumed\" is doing a lot of work. Theoretically, there could\nbe three categories:\n\n1. This pack was just repacked and will be removed because all of its\n   objects were placed into new objects.\n\n2. Either this pack was repacked and contains important reachable objects\n   OR we did a repack of reachable objects and this pack contained some\n   extra, unreachable objects.\n\n3. This pack was added to the repository while creating those repacked\n   packs from category 2, so we don't know if things are reachable or\n   not.\n\nSo, the packs that we discover on-disk but are not specified over stdin\nare in this third category, but these are grouped with category 1 as we\nwill treat them the same.\n\n>   - Then, we enumerate all objects in the repository, and add them to\n>     our packing list if they do not appear in an in-core kept pack.\n\nHere, we are looking at all of the objects in category 2 as well as\nloose objects.\n\n> This results in a new cruft pack which contains all known objects that\n> aren't included in the kept packs. When the kept pack is the result of\n> `git repack -A`, the resulting pack contains all unreachable objects.\n\nThis now describes how 'git repack' will interface with this new change\nto pack-objects. I'll keep an eye out for that.\n\n> +--cruft::\n\nNow getting to this description.\n\n> +\tPacks unreachable objects into a separate \"cruft\" pack, denoted\n> +\tby the existence of a `.mtimes` file. Pack names provided over\n> +\tstdin indicate which packs will remain after a `git repack`.\n> +\tPack names prefixed with a `-` indicate those which will be\n> +\tremoved. (...)\n\nThis description is too tied to 'git repack'. Can we describe the\ninput using terms independent of the 'git repack' operation? I need\nto keep reading.\n\n> (...) The contents of the cruft pack are all objects not\n> +\tcontained in the surviving packs specified by `--keep-pack`)\n\nNow you use --keep-pack, which is a way of specifying a pack as\n\"in-core keep\" which was not in your commit message. Here, we also\ndon't link the packs over stdin to the concept of keep packs.\n\n> +\twhich have not exceeded the grace period (see\n> +\t`--cruft-expiration` below), or which have exceeded the grace\n> +\tperiod, but are reachable from an other object which hasn't.\n\nAnd now we think about the grace period! There is so much going on\nthat I need to break it down to understand.\n\n  An object is _excluded_ from the new cruft pack if\n\n  1. It is reachable from at least one reference.\n  2. It is in a pack from stdin prefixed with \"-\"\n  3. It is in a pack specified by `--keep-pack`\n  4. It is in an existing cruft pack and the .mtimes file states\n     that its mtime is at least as recent as the time specified by\n     the --cruft-expiration option.\n\nBreaking it down into a list like this helps me, at least. I'm not\nsure what the best way would look like.\n\n(Needing to pause here and look at the implementation later.)\n\nThanks,\n-Stolee\n"},{"id":"443287","messageId":"38198f38-ca06-1ab3-344b-29e7b6857ed0@gmail.com","threadId":"56996","inReplyTo":"66165917a4660f63ce60b820d178d52a51304d20.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-07T15:17:28Z","receivedAt":"2021-12-07T15:17:32Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/29/2021 5:25 PM, Taylor Blau wrote:\n> diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> +static int add_cruft_object_entry(const struct object_id *oid, enum object_type type,\n> +\t\t\t\t  struct packed_git *pack, off_t offset,\n> +\t\t\t\t  const char *name, uint32_t mtime)\n> +{\n> +\tstruct object_entry *entry;\n> +\n> +\tdisplay_progress(progress_state, ++nr_seen);\n\nI don't love the global nr_seen here, but it is pervasive through the\nfile. OK.\n\n> +\tentry = packlist_find(&to_pack, oid);\n> +\tif (entry) {\n> +\t\tif (name) {\n> +\t\t\tentry->hash = pack_name_hash(name);\n> +\t\t\tentry->no_try_delta = name && no_try_delta(name);\n\nThis is already in an \"if (name)\" block, so \"name &&\" isn't needed.\n\n> +\t\t}\n> +\t} else {\n> +\t\tif (!want_object_in_pack(oid, 0, &pack, &offset))\n> +\t\t\treturn 0;\n> +\t\tif (!pack && type == OBJ_BLOB && !has_loose_object(oid)) {\n> +\t\t\t/*\n> +\t\t\t * If a traversed tree has a missing blob then we want\n> +\t\t\t * to avoid adding that missing object to our pack.\n> +\t\t\t *\n> +\t\t\t * This only applies to missing blobs, not trees,\n> +\t\t\t * because the traversal needs to parse sub-trees but\n> +\t\t\t * not blobs.\n> +\t\t\t *\n> +\t\t\t * Note we only perform this check when we couldn't\n> +\t\t\t * already find the object in a pack, so we're really\n> +\t\t\t * limited to \"ensure non-tip blobs which don't exist in\n> +\t\t\t * packs do exist via loose objects\". Confused?\n> +\t\t\t */\n> +\t\t\treturn 0;\n> +\t\t}\n> +\n> +\t\tentry = create_object_entry(oid, type, pack_name_hash(name),\n> +\t\t\t\t\t    0, name && no_try_delta(name),\n> +\t\t\t\t\t    pack, offset);\n> +\t}\n> +\n> +\tif (mtime > oe_cruft_mtime(&to_pack, entry))\n> +\t\toe_set_cruft_mtime(&to_pack, entry, mtime);\n> +\treturn 1;\n\nI was confused at this \"return 1\" here, while other cases return 0.\n\nIt turns out that there are multiple methods in this file that have\ndifferent semantics: add_loose_object() and add_object_entry_from_pack()\nare both called from iterators where \"return 1\" means \"stop iterating\"\nso they return 0 always. add_object_entry_from_bitmap() is used to\niterate over a bitmap and \"return 1\" means \"include this object\".\n\nHowever, the return code for add_cruft_object_entry() is never used,\nso it should probably return void or swap the meanings to have nonzero\nmean an error occurred.\n\n> +static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n> +{\n> +\tstruct string_list_item *item = NULL;\n> +\tfor_each_string_list_item(item, packs) {\n> +\t\tstruct packed_git *p = item->util;\n> +\t\tif (!p)\n> +\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n\nInteresting that this is a potential issue. We are expecting the pack\nto be loaded before we get here. Is this more because some packs might\nnot actually load, but it's fine as long as we don't mark them as kept?\n\n> +\t\tp->pack_keep_in_core = keep;\n> +\t}\n> +}\n...\n> +static void read_cruft_objects(void)\n> +{\n> +\tstruct strbuf buf = STRBUF_INIT;\n> +\tstruct string_list discard_packs = STRING_LIST_INIT_DUP;\n> +\tstruct string_list fresh_packs = STRING_LIST_INIT_DUP;\n> +\tstruct packed_git *p;\n> +\n> +\tignore_packed_keep_in_core = 1;\n\nHere is a global that we are suddenly changing. Should we not be\nreturning it to its initial state when this method is complete?\n\n> +static int option_parse_cruft_expiration(const struct option *opt,\n> +\t\t\t\t\t const char *arg, int unset)\n> +{\n> +\tif (unset) {\n> +\t\tcruft = 0;\n\nThis unassignment of 'cruft' when cruft-expiration is unset with\n--no-cruft-expiration seems odd. I would expect\n\n\tgit pack-objects --cruft --no-cruft-expiration\n\nto still make a cruft pack, but not expire anything. It seems that\nyour code here makes --no-cruft-expiration disable the --cruft option.\n\n> +\t\tcruft_expiration = 0;\n> +\t} else {\n> +\t\tcruft = 1;\n> +\t\tif (arg)\n> +\t\t\tcruft_expiration = approxidate(arg);\n> +\t}\n> +\treturn 0;\n> +}\n..\n> +\t\tOPT_BOOL(0, \"cruft\", &cruft, N_(\"create a cruft pack\")),\n> +\t\tOPT_CALLBACK_F(0, \"cruft-expiration\", NULL, N_(\"time\"),\n> +\t\t  N_(\"expire cruft objects older than <time>\"),\n> +\t\t  PARSE_OPT_OPTARG, option_parse_cruft_expiration),\n\n> -static int has_loose_object(const struct object_id *oid)\n> +int has_loose_object(const struct object_id *oid)\n>  {\n>  \treturn check_and_freshen(oid, 0);\n>  }\n\nI'm surprised this hasn't been modified to use a repository pointer.\nAdding another caller here isn't too much debt, though.\n\n> diff --git a/object-store.h b/object-store.h\n> index d87481f101..a79c1c91ab 100644\n> --- a/object-store.h\n> +++ b/object-store.h\n> @@ -308,6 +308,8 @@ int repo_has_object_file_with_flags(struct repository *r,\n>   */\n>  int has_loose_object_nonlocal(const struct object_id *);\n\nOf course, here is another example that is already more widely used.\n\n> +int has_loose_object(const struct object_id *);\n> +\n>  void assert_oid_type(const struct object_id *oid, enum object_type expect);\n\n...\n\n> +\ttest_expect_success \"unreachable packed objects are packed (expire $expire)\" '\n> +\t\tgit init repo &&\n> +\t\ttest_when_finished \"rm -fr repo\" &&\n> +\t\t(\n> +\t\t\tcd repo &&\n> +\n> +\t\t\ttest_commit packed &&\n> +\t\t\tgit repack -Ad &&\n> +\t\t\ttest_commit other &&\n> +\n> +\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n> +\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n> +\t\t\tother=\"$(git pack-objects --delta-base-offset \\\n> +\t\t\t\t$packdir/pack <objects)\" &&\n> +\t\t\tgit prune-packed &&\n> +\n> +\t\t\ttest-tool chmtime --get -100 \"$packdir/pack-$other.pack\" >expect &&\n\nI am missing how this test creates _unreachable_ objects. I would expect removal of\nsome refs or a 'git reset --hard' somewhere. What am I missing?\n\n> +\t\t\tcruft=\"$(git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack <<-EOF\n> +\t\t\t$keep\n> +\t\t\t-pack-$other.pack\n> +\t\t\tEOF\n> +\t\t\t)\" &&\n> +\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n> +\n> +\t\t\tcut -d\" \" -f2 <actual.raw | sort -u >actual &&\n> +\n> +\t\t\ttest_cmp expect actual\n> +\t\t)\n> +\t'\n> +\n> +\ttest_expect_success \"unreachable cruft objects are repacked (expire $expire)\" '\n\nI have the same question for all of the tests, really.\n\n> +\t\t\t# remove the unreachable tree, but leave the commit\n> +\t\t\t# which has it as its root tree in-tact\n\nnit: \"intact\" is one word.\n\n> +\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$tree\")\" &&\n> +\n> +\t\t\tgit repack -Ad &&\n> +\t\t\tbasename $(ls $packdir/pack-*.pack) >in &&\n> +\t\t\tgit pack-objects --cruft --cruft-expiration=\"$expire\" \\\n> +\t\t\t\t$packdir/pack <in\n> +\t\t)\n> +\t'\n\n...\n\n> +basic_cruft_pack_tests never\n\nI look forward to seeing how this changes with additional expiration values.\n\nThanks,\n-Stolee\n\n"},{"id":"443288","messageId":"865b99dd-0b18-9a07-49c1-3959a777c685@gmail.com","threadId":"56996","inReplyTo":"37fda94785f1e689c7a7c32e69c6ff16fee7da4f.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 11/17] builtin/pack-objects.c: --cruft with expiration","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-07T15:30:52Z","receivedAt":"2021-12-07T15:30:56Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/29/2021 5:25 PM, Taylor Blau wrote:\n\n> +static void enumerate_and_traverse_cruft_objects(struct string_list *fresh_packs)\n> +{\n...\n> +\t/*\n> +\t * Re-mark only the fresh packs as kept so that objects in\n> +\t * unknown packs do not halt the reachability traversal early.\n> +\t */\n> +\tfor (p = get_all_packs(the_repository); p; p = p->next)\n> +\t\tp->pack_keep_in_core = 0;\n> +\tmark_pack_kept_in_core(fresh_packs, 1);\n\nAre we ever going to recover this pack_keep_in_core state? Should we\nbe saving it somewhere so we can return without mutating this state\npermanently?\n\n> +\tif (prepare_revision_walk(&revs))\n> +\t\tdie(_(\"revision walk setup failed\"));\n> +\tif (progress)\n> +\t\tprogress_state = start_progress(_(\"Traversing cruft objects\"), 0);\n> +\tnr_seen = 0;\n> +\ttraverse_commit_list(&revs, show_cruft_commit, show_cruft_object, NULL);\n> +\n> +\tstop_progress(&progress_state);\n> +}\n> +\n>  static void read_cruft_objects(void)\n>  {\n>  \tstruct strbuf buf = STRBUF_INIT;\n> @@ -3515,7 +3597,7 @@ static void read_cruft_objects(void)\n>  \tmark_pack_kept_in_core(&discard_packs, 0);\n>  \n>  \tif (cruft_expiration)\n> -\t\tdie(\"--cruft-expiration not yet implemented\");\n> +\t\tenumerate_and_traverse_cruft_objects(&fresh_packs);\n>  \telse\n>  \t\tenumerate_cruft_objects();\n\n>  basic_cruft_pack_tests never\n> +basic_cruft_pack_tests 2.weeks.ago\n\nI'm surprised these tests didn't require any changes to adapt to the\nnew expiration date. But I suppose none of the mtimes were older than\ntwo weeks ago?\n\nI continue to miss something in these tests, because I don't see how\nthings are becoming unreachable.\n\nThanks,\n-Stolee\n"},{"id":"443289","messageId":"c9437c89-9258-4034-9886-8a2aec46aa6b@gmail.com","threadId":"56996","inReplyTo":"a05675ab834ac5e8bc3ab72847b0621a563e0e1b.1638224692.git.me@ttaylorr.com","subject":"Re: [PATCH 12/17] builtin/repack.c: support generating a cruft pack","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-07T15:38:05Z","receivedAt":"2021-12-07T15:38:09Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/29/2021 5:25 PM, Taylor Blau wrote:\n\n> +static int write_cruft_pack(const struct pack_objects_args *args,\n> +\t\t\t    const char *pack_prefix,\n> +\t\t\t    struct string_list *names,\n> +\t\t\t    struct string_list *existing_packs,\n> +\t\t\t    struct string_list *existing_kept_packs)\n> +{\n> +\tstruct child_process cmd = CHILD_PROCESS_INIT;\n> +\tstruct strbuf line = STRBUF_INIT;\n> +\tstruct string_list_item *item;\n> +\tFILE *in, *out;\n> +\tint ret;\n> +\n> +\tprepare_pack_objects(&cmd, args);\n> +\n> +\tstrvec_push(&cmd.args, \"--cruft\");\n> +\tif (cruft_expiration)\n> +\t\tstrvec_pushf(&cmd.args, \"--cruft-expiration=%s\",\n> +\t\t\t     cruft_expiration);\n> +\n> +\tstrvec_push(&cmd.args, \"--honor-pack-keep\");\n> +\tstrvec_push(&cmd.args, \"--non-empty\");\n> +\tstrvec_push(&cmd.args, \"--max-pack-size=0\");\n\nThis --max-pack-size is meaningless, right? The config that would change\nthis is already ignored by 'git pack-objects'.\n\n> +\t\tOPT_BIT(0, \"cruft\", &pack_everything,\n> +\t\t\t\tN_(\"same as -a, pack unreachable cruft objects separately\"),\n> +\t\t\t\t   PACK_CRUFT | ALL_INTO_ONE),\n\nI can understand the use of OPT_BIT here. Keep in mind that --no-cruft would\nremove the '-a' option, if it already existed. Perhaps we should just use\nOPT_BOOL and update to add the ALL_INTO_ONE if PACK_CRUFT exists?\n\n> +\t\tOPT_STRING(0, \"cruft-expiration\", &cruft_expiration, N_(\"approxidate\"),\n> +\t\t\t\tN_(\"with -C, expire objects older than this\")),\n\nHere, --no-cruft-expiration will set cruft_expiration to NULL and not overwrite\nthe --cruft option, as expected. Just pointing out that this is different than\nthe option in 'git pack-objects'.\n\n> --- a/t/t5327-pack-objects-cruft.sh\n> +++ b/t/t5327-pack-objects-cruft.sh\n> @@ -358,4 +358,157 @@ test_expect_success 'expired objects are pruned' '\n>  \t)\n>  '\n>  \n> +test_expect_success 'repack --cruft generates a cruft pack' '\n> +\tgit init repo &&\n> +\ttest_when_finished \"rm -fr repo\" &&\n> +\t(\n> +\t\tcd repo &&\n> +\n> +\t\ttest_commit reachable &&\n> +\t\tgit branch -M main &&\n> +\t\tgit checkout --orphan other &&\n\nHere is a way to make objects unreachable!\n\n> +\t\ttest_commit unreachable &&\n> +\n> +\t\tgit checkout main &&\n> +\t\tgit branch -D other &&\n> +\t\tgit tag -d unreachable &&\n\nThanks,\n-Stolee\n"},{"id":"445741","messageId":"YdiXecK6fAKl8++G@nand.local","threadId":"56996","inReplyTo":"YaqZA02FsCFA9qBi@nand.local","subject":"Re: [PATCH 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-01-07T19:41:45Z","receivedAt":"2022-01-07T19:41:49Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Dec 03, 2021 at 05:24:03PM -0500, Taylor Blau wrote:\n> On Thu, Dec 02, 2021 at 10:06:07AM -0500, Derrick Stolee wrote:\n>     - A table of 4-byte unsigned integers in network order. The ith\n>       value is the modification time (mtime) of the ith object in the\n>       corresponding pack by lexicographic (index) order. The mtimes\n>       count standard epoch seconds.\n>\n> > Storing these mtimes in 32-bits means we will hit the 2038 problem.\n> > The commit-graph stores commit times with an extra two bits to extend\n> > the lifetime by another hundred years or so.\n> >\n> > Could we extend the lifetime of cruft packs by decreasing the granularity\n> > here? Should 'mtime' store a number of _minutes_ instead of seconds? That\n> > should be enough granularity for these purposes.\n>\n> Perhaps, though it does add some complexity to the code that deals with\n> this format at the expense of some future-proofing. I'm open to it,\n> though.\n\nI still have quite a bit of review from this topic sitting in my inbox.\n\nBut this had been lingering on my mind, and I realized I said something\nincorrect. 32-bit mtimes won't cause us to run into the \"2038\" problem,\nsince these aren't signed values. So storing epoch seconds in a uint32_t\nshould get us into the year 2106.\n\nIf anybody is still using cruft packs by then, I'll call this project a\nwild success ;-). So in the meantime, I don't think it makes sense to\nreduce the granularity and/or use extra bits to store the timestamps.\n\nThanks,\nTaylor\n"},{"id":"449359","messageId":"Yha0M1bKxCb1/MrJ@nand.local","threadId":"56996","inReplyTo":"15f1bbc6-7ae6-0ed3-872a-51feebd1296c@gmail.com","subject":"Re: [PATCH 06/17] t/helper: add 'pack-mtimes' test-tool","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-02-23T22:24:51Z","receivedAt":"2022-02-23T22:24:56Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Mon, Dec 06, 2021 at 04:16:04PM -0500, Derrick Stolee wrote:\n> On 11/29/2021 5:25 PM, Taylor Blau wrote:\n> > +static int dump_mtimes(struct packed_git *p)\n>\n> nit: you return an int here so you can use it as an error code...\n>\n> > +{\n> > +\tuint32_t i;\n> > +\tif (load_pack_mtimes(p) < 0)\n> > +\t\tdie(\"could not load pack .mtimes\");\n> > +\n> > +\tfor (i = 0; i < p->num_objects; i++) {\n> > +\t\tstruct object_id oid;\n> > +\t\tif (nth_packed_object_id(&oid, p, i) < 0)\n> > +\t\t\tdie(\"could not load object id at position %\"PRIu32, i);\n> > +\n> > +\t\tprintf(\"%s %\"PRIu32\"\\n\",\n> > +\t\t       oid_to_hex(&oid), nth_packed_mtime(p, i));\n> > +\t}\n> > +\n> > +\treturn 0;\n>\n> But always return 0 unless you die().\n>\n> > +\treturn p ? dump_mtimes(p) : 1;\n>\n> It makes this line concise, I suppose.\n>\n> Perhaps just use \"return dump_mtimes(p)\" and have dump_mtimes()\n> return 1 if the given pack is NULL?\n\nI think just dying in the case we have a NULL pack is fine, and it\nshould be OK to lump it in the same case as \"could not load pack .mtimes\".\n\nBut we may want to catch the case a little earlier while we still have\nthe pack name handy. Perhaps something like this on top:\n\n--- 8< ---\ndiff --git a/t/helper/test-pack-mtimes.c b/t/helper/test-pack-mtimes.c\nindex b143f62520..f7b79daf4c 100644\n--- a/t/helper/test-pack-mtimes.c\n+++ b/t/helper/test-pack-mtimes.c\n@@ -5,7 +5,7 @@\n #include \"packfile.h\"\n #include \"pack-mtimes.h\"\n\n-static int dump_mtimes(struct packed_git *p)\n+static void dump_mtimes(struct packed_git *p)\n {\n \tuint32_t i;\n \tif (load_pack_mtimes(p) < 0)\n@@ -19,8 +19,6 @@ static int dump_mtimes(struct packed_git *p)\n \t\tprintf(\"%s %\"PRIu32\"\\n\",\n \t\t       oid_to_hex(&oid), nth_packed_mtime(p, i));\n \t}\n-\n-\treturn 0;\n }\n\n static const char *pack_mtimes_usage = \"\\n\"\n@@ -49,5 +47,10 @@ int cmd__pack_mtimes(int argc, const char **argv)\n\n \tstrbuf_release(&buf);\n\n-\treturn p ? dump_mtimes(p) : 1;\n+\tif (!p)\n+\t\tdie(\"could not find pack '%s'\", argv[1]);\n+\n+\tdump_mtimes(p);\n+\n+\treturn 0;\n }\n--- >8 ---\n\nThanks,\nTaylor\n"},{"id":"449370","messageId":"YhbEiLAX06LekNiR@nand.local","threadId":"56996","inReplyTo":"38198f38-ca06-1ab3-344b-29e7b6857ed0@gmail.com","subject":"Re: [PATCH 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-02-23T23:34:32Z","receivedAt":"2022-02-23T23:34:36Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Dec 07, 2021 at 10:17:28AM -0500, Derrick Stolee wrote:\n> On 11/29/2021 5:25 PM, Taylor Blau wrote:\n> > diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> > +static int add_cruft_object_entry(const struct object_id *oid, enum object_type type,\n> > +\t\t\t\t  struct packed_git *pack, off_t offset,\n> > +\t\t\t\t  const char *name, uint32_t mtime)\n> > +{\n> > +\tstruct object_entry *entry;\n> > +\n> > +\tdisplay_progress(progress_state, ++nr_seen);\n>\n> I don't love the global nr_seen here, but it is pervasive through the\n> file. OK.\n\nYeah; this is how all of the existing progress code works in\npack-objects.\n\n> > +\tentry = packlist_find(&to_pack, oid);\n> > +\tif (entry) {\n> > +\t\tif (name) {\n> > +\t\t\tentry->hash = pack_name_hash(name);\n> > +\t\t\tentry->no_try_delta = name && no_try_delta(name);\n>\n> This is already in an \"if (name)\" block, so \"name &&\" isn't needed.\n\nThanks; this is a copy-and-paste from add_object_entry(), where we\naren't in a conditional on \"name\". We could also fold the conditional on\nwhether or not name is NULL into no_try_delta itself, since all existing\ncalls look like \"name && no_try_delta(name)\".\n\nSo adding something like:\n\n    if (!name)\n      return 0;\n\nto the beginning of no_try_delta()'s implementation would allow us to\nget rid of the handful of \"name &&\"s. But I'm trying to avoid touching\nother parts of pack-objects as much as I can, so I'll hold off for now.\n\n> > +\t\t}\n> > +\t} else {\n> > +\t\tif (!want_object_in_pack(oid, 0, &pack, &offset))\n> > +\t\t\treturn 0;\n> > +\t\tif (!pack && type == OBJ_BLOB && !has_loose_object(oid)) {\n> > +\t\t\t/*\n> > +\t\t\t * If a traversed tree has a missing blob then we want\n> > +\t\t\t * to avoid adding that missing object to our pack.\n> > +\t\t\t *\n> > +\t\t\t * This only applies to missing blobs, not trees,\n> > +\t\t\t * because the traversal needs to parse sub-trees but\n> > +\t\t\t * not blobs.\n> > +\t\t\t *\n> > +\t\t\t * Note we only perform this check when we couldn't\n> > +\t\t\t * already find the object in a pack, so we're really\n> > +\t\t\t * limited to \"ensure non-tip blobs which don't exist in\n> > +\t\t\t * packs do exist via loose objects\". Confused?\n> > +\t\t\t */\n> > +\t\t\treturn 0;\n> > +\t\t}\n> > +\n> > +\t\tentry = create_object_entry(oid, type, pack_name_hash(name),\n> > +\t\t\t\t\t    0, name && no_try_delta(name),\n> > +\t\t\t\t\t    pack, offset);\n> > +\t}\n> > +\n> > +\tif (mtime > oe_cruft_mtime(&to_pack, entry))\n> > +\t\toe_set_cruft_mtime(&to_pack, entry, mtime);\n> > +\treturn 1;\n>\n> I was confused at this \"return 1\" here, while other cases return 0.\n>\n> It turns out that there are multiple methods in this file that have\n> different semantics: add_loose_object() and add_object_entry_from_pack()\n> are both called from iterators where \"return 1\" means \"stop iterating\"\n> so they return 0 always. add_object_entry_from_bitmap() is used to\n> iterate over a bitmap and \"return 1\" means \"include this object\".\n>\n> However, the return code for add_cruft_object_entry() is never used,\n> so it should probably return void or swap the meanings to have nonzero\n> mean an error occurred.\n\nYes, exactly. And thanks for tracing out both of the different\nmeanings/interpretations of these add_xyz_entry() functions. As you can\nimagine, this implementation is copy-and-pasted from add_object_entry(),\nwhich was specialized for this use here. At the time, I gave some effort\ntowards trying to share more code with add_object_entry() for this\nspecial case, but it ended up being pretty awkward, hence the separate\nimplementation.\n\nIronically, add_object_entry()'s return code is also unused, so we could\nprobably clean that up, too. But like the above, I'll avoid it for now\nin an effort to touch as little of pack-objects in this patch as I can.\n\n> > +static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n> > +{\n> > +\tstruct string_list_item *item = NULL;\n> > +\tfor_each_string_list_item(item, packs) {\n> > +\t\tstruct packed_git *p = item->util;\n> > +\t\tif (!p)\n> > +\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n>\n> Interesting that this is a potential issue. We are expecting the pack\n> to be loaded before we get here. Is this more because some packs might\n> not actually load, but it's fine as long as we don't mark them as kept?\n\nNot quite \"loaded\" (though any pack structures that we look at by this\npoint will be fully \"loaded\"). Instead, we're making sure that all of\nthe packs names we read from stdin could be matched to packs that we\nfound in the repository (i.e., that we produce an appropriate error\nmessage if we found \"pack-does-not-exist.pack\" on stdin).\n\nThis is all because we process input from stdin in two phases:\n\n  - First, read all of the input into two string_lists, one for the\n    packs we're about to discard (anything that start with '-'), and\n    another for all of the \"fresh\" packs (i.e., anything that we're not\n    going to discard).\n\n  - Then, loop through all of the packed_git structs we have, querying\n    both of the aforementioned string lists for input that matches each\n    pack's `pack_name` field, and setting the `->util` pointer of the\n    matching string_list_entry appropriately.\n\nFollowing those two steps, any list entries that have a NULL util\npointer correspond with bogus input, so we want to call die() there.\n\n> > +\t\tp->pack_keep_in_core = keep;\n> > +\t}\n> > +}\n> ...\n> > +static void read_cruft_objects(void)\n> > +{\n> > +\tstruct strbuf buf = STRBUF_INIT;\n> > +\tstruct string_list discard_packs = STRING_LIST_INIT_DUP;\n> > +\tstruct string_list fresh_packs = STRING_LIST_INIT_DUP;\n> > +\tstruct packed_git *p;\n> > +\n> > +\tignore_packed_keep_in_core = 1;\n>\n> Here is a global that we are suddenly changing. Should we not be\n> returning it to its initial state when this method is complete?\n\nWe could, although it won't matter in practice, because we'll want to\nkeep that setting around for our traversal, after which point\npack-objects will exit.\n\n> > +static int option_parse_cruft_expiration(const struct option *opt,\n> > +\t\t\t\t\t const char *arg, int unset)\n> > +{\n> > +\tif (unset) {\n> > +\t\tcruft = 0;\n>\n> This unassignment of 'cruft' when cruft-expiration is unset with\n> --no-cruft-expiration seems odd. I would expect\n>\n> \tgit pack-objects --cruft --no-cruft-expiration\n>\n> to still make a cruft pack, but not expire anything. It seems that\n> your code here makes --no-cruft-expiration disable the --cruft option.\n\nHmm. I could see compelling reasoning that goes both ways. On the one\nhand, `--no-cruft-expiration` (to me, at least) seems to imply \"set\n`--cruft-expiration` to \"never\"). On the other hand, it also matches our\nconvention of `--no`-prefixed options to unset some value. This\nimplementation takes the latter approach, though we could easily change\nit to set the cruft expiration to \"never\".\n\nI don't have a strong opinion about which is better, so I'm happy to do\neither if you have a better sense about which has more expected\nbehavior.\n\n> > +\t\tcruft_expiration = 0;\n> > +\t} else {\n> > +\t\tcruft = 1;\n> > +\t\tif (arg)\n> > +\t\t\tcruft_expiration = approxidate(arg);\n> > +\t}\n> > +\treturn 0;\n> > +}\n> ..\n> > +\t\tOPT_BOOL(0, \"cruft\", &cruft, N_(\"create a cruft pack\")),\n> > +\t\tOPT_CALLBACK_F(0, \"cruft-expiration\", NULL, N_(\"time\"),\n> > +\t\t  N_(\"expire cruft objects older than <time>\"),\n> > +\t\t  PARSE_OPT_OPTARG, option_parse_cruft_expiration),\n>\n> > -static int has_loose_object(const struct object_id *oid)\n> > +int has_loose_object(const struct object_id *oid)\n> >  {\n> >  \treturn check_and_freshen(oid, 0);\n> >  }\n>\n> I'm surprised this hasn't been modified to use a repository pointer.\n> Adding another caller here isn't too much debt, though.\n\nYeah, check_and_freshen() doesn't have a variant that takes a\nrepository pointer. Good #leftoverbits, I guess!\n\n> > +int has_loose_object(const struct object_id *);\n> > +\n> >  void assert_oid_type(const struct object_id *oid, enum object_type expect);\n>\n> ...\n>\n> > +\ttest_expect_success \"unreachable packed objects are packed (expire $expire)\" '\n> > +\t\tgit init repo &&\n> > +\t\ttest_when_finished \"rm -fr repo\" &&\n> > +\t\t(\n> > +\t\t\tcd repo &&\n> > +\n> > +\t\t\ttest_commit packed &&\n> > +\t\t\tgit repack -Ad &&\n> > +\t\t\ttest_commit other &&\n> > +\n> > +\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n> > +\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n> > +\t\t\tother=\"$(git pack-objects --delta-base-offset \\\n> > +\t\t\t\t$packdir/pack <objects)\" &&\n> > +\t\t\tgit prune-packed &&\n> > +\n> > +\t\t\ttest-tool chmtime --get -100 \"$packdir/pack-$other.pack\" >expect &&\n>\n> I am missing how this test creates _unreachable_ objects. I would expect removal of\n> some refs or a 'git reset --hard' somewhere. What am I missing?\n\nFor this and the other tests the so-called \"unreachable\" objects are\ntechnically reachable, but we can treat them as unreachable by putting\nthem in the \"discard\" packs list (or by not mentioning them at all to\n`git pack-objects --cruft`).\n\n> > +\t\t\t# remove the unreachable tree, but leave the commit\n> > +\t\t\t# which has it as its root tree in-tact\n>\n> nit: \"intact\" is one word.\n\nThanks; fixed here and in the other test which was added by this commit.\n\nThanks,\nTaylor\n"},{"id":"449371","messageId":"YhbE3YJPN+oj3LJJ@nand.local","threadId":"56996","inReplyTo":"865b99dd-0b18-9a07-49c1-3959a777c685@gmail.com","subject":"Re: [PATCH 11/17] builtin/pack-objects.c: --cruft with expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-02-23T23:35:57Z","receivedAt":"2022-02-23T23:36:01Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Dec 07, 2021 at 10:30:52AM -0500, Derrick Stolee wrote:\n> On 11/29/2021 5:25 PM, Taylor Blau wrote:\n>\n> > +static void enumerate_and_traverse_cruft_objects(struct string_list *fresh_packs)\n> > +{\n> ...\n> > +\t/*\n> > +\t * Re-mark only the fresh packs as kept so that objects in\n> > +\t * unknown packs do not halt the reachability traversal early.\n> > +\t */\n> > +\tfor (p = get_all_packs(the_repository); p; p = p->next)\n> > +\t\tp->pack_keep_in_core = 0;\n> > +\tmark_pack_kept_in_core(fresh_packs, 1);\n>\n> Are we ever going to recover this pack_keep_in_core state? Should we\n> be saving it somewhere so we can return without mutating this state\n> permanently?\n\nIn the same sense that we are free to modify the global\nignore_packed_keep_in_core variable (because we only stop caring about\nthe modified state right before the program is about to exist) we can\nfreely mutate these variables, too.\n\n> > +\tif (prepare_revision_walk(&revs))\n> > +\t\tdie(_(\"revision walk setup failed\"));\n> > +\tif (progress)\n> > +\t\tprogress_state = start_progress(_(\"Traversing cruft objects\"), 0);\n> > +\tnr_seen = 0;\n> > +\ttraverse_commit_list(&revs, show_cruft_commit, show_cruft_object, NULL);\n> > +\n> > +\tstop_progress(&progress_state);\n> > +}\n> > +\n> >  static void read_cruft_objects(void)\n> >  {\n> >  \tstruct strbuf buf = STRBUF_INIT;\n> > @@ -3515,7 +3597,7 @@ static void read_cruft_objects(void)\n> >  \tmark_pack_kept_in_core(&discard_packs, 0);\n> >\n> >  \tif (cruft_expiration)\n> > -\t\tdie(\"--cruft-expiration not yet implemented\");\n> > +\t\tenumerate_and_traverse_cruft_objects(&fresh_packs);\n> >  \telse\n> >  \t\tenumerate_cruft_objects();\n>\n> >  basic_cruft_pack_tests never\n> > +basic_cruft_pack_tests 2.weeks.ago\n>\n> I'm surprised these tests didn't require any changes to adapt to the\n> new expiration date. But I suppose none of the mtimes were older than\n> two weeks ago?\n\nExactly.\n\nThanks,\nTaylor\n"},{"id":"449372","messageId":"YhbFTYmjdci86oY7@nand.local","threadId":"56996","inReplyTo":"c9437c89-9258-4034-9886-8a2aec46aa6b@gmail.com","subject":"Re: [PATCH 12/17] builtin/repack.c: support generating a cruft pack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-02-23T23:37:49Z","receivedAt":"2022-02-23T23:37:54Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"(Jumping forward a little bit while responding to your review to finish\nmy train of though before I log off for today...)\n\nOn Tue, Dec 07, 2021 at 10:38:05AM -0500, Derrick Stolee wrote:\n> > --- a/t/t5327-pack-objects-cruft.sh\n> > +++ b/t/t5327-pack-objects-cruft.sh\n> > @@ -358,4 +358,157 @@ test_expect_success 'expired objects are pruned' '\n> >  \t)\n> >  '\n> >\n> > +test_expect_success 'repack --cruft generates a cruft pack' '\n> > +\tgit init repo &&\n> > +\ttest_when_finished \"rm -fr repo\" &&\n> > +\t(\n> > +\t\tcd repo &&\n> > +\n> > +\t\ttest_commit reachable &&\n> > +\t\tgit branch -M main &&\n> > +\t\tgit checkout --orphan other &&\n>\n> Here is a way to make objects unreachable!\n\nYes, indeed. And this is the first spot where we *need* to care about\nobject reachability, because the set of packs that `git repack` passes\nover stdin to `git pack-objects --cruft` depends on which objects are\nand aren't reachable.\n\nIn the tests that exercise `pack-objects --cruft` directly, we can\npretend that certain packs contain only unreachable objects by marking\nthem as \"discarded\".\n\nThanks,\nTaylor\n"},{"id":"449842","messageId":"Yh1+WQcs71rQYPEg@nand.local","threadId":"56996","inReplyTo":"xmqqk0gikcf8.fsf@gitster.g","subject":"Re: [PATCH 12/17] builtin/repack.c: support generating a cruft pack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-01T02:00:57Z","receivedAt":"2022-03-01T02:01:02Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Sun, Dec 05, 2021 at 12:46:19PM -0800, Junio C Hamano wrote:\n> Various thoughts on just this part, as the hunk got my attention\n> while merging with other topics in 'seen'.\n>\n> > +\tif (pack_everything & PACK_CRUFT && delete_redundant) {\n> > +\t\tif (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE))\n> > +\t\t\tdie(_(\"--cruft and -A are incompatible\"));\n> > +\t\tif (keep_unreachable)\n> > +\t\t\tdie(_(\"--cruft and -k are incompatible\"));\n> > +\t\tif (!(pack_everything & ALL_INTO_ONE))\n> > +\t\t\tdie(_(\"--cruft must be combined with all-into-one\"));\n> > +\t}\n>\n> The \"reuse similar messages for i18n\" topic will encourage us to\n> turn this part into:\n>\n> \tif (pack_everything & PACK_CRUFT && delete_redundant) {\n> \t\tif (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE))\n> \t\t\tdie(_(\"%s and %s are mutually exclusive\"),\n> \t\t\t    \"--cruft\", \"-A\");\n> \t\tif (keep_unreachable)\n> \t\t\tdie(_(\"%s and %s are mutually exclusive\"),\n> \t\t\t    \"--cruft\", \"-k\");\n> \t\tif (!(pack_everything & ALL_INTO_ONE))\n> \t\t\tdie(_(\"--cruft must be combined with all-into-one\"));\n> \t}\n\nThanks, done.\n\n> The conditionals are a bit unpleasant to read and maintain, but I\n> guess we cannot help it?\n\nI don't know that I find them unpleasant to read, but perhaps they are a\nhassle to maintain (as we add new, mutually-exclusive options). But I\ncan't seem to think of a better alternative...\n\n> Saying ALL_INTO_ONE is a bit unfriendly to the end user, who would\n> probably not know that it is the name the code gave to the bit that\n> is turned on when given an option externally known under a different\n> name (is that \"-a\"?).\n>\n> If \"--cruft\" must be used with \"all into one\", I wonder if it makes\n> sense to make it imply that?  Not in the sense that OPT_BIT()\n> initially flips the ALL_INTO_ONE bit on upon seeing \"--cruft\", but\n> after parse_options() returns, we check PACK_CRUFT and if it is on\n> turn ALL_INTO_ONE also on (so even if '-a' gains '--all-into-one'\n> option, the user won't break us by giving \"--no-all-into-one\" after\n> they gave us \"--cruft\")?  I didn't think about this part thoroughly\n> enough, though.\n\nYes, `--cruft` must be used with an option that sets ALL_INTO_ONE. Since\nwe don't have any automatic '--no-' versions of single character\noptions, I think that this conditional is currently redundant, but I\nagree that this code would break if we (a) removed the conditional\nyou're talking about and (b) allowed passing something like\n`--no-all-into-one` which unsets the ALL_INTO_ONE bit.\n\nSo setting ALL_INTO_ONE ourselves _after_ option parsing is done makes\nsense to me, thanks.\n\nThanks,\nTaylor\n"},{"id":"449846","messageId":"Yh2JgPzR4Tg3PmNL@nand.local","threadId":"56996","inReplyTo":"b3a30e27-7821-1fcb-bacc-07a6d2b3df76@gmail.com","subject":"Re: [PATCH 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-01T02:48:32Z","receivedAt":"2022-03-01T02:48:37Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Mon, Dec 06, 2021 at 04:44:31PM -0500, Derrick Stolee wrote:\n> On 11/29/2021 5:25 PM, Taylor Blau wrote:\n> > Generating a non-expiring cruft packs works as follows:\n>\n> I had trouble parsing the documentation changes below, so I came back\n> to this commit message to see if that helps.\n>\n> >   - Callers provide a list of every pack they know about, and indicate\n> >     which packs are about to be removed.\n>\n> This corresponds to the list over stdin.\n>\n> >   - All packs which are going to be removed (we'll call these the\n> >     redundant ones) are marked as kept in-core, as well as any packs\n> >     that `pack-objects` found but the caller did not specify.\n>\n> Ok, so as an implementation detail we mark these as keep packs.\n\n\n> >     These packs are presumed to have entered the repository between\n> >     the caller collecting packs and invoking `pack-objects`. Since we\n> >     do not want to include objects in these packs (because we don't know\n> >     which of their objects are or aren't reachable), these are also\n> >     marked as kept in-core.\n>\n> Here, \"are presumed\" is doing a lot of work. Theoretically, there could\n> be three categories:\n>\n> 1. This pack was just repacked and will be removed because all of its\n>    objects were placed into new objects.\n>\n> 2. Either this pack was repacked and contains important reachable objects\n>    OR we did a repack of reachable objects and this pack contained some\n>    extra, unreachable objects.\n>\n> 3. This pack was added to the repository while creating those repacked\n>    packs from category 2, so we don't know if things are reachable or\n>    not.\n>\n> So, the packs that we discover on-disk but are not specified over stdin\n> are in this third category, but these are grouped with category 1 as we\n> will treat them the same.\n\nAh, I think I caused some unintentional confusion by attaching \"are\npresumed\" to \"these packs\", when it wasn't clear that \"these packs\"\nmeant \"ones that aren't listed over stdin\".\n\nSince the caller is supposed to provide a complete picture of the\nrepository as they see it, any packs known to the pack-objects process\nthat aren't mentioned over stdin are assumed to have entered the\nrepository after the caller was spun up.\n\nI'll clarify this section of the commit message, since I agree it is\nunnecessarily confusing.\n\n> >   - Then, we enumerate all objects in the repository, and add them to\n> >     our packing list if they do not appear in an in-core kept pack.\n>\n> Here, we are looking at all of the objects in category 2 as well as\n> loose objects.\n\nWe're enumerating any objects that aren't in packs which are marked as\nkept in-core (along with loose objects which don't appear in packs that\nare marked as kept in-core).\n\nThe in-core kept packs are ones that the caller (and I find it's helpful\nto read \"the caller\" as \"git repack\") has marked as \"will delete\". So\nthe non in-core pack(s) that we're looking at here contain all reachable\nobjects (e.g., like you would get with `git repack -A`).\n\n> > +\tPacks unreachable objects into a separate \"cruft\" pack, denoted\n> > +\tby the existence of a `.mtimes` file. Pack names provided over\n> > +\tstdin indicate which packs will remain after a `git repack`.\n> > +\tPack names prefixed with a `-` indicate those which will be\n> > +\tremoved. (...)\n>\n> This description is too tied to 'git repack'. Can we describe the\n> input using terms independent of the 'git repack' operation? I need\n> to keep reading.\n>\n> > (...) The contents of the cruft pack are all objects not\n> > +\tcontained in the surviving packs specified by `--keep-pack`)\n>\n> Now you use --keep-pack, which is a way of specifying a pack as\n> \"in-core keep\" which was not in your commit message. Here, we also\n> don't link the packs over stdin to the concept of keep packs.\n\nThe mention of `--keep-pack` is a mistake left over from a previous\nversion; thanks for spotting. Here's a version of the first paragraph\nfrom this piece of documentation which is less tied to `git repack` and\nhopefully a little clearer:\n\n    --cruft::\n            Packs unreachable objects into a separate \"cruft\" pack, denoted\n            by the existence of a `.mtimes` file. Typically used by `git\n            repack --cruft`. Callers provide a list of pack names and\n            indicate which packs will remain in the repository, along with\n            which packs will be deleted (indicated by the `-` prefix). The\n            contents of the cruft pack are all objects not contained in the\n            surviving packs which have not exceeded the grace period (see\n            `--cruft-expiration` below), or which have exceeded the grace\n            period, but are reachable from an other object which hasn't.\n\n> > +\twhich have not exceeded the grace period (see\n> > +\t`--cruft-expiration` below), or which have exceeded the grace\n> > +\tperiod, but are reachable from an other object which hasn't.\n>\n> And now we think about the grace period! There is so much going on\n> that I need to break it down to understand.\n>\n>   An object is _excluded_ from the new cruft pack if\n>\n>   1. It is reachable from at least one reference.\n>   2. It is in a pack from stdin prefixed with \"-\"\n>   3. It is in a pack specified by `--keep-pack`\n>   4. It is in an existing cruft pack and the .mtimes file states\n>      that its mtime is at least as recent as the time specified by\n>      the --cruft-expiration option.\n>\n> Breaking it down into a list like this helps me, at least. I'm not\n> sure what the best way would look like.\n\nGiven some expiration T, cruft packs contain all unreachable objects\nwhich are newer than T, along with any cruft objects (i.e., those not\ndirectly reachable from any ref) which are older than T, but reachable\nfrom another cruft object newer than T.\n\nThanks,\nTaylor\n"},{"id":"450030","messageId":"cover.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH v2 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:57:57Z","receivedAt":"2022-03-02T00:58:04Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Here is a reroll of my series to implement \"cruft packs\", a pack which\nstores accumulated unreachable objects, along with a new \".mtimes\" file\nwhich tracks each object's last known modification time.\n\nThis was on the list towards the end of 2021[1], and I have been\naccumulating small changes to it locally for a couple of months now.\nMajor changes since last time include:\n\n  - Clearer documentation and commit message(s) to better illustrate how\n    the feature works and is supposed to be used.\n\n  - Some minor documentation updates to pack-format.txt, which make some\n    ambiguous details more explicit.\n\n  - Minor code movement / tweaks to make things easier to read, ensure\n    that functions aren't introduced in patches before they are used /\n    etc.\n\n  - Moved the new test script to t5328 (instead of t5327, which happens\n    to be taken up by a new MIDX bitmap-related test), and purged it of\n    all \"rm -fr .git/logs\" (replacing them with \"git reflog --expire\n    --all --expire=all\" instead).\n\n  - A new test which fixes a bug where loose objects which have copies\n    that appear in a cruft pack would not get accumulated when doing a\n    `--geometric` repack.\n\nFor convenience, a range-diff is below. Thanks in advance for taking\nanother look!\n\n[1]: https://lore.kernel.org/git/cover.1638224692.git.me@ttaylorr.com/\n\nTaylor Blau (17):\n  Documentation/technical: add cruft-packs.txt\n  pack-mtimes: support reading .mtimes files\n  pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'\n  chunk-format.h: extract oid_version()\n  pack-mtimes: support writing pack .mtimes files\n  t/helper: add 'pack-mtimes' test-tool\n  builtin/pack-objects.c: return from create_object_entry()\n  builtin/pack-objects.c: --cruft without expiration\n  reachable: add options to add_unseen_recent_objects_to_traversal\n  reachable: report precise timestamps from objects in cruft packs\n  builtin/pack-objects.c: --cruft with expiration\n  builtin/repack.c: support generating a cruft pack\n  builtin/repack.c: allow configuring cruft pack generation\n  builtin/repack.c: use named flags for existing_packs\n  builtin/repack.c: add cruft packs to MIDX during geometric repack\n  builtin/gc.c: conditionally avoid pruning objects via loose\n  sha1-file.c: don't freshen cruft packs\n\n Documentation/Makefile                  |   1 +\n Documentation/config/gc.txt             |  21 +-\n Documentation/config/repack.txt         |   9 +\n Documentation/git-gc.txt                |   5 +\n Documentation/git-pack-objects.txt      |  30 +\n Documentation/git-repack.txt            |  11 +\n Documentation/technical/cruft-packs.txt |  97 ++++\n Documentation/technical/pack-format.txt |  19 +\n Makefile                                |   2 +\n builtin/gc.c                            |  10 +-\n builtin/pack-objects.c                  | 304 +++++++++-\n builtin/repack.c                        | 183 +++++-\n bulk-checkin.c                          |   2 +-\n chunk-format.c                          |  12 +\n chunk-format.h                          |   3 +\n commit-graph.c                          |  18 +-\n midx.c                                  |  18 +-\n object-file.c                           |   4 +-\n object-store.h                          |   7 +-\n pack-mtimes.c                           | 129 +++++\n pack-mtimes.h                           |  15 +\n pack-objects.c                          |   6 +\n pack-objects.h                          |  25 +\n pack-write.c                            |  93 ++-\n pack.h                                  |   4 +\n packfile.c                              |  19 +-\n reachable.c                             |  58 +-\n reachable.h                             |   9 +-\n t/helper/test-pack-mtimes.c             |  56 ++\n t/helper/test-tool.c                    |   1 +\n t/helper/test-tool.h                    |   1 +\n t/t5328-pack-objects-cruft.sh           | 739 ++++++++++++++++++++++++\n 32 files changed, 1810 insertions(+), 101 deletions(-)\n create mode 100644 Documentation/technical/cruft-packs.txt\n create mode 100644 pack-mtimes.c\n create mode 100644 pack-mtimes.h\n create mode 100644 t/helper/test-pack-mtimes.c\n create mode 100755 t/t5328-pack-objects-cruft.sh\n\nRange-diff against v1:\n 1:  a9f7c738e0 !  1:  784ee7e0ee Documentation/technical: add cruft-packs.txt\n    @@ Documentation/technical/cruft-packs.txt (new)\n     @@\n     += Cruft packs\n     +\n    -+Cruft packs offer an alternative to Git's traditional mechanism of removing\n    -+unreachable objects. This document provides an overview of Git's pruning\n    -+mechanism, and how cruft packs can be used instead to accomplish the same.\n    ++The cruft packs feature offer an alternative to Git's traditional mechanism of\n    ++removing unreachable objects. This document provides an overview of Git's\n    ++pruning mechanism, and how a cruft pack can be used instead to accomplish the\n    ++same.\n     +\n     +== Background\n     +\n    @@ Documentation/technical/cruft-packs.txt (new)\n     +\n     +== Cruft packs\n     +\n    -+Cruft packs are designed to eliminate the need for storing unreachable objects\n    -+in a loose state by including the per-object mtimes in a separate file alongside\n    -+a single pack containing all loose objects.\n    ++A cruft pack eliminates the need for storing unreachable objects in a loose\n    ++state by including the per-object mtimes in a separate file alongside a single\n    ++pack containing all loose objects.\n     +\n     +A cruft pack is written by `git repack --cruft` when generating a new pack.\n     +linkgit:git-pack-objects[1]'s `--cruft` option. Note that `git repack --cruft`\n    @@ Documentation/technical/cruft-packs.txt (new)\n     +Notable alternatives to this design include:\n     +\n     +  - The location of the per-object mtime data, and\n    -+  - Whether cruft packs should be incremental or not.\n    ++  - Storing unreachable objects in multiple cruft packs.\n     +\n     +On the location of mtime data, a new auxiliary file tied to the pack was chosen\n     +to avoid complicating the `.idx` format. If the `.idx` format were ever to gain\n     +support for optional chunks of data, it may make sense to consolidate the\n     +`.mtimes` format into the `.idx` itself.\n     +\n    -+Incremental cruft packs (i.e., where each time a repository is repacked a new\n    -+cruft pack is generated containing only the unreachable objects introduced since\n    -+the last time a cruft pack was written) are significantly more complicated to\n    -+construct, and so aren't pursued here. The obvious drawback to the current\n    -+implementation is that the entire cruft pack must be re-written from scratch.\n    ++Storing unreachable objects among multiple cruft packs (e.g., creating a new\n    ++cruft pack during each repacking operation including only unreachable objects\n    ++which aren't already stored in an earlier cruft pack) is significantly more\n    ++complicated to construct, and so aren't pursued here. The obvious drawback to\n    ++the current implementation is that the entire cruft pack must be re-written from\n    ++scratch.\n 2:  7d4ae7bd3e !  2:  101b34660c pack-mtimes: support reading .mtimes files\n    @@ Documentation/technical/pack-format.txt: Pack file entry: <+\n     +\n     +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n     +\n    -+  - A table of mtimes (one per packed object, num_objects in total, each\n    -+    a 4-byte unsigned integer in network order), in the same order as\n    -+    objects appear in the index file (e.g., the first entry in the mtime\n    -+    table corresponds to the object with the lowest lexically-sorted\n    -+    oid). The mtimes count standard epoch seconds.\n    ++  - A table of 4-byte unsigned integers in network order. The ith\n    ++    value is the modification time (mtime) of the ith object in the\n    ++    corresponding pack by lexicographic (index) order. The mtimes\n    ++    count standard epoch seconds.\n     +\n    -+  - A trailer, containing a:\n    -+\n    -+    checksum of the corresponding packfile, and\n    -+\n    -+    a checksum of all of the above.\n    ++  - A trailer, containing a checksum of the corresponding packfile,\n    ++    and a checksum of all of the above (each having length according\n    ++    to the specified hash function).\n     +\n     +All 4-byte numbers are in network order.\n     +\n    @@ pack-mtimes.c (new)\n     +\treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n     +}\n     +\n    -+int pack_has_mtimes(struct packed_git *p)\n    -+{\n    -+\tstruct stat st;\n    -+\tchar *fname = pack_mtimes_filename(p);\n    -+\n    -+\tif (stat(fname, &st) < 0) {\n    -+\t\tif (errno == ENOENT)\n    -+\t\t\treturn 0;\n    -+\t\tdie_errno(_(\"could not stat %s\"), fname);\n    -+\t}\n    -+\n    -+\tfree(fname);\n    -+\treturn 1;\n    -+}\n    -+\n     +#define MTIMES_HEADER_SIZE (12)\n     +#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n     +\n    @@ pack-mtimes.c (new)\n     +\tstruct stat st;\n     +\tvoid *data = NULL;\n     +\tsize_t mtimes_size;\n    ++\tstruct mtimes_header header;\n     +\tuint32_t *hdr;\n     +\n     +\tfd = git_open(mtimes_file);\n    @@ pack-mtimes.c (new)\n     +\n     +\tdata = hdr = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n     +\n    -+\tif (ntohl(*hdr) != MTIMES_SIGNATURE) {\n    ++\theader.signature = ntohl(hdr[0]);\n    ++\theader.version = ntohl(hdr[1]);\n    ++\theader.hash_id = ntohl(hdr[2]);\n    ++\n    ++\tif (header.signature != MTIMES_SIGNATURE) {\n     +\t\tret = error(_(\"mtimes file %s has unknown signature\"), mtimes_file);\n     +\t\tgoto cleanup;\n     +\t}\n     +\n    -+\tif (ntohl(*++hdr) != 1) {\n    ++\tif (header.version != 1) {\n     +\t\tret = error(_(\"mtimes file %s has unsupported version %\"PRIu32),\n    -+\t\t\t    mtimes_file, ntohl(*hdr));\n    ++\t\t\t    mtimes_file, header.version);\n     +\t\tgoto cleanup;\n     +\t}\n    -+\thdr++;\n    -+\tif (!(ntohl(*hdr) == 1 || ntohl(*hdr) == 2)) {\n    ++\n    ++\tif (!(header.hash_id == 1 || header.hash_id == 2)) {\n     +\t\tret = error(_(\"mtimes file %s has unsupported hash id %\"PRIu32),\n    -+\t\t\t    mtimes_file, ntohl(*hdr));\n    ++\t\t\t    mtimes_file, header.hash_id);\n     +\t\tgoto cleanup;\n     +\t}\n     +\n    @@ pack-mtimes.h (new)\n     +\n     +struct packed_git;\n     +\n    -+int pack_has_mtimes(struct packed_git *p);\n     +int load_pack_mtimes(struct packed_git *p);\n     +\n     +uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos);\n    @@ pack-mtimes.h (new)\n     +#endif\n     \n      ## packfile.c ##\n    -@@ packfile.c: void close_pack_revindex(struct packed_git *p) {\n    +@@ packfile.c: static void close_pack_revindex(struct packed_git *p)\n      \tp->revindex_data = NULL;\n      }\n      \n    -+void close_pack_mtimes(struct packed_git *p) {\n    ++static void close_pack_mtimes(struct packed_git *p)\n    ++{\n     +\tif (!p->mtimes_map)\n     +\t\treturn;\n     +\n    @@ packfile.c: static void prepare_pack(const char *full_name, size_t full_name_len\n      \t\tstring_list_append(data->garbage, full_name);\n      \telse\n      \t\treport_garbage(PACKDIR_FILE_GARBAGE, full_name);\n    -\n    - ## packfile.h ##\n    -@@ packfile.h: uint32_t get_pack_fanout(struct packed_git *p, uint32_t value);\n    - unsigned char *use_pack(struct packed_git *, struct pack_window **, off_t, unsigned long *);\n    - void close_pack_windows(struct packed_git *);\n    - void close_pack_revindex(struct packed_git *);\n    -+void close_pack_mtimes(struct packed_git *p);\n    - void close_pack(struct packed_git *);\n    - void close_object_store(struct raw_object_store *o);\n    - void unuse_pack(struct pack_window **);\n 3:  7f4612e859 =  3:  a94d7dfeb3 pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'\n 4:  ea245b7216 =  4:  1e0ed363ae chunk-format.h: extract oid_version()\n 5:  deece9eb70 !  5:  5236490688 pack-mtimes: support writing pack .mtimes files\n    @@ pack-objects.h: struct packing_data {\n      \tunsigned int *tree_depth;\n      \tunsigned char *layer;\n     +\n    -+\t/* cruft packs */\n    ++\t/*\n    ++\t * Used when writing cruft packs.\n    ++\t *\n    ++\t * Object mtimes are stored in pack order when writing, but\n    ++\t * written out in lexicographic (index) order.\n    ++\t */\n     +\tuint32_t *cruft_mtime;\n      };\n      \n    @@ pack-write.c: const char *write_rev_file_order(const char *rev_name,\n     +\thashwrite_be32(f, oid_version(the_hash_algo));\n     +}\n     +\n    ++/*\n    ++ * Writes the object mtimes of \"objects\" for use in a .mtimes file.\n    ++ * Note that objects must be in lexicographic (index) order, which is\n    ++ * the expected ordering of these values in the .mtimes file.\n    ++ */\n     +static void write_mtimes_objects(struct hashfile *f,\n     +\t\t\t\t struct packing_data *to_pack,\n     +\t\t\t\t struct pack_idx_entry **objects,\n    @@ pack-write.c: const char *write_rev_file_order(const char *rev_name,\n     +\twrite_mtimes_objects(f, to_pack, objects, nr_objects);\n     +\twrite_mtimes_trailer(f, hash);\n     +\n    -+\tif (mtimes_name && adjust_shared_perm(mtimes_name) < 0)\n    ++\tif (adjust_shared_perm(mtimes_name) < 0)\n     +\t\tdie(_(\"failed to make %s readable\"), mtimes_name);\n     +\n     +\tfinalize_hashfile(f, NULL,\n    @@ pack-write.c: void stage_tmp_packfiles(struct strbuf *name_buffer,\n     +\t\tmtimes_tmp_name = write_mtimes_file(NULL, to_pack, written_list,\n     +\t\t\t\t\t\t    nr_written,\n     +\t\t\t\t\t\t    hash);\n    -+\t\tif (adjust_shared_perm(mtimes_tmp_name))\n    -+\t\t\tdie_errno(\"unable to make temporary mtimes file readable\");\n     +\t}\n     +\n      \trename_tmp_packfile(name_buffer, pack_tmp_name, \"pack\");\n 6:  e0a7b3b310 !  6:  78313bc441 t/helper: add 'pack-mtimes' test-tool\n    @@ t/helper/test-pack-mtimes.c (new)\n     +#include \"packfile.h\"\n     +#include \"pack-mtimes.h\"\n     +\n    -+static int dump_mtimes(struct packed_git *p)\n    ++static void dump_mtimes(struct packed_git *p)\n     +{\n     +\tuint32_t i;\n     +\tif (load_pack_mtimes(p) < 0)\n    @@ t/helper/test-pack-mtimes.c (new)\n     +\t\tprintf(\"%s %\"PRIu32\"\\n\",\n     +\t\t       oid_to_hex(&oid), nth_packed_mtime(p, i));\n     +\t}\n    -+\n    -+\treturn 0;\n     +}\n     +\n     +static const char *pack_mtimes_usage = \"\\n\"\n    @@ t/helper/test-pack-mtimes.c (new)\n     +\n     +\tstrbuf_release(&buf);\n     +\n    -+\treturn p ? dump_mtimes(p) : 1;\n    ++\tif (!p)\n    ++\t\tdie(\"could not find pack '%s'\", argv[1]);\n    ++\n    ++\tdump_mtimes(p);\n    ++\n    ++\treturn 0;\n     +}\n     \n      ## t/helper/test-tool.c ##\n 7:  5710933127 =  7:  142098668d builtin/pack-objects.c: return from create_object_entry()\n 8:  66165917a4 !  8:  2517a6be3d builtin/pack-objects.c: --cruft without expiration\n    @@ Commit message\n             which packs are about to be removed.\n     \n           - All packs which are going to be removed (we'll call these the\n    -        redundant ones) are marked as kept in-core, as well as any packs\n    -        that `pack-objects` found but the caller did not specify.\n    +        redundant ones) are marked as kept in-core.\n     \n    -        These packs are presumed to have entered the repository between\n    -        the caller collecting packs and invoking `pack-objects`. Since we\n    -        do not want to include objects in these packs (because we don't know\n    -        which of their objects are or aren't reachable), these are also\n    -        marked as kept in-core.\n    +        Any packs the caller did not mention (but are known to the\n    +        `pack-objects` process) are also marked as kept in-core. Packs not\n    +        mentioned by the caller are assumed to be unknown to them, i.e.,\n    +        they entered the repository after the caller decided which packs\n    +        should be kept and which should be discarded.\n    +\n    +        Since we do not want to include objects in these \"unknown\" packs\n    +        (because we don't know which of their objects are or aren't\n    +        reachable), these are also marked as kept in-core.\n     \n           - Then, we enumerate all objects in the repository, and add them to\n             our packing list if they do not appear in an in-core kept pack.\n    @@ Documentation/git-pack-objects.txt: SYNOPSIS\n      \t[--local] [--incremental] [--window=<n>] [--depth=<n>]\n      \t[--revs [--unpacked | --all]] [--keep-pack=<pack-name>]\n     +\t[--cruft] [--cruft-expiration=<time>]\n    - \t[--stdout [--filter=<filter-spec>] | base-name]\n    - \t[--shallow] [--keep-true-parents] [--[no-]sparse] < object-list\n    + \t[--stdout [--filter=<filter-spec>] | <base-name>]\n    + \t[--shallow] [--keep-true-parents] [--[no-]sparse] < <object-list>\n      \n     @@ Documentation/git-pack-objects.txt: base-name::\n      Incompatible with `--revs`, or options that imply `--revs` (such as\n    @@ Documentation/git-pack-objects.txt: base-name::\n      \n     +--cruft::\n     +\tPacks unreachable objects into a separate \"cruft\" pack, denoted\n    -+\tby the existence of a `.mtimes` file. Pack names provided over\n    -+\tstdin indicate which packs will remain after a `git repack`.\n    -+\tPack names prefixed with a `-` indicate those which will be\n    -+\tremoved. The contents of the cruft pack are all objects not\n    -+\tcontained in the surviving packs specified by `--keep-pack`)\n    -+\twhich have not exceeded the grace period (see\n    ++\tby the existence of a `.mtimes` file. Typically used by `git\n    ++\trepack --cruft`. Callers provide a list of pack names and\n    ++\tindicate which packs will remain in the repository, along with\n    ++\twhich packs will be deleted (indicated by the `-` prefix). The\n    ++\tcontents of the cruft pack are all objects not contained in the\n    ++\tsurviving packs which have not exceeded the grace period (see\n     +\t`--cruft-expiration` below), or which have exceeded the grace\n     +\tperiod, but are reachable from an other object which hasn't.\n     ++\n    ++When the input lists a pack containing all reachable objects (and lists\n    ++all other packs as pending deletion), the corresponding cruft pack will\n    ++contain all unreachable objects (with mtime newer than the\n    ++`--cruft-expiration`) along with any unreachable objects whose mtime is\n    ++older than the `--cruft-expiration`, but are reachable from an\n    ++unreachable object whose mtime is newer than the `--cruft-expiration`).\n    +++\n     +Incompatible with `--unpack-unreachable`, `--keep-unreachable`,\n     +`--pack-loose-unreachable`, `--stdin-packs`, as well as any other\n     +options which imply `--revs`. Also incompatible with `--max-pack-size`;\n    @@ builtin/pack-objects.c: static void read_packs_list_from_stdin(void)\n      \tstring_list_clear(&exclude_packs, 0);\n      }\n      \n    -+static int add_cruft_object_entry(const struct object_id *oid, enum object_type type,\n    -+\t\t\t\t  struct packed_git *pack, off_t offset,\n    -+\t\t\t\t  const char *name, uint32_t mtime)\n    ++static void add_cruft_object_entry(const struct object_id *oid, enum object_type type,\n    ++\t\t\t\t   struct packed_git *pack, off_t offset,\n    ++\t\t\t\t   const char *name, uint32_t mtime)\n     +{\n     +\tstruct object_entry *entry;\n     +\n    @@ builtin/pack-objects.c: static void read_packs_list_from_stdin(void)\n     +\tif (entry) {\n     +\t\tif (name) {\n     +\t\t\tentry->hash = pack_name_hash(name);\n    -+\t\t\tentry->no_try_delta = name && no_try_delta(name);\n    ++\t\t\tentry->no_try_delta = no_try_delta(name);\n     +\t\t}\n     +\t} else {\n     +\t\tif (!want_object_in_pack(oid, 0, &pack, &offset))\n    -+\t\t\treturn 0;\n    ++\t\t\treturn;\n     +\t\tif (!pack && type == OBJ_BLOB && !has_loose_object(oid)) {\n     +\t\t\t/*\n     +\t\t\t * If a traversed tree has a missing blob then we want\n    @@ builtin/pack-objects.c: static void read_packs_list_from_stdin(void)\n     +\t\t\t * limited to \"ensure non-tip blobs which don't exist in\n     +\t\t\t * packs do exist via loose objects\". Confused?\n     +\t\t\t */\n    -+\t\t\treturn 0;\n    ++\t\t\treturn;\n     +\t\t}\n     +\n     +\t\tentry = create_object_entry(oid, type, pack_name_hash(name),\n    @@ builtin/pack-objects.c: static void read_packs_list_from_stdin(void)\n     +\n     +\tif (mtime > oe_cruft_mtime(&to_pack, entry))\n     +\t\toe_set_cruft_mtime(&to_pack, entry, mtime);\n    -+\treturn 1;\n    ++\treturn;\n     +}\n     +\n     +static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n    @@ builtin/pack-objects.c: int cmd_pack_objects(int argc, const char **argv, const\n      \t\tread_packs_list_from_stdin();\n      \t\tif (rev_list_unpacked)\n      \t\t\tadd_unreachable_loose_objects();\n    --\t} else if (!use_internal_rev_list)\n    -+\t} else if (cruft)\n    ++\t} else if (cruft) {\n     +\t\tread_cruft_objects();\n    -+\telse if (!use_internal_rev_list)\n    + \t} else if (!use_internal_rev_list) {\n      \t\tread_object_list_from_stdin();\n    - \telse {\n    - \t\tget_object_list(rp.nr, rp.v);\n    + \t} else {\n     \n      ## object-file.c ##\n     @@ object-file.c: int has_loose_object_nonlocal(const struct object_id *oid)\n    @@ object-store.h: int repo_has_object_file_with_flags(struct repository *r,\n      \n      /*\n     \n    - ## t/t5327-pack-objects-cruft.sh (new) ##\n    + ## t/t5328-pack-objects-cruft.sh (new) ##\n     @@\n     +#!/bin/sh\n     +\n    @@ t/t5327-pack-objects-cruft.sh (new)\n     +\n     +\t\t\tgit reset --hard reachable &&\n     +\t\t\tgit tag -d cruft &&\n    -+\t\t\trm -fr .git/logs &&\n    ++\t\t\tgit reflog expire --all --expire=all &&\n     +\n     +\t\t\t# remove the unreachable tree, but leave the commit\n    -+\t\t\t# which has it as its root tree in-tact\n    ++\t\t\t# which has it as its root tree intact\n     +\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$tree\")\" &&\n     +\n     +\t\t\tgit repack -Ad &&\n    @@ t/t5327-pack-objects-cruft.sh (new)\n     +\n     +\t\t\tgit reset --hard reachable &&\n     +\t\t\tgit tag -d cruft &&\n    -+\t\t\trm -fr .git/logs &&\n    ++\t\t\tgit reflog expire --all --expire=all &&\n     +\n     +\t\t\t# remove the unreachable blob, but leave the commit (and\n    -+\t\t\t# the root tree of that commit) in-tact\n    ++\t\t\t# the root tree of that commit) intact\n     +\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$blob\")\" &&\n     +\n     +\t\t\tgit repack -Ad &&\n 9:  02f7fce788 =  9:  6f0e84273f reachable: add options to add_unseen_recent_objects_to_traversal\n10:  52e9ac5710 = 10:  a8bde361f9 reachable: report precise timestamps from objects in cruft packs\n11:  37fda94785 ! 11:  d68ce28132 builtin/pack-objects.c: --cruft with expiration\n    @@ Commit message\n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n     \n      ## builtin/pack-objects.c ##\n    -@@ builtin/pack-objects.c: static int add_cruft_object_entry(const struct object_id *oid, enum object_type\n    - \treturn 1;\n    +@@ builtin/pack-objects.c: static void add_cruft_object_entry(const struct object_id *oid, enum object_type\n    + \treturn;\n      }\n      \n     +static void show_cruft_object(struct object *obj, const char *name, void *data)\n    @@ builtin/pack-objects.c: static void read_cruft_objects(void)\n      \t\tenumerate_cruft_objects();\n      \n     \n    - ## t/t5327-pack-objects-cruft.sh ##\n    -@@ t/t5327-pack-objects-cruft.sh: basic_cruft_pack_tests () {\n    + ## t/t5328-pack-objects-cruft.sh ##\n    +@@ t/t5328-pack-objects-cruft.sh: basic_cruft_pack_tests () {\n      }\n      \n      basic_cruft_pack_tests never\n12:  a05675ab83 ! 12:  e5317cd472 builtin/repack.c: support generating a cruft pack\n    @@ builtin/repack.c: static int write_midx_included_packs(struct string_list *inclu\n      {\n      \tstruct child_process cmd = CHILD_PROCESS_INIT;\n     @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix)\n    - \tint show_progress = isatty(2);\n    + \tint show_progress;\n      \n      \t/* variables to be filled by option parsing */\n     -\tint pack_everything = 0;\n    @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix\n      \t\t\t\t   LOOSEN_UNREACHABLE | ALL_INTO_ONE),\n     +\t\tOPT_BIT(0, \"cruft\", &pack_everything,\n     +\t\t\t\tN_(\"same as -a, pack unreachable cruft objects separately\"),\n    -+\t\t\t\t   PACK_CRUFT | ALL_INTO_ONE),\n    ++\t\t\t\t   PACK_CRUFT),\n     +\t\tOPT_STRING(0, \"cruft-expiration\", &cruft_expiration, N_(\"approxidate\"),\n     +\t\t\t\tN_(\"with -C, expire objects older than this\")),\n      \t\tOPT_BOOL('d', NULL, &delete_redundant,\n      \t\t\t\tN_(\"remove redundant packs, and run git-prune-packed\")),\n      \t\tOPT_BOOL('f', NULL, &po_args.no_reuse_delta,\n     @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix)\n    - \tif (keep_unreachable &&\n      \t    (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE)))\n    - \t\tdie(_(\"--keep-unreachable and -A are incompatible\"));\n    -+\tif (pack_everything & PACK_CRUFT && delete_redundant) {\n    + \t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--keep-unreachable\", \"-A\");\n    + \n    ++\tif (pack_everything & PACK_CRUFT) {\n    ++\t\tpack_everything |= ALL_INTO_ONE;\n    ++\n     +\t\tif (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE))\n    -+\t\t\tdie(_(\"--cruft and -A are incompatible\"));\n    ++\t\t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--cruft\", \"-A\");\n     +\t\tif (keep_unreachable)\n    -+\t\t\tdie(_(\"--cruft and -k are incompatible\"));\n    -+\t\tif (!(pack_everything & ALL_INTO_ONE))\n    -+\t\t\tdie(_(\"--cruft must be combined with all-into-one\"));\n    ++\t\t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--cruft\", \"-k\");\n     +\t}\n    - \n    ++\n      \tif (write_bitmaps < 0) {\n      \t\tif (!write_midx &&\n    + \t\t    (!(pack_everything & ALL_INTO_ONE) || !is_bare_repository()))\n     @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix)\n      \tif (pack_everything & ALL_INTO_ONE) {\n      \t\trepack_promisor_objects(&po_args, &names);\n    @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix\n      \t\t\tfor_each_string_list_item(item, &names) {\n      \t\t\t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s-%s.pack\",\n      \t\t\t\t\t     packtmp_name, item->string);\n    -@@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix)\n    - \t\treturn ret;\n    - \n    - \tif (geometry) {\n    -+\t\tstruct packed_git *p;\n    - \t\tFILE *in = xfdopen(cmd.in, \"w\");\n    - \t\t/*\n    - \t\t * The resulting pack should contain all objects in packs that\n    -@@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix)\n    - \t\t\tfprintf(in, \"%s\\n\", pack_basename(geometry->pack[i]));\n    - \t\tfor (i = geometry->split; i < geometry->pack_nr; i++)\n    - \t\t\tfprintf(in, \"^%s\\n\", pack_basename(geometry->pack[i]));\n    -+\n    -+\t\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n    -+\t\t\tif (!p->is_cruft)\n    -+\t\t\t\tcontinue;\n    -+\t\t\tfprintf(in, \"^%s\\n\", pack_basename(p));\n    -+\t\t}\n    - \t\tfclose(in);\n    - \t}\n    - \n     @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix)\n      \tif (!names.nr && !po_args.quiet)\n      \t\tprintf_ln(_(\"Nothing new to pack.\"));\n    @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix\n      \t\titem->util = (void *)(uintptr_t)populate_pack_exts(item->string);\n      \t}\n     \n    - ## t/t5327-pack-objects-cruft.sh ##\n    -@@ t/t5327-pack-objects-cruft.sh: test_expect_success 'expired objects are pruned' '\n    + ## t/t5328-pack-objects-cruft.sh ##\n    +@@ t/t5328-pack-objects-cruft.sh: test_expect_success 'expired objects are pruned' '\n      \t)\n      '\n      \n    @@ t/t5327-pack-objects-cruft.sh: test_expect_success 'expired objects are pruned'\n     +\t\tgit branch -D other &&\n     +\t\tgit tag -d unreachable &&\n     +\t\t# objects are not cruft if they are contained in the reflogs\n    -+\t\trm -fr .git/logs &&\n    ++\t\tgit reflog expire --all --expire=all &&\n     +\n     +\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n     +\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n    @@ t/t5327-pack-objects-cruft.sh: test_expect_success 'expired objects are pruned'\n     +\t\tgit checkout main &&\n     +\t\tgit branch -D other &&\n     +\t\tgit tag -d cruft &&\n    -+\t\trm -fr .git/logs &&\n    ++\t\tgit reflog expire --all --expire=all &&\n     +\n     +\t\tgit repack --cruft -d &&\n     +\n    @@ t/t5327-pack-objects-cruft.sh: test_expect_success 'expired objects are pruned'\n     +\t\tgit checkout main &&\n     +\t\tgit branch -D other &&\n     +\t\tgit tag -d cruft &&\n    -+\t\trm -fr .git/logs &&\n    ++\t\tgit reflog expire --all --expire=all &&\n     +\n     +\t\tgit repack --cruft &&\n     +\n    @@ t/t5327-pack-objects-cruft.sh: test_expect_success 'expired objects are pruned'\n     +\t\ttest_cmp before after\n     +\t)\n     +'\n    ++\n    ++test_expect_success 'repack --geometric collects once-cruft objects' '\n    ++\tgit init repo &&\n    ++\ttest_when_finished \"rm -fr repo\" &&\n    ++\t(\n    ++\t\tcd repo &&\n    ++\n    ++\t\ttest_commit reachable &&\n    ++\t\tgit repack -Ad &&\n    ++\t\tgit branch -M main &&\n    ++\n    ++\t\tgit checkout --orphan other &&\n    ++\t\tgit rm -rf . &&\n    ++\t\ttest_commit --no-tag cruft &&\n    ++\t\tcruft=\"$(git rev-parse HEAD)\" &&\n    ++\n    ++\t\tgit checkout main &&\n    ++\t\tgit branch -D other &&\n    ++\t\tgit reflog expire --all --expire=all &&\n    ++\n    ++\t\t# Pack the objects created in the previous step into a cruft\n    ++\t\t# pack. Intentionally leave loose copies of those objects\n    ++\t\t# around so we can pick them up in a subsequent --geometric\n    ++\t\t# reapack.\n    ++\t\tgit repack --cruft &&\n    ++\n    ++\t\t# Now make those objects reachable, and ensure that they are\n    ++\t\t# packed into the new pack created via a --geometric repack.\n    ++\t\tgit update-ref refs/heads/other $cruft &&\n    ++\n    ++\t\t# Without this object, the set of unpacked objects is exactly\n    ++\t\t# the set of objects already in the cruft pack. Tweak that set\n    ++\t\t# to ensure we do not overwrite the cruft pack entirely.\n    ++\t\ttest_commit reachable2 &&\n    ++\n    ++\t\tfind $packdir -name \"pack-*.idx\" | sort >before &&\n    ++\t\tgit repack --geometric=2 -d &&\n    ++\t\tfind $packdir -name \"pack-*.idx\" | sort >after &&\n    ++\n    ++\t\t{\n    ++\t\t\tgit rev-list --objects --no-object-names $cruft &&\n    ++\t\t\tgit rev-list --objects --no-object-names reachable..reachable2\n    ++\t\t} >want.raw &&\n    ++\t\tsort want.raw >want &&\n    ++\n    ++\t\tpack=$(comm -13 before after) &&\n    ++\t\tgit show-index <$pack >objects.raw &&\n    ++\n    ++\t\tcut -d\" \" -f2 objects.raw | sort >got &&\n    ++\n    ++\t\ttest_cmp want got\n    ++\t)\n    ++'\n    ++\n     +test_expect_success 'cruft repack with no reachable objects' '\n     +\tgit init repo &&\n     +\ttest_when_finished \"rm -fr repo\" &&\n    @@ t/t5327-pack-objects-cruft.sh: test_expect_success 'expired objects are pruned'\n     +\n     +\t\tgit for-each-ref --format=\"delete %(refname)\" >in &&\n     +\t\tgit update-ref --stdin <in &&\n    -+\t\trm -fr .git/logs &&\n    ++\t\tgit reflog expire --all --expire=all &&\n     +\t\trm -fr .git/index &&\n     +\n     +\t\tgit repack --cruft -d &&\n13:  0d2dfaa062 ! 13:  b548dbbf80 builtin/repack.c: allow configuring cruft pack generation\n    @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix\n      \t\t\t\t       &existing_kept_packs);\n      \t\tif (ret)\n     \n    - ## t/t5327-pack-objects-cruft.sh ##\n    -@@ t/t5327-pack-objects-cruft.sh: test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n    + ## t/t5328-pack-objects-cruft.sh ##\n    +@@ t/t5328-pack-objects-cruft.sh: test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n      \t)\n      '\n      \n14:  fd50c39657 = 14:  e6eee7f15c builtin/repack.c: use named flags for existing_packs\n15:  b2937ceda7 ! 15:  b09dbc9fe5 builtin/repack.c: add cruft packs to MIDX during geometric repack\n    @@ builtin/repack.c: static void midx_included_packs(struct string_list *include,\n      \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n      \t\t\tif ((uintptr_t)item->util & DELETE_PACK)\n     \n    - ## t/t5327-pack-objects-cruft.sh ##\n    -@@ t/t5327-pack-objects-cruft.sh: test_expect_success 'cruft --local drops unreachable objects' '\n    + ## t/t5328-pack-objects-cruft.sh ##\n    +@@ t/t5328-pack-objects-cruft.sh: test_expect_success 'cruft --local drops unreachable objects' '\n      \t)\n      '\n      \n    @@ t/t5327-pack-objects-cruft.sh: test_expect_success 'cruft --local drops unreacha\n     +\n     +\t\tgit reset --hard $unreachable^ &&\n     +\t\tgit tag -d cruft &&\n    -+\t\trm -fr .git/logs &&\n    ++\t\tgit reflog expire --all --expire=all &&\n     +\n     +\t\tgit repack --cruft -d &&\n     +\n16:  394de0199f ! 16:  7a21ae1494 builtin/gc.c: conditionally avoid pruning objects via loose\n    @@ builtin/gc.c: int cmd_gc(int argc, const char **argv, const char *prefix)\n      \t\t\tif (quiet)\n      \t\t\t\tstrvec_push(&prune, \"--no-progress\");\n     \n    - ## t/t5327-pack-objects-cruft.sh ##\n    -@@ t/t5327-pack-objects-cruft.sh: test_expect_success 'loose objects mtimes upsert others' '\n    + ## t/t5328-pack-objects-cruft.sh ##\n    +@@ t/t5328-pack-objects-cruft.sh: test_expect_success 'loose objects mtimes upsert others' '\n      \t)\n      '\n      \n    @@ t/t5327-pack-objects-cruft.sh: test_expect_success 'loose objects mtimes upsert\n     +\t\tgit branch -D other &&\n     +\t\tgit tag -d unreachable &&\n     +\t\t# objects are not cruft if they are contained in the reflogs\n    -+\t\trm -fr .git/logs &&\n    ++\t\tgit reflog expire --all --expire=all &&\n     +\n     +\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n     +\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n17:  99aace8e16 ! 17:  b729b80963 sha1-file.c: don't freshen cruft packs\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n      \t\treturn 1;\n      \tif (!freshen_file(e.p->pack_name))\n     \n    - ## t/t5327-pack-objects-cruft.sh ##\n    -@@ t/t5327-pack-objects-cruft.sh: test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n    + ## t/t5328-pack-objects-cruft.sh ##\n    +@@ t/t5328-pack-objects-cruft.sh: test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n      \t)\n      '\n      \n-- \n2.35.1.73.gccc5557600\n"},{"id":"450031","messageId":"784ee7e0eec9ba520ebaaa27de2de810e2f6798a.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:00Z","receivedAt":"2022-03-02T00:58:11Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Create a technical document to explain cruft packs. It contains a brief\noverview of the problem, some background, details on the implementation,\nand a couple of alternative approaches not considered here.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/Makefile                  |  1 +\n Documentation/technical/cruft-packs.txt | 97 +++++++++++++++++++++++++\n 2 files changed, 98 insertions(+)\n create mode 100644 Documentation/technical/cruft-packs.txt\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex ed656db2ae..0b01c9408e 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -91,6 +91,7 @@ TECH_DOCS += MyFirstContribution\n TECH_DOCS += MyFirstObjectWalk\n TECH_DOCS += SubmittingPatches\n TECH_DOCS += technical/bundle-format\n+TECH_DOCS += technical/cruft-packs\n TECH_DOCS += technical/hash-function-transition\n TECH_DOCS += technical/http-protocol\n TECH_DOCS += technical/index-format\ndiff --git a/Documentation/technical/cruft-packs.txt b/Documentation/technical/cruft-packs.txt\nnew file mode 100644\nindex 0000000000..2c3c5d93f8\n--- /dev/null\n+++ b/Documentation/technical/cruft-packs.txt\n@@ -0,0 +1,97 @@\n+= Cruft packs\n+\n+The cruft packs feature offer an alternative to Git's traditional mechanism of\n+removing unreachable objects. This document provides an overview of Git's\n+pruning mechanism, and how a cruft pack can be used instead to accomplish the\n+same.\n+\n+== Background\n+\n+To remove unreachable objects from your repository, Git offers `git repack -Ad`\n+(see linkgit:git-repack[1]). Quoting from the documentation:\n+\n+[quote]\n+[...] unreachable objects in a previous pack become loose, unpacked objects,\n+instead of being left in the old pack. [...] loose unreachable objects will be\n+pruned according to normal expiry rules with the next 'git gc' invocation.\n+\n+Unreachable objects aren't removed immediately, since doing so could race with\n+an incoming push which may reference an object which is about to be deleted.\n+Instead, those unreachable objects are stored as loose object and stay that way\n+until they are older than the expiration window, at which point they are removed\n+by linkgit:git-prune[1].\n+\n+Git must store these unreachable objects loose in order to keep track of their\n+per-object mtimes. If these unreachable objects were written into one big pack,\n+then either freshening that pack (because an object contained within it was\n+re-written) or creating a new pack of unreachable objects would cause the pack's\n+mtime to get updated, and the objects within it would never leave the expiration\n+window. Instead, objects are stored loose in order to keep track of the\n+individual object mtimes and avoid a situation where all cruft objects are\n+freshened at once.\n+\n+This can lead to undesirable situations when a repository contains many\n+unreachable objects which have not yet left the grace period. Having large\n+directories in the shards of `.git/objects` can lead to decreased performance in\n+the repository. But given enough unreachable objects, this can lead to inode\n+starvation and degrade the performance of the whole system. Since we\n+can never pack those objects, these repositories often take up a large amount of\n+disk space, since we can only zlib compress them, but not store them in delta\n+chains.\n+\n+== Cruft packs\n+\n+A cruft pack eliminates the need for storing unreachable objects in a loose\n+state by including the per-object mtimes in a separate file alongside a single\n+pack containing all loose objects.\n+\n+A cruft pack is written by `git repack --cruft` when generating a new pack.\n+linkgit:git-pack-objects[1]'s `--cruft` option. Note that `git repack --cruft`\n+is a classic all-into-one repack, meaning that everything in the resulting pack is\n+reachable, and everything else is unreachable. Once written, the `--cruft`\n+option instructs `git repack` to generate another pack containing only objects\n+not packed in the previous step (which equates to packing all unreachable\n+objects together). This progresses as follows:\n+\n+  1. Enumerate every object, marking any object which is (a) not contained in a\n+     kept-pack, and (b) whose mtime is within the grace period as a traversal\n+     tip.\n+\n+  2. Perform a reachability traversal based on the tips gathered in the previous\n+     step, adding every object along the way to the pack.\n+\n+  3. Write the pack out, along with a `.mtimes` file that records the per-object\n+     timestamps.\n+\n+This mode is invoked internally by linkgit:git-repack[1] when instructed to\n+write a cruft pack. Crucially, the set of in-core kept packs is exactly the set\n+of packs which will not be deleted by the repack; in other words, they contain\n+all of the repository's reachable objects.\n+\n+When a repository already has a cruft pack, `git repack --cruft` typically only\n+adds objects to it. An exception to this is when `git repack` is given the\n+`--cruft-expiration` option, which allows the generated cruft pack to omit\n+expired objects instead of waiting for linkgit:git-gc[1] to expire those objects\n+later on.\n+\n+It is linkgit:git-gc[1] that is typically responsible for removing expired\n+unreachable objects.\n+\n+== Alternatives\n+\n+Notable alternatives to this design include:\n+\n+  - The location of the per-object mtime data, and\n+  - Storing unreachable objects in multiple cruft packs.\n+\n+On the location of mtime data, a new auxiliary file tied to the pack was chosen\n+to avoid complicating the `.idx` format. If the `.idx` format were ever to gain\n+support for optional chunks of data, it may make sense to consolidate the\n+`.mtimes` format into the `.idx` itself.\n+\n+Storing unreachable objects among multiple cruft packs (e.g., creating a new\n+cruft pack during each repacking operation including only unreachable objects\n+which aren't already stored in an earlier cruft pack) is significantly more\n+complicated to construct, and so aren't pursued here. The obvious drawback to\n+the current implementation is that the entire cruft pack must be re-written from\n+scratch.\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450032","messageId":"a94d7dfeb3fb847c3ad10e5b227344c5980defeb.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 03/17] pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:05Z","receivedAt":"2022-03-02T00:58:13Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This structure will be used to communicate the per-object mtimes when\nwriting a cruft pack. Here, we need the full packing_data structure\nbecause the mtime information is stored in an array there, not on the\nindividual object_entry's themselves (to avoid paying the overhead in\nstructure width for operations which do not generate a cruft pack).\n\nWe haven't passed this information down before because one of the two\ncallers (in bulk-checkin.c) does not have a packing_data structure at\nall. In that case (where no cruft pack will be generated), NULL is\npassed instead.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 3 ++-\n bulk-checkin.c         | 2 +-\n pack-write.c           | 1 +\n pack.h                 | 3 +++\n 4 files changed, 7 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 178e611f09..385970cb7b 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1254,7 +1254,8 @@ static void write_pack_file(void)\n \n \t\t\tstage_tmp_packfiles(&tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n-\t\t\t\t\t    &pack_idx_opts, hash, &idx_tmp_name);\n+\t\t\t\t\t    &to_pack, &pack_idx_opts, hash,\n+\t\t\t\t\t    &idx_tmp_name);\n \n \t\t\tif (write_bitmap_index) {\n \t\t\t\tsize_t tmpname_len = tmpname.len;\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac80..99f7596c4e 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -33,7 +33,7 @@ static void finish_tmp_packfile(struct strbuf *basename,\n \tchar *idx_tmp_name = NULL;\n \n \tstage_tmp_packfiles(basename, pack_tmp_name, written_list, nr_written,\n-\t\t\t    pack_idx_opts, hash, &idx_tmp_name);\n+\t\t\t    NULL, pack_idx_opts, hash, &idx_tmp_name);\n \trename_tmp_packfile_idx(basename, &idx_tmp_name);\n \n \tfree(idx_tmp_name);\ndiff --git a/pack-write.c b/pack-write.c\nindex a5846f3a34..d594e3008e 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -483,6 +483,7 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t const char *pack_tmp_name,\n \t\t\t struct pack_idx_entry **written_list,\n \t\t\t uint32_t nr_written,\n+\t\t\t struct packing_data *to_pack,\n \t\t\t struct pack_idx_option *pack_idx_opts,\n \t\t\t unsigned char hash[],\n \t\t\t char **idx_tmp_name)\ndiff --git a/pack.h b/pack.h\nindex b22bfc4a18..fd27cfdfd7 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -109,11 +109,14 @@ int encode_in_pack_object_header(unsigned char *hdr, int hdr_len,\n #define PH_ERROR_PROTOCOL\t(-3)\n int read_pack_header(int fd, struct pack_header *);\n \n+struct packing_data;\n+\n struct hashfile *create_tmp_packfile(char **pack_tmp_name);\n void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t const char *pack_tmp_name,\n \t\t\t struct pack_idx_entry **written_list,\n \t\t\t uint32_t nr_written,\n+\t\t\t struct packing_data *to_pack,\n \t\t\t struct pack_idx_option *pack_idx_opts,\n \t\t\t unsigned char hash[],\n \t\t\t char **idx_tmp_name);\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450033","messageId":"101b34660c0c5028ba591d052dc587bb8918ccb2.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:02Z","receivedAt":"2022-03-02T00:58:14Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"To store the individual mtimes of objects in a cruft pack, introduce a\nnew `.mtimes` format that can optionally accompany a single pack in the\nrepository.\n\nThe format is defined in Documentation/technical/pack-format.txt, and\nstores a 4-byte network order timestamp for each object in name (index)\norder.\n\nThis patch prepares for cruft packs by defining the `.mtimes` format,\nand introducing a basic API that callers can use to read out individual\nmtimes.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/technical/pack-format.txt |  19 ++++\n Makefile                                |   1 +\n builtin/repack.c                        |   1 +\n object-store.h                          |   5 +-\n pack-mtimes.c                           | 129 ++++++++++++++++++++++++\n pack-mtimes.h                           |  15 +++\n packfile.c                              |  19 +++-\n 7 files changed, 186 insertions(+), 3 deletions(-)\n create mode 100644 pack-mtimes.c\n create mode 100644 pack-mtimes.h\n\ndiff --git a/Documentation/technical/pack-format.txt b/Documentation/technical/pack-format.txt\nindex 6d3efb7d16..c443dbb526 100644\n--- a/Documentation/technical/pack-format.txt\n+++ b/Documentation/technical/pack-format.txt\n@@ -294,6 +294,25 @@ Pack file entry: <+\n \n All 4-byte numbers are in network order.\n \n+== pack-*.mtimes files have the format:\n+\n+  - A 4-byte magic number '0x4d544d45' ('MTME').\n+\n+  - A 4-byte version identifier (= 1).\n+\n+  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n+\n+  - A table of 4-byte unsigned integers in network order. The ith\n+    value is the modification time (mtime) of the ith object in the\n+    corresponding pack by lexicographic (index) order. The mtimes\n+    count standard epoch seconds.\n+\n+  - A trailer, containing a checksum of the corresponding packfile,\n+    and a checksum of all of the above (each having length according\n+    to the specified hash function).\n+\n+All 4-byte numbers are in network order.\n+\n == multi-pack-index (MIDX) files have the following format:\n \n The multi-pack-index files refer to multiple pack-files and loose objects.\ndiff --git a/Makefile b/Makefile\nindex 6f0b4b775f..1b186f4fd7 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -959,6 +959,7 @@ LIB_OBJS += oidtree.o\n LIB_OBJS += pack-bitmap-write.o\n LIB_OBJS += pack-bitmap.o\n LIB_OBJS += pack-check.o\n+LIB_OBJS += pack-mtimes.o\n LIB_OBJS += pack-objects.o\n LIB_OBJS += pack-revindex.o\n LIB_OBJS += pack-write.o\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex da1e364a75..f908f7d5dd 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -212,6 +212,7 @@ static struct {\n } exts[] = {\n \t{\".pack\"},\n \t{\".rev\", 1},\n+\t{\".mtimes\", 1},\n \t{\".bitmap\", 1},\n \t{\".promisor\", 1},\n \t{\".idx\"},\ndiff --git a/object-store.h b/object-store.h\nindex 6f89482df0..9b227661f2 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -115,12 +115,15 @@ struct packed_git {\n \t\t freshened:1,\n \t\t do_not_close:1,\n \t\t pack_promisor:1,\n-\t\t multi_pack_index:1;\n+\t\t multi_pack_index:1,\n+\t\t is_cruft:1;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n \tstruct revindex_entry *revindex;\n \tconst uint32_t *revindex_data;\n \tconst uint32_t *revindex_map;\n \tsize_t revindex_size;\n+\tconst uint32_t *mtimes_map;\n+\tsize_t mtimes_size;\n \t/* something like \".git/objects/pack/xxxxx.pack\" */\n \tchar pack_name[FLEX_ARRAY]; /* more */\n };\ndiff --git a/pack-mtimes.c b/pack-mtimes.c\nnew file mode 100644\nindex 0000000000..50caa34381\n--- /dev/null\n+++ b/pack-mtimes.c\n@@ -0,0 +1,129 @@\n+#include \"pack-mtimes.h\"\n+#include \"object-store.h\"\n+#include \"packfile.h\"\n+\n+static char *pack_mtimes_filename(struct packed_git *p)\n+{\n+\tsize_t len;\n+\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n+\t\tBUG(\"pack_name does not end in .pack\");\n+\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n+\treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n+}\n+\n+#define MTIMES_HEADER_SIZE (12)\n+#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n+\n+struct mtimes_header {\n+\tuint32_t signature;\n+\tuint32_t version;\n+\tuint32_t hash_id;\n+};\n+\n+static int load_pack_mtimes_file(char *mtimes_file,\n+\t\t\t\t uint32_t num_objects,\n+\t\t\t\t const uint32_t **data_p, size_t *len_p)\n+{\n+\tint fd, ret = 0;\n+\tstruct stat st;\n+\tvoid *data = NULL;\n+\tsize_t mtimes_size;\n+\tstruct mtimes_header header;\n+\tuint32_t *hdr;\n+\n+\tfd = git_open(mtimes_file);\n+\n+\tif (fd < 0) {\n+\t\tret = -1;\n+\t\tgoto cleanup;\n+\t}\n+\tif (fstat(fd, &st)) {\n+\t\tret = error_errno(_(\"failed to read %s\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tmtimes_size = xsize_t(st.st_size);\n+\n+\tif (mtimes_size < MTIMES_MIN_SIZE) {\n+\t\tret = error(_(\"mtimes file %s is too small\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t), num_objects)) {\n+\t\tret = error(_(\"mtimes file %s is corrupt\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tdata = hdr = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\n+\theader.signature = ntohl(hdr[0]);\n+\theader.version = ntohl(hdr[1]);\n+\theader.hash_id = ntohl(hdr[2]);\n+\n+\tif (header.signature != MTIMES_SIGNATURE) {\n+\t\tret = error(_(\"mtimes file %s has unknown signature\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (header.version != 1) {\n+\t\tret = error(_(\"mtimes file %s has unsupported version %\"PRIu32),\n+\t\t\t    mtimes_file, header.version);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (!(header.hash_id == 1 || header.hash_id == 2)) {\n+\t\tret = error(_(\"mtimes file %s has unsupported hash id %\"PRIu32),\n+\t\t\t    mtimes_file, header.hash_id);\n+\t\tgoto cleanup;\n+\t}\n+\n+cleanup:\n+\tif (ret) {\n+\t\tif (data)\n+\t\t\tmunmap(data, mtimes_size);\n+\t} else {\n+\t\t*len_p = mtimes_size;\n+\t\t*data_p = (const uint32_t *)data;\n+\t}\n+\n+\tclose(fd);\n+\treturn ret;\n+}\n+\n+int load_pack_mtimes(struct packed_git *p)\n+{\n+\tchar *mtimes_name = NULL;\n+\tint ret = 0;\n+\n+\tif (!p->is_cruft)\n+\t\treturn ret; /* not a cruft pack */\n+\tif (p->mtimes_map)\n+\t\treturn ret; /* already loaded */\n+\n+\tret = open_pack_index(p);\n+\tif (ret < 0)\n+\t\tgoto cleanup;\n+\n+\tmtimes_name = pack_mtimes_filename(p);\n+\tret = load_pack_mtimes_file(mtimes_name,\n+\t\t\t\t    p->num_objects,\n+\t\t\t\t    &p->mtimes_map,\n+\t\t\t\t    &p->mtimes_size);\n+\tif (ret)\n+\t\tgoto cleanup;\n+\n+cleanup:\n+\tfree(mtimes_name);\n+\treturn ret;\n+}\n+\n+uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos)\n+{\n+\tif (!p->mtimes_map)\n+\t\tBUG(\"pack .mtimes file not loaded for %s\", p->pack_name);\n+\tif (p->num_objects <= pos)\n+\t\tBUG(\"pack .mtimes out-of-bounds (%\"PRIu32\" vs %\"PRIu32\")\",\n+\t\t    pos, p->num_objects);\n+\n+\treturn get_be32(p->mtimes_map + pos + 3);\n+}\ndiff --git a/pack-mtimes.h b/pack-mtimes.h\nnew file mode 100644\nindex 0000000000..38ddb9f893\n--- /dev/null\n+++ b/pack-mtimes.h\n@@ -0,0 +1,15 @@\n+#ifndef PACK_MTIMES_H\n+#define PACK_MTIMES_H\n+\n+#include \"git-compat-util.h\"\n+\n+#define MTIMES_SIGNATURE 0x4d544d45 /* \"MTME\" */\n+#define MTIMES_VERSION 1\n+\n+struct packed_git;\n+\n+int load_pack_mtimes(struct packed_git *p);\n+\n+uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos);\n+\n+#endif\ndiff --git a/packfile.c b/packfile.c\nindex 835b2d2716..fc0245fbab 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -334,12 +334,22 @@ static void close_pack_revindex(struct packed_git *p)\n \tp->revindex_data = NULL;\n }\n \n+static void close_pack_mtimes(struct packed_git *p)\n+{\n+\tif (!p->mtimes_map)\n+\t\treturn;\n+\n+\tmunmap((void *)p->mtimes_map, p->mtimes_size);\n+\tp->mtimes_map = NULL;\n+}\n+\n void close_pack(struct packed_git *p)\n {\n \tclose_pack_windows(p);\n \tclose_pack_fd(p);\n \tclose_pack_index(p);\n \tclose_pack_revindex(p);\n+\tclose_pack_mtimes(p);\n \toidset_clear(&p->bad_objects);\n }\n \n@@ -363,7 +373,7 @@ void close_object_store(struct raw_object_store *o)\n \n void unlink_pack_path(const char *pack_name, int force_delete)\n {\n-\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\"};\n+\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\", \".mtimes\"};\n \tint i;\n \tstruct strbuf buf = STRBUF_INIT;\n \tsize_t plen;\n@@ -718,6 +728,10 @@ struct packed_git *add_packed_git(const char *path, size_t path_len, int local)\n \tif (!access(p->pack_name, F_OK))\n \t\tp->pack_promisor = 1;\n \n+\txsnprintf(p->pack_name + path_len, alloc - path_len, \".mtimes\");\n+\tif (!access(p->pack_name, F_OK))\n+\t\tp->is_cruft = 1;\n+\n \txsnprintf(p->pack_name + path_len, alloc - path_len, \".pack\");\n \tif (stat(p->pack_name, &st) || !S_ISREG(st.st_mode)) {\n \t\tfree(p);\n@@ -869,7 +883,8 @@ static void prepare_pack(const char *full_name, size_t full_name_len,\n \t    ends_with(file_name, \".pack\") ||\n \t    ends_with(file_name, \".bitmap\") ||\n \t    ends_with(file_name, \".keep\") ||\n-\t    ends_with(file_name, \".promisor\"))\n+\t    ends_with(file_name, \".promisor\") ||\n+\t    ends_with(file_name, \".mtimes\"))\n \t\tstring_list_append(data->garbage, full_name);\n \telse\n \t\treport_garbage(PACKDIR_FILE_GARBAGE, full_name);\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450034","messageId":"1e0ed363ae93099444b6626ff0a2043e8d88771d.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 04/17] chunk-format.h: extract oid_version()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:07Z","receivedAt":"2022-03-02T00:58:16Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"There are three definitions of an identical function which converts\n`the_hash_algo` into either 1 (for SHA-1) or 2 (for SHA-256). There is a\ncopy of this function for writing both the commit-graph and\nmulti-pack-index file, and another inline definition used to write the\n.rev header.\n\nConsolidate these into a single definition in chunk-format.h. It's not\nclear that this is the best header to define this function in, but it\nshould do for now.\n\n(Worth noting, the .rev caller expects a 4-byte unsigned, but the other\ntwo callers work with a single unsigned byte. The consolidated version\nuses the latter type, and lets the compiler widen it when required).\n\nAnother caller will be added in a subsequent patch.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n chunk-format.c | 12 ++++++++++++\n chunk-format.h |  3 +++\n commit-graph.c | 18 +++---------------\n midx.c         | 18 +++---------------\n pack-write.c   | 15 ++-------------\n 5 files changed, 23 insertions(+), 43 deletions(-)\n\ndiff --git a/chunk-format.c b/chunk-format.c\nindex 1c3dca62e2..0275b74a89 100644\n--- a/chunk-format.c\n+++ b/chunk-format.c\n@@ -181,3 +181,15 @@ int read_chunk(struct chunkfile *cf,\n \n \treturn CHUNK_NOT_FOUND;\n }\n+\n+uint8_t oid_version(const struct git_hash_algo *algop)\n+{\n+\tswitch (hash_algo_by_ptr(algop)) {\n+\tcase GIT_HASH_SHA1:\n+\t\treturn 1;\n+\tcase GIT_HASH_SHA256:\n+\t\treturn 2;\n+\tdefault:\n+\t\tdie(_(\"invalid hash version\"));\n+\t}\n+}\ndiff --git a/chunk-format.h b/chunk-format.h\nindex 9ccbe00377..7885aa0848 100644\n--- a/chunk-format.h\n+++ b/chunk-format.h\n@@ -2,6 +2,7 @@\n #define CHUNK_FORMAT_H\n \n #include \"git-compat-util.h\"\n+#include \"hash.h\"\n \n struct hashfile;\n struct chunkfile;\n@@ -65,4 +66,6 @@ int read_chunk(struct chunkfile *cf,\n \t       chunk_read_fn fn,\n \t       void *data);\n \n+uint8_t oid_version(const struct git_hash_algo *algop);\n+\n #endif\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 265c010122..f678d2c4a1 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -193,18 +193,6 @@ char *get_commit_graph_chain_filename(struct object_directory *odb)\n \treturn xstrfmt(\"%s/info/commit-graphs/commit-graph-chain\", odb->path);\n }\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n static struct commit_graph *alloc_commit_graph(void)\n {\n \tstruct commit_graph *g = xcalloc(1, sizeof(*g));\n@@ -365,9 +353,9 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n \t}\n \n \thash_version = *(unsigned char*)(data + 5);\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != oid_version(the_hash_algo)) {\n \t\terror(_(\"commit-graph hash version %X does not match version %X\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, oid_version(the_hash_algo));\n \t\treturn NULL;\n \t}\n \n@@ -1911,7 +1899,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \thashwrite_be32(f, GRAPH_SIGNATURE);\n \n \thashwrite_u8(f, GRAPH_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, oid_version(the_hash_algo));\n \thashwrite_u8(f, get_num_chunks(cf));\n \thashwrite_u8(f, ctx->num_commit_graphs_after - 1);\n \ndiff --git a/midx.c b/midx.c\nindex 865170bad0..65e670c5e2 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -41,18 +41,6 @@\n \n #define PACK_EXPIRED UINT_MAX\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n const unsigned char *get_midx_checksum(struct multi_pack_index *m)\n {\n \treturn m->data + m->data_len - the_hash_algo->rawsz;\n@@ -134,9 +122,9 @@ struct multi_pack_index *load_multi_pack_index(const char *object_dir, int local\n \t\t      m->version);\n \n \thash_version = m->data[MIDX_BYTE_HASH_VERSION];\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != oid_version(the_hash_algo)) {\n \t\terror(_(\"multi-pack-index hash version %u does not match version %u\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, oid_version(the_hash_algo));\n \t\tgoto cleanup_fail;\n \t}\n \tm->hash_len = the_hash_algo->rawsz;\n@@ -420,7 +408,7 @@ static size_t write_midx_header(struct hashfile *f,\n {\n \thashwrite_be32(f, MIDX_SIGNATURE);\n \thashwrite_u8(f, MIDX_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, oid_version(the_hash_algo));\n \thashwrite_u8(f, num_chunks);\n \thashwrite_u8(f, 0); /* unused */\n \thashwrite_be32(f, num_packs);\ndiff --git a/pack-write.c b/pack-write.c\nindex d594e3008e..ff305b404c 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -2,6 +2,7 @@\n #include \"pack.h\"\n #include \"csum-file.h\"\n #include \"remote.h\"\n+#include \"chunk-format.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -181,21 +182,9 @@ static int pack_order_cmp(const void *va, const void *vb, void *ctx)\n \n static void write_rev_header(struct hashfile *f)\n {\n-\tuint32_t oid_version;\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\toid_version = 1;\n-\t\tbreak;\n-\tcase GIT_HASH_SHA256:\n-\t\toid_version = 2;\n-\t\tbreak;\n-\tdefault:\n-\t\tdie(\"write_rev_header: unknown hash version\");\n-\t}\n-\n \thashwrite_be32(f, RIDX_SIGNATURE);\n \thashwrite_be32(f, RIDX_VERSION);\n-\thashwrite_be32(f, oid_version);\n+\thashwrite_be32(f, oid_version(the_hash_algo));\n }\n \n static void write_rev_index_positions(struct hashfile *f,\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450035","messageId":"78313bc4412fd480c91cc36d6032914ee79368c3.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 06/17] t/helper: add 'pack-mtimes' test-tool","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:12Z","receivedAt":"2022-03-02T00:58:18Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In the next patch, we will implement and test support for writing a\ncruft pack via a special mode of `git pack-objects`. To make sure that\nobjects are written with the correct timestamps, and a new test-tool\nthat can dump the object names and corresponding timestamps from a given\n`.mtimes` file.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Makefile                    |  1 +\n t/helper/test-pack-mtimes.c | 56 +++++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c        |  1 +\n t/helper/test-tool.h        |  1 +\n 4 files changed, 59 insertions(+)\n create mode 100644 t/helper/test-pack-mtimes.c\n\ndiff --git a/Makefile b/Makefile\nindex 1b186f4fd7..5c0ed1ade7 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -727,6 +727,7 @@ TEST_BUILTINS_OBJS += test-oid-array.o\n TEST_BUILTINS_OBJS += test-oidmap.o\n TEST_BUILTINS_OBJS += test-oidtree.o\n TEST_BUILTINS_OBJS += test-online-cpus.o\n+TEST_BUILTINS_OBJS += test-pack-mtimes.o\n TEST_BUILTINS_OBJS += test-parse-options.o\n TEST_BUILTINS_OBJS += test-parse-pathspec-file.o\n TEST_BUILTINS_OBJS += test-partial-clone.o\ndiff --git a/t/helper/test-pack-mtimes.c b/t/helper/test-pack-mtimes.c\nnew file mode 100644\nindex 0000000000..f7b79daf4c\n--- /dev/null\n+++ b/t/helper/test-pack-mtimes.c\n@@ -0,0 +1,56 @@\n+#include \"git-compat-util.h\"\n+#include \"test-tool.h\"\n+#include \"strbuf.h\"\n+#include \"object-store.h\"\n+#include \"packfile.h\"\n+#include \"pack-mtimes.h\"\n+\n+static void dump_mtimes(struct packed_git *p)\n+{\n+\tuint32_t i;\n+\tif (load_pack_mtimes(p) < 0)\n+\t\tdie(\"could not load pack .mtimes\");\n+\n+\tfor (i = 0; i < p->num_objects; i++) {\n+\t\tstruct object_id oid;\n+\t\tif (nth_packed_object_id(&oid, p, i) < 0)\n+\t\t\tdie(\"could not load object id at position %\"PRIu32, i);\n+\n+\t\tprintf(\"%s %\"PRIu32\"\\n\",\n+\t\t       oid_to_hex(&oid), nth_packed_mtime(p, i));\n+\t}\n+}\n+\n+static const char *pack_mtimes_usage = \"\\n\"\n+\"  test-tool pack-mtimes <pack-name.mtimes>\";\n+\n+int cmd__pack_mtimes(int argc, const char **argv)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct packed_git *p;\n+\n+\tsetup_git_directory();\n+\n+\tif (argc != 2)\n+\t\tusage(pack_mtimes_usage);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tstrbuf_addstr(&buf, basename(p->pack_name));\n+\t\tstrbuf_strip_suffix(&buf, \".pack\");\n+\t\tstrbuf_addstr(&buf, \".mtimes\");\n+\n+\t\tif (!strcmp(buf.buf, argv[1]))\n+\t\t\tbreak;\n+\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstrbuf_release(&buf);\n+\n+\tif (!p)\n+\t\tdie(\"could not find pack '%s'\", argv[1]);\n+\n+\tdump_mtimes(p);\n+\n+\treturn 0;\n+}\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex e6ec69cf32..7d472b31fd 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -47,6 +47,7 @@ static struct test_cmd cmds[] = {\n \t{ \"oidmap\", cmd__oidmap },\n \t{ \"oidtree\", cmd__oidtree },\n \t{ \"online-cpus\", cmd__online_cpus },\n+\t{ \"pack-mtimes\", cmd__pack_mtimes },\n \t{ \"parse-options\", cmd__parse_options },\n \t{ \"parse-pathspec-file\", cmd__parse_pathspec_file },\n \t{ \"partial-clone\", cmd__partial_clone },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex 20756eefdd..0ac4f32955 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -37,6 +37,7 @@ int cmd__mktemp(int argc, const char **argv);\n int cmd__oidmap(int argc, const char **argv);\n int cmd__oidtree(int argc, const char **argv);\n int cmd__online_cpus(int argc, const char **argv);\n+int cmd__pack_mtimes(int argc, const char **argv);\n int cmd__parse_options(int argc, const char **argv);\n int cmd__parse_pathspec_file(int argc, const char** argv);\n int cmd__partial_clone(int argc, const char **argv);\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450036","messageId":"5236490688213ff350b38f618ccb27f055300464.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 05/17] pack-mtimes: support writing pack .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:10Z","receivedAt":"2022-03-02T00:58:19Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Now that the `.mtimes` format is defined, supplement the pack-write API\nto be able to conditionally write an `.mtimes` file along with a pack by\nsetting an additional flag and passing an oidmap that contains the\ntimestamps corresponding to each object in the pack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n pack-objects.c |  6 ++++\n pack-objects.h | 25 ++++++++++++++++\n pack-write.c   | 77 ++++++++++++++++++++++++++++++++++++++++++++++++++\n pack.h         |  1 +\n 4 files changed, 109 insertions(+)\n\ndiff --git a/pack-objects.c b/pack-objects.c\nindex fe2a4eace9..272e8d4517 100644\n--- a/pack-objects.c\n+++ b/pack-objects.c\n@@ -170,6 +170,9 @@ struct object_entry *packlist_alloc(struct packing_data *pdata,\n \n \t\tif (pdata->layer)\n \t\t\tREALLOC_ARRAY(pdata->layer, pdata->nr_alloc);\n+\n+\t\tif (pdata->cruft_mtime)\n+\t\t\tREALLOC_ARRAY(pdata->cruft_mtime, pdata->nr_alloc);\n \t}\n \n \tnew_entry = pdata->objects + pdata->nr_objects++;\n@@ -198,6 +201,9 @@ struct object_entry *packlist_alloc(struct packing_data *pdata,\n \tif (pdata->layer)\n \t\tpdata->layer[pdata->nr_objects - 1] = 0;\n \n+\tif (pdata->cruft_mtime)\n+\t\tpdata->cruft_mtime[pdata->nr_objects - 1] = 0;\n+\n \treturn new_entry;\n }\n \ndiff --git a/pack-objects.h b/pack-objects.h\nindex dca2351ef9..393b9db546 100644\n--- a/pack-objects.h\n+++ b/pack-objects.h\n@@ -168,6 +168,14 @@ struct packing_data {\n \t/* delta islands */\n \tunsigned int *tree_depth;\n \tunsigned char *layer;\n+\n+\t/*\n+\t * Used when writing cruft packs.\n+\t *\n+\t * Object mtimes are stored in pack order when writing, but\n+\t * written out in lexicographic (index) order.\n+\t */\n+\tuint32_t *cruft_mtime;\n };\n \n void prepare_packing_data(struct repository *r, struct packing_data *pdata);\n@@ -289,4 +297,21 @@ static inline void oe_set_layer(struct packing_data *pack,\n \tpack->layer[e - pack->objects] = layer;\n }\n \n+static inline uint32_t oe_cruft_mtime(struct packing_data *pack,\n+\t\t\t\t      struct object_entry *e)\n+{\n+\tif (!pack->cruft_mtime)\n+\t\treturn 0;\n+\treturn pack->cruft_mtime[e - pack->objects];\n+}\n+\n+static inline void oe_set_cruft_mtime(struct packing_data *pack,\n+\t\t\t\t      struct object_entry *e,\n+\t\t\t\t      uint32_t mtime)\n+{\n+\tif (!pack->cruft_mtime)\n+\t\tCALLOC_ARRAY(pack->cruft_mtime, pack->nr_alloc);\n+\tpack->cruft_mtime[e - pack->objects] = mtime;\n+}\n+\n #endif\ndiff --git a/pack-write.c b/pack-write.c\nindex ff305b404c..270280c4df 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -3,6 +3,10 @@\n #include \"csum-file.h\"\n #include \"remote.h\"\n #include \"chunk-format.h\"\n+#include \"pack-mtimes.h\"\n+#include \"oidmap.h\"\n+#include \"chunk-format.h\"\n+#include \"pack-objects.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -276,6 +280,70 @@ const char *write_rev_file_order(const char *rev_name,\n \treturn rev_name;\n }\n \n+static void write_mtimes_header(struct hashfile *f)\n+{\n+\thashwrite_be32(f, MTIMES_SIGNATURE);\n+\thashwrite_be32(f, MTIMES_VERSION);\n+\thashwrite_be32(f, oid_version(the_hash_algo));\n+}\n+\n+/*\n+ * Writes the object mtimes of \"objects\" for use in a .mtimes file.\n+ * Note that objects must be in lexicographic (index) order, which is\n+ * the expected ordering of these values in the .mtimes file.\n+ */\n+static void write_mtimes_objects(struct hashfile *f,\n+\t\t\t\t struct packing_data *to_pack,\n+\t\t\t\t struct pack_idx_entry **objects,\n+\t\t\t\t uint32_t nr_objects)\n+{\n+\tuint32_t i;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\tstruct object_entry *e = (struct object_entry*)objects[i];\n+\t\thashwrite_be32(f, oe_cruft_mtime(to_pack, e));\n+\t}\n+}\n+\n+static void write_mtimes_trailer(struct hashfile *f, const unsigned char *hash)\n+{\n+\thashwrite(f, hash, the_hash_algo->rawsz);\n+}\n+\n+static const char *write_mtimes_file(const char *mtimes_name,\n+\t\t\t\t     struct packing_data *to_pack,\n+\t\t\t\t     struct pack_idx_entry **objects,\n+\t\t\t\t     uint32_t nr_objects,\n+\t\t\t\t     const unsigned char *hash)\n+{\n+\tstruct hashfile *f;\n+\tint fd;\n+\n+\tif (!to_pack)\n+\t\tBUG(\"cannot call write_mtimes_file with NULL packing_data\");\n+\n+\tif (!mtimes_name) {\n+\t\tstruct strbuf tmp_file = STRBUF_INIT;\n+\t\tfd = odb_mkstemp(&tmp_file, \"pack/tmp_mtimes_XXXXXX\");\n+\t\tmtimes_name = strbuf_detach(&tmp_file, NULL);\n+\t} else {\n+\t\tunlink(mtimes_name);\n+\t\tfd = xopen(mtimes_name, O_CREAT|O_EXCL|O_WRONLY, 0600);\n+\t}\n+\tf = hashfd(fd, mtimes_name);\n+\n+\twrite_mtimes_header(f);\n+\twrite_mtimes_objects(f, to_pack, objects, nr_objects);\n+\twrite_mtimes_trailer(f, hash);\n+\n+\tif (adjust_shared_perm(mtimes_name) < 0)\n+\t\tdie(_(\"failed to make %s readable\"), mtimes_name);\n+\n+\tfinalize_hashfile(f, NULL,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE | CSUM_FSYNC);\n+\n+\treturn mtimes_name;\n+}\n+\n off_t write_pack_header(struct hashfile *f, uint32_t nr_entries)\n {\n \tstruct pack_header hdr;\n@@ -478,6 +546,7 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t char **idx_tmp_name)\n {\n \tconst char *rev_tmp_name = NULL;\n+\tconst char *mtimes_tmp_name = NULL;\n \n \tif (adjust_shared_perm(pack_tmp_name))\n \t\tdie_errno(\"unable to make temporary pack file readable\");\n@@ -490,9 +559,17 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \trev_tmp_name = write_rev_file(NULL, written_list, nr_written, hash,\n \t\t\t\t      pack_idx_opts->flags);\n \n+\tif (pack_idx_opts->flags & WRITE_MTIMES) {\n+\t\tmtimes_tmp_name = write_mtimes_file(NULL, to_pack, written_list,\n+\t\t\t\t\t\t    nr_written,\n+\t\t\t\t\t\t    hash);\n+\t}\n+\n \trename_tmp_packfile(name_buffer, pack_tmp_name, \"pack\");\n \tif (rev_tmp_name)\n \t\trename_tmp_packfile(name_buffer, rev_tmp_name, \"rev\");\n+\tif (mtimes_tmp_name)\n+\t\trename_tmp_packfile(name_buffer, mtimes_tmp_name, \"mtimes\");\n }\n \n void write_promisor_file(const char *promisor_name, struct ref **sought, int nr_sought)\ndiff --git a/pack.h b/pack.h\nindex fd27cfdfd7..01d385903a 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -44,6 +44,7 @@ struct pack_idx_option {\n #define WRITE_IDX_STRICT 02\n #define WRITE_REV 04\n #define WRITE_REV_VERIFY 010\n+#define WRITE_MTIMES 020\n \n \tuint32_t version;\n \tuint32_t off32_limit;\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450037","messageId":"142098668d1ff6feea69be328ccc55119d14bf13.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 07/17] builtin/pack-objects.c: return from create_object_entry()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:15Z","receivedAt":"2022-03-02T00:58:28Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"A new caller in the next commit will want to immediately modify the\nobject_entry structure created by create_object_entry(). Instead of\nforcing that caller to wastefully look-up the entry we just created,\nreturn it from create_object_entry() instead.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 16 +++++++++-------\n 1 file changed, 9 insertions(+), 7 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 385970cb7b..3f08a3c63a 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1508,13 +1508,13 @@ static int want_object_in_pack(const struct object_id *oid,\n \treturn 1;\n }\n \n-static void create_object_entry(const struct object_id *oid,\n-\t\t\t\tenum object_type type,\n-\t\t\t\tuint32_t hash,\n-\t\t\t\tint exclude,\n-\t\t\t\tint no_try_delta,\n-\t\t\t\tstruct packed_git *found_pack,\n-\t\t\t\toff_t found_offset)\n+static struct object_entry *create_object_entry(const struct object_id *oid,\n+\t\t\t\t\t\tenum object_type type,\n+\t\t\t\t\t\tuint32_t hash,\n+\t\t\t\t\t\tint exclude,\n+\t\t\t\t\t\tint no_try_delta,\n+\t\t\t\t\t\tstruct packed_git *found_pack,\n+\t\t\t\t\t\toff_t found_offset)\n {\n \tstruct object_entry *entry;\n \n@@ -1531,6 +1531,8 @@ static void create_object_entry(const struct object_id *oid,\n \t}\n \n \tentry->no_try_delta = no_try_delta;\n+\n+\treturn entry;\n }\n \n static const char no_closure_warning[] = N_(\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450038","messageId":"2517a6be3d48a721dee6b5aa54f73b64e6abd1d6.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:17Z","receivedAt":"2022-03-02T00:58:36Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Teach `pack-objects` how to generate a cruft pack when no objects are\ndropped (i.e., `--cruft-expiration=never`). Later patches will teach\n`pack-objects` how to generate a cruft pack that prunes objects.\n\nWhen generating a cruft pack which does not prune objects, we want to\ncollect all unreachable objects into a single pack (noting and updating\ntheir mtimes as we accumulate them). Ordinary use will pass the result\nof a `git repack -A` as a kept pack, so when this patch says \"kept\npack\", readers should think \"reachable objects\".\n\nGenerating a non-expiring cruft packs works as follows:\n\n  - Callers provide a list of every pack they know about, and indicate\n    which packs are about to be removed.\n\n  - All packs which are going to be removed (we'll call these the\n    redundant ones) are marked as kept in-core.\n\n    Any packs the caller did not mention (but are known to the\n    `pack-objects` process) are also marked as kept in-core. Packs not\n    mentioned by the caller are assumed to be unknown to them, i.e.,\n    they entered the repository after the caller decided which packs\n    should be kept and which should be discarded.\n\n    Since we do not want to include objects in these \"unknown\" packs\n    (because we don't know which of their objects are or aren't\n    reachable), these are also marked as kept in-core.\n\n  - Then, we enumerate all objects in the repository, and add them to\n    our packing list if they do not appear in an in-core kept pack.\n\nThis results in a new cruft pack which contains all known objects that\naren't included in the kept packs. When the kept pack is the result of\n`git repack -A`, the resulting pack contains all unreachable objects.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-pack-objects.txt |  30 ++++\n builtin/pack-objects.c             | 201 +++++++++++++++++++++++++-\n object-file.c                      |   2 +-\n object-store.h                     |   2 +\n t/t5328-pack-objects-cruft.sh      | 218 +++++++++++++++++++++++++++++\n 5 files changed, 448 insertions(+), 5 deletions(-)\n create mode 100755 t/t5328-pack-objects-cruft.sh\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex f8344e1e5b..a9995a932c 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -13,6 +13,7 @@ SYNOPSIS\n \t[--no-reuse-delta] [--delta-base-offset] [--non-empty]\n \t[--local] [--incremental] [--window=<n>] [--depth=<n>]\n \t[--revs [--unpacked | --all]] [--keep-pack=<pack-name>]\n+\t[--cruft] [--cruft-expiration=<time>]\n \t[--stdout [--filter=<filter-spec>] | <base-name>]\n \t[--shallow] [--keep-true-parents] [--[no-]sparse] < <object-list>\n \n@@ -95,6 +96,35 @@ base-name::\n Incompatible with `--revs`, or options that imply `--revs` (such as\n `--all`), with the exception of `--unpacked`, which is compatible.\n \n+--cruft::\n+\tPacks unreachable objects into a separate \"cruft\" pack, denoted\n+\tby the existence of a `.mtimes` file. Typically used by `git\n+\trepack --cruft`. Callers provide a list of pack names and\n+\tindicate which packs will remain in the repository, along with\n+\twhich packs will be deleted (indicated by the `-` prefix). The\n+\tcontents of the cruft pack are all objects not contained in the\n+\tsurviving packs which have not exceeded the grace period (see\n+\t`--cruft-expiration` below), or which have exceeded the grace\n+\tperiod, but are reachable from an other object which hasn't.\n++\n+When the input lists a pack containing all reachable objects (and lists\n+all other packs as pending deletion), the corresponding cruft pack will\n+contain all unreachable objects (with mtime newer than the\n+`--cruft-expiration`) along with any unreachable objects whose mtime is\n+older than the `--cruft-expiration`, but are reachable from an\n+unreachable object whose mtime is newer than the `--cruft-expiration`).\n++\n+Incompatible with `--unpack-unreachable`, `--keep-unreachable`,\n+`--pack-loose-unreachable`, `--stdin-packs`, as well as any other\n+options which imply `--revs`. Also incompatible with `--max-pack-size`;\n+when this option is set, the maximum pack size is not inferred from\n+`pack.packSizeLimit`.\n+\n+--cruft-expiration=<approxidate>::\n+\tIf specified, objects are eliminated from the cruft pack if they\n+\thave an mtime older than `<approxidate>`. If unspecified (and\n+\tgiven `--cruft`), then no objects are eliminated.\n+\n --window=<n>::\n --depth=<n>::\n \tThese two options affect how the objects contained in\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 3f08a3c63a..5ba4fc9c2c 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -36,6 +36,7 @@\n #include \"trace2.h\"\n #include \"shallow.h\"\n #include \"promisor-remote.h\"\n+#include \"pack-mtimes.h\"\n \n /*\n  * Objects we are going to pack are collected in the `to_pack` structure.\n@@ -194,6 +195,8 @@ static int reuse_delta = 1, reuse_object = 1;\n static int keep_unreachable, unpack_unreachable, include_tag;\n static timestamp_t unpack_unreachable_expiration;\n static int pack_loose_unreachable;\n+static int cruft;\n+static timestamp_t cruft_expiration;\n static int local;\n static int have_non_local_packs;\n static int incremental;\n@@ -1252,6 +1255,9 @@ static void write_pack_file(void)\n \t\t\t\t\t&to_pack, written_list, nr_written);\n \t\t\t}\n \n+\t\t\tif (cruft)\n+\t\t\t\tpack_idx_opts.flags |= WRITE_MTIMES;\n+\n \t\t\tstage_tmp_packfiles(&tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n \t\t\t\t\t    &to_pack, &pack_idx_opts, hash,\n@@ -3389,6 +3395,135 @@ static void read_packs_list_from_stdin(void)\n \tstring_list_clear(&exclude_packs, 0);\n }\n \n+static void add_cruft_object_entry(const struct object_id *oid, enum object_type type,\n+\t\t\t\t   struct packed_git *pack, off_t offset,\n+\t\t\t\t   const char *name, uint32_t mtime)\n+{\n+\tstruct object_entry *entry;\n+\n+\tdisplay_progress(progress_state, ++nr_seen);\n+\n+\tentry = packlist_find(&to_pack, oid);\n+\tif (entry) {\n+\t\tif (name) {\n+\t\t\tentry->hash = pack_name_hash(name);\n+\t\t\tentry->no_try_delta = no_try_delta(name);\n+\t\t}\n+\t} else {\n+\t\tif (!want_object_in_pack(oid, 0, &pack, &offset))\n+\t\t\treturn;\n+\t\tif (!pack && type == OBJ_BLOB && !has_loose_object(oid)) {\n+\t\t\t/*\n+\t\t\t * If a traversed tree has a missing blob then we want\n+\t\t\t * to avoid adding that missing object to our pack.\n+\t\t\t *\n+\t\t\t * This only applies to missing blobs, not trees,\n+\t\t\t * because the traversal needs to parse sub-trees but\n+\t\t\t * not blobs.\n+\t\t\t *\n+\t\t\t * Note we only perform this check when we couldn't\n+\t\t\t * already find the object in a pack, so we're really\n+\t\t\t * limited to \"ensure non-tip blobs which don't exist in\n+\t\t\t * packs do exist via loose objects\". Confused?\n+\t\t\t */\n+\t\t\treturn;\n+\t\t}\n+\n+\t\tentry = create_object_entry(oid, type, pack_name_hash(name),\n+\t\t\t\t\t    0, name && no_try_delta(name),\n+\t\t\t\t\t    pack, offset);\n+\t}\n+\n+\tif (mtime > oe_cruft_mtime(&to_pack, entry))\n+\t\toe_set_cruft_mtime(&to_pack, entry, mtime);\n+\treturn;\n+}\n+\n+static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n+{\n+\tstruct string_list_item *item = NULL;\n+\tfor_each_string_list_item(item, packs) {\n+\t\tstruct packed_git *p = item->util;\n+\t\tif (!p)\n+\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n+\t\tp->pack_keep_in_core = keep;\n+\t}\n+}\n+\n+static void add_unreachable_loose_objects(void);\n+static void add_objects_in_unpacked_packs(void);\n+\n+static void enumerate_cruft_objects(void)\n+{\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Enumerating cruft objects\"), 0);\n+\n+\tadd_objects_in_unpacked_packs();\n+\tadd_unreachable_loose_objects();\n+\n+\tstop_progress(&progress_state);\n+}\n+\n+static void read_cruft_objects(void)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct string_list discard_packs = STRING_LIST_INIT_DUP;\n+\tstruct string_list fresh_packs = STRING_LIST_INIT_DUP;\n+\tstruct packed_git *p;\n+\n+\tignore_packed_keep_in_core = 1;\n+\n+\twhile (strbuf_getline(&buf, stdin) != EOF) {\n+\t\tif (!buf.len)\n+\t\t\tcontinue;\n+\n+\t\tif (*buf.buf == '-')\n+\t\t\tstring_list_append(&discard_packs, buf.buf + 1);\n+\t\telse\n+\t\t\tstring_list_append(&fresh_packs, buf.buf);\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstring_list_sort(&discard_packs);\n+\tstring_list_sort(&fresh_packs);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tconst char *pack_name = pack_basename(p);\n+\t\tstruct string_list_item *item;\n+\n+\t\titem = string_list_lookup(&fresh_packs, pack_name);\n+\t\tif (!item)\n+\t\t\titem = string_list_lookup(&discard_packs, pack_name);\n+\n+\t\tif (item) {\n+\t\t\titem->util = p;\n+\t\t} else {\n+\t\t\t/*\n+\t\t\t * This pack wasn't mentioned in either the \"fresh\" or\n+\t\t\t * \"discard\" list, so the caller didn't know about it.\n+\t\t\t *\n+\t\t\t * Mark it as kept so that its objects are ignored by\n+\t\t\t * add_unseen_recent_objects_to_traversal(). We'll\n+\t\t\t * unmark it before starting the traversal so it doesn't\n+\t\t\t * halt the traversal early.\n+\t\t\t */\n+\t\t\tp->pack_keep_in_core = 1;\n+\t\t}\n+\t}\n+\n+\tmark_pack_kept_in_core(&fresh_packs, 1);\n+\tmark_pack_kept_in_core(&discard_packs, 0);\n+\n+\tif (cruft_expiration)\n+\t\tdie(\"--cruft-expiration not yet implemented\");\n+\telse\n+\t\tenumerate_cruft_objects();\n+\n+\tstrbuf_release(&buf);\n+\tstring_list_clear(&discard_packs, 0);\n+\tstring_list_clear(&fresh_packs, 0);\n+}\n+\n static void read_object_list_from_stdin(void)\n {\n \tchar line[GIT_MAX_HEXSZ + 1 + PATH_MAX + 2];\n@@ -3521,7 +3656,24 @@ static int add_object_in_unpacked_pack(const struct object_id *oid,\n \t\t\t\t       uint32_t pos,\n \t\t\t\t       void *_data)\n {\n-\tadd_object_entry(oid, OBJ_NONE, \"\", 0);\n+\tif (cruft) {\n+\t\toff_t offset;\n+\t\ttime_t mtime;\n+\n+\t\tif (pack->is_cruft) {\n+\t\t\tif (load_pack_mtimes(pack) < 0)\n+\t\t\t\tdie(_(\"could not load cruft pack .mtimes\"));\n+\t\t\tmtime = nth_packed_mtime(pack, pos);\n+\t\t} else {\n+\t\t\tmtime = pack->mtime;\n+\t\t}\n+\t\toffset = nth_packed_object_offset(pack, pos);\n+\n+\t\tadd_cruft_object_entry(oid, OBJ_NONE, pack, offset,\n+\t\t\t\t       NULL, mtime);\n+\t} else {\n+\t\tadd_object_entry(oid, OBJ_NONE, \"\", 0);\n+\t}\n \treturn 0;\n }\n \n@@ -3545,7 +3697,19 @@ static int add_loose_object(const struct object_id *oid, const char *path,\n \t\treturn 0;\n \t}\n \n-\tadd_object_entry(oid, type, \"\", 0);\n+\tif (cruft) {\n+\t\tstruct stat st;\n+\t\tif (stat(path, &st) < 0) {\n+\t\t\tif (errno == ENOENT)\n+\t\t\t\treturn 0;\n+\t\t\treturn error_errno(\"unable to stat %s\", oid_to_hex(oid));\n+\t\t}\n+\n+\t\tadd_cruft_object_entry(oid, type, NULL, 0, NULL,\n+\t\t\t\t       st.st_mtime);\n+\t} else {\n+\t\tadd_object_entry(oid, type, \"\", 0);\n+\t}\n \treturn 0;\n }\n \n@@ -3864,6 +4028,20 @@ static int option_parse_unpack_unreachable(const struct option *opt,\n \treturn 0;\n }\n \n+static int option_parse_cruft_expiration(const struct option *opt,\n+\t\t\t\t\t const char *arg, int unset)\n+{\n+\tif (unset) {\n+\t\tcruft = 0;\n+\t\tcruft_expiration = 0;\n+\t} else {\n+\t\tcruft = 1;\n+\t\tif (arg)\n+\t\t\tcruft_expiration = approxidate(arg);\n+\t}\n+\treturn 0;\n+}\n+\n int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n {\n \tint use_internal_rev_list = 0;\n@@ -3936,6 +4114,10 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tOPT_CALLBACK_F(0, \"unpack-unreachable\", NULL, N_(\"time\"),\n \t\t  N_(\"unpack unreachable objects newer than <time>\"),\n \t\t  PARSE_OPT_OPTARG, option_parse_unpack_unreachable),\n+\t\tOPT_BOOL(0, \"cruft\", &cruft, N_(\"create a cruft pack\")),\n+\t\tOPT_CALLBACK_F(0, \"cruft-expiration\", NULL, N_(\"time\"),\n+\t\t  N_(\"expire cruft objects older than <time>\"),\n+\t\t  PARSE_OPT_OPTARG, option_parse_cruft_expiration),\n \t\tOPT_BOOL(0, \"sparse\", &sparse,\n \t\t\t N_(\"use the sparse reachability algorithm\")),\n \t\tOPT_BOOL(0, \"thin\", &thin,\n@@ -4062,7 +4244,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \n \tif (!HAVE_THREADS && delta_search_threads != 1)\n \t\twarning(_(\"no threads support, ignoring --threads\"));\n-\tif (!pack_to_stdout && !pack_size_limit)\n+\tif (!pack_to_stdout && !pack_size_limit && !cruft)\n \t\tpack_size_limit = pack_size_limit_cfg;\n \tif (pack_to_stdout && pack_size_limit)\n \t\tdie(_(\"--max-pack-size cannot be used to build a pack for transfer\"));\n@@ -4089,6 +4271,15 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (stdin_packs && use_internal_rev_list)\n \t\tdie(_(\"cannot use internal rev list with --stdin-packs\"));\n \n+\tif (cruft) {\n+\t\tif (use_internal_rev_list)\n+\t\t\tdie(_(\"cannot use internal rev list with --cruft\"));\n+\t\tif (stdin_packs)\n+\t\t\tdie(_(\"cannot use --stdin-packs with --cruft\"));\n+\t\tif (pack_size_limit)\n+\t\t\tdie(_(\"cannot use --max-pack-size with --cruft\"));\n+\t}\n+\n \t/*\n \t * \"soft\" reasons not to use bitmaps - for on-disk repack by default we want\n \t *\n@@ -4145,7 +4336,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t    the_repository);\n \tprepare_packing_data(the_repository, &to_pack);\n \n-\tif (progress)\n+\tif (progress && !cruft)\n \t\tprogress_state = start_progress(_(\"Enumerating objects\"), 0);\n \tif (stdin_packs) {\n \t\t/* avoids adding objects in excluded packs */\n@@ -4153,6 +4344,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tread_packs_list_from_stdin();\n \t\tif (rev_list_unpacked)\n \t\t\tadd_unreachable_loose_objects();\n+\t} else if (cruft) {\n+\t\tread_cruft_objects();\n \t} else if (!use_internal_rev_list) {\n \t\tread_object_list_from_stdin();\n \t} else {\ndiff --git a/object-file.c b/object-file.c\nindex 8be57f48de..e80da1368d 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -996,7 +996,7 @@ int has_loose_object_nonlocal(const struct object_id *oid)\n \treturn check_and_freshen_nonlocal(oid, 0);\n }\n \n-static int has_loose_object(const struct object_id *oid)\n+int has_loose_object(const struct object_id *oid)\n {\n \treturn check_and_freshen(oid, 0);\n }\ndiff --git a/object-store.h b/object-store.h\nindex 9b227661f2..6b025dc670 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -334,6 +334,8 @@ int repo_has_object_file_with_flags(struct repository *r,\n  */\n int has_loose_object_nonlocal(const struct object_id *);\n \n+int has_loose_object(const struct object_id *);\n+\n void assert_oid_type(const struct object_id *oid, enum object_type expect);\n \n /*\ndiff --git a/t/t5328-pack-objects-cruft.sh b/t/t5328-pack-objects-cruft.sh\nnew file mode 100755\nindex 0000000000..003ca7344e\n--- /dev/null\n+++ b/t/t5328-pack-objects-cruft.sh\n@@ -0,0 +1,218 @@\n+#!/bin/sh\n+\n+test_description='cruft pack related pack-objects tests'\n+. ./test-lib.sh\n+\n+objdir=.git/objects\n+packdir=$objdir/pack\n+\n+basic_cruft_pack_tests () {\n+\texpire=\"$1\"\n+\n+\ttest_expect_success \"unreachable loose objects are packed (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit base &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit loose &&\n+\n+\t\t\ttest-tool chmtime +2000 \"$objdir/$(test_oid_to_path \\\n+\t\t\t\t$(git rev-parse loose:loose.t))\" &&\n+\t\t\ttest-tool chmtime +1000 \"$objdir/$(test_oid_to_path \\\n+\t\t\t\t$(git rev-parse loose^{tree}))\" &&\n+\n+\t\t\t(\n+\t\t\t\tgit rev-list --objects --no-object-names base..loose |\n+\t\t\t\twhile read oid\n+\t\t\t\tdo\n+\t\t\t\t\tpath=\"$objdir/$(test_oid_to_path \"$oid\")\" &&\n+\t\t\t\t\tprintf \"%s %d\\n\" \"$oid\" \"$(test-tool chmtime --get \"$path\")\"\n+\t\t\t\tdone |\n+\t\t\t\tsort -k1\n+\t\t\t) >expect &&\n+\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t\t--cruft-expiration=\"$expire\" $packdir/pack)\" &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"unreachable packed objects are packed (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit packed &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit other &&\n+\n+\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\t\tother=\"$(git pack-objects --delta-base-offset \\\n+\t\t\t\t$packdir/pack <objects)\" &&\n+\t\t\tgit prune-packed &&\n+\n+\t\t\ttest-tool chmtime --get -100 \"$packdir/pack-$other.pack\" >expect &&\n+\n+\t\t\tcruft=\"$(git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-pack-$other.pack\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\t\tcut -d\" \" -f2 <actual.raw | sort -u >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"unreachable cruft objects are repacked (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit packed &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit other &&\n+\n+\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\n+\t\t\tcruft_a=\"$(echo $keep | git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack)\" &&\n+\t\t\tgit prune-packed &&\n+\t\t\tcruft_b=\"$(git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-pack-$cruft_a.pack\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft_a.mtimes\" >expect.raw &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft_b.mtimes\" >actual.raw &&\n+\n+\t\t\tsort <expect.raw >expect &&\n+\t\t\tsort <actual.raw >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"multiple cruft packs (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\tgit repack -Ad &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\n+\t\t\ttest_commit cruft &&\n+\t\t\tloose=\"$objdir/$(test_oid_to_path $(git rev-parse cruft))\" &&\n+\n+\t\t\t# generate three copies of the cruft object in different\n+\t\t\t# cruft packs, each with a unique mtime:\n+\t\t\t#   - one expired (1000 seconds ago)\n+\t\t\t#   - two non-expired (one 1000 seconds in the future,\n+\t\t\t#     one 1500 seconds in the future)\n+\t\t\ttest-tool chmtime =-1000 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-A <<-EOF &&\n+\t\t\t$keep\n+\t\t\tEOF\n+\t\t\ttest-tool chmtime =+1000 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-B <<-EOF &&\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\tEOF\n+\t\t\ttest-tool chmtime =+1500 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-C <<-EOF &&\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-B-*.pack))\n+\t\t\tEOF\n+\n+\t\t\t# ensure the resulting cruft pack takes the most recent\n+\t\t\t# mtime among all copies\n+\t\t\tcruft=\"$(git pack-objects --cruft \\\n+\t\t\t\t--cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-B-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-C-*.pack))\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\n+\t\t\ttest-tool pack-mtimes \"$(basename $(ls $packdir/pack-C-*.mtimes))\" >expect.raw &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\t\tsort expect.raw >expect &&\n+\t\t\tsort actual.raw >actual &&\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"cruft packs tolerate missing trees (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\ttest_commit cruft &&\n+\n+\t\t\ttree=\"$(git rev-parse cruft^{tree})\" &&\n+\n+\t\t\tgit reset --hard reachable &&\n+\t\t\tgit tag -d cruft &&\n+\t\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t\t# remove the unreachable tree, but leave the commit\n+\t\t\t# which has it as its root tree intact\n+\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$tree\")\" &&\n+\n+\t\t\tgit repack -Ad &&\n+\t\t\tbasename $(ls $packdir/pack-*.pack) >in &&\n+\t\t\tgit pack-objects --cruft --cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <in\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"cruft packs tolerate missing blobs (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\ttest_commit cruft &&\n+\n+\t\t\tblob=\"$(git rev-parse cruft:cruft.t)\" &&\n+\n+\t\t\tgit reset --hard reachable &&\n+\t\t\tgit tag -d cruft &&\n+\t\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t\t# remove the unreachable blob, but leave the commit (and\n+\t\t\t# the root tree of that commit) intact\n+\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$blob\")\" &&\n+\n+\t\t\tgit repack -Ad &&\n+\t\t\tbasename $(ls $packdir/pack-*.pack) >in &&\n+\t\t\tgit pack-objects --cruft --cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <in\n+\t\t)\n+\t'\n+}\n+\n+basic_cruft_pack_tests never\n+\n+test_done\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450039","messageId":"6f0e84273f78797c728058521969e73f8817b49c.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 09/17] reachable: add options to add_unseen_recent_objects_to_traversal","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:20Z","receivedAt":"2022-03-02T00:58:37Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This function behaves very similarly to what we will need in\npack-objects in order to implement cruft packs with expiration. But it\nis lacking a couple of things. Namely, it needs:\n\n  - a mechanism to communicate the timestamps of individual recent\n    objects to some external caller\n\n  - and, in the case of packed objects, our future caller will also want\n    to know the originating pack, as well as the offset within that pack\n    at which the object can be found\n\n  - finally, it needs a way to skip over packs which are marked as kept\n    in-core.\n\nTo address the first two, add a callback interface in this patch which\nreports the time of each recent object, as well as a (packed_git,\noff_t) pair for packed objects.\n\nLikewise, add a new option to the packed object iterators to skip over\npacks which are marked as kept in core. This option will become\nimplicitly tested in a future patch.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c |  2 +-\n reachable.c            | 51 +++++++++++++++++++++++++++++++++++-------\n reachable.h            |  9 +++++++-\n 3 files changed, 52 insertions(+), 10 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 5ba4fc9c2c..1ef333717d 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -3951,7 +3951,7 @@ static void get_object_list(int ac, const char **av)\n \tif (unpack_unreachable_expiration) {\n \t\trevs.ignore_missing_links = 1;\n \t\tif (add_unseen_recent_objects_to_traversal(&revs,\n-\t\t\t\tunpack_unreachable_expiration))\n+\t\t\t\tunpack_unreachable_expiration, NULL, 0))\n \t\t\tdie(_(\"unable to add recent objects\"));\n \t\tif (prepare_revision_walk(&revs))\n \t\t\tdie(_(\"revision walk setup failed\"));\ndiff --git a/reachable.c b/reachable.c\nindex 84e3d0d75e..0eb9909f47 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -60,9 +60,13 @@ static void mark_commit(struct commit *c, void *data)\n struct recent_data {\n \tstruct rev_info *revs;\n \ttimestamp_t timestamp;\n+\treport_recent_object_fn *cb;\n+\tint ignore_in_core_kept_packs;\n };\n \n static void add_recent_object(const struct object_id *oid,\n+\t\t\t      struct packed_git *pack,\n+\t\t\t      off_t offset,\n \t\t\t      timestamp_t mtime,\n \t\t\t      struct recent_data *data)\n {\n@@ -103,13 +107,29 @@ static void add_recent_object(const struct object_id *oid,\n \t\tdie(\"unable to lookup %s\", oid_to_hex(oid));\n \n \tadd_pending_object(data->revs, obj, \"\");\n+\tif (data->cb)\n+\t\tdata->cb(obj, pack, offset, mtime);\n+}\n+\n+static int want_recent_object(struct recent_data *data,\n+\t\t\t      const struct object_id *oid)\n+{\n+\tif (data->ignore_in_core_kept_packs &&\n+\t    has_object_kept_pack(oid, IN_CORE_KEEP_PACKS))\n+\t\treturn 0;\n+\treturn 1;\n }\n \n static int add_recent_loose(const struct object_id *oid,\n \t\t\t    const char *path, void *data)\n {\n \tstruct stat st;\n-\tstruct object *obj = lookup_object(the_repository, oid);\n+\tstruct object *obj;\n+\n+\tif (!want_recent_object(data, oid))\n+\t\treturn 0;\n+\n+\tobj = lookup_object(the_repository, oid);\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n@@ -126,7 +146,7 @@ static int add_recent_loose(const struct object_id *oid,\n \t\treturn error_errno(\"unable to stat %s\", oid_to_hex(oid));\n \t}\n \n-\tadd_recent_object(oid, st.st_mtime, data);\n+\tadd_recent_object(oid, NULL, 0, st.st_mtime, data);\n \treturn 0;\n }\n \n@@ -134,29 +154,43 @@ static int add_recent_packed(const struct object_id *oid,\n \t\t\t     struct packed_git *p, uint32_t pos,\n \t\t\t     void *data)\n {\n-\tstruct object *obj = lookup_object(the_repository, oid);\n+\tstruct object *obj;\n+\n+\tif (!want_recent_object(data, oid))\n+\t\treturn 0;\n+\n+\tobj = lookup_object(the_repository, oid);\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n-\tadd_recent_object(oid, p->mtime, data);\n+\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), p->mtime, data);\n \treturn 0;\n }\n \n int add_unseen_recent_objects_to_traversal(struct rev_info *revs,\n-\t\t\t\t\t   timestamp_t timestamp)\n+\t\t\t\t\t   timestamp_t timestamp,\n+\t\t\t\t\t   report_recent_object_fn *cb,\n+\t\t\t\t\t   int ignore_in_core_kept_packs)\n {\n \tstruct recent_data data;\n+\tenum for_each_object_flags flags;\n \tint r;\n \n \tdata.revs = revs;\n \tdata.timestamp = timestamp;\n+\tdata.cb = cb;\n+\tdata.ignore_in_core_kept_packs = ignore_in_core_kept_packs;\n \n \tr = for_each_loose_object(add_recent_loose, &data,\n \t\t\t\t  FOR_EACH_OBJECT_LOCAL_ONLY);\n \tif (r)\n \t\treturn r;\n-\treturn for_each_packed_object(add_recent_packed, &data,\n-\t\t\t\t      FOR_EACH_OBJECT_LOCAL_ONLY);\n+\n+\tflags = FOR_EACH_OBJECT_LOCAL_ONLY | FOR_EACH_OBJECT_PACK_ORDER;\n+\tif (ignore_in_core_kept_packs)\n+\t\tflags |= FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS;\n+\n+\treturn for_each_packed_object(add_recent_packed, &data, flags);\n }\n \n static int mark_object_seen(const struct object_id *oid,\n@@ -217,7 +251,8 @@ void mark_reachable_objects(struct rev_info *revs, int mark_reflog,\n \n \tif (mark_recent) {\n \t\trevs->ignore_missing_links = 1;\n-\t\tif (add_unseen_recent_objects_to_traversal(revs, mark_recent))\n+\t\tif (add_unseen_recent_objects_to_traversal(revs, mark_recent,\n+\t\t\t\t\t\t\t   NULL, 0))\n \t\t\tdie(\"unable to mark recent objects\");\n \t\tif (prepare_revision_walk(revs))\n \t\t\tdie(\"revision walk setup failed\");\ndiff --git a/reachable.h b/reachable.h\nindex 5df932ad8f..b776761baa 100644\n--- a/reachable.h\n+++ b/reachable.h\n@@ -1,11 +1,18 @@\n #ifndef REACHEABLE_H\n #define REACHEABLE_H\n \n+#include \"object.h\"\n+\n struct progress;\n struct rev_info;\n \n+typedef void report_recent_object_fn(const struct object *, struct packed_git *,\n+\t\t\t\t     off_t, time_t);\n+\n int add_unseen_recent_objects_to_traversal(struct rev_info *revs,\n-\t\t\t\t\t   timestamp_t timestamp);\n+\t\t\t\t\t   timestamp_t timestamp,\n+\t\t\t\t\t   report_recent_object_fn cb,\n+\t\t\t\t\t   int ignore_in_core_kept_packs);\n void mark_reachable_objects(struct rev_info *revs, int mark_reflog,\n \t\t\t    timestamp_t mark_recent, struct progress *);\n \n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450040","messageId":"d68ce281324097e10e4c1921d84c577bed6943e7.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 11/17] builtin/pack-objects.c: --cruft with expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:24Z","receivedAt":"2022-03-02T00:58:39Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a previous patch, pack-objects learned how to generate a cruft pack\nso long as no objects are dropped.\n\nThis patch teaches pack-objects to handle the case where a non-never\n`--cruft-expiration` value is passed. This case is slightly more\ncomplicated than before, because we want pack-objects to save\nunreachable objects which would have been pruned when there is another\nrecent (i.e., non-prunable) unreachable object which reaches the other.\nWe'll call these objects \"unreachable but reachable-from-recent\".\n\nHere is how pack-objects handles `--cruft-expiration`:\n\n  - Instead of adding all objects outside of the kept pack(s) into the\n    packing list, only handle the ones whose mtime is within the grace\n    period.\n\n  - Construct a reachability traversal whose tips are the\n    unreachable-but-recent objects.\n\n  - Then, walk along that traversal, stopping if we reach an object in\n    the kept pack. At each step along the traversal, we add the object\n    we are visiting to the packing list.\n\nIn the majority of these cases, any object we visit in this traversal\nwill already be in our packing list. But we will sometimes encounter\nreachable-from-recent cruft objects, which we want to retain even if\nthey aged out of the grace period.\n\nThe most subtle point of this process is that we actually don't need to\nbother to update the rescued object's mtime. Even though we will write\nan .mtimes file with a value that is older than the expiration window,\nit will continue to survive cruft repacks so long as any objects which\nreach it haven't aged out.\n\nThat is, a future repack will also exclude that object from the initial\npacking list, only to discover it later on when doing the reachability\ntraversal.\n\nFinally, stopping early once an object is found in a kept pack is safe\nto do because the kept packs ordinarily represent which packs will\nsurvive after repacking. Assuming that it _isn't_ safe to halt a\ntraversal early would mean that there is some ancestor object which is\nmissing, which implies repository corruption (i.e., the complete set of\nreachable objects isn't present).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c        |  84 +++++++++++++++++++-\n t/t5328-pack-objects-cruft.sh | 143 ++++++++++++++++++++++++++++++++++\n 2 files changed, 226 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 1ef333717d..fcac0b5c91 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -3439,6 +3439,44 @@ static void add_cruft_object_entry(const struct object_id *oid, enum object_type\n \treturn;\n }\n \n+static void show_cruft_object(struct object *obj, const char *name, void *data)\n+{\n+\t/*\n+\t * if we did not record it earlier, it's at least as old as our\n+\t * expiration value. Rather than find it exactly, just use that\n+\t * value.  This may bump it forward from its real mtime, but it\n+\t * will still be \"too old\" next time we run with the same\n+\t * expiration.\n+\t *\n+\t * if obj does appear in the packing list, this call is a noop (or may\n+\t * set the namehash).\n+\t */\n+\tadd_cruft_object_entry(&obj->oid, obj->type, NULL, 0, name, cruft_expiration);\n+}\n+\n+static void show_cruft_commit(struct commit *commit, void *data)\n+{\n+\tshow_cruft_object((struct object*)commit, NULL, data);\n+}\n+\n+static int cruft_include_check_obj(struct object *obj, void *data)\n+{\n+\treturn !has_object_kept_pack(&obj->oid, IN_CORE_KEEP_PACKS);\n+}\n+\n+static int cruft_include_check(struct commit *commit, void *data)\n+{\n+\treturn cruft_include_check_obj((struct object*)commit, data);\n+}\n+\n+static void set_cruft_mtime(const struct object *object,\n+\t\t\t    struct packed_git *pack,\n+\t\t\t    off_t offset, time_t mtime)\n+{\n+\tadd_cruft_object_entry(&object->oid, object->type, pack, offset, NULL,\n+\t\t\t       mtime);\n+}\n+\n static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n {\n \tstruct string_list_item *item = NULL;\n@@ -3464,6 +3502,50 @@ static void enumerate_cruft_objects(void)\n \tstop_progress(&progress_state);\n }\n \n+static void enumerate_and_traverse_cruft_objects(struct string_list *fresh_packs)\n+{\n+\tstruct packed_git *p;\n+\tstruct rev_info revs;\n+\tint ret;\n+\n+\trepo_init_revisions(the_repository, &revs, NULL);\n+\n+\trevs.tag_objects = 1;\n+\trevs.tree_objects = 1;\n+\trevs.blob_objects = 1;\n+\n+\trevs.include_check = cruft_include_check;\n+\trevs.include_check_obj = cruft_include_check_obj;\n+\n+\trevs.ignore_missing_links = 1;\n+\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Enumerating cruft objects\"), 0);\n+\tret = add_unseen_recent_objects_to_traversal(&revs, cruft_expiration,\n+\t\t\t\t\t\t     set_cruft_mtime, 1);\n+\tstop_progress(&progress_state);\n+\n+\tif (ret)\n+\t\tdie(_(\"unable to add cruft objects\"));\n+\n+\t/*\n+\t * Re-mark only the fresh packs as kept so that objects in\n+\t * unknown packs do not halt the reachability traversal early.\n+\t */\n+\tfor (p = get_all_packs(the_repository); p; p = p->next)\n+\t\tp->pack_keep_in_core = 0;\n+\tmark_pack_kept_in_core(fresh_packs, 1);\n+\n+\tif (prepare_revision_walk(&revs))\n+\t\tdie(_(\"revision walk setup failed\"));\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Traversing cruft objects\"), 0);\n+\tnr_seen = 0;\n+\ttraverse_commit_list(&revs, show_cruft_commit, show_cruft_object, NULL);\n+\n+\tstop_progress(&progress_state);\n+}\n+\n static void read_cruft_objects(void)\n {\n \tstruct strbuf buf = STRBUF_INIT;\n@@ -3515,7 +3597,7 @@ static void read_cruft_objects(void)\n \tmark_pack_kept_in_core(&discard_packs, 0);\n \n \tif (cruft_expiration)\n-\t\tdie(\"--cruft-expiration not yet implemented\");\n+\t\tenumerate_and_traverse_cruft_objects(&fresh_packs);\n \telse\n \t\tenumerate_cruft_objects();\n \ndiff --git a/t/t5328-pack-objects-cruft.sh b/t/t5328-pack-objects-cruft.sh\nindex 003ca7344e..939cdc297a 100755\n--- a/t/t5328-pack-objects-cruft.sh\n+++ b/t/t5328-pack-objects-cruft.sh\n@@ -214,5 +214,148 @@ basic_cruft_pack_tests () {\n }\n \n basic_cruft_pack_tests never\n+basic_cruft_pack_tests 2.weeks.ago\n+\n+test_expect_success 'cruft tags rescue tagged objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit tagged &&\n+\t\tgit tag -a annotated -m tag &&\n+\n+\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\twhile read oid\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $oid)\"\n+\t\tdone <objects &&\n+\n+\t\ttest-tool chmtime -500 \\\n+\t\t\t\"$objdir/$(test_oid_to_path $(git rev-parse annotated))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\t\tcut -f1 -d\" \" <actual.raw | sort >actual &&\n+\n+\t\t(\n+\t\t\tcat objects &&\n+\t\t\tgit rev-parse annotated\n+\t\t) >expect.raw &&\n+\t\tsort <expect.raw >expect &&\n+\n+\t\ttest_cmp expect actual &&\n+\t\tcat actual\n+\t)\n+'\n+\n+test_expect_success 'cruft commits rescue parents, trees' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit old &&\n+\t\ttest_commit new &&\n+\n+\t\tgit rev-list --objects --no-object-names packed..new >objects &&\n+\t\twhile read object\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $object)\"\n+\t\tdone <objects &&\n+\t\ttest-tool chmtime +500 \"$objdir/$(test_oid_to_path \\\n+\t\t\t$(git rev-parse HEAD))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\tcut -d\" \" -f1 <actual.raw | sort >actual &&\n+\t\tsort <objects >expect &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'cruft trees rescue sub-trees, blobs' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\tmkdir -p dir/sub &&\n+\t\techo foo >foo &&\n+\t\techo bar >dir/bar &&\n+\t\techo baz >dir/sub/baz &&\n+\n+\t\ttest_tick &&\n+\t\tgit add . &&\n+\t\tgit commit -m \"pruned\" &&\n+\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD^{tree}))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:foo))\" &&\n+\t\ttest-tool chmtime  -500 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/bar))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/sub))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/sub/baz))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\t\tcut -f1 -d\" \" <actual.raw | sort >actual &&\n+\n+\t\tgit rev-parse HEAD:dir HEAD:dir/bar HEAD:dir/sub HEAD:dir/sub/baz >expect.raw &&\n+\t\tsort <expect.raw >expect &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'expired objects are pruned' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit pruned &&\n+\n+\t\tgit rev-list --objects --no-object-names packed..pruned >objects &&\n+\t\twhile read object\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $object)\"\n+\t\tdone <objects &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual &&\n+\t\ttest_must_be_empty actual\n+\t)\n+'\n \n test_done\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450041","messageId":"a8bde361f9b99abe90727959208413fddea602e3.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 10/17] reachable: report precise timestamps from objects in cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:22Z","receivedAt":"2022-03-02T00:58:42Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When generating a cruft pack, the caller within pack-objects will want\nto know the precise timestamps of cruft objects (i.e., their\ncorresponding values in the .mtimes table) rather than the mtime of the\ncruft pack itself.\n\nTeach add_recent_packed() to lookup each object's precise mtime from the\n.mtimes file if one exists (indicated by the is_cruft bit on the\npacked_git structure).\n\nA couple of small things worth noting here:\n\n  - load_pack_mtimes() needs to be called before asking for\n    nth_packed_mtime(), and that call is done lazily here. That function\n    exits early if the .mtimes file has already been opened and parsed,\n    so only the first call is slow.\n\n  - Checking the is_cruft bit can be done without any extra work on the\n    caller's behalf, since it is set up for us automatically as a\n    side-effect of calling add_packed_git() (just like the 'pack_keep'\n    and 'pack_promisor' bits).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n reachable.c | 9 ++++++++-\n 1 file changed, 8 insertions(+), 1 deletion(-)\n\ndiff --git a/reachable.c b/reachable.c\nindex 0eb9909f47..9ec8e6bd5b 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -13,6 +13,7 @@\n #include \"worktree.h\"\n #include \"object-store.h\"\n #include \"pack-bitmap.h\"\n+#include \"pack-mtimes.h\"\n \n struct connectivity_progress {\n \tstruct progress *progress;\n@@ -155,6 +156,7 @@ static int add_recent_packed(const struct object_id *oid,\n \t\t\t     void *data)\n {\n \tstruct object *obj;\n+\ttimestamp_t mtime = p->mtime;\n \n \tif (!want_recent_object(data, oid))\n \t\treturn 0;\n@@ -163,7 +165,12 @@ static int add_recent_packed(const struct object_id *oid,\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n-\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), p->mtime, data);\n+\tif (p->is_cruft) {\n+\t\tif (load_pack_mtimes(p) < 0)\n+\t\t\tdie(_(\"could not load cruft pack .mtimes\"));\n+\t\tmtime = nth_packed_mtime(p, pos);\n+\t}\n+\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), mtime, data);\n \treturn 0;\n }\n \n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450042","messageId":"b548dbbf80962ba5168de26743fc9972af121c66.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 13/17] builtin/repack.c: allow configuring cruft pack generation","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:29Z","receivedAt":"2022-03-02T00:58:43Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In servers which set the pack.window configuration to a large value, we\ncan wind up spending quite a lot of time finding new bases when breaking\ndelta chains between reachable and unreachable objects while generating\na cruft pack.\n\nIntroduce a handful of `repack.cruft*` configuration variables to\ncontrol the parameters used by pack-objects when generating a cruft\npack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/repack.txt |  9 ++++\n builtin/repack.c                | 50 ++++++++++++++------\n t/t5328-pack-objects-cruft.sh   | 83 +++++++++++++++++++++++++++++++++\n 3 files changed, 128 insertions(+), 14 deletions(-)\n\ndiff --git a/Documentation/config/repack.txt b/Documentation/config/repack.txt\nindex 9c413e177e..fd18d1fb89 100644\n--- a/Documentation/config/repack.txt\n+++ b/Documentation/config/repack.txt\n@@ -25,3 +25,12 @@ repack.writeBitmaps::\n \tspace and extra time spent on the initial repack.  This has\n \tno effect if multiple packfiles are created.\n \tDefaults to true on bare repos, false otherwise.\n+\n+repack.cruftWindow::\n+repack.cruftWindowMemory::\n+repack.cruftDepth::\n+repack.cruftThreads::\n+\tParameters used by linkgit:git-pack-objects[1] when generating\n+\ta cruft pack and the respective parameters are not given over\n+\tthe command line. See similarly named `pack.*` configuration\n+\tvariables for defaults and meaning.\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex f7fb88bcf1..d61c78e94e 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -40,9 +40,21 @@ static const char incremental_bitmap_conflict_error[] = N_(\n \"--no-write-bitmap-index or disable the pack.writebitmaps configuration.\"\n );\n \n+struct pack_objects_args {\n+\tconst char *window;\n+\tconst char *window_memory;\n+\tconst char *depth;\n+\tconst char *threads;\n+\tconst char *max_pack_size;\n+\tint no_reuse_delta;\n+\tint no_reuse_object;\n+\tint quiet;\n+\tint local;\n+};\n \n static int repack_config(const char *var, const char *value, void *cb)\n {\n+\tstruct pack_objects_args *cruft_po_args = cb;\n \tif (!strcmp(var, \"repack.usedeltabaseoffset\")) {\n \t\tdelta_base_offset = git_config_bool(var, value);\n \t\treturn 0;\n@@ -61,6 +73,15 @@ static int repack_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"repack.cruftwindow\"))\n+\t\treturn git_config_string(&cruft_po_args->window, var, value);\n+\tif (!strcmp(var, \"repack.cruftwindowmemory\"))\n+\t\treturn git_config_string(&cruft_po_args->window_memory, var, value);\n+\tif (!strcmp(var, \"repack.cruftdepth\"))\n+\t\treturn git_config_string(&cruft_po_args->depth, var, value);\n+\tif (!strcmp(var, \"repack.cruftthreads\"))\n+\t\treturn git_config_string(&cruft_po_args->threads, var, value);\n+\n \treturn git_default_config(var, value, cb);\n }\n \n@@ -153,18 +174,6 @@ static void remove_redundant_pack(const char *dir_name, const char *base_name)\n \tstrbuf_release(&buf);\n }\n \n-struct pack_objects_args {\n-\tconst char *window;\n-\tconst char *window_memory;\n-\tconst char *depth;\n-\tconst char *threads;\n-\tconst char *max_pack_size;\n-\tint no_reuse_delta;\n-\tint no_reuse_object;\n-\tint quiet;\n-\tint local;\n-};\n-\n static void prepare_pack_objects(struct child_process *cmd,\n \t\t\t\t const struct pack_objects_args *args)\n {\n@@ -689,6 +698,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tint no_update_server_info = 0;\n \tstruct pack_objects_args po_args = {NULL};\n+\tstruct pack_objects_args cruft_po_args = {NULL};\n \tint geometric_factor = 0;\n \tint write_midx = 0;\n \n@@ -743,7 +753,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tOPT_END()\n \t};\n \n-\tgit_config(repack_config, NULL);\n+\tgit_config(repack_config, &cruft_po_args);\n \n \targc = parse_options(argc, argv, prefix, builtin_repack_options,\n \t\t\t\tgit_repack_usage, 0);\n@@ -918,7 +928,19 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tif (*pack_prefix == '/')\n \t\t\tpack_prefix++;\n \n-\t\tret = write_cruft_pack(&po_args, pack_prefix, &names,\n+\t\tif (!cruft_po_args.window)\n+\t\t\tcruft_po_args.window = po_args.window;\n+\t\tif (!cruft_po_args.window_memory)\n+\t\t\tcruft_po_args.window_memory = po_args.window_memory;\n+\t\tif (!cruft_po_args.depth)\n+\t\t\tcruft_po_args.depth = po_args.depth;\n+\t\tif (!cruft_po_args.threads)\n+\t\t\tcruft_po_args.threads = po_args.threads;\n+\n+\t\tcruft_po_args.local = po_args.local;\n+\t\tcruft_po_args.quiet = po_args.quiet;\n+\n+\t\tret = write_cruft_pack(&cruft_po_args, pack_prefix, &names,\n \t\t\t\t       &existing_nonkept_packs,\n \t\t\t\t       &existing_kept_packs);\n \t\tif (ret)\ndiff --git a/t/t5328-pack-objects-cruft.sh b/t/t5328-pack-objects-cruft.sh\nindex 06c550c958..e4744e4465 100755\n--- a/t/t5328-pack-objects-cruft.sh\n+++ b/t/t5328-pack-objects-cruft.sh\n@@ -565,4 +565,87 @@ test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n \t)\n '\n \n+test_expect_success 'cruft repack respects repack.cruftWindow' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tGIT_TRACE2_EVENT=$(pwd)/event.trace \\\n+\t\tgit -c pack.window=1 -c repack.cruftWindow=2 repack \\\n+\t\t       --cruft --window=3 &&\n+\n+\t\tgrep \"pack-objects.*--window=2.*--cruft\" event.trace\n+\t)\n+'\n+\n+test_expect_success 'cruft repack respects --window by default' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tGIT_TRACE2_EVENT=$(pwd)/event.trace \\\n+\t\tgit -c pack.window=2 repack --cruft --window=3 &&\n+\n+\t\tgrep \"pack-objects.*--window=3.*--cruft\" event.trace\n+\t)\n+'\n+\n+test_expect_success 'cruft repack respects --quiet' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\t\tGIT_PROGRESS_DELAY=0 git repack --cruft --quiet 2>err &&\n+\t\ttest_must_be_empty err\n+\t)\n+'\n+\n+test_expect_success 'cruft --local drops unreachable objects' '\n+\tgit init alternate &&\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr alternate repo\" &&\n+\n+\ttest_commit -C alternate base &&\n+\t# Pack all objects in alterate so that the cruft repack in \"repo\" sees\n+\t# the object it dropped due to `--local` as packed. Otherwise this\n+\t# object would not appear packed anywhere (since it is not packed in\n+\t# alternate and likewise not part of the cruft pack in the other repo\n+\t# because of `--local`).\n+\tgit -C alternate repack -ad &&\n+\n+\t(\n+\t\tcd repo &&\n+\n+\t\tobject=\"$(git -C ../alternate rev-parse HEAD:base.t)\" &&\n+\t\tgit -C ../alternate cat-file -p $object >contents &&\n+\n+\t\t# Write some reachable objects and two unreachable ones: one\n+\t\t# that the alternate has and another that is unique.\n+\t\ttest_commit other &&\n+\t\tgit hash-object -w -t blob contents &&\n+\t\tcruft=\"$(echo cruft | git hash-object -w -t blob --stdin)\" &&\n+\n+\t\t( cd ../alternate/.git/objects && pwd ) \\\n+\t\t       >.git/objects/info/alternates &&\n+\n+\t\ttest_path_is_file $objdir/$(test_oid_to_path $cruft) &&\n+\t\ttest_path_is_file $objdir/$(test_oid_to_path $object) &&\n+\n+\t\tgit repack -d --cruft --local &&\n+\n+\t\ttest-tool pack-mtimes \"$(basename $(ls $packdir/pack-*.mtimes))\" \\\n+\t\t       >objects &&\n+\t\t! grep $object objects &&\n+\t\tgrep $cruft objects\n+\t)\n+'\n+\n test_done\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450044","messageId":"e5317cd472999faf3fbacdfe122fe5a794e98725.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 12/17] builtin/repack.c: support generating a cruft pack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:27Z","receivedAt":"2022-03-02T00:58:44Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Expose a way to split the contents of a repository into a main and cruft\npack when doing an all-into-one repack with `git repack --cruft -d`, and\na complementary configuration variable.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-repack.txt            |  11 ++\n Documentation/technical/cruft-packs.txt |   2 +-\n builtin/repack.c                        | 106 +++++++++++-\n t/t5328-pack-objects-cruft.sh           | 207 ++++++++++++++++++++++++\n 4 files changed, 320 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/git-repack.txt b/Documentation/git-repack.txt\nindex ee30edc178..0bf13893d8 100644\n--- a/Documentation/git-repack.txt\n+++ b/Documentation/git-repack.txt\n@@ -63,6 +63,17 @@ to the new separate pack will be written.\n \tAlso run  'git prune-packed' to remove redundant\n \tloose object files.\n \n+--cruft::\n+\tSame as `-a`, unless `-d` is used. Then any unreachable objects\n+\tare packed into a separate cruft pack. Unreachable objects can\n+\tbe pruned using the normal expiry rules with the next `git gc`\n+\tinvocation (see linkgit:git-gc[1]). Incompatible with `-k`.\n+\n+--cruft-expiration=<approxidate>::\n+\tExpire unreachable objects older than `<approxidate>`\n+\timmediately instead of waiting for the next `git gc` invocation.\n+\tOnly useful with `--cruft -d`.\n+\n -l::\n \tPass the `--local` option to 'git pack-objects'. See\n \tlinkgit:git-pack-objects[1].\ndiff --git a/Documentation/technical/cruft-packs.txt b/Documentation/technical/cruft-packs.txt\nindex 2c3c5d93f8..f80e975a47 100644\n--- a/Documentation/technical/cruft-packs.txt\n+++ b/Documentation/technical/cruft-packs.txt\n@@ -17,7 +17,7 @@ pruned according to normal expiry rules with the next 'git gc' invocation.\n \n Unreachable objects aren't removed immediately, since doing so could race with\n an incoming push which may reference an object which is about to be deleted.\n-Instead, those unreachable objects are stored as loose object and stay that way\n+Instead, those unreachable objects are stored as loose objects and stay that way\n until they are older than the expiration window, at which point they are removed\n by linkgit:git-prune[1].\n \ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex f908f7d5dd..f7fb88bcf1 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -18,11 +18,17 @@\n #include \"pack-bitmap.h\"\n #include \"refs.h\"\n \n+#define ALL_INTO_ONE 1\n+#define LOOSEN_UNREACHABLE 2\n+#define PACK_CRUFT 4\n+\n+static int pack_everything;\n static int delta_base_offset = 1;\n static int pack_kept_objects = -1;\n static int write_bitmaps = -1;\n static int use_delta_islands;\n static char *packdir, *packtmp_name, *packtmp;\n+static char *cruft_expiration;\n \n static const char *const git_repack_usage[] = {\n \tN_(\"git repack [<options>]\"),\n@@ -54,6 +60,7 @@ static int repack_config(const char *var, const char *value, void *cb)\n \t\tuse_delta_islands = git_config_bool(var, value);\n \t\treturn 0;\n \t}\n+\n \treturn git_default_config(var, value, cb);\n }\n \n@@ -300,9 +307,6 @@ static void repack_promisor_objects(const struct pack_objects_args *args,\n \t\tdie(_(\"could not finish pack-objects to repack promisor objects\"));\n }\n \n-#define ALL_INTO_ONE 1\n-#define LOOSEN_UNREACHABLE 2\n-\n struct pack_geometry {\n \tstruct packed_git **pack;\n \tuint32_t pack_nr, pack_alloc;\n@@ -339,6 +343,8 @@ static void init_pack_geometry(struct pack_geometry **geometry_p)\n \tfor (p = get_all_packs(the_repository); p; p = p->next) {\n \t\tif (!pack_kept_objects && p->pack_keep)\n \t\t\tcontinue;\n+\t\tif (p->is_cruft)\n+\t\t\tcontinue;\n \n \t\tALLOC_GROW(geometry->pack,\n \t\t\t   geometry->pack_nr + 1,\n@@ -600,6 +606,67 @@ static int write_midx_included_packs(struct string_list *include,\n \treturn finish_command(&cmd);\n }\n \n+static int write_cruft_pack(const struct pack_objects_args *args,\n+\t\t\t    const char *pack_prefix,\n+\t\t\t    struct string_list *names,\n+\t\t\t    struct string_list *existing_packs,\n+\t\t\t    struct string_list *existing_kept_packs)\n+{\n+\tstruct child_process cmd = CHILD_PROCESS_INIT;\n+\tstruct strbuf line = STRBUF_INIT;\n+\tstruct string_list_item *item;\n+\tFILE *in, *out;\n+\tint ret;\n+\n+\tprepare_pack_objects(&cmd, args);\n+\n+\tstrvec_push(&cmd.args, \"--cruft\");\n+\tif (cruft_expiration)\n+\t\tstrvec_pushf(&cmd.args, \"--cruft-expiration=%s\",\n+\t\t\t     cruft_expiration);\n+\n+\tstrvec_push(&cmd.args, \"--honor-pack-keep\");\n+\tstrvec_push(&cmd.args, \"--non-empty\");\n+\tstrvec_push(&cmd.args, \"--max-pack-size=0\");\n+\n+\tcmd.in = -1;\n+\n+\tret = start_command(&cmd);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\t/*\n+\t * names has a confusing double use: it both provides the list\n+\t * of just-written new packs, and accepts the name of the cruft\n+\t * pack we are writing.\n+\t *\n+\t * By the time it is read here, it contains only the pack(s)\n+\t * that were just written, which is exactly the set of packs we\n+\t * want to consider kept.\n+\t */\n+\tin = xfdopen(cmd.in, \"w\");\n+\tfor_each_string_list_item(item, names)\n+\t\tfprintf(in, \"%s-%s.pack\\n\", pack_prefix, item->string);\n+\tfor_each_string_list_item(item, existing_packs)\n+\t\tfprintf(in, \"-%s.pack\\n\", item->string);\n+\tfor_each_string_list_item(item, existing_kept_packs)\n+\t\tfprintf(in, \"%s.pack\\n\", item->string);\n+\tfclose(in);\n+\n+\tout = xfdopen(cmd.out, \"r\");\n+\twhile (strbuf_getline_lf(&line, out) != EOF) {\n+\t\tif (line.len != the_hash_algo->hexsz)\n+\t\t\tdie(_(\"repack: Expecting full hex object ID lines only \"\n+\t\t\t      \"from pack-objects.\"));\n+\t\tstring_list_append(names, line.buf);\n+\t}\n+\tfclose(out);\n+\n+\tstrbuf_release(&line);\n+\n+\treturn finish_command(&cmd);\n+}\n+\n int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n \tstruct child_process cmd = CHILD_PROCESS_INIT;\n@@ -616,7 +683,6 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tint show_progress;\n \n \t/* variables to be filled by option parsing */\n-\tint pack_everything = 0;\n \tint delete_redundant = 0;\n \tconst char *unpack_unreachable = NULL;\n \tint keep_unreachable = 0;\n@@ -632,6 +698,11 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tOPT_BIT('A', NULL, &pack_everything,\n \t\t\t\tN_(\"same as -a, and turn unreachable objects loose\"),\n \t\t\t\t   LOOSEN_UNREACHABLE | ALL_INTO_ONE),\n+\t\tOPT_BIT(0, \"cruft\", &pack_everything,\n+\t\t\t\tN_(\"same as -a, pack unreachable cruft objects separately\"),\n+\t\t\t\t   PACK_CRUFT),\n+\t\tOPT_STRING(0, \"cruft-expiration\", &cruft_expiration, N_(\"approxidate\"),\n+\t\t\t\tN_(\"with -C, expire objects older than this\")),\n \t\tOPT_BOOL('d', NULL, &delete_redundant,\n \t\t\t\tN_(\"remove redundant packs, and run git-prune-packed\")),\n \t\tOPT_BOOL('f', NULL, &po_args.no_reuse_delta,\n@@ -684,6 +755,15 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t    (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE)))\n \t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--keep-unreachable\", \"-A\");\n \n+\tif (pack_everything & PACK_CRUFT) {\n+\t\tpack_everything |= ALL_INTO_ONE;\n+\n+\t\tif (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE))\n+\t\t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--cruft\", \"-A\");\n+\t\tif (keep_unreachable)\n+\t\t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--cruft\", \"-k\");\n+\t}\n+\n \tif (write_bitmaps < 0) {\n \t\tif (!write_midx &&\n \t\t    (!(pack_everything & ALL_INTO_ONE) || !is_bare_repository()))\n@@ -767,7 +847,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (pack_everything & ALL_INTO_ONE) {\n \t\trepack_promisor_objects(&po_args, &names);\n \n-\t\tif (existing_nonkept_packs.nr && delete_redundant) {\n+\t\tif (existing_nonkept_packs.nr && delete_redundant &&\n+\t\t    !(pack_everything & PACK_CRUFT)) {\n \t\t\tfor_each_string_list_item(item, &names) {\n \t\t\t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s-%s.pack\",\n \t\t\t\t\t     packtmp_name, item->string);\n@@ -829,6 +910,21 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (!names.nr && !po_args.quiet)\n \t\tprintf_ln(_(\"Nothing new to pack.\"));\n \n+\tif (pack_everything & PACK_CRUFT) {\n+\t\tconst char *pack_prefix;\n+\t\tif (!skip_prefix(packtmp, packdir, &pack_prefix))\n+\t\t\tdie(_(\"pack prefix %s does not begin with objdir %s\"),\n+\t\t\t    packtmp, packdir);\n+\t\tif (*pack_prefix == '/')\n+\t\t\tpack_prefix++;\n+\n+\t\tret = write_cruft_pack(&po_args, pack_prefix, &names,\n+\t\t\t\t       &existing_nonkept_packs,\n+\t\t\t\t       &existing_kept_packs);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t}\n+\n \tfor_each_string_list_item(item, &names) {\n \t\titem->util = (void *)(uintptr_t)populate_pack_exts(item->string);\n \t}\ndiff --git a/t/t5328-pack-objects-cruft.sh b/t/t5328-pack-objects-cruft.sh\nindex 939cdc297a..06c550c958 100755\n--- a/t/t5328-pack-objects-cruft.sh\n+++ b/t/t5328-pack-objects-cruft.sh\n@@ -358,4 +358,211 @@ test_expect_success 'expired objects are pruned' '\n \t)\n '\n \n+test_expect_success 'repack --cruft generates a cruft pack' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit branch -M main &&\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit unreachable &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d unreachable &&\n+\t\t# objects are not cruft if they are contained in the reflogs\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\t\tsort <reachable.raw >reachable &&\n+\t\tcomm -13 reachable objects >unreachable &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tcruft=$(basename $(ls $packdir/pack-*.mtimes) .mtimes) &&\n+\t\tpack=$(basename $(ls $packdir/pack-*.pack | grep -v $cruft) .pack) &&\n+\n+\t\tgit show-index <$packdir/$pack.idx >actual.raw &&\n+\t\tcut -f2 -d\" \" actual.raw | sort >actual &&\n+\t\ttest_cmp reachable actual &&\n+\n+\t\tgit show-index <$packdir/$cruft.idx >actual.raw &&\n+\t\tcut -f2 -d\" \" actual.raw | sort >actual &&\n+\t\ttest_cmp unreachable actual\n+\t)\n+'\n+\n+test_expect_success 'loose objects mtimes upsert others' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit cruft &&\n+\t\t# incremental repack, leaving existing objects loose (so\n+\t\t# they can be \"freshened\")\n+\t\tgit repack &&\n+\n+\t\ttip=\"$(git rev-parse cruft)\" &&\n+\t\tpath=\"$objdir/$(test_oid_to_path \"$(git rev-parse cruft)\")\" &&\n+\t\ttest-tool chmtime --get +1000 \"$path\" >expect &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tmtimes=\"$(basename $(ls $packdir/pack-*.mtimes))\" &&\n+\t\ttest-tool pack-mtimes \"$mtimes\" >actual.raw &&\n+\t\tgrep \"$tip\" actual.raw | cut -d\" \" -f2 >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'cruft packs are not included in geometric repack' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit cruft &&\n+\t\tgit repack -d &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft &&\n+\n+\t\tfind $packdir -type f | sort >before &&\n+\t\tgit repack --geometric=2 -d &&\n+\t\tfind $packdir -type f | sort >after &&\n+\n+\t\ttest_cmp before after\n+\t)\n+'\n+\n+test_expect_success 'repack --geometric collects once-cruft objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\tgit rm -rf . &&\n+\t\ttest_commit --no-tag cruft &&\n+\t\tcruft=\"$(git rev-parse HEAD)\" &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t# Pack the objects created in the previous step into a cruft\n+\t\t# pack. Intentionally leave loose copies of those objects\n+\t\t# around so we can pick them up in a subsequent --geometric\n+\t\t# reapack.\n+\t\tgit repack --cruft &&\n+\n+\t\t# Now make those objects reachable, and ensure that they are\n+\t\t# packed into the new pack created via a --geometric repack.\n+\t\tgit update-ref refs/heads/other $cruft &&\n+\n+\t\t# Without this object, the set of unpacked objects is exactly\n+\t\t# the set of objects already in the cruft pack. Tweak that set\n+\t\t# to ensure we do not overwrite the cruft pack entirely.\n+\t\ttest_commit reachable2 &&\n+\n+\t\tfind $packdir -name \"pack-*.idx\" | sort >before &&\n+\t\tgit repack --geometric=2 -d &&\n+\t\tfind $packdir -name \"pack-*.idx\" | sort >after &&\n+\n+\t\t{\n+\t\t\tgit rev-list --objects --no-object-names $cruft &&\n+\t\t\tgit rev-list --objects --no-object-names reachable..reachable2\n+\t\t} >want.raw &&\n+\t\tsort want.raw >want &&\n+\n+\t\tpack=$(comm -13 before after) &&\n+\t\tgit show-index <$pack >objects.raw &&\n+\n+\t\tcut -d\" \" -f2 objects.raw | sort >got &&\n+\n+\t\ttest_cmp want got\n+\t)\n+'\n+\n+test_expect_success 'cruft repack with no reachable objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\t\tgit repack -ad &&\n+\n+\t\tbase=\"$(git rev-parse base)\" &&\n+\n+\t\tgit for-each-ref --format=\"delete %(refname)\" >in &&\n+\t\tgit update-ref --stdin <in &&\n+\t\tgit reflog expire --all --expire=all &&\n+\t\trm -fr .git/index &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tgit cat-file -t $base\n+\t)\n+'\n+\n+test_expect_success 'cruft repack ignores --max-pack-size' '\n+\tgit init max-pack-size &&\n+\t(\n+\t\tcd max-pack-size &&\n+\t\ttest_commit base &&\n+\t\t# two cruft objects which exceed the maximum pack size\n+\t\ttest-tool genrandom foo 1048576 | git hash-object --stdin -w &&\n+\t\ttest-tool genrandom bar 1048576 | git hash-object --stdin -w &&\n+\t\tgit repack --cruft --max-pack-size=1M &&\n+\t\tfind $packdir -name \"*.mtimes\" >cruft &&\n+\t\ttest_line_count = 1 cruft &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(cat cruft)\")\" >objects &&\n+\t\ttest_line_count = 2 objects\n+\t)\n+'\n+\n+test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n+\t(\n+\t\tcd max-pack-size &&\n+\t\t# repack everything back together to remove the existing cruft\n+\t\t# pack (but to keep its objects)\n+\t\tgit repack -adk &&\n+\t\tgit -c pack.packSizeLimit=1M repack --cruft &&\n+\t\t# ensure the same post condition is met when --max-pack-size\n+\t\t# would otherwise be inferred from the configuration\n+\t\tfind $packdir -name \"*.mtimes\" >cruft &&\n+\t\ttest_line_count = 1 cruft &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(cat cruft)\")\" >objects &&\n+\t\ttest_line_count = 2 objects\n+\t)\n+'\n+\n test_done\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450043","messageId":"e6eee7f15c25a19e6b6c78aa1742df6d7b3d4faa.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 14/17] builtin/repack.c: use named flags for existing_packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:31Z","receivedAt":"2022-03-02T00:58:45Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"We use the `util` pointer for items in the `existing_packs` string list\nto indicate which packs are going to be deleted. Since that has so far\nbeen the only use of that `util` pointer, we just set it to 0 or 1.\n\nBut we're going to add an additional state to this field in the next\npatch, so prepare for that by adding a #define for the first bit so we\ncan more expressively inspect the flags state.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/repack.c | 9 ++++++---\n 1 file changed, 6 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex d61c78e94e..afa4d51a22 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -22,6 +22,8 @@\n #define LOOSEN_UNREACHABLE 2\n #define PACK_CRUFT 4\n \n+#define DELETE_PACK 1\n+\n static int pack_everything;\n static int delta_base_offset = 1;\n static int pack_kept_objects = -1;\n@@ -561,7 +563,7 @@ static void midx_included_packs(struct string_list *include,\n \t\t}\n \t} else {\n \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n-\t\t\tif (item->util)\n+\t\t\tif ((uintptr_t)item->util & DELETE_PACK)\n \t\t\t\tcontinue;\n \t\t\tstring_list_insert(include, xstrfmt(\"%s.idx\", item->string));\n \t\t}\n@@ -1000,7 +1002,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t * was given) and that we will actually delete this pack\n \t\t\t * (if `-d` was given).\n \t\t\t */\n-\t\t\titem->util = (void*)(intptr_t)!string_list_has_string(&names, sha1);\n+\t\t\tif (!string_list_has_string(&names, sha1))\n+\t\t\t\titem->util = (void*)(uintptr_t)((size_t)item->util | DELETE_PACK);\n \t\t}\n \t}\n \n@@ -1024,7 +1027,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (delete_redundant) {\n \t\tint opts = 0;\n \t\tfor_each_string_list_item(item, &existing_nonkept_packs) {\n-\t\t\tif (!item->util)\n+\t\t\tif (!((uintptr_t)item->util & DELETE_PACK))\n \t\t\t\tcontinue;\n \t\t\tremove_redundant_pack(packdir, item->string);\n \t\t}\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450045","messageId":"b09dbc9fe5f02b07f4b20503c4d8f427c6edb6fa.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 15/17] builtin/repack.c: add cruft packs to MIDX during geometric repack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:34Z","receivedAt":"2022-03-02T00:59:04Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When using cruft packs, the following race can occur when a geometric\nrepack that writes a MIDX bitmap takes place afterwords:\n\n  - First, create an unreachable object and do an all-into-one cruft\n    repack which stores that object in the repository's cruft pack.\n  - Then make that object reachable.\n  - Finally, do a geometric repack and write a MIDX bitmap.\n\nAssuming that we are sufficiently unlucky as to select a commit from the\nMIDX which reaches that object for bitmapping, then the `git\nmulti-pack-index` process will complain that that object is missing.\n\nThe reason is because we don't include cruft packs in the MIDX when\ndoing a geometric repack. Since the \"make that object reachable\" doesn't\nnecessarily mean that we'll create a new copy of that object in one of\nthe packs that will get rolled up as part of a geometric repack, it's\npossible that the MIDX won't see any copies of that now-reachable\nobject.\n\nOf course, it's desirable to avoid including cruft packs in the MIDX\nbecause it causes the MIDX to store a bunch of objects which are likely\nto get thrown away. But excluding that pack does open us up to the above\nrace.\n\nThis patch demonstrates the bug, and resolves it by including cruft\npacks in the MIDX even when doing a geometric repack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/repack.c              | 19 +++++++++++++++++--\n t/t5328-pack-objects-cruft.sh | 26 ++++++++++++++++++++++++++\n 2 files changed, 43 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex afa4d51a22..59b60cd309 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -23,6 +23,7 @@\n #define PACK_CRUFT 4\n \n #define DELETE_PACK 1\n+#define CRUFT_PACK 2\n \n static int pack_everything;\n static int delta_base_offset = 1;\n@@ -158,8 +159,11 @@ static void collect_pack_filenames(struct string_list *fname_nonkept_list,\n \t\tif ((extra_keep->nr > 0 && i < extra_keep->nr) ||\n \t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname))))\n \t\t\tstring_list_append_nodup(fname_kept_list, fname);\n-\t\telse\n-\t\t\tstring_list_append_nodup(fname_nonkept_list, fname);\n+\t\telse {\n+\t\t\tstruct string_list_item *item = string_list_append_nodup(fname_nonkept_list, fname);\n+\t\t\tif (file_exists(mkpath(\"%s/%s.mtimes\", packdir, fname)))\n+\t\t\t\titem->util = (void*)(uintptr_t)CRUFT_PACK;\n+\t\t}\n \t}\n \tclosedir(dir);\n }\n@@ -561,6 +565,17 @@ static void midx_included_packs(struct string_list *include,\n \n \t\t\tstring_list_insert(include, strbuf_detach(&buf, NULL));\n \t\t}\n+\n+\t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n+\t\t\tif (!((uintptr_t)item->util & CRUFT_PACK)) {\n+\t\t\t\t/*\n+\t\t\t\t * no need to check DELETE_PACK, since we're not\n+\t\t\t\t * doing an ALL_INTO_ONE repack\n+\t\t\t\t */\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tstring_list_insert(include, xstrfmt(\"%s.idx\", item->string));\n+\t\t}\n \t} else {\n \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n \t\t\tif ((uintptr_t)item->util & DELETE_PACK)\ndiff --git a/t/t5328-pack-objects-cruft.sh b/t/t5328-pack-objects-cruft.sh\nindex e4744e4465..13158e4ab7 100755\n--- a/t/t5328-pack-objects-cruft.sh\n+++ b/t/t5328-pack-objects-cruft.sh\n@@ -648,4 +648,30 @@ test_expect_success 'cruft --local drops unreachable objects' '\n \t)\n '\n \n+test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\ttest_commit cruft &&\n+\t\tunreachable=\"$(git rev-parse cruft)\" &&\n+\n+\t\tgit reset --hard $unreachable^ &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\t# resurrect the unreachable object via a new commit. the\n+\t\t# new commit will get selected for a bitmap, but be\n+\t\t# missing one of its parents from the selected packs.\n+\t\tgit reset --hard $unreachable &&\n+\t\ttest_commit resurrect &&\n+\n+\t\tgit repack --write-midx --write-bitmap-index --geometric=2 -d\n+\t)\n+'\n+\n test_done\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450046","messageId":"7a21ae1494eb59ab291b1c9cbdc2dcff93c4df9b.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 16/17] builtin/gc.c: conditionally avoid pruning objects via loose","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:36Z","receivedAt":"2022-03-02T00:59:05Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Expose the new `git repack --cruft` mode from `git gc` via a new opt-in\nflag. When invoked like `git gc --cruft`, `git gc` will avoid exploding\nunreachable objects as loose ones, and instead create a cruft pack and\n`.mtimes` file.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/gc.txt   | 21 +++++++++++++-------\n Documentation/git-gc.txt      |  5 +++++\n builtin/gc.c                  | 10 +++++++++-\n t/t5328-pack-objects-cruft.sh | 37 +++++++++++++++++++++++++++++++++++\n 4 files changed, 65 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/config/gc.txt b/Documentation/config/gc.txt\nindex c834e07991..38fea076a2 100644\n--- a/Documentation/config/gc.txt\n+++ b/Documentation/config/gc.txt\n@@ -81,14 +81,21 @@ gc.packRefs::\n \tto enable it within all non-bare repos or it can be set to a\n \tboolean value.  The default is `true`.\n \n+gc.cruftPacks::\n+\tStore unreachable objects in a cruft pack (see\n+\tlinkgit:git-repack[1]) instead of as loose objects. The default\n+\tis `false`.\n+\n gc.pruneExpire::\n-\tWhen 'git gc' is run, it will call 'prune --expire 2.weeks.ago'.\n-\tOverride the grace period with this config variable.  The value\n-\t\"now\" may be used to disable this grace period and always prune\n-\tunreachable objects immediately, or \"never\" may be used to\n-\tsuppress pruning.  This feature helps prevent corruption when\n-\t'git gc' runs concurrently with another process writing to the\n-\trepository; see the \"NOTES\" section of linkgit:git-gc[1].\n+\tWhen 'git gc' is run, it will call 'prune --expire 2.weeks.ago'\n+\t(and 'repack --cruft --cruft-expiration 2.weeks.ago' if using\n+\tcruft packs via `gc.cruftPacks` or `--cruft`).  Override the\n+\tgrace period with this config variable.  The value \"now\" may be\n+\tused to disable this grace period and always prune unreachable\n+\tobjects immediately, or \"never\" may be used to suppress pruning.\n+\tThis feature helps prevent corruption when 'git gc' runs\n+\tconcurrently with another process writing to the repository; see\n+\tthe \"NOTES\" section of linkgit:git-gc[1].\n \n gc.worktreePruneExpire::\n \tWhen 'git gc' is run, it calls\ndiff --git a/Documentation/git-gc.txt b/Documentation/git-gc.txt\nindex 853967dea0..ba4e67700e 100644\n--- a/Documentation/git-gc.txt\n+++ b/Documentation/git-gc.txt\n@@ -54,6 +54,11 @@ other housekeeping tasks (e.g. rerere, working trees, reflog...) will\n be performed as well.\n \n \n+--cruft::\n+\tWhen expiring unreachable objects, pack them separately into a\n+\tcruft pack instead of storing the loose objects as loose\n+\tobjects.\n+\n --prune=<date>::\n \tPrune loose objects older than date (default is 2 weeks ago,\n \toverridable by the config variable `gc.pruneExpire`).\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex ffaf0daf5d..11f5150234 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -43,6 +43,7 @@ static const char * const builtin_gc_usage[] = {\n \n static int pack_refs = 1;\n static int prune_reflogs = 1;\n+static int cruft_packs = 0;\n static int aggressive_depth = 50;\n static int aggressive_window = 250;\n static int gc_auto_threshold = 6700;\n@@ -153,6 +154,7 @@ static void gc_config(void)\n \tgit_config_get_int(\"gc.auto\", &gc_auto_threshold);\n \tgit_config_get_int(\"gc.autopacklimit\", &gc_auto_pack_limit);\n \tgit_config_get_bool(\"gc.autodetach\", &detach_auto);\n+\tgit_config_get_bool(\"gc.cruftpacks\", &cruft_packs);\n \tgit_config_get_expiry(\"gc.pruneexpire\", &prune_expire);\n \tgit_config_get_expiry(\"gc.worktreepruneexpire\", &prune_worktrees_expire);\n \tgit_config_get_expiry(\"gc.logexpiry\", &gc_log_expire);\n@@ -332,7 +334,11 @@ static void add_repack_all_option(struct string_list *keep_pack)\n {\n \tif (prune_expire && !strcmp(prune_expire, \"now\"))\n \t\tstrvec_push(&repack, \"-a\");\n-\telse {\n+\telse if (cruft_packs) {\n+\t\tstrvec_push(&repack, \"--cruft\");\n+\t\tif (prune_expire)\n+\t\t\tstrvec_pushf(&repack, \"--cruft-expiration=%s\", prune_expire);\n+\t} else {\n \t\tstrvec_push(&repack, \"-A\");\n \t\tif (prune_expire)\n \t\t\tstrvec_pushf(&repack, \"--unpack-unreachable=%s\", prune_expire);\n@@ -552,6 +558,7 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t{ OPTION_STRING, 0, \"prune\", &prune_expire, N_(\"date\"),\n \t\t\tN_(\"prune unreferenced objects\"),\n \t\t\tPARSE_OPT_OPTARG, NULL, (intptr_t)prune_expire },\n+\t\tOPT_BOOL(0, \"cruft\", &cruft_packs, N_(\"pack unreferenced objects separately\")),\n \t\tOPT_BOOL(0, \"aggressive\", &aggressive, N_(\"be more thorough (increased runtime)\")),\n \t\tOPT_BOOL_F(0, \"auto\", &auto_gc, N_(\"enable auto-gc mode\"),\n \t\t\t   PARSE_OPT_NOCOMPLETE),\n@@ -671,6 +678,7 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t\tdie(FAILED_RUN, repack.v[0]);\n \n \t\tif (prune_expire) {\n+\t\t\t/* run `git prune` even if using cruft packs */\n \t\t\tstrvec_push(&prune, prune_expire);\n \t\t\tif (quiet)\n \t\t\t\tstrvec_push(&prune, \"--no-progress\");\ndiff --git a/t/t5328-pack-objects-cruft.sh b/t/t5328-pack-objects-cruft.sh\nindex 13158e4ab7..3910e186ef 100755\n--- a/t/t5328-pack-objects-cruft.sh\n+++ b/t/t5328-pack-objects-cruft.sh\n@@ -429,6 +429,43 @@ test_expect_success 'loose objects mtimes upsert others' '\n \t)\n '\n \n+test_expect_success 'expiring cruft objects with git gc' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit branch -M main &&\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit unreachable &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d unreachable &&\n+\t\t# objects are not cruft if they are contained in the reflogs\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\t\tsort <reachable.raw >reachable &&\n+\t\tcomm -13 reachable objects >unreachable &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tmtimes=$(ls .git/objects/pack/pack-*.mtimes) &&\n+\t\ttest_path_is_file $mtimes &&\n+\n+\t\tgit gc --cruft --prune=now &&\n+\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\n+\t\tcomm -23 unreachable objects >removed &&\n+\t\ttest_cmp unreachable removed &&\n+\t\ttest_path_is_missing $mtimes\n+\t)\n+'\n+\n test_expect_success 'cruft packs are not included in geometric repack' '\n \tgit init repo &&\n \ttest_when_finished \"rm -fr repo\" &&\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450047","messageId":"b729b8096313e44d988db735218d4bc98ce5b6fb.1646182671.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"[PATCH v2 17/17] sha1-file.c: don't freshen cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T00:58:38Z","receivedAt":"2022-03-02T00:59:06Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"We don't bother to freshen objects stored in a cruft pack individually\nby updating the `.mtimes` file. This is because we can't portably `mmap`\nand write into the middle of a file (i.e., to update the mtime of just\none object). Instead, we would have to rewrite the entire `.mtimes` file\nwhich may incur some wasted effort especially if there a lot of cruft\nobjects and they are freshened infrequently.\n\nInstead, force the freshening code to avoid an optimizing write by\nwriting out the object loose and letting it pick up a current mtime.\n\nThis works because we prefer the mtime of the loose copy of an object\nwhen both a loose and packed one exist (whether or not the packed copy\ncomes from a cruft pack or not).\n\nThis could certainly do with a test and/or be included earlier in this\nseries/PR, but I want to wait until after I have a chance to clean up\nthe overly-repetitive nature of the cruft pack tests in general.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n object-file.c                 |  2 ++\n t/t5328-pack-objects-cruft.sh | 25 +++++++++++++++++++++++++\n 2 files changed, 27 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex e80da1368d..65b8df7fb6 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1989,6 +1989,8 @@ static int freshen_packed_object(const struct object_id *oid)\n \tstruct pack_entry e;\n \tif (!find_pack_entry(the_repository, oid, &e))\n \t\treturn 0;\n+\tif (e.p->is_cruft)\n+\t\treturn 0;\n \tif (e.p->freshened)\n \t\treturn 1;\n \tif (!freshen_file(e.p->pack_name))\ndiff --git a/t/t5328-pack-objects-cruft.sh b/t/t5328-pack-objects-cruft.sh\nindex 3910e186ef..4681558612 100755\n--- a/t/t5328-pack-objects-cruft.sh\n+++ b/t/t5328-pack-objects-cruft.sh\n@@ -711,4 +711,29 @@ test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n \t)\n '\n \n+test_expect_success 'cruft objects are freshend via loose' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\techo \"cruft\" >contents &&\n+\t\tblob=\"$(git hash-object -w -t blob contents)\" &&\n+\t\tloose=\"$objdir/$(test_oid_to_path $blob)\" &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\ttest_path_is_missing \"$loose\" &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(ls $packdir/pack-*.mtimes)\")\" >cruft &&\n+\t\tgrep \"$blob\" cruft &&\n+\n+\t\t# write the same object again\n+\t\tgit hash-object -w -t blob contents &&\n+\n+\t\ttest_path_is_file \"$loose\"\n+\t)\n+'\n+\n test_done\n-- \n2.35.1.73.gccc5557600\n"},{"id":"450064","messageId":"xmqqwnhcn6ke.fsf@gitster.g","threadId":"56996","inReplyTo":"d68ce281324097e10e4c1921d84c577bed6943e7.1646182671.git.me@ttaylorr.com","subject":"Re: [PATCH v2 11/17] builtin/pack-objects.c: --cruft with expiration","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-02T07:42:57Z","receivedAt":"2022-03-02T07:43:04Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n>  builtin/pack-objects.c        |  84 +++++++++++++++++++-\n>  t/t5328-pack-objects-cruft.sh | 143 ++++++++++++++++++++++++++++++++++\n>  2 files changed, 226 insertions(+), 1 deletion(-)\n\nI'd renumber this to 5329, as the latest iteration of generation\nnumber v2 series took 5328, while queuing.\n\n\n"},{"id":"450111","messageId":"Yh+TPppXFoBU2zbN@nand.local","threadId":"56996","inReplyTo":"xmqqwnhcn6ke.fsf@gitster.g","subject":"Re: [PATCH v2 11/17] builtin/pack-objects.c: --cruft with expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T15:54:38Z","receivedAt":"2022-03-02T15:54:43Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Mar 01, 2022 at 11:42:57PM -0800, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> >  builtin/pack-objects.c        |  84 +++++++++++++++++++-\n> >  t/t5328-pack-objects-cruft.sh | 143 ++++++++++++++++++++++++++++++++++\n> >  2 files changed, 226 insertions(+), 1 deletion(-)\n>\n> I'd renumber this to 5329, as the latest iteration of generation\n> number v2 series took 5328, while queuing.\n\nOops. I had scanned that series, but glossed over the new test number.\n\nThanks for renaming (I'll do the same, in case we end up accumulating\nmore reroll-able bits).\n\nThanks,\nTaylor\n"},{"id":"450161","messageId":"5e4d195b-5a4f-48d5-10a6-631501ca466e@github.com","threadId":"56996","inReplyTo":"Yh+TPppXFoBU2zbN@nand.local","subject":"Re: [PATCH v2 11/17] builtin/pack-objects.c: --cruft with expiration","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-03-02T19:57:35Z","receivedAt":"2022-03-02T19:57:44Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 3/2/2022 10:54 AM, Taylor Blau wrote:\n> On Tue, Mar 01, 2022 at 11:42:57PM -0800, Junio C Hamano wrote:\n>> Taylor Blau <me@ttaylorr.com> writes:\n>>\n>>>  builtin/pack-objects.c        |  84 +++++++++++++++++++-\n>>>  t/t5328-pack-objects-cruft.sh | 143 ++++++++++++++++++++++++++++++++++\n>>>  2 files changed, 226 insertions(+), 1 deletion(-)\n>>\n>> I'd renumber this to 5329, as the latest iteration of generation\n>> number v2 series took 5328, while queuing.\n> \n> Oops. I had scanned that series, but glossed over the new test number.\n> \n> Thanks for renaming (I'll do the same, in case we end up accumulating\n> more reroll-able bits).\n\nSorry for the collision! Had I realized this was already used here,\nI would have changed the number myself.\n\nThanks,\n-Stolee\n"},{"id":"450163","messageId":"245e34c3-ba85-2bb3-d17a-e48ee5b142bd@github.com","threadId":"56996","inReplyTo":"6f0e84273f78797c728058521969e73f8817b49c.1646182671.git.me@ttaylorr.com","subject":"Re: [PATCH v2 09/17] reachable: add options to add_unseen_recent_objects_to_traversal","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-03-02T20:19:57Z","receivedAt":"2022-03-02T20:20:02Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 3/1/2022 7:58 PM, Taylor Blau wrote:\n\n> diff --git a/reachable.h b/reachable.h\n> index 5df932ad8f..b776761baa 100644\n> --- a/reachable.h\n> +++ b/reachable.h\n> @@ -1,11 +1,18 @@\n>  #ifndef REACHEABLE_H\n>  #define REACHEABLE_H\n>  \n> +#include \"object.h\"\n> +\n\nNit: just realized this include could be replaced by a struct\ndeclaration:\n\n>  struct progress;\n>  struct rev_info;\n\nLike these. 'struct object;' should be enough for the typedef.\n>  \n> +typedef void report_recent_object_fn(const struct object *, struct packed_git *,\n> +\t\t\t\t     off_t, time_t);\n> +\n>  int add_unseen_recent_objects_to_traversal(struct rev_info *revs,\n> -\t\t\t\t\t   timestamp_t timestamp);\n> +\t\t\t\t\t   timestamp_t timestamp,\n> +\t\t\t\t\t   report_recent_object_fn cb,\n> +\t\t\t\t\t   int ignore_in_core_kept_packs);\n\nThanks,\n-Stolee\n"},{"id":"450164","messageId":"66eeada0-f2b4-6849-bf14-029bb6c6083d@github.com","threadId":"56996","inReplyTo":"101b34660c0c5028ba591d052dc587bb8918ccb2.1646182671.git.me@ttaylorr.com","subject":"Re: [PATCH v2 02/17] pack-mtimes: support reading .mtimes files","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-03-02T20:22:18Z","receivedAt":"2022-03-02T20:22:22Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 3/1/2022 7:58 PM, Taylor Blau wrote:\n> To store the individual mtimes of objects in a cruft pack, introduce a\n> new `.mtimes` format that can optionally accompany a single pack in the\n> repository.\n> \n> The format is defined in Documentation/technical/pack-format.txt, and\n> stores a 4-byte network order timestamp for each object in name (index)\n> order.\n> \n> This patch prepares for cruft packs by defining the `.mtimes` format,\n> and introducing a basic API that callers can use to read out individual\n> mtimes.\n...\n> +int load_pack_mtimes(struct packed_git *p)\n> +{\n> +\tchar *mtimes_name = NULL;\n> +\tint ret = 0;\n> +\n> +\tif (!p->is_cruft)\n> +\t\treturn ret; /* not a cruft pack */\n> +\tif (p->mtimes_map)\n> +\t\treturn ret; /* already loaded */\n> +\n> +\tret = open_pack_index(p);\n> +\tif (ret < 0)\n> +\t\tgoto cleanup;\n> +\n> +\tmtimes_name = pack_mtimes_filename(p);\n> +\tret = load_pack_mtimes_file(mtimes_name,\n> +\t\t\t\t    p->num_objects,\n> +\t\t\t\t    &p->mtimes_map,\n> +\t\t\t\t    &p->mtimes_size);\n> +\tif (ret)\n> +\t\tgoto cleanup;\n\nThis looked odd to me, so I supposed that you had some code\nthat would be inserted between this 'goto cleanup' and the\n'cleanup:' label, but I did not find such an insertion in\nthe remaining patchs. This 'if' can be deleted.\n\n> +cleanup:\n> +\tfree(mtimes_name);\n> +\treturn ret;\n> +}\n\nThanks,\n-Stolee\n"},{"id":"450165","messageId":"138d98bd-928d-1708-128f-217bfe9a2788@github.com","threadId":"56996","inReplyTo":"cover.1646182671.git.me@ttaylorr.com","subject":"Re: [PATCH v2 00/17] cruft packs","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-03-02T20:23:05Z","receivedAt":"2022-03-02T20:23:12Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 3/1/2022 7:57 PM, Taylor Blau wrote:\n> Here is a reroll of my series to implement \"cruft packs\", a pack which\n> stores accumulated unreachable objects, along with a new \".mtimes\" file\n> which tracks each object's last known modification time.\n> \n> This was on the list towards the end of 2021[1], and I have been\n> accumulating small changes to it locally for a couple of months now.\n> Major changes since last time include:\n> \n>   - Clearer documentation and commit message(s) to better illustrate how\n>     the feature works and is supposed to be used.\n> \n>   - Some minor documentation updates to pack-format.txt, which make some\n>     ambiguous details more explicit.\n> \n>   - Minor code movement / tweaks to make things easier to read, ensure\n>     that functions aren't introduced in patches before they are used /\n>     etc.\n> \n>   - Moved the new test script to t5328 (instead of t5327, which happens\n>     to be taken up by a new MIDX bitmap-related test), and purged it of\n>     all \"rm -fr .git/logs\" (replacing them with \"git reflog --expire\n>     --all --expire=all\" instead).\n> \n>   - A new test which fixes a bug where loose objects which have copies\n>     that appear in a cruft pack would not get accumulated when doing a\n>     `--geometric` repack.\n> \n> For convenience, a range-diff is below. Thanks in advance for taking\n> another look!\n\nIt had been a while since my last read, so I read the patches\nin full one more time. I found a couple nitpicks, but otherwise\neverything is looking good.\n\nThanks,\n-Stolee\n"},{"id":"450177","messageId":"Yh/hctKsg1gzmo3o@nand.local","threadId":"56996","inReplyTo":"245e34c3-ba85-2bb3-d17a-e48ee5b142bd@github.com","subject":"Re: [PATCH v2 09/17] reachable: add options to add_unseen_recent_objects_to_traversal","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T21:28:18Z","receivedAt":"2022-03-02T21:28:23Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Mar 02, 2022 at 03:19:57PM -0500, Derrick Stolee wrote:\n> Nit: just realized this include could be replaced by a struct\n> declaration:\n>\n> >  struct progress;\n> >  struct rev_info;\n>\n> Like these. 'struct object;' should be enough for the typedef.\n\nGood catch. We would need one for the packed_git struct, too. I don't\nhave a strong opinion about including object.h or not, though needing\ntwo stubs pushes me slightly in the direction of leaving the include\nalone.\n\nThanks,\nTaylor\n"},{"id":"450179","messageId":"Yh/isYFFmmOdpNa7@nand.local","threadId":"56996","inReplyTo":"66eeada0-f2b4-6849-bf14-029bb6c6083d@github.com","subject":"Re: [PATCH v2 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T21:33:37Z","receivedAt":"2022-03-02T21:33:42Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Mar 02, 2022 at 03:22:18PM -0500, Derrick Stolee wrote:\n> > +\tret = load_pack_mtimes_file(mtimes_name,\n> > +\t\t\t\t    p->num_objects,\n> > +\t\t\t\t    &p->mtimes_map,\n> > +\t\t\t\t    &p->mtimes_size);\n> > +\tif (ret)\n> > +\t\tgoto cleanup;\n>\n> This looked odd to me, so I supposed that you had some code\n> that would be inserted between this 'goto cleanup' and the\n> 'cleanup:' label, but I did not find such an insertion in\n> the remaining patchs. This 'if' can be deleted.\n\nThanks for spotting. My gut was that there must be something in the\nrange-diff between this and the previous round, but there isn't. So this\ncode has always been there.\n\nIt likely comes from load_pack_revindex_from_disk(), which assigns the\n`revindex_data` member of `struct packed_git` after calling\nload_revindex_from_disk(), but only if it returned zero.\n\nWe don't have to assign mtimes_data here (since it doesn't exist, and)\nbecause all of our reads into mtimes_map are offset by 3 to adjust for\nthe width of the header.\n\nAnyway, we don't need this if statement here, so I'll drop it.\n\nThanks,\nTaylor\n"},{"id":"450180","messageId":"Yh/jVeXfohXJBu6t@nand.local","threadId":"56996","inReplyTo":"138d98bd-928d-1708-128f-217bfe9a2788@github.com","subject":"Re: [PATCH v2 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-02T21:36:21Z","receivedAt":"2022-03-02T21:36:39Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Mar 02, 2022 at 03:23:05PM -0500, Derrick Stolee wrote:\n> > For convenience, a range-diff is below. Thanks in advance for taking\n> > another look!\n>\n> It had been a while since my last read, so I read the patches\n> in full one more time. I found a couple nitpicks, but otherwise\n> everything is looking good.\n\nThanks for reading! I took both of your suggestions (along with Junio's\nto rename the test script to t5329 to avoid a clash with your series)\nand will re-submit a tiny reroll shortly.\n\nThanks,\nTaylor\n"},{"id":"450202","messageId":"cover.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH v3 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:20:41Z","receivedAt":"2022-03-03T00:20:47Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Here is a small reroll of my series to implement \"cruft packs\", based on\nStolee's review.\n\nThe changes here are minor, and mostly are limited to removing a\nredundant \"if\" statement, avoiding an unnecessary header include, and\nmoving the tests (again!) to t5329's territory.\n\nAs always, a range-diff is below. Thanks in advance for taking another\nlook!\n\nTaylor Blau (17):\n  Documentation/technical: add cruft-packs.txt\n  pack-mtimes: support reading .mtimes files\n  pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'\n  chunk-format.h: extract oid_version()\n  pack-mtimes: support writing pack .mtimes files\n  t/helper: add 'pack-mtimes' test-tool\n  builtin/pack-objects.c: return from create_object_entry()\n  builtin/pack-objects.c: --cruft without expiration\n  reachable: add options to add_unseen_recent_objects_to_traversal\n  reachable: report precise timestamps from objects in cruft packs\n  builtin/pack-objects.c: --cruft with expiration\n  builtin/repack.c: support generating a cruft pack\n  builtin/repack.c: allow configuring cruft pack generation\n  builtin/repack.c: use named flags for existing_packs\n  builtin/repack.c: add cruft packs to MIDX during geometric repack\n  builtin/gc.c: conditionally avoid pruning objects via loose\n  sha1-file.c: don't freshen cruft packs\n\n Documentation/Makefile                  |   1 +\n Documentation/config/gc.txt             |  21 +-\n Documentation/config/repack.txt         |   9 +\n Documentation/git-gc.txt                |   5 +\n Documentation/git-pack-objects.txt      |  30 +\n Documentation/git-repack.txt            |  11 +\n Documentation/technical/cruft-packs.txt |  97 ++++\n Documentation/technical/pack-format.txt |  19 +\n Makefile                                |   2 +\n builtin/gc.c                            |  10 +-\n builtin/pack-objects.c                  | 304 +++++++++-\n builtin/repack.c                        | 183 +++++-\n bulk-checkin.c                          |   2 +-\n chunk-format.c                          |  12 +\n chunk-format.h                          |   3 +\n commit-graph.c                          |  18 +-\n midx.c                                  |  18 +-\n object-file.c                           |   4 +-\n object-store.h                          |   7 +-\n pack-mtimes.c                           | 126 ++++\n pack-mtimes.h                           |  15 +\n pack-objects.c                          |   6 +\n pack-objects.h                          |  25 +\n pack-write.c                            |  93 ++-\n pack.h                                  |   4 +\n packfile.c                              |  19 +-\n reachable.c                             |  58 +-\n reachable.h                             |   9 +-\n t/helper/test-pack-mtimes.c             |  56 ++\n t/helper/test-tool.c                    |   1 +\n t/helper/test-tool.h                    |   1 +\n t/t5329-pack-objects-cruft.sh           | 739 ++++++++++++++++++++++++\n 32 files changed, 1807 insertions(+), 101 deletions(-)\n create mode 100644 Documentation/technical/cruft-packs.txt\n create mode 100644 pack-mtimes.c\n create mode 100644 pack-mtimes.h\n create mode 100644 t/helper/test-pack-mtimes.c\n create mode 100755 t/t5329-pack-objects-cruft.sh\n\nRange-diff against v2:\n -:  ---------- >  1:  784ee7e0ee Documentation/technical: add cruft-packs.txt\n 1:  101b34660c !  2:  1ec754ad1b pack-mtimes: support reading .mtimes files\n    @@ pack-mtimes.c (new)\n     +\t\t\t\t    p->num_objects,\n     +\t\t\t\t    &p->mtimes_map,\n     +\t\t\t\t    &p->mtimes_size);\n    -+\tif (ret)\n    -+\t\tgoto cleanup;\n    -+\n     +cleanup:\n     +\tfree(mtimes_name);\n     +\treturn ret;\n 2:  a94d7dfeb3 =  3:  0f5d6d6492 pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'\n 3:  1e0ed363ae =  4:  135a07276b chunk-format.h: extract oid_version()\n 4:  5236490688 =  5:  0600503856 pack-mtimes: support writing pack .mtimes files\n 5:  78313bc441 =  6:  4780c8437b t/helper: add 'pack-mtimes' test-tool\n 6:  142098668d =  7:  33862a07c9 builtin/pack-objects.c: return from create_object_entry()\n 7:  2517a6be3d !  8:  22705e4887 builtin/pack-objects.c: --cruft without expiration\n    @@ object-store.h: int repo_has_object_file_with_flags(struct repository *r,\n      \n      /*\n     \n    - ## t/t5328-pack-objects-cruft.sh (new) ##\n    + ## t/t5329-pack-objects-cruft.sh (new) ##\n     @@\n     +#!/bin/sh\n     +\n 8:  6f0e84273f =  9:  cebb30b667 reachable: add options to add_unseen_recent_objects_to_traversal\n 9:  a8bde361f9 = 10:  fa4de8859d reachable: report precise timestamps from objects in cruft packs\n10:  d68ce28132 ! 11:  92318f8700 builtin/pack-objects.c: --cruft with expiration\n    @@ builtin/pack-objects.c: static void read_cruft_objects(void)\n      \t\tenumerate_cruft_objects();\n      \n     \n    - ## t/t5328-pack-objects-cruft.sh ##\n    -@@ t/t5328-pack-objects-cruft.sh: basic_cruft_pack_tests () {\n    + ## reachable.h ##\n    +@@\n    + #ifndef REACHEABLE_H\n    + #define REACHEABLE_H\n    + \n    +-#include \"object.h\"\n    +-\n    + struct progress;\n    + struct rev_info;\n    ++struct object;\n    ++struct packed_git;\n    + \n    + typedef void report_recent_object_fn(const struct object *, struct packed_git *,\n    + \t\t\t\t     off_t, time_t);\n    +\n    + ## t/t5329-pack-objects-cruft.sh ##\n    +@@ t/t5329-pack-objects-cruft.sh: basic_cruft_pack_tests () {\n      }\n      \n      basic_cruft_pack_tests never\n11:  e5317cd472 ! 12:  1e94b33cb4 builtin/repack.c: support generating a cruft pack\n    @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix\n      \t\titem->util = (void *)(uintptr_t)populate_pack_exts(item->string);\n      \t}\n     \n    - ## t/t5328-pack-objects-cruft.sh ##\n    -@@ t/t5328-pack-objects-cruft.sh: test_expect_success 'expired objects are pruned' '\n    + ## t/t5329-pack-objects-cruft.sh ##\n    +@@ t/t5329-pack-objects-cruft.sh: test_expect_success 'expired objects are pruned' '\n      \t)\n      '\n      \n12:  b548dbbf80 ! 13:  9cfcd123bd builtin/repack.c: allow configuring cruft pack generation\n    @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix\n      \t\t\t\t       &existing_kept_packs);\n      \t\tif (ret)\n     \n    - ## t/t5328-pack-objects-cruft.sh ##\n    -@@ t/t5328-pack-objects-cruft.sh: test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n    + ## t/t5329-pack-objects-cruft.sh ##\n    +@@ t/t5329-pack-objects-cruft.sh: test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n      \t)\n      '\n      \n13:  e6eee7f15c = 14:  1a58807df0 builtin/repack.c: use named flags for existing_packs\n14:  b09dbc9fe5 ! 15:  ed05cf536b builtin/repack.c: add cruft packs to MIDX during geometric repack\n    @@ builtin/repack.c: static void midx_included_packs(struct string_list *include,\n      \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n      \t\t\tif ((uintptr_t)item->util & DELETE_PACK)\n     \n    - ## t/t5328-pack-objects-cruft.sh ##\n    -@@ t/t5328-pack-objects-cruft.sh: test_expect_success 'cruft --local drops unreachable objects' '\n    + ## t/t5329-pack-objects-cruft.sh ##\n    +@@ t/t5329-pack-objects-cruft.sh: test_expect_success 'cruft --local drops unreachable objects' '\n      \t)\n      '\n      \n15:  7a21ae1494 ! 16:  1d5f334138 builtin/gc.c: conditionally avoid pruning objects via loose\n    @@ builtin/gc.c: int cmd_gc(int argc, const char **argv, const char *prefix)\n      \t\t\tif (quiet)\n      \t\t\t\tstrvec_push(&prune, \"--no-progress\");\n     \n    - ## t/t5328-pack-objects-cruft.sh ##\n    -@@ t/t5328-pack-objects-cruft.sh: test_expect_success 'loose objects mtimes upsert others' '\n    + ## t/t5329-pack-objects-cruft.sh ##\n    +@@ t/t5329-pack-objects-cruft.sh: test_expect_success 'loose objects mtimes upsert others' '\n      \t)\n      '\n      \n16:  b729b80963 ! 17:  f74b425872 sha1-file.c: don't freshen cruft packs\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n      \t\treturn 1;\n      \tif (!freshen_file(e.p->pack_name))\n     \n    - ## t/t5328-pack-objects-cruft.sh ##\n    -@@ t/t5328-pack-objects-cruft.sh: test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n    + ## t/t5329-pack-objects-cruft.sh ##\n    +@@ t/t5329-pack-objects-cruft.sh: test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n      \t)\n      '\n      \n-- \n2.35.1.73.gccc5557600\n"},{"id":"450203","messageId":"784ee7e0eec9ba520ebaaa27de2de810e2f6798a.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:20:44Z","receivedAt":"2022-03-03T00:20:53Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Create a technical document to explain cruft packs. It contains a brief\noverview of the problem, some background, details on the implementation,\nand a couple of alternative approaches not considered here.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/Makefile                  |  1 +\n Documentation/technical/cruft-packs.txt | 97 +++++++++++++++++++++++++\n 2 files changed, 98 insertions(+)\n create mode 100644 Documentation/technical/cruft-packs.txt\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex ed656db2ae..0b01c9408e 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -91,6 +91,7 @@ TECH_DOCS += MyFirstContribution\n TECH_DOCS += MyFirstObjectWalk\n TECH_DOCS += SubmittingPatches\n TECH_DOCS += technical/bundle-format\n+TECH_DOCS += technical/cruft-packs\n TECH_DOCS += technical/hash-function-transition\n TECH_DOCS += technical/http-protocol\n TECH_DOCS += technical/index-format\ndiff --git a/Documentation/technical/cruft-packs.txt b/Documentation/technical/cruft-packs.txt\nnew file mode 100644\nindex 0000000000..2c3c5d93f8\n--- /dev/null\n+++ b/Documentation/technical/cruft-packs.txt\n@@ -0,0 +1,97 @@\n+= Cruft packs\n+\n+The cruft packs feature offer an alternative to Git's traditional mechanism of\n+removing unreachable objects. This document provides an overview of Git's\n+pruning mechanism, and how a cruft pack can be used instead to accomplish the\n+same.\n+\n+== Background\n+\n+To remove unreachable objects from your repository, Git offers `git repack -Ad`\n+(see linkgit:git-repack[1]). Quoting from the documentation:\n+\n+[quote]\n+[...] unreachable objects in a previous pack become loose, unpacked objects,\n+instead of being left in the old pack. [...] loose unreachable objects will be\n+pruned according to normal expiry rules with the next 'git gc' invocation.\n+\n+Unreachable objects aren't removed immediately, since doing so could race with\n+an incoming push which may reference an object which is about to be deleted.\n+Instead, those unreachable objects are stored as loose object and stay that way\n+until they are older than the expiration window, at which point they are removed\n+by linkgit:git-prune[1].\n+\n+Git must store these unreachable objects loose in order to keep track of their\n+per-object mtimes. If these unreachable objects were written into one big pack,\n+then either freshening that pack (because an object contained within it was\n+re-written) or creating a new pack of unreachable objects would cause the pack's\n+mtime to get updated, and the objects within it would never leave the expiration\n+window. Instead, objects are stored loose in order to keep track of the\n+individual object mtimes and avoid a situation where all cruft objects are\n+freshened at once.\n+\n+This can lead to undesirable situations when a repository contains many\n+unreachable objects which have not yet left the grace period. Having large\n+directories in the shards of `.git/objects` can lead to decreased performance in\n+the repository. But given enough unreachable objects, this can lead to inode\n+starvation and degrade the performance of the whole system. Since we\n+can never pack those objects, these repositories often take up a large amount of\n+disk space, since we can only zlib compress them, but not store them in delta\n+chains.\n+\n+== Cruft packs\n+\n+A cruft pack eliminates the need for storing unreachable objects in a loose\n+state by including the per-object mtimes in a separate file alongside a single\n+pack containing all loose objects.\n+\n+A cruft pack is written by `git repack --cruft` when generating a new pack.\n+linkgit:git-pack-objects[1]'s `--cruft` option. Note that `git repack --cruft`\n+is a classic all-into-one repack, meaning that everything in the resulting pack is\n+reachable, and everything else is unreachable. Once written, the `--cruft`\n+option instructs `git repack` to generate another pack containing only objects\n+not packed in the previous step (which equates to packing all unreachable\n+objects together). This progresses as follows:\n+\n+  1. Enumerate every object, marking any object which is (a) not contained in a\n+     kept-pack, and (b) whose mtime is within the grace period as a traversal\n+     tip.\n+\n+  2. Perform a reachability traversal based on the tips gathered in the previous\n+     step, adding every object along the way to the pack.\n+\n+  3. Write the pack out, along with a `.mtimes` file that records the per-object\n+     timestamps.\n+\n+This mode is invoked internally by linkgit:git-repack[1] when instructed to\n+write a cruft pack. Crucially, the set of in-core kept packs is exactly the set\n+of packs which will not be deleted by the repack; in other words, they contain\n+all of the repository's reachable objects.\n+\n+When a repository already has a cruft pack, `git repack --cruft` typically only\n+adds objects to it. An exception to this is when `git repack` is given the\n+`--cruft-expiration` option, which allows the generated cruft pack to omit\n+expired objects instead of waiting for linkgit:git-gc[1] to expire those objects\n+later on.\n+\n+It is linkgit:git-gc[1] that is typically responsible for removing expired\n+unreachable objects.\n+\n+== Alternatives\n+\n+Notable alternatives to this design include:\n+\n+  - The location of the per-object mtime data, and\n+  - Storing unreachable objects in multiple cruft packs.\n+\n+On the location of mtime data, a new auxiliary file tied to the pack was chosen\n+to avoid complicating the `.idx` format. If the `.idx` format were ever to gain\n+support for optional chunks of data, it may make sense to consolidate the\n+`.mtimes` format into the `.idx` itself.\n+\n+Storing unreachable objects among multiple cruft packs (e.g., creating a new\n+cruft pack during each repacking operation including only unreachable objects\n+which aren't already stored in an earlier cruft pack) is significantly more\n+complicated to construct, and so aren't pursued here. The obvious drawback to\n+the current implementation is that the entire cruft pack must be re-written from\n+scratch.\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450204","messageId":"0f5d6d64924bcc1c81853ae246327338f7679a5e.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 03/17] pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:20:49Z","receivedAt":"2022-03-03T00:20:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This structure will be used to communicate the per-object mtimes when\nwriting a cruft pack. Here, we need the full packing_data structure\nbecause the mtime information is stored in an array there, not on the\nindividual object_entry's themselves (to avoid paying the overhead in\nstructure width for operations which do not generate a cruft pack).\n\nWe haven't passed this information down before because one of the two\ncallers (in bulk-checkin.c) does not have a packing_data structure at\nall. In that case (where no cruft pack will be generated), NULL is\npassed instead.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 3 ++-\n bulk-checkin.c         | 2 +-\n pack-write.c           | 1 +\n pack.h                 | 3 +++\n 4 files changed, 7 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 178e611f09..385970cb7b 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1254,7 +1254,8 @@ static void write_pack_file(void)\n \n \t\t\tstage_tmp_packfiles(&tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n-\t\t\t\t\t    &pack_idx_opts, hash, &idx_tmp_name);\n+\t\t\t\t\t    &to_pack, &pack_idx_opts, hash,\n+\t\t\t\t\t    &idx_tmp_name);\n \n \t\t\tif (write_bitmap_index) {\n \t\t\t\tsize_t tmpname_len = tmpname.len;\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac80..99f7596c4e 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -33,7 +33,7 @@ static void finish_tmp_packfile(struct strbuf *basename,\n \tchar *idx_tmp_name = NULL;\n \n \tstage_tmp_packfiles(basename, pack_tmp_name, written_list, nr_written,\n-\t\t\t    pack_idx_opts, hash, &idx_tmp_name);\n+\t\t\t    NULL, pack_idx_opts, hash, &idx_tmp_name);\n \trename_tmp_packfile_idx(basename, &idx_tmp_name);\n \n \tfree(idx_tmp_name);\ndiff --git a/pack-write.c b/pack-write.c\nindex a5846f3a34..d594e3008e 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -483,6 +483,7 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t const char *pack_tmp_name,\n \t\t\t struct pack_idx_entry **written_list,\n \t\t\t uint32_t nr_written,\n+\t\t\t struct packing_data *to_pack,\n \t\t\t struct pack_idx_option *pack_idx_opts,\n \t\t\t unsigned char hash[],\n \t\t\t char **idx_tmp_name)\ndiff --git a/pack.h b/pack.h\nindex b22bfc4a18..fd27cfdfd7 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -109,11 +109,14 @@ int encode_in_pack_object_header(unsigned char *hdr, int hdr_len,\n #define PH_ERROR_PROTOCOL\t(-3)\n int read_pack_header(int fd, struct pack_header *);\n \n+struct packing_data;\n+\n struct hashfile *create_tmp_packfile(char **pack_tmp_name);\n void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t const char *pack_tmp_name,\n \t\t\t struct pack_idx_entry **written_list,\n \t\t\t uint32_t nr_written,\n+\t\t\t struct packing_data *to_pack,\n \t\t\t struct pack_idx_option *pack_idx_opts,\n \t\t\t unsigned char hash[],\n \t\t\t char **idx_tmp_name);\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450205","messageId":"1ec754ad1b5c1051b52acef6ec72c0464c0eabf0.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:20:46Z","receivedAt":"2022-03-03T00:20:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"To store the individual mtimes of objects in a cruft pack, introduce a\nnew `.mtimes` format that can optionally accompany a single pack in the\nrepository.\n\nThe format is defined in Documentation/technical/pack-format.txt, and\nstores a 4-byte network order timestamp for each object in name (index)\norder.\n\nThis patch prepares for cruft packs by defining the `.mtimes` format,\nand introducing a basic API that callers can use to read out individual\nmtimes.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/technical/pack-format.txt |  19 ++++\n Makefile                                |   1 +\n builtin/repack.c                        |   1 +\n object-store.h                          |   5 +-\n pack-mtimes.c                           | 126 ++++++++++++++++++++++++\n pack-mtimes.h                           |  15 +++\n packfile.c                              |  19 +++-\n 7 files changed, 183 insertions(+), 3 deletions(-)\n create mode 100644 pack-mtimes.c\n create mode 100644 pack-mtimes.h\n\ndiff --git a/Documentation/technical/pack-format.txt b/Documentation/technical/pack-format.txt\nindex 6d3efb7d16..c443dbb526 100644\n--- a/Documentation/technical/pack-format.txt\n+++ b/Documentation/technical/pack-format.txt\n@@ -294,6 +294,25 @@ Pack file entry: <+\n \n All 4-byte numbers are in network order.\n \n+== pack-*.mtimes files have the format:\n+\n+  - A 4-byte magic number '0x4d544d45' ('MTME').\n+\n+  - A 4-byte version identifier (= 1).\n+\n+  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n+\n+  - A table of 4-byte unsigned integers in network order. The ith\n+    value is the modification time (mtime) of the ith object in the\n+    corresponding pack by lexicographic (index) order. The mtimes\n+    count standard epoch seconds.\n+\n+  - A trailer, containing a checksum of the corresponding packfile,\n+    and a checksum of all of the above (each having length according\n+    to the specified hash function).\n+\n+All 4-byte numbers are in network order.\n+\n == multi-pack-index (MIDX) files have the following format:\n \n The multi-pack-index files refer to multiple pack-files and loose objects.\ndiff --git a/Makefile b/Makefile\nindex 6f0b4b775f..1b186f4fd7 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -959,6 +959,7 @@ LIB_OBJS += oidtree.o\n LIB_OBJS += pack-bitmap-write.o\n LIB_OBJS += pack-bitmap.o\n LIB_OBJS += pack-check.o\n+LIB_OBJS += pack-mtimes.o\n LIB_OBJS += pack-objects.o\n LIB_OBJS += pack-revindex.o\n LIB_OBJS += pack-write.o\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex da1e364a75..f908f7d5dd 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -212,6 +212,7 @@ static struct {\n } exts[] = {\n \t{\".pack\"},\n \t{\".rev\", 1},\n+\t{\".mtimes\", 1},\n \t{\".bitmap\", 1},\n \t{\".promisor\", 1},\n \t{\".idx\"},\ndiff --git a/object-store.h b/object-store.h\nindex 6f89482df0..9b227661f2 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -115,12 +115,15 @@ struct packed_git {\n \t\t freshened:1,\n \t\t do_not_close:1,\n \t\t pack_promisor:1,\n-\t\t multi_pack_index:1;\n+\t\t multi_pack_index:1,\n+\t\t is_cruft:1;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n \tstruct revindex_entry *revindex;\n \tconst uint32_t *revindex_data;\n \tconst uint32_t *revindex_map;\n \tsize_t revindex_size;\n+\tconst uint32_t *mtimes_map;\n+\tsize_t mtimes_size;\n \t/* something like \".git/objects/pack/xxxxx.pack\" */\n \tchar pack_name[FLEX_ARRAY]; /* more */\n };\ndiff --git a/pack-mtimes.c b/pack-mtimes.c\nnew file mode 100644\nindex 0000000000..46ad584af1\n--- /dev/null\n+++ b/pack-mtimes.c\n@@ -0,0 +1,126 @@\n+#include \"pack-mtimes.h\"\n+#include \"object-store.h\"\n+#include \"packfile.h\"\n+\n+static char *pack_mtimes_filename(struct packed_git *p)\n+{\n+\tsize_t len;\n+\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n+\t\tBUG(\"pack_name does not end in .pack\");\n+\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n+\treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n+}\n+\n+#define MTIMES_HEADER_SIZE (12)\n+#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n+\n+struct mtimes_header {\n+\tuint32_t signature;\n+\tuint32_t version;\n+\tuint32_t hash_id;\n+};\n+\n+static int load_pack_mtimes_file(char *mtimes_file,\n+\t\t\t\t uint32_t num_objects,\n+\t\t\t\t const uint32_t **data_p, size_t *len_p)\n+{\n+\tint fd, ret = 0;\n+\tstruct stat st;\n+\tvoid *data = NULL;\n+\tsize_t mtimes_size;\n+\tstruct mtimes_header header;\n+\tuint32_t *hdr;\n+\n+\tfd = git_open(mtimes_file);\n+\n+\tif (fd < 0) {\n+\t\tret = -1;\n+\t\tgoto cleanup;\n+\t}\n+\tif (fstat(fd, &st)) {\n+\t\tret = error_errno(_(\"failed to read %s\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tmtimes_size = xsize_t(st.st_size);\n+\n+\tif (mtimes_size < MTIMES_MIN_SIZE) {\n+\t\tret = error(_(\"mtimes file %s is too small\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t), num_objects)) {\n+\t\tret = error(_(\"mtimes file %s is corrupt\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tdata = hdr = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\n+\theader.signature = ntohl(hdr[0]);\n+\theader.version = ntohl(hdr[1]);\n+\theader.hash_id = ntohl(hdr[2]);\n+\n+\tif (header.signature != MTIMES_SIGNATURE) {\n+\t\tret = error(_(\"mtimes file %s has unknown signature\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (header.version != 1) {\n+\t\tret = error(_(\"mtimes file %s has unsupported version %\"PRIu32),\n+\t\t\t    mtimes_file, header.version);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (!(header.hash_id == 1 || header.hash_id == 2)) {\n+\t\tret = error(_(\"mtimes file %s has unsupported hash id %\"PRIu32),\n+\t\t\t    mtimes_file, header.hash_id);\n+\t\tgoto cleanup;\n+\t}\n+\n+cleanup:\n+\tif (ret) {\n+\t\tif (data)\n+\t\t\tmunmap(data, mtimes_size);\n+\t} else {\n+\t\t*len_p = mtimes_size;\n+\t\t*data_p = (const uint32_t *)data;\n+\t}\n+\n+\tclose(fd);\n+\treturn ret;\n+}\n+\n+int load_pack_mtimes(struct packed_git *p)\n+{\n+\tchar *mtimes_name = NULL;\n+\tint ret = 0;\n+\n+\tif (!p->is_cruft)\n+\t\treturn ret; /* not a cruft pack */\n+\tif (p->mtimes_map)\n+\t\treturn ret; /* already loaded */\n+\n+\tret = open_pack_index(p);\n+\tif (ret < 0)\n+\t\tgoto cleanup;\n+\n+\tmtimes_name = pack_mtimes_filename(p);\n+\tret = load_pack_mtimes_file(mtimes_name,\n+\t\t\t\t    p->num_objects,\n+\t\t\t\t    &p->mtimes_map,\n+\t\t\t\t    &p->mtimes_size);\n+cleanup:\n+\tfree(mtimes_name);\n+\treturn ret;\n+}\n+\n+uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos)\n+{\n+\tif (!p->mtimes_map)\n+\t\tBUG(\"pack .mtimes file not loaded for %s\", p->pack_name);\n+\tif (p->num_objects <= pos)\n+\t\tBUG(\"pack .mtimes out-of-bounds (%\"PRIu32\" vs %\"PRIu32\")\",\n+\t\t    pos, p->num_objects);\n+\n+\treturn get_be32(p->mtimes_map + pos + 3);\n+}\ndiff --git a/pack-mtimes.h b/pack-mtimes.h\nnew file mode 100644\nindex 0000000000..38ddb9f893\n--- /dev/null\n+++ b/pack-mtimes.h\n@@ -0,0 +1,15 @@\n+#ifndef PACK_MTIMES_H\n+#define PACK_MTIMES_H\n+\n+#include \"git-compat-util.h\"\n+\n+#define MTIMES_SIGNATURE 0x4d544d45 /* \"MTME\" */\n+#define MTIMES_VERSION 1\n+\n+struct packed_git;\n+\n+int load_pack_mtimes(struct packed_git *p);\n+\n+uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos);\n+\n+#endif\ndiff --git a/packfile.c b/packfile.c\nindex 835b2d2716..fc0245fbab 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -334,12 +334,22 @@ static void close_pack_revindex(struct packed_git *p)\n \tp->revindex_data = NULL;\n }\n \n+static void close_pack_mtimes(struct packed_git *p)\n+{\n+\tif (!p->mtimes_map)\n+\t\treturn;\n+\n+\tmunmap((void *)p->mtimes_map, p->mtimes_size);\n+\tp->mtimes_map = NULL;\n+}\n+\n void close_pack(struct packed_git *p)\n {\n \tclose_pack_windows(p);\n \tclose_pack_fd(p);\n \tclose_pack_index(p);\n \tclose_pack_revindex(p);\n+\tclose_pack_mtimes(p);\n \toidset_clear(&p->bad_objects);\n }\n \n@@ -363,7 +373,7 @@ void close_object_store(struct raw_object_store *o)\n \n void unlink_pack_path(const char *pack_name, int force_delete)\n {\n-\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\"};\n+\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\", \".mtimes\"};\n \tint i;\n \tstruct strbuf buf = STRBUF_INIT;\n \tsize_t plen;\n@@ -718,6 +728,10 @@ struct packed_git *add_packed_git(const char *path, size_t path_len, int local)\n \tif (!access(p->pack_name, F_OK))\n \t\tp->pack_promisor = 1;\n \n+\txsnprintf(p->pack_name + path_len, alloc - path_len, \".mtimes\");\n+\tif (!access(p->pack_name, F_OK))\n+\t\tp->is_cruft = 1;\n+\n \txsnprintf(p->pack_name + path_len, alloc - path_len, \".pack\");\n \tif (stat(p->pack_name, &st) || !S_ISREG(st.st_mode)) {\n \t\tfree(p);\n@@ -869,7 +883,8 @@ static void prepare_pack(const char *full_name, size_t full_name_len,\n \t    ends_with(file_name, \".pack\") ||\n \t    ends_with(file_name, \".bitmap\") ||\n \t    ends_with(file_name, \".keep\") ||\n-\t    ends_with(file_name, \".promisor\"))\n+\t    ends_with(file_name, \".promisor\") ||\n+\t    ends_with(file_name, \".mtimes\"))\n \t\tstring_list_append(data->garbage, full_name);\n \telse\n \t\treport_garbage(PACKDIR_FILE_GARBAGE, full_name);\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450206","messageId":"135a07276b0a40b04f2c28d4f48c26b1af76c12c.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 04/17] chunk-format.h: extract oid_version()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:20:52Z","receivedAt":"2022-03-03T00:21:02Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"There are three definitions of an identical function which converts\n`the_hash_algo` into either 1 (for SHA-1) or 2 (for SHA-256). There is a\ncopy of this function for writing both the commit-graph and\nmulti-pack-index file, and another inline definition used to write the\n.rev header.\n\nConsolidate these into a single definition in chunk-format.h. It's not\nclear that this is the best header to define this function in, but it\nshould do for now.\n\n(Worth noting, the .rev caller expects a 4-byte unsigned, but the other\ntwo callers work with a single unsigned byte. The consolidated version\nuses the latter type, and lets the compiler widen it when required).\n\nAnother caller will be added in a subsequent patch.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n chunk-format.c | 12 ++++++++++++\n chunk-format.h |  3 +++\n commit-graph.c | 18 +++---------------\n midx.c         | 18 +++---------------\n pack-write.c   | 15 ++-------------\n 5 files changed, 23 insertions(+), 43 deletions(-)\n\ndiff --git a/chunk-format.c b/chunk-format.c\nindex 1c3dca62e2..0275b74a89 100644\n--- a/chunk-format.c\n+++ b/chunk-format.c\n@@ -181,3 +181,15 @@ int read_chunk(struct chunkfile *cf,\n \n \treturn CHUNK_NOT_FOUND;\n }\n+\n+uint8_t oid_version(const struct git_hash_algo *algop)\n+{\n+\tswitch (hash_algo_by_ptr(algop)) {\n+\tcase GIT_HASH_SHA1:\n+\t\treturn 1;\n+\tcase GIT_HASH_SHA256:\n+\t\treturn 2;\n+\tdefault:\n+\t\tdie(_(\"invalid hash version\"));\n+\t}\n+}\ndiff --git a/chunk-format.h b/chunk-format.h\nindex 9ccbe00377..7885aa0848 100644\n--- a/chunk-format.h\n+++ b/chunk-format.h\n@@ -2,6 +2,7 @@\n #define CHUNK_FORMAT_H\n \n #include \"git-compat-util.h\"\n+#include \"hash.h\"\n \n struct hashfile;\n struct chunkfile;\n@@ -65,4 +66,6 @@ int read_chunk(struct chunkfile *cf,\n \t       chunk_read_fn fn,\n \t       void *data);\n \n+uint8_t oid_version(const struct git_hash_algo *algop);\n+\n #endif\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 265c010122..f678d2c4a1 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -193,18 +193,6 @@ char *get_commit_graph_chain_filename(struct object_directory *odb)\n \treturn xstrfmt(\"%s/info/commit-graphs/commit-graph-chain\", odb->path);\n }\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n static struct commit_graph *alloc_commit_graph(void)\n {\n \tstruct commit_graph *g = xcalloc(1, sizeof(*g));\n@@ -365,9 +353,9 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n \t}\n \n \thash_version = *(unsigned char*)(data + 5);\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != oid_version(the_hash_algo)) {\n \t\terror(_(\"commit-graph hash version %X does not match version %X\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, oid_version(the_hash_algo));\n \t\treturn NULL;\n \t}\n \n@@ -1911,7 +1899,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \thashwrite_be32(f, GRAPH_SIGNATURE);\n \n \thashwrite_u8(f, GRAPH_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, oid_version(the_hash_algo));\n \thashwrite_u8(f, get_num_chunks(cf));\n \thashwrite_u8(f, ctx->num_commit_graphs_after - 1);\n \ndiff --git a/midx.c b/midx.c\nindex 865170bad0..65e670c5e2 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -41,18 +41,6 @@\n \n #define PACK_EXPIRED UINT_MAX\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n const unsigned char *get_midx_checksum(struct multi_pack_index *m)\n {\n \treturn m->data + m->data_len - the_hash_algo->rawsz;\n@@ -134,9 +122,9 @@ struct multi_pack_index *load_multi_pack_index(const char *object_dir, int local\n \t\t      m->version);\n \n \thash_version = m->data[MIDX_BYTE_HASH_VERSION];\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != oid_version(the_hash_algo)) {\n \t\terror(_(\"multi-pack-index hash version %u does not match version %u\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, oid_version(the_hash_algo));\n \t\tgoto cleanup_fail;\n \t}\n \tm->hash_len = the_hash_algo->rawsz;\n@@ -420,7 +408,7 @@ static size_t write_midx_header(struct hashfile *f,\n {\n \thashwrite_be32(f, MIDX_SIGNATURE);\n \thashwrite_u8(f, MIDX_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, oid_version(the_hash_algo));\n \thashwrite_u8(f, num_chunks);\n \thashwrite_u8(f, 0); /* unused */\n \thashwrite_be32(f, num_packs);\ndiff --git a/pack-write.c b/pack-write.c\nindex d594e3008e..ff305b404c 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -2,6 +2,7 @@\n #include \"pack.h\"\n #include \"csum-file.h\"\n #include \"remote.h\"\n+#include \"chunk-format.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -181,21 +182,9 @@ static int pack_order_cmp(const void *va, const void *vb, void *ctx)\n \n static void write_rev_header(struct hashfile *f)\n {\n-\tuint32_t oid_version;\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\toid_version = 1;\n-\t\tbreak;\n-\tcase GIT_HASH_SHA256:\n-\t\toid_version = 2;\n-\t\tbreak;\n-\tdefault:\n-\t\tdie(\"write_rev_header: unknown hash version\");\n-\t}\n-\n \thashwrite_be32(f, RIDX_SIGNATURE);\n \thashwrite_be32(f, RIDX_VERSION);\n-\thashwrite_be32(f, oid_version);\n+\thashwrite_be32(f, oid_version(the_hash_algo));\n }\n \n static void write_rev_index_positions(struct hashfile *f,\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450207","messageId":"0600503856dbccb135aaead27693b6815a774b4f.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 05/17] pack-mtimes: support writing pack .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:20:55Z","receivedAt":"2022-03-03T00:21:05Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Now that the `.mtimes` format is defined, supplement the pack-write API\nto be able to conditionally write an `.mtimes` file along with a pack by\nsetting an additional flag and passing an oidmap that contains the\ntimestamps corresponding to each object in the pack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n pack-objects.c |  6 ++++\n pack-objects.h | 25 ++++++++++++++++\n pack-write.c   | 77 ++++++++++++++++++++++++++++++++++++++++++++++++++\n pack.h         |  1 +\n 4 files changed, 109 insertions(+)\n\ndiff --git a/pack-objects.c b/pack-objects.c\nindex fe2a4eace9..272e8d4517 100644\n--- a/pack-objects.c\n+++ b/pack-objects.c\n@@ -170,6 +170,9 @@ struct object_entry *packlist_alloc(struct packing_data *pdata,\n \n \t\tif (pdata->layer)\n \t\t\tREALLOC_ARRAY(pdata->layer, pdata->nr_alloc);\n+\n+\t\tif (pdata->cruft_mtime)\n+\t\t\tREALLOC_ARRAY(pdata->cruft_mtime, pdata->nr_alloc);\n \t}\n \n \tnew_entry = pdata->objects + pdata->nr_objects++;\n@@ -198,6 +201,9 @@ struct object_entry *packlist_alloc(struct packing_data *pdata,\n \tif (pdata->layer)\n \t\tpdata->layer[pdata->nr_objects - 1] = 0;\n \n+\tif (pdata->cruft_mtime)\n+\t\tpdata->cruft_mtime[pdata->nr_objects - 1] = 0;\n+\n \treturn new_entry;\n }\n \ndiff --git a/pack-objects.h b/pack-objects.h\nindex dca2351ef9..393b9db546 100644\n--- a/pack-objects.h\n+++ b/pack-objects.h\n@@ -168,6 +168,14 @@ struct packing_data {\n \t/* delta islands */\n \tunsigned int *tree_depth;\n \tunsigned char *layer;\n+\n+\t/*\n+\t * Used when writing cruft packs.\n+\t *\n+\t * Object mtimes are stored in pack order when writing, but\n+\t * written out in lexicographic (index) order.\n+\t */\n+\tuint32_t *cruft_mtime;\n };\n \n void prepare_packing_data(struct repository *r, struct packing_data *pdata);\n@@ -289,4 +297,21 @@ static inline void oe_set_layer(struct packing_data *pack,\n \tpack->layer[e - pack->objects] = layer;\n }\n \n+static inline uint32_t oe_cruft_mtime(struct packing_data *pack,\n+\t\t\t\t      struct object_entry *e)\n+{\n+\tif (!pack->cruft_mtime)\n+\t\treturn 0;\n+\treturn pack->cruft_mtime[e - pack->objects];\n+}\n+\n+static inline void oe_set_cruft_mtime(struct packing_data *pack,\n+\t\t\t\t      struct object_entry *e,\n+\t\t\t\t      uint32_t mtime)\n+{\n+\tif (!pack->cruft_mtime)\n+\t\tCALLOC_ARRAY(pack->cruft_mtime, pack->nr_alloc);\n+\tpack->cruft_mtime[e - pack->objects] = mtime;\n+}\n+\n #endif\ndiff --git a/pack-write.c b/pack-write.c\nindex ff305b404c..270280c4df 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -3,6 +3,10 @@\n #include \"csum-file.h\"\n #include \"remote.h\"\n #include \"chunk-format.h\"\n+#include \"pack-mtimes.h\"\n+#include \"oidmap.h\"\n+#include \"chunk-format.h\"\n+#include \"pack-objects.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -276,6 +280,70 @@ const char *write_rev_file_order(const char *rev_name,\n \treturn rev_name;\n }\n \n+static void write_mtimes_header(struct hashfile *f)\n+{\n+\thashwrite_be32(f, MTIMES_SIGNATURE);\n+\thashwrite_be32(f, MTIMES_VERSION);\n+\thashwrite_be32(f, oid_version(the_hash_algo));\n+}\n+\n+/*\n+ * Writes the object mtimes of \"objects\" for use in a .mtimes file.\n+ * Note that objects must be in lexicographic (index) order, which is\n+ * the expected ordering of these values in the .mtimes file.\n+ */\n+static void write_mtimes_objects(struct hashfile *f,\n+\t\t\t\t struct packing_data *to_pack,\n+\t\t\t\t struct pack_idx_entry **objects,\n+\t\t\t\t uint32_t nr_objects)\n+{\n+\tuint32_t i;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\tstruct object_entry *e = (struct object_entry*)objects[i];\n+\t\thashwrite_be32(f, oe_cruft_mtime(to_pack, e));\n+\t}\n+}\n+\n+static void write_mtimes_trailer(struct hashfile *f, const unsigned char *hash)\n+{\n+\thashwrite(f, hash, the_hash_algo->rawsz);\n+}\n+\n+static const char *write_mtimes_file(const char *mtimes_name,\n+\t\t\t\t     struct packing_data *to_pack,\n+\t\t\t\t     struct pack_idx_entry **objects,\n+\t\t\t\t     uint32_t nr_objects,\n+\t\t\t\t     const unsigned char *hash)\n+{\n+\tstruct hashfile *f;\n+\tint fd;\n+\n+\tif (!to_pack)\n+\t\tBUG(\"cannot call write_mtimes_file with NULL packing_data\");\n+\n+\tif (!mtimes_name) {\n+\t\tstruct strbuf tmp_file = STRBUF_INIT;\n+\t\tfd = odb_mkstemp(&tmp_file, \"pack/tmp_mtimes_XXXXXX\");\n+\t\tmtimes_name = strbuf_detach(&tmp_file, NULL);\n+\t} else {\n+\t\tunlink(mtimes_name);\n+\t\tfd = xopen(mtimes_name, O_CREAT|O_EXCL|O_WRONLY, 0600);\n+\t}\n+\tf = hashfd(fd, mtimes_name);\n+\n+\twrite_mtimes_header(f);\n+\twrite_mtimes_objects(f, to_pack, objects, nr_objects);\n+\twrite_mtimes_trailer(f, hash);\n+\n+\tif (adjust_shared_perm(mtimes_name) < 0)\n+\t\tdie(_(\"failed to make %s readable\"), mtimes_name);\n+\n+\tfinalize_hashfile(f, NULL,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE | CSUM_FSYNC);\n+\n+\treturn mtimes_name;\n+}\n+\n off_t write_pack_header(struct hashfile *f, uint32_t nr_entries)\n {\n \tstruct pack_header hdr;\n@@ -478,6 +546,7 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t char **idx_tmp_name)\n {\n \tconst char *rev_tmp_name = NULL;\n+\tconst char *mtimes_tmp_name = NULL;\n \n \tif (adjust_shared_perm(pack_tmp_name))\n \t\tdie_errno(\"unable to make temporary pack file readable\");\n@@ -490,9 +559,17 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \trev_tmp_name = write_rev_file(NULL, written_list, nr_written, hash,\n \t\t\t\t      pack_idx_opts->flags);\n \n+\tif (pack_idx_opts->flags & WRITE_MTIMES) {\n+\t\tmtimes_tmp_name = write_mtimes_file(NULL, to_pack, written_list,\n+\t\t\t\t\t\t    nr_written,\n+\t\t\t\t\t\t    hash);\n+\t}\n+\n \trename_tmp_packfile(name_buffer, pack_tmp_name, \"pack\");\n \tif (rev_tmp_name)\n \t\trename_tmp_packfile(name_buffer, rev_tmp_name, \"rev\");\n+\tif (mtimes_tmp_name)\n+\t\trename_tmp_packfile(name_buffer, mtimes_tmp_name, \"mtimes\");\n }\n \n void write_promisor_file(const char *promisor_name, struct ref **sought, int nr_sought)\ndiff --git a/pack.h b/pack.h\nindex fd27cfdfd7..01d385903a 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -44,6 +44,7 @@ struct pack_idx_option {\n #define WRITE_IDX_STRICT 02\n #define WRITE_REV 04\n #define WRITE_REV_VERIFY 010\n+#define WRITE_MTIMES 020\n \n \tuint32_t version;\n \tuint32_t off32_limit;\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450208","messageId":"4780c8437bd2dcdf2c038d84160a4c575e92e58d.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 06/17] t/helper: add 'pack-mtimes' test-tool","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:20:57Z","receivedAt":"2022-03-03T00:21:05Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In the next patch, we will implement and test support for writing a\ncruft pack via a special mode of `git pack-objects`. To make sure that\nobjects are written with the correct timestamps, and a new test-tool\nthat can dump the object names and corresponding timestamps from a given\n`.mtimes` file.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Makefile                    |  1 +\n t/helper/test-pack-mtimes.c | 56 +++++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c        |  1 +\n t/helper/test-tool.h        |  1 +\n 4 files changed, 59 insertions(+)\n create mode 100644 t/helper/test-pack-mtimes.c\n\ndiff --git a/Makefile b/Makefile\nindex 1b186f4fd7..5c0ed1ade7 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -727,6 +727,7 @@ TEST_BUILTINS_OBJS += test-oid-array.o\n TEST_BUILTINS_OBJS += test-oidmap.o\n TEST_BUILTINS_OBJS += test-oidtree.o\n TEST_BUILTINS_OBJS += test-online-cpus.o\n+TEST_BUILTINS_OBJS += test-pack-mtimes.o\n TEST_BUILTINS_OBJS += test-parse-options.o\n TEST_BUILTINS_OBJS += test-parse-pathspec-file.o\n TEST_BUILTINS_OBJS += test-partial-clone.o\ndiff --git a/t/helper/test-pack-mtimes.c b/t/helper/test-pack-mtimes.c\nnew file mode 100644\nindex 0000000000..f7b79daf4c\n--- /dev/null\n+++ b/t/helper/test-pack-mtimes.c\n@@ -0,0 +1,56 @@\n+#include \"git-compat-util.h\"\n+#include \"test-tool.h\"\n+#include \"strbuf.h\"\n+#include \"object-store.h\"\n+#include \"packfile.h\"\n+#include \"pack-mtimes.h\"\n+\n+static void dump_mtimes(struct packed_git *p)\n+{\n+\tuint32_t i;\n+\tif (load_pack_mtimes(p) < 0)\n+\t\tdie(\"could not load pack .mtimes\");\n+\n+\tfor (i = 0; i < p->num_objects; i++) {\n+\t\tstruct object_id oid;\n+\t\tif (nth_packed_object_id(&oid, p, i) < 0)\n+\t\t\tdie(\"could not load object id at position %\"PRIu32, i);\n+\n+\t\tprintf(\"%s %\"PRIu32\"\\n\",\n+\t\t       oid_to_hex(&oid), nth_packed_mtime(p, i));\n+\t}\n+}\n+\n+static const char *pack_mtimes_usage = \"\\n\"\n+\"  test-tool pack-mtimes <pack-name.mtimes>\";\n+\n+int cmd__pack_mtimes(int argc, const char **argv)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct packed_git *p;\n+\n+\tsetup_git_directory();\n+\n+\tif (argc != 2)\n+\t\tusage(pack_mtimes_usage);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tstrbuf_addstr(&buf, basename(p->pack_name));\n+\t\tstrbuf_strip_suffix(&buf, \".pack\");\n+\t\tstrbuf_addstr(&buf, \".mtimes\");\n+\n+\t\tif (!strcmp(buf.buf, argv[1]))\n+\t\t\tbreak;\n+\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstrbuf_release(&buf);\n+\n+\tif (!p)\n+\t\tdie(\"could not find pack '%s'\", argv[1]);\n+\n+\tdump_mtimes(p);\n+\n+\treturn 0;\n+}\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex e6ec69cf32..7d472b31fd 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -47,6 +47,7 @@ static struct test_cmd cmds[] = {\n \t{ \"oidmap\", cmd__oidmap },\n \t{ \"oidtree\", cmd__oidtree },\n \t{ \"online-cpus\", cmd__online_cpus },\n+\t{ \"pack-mtimes\", cmd__pack_mtimes },\n \t{ \"parse-options\", cmd__parse_options },\n \t{ \"parse-pathspec-file\", cmd__parse_pathspec_file },\n \t{ \"partial-clone\", cmd__partial_clone },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex 20756eefdd..0ac4f32955 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -37,6 +37,7 @@ int cmd__mktemp(int argc, const char **argv);\n int cmd__oidmap(int argc, const char **argv);\n int cmd__oidtree(int argc, const char **argv);\n int cmd__online_cpus(int argc, const char **argv);\n+int cmd__pack_mtimes(int argc, const char **argv);\n int cmd__parse_options(int argc, const char **argv);\n int cmd__parse_pathspec_file(int argc, const char** argv);\n int cmd__partial_clone(int argc, const char **argv);\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450209","messageId":"33862a07c927184a40ccbfe5182404923a392c4a.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 07/17] builtin/pack-objects.c: return from create_object_entry()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:21:00Z","receivedAt":"2022-03-03T00:21:16Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"A new caller in the next commit will want to immediately modify the\nobject_entry structure created by create_object_entry(). Instead of\nforcing that caller to wastefully look-up the entry we just created,\nreturn it from create_object_entry() instead.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 16 +++++++++-------\n 1 file changed, 9 insertions(+), 7 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 385970cb7b..3f08a3c63a 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1508,13 +1508,13 @@ static int want_object_in_pack(const struct object_id *oid,\n \treturn 1;\n }\n \n-static void create_object_entry(const struct object_id *oid,\n-\t\t\t\tenum object_type type,\n-\t\t\t\tuint32_t hash,\n-\t\t\t\tint exclude,\n-\t\t\t\tint no_try_delta,\n-\t\t\t\tstruct packed_git *found_pack,\n-\t\t\t\toff_t found_offset)\n+static struct object_entry *create_object_entry(const struct object_id *oid,\n+\t\t\t\t\t\tenum object_type type,\n+\t\t\t\t\t\tuint32_t hash,\n+\t\t\t\t\t\tint exclude,\n+\t\t\t\t\t\tint no_try_delta,\n+\t\t\t\t\t\tstruct packed_git *found_pack,\n+\t\t\t\t\t\toff_t found_offset)\n {\n \tstruct object_entry *entry;\n \n@@ -1531,6 +1531,8 @@ static void create_object_entry(const struct object_id *oid,\n \t}\n \n \tentry->no_try_delta = no_try_delta;\n+\n+\treturn entry;\n }\n \n static const char no_closure_warning[] = N_(\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450210","messageId":"22705e4887b5c9e3d7ef9ff1eadaabeeac0d57da.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:21:02Z","receivedAt":"2022-03-03T00:21:17Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Teach `pack-objects` how to generate a cruft pack when no objects are\ndropped (i.e., `--cruft-expiration=never`). Later patches will teach\n`pack-objects` how to generate a cruft pack that prunes objects.\n\nWhen generating a cruft pack which does not prune objects, we want to\ncollect all unreachable objects into a single pack (noting and updating\ntheir mtimes as we accumulate them). Ordinary use will pass the result\nof a `git repack -A` as a kept pack, so when this patch says \"kept\npack\", readers should think \"reachable objects\".\n\nGenerating a non-expiring cruft packs works as follows:\n\n  - Callers provide a list of every pack they know about, and indicate\n    which packs are about to be removed.\n\n  - All packs which are going to be removed (we'll call these the\n    redundant ones) are marked as kept in-core.\n\n    Any packs the caller did not mention (but are known to the\n    `pack-objects` process) are also marked as kept in-core. Packs not\n    mentioned by the caller are assumed to be unknown to them, i.e.,\n    they entered the repository after the caller decided which packs\n    should be kept and which should be discarded.\n\n    Since we do not want to include objects in these \"unknown\" packs\n    (because we don't know which of their objects are or aren't\n    reachable), these are also marked as kept in-core.\n\n  - Then, we enumerate all objects in the repository, and add them to\n    our packing list if they do not appear in an in-core kept pack.\n\nThis results in a new cruft pack which contains all known objects that\naren't included in the kept packs. When the kept pack is the result of\n`git repack -A`, the resulting pack contains all unreachable objects.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-pack-objects.txt |  30 ++++\n builtin/pack-objects.c             | 201 +++++++++++++++++++++++++-\n object-file.c                      |   2 +-\n object-store.h                     |   2 +\n t/t5329-pack-objects-cruft.sh      | 218 +++++++++++++++++++++++++++++\n 5 files changed, 448 insertions(+), 5 deletions(-)\n create mode 100755 t/t5329-pack-objects-cruft.sh\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex f8344e1e5b..a9995a932c 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -13,6 +13,7 @@ SYNOPSIS\n \t[--no-reuse-delta] [--delta-base-offset] [--non-empty]\n \t[--local] [--incremental] [--window=<n>] [--depth=<n>]\n \t[--revs [--unpacked | --all]] [--keep-pack=<pack-name>]\n+\t[--cruft] [--cruft-expiration=<time>]\n \t[--stdout [--filter=<filter-spec>] | <base-name>]\n \t[--shallow] [--keep-true-parents] [--[no-]sparse] < <object-list>\n \n@@ -95,6 +96,35 @@ base-name::\n Incompatible with `--revs`, or options that imply `--revs` (such as\n `--all`), with the exception of `--unpacked`, which is compatible.\n \n+--cruft::\n+\tPacks unreachable objects into a separate \"cruft\" pack, denoted\n+\tby the existence of a `.mtimes` file. Typically used by `git\n+\trepack --cruft`. Callers provide a list of pack names and\n+\tindicate which packs will remain in the repository, along with\n+\twhich packs will be deleted (indicated by the `-` prefix). The\n+\tcontents of the cruft pack are all objects not contained in the\n+\tsurviving packs which have not exceeded the grace period (see\n+\t`--cruft-expiration` below), or which have exceeded the grace\n+\tperiod, but are reachable from an other object which hasn't.\n++\n+When the input lists a pack containing all reachable objects (and lists\n+all other packs as pending deletion), the corresponding cruft pack will\n+contain all unreachable objects (with mtime newer than the\n+`--cruft-expiration`) along with any unreachable objects whose mtime is\n+older than the `--cruft-expiration`, but are reachable from an\n+unreachable object whose mtime is newer than the `--cruft-expiration`).\n++\n+Incompatible with `--unpack-unreachable`, `--keep-unreachable`,\n+`--pack-loose-unreachable`, `--stdin-packs`, as well as any other\n+options which imply `--revs`. Also incompatible with `--max-pack-size`;\n+when this option is set, the maximum pack size is not inferred from\n+`pack.packSizeLimit`.\n+\n+--cruft-expiration=<approxidate>::\n+\tIf specified, objects are eliminated from the cruft pack if they\n+\thave an mtime older than `<approxidate>`. If unspecified (and\n+\tgiven `--cruft`), then no objects are eliminated.\n+\n --window=<n>::\n --depth=<n>::\n \tThese two options affect how the objects contained in\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 3f08a3c63a..5ba4fc9c2c 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -36,6 +36,7 @@\n #include \"trace2.h\"\n #include \"shallow.h\"\n #include \"promisor-remote.h\"\n+#include \"pack-mtimes.h\"\n \n /*\n  * Objects we are going to pack are collected in the `to_pack` structure.\n@@ -194,6 +195,8 @@ static int reuse_delta = 1, reuse_object = 1;\n static int keep_unreachable, unpack_unreachable, include_tag;\n static timestamp_t unpack_unreachable_expiration;\n static int pack_loose_unreachable;\n+static int cruft;\n+static timestamp_t cruft_expiration;\n static int local;\n static int have_non_local_packs;\n static int incremental;\n@@ -1252,6 +1255,9 @@ static void write_pack_file(void)\n \t\t\t\t\t&to_pack, written_list, nr_written);\n \t\t\t}\n \n+\t\t\tif (cruft)\n+\t\t\t\tpack_idx_opts.flags |= WRITE_MTIMES;\n+\n \t\t\tstage_tmp_packfiles(&tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n \t\t\t\t\t    &to_pack, &pack_idx_opts, hash,\n@@ -3389,6 +3395,135 @@ static void read_packs_list_from_stdin(void)\n \tstring_list_clear(&exclude_packs, 0);\n }\n \n+static void add_cruft_object_entry(const struct object_id *oid, enum object_type type,\n+\t\t\t\t   struct packed_git *pack, off_t offset,\n+\t\t\t\t   const char *name, uint32_t mtime)\n+{\n+\tstruct object_entry *entry;\n+\n+\tdisplay_progress(progress_state, ++nr_seen);\n+\n+\tentry = packlist_find(&to_pack, oid);\n+\tif (entry) {\n+\t\tif (name) {\n+\t\t\tentry->hash = pack_name_hash(name);\n+\t\t\tentry->no_try_delta = no_try_delta(name);\n+\t\t}\n+\t} else {\n+\t\tif (!want_object_in_pack(oid, 0, &pack, &offset))\n+\t\t\treturn;\n+\t\tif (!pack && type == OBJ_BLOB && !has_loose_object(oid)) {\n+\t\t\t/*\n+\t\t\t * If a traversed tree has a missing blob then we want\n+\t\t\t * to avoid adding that missing object to our pack.\n+\t\t\t *\n+\t\t\t * This only applies to missing blobs, not trees,\n+\t\t\t * because the traversal needs to parse sub-trees but\n+\t\t\t * not blobs.\n+\t\t\t *\n+\t\t\t * Note we only perform this check when we couldn't\n+\t\t\t * already find the object in a pack, so we're really\n+\t\t\t * limited to \"ensure non-tip blobs which don't exist in\n+\t\t\t * packs do exist via loose objects\". Confused?\n+\t\t\t */\n+\t\t\treturn;\n+\t\t}\n+\n+\t\tentry = create_object_entry(oid, type, pack_name_hash(name),\n+\t\t\t\t\t    0, name && no_try_delta(name),\n+\t\t\t\t\t    pack, offset);\n+\t}\n+\n+\tif (mtime > oe_cruft_mtime(&to_pack, entry))\n+\t\toe_set_cruft_mtime(&to_pack, entry, mtime);\n+\treturn;\n+}\n+\n+static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n+{\n+\tstruct string_list_item *item = NULL;\n+\tfor_each_string_list_item(item, packs) {\n+\t\tstruct packed_git *p = item->util;\n+\t\tif (!p)\n+\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n+\t\tp->pack_keep_in_core = keep;\n+\t}\n+}\n+\n+static void add_unreachable_loose_objects(void);\n+static void add_objects_in_unpacked_packs(void);\n+\n+static void enumerate_cruft_objects(void)\n+{\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Enumerating cruft objects\"), 0);\n+\n+\tadd_objects_in_unpacked_packs();\n+\tadd_unreachable_loose_objects();\n+\n+\tstop_progress(&progress_state);\n+}\n+\n+static void read_cruft_objects(void)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct string_list discard_packs = STRING_LIST_INIT_DUP;\n+\tstruct string_list fresh_packs = STRING_LIST_INIT_DUP;\n+\tstruct packed_git *p;\n+\n+\tignore_packed_keep_in_core = 1;\n+\n+\twhile (strbuf_getline(&buf, stdin) != EOF) {\n+\t\tif (!buf.len)\n+\t\t\tcontinue;\n+\n+\t\tif (*buf.buf == '-')\n+\t\t\tstring_list_append(&discard_packs, buf.buf + 1);\n+\t\telse\n+\t\t\tstring_list_append(&fresh_packs, buf.buf);\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstring_list_sort(&discard_packs);\n+\tstring_list_sort(&fresh_packs);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tconst char *pack_name = pack_basename(p);\n+\t\tstruct string_list_item *item;\n+\n+\t\titem = string_list_lookup(&fresh_packs, pack_name);\n+\t\tif (!item)\n+\t\t\titem = string_list_lookup(&discard_packs, pack_name);\n+\n+\t\tif (item) {\n+\t\t\titem->util = p;\n+\t\t} else {\n+\t\t\t/*\n+\t\t\t * This pack wasn't mentioned in either the \"fresh\" or\n+\t\t\t * \"discard\" list, so the caller didn't know about it.\n+\t\t\t *\n+\t\t\t * Mark it as kept so that its objects are ignored by\n+\t\t\t * add_unseen_recent_objects_to_traversal(). We'll\n+\t\t\t * unmark it before starting the traversal so it doesn't\n+\t\t\t * halt the traversal early.\n+\t\t\t */\n+\t\t\tp->pack_keep_in_core = 1;\n+\t\t}\n+\t}\n+\n+\tmark_pack_kept_in_core(&fresh_packs, 1);\n+\tmark_pack_kept_in_core(&discard_packs, 0);\n+\n+\tif (cruft_expiration)\n+\t\tdie(\"--cruft-expiration not yet implemented\");\n+\telse\n+\t\tenumerate_cruft_objects();\n+\n+\tstrbuf_release(&buf);\n+\tstring_list_clear(&discard_packs, 0);\n+\tstring_list_clear(&fresh_packs, 0);\n+}\n+\n static void read_object_list_from_stdin(void)\n {\n \tchar line[GIT_MAX_HEXSZ + 1 + PATH_MAX + 2];\n@@ -3521,7 +3656,24 @@ static int add_object_in_unpacked_pack(const struct object_id *oid,\n \t\t\t\t       uint32_t pos,\n \t\t\t\t       void *_data)\n {\n-\tadd_object_entry(oid, OBJ_NONE, \"\", 0);\n+\tif (cruft) {\n+\t\toff_t offset;\n+\t\ttime_t mtime;\n+\n+\t\tif (pack->is_cruft) {\n+\t\t\tif (load_pack_mtimes(pack) < 0)\n+\t\t\t\tdie(_(\"could not load cruft pack .mtimes\"));\n+\t\t\tmtime = nth_packed_mtime(pack, pos);\n+\t\t} else {\n+\t\t\tmtime = pack->mtime;\n+\t\t}\n+\t\toffset = nth_packed_object_offset(pack, pos);\n+\n+\t\tadd_cruft_object_entry(oid, OBJ_NONE, pack, offset,\n+\t\t\t\t       NULL, mtime);\n+\t} else {\n+\t\tadd_object_entry(oid, OBJ_NONE, \"\", 0);\n+\t}\n \treturn 0;\n }\n \n@@ -3545,7 +3697,19 @@ static int add_loose_object(const struct object_id *oid, const char *path,\n \t\treturn 0;\n \t}\n \n-\tadd_object_entry(oid, type, \"\", 0);\n+\tif (cruft) {\n+\t\tstruct stat st;\n+\t\tif (stat(path, &st) < 0) {\n+\t\t\tif (errno == ENOENT)\n+\t\t\t\treturn 0;\n+\t\t\treturn error_errno(\"unable to stat %s\", oid_to_hex(oid));\n+\t\t}\n+\n+\t\tadd_cruft_object_entry(oid, type, NULL, 0, NULL,\n+\t\t\t\t       st.st_mtime);\n+\t} else {\n+\t\tadd_object_entry(oid, type, \"\", 0);\n+\t}\n \treturn 0;\n }\n \n@@ -3864,6 +4028,20 @@ static int option_parse_unpack_unreachable(const struct option *opt,\n \treturn 0;\n }\n \n+static int option_parse_cruft_expiration(const struct option *opt,\n+\t\t\t\t\t const char *arg, int unset)\n+{\n+\tif (unset) {\n+\t\tcruft = 0;\n+\t\tcruft_expiration = 0;\n+\t} else {\n+\t\tcruft = 1;\n+\t\tif (arg)\n+\t\t\tcruft_expiration = approxidate(arg);\n+\t}\n+\treturn 0;\n+}\n+\n int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n {\n \tint use_internal_rev_list = 0;\n@@ -3936,6 +4114,10 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tOPT_CALLBACK_F(0, \"unpack-unreachable\", NULL, N_(\"time\"),\n \t\t  N_(\"unpack unreachable objects newer than <time>\"),\n \t\t  PARSE_OPT_OPTARG, option_parse_unpack_unreachable),\n+\t\tOPT_BOOL(0, \"cruft\", &cruft, N_(\"create a cruft pack\")),\n+\t\tOPT_CALLBACK_F(0, \"cruft-expiration\", NULL, N_(\"time\"),\n+\t\t  N_(\"expire cruft objects older than <time>\"),\n+\t\t  PARSE_OPT_OPTARG, option_parse_cruft_expiration),\n \t\tOPT_BOOL(0, \"sparse\", &sparse,\n \t\t\t N_(\"use the sparse reachability algorithm\")),\n \t\tOPT_BOOL(0, \"thin\", &thin,\n@@ -4062,7 +4244,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \n \tif (!HAVE_THREADS && delta_search_threads != 1)\n \t\twarning(_(\"no threads support, ignoring --threads\"));\n-\tif (!pack_to_stdout && !pack_size_limit)\n+\tif (!pack_to_stdout && !pack_size_limit && !cruft)\n \t\tpack_size_limit = pack_size_limit_cfg;\n \tif (pack_to_stdout && pack_size_limit)\n \t\tdie(_(\"--max-pack-size cannot be used to build a pack for transfer\"));\n@@ -4089,6 +4271,15 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (stdin_packs && use_internal_rev_list)\n \t\tdie(_(\"cannot use internal rev list with --stdin-packs\"));\n \n+\tif (cruft) {\n+\t\tif (use_internal_rev_list)\n+\t\t\tdie(_(\"cannot use internal rev list with --cruft\"));\n+\t\tif (stdin_packs)\n+\t\t\tdie(_(\"cannot use --stdin-packs with --cruft\"));\n+\t\tif (pack_size_limit)\n+\t\t\tdie(_(\"cannot use --max-pack-size with --cruft\"));\n+\t}\n+\n \t/*\n \t * \"soft\" reasons not to use bitmaps - for on-disk repack by default we want\n \t *\n@@ -4145,7 +4336,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t    the_repository);\n \tprepare_packing_data(the_repository, &to_pack);\n \n-\tif (progress)\n+\tif (progress && !cruft)\n \t\tprogress_state = start_progress(_(\"Enumerating objects\"), 0);\n \tif (stdin_packs) {\n \t\t/* avoids adding objects in excluded packs */\n@@ -4153,6 +4344,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tread_packs_list_from_stdin();\n \t\tif (rev_list_unpacked)\n \t\t\tadd_unreachable_loose_objects();\n+\t} else if (cruft) {\n+\t\tread_cruft_objects();\n \t} else if (!use_internal_rev_list) {\n \t\tread_object_list_from_stdin();\n \t} else {\ndiff --git a/object-file.c b/object-file.c\nindex 8be57f48de..e80da1368d 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -996,7 +996,7 @@ int has_loose_object_nonlocal(const struct object_id *oid)\n \treturn check_and_freshen_nonlocal(oid, 0);\n }\n \n-static int has_loose_object(const struct object_id *oid)\n+int has_loose_object(const struct object_id *oid)\n {\n \treturn check_and_freshen(oid, 0);\n }\ndiff --git a/object-store.h b/object-store.h\nindex 9b227661f2..6b025dc670 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -334,6 +334,8 @@ int repo_has_object_file_with_flags(struct repository *r,\n  */\n int has_loose_object_nonlocal(const struct object_id *);\n \n+int has_loose_object(const struct object_id *);\n+\n void assert_oid_type(const struct object_id *oid, enum object_type expect);\n \n /*\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nnew file mode 100755\nindex 0000000000..003ca7344e\n--- /dev/null\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -0,0 +1,218 @@\n+#!/bin/sh\n+\n+test_description='cruft pack related pack-objects tests'\n+. ./test-lib.sh\n+\n+objdir=.git/objects\n+packdir=$objdir/pack\n+\n+basic_cruft_pack_tests () {\n+\texpire=\"$1\"\n+\n+\ttest_expect_success \"unreachable loose objects are packed (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit base &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit loose &&\n+\n+\t\t\ttest-tool chmtime +2000 \"$objdir/$(test_oid_to_path \\\n+\t\t\t\t$(git rev-parse loose:loose.t))\" &&\n+\t\t\ttest-tool chmtime +1000 \"$objdir/$(test_oid_to_path \\\n+\t\t\t\t$(git rev-parse loose^{tree}))\" &&\n+\n+\t\t\t(\n+\t\t\t\tgit rev-list --objects --no-object-names base..loose |\n+\t\t\t\twhile read oid\n+\t\t\t\tdo\n+\t\t\t\t\tpath=\"$objdir/$(test_oid_to_path \"$oid\")\" &&\n+\t\t\t\t\tprintf \"%s %d\\n\" \"$oid\" \"$(test-tool chmtime --get \"$path\")\"\n+\t\t\t\tdone |\n+\t\t\t\tsort -k1\n+\t\t\t) >expect &&\n+\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t\t--cruft-expiration=\"$expire\" $packdir/pack)\" &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"unreachable packed objects are packed (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit packed &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit other &&\n+\n+\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\t\tother=\"$(git pack-objects --delta-base-offset \\\n+\t\t\t\t$packdir/pack <objects)\" &&\n+\t\t\tgit prune-packed &&\n+\n+\t\t\ttest-tool chmtime --get -100 \"$packdir/pack-$other.pack\" >expect &&\n+\n+\t\t\tcruft=\"$(git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-pack-$other.pack\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\t\tcut -d\" \" -f2 <actual.raw | sort -u >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"unreachable cruft objects are repacked (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit packed &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit other &&\n+\n+\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\n+\t\t\tcruft_a=\"$(echo $keep | git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack)\" &&\n+\t\t\tgit prune-packed &&\n+\t\t\tcruft_b=\"$(git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-pack-$cruft_a.pack\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft_a.mtimes\" >expect.raw &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft_b.mtimes\" >actual.raw &&\n+\n+\t\t\tsort <expect.raw >expect &&\n+\t\t\tsort <actual.raw >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"multiple cruft packs (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\tgit repack -Ad &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\n+\t\t\ttest_commit cruft &&\n+\t\t\tloose=\"$objdir/$(test_oid_to_path $(git rev-parse cruft))\" &&\n+\n+\t\t\t# generate three copies of the cruft object in different\n+\t\t\t# cruft packs, each with a unique mtime:\n+\t\t\t#   - one expired (1000 seconds ago)\n+\t\t\t#   - two non-expired (one 1000 seconds in the future,\n+\t\t\t#     one 1500 seconds in the future)\n+\t\t\ttest-tool chmtime =-1000 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-A <<-EOF &&\n+\t\t\t$keep\n+\t\t\tEOF\n+\t\t\ttest-tool chmtime =+1000 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-B <<-EOF &&\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\tEOF\n+\t\t\ttest-tool chmtime =+1500 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-C <<-EOF &&\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-B-*.pack))\n+\t\t\tEOF\n+\n+\t\t\t# ensure the resulting cruft pack takes the most recent\n+\t\t\t# mtime among all copies\n+\t\t\tcruft=\"$(git pack-objects --cruft \\\n+\t\t\t\t--cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-B-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-C-*.pack))\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\n+\t\t\ttest-tool pack-mtimes \"$(basename $(ls $packdir/pack-C-*.mtimes))\" >expect.raw &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\t\tsort expect.raw >expect &&\n+\t\t\tsort actual.raw >actual &&\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"cruft packs tolerate missing trees (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\ttest_commit cruft &&\n+\n+\t\t\ttree=\"$(git rev-parse cruft^{tree})\" &&\n+\n+\t\t\tgit reset --hard reachable &&\n+\t\t\tgit tag -d cruft &&\n+\t\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t\t# remove the unreachable tree, but leave the commit\n+\t\t\t# which has it as its root tree intact\n+\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$tree\")\" &&\n+\n+\t\t\tgit repack -Ad &&\n+\t\t\tbasename $(ls $packdir/pack-*.pack) >in &&\n+\t\t\tgit pack-objects --cruft --cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <in\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"cruft packs tolerate missing blobs (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\ttest_commit cruft &&\n+\n+\t\t\tblob=\"$(git rev-parse cruft:cruft.t)\" &&\n+\n+\t\t\tgit reset --hard reachable &&\n+\t\t\tgit tag -d cruft &&\n+\t\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t\t# remove the unreachable blob, but leave the commit (and\n+\t\t\t# the root tree of that commit) intact\n+\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$blob\")\" &&\n+\n+\t\t\tgit repack -Ad &&\n+\t\t\tbasename $(ls $packdir/pack-*.pack) >in &&\n+\t\t\tgit pack-objects --cruft --cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <in\n+\t\t)\n+\t'\n+}\n+\n+basic_cruft_pack_tests never\n+\n+test_done\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450211","messageId":"cebb30b6678497529f949d539e9d8a2910662067.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 09/17] reachable: add options to add_unseen_recent_objects_to_traversal","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:21:05Z","receivedAt":"2022-03-03T00:21:19Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This function behaves very similarly to what we will need in\npack-objects in order to implement cruft packs with expiration. But it\nis lacking a couple of things. Namely, it needs:\n\n  - a mechanism to communicate the timestamps of individual recent\n    objects to some external caller\n\n  - and, in the case of packed objects, our future caller will also want\n    to know the originating pack, as well as the offset within that pack\n    at which the object can be found\n\n  - finally, it needs a way to skip over packs which are marked as kept\n    in-core.\n\nTo address the first two, add a callback interface in this patch which\nreports the time of each recent object, as well as a (packed_git,\noff_t) pair for packed objects.\n\nLikewise, add a new option to the packed object iterators to skip over\npacks which are marked as kept in core. This option will become\nimplicitly tested in a future patch.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c |  2 +-\n reachable.c            | 51 +++++++++++++++++++++++++++++++++++-------\n reachable.h            |  9 +++++++-\n 3 files changed, 52 insertions(+), 10 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 5ba4fc9c2c..1ef333717d 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -3951,7 +3951,7 @@ static void get_object_list(int ac, const char **av)\n \tif (unpack_unreachable_expiration) {\n \t\trevs.ignore_missing_links = 1;\n \t\tif (add_unseen_recent_objects_to_traversal(&revs,\n-\t\t\t\tunpack_unreachable_expiration))\n+\t\t\t\tunpack_unreachable_expiration, NULL, 0))\n \t\t\tdie(_(\"unable to add recent objects\"));\n \t\tif (prepare_revision_walk(&revs))\n \t\t\tdie(_(\"revision walk setup failed\"));\ndiff --git a/reachable.c b/reachable.c\nindex 84e3d0d75e..0eb9909f47 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -60,9 +60,13 @@ static void mark_commit(struct commit *c, void *data)\n struct recent_data {\n \tstruct rev_info *revs;\n \ttimestamp_t timestamp;\n+\treport_recent_object_fn *cb;\n+\tint ignore_in_core_kept_packs;\n };\n \n static void add_recent_object(const struct object_id *oid,\n+\t\t\t      struct packed_git *pack,\n+\t\t\t      off_t offset,\n \t\t\t      timestamp_t mtime,\n \t\t\t      struct recent_data *data)\n {\n@@ -103,13 +107,29 @@ static void add_recent_object(const struct object_id *oid,\n \t\tdie(\"unable to lookup %s\", oid_to_hex(oid));\n \n \tadd_pending_object(data->revs, obj, \"\");\n+\tif (data->cb)\n+\t\tdata->cb(obj, pack, offset, mtime);\n+}\n+\n+static int want_recent_object(struct recent_data *data,\n+\t\t\t      const struct object_id *oid)\n+{\n+\tif (data->ignore_in_core_kept_packs &&\n+\t    has_object_kept_pack(oid, IN_CORE_KEEP_PACKS))\n+\t\treturn 0;\n+\treturn 1;\n }\n \n static int add_recent_loose(const struct object_id *oid,\n \t\t\t    const char *path, void *data)\n {\n \tstruct stat st;\n-\tstruct object *obj = lookup_object(the_repository, oid);\n+\tstruct object *obj;\n+\n+\tif (!want_recent_object(data, oid))\n+\t\treturn 0;\n+\n+\tobj = lookup_object(the_repository, oid);\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n@@ -126,7 +146,7 @@ static int add_recent_loose(const struct object_id *oid,\n \t\treturn error_errno(\"unable to stat %s\", oid_to_hex(oid));\n \t}\n \n-\tadd_recent_object(oid, st.st_mtime, data);\n+\tadd_recent_object(oid, NULL, 0, st.st_mtime, data);\n \treturn 0;\n }\n \n@@ -134,29 +154,43 @@ static int add_recent_packed(const struct object_id *oid,\n \t\t\t     struct packed_git *p, uint32_t pos,\n \t\t\t     void *data)\n {\n-\tstruct object *obj = lookup_object(the_repository, oid);\n+\tstruct object *obj;\n+\n+\tif (!want_recent_object(data, oid))\n+\t\treturn 0;\n+\n+\tobj = lookup_object(the_repository, oid);\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n-\tadd_recent_object(oid, p->mtime, data);\n+\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), p->mtime, data);\n \treturn 0;\n }\n \n int add_unseen_recent_objects_to_traversal(struct rev_info *revs,\n-\t\t\t\t\t   timestamp_t timestamp)\n+\t\t\t\t\t   timestamp_t timestamp,\n+\t\t\t\t\t   report_recent_object_fn *cb,\n+\t\t\t\t\t   int ignore_in_core_kept_packs)\n {\n \tstruct recent_data data;\n+\tenum for_each_object_flags flags;\n \tint r;\n \n \tdata.revs = revs;\n \tdata.timestamp = timestamp;\n+\tdata.cb = cb;\n+\tdata.ignore_in_core_kept_packs = ignore_in_core_kept_packs;\n \n \tr = for_each_loose_object(add_recent_loose, &data,\n \t\t\t\t  FOR_EACH_OBJECT_LOCAL_ONLY);\n \tif (r)\n \t\treturn r;\n-\treturn for_each_packed_object(add_recent_packed, &data,\n-\t\t\t\t      FOR_EACH_OBJECT_LOCAL_ONLY);\n+\n+\tflags = FOR_EACH_OBJECT_LOCAL_ONLY | FOR_EACH_OBJECT_PACK_ORDER;\n+\tif (ignore_in_core_kept_packs)\n+\t\tflags |= FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS;\n+\n+\treturn for_each_packed_object(add_recent_packed, &data, flags);\n }\n \n static int mark_object_seen(const struct object_id *oid,\n@@ -217,7 +251,8 @@ void mark_reachable_objects(struct rev_info *revs, int mark_reflog,\n \n \tif (mark_recent) {\n \t\trevs->ignore_missing_links = 1;\n-\t\tif (add_unseen_recent_objects_to_traversal(revs, mark_recent))\n+\t\tif (add_unseen_recent_objects_to_traversal(revs, mark_recent,\n+\t\t\t\t\t\t\t   NULL, 0))\n \t\t\tdie(\"unable to mark recent objects\");\n \t\tif (prepare_revision_walk(revs))\n \t\t\tdie(\"revision walk setup failed\");\ndiff --git a/reachable.h b/reachable.h\nindex 5df932ad8f..b776761baa 100644\n--- a/reachable.h\n+++ b/reachable.h\n@@ -1,11 +1,18 @@\n #ifndef REACHEABLE_H\n #define REACHEABLE_H\n \n+#include \"object.h\"\n+\n struct progress;\n struct rev_info;\n \n+typedef void report_recent_object_fn(const struct object *, struct packed_git *,\n+\t\t\t\t     off_t, time_t);\n+\n int add_unseen_recent_objects_to_traversal(struct rev_info *revs,\n-\t\t\t\t\t   timestamp_t timestamp);\n+\t\t\t\t\t   timestamp_t timestamp,\n+\t\t\t\t\t   report_recent_object_fn cb,\n+\t\t\t\t\t   int ignore_in_core_kept_packs);\n void mark_reachable_objects(struct rev_info *revs, int mark_reflog,\n \t\t\t    timestamp_t mark_recent, struct progress *);\n \n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450212","messageId":"fa4de8859d1c067d5773d023b1f434031abacc3a.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 10/17] reachable: report precise timestamps from objects in cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:21:07Z","receivedAt":"2022-03-03T00:21:27Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When generating a cruft pack, the caller within pack-objects will want\nto know the precise timestamps of cruft objects (i.e., their\ncorresponding values in the .mtimes table) rather than the mtime of the\ncruft pack itself.\n\nTeach add_recent_packed() to lookup each object's precise mtime from the\n.mtimes file if one exists (indicated by the is_cruft bit on the\npacked_git structure).\n\nA couple of small things worth noting here:\n\n  - load_pack_mtimes() needs to be called before asking for\n    nth_packed_mtime(), and that call is done lazily here. That function\n    exits early if the .mtimes file has already been opened and parsed,\n    so only the first call is slow.\n\n  - Checking the is_cruft bit can be done without any extra work on the\n    caller's behalf, since it is set up for us automatically as a\n    side-effect of calling add_packed_git() (just like the 'pack_keep'\n    and 'pack_promisor' bits).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n reachable.c | 9 ++++++++-\n 1 file changed, 8 insertions(+), 1 deletion(-)\n\ndiff --git a/reachable.c b/reachable.c\nindex 0eb9909f47..9ec8e6bd5b 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -13,6 +13,7 @@\n #include \"worktree.h\"\n #include \"object-store.h\"\n #include \"pack-bitmap.h\"\n+#include \"pack-mtimes.h\"\n \n struct connectivity_progress {\n \tstruct progress *progress;\n@@ -155,6 +156,7 @@ static int add_recent_packed(const struct object_id *oid,\n \t\t\t     void *data)\n {\n \tstruct object *obj;\n+\ttimestamp_t mtime = p->mtime;\n \n \tif (!want_recent_object(data, oid))\n \t\treturn 0;\n@@ -163,7 +165,12 @@ static int add_recent_packed(const struct object_id *oid,\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n-\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), p->mtime, data);\n+\tif (p->is_cruft) {\n+\t\tif (load_pack_mtimes(p) < 0)\n+\t\t\tdie(_(\"could not load cruft pack .mtimes\"));\n+\t\tmtime = nth_packed_mtime(p, pos);\n+\t}\n+\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), mtime, data);\n \treturn 0;\n }\n \n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450213","messageId":"92318f870097a9c164896043ada24ba819e296d1.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 11/17] builtin/pack-objects.c: --cruft with expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:21:10Z","receivedAt":"2022-03-03T00:21:29Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a previous patch, pack-objects learned how to generate a cruft pack\nso long as no objects are dropped.\n\nThis patch teaches pack-objects to handle the case where a non-never\n`--cruft-expiration` value is passed. This case is slightly more\ncomplicated than before, because we want pack-objects to save\nunreachable objects which would have been pruned when there is another\nrecent (i.e., non-prunable) unreachable object which reaches the other.\nWe'll call these objects \"unreachable but reachable-from-recent\".\n\nHere is how pack-objects handles `--cruft-expiration`:\n\n  - Instead of adding all objects outside of the kept pack(s) into the\n    packing list, only handle the ones whose mtime is within the grace\n    period.\n\n  - Construct a reachability traversal whose tips are the\n    unreachable-but-recent objects.\n\n  - Then, walk along that traversal, stopping if we reach an object in\n    the kept pack. At each step along the traversal, we add the object\n    we are visiting to the packing list.\n\nIn the majority of these cases, any object we visit in this traversal\nwill already be in our packing list. But we will sometimes encounter\nreachable-from-recent cruft objects, which we want to retain even if\nthey aged out of the grace period.\n\nThe most subtle point of this process is that we actually don't need to\nbother to update the rescued object's mtime. Even though we will write\nan .mtimes file with a value that is older than the expiration window,\nit will continue to survive cruft repacks so long as any objects which\nreach it haven't aged out.\n\nThat is, a future repack will also exclude that object from the initial\npacking list, only to discover it later on when doing the reachability\ntraversal.\n\nFinally, stopping early once an object is found in a kept pack is safe\nto do because the kept packs ordinarily represent which packs will\nsurvive after repacking. Assuming that it _isn't_ safe to halt a\ntraversal early would mean that there is some ancestor object which is\nmissing, which implies repository corruption (i.e., the complete set of\nreachable objects isn't present).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c        |  84 +++++++++++++++++++-\n reachable.h                   |   4 +-\n t/t5329-pack-objects-cruft.sh | 143 ++++++++++++++++++++++++++++++++++\n 3 files changed, 228 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 1ef333717d..fcac0b5c91 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -3439,6 +3439,44 @@ static void add_cruft_object_entry(const struct object_id *oid, enum object_type\n \treturn;\n }\n \n+static void show_cruft_object(struct object *obj, const char *name, void *data)\n+{\n+\t/*\n+\t * if we did not record it earlier, it's at least as old as our\n+\t * expiration value. Rather than find it exactly, just use that\n+\t * value.  This may bump it forward from its real mtime, but it\n+\t * will still be \"too old\" next time we run with the same\n+\t * expiration.\n+\t *\n+\t * if obj does appear in the packing list, this call is a noop (or may\n+\t * set the namehash).\n+\t */\n+\tadd_cruft_object_entry(&obj->oid, obj->type, NULL, 0, name, cruft_expiration);\n+}\n+\n+static void show_cruft_commit(struct commit *commit, void *data)\n+{\n+\tshow_cruft_object((struct object*)commit, NULL, data);\n+}\n+\n+static int cruft_include_check_obj(struct object *obj, void *data)\n+{\n+\treturn !has_object_kept_pack(&obj->oid, IN_CORE_KEEP_PACKS);\n+}\n+\n+static int cruft_include_check(struct commit *commit, void *data)\n+{\n+\treturn cruft_include_check_obj((struct object*)commit, data);\n+}\n+\n+static void set_cruft_mtime(const struct object *object,\n+\t\t\t    struct packed_git *pack,\n+\t\t\t    off_t offset, time_t mtime)\n+{\n+\tadd_cruft_object_entry(&object->oid, object->type, pack, offset, NULL,\n+\t\t\t       mtime);\n+}\n+\n static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n {\n \tstruct string_list_item *item = NULL;\n@@ -3464,6 +3502,50 @@ static void enumerate_cruft_objects(void)\n \tstop_progress(&progress_state);\n }\n \n+static void enumerate_and_traverse_cruft_objects(struct string_list *fresh_packs)\n+{\n+\tstruct packed_git *p;\n+\tstruct rev_info revs;\n+\tint ret;\n+\n+\trepo_init_revisions(the_repository, &revs, NULL);\n+\n+\trevs.tag_objects = 1;\n+\trevs.tree_objects = 1;\n+\trevs.blob_objects = 1;\n+\n+\trevs.include_check = cruft_include_check;\n+\trevs.include_check_obj = cruft_include_check_obj;\n+\n+\trevs.ignore_missing_links = 1;\n+\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Enumerating cruft objects\"), 0);\n+\tret = add_unseen_recent_objects_to_traversal(&revs, cruft_expiration,\n+\t\t\t\t\t\t     set_cruft_mtime, 1);\n+\tstop_progress(&progress_state);\n+\n+\tif (ret)\n+\t\tdie(_(\"unable to add cruft objects\"));\n+\n+\t/*\n+\t * Re-mark only the fresh packs as kept so that objects in\n+\t * unknown packs do not halt the reachability traversal early.\n+\t */\n+\tfor (p = get_all_packs(the_repository); p; p = p->next)\n+\t\tp->pack_keep_in_core = 0;\n+\tmark_pack_kept_in_core(fresh_packs, 1);\n+\n+\tif (prepare_revision_walk(&revs))\n+\t\tdie(_(\"revision walk setup failed\"));\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Traversing cruft objects\"), 0);\n+\tnr_seen = 0;\n+\ttraverse_commit_list(&revs, show_cruft_commit, show_cruft_object, NULL);\n+\n+\tstop_progress(&progress_state);\n+}\n+\n static void read_cruft_objects(void)\n {\n \tstruct strbuf buf = STRBUF_INIT;\n@@ -3515,7 +3597,7 @@ static void read_cruft_objects(void)\n \tmark_pack_kept_in_core(&discard_packs, 0);\n \n \tif (cruft_expiration)\n-\t\tdie(\"--cruft-expiration not yet implemented\");\n+\t\tenumerate_and_traverse_cruft_objects(&fresh_packs);\n \telse\n \t\tenumerate_cruft_objects();\n \ndiff --git a/reachable.h b/reachable.h\nindex b776761baa..020a887b99 100644\n--- a/reachable.h\n+++ b/reachable.h\n@@ -1,10 +1,10 @@\n #ifndef REACHEABLE_H\n #define REACHEABLE_H\n \n-#include \"object.h\"\n-\n struct progress;\n struct rev_info;\n+struct object;\n+struct packed_git;\n \n typedef void report_recent_object_fn(const struct object *, struct packed_git *,\n \t\t\t\t     off_t, time_t);\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 003ca7344e..939cdc297a 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -214,5 +214,148 @@ basic_cruft_pack_tests () {\n }\n \n basic_cruft_pack_tests never\n+basic_cruft_pack_tests 2.weeks.ago\n+\n+test_expect_success 'cruft tags rescue tagged objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit tagged &&\n+\t\tgit tag -a annotated -m tag &&\n+\n+\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\twhile read oid\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $oid)\"\n+\t\tdone <objects &&\n+\n+\t\ttest-tool chmtime -500 \\\n+\t\t\t\"$objdir/$(test_oid_to_path $(git rev-parse annotated))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\t\tcut -f1 -d\" \" <actual.raw | sort >actual &&\n+\n+\t\t(\n+\t\t\tcat objects &&\n+\t\t\tgit rev-parse annotated\n+\t\t) >expect.raw &&\n+\t\tsort <expect.raw >expect &&\n+\n+\t\ttest_cmp expect actual &&\n+\t\tcat actual\n+\t)\n+'\n+\n+test_expect_success 'cruft commits rescue parents, trees' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit old &&\n+\t\ttest_commit new &&\n+\n+\t\tgit rev-list --objects --no-object-names packed..new >objects &&\n+\t\twhile read object\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $object)\"\n+\t\tdone <objects &&\n+\t\ttest-tool chmtime +500 \"$objdir/$(test_oid_to_path \\\n+\t\t\t$(git rev-parse HEAD))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\tcut -d\" \" -f1 <actual.raw | sort >actual &&\n+\t\tsort <objects >expect &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'cruft trees rescue sub-trees, blobs' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\tmkdir -p dir/sub &&\n+\t\techo foo >foo &&\n+\t\techo bar >dir/bar &&\n+\t\techo baz >dir/sub/baz &&\n+\n+\t\ttest_tick &&\n+\t\tgit add . &&\n+\t\tgit commit -m \"pruned\" &&\n+\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD^{tree}))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:foo))\" &&\n+\t\ttest-tool chmtime  -500 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/bar))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/sub))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/sub/baz))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\t\tcut -f1 -d\" \" <actual.raw | sort >actual &&\n+\n+\t\tgit rev-parse HEAD:dir HEAD:dir/bar HEAD:dir/sub HEAD:dir/sub/baz >expect.raw &&\n+\t\tsort <expect.raw >expect &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'expired objects are pruned' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit pruned &&\n+\n+\t\tgit rev-list --objects --no-object-names packed..pruned >objects &&\n+\t\twhile read object\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $object)\"\n+\t\tdone <objects &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual &&\n+\t\ttest_must_be_empty actual\n+\t)\n+'\n \n test_done\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450214","messageId":"1e94b33cb4d9f3e1f29ef754a9cc09898779836a.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 12/17] builtin/repack.c: support generating a cruft pack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:21:12Z","receivedAt":"2022-03-03T00:21:33Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Expose a way to split the contents of a repository into a main and cruft\npack when doing an all-into-one repack with `git repack --cruft -d`, and\na complementary configuration variable.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-repack.txt            |  11 ++\n Documentation/technical/cruft-packs.txt |   2 +-\n builtin/repack.c                        | 106 +++++++++++-\n t/t5329-pack-objects-cruft.sh           | 207 ++++++++++++++++++++++++\n 4 files changed, 320 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/git-repack.txt b/Documentation/git-repack.txt\nindex ee30edc178..0bf13893d8 100644\n--- a/Documentation/git-repack.txt\n+++ b/Documentation/git-repack.txt\n@@ -63,6 +63,17 @@ to the new separate pack will be written.\n \tAlso run  'git prune-packed' to remove redundant\n \tloose object files.\n \n+--cruft::\n+\tSame as `-a`, unless `-d` is used. Then any unreachable objects\n+\tare packed into a separate cruft pack. Unreachable objects can\n+\tbe pruned using the normal expiry rules with the next `git gc`\n+\tinvocation (see linkgit:git-gc[1]). Incompatible with `-k`.\n+\n+--cruft-expiration=<approxidate>::\n+\tExpire unreachable objects older than `<approxidate>`\n+\timmediately instead of waiting for the next `git gc` invocation.\n+\tOnly useful with `--cruft -d`.\n+\n -l::\n \tPass the `--local` option to 'git pack-objects'. See\n \tlinkgit:git-pack-objects[1].\ndiff --git a/Documentation/technical/cruft-packs.txt b/Documentation/technical/cruft-packs.txt\nindex 2c3c5d93f8..f80e975a47 100644\n--- a/Documentation/technical/cruft-packs.txt\n+++ b/Documentation/technical/cruft-packs.txt\n@@ -17,7 +17,7 @@ pruned according to normal expiry rules with the next 'git gc' invocation.\n \n Unreachable objects aren't removed immediately, since doing so could race with\n an incoming push which may reference an object which is about to be deleted.\n-Instead, those unreachable objects are stored as loose object and stay that way\n+Instead, those unreachable objects are stored as loose objects and stay that way\n until they are older than the expiration window, at which point they are removed\n by linkgit:git-prune[1].\n \ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex f908f7d5dd..f7fb88bcf1 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -18,11 +18,17 @@\n #include \"pack-bitmap.h\"\n #include \"refs.h\"\n \n+#define ALL_INTO_ONE 1\n+#define LOOSEN_UNREACHABLE 2\n+#define PACK_CRUFT 4\n+\n+static int pack_everything;\n static int delta_base_offset = 1;\n static int pack_kept_objects = -1;\n static int write_bitmaps = -1;\n static int use_delta_islands;\n static char *packdir, *packtmp_name, *packtmp;\n+static char *cruft_expiration;\n \n static const char *const git_repack_usage[] = {\n \tN_(\"git repack [<options>]\"),\n@@ -54,6 +60,7 @@ static int repack_config(const char *var, const char *value, void *cb)\n \t\tuse_delta_islands = git_config_bool(var, value);\n \t\treturn 0;\n \t}\n+\n \treturn git_default_config(var, value, cb);\n }\n \n@@ -300,9 +307,6 @@ static void repack_promisor_objects(const struct pack_objects_args *args,\n \t\tdie(_(\"could not finish pack-objects to repack promisor objects\"));\n }\n \n-#define ALL_INTO_ONE 1\n-#define LOOSEN_UNREACHABLE 2\n-\n struct pack_geometry {\n \tstruct packed_git **pack;\n \tuint32_t pack_nr, pack_alloc;\n@@ -339,6 +343,8 @@ static void init_pack_geometry(struct pack_geometry **geometry_p)\n \tfor (p = get_all_packs(the_repository); p; p = p->next) {\n \t\tif (!pack_kept_objects && p->pack_keep)\n \t\t\tcontinue;\n+\t\tif (p->is_cruft)\n+\t\t\tcontinue;\n \n \t\tALLOC_GROW(geometry->pack,\n \t\t\t   geometry->pack_nr + 1,\n@@ -600,6 +606,67 @@ static int write_midx_included_packs(struct string_list *include,\n \treturn finish_command(&cmd);\n }\n \n+static int write_cruft_pack(const struct pack_objects_args *args,\n+\t\t\t    const char *pack_prefix,\n+\t\t\t    struct string_list *names,\n+\t\t\t    struct string_list *existing_packs,\n+\t\t\t    struct string_list *existing_kept_packs)\n+{\n+\tstruct child_process cmd = CHILD_PROCESS_INIT;\n+\tstruct strbuf line = STRBUF_INIT;\n+\tstruct string_list_item *item;\n+\tFILE *in, *out;\n+\tint ret;\n+\n+\tprepare_pack_objects(&cmd, args);\n+\n+\tstrvec_push(&cmd.args, \"--cruft\");\n+\tif (cruft_expiration)\n+\t\tstrvec_pushf(&cmd.args, \"--cruft-expiration=%s\",\n+\t\t\t     cruft_expiration);\n+\n+\tstrvec_push(&cmd.args, \"--honor-pack-keep\");\n+\tstrvec_push(&cmd.args, \"--non-empty\");\n+\tstrvec_push(&cmd.args, \"--max-pack-size=0\");\n+\n+\tcmd.in = -1;\n+\n+\tret = start_command(&cmd);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\t/*\n+\t * names has a confusing double use: it both provides the list\n+\t * of just-written new packs, and accepts the name of the cruft\n+\t * pack we are writing.\n+\t *\n+\t * By the time it is read here, it contains only the pack(s)\n+\t * that were just written, which is exactly the set of packs we\n+\t * want to consider kept.\n+\t */\n+\tin = xfdopen(cmd.in, \"w\");\n+\tfor_each_string_list_item(item, names)\n+\t\tfprintf(in, \"%s-%s.pack\\n\", pack_prefix, item->string);\n+\tfor_each_string_list_item(item, existing_packs)\n+\t\tfprintf(in, \"-%s.pack\\n\", item->string);\n+\tfor_each_string_list_item(item, existing_kept_packs)\n+\t\tfprintf(in, \"%s.pack\\n\", item->string);\n+\tfclose(in);\n+\n+\tout = xfdopen(cmd.out, \"r\");\n+\twhile (strbuf_getline_lf(&line, out) != EOF) {\n+\t\tif (line.len != the_hash_algo->hexsz)\n+\t\t\tdie(_(\"repack: Expecting full hex object ID lines only \"\n+\t\t\t      \"from pack-objects.\"));\n+\t\tstring_list_append(names, line.buf);\n+\t}\n+\tfclose(out);\n+\n+\tstrbuf_release(&line);\n+\n+\treturn finish_command(&cmd);\n+}\n+\n int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n \tstruct child_process cmd = CHILD_PROCESS_INIT;\n@@ -616,7 +683,6 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tint show_progress;\n \n \t/* variables to be filled by option parsing */\n-\tint pack_everything = 0;\n \tint delete_redundant = 0;\n \tconst char *unpack_unreachable = NULL;\n \tint keep_unreachable = 0;\n@@ -632,6 +698,11 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tOPT_BIT('A', NULL, &pack_everything,\n \t\t\t\tN_(\"same as -a, and turn unreachable objects loose\"),\n \t\t\t\t   LOOSEN_UNREACHABLE | ALL_INTO_ONE),\n+\t\tOPT_BIT(0, \"cruft\", &pack_everything,\n+\t\t\t\tN_(\"same as -a, pack unreachable cruft objects separately\"),\n+\t\t\t\t   PACK_CRUFT),\n+\t\tOPT_STRING(0, \"cruft-expiration\", &cruft_expiration, N_(\"approxidate\"),\n+\t\t\t\tN_(\"with -C, expire objects older than this\")),\n \t\tOPT_BOOL('d', NULL, &delete_redundant,\n \t\t\t\tN_(\"remove redundant packs, and run git-prune-packed\")),\n \t\tOPT_BOOL('f', NULL, &po_args.no_reuse_delta,\n@@ -684,6 +755,15 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t    (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE)))\n \t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--keep-unreachable\", \"-A\");\n \n+\tif (pack_everything & PACK_CRUFT) {\n+\t\tpack_everything |= ALL_INTO_ONE;\n+\n+\t\tif (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE))\n+\t\t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--cruft\", \"-A\");\n+\t\tif (keep_unreachable)\n+\t\t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--cruft\", \"-k\");\n+\t}\n+\n \tif (write_bitmaps < 0) {\n \t\tif (!write_midx &&\n \t\t    (!(pack_everything & ALL_INTO_ONE) || !is_bare_repository()))\n@@ -767,7 +847,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (pack_everything & ALL_INTO_ONE) {\n \t\trepack_promisor_objects(&po_args, &names);\n \n-\t\tif (existing_nonkept_packs.nr && delete_redundant) {\n+\t\tif (existing_nonkept_packs.nr && delete_redundant &&\n+\t\t    !(pack_everything & PACK_CRUFT)) {\n \t\t\tfor_each_string_list_item(item, &names) {\n \t\t\t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s-%s.pack\",\n \t\t\t\t\t     packtmp_name, item->string);\n@@ -829,6 +910,21 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (!names.nr && !po_args.quiet)\n \t\tprintf_ln(_(\"Nothing new to pack.\"));\n \n+\tif (pack_everything & PACK_CRUFT) {\n+\t\tconst char *pack_prefix;\n+\t\tif (!skip_prefix(packtmp, packdir, &pack_prefix))\n+\t\t\tdie(_(\"pack prefix %s does not begin with objdir %s\"),\n+\t\t\t    packtmp, packdir);\n+\t\tif (*pack_prefix == '/')\n+\t\t\tpack_prefix++;\n+\n+\t\tret = write_cruft_pack(&po_args, pack_prefix, &names,\n+\t\t\t\t       &existing_nonkept_packs,\n+\t\t\t\t       &existing_kept_packs);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t}\n+\n \tfor_each_string_list_item(item, &names) {\n \t\titem->util = (void *)(uintptr_t)populate_pack_exts(item->string);\n \t}\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 939cdc297a..06c550c958 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -358,4 +358,211 @@ test_expect_success 'expired objects are pruned' '\n \t)\n '\n \n+test_expect_success 'repack --cruft generates a cruft pack' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit branch -M main &&\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit unreachable &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d unreachable &&\n+\t\t# objects are not cruft if they are contained in the reflogs\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\t\tsort <reachable.raw >reachable &&\n+\t\tcomm -13 reachable objects >unreachable &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tcruft=$(basename $(ls $packdir/pack-*.mtimes) .mtimes) &&\n+\t\tpack=$(basename $(ls $packdir/pack-*.pack | grep -v $cruft) .pack) &&\n+\n+\t\tgit show-index <$packdir/$pack.idx >actual.raw &&\n+\t\tcut -f2 -d\" \" actual.raw | sort >actual &&\n+\t\ttest_cmp reachable actual &&\n+\n+\t\tgit show-index <$packdir/$cruft.idx >actual.raw &&\n+\t\tcut -f2 -d\" \" actual.raw | sort >actual &&\n+\t\ttest_cmp unreachable actual\n+\t)\n+'\n+\n+test_expect_success 'loose objects mtimes upsert others' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit cruft &&\n+\t\t# incremental repack, leaving existing objects loose (so\n+\t\t# they can be \"freshened\")\n+\t\tgit repack &&\n+\n+\t\ttip=\"$(git rev-parse cruft)\" &&\n+\t\tpath=\"$objdir/$(test_oid_to_path \"$(git rev-parse cruft)\")\" &&\n+\t\ttest-tool chmtime --get +1000 \"$path\" >expect &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tmtimes=\"$(basename $(ls $packdir/pack-*.mtimes))\" &&\n+\t\ttest-tool pack-mtimes \"$mtimes\" >actual.raw &&\n+\t\tgrep \"$tip\" actual.raw | cut -d\" \" -f2 >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'cruft packs are not included in geometric repack' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit cruft &&\n+\t\tgit repack -d &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft &&\n+\n+\t\tfind $packdir -type f | sort >before &&\n+\t\tgit repack --geometric=2 -d &&\n+\t\tfind $packdir -type f | sort >after &&\n+\n+\t\ttest_cmp before after\n+\t)\n+'\n+\n+test_expect_success 'repack --geometric collects once-cruft objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\tgit rm -rf . &&\n+\t\ttest_commit --no-tag cruft &&\n+\t\tcruft=\"$(git rev-parse HEAD)\" &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t# Pack the objects created in the previous step into a cruft\n+\t\t# pack. Intentionally leave loose copies of those objects\n+\t\t# around so we can pick them up in a subsequent --geometric\n+\t\t# reapack.\n+\t\tgit repack --cruft &&\n+\n+\t\t# Now make those objects reachable, and ensure that they are\n+\t\t# packed into the new pack created via a --geometric repack.\n+\t\tgit update-ref refs/heads/other $cruft &&\n+\n+\t\t# Without this object, the set of unpacked objects is exactly\n+\t\t# the set of objects already in the cruft pack. Tweak that set\n+\t\t# to ensure we do not overwrite the cruft pack entirely.\n+\t\ttest_commit reachable2 &&\n+\n+\t\tfind $packdir -name \"pack-*.idx\" | sort >before &&\n+\t\tgit repack --geometric=2 -d &&\n+\t\tfind $packdir -name \"pack-*.idx\" | sort >after &&\n+\n+\t\t{\n+\t\t\tgit rev-list --objects --no-object-names $cruft &&\n+\t\t\tgit rev-list --objects --no-object-names reachable..reachable2\n+\t\t} >want.raw &&\n+\t\tsort want.raw >want &&\n+\n+\t\tpack=$(comm -13 before after) &&\n+\t\tgit show-index <$pack >objects.raw &&\n+\n+\t\tcut -d\" \" -f2 objects.raw | sort >got &&\n+\n+\t\ttest_cmp want got\n+\t)\n+'\n+\n+test_expect_success 'cruft repack with no reachable objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\t\tgit repack -ad &&\n+\n+\t\tbase=\"$(git rev-parse base)\" &&\n+\n+\t\tgit for-each-ref --format=\"delete %(refname)\" >in &&\n+\t\tgit update-ref --stdin <in &&\n+\t\tgit reflog expire --all --expire=all &&\n+\t\trm -fr .git/index &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tgit cat-file -t $base\n+\t)\n+'\n+\n+test_expect_success 'cruft repack ignores --max-pack-size' '\n+\tgit init max-pack-size &&\n+\t(\n+\t\tcd max-pack-size &&\n+\t\ttest_commit base &&\n+\t\t# two cruft objects which exceed the maximum pack size\n+\t\ttest-tool genrandom foo 1048576 | git hash-object --stdin -w &&\n+\t\ttest-tool genrandom bar 1048576 | git hash-object --stdin -w &&\n+\t\tgit repack --cruft --max-pack-size=1M &&\n+\t\tfind $packdir -name \"*.mtimes\" >cruft &&\n+\t\ttest_line_count = 1 cruft &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(cat cruft)\")\" >objects &&\n+\t\ttest_line_count = 2 objects\n+\t)\n+'\n+\n+test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n+\t(\n+\t\tcd max-pack-size &&\n+\t\t# repack everything back together to remove the existing cruft\n+\t\t# pack (but to keep its objects)\n+\t\tgit repack -adk &&\n+\t\tgit -c pack.packSizeLimit=1M repack --cruft &&\n+\t\t# ensure the same post condition is met when --max-pack-size\n+\t\t# would otherwise be inferred from the configuration\n+\t\tfind $packdir -name \"*.mtimes\" >cruft &&\n+\t\ttest_line_count = 1 cruft &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(cat cruft)\")\" >objects &&\n+\t\ttest_line_count = 2 objects\n+\t)\n+'\n+\n test_done\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450215","messageId":"9cfcd123bd107357bf36652976cad16a56c9e366.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 13/17] builtin/repack.c: allow configuring cruft pack generation","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:21:15Z","receivedAt":"2022-03-03T00:21:34Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In servers which set the pack.window configuration to a large value, we\ncan wind up spending quite a lot of time finding new bases when breaking\ndelta chains between reachable and unreachable objects while generating\na cruft pack.\n\nIntroduce a handful of `repack.cruft*` configuration variables to\ncontrol the parameters used by pack-objects when generating a cruft\npack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/repack.txt |  9 ++++\n builtin/repack.c                | 50 ++++++++++++++------\n t/t5329-pack-objects-cruft.sh   | 83 +++++++++++++++++++++++++++++++++\n 3 files changed, 128 insertions(+), 14 deletions(-)\n\ndiff --git a/Documentation/config/repack.txt b/Documentation/config/repack.txt\nindex 9c413e177e..fd18d1fb89 100644\n--- a/Documentation/config/repack.txt\n+++ b/Documentation/config/repack.txt\n@@ -25,3 +25,12 @@ repack.writeBitmaps::\n \tspace and extra time spent on the initial repack.  This has\n \tno effect if multiple packfiles are created.\n \tDefaults to true on bare repos, false otherwise.\n+\n+repack.cruftWindow::\n+repack.cruftWindowMemory::\n+repack.cruftDepth::\n+repack.cruftThreads::\n+\tParameters used by linkgit:git-pack-objects[1] when generating\n+\ta cruft pack and the respective parameters are not given over\n+\tthe command line. See similarly named `pack.*` configuration\n+\tvariables for defaults and meaning.\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex f7fb88bcf1..d61c78e94e 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -40,9 +40,21 @@ static const char incremental_bitmap_conflict_error[] = N_(\n \"--no-write-bitmap-index or disable the pack.writebitmaps configuration.\"\n );\n \n+struct pack_objects_args {\n+\tconst char *window;\n+\tconst char *window_memory;\n+\tconst char *depth;\n+\tconst char *threads;\n+\tconst char *max_pack_size;\n+\tint no_reuse_delta;\n+\tint no_reuse_object;\n+\tint quiet;\n+\tint local;\n+};\n \n static int repack_config(const char *var, const char *value, void *cb)\n {\n+\tstruct pack_objects_args *cruft_po_args = cb;\n \tif (!strcmp(var, \"repack.usedeltabaseoffset\")) {\n \t\tdelta_base_offset = git_config_bool(var, value);\n \t\treturn 0;\n@@ -61,6 +73,15 @@ static int repack_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"repack.cruftwindow\"))\n+\t\treturn git_config_string(&cruft_po_args->window, var, value);\n+\tif (!strcmp(var, \"repack.cruftwindowmemory\"))\n+\t\treturn git_config_string(&cruft_po_args->window_memory, var, value);\n+\tif (!strcmp(var, \"repack.cruftdepth\"))\n+\t\treturn git_config_string(&cruft_po_args->depth, var, value);\n+\tif (!strcmp(var, \"repack.cruftthreads\"))\n+\t\treturn git_config_string(&cruft_po_args->threads, var, value);\n+\n \treturn git_default_config(var, value, cb);\n }\n \n@@ -153,18 +174,6 @@ static void remove_redundant_pack(const char *dir_name, const char *base_name)\n \tstrbuf_release(&buf);\n }\n \n-struct pack_objects_args {\n-\tconst char *window;\n-\tconst char *window_memory;\n-\tconst char *depth;\n-\tconst char *threads;\n-\tconst char *max_pack_size;\n-\tint no_reuse_delta;\n-\tint no_reuse_object;\n-\tint quiet;\n-\tint local;\n-};\n-\n static void prepare_pack_objects(struct child_process *cmd,\n \t\t\t\t const struct pack_objects_args *args)\n {\n@@ -689,6 +698,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tint no_update_server_info = 0;\n \tstruct pack_objects_args po_args = {NULL};\n+\tstruct pack_objects_args cruft_po_args = {NULL};\n \tint geometric_factor = 0;\n \tint write_midx = 0;\n \n@@ -743,7 +753,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tOPT_END()\n \t};\n \n-\tgit_config(repack_config, NULL);\n+\tgit_config(repack_config, &cruft_po_args);\n \n \targc = parse_options(argc, argv, prefix, builtin_repack_options,\n \t\t\t\tgit_repack_usage, 0);\n@@ -918,7 +928,19 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tif (*pack_prefix == '/')\n \t\t\tpack_prefix++;\n \n-\t\tret = write_cruft_pack(&po_args, pack_prefix, &names,\n+\t\tif (!cruft_po_args.window)\n+\t\t\tcruft_po_args.window = po_args.window;\n+\t\tif (!cruft_po_args.window_memory)\n+\t\t\tcruft_po_args.window_memory = po_args.window_memory;\n+\t\tif (!cruft_po_args.depth)\n+\t\t\tcruft_po_args.depth = po_args.depth;\n+\t\tif (!cruft_po_args.threads)\n+\t\t\tcruft_po_args.threads = po_args.threads;\n+\n+\t\tcruft_po_args.local = po_args.local;\n+\t\tcruft_po_args.quiet = po_args.quiet;\n+\n+\t\tret = write_cruft_pack(&cruft_po_args, pack_prefix, &names,\n \t\t\t\t       &existing_nonkept_packs,\n \t\t\t\t       &existing_kept_packs);\n \t\tif (ret)\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 06c550c958..e4744e4465 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -565,4 +565,87 @@ test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n \t)\n '\n \n+test_expect_success 'cruft repack respects repack.cruftWindow' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tGIT_TRACE2_EVENT=$(pwd)/event.trace \\\n+\t\tgit -c pack.window=1 -c repack.cruftWindow=2 repack \\\n+\t\t       --cruft --window=3 &&\n+\n+\t\tgrep \"pack-objects.*--window=2.*--cruft\" event.trace\n+\t)\n+'\n+\n+test_expect_success 'cruft repack respects --window by default' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tGIT_TRACE2_EVENT=$(pwd)/event.trace \\\n+\t\tgit -c pack.window=2 repack --cruft --window=3 &&\n+\n+\t\tgrep \"pack-objects.*--window=3.*--cruft\" event.trace\n+\t)\n+'\n+\n+test_expect_success 'cruft repack respects --quiet' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\t\tGIT_PROGRESS_DELAY=0 git repack --cruft --quiet 2>err &&\n+\t\ttest_must_be_empty err\n+\t)\n+'\n+\n+test_expect_success 'cruft --local drops unreachable objects' '\n+\tgit init alternate &&\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr alternate repo\" &&\n+\n+\ttest_commit -C alternate base &&\n+\t# Pack all objects in alterate so that the cruft repack in \"repo\" sees\n+\t# the object it dropped due to `--local` as packed. Otherwise this\n+\t# object would not appear packed anywhere (since it is not packed in\n+\t# alternate and likewise not part of the cruft pack in the other repo\n+\t# because of `--local`).\n+\tgit -C alternate repack -ad &&\n+\n+\t(\n+\t\tcd repo &&\n+\n+\t\tobject=\"$(git -C ../alternate rev-parse HEAD:base.t)\" &&\n+\t\tgit -C ../alternate cat-file -p $object >contents &&\n+\n+\t\t# Write some reachable objects and two unreachable ones: one\n+\t\t# that the alternate has and another that is unique.\n+\t\ttest_commit other &&\n+\t\tgit hash-object -w -t blob contents &&\n+\t\tcruft=\"$(echo cruft | git hash-object -w -t blob --stdin)\" &&\n+\n+\t\t( cd ../alternate/.git/objects && pwd ) \\\n+\t\t       >.git/objects/info/alternates &&\n+\n+\t\ttest_path_is_file $objdir/$(test_oid_to_path $cruft) &&\n+\t\ttest_path_is_file $objdir/$(test_oid_to_path $object) &&\n+\n+\t\tgit repack -d --cruft --local &&\n+\n+\t\ttest-tool pack-mtimes \"$(basename $(ls $packdir/pack-*.mtimes))\" \\\n+\t\t       >objects &&\n+\t\t! grep $object objects &&\n+\t\tgrep $cruft objects\n+\t)\n+'\n+\n test_done\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450216","messageId":"1a58807df02aa6cf487599b2d875a7a24dd16a32.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 14/17] builtin/repack.c: use named flags for existing_packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:21:18Z","receivedAt":"2022-03-03T00:21:43Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"We use the `util` pointer for items in the `existing_packs` string list\nto indicate which packs are going to be deleted. Since that has so far\nbeen the only use of that `util` pointer, we just set it to 0 or 1.\n\nBut we're going to add an additional state to this field in the next\npatch, so prepare for that by adding a #define for the first bit so we\ncan more expressively inspect the flags state.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/repack.c | 9 ++++++---\n 1 file changed, 6 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex d61c78e94e..afa4d51a22 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -22,6 +22,8 @@\n #define LOOSEN_UNREACHABLE 2\n #define PACK_CRUFT 4\n \n+#define DELETE_PACK 1\n+\n static int pack_everything;\n static int delta_base_offset = 1;\n static int pack_kept_objects = -1;\n@@ -561,7 +563,7 @@ static void midx_included_packs(struct string_list *include,\n \t\t}\n \t} else {\n \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n-\t\t\tif (item->util)\n+\t\t\tif ((uintptr_t)item->util & DELETE_PACK)\n \t\t\t\tcontinue;\n \t\t\tstring_list_insert(include, xstrfmt(\"%s.idx\", item->string));\n \t\t}\n@@ -1000,7 +1002,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t * was given) and that we will actually delete this pack\n \t\t\t * (if `-d` was given).\n \t\t\t */\n-\t\t\titem->util = (void*)(intptr_t)!string_list_has_string(&names, sha1);\n+\t\t\tif (!string_list_has_string(&names, sha1))\n+\t\t\t\titem->util = (void*)(uintptr_t)((size_t)item->util | DELETE_PACK);\n \t\t}\n \t}\n \n@@ -1024,7 +1027,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (delete_redundant) {\n \t\tint opts = 0;\n \t\tfor_each_string_list_item(item, &existing_nonkept_packs) {\n-\t\t\tif (!item->util)\n+\t\t\tif (!((uintptr_t)item->util & DELETE_PACK))\n \t\t\t\tcontinue;\n \t\t\tremove_redundant_pack(packdir, item->string);\n \t\t}\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450217","messageId":"1d5f334138998a3e078aba9c105bc89b71045bd1.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 16/17] builtin/gc.c: conditionally avoid pruning objects via loose","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:21:23Z","receivedAt":"2022-03-03T00:21:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Expose the new `git repack --cruft` mode from `git gc` via a new opt-in\nflag. When invoked like `git gc --cruft`, `git gc` will avoid exploding\nunreachable objects as loose ones, and instead create a cruft pack and\n`.mtimes` file.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/gc.txt   | 21 +++++++++++++-------\n Documentation/git-gc.txt      |  5 +++++\n builtin/gc.c                  | 10 +++++++++-\n t/t5329-pack-objects-cruft.sh | 37 +++++++++++++++++++++++++++++++++++\n 4 files changed, 65 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/config/gc.txt b/Documentation/config/gc.txt\nindex c834e07991..38fea076a2 100644\n--- a/Documentation/config/gc.txt\n+++ b/Documentation/config/gc.txt\n@@ -81,14 +81,21 @@ gc.packRefs::\n \tto enable it within all non-bare repos or it can be set to a\n \tboolean value.  The default is `true`.\n \n+gc.cruftPacks::\n+\tStore unreachable objects in a cruft pack (see\n+\tlinkgit:git-repack[1]) instead of as loose objects. The default\n+\tis `false`.\n+\n gc.pruneExpire::\n-\tWhen 'git gc' is run, it will call 'prune --expire 2.weeks.ago'.\n-\tOverride the grace period with this config variable.  The value\n-\t\"now\" may be used to disable this grace period and always prune\n-\tunreachable objects immediately, or \"never\" may be used to\n-\tsuppress pruning.  This feature helps prevent corruption when\n-\t'git gc' runs concurrently with another process writing to the\n-\trepository; see the \"NOTES\" section of linkgit:git-gc[1].\n+\tWhen 'git gc' is run, it will call 'prune --expire 2.weeks.ago'\n+\t(and 'repack --cruft --cruft-expiration 2.weeks.ago' if using\n+\tcruft packs via `gc.cruftPacks` or `--cruft`).  Override the\n+\tgrace period with this config variable.  The value \"now\" may be\n+\tused to disable this grace period and always prune unreachable\n+\tobjects immediately, or \"never\" may be used to suppress pruning.\n+\tThis feature helps prevent corruption when 'git gc' runs\n+\tconcurrently with another process writing to the repository; see\n+\tthe \"NOTES\" section of linkgit:git-gc[1].\n \n gc.worktreePruneExpire::\n \tWhen 'git gc' is run, it calls\ndiff --git a/Documentation/git-gc.txt b/Documentation/git-gc.txt\nindex 853967dea0..ba4e67700e 100644\n--- a/Documentation/git-gc.txt\n+++ b/Documentation/git-gc.txt\n@@ -54,6 +54,11 @@ other housekeeping tasks (e.g. rerere, working trees, reflog...) will\n be performed as well.\n \n \n+--cruft::\n+\tWhen expiring unreachable objects, pack them separately into a\n+\tcruft pack instead of storing the loose objects as loose\n+\tobjects.\n+\n --prune=<date>::\n \tPrune loose objects older than date (default is 2 weeks ago,\n \toverridable by the config variable `gc.pruneExpire`).\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex ffaf0daf5d..11f5150234 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -43,6 +43,7 @@ static const char * const builtin_gc_usage[] = {\n \n static int pack_refs = 1;\n static int prune_reflogs = 1;\n+static int cruft_packs = 0;\n static int aggressive_depth = 50;\n static int aggressive_window = 250;\n static int gc_auto_threshold = 6700;\n@@ -153,6 +154,7 @@ static void gc_config(void)\n \tgit_config_get_int(\"gc.auto\", &gc_auto_threshold);\n \tgit_config_get_int(\"gc.autopacklimit\", &gc_auto_pack_limit);\n \tgit_config_get_bool(\"gc.autodetach\", &detach_auto);\n+\tgit_config_get_bool(\"gc.cruftpacks\", &cruft_packs);\n \tgit_config_get_expiry(\"gc.pruneexpire\", &prune_expire);\n \tgit_config_get_expiry(\"gc.worktreepruneexpire\", &prune_worktrees_expire);\n \tgit_config_get_expiry(\"gc.logexpiry\", &gc_log_expire);\n@@ -332,7 +334,11 @@ static void add_repack_all_option(struct string_list *keep_pack)\n {\n \tif (prune_expire && !strcmp(prune_expire, \"now\"))\n \t\tstrvec_push(&repack, \"-a\");\n-\telse {\n+\telse if (cruft_packs) {\n+\t\tstrvec_push(&repack, \"--cruft\");\n+\t\tif (prune_expire)\n+\t\t\tstrvec_pushf(&repack, \"--cruft-expiration=%s\", prune_expire);\n+\t} else {\n \t\tstrvec_push(&repack, \"-A\");\n \t\tif (prune_expire)\n \t\t\tstrvec_pushf(&repack, \"--unpack-unreachable=%s\", prune_expire);\n@@ -552,6 +558,7 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t{ OPTION_STRING, 0, \"prune\", &prune_expire, N_(\"date\"),\n \t\t\tN_(\"prune unreferenced objects\"),\n \t\t\tPARSE_OPT_OPTARG, NULL, (intptr_t)prune_expire },\n+\t\tOPT_BOOL(0, \"cruft\", &cruft_packs, N_(\"pack unreferenced objects separately\")),\n \t\tOPT_BOOL(0, \"aggressive\", &aggressive, N_(\"be more thorough (increased runtime)\")),\n \t\tOPT_BOOL_F(0, \"auto\", &auto_gc, N_(\"enable auto-gc mode\"),\n \t\t\t   PARSE_OPT_NOCOMPLETE),\n@@ -671,6 +678,7 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t\tdie(FAILED_RUN, repack.v[0]);\n \n \t\tif (prune_expire) {\n+\t\t\t/* run `git prune` even if using cruft packs */\n \t\t\tstrvec_push(&prune, prune_expire);\n \t\t\tif (quiet)\n \t\t\t\tstrvec_push(&prune, \"--no-progress\");\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 13158e4ab7..3910e186ef 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -429,6 +429,43 @@ test_expect_success 'loose objects mtimes upsert others' '\n \t)\n '\n \n+test_expect_success 'expiring cruft objects with git gc' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit branch -M main &&\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit unreachable &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d unreachable &&\n+\t\t# objects are not cruft if they are contained in the reflogs\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\t\tsort <reachable.raw >reachable &&\n+\t\tcomm -13 reachable objects >unreachable &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tmtimes=$(ls .git/objects/pack/pack-*.mtimes) &&\n+\t\ttest_path_is_file $mtimes &&\n+\n+\t\tgit gc --cruft --prune=now &&\n+\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\n+\t\tcomm -23 unreachable objects >removed &&\n+\t\ttest_cmp unreachable removed &&\n+\t\ttest_path_is_missing $mtimes\n+\t)\n+'\n+\n test_expect_success 'cruft packs are not included in geometric repack' '\n \tgit init repo &&\n \ttest_when_finished \"rm -fr repo\" &&\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450218","messageId":"f74b42587208da364e647cc4847541cacc842753.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 17/17] sha1-file.c: don't freshen cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:21:25Z","receivedAt":"2022-03-03T00:21:57Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"We don't bother to freshen objects stored in a cruft pack individually\nby updating the `.mtimes` file. This is because we can't portably `mmap`\nand write into the middle of a file (i.e., to update the mtime of just\none object). Instead, we would have to rewrite the entire `.mtimes` file\nwhich may incur some wasted effort especially if there a lot of cruft\nobjects and they are freshened infrequently.\n\nInstead, force the freshening code to avoid an optimizing write by\nwriting out the object loose and letting it pick up a current mtime.\n\nThis works because we prefer the mtime of the loose copy of an object\nwhen both a loose and packed one exist (whether or not the packed copy\ncomes from a cruft pack or not).\n\nThis could certainly do with a test and/or be included earlier in this\nseries/PR, but I want to wait until after I have a chance to clean up\nthe overly-repetitive nature of the cruft pack tests in general.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n object-file.c                 |  2 ++\n t/t5329-pack-objects-cruft.sh | 25 +++++++++++++++++++++++++\n 2 files changed, 27 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex e80da1368d..65b8df7fb6 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1989,6 +1989,8 @@ static int freshen_packed_object(const struct object_id *oid)\n \tstruct pack_entry e;\n \tif (!find_pack_entry(the_repository, oid, &e))\n \t\treturn 0;\n+\tif (e.p->is_cruft)\n+\t\treturn 0;\n \tif (e.p->freshened)\n \t\treturn 1;\n \tif (!freshen_file(e.p->pack_name))\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 3910e186ef..4681558612 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -711,4 +711,29 @@ test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n \t)\n '\n \n+test_expect_success 'cruft objects are freshend via loose' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\techo \"cruft\" >contents &&\n+\t\tblob=\"$(git hash-object -w -t blob contents)\" &&\n+\t\tloose=\"$objdir/$(test_oid_to_path $blob)\" &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\ttest_path_is_missing \"$loose\" &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(ls $packdir/pack-*.mtimes)\")\" >cruft &&\n+\t\tgrep \"$blob\" cruft &&\n+\n+\t\t# write the same object again\n+\t\tgit hash-object -w -t blob contents &&\n+\n+\t\ttest_path_is_file \"$loose\"\n+\t)\n+'\n+\n test_done\n-- \n2.35.1.73.gccc5557600\n"},{"id":"450219","messageId":"ed05cf536bdea62a2b512bb3a610ec7861ff68b5.1646266835.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"[PATCH v3 15/17] builtin/repack.c: add cruft packs to MIDX during geometric repack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T00:21:20Z","receivedAt":"2022-03-03T00:21:58Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When using cruft packs, the following race can occur when a geometric\nrepack that writes a MIDX bitmap takes place afterwords:\n\n  - First, create an unreachable object and do an all-into-one cruft\n    repack which stores that object in the repository's cruft pack.\n  - Then make that object reachable.\n  - Finally, do a geometric repack and write a MIDX bitmap.\n\nAssuming that we are sufficiently unlucky as to select a commit from the\nMIDX which reaches that object for bitmapping, then the `git\nmulti-pack-index` process will complain that that object is missing.\n\nThe reason is because we don't include cruft packs in the MIDX when\ndoing a geometric repack. Since the \"make that object reachable\" doesn't\nnecessarily mean that we'll create a new copy of that object in one of\nthe packs that will get rolled up as part of a geometric repack, it's\npossible that the MIDX won't see any copies of that now-reachable\nobject.\n\nOf course, it's desirable to avoid including cruft packs in the MIDX\nbecause it causes the MIDX to store a bunch of objects which are likely\nto get thrown away. But excluding that pack does open us up to the above\nrace.\n\nThis patch demonstrates the bug, and resolves it by including cruft\npacks in the MIDX even when doing a geometric repack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/repack.c              | 19 +++++++++++++++++--\n t/t5329-pack-objects-cruft.sh | 26 ++++++++++++++++++++++++++\n 2 files changed, 43 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex afa4d51a22..59b60cd309 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -23,6 +23,7 @@\n #define PACK_CRUFT 4\n \n #define DELETE_PACK 1\n+#define CRUFT_PACK 2\n \n static int pack_everything;\n static int delta_base_offset = 1;\n@@ -158,8 +159,11 @@ static void collect_pack_filenames(struct string_list *fname_nonkept_list,\n \t\tif ((extra_keep->nr > 0 && i < extra_keep->nr) ||\n \t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname))))\n \t\t\tstring_list_append_nodup(fname_kept_list, fname);\n-\t\telse\n-\t\t\tstring_list_append_nodup(fname_nonkept_list, fname);\n+\t\telse {\n+\t\t\tstruct string_list_item *item = string_list_append_nodup(fname_nonkept_list, fname);\n+\t\t\tif (file_exists(mkpath(\"%s/%s.mtimes\", packdir, fname)))\n+\t\t\t\titem->util = (void*)(uintptr_t)CRUFT_PACK;\n+\t\t}\n \t}\n \tclosedir(dir);\n }\n@@ -561,6 +565,17 @@ static void midx_included_packs(struct string_list *include,\n \n \t\t\tstring_list_insert(include, strbuf_detach(&buf, NULL));\n \t\t}\n+\n+\t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n+\t\t\tif (!((uintptr_t)item->util & CRUFT_PACK)) {\n+\t\t\t\t/*\n+\t\t\t\t * no need to check DELETE_PACK, since we're not\n+\t\t\t\t * doing an ALL_INTO_ONE repack\n+\t\t\t\t */\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tstring_list_insert(include, xstrfmt(\"%s.idx\", item->string));\n+\t\t}\n \t} else {\n \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n \t\t\tif ((uintptr_t)item->util & DELETE_PACK)\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex e4744e4465..13158e4ab7 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -648,4 +648,30 @@ test_expect_success 'cruft --local drops unreachable objects' '\n \t)\n '\n \n+test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\ttest_commit cruft &&\n+\t\tunreachable=\"$(git rev-parse cruft)\" &&\n+\n+\t\tgit reset --hard $unreachable^ &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\t# resurrect the unreachable object via a new commit. the\n+\t\t# new commit will get selected for a bitmap, but be\n+\t\t# missing one of its parents from the selected packs.\n+\t\tgit reset --hard $unreachable &&\n+\t\ttest_commit resurrect &&\n+\n+\t\tgit repack --write-midx --write-bitmap-index --geometric=2 -d\n+\t)\n+'\n+\n test_done\n-- \n2.35.1.73.gccc5557600\n\n"},{"id":"450238","messageId":"5063be12-fb66-9936-9ec3-df02d4c9cfd9@github.com","threadId":"56996","inReplyTo":"cover.1646266835.git.me@ttaylorr.com","subject":"Re: [PATCH v3 00/17] cruft packs","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-03-03T01:29:37Z","receivedAt":"2022-03-03T01:29:42Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 3/2/2022 7:20 PM, Taylor Blau wrote:\n> Here is a small reroll of my series to implement \"cruft packs\", based on\n> Stolee's review.\n> \n> The changes here are minor, and mostly are limited to removing a\n> redundant \"if\" statement, avoiding an unnecessary header include, and\n> moving the tests (again!) to t5329's territory.\n> \n> As always, a range-diff is below. Thanks in advance for taking another\n> look!\n\nThis range-diff satisfies my comments. Thanks!\n-Stolee\n"},{"id":"450288","messageId":"220303.86ilsv2dke.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"135a07276b0a40b04f2c28d4f48c26b1af76c12c.1646266835.git.me@ttaylorr.com","subject":"Re: [PATCH v3 04/17] chunk-format.h: extract oid_version()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-03T16:30:44Z","receivedAt":"2022-03-03T16:41:59Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Mar 02 2022, Taylor Blau wrote:\n\n> Consolidate these into a single definition in chunk-format.h. It's not\n> clear that this is the best header to define this function in, but it\n> should do for now.\n> [...]\n> +\n> +uint8_t oid_version(const struct git_hash_algo *algop)\n> +{\n> +\tswitch (hash_algo_by_ptr(algop)) {\n> +\tcase GIT_HASH_SHA1:\n> +\t\treturn 1;\n> +\tcase GIT_HASH_SHA256:\n> +\t\treturn 2;\n\nNot a new issue, but I wonder why these don't return hash_algo_by_ptr\naka GIT_HASH_WHATEVER here. I.e. this is the same as this more\nstraightforward & obvious code that avoids re-hardcoding the magic\nconstants:\n\n\tconst int algo = hash_algo_by_ptr(algop)\n\n\tswitch (algo) {\n\tcase GIT_HASH_SHA1:\n\tcase GIT_HASH_SHA256:\n\t\treturn algo;\n\tdefault:\n        [...]\n        }\n\nProbably best left as a later cleanup. FWIW I came up with this on top\nof my designated init series:\n\ndiff --git a/hash.h b/hash.h\nindex 5d40368f18a..fd710ec6ae8 100644\n--- a/hash.h\n+++ b/hash.h\n@@ -86,14 +86,18 @@ static inline void git_SHA256_Clone(git_SHA256_CTX *dst, const git_SHA256_CTX *s\n  * field for being non-zero.  Use the name field for user-visible situations and\n  * the format_id field for fixed-length fields on disk.\n  */\n-/* An unknown hash function. */\n-#define GIT_HASH_UNKNOWN 0\n-/* SHA-1 */\n-#define GIT_HASH_SHA1 1\n-/* SHA-256  */\n-#define GIT_HASH_SHA256 2\n-/* Number of algorithms supported (including unknown). */\n-#define GIT_HASH_NALGOS (GIT_HASH_SHA256 + 1)\n+enum git_hash_algo_name {\n+\t/* An unknown hash function. */\n+\tGIT_HASH_UNKNOWN,\n+\t/* SHA-1 */\n+\tGIT_HASH_SHA1,\n+\tGIT_HASH_SHA256,\n+\t/*\n+\t * Number of algorithms supported (including unknown). This\n+\t * must be kept last!\n+\t */\n+\tGIT_HASH_NALGOS,\n+};\n \n /* \"sha1\", big-endian */\n #define GIT_SHA1_FORMAT_ID 0x73686131\ndiff --git a/object-file.c b/object-file.c\nindex 5074471b471..f2d54a86969 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -166,7 +166,7 @@ static void git_hash_unknown_final_oid(struct object_id *oid, git_hash_ctx *ctx)\n }\n \n const struct git_hash_algo hash_algos[GIT_HASH_NALGOS] = {\n-\t{\n+\t[GIT_HASH_UNKNOWN] = {\n \t\t.name = NULL,\n \t\t.format_id = 0x00000000,\n \t\t.rawsz = 0,\n@@ -181,7 +181,7 @@ const struct git_hash_algo hash_algos[GIT_HASH_NALGOS] = {\n \t\t.empty_blob = NULL,\n \t\t.null_oid = NULL,\n \t},\n-\t{\n+\t[GIT_HASH_SHA1] = {\n \t\t.name = \"sha1\",\n \t\t.format_id = GIT_SHA1_FORMAT_ID,\n \t\t.rawsz = GIT_SHA1_RAWSZ,\n@@ -196,7 +196,7 @@ const struct git_hash_algo hash_algos[GIT_HASH_NALGOS] = {\n \t\t.empty_blob = &empty_blob_oid,\n \t\t.null_oid = &null_oid_sha1,\n \t},\n-\t{\n+\t[GIT_HASH_SHA256] = {\n \t\t.name = \"sha256\",\n \t\t.format_id = GIT_SHA256_FORMAT_ID,\n \t\t.rawsz = GIT_SHA256_RAWSZ,\n"},{"id":"450289","messageId":"220303.86ee3j2dae.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"0600503856dbccb135aaead27693b6815a774b4f.1646266835.git.me@ttaylorr.com","subject":"Re: [PATCH v3 05/17] pack-mtimes: support writing pack .mtimes files","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-03T16:45:23Z","receivedAt":"2022-03-03T16:48:00Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Mar 02 2022, Taylor Blau wrote:\n\n> Now that the `.mtimes` format is defined, supplement the pack-write API\n> to be able to conditionally write an `.mtimes` file along with a pack by\n> setting an additional flag and passing an oidmap that contains the\n> timestamps corresponding to each object in the pack.\n> [...]\n>  void write_promisor_file(const char *promisor_name, struct ref **sought, int nr_sought)\n> diff --git a/pack.h b/pack.h\n> index fd27cfdfd7..01d385903a 100644\n> --- a/pack.h\n> +++ b/pack.h\n> @@ -44,6 +44,7 @@ struct pack_idx_option {\n>  #define WRITE_IDX_STRICT 02\n>  #define WRITE_REV 04\n>  #define WRITE_REV_VERIFY 010\n> +#define WRITE_MTIMES 020\n>  \n>  \tuint32_t version;\n>  \tuint32_t off32_limit;\n\nWhy the hardcoding? The 010 was added in your 8ef50d9958f (pack-write.c:\nprepare to write 'pack-*.rev' files, 2021-01-25). That would be the same\nas 8|2, but there's no 8 there., ditto this new 020 that's the same as\n1<<4 | 1<<2, but there's no \"16\", just WRITE_REV=4.\n"},{"id":"450322","messageId":"YiFQBobcLT2m/KMx@nand.local","threadId":"56996","inReplyTo":"220303.86ilsv2dke.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v3 04/17] chunk-format.h: extract oid_version()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T23:32:22Z","receivedAt":"2022-03-03T23:32:27Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, Mar 03, 2022 at 05:30:44PM +0100, Ævar Arnfjörð Bjarmason wrote:\n>\n> On Wed, Mar 02 2022, Taylor Blau wrote:\n>\n> > Consolidate these into a single definition in chunk-format.h. It's not\n> > clear that this is the best header to define this function in, but it\n> > should do for now.\n> > [...]\n> > +\n> > +uint8_t oid_version(const struct git_hash_algo *algop)\n> > +{\n> > +\tswitch (hash_algo_by_ptr(algop)) {\n> > +\tcase GIT_HASH_SHA1:\n> > +\t\treturn 1;\n> > +\tcase GIT_HASH_SHA256:\n> > +\t\treturn 2;\n>\n> Not a new issue, but I wonder why these don't return hash_algo_by_ptr\n> aka GIT_HASH_WHATEVER here. I.e. this is the same as this more\n> straightforward & obvious code that avoids re-hardcoding the magic\n> constants:\n\nHmm. Certainly the value returned by hash_algo_by_ptr() works for SHA-1\nand SHA-256, but writes may want to use a different value for future\nhashes. Not that this couldn't be changed then, but my feeling is that\nthe existing code is clearer since it avoids the reader having to jump\nto hash_algo_by_ptr()'s implementation to figure out what it returns.\n\nThanks,\nTaylor\n"},{"id":"450323","messageId":"YiFQxsmkcqb63azh@nand.local","threadId":"56996","inReplyTo":"220303.86ee3j2dae.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v3 05/17] pack-mtimes: support writing pack .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-03T23:35:34Z","receivedAt":"2022-03-03T23:35:40Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, Mar 03, 2022 at 05:45:23PM +0100, Ævar Arnfjörð Bjarmason wrote:\n>\n> On Wed, Mar 02 2022, Taylor Blau wrote:\n>\n> > Now that the `.mtimes` format is defined, supplement the pack-write API\n> > to be able to conditionally write an `.mtimes` file along with a pack by\n> > setting an additional flag and passing an oidmap that contains the\n> > timestamps corresponding to each object in the pack.\n> > [...]\n> >  void write_promisor_file(const char *promisor_name, struct ref **sought, int nr_sought)\n> > diff --git a/pack.h b/pack.h\n> > index fd27cfdfd7..01d385903a 100644\n> > --- a/pack.h\n> > +++ b/pack.h\n> > @@ -44,6 +44,7 @@ struct pack_idx_option {\n> >  #define WRITE_IDX_STRICT 02\n> >  #define WRITE_REV 04\n> >  #define WRITE_REV_VERIFY 010\n> > +#define WRITE_MTIMES 020\n> >\n> >  \tuint32_t version;\n> >  \tuint32_t off32_limit;\n>\n> Why the hardcoding? The 010 was added in your 8ef50d9958f (pack-write.c:\n> prepare to write 'pack-*.rev' files, 2021-01-25). That would be the same\n> as 8|2, but there's no 8 there., ditto this new 020 that's the same as\n> 1<<4 | 1<<2, but there's no \"16\", just WRITE_REV=4.\n\nI'm not sure I understand. These are octals, so octal \"20\" (or decimal\n16) just gives us bit 5 -- the next available -- by itself.\n\nThanks,\nTaylor\n"},{"id":"450328","messageId":"xmqq4k4e7esn.fsf@gitster.g","threadId":"56996","inReplyTo":"YiFQBobcLT2m/KMx@nand.local","subject":"Re: [PATCH v3 04/17] chunk-format.h: extract oid_version()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-04T00:16:24Z","receivedAt":"2022-03-04T00:16:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> On Thu, Mar 03, 2022 at 05:30:44PM +0100, Ævar Arnfjörð Bjarmason wrote:\n>>\n>> On Wed, Mar 02 2022, Taylor Blau wrote:\n>>\n>> > Consolidate these into a single definition in chunk-format.h. It's not\n>> > clear that this is the best header to define this function in, but it\n>> > should do for now.\n>> > [...]\n>> > +\n>> > +uint8_t oid_version(const struct git_hash_algo *algop)\n>> > +{\n>> > +\tswitch (hash_algo_by_ptr(algop)) {\n>> > +\tcase GIT_HASH_SHA1:\n>> > +\t\treturn 1;\n>> > +\tcase GIT_HASH_SHA256:\n>> > +\t\treturn 2;\n>>\n>> Not a new issue, but I wonder why these don't return hash_algo_by_ptr\n>> aka GIT_HASH_WHATEVER here. I.e. this is the same as this more\n>> straightforward & obvious code that avoids re-hardcoding the magic\n>> constants:\n>\n> Hmm. Certainly the value returned by hash_algo_by_ptr() works for SHA-1\n> and SHA-256, but writes may want to use a different value for future\n> hashes. Not that this couldn't be changed then, but my feeling is that\n> the existing code is clearer since it avoids the reader having to jump\n> to hash_algo_by_ptr()'s implementation to figure out what it returns.\n\nIf we promise that everywhere in file formats where we identify what\nhash is used, we write \"1\" for SHA1 and \"2\" for SHA256, it would be\nnatural to define GIT_HASH_SHA1 to \"1\" and GIT_HASH_SHA256 to \"2\".\n\nAnd readers do not have to \"figure out\", if that is a clearly\nwritten guideline to represent the hash used in file formats.  As\nwritten, the readers who -assumes- such a guideline is there must\nfigure out from hash.h that GIT_HASH_SHA1 is 1 and GIT_HASH_SHA256\nis 2 to be convinced that the above code is correct.\n\nNow, hash.h says GIT_HASH_SHA1 is 1 and GIT_HASH_SHA256 is 2.  So\n\n\tint oidv = hash_algo_by_ptr(algop)\n\tswitch (oidv) {\n\tcase GIT_HASH_SHA1:\n\tcase GIT_HASH_SHA256:\n\t\treturn oidv;\n\tdefault:\n\t\tdie();\n\t}\n\nshould work already.  To put it differently, if this didn't work, we\nshould renumber GIT_HASH_SHA1 and GIT_HASH_SHA256 to make it work, I\nwould think.  If not, we have a huge mess on our hands, as constants\nused in on-disk file formats is hard (almost impossible) to change.\n\nAn overly generic function name oid_version() cannot be justified\nunless the same constants are used everywhere.  I see hits from 'git\ngrep oid_version' in\n\n    chunk-format.c (obviously)\n    commit-graph.c\n    midx.c\n    pack-write.c\n\nso presumably these types of files are using the \"canonical\"\nnumbering.\n\nAnd when we introduce GIT_HASH_SHA3 or whatever, we should give it a\nnumber that this function can return (i.e. from the range 3..255).\n"},{"id":"450385","messageId":"220304.86ee3if13q.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"YiFQxsmkcqb63azh@nand.local","subject":"Re: [PATCH v3 05/17] pack-mtimes: support writing pack .mtimes files","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-04T10:40:10Z","receivedAt":"2022-03-04T10:45:03Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Mar 03 2022, Taylor Blau wrote:\n\n> On Thu, Mar 03, 2022 at 05:45:23PM +0100, Ævar Arnfjörð Bjarmason wrote:\n>>\n>> On Wed, Mar 02 2022, Taylor Blau wrote:\n>>\n>> > Now that the `.mtimes` format is defined, supplement the pack-write API\n>> > to be able to conditionally write an `.mtimes` file along with a pack by\n>> > setting an additional flag and passing an oidmap that contains the\n>> > timestamps corresponding to each object in the pack.\n>> > [...]\n>> >  void write_promisor_file(const char *promisor_name, struct ref **sought, int nr_sought)\n>> > diff --git a/pack.h b/pack.h\n>> > index fd27cfdfd7..01d385903a 100644\n>> > --- a/pack.h\n>> > +++ b/pack.h\n>> > @@ -44,6 +44,7 @@ struct pack_idx_option {\n>> >  #define WRITE_IDX_STRICT 02\n>> >  #define WRITE_REV 04\n>> >  #define WRITE_REV_VERIFY 010\n>> > +#define WRITE_MTIMES 020\n>> >\n>> >  \tuint32_t version;\n>> >  \tuint32_t off32_limit;\n>>\n>> Why the hardcoding? The 010 was added in your 8ef50d9958f (pack-write.c:\n>> prepare to write 'pack-*.rev' files, 2021-01-25). That would be the same\n>> as 8|2, but there's no 8 there., ditto this new 020 that's the same as\n>> 1<<4 | 1<<2, but there's no \"16\", just WRITE_REV=4.\n>\n> I'm not sure I understand. These are octals, so octal \"20\" (or decimal\n> 16) just gives us bit 5 -- the next available -- by itself.\n\nUrgh, tired/rushed eyes yesterday. I managed to read these as decimals,\nsorry.\n\nI see from:\n\n    git grep 'define[^0-9]*(\\b020\\b|\\b16\\b|1.*<<.*\\b4\\b)[^0-9]*$'\n\nThat I managed to patch what seems to be one of two other places in the\ncodebase using it recently (that goes >=020) in 245b9488150 (cat-file:\nuse GET_OID_ONLY_TO_DIE in --(textconv|filters), 2021-12-28).\n\nAnyway, I think nothing needs to be done here. If you ever feel like\nsome churn here I think converting it to the almost ubiquitous \"1 << N\"\nstyle we use almost everywhere else would be an improvement :)\n\nSorry!\n"},{"id":"450626","messageId":"YiZI99yeijQe5Jaq@google.com","threadId":"56996","inReplyTo":"784ee7e0eec9ba520ebaaa27de2de810e2f6798a.1646266835.git.me@ttaylorr.com","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2022-03-07T18:03:35Z","receivedAt":"2022-03-07T18:03:41Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nTaylor Blau wrote:\n\n> Create a technical document to explain cruft packs. It contains a brief\n> overview of the problem, some background, details on the implementation,\n> and a couple of alternative approaches not considered here.\n\nSorry for the very slow review!  I've mentioned a few times that this\noverlaps in interesting ways with the gc mechanism described in\nhash-function-transition.txt, so I'd like to compare and see how they\ninteract.\n\n[...]\n> --- /dev/null\n> +++ b/Documentation/technical/cruft-packs.txt\n> @@ -0,0 +1,97 @@\n[...]\n> +Unreachable objects aren't removed immediately, since doing so could race with\n> +an incoming push which may reference an object which is about to be deleted.\n> +Instead, those unreachable objects are stored as loose object and stay that way\n> +until they are older than the expiration window, at which point they are removed\n> +by linkgit:git-prune[1].\n> +\n> +Git must store these unreachable objects loose in order to keep track of their\n> +per-object mtimes.\n\nIt's worth noting that this behavior is already racy.  That is because\nwhen an unreachable object becomes newly reachable, we do not update\nits mtime and the mtimes of every object reachable from it, so if it\nthen becomes transiently unreachable again then it can be wrongly\ncollected.\n\n[...]\n>                                these repositories often take up a large amount of\n> +disk space, since we can only zlib compress them, but not store them in delta\n> +chains.\n\nYes!  I'm happy we're making progress on this.\n\n> +\n> +== Cruft packs\n> +\n> +A cruft pack eliminates the need for storing unreachable objects in a loose\n> +state by including the per-object mtimes in a separate file alongside a single\n> +pack containing all loose objects.\n\nCan this doc say a little about how \"git prune\" handles these files?\nIn particular, does a non cruft pack aware copy of Git (or JGit,\nlibgit2, etc) do the right thing or does it fight with this mechanism?\nIf the latter, do we have a repository extension (extensions.*) to\nprevent that?\n\n[...]\n> +  3. Write the pack out, along with a `.mtimes` file that records the per-object\n> +     timestamps.\n\nAs a point of comparison, the design in hash-function-transition uses\na single timestamp for the whole pack.  During read operations, objects\nin a cruft pack are considered present; during writes, they are\nconsidered _not present_ so that if we want to make a cruft object\nnewly present then we put a copy of it in a new pack.\n\nAdvantage of the mtimes file approach:\n- less duplication of storage: a revived object is only stored once,\n  in a cruft pack, and then the next gc can \"graduate\" it out of the\n  cruft pack and shrink the cruft pack\n- less affect on non-gc Git code: writes don't need to know that any\n  cruft objects referenced need to be copied into a new pack\n\nAdvantages of the mtime per cruft pack approach:\n- easy expiration: once a cruft pack has reached its expiration date,\n  it can be deleted as a whole\n- less I/O churn: a cruft pack stays as-is until combined into another\n  cruft pack or deleted.  There is no frequently-modified mtimes file\n  associated to it\n- informs the storage layer about what is likely to be accessed: cruft\n  packs can get filesystem attributes to put them in less-optimized\n  storage since they are likely to be less frequently read\n\n[...]\n> +Notable alternatives to this design include:\n> +\n> +  - The location of the per-object mtime data, and\n> +  - Storing unreachable objects in multiple cruft packs.\n> +\n> +On the location of mtime data, a new auxiliary file tied to the pack was chosen\n> +to avoid complicating the `.idx` format. If the `.idx` format were ever to gain\n> +support for optional chunks of data, it may make sense to consolidate the\n> +`.mtimes` format into the `.idx` itself.\n> +\n> +Storing unreachable objects among multiple cruft packs (e.g., creating a new\n> +cruft pack during each repacking operation including only unreachable objects\n> +which aren't already stored in an earlier cruft pack) is significantly more\n> +complicated to construct, and so aren't pursued here. The obvious drawback to\n> +the current implementation is that the entire cruft pack must be re-written from\n> +scratch.\n\nThis doesn't mention the approach described in\nhash-function-transition.txt (and that's already implemented and has\nbeen in use for many years in JGit's DfsRepository).  Does that mean\nyou aren't aware of it?\n\nThanks,\nJonathan\n"},{"id":"451825","messageId":"YjkjaH61dMLHXr0d@nand.local","threadId":"56996","inReplyTo":"YiZI99yeijQe5Jaq@google.com","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-22T01:16:24Z","receivedAt":"2022-03-22T01:16:29Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Mon, Mar 07, 2022 at 10:03:35AM -0800, Jonathan Nieder wrote:\n> Sorry for the very slow review!  I've mentioned a few times that this\n> overlaps in interesting ways with the gc mechanism described in\n> hash-function-transition.txt, so I'd like to compare and see how they\n> interact.\n\nSorry for my equally-slow reply ;). I was on vacation last week and\nwasn't following the list closely.\n\n> > +Unreachable objects aren't removed immediately, since doing so could race with\n> > +an incoming push which may reference an object which is about to be deleted.\n> > +Instead, those unreachable objects are stored as loose object and stay that way\n> > +until they are older than the expiration window, at which point they are removed\n> > +by linkgit:git-prune[1].\n> > +\n> > +Git must store these unreachable objects loose in order to keep track of their\n> > +per-object mtimes.\n>\n> It's worth noting that this behavior is already racy.  That is because\n> when an unreachable object becomes newly reachable, we do not update\n> its mtime and the mtimes of every object reachable from it, so if it\n> then becomes transiently unreachable again then it can be wrongly\n> collected.\n\nJust to be clear, the race here only happens if the object in question\nbecomes reachable _after_ a pruning GC determines its mtime. If that's\nthe case, then the object will indeed be wrongly collected. This is\nconsistent with the existing behavior (which is racy in the exact same\nway).\n\n(After re-reading what you wrote and my response, I think we are saying\nthe exact same thing, but it doesn't hurt to think aloud).\n\n> > +\n> > +== Cruft packs\n> > +\n> > +A cruft pack eliminates the need for storing unreachable objects in a loose\n> > +state by including the per-object mtimes in a separate file alongside a single\n> > +pack containing all loose objects.\n>\n> Can this doc say a little about how \"git prune\" handles these files?\n> In particular, does a non cruft pack aware copy of Git (or JGit,\n> libgit2, etc) do the right thing or does it fight with this mechanism?\n> If the latter, do we have a repository extension (extensions.*) to\n> prevent that?\n\nI mentioned this in much more detail in [1], but the answer is that the\ncruft pack looks like any other pack, it just happens to have another\nmetadata file (the .mtimes one) attached to it. So other implementations\nof Git should treat it as they would any other pack. Like I mentioned in\n[1], cruft packs were designed with the explicit goal of not requiring a\nrepository extension.\n\n> > +  3. Write the pack out, along with a `.mtimes` file that records the per-object\n> > +     timestamps.\n>\n> As a point of comparison, the design in hash-function-transition uses\n> a single timestamp for the whole pack.  During read operations, objects\n> in a cruft pack are considered present; during writes, they are\n> considered _not present_ so that if we want to make a cruft object\n> newly present then we put a copy of it in a new pack.\n>\n> Advantage of the mtimes file approach:\n> - less duplication of storage: a revived object is only stored once,\n>   in a cruft pack, and then the next gc can \"graduate\" it out of the\n>   cruft pack and shrink the cruft pack\n> - less affect on non-gc Git code: writes don't need to know that any\n>   cruft objects referenced need to be copied into a new pack\n>\n> Advantages of the mtime per cruft pack approach:\n> - easy expiration: once a cruft pack has reached its expiration date,\n>   it can be deleted as a whole\n> - less I/O churn: a cruft pack stays as-is until combined into another\n>   cruft pack or deleted.  There is no frequently-modified mtimes file\n>   associated to it\n> - informs the storage layer about what is likely to be accessed: cruft\n>   packs can get filesystem attributes to put them in less-optimized\n>   storage since they are likely to be less frequently read\n>\n> [...]\n\nThe key advantage of cruft packs is that you can expire unreachable\nobjects in piecemeal while still retaining the benefit of being able to\nde-duplicate cruft objects and store them packed against each other.\n\n> > +Notable alternatives to this design include:\n>\n> This doesn't mention the approach described in\n> hash-function-transition.txt (and that's already implemented and has\n> been in use for many years in JGit's DfsRepository).  Does that mean\n> you aren't aware of it?\n\nImplementing the UNREACHABLE_GARBAGE concept from\nhash-function-transition.txt in cruft pack-terms would be equivalent to\nnot writing the mtimes file at all. This follows from the fact that a\npre-cruft packs implementation of Git considers a packed object's mtime\nto be the same as the pack it's contained in. (I'm deliberately\navoiding any details from the h-f-t document regarding re-writing\nobjects contained in a garbage pack here, since this is separate from\nthe pack structure itself (and could easily be implemented on top of\ncruft packs)).\n\nSo I'm not sure what the alternative we'd list would be, since it\nremoves the key feature of the design of cruft packs.\n\nThanks,\nTaylor\n\n[1]: https://lore.kernel.org/git/YiZMhuI%2FDdpvQ%2FED@nand.local/\n"},{"id":"451931","messageId":"YjpDbHmKY9XA2p0K@google.com","threadId":"56996","inReplyTo":"YjkjaH61dMLHXr0d@nand.local","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2022-03-22T21:45:16Z","receivedAt":"2022-03-22T21:46:27Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nTaylor Blau wrote:\n> On Mon, Mar 07, 2022 at 10:03:35AM -0800, Jonathan Nieder wrote:\n\n>> Sorry for the very slow review!  I've mentioned a few times that this\n>> overlaps in interesting ways with the gc mechanism described in\n>> hash-function-transition.txt, so I'd like to compare and see how they\n>> interact.\n>\n> Sorry for my equally-slow reply ;). I was on vacation last week and\n> wasn't following the list closely.\n\nNo problem --- thanks for getting back to me.\n\n[...]\n> (After re-reading what you wrote and my response, I think we are saying\n> the exact same thing, but it doesn't hurt to think aloud).\n\nGreat.  Can the doc cover this?  I think it would be helpful to make\nthat easy to find for others with similar questions.\n\nIf it's a matter of finding enough time to write some text, let me\nknow and I can try to find some time to help.\n\n[...]\n>> Can this doc say a little about how \"git prune\" handles these files?\n>> In particular, does a non cruft pack aware copy of Git (or JGit,\n>> libgit2, etc) do the right thing or does it fight with this mechanism?\n>> If the latter, do we have a repository extension (extensions.*) to\n>> prevent that?\n>\n> I mentioned this in much more detail in [1], but the answer is that the\n> cruft pack looks like any other pack, it just happens to have another\n> metadata file (the .mtimes one) attached to it. So other implementations\n> of Git should treat it as they would any other pack. Like I mentioned in\n> [1], cruft packs were designed with the explicit goal of not requiring a\n> repository extension.\n\nSorry, the above seems like it's answering a different question than I\nasked.  The doc in Documentation/technical/ seems like a natural place\nto describe what semantics the new .mtimes file has, and I didn't find\nthat there.  Is there a different piece of documentation I should have\nbeen looking at?\n\nCan you tell me a little more about why we would want _not_ to have a\nrepository format extension?  To me, it seems like a fairly simple\naddition that would drastically reduce the cognitive overload for\npeople considering making use of this feature.\n\n[...]\n> The key advantage of cruft packs is that you can expire unreachable\n> objects in piecemeal while still retaining the benefit of being able to\n> de-duplicate cruft objects and store them packed against each other.\n\nCan you say a little more about this?  My experience with the similar\nfeature in JGit is that it has been helpful to be able to expire a\ncruft pack altogether; since objects that became reachable around the\nsame time get packed at the same time, it's not obvious to me what\nbenefit this extra piecemeal capability brings.\n\nThat doesn't mean the benefit doesn't exist, just that it seems like\nthere's a piece of context I'm still missing.\n\n>>> +Notable alternatives to this design include:\n>>\n>> This doesn't mention the approach described in\n>> hash-function-transition.txt (and that's already implemented and has\n>> been in use for many years in JGit's DfsRepository).  Does that mean\n>> you aren't aware of it?\n>\n> Implementing the UNREACHABLE_GARBAGE concept from\n> hash-function-transition.txt in cruft pack-terms would be equivalent to\n> not writing the mtimes file at all. This follows from the fact that a\n> pre-cruft packs implementation of Git considers a packed object's mtime\n> to be the same as the pack it's contained in. (I'm deliberately\n> avoiding any details from the h-f-t document regarding re-writing\n> objects contained in a garbage pack here, since this is separate from\n> the pack structure itself (and could easily be implemented on top of\n> cruft packs)).\n>\n> So I'm not sure what the alternative we'd list would be, since it\n> removes the key feature of the design of cruft packs.\n\nSorry, I don't understand this answer either.  Do you mean to say that\nJGit's DfsRepository does not in fact have a cruft packs like feature\nthat is live in the wild?  Or that that feature is equivalent to not\nhaving such a feature?  Or something else?\n\nTo be clear, I'm not trying to say that that's superior to what you've\nproposed here --- only that documenting the comparison would be\nuseful.\n\nPuzzled,\nJonathan\n"},{"id":"451932","messageId":"YjpHbaBspUasDdEy@nand.local","threadId":"56996","inReplyTo":"YjpDbHmKY9XA2p0K@google.com","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-22T22:02:21Z","receivedAt":"2022-03-22T22:02:31Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Mar 22, 2022 at 02:45:16PM -0700, Jonathan Nieder wrote:\n> Hi,\n>\n> Taylor Blau wrote:\n> > On Mon, Mar 07, 2022 at 10:03:35AM -0800, Jonathan Nieder wrote:\n>\n> >> Sorry for the very slow review!  I've mentioned a few times that this\n> >> overlaps in interesting ways with the gc mechanism described in\n> >> hash-function-transition.txt, so I'd like to compare and see how they\n> >> interact.\n> >\n> > Sorry for my equally-slow reply ;). I was on vacation last week and\n> > wasn't following the list closely.\n>\n> No problem --- thanks for getting back to me.\n>\n> [...]\n> > (After re-reading what you wrote and my response, I think we are saying\n> > the exact same thing, but it doesn't hurt to think aloud).\n>\n> Great.  Can the doc cover this?  I think it would be helpful to make\n> that easy to find for others with similar questions.\n\nI believe the doc covers this already, see the paragraph beginning with\n\"Unreachable objects aren't removed immediately...\".\n\n> >> Can this doc say a little about how \"git prune\" handles these files?\n> >> In particular, does a non cruft pack aware copy of Git (or JGit,\n> >> libgit2, etc) do the right thing or does it fight with this mechanism?\n> >> If the latter, do we have a repository extension (extensions.*) to\n> >> prevent that?\n> >\n> > I mentioned this in much more detail in [1], but the answer is that the\n> > cruft pack looks like any other pack, it just happens to have another\n> > metadata file (the .mtimes one) attached to it. So other implementations\n> > of Git should treat it as they would any other pack. Like I mentioned in\n> > [1], cruft packs were designed with the explicit goal of not requiring a\n> > repository extension.\n>\n> Sorry, the above seems like it's answering a different question than I\n> asked.  The doc in Documentation/technical/ seems like a natural place\n> to describe what semantics the new .mtimes file has, and I didn't find\n> that there.  Is there a different piece of documentation I should have\n> been looking at?\n\nAre you looking for a technical description of the mtimes file? If so,\nthere is a section in Documentation/technical/pack-format.txt (added in\n\"pack-mtimes: support reading .mtimes files\") that explains this.\n\n> Can you tell me a little more about why we would want _not_ to have a\n> repository format extension?  To me, it seems like a fairly simple\n> addition that would drastically reduce the cognitive overload for\n> people considering making use of this feature.\n\nThere is no reason to prevent a pre-cruft packs version of Git from\nreading/writing a repository that uses cruft packs, since the two\nversions will still function as normal. Since there's no need to prevent\nthe old version from interacting with a repository that has cruft packs,\nwe wouldn't want to enforce an unnecessary boundary with an extension.\n\n> [...]\n> > The key advantage of cruft packs is that you can expire unreachable\n> > objects in piecemeal while still retaining the benefit of being able to\n> > de-duplicate cruft objects and store them packed against each other.\n>\n> Can you say a little more about this?  My experience with the similar\n> feature in JGit is that it has been helpful to be able to expire a\n> cruft pack altogether; since objects that became reachable around the\n> same time get packed at the same time, it's not obvious to me what\n> benefit this extra piecemeal capability brings.\n>\n> That doesn't mean the benefit doesn't exist, just that it seems like\n> there's a piece of context I'm still missing.\n\nExpiring objects in piecemeal is somewhat interesting, but I think I was\nreaching a little too far when I said it was the \"key benefit\". It does\nhave some nice properties, like being able to store cruft objects as\ndeltas against other cruft objects which might get pruned at a different\ntime (though, of course, you'll need to re-delta them in the case you do\nprune an object which is the base of another cruft object).\n\nBut the issue with having multiple cruft packs is that the semantics get\nsignificantly more complicated. E.g., if you have an object represented\nin multiple cruft packs, which mtime do you use? If you want to prune\nit, you suddenly may have many packs you need to update and keep track\nof.\n\n> >>> +Notable alternatives to this design include:\n> >>\n> >> This doesn't mention the approach described in\n> >> hash-function-transition.txt (and that's already implemented and has\n> >> been in use for many years in JGit's DfsRepository).  Does that mean\n> >> you aren't aware of it?\n> >\n> > Implementing the UNREACHABLE_GARBAGE concept from\n> > hash-function-transition.txt in cruft pack-terms would be equivalent to\n> > not writing the mtimes file at all. This follows from the fact that a\n> > pre-cruft packs implementation of Git considers a packed object's mtime\n> > to be the same as the pack it's contained in. (I'm deliberately\n> > avoiding any details from the h-f-t document regarding re-writing\n> > objects contained in a garbage pack here, since this is separate from\n> > the pack structure itself (and could easily be implemented on top of\n> > cruft packs)).\n> >\n> > So I'm not sure what the alternative we'd list would be, since it\n> > removes the key feature of the design of cruft packs.\n>\n> Sorry, I don't understand this answer either.  Do you mean to say that\n> JGit's DfsRepository does not in fact have a cruft packs like feature\n> that is live in the wild?  Or that that feature is equivalent to not\n> having such a feature?  Or something else?\n>\n> To be clear, I'm not trying to say that that's superior to what you've\n> proposed here --- only that documenting the comparison would be\n> useful.\n\nI'm not familiar enough with JGit (or its DfsRepository class) to know\nhow to answer this. I was comparing cruft packs to the\nUNREACHABLE_GARBAGE concept mentioned in the hash-function-transition\ndoc, and noting the differences there.\n\nThanks,\nTaylor\n"},{"id":"451933","messageId":"YjpWFZ95OL7joFa4@google.com","threadId":"56996","inReplyTo":"YjpHbaBspUasDdEy@nand.local","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2022-03-22T23:04:53Z","receivedAt":"2022-03-22T23:04:59Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nTaylor Blau wrote:\n> On Tue, Mar 22, 2022 at 02:45:16PM -0700, Jonathan Nieder wrote:\n\n>> Great.  Can the doc cover this?  I think it would be helpful to make\n>> that easy to find for others with similar questions.\n>\n> I believe the doc covers this already, see the paragraph beginning with\n> \"Unreachable objects aren't removed immediately...\".\n\nThanks.  I just reread that section and it didn't say anything obvious\nabout the race that continues to exist and whether cruft packs address\nit.\n\n[...]\n>> Sorry, the above seems like it's answering a different question than I\n>> asked.  The doc in Documentation/technical/ seems like a natural place\n>> to describe what semantics the new .mtimes file has, and I didn't find\n>> that there.  Is there a different piece of documentation I should have\n>> been looking at?\n>\n> Are you looking for a technical description of the mtimes file? If so,\n> there is a section in Documentation/technical/pack-format.txt (added in\n> \"pack-mtimes: support reading .mtimes files\") that explains this.\n\nI see --- is the idea that cruft-packs.txt means to refer to\npack-format.txt for the details, and cruft-packs is an overview of some\nother, non-detail aspects?\n\nI just checked pack-format.txt and it didn't describe the semantics\n(what a Git implementation is expected to do when it sees an mtimes\nfile).  For example, in Documentation/technical/cruft-packs.txt, the\nkind of thing I'd expect to see is\n\n- what does an mtime value in the mtimes file represent?  When is it\n  meant to be updated?\n- what guarantees are present about when an object is safe to be\n  pruned?\n\n[...]\n>> Can you tell me a little more about why we would want _not_ to have a\n>> repository format extension?  To me, it seems like a fairly simple\n>> addition that would drastically reduce the cognitive overload for\n>> people considering making use of this feature.\n>\n>There is no reason to prevent a pre-cruft packs version of Git from\n> reading/writing a repository that uses cruft packs, since the two\n> versions will still function as normal. Since there's no need to prevent\n> the old version from interacting with a repository that has cruft packs,\n> we wouldn't want to enforce an unnecessary boundary with an extension.\n\nDoes \"function as normal\" include in repository maintenance operations\nlike \"git maintenance\", \"git gc\", and \"git prune\"?  If so, this seems\nlike something very useful to describe in the cruft-packs.txt\ndocument, since what happens when we bounce back and forth between old\nand new versions of Git operating on the same NFS mounted repository\nwould not be obvious without such a discussion.\n\nI'm still interested in the _downsides_ of using a repository format\nextension.  \"There is no reason\" is not a downside, unless you mean\nthat it requires adding a line of code. :)  The main downside I can\nimagine is that it prevents accessing the repository _that has enabled\nthis feature_ with an older version of Git, but I (perhaps due to a\nfailure of imagination) haven't put two and two together yet about\nwhen I would want to do so.\n\n[...]\n> Expiring objects in piecemeal is somewhat interesting, but I think I was\n> reaching a little too far when I said it was the \"key benefit\". It does\n> have some nice properties, like being able to store cruft objects as\n> deltas against other cruft objects which might get pruned at a different\n> time (though, of course, you'll need to re-delta them in the case you do\n> prune an object which is the base of another cruft object).\n>\n> But the issue with having multiple cruft packs is that the semantics get\n> significantly more complicated. E.g., if you have an object represented\n> in multiple cruft packs, which mtime do you use? If you want to prune\n> it, you suddenly may have many packs you need to update and keep track\n> of.\n\nThanks for this explanation.  In hash-function-transition.txt, I see\n\n\t\"git gc\" currently expels any unreachable objects it encounters in\n\tpack files to loose objects in an attempt to prevent a race when\n\tpruning them (in case another process is simultaneously writing a new\n\tobject that refers to the about-to-be-deleted object). This leads to\n\tan explosion in the number of loose objects present and disk space\n\tusage due to the objects in delta form being replaced with independent\n\tloose objects.  Worse, the race is still present for loose objects.\n\n\tInstead, \"git gc\" will need to move unreachable objects to a new\n\tpackfile marked as UNREACHABLE_GARBAGE (using the PSRC field; see\n\tbelow). To avoid the race when writing new objects referring to an\n\tabout-to-be-deleted object, code paths that write new objects will\n\tneed to copy any objects from UNREACHABLE_GARBAGE packs that they\n\trefer to new, non-UNREACHABLE_GARBAGE packs (or loose objects).\n\tUNREACHABLE_GARBAGE are then safe to delete if their creation time (as\n\tindicated by the file's mtime) is long enough ago.\n\n\tTo avoid a proliferation of UNREACHABLE_GARBAGE packs, they can be\n\tcombined under certain circumstances. [etc]\n\nSo the proposal there is that the file mtime for an UNREACHABLE_GARBAGE\npack refers to when that pack was written and governs when that pack\ncan be deleted.  If an object is present in multiple packs, then newer\npacks with the object have a newer mtime and thus cause the object to\nbe kept around for longer.\n\n[...]\n>> Sorry, I don't understand this answer either.  Do you mean to say that\n>> JGit's DfsRepository does not in fact have a cruft packs like feature\n>> that is live in the wild?  Or that that feature is equivalent to not\n>> having such a feature?  Or something else?\n>>\n>> To be clear, I'm not trying to say that that's superior to what you've\n>> proposed here --- only that documenting the comparison would be\n>> useful.\n>\n> I'm not familiar enough with JGit (or its DfsRepository class) to know\n> how to answer this. I was comparing cruft packs to the\n> UNREACHABLE_GARBAGE concept mentioned in the hash-function-transition\n> doc, and noting the differences there.\n\nThanks.  I think there's some implied feedback about the documentation\nof UNREACHABLE_GARBAGE there, because if I understand then you're\nsaying that it does not describe maintaining cruft packs.  Perhaps a\npointer to the particular sentence that led you to that conclusion\nwould help.\n\nSincerely,\nJonathan\n"},{"id":"451936","messageId":"Yjpxd8qhwnAIJJma@nand.local","threadId":"56996","inReplyTo":"YjpWFZ95OL7joFa4@google.com","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-23T01:01:43Z","receivedAt":"2022-03-23T01:01:49Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Mar 22, 2022 at 04:04:53PM -0700, Jonathan Nieder wrote:\n> Hi,\n>\n> Taylor Blau wrote:\n> > On Tue, Mar 22, 2022 at 02:45:16PM -0700, Jonathan Nieder wrote:\n>\n> >> Great.  Can the doc cover this?  I think it would be helpful to make\n> >> that easy to find for others with similar questions.\n> >\n> > I believe the doc covers this already, see the paragraph beginning with\n> > \"Unreachable objects aren't removed immediately...\".\n>\n> Thanks.  I just reread that section and it didn't say anything obvious\n> about the race that continues to exist and whether cruft packs address\n> it.\n\nYeah, there isn't an explicit \"and cruft packs addresses the loose\nobject explosion but does not address the race\" sentence. I'm not\nopposed to adding something like that to clarify (though TBH, I would\nrather do it as a clean-up on top rather than send out a bazillion\nmostly-unchanged patches).\n\n> [...]\n> >> Sorry, the above seems like it's answering a different question than I\n> >> asked.  The doc in Documentation/technical/ seems like a natural place\n> >> to describe what semantics the new .mtimes file has, and I didn't find\n> >> that there.  Is there a different piece of documentation I should have\n> >> been looking at?\n> >\n> > Are you looking for a technical description of the mtimes file? If so,\n> > there is a section in Documentation/technical/pack-format.txt (added in\n> > \"pack-mtimes: support reading .mtimes files\") that explains this.\n>\n> I see --- is the idea that cruft-packs.txt means to refer to\n> pack-format.txt for the details, and cruft-packs is an overview of some\n> other, non-detail aspects?\n>\n> I just checked pack-format.txt and it didn't describe the semantics\n> (what a Git implementation is expected to do when it sees an mtimes\n> file).  For example, in Documentation/technical/cruft-packs.txt, the\n> kind of thing I'd expect to see is\n\nRight; I have always considered the files in Documentation/technical to\nprimarily be about the file format itself.\n\n> - what does an mtime value in the mtimes file represent?  When is it\n>   meant to be updated?\n> - what guarantees are present about when an object is safe to be\n>   pruned?\n\nThe cruft-packs.txt document covers these, though I think somewhat\nimplicitly. Again, I'm not opposed to more clarification, but I again\nwould like to do so on top.\n\nI think many of these are discussed within the threads above, but to\nanswer your questions in order:\n\n  - The mtime of an object in a cruft pack represents the last time that\n    object was known to be reachable, and it's updated when generating a\n    cruft pack or pruning.\n\n  - The same guarantees are made in the cruft pack case as in the\n    non-cruft case (i.e., \"none\", and so a grace period is recommended).\n\n> [...]\n> >> Can you tell me a little more about why we would want _not_ to have a\n> >> repository format extension?  To me, it seems like a fairly simple\n> >> addition that would drastically reduce the cognitive overload for\n> >> people considering making use of this feature.\n> >\n> >There is no reason to prevent a pre-cruft packs version of Git from\n> > reading/writing a repository that uses cruft packs, since the two\n> > versions will still function as normal. Since there's no need to prevent\n> > the old version from interacting with a repository that has cruft packs,\n> > we wouldn't want to enforce an unnecessary boundary with an extension.\n>\n> Does \"function as normal\" include in repository maintenance operations\n> like \"git maintenance\", \"git gc\", and \"git prune\"?  If so, this seems\n> like something very useful to describe in the cruft-packs.txt\n> document, since what happens when we bounce back and forth between old\n> and new versions of Git operating on the same NFS mounted repository\n> would not be obvious without such a discussion.\n\nYes, all of those commands will simply ignore the .mtimes file and treat\nthe unreachable objects as normal (where \"normal\" means in the exact\nsame way as they currently do without cruft packs). I think adding a\nsection that summarizes our discussion would be useful.\n\n> I'm still interested in the _downsides_ of using a repository format\n> extension.  \"There is no reason\" is not a downside, unless you mean\n> that it requires adding a line of code. :)  The main downside I can\n> imagine is that it prevents accessing the repository _that has enabled\n> this feature_ with an older version of Git, but I (perhaps due to a\n> failure of imagination) haven't put two and two together yet about\n> when I would want to do so.\n\nSorry for not being clear; I meant: \"There is no reason [to prohibit\ntwo versions of Git from interacting with each other when they are\ncompatible to do so]\".\n\n> [...]\n> > Expiring objects in piecemeal is somewhat interesting, but I think I was\n> > reaching a little too far when I said it was the \"key benefit\". It does\n> > have some nice properties, like being able to store cruft objects as\n> > deltas against other cruft objects which might get pruned at a different\n> > time (though, of course, you'll need to re-delta them in the case you do\n> > prune an object which is the base of another cruft object).\n> >\n> > But the issue with having multiple cruft packs is that the semantics get\n> > significantly more complicated. E.g., if you have an object represented\n> > in multiple cruft packs, which mtime do you use? If you want to prune\n> > it, you suddenly may have many packs you need to update and keep track\n> > of.\n>\n> Thanks for this explanation.  In hash-function-transition.txt, I see\n>\n> \t\"git gc\" currently expels any unreachable objects it encounters in\n> \tpack files to loose objects in an attempt to prevent a race when\n> \tpruning them (in case another process is simultaneously writing a new\n> \tobject that refers to the about-to-be-deleted object). This leads to\n> \tan explosion in the number of loose objects present and disk space\n> \tusage due to the objects in delta form being replaced with independent\n> \tloose objects.  Worse, the race is still present for loose objects.\n>\n> \tInstead, \"git gc\" will need to move unreachable objects to a new\n> \tpackfile marked as UNREACHABLE_GARBAGE (using the PSRC field; see\n> \tbelow). To avoid the race when writing new objects referring to an\n> \tabout-to-be-deleted object, code paths that write new objects will\n> \tneed to copy any objects from UNREACHABLE_GARBAGE packs that they\n> \trefer to new, non-UNREACHABLE_GARBAGE packs (or loose objects).\n> \tUNREACHABLE_GARBAGE are then safe to delete if their creation time (as\n> \tindicated by the file's mtime) is long enough ago.\n>\n> \tTo avoid a proliferation of UNREACHABLE_GARBAGE packs, they can be\n> \tcombined under certain circumstances. [etc]\n>\n> So the proposal there is that the file mtime for an UNREACHABLE_GARBAGE\n> pack refers to when that pack was written and governs when that pack\n> can be deleted.  If an object is present in multiple packs, then newer\n> packs with the object have a newer mtime and thus cause the object to\n> be kept around for longer.\n\nThat matches my understanding.\n\nThanks,\nTaylor\n"},{"id":"452491","messageId":"YkICkpttOujOKeT3@nand.local","threadId":"56996","inReplyTo":"Yjpxd8qhwnAIJJma@nand.local","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-28T18:46:42Z","receivedAt":"2022-03-28T18:46:48Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Mar 22, 2022 at 09:01:43PM -0400, Taylor Blau wrote:\n> > >> Can you tell me a little more about why we would want _not_ to have a\n> > >> repository format extension?  To me, it seems like a fairly simple\n> > >> addition that would drastically reduce the cognitive overload for\n> > >> people considering making use of this feature.\n> > >\n> > >There is no reason to prevent a pre-cruft packs version of Git from\n> > > reading/writing a repository that uses cruft packs, since the two\n> > > versions will still function as normal. Since there's no need to prevent\n> > > the old version from interacting with a repository that has cruft packs,\n> > > we wouldn't want to enforce an unnecessary boundary with an extension.\n> >\n> > Does \"function as normal\" include in repository maintenance operations\n> > like \"git maintenance\", \"git gc\", and \"git prune\"?  If so, this seems\n> > like something very useful to describe in the cruft-packs.txt\n> > document, since what happens when we bounce back and forth between old\n> > and new versions of Git operating on the same NFS mounted repository\n> > would not be obvious without such a discussion.\n>\n> Yes, all of those commands will simply ignore the .mtimes file and treat\n> the unreachable objects as normal (where \"normal\" means in the exact\n> same way as they currently do without cruft packs). I think adding a\n> section that summarizes our discussion would be useful.\n>\n> > I'm still interested in the _downsides_ of using a repository format\n> > extension.  \"There is no reason\" is not a downside, unless you mean\n> > that it requires adding a line of code. :)  The main downside I can\n> > imagine is that it prevents accessing the repository _that has enabled\n> > this feature_ with an older version of Git, but I (perhaps due to a\n> > failure of imagination) haven't put two and two together yet about\n> > when I would want to do so.\n>\n> Sorry for not being clear; I meant: \"There is no reason [to prohibit\n> two versions of Git from interacting with each other when they are\n> compatible to do so]\".\n\nJonathan, myself, and others discussed this extensively in today's\nstandup.\n\nTo summarize Jonathan's point (as I think I severely misunderstood it\nbefore), if two writers are repacking a repository with unreachable\nobjects. The following can happen:\n\n  - $NEWGIT packs the repository and writes a cruft pack and .mtimes\n    file.\n\n  - $OLDGIT packs the repository, exploding unreachable objects from the\n    cruft pack as loose, setting their mtimes to \"now\".\n\nThis causes the repository to lose information about the unreachable\nmtimes, which would cause the repository to never prune objects (except\nfor when`--unpack-unreachable=now` is passed).\n\nOne approach (that Jonathan suggested) is to prevent the above situation\nby introducing a format extension, so $OLDGIT could not touch the\nrepository. But this comes at a (in my view, significant) cost which is\nthat $OLDGIT can't touch the repository _at all_. An extension would be\ndesirable if cross-version interaction resulted in repository\ncorruption, but this scenario does not lead to corruption at all.\n\nAnother approach (courtesy Stolee, in an off-list discussion) is that we\ncould introduce an optional extension available as an opt-in to prevent\nolder versions of Git from interacting in a repository that contains\ncruft packs, but is not required to write them.\n\nA third approach (and probably my preferred direction) is to indicate\nclearly via a combination of updates to Documentation/cruft-packs.txt\nand the release notes that say something along the lines of:\n\n    If you use are repacking a repository using both a pre- and\n    post-cruft packs version of Git, please be aware that you will lose\n    information about the mtimes of unreachable objects.\n\nI imagine that would probably be sufficient, but we could also introduce\nthe opt-in extension as an easy alternative to avoid forcing an upgrade\nof Git.\n\nThanks,\nTaylor\n"},{"id":"452512","messageId":"xmqq8rst23w0.fsf@gitster.g","threadId":"56996","inReplyTo":"YkICkpttOujOKeT3@nand.local","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-28T20:55:43Z","receivedAt":"2022-03-28T20:55:51Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> To summarize Jonathan's point (as I think I severely misunderstood it\n> before), if two writers are repacking a repository with unreachable\n> objects. The following can happen:\n>\n>   - $NEWGIT packs the repository and writes a cruft pack and .mtimes\n>     file.\n>\n>   - $OLDGIT packs the repository, exploding unreachable objects from the\n>     cruft pack as loose, setting their mtimes to \"now\".\n\nAnd if these repeat, alternating new and old versions of Git, we\nwill keep refreshing the unreachable objects' mtimes forever.\n\nBut once you stop using old versions of Git, perhaps in 3 release\ncycles or so, we'll eventually be able to purge them, right?\n\n> One approach (that Jonathan suggested) is to prevent the above situation\n> by introducing a format extension, so $OLDGIT could not touch the\n> repository. But this comes at a (in my view, significant) cost which is\n> that $OLDGIT can't touch the repository _at all_. An extension would be\n> desirable if cross-version interaction resulted in repository\n> corruption, but this scenario does not lead to corruption at all.\n\nA repository may not be in a healthy state, when tons of unreachable\nobjects stay around forever, but it probably is a bit too harsh to\ncall it \"corrupt\".\n\n> Another approach (courtesy Stolee, in an off-list discussion) is that we\n> could introduce an optional extension available as an opt-in to prevent\n> older versions of Git from interacting in a repository that contains\n> cruft packs, but is not required to write them.\n\nThat smells too magic; let's not go there.\n\n> A third approach (and probably my preferred direction) is to indicate\n> clearly via a combination of updates to Documentation/cruft-packs.txt\n> and the release notes that say something along the lines of:\n>\n>     If you use are repacking a repository using both a pre- and\n>     post-cruft packs version of Git, please be aware that you will lose\n>     information about the mtimes of unreachable objects.\n\nI do not quite see how it helps.  After hearing \"... will lose\ninformation about the mtimes ...\", what concrete action can a user\ntake?  Or a sys-admin?\n\nIt's not like use of cruft-pack is mandatory when you upgrade the\nnew version of Git, right?  Perhaps use of cruft-pack should be\nguarded behind a configuration variable so that users who might want\nto use mixed versions of Git will be protected against accidental\nuse of new version of Git that introduces the forever-renewing\nuntracked objects problem?  \n\nPerhaps a configuration variable, repack.cruftPackEnabled, that is\nby default disabled, can be used to protect people who do not want\nto get into the \"keep refreshing mtime\" loop from using the cruft\npacks by mistake?  repack.cruftPackEnabled can probably be part of\nthe \"experimental\" feature set, if we think it is the direction in\nthe future.\n\n\n\n\n"},{"id":"452514","messageId":"YkIm7lnQsUT0JnvS@nand.local","threadId":"56996","inReplyTo":"xmqq8rst23w0.fsf@gitster.g","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-28T21:21:50Z","receivedAt":"2022-03-28T21:22:01Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Mon, Mar 28, 2022 at 01:55:43PM -0700, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> > To summarize Jonathan's point (as I think I severely misunderstood it\n> > before), if two writers are repacking a repository with unreachable\n> > objects. The following can happen:\n> >\n> >   - $NEWGIT packs the repository and writes a cruft pack and .mtimes\n> >     file.\n> >\n> >   - $OLDGIT packs the repository, exploding unreachable objects from the\n> >     cruft pack as loose, setting their mtimes to \"now\".\n>\n> And if these repeat, alternating new and old versions of Git, we\n> will keep refreshing the unreachable objects' mtimes forever.\n>\n> But once you stop using old versions of Git, perhaps in 3 release\n> cycles or so, we'll eventually be able to purge them, right?\n\nAs soon as all of the repackers understand cruft packs, yes.\n\n> > One approach (that Jonathan suggested) is to prevent the above situation\n> > by introducing a format extension, so $OLDGIT could not touch the\n> > repository. But this comes at a (in my view, significant) cost which is\n> > that $OLDGIT can't touch the repository _at all_. An extension would be\n> > desirable if cross-version interaction resulted in repository\n> > corruption, but this scenario does not lead to corruption at all.\n>\n> A repository may not be in a healthy state, when tons of unreachable\n> objects stay around forever, but it probably is a bit too harsh to\n> call it \"corrupt\".\n\nI agree, though I would note that this is no worse than the situation\ntoday, where unreachable-but-recent objects are already exploded as\nloose can already cause the kinds of issues that this series is designed\nto prevent.\n\n> > Another approach (courtesy Stolee, in an off-list discussion) is that we\n> > could introduce an optional extension available as an opt-in to prevent\n> > older versions of Git from interacting in a repository that contains\n> > cruft packs, but is not required to write them.\n>\n> That smells too magic; let's not go there.\n\nI'm not sure... if we did:\n\n--- 8< ---\n\ndiff --git a/setup.c b/setup.c\nindex 04ce33cdcd..fa54c9baa4 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -565,2 +565,4 @@ static enum extension_result handle_extension(const char *var,\n \t\treturn EXTENSION_OK;\n+\t} else if (!strcmp(ext, \"cruftpacks\")) {\n+\t\treturn EXTENSION_OK;\n \t}\n\n--- >8 ---\n\nbut nothing more, then a hypothetical `extensions.cruftPacks` could be\nused to prevent older writers in a mixed version environment. But if you\ndon't have or care about older versions of Git, you can avoid setting it\naltogether.\n\nThe key bit is that we don't have a check along the lines of \"only allow\nwriting a cruft pack when extensions.cruftPacks\" is set, so it's opt-in\nas far as the new code is concerned.\n\n> > A third approach (and probably my preferred direction) is to indicate\n> > clearly via a combination of updates to Documentation/cruft-packs.txt\n> > and the release notes that say something along the lines of:\n> >\n> >     If you use are repacking a repository using both a pre- and\n> >     post-cruft packs version of Git, please be aware that you will lose\n> >     information about the mtimes of unreachable objects.\n>\n> I do not quite see how it helps.  After hearing \"... will lose\n> information about the mtimes ...\", what concrete action can a user\n> take?  Or a sys-admin?\n>\n> It's not like use of cruft-pack is mandatory when you upgrade the\n> new version of Git, right?  Perhaps use of cruft-pack should be\n> guarded behind a configuration variable so that users who might want\n> to use mixed versions of Git will be protected against accidental\n> use of new version of Git that introduces the forever-renewing\n> untracked objects problem?\n\nI don't think we would have much to offer a user in that case; if the\nmtimes are gone, then I couldn't think of anything to bring them back\noutside of setting them manually.\n\nBut cruft packs are already guarded in two places:\n\n  - `git repack` won't write a cruft pack unless given the `--cruft`\n    flag (i.e., `git repack -A` doesn't suddenly start generating cruft\n    packs upon upgrade).\n\n  - `git gc` won't write cruft packs unless the `gc.cruftPacks`\n    configuration is set, or `--cruft` is given as a flag.\n\nI'd be curious what Jonathan and others think of that approach (which,\nto be clear, is what this series already implements). We could make it\nclear to say:\n\n    If you have mixed versions of Git which both repack a repository\n    (either manually or by auto-GC / background maintenance), consider\n    leaving `gc.cruftPacks` unset and avoiding passing `--cruft` as a\n    command-line argument to `git repack` and `git gc`, since doing so\n    can lead to [...]\n\n> Perhaps a configuration variable, repack.cruftPackEnabled, that is\n> by default disabled, can be used to protect people who do not want\n> to get into the \"keep refreshing mtime\" loop from using the cruft\n> packs by mistake?  repack.cruftPackEnabled can probably be part of\n> the \"experimental\" feature set, if we think it is the direction in\n> the future.\n\nI'd probably want to leave `-A` separate from `--cruft`, since something\nabout setting `repack.cruftPackEnabled` having the effect of causing\n`-A` to produce a cruft pack feels strange to me.\n\nThanks,\nTaylor\n"},{"id":"452576","messageId":"xmqqa6d8yckj.fsf@gitster.g","threadId":"56996","inReplyTo":"YkIm7lnQsUT0JnvS@nand.local","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-29T15:59:24Z","receivedAt":"2022-03-29T15:59:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> I'm not sure... if we did:\n>\n> --- 8< ---\n>\n> diff --git a/setup.c b/setup.c\n> index 04ce33cdcd..fa54c9baa4 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -565,2 +565,4 @@ static enum extension_result handle_extension(const char *var,\n>  \t\treturn EXTENSION_OK;\n> +\t} else if (!strcmp(ext, \"cruftpacks\")) {\n> +\t\treturn EXTENSION_OK;\n>  \t}\n>\n> --- >8 ---\n>\n> but nothing more, then a hypothetical `extensions.cruftPacks` could be\n> used to prevent older writers in a mixed version environment. But if you\n> don't have or care about older versions of Git, you can avoid setting it\n> altogether.\n\nSmells like \"unsafe by default, but you can opt into safety\", which\nis backwards, isn't it?\n\n>> I do not quite see how it helps.  After hearing \"... will lose\n>> information about the mtimes ...\", what concrete action can a user\n>> take?  Or a sys-admin?\n>>\n>> It's not like use of cruft-pack is mandatory when you upgrade the\n>> new version of Git, right?  Perhaps use of cruft-pack should be\n>> guarded behind a configuration variable so that users who might want\n>> to use mixed versions of Git will be protected against accidental\n>> use of new version of Git that introduces the forever-renewing\n>> untracked objects problem?\n>\n> I don't think we would have much to offer a user in that case; if the\n> mtimes are gone, then I couldn't think of anything to bring them back\n> outside of setting them manually.\n\nYes, so rambling about losing mtimes in documentation or release\nnotes would not help users all that much.  Let's not do that.\n\n> But cruft packs are already guarded in two places:\n>\n>   - `git repack` won't write a cruft pack unless given the `--cruft`\n>     flag (i.e., `git repack -A` doesn't suddenly start generating cruft\n>     packs upon upgrade).\n>\n>   - `git gc` won't write cruft packs unless the `gc.cruftPacks`\n>     configuration is set, or `--cruft` is given as a flag.\n\nHmph, OK.  So individuals can sort-of protect from hurting\nthemselves by refraining from running these with --cruft or writing\n--cruft in their maintenance scripts.  An organization that wants to\nlet the more adventurous types to early opt-in can prepare two\nversions of the maintenance scripts they distribute to their users,\none with and the other without --cruft, and use the mechanism they\nuse for gradual rollouts to control the population.  Perhaps that\nwould make sufficient protection?  I dunno.\n\nJonathan, what do you think?\n\n> I'd be curious what Jonathan and others think of that approach (which,\n> to be clear, is what this series already implements). We could make it\n> clear to say:\n>\n>     If you have mixed versions of Git which both repack a repository\n>     (either manually or by auto-GC / background maintenance), consider\n>     leaving `gc.cruftPacks` unset and avoiding passing `--cruft` as a\n>     command-line argument to `git repack` and `git gc`, since doing so\n>     can lead to [...]\n\nThat message is (depending on what comes in [...]) much more helpful\nthan just throwing a word \"mtime\" out and letting the reader figure\nout the rest ;-)\n"},{"id":"452633","messageId":"YkO/P3xgGYmhAz2O@nand.local","threadId":"56996","inReplyTo":"xmqqa6d8yckj.fsf@gitster.g","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-30T02:23:59Z","receivedAt":"2022-03-30T02:24:05Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Mar 29, 2022 at 08:59:24AM -0700, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> > I'm not sure... if we did:\n> >\n> > --- 8< ---\n> >\n> > diff --git a/setup.c b/setup.c\n> > index 04ce33cdcd..fa54c9baa4 100644\n> > --- a/setup.c\n> > +++ b/setup.c\n> > @@ -565,2 +565,4 @@ static enum extension_result handle_extension(const char *var,\n> >  \t\treturn EXTENSION_OK;\n> > +\t} else if (!strcmp(ext, \"cruftpacks\")) {\n> > +\t\treturn EXTENSION_OK;\n> >  \t}\n> >\n> > --- >8 ---\n> >\n> > but nothing more, then a hypothetical `extensions.cruftPacks` could be\n> > used to prevent older writers in a mixed version environment. But if you\n> > don't have or care about older versions of Git, you can avoid setting it\n> > altogether.\n>\n> Smells like \"unsafe by default, but you can opt into safety\", which\n> is backwards, isn't it?\n\nI see it a little differently. The default (not writing cruft packs at\nall) is safe, even in a mixed-version environment. If a user (a) wants\nto use cruft packs, and (b) has older versions of Git also gc'ing the\nrepository, and (c) can't get rid of them, _then_ an opt-in extension\nwould make it impossible for those older versions to interact with the\nrepository.\n\nI still can't shake the feeling that this is a pretty fringe and\ntiming-dependent scenario, which at worst keeps too many unreachable\nobjects around.\n\nBut I think this in conjunction with the already opt-in nature of cruft\npacks would be a nice way to create safeguards for the situation\nJonathan described. There may be a simpler way, but I'm not sure I see\nit (i.e., if you control whether or not `--cruft` is passed when\ndoing maintenance with newer versions of Git, but not whether older\nversions are running around doing their own maintenance, then an\nextension would be necessary to lock the old versions out).\n\n> > But cruft packs are already guarded in two places:\n> >\n> >   - `git repack` won't write a cruft pack unless given the `--cruft`\n> >     flag (i.e., `git repack -A` doesn't suddenly start generating cruft\n> >     packs upon upgrade).\n> >\n> >   - `git gc` won't write cruft packs unless the `gc.cruftPacks`\n> >     configuration is set, or `--cruft` is given as a flag.\n>\n> Hmph, OK.  So individuals can sort-of protect from hurting\n> themselves by refraining from running these with --cruft or writing\n> --cruft in their maintenance scripts.  An organization that wants to\n> let the more adventurous types to early opt-in can prepare two\n> versions of the maintenance scripts they distribute to their users,\n> one with and the other without --cruft, and use the mechanism they\n> use for gradual rollouts to control the population.  Perhaps that\n> would make sufficient protection?  I dunno.\n>\n> Jonathan, what do you think?\n\nI'm confused: if newer versions of Git are writing cruft packs, then\nhaving the older versions gc'ing in the same repository runs into the\nsame scenario Jonathan originally describes.\n\nThe thing I think Jonathan seeks to prevent is older versions of Git\ngc'ing a repo that has cruft packs. I think I may need you to clarify a\nlittle, sorry :-(.\n\n> > I'd be curious what Jonathan and others think of that approach (which,\n> > to be clear, is what this series already implements). We could make it\n> > clear to say:\n> >\n> >     If you have mixed versions of Git which both repack a repository\n> >     (either manually or by auto-GC / background maintenance), consider\n> >     leaving `gc.cruftPacks` unset and avoiding passing `--cruft` as a\n> >     command-line argument to `git repack` and `git gc`, since doing so\n> >     can lead to [...]\n>\n> That message is (depending on what comes in [...]) much more helpful\n> than just throwing a word \"mtime\" out and letting the reader figure\n> out the rest ;-)\n\nYes, totally agreed.\n\nThanks,\nTaylor\n"},{"id":"452664","messageId":"xmqq1qyjblxp.fsf@gitster.g","threadId":"56996","inReplyTo":"YkO/P3xgGYmhAz2O@nand.local","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-30T13:37:54Z","receivedAt":"2022-03-30T13:38:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> The thing I think Jonathan seeks to prevent is older versions of Git\n> gc'ing a repo that has cruft packs. I think I may need you to clarify a\n> little, sorry :-(.\n\nBy making controlled rollout of the use of \"--cruft\" option (and the\nassumption here is that a large organization setting people do not\nmanually say \"gc --cruft\", and they can ship their maintenance\nscripts that may be run via cron or whatever with and without\n\"--cruft\"), you can control the number of repositories that can\npotentially see older versions of Git running gc on with cruft\npacks.  Those users, for whom it is not their turn to start using\n\"--cruft\" enabled version of the script, will not have cruft packs,\nso it does not matter if they keep an older version of Git somewhere\nhidden in a hermetic build of an IDE that bundles Git and gc kicks\nin for them.\n\n"},{"id":"452676","messageId":"YkSTvr42wFVAgHvp@nand.local","threadId":"56996","inReplyTo":"xmqq1qyjblxp.fsf@gitster.g","subject":"Re: [PATCH v3 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-03-30T17:30:38Z","receivedAt":"2022-03-30T17:30:47Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Mar 30, 2022 at 06:37:54AM -0700, Junio C Hamano wrote:\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n> > The thing I think Jonathan seeks to prevent is older versions of Git\n> > gc'ing a repo that has cruft packs. I think I may need you to clarify a\n> > little, sorry :-(.\n>\n> By making controlled rollout of the use of \"--cruft\" option (and the\n> assumption here is that a large organization setting people do not\n> manually say \"gc --cruft\", and they can ship their maintenance\n> scripts that may be run via cron or whatever with and without\n> \"--cruft\"), you can control the number of repositories that can\n> potentially see older versions of Git running gc on with cruft\n> packs.  Those users, for whom it is not their turn to start using\n> \"--cruft\" enabled version of the script, will not have cruft packs,\n> so it does not matter if they keep an older version of Git somewhere\n> hidden in a hermetic build of an IDE that bundles Git and gc kicks\n> in for them.\n\nAhh, OK. Thanks for explaining: this is what I was pretty sure you\nmeant, but I wanted to make sure before agreeing to it.\n\nYes, this solution amounts to: \"if you have mixed-versions of Git\nmutually gc'ing a repository, then use the same rollout method used for\ncontrolling Git itself to guard when to start creating cruft packs\".\n\nI would be very eager to hear if this works for Jonathan's case. It\nshould do the trick, I'd think.\n\nThanks,\nTaylor\n"},{"id":"455432","messageId":"cover.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH v4 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:10:26Z","receivedAt":"2022-05-18T23:10:35Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Here is another reroll of my series to implement \"cruft packs\", which is based\non the v2.36 tree, and incorporates feedback from the discussion we had about\nmixed-version GCs with cruft packs in [1].\n\nThe changes here are limited to:\n\n  - a cautionary note in Documentation/technical/cruft-packs.txt\n    describing the potential interaction between pruning GCs across pre-\n    and post-cruft pack versions of Git, as discussed towards the bottom\n    of [2]\n\n  - updating the `finalize_hashfile()` calls for writing `.mtimes` files\n    to indicate that they are `FSYNC_COMPONENT_PACK_METADATA`, since the\n    original version of this series predates the fine-grained fsync\n    configuration in 2.36.\n\nAs always, a range-diff is below. Thanks in advance for taking another\nlook!\n\n[1]: https://lore.kernel.org/git/YiZI99yeijQe5Jaq@google.com/\n[2]: https://lore.kernel.org/git/YkIm7lnQsUT0JnvS@nand.local/\n\nTaylor Blau (17):\n  Documentation/technical: add cruft-packs.txt\n  pack-mtimes: support reading .mtimes files\n  pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'\n  chunk-format.h: extract oid_version()\n  pack-mtimes: support writing pack .mtimes files\n  t/helper: add 'pack-mtimes' test-tool\n  builtin/pack-objects.c: return from create_object_entry()\n  builtin/pack-objects.c: --cruft without expiration\n  reachable: add options to add_unseen_recent_objects_to_traversal\n  reachable: report precise timestamps from objects in cruft packs\n  builtin/pack-objects.c: --cruft with expiration\n  builtin/repack.c: support generating a cruft pack\n  builtin/repack.c: allow configuring cruft pack generation\n  builtin/repack.c: use named flags for existing_packs\n  builtin/repack.c: add cruft packs to MIDX during geometric repack\n  builtin/gc.c: conditionally avoid pruning objects via loose\n  sha1-file.c: don't freshen cruft packs\n\n Documentation/Makefile                  |   1 +\n Documentation/config/gc.txt             |  21 +-\n Documentation/config/repack.txt         |   9 +\n Documentation/git-gc.txt                |   5 +\n Documentation/git-pack-objects.txt      |  30 +\n Documentation/git-repack.txt            |  11 +\n Documentation/technical/cruft-packs.txt | 123 ++++\n Documentation/technical/pack-format.txt |  19 +\n Makefile                                |   2 +\n builtin/gc.c                            |  10 +-\n builtin/pack-objects.c                  | 304 +++++++++-\n builtin/repack.c                        | 181 +++++-\n bulk-checkin.c                          |   2 +-\n chunk-format.c                          |  12 +\n chunk-format.h                          |   3 +\n commit-graph.c                          |  18 +-\n midx.c                                  |  18 +-\n object-file.c                           |   4 +-\n object-store.h                          |   7 +-\n pack-mtimes.c                           | 126 ++++\n pack-mtimes.h                           |  15 +\n pack-objects.c                          |   6 +\n pack-objects.h                          |  25 +\n pack-write.c                            |  93 ++-\n pack.h                                  |   4 +\n packfile.c                              |  19 +-\n reachable.c                             |  58 +-\n reachable.h                             |   9 +-\n t/helper/test-pack-mtimes.c             |  56 ++\n t/helper/test-tool.c                    |   1 +\n t/helper/test-tool.h                    |   1 +\n t/t5329-pack-objects-cruft.sh           | 739 ++++++++++++++++++++++++\n 32 files changed, 1831 insertions(+), 101 deletions(-)\n create mode 100644 Documentation/technical/cruft-packs.txt\n create mode 100644 pack-mtimes.c\n create mode 100644 pack-mtimes.h\n create mode 100644 t/helper/test-pack-mtimes.c\n create mode 100755 t/t5329-pack-objects-cruft.sh\n\nRange-diff against v3:\n 1:  784ee7e0ee !  1:  f494ef7377 Documentation/technical: add cruft-packs.txt\n    @@ Documentation/technical/cruft-packs.txt (new)\n     +It is linkgit:git-gc[1] that is typically responsible for removing expired\n     +unreachable objects.\n     +\n    ++== Caution for mixed-version environments\n    ++\n    ++Repositories that have cruft packs in them will continue to work with any older\n    ++version of Git. Note, however, that previous versions of Git which do not\n    ++understand the `.mtimes` file will use the cruft pack's mtime as the mtime for\n    ++all of the objects in it. In other words, do not expect older (pre-cruft pack)\n    ++versions of Git to interpret or even read the contents of the `.mtimes` file.\n    ++\n    ++Note that having mixed versions of Git GC-ing the same repository can lead to\n    ++unreachable objects never being completely pruned. This can happen under the\n    ++following circumstances:\n    ++\n    ++  - An older version of Git running GC explodes the contents of an existing\n    ++    cruft pack loose, using the cruft pack's mtime.\n    ++  - A newer version running GC collects those loose objects into a cruft pack,\n    ++    where the .mtime file reflects the loose object's actual mtimes, but the\n    ++    cruft pack mtime is \"now\".\n    ++\n    ++Repeating this process will lead to unreachable objects not getting pruned as a\n    ++result of repeatedly resetting the objects' mtimes to the present time.\n    ++\n    ++If you are GC-ing repositories in a mixed version environment, consider omitting\n    ++the `--cruft` option when using linkgit:git-repack[1] and linkgit:git-gc[1], and\n    ++leaving the `gc.cruftPacks` configuration unset until all writers understand\n    ++cruft packs.\n    ++\n     +== Alternatives\n     +\n     +Notable alternatives to this design include:\n 2:  1ec754ad1b =  2:  8f9fd21be9 pack-mtimes: support reading .mtimes files\n 3:  0f5d6d6492 =  3:  cdb21236e1 pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'\n 4:  135a07276b =  4:  1d775f9850 chunk-format.h: extract oid_version()\n 5:  0600503856 !  5:  6172861bd9 pack-mtimes: support writing pack .mtimes files\n    @@ pack-write.c: const char *write_rev_file_order(const char *rev_name,\n     +\tif (adjust_shared_perm(mtimes_name) < 0)\n     +\t\tdie(_(\"failed to make %s readable\"), mtimes_name);\n     +\n    -+\tfinalize_hashfile(f, NULL,\n    ++\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n     +\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE | CSUM_FSYNC);\n     +\n     +\treturn mtimes_name;\n 6:  4780c8437b =  6:  5f9a9a5b7b t/helper: add 'pack-mtimes' test-tool\n 7:  33862a07c9 =  7:  b8a38fe2e4 builtin/pack-objects.c: return from create_object_entry()\n 8:  22705e4887 !  8:  94fe03cc65 builtin/pack-objects.c: --cruft without expiration\n    @@ builtin/pack-objects.c: static int option_parse_unpack_unreachable(const struct\n     +\treturn 0;\n     +}\n     +\n    - int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n    - {\n    - \tint use_internal_rev_list = 0;\n    + struct po_filter_data {\n    + \tunsigned have_revs:1;\n    + \tstruct rev_info revs;\n     @@ builtin/pack-objects.c: int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n      \t\tOPT_CALLBACK_F(0, \"unpack-unreachable\", NULL, N_(\"time\"),\n      \t\t  N_(\"unpack unreachable objects newer than <time>\"),\n    @@ builtin/pack-objects.c: int cmd_pack_objects(int argc, const char **argv, const\n     +\t\tread_cruft_objects();\n      \t} else if (!use_internal_rev_list) {\n      \t\tread_object_list_from_stdin();\n    - \t} else {\n    + \t} else if (pfd.have_revs) {\n     \n      ## object-file.c ##\n     @@ object-file.c: int has_loose_object_nonlocal(const struct object_id *oid)\n    @@ object-store.h: int repo_has_object_file_with_flags(struct repository *r,\n      \n     +int has_loose_object(const struct object_id *);\n     +\n    - void assert_oid_type(const struct object_id *oid, enum object_type expect);\n    - \n    - /*\n    + /**\n    +  * format_object_header() is a thin wrapper around s xsnprintf() that\n    +  * writes the initial \"<type> <obj-len>\" part of the loose object\n     \n      ## t/t5329-pack-objects-cruft.sh (new) ##\n     @@\n 9:  cebb30b667 !  9:  da7273f41f reachable: add options to add_unseen_recent_objects_to_traversal\n    @@ Commit message\n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n     \n      ## builtin/pack-objects.c ##\n    -@@ builtin/pack-objects.c: static void get_object_list(int ac, const char **av)\n    +@@ builtin/pack-objects.c: static void get_object_list(struct rev_info *revs, int ac, const char **av)\n      \tif (unpack_unreachable_expiration) {\n    - \t\trevs.ignore_missing_links = 1;\n    - \t\tif (add_unseen_recent_objects_to_traversal(&revs,\n    + \t\trevs->ignore_missing_links = 1;\n    + \t\tif (add_unseen_recent_objects_to_traversal(revs,\n     -\t\t\t\tunpack_unreachable_expiration))\n     +\t\t\t\tunpack_unreachable_expiration, NULL, 0))\n      \t\t\tdie(_(\"unable to add recent objects\"));\n    - \t\tif (prepare_revision_walk(&revs))\n    + \t\tif (prepare_revision_walk(revs))\n      \t\t\tdie(_(\"revision walk setup failed\"));\n     \n      ## reachable.c ##\n10:  fa4de8859d = 10:  58fecd1747 reachable: report precise timestamps from objects in cruft packs\n11:  92318f8700 = 11:  1740b8ef01 builtin/pack-objects.c: --cruft with expiration\n12:  1e94b33cb4 ! 12:  5992a72cbf builtin/repack.c: support generating a cruft pack\n    @@ builtin/repack.c\n      static int pack_kept_objects = -1;\n      static int write_bitmaps = -1;\n      static int use_delta_islands;\n    + static int run_update_server_info = 1;\n      static char *packdir, *packtmp_name, *packtmp;\n     +static char *cruft_expiration;\n      \n      static const char *const git_repack_usage[] = {\n      \tN_(\"git repack [<options>]\"),\n    -@@ builtin/repack.c: static int repack_config(const char *var, const char *value, void *cb)\n    - \t\tuse_delta_islands = git_config_bool(var, value);\n    - \t\treturn 0;\n    - \t}\n    -+\n    - \treturn git_default_config(var, value, cb);\n    - }\n    - \n     @@ builtin/repack.c: static void repack_promisor_objects(const struct pack_objects_args *args,\n      \t\tdie(_(\"could not finish pack-objects to repack promisor objects\"));\n      }\n13:  9cfcd123bd ! 13:  1b241f8f91 builtin/repack.c: allow configuring cruft pack generation\n    @@ Commit message\n         Signed-off-by: Taylor Blau <me@ttaylorr.com>\n     \n      ## Documentation/config/repack.txt ##\n    -@@ Documentation/config/repack.txt: repack.writeBitmaps::\n    - \tspace and extra time spent on the initial repack.  This has\n    - \tno effect if multiple packfiles are created.\n    - \tDefaults to true on bare repos, false otherwise.\n    +@@ Documentation/config/repack.txt: repack.updateServerInfo::\n    + \tIf set to false, linkgit:git-repack[1] will not run\n    + \tlinkgit:git-update-server-info[1]. Defaults to true. Can be overridden\n    + \twhen true by the `-n` option of linkgit:git-repack[1].\n     +\n     +repack.cruftWindow::\n     +repack.cruftWindowMemory::\n    @@ builtin/repack.c: static const char incremental_bitmap_conflict_error[] = N_(\n      \t\tdelta_base_offset = git_config_bool(var, value);\n      \t\treturn 0;\n     @@ builtin/repack.c: static int repack_config(const char *var, const char *value, void *cb)\n    + \t\trun_update_server_info = git_config_bool(var, value);\n      \t\treturn 0;\n      \t}\n    - \n     +\tif (!strcmp(var, \"repack.cruftwindow\"))\n     +\t\treturn git_config_string(&cruft_po_args->window, var, value);\n     +\tif (!strcmp(var, \"repack.cruftwindowmemory\"))\n    @@ builtin/repack.c: static int repack_config(const char *var, const char *value, v\n     +\t\treturn git_config_string(&cruft_po_args->depth, var, value);\n     +\tif (!strcmp(var, \"repack.cruftthreads\"))\n     +\t\treturn git_config_string(&cruft_po_args->threads, var, value);\n    -+\n      \treturn git_default_config(var, value, cb);\n      }\n      \n    @@ builtin/repack.c: static void remove_redundant_pack(const char *dir_name, const\n      \t\t\t\t const struct pack_objects_args *args)\n      {\n     @@ builtin/repack.c: int cmd_repack(int argc, const char **argv, const char *prefix)\n    + \tint keep_unreachable = 0;\n      \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n    - \tint no_update_server_info = 0;\n      \tstruct pack_objects_args po_args = {NULL};\n     +\tstruct pack_objects_args cruft_po_args = {NULL};\n      \tint geometric_factor = 0;\n14:  1a58807df0 = 14:  ffae78852c builtin/repack.c: use named flags for existing_packs\n15:  ed05cf536b = 15:  0743e373ba builtin/repack.c: add cruft packs to MIDX during geometric repack\n16:  1d5f334138 = 16:  9f7e0acac6 builtin/gc.c: conditionally avoid pruning objects via loose\n17:  f74b425872 = 17:  07fa9d4b47 sha1-file.c: don't freshen cruft packs\n-- \n2.36.1.94.gb0d54bedca\n"},{"id":"455433","messageId":"f494ef7377bf8fb14d96e860106033d1bd1c9ec1.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:10:47Z","receivedAt":"2022-05-18T23:10:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Create a technical document to explain cruft packs. It contains a brief\noverview of the problem, some background, details on the implementation,\nand a couple of alternative approaches not considered here.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/Makefile                  |   1 +\n Documentation/technical/cruft-packs.txt | 123 ++++++++++++++++++++++++\n 2 files changed, 124 insertions(+)\n create mode 100644 Documentation/technical/cruft-packs.txt\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex adb2f1b50a..2faffb52ab 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -94,6 +94,7 @@ TECH_DOCS += MyFirstContribution\n TECH_DOCS += MyFirstObjectWalk\n TECH_DOCS += SubmittingPatches\n TECH_DOCS += technical/bundle-format\n+TECH_DOCS += technical/cruft-packs\n TECH_DOCS += technical/hash-function-transition\n TECH_DOCS += technical/http-protocol\n TECH_DOCS += technical/index-format\ndiff --git a/Documentation/technical/cruft-packs.txt b/Documentation/technical/cruft-packs.txt\nnew file mode 100644\nindex 0000000000..c0f583cd48\n--- /dev/null\n+++ b/Documentation/technical/cruft-packs.txt\n@@ -0,0 +1,123 @@\n+= Cruft packs\n+\n+The cruft packs feature offer an alternative to Git's traditional mechanism of\n+removing unreachable objects. This document provides an overview of Git's\n+pruning mechanism, and how a cruft pack can be used instead to accomplish the\n+same.\n+\n+== Background\n+\n+To remove unreachable objects from your repository, Git offers `git repack -Ad`\n+(see linkgit:git-repack[1]). Quoting from the documentation:\n+\n+[quote]\n+[...] unreachable objects in a previous pack become loose, unpacked objects,\n+instead of being left in the old pack. [...] loose unreachable objects will be\n+pruned according to normal expiry rules with the next 'git gc' invocation.\n+\n+Unreachable objects aren't removed immediately, since doing so could race with\n+an incoming push which may reference an object which is about to be deleted.\n+Instead, those unreachable objects are stored as loose object and stay that way\n+until they are older than the expiration window, at which point they are removed\n+by linkgit:git-prune[1].\n+\n+Git must store these unreachable objects loose in order to keep track of their\n+per-object mtimes. If these unreachable objects were written into one big pack,\n+then either freshening that pack (because an object contained within it was\n+re-written) or creating a new pack of unreachable objects would cause the pack's\n+mtime to get updated, and the objects within it would never leave the expiration\n+window. Instead, objects are stored loose in order to keep track of the\n+individual object mtimes and avoid a situation where all cruft objects are\n+freshened at once.\n+\n+This can lead to undesirable situations when a repository contains many\n+unreachable objects which have not yet left the grace period. Having large\n+directories in the shards of `.git/objects` can lead to decreased performance in\n+the repository. But given enough unreachable objects, this can lead to inode\n+starvation and degrade the performance of the whole system. Since we\n+can never pack those objects, these repositories often take up a large amount of\n+disk space, since we can only zlib compress them, but not store them in delta\n+chains.\n+\n+== Cruft packs\n+\n+A cruft pack eliminates the need for storing unreachable objects in a loose\n+state by including the per-object mtimes in a separate file alongside a single\n+pack containing all loose objects.\n+\n+A cruft pack is written by `git repack --cruft` when generating a new pack.\n+linkgit:git-pack-objects[1]'s `--cruft` option. Note that `git repack --cruft`\n+is a classic all-into-one repack, meaning that everything in the resulting pack is\n+reachable, and everything else is unreachable. Once written, the `--cruft`\n+option instructs `git repack` to generate another pack containing only objects\n+not packed in the previous step (which equates to packing all unreachable\n+objects together). This progresses as follows:\n+\n+  1. Enumerate every object, marking any object which is (a) not contained in a\n+     kept-pack, and (b) whose mtime is within the grace period as a traversal\n+     tip.\n+\n+  2. Perform a reachability traversal based on the tips gathered in the previous\n+     step, adding every object along the way to the pack.\n+\n+  3. Write the pack out, along with a `.mtimes` file that records the per-object\n+     timestamps.\n+\n+This mode is invoked internally by linkgit:git-repack[1] when instructed to\n+write a cruft pack. Crucially, the set of in-core kept packs is exactly the set\n+of packs which will not be deleted by the repack; in other words, they contain\n+all of the repository's reachable objects.\n+\n+When a repository already has a cruft pack, `git repack --cruft` typically only\n+adds objects to it. An exception to this is when `git repack` is given the\n+`--cruft-expiration` option, which allows the generated cruft pack to omit\n+expired objects instead of waiting for linkgit:git-gc[1] to expire those objects\n+later on.\n+\n+It is linkgit:git-gc[1] that is typically responsible for removing expired\n+unreachable objects.\n+\n+== Caution for mixed-version environments\n+\n+Repositories that have cruft packs in them will continue to work with any older\n+version of Git. Note, however, that previous versions of Git which do not\n+understand the `.mtimes` file will use the cruft pack's mtime as the mtime for\n+all of the objects in it. In other words, do not expect older (pre-cruft pack)\n+versions of Git to interpret or even read the contents of the `.mtimes` file.\n+\n+Note that having mixed versions of Git GC-ing the same repository can lead to\n+unreachable objects never being completely pruned. This can happen under the\n+following circumstances:\n+\n+  - An older version of Git running GC explodes the contents of an existing\n+    cruft pack loose, using the cruft pack's mtime.\n+  - A newer version running GC collects those loose objects into a cruft pack,\n+    where the .mtime file reflects the loose object's actual mtimes, but the\n+    cruft pack mtime is \"now\".\n+\n+Repeating this process will lead to unreachable objects not getting pruned as a\n+result of repeatedly resetting the objects' mtimes to the present time.\n+\n+If you are GC-ing repositories in a mixed version environment, consider omitting\n+the `--cruft` option when using linkgit:git-repack[1] and linkgit:git-gc[1], and\n+leaving the `gc.cruftPacks` configuration unset until all writers understand\n+cruft packs.\n+\n+== Alternatives\n+\n+Notable alternatives to this design include:\n+\n+  - The location of the per-object mtime data, and\n+  - Storing unreachable objects in multiple cruft packs.\n+\n+On the location of mtime data, a new auxiliary file tied to the pack was chosen\n+to avoid complicating the `.idx` format. If the `.idx` format were ever to gain\n+support for optional chunks of data, it may make sense to consolidate the\n+`.mtimes` format into the `.idx` itself.\n+\n+Storing unreachable objects among multiple cruft packs (e.g., creating a new\n+cruft pack during each repacking operation including only unreachable objects\n+which aren't already stored in an earlier cruft pack) is significantly more\n+complicated to construct, and so aren't pursued here. The obvious drawback to\n+the current implementation is that the entire cruft pack must be re-written from\n+scratch.\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455434","messageId":"8f9fd21be9fcdda5c73d800fc66d1087d61a6888.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:10:50Z","receivedAt":"2022-05-18T23:10:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"To store the individual mtimes of objects in a cruft pack, introduce a\nnew `.mtimes` format that can optionally accompany a single pack in the\nrepository.\n\nThe format is defined in Documentation/technical/pack-format.txt, and\nstores a 4-byte network order timestamp for each object in name (index)\norder.\n\nThis patch prepares for cruft packs by defining the `.mtimes` format,\nand introducing a basic API that callers can use to read out individual\nmtimes.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/technical/pack-format.txt |  19 ++++\n Makefile                                |   1 +\n builtin/repack.c                        |   1 +\n object-store.h                          |   5 +-\n pack-mtimes.c                           | 126 ++++++++++++++++++++++++\n pack-mtimes.h                           |  15 +++\n packfile.c                              |  19 +++-\n 7 files changed, 183 insertions(+), 3 deletions(-)\n create mode 100644 pack-mtimes.c\n create mode 100644 pack-mtimes.h\n\ndiff --git a/Documentation/technical/pack-format.txt b/Documentation/technical/pack-format.txt\nindex 6d3efb7d16..c443dbb526 100644\n--- a/Documentation/technical/pack-format.txt\n+++ b/Documentation/technical/pack-format.txt\n@@ -294,6 +294,25 @@ Pack file entry: <+\n \n All 4-byte numbers are in network order.\n \n+== pack-*.mtimes files have the format:\n+\n+  - A 4-byte magic number '0x4d544d45' ('MTME').\n+\n+  - A 4-byte version identifier (= 1).\n+\n+  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n+\n+  - A table of 4-byte unsigned integers in network order. The ith\n+    value is the modification time (mtime) of the ith object in the\n+    corresponding pack by lexicographic (index) order. The mtimes\n+    count standard epoch seconds.\n+\n+  - A trailer, containing a checksum of the corresponding packfile,\n+    and a checksum of all of the above (each having length according\n+    to the specified hash function).\n+\n+All 4-byte numbers are in network order.\n+\n == multi-pack-index (MIDX) files have the following format:\n \n The multi-pack-index files refer to multiple pack-files and loose objects.\ndiff --git a/Makefile b/Makefile\nindex 61aadf3ce8..a299580b7c 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -993,6 +993,7 @@ LIB_OBJS += oidtree.o\n LIB_OBJS += pack-bitmap-write.o\n LIB_OBJS += pack-bitmap.o\n LIB_OBJS += pack-check.o\n+LIB_OBJS += pack-mtimes.o\n LIB_OBJS += pack-objects.o\n LIB_OBJS += pack-revindex.o\n LIB_OBJS += pack-write.o\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex d1a563d5b6..e7a3920c6d 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -217,6 +217,7 @@ static struct {\n } exts[] = {\n \t{\".pack\"},\n \t{\".rev\", 1},\n+\t{\".mtimes\", 1},\n \t{\".bitmap\", 1},\n \t{\".promisor\", 1},\n \t{\".idx\"},\ndiff --git a/object-store.h b/object-store.h\nindex 53996018c1..2c4671ed7a 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -115,12 +115,15 @@ struct packed_git {\n \t\t freshened:1,\n \t\t do_not_close:1,\n \t\t pack_promisor:1,\n-\t\t multi_pack_index:1;\n+\t\t multi_pack_index:1,\n+\t\t is_cruft:1;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n \tstruct revindex_entry *revindex;\n \tconst uint32_t *revindex_data;\n \tconst uint32_t *revindex_map;\n \tsize_t revindex_size;\n+\tconst uint32_t *mtimes_map;\n+\tsize_t mtimes_size;\n \t/* something like \".git/objects/pack/xxxxx.pack\" */\n \tchar pack_name[FLEX_ARRAY]; /* more */\n };\ndiff --git a/pack-mtimes.c b/pack-mtimes.c\nnew file mode 100644\nindex 0000000000..46ad584af1\n--- /dev/null\n+++ b/pack-mtimes.c\n@@ -0,0 +1,126 @@\n+#include \"pack-mtimes.h\"\n+#include \"object-store.h\"\n+#include \"packfile.h\"\n+\n+static char *pack_mtimes_filename(struct packed_git *p)\n+{\n+\tsize_t len;\n+\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n+\t\tBUG(\"pack_name does not end in .pack\");\n+\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n+\treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n+}\n+\n+#define MTIMES_HEADER_SIZE (12)\n+#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n+\n+struct mtimes_header {\n+\tuint32_t signature;\n+\tuint32_t version;\n+\tuint32_t hash_id;\n+};\n+\n+static int load_pack_mtimes_file(char *mtimes_file,\n+\t\t\t\t uint32_t num_objects,\n+\t\t\t\t const uint32_t **data_p, size_t *len_p)\n+{\n+\tint fd, ret = 0;\n+\tstruct stat st;\n+\tvoid *data = NULL;\n+\tsize_t mtimes_size;\n+\tstruct mtimes_header header;\n+\tuint32_t *hdr;\n+\n+\tfd = git_open(mtimes_file);\n+\n+\tif (fd < 0) {\n+\t\tret = -1;\n+\t\tgoto cleanup;\n+\t}\n+\tif (fstat(fd, &st)) {\n+\t\tret = error_errno(_(\"failed to read %s\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tmtimes_size = xsize_t(st.st_size);\n+\n+\tif (mtimes_size < MTIMES_MIN_SIZE) {\n+\t\tret = error(_(\"mtimes file %s is too small\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t), num_objects)) {\n+\t\tret = error(_(\"mtimes file %s is corrupt\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tdata = hdr = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\n+\theader.signature = ntohl(hdr[0]);\n+\theader.version = ntohl(hdr[1]);\n+\theader.hash_id = ntohl(hdr[2]);\n+\n+\tif (header.signature != MTIMES_SIGNATURE) {\n+\t\tret = error(_(\"mtimes file %s has unknown signature\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (header.version != 1) {\n+\t\tret = error(_(\"mtimes file %s has unsupported version %\"PRIu32),\n+\t\t\t    mtimes_file, header.version);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (!(header.hash_id == 1 || header.hash_id == 2)) {\n+\t\tret = error(_(\"mtimes file %s has unsupported hash id %\"PRIu32),\n+\t\t\t    mtimes_file, header.hash_id);\n+\t\tgoto cleanup;\n+\t}\n+\n+cleanup:\n+\tif (ret) {\n+\t\tif (data)\n+\t\t\tmunmap(data, mtimes_size);\n+\t} else {\n+\t\t*len_p = mtimes_size;\n+\t\t*data_p = (const uint32_t *)data;\n+\t}\n+\n+\tclose(fd);\n+\treturn ret;\n+}\n+\n+int load_pack_mtimes(struct packed_git *p)\n+{\n+\tchar *mtimes_name = NULL;\n+\tint ret = 0;\n+\n+\tif (!p->is_cruft)\n+\t\treturn ret; /* not a cruft pack */\n+\tif (p->mtimes_map)\n+\t\treturn ret; /* already loaded */\n+\n+\tret = open_pack_index(p);\n+\tif (ret < 0)\n+\t\tgoto cleanup;\n+\n+\tmtimes_name = pack_mtimes_filename(p);\n+\tret = load_pack_mtimes_file(mtimes_name,\n+\t\t\t\t    p->num_objects,\n+\t\t\t\t    &p->mtimes_map,\n+\t\t\t\t    &p->mtimes_size);\n+cleanup:\n+\tfree(mtimes_name);\n+\treturn ret;\n+}\n+\n+uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos)\n+{\n+\tif (!p->mtimes_map)\n+\t\tBUG(\"pack .mtimes file not loaded for %s\", p->pack_name);\n+\tif (p->num_objects <= pos)\n+\t\tBUG(\"pack .mtimes out-of-bounds (%\"PRIu32\" vs %\"PRIu32\")\",\n+\t\t    pos, p->num_objects);\n+\n+\treturn get_be32(p->mtimes_map + pos + 3);\n+}\ndiff --git a/pack-mtimes.h b/pack-mtimes.h\nnew file mode 100644\nindex 0000000000..38ddb9f893\n--- /dev/null\n+++ b/pack-mtimes.h\n@@ -0,0 +1,15 @@\n+#ifndef PACK_MTIMES_H\n+#define PACK_MTIMES_H\n+\n+#include \"git-compat-util.h\"\n+\n+#define MTIMES_SIGNATURE 0x4d544d45 /* \"MTME\" */\n+#define MTIMES_VERSION 1\n+\n+struct packed_git;\n+\n+int load_pack_mtimes(struct packed_git *p);\n+\n+uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos);\n+\n+#endif\ndiff --git a/packfile.c b/packfile.c\nindex 835b2d2716..fc0245fbab 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -334,12 +334,22 @@ static void close_pack_revindex(struct packed_git *p)\n \tp->revindex_data = NULL;\n }\n \n+static void close_pack_mtimes(struct packed_git *p)\n+{\n+\tif (!p->mtimes_map)\n+\t\treturn;\n+\n+\tmunmap((void *)p->mtimes_map, p->mtimes_size);\n+\tp->mtimes_map = NULL;\n+}\n+\n void close_pack(struct packed_git *p)\n {\n \tclose_pack_windows(p);\n \tclose_pack_fd(p);\n \tclose_pack_index(p);\n \tclose_pack_revindex(p);\n+\tclose_pack_mtimes(p);\n \toidset_clear(&p->bad_objects);\n }\n \n@@ -363,7 +373,7 @@ void close_object_store(struct raw_object_store *o)\n \n void unlink_pack_path(const char *pack_name, int force_delete)\n {\n-\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\"};\n+\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\", \".mtimes\"};\n \tint i;\n \tstruct strbuf buf = STRBUF_INIT;\n \tsize_t plen;\n@@ -718,6 +728,10 @@ struct packed_git *add_packed_git(const char *path, size_t path_len, int local)\n \tif (!access(p->pack_name, F_OK))\n \t\tp->pack_promisor = 1;\n \n+\txsnprintf(p->pack_name + path_len, alloc - path_len, \".mtimes\");\n+\tif (!access(p->pack_name, F_OK))\n+\t\tp->is_cruft = 1;\n+\n \txsnprintf(p->pack_name + path_len, alloc - path_len, \".pack\");\n \tif (stat(p->pack_name, &st) || !S_ISREG(st.st_mode)) {\n \t\tfree(p);\n@@ -869,7 +883,8 @@ static void prepare_pack(const char *full_name, size_t full_name_len,\n \t    ends_with(file_name, \".pack\") ||\n \t    ends_with(file_name, \".bitmap\") ||\n \t    ends_with(file_name, \".keep\") ||\n-\t    ends_with(file_name, \".promisor\"))\n+\t    ends_with(file_name, \".promisor\") ||\n+\t    ends_with(file_name, \".mtimes\"))\n \t\tstring_list_append(data->garbage, full_name);\n \telse\n \t\treport_garbage(PACKDIR_FILE_GARBAGE, full_name);\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455435","messageId":"cdb21236e16ae72af5f234c0138cbd9ea725ab47.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 03/17] pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:10:53Z","receivedAt":"2022-05-18T23:11:20Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This structure will be used to communicate the per-object mtimes when\nwriting a cruft pack. Here, we need the full packing_data structure\nbecause the mtime information is stored in an array there, not on the\nindividual object_entry's themselves (to avoid paying the overhead in\nstructure width for operations which do not generate a cruft pack).\n\nWe haven't passed this information down before because one of the two\ncallers (in bulk-checkin.c) does not have a packing_data structure at\nall. In that case (where no cruft pack will be generated), NULL is\npassed instead.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 3 ++-\n bulk-checkin.c         | 2 +-\n pack-write.c           | 1 +\n pack.h                 | 3 +++\n 4 files changed, 7 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 014dcd4bc9..6ac927047c 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1262,7 +1262,8 @@ static void write_pack_file(void)\n \n \t\t\tstage_tmp_packfiles(&tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n-\t\t\t\t\t    &pack_idx_opts, hash, &idx_tmp_name);\n+\t\t\t\t\t    &to_pack, &pack_idx_opts, hash,\n+\t\t\t\t\t    &idx_tmp_name);\n \n \t\t\tif (write_bitmap_index) {\n \t\t\t\tsize_t tmpname_len = tmpname.len;\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6d6c37171c..e988a388b6 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -33,7 +33,7 @@ static void finish_tmp_packfile(struct strbuf *basename,\n \tchar *idx_tmp_name = NULL;\n \n \tstage_tmp_packfiles(basename, pack_tmp_name, written_list, nr_written,\n-\t\t\t    pack_idx_opts, hash, &idx_tmp_name);\n+\t\t\t    NULL, pack_idx_opts, hash, &idx_tmp_name);\n \trename_tmp_packfile_idx(basename, &idx_tmp_name);\n \n \tfree(idx_tmp_name);\ndiff --git a/pack-write.c b/pack-write.c\nindex 51812cb129..a2adc565f4 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -484,6 +484,7 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t const char *pack_tmp_name,\n \t\t\t struct pack_idx_entry **written_list,\n \t\t\t uint32_t nr_written,\n+\t\t\t struct packing_data *to_pack,\n \t\t\t struct pack_idx_option *pack_idx_opts,\n \t\t\t unsigned char hash[],\n \t\t\t char **idx_tmp_name)\ndiff --git a/pack.h b/pack.h\nindex b22bfc4a18..fd27cfdfd7 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -109,11 +109,14 @@ int encode_in_pack_object_header(unsigned char *hdr, int hdr_len,\n #define PH_ERROR_PROTOCOL\t(-3)\n int read_pack_header(int fd, struct pack_header *);\n \n+struct packing_data;\n+\n struct hashfile *create_tmp_packfile(char **pack_tmp_name);\n void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t const char *pack_tmp_name,\n \t\t\t struct pack_idx_entry **written_list,\n \t\t\t uint32_t nr_written,\n+\t\t\t struct packing_data *to_pack,\n \t\t\t struct pack_idx_option *pack_idx_opts,\n \t\t\t unsigned char hash[],\n \t\t\t char **idx_tmp_name);\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455436","messageId":"5f9a9a5b7b49a3f6e246ea8cb51aa60d204e0f49.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 06/17] t/helper: add 'pack-mtimes' test-tool","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:06Z","receivedAt":"2022-05-18T23:11:22Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In the next patch, we will implement and test support for writing a\ncruft pack via a special mode of `git pack-objects`. To make sure that\nobjects are written with the correct timestamps, and a new test-tool\nthat can dump the object names and corresponding timestamps from a given\n`.mtimes` file.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Makefile                    |  1 +\n t/helper/test-pack-mtimes.c | 56 +++++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c        |  1 +\n t/helper/test-tool.h        |  1 +\n 4 files changed, 59 insertions(+)\n create mode 100644 t/helper/test-pack-mtimes.c\n\ndiff --git a/Makefile b/Makefile\nindex a299580b7c..0b6eab0453 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -738,6 +738,7 @@ TEST_BUILTINS_OBJS += test-oid-array.o\n TEST_BUILTINS_OBJS += test-oidmap.o\n TEST_BUILTINS_OBJS += test-oidtree.o\n TEST_BUILTINS_OBJS += test-online-cpus.o\n+TEST_BUILTINS_OBJS += test-pack-mtimes.o\n TEST_BUILTINS_OBJS += test-parse-options.o\n TEST_BUILTINS_OBJS += test-parse-pathspec-file.o\n TEST_BUILTINS_OBJS += test-partial-clone.o\ndiff --git a/t/helper/test-pack-mtimes.c b/t/helper/test-pack-mtimes.c\nnew file mode 100644\nindex 0000000000..f7b79daf4c\n--- /dev/null\n+++ b/t/helper/test-pack-mtimes.c\n@@ -0,0 +1,56 @@\n+#include \"git-compat-util.h\"\n+#include \"test-tool.h\"\n+#include \"strbuf.h\"\n+#include \"object-store.h\"\n+#include \"packfile.h\"\n+#include \"pack-mtimes.h\"\n+\n+static void dump_mtimes(struct packed_git *p)\n+{\n+\tuint32_t i;\n+\tif (load_pack_mtimes(p) < 0)\n+\t\tdie(\"could not load pack .mtimes\");\n+\n+\tfor (i = 0; i < p->num_objects; i++) {\n+\t\tstruct object_id oid;\n+\t\tif (nth_packed_object_id(&oid, p, i) < 0)\n+\t\t\tdie(\"could not load object id at position %\"PRIu32, i);\n+\n+\t\tprintf(\"%s %\"PRIu32\"\\n\",\n+\t\t       oid_to_hex(&oid), nth_packed_mtime(p, i));\n+\t}\n+}\n+\n+static const char *pack_mtimes_usage = \"\\n\"\n+\"  test-tool pack-mtimes <pack-name.mtimes>\";\n+\n+int cmd__pack_mtimes(int argc, const char **argv)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct packed_git *p;\n+\n+\tsetup_git_directory();\n+\n+\tif (argc != 2)\n+\t\tusage(pack_mtimes_usage);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tstrbuf_addstr(&buf, basename(p->pack_name));\n+\t\tstrbuf_strip_suffix(&buf, \".pack\");\n+\t\tstrbuf_addstr(&buf, \".mtimes\");\n+\n+\t\tif (!strcmp(buf.buf, argv[1]))\n+\t\t\tbreak;\n+\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstrbuf_release(&buf);\n+\n+\tif (!p)\n+\t\tdie(\"could not find pack '%s'\", argv[1]);\n+\n+\tdump_mtimes(p);\n+\n+\treturn 0;\n+}\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex 0424f7adf5..d2eacd302d 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -48,6 +48,7 @@ static struct test_cmd cmds[] = {\n \t{ \"oidmap\", cmd__oidmap },\n \t{ \"oidtree\", cmd__oidtree },\n \t{ \"online-cpus\", cmd__online_cpus },\n+\t{ \"pack-mtimes\", cmd__pack_mtimes },\n \t{ \"parse-options\", cmd__parse_options },\n \t{ \"parse-pathspec-file\", cmd__parse_pathspec_file },\n \t{ \"partial-clone\", cmd__partial_clone },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex c876e8246f..960cc27ef7 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -38,6 +38,7 @@ int cmd__mktemp(int argc, const char **argv);\n int cmd__oidmap(int argc, const char **argv);\n int cmd__oidtree(int argc, const char **argv);\n int cmd__online_cpus(int argc, const char **argv);\n+int cmd__pack_mtimes(int argc, const char **argv);\n int cmd__parse_options(int argc, const char **argv);\n int cmd__parse_pathspec_file(int argc, const char** argv);\n int cmd__partial_clone(int argc, const char **argv);\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455437","messageId":"b8a38fe2e48cf0ccc3b40b97f396b26245d183cb.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 07/17] builtin/pack-objects.c: return from create_object_entry()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:09Z","receivedAt":"2022-05-18T23:11:24Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"A new caller in the next commit will want to immediately modify the\nobject_entry structure created by create_object_entry(). Instead of\nforcing that caller to wastefully look-up the entry we just created,\nreturn it from create_object_entry() instead.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 16 +++++++++-------\n 1 file changed, 9 insertions(+), 7 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 6ac927047c..c6d16872ee 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1516,13 +1516,13 @@ static int want_object_in_pack(const struct object_id *oid,\n \treturn 1;\n }\n \n-static void create_object_entry(const struct object_id *oid,\n-\t\t\t\tenum object_type type,\n-\t\t\t\tuint32_t hash,\n-\t\t\t\tint exclude,\n-\t\t\t\tint no_try_delta,\n-\t\t\t\tstruct packed_git *found_pack,\n-\t\t\t\toff_t found_offset)\n+static struct object_entry *create_object_entry(const struct object_id *oid,\n+\t\t\t\t\t\tenum object_type type,\n+\t\t\t\t\t\tuint32_t hash,\n+\t\t\t\t\t\tint exclude,\n+\t\t\t\t\t\tint no_try_delta,\n+\t\t\t\t\t\tstruct packed_git *found_pack,\n+\t\t\t\t\t\toff_t found_offset)\n {\n \tstruct object_entry *entry;\n \n@@ -1539,6 +1539,8 @@ static void create_object_entry(const struct object_id *oid,\n \t}\n \n \tentry->no_try_delta = no_try_delta;\n+\n+\treturn entry;\n }\n \n static const char no_closure_warning[] = N_(\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455438","messageId":"58fecd1747462faf323d861662f588590ece3376.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 10/17] reachable: report precise timestamps from objects in cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:17Z","receivedAt":"2022-05-18T23:11:26Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When generating a cruft pack, the caller within pack-objects will want\nto know the precise timestamps of cruft objects (i.e., their\ncorresponding values in the .mtimes table) rather than the mtime of the\ncruft pack itself.\n\nTeach add_recent_packed() to lookup each object's precise mtime from the\n.mtimes file if one exists (indicated by the is_cruft bit on the\npacked_git structure).\n\nA couple of small things worth noting here:\n\n  - load_pack_mtimes() needs to be called before asking for\n    nth_packed_mtime(), and that call is done lazily here. That function\n    exits early if the .mtimes file has already been opened and parsed,\n    so only the first call is slow.\n\n  - Checking the is_cruft bit can be done without any extra work on the\n    caller's behalf, since it is set up for us automatically as a\n    side-effect of calling add_packed_git() (just like the 'pack_keep'\n    and 'pack_promisor' bits).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n reachable.c | 9 ++++++++-\n 1 file changed, 8 insertions(+), 1 deletion(-)\n\ndiff --git a/reachable.c b/reachable.c\nindex d4507c4270..aba63ebeb3 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -13,6 +13,7 @@\n #include \"worktree.h\"\n #include \"object-store.h\"\n #include \"pack-bitmap.h\"\n+#include \"pack-mtimes.h\"\n \n struct connectivity_progress {\n \tstruct progress *progress;\n@@ -155,6 +156,7 @@ static int add_recent_packed(const struct object_id *oid,\n \t\t\t     void *data)\n {\n \tstruct object *obj;\n+\ttimestamp_t mtime = p->mtime;\n \n \tif (!want_recent_object(data, oid))\n \t\treturn 0;\n@@ -163,7 +165,12 @@ static int add_recent_packed(const struct object_id *oid,\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n-\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), p->mtime, data);\n+\tif (p->is_cruft) {\n+\t\tif (load_pack_mtimes(p) < 0)\n+\t\t\tdie(_(\"could not load cruft pack .mtimes\"));\n+\t\tmtime = nth_packed_mtime(p, pos);\n+\t}\n+\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), mtime, data);\n \treturn 0;\n }\n \n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455439","messageId":"da7273f41f61e7b87cecf5a9eb0ab6a4eef4353b.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 09/17] reachable: add options to add_unseen_recent_objects_to_traversal","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:14Z","receivedAt":"2022-05-18T23:11:29Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This function behaves very similarly to what we will need in\npack-objects in order to implement cruft packs with expiration. But it\nis lacking a couple of things. Namely, it needs:\n\n  - a mechanism to communicate the timestamps of individual recent\n    objects to some external caller\n\n  - and, in the case of packed objects, our future caller will also want\n    to know the originating pack, as well as the offset within that pack\n    at which the object can be found\n\n  - finally, it needs a way to skip over packs which are marked as kept\n    in-core.\n\nTo address the first two, add a callback interface in this patch which\nreports the time of each recent object, as well as a (packed_git,\noff_t) pair for packed objects.\n\nLikewise, add a new option to the packed object iterators to skip over\npacks which are marked as kept in core. This option will become\nimplicitly tested in a future patch.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c |  2 +-\n reachable.c            | 51 +++++++++++++++++++++++++++++++++++-------\n reachable.h            |  9 +++++++-\n 3 files changed, 52 insertions(+), 10 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 9cf89be673..3b8bf6a3dd 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -3957,7 +3957,7 @@ static void get_object_list(struct rev_info *revs, int ac, const char **av)\n \tif (unpack_unreachable_expiration) {\n \t\trevs->ignore_missing_links = 1;\n \t\tif (add_unseen_recent_objects_to_traversal(revs,\n-\t\t\t\tunpack_unreachable_expiration))\n+\t\t\t\tunpack_unreachable_expiration, NULL, 0))\n \t\t\tdie(_(\"unable to add recent objects\"));\n \t\tif (prepare_revision_walk(revs))\n \t\t\tdie(_(\"revision walk setup failed\"));\ndiff --git a/reachable.c b/reachable.c\nindex b9f4ad886e..d4507c4270 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -60,9 +60,13 @@ static void mark_commit(struct commit *c, void *data)\n struct recent_data {\n \tstruct rev_info *revs;\n \ttimestamp_t timestamp;\n+\treport_recent_object_fn *cb;\n+\tint ignore_in_core_kept_packs;\n };\n \n static void add_recent_object(const struct object_id *oid,\n+\t\t\t      struct packed_git *pack,\n+\t\t\t      off_t offset,\n \t\t\t      timestamp_t mtime,\n \t\t\t      struct recent_data *data)\n {\n@@ -103,13 +107,29 @@ static void add_recent_object(const struct object_id *oid,\n \t\tdie(\"unable to lookup %s\", oid_to_hex(oid));\n \n \tadd_pending_object(data->revs, obj, \"\");\n+\tif (data->cb)\n+\t\tdata->cb(obj, pack, offset, mtime);\n+}\n+\n+static int want_recent_object(struct recent_data *data,\n+\t\t\t      const struct object_id *oid)\n+{\n+\tif (data->ignore_in_core_kept_packs &&\n+\t    has_object_kept_pack(oid, IN_CORE_KEEP_PACKS))\n+\t\treturn 0;\n+\treturn 1;\n }\n \n static int add_recent_loose(const struct object_id *oid,\n \t\t\t    const char *path, void *data)\n {\n \tstruct stat st;\n-\tstruct object *obj = lookup_object(the_repository, oid);\n+\tstruct object *obj;\n+\n+\tif (!want_recent_object(data, oid))\n+\t\treturn 0;\n+\n+\tobj = lookup_object(the_repository, oid);\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n@@ -126,7 +146,7 @@ static int add_recent_loose(const struct object_id *oid,\n \t\treturn error_errno(\"unable to stat %s\", oid_to_hex(oid));\n \t}\n \n-\tadd_recent_object(oid, st.st_mtime, data);\n+\tadd_recent_object(oid, NULL, 0, st.st_mtime, data);\n \treturn 0;\n }\n \n@@ -134,29 +154,43 @@ static int add_recent_packed(const struct object_id *oid,\n \t\t\t     struct packed_git *p, uint32_t pos,\n \t\t\t     void *data)\n {\n-\tstruct object *obj = lookup_object(the_repository, oid);\n+\tstruct object *obj;\n+\n+\tif (!want_recent_object(data, oid))\n+\t\treturn 0;\n+\n+\tobj = lookup_object(the_repository, oid);\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n-\tadd_recent_object(oid, p->mtime, data);\n+\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), p->mtime, data);\n \treturn 0;\n }\n \n int add_unseen_recent_objects_to_traversal(struct rev_info *revs,\n-\t\t\t\t\t   timestamp_t timestamp)\n+\t\t\t\t\t   timestamp_t timestamp,\n+\t\t\t\t\t   report_recent_object_fn *cb,\n+\t\t\t\t\t   int ignore_in_core_kept_packs)\n {\n \tstruct recent_data data;\n+\tenum for_each_object_flags flags;\n \tint r;\n \n \tdata.revs = revs;\n \tdata.timestamp = timestamp;\n+\tdata.cb = cb;\n+\tdata.ignore_in_core_kept_packs = ignore_in_core_kept_packs;\n \n \tr = for_each_loose_object(add_recent_loose, &data,\n \t\t\t\t  FOR_EACH_OBJECT_LOCAL_ONLY);\n \tif (r)\n \t\treturn r;\n-\treturn for_each_packed_object(add_recent_packed, &data,\n-\t\t\t\t      FOR_EACH_OBJECT_LOCAL_ONLY);\n+\n+\tflags = FOR_EACH_OBJECT_LOCAL_ONLY | FOR_EACH_OBJECT_PACK_ORDER;\n+\tif (ignore_in_core_kept_packs)\n+\t\tflags |= FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS;\n+\n+\treturn for_each_packed_object(add_recent_packed, &data, flags);\n }\n \n static int mark_object_seen(const struct object_id *oid,\n@@ -217,7 +251,8 @@ void mark_reachable_objects(struct rev_info *revs, int mark_reflog,\n \n \tif (mark_recent) {\n \t\trevs->ignore_missing_links = 1;\n-\t\tif (add_unseen_recent_objects_to_traversal(revs, mark_recent))\n+\t\tif (add_unseen_recent_objects_to_traversal(revs, mark_recent,\n+\t\t\t\t\t\t\t   NULL, 0))\n \t\t\tdie(\"unable to mark recent objects\");\n \t\tif (prepare_revision_walk(revs))\n \t\t\tdie(\"revision walk setup failed\");\ndiff --git a/reachable.h b/reachable.h\nindex 5df932ad8f..b776761baa 100644\n--- a/reachable.h\n+++ b/reachable.h\n@@ -1,11 +1,18 @@\n #ifndef REACHEABLE_H\n #define REACHEABLE_H\n \n+#include \"object.h\"\n+\n struct progress;\n struct rev_info;\n \n+typedef void report_recent_object_fn(const struct object *, struct packed_git *,\n+\t\t\t\t     off_t, time_t);\n+\n int add_unseen_recent_objects_to_traversal(struct rev_info *revs,\n-\t\t\t\t\t   timestamp_t timestamp);\n+\t\t\t\t\t   timestamp_t timestamp,\n+\t\t\t\t\t   report_recent_object_fn cb,\n+\t\t\t\t\t   int ignore_in_core_kept_packs);\n void mark_reachable_objects(struct rev_info *revs, int mark_reflog,\n \t\t\t    timestamp_t mark_recent, struct progress *);\n \n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455440","messageId":"94fe03cc65716b6102e2d71df49d4ae5a1a60dc7.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:12Z","receivedAt":"2022-05-18T23:11:30Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Teach `pack-objects` how to generate a cruft pack when no objects are\ndropped (i.e., `--cruft-expiration=never`). Later patches will teach\n`pack-objects` how to generate a cruft pack that prunes objects.\n\nWhen generating a cruft pack which does not prune objects, we want to\ncollect all unreachable objects into a single pack (noting and updating\ntheir mtimes as we accumulate them). Ordinary use will pass the result\nof a `git repack -A` as a kept pack, so when this patch says \"kept\npack\", readers should think \"reachable objects\".\n\nGenerating a non-expiring cruft packs works as follows:\n\n  - Callers provide a list of every pack they know about, and indicate\n    which packs are about to be removed.\n\n  - All packs which are going to be removed (we'll call these the\n    redundant ones) are marked as kept in-core.\n\n    Any packs the caller did not mention (but are known to the\n    `pack-objects` process) are also marked as kept in-core. Packs not\n    mentioned by the caller are assumed to be unknown to them, i.e.,\n    they entered the repository after the caller decided which packs\n    should be kept and which should be discarded.\n\n    Since we do not want to include objects in these \"unknown\" packs\n    (because we don't know which of their objects are or aren't\n    reachable), these are also marked as kept in-core.\n\n  - Then, we enumerate all objects in the repository, and add them to\n    our packing list if they do not appear in an in-core kept pack.\n\nThis results in a new cruft pack which contains all known objects that\naren't included in the kept packs. When the kept pack is the result of\n`git repack -A`, the resulting pack contains all unreachable objects.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-pack-objects.txt |  30 ++++\n builtin/pack-objects.c             | 201 +++++++++++++++++++++++++-\n object-file.c                      |   2 +-\n object-store.h                     |   2 +\n t/t5329-pack-objects-cruft.sh      | 218 +++++++++++++++++++++++++++++\n 5 files changed, 448 insertions(+), 5 deletions(-)\n create mode 100755 t/t5329-pack-objects-cruft.sh\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex f8344e1e5b..a9995a932c 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -13,6 +13,7 @@ SYNOPSIS\n \t[--no-reuse-delta] [--delta-base-offset] [--non-empty]\n \t[--local] [--incremental] [--window=<n>] [--depth=<n>]\n \t[--revs [--unpacked | --all]] [--keep-pack=<pack-name>]\n+\t[--cruft] [--cruft-expiration=<time>]\n \t[--stdout [--filter=<filter-spec>] | <base-name>]\n \t[--shallow] [--keep-true-parents] [--[no-]sparse] < <object-list>\n \n@@ -95,6 +96,35 @@ base-name::\n Incompatible with `--revs`, or options that imply `--revs` (such as\n `--all`), with the exception of `--unpacked`, which is compatible.\n \n+--cruft::\n+\tPacks unreachable objects into a separate \"cruft\" pack, denoted\n+\tby the existence of a `.mtimes` file. Typically used by `git\n+\trepack --cruft`. Callers provide a list of pack names and\n+\tindicate which packs will remain in the repository, along with\n+\twhich packs will be deleted (indicated by the `-` prefix). The\n+\tcontents of the cruft pack are all objects not contained in the\n+\tsurviving packs which have not exceeded the grace period (see\n+\t`--cruft-expiration` below), or which have exceeded the grace\n+\tperiod, but are reachable from an other object which hasn't.\n++\n+When the input lists a pack containing all reachable objects (and lists\n+all other packs as pending deletion), the corresponding cruft pack will\n+contain all unreachable objects (with mtime newer than the\n+`--cruft-expiration`) along with any unreachable objects whose mtime is\n+older than the `--cruft-expiration`, but are reachable from an\n+unreachable object whose mtime is newer than the `--cruft-expiration`).\n++\n+Incompatible with `--unpack-unreachable`, `--keep-unreachable`,\n+`--pack-loose-unreachable`, `--stdin-packs`, as well as any other\n+options which imply `--revs`. Also incompatible with `--max-pack-size`;\n+when this option is set, the maximum pack size is not inferred from\n+`pack.packSizeLimit`.\n+\n+--cruft-expiration=<approxidate>::\n+\tIf specified, objects are eliminated from the cruft pack if they\n+\thave an mtime older than `<approxidate>`. If unspecified (and\n+\tgiven `--cruft`), then no objects are eliminated.\n+\n --window=<n>::\n --depth=<n>::\n \tThese two options affect how the objects contained in\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex c6d16872ee..9cf89be673 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -36,6 +36,7 @@\n #include \"trace2.h\"\n #include \"shallow.h\"\n #include \"promisor-remote.h\"\n+#include \"pack-mtimes.h\"\n \n /*\n  * Objects we are going to pack are collected in the `to_pack` structure.\n@@ -194,6 +195,8 @@ static int reuse_delta = 1, reuse_object = 1;\n static int keep_unreachable, unpack_unreachable, include_tag;\n static timestamp_t unpack_unreachable_expiration;\n static int pack_loose_unreachable;\n+static int cruft;\n+static timestamp_t cruft_expiration;\n static int local;\n static int have_non_local_packs;\n static int incremental;\n@@ -1260,6 +1263,9 @@ static void write_pack_file(void)\n \t\t\t\t\t&to_pack, written_list, nr_written);\n \t\t\t}\n \n+\t\t\tif (cruft)\n+\t\t\t\tpack_idx_opts.flags |= WRITE_MTIMES;\n+\n \t\t\tstage_tmp_packfiles(&tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n \t\t\t\t\t    &to_pack, &pack_idx_opts, hash,\n@@ -3397,6 +3403,135 @@ static void read_packs_list_from_stdin(void)\n \tstring_list_clear(&exclude_packs, 0);\n }\n \n+static void add_cruft_object_entry(const struct object_id *oid, enum object_type type,\n+\t\t\t\t   struct packed_git *pack, off_t offset,\n+\t\t\t\t   const char *name, uint32_t mtime)\n+{\n+\tstruct object_entry *entry;\n+\n+\tdisplay_progress(progress_state, ++nr_seen);\n+\n+\tentry = packlist_find(&to_pack, oid);\n+\tif (entry) {\n+\t\tif (name) {\n+\t\t\tentry->hash = pack_name_hash(name);\n+\t\t\tentry->no_try_delta = no_try_delta(name);\n+\t\t}\n+\t} else {\n+\t\tif (!want_object_in_pack(oid, 0, &pack, &offset))\n+\t\t\treturn;\n+\t\tif (!pack && type == OBJ_BLOB && !has_loose_object(oid)) {\n+\t\t\t/*\n+\t\t\t * If a traversed tree has a missing blob then we want\n+\t\t\t * to avoid adding that missing object to our pack.\n+\t\t\t *\n+\t\t\t * This only applies to missing blobs, not trees,\n+\t\t\t * because the traversal needs to parse sub-trees but\n+\t\t\t * not blobs.\n+\t\t\t *\n+\t\t\t * Note we only perform this check when we couldn't\n+\t\t\t * already find the object in a pack, so we're really\n+\t\t\t * limited to \"ensure non-tip blobs which don't exist in\n+\t\t\t * packs do exist via loose objects\". Confused?\n+\t\t\t */\n+\t\t\treturn;\n+\t\t}\n+\n+\t\tentry = create_object_entry(oid, type, pack_name_hash(name),\n+\t\t\t\t\t    0, name && no_try_delta(name),\n+\t\t\t\t\t    pack, offset);\n+\t}\n+\n+\tif (mtime > oe_cruft_mtime(&to_pack, entry))\n+\t\toe_set_cruft_mtime(&to_pack, entry, mtime);\n+\treturn;\n+}\n+\n+static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n+{\n+\tstruct string_list_item *item = NULL;\n+\tfor_each_string_list_item(item, packs) {\n+\t\tstruct packed_git *p = item->util;\n+\t\tif (!p)\n+\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n+\t\tp->pack_keep_in_core = keep;\n+\t}\n+}\n+\n+static void add_unreachable_loose_objects(void);\n+static void add_objects_in_unpacked_packs(void);\n+\n+static void enumerate_cruft_objects(void)\n+{\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Enumerating cruft objects\"), 0);\n+\n+\tadd_objects_in_unpacked_packs();\n+\tadd_unreachable_loose_objects();\n+\n+\tstop_progress(&progress_state);\n+}\n+\n+static void read_cruft_objects(void)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct string_list discard_packs = STRING_LIST_INIT_DUP;\n+\tstruct string_list fresh_packs = STRING_LIST_INIT_DUP;\n+\tstruct packed_git *p;\n+\n+\tignore_packed_keep_in_core = 1;\n+\n+\twhile (strbuf_getline(&buf, stdin) != EOF) {\n+\t\tif (!buf.len)\n+\t\t\tcontinue;\n+\n+\t\tif (*buf.buf == '-')\n+\t\t\tstring_list_append(&discard_packs, buf.buf + 1);\n+\t\telse\n+\t\t\tstring_list_append(&fresh_packs, buf.buf);\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstring_list_sort(&discard_packs);\n+\tstring_list_sort(&fresh_packs);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tconst char *pack_name = pack_basename(p);\n+\t\tstruct string_list_item *item;\n+\n+\t\titem = string_list_lookup(&fresh_packs, pack_name);\n+\t\tif (!item)\n+\t\t\titem = string_list_lookup(&discard_packs, pack_name);\n+\n+\t\tif (item) {\n+\t\t\titem->util = p;\n+\t\t} else {\n+\t\t\t/*\n+\t\t\t * This pack wasn't mentioned in either the \"fresh\" or\n+\t\t\t * \"discard\" list, so the caller didn't know about it.\n+\t\t\t *\n+\t\t\t * Mark it as kept so that its objects are ignored by\n+\t\t\t * add_unseen_recent_objects_to_traversal(). We'll\n+\t\t\t * unmark it before starting the traversal so it doesn't\n+\t\t\t * halt the traversal early.\n+\t\t\t */\n+\t\t\tp->pack_keep_in_core = 1;\n+\t\t}\n+\t}\n+\n+\tmark_pack_kept_in_core(&fresh_packs, 1);\n+\tmark_pack_kept_in_core(&discard_packs, 0);\n+\n+\tif (cruft_expiration)\n+\t\tdie(\"--cruft-expiration not yet implemented\");\n+\telse\n+\t\tenumerate_cruft_objects();\n+\n+\tstrbuf_release(&buf);\n+\tstring_list_clear(&discard_packs, 0);\n+\tstring_list_clear(&fresh_packs, 0);\n+}\n+\n static void read_object_list_from_stdin(void)\n {\n \tchar line[GIT_MAX_HEXSZ + 1 + PATH_MAX + 2];\n@@ -3529,7 +3664,24 @@ static int add_object_in_unpacked_pack(const struct object_id *oid,\n \t\t\t\t       uint32_t pos,\n \t\t\t\t       void *_data)\n {\n-\tadd_object_entry(oid, OBJ_NONE, \"\", 0);\n+\tif (cruft) {\n+\t\toff_t offset;\n+\t\ttime_t mtime;\n+\n+\t\tif (pack->is_cruft) {\n+\t\t\tif (load_pack_mtimes(pack) < 0)\n+\t\t\t\tdie(_(\"could not load cruft pack .mtimes\"));\n+\t\t\tmtime = nth_packed_mtime(pack, pos);\n+\t\t} else {\n+\t\t\tmtime = pack->mtime;\n+\t\t}\n+\t\toffset = nth_packed_object_offset(pack, pos);\n+\n+\t\tadd_cruft_object_entry(oid, OBJ_NONE, pack, offset,\n+\t\t\t\t       NULL, mtime);\n+\t} else {\n+\t\tadd_object_entry(oid, OBJ_NONE, \"\", 0);\n+\t}\n \treturn 0;\n }\n \n@@ -3553,7 +3705,19 @@ static int add_loose_object(const struct object_id *oid, const char *path,\n \t\treturn 0;\n \t}\n \n-\tadd_object_entry(oid, type, \"\", 0);\n+\tif (cruft) {\n+\t\tstruct stat st;\n+\t\tif (stat(path, &st) < 0) {\n+\t\t\tif (errno == ENOENT)\n+\t\t\t\treturn 0;\n+\t\t\treturn error_errno(\"unable to stat %s\", oid_to_hex(oid));\n+\t\t}\n+\n+\t\tadd_cruft_object_entry(oid, type, NULL, 0, NULL,\n+\t\t\t\t       st.st_mtime);\n+\t} else {\n+\t\tadd_object_entry(oid, type, \"\", 0);\n+\t}\n \treturn 0;\n }\n \n@@ -3870,6 +4034,20 @@ static int option_parse_unpack_unreachable(const struct option *opt,\n \treturn 0;\n }\n \n+static int option_parse_cruft_expiration(const struct option *opt,\n+\t\t\t\t\t const char *arg, int unset)\n+{\n+\tif (unset) {\n+\t\tcruft = 0;\n+\t\tcruft_expiration = 0;\n+\t} else {\n+\t\tcruft = 1;\n+\t\tif (arg)\n+\t\t\tcruft_expiration = approxidate(arg);\n+\t}\n+\treturn 0;\n+}\n+\n struct po_filter_data {\n \tunsigned have_revs:1;\n \tstruct rev_info revs;\n@@ -3959,6 +4137,10 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tOPT_CALLBACK_F(0, \"unpack-unreachable\", NULL, N_(\"time\"),\n \t\t  N_(\"unpack unreachable objects newer than <time>\"),\n \t\t  PARSE_OPT_OPTARG, option_parse_unpack_unreachable),\n+\t\tOPT_BOOL(0, \"cruft\", &cruft, N_(\"create a cruft pack\")),\n+\t\tOPT_CALLBACK_F(0, \"cruft-expiration\", NULL, N_(\"time\"),\n+\t\t  N_(\"expire cruft objects older than <time>\"),\n+\t\t  PARSE_OPT_OPTARG, option_parse_cruft_expiration),\n \t\tOPT_BOOL(0, \"sparse\", &sparse,\n \t\t\t N_(\"use the sparse reachability algorithm\")),\n \t\tOPT_BOOL(0, \"thin\", &thin,\n@@ -4085,7 +4267,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \n \tif (!HAVE_THREADS && delta_search_threads != 1)\n \t\twarning(_(\"no threads support, ignoring --threads\"));\n-\tif (!pack_to_stdout && !pack_size_limit)\n+\tif (!pack_to_stdout && !pack_size_limit && !cruft)\n \t\tpack_size_limit = pack_size_limit_cfg;\n \tif (pack_to_stdout && pack_size_limit)\n \t\tdie(_(\"--max-pack-size cannot be used to build a pack for transfer\"));\n@@ -4112,6 +4294,15 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (stdin_packs && use_internal_rev_list)\n \t\tdie(_(\"cannot use internal rev list with --stdin-packs\"));\n \n+\tif (cruft) {\n+\t\tif (use_internal_rev_list)\n+\t\t\tdie(_(\"cannot use internal rev list with --cruft\"));\n+\t\tif (stdin_packs)\n+\t\t\tdie(_(\"cannot use --stdin-packs with --cruft\"));\n+\t\tif (pack_size_limit)\n+\t\t\tdie(_(\"cannot use --max-pack-size with --cruft\"));\n+\t}\n+\n \t/*\n \t * \"soft\" reasons not to use bitmaps - for on-disk repack by default we want\n \t *\n@@ -4168,7 +4359,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t    the_repository);\n \tprepare_packing_data(the_repository, &to_pack);\n \n-\tif (progress)\n+\tif (progress && !cruft)\n \t\tprogress_state = start_progress(_(\"Enumerating objects\"), 0);\n \tif (stdin_packs) {\n \t\t/* avoids adding objects in excluded packs */\n@@ -4176,6 +4367,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tread_packs_list_from_stdin();\n \t\tif (rev_list_unpacked)\n \t\t\tadd_unreachable_loose_objects();\n+\t} else if (cruft) {\n+\t\tread_cruft_objects();\n \t} else if (!use_internal_rev_list) {\n \t\tread_object_list_from_stdin();\n \t} else if (pfd.have_revs) {\ndiff --git a/object-file.c b/object-file.c\nindex 5ffbf3d4fd..ff0cffe68e 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -997,7 +997,7 @@ int has_loose_object_nonlocal(const struct object_id *oid)\n \treturn check_and_freshen_nonlocal(oid, 0);\n }\n \n-static int has_loose_object(const struct object_id *oid)\n+int has_loose_object(const struct object_id *oid)\n {\n \treturn check_and_freshen(oid, 0);\n }\ndiff --git a/object-store.h b/object-store.h\nindex 2c4671ed7a..c41609e8db 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -330,6 +330,8 @@ int repo_has_object_file_with_flags(struct repository *r,\n  */\n int has_loose_object_nonlocal(const struct object_id *);\n \n+int has_loose_object(const struct object_id *);\n+\n /**\n  * format_object_header() is a thin wrapper around s xsnprintf() that\n  * writes the initial \"<type> <obj-len>\" part of the loose object\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nnew file mode 100755\nindex 0000000000..003ca7344e\n--- /dev/null\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -0,0 +1,218 @@\n+#!/bin/sh\n+\n+test_description='cruft pack related pack-objects tests'\n+. ./test-lib.sh\n+\n+objdir=.git/objects\n+packdir=$objdir/pack\n+\n+basic_cruft_pack_tests () {\n+\texpire=\"$1\"\n+\n+\ttest_expect_success \"unreachable loose objects are packed (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit base &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit loose &&\n+\n+\t\t\ttest-tool chmtime +2000 \"$objdir/$(test_oid_to_path \\\n+\t\t\t\t$(git rev-parse loose:loose.t))\" &&\n+\t\t\ttest-tool chmtime +1000 \"$objdir/$(test_oid_to_path \\\n+\t\t\t\t$(git rev-parse loose^{tree}))\" &&\n+\n+\t\t\t(\n+\t\t\t\tgit rev-list --objects --no-object-names base..loose |\n+\t\t\t\twhile read oid\n+\t\t\t\tdo\n+\t\t\t\t\tpath=\"$objdir/$(test_oid_to_path \"$oid\")\" &&\n+\t\t\t\t\tprintf \"%s %d\\n\" \"$oid\" \"$(test-tool chmtime --get \"$path\")\"\n+\t\t\t\tdone |\n+\t\t\t\tsort -k1\n+\t\t\t) >expect &&\n+\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t\t--cruft-expiration=\"$expire\" $packdir/pack)\" &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"unreachable packed objects are packed (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit packed &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit other &&\n+\n+\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\t\tother=\"$(git pack-objects --delta-base-offset \\\n+\t\t\t\t$packdir/pack <objects)\" &&\n+\t\t\tgit prune-packed &&\n+\n+\t\t\ttest-tool chmtime --get -100 \"$packdir/pack-$other.pack\" >expect &&\n+\n+\t\t\tcruft=\"$(git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-pack-$other.pack\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\t\tcut -d\" \" -f2 <actual.raw | sort -u >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"unreachable cruft objects are repacked (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit packed &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit other &&\n+\n+\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\n+\t\t\tcruft_a=\"$(echo $keep | git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack)\" &&\n+\t\t\tgit prune-packed &&\n+\t\t\tcruft_b=\"$(git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-pack-$cruft_a.pack\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft_a.mtimes\" >expect.raw &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft_b.mtimes\" >actual.raw &&\n+\n+\t\t\tsort <expect.raw >expect &&\n+\t\t\tsort <actual.raw >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"multiple cruft packs (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\tgit repack -Ad &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\n+\t\t\ttest_commit cruft &&\n+\t\t\tloose=\"$objdir/$(test_oid_to_path $(git rev-parse cruft))\" &&\n+\n+\t\t\t# generate three copies of the cruft object in different\n+\t\t\t# cruft packs, each with a unique mtime:\n+\t\t\t#   - one expired (1000 seconds ago)\n+\t\t\t#   - two non-expired (one 1000 seconds in the future,\n+\t\t\t#     one 1500 seconds in the future)\n+\t\t\ttest-tool chmtime =-1000 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-A <<-EOF &&\n+\t\t\t$keep\n+\t\t\tEOF\n+\t\t\ttest-tool chmtime =+1000 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-B <<-EOF &&\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\tEOF\n+\t\t\ttest-tool chmtime =+1500 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-C <<-EOF &&\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-B-*.pack))\n+\t\t\tEOF\n+\n+\t\t\t# ensure the resulting cruft pack takes the most recent\n+\t\t\t# mtime among all copies\n+\t\t\tcruft=\"$(git pack-objects --cruft \\\n+\t\t\t\t--cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-B-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-C-*.pack))\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\n+\t\t\ttest-tool pack-mtimes \"$(basename $(ls $packdir/pack-C-*.mtimes))\" >expect.raw &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\t\tsort expect.raw >expect &&\n+\t\t\tsort actual.raw >actual &&\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"cruft packs tolerate missing trees (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\ttest_commit cruft &&\n+\n+\t\t\ttree=\"$(git rev-parse cruft^{tree})\" &&\n+\n+\t\t\tgit reset --hard reachable &&\n+\t\t\tgit tag -d cruft &&\n+\t\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t\t# remove the unreachable tree, but leave the commit\n+\t\t\t# which has it as its root tree intact\n+\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$tree\")\" &&\n+\n+\t\t\tgit repack -Ad &&\n+\t\t\tbasename $(ls $packdir/pack-*.pack) >in &&\n+\t\t\tgit pack-objects --cruft --cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <in\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"cruft packs tolerate missing blobs (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\ttest_commit cruft &&\n+\n+\t\t\tblob=\"$(git rev-parse cruft:cruft.t)\" &&\n+\n+\t\t\tgit reset --hard reachable &&\n+\t\t\tgit tag -d cruft &&\n+\t\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t\t# remove the unreachable blob, but leave the commit (and\n+\t\t\t# the root tree of that commit) intact\n+\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$blob\")\" &&\n+\n+\t\t\tgit repack -Ad &&\n+\t\t\tbasename $(ls $packdir/pack-*.pack) >in &&\n+\t\t\tgit pack-objects --cruft --cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <in\n+\t\t)\n+\t'\n+}\n+\n+basic_cruft_pack_tests never\n+\n+test_done\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455441","messageId":"1740b8ef019d5ccae4ec87a96a5db085c8448178.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 11/17] builtin/pack-objects.c: --cruft with expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:19Z","receivedAt":"2022-05-18T23:11:40Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a previous patch, pack-objects learned how to generate a cruft pack\nso long as no objects are dropped.\n\nThis patch teaches pack-objects to handle the case where a non-never\n`--cruft-expiration` value is passed. This case is slightly more\ncomplicated than before, because we want pack-objects to save\nunreachable objects which would have been pruned when there is another\nrecent (i.e., non-prunable) unreachable object which reaches the other.\nWe'll call these objects \"unreachable but reachable-from-recent\".\n\nHere is how pack-objects handles `--cruft-expiration`:\n\n  - Instead of adding all objects outside of the kept pack(s) into the\n    packing list, only handle the ones whose mtime is within the grace\n    period.\n\n  - Construct a reachability traversal whose tips are the\n    unreachable-but-recent objects.\n\n  - Then, walk along that traversal, stopping if we reach an object in\n    the kept pack. At each step along the traversal, we add the object\n    we are visiting to the packing list.\n\nIn the majority of these cases, any object we visit in this traversal\nwill already be in our packing list. But we will sometimes encounter\nreachable-from-recent cruft objects, which we want to retain even if\nthey aged out of the grace period.\n\nThe most subtle point of this process is that we actually don't need to\nbother to update the rescued object's mtime. Even though we will write\nan .mtimes file with a value that is older than the expiration window,\nit will continue to survive cruft repacks so long as any objects which\nreach it haven't aged out.\n\nThat is, a future repack will also exclude that object from the initial\npacking list, only to discover it later on when doing the reachability\ntraversal.\n\nFinally, stopping early once an object is found in a kept pack is safe\nto do because the kept packs ordinarily represent which packs will\nsurvive after repacking. Assuming that it _isn't_ safe to halt a\ntraversal early would mean that there is some ancestor object which is\nmissing, which implies repository corruption (i.e., the complete set of\nreachable objects isn't present).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c        |  84 +++++++++++++++++++-\n reachable.h                   |   4 +-\n t/t5329-pack-objects-cruft.sh | 143 ++++++++++++++++++++++++++++++++++\n 3 files changed, 228 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 3b8bf6a3dd..8decc9dc0c 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -3447,6 +3447,44 @@ static void add_cruft_object_entry(const struct object_id *oid, enum object_type\n \treturn;\n }\n \n+static void show_cruft_object(struct object *obj, const char *name, void *data)\n+{\n+\t/*\n+\t * if we did not record it earlier, it's at least as old as our\n+\t * expiration value. Rather than find it exactly, just use that\n+\t * value.  This may bump it forward from its real mtime, but it\n+\t * will still be \"too old\" next time we run with the same\n+\t * expiration.\n+\t *\n+\t * if obj does appear in the packing list, this call is a noop (or may\n+\t * set the namehash).\n+\t */\n+\tadd_cruft_object_entry(&obj->oid, obj->type, NULL, 0, name, cruft_expiration);\n+}\n+\n+static void show_cruft_commit(struct commit *commit, void *data)\n+{\n+\tshow_cruft_object((struct object*)commit, NULL, data);\n+}\n+\n+static int cruft_include_check_obj(struct object *obj, void *data)\n+{\n+\treturn !has_object_kept_pack(&obj->oid, IN_CORE_KEEP_PACKS);\n+}\n+\n+static int cruft_include_check(struct commit *commit, void *data)\n+{\n+\treturn cruft_include_check_obj((struct object*)commit, data);\n+}\n+\n+static void set_cruft_mtime(const struct object *object,\n+\t\t\t    struct packed_git *pack,\n+\t\t\t    off_t offset, time_t mtime)\n+{\n+\tadd_cruft_object_entry(&object->oid, object->type, pack, offset, NULL,\n+\t\t\t       mtime);\n+}\n+\n static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n {\n \tstruct string_list_item *item = NULL;\n@@ -3472,6 +3510,50 @@ static void enumerate_cruft_objects(void)\n \tstop_progress(&progress_state);\n }\n \n+static void enumerate_and_traverse_cruft_objects(struct string_list *fresh_packs)\n+{\n+\tstruct packed_git *p;\n+\tstruct rev_info revs;\n+\tint ret;\n+\n+\trepo_init_revisions(the_repository, &revs, NULL);\n+\n+\trevs.tag_objects = 1;\n+\trevs.tree_objects = 1;\n+\trevs.blob_objects = 1;\n+\n+\trevs.include_check = cruft_include_check;\n+\trevs.include_check_obj = cruft_include_check_obj;\n+\n+\trevs.ignore_missing_links = 1;\n+\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Enumerating cruft objects\"), 0);\n+\tret = add_unseen_recent_objects_to_traversal(&revs, cruft_expiration,\n+\t\t\t\t\t\t     set_cruft_mtime, 1);\n+\tstop_progress(&progress_state);\n+\n+\tif (ret)\n+\t\tdie(_(\"unable to add cruft objects\"));\n+\n+\t/*\n+\t * Re-mark only the fresh packs as kept so that objects in\n+\t * unknown packs do not halt the reachability traversal early.\n+\t */\n+\tfor (p = get_all_packs(the_repository); p; p = p->next)\n+\t\tp->pack_keep_in_core = 0;\n+\tmark_pack_kept_in_core(fresh_packs, 1);\n+\n+\tif (prepare_revision_walk(&revs))\n+\t\tdie(_(\"revision walk setup failed\"));\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Traversing cruft objects\"), 0);\n+\tnr_seen = 0;\n+\ttraverse_commit_list(&revs, show_cruft_commit, show_cruft_object, NULL);\n+\n+\tstop_progress(&progress_state);\n+}\n+\n static void read_cruft_objects(void)\n {\n \tstruct strbuf buf = STRBUF_INIT;\n@@ -3523,7 +3605,7 @@ static void read_cruft_objects(void)\n \tmark_pack_kept_in_core(&discard_packs, 0);\n \n \tif (cruft_expiration)\n-\t\tdie(\"--cruft-expiration not yet implemented\");\n+\t\tenumerate_and_traverse_cruft_objects(&fresh_packs);\n \telse\n \t\tenumerate_cruft_objects();\n \ndiff --git a/reachable.h b/reachable.h\nindex b776761baa..020a887b99 100644\n--- a/reachable.h\n+++ b/reachable.h\n@@ -1,10 +1,10 @@\n #ifndef REACHEABLE_H\n #define REACHEABLE_H\n \n-#include \"object.h\"\n-\n struct progress;\n struct rev_info;\n+struct object;\n+struct packed_git;\n \n typedef void report_recent_object_fn(const struct object *, struct packed_git *,\n \t\t\t\t     off_t, time_t);\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 003ca7344e..939cdc297a 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -214,5 +214,148 @@ basic_cruft_pack_tests () {\n }\n \n basic_cruft_pack_tests never\n+basic_cruft_pack_tests 2.weeks.ago\n+\n+test_expect_success 'cruft tags rescue tagged objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit tagged &&\n+\t\tgit tag -a annotated -m tag &&\n+\n+\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\twhile read oid\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $oid)\"\n+\t\tdone <objects &&\n+\n+\t\ttest-tool chmtime -500 \\\n+\t\t\t\"$objdir/$(test_oid_to_path $(git rev-parse annotated))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\t\tcut -f1 -d\" \" <actual.raw | sort >actual &&\n+\n+\t\t(\n+\t\t\tcat objects &&\n+\t\t\tgit rev-parse annotated\n+\t\t) >expect.raw &&\n+\t\tsort <expect.raw >expect &&\n+\n+\t\ttest_cmp expect actual &&\n+\t\tcat actual\n+\t)\n+'\n+\n+test_expect_success 'cruft commits rescue parents, trees' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit old &&\n+\t\ttest_commit new &&\n+\n+\t\tgit rev-list --objects --no-object-names packed..new >objects &&\n+\t\twhile read object\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $object)\"\n+\t\tdone <objects &&\n+\t\ttest-tool chmtime +500 \"$objdir/$(test_oid_to_path \\\n+\t\t\t$(git rev-parse HEAD))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\tcut -d\" \" -f1 <actual.raw | sort >actual &&\n+\t\tsort <objects >expect &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'cruft trees rescue sub-trees, blobs' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\tmkdir -p dir/sub &&\n+\t\techo foo >foo &&\n+\t\techo bar >dir/bar &&\n+\t\techo baz >dir/sub/baz &&\n+\n+\t\ttest_tick &&\n+\t\tgit add . &&\n+\t\tgit commit -m \"pruned\" &&\n+\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD^{tree}))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:foo))\" &&\n+\t\ttest-tool chmtime  -500 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/bar))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/sub))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/sub/baz))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\t\tcut -f1 -d\" \" <actual.raw | sort >actual &&\n+\n+\t\tgit rev-parse HEAD:dir HEAD:dir/bar HEAD:dir/sub HEAD:dir/sub/baz >expect.raw &&\n+\t\tsort <expect.raw >expect &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'expired objects are pruned' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit pruned &&\n+\n+\t\tgit rev-list --objects --no-object-names packed..pruned >objects &&\n+\t\twhile read object\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $object)\"\n+\t\tdone <objects &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual &&\n+\t\ttest_must_be_empty actual\n+\t)\n+'\n \n test_done\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455442","messageId":"5992a72cbf9e8d076f1e312a789b40d52656ad3c.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 12/17] builtin/repack.c: support generating a cruft pack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:22Z","receivedAt":"2022-05-18T23:11:58Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Expose a way to split the contents of a repository into a main and cruft\npack when doing an all-into-one repack with `git repack --cruft -d`, and\na complementary configuration variable.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-repack.txt            |  11 ++\n Documentation/technical/cruft-packs.txt |   2 +-\n builtin/repack.c                        | 105 +++++++++++-\n t/t5329-pack-objects-cruft.sh           | 207 ++++++++++++++++++++++++\n 4 files changed, 319 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/git-repack.txt b/Documentation/git-repack.txt\nindex ee30edc178..0bf13893d8 100644\n--- a/Documentation/git-repack.txt\n+++ b/Documentation/git-repack.txt\n@@ -63,6 +63,17 @@ to the new separate pack will be written.\n \tAlso run  'git prune-packed' to remove redundant\n \tloose object files.\n \n+--cruft::\n+\tSame as `-a`, unless `-d` is used. Then any unreachable objects\n+\tare packed into a separate cruft pack. Unreachable objects can\n+\tbe pruned using the normal expiry rules with the next `git gc`\n+\tinvocation (see linkgit:git-gc[1]). Incompatible with `-k`.\n+\n+--cruft-expiration=<approxidate>::\n+\tExpire unreachable objects older than `<approxidate>`\n+\timmediately instead of waiting for the next `git gc` invocation.\n+\tOnly useful with `--cruft -d`.\n+\n -l::\n \tPass the `--local` option to 'git pack-objects'. See\n \tlinkgit:git-pack-objects[1].\ndiff --git a/Documentation/technical/cruft-packs.txt b/Documentation/technical/cruft-packs.txt\nindex c0f583cd48..d81f3a8982 100644\n--- a/Documentation/technical/cruft-packs.txt\n+++ b/Documentation/technical/cruft-packs.txt\n@@ -17,7 +17,7 @@ pruned according to normal expiry rules with the next 'git gc' invocation.\n \n Unreachable objects aren't removed immediately, since doing so could race with\n an incoming push which may reference an object which is about to be deleted.\n-Instead, those unreachable objects are stored as loose object and stay that way\n+Instead, those unreachable objects are stored as loose objects and stay that way\n until they are older than the expiration window, at which point they are removed\n by linkgit:git-prune[1].\n \ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex e7a3920c6d..593c18d4e8 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -18,12 +18,18 @@\n #include \"pack-bitmap.h\"\n #include \"refs.h\"\n \n+#define ALL_INTO_ONE 1\n+#define LOOSEN_UNREACHABLE 2\n+#define PACK_CRUFT 4\n+\n+static int pack_everything;\n static int delta_base_offset = 1;\n static int pack_kept_objects = -1;\n static int write_bitmaps = -1;\n static int use_delta_islands;\n static int run_update_server_info = 1;\n static char *packdir, *packtmp_name, *packtmp;\n+static char *cruft_expiration;\n \n static const char *const git_repack_usage[] = {\n \tN_(\"git repack [<options>]\"),\n@@ -305,9 +311,6 @@ static void repack_promisor_objects(const struct pack_objects_args *args,\n \t\tdie(_(\"could not finish pack-objects to repack promisor objects\"));\n }\n \n-#define ALL_INTO_ONE 1\n-#define LOOSEN_UNREACHABLE 2\n-\n struct pack_geometry {\n \tstruct packed_git **pack;\n \tuint32_t pack_nr, pack_alloc;\n@@ -344,6 +347,8 @@ static void init_pack_geometry(struct pack_geometry **geometry_p)\n \tfor (p = get_all_packs(the_repository); p; p = p->next) {\n \t\tif (!pack_kept_objects && p->pack_keep)\n \t\t\tcontinue;\n+\t\tif (p->is_cruft)\n+\t\t\tcontinue;\n \n \t\tALLOC_GROW(geometry->pack,\n \t\t\t   geometry->pack_nr + 1,\n@@ -605,6 +610,67 @@ static int write_midx_included_packs(struct string_list *include,\n \treturn finish_command(&cmd);\n }\n \n+static int write_cruft_pack(const struct pack_objects_args *args,\n+\t\t\t    const char *pack_prefix,\n+\t\t\t    struct string_list *names,\n+\t\t\t    struct string_list *existing_packs,\n+\t\t\t    struct string_list *existing_kept_packs)\n+{\n+\tstruct child_process cmd = CHILD_PROCESS_INIT;\n+\tstruct strbuf line = STRBUF_INIT;\n+\tstruct string_list_item *item;\n+\tFILE *in, *out;\n+\tint ret;\n+\n+\tprepare_pack_objects(&cmd, args);\n+\n+\tstrvec_push(&cmd.args, \"--cruft\");\n+\tif (cruft_expiration)\n+\t\tstrvec_pushf(&cmd.args, \"--cruft-expiration=%s\",\n+\t\t\t     cruft_expiration);\n+\n+\tstrvec_push(&cmd.args, \"--honor-pack-keep\");\n+\tstrvec_push(&cmd.args, \"--non-empty\");\n+\tstrvec_push(&cmd.args, \"--max-pack-size=0\");\n+\n+\tcmd.in = -1;\n+\n+\tret = start_command(&cmd);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\t/*\n+\t * names has a confusing double use: it both provides the list\n+\t * of just-written new packs, and accepts the name of the cruft\n+\t * pack we are writing.\n+\t *\n+\t * By the time it is read here, it contains only the pack(s)\n+\t * that were just written, which is exactly the set of packs we\n+\t * want to consider kept.\n+\t */\n+\tin = xfdopen(cmd.in, \"w\");\n+\tfor_each_string_list_item(item, names)\n+\t\tfprintf(in, \"%s-%s.pack\\n\", pack_prefix, item->string);\n+\tfor_each_string_list_item(item, existing_packs)\n+\t\tfprintf(in, \"-%s.pack\\n\", item->string);\n+\tfor_each_string_list_item(item, existing_kept_packs)\n+\t\tfprintf(in, \"%s.pack\\n\", item->string);\n+\tfclose(in);\n+\n+\tout = xfdopen(cmd.out, \"r\");\n+\twhile (strbuf_getline_lf(&line, out) != EOF) {\n+\t\tif (line.len != the_hash_algo->hexsz)\n+\t\t\tdie(_(\"repack: Expecting full hex object ID lines only \"\n+\t\t\t      \"from pack-objects.\"));\n+\t\tstring_list_append(names, line.buf);\n+\t}\n+\tfclose(out);\n+\n+\tstrbuf_release(&line);\n+\n+\treturn finish_command(&cmd);\n+}\n+\n int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n \tstruct child_process cmd = CHILD_PROCESS_INIT;\n@@ -621,7 +687,6 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tint show_progress;\n \n \t/* variables to be filled by option parsing */\n-\tint pack_everything = 0;\n \tint delete_redundant = 0;\n \tconst char *unpack_unreachable = NULL;\n \tint keep_unreachable = 0;\n@@ -636,6 +701,11 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tOPT_BIT('A', NULL, &pack_everything,\n \t\t\t\tN_(\"same as -a, and turn unreachable objects loose\"),\n \t\t\t\t   LOOSEN_UNREACHABLE | ALL_INTO_ONE),\n+\t\tOPT_BIT(0, \"cruft\", &pack_everything,\n+\t\t\t\tN_(\"same as -a, pack unreachable cruft objects separately\"),\n+\t\t\t\t   PACK_CRUFT),\n+\t\tOPT_STRING(0, \"cruft-expiration\", &cruft_expiration, N_(\"approxidate\"),\n+\t\t\t\tN_(\"with -C, expire objects older than this\")),\n \t\tOPT_BOOL('d', NULL, &delete_redundant,\n \t\t\t\tN_(\"remove redundant packs, and run git-prune-packed\")),\n \t\tOPT_BOOL('f', NULL, &po_args.no_reuse_delta,\n@@ -688,6 +758,15 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t    (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE)))\n \t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--keep-unreachable\", \"-A\");\n \n+\tif (pack_everything & PACK_CRUFT) {\n+\t\tpack_everything |= ALL_INTO_ONE;\n+\n+\t\tif (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE))\n+\t\t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--cruft\", \"-A\");\n+\t\tif (keep_unreachable)\n+\t\t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--cruft\", \"-k\");\n+\t}\n+\n \tif (write_bitmaps < 0) {\n \t\tif (!write_midx &&\n \t\t    (!(pack_everything & ALL_INTO_ONE) || !is_bare_repository()))\n@@ -771,7 +850,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (pack_everything & ALL_INTO_ONE) {\n \t\trepack_promisor_objects(&po_args, &names);\n \n-\t\tif (existing_nonkept_packs.nr && delete_redundant) {\n+\t\tif (existing_nonkept_packs.nr && delete_redundant &&\n+\t\t    !(pack_everything & PACK_CRUFT)) {\n \t\t\tfor_each_string_list_item(item, &names) {\n \t\t\t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s-%s.pack\",\n \t\t\t\t\t     packtmp_name, item->string);\n@@ -833,6 +913,21 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (!names.nr && !po_args.quiet)\n \t\tprintf_ln(_(\"Nothing new to pack.\"));\n \n+\tif (pack_everything & PACK_CRUFT) {\n+\t\tconst char *pack_prefix;\n+\t\tif (!skip_prefix(packtmp, packdir, &pack_prefix))\n+\t\t\tdie(_(\"pack prefix %s does not begin with objdir %s\"),\n+\t\t\t    packtmp, packdir);\n+\t\tif (*pack_prefix == '/')\n+\t\t\tpack_prefix++;\n+\n+\t\tret = write_cruft_pack(&po_args, pack_prefix, &names,\n+\t\t\t\t       &existing_nonkept_packs,\n+\t\t\t\t       &existing_kept_packs);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t}\n+\n \tfor_each_string_list_item(item, &names) {\n \t\titem->util = (void *)(uintptr_t)populate_pack_exts(item->string);\n \t}\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 939cdc297a..06c550c958 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -358,4 +358,211 @@ test_expect_success 'expired objects are pruned' '\n \t)\n '\n \n+test_expect_success 'repack --cruft generates a cruft pack' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit branch -M main &&\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit unreachable &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d unreachable &&\n+\t\t# objects are not cruft if they are contained in the reflogs\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\t\tsort <reachable.raw >reachable &&\n+\t\tcomm -13 reachable objects >unreachable &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tcruft=$(basename $(ls $packdir/pack-*.mtimes) .mtimes) &&\n+\t\tpack=$(basename $(ls $packdir/pack-*.pack | grep -v $cruft) .pack) &&\n+\n+\t\tgit show-index <$packdir/$pack.idx >actual.raw &&\n+\t\tcut -f2 -d\" \" actual.raw | sort >actual &&\n+\t\ttest_cmp reachable actual &&\n+\n+\t\tgit show-index <$packdir/$cruft.idx >actual.raw &&\n+\t\tcut -f2 -d\" \" actual.raw | sort >actual &&\n+\t\ttest_cmp unreachable actual\n+\t)\n+'\n+\n+test_expect_success 'loose objects mtimes upsert others' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit cruft &&\n+\t\t# incremental repack, leaving existing objects loose (so\n+\t\t# they can be \"freshened\")\n+\t\tgit repack &&\n+\n+\t\ttip=\"$(git rev-parse cruft)\" &&\n+\t\tpath=\"$objdir/$(test_oid_to_path \"$(git rev-parse cruft)\")\" &&\n+\t\ttest-tool chmtime --get +1000 \"$path\" >expect &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tmtimes=\"$(basename $(ls $packdir/pack-*.mtimes))\" &&\n+\t\ttest-tool pack-mtimes \"$mtimes\" >actual.raw &&\n+\t\tgrep \"$tip\" actual.raw | cut -d\" \" -f2 >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'cruft packs are not included in geometric repack' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit cruft &&\n+\t\tgit repack -d &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft &&\n+\n+\t\tfind $packdir -type f | sort >before &&\n+\t\tgit repack --geometric=2 -d &&\n+\t\tfind $packdir -type f | sort >after &&\n+\n+\t\ttest_cmp before after\n+\t)\n+'\n+\n+test_expect_success 'repack --geometric collects once-cruft objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\tgit rm -rf . &&\n+\t\ttest_commit --no-tag cruft &&\n+\t\tcruft=\"$(git rev-parse HEAD)\" &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t# Pack the objects created in the previous step into a cruft\n+\t\t# pack. Intentionally leave loose copies of those objects\n+\t\t# around so we can pick them up in a subsequent --geometric\n+\t\t# reapack.\n+\t\tgit repack --cruft &&\n+\n+\t\t# Now make those objects reachable, and ensure that they are\n+\t\t# packed into the new pack created via a --geometric repack.\n+\t\tgit update-ref refs/heads/other $cruft &&\n+\n+\t\t# Without this object, the set of unpacked objects is exactly\n+\t\t# the set of objects already in the cruft pack. Tweak that set\n+\t\t# to ensure we do not overwrite the cruft pack entirely.\n+\t\ttest_commit reachable2 &&\n+\n+\t\tfind $packdir -name \"pack-*.idx\" | sort >before &&\n+\t\tgit repack --geometric=2 -d &&\n+\t\tfind $packdir -name \"pack-*.idx\" | sort >after &&\n+\n+\t\t{\n+\t\t\tgit rev-list --objects --no-object-names $cruft &&\n+\t\t\tgit rev-list --objects --no-object-names reachable..reachable2\n+\t\t} >want.raw &&\n+\t\tsort want.raw >want &&\n+\n+\t\tpack=$(comm -13 before after) &&\n+\t\tgit show-index <$pack >objects.raw &&\n+\n+\t\tcut -d\" \" -f2 objects.raw | sort >got &&\n+\n+\t\ttest_cmp want got\n+\t)\n+'\n+\n+test_expect_success 'cruft repack with no reachable objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\t\tgit repack -ad &&\n+\n+\t\tbase=\"$(git rev-parse base)\" &&\n+\n+\t\tgit for-each-ref --format=\"delete %(refname)\" >in &&\n+\t\tgit update-ref --stdin <in &&\n+\t\tgit reflog expire --all --expire=all &&\n+\t\trm -fr .git/index &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tgit cat-file -t $base\n+\t)\n+'\n+\n+test_expect_success 'cruft repack ignores --max-pack-size' '\n+\tgit init max-pack-size &&\n+\t(\n+\t\tcd max-pack-size &&\n+\t\ttest_commit base &&\n+\t\t# two cruft objects which exceed the maximum pack size\n+\t\ttest-tool genrandom foo 1048576 | git hash-object --stdin -w &&\n+\t\ttest-tool genrandom bar 1048576 | git hash-object --stdin -w &&\n+\t\tgit repack --cruft --max-pack-size=1M &&\n+\t\tfind $packdir -name \"*.mtimes\" >cruft &&\n+\t\ttest_line_count = 1 cruft &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(cat cruft)\")\" >objects &&\n+\t\ttest_line_count = 2 objects\n+\t)\n+'\n+\n+test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n+\t(\n+\t\tcd max-pack-size &&\n+\t\t# repack everything back together to remove the existing cruft\n+\t\t# pack (but to keep its objects)\n+\t\tgit repack -adk &&\n+\t\tgit -c pack.packSizeLimit=1M repack --cruft &&\n+\t\t# ensure the same post condition is met when --max-pack-size\n+\t\t# would otherwise be inferred from the configuration\n+\t\tfind $packdir -name \"*.mtimes\" >cruft &&\n+\t\ttest_line_count = 1 cruft &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(cat cruft)\")\" >objects &&\n+\t\ttest_line_count = 2 objects\n+\t)\n+'\n+\n test_done\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455443","messageId":"1d775f9850f00b0c3d1e9133669a6365c8d7bbba.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 04/17] chunk-format.h: extract oid_version()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:01Z","receivedAt":"2022-05-18T23:12:00Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"There are three definitions of an identical function which converts\n`the_hash_algo` into either 1 (for SHA-1) or 2 (for SHA-256). There is a\ncopy of this function for writing both the commit-graph and\nmulti-pack-index file, and another inline definition used to write the\n.rev header.\n\nConsolidate these into a single definition in chunk-format.h. It's not\nclear that this is the best header to define this function in, but it\nshould do for now.\n\n(Worth noting, the .rev caller expects a 4-byte unsigned, but the other\ntwo callers work with a single unsigned byte. The consolidated version\nuses the latter type, and lets the compiler widen it when required).\n\nAnother caller will be added in a subsequent patch.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n chunk-format.c | 12 ++++++++++++\n chunk-format.h |  3 +++\n commit-graph.c | 18 +++---------------\n midx.c         | 18 +++---------------\n pack-write.c   | 15 ++-------------\n 5 files changed, 23 insertions(+), 43 deletions(-)\n\ndiff --git a/chunk-format.c b/chunk-format.c\nindex 1c3dca62e2..0275b74a89 100644\n--- a/chunk-format.c\n+++ b/chunk-format.c\n@@ -181,3 +181,15 @@ int read_chunk(struct chunkfile *cf,\n \n \treturn CHUNK_NOT_FOUND;\n }\n+\n+uint8_t oid_version(const struct git_hash_algo *algop)\n+{\n+\tswitch (hash_algo_by_ptr(algop)) {\n+\tcase GIT_HASH_SHA1:\n+\t\treturn 1;\n+\tcase GIT_HASH_SHA256:\n+\t\treturn 2;\n+\tdefault:\n+\t\tdie(_(\"invalid hash version\"));\n+\t}\n+}\ndiff --git a/chunk-format.h b/chunk-format.h\nindex 9ccbe00377..7885aa0848 100644\n--- a/chunk-format.h\n+++ b/chunk-format.h\n@@ -2,6 +2,7 @@\n #define CHUNK_FORMAT_H\n \n #include \"git-compat-util.h\"\n+#include \"hash.h\"\n \n struct hashfile;\n struct chunkfile;\n@@ -65,4 +66,6 @@ int read_chunk(struct chunkfile *cf,\n \t       chunk_read_fn fn,\n \t       void *data);\n \n+uint8_t oid_version(const struct git_hash_algo *algop);\n+\n #endif\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 06107beedc..066d82ed6a 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -193,18 +193,6 @@ char *get_commit_graph_chain_filename(struct object_directory *odb)\n \treturn xstrfmt(\"%s/info/commit-graphs/commit-graph-chain\", odb->path);\n }\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n static struct commit_graph *alloc_commit_graph(void)\n {\n \tstruct commit_graph *g = xcalloc(1, sizeof(*g));\n@@ -365,9 +353,9 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n \t}\n \n \thash_version = *(unsigned char*)(data + 5);\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != oid_version(the_hash_algo)) {\n \t\terror(_(\"commit-graph hash version %X does not match version %X\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, oid_version(the_hash_algo));\n \t\treturn NULL;\n \t}\n \n@@ -1924,7 +1912,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \thashwrite_be32(f, GRAPH_SIGNATURE);\n \n \thashwrite_u8(f, GRAPH_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, oid_version(the_hash_algo));\n \thashwrite_u8(f, get_num_chunks(cf));\n \thashwrite_u8(f, ctx->num_commit_graphs_after - 1);\n \ndiff --git a/midx.c b/midx.c\nindex 3db0e47735..c617c51cd0 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -41,18 +41,6 @@\n \n #define PACK_EXPIRED UINT_MAX\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n const unsigned char *get_midx_checksum(struct multi_pack_index *m)\n {\n \treturn m->data + m->data_len - the_hash_algo->rawsz;\n@@ -134,9 +122,9 @@ struct multi_pack_index *load_multi_pack_index(const char *object_dir, int local\n \t\t      m->version);\n \n \thash_version = m->data[MIDX_BYTE_HASH_VERSION];\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != oid_version(the_hash_algo)) {\n \t\terror(_(\"multi-pack-index hash version %u does not match version %u\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, oid_version(the_hash_algo));\n \t\tgoto cleanup_fail;\n \t}\n \tm->hash_len = the_hash_algo->rawsz;\n@@ -420,7 +408,7 @@ static size_t write_midx_header(struct hashfile *f,\n {\n \thashwrite_be32(f, MIDX_SIGNATURE);\n \thashwrite_u8(f, MIDX_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, oid_version(the_hash_algo));\n \thashwrite_u8(f, num_chunks);\n \thashwrite_u8(f, 0); /* unused */\n \thashwrite_be32(f, num_packs);\ndiff --git a/pack-write.c b/pack-write.c\nindex a2adc565f4..27b171e440 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -2,6 +2,7 @@\n #include \"pack.h\"\n #include \"csum-file.h\"\n #include \"remote.h\"\n+#include \"chunk-format.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -181,21 +182,9 @@ static int pack_order_cmp(const void *va, const void *vb, void *ctx)\n \n static void write_rev_header(struct hashfile *f)\n {\n-\tuint32_t oid_version;\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\toid_version = 1;\n-\t\tbreak;\n-\tcase GIT_HASH_SHA256:\n-\t\toid_version = 2;\n-\t\tbreak;\n-\tdefault:\n-\t\tdie(\"write_rev_header: unknown hash version\");\n-\t}\n-\n \thashwrite_be32(f, RIDX_SIGNATURE);\n \thashwrite_be32(f, RIDX_VERSION);\n-\thashwrite_be32(f, oid_version);\n+\thashwrite_be32(f, oid_version(the_hash_algo));\n }\n \n static void write_rev_index_positions(struct hashfile *f,\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455444","messageId":"6172861bd9caa3036d0042293b54a9de841ae2b5.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 05/17] pack-mtimes: support writing pack .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:04Z","receivedAt":"2022-05-18T23:12:04Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Now that the `.mtimes` format is defined, supplement the pack-write API\nto be able to conditionally write an `.mtimes` file along with a pack by\nsetting an additional flag and passing an oidmap that contains the\ntimestamps corresponding to each object in the pack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n pack-objects.c |  6 ++++\n pack-objects.h | 25 ++++++++++++++++\n pack-write.c   | 77 ++++++++++++++++++++++++++++++++++++++++++++++++++\n pack.h         |  1 +\n 4 files changed, 109 insertions(+)\n\ndiff --git a/pack-objects.c b/pack-objects.c\nindex fe2a4eace9..272e8d4517 100644\n--- a/pack-objects.c\n+++ b/pack-objects.c\n@@ -170,6 +170,9 @@ struct object_entry *packlist_alloc(struct packing_data *pdata,\n \n \t\tif (pdata->layer)\n \t\t\tREALLOC_ARRAY(pdata->layer, pdata->nr_alloc);\n+\n+\t\tif (pdata->cruft_mtime)\n+\t\t\tREALLOC_ARRAY(pdata->cruft_mtime, pdata->nr_alloc);\n \t}\n \n \tnew_entry = pdata->objects + pdata->nr_objects++;\n@@ -198,6 +201,9 @@ struct object_entry *packlist_alloc(struct packing_data *pdata,\n \tif (pdata->layer)\n \t\tpdata->layer[pdata->nr_objects - 1] = 0;\n \n+\tif (pdata->cruft_mtime)\n+\t\tpdata->cruft_mtime[pdata->nr_objects - 1] = 0;\n+\n \treturn new_entry;\n }\n \ndiff --git a/pack-objects.h b/pack-objects.h\nindex dca2351ef9..393b9db546 100644\n--- a/pack-objects.h\n+++ b/pack-objects.h\n@@ -168,6 +168,14 @@ struct packing_data {\n \t/* delta islands */\n \tunsigned int *tree_depth;\n \tunsigned char *layer;\n+\n+\t/*\n+\t * Used when writing cruft packs.\n+\t *\n+\t * Object mtimes are stored in pack order when writing, but\n+\t * written out in lexicographic (index) order.\n+\t */\n+\tuint32_t *cruft_mtime;\n };\n \n void prepare_packing_data(struct repository *r, struct packing_data *pdata);\n@@ -289,4 +297,21 @@ static inline void oe_set_layer(struct packing_data *pack,\n \tpack->layer[e - pack->objects] = layer;\n }\n \n+static inline uint32_t oe_cruft_mtime(struct packing_data *pack,\n+\t\t\t\t      struct object_entry *e)\n+{\n+\tif (!pack->cruft_mtime)\n+\t\treturn 0;\n+\treturn pack->cruft_mtime[e - pack->objects];\n+}\n+\n+static inline void oe_set_cruft_mtime(struct packing_data *pack,\n+\t\t\t\t      struct object_entry *e,\n+\t\t\t\t      uint32_t mtime)\n+{\n+\tif (!pack->cruft_mtime)\n+\t\tCALLOC_ARRAY(pack->cruft_mtime, pack->nr_alloc);\n+\tpack->cruft_mtime[e - pack->objects] = mtime;\n+}\n+\n #endif\ndiff --git a/pack-write.c b/pack-write.c\nindex 27b171e440..23c0342018 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -3,6 +3,10 @@\n #include \"csum-file.h\"\n #include \"remote.h\"\n #include \"chunk-format.h\"\n+#include \"pack-mtimes.h\"\n+#include \"oidmap.h\"\n+#include \"chunk-format.h\"\n+#include \"pack-objects.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -277,6 +281,70 @@ const char *write_rev_file_order(const char *rev_name,\n \treturn rev_name;\n }\n \n+static void write_mtimes_header(struct hashfile *f)\n+{\n+\thashwrite_be32(f, MTIMES_SIGNATURE);\n+\thashwrite_be32(f, MTIMES_VERSION);\n+\thashwrite_be32(f, oid_version(the_hash_algo));\n+}\n+\n+/*\n+ * Writes the object mtimes of \"objects\" for use in a .mtimes file.\n+ * Note that objects must be in lexicographic (index) order, which is\n+ * the expected ordering of these values in the .mtimes file.\n+ */\n+static void write_mtimes_objects(struct hashfile *f,\n+\t\t\t\t struct packing_data *to_pack,\n+\t\t\t\t struct pack_idx_entry **objects,\n+\t\t\t\t uint32_t nr_objects)\n+{\n+\tuint32_t i;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\tstruct object_entry *e = (struct object_entry*)objects[i];\n+\t\thashwrite_be32(f, oe_cruft_mtime(to_pack, e));\n+\t}\n+}\n+\n+static void write_mtimes_trailer(struct hashfile *f, const unsigned char *hash)\n+{\n+\thashwrite(f, hash, the_hash_algo->rawsz);\n+}\n+\n+static const char *write_mtimes_file(const char *mtimes_name,\n+\t\t\t\t     struct packing_data *to_pack,\n+\t\t\t\t     struct pack_idx_entry **objects,\n+\t\t\t\t     uint32_t nr_objects,\n+\t\t\t\t     const unsigned char *hash)\n+{\n+\tstruct hashfile *f;\n+\tint fd;\n+\n+\tif (!to_pack)\n+\t\tBUG(\"cannot call write_mtimes_file with NULL packing_data\");\n+\n+\tif (!mtimes_name) {\n+\t\tstruct strbuf tmp_file = STRBUF_INIT;\n+\t\tfd = odb_mkstemp(&tmp_file, \"pack/tmp_mtimes_XXXXXX\");\n+\t\tmtimes_name = strbuf_detach(&tmp_file, NULL);\n+\t} else {\n+\t\tunlink(mtimes_name);\n+\t\tfd = xopen(mtimes_name, O_CREAT|O_EXCL|O_WRONLY, 0600);\n+\t}\n+\tf = hashfd(fd, mtimes_name);\n+\n+\twrite_mtimes_header(f);\n+\twrite_mtimes_objects(f, to_pack, objects, nr_objects);\n+\twrite_mtimes_trailer(f, hash);\n+\n+\tif (adjust_shared_perm(mtimes_name) < 0)\n+\t\tdie(_(\"failed to make %s readable\"), mtimes_name);\n+\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE | CSUM_FSYNC);\n+\n+\treturn mtimes_name;\n+}\n+\n off_t write_pack_header(struct hashfile *f, uint32_t nr_entries)\n {\n \tstruct pack_header hdr;\n@@ -479,6 +547,7 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t char **idx_tmp_name)\n {\n \tconst char *rev_tmp_name = NULL;\n+\tconst char *mtimes_tmp_name = NULL;\n \n \tif (adjust_shared_perm(pack_tmp_name))\n \t\tdie_errno(\"unable to make temporary pack file readable\");\n@@ -491,9 +560,17 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \trev_tmp_name = write_rev_file(NULL, written_list, nr_written, hash,\n \t\t\t\t      pack_idx_opts->flags);\n \n+\tif (pack_idx_opts->flags & WRITE_MTIMES) {\n+\t\tmtimes_tmp_name = write_mtimes_file(NULL, to_pack, written_list,\n+\t\t\t\t\t\t    nr_written,\n+\t\t\t\t\t\t    hash);\n+\t}\n+\n \trename_tmp_packfile(name_buffer, pack_tmp_name, \"pack\");\n \tif (rev_tmp_name)\n \t\trename_tmp_packfile(name_buffer, rev_tmp_name, \"rev\");\n+\tif (mtimes_tmp_name)\n+\t\trename_tmp_packfile(name_buffer, mtimes_tmp_name, \"mtimes\");\n }\n \n void write_promisor_file(const char *promisor_name, struct ref **sought, int nr_sought)\ndiff --git a/pack.h b/pack.h\nindex fd27cfdfd7..01d385903a 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -44,6 +44,7 @@ struct pack_idx_option {\n #define WRITE_IDX_STRICT 02\n #define WRITE_REV 04\n #define WRITE_REV_VERIFY 010\n+#define WRITE_MTIMES 020\n \n \tuint32_t version;\n \tuint32_t off32_limit;\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455445","messageId":"1b241f8f91a5ea22ae9509b90ffdc596a1216a08.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 13/17] builtin/repack.c: allow configuring cruft pack generation","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:25Z","receivedAt":"2022-05-18T23:12:07Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In servers which set the pack.window configuration to a large value, we\ncan wind up spending quite a lot of time finding new bases when breaking\ndelta chains between reachable and unreachable objects while generating\na cruft pack.\n\nIntroduce a handful of `repack.cruft*` configuration variables to\ncontrol the parameters used by pack-objects when generating a cruft\npack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/repack.txt |  9 ++++\n builtin/repack.c                | 49 +++++++++++++------\n t/t5329-pack-objects-cruft.sh   | 83 +++++++++++++++++++++++++++++++++\n 3 files changed, 127 insertions(+), 14 deletions(-)\n\ndiff --git a/Documentation/config/repack.txt b/Documentation/config/repack.txt\nindex 41ac6953c8..c79af6d7b8 100644\n--- a/Documentation/config/repack.txt\n+++ b/Documentation/config/repack.txt\n@@ -30,3 +30,12 @@ repack.updateServerInfo::\n \tIf set to false, linkgit:git-repack[1] will not run\n \tlinkgit:git-update-server-info[1]. Defaults to true. Can be overridden\n \twhen true by the `-n` option of linkgit:git-repack[1].\n+\n+repack.cruftWindow::\n+repack.cruftWindowMemory::\n+repack.cruftDepth::\n+repack.cruftThreads::\n+\tParameters used by linkgit:git-pack-objects[1] when generating\n+\ta cruft pack and the respective parameters are not given over\n+\tthe command line. See similarly named `pack.*` configuration\n+\tvariables for defaults and meaning.\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 593c18d4e8..b85483a148 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -41,9 +41,21 @@ static const char incremental_bitmap_conflict_error[] = N_(\n \"--no-write-bitmap-index or disable the pack.writebitmaps configuration.\"\n );\n \n+struct pack_objects_args {\n+\tconst char *window;\n+\tconst char *window_memory;\n+\tconst char *depth;\n+\tconst char *threads;\n+\tconst char *max_pack_size;\n+\tint no_reuse_delta;\n+\tint no_reuse_object;\n+\tint quiet;\n+\tint local;\n+};\n \n static int repack_config(const char *var, const char *value, void *cb)\n {\n+\tstruct pack_objects_args *cruft_po_args = cb;\n \tif (!strcmp(var, \"repack.usedeltabaseoffset\")) {\n \t\tdelta_base_offset = git_config_bool(var, value);\n \t\treturn 0;\n@@ -65,6 +77,14 @@ static int repack_config(const char *var, const char *value, void *cb)\n \t\trun_update_server_info = git_config_bool(var, value);\n \t\treturn 0;\n \t}\n+\tif (!strcmp(var, \"repack.cruftwindow\"))\n+\t\treturn git_config_string(&cruft_po_args->window, var, value);\n+\tif (!strcmp(var, \"repack.cruftwindowmemory\"))\n+\t\treturn git_config_string(&cruft_po_args->window_memory, var, value);\n+\tif (!strcmp(var, \"repack.cruftdepth\"))\n+\t\treturn git_config_string(&cruft_po_args->depth, var, value);\n+\tif (!strcmp(var, \"repack.cruftthreads\"))\n+\t\treturn git_config_string(&cruft_po_args->threads, var, value);\n \treturn git_default_config(var, value, cb);\n }\n \n@@ -157,18 +177,6 @@ static void remove_redundant_pack(const char *dir_name, const char *base_name)\n \tstrbuf_release(&buf);\n }\n \n-struct pack_objects_args {\n-\tconst char *window;\n-\tconst char *window_memory;\n-\tconst char *depth;\n-\tconst char *threads;\n-\tconst char *max_pack_size;\n-\tint no_reuse_delta;\n-\tint no_reuse_object;\n-\tint quiet;\n-\tint local;\n-};\n-\n static void prepare_pack_objects(struct child_process *cmd,\n \t\t\t\t const struct pack_objects_args *args)\n {\n@@ -692,6 +700,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tint keep_unreachable = 0;\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tstruct pack_objects_args po_args = {NULL};\n+\tstruct pack_objects_args cruft_po_args = {NULL};\n \tint geometric_factor = 0;\n \tint write_midx = 0;\n \n@@ -746,7 +755,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tOPT_END()\n \t};\n \n-\tgit_config(repack_config, NULL);\n+\tgit_config(repack_config, &cruft_po_args);\n \n \targc = parse_options(argc, argv, prefix, builtin_repack_options,\n \t\t\t\tgit_repack_usage, 0);\n@@ -921,7 +930,19 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tif (*pack_prefix == '/')\n \t\t\tpack_prefix++;\n \n-\t\tret = write_cruft_pack(&po_args, pack_prefix, &names,\n+\t\tif (!cruft_po_args.window)\n+\t\t\tcruft_po_args.window = po_args.window;\n+\t\tif (!cruft_po_args.window_memory)\n+\t\t\tcruft_po_args.window_memory = po_args.window_memory;\n+\t\tif (!cruft_po_args.depth)\n+\t\t\tcruft_po_args.depth = po_args.depth;\n+\t\tif (!cruft_po_args.threads)\n+\t\t\tcruft_po_args.threads = po_args.threads;\n+\n+\t\tcruft_po_args.local = po_args.local;\n+\t\tcruft_po_args.quiet = po_args.quiet;\n+\n+\t\tret = write_cruft_pack(&cruft_po_args, pack_prefix, &names,\n \t\t\t\t       &existing_nonkept_packs,\n \t\t\t\t       &existing_kept_packs);\n \t\tif (ret)\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 06c550c958..e4744e4465 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -565,4 +565,87 @@ test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n \t)\n '\n \n+test_expect_success 'cruft repack respects repack.cruftWindow' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tGIT_TRACE2_EVENT=$(pwd)/event.trace \\\n+\t\tgit -c pack.window=1 -c repack.cruftWindow=2 repack \\\n+\t\t       --cruft --window=3 &&\n+\n+\t\tgrep \"pack-objects.*--window=2.*--cruft\" event.trace\n+\t)\n+'\n+\n+test_expect_success 'cruft repack respects --window by default' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tGIT_TRACE2_EVENT=$(pwd)/event.trace \\\n+\t\tgit -c pack.window=2 repack --cruft --window=3 &&\n+\n+\t\tgrep \"pack-objects.*--window=3.*--cruft\" event.trace\n+\t)\n+'\n+\n+test_expect_success 'cruft repack respects --quiet' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\t\tGIT_PROGRESS_DELAY=0 git repack --cruft --quiet 2>err &&\n+\t\ttest_must_be_empty err\n+\t)\n+'\n+\n+test_expect_success 'cruft --local drops unreachable objects' '\n+\tgit init alternate &&\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr alternate repo\" &&\n+\n+\ttest_commit -C alternate base &&\n+\t# Pack all objects in alterate so that the cruft repack in \"repo\" sees\n+\t# the object it dropped due to `--local` as packed. Otherwise this\n+\t# object would not appear packed anywhere (since it is not packed in\n+\t# alternate and likewise not part of the cruft pack in the other repo\n+\t# because of `--local`).\n+\tgit -C alternate repack -ad &&\n+\n+\t(\n+\t\tcd repo &&\n+\n+\t\tobject=\"$(git -C ../alternate rev-parse HEAD:base.t)\" &&\n+\t\tgit -C ../alternate cat-file -p $object >contents &&\n+\n+\t\t# Write some reachable objects and two unreachable ones: one\n+\t\t# that the alternate has and another that is unique.\n+\t\ttest_commit other &&\n+\t\tgit hash-object -w -t blob contents &&\n+\t\tcruft=\"$(echo cruft | git hash-object -w -t blob --stdin)\" &&\n+\n+\t\t( cd ../alternate/.git/objects && pwd ) \\\n+\t\t       >.git/objects/info/alternates &&\n+\n+\t\ttest_path_is_file $objdir/$(test_oid_to_path $cruft) &&\n+\t\ttest_path_is_file $objdir/$(test_oid_to_path $object) &&\n+\n+\t\tgit repack -d --cruft --local &&\n+\n+\t\ttest-tool pack-mtimes \"$(basename $(ls $packdir/pack-*.mtimes))\" \\\n+\t\t       >objects &&\n+\t\t! grep $object objects &&\n+\t\tgrep $cruft objects\n+\t)\n+'\n+\n test_done\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455446","messageId":"ffae78852c54682ed4240f3b2f8cb9c46e2756ec.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 14/17] builtin/repack.c: use named flags for existing_packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:27Z","receivedAt":"2022-05-18T23:12:17Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"We use the `util` pointer for items in the `existing_packs` string list\nto indicate which packs are going to be deleted. Since that has so far\nbeen the only use of that `util` pointer, we just set it to 0 or 1.\n\nBut we're going to add an additional state to this field in the next\npatch, so prepare for that by adding a #define for the first bit so we\ncan more expressively inspect the flags state.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/repack.c | 9 ++++++---\n 1 file changed, 6 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex b85483a148..36d1f03671 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -22,6 +22,8 @@\n #define LOOSEN_UNREACHABLE 2\n #define PACK_CRUFT 4\n \n+#define DELETE_PACK 1\n+\n static int pack_everything;\n static int delta_base_offset = 1;\n static int pack_kept_objects = -1;\n@@ -564,7 +566,7 @@ static void midx_included_packs(struct string_list *include,\n \t\t}\n \t} else {\n \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n-\t\t\tif (item->util)\n+\t\t\tif ((uintptr_t)item->util & DELETE_PACK)\n \t\t\t\tcontinue;\n \t\t\tstring_list_insert(include, xstrfmt(\"%s.idx\", item->string));\n \t\t}\n@@ -1002,7 +1004,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t * was given) and that we will actually delete this pack\n \t\t\t * (if `-d` was given).\n \t\t\t */\n-\t\t\titem->util = (void*)(intptr_t)!string_list_has_string(&names, sha1);\n+\t\t\tif (!string_list_has_string(&names, sha1))\n+\t\t\t\titem->util = (void*)(uintptr_t)((size_t)item->util | DELETE_PACK);\n \t\t}\n \t}\n \n@@ -1026,7 +1029,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (delete_redundant) {\n \t\tint opts = 0;\n \t\tfor_each_string_list_item(item, &existing_nonkept_packs) {\n-\t\t\tif (!item->util)\n+\t\t\tif (!((uintptr_t)item->util & DELETE_PACK))\n \t\t\t\tcontinue;\n \t\t\tremove_redundant_pack(packdir, item->string);\n \t\t}\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455447","messageId":"0743e373baaaaee2bc5f8664b1e64038e6dbf4c2.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 15/17] builtin/repack.c: add cruft packs to MIDX during geometric repack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:30Z","receivedAt":"2022-05-18T23:12:27Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When using cruft packs, the following race can occur when a geometric\nrepack that writes a MIDX bitmap takes place afterwords:\n\n  - First, create an unreachable object and do an all-into-one cruft\n    repack which stores that object in the repository's cruft pack.\n  - Then make that object reachable.\n  - Finally, do a geometric repack and write a MIDX bitmap.\n\nAssuming that we are sufficiently unlucky as to select a commit from the\nMIDX which reaches that object for bitmapping, then the `git\nmulti-pack-index` process will complain that that object is missing.\n\nThe reason is because we don't include cruft packs in the MIDX when\ndoing a geometric repack. Since the \"make that object reachable\" doesn't\nnecessarily mean that we'll create a new copy of that object in one of\nthe packs that will get rolled up as part of a geometric repack, it's\npossible that the MIDX won't see any copies of that now-reachable\nobject.\n\nOf course, it's desirable to avoid including cruft packs in the MIDX\nbecause it causes the MIDX to store a bunch of objects which are likely\nto get thrown away. But excluding that pack does open us up to the above\nrace.\n\nThis patch demonstrates the bug, and resolves it by including cruft\npacks in the MIDX even when doing a geometric repack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/repack.c              | 19 +++++++++++++++++--\n t/t5329-pack-objects-cruft.sh | 26 ++++++++++++++++++++++++++\n 2 files changed, 43 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 36d1f03671..e9e3a2b4e3 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -23,6 +23,7 @@\n #define PACK_CRUFT 4\n \n #define DELETE_PACK 1\n+#define CRUFT_PACK 2\n \n static int pack_everything;\n static int delta_base_offset = 1;\n@@ -161,8 +162,11 @@ static void collect_pack_filenames(struct string_list *fname_nonkept_list,\n \t\tif ((extra_keep->nr > 0 && i < extra_keep->nr) ||\n \t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname))))\n \t\t\tstring_list_append_nodup(fname_kept_list, fname);\n-\t\telse\n-\t\t\tstring_list_append_nodup(fname_nonkept_list, fname);\n+\t\telse {\n+\t\t\tstruct string_list_item *item = string_list_append_nodup(fname_nonkept_list, fname);\n+\t\t\tif (file_exists(mkpath(\"%s/%s.mtimes\", packdir, fname)))\n+\t\t\t\titem->util = (void*)(uintptr_t)CRUFT_PACK;\n+\t\t}\n \t}\n \tclosedir(dir);\n }\n@@ -564,6 +568,17 @@ static void midx_included_packs(struct string_list *include,\n \n \t\t\tstring_list_insert(include, strbuf_detach(&buf, NULL));\n \t\t}\n+\n+\t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n+\t\t\tif (!((uintptr_t)item->util & CRUFT_PACK)) {\n+\t\t\t\t/*\n+\t\t\t\t * no need to check DELETE_PACK, since we're not\n+\t\t\t\t * doing an ALL_INTO_ONE repack\n+\t\t\t\t */\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tstring_list_insert(include, xstrfmt(\"%s.idx\", item->string));\n+\t\t}\n \t} else {\n \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n \t\t\tif ((uintptr_t)item->util & DELETE_PACK)\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex e4744e4465..13158e4ab7 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -648,4 +648,30 @@ test_expect_success 'cruft --local drops unreachable objects' '\n \t)\n '\n \n+test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\ttest_commit cruft &&\n+\t\tunreachable=\"$(git rev-parse cruft)\" &&\n+\n+\t\tgit reset --hard $unreachable^ &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\t# resurrect the unreachable object via a new commit. the\n+\t\t# new commit will get selected for a bitmap, but be\n+\t\t# missing one of its parents from the selected packs.\n+\t\tgit reset --hard $unreachable &&\n+\t\ttest_commit resurrect &&\n+\n+\t\tgit repack --write-midx --write-bitmap-index --geometric=2 -d\n+\t)\n+'\n+\n test_done\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455448","messageId":"9f7e0acac63a9875895aec7d3ef819323255ec74.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 16/17] builtin/gc.c: conditionally avoid pruning objects via loose","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:32Z","receivedAt":"2022-05-18T23:12:43Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Expose the new `git repack --cruft` mode from `git gc` via a new opt-in\nflag. When invoked like `git gc --cruft`, `git gc` will avoid exploding\nunreachable objects as loose ones, and instead create a cruft pack and\n`.mtimes` file.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/gc.txt   | 21 +++++++++++++-------\n Documentation/git-gc.txt      |  5 +++++\n builtin/gc.c                  | 10 +++++++++-\n t/t5329-pack-objects-cruft.sh | 37 +++++++++++++++++++++++++++++++++++\n 4 files changed, 65 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/config/gc.txt b/Documentation/config/gc.txt\nindex c834e07991..38fea076a2 100644\n--- a/Documentation/config/gc.txt\n+++ b/Documentation/config/gc.txt\n@@ -81,14 +81,21 @@ gc.packRefs::\n \tto enable it within all non-bare repos or it can be set to a\n \tboolean value.  The default is `true`.\n \n+gc.cruftPacks::\n+\tStore unreachable objects in a cruft pack (see\n+\tlinkgit:git-repack[1]) instead of as loose objects. The default\n+\tis `false`.\n+\n gc.pruneExpire::\n-\tWhen 'git gc' is run, it will call 'prune --expire 2.weeks.ago'.\n-\tOverride the grace period with this config variable.  The value\n-\t\"now\" may be used to disable this grace period and always prune\n-\tunreachable objects immediately, or \"never\" may be used to\n-\tsuppress pruning.  This feature helps prevent corruption when\n-\t'git gc' runs concurrently with another process writing to the\n-\trepository; see the \"NOTES\" section of linkgit:git-gc[1].\n+\tWhen 'git gc' is run, it will call 'prune --expire 2.weeks.ago'\n+\t(and 'repack --cruft --cruft-expiration 2.weeks.ago' if using\n+\tcruft packs via `gc.cruftPacks` or `--cruft`).  Override the\n+\tgrace period with this config variable.  The value \"now\" may be\n+\tused to disable this grace period and always prune unreachable\n+\tobjects immediately, or \"never\" may be used to suppress pruning.\n+\tThis feature helps prevent corruption when 'git gc' runs\n+\tconcurrently with another process writing to the repository; see\n+\tthe \"NOTES\" section of linkgit:git-gc[1].\n \n gc.worktreePruneExpire::\n \tWhen 'git gc' is run, it calls\ndiff --git a/Documentation/git-gc.txt b/Documentation/git-gc.txt\nindex 853967dea0..ba4e67700e 100644\n--- a/Documentation/git-gc.txt\n+++ b/Documentation/git-gc.txt\n@@ -54,6 +54,11 @@ other housekeeping tasks (e.g. rerere, working trees, reflog...) will\n be performed as well.\n \n \n+--cruft::\n+\tWhen expiring unreachable objects, pack them separately into a\n+\tcruft pack instead of storing the loose objects as loose\n+\tobjects.\n+\n --prune=<date>::\n \tPrune loose objects older than date (default is 2 weeks ago,\n \toverridable by the config variable `gc.pruneExpire`).\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex b335cffa33..4d995e85e9 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -42,6 +42,7 @@ static const char * const builtin_gc_usage[] = {\n \n static int pack_refs = 1;\n static int prune_reflogs = 1;\n+static int cruft_packs = 0;\n static int aggressive_depth = 50;\n static int aggressive_window = 250;\n static int gc_auto_threshold = 6700;\n@@ -152,6 +153,7 @@ static void gc_config(void)\n \tgit_config_get_int(\"gc.auto\", &gc_auto_threshold);\n \tgit_config_get_int(\"gc.autopacklimit\", &gc_auto_pack_limit);\n \tgit_config_get_bool(\"gc.autodetach\", &detach_auto);\n+\tgit_config_get_bool(\"gc.cruftpacks\", &cruft_packs);\n \tgit_config_get_expiry(\"gc.pruneexpire\", &prune_expire);\n \tgit_config_get_expiry(\"gc.worktreepruneexpire\", &prune_worktrees_expire);\n \tgit_config_get_expiry(\"gc.logexpiry\", &gc_log_expire);\n@@ -331,7 +333,11 @@ static void add_repack_all_option(struct string_list *keep_pack)\n {\n \tif (prune_expire && !strcmp(prune_expire, \"now\"))\n \t\tstrvec_push(&repack, \"-a\");\n-\telse {\n+\telse if (cruft_packs) {\n+\t\tstrvec_push(&repack, \"--cruft\");\n+\t\tif (prune_expire)\n+\t\t\tstrvec_pushf(&repack, \"--cruft-expiration=%s\", prune_expire);\n+\t} else {\n \t\tstrvec_push(&repack, \"-A\");\n \t\tif (prune_expire)\n \t\t\tstrvec_pushf(&repack, \"--unpack-unreachable=%s\", prune_expire);\n@@ -551,6 +557,7 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t{ OPTION_STRING, 0, \"prune\", &prune_expire, N_(\"date\"),\n \t\t\tN_(\"prune unreferenced objects\"),\n \t\t\tPARSE_OPT_OPTARG, NULL, (intptr_t)prune_expire },\n+\t\tOPT_BOOL(0, \"cruft\", &cruft_packs, N_(\"pack unreferenced objects separately\")),\n \t\tOPT_BOOL(0, \"aggressive\", &aggressive, N_(\"be more thorough (increased runtime)\")),\n \t\tOPT_BOOL_F(0, \"auto\", &auto_gc, N_(\"enable auto-gc mode\"),\n \t\t\t   PARSE_OPT_NOCOMPLETE),\n@@ -670,6 +677,7 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t\tdie(FAILED_RUN, repack.v[0]);\n \n \t\tif (prune_expire) {\n+\t\t\t/* run `git prune` even if using cruft packs */\n \t\t\tstrvec_push(&prune, prune_expire);\n \t\t\tif (quiet)\n \t\t\t\tstrvec_push(&prune, \"--no-progress\");\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 13158e4ab7..3910e186ef 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -429,6 +429,43 @@ test_expect_success 'loose objects mtimes upsert others' '\n \t)\n '\n \n+test_expect_success 'expiring cruft objects with git gc' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit branch -M main &&\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit unreachable &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d unreachable &&\n+\t\t# objects are not cruft if they are contained in the reflogs\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\t\tsort <reachable.raw >reachable &&\n+\t\tcomm -13 reachable objects >unreachable &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tmtimes=$(ls .git/objects/pack/pack-*.mtimes) &&\n+\t\ttest_path_is_file $mtimes &&\n+\n+\t\tgit gc --cruft --prune=now &&\n+\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\n+\t\tcomm -23 unreachable objects >removed &&\n+\t\ttest_cmp unreachable removed &&\n+\t\ttest_path_is_missing $mtimes\n+\t)\n+'\n+\n test_expect_success 'cruft packs are not included in geometric repack' '\n \tgit init repo &&\n \ttest_when_finished \"rm -fr repo\" &&\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455449","messageId":"07fa9d4b475d37189e978a42a3939b0a20834af3.1652915424.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[PATCH v4 17/17] sha1-file.c: don't freshen cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-18T23:11:35Z","receivedAt":"2022-05-18T23:12:47Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"We don't bother to freshen objects stored in a cruft pack individually\nby updating the `.mtimes` file. This is because we can't portably `mmap`\nand write into the middle of a file (i.e., to update the mtime of just\none object). Instead, we would have to rewrite the entire `.mtimes` file\nwhich may incur some wasted effort especially if there a lot of cruft\nobjects and they are freshened infrequently.\n\nInstead, force the freshening code to avoid an optimizing write by\nwriting out the object loose and letting it pick up a current mtime.\n\nThis works because we prefer the mtime of the loose copy of an object\nwhen both a loose and packed one exist (whether or not the packed copy\ncomes from a cruft pack or not).\n\nThis could certainly do with a test and/or be included earlier in this\nseries/PR, but I want to wait until after I have a chance to clean up\nthe overly-repetitive nature of the cruft pack tests in general.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n object-file.c                 |  2 ++\n t/t5329-pack-objects-cruft.sh | 25 +++++++++++++++++++++++++\n 2 files changed, 27 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex ff0cffe68e..495a359200 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2035,6 +2035,8 @@ static int freshen_packed_object(const struct object_id *oid)\n \tstruct pack_entry e;\n \tif (!find_pack_entry(the_repository, oid, &e))\n \t\treturn 0;\n+\tif (e.p->is_cruft)\n+\t\treturn 0;\n \tif (e.p->freshened)\n \t\treturn 1;\n \tif (!freshen_file(e.p->pack_name))\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 3910e186ef..4681558612 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -711,4 +711,29 @@ test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n \t)\n '\n \n+test_expect_success 'cruft objects are freshend via loose' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\techo \"cruft\" >contents &&\n+\t\tblob=\"$(git hash-object -w -t blob contents)\" &&\n+\t\tloose=\"$objdir/$(test_oid_to_path $blob)\" &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\ttest_path_is_missing \"$loose\" &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(ls $packdir/pack-*.mtimes)\")\" >cruft &&\n+\t\tgrep \"$blob\" cruft &&\n+\n+\t\t# write the same object again\n+\t\tgit hash-object -w -t blob contents &&\n+\n+\t\ttest_path_is_file \"$loose\"\n+\t)\n+'\n+\n test_done\n-- \n2.36.1.94.gb0d54bedca\n"},{"id":"455450","messageId":"98d9bbe5-1902-0dc4-e41e-33020d0396ad@github.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"Re: [PATCH v4 00/17] cruft packs","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-05-18T23:48:37Z","receivedAt":"2022-05-18T23:48:47Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 5/18/2022 7:10 PM, Taylor Blau wrote:\n> Here is another reroll of my series to implement \"cruft packs\", which is based\n> on the v2.36 tree, and incorporates feedback from the discussion we had about\n> mixed-version GCs with cruft packs in [1].\n> \n> The changes here are limited to:\n> \n>   - a cautionary note in Documentation/technical/cruft-packs.txt\n>     describing the potential interaction between pruning GCs across pre-\n>     and post-cruft pack versions of Git, as discussed towards the bottom\n>     of [2]\n\nI think this documentation is sufficient guarding against this issue,\nwhich is not so critical as to do something more involved. When users\nopt-in to using cruft packs, they should know about their scenario\nenough to know if they would stumble into this issue.\n\n>   - updating the `finalize_hashfile()` calls for writing `.mtimes` files\n>     to indicate that they are `FSYNC_COMPONENT_PACK_METADATA`, since the\n>     original version of this series predates the fine-grained fsync\n>     configuration in 2.36.\n\nGood to have this update and not require it to be handled at merge\ntime by the maintainer.\n\n> As always, a range-diff is below. Thanks in advance for taking another\n> look!\n\nLooking at the range-diff, I'm happy with this version.\n\nThanks,\n-Stolee\n"},{"id":"455465","messageId":"xmqq7d6hsusd.fsf@gitster.g","threadId":"56996","inReplyTo":"94fe03cc65716b6102e2d71df49d4ae5a1a60dc7.1652915424.git.me@ttaylorr.com","subject":"Re: [PATCH v4 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-19T10:04:18Z","receivedAt":"2022-05-19T10:04:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> @@ -3870,6 +4034,20 @@ static int option_parse_unpack_unreachable(const struct option *opt,\n>  \treturn 0;\n>  }\n>  \n> +static int option_parse_cruft_expiration(const struct option *opt,\n> +\t\t\t\t\t const char *arg, int unset)\n> +{\n> +\tif (unset) {\n> +\t\tcruft = 0;\n> +\t\tcruft_expiration = 0;\n> +\t} else {\n> +\t\tcruft = 1;\n> +\t\tif (arg)\n> +\t\t\tcruft_expiration = approxidate(arg);\n> +\t}\n> +\treturn 0;\n> +}\n\nIt is somewhat sad that we have to invent this function, instead of\nusing parse_opt_expiry_date_cb().\n"},{"id":"455470","messageId":"220519.86zgjd4wvk.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"8f9fd21be9fcdda5c73d800fc66d1087d61a6888.1652915424.git.me@ttaylorr.com","subject":"Re: [PATCH v4 02/17] pack-mtimes: support reading .mtimes files","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-19T10:40:56Z","receivedAt":"2022-05-19T10:53:14Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 18 2022, Taylor Blau wrote:\n\nNit:\n\n> +  - A 4-byte magic number '0x4d544d45' ('MTME').\n> +\n> +  - A 4-byte version identifier (= 1).\n> +\n> +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n\nHere we let it suffice that later we'll say \"All 4-byte numbers are in\nnetwork order\".\n\n> +  - A table of 4-byte unsigned integers in network order. The ith\n\nBut here we call out \"network order\" explicitly, shouldn't this just be\ns/ in network order//?\n\n> +    value is the modification time (mtime) of the ith object in the\n> +    corresponding pack by lexicographic (index) order. The mtimes\n> +    count standard epoch seconds.\n> +\n> +  - A trailer, containing a checksum of the corresponding packfile,\n> +    and a checksum of all of the above (each having length according\n> +    to the specified hash function).\n> +\n> +All 4-byte numbers are in network order.\n\nI.e. this is sufficient.\n"},{"id":"455472","messageId":"220519.86v8u14v52.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"5992a72cbf9e8d076f1e312a789b40d52656ad3c.1652915424.git.me@ttaylorr.com","subject":"Re: [PATCH v4 12/17] builtin/repack.c: support generating a cruft pack","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-19T11:29:26Z","receivedAt":"2022-05-19T11:31:12Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 18 2022, Taylor Blau wrote:\n\n> +\t\ttip=\"$(git rev-parse cruft)\" &&\n\nHere we don't hide the exit status of \"git\", as it'll be reflected in what's &&-chained.\n\n> +\t\tpath=\"$objdir/$(test_oid_to_path \"$(git rev-parse cruft)\")\" &&\n\nBut here we do, as we'll get the exit status of test_oid_to_path. But as\nwe just rev parsed it shouldn't this be $tip in any case?\n"},{"id":"455473","messageId":"220519.86r14p4v0q.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"0743e373baaaaee2bc5f8664b1e64038e6dbf4c2.1652915424.git.me@ttaylorr.com","subject":"Re: [PATCH v4 15/17] builtin/repack.c: add cruft packs to MIDX during geometric repack","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-19T11:32:26Z","receivedAt":"2022-05-19T11:33:15Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 18 2022, Taylor Blau wrote:\n\n> When using cruft packs, the following race can occur when a geometric\n> repack that writes a MIDX bitmap takes place afterwords:\n>\n>   - First, create an unreachable object and do an all-into-one cruft\n>     repack which stores that object in the repository's cruft pack.\n>   - Then make that object reachable.\n>   - Finally, do a geometric repack and write a MIDX bitmap.\n>\n> Assuming that we are sufficiently unlucky as to select a commit from the\n> MIDX which reaches that object for bitmapping, then the `git\n> multi-pack-index` process will complain that that object is missing.\n>\n> The reason is because we don't include cruft packs in the MIDX when\n> doing a geometric repack. Since the \"make that object reachable\" doesn't\n> necessarily mean that we'll create a new copy of that object in one of\n> the packs that will get rolled up as part of a geometric repack, it's\n> possible that the MIDX won't see any copies of that now-reachable\n> object.\n>\n> Of course, it's desirable to avoid including cruft packs in the MIDX\n> because it causes the MIDX to store a bunch of objects which are likely\n> to get thrown away. But excluding that pack does open us up to the above\n> race.\n>\n> This patch demonstrates the bug, and resolves it by including cruft\n> packs in the MIDX even when doing a geometric repack.\n>\n> Signed-off-by: Taylor Blau <me@ttaylorr.com>\n> ---\n>  builtin/repack.c              | 19 +++++++++++++++++--\n>  t/t5329-pack-objects-cruft.sh | 26 ++++++++++++++++++++++++++\n>  2 files changed, 43 insertions(+), 2 deletions(-)\n>\n> diff --git a/builtin/repack.c b/builtin/repack.c\n> index 36d1f03671..e9e3a2b4e3 100644\n> --- a/builtin/repack.c\n> +++ b/builtin/repack.c\n> @@ -23,6 +23,7 @@\n>  #define PACK_CRUFT 4\n>  \n>  #define DELETE_PACK 1\n> +#define CRUFT_PACK 2\n>  \n>  static int pack_everything;\n>  static int delta_base_offset = 1;\n> @@ -161,8 +162,11 @@ static void collect_pack_filenames(struct string_list *fname_nonkept_list,\n>  \t\tif ((extra_keep->nr > 0 && i < extra_keep->nr) ||\n>  \t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname))))\n>  \t\t\tstring_list_append_nodup(fname_kept_list, fname);\n> -\t\telse\n> -\t\t\tstring_list_append_nodup(fname_nonkept_list, fname);\n> +\t\telse {\n> +\t\t\tstruct string_list_item *item = string_list_append_nodup(fname_nonkept_list, fname);\n\nNit: very long line, and we end up with {} just on the else, not the if.\n"},{"id":"455474","messageId":"RFC-cover-0.2-00000000000-20220519T113538Z-avarab@gmail.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"[RFC PATCH 0/2] Utility functions for duplicated pack(write) code","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-19T11:42:26Z","receivedAt":"2022-05-19T11:42:44Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Minor cleanups thath would semantically & textually conflict with\nTaylor's\nhttps://lore.kernel.org/git/cover.1652915424.git.me@ttaylorr.com/; but\nwhich I noted while reading through it.\n\nThe 2/2 here is something I wrote before spotting\nhttps://lore.kernel.org/git/1d775f9850f00b0c3d1e9133669a6365c8d7bbba.1652915424.git.me@ttaylorr.com/;\nwhich does pretty much the same thing. but IMO it's better to put this\nin hash.h than chunk-format.h.\n\nThe 1/2 then fixes the minor NEEDSWORK in that series:\nhttps://lore.kernel.org/git/8f9fd21be9fcdda5c73d800fc66d1087d61a6888.1652915424.git.me@ttaylorr.com/\n\nAll of this can be ignored for now, I can submit it after cruft packs\nland (if I remember), or if Taylor's interested in picking it up in\nsome way...\n\nBut I figured it was useful to send it along in liue of \"maybe do it\nthis way\" (2/2) or \"can we just create a utility function for this?\"\n(1/2) comments on the series itself.\n\nÆvar Arnfjörð Bjarmason (2):\n  packfile API: add and use a pack_name_to_ext() utility function\n  hash API: add and use a hash_short_id_by_algo() function\n\n commit-graph.c  | 18 +++---------------\n hash.h          | 26 ++++++++++++++++++++++++--\n midx.c          | 18 +++---------------\n pack-bitmap.c   |  6 +-----\n pack-revindex.c |  5 +----\n pack-write.c    | 12 +-----------\n packfile.c      | 14 ++++++++++----\n packfile.h      |  9 +++++++++\n 8 files changed, 52 insertions(+), 56 deletions(-)\n\n-- \n2.36.1.952.g6652f7f0e6b\n\n"},{"id":"455475","messageId":"RFC-patch-1.2-d8c3c03e90f-20220519T113538Z-avarab@gmail.com","threadId":"56996","inReplyTo":"RFC-cover-0.2-00000000000-20220519T113538Z-avarab@gmail.com","subject":"[RFC PATCH 1/2] packfile API: add and use a pack_name_to_ext() utility function","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-19T11:42:27Z","receivedAt":"2022-05-19T11:42:47Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Add and use a pack_name_to_ext() utility function for the copy/pasted\ncases of creating a FOO.ext file given a string like FOO.pack.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n pack-bitmap.c   |  6 +-----\n pack-revindex.c |  5 +----\n packfile.c      | 14 ++++++++++----\n packfile.h      |  9 +++++++++\n 4 files changed, 21 insertions(+), 13 deletions(-)\n\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nindex 97909d48da3..0c3770d038d 100644\n--- a/pack-bitmap.c\n+++ b/pack-bitmap.c\n@@ -302,11 +302,7 @@ char *midx_bitmap_filename(struct multi_pack_index *midx)\n \n char *pack_bitmap_filename(struct packed_git *p)\n {\n-\tsize_t len;\n-\n-\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n-\t\tBUG(\"pack_name does not end in .pack\");\n-\treturn xstrfmt(\"%.*s.bitmap\", (int)len, p->pack_name);\n+\treturn pack_name_to_ext(p->pack_name, \"bitmap\");\n }\n \n static int open_midx_bitmap_1(struct bitmap_index *bitmap_git,\ndiff --git a/pack-revindex.c b/pack-revindex.c\nindex 08dc1601679..69dc5688796 100644\n--- a/pack-revindex.c\n+++ b/pack-revindex.c\n@@ -179,10 +179,7 @@ static int create_pack_revindex_in_memory(struct packed_git *p)\n \n static char *pack_revindex_filename(struct packed_git *p)\n {\n-\tsize_t len;\n-\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n-\t\tBUG(\"pack_name does not end in .pack\");\n-\treturn xstrfmt(\"%.*s.rev\", (int)len, p->pack_name);\n+\treturn pack_name_to_ext(p->pack_name, \"rev\");\n }\n \n #define RIDX_HEADER_SIZE (12)\ndiff --git a/packfile.c b/packfile.c\nindex 835b2d27164..bd6ad441bf5 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -191,15 +191,12 @@ int load_idx(const char *path, const unsigned int hashsz, void *idx_map,\n int open_pack_index(struct packed_git *p)\n {\n \tchar *idx_name;\n-\tsize_t len;\n \tint ret;\n \n \tif (p->index_data)\n \t\treturn 0;\n \n-\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n-\t\tBUG(\"pack_name does not end in .pack\");\n-\tidx_name = xstrfmt(\"%.*s.idx\", (int)len, p->pack_name);\n+\tidx_name = pack_name_to_ext(p->pack_name, \"idx\");\n \tret = check_packed_git_idx(idx_name, p);\n \tfree(idx_name);\n \treturn ret;\n@@ -2266,3 +2263,12 @@ int is_promisor_object(const struct object_id *oid)\n \t}\n \treturn oidset_contains(&promisor_objects, oid);\n }\n+\n+char *pack_name_to_ext(const char *pack_name, const char *ext)\n+{\n+\tsize_t len;\n+\n+\tif (!strip_suffix(pack_name, \".pack\", &len))\n+\t\tBUG(\"pack_name does not end in .pack\");\n+\treturn xstrfmt(\"%.*s.%s\", (int)len, pack_name, ext);\n+}\ndiff --git a/packfile.h b/packfile.h\nindex a3f6723857b..6890c57ebdb 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -195,4 +195,13 @@ int is_promisor_object(const struct object_id *oid);\n int load_idx(const char *path, const unsigned int hashsz, void *idx_map,\n \t     size_t idx_size, struct packed_git *p);\n \n+/**\n+ * Given a string like \"foo.pack\" and \"ext\" returns an xstrdup()'d\n+ * \"foo.ext\" string. Used for creating e.g. PACK.{bitmap,rev,...}\n+ * filenames from PACK.pack.\n+ *\n+ * Will BUG() if the expected string can't be created from the\n+ * \"pack_name\" argument.\n+ */\n+char *pack_name_to_ext(const char *pack_name, const char *ext);\n #endif\n-- \n2.36.1.952.g6652f7f0e6b\n\n"},{"id":"455476","messageId":"RFC-patch-2.2-051f0612ab9-20220519T113538Z-avarab@gmail.com","threadId":"56996","inReplyTo":"RFC-cover-0.2-00000000000-20220519T113538Z-avarab@gmail.com","subject":"[RFC PATCH 2/2] hash API: add and use a hash_short_id_by_algo() function","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-19T11:42:28Z","receivedAt":"2022-05-19T11:42:52Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Add and use a hash_short_id_by_algo() function. As noted in the\ncomment being modified here (added in [1]) the intention wasn't to\nhave these end up in on-disk formats, but since [2], [3] and [3]\nthat's been the case, and there's an outstanding patch to add another\nformat that uses these[5].\n\nSo let's expose this functionality as a documented utility function,\ninstead of copy/pasting this code in various places.\n\nReplacing the die() in the existing functions with a BUG() might be\noverzelous, it's correct for the case of\ne.g. write_commit_graph_file() and write_midx_header(), but we also\nuse this for parsing on-disk files, e.g. in parse_commit_graph().\n\nWe could add a \"gently\" version of this, but for now I think that\nworrying about the distinction would be worrying too much. If we ever\nend up parsing such files that'll almost certainly be a bug in our own\nwriting code, so the distinction would be rather academic, even though\nsuch files could theoretically occur without a bug of ours.\n\n1. f50e766b7b3 (Add structure representing hash algorithm, 2017-11-12)\n2. 665d70ad033 (commit-graph: use the \"hash version\" byte, 2020-08-17)\n3. d96075428a9 (multi-pack-index: use hash version byte, 2020-08-17)\n4. 8ef50d9958f (pack-write.c: prepare to write 'pack-*.rev' files, 2021-01-25)\n5. https://lore.kernel.org/git/1d775f9850f00b0c3d1e9133669a6365c8d7bbba.1652915424.git.me@ttaylorr.com/\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n commit-graph.c | 18 +++---------------\n hash.h         | 26 ++++++++++++++++++++++++--\n midx.c         | 18 +++---------------\n pack-write.c   | 12 +-----------\n 4 files changed, 31 insertions(+), 43 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 06107beedcb..157de4dd717 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -193,18 +193,6 @@ char *get_commit_graph_chain_filename(struct object_directory *odb)\n \treturn xstrfmt(\"%s/info/commit-graphs/commit-graph-chain\", odb->path);\n }\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n static struct commit_graph *alloc_commit_graph(void)\n {\n \tstruct commit_graph *g = xcalloc(1, sizeof(*g));\n@@ -365,9 +353,9 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n \t}\n \n \thash_version = *(unsigned char*)(data + 5);\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != hash_short_id_by_algo()) {\n \t\terror(_(\"commit-graph hash version %X does not match version %X\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, hash_short_id_by_algo());\n \t\treturn NULL;\n \t}\n \n@@ -1924,7 +1912,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \thashwrite_be32(f, GRAPH_SIGNATURE);\n \n \thashwrite_u8(f, GRAPH_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, hash_short_id_by_algo());\n \thashwrite_u8(f, get_num_chunks(cf));\n \thashwrite_u8(f, ctx->num_commit_graphs_after - 1);\n \ndiff --git a/hash.h b/hash.h\nindex 5d40368f18a..31293401809 100644\n--- a/hash.h\n+++ b/hash.h\n@@ -80,8 +80,7 @@ static inline void git_SHA256_Clone(git_SHA256_CTX *dst, const git_SHA256_CTX *s\n \n /*\n  * Note that these constants are suitable for indexing the hash_algos array and\n- * comparing against each other, but are otherwise arbitrary, so they should not\n- * be exposed to the user or serialized to disk.  To know whether a\n+ * comparing against each other, but are otherwise arbitrary. To know whether a\n  * git_hash_algo struct points to some usable hash function, test the format_id\n  * field for being non-zero.  Use the name field for user-visible situations and\n  * the format_id field for fixed-length fields on disk.\n@@ -337,4 +336,27 @@ static inline void oid_set_algo(struct object_id *oid, const struct git_hash_alg\n const char *empty_tree_oid_hex(void);\n const char *empty_blob_oid_hex(void);\n \n+/**\n+ * Convert GIT_HASH_SHA1 to 1, GIT_HASH_SHA256 to 2 etc.\n+ *\n+ * It's preferable to use GIT_{SHA1,SHA256}_FORMAT_ID instead for file\n+ * formats. The original intention was not to make these short\n+ * constants part of any file format.\n+ * \n+ * But since that ship has sailed for various on-disk formats this\n+ * utility function allows us to do that consistently in one place.\n+ */\n+static inline int hash_short_id_by_algo(void)\n+{\n+\tint hash_algo = hash_algo_by_ptr(the_hash_algo);\n+\n+\tswitch (hash_algo) {\n+\tcase GIT_HASH_SHA1:\n+\tcase GIT_HASH_SHA256:\n+\t\treturn hash_algo;\n+\tdefault:\n+\t\tBUG(\"invalid hash version\");\n+\t}\n+}\n+\n #endif\ndiff --git a/midx.c b/midx.c\nindex 3db0e47735f..2e42afa5f00 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -41,18 +41,6 @@\n \n #define PACK_EXPIRED UINT_MAX\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n const unsigned char *get_midx_checksum(struct multi_pack_index *m)\n {\n \treturn m->data + m->data_len - the_hash_algo->rawsz;\n@@ -134,9 +122,9 @@ struct multi_pack_index *load_multi_pack_index(const char *object_dir, int local\n \t\t      m->version);\n \n \thash_version = m->data[MIDX_BYTE_HASH_VERSION];\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != hash_short_id_by_algo()) {\n \t\terror(_(\"multi-pack-index hash version %u does not match version %u\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, hash_short_id_by_algo());\n \t\tgoto cleanup_fail;\n \t}\n \tm->hash_len = the_hash_algo->rawsz;\n@@ -420,7 +408,7 @@ static size_t write_midx_header(struct hashfile *f,\n {\n \thashwrite_be32(f, MIDX_SIGNATURE);\n \thashwrite_u8(f, MIDX_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, hash_short_id_by_algo());\n \thashwrite_u8(f, num_chunks);\n \thashwrite_u8(f, 0); /* unused */\n \thashwrite_be32(f, num_packs);\ndiff --git a/pack-write.c b/pack-write.c\nindex 51812cb1299..c1ce8f6df8f 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -181,17 +181,7 @@ static int pack_order_cmp(const void *va, const void *vb, void *ctx)\n \n static void write_rev_header(struct hashfile *f)\n {\n-\tuint32_t oid_version;\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\toid_version = 1;\n-\t\tbreak;\n-\tcase GIT_HASH_SHA256:\n-\t\toid_version = 2;\n-\t\tbreak;\n-\tdefault:\n-\t\tdie(\"write_rev_header: unknown hash version\");\n-\t}\n+\tuint32_t oid_version = hash_short_id_by_algo();\n \n \thashwrite_be32(f, RIDX_SIGNATURE);\n \thashwrite_be32(f, RIDX_VERSION);\n-- \n2.36.1.952.g6652f7f0e6b\n\n"},{"id":"455477","messageId":"220519.86mtfd4uc0.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"1d775f9850f00b0c3d1e9133669a6365c8d7bbba.1652915424.git.me@ttaylorr.com","subject":"Re: [PATCH v4 04/17] chunk-format.h: extract oid_version()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-19T11:44:14Z","receivedAt":"2022-05-19T11:48:05Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 18 2022, Taylor Blau wrote:\n\n> There are three definitions of an identical function which converts\n> `the_hash_algo` into either 1 (for SHA-1) or 2 (for SHA-256). There is a\n> copy of this function for writing both the commit-graph and\n> multi-pack-index file, and another inline definition used to write the\n> .rev header.\n>\n> Consolidate these into a single definition in chunk-format.h. It's not\n> clear that this is the best header to define this function in, but it\n> should do for now.\n\nMaybe hash.h? :)\nhttps://lore.kernel.org/git/RFC-patch-2.2-051f0612ab9-20220519T113538Z-avarab@gmail.com/\n\n> (Worth noting, the .rev caller expects a 4-byte unsigned, but the other\n> two callers work with a single unsigned byte. The consolidated version\n> uses the latter type, and lets the compiler widen it when required).\n\nI just went for \"int\" and had the compiler similarly cast that, which\nseems simpler & more obvious, no?\n\nI.e. it seems to me that we really only need these more narrow types at\nthe time that we write this data, which we alredy have casts for.\n\n> +uint8_t oid_version(const struct git_hash_algo *algop)\n> +{\n> +\tswitch (hash_algo_by_ptr(algop)) {\n> +\tcase GIT_HASH_SHA1:\n> +\t\treturn 1;\n> +\tcase GIT_HASH_SHA256:\n> +\t\treturn 2;\n> +\tdefault:\n> +\t\tdie(_(\"invalid hash version\"));\n> +\t}\n\nAs noted in the 2/2 I posted above we have some cases where we really\nshould have BUG here, and others (reading) which are arguably die(). I\nthink just going for BUG() makes sense in this case.\n\nBut if you're just unifying existing code we can also just keep it\nas-is.\n\nFWIW I struggled to come up with a name for this, and ended up with\nhash_short_id_by_algo(). Somewhat bikesheddy, but I'd prefer if we fixed\nthat \"oid_version\" name while at it, since this really has nothing do do\nwith an \"OID version\" (whatever that is).\n\nWe only refer to hash versions elsewhere, which collectively describe\nthe versions of all OIDs we need to handle.\n"},{"id":"455478","messageId":"220519.86ee0p4ru6.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"cover.1652915424.git.me@ttaylorr.com","subject":"Re: [PATCH v4 00/17] cruft packs","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-19T11:54:41Z","receivedAt":"2022-05-19T12:42:01Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 18 2022, Taylor Blau wrote:\n\n> Here is another reroll of my series to implement \"cruft packs\", which is based\n> on the v2.36 tree, and incorporates feedback from the discussion we had about\n> mixed-version GCs with cruft packs in [1].\n>\n> The changes here are limited to:\n>\n>   - a cautionary note in Documentation/technical/cruft-packs.txt\n>     describing the potential interaction between pruning GCs across pre-\n>     and post-cruft pack versions of Git, as discussed towards the bottom\n>     of [2]\n>\n>   - updating the `finalize_hashfile()` calls for writing `.mtimes` files\n>     to indicate that they are `FSYNC_COMPONENT_PACK_METADATA`, since the\n>     original version of this series predates the fine-grained fsync\n>     configuration in 2.36.\n>\n> As always, a range-diff is below. Thanks in advance for taking another\n> look!\n\nI left some minor & nit-y comments on this v4, but overall I think this\nlooks really good with not much to add.\n"},{"id":"455483","messageId":"xmqq1qwpsjod.fsf@gitster.g","threadId":"56996","inReplyTo":"f494ef7377bf8fb14d96e860106033d1bd1c9ec1.1652915424.git.me@ttaylorr.com","subject":"Re: [PATCH v4 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-19T14:04:18Z","receivedAt":"2022-05-19T14:04:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> +== Caution for mixed-version environments\n> +\n> +Repositories that have cruft packs in them will continue to work with any older\n> +version of Git. Note, however, that previous versions of Git which do not\n> ...\n> +cruft packs.\n\nI've compared a rebase of the previous iteration on top of v2.36.0\nwith the result of application of this iteration on the same commit,\nand the above additional documentation seems to be the only real\ndifference.\n\nWill replace and queue.\n\nThanks.\n"},{"id":"455488","messageId":"xmqqv8u1r1r2.fsf@gitster.g","threadId":"56996","inReplyTo":"xmqq7d6hsusd.fsf@gitster.g","subject":"Re: [PATCH v4 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-19T15:16:49Z","receivedAt":"2022-05-19T15:17:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Taylor Blau <me@ttaylorr.com> writes:\n>\n>> @@ -3870,6 +4034,20 @@ static int option_parse_unpack_unreachable(const struct option *opt,\n>>  \treturn 0;\n>>  }\n>>  \n>> +static int option_parse_cruft_expiration(const struct option *opt,\n>> +\t\t\t\t\t const char *arg, int unset)\n>> +{\n>> +\tif (unset) {\n>> +\t\tcruft = 0;\n>> +\t\tcruft_expiration = 0;\n>> +\t} else {\n>> +\t\tcruft = 1;\n>> +\t\tif (arg)\n>> +\t\t\tcruft_expiration = approxidate(arg);\n>> +\t}\n>> +\treturn 0;\n>> +}\n>\n> It is somewhat sad that we have to invent this function, instead of\n> using parse_opt_expiry_date_cb().\n\nI failed to mention that this one does more than the bog-standard\ncallback so the latter cannot be reused as-is, and that is what I\nmeant by \"somewhat sad\".  If we can find a way to reuse the\nparse_opt_expiry_date_cb() for the purpose of the user of this\nfunction that would be ideal, but only if we can do so without\nmaking the caller too unnatural.  Having two separate values, \"did\nwe get --cruft-expiration option?\" and \"what's the value of it?\",\ndoes benefit the current caller and we do not want to twist it just\nfor not adding a similar callback---that's a tail wagging a dog.\n\nThanks.\n\n"},{"id":"455489","messageId":"xmqqr14pr1jt.fsf@gitster.g","threadId":"56996","inReplyTo":"220519.86zgjd4wvk.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 02/17] pack-mtimes: support reading .mtimes files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-19T15:21:10Z","receivedAt":"2022-05-19T15:21:21Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n> On Wed, May 18 2022, Taylor Blau wrote:\n>\n> Nit:\n>\n>> +  - A 4-byte magic number '0x4d544d45' ('MTME').\n>> +\n>> +  - A 4-byte version identifier (= 1).\n>> +\n>> +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n>\n> Here we let it suffice that later we'll say \"All 4-byte numbers are in\n> network order\".\n>\n>> +  - A table of 4-byte unsigned integers in network order. The ith\n>\n> But here we call out \"network order\" explicitly, shouldn't this just be\n> s/ in network order//?\n>\n>> +    value is the modification time (mtime) of the ith object in the\n>> +    corresponding pack by lexicographic (index) order. The mtimes\n>> +    count standard epoch seconds.\n>> +\n>> +  - A trailer, containing a checksum of the corresponding packfile,\n>> +    and a checksum of all of the above (each having length according\n>> +    to the specified hash function).\n>> +\n>> +All 4-byte numbers are in network order.\n>\n> I.e. this is sufficient.\n\nVery good eyes.  One explicit mention among several others can\nindeed be misleading the readers.\n\nWhen asked for \"network order\", all your search engines show are\nentries about \"network byte order\", so let's use that longer form of\nspelling.\n\nThanks.\n"},{"id":"455491","messageId":"xmqqmtfdr12r.fsf@gitster.g","threadId":"56996","inReplyTo":"RFC-cover-0.2-00000000000-20220519T113538Z-avarab@gmail.com","subject":"Re: [RFC PATCH 0/2] Utility functions for duplicated pack(write) code","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-19T15:31:24Z","receivedAt":"2022-05-19T15:31:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> Minor cleanups thath would semantically & textually conflict with\n> Taylor's\n> https://lore.kernel.org/git/cover.1652915424.git.me@ttaylorr.com/; but\n> which I noted while reading through it.\n\nVery much appreciated that you marked this as RFC.\n\nIt is natural and easy to notice problems in the code that is in\nflux, because it is inevitable for anybody working on the codebase\nto see the changes in-flight and the original code as they review,\nor as they make trial merges of their own work and see conflicts.\n\nBut making patches to address them immediately out of spinal reflex\nwould not help anybody.  Marking them as RFC and calling attention\nby those involved in the \"other topic\" while the code being cleaned\nup is still fresh in their mind makes it efficient to review the\nclean-up while letting the \"other topic\" to either proceed without\nclean-up or with it rolled in.\n\n"},{"id":"455492","messageId":"xmqqilq1r0no.fsf@gitster.g","threadId":"56996","inReplyTo":"RFC-patch-1.2-d8c3c03e90f-20220519T113538Z-avarab@gmail.com","subject":"Re: [RFC PATCH 1/2] packfile API: add and use a pack_name_to_ext() utility function","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-19T15:40:27Z","receivedAt":"2022-05-19T15:40:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> Add and use a pack_name_to_ext() utility function for the copy/pasted\n> cases of creating a FOO.ext file given a string like FOO.pack.\n\nI agree that \"remove .pack extension and replace it with .foo\nextention\" is a common thing to do and it would be a welcome\nsimplification.\n\nBut pack_name_to_ext() sounds like taking pack-123456...9876.pack\nand returning its extension (i.e. \".pack\") that will become useful\nwhen we introduce a drastically different naming convention and\nstart calling packfiles in newer format with different extensions\nlike \".pac4\".\n\nI wonder if this has easier-to-understand name that would be equally\n(or slightly more) useful?\n\n\tchar *replace_ext(const char *name, const char *src, const char *dst)\n\t{\n\t\tsize_t len;\n\n\t\tif (!strip_suffix(name, src, &len))\n\t\t\tBUG(\"name '%s' does not end in suffix '%s'\", name, src);\n\t\treturn xstrfmt(\"%.*s.%s\", (int)len, name, dst);\n\t}\n\n\n\n\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  pack-bitmap.c   |  6 +-----\n>  pack-revindex.c |  5 +----\n>  packfile.c      | 14 ++++++++++----\n>  packfile.h      |  9 +++++++++\n>  4 files changed, 21 insertions(+), 13 deletions(-)\n>\n> diff --git a/pack-bitmap.c b/pack-bitmap.c\n> index 97909d48da3..0c3770d038d 100644\n> --- a/pack-bitmap.c\n> +++ b/pack-bitmap.c\n> @@ -302,11 +302,7 @@ char *midx_bitmap_filename(struct multi_pack_index *midx)\n>  \n>  char *pack_bitmap_filename(struct packed_git *p)\n>  {\n> -\tsize_t len;\n> -\n> -\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n> -\t\tBUG(\"pack_name does not end in .pack\");\n> -\treturn xstrfmt(\"%.*s.bitmap\", (int)len, p->pack_name);\n> +\treturn pack_name_to_ext(p->pack_name, \"bitmap\");\n>  }\n>  \n>  static int open_midx_bitmap_1(struct bitmap_index *bitmap_git,\n> diff --git a/pack-revindex.c b/pack-revindex.c\n> index 08dc1601679..69dc5688796 100644\n> --- a/pack-revindex.c\n> +++ b/pack-revindex.c\n> @@ -179,10 +179,7 @@ static int create_pack_revindex_in_memory(struct packed_git *p)\n>  \n>  static char *pack_revindex_filename(struct packed_git *p)\n>  {\n> -\tsize_t len;\n> -\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n> -\t\tBUG(\"pack_name does not end in .pack\");\n> -\treturn xstrfmt(\"%.*s.rev\", (int)len, p->pack_name);\n> +\treturn pack_name_to_ext(p->pack_name, \"rev\");\n>  }\n>  \n>  #define RIDX_HEADER_SIZE (12)\n> diff --git a/packfile.c b/packfile.c\n> index 835b2d27164..bd6ad441bf5 100644\n> --- a/packfile.c\n> +++ b/packfile.c\n> @@ -191,15 +191,12 @@ int load_idx(const char *path, const unsigned int hashsz, void *idx_map,\n>  int open_pack_index(struct packed_git *p)\n>  {\n>  \tchar *idx_name;\n> -\tsize_t len;\n>  \tint ret;\n>  \n>  \tif (p->index_data)\n>  \t\treturn 0;\n>  \n> -\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n> -\t\tBUG(\"pack_name does not end in .pack\");\n> -\tidx_name = xstrfmt(\"%.*s.idx\", (int)len, p->pack_name);\n> +\tidx_name = pack_name_to_ext(p->pack_name, \"idx\");\n>  \tret = check_packed_git_idx(idx_name, p);\n>  \tfree(idx_name);\n>  \treturn ret;\n> @@ -2266,3 +2263,12 @@ int is_promisor_object(const struct object_id *oid)\n>  \t}\n>  \treturn oidset_contains(&promisor_objects, oid);\n>  }\n> +\n> +char *pack_name_to_ext(const char *pack_name, const char *ext)\n> +{\n> +\tsize_t len;\n> +\n> +\tif (!strip_suffix(pack_name, \".pack\", &len))\n> +\t\tBUG(\"pack_name does not end in .pack\");\n> +\treturn xstrfmt(\"%.*s.%s\", (int)len, pack_name, ext);\n> +}\n> diff --git a/packfile.h b/packfile.h\n> index a3f6723857b..6890c57ebdb 100644\n> --- a/packfile.h\n> +++ b/packfile.h\n> @@ -195,4 +195,13 @@ int is_promisor_object(const struct object_id *oid);\n>  int load_idx(const char *path, const unsigned int hashsz, void *idx_map,\n>  \t     size_t idx_size, struct packed_git *p);\n>  \n> +/**\n> + * Given a string like \"foo.pack\" and \"ext\" returns an xstrdup()'d\n> + * \"foo.ext\" string. Used for creating e.g. PACK.{bitmap,rev,...}\n> + * filenames from PACK.pack.\n> + *\n> + * Will BUG() if the expected string can't be created from the\n> + * \"pack_name\" argument.\n> + */\n> +char *pack_name_to_ext(const char *pack_name, const char *ext);\n>  #endif\n"},{"id":"455493","messageId":"xmqqbkvtr079.fsf@gitster.g","threadId":"56996","inReplyTo":"RFC-patch-2.2-051f0612ab9-20220519T113538Z-avarab@gmail.com","subject":"Re: [RFC PATCH 2/2] hash API: add and use a hash_short_id_by_algo() function","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-19T15:50:18Z","receivedAt":"2022-05-19T15:51:51Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> Add and use a hash_short_id_by_algo() function. As noted in the\n> comment being modified here (added in [1]) the intention wasn't to\n> have these end up in on-disk formats, but since [2], [3] and [3]\n\ndouble [3]?\n\n> that's been the case, and there's an outstanding patch to add another\n> format that uses these[5].\n>\n> So let's expose this functionality as a documented utility function,\n> instead of copy/pasting this code in various places.\n>\n> Replacing the die() in the existing functions with a BUG() might be\n> overzelous, it's correct for the case of\n> e.g. write_commit_graph_file() and write_midx_header(), but we also\n> use this for parsing on-disk files, e.g. in parse_commit_graph().\n\nIf we know the offending data can come from outside, not from\nliterals in the code, then there is no \"might be\", such a use of\nBUG() is simply wrong.\n\n> We could add a \"gently\" version of this, but for now I think that\n> worrying about the distinction would be worrying too much. If we ever\n> end up parsing such files that'll almost certainly be a bug in our own\n> writing code, so the distinction would be rather academic, even though\n\nIt would be a file written by an ancient buggy version of our code,\nor a buggy third-party reimplementation of Git.  It could be that a\nnew version of Git is using a yet-to-be-invented algorithm this\nversion of Git does not know about.\n\nThe distinction matters in that \"The version of Git I downloaded and\nbuilt last week out of the latest release tag said BUG\" should mean\nonly one thing: that version that reports BUG() is the culprit, not\nsome random other thing we do not even know where it came from that\nleft a corrupt data on disk.\n"},{"id":"455527","messageId":"220519.86wneh2v83.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"xmqqbkvtr079.fsf@gitster.g","subject":"Re: [RFC PATCH 2/2] hash API: add and use a hash_short_id_by_algo() function","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-19T19:07:51Z","receivedAt":"2022-05-19T19:11:51Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, May 19 2022, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n>\n>> Add and use a hash_short_id_by_algo() function. As noted in the\n>> comment being modified here (added in [1]) the intention wasn't to\n>> have these end up in on-disk formats, but since [2], [3] and [3]\n>\n> double [3]?\n\n*nod*, sorry.\n\n>> that's been the case, and there's an outstanding patch to add another\n>> format that uses these[5].\n>>\n>> So let's expose this functionality as a documented utility function,\n>> instead of copy/pasting this code in various places.\n>>\n>> Replacing the die() in the existing functions with a BUG() might be\n>> overzelous, it's correct for the case of\n>> e.g. write_commit_graph_file() and write_midx_header(), but we also\n>> use this for parsing on-disk files, e.g. in parse_commit_graph().\n>\n> If we know the offending data can come from outside, not from\n> literals in the code, then there is no \"might be\", such a use of\n> BUG() is simply wrong.\n\nFair enough, we could change it to have 1/2 of this (the writing) use\nBUG(), and die() for the other part, or just leave it as die() for both.\n\n>> We could add a \"gently\" version of this, but for now I think that\n>> worrying about the distinction would be worrying too much. If we ever\n>> end up parsing such files that'll almost certainly be a bug in our own\n>> writing code, so the distinction would be rather academic, even though\n>\n> It would be a file written by an ancient buggy version of our code,\n> or a buggy third-party reimplementation of Git.  It could be that a\n> new version of Git is using a yet-to-be-invented algorithm this\n> version of Git does not know about.\n>\n> The distinction matters in that \"The version of Git I downloaded and\n> built last week out of the latest release tag said BUG\" should mean\n> only one thing: that version that reports BUG() is the culprit, not\n> some random other thing we do not even know where it came from that\n> left a corrupt data on disk.\n\nMakes sense.\n"},{"id":"455554","messageId":"220520.86bkvs3bfm.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"xmqqr14pr1jt.fsf@gitster.g","subject":"Re: [PATCH v4 02/17] pack-mtimes: support reading .mtimes files","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-20T07:32:50Z","receivedAt":"2022-05-20T07:34:01Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, May 19 2022, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>\n>> On Wed, May 18 2022, Taylor Blau wrote:\n>>\n>> Nit:\n>>\n>>> +  - A 4-byte magic number '0x4d544d45' ('MTME').\n>>> +\n>>> +  - A 4-byte version identifier (= 1).\n>>> +\n>>> +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n>>\n>> Here we let it suffice that later we'll say \"All 4-byte numbers are in\n>> network order\".\n>>\n>>> +  - A table of 4-byte unsigned integers in network order. The ith\n>>\n>> But here we call out \"network order\" explicitly, shouldn't this just be\n>> s/ in network order//?\n>>\n>>> +    value is the modification time (mtime) of the ith object in the\n>>> +    corresponding pack by lexicographic (index) order. The mtimes\n>>> +    count standard epoch seconds.\n>>> +\n>>> +  - A trailer, containing a checksum of the corresponding packfile,\n>>> +    and a checksum of all of the above (each having length according\n>>> +    to the specified hash function).\n>>> +\n>>> +All 4-byte numbers are in network order.\n>>\n>> I.e. this is sufficient.\n>\n> Very good eyes.  One explicit mention among several others can\n> indeed be misleading the readers.\n>\n> When asked for \"network order\", all your search engines show are\n> entries about \"network byte order\", so let's use that longer form of\n> spelling.\n\n*Nod*, note that \"network order\" is on \"master\" already though,\ni.e. this section re-used a template introduced in 2f4ba2a867f\n(packfile: prepare for the existence of '*.rev' files, 2021-01-25) just\nabove this hunk.\n\nBefore that change the rest of the file used \"network byte order\"\nconsistently.\n"},{"id":"455635","messageId":"YogYP03hDH8hZqUV@nand.local","threadId":"56996","inReplyTo":"220520.86bkvs3bfm.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T22:37:51Z","receivedAt":"2022-05-20T22:37:56Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, May 20, 2022 at 09:32:50AM +0200, Ævar Arnfjörð Bjarmason wrote:\n>\n> On Thu, May 19 2022, Junio C Hamano wrote:\n>\n> > Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n> >\n> >> On Wed, May 18 2022, Taylor Blau wrote:\n> >>\n> >> Nit:\n> >>\n> >>> +  - A 4-byte magic number '0x4d544d45' ('MTME').\n> >>> +\n> >>> +  - A 4-byte version identifier (= 1).\n> >>> +\n> >>> +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n> >>\n> >> Here we let it suffice that later we'll say \"All 4-byte numbers are in\n> >> network order\".\n> >>\n> >>> +  - A table of 4-byte unsigned integers in network order. The ith\n> >>\n> >> But here we call out \"network order\" explicitly, shouldn't this just be\n> >> s/ in network order//?\n> >>\n> >>> +    value is the modification time (mtime) of the ith object in the\n> >>> +    corresponding pack by lexicographic (index) order. The mtimes\n> >>> +    count standard epoch seconds.\n> >>> +\n> >>> +  - A trailer, containing a checksum of the corresponding packfile,\n> >>> +    and a checksum of all of the above (each having length according\n> >>> +    to the specified hash function).\n> >>> +\n> >>> +All 4-byte numbers are in network order.\n> >>\n> >> I.e. this is sufficient.\n> >\n> > Very good eyes.  One explicit mention among several others can\n> > indeed be misleading the readers.\n> >\n> > When asked for \"network order\", all your search engines show are\n> > entries about \"network byte order\", so let's use that longer form of\n> > spelling.\n>\n> *Nod*, note that \"network order\" is on \"master\" already though,\n> i.e. this section re-used a template introduced in 2f4ba2a867f\n> (packfile: prepare for the existence of '*.rev' files, 2021-01-25) just\n> above this hunk.\n>\n> Before that change the rest of the file used \"network byte order\"\n> consistently.\n\nHmm. e0d1bcf825 (multi-pack-index: add format details, 2018-07-12)\n(which predates 2f4ba2a867f by a few years) introduced the first use of\n\"network order\" as opposed to \"network byte order\".\n\nI think it's worth cleaning this up, but let's do it in two parts.\nI'll send a rerolled version of tb/cruft-packs that moves the \"All\n4-byte numbers are in network order\" to the top of that section,\nswitching \"network order\" for \"network byte order\", and dropping other\nmentions of \"network [byte] order\" from that section.\n\nThen, we can come back later and perhaps do something like the following\n(but I don't want to do it now and tie up this series with semi-related\ncleanups):\n\n--- 8< ---\n\ndiff --git a/Documentation/technical/pack-format.txt b/Documentation/technical/pack-format.txt\nindex b520aa9c45..2591a410fd 100644\n--- a/Documentation/technical/pack-format.txt\n+++ b/Documentation/technical/pack-format.txt\n@@ -276,6 +276,8 @@ Pack file entry: <+\n\n == pack-*.rev files have the format:\n\n+All 4-byte numbers are in network byte order.\n+\n   - A 4-byte magic number '0x52494458' ('RIDX').\n\n   - A 4-byte version identifier (= 1).\n@@ -283,8 +285,8 @@ Pack file entry: <+\n   - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n\n   - A table of index positions (one per packed object, num_objects in\n-    total, each a 4-byte unsigned integer in network order), sorted by\n-    their corresponding offsets in the packfile.\n+    total, each a 4-byte unsigned integer), sorted by their\n+    corresponding offsets in the packfile.\n\n   - A trailer, containing a:\n\n@@ -292,8 +294,6 @@ Pack file entry: <+\n\n     a checksum of all of the above.\n\n-All 4-byte numbers are in network order.\n-\n == pack-*.mtimes files have the format:\n\n All 4-byte numbers are in network byte order.\n@@ -322,7 +322,7 @@ the body into \"chunks\" and provide a lookup table at the beginning of the\n body. The header includes certain length values, such as the number of packs,\n the number of base MIDX files, hash lengths and types.\n\n-All 4-byte numbers are in network order.\n+All 4-byte numbers are in network byte order.\n\n HEADER:\n\n@@ -397,8 +397,8 @@ CHUNK DATA:\n\n \t[Optional] Bitmap pack order (ID: {'R', 'I', 'D', 'X'})\n \t    A list of MIDX positions (one per object in the MIDX, num_objects in\n-\t    total, each a 4-byte unsigned integer in network byte order), sorted\n-\t    according to their relative bitmap/pseudo-pack positions.\n+\t    total, each a 4-byte unsigned integer), sorted according to their\n+\t    relative bitmap/pseudo-pack positions.\n\n TRAILER:\n\n\n--- >8 ---\n\nThanks,\nTaylor\n"},{"id":"455636","messageId":"YogYm50szub90Ift@nand.local","threadId":"56996","inReplyTo":"220519.86v8u14v52.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 12/17] builtin/repack.c: support generating a cruft pack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T22:39:23Z","receivedAt":"2022-05-20T22:39:29Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, May 19, 2022 at 01:29:26PM +0200, Ævar Arnfjörð Bjarmason wrote:\n>\n> On Wed, May 18 2022, Taylor Blau wrote:\n>\n> > +\t\ttip=\"$(git rev-parse cruft)\" &&\n>\n> Here we don't hide the exit status of \"git\", as it'll be reflected in what's &&-chained.\n\nOops! Nice catch.\n\n> > +\t\tpath=\"$objdir/$(test_oid_to_path \"$(git rev-parse cruft)\")\" &&\n>\n> But here we do, as we'll get the exit status of test_oid_to_path. But as\n> we just rev parsed it shouldn't this be $tip in any case?\n\nIndeed, this can just be:\n\n    path=\"$objdir/$(test_oid_to_path \"$tip\")\" &&\n\nWill include in a reroll shortly.\n\nThanks,\nTaylor\n"},{"id":"455637","messageId":"YogZPpWAamW1qREE@nand.local","threadId":"56996","inReplyTo":"220519.86r14p4v0q.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 15/17] builtin/repack.c: add cruft packs to MIDX during geometric repack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T22:42:06Z","receivedAt":"2022-05-20T22:42:12Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, May 19, 2022 at 01:32:26PM +0200, Ævar Arnfjörð Bjarmason wrote:\n> > @@ -161,8 +162,11 @@ static void collect_pack_filenames(struct string_list *fname_nonkept_list,\n> >  \t\tif ((extra_keep->nr > 0 && i < extra_keep->nr) ||\n> >  \t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname))))\n> >  \t\t\tstring_list_append_nodup(fname_kept_list, fname);\n> > -\t\telse\n> > -\t\t\tstring_list_append_nodup(fname_nonkept_list, fname);\n> > +\t\telse {\n> > +\t\t\tstruct string_list_item *item = string_list_append_nodup(fname_nonkept_list, fname);\n>\n> Nit: very long line, and we end up with {} just on the else, not the if.\n\nThanks for spotting, I split this line up and added braces to the other\nhalf of this conditional.\n\nThanks,\nTaylor\n"},{"id":"455638","messageId":"YogbupkkDSysm6Yw@nand.local","threadId":"56996","inReplyTo":"xmqqv8u1r1r2.fsf@gitster.g","subject":"Re: [PATCH v4 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T22:52:42Z","receivedAt":"2022-05-20T22:52:47Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, May 19, 2022 at 08:16:49AM -0700, Junio C Hamano wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n> > Taylor Blau <me@ttaylorr.com> writes:\n> >\n> >> @@ -3870,6 +4034,20 @@ static int option_parse_unpack_unreachable(const struct option *opt,\n> >>  \treturn 0;\n> >>  }\n> >>\n> >> +static int option_parse_cruft_expiration(const struct option *opt,\n> >> +\t\t\t\t\t const char *arg, int unset)\n> >> +{\n> >> +\tif (unset) {\n> >> +\t\tcruft = 0;\n> >> +\t\tcruft_expiration = 0;\n> >> +\t} else {\n> >> +\t\tcruft = 1;\n> >> +\t\tif (arg)\n> >> +\t\t\tcruft_expiration = approxidate(arg);\n> >> +\t}\n> >> +\treturn 0;\n> >> +}\n> >\n> > It is somewhat sad that we have to invent this function, instead of\n> > using parse_opt_expiry_date_cb().\n>\n> I failed to mention that this one does more than the bog-standard\n> callback so the latter cannot be reused as-is, and that is what I\n> meant by \"somewhat sad\".  If we can find a way to reuse the\n> parse_opt_expiry_date_cb() for the purpose of the user of this\n> function that would be ideal, but only if we can do so without\n> making the caller too unnatural.  Having two separate values, \"did\n> we get --cruft-expiration option?\" and \"what's the value of it?\",\n> does benefit the current caller and we do not want to twist it just\n> for not adding a similar callback---that's a tail wagging a dog.\n\nI agree, though I'm not sure such a cleanup is possible: if the caller\nspecified `--cruft-expiration=never`, how would we distinguish that from\n\"the caller did not want to generate cruft packs\"?\n\nIn that case, approxidate() would set our `cruft_expiration` variable to\n`0` there, which makes \"generate a cruft pack without expiring any\nobjects\" indistinguishable from \"do not generate a cruft pack\" without\nthe additional bit of information stored in the \"cruft\" variable.\n\nFor now we're just recycling the pattern from the callback immediately\nabove this one: option_parse_unpack_unreachable().\n\nThanks,\nTaylor\n"},{"id":"455639","messageId":"cover.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1638224692.git.me@ttaylorr.com","subject":"[PATCH v5 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:17:29Z","receivedAt":"2022-05-20T23:17:41Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Here is another reroll of my series to implement \"cruft packs\". This is really\nmore like v4.1, since it has only cosmetic changes incorporated from the review\non v4 of this topic.\n\nSince last time:\n\n  - The new section in pack-format.txt (describing the \".mtimes\" format) now\n    says at the top \"all 4-byte numbers are in network byte order\", and avoids\n    repeating \"network [byte] order\" throughout that section to reduce\n    confusion.\n\n  - A sub-shell in t5329 which incorrectly masked over the exit code of a \"git\"\n    process was removed.\n\n  - An overly-long line in builtin/repack.c::collect_pack_filenames() was\n    eliminated, and matching braces are added.\n\n...and that's pretty much it. In any case, a range-diff is included below.\nThanks again for all of the thoughtful feedback on this series.\n\nTaylor Blau (17):\n  Documentation/technical: add cruft-packs.txt\n  pack-mtimes: support reading .mtimes files\n  pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'\n  chunk-format.h: extract oid_version()\n  pack-mtimes: support writing pack .mtimes files\n  t/helper: add 'pack-mtimes' test-tool\n  builtin/pack-objects.c: return from create_object_entry()\n  builtin/pack-objects.c: --cruft without expiration\n  reachable: add options to add_unseen_recent_objects_to_traversal\n  reachable: report precise timestamps from objects in cruft packs\n  builtin/pack-objects.c: --cruft with expiration\n  builtin/repack.c: support generating a cruft pack\n  builtin/repack.c: allow configuring cruft pack generation\n  builtin/repack.c: use named flags for existing_packs\n  builtin/repack.c: add cruft packs to MIDX during geometric repack\n  builtin/gc.c: conditionally avoid pruning objects via loose\n  sha1-file.c: don't freshen cruft packs\n\n Documentation/Makefile                  |   1 +\n Documentation/config/gc.txt             |  21 +-\n Documentation/config/repack.txt         |   9 +\n Documentation/git-gc.txt                |   5 +\n Documentation/git-pack-objects.txt      |  30 +\n Documentation/git-repack.txt            |  11 +\n Documentation/technical/cruft-packs.txt | 123 ++++\n Documentation/technical/pack-format.txt |  19 +\n Makefile                                |   2 +\n builtin/gc.c                            |  10 +-\n builtin/pack-objects.c                  | 304 +++++++++-\n builtin/repack.c                        | 185 +++++-\n bulk-checkin.c                          |   2 +-\n chunk-format.c                          |  12 +\n chunk-format.h                          |   3 +\n commit-graph.c                          |  18 +-\n midx.c                                  |  18 +-\n object-file.c                           |   4 +-\n object-store.h                          |   7 +-\n pack-mtimes.c                           | 126 ++++\n pack-mtimes.h                           |  15 +\n pack-objects.c                          |   6 +\n pack-objects.h                          |  25 +\n pack-write.c                            |  93 ++-\n pack.h                                  |   4 +\n packfile.c                              |  19 +-\n reachable.c                             |  58 +-\n reachable.h                             |   9 +-\n t/helper/test-pack-mtimes.c             |  56 ++\n t/helper/test-tool.c                    |   1 +\n t/helper/test-tool.h                    |   1 +\n t/t5329-pack-objects-cruft.sh           | 739 ++++++++++++++++++++++++\n 32 files changed, 1834 insertions(+), 102 deletions(-)\n create mode 100644 Documentation/technical/cruft-packs.txt\n create mode 100644 pack-mtimes.c\n create mode 100644 pack-mtimes.h\n create mode 100644 t/helper/test-pack-mtimes.c\n create mode 100755 t/t5329-pack-objects-cruft.sh\n\nRange-diff against v4:\n -:  ---------- >  1:  f494ef7377 Documentation/technical: add cruft-packs.txt\n 1:  8f9fd21be9 !  2:  91a9d21b0b pack-mtimes: support reading .mtimes files\n    @@ Documentation/technical/pack-format.txt: Pack file entry: <+\n      \n     +== pack-*.mtimes files have the format:\n     +\n    ++All 4-byte numbers are in network byte order.\n    ++\n     +  - A 4-byte magic number '0x4d544d45' ('MTME').\n     +\n     +  - A 4-byte version identifier (= 1).\n     +\n     +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n     +\n    -+  - A table of 4-byte unsigned integers in network order. The ith\n    -+    value is the modification time (mtime) of the ith object in the\n    -+    corresponding pack by lexicographic (index) order. The mtimes\n    -+    count standard epoch seconds.\n    ++  - A table of 4-byte unsigned integers. The ith value is the\n    ++    modification time (mtime) of the ith object in the corresponding\n    ++    pack by lexicographic (index) order. The mtimes count standard\n    ++    epoch seconds.\n     +\n     +  - A trailer, containing a checksum of the corresponding packfile,\n     +    and a checksum of all of the above (each having length according\n     +    to the specified hash function).\n    -+\n    -+All 4-byte numbers are in network order.\n     +\n      == multi-pack-index (MIDX) files have the following format:\n      \n 2:  cdb21236e1 =  3:  67c4e7209d pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'\n 3:  1d775f9850 =  4:  fc86506881 chunk-format.h: extract oid_version()\n 4:  6172861bd9 =  5:  788d1f96f2 pack-mtimes: support writing pack .mtimes files\n 5:  5f9a9a5b7b =  6:  2a6cfb00bf t/helper: add 'pack-mtimes' test-tool\n 6:  b8a38fe2e4 =  7:  edb6fcd5ec builtin/pack-objects.c: return from create_object_entry()\n 7:  94fe03cc65 =  8:  e3185741f2 builtin/pack-objects.c: --cruft without expiration\n 8:  da7273f41f =  9:  1cf00d462c reachable: add options to add_unseen_recent_objects_to_traversal\n 9:  58fecd1747 = 10:  d66be44d9a reachable: report precise timestamps from objects in cruft packs\n10:  1740b8ef01 = 11:  1434e37623 builtin/pack-objects.c: --cruft with expiration\n11:  5992a72cbf ! 12:  0d3555d595 builtin/repack.c: support generating a cruft pack\n    @@ t/t5329-pack-objects-cruft.sh: test_expect_success 'expired objects are pruned'\n     +\t\tgit repack &&\n     +\n     +\t\ttip=\"$(git rev-parse cruft)\" &&\n    -+\t\tpath=\"$objdir/$(test_oid_to_path \"$(git rev-parse cruft)\")\" &&\n    ++\t\tpath=\"$objdir/$(test_oid_to_path \"$tip\")\" &&\n     +\t\ttest-tool chmtime --get +1000 \"$path\" >expect &&\n     +\n     +\t\tgit checkout main &&\n12:  1b241f8f91 = 13:  4b721d3ee9 builtin/repack.c: allow configuring cruft pack generation\n13:  ffae78852c = 14:  f9e3ab56b1 builtin/repack.c: use named flags for existing_packs\n14:  0743e373ba ! 15:  e9f46e7b5e builtin/repack.c: add cruft packs to MIDX during geometric repack\n    @@ builtin/repack.c\n      static int pack_everything;\n      static int delta_base_offset = 1;\n     @@ builtin/repack.c: static void collect_pack_filenames(struct string_list *fname_nonkept_list,\n    + \t\tfname = xmemdupz(e->d_name, len);\n    + \n      \t\tif ((extra_keep->nr > 0 && i < extra_keep->nr) ||\n    - \t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname))))\n    +-\t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname))))\n    ++\t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname)))) {\n      \t\t\tstring_list_append_nodup(fname_kept_list, fname);\n     -\t\telse\n     -\t\t\tstring_list_append_nodup(fname_nonkept_list, fname);\n    -+\t\telse {\n    -+\t\t\tstruct string_list_item *item = string_list_append_nodup(fname_nonkept_list, fname);\n    ++\t\t} else {\n    ++\t\t\tstruct string_list_item *item;\n    ++\t\t\titem = string_list_append_nodup(fname_nonkept_list,\n    ++\t\t\t\t\t\t\tfname);\n     +\t\t\tif (file_exists(mkpath(\"%s/%s.mtimes\", packdir, fname)))\n     +\t\t\t\titem->util = (void*)(uintptr_t)CRUFT_PACK;\n     +\t\t}\n15:  9f7e0acac6 = 16:  43c14eec07 builtin/gc.c: conditionally avoid pruning objects via loose\n16:  07fa9d4b47 = 17:  1e313b89e8 sha1-file.c: don't freshen cruft packs\n-- \n2.36.1.94.gb0d54bedca\n"},{"id":"455640","messageId":"f494ef7377bf8fb14d96e860106033d1bd1c9ec1.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 01/17] Documentation/technical: add cruft-packs.txt","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:17:32Z","receivedAt":"2022-05-20T23:17:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Create a technical document to explain cruft packs. It contains a brief\noverview of the problem, some background, details on the implementation,\nand a couple of alternative approaches not considered here.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/Makefile                  |   1 +\n Documentation/technical/cruft-packs.txt | 123 ++++++++++++++++++++++++\n 2 files changed, 124 insertions(+)\n create mode 100644 Documentation/technical/cruft-packs.txt\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex adb2f1b50a..2faffb52ab 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -94,6 +94,7 @@ TECH_DOCS += MyFirstContribution\n TECH_DOCS += MyFirstObjectWalk\n TECH_DOCS += SubmittingPatches\n TECH_DOCS += technical/bundle-format\n+TECH_DOCS += technical/cruft-packs\n TECH_DOCS += technical/hash-function-transition\n TECH_DOCS += technical/http-protocol\n TECH_DOCS += technical/index-format\ndiff --git a/Documentation/technical/cruft-packs.txt b/Documentation/technical/cruft-packs.txt\nnew file mode 100644\nindex 0000000000..c0f583cd48\n--- /dev/null\n+++ b/Documentation/technical/cruft-packs.txt\n@@ -0,0 +1,123 @@\n+= Cruft packs\n+\n+The cruft packs feature offer an alternative to Git's traditional mechanism of\n+removing unreachable objects. This document provides an overview of Git's\n+pruning mechanism, and how a cruft pack can be used instead to accomplish the\n+same.\n+\n+== Background\n+\n+To remove unreachable objects from your repository, Git offers `git repack -Ad`\n+(see linkgit:git-repack[1]). Quoting from the documentation:\n+\n+[quote]\n+[...] unreachable objects in a previous pack become loose, unpacked objects,\n+instead of being left in the old pack. [...] loose unreachable objects will be\n+pruned according to normal expiry rules with the next 'git gc' invocation.\n+\n+Unreachable objects aren't removed immediately, since doing so could race with\n+an incoming push which may reference an object which is about to be deleted.\n+Instead, those unreachable objects are stored as loose object and stay that way\n+until they are older than the expiration window, at which point they are removed\n+by linkgit:git-prune[1].\n+\n+Git must store these unreachable objects loose in order to keep track of their\n+per-object mtimes. If these unreachable objects were written into one big pack,\n+then either freshening that pack (because an object contained within it was\n+re-written) or creating a new pack of unreachable objects would cause the pack's\n+mtime to get updated, and the objects within it would never leave the expiration\n+window. Instead, objects are stored loose in order to keep track of the\n+individual object mtimes and avoid a situation where all cruft objects are\n+freshened at once.\n+\n+This can lead to undesirable situations when a repository contains many\n+unreachable objects which have not yet left the grace period. Having large\n+directories in the shards of `.git/objects` can lead to decreased performance in\n+the repository. But given enough unreachable objects, this can lead to inode\n+starvation and degrade the performance of the whole system. Since we\n+can never pack those objects, these repositories often take up a large amount of\n+disk space, since we can only zlib compress them, but not store them in delta\n+chains.\n+\n+== Cruft packs\n+\n+A cruft pack eliminates the need for storing unreachable objects in a loose\n+state by including the per-object mtimes in a separate file alongside a single\n+pack containing all loose objects.\n+\n+A cruft pack is written by `git repack --cruft` when generating a new pack.\n+linkgit:git-pack-objects[1]'s `--cruft` option. Note that `git repack --cruft`\n+is a classic all-into-one repack, meaning that everything in the resulting pack is\n+reachable, and everything else is unreachable. Once written, the `--cruft`\n+option instructs `git repack` to generate another pack containing only objects\n+not packed in the previous step (which equates to packing all unreachable\n+objects together). This progresses as follows:\n+\n+  1. Enumerate every object, marking any object which is (a) not contained in a\n+     kept-pack, and (b) whose mtime is within the grace period as a traversal\n+     tip.\n+\n+  2. Perform a reachability traversal based on the tips gathered in the previous\n+     step, adding every object along the way to the pack.\n+\n+  3. Write the pack out, along with a `.mtimes` file that records the per-object\n+     timestamps.\n+\n+This mode is invoked internally by linkgit:git-repack[1] when instructed to\n+write a cruft pack. Crucially, the set of in-core kept packs is exactly the set\n+of packs which will not be deleted by the repack; in other words, they contain\n+all of the repository's reachable objects.\n+\n+When a repository already has a cruft pack, `git repack --cruft` typically only\n+adds objects to it. An exception to this is when `git repack` is given the\n+`--cruft-expiration` option, which allows the generated cruft pack to omit\n+expired objects instead of waiting for linkgit:git-gc[1] to expire those objects\n+later on.\n+\n+It is linkgit:git-gc[1] that is typically responsible for removing expired\n+unreachable objects.\n+\n+== Caution for mixed-version environments\n+\n+Repositories that have cruft packs in them will continue to work with any older\n+version of Git. Note, however, that previous versions of Git which do not\n+understand the `.mtimes` file will use the cruft pack's mtime as the mtime for\n+all of the objects in it. In other words, do not expect older (pre-cruft pack)\n+versions of Git to interpret or even read the contents of the `.mtimes` file.\n+\n+Note that having mixed versions of Git GC-ing the same repository can lead to\n+unreachable objects never being completely pruned. This can happen under the\n+following circumstances:\n+\n+  - An older version of Git running GC explodes the contents of an existing\n+    cruft pack loose, using the cruft pack's mtime.\n+  - A newer version running GC collects those loose objects into a cruft pack,\n+    where the .mtime file reflects the loose object's actual mtimes, but the\n+    cruft pack mtime is \"now\".\n+\n+Repeating this process will lead to unreachable objects not getting pruned as a\n+result of repeatedly resetting the objects' mtimes to the present time.\n+\n+If you are GC-ing repositories in a mixed version environment, consider omitting\n+the `--cruft` option when using linkgit:git-repack[1] and linkgit:git-gc[1], and\n+leaving the `gc.cruftPacks` configuration unset until all writers understand\n+cruft packs.\n+\n+== Alternatives\n+\n+Notable alternatives to this design include:\n+\n+  - The location of the per-object mtime data, and\n+  - Storing unreachable objects in multiple cruft packs.\n+\n+On the location of mtime data, a new auxiliary file tied to the pack was chosen\n+to avoid complicating the `.idx` format. If the `.idx` format were ever to gain\n+support for optional chunks of data, it may make sense to consolidate the\n+`.mtimes` format into the `.idx` itself.\n+\n+Storing unreachable objects among multiple cruft packs (e.g., creating a new\n+cruft pack during each repacking operation including only unreachable objects\n+which aren't already stored in an earlier cruft pack) is significantly more\n+complicated to construct, and so aren't pursued here. The obvious drawback to\n+the current implementation is that the entire cruft pack must be re-written from\n+scratch.\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455641","messageId":"788d1f96f22323f93dec8cf37385865984eb246b.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 05/17] pack-mtimes: support writing pack .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:17:43Z","receivedAt":"2022-05-20T23:17:57Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Now that the `.mtimes` format is defined, supplement the pack-write API\nto be able to conditionally write an `.mtimes` file along with a pack by\nsetting an additional flag and passing an oidmap that contains the\ntimestamps corresponding to each object in the pack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n pack-objects.c |  6 ++++\n pack-objects.h | 25 ++++++++++++++++\n pack-write.c   | 77 ++++++++++++++++++++++++++++++++++++++++++++++++++\n pack.h         |  1 +\n 4 files changed, 109 insertions(+)\n\ndiff --git a/pack-objects.c b/pack-objects.c\nindex fe2a4eace9..272e8d4517 100644\n--- a/pack-objects.c\n+++ b/pack-objects.c\n@@ -170,6 +170,9 @@ struct object_entry *packlist_alloc(struct packing_data *pdata,\n \n \t\tif (pdata->layer)\n \t\t\tREALLOC_ARRAY(pdata->layer, pdata->nr_alloc);\n+\n+\t\tif (pdata->cruft_mtime)\n+\t\t\tREALLOC_ARRAY(pdata->cruft_mtime, pdata->nr_alloc);\n \t}\n \n \tnew_entry = pdata->objects + pdata->nr_objects++;\n@@ -198,6 +201,9 @@ struct object_entry *packlist_alloc(struct packing_data *pdata,\n \tif (pdata->layer)\n \t\tpdata->layer[pdata->nr_objects - 1] = 0;\n \n+\tif (pdata->cruft_mtime)\n+\t\tpdata->cruft_mtime[pdata->nr_objects - 1] = 0;\n+\n \treturn new_entry;\n }\n \ndiff --git a/pack-objects.h b/pack-objects.h\nindex dca2351ef9..393b9db546 100644\n--- a/pack-objects.h\n+++ b/pack-objects.h\n@@ -168,6 +168,14 @@ struct packing_data {\n \t/* delta islands */\n \tunsigned int *tree_depth;\n \tunsigned char *layer;\n+\n+\t/*\n+\t * Used when writing cruft packs.\n+\t *\n+\t * Object mtimes are stored in pack order when writing, but\n+\t * written out in lexicographic (index) order.\n+\t */\n+\tuint32_t *cruft_mtime;\n };\n \n void prepare_packing_data(struct repository *r, struct packing_data *pdata);\n@@ -289,4 +297,21 @@ static inline void oe_set_layer(struct packing_data *pack,\n \tpack->layer[e - pack->objects] = layer;\n }\n \n+static inline uint32_t oe_cruft_mtime(struct packing_data *pack,\n+\t\t\t\t      struct object_entry *e)\n+{\n+\tif (!pack->cruft_mtime)\n+\t\treturn 0;\n+\treturn pack->cruft_mtime[e - pack->objects];\n+}\n+\n+static inline void oe_set_cruft_mtime(struct packing_data *pack,\n+\t\t\t\t      struct object_entry *e,\n+\t\t\t\t      uint32_t mtime)\n+{\n+\tif (!pack->cruft_mtime)\n+\t\tCALLOC_ARRAY(pack->cruft_mtime, pack->nr_alloc);\n+\tpack->cruft_mtime[e - pack->objects] = mtime;\n+}\n+\n #endif\ndiff --git a/pack-write.c b/pack-write.c\nindex 27b171e440..23c0342018 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -3,6 +3,10 @@\n #include \"csum-file.h\"\n #include \"remote.h\"\n #include \"chunk-format.h\"\n+#include \"pack-mtimes.h\"\n+#include \"oidmap.h\"\n+#include \"chunk-format.h\"\n+#include \"pack-objects.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -277,6 +281,70 @@ const char *write_rev_file_order(const char *rev_name,\n \treturn rev_name;\n }\n \n+static void write_mtimes_header(struct hashfile *f)\n+{\n+\thashwrite_be32(f, MTIMES_SIGNATURE);\n+\thashwrite_be32(f, MTIMES_VERSION);\n+\thashwrite_be32(f, oid_version(the_hash_algo));\n+}\n+\n+/*\n+ * Writes the object mtimes of \"objects\" for use in a .mtimes file.\n+ * Note that objects must be in lexicographic (index) order, which is\n+ * the expected ordering of these values in the .mtimes file.\n+ */\n+static void write_mtimes_objects(struct hashfile *f,\n+\t\t\t\t struct packing_data *to_pack,\n+\t\t\t\t struct pack_idx_entry **objects,\n+\t\t\t\t uint32_t nr_objects)\n+{\n+\tuint32_t i;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\tstruct object_entry *e = (struct object_entry*)objects[i];\n+\t\thashwrite_be32(f, oe_cruft_mtime(to_pack, e));\n+\t}\n+}\n+\n+static void write_mtimes_trailer(struct hashfile *f, const unsigned char *hash)\n+{\n+\thashwrite(f, hash, the_hash_algo->rawsz);\n+}\n+\n+static const char *write_mtimes_file(const char *mtimes_name,\n+\t\t\t\t     struct packing_data *to_pack,\n+\t\t\t\t     struct pack_idx_entry **objects,\n+\t\t\t\t     uint32_t nr_objects,\n+\t\t\t\t     const unsigned char *hash)\n+{\n+\tstruct hashfile *f;\n+\tint fd;\n+\n+\tif (!to_pack)\n+\t\tBUG(\"cannot call write_mtimes_file with NULL packing_data\");\n+\n+\tif (!mtimes_name) {\n+\t\tstruct strbuf tmp_file = STRBUF_INIT;\n+\t\tfd = odb_mkstemp(&tmp_file, \"pack/tmp_mtimes_XXXXXX\");\n+\t\tmtimes_name = strbuf_detach(&tmp_file, NULL);\n+\t} else {\n+\t\tunlink(mtimes_name);\n+\t\tfd = xopen(mtimes_name, O_CREAT|O_EXCL|O_WRONLY, 0600);\n+\t}\n+\tf = hashfd(fd, mtimes_name);\n+\n+\twrite_mtimes_header(f);\n+\twrite_mtimes_objects(f, to_pack, objects, nr_objects);\n+\twrite_mtimes_trailer(f, hash);\n+\n+\tif (adjust_shared_perm(mtimes_name) < 0)\n+\t\tdie(_(\"failed to make %s readable\"), mtimes_name);\n+\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE | CSUM_FSYNC);\n+\n+\treturn mtimes_name;\n+}\n+\n off_t write_pack_header(struct hashfile *f, uint32_t nr_entries)\n {\n \tstruct pack_header hdr;\n@@ -479,6 +547,7 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t char **idx_tmp_name)\n {\n \tconst char *rev_tmp_name = NULL;\n+\tconst char *mtimes_tmp_name = NULL;\n \n \tif (adjust_shared_perm(pack_tmp_name))\n \t\tdie_errno(\"unable to make temporary pack file readable\");\n@@ -491,9 +560,17 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \trev_tmp_name = write_rev_file(NULL, written_list, nr_written, hash,\n \t\t\t\t      pack_idx_opts->flags);\n \n+\tif (pack_idx_opts->flags & WRITE_MTIMES) {\n+\t\tmtimes_tmp_name = write_mtimes_file(NULL, to_pack, written_list,\n+\t\t\t\t\t\t    nr_written,\n+\t\t\t\t\t\t    hash);\n+\t}\n+\n \trename_tmp_packfile(name_buffer, pack_tmp_name, \"pack\");\n \tif (rev_tmp_name)\n \t\trename_tmp_packfile(name_buffer, rev_tmp_name, \"rev\");\n+\tif (mtimes_tmp_name)\n+\t\trename_tmp_packfile(name_buffer, mtimes_tmp_name, \"mtimes\");\n }\n \n void write_promisor_file(const char *promisor_name, struct ref **sought, int nr_sought)\ndiff --git a/pack.h b/pack.h\nindex fd27cfdfd7..01d385903a 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -44,6 +44,7 @@ struct pack_idx_option {\n #define WRITE_IDX_STRICT 02\n #define WRITE_REV 04\n #define WRITE_REV_VERIFY 010\n+#define WRITE_MTIMES 020\n \n \tuint32_t version;\n \tuint32_t off32_limit;\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455642","messageId":"fc8650688193d5cf519281bf4ad4528f86e4abe1.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 04/17] chunk-format.h: extract oid_version()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:17:41Z","receivedAt":"2022-05-20T23:18:00Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"There are three definitions of an identical function which converts\n`the_hash_algo` into either 1 (for SHA-1) or 2 (for SHA-256). There is a\ncopy of this function for writing both the commit-graph and\nmulti-pack-index file, and another inline definition used to write the\n.rev header.\n\nConsolidate these into a single definition in chunk-format.h. It's not\nclear that this is the best header to define this function in, but it\nshould do for now.\n\n(Worth noting, the .rev caller expects a 4-byte unsigned, but the other\ntwo callers work with a single unsigned byte. The consolidated version\nuses the latter type, and lets the compiler widen it when required).\n\nAnother caller will be added in a subsequent patch.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n chunk-format.c | 12 ++++++++++++\n chunk-format.h |  3 +++\n commit-graph.c | 18 +++---------------\n midx.c         | 18 +++---------------\n pack-write.c   | 15 ++-------------\n 5 files changed, 23 insertions(+), 43 deletions(-)\n\ndiff --git a/chunk-format.c b/chunk-format.c\nindex 1c3dca62e2..0275b74a89 100644\n--- a/chunk-format.c\n+++ b/chunk-format.c\n@@ -181,3 +181,15 @@ int read_chunk(struct chunkfile *cf,\n \n \treturn CHUNK_NOT_FOUND;\n }\n+\n+uint8_t oid_version(const struct git_hash_algo *algop)\n+{\n+\tswitch (hash_algo_by_ptr(algop)) {\n+\tcase GIT_HASH_SHA1:\n+\t\treturn 1;\n+\tcase GIT_HASH_SHA256:\n+\t\treturn 2;\n+\tdefault:\n+\t\tdie(_(\"invalid hash version\"));\n+\t}\n+}\ndiff --git a/chunk-format.h b/chunk-format.h\nindex 9ccbe00377..7885aa0848 100644\n--- a/chunk-format.h\n+++ b/chunk-format.h\n@@ -2,6 +2,7 @@\n #define CHUNK_FORMAT_H\n \n #include \"git-compat-util.h\"\n+#include \"hash.h\"\n \n struct hashfile;\n struct chunkfile;\n@@ -65,4 +66,6 @@ int read_chunk(struct chunkfile *cf,\n \t       chunk_read_fn fn,\n \t       void *data);\n \n+uint8_t oid_version(const struct git_hash_algo *algop);\n+\n #endif\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 06107beedc..066d82ed6a 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -193,18 +193,6 @@ char *get_commit_graph_chain_filename(struct object_directory *odb)\n \treturn xstrfmt(\"%s/info/commit-graphs/commit-graph-chain\", odb->path);\n }\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n static struct commit_graph *alloc_commit_graph(void)\n {\n \tstruct commit_graph *g = xcalloc(1, sizeof(*g));\n@@ -365,9 +353,9 @@ struct commit_graph *parse_commit_graph(struct repository *r,\n \t}\n \n \thash_version = *(unsigned char*)(data + 5);\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != oid_version(the_hash_algo)) {\n \t\terror(_(\"commit-graph hash version %X does not match version %X\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, oid_version(the_hash_algo));\n \t\treturn NULL;\n \t}\n \n@@ -1924,7 +1912,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \thashwrite_be32(f, GRAPH_SIGNATURE);\n \n \thashwrite_u8(f, GRAPH_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, oid_version(the_hash_algo));\n \thashwrite_u8(f, get_num_chunks(cf));\n \thashwrite_u8(f, ctx->num_commit_graphs_after - 1);\n \ndiff --git a/midx.c b/midx.c\nindex 3db0e47735..c617c51cd0 100644\n--- a/midx.c\n+++ b/midx.c\n@@ -41,18 +41,6 @@\n \n #define PACK_EXPIRED UINT_MAX\n \n-static uint8_t oid_version(void)\n-{\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\treturn 1;\n-\tcase GIT_HASH_SHA256:\n-\t\treturn 2;\n-\tdefault:\n-\t\tdie(_(\"invalid hash version\"));\n-\t}\n-}\n-\n const unsigned char *get_midx_checksum(struct multi_pack_index *m)\n {\n \treturn m->data + m->data_len - the_hash_algo->rawsz;\n@@ -134,9 +122,9 @@ struct multi_pack_index *load_multi_pack_index(const char *object_dir, int local\n \t\t      m->version);\n \n \thash_version = m->data[MIDX_BYTE_HASH_VERSION];\n-\tif (hash_version != oid_version()) {\n+\tif (hash_version != oid_version(the_hash_algo)) {\n \t\terror(_(\"multi-pack-index hash version %u does not match version %u\"),\n-\t\t      hash_version, oid_version());\n+\t\t      hash_version, oid_version(the_hash_algo));\n \t\tgoto cleanup_fail;\n \t}\n \tm->hash_len = the_hash_algo->rawsz;\n@@ -420,7 +408,7 @@ static size_t write_midx_header(struct hashfile *f,\n {\n \thashwrite_be32(f, MIDX_SIGNATURE);\n \thashwrite_u8(f, MIDX_VERSION);\n-\thashwrite_u8(f, oid_version());\n+\thashwrite_u8(f, oid_version(the_hash_algo));\n \thashwrite_u8(f, num_chunks);\n \thashwrite_u8(f, 0); /* unused */\n \thashwrite_be32(f, num_packs);\ndiff --git a/pack-write.c b/pack-write.c\nindex a2adc565f4..27b171e440 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -2,6 +2,7 @@\n #include \"pack.h\"\n #include \"csum-file.h\"\n #include \"remote.h\"\n+#include \"chunk-format.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -181,21 +182,9 @@ static int pack_order_cmp(const void *va, const void *vb, void *ctx)\n \n static void write_rev_header(struct hashfile *f)\n {\n-\tuint32_t oid_version;\n-\tswitch (hash_algo_by_ptr(the_hash_algo)) {\n-\tcase GIT_HASH_SHA1:\n-\t\toid_version = 1;\n-\t\tbreak;\n-\tcase GIT_HASH_SHA256:\n-\t\toid_version = 2;\n-\t\tbreak;\n-\tdefault:\n-\t\tdie(\"write_rev_header: unknown hash version\");\n-\t}\n-\n \thashwrite_be32(f, RIDX_SIGNATURE);\n \thashwrite_be32(f, RIDX_VERSION);\n-\thashwrite_be32(f, oid_version);\n+\thashwrite_be32(f, oid_version(the_hash_algo));\n }\n \n static void write_rev_index_positions(struct hashfile *f,\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455643","messageId":"67c4e7209d510c09b5ddf018204f84ad0d0ebf9b.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 03/17] pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:17:38Z","receivedAt":"2022-05-20T23:18:03Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This structure will be used to communicate the per-object mtimes when\nwriting a cruft pack. Here, we need the full packing_data structure\nbecause the mtime information is stored in an array there, not on the\nindividual object_entry's themselves (to avoid paying the overhead in\nstructure width for operations which do not generate a cruft pack).\n\nWe haven't passed this information down before because one of the two\ncallers (in bulk-checkin.c) does not have a packing_data structure at\nall. In that case (where no cruft pack will be generated), NULL is\npassed instead.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 3 ++-\n bulk-checkin.c         | 2 +-\n pack-write.c           | 1 +\n pack.h                 | 3 +++\n 4 files changed, 7 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 014dcd4bc9..6ac927047c 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1262,7 +1262,8 @@ static void write_pack_file(void)\n \n \t\t\tstage_tmp_packfiles(&tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n-\t\t\t\t\t    &pack_idx_opts, hash, &idx_tmp_name);\n+\t\t\t\t\t    &to_pack, &pack_idx_opts, hash,\n+\t\t\t\t\t    &idx_tmp_name);\n \n \t\t\tif (write_bitmap_index) {\n \t\t\t\tsize_t tmpname_len = tmpname.len;\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6d6c37171c..e988a388b6 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -33,7 +33,7 @@ static void finish_tmp_packfile(struct strbuf *basename,\n \tchar *idx_tmp_name = NULL;\n \n \tstage_tmp_packfiles(basename, pack_tmp_name, written_list, nr_written,\n-\t\t\t    pack_idx_opts, hash, &idx_tmp_name);\n+\t\t\t    NULL, pack_idx_opts, hash, &idx_tmp_name);\n \trename_tmp_packfile_idx(basename, &idx_tmp_name);\n \n \tfree(idx_tmp_name);\ndiff --git a/pack-write.c b/pack-write.c\nindex 51812cb129..a2adc565f4 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -484,6 +484,7 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t const char *pack_tmp_name,\n \t\t\t struct pack_idx_entry **written_list,\n \t\t\t uint32_t nr_written,\n+\t\t\t struct packing_data *to_pack,\n \t\t\t struct pack_idx_option *pack_idx_opts,\n \t\t\t unsigned char hash[],\n \t\t\t char **idx_tmp_name)\ndiff --git a/pack.h b/pack.h\nindex b22bfc4a18..fd27cfdfd7 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -109,11 +109,14 @@ int encode_in_pack_object_header(unsigned char *hdr, int hdr_len,\n #define PH_ERROR_PROTOCOL\t(-3)\n int read_pack_header(int fd, struct pack_header *);\n \n+struct packing_data;\n+\n struct hashfile *create_tmp_packfile(char **pack_tmp_name);\n void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t const char *pack_tmp_name,\n \t\t\t struct pack_idx_entry **written_list,\n \t\t\t uint32_t nr_written,\n+\t\t\t struct packing_data *to_pack,\n \t\t\t struct pack_idx_option *pack_idx_opts,\n \t\t\t unsigned char hash[],\n \t\t\t char **idx_tmp_name);\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455644","messageId":"2a6cfb00bf287380f44f324549db274a1b022499.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 06/17] t/helper: add 'pack-mtimes' test-tool","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:17:46Z","receivedAt":"2022-05-20T23:18:05Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In the next patch, we will implement and test support for writing a\ncruft pack via a special mode of `git pack-objects`. To make sure that\nobjects are written with the correct timestamps, and a new test-tool\nthat can dump the object names and corresponding timestamps from a given\n`.mtimes` file.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Makefile                    |  1 +\n t/helper/test-pack-mtimes.c | 56 +++++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c        |  1 +\n t/helper/test-tool.h        |  1 +\n 4 files changed, 59 insertions(+)\n create mode 100644 t/helper/test-pack-mtimes.c\n\ndiff --git a/Makefile b/Makefile\nindex a299580b7c..0b6eab0453 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -738,6 +738,7 @@ TEST_BUILTINS_OBJS += test-oid-array.o\n TEST_BUILTINS_OBJS += test-oidmap.o\n TEST_BUILTINS_OBJS += test-oidtree.o\n TEST_BUILTINS_OBJS += test-online-cpus.o\n+TEST_BUILTINS_OBJS += test-pack-mtimes.o\n TEST_BUILTINS_OBJS += test-parse-options.o\n TEST_BUILTINS_OBJS += test-parse-pathspec-file.o\n TEST_BUILTINS_OBJS += test-partial-clone.o\ndiff --git a/t/helper/test-pack-mtimes.c b/t/helper/test-pack-mtimes.c\nnew file mode 100644\nindex 0000000000..f7b79daf4c\n--- /dev/null\n+++ b/t/helper/test-pack-mtimes.c\n@@ -0,0 +1,56 @@\n+#include \"git-compat-util.h\"\n+#include \"test-tool.h\"\n+#include \"strbuf.h\"\n+#include \"object-store.h\"\n+#include \"packfile.h\"\n+#include \"pack-mtimes.h\"\n+\n+static void dump_mtimes(struct packed_git *p)\n+{\n+\tuint32_t i;\n+\tif (load_pack_mtimes(p) < 0)\n+\t\tdie(\"could not load pack .mtimes\");\n+\n+\tfor (i = 0; i < p->num_objects; i++) {\n+\t\tstruct object_id oid;\n+\t\tif (nth_packed_object_id(&oid, p, i) < 0)\n+\t\t\tdie(\"could not load object id at position %\"PRIu32, i);\n+\n+\t\tprintf(\"%s %\"PRIu32\"\\n\",\n+\t\t       oid_to_hex(&oid), nth_packed_mtime(p, i));\n+\t}\n+}\n+\n+static const char *pack_mtimes_usage = \"\\n\"\n+\"  test-tool pack-mtimes <pack-name.mtimes>\";\n+\n+int cmd__pack_mtimes(int argc, const char **argv)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct packed_git *p;\n+\n+\tsetup_git_directory();\n+\n+\tif (argc != 2)\n+\t\tusage(pack_mtimes_usage);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tstrbuf_addstr(&buf, basename(p->pack_name));\n+\t\tstrbuf_strip_suffix(&buf, \".pack\");\n+\t\tstrbuf_addstr(&buf, \".mtimes\");\n+\n+\t\tif (!strcmp(buf.buf, argv[1]))\n+\t\t\tbreak;\n+\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstrbuf_release(&buf);\n+\n+\tif (!p)\n+\t\tdie(\"could not find pack '%s'\", argv[1]);\n+\n+\tdump_mtimes(p);\n+\n+\treturn 0;\n+}\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex 0424f7adf5..d2eacd302d 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -48,6 +48,7 @@ static struct test_cmd cmds[] = {\n \t{ \"oidmap\", cmd__oidmap },\n \t{ \"oidtree\", cmd__oidtree },\n \t{ \"online-cpus\", cmd__online_cpus },\n+\t{ \"pack-mtimes\", cmd__pack_mtimes },\n \t{ \"parse-options\", cmd__parse_options },\n \t{ \"parse-pathspec-file\", cmd__parse_pathspec_file },\n \t{ \"partial-clone\", cmd__partial_clone },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex c876e8246f..960cc27ef7 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -38,6 +38,7 @@ int cmd__mktemp(int argc, const char **argv);\n int cmd__oidmap(int argc, const char **argv);\n int cmd__oidtree(int argc, const char **argv);\n int cmd__online_cpus(int argc, const char **argv);\n+int cmd__pack_mtimes(int argc, const char **argv);\n int cmd__parse_options(int argc, const char **argv);\n int cmd__parse_pathspec_file(int argc, const char** argv);\n int cmd__partial_clone(int argc, const char **argv);\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455645","messageId":"edb6fcd5ec72159f2b1fdef41077a841a89cd6a4.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 07/17] builtin/pack-objects.c: return from create_object_entry()","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:17:49Z","receivedAt":"2022-05-20T23:18:07Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"A new caller in the next commit will want to immediately modify the\nobject_entry structure created by create_object_entry(). Instead of\nforcing that caller to wastefully look-up the entry we just created,\nreturn it from create_object_entry() instead.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c | 16 +++++++++-------\n 1 file changed, 9 insertions(+), 7 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 6ac927047c..c6d16872ee 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1516,13 +1516,13 @@ static int want_object_in_pack(const struct object_id *oid,\n \treturn 1;\n }\n \n-static void create_object_entry(const struct object_id *oid,\n-\t\t\t\tenum object_type type,\n-\t\t\t\tuint32_t hash,\n-\t\t\t\tint exclude,\n-\t\t\t\tint no_try_delta,\n-\t\t\t\tstruct packed_git *found_pack,\n-\t\t\t\toff_t found_offset)\n+static struct object_entry *create_object_entry(const struct object_id *oid,\n+\t\t\t\t\t\tenum object_type type,\n+\t\t\t\t\t\tuint32_t hash,\n+\t\t\t\t\t\tint exclude,\n+\t\t\t\t\t\tint no_try_delta,\n+\t\t\t\t\t\tstruct packed_git *found_pack,\n+\t\t\t\t\t\toff_t found_offset)\n {\n \tstruct object_entry *entry;\n \n@@ -1539,6 +1539,8 @@ static void create_object_entry(const struct object_id *oid,\n \t}\n \n \tentry->no_try_delta = no_try_delta;\n+\n+\treturn entry;\n }\n \n static const char no_closure_warning[] = N_(\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455646","messageId":"91a9d21b0b7d99023083c0bbb6f91ccdc1782736.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:17:35Z","receivedAt":"2022-05-20T23:18:09Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"To store the individual mtimes of objects in a cruft pack, introduce a\nnew `.mtimes` format that can optionally accompany a single pack in the\nrepository.\n\nThe format is defined in Documentation/technical/pack-format.txt, and\nstores a 4-byte network order timestamp for each object in name (index)\norder.\n\nThis patch prepares for cruft packs by defining the `.mtimes` format,\nand introducing a basic API that callers can use to read out individual\nmtimes.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/technical/pack-format.txt |  19 ++++\n Makefile                                |   1 +\n builtin/repack.c                        |   1 +\n object-store.h                          |   5 +-\n pack-mtimes.c                           | 126 ++++++++++++++++++++++++\n pack-mtimes.h                           |  15 +++\n packfile.c                              |  19 +++-\n 7 files changed, 183 insertions(+), 3 deletions(-)\n create mode 100644 pack-mtimes.c\n create mode 100644 pack-mtimes.h\n\ndiff --git a/Documentation/technical/pack-format.txt b/Documentation/technical/pack-format.txt\nindex 6d3efb7d16..b520aa9c45 100644\n--- a/Documentation/technical/pack-format.txt\n+++ b/Documentation/technical/pack-format.txt\n@@ -294,6 +294,25 @@ Pack file entry: <+\n \n All 4-byte numbers are in network order.\n \n+== pack-*.mtimes files have the format:\n+\n+All 4-byte numbers are in network byte order.\n+\n+  - A 4-byte magic number '0x4d544d45' ('MTME').\n+\n+  - A 4-byte version identifier (= 1).\n+\n+  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n+\n+  - A table of 4-byte unsigned integers. The ith value is the\n+    modification time (mtime) of the ith object in the corresponding\n+    pack by lexicographic (index) order. The mtimes count standard\n+    epoch seconds.\n+\n+  - A trailer, containing a checksum of the corresponding packfile,\n+    and a checksum of all of the above (each having length according\n+    to the specified hash function).\n+\n == multi-pack-index (MIDX) files have the following format:\n \n The multi-pack-index files refer to multiple pack-files and loose objects.\ndiff --git a/Makefile b/Makefile\nindex 61aadf3ce8..a299580b7c 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -993,6 +993,7 @@ LIB_OBJS += oidtree.o\n LIB_OBJS += pack-bitmap-write.o\n LIB_OBJS += pack-bitmap.o\n LIB_OBJS += pack-check.o\n+LIB_OBJS += pack-mtimes.o\n LIB_OBJS += pack-objects.o\n LIB_OBJS += pack-revindex.o\n LIB_OBJS += pack-write.o\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex d1a563d5b6..e7a3920c6d 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -217,6 +217,7 @@ static struct {\n } exts[] = {\n \t{\".pack\"},\n \t{\".rev\", 1},\n+\t{\".mtimes\", 1},\n \t{\".bitmap\", 1},\n \t{\".promisor\", 1},\n \t{\".idx\"},\ndiff --git a/object-store.h b/object-store.h\nindex 53996018c1..2c4671ed7a 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -115,12 +115,15 @@ struct packed_git {\n \t\t freshened:1,\n \t\t do_not_close:1,\n \t\t pack_promisor:1,\n-\t\t multi_pack_index:1;\n+\t\t multi_pack_index:1,\n+\t\t is_cruft:1;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n \tstruct revindex_entry *revindex;\n \tconst uint32_t *revindex_data;\n \tconst uint32_t *revindex_map;\n \tsize_t revindex_size;\n+\tconst uint32_t *mtimes_map;\n+\tsize_t mtimes_size;\n \t/* something like \".git/objects/pack/xxxxx.pack\" */\n \tchar pack_name[FLEX_ARRAY]; /* more */\n };\ndiff --git a/pack-mtimes.c b/pack-mtimes.c\nnew file mode 100644\nindex 0000000000..46ad584af1\n--- /dev/null\n+++ b/pack-mtimes.c\n@@ -0,0 +1,126 @@\n+#include \"pack-mtimes.h\"\n+#include \"object-store.h\"\n+#include \"packfile.h\"\n+\n+static char *pack_mtimes_filename(struct packed_git *p)\n+{\n+\tsize_t len;\n+\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n+\t\tBUG(\"pack_name does not end in .pack\");\n+\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n+\treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n+}\n+\n+#define MTIMES_HEADER_SIZE (12)\n+#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n+\n+struct mtimes_header {\n+\tuint32_t signature;\n+\tuint32_t version;\n+\tuint32_t hash_id;\n+};\n+\n+static int load_pack_mtimes_file(char *mtimes_file,\n+\t\t\t\t uint32_t num_objects,\n+\t\t\t\t const uint32_t **data_p, size_t *len_p)\n+{\n+\tint fd, ret = 0;\n+\tstruct stat st;\n+\tvoid *data = NULL;\n+\tsize_t mtimes_size;\n+\tstruct mtimes_header header;\n+\tuint32_t *hdr;\n+\n+\tfd = git_open(mtimes_file);\n+\n+\tif (fd < 0) {\n+\t\tret = -1;\n+\t\tgoto cleanup;\n+\t}\n+\tif (fstat(fd, &st)) {\n+\t\tret = error_errno(_(\"failed to read %s\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tmtimes_size = xsize_t(st.st_size);\n+\n+\tif (mtimes_size < MTIMES_MIN_SIZE) {\n+\t\tret = error(_(\"mtimes file %s is too small\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t), num_objects)) {\n+\t\tret = error(_(\"mtimes file %s is corrupt\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tdata = hdr = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\n+\theader.signature = ntohl(hdr[0]);\n+\theader.version = ntohl(hdr[1]);\n+\theader.hash_id = ntohl(hdr[2]);\n+\n+\tif (header.signature != MTIMES_SIGNATURE) {\n+\t\tret = error(_(\"mtimes file %s has unknown signature\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (header.version != 1) {\n+\t\tret = error(_(\"mtimes file %s has unsupported version %\"PRIu32),\n+\t\t\t    mtimes_file, header.version);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (!(header.hash_id == 1 || header.hash_id == 2)) {\n+\t\tret = error(_(\"mtimes file %s has unsupported hash id %\"PRIu32),\n+\t\t\t    mtimes_file, header.hash_id);\n+\t\tgoto cleanup;\n+\t}\n+\n+cleanup:\n+\tif (ret) {\n+\t\tif (data)\n+\t\t\tmunmap(data, mtimes_size);\n+\t} else {\n+\t\t*len_p = mtimes_size;\n+\t\t*data_p = (const uint32_t *)data;\n+\t}\n+\n+\tclose(fd);\n+\treturn ret;\n+}\n+\n+int load_pack_mtimes(struct packed_git *p)\n+{\n+\tchar *mtimes_name = NULL;\n+\tint ret = 0;\n+\n+\tif (!p->is_cruft)\n+\t\treturn ret; /* not a cruft pack */\n+\tif (p->mtimes_map)\n+\t\treturn ret; /* already loaded */\n+\n+\tret = open_pack_index(p);\n+\tif (ret < 0)\n+\t\tgoto cleanup;\n+\n+\tmtimes_name = pack_mtimes_filename(p);\n+\tret = load_pack_mtimes_file(mtimes_name,\n+\t\t\t\t    p->num_objects,\n+\t\t\t\t    &p->mtimes_map,\n+\t\t\t\t    &p->mtimes_size);\n+cleanup:\n+\tfree(mtimes_name);\n+\treturn ret;\n+}\n+\n+uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos)\n+{\n+\tif (!p->mtimes_map)\n+\t\tBUG(\"pack .mtimes file not loaded for %s\", p->pack_name);\n+\tif (p->num_objects <= pos)\n+\t\tBUG(\"pack .mtimes out-of-bounds (%\"PRIu32\" vs %\"PRIu32\")\",\n+\t\t    pos, p->num_objects);\n+\n+\treturn get_be32(p->mtimes_map + pos + 3);\n+}\ndiff --git a/pack-mtimes.h b/pack-mtimes.h\nnew file mode 100644\nindex 0000000000..38ddb9f893\n--- /dev/null\n+++ b/pack-mtimes.h\n@@ -0,0 +1,15 @@\n+#ifndef PACK_MTIMES_H\n+#define PACK_MTIMES_H\n+\n+#include \"git-compat-util.h\"\n+\n+#define MTIMES_SIGNATURE 0x4d544d45 /* \"MTME\" */\n+#define MTIMES_VERSION 1\n+\n+struct packed_git;\n+\n+int load_pack_mtimes(struct packed_git *p);\n+\n+uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos);\n+\n+#endif\ndiff --git a/packfile.c b/packfile.c\nindex 835b2d2716..fc0245fbab 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -334,12 +334,22 @@ static void close_pack_revindex(struct packed_git *p)\n \tp->revindex_data = NULL;\n }\n \n+static void close_pack_mtimes(struct packed_git *p)\n+{\n+\tif (!p->mtimes_map)\n+\t\treturn;\n+\n+\tmunmap((void *)p->mtimes_map, p->mtimes_size);\n+\tp->mtimes_map = NULL;\n+}\n+\n void close_pack(struct packed_git *p)\n {\n \tclose_pack_windows(p);\n \tclose_pack_fd(p);\n \tclose_pack_index(p);\n \tclose_pack_revindex(p);\n+\tclose_pack_mtimes(p);\n \toidset_clear(&p->bad_objects);\n }\n \n@@ -363,7 +373,7 @@ void close_object_store(struct raw_object_store *o)\n \n void unlink_pack_path(const char *pack_name, int force_delete)\n {\n-\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\"};\n+\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\", \".mtimes\"};\n \tint i;\n \tstruct strbuf buf = STRBUF_INIT;\n \tsize_t plen;\n@@ -718,6 +728,10 @@ struct packed_git *add_packed_git(const char *path, size_t path_len, int local)\n \tif (!access(p->pack_name, F_OK))\n \t\tp->pack_promisor = 1;\n \n+\txsnprintf(p->pack_name + path_len, alloc - path_len, \".mtimes\");\n+\tif (!access(p->pack_name, F_OK))\n+\t\tp->is_cruft = 1;\n+\n \txsnprintf(p->pack_name + path_len, alloc - path_len, \".pack\");\n \tif (stat(p->pack_name, &st) || !S_ISREG(st.st_mode)) {\n \t\tfree(p);\n@@ -869,7 +883,8 @@ static void prepare_pack(const char *full_name, size_t full_name_len,\n \t    ends_with(file_name, \".pack\") ||\n \t    ends_with(file_name, \".bitmap\") ||\n \t    ends_with(file_name, \".keep\") ||\n-\t    ends_with(file_name, \".promisor\"))\n+\t    ends_with(file_name, \".promisor\") ||\n+\t    ends_with(file_name, \".mtimes\"))\n \t\tstring_list_append(data->garbage, full_name);\n \telse\n \t\treport_garbage(PACKDIR_FILE_GARBAGE, full_name);\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455647","messageId":"e3185741f21a9271aa3f0eb191b30d84d3d2e975.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 08/17] builtin/pack-objects.c: --cruft without expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:17:52Z","receivedAt":"2022-05-20T23:18:11Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Teach `pack-objects` how to generate a cruft pack when no objects are\ndropped (i.e., `--cruft-expiration=never`). Later patches will teach\n`pack-objects` how to generate a cruft pack that prunes objects.\n\nWhen generating a cruft pack which does not prune objects, we want to\ncollect all unreachable objects into a single pack (noting and updating\ntheir mtimes as we accumulate them). Ordinary use will pass the result\nof a `git repack -A` as a kept pack, so when this patch says \"kept\npack\", readers should think \"reachable objects\".\n\nGenerating a non-expiring cruft packs works as follows:\n\n  - Callers provide a list of every pack they know about, and indicate\n    which packs are about to be removed.\n\n  - All packs which are going to be removed (we'll call these the\n    redundant ones) are marked as kept in-core.\n\n    Any packs the caller did not mention (but are known to the\n    `pack-objects` process) are also marked as kept in-core. Packs not\n    mentioned by the caller are assumed to be unknown to them, i.e.,\n    they entered the repository after the caller decided which packs\n    should be kept and which should be discarded.\n\n    Since we do not want to include objects in these \"unknown\" packs\n    (because we don't know which of their objects are or aren't\n    reachable), these are also marked as kept in-core.\n\n  - Then, we enumerate all objects in the repository, and add them to\n    our packing list if they do not appear in an in-core kept pack.\n\nThis results in a new cruft pack which contains all known objects that\naren't included in the kept packs. When the kept pack is the result of\n`git repack -A`, the resulting pack contains all unreachable objects.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-pack-objects.txt |  30 ++++\n builtin/pack-objects.c             | 201 +++++++++++++++++++++++++-\n object-file.c                      |   2 +-\n object-store.h                     |   2 +\n t/t5329-pack-objects-cruft.sh      | 218 +++++++++++++++++++++++++++++\n 5 files changed, 448 insertions(+), 5 deletions(-)\n create mode 100755 t/t5329-pack-objects-cruft.sh\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex f8344e1e5b..a9995a932c 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -13,6 +13,7 @@ SYNOPSIS\n \t[--no-reuse-delta] [--delta-base-offset] [--non-empty]\n \t[--local] [--incremental] [--window=<n>] [--depth=<n>]\n \t[--revs [--unpacked | --all]] [--keep-pack=<pack-name>]\n+\t[--cruft] [--cruft-expiration=<time>]\n \t[--stdout [--filter=<filter-spec>] | <base-name>]\n \t[--shallow] [--keep-true-parents] [--[no-]sparse] < <object-list>\n \n@@ -95,6 +96,35 @@ base-name::\n Incompatible with `--revs`, or options that imply `--revs` (such as\n `--all`), with the exception of `--unpacked`, which is compatible.\n \n+--cruft::\n+\tPacks unreachable objects into a separate \"cruft\" pack, denoted\n+\tby the existence of a `.mtimes` file. Typically used by `git\n+\trepack --cruft`. Callers provide a list of pack names and\n+\tindicate which packs will remain in the repository, along with\n+\twhich packs will be deleted (indicated by the `-` prefix). The\n+\tcontents of the cruft pack are all objects not contained in the\n+\tsurviving packs which have not exceeded the grace period (see\n+\t`--cruft-expiration` below), or which have exceeded the grace\n+\tperiod, but are reachable from an other object which hasn't.\n++\n+When the input lists a pack containing all reachable objects (and lists\n+all other packs as pending deletion), the corresponding cruft pack will\n+contain all unreachable objects (with mtime newer than the\n+`--cruft-expiration`) along with any unreachable objects whose mtime is\n+older than the `--cruft-expiration`, but are reachable from an\n+unreachable object whose mtime is newer than the `--cruft-expiration`).\n++\n+Incompatible with `--unpack-unreachable`, `--keep-unreachable`,\n+`--pack-loose-unreachable`, `--stdin-packs`, as well as any other\n+options which imply `--revs`. Also incompatible with `--max-pack-size`;\n+when this option is set, the maximum pack size is not inferred from\n+`pack.packSizeLimit`.\n+\n+--cruft-expiration=<approxidate>::\n+\tIf specified, objects are eliminated from the cruft pack if they\n+\thave an mtime older than `<approxidate>`. If unspecified (and\n+\tgiven `--cruft`), then no objects are eliminated.\n+\n --window=<n>::\n --depth=<n>::\n \tThese two options affect how the objects contained in\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex c6d16872ee..9cf89be673 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -36,6 +36,7 @@\n #include \"trace2.h\"\n #include \"shallow.h\"\n #include \"promisor-remote.h\"\n+#include \"pack-mtimes.h\"\n \n /*\n  * Objects we are going to pack are collected in the `to_pack` structure.\n@@ -194,6 +195,8 @@ static int reuse_delta = 1, reuse_object = 1;\n static int keep_unreachable, unpack_unreachable, include_tag;\n static timestamp_t unpack_unreachable_expiration;\n static int pack_loose_unreachable;\n+static int cruft;\n+static timestamp_t cruft_expiration;\n static int local;\n static int have_non_local_packs;\n static int incremental;\n@@ -1260,6 +1263,9 @@ static void write_pack_file(void)\n \t\t\t\t\t&to_pack, written_list, nr_written);\n \t\t\t}\n \n+\t\t\tif (cruft)\n+\t\t\t\tpack_idx_opts.flags |= WRITE_MTIMES;\n+\n \t\t\tstage_tmp_packfiles(&tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n \t\t\t\t\t    &to_pack, &pack_idx_opts, hash,\n@@ -3397,6 +3403,135 @@ static void read_packs_list_from_stdin(void)\n \tstring_list_clear(&exclude_packs, 0);\n }\n \n+static void add_cruft_object_entry(const struct object_id *oid, enum object_type type,\n+\t\t\t\t   struct packed_git *pack, off_t offset,\n+\t\t\t\t   const char *name, uint32_t mtime)\n+{\n+\tstruct object_entry *entry;\n+\n+\tdisplay_progress(progress_state, ++nr_seen);\n+\n+\tentry = packlist_find(&to_pack, oid);\n+\tif (entry) {\n+\t\tif (name) {\n+\t\t\tentry->hash = pack_name_hash(name);\n+\t\t\tentry->no_try_delta = no_try_delta(name);\n+\t\t}\n+\t} else {\n+\t\tif (!want_object_in_pack(oid, 0, &pack, &offset))\n+\t\t\treturn;\n+\t\tif (!pack && type == OBJ_BLOB && !has_loose_object(oid)) {\n+\t\t\t/*\n+\t\t\t * If a traversed tree has a missing blob then we want\n+\t\t\t * to avoid adding that missing object to our pack.\n+\t\t\t *\n+\t\t\t * This only applies to missing blobs, not trees,\n+\t\t\t * because the traversal needs to parse sub-trees but\n+\t\t\t * not blobs.\n+\t\t\t *\n+\t\t\t * Note we only perform this check when we couldn't\n+\t\t\t * already find the object in a pack, so we're really\n+\t\t\t * limited to \"ensure non-tip blobs which don't exist in\n+\t\t\t * packs do exist via loose objects\". Confused?\n+\t\t\t */\n+\t\t\treturn;\n+\t\t}\n+\n+\t\tentry = create_object_entry(oid, type, pack_name_hash(name),\n+\t\t\t\t\t    0, name && no_try_delta(name),\n+\t\t\t\t\t    pack, offset);\n+\t}\n+\n+\tif (mtime > oe_cruft_mtime(&to_pack, entry))\n+\t\toe_set_cruft_mtime(&to_pack, entry, mtime);\n+\treturn;\n+}\n+\n+static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n+{\n+\tstruct string_list_item *item = NULL;\n+\tfor_each_string_list_item(item, packs) {\n+\t\tstruct packed_git *p = item->util;\n+\t\tif (!p)\n+\t\t\tdie(_(\"could not find pack '%s'\"), item->string);\n+\t\tp->pack_keep_in_core = keep;\n+\t}\n+}\n+\n+static void add_unreachable_loose_objects(void);\n+static void add_objects_in_unpacked_packs(void);\n+\n+static void enumerate_cruft_objects(void)\n+{\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Enumerating cruft objects\"), 0);\n+\n+\tadd_objects_in_unpacked_packs();\n+\tadd_unreachable_loose_objects();\n+\n+\tstop_progress(&progress_state);\n+}\n+\n+static void read_cruft_objects(void)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct string_list discard_packs = STRING_LIST_INIT_DUP;\n+\tstruct string_list fresh_packs = STRING_LIST_INIT_DUP;\n+\tstruct packed_git *p;\n+\n+\tignore_packed_keep_in_core = 1;\n+\n+\twhile (strbuf_getline(&buf, stdin) != EOF) {\n+\t\tif (!buf.len)\n+\t\t\tcontinue;\n+\n+\t\tif (*buf.buf == '-')\n+\t\t\tstring_list_append(&discard_packs, buf.buf + 1);\n+\t\telse\n+\t\t\tstring_list_append(&fresh_packs, buf.buf);\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\n+\tstring_list_sort(&discard_packs);\n+\tstring_list_sort(&fresh_packs);\n+\n+\tfor (p = get_all_packs(the_repository); p; p = p->next) {\n+\t\tconst char *pack_name = pack_basename(p);\n+\t\tstruct string_list_item *item;\n+\n+\t\titem = string_list_lookup(&fresh_packs, pack_name);\n+\t\tif (!item)\n+\t\t\titem = string_list_lookup(&discard_packs, pack_name);\n+\n+\t\tif (item) {\n+\t\t\titem->util = p;\n+\t\t} else {\n+\t\t\t/*\n+\t\t\t * This pack wasn't mentioned in either the \"fresh\" or\n+\t\t\t * \"discard\" list, so the caller didn't know about it.\n+\t\t\t *\n+\t\t\t * Mark it as kept so that its objects are ignored by\n+\t\t\t * add_unseen_recent_objects_to_traversal(). We'll\n+\t\t\t * unmark it before starting the traversal so it doesn't\n+\t\t\t * halt the traversal early.\n+\t\t\t */\n+\t\t\tp->pack_keep_in_core = 1;\n+\t\t}\n+\t}\n+\n+\tmark_pack_kept_in_core(&fresh_packs, 1);\n+\tmark_pack_kept_in_core(&discard_packs, 0);\n+\n+\tif (cruft_expiration)\n+\t\tdie(\"--cruft-expiration not yet implemented\");\n+\telse\n+\t\tenumerate_cruft_objects();\n+\n+\tstrbuf_release(&buf);\n+\tstring_list_clear(&discard_packs, 0);\n+\tstring_list_clear(&fresh_packs, 0);\n+}\n+\n static void read_object_list_from_stdin(void)\n {\n \tchar line[GIT_MAX_HEXSZ + 1 + PATH_MAX + 2];\n@@ -3529,7 +3664,24 @@ static int add_object_in_unpacked_pack(const struct object_id *oid,\n \t\t\t\t       uint32_t pos,\n \t\t\t\t       void *_data)\n {\n-\tadd_object_entry(oid, OBJ_NONE, \"\", 0);\n+\tif (cruft) {\n+\t\toff_t offset;\n+\t\ttime_t mtime;\n+\n+\t\tif (pack->is_cruft) {\n+\t\t\tif (load_pack_mtimes(pack) < 0)\n+\t\t\t\tdie(_(\"could not load cruft pack .mtimes\"));\n+\t\t\tmtime = nth_packed_mtime(pack, pos);\n+\t\t} else {\n+\t\t\tmtime = pack->mtime;\n+\t\t}\n+\t\toffset = nth_packed_object_offset(pack, pos);\n+\n+\t\tadd_cruft_object_entry(oid, OBJ_NONE, pack, offset,\n+\t\t\t\t       NULL, mtime);\n+\t} else {\n+\t\tadd_object_entry(oid, OBJ_NONE, \"\", 0);\n+\t}\n \treturn 0;\n }\n \n@@ -3553,7 +3705,19 @@ static int add_loose_object(const struct object_id *oid, const char *path,\n \t\treturn 0;\n \t}\n \n-\tadd_object_entry(oid, type, \"\", 0);\n+\tif (cruft) {\n+\t\tstruct stat st;\n+\t\tif (stat(path, &st) < 0) {\n+\t\t\tif (errno == ENOENT)\n+\t\t\t\treturn 0;\n+\t\t\treturn error_errno(\"unable to stat %s\", oid_to_hex(oid));\n+\t\t}\n+\n+\t\tadd_cruft_object_entry(oid, type, NULL, 0, NULL,\n+\t\t\t\t       st.st_mtime);\n+\t} else {\n+\t\tadd_object_entry(oid, type, \"\", 0);\n+\t}\n \treturn 0;\n }\n \n@@ -3870,6 +4034,20 @@ static int option_parse_unpack_unreachable(const struct option *opt,\n \treturn 0;\n }\n \n+static int option_parse_cruft_expiration(const struct option *opt,\n+\t\t\t\t\t const char *arg, int unset)\n+{\n+\tif (unset) {\n+\t\tcruft = 0;\n+\t\tcruft_expiration = 0;\n+\t} else {\n+\t\tcruft = 1;\n+\t\tif (arg)\n+\t\t\tcruft_expiration = approxidate(arg);\n+\t}\n+\treturn 0;\n+}\n+\n struct po_filter_data {\n \tunsigned have_revs:1;\n \tstruct rev_info revs;\n@@ -3959,6 +4137,10 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tOPT_CALLBACK_F(0, \"unpack-unreachable\", NULL, N_(\"time\"),\n \t\t  N_(\"unpack unreachable objects newer than <time>\"),\n \t\t  PARSE_OPT_OPTARG, option_parse_unpack_unreachable),\n+\t\tOPT_BOOL(0, \"cruft\", &cruft, N_(\"create a cruft pack\")),\n+\t\tOPT_CALLBACK_F(0, \"cruft-expiration\", NULL, N_(\"time\"),\n+\t\t  N_(\"expire cruft objects older than <time>\"),\n+\t\t  PARSE_OPT_OPTARG, option_parse_cruft_expiration),\n \t\tOPT_BOOL(0, \"sparse\", &sparse,\n \t\t\t N_(\"use the sparse reachability algorithm\")),\n \t\tOPT_BOOL(0, \"thin\", &thin,\n@@ -4085,7 +4267,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \n \tif (!HAVE_THREADS && delta_search_threads != 1)\n \t\twarning(_(\"no threads support, ignoring --threads\"));\n-\tif (!pack_to_stdout && !pack_size_limit)\n+\tif (!pack_to_stdout && !pack_size_limit && !cruft)\n \t\tpack_size_limit = pack_size_limit_cfg;\n \tif (pack_to_stdout && pack_size_limit)\n \t\tdie(_(\"--max-pack-size cannot be used to build a pack for transfer\"));\n@@ -4112,6 +4294,15 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (stdin_packs && use_internal_rev_list)\n \t\tdie(_(\"cannot use internal rev list with --stdin-packs\"));\n \n+\tif (cruft) {\n+\t\tif (use_internal_rev_list)\n+\t\t\tdie(_(\"cannot use internal rev list with --cruft\"));\n+\t\tif (stdin_packs)\n+\t\t\tdie(_(\"cannot use --stdin-packs with --cruft\"));\n+\t\tif (pack_size_limit)\n+\t\t\tdie(_(\"cannot use --max-pack-size with --cruft\"));\n+\t}\n+\n \t/*\n \t * \"soft\" reasons not to use bitmaps - for on-disk repack by default we want\n \t *\n@@ -4168,7 +4359,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t    the_repository);\n \tprepare_packing_data(the_repository, &to_pack);\n \n-\tif (progress)\n+\tif (progress && !cruft)\n \t\tprogress_state = start_progress(_(\"Enumerating objects\"), 0);\n \tif (stdin_packs) {\n \t\t/* avoids adding objects in excluded packs */\n@@ -4176,6 +4367,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tread_packs_list_from_stdin();\n \t\tif (rev_list_unpacked)\n \t\t\tadd_unreachable_loose_objects();\n+\t} else if (cruft) {\n+\t\tread_cruft_objects();\n \t} else if (!use_internal_rev_list) {\n \t\tread_object_list_from_stdin();\n \t} else if (pfd.have_revs) {\ndiff --git a/object-file.c b/object-file.c\nindex 5ffbf3d4fd..ff0cffe68e 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -997,7 +997,7 @@ int has_loose_object_nonlocal(const struct object_id *oid)\n \treturn check_and_freshen_nonlocal(oid, 0);\n }\n \n-static int has_loose_object(const struct object_id *oid)\n+int has_loose_object(const struct object_id *oid)\n {\n \treturn check_and_freshen(oid, 0);\n }\ndiff --git a/object-store.h b/object-store.h\nindex 2c4671ed7a..c41609e8db 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -330,6 +330,8 @@ int repo_has_object_file_with_flags(struct repository *r,\n  */\n int has_loose_object_nonlocal(const struct object_id *);\n \n+int has_loose_object(const struct object_id *);\n+\n /**\n  * format_object_header() is a thin wrapper around s xsnprintf() that\n  * writes the initial \"<type> <obj-len>\" part of the loose object\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nnew file mode 100755\nindex 0000000000..003ca7344e\n--- /dev/null\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -0,0 +1,218 @@\n+#!/bin/sh\n+\n+test_description='cruft pack related pack-objects tests'\n+. ./test-lib.sh\n+\n+objdir=.git/objects\n+packdir=$objdir/pack\n+\n+basic_cruft_pack_tests () {\n+\texpire=\"$1\"\n+\n+\ttest_expect_success \"unreachable loose objects are packed (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit base &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit loose &&\n+\n+\t\t\ttest-tool chmtime +2000 \"$objdir/$(test_oid_to_path \\\n+\t\t\t\t$(git rev-parse loose:loose.t))\" &&\n+\t\t\ttest-tool chmtime +1000 \"$objdir/$(test_oid_to_path \\\n+\t\t\t\t$(git rev-parse loose^{tree}))\" &&\n+\n+\t\t\t(\n+\t\t\t\tgit rev-list --objects --no-object-names base..loose |\n+\t\t\t\twhile read oid\n+\t\t\t\tdo\n+\t\t\t\t\tpath=\"$objdir/$(test_oid_to_path \"$oid\")\" &&\n+\t\t\t\t\tprintf \"%s %d\\n\" \"$oid\" \"$(test-tool chmtime --get \"$path\")\"\n+\t\t\t\tdone |\n+\t\t\t\tsort -k1\n+\t\t\t) >expect &&\n+\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t\t--cruft-expiration=\"$expire\" $packdir/pack)\" &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"unreachable packed objects are packed (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit packed &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit other &&\n+\n+\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\t\tother=\"$(git pack-objects --delta-base-offset \\\n+\t\t\t\t$packdir/pack <objects)\" &&\n+\t\t\tgit prune-packed &&\n+\n+\t\t\ttest-tool chmtime --get -100 \"$packdir/pack-$other.pack\" >expect &&\n+\n+\t\t\tcruft=\"$(git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-pack-$other.pack\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\t\tcut -d\" \" -f2 <actual.raw | sort -u >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"unreachable cruft objects are repacked (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit packed &&\n+\t\t\tgit repack -Ad &&\n+\t\t\ttest_commit other &&\n+\n+\t\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\n+\t\t\tcruft_a=\"$(echo $keep | git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack)\" &&\n+\t\t\tgit prune-packed &&\n+\t\t\tcruft_b=\"$(git pack-objects --cruft --cruft-expiration=\"$expire\" $packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-pack-$cruft_a.pack\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft_a.mtimes\" >expect.raw &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft_b.mtimes\" >actual.raw &&\n+\n+\t\t\tsort <expect.raw >expect &&\n+\t\t\tsort <actual.raw >actual &&\n+\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"multiple cruft packs (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\tgit repack -Ad &&\n+\t\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\n+\t\t\ttest_commit cruft &&\n+\t\t\tloose=\"$objdir/$(test_oid_to_path $(git rev-parse cruft))\" &&\n+\n+\t\t\t# generate three copies of the cruft object in different\n+\t\t\t# cruft packs, each with a unique mtime:\n+\t\t\t#   - one expired (1000 seconds ago)\n+\t\t\t#   - two non-expired (one 1000 seconds in the future,\n+\t\t\t#     one 1500 seconds in the future)\n+\t\t\ttest-tool chmtime =-1000 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-A <<-EOF &&\n+\t\t\t$keep\n+\t\t\tEOF\n+\t\t\ttest-tool chmtime =+1000 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-B <<-EOF &&\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\tEOF\n+\t\t\ttest-tool chmtime =+1500 \"$loose\" &&\n+\t\t\tgit pack-objects --cruft $packdir/pack-C <<-EOF &&\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-B-*.pack))\n+\t\t\tEOF\n+\n+\t\t\t# ensure the resulting cruft pack takes the most recent\n+\t\t\t# mtime among all copies\n+\t\t\tcruft=\"$(git pack-objects --cruft \\\n+\t\t\t\t--cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <<-EOF\n+\t\t\t$keep\n+\t\t\t-$(basename $(ls $packdir/pack-A-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-B-*.pack))\n+\t\t\t-$(basename $(ls $packdir/pack-C-*.pack))\n+\t\t\tEOF\n+\t\t\t)\" &&\n+\n+\t\t\ttest-tool pack-mtimes \"$(basename $(ls $packdir/pack-C-*.mtimes))\" >expect.raw &&\n+\t\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\t\tsort expect.raw >expect &&\n+\t\t\tsort actual.raw >actual &&\n+\t\t\ttest_cmp expect actual\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"cruft packs tolerate missing trees (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\ttest_commit cruft &&\n+\n+\t\t\ttree=\"$(git rev-parse cruft^{tree})\" &&\n+\n+\t\t\tgit reset --hard reachable &&\n+\t\t\tgit tag -d cruft &&\n+\t\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t\t# remove the unreachable tree, but leave the commit\n+\t\t\t# which has it as its root tree intact\n+\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$tree\")\" &&\n+\n+\t\t\tgit repack -Ad &&\n+\t\t\tbasename $(ls $packdir/pack-*.pack) >in &&\n+\t\t\tgit pack-objects --cruft --cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <in\n+\t\t)\n+\t'\n+\n+\ttest_expect_success \"cruft packs tolerate missing blobs (expire $expire)\" '\n+\t\tgit init repo &&\n+\t\ttest_when_finished \"rm -fr repo\" &&\n+\t\t(\n+\t\t\tcd repo &&\n+\n+\t\t\ttest_commit reachable &&\n+\t\t\ttest_commit cruft &&\n+\n+\t\t\tblob=\"$(git rev-parse cruft:cruft.t)\" &&\n+\n+\t\t\tgit reset --hard reachable &&\n+\t\t\tgit tag -d cruft &&\n+\t\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t\t# remove the unreachable blob, but leave the commit (and\n+\t\t\t# the root tree of that commit) intact\n+\t\t\trm -fr \"$objdir/$(test_oid_to_path \"$blob\")\" &&\n+\n+\t\t\tgit repack -Ad &&\n+\t\t\tbasename $(ls $packdir/pack-*.pack) >in &&\n+\t\t\tgit pack-objects --cruft --cruft-expiration=\"$expire\" \\\n+\t\t\t\t$packdir/pack <in\n+\t\t)\n+\t'\n+}\n+\n+basic_cruft_pack_tests never\n+\n+test_done\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455648","messageId":"d66be44d9a08bb97761b8eb23861caec638025e3.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 10/17] reachable: report precise timestamps from objects in cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:17:57Z","receivedAt":"2022-05-20T23:18:15Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When generating a cruft pack, the caller within pack-objects will want\nto know the precise timestamps of cruft objects (i.e., their\ncorresponding values in the .mtimes table) rather than the mtime of the\ncruft pack itself.\n\nTeach add_recent_packed() to lookup each object's precise mtime from the\n.mtimes file if one exists (indicated by the is_cruft bit on the\npacked_git structure).\n\nA couple of small things worth noting here:\n\n  - load_pack_mtimes() needs to be called before asking for\n    nth_packed_mtime(), and that call is done lazily here. That function\n    exits early if the .mtimes file has already been opened and parsed,\n    so only the first call is slow.\n\n  - Checking the is_cruft bit can be done without any extra work on the\n    caller's behalf, since it is set up for us automatically as a\n    side-effect of calling add_packed_git() (just like the 'pack_keep'\n    and 'pack_promisor' bits).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n reachable.c | 9 ++++++++-\n 1 file changed, 8 insertions(+), 1 deletion(-)\n\ndiff --git a/reachable.c b/reachable.c\nindex d4507c4270..aba63ebeb3 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -13,6 +13,7 @@\n #include \"worktree.h\"\n #include \"object-store.h\"\n #include \"pack-bitmap.h\"\n+#include \"pack-mtimes.h\"\n \n struct connectivity_progress {\n \tstruct progress *progress;\n@@ -155,6 +156,7 @@ static int add_recent_packed(const struct object_id *oid,\n \t\t\t     void *data)\n {\n \tstruct object *obj;\n+\ttimestamp_t mtime = p->mtime;\n \n \tif (!want_recent_object(data, oid))\n \t\treturn 0;\n@@ -163,7 +165,12 @@ static int add_recent_packed(const struct object_id *oid,\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n-\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), p->mtime, data);\n+\tif (p->is_cruft) {\n+\t\tif (load_pack_mtimes(p) < 0)\n+\t\t\tdie(_(\"could not load cruft pack .mtimes\"));\n+\t\tmtime = nth_packed_mtime(p, pos);\n+\t}\n+\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), mtime, data);\n \treturn 0;\n }\n \n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455649","messageId":"1cf00d462cc94775825a758558a5428e30919fee.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 09/17] reachable: add options to add_unseen_recent_objects_to_traversal","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:17:54Z","receivedAt":"2022-05-20T23:18:17Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"This function behaves very similarly to what we will need in\npack-objects in order to implement cruft packs with expiration. But it\nis lacking a couple of things. Namely, it needs:\n\n  - a mechanism to communicate the timestamps of individual recent\n    objects to some external caller\n\n  - and, in the case of packed objects, our future caller will also want\n    to know the originating pack, as well as the offset within that pack\n    at which the object can be found\n\n  - finally, it needs a way to skip over packs which are marked as kept\n    in-core.\n\nTo address the first two, add a callback interface in this patch which\nreports the time of each recent object, as well as a (packed_git,\noff_t) pair for packed objects.\n\nLikewise, add a new option to the packed object iterators to skip over\npacks which are marked as kept in core. This option will become\nimplicitly tested in a future patch.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c |  2 +-\n reachable.c            | 51 +++++++++++++++++++++++++++++++++++-------\n reachable.h            |  9 +++++++-\n 3 files changed, 52 insertions(+), 10 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 9cf89be673..3b8bf6a3dd 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -3957,7 +3957,7 @@ static void get_object_list(struct rev_info *revs, int ac, const char **av)\n \tif (unpack_unreachable_expiration) {\n \t\trevs->ignore_missing_links = 1;\n \t\tif (add_unseen_recent_objects_to_traversal(revs,\n-\t\t\t\tunpack_unreachable_expiration))\n+\t\t\t\tunpack_unreachable_expiration, NULL, 0))\n \t\t\tdie(_(\"unable to add recent objects\"));\n \t\tif (prepare_revision_walk(revs))\n \t\t\tdie(_(\"revision walk setup failed\"));\ndiff --git a/reachable.c b/reachable.c\nindex b9f4ad886e..d4507c4270 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -60,9 +60,13 @@ static void mark_commit(struct commit *c, void *data)\n struct recent_data {\n \tstruct rev_info *revs;\n \ttimestamp_t timestamp;\n+\treport_recent_object_fn *cb;\n+\tint ignore_in_core_kept_packs;\n };\n \n static void add_recent_object(const struct object_id *oid,\n+\t\t\t      struct packed_git *pack,\n+\t\t\t      off_t offset,\n \t\t\t      timestamp_t mtime,\n \t\t\t      struct recent_data *data)\n {\n@@ -103,13 +107,29 @@ static void add_recent_object(const struct object_id *oid,\n \t\tdie(\"unable to lookup %s\", oid_to_hex(oid));\n \n \tadd_pending_object(data->revs, obj, \"\");\n+\tif (data->cb)\n+\t\tdata->cb(obj, pack, offset, mtime);\n+}\n+\n+static int want_recent_object(struct recent_data *data,\n+\t\t\t      const struct object_id *oid)\n+{\n+\tif (data->ignore_in_core_kept_packs &&\n+\t    has_object_kept_pack(oid, IN_CORE_KEEP_PACKS))\n+\t\treturn 0;\n+\treturn 1;\n }\n \n static int add_recent_loose(const struct object_id *oid,\n \t\t\t    const char *path, void *data)\n {\n \tstruct stat st;\n-\tstruct object *obj = lookup_object(the_repository, oid);\n+\tstruct object *obj;\n+\n+\tif (!want_recent_object(data, oid))\n+\t\treturn 0;\n+\n+\tobj = lookup_object(the_repository, oid);\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n@@ -126,7 +146,7 @@ static int add_recent_loose(const struct object_id *oid,\n \t\treturn error_errno(\"unable to stat %s\", oid_to_hex(oid));\n \t}\n \n-\tadd_recent_object(oid, st.st_mtime, data);\n+\tadd_recent_object(oid, NULL, 0, st.st_mtime, data);\n \treturn 0;\n }\n \n@@ -134,29 +154,43 @@ static int add_recent_packed(const struct object_id *oid,\n \t\t\t     struct packed_git *p, uint32_t pos,\n \t\t\t     void *data)\n {\n-\tstruct object *obj = lookup_object(the_repository, oid);\n+\tstruct object *obj;\n+\n+\tif (!want_recent_object(data, oid))\n+\t\treturn 0;\n+\n+\tobj = lookup_object(the_repository, oid);\n \n \tif (obj && obj->flags & SEEN)\n \t\treturn 0;\n-\tadd_recent_object(oid, p->mtime, data);\n+\tadd_recent_object(oid, p, nth_packed_object_offset(p, pos), p->mtime, data);\n \treturn 0;\n }\n \n int add_unseen_recent_objects_to_traversal(struct rev_info *revs,\n-\t\t\t\t\t   timestamp_t timestamp)\n+\t\t\t\t\t   timestamp_t timestamp,\n+\t\t\t\t\t   report_recent_object_fn *cb,\n+\t\t\t\t\t   int ignore_in_core_kept_packs)\n {\n \tstruct recent_data data;\n+\tenum for_each_object_flags flags;\n \tint r;\n \n \tdata.revs = revs;\n \tdata.timestamp = timestamp;\n+\tdata.cb = cb;\n+\tdata.ignore_in_core_kept_packs = ignore_in_core_kept_packs;\n \n \tr = for_each_loose_object(add_recent_loose, &data,\n \t\t\t\t  FOR_EACH_OBJECT_LOCAL_ONLY);\n \tif (r)\n \t\treturn r;\n-\treturn for_each_packed_object(add_recent_packed, &data,\n-\t\t\t\t      FOR_EACH_OBJECT_LOCAL_ONLY);\n+\n+\tflags = FOR_EACH_OBJECT_LOCAL_ONLY | FOR_EACH_OBJECT_PACK_ORDER;\n+\tif (ignore_in_core_kept_packs)\n+\t\tflags |= FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS;\n+\n+\treturn for_each_packed_object(add_recent_packed, &data, flags);\n }\n \n static int mark_object_seen(const struct object_id *oid,\n@@ -217,7 +251,8 @@ void mark_reachable_objects(struct rev_info *revs, int mark_reflog,\n \n \tif (mark_recent) {\n \t\trevs->ignore_missing_links = 1;\n-\t\tif (add_unseen_recent_objects_to_traversal(revs, mark_recent))\n+\t\tif (add_unseen_recent_objects_to_traversal(revs, mark_recent,\n+\t\t\t\t\t\t\t   NULL, 0))\n \t\t\tdie(\"unable to mark recent objects\");\n \t\tif (prepare_revision_walk(revs))\n \t\t\tdie(\"revision walk setup failed\");\ndiff --git a/reachable.h b/reachable.h\nindex 5df932ad8f..b776761baa 100644\n--- a/reachable.h\n+++ b/reachable.h\n@@ -1,11 +1,18 @@\n #ifndef REACHEABLE_H\n #define REACHEABLE_H\n \n+#include \"object.h\"\n+\n struct progress;\n struct rev_info;\n \n+typedef void report_recent_object_fn(const struct object *, struct packed_git *,\n+\t\t\t\t     off_t, time_t);\n+\n int add_unseen_recent_objects_to_traversal(struct rev_info *revs,\n-\t\t\t\t\t   timestamp_t timestamp);\n+\t\t\t\t\t   timestamp_t timestamp,\n+\t\t\t\t\t   report_recent_object_fn cb,\n+\t\t\t\t\t   int ignore_in_core_kept_packs);\n void mark_reachable_objects(struct rev_info *revs, int mark_reflog,\n \t\t\t    timestamp_t mark_recent, struct progress *);\n \n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455650","messageId":"1434e3762389a6f5cd4236aada6d3ab1afad5681.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 11/17] builtin/pack-objects.c: --cruft with expiration","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:18:00Z","receivedAt":"2022-05-20T23:18:25Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a previous patch, pack-objects learned how to generate a cruft pack\nso long as no objects are dropped.\n\nThis patch teaches pack-objects to handle the case where a non-never\n`--cruft-expiration` value is passed. This case is slightly more\ncomplicated than before, because we want pack-objects to save\nunreachable objects which would have been pruned when there is another\nrecent (i.e., non-prunable) unreachable object which reaches the other.\nWe'll call these objects \"unreachable but reachable-from-recent\".\n\nHere is how pack-objects handles `--cruft-expiration`:\n\n  - Instead of adding all objects outside of the kept pack(s) into the\n    packing list, only handle the ones whose mtime is within the grace\n    period.\n\n  - Construct a reachability traversal whose tips are the\n    unreachable-but-recent objects.\n\n  - Then, walk along that traversal, stopping if we reach an object in\n    the kept pack. At each step along the traversal, we add the object\n    we are visiting to the packing list.\n\nIn the majority of these cases, any object we visit in this traversal\nwill already be in our packing list. But we will sometimes encounter\nreachable-from-recent cruft objects, which we want to retain even if\nthey aged out of the grace period.\n\nThe most subtle point of this process is that we actually don't need to\nbother to update the rescued object's mtime. Even though we will write\nan .mtimes file with a value that is older than the expiration window,\nit will continue to survive cruft repacks so long as any objects which\nreach it haven't aged out.\n\nThat is, a future repack will also exclude that object from the initial\npacking list, only to discover it later on when doing the reachability\ntraversal.\n\nFinally, stopping early once an object is found in a kept pack is safe\nto do because the kept packs ordinarily represent which packs will\nsurvive after repacking. Assuming that it _isn't_ safe to halt a\ntraversal early would mean that there is some ancestor object which is\nmissing, which implies repository corruption (i.e., the complete set of\nreachable objects isn't present).\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/pack-objects.c        |  84 +++++++++++++++++++-\n reachable.h                   |   4 +-\n t/t5329-pack-objects-cruft.sh | 143 ++++++++++++++++++++++++++++++++++\n 3 files changed, 228 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 3b8bf6a3dd..8decc9dc0c 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -3447,6 +3447,44 @@ static void add_cruft_object_entry(const struct object_id *oid, enum object_type\n \treturn;\n }\n \n+static void show_cruft_object(struct object *obj, const char *name, void *data)\n+{\n+\t/*\n+\t * if we did not record it earlier, it's at least as old as our\n+\t * expiration value. Rather than find it exactly, just use that\n+\t * value.  This may bump it forward from its real mtime, but it\n+\t * will still be \"too old\" next time we run with the same\n+\t * expiration.\n+\t *\n+\t * if obj does appear in the packing list, this call is a noop (or may\n+\t * set the namehash).\n+\t */\n+\tadd_cruft_object_entry(&obj->oid, obj->type, NULL, 0, name, cruft_expiration);\n+}\n+\n+static void show_cruft_commit(struct commit *commit, void *data)\n+{\n+\tshow_cruft_object((struct object*)commit, NULL, data);\n+}\n+\n+static int cruft_include_check_obj(struct object *obj, void *data)\n+{\n+\treturn !has_object_kept_pack(&obj->oid, IN_CORE_KEEP_PACKS);\n+}\n+\n+static int cruft_include_check(struct commit *commit, void *data)\n+{\n+\treturn cruft_include_check_obj((struct object*)commit, data);\n+}\n+\n+static void set_cruft_mtime(const struct object *object,\n+\t\t\t    struct packed_git *pack,\n+\t\t\t    off_t offset, time_t mtime)\n+{\n+\tadd_cruft_object_entry(&object->oid, object->type, pack, offset, NULL,\n+\t\t\t       mtime);\n+}\n+\n static void mark_pack_kept_in_core(struct string_list *packs, unsigned keep)\n {\n \tstruct string_list_item *item = NULL;\n@@ -3472,6 +3510,50 @@ static void enumerate_cruft_objects(void)\n \tstop_progress(&progress_state);\n }\n \n+static void enumerate_and_traverse_cruft_objects(struct string_list *fresh_packs)\n+{\n+\tstruct packed_git *p;\n+\tstruct rev_info revs;\n+\tint ret;\n+\n+\trepo_init_revisions(the_repository, &revs, NULL);\n+\n+\trevs.tag_objects = 1;\n+\trevs.tree_objects = 1;\n+\trevs.blob_objects = 1;\n+\n+\trevs.include_check = cruft_include_check;\n+\trevs.include_check_obj = cruft_include_check_obj;\n+\n+\trevs.ignore_missing_links = 1;\n+\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Enumerating cruft objects\"), 0);\n+\tret = add_unseen_recent_objects_to_traversal(&revs, cruft_expiration,\n+\t\t\t\t\t\t     set_cruft_mtime, 1);\n+\tstop_progress(&progress_state);\n+\n+\tif (ret)\n+\t\tdie(_(\"unable to add cruft objects\"));\n+\n+\t/*\n+\t * Re-mark only the fresh packs as kept so that objects in\n+\t * unknown packs do not halt the reachability traversal early.\n+\t */\n+\tfor (p = get_all_packs(the_repository); p; p = p->next)\n+\t\tp->pack_keep_in_core = 0;\n+\tmark_pack_kept_in_core(fresh_packs, 1);\n+\n+\tif (prepare_revision_walk(&revs))\n+\t\tdie(_(\"revision walk setup failed\"));\n+\tif (progress)\n+\t\tprogress_state = start_progress(_(\"Traversing cruft objects\"), 0);\n+\tnr_seen = 0;\n+\ttraverse_commit_list(&revs, show_cruft_commit, show_cruft_object, NULL);\n+\n+\tstop_progress(&progress_state);\n+}\n+\n static void read_cruft_objects(void)\n {\n \tstruct strbuf buf = STRBUF_INIT;\n@@ -3523,7 +3605,7 @@ static void read_cruft_objects(void)\n \tmark_pack_kept_in_core(&discard_packs, 0);\n \n \tif (cruft_expiration)\n-\t\tdie(\"--cruft-expiration not yet implemented\");\n+\t\tenumerate_and_traverse_cruft_objects(&fresh_packs);\n \telse\n \t\tenumerate_cruft_objects();\n \ndiff --git a/reachable.h b/reachable.h\nindex b776761baa..020a887b99 100644\n--- a/reachable.h\n+++ b/reachable.h\n@@ -1,10 +1,10 @@\n #ifndef REACHEABLE_H\n #define REACHEABLE_H\n \n-#include \"object.h\"\n-\n struct progress;\n struct rev_info;\n+struct object;\n+struct packed_git;\n \n typedef void report_recent_object_fn(const struct object *, struct packed_git *,\n \t\t\t\t     off_t, time_t);\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 003ca7344e..939cdc297a 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -214,5 +214,148 @@ basic_cruft_pack_tests () {\n }\n \n basic_cruft_pack_tests never\n+basic_cruft_pack_tests 2.weeks.ago\n+\n+test_expect_success 'cruft tags rescue tagged objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit tagged &&\n+\t\tgit tag -a annotated -m tag &&\n+\n+\t\tgit rev-list --objects --no-object-names packed.. >objects &&\n+\t\twhile read oid\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $oid)\"\n+\t\tdone <objects &&\n+\n+\t\ttest-tool chmtime -500 \\\n+\t\t\t\"$objdir/$(test_oid_to_path $(git rev-parse annotated))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\t\tcut -f1 -d\" \" <actual.raw | sort >actual &&\n+\n+\t\t(\n+\t\t\tcat objects &&\n+\t\t\tgit rev-parse annotated\n+\t\t) >expect.raw &&\n+\t\tsort <expect.raw >expect &&\n+\n+\t\ttest_cmp expect actual &&\n+\t\tcat actual\n+\t)\n+'\n+\n+test_expect_success 'cruft commits rescue parents, trees' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit old &&\n+\t\ttest_commit new &&\n+\n+\t\tgit rev-list --objects --no-object-names packed..new >objects &&\n+\t\twhile read object\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $object)\"\n+\t\tdone <objects &&\n+\t\ttest-tool chmtime +500 \"$objdir/$(test_oid_to_path \\\n+\t\t\t$(git rev-parse HEAD))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\n+\t\tcut -d\" \" -f1 <actual.raw | sort >actual &&\n+\t\tsort <objects >expect &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'cruft trees rescue sub-trees, blobs' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\tmkdir -p dir/sub &&\n+\t\techo foo >foo &&\n+\t\techo bar >dir/bar &&\n+\t\techo baz >dir/sub/baz &&\n+\n+\t\ttest_tick &&\n+\t\tgit add . &&\n+\t\tgit commit -m \"pruned\" &&\n+\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD^{tree}))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:foo))\" &&\n+\t\ttest-tool chmtime  -500 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/bar))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/sub))\" &&\n+\t\ttest-tool chmtime -1000 \"$objdir/$(test_oid_to_path $(git rev-parse HEAD:dir/sub/baz))\" &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual.raw &&\n+\t\tcut -f1 -d\" \" <actual.raw | sort >actual &&\n+\n+\t\tgit rev-parse HEAD:dir HEAD:dir/bar HEAD:dir/sub HEAD:dir/sub/baz >expect.raw &&\n+\t\tsort <expect.raw >expect &&\n+\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'expired objects are pruned' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit packed &&\n+\t\tgit repack -Ad &&\n+\n+\t\ttest_commit pruned &&\n+\n+\t\tgit rev-list --objects --no-object-names packed..pruned >objects &&\n+\t\twhile read object\n+\t\tdo\n+\t\t\ttest-tool chmtime -1000 \\\n+\t\t\t\t\"$objdir/$(test_oid_to_path $object)\"\n+\t\tdone <objects &&\n+\n+\t\tkeep=\"$(basename \"$(ls $packdir/pack-*.pack)\")\" &&\n+\t\tcruft=\"$(echo $keep | git pack-objects --cruft \\\n+\t\t\t--cruft-expiration=750.seconds.ago \\\n+\t\t\t$packdir/pack)\" &&\n+\n+\t\ttest-tool pack-mtimes \"pack-$cruft.mtimes\" >actual &&\n+\t\ttest_must_be_empty actual\n+\t)\n+'\n \n test_done\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455651","messageId":"0d3555d595523dc3e8d4c064038ed11040f2a383.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 12/17] builtin/repack.c: support generating a cruft pack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:18:03Z","receivedAt":"2022-05-20T23:18:27Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Expose a way to split the contents of a repository into a main and cruft\npack when doing an all-into-one repack with `git repack --cruft -d`, and\na complementary configuration variable.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/git-repack.txt            |  11 ++\n Documentation/technical/cruft-packs.txt |   2 +-\n builtin/repack.c                        | 105 +++++++++++-\n t/t5329-pack-objects-cruft.sh           | 207 ++++++++++++++++++++++++\n 4 files changed, 319 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/git-repack.txt b/Documentation/git-repack.txt\nindex ee30edc178..0bf13893d8 100644\n--- a/Documentation/git-repack.txt\n+++ b/Documentation/git-repack.txt\n@@ -63,6 +63,17 @@ to the new separate pack will be written.\n \tAlso run  'git prune-packed' to remove redundant\n \tloose object files.\n \n+--cruft::\n+\tSame as `-a`, unless `-d` is used. Then any unreachable objects\n+\tare packed into a separate cruft pack. Unreachable objects can\n+\tbe pruned using the normal expiry rules with the next `git gc`\n+\tinvocation (see linkgit:git-gc[1]). Incompatible with `-k`.\n+\n+--cruft-expiration=<approxidate>::\n+\tExpire unreachable objects older than `<approxidate>`\n+\timmediately instead of waiting for the next `git gc` invocation.\n+\tOnly useful with `--cruft -d`.\n+\n -l::\n \tPass the `--local` option to 'git pack-objects'. See\n \tlinkgit:git-pack-objects[1].\ndiff --git a/Documentation/technical/cruft-packs.txt b/Documentation/technical/cruft-packs.txt\nindex c0f583cd48..d81f3a8982 100644\n--- a/Documentation/technical/cruft-packs.txt\n+++ b/Documentation/technical/cruft-packs.txt\n@@ -17,7 +17,7 @@ pruned according to normal expiry rules with the next 'git gc' invocation.\n \n Unreachable objects aren't removed immediately, since doing so could race with\n an incoming push which may reference an object which is about to be deleted.\n-Instead, those unreachable objects are stored as loose object and stay that way\n+Instead, those unreachable objects are stored as loose objects and stay that way\n until they are older than the expiration window, at which point they are removed\n by linkgit:git-prune[1].\n \ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex e7a3920c6d..593c18d4e8 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -18,12 +18,18 @@\n #include \"pack-bitmap.h\"\n #include \"refs.h\"\n \n+#define ALL_INTO_ONE 1\n+#define LOOSEN_UNREACHABLE 2\n+#define PACK_CRUFT 4\n+\n+static int pack_everything;\n static int delta_base_offset = 1;\n static int pack_kept_objects = -1;\n static int write_bitmaps = -1;\n static int use_delta_islands;\n static int run_update_server_info = 1;\n static char *packdir, *packtmp_name, *packtmp;\n+static char *cruft_expiration;\n \n static const char *const git_repack_usage[] = {\n \tN_(\"git repack [<options>]\"),\n@@ -305,9 +311,6 @@ static void repack_promisor_objects(const struct pack_objects_args *args,\n \t\tdie(_(\"could not finish pack-objects to repack promisor objects\"));\n }\n \n-#define ALL_INTO_ONE 1\n-#define LOOSEN_UNREACHABLE 2\n-\n struct pack_geometry {\n \tstruct packed_git **pack;\n \tuint32_t pack_nr, pack_alloc;\n@@ -344,6 +347,8 @@ static void init_pack_geometry(struct pack_geometry **geometry_p)\n \tfor (p = get_all_packs(the_repository); p; p = p->next) {\n \t\tif (!pack_kept_objects && p->pack_keep)\n \t\t\tcontinue;\n+\t\tif (p->is_cruft)\n+\t\t\tcontinue;\n \n \t\tALLOC_GROW(geometry->pack,\n \t\t\t   geometry->pack_nr + 1,\n@@ -605,6 +610,67 @@ static int write_midx_included_packs(struct string_list *include,\n \treturn finish_command(&cmd);\n }\n \n+static int write_cruft_pack(const struct pack_objects_args *args,\n+\t\t\t    const char *pack_prefix,\n+\t\t\t    struct string_list *names,\n+\t\t\t    struct string_list *existing_packs,\n+\t\t\t    struct string_list *existing_kept_packs)\n+{\n+\tstruct child_process cmd = CHILD_PROCESS_INIT;\n+\tstruct strbuf line = STRBUF_INIT;\n+\tstruct string_list_item *item;\n+\tFILE *in, *out;\n+\tint ret;\n+\n+\tprepare_pack_objects(&cmd, args);\n+\n+\tstrvec_push(&cmd.args, \"--cruft\");\n+\tif (cruft_expiration)\n+\t\tstrvec_pushf(&cmd.args, \"--cruft-expiration=%s\",\n+\t\t\t     cruft_expiration);\n+\n+\tstrvec_push(&cmd.args, \"--honor-pack-keep\");\n+\tstrvec_push(&cmd.args, \"--non-empty\");\n+\tstrvec_push(&cmd.args, \"--max-pack-size=0\");\n+\n+\tcmd.in = -1;\n+\n+\tret = start_command(&cmd);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\t/*\n+\t * names has a confusing double use: it both provides the list\n+\t * of just-written new packs, and accepts the name of the cruft\n+\t * pack we are writing.\n+\t *\n+\t * By the time it is read here, it contains only the pack(s)\n+\t * that were just written, which is exactly the set of packs we\n+\t * want to consider kept.\n+\t */\n+\tin = xfdopen(cmd.in, \"w\");\n+\tfor_each_string_list_item(item, names)\n+\t\tfprintf(in, \"%s-%s.pack\\n\", pack_prefix, item->string);\n+\tfor_each_string_list_item(item, existing_packs)\n+\t\tfprintf(in, \"-%s.pack\\n\", item->string);\n+\tfor_each_string_list_item(item, existing_kept_packs)\n+\t\tfprintf(in, \"%s.pack\\n\", item->string);\n+\tfclose(in);\n+\n+\tout = xfdopen(cmd.out, \"r\");\n+\twhile (strbuf_getline_lf(&line, out) != EOF) {\n+\t\tif (line.len != the_hash_algo->hexsz)\n+\t\t\tdie(_(\"repack: Expecting full hex object ID lines only \"\n+\t\t\t      \"from pack-objects.\"));\n+\t\tstring_list_append(names, line.buf);\n+\t}\n+\tfclose(out);\n+\n+\tstrbuf_release(&line);\n+\n+\treturn finish_command(&cmd);\n+}\n+\n int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n \tstruct child_process cmd = CHILD_PROCESS_INIT;\n@@ -621,7 +687,6 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tint show_progress;\n \n \t/* variables to be filled by option parsing */\n-\tint pack_everything = 0;\n \tint delete_redundant = 0;\n \tconst char *unpack_unreachable = NULL;\n \tint keep_unreachable = 0;\n@@ -636,6 +701,11 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tOPT_BIT('A', NULL, &pack_everything,\n \t\t\t\tN_(\"same as -a, and turn unreachable objects loose\"),\n \t\t\t\t   LOOSEN_UNREACHABLE | ALL_INTO_ONE),\n+\t\tOPT_BIT(0, \"cruft\", &pack_everything,\n+\t\t\t\tN_(\"same as -a, pack unreachable cruft objects separately\"),\n+\t\t\t\t   PACK_CRUFT),\n+\t\tOPT_STRING(0, \"cruft-expiration\", &cruft_expiration, N_(\"approxidate\"),\n+\t\t\t\tN_(\"with -C, expire objects older than this\")),\n \t\tOPT_BOOL('d', NULL, &delete_redundant,\n \t\t\t\tN_(\"remove redundant packs, and run git-prune-packed\")),\n \t\tOPT_BOOL('f', NULL, &po_args.no_reuse_delta,\n@@ -688,6 +758,15 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t    (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE)))\n \t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--keep-unreachable\", \"-A\");\n \n+\tif (pack_everything & PACK_CRUFT) {\n+\t\tpack_everything |= ALL_INTO_ONE;\n+\n+\t\tif (unpack_unreachable || (pack_everything & LOOSEN_UNREACHABLE))\n+\t\t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--cruft\", \"-A\");\n+\t\tif (keep_unreachable)\n+\t\t\tdie(_(\"options '%s' and '%s' cannot be used together\"), \"--cruft\", \"-k\");\n+\t}\n+\n \tif (write_bitmaps < 0) {\n \t\tif (!write_midx &&\n \t\t    (!(pack_everything & ALL_INTO_ONE) || !is_bare_repository()))\n@@ -771,7 +850,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (pack_everything & ALL_INTO_ONE) {\n \t\trepack_promisor_objects(&po_args, &names);\n \n-\t\tif (existing_nonkept_packs.nr && delete_redundant) {\n+\t\tif (existing_nonkept_packs.nr && delete_redundant &&\n+\t\t    !(pack_everything & PACK_CRUFT)) {\n \t\t\tfor_each_string_list_item(item, &names) {\n \t\t\t\tstrvec_pushf(&cmd.args, \"--keep-pack=%s-%s.pack\",\n \t\t\t\t\t     packtmp_name, item->string);\n@@ -833,6 +913,21 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (!names.nr && !po_args.quiet)\n \t\tprintf_ln(_(\"Nothing new to pack.\"));\n \n+\tif (pack_everything & PACK_CRUFT) {\n+\t\tconst char *pack_prefix;\n+\t\tif (!skip_prefix(packtmp, packdir, &pack_prefix))\n+\t\t\tdie(_(\"pack prefix %s does not begin with objdir %s\"),\n+\t\t\t    packtmp, packdir);\n+\t\tif (*pack_prefix == '/')\n+\t\t\tpack_prefix++;\n+\n+\t\tret = write_cruft_pack(&po_args, pack_prefix, &names,\n+\t\t\t\t       &existing_nonkept_packs,\n+\t\t\t\t       &existing_kept_packs);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t}\n+\n \tfor_each_string_list_item(item, &names) {\n \t\titem->util = (void *)(uintptr_t)populate_pack_exts(item->string);\n \t}\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 939cdc297a..067c50af38 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -358,4 +358,211 @@ test_expect_success 'expired objects are pruned' '\n \t)\n '\n \n+test_expect_success 'repack --cruft generates a cruft pack' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit branch -M main &&\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit unreachable &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d unreachable &&\n+\t\t# objects are not cruft if they are contained in the reflogs\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\t\tsort <reachable.raw >reachable &&\n+\t\tcomm -13 reachable objects >unreachable &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tcruft=$(basename $(ls $packdir/pack-*.mtimes) .mtimes) &&\n+\t\tpack=$(basename $(ls $packdir/pack-*.pack | grep -v $cruft) .pack) &&\n+\n+\t\tgit show-index <$packdir/$pack.idx >actual.raw &&\n+\t\tcut -f2 -d\" \" actual.raw | sort >actual &&\n+\t\ttest_cmp reachable actual &&\n+\n+\t\tgit show-index <$packdir/$cruft.idx >actual.raw &&\n+\t\tcut -f2 -d\" \" actual.raw | sort >actual &&\n+\t\ttest_cmp unreachable actual\n+\t)\n+'\n+\n+test_expect_success 'loose objects mtimes upsert others' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit cruft &&\n+\t\t# incremental repack, leaving existing objects loose (so\n+\t\t# they can be \"freshened\")\n+\t\tgit repack &&\n+\n+\t\ttip=\"$(git rev-parse cruft)\" &&\n+\t\tpath=\"$objdir/$(test_oid_to_path \"$tip\")\" &&\n+\t\ttest-tool chmtime --get +1000 \"$path\" >expect &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tmtimes=\"$(basename $(ls $packdir/pack-*.mtimes))\" &&\n+\t\ttest-tool pack-mtimes \"$mtimes\" >actual.raw &&\n+\t\tgrep \"$tip\" actual.raw | cut -d\" \" -f2 >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'cruft packs are not included in geometric repack' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit cruft &&\n+\t\tgit repack -d &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft &&\n+\n+\t\tfind $packdir -type f | sort >before &&\n+\t\tgit repack --geometric=2 -d &&\n+\t\tfind $packdir -type f | sort >after &&\n+\n+\t\ttest_cmp before after\n+\t)\n+'\n+\n+test_expect_success 'repack --geometric collects once-cruft objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit repack -Ad &&\n+\t\tgit branch -M main &&\n+\n+\t\tgit checkout --orphan other &&\n+\t\tgit rm -rf . &&\n+\t\ttest_commit --no-tag cruft &&\n+\t\tcruft=\"$(git rev-parse HEAD)\" &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\t# Pack the objects created in the previous step into a cruft\n+\t\t# pack. Intentionally leave loose copies of those objects\n+\t\t# around so we can pick them up in a subsequent --geometric\n+\t\t# reapack.\n+\t\tgit repack --cruft &&\n+\n+\t\t# Now make those objects reachable, and ensure that they are\n+\t\t# packed into the new pack created via a --geometric repack.\n+\t\tgit update-ref refs/heads/other $cruft &&\n+\n+\t\t# Without this object, the set of unpacked objects is exactly\n+\t\t# the set of objects already in the cruft pack. Tweak that set\n+\t\t# to ensure we do not overwrite the cruft pack entirely.\n+\t\ttest_commit reachable2 &&\n+\n+\t\tfind $packdir -name \"pack-*.idx\" | sort >before &&\n+\t\tgit repack --geometric=2 -d &&\n+\t\tfind $packdir -name \"pack-*.idx\" | sort >after &&\n+\n+\t\t{\n+\t\t\tgit rev-list --objects --no-object-names $cruft &&\n+\t\t\tgit rev-list --objects --no-object-names reachable..reachable2\n+\t\t} >want.raw &&\n+\t\tsort want.raw >want &&\n+\n+\t\tpack=$(comm -13 before after) &&\n+\t\tgit show-index <$pack >objects.raw &&\n+\n+\t\tcut -d\" \" -f2 objects.raw | sort >got &&\n+\n+\t\ttest_cmp want got\n+\t)\n+'\n+\n+test_expect_success 'cruft repack with no reachable objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\t\tgit repack -ad &&\n+\n+\t\tbase=\"$(git rev-parse base)\" &&\n+\n+\t\tgit for-each-ref --format=\"delete %(refname)\" >in &&\n+\t\tgit update-ref --stdin <in &&\n+\t\tgit reflog expire --all --expire=all &&\n+\t\trm -fr .git/index &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tgit cat-file -t $base\n+\t)\n+'\n+\n+test_expect_success 'cruft repack ignores --max-pack-size' '\n+\tgit init max-pack-size &&\n+\t(\n+\t\tcd max-pack-size &&\n+\t\ttest_commit base &&\n+\t\t# two cruft objects which exceed the maximum pack size\n+\t\ttest-tool genrandom foo 1048576 | git hash-object --stdin -w &&\n+\t\ttest-tool genrandom bar 1048576 | git hash-object --stdin -w &&\n+\t\tgit repack --cruft --max-pack-size=1M &&\n+\t\tfind $packdir -name \"*.mtimes\" >cruft &&\n+\t\ttest_line_count = 1 cruft &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(cat cruft)\")\" >objects &&\n+\t\ttest_line_count = 2 objects\n+\t)\n+'\n+\n+test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n+\t(\n+\t\tcd max-pack-size &&\n+\t\t# repack everything back together to remove the existing cruft\n+\t\t# pack (but to keep its objects)\n+\t\tgit repack -adk &&\n+\t\tgit -c pack.packSizeLimit=1M repack --cruft &&\n+\t\t# ensure the same post condition is met when --max-pack-size\n+\t\t# would otherwise be inferred from the configuration\n+\t\tfind $packdir -name \"*.mtimes\" >cruft &&\n+\t\ttest_line_count = 1 cruft &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(cat cruft)\")\" >objects &&\n+\t\ttest_line_count = 2 objects\n+\t)\n+'\n+\n test_done\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455652","messageId":"4b721d3ee962257fa28ac07fde6456895a7beae9.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 13/17] builtin/repack.c: allow configuring cruft pack generation","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:18:06Z","receivedAt":"2022-05-20T23:18:43Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In servers which set the pack.window configuration to a large value, we\ncan wind up spending quite a lot of time finding new bases when breaking\ndelta chains between reachable and unreachable objects while generating\na cruft pack.\n\nIntroduce a handful of `repack.cruft*` configuration variables to\ncontrol the parameters used by pack-objects when generating a cruft\npack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/repack.txt |  9 ++++\n builtin/repack.c                | 49 +++++++++++++------\n t/t5329-pack-objects-cruft.sh   | 83 +++++++++++++++++++++++++++++++++\n 3 files changed, 127 insertions(+), 14 deletions(-)\n\ndiff --git a/Documentation/config/repack.txt b/Documentation/config/repack.txt\nindex 41ac6953c8..c79af6d7b8 100644\n--- a/Documentation/config/repack.txt\n+++ b/Documentation/config/repack.txt\n@@ -30,3 +30,12 @@ repack.updateServerInfo::\n \tIf set to false, linkgit:git-repack[1] will not run\n \tlinkgit:git-update-server-info[1]. Defaults to true. Can be overridden\n \twhen true by the `-n` option of linkgit:git-repack[1].\n+\n+repack.cruftWindow::\n+repack.cruftWindowMemory::\n+repack.cruftDepth::\n+repack.cruftThreads::\n+\tParameters used by linkgit:git-pack-objects[1] when generating\n+\ta cruft pack and the respective parameters are not given over\n+\tthe command line. See similarly named `pack.*` configuration\n+\tvariables for defaults and meaning.\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 593c18d4e8..b85483a148 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -41,9 +41,21 @@ static const char incremental_bitmap_conflict_error[] = N_(\n \"--no-write-bitmap-index or disable the pack.writebitmaps configuration.\"\n );\n \n+struct pack_objects_args {\n+\tconst char *window;\n+\tconst char *window_memory;\n+\tconst char *depth;\n+\tconst char *threads;\n+\tconst char *max_pack_size;\n+\tint no_reuse_delta;\n+\tint no_reuse_object;\n+\tint quiet;\n+\tint local;\n+};\n \n static int repack_config(const char *var, const char *value, void *cb)\n {\n+\tstruct pack_objects_args *cruft_po_args = cb;\n \tif (!strcmp(var, \"repack.usedeltabaseoffset\")) {\n \t\tdelta_base_offset = git_config_bool(var, value);\n \t\treturn 0;\n@@ -65,6 +77,14 @@ static int repack_config(const char *var, const char *value, void *cb)\n \t\trun_update_server_info = git_config_bool(var, value);\n \t\treturn 0;\n \t}\n+\tif (!strcmp(var, \"repack.cruftwindow\"))\n+\t\treturn git_config_string(&cruft_po_args->window, var, value);\n+\tif (!strcmp(var, \"repack.cruftwindowmemory\"))\n+\t\treturn git_config_string(&cruft_po_args->window_memory, var, value);\n+\tif (!strcmp(var, \"repack.cruftdepth\"))\n+\t\treturn git_config_string(&cruft_po_args->depth, var, value);\n+\tif (!strcmp(var, \"repack.cruftthreads\"))\n+\t\treturn git_config_string(&cruft_po_args->threads, var, value);\n \treturn git_default_config(var, value, cb);\n }\n \n@@ -157,18 +177,6 @@ static void remove_redundant_pack(const char *dir_name, const char *base_name)\n \tstrbuf_release(&buf);\n }\n \n-struct pack_objects_args {\n-\tconst char *window;\n-\tconst char *window_memory;\n-\tconst char *depth;\n-\tconst char *threads;\n-\tconst char *max_pack_size;\n-\tint no_reuse_delta;\n-\tint no_reuse_object;\n-\tint quiet;\n-\tint local;\n-};\n-\n static void prepare_pack_objects(struct child_process *cmd,\n \t\t\t\t const struct pack_objects_args *args)\n {\n@@ -692,6 +700,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tint keep_unreachable = 0;\n \tstruct string_list keep_pack_list = STRING_LIST_INIT_NODUP;\n \tstruct pack_objects_args po_args = {NULL};\n+\tstruct pack_objects_args cruft_po_args = {NULL};\n \tint geometric_factor = 0;\n \tint write_midx = 0;\n \n@@ -746,7 +755,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tOPT_END()\n \t};\n \n-\tgit_config(repack_config, NULL);\n+\tgit_config(repack_config, &cruft_po_args);\n \n \targc = parse_options(argc, argv, prefix, builtin_repack_options,\n \t\t\t\tgit_repack_usage, 0);\n@@ -921,7 +930,19 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tif (*pack_prefix == '/')\n \t\t\tpack_prefix++;\n \n-\t\tret = write_cruft_pack(&po_args, pack_prefix, &names,\n+\t\tif (!cruft_po_args.window)\n+\t\t\tcruft_po_args.window = po_args.window;\n+\t\tif (!cruft_po_args.window_memory)\n+\t\t\tcruft_po_args.window_memory = po_args.window_memory;\n+\t\tif (!cruft_po_args.depth)\n+\t\t\tcruft_po_args.depth = po_args.depth;\n+\t\tif (!cruft_po_args.threads)\n+\t\t\tcruft_po_args.threads = po_args.threads;\n+\n+\t\tcruft_po_args.local = po_args.local;\n+\t\tcruft_po_args.quiet = po_args.quiet;\n+\n+\t\tret = write_cruft_pack(&cruft_po_args, pack_prefix, &names,\n \t\t\t\t       &existing_nonkept_packs,\n \t\t\t\t       &existing_kept_packs);\n \t\tif (ret)\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 067c50af38..c82f973b41 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -565,4 +565,87 @@ test_expect_success 'cruft repack ignores pack.packSizeLimit' '\n \t)\n '\n \n+test_expect_success 'cruft repack respects repack.cruftWindow' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tGIT_TRACE2_EVENT=$(pwd)/event.trace \\\n+\t\tgit -c pack.window=1 -c repack.cruftWindow=2 repack \\\n+\t\t       --cruft --window=3 &&\n+\n+\t\tgrep \"pack-objects.*--window=2.*--cruft\" event.trace\n+\t)\n+'\n+\n+test_expect_success 'cruft repack respects --window by default' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tGIT_TRACE2_EVENT=$(pwd)/event.trace \\\n+\t\tgit -c pack.window=2 repack --cruft --window=3 &&\n+\n+\t\tgrep \"pack-objects.*--window=3.*--cruft\" event.trace\n+\t)\n+'\n+\n+test_expect_success 'cruft repack respects --quiet' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit base &&\n+\t\tGIT_PROGRESS_DELAY=0 git repack --cruft --quiet 2>err &&\n+\t\ttest_must_be_empty err\n+\t)\n+'\n+\n+test_expect_success 'cruft --local drops unreachable objects' '\n+\tgit init alternate &&\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr alternate repo\" &&\n+\n+\ttest_commit -C alternate base &&\n+\t# Pack all objects in alterate so that the cruft repack in \"repo\" sees\n+\t# the object it dropped due to `--local` as packed. Otherwise this\n+\t# object would not appear packed anywhere (since it is not packed in\n+\t# alternate and likewise not part of the cruft pack in the other repo\n+\t# because of `--local`).\n+\tgit -C alternate repack -ad &&\n+\n+\t(\n+\t\tcd repo &&\n+\n+\t\tobject=\"$(git -C ../alternate rev-parse HEAD:base.t)\" &&\n+\t\tgit -C ../alternate cat-file -p $object >contents &&\n+\n+\t\t# Write some reachable objects and two unreachable ones: one\n+\t\t# that the alternate has and another that is unique.\n+\t\ttest_commit other &&\n+\t\tgit hash-object -w -t blob contents &&\n+\t\tcruft=\"$(echo cruft | git hash-object -w -t blob --stdin)\" &&\n+\n+\t\t( cd ../alternate/.git/objects && pwd ) \\\n+\t\t       >.git/objects/info/alternates &&\n+\n+\t\ttest_path_is_file $objdir/$(test_oid_to_path $cruft) &&\n+\t\ttest_path_is_file $objdir/$(test_oid_to_path $object) &&\n+\n+\t\tgit repack -d --cruft --local &&\n+\n+\t\ttest-tool pack-mtimes \"$(basename $(ls $packdir/pack-*.mtimes))\" \\\n+\t\t       >objects &&\n+\t\t! grep $object objects &&\n+\t\tgrep $cruft objects\n+\t)\n+'\n+\n test_done\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455653","messageId":"e9f46e7b5ed6564f5eb6988038bb34df4d2ad2ca.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 15/17] builtin/repack.c: add cruft packs to MIDX during geometric repack","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:18:11Z","receivedAt":"2022-05-20T23:18:48Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"When using cruft packs, the following race can occur when a geometric\nrepack that writes a MIDX bitmap takes place afterwords:\n\n  - First, create an unreachable object and do an all-into-one cruft\n    repack which stores that object in the repository's cruft pack.\n  - Then make that object reachable.\n  - Finally, do a geometric repack and write a MIDX bitmap.\n\nAssuming that we are sufficiently unlucky as to select a commit from the\nMIDX which reaches that object for bitmapping, then the `git\nmulti-pack-index` process will complain that that object is missing.\n\nThe reason is because we don't include cruft packs in the MIDX when\ndoing a geometric repack. Since the \"make that object reachable\" doesn't\nnecessarily mean that we'll create a new copy of that object in one of\nthe packs that will get rolled up as part of a geometric repack, it's\npossible that the MIDX won't see any copies of that now-reachable\nobject.\n\nOf course, it's desirable to avoid including cruft packs in the MIDX\nbecause it causes the MIDX to store a bunch of objects which are likely\nto get thrown away. But excluding that pack does open us up to the above\nrace.\n\nThis patch demonstrates the bug, and resolves it by including cruft\npacks in the MIDX even when doing a geometric repack.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/repack.c              | 23 ++++++++++++++++++++---\n t/t5329-pack-objects-cruft.sh | 26 ++++++++++++++++++++++++++\n 2 files changed, 46 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 36d1f03671..15071fadbe 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -23,6 +23,7 @@\n #define PACK_CRUFT 4\n \n #define DELETE_PACK 1\n+#define CRUFT_PACK 2\n \n static int pack_everything;\n static int delta_base_offset = 1;\n@@ -159,10 +160,15 @@ static void collect_pack_filenames(struct string_list *fname_nonkept_list,\n \t\tfname = xmemdupz(e->d_name, len);\n \n \t\tif ((extra_keep->nr > 0 && i < extra_keep->nr) ||\n-\t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname))))\n+\t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname)))) {\n \t\t\tstring_list_append_nodup(fname_kept_list, fname);\n-\t\telse\n-\t\t\tstring_list_append_nodup(fname_nonkept_list, fname);\n+\t\t} else {\n+\t\t\tstruct string_list_item *item;\n+\t\t\titem = string_list_append_nodup(fname_nonkept_list,\n+\t\t\t\t\t\t\tfname);\n+\t\t\tif (file_exists(mkpath(\"%s/%s.mtimes\", packdir, fname)))\n+\t\t\t\titem->util = (void*)(uintptr_t)CRUFT_PACK;\n+\t\t}\n \t}\n \tclosedir(dir);\n }\n@@ -564,6 +570,17 @@ static void midx_included_packs(struct string_list *include,\n \n \t\t\tstring_list_insert(include, strbuf_detach(&buf, NULL));\n \t\t}\n+\n+\t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n+\t\t\tif (!((uintptr_t)item->util & CRUFT_PACK)) {\n+\t\t\t\t/*\n+\t\t\t\t * no need to check DELETE_PACK, since we're not\n+\t\t\t\t * doing an ALL_INTO_ONE repack\n+\t\t\t\t */\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tstring_list_insert(include, xstrfmt(\"%s.idx\", item->string));\n+\t\t}\n \t} else {\n \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n \t\t\tif ((uintptr_t)item->util & DELETE_PACK)\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex c82f973b41..8de87afce2 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -648,4 +648,30 @@ test_expect_success 'cruft --local drops unreachable objects' '\n \t)\n '\n \n+test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\ttest_commit cruft &&\n+\t\tunreachable=\"$(git rev-parse cruft)\" &&\n+\n+\t\tgit reset --hard $unreachable^ &&\n+\t\tgit tag -d cruft &&\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\t# resurrect the unreachable object via a new commit. the\n+\t\t# new commit will get selected for a bitmap, but be\n+\t\t# missing one of its parents from the selected packs.\n+\t\tgit reset --hard $unreachable &&\n+\t\ttest_commit resurrect &&\n+\n+\t\tgit repack --write-midx --write-bitmap-index --geometric=2 -d\n+\t)\n+'\n+\n test_done\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455654","messageId":"f9e3ab56b1a8801f95a46b42084d5c7f923a8fc9.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 14/17] builtin/repack.c: use named flags for existing_packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:18:08Z","receivedAt":"2022-05-20T23:18:52Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"We use the `util` pointer for items in the `existing_packs` string list\nto indicate which packs are going to be deleted. Since that has so far\nbeen the only use of that `util` pointer, we just set it to 0 or 1.\n\nBut we're going to add an additional state to this field in the next\npatch, so prepare for that by adding a #define for the first bit so we\ncan more expressively inspect the flags state.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n builtin/repack.c | 9 ++++++---\n 1 file changed, 6 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex b85483a148..36d1f03671 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -22,6 +22,8 @@\n #define LOOSEN_UNREACHABLE 2\n #define PACK_CRUFT 4\n \n+#define DELETE_PACK 1\n+\n static int pack_everything;\n static int delta_base_offset = 1;\n static int pack_kept_objects = -1;\n@@ -564,7 +566,7 @@ static void midx_included_packs(struct string_list *include,\n \t\t}\n \t} else {\n \t\tfor_each_string_list_item(item, existing_nonkept_packs) {\n-\t\t\tif (item->util)\n+\t\t\tif ((uintptr_t)item->util & DELETE_PACK)\n \t\t\t\tcontinue;\n \t\t\tstring_list_insert(include, xstrfmt(\"%s.idx\", item->string));\n \t\t}\n@@ -1002,7 +1004,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\t * was given) and that we will actually delete this pack\n \t\t\t * (if `-d` was given).\n \t\t\t */\n-\t\t\titem->util = (void*)(intptr_t)!string_list_has_string(&names, sha1);\n+\t\t\tif (!string_list_has_string(&names, sha1))\n+\t\t\t\titem->util = (void*)(uintptr_t)((size_t)item->util | DELETE_PACK);\n \t\t}\n \t}\n \n@@ -1026,7 +1029,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tif (delete_redundant) {\n \t\tint opts = 0;\n \t\tfor_each_string_list_item(item, &existing_nonkept_packs) {\n-\t\t\tif (!item->util)\n+\t\t\tif (!((uintptr_t)item->util & DELETE_PACK))\n \t\t\t\tcontinue;\n \t\t\tremove_redundant_pack(packdir, item->string);\n \t\t}\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455655","messageId":"1e313b89e85ce0a5cc6fa6cb93127c13ae1e9e19.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 17/17] sha1-file.c: don't freshen cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:18:17Z","receivedAt":"2022-05-20T23:19:11Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"We don't bother to freshen objects stored in a cruft pack individually\nby updating the `.mtimes` file. This is because we can't portably `mmap`\nand write into the middle of a file (i.e., to update the mtime of just\none object). Instead, we would have to rewrite the entire `.mtimes` file\nwhich may incur some wasted effort especially if there a lot of cruft\nobjects and they are freshened infrequently.\n\nInstead, force the freshening code to avoid an optimizing write by\nwriting out the object loose and letting it pick up a current mtime.\n\nThis works because we prefer the mtime of the loose copy of an object\nwhen both a loose and packed one exist (whether or not the packed copy\ncomes from a cruft pack or not).\n\nThis could certainly do with a test and/or be included earlier in this\nseries/PR, but I want to wait until after I have a chance to clean up\nthe overly-repetitive nature of the cruft pack tests in general.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n object-file.c                 |  2 ++\n t/t5329-pack-objects-cruft.sh | 25 +++++++++++++++++++++++++\n 2 files changed, 27 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex ff0cffe68e..495a359200 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2035,6 +2035,8 @@ static int freshen_packed_object(const struct object_id *oid)\n \tstruct pack_entry e;\n \tif (!find_pack_entry(the_repository, oid, &e))\n \t\treturn 0;\n+\tif (e.p->is_cruft)\n+\t\treturn 0;\n \tif (e.p->freshened)\n \t\treturn 1;\n \tif (!freshen_file(e.p->pack_name))\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 70a6a9553c..b481224b93 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -711,4 +711,29 @@ test_expect_success 'MIDX bitmaps tolerate reachable cruft objects' '\n \t)\n '\n \n+test_expect_success 'cruft objects are freshend via loose' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\techo \"cruft\" >contents &&\n+\t\tblob=\"$(git hash-object -w -t blob contents)\" &&\n+\t\tloose=\"$objdir/$(test_oid_to_path $blob)\" &&\n+\n+\t\ttest_commit base &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\ttest_path_is_missing \"$loose\" &&\n+\t\ttest-tool pack-mtimes \"$(basename \"$(ls $packdir/pack-*.mtimes)\")\" >cruft &&\n+\t\tgrep \"$blob\" cruft &&\n+\n+\t\t# write the same object again\n+\t\tgit hash-object -w -t blob contents &&\n+\n+\t\ttest_path_is_file \"$loose\"\n+\t)\n+'\n+\n test_done\n-- \n2.36.1.94.gb0d54bedca\n"},{"id":"455656","messageId":"43c14eec0762170393c5e9681c3d5ef8fa60c96c.1653088640.git.me@ttaylorr.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"[PATCH v5 16/17] builtin/gc.c: conditionally avoid pruning objects via loose","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:18:14Z","receivedAt":"2022-05-20T23:19:13Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Expose the new `git repack --cruft` mode from `git gc` via a new opt-in\nflag. When invoked like `git gc --cruft`, `git gc` will avoid exploding\nunreachable objects as loose ones, and instead create a cruft pack and\n`.mtimes` file.\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n Documentation/config/gc.txt   | 21 +++++++++++++-------\n Documentation/git-gc.txt      |  5 +++++\n builtin/gc.c                  | 10 +++++++++-\n t/t5329-pack-objects-cruft.sh | 37 +++++++++++++++++++++++++++++++++++\n 4 files changed, 65 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/config/gc.txt b/Documentation/config/gc.txt\nindex c834e07991..38fea076a2 100644\n--- a/Documentation/config/gc.txt\n+++ b/Documentation/config/gc.txt\n@@ -81,14 +81,21 @@ gc.packRefs::\n \tto enable it within all non-bare repos or it can be set to a\n \tboolean value.  The default is `true`.\n \n+gc.cruftPacks::\n+\tStore unreachable objects in a cruft pack (see\n+\tlinkgit:git-repack[1]) instead of as loose objects. The default\n+\tis `false`.\n+\n gc.pruneExpire::\n-\tWhen 'git gc' is run, it will call 'prune --expire 2.weeks.ago'.\n-\tOverride the grace period with this config variable.  The value\n-\t\"now\" may be used to disable this grace period and always prune\n-\tunreachable objects immediately, or \"never\" may be used to\n-\tsuppress pruning.  This feature helps prevent corruption when\n-\t'git gc' runs concurrently with another process writing to the\n-\trepository; see the \"NOTES\" section of linkgit:git-gc[1].\n+\tWhen 'git gc' is run, it will call 'prune --expire 2.weeks.ago'\n+\t(and 'repack --cruft --cruft-expiration 2.weeks.ago' if using\n+\tcruft packs via `gc.cruftPacks` or `--cruft`).  Override the\n+\tgrace period with this config variable.  The value \"now\" may be\n+\tused to disable this grace period and always prune unreachable\n+\tobjects immediately, or \"never\" may be used to suppress pruning.\n+\tThis feature helps prevent corruption when 'git gc' runs\n+\tconcurrently with another process writing to the repository; see\n+\tthe \"NOTES\" section of linkgit:git-gc[1].\n \n gc.worktreePruneExpire::\n \tWhen 'git gc' is run, it calls\ndiff --git a/Documentation/git-gc.txt b/Documentation/git-gc.txt\nindex 853967dea0..ba4e67700e 100644\n--- a/Documentation/git-gc.txt\n+++ b/Documentation/git-gc.txt\n@@ -54,6 +54,11 @@ other housekeeping tasks (e.g. rerere, working trees, reflog...) will\n be performed as well.\n \n \n+--cruft::\n+\tWhen expiring unreachable objects, pack them separately into a\n+\tcruft pack instead of storing the loose objects as loose\n+\tobjects.\n+\n --prune=<date>::\n \tPrune loose objects older than date (default is 2 weeks ago,\n \toverridable by the config variable `gc.pruneExpire`).\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex b335cffa33..4d995e85e9 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -42,6 +42,7 @@ static const char * const builtin_gc_usage[] = {\n \n static int pack_refs = 1;\n static int prune_reflogs = 1;\n+static int cruft_packs = 0;\n static int aggressive_depth = 50;\n static int aggressive_window = 250;\n static int gc_auto_threshold = 6700;\n@@ -152,6 +153,7 @@ static void gc_config(void)\n \tgit_config_get_int(\"gc.auto\", &gc_auto_threshold);\n \tgit_config_get_int(\"gc.autopacklimit\", &gc_auto_pack_limit);\n \tgit_config_get_bool(\"gc.autodetach\", &detach_auto);\n+\tgit_config_get_bool(\"gc.cruftpacks\", &cruft_packs);\n \tgit_config_get_expiry(\"gc.pruneexpire\", &prune_expire);\n \tgit_config_get_expiry(\"gc.worktreepruneexpire\", &prune_worktrees_expire);\n \tgit_config_get_expiry(\"gc.logexpiry\", &gc_log_expire);\n@@ -331,7 +333,11 @@ static void add_repack_all_option(struct string_list *keep_pack)\n {\n \tif (prune_expire && !strcmp(prune_expire, \"now\"))\n \t\tstrvec_push(&repack, \"-a\");\n-\telse {\n+\telse if (cruft_packs) {\n+\t\tstrvec_push(&repack, \"--cruft\");\n+\t\tif (prune_expire)\n+\t\t\tstrvec_pushf(&repack, \"--cruft-expiration=%s\", prune_expire);\n+\t} else {\n \t\tstrvec_push(&repack, \"-A\");\n \t\tif (prune_expire)\n \t\t\tstrvec_pushf(&repack, \"--unpack-unreachable=%s\", prune_expire);\n@@ -551,6 +557,7 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t{ OPTION_STRING, 0, \"prune\", &prune_expire, N_(\"date\"),\n \t\t\tN_(\"prune unreferenced objects\"),\n \t\t\tPARSE_OPT_OPTARG, NULL, (intptr_t)prune_expire },\n+\t\tOPT_BOOL(0, \"cruft\", &cruft_packs, N_(\"pack unreferenced objects separately\")),\n \t\tOPT_BOOL(0, \"aggressive\", &aggressive, N_(\"be more thorough (increased runtime)\")),\n \t\tOPT_BOOL_F(0, \"auto\", &auto_gc, N_(\"enable auto-gc mode\"),\n \t\t\t   PARSE_OPT_NOCOMPLETE),\n@@ -670,6 +677,7 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t\tdie(FAILED_RUN, repack.v[0]);\n \n \t\tif (prune_expire) {\n+\t\t\t/* run `git prune` even if using cruft packs */\n \t\t\tstrvec_push(&prune, prune_expire);\n \t\t\tif (quiet)\n \t\t\t\tstrvec_push(&prune, \"--no-progress\");\ndiff --git a/t/t5329-pack-objects-cruft.sh b/t/t5329-pack-objects-cruft.sh\nindex 8de87afce2..70a6a9553c 100755\n--- a/t/t5329-pack-objects-cruft.sh\n+++ b/t/t5329-pack-objects-cruft.sh\n@@ -429,6 +429,43 @@ test_expect_success 'loose objects mtimes upsert others' '\n \t)\n '\n \n+test_expect_success 'expiring cruft objects with git gc' '\n+\tgit init repo &&\n+\ttest_when_finished \"rm -fr repo\" &&\n+\t(\n+\t\tcd repo &&\n+\n+\t\ttest_commit reachable &&\n+\t\tgit branch -M main &&\n+\t\tgit checkout --orphan other &&\n+\t\ttest_commit unreachable &&\n+\n+\t\tgit checkout main &&\n+\t\tgit branch -D other &&\n+\t\tgit tag -d unreachable &&\n+\t\t# objects are not cruft if they are contained in the reflogs\n+\t\tgit reflog expire --all --expire=all &&\n+\n+\t\tgit rev-list --objects --all --no-object-names >reachable.raw &&\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\t\tsort <reachable.raw >reachable &&\n+\t\tcomm -13 reachable objects >unreachable &&\n+\n+\t\tgit repack --cruft -d &&\n+\n+\t\tmtimes=$(ls .git/objects/pack/pack-*.mtimes) &&\n+\t\ttest_path_is_file $mtimes &&\n+\n+\t\tgit gc --cruft --prune=now &&\n+\n+\t\tgit cat-file --batch-all-objects --batch-check=\"%(objectname)\" >objects &&\n+\n+\t\tcomm -23 unreachable objects >removed &&\n+\t\ttest_cmp unreachable removed &&\n+\t\ttest_path_is_missing $mtimes\n+\t)\n+'\n+\n test_expect_success 'cruft packs are not included in geometric repack' '\n \tgit init repo &&\n \ttest_when_finished \"rm -fr repo\" &&\n-- \n2.36.1.94.gb0d54bedca\n\n"},{"id":"455657","messageId":"xmqqa6bbg5cc.fsf@gitster.g","threadId":"56996","inReplyTo":"98d9bbe5-1902-0dc4-e41e-33020d0396ad@github.com","subject":"Re: [PATCH v4 00/17] cruft packs","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-20T23:19:15Z","receivedAt":"2022-05-20T23:19:25Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <derrickstolee@github.com> writes:\n\n>>   - updating the `finalize_hashfile()` calls for writing `.mtimes` files\n>>     to indicate that they are `FSYNC_COMPONENT_PACK_METADATA`, since the\n>>     original version of this series predates the fine-grained fsync\n>>     configuration in 2.36.\n>\n> Good to have this update and not require it to be handled at merge\n> time by the maintainer.\n\nHeh, my rerere database is good enough to make it a non-issue ;-)\n\n>> As always, a range-diff is below. Thanks in advance for taking another\n>> look!\n>\n> Looking at the range-diff, I'm happy with this version.\n\nThanks.  I am tempted to mark the topic as \"expecting (hopefully the\nfinal) reroll\", to be merged down to 'next' soonish.\n\n\n"},{"id":"455661","messageId":"Yogkj5ypIdLdLnIC@nand.local","threadId":"56996","inReplyTo":"xmqqa6bbg5cc.fsf@gitster.g","subject":"Re: [PATCH v4 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-20T23:30:23Z","receivedAt":"2022-05-20T23:30:33Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, May 20, 2022 at 04:19:15PM -0700, Junio C Hamano wrote:\n> >> As always, a range-diff is below. Thanks in advance for taking another\n> >> look!\n> >\n> > Looking at the range-diff, I'm happy with this version.\n>\n> Thanks.  I am tempted to mark the topic as \"expecting (hopefully the\n> final) reroll\", to be merged down to 'next' soonish.\n\nHere it is:\n\n    https://lore.kernel.org/git/cover.1653088640.git.me@ttaylorr.com/.\n\nThanks,\nTaylor\n"},{"id":"455670","messageId":"220521.868rqv15tj.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"cover.1653088640.git.me@ttaylorr.com","subject":"Re: [PATCH v5 00/17] cruft packs","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-21T11:17:20Z","receivedAt":"2022-05-21T11:33:01Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, May 20 2022, Taylor Blau wrote:\n\n>   - The new section in pack-format.txt (describing the \".mtimes\" format) now\n>     says at the top \"all 4-byte numbers are in network byte order\", and avoids\n>     repeating \"network [byte] order\" throughout that section to reduce\n>     confusion.\n\nSuggestion (outside this series) since that's fixed perhaps a small\nstand-alone patch to fix the existing \"network order\" occurances in the\nsame file?\n\n> ...and that's pretty much it. In any case, a range-diff is included below.\n> Thanks again for all of the thoughtful feedback on this series.\n\nIt seems this didn't make it on-list, but for the last round (well, it\nseems I replied to v1 by accident) A sent this (at\nhttps://lore.kernel.org/git/220519.86ilq14u1a.gmgdl@evledraar.gmail.com/\nif it eventually shows up):\n\t\n\tReturn-Path: <avarab@gmail.com>\n\tReceived: from gmgdl (dhcp-077-248-183-071.chello.nl. [77.248.183.71])\n\t        by smtp.gmail.com with ESMTPSA id en21-20020a17090728d500b006fa9820b4a2sm1979775ejc.165.2022.05.19.04.54.26\n\t        (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256);\n\t        Thu, 19 May 2022 04:54:26 -0700 (PDT)\n\tReceived: from avar by gmgdl with local (Exim 4.95)\n\t\t(envelope-from <avarab@gmail.com>)\n\t\tid 1nrejZ-0026V8-AE;\n\t\tThu, 19 May 2022 13:54:25 +0200\n\tFrom: =?utf-8?B?w4Z2YXIgQXJuZmrDtnLDsA==?= Bjarmason <avarab@gmail.com>\n\tTo: Taylor Blau <me@ttaylorr.com>\n\tCc: git@vger.kernel.org, gitster@pobox.com, larsxschneider@gmail.com,\n\t peff@peff.net, tytso@mit.edu, brian m. carlson <bk2204@github.com>\n\tSubject: Re: [PATCH 05/17] pack-mtimes: support writing pack .mtimes files\n\tDate: Thu, 19 May 2022 13:48:42 +0200\n\tReferences: <cover.1638224692.git.me@ttaylorr.com>\n\t <deece9eb70e9750bb8350946679b521e59139fe2.1638224692.git.me@ttaylorr.com>\n\tUser-agent: Debian GNU/Linux bookworm/sid; Emacs 27.1; mu4e 1.7.12\n\tIn-reply-to: <deece9eb70e9750bb8350946679b521e59139fe2.1638224692.git.me@ttaylorr.com>\n\tMessage-ID: <220519.86ilq14u1a.gmgdl@evledraar.gmail.com>\n\tMIME-Version: 1.0\n\tContent-Type: text/plain\n\tX-TUID: ta0yc5HgmrrD\n\t\n\t\n\tOn Mon, Nov 29 2021, Taylor Blau wrote:\n\t\n\t> +static void write_mtimes_header(struct hashfile *f)\n\t> +{\n\t> +\thashwrite_be32(f, MTIMES_SIGNATURE);\n\t> +\thashwrite_be32(f, MTIMES_VERSION);\n\t> +\thashwrite_be32(f, oid_version(the_hash_algo));\n\t> +}\n\t\n\tGiven the history noted in\n\thttps://lore.kernel.org/git/RFC-patch-2.2-051f0612ab9-20220519T113538Z-avarab@gmail.com/\n\tmaybe we can say this ship has just sailed at this point.\n\t\n\tBut since this is a new format I think it's worth considering not using\n\tthe 1 or 2 you get from oid_version(), but the \"format_id\",\n\ti.e. GIT_SHA1_FORMAT_ID or GIT_SHA256_FORMAT_ID.\n\t\n\tYou'll use the same space in the format for it, but we'll end up with\n\tsomething more obvious (as the integer encodes the sha1 or sha256 name).\n\t\n\tAFAICT this code is just copied from your earlier work on *.rev, which\n\tin turn seems copied from earlier work on midx & commit-graph, which\n\tseems to have used this way of referring to the hash version more as an\n\taccident than anything explicitly indended...\n\t\n\tThen again we could just say that both are equally valid at this point,\n\tespecially given the use in adjacent formats.\n\nI.e. do we think continuing to use 1 v.s. 2 in new formats over instead\nof 0x73686131 and 0x73323536 is the right choice?\n\nOther than that the only question I have (I think) on this series is if\nJonathan Nieder is happy with it. I looked back in my logs and there was\nan extensive on-IRC discussion about it at the end of March, which ended\nin you sending: https://lore.kernel.org/git/YkICkpttOujOKeT3@nand.local/\n\nBut it seems Jonathan didn't chime in since then, and he had some major\nissues with the approach here. I think those should have been addressed\nby that discussion, but it would be nice to get a confirmation.\n\t\n> Taylor Blau (17):\n>   Documentation/technical: add cruft-packs.txt\n>   pack-mtimes: support reading .mtimes files\n>   pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'\n>   chunk-format.h: extract oid_version()\n>   pack-mtimes: support writing pack .mtimes files\n>   t/helper: add 'pack-mtimes' test-tool\n>   builtin/pack-objects.c: return from create_object_entry()\n>   builtin/pack-objects.c: --cruft without expiration\n>   reachable: add options to add_unseen_recent_objects_to_traversal\n>   reachable: report precise timestamps from objects in cruft packs\n>   builtin/pack-objects.c: --cruft with expiration\n>   builtin/repack.c: support generating a cruft pack\n>   builtin/repack.c: allow configuring cruft pack generation\n>   builtin/repack.c: use named flags for existing_packs\n>   builtin/repack.c: add cruft packs to MIDX during geometric repack\n>   builtin/gc.c: conditionally avoid pruning objects via loose\n>   sha1-file.c: don't freshen cruft packs\n>\n>  Documentation/Makefile                  |   1 +\n>  Documentation/config/gc.txt             |  21 +-\n>  Documentation/config/repack.txt         |   9 +\n>  Documentation/git-gc.txt                |   5 +\n>  Documentation/git-pack-objects.txt      |  30 +\n>  Documentation/git-repack.txt            |  11 +\n>  Documentation/technical/cruft-packs.txt | 123 ++++\n>  Documentation/technical/pack-format.txt |  19 +\n>  Makefile                                |   2 +\n>  builtin/gc.c                            |  10 +-\n>  builtin/pack-objects.c                  | 304 +++++++++-\n>  builtin/repack.c                        | 185 +++++-\n>  bulk-checkin.c                          |   2 +-\n>  chunk-format.c                          |  12 +\n>  chunk-format.h                          |   3 +\n>  commit-graph.c                          |  18 +-\n>  midx.c                                  |  18 +-\n>  object-file.c                           |   4 +-\n>  object-store.h                          |   7 +-\n>  pack-mtimes.c                           | 126 ++++\n>  pack-mtimes.h                           |  15 +\n>  pack-objects.c                          |   6 +\n>  pack-objects.h                          |  25 +\n>  pack-write.c                            |  93 ++-\n>  pack.h                                  |   4 +\n>  packfile.c                              |  19 +-\n>  reachable.c                             |  58 +-\n>  reachable.h                             |   9 +-\n>  t/helper/test-pack-mtimes.c             |  56 ++\n>  t/helper/test-tool.c                    |   1 +\n>  t/helper/test-tool.h                    |   1 +\n>  t/t5329-pack-objects-cruft.sh           | 739 ++++++++++++++++++++++++\n>  32 files changed, 1834 insertions(+), 102 deletions(-)\n>  create mode 100644 Documentation/technical/cruft-packs.txt\n>  create mode 100644 pack-mtimes.c\n>  create mode 100644 pack-mtimes.h\n>  create mode 100644 t/helper/test-pack-mtimes.c\n>  create mode 100755 t/t5329-pack-objects-cruft.sh\n>\n> Range-diff against v4:\n>  -:  ---------- >  1:  f494ef7377 Documentation/technical: add cruft-packs.txt\n>  1:  8f9fd21be9 !  2:  91a9d21b0b pack-mtimes: support reading .mtimes files\n>     @@ Documentation/technical/pack-format.txt: Pack file entry: <+\n>       \n>      +== pack-*.mtimes files have the format:\n>      +\n>     ++All 4-byte numbers are in network byte order.\n>     ++\n>      +  - A 4-byte magic number '0x4d544d45' ('MTME').\n>      +\n>      +  - A 4-byte version identifier (= 1).\n>      +\n>      +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n>      +\n>     -+  - A table of 4-byte unsigned integers in network order. The ith\n>     -+    value is the modification time (mtime) of the ith object in the\n>     -+    corresponding pack by lexicographic (index) order. The mtimes\n>     -+    count standard epoch seconds.\n>     ++  - A table of 4-byte unsigned integers. The ith value is the\n>     ++    modification time (mtime) of the ith object in the corresponding\n>     ++    pack by lexicographic (index) order. The mtimes count standard\n>     ++    epoch seconds.\n>      +\n>      +  - A trailer, containing a checksum of the corresponding packfile,\n>      +    and a checksum of all of the above (each having length according\n>      +    to the specified hash function).\n>     -+\n>     -+All 4-byte numbers are in network order.\n>      +\n>       == multi-pack-index (MIDX) files have the following format:\n>       \n>  2:  cdb21236e1 =  3:  67c4e7209d pack-write: pass 'struct packing_data' to 'stage_tmp_packfiles'\n>  3:  1d775f9850 =  4:  fc86506881 chunk-format.h: extract oid_version()\n>  4:  6172861bd9 =  5:  788d1f96f2 pack-mtimes: support writing pack .mtimes files\n>  5:  5f9a9a5b7b =  6:  2a6cfb00bf t/helper: add 'pack-mtimes' test-tool\n>  6:  b8a38fe2e4 =  7:  edb6fcd5ec builtin/pack-objects.c: return from create_object_entry()\n>  7:  94fe03cc65 =  8:  e3185741f2 builtin/pack-objects.c: --cruft without expiration\n>  8:  da7273f41f =  9:  1cf00d462c reachable: add options to add_unseen_recent_objects_to_traversal\n>  9:  58fecd1747 = 10:  d66be44d9a reachable: report precise timestamps from objects in cruft packs\n> 10:  1740b8ef01 = 11:  1434e37623 builtin/pack-objects.c: --cruft with expiration\n> 11:  5992a72cbf ! 12:  0d3555d595 builtin/repack.c: support generating a cruft pack\n>     @@ t/t5329-pack-objects-cruft.sh: test_expect_success 'expired objects are pruned'\n>      +\t\tgit repack &&\n>      +\n>      +\t\ttip=\"$(git rev-parse cruft)\" &&\n>     -+\t\tpath=\"$objdir/$(test_oid_to_path \"$(git rev-parse cruft)\")\" &&\n>     ++\t\tpath=\"$objdir/$(test_oid_to_path \"$tip\")\" &&\n>      +\t\ttest-tool chmtime --get +1000 \"$path\" >expect &&\n>      +\n>      +\t\tgit checkout main &&\n> 12:  1b241f8f91 = 13:  4b721d3ee9 builtin/repack.c: allow configuring cruft pack generation\n> 13:  ffae78852c = 14:  f9e3ab56b1 builtin/repack.c: use named flags for existing_packs\n> 14:  0743e373ba ! 15:  e9f46e7b5e builtin/repack.c: add cruft packs to MIDX during geometric repack\n>     @@ builtin/repack.c\n>       static int pack_everything;\n>       static int delta_base_offset = 1;\n>      @@ builtin/repack.c: static void collect_pack_filenames(struct string_list *fname_nonkept_list,\n>     + \t\tfname = xmemdupz(e->d_name, len);\n>     + \n>       \t\tif ((extra_keep->nr > 0 && i < extra_keep->nr) ||\n>     - \t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname))))\n>     +-\t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname))))\n>     ++\t\t    (file_exists(mkpath(\"%s/%s.keep\", packdir, fname)))) {\n>       \t\t\tstring_list_append_nodup(fname_kept_list, fname);\n>      -\t\telse\n>      -\t\t\tstring_list_append_nodup(fname_nonkept_list, fname);\n>     -+\t\telse {\n>     -+\t\t\tstruct string_list_item *item = string_list_append_nodup(fname_nonkept_list, fname);\n>     ++\t\t} else {\n>     ++\t\t\tstruct string_list_item *item;\n>     ++\t\t\titem = string_list_append_nodup(fname_nonkept_list,\n>     ++\t\t\t\t\t\t\tfname);\n>      +\t\t\tif (file_exists(mkpath(\"%s/%s.mtimes\", packdir, fname)))\n>      +\t\t\t\titem->util = (void*)(uintptr_t)CRUFT_PACK;\n>      +\t\t}\n> 15:  9f7e0acac6 = 16:  43c14eec07 builtin/gc.c: conditionally avoid pruning objects via loose\n> 16:  07fa9d4b47 = 17:  1e313b89e8 sha1-file.c: don't freshen cruft packs\n\n"},{"id":"455937","messageId":"Yo0ysWZKFJoiCSqv@google.com","threadId":"56996","inReplyTo":"91a9d21b0b7d99023083c0bbb6f91ccdc1782736.1653088640.git.me@ttaylorr.com","subject":"Re: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2022-05-24T19:32:01Z","receivedAt":"2022-05-24T19:32:12Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nTaylor Blau wrote:\n\n> This patch prepares for cruft packs by defining the `.mtimes` format,\n> and introducing a basic API that callers can use to read out individual\n> mtimes.\n\nMakes sense.  Does this intend to produce any functional change?  I'm\nguessing not (and the lack of tests agrees), but the commit message\ndoesn't say so.\n\nBy the way, is this something we could cover in tests, e.g. using a\ntest helper that exercises the new code?\n\n[...]\n> --- a/Documentation/technical/pack-format.txt\n> +++ b/Documentation/technical/pack-format.txt\n> @@ -294,6 +294,25 @@ Pack file entry: <+\n>  \n>  All 4-byte numbers are in network order.\n>  \n> +== pack-*.mtimes files have the format:\n> +\n> +All 4-byte numbers are in network byte order.\n> +\n> +  - A 4-byte magic number '0x4d544d45' ('MTME').\n> +\n> +  - A 4-byte version identifier (= 1).\n> +\n> +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n> +\n> +  - A table of 4-byte unsigned integers. The ith value is the\n> +    modification time (mtime) of the ith object in the corresponding\n> +    pack by lexicographic (index) order. The mtimes count standard\n> +    epoch seconds.\n> +\n> +  - A trailer, containing a checksum of the corresponding packfile,\n> +    and a checksum of all of the above (each having length according\n> +    to the specified hash function).\n> +\n\nThis describes the \"syntax\" but not the \"semantics\" of the file.\nShould I look to a separate piece of documentation for the semantics?\nIf so, can this one include a mention of that piece of documentation\nto make it easier to find?\n\n[...]\n> --- a/object-store.h\n> +++ b/object-store.h\n> @@ -115,12 +115,15 @@ struct packed_git {\n>  \t\t freshened:1,\n>  \t\t do_not_close:1,\n>  \t\t pack_promisor:1,\n> -\t\t multi_pack_index:1;\n> +\t\t multi_pack_index:1,\n> +\t\t is_cruft:1;\n>  \tunsigned char hash[GIT_MAX_RAWSZ];\n>  \tstruct revindex_entry *revindex;\n>  \tconst uint32_t *revindex_data;\n>  \tconst uint32_t *revindex_map;\n>  \tsize_t revindex_size;\n> +\tconst uint32_t *mtimes_map;\n> +\tsize_t mtimes_size;\n\nWhat does mtimes_map contain?  A comment would help.\n\n\n> --- /dev/null\n> +++ b/pack-mtimes.c\n> @@ -0,0 +1,126 @@\n> +#include \"pack-mtimes.h\"\n> +#include \"object-store.h\"\n> +#include \"packfile.h\"\n\nMissing #include of git-compat-util.h.\n\n> +\n> +static char *pack_mtimes_filename(struct packed_git *p)\n> +{\n> +\tsize_t len;\n> +\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n> +\t\tBUG(\"pack_name does not end in .pack\");\n> +\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n> +\treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n> +}\n\nThis seems simple enough that it's not obvious we need more code\nsharing.  Do you agree?  If so, I'd suggest just removing the\nNEEDSWORK comment.\n\n> +\n> +#define MTIMES_HEADER_SIZE (12)\n> +#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n\nHm, the all-caps name makes this feel like a compile-time constant but\nit contains a reference to the_hash_algo.  Could it be an inline\nfunction instead?\n\n> +\n> +struct mtimes_header {\n> +\tuint32_t signature;\n> +\tuint32_t version;\n> +\tuint32_t hash_id;\n> +};\n> +\n> +static int load_pack_mtimes_file(char *mtimes_file,\n> +\t\t\t\t uint32_t num_objects,\n> +\t\t\t\t const uint32_t **data_p, size_t *len_p)\n\nWhat does this function do?  A comment would help.\n\n> +{\n> +\tint fd, ret = 0;\n> +\tstruct stat st;\n> +\tvoid *data = NULL;\n> +\tsize_t mtimes_size;\n> +\tstruct mtimes_header header;\n> +\tuint32_t *hdr;\n> +\n> +\tfd = git_open(mtimes_file);\n> +\n> +\tif (fd < 0) {\n\nnit: this would be more readable without the blank line between\nsetting and checking fd (likewise for the other examples below).\n> +\t\tret = -1;\n> +\t\tgoto cleanup;\n> +\t}\n\n[...]\n> +\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t), num_objects)) {\n\nThis presupposes that the hash_id matches the_hash_algo.  Maybe worth\na NEEDSWORK comment.\n\n[...]\n> +cleanup:\n> +\tif (ret) {\n> +\t\tif (data)\n> +\t\t\tmunmap(data, mtimes_size);\n> +\t} else {\n> +\t\t*len_p = mtimes_size;\n> +\t\t*data_p = (const uint32_t *)data;\n\nDo we know that 'data' is uint32_t aligned?  Casting earlier in the\nfunction could make that more obvious.\n\n[...]\n> +int load_pack_mtimes(struct packed_git *p)\n\nThis could use a doc comment in the header file.  For example, what\nrequirements do we have on what the caller passes as 'p'?\n\n[...]\n> +uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos)\n\nLikewise.\n\n[...]\n> --- a/packfile.c\n> +++ b/packfile.c\n[...]\n> @@ -363,7 +373,7 @@ void close_object_store(struct raw_object_store *o)\n>  \n>  void unlink_pack_path(const char *pack_name, int force_delete)\n>  {\n> -\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\"};\n> +\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\", \".mtimes\"};\n\nAre these in any particular order?  Should they be?\n\n[...]\n> @@ -718,6 +728,10 @@ struct packed_git *add_packed_git(const char *path, size_t path_len, int local)\n>  \tif (!access(p->pack_name, F_OK))\n>  \t\tp->pack_promisor = 1;\n>  \n> +\txsnprintf(p->pack_name + path_len, alloc - path_len, \".mtimes\");\n> +\tif (!access(p->pack_name, F_OK))\n> +\t\tp->is_cruft = 1;\n> +\n>  \txsnprintf(p->pack_name + path_len, alloc - path_len, \".pack\");\n>  \tif (stat(p->pack_name, &st) || !S_ISREG(st.st_mode)) {\n>  \t\tfree(p);\n> @@ -869,7 +883,8 @@ static void prepare_pack(const char *full_name, size_t full_name_len,\n>  \t    ends_with(file_name, \".pack\") ||\n>  \t    ends_with(file_name, \".bitmap\") ||\n>  \t    ends_with(file_name, \".keep\") ||\n> -\t    ends_with(file_name, \".promisor\"))\n> +\t    ends_with(file_name, \".promisor\") ||\n> +\t    ends_with(file_name, \".mtimes\"))\n\nlikewise\n\nThanks,\nJonathan\n"},{"id":"455939","messageId":"Yo00X0NEu8N0MnZV@google.com","threadId":"56996","inReplyTo":"220521.868rqv15tj.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v5 00/17] cruft packs","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2022-05-24T19:39:11Z","receivedAt":"2022-05-24T19:39:20Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nÆvar Arnfjörð Bjarmason wrote:\n> On Fri, May 20 2022, Taylor Blau wrote:\n> \tOn Mon, Nov 29 2021, Taylor Blau wrote:\n\n> \t> +static void write_mtimes_header(struct hashfile *f)\n> \t> +{\n> \t> +\thashwrite_be32(f, MTIMES_SIGNATURE);\n> \t> +\thashwrite_be32(f, MTIMES_VERSION);\n> \t> +\thashwrite_be32(f, oid_version(the_hash_algo));\n> \t> +}\n[...]\n> \tBut since this is a new format I think it's worth considering not using\n> \tthe 1 or 2 you get from oid_version(), but the \"format_id\",\n> \ti.e. GIT_SHA1_FORMAT_ID or GIT_SHA256_FORMAT_ID.\n>\n> \tYou'll use the same space in the format for it, but we'll end up with\n> \tsomething more obvious (as the integer encodes the sha1 or sha256 name).\n\nAgreed.\n\n[...]\n> Other than that the only question I have (I think) on this series is if\n> Jonathan Nieder is happy with it. I looked back in my logs and there was\n> an extensive on-IRC discussion about it at the end of March, which ended\n> in you sending: https://lore.kernel.org/git/YkICkpttOujOKeT3@nand.local/\n>\n> But it seems Jonathan didn't chime in since then, and he had some major\n> issues with the approach here. I think those should have been addressed\n> by that discussion, but it would be nice to get a confirmation.\n\nI would still prefer if this used a repository format extension, but\nthat preference is not strong enough that I'd say \"this must not go in\nwithout one\".  What I think would help would be some information in\nthe user-facing documentation for commands that create and work with\ncruft packs.  In other words, if our take on people sharing\nrepositories between implementations that understand and don't\nunderstand cruft packs and get objects moving back and forth between\npacked and loose objects is \"you should have known you were doing\nsomething strange\", the least we can do is to warn them.\n\nI don't see a config to enable PACK_CRUFT by default yet in this\nseries.  I'd like one, so that people can turn it on and get the good\nnew behavior. :)\n\nThanks,\nJonathan\n"},{"id":"455941","messageId":"015d01d86fa6$a10519f0$e30f4dd0$@nexbridge.com","threadId":"56996","inReplyTo":"Yo0ysWZKFJoiCSqv@google.com","subject":"RE: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-05-24T19:44:00Z","receivedAt":"2022-05-24T19:44:11Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On May 24, 2022 3:32 PM, Taylor Blau wrote:\n>Taylor Blau wrote:\n>\n>> This patch prepares for cruft packs by defining the `.mtimes` format,\n>> and introducing a basic API that callers can use to read out\n>> individual mtimes.\n>\n>Makes sense.  Does this intend to produce any functional change?  I'm\nguessing\n>not (and the lack of tests agrees), but the commit message doesn't say so.\n>\n>By the way, is this something we could cover in tests, e.g. using a test\nhelper that\n>exercises the new code?\n>\n>[...]\n>> --- a/Documentation/technical/pack-format.txt\n>> +++ b/Documentation/technical/pack-format.txt\n>> @@ -294,6 +294,25 @@ Pack file entry: <+\n>>\n>>  All 4-byte numbers are in network order.\n>>\n>> +== pack-*.mtimes files have the format:\n>> +\n>> +All 4-byte numbers are in network byte order.\n>> +\n>> +  - A 4-byte magic number '0x4d544d45' ('MTME').\n>> +\n>> +  - A 4-byte version identifier (= 1).\n>> +\n>> +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n>> +\n>> +  - A table of 4-byte unsigned integers. The ith value is the\n>> +    modification time (mtime) of the ith object in the corresponding\n>> +    pack by lexicographic (index) order. The mtimes count standard\n>> +    epoch seconds.\n>> +\n>> +  - A trailer, containing a checksum of the corresponding packfile,\n>> +    and a checksum of all of the above (each having length according\n>> +    to the specified hash function).\n>> +\n>\n>This describes the \"syntax\" but not the \"semantics\" of the file.\n>Should I look to a separate piece of documentation for the semantics?\n>If so, can this one include a mention of that piece of documentation to\nmake it\n>easier to find?\n>\n>[...]\n>> --- a/object-store.h\n>> +++ b/object-store.h\n>> @@ -115,12 +115,15 @@ struct packed_git {\n>>  \t\t freshened:1,\n>>  \t\t do_not_close:1,\n>>  \t\t pack_promisor:1,\n>> -\t\t multi_pack_index:1;\n>> +\t\t multi_pack_index:1,\n>> +\t\t is_cruft:1;\n>>  \tunsigned char hash[GIT_MAX_RAWSZ];\n>>  \tstruct revindex_entry *revindex;\n>>  \tconst uint32_t *revindex_data;\n>>  \tconst uint32_t *revindex_map;\n>>  \tsize_t revindex_size;\n>> +\tconst uint32_t *mtimes_map;\n>> +\tsize_t mtimes_size;\n>\n>What does mtimes_map contain?  A comment would help.\n>\n>\n>> --- /dev/null\n>> +++ b/pack-mtimes.c\n>> @@ -0,0 +1,126 @@\n>> +#include \"pack-mtimes.h\"\n>> +#include \"object-store.h\"\n>> +#include \"packfile.h\"\n>\n>Missing #include of git-compat-util.h.\n>\n>> +\n>> +static char *pack_mtimes_filename(struct packed_git *p) {\n>> +\tsize_t len;\n>> +\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n>> +\t\tBUG(\"pack_name does not end in .pack\");\n>> +\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n>> +\treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name); }\n>\n>This seems simple enough that it's not obvious we need more code sharing.\nDo\n>you agree?  If so, I'd suggest just removing the NEEDSWORK comment.\n>\n>> +\n>> +#define MTIMES_HEADER_SIZE (12)\n>> +#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 *\n>> +the_hash_algo->rawsz))\n>\n>Hm, the all-caps name makes this feel like a compile-time constant but it\ncontains\n>a reference to the_hash_algo.  Could it be an inline function instead?\n>\n>> +\n>> +struct mtimes_header {\n>> +\tuint32_t signature;\n>> +\tuint32_t version;\n>> +\tuint32_t hash_id;\n>> +};\n>> +\n>> +static int load_pack_mtimes_file(char *mtimes_file,\n>> +\t\t\t\t uint32_t num_objects,\n>> +\t\t\t\t const uint32_t **data_p, size_t *len_p)\n>\n>What does this function do?  A comment would help.\n>\n>> +{\n>> +\tint fd, ret = 0;\n>> +\tstruct stat st;\n>> +\tvoid *data = NULL;\n>> +\tsize_t mtimes_size;\n>> +\tstruct mtimes_header header;\n>> +\tuint32_t *hdr;\n>> +\n>> +\tfd = git_open(mtimes_file);\n>> +\n>> +\tif (fd < 0) {\n>\n>nit: this would be more readable without the blank line between setting and\n>checking fd (likewise for the other examples below).\n>> +\t\tret = -1;\n>> +\t\tgoto cleanup;\n>> +\t}\n>\n>[...]\n>> +\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t),\n>> +num_objects)) {\n>\n>This presupposes that the hash_id matches the_hash_algo.  Maybe worth a\n>NEEDSWORK comment.\n>\n>[...]\n>> +cleanup:\n>> +\tif (ret) {\n>> +\t\tif (data)\n>> +\t\t\tmunmap(data, mtimes_size);\n>> +\t} else {\n>> +\t\t*len_p = mtimes_size;\n>> +\t\t*data_p = (const uint32_t *)data;\n>\n>Do we know that 'data' is uint32_t aligned?  Casting earlier in the\nfunction could\n>make that more obvious.\n>\n>[...]\n>> +int load_pack_mtimes(struct packed_git *p)\n>\n>This could use a doc comment in the header file.  For example, what\nrequirements\n>do we have on what the caller passes as 'p'?\n>\n>[...]\n>> +uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos)\n>\n>Likewise.\n>\n>[...]\n>> --- a/packfile.c\n>> +++ b/packfile.c\n>[...]\n>> @@ -363,7 +373,7 @@ void close_object_store(struct raw_object_store\n>> *o)\n>>\n>>  void unlink_pack_path(const char *pack_name, int force_delete)  {\n>> -\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\",\n\".bitmap\",\n>\".promisor\"};\n>> +\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\",\n>> +\".bitmap\", \".promisor\", \".mtimes\"};\n>\n>Are these in any particular order?  Should they be?\n>\n>[...]\n>> @@ -718,6 +728,10 @@ struct packed_git *add_packed_git(const char *path,\n>size_t path_len, int local)\n>>  \tif (!access(p->pack_name, F_OK))\n>>  \t\tp->pack_promisor = 1;\n>>\n>> +\txsnprintf(p->pack_name + path_len, alloc - path_len, \".mtimes\");\n>> +\tif (!access(p->pack_name, F_OK))\n>> +\t\tp->is_cruft = 1;\n>> +\n>>  \txsnprintf(p->pack_name + path_len, alloc - path_len, \".pack\");\n>>  \tif (stat(p->pack_name, &st) || !S_ISREG(st.st_mode)) {\n>>  \t\tfree(p);\n>> @@ -869,7 +883,8 @@ static void prepare_pack(const char *full_name,\nsize_t\n>full_name_len,\n>>  \t    ends_with(file_name, \".pack\") ||\n>>  \t    ends_with(file_name, \".bitmap\") ||\n>>  \t    ends_with(file_name, \".keep\") ||\n>> -\t    ends_with(file_name, \".promisor\"))\n>> +\t    ends_with(file_name, \".promisor\") ||\n>> +\t    ends_with(file_name, \".mtimes\"))\n>\n>likewise\n\nI am again concerned about 32-bit time_t assumptions. time_t is 32-bit on\nsome platforms, signed/unsigned, and sometimes 64-bit. We are talking about\npotentially long-persistent files, as I understand this series, so we should\nnot be limiting times to end at 2038. That's only 16 years off and I would\nwager that many clones that exist today will exist then.\n--Randall\n\n"},{"id":"455970","messageId":"Yo1TIQqvlxhvLZ58@nand.local","threadId":"56996","inReplyTo":"Yo00X0NEu8N0MnZV@google.com","subject":"Re: [PATCH v5 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-24T21:50:25Z","receivedAt":"2022-05-24T21:50:43Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, May 24, 2022 at 12:39:11PM -0700, Jonathan Nieder wrote:\n> Ævar Arnfjörð Bjarmason wrote:\n> > On Fri, May 20 2022, Taylor Blau wrote:\n> > \tOn Mon, Nov 29 2021, Taylor Blau wrote:\n>\n> > \t> +static void write_mtimes_header(struct hashfile *f)\n> > \t> +{\n> > \t> +\thashwrite_be32(f, MTIMES_SIGNATURE);\n> > \t> +\thashwrite_be32(f, MTIMES_VERSION);\n> > \t> +\thashwrite_be32(f, oid_version(the_hash_algo));\n> > \t> +}\n> [...]\n> > \tBut since this is a new format I think it's worth considering not using\n> > \tthe 1 or 2 you get from oid_version(), but the \"format_id\",\n> > \ti.e. GIT_SHA1_FORMAT_ID or GIT_SHA256_FORMAT_ID.\n> >\n> > \tYou'll use the same space in the format for it, but we'll end up with\n> > \tsomething more obvious (as the integer encodes the sha1 or sha256 name).\n>\n> Agreed.\n\nI know we recommend using the format_id for on-disk formats, but I think\nthere is enough existing uses of \"1\" or \"2\" that either are acceptable\nin practice.\n\nE.g., grepping around for \"hashwrite.*oid_version\", there are three\nexisting formats that use \"1\" or \"2\" instead of the format_id. They are:\n\n  - the commit-graph format\n  - the midx format\n  - the .rev format\n\nMoreover, I can't seem to find any formats that _don't_ use that\nconvention. So I have a vague preference towards using the values \"1\"\nand \"2\" as we currently do in these patches. (TBH, I don't find \"sha1\"\nsignificantly more interpretable than just \"1\", so I would be just as\nhappy leaving it as-is).\n\n> [...]\n> > Other than that the only question I have (I think) on this series is if\n> > Jonathan Nieder is happy with it. I looked back in my logs and there was\n> > an extensive on-IRC discussion about it at the end of March, which ended\n> > in you sending: https://lore.kernel.org/git/YkICkpttOujOKeT3@nand.local/\n> >\n> > But it seems Jonathan didn't chime in since then, and he had some major\n> > issues with the approach here. I think those should have been addressed\n> > by that discussion, but it would be nice to get a confirmation.\n>\n> I would still prefer if this used a repository format extension, but\n> that preference is not strong enough that I'd say \"this must not go in\n> without one\".  What I think would help would be some information in\n> the user-facing documentation for commands that create and work with\n> cruft packs.  In other words, if our take on people sharing\n> repositories between implementations that understand and don't\n> understand cruft packs and get objects moving back and forth between\n> packed and loose objects is \"you should have known you were doing\n> something strange\", the least we can do is to warn them.\n\nI think that's a good suggestion. We already have some documentation in\nDocumentation/technical/cruft-packs.txt, but I think it could be helpful\nto add user-facing documentation, too.\n\nWould you be opposed to doing that outside of this series? ISTM that the\ntechnical discussion has mostly settled, so I'd rather wordsmith the\nuser-facing documentation separately.\n\n> I don't see a config to enable PACK_CRUFT by default yet in this\n> series.  I'd like one, so that people can turn it on and get the good\n> new behavior. :)\n\n`git gc` has support for this (c.f., \"gc.cruftPacks\"). `git repack`\nrequires you to pass `--cruft`; IIRC I originally had a similar\nconfiguration in `git repack` which would change the behavior of `-A` /\n`-a` when set, but I found it too confusing and scrapped it.\n\nThanks,\nTaylor\n"},{"id":"455975","messageId":"220525.86sfoytwjn.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"Yo1TIQqvlxhvLZ58@nand.local","subject":"Re: [PATCH v5 00/17] cruft packs","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-24T21:55:02Z","receivedAt":"2022-05-24T22:07:04Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, May 24 2022, Taylor Blau wrote:\n\n> On Tue, May 24, 2022 at 12:39:11PM -0700, Jonathan Nieder wrote:\n>> Ævar Arnfjörð Bjarmason wrote:\n>> > On Fri, May 20 2022, Taylor Blau wrote:\n>> > \tOn Mon, Nov 29 2021, Taylor Blau wrote:\n>>\n>> > \t> +static void write_mtimes_header(struct hashfile *f)\n>> > \t> +{\n>> > \t> +\thashwrite_be32(f, MTIMES_SIGNATURE);\n>> > \t> +\thashwrite_be32(f, MTIMES_VERSION);\n>> > \t> +\thashwrite_be32(f, oid_version(the_hash_algo));\n>> > \t> +}\n>> [...]\n>> > \tBut since this is a new format I think it's worth considering not using\n>> > \tthe 1 or 2 you get from oid_version(), but the \"format_id\",\n>> > \ti.e. GIT_SHA1_FORMAT_ID or GIT_SHA256_FORMAT_ID.\n>> >\n>> > \tYou'll use the same space in the format for it, but we'll end up with\n>> > \tsomething more obvious (as the integer encodes the sha1 or sha256 name).\n>>\n>> Agreed.\n>\n> I know we recommend using the format_id for on-disk formats, but I think\n> there is enough existing uses of \"1\" or \"2\" that either are acceptable\n> in practice.\n>\n> E.g., grepping around for \"hashwrite.*oid_version\", there are three\n> existing formats that use \"1\" or \"2\" instead of the format_id. They are:\n>\n>   - the commit-graph format\n>   - the midx format\n>   - the .rev format\n>\n> Moreover, I can't seem to find any formats that _don't_ use that\n> convention.\n\nIt's used in the reftable format.\n\n> So I have a vague preference towards using the values \"1\"\n> and \"2\" as we currently do in these patches.\n\nI suspect that's less \"vague\" and more \"c'mon, I'm using it in\nproduction already\" :)\n\nAnyway, I'm fine with leaving it be as you have it currently. I'd first\nencountered this magic with reftable I think, and only recently found\nthat we use 1 and 2 in these other more recent places.\n\n> (TBH, I don't find \"sha1\"\n> significantly more interpretable than just \"1\", so I would be just as\n> happy leaving it as-is).\n\nHrm, I'd think having it sha1 or s256 in big-endian would be a bit more\nself-explanatory. I.e. SHA-256 is 2, not 256, and our 3 (if that ever\narrives) is likely not to be SHA-3 (but probably some successor).\n"},{"id":"455976","messageId":"Yo1YZM2dI6t+RsWv@nand.local","threadId":"56996","inReplyTo":"220525.86sfoytwjn.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v5 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-24T22:12:52Z","receivedAt":"2022-05-24T22:12:59Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, May 24, 2022 at 11:55:02PM +0200, Ævar Arnfjörð Bjarmason wrote:\n>\n> On Tue, May 24 2022, Taylor Blau wrote:\n>\n> > On Tue, May 24, 2022 at 12:39:11PM -0700, Jonathan Nieder wrote:\n> >> Ævar Arnfjörð Bjarmason wrote:\n> >> > On Fri, May 20 2022, Taylor Blau wrote:\n> >> > \tOn Mon, Nov 29 2021, Taylor Blau wrote:\n> >>\n> >> > \t> +static void write_mtimes_header(struct hashfile *f)\n> >> > \t> +{\n> >> > \t> +\thashwrite_be32(f, MTIMES_SIGNATURE);\n> >> > \t> +\thashwrite_be32(f, MTIMES_VERSION);\n> >> > \t> +\thashwrite_be32(f, oid_version(the_hash_algo));\n> >> > \t> +}\n> >> [...]\n> >> > \tBut since this is a new format I think it's worth considering not using\n> >> > \tthe 1 or 2 you get from oid_version(), but the \"format_id\",\n> >> > \ti.e. GIT_SHA1_FORMAT_ID or GIT_SHA256_FORMAT_ID.\n> >> >\n> >> > \tYou'll use the same space in the format for it, but we'll end up with\n> >> > \tsomething more obvious (as the integer encodes the sha1 or sha256 name).\n> >>\n> >> Agreed.\n> >\n> > I know we recommend using the format_id for on-disk formats, but I think\n> > there is enough existing uses of \"1\" or \"2\" that either are acceptable\n> > in practice.\n> >\n> > E.g., grepping around for \"hashwrite.*oid_version\", there are three\n> > existing formats that use \"1\" or \"2\" instead of the format_id. They are:\n> >\n> >   - the commit-graph format\n> >   - the midx format\n> >   - the .rev format\n> >\n> > Moreover, I can't seem to find any formats that _don't_ use that\n> > convention.\n>\n> It's used in the reftable format.\n\nAh, thanks for pointing it out. Still, I think there's enough uses of\n\"1\" and \"2\" over format_id that I'm not convinced here.\n\n> > So I have a vague preference towards using the values \"1\"\n> > and \"2\" as we currently do in these patches.\n>\n> I suspect that's less \"vague\" and more \"c'mon, I'm using it in\n> production already\" :)\n\nNo, this wasn't a veil over anything. Yes, GitHub is using this in\nproduction already, but that isn't why I'm opposed here. I'm opposed for\nthe reasons I explained in the quoted bits (and would happily carry a\nsmall amount of custom code in GitHub's fork to continue to recognize\nthe \"1\" or \"2\" values if this ever changed to use format_id).\n\n> Anyway, I'm fine with leaving it be as you have it currently. I'd first\n> encountered this magic with reftable I think, and only recently found\n> that we use 1 and 2 in these other more recent places.\n\nSounds good. Unless others have a very strong opinion, let's leave it as\nis.\n\nThanks,\nTaylor\n"},{"id":"455979","messageId":"Yo1aaLDmPKJ5/rh5@nand.local","threadId":"56996","inReplyTo":"Yo0ysWZKFJoiCSqv@google.com","subject":"Re: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-24T22:21:28Z","receivedAt":"2022-05-24T22:21:35Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, May 24, 2022 at 12:32:01PM -0700, Jonathan Nieder wrote:\n> Hi,\n>\n> Taylor Blau wrote:\n>\n> > This patch prepares for cruft packs by defining the `.mtimes` format,\n> > and introducing a basic API that callers can use to read out individual\n> > mtimes.\n>\n> Makes sense.  Does this intend to produce any functional change?  I'm\n> guessing not (and the lack of tests agrees), but the commit message\n> doesn't say so.\n>\n> By the way, is this something we could cover in tests, e.g. using a\n> test helper that exercises the new code?\n\nThis does not produce a functional change, no. This commit in isolation\nadds a bunch of dead code that will be used (and tested) in the\nfollowing patches.\n\nThere is a test helper that is added (and then used extensively further\non in the series) four patches later, c.f., \"t/helper: add 'pack-mtimes'\ntest-tool\".\n\n> [...]\n> > --- a/Documentation/technical/pack-format.txt\n> > +++ b/Documentation/technical/pack-format.txt\n> > @@ -294,6 +294,25 @@ Pack file entry: <+\n> >\n> >  All 4-byte numbers are in network order.\n> >\n> > +== pack-*.mtimes files have the format:\n> > +\n> > +All 4-byte numbers are in network byte order.\n> > +\n> > +  - A 4-byte magic number '0x4d544d45' ('MTME').\n> > +\n> > +  - A 4-byte version identifier (= 1).\n> > +\n> > +  - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).\n> > +\n> > +  - A table of 4-byte unsigned integers. The ith value is the\n> > +    modification time (mtime) of the ith object in the corresponding\n> > +    pack by lexicographic (index) order. The mtimes count standard\n> > +    epoch seconds.\n> > +\n> > +  - A trailer, containing a checksum of the corresponding packfile,\n> > +    and a checksum of all of the above (each having length according\n> > +    to the specified hash function).\n> > +\n>\n> This describes the \"syntax\" but not the \"semantics\" of the file.\n> Should I look to a separate piece of documentation for the semantics?\n> If so, can this one include a mention of that piece of documentation\n> to make it easier to find?\n>\n> [...]\n> > --- a/object-store.h\n> > +++ b/object-store.h\n> > @@ -115,12 +115,15 @@ struct packed_git {\n> >  \t\t freshened:1,\n> >  \t\t do_not_close:1,\n> >  \t\t pack_promisor:1,\n> > -\t\t multi_pack_index:1;\n> > +\t\t multi_pack_index:1,\n> > +\t\t is_cruft:1;\n> >  \tunsigned char hash[GIT_MAX_RAWSZ];\n> >  \tstruct revindex_entry *revindex;\n> >  \tconst uint32_t *revindex_data;\n> >  \tconst uint32_t *revindex_map;\n> >  \tsize_t revindex_size;\n> > +\tconst uint32_t *mtimes_map;\n> > +\tsize_t mtimes_size;\n>\n> What does mtimes_map contain?  A comment would help.\n\nIt contains a pointer at the beginning of the mmapped region of the\n.mtimes file, similar to revindex_map above it.\n\n>\n> > --- /dev/null\n> > +++ b/pack-mtimes.c\n> > @@ -0,0 +1,126 @@\n> > +#include \"pack-mtimes.h\"\n> > +#include \"object-store.h\"\n> > +#include \"packfile.h\"\n>\n> Missing #include of git-compat-util.h.\n\nAh, good eyes: thanks.\n\nJunio: would you like a replacement patch / a whole new copy of the\nseries / or can you amend this locally when queuing? Whatever is lowest\neffort for you works for me.\n\n> > +\n> > +static char *pack_mtimes_filename(struct packed_git *p)\n> > +{\n> > +\tsize_t len;\n> > +\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n> > +\t\tBUG(\"pack_name does not end in .pack\");\n> > +\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n> > +\treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n> > +}\n>\n> This seems simple enough that it's not obvious we need more code\n> sharing.  Do you agree?  If so, I'd suggest just removing the\n> NEEDSWORK comment.\n\nYeah, it is conceptually simple, though it feels like the sort of thing\nthat could benefit from not having to be written once for each\nextension (hence the comment).\n\n> > +\n> > +#define MTIMES_HEADER_SIZE (12)\n> > +#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n>\n> Hm, the all-caps name makes this feel like a compile-time constant but\n> it contains a reference to the_hash_algo.  Could it be an inline\n> function instead?\n\nYes, it could be an inline function, but I don't think there is\nnecessarily anything wrong with it being a #define'd macro. There are\nsome other examples, e.g., RIDX_MIN_SIZE, MIDX_MIN_SIZE,\nGRAPH_DATA_WIDTH, and PACK_SIZE_THRESHOLD (to name a few) which also use\nthe_hash_algo on the right-hand side of a `#define`.\n\n> > +\n> > +struct mtimes_header {\n> > +\tuint32_t signature;\n> > +\tuint32_t version;\n> > +\tuint32_t hash_id;\n> > +};\n> > +\n> > +static int load_pack_mtimes_file(char *mtimes_file,\n> > +\t\t\t\t uint32_t num_objects,\n> > +\t\t\t\t const uint32_t **data_p, size_t *len_p)\n>\n> What does this function do?  A comment would help.\n\nI know that I'm biased as the author of this code, but I think the\nsignature is clear here. At least, I'm not sure what information a\ncomment would add that the function name and its arguments don't already\nconvey.\n\n> > +\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t), num_objects)) {\n>\n> This presupposes that the hash_id matches the_hash_algo.  Maybe worth\n> a NEEDSWORK comment.\n\nGood catch.\n\n> [...]\n> > +cleanup:\n> > +\tif (ret) {\n> > +\t\tif (data)\n> > +\t\t\tmunmap(data, mtimes_size);\n> > +\t} else {\n> > +\t\t*len_p = mtimes_size;\n> > +\t\t*data_p = (const uint32_t *)data;\n>\n> Do we know that 'data' is uint32_t aligned?  Casting earlier in the\n> function could make that more obvious.\n\n`data` is definitely uint32_t aligned, but this is a tradeoff, since if\nwe wrote:\n\n    uint32_t *data = xmmap(...);\n\nthen I think we would have to change the case where ret is non-zero to be:\n\n    if (data)\n        munmap((void*)data, ...);\n\nand likewise, data_p is const.\n\n> > +int load_pack_mtimes(struct packed_git *p)\n>\n> This could use a doc comment in the header file.  For example, what\n> requirements do we have on what the caller passes as 'p'?\n>\n> [...]\n> > +uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos)\n>\n> Likewise.\n\nSure. I wonder when we should do that, though. I'm not trying to be\nimpatient to get this merged, but iterating on the documentation feels\nlike it could be done on top without having to re-send the substantive\nparts of this series over and over.\n\n> [...]\n> > --- a/packfile.c\n> > +++ b/packfile.c\n> [...]\n> > @@ -363,7 +373,7 @@ void close_object_store(struct raw_object_store *o)\n> >\n> >  void unlink_pack_path(const char *pack_name, int force_delete)\n> >  {\n> > -\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\"};\n> > +\tstatic const char *exts[] = {\".pack\", \".idx\", \".rev\", \".keep\", \".bitmap\", \".promisor\", \".mtimes\"};\n>\n> Are these in any particular order?  Should they be?\n>\n> [...]\n\nThese aren't technically in any particular order (nor should they be),\nthough .idx should be first. I'm leaving it alone here\n(semi-intentionally, since the race it opens up isn't related to this\nseries, and it's on my list to deal with after this code has settled).\n\n> > @@ -718,6 +728,10 @@ struct packed_git *add_packed_git(const char *path, size_t path_len, int local)\n> >  \tif (!access(p->pack_name, F_OK))\n> >  \t\tp->pack_promisor = 1;\n> >\n> > +\txsnprintf(p->pack_name + path_len, alloc - path_len, \".mtimes\");\n> > +\tif (!access(p->pack_name, F_OK))\n> > +\t\tp->is_cruft = 1;\n> > +\n> >  \txsnprintf(p->pack_name + path_len, alloc - path_len, \".pack\");\n> >  \tif (stat(p->pack_name, &st) || !S_ISREG(st.st_mode)) {\n> >  \t\tfree(p);\n> > @@ -869,7 +883,8 @@ static void prepare_pack(const char *full_name, size_t full_name_len,\n> >  \t    ends_with(file_name, \".pack\") ||\n> >  \t    ends_with(file_name, \".bitmap\") ||\n> >  \t    ends_with(file_name, \".keep\") ||\n> > -\t    ends_with(file_name, \".promisor\"))\n> > +\t    ends_with(file_name, \".promisor\") ||\n> > +\t    ends_with(file_name, \".mtimes\"))\n>\n> likewise\n\nNo specific order here (since these are all OR'd together).\n\nThanks,\nTaylor\n"},{"id":"455980","messageId":"Yo1bUbys+Fz7g+6h@nand.local","threadId":"56996","inReplyTo":"015d01d86fa6$a10519f0$e30f4dd0$@nexbridge.com","subject":"Re: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-24T22:25:21Z","receivedAt":"2022-05-24T22:25:33Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, May 24, 2022 at 03:44:00PM -0400, rsbecker@nexbridge.com wrote:\n> I am again concerned about 32-bit time_t assumptions. time_t is 32-bit on\n> some platforms, signed/unsigned, and sometimes 64-bit. We are talking about\n> potentially long-persistent files, as I understand this series, so we should\n> not be limiting times to end at 2038. That's only 16 years off and I would\n> wager that many clones that exist today will exist then.\n\nNote that we're using unsigned fields here, so we have until 2106 (see\nmy earlier response on this in\nhttps://lore.kernel.org/git/YdiXecK6fAKl8++G@nand.local/).\n\nThanks,\nTaylor\n"},{"id":"455987","messageId":"016e01d86fc5$64ecf180$2ec6d480$@nexbridge.com","threadId":"56996","inReplyTo":"Yo1bUbys+Fz7g+6h@nand.local","subject":"RE: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-05-24T23:24:14Z","receivedAt":"2022-05-24T23:24:23Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On May 24, 2022 6:25 PM ,Taylor Blau write:\n>On Tue, May 24, 2022 at 03:44:00PM -0400, rsbecker@nexbridge.com wrote:\n>> I am again concerned about 32-bit time_t assumptions. time_t is 32-bit\n>> on some platforms, signed/unsigned, and sometimes 64-bit. We are\n>> talking about potentially long-persistent files, as I understand this\n>> series, so we should not be limiting times to end at 2038. That's only\n>> 16 years off and I would wager that many clones that exist today will exist then.\n>\n>Note that we're using unsigned fields here, so we have until 2106 (see my earlier\n>response on this in https://lore.kernel.org/git/YdiXecK6fAKl8++G@nand.local/).\n\nI appreciate that, but 32-bit time_t is still signed on many platforms, so when cast, it still might, at some point in another series, cause issues. Please be cautious. I expect that this is the particular hill on which I will die. 😉\n--Randall\n\n"},{"id":"455989","messageId":"Yo1zW7ntTuNakpOD@nand.local","threadId":"56996","inReplyTo":"016e01d86fc5$64ecf180$2ec6d480$@nexbridge.com","subject":"Re: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-25T00:07:55Z","receivedAt":"2022-05-25T00:14:12Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, May 24, 2022 at 07:24:14PM -0400, rsbecker@nexbridge.com wrote:\n> On May 24, 2022 6:25 PM ,Taylor Blau write:\n> >On Tue, May 24, 2022 at 03:44:00PM -0400, rsbecker@nexbridge.com wrote:\n> >> I am again concerned about 32-bit time_t assumptions. time_t is 32-bit\n> >> on some platforms, signed/unsigned, and sometimes 64-bit. We are\n> >> talking about potentially long-persistent files, as I understand this\n> >> series, so we should not be limiting times to end at 2038. That's only\n> >> 16 years off and I would wager that many clones that exist today will exist then.\n> >\n> >Note that we're using unsigned fields here, so we have until 2106 (see my earlier\n> >response on this in https://lore.kernel.org/git/YdiXecK6fAKl8++G@nand.local/).\n>\n> I appreciate that, but 32-bit time_t is still signed on many\n> platforms, so when cast, it still might, at some point in another\n> series, cause issues. Please be cautious. I expect that this is the\n> particular hill on which I will die. 😉\n> --Randall\n\nYes, definitely. There is only one spot that we turn the result of\nnth_packed_mtime() into a time_t, and that's in\nadd_object_in_unpacked_pack(). The code there is something like:\n\n    time_t mtime;\n    if (pack->is_cruft)\n      mtime = nth_packed_mtime(pack, object_pos);\n    else\n      mtime = pack->mtime;\n\n    ...\n\n    add_cruft_object_entry(oid, ..., mtime);\n\n...and the reason mtime is a time_t is because that's the type of\npack->mtime.\n\nAnd we quickly convert that back to a uint32_t in\nadd_cruft_object_entry(). If time_t is signed, then we'll truncate any\nvalues beyond 2106, and pre-epoch values will become large positive\nvalues. That means our error is one-sided in the favorable direction,\ni.e., that we'll keep objects around for longer instead of pruning\nsomething that we shouldn't have.\n\nThanks,\nTaylor\n"},{"id":"455992","messageId":"017201d86fcd$485839f0$d908add0$@nexbridge.com","threadId":"56996","inReplyTo":"Yo1zW7ntTuNakpOD@nand.local","subject":"RE: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-05-25T00:20:42Z","receivedAt":"2022-05-25T00:20:51Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On May 24, 2022 8:08 PM, Taylor Blau wrote:\n>On Tue, May 24, 2022 at 07:24:14PM -0400, rsbecker@nexbridge.com wrote:\n>> On May 24, 2022 6:25 PM ,Taylor Blau write:\n>> >On Tue, May 24, 2022 at 03:44:00PM -0400, rsbecker@nexbridge.com wrote:\n>> >> I am again concerned about 32-bit time_t assumptions. time_t is\n>> >> 32-bit on some platforms, signed/unsigned, and sometimes 64-bit. We\n>> >> are talking about potentially long-persistent files, as I\n>> >> understand this series, so we should not be limiting times to end\n>> >> at 2038. That's only\n>> >> 16 years off and I would wager that many clones that exist today will exist\n>then.\n>> >\n>> >Note that we're using unsigned fields here, so we have until 2106\n>> >(see my earlier response on this in\n>https://lore.kernel.org/git/YdiXecK6fAKl8++G@nand.local/).\n>>\n>> I appreciate that, but 32-bit time_t is still signed on many\n>> platforms, so when cast, it still might, at some point in another\n>> series, cause issues. Please be cautious. I expect that this is the\n>> particular hill on which I will die. 😉\n>> --Randall\n>\n>Yes, definitely. There is only one spot that we turn the result of\n>nth_packed_mtime() into a time_t, and that's in\n>add_object_in_unpacked_pack(). The code there is something like:\n>\n>    time_t mtime;\n>    if (pack->is_cruft)\n>      mtime = nth_packed_mtime(pack, object_pos);\n>    else\n>      mtime = pack->mtime;\n>\n>    ...\n>\n>    add_cruft_object_entry(oid, ..., mtime);\n>\n>...and the reason mtime is a time_t is because that's the type of\n>pack->mtime.\n>\n>And we quickly convert that back to a uint32_t in add_cruft_object_entry(). If\n>time_t is signed, then we'll truncate any values beyond 2106, and pre-epoch\n>values will become large positive values. That means our error is one-sided in the\n>favorable direction, i.e., that we'll keep objects around for longer instead of\n>pruning something that we shouldn't have.\n\nI can only hope. I am working with the platform compiler team on time_t issues. Hoping we can get to 64-bit builds within two years, but out of my control. That would make time_t an int64_t, which puts failure well outside my own lifespan and anyone else in my company. Provisioning for signed 64-bit time values would be prudent even if unsupported in a specific build. We are almost 1/4 through 20xx.\n--Randall\n\n"},{"id":"456003","messageId":"Yo3fZkpkCLPbAC8B@google.com","threadId":"56996","inReplyTo":"Yo1aaLDmPKJ5/rh5@nand.local","subject":"Re: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2022-05-25T07:48:54Z","receivedAt":"2022-05-25T07:49:10Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nTaylor Blau wrote:\n> On Tue, May 24, 2022 at 12:32:01PM -0700, Jonathan Nieder wrote:\n\n>> Makes sense.  Does this intend to produce any functional change?  I'm\n>> guessing not (and the lack of tests agrees), but the commit message\n>> doesn't say so.\n[...]\n> This does not produce a functional change, no. This commit in isolation\n> adds a bunch of dead code that will be used (and tested) in the\n> following patches.\n[...]\n>> What does mtimes_map contain?  A comment would help.\n>\n> It contains a pointer at the beginning of the mmapped region of the\n> .mtimes file, similar to revindex_map above it.\n\nTo be clear, in cases like this by \"comment\" I mean \"in-code comment\".\nI.e., my interest is not that _I_ find out the answer but that the\ncode becomes more maintainable via the answer becoming easier to find.\n\n[...]\n>> This seems simple enough that it's not obvious we need more code\n>> sharing.  Do you agree?  If so, I'd suggest just removing the\n>> NEEDSWORK comment.\n>\n> Yeah, it is conceptually simple, though it feels like the sort of thing\n> that could benefit from not having to be written once for each\n> extension (hence the comment).\n\nThe reason I asked is that the NEEDSWORK here actually got in the way\nof comprehension for me --- it made me wonder \"is there some\ncomplexity here I'm missing?\"\n\nThat's why I'd suggest one of\n- removing the NEEDSWORK comment\n- going ahead and implementing the code sharing you mean, or\n- fleshing out the NEEDSWORK comment so the reader can wonder less\n\n>>> +\n>>> +#define MTIMES_HEADER_SIZE (12)\n>>> +#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n>>\n>> Hm, the all-caps name makes this feel like a compile-time constant but\n>> it contains a reference to the_hash_algo.  Could it be an inline\n>> function instead?\n>\n> Yes, it could be an inline function, but I don't think there is\n> necessarily anything wrong with it being a #define'd macro. There are\n> some other examples, e.g., RIDX_MIN_SIZE, MIDX_MIN_SIZE,\n> GRAPH_DATA_WIDTH, and PACK_SIZE_THRESHOLD (to name a few) which also use\n> the_hash_algo on the right-hand side of a `#define`.\n\nThose are due to an incomplete migration from use of the true constant\nGIT_SHA1_RAWSZ to use of the dynamic value the_hash_algo->rawsz, no?\nIn other words, \"other examples do it wrong\" doesn't feel like a great\njustification for making it worse in new code.\n\n[...]\n>>> +static int load_pack_mtimes_file(char *mtimes_file,\n>>> +\t\t\t\t uint32_t num_objects,\n>>> +\t\t\t\t const uint32_t **data_p, size_t *len_p)\n>>\n>> What does this function do?  A comment would help.\n>\n> I know that I'm biased as the author of this code, but I think the\n> signature is clear here. At least, I'm not sure what information a\n> comment would add that the function name and its arguments don't already\n> convey.\n\nAh, thanks for this point of clarification.  What isn't clear from the\nsignature is\n- when should I call this function?\n- what does its return value represent?\n- how does it handle errors?\n\nI agree that the parameters are self-explanatory.\n\n>>> +cleanup:\n>>> +\tif (ret) {\n>>> +\t\tif (data)\n>>> +\t\t\tmunmap(data, mtimes_size);\n>>> +\t} else {\n>>> +\t\t*len_p = mtimes_size;\n>>> +\t\t*data_p = (const uint32_t *)data;\n>>\n>> Do we know that 'data' is uint32_t aligned?  Casting earlier in the\n>> function could make that more obvious.\n>\n> `data` is definitely uint32_t aligned, but this is a tradeoff, since if\n> we wrote:\n>\n>     uint32_t *data = xmmap(...);\n>\n> then I think we would have to change the case where ret is non-zero to be:\n>\n>     if (data)\n>         munmap((void*)data, ...);\n>\n> and likewise, data_p is const.\n\nDoing it that way sounds great to me.  That way, the type contains the\ninformation we need up-front and the safety of the cast is obvious in\nthe place where the cast is needed.\n\n(Although my understanding is also that in C it's fine to pass a\nuint32_t* to a function expecting a void*, so the second cast would\nalso not be needed.)\n\n[...]\n>>> +int load_pack_mtimes(struct packed_git *p)\n>>\n>> This could use a doc comment in the header file.  For example, what\n>> requirements do we have on what the caller passes as 'p'?\n>>\n>> [...]\n>>> +uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos)\n>>\n>> Likewise.\n>\n> Sure. I wonder when we should do that, though. I'm not trying to be\n> impatient to get this merged, but iterating on the documentation feels\n> like it could be done on top without having to re-send the substantive\n> parts of this series over and over.\n\nIn terms of re-sending patches, sending a \"fixup!\" patch with the\nminor changes you want to make doesn't seem too problematic to me.  In\ngeneral a major benefit of code review is getting others' eyes on new\ncode from the standpoint of readability and maintainability; including\ncomments like this up front doesn't seem like a huge amount to ask\n(versus getting those comments to be perfect, which would be\nunreasonable to expect since it's not hard to update them over time).\n\n> Thanks,\n> Taylor\n\nThanks for looking it through.\n\nSincerely,\nJonathan\n"},{"id":"456004","messageId":"Yo3gl5Wv82mTZQb2@google.com","threadId":"56996","inReplyTo":"Yo1YZM2dI6t+RsWv@nand.local","subject":"Re: [PATCH v5 00/17] cruft packs","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2022-05-25T07:53:59Z","receivedAt":"2022-05-25T07:54:08Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nTaylor Blau wrote:\n> On Tue, May 24, 2022 at 11:55:02PM +0200, Ævar Arnfjörð Bjarmason wrote:\n\n>>> Moreover, I can't seem to find any formats that _don't_ use that\n>>> convention.\n>>\n>> It's used in the reftable format.\n\nIt's also used in the formats described in\nDocumentation/technical/hash-function-transition.\n\n[...]\n> Sounds good. Unless others have a very strong opinion, let's leave it as\n> is.\n\nFile formats are one of those things where a little time early can save\na lot of work later.  If there were a strong reason to use \"1\" and \"2\"\nhere then I'd be okay with living with it --- I'm a pragmatic person.\nBut in general, using the magic numbers instead of a sequential value is\nreally helpful both in making the file formats more self-explanatory and\nin making it possible to experiment with multiple new hash_algos at the\nsame time.\n\nThe main argument I'm hearing for using \"1\" and \"2\" is \"because some\nother formats got that wrong\".  That reason is the opposite of\ncompelling to me: it makes me suspect that as a project we should more\neagerly break the old bad habits and form new ones.  I guess this\nqualifies as a very strong opinion.\n\nThanks,\nJonathan\n"},{"id":"456006","messageId":"220525.86o7zmt0l0.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"Yo1zW7ntTuNakpOD@nand.local","subject":"adding new 32-bit on-disk (unsigned) timestamp formats (was: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files)","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-25T09:11:05Z","receivedAt":"2022-05-25T09:37:24Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, May 24 2022, Taylor Blau wrote:\n\n> On Tue, May 24, 2022 at 07:24:14PM -0400, rsbecker@nexbridge.com wrote:\n>> On May 24, 2022 6:25 PM ,Taylor Blau write:\n>> >On Tue, May 24, 2022 at 03:44:00PM -0400, rsbecker@nexbridge.com wrote:\n>> >> I am again concerned about 32-bit time_t assumptions. time_t is 32-bit\n>> >> on some platforms, signed/unsigned, and sometimes 64-bit. We are\n>> >> talking about potentially long-persistent files, as I understand this\n>> >> series, so we should not be limiting times to end at 2038. That's only\n>> >> 16 years off and I would wager that many clones that exist today will exist then.\n>> >\n>> >Note that we're using unsigned fields here, so we have until 2106 (see my earlier\n>> >response on this in https://lore.kernel.org/git/YdiXecK6fAKl8++G@nand.local/).\n>>\n>> I appreciate that, but 32-bit time_t is still signed on many\n>> platforms, so when cast, it still might, at some point in another\n>> series, cause issues. Please be cautious. I expect that this is the\n>> particular hill on which I will die. 😉\n>> --Randall\n>\n> Yes, definitely. There is only one spot that we turn the result of\n> nth_packed_mtime() into a time_t, and that's in\n> add_object_in_unpacked_pack(). The code there is something like:\n>\n>     time_t mtime;\n>     if (pack->is_cruft)\n>       mtime = nth_packed_mtime(pack, object_pos);\n>     else\n>       mtime = pack->mtime;\n>\n>     ...\n>\n>     add_cruft_object_entry(oid, ..., mtime);\n>\n> ...and the reason mtime is a time_t is because that's the type of\n> pack->mtime.\n>\n> And we quickly convert that back to a uint32_t in\n> add_cruft_object_entry(). If time_t is signed, then we'll truncate any\n> values beyond 2106, and pre-epoch values will become large positive\n> values. That means our error is one-sided in the favorable direction,\n> i.e., that we'll keep objects around for longer instead of pruning\n> something that we shouldn't have.\n\nI must say that I really don't like this part of the format. Is it\nreally necessary to optimize the storage space here in a way that leaves\nopen questions about future time_t compatibility, and having to\nintroduce the first use of unsigned 32 bit timestamps to git's codebase?\n\nYes, this is its own self-contained format, so we don't *need* time_t\nhere, but it's also really handy if we can eventually consistently use\n64 time_t everywhere and not worry about any compatibility issues, or\nunsigned v.s. signed, or to create our own little ext4-like signed 32\nbit timestamp format.\n\nOnce we hit 2038 (or near that date) this would be the only part of our\ncodebase & on-disk formats that I'm aware of that would differ from\ntime_t's signedness, but perhaps there's some I've missed.\n\nIf there isn't a demonstrable reason (as in some real numbers, or\naccompanying benchmark etc.) to special-snowflake this I really think we\nshould just go for signed 64 bit here, i.e. matching time_t on 64 bit\nsystems.\n\nIf we really are trying to micro-optimize storage space here I'm willing\nto bet that this is still a bad/premature optimization. There's much\nbetter ways to store this sort of data in a compact way if that's the\nconcern. E.g. you'd store a 64 bit \"base\" timestamp in the header for\nthe first entry, and have smaller (signed) \"delta\" timestamps storing\noffsets from that \"base\" timestamp.\n\nThis would take advantage of the fact that when we find loose objects\nwe're vanishingly unlikely to have them splayed over more than a\ndays/weeks/months or in the worst case small number of years from the\n\"base\" (and if we ever do we could simply shrug and leave such objects\nout of the pack entirely).\n\nWe could thus keep the 32 bit second-resolution timestamps you have\nhere, they'd just be signed deltas to the 64 bit signed \"base\" in a\nheader.\n\nEven better (again, if micro-optimizing this is really needed) would be\nto store a 64 bit signed base and a table of 16 bit signed offsets.\n\nWe'd simply declare that for our expiry times we'd \"snap\" any such\nvalues to the next day. Our current GC config exposes down-to-the-second\nexpiry times, but in practice nobody needs that. A 16 bit signed \"day\noffset\" would give you 2^15/365 = 89 years +/- of day-resolution expiry\nfor objects. To avoid thundering herds we could even fake up an exact\ndown-to-the-second expiry on the computed day by combining the expiry\ntime & the first few bits of the OID.\n\n== BREAK\n\nAside about time_t being signed v.s. unsigned. This is edited from an\nolder off-list E-Mail of mine (from git-security): For time_t itself no\nstandard says that time_t must be signed, but in practice it's\nubiquitous\n\nThis thread is informative\nhttp://mm.icann.org/pipermail/tz/2004-July/012503.html it continues the\nmonth after: http://mm.icann.org/pipermail/tz/2004-August/thread.html\n\nSummary: Yeah it can be unsigned in theory, but it seems like nobody's\nbeen crazy enough to try it, so it's de-facto standardized to\nsigned. Everyone has a Y2038 problem, nobody has a Y2106 problem. Well,\nwith time_t, e.g. Linux filesystems tend to use unsigned 32 bit epochs:\nhttps://kernelnewbies.org/y2038/vfs\n"},{"id":"456064","messageId":"32db3720-e9c8-e192-6278-c55855ce1d3e@github.com","threadId":"56996","inReplyTo":"220525.86o7zmt0l0.gmgdl@evledraar.gmail.com","subject":"Re: adding new 32-bit on-disk (unsigned) timestamp formats (was: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files)","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-05-25T13:30:55Z","receivedAt":"2022-05-25T13:31:00Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 5/25/2022 5:11 AM, Ævar Arnfjörð Bjarmason wrote:\n> I must say that I really don't like this part of the format. Is it\n> really necessary to optimize the storage space here in a way that leaves\n> open questions about future time_t compatibility, and having to\n> introduce the first use of unsigned 32 bit timestamps to git's codebase?\n\nThe commit-graph file format uses unsigned 34-bit timestamps (packed\nwith 30-bit topological levels in the CDAT chunk), so this \"not-64-bit\nsigned timestamps\" thing is something we've done before.\n \n> Yes, this is its own self-contained format, so we don't *need* time_t\n> here, but it's also really handy if we can eventually consistently use\n> 64 time_t everywhere and not worry about any compatibility issues, or\n> unsigned v.s. signed, or to create our own little ext4-like signed 32\n> bit timestamp format.\n\nWe can also use a new file format version when it is necessary. We\nhave a lot of time to add that detail without overly complicating the\nformat right now.\n\n> If we really are trying to micro-optimize storage space here I'm willing\n> to bet that this is still a bad/premature optimization. There's much\n> better ways to store this sort of data in a compact way if that's the\n> concern. E.g. you'd store a 64 bit \"base\" timestamp in the header for\n> the first entry, and have smaller (signed) \"delta\" timestamps storing\n> offsets from that \"base\" timestamp.\n\nThis is a good idea for a v2 format when that is necessary.\n\nThanks,\n-Stolee\n"},{"id":"456113","messageId":"7f5a6a6a-c554-c659-72a8-404bc39e08c7@github.com","threadId":"56996","inReplyTo":"Yo3gl5Wv82mTZQb2@google.com","subject":"Re: [PATCH v5 00/17] cruft packs","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-05-25T19:59:24Z","receivedAt":"2022-05-25T19:59:31Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 5/25/2022 3:53 AM, Jonathan Nieder wrote:\n> Taylor Blau wrote:\n>> On Tue, May 24, 2022 at 11:55:02PM +0200, Ævar Arnfjörð Bjarmason wrote:\n> \n>>>> Moreover, I can't seem to find any formats that _don't_ use that\n>>>> convention.\n>>>\n>>> It's used in the reftable format.\n\nThe use in reftable is the only one I can find and that implementation\nis not idiomatic. Specifically, the way the four-byte header was\nimplemented is not easy to extract and share in other formats.\n\nThis series does the good work of extracting oid_version() as a\ncommon method across these formats so it is easier to share.\n\n> It's also used in the formats described in\n> Documentation/technical/hash-function-transition.\n\nIt documents things that have not been implemented, such as the v3\npack-index format:\n\n  Pack index (.idx) files use a new v3 format that supports multiple\n  hash functions. They have the following format (all integers are in\n  network byte order):\n(...)\n  * 4-byte number of object formats in this pack index: 2\n  * For each object format:\n    ** 4-byte format identifier (e.g., 'sha1' for SHA-1)\n    ** 4-byte length in bytes of shortened object names. This is the\n      shortest possible length needed to make names in the shortened\n      object name table unambiguous.\n    ** 4-byte integer, recording where tables relating to this format\n      are stored in this index file, as an offset from the beginning.\n\nThis was added in your 752414ae431 (technical doc: add a design doc\nfor hash function transition, 2017-09-27), but has not been acted upon\nyet.\n\n> [...]\n>> Sounds good. Unless others have a very strong opinion, let's leave it as\n>> is.\n> \n> File formats are one of those things where a little time early can save\n> a lot of work later.  If there were a strong reason to use \"1\" and \"2\"\n> here then I'd be okay with living with it --- I'm a pragmatic person.\n> But in general, using the magic numbers instead of a sequential value is\n> really helpful both in making the file formats more self-explanatory and\n> in making it possible to experiment with multiple new hash_algos at the\n> same time.\n> \n> The main argument I'm hearing for using \"1\" and \"2\" is \"because some\n> other formats got that wrong\".  That reason is the opposite of\n> compelling to me: it makes me suspect that as a project we should more\n> eagerly break the old bad habits and form new ones.  I guess this\n> qualifies as a very strong opinion.\n\nEither way, these are magic numbers. One happens to somewhat spell\nout something when looking at the file in a hex editor with ASCII\npreviews, but that doesn't change the fact that it is most important\nthat the hash function is correctly indicated by the file format and\nparsed by the Git executable (not a human).\n\nI'd much rather have a consistent and proven way of specifying the\nhash value (using the oid_version() helper) than to try and make a\nnew mechanism.\n\nThanks,\n-Stolee\n"},{"id":"456118","messageId":"Yo6bDC8uivC3gM2o@nand.local","threadId":"56996","inReplyTo":"7f5a6a6a-c554-c659-72a8-404bc39e08c7@github.com","subject":"Re: [PATCH v5 00/17] cruft packs","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-25T21:09:32Z","receivedAt":"2022-05-25T21:09:45Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, May 25, 2022 at 03:59:24PM -0400, Derrick Stolee wrote:\n> I'd much rather have a consistent and proven way of specifying the\n> hash value (using the oid_version() helper) than to try and make a\n> new mechanism.\n\nTo be clear, I absolutely don't think any of us should have the attitude\nof repeating past bad decisions for the sake of consistency.\n\nAs best I can tell, our (Jonathan and I's) disagreement is on whether\nusing \"1\" and \"2\" to identify which hash function is used by the .mtimes\nfile is OK or not. I happen to think that it is acceptable, so the\nchoice to continue to adopt this pattern was motivated by being\nconsistent with a pattern that is good and works.\n\nThanks,\nTaylor\n"},{"id":"456120","messageId":"Yo6b+8sixGAqMm/x@nand.local","threadId":"56996","inReplyTo":"32db3720-e9c8-e192-6278-c55855ce1d3e@github.com","subject":"Re: adding new 32-bit on-disk (unsigned) timestamp formats (was: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files)","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-25T21:13:31Z","receivedAt":"2022-05-25T21:13:40Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, May 25, 2022 at 09:30:55AM -0400, Derrick Stolee wrote:\n> On 5/25/2022 5:11 AM, Ævar Arnfjörð Bjarmason wrote:\n> > I must say that I really don't like this part of the format. Is it\n> > really necessary to optimize the storage space here in a way that leaves\n> > open questions about future time_t compatibility, and having to\n> > introduce the first use of unsigned 32 bit timestamps to git's codebase?\n>\n> The commit-graph file format uses unsigned 34-bit timestamps (packed\n> with 30-bit topological levels in the CDAT chunk), so this \"not-64-bit\n> signed timestamps\" thing is something we've done before.\n>\n> > Yes, this is its own self-contained format, so we don't *need* time_t\n> > here, but it's also really handy if we can eventually consistently use\n> > 64 time_t everywhere and not worry about any compatibility issues, or\n> > unsigned v.s. signed, or to create our own little ext4-like signed 32\n> > bit timestamp format.\n>\n> We can also use a new file format version when it is necessary. We\n> have a lot of time to add that detail without overly complicating the\n> format right now.\n>\n> > If we really are trying to micro-optimize storage space here I'm willing\n> > to bet that this is still a bad/premature optimization. There's much\n> > better ways to store this sort of data in a compact way if that's the\n> > concern. E.g. you'd store a 64 bit \"base\" timestamp in the header for\n> > the first entry, and have smaller (signed) \"delta\" timestamps storing\n> > offsets from that \"base\" timestamp.\n>\n> This is a good idea for a v2 format when that is necessary.\n\nI agree here.\n\nI'm not opposed to such a change (or even being the one to work on it!),\nbut I would encourage us to pursue that change outside of this series,\nsince it can easily be done on top.\n\nOf course, if we ever did decide to implement 64-bit mtimes, we would\nhave to maintain support for reading both the 32-bit and 64-bit values.\nBut I think the code is well-equipped to do that, and it could be done\non top without significant additional complexity.\n\nThanks,\nTaylor\n"},{"id":"456123","messageId":"Yo6hcOjIlYglqdxs@nand.local","threadId":"56996","inReplyTo":"Yo3fZkpkCLPbAC8B@google.com","subject":"Re: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-25T21:36:48Z","receivedAt":"2022-05-25T21:36:55Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, May 25, 2022 at 12:48:54AM -0700, Jonathan Nieder wrote:\n> >> What does mtimes_map contain?  A comment would help.\n> >\n> > It contains a pointer at the beginning of the mmapped region of the\n> > .mtimes file, similar to revindex_map above it.\n>\n> To be clear, in cases like this by \"comment\" I mean \"in-code comment\".\n> I.e., my interest is not that _I_ find out the answer but that the\n> code becomes more maintainable via the answer becoming easier to find.\n\nOK. I'll add a comment in the fixup! patch which I'm about to send.\n\n> [...]\n> >> This seems simple enough that it's not obvious we need more code\n> >> sharing.  Do you agree?  If so, I'd suggest just removing the\n> >> NEEDSWORK comment.\n> >\n> > Yeah, it is conceptually simple, though it feels like the sort of thing\n> > that could benefit from not having to be written once for each\n> > extension (hence the comment).\n>\n> The reason I asked is that the NEEDSWORK here actually got in the way\n> of comprehension for me --- it made me wonder \"is there some\n> complexity here I'm missing?\"\n>\n> That's why I'd suggest one of\n> - removing the NEEDSWORK comment\n> - going ahead and implementing the code sharing you mean, or\n> - fleshing out the NEEDSWORK comment so the reader can wonder less\n\nI am a little sad to remove it, since I thought it was useful as-is. But\nI can just as easily remember to come back to this myself in the future,\nso if it is distracting to you in the meantime, then I don't mind\nholding onto it in my own head.\n\n> >>> +\n> >>> +#define MTIMES_HEADER_SIZE (12)\n> >>> +#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n> >>\n> >> Hm, the all-caps name makes this feel like a compile-time constant but\n> >> it contains a reference to the_hash_algo.  Could it be an inline\n> >> function instead?\n> >\n> > Yes, it could be an inline function, but I don't think there is\n> > necessarily anything wrong with it being a #define'd macro. There are\n> > some other examples, e.g., RIDX_MIN_SIZE, MIDX_MIN_SIZE,\n> > GRAPH_DATA_WIDTH, and PACK_SIZE_THRESHOLD (to name a few) which also use\n> > the_hash_algo on the right-hand side of a `#define`.\n>\n> Those are due to an incomplete migration from use of the true constant\n> GIT_SHA1_RAWSZ to use of the dynamic value the_hash_algo->rawsz, no?\n> In other words, \"other examples do it wrong\" doesn't feel like a great\n> justification for making it worse in new code.\n\nFair point. I can imagine reasons for the existing pattern, but updating\nit to handle the variable rawsz is easy to do (and it probably should\nhave been that way since the beginning).\n\n> [...]\n> >>> +static int load_pack_mtimes_file(char *mtimes_file,\n> >>> +\t\t\t\t uint32_t num_objects,\n> >>> +\t\t\t\t const uint32_t **data_p, size_t *len_p)\n> >>\n> >> What does this function do?  A comment would help.\n> >\n> > I know that I'm biased as the author of this code, but I think the\n> > signature is clear here. At least, I'm not sure what information a\n> > comment would add that the function name and its arguments don't already\n> > convey.\n>\n> Ah, thanks for this point of clarification.  What isn't clear from the\n> signature is\n> - when should I call this function?\n> - what does its return value represent?\n> - how does it handle errors?\n>\n> I agree that the parameters are self-explanatory.\n\nI'm hesitant to over-document a static function with a single caller,\nbut when looking at this, I think there is an opportunity to document\n_its_ caller (`load_pack_mtimes()`) which isn't static, but was also\nmissing documentation.\n\n> >>> +cleanup:\n> >>> +\tif (ret) {\n> >>> +\t\tif (data)\n> >>> +\t\t\tmunmap(data, mtimes_size);\n> >>> +\t} else {\n> >>> +\t\t*len_p = mtimes_size;\n> >>> +\t\t*data_p = (const uint32_t *)data;\n> >>\n> >> Do we know that 'data' is uint32_t aligned?  Casting earlier in the\n> >> function could make that more obvious.\n> >\n> > `data` is definitely uint32_t aligned, but this is a tradeoff, since if\n> > we wrote:\n> >\n> >     uint32_t *data = xmmap(...);\n> >\n> > then I think we would have to change the case where ret is non-zero to be:\n> >\n> >     if (data)\n> >         munmap((void*)data, ...);\n> >\n> > and likewise, data_p is const.\n>\n> Doing it that way sounds great to me.  That way, the type contains the\n> information we need up-front and the safety of the cast is obvious in\n> the place where the cast is needed.\n>\n> (Although my understanding is also that in C it's fine to pass a\n> uint32_t* to a function expecting a void*, so the second cast would\n> also not be needed.)\n>\n> [...]\n\nDone, thanks for the suggestion.\n\nThanks,\nTaylor\n"},{"id":"456125","messageId":"01df01d87082$a0757020$e1605060$@nexbridge.com","threadId":"56996","inReplyTo":"Yo6hcOjIlYglqdxs@nand.local","subject":"RE: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-05-25T21:58:49Z","receivedAt":"2022-05-25T21:58:59Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On May 25, 2022 5:37 PM, Taylor Blau wrote:\n>On Wed, May 25, 2022 at 12:48:54AM -0700, Jonathan Nieder wrote:\n>> >> What does mtimes_map contain?  A comment would help.\n>> >\n>> > It contains a pointer at the beginning of the mmapped region of the\n>> > .mtimes file, similar to revindex_map above it.\n>>\n>> To be clear, in cases like this by \"comment\" I mean \"in-code comment\".\n>> I.e., my interest is not that _I_ find out the answer but that the\n>> code becomes more maintainable via the answer becoming easier to find.\n>\n>OK. I'll add a comment in the fixup! patch which I'm about to send.\n>\n>> [...]\n>> >> This seems simple enough that it's not obvious we need more code\n>> >> sharing.  Do you agree?  If so, I'd suggest just removing the\n>> >> NEEDSWORK comment.\n>> >\n>> > Yeah, it is conceptually simple, though it feels like the sort of\n>> > thing that could benefit from not having to be written once for each\n>> > extension (hence the comment).\n>>\n>> The reason I asked is that the NEEDSWORK here actually got in the way\n>> of comprehension for me --- it made me wonder \"is there some\n>> complexity here I'm missing?\"\n>>\n>> That's why I'd suggest one of\n>> - removing the NEEDSWORK comment\n>> - going ahead and implementing the code sharing you mean, or\n>> - fleshing out the NEEDSWORK comment so the reader can wonder less\n>\n>I am a little sad to remove it, since I thought it was useful as-is. But I can just as\n>easily remember to come back to this myself in the future, so if it is distracting to\n>you in the meantime, then I don't mind holding onto it in my own head.\n>\n>> >>> +\n>> >>> +#define MTIMES_HEADER_SIZE (12)\n>> >>> +#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 *\n>> >>> +the_hash_algo->rawsz))\n>> >>\n>> >> Hm, the all-caps name makes this feel like a compile-time constant\n>> >> but it contains a reference to the_hash_algo.  Could it be an\n>> >> inline function instead?\n>> >\n>> > Yes, it could be an inline function, but I don't think there is\n>> > necessarily anything wrong with it being a #define'd macro. There\n>> > are some other examples, e.g., RIDX_MIN_SIZE, MIDX_MIN_SIZE,\n>> > GRAPH_DATA_WIDTH, and PACK_SIZE_THRESHOLD (to name a few) which\n>also\n>> > use the_hash_algo on the right-hand side of a `#define`.\n>>\n>> Those are due to an incomplete migration from use of the true constant\n>> GIT_SHA1_RAWSZ to use of the dynamic value the_hash_algo->rawsz, no?\n>> In other words, \"other examples do it wrong\" doesn't feel like a great\n>> justification for making it worse in new code.\n>\n>Fair point. I can imagine reasons for the existing pattern, but updating it to handle\n>the variable rawsz is easy to do (and it probably should have been that way since\n>the beginning).\n>\n>> [...]\n>> >>> +static int load_pack_mtimes_file(char *mtimes_file,\n>> >>> +\t\t\t\t uint32_t num_objects,\n>> >>> +\t\t\t\t const uint32_t **data_p, size_t *len_p)\n>> >>\n>> >> What does this function do?  A comment would help.\n>> >\n>> > I know that I'm biased as the author of this code, but I think the\n>> > signature is clear here. At least, I'm not sure what information a\n>> > comment would add that the function name and its arguments don't\n>> > already convey.\n>>\n>> Ah, thanks for this point of clarification.  What isn't clear from the\n>> signature is\n>> - when should I call this function?\n>> - what does its return value represent?\n>> - how does it handle errors?\n>>\n>> I agree that the parameters are self-explanatory.\n>\n>I'm hesitant to over-document a static function with a single caller, but when\n>looking at this, I think there is an opportunity to document _its_ caller\n>(`load_pack_mtimes()`) which isn't static, but was also missing documentation.\n>\n>> >>> +cleanup:\n>> >>> +\tif (ret) {\n>> >>> +\t\tif (data)\n>> >>> +\t\t\tmunmap(data, mtimes_size);\n>> >>> +\t} else {\n>> >>> +\t\t*len_p = mtimes_size;\n>> >>> +\t\t*data_p = (const uint32_t *)data;\n>> >>\n>> >> Do we know that 'data' is uint32_t aligned?  Casting earlier in the\n>> >> function could make that more obvious.\n>> >\n>> > `data` is definitely uint32_t aligned, but this is a tradeoff, since\n>> > if we wrote:\n>> >\n>> >     uint32_t *data = xmmap(...);\n>> >\n>> > then I think we would have to change the case where ret is non-zero to be:\n>> >\n>> >     if (data)\n>> >         munmap((void*)data, ...);\n>> >\n>> > and likewise, data_p is const.\n>>\n>> Doing it that way sounds great to me.  That way, the type contains the\n>> information we need up-front and the safety of the cast is obvious in\n>> the place where the cast is needed.\n>>\n>> (Although my understanding is also that in C it's fine to pass a\n>> uint32_t* to a function expecting a void*, so the second cast would\n>> also not be needed.)\n\nI do not think c99 allows this in 100% of cases - specifically if there a const void * involved. gcc does not care. I do not think c89 cares either. I will watch out for it when this is merged.\n--Randall\n\n"},{"id":"456129","messageId":"Yo605oy4gfQmJ+VE@nand.local","threadId":"56996","inReplyTo":"01df01d87082$a0757020$e1605060$@nexbridge.com","subject":"Re: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-25T22:59:50Z","receivedAt":"2022-05-25T23:00:02Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, May 25, 2022 at 05:58:49PM -0400, rsbecker@nexbridge.com wrote:\n> >> > `data` is definitely uint32_t aligned, but this is a tradeoff, since\n> >> > if we wrote:\n> >> >\n> >> >     uint32_t *data = xmmap(...);\n> >> >\n> >> > then I think we would have to change the case where ret is non-zero to be:\n> >> >\n> >> >     if (data)\n> >> >         munmap((void*)data, ...);\n> >> >\n> >> > and likewise, data_p is const.\n> >>\n> >> Doing it that way sounds great to me.  That way, the type contains the\n> >> information we need up-front and the safety of the cast is obvious in\n> >> the place where the cast is needed.\n> >>\n> >> (Although my understanding is also that in C it's fine to pass a\n> >> uint32_t* to a function expecting a void*, so the second cast would\n> >> also not be needed.)\n>\n> I do not think c99 allows this in 100% of cases - specifically if\n> there a const void * involved. gcc does not care. I do not think c89\n> cares either. I will watch out for it when this is merged.\n\nThanks for the heads up. I looked through the results of \"git grep '=\nxmmap'\" to see if we had contemporary examples of either assigning to a\nnon-'void *', or passing a non-'void *' variable to munmap.\n\nLuckily, we have both, so this shouldn't cause a problem. fixup! patch\nincoming shortly...\n\nThanks,\nTaylor\n"},{"id":"456130","messageId":"Yo61aqaQ/tXh+moi@nand.local","threadId":"56996","inReplyTo":"91a9d21b0b7d99023083c0bbb6f91ccdc1782736.1653088640.git.me@ttaylorr.com","subject":"Re: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-25T23:02:02Z","receivedAt":"2022-05-25T23:02:10Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Junio,\n\nOn Fri, May 20, 2022 at 07:17:35PM -0400, Taylor Blau wrote:\n> To store the individual mtimes of objects in a cruft pack, introduce a\n> new `.mtimes` format that can optionally accompany a single pack in the\n> repository.\n\nLike I mentioned in this sub-thread, here is a small fixup! to apply on\ntop of this patch when queueing. I'm hoping this will be easier than\nreapplying the dozen+ or so patches in this series (the rest of which\nare unchanged). But if it isn't, please let me know and I can send you a\nreroll of the whole thing.\n\nIn the meantime, here's the fixup...\n\n--- 8< ---\nSubject: [PATCH] fixup! pack-mtimes: support reading .mtimes files\n\nSigned-off-by: Taylor Blau <me@ttaylorr.com>\n---\n object-store.h |  5 +++++\n pack-mtimes.c  | 35 +++++++++++++++++++----------------\n pack-mtimes.h  | 11 +++++++++++\n 3 files changed, 35 insertions(+), 16 deletions(-)\n\ndiff --git a/object-store.h b/object-store.h\nindex 2c4671ed7a..05cc9a33ed 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -122,6 +122,11 @@ struct packed_git {\n \tconst uint32_t *revindex_data;\n \tconst uint32_t *revindex_map;\n \tsize_t revindex_size;\n+\t/*\n+\t * mtimes_map points at the beginning of the memory mapped region of\n+\t * this pack's corresponding .mtimes file, and mtimes_size is the size\n+\t * of that .mtimes file\n+\t */\n \tconst uint32_t *mtimes_map;\n \tsize_t mtimes_size;\n \t/* something like \".git/objects/pack/xxxxx.pack\" */\ndiff --git a/pack-mtimes.c b/pack-mtimes.c\nindex 46ad584af1..0e0aafdcb0 100644\n--- a/pack-mtimes.c\n+++ b/pack-mtimes.c\n@@ -1,3 +1,4 @@\n+#include \"git-compat-util.h\"\n #include \"pack-mtimes.h\"\n #include \"object-store.h\"\n #include \"packfile.h\"\n@@ -7,12 +8,10 @@ static char *pack_mtimes_filename(struct packed_git *p)\n \tsize_t len;\n \tif (!strip_suffix(p->pack_name, \".pack\", &len))\n \t\tBUG(\"pack_name does not end in .pack\");\n-\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n \treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n }\n\n #define MTIMES_HEADER_SIZE (12)\n-#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n\n struct mtimes_header {\n \tuint32_t signature;\n@@ -26,10 +25,9 @@ static int load_pack_mtimes_file(char *mtimes_file,\n {\n \tint fd, ret = 0;\n \tstruct stat st;\n-\tvoid *data = NULL;\n-\tsize_t mtimes_size;\n+\tuint32_t *data = NULL;\n+\tsize_t mtimes_size, expected_size;\n \tstruct mtimes_header header;\n-\tuint32_t *hdr;\n\n \tfd = git_open(mtimes_file);\n\n@@ -44,21 +42,16 @@ static int load_pack_mtimes_file(char *mtimes_file,\n\n \tmtimes_size = xsize_t(st.st_size);\n\n-\tif (mtimes_size < MTIMES_MIN_SIZE) {\n+\tif (mtimes_size < MTIMES_HEADER_SIZE) {\n \t\tret = error(_(\"mtimes file %s is too small\"), mtimes_file);\n \t\tgoto cleanup;\n \t}\n\n-\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t), num_objects)) {\n-\t\tret = error(_(\"mtimes file %s is corrupt\"), mtimes_file);\n-\t\tgoto cleanup;\n-\t}\n+\tdata = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n\n-\tdata = hdr = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n-\n-\theader.signature = ntohl(hdr[0]);\n-\theader.version = ntohl(hdr[1]);\n-\theader.hash_id = ntohl(hdr[2]);\n+\theader.signature = ntohl(data[0]);\n+\theader.version = ntohl(data[1]);\n+\theader.hash_id = ntohl(data[2]);\n\n \tif (header.signature != MTIMES_SIGNATURE) {\n \t\tret = error(_(\"mtimes file %s has unknown signature\"), mtimes_file);\n@@ -77,13 +70,23 @@ static int load_pack_mtimes_file(char *mtimes_file,\n \t\tgoto cleanup;\n \t}\n\n+\n+\texpected_size = MTIMES_HEADER_SIZE;\n+\texpected_size = st_add(expected_size, st_mult(sizeof(uint32_t), num_objects));\n+\texpected_size = st_add(expected_size, 2 * (header.hash_id == 1 ? GIT_SHA1_RAWSZ : GIT_SHA256_RAWSZ));\n+\n+\tif (mtimes_size != expected_size) {\n+\t\tret = error(_(\"mtimes file %s is corrupt\"), mtimes_file);\n+\t\tgoto cleanup;\n+\t}\n+\n cleanup:\n \tif (ret) {\n \t\tif (data)\n \t\t\tmunmap(data, mtimes_size);\n \t} else {\n \t\t*len_p = mtimes_size;\n-\t\t*data_p = (const uint32_t *)data;\n+\t\t*data_p = data;\n \t}\n\n \tclose(fd);\ndiff --git a/pack-mtimes.h b/pack-mtimes.h\nindex 38ddb9f893..cc957b3e85 100644\n--- a/pack-mtimes.h\n+++ b/pack-mtimes.h\n@@ -8,8 +8,19 @@\n\n struct packed_git;\n\n+/*\n+ * Loads the .mtimes file corresponding to \"p\", if any, returning zero\n+ * on success.\n+ */\n int load_pack_mtimes(struct packed_git *p);\n\n+/* Returns the mtime associated with the object at position \"pos\" (in\n+ * lexicographic/index order) in pack \"p\".\n+ *\n+ * Note that it is a BUG() to call this function if either (a) \"p\" does\n+ * not have a corresponding .mtimes file, or (b) it does, but it hasn't\n+ * been loaded\n+ */\n uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos);\n\n #endif\n--\n2.36.1.94.gb0d54bedca\n\n--- >8 ---\n"},{"id":"456132","messageId":"220526.86sfox6tvp.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"Yo6b+8sixGAqMm/x@nand.local","subject":"Re: adding new 32-bit on-disk (unsigned) timestamp formats (was: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files)","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-26T00:02:39Z","receivedAt":"2022-05-26T00:06:58Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 25 2022, Taylor Blau wrote:\n\n> On Wed, May 25, 2022 at 09:30:55AM -0400, Derrick Stolee wrote:\n>> On 5/25/2022 5:11 AM, Ævar Arnfjörð Bjarmason wrote:\n>> > I must say that I really don't like this part of the format. Is it\n>> > really necessary to optimize the storage space here in a way that leaves\n>> > open questions about future time_t compatibility, and having to\n>> > introduce the first use of unsigned 32 bit timestamps to git's codebase?\n>>\n>> The commit-graph file format uses unsigned 34-bit timestamps (packed\n>> with 30-bit topological levels in the CDAT chunk), so this \"not-64-bit\n>> signed timestamps\" thing is something we've done before.\n>>\n>> > Yes, this is its own self-contained format, so we don't *need* time_t\n>> > here, but it's also really handy if we can eventually consistently use\n>> > 64 time_t everywhere and not worry about any compatibility issues, or\n>> > unsigned v.s. signed, or to create our own little ext4-like signed 32\n>> > bit timestamp format.\n>>\n>> We can also use a new file format version when it is necessary. We\n>> have a lot of time to add that detail without overly complicating the\n>> format right now.\n>>\n>> > If we really are trying to micro-optimize storage space here I'm willing\n>> > to bet that this is still a bad/premature optimization. There's much\n>> > better ways to store this sort of data in a compact way if that's the\n>> > concern. E.g. you'd store a 64 bit \"base\" timestamp in the header for\n>> > the first entry, and have smaller (signed) \"delta\" timestamps storing\n>> > offsets from that \"base\" timestamp.\n>>\n>> This is a good idea for a v2 format when that is necessary.\n>\n> I agree here.\n>\n> I'm not opposed to such a change (or even being the one to work on it!),\n> but I would encourage us to pursue that change outside of this series,\n> since it can easily be done on top.\n>\n> Of course, if we ever did decide to implement 64-bit mtimes, we would\n> have to maintain support for reading both the 32-bit and 64-bit values.\n> But I think the code is well-equipped to do that, and it could be done\n> on top without significant additional complexity.\n\nDo you mean \"on top\" in the sense that we'd expect that before the next\nrelease, so that we wouldn't need to deal with bumping the format, and\nhave some phase-out period for the older version etc.\n\nOr that we would need to treat what's landing here as something we'll\nneed to support going forward?\n\nI think if a format change is worthwhile doing at all that it's worth\njust doing it now if it's going to be the latter of those, as changing\nfile formats before they're in the wild is easy, but after that it's at\nbest a bit tedious. E.g. we'll need testing to see how we deal with\nmixed new/old format files etc. etc.\n"},{"id":"456133","messageId":"220526.86o7zl6tpo.gmgdl@evledraar.gmail.com","threadId":"56996","inReplyTo":"Yo6bDC8uivC3gM2o@nand.local","subject":"Re: [PATCH v5 00/17] cruft packs","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-26T00:06:09Z","receivedAt":"2022-05-26T00:09:16Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 25 2022, Taylor Blau wrote:\n\n> On Wed, May 25, 2022 at 03:59:24PM -0400, Derrick Stolee wrote:\n>> I'd much rather have a consistent and proven way of specifying the\n>> hash value (using the oid_version() helper) than to try and make a\n>> new mechanism.\n>\n> To be clear, I absolutely don't think any of us should have the attitude\n> of repeating past bad decisions for the sake of consistency.\n>\n> As best I can tell, our (Jonathan and I's) disagreement is on whether\n> using \"1\" and \"2\" to identify which hash function is used by the .mtimes\n> file is OK or not. I happen to think that it is acceptable, so the\n> choice to continue to adopt this pattern was motivated by being\n> consistent with a pattern that is good and works.\n\nI don't have a strong opinion on whether we \"bless\" that or not, and say\nthat we should just use 1, 2 etc. going forward or not.\n\nBut I do think that us doing so initially wasn't intentional, and has\nbeen in opposition to a strongly worded claim in a comment in hash.h\n(which I modified in my earlier related RFC series).\n\nSo maybe not part of this series, but it seems prudent if you feel\nstrongly about using this for new formats over what hash.h is currently\nrecommending that we have some patch sooner than later to update it\naccordingly.\n"},{"id":"456134","messageId":"Yo7F4tj8aGMvwM7/@nand.local","threadId":"56996","inReplyTo":"220526.86sfox6tvp.gmgdl@evledraar.gmail.com","subject":"Re: adding new 32-bit on-disk (unsigned) timestamp formats (was: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files)","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-05-26T00:12:18Z","receivedAt":"2022-05-26T00:12:29Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Thu, May 26, 2022 at 02:02:39AM +0200, Ævar Arnfjörð Bjarmason wrote:\n> >> > If we really are trying to micro-optimize storage space here I'm willing\n> >> > to bet that this is still a bad/premature optimization. There's much\n> >> > better ways to store this sort of data in a compact way if that's the\n> >> > concern. E.g. you'd store a 64 bit \"base\" timestamp in the header for\n> >> > the first entry, and have smaller (signed) \"delta\" timestamps storing\n> >> > offsets from that \"base\" timestamp.\n> >>\n> >> This is a good idea for a v2 format when that is necessary.\n> >\n> > I agree here.\n> >\n> > I'm not opposed to such a change (or even being the one to work on it!),\n> > but I would encourage us to pursue that change outside of this series,\n> > since it can easily be done on top.\n> >\n> > Of course, if we ever did decide to implement 64-bit mtimes, we would\n> > have to maintain support for reading both the 32-bit and 64-bit values.\n> > But I think the code is well-equipped to do that, and it could be done\n> > on top without significant additional complexity.\n>\n> Do you mean \"on top\" in the sense that we'd expect that before the next\n> release, so that we wouldn't need to deal with bumping the format, and\n> have some phase-out period for the older version etc.\n>\n> Or that we would need to treat what's landing here as something we'll\n> need to support going forward?\n\nMy plan is to treat what will hopefully land here as something we're\ngoing to support.\n\nI meant \"on top\" in the sense that the format implemented here does not\nrestrict us against making changes (like adding support for wider\nrecords) in the future. IOW, I did not mean to suggest that we should\nexpect more patches from me in this cycle to deprecate parts of the v1\nformat.\n\nIn other words (again ;-)), I would like to see us ship this format with\nthe existing 32-bit records.\n\n> I think if a format change is worthwhile doing at all that it's worth\n> just doing it now if it's going to be the latter of those, as changing\n> file formats before they're in the wild is easy, but after that it's at\n> best a bit tedious. E.g. we'll need testing to see how we deal with\n> mixed new/old format files etc. etc.\n\nI can understand where you're coming from, though as I noted earlier in\nthe thread, I don't think changing the format in the manner you suggest\nwould be that difficult in practice.\n\nBut in the meantime, the existing format is useful and works, and I\ndon't think we should go back to the drawing board for something that we\ncan do later if we decide to.\n\nThanks,\nTaylor\n"},{"id":"456139","messageId":"xmqqleup3zkt.fsf@gitster.g","threadId":"56996","inReplyTo":"Yo61aqaQ/tXh+moi@nand.local","subject":"Re: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-26T00:30:26Z","receivedAt":"2022-05-26T00:31:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n> diff --git a/pack-mtimes.c b/pack-mtimes.c\n> index 46ad584af1..0e0aafdcb0 100644\n> --- a/pack-mtimes.c\n> +++ b/pack-mtimes.c\n> @@ -1,3 +1,4 @@\n> +#include \"git-compat-util.h\"\n>  #include \"pack-mtimes.h\"\n>  #include \"object-store.h\"\n>  #include \"packfile.h\"\n> @@ -7,12 +8,10 @@ static char *pack_mtimes_filename(struct packed_git *p)\n>  \tsize_t len;\n>  \tif (!strip_suffix(p->pack_name, \".pack\", &len))\n>  \t\tBUG(\"pack_name does not end in .pack\");\n> -\t/* NEEDSWORK: this could reuse code from pack-revindex.c. */\n>  \treturn xstrfmt(\"%.*s.mtimes\", (int)len, p->pack_name);\n>  }\n>\n>  #define MTIMES_HEADER_SIZE (12)\n> -#define MTIMES_MIN_SIZE (MTIMES_HEADER_SIZE + (2 * the_hash_algo->rawsz))\n>\n>  struct mtimes_header {\n>  \tuint32_t signature;\n> @@ -26,10 +25,9 @@ static int load_pack_mtimes_file(char *mtimes_file,\n>  {\n>  \tint fd, ret = 0;\n>  \tstruct stat st;\n> -\tvoid *data = NULL;\n> -\tsize_t mtimes_size;\n> +\tuint32_t *data = NULL;\n> +\tsize_t mtimes_size, expected_size;\n>  \tstruct mtimes_header header;\n> -\tuint32_t *hdr;\n>\n>  \tfd = git_open(mtimes_file);\n>\n> @@ -44,21 +42,16 @@ static int load_pack_mtimes_file(char *mtimes_file,\n>\n>  \tmtimes_size = xsize_t(st.st_size);\n>\n> -\tif (mtimes_size < MTIMES_MIN_SIZE) {\n> +\tif (mtimes_size < MTIMES_HEADER_SIZE) {\n>  \t\tret = error(_(\"mtimes file %s is too small\"), mtimes_file);\n>  \t\tgoto cleanup;\n>  \t}\n>\n> -\tif (mtimes_size - MTIMES_MIN_SIZE != st_mult(sizeof(uint32_t), num_objects)) {\n> -\t\tret = error(_(\"mtimes file %s is corrupt\"), mtimes_file);\n> -\t\tgoto cleanup;\n> -\t}\n> +\tdata = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n>\n> -\tdata = hdr = xmmap(NULL, mtimes_size, PROT_READ, MAP_PRIVATE, fd, 0);\n> -\n> -\theader.signature = ntohl(hdr[0]);\n> -\theader.version = ntohl(hdr[1]);\n> -\theader.hash_id = ntohl(hdr[2]);\n> +\theader.signature = ntohl(data[0]);\n> +\theader.version = ntohl(data[1]);\n> +\theader.hash_id = ntohl(data[2]);\n\nSo, instead of assuming that the size of the file is cast in stone\nto match the size the current implementation happens to give and\nreject a file from a future version, we check the header first to\ngive a more readable error when we see a version of the file that\nwe do not understand.\n\nMakes sense.\n\nAt least, \"here is a small fixup!\" should have been accompanied by a\nbrief explanation to say something like that, i.e. why a fixup is\nneeded, what shortcoming in the original it is meant to address,\netc.\n\nWill queue between 2/17 and 3/17 without squashing (yet).\n\nThanks.\n\n>  \tif (header.signature != MTIMES_SIGNATURE) {\n>  \t\tret = error(_(\"mtimes file %s has unknown signature\"), mtimes_file);\n> @@ -77,13 +70,23 @@ static int load_pack_mtimes_file(char *mtimes_file,\n>  \t\tgoto cleanup;\n>  \t}\n>\n> +\n> +\texpected_size = MTIMES_HEADER_SIZE;\n> +\texpected_size = st_add(expected_size, st_mult(sizeof(uint32_t), num_objects));\n> +\texpected_size = st_add(expected_size, 2 * (header.hash_id == 1 ? GIT_SHA1_RAWSZ : GIT_SHA256_RAWSZ));\n> +\n> +\tif (mtimes_size != expected_size) {\n> +\t\tret = error(_(\"mtimes file %s is corrupt\"), mtimes_file);\n> +\t\tgoto cleanup;\n> +\t}\n> +\n>  cleanup:\n>  \tif (ret) {\n>  \t\tif (data)\n>  \t\t\tmunmap(data, mtimes_size);\n>  \t} else {\n>  \t\t*len_p = mtimes_size;\n> -\t\t*data_p = (const uint32_t *)data;\n> +\t\t*data_p = data;\n>  \t}\n>\n>  \tclose(fd);\n> diff --git a/pack-mtimes.h b/pack-mtimes.h\n> index 38ddb9f893..cc957b3e85 100644\n> --- a/pack-mtimes.h\n> +++ b/pack-mtimes.h\n> @@ -8,8 +8,19 @@\n>\n>  struct packed_git;\n>\n> +/*\n> + * Loads the .mtimes file corresponding to \"p\", if any, returning zero\n> + * on success.\n> + */\n>  int load_pack_mtimes(struct packed_git *p);\n>\n> +/* Returns the mtime associated with the object at position \"pos\" (in\n> + * lexicographic/index order) in pack \"p\".\n> + *\n> + * Note that it is a BUG() to call this function if either (a) \"p\" does\n> + * not have a corresponding .mtimes file, or (b) it does, but it hasn't\n> + * been loaded\n> + */\n>  uint32_t nth_packed_mtime(struct packed_git *p, uint32_t pos);\n>\n>  #endif\n> --\n> 2.36.1.94.gb0d54bedca\n>\n> --- >8 ---\n"},{"id":"457521","messageId":"157741e2-cd06-9304-bb21-c67c2cbd923e@web.de","threadId":"56996","inReplyTo":"43c14eec0762170393c5e9681c3d5ef8fa60c96c.1653088640.git.me@ttaylorr.com","subject":"Re: [PATCH v5 16/17] builtin/gc.c: conditionally avoid pruning objects via loose","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-06-19T05:38:50Z","receivedAt":"2022-06-19T05:39:16Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 21.05.22 um 01:18 schrieb Taylor Blau:\n> diff --git a/Documentation/git-gc.txt b/Documentation/git-gc.txt\n> index 853967dea0..ba4e67700e 100644\n> --- a/Documentation/git-gc.txt\n> +++ b/Documentation/git-gc.txt\n> @@ -54,6 +54,11 @@ other housekeeping tasks (e.g. rerere, working trees, reflog...) will\n>  be performed as well.\n>\n>\n> +--cruft::\n> +\tWhen expiring unreachable objects, pack them separately into a\n> +\tcruft pack instead of storing the loose objects as loose\n> +\tobjects.\n\nThe last part looks tautological.  How about:\n\n--- >8 ---\nSubject: [PATCH] gc: simplify --cruft description\n\nRemove duplicate \"loose objects\".\n\nSigned-off-by: René Scharfe <l.s.r@web.de>\n---\n Documentation/git-gc.txt | 3 +--\n 1 file changed, 1 insertion(+), 2 deletions(-)\n\ndiff --git a/Documentation/git-gc.txt b/Documentation/git-gc.txt\nindex ba4e67700e..0af7540a0c 100644\n--- a/Documentation/git-gc.txt\n+++ b/Documentation/git-gc.txt\n@@ -56,8 +56,7 @@ be performed as well.\n\n --cruft::\n \tWhen expiring unreachable objects, pack them separately into a\n-\tcruft pack instead of storing the loose objects as loose\n-\tobjects.\n+\tcruft pack instead of storing them as loose objects.\n\n --prune=<date>::\n \tPrune loose objects older than date (default is 2 weeks ago,\n--\n2.36.1\n"},{"id":"457624","messageId":"xmqqv8suhuty.fsf@gitster.g","threadId":"56996","inReplyTo":"157741e2-cd06-9304-bb21-c67c2cbd923e@web.de","subject":"Re: [PATCH v5 16/17] builtin/gc.c: conditionally avoid pruning objects via loose","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-21T15:58:33Z","receivedAt":"2022-06-21T15:59:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> Am 21.05.22 um 01:18 schrieb Taylor Blau:\n>> diff --git a/Documentation/git-gc.txt b/Documentation/git-gc.txt\n>> index 853967dea0..ba4e67700e 100644\n>> --- a/Documentation/git-gc.txt\n>> +++ b/Documentation/git-gc.txt\n>> @@ -54,6 +54,11 @@ other housekeeping tasks (e.g. rerere, working trees, reflog...) will\n>>  be performed as well.\n>>\n>>\n>> +--cruft::\n>> +\tWhen expiring unreachable objects, pack them separately into a\n>> +\tcruft pack instead of storing the loose objects as loose\n>> +\tobjects.\n>\n> The last part looks tautological.  How about:\n>\n> --- >8 ---\n> Subject: [PATCH] gc: simplify --cruft description\n>\n> Remove duplicate \"loose objects\".\n>\n> Signed-off-by: René Scharfe <l.s.r@web.de>\n> ---\n\nSounds good.  Will apply.\n\nThanks.\n\n>  Documentation/git-gc.txt | 3 +--\n>  1 file changed, 1 insertion(+), 2 deletions(-)\n>\n> diff --git a/Documentation/git-gc.txt b/Documentation/git-gc.txt\n> index ba4e67700e..0af7540a0c 100644\n> --- a/Documentation/git-gc.txt\n> +++ b/Documentation/git-gc.txt\n> @@ -56,8 +56,7 @@ be performed as well.\n>\n>  --cruft::\n>  \tWhen expiring unreachable objects, pack them separately into a\n> -\tcruft pack instead of storing the loose objects as loose\n> -\tobjects.\n> +\tcruft pack instead of storing them as loose objects.\n>\n>  --prune=<date>::\n>  \tPrune loose objects older than date (default is 2 weeks ago,\n> --\n> 2.36.1\n"},{"id":"477873","messageId":"mvmv8g78hfo.fsf@suse.de","threadId":"56996","inReplyTo":"91a9d21b0b7d99023083c0bbb6f91ccdc1782736.1653088640.git.me@ttaylorr.com","subject":"Re: [PATCH v5 02/17] pack-mtimes: support reading .mtimes files","fromName":"Andreas Schwab","fromEmail":"schwab@suse.de","sentAt":"2023-06-01T13:01:47Z","receivedAt":"2023-06-01T13:01:52Z","isPatch":true,"sender":{"key":"schwab@suse.de","avatar":"https://avatars.githubusercontent.com/u/2175493?v=4"},"body":"On Mai 20 2022, Taylor Blau wrote:\n\n> diff --git a/Documentation/technical/pack-format.txt b/Documentation/technical/pack-format.txt\n> index 6d3efb7d16..b520aa9c45 100644\n> --- a/Documentation/technical/pack-format.txt\n> +++ b/Documentation/technical/pack-format.txt\n> @@ -294,6 +294,25 @@ Pack file entry: <+\n>  \n>  All 4-byte numbers are in network order.\n>  \n> +== pack-*.mtimes files have the format:\n> +\n> +All 4-byte numbers are in network byte order.\n> +\n> +  - A 4-byte magic number '0x4d544d45' ('MTME').\n\nThis is identified by file(1) as \"Multitracker Version 4.05\". ;-)\n\n-- \nAndreas Schwab, SUSE Labs, schwab@suse.de\nGPG Key fingerprint = 0196 BAD8 1CE9 1970 F4BE  1748 E4D4 88E3 0EEA B9D7\n\"And now for something completely different.\"\n"}]}