{"thread":{"id":"35335","subject":"[PATCH v3 0/21] pack bitmaps","startedAt":"2013-11-14T12:41:58Z","lastAt":"2013-12-21T13:40:17Z","messageCount":55,"participants":["Jeff King","Ramsay Jones","Thomas Rast","Karsten Blees","Junio C Hamano"],"isPatch":true,"patchVersion":3,"patchTotal":21},"messages":[{"id":"230598","messageId":"20131114124157.GA23784@sigill.intra.peff.net","threadId":"35335","inReplyTo":null,"subject":"[PATCH v3 0/21] pack bitmaps","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:41:58Z","receivedAt":"2013-11-14T12:41:58Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"Here's another iteration of the pack bitmaps series. Compared to v2, it\nchanges:\n\n - misc style/typo fixes\n\n - portability fixes from Ramsay and Torsten\n\n - count-objects garbage-reporting patch from Duy\n\n - disable bitmaps when is_repository_shallow(); this also covers the\n   case where the client is shallow, since we feed pack-objects a\n   --shallow-file in that case. This used to done by checking\n   !internal_rev_list, but that doesn't apply after cdab485.\n\n - ewah sources now properly use git-compat-util.h and do not include\n   system headers\n\n - the ewah code uses ewah_malloc, ewah_realloc, and so forth to let the\n   project use a particular allocator (and we want to use xmalloc and\n   friends). And we defined those in pack-bitmap.h, but of course that\n   had no effect on the ewah/*.c files that did not include\n   pack-bitmap.h.  Since we are hacking up and git-ifying libewok\n   anyway, we can just set the hardcoded fallback to xmalloc instead of\n   malloc.\n\n  - the ewah code used gcc's __builtin_ctzll, but did not provide a\n    suitable fallback. We now provide a fallback in C.\n\n  - The bitmap reading code only handles a single bitmapped pack (since\n    they must be fully closed, there is not much point in having\n    multiple). It used to silently ignore extra bitmap indices it found,\n    but will now warn that they are being ignored.\n\n  - The name-hash cache is now optional, controlled by\n    pack.writeBitmapHashCache.\n\n  - The test script will now do basic interoperability testing with jgit\n    (if you have jgit in your $PATH).\n\n  - There are now perf tests. Spoiler alert: bitmaps make clones faster.\n    See patch 20 for details. We can also measure the speedup from the\n    hash cache (see patch 21).\n\nNot addressed:\n\n  - I did not include the NEEDS_ALIGNED_ACCESS patch. I note that we do\n    not even have a Makefile knob for this, and the code in read-cache.c\n    has probably never actually been used. Are there real systems that\n    have a problem? The read-cache code was in support of the index v4\n    experiment, which did away with the 8-byte padding. So it could be\n    that we simply don't see it, because everything is currently\n    aligned.\n\n  - On a related note, we do some cast-buffer-to-struct magic on the\n    mmap'd file. I note that the regular packfile reader also does this.\n    How careful do we want to be?\n\n  - We still assume that reusing a slice from the front of the pack will\n    never miss delta bases. This is the case currently for packs\n    generated by both git and JGit, but it would be nice to mark the\n    property in the bitmap index. Adding a new flag would break JGit\n    compatibility, though. We can either make it an option, or assume\n    it's good enough for now and worry about it in v2.\n\n  [01/21]: sha1write: make buffer const-correct\n  [02/21]: revindex: Export new APIs\n  [03/21]: pack-objects: Refactor the packing list\n  [04/21]: pack-objects: factor out name_hash\n  [05/21]: revision: allow setting custom limiter function\n  [06/21]: sha1_file: export `git_open_noatime`\n  [07/21]: compat: add endianness helpers\n  [08/21]: ewah: compressed bitmap implementation\n  [09/21]: documentation: add documentation for the bitmap format\n  [10/21]: pack-bitmap: add support for bitmap indexes\n  [11/21]: pack-objects: use bitmaps when packing objects\n  [12/21]: rev-list: add bitmap mode to speed up object lists\n  [13/21]: pack-objects: implement bitmap writing\n  [14/21]: repack: stop using magic number for ARRAY_SIZE(exts)\n  [15/21]: repack: turn exts array into array-of-struct\n  [16/21]: repack: handle optional files created by pack-objects\n  [17/21]: repack: consider bitmaps when performing repacks\n  [18/21]: count-objects: recognize .bitmap in garbage-checking\n  [19/21]: t: add basic bitmap functionality tests\n  [20/21]: t/perf: add tests for pack bitmaps\n  [21/21]: pack-bitmap: implement optional name_hash cache\n\n-Peff\n"},{"id":"230599","messageId":"20131114124249.GA10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 01/21] sha1write: make buffer const-correct","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:42:50Z","receivedAt":"2013-11-14T12:42:50Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"We are passed a \"void *\" and write it out without ever\ntouching it; let's indicate that by using \"const\".\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n csum-file.c | 6 +++---\n csum-file.h | 2 +-\n 2 files changed, 4 insertions(+), 4 deletions(-)\n\ndiff --git a/csum-file.c b/csum-file.c\nindex 53f5375..465971c 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -11,7 +11,7 @@\n #include \"progress.h\"\n #include \"csum-file.h\"\n \n-static void flush(struct sha1file *f, void *buf, unsigned int count)\n+static void flush(struct sha1file *f, const void *buf, unsigned int count)\n {\n \tif (0 <= f->check_fd && count)  {\n \t\tunsigned char check_buffer[8192];\n@@ -86,13 +86,13 @@ int sha1close(struct sha1file *f, unsigned char *result, unsigned int flags)\n \treturn fd;\n }\n \n-int sha1write(struct sha1file *f, void *buf, unsigned int count)\n+int sha1write(struct sha1file *f, const void *buf, unsigned int count)\n {\n \twhile (count) {\n \t\tunsigned offset = f->offset;\n \t\tunsigned left = sizeof(f->buffer) - offset;\n \t\tunsigned nr = count > left ? left : count;\n-\t\tvoid *data;\n+\t\tconst void *data;\n \n \t\tif (f->do_crc)\n \t\t\tf->crc32 = crc32(f->crc32, buf, nr);\ndiff --git a/csum-file.h b/csum-file.h\nindex 3b540bd..9dedb03 100644\n--- a/csum-file.h\n+++ b/csum-file.h\n@@ -34,7 +34,7 @@ extern struct sha1file *sha1fd(int fd, const char *name);\n extern struct sha1file *sha1fd_check(const char *name);\n extern struct sha1file *sha1fd_throughput(int fd, const char *name, struct progress *tp);\n extern int sha1close(struct sha1file *, unsigned char *, unsigned int);\n-extern int sha1write(struct sha1file *, void *, unsigned int);\n+extern int sha1write(struct sha1file *, const void *, unsigned int);\n extern void sha1flush(struct sha1file *f);\n extern void crc32_begin(struct sha1file *);\n extern uint32_t crc32_end(struct sha1file *);\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230600","messageId":"20131114124254.GB10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 02/21] revindex: Export new APIs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:42:54Z","receivedAt":"2013-11-14T12:42:54Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nAllow users to efficiently lookup consecutive entries that are expected\nto be found on the same revindex by exporting `find_revindex_position`:\nthis function takes a pointer to revindex itself, instead of looking up\nthe proper revindex for a given packfile on each call.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n pack-revindex.c | 38 +++++++++++++++++++++++++-------------\n pack-revindex.h |  8 ++++++++\n 2 files changed, 33 insertions(+), 13 deletions(-)\n\ndiff --git a/pack-revindex.c b/pack-revindex.c\nindex b4d2b35..0bb13b1 100644\n--- a/pack-revindex.c\n+++ b/pack-revindex.c\n@@ -16,11 +16,6 @@\n  * get the object sha1 from the main index.\n  */\n \n-struct pack_revindex {\n-\tstruct packed_git *p;\n-\tstruct revindex_entry *revindex;\n-};\n-\n static struct pack_revindex *pack_revindex;\n static int pack_revindex_hashsz;\n \n@@ -201,15 +196,14 @@ static void create_pack_revindex(struct pack_revindex *rix)\n \tsort_revindex(rix->revindex, num_ent, p->pack_size);\n }\n \n-struct revindex_entry *find_pack_revindex(struct packed_git *p, off_t ofs)\n+struct pack_revindex *revindex_for_pack(struct packed_git *p)\n {\n \tint num;\n-\tunsigned lo, hi;\n \tstruct pack_revindex *rix;\n-\tstruct revindex_entry *revindex;\n \n \tif (!pack_revindex_hashsz)\n \t\tinit_pack_revindex();\n+\n \tnum = pack_revindex_ix(p);\n \tif (num < 0)\n \t\tdie(\"internal error: pack revindex fubar\");\n@@ -217,21 +211,39 @@ struct revindex_entry *find_pack_revindex(struct packed_git *p, off_t ofs)\n \trix = &pack_revindex[num];\n \tif (!rix->revindex)\n \t\tcreate_pack_revindex(rix);\n-\trevindex = rix->revindex;\n \n-\tlo = 0;\n-\thi = p->num_objects + 1;\n+\treturn rix;\n+}\n+\n+int find_revindex_position(struct pack_revindex *pridx, off_t ofs)\n+{\n+\tint lo = 0;\n+\tint hi = pridx->p->num_objects + 1;\n+\tstruct revindex_entry *revindex = pridx->revindex;\n+\n \tdo {\n \t\tunsigned mi = lo + (hi - lo) / 2;\n \t\tif (revindex[mi].offset == ofs) {\n-\t\t\treturn revindex + mi;\n+\t\t\treturn mi;\n \t\t} else if (ofs < revindex[mi].offset)\n \t\t\thi = mi;\n \t\telse\n \t\t\tlo = mi + 1;\n \t} while (lo < hi);\n+\n \terror(\"bad offset for revindex\");\n-\treturn NULL;\n+\treturn -1;\n+}\n+\n+struct revindex_entry *find_pack_revindex(struct packed_git *p, off_t ofs)\n+{\n+\tstruct pack_revindex *pridx = revindex_for_pack(p);\n+\tint pos = find_revindex_position(pridx, ofs);\n+\n+\tif (pos < 0)\n+\t\treturn NULL;\n+\n+\treturn pridx->revindex + pos;\n }\n \n void discard_revindex(void)\ndiff --git a/pack-revindex.h b/pack-revindex.h\nindex 8d5027a..866ca9c 100644\n--- a/pack-revindex.h\n+++ b/pack-revindex.h\n@@ -6,6 +6,14 @@ struct revindex_entry {\n \tunsigned int nr;\n };\n \n+struct pack_revindex {\n+\tstruct packed_git *p;\n+\tstruct revindex_entry *revindex;\n+};\n+\n+struct pack_revindex *revindex_for_pack(struct packed_git *p);\n+int find_revindex_position(struct pack_revindex *pridx, off_t ofs);\n+\n struct revindex_entry *find_pack_revindex(struct packed_git *p, off_t ofs);\n void discard_revindex(void);\n \n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230601","messageId":"20131114124258.GC10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 03/21] pack-objects: Refactor the packing list","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:42:59Z","receivedAt":"2013-11-14T12:42:59Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nThe hash table that stores the packing list for a given `pack-objects`\nrun was tightly coupled to the pack-objects code.\n\nIn this commit, we refactor the hash table and the underlying storage\narray into a `packing_data` struct. The functionality for accessing and\nadding entries to the packing list is hence accessible from other parts\nof Git besides the `pack-objects` builtin.\n\nThis refactoring is a requirement for further patches in this series\nthat will require accessing the commit packing list from outside of\n`pack-objects`.\n\nThe hash table implementation has been minimally altered: we now\nuse table sizes which are always a power of two, to ensure a uniform\nindex distribution in the array.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Makefile               |   2 +\n builtin/pack-objects.c | 175 +++++++++++--------------------------------------\n pack-objects.c         | 111 +++++++++++++++++++++++++++++++\n pack-objects.h         |  47 +++++++++++++\n 4 files changed, 200 insertions(+), 135 deletions(-)\n create mode 100644 pack-objects.c\n create mode 100644 pack-objects.h\n\ndiff --git a/Makefile b/Makefile\nindex af847f8..48ff0bd 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -694,6 +694,7 @@ LIB_H += notes-merge.h\n LIB_H += notes-utils.h\n LIB_H += notes.h\n LIB_H += object.h\n+LIB_H += pack-objects.h\n LIB_H += pack-revindex.h\n LIB_H += pack.h\n LIB_H += parse-options.h\n@@ -831,6 +832,7 @@ LIB_OBJS += notes-merge.o\n LIB_OBJS += notes-utils.o\n LIB_OBJS += object.o\n LIB_OBJS += pack-check.o\n+LIB_OBJS += pack-objects.o\n LIB_OBJS += pack-revindex.o\n LIB_OBJS += pack-write.o\n LIB_OBJS += pager.o\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 36273dd..f3f0cf9 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -14,6 +14,7 @@\n #include \"diff.h\"\n #include \"revision.h\"\n #include \"list-objects.h\"\n+#include \"pack-objects.h\"\n #include \"progress.h\"\n #include \"refs.h\"\n #include \"streaming.h\"\n@@ -25,42 +26,15 @@ static const char *pack_usage[] = {\n \tNULL\n };\n \n-struct object_entry {\n-\tstruct pack_idx_entry idx;\n-\tunsigned long size;\t/* uncompressed size */\n-\tstruct packed_git *in_pack; \t/* already in pack */\n-\toff_t in_pack_offset;\n-\tstruct object_entry *delta;\t/* delta base object */\n-\tstruct object_entry *delta_child; /* deltified objects who bases me */\n-\tstruct object_entry *delta_sibling; /* other deltified objects who\n-\t\t\t\t\t     * uses the same base as me\n-\t\t\t\t\t     */\n-\tvoid *delta_data;\t/* cached delta (uncompressed) */\n-\tunsigned long delta_size;\t/* delta data size (uncompressed) */\n-\tunsigned long z_delta_size;\t/* delta data size (compressed) */\n-\tenum object_type type;\n-\tenum object_type in_pack_type;\t/* could be delta */\n-\tuint32_t hash;\t\t\t/* name hint hash */\n-\tunsigned char in_pack_header_size;\n-\tunsigned preferred_base:1; /*\n-\t\t\t\t    * we do not pack this, but is available\n-\t\t\t\t    * to be used as the base object to delta\n-\t\t\t\t    * objects against.\n-\t\t\t\t    */\n-\tunsigned no_try_delta:1;\n-\tunsigned tagged:1; /* near the very tip of refs */\n-\tunsigned filled:1; /* assigned write-order */\n-};\n-\n /*\n- * Objects we are going to pack are collected in objects array (dynamically\n- * expanded).  nr_objects & nr_alloc controls this array.  They are stored\n- * in the order we see -- typically rev-list --objects order that gives us\n- * nice \"minimum seek\" order.\n+ * Objects we are going to pack are collected in the `to_pack` structure.\n+ * It contains an array (dynamically expanded) of the object data, and a map\n+ * that can resolve SHA1s to their position in the array.\n  */\n-static struct object_entry *objects;\n+static struct packing_data to_pack;\n+\n static struct pack_idx_entry **written_list;\n-static uint32_t nr_objects, nr_alloc, nr_result, nr_written;\n+static uint32_t nr_result, nr_written;\n \n static int non_empty;\n static int reuse_delta = 1, reuse_object = 1;\n@@ -90,21 +64,11 @@ static unsigned long cache_max_small_delta_size = 1000;\n static unsigned long window_memory_limit = 0;\n \n /*\n- * The object names in objects array are hashed with this hashtable,\n- * to help looking up the entry by object name.\n- * This hashtable is built after all the objects are seen.\n- */\n-static int *object_ix;\n-static int object_ix_hashsz;\n-static struct object_entry *locate_object_entry(const unsigned char *sha1);\n-\n-/*\n  * stats\n  */\n static uint32_t written, written_delta;\n static uint32_t reused, reused_delta;\n \n-\n static void *get_delta(struct object_entry *entry)\n {\n \tunsigned long size, base_size, delta_size;\n@@ -553,12 +517,12 @@ static int mark_tagged(const char *path, const unsigned char *sha1, int flag,\n \t\t       void *cb_data)\n {\n \tunsigned char peeled[20];\n-\tstruct object_entry *entry = locate_object_entry(sha1);\n+\tstruct object_entry *entry = packlist_find(&to_pack, sha1, NULL);\n \n \tif (entry)\n \t\tentry->tagged = 1;\n \tif (!peel_ref(path, peeled)) {\n-\t\tentry = locate_object_entry(peeled);\n+\t\tentry = packlist_find(&to_pack, peeled, NULL);\n \t\tif (entry)\n \t\t\tentry->tagged = 1;\n \t}\n@@ -633,9 +597,10 @@ static struct object_entry **compute_write_order(void)\n {\n \tunsigned int i, wo_end, last_untagged;\n \n-\tstruct object_entry **wo = xmalloc(nr_objects * sizeof(*wo));\n+\tstruct object_entry **wo = xmalloc(to_pack.nr_objects * sizeof(*wo));\n+\tstruct object_entry *objects = to_pack.objects;\n \n-\tfor (i = 0; i < nr_objects; i++) {\n+\tfor (i = 0; i < to_pack.nr_objects; i++) {\n \t\tobjects[i].tagged = 0;\n \t\tobjects[i].filled = 0;\n \t\tobjects[i].delta_child = NULL;\n@@ -647,7 +612,7 @@ static struct object_entry **compute_write_order(void)\n \t * Make sure delta_sibling is sorted in the original\n \t * recency order.\n \t */\n-\tfor (i = nr_objects; i > 0;) {\n+\tfor (i = to_pack.nr_objects; i > 0;) {\n \t\tstruct object_entry *e = &objects[--i];\n \t\tif (!e->delta)\n \t\t\tcontinue;\n@@ -665,7 +630,7 @@ static struct object_entry **compute_write_order(void)\n \t * Give the objects in the original recency order until\n \t * we see a tagged tip.\n \t */\n-\tfor (i = wo_end = 0; i < nr_objects; i++) {\n+\tfor (i = wo_end = 0; i < to_pack.nr_objects; i++) {\n \t\tif (objects[i].tagged)\n \t\t\tbreak;\n \t\tadd_to_write_order(wo, &wo_end, &objects[i]);\n@@ -675,7 +640,7 @@ static struct object_entry **compute_write_order(void)\n \t/*\n \t * Then fill all the tagged tips.\n \t */\n-\tfor (; i < nr_objects; i++) {\n+\tfor (; i < to_pack.nr_objects; i++) {\n \t\tif (objects[i].tagged)\n \t\t\tadd_to_write_order(wo, &wo_end, &objects[i]);\n \t}\n@@ -683,7 +648,7 @@ static struct object_entry **compute_write_order(void)\n \t/*\n \t * And then all remaining commits and tags.\n \t */\n-\tfor (i = last_untagged; i < nr_objects; i++) {\n+\tfor (i = last_untagged; i < to_pack.nr_objects; i++) {\n \t\tif (objects[i].type != OBJ_COMMIT &&\n \t\t    objects[i].type != OBJ_TAG)\n \t\t\tcontinue;\n@@ -693,7 +658,7 @@ static struct object_entry **compute_write_order(void)\n \t/*\n \t * And then all the trees.\n \t */\n-\tfor (i = last_untagged; i < nr_objects; i++) {\n+\tfor (i = last_untagged; i < to_pack.nr_objects; i++) {\n \t\tif (objects[i].type != OBJ_TREE)\n \t\t\tcontinue;\n \t\tadd_to_write_order(wo, &wo_end, &objects[i]);\n@@ -702,13 +667,13 @@ static struct object_entry **compute_write_order(void)\n \t/*\n \t * Finally all the rest in really tight order\n \t */\n-\tfor (i = last_untagged; i < nr_objects; i++) {\n+\tfor (i = last_untagged; i < to_pack.nr_objects; i++) {\n \t\tif (!objects[i].filled)\n \t\t\tadd_family_to_write_order(wo, &wo_end, &objects[i]);\n \t}\n \n-\tif (wo_end != nr_objects)\n-\t\tdie(\"ordered %u objects, expected %\"PRIu32, wo_end, nr_objects);\n+\tif (wo_end != to_pack.nr_objects)\n+\t\tdie(\"ordered %u objects, expected %\"PRIu32, wo_end, to_pack.nr_objects);\n \n \treturn wo;\n }\n@@ -724,7 +689,7 @@ static void write_pack_file(void)\n \n \tif (progress > pack_to_stdout)\n \t\tprogress_state = start_progress(\"Writing objects\", nr_result);\n-\twritten_list = xmalloc(nr_objects * sizeof(*written_list));\n+\twritten_list = xmalloc(to_pack.nr_objects * sizeof(*written_list));\n \twrite_order = compute_write_order();\n \n \tdo {\n@@ -740,7 +705,7 @@ static void write_pack_file(void)\n \t\tif (!offset)\n \t\t\tdie_errno(\"unable to write pack header\");\n \t\tnr_written = 0;\n-\t\tfor (; i < nr_objects; i++) {\n+\t\tfor (; i < to_pack.nr_objects; i++) {\n \t\t\tstruct object_entry *e = write_order[i];\n \t\t\tif (write_one(f, e, &offset) == WRITE_ONE_BREAK)\n \t\t\t\tbreak;\n@@ -803,7 +768,7 @@ static void write_pack_file(void)\n \t\t\twritten_list[j]->offset = (off_t)-1;\n \t\t}\n \t\tnr_remaining -= nr_written;\n-\t} while (nr_remaining && i < nr_objects);\n+\t} while (nr_remaining && i < to_pack.nr_objects);\n \n \tfree(written_list);\n \tfree(write_order);\n@@ -813,53 +778,6 @@ static void write_pack_file(void)\n \t\t\twritten, nr_result);\n }\n \n-static int locate_object_entry_hash(const unsigned char *sha1)\n-{\n-\tint i;\n-\tunsigned int ui;\n-\tmemcpy(&ui, sha1, sizeof(unsigned int));\n-\ti = ui % object_ix_hashsz;\n-\twhile (0 < object_ix[i]) {\n-\t\tif (!hashcmp(sha1, objects[object_ix[i] - 1].idx.sha1))\n-\t\t\treturn i;\n-\t\tif (++i == object_ix_hashsz)\n-\t\t\ti = 0;\n-\t}\n-\treturn -1 - i;\n-}\n-\n-static struct object_entry *locate_object_entry(const unsigned char *sha1)\n-{\n-\tint i;\n-\n-\tif (!object_ix_hashsz)\n-\t\treturn NULL;\n-\n-\ti = locate_object_entry_hash(sha1);\n-\tif (0 <= i)\n-\t\treturn &objects[object_ix[i]-1];\n-\treturn NULL;\n-}\n-\n-static void rehash_objects(void)\n-{\n-\tuint32_t i;\n-\tstruct object_entry *oe;\n-\n-\tobject_ix_hashsz = nr_objects * 3;\n-\tif (object_ix_hashsz < 1024)\n-\t\tobject_ix_hashsz = 1024;\n-\tobject_ix = xrealloc(object_ix, sizeof(int) * object_ix_hashsz);\n-\tmemset(object_ix, 0, sizeof(int) * object_ix_hashsz);\n-\tfor (i = 0, oe = objects; i < nr_objects; i++, oe++) {\n-\t\tint ix = locate_object_entry_hash(oe->idx.sha1);\n-\t\tif (0 <= ix)\n-\t\t\tcontinue;\n-\t\tix = -1 - ix;\n-\t\tobject_ix[ix] = i + 1;\n-\t}\n-}\n-\n static uint32_t name_hash(const char *name)\n {\n \tuint32_t c, hash = 0;\n@@ -908,13 +826,12 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n \tstruct object_entry *entry;\n \tstruct packed_git *p, *found_pack = NULL;\n \toff_t found_offset = 0;\n-\tint ix;\n \tuint32_t hash = name_hash(name);\n+\tuint32_t index_pos;\n \n-\tix = nr_objects ? locate_object_entry_hash(sha1) : -1;\n-\tif (ix >= 0) {\n+\tentry = packlist_find(&to_pack, sha1, &index_pos);\n+\tif (entry) {\n \t\tif (exclude) {\n-\t\t\tentry = objects + object_ix[ix] - 1;\n \t\t\tif (!entry->preferred_base)\n \t\t\t\tnr_result--;\n \t\t\tentry->preferred_base = 1;\n@@ -947,14 +864,7 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n \t\t}\n \t}\n \n-\tif (nr_objects >= nr_alloc) {\n-\t\tnr_alloc = (nr_alloc  + 1024) * 3 / 2;\n-\t\tobjects = xrealloc(objects, nr_alloc * sizeof(*entry));\n-\t}\n-\n-\tentry = objects + nr_objects++;\n-\tmemset(entry, 0, sizeof(*entry));\n-\thashcpy(entry->idx.sha1, sha1);\n+\tentry = packlist_alloc(&to_pack, sha1, index_pos);\n \tentry->hash = hash;\n \tif (type)\n \t\tentry->type = type;\n@@ -967,12 +877,7 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n \t\tentry->in_pack_offset = found_offset;\n \t}\n \n-\tif (object_ix_hashsz * 3 <= nr_objects * 4)\n-\t\trehash_objects();\n-\telse\n-\t\tobject_ix[-1 - ix] = nr_objects;\n-\n-\tdisplay_progress(progress_state, nr_objects);\n+\tdisplay_progress(progress_state, to_pack.nr_objects);\n \n \tif (name && no_try_delta(name))\n \t\tentry->no_try_delta = 1;\n@@ -1329,7 +1234,7 @@ static void check_object(struct object_entry *entry)\n \t\t\tbreak;\n \t\t}\n \n-\t\tif (base_ref && (base_entry = locate_object_entry(base_ref))) {\n+\t\tif (base_ref && (base_entry = packlist_find(&to_pack, base_ref, NULL))) {\n \t\t\t/*\n \t\t\t * If base_ref was set above that means we wish to\n \t\t\t * reuse delta data, and we even found that base\n@@ -1403,12 +1308,12 @@ static void get_object_details(void)\n \tuint32_t i;\n \tstruct object_entry **sorted_by_offset;\n \n-\tsorted_by_offset = xcalloc(nr_objects, sizeof(struct object_entry *));\n-\tfor (i = 0; i < nr_objects; i++)\n-\t\tsorted_by_offset[i] = objects + i;\n-\tqsort(sorted_by_offset, nr_objects, sizeof(*sorted_by_offset), pack_offset_sort);\n+\tsorted_by_offset = xcalloc(to_pack.nr_objects, sizeof(struct object_entry *));\n+\tfor (i = 0; i < to_pack.nr_objects; i++)\n+\t\tsorted_by_offset[i] = to_pack.objects + i;\n+\tqsort(sorted_by_offset, to_pack.nr_objects, sizeof(*sorted_by_offset), pack_offset_sort);\n \n-\tfor (i = 0; i < nr_objects; i++) {\n+\tfor (i = 0; i < to_pack.nr_objects; i++) {\n \t\tstruct object_entry *entry = sorted_by_offset[i];\n \t\tcheck_object(entry);\n \t\tif (big_file_threshold < entry->size)\n@@ -2034,7 +1939,7 @@ static int add_ref_tag(const char *path, const unsigned char *sha1, int flag, vo\n \n \tif (!prefixcmp(path, \"refs/tags/\") && /* is a tag? */\n \t    !peel_ref(path, peeled)        && /* peelable? */\n-\t    locate_object_entry(peeled))      /* object packed? */\n+\t    packlist_find(&to_pack, peeled, NULL))      /* object packed? */\n \t\tadd_object_entry(sha1, OBJ_TAG, NULL, 0);\n \treturn 0;\n }\n@@ -2057,14 +1962,14 @@ static void prepare_pack(int window, int depth)\n \tif (!pack_to_stdout)\n \t\tdo_check_packed_object_crc = 1;\n \n-\tif (!nr_objects || !window || !depth)\n+\tif (!to_pack.nr_objects || !window || !depth)\n \t\treturn;\n \n-\tdelta_list = xmalloc(nr_objects * sizeof(*delta_list));\n+\tdelta_list = xmalloc(to_pack.nr_objects * sizeof(*delta_list));\n \tnr_deltas = n = 0;\n \n-\tfor (i = 0; i < nr_objects; i++) {\n-\t\tstruct object_entry *entry = objects + i;\n+\tfor (i = 0; i < to_pack.nr_objects; i++) {\n+\t\tstruct object_entry *entry = to_pack.objects + i;\n \n \t\tif (entry->delta)\n \t\t\t/* This happens if we decided to reuse existing\n@@ -2342,7 +2247,7 @@ static void loosen_unused_packed_objects(struct rev_info *revs)\n \n \t\tfor (i = 0; i < p->num_objects; i++) {\n \t\t\tsha1 = nth_packed_object_sha1(p, i);\n-\t\t\tif (!locate_object_entry(sha1) &&\n+\t\t\tif (!packlist_find(&to_pack, sha1, NULL) &&\n \t\t\t\t!has_sha1_pack_kept_or_nonlocal(sha1))\n \t\t\t\tif (force_object_loose(sha1, p->mtime))\n \t\t\t\t\tdie(\"unable to force loose object\");\ndiff --git a/pack-objects.c b/pack-objects.c\nnew file mode 100644\nindex 0000000..d01d851\n--- /dev/null\n+++ b/pack-objects.c\n@@ -0,0 +1,111 @@\n+#include \"cache.h\"\n+#include \"object.h\"\n+#include \"pack.h\"\n+#include \"pack-objects.h\"\n+\n+static uint32_t locate_object_entry_hash(struct packing_data *pdata,\n+\t\t\t\t\t const unsigned char *sha1,\n+\t\t\t\t\t int *found)\n+{\n+\tuint32_t i, hash, mask = (pdata->index_size - 1);\n+\n+\tmemcpy(&hash, sha1, sizeof(uint32_t));\n+\ti = hash & mask;\n+\n+\twhile (pdata->index[i] > 0) {\n+\t\tuint32_t pos = pdata->index[i] - 1;\n+\n+\t\tif (!hashcmp(sha1, pdata->objects[pos].idx.sha1)) {\n+\t\t\t*found = 1;\n+\t\t\treturn i;\n+\t\t}\n+\n+\t\ti = (i + 1) & mask;\n+\t}\n+\n+\t*found = 0;\n+\treturn i;\n+}\n+\n+static inline uint32_t closest_pow2(uint32_t v)\n+{\n+\tv = v - 1;\n+\tv |= v >> 1;\n+\tv |= v >> 2;\n+\tv |= v >> 4;\n+\tv |= v >> 8;\n+\tv |= v >> 16;\n+\treturn v + 1;\n+}\n+\n+static void rehash_objects(struct packing_data *pdata)\n+{\n+\tuint32_t i;\n+\tstruct object_entry *entry;\n+\n+\tpdata->index_size = closest_pow2(pdata->nr_objects * 3);\n+\tif (pdata->index_size < 1024)\n+\t\tpdata->index_size = 1024;\n+\n+\tpdata->index = xrealloc(pdata->index, sizeof(uint32_t) * pdata->index_size);\n+\tmemset(pdata->index, 0, sizeof(int) * pdata->index_size);\n+\n+\tentry = pdata->objects;\n+\n+\tfor (i = 0; i < pdata->nr_objects; i++) {\n+\t\tint found;\n+\t\tuint32_t ix = locate_object_entry_hash(pdata, entry->idx.sha1, &found);\n+\n+\t\tif (found)\n+\t\t\tdie(\"BUG: Duplicate object in hash\");\n+\n+\t\tpdata->index[ix] = i + 1;\n+\t\tentry++;\n+\t}\n+}\n+\n+struct object_entry *packlist_find(struct packing_data *pdata,\n+\t\t\t\t   const unsigned char *sha1,\n+\t\t\t\t   uint32_t *index_pos)\n+{\n+\tuint32_t i;\n+\tint found;\n+\n+\tif (!pdata->index_size)\n+\t\treturn NULL;\n+\n+\ti = locate_object_entry_hash(pdata, sha1, &found);\n+\n+\tif (index_pos)\n+\t\t*index_pos = i;\n+\n+\tif (!found)\n+\t\treturn NULL;\n+\n+\treturn &pdata->objects[pdata->index[i] - 1];\n+}\n+\n+struct object_entry *packlist_alloc(struct packing_data *pdata,\n+\t\t\t\t    const unsigned char *sha1,\n+\t\t\t\t    uint32_t index_pos)\n+{\n+\tstruct object_entry *new_entry;\n+\n+\tif (pdata->nr_objects >= pdata->nr_alloc) {\n+\t\tpdata->nr_alloc = (pdata->nr_alloc  + 1024) * 3 / 2;\n+\t\tpdata->objects = xrealloc(pdata->objects,\n+\t\t\t\t\t  pdata->nr_alloc * sizeof(*new_entry));\n+\t}\n+\n+\tnew_entry = pdata->objects + pdata->nr_objects++;\n+\n+\tmemset(new_entry, 0, sizeof(*new_entry));\n+\thashcpy(new_entry->idx.sha1, sha1);\n+\n+\tif (pdata->index_size * 3 <= pdata->nr_objects * 4)\n+\t\trehash_objects(pdata);\n+\telse\n+\t\tpdata->index[index_pos] = pdata->nr_objects;\n+\n+\treturn new_entry;\n+}\ndiff --git a/pack-objects.h b/pack-objects.h\nnew file mode 100644\nindex 0000000..f528215\n--- /dev/null\n+++ b/pack-objects.h\n@@ -0,0 +1,47 @@\n+#ifndef PACK_OBJECTS_H\n+#define PACK_OBJECTS_H\n+\n+struct object_entry {\n+\tstruct pack_idx_entry idx;\n+\tunsigned long size;\t/* uncompressed size */\n+\tstruct packed_git *in_pack;\t/* already in pack */\n+\toff_t in_pack_offset;\n+\tstruct object_entry *delta;\t/* delta base object */\n+\tstruct object_entry *delta_child; /* deltified objects who bases me */\n+\tstruct object_entry *delta_sibling; /* other deltified objects who\n+\t\t\t\t\t     * uses the same base as me\n+\t\t\t\t\t     */\n+\tvoid *delta_data;\t/* cached delta (uncompressed) */\n+\tunsigned long delta_size;\t/* delta data size (uncompressed) */\n+\tunsigned long z_delta_size;\t/* delta data size (compressed) */\n+\tenum object_type type;\n+\tenum object_type in_pack_type;\t/* could be delta */\n+\tuint32_t hash;\t\t\t/* name hint hash */\n+\tunsigned char in_pack_header_size;\n+\tunsigned preferred_base:1; /*\n+\t\t\t\t    * we do not pack this, but is available\n+\t\t\t\t    * to be used as the base object to delta\n+\t\t\t\t    * objects against.\n+\t\t\t\t    */\n+\tunsigned no_try_delta:1;\n+\tunsigned tagged:1; /* near the very tip of refs */\n+\tunsigned filled:1; /* assigned write-order */\n+};\n+\n+struct packing_data {\n+\tstruct object_entry *objects;\n+\tuint32_t nr_objects, nr_alloc;\n+\n+\tint32_t *index;\n+\tuint32_t index_size;\n+};\n+\n+struct object_entry *packlist_alloc(struct packing_data *pdata,\n+\t\t\t\t    const unsigned char *sha1,\n+\t\t\t\t    uint32_t index_pos);\n+\n+struct object_entry *packlist_find(struct packing_data *pdata,\n+\t\t\t\t   const unsigned char *sha1,\n+\t\t\t\t   uint32_t *index_pos);\n+\n+#endif\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230602","messageId":"20131114124306.GD10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 04/21] pack-objects: factor out name_hash","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:43:06Z","receivedAt":"2013-11-14T12:43:06Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nAs the pack-objects system grows beyond the single\npack-objects.c file, more parts (like the soon-to-exist\nbitmap code) will need to compute hashes for matching\ndeltas. Factor out name_hash to make it available to other\nfiles.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n builtin/pack-objects.c | 24 ++----------------------\n pack-objects.h         | 20 ++++++++++++++++++++\n 2 files changed, 22 insertions(+), 22 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex f3f0cf9..faf746b 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -778,26 +778,6 @@ static void write_pack_file(void)\n \t\t\twritten, nr_result);\n }\n \n-static uint32_t name_hash(const char *name)\n-{\n-\tuint32_t c, hash = 0;\n-\n-\tif (!name)\n-\t\treturn 0;\n-\n-\t/*\n-\t * This effectively just creates a sortable number from the\n-\t * last sixteen non-whitespace characters. Last characters\n-\t * count \"most\", so things that end in \".c\" sort together.\n-\t */\n-\twhile ((c = *name++) != 0) {\n-\t\tif (isspace(c))\n-\t\t\tcontinue;\n-\t\thash = (hash >> 2) + (c << 24);\n-\t}\n-\treturn hash;\n-}\n-\n static void setup_delta_attr_check(struct git_attr_check *check)\n {\n \tstatic struct git_attr *attr_delta;\n@@ -826,7 +806,7 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n \tstruct object_entry *entry;\n \tstruct packed_git *p, *found_pack = NULL;\n \toff_t found_offset = 0;\n-\tuint32_t hash = name_hash(name);\n+\tuint32_t hash = pack_name_hash(name);\n \tuint32_t index_pos;\n \n \tentry = packlist_find(&to_pack, sha1, &index_pos);\n@@ -1082,7 +1062,7 @@ static void add_preferred_base_object(const char *name)\n {\n \tstruct pbase_tree *it;\n \tint cmplen;\n-\tunsigned hash = name_hash(name);\n+\tunsigned hash = pack_name_hash(name);\n \n \tif (!num_preferred_base || check_pbase_path(hash))\n \t\treturn;\ndiff --git a/pack-objects.h b/pack-objects.h\nindex f528215..90ad0a8 100644\n--- a/pack-objects.h\n+++ b/pack-objects.h\n@@ -44,4 +44,24 @@ struct object_entry *packlist_find(struct packing_data *pdata,\n \t\t\t\t   const unsigned char *sha1,\n \t\t\t\t   uint32_t *index_pos);\n \n+static inline uint32_t pack_name_hash(const char *name)\n+{\n+\tuint32_t c, hash = 0;\n+\n+\tif (!name)\n+\t\treturn 0;\n+\n+\t/*\n+\t * This effectively just creates a sortable number from the\n+\t * last sixteen non-whitespace characters. Last characters\n+\t * count \"most\", so things that end in \".c\" sort together.\n+\t */\n+\twhile ((c = *name++) != 0) {\n+\t\tif (isspace(c))\n+\t\t\tcontinue;\n+\t\thash = (hash >> 2) + (c << 24);\n+\t}\n+\treturn hash;\n+}\n+\n #endif\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230603","messageId":"20131114124311.GE10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 05/21] revision: allow setting custom limiter function","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:43:11Z","receivedAt":"2013-11-14T12:43:11Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nThis commit enables users of `struct rev_info` to peform custom limiting\nduring a revision walk (i.e. `get_revision`).\n\nIf the field `include_check` has been set to a callback, this callback\nwill be issued once for each commit before it is added to the \"pending\"\nlist of the revwalk. If the include check returns 0, the commit will be\nmarked as added but won't be pushed to the pending list, effectively\nlimiting the walk.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n revision.c | 4 ++++\n revision.h | 2 ++\n 2 files changed, 6 insertions(+)\n\ndiff --git a/revision.c b/revision.c\nindex 956040c..260f4a1 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -779,6 +779,10 @@ static int add_parents_to_list(struct rev_info *revs, struct commit *commit,\n \t\treturn 0;\n \tcommit->object.flags |= ADDED;\n \n+\tif (revs->include_check &&\n+\t    !revs->include_check(commit, revs->include_check_data))\n+\t\treturn 0;\n+\n \t/*\n \t * If the commit is uninteresting, don't try to\n \t * prune parents - we want the maximal uninteresting\ndiff --git a/revision.h b/revision.h\nindex 89132df..5658ddd 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -169,6 +169,8 @@ struct rev_info {\n \tunsigned long min_age;\n \tint min_parents;\n \tint max_parents;\n+\tint (*include_check)(struct commit *, void *);\n+\tvoid *include_check_data;\n \n \t/* diff info for patches and for paths limiting */\n \tstruct diff_options diffopt;\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230604","messageId":"20131114124316.GF10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 06/21] sha1_file: export `git_open_noatime`","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:43:16Z","receivedAt":"2013-11-14T12:43:16Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nThe `git_open_noatime` helper can be of general interest for other\nconsumers of git's different on-disk formats.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n cache.h     | 1 +\n sha1_file.c | 4 +---\n 2 files changed, 2 insertions(+), 3 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex ce377e1..6d4ef65 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -780,6 +780,7 @@ extern int hash_sha1_file(const void *buf, unsigned long len, const char *type,\n extern int write_sha1_file(const void *buf, unsigned long len, const char *type, unsigned char *return_sha1);\n extern int pretend_sha1_file(void *, unsigned long, enum object_type, unsigned char *);\n extern int force_object_loose(const unsigned char *sha1, time_t mtime);\n+extern int git_open_noatime(const char *name);\n extern void *map_sha1_file(const unsigned char *sha1, unsigned long *size);\n extern int unpack_sha1_header(git_zstream *stream, unsigned char *map, unsigned long mapsize, void *buffer, unsigned long bufsiz);\n extern int parse_sha1_header(const char *hdr, unsigned long *sizep);\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 7dadd04..5557bd9 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -239,8 +239,6 @@ char *sha1_pack_index_name(const unsigned char *sha1)\n struct alternate_object_database *alt_odb_list;\n static struct alternate_object_database **alt_odb_tail;\n \n-static int git_open_noatime(const char *name);\n-\n /*\n  * Prepare alternate object database registry.\n  *\n@@ -1357,7 +1355,7 @@ int check_sha1_signature(const unsigned char *sha1, void *map,\n \treturn hashcmp(sha1, real_sha1) ? -1 : 0;\n }\n \n-static int git_open_noatime(const char *name)\n+int git_open_noatime(const char *name)\n {\n \tstatic int sha1_file_open_flag = O_NOATIME;\n \n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230605","messageId":"20131114124336.GG10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 07/21] compat: add endianness helpers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:43:36Z","receivedAt":"2013-11-14T12:43:36Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nThe POSIX standard doesn't currently define a `ntohll`/`htonll`\nfunction pair to perform network-to-host and host-to-network\nswaps of 64-bit data. These 64-bit swaps are necessary for the on-disk\nstorage of EWAH bitmaps if they are not in native byte order.\n\nMany thanks to Ramsay Jones <ramsay@ramsay1.demon.co.uk> and\nTorsten Bögershausen <tboegi@web.de> for cygwin/mingw/msvc\nportability fixes.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n compat/bswap.h | 76 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 75 insertions(+), 1 deletion(-)\n\ndiff --git a/compat/bswap.h b/compat/bswap.h\nindex 5061214..c18a78e 100644\n--- a/compat/bswap.h\n+++ b/compat/bswap.h\n@@ -17,7 +17,20 @@ static inline uint32_t default_swab32(uint32_t val)\n \t\t((val & 0x000000ff) << 24));\n }\n \n+static inline uint64_t default_bswap64(uint64_t val)\n+{\n+\treturn (((val & (uint64_t)0x00000000000000ffULL) << 56) |\n+\t\t((val & (uint64_t)0x000000000000ff00ULL) << 40) |\n+\t\t((val & (uint64_t)0x0000000000ff0000ULL) << 24) |\n+\t\t((val & (uint64_t)0x00000000ff000000ULL) <<  8) |\n+\t\t((val & (uint64_t)0x000000ff00000000ULL) >>  8) |\n+\t\t((val & (uint64_t)0x0000ff0000000000ULL) >> 24) |\n+\t\t((val & (uint64_t)0x00ff000000000000ULL) >> 40) |\n+\t\t((val & (uint64_t)0xff00000000000000ULL) >> 56));\n+}\n+\n #undef bswap32\n+#undef bswap64\n \n #if defined(__GNUC__) && (defined(__i386__) || defined(__x86_64__))\n \n@@ -32,15 +45,42 @@ static inline uint32_t git_bswap32(uint32_t x)\n \treturn result;\n }\n \n+#define bswap64 git_bswap64\n+#if defined(__x86_64__)\n+static inline uint64_t git_bswap64(uint64_t x)\n+{\n+\tuint64_t result;\n+\tif (__builtin_constant_p(x))\n+\t\tresult = default_bswap64(x);\n+\telse\n+\t\t__asm__(\"bswap %q0\" : \"=r\" (result) : \"0\" (x));\n+\treturn result;\n+}\n+#else\n+static inline uint64_t git_bswap64(uint64_t x)\n+{\n+\tunion { uint64_t i64; uint32_t i32[2]; } tmp, result;\n+\tif (__builtin_constant_p(x))\n+\t\tresult.i64 = default_bswap64(x);\n+\telse {\n+\t\ttmp.i64 = x;\n+\t\tresult.i32[0] = git_bswap32(tmp.i32[1]);\n+\t\tresult.i32[1] = git_bswap32(tmp.i32[0]);\n+\t}\n+\treturn result.i64;\n+}\n+#endif\n+\n #elif defined(_MSC_VER) && (defined(_M_IX86) || defined(_M_X64))\n \n #include <stdlib.h>\n \n #define bswap32(x) _byteswap_ulong(x)\n+#define bswap64(x) _byteswap_uint64(x)\n \n #endif\n \n-#ifdef bswap32\n+#if defined(bswap32)\n \n #undef ntohl\n #undef htonl\n@@ -48,3 +88,37 @@ static inline uint32_t git_bswap32(uint32_t x)\n #define htonl(x) bswap32(x)\n \n #endif\n+\n+#if defined(bswap64)\n+\n+#undef ntohll\n+#undef htonll\n+#define ntohll(x) bswap64(x)\n+#define htonll(x) bswap64(x)\n+\n+#else\n+\n+#undef ntohll\n+#undef htonll\n+\n+#if !defined(__BYTE_ORDER)\n+# if defined(BYTE_ORDER) && defined(LITTLE_ENDIAN) && defined(BIG_ENDIAN)\n+#  define __BYTE_ORDER BYTE_ORDER\n+#  define __LITTLE_ENDIAN LITTLE_ENDIAN\n+#  define __BIG_ENDIAN BIG_ENDIAN\n+# endif\n+#endif\n+\n+#if !defined(__BYTE_ORDER)\n+# error \"Cannot determine endianness\"\n+#endif\n+\n+#if __BYTE_ORDER == __BIG_ENDIAN\n+# define ntohll(n) (n)\n+# define htonll(n) (n)\n+#else\n+# define ntohll(n) default_bswap64(n)\n+# define htonll(n) default_bswap64(n)\n+#endif\n+\n+#endif\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230606","messageId":"20131114124351.GH10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 08/21] ewah: compressed bitmap implementation","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:43:51Z","receivedAt":"2013-11-14T12:43:51Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nEWAH is a word-aligned compressed variant of a bitset (i.e. a data\nstructure that acts as a 0-indexed boolean array for many entries).\n\nIt uses a 64-bit run-length encoding (RLE) compression scheme,\ntrading some compression for better processing speed.\n\nThe goal of this word-aligned implementation is not to achieve\nthe best compression, but rather to improve query processing time.\nAs it stands right now, this EWAH implementation will always be more\nefficient storage-wise than its uncompressed alternative.\n\nEWAH arrays will be used as the on-disk format to store reachability\nbitmaps for all objects in a repository while keeping reasonable sizes,\nin the same way that JGit does.\n\nThis EWAH implementation is a mostly straightforward port of the\noriginal `javaewah` library that JGit currently uses. The library is\nself-contained and has been embedded whole (4 files) inside the `ewah`\nfolder to ease redistribution.\n\nThe library is re-licensed under the GPLv2 with the permission of Daniel\nLemire, the original author. The source code for the C version can\nbe found on GitHub:\n\n\thttps://github.com/vmg/libewok\n\nThe original Java implementation can also be found on GitHub:\n\n\thttps://github.com/lemire/javaewah\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Makefile           |  11 +-\n ewah/bitmap.c      | 221 ++++++++++++++++\n ewah/ewah_bitmap.c | 726 +++++++++++++++++++++++++++++++++++++++++++++++++++++\n ewah/ewah_io.c     | 193 ++++++++++++++\n ewah/ewah_rlw.c    | 115 +++++++++\n ewah/ewok.h        | 235 +++++++++++++++++\n ewah/ewok_rlw.h    | 114 +++++++++\n 7 files changed, 1613 insertions(+), 2 deletions(-)\n create mode 100644 ewah/bitmap.c\n create mode 100644 ewah/ewah_bitmap.c\n create mode 100644 ewah/ewah_io.c\n create mode 100644 ewah/ewah_rlw.c\n create mode 100644 ewah/ewok.h\n create mode 100644 ewah/ewok_rlw.h\n\ndiff --git a/Makefile b/Makefile\nindex 48ff0bd..64a1ed7 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -667,6 +667,8 @@ LIB_H += diff.h\n LIB_H += diffcore.h\n LIB_H += dir.h\n LIB_H += exec_cmd.h\n+LIB_H += ewah/ewok.h\n+LIB_H += ewah/ewok_rlw.h\n LIB_H += fetch-pack.h\n LIB_H += fmt-merge-msg.h\n LIB_H += fsck.h\n@@ -800,6 +802,10 @@ LIB_OBJS += dir.o\n LIB_OBJS += editor.o\n LIB_OBJS += entry.o\n LIB_OBJS += environment.o\n+LIB_OBJS += ewah/bitmap.o\n+LIB_OBJS += ewah/ewah_bitmap.o\n+LIB_OBJS += ewah/ewah_io.o\n+LIB_OBJS += ewah/ewah_rlw.o\n LIB_OBJS += exec_cmd.o\n LIB_OBJS += fetch-pack.o\n LIB_OBJS += fsck.o\n@@ -2474,8 +2480,9 @@ profile-clean:\n \t$(RM) $(addsuffix *.gcno,$(addprefix $(PROFILE_DIR)/, $(object_dirs)))\n \n clean: profile-clean coverage-clean\n-\t$(RM) *.o *.res block-sha1/*.o ppc/*.o compat/*.o compat/*/*.o xdiff/*.o vcs-svn/*.o \\\n-\t\tbuiltin/*.o $(LIB_FILE) $(XDIFF_LIB) $(VCSSVN_LIB)\n+\t$(RM) *.o *.res block-sha1/*.o ppc/*.o compat/*.o compat/*/*.o\n+\t$(RM) xdiff/*.o vcs-svn/*.o ewah/*.o builtin/*.o\n+\t$(RM) $(LIB_FILE) $(XDIFF_LIB) $(VCSSVN_LIB)\n \t$(RM) $(ALL_PROGRAMS) $(SCRIPT_LIB) $(BUILT_INS) git$X\n \t$(RM) $(TEST_PROGRAMS) $(NO_INSTALL)\n \t$(RM) -r bin-wrappers $(dep_dirs)\ndiff --git a/ewah/bitmap.c b/ewah/bitmap.c\nnew file mode 100644\nindex 0000000..710e58c\n--- /dev/null\n+++ b/ewah/bitmap.c\n@@ -0,0 +1,221 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#include \"git-compat-util.h\"\n+#include \"ewok.h\"\n+\n+#define MASK(x) ((eword_t)1 << (x % BITS_IN_WORD))\n+#define BLOCK(x) (x / BITS_IN_WORD)\n+\n+struct bitmap *bitmap_new(void)\n+{\n+\tstruct bitmap *bitmap = ewah_malloc(sizeof(struct bitmap));\n+\tbitmap->words = ewah_calloc(32, sizeof(eword_t));\n+\tbitmap->word_alloc = 32;\n+\treturn bitmap;\n+}\n+\n+void bitmap_set(struct bitmap *self, size_t pos)\n+{\n+\tsize_t block = BLOCK(pos);\n+\n+\tif (block >= self->word_alloc) {\n+\t\tsize_t old_size = self->word_alloc;\n+\t\tself->word_alloc = block * 2;\n+\t\tself->words = ewah_realloc(self->words,\n+\t\t\tself->word_alloc * sizeof(eword_t));\n+\n+\t\tmemset(self->words + old_size, 0x0,\n+\t\t\t(self->word_alloc - old_size) * sizeof(eword_t));\n+\t}\n+\n+\tself->words[block] |= MASK(pos);\n+}\n+\n+void bitmap_clear(struct bitmap *self, size_t pos)\n+{\n+\tsize_t block = BLOCK(pos);\n+\n+\tif (block < self->word_alloc)\n+\t\tself->words[block] &= ~MASK(pos);\n+}\n+\n+int bitmap_get(struct bitmap *self, size_t pos)\n+{\n+\tsize_t block = BLOCK(pos);\n+\treturn block < self->word_alloc &&\n+\t\t(self->words[block] & MASK(pos)) != 0;\n+}\n+\n+struct ewah_bitmap *bitmap_to_ewah(struct bitmap *bitmap)\n+{\n+\tstruct ewah_bitmap *ewah = ewah_new();\n+\tsize_t i, running_empty_words = 0;\n+\teword_t last_word = 0;\n+\n+\tfor (i = 0; i < bitmap->word_alloc; ++i) {\n+\t\tif (bitmap->words[i] == 0) {\n+\t\t\trunning_empty_words++;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tif (last_word != 0)\n+\t\t\tewah_add(ewah, last_word);\n+\n+\t\tif (running_empty_words > 0) {\n+\t\t\tewah_add_empty_words(ewah, 0, running_empty_words);\n+\t\t\trunning_empty_words = 0;\n+\t\t}\n+\n+\t\tlast_word = bitmap->words[i];\n+\t}\n+\n+\tewah_add(ewah, last_word);\n+\treturn ewah;\n+}\n+\n+struct bitmap *ewah_to_bitmap(struct ewah_bitmap *ewah)\n+{\n+\tstruct bitmap *bitmap = bitmap_new();\n+\tstruct ewah_iterator it;\n+\teword_t blowup;\n+\tsize_t i = 0;\n+\n+\tewah_iterator_init(&it, ewah);\n+\n+\twhile (ewah_iterator_next(&blowup, &it)) {\n+\t\tif (i >= bitmap->word_alloc) {\n+\t\t\tbitmap->word_alloc *= 1.5;\n+\t\t\tbitmap->words = ewah_realloc(\n+\t\t\t\tbitmap->words, bitmap->word_alloc * sizeof(eword_t));\n+\t\t}\n+\n+\t\tbitmap->words[i++] = blowup;\n+\t}\n+\n+\tbitmap->word_alloc = i;\n+\treturn bitmap;\n+}\n+\n+void bitmap_and_not(struct bitmap *self, struct bitmap *other)\n+{\n+\tconst size_t count = (self->word_alloc < other->word_alloc) ?\n+\t\tself->word_alloc : other->word_alloc;\n+\n+\tsize_t i;\n+\n+\tfor (i = 0; i < count; ++i)\n+\t\tself->words[i] &= ~other->words[i];\n+}\n+\n+void bitmap_or_ewah(struct bitmap *self, struct ewah_bitmap *other)\n+{\n+\tsize_t original_size = self->word_alloc;\n+\tsize_t other_final = (other->bit_size / BITS_IN_WORD) + 1;\n+\tsize_t i = 0;\n+\tstruct ewah_iterator it;\n+\teword_t word;\n+\n+\tif (self->word_alloc < other_final) {\n+\t\tself->word_alloc = other_final;\n+\t\tself->words = ewah_realloc(self->words,\n+\t\t\tself->word_alloc * sizeof(eword_t));\n+\t\tmemset(self->words + original_size, 0x0,\n+\t\t\t(self->word_alloc - original_size) * sizeof(eword_t));\n+\t}\n+\n+\tewah_iterator_init(&it, other);\n+\n+\twhile (ewah_iterator_next(&word, &it))\n+\t\tself->words[i++] |= word;\n+}\n+\n+void bitmap_each_bit(struct bitmap *self, ewah_callback callback, void *data)\n+{\n+\tsize_t pos = 0, i;\n+\n+\tfor (i = 0; i < self->word_alloc; ++i) {\n+\t\teword_t word = self->words[i];\n+\t\tuint32_t offset;\n+\n+\t\tif (word == (eword_t)~0) {\n+\t\t\tfor (offset = 0; offset < BITS_IN_WORD; ++offset)\n+\t\t\t\tcallback(pos++, data);\n+\t\t} else {\n+\t\t\tfor (offset = 0; offset < BITS_IN_WORD; ++offset) {\n+\t\t\t\tif ((word >> offset) == 0)\n+\t\t\t\t\tbreak;\n+\n+\t\t\t\toffset += ewah_bit_ctz64(word >> offset);\n+\t\t\t\tcallback(pos + offset, data);\n+\t\t\t}\n+\t\t\tpos += BITS_IN_WORD;\n+\t\t}\n+\t}\n+}\n+\n+size_t bitmap_popcount(struct bitmap *self)\n+{\n+\tsize_t i, count = 0;\n+\n+\tfor (i = 0; i < self->word_alloc; ++i)\n+\t\tcount += ewah_bit_popcount64(self->words[i]);\n+\n+\treturn count;\n+}\n+\n+int bitmap_equals(struct bitmap *self, struct bitmap *other)\n+{\n+\tstruct bitmap *big, *small;\n+\tsize_t i;\n+\n+\tif (self->word_alloc < other->word_alloc) {\n+\t\tsmall = self;\n+\t\tbig = other;\n+\t} else {\n+\t\tsmall = other;\n+\t\tbig = self;\n+\t}\n+\n+\tfor (i = 0; i < small->word_alloc; ++i) {\n+\t\tif (small->words[i] != big->words[i])\n+\t\t\treturn 0;\n+\t}\n+\n+\tfor (; i < big->word_alloc; ++i) {\n+\t\tif (big->words[i] != 0)\n+\t\t\treturn 0;\n+\t}\n+\n+\treturn 1;\n+}\n+\n+void bitmap_reset(struct bitmap *bitmap)\n+{\n+\tmemset(bitmap->words, 0x0, bitmap->word_alloc * sizeof(eword_t));\n+}\n+\n+void bitmap_free(struct bitmap *bitmap)\n+{\n+\tif (bitmap == NULL)\n+\t\treturn;\n+\n+\tfree(bitmap->words);\n+\tfree(bitmap);\n+}\ndiff --git a/ewah/ewah_bitmap.c b/ewah/ewah_bitmap.c\nnew file mode 100644\nindex 0000000..f104b87\n--- /dev/null\n+++ b/ewah/ewah_bitmap.c\n@@ -0,0 +1,726 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#include \"git-compat-util.h\"\n+#include \"ewok.h\"\n+#include \"ewok_rlw.h\"\n+\n+static inline size_t min_size(size_t a, size_t b)\n+{\n+\treturn a < b ? a : b;\n+}\n+\n+static inline size_t max_size(size_t a, size_t b)\n+{\n+\treturn a > b ? a : b;\n+}\n+\n+static inline void buffer_grow(struct ewah_bitmap *self, size_t new_size)\n+{\n+\tsize_t rlw_offset = (uint8_t *)self->rlw - (uint8_t *)self->buffer;\n+\n+\tif (self->alloc_size >= new_size)\n+\t\treturn;\n+\n+\tself->alloc_size = new_size;\n+\tself->buffer = ewah_realloc(self->buffer,\n+\t\tself->alloc_size * sizeof(eword_t));\n+\tself->rlw = self->buffer + (rlw_offset / sizeof(size_t));\n+}\n+\n+static inline void buffer_push(struct ewah_bitmap *self, eword_t value)\n+{\n+\tif (self->buffer_size + 1 >= self->alloc_size)\n+\t\tbuffer_grow(self, self->buffer_size * 3 / 2);\n+\n+\tself->buffer[self->buffer_size++] = value;\n+}\n+\n+static void buffer_push_rlw(struct ewah_bitmap *self, eword_t value)\n+{\n+\tbuffer_push(self, value);\n+\tself->rlw = self->buffer + self->buffer_size - 1;\n+}\n+\n+static size_t add_empty_words(struct ewah_bitmap *self, int v, size_t number)\n+{\n+\tsize_t added = 0;\n+\teword_t runlen, can_add;\n+\n+\tif (rlw_get_run_bit(self->rlw) != v && rlw_size(self->rlw) == 0) {\n+\t\trlw_set_run_bit(self->rlw, v);\n+\t} else if (rlw_get_literal_words(self->rlw) != 0 ||\n+\t\t\trlw_get_run_bit(self->rlw) != v) {\n+\t\tbuffer_push_rlw(self, 0);\n+\t\tif (v) rlw_set_run_bit(self->rlw, v);\n+\t\tadded++;\n+\t}\n+\n+\trunlen = rlw_get_running_len(self->rlw);\n+\tcan_add = min_size(number, RLW_LARGEST_RUNNING_COUNT - runlen);\n+\n+\trlw_set_running_len(self->rlw, runlen + can_add);\n+\tnumber -= can_add;\n+\n+\twhile (number >= RLW_LARGEST_RUNNING_COUNT) {\n+\t\tbuffer_push_rlw(self, 0);\n+\t\tadded++;\n+\t\tif (v) rlw_set_run_bit(self->rlw, v);\n+\t\trlw_set_running_len(self->rlw, RLW_LARGEST_RUNNING_COUNT);\n+\t\tnumber -= RLW_LARGEST_RUNNING_COUNT;\n+\t}\n+\n+\tif (number > 0) {\n+\t\tbuffer_push_rlw(self, 0);\n+\t\tadded++;\n+\n+\t\tif (v) rlw_set_run_bit(self->rlw, v);\n+\t\trlw_set_running_len(self->rlw, number);\n+\t}\n+\n+\treturn added;\n+}\n+\n+size_t ewah_add_empty_words(struct ewah_bitmap *self, int v, size_t number)\n+{\n+\tif (number == 0)\n+\t\treturn 0;\n+\n+\tself->bit_size += number * BITS_IN_WORD;\n+\treturn add_empty_words(self, v, number);\n+}\n+\n+static size_t add_literal(struct ewah_bitmap *self, eword_t new_data)\n+{\n+\teword_t current_num = rlw_get_literal_words(self->rlw);\n+\n+\tif (current_num >= RLW_LARGEST_LITERAL_COUNT) {\n+\t\tbuffer_push_rlw(self, 0);\n+\n+\t\trlw_set_literal_words(self->rlw, 1);\n+\t\tbuffer_push(self, new_data);\n+\t\treturn 2;\n+\t}\n+\n+\trlw_set_literal_words(self->rlw, current_num + 1);\n+\n+\t/* sanity check */\n+\tassert(rlw_get_literal_words(self->rlw) == current_num + 1);\n+\n+\tbuffer_push(self, new_data);\n+\treturn 1;\n+}\n+\n+void ewah_add_dirty_words(\n+\tstruct ewah_bitmap *self, const eword_t *buffer,\n+\tsize_t number, int negate)\n+{\n+\tsize_t literals, can_add;\n+\n+\twhile (1) {\n+\t\tliterals = rlw_get_literal_words(self->rlw);\n+\t\tcan_add = min_size(number, RLW_LARGEST_LITERAL_COUNT - literals);\n+\n+\t\trlw_set_literal_words(self->rlw, literals + can_add);\n+\n+\t\tif (self->buffer_size + can_add >= self->alloc_size)\n+\t\t\tbuffer_grow(self, (self->buffer_size + can_add) * 3 / 2);\n+\n+\t\tif (negate) {\n+\t\t\tsize_t i;\n+\t\t\tfor (i = 0; i < can_add; ++i)\n+\t\t\t\tself->buffer[self->buffer_size++] = ~buffer[i];\n+\t\t} else {\n+\t\t\tmemcpy(self->buffer + self->buffer_size,\n+\t\t\t\tbuffer, can_add * sizeof(eword_t));\n+\t\t\tself->buffer_size += can_add;\n+\t\t}\n+\n+\t\tself->bit_size += can_add * BITS_IN_WORD;\n+\n+\t\tif (number - can_add == 0)\n+\t\t\tbreak;\n+\n+\t\tbuffer_push_rlw(self, 0);\n+\t\tbuffer += can_add;\n+\t\tnumber -= can_add;\n+\t}\n+}\n+\n+static size_t add_empty_word(struct ewah_bitmap *self, int v)\n+{\n+\tint no_literal = (rlw_get_literal_words(self->rlw) == 0);\n+\teword_t run_len = rlw_get_running_len(self->rlw);\n+\n+\tif (no_literal && run_len == 0) {\n+\t\trlw_set_run_bit(self->rlw, v);\n+\t\tassert(rlw_get_run_bit(self->rlw) == v);\n+\t}\n+\n+\tif (no_literal && rlw_get_run_bit(self->rlw) == v &&\n+\t\trun_len < RLW_LARGEST_RUNNING_COUNT) {\n+\t\trlw_set_running_len(self->rlw, run_len + 1);\n+\t\tassert(rlw_get_running_len(self->rlw) == run_len + 1);\n+\t\treturn 0;\n+\t} else {\n+\t\tbuffer_push_rlw(self, 0);\n+\n+\t\tassert(rlw_get_running_len(self->rlw) == 0);\n+\t\tassert(rlw_get_run_bit(self->rlw) == 0);\n+\t\tassert(rlw_get_literal_words(self->rlw) == 0);\n+\n+\t\trlw_set_run_bit(self->rlw, v);\n+\t\tassert(rlw_get_run_bit(self->rlw) == v);\n+\n+\t\trlw_set_running_len(self->rlw, 1);\n+\t\tassert(rlw_get_running_len(self->rlw) == 1);\n+\t\tassert(rlw_get_literal_words(self->rlw) == 0);\n+\t\treturn 1;\n+\t}\n+}\n+\n+size_t ewah_add(struct ewah_bitmap *self, eword_t word)\n+{\n+\tself->bit_size += BITS_IN_WORD;\n+\n+\tif (word == 0)\n+\t\treturn add_empty_word(self, 0);\n+\n+\tif (word == (eword_t)(~0))\n+\t\treturn add_empty_word(self, 1);\n+\n+\treturn add_literal(self, word);\n+}\n+\n+void ewah_set(struct ewah_bitmap *self, size_t i)\n+{\n+\tconst size_t dist =\n+\t\t(i + BITS_IN_WORD) / BITS_IN_WORD -\n+\t\t(self->bit_size + BITS_IN_WORD - 1) / BITS_IN_WORD;\n+\n+\tassert(i >= self->bit_size);\n+\n+\tself->bit_size = i + 1;\n+\n+\tif (dist > 0) {\n+\t\tif (dist > 1)\n+\t\t\tadd_empty_words(self, 0, dist - 1);\n+\n+\t\tadd_literal(self, (eword_t)1 << (i % BITS_IN_WORD));\n+\t\treturn;\n+\t}\n+\n+\tif (rlw_get_literal_words(self->rlw) == 0) {\n+\t\trlw_set_running_len(self->rlw,\n+\t\t\trlw_get_running_len(self->rlw) - 1);\n+\t\tadd_literal(self, (eword_t)1 << (i % BITS_IN_WORD));\n+\t\treturn;\n+\t}\n+\n+\tself->buffer[self->buffer_size - 1] |=\n+\t\t((eword_t)1 << (i % BITS_IN_WORD));\n+\n+\t/* check if we just completed a stream of 1s */\n+\tif (self->buffer[self->buffer_size - 1] == (eword_t)(~0)) {\n+\t\tself->buffer[--self->buffer_size] = 0;\n+\t\trlw_set_literal_words(self->rlw,\n+\t\t\trlw_get_literal_words(self->rlw) - 1);\n+\t\tadd_empty_word(self, 1);\n+\t}\n+}\n+\n+void ewah_each_bit(struct ewah_bitmap *self, void (*callback)(size_t, void*), void *payload)\n+{\n+\tsize_t pos = 0;\n+\tsize_t pointer = 0;\n+\tsize_t k;\n+\n+\twhile (pointer < self->buffer_size) {\n+\t\teword_t *word = &self->buffer[pointer];\n+\n+\t\tif (rlw_get_run_bit(word)) {\n+\t\t\tsize_t len = rlw_get_running_len(word) * BITS_IN_WORD;\n+\t\t\tfor (k = 0; k < len; ++k, ++pos)\n+\t\t\t\tcallback(pos, payload);\n+\t\t} else {\n+\t\t\tpos += rlw_get_running_len(word) * BITS_IN_WORD;\n+\t\t}\n+\n+\t\t++pointer;\n+\n+\t\tfor (k = 0; k < rlw_get_literal_words(word); ++k) {\n+\t\t\tint c;\n+\n+\t\t\t/* todo: zero count optimization */\n+\t\t\tfor (c = 0; c < BITS_IN_WORD; ++c, ++pos) {\n+\t\t\t\tif ((self->buffer[pointer] & ((eword_t)1 << c)) != 0)\n+\t\t\t\t\tcallback(pos, payload);\n+\t\t\t}\n+\n+\t\t\t++pointer;\n+\t\t}\n+\t}\n+}\n+\n+struct ewah_bitmap *ewah_new(void)\n+{\n+\tstruct ewah_bitmap *self;\n+\n+\tself = ewah_malloc(sizeof(struct ewah_bitmap));\n+\tif (self == NULL)\n+\t\treturn NULL;\n+\n+\tself->buffer = ewah_malloc(32 * sizeof(eword_t));\n+\tself->alloc_size = 32;\n+\n+\tewah_clear(self);\n+\treturn self;\n+}\n+\n+void ewah_clear(struct ewah_bitmap *self)\n+{\n+\tself->buffer_size = 1;\n+\tself->buffer[0] = 0;\n+\tself->bit_size = 0;\n+\tself->rlw = self->buffer;\n+}\n+\n+void ewah_free(struct ewah_bitmap *self)\n+{\n+\tif (!self)\n+\t\treturn;\n+\n+\tif (self->alloc_size)\n+\t\tfree(self->buffer);\n+\n+\tfree(self);\n+}\n+\n+static void read_new_rlw(struct ewah_iterator *it)\n+{\n+\tconst eword_t *word = NULL;\n+\n+\tit->literals = 0;\n+\tit->compressed = 0;\n+\n+\twhile (1) {\n+\t\tword = &it->buffer[it->pointer];\n+\n+\t\tit->rl = rlw_get_running_len(word);\n+\t\tit->lw = rlw_get_literal_words(word);\n+\t\tit->b = rlw_get_run_bit(word);\n+\n+\t\tif (it->rl || it->lw)\n+\t\t\treturn;\n+\n+\t\tif (it->pointer < it->buffer_size - 1) {\n+\t\t\tit->pointer++;\n+\t\t} else {\n+\t\t\tit->pointer = it->buffer_size;\n+\t\t\treturn;\n+\t\t}\n+\t}\n+}\n+\n+int ewah_iterator_next(eword_t *next, struct ewah_iterator *it)\n+{\n+\tif (it->pointer >= it->buffer_size)\n+\t\treturn 0;\n+\n+\tif (it->compressed < it->rl) {\n+\t\tit->compressed++;\n+\t\t*next = it->b ? (eword_t)(~0) : 0;\n+\t} else {\n+\t\tassert(it->literals < it->lw);\n+\n+\t\tit->literals++;\n+\t\tit->pointer++;\n+\n+\t\tassert(it->pointer < it->buffer_size);\n+\n+\t\t*next = it->buffer[it->pointer];\n+\t}\n+\n+\tif (it->compressed == it->rl && it->literals == it->lw) {\n+\t\tif (++it->pointer < it->buffer_size)\n+\t\t\tread_new_rlw(it);\n+\t}\n+\n+\treturn 1;\n+}\n+\n+void ewah_iterator_init(struct ewah_iterator *it, struct ewah_bitmap *parent)\n+{\n+\tit->buffer = parent->buffer;\n+\tit->buffer_size = parent->buffer_size;\n+\tit->pointer = 0;\n+\n+\tit->lw = 0;\n+\tit->rl = 0;\n+\tit->compressed = 0;\n+\tit->literals = 0;\n+\tit->b = 0;\n+\n+\tif (it->pointer < it->buffer_size)\n+\t\tread_new_rlw(it);\n+}\n+\n+void ewah_dump(struct ewah_bitmap *self)\n+{\n+\tsize_t i;\n+\tfprintf(stderr, \"%\"PRIuMAX\" bits | %\"PRIuMAX\" words | \",\n+\t\t(uintmax_t)self->bit_size, (uintmax_t)self->buffer_size);\n+\n+\tfor (i = 0; i < self->buffer_size; ++i)\n+\t\tfprintf(stderr, \"%016\"PRIx64\" \", (uint64_t)self->buffer[i]);\n+\n+\tfprintf(stderr, \"\\n\");\n+}\n+\n+void ewah_not(struct ewah_bitmap *self)\n+{\n+\tsize_t pointer = 0;\n+\n+\twhile (pointer < self->buffer_size) {\n+\t\teword_t *word = &self->buffer[pointer];\n+\t\tsize_t literals, k;\n+\n+\t\trlw_xor_run_bit(word);\n+\t\t++pointer;\n+\n+\t\tliterals = rlw_get_literal_words(word);\n+\t\tfor (k = 0; k < literals; ++k) {\n+\t\t\tself->buffer[pointer] = ~self->buffer[pointer];\n+\t\t\t++pointer;\n+\t\t}\n+\t}\n+}\n+\n+void ewah_xor(\n+\tstruct ewah_bitmap *ewah_i,\n+\tstruct ewah_bitmap *ewah_j,\n+\tstruct ewah_bitmap *out)\n+{\n+\tstruct rlw_iterator rlw_i;\n+\tstruct rlw_iterator rlw_j;\n+\tsize_t literals;\n+\n+\trlwit_init(&rlw_i, ewah_i);\n+\trlwit_init(&rlw_j, ewah_j);\n+\n+\twhile (rlwit_word_size(&rlw_i) > 0 && rlwit_word_size(&rlw_j) > 0) {\n+\t\twhile (rlw_i.rlw.running_len > 0 || rlw_j.rlw.running_len > 0) {\n+\t\t\tstruct rlw_iterator *prey, *predator;\n+\t\t\tsize_t index;\n+\t\t\tint negate_words;\n+\n+\t\t\tif (rlw_i.rlw.running_len < rlw_j.rlw.running_len) {\n+\t\t\t\tprey = &rlw_i;\n+\t\t\t\tpredator = &rlw_j;\n+\t\t\t} else {\n+\t\t\t\tprey = &rlw_j;\n+\t\t\t\tpredator = &rlw_i;\n+\t\t\t}\n+\n+\t\t\tnegate_words = !!predator->rlw.running_bit;\n+\t\t\tindex = rlwit_discharge(prey, out,\n+\t\t\t\tpredator->rlw.running_len, negate_words);\n+\n+\t\t\tewah_add_empty_words(out, negate_words,\n+\t\t\t\tpredator->rlw.running_len - index);\n+\n+\t\t\trlwit_discard_first_words(predator,\n+\t\t\t\tpredator->rlw.running_len);\n+\t\t}\n+\n+\t\tliterals = min_size(\n+\t\t\trlw_i.rlw.literal_words,\n+\t\t\trlw_j.rlw.literal_words);\n+\n+\t\tif (literals) {\n+\t\t\tsize_t k;\n+\n+\t\t\tfor (k = 0; k < literals; ++k) {\n+\t\t\t\tewah_add(out,\n+\t\t\t\t\trlw_i.buffer[rlw_i.literal_word_start + k] ^\n+\t\t\t\t\trlw_j.buffer[rlw_j.literal_word_start + k]\n+\t\t\t\t);\n+\t\t\t}\n+\n+\t\t\trlwit_discard_first_words(&rlw_i, literals);\n+\t\t\trlwit_discard_first_words(&rlw_j, literals);\n+\t\t}\n+\t}\n+\n+\tif (rlwit_word_size(&rlw_i) > 0)\n+\t\trlwit_discharge(&rlw_i, out, ~0, 0);\n+\telse\n+\t\trlwit_discharge(&rlw_j, out, ~0, 0);\n+\n+\tout->bit_size = max_size(ewah_i->bit_size, ewah_j->bit_size);\n+}\n+\n+void ewah_and(\n+\tstruct ewah_bitmap *ewah_i,\n+\tstruct ewah_bitmap *ewah_j,\n+\tstruct ewah_bitmap *out)\n+{\n+\tstruct rlw_iterator rlw_i;\n+\tstruct rlw_iterator rlw_j;\n+\tsize_t literals;\n+\n+\trlwit_init(&rlw_i, ewah_i);\n+\trlwit_init(&rlw_j, ewah_j);\n+\n+\twhile (rlwit_word_size(&rlw_i) > 0 && rlwit_word_size(&rlw_j) > 0) {\n+\t\twhile (rlw_i.rlw.running_len > 0 || rlw_j.rlw.running_len > 0) {\n+\t\t\tstruct rlw_iterator *prey, *predator;\n+\n+\t\t\tif (rlw_i.rlw.running_len < rlw_j.rlw.running_len) {\n+\t\t\t\tprey = &rlw_i;\n+\t\t\t\tpredator = &rlw_j;\n+\t\t\t} else {\n+\t\t\t\tprey = &rlw_j;\n+\t\t\t\tpredator = &rlw_i;\n+\t\t\t}\n+\n+\t\t\tif (predator->rlw.running_bit == 0) {\n+\t\t\t\tewah_add_empty_words(out, 0,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(prey,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(predator,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t} else {\n+\t\t\t\tsize_t index = rlwit_discharge(prey, out,\n+\t\t\t\t\tpredator->rlw.running_len, 0);\n+\t\t\t\tewah_add_empty_words(out, 0,\n+\t\t\t\t\tpredator->rlw.running_len - index);\n+\t\t\t\trlwit_discard_first_words(predator,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t}\n+\t\t}\n+\n+\t\tliterals = min_size(\n+\t\t\trlw_i.rlw.literal_words,\n+\t\t\trlw_j.rlw.literal_words);\n+\n+\t\tif (literals) {\n+\t\t\tsize_t k;\n+\n+\t\t\tfor (k = 0; k < literals; ++k) {\n+\t\t\t\tewah_add(out,\n+\t\t\t\t\trlw_i.buffer[rlw_i.literal_word_start + k] &\n+\t\t\t\t\trlw_j.buffer[rlw_j.literal_word_start + k]\n+\t\t\t\t);\n+\t\t\t}\n+\n+\t\t\trlwit_discard_first_words(&rlw_i, literals);\n+\t\t\trlwit_discard_first_words(&rlw_j, literals);\n+\t\t}\n+\t}\n+\n+\tif (rlwit_word_size(&rlw_i) > 0)\n+\t\trlwit_discharge_empty(&rlw_i, out);\n+\telse\n+\t\trlwit_discharge_empty(&rlw_j, out);\n+\n+\tout->bit_size = max_size(ewah_i->bit_size, ewah_j->bit_size);\n+}\n+\n+void ewah_and_not(\n+\tstruct ewah_bitmap *ewah_i,\n+\tstruct ewah_bitmap *ewah_j,\n+\tstruct ewah_bitmap *out)\n+{\n+\tstruct rlw_iterator rlw_i;\n+\tstruct rlw_iterator rlw_j;\n+\tsize_t literals;\n+\n+\trlwit_init(&rlw_i, ewah_i);\n+\trlwit_init(&rlw_j, ewah_j);\n+\n+\twhile (rlwit_word_size(&rlw_i) > 0 && rlwit_word_size(&rlw_j) > 0) {\n+\t\twhile (rlw_i.rlw.running_len > 0 || rlw_j.rlw.running_len > 0) {\n+\t\t\tstruct rlw_iterator *prey, *predator;\n+\n+\t\t\tif (rlw_i.rlw.running_len < rlw_j.rlw.running_len) {\n+\t\t\t\tprey = &rlw_i;\n+\t\t\t\tpredator = &rlw_j;\n+\t\t\t} else {\n+\t\t\t\tprey = &rlw_j;\n+\t\t\t\tpredator = &rlw_i;\n+\t\t\t}\n+\n+\t\t\tif ((predator->rlw.running_bit && prey == &rlw_i) ||\n+\t\t\t\t(!predator->rlw.running_bit && prey != &rlw_i)) {\n+\t\t\t\tewah_add_empty_words(out, 0,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(prey,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(predator,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t} else {\n+\t\t\t\tsize_t index;\n+\t\t\t\tint negate_words;\n+\n+\t\t\t\tnegate_words = (&rlw_i != prey);\n+\t\t\t\tindex = rlwit_discharge(prey, out,\n+\t\t\t\t\tpredator->rlw.running_len, negate_words);\n+\t\t\t\tewah_add_empty_words(out, negate_words,\n+\t\t\t\t\tpredator->rlw.running_len - index);\n+\t\t\t\trlwit_discard_first_words(predator,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t}\n+\t\t}\n+\n+\t\tliterals = min_size(\n+\t\t\trlw_i.rlw.literal_words,\n+\t\t\trlw_j.rlw.literal_words);\n+\n+\t\tif (literals) {\n+\t\t\tsize_t k;\n+\n+\t\t\tfor (k = 0; k < literals; ++k) {\n+\t\t\t\tewah_add(out,\n+\t\t\t\t\trlw_i.buffer[rlw_i.literal_word_start + k] &\n+\t\t\t\t\t~(rlw_j.buffer[rlw_j.literal_word_start + k])\n+\t\t\t\t);\n+\t\t\t}\n+\n+\t\t\trlwit_discard_first_words(&rlw_i, literals);\n+\t\t\trlwit_discard_first_words(&rlw_j, literals);\n+\t\t}\n+\t}\n+\n+\tif (rlwit_word_size(&rlw_i) > 0)\n+\t\trlwit_discharge(&rlw_i, out, ~0, 0);\n+\telse\n+\t\trlwit_discharge_empty(&rlw_j, out);\n+\n+\tout->bit_size = max_size(ewah_i->bit_size, ewah_j->bit_size);\n+}\n+\n+void ewah_or(\n+\tstruct ewah_bitmap *ewah_i,\n+\tstruct ewah_bitmap *ewah_j,\n+\tstruct ewah_bitmap *out)\n+{\n+\tstruct rlw_iterator rlw_i;\n+\tstruct rlw_iterator rlw_j;\n+\tsize_t literals;\n+\n+\trlwit_init(&rlw_i, ewah_i);\n+\trlwit_init(&rlw_j, ewah_j);\n+\n+\twhile (rlwit_word_size(&rlw_i) > 0 && rlwit_word_size(&rlw_j) > 0) {\n+\t\twhile (rlw_i.rlw.running_len > 0 || rlw_j.rlw.running_len > 0) {\n+\t\t\tstruct rlw_iterator *prey, *predator;\n+\n+\t\t\tif (rlw_i.rlw.running_len < rlw_j.rlw.running_len) {\n+\t\t\t\tprey = &rlw_i;\n+\t\t\t\tpredator = &rlw_j;\n+\t\t\t} else {\n+\t\t\t\tprey = &rlw_j;\n+\t\t\t\tpredator = &rlw_i;\n+\t\t\t}\n+\n+\t\t\tif (predator->rlw.running_bit) {\n+\t\t\t\tewah_add_empty_words(out, 0,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(prey,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(predator,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t} else {\n+\t\t\t\tsize_t index = rlwit_discharge(prey, out,\n+\t\t\t\t\tpredator->rlw.running_len, 0);\n+\t\t\t\tewah_add_empty_words(out, 0,\n+\t\t\t\t\tpredator->rlw.running_len - index);\n+\t\t\t\trlwit_discard_first_words(predator,\n+\t\t\t\t\tpredator->rlw.running_len);\n+\t\t\t}\n+\t\t}\n+\n+\t\tliterals = min_size(\n+\t\t\trlw_i.rlw.literal_words,\n+\t\t\trlw_j.rlw.literal_words);\n+\n+\t\tif (literals) {\n+\t\t\tsize_t k;\n+\n+\t\t\tfor (k = 0; k < literals; ++k) {\n+\t\t\t\tewah_add(out,\n+\t\t\t\t\trlw_i.buffer[rlw_i.literal_word_start + k] |\n+\t\t\t\t\trlw_j.buffer[rlw_j.literal_word_start + k]\n+\t\t\t\t);\n+\t\t\t}\n+\n+\t\t\trlwit_discard_first_words(&rlw_i, literals);\n+\t\t\trlwit_discard_first_words(&rlw_j, literals);\n+\t\t}\n+\t}\n+\n+\tif (rlwit_word_size(&rlw_i) > 0)\n+\t\trlwit_discharge(&rlw_i, out, ~0, 0);\n+\telse\n+\t\trlwit_discharge(&rlw_j, out, ~0, 0);\n+\n+\tout->bit_size = max_size(ewah_i->bit_size, ewah_j->bit_size);\n+}\n+\n+\n+#define BITMAP_POOL_MAX 16\n+static struct ewah_bitmap *bitmap_pool[BITMAP_POOL_MAX];\n+static size_t bitmap_pool_size;\n+\n+struct ewah_bitmap *ewah_pool_new(void)\n+{\n+\tif (bitmap_pool_size)\n+\t\treturn bitmap_pool[--bitmap_pool_size];\n+\n+\treturn ewah_new();\n+}\n+\n+void ewah_pool_free(struct ewah_bitmap *self)\n+{\n+\tif (self == NULL)\n+\t\treturn;\n+\n+\tif (bitmap_pool_size == BITMAP_POOL_MAX ||\n+\t\tself->alloc_size == 0) {\n+\t\tewah_free(self);\n+\t\treturn;\n+\t}\n+\n+\tewah_clear(self);\n+\tbitmap_pool[bitmap_pool_size++] = self;\n+}\n+\n+uint32_t ewah_checksum(struct ewah_bitmap *self)\n+{\n+\tconst uint8_t *p = (uint8_t *)self->buffer;\n+\tuint32_t crc = (uint32_t)self->bit_size;\n+\tsize_t size = self->buffer_size * sizeof(eword_t);\n+\n+\twhile (size--)\n+\t\tcrc = (crc << 5) - crc + (uint32_t)*p++;\n+\n+\treturn crc;\n+}\ndiff --git a/ewah/ewah_io.c b/ewah/ewah_io.c\nnew file mode 100644\nindex 0000000..aed0da6\n--- /dev/null\n+++ b/ewah/ewah_io.c\n@@ -0,0 +1,193 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#include \"git-compat-util.h\"\n+#include \"ewok.h\"\n+\n+int ewah_serialize_native(struct ewah_bitmap *self, int fd)\n+{\n+\tuint32_t write32;\n+\tsize_t to_write = self->buffer_size * 8;\n+\n+\t/* 32 bit -- bit size for the map */\n+\twrite32 = (uint32_t)self->bit_size;\n+\tif (write(fd, &write32, 4) != 4)\n+\t\treturn -1;\n+\n+\t/** 32 bit -- number of compressed 64-bit words */\n+\twrite32 = (uint32_t)self->buffer_size;\n+\tif (write(fd, &write32, 4) != 4)\n+\t\treturn -1;\n+\n+\tif (write(fd, self->buffer, to_write) != to_write)\n+\t\treturn -1;\n+\n+\t/** 32 bit -- position for the RLW */\n+\twrite32 = self->rlw - self->buffer;\n+\tif (write(fd, &write32, 4) != 4)\n+\t\treturn -1;\n+\n+\treturn (3 * 4) + to_write;\n+}\n+\n+int ewah_serialize_to(struct ewah_bitmap *self,\n+\t\t      int (*write_fun)(void *, const void *, size_t),\n+\t\t      void *data)\n+{\n+\tsize_t i;\n+\teword_t dump[2048];\n+\tconst size_t words_per_dump = sizeof(dump) / sizeof(eword_t);\n+\tuint32_t bitsize, word_count, rlw_pos;\n+\n+\tconst eword_t *buffer;\n+\tsize_t words_left;\n+\n+\t/* 32 bit -- bit size for the map */\n+\tbitsize =  htonl((uint32_t)self->bit_size);\n+\tif (write_fun(data, &bitsize, 4) != 4)\n+\t\treturn -1;\n+\n+\t/** 32 bit -- number of compressed 64-bit words */\n+\tword_count =  htonl((uint32_t)self->buffer_size);\n+\tif (write_fun(data, &word_count, 4) != 4)\n+\t\treturn -1;\n+\n+\t/** 64 bit x N -- compressed words */\n+\tbuffer = self->buffer;\n+\twords_left = self->buffer_size;\n+\n+\twhile (words_left >= words_per_dump) {\n+\t\tfor (i = 0; i < words_per_dump; ++i, ++buffer)\n+\t\t\tdump[i] = htonll(*buffer);\n+\n+\t\tif (write_fun(data, dump, sizeof(dump)) != sizeof(dump))\n+\t\t\treturn -1;\n+\n+\t\twords_left -= words_per_dump;\n+\t}\n+\n+\tif (words_left) {\n+\t\tfor (i = 0; i < words_left; ++i, ++buffer)\n+\t\t\tdump[i] = htonll(*buffer);\n+\n+\t\tif (write_fun(data, dump, words_left * 8) != words_left * 8)\n+\t\t\treturn -1;\n+\t}\n+\n+\t/** 32 bit -- position for the RLW */\n+\trlw_pos = (uint8_t*)self->rlw - (uint8_t *)self->buffer;\n+\trlw_pos = htonl(rlw_pos / sizeof(eword_t));\n+\n+\tif (write_fun(data, &rlw_pos, 4) != 4)\n+\t\treturn -1;\n+\n+\treturn (3 * 4) + (self->buffer_size * 8);\n+}\n+\n+static int write_helper(void *fd, const void *buf, size_t len)\n+{\n+\treturn write((intptr_t)fd, buf, len);\n+}\n+\n+int ewah_serialize(struct ewah_bitmap *self, int fd)\n+{\n+\treturn ewah_serialize_to(self, write_helper, (void *)(intptr_t)fd);\n+}\n+\n+int ewah_read_mmap(struct ewah_bitmap *self, void *map, size_t len)\n+{\n+\tuint32_t *read32 = map;\n+\teword_t *read64;\n+\tsize_t i;\n+\n+\tself->bit_size = ntohl(*read32++);\n+\tself->buffer_size = self->alloc_size = ntohl(*read32++);\n+\tself->buffer = ewah_realloc(self->buffer,\n+\t\tself->alloc_size * sizeof(eword_t));\n+\n+\tif (!self->buffer)\n+\t\treturn -1;\n+\n+\tfor (i = 0, read64 = (void *)read32; i < self->buffer_size; ++i)\n+\t\tself->buffer[i] = ntohll(*read64++);\n+\n+\tread32 = (void *)read64;\n+\tself->rlw = self->buffer + ntohl(*read32++);\n+\n+\treturn (3 * 4) + (self->buffer_size * 8);\n+}\n+\n+int ewah_deserialize(struct ewah_bitmap *self, int fd)\n+{\n+\tsize_t i;\n+\teword_t dump[2048];\n+\tconst size_t words_per_dump = sizeof(dump) / sizeof(eword_t);\n+\tuint32_t bitsize, word_count, rlw_pos;\n+\n+\teword_t *buffer = NULL;\n+\tsize_t words_left;\n+\n+\tewah_clear(self);\n+\n+\t/* 32 bit -- bit size for the map */\n+\tif (read(fd, &bitsize, 4) != 4)\n+\t\treturn -1;\n+\n+\tself->bit_size = (size_t)ntohl(bitsize);\n+\n+\t/** 32 bit -- number of compressed 64-bit words */\n+\tif (read(fd, &word_count, 4) != 4)\n+\t\treturn -1;\n+\n+\tself->buffer_size = self->alloc_size = (size_t)ntohl(word_count);\n+\tself->buffer = ewah_realloc(self->buffer,\n+\t\tself->alloc_size * sizeof(eword_t));\n+\n+\tif (!self->buffer)\n+\t\treturn -1;\n+\n+\t/** 64 bit x N -- compressed words */\n+\tbuffer = self->buffer;\n+\twords_left = self->buffer_size;\n+\n+\twhile (words_left >= words_per_dump) {\n+\t\tif (read(fd, dump, sizeof(dump)) != sizeof(dump))\n+\t\t\treturn -1;\n+\n+\t\tfor (i = 0; i < words_per_dump; ++i, ++buffer)\n+\t\t\t*buffer = ntohll(dump[i]);\n+\n+\t\twords_left -= words_per_dump;\n+\t}\n+\n+\tif (words_left) {\n+\t\tif (read(fd, dump, words_left * 8) != words_left * 8)\n+\t\t\treturn -1;\n+\n+\t\tfor (i = 0; i < words_left; ++i, ++buffer)\n+\t\t\t*buffer = ntohll(dump[i]);\n+\t}\n+\n+\t/** 32 bit -- position for the RLW */\n+\tif (read(fd, &rlw_pos, 4) != 4)\n+\t\treturn -1;\n+\n+\tself->rlw = self->buffer + ntohl(rlw_pos);\n+\treturn 0;\n+}\ndiff --git a/ewah/ewah_rlw.c b/ewah/ewah_rlw.c\nnew file mode 100644\nindex 0000000..c723f1a\n--- /dev/null\n+++ b/ewah/ewah_rlw.c\n@@ -0,0 +1,115 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#include \"git-compat-util.h\"\n+#include \"ewok.h\"\n+#include \"ewok_rlw.h\"\n+\n+static inline int next_word(struct rlw_iterator *it)\n+{\n+\tif (it->pointer >= it->size)\n+\t\treturn 0;\n+\n+\tit->rlw.word = &it->buffer[it->pointer];\n+\tit->pointer += rlw_get_literal_words(it->rlw.word) + 1;\n+\n+\tit->rlw.literal_words = rlw_get_literal_words(it->rlw.word);\n+\tit->rlw.running_len = rlw_get_running_len(it->rlw.word);\n+\tit->rlw.running_bit = rlw_get_run_bit(it->rlw.word);\n+\tit->rlw.literal_word_offset = 0;\n+\n+\treturn 1;\n+}\n+\n+void rlwit_init(struct rlw_iterator *it, struct ewah_bitmap *from_ewah)\n+{\n+\tit->buffer = from_ewah->buffer;\n+\tit->size = from_ewah->buffer_size;\n+\tit->pointer = 0;\n+\n+\tnext_word(it);\n+\n+\tit->literal_word_start = rlwit_literal_words(it) +\n+\t\tit->rlw.literal_word_offset;\n+}\n+\n+void rlwit_discard_first_words(struct rlw_iterator *it, size_t x)\n+{\n+\twhile (x > 0) {\n+\t\tsize_t discard;\n+\n+\t\tif (it->rlw.running_len > x) {\n+\t\t\tit->rlw.running_len -= x;\n+\t\t\treturn;\n+\t\t}\n+\n+\t\tx -= it->rlw.running_len;\n+\t\tit->rlw.running_len = 0;\n+\n+\t\tdiscard = (x > it->rlw.literal_words) ? it->rlw.literal_words : x;\n+\n+\t\tit->literal_word_start += discard;\n+\t\tit->rlw.literal_words -= discard;\n+\t\tx -= discard;\n+\n+\t\tif (x > 0 || rlwit_word_size(it) == 0) {\n+\t\t\tif (!next_word(it))\n+\t\t\t\tbreak;\n+\n+\t\t\tit->literal_word_start =\n+\t\t\t\trlwit_literal_words(it) + it->rlw.literal_word_offset;\n+\t\t}\n+\t}\n+}\n+\n+size_t rlwit_discharge(\n+\tstruct rlw_iterator *it, struct ewah_bitmap *out, size_t max, int negate)\n+{\n+\tsize_t index = 0;\n+\n+\twhile (index < max && rlwit_word_size(it) > 0) {\n+\t\tsize_t pd, pl = it->rlw.running_len;\n+\n+\t\tif (index + pl > max)\n+\t\t\tpl = max - index;\n+\n+\t\tewah_add_empty_words(out, it->rlw.running_bit ^ negate, pl);\n+\t\tindex += pl;\n+\n+\t\tpd = it->rlw.literal_words;\n+\t\tif (pd + index > max)\n+\t\t\tpd = max - index;\n+\n+\t\tewah_add_dirty_words(out,\n+\t\t\tit->buffer + it->literal_word_start, pd, negate);\n+\n+\t\trlwit_discard_first_words(it, pd + pl);\n+\t\tindex += pd;\n+\t}\n+\n+\treturn index;\n+}\n+\n+void rlwit_discharge_empty(struct rlw_iterator *it, struct ewah_bitmap *out)\n+{\n+\twhile (rlwit_word_size(it) > 0) {\n+\t\tewah_add_empty_words(out, 0, rlwit_word_size(it));\n+\t\trlwit_discard_first_words(it, rlwit_word_size(it));\n+\t}\n+}\ndiff --git a/ewah/ewok.h b/ewah/ewok.h\nnew file mode 100644\nindex 0000000..619afaa\n--- /dev/null\n+++ b/ewah/ewok.h\n@@ -0,0 +1,235 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#ifndef __EWOK_BITMAP_H__\n+#define __EWOK_BITMAP_H__\n+\n+#ifndef ewah_malloc\n+#\tdefine ewah_malloc xmalloc\n+#endif\n+#ifndef ewah_realloc\n+#\tdefine ewah_realloc xrealloc\n+#endif\n+#ifndef ewah_calloc\n+#\tdefine ewah_calloc xcalloc\n+#endif\n+\n+typedef uint64_t eword_t;\n+#define BITS_IN_WORD (sizeof(eword_t) * 8)\n+\n+/**\n+ * Do not use __builtin_popcountll. The GCC implementation\n+ * is notoriously slow on all platforms.\n+ *\n+ * See: http://gcc.gnu.org/bugzilla/show_bug.cgi?id=36041\n+ */\n+static inline uint32_t ewah_bit_popcount64(uint64_t x)\n+{\n+\tx = (x & 0x5555555555555555ULL) + ((x >>  1) & 0x5555555555555555ULL);\n+\tx = (x & 0x3333333333333333ULL) + ((x >>  2) & 0x3333333333333333ULL);\n+\tx = (x & 0x0F0F0F0F0F0F0F0FULL) + ((x >>  4) & 0x0F0F0F0F0F0F0F0FULL);\n+\treturn (x * 0x0101010101010101ULL) >> 56;\n+}\n+\n+#ifdef __GNUC__\n+#define ewah_bit_ctz64(x) __builtin_ctzll(x)\n+#else\n+static inline int ewah_bit_ctz64(uint64_t x)\n+{\n+\tint n = 0;\n+\tif ((x & 0xffffffff) == 0) { x >>= 32; n += 32; }\n+\tif ((x &     0xffff) == 0) { x >>= 16; n += 16; }\n+\tif ((x &       0xff) == 0) { x >>=  8; n +=  8; }\n+\tif ((x &        0xf) == 0) { x >>=  4; n +=  4; }\n+\tif ((x &        0x3) == 0) { x >>=  2; n +=  2; }\n+\tif ((x &        0x1) == 0) { x >>=  1; n +=  1; }\n+\treturn n + !x;\n+}\n+#endif\n+\n+struct ewah_bitmap {\n+\teword_t *buffer;\n+\tsize_t buffer_size;\n+\tsize_t alloc_size;\n+\tsize_t bit_size;\n+\teword_t *rlw;\n+};\n+\n+typedef void (*ewah_callback)(size_t pos, void *);\n+\n+struct ewah_bitmap *ewah_pool_new(void);\n+void ewah_pool_free(struct ewah_bitmap *self);\n+\n+/**\n+ * Allocate a new EWAH Compressed bitmap\n+ */\n+struct ewah_bitmap *ewah_new(void);\n+\n+/**\n+ * Clear all the bits in the bitmap. Does not free or resize\n+ * memory.\n+ */\n+void ewah_clear(struct ewah_bitmap *self);\n+\n+/**\n+ * Free all the memory of the bitmap\n+ */\n+void ewah_free(struct ewah_bitmap *self);\n+\n+int ewah_serialize_to(struct ewah_bitmap *self,\n+\t\t      int (*write_fun)(void *out, const void *buf, size_t len),\n+\t\t      void *out);\n+int ewah_serialize(struct ewah_bitmap *self, int fd);\n+int ewah_serialize_native(struct ewah_bitmap *self, int fd);\n+\n+int ewah_deserialize(struct ewah_bitmap *self, int fd);\n+int ewah_read_mmap(struct ewah_bitmap *self, void *map, size_t len);\n+int ewah_read_mmap_native(struct ewah_bitmap *self, void *map, size_t len);\n+\n+uint32_t ewah_checksum(struct ewah_bitmap *self);\n+\n+/**\n+ * Logical not (bitwise negation) in-place on the bitmap\n+ *\n+ * This operation is linear time based on the size of the bitmap.\n+ */\n+void ewah_not(struct ewah_bitmap *self);\n+\n+/**\n+ * Call the given callback with the position of every single bit\n+ * that has been set on the bitmap.\n+ *\n+ * This is an efficient operation that does not fully decompress\n+ * the bitmap.\n+ */\n+void ewah_each_bit(struct ewah_bitmap *self, ewah_callback callback, void *payload);\n+\n+/**\n+ * Set a given bit on the bitmap.\n+ *\n+ * The bit at position `pos` will be set to true. Because of the\n+ * way that the bitmap is compressed, a set bit cannot be unset\n+ * later on.\n+ *\n+ * Furthermore, since the bitmap uses streaming compression, bits\n+ * can only set incrementally.\n+ *\n+ * E.g.\n+ *\t\tewah_set(bitmap, 1); // ok\n+ *\t\tewah_set(bitmap, 76); // ok\n+ *\t\tewah_set(bitmap, 77); // ok\n+ *\t\tewah_set(bitmap, 8712800127); // ok\n+ *\t\tewah_set(bitmap, 25); // failed, assert raised\n+ */\n+void ewah_set(struct ewah_bitmap *self, size_t i);\n+\n+struct ewah_iterator {\n+\tconst eword_t *buffer;\n+\tsize_t buffer_size;\n+\n+\tsize_t pointer;\n+\teword_t compressed, literals;\n+\teword_t rl, lw;\n+\tint b;\n+};\n+\n+/**\n+ * Initialize a new iterator to run through the bitmap in uncompressed form.\n+ *\n+ * The iterator can be stack allocated. The underlying bitmap must not be freed\n+ * before the iteration is over.\n+ *\n+ * E.g.\n+ *\n+ *\t\tstruct ewah_bitmap *bitmap = ewah_new();\n+ *\t\tstruct ewah_iterator it;\n+ *\n+ *\t\tewah_iterator_init(&it, bitmap);\n+ */\n+void ewah_iterator_init(struct ewah_iterator *it, struct ewah_bitmap *parent);\n+\n+/**\n+ * Yield every single word in the bitmap in uncompressed form. This is:\n+ * yield single words (32-64 bits) where each bit represents an actual\n+ * bit from the bitmap.\n+ *\n+ * Return: true if a word was yield, false if there are no words left\n+ */\n+int ewah_iterator_next(eword_t *next, struct ewah_iterator *it);\n+\n+void ewah_or(\n+\tstruct ewah_bitmap *ewah_i,\n+\tstruct ewah_bitmap *ewah_j,\n+\tstruct ewah_bitmap *out);\n+\n+void ewah_and_not(\n+\tstruct ewah_bitmap *ewah_i,\n+\tstruct ewah_bitmap *ewah_j,\n+\tstruct ewah_bitmap *out);\n+\n+void ewah_xor(\n+\tstruct ewah_bitmap *ewah_i,\n+\tstruct ewah_bitmap *ewah_j,\n+\tstruct ewah_bitmap *out);\n+\n+void ewah_and(\n+\tstruct ewah_bitmap *ewah_i,\n+\tstruct ewah_bitmap *ewah_j,\n+\tstruct ewah_bitmap *out);\n+\n+void ewah_dump(struct ewah_bitmap *self);\n+\n+/**\n+ * Direct word access\n+ */\n+size_t ewah_add_empty_words(struct ewah_bitmap *self, int v, size_t number);\n+void ewah_add_dirty_words(\n+\tstruct ewah_bitmap *self, const eword_t *buffer, size_t number, int negate);\n+size_t ewah_add(struct ewah_bitmap *self, eword_t word);\n+\n+\n+/**\n+ * Uncompressed, old-school bitmap that can be efficiently compressed\n+ * into an `ewah_bitmap`.\n+ */\n+struct bitmap {\n+\teword_t *words;\n+\tsize_t word_alloc;\n+};\n+\n+struct bitmap *bitmap_new(void);\n+void bitmap_set(struct bitmap *self, size_t pos);\n+void bitmap_clear(struct bitmap *self, size_t pos);\n+int bitmap_get(struct bitmap *self, size_t pos);\n+void bitmap_reset(struct bitmap *self);\n+void bitmap_free(struct bitmap *self);\n+int bitmap_equals(struct bitmap *self, struct bitmap *other);\n+int bitmap_is_subset(struct bitmap *self, struct bitmap *super);\n+\n+struct ewah_bitmap * bitmap_to_ewah(struct bitmap *bitmap);\n+struct bitmap *ewah_to_bitmap(struct ewah_bitmap *ewah);\n+\n+void bitmap_and_not(struct bitmap *self, struct bitmap *other);\n+void bitmap_or_ewah(struct bitmap *self, struct ewah_bitmap *other);\n+void bitmap_or(struct bitmap *self, const struct bitmap *other);\n+\n+void bitmap_each_bit(struct bitmap *self, ewah_callback callback, void *data);\n+size_t bitmap_popcount(struct bitmap *self);\n+\n+#endif\ndiff --git a/ewah/ewok_rlw.h b/ewah/ewok_rlw.h\nnew file mode 100644\nindex 0000000..63efdf9\n--- /dev/null\n+++ b/ewah/ewok_rlw.h\n@@ -0,0 +1,114 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#ifndef __EWOK_RLW_H__\n+#define __EWOK_RLW_H__\n+\n+#define RLW_RUNNING_BITS (sizeof(eword_t) * 4)\n+#define RLW_LITERAL_BITS (sizeof(eword_t) * 8 - 1 - RLW_RUNNING_BITS)\n+\n+#define RLW_LARGEST_RUNNING_COUNT (((eword_t)1 << RLW_RUNNING_BITS) - 1)\n+#define RLW_LARGEST_LITERAL_COUNT (((eword_t)1 << RLW_LITERAL_BITS) - 1)\n+\n+#define RLW_LARGEST_RUNNING_COUNT_SHIFT (RLW_LARGEST_RUNNING_COUNT << 1)\n+\n+#define RLW_RUNNING_LEN_PLUS_BIT (((eword_t)1 << (RLW_RUNNING_BITS + 1)) - 1)\n+\n+static int rlw_get_run_bit(const eword_t *word)\n+{\n+\treturn *word & (eword_t)1;\n+}\n+\n+static inline void rlw_set_run_bit(eword_t *word, int b)\n+{\n+\tif (b) {\n+\t\t*word |= (eword_t)1;\n+\t} else {\n+\t\t*word &= (eword_t)(~1);\n+\t}\n+}\n+\n+static inline void rlw_xor_run_bit(eword_t *word)\n+{\n+\tif (*word & 1) {\n+\t\t*word &= (eword_t)(~1);\n+\t} else {\n+\t\t*word |= (eword_t)1;\n+\t}\n+}\n+\n+static inline void rlw_set_running_len(eword_t *word, eword_t l)\n+{\n+\t*word |= RLW_LARGEST_RUNNING_COUNT_SHIFT;\n+\t*word &= (l << 1) | (~RLW_LARGEST_RUNNING_COUNT_SHIFT);\n+}\n+\n+static inline eword_t rlw_get_running_len(const eword_t *word)\n+{\n+\treturn (*word >> 1) & RLW_LARGEST_RUNNING_COUNT;\n+}\n+\n+static inline eword_t rlw_get_literal_words(const eword_t *word)\n+{\n+\treturn *word >> (1 + RLW_RUNNING_BITS);\n+}\n+\n+static inline void rlw_set_literal_words(eword_t *word, eword_t l)\n+{\n+\t*word |= ~RLW_RUNNING_LEN_PLUS_BIT;\n+\t*word &= (l << (RLW_RUNNING_BITS + 1)) | RLW_RUNNING_LEN_PLUS_BIT;\n+}\n+\n+static inline eword_t rlw_size(const eword_t *self)\n+{\n+\treturn rlw_get_running_len(self) + rlw_get_literal_words(self);\n+}\n+\n+struct rlw_iterator {\n+\tconst eword_t *buffer;\n+\tsize_t size;\n+\tsize_t pointer;\n+\tsize_t literal_word_start;\n+\n+\tstruct {\n+\t\tconst eword_t *word;\n+\t\tint literal_words;\n+\t\tint running_len;\n+\t\tint literal_word_offset;\n+\t\tint running_bit;\n+\t} rlw;\n+};\n+\n+void rlwit_init(struct rlw_iterator *it, struct ewah_bitmap *bitmap);\n+void rlwit_discard_first_words(struct rlw_iterator *it, size_t x);\n+size_t rlwit_discharge(\n+\tstruct rlw_iterator *it, struct ewah_bitmap *out, size_t max, int negate);\n+void rlwit_discharge_empty(struct rlw_iterator *it, struct ewah_bitmap *out);\n+\n+static inline size_t rlwit_word_size(struct rlw_iterator *it)\n+{\n+\treturn it->rlw.running_len + it->rlw.literal_words;\n+}\n+\n+static inline size_t rlwit_literal_words(struct rlw_iterator *it)\n+{\n+\treturn it->pointer - it->rlw.literal_words;\n+}\n+\n+#endif\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230607","messageId":"20131114124402.GI10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 09/21] documentation: add documentation for the bitmap format","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:44:02Z","receivedAt":"2013-11-14T12:44:02Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nThis is the technical documentation for the JGit-compatible Bitmap v1\non-disk format.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Documentation/technical/bitmap-format.txt | 131 ++++++++++++++++++++++++++++++\n 1 file changed, 131 insertions(+)\n create mode 100644 Documentation/technical/bitmap-format.txt\n\ndiff --git a/Documentation/technical/bitmap-format.txt b/Documentation/technical/bitmap-format.txt\nnew file mode 100644\nindex 0000000..7a86bd7\n--- /dev/null\n+++ b/Documentation/technical/bitmap-format.txt\n@@ -0,0 +1,131 @@\n+GIT bitmap v1 format\n+====================\n+\n+\t- A header appears at the beginning:\n+\n+\t\t4-byte signature: {'B', 'I', 'T', 'M'}\n+\n+\t\t2-byte version number (network byte order)\n+\t\t\tThe current implementation only supports version 1\n+\t\t\tof the bitmap index (the same one as JGit).\n+\n+\t\t2-byte flags (network byte order)\n+\n+\t\t\tThe following flags are supported:\n+\n+\t\t\t- BITMAP_OPT_FULL_DAG (0x1) REQUIRED\n+\t\t\tThis flag must always be present. It implies that the bitmap\n+\t\t\tindex has been generated for a packfile with full closure\n+\t\t\t(i.e. where every single object in the packfile can find\n+\t\t\t its parent links inside the same packfile). This is a\n+\t\t\trequirement for the bitmap index format, also present in JGit,\n+\t\t\tthat greatly reduces the complexity of the implementation.\n+\n+\t\t4-byte entry count (network byte order)\n+\n+\t\t\tThe total count of entries (bitmapped commits) in this bitmap index.\n+\n+\t\t20-byte checksum\n+\n+\t\t\tThe SHA1 checksum of the pack this bitmap index belongs to.\n+\n+\t- 4 EWAH bitmaps that act as type indexes\n+\n+\t\tType indexes are serialized after the hash cache in the shape\n+\t\tof four EWAH bitmaps stored consecutively (see Appendix A for\n+\t\tthe serialization format of an EWAH bitmap).\n+\n+\t\tThere is a bitmap for each Git object type, stored in the following\n+\t\torder:\n+\n+\t\t\t- Commits\n+\t\t\t- Trees\n+\t\t\t- Blobs\n+\t\t\t- Tags\n+\n+\t\tIn each bitmap, the `n`th bit is set to true if the `n`th object\n+\t\tin the packfile is of that type.\n+\n+\t\tThe obvious consequence is that the OR of all 4 bitmaps will result\n+\t\tin a full set (all bits set), and the AND of all 4 bitmaps will\n+\t\tresult in an empty bitmap (no bits set).\n+\n+\t- N entries with compressed bitmaps, one for each indexed commit\n+\n+\t\tWhere `N` is the total amount of entries in this bitmap index.\n+\t\tEach entry contains the following:\n+\n+\t\t- 4-byte object position (network byte order)\n+\t\t\tThe position **in the index for the packfile** where the\n+\t\t\tbitmap for this commit is found.\n+\n+\t\t- 1-byte XOR-offset\n+\t\t\tThe xor offset used to compress this bitmap. For an entry\n+\t\t\tin position `x`, a XOR offset of `y` means that the actual\n+\t\t\tbitmap representing this commit is composed by XORing the\n+\t\t\tbitmap for this entry with the bitmap in entry `x-y` (i.e.\n+\t\t\tthe bitmap `y` entries before this one).\n+\n+\t\t\tNote that this compression can be recursive. In order to\n+\t\t\tXOR this entry with a previous one, the previous entry needs\n+\t\t\tto be decompressed first, and so on.\n+\n+\t\t\tThe hard-limit for this offset is 160 (an entry can only be\n+\t\t\txor'ed against one of the 160 entries preceding it). This\n+\t\t\tnumber is always positive, and hence entries are always xor'ed\n+\t\t\twith **previous** bitmaps, not bitmaps that will come afterwards\n+\t\t\tin the index.\n+\n+\t\t- 1-byte flags for this bitmap\n+\t\t\tAt the moment the only available flag is `0x1`, which hints\n+\t\t\tthat this bitmap can be re-used when rebuilding bitmap indexes\n+\t\t\tfor the repository.\n+\n+\t\t- The compressed bitmap itself, see Appendix A.\n+\n+== Appendix A: Serialization format for an EWAH bitmap\n+\n+Ewah bitmaps are serialized in the same protocol as the JAVAEWAH\n+library, making them backwards compatible with the JGit\n+implementation:\n+\n+\t- 4-byte number of bits of the resulting UNCOMPRESSED bitmap\n+\n+\t- 4-byte number of words of the COMPRESSED bitmap, when stored\n+\n+\t- N x 8-byte words, as specified by the previous field\n+\n+\t\tThis is the actual content of the compressed bitmap.\n+\n+\t- 4-byte position of the current RLW for the compressed\n+\t\tbitmap\n+\n+All words are stored in network byte order for their corresponding\n+sizes.\n+\n+The compressed bitmap is stored in a form of run-length encoding, as\n+follows.  It consists of a concatenation of an arbitrary number of\n+chunks.  Each chunk consists of one or more 64-bit words\n+\n+     H  L_1  L_2  L_3 .... L_M\n+\n+H is called RLW (run length word).  It consists of (from lower to higher\n+order bits):\n+\n+     - 1 bit: the repeated bit B\n+\n+     - 32 bits: repetition count K (unsigned)\n+\n+     - 31 bits: literal word count M (unsigned)\n+\n+The bitstream represented by the above chunk is then:\n+\n+     - K repetitions of B\n+\n+     - The bits stored in `L_1` through `L_M`.  Within a word, bits at\n+       lower order come earlier in the stream than those at higher\n+       order.\n+\n+The next word after `L_M` (if any) must again be a RLW, for the next\n+chunk.  For efficient appending to the bitstream, the EWAH stores a\n+pointer to the last RLW in the stream.\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230608","messageId":"20131114124432.GJ10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:44:32Z","receivedAt":"2013-11-14T12:44:32Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nA bitmap index is a `.bitmap` file that can be found inside\n`$GIT_DIR/objects/pack/`, next to its corresponding packfile, and\ncontains precalculated reachability information for selected commits.\nThe full specification of the format for these bitmap indexes can be found\nin `Documentation/technical/bitmap-format.txt`.\n\nFor a given commit SHA1, if it happens to be available in the bitmap\nindex, its bitmap will represent every single object that is reachable\nfrom the commit itself. The nth bit in the bitmap is the nth object in\nthe packfile; if it's set to 1, the object is reachable.\n\nBy using the bitmaps available in the index, this commit implements\nseveral new functions:\n\n\t- `prepare_bitmap_git`\n\t- `prepare_bitmap_walk`\n\t- `traverse_bitmap_commit_list`\n\t- `reuse_partial_packfile_from_bitmap`\n\nThe `prepare_bitmap_walk` function tries to build a bitmap of all the\nobjects that can be reached from the commit roots of a given `rev_info`\nstruct by using the following algorithm:\n\n- If all the interesting commits for a revision walk are available in\nthe index, the resulting reachability bitmap is the bitwise OR of all\nthe individual bitmaps.\n\n- When the full set of WANTs is not available in the index, we perform a\npartial revision walk using the commits that don't have bitmaps as\nroots, and limiting the revision walk as soon as we reach a commit that\nhas a corresponding bitmap. The earlier OR'ed bitmap with all the\nindexed commits can now be completed as this walk progresses, so the end\nresult is the full reachability list.\n\n- For revision walks with a HAVEs set (a set of commits that are deemed\nuninteresting), first we perform the same method as for the WANTs, but\nusing our HAVEs as roots, in order to obtain a full reachability bitmap\nof all the uninteresting commits. This bitmap then can be used to:\n\n\ta) limit the subsequent walk when building the WANTs bitmap\n\tb) finding the final set of interesting commits by performing an\n\t   AND-NOT of the WANTs and the HAVEs.\n\nIf `prepare_bitmap_walk` runs successfully, the resulting bitmap is\nstored and the equivalent of a `traverse_commit_list` call can be\nperformed by using `traverse_bitmap_commit_list`; the bitmap version\nof this call yields the objects straight from the packfile index\n(without having to look them up or parse them) and hence is several\norders of magnitude faster.\n\nAs an extra optimization, when `prepare_bitmap_walk` succeeds, the\n`reuse_partial_packfile_from_bitmap` call can be attempted: it will find\nthe amount of objects at the beginning of the on-disk packfile that can\nbe reused as-is, and return an offset into the packfile. The source\npackfile can then be loaded and the bytes up to `offset` can be written\ndirectly to the result without having to consider the entires inside the\npackfile individually.\n\nIf the `prepare_bitmap_walk` call fails (e.g. because no bitmap files\nare available), the `rev_info` struct is left untouched, and can be used\nto perform a manual rev-walk using `traverse_commit_list`.\n\nHence, this new set of functions are a generic API that allows to\nperform the equivalent of\n\n\tgit rev-list --objects [roots...] [^uninteresting...]\n\nfor any set of commits, even if they don't have specific bitmaps\ngenerated for them.\n\nIn further patches, we'll use this bitmap traversal optimization to\nspeed up the `pack-objects` and `rev-list` commands.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Makefile      |   2 +\n khash.h       | 338 ++++++++++++++++++++\n pack-bitmap.c | 970 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n pack-bitmap.h |  43 +++\n 4 files changed, 1353 insertions(+)\n create mode 100644 khash.h\n create mode 100644 pack-bitmap.c\n create mode 100644 pack-bitmap.h\n\ndiff --git a/Makefile b/Makefile\nindex 64a1ed7..b983d78 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -699,6 +699,7 @@ LIB_H += object.h\n LIB_H += pack-objects.h\n LIB_H += pack-revindex.h\n LIB_H += pack.h\n+LIB_H += pack-bitmap.h\n LIB_H += parse-options.h\n LIB_H += patch-ids.h\n LIB_H += pathspec.h\n@@ -837,6 +838,7 @@ LIB_OBJS += notes-cache.o\n LIB_OBJS += notes-merge.o\n LIB_OBJS += notes-utils.o\n LIB_OBJS += object.o\n+LIB_OBJS += pack-bitmap.o\n LIB_OBJS += pack-check.o\n LIB_OBJS += pack-objects.o\n LIB_OBJS += pack-revindex.o\ndiff --git a/khash.h b/khash.h\nnew file mode 100644\nindex 0000000..57ff603\n--- /dev/null\n+++ b/khash.h\n@@ -0,0 +1,338 @@\n+/* The MIT License\n+\n+   Copyright (c) 2008, 2009, 2011 by Attractive Chaos <attractor@live.co.uk>\n+\n+   Permission is hereby granted, free of charge, to any person obtaining\n+   a copy of this software and associated documentation files (the\n+   \"Software\"), to deal in the Software without restriction, including\n+   without limitation the rights to use, copy, modify, merge, publish,\n+   distribute, sublicense, and/or sell copies of the Software, and to\n+   permit persons to whom the Software is furnished to do so, subject to\n+   the following conditions:\n+\n+   The above copyright notice and this permission notice shall be\n+   included in all copies or substantial portions of the Software.\n+\n+   THE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND,\n+   EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF\n+   MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND\n+   NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS\n+   BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN\n+   ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN\n+   CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\n+   SOFTWARE.\n+*/\n+\n+#ifndef __AC_KHASH_H\n+#define __AC_KHASH_H\n+\n+#define AC_VERSION_KHASH_H \"0.2.8\"\n+\n+typedef uint32_t khint32_t;\n+typedef uint64_t khint64_t;\n+\n+typedef khint32_t khint_t;\n+typedef khint_t khiter_t;\n+\n+#define __ac_isempty(flag, i) ((flag[i>>4]>>((i&0xfU)<<1))&2)\n+#define __ac_isdel(flag, i) ((flag[i>>4]>>((i&0xfU)<<1))&1)\n+#define __ac_iseither(flag, i) ((flag[i>>4]>>((i&0xfU)<<1))&3)\n+#define __ac_set_isdel_false(flag, i) (flag[i>>4]&=~(1ul<<((i&0xfU)<<1)))\n+#define __ac_set_isempty_false(flag, i) (flag[i>>4]&=~(2ul<<((i&0xfU)<<1)))\n+#define __ac_set_isboth_false(flag, i) (flag[i>>4]&=~(3ul<<((i&0xfU)<<1)))\n+#define __ac_set_isdel_true(flag, i) (flag[i>>4]|=1ul<<((i&0xfU)<<1))\n+\n+#define __ac_fsize(m) ((m) < 16? 1 : (m)>>4)\n+\n+#define kroundup32(x) (--(x), (x)|=(x)>>1, (x)|=(x)>>2, (x)|=(x)>>4, (x)|=(x)>>8, (x)|=(x)>>16, ++(x))\n+\n+static inline khint_t __ac_X31_hash_string(const char *s)\n+{\n+\tkhint_t h = (khint_t)*s;\n+\tif (h) for (++s ; *s; ++s) h = (h << 5) - h + (khint_t)*s;\n+\treturn h;\n+}\n+\n+#define kh_str_hash_func(key) __ac_X31_hash_string(key)\n+#define kh_str_hash_equal(a, b) (strcmp(a, b) == 0)\n+\n+static const double __ac_HASH_UPPER = 0.77;\n+\n+#define __KHASH_TYPE(name, khkey_t, khval_t) \\\n+\ttypedef struct { \\\n+\t\tkhint_t n_buckets, size, n_occupied, upper_bound; \\\n+\t\tkhint32_t *flags; \\\n+\t\tkhkey_t *keys; \\\n+\t\tkhval_t *vals; \\\n+\t} kh_##name##_t;\n+\n+#define __KHASH_PROTOTYPES(name, khkey_t, khval_t)\t \t\t\t\t\t\\\n+\textern kh_##name##_t *kh_init_##name(void);\t\t\t\t\t\t\t\\\n+\textern void kh_destroy_##name(kh_##name##_t *h);\t\t\t\t\t\\\n+\textern void kh_clear_##name(kh_##name##_t *h);\t\t\t\t\t\t\\\n+\textern khint_t kh_get_##name(const kh_##name##_t *h, khkey_t key); \t\\\n+\textern int kh_resize_##name(kh_##name##_t *h, khint_t new_n_buckets); \\\n+\textern khint_t kh_put_##name(kh_##name##_t *h, khkey_t key, int *ret); \\\n+\textern void kh_del_##name(kh_##name##_t *h, khint_t x);\n+\n+#define __KHASH_IMPL(name, SCOPE, khkey_t, khval_t, kh_is_map, __hash_func, __hash_equal) \\\n+\tSCOPE kh_##name##_t *kh_init_##name(void) {\t\t\t\t\t\t\t\\\n+\t\treturn (kh_##name##_t*)xcalloc(1, sizeof(kh_##name##_t));\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE void kh_destroy_##name(kh_##name##_t *h)\t\t\t\t\t\t\\\n+\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (h) {\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tfree((void *)h->keys); free(h->flags);\t\t\t\t\t\\\n+\t\t\tfree((void *)h->vals);\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tfree(h);\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE void kh_clear_##name(kh_##name##_t *h)\t\t\t\t\t\t\\\n+\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (h && h->flags) {\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tmemset(h->flags, 0xaa, __ac_fsize(h->n_buckets) * sizeof(khint32_t)); \\\n+\t\t\th->size = h->n_occupied = 0;\t\t\t\t\t\t\t\t\\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE khint_t kh_get_##name(const kh_##name##_t *h, khkey_t key) \t\\\n+\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (h->n_buckets) {\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tkhint_t k, i, last, mask, step = 0; \\\n+\t\t\tmask = h->n_buckets - 1;\t\t\t\t\t\t\t\t\t\\\n+\t\t\tk = __hash_func(key); i = k & mask;\t\t\t\t\t\t\t\\\n+\t\t\tlast = i; \\\n+\t\t\twhile (!__ac_isempty(h->flags, i) && (__ac_isdel(h->flags, i) || !__hash_equal(h->keys[i], key))) { \\\n+\t\t\t\ti = (i + (++step)) & mask; \\\n+\t\t\t\tif (i == last) return h->n_buckets;\t\t\t\t\t\t\\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\treturn __ac_iseither(h->flags, i)? h->n_buckets : i;\t\t\\\n+\t\t} else return 0;\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE int kh_resize_##name(kh_##name##_t *h, khint_t new_n_buckets) \\\n+\t{ /* This function uses 0.25*n_buckets bytes of working space instead of [sizeof(key_t+val_t)+.25]*n_buckets. */ \\\n+\t\tkhint32_t *new_flags = NULL;\t\t\t\t\t\t\t\t\t\t\\\n+\t\tkhint_t j = 1;\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tkroundup32(new_n_buckets); \t\t\t\t\t\t\t\t\t\\\n+\t\t\tif (new_n_buckets < 4) new_n_buckets = 4;\t\t\t\t\t\\\n+\t\t\tif (h->size >= (khint_t)(new_n_buckets * __ac_HASH_UPPER + 0.5)) j = 0;\t/* requested size is too small */ \\\n+\t\t\telse { /* hash table size to be changed (shrink or expand); rehash */ \\\n+\t\t\t\tnew_flags = (khint32_t*)xmalloc(__ac_fsize(new_n_buckets) * sizeof(khint32_t));\t\\\n+\t\t\t\tif (!new_flags) return -1;\t\t\t\t\t\t\t\t\\\n+\t\t\t\tmemset(new_flags, 0xaa, __ac_fsize(new_n_buckets) * sizeof(khint32_t)); \\\n+\t\t\t\tif (h->n_buckets < new_n_buckets) {\t/* expand */\t\t\\\n+\t\t\t\t\tkhkey_t *new_keys = (khkey_t*)xrealloc((void *)h->keys, new_n_buckets * sizeof(khkey_t)); \\\n+\t\t\t\t\tif (!new_keys) return -1;\t\t\t\t\t\t\t\\\n+\t\t\t\t\th->keys = new_keys;\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\tif (kh_is_map) {\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\t\tkhval_t *new_vals = (khval_t*)xrealloc((void *)h->vals, new_n_buckets * sizeof(khval_t)); \\\n+\t\t\t\t\t\tif (!new_vals) return -1;\t\t\t\t\t\t\\\n+\t\t\t\t\t\th->vals = new_vals;\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t} /* otherwise shrink */\t\t\t\t\t\t\t\t\\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (j) { /* rehashing is needed */\t\t\t\t\t\t\t\t\\\n+\t\t\tfor (j = 0; j != h->n_buckets; ++j) {\t\t\t\t\t\t\\\n+\t\t\t\tif (__ac_iseither(h->flags, j) == 0) {\t\t\t\t\t\\\n+\t\t\t\t\tkhkey_t key = h->keys[j];\t\t\t\t\t\t\t\\\n+\t\t\t\t\tkhval_t val;\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\tkhint_t new_mask;\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\tnew_mask = new_n_buckets - 1; \t\t\t\t\t\t\\\n+\t\t\t\t\tif (kh_is_map) val = h->vals[j];\t\t\t\t\t\\\n+\t\t\t\t\t__ac_set_isdel_true(h->flags, j);\t\t\t\t\t\\\n+\t\t\t\t\twhile (1) { /* kick-out process; sort of like in Cuckoo hashing */ \\\n+\t\t\t\t\t\tkhint_t k, i, step = 0; \\\n+\t\t\t\t\t\tk = __hash_func(key);\t\t\t\t\t\t\t\\\n+\t\t\t\t\t\ti = k & new_mask;\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\t\twhile (!__ac_isempty(new_flags, i)) i = (i + (++step)) & new_mask; \\\n+\t\t\t\t\t\t__ac_set_isempty_false(new_flags, i);\t\t\t\\\n+\t\t\t\t\t\tif (i < h->n_buckets && __ac_iseither(h->flags, i) == 0) { /* kick out the existing element */ \\\n+\t\t\t\t\t\t\t{ khkey_t tmp = h->keys[i]; h->keys[i] = key; key = tmp; } \\\n+\t\t\t\t\t\t\tif (kh_is_map) { khval_t tmp = h->vals[i]; h->vals[i] = val; val = tmp; } \\\n+\t\t\t\t\t\t\t__ac_set_isdel_true(h->flags, i); /* mark it as deleted in the old hash table */ \\\n+\t\t\t\t\t\t} else { /* write the element and jump out of the loop */ \\\n+\t\t\t\t\t\t\th->keys[i] = key;\t\t\t\t\t\t\t\\\n+\t\t\t\t\t\t\tif (kh_is_map) h->vals[i] = val;\t\t\t\\\n+\t\t\t\t\t\t\tbreak;\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tif (h->n_buckets > new_n_buckets) { /* shrink the hash table */ \\\n+\t\t\t\th->keys = (khkey_t*)xrealloc((void *)h->keys, new_n_buckets * sizeof(khkey_t)); \\\n+\t\t\t\tif (kh_is_map) h->vals = (khval_t*)xrealloc((void *)h->vals, new_n_buckets * sizeof(khval_t)); \\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tfree(h->flags); /* free the working space */\t\t\t\t\\\n+\t\t\th->flags = new_flags;\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\th->n_buckets = new_n_buckets;\t\t\t\t\t\t\t\t\\\n+\t\t\th->n_occupied = h->size;\t\t\t\t\t\t\t\t\t\\\n+\t\t\th->upper_bound = (khint_t)(h->n_buckets * __ac_HASH_UPPER + 0.5); \\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\treturn 0;\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE khint_t kh_put_##name(kh_##name##_t *h, khkey_t key, int *ret) \\\n+\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tkhint_t x;\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (h->n_occupied >= h->upper_bound) { /* update the hash table */ \\\n+\t\t\tif (h->n_buckets > (h->size<<1)) {\t\t\t\t\t\t\t\\\n+\t\t\t\tif (kh_resize_##name(h, h->n_buckets - 1) < 0) { /* clear \"deleted\" elements */ \\\n+\t\t\t\t\t*ret = -1; return h->n_buckets;\t\t\t\t\t\t\\\n+\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t} else if (kh_resize_##name(h, h->n_buckets + 1) < 0) { /* expand the hash table */ \\\n+\t\t\t\t*ret = -1; return h->n_buckets;\t\t\t\t\t\t\t\\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t} /* TODO: to implement automatically shrinking; resize() already support shrinking */ \\\n+\t\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tkhint_t k, i, site, last, mask = h->n_buckets - 1, step = 0; \\\n+\t\t\tx = site = h->n_buckets; k = __hash_func(key); i = k & mask; \\\n+\t\t\tif (__ac_isempty(h->flags, i)) x = i; /* for speed up */\t\\\n+\t\t\telse {\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\tlast = i; \\\n+\t\t\t\twhile (!__ac_isempty(h->flags, i) && (__ac_isdel(h->flags, i) || !__hash_equal(h->keys[i], key))) { \\\n+\t\t\t\t\tif (__ac_isdel(h->flags, i)) site = i;\t\t\t\t\\\n+\t\t\t\t\ti = (i + (++step)) & mask; \\\n+\t\t\t\t\tif (i == last) { x = site; break; }\t\t\t\t\t\\\n+\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\tif (x == h->n_buckets) {\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\tif (__ac_isempty(h->flags, i) && site != h->n_buckets) x = site; \\\n+\t\t\t\t\telse x = i;\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (__ac_isempty(h->flags, x)) { /* not present at all */\t\t\\\n+\t\t\th->keys[x] = key;\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t__ac_set_isboth_false(h->flags, x);\t\t\t\t\t\t\t\\\n+\t\t\t++h->size; ++h->n_occupied;\t\t\t\t\t\t\t\t\t\\\n+\t\t\t*ret = 1;\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t} else if (__ac_isdel(h->flags, x)) { /* deleted */\t\t\t\t\\\n+\t\t\th->keys[x] = key;\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t__ac_set_isboth_false(h->flags, x);\t\t\t\t\t\t\t\\\n+\t\t\t++h->size;\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t*ret = 2;\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t} else *ret = 0; /* Don't touch h->keys[x] if present and not deleted */ \\\n+\t\treturn x;\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE void kh_del_##name(kh_##name##_t *h, khint_t x)\t\t\t\t\\\n+\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (x != h->n_buckets && !__ac_iseither(h->flags, x)) {\t\t\t\\\n+\t\t\t__ac_set_isdel_true(h->flags, x);\t\t\t\t\t\t\t\\\n+\t\t\t--h->size;\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\n+\n+#define KHASH_DECLARE(name, khkey_t, khval_t)\t\t \t\t\t\t\t\\\n+\t__KHASH_TYPE(name, khkey_t, khval_t) \t\t\t\t\t\t\t\t\\\n+\t__KHASH_PROTOTYPES(name, khkey_t, khval_t)\n+\n+#define KHASH_INIT2(name, SCOPE, khkey_t, khval_t, kh_is_map, __hash_func, __hash_equal) \\\n+\t__KHASH_TYPE(name, khkey_t, khval_t) \t\t\t\t\t\t\t\t\\\n+\t__KHASH_IMPL(name, SCOPE, khkey_t, khval_t, kh_is_map, __hash_func, __hash_equal)\n+\n+#define KHASH_INIT(name, khkey_t, khval_t, kh_is_map, __hash_func, __hash_equal) \\\n+\tKHASH_INIT2(name, static inline, khkey_t, khval_t, kh_is_map, __hash_func, __hash_equal)\n+\n+/* Other convenient macros... */\n+\n+/*! @function\n+  @abstract     Test whether a bucket contains data.\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @param  x     Iterator to the bucket [khint_t]\n+  @return       1 if containing data; 0 otherwise [int]\n+ */\n+#define kh_exist(h, x) (!__ac_iseither((h)->flags, (x)))\n+\n+/*! @function\n+  @abstract     Get key given an iterator\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @param  x     Iterator to the bucket [khint_t]\n+  @return       Key [type of keys]\n+ */\n+#define kh_key(h, x) ((h)->keys[x])\n+\n+/*! @function\n+  @abstract     Get value given an iterator\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @param  x     Iterator to the bucket [khint_t]\n+  @return       Value [type of values]\n+  @discussion   For hash sets, calling this results in segfault.\n+ */\n+#define kh_val(h, x) ((h)->vals[x])\n+\n+/*! @function\n+  @abstract     Alias of kh_val()\n+ */\n+#define kh_value(h, x) ((h)->vals[x])\n+\n+/*! @function\n+  @abstract     Get the start iterator\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @return       The start iterator [khint_t]\n+ */\n+#define kh_begin(h) (khint_t)(0)\n+\n+/*! @function\n+  @abstract     Get the end iterator\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @return       The end iterator [khint_t]\n+ */\n+#define kh_end(h) ((h)->n_buckets)\n+\n+/*! @function\n+  @abstract     Get the number of elements in the hash table\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @return       Number of elements in the hash table [khint_t]\n+ */\n+#define kh_size(h) ((h)->size)\n+\n+/*! @function\n+  @abstract     Get the number of buckets in the hash table\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @return       Number of buckets in the hash table [khint_t]\n+ */\n+#define kh_n_buckets(h) ((h)->n_buckets)\n+\n+/*! @function\n+  @abstract     Iterate over the entries in the hash table\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @param  kvar  Variable to which key will be assigned\n+  @param  vvar  Variable to which value will be assigned\n+  @param  code  Block of code to execute\n+ */\n+#define kh_foreach(h, kvar, vvar, code) { khint_t __i;\t\t\\\n+\tfor (__i = kh_begin(h); __i != kh_end(h); ++__i) {\t\t\\\n+\t\tif (!kh_exist(h,__i)) continue;\t\t\t\t\t\t\\\n+\t\t(kvar) = kh_key(h,__i);\t\t\t\t\t\t\t\t\\\n+\t\t(vvar) = kh_val(h,__i);\t\t\t\t\t\t\t\t\\\n+\t\tcode;\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t} }\n+\n+/*! @function\n+  @abstract     Iterate over the values in the hash table\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @param  vvar  Variable to which value will be assigned\n+  @param  code  Block of code to execute\n+ */\n+#define kh_foreach_value(h, vvar, code) { khint_t __i;\t\t\\\n+\tfor (__i = kh_begin(h); __i != kh_end(h); ++__i) {\t\t\\\n+\t\tif (!kh_exist(h,__i)) continue;\t\t\t\t\t\t\\\n+\t\t(vvar) = kh_val(h,__i);\t\t\t\t\t\t\t\t\\\n+\t\tcode;\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t} }\n+\n+static inline khint_t __kh_oid_hash(const unsigned char *oid)\n+{\n+\tkhint_t hash;\n+\tmemcpy(&hash, oid, sizeof(hash));\n+\treturn hash;\n+}\n+\n+#define __kh_oid_cmp(a, b) (hashcmp(a, b) == 0)\n+\n+KHASH_INIT(sha1, const unsigned char *, void *, 1, __kh_oid_hash, __kh_oid_cmp)\n+typedef kh_sha1_t khash_sha1;\n+\n+KHASH_INIT(sha1_pos, const unsigned char *, int, 1, __kh_oid_hash, __kh_oid_cmp)\n+typedef kh_sha1_pos_t khash_sha1_pos;\n+\n+#endif /* __AC_KHASH_H */\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nnew file mode 100644\nindex 0000000..340f02b\n--- /dev/null\n+++ b/pack-bitmap.c\n@@ -0,0 +1,970 @@\n+#include \"cache.h\"\n+#include \"commit.h\"\n+#include \"tag.h\"\n+#include \"diff.h\"\n+#include \"revision.h\"\n+#include \"progress.h\"\n+#include \"list-objects.h\"\n+#include \"pack.h\"\n+#include \"pack-bitmap.h\"\n+#include \"pack-revindex.h\"\n+#include \"pack-objects.h\"\n+\n+/*\n+ * An entry on the bitmap index, representing the bitmap for a given\n+ * commit.\n+ */\n+struct stored_bitmap {\n+\tunsigned char sha1[20];\n+\tstruct ewah_bitmap *root;\n+\tstruct stored_bitmap *xor;\n+\tint flags;\n+};\n+\n+/*\n+ * The currently active bitmap index. By design, repositories only have\n+ * a single bitmap index available (the index for the biggest packfile in\n+ * the repository), since bitmap indexes need full closure.\n+ *\n+ * If there is more than one bitmap index available (e.g. because of alternates),\n+ * the active bitmap index is the largest one.\n+ */\n+static struct bitmap_index {\n+\t/* Packfile to which this bitmap index belongs to */\n+\tstruct packed_git *pack;\n+\n+\t/* reverse index for the packfile */\n+\tstruct pack_revindex *reverse_index;\n+\n+\t/*\n+\t * Mark the first `reuse_objects` in the packfile as reused:\n+\t * they will be sent as-is without using them for repacking\n+\t * calculations\n+\t */\n+\tuint32_t reuse_objects;\n+\n+\t/* mmapped buffer of the whole bitmap index */\n+\tunsigned char *map;\n+\tsize_t map_size; /* size of the mmaped buffer */\n+\tsize_t map_pos; /* current position when loading the index */\n+\n+\t/*\n+\t * Type indexes.\n+\t *\n+\t * Each bitmap marks which objects in the packfile  are of the given\n+\t * type. This provides type information when yielding the objects from\n+\t * the packfile during a walk, which allows for better delta bases.\n+\t */\n+\tstruct ewah_bitmap *commits;\n+\tstruct ewah_bitmap *trees;\n+\tstruct ewah_bitmap *blobs;\n+\tstruct ewah_bitmap *tags;\n+\n+\t/* Map from SHA1 -> `stored_bitmap` for all the bitmapped comits */\n+\tkhash_sha1 *bitmaps;\n+\n+\t/* Number of bitmapped commits */\n+\tuint32_t entry_count;\n+\n+\t/*\n+\t * Extended index.\n+\t *\n+\t * When trying to perform bitmap operations with objects that are not\n+\t * packed in `pack`, these objects are added to this \"fake index\" and\n+\t * are assumed to appear at the end of the packfile for all operations\n+\t */\n+\tstruct eindex {\n+\t\tstruct object **objects;\n+\t\tuint32_t *hashes;\n+\t\tuint32_t count, alloc;\n+\t\tkhash_sha1_pos *positions;\n+\t} ext_index;\n+\n+\t/* Bitmap result of the last performed walk */\n+\tstruct bitmap *result;\n+\n+\t/* Version of the bitmap index */\n+\tunsigned int version;\n+\n+\tunsigned loaded : 1;\n+\n+} bitmap_git;\n+\n+static struct ewah_bitmap *lookup_stored_bitmap(struct stored_bitmap *st)\n+{\n+\tstruct ewah_bitmap *parent;\n+\tstruct ewah_bitmap *composed;\n+\n+\tif (st->xor == NULL)\n+\t\treturn st->root;\n+\n+\tcomposed = ewah_pool_new();\n+\tparent = lookup_stored_bitmap(st->xor);\n+\tewah_xor(st->root, parent, composed);\n+\n+\tewah_pool_free(st->root);\n+\tst->root = composed;\n+\tst->xor = NULL;\n+\n+\treturn composed;\n+}\n+\n+/*\n+ * Read a bitmap from the current read position on the mmaped\n+ * index, and increase the read position accordingly\n+ */\n+static struct ewah_bitmap *read_bitmap_1(struct bitmap_index *index)\n+{\n+\tstruct ewah_bitmap *b = ewah_pool_new();\n+\n+\tint bitmap_size = ewah_read_mmap(b,\n+\t\tindex->map + index->map_pos,\n+\t\tindex->map_size - index->map_pos);\n+\n+\tif (bitmap_size < 0) {\n+\t\terror(\"Failed to load bitmap index (corrupted?)\");\n+\t\tewah_pool_free(b);\n+\t\treturn NULL;\n+\t}\n+\n+\tindex->map_pos += bitmap_size;\n+\treturn b;\n+}\n+\n+static int load_bitmap_header(struct bitmap_index *index)\n+{\n+\tstruct bitmap_disk_header *header = (void *)index->map;\n+\n+\tif (index->map_size < sizeof(*header) + 20)\n+\t\treturn error(\"Corrupted bitmap index (missing header data)\");\n+\n+\tif (memcmp(header->magic, BITMAP_IDX_SIGNATURE, sizeof(BITMAP_IDX_SIGNATURE)) != 0)\n+\t\treturn error(\"Corrupted bitmap index file (wrong header)\");\n+\n+\tindex->version = ntohs(header->version);\n+\tif (index->version != 1)\n+\t\treturn error(\"Unsupported version for bitmap index file (%d)\", index->version);\n+\n+\t/* Parse known bitmap format options */\n+\t{\n+\t\tuint32_t flags = ntohs(header->options);\n+\n+\t\tif ((flags & BITMAP_OPT_FULL_DAG) == 0)\n+\t\t\treturn error(\"Unsupported options for bitmap index file \"\n+\t\t\t\t\"(Git requires BITMAP_OPT_FULL_DAG)\");\n+\t}\n+\n+\tindex->entry_count = ntohl(header->entry_count);\n+\tindex->map_pos += sizeof(*header);\n+\treturn 0;\n+}\n+\n+static struct stored_bitmap *store_bitmap(struct bitmap_index *index,\n+\t\t\t\t\t  struct ewah_bitmap *root,\n+\t\t\t\t\t  const unsigned char *sha1,\n+\t\t\t\t\t  struct stored_bitmap *xor_with,\n+\t\t\t\t\t  int flags)\n+{\n+\tstruct stored_bitmap *stored;\n+\tkhiter_t hash_pos;\n+\tint ret;\n+\n+\tstored = xmalloc(sizeof(struct stored_bitmap));\n+\tstored->root = root;\n+\tstored->xor = xor_with;\n+\tstored->flags = flags;\n+\thashcpy(stored->sha1, sha1);\n+\n+\thash_pos = kh_put_sha1(index->bitmaps, stored->sha1, &ret);\n+\n+\t/* a 0 return code means the insertion succeeded with no changes,\n+\t * because the SHA1 already existed on the map. this is bad, there\n+\t * shouldn't be duplicated commits in the index */\n+\tif (ret == 0) {\n+\t\terror(\"Duplicate entry in bitmap index: %s\", sha1_to_hex(sha1));\n+\t\treturn NULL;\n+\t}\n+\n+\tkh_value(index->bitmaps, hash_pos) = stored;\n+\treturn stored;\n+}\n+\n+static int load_bitmap_entries_v1(struct bitmap_index *index)\n+{\n+\tstatic const size_t MAX_XOR_OFFSET = 160;\n+\n+\tuint32_t i;\n+\tstruct stored_bitmap **recent_bitmaps;\n+\tstruct bitmap_disk_entry *entry;\n+\n+\trecent_bitmaps = xcalloc(MAX_XOR_OFFSET, sizeof(struct stored_bitmap));\n+\n+\tfor (i = 0; i < index->entry_count; ++i) {\n+\t\tint xor_offset, flags;\n+\t\tstruct ewah_bitmap *bitmap = NULL;\n+\t\tstruct stored_bitmap *xor_bitmap = NULL;\n+\t\tuint32_t commit_idx_pos;\n+\t\tconst unsigned char *sha1;\n+\n+\t\tentry = (struct bitmap_disk_entry *)(index->map + index->map_pos);\n+\t\tindex->map_pos += sizeof(struct bitmap_disk_entry);\n+\n+\t\tcommit_idx_pos = ntohl(entry->object_pos);\n+\t\tsha1 = nth_packed_object_sha1(index->pack, commit_idx_pos);\n+\n+\t\txor_offset = (int)entry->xor_offset;\n+\t\tflags = (int)entry->flags;\n+\n+\t\tbitmap = read_bitmap_1(index);\n+\t\tif (!bitmap)\n+\t\t\treturn -1;\n+\n+\t\tif (xor_offset > MAX_XOR_OFFSET || xor_offset > i)\n+\t\t\treturn error(\"Corrupted bitmap pack index\");\n+\n+\t\tif (xor_offset > 0) {\n+\t\t\txor_bitmap = recent_bitmaps[(i - xor_offset) % MAX_XOR_OFFSET];\n+\n+\t\t\tif (xor_bitmap == NULL)\n+\t\t\t\treturn error(\"Invalid XOR offset in bitmap pack index\");\n+\t\t}\n+\n+\t\trecent_bitmaps[i % MAX_XOR_OFFSET] = store_bitmap(\n+\t\t\tindex, bitmap, sha1, xor_bitmap, flags);\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int open_pack_bitmap_1(struct packed_git *packfile)\n+{\n+\tint fd;\n+\tstruct stat st;\n+\tchar *idx_name;\n+\n+\tif (open_pack_index(packfile))\n+\t\treturn -1;\n+\n+\tidx_name = pack_bitmap_filename(packfile);\n+\tfd = git_open_noatime(idx_name);\n+\tfree(idx_name);\n+\n+\tif (fd < 0)\n+\t\treturn -1;\n+\n+\tif (fstat(fd, &st)) {\n+\t\tclose(fd);\n+\t\treturn -1;\n+\t}\n+\n+\tif (bitmap_git.pack) {\n+\t\twarning(\"ignoring extra bitmap file: %s\", idx_name);\n+\t\tclose(fd);\n+\t\treturn -1;\n+\t}\n+\n+\tbitmap_git.pack = packfile;\n+\tbitmap_git.map_size = xsize_t(st.st_size);\n+\tbitmap_git.map = xmmap(NULL, bitmap_git.map_size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\tbitmap_git.map_pos = 0;\n+\tclose(fd);\n+\n+\tif (load_bitmap_header(&bitmap_git) < 0) {\n+\t\tmunmap(bitmap_git.map, bitmap_git.map_size);\n+\t\tbitmap_git.map = NULL;\n+\t\tbitmap_git.map_size = 0;\n+\t\treturn -1;\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int load_pack_bitmap(void)\n+{\n+\tassert(bitmap_git.map && !bitmap_git.loaded);\n+\n+\tbitmap_git.bitmaps = kh_init_sha1();\n+\tbitmap_git.ext_index.positions = kh_init_sha1_pos();\n+\tbitmap_git.reverse_index = revindex_for_pack(bitmap_git.pack);\n+\n+\tif (!(bitmap_git.commits = read_bitmap_1(&bitmap_git)) ||\n+\t\t!(bitmap_git.trees = read_bitmap_1(&bitmap_git)) ||\n+\t\t!(bitmap_git.blobs = read_bitmap_1(&bitmap_git)) ||\n+\t\t!(bitmap_git.tags = read_bitmap_1(&bitmap_git)))\n+\t\tgoto failed;\n+\n+\tif (load_bitmap_entries_v1(&bitmap_git) < 0)\n+\t\tgoto failed;\n+\n+\tbitmap_git.loaded = 1;\n+\treturn 0;\n+\n+failed:\n+\tmunmap(bitmap_git.map, bitmap_git.map_size);\n+\tbitmap_git.map = NULL;\n+\tbitmap_git.map_size = 0;\n+\treturn -1;\n+}\n+\n+char *pack_bitmap_filename(struct packed_git *p)\n+{\n+\tchar *idx_name;\n+\tint len;\n+\n+\tlen = strlen(p->pack_name) - strlen(\".pack\");\n+\tidx_name = xmalloc(len + strlen(\".bitmap\") + 1);\n+\n+\tmemcpy(idx_name, p->pack_name, len);\n+\tmemcpy(idx_name + len, \".bitmap\", strlen(\".bitmap\") + 1);\n+\n+\treturn idx_name;\n+}\n+\n+static int open_pack_bitmap(void)\n+{\n+\tstruct packed_git *p;\n+\tint ret = -1;\n+\n+\tassert(!bitmap_git.map && !bitmap_git.loaded);\n+\n+\tprepare_packed_git();\n+\tfor (p = packed_git; p; p = p->next) {\n+\t\tif (open_pack_bitmap_1(p) == 0)\n+\t\t\tret = 0;\n+\t}\n+\n+\treturn ret;\n+}\n+\n+int prepare_bitmap_git(void)\n+{\n+\tif (bitmap_git.loaded)\n+\t\treturn 0;\n+\n+\tif (!open_pack_bitmap())\n+\t\treturn load_pack_bitmap();\n+\n+\treturn -1;\n+}\n+\n+struct include_data {\n+\tstruct bitmap *base;\n+\tstruct bitmap *seen;\n+};\n+\n+static inline int bitmap_position_extended(const unsigned char *sha1)\n+{\n+\tkhash_sha1_pos *positions = bitmap_git.ext_index.positions;\n+\tkhiter_t pos = kh_get_sha1_pos(positions, sha1);\n+\n+\tif (pos < kh_end(positions)) {\n+\t\tint bitmap_pos = kh_value(positions, pos);\n+\t\treturn bitmap_pos + bitmap_git.pack->num_objects;\n+\t}\n+\n+\treturn -1;\n+}\n+\n+static inline int bitmap_position_packfile(const unsigned char *sha1)\n+{\n+\toff_t offset = find_pack_entry_one(sha1, bitmap_git.pack);\n+\tif (!offset)\n+\t\treturn -1;\n+\n+\treturn find_revindex_position(bitmap_git.reverse_index, offset);\n+}\n+\n+static int bitmap_position(const unsigned char *sha1)\n+{\n+\tint pos = bitmap_position_packfile(sha1);\n+\treturn (pos >= 0) ? pos : bitmap_position_extended(sha1);\n+}\n+\n+static int ext_index_add_object(struct object *object, const char *name)\n+{\n+\tstruct eindex *eindex = &bitmap_git.ext_index;\n+\n+\tkhiter_t hash_pos;\n+\tint hash_ret;\n+\tint bitmap_pos;\n+\n+\thash_pos = kh_put_sha1_pos(eindex->positions, object->sha1, &hash_ret);\n+\tif (hash_ret > 0) {\n+\t\tif (eindex->count >= eindex->alloc) {\n+\t\t\teindex->alloc = (eindex->alloc + 16) * 3 / 2;\n+\t\t\teindex->objects = xrealloc(eindex->objects,\n+\t\t\t\teindex->alloc * sizeof(struct object *));\n+\t\t\teindex->hashes = xrealloc(eindex->hashes,\n+\t\t\t\teindex->alloc * sizeof(uint32_t));\n+\t\t}\n+\n+\t\tbitmap_pos = eindex->count;\n+\t\teindex->objects[eindex->count] = object;\n+\t\teindex->hashes[eindex->count] = pack_name_hash(name);\n+\t\tkh_value(eindex->positions, hash_pos) = bitmap_pos;\n+\t\teindex->count++;\n+\t} else {\n+\t\tbitmap_pos = kh_value(eindex->positions, hash_pos);\n+\t}\n+\n+\treturn bitmap_pos + bitmap_git.pack->num_objects;\n+}\n+\n+static void show_object(struct object *object, const struct name_path *path,\n+\t\t\tconst char *last, void *data)\n+{\n+\tstruct bitmap *base = data;\n+\tint bitmap_pos;\n+\n+\tbitmap_pos = bitmap_position(object->sha1);\n+\n+\tif (bitmap_pos < 0) {\n+\t\tchar *name = path_name(path, last);\n+\t\tbitmap_pos = ext_index_add_object(object, name);\n+\t\tfree(name);\n+\t}\n+\n+\tbitmap_set(base, bitmap_pos);\n+}\n+\n+static void show_commit(struct commit *commit, void *data)\n+{\n+}\n+\n+static int add_to_include_set(struct include_data *data,\n+\t\t\t      const unsigned char *sha1,\n+\t\t\t      int bitmap_pos)\n+{\n+\tkhiter_t hash_pos;\n+\n+\tif (data->seen && bitmap_get(data->seen, bitmap_pos))\n+\t\treturn 0;\n+\n+\tif (bitmap_get(data->base, bitmap_pos))\n+\t\treturn 0;\n+\n+\thash_pos = kh_get_sha1(bitmap_git.bitmaps, sha1);\n+\tif (hash_pos < kh_end(bitmap_git.bitmaps)) {\n+\t\tstruct stored_bitmap *st = kh_value(bitmap_git.bitmaps, hash_pos);\n+\t\tbitmap_or_ewah(data->base, lookup_stored_bitmap(st));\n+\t\treturn 0;\n+\t}\n+\n+\tbitmap_set(data->base, bitmap_pos);\n+\treturn 1;\n+}\n+\n+static int should_include(struct commit *commit, void *_data)\n+{\n+\tstruct include_data *data = _data;\n+\tint bitmap_pos;\n+\n+\tbitmap_pos = bitmap_position(commit->object.sha1);\n+\tif (bitmap_pos < 0)\n+\t\tbitmap_pos = ext_index_add_object((struct object *)commit, NULL);\n+\n+\tif (!add_to_include_set(data, commit->object.sha1, bitmap_pos)) {\n+\t\tstruct commit_list *parent = commit->parents;\n+\n+\t\twhile (parent) {\n+\t\t\tparent->item->object.flags |= SEEN;\n+\t\t\tparent = parent->next;\n+\t\t}\n+\n+\t\treturn 0;\n+\t}\n+\n+\treturn 1;\n+}\n+\n+static struct bitmap *find_objects(struct rev_info *revs,\n+\t\t\t\t   struct object_list *roots,\n+\t\t\t\t   struct bitmap *seen)\n+{\n+\tstruct bitmap *base = NULL;\n+\tint needs_walk = 0;\n+\n+\tstruct object_list *not_mapped = NULL;\n+\n+\t/*\n+\t * Go through all the roots for the walk. The ones that have bitmaps\n+\t * on the bitmap index will be `or`ed together to form an initial\n+\t * global reachability analysis.\n+\t *\n+\t * The ones without bitmaps in the index will be stored in the\n+\t * `not_mapped_list` for further processing.\n+\t */\n+\twhile (roots) {\n+\t\tstruct object *object = roots->item;\n+\t\troots = roots->next;\n+\n+\t\tif (object->type == OBJ_COMMIT) {\n+\t\t\tkhiter_t pos = kh_get_sha1(bitmap_git.bitmaps, object->sha1);\n+\n+\t\t\tif (pos < kh_end(bitmap_git.bitmaps)) {\n+\t\t\t\tstruct stored_bitmap *st = kh_value(bitmap_git.bitmaps, pos);\n+\t\t\t\tstruct ewah_bitmap *or_with = lookup_stored_bitmap(st);\n+\n+\t\t\t\tif (base == NULL)\n+\t\t\t\t\tbase = ewah_to_bitmap(or_with);\n+\t\t\t\telse\n+\t\t\t\t\tbitmap_or_ewah(base, or_with);\n+\n+\t\t\t\tobject->flags |= SEEN;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\n+\t\tobject_list_insert(object, &not_mapped);\n+\t}\n+\n+\t/*\n+\t * Best case scenario: We found bitmaps for all the roots,\n+\t * so the resulting `or` bitmap has the full reachability analysis\n+\t */\n+\tif (not_mapped == NULL)\n+\t\treturn base;\n+\n+\troots = not_mapped;\n+\n+\t/*\n+\t * Let's iterate through all the roots that don't have bitmaps to\n+\t * check if we can determine them to be reachable from the existing\n+\t * global bitmap.\n+\t *\n+\t * If we cannot find them in the existing global bitmap, we'll need\n+\t * to push them to an actual walk and run it until we can confirm\n+\t * they are reachable\n+\t */\n+\twhile (roots) {\n+\t\tstruct object *object = roots->item;\n+\t\tint pos;\n+\n+\t\troots = roots->next;\n+\t\tpos = bitmap_position(object->sha1);\n+\n+\t\tif (pos < 0 || base == NULL || !bitmap_get(base, pos)) {\n+\t\t\tobject->flags &= ~UNINTERESTING;\n+\t\t\tadd_pending_object(revs, object, \"\");\n+\t\t\tneeds_walk = 1;\n+\t\t} else {\n+\t\t\tobject->flags |= SEEN;\n+\t\t}\n+\t}\n+\n+\tif (needs_walk) {\n+\t\tstruct include_data incdata;\n+\n+\t\tif (base == NULL)\n+\t\t\tbase = bitmap_new();\n+\n+\t\tincdata.base = base;\n+\t\tincdata.seen = seen;\n+\n+\t\trevs->include_check = should_include;\n+\t\trevs->include_check_data = &incdata;\n+\n+\t\tif (prepare_revision_walk(revs))\n+\t\t\tdie(\"revision walk setup failed\");\n+\n+\t\ttraverse_commit_list(revs, show_commit, show_object, base);\n+\t}\n+\n+\treturn base;\n+}\n+\n+static void show_extended_objects(struct bitmap *objects,\n+\t\t\t\t  show_reachable_fn show_reach)\n+{\n+\tstruct eindex *eindex = &bitmap_git.ext_index;\n+\tuint32_t i;\n+\n+\tfor (i = 0; i < eindex->count; ++i) {\n+\t\tstruct object *obj;\n+\n+\t\tif (!bitmap_get(objects, bitmap_git.pack->num_objects + i))\n+\t\t\tcontinue;\n+\n+\t\tobj = eindex->objects[i];\n+\t\tshow_reach(obj->sha1, obj->type, 0, eindex->hashes[i], NULL, 0);\n+\t}\n+}\n+\n+static void show_objects_for_type(\n+\tstruct bitmap *objects,\n+\tstruct ewah_bitmap *type_filter,\n+\tenum object_type object_type,\n+\tshow_reachable_fn show_reach)\n+{\n+\tsize_t pos = 0, i = 0;\n+\tuint32_t offset;\n+\n+\tstruct ewah_iterator it;\n+\teword_t filter;\n+\n+\tif (bitmap_git.reuse_objects == bitmap_git.pack->num_objects)\n+\t\treturn;\n+\n+\tewah_iterator_init(&it, type_filter);\n+\n+\twhile (i < objects->word_alloc && ewah_iterator_next(&filter, &it)) {\n+\t\teword_t word = objects->words[i] & filter;\n+\n+\t\tfor (offset = 0; offset < BITS_IN_WORD; ++offset) {\n+\t\t\tconst unsigned char *sha1;\n+\t\t\tstruct revindex_entry *entry;\n+\t\t\tuint32_t hash = 0;\n+\n+\t\t\tif ((word >> offset) == 0)\n+\t\t\t\tbreak;\n+\n+\t\t\toffset += ewah_bit_ctz64(word >> offset);\n+\n+\t\t\tif (pos + offset < bitmap_git.reuse_objects)\n+\t\t\t\tcontinue;\n+\n+\t\t\tentry = &bitmap_git.reverse_index->revindex[pos + offset];\n+\t\t\tsha1 = nth_packed_object_sha1(bitmap_git.pack, entry->nr);\n+\n+\t\t\tshow_reach(sha1, object_type, 0, hash, bitmap_git.pack, entry->offset);\n+\t\t}\n+\n+\t\tpos += BITS_IN_WORD;\n+\t\ti++;\n+\t}\n+}\n+\n+static int in_bitmapped_pack(struct object_list *roots)\n+{\n+\twhile (roots) {\n+\t\tstruct object *object = roots->item;\n+\t\troots = roots->next;\n+\n+\t\tif (find_pack_entry_one(object->sha1, bitmap_git.pack) > 0)\n+\t\t\treturn 1;\n+\t}\n+\n+\treturn 0;\n+}\n+\n+int prepare_bitmap_walk(struct rev_info *revs)\n+{\n+\tunsigned int i;\n+\tunsigned int pending_nr = revs->pending.nr;\n+\tstruct object_array_entry *pending_e = revs->pending.objects;\n+\n+\tstruct object_list *wants = NULL;\n+\tstruct object_list *haves = NULL;\n+\n+\tstruct bitmap *wants_bitmap = NULL;\n+\tstruct bitmap *haves_bitmap = NULL;\n+\n+\tif (!bitmap_git.loaded) {\n+\t\t/* try to open a bitmapped pack, but don't parse it yet\n+\t\t * because we may not need to use it */\n+\t\tif (open_pack_bitmap() < 0)\n+\t\t\treturn -1;\n+\t}\n+\n+\tfor (i = 0; i < pending_nr; ++i) {\n+\t\tstruct object *object = pending_e[i].item;\n+\n+\t\tif (object->type == OBJ_NONE)\n+\t\t\tparse_object_or_die(object->sha1, NULL);\n+\n+\t\twhile (object->type == OBJ_TAG) {\n+\t\t\tstruct tag *tag = (struct tag *) object;\n+\n+\t\t\tif (object->flags & UNINTERESTING)\n+\t\t\t\tobject_list_insert(object, &haves);\n+\t\t\telse\n+\t\t\t\tobject_list_insert(object, &wants);\n+\n+\t\t\tif (!tag->tagged)\n+\t\t\t\tdie(\"bad tag\");\n+\t\t\tobject = parse_object_or_die(tag->tagged->sha1, NULL);\n+\t\t}\n+\n+\t\tif (object->flags & UNINTERESTING)\n+\t\t\tobject_list_insert(object, &haves);\n+\t\telse\n+\t\t\tobject_list_insert(object, &wants);\n+\t}\n+\n+\t/*\n+\t * if we have a HAVES list, but none of those haves is contained\n+\t * in the packfile that has a bitmap, we don't have anything to\n+\t * optimize here\n+\t */\n+\tif (haves && !in_bitmapped_pack(haves))\n+\t\treturn -1;\n+\n+\t/* if we don't want anything, we're done here */\n+\tif (!wants)\n+\t\treturn -1;\n+\n+\t/*\n+\t * now we're going to use bitmaps, so load the actual bitmap entries\n+\t * from disk. this is the point of no return; after this the rev_list\n+\t * becomes invalidated and we must perform the revwalk through bitmaps\n+\t */\n+\tif (!bitmap_git.loaded && load_pack_bitmap() < 0)\n+\t\treturn -1;\n+\n+\trevs->pending.nr = 0;\n+\trevs->pending.alloc = 0;\n+\trevs->pending.objects = NULL;\n+\n+\tif (haves) {\n+\t\thaves_bitmap = find_objects(revs, haves, NULL);\n+\t\treset_revision_walk();\n+\n+\t\tif (haves_bitmap == NULL)\n+\t\t\tdie(\"BUG: failed to perform bitmap walk\");\n+\t}\n+\n+\twants_bitmap = find_objects(revs, wants, haves_bitmap);\n+\n+\tif (!wants_bitmap)\n+\t\tdie(\"BUG: failed to perform bitmap walk\");\n+\n+\tif (haves_bitmap)\n+\t\tbitmap_and_not(wants_bitmap, haves_bitmap);\n+\n+\tbitmap_git.result = wants_bitmap;\n+\n+\tbitmap_free(haves_bitmap);\n+\treturn 0;\n+}\n+\n+int reuse_partial_packfile_from_bitmap(struct packed_git **packfile,\n+\t\t\t\t       uint32_t *entries,\n+\t\t\t\t       off_t *up_to)\n+{\n+\t/*\n+\t * Reuse the packfile content if we need more than\n+\t * 90% of its objects\n+\t */\n+\tstatic const double REUSE_PERCENT = 0.9;\n+\n+\tstruct bitmap *result = bitmap_git.result;\n+\tuint32_t reuse_threshold;\n+\tuint32_t i, reuse_objects = 0;\n+\n+\tassert(result);\n+\n+\tfor (i = 0; i < result->word_alloc; ++i) {\n+\t\tif (result->words[i] != (eword_t)~0) {\n+\t\t\treuse_objects += ewah_bit_ctz64(~result->words[i]);\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\treuse_objects += BITS_IN_WORD;\n+\t}\n+\n+#ifdef GIT_BITMAP_DEBUG\n+\t{\n+\t\tconst unsigned char *sha1;\n+\t\tstruct revindex_entry *entry;\n+\n+\t\tentry = &bitmap_git.reverse_index->revindex[reuse_objects];\n+\t\tsha1 = nth_packed_object_sha1(bitmap_git.pack, entry->nr);\n+\n+\t\tfprintf(stderr, \"Failed to reuse at %d (%016llx)\\n\",\n+\t\t\treuse_objects, result->words[i]);\n+\t\tfprintf(stderr, \" %s\\n\", sha1_to_hex(sha1));\n+\t}\n+#endif\n+\n+\tif (!reuse_objects)\n+\t\treturn -1;\n+\n+\tif (reuse_objects >= bitmap_git.pack->num_objects) {\n+\t\tbitmap_git.reuse_objects = *entries = bitmap_git.pack->num_objects;\n+\t\t*up_to = -1; /* reuse the full pack */\n+\t\t*packfile = bitmap_git.pack;\n+\t\treturn 0;\n+\t}\n+\n+\treuse_threshold = bitmap_popcount(bitmap_git.result) * REUSE_PERCENT;\n+\n+\tif (reuse_objects < reuse_threshold)\n+\t\treturn -1;\n+\n+\tbitmap_git.reuse_objects = *entries = reuse_objects;\n+\t*up_to = bitmap_git.reverse_index->revindex[reuse_objects].offset;\n+\t*packfile = bitmap_git.pack;\n+\n+\treturn 0;\n+}\n+\n+void traverse_bitmap_commit_list(show_reachable_fn show_reachable)\n+{\n+\tassert(bitmap_git.result);\n+\n+\tshow_objects_for_type(bitmap_git.result, bitmap_git.commits,\n+\t\tOBJ_COMMIT, show_reachable);\n+\tshow_objects_for_type(bitmap_git.result, bitmap_git.trees,\n+\t\tOBJ_TREE, show_reachable);\n+\tshow_objects_for_type(bitmap_git.result, bitmap_git.blobs,\n+\t\tOBJ_BLOB, show_reachable);\n+\tshow_objects_for_type(bitmap_git.result, bitmap_git.tags,\n+\t\tOBJ_TAG, show_reachable);\n+\n+\tshow_extended_objects(bitmap_git.result, show_reachable);\n+\n+\tbitmap_free(bitmap_git.result);\n+\tbitmap_git.result = NULL;\n+}\n+\n+static uint32_t count_object_type(struct bitmap *objects,\n+\t\t\t\t  enum object_type type)\n+{\n+\tstruct eindex *eindex = &bitmap_git.ext_index;\n+\n+\tuint32_t i = 0, count = 0;\n+\tstruct ewah_iterator it;\n+\teword_t filter;\n+\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\t\tewah_iterator_init(&it, bitmap_git.commits);\n+\t\tbreak;\n+\n+\tcase OBJ_TREE:\n+\t\tewah_iterator_init(&it, bitmap_git.trees);\n+\t\tbreak;\n+\n+\tcase OBJ_BLOB:\n+\t\tewah_iterator_init(&it, bitmap_git.blobs);\n+\t\tbreak;\n+\n+\tcase OBJ_TAG:\n+\t\tewah_iterator_init(&it, bitmap_git.tags);\n+\t\tbreak;\n+\n+\tdefault:\n+\t\treturn 0;\n+\t}\n+\n+\twhile (i < objects->word_alloc && ewah_iterator_next(&filter, &it)) {\n+\t\teword_t word = objects->words[i++] & filter;\n+\t\tcount += ewah_bit_popcount64(word);\n+\t}\n+\n+\tfor (i = 0; i < eindex->count; ++i) {\n+\t\tif (eindex->objects[i]->type == type &&\n+\t\t\tbitmap_get(objects, bitmap_git.pack->num_objects + i))\n+\t\t\tcount++;\n+\t}\n+\n+\treturn count;\n+}\n+\n+void count_bitmap_commit_list(uint32_t *commits, uint32_t *trees,\n+\t\t\t      uint32_t *blobs, uint32_t *tags)\n+{\n+\tassert(bitmap_git.result);\n+\n+\tif (commits)\n+\t\t*commits = count_object_type(bitmap_git.result, OBJ_COMMIT);\n+\n+\tif (trees)\n+\t\t*trees = count_object_type(bitmap_git.result, OBJ_TREE);\n+\n+\tif (blobs)\n+\t\t*blobs = count_object_type(bitmap_git.result, OBJ_BLOB);\n+\n+\tif (tags)\n+\t\t*tags = count_object_type(bitmap_git.result, OBJ_TAG);\n+}\n+\n+struct bitmap_test_data {\n+\tstruct bitmap *base;\n+\tstruct progress *prg;\n+\tsize_t seen;\n+};\n+\n+static void test_show_object(struct object *object,\n+\t\t\t     const struct name_path *path,\n+\t\t\t     const char *last, void *data)\n+{\n+\tstruct bitmap_test_data *tdata = data;\n+\tint bitmap_pos;\n+\n+\tbitmap_pos = bitmap_position(object->sha1);\n+\tif (bitmap_pos < 0)\n+\t\tdie(\"Object not in bitmap: %s\\n\", sha1_to_hex(object->sha1));\n+\n+\tbitmap_set(tdata->base, bitmap_pos);\n+\tdisplay_progress(tdata->prg, ++tdata->seen);\n+}\n+\n+static void test_show_commit(struct commit *commit, void *data)\n+{\n+\tstruct bitmap_test_data *tdata = data;\n+\tint bitmap_pos;\n+\n+\tbitmap_pos = bitmap_position(commit->object.sha1);\n+\tif (bitmap_pos < 0)\n+\t\tdie(\"Object not in bitmap: %s\\n\", sha1_to_hex(commit->object.sha1));\n+\n+\tbitmap_set(tdata->base, bitmap_pos);\n+\tdisplay_progress(tdata->prg, ++tdata->seen);\n+}\n+\n+void test_bitmap_walk(struct rev_info *revs)\n+{\n+\tstruct object *root;\n+\tstruct bitmap *result = NULL;\n+\tkhiter_t pos;\n+\tsize_t result_popcnt;\n+\tstruct bitmap_test_data tdata;\n+\n+\tif (prepare_bitmap_git())\n+\t\tdie(\"failed to load bitmap indexes\");\n+\n+\tif (revs->pending.nr != 1)\n+\t\tdie(\"you must specify exactly one commit to test\");\n+\n+\tfprintf(stderr, \"Bitmap v%d test (%d entries loaded)\\n\",\n+\t\tbitmap_git.version, bitmap_git.entry_count);\n+\n+\troot = revs->pending.objects[0].item;\n+\tpos = kh_get_sha1(bitmap_git.bitmaps, root->sha1);\n+\n+\tif (pos < kh_end(bitmap_git.bitmaps)) {\n+\t\tstruct stored_bitmap *st = kh_value(bitmap_git.bitmaps, pos);\n+\t\tstruct ewah_bitmap *bm = lookup_stored_bitmap(st);\n+\n+\t\tfprintf(stderr, \"Found bitmap for %s. %d bits / %08x checksum\\n\",\n+\t\t\tsha1_to_hex(root->sha1), (int)bm->bit_size, ewah_checksum(bm));\n+\n+\t\tresult = ewah_to_bitmap(bm);\n+\t}\n+\n+\tif (result == NULL)\n+\t\tdie(\"Commit %s doesn't have an indexed bitmap\", sha1_to_hex(root->sha1));\n+\n+\trevs->tag_objects = 1;\n+\trevs->tree_objects = 1;\n+\trevs->blob_objects = 1;\n+\n+\tresult_popcnt = bitmap_popcount(result);\n+\n+\tif (prepare_revision_walk(revs))\n+\t\tdie(\"revision walk setup failed\");\n+\n+\ttdata.base = bitmap_new();\n+\ttdata.prg = start_progress(\"Verifying bitmap entries\", result_popcnt);\n+\ttdata.seen = 0;\n+\n+\ttraverse_commit_list(revs, &test_show_commit, &test_show_object, &tdata);\n+\n+\tstop_progress(&tdata.prg);\n+\n+\tif (bitmap_equals(result, tdata.base))\n+\t\tfprintf(stderr, \"OK!\\n\");\n+\telse\n+\t\tfprintf(stderr, \"Mismatch!\\n\");\n+}\ndiff --git a/pack-bitmap.h b/pack-bitmap.h\nnew file mode 100644\nindex 0000000..9351aa5\n--- /dev/null\n+++ b/pack-bitmap.h\n@@ -0,0 +1,43 @@\n+#ifndef PACK_BITMAP_H\n+#define PACK_BITMAP_H\n+\n+#include \"ewah/ewok.h\"\n+#include \"khash.h\"\n+\n+struct bitmap_disk_entry {\n+\tuint32_t object_pos;\n+\tuint8_t xor_offset;\n+\tuint8_t flags;\n+} __attribute__((packed));\n+\n+struct bitmap_disk_header {\n+\tchar magic[4];\n+\tuint16_t version;\n+\tuint16_t options;\n+\tuint32_t entry_count;\n+\tunsigned char checksum[20];\n+};\n+\n+static const char BITMAP_IDX_SIGNATURE[] = {'B', 'I', 'T', 'M'};;\n+\n+enum pack_bitmap_opts {\n+\tBITMAP_OPT_FULL_DAG = 1,\n+};\n+\n+typedef int (*show_reachable_fn)(\n+\tconst unsigned char *sha1,\n+\tenum object_type type,\n+\tint flags,\n+\tuint32_t hash,\n+\tstruct packed_git *found_pack,\n+\toff_t found_offset);\n+\n+int prepare_bitmap_git(void);\n+void count_bitmap_commit_list(uint32_t *commits, uint32_t *trees, uint32_t *blobs, uint32_t *tags);\n+void traverse_bitmap_commit_list(show_reachable_fn show_reachable);\n+void test_bitmap_walk(struct rev_info *revs);\n+char *pack_bitmap_filename(struct packed_git *p);\n+int prepare_bitmap_walk(struct rev_info *revs);\n+int reuse_partial_packfile_from_bitmap(struct packed_git **packfile, uint32_t *entries, off_t *up_to);\n+\n+#endif\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230609","messageId":"20131114124510.GK10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 11/21] pack-objects: use bitmaps when packing objects","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:45:10Z","receivedAt":"2013-11-14T12:45:10Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nIn this patch, we use the bitmap API to perform the `Counting Objects`\nphase in pack-objects, rather than a traditional walk through the object\ngraph. For a reasonably-packed large repo, the time to fetch and clone\nis often dominated by the full-object revision walk during the Counting\nObjects phase. Using bitmaps can reduce the CPU time required on the\nserver (and therefore start sending the actual pack data with less\ndelay).\n\nFor bitmaps to be used, the following must be true:\n\n  1. We must be packing to stdout (as a normal `pack-objects` from\n     `upload-pack` would do).\n\n  2. There must be a .bitmap index containing at least one of the\n     \"have\" objects that the client is asking for.\n\n  3. Bitmaps must be enabled (they are enabled by default, but can be\n     disabled by setting `pack.usebitmaps` to false, or by using\n     `--no-use-bitmap-index` on the command-line).\n\nIf any of these is not true, we fall back to doing a normal walk of the\nobject graph.\n\nHere are some sample timings from a full pack of `torvalds/linux` (i.e.\nsomething very similar to what would be generated for a clone of the\nrepository) that show the speedup produced by various\nmethods:\n\n    [existing graph traversal]\n    $ time git pack-objects --all --stdout --no-use-bitmap-index \\\n\t\t\t    </dev/null >/dev/null\n    Counting objects: 3237103, done.\n    Compressing objects: 100% (508752/508752), done.\n    Total 3237103 (delta 2699584), reused 3237103 (delta 2699584)\n\n    real    0m44.111s\n    user    0m42.396s\n    sys     0m3.544s\n\n    [bitmaps only, without partial pack reuse; note that\n     pack reuse is automatic, so timing this required a\n     patch to disable it]\n    $ time git pack-objects --all --stdout </dev/null >/dev/null\n    Counting objects: 3237103, done.\n    Compressing objects: 100% (508752/508752), done.\n    Total 3237103 (delta 2699584), reused 3237103 (delta 2699584)\n\n    real    0m5.413s\n    user    0m5.604s\n    sys     0m1.804s\n\n    [bitmaps with pack reuse (what you get with this patch)]\n    $ time git pack-objects --all --stdout </dev/null >/dev/null\n    Reusing existing pack: 3237103, done.\n    Total 3237103 (delta 0), reused 0 (delta 0)\n\n    real    0m1.636s\n    user    0m1.460s\n    sys     0m0.172s\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Documentation/config.txt |   6 ++\n builtin/pack-objects.c   | 169 ++++++++++++++++++++++++++++++++++++++++-------\n 2 files changed, 150 insertions(+), 25 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex ab26963..a981369 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -1858,6 +1858,12 @@ pack.packSizeLimit::\n \tCommon unit suffixes of 'k', 'm', or 'g' are\n \tsupported.\n \n+pack.useBitmaps::\n+\tWhen true, git will use pack bitmaps (if available) when packing\n+\tto stdout (e.g., during the server side of a fetch). Defaults to\n+\ttrue. You should not generally need to turn this off unless\n+\tyou are debugging pack bitmaps.\n+\n pager.<cmd>::\n \tIf the value is boolean, turns on or off pagination of the\n \toutput of a particular Git subcommand when writing to a tty.\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex faf746b..f04cce5 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -19,6 +19,7 @@\n #include \"refs.h\"\n #include \"streaming.h\"\n #include \"thread-utils.h\"\n+#include \"pack-bitmap.h\"\n \n static const char *pack_usage[] = {\n \tN_(\"git pack-objects --stdout [options...] [< ref-list | < object-list]\"),\n@@ -57,12 +58,23 @@ static struct progress *progress_state;\n static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n static int pack_compression_seen;\n \n+static struct packed_git *reuse_packfile;\n+static uint32_t reuse_packfile_objects;\n+static off_t reuse_packfile_offset;\n+\n+static int use_bitmap_index = 1;\n+\n static unsigned long delta_cache_size = 0;\n static unsigned long max_delta_cache_size = 256 * 1024 * 1024;\n static unsigned long cache_max_small_delta_size = 1000;\n \n static unsigned long window_memory_limit = 0;\n \n+enum {\n+\tOBJECT_ENTRY_EXCLUDE = (1 << 0),\n+\tOBJECT_ENTRY_NO_TRY_DELTA = (1 << 1)\n+};\n+\n /*\n  * stats\n  */\n@@ -678,6 +690,49 @@ static struct object_entry **compute_write_order(void)\n \treturn wo;\n }\n \n+static off_t write_reused_pack(struct sha1file *f)\n+{\n+\tuint8_t buffer[8192];\n+\toff_t to_write;\n+\tint fd;\n+\n+\tif (!is_pack_valid(reuse_packfile))\n+\t\treturn 0;\n+\n+\tfd = git_open_noatime(reuse_packfile->pack_name);\n+\tif (fd < 0)\n+\t\treturn 0;\n+\n+\tif (lseek(fd, sizeof(struct pack_header), SEEK_SET) == -1) {\n+\t\tclose(fd);\n+\t\treturn 0;\n+\t}\n+\n+\tif (reuse_packfile_offset < 0)\n+\t\treuse_packfile_offset = reuse_packfile->pack_size - 20;\n+\n+\tto_write = reuse_packfile_offset - sizeof(struct pack_header);\n+\n+\twhile (to_write) {\n+\t\tint read_pack = xread(fd, buffer, sizeof(buffer));\n+\n+\t\tif (read_pack <= 0) {\n+\t\t\tclose(fd);\n+\t\t\treturn 0;\n+\t\t}\n+\n+\t\tif (read_pack > to_write)\n+\t\t\tread_pack = to_write;\n+\n+\t\tsha1write(f, buffer, read_pack);\n+\t\tto_write -= read_pack;\n+\t}\n+\n+\tclose(fd);\n+\twritten += reuse_packfile_objects;\n+\treturn reuse_packfile_offset - sizeof(struct pack_header);\n+}\n+\n static void write_pack_file(void)\n {\n \tuint32_t i = 0, j;\n@@ -704,6 +759,18 @@ static void write_pack_file(void)\n \t\toffset = write_pack_header(f, nr_remaining);\n \t\tif (!offset)\n \t\t\tdie_errno(\"unable to write pack header\");\n+\n+\t\tif (reuse_packfile) {\n+\t\t\toff_t packfile_size;\n+\t\t\tassert(pack_to_stdout);\n+\n+\t\t\tpackfile_size = write_reused_pack(f);\n+\t\t\tif (!packfile_size)\n+\t\t\t\tdie_errno(\"failed to re-use existing pack\");\n+\n+\t\t\toffset += packfile_size;\n+\t\t}\n+\n \t\tnr_written = 0;\n \t\tfor (; i < to_pack.nr_objects; i++) {\n \t\t\tstruct object_entry *e = write_order[i];\n@@ -800,14 +867,14 @@ static int no_try_delta(const char *path)\n \treturn 0;\n }\n \n-static int add_object_entry(const unsigned char *sha1, enum object_type type,\n-\t\t\t    const char *name, int exclude)\n+static int add_object_entry_1(const unsigned char *sha1, enum object_type type,\n+\t\t\t      int flags, uint32_t name_hash,\n+\t\t\t      struct packed_git *found_pack, off_t found_offset)\n {\n \tstruct object_entry *entry;\n-\tstruct packed_git *p, *found_pack = NULL;\n-\toff_t found_offset = 0;\n-\tuint32_t hash = pack_name_hash(name);\n+\tstruct packed_git *p;\n \tuint32_t index_pos;\n+\tint exclude = (flags & OBJECT_ENTRY_EXCLUDE);\n \n \tentry = packlist_find(&to_pack, sha1, &index_pos);\n \tif (entry) {\n@@ -822,36 +889,42 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n \tif (!exclude && local && has_loose_object_nonlocal(sha1))\n \t\treturn 0;\n \n-\tfor (p = packed_git; p; p = p->next) {\n-\t\toff_t offset = find_pack_entry_one(sha1, p);\n-\t\tif (offset) {\n-\t\t\tif (!found_pack) {\n-\t\t\t\tif (!is_pack_valid(p)) {\n-\t\t\t\t\twarning(\"packfile %s cannot be accessed\", p->pack_name);\n-\t\t\t\t\tcontinue;\n+\tif (!found_pack) {\n+\t\tfor (p = packed_git; p; p = p->next) {\n+\t\t\toff_t offset = find_pack_entry_one(sha1, p);\n+\t\t\tif (offset) {\n+\t\t\t\tif (!found_pack) {\n+\t\t\t\t\tif (!is_pack_valid(p)) {\n+\t\t\t\t\t\twarning(\"packfile %s cannot be accessed\", p->pack_name);\n+\t\t\t\t\t\tcontinue;\n+\t\t\t\t\t}\n+\t\t\t\t\tfound_offset = offset;\n+\t\t\t\t\tfound_pack = p;\n \t\t\t\t}\n-\t\t\t\tfound_offset = offset;\n-\t\t\t\tfound_pack = p;\n+\t\t\t\tif (exclude)\n+\t\t\t\t\tbreak;\n+\t\t\t\tif (incremental)\n+\t\t\t\t\treturn 0;\n+\t\t\t\tif (local && !p->pack_local)\n+\t\t\t\t\treturn 0;\n+\t\t\t\tif (ignore_packed_keep && p->pack_local && p->pack_keep)\n+\t\t\t\t\treturn 0;\n \t\t\t}\n-\t\t\tif (exclude)\n-\t\t\t\tbreak;\n-\t\t\tif (incremental)\n-\t\t\t\treturn 0;\n-\t\t\tif (local && !p->pack_local)\n-\t\t\t\treturn 0;\n-\t\t\tif (ignore_packed_keep && p->pack_local && p->pack_keep)\n-\t\t\t\treturn 0;\n \t\t}\n \t}\n \n \tentry = packlist_alloc(&to_pack, sha1, index_pos);\n-\tentry->hash = hash;\n+\tentry->hash = name_hash;\n \tif (type)\n \t\tentry->type = type;\n \tif (exclude)\n \t\tentry->preferred_base = 1;\n \telse\n \t\tnr_result++;\n+\n+\tif (flags & OBJECT_ENTRY_NO_TRY_DELTA)\n+\t\tentry->no_try_delta = 1;\n+\n \tif (found_pack) {\n \t\tentry->in_pack = found_pack;\n \t\tentry->in_pack_offset = found_offset;\n@@ -859,10 +932,21 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n \n \tdisplay_progress(progress_state, to_pack.nr_objects);\n \n+\treturn 1;\n+}\n+\n+static int add_object_entry(const unsigned char *sha1, enum object_type type,\n+\t\t\t    const char *name, int exclude)\n+{\n+\tint flags = 0;\n+\n+\tif (exclude)\n+\t\tflags |= OBJECT_ENTRY_EXCLUDE;\n+\n \tif (name && no_try_delta(name))\n-\t\tentry->no_try_delta = 1;\n+\t\tflags |= OBJECT_ENTRY_NO_TRY_DELTA;\n \n-\treturn 1;\n+\treturn add_object_entry_1(sha1, type, flags, pack_name_hash(name), NULL, 0);\n }\n \n struct pbase_tree_cache {\n@@ -2027,6 +2111,10 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n \t\tcache_max_small_delta_size = git_config_int(k, v);\n \t\treturn 0;\n \t}\n+\tif (!strcmp(k, \"pack.usebitmaps\")) {\n+\t\tuse_bitmap_index = git_config_bool(k, v);\n+\t\treturn 0;\n+\t}\n \tif (!strcmp(k, \"pack.threads\")) {\n \t\tdelta_search_threads = git_config_int(k, v);\n \t\tif (delta_search_threads < 0)\n@@ -2235,6 +2323,29 @@ static void loosen_unused_packed_objects(struct rev_info *revs)\n \t}\n }\n \n+static int get_object_list_from_bitmap(struct rev_info *revs)\n+{\n+\tif (prepare_bitmap_walk(revs) < 0)\n+\t\treturn -1;\n+\n+\tif (!reuse_partial_packfile_from_bitmap(\n+\t\t\t&reuse_packfile,\n+\t\t\t&reuse_packfile_objects,\n+\t\t\t&reuse_packfile_offset)) {\n+\t\tassert(reuse_packfile_objects);\n+\t\tnr_result += reuse_packfile_objects;\n+\n+\t\tif (progress) {\n+\t\t\tfprintf(stderr, \"Reusing existing pack: %d, done.\\n\",\n+\t\t\t\treuse_packfile_objects);\n+\t\t\tfflush(stderr);\n+\t\t}\n+\t}\n+\n+\ttraverse_bitmap_commit_list(&add_object_entry_1);\n+\treturn 0;\n+}\n+\n static void get_object_list(int ac, const char **av)\n {\n \tstruct rev_info revs;\n@@ -2262,6 +2373,9 @@ static void get_object_list(int ac, const char **av)\n \t\t\tdie(\"bad revision '%s'\", line);\n \t}\n \n+\tif (use_bitmap_index && !get_object_list_from_bitmap(&revs))\n+\t\treturn;\n+\n \tif (prepare_revision_walk(&revs))\n \t\tdie(\"revision walk setup failed\");\n \tmark_edges_uninteresting(&revs, show_edge);\n@@ -2391,6 +2505,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t    N_(\"pack compression level\")),\n \t\tOPT_SET_INT(0, \"keep-true-parents\", &grafts_replace_parents,\n \t\t\t    N_(\"do not hide commits by grafts\"), 0),\n+\t\tOPT_BOOL(0, \"use-bitmap-index\", &use_bitmap_index,\n+\t\t\t N_(\"use a bitmap index if available to speed up counting objects\")),\n \t\tOPT_END(),\n \t};\n \n@@ -2457,6 +2573,9 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (keep_unreachable && unpack_unreachable)\n \t\tdie(\"--keep-unreachable and --unpack-unreachable are incompatible.\");\n \n+\tif (!use_internal_rev_list || !pack_to_stdout || is_repository_shallow())\n+\t\tuse_bitmap_index = 0;\n+\n \tif (progress && all_progress_implied)\n \t\tprogress = 2;\n \n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230610","messageId":"20131114124523.GL10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 12/21] rev-list: add bitmap mode to speed up object lists","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:45:23Z","receivedAt":"2013-11-14T12:45:23Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nThe bitmap reachability index used to speed up the counting objects\nphase during `pack-objects` can also be used to optimize a normal\nrev-list if the only thing required are the SHA1s of the objects during\nthe list (i.e., not the path names at which trees and blobs were found).\n\nCalling `git rev-list --objects --use-bitmap-index [committish]` will\nperform an object iteration based on a bitmap result instead of actually\nwalking the object graph.\n\nThese are some example timings for `torvalds/linux` (warm cache,\nbest-of-five):\n\n    $ time git rev-list --objects master > /dev/null\n\n    real    0m34.191s\n    user    0m33.904s\n    sys     0m0.268s\n\n    $ time git rev-list --objects --use-bitmap-index master > /dev/null\n\n    real    0m1.041s\n    user    0m0.976s\n    sys     0m0.064s\n\nLikewise, using `git rev-list --count --use-bitmap-index` will speed up\nthe counting operation by building the resulting bitmap and performing a\nfast popcount (number of bits set on the bitmap) on the result.\n\nHere are some sample timings of different ways to count commits in\n`torvalds/linux`:\n\n    $ time git rev-list master | wc -l\n        399882\n\n        real    0m6.524s\n        user    0m6.060s\n        sys     0m3.284s\n\n    $ time git rev-list --count master\n        399882\n\n        real    0m4.318s\n        user    0m4.236s\n        sys     0m0.076s\n\n    $ time git rev-list --use-bitmap-index --count master\n        399882\n\n        real    0m0.217s\n        user    0m0.176s\n        sys     0m0.040s\n\nThis also respects negative refs, so you can use it to count\na slice of history:\n\n        $ time git rev-list --count v3.0..master\n        144843\n\n        real    0m1.971s\n        user    0m1.932s\n        sys     0m0.036s\n\n        $ time git rev-list --use-bitmap-index --count v3.0..master\n        real    0m0.280s\n        user    0m0.220s\n        sys     0m0.056s\n\nThough note that the closer the endpoints, the less it helps. In the\ntraversal case, we have fewer commits to cross, so we take less time.\nBut the bitmap time is dominated by generating the pack revindex, which\nis constant with respect to the refs given.\n\nNote that you cannot yet get a fast --left-right count of a symmetric\ndifference (e.g., \"--count --left-right master...topic\"). The slow part\nof that walk actually happens during the merge-base determination when\nwe parse \"master...topic\". Even though a count does not actually need to\nknow the real merge base (it only needs to take the symmetric difference\nof the bitmaps), the revision code would require some refactoring to\nhandle this case.\n\nAdditionally, a `--test-bitmap` flag has been added that will perform\nthe same rev-list manually (i.e. using a normal revwalk) and using\nbitmaps, and verify that the results are the same. This can be used to\nexercise the bitmap code, and also to verify that the contents of the\n.bitmap file are sane.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Documentation/git-rev-list.txt     |  1 +\n Documentation/rev-list-options.txt |  8 ++++++++\n builtin/rev-list.c                 | 39 ++++++++++++++++++++++++++++++++++++++\n 3 files changed, 48 insertions(+)\n\ndiff --git a/Documentation/git-rev-list.txt b/Documentation/git-rev-list.txt\nindex 045b37b..7a1585d 100644\n--- a/Documentation/git-rev-list.txt\n+++ b/Documentation/git-rev-list.txt\n@@ -55,6 +55,7 @@ SYNOPSIS\n \t     [ \\--reverse ]\n \t     [ \\--walk-reflogs ]\n \t     [ \\--no-walk ] [ \\--do-walk ]\n+\t     [ \\--use-bitmap-index ]\n \t     <commit>... [ \\-- <paths>... ]\n \n DESCRIPTION\ndiff --git a/Documentation/rev-list-options.txt b/Documentation/rev-list-options.txt\nindex ec86d09..f7c8a4d 100644\n--- a/Documentation/rev-list-options.txt\n+++ b/Documentation/rev-list-options.txt\n@@ -274,6 +274,14 @@ See also linkgit:git-reflog[1].\n \tOutput excluded boundary commits. Boundary commits are\n \tprefixed with `-`.\n \n+ifdef::git-rev-list[]\n+--use-bitmap-index::\n+\n+\tTry to speed up the traversal using the pack bitmap index (if\n+\tone is available). Note that when traversing with `--objects`,\n+\ttrees and blobs will not have their associated path printed.\n+endif::git-rev-list[]\n+\n --\n \n History Simplification\ndiff --git a/builtin/rev-list.c b/builtin/rev-list.c\nindex 0745e2d..9f92905 100644\n--- a/builtin/rev-list.c\n+++ b/builtin/rev-list.c\n@@ -3,6 +3,8 @@\n #include \"diff.h\"\n #include \"revision.h\"\n #include \"list-objects.h\"\n+#include \"pack.h\"\n+#include \"pack-bitmap.h\"\n #include \"builtin.h\"\n #include \"log-tree.h\"\n #include \"graph.h\"\n@@ -257,6 +259,18 @@ static int show_bisect_vars(struct rev_list_info *info, int reaches, int all)\n \treturn 0;\n }\n \n+static int show_object_fast(\n+\tconst unsigned char *sha1,\n+\tenum object_type type,\n+\tint exclude,\n+\tuint32_t name_hash,\n+\tstruct packed_git *found_pack,\n+\toff_t found_offset)\n+{\n+\tfprintf(stdout, \"%s\\n\", sha1_to_hex(sha1));\n+\treturn 1;\n+}\n+\n int cmd_rev_list(int argc, const char **argv, const char *prefix)\n {\n \tstruct rev_info revs;\n@@ -265,6 +279,7 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \tint bisect_list = 0;\n \tint bisect_show_vars = 0;\n \tint bisect_find_all = 0;\n+\tint use_bitmap_index = 0;\n \n \tgit_config(git_default_config, NULL);\n \tinit_revisions(&revs, prefix);\n@@ -306,6 +321,14 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \t\t\tbisect_show_vars = 1;\n \t\t\tcontinue;\n \t\t}\n+\t\tif (!strcmp(arg, \"--use-bitmap-index\")) {\n+\t\t\tuse_bitmap_index = 1;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (!strcmp(arg, \"--test-bitmap\")) {\n+\t\t\ttest_bitmap_walk(&revs);\n+\t\t\treturn 0;\n+\t\t}\n \t\tusage(rev_list_usage);\n \n \t}\n@@ -333,6 +356,22 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \tif (bisect_list)\n \t\trevs.limited = 1;\n \n+\tif (use_bitmap_index) {\n+\t\tif (revs.count && !revs.left_right && !revs.cherry_mark) {\n+\t\t\tuint32_t commit_count;\n+\t\t\tif (!prepare_bitmap_walk(&revs)) {\n+\t\t\t\tcount_bitmap_commit_list(&commit_count, NULL, NULL, NULL);\n+\t\t\t\tprintf(\"%d\\n\", commit_count);\n+\t\t\t\treturn 0;\n+\t\t\t}\n+\t\t} else if (revs.tag_objects && revs.tree_objects && revs.blob_objects) {\n+\t\t\tif (!prepare_bitmap_walk(&revs)) {\n+\t\t\t\ttraverse_bitmap_commit_list(&show_object_fast);\n+\t\t\t\treturn 0;\n+\t\t\t}\n+\t\t}\n+\t}\n+\n \tif (prepare_revision_walk(&revs))\n \t\tdie(\"revision walk setup failed\");\n \tif (revs.tree_objects)\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230611","messageId":"20131114124544.GM10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 13/21] pack-objects: implement bitmap writing","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:45:45Z","receivedAt":"2013-11-14T12:45:45Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nThis commit extends more the functionality of `pack-objects` by allowing\nit to write out a `.bitmap` index next to any written packs, together\nwith the `.idx` index that currently gets written.\n\nIf bitmap writing is enabled for a given repository (either by calling\n`pack-objects` with the `--write-bitmap-index` flag or by having\n`pack.writebitmaps` set to `true` in the config) and pack-objects is\nwriting a packfile that would normally be indexed (i.e. not piping to\nstdout), we will attempt to write the corresponding bitmap index for the\npackfile.\n\nBitmap index writing happens after the packfile and its index has been\nsuccessfully written to disk (`finish_tmp_packfile`). The process is\nperformed in several steps:\n\n    1. `bitmap_writer_set_checksum`: this call stores the partial\n       checksum for the packfile being written; the checksum will be\n       written in the resulting bitmap index to verify its integrity\n\n    2. `bitmap_writer_build_type_index`: this call uses the array of\n       `struct object_entry` that has just been sorted when writing out\n       the actual packfile index to disk to generate 4 type-index bitmaps\n       (one for each object type).\n\n       These bitmaps have their nth bit set if the given object is of\n       the bitmap's type. E.g. the nth bit of the Commits bitmap will be\n       1 if the nth object in the packfile index is a commit.\n\n       This is a very cheap operation because the bitmap writing code has\n       access to the metadata stored in the `struct object_entry` array,\n       and hence the real type for each object in the packfile.\n\n    3. `bitmap_writer_reuse_bitmaps`: if there exists an existing bitmap\n       index for one of the packfiles we're trying to repack, this call\n       will efficiently rebuild the existing bitmaps so they can be\n       reused on the new index. All the existing bitmaps will be stored\n       in a `reuse` hash table, and the commit selection phase will\n       prioritize these when selecting, as they can be written directly\n       to the new index without having to perform a revision walk to\n       fill the bitmap. This can greatly speed up the repack of a\n       repository that already has bitmaps.\n\n    4. `bitmap_writer_select_commits`: if bitmap writing is enabled for\n       a given `pack-objects` run, the sequence of commits generated\n       during the Counting Objects phase will be stored in an array.\n\n       We then use that array to build up the list of selected commits.\n       Writing a bitmap in the index for each object in the repository\n       would be cost-prohibitive, so we use a simple heuristic to pick\n       the commits that will be indexed with bitmaps.\n\n       The current heuristics are a simplified version of JGit's\n       original implementation. We select a higher density of commits\n       depending on their age: the 100 most recent commits are always\n       selected, after that we pick 1 commit of each 100, and the gap\n       increases as the commits grow older. On top of that, we make sure\n       that every single branch that has not been merged (all the tips\n       that would be required from a clone) gets their own bitmap, and\n       when selecting commits between a gap, we tend to prioritize the\n       commit with the most parents.\n\n       Do note that there is no right/wrong way to perform commit\n       selection; different selection algorithms will result in\n       different commits being selected, but there's no such thing as\n       \"missing a commit\". The bitmap walker algorithm implemented in\n       `prepare_bitmap_walk` is able to adapt to missing bitmaps by\n       performing manual walks that complete the bitmap: the ideal\n       selection algorithm, however, would select the commits that are\n       more likely to be used as roots for a walk in the future (e.g.\n       the tips of each branch, and so on) to ensure a bitmap for them\n       is always available.\n\n    5. `bitmap_writer_build`: this is the computationally expensive part\n       of bitmap generation. Based on the list of commits that were\n       selected in the previous step, we perform several incremental\n       walks to generate the bitmap for each commit.\n\n       The walks begin from the oldest commit, and are built up\n       incrementally for each branch. E.g. consider this dag where A, B,\n       C, D, E, F are the selected commits, and a, b, c, e are a chunk\n       of simplified history that will not receive bitmaps.\n\n            A---a---B--b--C--c--D\n                     \\\n                      E--e--F\n\n       We start by building the bitmap for A, using A as the root for a\n       revision walk and marking all the objects that are reachable\n       until the walk is over. Once this bitmap is stored, we reuse the\n       bitmap walker to perform the walk for B, assuming that once we\n       reach A again, the walk will be terminated because A has already\n       been SEEN on the previous walk.\n\n       This process is repeated for C, and D, but when we try to\n       generate the bitmaps for E, we can reuse neither the current walk\n       nor the bitmap we have generated so far.\n\n       What we do now is resetting both the walk and clearing the\n       bitmap, and performing the walk from scratch using E as the\n       origin. This new walk, however, does not need to be completed.\n       Once we hit B, we can lookup the bitmap we have already stored\n       for that commit and OR it with the existing bitmap we've composed\n       so far, allowing us to limit the walk early.\n\n       After all the bitmaps have been generated, another iteration\n       through the list of commits is performed to find the best XOR\n       offsets for compression before writing them to disk. Because of\n       the incremental nature of these bitmaps, XORing one of them with\n       its predecesor results in a minimal \"bitmap delta\" most of the\n       time. We can write this delta to the on-disk bitmap index, and\n       then re-compose the original bitmaps by XORing them again when\n       loaded.\n\n       This is a phase very similar to pack-object's `find_delta` (using\n       bitmaps instead of objects, of course), except the heuristics\n       have been greatly simplified: we only check the 10 bitmaps before\n       any given one to find best compressing one. This gives good\n       results in practice, because there is locality in the ordering of\n       the objects (and therefore bitmaps) in the packfile.\n\n     6. `bitmap_writer_finish`: the last step in the process is\n\tserializing to disk all the bitmap data that has been generated\n\tin the two previous steps.\n\n\tThe bitmap is written to a tmp file and then moved atomically to\n\tits final destination, using the same process as\n\t`pack-write.c:write_idx_file`.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Documentation/config.txt |   8 +\n Makefile                 |   1 +\n builtin/pack-objects.c   |  53 +++++\n pack-bitmap-write.c      | 535 +++++++++++++++++++++++++++++++++++++++++++++++\n pack-bitmap.c            |  92 ++++++++\n pack-bitmap.h            |  19 ++\n pack-objects.h           |   1 +\n pack-write.c             |   2 +\n 8 files changed, 711 insertions(+)\n create mode 100644 pack-bitmap-write.c\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex a981369..a439a4c 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -1864,6 +1864,14 @@ pack.useBitmaps::\n \ttrue. You should not generally need to turn this off unless\n \tyou are debugging pack bitmaps.\n \n+pack.writebitmaps::\n+\tWhen true, git will write a bitmap index when packing all\n+\tobjects to disk (e.g., as when `git repack -a` is run).  This\n+\tindex can speed up the \"counting objects\" phase of subsequent\n+\tpacks created for clones and fetches, at the cost of some disk\n+\tspace and extra time spent on the initial repack.  Defaults to\n+\tfalse.\n+\n pager.<cmd>::\n \tIf the value is boolean, turns on or off pagination of the\n \toutput of a particular Git subcommand when writing to a tty.\ndiff --git a/Makefile b/Makefile\nindex b983d78..555d44c 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -839,6 +839,7 @@ LIB_OBJS += notes-merge.o\n LIB_OBJS += notes-utils.o\n LIB_OBJS += object.o\n LIB_OBJS += pack-bitmap.o\n+LIB_OBJS += pack-bitmap-write.o\n LIB_OBJS += pack-check.o\n LIB_OBJS += pack-objects.o\n LIB_OBJS += pack-revindex.o\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex f04cce5..26646e7 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -63,6 +63,7 @@ static uint32_t reuse_packfile_objects;\n static off_t reuse_packfile_offset;\n \n static int use_bitmap_index = 1;\n+static int write_bitmap_index;\n \n static unsigned long delta_cache_size = 0;\n static unsigned long max_delta_cache_size = 256 * 1024 * 1024;\n@@ -81,6 +82,24 @@ enum {\n static uint32_t written, written_delta;\n static uint32_t reused, reused_delta;\n \n+/*\n+ * Indexed commits\n+ */\n+static struct commit **indexed_commits;\n+static unsigned int indexed_commits_nr;\n+static unsigned int indexed_commits_alloc;\n+\n+static void index_commit_for_bitmap(struct commit *commit)\n+{\n+\tif (indexed_commits_nr >= indexed_commits_alloc) {\n+\t\tindexed_commits_alloc = (indexed_commits_alloc + 32) * 2;\n+\t\tindexed_commits = xrealloc(indexed_commits,\n+\t\t\tindexed_commits_alloc * sizeof(struct commit *));\n+\t}\n+\n+\tindexed_commits[indexed_commits_nr++] = commit;\n+}\n+\n static void *get_delta(struct object_entry *entry)\n {\n \tunsigned long size, base_size, delta_size;\n@@ -823,9 +842,30 @@ static void write_pack_file(void)\n \t\t\tif (sizeof(tmpname) <= strlen(base_name) + 50)\n \t\t\t\tdie(\"pack base name '%s' too long\", base_name);\n \t\t\tsnprintf(tmpname, sizeof(tmpname), \"%s-\", base_name);\n+\n+\t\t\tif (write_bitmap_index) {\n+\t\t\t\tbitmap_writer_set_checksum(sha1);\n+\t\t\t\tbitmap_writer_build_type_index(written_list, nr_written);\n+\t\t\t}\n+\n \t\t\tfinish_tmp_packfile(tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n \t\t\t\t\t    &pack_idx_opts, sha1);\n+\n+\t\t\tif (write_bitmap_index) {\n+\t\t\t\tchar *end_of_name_prefix = strrchr(tmpname, 0);\n+\t\t\t\tsprintf(end_of_name_prefix, \"%s.bitmap\", sha1_to_hex(sha1));\n+\n+\t\t\t\tstop_progress(&progress_state);\n+\n+\t\t\t\tbitmap_writer_show_progress(progress);\n+\t\t\t\tbitmap_writer_reuse_bitmaps(&to_pack);\n+\t\t\t\tbitmap_writer_select_commits(indexed_commits, indexed_commits_nr, -1);\n+\t\t\t\tbitmap_writer_build(&to_pack);\n+\t\t\t\tbitmap_writer_finish(written_list, nr_written, tmpname);\n+\t\t\t\twrite_bitmap_index = 0;\n+\t\t\t}\n+\n \t\t\tfree(pack_tmp_name);\n \t\t\tputs(sha1_to_hex(sha1));\n \t\t}\n@@ -2111,6 +2151,10 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n \t\tcache_max_small_delta_size = git_config_int(k, v);\n \t\treturn 0;\n \t}\n+\tif (!strcmp(k, \"pack.writebitmaps\")) {\n+\t\twrite_bitmap_index = git_config_bool(k, v);\n+\t\treturn 0;\n+\t}\n \tif (!strcmp(k, \"pack.usebitmaps\")) {\n \t\tuse_bitmap_index = git_config_bool(k, v);\n \t\treturn 0;\n@@ -2173,6 +2217,9 @@ static void show_commit(struct commit *commit, void *data)\n {\n \tadd_object_entry(commit->object.sha1, OBJ_COMMIT, NULL, 0);\n \tcommit->object.flags |= OBJECT_ADDED;\n+\n+\tif (write_bitmap_index)\n+\t\tindex_commit_for_bitmap(commit);\n }\n \n static void show_object(struct object *obj,\n@@ -2365,6 +2412,7 @@ static void get_object_list(int ac, const char **av)\n \t\tif (*line == '-') {\n \t\t\tif (!strcmp(line, \"--not\")) {\n \t\t\t\tflags ^= UNINTERESTING;\n+\t\t\t\twrite_bitmap_index = 0;\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tdie(\"not a rev '%s'\", line);\n@@ -2507,6 +2555,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t    N_(\"do not hide commits by grafts\"), 0),\n \t\tOPT_BOOL(0, \"use-bitmap-index\", &use_bitmap_index,\n \t\t\t N_(\"use a bitmap index if available to speed up counting objects\")),\n+\t\tOPT_BOOL(0, \"write-bitmap-index\", &write_bitmap_index,\n+\t\t\t N_(\"write a bitmap index together with the pack index\")),\n \t\tOPT_END(),\n \t};\n \n@@ -2576,6 +2626,9 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (!use_internal_rev_list || !pack_to_stdout || is_repository_shallow())\n \t\tuse_bitmap_index = 0;\n \n+\tif (pack_to_stdout || !rev_list_all)\n+\t\twrite_bitmap_index = 0;\n+\n \tif (progress && all_progress_implied)\n \t\tprogress = 2;\n \ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nnew file mode 100644\nindex 0000000..954a74d\n--- /dev/null\n+++ b/pack-bitmap-write.c\n@@ -0,0 +1,535 @@\n+#include \"cache.h\"\n+#include \"commit.h\"\n+#include \"tag.h\"\n+#include \"diff.h\"\n+#include \"revision.h\"\n+#include \"list-objects.h\"\n+#include \"progress.h\"\n+#include \"pack-revindex.h\"\n+#include \"pack.h\"\n+#include \"pack-bitmap.h\"\n+#include \"sha1-lookup.h\"\n+#include \"pack-objects.h\"\n+\n+struct bitmapped_commit {\n+\tstruct commit *commit;\n+\tstruct ewah_bitmap *bitmap;\n+\tstruct ewah_bitmap *write_as;\n+\tint flags;\n+\tint xor_offset;\n+\tuint32_t commit_pos;\n+};\n+\n+struct bitmap_writer {\n+\tstruct ewah_bitmap *commits;\n+\tstruct ewah_bitmap *trees;\n+\tstruct ewah_bitmap *blobs;\n+\tstruct ewah_bitmap *tags;\n+\n+\tkhash_sha1 *bitmaps;\n+\tkhash_sha1 *reused;\n+\tstruct packing_data *to_pack;\n+\n+\tstruct bitmapped_commit *selected;\n+\tunsigned int selected_nr, selected_alloc;\n+\n+\tstruct progress *progress;\n+\tint show_progress;\n+\tunsigned char pack_checksum[20];\n+};\n+\n+static struct bitmap_writer writer;\n+\n+void bitmap_writer_show_progress(int show)\n+{\n+\twriter.show_progress = show;\n+}\n+\n+/**\n+ * Build the initial type index for the packfile\n+ */\n+void bitmap_writer_build_type_index(struct pack_idx_entry **index,\n+\t\t\t\t    uint32_t index_nr)\n+{\n+\tuint32_t i;\n+\n+\twriter.commits = ewah_new();\n+\twriter.trees = ewah_new();\n+\twriter.blobs = ewah_new();\n+\twriter.tags = ewah_new();\n+\n+\tfor (i = 0; i < index_nr; ++i) {\n+\t\tstruct object_entry *entry = (struct object_entry *)index[i];\n+\t\tenum object_type real_type;\n+\n+\t\tentry->in_pack_pos = i;\n+\n+\t\tswitch (entry->type) {\n+\t\tcase OBJ_COMMIT:\n+\t\tcase OBJ_TREE:\n+\t\tcase OBJ_BLOB:\n+\t\tcase OBJ_TAG:\n+\t\t\treal_type = entry->type;\n+\t\t\tbreak;\n+\n+\t\tdefault:\n+\t\t\treal_type = sha1_object_info(entry->idx.sha1, NULL);\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\tswitch (real_type) {\n+\t\tcase OBJ_COMMIT:\n+\t\t\tewah_set(writer.commits, i);\n+\t\t\tbreak;\n+\n+\t\tcase OBJ_TREE:\n+\t\t\tewah_set(writer.trees, i);\n+\t\t\tbreak;\n+\n+\t\tcase OBJ_BLOB:\n+\t\t\tewah_set(writer.blobs, i);\n+\t\t\tbreak;\n+\n+\t\tcase OBJ_TAG:\n+\t\t\tewah_set(writer.tags, i);\n+\t\t\tbreak;\n+\n+\t\tdefault:\n+\t\t\tdie(\"Missing type information for %s (%d/%d)\",\n+\t\t\t    sha1_to_hex(entry->idx.sha1), real_type, entry->type);\n+\t\t}\n+\t}\n+}\n+\n+/**\n+ * Compute the actual bitmaps\n+ */\n+static struct object **seen_objects;\n+static unsigned int seen_objects_nr, seen_objects_alloc;\n+\n+static inline void push_bitmapped_commit(struct commit *commit, struct ewah_bitmap *reused)\n+{\n+\tif (writer.selected_nr >= writer.selected_alloc) {\n+\t\twriter.selected_alloc = (writer.selected_alloc + 32) * 2;\n+\t\twriter.selected = xrealloc(writer.selected,\n+\t\t\t\t\t   writer.selected_alloc * sizeof(struct bitmapped_commit));\n+\t}\n+\n+\twriter.selected[writer.selected_nr].commit = commit;\n+\twriter.selected[writer.selected_nr].bitmap = reused;\n+\twriter.selected[writer.selected_nr].flags = 0;\n+\n+\twriter.selected_nr++;\n+}\n+\n+static inline void mark_as_seen(struct object *object)\n+{\n+\tALLOC_GROW(seen_objects, seen_objects_nr + 1, seen_objects_alloc);\n+\tseen_objects[seen_objects_nr++] = object;\n+}\n+\n+static inline void reset_all_seen(void)\n+{\n+\tunsigned int i;\n+\tfor (i = 0; i < seen_objects_nr; ++i) {\n+\t\tseen_objects[i]->flags &= ~(SEEN | ADDED | SHOWN);\n+\t}\n+\tseen_objects_nr = 0;\n+}\n+\n+static uint32_t find_object_pos(const unsigned char *sha1)\n+{\n+\tstruct object_entry *entry = packlist_find(writer.to_pack, sha1, NULL);\n+\n+\tif (!entry) {\n+\t\tdie(\"Failed to write bitmap index. Packfile doesn't have full closure \"\n+\t\t\t\"(object %s is missing)\", sha1_to_hex(sha1));\n+\t}\n+\n+\treturn entry->in_pack_pos;\n+}\n+\n+static void show_object(struct object *object, const struct name_path *path,\n+\t\t\tconst char *last, void *data)\n+{\n+\tstruct bitmap *base = data;\n+\tbitmap_set(base, find_object_pos(object->sha1));\n+\tmark_as_seen(object);\n+}\n+\n+static void show_commit(struct commit *commit, void *data)\n+{\n+\tmark_as_seen((struct object *)commit);\n+}\n+\n+static int\n+add_to_include_set(struct bitmap *base, struct commit *commit)\n+{\n+\tkhiter_t hash_pos;\n+\tuint32_t bitmap_pos = find_object_pos(commit->object.sha1);\n+\n+\tif (bitmap_get(base, bitmap_pos))\n+\t\treturn 0;\n+\n+\thash_pos = kh_get_sha1(writer.bitmaps, commit->object.sha1);\n+\tif (hash_pos < kh_end(writer.bitmaps)) {\n+\t\tstruct bitmapped_commit *bc = kh_value(writer.bitmaps, hash_pos);\n+\t\tbitmap_or_ewah(base, bc->bitmap);\n+\t\treturn 0;\n+\t}\n+\n+\tbitmap_set(base, bitmap_pos);\n+\treturn 1;\n+}\n+\n+static int\n+should_include(struct commit *commit, void *_data)\n+{\n+\tstruct bitmap *base = _data;\n+\n+\tif (!add_to_include_set(base, commit)) {\n+\t\tstruct commit_list *parent = commit->parents;\n+\n+\t\tmark_as_seen((struct object *)commit);\n+\n+\t\twhile (parent) {\n+\t\t\tparent->item->object.flags |= SEEN;\n+\t\t\tmark_as_seen((struct object *)parent->item);\n+\t\t\tparent = parent->next;\n+\t\t}\n+\n+\t\treturn 0;\n+\t}\n+\n+\treturn 1;\n+}\n+\n+static void compute_xor_offsets(void)\n+{\n+\tstatic const int MAX_XOR_OFFSET_SEARCH = 10;\n+\n+\tint i, next = 0;\n+\n+\twhile (next < writer.selected_nr) {\n+\t\tstruct bitmapped_commit *stored = &writer.selected[next];\n+\n+\t\tint best_offset = 0;\n+\t\tstruct ewah_bitmap *best_bitmap = stored->bitmap;\n+\t\tstruct ewah_bitmap *test_xor;\n+\n+\t\tfor (i = 1; i <= MAX_XOR_OFFSET_SEARCH; ++i) {\n+\t\t\tint curr = next - i;\n+\n+\t\t\tif (curr < 0)\n+\t\t\t\tbreak;\n+\n+\t\t\ttest_xor = ewah_pool_new();\n+\t\t\tewah_xor(writer.selected[curr].bitmap, stored->bitmap, test_xor);\n+\n+\t\t\tif (test_xor->buffer_size < best_bitmap->buffer_size) {\n+\t\t\t\tif (best_bitmap != stored->bitmap)\n+\t\t\t\t\tewah_pool_free(best_bitmap);\n+\n+\t\t\t\tbest_bitmap = test_xor;\n+\t\t\t\tbest_offset = i;\n+\t\t\t} else {\n+\t\t\t\tewah_pool_free(test_xor);\n+\t\t\t}\n+\t\t}\n+\n+\t\tstored->xor_offset = best_offset;\n+\t\tstored->write_as = best_bitmap;\n+\n+\t\tnext++;\n+\t}\n+}\n+\n+void bitmap_writer_build(struct packing_data *to_pack)\n+{\n+\tstatic const double REUSE_BITMAP_THRESHOLD = 0.2;\n+\n+\tint i, reuse_after, need_reset;\n+\tstruct bitmap *base = bitmap_new();\n+\tstruct rev_info revs;\n+\n+\twriter.bitmaps = kh_init_sha1();\n+\twriter.to_pack = to_pack;\n+\n+\tif (writer.show_progress)\n+\t\twriter.progress = start_progress(\"Building bitmaps\", writer.selected_nr);\n+\n+\tinit_revisions(&revs, NULL);\n+\trevs.tag_objects = 1;\n+\trevs.tree_objects = 1;\n+\trevs.blob_objects = 1;\n+\trevs.no_walk = 0;\n+\n+\trevs.include_check = should_include;\n+\treset_revision_walk();\n+\n+\treuse_after = writer.selected_nr * REUSE_BITMAP_THRESHOLD;\n+\tneed_reset = 0;\n+\n+\tfor (i = writer.selected_nr - 1; i >= 0; --i) {\n+\t\tstruct bitmapped_commit *stored;\n+\t\tstruct object *object;\n+\n+\t\tkhiter_t hash_pos;\n+\t\tint hash_ret;\n+\n+\t\tstored = &writer.selected[i];\n+\t\tobject = (struct object *)stored->commit;\n+\n+\t\tif (stored->bitmap == NULL) {\n+\t\t\tif (i < writer.selected_nr - 1 &&\n+\t\t\t    (need_reset ||\n+\t\t\t     !in_merge_bases(writer.selected[i + 1].commit,\n+\t\t\t\t\t     stored->commit))) {\n+\t\t\t    bitmap_reset(base);\n+\t\t\t    reset_all_seen();\n+\t\t\t}\n+\n+\t\t\tadd_pending_object(&revs, object, \"\");\n+\t\t\trevs.include_check_data = base;\n+\n+\t\t\tif (prepare_revision_walk(&revs))\n+\t\t\t\tdie(\"revision walk setup failed\");\n+\n+\t\t\ttraverse_commit_list(&revs, show_commit, show_object, base);\n+\n+\t\t\trevs.pending.nr = 0;\n+\t\t\trevs.pending.alloc = 0;\n+\t\t\trevs.pending.objects = NULL;\n+\n+\t\t\tstored->bitmap = bitmap_to_ewah(base);\n+\t\t\tneed_reset = 0;\n+\t\t} else\n+\t\t\tneed_reset = 1;\n+\n+\t\tif (i >= reuse_after)\n+\t\t\tstored->flags |= BITMAP_FLAG_REUSE;\n+\n+\t\thash_pos = kh_put_sha1(writer.bitmaps, object->sha1, &hash_ret);\n+\t\tif (hash_ret == 0)\n+\t\t\tdie(\"Duplicate entry when writing index: %s\",\n+\t\t\t    sha1_to_hex(object->sha1));\n+\n+\t\tkh_value(writer.bitmaps, hash_pos) = stored;\n+\t\tdisplay_progress(writer.progress, writer.selected_nr - i);\n+\t}\n+\n+\tbitmap_free(base);\n+\tstop_progress(&writer.progress);\n+\n+\tcompute_xor_offsets();\n+}\n+\n+/**\n+ * Select the commits that will be bitmapped\n+ */\n+static inline unsigned int next_commit_index(unsigned int idx)\n+{\n+\tstatic const unsigned int MIN_COMMITS = 100;\n+\tstatic const unsigned int MAX_COMMITS = 5000;\n+\n+\tstatic const unsigned int MUST_REGION = 100;\n+\tstatic const unsigned int MIN_REGION = 20000;\n+\n+\tunsigned int offset, next;\n+\n+\tif (idx <= MUST_REGION)\n+\t\treturn 0;\n+\n+\tif (idx <= MIN_REGION) {\n+\t\toffset = idx - MUST_REGION;\n+\t\treturn (offset < MIN_COMMITS) ? offset : MIN_COMMITS;\n+\t}\n+\n+\toffset = idx - MIN_REGION;\n+\tnext = (offset < MAX_COMMITS) ? offset : MAX_COMMITS;\n+\n+\treturn (next > MIN_COMMITS) ? next : MIN_COMMITS;\n+}\n+\n+static int date_compare(const void *_a, const void *_b)\n+{\n+\tstruct commit *a = *(struct commit **)_a;\n+\tstruct commit *b = *(struct commit **)_b;\n+\treturn (long)b->date - (long)a->date;\n+}\n+\n+void bitmap_writer_reuse_bitmaps(struct packing_data *to_pack)\n+{\n+\tif (prepare_bitmap_git() < 0)\n+\t\treturn;\n+\n+\twriter.reused = kh_init_sha1();\n+\trebuild_existing_bitmaps(to_pack, writer.reused, writer.show_progress);\n+}\n+\n+static struct ewah_bitmap *find_reused_bitmap(const unsigned char *sha1)\n+{\n+\tkhiter_t hash_pos;\n+\n+\tif (!writer.reused)\n+\t\treturn NULL;\n+\n+\thash_pos = kh_get_sha1(writer.reused, sha1);\n+\tif (hash_pos >= kh_end(writer.reused))\n+\t\treturn NULL;\n+\n+\treturn kh_value(writer.reused, hash_pos);\n+}\n+\n+void bitmap_writer_select_commits(struct commit **indexed_commits,\n+\t\t\t\t  unsigned int indexed_commits_nr,\n+\t\t\t\t  int max_bitmaps)\n+{\n+\tunsigned int i = 0, j, next;\n+\n+\tqsort(indexed_commits, indexed_commits_nr, sizeof(indexed_commits[0]),\n+\t      date_compare);\n+\n+\tif (writer.show_progress)\n+\t\twriter.progress = start_progress(\"Selecting bitmap commits\", 0);\n+\n+\tif (indexed_commits_nr < 100) {\n+\t\tfor (i = 0; i < indexed_commits_nr; ++i)\n+\t\t\tpush_bitmapped_commit(indexed_commits[i], NULL);\n+\t\treturn;\n+\t}\n+\n+\tfor (;;) {\n+\t\tstruct ewah_bitmap *reused_bitmap = NULL;\n+\t\tstruct commit *chosen = NULL;\n+\n+\t\tnext = next_commit_index(i);\n+\n+\t\tif (i + next >= indexed_commits_nr)\n+\t\t\tbreak;\n+\n+\t\tif (max_bitmaps > 0 && writer.selected_nr >= max_bitmaps) {\n+\t\t\twriter.selected_nr = max_bitmaps;\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\tif (next == 0) {\n+\t\t\tchosen = indexed_commits[i];\n+\t\t\treused_bitmap = find_reused_bitmap(chosen->object.sha1);\n+\t\t} else {\n+\t\t\tchosen = indexed_commits[i + next];\n+\n+\t\t\tfor (j = 0; j <= next; ++j) {\n+\t\t\t\tstruct commit *cm = indexed_commits[i + j];\n+\n+\t\t\t\treused_bitmap = find_reused_bitmap(cm->object.sha1);\n+\t\t\t\tif (reused_bitmap || (cm->object.flags & NEEDS_BITMAP) != 0) {\n+\t\t\t\t\tchosen = cm;\n+\t\t\t\t\tbreak;\n+\t\t\t\t}\n+\n+\t\t\t\tif (cm->parents && cm->parents->next)\n+\t\t\t\t\tchosen = cm;\n+\t\t\t}\n+\t\t}\n+\n+\t\tpush_bitmapped_commit(chosen, reused_bitmap);\n+\n+\t\ti += next + 1;\n+\t\tdisplay_progress(writer.progress, i);\n+\t}\n+\n+\tstop_progress(&writer.progress);\n+}\n+\n+\n+static int sha1write_ewah_helper(void *f, const void *buf, size_t len)\n+{\n+\t/* sha1write will die on error */\n+\tsha1write(f, buf, len);\n+\treturn len;\n+}\n+\n+/**\n+ * Write the bitmap index to disk\n+ */\n+static inline void dump_bitmap(struct sha1file *f, struct ewah_bitmap *bitmap)\n+{\n+\tif (ewah_serialize_to(bitmap, sha1write_ewah_helper, f) < 0)\n+\t\tdie(\"Failed to write bitmap index\");\n+}\n+\n+static const unsigned char *sha1_access(size_t pos, void *table)\n+{\n+\tstruct pack_idx_entry **index = table;\n+\treturn index[pos]->sha1;\n+}\n+\n+static void write_selected_commits_v1(struct sha1file *f,\n+\t\t\t\t      struct pack_idx_entry **index,\n+\t\t\t\t      uint32_t index_nr)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i < writer.selected_nr; ++i) {\n+\t\tstruct bitmapped_commit *stored = &writer.selected[i];\n+\t\tstruct bitmap_disk_entry on_disk;\n+\n+\t\tint commit_pos =\n+\t\t\tsha1_pos(stored->commit->object.sha1, index, index_nr, sha1_access);\n+\n+\t\tif (commit_pos < 0)\n+\t\t\tdie(\"BUG: trying to write commit not in index\");\n+\n+\t\ton_disk.object_pos = htonl(commit_pos);\n+\t\ton_disk.xor_offset = stored->xor_offset;\n+\t\ton_disk.flags = stored->flags;\n+\n+\t\tsha1write(f, &on_disk, sizeof(on_disk));\n+\t\tdump_bitmap(f, stored->write_as);\n+\t}\n+}\n+\n+void bitmap_writer_set_checksum(unsigned char *sha1)\n+{\n+\thashcpy(writer.pack_checksum, sha1);\n+}\n+\n+void bitmap_writer_finish(struct pack_idx_entry **index,\n+\t\t\t  uint32_t index_nr,\n+\t\t\t  const char *filename)\n+{\n+\tstatic char tmp_file[PATH_MAX];\n+\tstatic uint16_t default_version = 1;\n+\tstatic uint16_t flags = BITMAP_OPT_FULL_DAG;\n+\tstruct sha1file *f;\n+\n+\tstruct bitmap_disk_header header;\n+\n+\tint fd = odb_mkstemp(tmp_file, sizeof(tmp_file), \"pack/tmp_bitmap_XXXXXX\");\n+\n+\tif (fd < 0)\n+\t\tdie_errno(\"unable to create '%s'\", tmp_file);\n+\tf = sha1fd(fd, tmp_file);\n+\n+\tmemcpy(header.magic, BITMAP_IDX_SIGNATURE, sizeof(BITMAP_IDX_SIGNATURE));\n+\theader.version = htons(default_version);\n+\theader.options = htons(flags);\n+\theader.entry_count = htonl(writer.selected_nr);\n+\tmemcpy(header.checksum, writer.pack_checksum, 20);\n+\n+\tsha1write(f, &header, sizeof(header));\n+\tdump_bitmap(f, writer.commits);\n+\tdump_bitmap(f, writer.trees);\n+\tdump_bitmap(f, writer.blobs);\n+\tdump_bitmap(f, writer.tags);\n+\twrite_selected_commits_v1(f, index, index_nr);\n+\n+\tsha1close(f, NULL, CSUM_FSYNC);\n+\n+\tif (adjust_shared_perm(tmp_file))\n+\t\tdie_errno(\"unable to make temporary bitmap file readable\");\n+\n+\tif (rename(tmp_file, filename))\n+\t\tdie_errno(\"unable to rename temporary bitmap file to '%s'\", filename);\n+}\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nindex 340f02b..7792dfd 100644\n--- a/pack-bitmap.c\n+++ b/pack-bitmap.c\n@@ -968,3 +968,95 @@ void test_bitmap_walk(struct rev_info *revs)\n \telse\n \t\tfprintf(stderr, \"Mismatch!\\n\");\n }\n+\n+static int rebuild_bitmap(uint32_t *reposition,\n+\t\t\t  struct ewah_bitmap *source,\n+\t\t\t  struct bitmap *dest)\n+{\n+\tuint32_t pos = 0;\n+\tstruct ewah_iterator it;\n+\teword_t word;\n+\n+\tewah_iterator_init(&it, source);\n+\n+\twhile (ewah_iterator_next(&word, &it)) {\n+\t\tuint32_t offset, bit_pos;\n+\n+\t\tfor (offset = 0; offset < BITS_IN_WORD; ++offset) {\n+\t\t\tif ((word >> offset) == 0)\n+\t\t\t\tbreak;\n+\n+\t\t\toffset += ewah_bit_ctz64(word >> offset);\n+\n+\t\t\tbit_pos = reposition[pos + offset];\n+\t\t\tif (bit_pos > 0)\n+\t\t\t\tbitmap_set(dest, bit_pos - 1);\n+\t\t\telse /* can't reuse, we don't have the object */\n+\t\t\t\treturn -1;\n+\t\t}\n+\n+\t\tpos += BITS_IN_WORD;\n+\t}\n+\treturn 0;\n+}\n+\n+int rebuild_existing_bitmaps(struct packing_data *mapping,\n+\t\t\t     khash_sha1 *reused_bitmaps,\n+\t\t\t     int show_progress)\n+{\n+\tuint32_t i, num_objects;\n+\tuint32_t *reposition;\n+\tstruct bitmap *rebuild;\n+\tstruct stored_bitmap *stored;\n+\tstruct progress *progress = NULL;\n+\n+\tkhiter_t hash_pos;\n+\tint hash_ret;\n+\n+\tif (prepare_bitmap_git() < 0)\n+\t\treturn -1;\n+\n+\tnum_objects = bitmap_git.pack->num_objects;\n+\treposition = xcalloc(num_objects, sizeof(uint32_t));\n+\n+\tfor (i = 0; i < num_objects; ++i) {\n+\t\tconst unsigned char *sha1;\n+\t\tstruct revindex_entry *entry;\n+\t\tstruct object_entry *oe;\n+\n+\t\tentry = &bitmap_git.reverse_index->revindex[i];\n+\t\tsha1 = nth_packed_object_sha1(bitmap_git.pack, entry->nr);\n+\t\toe = packlist_find(mapping, sha1, NULL);\n+\n+\t\tif (oe)\n+\t\t\treposition[i] = oe->in_pack_pos + 1;\n+\t}\n+\n+\trebuild = bitmap_new();\n+\ti = 0;\n+\n+\tif (show_progress)\n+\t\tprogress = start_progress(\"Reusing bitmaps\", 0);\n+\n+\tkh_foreach_value(bitmap_git.bitmaps, stored, {\n+\t\tif (stored->flags & BITMAP_FLAG_REUSE) {\n+\t\t\tif (!rebuild_bitmap(reposition,\n+\t\t\t\t\t    lookup_stored_bitmap(stored),\n+\t\t\t\t\t    rebuild)) {\n+\t\t\t\thash_pos = kh_put_sha1(reused_bitmaps,\n+\t\t\t\t\t\t       stored->sha1,\n+\t\t\t\t\t\t       &hash_ret);\n+\t\t\t\tkh_value(reused_bitmaps, hash_pos) =\n+\t\t\t\t\tbitmap_to_ewah(rebuild);\n+\t\t\t}\n+\t\t\tbitmap_reset(rebuild);\n+\t\t\tdisplay_progress(progress, ++i);\n+\t\t}\n+\t});\n+\n+\tstop_progress(&progress);\n+\n+\tfree(reposition);\n+\tbitmap_free(rebuild);\n+\treturn 0;\n+}\ndiff --git a/pack-bitmap.h b/pack-bitmap.h\nindex 9351aa5..5e3fb9a 100644\n--- a/pack-bitmap.h\n+++ b/pack-bitmap.h\n@@ -3,6 +3,7 @@\n \n #include \"ewah/ewok.h\"\n #include \"khash.h\"\n+#include \"pack-objects.h\"\n \n struct bitmap_disk_entry {\n \tuint32_t object_pos;\n@@ -20,10 +21,16 @@ struct bitmap_disk_header {\n \n static const char BITMAP_IDX_SIGNATURE[] = {'B', 'I', 'T', 'M'};;\n \n+#define NEEDS_BITMAP (1u<<22)\n+\n enum pack_bitmap_opts {\n \tBITMAP_OPT_FULL_DAG = 1,\n };\n \n+enum pack_bitmap_flags {\n+\tBITMAP_FLAG_REUSE = 0x1\n+};\n+\n typedef int (*show_reachable_fn)(\n \tconst unsigned char *sha1,\n \tenum object_type type,\n@@ -39,5 +46,17 @@ void test_bitmap_walk(struct rev_info *revs);\n char *pack_bitmap_filename(struct packed_git *p);\n int prepare_bitmap_walk(struct rev_info *revs);\n int reuse_partial_packfile_from_bitmap(struct packed_git **packfile, uint32_t *entries, off_t *up_to);\n+int rebuild_existing_bitmaps(struct packing_data *mapping, khash_sha1 *reused_bitmaps, int show_progress);\n+\n+void bitmap_writer_show_progress(int show);\n+void bitmap_writer_set_checksum(unsigned char *sha1);\n+void bitmap_writer_build_type_index(struct pack_idx_entry **index, uint32_t index_nr);\n+void bitmap_writer_reuse_bitmaps(struct packing_data *to_pack);\n+void bitmap_writer_select_commits(struct commit **indexed_commits,\n+\t\tunsigned int indexed_commits_nr, int max_bitmaps);\n+void bitmap_writer_build(struct packing_data *to_pack);\n+void bitmap_writer_finish(struct pack_idx_entry **index,\n+\t\t\t  uint32_t index_nr,\n+\t\t\t  const char *filename);\n \n #endif\ndiff --git a/pack-objects.h b/pack-objects.h\nindex 90ad0a8..d1b98b3 100644\n--- a/pack-objects.h\n+++ b/pack-objects.h\n@@ -17,6 +17,7 @@ struct object_entry {\n \tenum object_type type;\n \tenum object_type in_pack_type;\t/* could be delta */\n \tuint32_t hash;\t\t\t/* name hint hash */\n+\tunsigned int in_pack_pos;\n \tunsigned char in_pack_header_size;\n \tunsigned preferred_base:1; /*\n \t\t\t\t    * we do not pack this, but is available\ndiff --git a/pack-write.c b/pack-write.c\nindex ca9e63b..6203d37 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -371,5 +371,7 @@ void finish_tmp_packfile(char *name_buffer,\n \tif (rename(idx_tmp_name, name_buffer))\n \t\tdie_errno(\"unable to rename temporary index file\");\n \n+\t*end_of_name_prefix = '\\0';\n+\n \tfree((void *)idx_tmp_name);\n }\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230612","messageId":"20131114124554.GN10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 14/21] repack: stop using magic number for ARRAY_SIZE(exts)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:45:54Z","receivedAt":"2013-11-14T12:45:54Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"We have a static array of extensions, but hardcode the size\nof the array in our loops. Let's pull out this magic number,\nwhich will make it easier to change.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n builtin/repack.c | 8 ++++----\n 1 file changed, 4 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex a0ff5c7..2e88975 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -115,7 +115,7 @@ static void remove_redundant_pack(const char *dir_name, const char *base_name)\n \n int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n-\tconst char *exts[2] = {\".pack\", \".idx\"};\n+\tconst char *exts[] = {\".pack\", \".idx\"};\n \tstruct child_process cmd;\n \tstruct string_list_item *item;\n \tstruct argv_array cmd_args = ARGV_ARRAY_INIT;\n@@ -258,7 +258,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t */\n \tfailed = 0;\n \tfor_each_string_list_item(item, &names) {\n-\t\tfor (ext = 0; ext < 2; ext++) {\n+\t\tfor (ext = 0; ext < ARRAY_SIZE(exts); ext++) {\n \t\t\tchar *fname, *fname_old;\n \t\t\tfname = mkpathdup(\"%s/%s%s\", packdir,\n \t\t\t\t\t\titem->string, exts[ext]);\n@@ -315,7 +315,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \n \t/* Now the ones with the same name are out of the way... */\n \tfor_each_string_list_item(item, &names) {\n-\t\tfor (ext = 0; ext < 2; ext++) {\n+\t\tfor (ext = 0; ext < ARRAY_SIZE(exts); ext++) {\n \t\t\tchar *fname, *fname_old;\n \t\t\tstruct stat statbuffer;\n \t\t\tfname = mkpathdup(\"%s/pack-%s%s\",\n@@ -335,7 +335,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \n \t/* Remove the \"old-\" files */\n \tfor_each_string_list_item(item, &names) {\n-\t\tfor (ext = 0; ext < 2; ext++) {\n+\t\tfor (ext = 0; ext < ARRAY_SIZE(exts); ext++) {\n \t\t\tchar *fname;\n \t\t\tfname = mkpath(\"%s/old-pack-%s%s\",\n \t\t\t\t\tpackdir,\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230613","messageId":"20131114124600.GO10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 15/21] repack: turn exts array into array-of-struct","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:46:01Z","receivedAt":"2013-11-14T12:46:01Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"This is slightly more verbose, but will let us annotate the\nextensions with further options in future commits.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n builtin/repack.c | 17 +++++++++++------\n 1 file changed, 11 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 2e88975..a176de2 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -115,7 +115,12 @@ static void remove_redundant_pack(const char *dir_name, const char *base_name)\n \n int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n-\tconst char *exts[] = {\".pack\", \".idx\"};\n+\tstruct {\n+\t\tconst char *name;\n+\t} exts[] = {\n+\t\t{\".pack\"},\n+\t\t{\".idx\"},\n+\t};\n \tstruct child_process cmd;\n \tstruct string_list_item *item;\n \tstruct argv_array cmd_args = ARGV_ARRAY_INIT;\n@@ -261,14 +266,14 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tfor (ext = 0; ext < ARRAY_SIZE(exts); ext++) {\n \t\t\tchar *fname, *fname_old;\n \t\t\tfname = mkpathdup(\"%s/%s%s\", packdir,\n-\t\t\t\t\t\titem->string, exts[ext]);\n+\t\t\t\t\t\titem->string, exts[ext].name);\n \t\t\tif (!file_exists(fname)) {\n \t\t\t\tfree(fname);\n \t\t\t\tcontinue;\n \t\t\t}\n \n \t\t\tfname_old = mkpath(\"%s/old-%s%s\", packdir,\n-\t\t\t\t\t\titem->string, exts[ext]);\n+\t\t\t\t\t\titem->string, exts[ext].name);\n \t\t\tif (file_exists(fname_old))\n \t\t\t\tif (unlink(fname_old))\n \t\t\t\t\tfailed = 1;\n@@ -319,9 +324,9 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\tchar *fname, *fname_old;\n \t\t\tstruct stat statbuffer;\n \t\t\tfname = mkpathdup(\"%s/pack-%s%s\",\n-\t\t\t\t\tpackdir, item->string, exts[ext]);\n+\t\t\t\t\tpackdir, item->string, exts[ext].name);\n \t\t\tfname_old = mkpathdup(\"%s-%s%s\",\n-\t\t\t\t\tpacktmp, item->string, exts[ext]);\n+\t\t\t\t\tpacktmp, item->string, exts[ext].name);\n \t\t\tif (!stat(fname_old, &statbuffer)) {\n \t\t\t\tstatbuffer.st_mode &= ~(S_IWUSR | S_IWGRP | S_IWOTH);\n \t\t\t\tchmod(fname_old, statbuffer.st_mode);\n@@ -340,7 +345,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\tfname = mkpath(\"%s/old-pack-%s%s\",\n \t\t\t\t\tpackdir,\n \t\t\t\t\titem->string,\n-\t\t\t\t\texts[ext]);\n+\t\t\t\t\texts[ext].name);\n \t\t\tif (remove_path(fname))\n \t\t\t\twarning(_(\"removing '%s' failed\"), fname);\n \t\t}\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230614","messageId":"20131114124605.GP10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 16/21] repack: handle optional files created by pack-objects","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:46:05Z","receivedAt":"2013-11-14T12:46:05Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"We ask pack-objects to pack to a set of temporary files, and\nthen rename them into place. Some files that pack-objects\ncreates may be optional (like a .bitmap file), in which case\nwe would not want to call rename(). We already call stat()\nand make the chmod optional if the file cannot be accessed.\nWe could simply skip the rename step in this case, but that\nwould be a minor regression in noticing problems with\nnon-optional files (like the .pack and .idx files).\n\nInstead, we can now annotate extensions as optional, and\nskip them if they don't exist (and otherwise rely on\nrename() to barf).\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n builtin/repack.c | 9 +++++++--\n 1 file changed, 7 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex a176de2..8b7dfd0 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -117,6 +117,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n {\n \tstruct {\n \t\tconst char *name;\n+\t\tunsigned optional:1;\n \t} exts[] = {\n \t\t{\".pack\"},\n \t\t{\".idx\"},\n@@ -323,6 +324,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tfor (ext = 0; ext < ARRAY_SIZE(exts); ext++) {\n \t\t\tchar *fname, *fname_old;\n \t\t\tstruct stat statbuffer;\n+\t\t\tint exists = 0;\n \t\t\tfname = mkpathdup(\"%s/pack-%s%s\",\n \t\t\t\t\tpackdir, item->string, exts[ext].name);\n \t\t\tfname_old = mkpathdup(\"%s-%s%s\",\n@@ -330,9 +332,12 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\t\tif (!stat(fname_old, &statbuffer)) {\n \t\t\t\tstatbuffer.st_mode &= ~(S_IWUSR | S_IWGRP | S_IWOTH);\n \t\t\t\tchmod(fname_old, statbuffer.st_mode);\n+\t\t\t\texists = 1;\n+\t\t\t}\n+\t\t\tif (exists || !exts[ext].optional) {\n+\t\t\t\tif (rename(fname_old, fname))\n+\t\t\t\t\tdie_errno(_(\"renaming '%s' failed\"), fname_old);\n \t\t\t}\n-\t\t\tif (rename(fname_old, fname))\n-\t\t\t\tdie_errno(_(\"renaming '%s' failed\"), fname_old);\n \t\t\tfree(fname);\n \t\t\tfree(fname_old);\n \t\t}\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230615","messageId":"20131114124610.GQ10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 17/21] repack: consider bitmaps when performing repacks","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:46:10Z","receivedAt":"2013-11-14T12:46:10Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nSince `pack-objects` will write a `.bitmap` file next to the `.pack` and\n`.idx` files, this commit teaches `git-repack` to consider the new\nbitmap indexes (if they exist) when performing repack operations.\n\nThis implies moving old bitmap indexes out of the way if we are\nrepacking a repository that already has them, and moving the newly\ngenerated bitmap indexes into the `objects/pack` directory, next to\ntheir corresponding packfiles.\n\nSince `git repack` is now capable of handling these `.bitmap` files,\na normal `git gc` run on a repository that has `pack.writebitmaps` set\nto true in its config file will generate bitmap indexes as part of the\ngarbage collection process.\n\nAlternatively, `git repack` can be called with the `-b` switch to\nexplicitly generate bitmap indexes if you are experimenting\nand don't want them on all the time.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Documentation/git-repack.txt | 9 ++++++++-\n builtin/repack.c             | 9 ++++++++-\n 2 files changed, 16 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/git-repack.txt b/Documentation/git-repack.txt\nindex 509cf73..002cfd5 100644\n--- a/Documentation/git-repack.txt\n+++ b/Documentation/git-repack.txt\n@@ -9,7 +9,7 @@ git-repack - Pack unpacked objects in a repository\n SYNOPSIS\n --------\n [verse]\n-'git repack' [-a] [-A] [-d] [-f] [-F] [-l] [-n] [-q] [--window=<n>] [--depth=<n>]\n+'git repack' [-a] [-A] [-d] [-f] [-F] [-l] [-n] [-q] [-b] [--window=<n>] [--depth=<n>]\n \n DESCRIPTION\n -----------\n@@ -110,6 +110,13 @@ other objects in that pack they already have locally.\n \tThe default is unlimited, unless the config variable\n \t`pack.packSizeLimit` is set.\n \n+-b::\n+--write-bitmap-index::\n+\tWrite a reachability bitmap index as part of the repack. This\n+\tonly makes sense when used with `-a` or `-A`, as the bitmaps\n+\tmust be able to refer to all reachable objects. This option\n+\toverrides the setting of `pack.writebitmaps`.\n+\n \n Configuration\n -------------\ndiff --git a/builtin/repack.c b/builtin/repack.c\nindex 8b7dfd0..239f278 100644\n--- a/builtin/repack.c\n+++ b/builtin/repack.c\n@@ -94,7 +94,7 @@ static void get_non_kept_pack_filenames(struct string_list *fname_list)\n \n static void remove_redundant_pack(const char *dir_name, const char *base_name)\n {\n-\tconst char *exts[] = {\".pack\", \".idx\", \".keep\"};\n+\tconst char *exts[] = {\".pack\", \".idx\", \".keep\", \".bitmap\"};\n \tint i;\n \tstruct strbuf buf = STRBUF_INIT;\n \tsize_t plen;\n@@ -121,6 +121,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t} exts[] = {\n \t\t{\".pack\"},\n \t\t{\".idx\"},\n+\t\t{\".bitmap\", 1},\n \t};\n \tstruct child_process cmd;\n \tstruct string_list_item *item;\n@@ -143,6 +144,7 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \tint no_update_server_info = 0;\n \tint quiet = 0;\n \tint local = 0;\n+\tint write_bitmap = -1;\n \n \tstruct option builtin_repack_options[] = {\n \t\tOPT_BIT('a', NULL, &pack_everything,\n@@ -161,6 +163,8 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\tOPT__QUIET(&quiet, N_(\"be quiet\")),\n \t\tOPT_BOOL('l', \"local\", &local,\n \t\t\t\tN_(\"pass --local to git-pack-objects\")),\n+\t\tOPT_BOOL('b', \"write-bitmap-index\", &write_bitmap,\n+\t\t\t\tN_(\"write bitmap index\")),\n \t\tOPT_STRING(0, \"unpack-unreachable\", &unpack_unreachable, N_(\"approxidate\"),\n \t\t\t\tN_(\"with -A, do not loosen objects older than this\")),\n \t\tOPT_INTEGER(0, \"window\", &window,\n@@ -202,6 +206,9 @@ int cmd_repack(int argc, const char **argv, const char *prefix)\n \t\targv_array_pushf(&cmd_args, \"--no-reuse-delta\");\n \tif (no_reuse_object)\n \t\targv_array_pushf(&cmd_args, \"--no-reuse-object\");\n+\tif (write_bitmap >= 0)\n+\t\targv_array_pushf(&cmd_args, \"--%swrite-bitmap-index\",\n+\t\t\t\t write_bitmap ? \"\" : \"no-\");\n \n \tif (pack_everything & ALL_INTO_ONE) {\n \t\tget_non_kept_pack_filenames(&existing_packs);\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230616","messageId":"20131114124618.GR10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 18/21] count-objects: recognize .bitmap in garbage-checking","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:46:18Z","receivedAt":"2013-11-14T12:46:18Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n\nCount-objects will report any \"garbage\" files in the packs\ndirectory, including files whose extensions it does not\nknow (case 1), and files whose matching \".pack\" file is\nmissing (case 2).  Without having learned about \".bitmap\"\nfiles, the current code reports all such files as garbage\n(case 1), even if their pack exists. Instead, they should be\ntreated as case 2.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1_file.c | 1 +\n 1 file changed, 1 insertion(+)\n\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 5557bd9..d810b58 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1194,6 +1194,7 @@ static void prepare_packed_git_one(char *objdir, int local)\n \n \t\tif (has_extension(de->d_name, \".idx\") ||\n \t\t    has_extension(de->d_name, \".pack\") ||\n+\t\t    has_extension(de->d_name, \".bitmap\") ||\n \t\t    has_extension(de->d_name, \".keep\"))\n \t\t\tstring_list_append(&garbage, path);\n \t\telse\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230617","messageId":"20131114124635.GS10757@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 19/21] t: add basic bitmap functionality tests","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:46:35Z","receivedAt":"2013-11-14T12:46:35Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"Now that we can read and write bitmaps, we can exercise them\nwith some basic functionality tests. These tests aren't\nparticularly useful for seeing the benefit, as the test\nrepo is too small for it to make a difference. However, we\ncan at least check that using bitmaps does not break anything.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n t/t5310-pack-bitmaps.sh | 138 ++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 138 insertions(+)\n create mode 100755 t/t5310-pack-bitmaps.sh\n\ndiff --git a/t/t5310-pack-bitmaps.sh b/t/t5310-pack-bitmaps.sh\nnew file mode 100755\nindex 0000000..bd743f4\n--- /dev/null\n+++ b/t/t5310-pack-bitmaps.sh\n@@ -0,0 +1,138 @@\n+#!/bin/sh\n+\n+test_description='exercise basic bitmap functionality'\n+. ./test-lib.sh\n+\n+test_expect_success 'setup repo with moderate-sized history' '\n+\tfor i in $(test_seq 1 10); do\n+\t\ttest_commit $i\n+\tdone &&\n+\tgit checkout -b other HEAD~5 &&\n+\tfor i in $(test_seq 1 10); do\n+\t\ttest_commit side-$i\n+\tdone &&\n+\tgit checkout master &&\n+\tblob=$(echo tagged-blob | git hash-object -w --stdin) &&\n+\tgit tag tagged-blob $blob &&\n+\tgit config pack.writebitmaps true\n+'\n+\n+test_expect_success 'full repack creates bitmaps' '\n+\tgit repack -ad &&\n+\tls .git/objects/pack/ | grep bitmap >output &&\n+\ttest_line_count = 1 output\n+'\n+\n+test_expect_success 'rev-list --test-bitmap verifies bitmaps' '\n+\tgit rev-list --test-bitmap HEAD\n+'\n+\n+rev_list_tests() {\n+\tstate=$1\n+\n+\ttest_expect_success \"counting commits via bitmap ($state)\" '\n+\t\tgit rev-list --count HEAD >expect &&\n+\t\tgit rev-list --use-bitmap-index --count HEAD >actual &&\n+\t\ttest_cmp expect actual\n+\t'\n+\n+\ttest_expect_success \"counting partial commits via bitmap ($state)\" '\n+\t\tgit rev-list --count HEAD~5..HEAD >expect &&\n+\t\tgit rev-list --use-bitmap-index --count HEAD~5..HEAD >actual &&\n+\t\ttest_cmp expect actual\n+\t'\n+\n+\ttest_expect_success \"counting non-linear history ($state)\" '\n+\t\tgit rev-list --count other...master >expect &&\n+\t\tgit rev-list --use-bitmap-index --count other...master >actual &&\n+\t\ttest_cmp expect actual\n+\t'\n+\n+\ttest_expect_success \"enumerate --objects ($state)\" '\n+\t\tgit rev-list --objects --use-bitmap-index HEAD >tmp &&\n+\t\tcut -d\" \" -f1 <tmp >tmp2 &&\n+\t\tsort <tmp2 >actual &&\n+\t\tgit rev-list --objects HEAD >tmp &&\n+\t\tcut -d\" \" -f1 <tmp >tmp2 &&\n+\t\tsort <tmp2 >expect &&\n+\t\ttest_cmp expect actual\n+\t'\n+\n+\ttest_expect_success \"bitmap --objects handles non-commit objects ($state)\" '\n+\t\tgit rev-list --objects --use-bitmap-index HEAD tagged-blob >actual &&\n+\t\tgrep $blob actual\n+\t'\n+}\n+\n+rev_list_tests 'full bitmap'\n+\n+test_expect_success 'clone from bitmapped repository' '\n+\tgit clone --no-local --bare . clone.git &&\n+\tgit rev-parse HEAD >expect &&\n+\tgit --git-dir=clone.git rev-parse HEAD >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'setup further non-bitmapped commits' '\n+\tfor i in $(test_seq 1 10); do\n+\t\ttest_commit further-$i\n+\tdone\n+'\n+\n+rev_list_tests 'partial bitmap'\n+\n+test_expect_success 'fetch (partial bitmap)' '\n+\tgit --git-dir=clone.git fetch origin master:master &&\n+\tgit rev-parse HEAD >expect &&\n+\tgit --git-dir=clone.git rev-parse HEAD >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'incremental repack cannot create bitmaps' '\n+\ttest_commit more-1 &&\n+\ttest_must_fail git repack -d\n+'\n+\n+test_expect_success 'incremental repack can disable bitmaps' '\n+\ttest_commit more-2 &&\n+\tgit repack -d --no-write-bitmap-index\n+'\n+\n+test_expect_success 'full repack, reusing previous bitmaps' '\n+\tgit repack -ad &&\n+\tls .git/objects/pack/ | grep bitmap >output &&\n+\ttest_line_count = 1 output\n+'\n+\n+test_expect_success 'fetch (full bitmap)' '\n+\tgit --git-dir=clone.git fetch origin master:master &&\n+\tgit rev-parse HEAD >expect &&\n+\tgit --git-dir=clone.git rev-parse HEAD >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_lazy_prereq JGIT '\n+\ttype jgit\n+'\n+\n+test_expect_success JGIT 'we can read jgit bitmaps' '\n+\tgit clone . compat-jgit &&\n+\t(\n+\t\tcd compat-jgit &&\n+\t\trm -f .git/objects/pack/*.bitmap &&\n+\t\tjgit gc &&\n+\t\tgit rev-list --test-bitmap HEAD\n+\t)\n+'\n+\n+test_expect_success JGIT 'jgit can read our bitmaps' '\n+\tgit clone . compat-us.git &&\n+\t(\n+\t\tcd compat-us.git &&\n+\t\tgit repack -adb &&\n+\t\t# jgit gc will barf if it does not like our bitmaps\n+\t\tjgit gc\n+\t)\n+'\n+\n+test_done\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230618","messageId":"20131114124834.GA11612@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 20/21] t/perf: add tests for pack bitmaps","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:48:34Z","receivedAt":"2013-11-14T12:48:34Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"This adds a few basic perf tests for the pack bitmap code to\nshow off its improvements. The tests are:\n\n  1. How long does it take to do a repack (it gets slower\n     with bitmaps, since we have to do extra work)?\n\n  2. How long does it take to do a clone (it gets faster\n     with bitmaps)?\n\n  3. How does a small fetch perform when we've just\n     repacked?\n\n  4. How does a clone perform when we haven't repacked since\n     a week of pushes?\n\nHere are results against linux.git:\n\nTest                      origin/master       this tree\n-----------------------------------------------------------------------\n5310.2: repack to disk    33.64(32.64+2.04)   67.67(66.75+1.84) +101.2%\n5310.3: simulated clone   30.49(29.47+2.05)   1.20(1.10+0.10) -96.1%\n5310.4: simulated fetch   3.49(6.79+0.06)     5.57(22.35+0.07) +59.6%\n5310.6: partial bitmap    36.70(43.87+1.81)   8.18(21.92+0.73) -77.7%\n\nYou can see that we do take longer to repack, but we do way\nbetter for further clones. A small fetch performs a bit\nworse, as we spend way more time on delta compression (note\nthe heavy user CPU time, as we have 8 threads) due to the\nlack of name hashes for the bitmapped objects.\n\nThe final test shows how the bitmaps degrade over time\nbetween packs. There's still a significant speedup over the\nnon-bitmap case, but we don't do quite as well (we have to\nspend time accessing the \"new\" objects the old fashioned\nway, including delta compression).\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n t/perf/p5310-pack-bitmaps.sh | 56 ++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 56 insertions(+)\n create mode 100755 t/perf/p5310-pack-bitmaps.sh\n\ndiff --git a/t/perf/p5310-pack-bitmaps.sh b/t/perf/p5310-pack-bitmaps.sh\nnew file mode 100755\nindex 0000000..ac036d0\n--- /dev/null\n+++ b/t/perf/p5310-pack-bitmaps.sh\n@@ -0,0 +1,56 @@\n+#!/bin/sh\n+\n+test_description='Tests pack performance using bitmaps'\n+. ./perf-lib.sh\n+\n+test_perf_large_repo\n+\n+# note that we do everything through config,\n+# since we want to be able to compare bitmap-aware\n+# git versus non-bitmap git\n+test_expect_success 'setup bitmap config' '\n+\tgit config pack.writebitmaps true\n+'\n+\n+test_perf 'repack to disk' '\n+\tgit repack -ad\n+'\n+\n+test_perf 'simulated clone' '\n+\tgit pack-objects --stdout --all </dev/null >/dev/null\n+'\n+\n+test_perf 'simulated fetch' '\n+\thave=$(git rev-list HEAD --until=1.week.ago -1) &&\n+\t{\n+\t\techo HEAD &&\n+\t\techo ^$have\n+\t} | git pack-objects --revs --stdout >/dev/null\n+'\n+\n+test_expect_success 'create partial bitmap state' '\n+\t# pick a commit to represent the repo tip in the past\n+\tcutoff=$(git rev-list HEAD --until=1.week.ago -1) &&\n+\torig_tip=$(git rev-parse HEAD) &&\n+\n+\t# now kill off all of the refs and pretend we had\n+\t# just the one tip\n+\trm -rf .git/logs .git/refs/* .git/packed-refs\n+\tgit update-ref HEAD $cutoff\n+\n+\t# and then repack, which will leave us with a nice\n+\t# big bitmap pack of the \"old\" history, and all of\n+\t# the new history will be loose, as if it had been pushed\n+\t# up incrementally and exploded via unpack-objects\n+\tgit repack -Ad\n+\n+\t# and now restore our original tip, as if the pushes\n+\t# had happened\n+\tgit update-ref HEAD $orig_tip\n+'\n+\n+test_perf 'partial bitmap' '\n+\tgit pack-objects --stdout --all </dev/null >/dev/null\n+'\n+\n+test_done\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230619","messageId":"20131114124855.GB11612@sigill.intra.peff.net","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"[PATCH v3 21/21] pack-bitmap: implement optional name_hash cache","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T12:48:55Z","receivedAt":"2013-11-14T12:48:55Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Vicent Marti <tanoku@gmail.com>\n\nWhen we use pack bitmaps rather than walking the object\ngraph, we end up with the list of objects to include in the\npackfile, but we do not know the path at which any tree or\nblob objects would be found.\n\nIn a recently packed repository, this is fine. A fetch would\nuse the paths only as a heuristic in the delta compression\nphase, and a fully packed repository should not need to do\nmuch delta compression.\n\nAs time passes, though, we may acquire more objects on top\nof our large bitmapped pack. If clients fetch frequently,\nthen they never even look at the bitmapped history, and all\nworks as usual. However, a client who has not fetched since\nthe last bitmap repack will have \"have\" tips in the\nbitmapped history, but \"want\" newer objects.\n\nThe bitmaps themselves degrade gracefully in this\ncircumstance. We manually walk the more recent bits of\nhistory, and then use bitmaps when we hit them.\n\nBut we would also like to perform delta compression between\nthe newer objects and the bitmapped objects (both to delta\nagainst what we know the user already has, but also between\n\"new\" and \"old\" objects that the user is fetching). The lack\nof pathnames makes our delta heuristics much less effective.\n\nThis patch adds an optional cache of the 32-bit name_hash\nvalues to the end of the bitmap file. If present, a reader\ncan use it to match bitmapped and non-bitmapped names during\ndelta compression.\n\nHere are perf results for p5310:\n\nTest                      origin/master       HEAD^                      HEAD\n-------------------------------------------------------------------------------------------------\n5310.2: repack to disk    36.81(37.82+1.43)   47.70(48.74+1.41) +29.6%   47.75(48.70+1.51) +29.7%\n5310.3: simulated clone   30.78(29.70+2.14)   1.08(0.97+0.10) -96.5%     1.07(0.94+0.12) -96.5%\n5310.4: simulated fetch   3.16(6.10+0.08)     3.54(10.65+0.06) +12.0%    1.70(3.07+0.06) -46.2%\n5310.6: partial bitmap    36.76(43.19+1.81)   6.71(11.25+0.76) -81.7%    4.08(6.26+0.46) -88.9%\n\nYou can see that the time spent on an incremental fetch goes\ndown, as our delta heuristics are able to do their work.\nAnd we save time on the partial bitmap clone for the same\nreason.\n\nSigned-off-by: Vicent Marti <tanoku@gmail.com>\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Documentation/config.txt                  | 11 +++++++++++\n Documentation/technical/bitmap-format.txt | 33 +++++++++++++++++++++++++++++++\n builtin/pack-objects.c                    | 10 +++++++++-\n pack-bitmap-write.c                       | 21 ++++++++++++++++++--\n pack-bitmap.c                             | 11 +++++++++++\n pack-bitmap.h                             |  4 +++-\n t/perf/p5310-pack-bitmaps.sh              |  3 ++-\n t/t5310-pack-bitmaps.sh                   |  3 ++-\n 8 files changed, 90 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex a439a4c..e6d3922 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -1872,6 +1872,17 @@ pack.writebitmaps::\n \tspace and extra time spent on the initial repack.  Defaults to\n \tfalse.\n \n+pack.writeBitmapHashCache::\n+\tWhen true, git will include a \"hash cache\" section in the bitmap\n+\tindex (if one is written). This cache can be used to feed git's\n+\tdelta heuristics, potentially leading to better deltas between\n+\tbitmapped and non-bitmapped objects (e.g., when serving a fetch\n+\tbetween an older, bitmapped pack and objects that have been\n+\tpushed since the last gc). The downside is that it consumes 4\n+\tbytes per object of disk space, and that JGit's bitmap\n+\timplementation does not understand it, causing it to complain if\n+\tGit and JGit are used on the same repository. Defaults to false.\n+\n pager.<cmd>::\n \tIf the value is boolean, turns on or off pagination of the\n \toutput of a particular Git subcommand when writing to a tty.\ndiff --git a/Documentation/technical/bitmap-format.txt b/Documentation/technical/bitmap-format.txt\nindex 7a86bd7..f8c18a0 100644\n--- a/Documentation/technical/bitmap-format.txt\n+++ b/Documentation/technical/bitmap-format.txt\n@@ -21,6 +21,12 @@ GIT bitmap v1 format\n \t\t\trequirement for the bitmap index format, also present in JGit,\n \t\t\tthat greatly reduces the complexity of the implementation.\n \n+\t\t\t- BITMAP_OPT_HASH_CACHE (0x4)\n+\t\t\tIf present, the end of the bitmap file contains\n+\t\t\t`N` 32-bit name-hash values, one per object in the\n+\t\t\tpack. The format and meaning of the name-hash is\n+\t\t\tdescribed below.\n+\n \t\t4-byte entry count (network byte order)\n \n \t\t\tThe total count of entries (bitmapped commits) in this bitmap index.\n@@ -129,3 +135,30 @@ The bitstream represented by the above chunk is then:\n The next word after `L_M` (if any) must again be a RLW, for the next\n chunk.  For efficient appending to the bitstream, the EWAH stores a\n pointer to the last RLW in the stream.\n+\n+\n+== Appendix B: Optional Bitmap Sections\n+\n+These sections may or may not be present in the `.bitmap` file; their\n+presence is indicated by the header flags section described above.\n+\n+Name-hash cache\n+---------------\n+\n+If the BITMAP_OPT_HASH_CACHE flag is set, the end of the bitmap contains\n+a cache of 32-bit values, one per object in the pack. The value at\n+position `i` is the hash of the pathname at which the `i`th object\n+(counting in index order) in the pack can be found.  This can be fed\n+into the delta heuristics to compare objects with similar pathnames.\n+\n+The hash algorithm used is:\n+\n+    hash = 0;\n+    while ((c = *name++))\n+\t    if (!isspace(c))\n+\t\t    hash = (hash >> 2) + (c << 24);\n+\n+Note that this hashing scheme is tied to the BITMAP_OPT_HASH_CACHE flag.\n+If implementations want to choose a different hashing scheme, they are\n+free to do so, but MUST allocate a new header flag (because comparing\n+hashes made under two different schemes would be pointless).\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 26646e7..4504789 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -64,6 +64,7 @@ static off_t reuse_packfile_offset;\n \n static int use_bitmap_index = 1;\n static int write_bitmap_index;\n+static uint16_t write_bitmap_options;\n \n static unsigned long delta_cache_size = 0;\n static unsigned long max_delta_cache_size = 256 * 1024 * 1024;\n@@ -862,7 +863,8 @@ static void write_pack_file(void)\n \t\t\t\tbitmap_writer_reuse_bitmaps(&to_pack);\n \t\t\t\tbitmap_writer_select_commits(indexed_commits, indexed_commits_nr, -1);\n \t\t\t\tbitmap_writer_build(&to_pack);\n-\t\t\t\tbitmap_writer_finish(written_list, nr_written, tmpname);\n+\t\t\t\tbitmap_writer_finish(written_list, nr_written,\n+\t\t\t\t\t\t     tmpname, write_bitmap_options);\n \t\t\t\twrite_bitmap_index = 0;\n \t\t\t}\n \n@@ -2155,6 +2157,12 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n \t\twrite_bitmap_index = git_config_bool(k, v);\n \t\treturn 0;\n \t}\n+\tif (!strcmp(k, \"pack.writebitmaphashcache\")) {\n+\t\tif (git_config_bool(k, v))\n+\t\t\twrite_bitmap_options |= BITMAP_OPT_HASH_CACHE;\n+\t\telse\n+\t\t\twrite_bitmap_options &= ~BITMAP_OPT_HASH_CACHE;\n+\t}\n \tif (!strcmp(k, \"pack.usebitmaps\")) {\n \t\tuse_bitmap_index = git_config_bool(k, v);\n \t\treturn 0;\ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nindex 954a74d..1218bef 100644\n--- a/pack-bitmap-write.c\n+++ b/pack-bitmap-write.c\n@@ -490,6 +490,19 @@ static void write_selected_commits_v1(struct sha1file *f,\n \t}\n }\n \n+static void write_hash_cache(struct sha1file *f,\n+\t\t\t     struct pack_idx_entry **index,\n+\t\t\t     uint32_t index_nr)\n+{\n+\tuint32_t i;\n+\n+\tfor (i = 0; i < index_nr; ++i) {\n+\t\tstruct object_entry *entry = (struct object_entry *)index[i];\n+\t\tuint32_t hash_value = htonl(entry->hash);\n+\t\tsha1write(f, &hash_value, sizeof(hash_value));\n+\t}\n+}\n+\n void bitmap_writer_set_checksum(unsigned char *sha1)\n {\n \thashcpy(writer.pack_checksum, sha1);\n@@ -497,7 +510,8 @@ void bitmap_writer_set_checksum(unsigned char *sha1)\n \n void bitmap_writer_finish(struct pack_idx_entry **index,\n \t\t\t  uint32_t index_nr,\n-\t\t\t  const char *filename)\n+\t\t\t  const char *filename,\n+\t\t\t  uint16_t options)\n {\n \tstatic char tmp_file[PATH_MAX];\n \tstatic uint16_t default_version = 1;\n@@ -514,7 +528,7 @@ void bitmap_writer_finish(struct pack_idx_entry **index,\n \n \tmemcpy(header.magic, BITMAP_IDX_SIGNATURE, sizeof(BITMAP_IDX_SIGNATURE));\n \theader.version = htons(default_version);\n-\theader.options = htons(flags);\n+\theader.options = htons(flags | options);\n \theader.entry_count = htonl(writer.selected_nr);\n \tmemcpy(header.checksum, writer.pack_checksum, 20);\n \n@@ -525,6 +539,9 @@ void bitmap_writer_finish(struct pack_idx_entry **index,\n \tdump_bitmap(f, writer.tags);\n \twrite_selected_commits_v1(f, index, index_nr);\n \n+\tif (options & BITMAP_OPT_HASH_CACHE)\n+\t\twrite_hash_cache(f, index, index_nr);\n+\n \tsha1close(f, NULL, CSUM_FSYNC);\n \n \tif (adjust_shared_perm(tmp_file))\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nindex 7792dfd..078f7c6 100644\n--- a/pack-bitmap.c\n+++ b/pack-bitmap.c\n@@ -66,6 +66,9 @@ static struct bitmap_index {\n \t/* Number of bitmapped commits */\n \tuint32_t entry_count;\n \n+\t/* Name-hash cache (or NULL if not present). */\n+\tuint32_t *hashes;\n+\n \t/*\n \t * Extended index.\n \t *\n@@ -152,6 +155,11 @@ static int load_bitmap_header(struct bitmap_index *index)\n \t\tif ((flags & BITMAP_OPT_FULL_DAG) == 0)\n \t\t\treturn error(\"Unsupported options for bitmap index file \"\n \t\t\t\t\"(Git requires BITMAP_OPT_FULL_DAG)\");\n+\n+\t\tif (flags & BITMAP_OPT_HASH_CACHE) {\n+\t\t\tunsigned char *end = index->map + index->map_size - 20;\n+\t\t\tindex->hashes = ((uint32_t *)end) - index->pack->num_objects;\n+\t\t}\n \t}\n \n \tindex->entry_count = ntohl(header->entry_count);\n@@ -626,6 +634,9 @@ static void show_objects_for_type(\n \t\t\tentry = &bitmap_git.reverse_index->revindex[pos + offset];\n \t\t\tsha1 = nth_packed_object_sha1(bitmap_git.pack, entry->nr);\n \n+\t\t\tif (bitmap_git.hashes)\n+\t\t\t\thash = ntohl(bitmap_git.hashes[entry->nr]);\n+\n \t\t\tshow_reach(sha1, object_type, 0, hash, bitmap_git.pack, entry->offset);\n \t\t}\n \ndiff --git a/pack-bitmap.h b/pack-bitmap.h\nindex 5e3fb9a..e4e1a57 100644\n--- a/pack-bitmap.h\n+++ b/pack-bitmap.h\n@@ -25,6 +25,7 @@ static const char BITMAP_IDX_SIGNATURE[] = {'B', 'I', 'T', 'M'};;\n \n enum pack_bitmap_opts {\n \tBITMAP_OPT_FULL_DAG = 1,\n+\tBITMAP_OPT_HASH_CACHE = 4,\n };\n \n enum pack_bitmap_flags {\n@@ -57,6 +58,7 @@ void bitmap_writer_select_commits(struct commit **indexed_commits,\n void bitmap_writer_build(struct packing_data *to_pack);\n void bitmap_writer_finish(struct pack_idx_entry **index,\n \t\t\t  uint32_t index_nr,\n-\t\t\t  const char *filename);\n+\t\t\t  const char *filename,\n+\t\t\t  uint16_t options);\n \n #endif\ndiff --git a/t/perf/p5310-pack-bitmaps.sh b/t/perf/p5310-pack-bitmaps.sh\nindex ac036d0..8811fc4 100755\n--- a/t/perf/p5310-pack-bitmaps.sh\n+++ b/t/perf/p5310-pack-bitmaps.sh\n@@ -9,7 +9,8 @@ test_perf_large_repo\n # since we want to be able to compare bitmap-aware\n # git versus non-bitmap git\n test_expect_success 'setup bitmap config' '\n-\tgit config pack.writebitmaps true\n+\tgit config pack.writebitmaps true &&\n+\tgit config pack.writebitmaphashcache true\n '\n \n test_perf 'repack to disk' '\ndiff --git a/t/t5310-pack-bitmaps.sh b/t/t5310-pack-bitmaps.sh\nindex bd743f4..2c2632f 100755\n--- a/t/t5310-pack-bitmaps.sh\n+++ b/t/t5310-pack-bitmaps.sh\n@@ -14,7 +14,8 @@ test_expect_success 'setup repo with moderate-sized history' '\n \tgit checkout master &&\n \tblob=$(echo tagged-blob | git hash-object -w --stdin) &&\n \tgit tag tagged-blob $blob &&\n-\tgit config pack.writebitmaps true\n+\tgit config pack.writebitmaps true &&\n+\tgit config pack.writebitmaphashcache true\n '\n \n test_expect_success 'full repack creates bitmaps' '\n-- \n1.8.5.rc0.443.g2df7f3f\n"},{"id":"230640","messageId":"5285224A.2070606@ramsay1.demon.co.uk","threadId":"35335","inReplyTo":"20131114124157.GA23784@sigill.intra.peff.net","subject":"Re: [PATCH v3 0/21] pack bitmaps","fromName":"Ramsay Jones","fromEmail":"ramsay@ramsay1.demon.co.uk","sentAt":"2013-11-14T19:19:38Z","receivedAt":"2013-11-14T19:19:38Z","isPatch":true,"sender":{"key":"ramsay@ramsayjones.plus.com","avatar":"https://avatars.githubusercontent.com/u/33702710?v=4"},"body":"On 14/11/13 12:41, Jeff King wrote:\n> Here's another iteration of the pack bitmaps series. Compared to v2, it\n> changes:\n> \n>  - misc style/typo fixes\n> \n>  - portability fixes from Ramsay and Torsten\n\nUnfortunately, I didn't find time this weekend to finish the msvc build\nfixes. However, after a quick squint at these patches, I think you have\nalmost done it for me! :-D\n\nI must have misunderstood the previous discussion, because my patch was\nwritten on the assumption that the ewah directory wouldn't be \"git-ified\"\n(e.g. #include git-compat-util.h).\n\nSo, most of my patch is no longer necessary, given the use of the git\ncompat header (and removal of system headers). I suspect that you only\nneed to add an '#define PRIx64 \"I64x\"' definition (Hmm, probably to the\ncompat/mingw.h header).\n\nI won't know for sure until I actually try them out, of course. I will\nwait until these patches land in pu.\n\n[Note: the msvc build is still broken, but the failure is not caused by\nthese patches. Unfortunately, the tests in t5310-*.sh fail. However, if\nI include some debug code, the tests pass ... :-P ]\n\nThe part of the patch I was still working on was ...\n\n> \n>  - count-objects garbage-reporting patch from Duy\n> \n>  - disable bitmaps when is_repository_shallow(); this also covers the\n>    case where the client is shallow, since we feed pack-objects a\n>    --shallow-file in that case. This used to done by checking\n>    !internal_rev_list, but that doesn't apply after cdab485.\n> \n>  - ewah sources now properly use git-compat-util.h and do not include\n>    system headers\n> \n>  - the ewah code uses ewah_malloc, ewah_realloc, and so forth to let the\n>    project use a particular allocator (and we want to use xmalloc and\n>    friends). And we defined those in pack-bitmap.h, but of course that\n>    had no effect on the ewah/*.c files that did not include\n>    pack-bitmap.h.  Since we are hacking up and git-ifying libewok\n>    anyway, we can just set the hardcoded fallback to xmalloc instead of\n>    malloc.\n> \n>   - the ewah code used gcc's __builtin_ctzll, but did not provide a\n>     suitable fallback. We now provide a fallback in C.\n\n... here.\n\nI was messing around with several implementations (including the use of\nmsvc compiler intrinsics) with the intention of doing some timing tests\netc. [I suspected my C fallback function (a different implementation to\nyours) would be slightly faster.]\n\nATB,\nRamsay Jones\n"},{"id":"230652","messageId":"20131114213320.GA16466@sigill.intra.peff.net","threadId":"35335","inReplyTo":"5285224A.2070606@ramsay1.demon.co.uk","subject":"Re: [PATCH v3 0/21] pack bitmaps","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-14T21:33:20Z","receivedAt":"2013-11-14T21:33:20Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Nov 14, 2013 at 07:19:38PM +0000, Ramsay Jones wrote:\n\n> Unfortunately, I didn't find time this weekend to finish the msvc build\n> fixes. However, after a quick squint at these patches, I think you have\n> almost done it for me! :-D\n> \n> I must have misunderstood the previous discussion, because my patch was\n> written on the assumption that the ewah directory wouldn't be \"git-ified\"\n> (e.g. #include git-compat-util.h).\n\nI think it was up for debate at some point, but we did decide to go\nahead and git-ify. Please feel free to submit further fixups if you need\nthem.\n\n> >   - the ewah code used gcc's __builtin_ctzll, but did not provide a\n> >     suitable fallback. We now provide a fallback in C.\n> \n> ... here.\n> \n> I was messing around with several implementations (including the use of\n> msvc compiler intrinsics) with the intention of doing some timing tests\n> etc. [I suspected my C fallback function (a different implementation to\n> yours) would be slightly faster.]\n\nYeah, I looked around for several implementations, and ultimately wrote\none that was the most readable to me. The one I found shortest and most\ninscrutable was:\n\n  return popcount((x & -x) - 1);\n\nthe details of which I still haven't worked through in my head. ;)\n\nI do think on most platforms that intrinsics or inline assembler are the\nway to go. My main goal was to get something correct that would let it\ncompile everywhere, and then people can use that as a base for\noptimizing. Patches welcome. :)\n\n-Peff\n"},{"id":"230656","messageId":"52855834.3080608@ramsay1.demon.co.uk","threadId":"35335","inReplyTo":"20131114213320.GA16466@sigill.intra.peff.net","subject":"Re: [PATCH v3 0/21] pack bitmaps","fromName":"Ramsay Jones","fromEmail":"ramsay@ramsay1.demon.co.uk","sentAt":"2013-11-14T23:09:40Z","receivedAt":"2013-11-14T23:09:40Z","isPatch":true,"sender":{"key":"ramsay@ramsayjones.plus.com","avatar":"https://avatars.githubusercontent.com/u/33702710?v=4"},"body":"On 14/11/13 21:33, Jeff King wrote:\n> On Thu, Nov 14, 2013 at 07:19:38PM +0000, Ramsay Jones wrote:\n> \n>> Unfortunately, I didn't find time this weekend to finish the msvc build\n>> fixes. However, after a quick squint at these patches, I think you have\n>> almost done it for me! :-D\n>>\n>> I must have misunderstood the previous discussion, because my patch was\n>> written on the assumption that the ewah directory wouldn't be \"git-ified\"\n>> (e.g. #include git-compat-util.h).\n> \n> I think it was up for debate at some point, but we did decide to go\n> ahead and git-ify. Please feel free to submit further fixups if you need\n> them.\n\nYep, will do; at present it looks like that one-liner.\n\n> \n>>>   - the ewah code used gcc's __builtin_ctzll, but did not provide a\n>>>     suitable fallback. We now provide a fallback in C.\n>>\n>> ... here.\n>>\n>> I was messing around with several implementations (including the use of\n>> msvc compiler intrinsics) with the intention of doing some timing tests\n>> etc. [I suspected my C fallback function (a different implementation to\n>> yours) would be slightly faster.]\n> \n> Yeah, I looked around for several implementations, and ultimately wrote\n> one that was the most readable to me. The one I found shortest and most\n> inscrutable was:\n> \n>   return popcount((x & -x) - 1);\n> \n> the details of which I still haven't worked through in my head. ;)\n\nYeah, I stumbled over that one too!\n\n> I do think on most platforms that intrinsics or inline assembler are the\n> way to go. My main goal was to get something correct that would let it\n> compile everywhere, and then people can use that as a base for\n> optimizing. Patches welcome. :)\n\nIndeed, I can happily leave that to another day (or to someone else\nmore motivated ;-)\n\nATB,\nRamsay Jones\n"},{"id":"230699","messageId":"87k3g8ljxv.fsf@linux-k42r.v.cablecom.net","threadId":"35335","inReplyTo":"20131114213320.GA16466@sigill.intra.peff.net","subject":"Re: [PATCH v3 0/21] pack bitmaps","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-11-16T10:28:28Z","receivedAt":"2013-11-16T10:28:28Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>> >   - the ewah code used gcc's __builtin_ctzll, but did not provide a\n>> >     suitable fallback. We now provide a fallback in C.\n>> \n>> I was messing around with several implementations (including the use of\n>> msvc compiler intrinsics) with the intention of doing some timing tests\n>> etc. [I suspected my C fallback function (a different implementation to\n>> yours) would be slightly faster.]\n>\n> Yeah, I looked around for several implementations, and ultimately wrote\n> one that was the most readable to me. The one I found shortest and most\n> inscrutable was:\n>\n>   return popcount((x & -x) - 1);\n\nIn two's complement, -x = ~x + 1 [1].  If you have a bunch of 0s at the\nend, as in (binary; a=~A etc)\n\n           x = abcdef1000\n\nthen\n\n          ~x = ABCDEF0111\n ~x + 1 = -x = ABCDEF1000\n\n      (x&-x) = 0000001000\n  (x&-x) - 1 = 0000000111\n\npopcount() of that is the number of trailing zeroes you started with.\n\nPlease don't ask me to work out what happens in border cases; my head\nhurts already.\n\n\n[1] because x + ~x is all one bits.  +1 makes it overflow to 0, so that\nx + -x = 0 as it should.\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"230781","messageId":"528A839F.10804@ramsay1.demon.co.uk","threadId":"35335","inReplyTo":"52855834.3080608@ramsay1.demon.co.uk","subject":"Re: [PATCH v3 0/21] pack bitmaps","fromName":"Ramsay Jones","fromEmail":"ramsay@ramsay1.demon.co.uk","sentAt":"2013-11-18T21:16:15Z","receivedAt":"2013-11-18T21:16:15Z","isPatch":true,"sender":{"key":"ramsay@ramsayjones.plus.com","avatar":"https://avatars.githubusercontent.com/u/33702710?v=4"},"body":"On 14/11/13 23:09, Ramsay Jones wrote:\n> On 14/11/13 21:33, Jeff King wrote:\n>> On Thu, Nov 14, 2013 at 07:19:38PM +0000, Ramsay Jones wrote:\n>>\n>>> Unfortunately, I didn't find time this weekend to finish the msvc build\n>>> fixes. However, after a quick squint at these patches, I think you have\n>>> almost done it for me! :-D\n>>>\n>>> I must have misunderstood the previous discussion, because my patch was\n>>> written on the assumption that the ewah directory wouldn't be \"git-ified\"\n>>> (e.g. #include git-compat-util.h).\n>>\n>> I think it was up for debate at some point, but we did decide to go\n>> ahead and git-ify. Please feel free to submit further fixups if you need\n>> them.\n> \n> Yep, will do; at present it looks like that one-liner.\n\nDespite saying I would wait for these patches to land in pu, I applied these\npatches to the next branch (@ 8721652 \"Sync with 1.8.5-rc2\") so that I could\ntry them out.\n\nAs expected, everything was fine on Linux and cygwin, and the msvc build needed\na one-liner fix. However, I didn't expect that MinGW would need the same fix! ;-)\n\nI had forgotten that \"git-compat-util.h\" does not include the <inttypes.h> header\non MinGW (my previous patch included that header directly), so that we have to\nadd the printf format macros directly to the compat/mingw.h header.\n\n[I assume that an earlier version of the library on MinGW did not have the\n<inttypes.h> and <stdint.h> headers. We could now include that header on MinGW\nand move the PRIuMAX, PRId64 and PRIx64 macros to compat/msvc.h, but I didn't\nthink it was worth the churn ... ]\n\nThe one-liner is given below ...\n\nATB,\nRamsay Jones\n\n-- >8 --\nSubject: [PATCH] compat/mingw.h: Fix the MinGW and msvc builds\n\nSigned-off-by: Ramsay Jones <ramsay@ramsay1.demon.co.uk>\n---\n compat/mingw.h | 1 +\n 1 file changed, 1 insertion(+)\n\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex 92cd728..8828ede 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -345,6 +345,7 @@ static inline char *mingw_find_last_dir_sep(const char *path)\n #define PATH_SEP ';'\n #define PRIuMAX \"I64u\"\n #define PRId64 \"I64d\"\n+#define PRIx64 \"I64x\"\n \n void mingw_open_html(const char *path);\n #define open_html mingw_open_html\n-- \n1.8.4\n"},{"id":"231023","messageId":"87fvqlfpmw.fsf@linux-k42r.v.cablecom.net","threadId":"35335","inReplyTo":"20131114124432.GJ10757@sigill.intra.peff.net","subject":"Re: [PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-11-24T21:36:55Z","receivedAt":"2013-11-24T21:36:55Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>  khash.h       | 338 ++++++++++++++++++++\n[...]\n> diff --git a/khash.h b/khash.h\n> new file mode 100644\n> index 0000000..57ff603\n> --- /dev/null\n> +++ b/khash.h\n> @@ -0,0 +1,338 @@\n> +/* The MIT License\n> +\n> +   Copyright (c) 2008, 2009, 2011 by Attractive Chaos <attractor@live.co.uk>\n> +\n> +   Permission is hereby granted, free of charge, to any person obtaining\n> +   a copy of this software and associated documentation files (the\n> +   \"Software\"), to deal in the Software without restriction, including\n> +   without limitation the rights to use, copy, modify, merge, publish,\n> +   distribute, sublicense, and/or sell copies of the Software, and to\n> +   permit persons to whom the Software is furnished to do so, subject to\n> +   the following conditions:\n> +\n> +   The above copyright notice and this permission notice shall be\n> +   included in all copies or substantial portions of the Software.\n> +\n> +   THE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND,\n> +   EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF\n> +   MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND\n> +   NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS\n> +   BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN\n> +   ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN\n> +   CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\n> +   SOFTWARE.\n> +*/\n> +\n> +#ifndef __AC_KHASH_H\n> +#define __AC_KHASH_H\n[...]\n> +static inline khint_t __kh_oid_hash(const unsigned char *oid)\n> +{\n> +\tkhint_t hash;\n> +\tmemcpy(&hash, oid, sizeof(hash));\n> +\treturn hash;\n> +}\n> +\n> +#define __kh_oid_cmp(a, b) (hashcmp(a, b) == 0)\n> +\n> +KHASH_INIT(sha1, const unsigned char *, void *, 1, __kh_oid_hash, __kh_oid_cmp)\n> +typedef kh_sha1_t khash_sha1;\n> +\n> +KHASH_INIT(sha1_pos, const unsigned char *, int, 1, __kh_oid_hash, __kh_oid_cmp)\n> +typedef kh_sha1_pos_t khash_sha1_pos;\n> +\n> +#endif /* __AC_KHASH_H */\n\nAFAICS, the part after the [...] are additions specific to git.  Is that\nright?\n\nCan we store them in a separate khash-sha1.h or some such, to make it\nclearer what's what?  As things stand, one has to look for an identifier\nthat is built from macros at the far end of a file that looks like it\nwas imported verbatim from klib(?).  I don't know about you, but that\njust took me far too long to get right.\n\nI think I'll also lend you a hand writing Documentation/technical/api-khash.txt\n(expect it tomorrow) so that we also have documentation in the git\nstyle, where gitters can be expected to find it on their own.\n\nAll that could then nicely fit into a commit that actually says where\nyou conjured khash.h from and such.  This information seems\nconspicuously absent from this commit's message :-)\n\nFurthermore, would it be a problem to name the second hash sha1_int\ninstead?  I have another use for such a hash, and I can't imagine I'm\nthe only one.  (That's not critical however, I can do the required\nediting in that other series.)\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231069","messageId":"ccce08818b6856f67a5316eba148d5a860d1d8f7.1385391631.git.tr@thomasrast.ch","threadId":"35335","inReplyTo":"87fvqlfpmw.fsf@linux-k42r.v.cablecom.net","subject":"[PATCH] Document khash","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-11-25T15:04:48Z","receivedAt":"2013-11-25T15:04:48Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"For squashing into a commit that adds khash.h.\n\nSigned-off-by: Thomas Rast <tr@thomasrast.ch>\n---\n\n> I think I'll also lend you a hand writing Documentation/technical/api-khash.txt\n> (expect it tomorrow) so that we also have documentation in the git\n> style, where gitters can be expected to find it on their own.\n\nHere goes.\n\n> Furthermore, would it be a problem to name the second hash sha1_int\n> instead?  I have another use for such a hash, and I can't imagine I'm\n> the only one.  (That's not critical however, I can do the required\n> editing in that other series.)\n\nActually, let's not do that.  Since everything is 'static inline'\nanyway, there's no cost to simply instantiating more hashes as needed.\n\n\n Documentation/technical/api-khash.txt | 109 ++++++++++++++++++++++++++++++++++\n 1 file changed, 109 insertions(+)\n create mode 100644 Documentation/technical/api-khash.txt\n\ndiff --git a/Documentation/technical/api-khash.txt b/Documentation/technical/api-khash.txt\nnew file mode 100644\nindex 0000000..e7ca883\n--- /dev/null\n+++ b/Documentation/technical/api-khash.txt\n@@ -0,0 +1,109 @@\n+khash\n+=====\n+\n+The khash API is a collection of macros that instantiate hash table\n+routines for arbitrary key/value types.\n+\n+It lives in 'khash.h', which we imported from\n+  https://github.com/attractivechaos/klib/blob/master/khash.h\n+and then gitified somewhat.\n+\n+Usage\n+-----\n+\n+------------\n+#import \"khash.h\"\n+KHASH_INIT(NAME, key_t, value_t, is_map, key_hash_fn, key_equal_fn)\n+------------\n+\n+The arguments are as follows:\n+\n+`NAME`::\n+\tUsed to give your hash and its API a unique name.  Spelled\n+\t`NAME` here to set it apart from the rest of the names, but\n+\tgenerally should be a lowercase C identifier.\n+\n+`key_t`::\n+\tType of the keys for your hashes.  Generally should be\n+\t`const`.\n+\n+`value_t`::\n+\tType of the values of your hashes.\n+\n+`is_map`::\n+\tWhether the hash holds values.  Set to 0 to create a hash for\n+\tuse as a set.\n+\n+`khint_t key_hash_fn(key_t)`::\n+\tHash function.\n+\n+`int key_equal_fn(key_t a, key_t b)`::\n+\tComparison function.  Return 1 if the two keys are the same, 0\n+\totherwise.\n+\n+These two functions may also be macros.\n+\n+\n+API\n+---\n+\n+The above instantiation defines a single type:\n+\n+`kh_NAME_t`::\n+\tA struct that holds your hash.  You should only ever use\n+\tpointers to it.\n+\n+After the above instantiation, the following functions and\n+function-like macros are defined:\n+\n+`kh_NAME_t *kh_init_NAME(void)`::\n+\tAllocate and initialize a new hash table.\n+\n+`void kh_destroy_NAME(kh_NAME_t *)`::\n+\tFree the hash table.\n+\n+`void kh_clear_NAME(kh_NAME_t *)`::\n+\tClear (but do not free) the hash table.\n+\n+`khint_t kh_get_NAME(const kh_NAME_t *hash, key_t key)`::\n+\tFind the given key in the hash table.  The returned khint_t\n+\tshould be treated as an opaque iterator token that indexes\n+\tinto the hash.  Use `kh_value` to get the value from the\n+\ttoken.  If the key does not exist, returns kh_end(hash).\n+\n+`khint_t kh_put_NAME(const kh_NAME_t *hash, key_t key, int *ret)`::\n+\tPut the given key in the hash table.  The returned khint_t\n+\tshould be treated as an opaque iterator token that indexes\n+\tinto the hash.  Use `kh_value(hash, token) = value` to assign\n+\ta value after putting the key.\n++\n+'ret' tells you whether the key already existed: it is -1 if the\n+operation failed; 0 if the key was already present; 1 if the bucket is\n+empty; 2 if the bucket has been deleted.\n+\n+`kh_key(kh_NAME_t *, khint_t)`::\n+\tThe key slot for the given iterator token.  This can be used\n+\tas an lvalue, but you must not change the key!\n+\n+`kh_value(kh_NAME_t *, khint_t)`::\n+`kh_val(kh_NAME_t *, khint_t)`::\n+\tThe value slot for the given iterator token.  This can be used\n+\tas an lvalue.\n+\n+`int kh_exist(const kh_NAME_t *, key_t)`::\n+\tTests if the given bucket contains data.\n+\n+`khint_t kh_begin(const kh_NAME_t *)`::\n+\tThe smallest iterator token.\n+\n+`khint_t kh_end(const kh_NAME_t *)`::\n+\tThe beyond-the-hash iterator token; it is one larger than the\n+\tlargest token that points to a hash member.\n+\n+`khint_t kh_size(const kh_NAME_t *)`::\n+\tThe number of elements in the hash.\n+\n+`kh_foreach(const kh_NAME_t *, keyvar, valuevar, { ... loop body ... })`::\n+\tIterate over every key-value pair in the hash.  This is a\n+\tmacro, and keyvar and valuevar must be pre-declared of type\n+\tkey_t and value_t, respectively.\n-- \n1.8.5.rc3.397.g2a3acd5\n"},{"id":"231160","messageId":"5295B6A8.70303@gmail.com","threadId":"35335","inReplyTo":"87fvqlfpmw.fsf@linux-k42r.v.cablecom.net","subject":"Re: [PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2013-11-27T09:08:56Z","receivedAt":"2013-11-27T09:08:56Z","isPatch":true,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 24.11.2013 22:36, schrieb Thomas Rast:\n> Jeff King <peff@peff.net> writes:\n[...]\n> \n> I think I'll also lend you a hand writing Documentation/technical/api-khash.txt\n> (expect it tomorrow) so that we also have documentation in the git\n> style, where gitters can be expected to find it on their own.\n> \n\nKhash is OK for sha1 keys, but I don't think it should be advertised as a second general purpose hash table implementation. Its far too easy to shoot yourself in the foot by using 'straightforward' hash- and comparison functions. Khash doesn't store the hash codes of the keys, so you have to take care of that yourself or live with the performance penalties (see [1]).\n\n[1] http://article.gmane.org/gmane.comp.version-control.git/237876\n"},{"id":"231239","messageId":"20131128103555.GA14615@sigill.intra.peff.net","threadId":"35335","inReplyTo":"ccce08818b6856f67a5316eba148d5a860d1d8f7.1385391631.git.tr@thomasrast.ch","subject":"Re: [PATCH] Document khash","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-28T10:35:55Z","receivedAt":"2013-11-28T10:35:55Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Nov 25, 2013 at 04:04:48PM +0100, Thomas Rast wrote:\n\n> > I think I'll also lend you a hand writing Documentation/technical/api-khash.txt\n> > (expect it tomorrow) so that we also have documentation in the git\n> > style, where gitters can be expected to find it on their own.\n> \n> Here goes.\n\nThanks. Some comments below.\n\n> > Furthermore, would it be a problem to name the second hash sha1_int\n> > instead?  I have another use for such a hash, and I can't imagine I'm\n> > the only one.  (That's not critical however, I can do the required\n> > editing in that other series.)\n> \n> Actually, let's not do that.  Since everything is 'static inline'\n> anyway, there's no cost to simply instantiating more hashes as needed.\n\nYeah. If we cared about code duplication, we'd have to split it into\ndeclaration and definition macros, and then collect the definitions in\none .c file (khash does support this, but we don't use it). I don't\nthink we are using enough repeated ones to make it matter, and most of\nthe functions benefit from inlining anyway.\n\n> +------------\n> +#import \"khash.h\"\n> +KHASH_INIT(NAME, key_t, value_t, is_map, key_hash_fn, key_equal_fn)\n> +------------\n\n#import?\n\n> +The arguments are as follows:\n> [...]\n> +`khint_t key_hash_fn(key_t)`::\n> +\tHash function.\n\nIt is true that this is a khint_t, but I do not think knowing that helps\nthe caller. They need to design a hash function, so knowing the size and\ntype of the hint helps. Maybe:\n\n  Return a hash value for a key. The khint_t is a 32-bit unsigned value.\n  Git provides __kh_oid_hash, which converts a sha1 into a hash value.\n\n> +`int key_equal_fn(key_t a, key_t b)`::\n> +\tComparison function.  Return 1 if the two keys are the same, 0\n> +\totherwise.\n\nHere we provide __kh_oid_cmp, and we should mention it. Based on recent\ndiscussions, this should probably be __kh_oid_equal, since it is not a\n\"cmp\" function in the ordering sense.\n\n> +`khint_t kh_get_NAME(const kh_NAME_t *hash, key_t key)`::\n> +\tFind the given key in the hash table.  The returned khint_t\n> +\tshould be treated as an opaque iterator token that indexes\n> +\tinto the hash.  Use `kh_value` to get the value from the\n> +\ttoken.  If the key does not exist, returns kh_end(hash).\n\nI do not know of the khash author's intention, but in the bitmap code we\nprefer khiter_t to represent an iterator. It is the same as a khint_t,\nbut I think that is an implementation detail (and the khash functions\nshould probably be returning a khiter_t).\n\n> [...]\n\nThe rest of it looks correct to me.\n\n-Peff\n"},{"id":"231240","messageId":"20131128103838.GB14615@sigill.intra.peff.net","threadId":"35335","inReplyTo":"5295B6A8.70303@gmail.com","subject":"Re: [PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-28T10:38:38Z","receivedAt":"2013-11-28T10:38:38Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Nov 27, 2013 at 10:08:56AM +0100, Karsten Blees wrote:\n\n> Khash is OK for sha1 keys, but I don't think it should be advertised\n> as a second general purpose hash table implementation. Its far too\n> easy to shoot yourself in the foot by using 'straightforward' hash-\n> and comparison functions. Khash doesn't store the hash codes of the\n> keys, so you have to take care of that yourself or live with the\n> performance penalties (see [1]).\n> \n> [1] http://article.gmane.org/gmane.comp.version-control.git/237876\n\nYes. I wonder if we should improve it in that respect. I haven't looked\ncarefully at the hash code you posted elsewhere, but I feel like many\nuses will want a macro implementation to let them store arbitrary types\nsmaller or larger than a pointer.\n\n-Peff\n"},{"id":"231281","messageId":"87siuedhvj.fsf@thomasrast.ch","threadId":"35335","inReplyTo":"20131114124432.GJ10757@sigill.intra.peff.net","subject":"Re: [PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-11-29T21:21:04Z","receivedAt":"2013-11-29T21:21:04Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"TLDR: nitpicks.  Thanks for a very nice read.\n\nI do think it's worth fixing the syntax pedantry at the end so that we\ncan keep supporting arcane compilers, but otherwise, meh.\n\n> +static int open_pack_bitmap_1(struct packed_git *packfile)\n\nThis goes somewhat against the naming convention (if you can call it\nthat) used elsewhere in git.  Usually foo_1() is an implementation\ndetail of foo(), used because it is convenient to wrap the main part in\nanother function, e.g. so that it can consistently free resources or\nsome such.  But this one operates on one pack file, so in the terms of\nthe rest of git, it should probably be called open_pack_bitmap_one().\n\n> +static void show_object(struct object *object, const struct name_path *path,\n> +\t\t\tconst char *last, void *data)\n> +{\n> +\tstruct bitmap *base = data;\n> +\tint bitmap_pos;\n> +\n> +\tbitmap_pos = bitmap_position(object->sha1);\n> +\n> +\tif (bitmap_pos < 0) {\n> +\t\tchar *name = path_name(path, last);\n> +\t\tbitmap_pos = ext_index_add_object(object, name);\n> +\t\tfree(name);\n> +\t}\n> +\n> +\tbitmap_set(base, bitmap_pos);\n> +}\n> +\n> +static void show_commit(struct commit *commit, void *data)\n> +{\n> +}\n\nA bit unfortunate that you inherit the strange show_* naming from\nbuiltin/pack-objects.c, which seems to have stolen some code from\nbuiltin/rev-list.c at some point without worrying about better naming...\n\n> +static void show_objects_for_type(\n> +\tstruct bitmap *objects,\n> +\tstruct ewah_bitmap *type_filter,\n> +\tenum object_type object_type,\n> +\tshow_reachable_fn show_reach)\n> +{\n[...]\n> +\twhile (i < objects->word_alloc && ewah_iterator_next(&filter, &it)) {\n> +\t\teword_t word = objects->words[i] & filter;\n> +\n> +\t\tfor (offset = 0; offset < BITS_IN_WORD; ++offset) {\n> +\t\t\tconst unsigned char *sha1;\n> +\t\t\tstruct revindex_entry *entry;\n> +\t\t\tuint32_t hash = 0;\n> +\n> +\t\t\tif ((word >> offset) == 0)\n> +\t\t\t\tbreak;\n> +\n> +\t\t\toffset += ewah_bit_ctz64(word >> offset);\n> +\n> +\t\t\tif (pos + offset < bitmap_git.reuse_objects)\n> +\t\t\t\tcontinue;\n> +\n> +\t\t\tentry = &bitmap_git.reverse_index->revindex[pos + offset];\n> +\t\t\tsha1 = nth_packed_object_sha1(bitmap_git.pack, entry->nr);\n> +\n> +\t\t\tshow_reach(sha1, object_type, 0, hash, bitmap_git.pack, entry->offset);\n> +\t\t}\n\nYou have a very nice bitmap_each_bit() function in ewah/bitmap.c, why\nnot use it here?\n\n> +int reuse_partial_packfile_from_bitmap(struct packed_git **packfile,\n> +\t\t\t\t       uint32_t *entries,\n> +\t\t\t\t       off_t *up_to)\n> +{\n> +\t/*\n> +\t * Reuse the packfile content if we need more than\n> +\t * 90% of its objects\n> +\t */\n> +\tstatic const double REUSE_PERCENT = 0.9;\n\nCurious: is this based on some measurements or just a guess?\n\n\n> diff --git a/pack-bitmap.h b/pack-bitmap.h\n[...]\n> +static const char BITMAP_IDX_SIGNATURE[] = {'B', 'I', 'T', 'M'};;\n\nThere's a stray ; at the end of the line that is technically not\npermitted:\n\npack-bitmap.h:22:65: warning: ISO C does not allow extra ‘;’ outside of a function [-Wpedantic]\n\n> +enum pack_bitmap_opts {\n> +\tBITMAP_OPT_FULL_DAG = 1,\n\nAnd I think this trailing comma on the last enum item is also strictly\nspeaking not allowed, even though it is very nice to have:\n\npack-bitmap.h:28:27: warning: comma at end of enumerator list [-Wpedantic]\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231374","messageId":"20131202161208.GB24202@sigill.intra.peff.net","threadId":"35335","inReplyTo":"87siuedhvj.fsf@thomasrast.ch","subject":"Re: [PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-02T16:12:08Z","receivedAt":"2013-12-02T16:12:08Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Nov 29, 2013 at 10:21:04PM +0100, Thomas Rast wrote:\n\n> I do think it's worth fixing the syntax pedantry at the end so that we\n> can keep supporting arcane compilers, but otherwise, meh.\n\nAgreed. I've picked up those changes in my tree.\n\n> > +static int open_pack_bitmap_1(struct packed_git *packfile)\n> \n> This goes somewhat against the naming convention (if you can call it\n> that) used elsewhere in git.  Usually foo_1() is an implementation\n> detail of foo(), used because it is convenient to wrap the main part in\n> another function, e.g. so that it can consistently free resources or\n> some such.  But this one operates on one pack file, so in the terms of\n> the rest of git, it should probably be called open_pack_bitmap_one().\n\nHmm. I see your point, but I think that my (and Vicent's) mental model\nwas that is _was_ a helper for open_pack_bitmap. It just happens to also\nfill the role of open_pack_bitmap_one(), but you would not want the\nlatter. We only support a single bitmap at a time; by calling the\nhelper, you would miss out on the assert which would catch the error.\n\nSo I don't care much, but I have a slight preference to leave it, as it\nsignals \"you should not be calling this directly\" more clearly.\n\n> A bit unfortunate that you inherit the strange show_* naming from\n> builtin/pack-objects.c, which seems to have stolen some code from\n> builtin/rev-list.c at some point without worrying about better naming...\n\nYes, I agree they're not very descriptive. Let's leave it for now to\nstay consistent with pack-objects, and I'd be happy to see a patch\ngiving all of them better names come later.\n\n> > +\twhile (i < objects->word_alloc && ewah_iterator_next(&filter, &it)) {\n> > +\t\teword_t word = objects->words[i] & filter;\n> > +\n> > +\t\tfor (offset = 0; offset < BITS_IN_WORD; ++offset) {\n> > +\t\t\tconst unsigned char *sha1;\n> > +\t\t\tstruct revindex_entry *entry;\n> > +\t\t\tuint32_t hash = 0;\n> > +\n> > +\t\t\tif ((word >> offset) == 0)\n> > +\t\t\t\tbreak;\n> > +\n> > +\t\t\toffset += ewah_bit_ctz64(word >> offset);\n> > +\n> > +\t\t\tif (pos + offset < bitmap_git.reuse_objects)\n> > +\t\t\t\tcontinue;\n> > +\n> > +\t\t\tentry = &bitmap_git.reverse_index->revindex[pos + offset];\n> > +\t\t\tsha1 = nth_packed_object_sha1(bitmap_git.pack, entry->nr);\n> > +\n> > +\t\t\tshow_reach(sha1, object_type, 0, hash, bitmap_git.pack, entry->offset);\n> > +\t\t}\n> \n> You have a very nice bitmap_each_bit() function in ewah/bitmap.c, why\n> not use it here?\n\nWe are bitwise-ANDing against an on-disk ewah bitmap to filter out\nobjects which do not match the desired type. bitmap_each_bit would make\nthis more complicated, because we wouldn't be able to move the\newah_iterator in single-word lockstep. And it would probably be slower\n(if you did it naively), because we'd end up checking each bit in the\newah, rather than AND-ing whole words.\n\nThe right, reusable way to do it would probably be to bitmap_and_ewah\nthe original and the filter together, and then bitmap_each_bit the\nresult. But you would have to write bitmap_and_ewah first. :)\n\n> > +\t/*\n> > +\t * Reuse the packfile content if we need more than\n> > +\t * 90% of its objects\n> > +\t */\n> > +\tstatic const double REUSE_PERCENT = 0.9;\n> \n> Curious: is this based on some measurements or just a guess?\n\nI think it's mostly a guess.\n\n> > +enum pack_bitmap_opts {\n> > +\tBITMAP_OPT_FULL_DAG = 1,\n> \n> And I think this trailing comma on the last enum item is also strictly\n> speaking not allowed, even though it is very nice to have:\n> \n> pack-bitmap.h:28:27: warning: comma at end of enumerator list [-Wpedantic]\n\nIt's allowed in C99, but was not in C89.  I've fixed this site for\nconsistency with the rest of git. But I wonder how relevant it still is.\nThe only data points I know of are:\n\n  http://article.gmane.org/gmane.comp.version-control.git/145739\n\nand\n\n  http://article.gmane.org/gmane.comp.version-control.git/145739\n\nIt sounds like an ancient IBM VisualAge is the only reported problem.\nAnd according to IBM, they stopped supporting it 10 years ago (well,\ntechnically we have a few more weeks to hit the 10-year mark):\n\n  http://www-01.ibm.com/common/ssi/cgi-bin/ssialias?infotype=an&subtype=ca&supplier=897&appname=IBMLinkRedirect&letternum=ENUS903-227\n\nI do wonder if at some point we should revisit our \"do not use any\nC99-isms\" philosophy. It was very good advice in 2005. I don't know how\ngood it is over 8 years later (it seems like even ancient systems should\nbe able to get gcc compiled as a last resort, but maybe there really are\npeople for whom that is a burden).\n\n-Peff\n"},{"id":"231388","messageId":"xmqqk3fnq9bh.fsf@gitster.dls.corp.google.com","threadId":"35335","inReplyTo":"20131202161208.GB24202@sigill.intra.peff.net","subject":"Re: [PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-12-02T20:36:34Z","receivedAt":"2013-12-02T20:36:34Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> I do wonder if at some point we should revisit our \"do not use any\n> C99-isms\" philosophy. It was very good advice in 2005. I don't know how\n> good it is over 8 years later (it seems like even ancient systems should\n> be able to get gcc compiled as a last resort, but maybe there really are\n> people for whom that is a burden).\n\nWell, we are not kernel where being able to precisely control\ngenerated machine code matters and enforcement of acceptable\ncompiler versions to achieve that goal is warranted, so I'd prefer\nto avoid anything that tells the users \"go get a newer gcc\".\n\nThere are certain things outside C89 that would make our code easier\nto read and maintain (e.g. named member initialization of\nstruct/union, cf. ANSI C99 s6.7.9, just to name one) that I would\nlove to be able to use in our codebase, but being able to leave an\nextra comma at the list of enums is very low on that list.  Other\nthan making a patch unnecessarily verbose when you introduce a new\nenum and append it at the end, I do not think it hurts us much (I\nmay be missing other reasons why you may want to leave an extra\ncomma at the end, though).\n\nEven that problem can easily be solved by having the \"a value of\nthis enum has to be lower than this\" at the end, e.g.\n\n        enum object_type {\n                OBJ_BAD = -1,\n                OBJ_NONE = 0,\n                OBJ_COMMIT = 1,\n                OBJ_TREE = 2,\n                OBJ_BLOB = 3,\n                OBJ_TAG = 4,\n                /* 5 for future expansion */\n                OBJ_OFS_DELTA = 6,\n                OBJ_REF_DELTA = 7,\n                OBJ_ANY,\n                OBJ_MAX\n        };\n\n\nwhich will never tempt anybody to append at the end, causing the\npatch bloat.\n"},{"id":"231389","messageId":"20131202204715.GA18842@sigill.intra.peff.net","threadId":"35335","inReplyTo":"xmqqk3fnq9bh.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-02T20:47:15Z","receivedAt":"2013-12-02T20:47:15Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Dec 02, 2013 at 12:36:34PM -0800, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > I do wonder if at some point we should revisit our \"do not use any\n> > C99-isms\" philosophy. It was very good advice in 2005. I don't know how\n> > good it is over 8 years later (it seems like even ancient systems should\n> > be able to get gcc compiled as a last resort, but maybe there really are\n> > people for whom that is a burden).\n> \n> Well, we are not kernel where being able to precisely control\n> generated machine code matters and enforcement of acceptable\n> compiler versions to achieve that goal is warranted, so I'd prefer\n> to avoid anything that tells the users \"go get a newer gcc\".\n\nSorry, I was not very clear about what I said. I do not think \"go get a\nnewer gcc\" is a good thing to be telling people. But I wonder:\n\n  a. if there are actually people on systems that have pre-c99 compilers\n     in 2013\n\n  b. if there are, do they actually _use_ the ancient system compiler,\n     and not just install gcc as the first step anyway?\n\nIn other words, I am questioning whether we would have to tell anybody\n\"go install gcc\" these days. I'm not sure of the best way to answer that\nquestion, though.\n\n> There are certain things outside C89 that would make our code easier\n> to read and maintain (e.g. named member initialization of\n> struct/union, cf. ANSI C99 s6.7.9, just to name one) that I would\n> love to be able to use in our codebase, but being able to leave an\n> extra comma at the list of enums is very low on that list.\n\nYes, I can live without trailing commas. I was musing more on the\ngeneral issue (of course, we don't _have_ to take C99 as a whole, and\ncan pick and choose features that even pre-C99 compilers got right, but\nI was wondering mainly when it would be time to say C99 is \"old enough\"\nthat everybody supports it).\n\n-Peff\n"},{"id":"231393","messageId":"xmqq7gbnq67k.fsf@gitster.dls.corp.google.com","threadId":"35335","inReplyTo":"20131202204715.GA18842@sigill.intra.peff.net","subject":"Re: [PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-12-02T21:43:43Z","receivedAt":"2013-12-02T21:43:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Sorry, I was not very clear about what I said. I do not think \"go get a\n> newer gcc\" is a good thing to be telling people. But I wonder:\n>\n>   a. if there are actually people on systems that have pre-c99 compilers\n>      in 2013\n>\n>   b. if there are, do they actually _use_ the ancient system compiler,\n>      and not just install gcc as the first step anyway?\n>\n> In other words, I am questioning whether we would have to tell anybody\n> \"go install gcc\" these days. I'm not sure of the best way to answer that\n> question, though.\n\nThanks. We are at the same wavelength.  I do not know what the best\nway to answer that question, either.  Unleashing such a change to\nthe wild is one obvious way to do so, and it will give us the most\naccurate answer, so it may be the \"best way\" if we primarily care\nabout the quality of the answer, but we value the experience of our\ndirect consumers even more when judging what is best, so...\n"},{"id":"231428","messageId":"529DED69.4080300@gmail.com","threadId":"35335","inReplyTo":"20131128103838.GB14615@sigill.intra.peff.net","subject":"Re: [PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2013-12-03T14:40:41Z","receivedAt":"2013-12-03T14:40:41Z","isPatch":true,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 28.11.2013 11:38, schrieb Jeff King:\n> On Wed, Nov 27, 2013 at 10:08:56AM +0100, Karsten Blees wrote:\n> \n>> Khash is OK for sha1 keys, but I don't think it should be advertised\n>> as a second general purpose hash table implementation. Its far too\n>> easy to shoot yourself in the foot by using 'straightforward' hash-\n>> and comparison functions. Khash doesn't store the hash codes of the\n>> keys, so you have to take care of that yourself or live with the\n>> performance penalties (see [1]).\n>>\n>> [1] http://article.gmane.org/gmane.comp.version-control.git/237876\n> \n> Yes. I wonder if we should improve it in that respect. I haven't looked\n> carefully at the hash code you posted elsewhere, but I feel like many\n> uses will want a macro implementation to let them store arbitrary types\n> smaller or larger than a pointer.\n> \n> -Peff\n> \n\nIMO, trying to improve khash isn't worth the trouble.\n\nWith smaller-than-pointer types, khash _may_ actually save a few bytes compared to hash.[ch] or hashmap.[ch]. E.g. a set of 'int's would be a perfect use case for khash. However, you're using bitmaps for that, and khash's predefined macros for int sets and maps have been removed from the git version.\n\nUsing khash with pointers and larger-than-pointer types is just a waste of memory and performance, though. Hash tables are sparsely filled by design, and using large types means lots of large empty buckets. E.g. kh_resize requires almost four times the size or your data, and copies everything at least three times.\n\nAdditionally, there are the obvious problems with khash's macro design (hard to read, impossible to debug, and each named instance increases executable size by ~4k).\n\n\nBelow is a patch that converts pack-bitmap.c to hashmap. Its not even longer than the khash version, and the hashmap API forces you to think about no-brainer improvements such as specifying an expected size or skipping duplicates checks where they aren't needed. I could do the same for pack-bitmap-write.c if you like.\n\nKarsten\n\n-----------8<----------\nSubject: [PATCH] pack-bitmap.c: convert to new hashmap implementation\n\nPreallocates bitmap_index.bitmaps with the expected size to prevent\nrehashing.\n\nCombines the five arrays of ext_index ('objects', 'hashes' and khash's\n'keys', 'flags' and 'vals') into a single structure for better locality.\n\nRemoves two unnecessary duplicates checks:\n - we don't expect a pack index file to contain duplicate sha1's\n - ext_index_add_object is only called after bitmap_position returned -1,\n   so it is impossible that the entry is already there\n\nSigned-off-by: Karsten Blees <blees@dcon.de>\n---\n pack-bitmap.c | 166 ++++++++++++++++++++++++++++------------------------------\n 1 file changed, 81 insertions(+), 85 deletions(-)\n\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nindex 078f7c6..115caed 100644\n--- a/pack-bitmap.c\n+++ b/pack-bitmap.c\n@@ -15,6 +15,7 @@\n  * commit.\n  */\n struct stored_bitmap {\n+\tstruct hashmap_entry ent;\n \tunsigned char sha1[20];\n \tstruct ewah_bitmap *root;\n \tstruct stored_bitmap *xor;\n@@ -22,6 +23,16 @@ struct stored_bitmap {\n };\n \n /*\n+ * An entry in the extended index.\n+ */\n+struct ext_entry {\n+\tstruct hashmap_entry ent;\n+\tuint32_t name_hash;\n+\tstruct object *object;\n+\tunsigned int nr;\n+};\n+\n+/*\n  * The currently active bitmap index. By design, repositories only have\n  * a single bitmap index available (the index for the biggest packfile in\n  * the repository), since bitmap indexes need full closure.\n@@ -61,7 +72,7 @@ static struct bitmap_index {\n \tstruct ewah_bitmap *tags;\n \n \t/* Map from SHA1 -> `stored_bitmap` for all the bitmapped comits */\n-\tkhash_sha1 *bitmaps;\n+\tstruct hashmap bitmaps;\n \n \t/* Number of bitmapped commits */\n \tuint32_t entry_count;\n@@ -76,12 +87,7 @@ static struct bitmap_index {\n \t * packed in `pack`, these objects are added to this \"fake index\" and\n \t * are assumed to appear at the end of the packfile for all operations\n \t */\n-\tstruct eindex {\n-\t\tstruct object **objects;\n-\t\tuint32_t *hashes;\n-\t\tuint32_t count, alloc;\n-\t\tkhash_sha1_pos *positions;\n-\t} ext_index;\n+\tstruct hashmap ext_index;\n \n \t/* Bitmap result of the last performed walk */\n \tstruct bitmap *result;\n@@ -93,6 +99,21 @@ static struct bitmap_index {\n \n } bitmap_git;\n \n+static int stored_bitmap_cmp(const struct stored_bitmap *entry,\n+\t\tconst struct stored_bitmap *entry_or_key,\n+\t\tconst unsigned char *sha1)\n+{\n+\treturn hashcmp(entry->sha1, sha1 ? sha1 : entry_or_key->sha1);\n+}\n+\n+static int ext_entry_cmp(const struct ext_entry *entry,\n+\t\tconst struct ext_entry *entry_or_key,\n+\t\tconst unsigned char *sha1)\n+{\n+\treturn hashcmp(entry->object->sha1,\n+\t\t       sha1 ? sha1 : entry_or_key->object->sha1);\n+}\n+\n static struct ewah_bitmap *lookup_stored_bitmap(struct stored_bitmap *st)\n {\n \tstruct ewah_bitmap *parent;\n@@ -112,6 +133,16 @@ static struct ewah_bitmap *lookup_stored_bitmap(struct stored_bitmap *st)\n \treturn composed;\n }\n \n+static struct ewah_bitmap *lookup_stored_bitmap_sha1(const unsigned char *sha1)\n+{\n+\tstruct stored_bitmap key, *st;\n+\thashmap_entry_init(&key, __kh_oid_hash(sha1));\n+\tst = hashmap_get(&bitmap_git.bitmaps, &key, sha1);\n+\tif (st)\n+\t\treturn lookup_stored_bitmap(st);\n+\treturn NULL;\n+}\n+\n /*\n  * Read a bitmap from the current read position on the mmaped\n  * index, and increase the read position accordingly\n@@ -174,26 +205,15 @@ static struct stored_bitmap *store_bitmap(struct bitmap_index *index,\n \t\t\t\t\t  int flags)\n {\n \tstruct stored_bitmap *stored;\n-\tkhiter_t hash_pos;\n-\tint ret;\n \n \tstored = xmalloc(sizeof(struct stored_bitmap));\n+\thashmap_entry_init(stored, __kh_oid_hash(sha1));\n \tstored->root = root;\n \tstored->xor = xor_with;\n \tstored->flags = flags;\n \thashcpy(stored->sha1, sha1);\n \n-\thash_pos = kh_put_sha1(index->bitmaps, stored->sha1, &ret);\n-\n-\t/* a 0 return code means the insertion succeeded with no changes,\n-\t * because the SHA1 already existed on the map. this is bad, there\n-\t * shouldn't be duplicated commits in the index */\n-\tif (ret == 0) {\n-\t\terror(\"Duplicate entry in bitmap index: %s\", sha1_to_hex(sha1));\n-\t\treturn NULL;\n-\t}\n-\n-\tkh_value(index->bitmaps, hash_pos) = stored;\n+\thashmap_add(&index->bitmaps, stored);\n \treturn stored;\n }\n \n@@ -291,8 +311,9 @@ static int load_pack_bitmap(void)\n {\n \tassert(bitmap_git.map && !bitmap_git.loaded);\n \n-\tbitmap_git.bitmaps = kh_init_sha1();\n-\tbitmap_git.ext_index.positions = kh_init_sha1_pos();\n+\thashmap_init(&bitmap_git.bitmaps, (hashmap_cmp_fn) stored_bitmap_cmp,\n+\t\t\tbitmap_git.entry_count);\n+\thashmap_init(&bitmap_git.ext_index, (hashmap_cmp_fn) ext_entry_cmp, 0);\n \tbitmap_git.reverse_index = revindex_for_pack(bitmap_git.pack);\n \n \tif (!(bitmap_git.commits = read_bitmap_1(&bitmap_git)) ||\n@@ -362,13 +383,12 @@ struct include_data {\n \n static inline int bitmap_position_extended(const unsigned char *sha1)\n {\n-\tkhash_sha1_pos *positions = bitmap_git.ext_index.positions;\n-\tkhiter_t pos = kh_get_sha1_pos(positions, sha1);\n+\tstruct ext_entry key, *e;\n+\thashmap_entry_init(&key, __kh_oid_hash(sha1));\n+\te = hashmap_get(&bitmap_git.bitmaps, &key, sha1);\n \n-\tif (pos < kh_end(positions)) {\n-\t\tint bitmap_pos = kh_value(positions, pos);\n-\t\treturn bitmap_pos + bitmap_git.pack->num_objects;\n-\t}\n+\tif (e)\n+\t\treturn e->nr + bitmap_git.pack->num_objects;\n \n \treturn -1;\n }\n@@ -390,32 +410,13 @@ static int bitmap_position(const unsigned char *sha1)\n \n static int ext_index_add_object(struct object *object, const char *name)\n {\n-\tstruct eindex *eindex = &bitmap_git.ext_index;\n-\n-\tkhiter_t hash_pos;\n-\tint hash_ret;\n-\tint bitmap_pos;\n-\n-\thash_pos = kh_put_sha1_pos(eindex->positions, object->sha1, &hash_ret);\n-\tif (hash_ret > 0) {\n-\t\tif (eindex->count >= eindex->alloc) {\n-\t\t\teindex->alloc = (eindex->alloc + 16) * 3 / 2;\n-\t\t\teindex->objects = xrealloc(eindex->objects,\n-\t\t\t\teindex->alloc * sizeof(struct object *));\n-\t\t\teindex->hashes = xrealloc(eindex->hashes,\n-\t\t\t\teindex->alloc * sizeof(uint32_t));\n-\t\t}\n-\n-\t\tbitmap_pos = eindex->count;\n-\t\teindex->objects[eindex->count] = object;\n-\t\teindex->hashes[eindex->count] = pack_name_hash(name);\n-\t\tkh_value(eindex->positions, hash_pos) = bitmap_pos;\n-\t\teindex->count++;\n-\t} else {\n-\t\tbitmap_pos = kh_value(eindex->positions, hash_pos);\n-\t}\n-\n-\treturn bitmap_pos + bitmap_git.pack->num_objects;\n+\tstruct ext_entry *e = xmalloc(sizeof(struct ext_entry));\n+\thashmap_entry_init(e, __kh_oid_hash(object->sha1));\n+\te->object = object;\n+\te->name_hash = pack_name_hash(name);\n+\te->nr = bitmap_git.ext_index.size;\n+\thashmap_add(&bitmap_git.ext_index, e);\n+\treturn e->nr + bitmap_git.pack->num_objects;\n }\n \n static void show_object(struct object *object, const struct name_path *path,\n@@ -443,7 +444,7 @@ static int add_to_include_set(struct include_data *data,\n \t\t\t      const unsigned char *sha1,\n \t\t\t      int bitmap_pos)\n {\n-\tkhiter_t hash_pos;\n+\tstruct ewah_bitmap *bm;\n \n \tif (data->seen && bitmap_get(data->seen, bitmap_pos))\n \t\treturn 0;\n@@ -451,10 +452,9 @@ static int add_to_include_set(struct include_data *data,\n \tif (bitmap_get(data->base, bitmap_pos))\n \t\treturn 0;\n \n-\thash_pos = kh_get_sha1(bitmap_git.bitmaps, sha1);\n-\tif (hash_pos < kh_end(bitmap_git.bitmaps)) {\n-\t\tstruct stored_bitmap *st = kh_value(bitmap_git.bitmaps, hash_pos);\n-\t\tbitmap_or_ewah(data->base, lookup_stored_bitmap(st));\n+\tbm = lookup_stored_bitmap_sha1(sha1);\n+\tif (bm) {\n+\t\tbitmap_or_ewah(data->base, bm);\n \t\treturn 0;\n \t}\n \n@@ -507,12 +507,9 @@ static struct bitmap *find_objects(struct rev_info *revs,\n \t\troots = roots->next;\n \n \t\tif (object->type == OBJ_COMMIT) {\n-\t\t\tkhiter_t pos = kh_get_sha1(bitmap_git.bitmaps, object->sha1);\n-\n-\t\t\tif (pos < kh_end(bitmap_git.bitmaps)) {\n-\t\t\t\tstruct stored_bitmap *st = kh_value(bitmap_git.bitmaps, pos);\n-\t\t\t\tstruct ewah_bitmap *or_with = lookup_stored_bitmap(st);\n-\n+\t\t\tstruct ewah_bitmap *or_with =\n+\t\t\t\tlookup_stored_bitmap_sha1(object->sha1);\n+\t\t\tif (or_with) {\n \t\t\t\tif (base == NULL)\n \t\t\t\t\tbase = ewah_to_bitmap(or_with);\n \t\t\t\telse\n@@ -584,17 +581,15 @@ static struct bitmap *find_objects(struct rev_info *revs,\n static void show_extended_objects(struct bitmap *objects,\n \t\t\t\t  show_reachable_fn show_reach)\n {\n-\tstruct eindex *eindex = &bitmap_git.ext_index;\n-\tuint32_t i;\n-\n-\tfor (i = 0; i < eindex->count; ++i) {\n-\t\tstruct object *obj;\n+\tstruct ext_entry *e;\n+\tstruct hashmap_iter iter;\n \n-\t\tif (!bitmap_get(objects, bitmap_git.pack->num_objects + i))\n+\tfor (e = hashmap_iter_first(&bitmap_git.ext_index, &iter); e;\n+\t     e = hashmap_iter_next(&iter)) {\n+\t\tif (!bitmap_get(objects, bitmap_git.pack->num_objects + e->nr))\n \t\t\tcontinue;\n \n-\t\tobj = eindex->objects[i];\n-\t\tshow_reach(obj->sha1, obj->type, 0, eindex->hashes[i], NULL, 0);\n+\t\tshow_reach(e->object->sha1, e->object->type, 0, e->name_hash, NULL, 0);\n \t}\n }\n \n@@ -831,7 +826,8 @@ void traverse_bitmap_commit_list(show_reachable_fn show_reachable)\n static uint32_t count_object_type(struct bitmap *objects,\n \t\t\t\t  enum object_type type)\n {\n-\tstruct eindex *eindex = &bitmap_git.ext_index;\n+\tstruct ext_entry *e;\n+\tstruct hashmap_iter iter;\n \n \tuint32_t i = 0, count = 0;\n \tstruct ewah_iterator it;\n@@ -863,9 +859,10 @@ static uint32_t count_object_type(struct bitmap *objects,\n \t\tcount += ewah_bit_popcount64(word);\n \t}\n \n-\tfor (i = 0; i < eindex->count; ++i) {\n-\t\tif (eindex->objects[i]->type == type &&\n-\t\t\tbitmap_get(objects, bitmap_git.pack->num_objects + i))\n+\tfor (e = hashmap_iter_first(&bitmap_git.ext_index, &iter); e;\n+\t     e = hashmap_iter_next(&iter)) {\n+\t\tif (e->object->type == type &&\n+\t\t\tbitmap_get(objects, bitmap_git.pack->num_objects + e->nr))\n \t\t\tcount++;\n \t}\n \n@@ -928,7 +925,7 @@ void test_bitmap_walk(struct rev_info *revs)\n {\n \tstruct object *root;\n \tstruct bitmap *result = NULL;\n-\tkhiter_t pos;\n+\tstruct ewah_bitmap *bm;\n \tsize_t result_popcnt;\n \tstruct bitmap_test_data tdata;\n \n@@ -942,12 +939,9 @@ void test_bitmap_walk(struct rev_info *revs)\n \t\tbitmap_git.version, bitmap_git.entry_count);\n \n \troot = revs->pending.objects[0].item;\n-\tpos = kh_get_sha1(bitmap_git.bitmaps, root->sha1);\n-\n-\tif (pos < kh_end(bitmap_git.bitmaps)) {\n-\t\tstruct stored_bitmap *st = kh_value(bitmap_git.bitmaps, pos);\n-\t\tstruct ewah_bitmap *bm = lookup_stored_bitmap(st);\n \n+\tbm = lookup_stored_bitmap_sha1(root->sha1);\n+\tif (bm) {\n \t\tfprintf(stderr, \"Found bitmap for %s. %d bits / %08x checksum\\n\",\n \t\t\tsha1_to_hex(root->sha1), (int)bm->bit_size, ewah_checksum(bm));\n \n@@ -1020,6 +1014,7 @@ int rebuild_existing_bitmaps(struct packing_data *mapping,\n \tstruct bitmap *rebuild;\n \tstruct stored_bitmap *stored;\n \tstruct progress *progress = NULL;\n+\tstruct hashmap_iter iter;\n \n \tkhiter_t hash_pos;\n \tint hash_ret;\n@@ -1049,7 +1044,8 @@ int rebuild_existing_bitmaps(struct packing_data *mapping,\n \tif (show_progress)\n \t\tprogress = start_progress(\"Reusing bitmaps\", 0);\n \n-\tkh_foreach_value(bitmap_git.bitmaps, stored, {\n+\tfor (stored = hashmap_iter_first(&bitmap_git.bitmaps, &iter); stored;\n+\t     stored = hashmap_iter_next(&iter)) {\n \t\tif (stored->flags & BITMAP_FLAG_REUSE) {\n \t\t\tif (!rebuild_bitmap(reposition,\n \t\t\t\t\t    lookup_stored_bitmap(stored),\n@@ -1063,7 +1059,7 @@ int rebuild_existing_bitmaps(struct packing_data *mapping,\n \t\t\tbitmap_reset(rebuild);\n \t\t\tdisplay_progress(progress, ++i);\n \t\t}\n-\t});\n+\t}\n \n \tstop_progress(&progress);\n \n-- \n"},{"id":"231442","messageId":"20131203182133.GA21296@sigill.intra.peff.net","threadId":"35335","inReplyTo":"529DED69.4080300@gmail.com","subject":"Re: [PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-03T18:21:33Z","receivedAt":"2013-12-03T18:21:33Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Dec 03, 2013 at 03:40:41PM +0100, Karsten Blees wrote:\n\n> IMO, trying to improve khash isn't worth the trouble.\n> \n> With smaller-than-pointer types, khash _may_ actually save a few bytes\n> compared to hash.[ch] or hashmap.[ch]. E.g. a set of 'int's would be a\n> perfect use case for khash. However, you're using bitmaps for that,\n> and khash's predefined macros for int sets and maps have been removed\n> from the git version.\n\nTrue, we are not using it for smaller-than-pointer sizes here. So it may\nnot be worth thinking about (I was considering more for the general\ncase). In most instances, though, we can shove the int bits into a\npointer with the right casting. So it probably isn't worth worrying\nabout (you may waste a few bytes, but probably not more than one word\nper entry).\n\n> Using khash with pointers and larger-than-pointer types is just a\n> waste of memory and performance, though. Hash tables are sparsely\n> filled by design, and using large types means lots of large empty\n> buckets. E.g. kh_resize requires almost four times the size or your\n> data, and copies everything at least three times.\n\nI think the analysis is more complicated than that, and depends on what\nyou are storing. If you are storing something that is 1.5 times the size\nof a pointer, it is more space efficient to just stick it in the hash\ntable than it is to have a separate pointer in the hash table (you pay\nthe load factor penalty only on the single pointer, but you've almost\ndoubled your total storage).\n\nBut again, that's a more general argument. We're not storing anything of\nthat size here, and in fact I think we are just storing pointers\neverywhere.\n\n> Additionally, there are the obvious problems with khash's macro design\n> (hard to read, impossible to debug, and each named instance increases\n> executable size by ~4k).\n\nYes, those are all downsides to macros. Type safety is one of the\nupsides, though.\n\nBesides macros, I think the major difference between the two\nimplementations is open-addressing versus chaining. Especially for sets,\nwe've had good experiences with open-addressing by keeping the load\nfactor low (e.g., the one in object.c). As you note, the resizing\noperation pays some penalty, but in most of our workloads it's largely\nirrelevant compared to lookup times.\n\nChaining typically adds an extra layer of pointer-following to the\nlookup, and I'd be worried about that loss of locality. It's true that\nwhen you are storing a pointer you are already hurting locality to\nfollow the pointer to the key, but I don't know whether that means an\n_extra_ layer doesn't still hurt more (and I really mean I don't know --\nI'd be interested to see measurements).\n\nIn your implementation, it looks like you break even there because you\nstore the hash directly in the entry, and do a single-word compare (so\nyou avoid having to follow a pointer to the key in the common case\nduring lookup). But that also means you're paying extra to store the\nhash. That probably makes sense for things like strings, where it takes\nsome effort to calculate the hash. But not necessarily for sha1s, where\nlooking at the hash is the same thing as looking at the key bytes (so\nyou are storing extra bytes, and when you do have a hit on the stored\nhash, which is just the first bytes of the sha1, you end up comparing\nthem again as part of the hashcmp. The tradeoff is that in the\nnon-matching cases, you avoid an extra level of indirection).\n\nAll of these are things I could see helping or hurting, depending on the\ncase. I'd really love to see more numbers. I tried timing your patch\nbelow, but there was no interesting change. Mostly because the code path\nyou changed is not heavily exercised (the \"ext_index\" is for objects\nthat are not in the pack, and frequent packing keeps that number low).\n\nI'd be really interested to see if you could make the hash in object.c\nfaster. That one is a prominent lookup bottle-neck (e.g., for \"git\nrev-list --objects --all\"); if you can make that faster, I would be very\nconvinced that your implementation is fast (note that I am not implying\nthe converse; if you cannot make it faster, that does not necessarily\nmean your implementation sucks, but perhaps only that the existing one\nhas lots of type-specific optimizations which add up).\n\nKhash also has a lot of bloat (e.g., flags) that the one in object.c\ndoes not have. If you do not care about deletion and are storing\nsomething with a sentinel value (e.g., NULL for pointers), you can trim\nquite a bit of fat.\n\n> Below is a patch that converts pack-bitmap.c to hashmap. Its not even\n> longer than the khash version, and the hashmap API forces you to think\n> about no-brainer improvements such as specifying an expected size or\n> skipping duplicates checks where they aren't needed. I could do the\n> same for pack-bitmap-write.c if you like.\n\nIf it's not too much trouble, I'd be curious to measure the performance\nimpact on pack-bitmap-write.\n\n> Removes two unnecessary duplicates checks:\n>  - we don't expect a pack index file to contain duplicate sha1's\n\nWe don't expect them to, but it has happened (and caused bugs not too\nlong ago). What happens after your patch when there are duplicates?\n\n> +static struct ewah_bitmap *lookup_stored_bitmap_sha1(const unsigned char *sha1)\n> +{\n> +\tstruct stored_bitmap key, *st;\n> +\thashmap_entry_init(&key, __kh_oid_hash(sha1));\n> +\tst = hashmap_get(&bitmap_git.bitmaps, &key, sha1);\n> +\tif (st)\n> +\t\treturn lookup_stored_bitmap(st);\n> +\treturn NULL;\n> +}\n\nThis interface looks odd to me. You create a fake stored_bitmap for the\nhashmap_entry part of it, and then fill in the \"hash\" field by hashing\nthe sha1. And then pass the same sha1 in. I guess you are trying to\navoid the hash table knowing about the hash function at all, since it\njust stores the hash for each entry already.\n\nI guess that makes sense, but would not work for a system where you\nwanted to get rid of the extra hash storage (but as I implied above, I\nam not sure if it is helping or hurting for the sha1 case).\n\n> -\thash_pos = kh_put_sha1_pos(eindex->positions, object->sha1, &hash_ret);\n> -\tif (hash_ret > 0) {\n> -\t\tif (eindex->count >= eindex->alloc) {\n> -\t\t\teindex->alloc = (eindex->alloc + 16) * 3 / 2;\n> -\t\t\teindex->objects = xrealloc(eindex->objects,\n> -\t\t\t\teindex->alloc * sizeof(struct object *));\n> -\t\t\teindex->hashes = xrealloc(eindex->hashes,\n> -\t\t\t\teindex->alloc * sizeof(uint32_t));\n> -\t\t}\n> -\n> -\t\tbitmap_pos = eindex->count;\n> -\t\teindex->objects[eindex->count] = object;\n> -\t\teindex->hashes[eindex->count] = pack_name_hash(name);\n> -\t\tkh_value(eindex->positions, hash_pos) = bitmap_pos;\n> -\t\teindex->count++;\n> -\t} else {\n> -\t\tbitmap_pos = kh_value(eindex->positions, hash_pos);\n> -\t}\n> -\n> -\treturn bitmap_pos + bitmap_git.pack->num_objects;\n> +\tstruct ext_entry *e = xmalloc(sizeof(struct ext_entry));\n> +\thashmap_entry_init(e, __kh_oid_hash(object->sha1));\n> +\te->object = object;\n> +\te->name_hash = pack_name_hash(name);\n> +\te->nr = bitmap_git.ext_index.size;\n> +\thashmap_add(&bitmap_git.ext_index, e);\n> +\treturn e->nr + bitmap_git.pack->num_objects;\n\nOne of the side effects of the current system is that the array\neffectively works as a custom allocator, and we do not pay per-entry\nmalloc overhead. As I mentioned above, this is not that\nheavily-exercised a code path, so it may not matter (and you can convert\nit by explicitly using a custom allocator, at which point you could drop\nthe e->nr field entirely, I'd think).\n\n> @@ -584,17 +581,15 @@ static struct bitmap *find_objects(struct rev_info *revs,\n>  static void show_extended_objects(struct bitmap *objects,\n>  \t\t\t\t  show_reachable_fn show_reach)\n>  {\n> -\tstruct eindex *eindex = &bitmap_git.ext_index;\n> -\tuint32_t i;\n> -\n> -\tfor (i = 0; i < eindex->count; ++i) {\n> -\t\tstruct object *obj;\n> +\tstruct ext_entry *e;\n> +\tstruct hashmap_iter iter;\n>  \n> -\t\tif (!bitmap_get(objects, bitmap_git.pack->num_objects + i))\n> +\tfor (e = hashmap_iter_first(&bitmap_git.ext_index, &iter); e;\n> +\t     e = hashmap_iter_next(&iter)) {\n> +\t\tif (!bitmap_get(objects, bitmap_git.pack->num_objects + e->nr))\n>  \t\t\tcontinue;\n>  \n> -\t\tobj = eindex->objects[i];\n> -\t\tshow_reach(obj->sha1, obj->type, 0, eindex->hashes[i], NULL, 0);\n> +\t\tshow_reach(e->object->sha1, e->object->type, 0, e->name_hash, NULL, 0);\n\nBefore we were iterating in eindex order. Now we are iterating in\nhashmap order. That effects our traversal order, and ultimately our\noutput for something like \"rev-list\". We'd want to sort on e->nr (or\nagain, a custom allocator would make this go away because we would\niterate over the array).\n\n> -\tfor (i = 0; i < eindex->count; ++i) {\n> -\t\tif (eindex->objects[i]->type == type &&\n> -\t\t\tbitmap_get(objects, bitmap_git.pack->num_objects + i))\n> +\tfor (e = hashmap_iter_first(&bitmap_git.ext_index, &iter); e;\n> +\t     e = hashmap_iter_next(&iter)) {\n> +\t\tif (e->object->type == type &&\n> +\t\t\tbitmap_get(objects, bitmap_git.pack->num_objects + e->nr))\n>  \t\t\tcount++;\n\nDitto here, but I do not think the order matters in this instance.\n\n-Peff\n"},{"id":"231708","messageId":"87zjock6if.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","threadId":"35335","inReplyTo":"20131114124510.GK10757@sigill.intra.peff.net","subject":"Re: [PATCH v3 11/21] pack-objects: use bitmaps when packing objects","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-12-07T15:47:20Z","receivedAt":"2013-12-07T15:47:20Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"This week's nits...\n\nI found this harder to read than the previous patch, but I think it's\nmostly because the existing code is already a bit tangled.  I think the\nsecond item below is worth fixing, though.\n\n\nJeff King <peff@peff.net> writes:\n\n> +static off_t write_reused_pack(struct sha1file *f)\n> +{\n> +\tuint8_t buffer[8192];\n\nWe usually just call this 'unsigned char'.  I can see why this would be\nmore portable, but git would already fall apart badly on an architecture\nwhere char is not 8 bits.\n\n> +\toff_t to_write;\n> +\tint fd;\n> +\n> +\tif (!is_pack_valid(reuse_packfile))\n> +\t\treturn 0;\n> +\n> +\tfd = git_open_noatime(reuse_packfile->pack_name);\n> +\tif (fd < 0)\n> +\t\treturn 0;\n> +\n> +\tif (lseek(fd, sizeof(struct pack_header), SEEK_SET) == -1) {\n> +\t\tclose(fd);\n> +\t\treturn 0;\n> +\t}\n\nYou do an error return if any of the syscalls in this routine fails, but\nthere is only one caller and it immediately dies:\n\n} +\t\t\tpackfile_size = write_reused_pack(f);\n} +\t\t\tif (!packfile_size)\n} +\t\t\t\tdie_errno(\"failed to re-use existing pack\");\n\nSo if you just died here, when the error happens, you could take the\nchance to tell the user _which_ syscall failed.\n\n> +\n> +\tif (reuse_packfile_offset < 0)\n> +\t\treuse_packfile_offset = reuse_packfile->pack_size - 20;\n> +\n> +\tto_write = reuse_packfile_offset - sizeof(struct pack_header);\n> +\n> +\twhile (to_write) {\n> +\t\tint read_pack = xread(fd, buffer, sizeof(buffer));\n> +\n> +\t\tif (read_pack <= 0) {\n> +\t\t\tclose(fd);\n> +\t\t\treturn 0;\n\nSimilar to the above, but this one may also clobber the 'errno' during\nclose(), which can lead to misleading messages.\n\n> +\t\t}\n> +\n> +\t\tif (read_pack > to_write)\n> +\t\t\tread_pack = to_write;\n> +\n> +\t\tsha1write(f, buffer, read_pack);\n\nNot your fault, but sha1write() is an odd function -- it purportedly is\n\n  int sha1write(struct sha1file *f, const void *buf, unsigned int count);\n\nbut it can only return 0.  This goes back all the way to c38138c\n(git-pack-objects: write the pack files with a SHA1 csum, 2005-06-26).\n\n> +\t\tto_write -= read_pack;\n> +\t}\n> +\n> +\tclose(fd);\n> +\twritten += reuse_packfile_objects;\n> +\treturn reuse_packfile_offset - sizeof(struct pack_header);\n> +}\n[...]\n> -static int add_object_entry(const unsigned char *sha1, enum object_type type,\n> -\t\t\t    const char *name, int exclude)\n> +static int add_object_entry_1(const unsigned char *sha1, enum object_type type,\n> +\t\t\t      int flags, uint32_t name_hash,\n> +\t\t\t      struct packed_git *found_pack, off_t found_offset)\n>  {\n[...]\n> -\tfor (p = packed_git; p; p = p->next) {\n> -\t\toff_t offset = find_pack_entry_one(sha1, p);\n> -\t\tif (offset) {\n> -\t\t\tif (!found_pack) {\n> -\t\t\t\tif (!is_pack_valid(p)) {\n> -\t\t\t\t\twarning(\"packfile %s cannot be accessed\", p->pack_name);\n> -\t\t\t\t\tcontinue;\n> +\tif (!found_pack) {\n> +\t\tfor (p = packed_git; p; p = p->next) {\n> +\t\t\toff_t offset = find_pack_entry_one(sha1, p);\n> +\t\t\tif (offset) {\n> +\t\t\t\tif (!found_pack) {\n> +\t\t\t\t\tif (!is_pack_valid(p)) {\n> +\t\t\t\t\t\twarning(\"packfile %s cannot be accessed\", p->pack_name);\n> +\t\t\t\t\t\tcontinue;\n> +\t\t\t\t\t}\n> +\t\t\t\t\tfound_offset = offset;\n> +\t\t\t\t\tfound_pack = p;\n>  \t\t\t\t}\n> -\t\t\t\tfound_offset = offset;\n> -\t\t\t\tfound_pack = p;\n> +\t\t\t\tif (exclude)\n> +\t\t\t\t\tbreak;\n> +\t\t\t\tif (incremental)\n> +\t\t\t\t\treturn 0;\n> +\t\t\t\tif (local && !p->pack_local)\n> +\t\t\t\t\treturn 0;\n> +\t\t\t\tif (ignore_packed_keep && p->pack_local && p->pack_keep)\n> +\t\t\t\t\treturn 0;\n>  \t\t\t}\n> -\t\t\tif (exclude)\n> -\t\t\t\tbreak;\n> -\t\t\tif (incremental)\n> -\t\t\t\treturn 0;\n> -\t\t\tif (local && !p->pack_local)\n> -\t\t\t\treturn 0;\n> -\t\t\tif (ignore_packed_keep && p->pack_local && p->pack_keep)\n> -\t\t\t\treturn 0;\n>  \t\t}\n>  \t}\n\nThis function makes my head spin, and you're indenting it yet another\nlevel.\n\nIf it's not too much work, can you split it into the three parts that it\nreally is?  IIUC it boils down to\n\n  do we have this already?\n      possibly apply 'exclude', then return\n  are we coming from a call path that doesn't tell us which pack to take\n  it from?\n      find _all_ instances in packs\n      check if any of them are local .keep packs\n          if so, return\n  construct a packlist entry to taste\n\n>  \tentry = packlist_alloc(&to_pack, sha1, index_pos);\n> -\tentry->hash = hash;\n> +\tentry->hash = name_hash;\n>  \tif (type)\n>  \t\tentry->type = type;\n>  \tif (exclude)\n>  \t\tentry->preferred_base = 1;\n>  \telse\n>  \t\tnr_result++;\n> +\n> +\tif (flags & OBJECT_ENTRY_NO_TRY_DELTA)\n> +\t\tentry->no_try_delta = 1;\n> +\n>  \tif (found_pack) {\n>  \t\tentry->in_pack = found_pack;\n>  \t\tentry->in_pack_offset = found_offset;\n> @@ -859,10 +932,21 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n>  \n>  \tdisplay_progress(progress_state, to_pack.nr_objects);\n>  \n> +\treturn 1;\n> +}\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231709","messageId":"87txekk5ne.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","threadId":"35335","inReplyTo":"20131114124523.GL10757@sigill.intra.peff.net","subject":"Re: [PATCH v3 12/21] rev-list: add bitmap mode to speed up object lists","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-12-07T16:05:57Z","receivedAt":"2013-12-07T16:05:57Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>  Documentation/git-rev-list.txt     |  1 +\n>  Documentation/rev-list-options.txt |  8 ++++++++\n>  builtin/rev-list.c                 | 39 ++++++++++++++++++++++++++++++++++++++\n>  3 files changed, 48 insertions(+)\n\nNice and short.  Cheating though, as a lot of the support code is from\n[10/21] ;-)\n\nReviewed-by: Thomas Rast <tr@thomasrast.ch>\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231721","messageId":"87d2l8k4es.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","threadId":"35335","inReplyTo":"20131114124544.GM10757@sigill.intra.peff.net","subject":"Re: [PATCH v3 13/21] pack-objects: implement bitmap writing","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-12-07T16:32:43Z","receivedAt":"2013-12-07T16:32:43Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Reviewed-by: Thomas Rast <tr@thomasrast.ch>\n\nYou could fix this:\n\n> +pack.writebitmaps::\n> +\tWhen true, git will write a bitmap index when packing all\n> +\tobjects to disk (e.g., as when `git repack -a` is run).  This\n                               ^^\n\nDoesn't sound right in my ears.  Remove the \"as\"?\n\n> +\tindex can speed up the \"counting objects\" phase of subsequent\n> +\tpacks created for clones and fetches, at the cost of some disk\n> +\tspace and extra time spent on the initial repack.  Defaults to\n> +\tfalse.\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231722","messageId":"878uvwk4cm.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","threadId":"35335","inReplyTo":"20131114124554.GN10757@sigill.intra.peff.net","subject":"Re: [PATCH v3 14/21] repack: stop using magic number for ARRAY_SIZE(exts)","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-12-07T16:34:01Z","receivedAt":"2013-12-07T16:34:01Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> We have a static array of extensions, but hardcode the size\n> of the array in our loops. Let's pull out this magic number,\n> which will make it easier to change.\n>\n> Signed-off-by: Jeff King <peff@peff.net>\n\nReviewed-by: Thomas Rast <tr@thomasrast.ch>\n\n(Ok, this one was easy.)\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231723","messageId":"874n6kk4br.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","threadId":"35335","inReplyTo":"20131114124600.GO10757@sigill.intra.peff.net","subject":"Re: [PATCH v3 15/21] repack: turn exts array into array-of-struct","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-12-07T16:34:32Z","receivedAt":"2013-12-07T16:34:32Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> This is slightly more verbose, but will let us annotate the\n> extensions with further options in future commits.\n>\n> Signed-off-by: Jeff King <peff@peff.net>\n\nReviewed-by: Thomas Rast <tr@thomasrast.ch>\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231724","messageId":"87zjocippb.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","threadId":"35335","inReplyTo":"20131114124605.GP10757@sigill.intra.peff.net","subject":"Re: [PATCH v3 16/21] repack: handle optional files created by pack-objects","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-12-07T16:35:44Z","receivedAt":"2013-12-07T16:35:44Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> We ask pack-objects to pack to a set of temporary files, and\n> then rename them into place. Some files that pack-objects\n> creates may be optional (like a .bitmap file), in which case\n> we would not want to call rename(). We already call stat()\n> and make the chmod optional if the file cannot be accessed.\n> We could simply skip the rename step in this case, but that\n> would be a minor regression in noticing problems with\n> non-optional files (like the .pack and .idx files).\n>\n> Instead, we can now annotate extensions as optional, and\n> skip them if they don't exist (and otherwise rely on\n> rename() to barf).\n>\n> Signed-off-by: Jeff King <peff@peff.net>\n\nReviewed-by: Thomas Rast <tr@thomasrast.ch>\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231725","messageId":"87vbz0ipmh.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","threadId":"35335","inReplyTo":"20131114124610.GQ10757@sigill.intra.peff.net","subject":"Re: [PATCH v3 17/21] repack: consider bitmaps when performing repacks","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-12-07T16:37:26Z","receivedAt":"2013-12-07T16:37:26Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> From: Vicent Marti <tanoku@gmail.com>\n>\n> Since `pack-objects` will write a `.bitmap` file next to the `.pack` and\n> `.idx` files, this commit teaches `git-repack` to consider the new\n> bitmap indexes (if they exist) when performing repack operations.\n>\n> This implies moving old bitmap indexes out of the way if we are\n> repacking a repository that already has them, and moving the newly\n> generated bitmap indexes into the `objects/pack` directory, next to\n> their corresponding packfiles.\n>\n> Since `git repack` is now capable of handling these `.bitmap` files,\n> a normal `git gc` run on a repository that has `pack.writebitmaps` set\n> to true in its config file will generate bitmap indexes as part of the\n> garbage collection process.\n>\n> Alternatively, `git repack` can be called with the `-b` switch to\n> explicitly generate bitmap indexes if you are experimenting\n> and don't want them on all the time.\n>\n> Signed-off-by: Vicent Marti <tanoku@gmail.com>\n> Signed-off-by: Jeff King <peff@peff.net>\n\nReviewed-by: Thomas Rast <tr@thomasrast.ch>\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231726","messageId":"87r49oipl8.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","threadId":"35335","inReplyTo":"20131114124618.GR10757@sigill.intra.peff.net","subject":"Re: [PATCH v3 18/21] count-objects: recognize .bitmap in garbage-checking","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-12-07T16:38:11Z","receivedAt":"2013-12-07T16:38:11Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> From: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n>\n> Count-objects will report any \"garbage\" files in the packs\n> directory, including files whose extensions it does not\n> know (case 1), and files whose matching \".pack\" file is\n> missing (case 2).  Without having learned about \".bitmap\"\n> files, the current code reports all such files as garbage\n> (case 1), even if their pack exists. Instead, they should be\n> treated as case 2.\n>\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> Signed-off-by: Jeff King <peff@peff.net>\n\nReviewed-by: Thomas Rast <tr@thomasrast.ch>\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231727","messageId":"87mwkcipce.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","threadId":"35335","inReplyTo":"20131114124635.GS10757@sigill.intra.peff.net","subject":"Re: [PATCH v3 19/21] t: add basic bitmap functionality tests","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-12-07T16:43:29Z","receivedAt":"2013-12-07T16:43:29Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Now that we can read and write bitmaps, we can exercise them\n> with some basic functionality tests. These tests aren't\n> particularly useful for seeing the benefit, as the test\n> repo is too small for it to make a difference. However, we\n> can at least check that using bitmaps does not break anything.\n>\n> Signed-off-by: Jeff King <peff@peff.net>\n\nReviewed-by: Thomas Rast <tr@thomasrast.ch>\n\nOne nit:\n\n> +test_expect_success JGIT 'jgit can read our bitmaps' '\n> +\tgit clone . compat-us.git &&\n> +\t(\n> +\t\tcd compat-us.git &&\n\nThe name suggests a bare repo, but it is a full clone.  Not that it\nmatters.\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231728","messageId":"8761r0ioyo.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","threadId":"35335","inReplyTo":"20131114124834.GA11612@sigill.intra.peff.net","subject":"Re: [PATCH v3 20/21] t/perf: add tests for pack bitmaps","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-12-07T16:51:43Z","receivedAt":"2013-12-07T16:51:43Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> +test_perf 'simulated fetch' '\n> +\thave=$(git rev-list HEAD --until=1.week.ago -1) &&\n\nThis will give you HEAD if your GIT_PERF_LARGE_REPO hasn't seen any\nactivity lately.  I'd prefer something that always takes a fixed commit,\ne.g. HEAD~1000, keeping the perf test reproducible over time (not over\nchanging GIT_PERF_LARGE_REPO, of course).\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231729","messageId":"87wqjgha1b.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","threadId":"35335","inReplyTo":"20131114124855.GB11612@sigill.intra.peff.net","subject":"Re: [PATCH v3 21/21] pack-bitmap: implement optional name_hash cache","fromName":"Thomas Rast","fromEmail":"tr@thomasrast.ch","sentAt":"2013-12-07T16:59:28Z","receivedAt":"2013-12-07T16:59:28Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Test                      origin/master       HEAD^                      HEAD\n> -------------------------------------------------------------------------------------------------\n> 5310.2: repack to disk    36.81(37.82+1.43)   47.70(48.74+1.41) +29.6%   47.75(48.70+1.51) +29.7%\n> 5310.3: simulated clone   30.78(29.70+2.14)   1.08(0.97+0.10) -96.5%     1.07(0.94+0.12) -96.5%\n> 5310.4: simulated fetch   3.16(6.10+0.08)     3.54(10.65+0.06) +12.0%    1.70(3.07+0.06) -46.2%\n> 5310.6: partial bitmap    36.76(43.19+1.81)   6.71(11.25+0.76) -81.7%    4.08(6.26+0.46) -88.9%\n>\n> You can see that the time spent on an incremental fetch goes\n> down, as our delta heuristics are able to do their work.\n> And we save time on the partial bitmap clone for the same\n> reason.\n\nThe time now goes down across the board compared to master.  Good job!\n\n> Signed-off-by: Vicent Marti <tanoku@gmail.com>\n> Signed-off-by: Jeff King <peff@peff.net>\n\nReviewed-by: Thomas Rast <tr@thomasrast.ch>\n\n-- \nThomas Rast\ntr@thomasrast.ch\n"},{"id":"231741","messageId":"52A38AAA.7040308@gmail.com","threadId":"35335","inReplyTo":"20131203182133.GA21296@sigill.intra.peff.net","subject":"Re: [PATCH v3 10/21] pack-bitmap: add support for bitmap indexes","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2013-12-07T20:52:58Z","receivedAt":"2013-12-07T20:52:58Z","isPatch":true,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 03.12.2013 19:21, schrieb Jeff King:\n> On Tue, Dec 03, 2013 at 03:40:41PM +0100, Karsten Blees wrote:\n> \n>> IMO, trying to improve khash isn't worth the trouble.\n>>\n>> With smaller-than-pointer types, khash _may_ actually save a few bytes\n>> compared to hash.[ch] or hashmap.[ch]. E.g. a set of 'int's would be a\n>> perfect use case for khash. However, you're using bitmaps for that,\n>> and khash's predefined macros for int sets and maps have been removed\n>> from the git version.\n> \n> True, we are not using it for smaller-than-pointer sizes here. So it may\n> not be worth thinking about (I was considering more for the general\n> case). In most instances, though, we can shove the int bits into a\n> pointer with the right casting. So it probably isn't worth worrying\n> about (you may waste a few bytes, but probably not more than one word\n> per entry).\n> \n>> Using khash with pointers and larger-than-pointer types is just a\n>> waste of memory and performance, though. Hash tables are sparsely\n>> filled by design, and using large types means lots of large empty\n>> buckets. E.g. kh_resize requires almost four times the size or your\n>> data, and copies everything at least three times.\n> \n> I think the analysis is more complicated than that, and depends on what\n> you are storing. If you are storing something that is 1.5 times the size\n> of a pointer, it is more space efficient to just stick it in the hash\n> table than it is to have a separate pointer in the hash table (you pay\n> the load factor penalty only on the single pointer, but you've almost\n> doubled your total storage).\n> \n> But again, that's a more general argument. We're not storing anything of\n> that size here, and in fact I think we are just storing pointers\n> everywhere.\n> \n\nYes. With pointers, its a very close call regarding space. Let's take bitmap_index.bitmaps / struct stored_bitmap as an example. I'm assuming 8 byte pointers and an average load factor of 0.5.\n\nstruct stored_bitmap is already kind of an entry structure (key + value) allocated on the heap. This is a very common szenario, all hash.[ch] -> hashmap.[ch] conversions were that way.\n\nThe average khash memory requirement per entry is (on top of the data on the heap):\n\n  (key + value + flags) / load-factor = (8 + 8 + .25) / .5 = 32.5 bytes\n\nFor the hashmap conversion, I simply injected struct hashmap_entry into the existing structure, adding 12 bytes per entry. So the total hashmap memory per entry is:\n\n  hashmap_entry + (entry-pointer / load-factor) = 12 + (8 / .5) = 28 bytes\n\n>> Additionally, there are the obvious problems with khash's macro design\n>> (hard to read, impossible to debug, and each named instance increases\n>> executable size by ~4k).\n> \n> Yes, those are all downsides to macros. Type safety is one of the\n> upsides, though.\n>\n\nWell, khash_sha1 maps to void *, so I guess its not that important :-) But if you care about type safety, a common core implementation with a few small wrapper macros would be hugely preferable, don't you think?\n \n> Besides macros, I think the major difference between the two\n> implementations is open-addressing versus chaining. Especially for sets,\n> we've had good experiences with open-addressing by keeping the load\n> factor low (e.g., the one in object.c). As you note, the resizing\n> operation pays some penalty, but in most of our workloads it's largely\n> irrelevant compared to lookup times.\n>\n\nResizing is reasonably fast and straightforward for both open addressing and chaining. My 'copy three times' remark above was specific to the rehash-in-place stunt in kh_resize (if you haven't wondered what these 64 multi-statement-lines do, you probably don't wanna know... ;-)\n\nBut its still best if you know the target size and can prevent resizing altogether.\n\n> Chaining typically adds an extra layer of pointer-following to the\n> lookup, and I'd be worried about that loss of locality. It's true that\n> when you are storing a pointer you are already hurting locality to\n> follow the pointer to the key, but I don't know whether that means an\n> _extra_ layer doesn't still hurt more (and I really mean I don't know --\n> I'd be interested to see measurements).\n>\n\nThe locality advantage of open addressing _only_ applies when storing the data directly in the table, but AFAIK we're not doing that anywhere in git. Compared to open addressing with pointers, chaining in fact has _better_ locality.\n\nIterating a linked list (i.e. chaining) accesses exactly one (1.0) memory location per element.\n\nIterating an array of pointers (i.e. open addressing with linear probing) accesses ~1.125 memory locations per element (assuming 8 byte pointers and 64 byte cache lines, array[8] will be in a different cache line than array[0], thus the +.125).\n\nOf course, any implementation can add more pointer-following to that minimum. E.g. khash uses separate flags and keys arrays and triangular numbers as probe sequence, so its up to 3 memory locations per element.\n\nHashmap, on the other hand, allows you to inject it's per-entry data into the user data structure to prevent locality penalties (well, the user data structure could add more pointer-following, but that's not hashmap's fault).\n\n> In your implementation, it looks like you break even there because you\n> store the hash directly in the entry, and do a single-word compare (so\n> you avoid having to follow a pointer to the key in the common case\n> during lookup).\n\nMore importantly, it saves calling a function pointer with three parameters unless its a _very_ probable match.\n\n> But that also means you're paying extra to store the\n> hash. That probably makes sense for things like strings, where it takes\n> some effort to calculate the hash. But not necessarily for sha1s, where\n> looking at the hash is the same thing as looking at the key bytes (so\n> you are storing extra bytes,\n\nTrue. Its 4 bytes overhead for sha1s and probably ints (but even for ints you may want to wiggle the bits for better distribution). However, its an utterly necessary optimization for e.g. strings with common prefixes. Think of a Java project where all paths start with com/mycompany/myou/productname/component/... i.e. strcmp has to compare ~40 bytes before finding a difference.\n\n> and when you do have a hit on the stored\n> hash, which is just the first bytes of the sha1, you end up comparing\n> them again as part of the hashcmp.\n\nNot necessarily - I could skip the part that I've used as hash code...\n\n> The tradeoff is that in the\n> non-matching cases, you avoid an extra level of indirection).\n> \n> All of these are things I could see helping or hurting, depending on the\n> case. I'd really love to see more numbers. I tried timing your patch\n> below, but there was no interesting change. Mostly because the code path\n> you changed is not heavily exercised (the \"ext_index\" is for objects\n> that are not in the pack, and frequent packing keeps that number low).\n> \n> I'd be really interested to see if you could make the hash in object.c\n> faster. That one is a prominent lookup bottle-neck (e.g., for \"git\n> rev-list --objects --all\"); if you can make that faster, I would be very\n> convinced that your implementation is fast (note that I am not implying\n> the converse; if you cannot make it faster, that does not necessarily\n> mean your implementation sucks, but perhaps only that the existing one\n> has lots of type-specific optimizations which add up).\n> \n\nOk, let's see...\n\nHash tables only compare for equality, not for sorting, so the first thing that comes to mind is to compare 4 bytes at a time.\n\nThe hashmap version cuts quite a bit of code, but uses slightly more memory (~6 bytes per entry). Still todo: get_indexed_object now only returns the list heads, so callers need to follow the chain...\n\nThen I wrote a custom chaining version (i.e. no cached hash code, inlined hashcmp() etc.). The increased load-factor of 0.8 should fully compensate the additional next pointer in struct object. Ditto for get_indexed_object.\n\nFinally, just for fun, the khash version...\n\nYour move-to-front optimization seems highly specialized to lookup_object access patterns, and also breaks the assumption that concurrent lookup is thread safe. So I didn't bother to pimp hashmap and khash internals, and instead did the other measurements with and without move-to-front enabled. The hashmap move-to-front is via the existing API (remove + add), i.e. its true LRU instead of swap (same for custom-chaining).\n\nNumbers are in seconds, best of 10 runs of 'git rev-list --all --objects >/dev/null' on the linux kernel repo.\n\n fast  | move  ||  next  |  hash  | custom | khash  |\n hash  |  to   ||        |  map   |        |        |\n cmp   | front || (open) |(chain) |(chain) | (open) |\n=======+=======++========+========+========+========+\n  no   |  no   || 45.593 | 42.924 | 41.033 | 62.974 |\n  no   |  yes  || 41.017 | 43.095 | 40.478 |        |\n  yes  |  no   || 43.172 | 40.965 | 39.564 |        |\n  yes  |  yes  || 39.332 | 41.008 | 38.869 |        |\n\nThe fastest versions of each variant can be found here:\n\nhttps://github.com/kblees/git/commits/kb/optimize-lookup-object-next\nhttps://github.com/kblees/git/commits/kb/optimize-lookup-object-hashmap\nhttps://github.com/kblees/git/commits/kb/optimize-lookup-object-custom-chaining\nhttps://github.com/kblees/git/commits/kb/optimize-lookup-object-khash\n\n\n> Khash also has a lot of bloat (e.g., flags) that the one in object.c\n> does not have. If you do not care about deletion and are storing\n> something with a sentinel value (e.g., NULL for pointers), you can trim\n> quite a bit of fat.\n> \n\nActually, the lack of deletion in hash.[ch] was the reason I started this...one advantage of chaining is that delete can be implemented efficiently (O(1)) without affecting lookup / insert performance.\n\n>> Below is a patch that converts pack-bitmap.c to hashmap. Its not even\n>> longer than the khash version, and the hashmap API forces you to think\n>> about no-brainer improvements such as specifying an expected size or\n>> skipping duplicates checks where they aren't needed. I could do the\n>> same for pack-bitmap-write.c if you like.\n> \n> If it's not too much trouble, I'd be curious to measure the performance\n> impact on pack-bitmap-write.\n> \n\nI'll see what I can do\n\n>> Removes two unnecessary duplicates checks:\n>>  - we don't expect a pack index file to contain duplicate sha1's\n> \n> We don't expect them to, but it has happened (and caused bugs not too\n> long ago). What happens after your patch when there are duplicates?\n> \n\nBoth entries are added, and hashmap_get returns one of them at random. Re-adding duplicates checks is simple, though:\n\n- hashmap_add(...);\n+ if (hashmap_put(...))\n+   die(\"foo has duplicate o's\");\n\n>> +static struct ewah_bitmap *lookup_stored_bitmap_sha1(const unsigned char *sha1)\n>> +{\n>> +\tstruct stored_bitmap key, *st;\n>> +\thashmap_entry_init(&key, __kh_oid_hash(sha1));\n>> +\tst = hashmap_get(&bitmap_git.bitmaps, &key, sha1);\n>> +\tif (st)\n>> +\t\treturn lookup_stored_bitmap(st);\n>> +\treturn NULL;\n>> +}\n> \n> This interface looks odd to me. You create a fake stored_bitmap for the\n> hashmap_entry part of it, and then fill in the \"hash\" field by hashing\n> the sha1. And then pass the same sha1 in. I guess you are trying to\n> avoid the hash table knowing about the hash function at all, since it\n> just stores the hash for each entry already.\n> \n\nThe general case would be:\n\n struct stored_bitmap key;\n /* fill in hash code */\n hashmap_entry_init(&key, __kh_oid_hash(sha1));\n /* fill in key data */\n hashcpy(key.sha1, sha1);\n key.another_key_field = foobar;\n ...\n return hashmap_get(&bitmap_git.bitmaps, &key, NULL);\n\nIn this case, the sha1 is the only key data, and its larger that a word, so passing it as keydata parameter and checking for this in the cmp-function is slightly more efficient.\n\n> I guess that makes sense, but would not work for a system where you\n> wanted to get rid of the extra hash storage (but as I implied above, I\n> am not sure if it is helping or hurting for the sha1 case).\n> \n>> -\thash_pos = kh_put_sha1_pos(eindex->positions, object->sha1, &hash_ret);\n>> -\tif (hash_ret > 0) {\n>> -\t\tif (eindex->count >= eindex->alloc) {\n>> -\t\t\teindex->alloc = (eindex->alloc + 16) * 3 / 2;\n>> -\t\t\teindex->objects = xrealloc(eindex->objects,\n>> -\t\t\t\teindex->alloc * sizeof(struct object *));\n>> -\t\t\teindex->hashes = xrealloc(eindex->hashes,\n>> -\t\t\t\teindex->alloc * sizeof(uint32_t));\n>> -\t\t}\n>> -\n>> -\t\tbitmap_pos = eindex->count;\n>> -\t\teindex->objects[eindex->count] = object;\n>> -\t\teindex->hashes[eindex->count] = pack_name_hash(name);\n>> -\t\tkh_value(eindex->positions, hash_pos) = bitmap_pos;\n>> -\t\teindex->count++;\n>> -\t} else {\n>> -\t\tbitmap_pos = kh_value(eindex->positions, hash_pos);\n>> -\t}\n>> -\n>> -\treturn bitmap_pos + bitmap_git.pack->num_objects;\n>> +\tstruct ext_entry *e = xmalloc(sizeof(struct ext_entry));\n>> +\thashmap_entry_init(e, __kh_oid_hash(object->sha1));\n>> +\te->object = object;\n>> +\te->name_hash = pack_name_hash(name);\n>> +\te->nr = bitmap_git.ext_index.size;\n>> +\thashmap_add(&bitmap_git.ext_index, e);\n>> +\treturn e->nr + bitmap_git.pack->num_objects;\n> \n> One of the side effects of the current system is that the array\n> effectively works as a custom allocator, and we do not pay per-entry\n> malloc overhead. As I mentioned above, this is not that\n> heavily-exercised a code path, so it may not matter (and you can convert\n> it by explicitly using a custom allocator, at which point you could drop\n> the e->nr field entirely, I'd think).\n> \n\nRealloc doesn't produce stable pointers, so we'd have to use a custom slab-allocator such as in alloc.c. However, I tend to think of malloc as a separate problem. There are quite efficient heap implementations with near-zero memory overhead even for small objects out there. I wouldn't be surprised if some day allocating many small objects will be faster and more space-efficient than re-allocating big blocks. Then you'd have all your code base littered with custom allocators that actually slow down the application...\n\n>> @@ -584,17 +581,15 @@ static struct bitmap *find_objects(struct rev_info *revs,\n>>  static void show_extended_objects(struct bitmap *objects,\n>>  \t\t\t\t  show_reachable_fn show_reach)\n>>  {\n>> -\tstruct eindex *eindex = &bitmap_git.ext_index;\n>> -\tuint32_t i;\n>> -\n>> -\tfor (i = 0; i < eindex->count; ++i) {\n>> -\t\tstruct object *obj;\n>> +\tstruct ext_entry *e;\n>> +\tstruct hashmap_iter iter;\n>>  \n>> -\t\tif (!bitmap_get(objects, bitmap_git.pack->num_objects + i))\n>> +\tfor (e = hashmap_iter_first(&bitmap_git.ext_index, &iter); e;\n>> +\t     e = hashmap_iter_next(&iter)) {\n>> +\t\tif (!bitmap_get(objects, bitmap_git.pack->num_objects + e->nr))\n>>  \t\t\tcontinue;\n>>  \n>> -\t\tobj = eindex->objects[i];\n>> -\t\tshow_reach(obj->sha1, obj->type, 0, eindex->hashes[i], NULL, 0);\n>> +\t\tshow_reach(e->object->sha1, e->object->type, 0, e->name_hash, NULL, 0);\n> \n> Before we were iterating in eindex order. Now we are iterating in\n> hashmap order. That effects our traversal order, and ultimately our\n> output for something like \"rev-list\". We'd want to sort on e->nr (or\n> again, a custom allocator would make this go away because we would\n> iterate over the array).\n> \n\nI didn't think it would matter, as the caller emits the objects from the bitmapped pack ordered by type and pack-index...well, I'll try to think of something (probably a linked list between entries?)\n\n>> -\tfor (i = 0; i < eindex->count; ++i) {\n>> -\t\tif (eindex->objects[i]->type == type &&\n>> -\t\t\tbitmap_get(objects, bitmap_git.pack->num_objects + i))\n>> +\tfor (e = hashmap_iter_first(&bitmap_git.ext_index, &iter); e;\n>> +\t     e = hashmap_iter_next(&iter)) {\n>> +\t\tif (e->object->type == type &&\n>> +\t\t\tbitmap_get(objects, bitmap_git.pack->num_objects + e->nr))\n>>  \t\t\tcount++;\n> \n> Ditto here, but I do not think the order matters in this instance.\n> \n> -Peff\n> \n"},{"id":"232301","messageId":"20131221131502.GA10123@sigill.intra.peff.net","threadId":"35335","inReplyTo":"87zjock6if.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","subject":"Re: [PATCH v3 11/21] pack-objects: use bitmaps when packing objects","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-21T13:15:03Z","receivedAt":"2013-12-21T13:15:03Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Dec 07, 2013 at 04:47:20PM +0100, Thomas Rast wrote:\n\n> > +static off_t write_reused_pack(struct sha1file *f)\n> > +{\n> > +\tuint8_t buffer[8192];\n> \n> We usually just call this 'unsigned char'.  I can see why this would be\n> more portable, but git would already fall apart badly on an architecture\n> where char is not 8 bits.\n\nI think it's worth switching just for consistency with the rest of git.\nFixed.\n\n> } +\t\t\tpackfile_size = write_reused_pack(f);\n> } +\t\t\tif (!packfile_size)\n> } +\t\t\t\tdie_errno(\"failed to re-use existing pack\");\n> \n> So if you just died here, when the error happens, you could take the\n> chance to tell the user _which_ syscall failed.\n\nYeah, agreed (and especially the fact that we may get bogus errno\nvalues). Fixed.\n\n> Not your fault, but sha1write() is an odd function -- it purportedly is\n> \n>   int sha1write(struct sha1file *f, const void *buf, unsigned int count);\n> \n> but it can only return 0.  This goes back all the way to c38138c\n> (git-pack-objects: write the pack files with a SHA1 csum, 2005-06-26).\n\nIt looks like there's exactly one site that checks its return value, and\nit's just to die. We should drop the return value from sha1write\nentirely to make it clear that it dies on error. But that's orthogonal\nto this series.\n\n> > -static int add_object_entry(const unsigned char *sha1, enum object_type type,\n> > -\t\t\t    const char *name, int exclude)\n> > +static int add_object_entry_1(const unsigned char *sha1, enum object_type type,\n> > +\t\t\t      int flags, uint32_t name_hash,\n> > +\t\t\t      struct packed_git *found_pack, off_t found_offset)\n> [...]\n> This function makes my head spin, and you're indenting it yet another\n> level.\n\nYeah. In addition, the use of the \"flags\" here is somewhat questionable.\nWe use them for internal values in the call from add_object_entry to\nadd_object_entry_1. But we also pass the latter as a traversal callback,\nmeaning that what it would get in \"flags\" is totally different. It's not\nactually a bug in the current code, since the bitmap traversal always\npasses empty flags, but it's still rather confusing.\n\n> If it's not too much work, can you split it into the three parts that it\n> really is?  IIUC it boils down to\n> \n>   do we have this already?\n>       possibly apply 'exclude', then return\n>   are we coming from a call path that doesn't tell us which pack to take\n>   it from?\n>       find _all_ instances in packs\n>       check if any of them are local .keep packs\n>           if so, return\n>   construct a packlist entry to taste\n\nI did this split. By itself, I was on the fence, as there are actually\nsome interdependencies between the three parts that make it hairy.\n\nBut then I realized that instead of making the weird relationship\nbetween add_object_entry and add_object_entry_1, we can simply make a\nnew function that composes the functions differently.\n\nSo we get:\n\n+static int add_object_entry_from_bitmap(const unsigned char *sha1,\n+\t\t\t\t\tenum object_type type,\n+\t\t\t\t\tint flags, uint32_t name_hash,\n+\t\t\t\t\tstruct packed_git *pack, off_t offset)\n+{\n+\tuint32_t index_pos;\n+\n+\tif (have_duplicate_entry(sha1, 0, &index_pos))\n+\t\treturn 0;\n+\n+\tcreate_object_entry(sha1, type, name_hash, 0, 0, index_pos, pack, offset);\n+\n+\tdisplay_progress(progress_state, to_pack.nr_objects);\n+\treturn 1;\n+}\n[...]\n+\ttraverse_bitmap_commit_list(&add_object_entry_from_bitmap);\n\nwhich makes much more sense.\n\n-Peff\n"},{"id":"232302","messageId":"20131221131705.GB10123@sigill.intra.peff.net","threadId":"35335","inReplyTo":"87d2l8k4es.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","subject":"Re: [PATCH v3 13/21] pack-objects: implement bitmap writing","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-21T13:17:05Z","receivedAt":"2013-12-21T13:17:05Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Dec 07, 2013 at 05:32:43PM +0100, Thomas Rast wrote:\n\n> > +pack.writebitmaps::\n> > +\tWhen true, git will write a bitmap index when packing all\n> > +\tobjects to disk (e.g., as when `git repack -a` is run).  This\n>                                ^^\n> \n> Doesn't sound right in my ears.  Remove the \"as\"?\n\nIt was meant to be \"such as when...\", but I think it is just as clear to\ndrop it. Will do.\n\n-Peff\n"},{"id":"232303","messageId":"20131221132246.GC10123@sigill.intra.peff.net","threadId":"35335","inReplyTo":"87mwkcipce.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","subject":"Re: [PATCH v3 19/21] t: add basic bitmap functionality tests","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-21T13:22:46Z","receivedAt":"2013-12-21T13:22:46Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Dec 07, 2013 at 05:43:29PM +0100, Thomas Rast wrote:\n\n> One nit:\n> \n> > +test_expect_success JGIT 'jgit can read our bitmaps' '\n> > +\tgit clone . compat-us.git &&\n> > +\t(\n> > +\t\tcd compat-us.git &&\n> \n> The name suggests a bare repo, but it is a full clone.  Not that it\n> matters.\n\nIt was originally supposed to be a bare repo, but I had trouble\nconvincing jgit to actually run in a bare repo. The solution was to\nswitch to a non-bare one. :)\n\nBut obviously I forgot to update the name. I agree it's better to use a\nmore obvious name. Fixed.\n\n-Peff\n"},{"id":"232304","messageId":"20131221134017.GD10123@sigill.intra.peff.net","threadId":"35335","inReplyTo":"8761r0ioyo.fsf@linux-1gf2.Speedport_W723_V_Typ_A_1_00_098","subject":"Re: [PATCH v3 20/21] t/perf: add tests for pack bitmaps","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-21T13:40:17Z","receivedAt":"2013-12-21T13:40:17Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Dec 07, 2013 at 05:51:43PM +0100, Thomas Rast wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > +test_perf 'simulated fetch' '\n> > +\thave=$(git rev-list HEAD --until=1.week.ago -1) &&\n> \n> This will give you HEAD if your GIT_PERF_LARGE_REPO hasn't seen any\n> activity lately.  I'd prefer something that always takes a fixed commit,\n> e.g. HEAD~1000, keeping the perf test reproducible over time (not over\n> changing GIT_PERF_LARGE_REPO, of course).\n\nGood point. I just did a \"git pull\" on the kernel before my tests, but\nof course the same test with my 3-week-stale repo would be a lot less\ninteresting.\n\nI tried to figure out what is the right N for HEAD~N to represent a\nweek's worth of commits. It turns out the kernel is really bursty\ndepending on where they are in the cycle. Sometimes N is 50, and\nsometimes 200. I'm picking 100. It doesn't have to be a week; the main\ngoal was to approximate \"here's how it performs if you don't repack for\na week\", so anything in the ballpark is fine.\n\n-Peff\n"}]}