{"thread":{"id":"34271","subject":"[PATCH 00/16] Speed up Counting Objects with bitmap data","startedAt":"2013-06-24T23:22:57Z","lastAt":"2013-07-07T17:27:01Z","messageCount":64,"participants":["Vicent Marti","Junio C Hamano","Shawn Pearce","Ramkumar Ramachandra","Peter Krefting","Vicent Martí","Thomas Rast","Jeff King","Colby Ranger"],"isPatch":true,"patchVersion":1,"patchTotal":16},"messages":[{"id":"221898","messageId":"1372116193-32762-1-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":null,"subject":"[PATCH 00/16] Speed up Counting Objects with bitmap data","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:22:57Z","receivedAt":"2013-06-24T23:22:57Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"Hello friends and enemies from the lovevely Git Mailing list.\n\nI bring to you a patch series that implement a quite interesting performance\noptimization: the removal of the \"Counting Objects\" phase during `pack-objects`\nby using a pre-computed bitmap to find the reachable objects in the packfile.\n\nAs you probably know, Shawn Pearce designed this approach a few months ago and\nimplemented it for JGit, with very exciting results.\n\nThis is a not-so-straightforward port of his original design: The general approach\nis the same, but unfortunately we were not able to re-use JGit's original on-disk\nformat for the `.bitmap` files.\n\nThere is a full technical spec for the new format (v2) in patch 09, including\nbenchmarks and rationale for the new design. The gist of it is that JGit's\noriginal format is not `mmap`able (JGit tends to not mmap anything), and that\nbecomes very costly in practice with `upload-pack`, which spawns a new process\nfor every upload.\n\nThe header and metadata for both formats are however compatible, so it should be\ntrivial to update JGit to read/write this format too. I intend to do this on the\ncoming weeks, and I also hope that the v2 implementation will be slightly faster\nthan the actual, even with the shortcomings of the JVM.\n\nThe patch series, although massive, is rather straightforward.\n\nMost of the patches are isolated refactorings that enable access to a few functions\nthat were previously hidden (re. packfile data). These functions are needed for\nreading and writing the bitmap indexes.\n\nPatch 03 is worth noting because it implements a performance optimization for\n`pack-objects` which isn't particularly good in normal invocations (~10% speed up)\nbut that will show great benefits in later patches when it comes to writing the\nbitmap indexes.\n\nPatch 10 is the core of the series, implementing the actual loading of bitmap indexes\nand optimizing the Counting Objects phase of `pack-objects`. Like with every other\npatch that offers performance improvements, sample benchmarks are provided (spoiler:\nthey are pretty fucking cool).\n\nPatch 11 and 16 are samples of using the new Bitmap traversal API to speed up other\nparts of Git (`rev-list --objects` and `rev-list --count`, respectively).\n\nPatch 12, 13 and 15 implement the actual writing of bitmap indexes. Like JGit, patch\n12 enables writing a bitmap index as part of the `pack-objects` process (and hence\nas part of a normal `gc` run). On top of that, I implemented a new plumbing command\nin patch 15 that allows to write bitmap indexes for already-existing packfiles.\n\nI'd love your feedback on the design and implementation of this feature. I deem it\nrather stable, as we've been testing it on production on the world's largest Git\nhost (Git Hub Dot Com The Web Site) with good results, so I'd love it to have it\nupstreamed on Core Git.\n\nStrawberry kisses,\nvmg\n\nJeff King (1):\n  list-objects: mark tree as unparsed when we free its buffer\n\nVicent Marti (15):\n  sha1_file: refactor into `find_pack_object_pos`\n  pack-objects: use a faster hash table\n  pack-objects: make `pack_name_hash` global\n  revision: allow setting custom limiter function\n  sha1_file: export `git_open_noatime`\n  compat: add endinanness helpers\n  ewah: compressed bitmap implementation\n  documentation: add documentation for the bitmap format\n  pack-objects: use bitmaps when packing objects\n  rev-list: add bitmap mode to speed up lists\n  pack-objects: implement bitmap writing\n  repack: consider bitmaps when performing repacks\n  sha1_file: implement `nth_packed_object_info`\n  write-bitmap: implement new git command to write bitmaps\n  rev-list: Optimize --count using bitmaps too\n\n Documentation/technical/bitmap-format.txt |  235 ++++++++\n Makefile                                  |   11 +\n builtin.h                                 |    1 +\n builtin/pack-objects.c                    |  362 +++++++-----\n builtin/pack-objects.h                    |   33 ++\n builtin/rev-list.c                        |   35 +-\n builtin/write-bitmap.c                    |  256 +++++++++\n cache.h                                   |    5 +\n ewah/bitmap.c                             |  229 ++++++++\n ewah/ewah_bitmap.c                        |  703 ++++++++++++++++++++++++\n ewah/ewah_io.c                            |  199 +++++++\n ewah/ewah_rlw.c                           |  124 +++++\n ewah/ewok.h                               |  194 +++++++\n ewah/ewok_rlw.h                           |  114 ++++\n git-compat-util.h                         |   28 +\n git-repack.sh                             |   10 +-\n git.c                                     |    1 +\n khash.h                                   |  329 +++++++++++\n list-objects.c                            |    1 +\n pack-bitmap-write.c                       |  520 ++++++++++++++++++\n pack-bitmap.c                             |  855 +++++++++++++++++++++++++++++\n pack-bitmap.h                             |   64 +++\n pack-write.c                              |    2 +\n revision.c                                |    5 +\n revision.h                                |    2 +\n sha1_file.c                               |   57 +-\n 26 files changed, 4212 insertions(+), 163 deletions(-)\n create mode 100644 Documentation/technical/bitmap-format.txt\n create mode 100644 builtin/pack-objects.h\n create mode 100644 builtin/write-bitmap.c\n create mode 100644 ewah/bitmap.c\n create mode 100644 ewah/ewah_bitmap.c\n create mode 100644 ewah/ewah_io.c\n create mode 100644 ewah/ewah_rlw.c\n create mode 100644 ewah/ewok.h\n create mode 100644 ewah/ewok_rlw.h\n create mode 100644 khash.h\n create mode 100644 pack-bitmap-write.c\n create mode 100644 pack-bitmap.c\n create mode 100644 pack-bitmap.h\n\n-- \n1.7.9.5\n"},{"id":"221899","messageId":"1372116193-32762-2-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 01/16] list-objects: mark tree as unparsed when we free its buffer","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:22:58Z","receivedAt":"2013-06-24T23:22:58Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"From: Jeff King <peff@peff.net>\n\nWe free the tree buffer during traversal to save memory.\nHowever, we do not reset the \"parsed\" flag, which leaves a\nlandmine for the next person to use the tree. When they call\nparse_tree it will do nothing, and they will segfault when\nthey try to access the buffer.\n\nThis hasn't mattered until now because most rev-list\ntraversals would exit the program immediately afterwards,\nbut the bitmap writer wants to access the trees twice.\n---\n list-objects.c |    1 +\n 1 file changed, 1 insertion(+)\n\ndiff --git a/list-objects.c b/list-objects.c\nindex 3dd4a96..1251180 100644\n--- a/list-objects.c\n+++ b/list-objects.c\n@@ -125,6 +125,7 @@ static void process_tree(struct rev_info *revs,\n \tstrbuf_setlen(base, baselen);\n \tfree(tree->buffer);\n \ttree->buffer = NULL;\n+\ttree->object.parsed = 0;\n }\n \n static void mark_edge_parents_uninteresting(struct commit *commit,\n-- \n1.7.9.5\n"},{"id":"221900","messageId":"1372116193-32762-3-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 02/16] sha1_file: refactor into `find_pack_object_pos`","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:22:59Z","receivedAt":"2013-06-24T23:22:59Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"Looking up the offset in the packfile for a given SHA1 involves the\nfollowing:\n\n\t- Finding the position in the index for the given SHA1\n\t- Accessing the offset cache in the index for the found position\n\nThere are cases however where we'd like to find the position of a SHA1\nin the index without looking up the packfile offset (e.g. when accessing\ninformation that has been indexed based on index offsets).\n\nThis refactoring implements `find_pack_object_pos`, returning the\nposition in the index, and re-implements `find_pack_entry_one`(returning\nthe actual offset in the packfile) to use the new function.\n---\n cache.h     |    1 +\n sha1_file.c |   27 +++++++++++++++++----------\n 2 files changed, 18 insertions(+), 10 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex ec8240f..a29645e 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1101,6 +1101,7 @@ extern void clear_delta_base_cache(void);\n extern struct packed_git *add_packed_git(const char *, int, int);\n extern const unsigned char *nth_packed_object_sha1(struct packed_git *, uint32_t);\n extern off_t nth_packed_object_offset(const struct packed_git *, uint32_t);\n+extern int find_pack_entry_pos(const unsigned char *sha1, struct packed_git *p);\n extern off_t find_pack_entry_one(const unsigned char *, struct packed_git *);\n extern int is_pack_valid(struct packed_git *);\n extern void *unpack_entry(struct packed_git *, off_t, enum object_type *, unsigned long *);\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 0af19c0..371e295 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -2205,8 +2205,7 @@ off_t nth_packed_object_offset(const struct packed_git *p, uint32_t n)\n \t}\n }\n \n-off_t find_pack_entry_one(const unsigned char *sha1,\n-\t\t\t\t  struct packed_git *p)\n+int find_pack_entry_pos(const unsigned char *sha1, struct packed_git *p)\n {\n \tconst uint32_t *level1_ofs = p->index_data;\n \tconst unsigned char *index = p->index_data;\n@@ -2219,7 +2218,7 @@ off_t find_pack_entry_one(const unsigned char *sha1,\n \n \tif (!index) {\n \t\tif (open_pack_index(p))\n-\t\t\treturn 0;\n+\t\t\treturn -1;\n \t\tlevel1_ofs = p->index_data;\n \t\tindex = p->index_data;\n \t}\n@@ -2243,12 +2242,9 @@ off_t find_pack_entry_one(const unsigned char *sha1,\n \n \tif (use_lookup < 0)\n \t\tuse_lookup = !!getenv(\"GIT_USE_LOOKUP\");\n+\n \tif (use_lookup) {\n-\t\tint pos = sha1_entry_pos(index, stride, 0,\n-\t\t\t\t\t lo, hi, p->num_objects, sha1);\n-\t\tif (pos < 0)\n-\t\t\treturn 0;\n-\t\treturn nth_packed_object_offset(p, pos);\n+\t\treturn sha1_entry_pos(index, stride, 0, lo, hi, p->num_objects, sha1);\n \t}\n \n \tdo {\n@@ -2259,13 +2255,24 @@ off_t find_pack_entry_one(const unsigned char *sha1,\n \t\t\tprintf(\"lo %u hi %u rg %u mi %u\\n\",\n \t\t\t       lo, hi, hi - lo, mi);\n \t\tif (!cmp)\n-\t\t\treturn nth_packed_object_offset(p, mi);\n+\t\t\treturn mi;\n \t\tif (cmp > 0)\n \t\t\thi = mi;\n \t\telse\n \t\t\tlo = mi+1;\n \t} while (lo < hi);\n-\treturn 0;\n+\n+\treturn -1;\n+}\n+\n+off_t find_pack_entry_one(const unsigned char *sha1, struct packed_git *p)\n+{\n+\tint pos;\n+\n+\tif ((pos = find_pack_entry_pos(sha1, p)) < 0)\n+\t\treturn 0;\n+\n+\treturn nth_packed_object_offset(p, (uint32_t)pos);\n }\n \n int is_pack_valid(struct packed_git *p)\n-- \n1.7.9.5\n"},{"id":"221902","messageId":"1372116193-32762-4-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 03/16] pack-objects: use a faster hash table","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:00Z","receivedAt":"2013-06-24T23:23:00Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"In a normal pack-objects invocations runtime is dominated by the\ncounting objects phase and the time spent performing hash operations\nwith the `object_entry` structs is not particularly relevant.\n\nIf the Counting Objects phase however gets optimized to perform batch\ninsertions (like it is the goal of this patch series), the current\nhash table implementation starts to bottleneck in both insertion times\nand memory reallocations caused by the fast growth of the hash table.\n\nFor instance, these are function timings in a hotspot profiling of a\npack-objects run:\n\n\tlocate_object_entry_hash :: 563.935ms\n\thashcmp :: 195.991ms\n\tadd_object_entry_1 :: 47.962ms\n\nThis commit brings `khash.h`, a header only hash table implementation\nthat while remaining rather simple (uses quadratic probing and a\nstandard hashing scheme) and self-contained, offers a significant\nperformance improvement in both insertion and lookup times.\n\n`khash` is a generic hash table implementation that can be 'templated'\nfor any given type while maintaining good performance by using preprocessor\nmacros. This specific version has been modified to define by default a\n`khash_sha1` type, a map of SHA1s (const unsigned char[20]) to void *\npointers.\n\nWhen replacing the old hash table implementation in `pack-objects` with\nthe khash_sha1 table, the insertion time is greatly reduced:\n\n\tkh_put_sha1 :: 284.011ms\n\tadd_object_entry_1 : 36.06ms\n\thashcmp :: 24.045ms\n\nThis reduction of more than 50% in the insertion and lookup times,\nalthough nice, is not particularly noticeable for normal `pack-objects`\noperation: `pack-objects` performs massive batch insertions and\nrelatively few lookups, so `khash` doesn't get a chance to shine here.\n\nThe big win here, however, is in the massively reduced amount of hash\ncollisions (as you can see from the huge reduction of time spent in\n`hashcmp` after the change). These greatly improved lookup times\nwill result critical once we implement the writing algorithm for bitmap\nindxes in a later patch of this series.\n\nThe bitmap writing phase for a repository like `linux` requires several\nmillion table lookups: using the new hash table saves 1min and 20s from\na `pack-objects` invocation that also writes out bitmaps.\n---\n Makefile               |    1 +\n builtin/pack-objects.c |  210 ++++++++++++++++---------------\n khash.h                |  329 ++++++++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 441 insertions(+), 99 deletions(-)\n create mode 100644 khash.h\n\ndiff --git a/Makefile b/Makefile\nindex 79f961e..e01506d 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -683,6 +683,7 @@ LIB_H += grep.h\n LIB_H += hash.h\n LIB_H += help.h\n LIB_H += http.h\n+LIB_H += khash.h\n LIB_H += kwset.h\n LIB_H += levenshtein.h\n LIB_H += line-log.h\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex f069462..fc12df8 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -18,6 +18,7 @@\n #include \"refs.h\"\n #include \"streaming.h\"\n #include \"thread-utils.h\"\n+#include \"khash.h\"\n \n static const char *pack_usage[] = {\n \tN_(\"git pack-objects --stdout [options...] [< ref-list | < object-list]\"),\n@@ -57,7 +58,7 @@ struct object_entry {\n  * in the order we see -- typically rev-list --objects order that gives us\n  * nice \"minimum seek\" order.\n  */\n-static struct object_entry *objects;\n+static struct object_entry **objects;\n static struct pack_idx_entry **written_list;\n static uint32_t nr_objects, nr_alloc, nr_result, nr_written;\n \n@@ -93,8 +94,8 @@ static unsigned long window_memory_limit = 0;\n  * to help looking up the entry by object name.\n  * This hashtable is built after all the objects are seen.\n  */\n-static int *object_ix;\n-static int object_ix_hashsz;\n+\n+static khash_sha1 *packed_objects;\n static struct object_entry *locate_object_entry(const unsigned char *sha1);\n \n /*\n@@ -104,6 +105,35 @@ static uint32_t written, written_delta;\n static uint32_t reused, reused_delta;\n \n \n+static struct object_slab {\n+\tstruct object_slab *next;\n+\tuint32_t count;\n+\tstruct object_entry data[0];\n+} *slab;\n+\n+#define OBJECTS_PER_SLAB 2048\n+\n+static void push_slab(void)\n+{\n+\tstatic const size_t slab_size =\n+\t\t(OBJECTS_PER_SLAB * sizeof(struct object_entry)) + sizeof(struct object_slab);\n+\n+\tstruct object_slab *new_slab = calloc(1, slab_size);\n+\n+\tnew_slab->count = 0;\n+\tnew_slab->next = slab;\n+\tslab = new_slab;\n+}\n+\n+static struct object_entry *alloc_object_entry(void)\n+{\n+\tif (!slab || slab->count == OBJECTS_PER_SLAB)\n+\t\tpush_slab();\n+\n+\treturn &slab->data[slab->count++];\n+}\n+\n+\n static void *get_delta(struct object_entry *entry)\n {\n \tunsigned long size, base_size, delta_size;\n@@ -635,10 +665,10 @@ static struct object_entry **compute_write_order(void)\n \tstruct object_entry **wo = xmalloc(nr_objects * sizeof(*wo));\n \n \tfor (i = 0; i < nr_objects; i++) {\n-\t\tobjects[i].tagged = 0;\n-\t\tobjects[i].filled = 0;\n-\t\tobjects[i].delta_child = NULL;\n-\t\tobjects[i].delta_sibling = NULL;\n+\t\tobjects[i]->tagged = 0;\n+\t\tobjects[i]->filled = 0;\n+\t\tobjects[i]->delta_child = NULL;\n+\t\tobjects[i]->delta_sibling = NULL;\n \t}\n \n \t/*\n@@ -647,7 +677,7 @@ static struct object_entry **compute_write_order(void)\n \t * recency order.\n \t */\n \tfor (i = nr_objects; i > 0;) {\n-\t\tstruct object_entry *e = &objects[--i];\n+\t\tstruct object_entry *e = objects[--i];\n \t\tif (!e->delta)\n \t\t\tcontinue;\n \t\t/* Mark me as the first child */\n@@ -665,9 +695,9 @@ static struct object_entry **compute_write_order(void)\n \t * we see a tagged tip.\n \t */\n \tfor (i = wo_end = 0; i < nr_objects; i++) {\n-\t\tif (objects[i].tagged)\n+\t\tif (objects[i]->tagged)\n \t\t\tbreak;\n-\t\tadd_to_write_order(wo, &wo_end, &objects[i]);\n+\t\tadd_to_write_order(wo, &wo_end, objects[i]);\n \t}\n \tlast_untagged = i;\n \n@@ -675,35 +705,35 @@ static struct object_entry **compute_write_order(void)\n \t * Then fill all the tagged tips.\n \t */\n \tfor (; i < nr_objects; i++) {\n-\t\tif (objects[i].tagged)\n-\t\t\tadd_to_write_order(wo, &wo_end, &objects[i]);\n+\t\tif (objects[i]->tagged)\n+\t\t\tadd_to_write_order(wo, &wo_end, objects[i]);\n \t}\n \n \t/*\n \t * And then all remaining commits and tags.\n \t */\n \tfor (i = last_untagged; i < nr_objects; i++) {\n-\t\tif (objects[i].type != OBJ_COMMIT &&\n-\t\t    objects[i].type != OBJ_TAG)\n+\t\tif (objects[i]->type != OBJ_COMMIT &&\n+\t\t    objects[i]->type != OBJ_TAG)\n \t\t\tcontinue;\n-\t\tadd_to_write_order(wo, &wo_end, &objects[i]);\n+\t\tadd_to_write_order(wo, &wo_end, objects[i]);\n \t}\n \n \t/*\n \t * And then all the trees.\n \t */\n \tfor (i = last_untagged; i < nr_objects; i++) {\n-\t\tif (objects[i].type != OBJ_TREE)\n+\t\tif (objects[i]->type != OBJ_TREE)\n \t\t\tcontinue;\n-\t\tadd_to_write_order(wo, &wo_end, &objects[i]);\n+\t\tadd_to_write_order(wo, &wo_end, objects[i]);\n \t}\n \n \t/*\n \t * Finally all the rest in really tight order\n \t */\n \tfor (i = last_untagged; i < nr_objects; i++) {\n-\t\tif (!objects[i].filled)\n-\t\t\tadd_family_to_write_order(wo, &wo_end, &objects[i]);\n+\t\tif (!objects[i]->filled)\n+\t\t\tadd_family_to_write_order(wo, &wo_end, objects[i]);\n \t}\n \n \tif (wo_end != nr_objects)\n@@ -804,61 +834,26 @@ static void write_pack_file(void)\n \t\tnr_remaining -= nr_written;\n \t} while (nr_remaining && i < nr_objects);\n \n+\tstop_progress(&progress_state);\n+\n \tfree(written_list);\n \tfree(write_order);\n-\tstop_progress(&progress_state);\n \tif (written != nr_result)\n \t\tdie(\"wrote %\"PRIu32\" objects while expecting %\"PRIu32,\n \t\t\twritten, nr_result);\n }\n \n-static int locate_object_entry_hash(const unsigned char *sha1)\n-{\n-\tint i;\n-\tunsigned int ui;\n-\tmemcpy(&ui, sha1, sizeof(unsigned int));\n-\ti = ui % object_ix_hashsz;\n-\twhile (0 < object_ix[i]) {\n-\t\tif (!hashcmp(sha1, objects[object_ix[i] - 1].idx.sha1))\n-\t\t\treturn i;\n-\t\tif (++i == object_ix_hashsz)\n-\t\t\ti = 0;\n-\t}\n-\treturn -1 - i;\n-}\n-\n static struct object_entry *locate_object_entry(const unsigned char *sha1)\n {\n-\tint i;\n+\tkhiter_t pos = kh_get_sha1(packed_objects, sha1);\n \n-\tif (!object_ix_hashsz)\n-\t\treturn NULL;\n+\tif (pos < kh_end(packed_objects)) {\n+\t\treturn kh_value(packed_objects, pos);\n+\t}\n \n-\ti = locate_object_entry_hash(sha1);\n-\tif (0 <= i)\n-\t\treturn &objects[object_ix[i]-1];\n \treturn NULL;\n }\n \n-static void rehash_objects(void)\n-{\n-\tuint32_t i;\n-\tstruct object_entry *oe;\n-\n-\tobject_ix_hashsz = nr_objects * 3;\n-\tif (object_ix_hashsz < 1024)\n-\t\tobject_ix_hashsz = 1024;\n-\tobject_ix = xrealloc(object_ix, sizeof(int) * object_ix_hashsz);\n-\tmemset(object_ix, 0, sizeof(int) * object_ix_hashsz);\n-\tfor (i = 0, oe = objects; i < nr_objects; i++, oe++) {\n-\t\tint ix = locate_object_entry_hash(oe->idx.sha1);\n-\t\tif (0 <= ix)\n-\t\t\tcontinue;\n-\t\tix = -1 - ix;\n-\t\tobject_ix[ix] = i + 1;\n-\t}\n-}\n-\n static unsigned name_hash(const char *name)\n {\n \tunsigned c, hash = 0;\n@@ -901,19 +896,19 @@ static int no_try_delta(const char *path)\n \treturn 0;\n }\n \n-static int add_object_entry(const unsigned char *sha1, enum object_type type,\n-\t\t\t    const char *name, int exclude)\n+static int add_object_entry_1(const unsigned char *sha1, enum object_type type,\n+\t\t\t    uint32_t hash, int exclude, struct packed_git *found_pack,\n+\t\t\t\toff_t found_offset)\n {\n \tstruct object_entry *entry;\n-\tstruct packed_git *p, *found_pack = NULL;\n-\toff_t found_offset = 0;\n-\tint ix;\n-\tunsigned hash = name_hash(name);\n+\tstruct packed_git *p;\n+\tkhiter_t ix;\n+\tint hash_ret;\n \n-\tix = nr_objects ? locate_object_entry_hash(sha1) : -1;\n-\tif (ix >= 0) {\n+\tix = kh_put_sha1(packed_objects, sha1, &hash_ret);\n+\tif (hash_ret == 0) {\n \t\tif (exclude) {\n-\t\t\tentry = objects + object_ix[ix] - 1;\n+\t\t\tentry = kh_value(packed_objects, ix);\n \t\t\tif (!entry->preferred_base)\n \t\t\t\tnr_result--;\n \t\t\tentry->preferred_base = 1;\n@@ -921,38 +916,42 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n \t\treturn 0;\n \t}\n \n-\tif (!exclude && local && has_loose_object_nonlocal(sha1))\n+\tif (!exclude && local && has_loose_object_nonlocal(sha1)) {\n+\t\tkh_del_sha1(packed_objects, ix);\n \t\treturn 0;\n+\t}\n \n-\tfor (p = packed_git; p; p = p->next) {\n-\t\toff_t offset = find_pack_entry_one(sha1, p);\n-\t\tif (offset) {\n-\t\t\tif (!found_pack) {\n-\t\t\t\tif (!is_pack_valid(p)) {\n-\t\t\t\t\twarning(\"packfile %s cannot be accessed\", p->pack_name);\n-\t\t\t\t\tcontinue;\n+\tif (!found_pack) {\n+\t\tfor (p = packed_git; p; p = p->next) {\n+\t\t\toff_t offset = find_pack_entry_one(sha1, p);\n+\t\t\tif (offset) {\n+\t\t\t\tif (!found_pack) {\n+\t\t\t\t\tif (!is_pack_valid(p)) {\n+\t\t\t\t\t\twarning(\"packfile %s cannot be accessed\", p->pack_name);\n+\t\t\t\t\t\tcontinue;\n+\t\t\t\t\t}\n+\t\t\t\t\tfound_offset = offset;\n+\t\t\t\t\tfound_pack = p;\n+\t\t\t\t}\n+\t\t\t\tif (exclude)\n+\t\t\t\t\tbreak;\n+\t\t\t\tif (incremental ||\n+\t\t\t\t\t(local && !p->pack_local) ||\n+\t\t\t\t\t(ignore_packed_keep && p->pack_local && p->pack_keep)) {\n+\t\t\t\t\tkh_del_sha1(packed_objects, ix);\n+\t\t\t\t\treturn 0;\n \t\t\t\t}\n-\t\t\t\tfound_offset = offset;\n-\t\t\t\tfound_pack = p;\n \t\t\t}\n-\t\t\tif (exclude)\n-\t\t\t\tbreak;\n-\t\t\tif (incremental)\n-\t\t\t\treturn 0;\n-\t\t\tif (local && !p->pack_local)\n-\t\t\t\treturn 0;\n-\t\t\tif (ignore_packed_keep && p->pack_local && p->pack_keep)\n-\t\t\t\treturn 0;\n \t\t}\n \t}\n \n \tif (nr_objects >= nr_alloc) {\n \t\tnr_alloc = (nr_alloc  + 1024) * 3 / 2;\n-\t\tobjects = xrealloc(objects, nr_alloc * sizeof(*entry));\n+\t\tobjects = xrealloc(objects, nr_alloc * sizeof(struct object_entry *));\n \t}\n \n-\tentry = objects + nr_objects++;\n-\tmemset(entry, 0, sizeof(*entry));\n+\tentry = alloc_object_entry();\n+\n \thashcpy(entry->idx.sha1, sha1);\n \tentry->hash = hash;\n \tif (type)\n@@ -966,19 +965,30 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n \t\tentry->in_pack_offset = found_offset;\n \t}\n \n-\tif (object_ix_hashsz * 3 <= nr_objects * 4)\n-\t\trehash_objects();\n-\telse\n-\t\tobject_ix[-1 - ix] = nr_objects;\n+\tkh_value(packed_objects, ix) = entry;\n+\tkh_key(packed_objects, ix) = entry->idx.sha1;\n+\tobjects[nr_objects++] = entry;\n \n \tdisplay_progress(progress_state, nr_objects);\n \n-\tif (name && no_try_delta(name))\n-\t\tentry->no_try_delta = 1;\n-\n \treturn 1;\n }\n \n+static int add_object_entry(const unsigned char *sha1, enum object_type type,\n+\t\t\t    const char *name, int exclude)\n+{\n+\tif (add_object_entry_1(sha1, type, name_hash(name), exclude, NULL, 0)) {\n+\t\tstruct object_entry *entry = objects[nr_objects - 1];\n+\n+\t\tif (name && no_try_delta(name))\n+\t\t\tentry->no_try_delta = 1;\n+\n+\t\treturn 1;\n+\t}\n+\n+\treturn 0;\n+}\n+\n struct pbase_tree_cache {\n \tunsigned char sha1[20];\n \tint ref;\n@@ -1404,7 +1414,7 @@ static void get_object_details(void)\n \n \tsorted_by_offset = xcalloc(nr_objects, sizeof(struct object_entry *));\n \tfor (i = 0; i < nr_objects; i++)\n-\t\tsorted_by_offset[i] = objects + i;\n+\t\tsorted_by_offset[i] = objects[i];\n \tqsort(sorted_by_offset, nr_objects, sizeof(*sorted_by_offset), pack_offset_sort);\n \n \tfor (i = 0; i < nr_objects; i++) {\n@@ -2063,7 +2073,7 @@ static void prepare_pack(int window, int depth)\n \tnr_deltas = n = 0;\n \n \tfor (i = 0; i < nr_objects; i++) {\n-\t\tstruct object_entry *entry = objects + i;\n+\t\tstruct object_entry *entry = objects[i];\n \n \t\tif (entry->delta)\n \t\t\t/* This happens if we decided to reuse existing\n@@ -2574,6 +2584,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (progress && all_progress_implied)\n \t\tprogress = 2;\n \n+\tpacked_objects = kh_init_sha1();\n \tprepare_packed_git();\n \n \tif (progress)\n@@ -2598,5 +2609,6 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tfprintf(stderr, \"Total %\"PRIu32\" (delta %\"PRIu32\"),\"\n \t\t\t\" reused %\"PRIu32\" (delta %\"PRIu32\")\\n\",\n \t\t\twritten, written_delta, reused, reused_delta);\n+\n \treturn 0;\n }\ndiff --git a/khash.h b/khash.h\nnew file mode 100644\nindex 0000000..e6cdb38\n--- /dev/null\n+++ b/khash.h\n@@ -0,0 +1,329 @@\n+/* The MIT License\n+\n+   Copyright (c) 2008, 2009, 2011 by Attractive Chaos <attractor@live.co.uk>\n+\n+   Permission is hereby granted, free of charge, to any person obtaining\n+   a copy of this software and associated documentation files (the\n+   \"Software\"), to deal in the Software without restriction, including\n+   without limitation the rights to use, copy, modify, merge, publish,\n+   distribute, sublicense, and/or sell copies of the Software, and to\n+   permit persons to whom the Software is furnished to do so, subject to\n+   the following conditions:\n+\n+   The above copyright notice and this permission notice shall be\n+   included in all copies or substantial portions of the Software.\n+\n+   THE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND,\n+   EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF\n+   MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND\n+   NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS\n+   BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN\n+   ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN\n+   CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\n+   SOFTWARE.\n+*/\n+\n+#ifndef __AC_KHASH_H\n+#define __AC_KHASH_H\n+\n+#define AC_VERSION_KHASH_H \"0.2.8\"\n+\n+#include <stdlib.h>\n+#include <string.h>\n+#include <limits.h>\n+\n+typedef uint32_t khint32_t;\n+typedef uint64_t khint64_t;\n+\n+typedef khint32_t khint_t;\n+typedef khint_t khiter_t;\n+\n+#define __ac_isempty(flag, i) ((flag[i>>4]>>((i&0xfU)<<1))&2)\n+#define __ac_isdel(flag, i) ((flag[i>>4]>>((i&0xfU)<<1))&1)\n+#define __ac_iseither(flag, i) ((flag[i>>4]>>((i&0xfU)<<1))&3)\n+#define __ac_set_isdel_false(flag, i) (flag[i>>4]&=~(1ul<<((i&0xfU)<<1)))\n+#define __ac_set_isempty_false(flag, i) (flag[i>>4]&=~(2ul<<((i&0xfU)<<1)))\n+#define __ac_set_isboth_false(flag, i) (flag[i>>4]&=~(3ul<<((i&0xfU)<<1)))\n+#define __ac_set_isdel_true(flag, i) (flag[i>>4]|=1ul<<((i&0xfU)<<1))\n+\n+#define __ac_fsize(m) ((m) < 16? 1 : (m)>>4)\n+\n+#define kroundup32(x) (--(x), (x)|=(x)>>1, (x)|=(x)>>2, (x)|=(x)>>4, (x)|=(x)>>8, (x)|=(x)>>16, ++(x))\n+\n+static const double __ac_HASH_UPPER = 0.77;\n+\n+#define __KHASH_TYPE(name, khkey_t, khval_t) \\\n+\ttypedef struct { \\\n+\t\tkhint_t n_buckets, size, n_occupied, upper_bound; \\\n+\t\tkhint32_t *flags; \\\n+\t\tkhkey_t *keys; \\\n+\t\tkhval_t *vals; \\\n+\t} kh_##name##_t;\n+\n+#define __KHASH_PROTOTYPES(name, khkey_t, khval_t)\t \t\t\t\t\t\\\n+\textern kh_##name##_t *kh_init_##name(void);\t\t\t\t\t\t\t\\\n+\textern void kh_destroy_##name(kh_##name##_t *h);\t\t\t\t\t\\\n+\textern void kh_clear_##name(kh_##name##_t *h);\t\t\t\t\t\t\\\n+\textern khint_t kh_get_##name(const kh_##name##_t *h, khkey_t key); \t\\\n+\textern int kh_resize_##name(kh_##name##_t *h, khint_t new_n_buckets); \\\n+\textern khint_t kh_put_##name(kh_##name##_t *h, khkey_t key, int *ret); \\\n+\textern void kh_del_##name(kh_##name##_t *h, khint_t x);\n+\n+#define __KHASH_IMPL(name, SCOPE, khkey_t, khval_t, kh_is_map, __hash_func, __hash_equal) \\\n+\tSCOPE kh_##name##_t *kh_init_##name(void) {\t\t\t\t\t\t\t\\\n+\t\treturn (kh_##name##_t*)xcalloc(1, sizeof(kh_##name##_t));\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE void kh_destroy_##name(kh_##name##_t *h)\t\t\t\t\t\t\\\n+\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (h) {\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tfree((void *)h->keys); free(h->flags);\t\t\t\t\t\\\n+\t\t\tfree((void *)h->vals);\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tfree(h);\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE void kh_clear_##name(kh_##name##_t *h)\t\t\t\t\t\t\\\n+\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (h && h->flags) {\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tmemset(h->flags, 0xaa, __ac_fsize(h->n_buckets) * sizeof(khint32_t)); \\\n+\t\t\th->size = h->n_occupied = 0;\t\t\t\t\t\t\t\t\\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE khint_t kh_get_##name(const kh_##name##_t *h, khkey_t key) \t\\\n+\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (h->n_buckets) {\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tkhint_t k, i, last, mask, step = 0; \\\n+\t\t\tmask = h->n_buckets - 1;\t\t\t\t\t\t\t\t\t\\\n+\t\t\tk = __hash_func(key); i = k & mask;\t\t\t\t\t\t\t\\\n+\t\t\tlast = i; \\\n+\t\t\twhile (!__ac_isempty(h->flags, i) && (__ac_isdel(h->flags, i) || !__hash_equal(h->keys[i], key))) { \\\n+\t\t\t\ti = (i + (++step)) & mask; \\\n+\t\t\t\tif (i == last) return h->n_buckets;\t\t\t\t\t\t\\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\treturn __ac_iseither(h->flags, i)? h->n_buckets : i;\t\t\\\n+\t\t} else return 0;\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE int kh_resize_##name(kh_##name##_t *h, khint_t new_n_buckets) \\\n+\t{ /* This function uses 0.25*n_buckets bytes of working space instead of [sizeof(key_t+val_t)+.25]*n_buckets. */ \\\n+\t\tkhint32_t *new_flags = 0;\t\t\t\t\t\t\t\t\t\t\\\n+\t\tkhint_t j = 1;\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tkroundup32(new_n_buckets); \t\t\t\t\t\t\t\t\t\\\n+\t\t\tif (new_n_buckets < 4) new_n_buckets = 4;\t\t\t\t\t\\\n+\t\t\tif (h->size >= (khint_t)(new_n_buckets * __ac_HASH_UPPER + 0.5)) j = 0;\t/* requested size is too small */ \\\n+\t\t\telse { /* hash table size to be changed (shrink or expand); rehash */ \\\n+\t\t\t\tnew_flags = (khint32_t*)xmalloc(__ac_fsize(new_n_buckets) * sizeof(khint32_t));\t\\\n+\t\t\t\tif (!new_flags) return -1;\t\t\t\t\t\t\t\t\\\n+\t\t\t\tmemset(new_flags, 0xaa, __ac_fsize(new_n_buckets) * sizeof(khint32_t)); \\\n+\t\t\t\tif (h->n_buckets < new_n_buckets) {\t/* expand */\t\t\\\n+\t\t\t\t\tkhkey_t *new_keys = (khkey_t*)xrealloc((void *)h->keys, new_n_buckets * sizeof(khkey_t)); \\\n+\t\t\t\t\tif (!new_keys) return -1;\t\t\t\t\t\t\t\\\n+\t\t\t\t\th->keys = new_keys;\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\tif (kh_is_map) {\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\t\tkhval_t *new_vals = (khval_t*)xrealloc((void *)h->vals, new_n_buckets * sizeof(khval_t)); \\\n+\t\t\t\t\t\tif (!new_vals) return -1;\t\t\t\t\t\t\\\n+\t\t\t\t\t\th->vals = new_vals;\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t} /* otherwise shrink */\t\t\t\t\t\t\t\t\\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (j) { /* rehashing is needed */\t\t\t\t\t\t\t\t\\\n+\t\t\tfor (j = 0; j != h->n_buckets; ++j) {\t\t\t\t\t\t\\\n+\t\t\t\tif (__ac_iseither(h->flags, j) == 0) {\t\t\t\t\t\\\n+\t\t\t\t\tkhkey_t key = h->keys[j];\t\t\t\t\t\t\t\\\n+\t\t\t\t\tkhval_t val;\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\tkhint_t new_mask;\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\tnew_mask = new_n_buckets - 1; \t\t\t\t\t\t\\\n+\t\t\t\t\tif (kh_is_map) val = h->vals[j];\t\t\t\t\t\\\n+\t\t\t\t\t__ac_set_isdel_true(h->flags, j);\t\t\t\t\t\\\n+\t\t\t\t\twhile (1) { /* kick-out process; sort of like in Cuckoo hashing */ \\\n+\t\t\t\t\t\tkhint_t k, i, step = 0; \\\n+\t\t\t\t\t\tk = __hash_func(key);\t\t\t\t\t\t\t\\\n+\t\t\t\t\t\ti = k & new_mask;\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\t\twhile (!__ac_isempty(new_flags, i)) i = (i + (++step)) & new_mask; \\\n+\t\t\t\t\t\t__ac_set_isempty_false(new_flags, i);\t\t\t\\\n+\t\t\t\t\t\tif (i < h->n_buckets && __ac_iseither(h->flags, i) == 0) { /* kick out the existing element */ \\\n+\t\t\t\t\t\t\t{ khkey_t tmp = h->keys[i]; h->keys[i] = key; key = tmp; } \\\n+\t\t\t\t\t\t\tif (kh_is_map) { khval_t tmp = h->vals[i]; h->vals[i] = val; val = tmp; } \\\n+\t\t\t\t\t\t\t__ac_set_isdel_true(h->flags, i); /* mark it as deleted in the old hash table */ \\\n+\t\t\t\t\t\t} else { /* write the element and jump out of the loop */ \\\n+\t\t\t\t\t\t\th->keys[i] = key;\t\t\t\t\t\t\t\\\n+\t\t\t\t\t\t\tif (kh_is_map) h->vals[i] = val;\t\t\t\\\n+\t\t\t\t\t\t\tbreak;\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tif (h->n_buckets > new_n_buckets) { /* shrink the hash table */ \\\n+\t\t\t\th->keys = (khkey_t*)xrealloc((void *)h->keys, new_n_buckets * sizeof(khkey_t)); \\\n+\t\t\t\tif (kh_is_map) h->vals = (khval_t*)xrealloc((void *)h->vals, new_n_buckets * sizeof(khval_t)); \\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tfree(h->flags); /* free the working space */\t\t\t\t\\\n+\t\t\th->flags = new_flags;\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\th->n_buckets = new_n_buckets;\t\t\t\t\t\t\t\t\\\n+\t\t\th->n_occupied = h->size;\t\t\t\t\t\t\t\t\t\\\n+\t\t\th->upper_bound = (khint_t)(h->n_buckets * __ac_HASH_UPPER + 0.5); \\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\treturn 0;\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE khint_t kh_put_##name(kh_##name##_t *h, khkey_t key, int *ret) \\\n+\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tkhint_t x;\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (h->n_occupied >= h->upper_bound) { /* update the hash table */ \\\n+\t\t\tif (h->n_buckets > (h->size<<1)) {\t\t\t\t\t\t\t\\\n+\t\t\t\tif (kh_resize_##name(h, h->n_buckets - 1) < 0) { /* clear \"deleted\" elements */ \\\n+\t\t\t\t\t*ret = -1; return h->n_buckets;\t\t\t\t\t\t\\\n+\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t} else if (kh_resize_##name(h, h->n_buckets + 1) < 0) { /* expand the hash table */ \\\n+\t\t\t\t*ret = -1; return h->n_buckets;\t\t\t\t\t\t\t\\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t} /* TODO: to implement automatically shrinking; resize() already support shrinking */ \\\n+\t\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\tkhint_t k, i, site, last, mask = h->n_buckets - 1, step = 0; \\\n+\t\t\tx = site = h->n_buckets; k = __hash_func(key); i = k & mask; \\\n+\t\t\tif (__ac_isempty(h->flags, i)) x = i; /* for speed up */\t\\\n+\t\t\telse {\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\tlast = i; \\\n+\t\t\t\twhile (!__ac_isempty(h->flags, i) && (__ac_isdel(h->flags, i) || !__hash_equal(h->keys[i], key))) { \\\n+\t\t\t\t\tif (__ac_isdel(h->flags, i)) site = i;\t\t\t\t\\\n+\t\t\t\t\ti = (i + (++step)) & mask; \\\n+\t\t\t\t\tif (i == last) { x = site; break; }\t\t\t\t\t\\\n+\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\tif (x == h->n_buckets) {\t\t\t\t\t\t\t\t\\\n+\t\t\t\t\tif (__ac_isempty(h->flags, i) && site != h->n_buckets) x = site; \\\n+\t\t\t\t\telse x = i;\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (__ac_isempty(h->flags, x)) { /* not present at all */\t\t\\\n+\t\t\th->keys[x] = key;\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t__ac_set_isboth_false(h->flags, x);\t\t\t\t\t\t\t\\\n+\t\t\t++h->size; ++h->n_occupied;\t\t\t\t\t\t\t\t\t\\\n+\t\t\t*ret = 1;\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t} else if (__ac_isdel(h->flags, x)) { /* deleted */\t\t\t\t\\\n+\t\t\th->keys[x] = key;\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t__ac_set_isboth_false(h->flags, x);\t\t\t\t\t\t\t\\\n+\t\t\t++h->size;\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t\t*ret = 2;\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t} else *ret = 0; /* Don't touch h->keys[x] if present and not deleted */ \\\n+\t\treturn x;\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\tSCOPE void kh_del_##name(kh_##name##_t *h, khint_t x)\t\t\t\t\\\n+\t{\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\tif (x != h->n_buckets && !__ac_iseither(h->flags, x)) {\t\t\t\\\n+\t\t\t__ac_set_isdel_true(h->flags, x);\t\t\t\t\t\t\t\\\n+\t\t\t--h->size;\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t\t}\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t}\n+\n+#define KHASH_DECLARE(name, khkey_t, khval_t)\t\t \t\t\t\t\t\\\n+\t__KHASH_TYPE(name, khkey_t, khval_t) \t\t\t\t\t\t\t\t\\\n+\t__KHASH_PROTOTYPES(name, khkey_t, khval_t)\n+\n+#define KHASH_INIT2(name, SCOPE, khkey_t, khval_t, kh_is_map, __hash_func, __hash_equal) \\\n+\t__KHASH_TYPE(name, khkey_t, khval_t) \t\t\t\t\t\t\t\t\\\n+\t__KHASH_IMPL(name, SCOPE, khkey_t, khval_t, kh_is_map, __hash_func, __hash_equal)\n+\n+#define KHASH_INIT(name, khkey_t, khval_t, kh_is_map, __hash_func, __hash_equal) \\\n+\tKHASH_INIT2(name, static inline, khkey_t, khval_t, kh_is_map, __hash_func, __hash_equal)\n+\n+/* Other convenient macros... */\n+\n+/*! @function\n+  @abstract     Test whether a bucket contains data.\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @param  x     Iterator to the bucket [khint_t]\n+  @return       1 if containing data; 0 otherwise [int]\n+ */\n+#define kh_exist(h, x) (!__ac_iseither((h)->flags, (x)))\n+\n+/*! @function\n+  @abstract     Get key given an iterator\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @param  x     Iterator to the bucket [khint_t]\n+  @return       Key [type of keys]\n+ */\n+#define kh_key(h, x) ((h)->keys[x])\n+\n+/*! @function\n+  @abstract     Get value given an iterator\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @param  x     Iterator to the bucket [khint_t]\n+  @return       Value [type of values]\n+  @discussion   For hash sets, calling this results in segfault.\n+ */\n+#define kh_val(h, x) ((h)->vals[x])\n+\n+/*! @function\n+  @abstract     Alias of kh_val()\n+ */\n+#define kh_value(h, x) ((h)->vals[x])\n+\n+/*! @function\n+  @abstract     Get the start iterator\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @return       The start iterator [khint_t]\n+ */\n+#define kh_begin(h) (khint_t)(0)\n+\n+/*! @function\n+  @abstract     Get the end iterator\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @return       The end iterator [khint_t]\n+ */\n+#define kh_end(h) ((h)->n_buckets)\n+\n+/*! @function\n+  @abstract     Get the number of elements in the hash table\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @return       Number of elements in the hash table [khint_t]\n+ */\n+#define kh_size(h) ((h)->size)\n+\n+/*! @function\n+  @abstract     Get the number of buckets in the hash table\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @return       Number of buckets in the hash table [khint_t]\n+ */\n+#define kh_n_buckets(h) ((h)->n_buckets)\n+\n+/*! @function\n+  @abstract     Iterate over the entries in the hash table\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @param  kvar  Variable to which key will be assigned\n+  @param  vvar  Variable to which value will be assigned\n+  @param  code  Block of code to execute\n+ */\n+#define kh_foreach(h, kvar, vvar, code) { khint_t __i;\t\t\\\n+\tfor (__i = kh_begin(h); __i != kh_end(h); ++__i) {\t\t\\\n+\t\tif (!kh_exist(h,__i)) continue;\t\t\t\t\t\t\\\n+\t\t(kvar) = kh_key(h,__i);\t\t\t\t\t\t\t\t\\\n+\t\t(vvar) = kh_val(h,__i);\t\t\t\t\t\t\t\t\\\n+\t\tcode;\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t} }\n+\n+/*! @function\n+  @abstract     Iterate over the values in the hash table\n+  @param  h     Pointer to the hash table [khash_t(name)*]\n+  @param  vvar  Variable to which value will be assigned\n+  @param  code  Block of code to execute\n+ */\n+#define kh_foreach_value(h, vvar, code) { khint_t __i;\t\t\\\n+\tfor (__i = kh_begin(h); __i != kh_end(h); ++__i) {\t\t\\\n+\t\tif (!kh_exist(h,__i)) continue;\t\t\t\t\t\t\\\n+\t\t(vvar) = kh_val(h,__i);\t\t\t\t\t\t\t\t\\\n+\t\tcode;\t\t\t\t\t\t\t\t\t\t\t\t\\\n+\t} }\n+\n+static inline khint_t __kh_oid_hash(const unsigned char *oid)\n+{\n+\tkhint_t hash;\n+\tmemcpy(&hash, oid, sizeof(hash));\n+\treturn hash;\n+}\n+\n+#define __kh_oid_cmp(a, b) (hashcmp(a, b) == 0)\n+\n+KHASH_INIT(sha1, const unsigned char *, void *, 1, __kh_oid_hash, __kh_oid_cmp)\n+typedef kh_sha1_t khash_sha1;\n+\n+#endif /* __AC_KHASH_H */\n-- \n1.7.9.5\n"},{"id":"221901","messageId":"1372116193-32762-5-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 04/16] pack-objects: make `pack_name_hash` global","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:01Z","receivedAt":"2013-06-24T23:23:01Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"The hash function used by `builtin/pack-objects.c` to efficiently find\ndelta bases when packing can be of interest for other parts of Git that\nalso have to deal with delta bases.\n---\n builtin/pack-objects.c |   24 ++----------------------\n cache.h                |    2 ++\n sha1_file.c            |   20 ++++++++++++++++++++\n 3 files changed, 24 insertions(+), 22 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex fc12df8..b7cab18 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -854,26 +854,6 @@ static struct object_entry *locate_object_entry(const unsigned char *sha1)\n \treturn NULL;\n }\n \n-static unsigned name_hash(const char *name)\n-{\n-\tunsigned c, hash = 0;\n-\n-\tif (!name)\n-\t\treturn 0;\n-\n-\t/*\n-\t * This effectively just creates a sortable number from the\n-\t * last sixteen non-whitespace characters. Last characters\n-\t * count \"most\", so things that end in \".c\" sort together.\n-\t */\n-\twhile ((c = *name++) != 0) {\n-\t\tif (isspace(c))\n-\t\t\tcontinue;\n-\t\thash = (hash >> 2) + (c << 24);\n-\t}\n-\treturn hash;\n-}\n-\n static void setup_delta_attr_check(struct git_attr_check *check)\n {\n \tstatic struct git_attr *attr_delta;\n@@ -977,7 +957,7 @@ static int add_object_entry_1(const unsigned char *sha1, enum object_type type,\n static int add_object_entry(const unsigned char *sha1, enum object_type type,\n \t\t\t    const char *name, int exclude)\n {\n-\tif (add_object_entry_1(sha1, type, name_hash(name), exclude, NULL, 0)) {\n+\tif (add_object_entry_1(sha1, type, pack_name_hash(name), exclude, NULL, 0)) {\n \t\tstruct object_entry *entry = objects[nr_objects - 1];\n \n \t\tif (name && no_try_delta(name))\n@@ -1186,7 +1166,7 @@ static void add_preferred_base_object(const char *name)\n {\n \tstruct pbase_tree *it;\n \tint cmplen;\n-\tunsigned hash = name_hash(name);\n+\tunsigned hash = pack_name_hash(name);\n \n \tif (!num_preferred_base || check_pbase_path(hash))\n \t\treturn;\ndiff --git a/cache.h b/cache.h\nindex a29645e..95ef14d 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -653,6 +653,8 @@ extern char *sha1_pack_index_name(const unsigned char *sha1);\n extern const char *find_unique_abbrev(const unsigned char *sha1, int);\n extern const unsigned char null_sha1[20];\n \n+extern uint32_t pack_name_hash(const char *name);\n+\n static inline int hashcmp(const unsigned char *sha1, const unsigned char *sha2)\n {\n \tint i;\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 371e295..44c7bca 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -60,6 +60,26 @@ static struct cached_object empty_tree = {\n \t0\n };\n \n+uint32_t pack_name_hash(const char *name)\n+{\n+\tunsigned c, hash = 0;\n+\n+\tif (!name)\n+\t\treturn 0;\n+\n+\t/*\n+\t * This effectively just creates a sortable number from the\n+\t * last sixteen non-whitespace characters. Last characters\n+\t * count \"most\", so things that end in \".c\" sort together.\n+\t */\n+\twhile ((c = *name++) != 0) {\n+\t\tif (isspace(c))\n+\t\t\tcontinue;\n+\t\thash = (hash >> 2) + (c << 24);\n+\t}\n+\treturn hash;\n+}\n+\n static struct packed_git *last_found_pack;\n \n static struct cached_object *find_cached_object(const unsigned char *sha1)\n-- \n1.7.9.5\n"},{"id":"221905","messageId":"1372116193-32762-6-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 05/16] revision: allow setting custom limiter function","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:02Z","receivedAt":"2013-06-24T23:23:02Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"This commit enables users of `struct rev_info` to peform custom limiting\nduring a revision walk (i.e. `get_revision`).\n\nIf the field `include_check` has been set to a callback, this callback\nwill be issued once for each commit before it is added to the \"pending\"\nlist of the revwalk. If the include check returns 0, the commit will be\nmarked as added but won't be pushed to the pending list, effectively\nlimiting the walk.\n---\n revision.c |    5 +++++\n revision.h |    2 ++\n 2 files changed, 7 insertions(+)\n\ndiff --git a/revision.c b/revision.c\nindex f1bb731..fa78c65 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -777,8 +777,13 @@ static int add_parents_to_list(struct rev_info *revs, struct commit *commit,\n \n \tif (commit->object.flags & ADDED)\n \t\treturn 0;\n+\n \tcommit->object.flags |= ADDED;\n \n+\tif (revs->include_check &&\n+\t\t!revs->include_check(commit, revs->include_check_data))\n+\t\treturn 0;\n+\n \t/*\n \t * If the commit is uninteresting, don't try to\n \t * prune parents - we want the maximal uninteresting\ndiff --git a/revision.h b/revision.h\nindex eeea6fb..997a093 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -162,6 +162,8 @@ struct rev_info {\n \tunsigned long min_age;\n \tint min_parents;\n \tint max_parents;\n+\tint (*include_check)(struct commit *, void *);\n+\tvoid *include_check_data;\n \n \t/* diff info for patches and for paths limiting */\n \tstruct diff_options diffopt;\n-- \n1.7.9.5\n"},{"id":"221903","messageId":"1372116193-32762-7-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 06/16] sha1_file: export `git_open_noatime`","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:03Z","receivedAt":"2013-06-24T23:23:03Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"The `git_open_noatime` helper can be of general interest for other\nconsumers of git's different on-disk formats.\n---\n cache.h     |    1 +\n sha1_file.c |    4 +---\n 2 files changed, 2 insertions(+), 3 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 95ef14d..bbe5e2a 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -769,6 +769,7 @@ extern int hash_sha1_file(const void *buf, unsigned long len, const char *type,\n extern int write_sha1_file(const void *buf, unsigned long len, const char *type, unsigned char *return_sha1);\n extern int pretend_sha1_file(void *, unsigned long, enum object_type, unsigned char *);\n extern int force_object_loose(const unsigned char *sha1, time_t mtime);\n+extern int git_open_noatime(const char *name);\n extern void *map_sha1_file(const unsigned char *sha1, unsigned long *size);\n extern int unpack_sha1_header(git_zstream *stream, unsigned char *map, unsigned long mapsize, void *buffer, unsigned long bufsiz);\n extern int parse_sha1_header(const char *hdr, unsigned long *sizep);\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 44c7bca..018a847 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -259,8 +259,6 @@ char *sha1_pack_index_name(const unsigned char *sha1)\n struct alternate_object_database *alt_odb_list;\n static struct alternate_object_database **alt_odb_tail;\n \n-static int git_open_noatime(const char *name);\n-\n /*\n  * Prepare alternate object database registry.\n  *\n@@ -1307,7 +1305,7 @@ int check_sha1_signature(const unsigned char *sha1, void *map,\n \treturn hashcmp(sha1, real_sha1) ? -1 : 0;\n }\n \n-static int git_open_noatime(const char *name)\n+int git_open_noatime(const char *name)\n {\n \tstatic int sha1_file_open_flag = O_NOATIME;\n \n-- \n1.7.9.5\n"},{"id":"221906","messageId":"1372116193-32762-8-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 07/16] compat: add endinanness helpers","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:04Z","receivedAt":"2013-06-24T23:23:04Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"The POSIX standard doesn't currently define a `nothll`/`htonll`\nfunction pair to perform network-to-host and host-to-network\nswaps of 64-bit data. These 64-bit swaps are necessary for the on-disk\nstorage of EWAH bitmaps if they are not in native byte order.\n---\n git-compat-util.h |   28 ++++++++++++++++++++++++++++\n 1 file changed, 28 insertions(+)\n\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex ff193f4..bc9b591 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -710,4 +710,32 @@ void warn_on_inaccessible(const char *path);\n /* Get the passwd entry for the UID of the current process. */\n struct passwd *xgetpwuid_self(void);\n \n+#include <endian.h>\n+\n+#ifndef __BYTE_ORDER\n+# if defined(BYTE_ORDER) && defined(LITTLE_ENDIAN) && defined(BIG_ENDIAN)\n+#  define __BYTE_ORDER BYTE_ORDER\n+#  define __LITTLE_ENDIAN LITTLE_ENDIAN\n+#  define __BIG_ENDIAN BIG_ENDIAN\n+# else\n+#  error \"Cannot determine endianness\"\n+# endif\n+#endif\n+\n+#if __BYTE_ORDER == __BIG_ENDIAN\n+# define ntohll(n) (n)\n+# define htonll(n) (n)\n+#elif __BYTE_ORDER == __LITTLE_ENDIAN\n+# if defined(__GNUC__) && defined(__GLIBC__)\n+#  include <byteswap.h>\n+#  define ntohll(n) bswap_64(n)\n+#  define htonll(n) bswap_64(n)\n+# else /* GNUC & GLIBC */\n+#  define ntohll(n) ( (((unsigned long long)ntohl(n)) << 32) + ntohl(n >> 32) )\n+#  define htonll(n) ( (((unsigned long long)htonl(n)) << 32) + htonl(n >> 32) )\n+# endif /* GNUC & GLIBC */\n+#else /* __BYTE_ORDER */\n+# error \"Can't define htonll or ntohll!\"\n+#endif\n+\n #endif\n-- \n1.7.9.5\n"},{"id":"221904","messageId":"1372116193-32762-9-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 08/16] ewah: compressed bitmap implementation","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:05Z","receivedAt":"2013-06-24T23:23:05Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"EWAH is a word-aligned compressed variant of a bitset (i.e. a data\nstructure that acts as a 0-indexed boolean array for many entries).\n\nIt uses a 64-bit run-length encoding (RLE) compression scheme,\ntrading some compression for better processing speed.\n\nThe goal of this word-aligned implementation is not to achieve\nthe best compression, but rather to improve query processing time.\nAs it stands right now, this EWAH implementation will always be more\nefficient storage-wise than its uncompressed alternative.\n\nEWAH arrays will be used as the on-disk format to store reachability\nbitmaps for all objects in a repository while keeping reasonable sizes,\nin the same way that JGit does.\n\nThis EWAH implementation is a mostly straightforward port of the\noriginal `javaewah` library that JGit currently uses. The library is\nself-contained and has been embedded whole (4 files) inside the `ewah`\nfolder to ease redistribution.\n\nThe library is re-licensed under the GPLv2 with the permission of Daniel\nLemire, the original author. The source code for the C version can\nbe found on GitHub:\n\n\thttps://github.com/vmg/libewok\n\nThe original Java implementation can also be found on GitHub:\n\n\thttps://github.com/lemire/javaewah\n---\n Makefile           |    6 +\n ewah/bitmap.c      |  229 +++++++++++++++++\n ewah/ewah_bitmap.c |  703 ++++++++++++++++++++++++++++++++++++++++++++++++++++\n ewah/ewah_io.c     |  199 +++++++++++++++\n ewah/ewah_rlw.c    |  124 +++++++++\n ewah/ewok.h        |  194 +++++++++++++++\n ewah/ewok_rlw.h    |  114 +++++++++\n 7 files changed, 1569 insertions(+)\n create mode 100644 ewah/bitmap.c\n create mode 100644 ewah/ewah_bitmap.c\n create mode 100644 ewah/ewah_io.c\n create mode 100644 ewah/ewah_rlw.c\n create mode 100644 ewah/ewok.h\n create mode 100644 ewah/ewok_rlw.h\n\ndiff --git a/Makefile b/Makefile\nindex e01506d..e03c773 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -672,6 +672,8 @@ LIB_H += diff.h\n LIB_H += diffcore.h\n LIB_H += dir.h\n LIB_H += exec_cmd.h\n+LIB_H += ewah/ewok.h\n+LIB_H += ewah/ewok_rlw.h\n LIB_H += fetch-pack.h\n LIB_H += fmt-merge-msg.h\n LIB_H += fsck.h\n@@ -802,6 +804,10 @@ LIB_OBJS += dir.o\n LIB_OBJS += editor.o\n LIB_OBJS += entry.o\n LIB_OBJS += environment.o\n+LIB_OBJS += ewah/bitmap.o\n+LIB_OBJS += ewah/ewah_bitmap.o\n+LIB_OBJS += ewah/ewah_io.o\n+LIB_OBJS += ewah/ewah_rlw.o\n LIB_OBJS += exec_cmd.o\n LIB_OBJS += fetch-pack.o\n LIB_OBJS += fsck.o\ndiff --git a/ewah/bitmap.c b/ewah/bitmap.c\nnew file mode 100644\nindex 0000000..75ca8fd\n--- /dev/null\n+++ b/ewah/bitmap.c\n@@ -0,0 +1,229 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#include <assert.h>\n+#include <stdlib.h>\n+#include <string.h>\n+\n+#include \"ewok.h\"\n+\n+#define MASK(x) ((eword_t)1 << (x % BITS_IN_WORD))\n+#define BLOCK(x) (x / BITS_IN_WORD)\n+\n+struct bitmap *bitmap_new(void)\n+{\n+\tstruct bitmap *bitmap = ewah_malloc(sizeof(struct bitmap));\n+\tbitmap->words = ewah_calloc(32, sizeof(eword_t));\n+\tbitmap->word_alloc = 32;\n+\treturn bitmap;\n+}\n+\n+void bitmap_set(struct bitmap *self, size_t pos)\n+{\n+\tsize_t block = BLOCK(pos);\n+\n+\tif (block >= self->word_alloc) {\n+\t\tsize_t old_size = self->word_alloc;\n+\t\tself->word_alloc = block * 2;\n+\t\tself->words = ewah_realloc(self->words, self->word_alloc * sizeof(eword_t));\n+\n+\t\tmemset(self->words + old_size, 0x0,\n+\t\t\t(self->word_alloc - old_size) * sizeof(eword_t));\n+\t}\n+\n+\tself->words[block] |= MASK(pos);\n+}\n+\n+void bitmap_clear(struct bitmap *self, size_t pos)\n+{\n+\tsize_t block = BLOCK(pos);\n+\n+\tif (block < self->word_alloc)\n+\t\tself->words[block] &= ~MASK(pos);\n+}\n+\n+bool bitmap_get(struct bitmap *self, size_t pos)\n+{\n+\tsize_t block = BLOCK(pos);\n+\treturn block < self->word_alloc && (self->words[block] & MASK(pos)) != 0;\n+}\n+\n+extern size_t ewah_add_empty_words(struct ewah_bitmap *self, bool v, size_t number);\n+extern size_t ewah_add(struct ewah_bitmap *self, eword_t word);\n+\n+struct ewah_bitmap *bitmap_to_ewah(struct bitmap *bitmap)\n+{\n+\tstruct ewah_bitmap *ewah = ewah_new();\n+\tsize_t i, running_empty_words = 0;\n+\teword_t last_word = 0;\n+\n+\tfor (i = 0; i < bitmap->word_alloc; ++i) {\n+\t\tif (bitmap->words[i] == 0) {\n+\t\t\trunning_empty_words++;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tif (last_word != 0) {\n+\t\t\tewah_add(ewah, last_word);\n+\t\t}\n+\n+\t\tif (running_empty_words > 0) {\n+\t\t\tewah_add_empty_words(ewah, false, running_empty_words);\n+\t\t\trunning_empty_words = 0;\n+\t\t}\n+\n+\t\tlast_word = bitmap->words[i];\n+\t}\n+\n+\tewah_add(ewah, last_word);\n+\treturn ewah;\n+}\n+\n+struct bitmap *ewah_to_bitmap(struct ewah_bitmap *ewah)\n+{\n+\tstruct bitmap *bitmap = bitmap_new();\n+\tstruct ewah_iterator it;\n+\teword_t blowup;\n+\tsize_t i = 0;\n+\n+\tewah_iterator_init(&it, ewah);\n+\n+\twhile (ewah_iterator_next(&blowup, &it)) {\n+\t\tif (i >= bitmap->word_alloc) {\n+\t\t\tbitmap->word_alloc *= 1.5;\n+\t\t\tbitmap->words = ewah_realloc(\n+\t\t\t\tbitmap->words, bitmap->word_alloc * sizeof(eword_t));\n+\t\t}\n+\n+\t\tbitmap->words[i++] = blowup;\n+\t}\n+\n+\tbitmap->word_alloc = i;\n+\treturn bitmap;\n+}\n+\n+void bitmap_and_not_inplace(struct bitmap *self, struct bitmap *other)\n+{\n+\tconst size_t count = (self->word_alloc < other->word_alloc) ?\n+\t\tself->word_alloc : other->word_alloc;\n+\n+\tsize_t i;\n+\n+\tfor (i = 0; i < count; ++i) {\n+\t\tself->words[i] &= ~other->words[i];\n+\t}\n+}\n+\n+void bitmap_or_inplace(struct bitmap *self, struct ewah_bitmap *other)\n+{\n+\tsize_t original_size = self->word_alloc;\n+\tsize_t other_final = (other->bit_size / BITS_IN_WORD) + 1;\n+\tsize_t i = 0;\n+\tstruct ewah_iterator it;\n+\teword_t word;\n+\n+\tif (self->word_alloc < other_final) {\n+\t\tself->word_alloc = other_final;\n+\t\tself->words = ewah_realloc(self->words, self->word_alloc * sizeof(eword_t));\n+\t\tmemset(self->words + original_size, 0x0,\n+\t\t\t(self->word_alloc - original_size) * sizeof(eword_t));\n+\t}\n+\n+\tewah_iterator_init(&it, other);\n+\n+\twhile (ewah_iterator_next(&word, &it)) {\n+\t\tself->words[i++] |= word;\n+\t}\n+}\n+\n+void bitmap_each_bit(struct bitmap *self, ewah_callback callback, void *data)\n+{\n+\tsize_t pos = 0, i;\n+\n+\tfor (i = 0; i < self->word_alloc; ++i) {\n+\t\teword_t word = self->words[i];\n+\t\tuint32_t offset;\n+\n+\t\tif (word == (eword_t)~0) {\n+\t\t\tfor (offset = 0; offset < BITS_IN_WORD; ++offset) {\n+\t\t\t\tcallback(pos++, data);\n+\t\t\t}\n+\t\t} else {\n+\t\t\tfor (offset = 0; offset < BITS_IN_WORD; ++offset) {\n+\t\t\t\tif ((word >> offset) == 0)\n+\t\t\t\t\tbreak;\n+\n+\t\t\t\toffset += __builtin_ctzll(word >> offset);\n+\t\t\t\tcallback(pos + offset, data);\n+\t\t\t}\n+\t\t\tpos += BITS_IN_WORD;\n+\t\t}\n+\t}\n+}\n+\n+size_t bitmap_popcount(struct bitmap *self)\n+{\n+\tsize_t i, count = 0;\n+\n+\tfor (i = 0; i < self->word_alloc; ++i) {\n+\t\tcount += __builtin_popcountll(self->words[i]);\n+\t}\n+\n+\treturn count;\n+}\n+\n+bool bitmap_equals(struct bitmap *self, struct bitmap *other)\n+{\n+\tstruct bitmap *big, *small;\n+\tsize_t i;\n+\n+\tif (self->word_alloc < other->word_alloc) {\n+\t\tsmall = self;\n+\t\tbig = other;\n+\t} else {\n+\t\tsmall = other;\n+\t\tbig = self;\n+\t}\n+\n+\tfor (i = 0; i < small->word_alloc; ++i) {\n+\t\tif (small->words[i] != big->words[i])\n+\t\t\treturn false;\n+\t}\n+\n+\tfor (; i < big->word_alloc; ++i) {\n+\t\tif (big->words[i] != 0)\n+\t\t\treturn false;\n+\t}\n+\n+\treturn true;\n+}\n+\n+void bitmap_reset(struct bitmap *bitmap)\n+{\n+\tmemset(bitmap->words, 0x0, bitmap->word_alloc * sizeof(eword_t));\n+}\n+\n+void bitmap_free(struct bitmap *bitmap)\n+{\n+\tif (bitmap == NULL)\n+\t\treturn;\n+\n+\tfree(bitmap->words);\n+\tfree(bitmap);\n+}\ndiff --git a/ewah/ewah_bitmap.c b/ewah/ewah_bitmap.c\nnew file mode 100644\nindex 0000000..8a23494\n--- /dev/null\n+++ b/ewah/ewah_bitmap.c\n@@ -0,0 +1,703 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#include <assert.h>\n+#include <stdlib.h>\n+#include <unistd.h>\n+#include <string.h>\n+#include <stdio.h>\n+\n+#include \"ewok.h\"\n+#include \"ewok_rlw.h\"\n+\n+static inline size_t min_size(size_t a, size_t b)\n+{\n+\treturn a < b ? a : b;\n+}\n+\n+static inline size_t max_size(size_t a, size_t b)\n+{\n+\treturn a < b ? a : b;\n+}\n+\n+static inline void buffer_grow(struct ewah_bitmap *self, size_t new_size)\n+{\n+\tsize_t rlw_offset = (uint8_t *)self->rlw - (uint8_t *)self->buffer;\n+\n+\tif (self->alloc_size >= new_size)\n+\t\treturn;\n+\n+\tself->alloc_size = new_size;\n+\tself->buffer = ewah_realloc(self->buffer, self->alloc_size * sizeof(eword_t));\n+\tself->rlw = self->buffer + (rlw_offset / sizeof(size_t));\n+}\n+\n+static inline void buffer_push(struct ewah_bitmap *self, eword_t value)\n+{\n+\tif (self->buffer_size + 1 >= self->alloc_size) {\n+\t\tbuffer_grow(self, self->buffer_size * 1.5);\n+\t}\n+\n+\tself->buffer[self->buffer_size++] = value;\n+}\n+\n+static void buffer_push_rlw(struct ewah_bitmap *self, eword_t value)\n+{\n+\tbuffer_push(self, value);\n+\tself->rlw = self->buffer + self->buffer_size - 1;\n+}\n+\n+static size_t add_empty_words(struct ewah_bitmap *self, bool v, size_t number)\n+{\n+\tsize_t added = 0;\n+\n+\tif (rlw_get_run_bit(self->rlw) != v && rlw_size(self->rlw) == 0) {\n+\t\trlw_set_run_bit(self->rlw, v);\n+\t}\n+\telse if (rlw_get_literal_words(self->rlw) != 0 || rlw_get_run_bit(self->rlw) != v) {\n+\t\tbuffer_push_rlw(self, 0);\n+\t\tif (v) rlw_set_run_bit(self->rlw, v);\n+\t\tadded++;\n+\t}\n+\n+\teword_t runlen = rlw_get_running_len(self->rlw);\n+\teword_t can_add = min_size(number, RLW_LARGEST_RUNNING_COUNT - runlen);\n+\n+\trlw_set_running_len(self->rlw, runlen + can_add);\n+\tnumber -= can_add;\n+\n+\twhile (number >= RLW_LARGEST_RUNNING_COUNT) {\n+\t\tbuffer_push_rlw(self, 0);\n+\t\tadded++;\n+\n+\t\tif (v) rlw_set_run_bit(self->rlw, v);\n+\t\trlw_set_running_len(self->rlw, RLW_LARGEST_RUNNING_COUNT);\n+\n+\t\tnumber -= RLW_LARGEST_RUNNING_COUNT;\n+\t}\n+\n+\tif (number > 0) {\n+\t\tbuffer_push_rlw(self, 0);\n+\t\tadded++;\n+\n+\t\tif (v) rlw_set_run_bit(self->rlw, v);\n+\t\trlw_set_running_len(self->rlw, number);\n+\t}\n+\n+\treturn added;\n+}\n+\n+size_t ewah_add_empty_words(struct ewah_bitmap *self, bool v, size_t number)\n+{\n+\tif (number == 0)\n+\t\treturn 0;\n+\n+\tself->bit_size += number * BITS_IN_WORD;\n+\treturn add_empty_words(self, v, number);\n+}\n+\n+static size_t add_literal(struct ewah_bitmap *self, eword_t new_data)\n+{\n+\teword_t current_num = rlw_get_literal_words(self->rlw);\n+\n+\tif (current_num >= RLW_LARGEST_LITERAL_COUNT) {\n+\t\tbuffer_push_rlw(self, 0);\n+\n+\t\trlw_set_literal_words(self->rlw, 1);\n+\t\tbuffer_push(self, new_data);\n+\t\treturn 2;\n+\t}\n+\n+\trlw_set_literal_words(self->rlw, current_num + 1);\n+\n+\t/* sanity check */\n+\tassert(rlw_get_literal_words(self->rlw) == current_num + 1);\n+\n+\tbuffer_push(self, new_data);\n+\treturn 1;\n+}\n+\n+void ewah_add_dirty_words(\n+\tstruct ewah_bitmap *self, const eword_t *buffer, size_t number, bool negate)\n+{\n+\tsize_t literals, can_add;\n+\n+\twhile (1) {\n+\t\tliterals = rlw_get_literal_words(self->rlw);\n+\t\tcan_add = min_size(number, RLW_LARGEST_LITERAL_COUNT - literals);\n+\n+\t\trlw_set_literal_words(self->rlw, literals + can_add);\n+\n+\t\tif (self->buffer_size + can_add >= self->alloc_size) {\n+\t\t\tbuffer_grow(self, (self->buffer_size + can_add) * 1.5);\n+\t\t}\n+\n+\t\tif (negate) {\n+\t\t\tsize_t i;\n+\t\t\tfor (i = 0; i < can_add; ++i)\n+\t\t\t\tself->buffer[self->buffer_size++] = ~buffer[i];\n+\t\t} else {\n+\t\t\tmemcpy(self->buffer + self->buffer_size, buffer, can_add * sizeof(eword_t));\n+\t\t\tself->buffer_size += can_add;\n+\t\t}\n+\n+\t\tself->bit_size += can_add * BITS_IN_WORD;\n+\n+\t\tif (number - can_add == 0)\n+\t\t\tbreak;\n+\n+\t\tbuffer_push_rlw(self, 0);\n+\t\tbuffer += can_add;\n+\t\tnumber -= can_add;\n+\t}\n+}\n+\n+static size_t add_empty_word(struct ewah_bitmap *self, bool v)\n+{\n+\tbool no_literal = (rlw_get_literal_words(self->rlw) == 0);\n+\teword_t run_len = rlw_get_running_len(self->rlw);\n+\n+\tif (no_literal && run_len == 0) {\n+\t\trlw_set_run_bit(self->rlw, v);\n+\t\tassert(rlw_get_run_bit(self->rlw) == v);\n+\t}\n+\n+\tif (no_literal && rlw_get_run_bit(self->rlw) == v &&\n+\t\trun_len < RLW_LARGEST_RUNNING_COUNT) {\n+\t\trlw_set_running_len(self->rlw, run_len + 1);\n+\t\tassert(rlw_get_running_len(self->rlw) == run_len + 1);\n+\t\treturn 0;\n+\t}\n+\n+\telse {\n+\t\tbuffer_push_rlw(self, 0);\n+\n+\t\tassert(rlw_get_running_len(self->rlw) == 0);\n+\t\tassert(rlw_get_run_bit(self->rlw) == 0);\n+\t\tassert(rlw_get_literal_words(self->rlw) == 0);\n+\n+\t\trlw_set_run_bit(self->rlw, v);\n+\t\tassert(rlw_get_run_bit(self->rlw) == v);\n+\n+\t\trlw_set_running_len(self->rlw, 1);\n+\t\tassert(rlw_get_running_len(self->rlw) == 1);\n+\t\tassert(rlw_get_literal_words(self->rlw) == 0);\n+\t\treturn 1;\n+\t}\n+}\n+\n+size_t ewah_add(struct ewah_bitmap *self, eword_t word)\n+{\n+\tself->bit_size += BITS_IN_WORD;\n+\n+\tif (word == 0)\n+\t\treturn add_empty_word(self, false);\n+\n+\tif (word == (eword_t)(~0))\n+\t\treturn add_empty_word(self, true);\n+\n+\treturn add_literal(self, word);\n+}\n+\n+void ewah_set(struct ewah_bitmap *self, size_t i)\n+{\n+\tconst size_t dist =\n+\t\t(i + BITS_IN_WORD) / BITS_IN_WORD -\n+\t\t(self->bit_size + BITS_IN_WORD - 1) / BITS_IN_WORD;\n+\n+\tassert(i >= self->bit_size);\n+\n+\tself->bit_size = i + 1;\n+\n+\tif (dist > 0) {\n+\t\tif (dist > 1)\n+\t\t\tadd_empty_words(self, false, dist - 1);\n+\n+\t\tadd_literal(self, (eword_t)1 << (i % BITS_IN_WORD));\n+\t\treturn;\n+\t}\n+\n+\tif (rlw_get_literal_words(self->rlw) == 0) {\n+\t\trlw_set_running_len(self->rlw, rlw_get_running_len(self->rlw) - 1);\n+\t\tadd_literal(self, (eword_t)1 << (i % BITS_IN_WORD));\n+\t\treturn;\n+\t}\n+\n+\tself->buffer[self->buffer_size - 1] |= ((eword_t)1 << (i % BITS_IN_WORD));\n+\n+\t/* check if we just completed a stream of 1s */\n+\tif (self->buffer[self->buffer_size - 1] == (eword_t)(~0)) {\n+\t\tself->buffer[--self->buffer_size] = 0;\n+\t\trlw_set_literal_words(self->rlw, rlw_get_literal_words(self->rlw) - 1);\n+\t\tadd_empty_word(self, true);\n+\t}\n+}\n+\n+void ewah_each_bit(struct ewah_bitmap *self, void (*callback)(size_t, void*), void *payload)\n+{\n+\tsize_t pos = 0;\n+\tsize_t pointer = 0;\n+\tsize_t k;\n+\n+\twhile (pointer < self->buffer_size) {\n+\t\teword_t *word = &self->buffer[pointer];\n+\n+\t\tif (rlw_get_run_bit(word)) {\n+\t\t\tsize_t len = rlw_get_running_len(word) * BITS_IN_WORD;\n+\t\t\tfor (k = 0; k < len; ++k, ++pos) {\n+\t\t\t\tcallback(pos, payload);\n+\t\t\t}\n+\t\t} else {\n+\t\t\tpos += rlw_get_running_len(word) * BITS_IN_WORD;\n+\t\t}\n+\n+\t\t++pointer;\n+\n+\t\tfor (k = 0; k < rlw_get_literal_words(word); ++k) {\n+\t\t\tint c;\n+\n+\t\t\t/* todo: zero count optimization */\n+\t\t\tfor (c = 0; c < BITS_IN_WORD; ++c, ++pos) {\n+\t\t\t\tif ((self->buffer[pointer] & ((eword_t)1 << c)) != 0) {\n+\t\t\t\t\tcallback(pos, payload);\n+\t\t\t\t}\n+\t\t\t}\n+\n+\t\t\t++pointer;\n+\t\t}\n+\t}\n+}\n+\n+struct ewah_bitmap *ewah_new(void)\n+{\n+\tstruct ewah_bitmap *bitmap;\n+\n+\tbitmap = ewah_malloc(sizeof(struct ewah_bitmap));\n+\tif (bitmap == NULL)\n+\t\treturn NULL;\n+\n+\tbitmap->buffer = ewah_malloc(32 * sizeof(eword_t));\n+\tbitmap->alloc_size = 32;\n+\n+\tewah_clear(bitmap);\n+\n+\treturn bitmap;\n+}\n+\n+void ewah_clear(struct ewah_bitmap *bitmap)\n+{\n+\tbitmap->buffer_size = 1;\n+\tbitmap->buffer[0] = 0;\n+\tbitmap->bit_size = 0;\n+\tbitmap->rlw = bitmap->buffer;\n+}\n+\n+void ewah_free(struct ewah_bitmap *bitmap)\n+{\n+\tif (bitmap->alloc_size)\n+\t\tfree(bitmap->buffer);\n+\n+\tfree(bitmap);\n+}\n+\n+static void read_new_rlw(struct ewah_iterator *it)\n+{\n+\tconst eword_t *word = NULL;\n+\n+\tit->literals = 0;\n+\tit->compressed = 0;\n+\n+\twhile (1) {\n+\t\tword = &it->buffer[it->pointer];\n+\n+\t\tit->rl = rlw_get_running_len(word);\n+\t\tit->lw = rlw_get_literal_words(word);\n+\t\tit->b = rlw_get_run_bit(word);\n+\n+\t\tif (it->rl || it->lw)\n+\t\t\treturn;\n+\n+\t\tif (it->pointer < it->buffer_size - 1) {\n+\t\t\tit->pointer++;\n+\t\t} else {\n+\t\t\tit->pointer = it->buffer_size;\n+\t\t\treturn;\n+\t\t}\n+\t}\n+}\n+\n+bool ewah_iterator_next(eword_t *next, struct ewah_iterator *it)\n+{\n+\tif (it->pointer >= it->buffer_size)\n+\t\treturn false;\n+\n+\tif (it->compressed < it->rl) {\n+\t\tit->compressed++;\n+\t\t*next = it->b ? (eword_t)(~0) : 0;\n+\t} else {\n+\t\tassert(it->literals < it->lw);\n+\n+\t\tit->literals++;\n+\t\tit->pointer++;\n+\n+\t\tassert(it->pointer < it->buffer_size);\n+\n+\t\t*next = it->buffer[it->pointer];\n+\t}\n+\n+\tif (it->compressed == it->rl && it->literals == it->lw) {\n+\t\tif (++it->pointer < it->buffer_size)\n+\t\t\tread_new_rlw(it);\n+\t}\n+\n+\treturn true;\n+}\n+\n+void ewah_iterator_init(struct ewah_iterator *it, struct ewah_bitmap *parent)\n+{\n+\tit->buffer = parent->buffer;\n+\tit->buffer_size = parent->buffer_size;\n+\tit->pointer = 0;\n+\n+\tit->lw = 0;\n+\tit->rl = 0;\n+\tit->compressed = 0;\n+\tit->literals = 0;\n+\tit->b = false;\n+\n+\tif (it->pointer < it->buffer_size)\n+\t\tread_new_rlw(it);\n+}\n+\n+void ewah_dump(struct ewah_bitmap *bitmap)\n+{\n+\tsize_t i;\n+\tfprintf(stderr, \"%zu bits | %zu words | \", bitmap->bit_size, bitmap->buffer_size);\n+\n+\tfor (i = 0; i < bitmap->buffer_size; ++i)\n+\t\tfprintf(stderr, \"%016llx \", (unsigned long long)bitmap->buffer[i]);\n+\n+\tfprintf(stderr, \"\\n\");\n+}\n+\n+void ewah_not(struct ewah_bitmap *self)\n+{\n+\tsize_t pointer = 0;\n+\n+\twhile (pointer < self->buffer_size) {\n+\t\teword_t *word = &self->buffer[pointer];\n+\t\tsize_t literals, k;\n+\n+\t\trlw_xor_run_bit(word);\n+\t\t++pointer;\n+\n+\t\tliterals = rlw_get_literal_words(word);\n+\t\tfor (k = 0; k < literals; ++k) {\n+\t\t\tself->buffer[pointer] = ~self->buffer[pointer];\n+\t\t\t++pointer;\n+\t\t}\n+\t}\n+}\n+\n+void ewah_xor(\n+\tstruct ewah_bitmap *bitmap_i,\n+\tstruct ewah_bitmap *bitmap_j,\n+\tstruct ewah_bitmap *out)\n+{\n+\tstruct rlw_iterator rlw_i;\n+\tstruct rlw_iterator rlw_j;\n+\n+\trlwit_init(&rlw_i, bitmap_i);\n+\trlwit_init(&rlw_j, bitmap_j);\n+\n+\twhile (rlwit_word_size(&rlw_i) > 0 && rlwit_word_size(&rlw_j) > 0) {\n+\t\twhile (rlw_i.rlw.running_len > 0 || rlw_j.rlw.running_len > 0) {\n+\t\t\tstruct rlw_iterator *prey, *predator;\n+\t\t\tsize_t index;\n+\t\t\tbool negate_words;\n+\n+\t\t\tif (rlw_i.rlw.running_len < rlw_j.rlw.running_len) {\n+\t\t\t\tprey = &rlw_i;\n+\t\t\t\tpredator = &rlw_j;\n+\t\t\t} else {\n+\t\t\t\tprey = &rlw_j;\n+\t\t\t\tpredator = &rlw_i;\n+\t\t\t}\n+\n+\t\t\tnegate_words = !!predator->rlw.running_bit;\n+\t\t\tindex = rlwit_discharge(prey, out, predator->rlw.running_len, negate_words);\n+\n+\t\t\tewah_add_empty_words(out, negate_words, predator->rlw.running_len - index);\n+\t\t\trlwit_discard_first_words(predator, predator->rlw.running_len);\n+\t\t}\n+\n+\t\tsize_t literals = min_size(rlw_i.rlw.literal_words, rlw_j.rlw.literal_words);\n+\n+\t\tif (literals) {\n+\t\t\tsize_t k;\n+\n+\t\t\tfor (k = 0; k < literals; ++k) {\n+\t\t\t\tewah_add(out,\n+\t\t\t\t\trlw_i.buffer[rlw_i.literal_word_start + k] ^\n+\t\t\t\t\trlw_j.buffer[rlw_j.literal_word_start + k]\n+\t\t\t\t);\n+\t\t\t}\n+\n+\t\t\trlwit_discard_first_words(&rlw_i, literals);\n+\t\t\trlwit_discard_first_words(&rlw_j, literals);\n+\t\t}\n+\t}\n+\n+\tif (rlwit_word_size(&rlw_i) > 0) {\n+\t\trlwit_discharge(&rlw_i, out, ~0, false);\n+\t} else {\n+\t\trlwit_discharge(&rlw_j, out, ~0, false);\n+\t}\n+\n+\tout->bit_size = max_size(bitmap_i->bit_size, bitmap_j->bit_size);\n+}\n+\n+void ewah_and(\n+\tstruct ewah_bitmap *bitmap_i,\n+\tstruct ewah_bitmap *bitmap_j,\n+\tstruct ewah_bitmap *out)\n+{\n+\tstruct rlw_iterator rlw_i;\n+\tstruct rlw_iterator rlw_j;\n+\n+\trlwit_init(&rlw_i, bitmap_i);\n+\trlwit_init(&rlw_j, bitmap_j);\n+\n+\twhile (rlwit_word_size(&rlw_i) > 0 && rlwit_word_size(&rlw_j) > 0) {\n+\t\twhile (rlw_i.rlw.running_len > 0 || rlw_j.rlw.running_len > 0) {\n+\t\t\tstruct rlw_iterator *prey, *predator;\n+\n+\t\t\tif (rlw_i.rlw.running_len < rlw_j.rlw.running_len) {\n+\t\t\t\tprey = &rlw_i;\n+\t\t\t\tpredator = &rlw_j;\n+\t\t\t} else {\n+\t\t\t\tprey = &rlw_j;\n+\t\t\t\tpredator = &rlw_i;\n+\t\t\t}\n+\n+\t\t\tif (predator->rlw.running_bit == 0) {\n+\t\t\t\tewah_add_empty_words(out, false, predator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(prey, predator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(predator, predator->rlw.running_len);\n+\t\t\t} else {\n+\t\t\t\tsize_t index;\n+\t\t\t\tindex = rlwit_discharge(prey, out, predator->rlw.running_len, false);\n+\t\t\t\tewah_add_empty_words(out, false, predator->rlw.running_len - index);\n+\t\t\t\trlwit_discard_first_words(predator, predator->rlw.running_len);\n+\t\t\t}\n+\t\t}\n+\n+\t\tsize_t literals = min_size(rlw_i.rlw.literal_words, rlw_j.rlw.literal_words);\n+\n+\t\tif (literals) {\n+\t\t\tsize_t k;\n+\n+\t\t\tfor (k = 0; k < literals; ++k) {\n+\t\t\t\tewah_add(out,\n+\t\t\t\t\trlw_i.buffer[rlw_i.literal_word_start + k] &\n+\t\t\t\t\trlw_j.buffer[rlw_j.literal_word_start + k]\n+\t\t\t\t);\n+\t\t\t}\n+\n+\t\t\trlwit_discard_first_words(&rlw_i, literals);\n+\t\t\trlwit_discard_first_words(&rlw_j, literals);\n+\t\t}\n+\t}\n+\n+\tif (rlwit_word_size(&rlw_i) > 0) {\n+\t\trlwit_discharge_empty(&rlw_i, out);\n+\t} else {\n+\t\trlwit_discharge_empty(&rlw_j, out);\n+\t}\n+\n+\tout->bit_size = max_size(bitmap_i->bit_size, bitmap_j->bit_size);\n+}\n+\n+void ewah_and_not(\n+\tstruct ewah_bitmap *bitmap_i,\n+\tstruct ewah_bitmap *bitmap_j,\n+\tstruct ewah_bitmap *out)\n+{\n+\tstruct rlw_iterator rlw_i;\n+\tstruct rlw_iterator rlw_j;\n+\n+\trlwit_init(&rlw_i, bitmap_i);\n+\trlwit_init(&rlw_j, bitmap_j);\n+\n+\twhile (rlwit_word_size(&rlw_i) > 0 && rlwit_word_size(&rlw_j) > 0) {\n+\t\twhile (rlw_i.rlw.running_len > 0 || rlw_j.rlw.running_len > 0) {\n+\t\t\tstruct rlw_iterator *prey, *predator;\n+\n+\t\t\tif (rlw_i.rlw.running_len < rlw_j.rlw.running_len) {\n+\t\t\t\tprey = &rlw_i;\n+\t\t\t\tpredator = &rlw_j;\n+\t\t\t} else {\n+\t\t\t\tprey = &rlw_j;\n+\t\t\t\tpredator = &rlw_i;\n+\t\t\t}\n+\n+\t\t\tif ((predator->rlw.running_bit && prey == &rlw_i) ||\n+\t\t\t\t(!predator->rlw.running_bit && prey != &rlw_i)) {\n+\t\t\t\tewah_add_empty_words(out, false, predator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(prey, predator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(predator, predator->rlw.running_len);\n+\t\t\t} else {\n+\t\t\t\tsize_t index;\n+\t\t\t\tbool negate_words;\n+\n+\t\t\t\tnegate_words = (&rlw_i != prey);\n+\t\t\t\tindex = rlwit_discharge(prey, out, predator->rlw.running_len, negate_words);\n+\t\t\t\tewah_add_empty_words(out, negate_words, predator->rlw.running_len - index);\n+\t\t\t\trlwit_discard_first_words(predator, predator->rlw.running_len);\n+\t\t\t}\n+\t\t}\n+\n+\t\tsize_t literals = min_size(rlw_i.rlw.literal_words, rlw_j.rlw.literal_words);\n+\n+\t\tif (literals) {\n+\t\t\tsize_t k;\n+\n+\t\t\tfor (k = 0; k < literals; ++k) {\n+\t\t\t\tewah_add(out,\n+\t\t\t\t\trlw_i.buffer[rlw_i.literal_word_start + k] &\n+\t\t\t\t\t~(rlw_j.buffer[rlw_j.literal_word_start + k])\n+\t\t\t\t);\n+\t\t\t}\n+\n+\t\t\trlwit_discard_first_words(&rlw_i, literals);\n+\t\t\trlwit_discard_first_words(&rlw_j, literals);\n+\t\t}\n+\t}\n+\n+\tif (rlwit_word_size(&rlw_i) > 0) {\n+\t\trlwit_discharge(&rlw_i, out, ~0, false);\n+\t} else {\n+\t\trlwit_discharge_empty(&rlw_j, out);\n+\t}\n+\n+\tout->bit_size = max_size(bitmap_i->bit_size, bitmap_j->bit_size);\n+}\n+\n+void ewah_or(\n+\tstruct ewah_bitmap *bitmap_i,\n+\tstruct ewah_bitmap *bitmap_j,\n+\tstruct ewah_bitmap *out)\n+{\n+\tstruct rlw_iterator rlw_i;\n+\tstruct rlw_iterator rlw_j;\n+\n+\trlwit_init(&rlw_i, bitmap_i);\n+\trlwit_init(&rlw_j, bitmap_j);\n+\n+\twhile (rlwit_word_size(&rlw_i) > 0 && rlwit_word_size(&rlw_j) > 0) {\n+\t\twhile (rlw_i.rlw.running_len > 0 || rlw_j.rlw.running_len > 0) {\n+\t\t\tstruct rlw_iterator *prey, *predator;\n+\n+\t\t\tif (rlw_i.rlw.running_len < rlw_j.rlw.running_len) {\n+\t\t\t\tprey = &rlw_i;\n+\t\t\t\tpredator = &rlw_j;\n+\t\t\t} else {\n+\t\t\t\tprey = &rlw_j;\n+\t\t\t\tpredator = &rlw_i;\n+\t\t\t}\n+\n+\n+\t\t\tif (predator->rlw.running_bit) {\n+\t\t\t\tewah_add_empty_words(out, false, predator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(prey, predator->rlw.running_len);\n+\t\t\t\trlwit_discard_first_words(predator, predator->rlw.running_len);\n+\t\t\t} else {\n+\t\t\t\tsize_t index;\n+\t\t\t\tindex = rlwit_discharge(prey, out, predator->rlw.running_len, false);\n+\t\t\t\tewah_add_empty_words(out, false, predator->rlw.running_len - index);\n+\t\t\t\trlwit_discard_first_words(predator, predator->rlw.running_len);\n+\t\t\t}\n+\t\t}\n+\n+\t\tsize_t literals = min_size(rlw_i.rlw.literal_words, rlw_j.rlw.literal_words);\n+\n+\t\tif (literals) {\n+\t\t\tsize_t k;\n+\n+\t\t\tfor (k = 0; k < literals; ++k) {\n+\t\t\t\tewah_add(out,\n+\t\t\t\t\trlw_i.buffer[rlw_i.literal_word_start + k] |\n+\t\t\t\t\trlw_j.buffer[rlw_j.literal_word_start + k]\n+\t\t\t\t);\n+\t\t\t}\n+\n+\t\t\trlwit_discard_first_words(&rlw_i, literals);\n+\t\t\trlwit_discard_first_words(&rlw_j, literals);\n+\t\t}\n+\t}\n+\n+\tif (rlwit_word_size(&rlw_i) > 0) {\n+\t\trlwit_discharge(&rlw_i, out, ~0, false);\n+\t} else {\n+\t\trlwit_discharge(&rlw_j, out, ~0, false);\n+\t}\n+\n+\tout->bit_size = max_size(bitmap_i->bit_size, bitmap_j->bit_size);\n+}\n+\n+\n+#define BITMAP_POOL_MAX 16\n+static struct ewah_bitmap *bitmap_pool[BITMAP_POOL_MAX];\n+static size_t bitmap_pool_size;\n+\n+struct ewah_bitmap *ewah_pool_new(void)\n+{\n+\tif (bitmap_pool_size)\n+\t\treturn bitmap_pool[--bitmap_pool_size];\n+\n+\treturn ewah_new();\n+}\n+\n+void ewah_pool_free(struct ewah_bitmap *bitmap)\n+{\n+\tif (bitmap == NULL)\n+\t\treturn;\n+\n+\tif (bitmap_pool_size == BITMAP_POOL_MAX ||\n+\t\tbitmap->alloc_size == 0) {\n+\t\tewah_free(bitmap);\n+\t\treturn;\n+\t}\n+\n+\tewah_clear(bitmap);\n+\tbitmap_pool[bitmap_pool_size++] = bitmap;\n+}\n+\n+uint32_t\n+ewah_checksum(struct ewah_bitmap *self)\n+{\n+\tconst uint8_t *p = (uint8_t *)self->buffer;\n+\tuint32_t crc = (uint32_t)self->bit_size;\n+\tsize_t size = self->buffer_size * sizeof(eword_t);\n+\n+\twhile (size--)\n+\t\tcrc = (crc << 5) - crc + (uint32_t)*p++;\n+\n+\treturn crc;\n+}\ndiff --git a/ewah/ewah_io.c b/ewah/ewah_io.c\nnew file mode 100644\nindex 0000000..b44c90e\n--- /dev/null\n+++ b/ewah/ewah_io.c\n@@ -0,0 +1,199 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#include <stdlib.h>\n+#include <unistd.h>\n+#include <stdio.h>\n+\n+#include \"git-compat-util.h\"\n+#include \"ewok.h\"\n+\n+int ewah_serialize_native(struct ewah_bitmap *self, int fd)\n+{\n+\tuint32_t write32;\n+\tsize_t to_write = self->buffer_size * 8;\n+\n+\t/* 32 bit -- bit size fr the map */\n+\twrite32 = (uint32_t)self->bit_size;\n+\tif (write(fd, &write32, 4) != 4)\n+\t\treturn -1;\n+\n+\t/** 32 bit -- number of compressed 64-bit words */\n+\twrite32 = (uint32_t)self->buffer_size;\n+\tif (write(fd, &write32, 4) != 4)\n+\t\treturn -1;\n+\n+\tif (write(fd, self->buffer, to_write) != to_write)\n+\t\treturn -1;\n+\n+\t/** 32 bit -- position for the RLW */\n+\twrite32 = self->rlw - self->buffer;\n+\tif (write(fd, &write32, 4) != 4)\n+\t\treturn -1;\n+\n+\treturn (3 * 4) + to_write;\n+}\n+\n+int ewah_serialize(struct ewah_bitmap *self, int fd)\n+{\n+\tsize_t i;\n+\teword_t dump[2048];\n+\tconst size_t words_per_dump = sizeof(dump) / sizeof(eword_t);\n+\n+\t/* 32 bit -- bit size fr the map */\n+\tuint32_t bitsize =  htonl((uint32_t)self->bit_size);\n+\tif (write(fd, &bitsize, 4) != 4)\n+\t\treturn -1;\n+\n+\t/** 32 bit -- number of compressed 64-bit words */\n+\tuint32_t word_count =  htonl((uint32_t)self->buffer_size);\n+\tif (write(fd, &word_count, 4) != 4)\n+\t\treturn -1;\n+\n+\t/** 64 bit x N -- compressed words */\n+\tconst eword_t *buffer = self->buffer;\n+\tsize_t words_left = self->buffer_size;\n+\n+\twhile (words_left >= words_per_dump) {\n+\t\tfor (i = 0; i < words_per_dump; ++i, ++buffer)\n+\t\t\tdump[i] = htonll(*buffer);\n+\n+\t\tif (write(fd, dump, sizeof(dump)) != sizeof(dump))\n+\t\t\treturn -1;\n+\n+\t\twords_left -= words_per_dump;\n+\t}\n+\n+\tif (words_left) {\n+\t\tfor (i = 0; i < words_left; ++i, ++buffer)\n+\t\t\tdump[i] = htonll(*buffer);\n+\n+\t\tif (write(fd, dump, words_left * 8) != words_left * 8)\n+\t\t\treturn -1;\n+\t}\n+\n+\t/** 32 bit -- position for the RLW */\n+\tuint32_t rlw_pos = (uint8_t*)self->rlw - (uint8_t *)self->buffer;\n+\trlw_pos = htonl(rlw_pos / sizeof(eword_t));\n+\n+\tif (write(fd, &rlw_pos, 4) != 4)\n+\t\treturn -1;\n+\n+\treturn 0;\n+}\n+\n+int ewah_read_mmap(struct ewah_bitmap *self, void *map, size_t len)\n+{\n+\tuint32_t *read32 = map;\n+\teword_t *read64;\n+\tsize_t i;\n+\n+\tself->bit_size = ntohl(*read32++);\n+\tself->buffer_size = self->alloc_size = ntohl(*read32++);\n+\tself->buffer = ewah_realloc(self->buffer, self->alloc_size * sizeof(eword_t));\n+\n+\tif (!self->buffer)\n+\t\treturn -1;\n+\n+\tfor (i = 0, read64 = (void *)read32; i < self->buffer_size; ++i) {\n+\t\tself->buffer[i] = ntohll(*read64++);\n+\t}\n+\n+\tread32 = (void *)read64;\n+\tself->rlw = self->buffer + ntohl(*read32++);\n+\n+\treturn (char *)read32 - (char *)map;\n+}\n+\n+int ewah_read_mmap_native(struct ewah_bitmap *self, void *map, size_t len)\n+{\n+\tuint32_t *read32 = map;\n+\n+\tself->bit_size = *read32++;\n+\tself->buffer_size = *read32++;\n+\n+\tif (self->alloc_size)\n+\t\tfree(self->buffer);\n+\n+\tself->alloc_size = 0;\n+\tself->buffer = (eword_t *)read32;\n+\n+\tread32 += self->buffer_size * 2;\n+\tself->rlw = self->buffer + *read32++;\n+\n+\treturn (char *)read32 - (char *)map;\n+}\n+\n+int ewah_deserialize(struct ewah_bitmap *self, int fd)\n+{\n+\tsize_t i;\n+\teword_t dump[2048];\n+\tconst size_t words_per_dump = sizeof(dump) / sizeof(eword_t);\n+\n+\tewah_clear(self);\n+\n+\t/* 32 bit -- bit size fr the map */\n+\tuint32_t bitsize;\n+\tif (read(fd, &bitsize, 4) != 4)\n+\t\treturn -1;\n+\n+\tself->bit_size = (size_t)ntohl(bitsize);\n+\n+\t/** 32 bit -- number of compressed 64-bit words */\n+\tuint32_t word_count;\n+\tif (read(fd, &word_count, 4) != 4)\n+\t\treturn -1;\n+\n+\tself->buffer_size = self->alloc_size = (size_t)ntohl(word_count);\n+\tself->buffer = ewah_realloc(self->buffer, self->alloc_size * sizeof(eword_t));\n+\n+\tif (!self->buffer)\n+\t\treturn -1;\n+\n+\t/** 64 bit x N -- compressed words */\n+\teword_t *buffer = self->buffer;\n+\tsize_t words_left = self->buffer_size;\n+\n+\twhile (words_left >= words_per_dump) {\n+\t\tif (read(fd, dump, sizeof(dump)) != sizeof(dump))\n+\t\t\treturn -1;\n+\n+\t\tfor (i = 0; i < words_per_dump; ++i, ++buffer)\n+\t\t\t*buffer = ntohll(dump[i]);\n+\n+\t\twords_left -= words_per_dump;\n+\t}\n+\n+\tif (words_left) {\n+\t\tif (read(fd, dump, words_left * 8) != words_left * 8)\n+\t\t\treturn -1;\n+\n+\t\tfor (i = 0; i < words_left; ++i, ++buffer)\n+\t\t\t*buffer = ntohll(dump[i]);\n+\t}\n+\n+\t/** 32 bit -- position for the RLW */\n+\tuint32_t rlw_pos;\n+\tif (read(fd, &rlw_pos, 4) != 4)\n+\t\treturn -1;\n+\n+\tself->rlw = self->buffer + ntohl(rlw_pos);\n+\n+\treturn 0;\n+}\ndiff --git a/ewah/ewah_rlw.c b/ewah/ewah_rlw.c\nnew file mode 100644\nindex 0000000..7e10fd4\n--- /dev/null\n+++ b/ewah/ewah_rlw.c\n@@ -0,0 +1,124 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#include <assert.h>\n+#include <stdlib.h>\n+#include <unistd.h>\n+#include <string.h>\n+\n+#include \"ewok.h\"\n+#include \"ewok_rlw.h\"\n+\n+extern size_t ewah_add_empty_words(struct ewah_bitmap *self, bool v, size_t number);\n+extern void ewah_add_dirty_words(\n+\tstruct ewah_bitmap *self, const eword_t *buffer, size_t number, bool negate);\n+\n+static inline bool next_word(struct rlw_iterator *it)\n+{\n+\tif (it->pointer >= it->size)\n+\t\treturn false;\n+\n+\tit->rlw.word = &it->buffer[it->pointer];\n+\tit->pointer += rlw_get_literal_words(it->rlw.word) + 1;\n+\n+\tit->rlw.literal_words = rlw_get_literal_words(it->rlw.word);\n+\tit->rlw.running_len = rlw_get_running_len(it->rlw.word);\n+\tit->rlw.running_bit = rlw_get_run_bit(it->rlw.word);\n+\tit->rlw.literal_word_offset = 0;\n+\n+\treturn true;\n+}\n+\n+void rlwit_init(struct rlw_iterator *it, struct ewah_bitmap *bitmap)\n+{\n+\tit->buffer = bitmap->buffer;\n+\tit->size = bitmap->buffer_size;\n+\tit->pointer = 0;\n+\n+\tnext_word(it);\n+\n+\tit->literal_word_start = rlwit_literal_words(it) + it->rlw.literal_word_offset;\n+}\n+\n+void rlwit_discard_first_words(struct rlw_iterator *it, size_t x)\n+{\n+\twhile (x > 0) {\n+\t\tsize_t discard;\n+\n+\t\tif (it->rlw.running_len > x) {\n+\t\t\tit->rlw.running_len -= x;\n+\t\t\treturn;\n+\t\t}\n+\n+\t\tx -= it->rlw.running_len;\n+\t\tit->rlw.running_len = 0;\n+\n+\t\tdiscard = (x > it->rlw.literal_words) ? it->rlw.literal_words : x;\n+\n+\t\tit->literal_word_start += discard;\n+\t\tit->rlw.literal_words -= discard;\n+\t\tx -= discard;\n+\n+\t\tif (x > 0 || rlwit_word_size(it) == 0) {\n+\t\t\tif (!next_word(it))\n+\t\t\t\tbreak;\n+\n+\t\t\tit->literal_word_start =\n+\t\t\t\trlwit_literal_words(it) + it->rlw.literal_word_offset;\n+\t\t}\n+\t}\n+}\n+\n+size_t rlwit_discharge(\n+\tstruct rlw_iterator *it, struct ewah_bitmap *out, size_t max, bool negate)\n+{\n+\tsize_t index = 0;\n+\n+\twhile (index < max && rlwit_word_size(it) > 0) {\n+\t\tsize_t pd, pl = it->rlw.running_len;\n+\n+\t\tif (index + pl > max) {\n+\t\t\tpl = max - index;\n+\t\t}\n+\n+\t\tewah_add_empty_words(out, it->rlw.running_bit ^ negate, pl);\n+\t\tindex += pl;\n+\n+\t\tpd = it->rlw.literal_words;\n+\t\tif (pd + index > max) {\n+\t\t\tpd = max - index;\n+\t\t}\n+\n+\t\tewah_add_dirty_words(out,\n+\t\t\tit->buffer + it->literal_word_start, pd, negate);\n+\n+\t\trlwit_discard_first_words(it, pd + pl);\n+\t\tindex += pd;\n+\t}\n+\n+\treturn index;\n+}\n+\n+void rlwit_discharge_empty(struct rlw_iterator *it, struct ewah_bitmap *out)\n+{\n+\twhile (rlwit_word_size(it) > 0) {\n+\t\tewah_add_empty_words(out, false, rlwit_word_size(it));\n+\t\trlwit_discard_first_words(it, rlwit_word_size(it));\n+\t}\n+}\ndiff --git a/ewah/ewok.h b/ewah/ewok.h\nnew file mode 100644\nindex 0000000..691e21e\n--- /dev/null\n+++ b/ewah/ewok.h\n@@ -0,0 +1,194 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#ifndef __EWOK_BITMAP_C__\n+#define __EWOK_BITMAP_C__\n+\n+#include <stdbool.h>\n+#include <stdint.h>\n+\n+#ifndef ewah_malloc\n+#\tdefine ewah_malloc malloc\n+#endif\n+#ifndef ewah_realloc\n+#\tdefine ewah_realloc realloc\n+#endif\n+#ifndef ewah_calloc\n+#\tdefine ewah_calloc calloc\n+#endif\n+\n+typedef uint64_t eword_t;\n+#define BITS_IN_WORD (sizeof(eword_t) * 8)\n+\n+struct ewah_bitmap {\n+\teword_t *buffer;\n+\tsize_t buffer_size;\n+\tsize_t alloc_size;\n+\tsize_t bit_size;\n+\teword_t *rlw;\n+};\n+\n+typedef void (*ewah_callback)(size_t pos, void *);\n+\n+struct ewah_bitmap *ewah_pool_new(void);\n+void ewah_pool_free(struct ewah_bitmap *bitmap);\n+\n+/**\n+ * Allocate a new EWAH Compressed bitmap\n+ */\n+struct ewah_bitmap *ewah_new(void);\n+\n+/**\n+ * Clear all the bits in the bitmap. Does not free or resize\n+ * memory.\n+ */\n+void ewah_clear(struct ewah_bitmap *bitmap);\n+\n+/**\n+ * Free all the memory of the bitmap\n+ */\n+void ewah_free(struct ewah_bitmap *bitmap);\n+\n+int ewah_serialize(struct ewah_bitmap *self, int fd);\n+int ewah_serialize_native(struct ewah_bitmap *self, int fd);\n+\n+int ewah_deserialize(struct ewah_bitmap *self, int fd);\n+int ewah_read_mmap(struct ewah_bitmap *self, void *map, size_t len);\n+int ewah_read_mmap_native(struct ewah_bitmap *self, void *map, size_t len);\n+\n+uint32_t ewah_checksum(struct ewah_bitmap *self);\n+\n+/**\n+ * Logical not (bitwise negation) in-place on the bitmap\n+ *\n+ * This operation is linear time based on the size of the bitmap.\n+ */\n+void ewah_not(struct ewah_bitmap *self);\n+\n+/**\n+ * Call the given callback with the position of every single bit\n+ * that has been set on the bitmap.\n+ *\n+ * This is an efficient operation that does not fully decompress\n+ * the bitmap.\n+ */\n+void ewah_each_bit(struct ewah_bitmap *self, ewah_callback callback, void *payload);\n+\n+/**\n+ * Set a given bit on the bitmap.\n+ *\n+ * The bit at position `pos` will be set to true. Because of the\n+ * way that the bitmap is compressed, a set bit cannot be unset\n+ * later on.\n+ *\n+ * Furthermore, since the bitmap uses streaming compression, bits\n+ * can only set incrementally.\n+ *\n+ * E.g.\n+ *\t\tewah_set(bitmap, 1); // ok\n+ *\t\tewah_set(bitmap, 76); // ok\n+ *\t\tewah_set(bitmap, 77); // ok\n+ *\t\tewah_set(bitmap, 8712800127); // ok\n+ *\t\tewah_set(bitmap, 25); // failed, assert raised\n+ */\n+void ewah_set(struct ewah_bitmap *self, size_t i);\n+\n+struct ewah_iterator {\n+\tconst eword_t *buffer;\n+\tsize_t buffer_size;\n+\n+\tsize_t pointer;\n+\teword_t compressed, literals;\n+\teword_t rl, lw;\n+\tbool b;\n+};\n+\n+/**\n+ * Initialize a new iterator to run through the bitmap in uncompressed form.\n+ *\n+ * The iterator can be stack allocated. The underlying bitmap must not be freed\n+ * before the iteration is over.\n+ *\n+ * E.g.\n+ *\n+ *\t\tstruct ewah_bitmap *bitmap = ewah_new();\n+ *\t\tstruct ewah_iterator it;\n+ *\n+ *\t\tewah_iterator_init(&it, bitmap);\n+ */\n+void ewah_iterator_init(struct ewah_iterator *it, struct ewah_bitmap *parent);\n+\n+/**\n+ * Yield every single word in the bitmap in uncompressed form. This is:\n+ * yield single words (32-64 bits) where each bit represents an actual\n+ * bit from the bitmap.\n+ *\n+ * Return: true if a word was yield, false if there are no words left\n+ */\n+bool ewah_iterator_next(eword_t *next, struct ewah_iterator *it);\n+\n+void ewah_or(\n+\tstruct ewah_bitmap *bitmap_i,\n+\tstruct ewah_bitmap *bitmap_j,\n+\tstruct ewah_bitmap *out);\n+\n+void ewah_and_not(\n+\tstruct ewah_bitmap *bitmap_i,\n+\tstruct ewah_bitmap *bitmap_j,\n+\tstruct ewah_bitmap *out);\n+\n+void ewah_xor(\n+\tstruct ewah_bitmap *bitmap_i,\n+\tstruct ewah_bitmap *bitmap_j,\n+\tstruct ewah_bitmap *out);\n+\n+void ewah_and(\n+\tstruct ewah_bitmap *bitmap_i,\n+\tstruct ewah_bitmap *bitmap_j,\n+\tstruct ewah_bitmap *out);\n+\n+void ewah_dump(struct ewah_bitmap *bitmap);\n+\n+/**\n+ * Uncompressed, old-school bitmap that can be efficiently compressed\n+ * into an `ewah_bitmap`.\n+ */\n+struct bitmap {\n+\teword_t *words;\n+\tsize_t word_alloc;\n+};\n+\n+struct bitmap *bitmap_new(void);\n+void bitmap_set(struct bitmap *self, size_t pos);\n+void bitmap_clear(struct bitmap *self, size_t pos);\n+bool bitmap_get(struct bitmap *self, size_t pos);\n+void bitmap_reset(struct bitmap *bitmap);\n+void bitmap_free(struct bitmap *self);\n+bool bitmap_equals(struct bitmap *self, struct bitmap *other);\n+\n+struct ewah_bitmap * bitmap_to_ewah(struct bitmap *bitmap);\n+struct bitmap *ewah_to_bitmap(struct ewah_bitmap *ewah);\n+\n+void bitmap_and_not_inplace(struct bitmap *self, struct bitmap *other);\n+void bitmap_or_inplace(struct bitmap *self, struct ewah_bitmap *other);\n+\n+void bitmap_each_bit(struct bitmap *self, ewah_callback callback, void *data);\n+size_t bitmap_popcount(struct bitmap *self);\n+\n+#endif\ndiff --git a/ewah/ewok_rlw.h b/ewah/ewok_rlw.h\nnew file mode 100644\nindex 0000000..2e31836\n--- /dev/null\n+++ b/ewah/ewok_rlw.h\n@@ -0,0 +1,114 @@\n+/**\n+ * Copyright 2013, GitHub, Inc\n+ * Copyright 2009-2013, Daniel Lemire, Cliff Moon,\n+ *\tDavid McIntosh, Robert Becho, Google Inc. and Veronika Zenz\n+ *\n+ * This program is free software; you can redistribute it and/or\n+ * modify it under the terms of the GNU General Public License\n+ * as published by the Free Software Foundation; either version 2\n+ * of the License, or (at your option) any later version.\n+ *\n+ * This program is distributed in the hope that it will be useful,\n+ * but WITHOUT ANY WARRANTY; without even the implied warranty of\n+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+ * GNU General Public License for more details.\n+ *\n+ * You should have received a copy of the GNU General Public License\n+ * along with this program; if not, write to the Free Software\n+ * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA  02110-1301, USA.\n+ */\n+#ifndef __EWOK_RLW_H__\n+#define __EWOK_RLW_H__\n+\n+#define RLW_RUNNING_BITS (sizeof(eword_t) * 4)\n+#define RLW_LITERAL_BITS (sizeof(eword_t) * 8 - 1 - RLW_RUNNING_BITS)\n+\n+#define RLW_LARGEST_RUNNING_COUNT (((eword_t)1 << RLW_RUNNING_BITS) - 1)\n+#define RLW_LARGEST_LITERAL_COUNT (((eword_t)1 << RLW_LITERAL_BITS) - 1)\n+\n+#define RLW_LARGEST_RUNNING_COUNT_SHIFT (RLW_LARGEST_RUNNING_COUNT << 1)\n+\n+#define RLW_RUNNING_LEN_PLUS_BIT (((eword_t)1 << (RLW_RUNNING_BITS + 1)) - 1)\n+\n+static bool rlw_get_run_bit(const eword_t *word)\n+{\n+\treturn *word & (eword_t)1;\n+}\n+\n+static inline void rlw_set_run_bit(eword_t *word, bool b)\n+{\n+\tif (b) {\n+\t\t*word |= (eword_t)1;\n+\t} else {\n+\t\t*word &= (eword_t)(~1);\n+\t}\n+}\n+\n+static inline void rlw_xor_run_bit(eword_t *word)\n+{\n+\tif (*word & 1) {\n+\t\t*word &= (eword_t)(~1);\n+\t} else {\n+\t\t*word |= (eword_t)1;\n+\t}\n+}\n+\n+static inline void rlw_set_running_len(eword_t *word, eword_t l)\n+{\n+\t*word |= RLW_LARGEST_RUNNING_COUNT_SHIFT;\n+\t*word &= (l << 1) | (~RLW_LARGEST_RUNNING_COUNT_SHIFT);\n+}\n+\n+static inline eword_t rlw_get_running_len(const eword_t *word)\n+{\n+\treturn (*word >> 1) & RLW_LARGEST_RUNNING_COUNT;\n+}\n+\n+static inline eword_t rlw_get_literal_words(const eword_t *word)\n+{\n+\treturn *word >> (1 + RLW_RUNNING_BITS);\n+}\n+\n+static inline void rlw_set_literal_words(eword_t *word, eword_t l)\n+{\n+\t*word |= ~RLW_RUNNING_LEN_PLUS_BIT;\n+\t*word &= (l << (RLW_RUNNING_BITS + 1)) | RLW_RUNNING_LEN_PLUS_BIT;\n+}\n+\n+static inline eword_t rlw_size(const eword_t *self)\n+{\n+\treturn rlw_get_running_len(self) + rlw_get_literal_words(self);\n+}\n+\n+struct rlw_iterator {\n+\tconst eword_t *buffer;\n+\tsize_t size;\n+\tsize_t pointer;\n+\tsize_t literal_word_start;\n+\n+\tstruct {\n+\t\tconst eword_t *word;\n+\t\tint literal_words;\n+\t\tint running_len;\n+\t\tint literal_word_offset;\n+\t\tint running_bit;\n+\t} rlw;\n+};\n+\n+void rlwit_init(struct rlw_iterator *it, struct ewah_bitmap *bitmap);\n+void rlwit_discard_first_words(struct rlw_iterator *it, size_t x);\n+size_t rlwit_discharge(\n+\tstruct rlw_iterator *it, struct ewah_bitmap *out, size_t max, bool negate);\n+void rlwit_discharge_empty(struct rlw_iterator *it, struct ewah_bitmap *out);\n+\n+static inline size_t rlwit_word_size(struct rlw_iterator *it)\n+{\n+\treturn it->rlw.running_len + it->rlw.literal_words;\n+}\n+\n+static inline size_t rlwit_literal_words(struct rlw_iterator *it)\n+{\n+\treturn it->pointer - it->rlw.literal_words;\n+}\n+\n+#endif\n-- \n1.7.9.5\n"},{"id":"221907","messageId":"1372116193-32762-10-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:06Z","receivedAt":"2013-06-24T23:23:06Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"This is the technical documentation and design rationale for the new\nBitmap v2 on-disk format.\n---\n Documentation/technical/bitmap-format.txt |  235 +++++++++++++++++++++++++++++\n 1 file changed, 235 insertions(+)\n create mode 100644 Documentation/technical/bitmap-format.txt\n\ndiff --git a/Documentation/technical/bitmap-format.txt b/Documentation/technical/bitmap-format.txt\nnew file mode 100644\nindex 0000000..5400082\n--- /dev/null\n+++ b/Documentation/technical/bitmap-format.txt\n@@ -0,0 +1,235 @@\n+GIT bitmap v2 format & rationale\n+================================\n+\n+\t- A header appears at the beginning, using the same format\n+\tas JGit's original bitmap indexes.\n+\n+\t\t4-byte signature: {'B', 'I', 'T', 'M'}\n+\n+\t\t2-byte version number (network byte order)\n+\t\t\tThe current implementation only supports version 2\n+\t\t\tof the bitmap index. The rationale for this is explained\n+\t\t\tin this document.\n+\n+\t\t2-byte flags (network byte order)\n+\n+\t\t\tThe folowing flags are supported:\n+\n+\t\t\t- BITMAP_OPT_FULL_DAG (0x1) REQUIRED\n+\t\t\tThis flag must always be present. It implies that the bitmap\n+\t\t\tindex has been generated for a packfile with full closure\n+\t\t\t(i.e. where every single object in the packfile can find\n+\t\t\t its parent links inside the same packfile). This is a\n+\t\t\trequirement for the bitmap index format, also present in JGit,\n+\t\t\tthat greatly reduces the complexity of the implementation.\n+\n+\t\t\t- BITMAP_OPT_LE_BITMAPS (0x2)\n+\t\t\tIf present, this implies that that the EWAH bitmaps in this\n+\t\t\tindex has been serialized to disk in little-endian byte order.\n+\t\t\tNote that this only applies to the actual bitmaps, not to the\n+\t\t\tGit data structures in the index, which are always in Network\n+\t\t\tByte order as it's costumary.\n+\n+\t\t\t- BITMAP_OPT_BE_BITMAPS (0x4)\n+\t\t\tIf present, this implies that the EWAH bitmaps have been serialized\n+\t\t\tusing big-endian byte order (NWO). If the flag is missing, **the\n+\t\t\tdefault is to assume that the bitmaps are in big-endian**.\n+\n+\t\t\t- BITMAP_OPT_HASH_CACHE (0x8)\n+\t\t\tIf present, a hash cache for finding delta bases will be available\n+\t\t\tright after the header block in this index. See the following\n+\t\t\tsection for details.\n+\n+\t\t4-byte entry count (network byte order)\n+\n+\t\t\tThe total count of entries (bitmapped commits) in this bitmap index.\n+\n+\t\t20-byte checksum\n+\n+\t\t\tThe SHA1 checksum of the pack this bitmap index belongs to.\n+\n+\t- An OPTIONAL delta cache follows the header.\n+\n+\t\tThe cache is formed by `n` 4-byte hashes in a row, where `n` is\n+\t\tthe amount of objects in the indexed packfile. Note that this amount\n+\t\tis the **total number of objects** and is not related to the\n+\t\tnumber of commits that have been selected and indexed in the\n+\t\tbitmap index.\n+\n+\t\tThe hashes are stored in Network Byte Order and they are the same\n+\t\tvalues generated by a normal revision walk during the `pack-objects`\n+\t\tphase.\n+\n+\t\tThe `n`nth hash in the cache is the name hash for the `n`th object\n+\t\tin the index for the indexed packfile.\n+\n+\t\t[RATIONALE]:\n+\n+\t\tThe bitmap index allows us to skip the Counting Objects phase\n+\t\tduring `pack-objects` and yield all the OIDs that would be reachable\n+\t\t(\"WANTS\") when generating the pack.\n+\n+\t\tThis optimization, however, means that we're adding objects to the\n+\t\tpackfile straight from the packfile index, and hence we are lacking\n+\t\tpath information for the objects that would normally be generated\n+\t\tduring the \"Counting Objects\" phase.\n+\n+\t\tThis path information for each object is hashed and used as a very\n+\t\teffective way to find good delta bases when compressing the packfile;\n+\t\twithout these hashes, the resulting packfiles are much less optimal.\n+\n+\t\tBy storing all the hashes in a cache together with the bitmapsin\n+\t\tthe bitmap index, we can yield not only the SHA1 of all the reachable\n+\t\tobjects, but also their hashes, and allow Git to be much smarter when\n+\t\tfinding delta bases for packing.\n+\n+\t\tIf the delta cache is not available, the bitmap index will obviously\n+\t\tbe smaller in disk, but the packfiles generated using this index will\n+\t\tbe between 20% and 30% bigger, because of the lack of name/path\n+\t\tinformation when finding delta bases.\n+\n+\t- 4 EWAH bitmaps that act as type indexes\n+\n+\t\tType indexes are serialized after the hash cache in the shape\n+\t\tof four EWAH bitmaps stored consecutively (see Appendix A for\n+\t\tthe serialization format of an EWAH bitmap).\n+\n+\t\tThere is a bitmap for each Git object type, stored in the following\n+\t\torder:\n+\n+\t\t\t- Commits\n+\t\t\t- Trees\n+\t\t\t- Blobs\n+\t\t\t- Tags\n+\n+\t\tIn each bitmap, the `n`th bit is set to true if the `n`th object\n+\t\tin the packfile index is of that type.\n+\n+\t\tThe obvious consequence is that the XOR of all 4 bitmaps will result\n+\t\tin a full set (all bits sets), and the AND of all 4 bitmaps will\n+\t\tresult in an empty bitmap (no bits set).\n+\n+\t- N EWAH bitmaps, one for each indexed commit\n+\n+\t\tWhere `N` is the total amount of entries in this bitmap index.\n+\t\tSee Appendix A for the serialization format of an EWAH bitmap.\n+\n+\t- An entry index with `N` entries for the indexed commits\n+\n+\t\tIndex entries are stored consecutively, and each entry has the\n+\t\tfollowing format:\n+\n+\t\t- 20-byte SHA1\n+\t\t\tThe SHA1 of the commit that this bitmap indexes\n+\n+\t\t- 4-byte offset (Network Byte Order)\n+\t\t\tThe offset **from the beginning of the file** where the\n+\t\t\tbitmap for this commit is stored.\n+\n+\t\t- 1-byte XOR-offset\n+\t\t\tThe xor offset used to compress this bitmap. For an entry\n+\t\t\tin position `x`, a XOR offset of `y` means that the actual\n+\t\t\tbitmap representing for this commit is composed by XORing the\n+\t\t\tbitmap for this entry with the bitmap in entry `x-y` (i.e.\n+\t\t\tthe bitmap `y` entries before this one).\n+\n+\t\t\tNote that this compression can be recursive. In order to\n+\t\t\tXOR this entry with a previous one, the previous entry needs\n+\t\t\tto be decompressed first, and so on.\n+\n+\t\t\tThe hard-limit for this offset is 160 (an entry can only be\n+\t\t\txor'ed against one of the 160 entries preceding it). This\n+\t\t\tnumber is always positivea, and hence entries are always xor'ed\n+\t\t\twith **previous** bitmaps, not bitmaps that will come afterwards\n+\t\t\tin the index.\n+\n+\t\t- 1-byte flags for this bitmap\n+\t\t\tAt the moment the only available flag is `0x1`, which hints\n+\t\t\tthat this bitmap can be re-used when rebuilding bitmap indexes\n+\t\t\tfor the repository.\n+\n+\t\t- 2 bytes of RESERVED data (used right now for better packing).\n+\n+== Rationale for changes from the Bitmap Format v1\n+\n+- Serialized EWAH bitmaps can be stored in Little-Endian byte order,\n+  if defined by the BITMAP_OPT_LE_BITMAPS flag in the header.\n+\n+  The original JGit implementation stored bitmaps in Big-Endian byte\n+  order (NWO) because it was unable to `mmap` the serialized format,\n+  and hence always required a full parse of the bitmap index to memory,\n+  where the BE->LE conversion could be performed.\n+\n+  This full parse, however, requires prohibitive loading times in LE\n+  machines (i.e. all modern server hardware): a repository like\n+  `torvalds/linux` can have about 8mb of bitmap indexes, resulting\n+  in roughly 400ms of parse time.\n+\n+  This is not an issue in JGit, which is capable of serving repositories\n+  from a single-process daemon running on the JVM, but `git-daemon` in\n+  git has been implemented with a process-based design (a new\n+  `pack-objects` is spawned for each request), and the boot times\n+  of parsing the bitmap index every time `pack-objects` is spawned can\n+  seriously slow down requests (particularly for small fetches, where we'd\n+  spend about 1.5s booting up and 300ms performing the Counting Objects\n+  phase).\n+\n+  By storing the bitmaps in Little-Endian, we're able to `mmap` their\n+  compressed data straight in memory without parsing it beforehand, and\n+  since most queries don't require accessing all the serialized bitmaps,\n+  we'll only page in the minimal amount of bitmaps necessary to perform\n+  the reachability analysis as they are accessed.\n+\n+- An index of all the bitmapped commits is written at the end of the packfile,\n+  instead of interpersed with the serialized bitmaps in the middle of the\n+  file.\n+\n+  Again, the old design implied a full parse of the whole bitmap index\n+  (which JGit can afford because its daemon is single-process), but it made\n+  impossible `mmaping` the bitmap index file and accessing only the parts\n+  required to actually solve the query.\n+\n+  With an index at the end of the file, we can load only this index in memory,\n+  allowing for very efficient access to all the available bitmaps lazily (we\n+  have their offsets in the mmaped file).\n+\n+- The ordering of the objects in each bitmap has changed from\n+  packfile-order (the nth bit in the bitmap is the nth object in the\n+  packfile) to index-order (the nth bit in the bitmap is the nth object\n+  in the INDEX of the packfile).\n+\n+  There is not a noticeable performance difference when actually converting\n+  from bitmap position to SHA1 and from SHA1 to bitmap position, but when\n+  using packfile ordering like JGit does, queries need to go through the\n+  reverse index (pack-revindex.c).\n+\n+  Generating this reverse index at runtime is **not** free (around 900ms\n+  generation time for a repository like `torvalds/linux`), and once again,\n+  this generation time needs to happen every time `pack-objects` is\n+  spawned.\n+\n+  With index-ordering, the only requirement for SHA1 -> Bitmap conversions\n+  is the packfile index, which we essentially load for free.\n+\n+\n+== Appendix A: Serialization format for an EWAH bitmap\n+\n+Ewah bitmaps are serialized in the protocol as the JAVAEWAH\n+library, making them backwards compatible with the JGit\n+implementation:\n+\n+\t- 4-byte number of bits of the resulting UNCOMPRESSED bitmap\n+\n+\t- 4-byte number of words of the COMPRESSED bitmap, when stored\n+\n+\t- N x 8-byte words, as specified by the previous field\n+\n+\t\tThis is the actual content of the compressed bitmap.\n+\n+\t- 4-byte position of the current RLW for the compressed\n+\t\tbitmap\n+\n+Note that the byte order for this serialization is not defined by\n+default. The byte order for all the content in a serialized EWAH\n+bitmap can be known by the byte order flags in the header of the\n+bitmap index file.\n-- \n1.7.9.5\n"},{"id":"221908","messageId":"1372116193-32762-11-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 10/16] pack-objects: use bitmaps when packing objects","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:07Z","receivedAt":"2013-06-24T23:23:07Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"A bitmap index is used, if available, to speed up the Counting Objects\nphase during `pack-objects`.\n\nThe bitmap index is a `.bitmap` file that can be found inside\n`$GIT_DIR/objects/pack/`, next to its corresponding packfile, and\ncontains precalculated reachability information for selected commits.\nThe full specification of the format for these bitmap indexes can be found\nin `Documentation/technical/bitmap-format.txt`.\n\nFor a given commit SHA1, if it happens to be available in the bitmap\nindex, its bitmap will represent every single object that is reachable\nfrom the commit itself. The nth bit in the bitmap is the nth object in\nthe index of the packfile; if it's set to 1, the object is reachable.\n\nBy using the bitmaps available in the index, this commit implements a new\npair of functions:\n\n\t- `prepare_bitmap_walk`\n\t- `traverse_bitmap_commit_list`\n\nThis first function tries to build a bitmap of all the objects that can be\nreached from the commit roots of a given `rev_info` struct by using\nthe following algorithm:\n\n- If all the interesting commits for a revision walk are available in\nthe index, the resulting reachability bitmap is the bitwise OR of all\nthe individual bitmaps.\n\n- When the full set of WANTs is not available in the index, we perform a\npartial revision walk using the commits that don't have bitmaps as\nroots, and limiting the revision walk as soon as we reach a commit that\nhas a corresponding bitmap. The earlier OR'ed bitmap with all the\nindexed commits can now be completed as this walk progresses, so the end\nresult is the full reachability list.\n\n- For revision walks with a HAVEs set (a set of commits that are deemed\nuninteresting), first we perform the same method as for the WANTs, but\nusing our HAVEs as roots, in order to obtain a full reachability bitmap\nof all the uninteresting commits. This bitmap then can be used to:\n\n\ta) limit the subsequent walk when building the WANTs bitmap\n\tb) finding the final set of interesting commits by performing an\n\t   AND-NOT of the WANTs and the HAVEs.\n\nIf `prepare_bitmap_walk` runs successfully, the resulting bitmap is\nstored and the equivalent of a `traverse_commit_list` call can be\nperformed by using `traverse_bitmap_commit_list`; the bitmap version\nof this call yields the objects straight from the packfile index\n(without having to look them up or parse them) and hence is several\norders of magnitude faster.\n\nIf the `prepare_bitmap_walk` call fails (e.g. because no bitmap files\nare available), the `rev_info` struct is left untouched, and can be used\nto perform a manual rev-walk using `traverse_commit_list`.\n\nHence, this new pair of functions are a generic API that allows to\nperform the equivalent of\n\n\tgit rev-list --objects [roots...] [^uninteresting...]\n\nfor any set of commits, even if they don't have specific bitmaps\ngenerated for them.\n\nIn this specific commit, we use the API to perform the\n`Counting Objects` phase in `builtin/pack-objects.c`, although it could\nbe used to speed up other parts of Git that use the same mechanism.\n\nIf the pack-objects invocation is being piped to `stdout` (like a normal\n`pack-objects` from `upload-pack` would be used) and bitmaps are\nenabled, the new `bitmap_walk` API will be used instead of\n`traverse_commit_list`.\n\nThere are two ways to enable bitmaps for pack-objecs:\n\n\t- Pass the `--use-bitmaps` flag when calling `pack-objects`\n\t- Set `pack.usebitmaps` to `true` in the git config for the\n\trepository.\n\nOf course, simply enabling the bitmaps is not enought to perform the\noptimization: a bitmap index must be available on disk. If no bitmap\nindex can be found, we'll silently fall back to the slow counting\nobjects phase.\n\nThe point of speeding up the Counting Objects phase of `pack-objects` is\nto reduce fetch and clone times for big repositories, which right now\nare definitely dominated by the rev-walk algorithm during the Counting\nObjects phase.\n\nHere are some sample timings from a full pack of `torvalds/linux` (i.e.\nsomething very similar to what would be generated for a clone of the\nrepository):\n\n\t$ time ../git/git pack-objects --all --stdout\n\tCounting objects: 3053537, done.\n\tCompressing objects: 100% (495706/495706), done.\n\tTotal 3053537 (delta 2529614), reused 3053537 (delta 2529614)\n\n\treal    0m36.686s\n\tuser    0m34.440s\n\tsys     0m2.184s\n\n\t$ time ../git/git pack-objects --all --stdout\n\tCounting objects: 3053537, done.\n\tCompressing objects: 100% (495706/495706), done.\n\tTotal 3053537 (delta 2529614), reused 3053537 (delta 2529614)\n\n\treal    0m7.255s\n\tuser    0m6.892s\n\tsys     0m0.444s\n\n>From a hotspot profiling run, we can see how the counting\nobjects phase has been reduced to about 400ms (down from 28s).\nThe remaining time is spent finding deltas and writing the packfile, the\noptimization of which is out of the scope of this topic.\n---\n Makefile               |    2 +\n builtin/pack-objects.c |   31 ++\n pack-bitmap.c          |  818 ++++++++++++++++++++++++++++++++++++++++++++++++\n pack-bitmap.h          |   53 ++++\n 4 files changed, 904 insertions(+)\n create mode 100644 pack-bitmap.c\n create mode 100644 pack-bitmap.h\n\ndiff --git a/Makefile b/Makefile\nindex e03c773..0f2e72b 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -703,6 +703,7 @@ LIB_H += notes.h\n LIB_H += object.h\n LIB_H += pack-revindex.h\n LIB_H += pack.h\n+LIB_H += pack-bitmap.h\n LIB_H += parse-options.h\n LIB_H += patch-ids.h\n LIB_H += pathspec.h\n@@ -838,6 +839,7 @@ LIB_OBJS += notes.o\n LIB_OBJS += notes-cache.o\n LIB_OBJS += notes-merge.o\n LIB_OBJS += object.o\n+LIB_OBJS += pack-bitmap.o\n LIB_OBJS += pack-check.o\n LIB_OBJS += pack-revindex.o\n LIB_OBJS += pack-write.o\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex b7cab18..469b8da 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -19,6 +19,7 @@\n #include \"streaming.h\"\n #include \"thread-utils.h\"\n #include \"khash.h\"\n+#include \"pack-bitmap.h\"\n \n static const char *pack_usage[] = {\n \tN_(\"git pack-objects --stdout [options...] [< ref-list | < object-list]\"),\n@@ -83,6 +84,9 @@ static struct progress *progress_state;\n static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n static int pack_compression_seen;\n \n+static int bitmap_support;\n+static int use_bitmap_index;\n+\n static unsigned long delta_cache_size = 0;\n static unsigned long max_delta_cache_size = 256 * 1024 * 1024;\n static unsigned long cache_max_small_delta_size = 1000;\n@@ -2131,6 +2135,10 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n \t\tcache_max_small_delta_size = git_config_int(k, v);\n \t\treturn 0;\n \t}\n+\tif (!strcmp(k, \"pack.usebitmaps\")) {\n+\t\tbitmap_support = git_config_bool(k, v);\n+\t\treturn 0;\n+\t}\n \tif (!strcmp(k, \"pack.threads\")) {\n \t\tdelta_search_threads = git_config_int(k, v);\n \t\tif (delta_search_threads < 0)\n@@ -2366,8 +2374,24 @@ static void get_object_list(int ac, const char **av)\n \t\t\tdie(\"bad revision '%s'\", line);\n \t}\n \n+\tif (use_bitmap_index) {\n+\t\tuint32_t size_hint;\n+\n+\t\tif (!prepare_bitmap_walk(&revs, &size_hint)) {\n+\t\t\tkhint_t new_hash_size = (size_hint * (1.0 / __ac_HASH_UPPER)) + 0.5;\n+\t\t\tkh_resize_sha1(packed_objects, new_hash_size);\n+\n+\t\t\tnr_alloc = (size_hint + 63) & ~63;\n+\t\t\tobjects = xrealloc(objects, nr_alloc * sizeof(struct object_entry *));\n+\n+\t\t\ttraverse_bitmap_commit_list(&add_object_entry_1);\n+\t\t\treturn;\n+\t\t}\n+\t}\n+\n \tif (prepare_revision_walk(&revs))\n \t\tdie(\"revision walk setup failed\");\n+\n \tmark_edges_uninteresting(revs.commits, &revs, show_edge);\n \ttraverse_commit_list(&revs, show_commit, show_object, NULL);\n \n@@ -2495,6 +2519,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t    N_(\"pack compression level\")),\n \t\tOPT_SET_INT(0, \"keep-true-parents\", &grafts_replace_parents,\n \t\t\t    N_(\"do not hide commits by grafts\"), 0),\n+\t\tOPT_BOOL(0, \"bitmaps\", &bitmap_support,\n+\t\t\t N_(\"enable support for bitmap optimizations\")),\n \t\tOPT_END(),\n \t};\n \n@@ -2561,6 +2587,11 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (keep_unreachable && unpack_unreachable)\n \t\tdie(\"--keep-unreachable and --unpack-unreachable are incompatible.\");\n \n+\tif (bitmap_support) {\n+\t\tif (use_internal_rev_list && pack_to_stdout)\n+\t\t\tuse_bitmap_index = 1;\n+\t}\n+\n \tif (progress && all_progress_implied)\n \t\tprogress = 2;\n \ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nnew file mode 100644\nindex 0000000..090db15\n--- /dev/null\n+++ b/pack-bitmap.c\n@@ -0,0 +1,818 @@\n+#include <stdlib.h>\n+\n+#include \"cache.h\"\n+#include \"commit.h\"\n+#include \"tag.h\"\n+#include \"diff.h\"\n+#include \"revision.h\"\n+#include \"progress.h\"\n+#include \"list-objects.h\"\n+#include \"pack.h\"\n+#include \"pack-bitmap.h\"\n+\n+struct stored_bitmap {\n+\tunsigned char sha1[20];\n+\tstruct ewah_bitmap *root;\n+\tstruct stored_bitmap *xor;\n+\tint flags;\n+};\n+\n+struct bitmap_index {\n+\tstruct ewah_bitmap *commits;\n+\tstruct ewah_bitmap *trees;\n+\tstruct ewah_bitmap *blobs;\n+\tstruct ewah_bitmap *tags;\n+\n+\tkhash_sha1 *bitmaps;\n+\n+\tstruct packed_git *pack;\n+\n+\tstruct {\n+\t\tstruct object_array entries;\n+\t\tkhash_sha1 *map;\n+\t} fake_index;\n+\n+\tstruct bitmap *result;\n+\n+\tint entry_count;\n+\tchar pack_checksum[20];\n+\n+\tint version;\n+\tunsigned loaded : 1,\n+\t\t\t native_bitmaps : 1,\n+\t\t\t has_hash_cache : 1;\n+\n+\tstruct ewah_bitmap *(*read_bitmap)(struct bitmap_index *index);\n+\n+\tvoid *map;\n+\tsize_t map_size, map_pos;\n+\n+\tuint32_t *delta_hashes;\n+};\n+\n+static struct bitmap_index bitmap_git;\n+\n+static struct ewah_bitmap *\n+lookup_stored_bitmap(struct stored_bitmap *st)\n+{\n+\tstruct ewah_bitmap *parent;\n+\tstruct ewah_bitmap *composed;\n+\n+\tif (st->xor == NULL)\n+\t\treturn st->root;\n+\n+\tcomposed = ewah_pool_new();\n+\tparent = lookup_stored_bitmap(st->xor);\n+\tewah_xor(st->root, parent, composed);\n+\n+\tewah_pool_free(st->root);\n+\tst->root = composed;\n+\tst->xor = NULL;\n+\n+\treturn composed;\n+}\n+\n+static struct ewah_bitmap *\n+_read_bitmap(struct bitmap_index *index)\n+{\n+\tstruct ewah_bitmap *b = ewah_pool_new();\n+\tint bitmap_size;\n+\n+\tbitmap_size = ewah_read_mmap(b,\n+\t\tindex->map + index->map_pos,\n+\t\tindex->map_size - index->map_pos);\n+\n+\tif (bitmap_size < 0) {\n+\t\terror(\"Failed to load bitmap index (corruped?)\");\n+\t\tewah_pool_free(b);\n+\t\treturn NULL;\n+\t}\n+\n+\tindex->map_pos += bitmap_size;\n+\treturn b;\n+}\n+\n+static struct ewah_bitmap *\n+_read_bitmap_native(struct bitmap_index *index)\n+{\n+\tstruct ewah_bitmap *b = calloc(1, sizeof(struct ewah_bitmap));\n+\tint bitmap_size;\n+\n+\tbitmap_size = ewah_read_mmap_native(b,\n+\t\tindex->map + index->map_pos,\n+\t\tindex->map_size - index->map_pos);\n+\n+\tif (bitmap_size < 0) {\n+\t\terror(\"Failed to load bitmap index (corruped?)\");\n+\t\tfree(b);\n+\t\treturn NULL;\n+\t}\n+\n+\tindex->map_pos += bitmap_size;\n+\treturn b;\n+}\n+\n+static int load_bitmap_header(struct bitmap_index *index)\n+{\n+\tstruct bitmap_disk_header *header = (void *)index->map;\n+\n+\tif (index->map_size < sizeof(*header))\n+\t\treturn error(\"Corrupted bitmap index (missing header data)\");\n+\n+\tif (memcmp(header->magic, BITMAP_MAGIC_PREFIX, sizeof(BITMAP_MAGIC_PREFIX)) != 0)\n+\t\treturn error(\"Corrupted bitmap index file (wrong header)\");\n+\n+\tindex->version = (int)ntohs(header->version);\n+\tif (index->version != 2)\n+\t\treturn error(\"Unsupported version for bitmap index file (%d)\", index->version);\n+\n+\t/* Parse known bitmap format options */\n+\t{\n+\t\tuint32_t flags = ntohs(header->options);\n+\n+\t\tif ((flags & BITMAP_OPT_FULL_DAG) == 0) {\n+\t\t\treturn error(\"Unsupported options for bitmap index file \"\n+\t\t\t\t\"(Git requires BITMAP_OPT_FULL_DAG)\");\n+\t\t}\n+\n+\t\tif (flags & BITMAP_OPT_HASH_CACHE)\n+\t\t\tindex->has_hash_cache = 1;\n+\n+\t\tindex->read_bitmap = &_read_bitmap;\n+\n+\t\t/*\n+\t\t * If we are in a little endian machine and the bitmap\n+\t\t * was written in LE, we can mmap it straight into memory\n+\t\t * without having to parse it\n+\t\t */\n+\t\tif ((flags & BITMAP_OPT_LE_BITMAPS)) {\n+#if __BYTE_ORDER == __LITTLE_ENDIAN\n+\t\t\tindex->native_bitmaps = 1;\n+\t\t\tindex->read_bitmap = &_read_bitmap_native;\n+#else\n+\t\t\tdie(\"The existing bitmap index is written in little-endian \"\n+\t\t\t\t\"byte order and cannot be read in this machine.\\n\"\n+\t\t\t\t\"Please re-build the bitmap indexes locally.\");\n+#endif\n+\t\t}\n+\t}\n+\n+\tindex->entry_count = ntohl(header->entry_count);\n+\tmemcpy(index->pack_checksum, header->checksum, sizeof(header->checksum));\n+\tindex->map_pos += sizeof(*header);\n+\n+\treturn 0;\n+}\n+\n+static struct stored_bitmap *\n+store_bitmap(struct bitmap_index *index,\n+\tconst unsigned char *sha1,\n+\tstruct ewah_bitmap *bitmap,\n+\tstruct stored_bitmap *xor_with, int flags)\n+{\n+\tstruct stored_bitmap *stored;\n+\tkhiter_t hash_pos;\n+\tint ret;\n+\n+\tstored = xmalloc(sizeof(struct stored_bitmap));\n+\tstored->root = bitmap;\n+\tstored->xor = xor_with;\n+\tstored->flags = flags;\n+\tmemcpy(stored->sha1, sha1, 20);\n+\n+\thash_pos = kh_put_sha1(index->bitmaps, stored->sha1, &ret);\n+\tif (ret == 0) {\n+\t\terror(\"Duplicate entry in bitmap index: %s\", sha1_to_hex(sha1));\n+\t\treturn NULL;\n+\t}\n+\n+\tkh_value(index->bitmaps, hash_pos) = stored;\n+\treturn stored;\n+}\n+\n+static int\n+load_bitmap_entries_v2(struct bitmap_index *index)\n+{\n+\tstatic const int MAX_XOR_OFFSET = 16;\n+\n+\tint i;\n+\tstruct stored_bitmap *recent_bitmaps[16];\n+\tstruct bitmap_disk_entry_v2 *entry;\n+\n+\tvoid *index_pos = index->map + index->map_size -\n+\t\t(index->entry_count * sizeof(struct bitmap_disk_entry_v2));\n+\n+\tfor (i = 0; i < index->entry_count; ++i) {\n+\t\tint xor_offset, flags, ret;\n+\t\tstruct stored_bitmap *xor_bitmap = NULL;\n+\t\tstruct ewah_bitmap *bitmap = NULL;\n+\t\tuint32_t bitmap_pos;\n+\n+\t\tentry = index_pos;\n+\t\tindex_pos += sizeof(struct bitmap_disk_entry_v2);\n+\n+\t\tbitmap_pos = ntohl(entry->bitmap_pos);\n+\t\txor_offset = (int)entry->xor_offset;\n+\t\tflags = (int)entry->flags;\n+\n+\t\tif (index->native_bitmaps) {\n+\t\t\tbitmap = calloc(1, sizeof(struct ewah_bitmap));\n+\t\t\tret = ewah_read_mmap_native(bitmap,\n+\t\t\t\tindex->map + bitmap_pos,\n+\t\t\t\tindex->map_size - bitmap_pos);\n+\t\t} else {\n+\t\t\tbitmap = ewah_pool_new();\n+\t\t\tret = ewah_read_mmap(bitmap,\n+\t\t\t\tindex->map + bitmap_pos,\n+\t\t\t\tindex->map_size - bitmap_pos);\n+\t\t}\n+\n+\t\tif (ret < 0 || xor_offset > MAX_XOR_OFFSET || xor_offset > i) {\n+\t\t\treturn error(\"Corrupted bitmap pack index\");\n+\t\t}\n+\n+\t\tif (xor_offset > 0) {\n+\t\t\txor_bitmap = recent_bitmaps[(i - xor_offset) % MAX_XOR_OFFSET];\n+\n+\t\t\tif (xor_bitmap == NULL)\n+\t\t\t\treturn error(\"Invalid XOR offset in bitmap pack index\");\n+\t\t}\n+\n+\t\trecent_bitmaps[i % MAX_XOR_OFFSET] = store_bitmap(\n+\t\t\tindex, entry->sha1, bitmap, xor_bitmap, flags);\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int load_bitmap_index(\n+\tstruct bitmap_index *index,\n+\tconst char *path,\n+\tstruct packed_git *packfile)\n+{\n+\tint fd = git_open_noatime(path);\n+\tstruct stat st;\n+\n+\tif (fd < 0) {\n+\t\treturn -1;\n+\t}\n+\n+\tif (fstat(fd, &st)) {\n+\t\tclose(fd);\n+\t\treturn -1;\n+\t}\n+\n+\tindex->map_size = xsize_t(st.st_size);\n+\tindex->map = xmmap(NULL, index->map_size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\tclose(fd);\n+\n+\tindex->bitmaps = kh_init_sha1();\n+\tindex->pack = packfile;\n+\tindex->fake_index.map = kh_init_sha1();\n+\n+\tif (load_bitmap_header(index) < 0)\n+\t\treturn -1;\n+\n+\tif (index->has_hash_cache) {\n+\t\tindex->delta_hashes = index->map + index->map_pos;\n+\t\tindex->map_pos += (packfile->num_objects * sizeof(uint32_t));\n+\t}\n+\n+\tif ((index->commits = index->read_bitmap(index)) == NULL ||\n+\t\t(index->trees = index->read_bitmap(index)) == NULL ||\n+\t\t(index->blobs = index->read_bitmap(index)) == NULL ||\n+\t\t(index->tags = index->read_bitmap(index)) == NULL)\n+\t\treturn -1;\n+\n+\tif (load_bitmap_entries_v2(index) < 0)\n+\t\treturn -1;\n+\n+\tindex->loaded = true;\n+\treturn 0;\n+}\n+\n+char *pack_bitmap_filename(struct packed_git *p)\n+{\n+\tchar *idx_name;\n+\tint len;\n+\n+\tlen = strlen(p->pack_name) - strlen(\".pack\");\n+\tidx_name = xmalloc(len + strlen(\".bitmap\") + 1);\n+\n+\tmemcpy(idx_name, p->pack_name, len);\n+\tmemcpy(idx_name + len, \".bitmap\", strlen(\".bitmap\") + 1);\n+\n+\treturn idx_name;\n+}\n+\n+int open_pack_bitmap(struct packed_git *p)\n+{\n+\tchar *idx_name;\n+\tint ret;\n+\n+\tif (open_pack_index(p))\n+\t\tdie(\"failed to open pack %s\", p->pack_name);\n+\n+\tidx_name = pack_bitmap_filename(p);\n+\tret = load_bitmap_index(&bitmap_git, idx_name, p);\n+\tfree(idx_name);\n+\n+\treturn ret;\n+}\n+\n+void prepare_bitmap_git(void)\n+{\n+\tstruct packed_git *p;\n+\n+\tif (bitmap_git.loaded)\n+\t\treturn;\n+\n+\tfor (p = packed_git; p; p = p->next) {\n+\t\tif (open_pack_bitmap(p) == 0)\n+\t\t\treturn;\n+\t}\n+}\n+\n+struct include_data {\n+\tstruct bitmap *base;\n+\tstruct bitmap *seen;\n+};\n+\n+static inline int bitmap_position_extended(const unsigned char *sha1)\n+{\n+\tstruct object_array *array = &bitmap_git.fake_index.entries;\n+\tstruct object_array_entry *entry ;\n+\tint bitmap_pos;\n+\n+\tkhiter_t pos = kh_get_sha1(bitmap_git.fake_index.map, sha1);\n+\n+\tif (pos < kh_end(bitmap_git.fake_index.map)) {\n+\t\tentry = kh_value(bitmap_git.fake_index.map, pos);\n+\n+\t\tbitmap_pos = (entry - array->objects);\n+\t\tbitmap_pos += bitmap_git.pack->num_objects;\n+\n+\t\treturn bitmap_pos;\n+\t}\n+\n+\treturn -1;\n+}\n+\n+static int bitmap_position(const unsigned char *sha1)\n+{\n+\tint pos = find_pack_entry_pos(sha1, bitmap_git.pack);\n+\treturn (pos >= 0) ? pos : bitmap_position_extended(sha1);\n+}\n+\n+static int fake_index_add_object(struct object *object, const char *name)\n+{\n+\tkhiter_t hash_pos;\n+\tint hash_ret;\n+\tint bitmap_pos;\n+\n+\tstruct object_array *array = &bitmap_git.fake_index.entries;\n+\tstruct object_array_entry *entry;\n+\n+\thash_pos = kh_put_sha1(bitmap_git.fake_index.map, object->sha1, &hash_ret);\n+\tif (hash_ret > 0) {\n+\t\tadd_object_array(object, name, array);\n+\t\tentry = &array->objects[array->nr - 1];\n+\t\tkh_value(bitmap_git.fake_index.map, hash_pos) = entry;\n+\t} else {\n+\t\tentry = kh_value(bitmap_git.fake_index.map, hash_pos);\n+\t}\n+\n+\tbitmap_pos = (entry - array->objects);\n+\tbitmap_pos += bitmap_git.pack->num_objects;\n+\n+\treturn bitmap_pos;\n+}\n+\n+static void show_object(struct object *object,\n+\tconst struct name_path *path, const char *last, void *data)\n+{\n+\tstruct bitmap *base = data;\n+\tint bitmap_pos;\n+\n+\tbitmap_pos = bitmap_position(object->sha1);\n+\tif (bitmap_pos < 0) {\n+\t\tbitmap_pos = fake_index_add_object(object, path_name(path, last));\n+\t}\n+\n+\tbitmap_set(base, bitmap_pos);\n+}\n+\n+static void show_commit(struct commit *commit, void *data)\n+{\n+\t/* Nothing to do here */\n+}\n+\n+static int\n+add_to_include_set(struct include_data *data, const unsigned char *sha1, int bitmap_pos)\n+{\n+\tkhiter_t hash_pos;\n+\n+\tif (data->seen && bitmap_get(data->seen, bitmap_pos))\n+\t\treturn 0;\n+\n+\tif (bitmap_get(data->base, bitmap_pos))\n+\t\treturn 0;\n+\n+\thash_pos = kh_get_sha1(bitmap_git.bitmaps, sha1);\n+\tif (hash_pos < kh_end(bitmap_git.bitmaps)) {\n+\t\tstruct stored_bitmap *st = kh_value(bitmap_git.bitmaps, hash_pos);\n+\t\tbitmap_or_inplace(data->base, lookup_stored_bitmap(st));\n+\t\treturn 0;\n+\t}\n+\n+\tbitmap_set(data->base, bitmap_pos);\n+\treturn 1;\n+}\n+\n+static int\n+should_include(struct commit *commit, void *_data)\n+{\n+\tstruct include_data *data = _data;\n+\tint bitmap_pos;\n+\n+\tbitmap_pos = bitmap_position(commit->object.sha1);\n+\tif (bitmap_pos < 0) {\n+\t\tbitmap_pos = fake_index_add_object((struct object *)commit, \"\");\n+\t}\n+\n+\tif (!add_to_include_set(data, commit->object.sha1, bitmap_pos)) {\n+\t\tstruct commit_list *parent = commit->parents;\n+\n+\t\twhile (parent) {\n+\t\t\tparent->item->object.flags |= SEEN;\n+\t\t\tparent = parent->next;\n+\t\t}\n+\n+\t\treturn 0;\n+\t}\n+\n+\treturn 1;\n+}\n+\n+static struct bitmap *\n+find_objects(\n+\tstruct rev_info *revs,\n+\tstruct object_list *roots,\n+\tstruct bitmap *seen)\n+{\n+\tstruct bitmap *base = NULL;\n+\tbool needs_walk = false;\n+\n+\tstruct object_list *not_mapped = NULL;\n+\n+\t/**\n+\t * Go through all the roots for the walk. The ones that have bitmaps\n+\t * on the bitmap index will be `or`ed together to form an initial\n+\t * global reachability analysis.\n+\t *\n+\t * The ones without bitmaps in the index will be stored in the\n+\t * `not_mapped_list` for further processing.\n+\t */\n+\twhile (roots) {\n+\t\tstruct object *object = roots->item;\n+\t\troots = roots->next;\n+\n+\t\tif (object->type == OBJ_COMMIT) {\n+\t\t\tkhiter_t pos = kh_get_sha1(bitmap_git.bitmaps, object->sha1);\n+\n+\t\t\tif (pos < kh_end(bitmap_git.bitmaps)) {\n+\t\t\t\tstruct stored_bitmap *st = kh_value(bitmap_git.bitmaps, pos);\n+\t\t\t\tstruct ewah_bitmap *or_with = lookup_stored_bitmap(st);\n+\n+\t\t\t\tif (base == NULL)\n+\t\t\t\t\tbase = ewah_to_bitmap(or_with);\n+\t\t\t\telse\n+\t\t\t\t\tbitmap_or_inplace(base, or_with);\n+\n+\t\t\t\tobject->flags |= SEEN;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\n+\t\tobject_list_insert(object, &not_mapped);\n+\t}\n+\n+\t/**\n+\t * Best case scenario: We found bitmaps for all the roots,\n+\t * so the resulting `or` bitmap has the full reachability analysis\n+\t */\n+\tif (not_mapped == NULL)\n+\t\treturn base;\n+\n+\troots = not_mapped;\n+\n+\t/**\n+\t * Let's iterate through all the roots that don't have bitmaps to\n+\t * check we can determine them to be reachable from the existing\n+\t * global bitmap.\n+\t *\n+\t * If we cannot find them in the existing global bitmap, we'll need\n+\t * to push them to an actual walk and run it until we can confirm\n+\t * they are reachable\n+\t */\n+\twhile (roots) {\n+\t\tstruct object *object = roots->item;\n+\t\tint pos;\n+\n+\t\troots = roots->next;\n+\t\tpos = bitmap_position(object->sha1);\n+\n+\t\tif (pos < 0 || base == NULL || !bitmap_get(base, pos)) {\n+\t\t\tobject->flags &= ~UNINTERESTING;\n+\t\t\tadd_pending_object(revs, object, \"\");\n+\t\t\tneeds_walk = true;\n+\t\t} else {\n+\t\t\tobject->flags |= SEEN;\n+\t\t}\n+\t}\n+\n+\tif (needs_walk) {\n+\t\tstruct include_data incdata;\n+\n+\t\tif (base == NULL)\n+\t\t\tbase = bitmap_new();\n+\n+\t\tincdata.base = base;\n+\t\tincdata.seen = seen;\n+\n+\t\trevs->include_check = should_include;\n+\t\trevs->include_check_data = &incdata;\n+\n+\t\tif (prepare_revision_walk(revs))\n+\t\t\tdie(\"revision walk setup failed\");\n+\n+\t\ttraverse_commit_list(revs, show_commit, show_object, base);\n+\t}\n+\n+\treturn base;\n+}\n+\n+static void show_extended_objects(\n+\tstruct bitmap *objects,\n+\tshow_reachable_fn show_reach)\n+{\n+\tstruct object_array_entry *entries = bitmap_git.fake_index.entries.objects;\n+\tunsigned int nr = bitmap_git.fake_index.entries.nr;\n+\tunsigned int i;\n+\n+\tfor (i = 0; i < nr; ++i) {\n+\t\tstruct object *obj;\n+\n+\t\tif (!bitmap_get(objects, bitmap_git.pack->num_objects + i))\n+\t\t\tcontinue;\n+\n+\t\tobj = entries[i].item;\n+\t\tshow_reach(obj->sha1, obj->type, pack_name_hash(entries[i].name), 0, NULL, 0);\n+\t}\n+}\n+\n+static void show_objects_for_type(\n+\tstruct bitmap *objects,\n+\tstruct ewah_bitmap *type_filter,\n+\tenum object_type object_type,\n+\tshow_reachable_fn show_reach)\n+{\n+\tsize_t pos = 0, i = 0;\n+\tuint32_t offset;\n+\n+\tstruct ewah_iterator it;\n+\teword_t filter;\n+\n+\tewah_iterator_init(&it, type_filter);\n+\n+\twhile (i < objects->word_alloc && ewah_iterator_next(&filter, &it)) {\n+\t\teword_t word = objects->words[i] & filter;\n+\n+\t\tfor (offset = 0; offset < BITS_IN_WORD; ++offset) {\n+\t\t\tconst unsigned char *sha1;\n+\t\t\toff_t pack_off;\n+\t\t\tuint32_t hash = 0;\n+\n+\t\t\tif ((word >> offset) == 0)\n+\t\t\t\tbreak;\n+\n+\t\t\toffset += __builtin_ctzll(word >> offset);\n+\n+\t\t\tsha1 = nth_packed_object_sha1(bitmap_git.pack, pos + offset);\n+\t\t\tpack_off = nth_packed_object_offset(bitmap_git.pack, pos + offset);\n+\n+\t\t\tif (bitmap_git.delta_hashes)\n+\t\t\t\thash = ntohl(bitmap_git.delta_hashes[pos + offset]);\n+\n+\t\t\tshow_reach(sha1, object_type, hash, 0, bitmap_git.pack, pack_off);\n+\t\t}\n+\n+\t\tpos += BITS_IN_WORD;\n+\t\ti++;\n+\t}\n+}\n+\n+int prepare_bitmap_walk(struct rev_info *revs, uint32_t *result_size)\n+{\n+\tunsigned int i;\n+\tunsigned int pending_nr = revs->pending.nr;\n+\tunsigned int pending_alloc = revs->pending.alloc;\n+\tstruct object_array_entry *pending_e = revs->pending.objects;\n+\n+\tstruct object_list *wants = NULL;\n+\tstruct object_list *haves = NULL;\n+\n+\tstruct bitmap *wants_bitmap = NULL;\n+\tstruct bitmap *haves_bitmap = NULL;\n+\n+\tprepare_bitmap_git();\n+\n+\tif (!bitmap_git.loaded)\n+\t\treturn -1;\n+\n+\trevs->pending.nr = 0;\n+\trevs->pending.alloc = 0;\n+\trevs->pending.objects = NULL;\n+\n+\tfor (i = 0; i < pending_nr; ++i) {\n+\t\tstruct object *object = pending_e[i].item;\n+\n+\t\tif (object->type == OBJ_NONE)\n+\t\t\tparse_object(object->sha1);\n+\n+\t\twhile (object->type == OBJ_TAG) {\n+\t\t\tstruct tag *tag = (struct tag *) object;\n+\n+\t\t\tif (object->flags & UNINTERESTING) {\n+\t\t\t\tobject_list_insert(object, &haves);\n+\t\t\t} else {\n+\t\t\t\tobject_list_insert(object, &wants);\n+\t\t\t}\n+\n+\t\t\tif (!tag->tagged)\n+\t\t\t\tdie(\"bad tag\");\n+\t\t\tobject = parse_object(tag->tagged->sha1);\n+\t\t\tif (!object)\n+\t\t\t\tdie(\"bad object %s\", sha1_to_hex(tag->tagged->sha1));\n+\t\t}\n+\n+\t\tif (object->flags & UNINTERESTING) {\n+\t\t\tobject_list_insert(object, &haves);\n+\t\t} else {\n+\t\t\tobject_list_insert(object, &wants);\n+\t\t}\n+\t}\n+\n+\tif (wants == NULL) {\n+\t\t/* we don't want anything! we're done! */\n+\t\treturn 0;\n+\t}\n+\n+\tif (haves != NULL) {\n+\t\thaves_bitmap = find_objects(revs, haves, NULL);\n+\t\treset_revision_walk();\n+\n+\t\tif (haves_bitmap == NULL)\n+\t\t\tgoto restore_revs;\n+\t}\n+\n+\twants_bitmap = find_objects(revs, wants, haves_bitmap);\n+\n+\tif (wants_bitmap == NULL) {\n+\t\tbitmap_free(haves_bitmap);\n+\t\treset_revision_walk();\n+\t\tgoto restore_revs;\n+\t}\n+\n+\tif (haves_bitmap) {\n+\t\tbitmap_and_not_inplace(wants_bitmap, haves_bitmap);\n+\t}\n+\n+\tbitmap_git.result = wants_bitmap;\n+\n+\tif (result_size) {\n+\t\t*result_size = bitmap_popcount(wants_bitmap);\n+\t}\n+\n+\tbitmap_free(haves_bitmap);\n+\treturn 0;\n+\n+restore_revs:\n+\trevs->pending.nr = pending_nr;\n+\trevs->pending.alloc = pending_alloc;\n+\trevs->pending.objects = pending_e;\n+\treturn -1;\n+}\n+\n+void traverse_bitmap_commit_list(show_reachable_fn show_reachable)\n+{\n+\tif (!bitmap_git.result)\n+\t\tdie(\"Tried to traverse bitmap commit without setting it up first\");\n+\n+\tshow_objects_for_type(bitmap_git.result, bitmap_git.commits, OBJ_COMMIT, show_reachable);\n+\tshow_objects_for_type(bitmap_git.result, bitmap_git.trees, OBJ_TREE, show_reachable);\n+\tshow_objects_for_type(bitmap_git.result, bitmap_git.blobs, OBJ_BLOB, show_reachable);\n+\tshow_objects_for_type(bitmap_git.result, bitmap_git.tags, OBJ_TAG, show_reachable);\n+\n+\tshow_extended_objects(bitmap_git.result, show_reachable);\n+\n+\tbitmap_free(bitmap_git.result);\n+\tbitmap_git.result = NULL;\n+}\n+\n+struct bitmap_test_data {\n+\tstruct bitmap *base;\n+\tstruct progress *prg;\n+\tsize_t seen;\n+};\n+\n+static void test_show_object(struct object *object,\n+\tconst struct name_path *path, const char *last, void *data)\n+{\n+\tstruct bitmap_test_data *tdata = data;\n+\tint bitmap_pos;\n+\n+\tbitmap_pos = bitmap_position(object->sha1);\n+\tif (bitmap_pos < 0) {\n+\t\tdie(\"Object not in bitmap: %s\\n\", sha1_to_hex(object->sha1));\n+\t}\n+\n+\tbitmap_set(tdata->base, bitmap_pos);\n+\tdisplay_progress(tdata->prg, ++tdata->seen);\n+}\n+\n+static void test_show_commit(struct commit *commit, void *data)\n+{\n+\tstruct bitmap_test_data *tdata = data;\n+\tint bitmap_pos;\n+\n+\tbitmap_pos = bitmap_position(commit->object.sha1);\n+\tif (bitmap_pos < 0) {\n+\t\tdie(\"Object not in bitmap: %s\\n\", sha1_to_hex(commit->object.sha1));\n+\t}\n+\n+\tbitmap_set(tdata->base, bitmap_pos);\n+\tdisplay_progress(tdata->prg, ++tdata->seen);\n+}\n+\n+void test_bitmap_walk(struct rev_info *revs)\n+{\n+\tstruct object *root;\n+\tstruct bitmap *result = NULL;\n+\tkhiter_t pos;\n+\tsize_t result_popcnt;\n+\tstruct bitmap_test_data tdata;\n+\n+\tprepare_bitmap_git();\n+\n+\tif (!bitmap_git.loaded) {\n+\t\tdie(\"failed to load bitmap indexes\");\n+\t}\n+\n+\tif (revs->pending.nr != 1) {\n+\t\tdie(\"only one bitmap can be tested at a time\");\n+\t}\n+\n+\tfprintf(stderr, \"Bitmap v%d test (%d entries loaded)\\n\",\n+\t\tbitmap_git.version, bitmap_git.entry_count);\n+\n+\troot = revs->pending.objects[0].item;\n+\tpos = kh_get_sha1(bitmap_git.bitmaps, root->sha1);\n+\n+\tif (pos < kh_end(bitmap_git.bitmaps)) {\n+\t\tstruct stored_bitmap *st = kh_value(bitmap_git.bitmaps, pos);\n+\t\tstruct ewah_bitmap *bm = lookup_stored_bitmap(st);\n+\n+\t\tfprintf(stderr, \"Found bitmap for %s. %d bits / %08x checksum\\n\",\n+\t\t\tsha1_to_hex(root->sha1), (int)bm->bit_size, ewah_checksum(bm));\n+\n+\t\tresult = ewah_to_bitmap(bm);\n+\t}\n+\n+\tif (result == NULL) {\n+\t\tdie(\"Commit %s doesn't have an indexed bitmap\", sha1_to_hex(root->sha1));\n+\t}\n+\n+\trevs->tag_objects = 1;\n+\trevs->tree_objects = 1;\n+\trevs->blob_objects = 1;\n+\n+\tresult_popcnt = bitmap_popcount(result);\n+\n+\tif (prepare_revision_walk(revs))\n+\t\tdie(\"revision walk setup failed\");\n+\n+\ttdata.base = bitmap_new();\n+\ttdata.prg = start_progress(\"Verifying bitmap entries\", result_popcnt);\n+\ttdata.seen = 0;\n+\n+\ttraverse_commit_list(revs, &test_show_commit, &test_show_object, &tdata);\n+\n+\tstop_progress(&tdata.prg);\n+\n+\tif (bitmap_equals(result, tdata.base)) {\n+\t\tfprintf(stderr, \"OK!\\n\");\n+\t} else {\n+\t\tfprintf(stderr, \"Mismatch!\\n\");\n+\t}\n+}\ndiff --git a/pack-bitmap.h b/pack-bitmap.h\nnew file mode 100644\nindex 0000000..b97bd46\n--- /dev/null\n+++ b/pack-bitmap.h\n@@ -0,0 +1,53 @@\n+#ifndef PACK_BITMAP_H\n+#define PACK_BITMAP_H\n+\n+#define ewah_malloc xmalloc\n+#define ewah_calloc xcalloc\n+#define ewah_realloc xrealloc\n+#include \"ewah/ewok.h\"\n+#include \"khash.h\"\n+\n+struct bitmap_disk_entry {\n+\tuint32_t object_pos;\n+\tuint8_t xor_offset;\n+\tuint8_t flags;\n+};\n+\n+struct bitmap_disk_entry_v2 {\n+\tunsigned char sha1[20];\n+\tuint32_t bitmap_pos;\n+\tuint8_t xor_offset;\n+\tuint8_t flags;\n+\tuint8_t __pad[2];\n+};\n+\n+struct bitmap_disk_header {\n+\tchar magic[4];\n+\tuint16_t version;\n+\tuint16_t options;\n+\tuint32_t entry_count;\n+\tchar checksum[20];\n+};\n+\n+static const char BITMAP_MAGIC_PREFIX[] = {'B', 'I', 'T', 'M'};;\n+\n+enum pack_bitmap_opts {\n+\tBITMAP_OPT_FULL_DAG = 1,\n+\tBITMAP_OPT_LE_BITMAPS = 2,\n+\tBITMAP_OPT_BE_BITMAPS = 4,\n+\tBITMAP_OPT_HASH_CACHE = 8\n+};\n+\n+typedef int (*show_reachable_fn)(\n+\tconst unsigned char *sha1,\n+\tenum object_type type,\n+\tuint32_t hash, int exclude,\n+\tstruct packed_git *found_pack,\n+\toff_t found_offset);\n+\n+void traverse_bitmap_commit_list(show_reachable_fn show_reachable);\n+int prepare_bitmap_walk(struct rev_info *revs, uint32_t *result_size);\n+void test_bitmap_walk(struct rev_info *revs);\n+char *pack_bitmap_filename(struct packed_git *p);\n+\n+#endif\n-- \n1.7.9.5\n"},{"id":"221910","messageId":"1372116193-32762-12-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 11/16] rev-list: add bitmap mode to speed up lists","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:08Z","receivedAt":"2013-06-24T23:23:08Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"The bitmap reachability index used to speed up the counting objects\nphase during `pack-objects` can also be used to optimize a normal\nrev-list if the only thing required are the SHA1s of the objects during\nthe list.\n\nCalling `git rev-list --use-bitmaps [committish]` is the equivalent\nof `git rev-list --objects`, but the rev list is performed based on\na bitmap result instead of using a manual counting objects phase.\n\nThese are some example timings for `torvalds/linux`:\n\n\t$ time ../git/git rev-list --objects master > /dev/null\n\n\treal    0m25.567s\n\tuser    0m25.148s\n\tsys     0m0.384s\n\n\t$ time ../git/git rev-list --use-bitmaps master > /dev/null\n\n\treal    0m0.393s\n\tuser    0m0.356s\n\tsys     0m0.036s\n\nAdditionally, a `--test-bitmap` flag has been added that will perform\nthe same rev-list manually (i.e. using a normal revwalk) and using\nbitmaps, and verify that the results are the same.\n---\n builtin/rev-list.c |   28 +++++++++++++++++++++++++++-\n 1 file changed, 27 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/rev-list.c b/builtin/rev-list.c\nindex 67701be..905ed08 100644\n--- a/builtin/rev-list.c\n+++ b/builtin/rev-list.c\n@@ -3,6 +3,8 @@\n #include \"diff.h\"\n #include \"revision.h\"\n #include \"list-objects.h\"\n+#include \"pack.h\"\n+#include \"pack-bitmap.h\"\n #include \"builtin.h\"\n #include \"log-tree.h\"\n #include \"graph.h\"\n@@ -256,6 +258,17 @@ static int show_bisect_vars(struct rev_list_info *info, int reaches, int all)\n \treturn 0;\n }\n \n+static int show_object_fast(\n+\tconst unsigned char *sha1,\n+\tenum object_type type,\n+\tuint32_t hash, int exclude,\n+\tstruct packed_git *found_pack,\n+\toff_t found_offset)\n+{\n+\tfprintf(stdout, \"%ss\\n\", sha1_to_hex(sha1));\n+\treturn 1;\n+}\n+\n int cmd_rev_list(int argc, const char **argv, const char *prefix)\n {\n \tstruct rev_info revs;\n@@ -264,6 +277,7 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \tint bisect_list = 0;\n \tint bisect_show_vars = 0;\n \tint bisect_find_all = 0;\n+\tint use_bitmaps = 0;\n \n \tgit_config(git_default_config, NULL);\n \tinit_revisions(&revs, prefix);\n@@ -305,8 +319,15 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \t\t\tbisect_show_vars = 1;\n \t\t\tcontinue;\n \t\t}\n+\t\tif (!strcmp(arg, \"--use-bitmaps\")) {\n+\t\t\tuse_bitmaps = 1;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (!strcmp(arg, \"--test-bitmap\")) {\n+\t\t\ttest_bitmap_walk(&revs);\n+\t\t\treturn 0;\n+\t\t}\n \t\tusage(rev_list_usage);\n-\n \t}\n \tif (revs.commit_format != CMIT_FMT_UNSPECIFIED) {\n \t\t/* The command line has a --pretty  */\n@@ -332,6 +353,11 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \tif (bisect_list)\n \t\trevs.limited = 1;\n \n+\tif (use_bitmaps && !prepare_bitmap_walk(&revs, NULL)) {\n+\t\ttraverse_bitmap_commit_list(&show_object_fast);\n+\t\treturn 0;\n+\t}\n+\n \tif (prepare_revision_walk(&revs))\n \t\tdie(\"revision walk setup failed\");\n \tif (revs.tree_objects)\n-- \n1.7.9.5\n"},{"id":"221909","messageId":"1372116193-32762-13-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 12/16] pack-objects: implement bitmap writing","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:09Z","receivedAt":"2013-06-24T23:23:09Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"This commit extends more the functionality of `pack-objects` by allowing\nit to write out a `.bitmap` index next to any written packs, together\nwith the `.idx` index that currently gets written.\n\nIf bitmaps are enabled for a given repository (either by calling\n`pack-objects` with the `--use-bitmaps` flag or by having\n`pack.usebitmaps` set to `true` in the config) and pack-objects is\nwriting a packfile that would normally be indexed (i.e. not piping to\nstdout), we will attempt to write the corresponding bitmap index for the\npackfile.\n\nBitmap index writing happens after the packfile and its index has been\nsuccessfully written to disk (`finish_tmp_packfile`). The process is\nperformed in several steps:\n\n\t1. `bitmap_writer_build_type_index`: this call uses the array of\n\t`struct object_entry`es that has just been sorted when writing out\n\tthe actual packfile index to disk to generate 4 type-index bitmaps\n\t(one for each object type).\n\n\tThese bitmaps have their nth bit set if the given object is of the\n\tbitmap's type. E.g. the nth bit of the Commits bitmap will be 1 if\n\tthe nth object in the packfile index is a commit.\n\n\tThis is a very cheap operation because the bitmap writing code has\n\taccess to the metadata stored in the `struct object_entry` array,\n\tand hence the real type for each object in the packfile.\n\n\t2. `bitmap_writer_select_commits`: if bitmap writing is enabled for\n\ta given `pack-objects` run, the sequence of commits generated during\n\tthe Counting Objects phase will be stored in an array.\n\n\tWe then use that array to build up the list of selected commits.\n\tWriting a bitmap in the index for each object in the repository\n\twould be cost-prohibitive, so we use a simple heuristic to pick the\n\tcommits that will be indexed with bitmaps.\n\n\tThe current heuristics are a simplified version of JGit's original\n\timplementation. We select a higher density of commits depending on\n\ttheir age: the 100 most recent commits are always selected, after\n\tthat we pick 1 commit of each 100, and the gap increases as the\n\tcommits grow older. On top of that, we make sure that every single\n\tbranch that has not been merged (all the tips that would be required\n\tfrom a clone) gets their own bitmap, and when selecting commits\n\tbetween a gap, we tend to prioritize the commit with the most\n\tparents.\n\n\tDo note that there is no right/wrong way to perform commit selection;\n\tdifferent selection algorithms will result in different commits\n\tbeing selected, but there's no such thing as \"missing a commit\". The\n\tbitmap walker algorithm implemented in `prepare_bitmap_walk` is able\n\tto adapt to missing bitmaps by performing manual walks that complete\n\tthe bitmap: the ideal selection algorithm, however, would select\n\tthe commits that are more likely to be used as roots for a walk in\n\tthe future (e.g. the tips of each branch, and so on) to ensure a\n\tbitmap for them is always available.\n\n\t3. `bitmap_writer_build`: this is the computationally expensive part\n\tof bitmap generation. Based on the list of commits that were\n\tselected in the previous step, we perform several incremental walks\n\tto generate the bitmap for each commit.\n\n\tThe walks begin from the oldest commit, and are built up\n\tincrementally for each branch. E.g. consider this dag where A, B, C,\n\tD, E, F are the selected commits, and a, b, c, e are a chunk of\n\tsimplified history that will not receive bitmaps.\n\n\t\tA---a---B--b--C--c--D\n\t\t         \\\n\t\t          E--e--F\n\n\tWe start by building the bitmap for A, using A as the root for a\n\trevision walk and marking all the objects that are reachable until\n\tthe walk is over. Once this bitmap is stored, we reuse the bitmap\n\twalker to perform the walk for B, assuming that once we reach A\n\tagain, the walk will be terminated because A has already been SEEN\n\ton the previous walk.\n\n\tThis process is repeated for C, and D, but when we try to generate\n\tthe bitmaps for E, we cannot reuse neither the current walk nor the\n\tbitmap we have generated so far.\n\n\tWhat we do now is resetting both the walk and clearing the bitmap,\n\tand performing the walk from scratch using E as the origin. This new\n\twalk, however, does not need to be completed. Once we hit B, we can\n\tlookup the bitmap we have already stored for that commit and OR it\n\twith the existing bitmap we've composed so far, allowing us to limit\n\tthe walk early.\n\n\tAfter all the bitmaps have been generated, another iteration through\n\tthe list of commits is performed to find the best XOR offsets for\n\tcompression before writing them to disk. Because of the incremental\n\tnature of these bitmaps, XORing one of them with its predecesor\n\tresults in a minimal \"bitmap delta\" most of the time. We can write\n\tthis delta to the on-disk bitmap index, and then re-compose the\n\toriginal bitmaps by XORing them again when loaded.\n\n\tThis is a phase very similar to pack-object's `find_delta` (using\n\tbitmaps instead of objects, of course), except the heuristics have\n\tbeen greatly simplified: we only check the 10 bitmaps before any\n\tgiven one to find best compressing one. This operation gives optimal\n\tresults (again, because of the incremental nature of the bitmaps)\n\tand has a very good runtime performance because of the way EWAH\n\tbitmaps are implemented.\n\n\t3. `bitmap_writer_finish`: the last step in the process is\n\tserializing to disk all the bitmap data that has been generated in\n\tthe two previous steps.\n\n\tThe bitmap is written to a tmp file and then moved atomically to its\n\tfinal destination, using the same process as `pack-write.c:write_idx_file`.\n---\n Makefile               |    1 +\n builtin/pack-objects.c |  117 +++++++----\n builtin/pack-objects.h |   33 +++\n pack-bitmap-write.c    |  520 ++++++++++++++++++++++++++++++++++++++++++++++++\n pack-bitmap.h          |    9 +\n pack-write.c           |    2 +\n 6 files changed, 646 insertions(+), 36 deletions(-)\n create mode 100644 builtin/pack-objects.h\n create mode 100644 pack-bitmap-write.c\n\ndiff --git a/Makefile b/Makefile\nindex 0f2e72b..599aa59 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -840,6 +840,7 @@ LIB_OBJS += notes-cache.o\n LIB_OBJS += notes-merge.o\n LIB_OBJS += object.o\n LIB_OBJS += pack-bitmap.o\n+LIB_OBJS += pack-bitmap-write.o\n LIB_OBJS += pack-check.o\n LIB_OBJS += pack-revindex.o\n LIB_OBJS += pack-write.o\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 469b8da..58003ec 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -20,6 +20,7 @@\n #include \"thread-utils.h\"\n #include \"khash.h\"\n #include \"pack-bitmap.h\"\n+#include \"builtin/pack-objects.h\"\n \n static const char *pack_usage[] = {\n \tN_(\"git pack-objects --stdout [options...] [< ref-list | < object-list]\"),\n@@ -27,32 +28,6 @@ static const char *pack_usage[] = {\n \tNULL\n };\n \n-struct object_entry {\n-\tstruct pack_idx_entry idx;\n-\tunsigned long size;\t/* uncompressed size */\n-\tstruct packed_git *in_pack; \t/* already in pack */\n-\toff_t in_pack_offset;\n-\tstruct object_entry *delta;\t/* delta base object */\n-\tstruct object_entry *delta_child; /* deltified objects who bases me */\n-\tstruct object_entry *delta_sibling; /* other deltified objects who\n-\t\t\t\t\t     * uses the same base as me\n-\t\t\t\t\t     */\n-\tvoid *delta_data;\t/* cached delta (uncompressed) */\n-\tunsigned long delta_size;\t/* delta data size (uncompressed) */\n-\tunsigned long z_delta_size;\t/* delta data size (compressed) */\n-\tunsigned int hash;\t/* name hint hash */\n-\tenum object_type type;\n-\tenum object_type in_pack_type;\t/* could be delta */\n-\tunsigned char in_pack_header_size;\n-\tunsigned char preferred_base; /* we do not pack this, but is available\n-\t\t\t\t       * to be used as the base object to delta\n-\t\t\t\t       * objects against.\n-\t\t\t\t       */\n-\tunsigned char no_try_delta;\n-\tunsigned char tagged; /* near the very tip of refs */\n-\tunsigned char filled; /* assigned write-order */\n-};\n-\n /*\n  * Objects we are going to pack are collected in objects array (dynamically\n  * expanded).  nr_objects & nr_alloc controls this array.  They are stored\n@@ -86,6 +61,7 @@ static int pack_compression_seen;\n \n static int bitmap_support;\n static int use_bitmap_index;\n+static int write_bitmap_index;\n \n static unsigned long delta_cache_size = 0;\n static unsigned long max_delta_cache_size = 256 * 1024 * 1024;\n@@ -108,6 +84,12 @@ static struct object_entry *locate_object_entry(const unsigned char *sha1);\n static uint32_t written, written_delta;\n static uint32_t reused, reused_delta;\n \n+/*\n+ * Indexed commits\n+ */\n+struct commit **indexed_commits;\n+unsigned int indexed_commits_nr;\n+unsigned int indexed_commits_alloc;\n \n static struct object_slab {\n \tstruct object_slab *next;\n@@ -137,6 +119,16 @@ static struct object_entry *alloc_object_entry(void)\n \treturn &slab->data[slab->count++];\n }\n \n+static void index_commit_for_bitmap(struct commit *commit)\n+{\n+\tif (indexed_commits_nr >= indexed_commits_alloc) {\n+\t\tindexed_commits_alloc = (indexed_commits_alloc + 32) * 2;\n+\t\tindexed_commits = xrealloc(indexed_commits,\n+\t\t\tindexed_commits_alloc * sizeof(struct commit *));\n+\t}\n+\n+\tindexed_commits[indexed_commits_nr++] = commit;\n+}\n \n static void *get_delta(struct object_entry *entry)\n {\n@@ -746,6 +738,29 @@ static struct object_entry **compute_write_order(void)\n \treturn wo;\n }\n \n+static void resolve_real_types(\n+\t struct pack_idx_entry **index, uint32_t index_nr)\n+{\n+\tuint32_t i;\n+\n+\tfor (i = 0; i < index_nr; ++i) {\n+\t\tstruct object_entry *entry = (struct object_entry *)index[i];\n+\n+\t\tswitch (entry->type) {\n+\t\tcase OBJ_COMMIT:\n+\t\tcase OBJ_TREE:\n+\t\tcase OBJ_BLOB:\n+\t\tcase OBJ_TAG:\n+\t\t\tentry->real_type = entry->type;\n+\t\t\tbreak;\n+\n+\t\tdefault:\n+\t\t\tentry->real_type = sha1_object_info(entry->idx.sha1, NULL);\n+\t\t\tbreak;\n+\t\t}\n+\t}\n+}\n+\n static void write_pack_file(void)\n {\n \tuint32_t i = 0, j;\n@@ -824,9 +839,27 @@ static void write_pack_file(void)\n \t\t\tif (sizeof(tmpname) <= strlen(base_name) + 50)\n \t\t\t\tdie(\"pack base name '%s' too long\", base_name);\n \t\t\tsnprintf(tmpname, sizeof(tmpname), \"%s-\", base_name);\n+\n+\t\t\tif (write_bitmap_index)\n+\t\t\t\tresolve_real_types(written_list, nr_written);\n+\n \t\t\tfinish_tmp_packfile(tmpname, pack_tmp_name,\n \t\t\t\t\t    written_list, nr_written,\n \t\t\t\t\t    &pack_idx_opts, sha1);\n+\n+\t\t\tif (write_bitmap_index && nr_remaining == nr_written) {\n+\t\t\t\tchar *end_of_name_prefix = strrchr(tmpname, 0);\n+\t\t\t\tsprintf(end_of_name_prefix, \"%s.bitmap\", sha1_to_hex(sha1));\n+\n+\t\t\t\tstop_progress(&progress_state);\n+\n+\t\t\t\tbitmap_writer_show_progress(progress);\n+\t\t\t\tbitmap_writer_build_type_index(written_list, nr_written);\n+\t\t\t\tbitmap_writer_select_commits(indexed_commits, indexed_commits_nr, -1);\n+\t\t\t\tbitmap_writer_build(packed_objects);\n+\t\t\t\tbitmap_writer_finish(tmpname, sha1, BITMAP_OPT_HASH_CACHE);\n+\t\t\t}\n+\n \t\t\tfree(pack_tmp_name);\n \t\t\tputs(sha1_to_hex(sha1));\n \t\t}\n@@ -900,10 +933,8 @@ static int add_object_entry_1(const unsigned char *sha1, enum object_type type,\n \t\treturn 0;\n \t}\n \n-\tif (!exclude && local && has_loose_object_nonlocal(sha1)) {\n-\t\tkh_del_sha1(packed_objects, ix);\n-\t\treturn 0;\n-\t}\n+\tif (!exclude && local && has_loose_object_nonlocal(sha1))\n+\t\tgoto skip_entry;\n \n \tif (!found_pack) {\n \t\tfor (p = packed_git; p; p = p->next) {\n@@ -919,12 +950,12 @@ static int add_object_entry_1(const unsigned char *sha1, enum object_type type,\n \t\t\t\t}\n \t\t\t\tif (exclude)\n \t\t\t\t\tbreak;\n-\t\t\t\tif (incremental ||\n-\t\t\t\t\t(local && !p->pack_local) ||\n-\t\t\t\t\t(ignore_packed_keep && p->pack_local && p->pack_keep)) {\n-\t\t\t\t\tkh_del_sha1(packed_objects, ix);\n-\t\t\t\t\treturn 0;\n-\t\t\t\t}\n+\t\t\t\tif (incremental)\n+\t\t\t\t\tgoto skip_entry;\n+\t\t\t\tif (local && !p->pack_local)\n+\t\t\t\t\tgoto skip_entry;\n+\t\t\t\tif (ignore_packed_keep && p->pack_local && p->pack_keep)\n+\t\t\t\t\tgoto skip_entry;\n \t\t\t}\n \t\t}\n \t}\n@@ -956,6 +987,11 @@ static int add_object_entry_1(const unsigned char *sha1, enum object_type type,\n \tdisplay_progress(progress_state, nr_objects);\n \n \treturn 1;\n+\n+skip_entry:\n+\tkh_del_sha1(packed_objects, ix);\n+\twrite_bitmap_index = 0;\n+\treturn 0;\n }\n \n static int add_object_entry(const unsigned char *sha1, enum object_type type,\n@@ -1266,6 +1302,7 @@ static void check_object(struct object_entry *entry)\n \t\tused = unpack_object_header_buffer(buf, avail,\n \t\t\t\t\t\t   &entry->in_pack_type,\n \t\t\t\t\t\t   &entry->size);\n+\n \t\tif (used == 0)\n \t\t\tgoto give_up;\n \n@@ -2197,6 +2234,10 @@ static void show_commit(struct commit *commit, void *data)\n {\n \tadd_object_entry(commit->object.sha1, OBJ_COMMIT, NULL, 0);\n \tcommit->object.flags |= OBJECT_ADDED;\n+\n+\tif (write_bitmap_index) {\n+\t\tindex_commit_for_bitmap(commit);\n+\t}\n }\n \n static void show_object(struct object *obj,\n@@ -2366,6 +2407,7 @@ static void get_object_list(int ac, const char **av)\n \t\tif (*line == '-') {\n \t\t\tif (!strcmp(line, \"--not\")) {\n \t\t\t\tflags ^= UNINTERESTING;\n+\t\t\t\twrite_bitmap_index = 0;\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tdie(\"not a rev '%s'\", line);\n@@ -2588,6 +2630,9 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tdie(\"--keep-unreachable and --unpack-unreachable are incompatible.\");\n \n \tif (bitmap_support) {\n+\t\tif (!pack_to_stdout && rev_list_all)\n+\t\t\twrite_bitmap_index = 1;\n+\n \t\tif (use_internal_rev_list && pack_to_stdout)\n \t\t\tuse_bitmap_index = 1;\n \t}\ndiff --git a/builtin/pack-objects.h b/builtin/pack-objects.h\nnew file mode 100644\nindex 0000000..e186161\n--- /dev/null\n+++ b/builtin/pack-objects.h\n@@ -0,0 +1,33 @@\n+#ifndef BUILTIN_PACK_OBJECTS_H\n+#define BUILTIN_PACK_OBJECTS_H\n+\n+struct object_entry {\n+\tstruct pack_idx_entry idx;\n+\tunsigned long size;\t/* uncompressed size */\n+\tstruct packed_git *in_pack; \t/* already in pack */\n+\toff_t in_pack_offset;\n+\tstruct object_entry *delta;\t/* delta base object */\n+\tstruct object_entry *delta_child; /* deltified objects who bases me */\n+\tstruct object_entry *delta_sibling; /* other deltified objects who\n+\t\t\t\t\t     * uses the same base as me\n+\t\t\t\t\t     */\n+\tvoid *delta_data;\t/* cached delta (uncompressed) */\n+\tunsigned long delta_size;\t/* delta data size (uncompressed) */\n+\tunsigned long z_delta_size;\t/* delta data size (compressed) */\n+\tunsigned int hash;\t/* name hint hash */\n+\n+\tenum object_type type;\n+\tenum object_type in_pack_type;\n+\tenum object_type real_type;\n+\n+\tunsigned int index_pos;\n+\n+\tunsigned char in_pack_header_size;\n+\tunsigned char preferred_base;\n+\tunsigned char no_try_delta;\n+\tunsigned char tagged;\n+\tunsigned char filled;\n+\tunsigned char refered;\n+};\n+\n+#endif\ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nnew file mode 100644\nindex 0000000..d232545\n--- /dev/null\n+++ b/pack-bitmap-write.c\n@@ -0,0 +1,520 @@\n+#include <stdlib.h>\n+\n+#include \"cache.h\"\n+#include \"commit.h\"\n+#include \"tag.h\"\n+#include \"diff.h\"\n+#include \"revision.h\"\n+#include \"list-objects.h\"\n+#include \"progress.h\"\n+#include \"pack-revindex.h\"\n+#include \"pack.h\"\n+#include \"pack-bitmap.h\"\n+#include \"builtin/pack-objects.h\"\n+\n+struct bitmapped_commit {\n+\tstruct commit *commit;\n+\tstruct ewah_bitmap *bitmap;\n+\tstruct ewah_bitmap *write_as;\n+\tint flags;\n+\tint xor_offset;\n+\tuint32_t write_pos;\n+};\n+\n+struct bitmap_writer {\n+\tstruct ewah_bitmap *commits;\n+\tstruct ewah_bitmap *trees;\n+\tstruct ewah_bitmap *blobs;\n+\tstruct ewah_bitmap *tags;\n+\n+\tkhash_sha1 *bitmaps;\n+\tkhash_sha1 *packed_objects;\n+\n+\tstruct bitmapped_commit *selected;\n+\tunsigned int selected_nr, selected_alloc;\n+\n+\tstruct object_entry **index;\n+\tuint32_t index_nr;\n+\n+\tint fd;\n+\tuint32_t written;\n+\n+\tstruct progress *progress;\n+\tint show_progress;\n+};\n+\n+static struct bitmap_writer writer;\n+\n+void bitmap_writer_show_progress(int show)\n+{\n+\twriter.show_progress = show;\n+}\n+\n+/**\n+ * Build the initial type index for the packfile\n+ */\n+void bitmap_writer_build_type_index(\n+\t struct pack_idx_entry **index, uint32_t index_nr)\n+{\n+\tuint32_t i = 0;\n+\n+\tif (writer.show_progress)\n+\t\twriter.progress = start_progress(\"Building bitmap type index\", index_nr);\n+\n+\twriter.commits = ewah_new();\n+\twriter.trees = ewah_new();\n+\twriter.blobs = ewah_new();\n+\twriter.tags = ewah_new();\n+\n+\twriter.index = (struct object_entry **)index;\n+\twriter.index_nr = index_nr;\n+\n+\twhile (i < index_nr) {\n+\t\tstruct object_entry *entry = (struct object_entry *)index[i];\n+\t\tentry->index_pos = i;\n+\n+\t\tswitch (entry->real_type) {\n+\t\tcase OBJ_COMMIT:\n+\t\t\tewah_set(writer.commits, i);\n+\t\t\tbreak;\n+\n+\t\tcase OBJ_TREE:\n+\t\t\tewah_set(writer.trees, i);\n+\t\t\tbreak;\n+\n+\t\tcase OBJ_BLOB:\n+\t\t\tewah_set(writer.blobs, i);\n+\t\t\tbreak;\n+\n+\t\tcase OBJ_TAG:\n+\t\t\tewah_set(writer.tags, i);\n+\t\t\tbreak;\n+\n+\t\tdefault:\n+\t\t\tdie(\"Missing type information for %s (%d/%d)\",\n+\t\t\t\t\tsha1_to_hex(entry->idx.sha1), entry->real_type, entry->type);\n+\t\t}\n+\n+\t\ti++;\n+\t\tdisplay_progress(writer.progress, i);\n+\t}\n+\n+\tstop_progress(&writer.progress);\n+}\n+\n+/**\n+ * Compute the actual bitmaps\n+ */\n+static struct object **seen_objects;\n+static unsigned int seen_objects_nr, seen_objects_alloc;\n+\n+static inline void push_bitmapped_commit(struct commit *commit)\n+{\n+\tif (writer.selected_nr >= writer.selected_alloc) {\n+\t\twriter.selected_alloc = (writer.selected_alloc + 32) * 2;\n+\t\twriter.selected = xrealloc(writer.selected,\n+\t\t\twriter.selected_alloc * sizeof(struct bitmapped_commit));\n+\t}\n+\n+\twriter.selected[writer.selected_nr].commit = commit;\n+\twriter.selected[writer.selected_nr].bitmap = NULL;\n+\twriter.selected[writer.selected_nr].flags = 0;\n+\n+\twriter.selected_nr++;\n+}\n+\n+static inline void mark_as_seen(struct object *object)\n+{\n+\tif (seen_objects_nr >= seen_objects_alloc) {\n+\t\tseen_objects_alloc = (seen_objects_alloc + 32) * 2;\n+\t\tseen_objects = xrealloc(seen_objects,\n+\t\t\tseen_objects_alloc * sizeof(struct object*));\n+\t}\n+\n+\tseen_objects[seen_objects_nr++] = object;\n+}\n+\n+static inline void reset_all_seen(void)\n+{\n+\tunsigned int i;\n+\tfor (i = 0; i < seen_objects_nr; ++i) {\n+\t\tseen_objects[i]->flags &= ~(SEEN | ADDED | SHOWN);\n+\t}\n+\tseen_objects_nr = 0;\n+}\n+\n+static uint32_t find_object_pos(const unsigned char *sha1)\n+{\n+\tkhiter_t pos = kh_get_sha1(writer.packed_objects, sha1);\n+\n+\tif (pos < kh_end(writer.packed_objects)) {\n+\t\tstruct object_entry *entry = kh_value(writer.packed_objects, pos);\n+\t\treturn entry->index_pos;\n+\t}\n+\n+\tdie(\"Failed to write bitmap index. Packfile doesn't have full closure \"\n+\t\t\"(object %s is missing)\", sha1_to_hex(sha1));\n+}\n+\n+static void show_object(struct object *object,\n+\tconst struct name_path *path, const char *last, void *data)\n+{\n+\tstruct bitmap *base = data;\n+\tbitmap_set(base, find_object_pos(object->sha1));\n+\tmark_as_seen(object);\n+}\n+\n+static void show_commit(struct commit *commit, void *data)\n+{\n+\tmark_as_seen((struct object *)commit);\n+}\n+\n+static int\n+add_to_include_set(struct bitmap *base, struct commit *commit)\n+{\n+\tkhiter_t hash_pos;\n+\tuint32_t bitmap_pos = find_object_pos(commit->object.sha1);\n+\n+\tif (bitmap_get(base, bitmap_pos))\n+\t\treturn 0;\n+\n+\thash_pos = kh_get_sha1(writer.bitmaps, commit->object.sha1);\n+\tif (hash_pos < kh_end(writer.bitmaps)) {\n+\t\tstruct bitmapped_commit *bc = kh_value(writer.bitmaps, hash_pos);\n+\t\tbitmap_or_inplace(base, bc->bitmap);\n+\t\treturn 0;\n+\t}\n+\n+\tbitmap_set(base, bitmap_pos);\n+\treturn 1;\n+}\n+\n+static int\n+should_include(struct commit *commit, void *_data)\n+{\n+\tstruct bitmap *base = _data;\n+\n+\tif (!add_to_include_set(base, commit)) {\n+\t\tstruct commit_list *parent = commit->parents;\n+\n+\t\tmark_as_seen((struct object *)commit);\n+\n+\t\twhile (parent) {\n+\t\t\tparent->item->object.flags |= SEEN;\n+\t\t\tmark_as_seen((struct object *)parent->item);\n+\t\t\tparent = parent->next;\n+\t\t}\n+\n+\t\treturn 0;\n+\t}\n+\n+\treturn 1;\n+}\n+\n+static void\n+compute_xor_offsets(void)\n+{\n+\tstatic const int MAX_XOR_OFFSET_SEARCH = 10;\n+\n+\tint i, next = 0;\n+\n+\twhile (next < writer.selected_nr) {\n+\t\tstruct bitmapped_commit *stored = &writer.selected[next];\n+\n+\t\tint best_offset = 0;\n+\t\tstruct ewah_bitmap *best_bitmap = stored->bitmap;\n+\t\tstruct ewah_bitmap *test_xor;\n+\n+\t\tfor (i = 1; i <= MAX_XOR_OFFSET_SEARCH; ++i) {\n+\t\t\tint curr = next - i;\n+\n+\t\t\tif (curr < 0)\n+\t\t\t\tbreak;\n+\n+\t\t\ttest_xor = ewah_pool_new();\n+\t\t\tewah_xor(writer.selected[curr].bitmap, stored->bitmap, test_xor);\n+\n+\t\t\tif (test_xor->buffer_size < best_bitmap->buffer_size) {\n+\t\t\t\tif (best_bitmap != stored->bitmap)\n+\t\t\t\t\tewah_pool_free(best_bitmap);\n+\n+\t\t\t\tbest_bitmap = test_xor;\n+\t\t\t\tbest_offset = i;\n+\t\t\t} else {\n+\t\t\t\tewah_pool_free(test_xor);\n+\t\t\t}\n+\t\t}\n+\n+\t\tstored->xor_offset = best_offset;\n+\t\tstored->write_as = best_bitmap;\n+\n+\t\tnext++;\n+\t}\n+}\n+\n+void\n+bitmap_writer_build(khash_sha1 *packed_objects)\n+{\n+\tint i;\n+\tstruct bitmap *base = bitmap_new();\n+\tstruct rev_info revs;\n+\n+\twriter.bitmaps = kh_init_sha1();\n+\twriter.packed_objects = packed_objects;\n+\n+\tif (writer.show_progress)\n+\t\twriter.progress = start_progress(\"Building bitmaps\", writer.selected_nr);\n+\n+\tinit_revisions(&revs, NULL);\n+\trevs.tag_objects = 1;\n+\trevs.tree_objects = 1;\n+\trevs.blob_objects = 1;\n+\trevs.no_walk = 0;\n+\n+\trevs.include_check = should_include;\n+\treset_revision_walk();\n+\n+\tfor (i = writer.selected_nr - 1; i >= 0; --i) {\n+\t\tstruct bitmapped_commit *stored;\n+\t\tstruct object *object;\n+\n+\t\tkhiter_t hash_pos;\n+\t\tint hash_ret;\n+\n+\t\tstored = &writer.selected[i];\n+\t\tobject = (struct object *)stored->commit;\n+\n+\t\tif (i < writer.selected_nr - 1) {\n+\t\t\tif (!in_merge_bases(writer.selected[i + 1].commit, stored->commit)) {\n+\t\t\t\tbitmap_reset(base);\n+\t\t\t\treset_all_seen();\n+\t\t\t}\n+\t\t}\n+\n+\t\tadd_pending_object(&revs, object, \"\");\n+\t\trevs.include_check_data = base;\n+\n+\t\tif (prepare_revision_walk(&revs))\n+\t\t\tdie(\"revision walk setup failed\");\n+\n+\t\ttraverse_commit_list(&revs, show_commit, show_object, base);\n+\n+\t\trevs.pending.nr = 0;\n+\t\trevs.pending.alloc = 0;\n+\t\trevs.pending.objects = NULL;\n+\n+\t\tstored->bitmap = bitmap_to_ewah(base);\n+\t\tstored->flags = object->flags;\n+\n+\t\thash_pos = kh_put_sha1(writer.bitmaps, object->sha1, &hash_ret);\n+\t\tif (hash_ret == 0)\n+\t\t\tdie(\"Duplicate entry when writing index: %s\",\n+\t\t\t\tsha1_to_hex(object->sha1));\n+\n+\t\tkh_value(writer.bitmaps, hash_pos) = stored;\n+\n+\t\tdisplay_progress(writer.progress, writer.selected_nr - i);\n+\t}\n+\n+\tbitmap_free(base);\n+\tstop_progress(&writer.progress);\n+\n+\tcompute_xor_offsets();\n+}\n+\n+/**\n+ * Select the commits that will be bitmapped\n+ */\n+static inline unsigned int next_commit_index(unsigned int idx)\n+{\n+\tstatic const unsigned int MIN_COMMITS = 100;\n+\tstatic const unsigned int MAX_COMMITS = 5000;\n+\n+\tstatic const unsigned int MUST_REGION = 100;\n+\tstatic const unsigned int MIN_REGION = 20000;\n+\n+\tunsigned int offset, next;\n+\n+\tif (idx <= MUST_REGION)\n+\t\treturn 0;\n+\n+\tif (idx <= MIN_REGION) {\n+\t\toffset = idx - MUST_REGION;\n+\t\treturn (offset < MIN_COMMITS) ? offset : MIN_COMMITS;\n+\t}\n+\n+\toffset = idx - MIN_REGION;\n+\tnext = (offset < MAX_COMMITS) ? offset : MAX_COMMITS;\n+\n+\treturn (next > MIN_COMMITS) ? next : MIN_COMMITS;\n+}\n+\n+void bitmap_writer_select_commits(\n+\t\tstruct commit **indexed_commits,\n+\t\tunsigned int indexed_commits_nr,\n+\t\tint max_bitmaps)\n+{\n+\tunsigned int i = 0, next;\n+\n+\tif (writer.show_progress)\n+\t\twriter.progress = start_progress(\"Selecting bitmap commits\", 0);\n+\n+\tif (indexed_commits_nr < 100) {\n+\t\tfor (i = 0; i < indexed_commits_nr; ++i) {\n+\t\t\tpush_bitmapped_commit(indexed_commits[i]);\n+\t\t}\n+\t\treturn;\n+\t}\n+\n+\tfor (;;) {\n+\t\tnext = next_commit_index(i);\n+\n+\t\tif (i + next >= indexed_commits_nr)\n+\t\t\tbreak;\n+\n+\t\tif (max_bitmaps > 0 && writer.selected_nr >= max_bitmaps) {\n+\t\t\twriter.selected_nr = max_bitmaps;\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\tif (next == 0) {\n+\t\t\tpush_bitmapped_commit(indexed_commits[i]);\n+\t\t} else {\n+\t\t\tunsigned int j;\n+\t\t\tstruct commit *chosen = indexed_commits[i + next];\n+\n+\t\t\tfor (j = 0; j <= next; ++j) {\n+\t\t\t\tstruct commit *cm = indexed_commits[i + j];\n+\t\t\t\tif (cm->parents && cm->parents->next)\n+\t\t\t\t\tchosen = cm;\n+\t\t\t}\n+\n+\t\t\tpush_bitmapped_commit(chosen);\n+\t\t}\n+\n+\t\ti += next + 1;\n+\t\tdisplay_progress(writer.progress, i);\n+\t}\n+\n+\tstop_progress(&writer.progress);\n+}\n+\n+/**\n+ * Write the bitmap index to disk\n+ */\n+static void write_hash_table(\n+\t struct object_entry **index, uint32_t index_nr)\n+{\n+\tuint32_t i, j = 0;\n+\tuint32_t buffer[1024];\n+\n+\tfor (i = 0; i < index_nr; ++i) {\n+\t\tstruct object_entry *entry = index[i];\n+\n+\t\tbuffer[j++] = htonl(entry->hash);\n+\t\tif (j == 1024) {\n+\t\t\twrite_or_die(writer.fd, buffer, sizeof(buffer));\n+\t\t\tj = 0;\n+\t\t}\n+\t}\n+\n+\tif (j > 0) {\n+\t\twrite_or_die(writer.fd, buffer, j * sizeof(uint32_t));\n+\t}\n+\n+\twriter.written += (index_nr * sizeof(uint32_t));\n+}\n+\n+static void dump_bitmap(struct ewah_bitmap *bitmap)\n+{\n+\tint written;\n+\n+#if __BYTE_ORDER == __LITTLE_ENDIAN\n+\twritten = ewah_serialize_native(bitmap, writer.fd);\n+#else\n+\twritten = ewah_serialize(bitmap, writer.fd);\n+#endif\n+\n+\tif (written < 0)\n+\t\tdie(\"Failed to write bitmap index\");\n+\n+\twriter.written += written;\n+}\n+\n+static void\n+write_selected_commits_v2(void)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i < writer.selected_nr; ++i) {\n+\t\tstruct bitmapped_commit *stored = &writer.selected[i];\n+\t\tstored->write_pos = writer.written;\n+\t\tdump_bitmap(stored->write_as);\n+\t}\n+\n+\tfor (i = 0; i < writer.selected_nr; ++i) {\n+\t\tstruct bitmapped_commit *stored = &writer.selected[i];\n+\t\tstruct bitmap_disk_entry_v2 on_disk;\n+\n+\t\tmemcpy(on_disk.sha1, stored->commit->object.sha1, 20);\n+\t\ton_disk.bitmap_pos = htonl(stored->write_pos);\n+\t\ton_disk.xor_offset = stored->xor_offset;\n+\t\ton_disk.flags = stored->flags;\n+\n+\t\twrite_or_die(writer.fd, &on_disk, sizeof(on_disk));\n+\t\twriter.written += sizeof(on_disk);\n+\t}\n+}\n+\n+void bitmap_writer_finish(\n+\t const char *filename, unsigned char sha1[], uint16_t flags)\n+{\n+\tstatic char tmp_file[PATH_MAX];\n+\tstatic uint16_t default_version = 2;\n+\n+\tstruct bitmap_disk_header header;\n+\n+\tflags |= BITMAP_OPT_FULL_DAG;\n+\n+#if __BYTE_ORDER == __LITTLE_ENDIAN\n+\t/*\n+\t * In little endian machines (i.e. most of them) we're\n+\t * going to dump the bitmaps straight from memory into\n+\t * disk, and tag the bitmap index as having LE bitmaps\n+\t */\n+\tflags |= BITMAP_OPT_LE_BITMAPS;\n+#else\n+\tflags |= BITMAP_OPT_BE_BITMAPS;\n+#endif\n+\n+\twriter.fd = odb_mkstemp(tmp_file, sizeof(tmp_file), \"pack/tmp_bitmap_XXXXXX\");\n+\n+\tif (writer.fd < 0)\n+\t\tdie_errno(\"unable to create '%s'\", tmp_file);\n+\n+\tmemcpy(header.magic, BITMAP_MAGIC_PREFIX, sizeof(BITMAP_MAGIC_PREFIX));\n+\theader.version = htons(default_version);\n+\theader.options = htons(flags);\n+\theader.entry_count = htonl(writer.selected_nr);\n+\tmemcpy(header.checksum, sha1, 20);\n+\n+\twrite_or_die(writer.fd, &header, sizeof(header));\n+\twriter.written += sizeof(header);\n+\n+\tif (flags & BITMAP_OPT_HASH_CACHE)\n+\t\twrite_hash_table(writer.index, writer.index_nr);\n+\n+\tdump_bitmap(writer.commits);\n+\tdump_bitmap(writer.trees);\n+\tdump_bitmap(writer.blobs);\n+\tdump_bitmap(writer.tags);\n+\twrite_selected_commits_v2();\n+\n+\tclose(writer.fd);\n+\n+\tif (adjust_shared_perm(tmp_file))\n+\t\tdie_errno(\"unable to make temporary bitmap file readable\");\n+\n+\tif (rename(tmp_file, filename))\n+\t\tdie_errno(\"unable to rename temporary bitmap file to '%s'\", filename);\n+}\ndiff --git a/pack-bitmap.h b/pack-bitmap.h\nindex b97bd46..8e7e3dc 100644\n--- a/pack-bitmap.h\n+++ b/pack-bitmap.h\n@@ -31,6 +31,8 @@ struct bitmap_disk_header {\n \n static const char BITMAP_MAGIC_PREFIX[] = {'B', 'I', 'T', 'M'};;\n \n+#define NEEDS_BITMAP (1u<<22)\n+\n enum pack_bitmap_opts {\n \tBITMAP_OPT_FULL_DAG = 1,\n \tBITMAP_OPT_LE_BITMAPS = 2,\n@@ -50,4 +52,11 @@ int prepare_bitmap_walk(struct rev_info *revs, uint32_t *result_size);\n void test_bitmap_walk(struct rev_info *revs);\n char *pack_bitmap_filename(struct packed_git *p);\n \n+void bitmap_writer_show_progress(int show);\n+void bitmap_writer_build_type_index(struct pack_idx_entry **index, uint32_t index_nr);\n+void bitmap_writer_select_commits(struct commit **indexed_commits,\n+\t\tunsigned int indexed_commits_nr, int max_bitmaps);\n+void bitmap_writer_build(khash_sha1 *packed_objects);\n+void bitmap_writer_finish(const char *filename, unsigned char sha1[], uint16_t flags);\n+\n #endif\ndiff --git a/pack-write.c b/pack-write.c\nindex ca9e63b..6203d37 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -371,5 +371,7 @@ void finish_tmp_packfile(char *name_buffer,\n \tif (rename(idx_tmp_name, name_buffer))\n \t\tdie_errno(\"unable to rename temporary index file\");\n \n+\t*end_of_name_prefix = '\\0';\n+\n \tfree((void *)idx_tmp_name);\n }\n-- \n1.7.9.5\n"},{"id":"221913","messageId":"1372116193-32762-14-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 13/16] repack: consider bitmaps when performing repacks","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:10Z","receivedAt":"2013-06-24T23:23:10Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"Since `pack-objects` will write a `.bitmap` file next to the `.pack` and\n`.idx` files, this commit teaches `git-repack` to consider the new\nbitmap indexes (if they exist) when performing repack operations.\n\nThis implies moving old bitmap indexes out of the way if we are\nrepacking a repository that already has them, and moving the newly\ngenerated bitmap indexes into the `objects/pack` directory, next to\ntheir corresponding packfiles.\n\nSince `git repack` is now capable of handling these `.bitmap` files,\na normal `git gc` run on a repository that has `pack.usebitmaps` set\nto true in its config file will generate bitmap indexes as part of the\ngarbage collection process.\n---\n git-repack.sh |   10 ++++++++--\n 1 file changed, 8 insertions(+), 2 deletions(-)\n\ndiff --git a/git-repack.sh b/git-repack.sh\nindex 7579331..d5355ae 100755\n--- a/git-repack.sh\n+++ b/git-repack.sh\n@@ -108,7 +108,7 @@ rollback=\n failed=\n for name in $names\n do\n-\tfor sfx in pack idx\n+\tfor sfx in pack idx bitmap\n \tdo\n \t\tfile=pack-$name.$sfx\n \t\ttest -f \"$PACKDIR/$file\" || continue\n@@ -156,6 +156,11 @@ do\n \tfullbases=\"$fullbases pack-$name\"\n \tchmod a-w \"$PACKTMP-$name.pack\"\n \tchmod a-w \"$PACKTMP-$name.idx\"\n+\n+\ttest -f \"$PACKTMP-$name.bitmap\" &&\n+\tchmod a-w \"$PACKTMP-$name.bitmap\" &&\n+\tmv -f \"$PACKTMP-$name.bitmap\" \"$PACKDIR/pack-$name.bitmap\"\n+\n \tmv -f \"$PACKTMP-$name.pack\" \"$PACKDIR/pack-$name.pack\" &&\n \tmv -f \"$PACKTMP-$name.idx\"  \"$PACKDIR/pack-$name.idx\" ||\n \texit\n@@ -166,6 +171,7 @@ for name in $names\n do\n \trm -f \"$PACKDIR/old-pack-$name.idx\"\n \trm -f \"$PACKDIR/old-pack-$name.pack\"\n+\trm -f \"$PACKDIR/old-pack-$name.bitmap\"\n done\n \n # End of pack replacement.\n@@ -180,7 +186,7 @@ then\n \t\t  do\n \t\t\tcase \" $fullbases \" in\n \t\t\t*\" $e \"*) ;;\n-\t\t\t*)\trm -f \"$e.pack\" \"$e.idx\" \"$e.keep\" ;;\n+\t\t\t*)\trm -f \"$e.pack\" \"$e.idx\" \"$e.keep\" \"$e.bitmap\" ;;\n \t\t\tesac\n \t\t  done\n \t\t)\n-- \n1.7.9.5\n"},{"id":"221911","messageId":"1372116193-32762-15-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 14/16] sha1_file: implement `nth_packed_object_info`","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:11Z","receivedAt":"2013-06-24T23:23:11Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"A new helper function allows to efficiently query the size and real type\nof an object in a packfile based on its position on the packfile index.\n\nThis is particularly useful when trying to parse all the information of\nan index in memory.\n---\n cache.h     |    1 +\n sha1_file.c |    6 ++++++\n 2 files changed, 7 insertions(+)\n\ndiff --git a/cache.h b/cache.h\nindex bbe5e2a..26e4567 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1104,6 +1104,7 @@ extern void clear_delta_base_cache(void);\n extern struct packed_git *add_packed_git(const char *, int, int);\n extern const unsigned char *nth_packed_object_sha1(struct packed_git *, uint32_t);\n extern off_t nth_packed_object_offset(const struct packed_git *, uint32_t);\n+extern int nth_packed_object_info(struct packed_git *p, uint32_t n, unsigned long *sizep);\n extern int find_pack_entry_pos(const unsigned char *sha1, struct packed_git *p);\n extern off_t find_pack_entry_one(const unsigned char *, struct packed_git *);\n extern int is_pack_valid(struct packed_git *);\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 018a847..fd5bd01 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -2223,6 +2223,12 @@ off_t nth_packed_object_offset(const struct packed_git *p, uint32_t n)\n \t}\n }\n \n+int nth_packed_object_info(struct packed_git *p, uint32_t n, unsigned long *sizep)\n+{\n+\toff_t offset = nth_packed_object_offset(p, n);\n+\treturn packed_object_info(p, offset, sizep, NULL);\n+}\n+\n int find_pack_entry_pos(const unsigned char *sha1, struct packed_git *p)\n {\n \tconst uint32_t *level1_ofs = p->index_data;\n-- \n1.7.9.5\n"},{"id":"221914","messageId":"1372116193-32762-16-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 15/16] write-bitmap: implement new git command to write bitmaps","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:12Z","receivedAt":"2013-06-24T23:23:12Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"The `pack-objects` builtin is capable of writing out bitmap indexes\n(.bitmap) next to the their corresponding packfile, as part of the\nprocess of actually generating the packfile.\n\nThis is a very efficient operation because all the required data for\nwriting the bitmap index (commit traversal list, list of all objects in\na packfile, sorted index for the packfile, and types for all objects in\nthe packfile) is readily available in memory as part of the process of\nbuilding the packfile itself.\n\nThere are however cases when we want to generate a bitmap index for a\npackfile that already exists on disk (i.e. one we're not writing from\nscratch). This new git builtin implements the bitmap index equivalent of\n`git index-pack`: it writes a `.bitmap` file given a pair of existing\n`.pack` and `.idx` files.\n\n\tNOTE that `write-bitmap` requires the packfile to have been indexed\n\tbeforehand. If the packfile doesn't have its corresponding `.idx`\n\tfile, `git index-pack` must be called before `write-bitmap` can\n\twork.\n\nThe process of generating bitmaps for an existing packfile is as\nfollows:\n\n\t1. Load the existing pack index in memory. The `.idx` for the\n\tpackfile is loaded into a hash table in mememory so it can be\n\tefficiently queried. As part of this loading process, the real type\n\tfor each object in the packfile is resolved (this implies resolving\n\tdeltas, which can make this process rather expensive).\n\n\t2. Find the full closure for the packfile. All the objects from the\n\tpackfile that have been loaded in memory are iterated, looking for\n\tcommits. These commits are parsed, and their parents are marked as\n\tsuch to ensure that\n\n\t\ta) there is a full closure in the packfile, and no commit has a\n\t\tdangling parent pointer\n\n\t\tb) we can find the set of \"tips\" for the packfile, i.e., the set\n\t\tof commits that don't have any commits pointing to them\n\n\t3. The \"tips\" of the packfile are then used as the roots to perform a\n\tnormal revision walk. The result of this revision walk is the list\n\tof commits that will be used by `bitmap_writer_select_commits` when\n\tselecting which commits are going to be bitmapped.\n\n\t4. We build and write the bitmap index in the same way that\n\t`pack-objects` does, given that we have all the required metadata:\n\n\t\t- an array of all the objects in the packfile, in index order\n\t\t- the types of all these objects\n\t\t- an array with a walk-ordering of all the commits in the\n\t\t\tpackfile, which will be used for selection\n\t\t- a hash table to efficiently look up objects in the index\n\n\tSee the previous patch \"pack-objects: implement bitmap writing\" for\n\tdetails on how the bitmap computation happens.\n---\n Makefile               |    1 +\n builtin.h              |    1 +\n builtin/write-bitmap.c |  256 ++++++++++++++++++++++++++++++++++++++++++++++++\n git.c                  |    1 +\n 4 files changed, 259 insertions(+)\n create mode 100644 builtin/write-bitmap.c\n\ndiff --git a/Makefile b/Makefile\nindex 599aa59..4a0a7dd 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1000,6 +1000,7 @@ BUILTIN_OBJS += builtin/upload-archive.o\n BUILTIN_OBJS += builtin/var.o\n BUILTIN_OBJS += builtin/verify-pack.o\n BUILTIN_OBJS += builtin/verify-tag.o\n+BUILTIN_OBJS += builtin/write-bitmap.o\n BUILTIN_OBJS += builtin/write-tree.o\n \n GITLIBS = $(LIB_FILE) $(XDIFF_LIB)\ndiff --git a/builtin.h b/builtin.h\nindex 64bab6b..e39685f 100644\n--- a/builtin.h\n+++ b/builtin.h\n@@ -144,6 +144,7 @@ extern int cmd_var(int argc, const char **argv, const char *prefix);\n extern int cmd_verify_tag(int argc, const char **argv, const char *prefix);\n extern int cmd_version(int argc, const char **argv, const char *prefix);\n extern int cmd_whatchanged(int argc, const char **argv, const char *prefix);\n+extern int cmd_write_bitmap(int argc, const char **argv, const char *prefix);\n extern int cmd_write_tree(int argc, const char **argv, const char *prefix);\n extern int cmd_verify_pack(int argc, const char **argv, const char *prefix);\n extern int cmd_show_ref(int argc, const char **argv, const char *prefix);\ndiff --git a/builtin/write-bitmap.c b/builtin/write-bitmap.c\nnew file mode 100644\nindex 0000000..0cc1c9e\n--- /dev/null\n+++ b/builtin/write-bitmap.c\n@@ -0,0 +1,256 @@\n+#include <stdlib.h>\n+\n+#include \"cache.h\"\n+#include \"commit.h\"\n+#include \"tag.h\"\n+#include \"diff.h\"\n+#include \"revision.h\"\n+#include \"progress.h\"\n+#include \"list-objects.h\"\n+#include \"pack.h\"\n+#include \"refs.h\"\n+#include \"pack-bitmap.h\"\n+\n+#include \"builtin/pack-objects.h\"\n+\n+static int progress = 1;\n+static struct progress *progress_state;\n+static int write_hash_cache;\n+\n+static struct object_entry **objects;\n+static uint32_t nr_objects;\n+\n+static struct commit **walked_commits;\n+static uint32_t nr_commits;\n+\n+static khash_sha1 *packed_objects;\n+\n+static struct object_entry *\n+allocate_entry(const unsigned char *sha1)\n+{\n+\tstruct object_entry *entry;\n+\tkhiter_t pos;\n+\tint hash_ret;\n+\n+\tentry = calloc(1, sizeof(struct object_entry));\n+\thashcpy(entry->idx.sha1, sha1);\n+\n+\tpos = kh_put_sha1(packed_objects, entry->idx.sha1, &hash_ret);\n+\tif (hash_ret == 0) {\n+\t\tdie(\"BUG: duplicate entry in packfile\");\n+\t}\n+\n+\tkh_value(packed_objects, pos) = entry;\n+\tobjects[nr_objects++] = entry;\n+\n+\treturn entry;\n+}\n+\n+static void\n+load_pack_index(struct packed_git *pack)\n+{\n+\tuint32_t i, commits_found = 0;\n+\tkhint_t new_hash_size, nr_alloc;\n+\n+\tif (open_pack_index(pack))\n+\t\tdie(\"Failed to load packfile\");\n+\n+\tnew_hash_size = (pack->num_objects * (1.0 / __ac_HASH_UPPER)) + 0.5;\n+\tkh_resize_sha1(packed_objects, new_hash_size);\n+\n+\tnr_alloc = (pack->num_objects + 63) & ~63;\n+\tobjects = xmalloc(nr_alloc * sizeof(struct object_entry *));\n+\n+\tif (progress)\n+\t\tprogress_state = start_progress(\"Loading existing index\", pack->num_objects);\n+\n+\tfor (i = 0; i < pack->num_objects; ++i) {\n+\t\tstruct object_entry *entry;\n+\t\tconst unsigned char *sha1;\n+\n+\t\tsha1 = nth_packed_object_sha1(pack, i);\n+\t\tentry = allocate_entry(sha1);\n+\n+\t\tentry->in_pack = pack;\n+\t\tentry->type = entry->real_type = nth_packed_object_info(pack, i, NULL);\n+\t\tentry->index_pos = i;\n+\n+\t\tdisplay_progress(progress_state, i + 1);\n+\t}\n+\n+\tstop_progress(&progress_state);\n+\tif (progress)\n+\t\tprogress_state = start_progress(\"Finding pack closure\", 0);\n+\n+\tfor (i = 0; i < nr_objects; ++i) {\n+\t\tstruct commit *commit;\n+\t\tstruct commit_list *parent;\n+\n+\t\tif (objects[i]->type != OBJ_COMMIT)\n+\t\t\tcontinue;\n+\n+\t\tcommit = lookup_commit(objects[i]->idx.sha1);\n+\t\tif (parse_commit(commit)) {\n+\t\t\tdie(\"Bad commit: %s\\n\", sha1_to_hex(objects[i]->idx.sha1));\n+\t\t}\n+\n+\t\tparent = commit->parents;\n+\n+\t\twhile (parent) {\n+\t\t\tkhiter_t pos = kh_get_sha1(packed_objects, parent->item->object.sha1);\n+\n+\t\t\tif (pos < kh_end(packed_objects)) {\n+\t\t\t\tstruct object_entry *entry = kh_value(packed_objects, pos);\n+\t\t\t\tentry->refered = 1;\n+\t\t\t} else {\n+\t\t\t\tdie(\"Failed to write bitmaps for packfile: No closure\");\n+\t\t\t}\n+\n+\t\t\tparent = parent->next;\n+\t\t}\n+\n+\t\tdisplay_progress(progress_state, ++commits_found);\n+\t}\n+\n+\tstop_progress(&progress_state);\n+}\n+\n+static void show_object(struct object *object,\n+\tconst struct name_path *path, const char *last, void *data)\n+{\n+\tchar *name = path_name(path, last);\n+\tkhiter_t pos = kh_get_sha1(packed_objects, object->sha1);\n+\n+\tif (pos < kh_end(packed_objects)) {\n+\t\tstruct object_entry *entry = kh_value(packed_objects, pos);\n+\t\tentry->hash = pack_name_hash(name);\n+\t}\n+\n+\tfree(name);\n+}\n+\n+static void show_commit(struct commit *commit, void *data)\n+{\n+\twalked_commits[nr_commits++] = commit;\n+\tdisplay_progress(progress_state, nr_commits);\n+}\n+\n+static void\n+find_all_objects(struct packed_git *pack)\n+{\n+\tstruct rev_info revs;\n+\tuint32_t i, found_commits = 0;\n+\n+\tinit_revisions(&revs, NULL);\n+\tif (write_hash_cache) {\n+\t\trevs.tag_objects = 1;\n+\t\trevs.tree_objects = 1;\n+\t\trevs.blob_objects = 1;\n+\t}\n+\trevs.no_walk = 0;\n+\n+\tfor (i = 0; i < nr_objects; ++i) {\n+\t\tif (objects[i]->type == OBJ_COMMIT) {\n+\t\t\tif (!objects[i]->refered) {\n+\t\t\t\tstruct object *object = parse_object(objects[i]->idx.sha1);\n+\t\t\t\tadd_pending_object(&revs, object, \"\");\n+\t\t\t}\n+\n+\t\t\tfound_commits++;\n+\t\t}\n+\t}\n+\n+\tif (progress)\n+\t\tprogress_state = start_progress(\"Computing walk order\", found_commits);\n+\n+\twalked_commits = xmalloc(found_commits * sizeof(struct commit *));\n+\n+\tif (prepare_revision_walk(&revs))\n+\t\tdie(\"revision walk setup failed\");\n+\n+\ttraverse_commit_list(&revs, show_commit, show_object, NULL);\n+\tstop_progress(&progress_state);\n+\n+\tif (found_commits != nr_commits)\n+\t\tdie(\"Missing commits in the walk? Got %d, expected %d\", i, nr_commits);\n+}\n+\n+static const char *write_bitmaps_usage[] = {\n+\tN_(\"git write-bitmap --hash-cache [options...] [pack-sha1]\"),\n+\tNULL\n+};\n+\n+int cmd_write_bitmap(int argc, const char **argv, const char *prefix)\n+{\n+\tint max_bitmaps = 0;\n+\n+\tstruct option write_bitmaps_options[] = {\n+\t\tOPT_SET_INT('q', \"quiet\", &progress,\n+\t\t\t    N_(\"do not show progress meter\"), 0),\n+\t\tOPT_SET_INT(0, \"progress\", &progress,\n+\t\t\t    N_(\"show progress meter\"), 1),\n+\t\tOPT_BOOL(0, \"hash-cache\", &write_hash_cache,\n+\t\t\t N_(\"Write a cache of hashes for delta resolution\")),\n+\t\tOPT_INTEGER(0, \"max\", &max_bitmaps,\n+\t\t\t    N_(\"max number of bitmaps to generate\")),\n+\t\tOPT_END(),\n+\t};\n+\n+\tstruct packed_git *p;\n+\tstruct packed_git *pack_to_index = NULL;\n+\tchar *bitmap_filename;\n+\tuint16_t write_flags;\n+\n+\tprogress = isatty(2);\n+\targc = parse_options(argc, argv, prefix,\n+\t\t\twrite_bitmaps_options, write_bitmaps_usage, 0);\n+\n+\tpacked_objects = kh_init_sha1();\n+\tprepare_packed_git();\n+\n+\tif (argc) {\n+\t\tunsigned char pack_sha[20];\n+\n+\t\tif (get_sha1_hex(argv[0], pack_sha))\n+\t\t\tdie(\"Invalid SHA1 for packfile\");\n+\n+\t\tfor (p = packed_git; p; p = p->next) {\n+\t\t\tif (hashcmp(p->sha1, pack_sha) == 0) {\n+\t\t\t\tpack_to_index = p;\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t}\n+\t} else {\n+\t\tpack_to_index = packed_git;\n+\n+\t\tfor (p = packed_git; p; p = p->next) {\n+\t\t\tif (p->pack_size > pack_to_index->pack_size)\n+\t\t\t\tpack_to_index = p;\n+\t\t}\n+\t}\n+\n+\tif (!pack_to_index)\n+\t\tdie(\"No packs found for indexing\");\n+\n+\tif (progress)\n+\t\tfprintf(stderr, \"Indexing 'pack-%s.pack'\\n\",\n+\t\t\tsha1_to_hex(pack_to_index->sha1));\n+\n+\tload_pack_index(pack_to_index);\n+\tfind_all_objects(pack_to_index);\n+\n+\tbitmap_filename = pack_bitmap_filename(pack_to_index);\n+\twrite_flags = 0;\n+\n+\tif (write_hash_cache)\n+\t\twrite_flags |= BITMAP_OPT_HASH_CACHE;\n+\n+\tbitmap_writer_show_progress(progress);\n+\tbitmap_writer_build_type_index((struct pack_idx_entry **)objects, nr_objects);\n+\tbitmap_writer_select_commits(walked_commits, nr_commits, max_bitmaps);\n+\tbitmap_writer_build(packed_objects);\n+\tbitmap_writer_finish(bitmap_filename, pack_to_index->sha1, write_flags);\n+\n+\tfree(bitmap_filename);\n+\treturn 0;\n+}\ndiff --git a/git.c b/git.c\nindex 4359086..66ceb2c 100644\n--- a/git.c\n+++ b/git.c\n@@ -426,6 +426,7 @@ static void handle_internal_command(int argc, const char **argv)\n \t\t{ \"verify-tag\", cmd_verify_tag, RUN_SETUP },\n \t\t{ \"version\", cmd_version },\n \t\t{ \"whatchanged\", cmd_whatchanged, RUN_SETUP },\n+\t\t{ \"write-bitmap\", cmd_write_bitmap, RUN_SETUP },\n \t\t{ \"write-tree\", cmd_write_tree, RUN_SETUP },\n \t};\n \tint i;\n-- \n1.7.9.5\n"},{"id":"221912","messageId":"1372116193-32762-17-git-send-email-tanoku@gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"[PATCH 16/16] rev-list: Optimize --count using bitmaps too","fromName":"Vicent Marti","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-24T23:23:13Z","receivedAt":"2013-06-24T23:23:13Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"If bitmap indexes are available, the process of counting reachable\ncommits with `git rev-list --count` can be greatly sped up. Instead of\nhaving to use callbacks that yield each object in the revision list, we\ncan build the reachable bitmap for the list and then use an efficient\npopcount to find the number of bits set in the bitmap.\n\nThis commit implements a `count_bitmap_commit_list` that can be used\nafter `prepare_bitmap_walk` has returned successfully to return the\nnumber of commits, trees, blobs or tags that have been found to be\nreachable during the walk.\n\n`git rev-list` is taught to use this function call when bitmaps are\nenabled instead of going through the old rev-list machinery. Do note,\nhowever, that counts with `left_right` and `cherry_mark` are not\noptimized by this patch.\n\nHere are some sample timings of different ways to count commits in\n`torvalds/linux`:\n\n\t$ time ../git/git rev-list master | wc -l\n\t376549\n\n\treal    0m6.973s\n\tuser    0m3.216s\n\tsys     0m5.316s\n\n\t$ time ../git/git rev-list --count master\n\t376549\n\n\treal    0m1.933s\n\tuser    0m1.744s\n\tsys     0m0.188s\n\n\t$ time ../git/git rev-list --use-bitmaps --count master\n\t376549\n\n\treal    0m0.005s\n\tuser    0m0.000s\n\tsys     0m0.004s\n\nNote that the time in the `--use-bitmaps` invocation is basically noise.\nIn my machine it ranges from 2ms to 6ms.\n---\n builtin/rev-list.c |   11 +++++++++--\n pack-bitmap.c      |   37 +++++++++++++++++++++++++++++++++++++\n pack-bitmap.h      |    2 ++\n 3 files changed, 48 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/rev-list.c b/builtin/rev-list.c\nindex 905ed08..097adb8 100644\n--- a/builtin/rev-list.c\n+++ b/builtin/rev-list.c\n@@ -354,8 +354,15 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \t\trevs.limited = 1;\n \n \tif (use_bitmaps && !prepare_bitmap_walk(&revs, NULL)) {\n-\t\ttraverse_bitmap_commit_list(&show_object_fast);\n-\t\treturn 0;\n+\t\tif (revs.count && !revs.left_right && !revs.cherry_mark) {\n+\t\t\tuint32_t commit_count;\n+\t\t\tcount_bitmap_commit_list(&commit_count, NULL, NULL, NULL);\n+\t\t\tprintf(\"%d\\n\", commit_count);\n+\t\t\treturn 0;\n+\t\t} else {\n+\t\t\ttraverse_bitmap_commit_list(&show_object_fast);\n+\t\t\treturn 0;\n+\t\t}\n \t}\n \n \tif (prepare_revision_walk(&revs))\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nindex 090db15..65fdce7 100644\n--- a/pack-bitmap.c\n+++ b/pack-bitmap.c\n@@ -720,6 +720,43 @@ void traverse_bitmap_commit_list(show_reachable_fn show_reachable)\n \tbitmap_git.result = NULL;\n }\n \n+static uint32_t count_object_type(\n+\tstruct bitmap *objects,\n+\tstruct ewah_bitmap *type_filter)\n+{\n+\tsize_t i = 0, count = 0;\n+\tstruct ewah_iterator it;\n+\teword_t filter;\n+\n+\tewah_iterator_init(&it, type_filter);\n+\n+\twhile (i < objects->word_alloc && ewah_iterator_next(&filter, &it)) {\n+\t\teword_t word = objects->words[i++] & filter;\n+\t\tcount += __builtin_popcountll(word);\n+\t}\n+\n+\treturn count;\n+}\n+\n+void count_bitmap_commit_list(\n+\tuint32_t *commits, uint32_t *trees, uint32_t *blobs, uint32_t *tags)\n+{\n+\tif (!bitmap_git.result)\n+\t\tdie(\"Tried to count bitmap without setting it up first\");\n+\n+\tif (commits)\n+\t\t*commits = count_object_type(bitmap_git.result, bitmap_git.commits);\n+\n+\tif (trees)\n+\t\t*trees = count_object_type(bitmap_git.result, bitmap_git.trees);\n+\n+\tif (blobs)\n+\t\t*blobs = count_object_type(bitmap_git.result, bitmap_git.blobs);\n+\n+\tif (tags)\n+\t\t*tags = count_object_type(bitmap_git.result, bitmap_git.tags);\n+}\n+\n struct bitmap_test_data {\n \tstruct bitmap *base;\n \tstruct progress *prg;\ndiff --git a/pack-bitmap.h b/pack-bitmap.h\nindex 8e7e3dc..816da6d 100644\n--- a/pack-bitmap.h\n+++ b/pack-bitmap.h\n@@ -47,6 +47,8 @@ typedef int (*show_reachable_fn)(\n \tstruct packed_git *found_pack,\n \toff_t found_offset);\n \n+void count_bitmap_commit_list(\n+\tuint32_t *commits, uint32_t *trees, uint32_t *blobs, uint32_t *tags);\n void traverse_bitmap_commit_list(show_reachable_fn show_reachable);\n int prepare_bitmap_walk(struct rev_info *revs, uint32_t *result_size);\n void test_bitmap_walk(struct rev_info *revs);\n-- \n1.7.9.5\n"},{"id":"221917","messageId":"7v7ghj571o.fsf@alter.siamese.dyndns.org","threadId":"34271","inReplyTo":"1372116193-32762-9-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 08/16] ewah: compressed bitmap implementation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-06-25T01:10:43Z","receivedAt":"2013-06-25T01:10:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Vicent Marti <tanoku@gmail.com> writes:\n\n> The library is re-licensed under the GPLv2 with the permission of Daniel\n> Lemire, the original author. The source code for the C version can\n> be found on GitHub:\n>\n> \thttps://github.com/vmg/libewok\n>\n> The original Java implementation can also be found on GitHub:\n>\n> \thttps://github.com/lemire/javaewah\n> ---\n\nPlease make sure that all patches are properly signed off.\n\n>  Makefile           |    6 +\n>  ewah/bitmap.c      |  229 +++++++++++++++++\n>  ewah/ewah_bitmap.c |  703 ++++++++++++++++++++++++++++++++++++++++++++++++++++\n>  ewah/ewah_io.c     |  199 +++++++++++++++\n>  ewah/ewah_rlw.c    |  124 +++++++++\n>  ewah/ewok.h        |  194 +++++++++++++++\n>  ewah/ewok_rlw.h    |  114 +++++++++\n\nThis is lovely.  A few comments after an initial quick scan-through.\n\n - The code and the headers are well commented, which is good.\n\n - What's __builtin_popcountll() doing there in a presumably generic\n   codepath?\n\n - Two variants of \"bitmap\" are given different and easy to\n   understand type names (vanilla one is \"bitmap\", the clever one is\n   \"ewah_bitmap\"), but at many places, a pointer to ewah_bitmap is\n   simply called \"bitmap\" or \"bitmap_i\" without \"ewah\" anywhere,\n   which waas confusing to read.  Especially, the \"NAND\" operation\n   for bitmap takes two bitmaps, while \"OR\" takes one bitmap and\n   ewah_bitmap.  That is fine as long as the combination is\n   convenient for callers, but I wished the ewah variables be called\n   with \"ewah\" somewhere in their names.\n\n - I compile with \"-Werror -Wdeclaration-after-statement\"; some\n   places seem to trigger it.\n\n - Some \"extern\" declarations in *.c sources were irritating;\n   shouldn't they be declared in *.h file and included?\n\n - There are some instances of \"if (condition) stmt;\" on a single\n   line; looked irritating.   \n\n - \"bool\" is not a C type we use (and not a particularly good type\n   in C++, either).\n\nThat is it for now. I am looking forward to read through the users\nof the library ;-)\n\nThanks for working on this.\n"},{"id":"221923","messageId":"CAJo=hJtcQwh-N-9_i84y1ZsL0mdREHcxhP2gepcrREiaxvxS6A@mail.gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-10-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-06-25T05:42:49Z","receivedAt":"2013-06-25T05:42:49Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Mon, Jun 24, 2013 at 5:23 PM, Vicent Marti <tanoku@gmail.com> wrote:\n> This is the technical documentation and design rationale for the new\n> Bitmap v2 on-disk format.\n> ---\n>  Documentation/technical/bitmap-format.txt |  235 +++++++++++++++++++++++++++++\n>  1 file changed, 235 insertions(+)\n>  create mode 100644 Documentation/technical/bitmap-format.txt\n>\n> diff --git a/Documentation/technical/bitmap-format.txt b/Documentation/technical/bitmap-format.txt\n> new file mode 100644\n> index 0000000..5400082\n> --- /dev/null\n> +++ b/Documentation/technical/bitmap-format.txt\n> @@ -0,0 +1,235 @@\n> +GIT bitmap v2 format & rationale\n> +================================\n> +\n> +       - A header appears at the beginning, using the same format\n> +       as JGit's original bitmap indexes.\n> +\n> +               4-byte signature: {'B', 'I', 'T', 'M'}\n> +\n> +               2-byte version number (network byte order)\n> +                       The current implementation only supports version 2\n> +                       of the bitmap index. The rationale for this is explained\n> +                       in this document.\n> +\n> +               2-byte flags (network byte order)\n> +\n> +                       The folowing flags are supported:\n> +\n> +                       - BITMAP_OPT_FULL_DAG (0x1) REQUIRED\n> +                       This flag must always be present. It implies that the bitmap\n> +                       index has been generated for a packfile with full closure\n> +                       (i.e. where every single object in the packfile can find\n> +                        its parent links inside the same packfile). This is a\n> +                       requirement for the bitmap index format, also present in JGit,\n> +                       that greatly reduces the complexity of the implementation.\n> +\n> +                       - BITMAP_OPT_LE_BITMAPS (0x2)\n> +                       If present, this implies that that the EWAH bitmaps in this\n> +                       index has been serialized to disk in little-endian byte order.\n> +                       Note that this only applies to the actual bitmaps, not to the\n> +                       Git data structures in the index, which are always in Network\n> +                       Byte order as it's costumary.\n> +\n> +                       - BITMAP_OPT_BE_BITMAPS (0x4)\n> +                       If present, this implies that the EWAH bitmaps have been serialized\n> +                       using big-endian byte order (NWO). If the flag is missing, **the\n> +                       default is to assume that the bitmaps are in big-endian**.\n\nI very much hate seeing a file format that is supposed to be portable\nthat supports both big-endian and little-endian encoding. Such a\nspecification forces everyone to implement two code paths to handle\nreading data from the file, on the off-chance they are on the wrong\nplatform. Or it forces one platform to be unable to use the file. In\nwhich case a repository may want to build two files, one for each\nplatform, but you blocked that off by allowing only one file.\n\nWhat is wrong with picking one encoding and sticking to it? The .idx\nfile and the dircache use big-endian format. Why not just use\nbig-endian here too and convert words on demand as they are accessed\nfrom the mmap region? That is what the .idx format does when accessing\nan offset.\n\n> +                       - BITMAP_OPT_HASH_CACHE (0x8)\n> +                       If present, a hash cache for finding delta bases will be available\n> +                       right after the header block in this index. See the following\n> +                       section for details.\n> +\n> +               4-byte entry count (network byte order)\n> +\n> +                       The total count of entries (bitmapped commits) in this bitmap index.\n> +\n> +               20-byte checksum\n> +\n> +                       The SHA1 checksum of the pack this bitmap index belongs to.\n> +\n> +       - An OPTIONAL delta cache follows the header.\n\nSome may find the name \"delta cache\" confusing as it does not cache\ndeltas of objects. May I suggest \"path hash cache\" as an alternative\nname?\n\n> +               The cache is formed by `n` 4-byte hashes in a row, where `n` is\n> +               the amount of objects in the indexed packfile. Note that this amount\n> +               is the **total number of objects** and is not related to the\n> +               number of commits that have been selected and indexed in the\n> +               bitmap index.\n> +\n> +               The hashes are stored in Network Byte Order and they are the same\n> +               values generated by a normal revision walk during the `pack-objects`\n> +               phase.\n\nI find it interesting this is network byte order and not big-endian or\nlittle-endian based on the flag in the header.\n\n> +               The `n`nth hash in the cache is the name hash for the `n`th object\n> +               in the index for the indexed packfile.\n> +\n> +               [RATIONALE]:\n> +\n> +               The bitmap index allows us to skip the Counting Objects phase\n> +               during `pack-objects` and yield all the OIDs that would be reachable\n> +               (\"WANTS\") when generating the pack.\n> +\n> +               This optimization, however, means that we're adding objects to the\n> +               packfile straight from the packfile index, and hence we are lacking\n> +               path information for the objects that would normally be generated\n> +               during the \"Counting Objects\" phase.\n> +\n> +               This path information for each object is hashed and used as a very\n> +               effective way to find good delta bases when compressing the packfile;\n> +               without these hashes, the resulting packfiles are much less optimal.\n> +\n> +               By storing all the hashes in a cache together with the bitmapsin\n> +               the bitmap index, we can yield not only the SHA1 of all the reachable\n> +               objects, but also their hashes, and allow Git to be much smarter when\n> +               finding delta bases for packing.\n> +\n> +               If the delta cache is not available, the bitmap index will obviously\n> +               be smaller in disk, but the packfiles generated using this index will\n> +               be between 20% and 30% bigger, because of the lack of name/path\n> +               information when finding delta bases.\n\nJGit does not encode this because we were afraid of freezing the hash\nfunction into the file format. Indeed we are not certain JGit even\nuses the same path hash function as C Git does, because C Git's\nimplementation is covered by the GPL and JGit prefers to license its\nwork under BSD.\n\nIf the path hash is going to become part of the format, the algorithm\nfor computing the hash should also be specified in the format so that\nnon-GPL implementations have an opportunity to be compatible.\n\nOne way we side-stepped the size inflation problem in JGit was to only\nuse the bitmap index information when sending data on the wire to a\nclient. Here delta reuse plays a significant factor in building the\npack, and we don't have to be as accurate on matching deltas. During\nthe equivalent of `git repack` bitmaps are not used, allowing the\ntraditional graph enumeration algorithm to generate path hash\ninformation.\n\n> +       - 4 EWAH bitmaps that act as type indexes\n> +\n> +               Type indexes are serialized after the hash cache in the shape\n> +               of four EWAH bitmaps stored consecutively (see Appendix A for\n> +               the serialization format of an EWAH bitmap).\n> +\n> +               There is a bitmap for each Git object type, stored in the following\n> +               order:\n> +\n> +                       - Commits\n> +                       - Trees\n> +                       - Blobs\n> +                       - Tags\n> +\n> +               In each bitmap, the `n`th bit is set to true if the `n`th object\n> +               in the packfile index is of that type.\n> +\n> +               The obvious consequence is that the XOR of all 4 bitmaps will result\n> +               in a full set (all bits sets), and the AND of all 4 bitmaps will\n> +               result in an empty bitmap (no bits set).\n\nInstead of XOR did you mean OR here?\n\n> +       - N EWAH bitmaps, one for each indexed commit\n> +\n> +               Where `N` is the total amount of entries in this bitmap index.\n> +               See Appendix A for the serialization format of an EWAH bitmap.\n> +\n> +       - An entry index with `N` entries for the indexed commits\n> +\n> +               Index entries are stored consecutively, and each entry has the\n> +               following format:\n> +\n> +               - 20-byte SHA1\n> +                       The SHA1 of the commit that this bitmap indexes\n> +\n> +               - 4-byte offset (Network Byte Order)\n> +                       The offset **from the beginning of the file** where the\n> +                       bitmap for this commit is stored.\n\nEh, another network byte order field in a file that also has selective\nordering. *sigh*\n\n> +               - 1-byte XOR-offset\n> +                       The xor offset used to compress this bitmap. For an entry\n> +                       in position `x`, a XOR offset of `y` means that the actual\n> +                       bitmap representing for this commit is composed by XORing the\n> +                       bitmap for this entry with the bitmap in entry `x-y` (i.e.\n> +                       the bitmap `y` entries before this one).\n> +\n> +                       Note that this compression can be recursive. In order to\n> +                       XOR this entry with a previous one, the previous entry needs\n> +                       to be decompressed first, and so on.\n> +\n> +                       The hard-limit for this offset is 160 (an entry can only be\n> +                       xor'ed against one of the 160 entries preceding it). This\n> +                       number is always positivea, and hence entries are always xor'ed\n> +                       with **previous** bitmaps, not bitmaps that will come afterwards\n> +                       in the index.\n\nWhat order are these entries in? Sorted by SHA-1 or random?\n\nColby found that doing an XOR against the descendant commit yielded\nvery small bitmaps, so JGit tries to XOR-compress bitmaps along common\nlinear slices of history. This is trivial in Linus' kernel tree where\nthere is effectively only one history, but its more relevant with\nlong-running side branches that have release tags that may not have\nfully merged into \"master\".\n\n> +               - 1-byte flags for this bitmap\n> +                       At the moment the only available flag is `0x1`, which hints\n> +                       that this bitmap can be re-used when rebuilding bitmap indexes\n> +                       for the repository.\n> +\n> +               - 2 bytes of RESERVED data (used right now for better packing).\n> +\n> +== Rationale for changes from the Bitmap Format v1\n> +\n> +- Serialized EWAH bitmaps can be stored in Little-Endian byte order,\n> +  if defined by the BITMAP_OPT_LE_BITMAPS flag in the header.\n> +\n> +  The original JGit implementation stored bitmaps in Big-Endian byte\n> +  order (NWO) because it was unable to `mmap` the serialized format,\n> +  and hence always required a full parse of the bitmap index to memory,\n> +  where the BE->LE conversion could be performed.\n\nYou can mmap a file and convert each word on access if your machine is\nnot using the same byte order. It is not necessary to convert the\nentire file before using it through a mmap region.\n\n> +  This full parse, however, requires prohibitive loading times in LE\n> +  machines (i.e. all modern server hardware): a repository like\n> +  `torvalds/linux` can have about 8mb of bitmap indexes, resulting\n> +  in roughly 400ms of parse time.\n\nThis makes me wonder what the JGit parse time is. It is ugly if we are\nspending 400ms to load the bitmap index for the kernel repository.\n\n> +  This is not an issue in JGit, which is capable of serving repositories\n> +  from a single-process daemon running on the JVM, but `git-daemon` in\n> +  git has been implemented with a process-based design (a new\n> +  `pack-objects` is spawned for each request), and the boot times\n> +  of parsing the bitmap index every time `pack-objects` is spawned can\n> +  seriously slow down requests (particularly for small fetches, where we'd\n> +  spend about 1.5s booting up and 300ms performing the Counting Objects\n> +  phase).\n\nThere are other strategies that Git could use to handle request\nprocessing at scale. But I guess its reasonable to assume these aren't\nviable for Git for a number of reasons. E.g. \"long tail\" access effect\nthat many servers have, where most requests are to a large number of\nrepositories that themselves receive very few requests, an environment\nthat does not lend itself to caching.\n\n> +  By storing the bitmaps in Little-Endian, we're able to `mmap` their\n> +  compressed data straight in memory without parsing it beforehand, and\n> +  since most queries don't require accessing all the serialized bitmaps,\n> +  we'll only page in the minimal amount of bitmaps necessary to perform\n> +  the reachability analysis as they are accessed.\n\nFWIW the .idx and .pack file formats `mmap` the compressed data\nstraight into memory without parsing it beforehand, and do not use\nlittle-endian byte order. It is possible to have a single compressed\nfile format definition that is portable to all architectures, and is\naccessed by mmap, at scale, with reasonable efficiency.\n\n> +- An index of all the bitmapped commits is written at the end of the packfile,\n> +  instead of interpersed with the serialized bitmaps in the middle of the\n> +  file.\n\nThis is probably a mistake in the JGit design. Your approach is\nslightly more complex, but in general I agree with having a table of\nthe SHA-1s isolated from the bitmaps themselves so that a reader can\naccess specific bitmaps at random without needing to wade through all\ncompressed bitmaps.\n\nI would have proposed putting the table at the start of the file, not\nthe end. The writer making the file can completely serialize the\nbitmaps into memory before writing them to disk, and thus knows the\nfull layout of the resulting file. If the bitmaps don't fit in RAM at\nwriting time, game over, the optimization of having a very compact\nrepresentation of the graph is no longer helping you.\n\n> +  Again, the old design implied a full parse of the whole bitmap index\n> +  (which JGit can afford because its daemon is single-process), but it made\n> +  impossible `mmaping` the bitmap index file and accessing only the parts\n> +  required to actually solve the query.\n> +\n> +  With an index at the end of the file, we can load only this index in memory,\n> +  allowing for very efficient access to all the available bitmaps lazily (we\n> +  have their offsets in the mmaped file).\n> +\n> +- The ordering of the objects in each bitmap has changed from\n> +  packfile-order (the nth bit in the bitmap is the nth object in the\n> +  packfile) to index-order (the nth bit in the bitmap is the nth object\n> +  in the INDEX of the packfile).\n\nDid you notice an increase in bitmap size when you did this? Colby\ntested both orderings and we observed the bitmaps were quantifiably\nsmaller when using the pack file ordering, due to the pack file\nlocality rules and the EWAH compression. Using the pack file ordering\nwas a very conscious design decision.\n\n> +  There is not a noticeable performance difference when actually converting\n> +  from bitmap position to SHA1 and from SHA1 to bitmap position, but when\n> +  using packfile ordering like JGit does, queries need to go through the\n> +  reverse index (pack-revindex.c).\n> +\n> +  Generating this reverse index at runtime is **not** free (around 900ms\n> +  generation time for a repository like `torvalds/linux`), and once again,\n> +  this generation time needs to happen every time `pack-objects` is\n> +  spawned.\n\nDid you know the packer needs the reverse index in order to compute\nthe end offset of an object it will copy as-is during delta reuse? How\nhave you avoided making the reverse index?\n\nAgain this is why we chose to pin the JGit bitmap on the reverse index\nbeing present. It already had to be present to support as-is reuse.\nOnce we knew we had to have that reverse index it was OK to rely on it\nto get better compression on the bitmaps, and thus make them take up\nless memory when loaded into a server. Even if you mmap a file you\nwant it to be small so it is more likely to retain in the kernel\nbuffer cache across process invocations.\n"},{"id":"221937","messageId":"CALkWK0kG8VJh9TC8Yh82fg8zCHzevjP0_yghcH7woEML3-MTog@mail.gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-11-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 10/16] pack-objects: use bitmaps when packing objects","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-06-25T12:48:41Z","receivedAt":"2013-06-25T12:48:41Z","isPatch":true,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Vicent Marti wrote:\n>         $ time ../git/git pack-objects --all --stdout\n>         Counting objects: 3053537, done.\n>         Compressing objects: 100% (495706/495706), done.\n>         Total 3053537 (delta 2529614), reused 3053537 (delta 2529614)\n>\n>         real    0m36.686s\n>         user    0m34.440s\n>         sys     0m2.184s\n>\n>         $ time ../git/git pack-objects --all --stdout\n>         Counting objects: 3053537, done.\n>         Compressing objects: 100% (495706/495706), done.\n>         Total 3053537 (delta 2529614), reused 3053537 (delta 2529614)\n>\n>         real    0m7.255s\n>         user    0m6.892s\n>         sys     0m0.444s\n\nAwesome work!  Can you put up this series on gh:vmg so I can try it\nout for myself?\n\n> diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> index b7cab18..469b8da 100644\n> --- a/builtin/pack-objects.c\n> +++ b/builtin/pack-objects.c\n> +       if (!strcmp(k, \"pack.usebitmaps\")) {\n> +               bitmap_support = git_config_bool(k, v);\n> +               return 0;\n> +       }\n\nNot using config_error_nonbool() to indicate an error?\n\n> +       if (use_bitmap_index) {\n> +               uint32_t size_hint;\n> +\n> +               if (!prepare_bitmap_walk(&revs, &size_hint)) {\n> +                       khint_t new_hash_size = (size_hint * (1.0 / __ac_HASH_UPPER)) + 0.5;\n\nHow does this work?  You've taken the inverse of __ac_HASH_UPPER,\nmultiplied it by the size_hint you get from prepare_bitmap_walk(), and\nadd 0.5?\n\n> +                       kh_resize_sha1(packed_objects, new_hash_size);\n\nSo packed_objects is a hashtable of type kh_sha1_t * (you introduced\nin [03/16]) that you're now resizing to new_hash_size.  To find out\nwhat the significance of this new_hash_size is, it looks like I have\nto read prepare_bitmap_walk().\n\n> +                       nr_alloc = (size_hint + 63) & ~63;\n> +                       objects = xrealloc(objects, nr_alloc * sizeof(struct object_entry *));\n\nInteresting.  The only other place where we realloc the objects in\nthis file is in pack-objects.c:949, and we do that because nr_object\n>= nr_alloc.  What is this 63 magic?\n\n>         if (prepare_revision_walk(&revs))\n>                 die(\"revision walk setup failed\");\n> +\n\nStray newline.\n\n> +       if (bitmap_support) {\n> +               if (use_internal_rev_list && pack_to_stdout)\n> +                       use_bitmap_index = 1;\n> +       }\n> +\n\nWait, what does pack_to_stdout have to do with deciding whether or not\nto walk the bitmap?\n\n> diff --git a/pack-bitmap.c b/pack-bitmap.c\n> new file mode 100644\n> index 0000000..090db15\n> --- /dev/null\n> +++ b/pack-bitmap.c\n> +struct stored_bitmap {\n> +       unsigned char sha1[20];\n> +       struct ewah_bitmap *root;\n> +       struct stored_bitmap *xor;\n> +       int flags;\n> +};\n\nWhat exactly is this?  What is stored_bitmap *xor?  It looks like some\nsort of next-pointer, but why is it named xor?\n\n> +struct bitmap_index {\n\nOkay, the bitmap index.\n\n> +       struct ewah_bitmap *commits;\n> +       struct ewah_bitmap *trees;\n> +       struct ewah_bitmap *blobs;\n> +       struct ewah_bitmap *tags;\n\nI might be asking a really stupid question here, but why do you have\ndifferent bitmaps for different object types?  Unless I'm mistaken,\nthe packfile index doesn't make this differentiation: it sorts and\nstores the SHA-1s of the various objects; you request a SHA-1, it does\na binary search and returns the object.\n\n> +       khash_sha1 *bitmaps;\n\nA hashmap keyed with the SHA-1, I presume.\n\n> +       struct packed_git *pack;\n\nYou're defining which pack this bitmap index is for, right?\n\n> +       struct {\n> +               struct object_array entries;\n> +               khash_sha1 *map;\n> +       } fake_index;\n\nWhat is this?\n\n> +       struct bitmap *result;\n\nNo clue what this is about.\n\n> +       int entry_count;\n\nNo clue what this is, but I'm assuming it can't be important because\nit's an int.\n\n> +       char pack_checksum[20];\n> +\n> +       int version;\n\nUse something invariant like uint32_t?  Also, there is no clear\nindication about where this information is going to go (header,\npresumably?).  Look at pack.h:\n\nstruct pack_idx_header {\n\tuint32_t idx_signature;\n\tuint32_t idx_version;\n};\n\n> +       unsigned loaded : 1,\n> +                        native_bitmaps : 1,\n> +                        has_hash_cache : 1;\n\nBooleans, but I don't know what they're doing here even after reading\nyour bitmap-format.txt.\n\n> +       struct ewah_bitmap *(*read_bitmap)(struct bitmap_index *index);\n\nI'm very confused now.  Each bitmap_index has a specialized read_bitmap()?\n\n> +       void *map;\n> +       size_t map_size, map_pos;\n> +\n> +       uint32_t *delta_hashes;\n\nI'll give up on the rest.\n\n> +static struct bitmap_index bitmap_git;\n\nYou could have made the struct static to begin with and ended it with\na } bitmap_git;\n\n> +static struct ewah_bitmap *\n> +lookup_stored_bitmap(struct stored_bitmap *st)\n\nPlease conform to Linux style and make it easier for us to grep by\nputting this in one line?\n\n> +{\n> +       struct ewah_bitmap *parent;\n> +       struct ewah_bitmap *composed;\n> +\n> +       if (st->xor == NULL)\n\nif (!st->xor)\n\n> +               return st->root;\n\nOkay, st->xor needs to be set to something for lookup_stored_bitmap()\nto do something useful.\n\n> +       composed = ewah_pool_new();\n> +       parent = lookup_stored_bitmap(st->xor);\n\nSo st->xor is a parent-pointer?  Still doesn't answer my question\nabout why it is named xor.\n\n> +       ewah_xor(st->root, parent, composed);\n\nI would have loved it if the prototype of this function made it clear\nwhat it was writing like: ewah_xor(st->root, parent, &composed); but\nthen it expects the caller to do memory allocation, so never mind.\n\n> +       ewah_pool_free(st->root);\n> +       st->root = composed;\n> +       st->xor = NULL;\n> +\n> +       return composed;\n\nSo lookup_stored_bitmap() just xors st->root with parent (determined\nby recursively looking up st->xor)?\n\n> +}\n> +\n> +static struct ewah_bitmap *\n> +_read_bitmap(struct bitmap_index *index)\n\nWe usually use the _1 suffix convention for internal functions, not _ prefix.\n\n> +       int bitmap_size;\n> +\n> +       bitmap_size = ewah_read_mmap(b,\n> +               index->map + index->map_pos,\n> +               index->map_size - index->map_pos);\n\nMake this a single statement.  Also, why can't I see this on\ngh:vmg/libework?  Your [08/16] has diverged from there :/\n\n> +       return b;\n\nSo _read_bitmap() mmaps the bitmap and returns it.\n\n> +static struct ewah_bitmap *\n> +_read_bitmap_native(struct bitmap_index *index)\n\nThe counterpart that calls *_mmap_native() in ewah.  Have to look at\nthe difference.  In any case, I hope you've used xmmap().\n\n> +static int load_bitmap_header(struct bitmap_index *index)\n> +{\n\nI'm going to compare this to sha1_file.c:check_packed_git_idx().\n\n> +       struct bitmap_disk_header *header = (void *)index->map;\n> +\n> +       if (index->map_size < sizeof(*header))\n> +               return error(\"Corrupted bitmap index (missing header data)\");\n\nNo munmap()?  Was abstracting out the mmap detail a good idea?\n\n> +       if (memcmp(header->magic, BITMAP_MAGIC_PREFIX, sizeof(BITMAP_MAGIC_PREFIX)) != 0)\n> +               return error(\"Corrupted bitmap index file (wrong header)\");\n\nPACK_SIGNATURE, PACK_IDX_SIGNATURE.  Name this BITMAP_IDX_SIGNATURE?\n\n> +       index->version = (int)ntohs(header->version);\n\nYou wouldn't have to coerce to int if version were a uint32_t in the\nfirst place.\n\n> +       /* Parse known bitmap format options */\n> +       {\n> +               uint32_t flags = ntohs(header->options);\n\nOkay.\n\n> +               if ((flags & BITMAP_OPT_FULL_DAG) == 0) {\n> +                       return error(\"Unsupported options for bitmap index file \"\n> +                               \"(Git requires BITMAP_OPT_FULL_DAG)\");\n> +               }\n\nUnnecessary braces for single statement.\n\n> +               if (flags & BITMAP_OPT_HASH_CACHE)\n> +                       index->has_hash_cache = 1;\n> +\n> +               index->read_bitmap = &_read_bitmap;\n\nSo you've set the read_bitmap() function to _read_bitmap().  Let's see why.\n\n> +               /*\n> +                * If we are in a little endian machine and the bitmap\n> +                * was written in LE, we can mmap it straight into memory\n> +                * without having to parse it\n> +                */\n> +               if ((flags & BITMAP_OPT_LE_BITMAPS)) {\n> +#if __BYTE_ORDER == __LITTLE_ENDIAN\n> +                       index->native_bitmaps = 1;\n> +                       index->read_bitmap = &_read_bitmap_native;\n> +#else\n> +                       die(\"The existing bitmap index is written in little-endian \"\n> +                               \"byte order and cannot be read in this machine.\\n\"\n> +                               \"Please re-build the bitmap indexes locally.\");\n> +#endif\n> +               }\n> +       }\n\nOkay.\n\n> +       index->entry_count = ntohl(header->entry_count);\n> +       memcpy(index->pack_checksum, header->checksum, sizeof(header->checksum));\n\nI might be asking (yet another) really stupid question here, but why\nisn't the checksum an unsigned char[20] (i.e. a simple 20-byte SHA-1)?\n We already have infrastructure to deal with SHA-1s, so might as well\nreuse it, right?\n\n> +       index->map_pos += sizeof(*header);\n\nYou've read the header successfully, and incremented map_pos for other callers.\n\n> +static struct stored_bitmap *\n> +store_bitmap(struct bitmap_index *index,\n> +       const unsigned char *sha1,\n> +       struct ewah_bitmap *bitmap,\n> +       struct stored_bitmap *xor_with, int flags)\n\nWhy don't you just prepare a struct and send it to this function to\nwrite instead of so many arguments?\n\n> +       stored = xmalloc(sizeof(struct stored_bitmap));\n> +       stored->root = bitmap;\n> +       stored->xor = xor_with;\n> +       stored->flags = flags;\n> +       memcpy(stored->sha1, sha1, 20);\n\nUse hashcpy().  You would have had to do none of this if the caller\nhad passed a readymade struct, no?\n\n> +       hash_pos = kh_put_sha1(index->bitmaps, stored->sha1, &ret);\n\nOkay, you store a SHA-1.\n\n> +       if (ret == 0) {\n> +               error(\"Duplicate entry in bitmap index: %s\", sha1_to_hex(sha1));\n> +               return NULL;\n> +       }\n\n0 is success by convention!\n\n> +       kh_value(index->bitmaps, hash_pos) = stored;\n\nOkay.\n\n> +       return stored;\n\nYou're returning allocated memory, that the caller must remember to free.\n\n> +static int\n> +load_bitmap_entries_v2(struct bitmap_index *index)\n\nI'm not sure it's a great idea to put a volatile version number in the\nfunction name.\n\n> +{\n> +       static const int MAX_XOR_OFFSET = 16;\n> +\n> +       int i;\n> +       struct stored_bitmap *recent_bitmaps[16];\n\nDoes this 16 have anything to do with MAX_XOR_OFFSET?\n\n> +       struct bitmap_disk_entry_v2 *entry;\n> +\n> +       void *index_pos = index->map + index->map_size -\n> +               (index->entry_count * sizeof(struct bitmap_disk_entry_v2));\n\nWait, why did we set map_pos earlier if you're recomputing it here?\nAnd why is this a void *?\n\n> +       for (i = 0; i < index->entry_count; ++i) {\n> +               int xor_offset, flags, ret;\n> +               struct stored_bitmap *xor_bitmap = NULL;\n> +               struct ewah_bitmap *bitmap = NULL;\n> +               uint32_t bitmap_pos;\n> +\n> +               entry = index_pos;\n> +               index_pos += sizeof(struct bitmap_disk_entry_v2);\n\nOkay, so I understand that you're parsing one bitmap_disk_entry_v2\nstruct at a time.\n\n> +               bitmap_pos = ntohl(entry->bitmap_pos);\n> +               xor_offset = (int)entry->xor_offset;\n> +               flags = (int)entry->flags;\n\nI have no clue why you're casting like this.\n\n> +               if (index->native_bitmaps) {\n\nWhat is this native versus non-native bitmaps?  Your\nbitmap-formats.txt has nothing to say on the matter.\n\n> +                       bitmap = calloc(1, sizeof(struct ewah_bitmap));\n> +                       ret = ewah_read_mmap_native(bitmap,\n> +                               index->map + bitmap_pos,\n> +                               index->map_size - bitmap_pos);\n\nWait a minute.  Isn't this what you wrapped in _read_bitmap_native()?\nTotally confused.\n\n> +               } else {\n> +                       bitmap = ewah_pool_new();\n> +                       ret = ewah_read_mmap(bitmap,\n> +                               index->map + bitmap_pos,\n> +                               index->map_size - bitmap_pos);\n\nDid you forget about _read_bitmap()?\n\n> +               if (ret < 0 || xor_offset > MAX_XOR_OFFSET || xor_offset > i) {\n> +                       return error(\"Corrupted bitmap pack index\");\n> +               }\n\nUnnecessary braces.\n\n> +               if (xor_offset > 0) {\n> +                       xor_bitmap = recent_bitmaps[(i - xor_offset) % MAX_XOR_OFFSET];\n> +\n> +                       if (xor_bitmap == NULL)\n\nif (!xor_bitmap)\n\n> +                               return error(\"Invalid XOR offset in bitmap pack index\");\n\nI haven't seen a single die() until now, and that's a Good sign.\n\n> +               recent_bitmaps[i % MAX_XOR_OFFSET] = store_bitmap(\n> +                       index, entry->sha1, bitmap, xor_bitmap, flags);\n\nSo you fill in the 16 recent bitmaps in this function?\n\n> +static int load_bitmap_index(\n> +       struct bitmap_index *index,\n> +       const char *path,\n> +       struct packed_git *packfile)\n> +{\n> +       int fd = git_open_noatime(path);\n\nI assume you exposed this static defined in sha1_file.c in an earlier\npatch, but I didn't check.\n\n> +       struct stat st;\n> +\n> +       if (fd < 0) {\n> +               return -1;\n> +       }\n\nUnnecessary braces.\n\n> +       index->map_size = xsize_t(st.st_size);\n> +       index->map = xmmap(NULL, index->map_size, PROT_READ, MAP_PRIVATE, fd, 0);\n> +       close(fd);\n\nI like how similar this is to check_packed_git_idx().  What happened\nto your ewah mapping abstractions though?\n\n> +       index->bitmaps = kh_init_sha1();\n> +       index->pack = packfile;\n> +       index->fake_index.map = kh_init_sha1();\n\nI'll hopefully get to find out what fake_index is here.\n\n> +       if (load_bitmap_header(index) < 0)\n> +               return -1;\n\nOkay.  Notice how the format we're parsing is documented tersely as\ninline comments in sha1_name.c: you might like to do that too.\n\n> +       if (index->has_hash_cache) {\n> +               index->delta_hashes = index->map + index->map_pos;\n> +               index->map_pos += (packfile->num_objects * sizeof(uint32_t));\n> +       }\n\nOkay.\n\n> +       if ((index->commits = index->read_bitmap(index)) == NULL ||\n> +               (index->trees = index->read_bitmap(index)) == NULL ||\n> +               (index->blobs = index->read_bitmap(index)) == NULL ||\n> +               (index->tags = index->read_bitmap(index)) == NULL)\n> +               return -1;\n\nAs usual, please use !() instead of explicitly comparing with NULL.\nIt looks like I'll get to find out why you have four different bitmaps\nset to the same thing (?) soon; exciting!\n\n> +       if (load_bitmap_entries_v2(index) < 0)\n> +               return -1;\n> +\n> +       index->loaded = true;\n\nFine.  This function calls out to various little parsing functions and\nsets index->loaded.  It returns -1 instead of error(), because those\nlittle functions report the errors.\n\n> +char *pack_bitmap_filename(struct packed_git *p)\n\nCompare with sha1_name.c:open_pack_index().\n\n> +int open_pack_bitmap(struct packed_git *p)\n\nOkay.\n\n> +void prepare_bitmap_git(void)\n\nOkay.\n\n> +struct include_data {\n> +       struct bitmap *base;\n> +       struct bitmap *seen;\n> +};\n\nI wonder what this is.\n\n> +static inline int bitmap_position_extended(const unsigned char *sha1)\n\nSorry, I'm stopping here.  It's impossible to review this gigantic\npatch in one sitting.\n\nThanks.\n"},{"id":"221938","messageId":"alpine.DEB.2.00.1306251404510.9929@ds9.cixit.se","threadId":"34271","inReplyTo":"1372116193-32762-8-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 07/16] compat: add endinanness helpers","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2013-06-25T13:08:39Z","receivedAt":"2013-06-25T13:08:39Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Vicent Marti:\n\n> The POSIX standard doesn't currently define a `nothll`/`htonll` \n> function pair to perform network-to-host and host-to-network swaps \n> of 64-bit data. These 64-bit swaps are necessary for the on-disk \n> storage of EWAH bitmaps if they are not in native byte order.\n\nendian(3) claims that glibc 2.9+ define be64toh() and htobe64() which \nshould do what you are looking for. The manual page does mention them \nbeing named differently across OSes, though, so you may need to be \ncareful with that.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"221939","messageId":"CAFFjANSNagvDgvrFNV1OLg=-4BPyQVjMDnfMPihdhVJR7o0TdQ@mail.gmail.com","threadId":"34271","inReplyTo":"alpine.DEB.2.00.1306251404510.9929@ds9.cixit.se","subject":"Re: [PATCH 07/16] compat: add endinanness helpers","fromName":"Vicent Martí","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-25T13:25:14Z","receivedAt":"2013-06-25T13:25:14Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"On Tue, Jun 25, 2013 at 3:08 PM, Peter Krefting <peter@softwolves.pp.se> wrote:\n> endian(3) claims that glibc 2.9+ define be64toh() and htobe64() which should\n> do what you are looking for. The manual page does mention them being named\n> differently across OSes, though, so you may need to be careful with that.\n\nI'm aware of that, but Git needs to build with glibc 2.7+ (or was it\n2.6?), hence the need for this compat layer.\n"},{"id":"221973","messageId":"87fvw5nae9.fsf@linux-k42r.v.cablecom.net","threadId":"34271","inReplyTo":"1372116193-32762-3-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 02/16] sha1_file: refactor into `find_pack_object_pos`","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-06-25T13:59:28Z","receivedAt":"2013-06-25T13:59:28Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Vicent Marti <tanoku@gmail.com> writes:\n\n>  \tif (use_lookup) {\n> -\t\tint pos = sha1_entry_pos(index, stride, 0,\n> -\t\t\t\t\t lo, hi, p->num_objects, sha1);\n> -\t\tif (pos < 0)\n> -\t\t\treturn 0;\n> -\t\treturn nth_packed_object_offset(p, pos);\n> +\t\treturn sha1_entry_pos(index, stride, 0, lo, hi, p->num_objects, sha1);\n>  \t}\n\nOur house style prefers not having the braces in a single-line conditional.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"221974","messageId":"87a9mdnae3.fsf@linux-k42r.v.cablecom.net","threadId":"34271","inReplyTo":"1372116193-32762-4-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 03/16] pack-objects: use a faster hash table","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-06-25T14:03:22Z","receivedAt":"2013-06-25T14:03:22Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Vicent Marti <tanoku@gmail.com> writes:\n\n> This commit brings `khash.h`, a header only hash table implementation\n> that while remaining rather simple (uses quadratic probing and a\n> standard hashing scheme) and self-contained, offers a significant\n> performance improvement in both insertion and lookup times.\n>\n> `khash` is a generic hash table implementation that can be 'templated'\n> for any given type while maintaining good performance by using preprocessor\n> macros. This specific version has been modified to define by default a\n> `khash_sha1` type, a map of SHA1s (const unsigned char[20]) to void *\n> pointers.\n>\n> When replacing the old hash table implementation in `pack-objects` with\n> the khash_sha1 table, the insertion time is greatly reduced:\n>\n> \tkh_put_sha1 :: 284.011ms\n> \tadd_object_entry_1 : 36.06ms\n> \thashcmp :: 24.045ms\n>\n> This reduction of more than 50% in the insertion and lookup times,\n> although nice, is not particularly noticeable for normal `pack-objects`\n> operation: `pack-objects` performs massive batch insertions and\n> relatively few lookups, so `khash` doesn't get a chance to shine here.\n>\n> The big win here, however, is in the massively reduced amount of hash\n> collisions (as you can see from the huge reduction of time spent in\n> `hashcmp` after the change). These greatly improved lookup times\n> will result critical once we implement the writing algorithm for bitmap\n> indxes in a later patch of this series.\n\nIs that reduction in collisions purely because it uses quadratic\nprobing, or is there some other magic trick involved?  Is the same also\napplicable to the other users of the \"big\" object hash table?  (I assume\nPeff has already tried applying it there, but I'm still curious...)\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"221975","messageId":"874nclnadu.fsf@linux-k42r.v.cablecom.net","threadId":"34271","inReplyTo":"1372116193-32762-9-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 08/16] ewah: compressed bitmap implementation","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-06-25T15:38:25Z","receivedAt":"2013-06-25T15:38:25Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Vicent Marti <tanoku@gmail.com> writes:\n\n> The library is re-licensed under the GPLv2 with the permission of Daniel\n> Lemire, the original author.\n\nThis says \"GPLv2\", but the license blurbs all say \"or (at your option)\nany later version\".  IANAL, does this cause any problems?  If so, can\nthey be GPLv2-only instead?\n\n>  Makefile           |    6 +\n>  ewah/bitmap.c      |  229 +++++++++++++++++\n>  ewah/ewah_bitmap.c |  703 ++++++++++++++++++++++++++++++++++++++++++++++++++++\n>  ewah/ewah_io.c     |  199 +++++++++++++++\n>  ewah/ewah_rlw.c    |  124 +++++++++\n>  ewah/ewok.h        |  194 +++++++++++++++\n>  ewah/ewok_rlw.h    |  114 +++++++++\n\nCan we have a Documentation/technical/api-ewah.txt?\n\n(Maybe if you insert all the comments I ask for in the below, it's not\nnecessary, but it would still be nice to have some central place where\nthe formats are documented.)\n\n[...]\n> +struct ewah_bitmap *bitmap_to_ewah(struct bitmap *bitmap)\n> +{\n> +\tstruct ewah_bitmap *ewah = ewah_new();\n> +\tsize_t i, running_empty_words = 0;\n> +\teword_t last_word = 0;\n> +\n> +\tfor (i = 0; i < bitmap->word_alloc; ++i) {\n> +\t\tif (bitmap->words[i] == 0) {\n> +\t\t\trunning_empty_words++;\n> +\t\t\tcontinue;\n> +\t\t}\n> +\n> +\t\tif (last_word != 0) {\n> +\t\t\tewah_add(ewah, last_word);\n> +\t\t}\n\nThere are a lot of \"noisy\" braces -- like in this instance -- if you\napply the git style to the files in ewah/.  I assume we'll give the\ndirectory its own style, so that it should always use braces even on\none-line blocks.\n\n[...]\n> +\tewah_add(ewah, last_word);\n> +\treturn ewah;\n> +}\n> +\n> +struct bitmap *ewah_to_bitmap(struct ewah_bitmap *ewah)\n> +{\n> +\tstruct bitmap *bitmap = bitmap_new();\n> +\tstruct ewah_iterator it;\n> +\teword_t blowup;\n> +\tsize_t i = 0;\n> +\n> +\tewah_iterator_init(&it, ewah);\n> +\n> +\twhile (ewah_iterator_next(&blowup, &it)) {\n> +\t\tif (i >= bitmap->word_alloc) {\n> +\t\t\tbitmap->word_alloc *= 1.5;\n\nAny reason that this uses a scale factor of 1.5, while the bitmap_set\noperation above uses 2?\n\n> +\t\t\tbitmap->words = ewah_realloc(\n> +\t\t\t\tbitmap->words, bitmap->word_alloc * sizeof(eword_t));\n> +\t\t}\n[...]\n> +\n> +void bitmap_each_bit(struct bitmap *self, ewah_callback callback, void *data)\n> +{\n[...]\n> +\t\t\tfor (offset = 0; offset < BITS_IN_WORD; ++offset) {\n> +\t\t\t\tif ((word >> offset) == 0)\n> +\t\t\t\t\tbreak;\n> +\n> +\t\t\t\toffset += __builtin_ctzll(word >> offset);\n\nHere and in the rest, you use __builtin_* within the code.  This needs\nto be either in a separate helper that reimplements the function in\nterms of C if it is not available (i.e. you don't use GCC).\n(Alternatively, the whole series could be conditional on some\nHAVE_GCC_BUILTINS macro.  I'd think that would be a bad tradeoff\nthough.)\n\n> +\t\t\t\tcallback(pos + offset, data);\n> +\t\t\t}\n> +\t\t\tpos += BITS_IN_WORD;\n> +\t\t}\n> +\t}\n> +}\n[...]\n\n> diff --git a/ewah/ewah_bitmap.c b/ewah/ewah_bitmap.c\n[...]\n> +void ewah_free(struct ewah_bitmap *bitmap)\n> +{\n> +\tif (bitmap->alloc_size)\n> +\t\tfree(bitmap->buffer);\n> +\n> +\tfree(bitmap);\n> +}\n\nMaybe first if (!bitmap) return, so that it behaves like other free()s?\n\n> diff --git a/ewah/ewah_io.c b/ewah/ewah_io.c\n[...]\n> +int ewah_serialize_native(struct ewah_bitmap *self, int fd)\n> +{\n> +\tuint32_t write32;\n> +\tsize_t to_write = self->buffer_size * 8;\n> +\n> +\t/* 32 bit -- bit size fr the map */\n\nYou cut&pasted the typo (\"for\") throughout the file :-)\n\n[...]\n> +\t/** 32 bit -- number of compressed 64-bit words */\n> +\twrite32 = (uint32_t)self->buffer_size;\n> +\tif (write(fd, &write32, 4) != 4)\n> +\t\treturn -1;\n> +\n> +\tif (write(fd, self->buffer, to_write) != to_write)\n> +\t\treturn -1;\n\nShouldn't you use our neat write_in_full() and read_in_full() helpers,\nthroughout the file?\n\n[...]\n> diff --git a/ewah/ewok.h b/ewah/ewok.h\n[...]\n> +#ifndef __EWOK_BITMAP_C__\n> +#define __EWOK_BITMAP_C__\n\n_H_?\n\n> +#ifndef ewah_malloc\n> +#\tdefine ewah_malloc malloc\n> +#endif\n> +#ifndef ewah_realloc\n> +#\tdefine ewah_realloc realloc\n> +#endif\n> +#ifndef ewah_calloc\n> +#\tdefine ewah_calloc calloc\n> +#endif\n\nI see you later #define them to the corresponding x*alloc version in\npack-bitmap.h.  Good.\n\n> +\n> +typedef uint64_t eword_t;\n\nI assume this isn't ifdef'd to help 32bit platforms because the on-disk\nformat depends on it?\n\n> +#define BITS_IN_WORD (sizeof(eword_t) * 8)\n> +\n> +struct ewah_bitmap {\n> +\teword_t *buffer;\n> +\tsize_t buffer_size;\n> +\tsize_t alloc_size;\n> +\tsize_t bit_size;\n> +\teword_t *rlw;\n> +};\n> +\n> +typedef void (*ewah_callback)(size_t pos, void *);\n> +\n> +struct ewah_bitmap *ewah_pool_new(void);\n> +void ewah_pool_free(struct ewah_bitmap *bitmap);\n\nHow do the pool versions differ from the non-pool versions below?  I\nwould have expected a memory pool argument somewhere.\n\n> +\n> +/**\n> + * Allocate a new EWAH Compressed bitmap\n> + */\n> +struct ewah_bitmap *ewah_new(void);\n> +\n> +/**\n> + * Clear all the bits in the bitmap. Does not free or resize\n> + * memory.\n> + */\n> +void ewah_clear(struct ewah_bitmap *bitmap);\n> +\n> +/**\n> + * Free all the memory of the bitmap\n> + */\n> +void ewah_free(struct ewah_bitmap *bitmap);\n> +\n> +int ewah_serialize(struct ewah_bitmap *self, int fd);\n> +int ewah_serialize_native(struct ewah_bitmap *self, int fd);\n> +\n> +int ewah_deserialize(struct ewah_bitmap *self, int fd);\n> +int ewah_read_mmap(struct ewah_bitmap *self, void *map, size_t len);\n> +int ewah_read_mmap_native(struct ewah_bitmap *self, void *map, size_t len);\n\nThe whole file is so neatly commented, and then you skimp on these? :-)\n\nIn particular, it would be nice to have a comment here on what the\n_native distinction means, and what (if any) the constraints are if you\nwant to use _mmap.  Also, if you read or deserialize, does the 'self'\nhave to be initialized first?\n\n[...]\n> +/**\n> + * Set a given bit on the bitmap.\n> + *\n> + * The bit at position `pos` will be set to true. Because of the\n> + * way that the bitmap is compressed, a set bit cannot be unset\n> + * later on.\n> + *\n> + * Furthermore, since the bitmap uses streaming compression, bits\n> + * can only set incrementally.\n\nI'm not a native speaker, but does 'incrementally' also mean 'in order\nof increasing indexes'?  That's what the example seems to say.\n\n> + *\n> + * E.g.\n> + *\t\tewah_set(bitmap, 1); // ok\n> + *\t\tewah_set(bitmap, 76); // ok\n> + *\t\tewah_set(bitmap, 77); // ok\n> + *\t\tewah_set(bitmap, 8712800127); // ok\n> + *\t\tewah_set(bitmap, 25); // failed, assert raised\n> + */\n> +void ewah_set(struct ewah_bitmap *self, size_t i);\n> +\n[...]\n> +struct ewah_bitmap * bitmap_to_ewah(struct bitmap *bitmap);\n\nStyle (around the *).\n\n> +struct bitmap *ewah_to_bitmap(struct ewah_bitmap *ewah);\n> +\n> +void bitmap_and_not_inplace(struct bitmap *self, struct bitmap *other);\n> +void bitmap_or_inplace(struct bitmap *self, struct ewah_bitmap *other);\n\nWhy does one of them take an ewah_bitmap for 'other', but the other\ntakes a straight 'bitmap'?\n\n> +\n> +void bitmap_each_bit(struct bitmap *self, ewah_callback callback, void *data);\n> +size_t bitmap_popcount(struct bitmap *self);\n> +\n> +#endif\n> diff --git a/ewah/ewok_rlw.h b/ewah/ewok_rlw.h\n> new file mode 100644\n> index 0000000..2e31836\n> --- /dev/null\n> +++ b/ewah/ewok_rlw.h\n> @@ -0,0 +1,114 @@\n[...]\n> +#define RLW_RUNNING_BITS (sizeof(eword_t) * 4)\n> +#define RLW_LITERAL_BITS (sizeof(eword_t) * 8 - 1 - RLW_RUNNING_BITS)\n\nIt would be nice to have some minimal documentation of the word format\nhere (or in ewok.h), in particular because you snip off 1 bit here for a\nreason that is not immediately obvious.\n\n> +#define RLW_LARGEST_RUNNING_COUNT (((eword_t)1 << RLW_RUNNING_BITS) - 1)\n> +#define RLW_LARGEST_LITERAL_COUNT (((eword_t)1 << RLW_LITERAL_BITS) - 1)\n> +\n> +#define RLW_LARGEST_RUNNING_COUNT_SHIFT (RLW_LARGEST_RUNNING_COUNT << 1)\n> +\n> +#define RLW_RUNNING_LEN_PLUS_BIT (((eword_t)1 << (RLW_RUNNING_BITS + 1)) - 1)\n\nThis one is doubly strange.  The name claims it's a bit(?), but the\ndefinition (if you expand the preceding macros) effectively makes it\n0x1ffffffff, i.e., a mask for RLW_RUNNING_BITS+1 number of bits.\n\n> +static inline void rlw_xor_run_bit(eword_t *word)\n> +{\n> +\tif (*word & 1) {\n> +\t\t*word &= (eword_t)(~1);\n> +\t} else {\n> +\t\t*word |= (eword_t)1;\n> +\t}\n> +}\n\nWhy is this called xor?  Looks a lot like a negation to me.\n\n> +static bool rlw_get_run_bit(const eword_t *word)\n> +{\n> +\treturn *word & (eword_t)1;\n> +}\n[...]\n> +static inline void rlw_set_running_len(eword_t *word, eword_t l)\n> +{\n> +\t*word |= RLW_LARGEST_RUNNING_COUNT_SHIFT;\n> +\t*word &= (l << 1) | (~RLW_LARGEST_RUNNING_COUNT_SHIFT);\n> +}\n> +\n> +static inline void rlw_set_literal_words(eword_t *word, eword_t l)\n> +{\n> +\t*word |= ~RLW_RUNNING_LEN_PLUS_BIT;\n> +\t*word &= (l << (RLW_RUNNING_BITS + 1)) | RLW_RUNNING_LEN_PLUS_BIT;\n> +}\n\n>From these I gather that the layout is, LSB first:\n\n  1 bit:   bit that will be repeated\n  32 bits: length of the run\n  31 bits: number of literal words to be read after the run\n\nIs that correct?  This took some figuring out for me, please add a\ncomment.\n\nAnd then from there I would extrapolate that the data format requires\none such \"specifier\" word in between of chunks of stuff, but it's not\nclear how exactly.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"221977","messageId":"87sj05lvsy.fsf@linux-k42r.v.cablecom.net","threadId":"34271","inReplyTo":"1372116193-32762-11-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 10/16] pack-objects: use bitmaps when packing objects","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-06-25T15:58:42Z","receivedAt":"2013-06-25T15:58:42Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Vicent Marti <tanoku@gmail.com> writes:\n\n> diff --git a/Makefile b/Makefile\n> index e03c773..0f2e72b 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -703,6 +703,7 @@ LIB_H += notes.h\n>  LIB_H += object.h\n>  LIB_H += pack-revindex.h\n>  LIB_H += pack.h\n> +LIB_H += pack-bitmap.h\n>  LIB_H += parse-options.h\n>  LIB_H += patch-ids.h\n>  LIB_H += pathspec.h\n> @@ -838,6 +839,7 @@ LIB_OBJS += notes.o\n>  LIB_OBJS += notes-cache.o\n>  LIB_OBJS += notes-merge.o\n>  LIB_OBJS += object.o\n> +LIB_OBJS += pack-bitmap.o\n>  LIB_OBJS += pack-check.o\n>  LIB_OBJS += pack-revindex.o\n>  LIB_OBJS += pack-write.o\n\nWhat does this apply on?  When starting with the series from\norigin/master, git-am fails, and 'git am -3' tells me I don't have the\nnecessary blobs (from the 'index' line above).\n\nNot that it's super hard to fix this up as long as it's in the Makefile\nonly, but still.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"221976","messageId":"87y59xlvt7.fsf@linux-k42r.v.cablecom.net","threadId":"34271","inReplyTo":"1372116193-32762-10-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-06-25T15:58:54Z","receivedAt":"2013-06-25T15:58:54Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Vicent Marti <tanoku@gmail.com> writes:\n\n> This is the technical documentation and design rationale for the new\n> Bitmap v2 on-disk format.\n\nHrmpf, that's what I get for reading the series in order...\n\n> +\t\t\tThe folowing flags are supported:\n                              ^^\n\ntypos marked by ^\n\n> +\t\tBy storing all the hashes in a cache together with the bitmapsin\n                                                                             ^^\n\n> +\t\tThe obvious consequence is that the XOR of all 4 bitmaps will result\n> +\t\tin a full set (all bits sets), and the AND of all 4 bitmaps will\n                                           ^\n\n> +\t\t- 1-byte XOR-offset\n> +\t\t\tThe xor offset used to compress this bitmap. For an entry\n> +\t\t\tin position `x`, a XOR offset of `y` means that the actual\n> +\t\t\tbitmap representing for this commit is composed by XORing the\n> +\t\t\tbitmap for this entry with the bitmap in entry `x-y` (i.e.\n> +\t\t\tthe bitmap `y` entries before this one).\n> +\n> +\t\t\tNote that this compression can be recursive. In order to\n> +\t\t\tXOR this entry with a previous one, the previous entry needs\n> +\t\t\tto be decompressed first, and so on.\n> +\n> +\t\t\tThe hard-limit for this offset is 160 (an entry can only be\n> +\t\t\txor'ed against one of the 160 entries preceding it). This\n> +\t\t\tnumber is always positivea, and hence entries are always xor'ed\n                                                 ^\n\n> +\t\t\twith **previous** bitmaps, not bitmaps that will come afterwards\n> +\t\t\tin the index.\n\nClever.  Why 160 though?\n\n> +\t\t- 2 bytes of RESERVED data (used right now for better packing).\n\nWhat do they mean?\n\n> +  With an index at the end of the file, we can load only this index in memory,\n> +  allowing for very efficient access to all the available bitmaps lazily (we\n> +  have their offsets in the mmaped file).\n\nIs there anything preventing you from mmap()ing the index also?\n\n> +== Appendix A: Serialization format for an EWAH bitmap\n> +\n> +Ewah bitmaps are serialized in the protocol as the JAVAEWAH\n> +library, making them backwards compatible with the JGit\n> +implementation:\n> +\n> +\t- 4-byte number of bits of the resulting UNCOMPRESSED bitmap\n> +\n> +\t- 4-byte number of words of the COMPRESSED bitmap, when stored\n> +\n> +\t- N x 8-byte words, as specified by the previous field\n> +\n> +\t\tThis is the actual content of the compressed bitmap.\n> +\n> +\t- 4-byte position of the current RLW for the compressed\n> +\t\tbitmap\n> +\n> +Note that the byte order for this serialization is not defined by\n> +default. The byte order for all the content in a serialized EWAH\n> +bitmap can be known by the byte order flags in the header of the\n> +bitmap index file.\n\nPlease document the RLW format here.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"221979","messageId":"87hagllvsk.fsf@linux-k42r.v.cablecom.net","threadId":"34271","inReplyTo":"1372116193-32762-1-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 00/16] Speed up Counting Objects with bitmap data","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-06-25T16:05:05Z","receivedAt":"2013-06-25T16:05:05Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Vicent Marti <tanoku@gmail.com> writes:\n\n> Like with every other patch that offers performance improvements,\n> sample benchmarks are provided (spoiler: they are pretty fucking\n> cool).\n\nGreat stuff.\n\nI read the first half, and skimmed the second half.  See the individual\nreplies for comments.\n\nHowever:\n\n>  Documentation/technical/bitmap-format.txt |  235 ++++++++\n>  Makefile                                  |   11 +\n>  builtin.h                                 |    1 +\n>  builtin/pack-objects.c                    |  362 +++++++-----\n>  builtin/pack-objects.h                    |   33 ++\n>  builtin/rev-list.c                        |   35 +-\n>  builtin/write-bitmap.c                    |  256 +++++++++\n>  cache.h                                   |    5 +\n>  ewah/bitmap.c                             |  229 ++++++++\n>  ewah/ewah_bitmap.c                        |  703 ++++++++++++++++++++++++\n>  ewah/ewah_io.c                            |  199 +++++++\n>  ewah/ewah_rlw.c                           |  124 +++++\n>  ewah/ewok.h                               |  194 +++++++\n>  ewah/ewok_rlw.h                           |  114 ++++\n>  git-compat-util.h                         |   28 +\n>  git-repack.sh                             |   10 +-\n>  git.c                                     |    1 +\n>  khash.h                                   |  329 +++++++++++\n>  list-objects.c                            |    1 +\n>  pack-bitmap-write.c                       |  520 ++++++++++++++++++\n>  pack-bitmap.c                             |  855 +++++++++++++++++++++++++++++\n>  pack-bitmap.h                             |   64 +++\n>  pack-write.c                              |    2 +\n>  revision.c                                |    5 +\n>  revision.h                                |    2 +\n>  sha1_file.c                               |   57 +-\n\nIt's pretty hard to miss that there isn't a single test in the entire\nseries.  It seems that the features you add depend on pack.usebitmaps,\nand since the tests run with empty config (unless of course they set\ntheir own) your feature is completely untested -- unless I'm missing\nsomething.\n\nI imagine the tests would be of the format\n\ntest_expect_success 'do <stuff> without bitmaps' '\n\tgit ... >expect\n'\n\ntest_expect_success 'do <stuff> with bitmaps' '\n\ttest_config pack.usebitmaps true &&\n\t# do something to ensure that we have bitmaps\n\tgit ... >actual &&\n\ttest_cmp expect actual\n'\n\nor some such.\n\nFor bonus points, you could also add some light performance tests in\nt/perf/, just to show off ;-)\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"221978","messageId":"87mwqdlvsq.fsf@linux-k42r.v.cablecom.net","threadId":"34271","inReplyTo":"1372116193-32762-12-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 11/16] rev-list: add bitmap mode to speed up lists","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-06-25T16:22:28Z","receivedAt":"2013-06-25T16:22:28Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Vicent Marti <tanoku@gmail.com> writes:\n\n> Calling `git rev-list --use-bitmaps [committish]` is the equivalent\n> of `git rev-list --objects`, but the rev list is performed based on\n> a bitmap result instead of using a manual counting objects phase.\n\nWhy would we ever want to not --use-bitmaps, once it actually works?\nI.e., shouldn't this be the default if pack.usebitmaps is set (or\npossibly even core.usebitmaps for these things)?\n\n> These are some example timings for `torvalds/linux`:\n>\n> \t$ time ../git/git rev-list --objects master > /dev/null\n>\n> \treal    0m25.567s\n> \tuser    0m25.148s\n> \tsys     0m0.384s\n>\n> \t$ time ../git/git rev-list --use-bitmaps master > /dev/null\n>\n> \treal    0m0.393s\n> \tuser    0m0.356s\n> \tsys     0m0.036s\n\nI see your badass numbers, and raise you a critical issue:\n\n  $ time git rev-list --use-bitmaps --count --left-right origin/pu...origin/next\n  Segmentation fault\n\n  real    0m0.408s\n  user    0m0.383s\n  sys     0m0.022s\n\nIt actually seems to be related solely to having negated commits in the\nwalk:\n\n  thomas@linux-k42r:~/g(next u+65)$ time git rev-list --use-bitmaps --count origin/pu\n  32315\n\n  real    0m0.041s\n  user    0m0.034s\n  sys     0m0.006s\n  thomas@linux-k42r:~/g(next u+65)$ time git rev-list --use-bitmaps --count origin/pu ^origin/next\n  Segmentation fault\n\n  real    0m0.460s\n  user    0m0.214s\n  sys     0m0.244s\n\nI also can't help noticing that the time spent generating the segfault\nwould have sufficed to generate the answer \"the old way\" as well:\n\n  $ time git rev-list --count --left-right origin/pu...origin/next\n  189     125\n\n  real    0m0.409s\n  user    0m0.386s\n  sys     0m0.022s\n\nCan we use the same trick to speed up merge base computation and then\n--left-right?  The latter is a component of __git_ps1 and can get\nsomewhat slow in some cases, so it would be nice to make it really fast,\ntoo.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"221961","messageId":"CALkWK0krR=PaGC6iX8abmSswr7HBtzU61-7DR00xA3CyW80dtA@mail.gmail.com","threadId":"34271","inReplyTo":"1372116193-32762-4-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 03/16] pack-objects: use a faster hash table","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-06-25T17:58:12Z","receivedAt":"2013-06-25T17:58:12Z","isPatch":true,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Vicent Marti wrote:\n> When replacing the old hash table implementation in `pack-objects` with\n> the khash_sha1 table, the insertion time is greatly reduced:\n\nWhy?  What is the exact change?\n\n> The big win here, however, is in the massively reduced amount of hash\n> collisions\n\nOkay, so there seems to be some problem with how collisions are\nhandled in the hashtable.\n\n> -static int locate_object_entry_hash(const unsigned char *sha1)\n> -{\n> -       int i;\n> -       unsigned int ui;\n> -       memcpy(&ui, sha1, sizeof(unsigned int));\n> -       i = ui % object_ix_hashsz;\n> -       while (0 < object_ix[i]) {\n> -               if (!hashcmp(sha1, objects[object_ix[i] - 1].idx.sha1))\n> -                       return i;\n> -               if (++i == object_ix_hashsz)\n> -                       i = 0;\n> -       }\n> -       return -1 - i;\n> -}\n\nClassical chaining to handle collisions: very naive.  Deserves to be thrown out.\n\n> -static void rehash_objects(void)\n> -{\n> -       uint32_t i;\n> -       struct object_entry *oe;\n> -\n> -       object_ix_hashsz = nr_objects * 3;\n> -       if (object_ix_hashsz < 1024)\n> -               object_ix_hashsz = 1024;\n> -       object_ix = xrealloc(object_ix, sizeof(int) * object_ix_hashsz);\n> -       memset(object_ix, 0, sizeof(int) * object_ix_hashsz);\n> -       for (i = 0, oe = objects; i < nr_objects; i++, oe++) {\n> -               int ix = locate_object_entry_hash(oe->idx.sha1);\n> -               if (0 <= ix)\n> -                       continue;\n> -               ix = -1 - ix;\n> -               object_ix[ix] = i + 1;\n> -       }\n> -}\n\nThis is called when the hashtable runs out of space.  It didn't appear\nin your profiler because it doesn't appear to be a bottleneck, right?\nGrowing aggressively to 3x times the number of objects probably\nexplains it.  Just for comparison, how does khash grow?\n\n>  static struct object_entry *locate_object_entry(const unsigned char *sha1)\n>  {\n> -       int i;\n> +       khiter_t pos = kh_get_sha1(packed_objects, sha1);\n>\n> -       if (!object_ix_hashsz)\n> -               return NULL;\n> +       if (pos < kh_end(packed_objects)) {\n\nWait, why is this required?  When will kh_get_sha1() return a position\nbeyond kh_end()?  What does that mean?\n\n> +               return kh_value(packed_objects, pos);\n> +       }\n>\n> -       i = locate_object_entry_hash(sha1);\n> -       if (0 <= i)\n> -               return &objects[object_ix[i]-1];\n>         return NULL;\n>  }\n\nOverall, replaced call to locate_object_entry_hash() with a call to\nkh_get_sha1().  Okay.\n\n> -static int add_object_entry(const unsigned char *sha1, enum object_type type,\n> -                           const char *name, int exclude)\n> +static int add_object_entry_1(const unsigned char *sha1, enum object_type type,\n> +                           uint32_t hash, int exclude, struct packed_git *found_pack,\n> +                               off_t found_offset)\n>  {\n>         struct object_entry *entry;\n> -       struct packed_git *p, *found_pack = NULL;\n> -       off_t found_offset = 0;\n> -       int ix;\n> -       unsigned hash = name_hash(name);\n> +       struct packed_git *p;\n> +       khiter_t ix;\n> +       int hash_ret;\n>\n> -       ix = nr_objects ? locate_object_entry_hash(sha1) : -1;\n> -       if (ix >= 0) {\n> +       ix = kh_put_sha1(packed_objects, sha1, &hash_ret);\n\nYou don't need to call locate_object_entry() to check for collisions\nbecause kh_put_sha1() takes care of that?\n\n> +       if (hash_ret == 0) {\n>                 if (exclude) {\n> -                       entry = objects + object_ix[ix] - 1;\n> +                       entry = kh_value(packed_objects, ix);\n\nSuperficial change: using kh_value(), because we stripped out the\nchaining logic.\n\n> @@ -966,19 +965,30 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n>                 entry->in_pack_offset = found_offset;\n>         }\n>\n> -       if (object_ix_hashsz * 3 <= nr_objects * 4)\n> -               rehash_objects();\n> -       else\n> -               object_ix[-1 - ix] = nr_objects;\n> +       kh_value(packed_objects, ix) = entry;\n> +       kh_key(packed_objects, ix) = entry->idx.sha1;\n> +       objects[nr_objects++] = entry;\n\nWait, what?  Why didn't you use kh_put_sha1()?\n\nI didn't look very carefully, but the patch seems to be okay overall.\nOn the issue of which hashtable replacement to use (why khash, and not\nsomething else?), I briefly looked at linux.git's linux/hashtable.h\nand git.git's hash.h; both of them are chaining hashes.  From a brief\nlook at khash.h, it seems to be somewhat less naive and sane: my only\nconcern is that it is written entirely in using CPP macros which is a\ngreat for syntax/performance, but not-so-great for debugging.  I don't\nknow if there's a better off-the-shelf implementation out there, but I\nhaven't been looking for one either.  By the way, it's MIT license\nauthored by an anonymous person (sources at:\nhttps://github.com/attractivechaos/klib/blob/master/khash.h), but I\ndon't know if that's a problem.\n\nThanks.\n"},{"id":"221965","messageId":"CAFFjANRwBBcORhu4mwjESBfr4GJ3zDrgYvUhY=VxK9abv7k2MA@mail.gmail.com","threadId":"34271","inReplyTo":"CAJo=hJtcQwh-N-9_i84y1ZsL0mdREHcxhP2gepcrREiaxvxS6A@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Vicent Martí","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-25T19:33:11Z","receivedAt":"2013-06-25T19:33:11Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"On Tue, Jun 25, 2013 at 7:42 AM, Shawn Pearce <spearce@spearce.org> wrote:\n> I very much hate seeing a file format that is supposed to be portable\n> that supports both big-endian and little-endian encoding.\n\nWell, the bitmap index is not supposed to be portable, as it doesn't\nget sent over the wire in any situation. Regardless, the format is\nportable because it supports both encodings and clearly defines which\none the current file is using. I think that's a good tradeoff!\n\n> Such a specification forces everyone to implement two code paths to handle\n> reading data from the file, on the off-chance they are on the wrong\n> platform.\n\nExtra code paths have never been an issue in the\nJGitAbstractFactoryGeneratorOfGit, har har har. Ah. I'm such a funny\nguy when it comes to Java.\n\nAnyway, I designed this keeping JGit in mind. In this specific case,\nit doesn't force you to add any new code paths. The endianness changes\nonly affect the serialization format of the bitmaps, which is not part\nof Git or JGit itself but of the Javaewah/libewok library. The\ninterface for reading on that library has already been wisely\nabstracted on JGit\n(https://github.com/eclipse/jgit/blob/master/org.eclipse.jgit/src/org/eclipse/jgit/internal/storage/file/PackBitmapIndexV1.java#L133),\nso changing the byte order simply means changing the SimpleDataInput\nto a LE one.\n\nI agree this is not ideal, or elegant, but I'm having a hard time\nmaking an argument for sacrificing objective speed for the sake of\nsubjective \"simplicity\".\n\n> What is wrong with picking one encoding and sticking to it?\n\nIt prevents us from making this optimally fast on the machines where\nit needs to be.\n\nRegardless, I must admit I haven't generated numbers for this in a\nwhile (the BE->LE switch is one of the first optimizations I did). I'm\ngoing to try to re-implement full NWO loading and see how much slower\nit is, before I continue arguing for/against it.\n\nIf I can get it within a reasonable margin (say, 15%) of the current\nimplementation, I'd definitely be in favor of sticking to only NWO on\nthe whole file. If it's slower than that, well, Git has never\ncompromised on speed, and I don't think there's a point to be made for\nstarting to do that now.\n\n>> +                       - BITMAP_OPT_HASH_CACHE (0x8)\n>> +                       If present, a hash cache for finding delta bases will be available\n>> +                       right after the header block in this index. See the following\n>> +                       section for details.\n>> +\n>> +               4-byte entry count (network byte order)\n>> +\n>> +                       The total count of entries (bitmapped commits) in this bitmap index.\n>> +\n>> +               20-byte checksum\n>> +\n>> +                       The SHA1 checksum of the pack this bitmap index belongs to.\n>> +\n>> +       - An OPTIONAL delta cache follows the header.\n>\n> Some may find the name \"delta cache\" confusing as it does not cache\n> deltas of objects. May I suggest \"path hash cache\" as an alternative\n> name?\n\nDefinitely, this is a typo.\n\n>> +               The cache is formed by `n` 4-byte hashes in a row, where `n` is\n>> +               the amount of objects in the indexed packfile. Note that this amount\n>> +               is the **total number of objects** and is not related to the\n>> +               number of commits that have been selected and indexed in the\n>> +               bitmap index.\n>> +\n>> +               The hashes are stored in Network Byte Order and they are the same\n>> +               values generated by a normal revision walk during the `pack-objects`\n>> +               phase.\n>\n> I find it interesting this is network byte order and not big-endian or\n> little-endian based on the flag in the header.\n\nAs stated before, the flag in the header only affects the\nJavaewah/libewok interface. Everything Git-related in the bitmap index\nis in NWO, like it is customary in Git.\n\n>\n>> +               The `n`nth hash in the cache is the name hash for the `n`th object\n>> +               in the index for the indexed packfile.\n>> +\n>> +               [RATIONALE]:\n>> +\n>> +               The bitmap index allows us to skip the Counting Objects phase\n>> +               during `pack-objects` and yield all the OIDs that would be reachable\n>> +               (\"WANTS\") when generating the pack.\n>> +\n>> +               This optimization, however, means that we're adding objects to the\n>> +               packfile straight from the packfile index, and hence we are lacking\n>> +               path information for the objects that would normally be generated\n>> +               during the \"Counting Objects\" phase.\n>> +\n>> +               This path information for each object is hashed and used as a very\n>> +               effective way to find good delta bases when compressing the packfile;\n>> +               without these hashes, the resulting packfiles are much less optimal.\n>> +\n>> +               By storing all the hashes in a cache together with the bitmapsin\n>> +               the bitmap index, we can yield not only the SHA1 of all the reachable\n>> +               objects, but also their hashes, and allow Git to be much smarter when\n>> +               finding delta bases for packing.\n>> +\n>> +               If the delta cache is not available, the bitmap index will obviously\n>> +               be smaller in disk, but the packfiles generated using this index will\n>> +               be between 20% and 30% bigger, because of the lack of name/path\n>> +               information when finding delta bases.\n>\n> JGit does not encode this because we were afraid of freezing the hash\n> function into the file format. Indeed we are not certain JGit even\n> uses the same path hash function as C Git does, because C Git's\n> implementation is covered by the GPL and JGit prefers to license its\n> work under BSD.\n>\n> If the path hash is going to become part of the format, the algorithm\n> for computing the hash should also be specified in the format so that\n> non-GPL implementations have an opportunity to be compatible.\n\nVery valid point. I would hope that whoever wrote the original hash\nwill let us re-license it under BSD. It would be nice to define this\nhash function as part of the format.\n\n> One way we side-stepped the size inflation problem in JGit was to only\n> use the bitmap index information when sending data on the wire to a\n> client. Here delta reuse plays a significant factor in building the\n> pack, and we don't have to be as accurate on matching deltas. During\n> the equivalent of `git repack` bitmaps are not used, allowing the\n> traditional graph enumeration algorithm to generate path hash\n> information.\n\nOH BOY HERE WE GO. This is worth its own thread, lots to discuss here.\nI think peff will have a patchset regarding this to upstream soon,\nwe'll get back to it later.\n\n>\n>> +       - 4 EWAH bitmaps that act as type indexes\n>> +\n>> +               Type indexes are serialized after the hash cache in the shape\n>> +               of four EWAH bitmaps stored consecutively (see Appendix A for\n>> +               the serialization format of an EWAH bitmap).\n>> +\n>> +               There is a bitmap for each Git object type, stored in the following\n>> +               order:\n>> +\n>> +                       - Commits\n>> +                       - Trees\n>> +                       - Blobs\n>> +                       - Tags\n>> +\n>> +               In each bitmap, the `n`th bit is set to true if the `n`th object\n>> +               in the packfile index is of that type.\n>> +\n>> +               The obvious consequence is that the XOR of all 4 bitmaps will result\n>> +               in a full set (all bits sets), and the AND of all 4 bitmaps will\n>> +               result in an empty bitmap (no bits set).\n>\n> Instead of XOR did you mean OR here?\n\nNope, I think XOR makes it more obvious: if the same bit is set on two\nbitmaps, it would be cleared when XORed together, and hence all the\nbits wouldn't be set. An OR would hide this case.\n\n>\n>> +       - N EWAH bitmaps, one for each indexed commit\n>> +\n>> +               Where `N` is the total amount of entries in this bitmap index.\n>> +               See Appendix A for the serialization format of an EWAH bitmap.\n>> +\n>> +       - An entry index with `N` entries for the indexed commits\n>> +\n>> +               Index entries are stored consecutively, and each entry has the\n>> +               following format:\n>> +\n>> +               - 20-byte SHA1\n>> +                       The SHA1 of the commit that this bitmap indexes\n>> +\n>> +               - 4-byte offset (Network Byte Order)\n>> +                       The offset **from the beginning of the file** where the\n>> +                       bitmap for this commit is stored.\n>\n> Eh, another network byte order field in a file that also has selective\n> ordering. *sigh*\n\n:D :D :D again only the Javaewah interface.\n\n>\n>> +               - 1-byte XOR-offset\n>> +                       The xor offset used to compress this bitmap. For an entry\n>> +                       in position `x`, a XOR offset of `y` means that the actual\n>> +                       bitmap representing for this commit is composed by XORing the\n>> +                       bitmap for this entry with the bitmap in entry `x-y` (i.e.\n>> +                       the bitmap `y` entries before this one).\n>> +\n>> +                       Note that this compression can be recursive. In order to\n>> +                       XOR this entry with a previous one, the previous entry needs\n>> +                       to be decompressed first, and so on.\n>> +\n>> +                       The hard-limit for this offset is 160 (an entry can only be\n>> +                       xor'ed against one of the 160 entries preceding it). This\n>> +                       number is always positivea, and hence entries are always xor'ed\n>> +                       with **previous** bitmaps, not bitmaps that will come afterwards\n>> +                       in the index.\n>\n> What order are these entries in? Sorted by SHA-1 or random?\n\nNot-specified, since the whole index will be loaded in a hash table\nlike JGit does. In practice, it's toposorted because that makes for\nbetter XOR bases.\n\n> Colby found that doing an XOR against the descendant commit yielded\n> very small bitmaps, so JGit tries to XOR-compress bitmaps along common\n> linear slices of history. This is trivial in Linus' kernel tree where\n> there is effectively only one history, but its more relevant with\n> long-running side branches that have release tags that may not have\n> fully merged into \"master\".\n\nIndeed, indeed. We try to do the same.\n\n>> +  This full parse, however, requires prohibitive loading times in LE\n>> +  machines (i.e. all modern server hardware): a repository like\n>> +  `torvalds/linux` can have about 8mb of bitmap indexes, resulting\n>> +  in roughly 400ms of parse time.\n>\n> This makes me wonder what the JGit parse time is. It is ugly if we are\n> spending 400ms to load the bitmap index for the kernel repository.\n\nIt's very bad. Of course you don't notice it because you're running in\n`daemon` mode. It takes about 3s to load indexes for the\n`torvalds/linux` network in my machine.\n\n>\n>> +  This is not an issue in JGit, which is capable of serving repositories\n>> +  from a single-process daemon running on the JVM, but `git-daemon` in\n>> +  git has been implemented with a process-based design (a new\n>> +  `pack-objects` is spawned for each request), and the boot times\n>> +  of parsing the bitmap index every time `pack-objects` is spawned can\n>> +  seriously slow down requests (particularly for small fetches, where we'd\n>> +  spend about 1.5s booting up and 300ms performing the Counting Objects\n>> +  phase).\n>\n> There are other strategies that Git could use to handle request\n> processing at scale. But I guess its reasonable to assume these aren't\n> viable for Git for a number of reasons. E.g. \"long tail\" access effect\n> that many servers have, where most requests are to a large number of\n> repositories that themselves receive very few requests, an environment\n> that does not lend itself to caching.\n>\n>> +  By storing the bitmaps in Little-Endian, we're able to `mmap` their\n>> +  compressed data straight in memory without parsing it beforehand, and\n>> +  since most queries don't require accessing all the serialized bitmaps,\n>> +  we'll only page in the minimal amount of bitmaps necessary to perform\n>> +  the reachability analysis as they are accessed.\n>\n> FWIW the .idx and .pack file formats `mmap` the compressed data\n> straight into memory without parsing it beforehand, and do not use\n> little-endian byte order. It is possible to have a single compressed\n> file format definition that is portable to all architectures, and is\n> accessed by mmap, at scale, with reasonable efficiency.\n>\n>> +- An index of all the bitmapped commits is written at the end of the packfile,\n>> +  instead of interpersed with the serialized bitmaps in the middle of the\n>> +  file.\n>\n> This is probably a mistake in the JGit design. Your approach is\n> slightly more complex, but in general I agree with having a table of\n> the SHA-1s isolated from the bitmaps themselves so that a reader can\n> access specific bitmaps at random without needing to wade through all\n> compressed bitmaps.\n>\n> I would have proposed putting the table at the start of the file, not\n> the end. The writer making the file can completely serialize the\n> bitmaps into memory before writing them to disk, and thus knows the\n> full layout of the resulting file. If the bitmaps don't fit in RAM at\n> writing time, game over, the optimization of having a very compact\n> representation of the graph is no longer helping you.\n>\n>> +  Again, the old design implied a full parse of the whole bitmap index\n>> +  (which JGit can afford because its daemon is single-process), but it made\n>> +  impossible `mmaping` the bitmap index file and accessing only the parts\n>> +  required to actually solve the query.\n>> +\n>> +  With an index at the end of the file, we can load only this index in memory,\n>> +  allowing for very efficient access to all the available bitmaps lazily (we\n>> +  have their offsets in the mmaped file).\n>> +\n>> +- The ordering of the objects in each bitmap has changed from\n>> +  packfile-order (the nth bit in the bitmap is the nth object in the\n>> +  packfile) to index-order (the nth bit in the bitmap is the nth object\n>> +  in the INDEX of the packfile).\n>\n> Did you notice an increase in bitmap size when you did this? Colby\n> tested both orderings and we observed the bitmaps were quantifiably\n> smaller when using the pack file ordering, due to the pack file\n> locality rules and the EWAH compression. Using the pack file ordering\n> was a very conscious design decision.\n>\n>> +  There is not a noticeable performance difference when actually converting\n>> +  from bitmap position to SHA1 and from SHA1 to bitmap position, but when\n>> +  using packfile ordering like JGit does, queries need to go through the\n>> +  reverse index (pack-revindex.c).\n>> +\n>> +  Generating this reverse index at runtime is **not** free (around 900ms\n>> +  generation time for a repository like `torvalds/linux`), and once again,\n>> +  this generation time needs to happen every time `pack-objects` is\n>> +  spawned.\n>\n> Did you know the packer needs the reverse index in order to compute\n> the end offset of an object it will copy as-is during delta reuse? How\n> have you avoided making the reverse index?\n\nI'm aware of that, but there are lots of other operations that can be\noptimized with bitmaps that don't require a reverse index.\n\n> Again this is why we chose to pin the JGit bitmap on the reverse index\n> being present. It already had to be present to support as-is reuse.\n> Once we knew we had to have that reverse index it was OK to rely on it\n> to get better compression on the bitmaps, and thus make them take up\n> less memory when loaded into a server. Even if you mmap a file you\n> want it to be small so it is more likely to retain in the kernel\n> buffer cache across process invocations.\n\nMaybe this applies to the JVM (where you have to load the whole\nindex), but bitmap indexes are consistently one order of magnitude\nsmaller than packfile indexes (regardless of whether they use index\nordering or packfile ordering), and the packfile indexes are always\nmapped on memory: This has never been an issue in Git, and I don't see\nwhy it would become an issue now.\n\nPinning the bitmap index on the reverse index adds complexity (lookups\nare two-step: first find the entry in the reverse index, and then find\nthe SHA1 in the index) and is measurably slower, in both loading and\nlookup times. Since Git doesn't have a memory problem, it's very hard\nto make an argument for design that is more complex and runs slower to\nsave memory.\n\nTo sum it up: I'd like to see this format be strictly in Network Byte\nOrder, and I'm going to try to make it run fast enough in that\nencoding. Having the entry index at the end of the file and having the\nbitmaps in index-order are Good Ideas (TM) because they are measurably\nsimpler and faster than their counterpoints. Do show code & benchmarks\nif you think otherwise, though.\n\nstrawberry and watermelon kisses,\nvmg\n"},{"id":"221970","messageId":"7vtxkl28m7.fsf@alter.siamese.dyndns.org","threadId":"34271","inReplyTo":"CAFFjANRwBBcORhu4mwjESBfr4GJ3zDrgYvUhY=VxK9abv7k2MA@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-06-25T21:17:20Z","receivedAt":"2013-06-25T21:17:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Vicent Martí <tanoku@gmail.com> writes:\n\n>>> +               There is a bitmap for each Git object type, stored in the following\n>>> +               order:\n>>> +\n>>> +                       - Commits\n>>> +                       - Trees\n>>> +                       - Blobs\n>>> +                       - Tags\n>>> +\n>>> +               In each bitmap, the `n`th bit is set to true if the `n`th object\n>>> +               in the packfile index is of that type.\n>>> +\n>>> +               The obvious consequence is that the XOR of all 4 bitmaps will result\n>>> +               in a full set (all bits sets), and the AND of all 4 bitmaps will\n>>> +               result in an empty bitmap (no bits set).\n>>\n>> Instead of XOR did you mean OR here?\n>\n> Nope, I think XOR makes it more obvious: if the same bit is set on two\n> bitmaps, it would be cleared when XORed together, and hence all the\n> bits wouldn't be set. An OR would hide this case.\n\nWhat case are you talking about?\n\nThe n-th object must be one of these four types and can never be of\nmore than one type at the same time, so a natural expectation from\nthe reader is \"If you OR them together, you will get the same set\".\nIf you say \"If you XOR them\", that forces the reader to wonder when\nthese bitmaps ever can overlap at the same bit position.\n\n> To sum it up: I'd like to see this format be strictly in Network Byte\n> Order,\n\nGood.\n\nI've been wondering what you meant by \"cannot be mmap-ed\" from the\nvery beginning.  We mmapped the index for a long time, and it is\ndefined in terms of network byte order.  Of course, pack .idx files\nare in network byte order, too, and we mmap them without problems.\nIt seems that it primarily came from your fear that using network\nbyte order may be unnecessarily hard to perform well, and it would\nbe a good thing to do to try to do so first instead of punting from\nthe beginning.\n\n> and I'm going to try to make it run fast enough in that\n> encoding.\n\nHmph.  Is it an option to start from what JGit does, so that people\ncan use both JGit and your code on the same repository?  And then if\nyou do not succeed, after trying to optimize in-core processing\nusing that on-disk format to make it fast enough, start thinking\nabout tweaking the on-disk format?\n\nThanks.\n"},{"id":"221984","messageId":"CAFFjANRqZ0U5tGhgjACUtquyVKCyuHiS3CC2Xxwo0J1UJVrf=g@mail.gmail.com","threadId":"34271","inReplyTo":"7vtxkl28m7.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Vicent Martí","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-25T22:08:20Z","receivedAt":"2013-06-25T22:08:20Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"On Tue, Jun 25, 2013 at 11:17 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> What case are you talking about?\n>\n> The n-th object must be one of these four types and can never be of\n> more than one type at the same time, so a natural expectation from\n> the reader is \"If you OR them together, you will get the same set\".\n> If you say \"If you XOR them\", that forces the reader to wonder when\n> these bitmaps ever can overlap at the same bit position.\n\nI guess this is just wording. I don't particularly care about the\ndistinction, but I'll change it to OR.\n\n>\n>> To sum it up: I'd like to see this format be strictly in Network Byte\n>> Order,\n>\n> Good.\n>\n> I've been wondering what you meant by \"cannot be mmap-ed\" from the\n> very beginning.  We mmapped the index for a long time, and it is\n> defined in terms of network byte order.  Of course, pack .idx files\n> are in network byte order, too, and we mmap them without problems.\n> It seems that it primarily came from your fear that using network\n> byte order may be unnecessarily hard to perform well, and it would\n> be a good thing to do to try to do so first instead of punting from\n> the beginning.\n\nIt cannot be mmapped not particularly because of endianness issues,\nbut because the original format is not indexed and requires a full\nparse of the whole index before it can be accessed programatically.\nThe wrong endianness just increases the parse time.\n\n>\n>> and I'm going to try to make it run fast enough in that\n>> encoding.\n>\n> Hmph.  Is it an option to start from what JGit does, so that people\n> can use both JGit and your code on the same repository?  And then if\n> you do not succeed, after trying to optimize in-core processing\n> using that on-disk format to make it fast enough, start thinking\n> about tweaking the on-disk format?\n\nI'm afraid this is not an option. I have an old patchset that\nimplements JGit v1 bitmap loading (and in fact that's how I initially\ndeveloped these series -- by loading the bitmaps from JGit for\ndebugging), but I discarded it because it simply doesn't pan out in\nproduction. ~3 seconds time to spawn `upload-pack` is not an option\nfor us. I did not develop a tweaked on-disk format out of boredom.\n\nI could dig up the patch if you're particularly interested in\nbackwards compatibility, but since it was several times slower than\nthe current iteration, I have no interest (time, actually) to maintain\nit, brush it up, and so on. I have already offered myself to port the\nv2 format to JGit as soon as it's settled. It sounds like a better\ninvestment of all our times.\n\nFollowing up on Shawn's comments, I removed the little-endian support\nfrom the on-disk format and implemented lazy loading of the bitmaps to\nmake up for it. The result is decent (slowed down from 250ms to 300ms)\nand it lets us keep the whole format as NWO on disk. I think it's a\ngood tradeback.\n\nThe relevant commits are available on my fork of Git (I'll be sending\nv2 of the patchset once I finish tackling the other reviews):\n\n    https://github.com/vmg/git/commit/d6cdd4329a547580bbc0143764c726c48b887271\n    https://github.com/vmg/git/commit/d8ec342fee87425e05c0db1e1630db8424612c71\n\nAs it stands right now, the only two changes from v1 of the on-disk format are:\n\n- There is an index at the end. This is a good idea.\n- The bitmaps are sorted in packfile-index order, not in packfile\norder. This is a good idea.\n\nAs always, all your feedback is appreciated, but please keep in mind I\nhave strict performance concerns.\n\nGerman kisses,\nvmg\n"},{"id":"221987","messageId":"CAFFjANQ_PoTT5bUrZ_0oARz=oZysJdMC1MAsHR2MCZVubfSbsw@mail.gmail.com","threadId":"34271","inReplyTo":"87y59xlvt7.fsf@linux-k42r.v.cablecom.net","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Vicent Martí","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-25T22:30:06Z","receivedAt":"2013-06-25T22:30:06Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"On Tue, Jun 25, 2013 at 5:58 PM, Thomas Rast <trast@inf.ethz.ch> wrote:\n>\n>> This is the technical documentation and design rationale for the new\n>> Bitmap v2 on-disk format.\n>\n> Hrmpf, that's what I get for reading the series in order...\n>\n>> +                     The folowing flags are supported:\n>                               ^^\n>\n> typos marked by ^\n>\n>> +             By storing all the hashes in a cache together with the bitmapsin\n>                                                                              ^^\n>\n>> +             The obvious consequence is that the XOR of all 4 bitmaps will result\n>> +             in a full set (all bits sets), and the AND of all 4 bitmaps will\n>                                            ^\n>\n>> +             - 1-byte XOR-offset\n>> +                     The xor offset used to compress this bitmap. For an entry\n>> +                     in position `x`, a XOR offset of `y` means that the actual\n>> +                     bitmap representing for this commit is composed by XORing the\n>> +                     bitmap for this entry with the bitmap in entry `x-y` (i.e.\n>> +                     the bitmap `y` entries before this one).\n>> +\n>> +                     Note that this compression can be recursive. In order to\n>> +                     XOR this entry with a previous one, the previous entry needs\n>> +                     to be decompressed first, and so on.\n>> +\n>> +                     The hard-limit for this offset is 160 (an entry can only be\n>> +                     xor'ed against one of the 160 entries preceding it). This\n>> +                     number is always positivea, and hence entries are always xor'ed\n>                                                  ^\n>\n>> +                     with **previous** bitmaps, not bitmaps that will come afterwards\n>> +                     in the index.\n>\n> Clever.  Why 160 though?\n\nJGit implementation detail. It's the equivalent of the delta-window in\n`pack-objects` for example.\n\nHINT HINT: in practice, JGit only looks 16 positions behind to find\ndeltas, and we do the same. So the practical limit is 16. harhar\n\n>\n>> +             - 2 bytes of RESERVED data (used right now for better packing).\n>\n> What do they mean?\n>\n>> +  With an index at the end of the file, we can load only this index in memory,\n>> +  allowing for very efficient access to all the available bitmaps lazily (we\n>> +  have their offsets in the mmaped file).\n>\n> Is there anything preventing you from mmap()ing the index also?\n\nYeah, this format allows you to easily do a SHA1 bsearch with custom\nstep to lookup entries on the bitmap index, except for the fact that\nthe index is not sorted by SHA1, so you'd need a linear search\ninstead. :)\n\nI decided against it because during most complex invocations of\n`pack-objects`, we perform a couple thousand commit lookups to see if\nthey have a bitmap in the index, so it makes a lot of sense to load\nthe index tightly in a hash table before hand (which takes very little\ntime, to be fair). We more-than-make up for the loading time by having\nmuch much faster lookups. I felt it was the right tradeoff (JGit does\nthe same, but in their case, because they cannot mmap. :p)\n\n>> +== Appendix A: Serialization format for an EWAH bitmap\n>> +\n>> +Ewah bitmaps are serialized in the protocol as the JAVAEWAH\n>> +library, making them backwards compatible with the JGit\n>> +implementation:\n>> +\n>> +     - 4-byte number of bits of the resulting UNCOMPRESSED bitmap\n>> +\n>> +     - 4-byte number of words of the COMPRESSED bitmap, when stored\n>> +\n>> +     - N x 8-byte words, as specified by the previous field\n>> +\n>> +             This is the actual content of the compressed bitmap.\n>> +\n>> +     - 4-byte position of the current RLW for the compressed\n>> +             bitmap\n>> +\n>> +Note that the byte order for this serialization is not defined by\n>> +default. The byte order for all the content in a serialized EWAH\n>> +bitmap can be known by the byte order flags in the header of the\n>> +bitmap index file.\n>\n> Please document the RLW format here.\n\nHar har. I was going to comment on your review of the Ewah patchset,\nbut might as well do it here: the only thing I know about Ewah bitmaps\nis that they work. And I know this because I did extensive fuzz\ntesting of my C port. Unfortunately, the original Java code I ported\nfrom has 0 comments, so any documentation here would have to be\nreverse-engineered.\n\nPersonally, I'd lean towards considering Ewah an external dependency\n(black box); the headers for the library are commented accordingly,\nclearly explaining the interfaces while hiding implementation details.\nOf course, you're welcome to help me reverse engineer the\nimplementation, but I'm not sure this would be of much value. It'd be\nbetter to make sure it passes the extensive test suite of the Java\nversion, and assume that Mr Lemire designed a sound format for the\nbitmaps.\n\nSwiss kisses,\nvmg\n"},{"id":"221988","messageId":"7vk3lhzu0q.fsf@alter.siamese.dyndns.org","threadId":"34271","inReplyTo":"1372116193-32762-4-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 03/16] pack-objects: use a faster hash table","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-06-25T22:48:37Z","receivedAt":"2013-06-25T22:48:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Vicent Marti <tanoku@gmail.com> writes:\n\n> @@ -901,19 +896,19 @@ static int no_try_delta(const char *path)\n>  \treturn 0;\n>  }\n>  \n> -static int add_object_entry(const unsigned char *sha1, enum object_type type,\n> -\t\t\t    const char *name, int exclude)\n> +static int add_object_entry_1(const unsigned char *sha1, enum object_type type,\n> +\t\t\t    uint32_t hash, int exclude, struct packed_git *found_pack,\n> +\t\t\t\toff_t found_offset)\n>  {\n>  \tstruct object_entry *entry;\n> -\tstruct packed_git *p, *found_pack = NULL;\n> -\toff_t found_offset = 0;\n> -\tint ix;\n> -\tunsigned hash = name_hash(name);\n> +\tstruct packed_git *p;\n> +\tkhiter_t ix;\n> +\tint hash_ret;\n>  \n> -\tix = nr_objects ? locate_object_entry_hash(sha1) : -1;\n> -\tif (ix >= 0) {\n> +\tix = kh_put_sha1(packed_objects, sha1, &hash_ret);\n> +\tif (hash_ret == 0) {\n>  \t\tif (exclude) {\n> -\t\t\tentry = objects + object_ix[ix] - 1;\n> +\t\t\tentry = kh_value(packed_objects, ix);\n>  \t\t\tif (!entry->preferred_base)\n>  \t\t\t\tnr_result--;\n>  \t\t\tentry->preferred_base = 1;\n\nAfter this, the function returns.  The original did not add to the\ntable the object name we are looking at, but the new code first adds\nit to the table with the unconditional kh_put_sha1() above.  Is a\ncall to kh_del_sha1() missing here ...\n\n> @@ -921,38 +916,42 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n>  \t\treturn 0;\n>  \t}\n>  \n> -\tif (!exclude && local && has_loose_object_nonlocal(sha1))\n> +\tif (!exclude && local && has_loose_object_nonlocal(sha1)) {\n> +\t\tkh_del_sha1(packed_objects, ix);\n>  \t\treturn 0;\n\n... like this one, which seems to compensate for \"ahh, after all we\nrealize we do not want to add this one to the table\"?\n\n> @@ -966,19 +965,30 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n>  \t\tentry->in_pack_offset = found_offset;\n>  \t}\n>  \n> -\tif (object_ix_hashsz * 3 <= nr_objects * 4)\n> -\t\trehash_objects();\n> -\telse\n> -\t\tobject_ix[-1 - ix] = nr_objects;\n> +\tkh_value(packed_objects, ix) = entry;\n> +\tkh_key(packed_objects, ix) = entry->idx.sha1;\n> +\tobjects[nr_objects++] = entry;\n>  \n>  \tdisplay_progress(progress_state, nr_objects);\n>  \n> -\tif (name && no_try_delta(name))\n> -\t\tentry->no_try_delta = 1;\n> -\n>  \treturn 1;\n>  }\n>  \n> +static int add_object_entry(const unsigned char *sha1, enum object_type type,\n> +\t\t\t    const char *name, int exclude)\n> +{\n> +\tif (add_object_entry_1(sha1, type, name_hash(name), exclude, NULL, 0)) {\n> +\t\tstruct object_entry *entry = objects[nr_objects - 1];\n> +\n> +\t\tif (name && no_try_delta(name))\n> +\t\t\tentry->no_try_delta = 1;\n> +\n> +\t\treturn 1;\n> +\t}\n> +\n> +\treturn 0;\n> +}\n\nIt is somewhat unclear what we are getting from the split of the\nmain part of this function into *_1(), other than the *_1() function\nnow has a very deep indentation inside \"if (!found_pack)\", which is\nalways true because the caller always passes NULL to found_pack.\nPerhaps this is an unrelated refactoring that is needed for later\nsteps and does not have anything to do with the use of new hash\nfunction?\n"},{"id":"221989","messageId":"7vfvw5ztvh.fsf@alter.siamese.dyndns.org","threadId":"34271","inReplyTo":"7v7ghj571o.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH 08/16] ewah: compressed bitmap implementation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-06-25T22:51:46Z","receivedAt":"2013-06-25T22:51:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Vicent Marti <tanoku@gmail.com> writes:\n>\n>> The library is re-licensed under the GPLv2 with the permission of Daniel\n>> Lemire, the original author. The source code for the C version can\n>> be found on GitHub:\n>>\n>> \thttps://github.com/vmg/libewok\n>>\n>> The original Java implementation can also be found on GitHub:\n>>\n>> \thttps://github.com/lemire/javaewah\n>> ---\n>\n> Please make sure that all patches are properly signed off.\n>\n>>  Makefile           |    6 +\n>>  ewah/bitmap.c      |  229 +++++++++++++++++\n>>  ewah/ewah_bitmap.c |  703 ++++++++++++++++++++++++++++++++++++++++++++++++++++\n>>  ewah/ewah_io.c     |  199 +++++++++++++++\n>>  ewah/ewah_rlw.c    |  124 +++++++++\n>>  ewah/ewok.h        |  194 +++++++++++++++\n>>  ewah/ewok_rlw.h    |  114 +++++++++\n>\n> This is lovely.  A few comments after an initial quick scan-through.\n>\n>  - The code and the headers are well commented, which is good.\n>\n>  - What's __builtin_popcountll() doing there in a presumably generic\n>    codepath?\n>\n>  - Two variants of \"bitmap\" are given different and easy to\n>    understand type names (vanilla one is \"bitmap\", the clever one is\n>    \"ewah_bitmap\"), but at many places, a pointer to ewah_bitmap is\n>    simply called \"bitmap\" or \"bitmap_i\" without \"ewah\" anywhere,\n>    which waas confusing to read.  Especially, the \"NAND\" operation\n>    for bitmap takes two bitmaps, while \"OR\" takes one bitmap and\n>    ewah_bitmap.  That is fine as long as the combination is\n>    convenient for callers, but I wished the ewah variables be called\n>    with \"ewah\" somewhere in their names.\n>\n>  - I compile with \"-Werror -Wdeclaration-after-statement\"; some\n>    places seem to trigger it.\n>\n>  - Some \"extern\" declarations in *.c sources were irritating;\n>    shouldn't they be declared in *.h file and included?\n>\n>  - There are some instances of \"if (condition) stmt;\" on a single\n>    line; looked irritating.   \n>\n>  - \"bool\" is not a C type we use (and not a particularly good type\n>    in C++, either).\n\nOne more.\n\n  - Use of unnecessary float (e.g. \"oldval *= 1.5\") were moderately\n    annoying.\n\n\n> That is it for now. I am looking forward to read through the users\n> of the library ;-)\n>\n> Thanks for working on this.\n"},{"id":"221990","messageId":"7vbo6tztgn.fsf@alter.siamese.dyndns.org","threadId":"34271","inReplyTo":"1372116193-32762-14-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 13/16] repack: consider bitmaps when performing repacks","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-06-25T23:00:40Z","receivedAt":"2013-06-25T23:00:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Vicent Marti <tanoku@gmail.com> writes:\n\n> @@ -156,6 +156,11 @@ do\n>  \tfullbases=\"$fullbases pack-$name\"\n>  \tchmod a-w \"$PACKTMP-$name.pack\"\n>  \tchmod a-w \"$PACKTMP-$name.idx\"\n> +\n> +\ttest -f \"$PACKTMP-$name.bitmap\" &&\n> +\tchmod a-w \"$PACKTMP-$name.bitmap\" &&\n> +\tmv -f \"$PACKTMP-$name.bitmap\" \"$PACKDIR/pack-$name.bitmap\"\n\nIf we see a temporary bitmap but somehow failed to move it to the\nfinal name, should we _ignore_ that error, or should we die, like\nthe next two lines do?\n\n>  \tmv -f \"$PACKTMP-$name.pack\" \"$PACKDIR/pack-$name.pack\" &&\n>  \tmv -f \"$PACKTMP-$name.idx\"  \"$PACKDIR/pack-$name.idx\" ||\n>  \texit\n"},{"id":"221992","messageId":"7v7ghhzt73.fsf@alter.siamese.dyndns.org","threadId":"34271","inReplyTo":"1372116193-32762-11-git-send-email-tanoku@gmail.com","subject":"Re: [PATCH 10/16] pack-objects: use bitmaps when packing objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-06-25T23:06:24Z","receivedAt":"2013-06-25T23:06:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Vicent Marti <tanoku@gmail.com> writes:\n\n> @@ -83,6 +84,9 @@ static struct progress *progress_state;\n>  static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n>  static int pack_compression_seen;\n>  \n> +static int bitmap_support;\n> +static int use_bitmap_index;\n\nOK.\n\n> @@ -2131,6 +2135,10 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n>  \t\tcache_max_small_delta_size = git_config_int(k, v);\n>  \t\treturn 0;\n>  \t}\n> +\tif (!strcmp(k, \"pack.usebitmaps\")) {\n> +\t\tbitmap_support = git_config_bool(k, v);\n> +\t\treturn 0;\n> +\t}\n\nHmph, so bitmap_support, not use_bitmap_index, keeps track of the\nuser request?  Somewhat confusing.\n\n>  \tif (!strcmp(k, \"pack.threads\")) {\n>  \t\tdelta_search_threads = git_config_int(k, v);\n>  \t\tif (delta_search_threads < 0)\n> @@ -2366,8 +2374,24 @@ static void get_object_list(int ac, const char **av)\n>  \t\t\tdie(\"bad revision '%s'\", line);\n>  \t}\n>  \n> +\tif (use_bitmap_index) {\n> +\t\tuint32_t size_hint;\n> +\n> +\t\tif (!prepare_bitmap_walk(&revs, &size_hint)) {\n> +\t\t\tkhint_t new_hash_size = (size_hint * (1.0 / __ac_HASH_UPPER)) + 0.5;\n\nWhat is __ac_HASH_UPPER?  That is a very unusual name for a variable\nor a constant.  Also it is mildly annoying to see unnecessary use of\nfloat like this.\n\n> +\t\t\tkh_resize_sha1(packed_objects, new_hash_size);\n> +\n> +\t\t\tnr_alloc = (size_hint + 63) & ~63;\n> +\t\t\tobjects = xrealloc(objects, nr_alloc * sizeof(struct object_entry *));\n> +\n> +\t\t\ttraverse_bitmap_commit_list(&add_object_entry_1);\n> +\t\t\treturn;\n> +\t\t}\n> +\t}\n> +\n>  \tif (prepare_revision_walk(&revs))\n>  \t\tdie(\"revision walk setup failed\");\n> +\n>  \tmark_edges_uninteresting(revs.commits, &revs, show_edge);\n>  \ttraverse_commit_list(&revs, show_commit, show_object, NULL);\n>  \n> @@ -2495,6 +2519,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n>  \t\t\t    N_(\"pack compression level\")),\n>  \t\tOPT_SET_INT(0, \"keep-true-parents\", &grafts_replace_parents,\n>  \t\t\t    N_(\"do not hide commits by grafts\"), 0),\n> +\t\tOPT_BOOL(0, \"bitmaps\", &bitmap_support,\n> +\t\t\t N_(\"enable support for bitmap optimizations\")),\n\nPlease match this with the name of configuration variable, i.e. --use-bitmaps\n\n>  \t\tOPT_END(),\n>  \t};\n>  \n> @@ -2561,6 +2587,11 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n>  \tif (keep_unreachable && unpack_unreachable)\n>  \t\tdie(\"--keep-unreachable and --unpack-unreachable are incompatible.\");\n>  \n> +\tif (bitmap_support) {\n> +\t\tif (use_internal_rev_list && pack_to_stdout)\n> +\t\t\tuse_bitmap_index = 1;\n\nOK, so only when some internal condition is met, the user request to\nuse bitmap is honored and the deision is kept in use_bitmap_index.\n\nIt may be easier to read if you get rid of bitmap_support, set\nuser_bitmap_index directly from the command line and config, and did\nthis here instead:\n\n\tif (!(use_internal_rev_list && pack_to_stdout))\n\t\tuse_bitmap_index = 0;\n"},{"id":"221993","messageId":"CAFFjANQWb8S4NJAGQYs2-O9abLKBCxE4M7SqWG4pB_CC1K5G4Q@mail.gmail.com","threadId":"34271","inReplyTo":"7vk3lhzu0q.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH 03/16] pack-objects: use a faster hash table","fromName":"Vicent Martí","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-25T23:09:35Z","receivedAt":"2013-06-25T23:09:35Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"On Wed, Jun 26, 2013 at 12:48 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> After this, the function returns.  The original did not add to the\n> table the object name we are looking at, but the new code first adds\n> it to the table with the unconditional kh_put_sha1() above.  Is a\n> call to kh_del_sha1() missing here ...\n\nNo, this is not the case. That's the return case for when *the object\nwas found because it already existed in the hash table* (hence we\naccess it if we're excluding it, to tag it as excluded). We don't want\nto remove it from the hash table because we're not the ones we\ninserted it.\n\nWe only call `kh_del_sha1` in the cases where:\n\n    1. The object wasn't found.\n    2. We inserted its key on the hash table.\n    3. We later learnt that we don't really want to pack this object.\n\n>\n>> @@ -921,38 +916,42 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n>>               return 0;\n>>       }\n>>\n>> -     if (!exclude && local && has_loose_object_nonlocal(sha1))\n>> +     if (!exclude && local && has_loose_object_nonlocal(sha1)) {\n>> +             kh_del_sha1(packed_objects, ix);\n>>               return 0;\n>\n> ... like this one, which seems to compensate for \"ahh, after all we\n> realize we do not want to add this one to the table\"?\n>\n>> @@ -966,19 +965,30 @@ static int add_object_entry(const unsigned char *sha1, enum object_type type,\n>>               entry->in_pack_offset = found_offset;\n>>       }\n>>\n>> -     if (object_ix_hashsz * 3 <= nr_objects * 4)\n>> -             rehash_objects();\n>> -     else\n>> -             object_ix[-1 - ix] = nr_objects;\n>> +     kh_value(packed_objects, ix) = entry;\n>> +     kh_key(packed_objects, ix) = entry->idx.sha1;\n>> +     objects[nr_objects++] = entry;\n>>\n>>       display_progress(progress_state, nr_objects);\n>>\n>> -     if (name && no_try_delta(name))\n>> -             entry->no_try_delta = 1;\n>> -\n>>       return 1;\n>>  }\n>>\n>> +static int add_object_entry(const unsigned char *sha1, enum object_type type,\n>> +                         const char *name, int exclude)\n>> +{\n>> +     if (add_object_entry_1(sha1, type, name_hash(name), exclude, NULL, 0)) {\n>> +             struct object_entry *entry = objects[nr_objects - 1];\n>> +\n>> +             if (name && no_try_delta(name))\n>> +                     entry->no_try_delta = 1;\n>> +\n>> +             return 1;\n>> +     }\n>> +\n>> +     return 0;\n>> +}\n>\n> It is somewhat unclear what we are getting from the split of the\n> main part of this function into *_1(), other than the *_1() function\n> now has a very deep indentation inside \"if (!found_pack)\", which is\n> always true because the caller always passes NULL to found_pack.\n> Perhaps this is an unrelated refactoring that is needed for later\n> steps and does not have anything to do with the use of new hash\n> function?\n\nYes, apologies for not making this clear. By refactoring into `_1`,\nyou can see how `traverse_bitmap_commit_list` can use the `_1` version\ndirectly as a callback, to insert objects straight into the packing\nlist without looking them up. This is very efficient because we can\npass the whole API straight from the bitmap code:\n\n1. The SHA1: we find it by simply looking up the `nth` sha1 on the\npack index (if we are yielding bit `n`)\n2. The object type: we find it because we have type indexes that let\nus know the type of any given bit in the bitmap by and-ing it with the\nindex.\n3. The hash for its name: we can look it up from the name hash cache\nin the new bitmap format.\n4. Exclude flag: we never exclude when working with bitmaps\n5. found_pack: all the bitmapped objects come from the same pack!\n6. found_offset: we find it by simply looking up the `nth` offset on\nthe pack index (if we are yielding bit `n`)\n\nBoom! We filled the callback just from the data in a bitmap. Ain't that nice?\n\nLet me amend the commit message.\n"},{"id":"221994","messageId":"CAFFjANQ2VD_Eir9-8u4xDD6NherudJa4xp3HN6FaTUsspnzt8g@mail.gmail.com","threadId":"34271","inReplyTo":"7v7ghhzt73.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH 10/16] pack-objects: use bitmaps when packing objects","fromName":"Vicent Martí","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-25T23:14:37Z","receivedAt":"2013-06-25T23:14:37Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"On Wed, Jun 26, 2013 at 1:06 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>> @@ -83,6 +84,9 @@ static struct progress *progress_state;\n>>  static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n>>  static int pack_compression_seen;\n>>\n>> +static int bitmap_support;\n>> +static int use_bitmap_index;\n>\n> OK.\n>\n>> @@ -2131,6 +2135,10 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n>>               cache_max_small_delta_size = git_config_int(k, v);\n>>               return 0;\n>>       }\n>> +     if (!strcmp(k, \"pack.usebitmaps\")) {\n>> +             bitmap_support = git_config_bool(k, v);\n>> +             return 0;\n>> +     }\n>\n> Hmph, so bitmap_support, not use_bitmap_index, keeps track of the\n> user request?  Somewhat confusing.\n>\n>>       if (!strcmp(k, \"pack.threads\")) {\n>>               delta_search_threads = git_config_int(k, v);\n>>               if (delta_search_threads < 0)\n>> @@ -2366,8 +2374,24 @@ static void get_object_list(int ac, const char **av)\n>>                       die(\"bad revision '%s'\", line);\n>>       }\n>>\n>> +     if (use_bitmap_index) {\n>> +             uint32_t size_hint;\n>> +\n>> +             if (!prepare_bitmap_walk(&revs, &size_hint)) {\n>> +                     khint_t new_hash_size = (size_hint * (1.0 / __ac_HASH_UPPER)) + 0.5;\n>\n> What is __ac_HASH_UPPER?  That is a very unusual name for a variable\n> or a constant.  Also it is mildly annoying to see unnecessary use of\n> float like this.\n\nSee the updated patch at:\n\nhttps://github.com/vmg/git/blob/vmg/bitmaps-master/builtin/pack-objects.c#L2422\n\n>\n>> +                     kh_resize_sha1(packed_objects, new_hash_size);\n>> +\n>> +                     nr_alloc = (size_hint + 63) & ~63;\n>> +                     objects = xrealloc(objects, nr_alloc * sizeof(struct object_entry *));\n>> +\n>> +                     traverse_bitmap_commit_list(&add_object_entry_1);\n>> +                     return;\n>> +             }\n>> +     }\n>> +\n>>       if (prepare_revision_walk(&revs))\n>>               die(\"revision walk setup failed\");\n>> +\n>>       mark_edges_uninteresting(revs.commits, &revs, show_edge);\n>>       traverse_commit_list(&revs, show_commit, show_object, NULL);\n>>\n>> @@ -2495,6 +2519,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n>>                           N_(\"pack compression level\")),\n>>               OPT_SET_INT(0, \"keep-true-parents\", &grafts_replace_parents,\n>>                           N_(\"do not hide commits by grafts\"), 0),\n>> +             OPT_BOOL(0, \"bitmaps\", &bitmap_support,\n>> +                      N_(\"enable support for bitmap optimizations\")),\n>\n> Please match this with the name of configuration variable, i.e. --use-bitmaps\n>\n>>               OPT_END(),\n>>       };\n>>\n>> @@ -2561,6 +2587,11 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n>>       if (keep_unreachable && unpack_unreachable)\n>>               die(\"--keep-unreachable and --unpack-unreachable are incompatible.\");\n>>\n>> +     if (bitmap_support) {\n>> +             if (use_internal_rev_list && pack_to_stdout)\n>> +                     use_bitmap_index = 1;\n>\n> OK, so only when some internal condition is met, the user request to\n> use bitmap is honored and the deision is kept in use_bitmap_index.\n>\n> It may be easier to read if you get rid of bitmap_support, set\n> user_bitmap_index directly from the command line and config, and did\n> this here instead:\n>\n>         if (!(use_internal_rev_list && pack_to_stdout))\n>                 use_bitmap_index = 0;\n\nYeah, I'm not particularly happy with the way these flags are\nimplemented. I'll update this.\n"},{"id":"221995","messageId":"CAFFjANQ4wbMZQO-Y++bzakpqKcD_Co4KPo8sj3i-wCC+730Sig@mail.gmail.com","threadId":"34271","inReplyTo":"7vbo6tztgn.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH 13/16] repack: consider bitmaps when performing repacks","fromName":"Vicent Martí","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-25T23:16:52Z","receivedAt":"2013-06-25T23:16:52Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"On Wed, Jun 26, 2013 at 1:00 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>> @@ -156,6 +156,11 @@ do\n>>       fullbases=\"$fullbases pack-$name\"\n>>       chmod a-w \"$PACKTMP-$name.pack\"\n>>       chmod a-w \"$PACKTMP-$name.idx\"\n>> +\n>> +     test -f \"$PACKTMP-$name.bitmap\" &&\n>> +     chmod a-w \"$PACKTMP-$name.bitmap\" &&\n>> +     mv -f \"$PACKTMP-$name.bitmap\" \"$PACKDIR/pack-$name.bitmap\"\n>\n> If we see a temporary bitmap but somehow failed to move it to the\n> final name, should we _ignore_ that error, or should we die, like\n> the next two lines do?\n\nI obviously decided against dying (as you can see on the patch, har\nhar), because the bitmap is not required for the proper operation of\nthe Git repository, unlike the packfile and the index.\n"},{"id":"221997","messageId":"CAFFjANSYoRGFDx109kMWJtYAO4TaTwSW0NCaemnrERuwakfpGg@mail.gmail.com","threadId":"34271","inReplyTo":"87mwqdlvsq.fsf@linux-k42r.v.cablecom.net","subject":"Re: [PATCH 11/16] rev-list: add bitmap mode to speed up lists","fromName":"Vicent Martí","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-26T01:45:26Z","receivedAt":"2013-06-26T01:45:26Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"I'm afraid I cannot reproduce the segfault locally (assuming you're\nperforming the rev-list on the git/git repository). Could you please\nsend me more information, and a core dump if possible?\n\nOn Tue, Jun 25, 2013 at 6:22 PM, Thomas Rast <trast@inf.ethz.ch> wrote:\n> Vicent Marti <tanoku@gmail.com> writes:\n>\n>> Calling `git rev-list --use-bitmaps [committish]` is the equivalent\n>> of `git rev-list --objects`, but the rev list is performed based on\n>> a bitmap result instead of using a manual counting objects phase.\n>\n> Why would we ever want to not --use-bitmaps, once it actually works?\n> I.e., shouldn't this be the default if pack.usebitmaps is set (or\n> possibly even core.usebitmaps for these things)?\n>\n>> These are some example timings for `torvalds/linux`:\n>>\n>>       $ time ../git/git rev-list --objects master > /dev/null\n>>\n>>       real    0m25.567s\n>>       user    0m25.148s\n>>       sys     0m0.384s\n>>\n>>       $ time ../git/git rev-list --use-bitmaps master > /dev/null\n>>\n>>       real    0m0.393s\n>>       user    0m0.356s\n>>       sys     0m0.036s\n>\n> I see your badass numbers, and raise you a critical issue:\n>\n>   $ time git rev-list --use-bitmaps --count --left-right origin/pu...origin/next\n>   Segmentation fault\n>\n>   real    0m0.408s\n>   user    0m0.383s\n>   sys     0m0.022s\n>\n> It actually seems to be related solely to having negated commits in the\n> walk:\n>\n>   thomas@linux-k42r:~/g(next u+65)$ time git rev-list --use-bitmaps --count origin/pu\n>   32315\n>\n>   real    0m0.041s\n>   user    0m0.034s\n>   sys     0m0.006s\n>   thomas@linux-k42r:~/g(next u+65)$ time git rev-list --use-bitmaps --count origin/pu ^origin/next\n>   Segmentation fault\n>\n>   real    0m0.460s\n>   user    0m0.214s\n>   sys     0m0.244s\n>\n> I also can't help noticing that the time spent generating the segfault\n> would have sufficed to generate the answer \"the old way\" as well:\n>\n>   $ time git rev-list --count --left-right origin/pu...origin/next\n>   189     125\n>\n>   real    0m0.409s\n>   user    0m0.386s\n>   sys     0m0.022s\n>\n> Can we use the same trick to speed up merge base computation and then\n> --left-right?  The latter is a component of __git_ps1 and can get\n> somewhat slow in some cases, so it would be nice to make it really fast,\n> too.\n>\n> --\n> Thomas Rast\n> trast@{inf,student}.ethz.ch\n"},{"id":"221998","messageId":"20130626021417.GB21212@sigill.intra.peff.net","threadId":"34271","inReplyTo":"87a9mdnae3.fsf@linux-k42r.v.cablecom.net","subject":"Re: [PATCH 03/16] pack-objects: use a faster hash table","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-06-26T02:14:17Z","receivedAt":"2013-06-26T02:14:17Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 25, 2013 at 04:03:22PM +0200, Thomas Rast wrote:\n\n> > The big win here, however, is in the massively reduced amount of hash\n> > collisions (as you can see from the huge reduction of time spent in\n> > `hashcmp` after the change). These greatly improved lookup times\n> > will result critical once we implement the writing algorithm for bitmap\n> > indxes in a later patch of this series.\n> \n> Is that reduction in collisions purely because it uses quadratic\n> probing, or is there some other magic trick involved?  Is the same also\n> applicable to the other users of the \"big\" object hash table?  (I assume\n> Peff has already tried applying it there, but I'm still curious...)\n\nI haven't done any actual timings yet.\n\nThe general code is quite similar to our object.c hash table, with the\nexception that it does quadratic probing.  I did try quadratic probing\non our object.c hash once and didn't see much improvement (similarly,\nJunio tried cuckoo hashing, but the numbers were not that exciting).\n\nIt's possible that the hash table in pack-objects did not behave as well\nas the one in object.c. It looks like we grow it when the table is 3/4\nfull, which is a little high (we grow at 1/2 in object.c).  Quadratic\nprobing should help when the hash table is close to full, so it would\nprobably help. However, I also note that khash keeps its hash tables\nonly half full, so that may be the real source of the performance\nimprovement.\n\nSo I suspect two things (but as I said, haven't verified):\n\n  1. You could speed up pack-objects just by keeping the table half full\n     rather than 3/4 full.\n\n  2. You would see little to no speedup by moving object.c to khash, as\n     it is adding only quadratic probing. With quadratic probing, you\n     could potentially tweak the kh_put_* to resize less aggressively\n     (say, 2/3) and save some memory without loss of performance.\n\n-Peff\n"},{"id":"222005","messageId":"20130626044752.GA26755@sigill.intra.peff.net","threadId":"34271","inReplyTo":"20130626021417.GB21212@sigill.intra.peff.net","subject":"Re: [PATCH 03/16] pack-objects: use a faster hash table","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-06-26T04:47:52Z","receivedAt":"2013-06-26T04:47:52Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 25, 2013 at 10:14:17PM -0400, Jeff King wrote:\n\n> So I suspect two things (but as I said, haven't verified):\n> \n>   1. You could speed up pack-objects just by keeping the table half full\n>      rather than 3/4 full.\n\nI wasn't able to show any measurable speedup with this. I tried to make\nas specific a measurement as I could, by adding a \"counting only\" option\nlike this:\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex fc12df8..a0438d0 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -2452,6 +2452,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tconst char *rp_av[6];\n \tint rp_ac = 0;\n \tint rev_list_unpacked = 0, rev_list_all = 0, rev_list_reflog = 0;\n+\tint counting_only = 0;\n \tstruct option pack_objects_options[] = {\n \t\tOPT_SET_INT('q', \"quiet\", &progress,\n \t\t\t    N_(\"do not show progress meter\"), 0),\n@@ -2515,6 +2516,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t    N_(\"pack compression level\")),\n \t\tOPT_SET_INT(0, \"keep-true-parents\", &grafts_replace_parents,\n \t\t\t    N_(\"do not hide commits by grafts\"), 0),\n+\t\tOPT_BOOL(0, \"counting-only\", &counting_only,\n+\t\t\t N_(\"exit after counting objects phase\")),\n \t\tOPT_END(),\n \t};\n \n@@ -2600,6 +2603,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tfor_each_ref(add_ref_tag, NULL);\n \tstop_progress(&progress_state);\n \n+\tif (counting_only)\n+\t\treturn 0;\n \tif (non_empty && !nr_result)\n \t\treturn 0;\n \tif (nr_result)\n\nand even doing the whole object traversal ahead of time to just focus on\nthe object-entry hash, like this:\n\n  git rev-list --objects --all >objects.out\n  time git pack-objects --counting-only --stdout <objects.out\n\nTweaking the hash size didn't have any effect, but using Vicent's khash\npatch actually made it about 5% slower. So I wonder if I'm even\nmeasuring the right thing. Vicent, how did you get the timings you\nshowed in the commit message?\n\n-Peff\n"},{"id":"222006","messageId":"20130626051117.GB26755@sigill.intra.peff.net","threadId":"34271","inReplyTo":"CAFFjANRwBBcORhu4mwjESBfr4GJ3zDrgYvUhY=VxK9abv7k2MA@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-06-26T05:11:17Z","receivedAt":"2013-06-26T05:11:17Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 25, 2013 at 09:33:11PM +0200, Vicent Martí wrote:\n\n> > One way we side-stepped the size inflation problem in JGit was to only\n> > use the bitmap index information when sending data on the wire to a\n> > client. Here delta reuse plays a significant factor in building the\n> > pack, and we don't have to be as accurate on matching deltas. During\n> > the equivalent of `git repack` bitmaps are not used, allowing the\n> > traditional graph enumeration algorithm to generate path hash\n> > information.\n> \n> OH BOY HERE WE GO. This is worth its own thread, lots to discuss here.\n> I think peff will have a patchset regarding this to upstream soon,\n> we'll get back to it later.\n\nWe do the same thing (only use bitmaps during on-the-wire fetches).  But\nthere a few problems with assuming delta reuse.\n\nFor us (GitHub), the foremost one is that we pack many \"forks\" of a\nrepository together into a single packfile. That means when you clone\ntorvalds/linux, an object you want may be stored in the on-disk pack\nwith a delta against an object that you are not going to get. So we have\nto throw out that delta and find a new one.\n\nI'm dealing with that by adding an option to respect \"islands\" during\npacking, where an island is a set of common objects (we split it by\nfork, since we expect those objects to be fetched together, but you\ncould use other criteria). The rule is that an object cannot delta\nagainst another object that is not in all of its islands. So everybody\ncan delta against shared history, but objects in your fork can only\ndelta against other objects in the fork.  You are guaranteed to be able\nto reuse such deltas during a full clone of a fork, and the on-disk pack\nsize does not suffer all that much (because there is usually a good\nalternate delta base within your reachable history).\n\nSo with that series, we can get good reuse for clones. But there are\nstill two cases worth considering:\n\n  1. When you fetch a subset of the commits, git marks only the edges as\n     preferred bases, and does not walk the full object graph down to\n     the roots. So any object you want that is delta'd against something\n     older will not get reused. If you have reachability bitmaps, I\n     don't think there is any reason that we cannot use the entire\n     object graph (starting at the \"have\" tips, of course) as preferred\n     bases.\n\n  2. The server is not necessarily fully packed. In an active repo, you\n     may have a large \"base\" pack with bitmaps, with several recently\n     pushed packs on top. You still need to delta the recently pushed\n     objects against the base objects.\n\nI don't have measurements on how much the deltas suffer in those two\ncases. I know they suffered quite badly for clones without the name\nhashes in our alternates repos, but that part should go away with my\npatch series.\n\n-Peff\n"},{"id":"222007","messageId":"20130626052226.GC26755@sigill.intra.peff.net","threadId":"34271","inReplyTo":"87mwqdlvsq.fsf@linux-k42r.v.cablecom.net","subject":"Re: [PATCH 11/16] rev-list: add bitmap mode to speed up lists","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-06-26T05:22:26Z","receivedAt":"2013-06-26T05:22:26Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 25, 2013 at 09:22:28AM -0700, Thomas Rast wrote:\n\n> Vicent Marti <tanoku@gmail.com> writes:\n> \n> > Calling `git rev-list --use-bitmaps [committish]` is the equivalent\n> > of `git rev-list --objects`, but the rev list is performed based on\n> > a bitmap result instead of using a manual counting objects phase.\n> \n> Why would we ever want to not --use-bitmaps, once it actually works?\n> I.e., shouldn't this be the default if pack.usebitmaps is set (or\n> possibly even core.usebitmaps for these things)?\n\nIf you are using bitmaps, you cannot produce the same output as\n\"--objects\"; the latter prints the path at which each object is found.\nIn the JGit bitmap format, we have no information at all; in Vicent's\n\"v2\", we have only a hash of that pathname.\n\n-Peff\n"},{"id":"222030","messageId":"CAFFbUKJZ1w2puKFLjPNZmMhSLo3_1kpfA1upv7K6qZV256vTyQ@mail.gmail.com","threadId":"34271","inReplyTo":"20130626051117.GB26755@sigill.intra.peff.net","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Colby Ranger","fromEmail":"cranger@google.com","sentAt":"2013-06-26T18:41:12Z","receivedAt":"2013-06-26T18:41:12Z","isPatch":true,"sender":{"key":"cranger@google.com","avatar":"https://avatars.githubusercontent.com/u/2480655?v=4"},"body":"> Pinning the bitmap index on the reverse index adds complexity (lookups\n> are two-step: first find the entry in the reverse index, and then find\n> the SHA1 in the index) and is measurably slower, in both loading and\n> lookup times. Since Git doesn't have a memory problem, it's very hard\n> to make an argument for design that is more complex and runs slower to\n> save memory.\n\nSorting by SHA1 will generate a random distribution. This will require\nyou to inflate the entire bitmap on every fetch request, in order to\ndo the \"contains\" operation.  Sorting by pack offset allows us to\ninflate only the bits we need as we are walking the graph, since they\nare usually at the start of the bitmap.\n\nWhat is the general size in bytes of the SHA1 sorted bitmaps?  If they\nare much larger, the size of the bitmap has an impact on how fast you\ncan perform bitwise operations on them, which is important for fetch\nwhen doing wants AND NOT haves.\n"},{"id":"222044","messageId":"CAFFbUK+emr44o_2EHW2Y4o5fs8Livif_5D=G=NLDzE=2MEx6NQ@mail.gmail.com","threadId":"34271","inReplyTo":"CAFFbUKJZ1w2puKFLjPNZmMhSLo3_1kpfA1upv7K6qZV256vTyQ@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Colby Ranger","fromEmail":"cranger@google.com","sentAt":"2013-06-26T22:33:00Z","receivedAt":"2013-06-26T22:33:00Z","isPatch":true,"sender":{"key":"cranger@google.com","avatar":"https://avatars.githubusercontent.com/u/2480655?v=4"},"body":">> Pinning the bitmap index on the reverse index adds complexity (lookups\n>> are two-step: first find the entry in the reverse index, and then find\n>> the SHA1 in the index) and is measurably slower, in both loading and\n>> lookup times. Since Git doesn't have a memory problem, it's very hard\n>> to make an argument for design that is more complex and runs slower to\n>> save memory.\n>\n> Sorting by SHA1 will generate a random distribution. This will require\n> you to inflate the entire bitmap on every fetch request, in order to\n> do the \"contains\" operation.  Sorting by pack offset allows us to\n> inflate only the bits we need as we are walking the graph, since they\n> are usually at the start of the bitmap.\n>\n> What is the general size in bytes of the SHA1 sorted bitmaps?  If they\n> are much larger, the size of the bitmap has an impact on how fast you\n> can perform bitwise operations on them, which is important for fetch\n> when doing wants AND NOT haves.\n\nFurthermore, JGit primarily operates on the bitmap representation,\nrarely converting bitmap id -> SHA1 during clone. When the bitmap of\nobjects to include in the output pack contains all of the objects in\nthe bitmap'd pack, we only do the translation of the bitmap ids of new\nobjects, not in the bitmap index, and it is just a lookup in an array.\nThose objects are put at the front of the stream. The rest of the\nobjects are streamed directly from the pack, with some header munging,\nsince it is guaranteed to be a fully connected pack. Most of the time\nthis works because JGit creates 2 packs during GC: a heads pack, which\nis bitmap'd, and an everything else pack.\n"},{"id":"222047","messageId":"87vc50lb4i.fsf@linux-k42r.v.cablecom.net","threadId":"34271","inReplyTo":"CAFFjANQ_PoTT5bUrZ_0oARz=oZysJdMC1MAsHR2MCZVubfSbsw@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-06-26T23:12:45Z","receivedAt":"2013-06-26T23:12:45Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Vicent Martí <tanoku@gmail.com> writes:\n\n> On Tue, Jun 25, 2013 at 5:58 PM, Thomas Rast <trast@inf.ethz.ch> wrote:\n>>\n>> Please document the RLW format here.\n>\n> Har har. I was going to comment on your review of the Ewah patchset,\n> but might as well do it here: the only thing I know about Ewah bitmaps\n> is that they work. And I know this because I did extensive fuzz\n> testing of my C port. Unfortunately, the original Java code I ported\n> from has 0 comments, so any documentation here would have to be\n> reverse-engineered.\n\nI think the below would be a reasonable documentation, to be appended\nafter your description of the EWAH format.  Maybe Colby can correct me\nif I got anything wrong.  You can basically read this off from the\nimplementation of ewah_each_bit() and the helper functions it uses.\n\n-- 8< --\nThe compressed bitmap is stored in a form of run-length encoding, as\nfollows.  It consists of a concatenation of an arbitrary number of\nchunks.  Each chunk consists of one or more 64-bit words\n\n     H  L_1  L_2  L_3 .... L_M\n\nH is called RLW (run length word).  It consists of (from lower to higher\norder bits):\n\n     - 1 bit: the repeated bit B\n\n     - 32 bits: repetition count K (unsigned)\n\n     - 31 bits: literal word count M (unsigned)\n\nThe bitstream represented by the above chunk is then:\n\n     - K repetitions of B\n\n     - The bits stored in `L_1` through `L_M`.  Within a word, bits at\n       lower order come earlier in the stream than those at higher\n       order.\n\nThe next word after `L_M` (if any) must again be a RLW, for the next\nchunk.  For efficient appending to the bitstream, the EWAH stores a\nformat to the last RLW in the stream.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"222048","messageId":"87obaslb2g.fsf@linux-k42r.v.cablecom.net","threadId":"34271","inReplyTo":"CAFFjANSYoRGFDx109kMWJtYAO4TaTwSW0NCaemnrERuwakfpGg@mail.gmail.com","subject":"Re: [PATCH 11/16] rev-list: add bitmap mode to speed up lists","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-06-26T23:13:59Z","receivedAt":"2013-06-26T23:13:59Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Vicent Martí <tanoku@gmail.com> writes:\n\n> I'm afraid I cannot reproduce the segfault locally (assuming you're\n> performing the rev-list on the git/git repository). Could you please\n> send me more information, and a core dump if possible?\n\nSure, but isn't the core dump useless if you don't have the same\nexecutable?  And since I'm building \"custom\" git, you won't have that.\n\nHere's a semi-full backtrace (I left out the spammy output in the\noutermost frames).  Some variables in #2 and #3 seem to have gone off\nthe rails.\n\n#0  0x00007ffff72b06fb in __memset_sse2 () from /lib64/libc.so.6\nNo symbol table info available.\n#1  0x000000000054c31c in bitmap_set (self=0x89c360, pos=18446744072278122040) at ewah/bitmap.c:46\n        old_size = 7666\n        block = 288230376129345656\n#2  0x00000000004e6c70 in add_to_include_set (data=0x7fffffffcd00, sha1=0x85c014 \"\\230\\062˝M\\311i\\373\\372\\317\\321\\370\\224\\017\\313\\336\\301\\213\\271\\060\", bitmap_pos=-1431429576) at pack-bitmap.c:428\n        hash_pos = 512\n#3  0x00000000004e6cd6 in should_include (commit=0x85c010, _data=0x7fffffffcd00) at pack-bitmap.c:443\n        data = 0x7fffffffcd00\n        bitmap_pos = -1431429576\n#4  0x000000000050cf1d in add_parents_to_list (revs=0x7fffffffce30, commit=0x85c010, list=0x7fffffffce30, cache_ptr=0x0) at revision.c:784\n        parent = 0x88c260\n        left_flag = 32767\n        cached_base = 0x0\n#5  0x0000000000512b66 in get_revision_1 (revs=0x7fffffffce30) at revision.c:2857\n        entry = 0x8f9ce0\n        commit = 0x85c010\n#6  0x0000000000512dcf in get_revision_internal (revs=0x7fffffffce30) at revision.c:2964\n        c = 0x0\n        l = 0x1000\n#7  0x0000000000512fe1 in get_revision (revs=0x7fffffffce30) at revision.c:3040\n        c = 0xb92608\n        reversed = 0x89c360\n#8  0x00000000004d2a24 in traverse_commit_list (revs=0x7fffffffce30, show_commit=0x4e6b72 <show_commit>, show_object=0x4e6afa <show_object>, data=0x89c360) at list-objects.c:179\n        i = -1\n        commit = 0xb92608\n        base = {\n          alloc = 4097, \n          len = 0, \n          buf = 0x87bbe0 \"\"\n        }\n#9  0x00000000004e6fa4 in find_objects (revs=0x7fffffffce30, roots=0x0, seen=0x85b760) at pack-bitmap.c:549\n        incdata = {\n          base = 0x89c360, \n          seen = 0x85b760\n        }\n        base = 0x89c360\n        needs_walk = true\n        not_mapped = 0x8f9dc0\n#10 0x00000000004e747b in prepare_bitmap_walk (revs=0x7fffffffce30, result_size=0x0) at pack-bitmap.c:679\n        i = 2\n        pending_nr = 2\n        pending_alloc = 64\n        pending_e = 0x853e10\n        wants = 0x8545b0\n        haves = 0x854820\n        wants_bitmap = 0x0\n        haves_bitmap = 0x85b760\n#11 0x0000000000474bb3 in cmd_rev_list (argc=2, argv=0x7fffffffd6e8, prefix=0x0) at builtin/rev-list.c:356\n#12 0x0000000000405820 in run_builtin (p=0x7c3ef8 <commands.20770+2040>, argc=4, argv=0x7fffffffd6e8) at git.c:291\n#13 0x00000000004059b3 in handle_internal_command (argc=4, argv=0x7fffffffd6e8) at git.c:454\n#14 0x0000000000405b87 in main (argc=4, av=0x7fffffffd6e8) at git.c:544\n\n\nThis is with a version of your series that you can find at\n\n  https://github.com/trast/git.git vm/ewah\n\nI am'd your patches on top of Junio's master at the time, except for the\nparts to the Makefile that did not apply, which I fixed up manually.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"222049","messageId":"87ehbolasl.fsf@linux-k42r.v.cablecom.net","threadId":"34271","inReplyTo":"87vc50lb4i.fsf@linux-k42r.v.cablecom.net","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-06-26T23:19:54Z","receivedAt":"2013-06-26T23:19:54Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Thomas Rast <trast@inf.ethz.ch> writes:\n\n[...]\n> The next word after `L_M` (if any) must again be a RLW, for the next\n> chunk.  For efficient appending to the bitstream, the EWAH stores a\n> format to the last RLW in the stream.\n  ^^^^^^\n\nI have no idea what Freud did there, but \"pointer\" or some such is\nprobably a saner choice.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"222053","messageId":"CAFFbUKJTh79-u-D89OAZxjCnfvyZ40DE_iwQ6e+zXSw88AY6PQ@mail.gmail.com","threadId":"34271","inReplyTo":"CAFFbUK+emr44o_2EHW2Y4o5fs8Livif_5D=G=NLDzE=2MEx6NQ@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Colby Ranger","fromEmail":"cranger@google.com","sentAt":"2013-06-27T00:53:26Z","receivedAt":"2013-06-27T00:53:26Z","isPatch":true,"sender":{"key":"cranger@google.com","avatar":"https://avatars.githubusercontent.com/u/2480655?v=4"},"body":"> +  Generating this reverse index at runtime is **not** free (around 900ms\n> +  generation time for a repository like `torvalds/linux`), and once again,\n> +  this generation time needs to happen every time `pack-objects` is\n> +  spawned.\n\nIf generating the reverse index is expensive, it is probably\nworthwhile to create a \".revidx\" or extend the \".idx\" with the\ninformation sorted by offset.\n"},{"id":"222056","messageId":"CAJo=hJtJoizQUubriTPvs2bsjvw+N82MCPvw263fUB8vv8_VVA@mail.gmail.com","threadId":"34271","inReplyTo":"CAFFjANRqZ0U5tGhgjACUtquyVKCyuHiS3CC2Xxwo0J1UJVrf=g@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-06-27T01:11:48Z","receivedAt":"2013-06-27T01:11:48Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Tue, Jun 25, 2013 at 4:08 PM, Vicent Martí <tanoku@gmail.com> wrote:\n> On Tue, Jun 25, 2013 at 11:17 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>> What case are you talking about?\n>>\n>> The n-th object must be one of these four types and can never be of\n>> more than one type at the same time, so a natural expectation from\n>> the reader is \"If you OR them together, you will get the same set\".\n>> If you say \"If you XOR them\", that forces the reader to wonder when\n>> these bitmaps ever can overlap at the same bit position.\n>\n> I guess this is just wording. I don't particularly care about the\n> distinction, but I'll change it to OR.\n\nHmm, OK. If you think XOR and OR are the same operation, I also have a\nbridge to sell you. Its in Brooklyn. Its a great value.\n\nThe correct operation is OR. Not XOR. OR. Drop the X.\n\n> It cannot be mmapped not particularly because of endianness issues,\n> but because the original format is not indexed and requires a full\n> parse of the whole index before it can be accessed programatically.\n> The wrong endianness just increases the parse time.\n\nWrong endianness has nothing to do with the parse time. Modern CPUs\ncan flip a word around very quickly. In JGit we chose to parse the\nfile at load time because its simpler than having an additional index\nsegment, and we do what you did which is to toss the object SHA-1s\ninto a hashtable for fast lookup. By the time we look for the SHA-1s\nand toss them into a hashtable we can stride through the file and find\nthe bitmap regions. Simple.\n\nIn other words, the least complex solution possible that still\nprovides good performance. I'd say we have pretty good performance.\n\n>>> and I'm going to try to make it run fast enough in that\n>>> encoding.\n>>\n>> Hmph.  Is it an option to start from what JGit does, so that people\n>> can use both JGit and your code on the same repository?\n\nI'm afraid I agree here with Junio. The JGit format is already\nshipping in JGit 3.0, Gerrit Code Review 2.6, and in heavy production\nuse for almost a year on android.googlesource.com, and Google's own\ninternal Git trees.\n\nI would prefer to see a series adding bitmap support to C Git start\nwith the existing format, make it run, taking advantage of the\noptimizations JGit uses (many of which you ignored and tried to \"fix\"\nin other ways), and then look at improving the file format itself if\nload time is still the largest low hanging fruit in upload-pack. I'm\nguessing its not. You futzed around with the object table, but JGit\nsped itself up considerably by simply not using the object table when\nthe bitmap is used. I think there are several such optimizations you\nmissed in your rush to redefine the file format.\n\n>>  And then if\n>> you do not succeed, after trying to optimize in-core processing\n>> using that on-disk format to make it fast enough, start thinking\n>> about tweaking the on-disk format?\n>\n> I'm afraid this is not an option. I have an old patchset that\n> implements JGit v1 bitmap loading (and in fact that's how I initially\n> developed these series -- by loading the bitmaps from JGit for\n> debugging), but I discarded it because it simply doesn't pan out in\n> production. ~3 seconds time to spawn `upload-pack` is not an option\n> for us. I did not develop a tweaked on-disk format out of boredom.\n\nI think your code or experiments are bogus. Even on our systems with\nJGit a cold start for the Linux kernel doesn't take 3s. And this is\nJGit where Java is slow because \"Jesus it has a lot of factories\", and\nwithout mmap'ing the file into the server's address space. Hell the\nfile has to come over the network from a remote disk array.\n\n> I could dig up the patch if you're particularly interested in\n> backwards compatibility, but since it was several times slower than\n> the current iteration, I have no interest (time, actually) to maintain\n> it, brush it up, and so on. I have already offered myself to port the\n> v2 format to JGit as soon as it's settled. It sounds like a better\n> investment of all our times.\n\nActually, I think the format you propose here is inferior to the JGit\nformat. In particular the idx-ordering means the EWAH code is useless.\nYou might as well not use the EWAH format and just store 2.6M bits per\ncommit. The idx-ordering also makes *much* harder to emit a pack file\na reasonable order for the client. Colby and I tried idx-ordering and\ndiscarded it when it didn't perform as well as the pack-ordering that\nJGit uses.\n\n> Following up on Shawn's comments, I removed the little-endian support\n> from the on-disk format and implemented lazy loading of the bitmaps to\n> make up for it. The result is decent (slowed down from 250ms to 300ms)\n> and it lets us keep the whole format as NWO on disk. I think it's a\n> good tradeback.\n\nThe maintenance burden of two endian formats in a single file is too\nhigh to justify. I'm glad to see you saw that.\n\n> As it stands right now, the only two changes from v1 of the on-disk format are:\n>\n> - There is an index at the end. This is a good idea.\n\nI don't think the index is necessary if you plan to build a hashtable\nat runtime anyway. If you mmap the file you can quickly skip over a\nbitmap and find the next SHA-1 using this thing called \"pointer\narithmetic\". I am not sure if you are familiar with the term, perhaps\nyou could search the web for it.\n\n> - The bitmaps are sorted in packfile-index order, not in packfile\n> order. This is a good idea.\n\nAs Colby and I have repeatedly tried to explain, this is not a good idea.\n\n> German kisses,\n\nStrawberry and now German kisses? What's next, Mango kisses?\n"},{"id":"222057","messageId":"CAJo=hJuH98sT5WNykxQ5JX+yKxOH-5p3CCRGa-WLAYtMGAj6oA@mail.gmail.com","threadId":"34271","inReplyTo":"20130626051117.GB26755@sigill.intra.peff.net","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-06-27T01:29:19Z","receivedAt":"2013-06-27T01:29:19Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Tue, Jun 25, 2013 at 11:11 PM, Jeff King <peff@peff.net> wrote:\n> On Tue, Jun 25, 2013 at 09:33:11PM +0200, Vicent Martí wrote:\n>\n>> > One way we side-stepped the size inflation problem in JGit was to only\n>> > use the bitmap index information when sending data on the wire to a\n>> > client. Here delta reuse plays a significant factor in building the\n>> > pack, and we don't have to be as accurate on matching deltas. During\n>> > the equivalent of `git repack` bitmaps are not used, allowing the\n>> > traditional graph enumeration algorithm to generate path hash\n>> > information.\n>>\n>> OH BOY HERE WE GO. This is worth its own thread, lots to discuss here.\n>> I think peff will have a patchset regarding this to upstream soon,\n>> we'll get back to it later.\n>\n> We do the same thing (only use bitmaps during on-the-wire fetches).  But\n> there a few problems with assuming delta reuse.\n>\n> For us (GitHub), the foremost one is that we pack many \"forks\" of a\n> repository together into a single packfile. That means when you clone\n> torvalds/linux, an object you want may be stored in the on-disk pack\n> with a delta against an object that you are not going to get. So we have\n> to throw out that delta and find a new one.\n\nGerrit Code Review ran into the same problem a few years ago with the\nrefs/changes namespace. Objects reachable from a branch were often\ndelta compressed against dropped code review revisions, making for\nsome slow transfers. We fixed this by creating a pack of everything\nreachable from refs/heads/* and then another pack of the other stuff.\n\nI would encourage you to do what you suggest...\n\n> I'm dealing with that by adding an option to respect \"islands\" during\n> packing, where an island is a set of common objects (we split it by\n> fork, since we expect those objects to be fetched together, but you\n> could use other criteria). The rule is that an object cannot delta\n> against another object that is not in all of its islands. So everybody\n> can delta against shared history, but objects in your fork can only\n> delta against other objects in the fork.  You are guaranteed to be able\n> to reuse such deltas during a full clone of a fork, and the on-disk pack\n> size does not suffer all that much (because there is usually a good\n> alternate delta base within your reachable history).\n\nYes, exactly. I want to do the same thing on our servers, as we have\nmany forks of some popular open source repositories that are also not\nsmall (Linux kernel, WebKit). Unfortunately Google has not had the\ntime to develop the necessary support into JGit.\n\n> So with that series, we can get good reuse for clones. But there are\n> still two cases worth considering:\n>\n>   1. When you fetch a subset of the commits, git marks only the edges as\n>      preferred bases, and does not walk the full object graph down to\n>      the roots. So any object you want that is delta'd against something\n>      older will not get reused. If you have reachability bitmaps, I\n>      don't think there is any reason that we cannot use the entire\n>      object graph (starting at the \"have\" tips, of course) as preferred\n>      bases.\n\nIn JGit we use the reachability bitmap to provide proof a client has\nan object. Even if its not in the edges. This allows us much better\ndelta reuse, as often frequently deltas will be available pointing to\nsomething behind the edge, but that the client certainly has given the\nedges we know about.\n\nWe also use the reachability bitmap to provide proof a client does not\nneed an object. We found a reduction in number of objects transferred\nbecause the \"want AND NOT have\" subtracted out a number of objects not\nin the edge. Apparently merges, reverts and cherry-picks happen often\nenough in the repositories we host that this particular optimization\nhelps reduce data transfer, and work at both server and client ends of\nthe connection. Its a nice freebie the bitmap algorithm gives us.\n\n>   2. The server is not necessarily fully packed. In an active repo, you\n>      may have a large \"base\" pack with bitmaps, with several recently\n>      pushed packs on top. You still need to delta the recently pushed\n>      objects against the base objects.\n\nYes, this is unfortunate. One way we avoid this in JGit is to keep\neverything in pack files, rather than exploding loose. The\nreachability bitmap often proves the client has the delta base the\npusher used to make the object, allowing us to reuse the delta. It may\nnot be the absolute best delta in the world, but reuse is faster than\ninflate()+delta()+deflate(), and the delta is probably \"good enough\"\nuntil the server can do a real GC in the background.\n\nWe combine small packs from pushes together by almost literally just\nconcat'ing the packs together and creating a new .idx. Newer pushed\ndata is put in front of the older data, the pack is clustered by\n\"commit, tree, blob\" ordering, duplicates are removed, and its written\nback to disk. Typically we complete this \"pack concat\" operation mere\nseconds after a push finishes, so readers have very few packs to deal\nwith.\n\n> I don't have measurements on how much the deltas suffer in those two\n> cases. I know they suffered quite badly for clones without the name\n> hashes in our alternates repos, but that part should go away with my\n> patch series.\n\nJGit doesn't poke objects into the object table (or even the object\nlist) when a bitmap is used. We spool the bits out of the bitmap in\nbitmap order and write them to the wire in that order. Its way faster,\nbut depends on the bitmap being in pack-ordering. So clones are crazy\nfast even though we don't have the path-hash table.\n\nJGit also has another optimization where we figure out based on the\nbitmap if the client needs *everything* in this pack. Which given a\npack created only for refs/heads/* is the common case for a clone. If\nthe client is getting all objects we essentially just do a sendfile()\nfor the region starting at offset 12 through end-20. It can't be a\nsendfile() syscall because it has to be computed into the trailer\nSHA-1 the client sees, but its a crazy tight IO copy loop with no Git\nsmarts beyond the SHA-1 updating.\n\nLike I said, there are a ton of optimizations you guys missed. And we\nthink they make a bigger difference than screwing around with\nlittle-endian format to favor x86 CPUs.\n"},{"id":"222058","messageId":"CAJo=hJuA_4zL-cyygqPkDA+6ipbTqsmMYHw-c6OTKCcb5GAs0g@mail.gmail.com","threadId":"34271","inReplyTo":"CAFFbUKJTh79-u-D89OAZxjCnfvyZ40DE_iwQ6e+zXSw88AY6PQ@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-06-27T01:32:21Z","receivedAt":"2013-06-27T01:32:21Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Wed, Jun 26, 2013 at 6:53 PM, Colby Ranger <cranger@google.com> wrote:\n>> +  Generating this reverse index at runtime is **not** free (around 900ms\n>> +  generation time for a repository like `torvalds/linux`), and once again,\n>> +  this generation time needs to happen every time `pack-objects` is\n>> +  spawned.\n\n900ms is fishy. Creating the revidx should not take that long. But if it is...\n\n> If generating the reverse index is expensive, it is probably\n> worthwhile to create a \".revidx\" or extend the \".idx\" with the\n> information sorted by offset.\n\nColby is probably right that a cached copy of the revidx would help.\nOr update the .idx format to have two additional sections that stores\nthe length of each packed object and the delta base of each packed\nobject, allowing pack-objects to avoid creating the revidx. This would\nbe an additional ~8 bytes per object, so ~19.8M for the Linux kernel\n(given ~2.6M objects).\n"},{"id":"222061","messageId":"CAFFjANSr2QRLE8DSPP2zZ_baEZUqR8dzkPzMwqyEqgFX=8cnog@mail.gmail.com","threadId":"34271","inReplyTo":"CAJo=hJtJoizQUubriTPvs2bsjvw+N82MCPvw263fUB8vv8_VVA@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Vicent Martí","fromEmail":"tanoku@gmail.com","sentAt":"2013-06-27T02:36:54Z","receivedAt":"2013-06-27T02:36:54Z","isPatch":true,"sender":{"key":"tanoku@gmail.com","avatar":"https://gravatar.com/avatar/271386991cb4c2b8f1e1ed1d059f3422cc3485de7a598f65043f70be021d095b?d=mp&s=160"},"body":"That was a very rude reply. :(\n\nPlease refrain from interacting with me in the ML in the future. I'l\ndo accordingly.\n\nThanks!\nvmg\n\nOn Thu, Jun 27, 2013 at 3:11 AM, Shawn Pearce <spearce@spearce.org> wrote:\n> On Tue, Jun 25, 2013 at 4:08 PM, Vicent Martí <tanoku@gmail.com> wrote:\n>> On Tue, Jun 25, 2013 at 11:17 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>>> What case are you talking about?\n>>>\n>>> The n-th object must be one of these four types and can never be of\n>>> more than one type at the same time, so a natural expectation from\n>>> the reader is \"If you OR them together, you will get the same set\".\n>>> If you say \"If you XOR them\", that forces the reader to wonder when\n>>> these bitmaps ever can overlap at the same bit position.\n>>\n>> I guess this is just wording. I don't particularly care about the\n>> distinction, but I'll change it to OR.\n>\n> Hmm, OK. If you think XOR and OR are the same operation, I also have a\n> bridge to sell you. Its in Brooklyn. Its a great value.\n>\n> The correct operation is OR. Not XOR. OR. Drop the X.\n>\n>> It cannot be mmapped not particularly because of endianness issues,\n>> but because the original format is not indexed and requires a full\n>> parse of the whole index before it can be accessed programatically.\n>> The wrong endianness just increases the parse time.\n>\n> Wrong endianness has nothing to do with the parse time. Modern CPUs\n> can flip a word around very quickly. In JGit we chose to parse the\n> file at load time because its simpler than having an additional index\n> segment, and we do what you did which is to toss the object SHA-1s\n> into a hashtable for fast lookup. By the time we look for the SHA-1s\n> and toss them into a hashtable we can stride through the file and find\n> the bitmap regions. Simple.\n>\n> In other words, the least complex solution possible that still\n> provides good performance. I'd say we have pretty good performance.\n>\n>>>> and I'm going to try to make it run fast enough in that\n>>>> encoding.\n>>>\n>>> Hmph.  Is it an option to start from what JGit does, so that people\n>>> can use both JGit and your code on the same repository?\n>\n> I'm afraid I agree here with Junio. The JGit format is already\n> shipping in JGit 3.0, Gerrit Code Review 2.6, and in heavy production\n> use for almost a year on android.googlesource.com, and Google's own\n> internal Git trees.\n>\n> I would prefer to see a series adding bitmap support to C Git start\n> with the existing format, make it run, taking advantage of the\n> optimizations JGit uses (many of which you ignored and tried to \"fix\"\n> in other ways), and then look at improving the file format itself if\n> load time is still the largest low hanging fruit in upload-pack. I'm\n> guessing its not. You futzed around with the object table, but JGit\n> sped itself up considerably by simply not using the object table when\n> the bitmap is used. I think there are several such optimizations you\n> missed in your rush to redefine the file format.\n>\n>>>  And then if\n>>> you do not succeed, after trying to optimize in-core processing\n>>> using that on-disk format to make it fast enough, start thinking\n>>> about tweaking the on-disk format?\n>>\n>> I'm afraid this is not an option. I have an old patchset that\n>> implements JGit v1 bitmap loading (and in fact that's how I initially\n>> developed these series -- by loading the bitmaps from JGit for\n>> debugging), but I discarded it because it simply doesn't pan out in\n>> production. ~3 seconds time to spawn `upload-pack` is not an option\n>> for us. I did not develop a tweaked on-disk format out of boredom.\n>\n> I think your code or experiments are bogus. Even on our systems with\n> JGit a cold start for the Linux kernel doesn't take 3s. And this is\n> JGit where Java is slow because \"Jesus it has a lot of factories\", and\n> without mmap'ing the file into the server's address space. Hell the\n> file has to come over the network from a remote disk array.\n>\n>> I could dig up the patch if you're particularly interested in\n>> backwards compatibility, but since it was several times slower than\n>> the current iteration, I have no interest (time, actually) to maintain\n>> it, brush it up, and so on. I have already offered myself to port the\n>> v2 format to JGit as soon as it's settled. It sounds like a better\n>> investment of all our times.\n>\n> Actually, I think the format you propose here is inferior to the JGit\n> format. In particular the idx-ordering means the EWAH code is useless.\n> You might as well not use the EWAH format and just store 2.6M bits per\n> commit. The idx-ordering also makes *much* harder to emit a pack file\n> a reasonable order for the client. Colby and I tried idx-ordering and\n> discarded it when it didn't perform as well as the pack-ordering that\n> JGit uses.\n>\n>> Following up on Shawn's comments, I removed the little-endian support\n>> from the on-disk format and implemented lazy loading of the bitmaps to\n>> make up for it. The result is decent (slowed down from 250ms to 300ms)\n>> and it lets us keep the whole format as NWO on disk. I think it's a\n>> good tradeback.\n>\n> The maintenance burden of two endian formats in a single file is too\n> high to justify. I'm glad to see you saw that.\n>\n>> As it stands right now, the only two changes from v1 of the on-disk format are:\n>>\n>> - There is an index at the end. This is a good idea.\n>\n> I don't think the index is necessary if you plan to build a hashtable\n> at runtime anyway. If you mmap the file you can quickly skip over a\n> bitmap and find the next SHA-1 using this thing called \"pointer\n> arithmetic\". I am not sure if you are familiar with the term, perhaps\n> you could search the web for it.\n>\n>> - The bitmaps are sorted in packfile-index order, not in packfile\n>> order. This is a good idea.\n>\n> As Colby and I have repeatedly tried to explain, this is not a good idea.\n>\n>> German kisses,\n>\n> Strawberry and now German kisses? What's next, Mango kisses?\n"},{"id":"222062","messageId":"20130627024521.GA6936@sigill.intra.peff.net","threadId":"34271","inReplyTo":"CAFFjANSr2QRLE8DSPP2zZ_baEZUqR8dzkPzMwqyEqgFX=8cnog@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-06-27T02:45:22Z","receivedAt":"2013-06-27T02:45:22Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jun 27, 2013 at 04:36:54AM +0200, Vicent Martí wrote:\n\n> That was a very rude reply. :(\n> \n> Please refrain from interacting with me in the ML in the future. I'l\n> do accordingly.\n\nI agree that the pointer arithmetic thing may have been a little much,\nbut I think there are some points we need to address in Shawn's email.\n\nIn particular, it seems like the slowness we saw with the v1 bitmap\nformat is not what Shawn and Colby have experienced. So it's possible\nthat our test setup is bad or different. Or maybe the C v1 reading\nimplementation had some problems that are fixable. It's hard to say\nbecause we haven't shown any code that can be timed and compared.\n\nAnd the pack-order versus idx-order for the bitmaps is still up in the\nair. Do we have numbers on the on-disk sizes of the resulting EWAHs? The\npack-order ones should be more amenable to run-length encoding,\nespecially as you get further down into history (the tip ones would\nmostly be 1's, no matter how you order them).\n\n-Peff\n"},{"id":"222067","messageId":"alpine.DEB.2.00.1306270655120.10263@ds9.cixit.se","threadId":"34271","inReplyTo":"CAFFjANSNagvDgvrFNV1OLg=-4BPyQVjMDnfMPihdhVJR7o0TdQ@mail.gmail.com","subject":"Re: [PATCH 07/16] compat: add endinanness helpers","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2013-06-27T05:56:58Z","receivedAt":"2013-06-27T05:56:58Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Vicent Martí:\n\n> I'm aware of that, but Git needs to build with glibc 2.7+ (or was it \n> 2.6?), hence the need for this compat layer.\n\nRight. But perhaps the compatibility layer could provide the \nfunctionality with the names available in the later glibc versions \n(and on *BSD)? That would make it easier to read the code that is \nusing it.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"222091","messageId":"CAJo=hJvOq=CATrDeYAwi+jgkPpqjywWhuKeC1TVYeCXr6NVM6w@mail.gmail.com","threadId":"34271","inReplyTo":"20130627024521.GA6936@sigill.intra.peff.net","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-06-27T16:07:38Z","receivedAt":"2013-06-27T16:07:38Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Wed, Jun 26, 2013 at 7:45 PM, Jeff King <peff@peff.net> wrote:\n>\n> In particular, it seems like the slowness we saw with the v1 bitmap\n> format is not what Shawn and Colby have experienced. So it's possible\n> that our test setup is bad or different. Or maybe the C v1 reading\n> implementation had some problems that are fixable. It's hard to say\n> because we haven't shown any code that can be timed and compared.\n\nRight, the format and implementation in JGit can do \"Counting objects\"\nin 87ms for the Linux kernel history. But I think we are comparing\napples to steaks here, Vincent is (rightfully) concerned about process\nstartup performance, whereas our timings were assuming the process was\nalready running.\n\nIt would help everyone to understand the issues involved if we are at\nleast looking at the same file format.\n\n> And the pack-order versus idx-order for the bitmaps is still up in the\n> air. Do we have numbers on the on-disk sizes of the resulting EWAHs?\n\nI did not see any presented in this thread, and I am very interested\nin this aspect of the series. The path hash cache should be taking\nabout 9.9M of disk space, but I recall reading the bitmap file is 8M.\nI don't understand.\n\nColby and I were very concerned about the size of the EWAH compressed\nbitmaps because we wanted hundreds of them for a large history like\nthe kernel, and we wanted to minimize the amount of memory consumed by\nthe bitmap index when loaded into the process.\n\nIn the JGit implementation our copy of Linus' tree has 3.1M objects,\nan 81.5 MiB idx file, and a 3.8 MiB bitmap file. We were trying to\nkeep the overhead below 10% of the idx file, and I think we have\nsucceeded on that. With 3.1M objects the v2 bitmap proposed in this\nthread needs at least 11.8M, or 14+% overhead just for the path hash\ncache.\n\nThe path hash cache may still be required, Colby and I have been\ndebating the merits of having the data available for delta compression\nvs. the increase in memory required to hold it.\n\n> The\n> pack-order ones should be more amenable to run-length encoding,\n> especially as you get further down into history (the tip ones would\n> mostly be 1's, no matter how you order them).\n\nThis is also true for the type bitmaps, but especially so in the JGit\nfile ordering where we always write all trees before any blobs. The\ntype bitmaps are very compact and basically amount to defining a\nsingle range in the file. This takes only a few words in the EWAH\ncompressed format.\n"},{"id":"222100","messageId":"20130627171733.GA17601@sigill.intra.peff.net","threadId":"34271","inReplyTo":"CAJo=hJvOq=CATrDeYAwi+jgkPpqjywWhuKeC1TVYeCXr6NVM6w@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-06-27T17:17:33Z","receivedAt":"2013-06-27T17:17:33Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jun 27, 2013 at 09:07:38AM -0700, Shawn O. Pearce wrote:\n\n> > And the pack-order versus idx-order for the bitmaps is still up in the\n> > air. Do we have numbers on the on-disk sizes of the resulting EWAHs?\n> \n> I did not see any presented in this thread, and I am very interested\n> in this aspect of the series. The path hash cache should be taking\n> about 9.9M of disk space, but I recall reading the bitmap file is 8M.\n> I don't understand.\n\nI don't know there the 8M number came from, or if it was on the kernel\nrepo. My bitmap-enabled pack of linux-2.6 (about 3.2M objects) using\nVicent's patches looks like:\n\n  $ du -sh *\n  42M     pack-9ea76831aec6c49c5ff42509a2a2ce97da13c5ad.bitmap\n  87M     pack-9ea76831aec6c49c5ff42509a2a2ce97da13c5ad.idx\n  630M    pack-9ea76831aec6c49c5ff42509a2a2ce97da13c5ad.pack\n\nPacking the same repo with \"jgit debug-gc\" (jgit 3.0.0) yields:\n\n  $ du -sh *\n  3.0M    pack-2478783825733a1f1012f0087a0b5a92aa7437d8.bitmap\n  82M     pack-2478783825733a1f1012f0087a0b5a92aa7437d8.idx\n  585M    pack-2478783825733a1f1012f0087a0b5a92aa7437d8.pack\n  4.8M    pack-f61fb76112372288923be7a0464476892dfebe3e.idx\n  97M     pack-f61fb76112372288923be7a0464476892dfebe3e.pack\n\nIf we assume that 12M of that is name-hash, that's still an order of\nmagnitude larger. For reference, jgit created 327 bitmaps (according to\nits progress eye candy), and Vicent's patches generated 385. So that\nexplains some of the increase, but the per-bitmap size is still much\nlarger.\n\n> The path hash cache may still be required, Colby and I have been\n> debating the merits of having the data available for delta compression\n> vs. the increase in memory required to hold it.\n\nI guess this is not an option for JGit, but for C git, an mmap-able\nname-hash file means we can just fault in the pages mentioning objects\nwe actually need it for. And its use can be completely optional; in\nfact, it doesn't even need to be inside the .bitmap file (though I\ncannot think of a reason it would be useful outside of having bitmaps).\n\n-Peff\n"},{"id":"222293","messageId":"CAFFbUKKm89n0HG6xUhYMLs_yjRJ8n0jFtOEEN=vXxJfWKLx5FA@mail.gmail.com","threadId":"34271","inReplyTo":"CAJo=hJvOq=CATrDeYAwi+jgkPpqjywWhuKeC1TVYeCXr6NVM6w@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Colby Ranger","fromEmail":"cranger@google.com","sentAt":"2013-07-01T18:47:32Z","receivedAt":"2013-07-01T18:47:32Z","isPatch":true,"sender":{"key":"cranger@google.com","avatar":"https://avatars.githubusercontent.com/u/2480655?v=4"},"body":"> Right, the format and implementation in JGit can do \"Counting objects\"\n> in 87ms for the Linux kernel history.\n\nActually, that was the timing when I first pushed the change. With the\nimprovements submitted throughout the year, we can do counting in\n50ms, on my same machine.\n\n> But I think we are comparing\n> apples to steaks here, Vincent is (rightfully) concerned about process\n> startup performance, whereas our timings were assuming the process was\n> already running.\n>\n\nI did some timing on loading the reverse index for the kernel and it\nis pretty slow (~1200ms). I just submitted a fix to do a bucket sort\nand reduced that to ~450ms, which is still slow but much better:\nhttps://eclipse.googlesource.com/jgit/jgit/+/6cc532a43cf28403cb623d3df8600a2542a40a43%5E%21/\n"},{"id":"222297","messageId":"CAJo=hJv=rb2S1DqERQ9GJRQkix=Z5pG2jEyghVDcCHPSF7AKDA@mail.gmail.com","threadId":"34271","inReplyTo":"CAFFbUKKm89n0HG6xUhYMLs_yjRJ8n0jFtOEEN=vXxJfWKLx5FA@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-07-01T19:13:25Z","receivedAt":"2013-07-01T19:13:25Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Mon, Jul 1, 2013 at 11:47 AM, Colby Ranger <cranger@google.com> wrote:\n>> But I think we are comparing\n>> apples to steaks here, Vincent is (rightfully) concerned about process\n>> startup performance, whereas our timings were assuming the process was\n>> already running.\n>>\n>\n> I did some timing on loading the reverse index for the kernel and it\n> is pretty slow (~1200ms). I just submitted a fix to do a bucket sort\n> and reduced that to ~450ms, which is still slow but much better:\n> https://eclipse.googlesource.com/jgit/jgit/+/6cc532a43cf28403cb623d3df8600a2542a40a43%5E%21/\n\nA reverse index that is hot in RAM would obviously load in about 0ms.\nBut a cold load of a reverse index that uses only 4 bytes per object\n(as Colby did here) for 3.1M objects could take ~590ms to read from\ndisk, assuming spinning media moving 20 MiB/s. If 8 byte offsets were\nalso stored this could be more like 1700ms.\n\nNumbers obviously get better if the spinning media can transfer at 40\nMiB/s, now its more like 295ms for 4 bytes/object and 885ms for 12\nbytes/object.\n\nI think its still reasonable to compute the reverse index on the fly.\nBut JGit certainly does have the benefit of reusing it across requests\nby relying on process memory based caches. C Git needs to rely on the\nkernel buffer cache, which requires this data be written out to a file\nto be shared.\n"},{"id":"222729","messageId":"20130707094646.GA18120@sigill.intra.peff.net","threadId":"34271","inReplyTo":"CAFFbUKKm89n0HG6xUhYMLs_yjRJ8n0jFtOEEN=vXxJfWKLx5FA@mail.gmail.com","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-07-07T09:46:46Z","receivedAt":"2013-07-07T09:46:46Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jul 01, 2013 at 11:47:32AM -0700, Colby Ranger wrote:\n\n> > But I think we are comparing\n> > apples to steaks here, Vincent is (rightfully) concerned about process\n> > startup performance, whereas our timings were assuming the process was\n> > already running.\n> >\n> \n> I did some timing on loading the reverse index for the kernel and it\n> is pretty slow (~1200ms). I just submitted a fix to do a bucket sort\n> and reduced that to ~450ms, which is still slow but much better:\n\nOn my machine, loading the kernel revidx in C git is about ~830ms. I\nswitched the qsort() call to a radix/bucket sort, and have it down to\n~200ms. So definitely much better, though that still leaves a bit to be\ndesired for quick commands. E.g., \"git rev-list --count A..B\" should\nbecome fairly instantaneous with bitmaps, but in many cases the revindex\nloading will take longer than it would have to simply do the actual\ntraversal.\n\n-Peff\n"},{"id":"222750","messageId":"CAJo=hJsUr1osy5rDZja1WBAgzczY5Yp0CJuLoAQs7J59PsX7Zw@mail.gmail.com","threadId":"34271","inReplyTo":"20130707094646.GA18120@sigill.intra.peff.net","subject":"Re: [PATCH 09/16] documentation: add documentation for the bitmap format","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-07-07T17:27:01Z","receivedAt":"2013-07-07T17:27:01Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Sun, Jul 7, 2013 at 2:46 AM, Jeff King <peff@peff.net> wrote:\n> On Mon, Jul 01, 2013 at 11:47:32AM -0700, Colby Ranger wrote:\n>\n>> > But I think we are comparing\n>> > apples to steaks here, Vincent is (rightfully) concerned about process\n>> > startup performance, whereas our timings were assuming the process was\n>> > already running.\n>> >\n>>\n>> I did some timing on loading the reverse index for the kernel and it\n>> is pretty slow (~1200ms). I just submitted a fix to do a bucket sort\n>> and reduced that to ~450ms, which is still slow but much better:\n>\n> On my machine, loading the kernel revidx in C git is about ~830ms. I\n> switched the qsort() call to a radix/bucket sort, and have it down to\n> ~200ms. So definitely much better,\n\nThis is a very nice reduction. pack-objects would benefit from it even\nwithout bitmaps. Since it doesn't require a data format change this is\na pretty harmless patch to include in Git. We may later conclude\ncaching the revidx is worthwhile, but until then a bucket sort doesn't\nhurt. :-)\n\n> though that still leaves a bit to be\n> desired for quick commands. E.g., \"git rev-list --count A..B\" should\n> become fairly instantaneous with bitmaps, but in many cases the revindex\n> loading will take longer than it would have to simply do the actual\n> traversal.\n\nYea, we don't know of a way around this. In a few cases the bitmap\ncode in JGit is slower than the naive traversal, but these are only on\nsmall segments of history. I wonder if you could guess which algorithm\nto use by looking at the offsets of A and B using the idx file. If\nthey are near each other in the pack, run the naive algorithm without\nbitmaps and revidx. If they are farther apart assume the bitmap would\nhelp more than traversal and use bitmap+revidx.\n\nWorking out what the correct \"distance\" should be before switching\nalgorithms is hard. A and B could be megabytes apart in the pack but A\ncould be B's grandparent and traversed in milliseconds. I wonder how\noften that is in practice, certainly if A and B are within a few\nhundred kilobytes of each other the naive traversal should be almost\ninstant.\n"}]}