{"thread":{"id":"34852","subject":"[PATCH 01/38] pack v4: initial pack dictionary structure and code","startedAt":"2013-09-05T06:19:23Z","lastAt":"2013-09-12T15:34:14Z","messageCount":124,"participants":["Nicolas Pitre","SZEDER Gábor","Duy Nguyen","Junio C Hamano","Nguyễn Thái Ngọc Duy"],"isPatch":true,"patchVersion":1,"patchTotal":38},"messages":[{"id":"226821","messageId":"1378362001-1738-1-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":null,"subject":"[PATCH 00/38] pack version 4 basic functionalities","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:23Z","receivedAt":"2013-09-05T06:19:23Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"After the initial posting here:\n\n  http://news.gmane.org/group/gmane.comp.version-control.git/thread=233061\n\nThis is a repost plus the basic read side working, at least to validate\nthe write side and the pack format itself.  And many many bug fixes.\n\nThis can also be fetched here:\n\n  git://git.linaro.org/people/nico/git\n\nI consider the actual pack format definition final as implemented\nby this code.\n\nTODO:\n\n- index-pack support\n\n- native tree walk support\n\n- native commit graph walk support\n\n- better heuristics when creating tree delta encoding\n\n- integration with pack-objects\n\n- transfer protocol backward compatibility\n\n- thin pack completion\n\n- figure out unexplained runtime performance issues\n\nHowever, as I mentioned already, I've put more time on this project lately\nthan I actually had available.  I really wanted to bring this project far\nenough to be able to kick it out the door for others to take over, and\nthere we are.\n\nI'm always available for design discussions and code review.  But don't\nexpect much additional code from me at this point.\n\n@junio: I'm hoping you can take this branch as is, and apply any ffurther\npatches on top.\n\nThe diffstat goes like this:\n\n Makefile        |    3 +\n cache.h         |   11 +\n hex.c           |   11 +\n pack-check.c    |    4 +-\n pack-revindex.c |    7 +-\n pack-write.c    |    6 +-\n packv4-create.c | 1105 +++++++++++++++++++++++++++++++++++++++++++++++++\n packv4-parse.c  |  408 ++++++++++++++++++\n packv4-parse.h  |    9 +\n sha1_file.c     |  110 ++++-\n 10 files changed, 1648 insertions(+), 26 deletions(-)\n\nEnjoy !\n"},{"id":"226785","messageId":"1378362001-1738-2-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 01/38] pack v4: initial pack dictionary structure and code","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:24Z","receivedAt":"2013-09-05T06:19:24Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 137 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 137 insertions(+)\n create mode 100644 packv4-create.c\n\ndiff --git a/packv4-create.c b/packv4-create.c\nnew file mode 100644\nindex 0000000..2de6d41\n--- /dev/null\n+++ b/packv4-create.c\n@@ -0,0 +1,137 @@\n+/*\n+ * packv4-create.c: management of dictionary tables used in pack v4\n+ *\n+ * (C) Nicolas Pitre <nico@fluxnic.net>\n+ *\n+ * This code is free software; you can redistribute it and/or modify\n+ * it under the terms of the GNU General Public License version 2 as\n+ * published by the Free Software Foundation.\n+ */\n+\n+#include \"cache.h\"\n+\n+struct data_entry {\n+\tunsigned offset;\n+\tunsigned hits;\n+};\n+\n+struct dict_table {\n+\tchar *data;\n+\tunsigned ptr;\n+\tunsigned size;\n+\tstruct data_entry *entry;\n+\tunsigned nb_entries;\n+\tunsigned max_entries;\n+\tunsigned *hash;\n+\tunsigned hash_size;\n+};\n+\n+struct dict_table *create_dict_table(void)\n+{\n+\treturn xcalloc(sizeof(struct dict_table), 1);\n+}\n+\n+void destroy_dict_table(struct dict_table *t)\n+{\n+\tfree(t->data);\n+\tfree(t->entry);\n+\tfree(t->hash);\n+\tfree(t);\n+}\n+\n+static int locate_entry(struct dict_table *t, const char *str)\n+{\n+\tint i = 0;\n+\tconst unsigned char *s = (const unsigned char *) str;\n+\n+\twhile (*s)\n+\t\ti = i * 111 + *s++;\n+\ti = (unsigned)i % t->hash_size;\n+\n+\twhile (t->hash[i]) {\n+\t\tunsigned n = t->hash[i] - 1;\n+\t\tif (!strcmp(str, t->data + t->entry[n].offset))\n+\t\t\treturn n;\n+\t\tif (++i >= t->hash_size)\n+\t\t\ti = 0;\n+\t}\n+\treturn -1 - i;\n+}\n+\n+static void rehash_entries(struct dict_table *t)\n+{\n+\tunsigned n;\n+\n+\tt->hash_size *= 2;\n+\tif (t->hash_size < 1024)\n+\t\tt->hash_size = 1024;\n+\tt->hash = xrealloc(t->hash, t->hash_size * sizeof(*t->hash));\n+\tmemset(t->hash, 0, t->hash_size * sizeof(*t->hash));\n+\n+\tfor (n = 0; n < t->nb_entries; n++) {\n+\t\tint i = locate_entry(t, t->data + t->entry[n].offset);\n+\t\tif (i < 0)\n+\t\t\tt->hash[-1 - i] = n + 1;\n+\t}\n+}\n+\n+int dict_add_entry(struct dict_table *t, const char *str)\n+{\n+\tint i, len = strlen(str) + 1;\n+\n+\tif (t->ptr + len >= t->size) {\n+\t\tt->size = (t->size + len + 1024) * 3 / 2;\n+\t\tt->data = xrealloc(t->data, t->size);\n+\t}\n+\tmemcpy(t->data + t->ptr, str, len);\n+\n+\ti = (t->nb_entries) ? locate_entry(t, t->data + t->ptr) : -1;\n+\tif (i >= 0) {\n+\t\tt->entry[i].hits++;\n+\t\treturn i;\n+\t}\n+\n+\tif (t->nb_entries >= t->max_entries) {\n+\t\tt->max_entries = (t->max_entries + 1024) * 3 / 2;\n+\t\tt->entry = xrealloc(t->entry, t->max_entries * sizeof(*t->entry));\n+\t}\n+\tt->entry[t->nb_entries].offset = t->ptr;\n+\tt->entry[t->nb_entries].hits = 1;\n+\tt->ptr += len + 1;\n+\tt->nb_entries++;\n+\n+\tif (t->hash_size * 3 <= t->nb_entries * 4)\n+\t\trehash_entries(t);\n+\telse\n+\t\tt->hash[-1 - i] = t->nb_entries;\n+\n+\treturn t->nb_entries - 1;\n+}\n+\n+static int cmp_dict_entries(const void *a_, const void *b_)\n+{\n+\tconst struct data_entry *a = a_;\n+\tconst struct data_entry *b = b_;\n+\tint diff = b->hits - a->hits;\n+\tif (!diff)\n+\t\tdiff = a->offset - b->offset;\n+\treturn diff;\n+}\n+\n+static void sort_dict_entries_by_hits(struct dict_table *t)\n+{\n+\tqsort(t->entry, t->nb_entries, sizeof(*t->entry), cmp_dict_entries);\n+\tt->hash_size = (t->nb_entries * 4 / 3) / 2;\n+\trehash_entries(t);\n+}\n+\n+void dict_dump(struct dict_table *t)\n+{\n+\tint i;\n+\n+\tsort_dict_entries_by_hits(t);\n+\tfor (i = 0; i < t->nb_entries; i++)\n+\t\tprintf(\"%d\\t%s\\n\",\n+\t\t\tt->entry[i].hits,\n+\t\t\tt->data + t->entry[i].offset);\n+}\n-- \n1.8.4.38.g317e65b\n"},{"id":"226823","messageId":"1378362001-1738-3-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 02/38] export packed_object_info()","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:25Z","receivedAt":"2013-09-05T06:19:25Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n cache.h     | 1 +\n sha1_file.c | 4 ++--\n 2 files changed, 3 insertions(+), 2 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 85b544f..b6634c4 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1160,6 +1160,7 @@ struct object_info {\n \t} u;\n };\n extern int sha1_object_info_extended(const unsigned char *, struct object_info *);\n+extern int packed_object_info(struct packed_git *, off_t, struct object_info *);\n \n /* Dumb servers support */\n extern int update_server_info(int);\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 8e27db1..c2020d0 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1782,8 +1782,8 @@ unwind:\n \tgoto out;\n }\n \n-static int packed_object_info(struct packed_git *p, off_t obj_offset,\n-\t\t\t      struct object_info *oi)\n+int packed_object_info(struct packed_git *p, off_t obj_offset,\n+\t\t       struct object_info *oi)\n {\n \tstruct pack_window *w_curs = NULL;\n \tunsigned long size;\n-- \n1.8.4.38.g317e65b\n"},{"id":"226822","messageId":"1378362001-1738-4-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 03/38] pack v4: scan tree objects","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:26Z","receivedAt":"2013-09-05T06:19:26Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Let's read a pack to feed our dictionary with all the path strings\ncontained in all the tree objects.\n\nDump the resulting dictionary sorted by frequency to stdout.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n Makefile        |   1 +\n packv4-create.c | 137 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 138 insertions(+)\n\ndiff --git a/Makefile b/Makefile\nindex 3588ca1..4716113 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -550,6 +550,7 @@ PROGRAM_OBJS += shell.o\n PROGRAM_OBJS += show-index.o\n PROGRAM_OBJS += upload-pack.o\n PROGRAM_OBJS += remote-testsvn.o\n+PROGRAM_OBJS += packv4-create.o\n \n # Binary suffix, set to .exe for Windows builds\n X =\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 2de6d41..00762a5 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -9,6 +9,8 @@\n  */\n \n #include \"cache.h\"\n+#include \"object.h\"\n+#include \"tree-walk.h\"\n \n struct data_entry {\n \tunsigned offset;\n@@ -125,6 +127,22 @@ static void sort_dict_entries_by_hits(struct dict_table *t)\n \trehash_entries(t);\n }\n \n+static struct dict_table *tree_path_table;\n+\n+static int add_tree_dict_entries(void *buf, unsigned long size)\n+{\n+\tstruct tree_desc desc;\n+\tstruct name_entry name_entry;\n+\n+\tif (!tree_path_table)\n+\t\ttree_path_table = create_dict_table();\n+\n+\tinit_tree_desc(&desc, buf, size);\n+\twhile (tree_entry(&desc, &name_entry))\n+\t\tdict_add_entry(tree_path_table, name_entry.path);\n+\treturn 0;\n+}\n+\n void dict_dump(struct dict_table *t)\n {\n \tint i;\n@@ -135,3 +153,122 @@ void dict_dump(struct dict_table *t)\n \t\t\tt->entry[i].hits,\n \t\t\tt->data + t->entry[i].offset);\n }\n+\n+struct idx_entry\n+{\n+\toff_t                offset;\n+\tconst unsigned char *sha1;\n+};\n+\n+static int sort_by_offset(const void *e1, const void *e2)\n+{\n+\tconst struct idx_entry *entry1 = e1;\n+\tconst struct idx_entry *entry2 = e2;\n+\tif (entry1->offset < entry2->offset)\n+\t\treturn -1;\n+\tif (entry1->offset > entry2->offset)\n+\t\treturn 1;\n+\treturn 0;\n+}\n+static int create_pack_dictionaries(struct packed_git *p)\n+{\n+\tuint32_t nr_objects, i;\n+\tstruct idx_entry *objects;\n+\n+\tnr_objects = p->num_objects;\n+\tobjects = xmalloc((nr_objects + 1) * sizeof(*objects));\n+\tobjects[nr_objects].offset = p->index_size - 40;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\tobjects[i].sha1 = nth_packed_object_sha1(p, i);\n+\t\tobjects[i].offset = nth_packed_object_offset(p, i);\n+\t}\n+\tqsort(objects, nr_objects, sizeof(*objects), sort_by_offset);\n+\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\tvoid *data;\n+\t\tenum object_type type;\n+\t\tunsigned long size;\n+\t\tstruct object_info oi = {};\n+\n+\t\toi.typep = &type;\n+\t\toi.sizep = &size;\n+\t\tif (packed_object_info(p, objects[i].offset, &oi) < 0)\n+\t\t\tdie(\"cannot get type of %s from %s\",\n+\t\t\t    sha1_to_hex(objects[i].sha1), p->pack_name);\n+\n+\t\tswitch (type) {\n+\t\tcase OBJ_TREE:\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tcontinue;\n+\t\t}\n+\t\tdata = unpack_entry(p, objects[i].offset, &type, &size);\n+\t\tif (!data)\n+\t\t\tdie(\"cannot unpack %s from %s\",\n+\t\t\t    sha1_to_hex(objects[i].sha1), p->pack_name);\n+\t\tif (check_sha1_signature(objects[i].sha1, data, size, typename(type)))\n+\t\t\tdie(\"packed %s from %s is corrupt\",\n+\t\t\t    sha1_to_hex(objects[i].sha1), p->pack_name);\n+\t\tif (add_tree_dict_entries(data, size) < 0)\n+\t\t\tdie(\"can't process %s object %s\",\n+\t\t\t\ttypename(type), sha1_to_hex(objects[i].sha1));\n+\t\tfree(data);\n+\t}\n+\tfree(objects);\n+\n+\treturn 0;\n+}\n+\n+static int process_one_pack(const char *path)\n+{\n+\tchar arg[PATH_MAX];\n+\tint len;\n+\tstruct packed_git *p;\n+\n+\tlen = strlcpy(arg, path, PATH_MAX);\n+\tif (len >= PATH_MAX)\n+\t\treturn error(\"name too long: %s\", path);\n+\n+\t/*\n+\t * In addition to \"foo.idx\" we accept \"foo.pack\" and \"foo\";\n+\t * normalize these forms to \"foo.idx\" for add_packed_git().\n+\t */\n+\tif (has_extension(arg, \".pack\")) {\n+\t\tstrcpy(arg + len - 5, \".idx\");\n+\t\tlen--;\n+\t} else if (!has_extension(arg, \".idx\")) {\n+\t\tif (len + 4 >= PATH_MAX)\n+\t\t\treturn error(\"name too long: %s.idx\", arg);\n+\t\tstrcpy(arg + len, \".idx\");\n+\t\tlen += 4;\n+\t}\n+\n+\t/*\n+\t * add_packed_git() uses our buffer (containing \"foo.idx\") to\n+\t * build the pack filename (\"foo.pack\").  Make sure it fits.\n+\t */\n+\tif (len + 1 >= PATH_MAX) {\n+\t\targ[len - 4] = '\\0';\n+\t\treturn error(\"name too long: %s.pack\", arg);\n+\t}\n+\n+\tp = add_packed_git(arg, len, 1);\n+\tif (!p)\n+\t\treturn error(\"packfile %s not found.\", arg);\n+\n+\tinstall_packed_git(p);\n+\tif (open_pack_index(p))\n+\t\treturn error(\"packfile %s index not opened\", p->pack_name);\n+\treturn create_pack_dictionaries(p);\n+}\n+\n+int main(int argc, char *argv[])\n+{\n+\tif (argc != 2) {\n+\t\tfprintf(stderr, \"Usage: %s <packfile>\\n\", argv[0]);\n+\t\texit(1);\n+\t}\n+\tprocess_one_pack(argv[1]);\n+\tdict_dump(tree_path_table);\n+\treturn 0;\n+}\n-- \n1.8.4.38.g317e65b\n"},{"id":"226817","messageId":"1378362001-1738-5-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 04/38] pack v4: add tree entry mode support to dictionary entries","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:27Z","receivedAt":"2013-09-05T06:19:27Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Augment dict entries with a 16-bit prefix in order to store the file\nmode value of tree entries.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 56 ++++++++++++++++++++++++++++++++++++--------------------\n 1 file changed, 36 insertions(+), 20 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 00762a5..eccd9fc 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -14,11 +14,12 @@\n \n struct data_entry {\n \tunsigned offset;\n+\tunsigned size;\n \tunsigned hits;\n };\n \n struct dict_table {\n-\tchar *data;\n+\tunsigned char *data;\n \tunsigned ptr;\n \tunsigned size;\n \tstruct data_entry *entry;\n@@ -41,18 +42,19 @@ void destroy_dict_table(struct dict_table *t)\n \tfree(t);\n }\n \n-static int locate_entry(struct dict_table *t, const char *str)\n+static int locate_entry(struct dict_table *t, const void *data, int size)\n {\n-\tint i = 0;\n-\tconst unsigned char *s = (const unsigned char *) str;\n+\tint i = 0, len = size;\n+\tconst unsigned char *p = data;\n \n-\twhile (*s)\n-\t\ti = i * 111 + *s++;\n+\twhile (len--)\n+\t\ti = i * 111 + *p++;\n \ti = (unsigned)i % t->hash_size;\n \n \twhile (t->hash[i]) {\n \t\tunsigned n = t->hash[i] - 1;\n-\t\tif (!strcmp(str, t->data + t->entry[n].offset))\n+\t\tif (t->entry[n].size == size &&\n+\t\t    memcmp(t->data + t->entry[n].offset, data, size) == 0)\n \t\t\treturn n;\n \t\tif (++i >= t->hash_size)\n \t\t\ti = 0;\n@@ -71,23 +73,28 @@ static void rehash_entries(struct dict_table *t)\n \tmemset(t->hash, 0, t->hash_size * sizeof(*t->hash));\n \n \tfor (n = 0; n < t->nb_entries; n++) {\n-\t\tint i = locate_entry(t, t->data + t->entry[n].offset);\n+\t\tint i = locate_entry(t, t->data + t->entry[n].offset,\n+\t\t\t\t\tt->entry[n].size);\n \t\tif (i < 0)\n \t\t\tt->hash[-1 - i] = n + 1;\n \t}\n }\n \n-int dict_add_entry(struct dict_table *t, const char *str)\n+int dict_add_entry(struct dict_table *t, int val, const char *str)\n {\n-\tint i, len = strlen(str) + 1;\n+\tint i, val_len = 2, str_len = strlen(str) + 1;\n \n-\tif (t->ptr + len >= t->size) {\n-\t\tt->size = (t->size + len + 1024) * 3 / 2;\n+\tif (t->ptr + val_len + str_len > t->size) {\n+\t\tt->size = (t->size + val_len + str_len + 1024) * 3 / 2;\n \t\tt->data = xrealloc(t->data, t->size);\n \t}\n-\tmemcpy(t->data + t->ptr, str, len);\n \n-\ti = (t->nb_entries) ? locate_entry(t, t->data + t->ptr) : -1;\n+\tt->data[t->ptr] = val >> 8;\n+\tt->data[t->ptr + 1] = val;\n+\tmemcpy(t->data + t->ptr + val_len, str, str_len);\n+\n+\ti = (t->nb_entries) ?\n+\t\tlocate_entry(t, t->data + t->ptr, val_len + str_len) : -1;\n \tif (i >= 0) {\n \t\tt->entry[i].hits++;\n \t\treturn i;\n@@ -98,8 +105,9 @@ int dict_add_entry(struct dict_table *t, const char *str)\n \t\tt->entry = xrealloc(t->entry, t->max_entries * sizeof(*t->entry));\n \t}\n \tt->entry[t->nb_entries].offset = t->ptr;\n+\tt->entry[t->nb_entries].size = val_len + str_len;\n \tt->entry[t->nb_entries].hits = 1;\n-\tt->ptr += len + 1;\n+\tt->ptr += val_len + str_len;\n \tt->nb_entries++;\n \n \tif (t->hash_size * 3 <= t->nb_entries * 4)\n@@ -139,7 +147,8 @@ static int add_tree_dict_entries(void *buf, unsigned long size)\n \n \tinit_tree_desc(&desc, buf, size);\n \twhile (tree_entry(&desc, &name_entry))\n-\t\tdict_add_entry(tree_path_table, name_entry.path);\n+\t\tdict_add_entry(tree_path_table, name_entry.mode,\n+\t\t\t       name_entry.path);\n \treturn 0;\n }\n \n@@ -148,10 +157,16 @@ void dict_dump(struct dict_table *t)\n \tint i;\n \n \tsort_dict_entries_by_hits(t);\n-\tfor (i = 0; i < t->nb_entries; i++)\n-\t\tprintf(\"%d\\t%s\\n\",\n-\t\t\tt->entry[i].hits,\n-\t\t\tt->data + t->entry[i].offset);\n+\tfor (i = 0; i < t->nb_entries; i++) {\n+\t\tint16_t val;\n+\t\tuint16_t uval;\n+\t\tval = t->data[t->entry[i].offset] << 8;\n+\t\tval |= t->data[t->entry[i].offset + 1];\n+\t\tuval = val;\n+\t\tprintf(\"%d\\t%d\\t%o\\t%s\\n\",\n+\t\t\tt->entry[i].hits, val, uval,\n+\t\t\tt->data + t->entry[i].offset + 2);\n+\t}\n }\n \n struct idx_entry\n@@ -170,6 +185,7 @@ static int sort_by_offset(const void *e1, const void *e2)\n \t\treturn 1;\n \treturn 0;\n }\n+\n static int create_pack_dictionaries(struct packed_git *p)\n {\n \tuint32_t nr_objects, i;\n-- \n1.8.4.38.g317e65b\n"},{"id":"226819","messageId":"1378362001-1738-6-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 05/38] pack v4: add commit object parsing","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:28Z","receivedAt":"2013-09-05T06:19:28Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Let's create another dictionary table to hold the author and committer\nentries.  We use the same table format used for tree entries where the\n16 bit data prefix is conveniently used to store the timezone value.\n\nIn order to copy straight from a commit object buffer, dict_add_entry()\nis modified to get the string length as the provided string pointer is\nnot always be null terminated.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 98 +++++++++++++++++++++++++++++++++++++++++++++++++++------\n 1 file changed, 89 insertions(+), 9 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex eccd9fc..5c08871 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -1,5 +1,5 @@\n /*\n- * packv4-create.c: management of dictionary tables used in pack v4\n+ * packv4-create.c: creation of dictionary tables and objects used in pack v4\n  *\n  * (C) Nicolas Pitre <nico@fluxnic.net>\n  *\n@@ -80,9 +80,9 @@ static void rehash_entries(struct dict_table *t)\n \t}\n }\n \n-int dict_add_entry(struct dict_table *t, int val, const char *str)\n+int dict_add_entry(struct dict_table *t, int val, const char *str, int str_len)\n {\n-\tint i, val_len = 2, str_len = strlen(str) + 1;\n+\tint i, val_len = 2;\n \n \tif (t->ptr + val_len + str_len > t->size) {\n \t\tt->size = (t->size + val_len + str_len + 1024) * 3 / 2;\n@@ -92,6 +92,7 @@ int dict_add_entry(struct dict_table *t, int val, const char *str)\n \tt->data[t->ptr] = val >> 8;\n \tt->data[t->ptr + 1] = val;\n \tmemcpy(t->data + t->ptr + val_len, str, str_len);\n+\tt->data[t->ptr + val_len + str_len] = 0;\n \n \ti = (t->nb_entries) ?\n \t\tlocate_entry(t, t->data + t->ptr, val_len + str_len) : -1;\n@@ -107,7 +108,7 @@ int dict_add_entry(struct dict_table *t, int val, const char *str)\n \tt->entry[t->nb_entries].offset = t->ptr;\n \tt->entry[t->nb_entries].size = val_len + str_len;\n \tt->entry[t->nb_entries].hits = 1;\n-\tt->ptr += val_len + str_len;\n+\tt->ptr += val_len + str_len + 1;\n \tt->nb_entries++;\n \n \tif (t->hash_size * 3 <= t->nb_entries * 4)\n@@ -135,8 +136,73 @@ static void sort_dict_entries_by_hits(struct dict_table *t)\n \trehash_entries(t);\n }\n \n+static struct dict_table *commit_name_table;\n static struct dict_table *tree_path_table;\n \n+/*\n+ * Parse the author/committer line from a canonical commit object.\n+ * The 'from' argument points right after the \"author \" or \"committer \"\n+ * string.  The time zone is parsed and stored in *tz_val.  The returned\n+ * pointer is right after the end of the email address which is also just\n+ * before the time value, or NULL if a parsing error is encountered.\n+ */\n+static char *get_nameend_and_tz(char *from, int *tz_val)\n+{\n+\tchar *end, *tz;\n+\n+\ttz = strchr(from, '\\n');\n+\t/* let's assume the smallest possible string to be \"x <x> 0 +0000\\n\" */\n+\tif (!tz || tz - from < 13)\n+\t\treturn NULL;\n+\ttz -= 4;\n+\tend = tz - 4;\n+\twhile (end - from > 5 && *end != ' ')\n+\t\tend--;\n+\tif (end[-1] != '>' || end[0] != ' ' || tz[-2] != ' ')\n+\t\treturn NULL;\n+\t*tz_val = (tz[0] - '0') * 1000 +\n+\t\t  (tz[1] - '0') * 100 +\n+\t\t  (tz[2] - '0') * 10 +\n+\t\t  (tz[3] - '0');\n+\tswitch (tz[-1]) {\n+\tdefault:\treturn NULL;\n+\tcase '+':\tbreak;\n+\tcase '-':\t*tz_val = -*tz_val;\n+\t}\n+\treturn end;\n+}\n+\n+static int add_commit_dict_entries(void *buf, unsigned long size)\n+{\n+\tchar *name, *end = NULL;\n+\tint tz_val;\n+\n+\tif (!commit_name_table)\n+\t\tcommit_name_table = create_dict_table();\n+\n+\t/* parse and add author info */\n+\tname = strstr(buf, \"\\nauthor \");\n+\tif (name) {\n+\t\tname += 8;\n+\t\tend = get_nameend_and_tz(name, &tz_val);\n+\t}\n+\tif (!name || !end)\n+\t\treturn -1;\n+\tdict_add_entry(commit_name_table, tz_val, name, end - name);\n+\n+\t/* parse and add committer info */\n+\tname = strstr(end, \"\\ncommitter \");\n+\tif (name) {\n+\t       name += 11;\n+\t       end = get_nameend_and_tz(name, &tz_val);\n+\t}\n+\tif (!name || !end)\n+\t\treturn -1;\n+\tdict_add_entry(commit_name_table, tz_val, name, end - name);\n+\n+\treturn 0;\n+}\n+\n static int add_tree_dict_entries(void *buf, unsigned long size)\n {\n \tstruct tree_desc desc;\n@@ -146,13 +212,16 @@ static int add_tree_dict_entries(void *buf, unsigned long size)\n \t\ttree_path_table = create_dict_table();\n \n \tinit_tree_desc(&desc, buf, size);\n-\twhile (tree_entry(&desc, &name_entry))\n+\twhile (tree_entry(&desc, &name_entry)) {\n+\t\tint pathlen = tree_entry_len(&name_entry);\n \t\tdict_add_entry(tree_path_table, name_entry.mode,\n-\t\t\t       name_entry.path);\n+\t\t\t\tname_entry.path, pathlen);\n+\t}\n+\n \treturn 0;\n }\n \n-void dict_dump(struct dict_table *t)\n+void dump_dict_table(struct dict_table *t)\n {\n \tint i;\n \n@@ -169,6 +238,12 @@ void dict_dump(struct dict_table *t)\n \t}\n }\n \n+static void dict_dump(void)\n+{\n+\tdump_dict_table(commit_name_table);\n+\tdump_dict_table(tree_path_table);\n+}\n+\n struct idx_entry\n {\n \toff_t                offset;\n@@ -205,6 +280,7 @@ static int create_pack_dictionaries(struct packed_git *p)\n \t\tenum object_type type;\n \t\tunsigned long size;\n \t\tstruct object_info oi = {};\n+\t\tint (*add_dict_entries)(void *, unsigned long);\n \n \t\toi.typep = &type;\n \t\toi.sizep = &size;\n@@ -213,7 +289,11 @@ static int create_pack_dictionaries(struct packed_git *p)\n \t\t\t    sha1_to_hex(objects[i].sha1), p->pack_name);\n \n \t\tswitch (type) {\n+\t\tcase OBJ_COMMIT:\n+\t\t\tadd_dict_entries = add_commit_dict_entries;\n+\t\t\tbreak;\n \t\tcase OBJ_TREE:\n+\t\t\tadd_dict_entries = add_tree_dict_entries;\n \t\t\tbreak;\n \t\tdefault:\n \t\t\tcontinue;\n@@ -225,7 +305,7 @@ static int create_pack_dictionaries(struct packed_git *p)\n \t\tif (check_sha1_signature(objects[i].sha1, data, size, typename(type)))\n \t\t\tdie(\"packed %s from %s is corrupt\",\n \t\t\t    sha1_to_hex(objects[i].sha1), p->pack_name);\n-\t\tif (add_tree_dict_entries(data, size) < 0)\n+\t\tif (add_dict_entries(data, size) < 0)\n \t\t\tdie(\"can't process %s object %s\",\n \t\t\t\ttypename(type), sha1_to_hex(objects[i].sha1));\n \t\tfree(data);\n@@ -285,6 +365,6 @@ int main(int argc, char *argv[])\n \t\texit(1);\n \t}\n \tprocess_one_pack(argv[1]);\n-\tdict_dump(tree_path_table);\n+\tdict_dump();\n \treturn 0;\n }\n-- \n1.8.4.38.g317e65b\n"},{"id":"226820","messageId":"1378362001-1738-7-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 06/38] pack v4: split the object list and dictionary creation","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:29Z","receivedAt":"2013-09-05T06:19:29Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 58 +++++++++++++++++++++++++++++++++++++++++++--------------\n 1 file changed, 44 insertions(+), 14 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 5c08871..20d97a4 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -261,7 +261,7 @@ static int sort_by_offset(const void *e1, const void *e2)\n \treturn 0;\n }\n \n-static int create_pack_dictionaries(struct packed_git *p)\n+static struct idx_entry *get_packed_object_list(struct packed_git *p)\n {\n \tuint32_t nr_objects, i;\n \tstruct idx_entry *objects;\n@@ -275,7 +275,15 @@ static int create_pack_dictionaries(struct packed_git *p)\n \t}\n \tqsort(objects, nr_objects, sizeof(*objects), sort_by_offset);\n \n-\tfor (i = 0; i < nr_objects; i++) {\n+\treturn objects;\n+}\n+\n+static int create_pack_dictionaries(struct packed_git *p,\n+\t\t\t\t    struct idx_entry *objects)\n+{\n+\tunsigned int i;\n+\n+\tfor (i = 0; i < p->num_objects; i++) {\n \t\tvoid *data;\n \t\tenum object_type type;\n \t\tunsigned long size;\n@@ -310,20 +318,21 @@ static int create_pack_dictionaries(struct packed_git *p)\n \t\t\t\ttypename(type), sha1_to_hex(objects[i].sha1));\n \t\tfree(data);\n \t}\n-\tfree(objects);\n \n \treturn 0;\n }\n \n-static int process_one_pack(const char *path)\n+static struct packed_git *open_pack(const char *path)\n {\n \tchar arg[PATH_MAX];\n \tint len;\n \tstruct packed_git *p;\n \n \tlen = strlcpy(arg, path, PATH_MAX);\n-\tif (len >= PATH_MAX)\n-\t\treturn error(\"name too long: %s\", path);\n+\tif (len >= PATH_MAX) {\n+\t\terror(\"name too long: %s\", path);\n+\t\treturn NULL;\n+\t}\n \n \t/*\n \t * In addition to \"foo.idx\" we accept \"foo.pack\" and \"foo\";\n@@ -333,8 +342,10 @@ static int process_one_pack(const char *path)\n \t\tstrcpy(arg + len - 5, \".idx\");\n \t\tlen--;\n \t} else if (!has_extension(arg, \".idx\")) {\n-\t\tif (len + 4 >= PATH_MAX)\n-\t\t\treturn error(\"name too long: %s.idx\", arg);\n+\t\tif (len + 4 >= PATH_MAX) {\n+\t\t\terror(\"name too long: %s.idx\", arg);\n+\t\t\treturn NULL;\n+\t\t}\n \t\tstrcpy(arg + len, \".idx\");\n \t\tlen += 4;\n \t}\n@@ -345,17 +356,36 @@ static int process_one_pack(const char *path)\n \t */\n \tif (len + 1 >= PATH_MAX) {\n \t\targ[len - 4] = '\\0';\n-\t\treturn error(\"name too long: %s.pack\", arg);\n+\t\terror(\"name too long: %s.pack\", arg);\n+\t\treturn NULL;\n \t}\n \n \tp = add_packed_git(arg, len, 1);\n-\tif (!p)\n-\t\treturn error(\"packfile %s not found.\", arg);\n+\tif (!p) {\n+\t\terror(\"packfile %s not found.\", arg);\n+\t\treturn NULL;\n+\t}\n \n \tinstall_packed_git(p);\n-\tif (open_pack_index(p))\n-\t\treturn error(\"packfile %s index not opened\", p->pack_name);\n-\treturn create_pack_dictionaries(p);\n+\tif (open_pack_index(p)) {\n+\t\terror(\"packfile %s index not opened\", p->pack_name);\n+\t\treturn NULL;\n+\t}\n+\n+\treturn p;\n+}\n+\n+static void process_one_pack(char *src_pack)\n+{\n+\tstruct packed_git *p;\n+\tstruct idx_entry *objs;\n+\n+\tp = open_pack(src_pack);\n+\tif (!p)\n+\t\tdie(\"unable to open source pack\");\n+\n+\tobjs = get_packed_object_list(p);\n+\tcreate_pack_dictionaries(p, objs);\n }\n \n int main(int argc, char *argv[])\n-- \n1.8.4.38.g317e65b\n"},{"id":"226814","messageId":"1378362001-1738-8-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 07/38] pack v4: move to struct pack_idx_entry and get rid of our own struct idx_entry","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:30Z","receivedAt":"2013-09-05T06:19:30Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Let's create a struct pack_idx_entry list with sorted sha1 which will\nbe useful later.  The offset sorted list is now a separate indirect\nlist.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 72 +++++++++++++++++++++++++++++++++------------------------\n 1 file changed, 42 insertions(+), 30 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 20d97a4..012129b 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -11,6 +11,7 @@\n #include \"cache.h\"\n #include \"object.h\"\n #include \"tree-walk.h\"\n+#include \"pack.h\"\n \n struct data_entry {\n \tunsigned offset;\n@@ -244,46 +245,53 @@ static void dict_dump(void)\n \tdump_dict_table(tree_path_table);\n }\n \n-struct idx_entry\n+static struct pack_idx_entry *get_packed_object_list(struct packed_git *p)\n {\n-\toff_t                offset;\n-\tconst unsigned char *sha1;\n-};\n+\tunsigned i, nr_objects = p->num_objects;\n+\tstruct pack_idx_entry *objects;\n+\n+\tobjects = xmalloc((nr_objects + 1) * sizeof(*objects));\n+\tobjects[nr_objects].offset = p->pack_size - 20;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\thashcpy(objects[i].sha1, nth_packed_object_sha1(p, i));\n+\t\tobjects[i].offset = nth_packed_object_offset(p, i);\n+\t}\n+\n+\treturn objects;\n+}\n \n static int sort_by_offset(const void *e1, const void *e2)\n {\n-\tconst struct idx_entry *entry1 = e1;\n-\tconst struct idx_entry *entry2 = e2;\n-\tif (entry1->offset < entry2->offset)\n+\tconst struct pack_idx_entry * const *entry1 = e1;\n+\tconst struct pack_idx_entry * const *entry2 = e2;\n+\tif ((*entry1)->offset < (*entry2)->offset)\n \t\treturn -1;\n-\tif (entry1->offset > entry2->offset)\n+\tif ((*entry1)->offset > (*entry2)->offset)\n \t\treturn 1;\n \treturn 0;\n }\n \n-static struct idx_entry *get_packed_object_list(struct packed_git *p)\n+static struct pack_idx_entry **sort_objs_by_offset(struct pack_idx_entry *list,\n+\t\t\t\t\t\t    unsigned nr_objects)\n {\n-\tuint32_t nr_objects, i;\n-\tstruct idx_entry *objects;\n+\tunsigned i;\n+\tstruct pack_idx_entry **sorted;\n \n-\tnr_objects = p->num_objects;\n-\tobjects = xmalloc((nr_objects + 1) * sizeof(*objects));\n-\tobjects[nr_objects].offset = p->index_size - 40;\n-\tfor (i = 0; i < nr_objects; i++) {\n-\t\tobjects[i].sha1 = nth_packed_object_sha1(p, i);\n-\t\tobjects[i].offset = nth_packed_object_offset(p, i);\n-\t}\n-\tqsort(objects, nr_objects, sizeof(*objects), sort_by_offset);\n+\tsorted = xmalloc((nr_objects + 1) * sizeof(*sorted));\n+\tfor (i = 0; i < nr_objects + 1; i++)\n+\t\tsorted[i] = &list[i];\n+\tqsort(sorted, nr_objects + 1, sizeof(*sorted), sort_by_offset);\n \n-\treturn objects;\n+\treturn sorted;\n }\n \n static int create_pack_dictionaries(struct packed_git *p,\n-\t\t\t\t    struct idx_entry *objects)\n+\t\t\t\t    struct pack_idx_entry **obj_list)\n {\n \tunsigned int i;\n \n \tfor (i = 0; i < p->num_objects; i++) {\n+\t\tstruct pack_idx_entry *obj = obj_list[i];\n \t\tvoid *data;\n \t\tenum object_type type;\n \t\tunsigned long size;\n@@ -292,9 +300,9 @@ static int create_pack_dictionaries(struct packed_git *p,\n \n \t\toi.typep = &type;\n \t\toi.sizep = &size;\n-\t\tif (packed_object_info(p, objects[i].offset, &oi) < 0)\n+\t\tif (packed_object_info(p, obj->offset, &oi) < 0)\n \t\t\tdie(\"cannot get type of %s from %s\",\n-\t\t\t    sha1_to_hex(objects[i].sha1), p->pack_name);\n+\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n \n \t\tswitch (type) {\n \t\tcase OBJ_COMMIT:\n@@ -306,16 +314,16 @@ static int create_pack_dictionaries(struct packed_git *p,\n \t\tdefault:\n \t\t\tcontinue;\n \t\t}\n-\t\tdata = unpack_entry(p, objects[i].offset, &type, &size);\n+\t\tdata = unpack_entry(p, obj->offset, &type, &size);\n \t\tif (!data)\n \t\t\tdie(\"cannot unpack %s from %s\",\n-\t\t\t    sha1_to_hex(objects[i].sha1), p->pack_name);\n-\t\tif (check_sha1_signature(objects[i].sha1, data, size, typename(type)))\n+\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\t\tif (check_sha1_signature(obj->sha1, data, size, typename(type)))\n \t\t\tdie(\"packed %s from %s is corrupt\",\n-\t\t\t    sha1_to_hex(objects[i].sha1), p->pack_name);\n+\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n \t\tif (add_dict_entries(data, size) < 0)\n \t\t\tdie(\"can't process %s object %s\",\n-\t\t\t\ttypename(type), sha1_to_hex(objects[i].sha1));\n+\t\t\t\ttypename(type), sha1_to_hex(obj->sha1));\n \t\tfree(data);\n \t}\n \n@@ -378,14 +386,18 @@ static struct packed_git *open_pack(const char *path)\n static void process_one_pack(char *src_pack)\n {\n \tstruct packed_git *p;\n-\tstruct idx_entry *objs;\n+\tstruct pack_idx_entry *objs, **p_objs;\n+\tunsigned nr_objects;\n \n \tp = open_pack(src_pack);\n \tif (!p)\n \t\tdie(\"unable to open source pack\");\n \n+\tnr_objects = p->num_objects;\n \tobjs = get_packed_object_list(p);\n-\tcreate_pack_dictionaries(p, objs);\n+\tp_objs = sort_objs_by_offset(objs, nr_objects);\n+\n+\tcreate_pack_dictionaries(p, p_objs);\n }\n \n int main(int argc, char *argv[])\n-- \n1.8.4.38.g317e65b\n"},{"id":"226818","messageId":"1378362001-1738-9-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 08/38] pack v4: basic SHA1 reference encoding","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:31Z","receivedAt":"2013-09-05T06:19:31Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"The SHA1 reference is either an index into a SHA1 table using the variable\nlength number encoding, or the literal 20 bytes SHA1 prefixed with a 0.\n\nThe index 0 discriminates between an actual index value or the literal\nSHA1.  Therefore when the index is used its value must be increased by 1.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 29 +++++++++++++++++++++++++++++\n 1 file changed, 29 insertions(+)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 012129b..12527c0 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -12,6 +12,7 @@\n #include \"object.h\"\n #include \"tree-walk.h\"\n #include \"pack.h\"\n+#include \"varint.h\"\n \n struct data_entry {\n \tunsigned offset;\n@@ -245,6 +246,34 @@ static void dict_dump(void)\n \tdump_dict_table(tree_path_table);\n }\n \n+/*\n+ * Encode an object SHA1 reference with either an object index into the\n+ * pack SHA1 table incremented by 1, or the literal SHA1 value prefixed\n+ * with a zero byte if the needed SHA1 is not available in the table.\n+ */\n+static struct pack_idx_entry *all_objs;\n+static unsigned all_objs_nr;\n+static int encode_sha1ref(const unsigned char *sha1, unsigned char *buf)\n+{\n+\tunsigned lo = 0, hi = all_objs_nr;\n+\n+\tdo {\n+\t\tunsigned mi = (lo + hi) / 2;\n+\t\tint cmp = hashcmp(all_objs[mi].sha1, sha1);\n+\n+\t\tif (cmp == 0)\n+\t\t\treturn encode_varint(mi + 1, buf);\n+\t\tif (cmp > 0)\n+\t\t\thi = mi;\n+\t\telse\n+\t\t\tlo = mi+1;\n+\t} while (lo < hi);\n+\n+\t*buf++ = 0;\n+\thashcpy(buf, sha1);\n+\treturn 1 + 20;\n+}\n+\n static struct pack_idx_entry *get_packed_object_list(struct packed_git *p)\n {\n \tunsigned i, nr_objects = p->num_objects;\n-- \n1.8.4.38.g317e65b\n"},{"id":"226812","messageId":"1378362001-1738-10-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 09/38] introduce get_sha1_lowhex()","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:32Z","receivedAt":"2013-09-05T06:19:32Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"This is like get_sha1_hex() but stricter in accepting lowercase letters\nonly.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n cache.h |  3 +++\n hex.c   | 11 +++++++++++\n 2 files changed, 14 insertions(+)\n\ndiff --git a/cache.h b/cache.h\nindex b6634c4..4231dfa 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -850,8 +850,11 @@ extern int for_each_abbrev(const char *prefix, each_abbrev_fn, void *);\n  * Return 0 on success.  Reading stops if a NUL is encountered in the\n  * input, so it is safe to pass this function an arbitrary\n  * null-terminated string.\n+ *\n+ * The \"low\" version accepts numbers and lowercase letters only.\n  */\n extern int get_sha1_hex(const char *hex, unsigned char *sha1);\n+extern int get_sha1_lowhex(const char *hex, unsigned char *sha1);\n \n extern char *sha1_to_hex(const unsigned char *sha1);\t/* static buffer result! */\n extern int read_ref_full(const char *refname, unsigned char *sha1,\ndiff --git a/hex.c b/hex.c\nindex 9ebc050..1d7eae1 100644\n--- a/hex.c\n+++ b/hex.c\n@@ -56,6 +56,17 @@ int get_sha1_hex(const char *hex, unsigned char *sha1)\n \treturn 0;\n }\n \n+int get_sha1_lowhex(const char *hex, unsigned char *sha1)\n+{\n+\tint i;\n+\n+\t/* uppercase letters (as well as '\\0') have bit 5 clear */\n+\tfor (i = 0; i < 20; i++)\n+\t\tif (!(hex[i] & 0x20))\n+\t\t\treturn -1;\n+\treturn get_sha1_hex(hex, sha1);\n+}\n+\n char *sha1_to_hex(const unsigned char *sha1)\n {\n \tstatic int bufno;\n-- \n1.8.4.38.g317e65b\n"},{"id":"226787","messageId":"1378362001-1738-11-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 10/38] pack v4: commit object encoding","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:33Z","receivedAt":"2013-09-05T06:19:33Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"This goes as follows:\n\n- Tree reference: either variable length encoding of the index\n  into the SHA1 table or the literal SHA1 prefixed by 0 (see\n  encode_sha1ref()).\n\n- Parent count: variable length encoding of the number of parents.\n  This is normally going to occupy a single byte but doesn't have to.\n\n- List of parent references: a list of encode_sha1ref() encoded\n  references, or nothing if the parent count was zero.\n\n- Author reference: variable length encoding of an index into the author\n  identifier dictionary table which also covers the time zone.  To make\n  the overall encoding efficient, the author table is sorted by usage\n  frequency so the most used names are first and require the shortest\n  index encoding.\n\n- Author time stamp: variable length encoded.  Year 2038 ready!\n\n- Committer reference: same as author reference.\n\n- Committer time stamp: same as author time stamp.\n\nThe remainder of the canonical commit object content is then zlib\ncompressed and appended to the above.\n\nRationale: The most important commit object data is densely encoded while\nrequiring no zlib inflate processing on access, and all SHA1 references\nare most likely to be direct indices into the pack index file requiring\nno SHA1 search into the pack index file.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 119 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 119 insertions(+)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 12527c0..d4a79f4 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -14,6 +14,9 @@\n #include \"pack.h\"\n #include \"varint.h\"\n \n+\n+static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n+\n struct data_entry {\n \tunsigned offset;\n \tunsigned size;\n@@ -274,6 +277,122 @@ static int encode_sha1ref(const unsigned char *sha1, unsigned char *buf)\n \treturn 1 + 20;\n }\n \n+/*\n+ * This converts a canonical commit object buffer into its\n+ * tightly packed representation using the already populated\n+ * and sorted commit_name_table dictionary.  The parsing is\n+ * strict so to ensure the canonical version may always be\n+ * regenerated and produce the same hash.\n+ */\n+void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n+{\n+\tunsigned long size = *sizep;\n+\tchar *in, *tail, *end;\n+\tunsigned char *out;\n+\tunsigned char sha1[20];\n+\tint nb_parents, index, tz_val;\n+\tunsigned long time;\n+\tz_stream stream;\n+\tint status;\n+\n+\t/*\n+\t * It is guaranteed that the output is always going to be smaller\n+\t * than the input.  We could even do this conversion in place.\n+\t */\n+\tin = buffer;\n+\ttail = in + size;\n+\tbuffer = xmalloc(size);\n+\tout = buffer;\n+\n+\t/* parse the \"tree\" line */\n+\tif (in + 46 >= tail || memcmp(in, \"tree \", 5) || in[45] != '\\n')\n+\t\tgoto bad_data;\n+\tif (get_sha1_lowhex(in + 5, sha1) < 0)\n+\t\tgoto bad_data;\n+\tin += 46;\n+\tout += encode_sha1ref(sha1, out);\n+\n+\t/* count how many \"parent\" lines */\n+\tnb_parents = 0;\n+\twhile (in + 48 < tail && !memcmp(in, \"parent \", 7) && in[47] == '\\n') {\n+\t\tnb_parents++;\n+\t\tin += 48;\n+\t}\n+\tout += encode_varint(nb_parents, out);\n+\n+\t/* rewind and parse the \"parent\" lines */\n+\tin -= 48 * nb_parents;\n+\twhile (nb_parents--) {\n+\t\tif (get_sha1_lowhex(in + 7, sha1))\n+\t\t\tgoto bad_data;\n+\t\tout += encode_sha1ref(sha1, out);\n+\t\tin += 48;\n+\t}\n+\n+\t/* parse the \"author\" line */\n+\t/* it must be at least \"author x <x> 0 +0000\\n\" i.e. 21 chars */\n+\tif (in + 21 >= tail || memcmp(in, \"author \", 7))\n+\t\tgoto bad_data;\n+\tin += 7;\n+\tend = get_nameend_and_tz(in, &tz_val);\n+\tif (!end)\n+\t\tgoto bad_data;\n+\tindex = dict_add_entry(commit_name_table, tz_val, in, end - in);\n+\tif (index < 0)\n+\t\tgoto bad_dict;\n+\tout += encode_varint(index, out);\n+\ttime = strtoul(end, &end, 10);\n+\tif (!end || end[0] != ' ' || end[6] != '\\n')\n+\t\tgoto bad_data;\n+\tout += encode_varint(time, out);\n+\tin = end + 7;\n+\n+\t/* parse the \"committer\" line */\n+\t/* it must be at least \"committer x <x> 0 +0000\\n\" i.e. 24 chars */\n+\tif (in + 24 >= tail || memcmp(in, \"committer \", 7))\n+\t\tgoto bad_data;\n+\tin += 10;\n+\tend = get_nameend_and_tz(in, &tz_val);\n+\tif (!end)\n+\t\tgoto bad_data;\n+\tindex = dict_add_entry(commit_name_table, tz_val, in, end - in);\n+\tif (index < 0)\n+\t\tgoto bad_dict;\n+\tout += encode_varint(index, out);\n+\ttime = strtoul(end, &end, 10);\n+\tif (!end || end[0] != ' ' || end[6] != '\\n')\n+\t\tgoto bad_data;\n+\tout += encode_varint(time, out);\n+\tin = end + 7;\n+\n+\t/* finally, deflate the remaining data */\n+\tmemset(&stream, 0, sizeof(stream));\n+\tdeflateInit(&stream, pack_compression_level);\n+\tstream.next_in = (unsigned char *)in;\n+\tstream.avail_in = tail - in;\n+\tstream.next_out = (unsigned char *)out;\n+\tstream.avail_out = size - (out - (unsigned char *)buffer);\n+\tstatus = deflate(&stream, Z_FINISH);\n+\tend = (char *)stream.next_out;\n+\tdeflateEnd(&stream);\n+\tif (status != Z_STREAM_END) {\n+\t\terror(\"deflate error status %d\", status);\n+\t\tgoto bad;\n+\t}\n+\n+\t*sizep = end - (char *)buffer;\n+\treturn buffer;\n+\n+bad_data:\n+\terror(\"bad commit data\");\n+\tgoto bad;\n+bad_dict:\n+\terror(\"bad dict entry\");\n+bad:\n+\tfree(buffer);\n+\treturn NULL;\n+}\n+\n static struct pack_idx_entry *get_packed_object_list(struct packed_git *p)\n {\n \tunsigned i, nr_objects = p->num_objects;\n-- \n1.8.4.38.g317e65b\n"},{"id":"226805","messageId":"1378362001-1738-12-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 11/38] pack v4: tree object encoding","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:34Z","receivedAt":"2013-09-05T06:19:34Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"This goes as follows:\n\n- Number of tree entries: variable length encoded.\n\nThen for each tree entry:\n\n- Path component reference: variable length encoded index into the path\n  dictionary table which also covers the entry mode. To make the overall\n  encoding efficient, the path table is already sorted by usage frequency\n  so the most used path names are first and require the shortest index\n  encoding.\n\n- SHA1 reference: either variable length encoding of the index into the\n  SHA1 table or the literal SHA1 prefixed by 0 (see encode_sha1ref()).\n\nRationale: all the tree object data is densely encoded while requiring\nno zlib inflate processing on access, and all SHA1 references are most\nlikely to be direct indices into the pack index file requiring no SHA1\nsearch.  Path filtering can be accomplished on the path index directly\nwithout any string comparison during the tree traversal.\n\nStill lacking is some kind of delta encoding for multiple tree objects\nwith only small differences between them.  But that'll come later.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 66 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 66 insertions(+)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex d4a79f4..b91ee0b 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -393,6 +393,72 @@ bad:\n \treturn NULL;\n }\n \n+/*\n+ * This converts a canonical tree object buffer into its\n+ * tightly packed representation using the already populated\n+ * and sorted tree_path_table dictionary.  The parsing is\n+ * strict so to ensure the canonical version may always be\n+ * regenerated and produce the same hash.\n+ */\n+void *pv4_encode_tree(void *_buffer, unsigned long *sizep)\n+{\n+\tunsigned long size = *sizep;\n+\tunsigned char *in, *out, *end, *buffer = _buffer;\n+\tstruct tree_desc desc;\n+\tstruct name_entry name_entry;\n+\tint nb_entries;\n+\n+\tif (!size)\n+\t\treturn NULL;\n+\n+\t/*\n+\t * We can't make sure the result will always be smaller than the\n+\t * input. The smallest possible entry is \"0 x\\0<40 byte SHA1>\"\n+\t * or 44 bytes.  The output entry may have a realistic path index\n+\t * encoding using up to 3 bytes, and a non indexable SHA1 meaning\n+\t * 41 bytes.  And the output data already has the nb_entries\n+\t * headers.  In practice the output size will be significantly\n+\t * smaller but for now let's make it simple.\n+\t */\n+\tin = buffer;\n+\tout = xmalloc(size + 48);\n+\tend = out + size + 48;\n+\tbuffer = out;\n+\n+\t/* let's count how many entries there are */\n+\tinit_tree_desc(&desc, in, size);\n+\tnb_entries = 0;\n+\twhile (tree_entry(&desc, &name_entry))\n+\t\tnb_entries++;\n+\tout += encode_varint(nb_entries, out);\n+\n+\tinit_tree_desc(&desc, in, size);\n+\twhile (tree_entry(&desc, &name_entry)) {\n+\t\tint pathlen, index;\n+\n+\t\tif (end - out < 48) {\n+\t\t\tunsigned long sofar = out - buffer;\n+\t\t\tbuffer = xrealloc(buffer, (sofar + 48)*2);\n+\t\t\tend = buffer + (sofar + 48)*2;\n+\t\t\tout = buffer + sofar;\n+\t\t}\n+\n+\t\tpathlen = tree_entry_len(&name_entry);\n+\t\tindex = dict_add_entry(tree_path_table, name_entry.mode,\n+\t\t\t\t       name_entry.path, pathlen);\n+\t\tif (index < 0) {\n+\t\t\terror(\"missing tree dict entry\");\n+\t\t\tfree(buffer);\n+\t\t\treturn NULL;\n+\t\t}\n+\t\tout += encode_varint(index, out);\n+\t\tout += encode_sha1ref(name_entry.sha1, out);\n+\t}\n+\n+\t*sizep = out - buffer;\n+\treturn buffer;\n+}\n+\n static struct pack_idx_entry *get_packed_object_list(struct packed_git *p)\n {\n \tunsigned i, nr_objects = p->num_objects;\n-- \n1.8.4.38.g317e65b\n"},{"id":"226786","messageId":"1378362001-1738-13-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 12/38] pack v4: dictionary table output","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:35Z","receivedAt":"2013-09-05T06:19:35Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Here's the code to dump a table into a pack.  Table entries are written\naccording to the current sort order. This is important as objects use\nthis order to index into the table.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 49 +++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 49 insertions(+)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex b91ee0b..92d3662 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -544,6 +544,55 @@ static int create_pack_dictionaries(struct packed_git *p,\n \treturn 0;\n }\n \n+static unsigned long write_dict_table(struct sha1file *f, struct dict_table *t)\n+{\n+\tunsigned char buffer[1024];\n+\tunsigned hdrlen;\n+\tunsigned long size, datalen;\n+\tz_stream stream;\n+\tint i, status;\n+\n+\t/*\n+\t * Stored dict table format: uncompressed data length followed by\n+\t * compressed content.\n+\t */\n+\n+\tdatalen = t->ptr;\n+\thdrlen = encode_varint(datalen, buffer);\n+\tsha1write(f, buffer, hdrlen);\n+\n+\tmemset(&stream, 0, sizeof(stream));\n+\tdeflateInit(&stream, pack_compression_level);\n+\n+\tfor (i = 0; i < t->nb_entries; i++) {\n+\t\tstream.next_in = t->data + t->entry[i].offset;\n+\t\tstream.avail_in = 2 + strlen((char *)t->data + t->entry[i].offset + 2) + 1;\n+\t\tdo {\n+\t\t\tstream.next_out = buffer;\n+\t\t\tstream.avail_out = sizeof(buffer);\n+\t\t\tstatus = deflate(&stream, 0);\n+\t\t\tsize = stream.next_out - (unsigned char *)buffer;\n+\t\t\tsha1write(f, buffer, size);\n+\t\t} while (status == Z_OK);\n+\t}\n+\tdo {\n+\t\tstream.next_out = buffer;\n+\t\tstream.avail_out = sizeof(buffer);\n+\t\tstatus = deflate(&stream, Z_FINISH);\n+\t\tsize = stream.next_out - (unsigned char *)buffer;\n+\t\tsha1write(f, buffer, size);\n+\t} while (status == Z_OK);\n+\tif (status != Z_STREAM_END)\n+\t\tdie(\"unable to deflate dictionary table (%d)\", status);\n+\tif (stream.total_in != datalen)\n+\t\tdie(\"dict data size mismatch (%ld vs %ld)\",\n+\t\t    stream.total_in, datalen);\n+\tdatalen = stream.total_out;\n+\tdeflateEnd(&stream);\n+\n+\treturn hdrlen + datalen;\n+}\n+\n static struct packed_git *open_pack(const char *path)\n {\n \tchar arg[PATH_MAX];\n-- \n1.8.4.38.g317e65b\n"},{"id":"226808","messageId":"1378362001-1738-14-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 13/38] pack v4: creation code","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:36Z","receivedAt":"2013-09-05T06:19:36Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Let's actually open the destination pack file and write the header and\nthe tables.\n\nThe header isn't much different from pack v3, except for the pack version\nnumber of course.\n\nThe first table is the sorted SHA1 table normally found in the pack index\nfile.  With pack v4 we write this table in the main pack file instead as\nit is index referenced by subsequent objects in the pack.  Doing so has\nmany advantages:\n\n- The SHA1 references used to be duplicated on disk: once in the pack\n  index file, and then at least once or more within commit and tree\n  objects referencing them.  The only SHA1 which is not being listed more\n  than once this way is the one for a branch tip commit object and those\n  are normally very few.  Now all that SHA1 data is represented only once.\n\n- The SHA1 references found in commit and tree objects can be obtained\n  on disk directly without having to deflate those objects first.\n\nThe SHA1 table size is obtained by multiplying the number of objects by 20.\n\nAnd then the commit and path dictionary tables are written right after\nthe SHA1 table.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 60 ++++++++++++++++++++++++++++++++++++++++++++++++++++-----\n 1 file changed, 55 insertions(+), 5 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 92d3662..61b70c8 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -593,6 +593,48 @@ static unsigned long write_dict_table(struct sha1file *f, struct dict_table *t)\n \treturn hdrlen + datalen;\n }\n \n+static struct sha1file * packv4_open(char *path)\n+{\n+\tint fd;\n+\n+\tfd = open(path, O_CREAT|O_EXCL|O_WRONLY, 0600);\n+\tif (fd < 0)\n+\t\tdie_errno(\"unable to create '%s'\", path);\n+\treturn sha1fd(fd, path);\n+}\n+\n+static unsigned int packv4_write_header(struct sha1file *f, unsigned nr_objects)\n+{\n+\tstruct pack_header hdr;\n+\n+\thdr.hdr_signature = htonl(PACK_SIGNATURE);\n+\thdr.hdr_version = htonl(4);\n+\thdr.hdr_entries = htonl(nr_objects);\n+\tsha1write(f, &hdr, sizeof(hdr));\n+\n+\treturn sizeof(hdr);\n+}\n+\n+static unsigned long packv4_write_tables(struct sha1file *f, unsigned nr_objects,\n+\t\t\t\t\t struct pack_idx_entry *objs)\n+{\n+\tunsigned i;\n+\tunsigned long written = 0;\n+\n+\t/* The sorted list of object SHA1's is always first */\n+\tfor (i = 0; i < nr_objects; i++)\n+\t\tsha1write(f, objs[i].sha1, 20);\n+\twritten = 20 * nr_objects;\n+\n+\t/* Then the commit dictionary table */\n+\twritten += write_dict_table(f, commit_name_table);\n+\n+\t/* Followed by the path component dictionary table */\n+\twritten += write_dict_table(f, tree_path_table);\n+\n+\treturn written;\n+}\n+\n static struct packed_git *open_pack(const char *path)\n {\n \tchar arg[PATH_MAX];\n@@ -646,9 +688,10 @@ static struct packed_git *open_pack(const char *path)\n \treturn p;\n }\n \n-static void process_one_pack(char *src_pack)\n+static void process_one_pack(char *src_pack, char *dst_pack)\n {\n \tstruct packed_git *p;\n+\tstruct sha1file *f;\n \tstruct pack_idx_entry *objs, **p_objs;\n \tunsigned nr_objects;\n \n@@ -661,15 +704,22 @@ static void process_one_pack(char *src_pack)\n \tp_objs = sort_objs_by_offset(objs, nr_objects);\n \n \tcreate_pack_dictionaries(p, p_objs);\n+\n+\tf = packv4_open(dst_pack);\n+\tif (!f)\n+\t\tdie(\"unable to open destination pack\");\n+\tpackv4_write_header(f, nr_objects);\n+\tpackv4_write_tables(f, nr_objects, objs);\n }\n \n int main(int argc, char *argv[])\n {\n-\tif (argc != 2) {\n-\t\tfprintf(stderr, \"Usage: %s <packfile>\\n\", argv[0]);\n+\tif (argc != 3) {\n+\t\tfprintf(stderr, \"Usage: %s <src_packfile> <dst_packfile>\\n\", argv[0]);\n \t\texit(1);\n \t}\n-\tprocess_one_pack(argv[1]);\n-\tdict_dump();\n+\tprocess_one_pack(argv[1], argv[2]);\n+\tif (0)\n+\t\tdict_dump();\n \treturn 0;\n }\n-- \n1.8.4.38.g317e65b\n"},{"id":"226797","messageId":"1378362001-1738-15-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 14/38] pack v4: object headers","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:37Z","receivedAt":"2013-09-05T06:19:37Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"In pack v4 the object size and type is encoded differently from pack v3.\nThe object size uses the same efficient variable length number encoding\nalready used elsewhere.\n\nThe object type has 4 bits allocated to it compared to 3 bits in pack v3.\nThis should be quite sufficient for the foreseeable future, especially\nsince pack v4 has only one type of delta object instead of two.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 27 +++++++++++++++++++++++++++\n 1 file changed, 27 insertions(+)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 61b70c8..6098062 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -635,6 +635,33 @@ static unsigned long packv4_write_tables(struct sha1file *f, unsigned nr_objects\n \treturn written;\n }\n \n+static int write_object_header(struct sha1file *f, enum object_type type, unsigned long size)\n+{\n+\tunsigned char buf[16];\n+\tuint64_t val;\n+\tint len;\n+\n+\t/*\n+\t * We really have only one kind of delta object.\n+\t */\n+\tif (type == OBJ_OFS_DELTA)\n+\t\ttype = OBJ_REF_DELTA;\n+\n+\t/*\n+\t * We allocate 4 bits in the LSB for the object type which should\n+\t * be good for quite a while, given that we effectively encodes\n+\t * only 5 object types: commit, tree, blob, delta, tag.\n+\t */\n+\tval = size;\n+\tif (MSB(val, 4))\n+\t\tdie(\"fixme: the code doesn't currently cope with big sizes\");\n+\tval <<= 4;\n+\tval |= type;\n+\tlen = encode_varint(val, buf);\n+\tsha1write(f, buf, len);\n+\treturn len;\n+}\n+\n static struct packed_git *open_pack(const char *path)\n {\n \tchar arg[PATH_MAX];\n-- \n1.8.4.38.g317e65b\n"},{"id":"226813","messageId":"1378362001-1738-16-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 15/38] pack v4: object data copy","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:38Z","receivedAt":"2013-09-05T06:19:38Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Blob and tag objects have no particular changes except for their object\nheader.\n\nDelta objects are also copied as is, except for their delta base reference\nwhich is converted to the new way as used elsewhere in pack v4 encoding\ni.e. an index into the SHA1 table or a literal SHA1 prefixed by 0 if not\nfound in the table (see encode_sha1ref).  This is true for both REF_DELTA\nas well as OFS_DELTA.\n\nObject payload is validated against the recorded CRC32 in the source\npack index file when possible before being copied.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 60 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 60 insertions(+)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 6098062..b0e344f 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -12,6 +12,7 @@\n #include \"object.h\"\n #include \"tree-walk.h\"\n #include \"pack.h\"\n+#include \"pack-revindex.h\"\n #include \"varint.h\"\n \n \n@@ -662,6 +663,65 @@ static int write_object_header(struct sha1file *f, enum object_type type, unsign\n \treturn len;\n }\n \n+static unsigned long copy_object_data(struct sha1file *f, struct packed_git *p,\n+\t\t\t\t      off_t offset)\n+{\n+\tstruct pack_window *w_curs = NULL;\n+\tstruct revindex_entry *revidx;\n+\tenum object_type type;\n+\tunsigned long avail, size, datalen, written;\n+\tint hdrlen, reflen, idx_nr;\n+\tunsigned char *src, buf[24];\n+\n+\trevidx = find_pack_revindex(p, offset);\n+\tidx_nr = revidx->nr;\n+\tdatalen = revidx[1].offset - offset;\n+\n+\tsrc = use_pack(p, &w_curs, offset, &avail);\n+\thdrlen = unpack_object_header_buffer(src, avail, &type, &size);\n+\n+\twritten = write_object_header(f, type, size);\n+\n+\tif (type == OBJ_OFS_DELTA) {\n+\t\tconst unsigned char *cp = src + hdrlen;\n+\t\toff_t base_offset = decode_varint(&cp);\n+\t\thdrlen = cp - src;\n+\t\tbase_offset = offset - base_offset;\n+\t\tif (base_offset <= 0 || base_offset >= offset)\n+\t\t\tdie(\"delta offset out of bound\");\n+\t\trevidx = find_pack_revindex(p, base_offset);\n+\t\treflen = encode_sha1ref(nth_packed_object_sha1(p, revidx->nr), buf);\n+\t\tsha1write(f, buf, reflen);\n+\t\twritten += reflen;\n+\t} else if (type == OBJ_REF_DELTA) {\n+\t\treflen = encode_sha1ref(src + hdrlen, buf);\n+\t\thdrlen += 20;\n+\t\tsha1write(f, buf, reflen);\n+\t\twritten += reflen;\n+\t}\n+\n+\tif (p->index_version > 1 &&\n+\t    check_pack_crc(p, &w_curs, offset, datalen, idx_nr))\n+\t\tdie(\"bad CRC for object at offset %\"PRIuMAX\" in %s\",\n+\t\t    (uintmax_t)offset, p->pack_name);\n+\n+\toffset += hdrlen;\n+\tdatalen -= hdrlen;\n+\n+\twhile (datalen) {\n+\t\tsrc = use_pack(p, &w_curs, offset, &avail);\n+\t\tif (avail > datalen)\n+\t\t\tavail = datalen;\n+\t\tsha1write(f, src, avail);\n+\t\twritten += avail;\n+\t\toffset += avail;\n+\t\tdatalen -= avail;\n+\t}\n+\tunuse_pack(&w_curs);\n+\n+\treturn written;\n+}\n+\n static struct packed_git *open_pack(const char *path)\n {\n \tchar arg[PATH_MAX];\n-- \n1.8.4.38.g317e65b\n"},{"id":"226815","messageId":"1378362001-1738-17-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 16/38] pack v4: object writing","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:39Z","receivedAt":"2013-09-05T06:19:39Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"This adds the missing code to finally be able to produce a complete\npack file version 4.  We trap commit and tree objects as those have\na completely new encoding.  Other object types are copied almost\nunchanged.\n\nAs we go the pack index entries are updated  in place to store the new\nobject offsets once they're written to the destination file.  This will\nbe needed later for writing the pack index file.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 74 ++++++++++++++++++++++++++++++++++++++++++++++++++++++---\n 1 file changed, 71 insertions(+), 3 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex b0e344f..5d76234 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -722,6 +722,59 @@ static unsigned long copy_object_data(struct sha1file *f, struct packed_git *p,\n \treturn written;\n }\n \n+static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n+\t\t\t\t struct pack_idx_entry *obj)\n+{\n+\tvoid *src, *result;\n+\tstruct object_info oi = {};\n+\tenum object_type type;\n+\tunsigned long size;\n+\tunsigned int hdrlen;\n+\n+\toi.typep = &type;\n+\toi.sizep = &size;\n+\tif (packed_object_info(p, obj->offset, &oi) < 0)\n+\t\tdie(\"cannot get type of %s from %s\",\n+\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\n+\t/* Some objects are copied without decompression */\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TREE:\n+\t\tbreak;\n+\tdefault:\n+\t\treturn copy_object_data(f, p, obj->offset);\n+\t}\n+\n+\t/* The rest is converted into their new format */\n+\tsrc = unpack_entry(p, obj->offset, &type, &size);\n+\tif (!src)\n+\t\tdie(\"cannot unpack %s from %s\",\n+\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\tif (check_sha1_signature(obj->sha1, src, size, typename(type)))\n+\t\tdie(\"packed %s from %s is corrupt\",\n+\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\n+\thdrlen = write_object_header(f, type, size);\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\t\tresult = pv4_encode_commit(src, &size);\n+\t\tbreak;\n+\tcase OBJ_TREE:\n+\t\tresult = pv4_encode_tree(src, &size);\n+\t\tbreak;\n+\tdefault:\n+\t\tdie(\"unexpected object type %d\", type);\n+\t}\n+\tfree(src);\n+\tif (!result)\n+\t\tdie(\"can't convert %s object %s\",\n+\t\t    typename(type), sha1_to_hex(obj->sha1));\n+\tsha1write(f, result, size);\n+\tfree(result);\n+\treturn hdrlen + size;\n+}\n+\n static struct packed_git *open_pack(const char *path)\n {\n \tchar arg[PATH_MAX];\n@@ -780,7 +833,8 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \tstruct packed_git *p;\n \tstruct sha1file *f;\n \tstruct pack_idx_entry *objs, **p_objs;\n-\tunsigned nr_objects;\n+\tunsigned i, nr_objects;\n+\toff_t written = 0;\n \n \tp = open_pack(src_pack);\n \tif (!p)\n@@ -791,12 +845,26 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \tp_objs = sort_objs_by_offset(objs, nr_objects);\n \n \tcreate_pack_dictionaries(p, p_objs);\n+\tsort_dict_entries_by_hits(commit_name_table);\n+\tsort_dict_entries_by_hits(tree_path_table);\n \n \tf = packv4_open(dst_pack);\n \tif (!f)\n \t\tdie(\"unable to open destination pack\");\n-\tpackv4_write_header(f, nr_objects);\n-\tpackv4_write_tables(f, nr_objects, objs);\n+\twritten += packv4_write_header(f, nr_objects);\n+\twritten += packv4_write_tables(f, nr_objects, objs);\n+\n+\t/* Let's write objects out, updating the object index list in place */\n+\tall_objs = objs;\n+\tall_objs_nr = nr_objects;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\toff_t obj_pos = written;\n+\t\tstruct pack_idx_entry *obj = p_objs[i];\n+\t\twritten += packv4_write_object(f, p, obj);\n+\t\tobj->offset = obj_pos;\n+\t}\n+\n+\tsha1close(f, NULL, CSUM_CLOSE | CSUM_FSYNC);\n }\n \n int main(int argc, char *argv[])\n-- \n1.8.4.38.g317e65b\n"},{"id":"226816","messageId":"1378362001-1738-18-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 17/38] pack v4: tree object delta encoding","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:40Z","receivedAt":"2013-09-05T06:19:40Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"In order to be able to quickly walk tree objects, let's encode their\n\"delta\" as a range of entries into another tree object.\n\nIn order to discriminate between a copy sequence from a regular entry,\nthe entry index LSB is reserved to indicate a copy sequence.  Therefore\nthe actual index of a path component is shifted left one bit.\n\nThe encoding allows for the base object to change so multiple base\nobjects can be borrowed from.  The code doesn't try to exploit this\npossibility at the moment though.\n\nThe code isn't optimal at the moment as it doesn't consider the case\nwhere a copy sequence could be larger than the local sequence it\nmeans to replace.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 108 +++++++++++++++++++++++++++++++++++++++++++++++++++++---\n 1 file changed, 103 insertions(+), 5 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 5d76234..6830a0a 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -394,24 +394,53 @@ bad:\n \treturn NULL;\n }\n \n+static int compare_tree_entries(struct name_entry *e1, struct name_entry *e2)\n+{\n+\tint len1 = tree_entry_len(e1);\n+\tint len2 = tree_entry_len(e2);\n+\tint len = len1 < len2 ? len1 : len2;\n+\tunsigned char c1, c2;\n+\tint cmp;\n+\n+\tcmp = memcmp(e1->path, e2->path, len);\n+\tif (cmp)\n+\t\treturn cmp;\n+\tc1 = e1->path[len];\n+\tc2 = e2->path[len];\n+\tif (!c1 && S_ISDIR(e1->mode))\n+\t\tc1 = '/';\n+\tif (!c2 && S_ISDIR(e2->mode))\n+\t\tc2 = '/';\n+\treturn c1 - c2;\n+}\n+\n /*\n  * This converts a canonical tree object buffer into its\n  * tightly packed representation using the already populated\n  * and sorted tree_path_table dictionary.  The parsing is\n  * strict so to ensure the canonical version may always be\n  * regenerated and produce the same hash.\n+ *\n+ * If a delta buffer is provided, we may encode multiple ranges of tree\n+ * entries against that buffer.\n  */\n-void *pv4_encode_tree(void *_buffer, unsigned long *sizep)\n+void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n+\t\t      void *delta, unsigned long delta_size,\n+\t\t      const unsigned char *delta_sha1)\n {\n \tunsigned long size = *sizep;\n \tunsigned char *in, *out, *end, *buffer = _buffer;\n-\tstruct tree_desc desc;\n-\tstruct name_entry name_entry;\n+\tstruct tree_desc desc, delta_desc;\n+\tstruct name_entry name_entry, delta_entry;\n \tint nb_entries;\n+\tunsigned int copy_start, copy_count = 0, delta_pos = 0, first_delta = 1;\n \n \tif (!size)\n \t\treturn NULL;\n \n+\tif (!delta_size)\n+\t\tdelta = NULL;\n+\n \t/*\n \t * We can't make sure the result will always be smaller than the\n \t * input. The smallest possible entry is \"0 x\\0<40 byte SHA1>\"\n@@ -434,9 +463,42 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep)\n \tout += encode_varint(nb_entries, out);\n \n \tinit_tree_desc(&desc, in, size);\n+\tif (delta) {\n+\t\tinit_tree_desc(&delta_desc, delta, delta_size);\n+\t\tif (!tree_entry(&delta_desc, &delta_entry))\n+\t\t\tdelta = NULL;\n+\t}\n+\n \twhile (tree_entry(&desc, &name_entry)) {\n \t\tint pathlen, index;\n \n+\t\t/*\n+\t\t * Try to match entries against our delta object.\n+\t\t */\n+\t\tif (delta) {\n+\t\t\tint ret;\n+\n+\t\t\tdo {\n+\t\t\t\tret = compare_tree_entries(&name_entry, &delta_entry);\n+\t\t\t\tif (ret <= 0 || copy_count != 0)\n+\t\t\t\t\tbreak;\n+\t\t\t\tdelta_pos++;\n+\t\t\t\tif (!tree_entry(&delta_desc, &delta_entry))\n+\t\t\t\t\tdelta = NULL;\n+\t\t\t} while (delta);\n+\n+\t\t\tif (ret == 0 && name_entry.mode == delta_entry.mode &&\n+\t\t\t    hashcmp(name_entry.sha1, delta_entry.sha1) == 0) {\n+\t\t\t\tif (!copy_count)\n+\t\t\t\t\tcopy_start = delta_pos;\n+\t\t\t\tcopy_count++;\n+\t\t\t\tdelta_pos++;\n+\t\t\t\tif (!tree_entry(&delta_desc, &delta_entry))\n+\t\t\t\t\tdelta = NULL;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\n \t\tif (end - out < 48) {\n \t\t\tunsigned long sofar = out - buffer;\n \t\t\tbuffer = xrealloc(buffer, (sofar + 48)*2);\n@@ -444,6 +506,32 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep)\n \t\t\tout = buffer + sofar;\n \t\t}\n \n+\t\tif (copy_count) {\n+\t\t\t/*\n+\t\t\t * Let's write a sequence indicating we're copying\n+\t\t\t * entries from another object:\n+\t\t\t *\n+\t\t\t * entry_start + entry_count + object_ref\n+\t\t\t *\n+\t\t\t * To distinguish between 'entry_start' and an actual\n+\t\t\t * entry index, we use the LSB = 1.\n+\t\t\t *\n+\t\t\t * Furthermore, if object_ref is the same as the\n+\t\t\t * preceding one, we can omit it and save some\n+\t\t\t * more space, especially if that ends up being a\n+\t\t\t * full sha1 reference.  Let's steal the LSB\n+\t\t\t * of entry_count for that purpose.\n+\t\t\t */\n+\t\t\tcopy_start = (copy_start << 1) | 1;\n+\t\t\tcopy_count = (copy_count << 1) | first_delta;\n+\t\t\tout += encode_varint(copy_start, out);\n+\t\t\tout += encode_varint(copy_count, out);\n+\t\t\tif (first_delta)\n+\t\t\t\tout += encode_sha1ref(delta_sha1, out);\n+\t\t\tcopy_count = 0;\n+\t\t\tfirst_delta = 0;\n+\t\t}\n+\n \t\tpathlen = tree_entry_len(&name_entry);\n \t\tindex = dict_add_entry(tree_path_table, name_entry.mode,\n \t\t\t\t       name_entry.path, pathlen);\n@@ -452,10 +540,20 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep)\n \t\t\tfree(buffer);\n \t\t\treturn NULL;\n \t\t}\n-\t\tout += encode_varint(index, out);\n+\t\tout += encode_varint(index << 1, out);\n \t\tout += encode_sha1ref(name_entry.sha1, out);\n \t}\n \n+\tif (copy_count) {\n+\t\t/* flush the trailing copy */\n+\t\tcopy_start = (copy_start << 1) | 1;\n+\t\tcopy_count = (copy_count << 1) | first_delta;\n+\t\tout += encode_varint(copy_start, out);\n+\t\tout += encode_varint(copy_count, out);\n+\t\tif (first_delta)\n+\t\t\tout += encode_sha1ref(delta_sha1, out);\n+\t}\n+\n \t*sizep = out - buffer;\n \treturn buffer;\n }\n@@ -761,7 +859,7 @@ static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \t\tresult = pv4_encode_commit(src, &size);\n \t\tbreak;\n \tcase OBJ_TREE:\n-\t\tresult = pv4_encode_tree(src, &size);\n+\t\tresult = pv4_encode_tree(src, &size, NULL, 0, NULL);\n \t\tbreak;\n \tdefault:\n \t\tdie(\"unexpected object type %d\", type);\n-- \n1.8.4.38.g317e65b\n"},{"id":"226809","messageId":"1378362001-1738-19-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 18/38] pack v4: load delta candidate for encoding tree objects","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:41Z","receivedAt":"2013-09-05T06:19:41Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"The SHA1 of the base object is retrieved and the corresponding object\nis loaded in memory for pv4_encode_tree() to look at.  Simple but\neffective.  Obviously this relies on the delta matching already performed\nduring the pack v3 delta search.  Some native delta search for pack v4\ncould be investigated eventually.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 63 ++++++++++++++++++++++++++++++++++++++++++++++++++++++---\n 1 file changed, 60 insertions(+), 3 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 6830a0a..15c5959 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -820,18 +820,56 @@ static unsigned long copy_object_data(struct sha1file *f, struct packed_git *p,\n \treturn written;\n }\n \n+static unsigned char *get_delta_base(struct packed_git *p, off_t offset,\n+\t\t\t\t     unsigned char *sha1_buf)\n+{\n+\tstruct pack_window *w_curs = NULL;\n+\tenum object_type type;\n+\tunsigned long avail, size;\n+\tint hdrlen;\n+\tunsigned char *src;\n+\tconst unsigned char *base_sha1 = NULL; ;\n+\n+\tsrc = use_pack(p, &w_curs, offset, &avail);\n+\thdrlen = unpack_object_header_buffer(src, avail, &type, &size);\n+\n+\tif (type == OBJ_OFS_DELTA) {\n+\t\tconst unsigned char *cp = src + hdrlen;\n+\t\toff_t base_offset = decode_varint(&cp);\n+\t\tbase_offset = offset - base_offset;\n+\t\tif (base_offset <= 0 || base_offset >= offset) {\n+\t\t\terror(\"delta offset out of bound\");\n+\t\t} else {\n+\t\t\tstruct revindex_entry *revidx;\n+\t\t\trevidx = find_pack_revindex(p, base_offset);\n+\t\t\tbase_sha1 = nth_packed_object_sha1(p, revidx->nr);\n+\t\t}\n+\t} else if (type == OBJ_REF_DELTA) {\n+\t\tbase_sha1 = src + hdrlen;\n+\t} else\n+\t\terror(\"expected to get a delta but got a %s\", typename(type));\n+\n+\tunuse_pack(&w_curs);\n+\n+\tif (!base_sha1)\n+\t\treturn NULL;\n+\thashcpy(sha1_buf, base_sha1);\n+\treturn sha1_buf;\n+}\n+\n static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \t\t\t\t struct pack_idx_entry *obj)\n {\n \tvoid *src, *result;\n \tstruct object_info oi = {};\n-\tenum object_type type;\n+\tenum object_type type, packed_type;\n \tunsigned long size;\n \tunsigned int hdrlen;\n \n \toi.typep = &type;\n \toi.sizep = &size;\n-\tif (packed_object_info(p, obj->offset, &oi) < 0)\n+\tpacked_type = packed_object_info(p, obj->offset, &oi);\n+\tif (packed_type < 0)\n \t\tdie(\"cannot get type of %s from %s\",\n \t\t    sha1_to_hex(obj->sha1), p->pack_name);\n \n@@ -859,7 +897,26 @@ static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \t\tresult = pv4_encode_commit(src, &size);\n \t\tbreak;\n \tcase OBJ_TREE:\n-\t\tresult = pv4_encode_tree(src, &size, NULL, 0, NULL);\n+\t\tif (packed_type != OBJ_TREE) {\n+\t\t\tunsigned char sha1_buf[20], *ref_sha1;\n+\t\t\tvoid *ref;\n+\t\t\tenum object_type ref_type;\n+\t\t\tunsigned long ref_size;\n+\n+\t\t\tref_sha1 = get_delta_base(p, obj->offset, sha1_buf);\n+\t\t\tif (!ref_sha1)\n+\t\t\t\tdie(\"unable to get delta base sha1 for %s\",\n+\t\t\t\t\t\tsha1_to_hex(obj->sha1));\n+\t\t\tref = read_sha1_file(ref_sha1, &ref_type, &ref_size);\n+\t\t\tif (!ref || ref_type != OBJ_TREE)\n+\t\t\t\tdie(\"cannot obtain delta base for %s\",\n+\t\t\t\t\t\tsha1_to_hex(obj->sha1));\n+\t\t\tresult = pv4_encode_tree(src, &size,\n+\t\t\t\t\t\t ref, ref_size, ref_sha1);\n+\t\t\tfree(ref);\n+\t\t} else {\n+\t\t\tresult = pv4_encode_tree(src, &size, NULL, 0, NULL);\n+\t\t}\n \t\tbreak;\n \tdefault:\n \t\tdie(\"unexpected object type %d\", type);\n-- \n1.8.4.38.g317e65b\n"},{"id":"226811","messageId":"1378362001-1738-20-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 19/38] packv4-create: optimize delta encoding","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:42Z","receivedAt":"2013-09-05T06:19:42Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Make sure the copy sequence is smaller than the list of tree entries it\nis meant to replace.  We do so by encoding tree entries in parallel with\nthe delta entry comparison, and then comparing the length of both\nsequences.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 65 +++++++++++++++++++++++++++++++++++++++------------------\n 1 file changed, 45 insertions(+), 20 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 15c5959..c8d3053 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -433,7 +433,8 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \tstruct tree_desc desc, delta_desc;\n \tstruct name_entry name_entry, delta_entry;\n \tint nb_entries;\n-\tunsigned int copy_start, copy_count = 0, delta_pos = 0, first_delta = 1;\n+\tunsigned int copy_start = 0, copy_count = 0, copy_pos = 0, copy_end = 0;\n+\tunsigned int delta_pos = 0, first_delta = 1;\n \n \tif (!size)\n \t\treturn NULL;\n@@ -489,24 +490,23 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \n \t\t\tif (ret == 0 && name_entry.mode == delta_entry.mode &&\n \t\t\t    hashcmp(name_entry.sha1, delta_entry.sha1) == 0) {\n-\t\t\t\tif (!copy_count)\n+\t\t\t\tif (!copy_count) {\n \t\t\t\t\tcopy_start = delta_pos;\n+\t\t\t\t\tcopy_pos = out - buffer;\n+\t\t\t\t\tcopy_end = 0;\n+\t\t\t\t}\n \t\t\t\tcopy_count++;\n \t\t\t\tdelta_pos++;\n \t\t\t\tif (!tree_entry(&delta_desc, &delta_entry))\n \t\t\t\t\tdelta = NULL;\n-\t\t\t\tcontinue;\n-\t\t\t}\n-\t\t}\n+\t\t\t} else\n+\t\t\t\tcopy_end = 1;\n+\t\t} else\n+\t\t\tcopy_end = 1;\n \n-\t\tif (end - out < 48) {\n-\t\t\tunsigned long sofar = out - buffer;\n-\t\t\tbuffer = xrealloc(buffer, (sofar + 48)*2);\n-\t\t\tend = buffer + (sofar + 48)*2;\n-\t\t\tout = buffer + sofar;\n-\t\t}\n+\t\tif (copy_count && copy_end) {\n+\t\t\tunsigned char copy_buf[48], *cp = copy_buf;\n \n-\t\tif (copy_count) {\n \t\t\t/*\n \t\t\t * Let's write a sequence indicating we're copying\n \t\t\t * entries from another object:\n@@ -524,12 +524,31 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\t\t */\n \t\t\tcopy_start = (copy_start << 1) | 1;\n \t\t\tcopy_count = (copy_count << 1) | first_delta;\n-\t\t\tout += encode_varint(copy_start, out);\n-\t\t\tout += encode_varint(copy_count, out);\n+\t\t\tcp += encode_varint(copy_start, cp);\n+\t\t\tcp += encode_varint(copy_count, cp);\n \t\t\tif (first_delta)\n-\t\t\t\tout += encode_sha1ref(delta_sha1, out);\n+\t\t\t\tcp += encode_sha1ref(delta_sha1, cp);\n \t\t\tcopy_count = 0;\n-\t\t\tfirst_delta = 0;\n+\n+\t\t\t/*\n+\t\t\t * Now let's make sure this is going to take less\n+\t\t\t * space than the corresponding direct entries we've\n+\t\t\t * created in parallel.  If so we dump the copy\n+\t\t\t * sequence over those entries in the output buffer.\n+\t\t\t */\n+\t\t\tif (cp - copy_buf < out - &buffer[copy_pos]) {\n+\t\t\t\tout = buffer + copy_pos;\n+\t\t\t\tmemcpy(out, copy_buf, cp - copy_buf);\n+\t\t\t\tout += cp - copy_buf;\n+\t\t\t\tfirst_delta = 0;\n+\t\t\t}\n+\t\t}\n+\n+\t\tif (end - out < 48) {\n+\t\t\tunsigned long sofar = out - buffer;\n+\t\t\tbuffer = xrealloc(buffer, (sofar + 48)*2);\n+\t\t\tend = buffer + (sofar + 48)*2;\n+\t\t\tout = buffer + sofar;\n \t\t}\n \n \t\tpathlen = tree_entry_len(&name_entry);\n@@ -545,13 +564,19 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t}\n \n \tif (copy_count) {\n-\t\t/* flush the trailing copy */\n+\t\t/* process the trailing copy */\n+\t\tunsigned char copy_buf[48], *cp = copy_buf;\n \t\tcopy_start = (copy_start << 1) | 1;\n \t\tcopy_count = (copy_count << 1) | first_delta;\n-\t\tout += encode_varint(copy_start, out);\n-\t\tout += encode_varint(copy_count, out);\n+\t\tcp += encode_varint(copy_start, cp);\n+\t\tcp += encode_varint(copy_count, cp);\n \t\tif (first_delta)\n-\t\t\tout += encode_sha1ref(delta_sha1, out);\n+\t\t\tcp += encode_sha1ref(delta_sha1, cp);\n+\t\tif (cp - copy_buf < out - &buffer[copy_pos]) {\n+\t\t\tout = buffer + copy_pos;\n+\t\t\tmemcpy(out, copy_buf, cp - copy_buf);\n+\t\t\tout += cp - copy_buf;\n+\t\t}\n \t}\n \n \t*sizep = out - buffer;\n-- \n1.8.4.38.g317e65b\n"},{"id":"226810","messageId":"1378362001-1738-21-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 20/38] pack v4: honor pack.compression config option","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:43Z","receivedAt":"2013-09-05T06:19:43Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 19 +++++++++++++++++++\n 1 file changed, 19 insertions(+)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex c8d3053..45f8427 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -16,6 +16,7 @@\n #include \"varint.h\"\n \n \n+static int pack_compression_seen;\n static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n \n struct data_entry {\n@@ -1047,12 +1048,30 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \tsha1close(f, NULL, CSUM_CLOSE | CSUM_FSYNC);\n }\n \n+static int git_pack_config(const char *k, const char *v, void *cb)\n+{\n+\tif (!strcmp(k, \"pack.compression\")) {\n+\t\tint level = git_config_int(k, v);\n+\t\tif (level == -1)\n+\t\t\tlevel = Z_DEFAULT_COMPRESSION;\n+\t\telse if (level < 0 || level > Z_BEST_COMPRESSION)\n+\t\t\tdie(\"bad pack compression level %d\", level);\n+\t\tpack_compression_level = level;\n+\t\tpack_compression_seen = 1;\n+\t\treturn 0;\n+\t}\n+\treturn git_default_config(k, v, cb);\n+}\n+\n int main(int argc, char *argv[])\n {\n \tif (argc != 3) {\n \t\tfprintf(stderr, \"Usage: %s <src_packfile> <dst_packfile>\\n\", argv[0]);\n \t\texit(1);\n \t}\n+\tgit_config(git_pack_config, NULL);\n+\tif (!pack_compression_seen && core_compression_seen)\n+\t\tpack_compression_level = core_compression_level;\n \tprocess_one_pack(argv[1], argv[2]);\n \tif (0)\n \t\tdict_dump();\n-- \n1.8.4.38.g317e65b\n"},{"id":"226807","messageId":"1378362001-1738-22-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 21/38] pack v4: relax commit parsing a bit","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:44Z","receivedAt":"2013-09-05T06:19:44Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"At least commit af25e94d4dcfb9608846242fabdd4e6014e5c9f0 in the Linux\nkernel repository has \"author  <> 1120285620 -0700\"\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 45f8427..a9e9002 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -158,12 +158,12 @@ static char *get_nameend_and_tz(char *from, int *tz_val)\n \tchar *end, *tz;\n \n \ttz = strchr(from, '\\n');\n-\t/* let's assume the smallest possible string to be \"x <x> 0 +0000\\n\" */\n-\tif (!tz || tz - from < 13)\n+\t/* let's assume the smallest possible string to be \" <> 0 +0000\\n\" */\n+\tif (!tz || tz - from < 11)\n \t\treturn NULL;\n \ttz -= 4;\n \tend = tz - 4;\n-\twhile (end - from > 5 && *end != ' ')\n+\twhile (end - from > 3 && *end != ' ')\n \t\tend--;\n \tif (end[-1] != '>' || end[0] != ' ' || tz[-2] != ' ')\n \t\treturn NULL;\n-- \n1.8.4.38.g317e65b\n"},{"id":"226804","messageId":"1378362001-1738-23-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 22/38] pack index v3","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:45Z","receivedAt":"2013-09-05T06:19:45Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"This is a minor change over pack index v2.  Since pack v4 already contains\nthe sorted SHA1 table, it is therefore ommitted from the index file.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n pack-write.c    |  6 +++++-\n packv4-create.c | 10 +++++++++-\n 2 files changed, 14 insertions(+), 2 deletions(-)\n\ndiff --git a/pack-write.c b/pack-write.c\nindex ca9e63b..631007e 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -87,6 +87,8 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \n \t/* if last object's offset is >= 2^31 we should use index V2 */\n \tindex_version = need_large_offset(last_obj_offset, opts) ? 2 : opts->version;\n+\tif (index_version < opts->version)\n+\t\tindex_version = opts->version;\n \n \t/* index versions 2 and above need a header */\n \tif (index_version >= 2) {\n@@ -127,7 +129,9 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \t\t\tuint32_t offset = htonl(obj->offset);\n \t\t\tsha1write(f, &offset, 4);\n \t\t}\n-\t\tsha1write(f, obj->sha1, 20);\n+\t\t/* Pack v4 (using index v3) carries the SHA1 table already */\n+\t\tif (index_version < 3)\n+\t\t\tsha1write(f, obj->sha1, 20);\n \t\tgit_SHA1_Update(&ctx, obj->sha1, 20);\n \t\tif ((opts->flags & WRITE_IDX_STRICT) &&\n \t\t    (i && !hashcmp(list[-2]->sha1, obj->sha1)))\ndiff --git a/packv4-create.c b/packv4-create.c\nindex a9e9002..22cdf8e 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -1014,8 +1014,10 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \tstruct packed_git *p;\n \tstruct sha1file *f;\n \tstruct pack_idx_entry *objs, **p_objs;\n+\tstruct pack_idx_option idx_opts;\n \tunsigned i, nr_objects;\n \toff_t written = 0;\n+\tunsigned char pack_sha1[20];\n \n \tp = open_pack(src_pack);\n \tif (!p)\n@@ -1041,11 +1043,17 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \tfor (i = 0; i < nr_objects; i++) {\n \t\toff_t obj_pos = written;\n \t\tstruct pack_idx_entry *obj = p_objs[i];\n+\t\tcrc32_begin(f);\n \t\twritten += packv4_write_object(f, p, obj);\n \t\tobj->offset = obj_pos;\n+\t\tobj->crc32 = crc32_end(f);\n \t}\n \n-\tsha1close(f, NULL, CSUM_CLOSE | CSUM_FSYNC);\n+\tsha1close(f, pack_sha1, CSUM_CLOSE | CSUM_FSYNC);\n+\n+\treset_pack_idx_option(&idx_opts);\n+\tidx_opts.version = 3;\n+\twrite_idx_file(dst_pack, p_objs, nr_objects, &idx_opts, pack_sha1);\n }\n \n static int git_pack_config(const char *k, const char *v, void *cb)\n-- \n1.8.4.38.g317e65b\n"},{"id":"226803","messageId":"1378362001-1738-24-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 23/38] packv4-create: normalize pack name to properly generate the pack index file name","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:46Z","receivedAt":"2013-09-05T06:19:46Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 73 +++++++++++++++++++++++++++------------------------------\n 1 file changed, 34 insertions(+), 39 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 22cdf8e..c23c791 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -956,56 +956,46 @@ static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \treturn hdrlen + size;\n }\n \n-static struct packed_git *open_pack(const char *path)\n+static char *normalize_pack_name(const char *path)\n {\n-\tchar arg[PATH_MAX];\n+\tchar buf[PATH_MAX];\n \tint len;\n-\tstruct packed_git *p;\n \n-\tlen = strlcpy(arg, path, PATH_MAX);\n-\tif (len >= PATH_MAX) {\n-\t\terror(\"name too long: %s\", path);\n-\t\treturn NULL;\n-\t}\n+\tlen = strlcpy(buf, path, PATH_MAX);\n+\tif (len >= PATH_MAX - 6)\n+\t\tdie(\"name too long: %s\", path);\n \n \t/*\n \t * In addition to \"foo.idx\" we accept \"foo.pack\" and \"foo\";\n-\t * normalize these forms to \"foo.idx\" for add_packed_git().\n+\t * normalize these forms to \"foo.pack\".\n \t */\n-\tif (has_extension(arg, \".pack\")) {\n-\t\tstrcpy(arg + len - 5, \".idx\");\n-\t\tlen--;\n-\t} else if (!has_extension(arg, \".idx\")) {\n-\t\tif (len + 4 >= PATH_MAX) {\n-\t\t\terror(\"name too long: %s.idx\", arg);\n-\t\t\treturn NULL;\n-\t\t}\n-\t\tstrcpy(arg + len, \".idx\");\n-\t\tlen += 4;\n+\tif (has_extension(buf, \".idx\")) {\n+\t\tstrcpy(buf + len - 4, \".pack\");\n+\t\tlen++;\n+\t} else if (!has_extension(buf, \".pack\")) {\n+\t\tstrcpy(buf + len, \".pack\");\n+\t\tlen += 5;\n \t}\n \n-\t/*\n-\t * add_packed_git() uses our buffer (containing \"foo.idx\") to\n-\t * build the pack filename (\"foo.pack\").  Make sure it fits.\n-\t */\n-\tif (len + 1 >= PATH_MAX) {\n-\t\targ[len - 4] = '\\0';\n-\t\terror(\"name too long: %s.pack\", arg);\n-\t\treturn NULL;\n-\t}\n+\treturn xstrdup(buf);\n+}\n \n-\tp = add_packed_git(arg, len, 1);\n-\tif (!p) {\n-\t\terror(\"packfile %s not found.\", arg);\n-\t\treturn NULL;\n-\t}\n+static struct packed_git *open_pack(const char *path)\n+{\n+\tchar *packname = normalize_pack_name(path);\n+\tint len = strlen(packname);\n+\tstruct packed_git *p;\n+\n+\tstrcpy(packname + len - 5, \".idx\");\n+\tp = add_packed_git(packname, len - 1, 1);\n+\tif (!p)\n+\t\tdie(\"packfile %s not found.\", packname);\n \n \tinstall_packed_git(p);\n-\tif (open_pack_index(p)) {\n-\t\terror(\"packfile %s index not opened\", p->pack_name);\n-\t\treturn NULL;\n-\t}\n+\tif (open_pack_index(p))\n+\t\tdie(\"packfile %s index not opened\", p->pack_name);\n \n+\tfree(packname);\n \treturn p;\n }\n \n@@ -1017,6 +1007,7 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \tstruct pack_idx_option idx_opts;\n \tunsigned i, nr_objects;\n \toff_t written = 0;\n+\tchar *packname;\n \tunsigned char pack_sha1[20];\n \n \tp = open_pack(src_pack);\n@@ -1031,7 +1022,8 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \tsort_dict_entries_by_hits(commit_name_table);\n \tsort_dict_entries_by_hits(tree_path_table);\n \n-\tf = packv4_open(dst_pack);\n+\tpackname = normalize_pack_name(dst_pack);\n+\tf = packv4_open(packname);\n \tif (!f)\n \t\tdie(\"unable to open destination pack\");\n \twritten += packv4_write_header(f, nr_objects);\n@@ -1053,7 +1045,10 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \n \treset_pack_idx_option(&idx_opts);\n \tidx_opts.version = 3;\n-\twrite_idx_file(dst_pack, p_objs, nr_objects, &idx_opts, pack_sha1);\n+\tstrcpy(packname + strlen(packname) - 5, \".idx\");\n+\twrite_idx_file(packname, p_objs, nr_objects, &idx_opts, pack_sha1);\n+\n+\tfree(packname);\n }\n \n static int git_pack_config(const char *k, const char *v, void *cb)\n-- \n1.8.4.38.g317e65b\n"},{"id":"226806","messageId":"1378362001-1738-25-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 24/38] packv4-create: add progress display","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:47Z","receivedAt":"2013-09-05T06:19:47Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 10 ++++++++++\n 1 file changed, 10 insertions(+)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex c23c791..fd16222 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -13,6 +13,7 @@\n #include \"tree-walk.h\"\n #include \"pack.h\"\n #include \"pack-revindex.h\"\n+#include \"progress.h\"\n #include \"varint.h\"\n \n \n@@ -627,8 +628,10 @@ static struct pack_idx_entry **sort_objs_by_offset(struct pack_idx_entry *list,\n static int create_pack_dictionaries(struct packed_git *p,\n \t\t\t\t    struct pack_idx_entry **obj_list)\n {\n+\tstruct progress *progress_state;\n \tunsigned int i;\n \n+\tprogress_state = start_progress(\"Scanning objects\", p->num_objects);\n \tfor (i = 0; i < p->num_objects; i++) {\n \t\tstruct pack_idx_entry *obj = obj_list[i];\n \t\tvoid *data;\n@@ -637,6 +640,8 @@ static int create_pack_dictionaries(struct packed_git *p,\n \t\tstruct object_info oi = {};\n \t\tint (*add_dict_entries)(void *, unsigned long);\n \n+\t\tdisplay_progress(progress_state, i+1);\n+\n \t\toi.typep = &type;\n \t\toi.sizep = &size;\n \t\tif (packed_object_info(p, obj->offset, &oi) < 0)\n@@ -666,6 +671,7 @@ static int create_pack_dictionaries(struct packed_git *p,\n \t\tfree(data);\n \t}\n \n+\tstop_progress(&progress_state);\n \treturn 0;\n }\n \n@@ -1009,6 +1015,7 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \toff_t written = 0;\n \tchar *packname;\n \tunsigned char pack_sha1[20];\n+\tstruct progress *progress_state;\n \n \tp = open_pack(src_pack);\n \tif (!p)\n@@ -1030,6 +1037,7 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \twritten += packv4_write_tables(f, nr_objects, objs);\n \n \t/* Let's write objects out, updating the object index list in place */\n+\tprogress_state = start_progress(\"Writing objects\", nr_objects);\n \tall_objs = objs;\n \tall_objs_nr = nr_objects;\n \tfor (i = 0; i < nr_objects; i++) {\n@@ -1039,7 +1047,9 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \t\twritten += packv4_write_object(f, p, obj);\n \t\tobj->offset = obj_pos;\n \t\tobj->crc32 = crc32_end(f);\n+\t\tdisplay_progress(progress_state, i+1);\n \t}\n+\tstop_progress(&progress_state);\n \n \tsha1close(f, pack_sha1, CSUM_CLOSE | CSUM_FSYNC);\n \n-- \n1.8.4.38.g317e65b\n"},{"id":"226799","messageId":"1378362001-1738-26-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 25/38] pack v4: initial pack index v3 support on the read side","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:48Z","receivedAt":"2013-09-05T06:19:48Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"A bit crud but good enough for now.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n cache.h         |  1 +\n pack-check.c    |  4 +++-\n pack-revindex.c |  7 ++++---\n sha1_file.c     | 56 ++++++++++++++++++++++++++++++++++++++++++++++++++------\n 4 files changed, 58 insertions(+), 10 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 4231dfa..c939b60 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1021,6 +1021,7 @@ extern struct packed_git {\n \toff_t pack_size;\n \tconst void *index_data;\n \tsize_t index_size;\n+\tconst unsigned char *sha1_table;\n \tuint32_t num_objects;\n \tuint32_t num_bad_objects;\n \tunsigned char *bad_object_sha1;\ndiff --git a/pack-check.c b/pack-check.c\nindex 63a595c..8200f24 100644\n--- a/pack-check.c\n+++ b/pack-check.c\n@@ -25,6 +25,7 @@ int check_pack_crc(struct packed_git *p, struct pack_window **w_curs,\n {\n \tconst uint32_t *index_crc;\n \tuint32_t data_crc = crc32(0, NULL, 0);\n+\tunsigned sha1_table;\n \n \tdo {\n \t\tunsigned long avail;\n@@ -36,8 +37,9 @@ int check_pack_crc(struct packed_git *p, struct pack_window **w_curs,\n \t\tlen -= avail;\n \t} while (len);\n \n+\tsha1_table = p->index_version < 3 ? (p->num_objects * (20/4)) : 0;\n \tindex_crc = p->index_data;\n-\tindex_crc += 2 + 256 + p->num_objects * (20/4) + nr;\n+\tindex_crc += 2 + 256 + sha1_table + nr;\n \n \treturn data_crc != ntohl(*index_crc);\n }\ndiff --git a/pack-revindex.c b/pack-revindex.c\nindex b4d2b35..739a568 100644\n--- a/pack-revindex.c\n+++ b/pack-revindex.c\n@@ -170,9 +170,10 @@ static void create_pack_revindex(struct pack_revindex *rix)\n \tindex += 4 * 256;\n \n \tif (p->index_version > 1) {\n-\t\tconst uint32_t *off_32 =\n-\t\t\t(uint32_t *)(index + 8 + p->num_objects * (20 + 4));\n-\t\tconst uint32_t *off_64 = off_32 + p->num_objects;\n+\t\tconst uint32_t *off_32, *off_64;\n+\t\tunsigned sha1 = p->index_version < 3 ? 20 : 0;\n+\t\toff_32 = (uint32_t *)(index + 8 + p->num_objects * (sha1 + 4));\n+\t\toff_64 = off_32 + p->num_objects;\n \t\tfor (i = 0; i < num_ent; i++) {\n \t\t\tuint32_t off = ntohl(*off_32++);\n \t\t\tif (!(off & 0x80000000)) {\ndiff --git a/sha1_file.c b/sha1_file.c\nindex c2020d0..5c63781 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -504,7 +504,7 @@ static int check_packed_git_idx(const char *path,  struct packed_git *p)\n \thdr = idx_map;\n \tif (hdr->idx_signature == htonl(PACK_IDX_SIGNATURE)) {\n \t\tversion = ntohl(hdr->idx_version);\n-\t\tif (version < 2 || version > 2) {\n+\t\tif (version < 2 || version > 3) {\n \t\t\tmunmap(idx_map, idx_size);\n \t\t\treturn error(\"index file %s is version %\"PRIu32\n \t\t\t\t     \" and is not supported by this binary\"\n@@ -539,12 +539,13 @@ static int check_packed_git_idx(const char *path,  struct packed_git *p)\n \t\t\tmunmap(idx_map, idx_size);\n \t\t\treturn error(\"wrong index v1 file size in %s\", path);\n \t\t}\n-\t} else if (version == 2) {\n+\t} else if (version == 2 || version == 3) {\n+\t\tunsigned long min_size, max_size;\n \t\t/*\n \t\t * Minimum size:\n \t\t *  - 8 bytes of header\n \t\t *  - 256 index entries 4 bytes each\n-\t\t *  - 20-byte sha1 entry * nr\n+\t\t *  - 20-byte sha1 entry * nr (version 2 only)\n \t\t *  - 4-byte crc entry * nr\n \t\t *  - 4-byte offset entry * nr\n \t\t *  - 20-byte SHA1 of the packfile\n@@ -553,8 +554,10 @@ static int check_packed_git_idx(const char *path,  struct packed_git *p)\n \t\t * variable sized table containing 8-byte entries\n \t\t * for offsets larger than 2^31.\n \t\t */\n-\t\tunsigned long min_size = 8 + 4*256 + nr*(20 + 4 + 4) + 20 + 20;\n-\t\tunsigned long max_size = min_size;\n+\t\tmin_size = 8 + 4*256 + nr*(20 + 4 + 4) + 20 + 20;\n+\t\tif (version != 2)\n+\t\t\tmin_size -= nr*20;\n+\t\tmax_size = min_size;\n \t\tif (nr)\n \t\t\tmax_size += (nr - 1)*8;\n \t\tif (idx_size < min_size || idx_size > max_size) {\n@@ -573,6 +576,36 @@ static int check_packed_git_idx(const char *path,  struct packed_git *p)\n \t\t}\n \t}\n \n+\tif (version >= 3) {\n+\t\t/* the SHA1 table is located in the main pack file */\n+\t\tvoid *pack_map;\n+\t\tstruct pack_header *pack_hdr;\n+\n+\t\tfd = git_open_noatime(p->pack_name);\n+\t\tif (fd < 0) {\n+\t\t\tmunmap(idx_map, idx_size);\n+\t\t\treturn error(\"unable to open %s\", p->pack_name);\n+\t\t}\n+\t\tif (fstat(fd, &st) != 0 || xsize_t(st.st_size) < 12 + nr*20) {\n+\t\t\tclose(fd);\n+\t\t\tmunmap(idx_map, idx_size);\n+\t\t\treturn error(\"size of %s is wrong\", p->pack_name);\n+\t\t}\n+\t\tpack_map = xmmap(NULL, 12 + nr*20, PROT_READ, MAP_PRIVATE, fd, 0);\n+\t\tclose(fd);\n+\t\tpack_hdr = pack_map;\n+\t\tif (pack_hdr->hdr_signature != htonl(PACK_SIGNATURE) ||\n+\t\t    pack_hdr->hdr_version != htonl(4) ||\n+\t\t    pack_hdr->hdr_entries != htonl(nr)) {\n+\t\t\tmunmap(idx_map, idx_size);\n+\t\t\tmunmap(pack_map, 12 + nr*20);\n+\t\t\treturn error(\"packfile for %s doesn't match expectations\", path);\n+\t\t}\n+\t\tp->sha1_table = pack_map;\n+\t\tp->sha1_table += 12;\n+\t} else\n+\t\tp->sha1_table = NULL;\n+\n \tp->index_version = version;\n \tp->index_data = idx_map;\n \tp->index_size = idx_size;\n@@ -697,6 +730,10 @@ void close_pack_index(struct packed_git *p)\n \t\tmunmap((void *)p->index_data, p->index_size);\n \t\tp->index_data = NULL;\n \t}\n+\tif (p->sha1_table) {\n+\t\tmunmap((void *)(p->sha1_table - 12), 12 + p->num_objects * 20);\n+\t\tp->sha1_table = NULL;\n+\t}\n }\n \n /*\n@@ -2226,9 +2263,12 @@ const unsigned char *nth_packed_object_sha1(struct packed_git *p,\n \tindex += 4 * 256;\n \tif (p->index_version == 1) {\n \t\treturn index + 24 * n + 4;\n-\t} else {\n+\t} else if (p->index_version == 2) {\n \t\tindex += 8;\n \t\treturn index + 20 * n;\n+\t} else {\n+\t\tindex = p->sha1_table;\n+\t\treturn index + 20 * n;\n \t}\n }\n \n@@ -2241,6 +2281,8 @@ off_t nth_packed_object_offset(const struct packed_git *p, uint32_t n)\n \t} else {\n \t\tuint32_t off;\n \t\tindex += 8 + p->num_objects * (20 + 4);\n+\t\tif (p->index_version != 2)\n+\t\t\tindex -= p->num_objects * 20;\n \t\toff = ntohl(*((uint32_t *)(index + 4 * n)));\n \t\tif (!(off & 0x80000000))\n \t\t\treturn off;\n@@ -2281,6 +2323,8 @@ off_t find_pack_entry_one(const unsigned char *sha1,\n \t\tstride = 24;\n \t\tindex += 4;\n \t}\n+\tif (p->index_version > 2)\n+\t\tindex = p->sha1_table;\n \n \tif (debug_lookup)\n \t\tprintf(\"%02x%02x%02x... lo %u hi %u nr %\"PRIu32\"\\n\",\n-- \n1.8.4.38.g317e65b\n"},{"id":"226802","messageId":"1378362001-1738-27-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 26/38] pack v4: object header decode","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:49Z","receivedAt":"2013-09-05T06:19:49Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"For this we need the pack version.  However only open_packed_git_1() has\nbeen audited for pack v4 so far, hence the version validation is not\nadded to pack_version_ok() just yet.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n cache.h     |  1 +\n sha1_file.c | 14 ++++++++++++--\n 2 files changed, 13 insertions(+), 2 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex c939b60..59d9ba7 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1025,6 +1025,7 @@ extern struct packed_git {\n \tuint32_t num_objects;\n \tuint32_t num_bad_objects;\n \tunsigned char *bad_object_sha1;\n+\tint version;\n \tint index_version;\n \ttime_t mtime;\n \tint pack_fd;\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 5c63781..a298933 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -10,6 +10,7 @@\n #include \"string-list.h\"\n #include \"delta.h\"\n #include \"pack.h\"\n+#include \"varint.h\"\n #include \"blob.h\"\n #include \"commit.h\"\n #include \"run-command.h\"\n@@ -845,10 +846,11 @@ static int open_packed_git_1(struct packed_git *p)\n \t\treturn error(\"file %s is far too short to be a packfile\", p->pack_name);\n \tif (hdr.hdr_signature != htonl(PACK_SIGNATURE))\n \t\treturn error(\"file %s is not a GIT packfile\", p->pack_name);\n-\tif (!pack_version_ok(hdr.hdr_version))\n+\tif (!pack_version_ok(hdr.hdr_version) && hdr.hdr_version != htonl(4))\n \t\treturn error(\"packfile %s is version %\"PRIu32\" and not\"\n \t\t\t\" supported (try upgrading GIT to a newer version)\",\n \t\t\tp->pack_name, ntohl(hdr.hdr_version));\n+\tp->version = ntohl(hdr.hdr_version);\n \n \t/* Verify the pack matches its index. */\n \tif (p->num_objects != ntohl(hdr.hdr_entries))\n@@ -1725,7 +1727,15 @@ int unpack_object_header(struct packed_git *p,\n \t * insane, so we know won't exceed what we have been given.\n \t */\n \tbase = use_pack(p, w_curs, *curpos, &left);\n-\tused = unpack_object_header_buffer(base, left, &type, sizep);\n+\tif (p->version < 4) {\n+\t\tused = unpack_object_header_buffer(base, left, &type, sizep);\n+\t} else {\n+\t\tconst unsigned char *cp = base;\n+\t\tuintmax_t val = decode_varint(&cp);\n+\t\tused = cp - base;\n+\t\ttype = val & 0xf;\n+\t\t*sizep = val >> 4;\n+\t}\n \tif (!used) {\n \t\ttype = OBJ_BAD;\n \t} else\n-- \n1.8.4.38.g317e65b\n"},{"id":"226801","messageId":"1378362001-1738-28-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 27/38] pack v4: code to obtain a SHA1 from a sha1ref","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:50Z","receivedAt":"2013-09-05T06:19:50Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Let's start actually parsing pack v4 data.  Here's the first item.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n Makefile       |  1 +\n packv4-parse.c | 30 ++++++++++++++++++++++++++++++\n 2 files changed, 31 insertions(+)\n create mode 100644 packv4-parse.c\n\ndiff --git a/Makefile b/Makefile\nindex 4716113..ba6cafc 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -838,6 +838,7 @@ LIB_OBJS += object.o\n LIB_OBJS += pack-check.o\n LIB_OBJS += pack-revindex.o\n LIB_OBJS += pack-write.o\n+LIB_OBJS += packv4-parse.o\n LIB_OBJS += pager.o\n LIB_OBJS += parse-options.o\n LIB_OBJS += parse-options-cb.o\ndiff --git a/packv4-parse.c b/packv4-parse.c\nnew file mode 100644\nindex 0000000..299fc48\n--- /dev/null\n+++ b/packv4-parse.c\n@@ -0,0 +1,30 @@\n+/*\n+ * Code to parse pack v4 object encoding\n+ *\n+ * (C) Nicolas Pitre <nico@fluxnic.net>\n+ *\n+ * This code is free software; you can redistribute it and/or modify\n+ * it under the terms of the GNU General Public License version 2 as\n+ * published by the Free Software Foundation.\n+ */\n+\n+#include \"cache.h\"\n+#include \"varint.h\"\n+\n+const unsigned char *get_sha1ref(struct packed_git *p,\n+\t\t\t\t const unsigned char **bufp)\n+{\n+\tconst unsigned char *sha1;\n+\n+\tif (!**bufp) {\n+\t\tsha1 = *bufp + 1;\n+\t\t*bufp += 21;\n+\t} else {\n+\t\tunsigned int index = decode_varint(bufp);\n+\t\tif (index < 1 || index - 1 > p->num_objects)\n+\t\t\tdie(\"bad index in %s\", __func__);\n+\t\tsha1 = p->sha1_table + (index - 1) * 20;\n+\t}\n+\n+\treturn sha1;\n+}\n-- \n1.8.4.38.g317e65b\n"},{"id":"226798","messageId":"1378362001-1738-29-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 28/38] pack v4: code to load and prepare a pack dictionary table for use","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:51Z","receivedAt":"2013-09-05T06:19:51Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-parse.c | 77 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 77 insertions(+)\n\ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex 299fc48..26894bc 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -28,3 +28,80 @@ const unsigned char *get_sha1ref(struct packed_git *p,\n \n \treturn sha1;\n }\n+\n+struct packv4_dict {\n+\tconst unsigned char *data;\n+\tunsigned int nb_entries;\n+\tunsigned int offsets[FLEX_ARRAY];\n+};\n+\n+static struct packv4_dict *load_dict(struct packed_git *p, off_t *offset)\n+{\n+\tstruct pack_window *w_curs = NULL;\n+\toff_t curpos = *offset;\n+\tunsigned long dict_size, avail;\n+\tunsigned char *src, *data;\n+\tconst unsigned char *cp;\n+\tgit_zstream stream;\n+\tstruct packv4_dict *dict;\n+\tint nb_entries, i, st;\n+\n+\t/* get uncompressed dictionary data size */\n+\tsrc = use_pack(p, &w_curs, curpos, &avail);\n+\tcp = src;\n+\tdict_size = decode_varint(&cp);\n+\tif (dict_size < 3) {\n+\t\terror(\"bad dict size\");\n+\t\treturn NULL;\n+\t}\n+\tcurpos += cp - src;\n+\n+\tdata = xmallocz(dict_size);\n+\tmemset(&stream, 0, sizeof(stream));\n+\tstream.next_out = data;\n+\tstream.avail_out = dict_size + 1;\n+\n+\tgit_inflate_init(&stream);\n+\tdo {\n+\t\tsrc = use_pack(p, &w_curs, curpos, &stream.avail_in);\n+\t\tstream.next_in = src;\n+\t\tst = git_inflate(&stream, Z_FINISH);\n+\t\tcurpos += stream.next_in - src;\n+\t} while ((st == Z_OK || st == Z_BUF_ERROR) && stream.avail_out);\n+\tgit_inflate_end(&stream);\n+\tunuse_pack(&w_curs);\n+\tif (st != Z_STREAM_END || stream.total_out != dict_size) {\n+\t\terror(\"pack dictionary bad\");\n+\t\tfree(data);\n+\t\treturn NULL;\n+\t}\n+\n+\t/* count number of entries */\n+\tnb_entries = 0;\n+\tcp = data;\n+\twhile (cp < data + dict_size - 3) {\n+\t\tcp += 2;  /* prefix bytes */\n+\t\tcp += strlen((const char *)cp);  /* entry string */\n+\t\tcp += 1;  /* terminating NUL */\n+\t\tnb_entries++;\n+\t}\n+\tif (cp - data != dict_size) {\n+\t\terror(\"dict size mismatch\");\n+\t\tfree(data);\n+\t\treturn NULL;\n+\t}\n+\n+\tdict = xmalloc(sizeof(*dict) + nb_entries * sizeof(dict->offsets[0]));\n+\tdict->data = data;\n+\tdict->nb_entries = nb_entries;\n+\n+\tcp = data;\n+\tfor (i = 0; i < nb_entries; i++) {\n+\t\tdict->offsets[i] = cp - data;\n+\t\tcp += 2;\n+\t\tcp += strlen((const char *)cp) + 1;\n+\t}\n+\n+\t*offset = curpos;\n+\treturn dict;\n+}\n-- \n1.8.4.38.g317e65b\n"},{"id":"226788","messageId":"1378362001-1738-30-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 29/38] pack v4: code to retrieve a name","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:52Z","receivedAt":"2013-09-05T06:19:52Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"The name dictionary is loaded if not already done.  We know it is\nlocated right after the SHA1 table (20 bytes per object) which is\nitself right after the 12-byte header.\n\nThen the index is parsed from the input buffer and a pointer to the\ncorresponding entry is returned.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n cache.h        |  3 +++\n packv4-parse.c | 24 ++++++++++++++++++++++++\n 2 files changed, 27 insertions(+)\n\ndiff --git a/cache.h b/cache.h\nindex 59d9ba7..6ce327e 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1015,6 +1015,8 @@ struct pack_window {\n \tunsigned int inuse_cnt;\n };\n \n+struct packv4_dict;\n+\n extern struct packed_git {\n \tstruct packed_git *next;\n \tstruct pack_window *windows;\n@@ -1027,6 +1029,7 @@ extern struct packed_git {\n \tunsigned char *bad_object_sha1;\n \tint version;\n \tint index_version;\n+\tstruct packv4_dict *name_dict;\n \ttime_t mtime;\n \tint pack_fd;\n \tunsigned pack_local:1,\ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex 26894bc..074e107 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -105,3 +105,27 @@ static struct packv4_dict *load_dict(struct packed_git *p, off_t *offset)\n \t*offset = curpos;\n \treturn dict;\n }\n+\n+static void load_name_dict(struct packed_git *p)\n+{\n+\toff_t offset = 12 + p->num_objects * 20;\n+\tstruct packv4_dict *names = load_dict(p, &offset);\n+\tif (!names)\n+\t\tdie(\"bad pack name dictionary in %s\", p->pack_name);\n+\tp->name_dict = names;\n+}\n+\n+const unsigned char *get_nameref(struct packed_git *p, const unsigned char **srcp)\n+{\n+\tunsigned int index;\n+\n+\tif (!p->name_dict)\n+\t\tload_name_dict(p);\n+\n+\tindex = decode_varint(srcp);\n+\tif (index >= p->name_dict->nb_entries) {\n+\t\terror(\"%s: index overflow\", __func__);\n+\t\treturn NULL;\n+\t}\n+\treturn p->name_dict->data + p->name_dict->offsets[index];\n+}\n-- \n1.8.4.38.g317e65b\n"},{"id":"226800","messageId":"1378362001-1738-31-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 30/38] pack v4: code to recreate a canonical commit object","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:53Z","receivedAt":"2013-09-05T06:19:53Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Usage of snprintf() is possibly not the most efficient approach.\nFor example we could simply copy the needed strings and generate\nthe SHA1 hex strings directly into the destination buffer.  But\nsuch optimizations may come later.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-parse.c | 74 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 74 insertions(+)\n\ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex 074e107..bca1a97 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -129,3 +129,77 @@ const unsigned char *get_nameref(struct packed_git *p, const unsigned char **src\n \t}\n \treturn p->name_dict->data + p->name_dict->offsets[index];\n }\n+\n+void *pv4_get_commit(struct packed_git *p, struct pack_window **w_curs,\n+\t\t     off_t offset, unsigned long size)\n+{\n+\tunsigned long avail;\n+\tgit_zstream stream;\n+\tint len, st;\n+\tunsigned int nb_parents;\n+\tunsigned char *dst, *dcp;\n+\tconst unsigned char *src, *scp, *sha1, *name;\n+\tunsigned long time;\n+\tint16_t tz;\n+\n+\tdst = xmallocz(size);\n+\tdcp = dst;\n+\n+\tsrc = use_pack(p, w_curs, offset, &avail);\n+\tscp = src;\n+\n+\tsha1 = get_sha1ref(p, &scp);\n+\tlen = snprintf((char *)dcp, size, \"tree %s\\n\", sha1_to_hex(sha1));\n+\tdcp += len;\n+\tsize -= len;\n+\n+\tnb_parents = decode_varint(&scp);\n+\twhile (nb_parents--) {\n+\t\tsha1 = get_sha1ref(p, &scp);\n+\t\tlen = snprintf((char *)dcp, size, \"parent %s\\n\", sha1_to_hex(sha1));\n+\t\tif (len >= size)\n+\t\t\tdie(\"overflow in %s\", __func__);\n+\t\tdcp += len;\n+\t\tsize -= len;\n+\t}\n+\n+\tname = get_nameref(p, &scp);\n+\ttz = (name[0] << 8) | name[1];\n+\ttime = decode_varint(&scp);\n+\tlen = snprintf((char *)dcp, size, \"author %s %lu %+05d\\n\", name+2, time, tz);\n+\tif (len >= size)\n+\t\tdie(\"overflow in %s\", __func__);\n+\tdcp += len;\n+\tsize -= len;\n+\n+\tname = get_nameref(p, &scp);\n+\ttz = (name[0] << 8) | name[1];\n+\ttime = decode_varint(&scp);\n+\tlen = snprintf((char *)dcp, size, \"committer %s %lu %+05d\\n\", name+2, time, tz);\n+\tif (len >= size)\n+\t\tdie(\"overflow in %s\", __func__);\n+\tdcp += len;\n+\tsize -= len;\n+\n+\tif (scp - src > avail)\n+\t\tdie(\"overflow in %s\", __func__);\n+\toffset += scp - src;\n+\n+\tmemset(&stream, 0, sizeof(stream));\n+\tstream.next_out = dcp;\n+\tstream.avail_out = size + 1;\n+\tgit_inflate_init(&stream);\n+\tdo {\n+\t\tsrc = use_pack(p, w_curs, offset, &stream.avail_in);\n+\t\tstream.next_in = (unsigned char *)src;\n+\t\tst = git_inflate(&stream, Z_FINISH);\n+\t\toffset += stream.next_in - src;\n+\t} while ((st == Z_OK || st == Z_BUF_ERROR) && stream.avail_out);\n+\tgit_inflate_end(&stream);\n+\tif (st != Z_STREAM_END || stream.total_out != size) {\n+\t\tfree(dst);\n+\t\treturn NULL;\n+\t}\n+\n+\treturn dst;\n+}\n-- \n1.8.4.38.g317e65b\n"},{"id":"226792","messageId":"1378362001-1738-32-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 31/38] sha1_file.c: make use of decode_varint()","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:54Z","receivedAt":"2013-09-05T06:19:54Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"... replacing the equivalent open coded loop.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n sha1_file.c | 14 +++-----------\n 1 file changed, 3 insertions(+), 11 deletions(-)\n\ndiff --git a/sha1_file.c b/sha1_file.c\nindex a298933..67eb903 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1687,20 +1687,12 @@ static off_t get_delta_base(struct packed_git *p,\n \t * is stupid, as then a REF_DELTA would be smaller to store.\n \t */\n \tif (type == OBJ_OFS_DELTA) {\n-\t\tunsigned used = 0;\n-\t\tunsigned char c = base_info[used++];\n-\t\tbase_offset = c & 127;\n-\t\twhile (c & 128) {\n-\t\t\tbase_offset += 1;\n-\t\t\tif (!base_offset || MSB(base_offset, 7))\n-\t\t\t\treturn 0;  /* overflow */\n-\t\t\tc = base_info[used++];\n-\t\t\tbase_offset = (base_offset << 7) + (c & 127);\n-\t\t}\n+\t\tconst unsigned char *cp = base_info;\n+\t\tbase_offset = decode_varint(&cp);\n \t\tbase_offset = delta_obj_offset - base_offset;\n \t\tif (base_offset <= 0 || base_offset >= delta_obj_offset)\n \t\t\treturn 0;  /* out of bound */\n-\t\t*curpos += used;\n+\t\t*curpos += cp - base_info;\n \t} else if (type == OBJ_REF_DELTA) {\n \t\t/* The base entry _must_ be in the same pack */\n \t\tbase_offset = find_pack_entry_one(base_info, p);\n-- \n1.8.4.38.g317e65b\n"},{"id":"226793","messageId":"1378362001-1738-33-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 32/38] pack v4: parse delta base reference","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:55Z","receivedAt":"2013-09-05T06:19:55Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"There is only one type of delta with pack v4.  The base reference\nencoding already handles either an offset (via the pack index) or a\nliteral SHA1.\n\nWe assume in the literal SHA1 case that the object lives in the same\npack, just like with previous pack versions.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n sha1_file.c | 14 +++++++++++++-\n 1 file changed, 13 insertions(+), 1 deletion(-)\n\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 67eb903..f3bfa28 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1686,7 +1686,19 @@ static off_t get_delta_base(struct packed_git *p,\n \t * that is assured.  An OFS_DELTA longer than the hash size\n \t * is stupid, as then a REF_DELTA would be smaller to store.\n \t */\n-\tif (type == OBJ_OFS_DELTA) {\n+\tif (p->version >= 4) {\n+\t\tif (base_info[0] != 0) {\n+\t\t\tconst unsigned char *cp = base_info;\n+\t\t\tunsigned int base_index = decode_varint(&cp);\n+\t\t\tif (!base_index || base_index - 1 >= p->num_objects)\n+\t\t\t\treturn 0;  /* out of bounds */\n+\t\t\t*curpos += cp - base_info;\n+\t\t\tbase_offset = nth_packed_object_offset(p, base_index - 1);\n+\t\t} else {\n+\t\t\tbase_offset = find_pack_entry_one(base_info+1, p);\n+\t\t\t*curpos += 21;\n+\t\t}\n+\t} else if (type == OBJ_OFS_DELTA) {\n \t\tconst unsigned char *cp = base_info;\n \t\tbase_offset = decode_varint(&cp);\n \t\tbase_offset = delta_obj_offset - base_offset;\n-- \n1.8.4.38.g317e65b\n"},{"id":"226796","messageId":"1378362001-1738-34-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 33/38] pack v4: we can read commit objects now","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:56Z","receivedAt":"2013-09-05T06:19:56Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n Makefile       |  1 +\n packv4-parse.c |  1 +\n packv4-parse.h |  7 +++++++\n sha1_file.c    | 10 ++++++++++\n 4 files changed, 19 insertions(+)\n create mode 100644 packv4-parse.h\n\ndiff --git a/Makefile b/Makefile\nindex ba6cafc..22fc276 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -702,6 +702,7 @@ LIB_H += notes.h\n LIB_H += object.h\n LIB_H += pack-revindex.h\n LIB_H += pack.h\n+LIB_H += packv4-parse.h\n LIB_H += parse-options.h\n LIB_H += patch-ids.h\n LIB_H += pathspec.h\ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex bca1a97..431f47e 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -9,6 +9,7 @@\n  */\n \n #include \"cache.h\"\n+#include \"packv4-parse.h\"\n #include \"varint.h\"\n \n const unsigned char *get_sha1ref(struct packed_git *p,\ndiff --git a/packv4-parse.h b/packv4-parse.h\nnew file mode 100644\nindex 0000000..40aa75a\n--- /dev/null\n+++ b/packv4-parse.h\n@@ -0,0 +1,7 @@\n+#ifndef PACKV4_PARSE_H\n+#define PACKV4_PARSE_H\n+\n+void *pv4_get_commit(struct packed_git *p, struct pack_window **w_curs,\n+\t\t     off_t offset, unsigned long size);\n+\n+#endif\ndiff --git a/sha1_file.c b/sha1_file.c\nindex f3bfa28..b57d9f8 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -19,6 +19,7 @@\n #include \"tree-walk.h\"\n #include \"refs.h\"\n #include \"pack-revindex.h\"\n+#include \"packv4-parse.h\"\n #include \"sha1-lookup.h\"\n #include \"bulk-checkin.h\"\n #include \"streaming.h\"\n@@ -2172,6 +2173,15 @@ void *unpack_entry(struct packed_git *p, off_t obj_offset,\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n+\t\tif (p->version >= 4 && !base_from_cache) {\n+\t\t\tif (type == OBJ_COMMIT) {\n+\t\t\t\tdata = pv4_get_commit(p, &w_curs, curpos, size);\n+\t\t\t} else {\n+\t\t\t\tdie(\"no pack v4 tree parsing yet\");\n+\t\t\t}\n+\t\t\tbreak;\n+\t\t}\n+\t\t/* fall through */\n \tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n \t\tif (!base_from_cache)\n-- \n1.8.4.38.g317e65b\n"},{"id":"226795","messageId":"1378362001-1738-35-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 34/38] pack v4: code to retrieve a path component","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:57Z","receivedAt":"2013-09-05T06:19:57Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Because the path dictionary table is located right after the name\ndictionary table, we currently need to load the later to find the\nformer.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n cache.h        |  2 ++\n packv4-parse.c | 36 ++++++++++++++++++++++++++++++++++++\n 2 files changed, 38 insertions(+)\n\ndiff --git a/cache.h b/cache.h\nindex 6ce327e..5f2147a 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1030,6 +1030,8 @@ extern struct packed_git {\n \tint version;\n \tint index_version;\n \tstruct packv4_dict *name_dict;\n+\toff_t name_dict_end;\n+\tstruct packv4_dict *path_dict;\n \ttime_t mtime;\n \tint pack_fd;\n \tunsigned pack_local:1,\ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex 431f47e..b80b73e 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -114,6 +114,7 @@ static void load_name_dict(struct packed_git *p)\n \tif (!names)\n \t\tdie(\"bad pack name dictionary in %s\", p->pack_name);\n \tp->name_dict = names;\n+\tp->name_dict_end = offset;\n }\n \n const unsigned char *get_nameref(struct packed_git *p, const unsigned char **srcp)\n@@ -131,6 +132,41 @@ const unsigned char *get_nameref(struct packed_git *p, const unsigned char **src\n \treturn p->name_dict->data + p->name_dict->offsets[index];\n }\n \n+static void load_path_dict(struct packed_git *p)\n+{\n+\toff_t offset;\n+\tstruct packv4_dict *paths;\n+\n+\t/*\n+\t * For now we need to load the name dictionary to know where\n+\t * it ends and therefore where the path dictionary starts.\n+\t */\n+\tif (!p->name_dict)\n+\t\tload_name_dict(p);\n+\n+\toffset = p->name_dict_end;\n+\tpaths = load_dict(p, &offset);\n+\tif (!paths)\n+\t\tdie(\"bad pack path dictionary in %s\", p->pack_name);\n+\tp->path_dict = paths;\n+}\n+\n+const unsigned char *get_pathref(struct packed_git *p, const unsigned char **srcp)\n+{\n+\tunsigned int index;\n+\n+\tif (!p->path_dict)\n+\t\tload_path_dict(p);\n+\n+\tindex = decode_varint(srcp);\n+\tif (index < 1 || index - 1 >= p->path_dict->nb_entries) {\n+\t\terror(\"%s: index overflow\", __func__);\n+\t\treturn NULL;\n+\t}\n+\tindex -= 1;\n+\treturn p->path_dict->data + p->path_dict->offsets[index];\n+}\n+\n void *pv4_get_commit(struct packed_git *p, struct pack_window **w_curs,\n \t\t     off_t offset, unsigned long size)\n {\n-- \n1.8.4.38.g317e65b\n"},{"id":"226794","messageId":"1378362001-1738-36-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 35/38] pack v4: decode tree objects","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:58Z","receivedAt":"2013-09-05T06:19:58Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"For now we recreate the whole tree object in its canonical form.\n\nEventually, the core code should grow some ability to walk packv4 tree\nentries directly which would be way more efficient.  Not only would that\navoid double tree entry parsing, but the pack v4 encoding allows for\ngetting at child objects without going through the SHA1 search.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-parse.c | 137 ++++++++++++++++++++++++++++++++++++++++++++++++++++++---\n 1 file changed, 131 insertions(+), 6 deletions(-)\n\ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex b80b73e..04eab46 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -151,19 +151,15 @@ static void load_path_dict(struct packed_git *p)\n \tp->path_dict = paths;\n }\n \n-const unsigned char *get_pathref(struct packed_git *p, const unsigned char **srcp)\n+const unsigned char *get_pathref(struct packed_git *p, unsigned int index)\n {\n-\tunsigned int index;\n-\n \tif (!p->path_dict)\n \t\tload_path_dict(p);\n \n-\tindex = decode_varint(srcp);\n-\tif (index < 1 || index - 1 >= p->path_dict->nb_entries) {\n+\tif (index >= p->path_dict->nb_entries) {\n \t\terror(\"%s: index overflow\", __func__);\n \t\treturn NULL;\n \t}\n-\tindex -= 1;\n \treturn p->path_dict->data + p->path_dict->offsets[index];\n }\n \n@@ -240,3 +236,132 @@ void *pv4_get_commit(struct packed_git *p, struct pack_window **w_curs,\n \n \treturn dst;\n }\n+\n+static int decode_entries(struct packed_git *p, struct pack_window **w_curs,\n+\t\t\t  off_t offset, unsigned int start, unsigned int count,\n+\t\t\t  unsigned char **dstp, unsigned long *sizep, int hdr)\n+{\n+\tunsigned long avail;\n+\tunsigned int nb_entries;\n+\tconst unsigned char *src, *scp;\n+\toff_t copy_objoffset = 0;\n+\n+\tsrc = use_pack(p, w_curs, offset, &avail);\n+\tscp = src;\n+\n+\tif (hdr) {\n+\t\t/* we need to skip over the object header */\n+\t\twhile (*scp & 128)\n+\t\t\tif (++scp - src >= avail - 20)\n+\t\t\t\treturn -1;\n+\t\t/* let's still make sure this is actually a tree */\n+\t\tif ((*scp++ & 0xf) != OBJ_TREE)\n+\t\t\treturn -1;\n+\t}\n+\n+\tnb_entries = decode_varint(&scp);\n+\tif (scp == src || start > nb_entries || count > nb_entries - start)\n+\t\treturn -1;\n+\toffset += scp - src;\n+\tavail -= scp - src;\n+\tsrc = scp;\n+\n+\twhile (count) {\n+\t\tunsigned int what;\n+\n+\t\tif (avail < 20) {\n+\t\t\tsrc = use_pack(p, w_curs, offset, &avail);\n+\t\t\tif (avail < 20)\n+\t\t\t\treturn -1;\n+\t\t}\n+\t\tscp = src;\n+\n+\t\twhat = decode_varint(&scp);\n+\t\tif (scp == src)\n+\t\t\treturn -1;\n+\n+\t\tif (!(what & 1) && start != 0) {\n+\t\t\t/*\n+\t\t\t * This is a single entry and we have to skip it.\n+\t\t\t * The path index was parsed and is in 'what'.\n+\t\t\t * Skip over the SHA1 index.\n+\t\t\t */\n+\t\t\twhile (*scp++ & 128);\n+\t\t\tstart--;\n+\t\t} else if (!(what & 1) && start == 0) {\n+\t\t\t/*\n+\t\t\t * This is an actual tree entry to recreate.\n+\t\t\t */\n+\t\t\tconst unsigned char *path, *sha1;\n+\t\t\tunsigned mode;\n+\t\t\tint len;\n+\n+\t\t\tpath = get_pathref(p, what >> 1);\n+\t\t\tsha1 = get_sha1ref(p, &scp);\n+\t\t\tif (!path || !sha1)\n+\t\t\t\treturn -1;\n+\t\t\tmode = (path[0] << 8) | path[1];\n+\t\t\tlen = snprintf((char *)*dstp, *sizep, \"%o %s%c\",\n+\t\t\t\t\t   mode, path+2, '\\0');\n+\t\t\tif (len + 20 > *sizep)\n+\t\t\t\treturn -1;\n+\t\t\thashcpy(*dstp + len, sha1);\n+\t\t\t*dstp += len + 20;\n+\t\t\t*sizep -= len + 20;\n+\t\t\tcount--;\n+\t\t} else if (what & 1) {\n+\t\t\t/*\n+\t\t\t * Copy from another tree object.\n+\t\t\t */\n+\t\t\tunsigned int copy_start, copy_count;\n+\n+\t\t\tcopy_start = what >> 1;\n+\t\t\tcopy_count = decode_varint(&scp);\n+\t\t\tif (!copy_count)\n+\t\t\t\treturn -1;\n+\n+\t\t\t/*\n+\t\t\t * The LSB of copy_count is a flag indicating if\n+\t\t\t * a third value is provided to specify the source\n+\t\t\t * object.  This may be omitted when it doesn't\n+\t\t\t * change, but has to be specified at least for the\n+\t\t\t * first copy sequence.\n+\t\t\t */\n+\t\t\tif (copy_count & 1) {\n+\t\t\t\tunsigned index = decode_varint(&scp);\n+\t\t\t\tif (!index)  /* thin pack */\n+\t\t\t\t\treturn -1;\n+\t\t\t\tcopy_objoffset =\n+\t\t\t\t\tnth_packed_object_offset(p, index - 1);\n+\t\t\t}\n+\t\t\tif (!copy_objoffset)\n+\t\t\t\treturn -1;\n+\t\t\tcopy_count >>= 1;\n+\n+\t\t\tif (start >= copy_count) {\n+\t\t\t\tstart -= copy_count;\n+\t\t\t} else {\n+\t\t\t\tint ret;\n+\t\t\t\tcopy_count -= start;\n+\t\t\t\tcopy_start += start;\n+\t\t\t\tstart = 0;\n+\t\t\t\tif (copy_count > count)\n+\t\t\t\t\tcopy_count = count;\n+\t\t\t\tcount -= copy_count;\n+\t\t\t\tret = decode_entries(p, w_curs,\n+\t\t\t\t\tcopy_objoffset, copy_start, copy_count,\n+\t\t\t\t\tdstp, sizep, 1);\n+\t\t\t\tif (ret)\n+\t\t\t\t\treturn ret;\n+\t\t\t\t/* force pack window readjustment */\n+\t\t\t\tavail = scp - src;\n+\t\t\t}\n+\t\t}\n+\n+\t\toffset += scp - src;\n+\t\tavail -= scp - src;\n+\t\tsrc = scp;\n+\t}\n+\n+\treturn 0;\n+}\n-- \n1.8.4.38.g317e65b\n"},{"id":"226789","messageId":"1378362001-1738-37-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 36/38] pack v4: get tree objects","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:19:59Z","receivedAt":"2013-09-05T06:19:59Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-parse.c | 25 +++++++++++++++++++++++++\n packv4-parse.h |  2 ++\n sha1_file.c    |  2 +-\n 3 files changed, 28 insertions(+), 1 deletion(-)\n\ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex 04eab46..4c218d2 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -365,3 +365,28 @@ static int decode_entries(struct packed_git *p, struct pack_window **w_curs,\n \n \treturn 0;\n }\n+\n+void *pv4_get_tree(struct packed_git *p, struct pack_window **w_curs,\n+\t\t   off_t offset, unsigned long size)\n+{\n+\tunsigned long avail;\n+\tunsigned int nb_entries;\n+\tunsigned char *dst, *dcp;\n+\tconst unsigned char *src, *scp;\n+\tint ret;\n+\n+\tsrc = use_pack(p, w_curs, offset, &avail);\n+\tscp = src;\n+\tnb_entries = decode_varint(&scp);\n+\tif (scp == src)\n+\t\treturn NULL;\n+\n+\tdst = xmallocz(size);\n+\tdcp = dst;\n+\tret = decode_entries(p, w_curs, offset, 0, nb_entries, &dcp, &size, 0);\n+\tif (ret < 0 || size != 0) {\n+\t\tfree(dst);\n+\t\treturn NULL;\n+\t}\n+\treturn dst;\n+}\ndiff --git a/packv4-parse.h b/packv4-parse.h\nindex 40aa75a..5f9d809 100644\n--- a/packv4-parse.h\n+++ b/packv4-parse.h\n@@ -3,5 +3,7 @@\n \n void *pv4_get_commit(struct packed_git *p, struct pack_window **w_curs,\n \t\t     off_t offset, unsigned long size);\n+void *pv4_get_tree(struct packed_git *p, struct pack_window **w_curs,\n+\t\t   off_t offset, unsigned long size);\n \n #endif\ndiff --git a/sha1_file.c b/sha1_file.c\nindex b57d9f8..79e1293 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -2177,7 +2177,7 @@ void *unpack_entry(struct packed_git *p, off_t obj_offset,\n \t\t\tif (type == OBJ_COMMIT) {\n \t\t\t\tdata = pv4_get_commit(p, &w_curs, curpos, size);\n \t\t\t} else {\n-\t\t\t\tdie(\"no pack v4 tree parsing yet\");\n+\t\t\t\tdata = pv4_get_tree(p, &w_curs, curpos, size);\n \t\t\t}\n \t\t\tbreak;\n \t\t}\n-- \n1.8.4.38.g317e65b\n"},{"id":"226791","messageId":"1378362001-1738-38-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 37/38] pack v4: introduce \"escape hatches\" in the name and path indexes","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:20:00Z","receivedAt":"2013-09-05T06:20:00Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"If the path or name index is zero, this means the entry data is to be\nfound inline rather than being located in the dictionary table. This is\nthere to allow easy completion of thin packs without having to add new\ntable entries which would have required a full rewrite of the pack data.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c |  6 +++---\n packv4-parse.c  | 28 ++++++++++++++++++++++------\n 2 files changed, 25 insertions(+), 9 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex fd16222..9d6ffc0 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -343,7 +343,7 @@ void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n \tindex = dict_add_entry(commit_name_table, tz_val, in, end - in);\n \tif (index < 0)\n \t\tgoto bad_dict;\n-\tout += encode_varint(index, out);\n+\tout += encode_varint(index + 1, out);\n \ttime = strtoul(end, &end, 10);\n \tif (!end || end[0] != ' ' || end[6] != '\\n')\n \t\tgoto bad_data;\n@@ -361,7 +361,7 @@ void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n \tindex = dict_add_entry(commit_name_table, tz_val, in, end - in);\n \tif (index < 0)\n \t\tgoto bad_dict;\n-\tout += encode_varint(index, out);\n+\tout += encode_varint(index + 1, out);\n \ttime = strtoul(end, &end, 10);\n \tif (!end || end[0] != ' ' || end[6] != '\\n')\n \t\tgoto bad_data;\n@@ -561,7 +561,7 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\t\tfree(buffer);\n \t\t\treturn NULL;\n \t\t}\n-\t\tout += encode_varint(index << 1, out);\n+\t\tout += encode_varint((index + 1) << 1, out);\n \t\tout += encode_sha1ref(name_entry.sha1, out);\n \t}\n \ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex 4c218d2..6db4ed3 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -125,11 +125,19 @@ const unsigned char *get_nameref(struct packed_git *p, const unsigned char **src\n \t\tload_name_dict(p);\n \n \tindex = decode_varint(srcp);\n-\tif (index >= p->name_dict->nb_entries) {\n+\n+\tif (!index) {\n+\t\t/* the entry data is inline */\n+\t\tconst unsigned char *data = *srcp;\n+\t\t*srcp += 2 + strlen((const char *)*srcp + 2) + 1;\n+\t\treturn data;\n+\t}\n+\n+\tif (index - 1 >= p->name_dict->nb_entries) {\n \t\terror(\"%s: index overflow\", __func__);\n \t\treturn NULL;\n \t}\n-\treturn p->name_dict->data + p->name_dict->offsets[index];\n+\treturn p->name_dict->data + p->name_dict->offsets[index - 1];\n }\n \n static void load_path_dict(struct packed_git *p)\n@@ -151,16 +159,24 @@ static void load_path_dict(struct packed_git *p)\n \tp->path_dict = paths;\n }\n \n-const unsigned char *get_pathref(struct packed_git *p, unsigned int index)\n+const unsigned char *get_pathref(struct packed_git *p, unsigned int index,\n+\t\t\t\t const unsigned char **srcp)\n {\n \tif (!p->path_dict)\n \t\tload_path_dict(p);\n \n-\tif (index >= p->path_dict->nb_entries) {\n+\tif (!index) {\n+\t\t/* the entry data is inline */\n+\t\tconst unsigned char *data = *srcp;\n+\t\t*srcp += 2 + strlen((const char *)*srcp + 2) + 1;\n+\t\treturn data;\n+\t}\n+\n+\tif (index - 1 >= p->path_dict->nb_entries) {\n \t\terror(\"%s: index overflow\", __func__);\n \t\treturn NULL;\n \t}\n-\treturn p->path_dict->data + p->path_dict->offsets[index];\n+\treturn p->path_dict->data + p->path_dict->offsets[index - 1];\n }\n \n void *pv4_get_commit(struct packed_git *p, struct pack_window **w_curs,\n@@ -296,7 +312,7 @@ static int decode_entries(struct packed_git *p, struct pack_window **w_curs,\n \t\t\tunsigned mode;\n \t\t\tint len;\n \n-\t\t\tpath = get_pathref(p, what >> 1);\n+\t\t\tpath = get_pathref(p, what >> 1, &scp);\n \t\t\tsha1 = get_sha1ref(p, &scp);\n \t\t\tif (!path || !sha1)\n \t\t\t\treturn -1;\n-- \n1.8.4.38.g317e65b\n"},{"id":"226790","messageId":"1378362001-1738-39-git-send-email-nico@fluxnic.net","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 38/38] packv4-create: add a command line argument to limit tree copy sequences","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T06:20:01Z","receivedAt":"2013-09-05T06:20:01Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"Because there is no delta object cache for tree objects yet, walking\ntree entries may result in a lot of recursion.\n\nLet's add --min-tree-copy=N where N is the minimum number of copied\nentries in a single copy sequence allowed for encoding tree deltas.\nThe default is 1. Specifying 0 disables tree deltas entirely.\n\nThis allows for experiments with the delta width and see the influence\non pack size vs runtime access cost.\n\nSigned-off-by: Nicolas Pitre <nico@fluxnic.net>\n---\n packv4-create.c | 27 ++++++++++++++++++++-------\n 1 file changed, 20 insertions(+), 7 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 9d6ffc0..34dcebf 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -19,6 +19,7 @@\n \n static int pack_compression_seen;\n static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n+static int min_tree_copy = 1;\n \n struct data_entry {\n \tunsigned offset;\n@@ -441,7 +442,7 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \tif (!size)\n \t\treturn NULL;\n \n-\tif (!delta_size)\n+\tif (!delta_size || !min_tree_copy)\n \t\tdelta = NULL;\n \n \t/*\n@@ -530,7 +531,6 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\t\tcp += encode_varint(copy_count, cp);\n \t\t\tif (first_delta)\n \t\t\t\tcp += encode_sha1ref(delta_sha1, cp);\n-\t\t\tcopy_count = 0;\n \n \t\t\t/*\n \t\t\t * Now let's make sure this is going to take less\n@@ -538,12 +538,14 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\t\t * created in parallel.  If so we dump the copy\n \t\t\t * sequence over those entries in the output buffer.\n \t\t\t */\n-\t\t\tif (cp - copy_buf < out - &buffer[copy_pos]) {\n+\t\t\tif (copy_count >= min_tree_copy &&\n+\t\t\t    cp - copy_buf < out - &buffer[copy_pos]) {\n \t\t\t\tout = buffer + copy_pos;\n \t\t\t\tmemcpy(out, copy_buf, cp - copy_buf);\n \t\t\t\tout += cp - copy_buf;\n \t\t\t\tfirst_delta = 0;\n \t\t\t}\n+\t\t\tcopy_count = 0;\n \t\t}\n \n \t\tif (end - out < 48) {\n@@ -574,7 +576,8 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\tcp += encode_varint(copy_count, cp);\n \t\tif (first_delta)\n \t\t\tcp += encode_sha1ref(delta_sha1, cp);\n-\t\tif (cp - copy_buf < out - &buffer[copy_pos]) {\n+\t\tif (copy_count >= min_tree_copy &&\n+\t\t    cp - copy_buf < out - &buffer[copy_pos]) {\n \t\t\tout = buffer + copy_pos;\n \t\t\tmemcpy(out, copy_buf, cp - copy_buf);\n \t\t\tout += cp - copy_buf;\n@@ -1078,14 +1081,24 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n \n int main(int argc, char *argv[])\n {\n-\tif (argc != 3) {\n-\t\tfprintf(stderr, \"Usage: %s <src_packfile> <dst_packfile>\\n\", argv[0]);\n+\tchar *src_pack, *dst_pack;\n+\n+\tif (argc == 3) {\n+\t\tsrc_pack = argv[1];\n+\t\tdst_pack = argv[2];\n+\t} else if (argc == 4 && !prefixcmp(argv[1], \"--min-tree-copy=\")) {\n+\t\tmin_tree_copy = atoi(argv[1] + strlen(\"--min-tree-copy=\"));\n+\t\tsrc_pack = argv[2];\n+\t\tdst_pack = argv[3];\n+\t} else {\n+\t\tfprintf(stderr, \"Usage: %s [--min-tree-copy=<n>] <src_packfile> <dst_packfile>\\n\", argv[0]);\n \t\texit(1);\n \t}\n+\n \tgit_config(git_pack_config, NULL);\n \tif (!pack_compression_seen && core_compression_seen)\n \t\tpack_compression_level = core_compression_level;\n-\tprocess_one_pack(argv[1], argv[2]);\n+\tprocess_one_pack(src_pack, dst_pack);\n \tif (0)\n \t\tdict_dump();\n \treturn 0;\n-- \n1.8.4.38.g317e65b\n"},{"id":"226828","messageId":"20130905073552.GD28959@goldbirke","threadId":"34852","inReplyTo":"1378362001-1738-32-git-send-email-nico@fluxnic.net","subject":"Re: [PATCH 31/38] sha1_file.c: make use of decode_varint()","fromName":"SZEDER Gábor","fromEmail":"szeder@ira.uka.de","sentAt":"2013-09-05T07:35:52Z","receivedAt":"2013-09-05T07:35:52Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Thu, Sep 05, 2013 at 02:19:54AM -0400, Nicolas Pitre wrote:\n> ... replacing the equivalent open coded loop.\n> \n> Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n> ---\n>  sha1_file.c | 14 +++-----------\n>  1 file changed, 3 insertions(+), 11 deletions(-)\n> \n> diff --git a/sha1_file.c b/sha1_file.c\n> index a298933..67eb903 100644\n> --- a/sha1_file.c\n> +++ b/sha1_file.c\n> @@ -1687,20 +1687,12 @@ static off_t get_delta_base(struct packed_git *p,\n>  \t * is stupid, as then a REF_DELTA would be smaller to store.\n>  \t */\n>  \tif (type == OBJ_OFS_DELTA) {\n> -\t\tunsigned used = 0;\n> -\t\tunsigned char c = base_info[used++];\n> -\t\tbase_offset = c & 127;\n> -\t\twhile (c & 128) {\n> -\t\t\tbase_offset += 1;\n> -\t\t\tif (!base_offset || MSB(base_offset, 7))\n> -\t\t\t\treturn 0;  /* overflow */\n> -\t\t\tc = base_info[used++];\n> -\t\t\tbase_offset = (base_offset << 7) + (c & 127);\n> -\t\t}\n> +\t\tconst unsigned char *cp = base_info;\n> +\t\tbase_offset = decode_varint(&cp);\n>  \t\tbase_offset = delta_obj_offset - base_offset;\n>  \t\tif (base_offset <= 0 || base_offset >= delta_obj_offset)\n>  \t\t\treturn 0;  /* out of bound */\n> -\t\t*curpos += used;\n> +\t\t*curpos += cp - base_info;\n>  \t} else if (type == OBJ_REF_DELTA) {\n>  \t\t/* The base entry _must_ be in the same pack */\n>  \t\tbase_offset = find_pack_entry_one(base_info, p);\n> -- \n> 1.8.4.38.g317e65b\n\nThis patch seems to be a cleanup independent from pack v4, it applies\ncleanly on master and passes all tests in itself.\n\nBest,\nGábor\n"},{"id":"226842","messageId":"20130905103011.GA20919@goldbirke","threadId":"34852","inReplyTo":"1378362001-1738-6-git-send-email-nico@fluxnic.net","subject":"Re: [PATCH 05/38] pack v4: add commit object parsing","fromName":"SZEDER Gábor","fromEmail":"szeder@ira.uka.de","sentAt":"2013-09-05T10:30:11Z","receivedAt":"2013-09-05T10:30:11Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"Hi,\n\n\nOn Thu, Sep 05, 2013 at 02:19:28AM -0400, Nicolas Pitre wrote:\n> Let's create another dictionary table to hold the author and committer\n> entries.  We use the same table format used for tree entries where the\n> 16 bit data prefix is conveniently used to store the timezone value.\n> \n> In order to copy straight from a commit object buffer, dict_add_entry()\n> is modified to get the string length as the provided string pointer is\n> not always be null terminated.\n> \n> Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n> ---\n>  packv4-create.c | 98 +++++++++++++++++++++++++++++++++++++++++++++++++++------\n>  1 file changed, 89 insertions(+), 9 deletions(-)\n> \n> diff --git a/packv4-create.c b/packv4-create.c\n> index eccd9fc..5c08871 100644\n> --- a/packv4-create.c\n> +++ b/packv4-create.c\n> @@ -1,5 +1,5 @@\n>  /*\n> - * packv4-create.c: management of dictionary tables used in pack v4\n> + * packv4-create.c: creation of dictionary tables and objects used in pack v4\n>   *\n>   * (C) Nicolas Pitre <nico@fluxnic.net>\n>   *\n> @@ -80,9 +80,9 @@ static void rehash_entries(struct dict_table *t)\n>  \t}\n>  }\n>  \n> -int dict_add_entry(struct dict_table *t, int val, const char *str)\n> +int dict_add_entry(struct dict_table *t, int val, const char *str, int str_len)\n>  {\n> -\tint i, val_len = 2, str_len = strlen(str) + 1;\n> +\tint i, val_len = 2;\n>  \n>  \tif (t->ptr + val_len + str_len > t->size) {\n\nWe need a +1 here on the left side, i.e.\n\n        if (t->ptr + val_len + str_len + 1 > t->size) {\n\nThe str_len variable accounted for the terminating null character\nbefore, but this patch removes str_len = strlen(str) + 1; above, and\ncallers specify the length of str without the terminating null in\nstr_len.  Thus it can lead to memory corruption, when the new entry\nhappens to end at 't->ptr + val_len + str_len' and the line added in\nthe next hunk writes the terminating null beyond the end of the\nbuffer.  I couldn't create a v4 pack from a current linux repo because\nof this; either glibc detected something or 'git packv4-create'\ncrashed.\n\nSidenote: couldn't we call the 'ptr' field something else, like\nend_offset or end_idx?  It took me some headscratching to figure out\nwhy is it OK to compare a pointer to an integer above, or use a\npointer without dereferencing as an index into an array below (because\nptr is, well, not a pointer after all).\n\n>  \t\tt->size = (t->size + val_len + str_len + 1024) * 3 / 2;\n> @@ -92,6 +92,7 @@ int dict_add_entry(struct dict_table *t, int val, const char *str)\n>  \tt->data[t->ptr] = val >> 8;\n>  \tt->data[t->ptr + 1] = val;\n>  \tmemcpy(t->data + t->ptr + val_len, str, str_len);\n> +\tt->data[t->ptr + val_len + str_len] = 0;\n>  \n>  \ti = (t->nb_entries) ?\n>  \t\tlocate_entry(t, t->data + t->ptr, val_len + str_len) : -1;\n> @@ -107,7 +108,7 @@ int dict_add_entry(struct dict_table *t, int val, const char *str)\n>  \tt->entry[t->nb_entries].offset = t->ptr;\n>  \tt->entry[t->nb_entries].size = val_len + str_len;\n>  \tt->entry[t->nb_entries].hits = 1;\n> -\tt->ptr += val_len + str_len;\n> +\tt->ptr += val_len + str_len + 1;\n\nGood.\n\n\nBest,\nGábor\n\n\n>  \tt->nb_entries++;\n>  \n>  \tif (t->hash_size * 3 <= t->nb_entries * 4)\n> @@ -135,8 +136,73 @@ static void sort_dict_entries_by_hits(struct dict_table *t)\n>  \trehash_entries(t);\n>  }\n>  \n> +static struct dict_table *commit_name_table;\n>  static struct dict_table *tree_path_table;\n>  \n> +/*\n> + * Parse the author/committer line from a canonical commit object.\n> + * The 'from' argument points right after the \"author \" or \"committer \"\n> + * string.  The time zone is parsed and stored in *tz_val.  The returned\n> + * pointer is right after the end of the email address which is also just\n> + * before the time value, or NULL if a parsing error is encountered.\n> + */\n> +static char *get_nameend_and_tz(char *from, int *tz_val)\n> +{\n> +\tchar *end, *tz;\n> +\n> +\ttz = strchr(from, '\\n');\n> +\t/* let's assume the smallest possible string to be \"x <x> 0 +0000\\n\" */\n> +\tif (!tz || tz - from < 13)\n> +\t\treturn NULL;\n> +\ttz -= 4;\n> +\tend = tz - 4;\n> +\twhile (end - from > 5 && *end != ' ')\n> +\t\tend--;\n> +\tif (end[-1] != '>' || end[0] != ' ' || tz[-2] != ' ')\n> +\t\treturn NULL;\n> +\t*tz_val = (tz[0] - '0') * 1000 +\n> +\t\t  (tz[1] - '0') * 100 +\n> +\t\t  (tz[2] - '0') * 10 +\n> +\t\t  (tz[3] - '0');\n> +\tswitch (tz[-1]) {\n> +\tdefault:\treturn NULL;\n> +\tcase '+':\tbreak;\n> +\tcase '-':\t*tz_val = -*tz_val;\n> +\t}\n> +\treturn end;\n> +}\n> +\n> +static int add_commit_dict_entries(void *buf, unsigned long size)\n> +{\n> +\tchar *name, *end = NULL;\n> +\tint tz_val;\n> +\n> +\tif (!commit_name_table)\n> +\t\tcommit_name_table = create_dict_table();\n> +\n> +\t/* parse and add author info */\n> +\tname = strstr(buf, \"\\nauthor \");\n> +\tif (name) {\n> +\t\tname += 8;\n> +\t\tend = get_nameend_and_tz(name, &tz_val);\n> +\t}\n> +\tif (!name || !end)\n> +\t\treturn -1;\n> +\tdict_add_entry(commit_name_table, tz_val, name, end - name);\n> +\n> +\t/* parse and add committer info */\n> +\tname = strstr(end, \"\\ncommitter \");\n> +\tif (name) {\n> +\t       name += 11;\n> +\t       end = get_nameend_and_tz(name, &tz_val);\n> +\t}\n> +\tif (!name || !end)\n> +\t\treturn -1;\n> +\tdict_add_entry(commit_name_table, tz_val, name, end - name);\n> +\n> +\treturn 0;\n> +}\n> +\n>  static int add_tree_dict_entries(void *buf, unsigned long size)\n>  {\n>  \tstruct tree_desc desc;\n> @@ -146,13 +212,16 @@ static int add_tree_dict_entries(void *buf, unsigned long size)\n>  \t\ttree_path_table = create_dict_table();\n>  \n>  \tinit_tree_desc(&desc, buf, size);\n> -\twhile (tree_entry(&desc, &name_entry))\n> +\twhile (tree_entry(&desc, &name_entry)) {\n> +\t\tint pathlen = tree_entry_len(&name_entry);\n>  \t\tdict_add_entry(tree_path_table, name_entry.mode,\n> -\t\t\t       name_entry.path);\n> +\t\t\t\tname_entry.path, pathlen);\n> +\t}\n> +\n>  \treturn 0;\n>  }\n>  \n> -void dict_dump(struct dict_table *t)\n> +void dump_dict_table(struct dict_table *t)\n>  {\n>  \tint i;\n>  \n> @@ -169,6 +238,12 @@ void dict_dump(struct dict_table *t)\n>  \t}\n>  }\n>  \n> +static void dict_dump(void)\n> +{\n> +\tdump_dict_table(commit_name_table);\n> +\tdump_dict_table(tree_path_table);\n> +}\n> +\n>  struct idx_entry\n>  {\n>  \toff_t                offset;\n> @@ -205,6 +280,7 @@ static int create_pack_dictionaries(struct packed_git *p)\n>  \t\tenum object_type type;\n>  \t\tunsigned long size;\n>  \t\tstruct object_info oi = {};\n> +\t\tint (*add_dict_entries)(void *, unsigned long);\n>  \n>  \t\toi.typep = &type;\n>  \t\toi.sizep = &size;\n> @@ -213,7 +289,11 @@ static int create_pack_dictionaries(struct packed_git *p)\n>  \t\t\t    sha1_to_hex(objects[i].sha1), p->pack_name);\n>  \n>  \t\tswitch (type) {\n> +\t\tcase OBJ_COMMIT:\n> +\t\t\tadd_dict_entries = add_commit_dict_entries;\n> +\t\t\tbreak;\n>  \t\tcase OBJ_TREE:\n> +\t\t\tadd_dict_entries = add_tree_dict_entries;\n>  \t\t\tbreak;\n>  \t\tdefault:\n>  \t\t\tcontinue;\n> @@ -225,7 +305,7 @@ static int create_pack_dictionaries(struct packed_git *p)\n>  \t\tif (check_sha1_signature(objects[i].sha1, data, size, typename(type)))\n>  \t\t\tdie(\"packed %s from %s is corrupt\",\n>  \t\t\t    sha1_to_hex(objects[i].sha1), p->pack_name);\n> -\t\tif (add_tree_dict_entries(data, size) < 0)\n> +\t\tif (add_dict_entries(data, size) < 0)\n>  \t\t\tdie(\"can't process %s object %s\",\n>  \t\t\t\ttypename(type), sha1_to_hex(objects[i].sha1));\n>  \t\tfree(data);\n> @@ -285,6 +365,6 @@ int main(int argc, char *argv[])\n>  \t\texit(1);\n>  \t}\n>  \tprocess_one_pack(argv[1]);\n> -\tdict_dump(tree_path_table);\n> +\tdict_dump();\n>  \treturn 0;\n>  }\n> -- \n> 1.8.4.38.g317e65b\n> \n> \n"},{"id":"226864","messageId":"alpine.LFD.2.03.1309051318550.14472@syhkavp.arg","threadId":"34852","inReplyTo":"20130905103011.GA20919@goldbirke","subject":"Re: [PATCH 05/38] pack v4: add commit object parsing","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T17:30:50Z","receivedAt":"2013-09-05T17:30:50Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 5 Sep 2013, SZEDER Gábor wrote:\n\n> Hi,\n> \n> \n> On Thu, Sep 05, 2013 at 02:19:28AM -0400, Nicolas Pitre wrote:\n> > Let's create another dictionary table to hold the author and committer\n> > entries.  We use the same table format used for tree entries where the\n> > 16 bit data prefix is conveniently used to store the timezone value.\n> > \n> > In order to copy straight from a commit object buffer, dict_add_entry()\n> > is modified to get the string length as the provided string pointer is\n> > not always be null terminated.\n> > \n> > Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n> > ---\n> >  packv4-create.c | 98 +++++++++++++++++++++++++++++++++++++++++++++++++++------\n> >  1 file changed, 89 insertions(+), 9 deletions(-)\n> > \n> > diff --git a/packv4-create.c b/packv4-create.c\n> > index eccd9fc..5c08871 100644\n> > --- a/packv4-create.c\n> > +++ b/packv4-create.c\n> > @@ -1,5 +1,5 @@\n> >  /*\n> > - * packv4-create.c: management of dictionary tables used in pack v4\n> > + * packv4-create.c: creation of dictionary tables and objects used in pack v4\n> >   *\n> >   * (C) Nicolas Pitre <nico@fluxnic.net>\n> >   *\n> > @@ -80,9 +80,9 @@ static void rehash_entries(struct dict_table *t)\n> >  \t}\n> >  }\n> >  \n> > -int dict_add_entry(struct dict_table *t, int val, const char *str)\n> > +int dict_add_entry(struct dict_table *t, int val, const char *str, int str_len)\n> >  {\n> > -\tint i, val_len = 2, str_len = strlen(str) + 1;\n> > +\tint i, val_len = 2;\n> >  \n> >  \tif (t->ptr + val_len + str_len > t->size) {\n> \n> We need a +1 here on the left side, i.e.\n> \n>         if (t->ptr + val_len + str_len + 1 > t->size) {\n\nAbsolutely, good catch.\n\n> Sidenote: couldn't we call the 'ptr' field something else, like\n> end_offset or end_idx?  It took me some headscratching to figure out\n> why is it OK to compare a pointer to an integer above, or use a\n> pointer without dereferencing as an index into an array below (because\n> ptr is, well, not a pointer after all).\n\nIndeed.  This is a remnant of an earlier implementation which didn't use \nrealloc() and therefore this used to be a real pointer.\n\nBoth issues now addressed in my tree.\n\nThanks\n\n\nNicolas\n"},{"id":"226870","messageId":"alpine.LFD.2.03.1309051445140.14472@syhkavp.arg","threadId":"34852","inReplyTo":"1378362001-1738-38-git-send-email-nico@fluxnic.net","subject":"Re: [PATCH 37/38] pack v4: introduce \"escape hatches\" in the name and path indexes","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T19:02:34Z","receivedAt":"2013-09-05T19:02:34Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 5 Sep 2013, Nicolas Pitre wrote:\n\n> If the path or name index is zero, this means the entry data is to be\n> found inline rather than being located in the dictionary table. This is\n> there to allow easy completion of thin packs without having to add new\n> table entries which would have required a full rewrite of the pack data.\n> \n> Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n\nI'm now dropping this patch.  Please also remove this from your \ndocumentation patch.\n\nI think that we've found a way to better support thin packs.\n\nYou said:\n\n> What if the sender prepares the sha-1 table to contain missing objects\n> in advance? The sender should know what base objects are missing. Then\n> we only need to append objects at the receiving end and verify that\n> all new objects are also present in the sha-1 table.\n\nSo the SHA1 table is covered.\n\nMissing objects in a thin pack cannot themselves be deltas.  We had \ntheir undeltified form at the end of a pack for the pack to be complete.  \nTherefore those missing objects serve only as base objects for other \ndeltas.\n\nAlthough this is possible to have deltified commit objects in pack v2, I \ndon't think this happens very often. There is no deltified commit \nobjects in pack v4.\n\nBlob objects are the same in pack v2 and pack v4.  No dictionary \nreferences are needed.\n\nThat leaves only tree objects.  And because we've also discussed the \nneed to have non transcoded object representations for those odd cases \nsuch as zero padded file modes, we might as well simply use that for the \nappended tree objects already needed to complete a thin pack.  At least \nthe strings in tree entries will be compressed that way.\n\nProblem solved, and one less special case in the code.\n\nWhat do you think?\n\n\nNicolas\n"},{"id":"226885","messageId":"alpine.LFD.2.03.1309051744490.20709@syhkavp.arg","threadId":"34852","inReplyTo":"alpine.LFD.2.03.1309051445140.14472@syhkavp.arg","subject":"Re: [PATCH 37/38] pack v4: introduce \"escape hatches\" in the name and path indexes","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-05T21:48:23Z","receivedAt":"2013-09-05T21:48:23Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 5 Sep 2013, Nicolas Pitre wrote:\n\n> On Thu, 5 Sep 2013, Nicolas Pitre wrote:\n> \n> > If the path or name index is zero, this means the entry data is to be\n> > found inline rather than being located in the dictionary table. This is\n> > there to allow easy completion of thin packs without having to add new\n> > table entries which would have required a full rewrite of the pack data.\n> > \n> > Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n> \n> I'm now dropping this patch.  Please also remove this from your \n> documentation patch.\n\nWell... I couldn't resist another little change that has been nagging me \nfor a while.\n\nBoth the author and committer time stamps are very closely related most \nof the time.  So the committer time stamp is now encoded as a difference \nagainst the author time stamp with the LSB indicating a negative \ndifference.\n\nOn git.git this saves 0.3% on the pack size.  Not much, but still \nimpressive for only a time stamp.\n\n\nNicolas\n"},{"id":"226897","messageId":"CACsJy8CyXNEc113U1cDbS3uR18sbALU4Uu_ULf_+EpYhwVdCmg@mail.gmail.com","threadId":"34852","inReplyTo":"alpine.LFD.2.03.1309051445140.14472@syhkavp.arg","subject":"Re: [PATCH 37/38] pack v4: introduce \"escape hatches\" in the name and path indexes","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-05T23:57:02Z","receivedAt":"2013-09-05T23:57:02Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Sep 6, 2013 at 2:02 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> I think that we've found a way to better support thin packs.\n>\n> You said:\n>\n>> What if the sender prepares the sha-1 table to contain missing objects\n>> in advance? The sender should know what base objects are missing. Then\n>> we only need to append objects at the receiving end and verify that\n>> all new objects are also present in the sha-1 table.\n>\n> So the SHA1 table is covered.\n>\n> Missing objects in a thin pack cannot themselves be deltas.  We had\n> their undeltified form at the end of a pack for the pack to be complete.\n> Therefore those missing objects serve only as base objects for other\n> deltas.\n>\n> Although this is possible to have deltified commit objects in pack v2, I\n> don't think this happens very often. There is no deltified commit\n> objects in pack v4.\n>\n> Blob objects are the same in pack v2 and pack v4.  No dictionary\n> references are needed.\n>\n> That leaves only tree objects.  And because we've also discussed the\n> need to have non transcoded object representations for those odd cases\n> such as zero padded file modes, we might as well simply use that for the\n> appended tree objects already needed to complete a thin pack.  At least\n> the strings in tree entries will be compressed that way.\n>\n> Problem solved, and one less special case in the code.\n>\n> What do you think?\n\nAgreed.\n\n>  Please also remove this from your documentation patch.\n\nWill do.\n-- \nDuy\n"},{"id":"226918","messageId":"xmqqvc2ezbbi.fsf@gitster.dls.corp.google.com","threadId":"34852","inReplyTo":"1378362001-1738-11-git-send-email-nico@fluxnic.net","subject":"Re: [PATCH 10/38] pack v4: commit object encoding","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-09-06T06:57:53Z","receivedAt":"2013-09-06T06:57:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@fluxnic.net> writes:\n\n> This goes as follows:\n>\n> - Tree reference: either variable length encoding of the index\n>   into the SHA1 table or the literal SHA1 prefixed by 0 (see\n>   encode_sha1ref()).\n>\n> - Parent count: variable length encoding of the number of parents.\n>   This is normally going to occupy a single byte but doesn't have to.\n>\n> - List of parent references: a list of encode_sha1ref() encoded\n>   references, or nothing if the parent count was zero.\n>\n> - Author reference: variable length encoding of an index into the author\n>   identifier dictionary table which also covers the time zone.  To make\n>   the overall encoding efficient, the author table is sorted by usage\n>   frequency so the most used names are first and require the shortest\n>   index encoding.\n>\n> - Author time stamp: variable length encoded.  Year 2038 ready!\n>\n> - Committer reference: same as author reference.\n>\n> - Committer time stamp: same as author time stamp.\n>\n> The remainder of the canonical commit object content is then zlib\n> compressed and appended to the above.\n>\n> Rationale: The most important commit object data is densely encoded while\n> requiring no zlib inflate processing on access, and all SHA1 references\n> are most likely to be direct indices into the pack index file requiring\n> no SHA1 search into the pack index file.\n\nMay I suggest a small change to the above, though.\n\nReorder the entries so that Parent count, list of parents and\ncommitter time stamp come first in this order, and then the rest.\n\nThat way, commit.c::parse_commit() could populate its field lazily\nwith parsing only the very minimum set of fields, and then the\nrevision walker, revision.c::add_parents_to_list(), can find where\nin the priority queue to add the parents to the list of commits to\nbe processed while still keeping the object partially parsed.  If a\ncommit is UNINTERESTING, no further parsing needs to be done.\n\n>\n> Signed-off-by: Nicolas Pitre <nico@fluxnic.net>\n> ---\n>  packv4-create.c | 119 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n>  1 file changed, 119 insertions(+)\n>\n> diff --git a/packv4-create.c b/packv4-create.c\n> index 12527c0..d4a79f4 100644\n> --- a/packv4-create.c\n> +++ b/packv4-create.c\n> @@ -14,6 +14,9 @@\n>  #include \"pack.h\"\n>  #include \"varint.h\"\n>  \n> +\n> +static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n> +\n>  struct data_entry {\n>  \tunsigned offset;\n>  \tunsigned size;\n> @@ -274,6 +277,122 @@ static int encode_sha1ref(const unsigned char *sha1, unsigned char *buf)\n>  \treturn 1 + 20;\n>  }\n>  \n> +/*\n> + * This converts a canonical commit object buffer into its\n> + * tightly packed representation using the already populated\n> + * and sorted commit_name_table dictionary.  The parsing is\n> + * strict so to ensure the canonical version may always be\n> + * regenerated and produce the same hash.\n> + */\n> +void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n> +{\n> +\tunsigned long size = *sizep;\n> +\tchar *in, *tail, *end;\n> +\tunsigned char *out;\n> +\tunsigned char sha1[20];\n> +\tint nb_parents, index, tz_val;\n> +\tunsigned long time;\n> +\tz_stream stream;\n> +\tint status;\n> +\n> +\t/*\n> +\t * It is guaranteed that the output is always going to be smaller\n> +\t * than the input.  We could even do this conversion in place.\n> +\t */\n> +\tin = buffer;\n> +\ttail = in + size;\n> +\tbuffer = xmalloc(size);\n> +\tout = buffer;\n> +\n> +\t/* parse the \"tree\" line */\n> +\tif (in + 46 >= tail || memcmp(in, \"tree \", 5) || in[45] != '\\n')\n> +\t\tgoto bad_data;\n> +\tif (get_sha1_lowhex(in + 5, sha1) < 0)\n> +\t\tgoto bad_data;\n> +\tin += 46;\n> +\tout += encode_sha1ref(sha1, out);\n> +\n> +\t/* count how many \"parent\" lines */\n> +\tnb_parents = 0;\n> +\twhile (in + 48 < tail && !memcmp(in, \"parent \", 7) && in[47] == '\\n') {\n> +\t\tnb_parents++;\n> +\t\tin += 48;\n> +\t}\n> +\tout += encode_varint(nb_parents, out);\n> +\n> +\t/* rewind and parse the \"parent\" lines */\n> +\tin -= 48 * nb_parents;\n> +\twhile (nb_parents--) {\n> +\t\tif (get_sha1_lowhex(in + 7, sha1))\n> +\t\t\tgoto bad_data;\n> +\t\tout += encode_sha1ref(sha1, out);\n> +\t\tin += 48;\n> +\t}\n> +\n> +\t/* parse the \"author\" line */\n> +\t/* it must be at least \"author x <x> 0 +0000\\n\" i.e. 21 chars */\n> +\tif (in + 21 >= tail || memcmp(in, \"author \", 7))\n> +\t\tgoto bad_data;\n> +\tin += 7;\n> +\tend = get_nameend_and_tz(in, &tz_val);\n> +\tif (!end)\n> +\t\tgoto bad_data;\n> +\tindex = dict_add_entry(commit_name_table, tz_val, in, end - in);\n> +\tif (index < 0)\n> +\t\tgoto bad_dict;\n> +\tout += encode_varint(index, out);\n> +\ttime = strtoul(end, &end, 10);\n> +\tif (!end || end[0] != ' ' || end[6] != '\\n')\n> +\t\tgoto bad_data;\n> +\tout += encode_varint(time, out);\n> +\tin = end + 7;\n> +\n> +\t/* parse the \"committer\" line */\n> +\t/* it must be at least \"committer x <x> 0 +0000\\n\" i.e. 24 chars */\n> +\tif (in + 24 >= tail || memcmp(in, \"committer \", 7))\n> +\t\tgoto bad_data;\n> +\tin += 10;\n> +\tend = get_nameend_and_tz(in, &tz_val);\n> +\tif (!end)\n> +\t\tgoto bad_data;\n> +\tindex = dict_add_entry(commit_name_table, tz_val, in, end - in);\n> +\tif (index < 0)\n> +\t\tgoto bad_dict;\n> +\tout += encode_varint(index, out);\n> +\ttime = strtoul(end, &end, 10);\n> +\tif (!end || end[0] != ' ' || end[6] != '\\n')\n> +\t\tgoto bad_data;\n> +\tout += encode_varint(time, out);\n> +\tin = end + 7;\n> +\n> +\t/* finally, deflate the remaining data */\n> +\tmemset(&stream, 0, sizeof(stream));\n> +\tdeflateInit(&stream, pack_compression_level);\n> +\tstream.next_in = (unsigned char *)in;\n> +\tstream.avail_in = tail - in;\n> +\tstream.next_out = (unsigned char *)out;\n> +\tstream.avail_out = size - (out - (unsigned char *)buffer);\n> +\tstatus = deflate(&stream, Z_FINISH);\n> +\tend = (char *)stream.next_out;\n> +\tdeflateEnd(&stream);\n> +\tif (status != Z_STREAM_END) {\n> +\t\terror(\"deflate error status %d\", status);\n> +\t\tgoto bad;\n> +\t}\n> +\n> +\t*sizep = end - (char *)buffer;\n> +\treturn buffer;\n> +\n> +bad_data:\n> +\terror(\"bad commit data\");\n> +\tgoto bad;\n> +bad_dict:\n> +\terror(\"bad dict entry\");\n> +bad:\n> +\tfree(buffer);\n> +\treturn NULL;\n> +}\n> +\n>  static struct pack_idx_entry *get_packed_object_list(struct packed_git *p)\n>  {\n>  \tunsigned i, nr_objects = p->num_objects;\n"},{"id":"226987","messageId":"alpine.LFD.2.03.1309061720330.20709@syhkavp.arg","threadId":"34852","inReplyTo":"xmqqvc2ezbbi.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH 10/38] pack v4: commit object encoding","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-06T21:28:04Z","receivedAt":"2013-09-06T21:28:04Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 5 Sep 2013, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@fluxnic.net> writes:\n> \n> > This goes as follows:\n> >\n> > - Tree reference: either variable length encoding of the index\n> >   into the SHA1 table or the literal SHA1 prefixed by 0 (see\n> >   encode_sha1ref()).\n> >\n> > - Parent count: variable length encoding of the number of parents.\n> >   This is normally going to occupy a single byte but doesn't have to.\n> >\n> > - List of parent references: a list of encode_sha1ref() encoded\n> >   references, or nothing if the parent count was zero.\n> >\n> > - Author reference: variable length encoding of an index into the author\n> >   identifier dictionary table which also covers the time zone.  To make\n> >   the overall encoding efficient, the author table is sorted by usage\n> >   frequency so the most used names are first and require the shortest\n> >   index encoding.\n> >\n> > - Author time stamp: variable length encoded.  Year 2038 ready!\n> >\n> > - Committer reference: same as author reference.\n> >\n> > - Committer time stamp: same as author time stamp.\n> >\n> > The remainder of the canonical commit object content is then zlib\n> > compressed and appended to the above.\n> >\n> > Rationale: The most important commit object data is densely encoded while\n> > requiring no zlib inflate processing on access, and all SHA1 references\n> > are most likely to be direct indices into the pack index file requiring\n> > no SHA1 search into the pack index file.\n> \n> May I suggest a small change to the above, though.\n> \n> Reorder the entries so that Parent count, list of parents and\n> committer time stamp come first in this order, and then the rest.\n> \n> That way, commit.c::parse_commit() could populate its field lazily\n> with parsing only the very minimum set of fields, and then the\n> revision walker, revision.c::add_parents_to_list(), can find where\n> in the priority queue to add the parents to the list of commits to\n> be processed while still keeping the object partially parsed.  If a\n> commit is UNINTERESTING, no further parsing needs to be done.\n\nOK.  If I understand correctly, the committer time stamp is more \nimportant than the author's, right?  Because my latest change in the \nformat was to make the former as a difference against the later and that \nwould obviously have to be reversed.\n\nAlso, to keep some kind of estetic symetry (if such thing may exist in a \nraw byte format) may I suggest keeping the tree reference first.  That \nis easy to skip over if you don't need it, something like:\n\n\tif (!*ptr)\n\t\tptr += 1 + 20;\n\telse\n\t\twhile (*ptr++ & 128);\n\nWhereas, for a checkout where only the tree info is needed, if it is \nlocated after the list of parents, then the above needs to be done for \nall those parents and the committer time.\n\n\nNicolas\n"},{"id":"226994","messageId":"xmqq61udwqlz.fsf@gitster.dls.corp.google.com","threadId":"34852","inReplyTo":"alpine.LFD.2.03.1309061720330.20709@syhkavp.arg","subject":"Re: [PATCH 10/38] pack v4: commit object encoding","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-09-06T22:08:08Z","receivedAt":"2013-09-06T22:08:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@fluxnic.net> writes:\n\n> OK.  If I understand correctly, the committer time stamp is more \n> important than the author's, right?\n\nYeah, it matters a lot more when doing timestamp based traversal\nwithout the reachability bitmaps.\n\n> ... may I suggest keeping the tree reference first.  That \n> is easy to skip over if you don't need it,...\n> ... Whereas, for a checkout where only the tree info is needed, if it is \n> located after the list of parents, then the above needs to be done for \n> all those parents and the committer time.\n\nHmm.  I wonder if that is a really good trade-off.\n\n\"checkout\" is to parse a single commit object and grab the \"tree\"\nfield, while \"log\" is to parse millions of commit objects to grab\ntheir \"parents\" and \"committer timestamp\" fields (\"log path/spec\"\nneeds to grab \"tree\", too, so that does not make \"tree\" extremely\nuncommon compared to the other two fields, though).\n\nI dunno.\n"},{"id":"227001","messageId":"alpine.LFD.2.03.1309070036450.20709@syhkavp.arg","threadId":"34852","inReplyTo":"xmqq61udwqlz.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH 10/38] pack v4: commit object encoding","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-07T04:41:29Z","receivedAt":"2013-09-07T04:41:29Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 6 Sep 2013, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@fluxnic.net> writes:\n> \n> > OK.  If I understand correctly, the committer time stamp is more \n> > important than the author's, right?\n> \n> Yeah, it matters a lot more when doing timestamp based traversal\n> without the reachability bitmaps.\n> \n> > ... may I suggest keeping the tree reference first.  That \n> > is easy to skip over if you don't need it,...\n> > ... Whereas, for a checkout where only the tree info is needed, if it is \n> > located after the list of parents, then the above needs to be done for \n> > all those parents and the committer time.\n> \n> Hmm.  I wonder if that is a really good trade-off.\n> \n> \"checkout\" is to parse a single commit object and grab the \"tree\"\n> field, while \"log\" is to parse millions of commit objects to grab\n> their \"parents\" and \"committer timestamp\" fields (\"log path/spec\"\n> needs to grab \"tree\", too, so that does not make \"tree\" extremely\n> uncommon compared to the other two fields, though).\n> \n> I dunno.\n\nI've therefore settled in the middle.  The patch description now looks \nlike:\n\n|    This goes as follows:\n|\n|    - Tree reference: either variable length encoding of the index\n|      into the SHA1 table or the literal SHA1 prefixed by 0 (see\n|      encode_sha1ref()).\n|\n|    - Parent count: variable length encoding of the number of parents.\n|      This is normally going to occupy a single byte but doesn't have to.\n|\n|    - List of parent references: a list of encode_sha1ref() encoded\n|      references, or nothing if the parent count was zero.\n|\n|    - Committer time stamp: variable length encoded.  Year 2038 ready!\n|      Unlike the canonical representation, this is stored close to the\n|      list of parents so the important data for history traversal can be\n|      retrieved without parsing the rest of the object.\n|\n|    - Committer reference: variable length encoding of an index into the\n|      ident dictionary table which also covers the time zone.  To make\n|      the overall encoding efficient, the ident table is sorted by usage\n|      frequency so the most used entries are first and require the shortest\n|      index encoding.\n|\n|    - Author time stamp: encoded as a difference against the committer\n|      time stamp, with the LSB used to indicate commit time is behind\n|      author time.\n|\n|    - Author reference: same as committer reference.\n|\n|    The remainder of the canonical commit object content is then zlib\n|    compressed and appended to the above.\n\nI also updated the documentation patch accordingly in my tree.\n\n\nNicolas\n"},{"id":"227006","messageId":"1378550599-25365-1-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 00/12] pack v4 support in index-pack","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:07Z","receivedAt":"2013-09-07T10:43:07Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This makes index-pack recognize pack v4. It still lacks:\n\n - the ability to walk through multi-base trees\n - thin pack support\n\nThe first is not easy to solve imo and but does not impact us in short\nterm because pack-objects probably will not learn to produce such\ntrees any time soon.\n\nThe second should be done after pack-objects can produce thin packs,\nelse it's hard to verify that the code works as expected.\n\nThis bases on Nico's tree, which does not really match the series this\npost is replied to due to some format changes. I don't know, maybe we\ncould share more code with packv4-parse.c. Right now I just need\nsomething that works and somewhat maintainable.\n\nNguyễn Thái Ngọc Duy (12):\n  pack v4: split pv4_create_dict() out of load_dict()\n  index-pack: split out varint decoding code\n  index-pack: do not allocate buffer for unpacking deltas in the first pass\n  index-pack: split inflate/digest code out of unpack_entry_data\n  index-pack: parse v4 header and dictionaries\n  index-pack: make sure all objects are registered in v4's SHA-1 table\n  index-pack: parse v4 commit format\n  index-pack: parse v4 tree format\n  index-pack: move delta base queuing code to unpack_raw_entry\n  index-pack: record all delta bases in v4 (tree and ref-delta)\n  index-pack: skip looking for ofs-deltas in v4 as they are not allowed\n  index-pack: resolve v4 one-base trees\n\n builtin/index-pack.c | 679 ++++++++++++++++++++++++++++++++++++++++++++-------\n packv4-parse.c       |  63 ++---\n packv4-parse.h       |   8 +\n 3 files changed, 627 insertions(+), 123 deletions(-)\n\n-- \n1.8.2.83.gc99314b\n"},{"id":"227007","messageId":"1378550599-25365-2-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 01/12] pack v4: split pv4_create_dict() out of load_dict()","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:08Z","receivedAt":"2013-09-07T10:43:08Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n packv4-parse.c | 63 ++++++++++++++++++++++++++++++++--------------------------\n packv4-parse.h |  8 ++++++++\n 2 files changed, 43 insertions(+), 28 deletions(-)\n\ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex 63bba03..82661ba 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -30,11 +30,38 @@ const unsigned char *get_sha1ref(struct packed_git *p,\n \treturn sha1;\n }\n \n-struct packv4_dict {\n-\tconst unsigned char *data;\n-\tunsigned int nb_entries;\n-\tunsigned int offsets[FLEX_ARRAY];\n-};\n+struct packv4_dict *pv4_create_dict(const unsigned char *data, int dict_size)\n+{\n+\tstruct packv4_dict *dict;\n+\tint i;\n+\n+\t/* count number of entries */\n+\tint nb_entries = 0;\n+\tconst unsigned char *cp = data;\n+\twhile (cp < data + dict_size - 3) {\n+\t\tcp += 2;  /* prefix bytes */\n+\t\tcp += strlen((const char *)cp);  /* entry string */\n+\t\tcp += 1;  /* terminating NUL */\n+\t\tnb_entries++;\n+\t}\n+\tif (cp - data != dict_size) {\n+\t\terror(\"dict size mismatch\");\n+\t\treturn NULL;\n+\t}\n+\n+\tdict = xmalloc(sizeof(*dict) + nb_entries * sizeof(dict->offsets[0]));\n+\tdict->data = data;\n+\tdict->nb_entries = nb_entries;\n+\n+\tcp = data;\n+\tfor (i = 0; i < nb_entries; i++) {\n+\t\tdict->offsets[i] = cp - data;\n+\t\tcp += 2;\n+\t\tcp += strlen((const char *)cp) + 1;\n+\t}\n+\n+\treturn dict;\n+}\n \n static struct packv4_dict *load_dict(struct packed_git *p, off_t *offset)\n {\n@@ -45,7 +72,7 @@ static struct packv4_dict *load_dict(struct packed_git *p, off_t *offset)\n \tconst unsigned char *cp;\n \tgit_zstream stream;\n \tstruct packv4_dict *dict;\n-\tint nb_entries, i, st;\n+\tint st;\n \n \t/* get uncompressed dictionary data size */\n \tsrc = use_pack(p, &w_curs, curpos, &avail);\n@@ -77,32 +104,12 @@ static struct packv4_dict *load_dict(struct packed_git *p, off_t *offset)\n \t\treturn NULL;\n \t}\n \n-\t/* count number of entries */\n-\tnb_entries = 0;\n-\tcp = data;\n-\twhile (cp < data + dict_size - 3) {\n-\t\tcp += 2;  /* prefix bytes */\n-\t\tcp += strlen((const char *)cp);  /* entry string */\n-\t\tcp += 1;  /* terminating NUL */\n-\t\tnb_entries++;\n-\t}\n-\tif (cp - data != dict_size) {\n-\t\terror(\"dict size mismatch\");\n+\tdict = pv4_create_dict(data, dict_size);\n+\tif (!dict) {\n \t\tfree(data);\n \t\treturn NULL;\n \t}\n \n-\tdict = xmalloc(sizeof(*dict) + nb_entries * sizeof(dict->offsets[0]));\n-\tdict->data = data;\n-\tdict->nb_entries = nb_entries;\n-\n-\tcp = data;\n-\tfor (i = 0; i < nb_entries; i++) {\n-\t\tdict->offsets[i] = cp - data;\n-\t\tcp += 2;\n-\t\tcp += strlen((const char *)cp) + 1;\n-\t}\n-\n \t*offset = curpos;\n \treturn dict;\n }\ndiff --git a/packv4-parse.h b/packv4-parse.h\nindex 5f9d809..0b2405a 100644\n--- a/packv4-parse.h\n+++ b/packv4-parse.h\n@@ -1,6 +1,14 @@\n #ifndef PACKV4_PARSE_H\n #define PACKV4_PARSE_H\n \n+struct packv4_dict {\n+\tconst unsigned char *data;\n+\tunsigned int nb_entries;\n+\tunsigned int offsets[FLEX_ARRAY];\n+};\n+\n+struct packv4_dict *pv4_create_dict(const unsigned char *data, int dict_size);\n+\n void *pv4_get_commit(struct packed_git *p, struct pack_window **w_curs,\n \t\t     off_t offset, unsigned long size);\n void *pv4_get_tree(struct packed_git *p, struct pack_window **w_curs,\n-- \n1.8.2.83.gc99314b\n"},{"id":"227008","messageId":"1378550599-25365-3-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 02/12] index-pack: split out varint decoding code","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:09Z","receivedAt":"2013-09-07T10:43:09Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 82 ++++++++++++++++++++++++++++------------------------\n 1 file changed, 45 insertions(+), 37 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 9c1cfac..5b1395d 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -275,6 +275,31 @@ static void use(int bytes)\n \tconsumed_bytes += bytes;\n }\n \n+static inline void *fill_and_use(int bytes)\n+{\n+\tvoid *p = fill(bytes);\n+\tuse(bytes);\n+\treturn p;\n+}\n+\n+static NORETURN void bad_object(unsigned long offset, const char *format,\n+\t\t       ...) __attribute__((format (printf, 2, 3)));\n+\n+static uintmax_t read_varint(void)\n+{\n+\tunsigned char c = *(char*)fill_and_use(1);\n+\tuintmax_t val = c & 127;\n+\twhile (c & 128) {\n+\t\tval += 1;\n+\t\tif (!val || MSB(val, 7))\n+\t\t\tbad_object(consumed_bytes,\n+\t\t\t\t   _(\"offset overflow in read_varint\"));\n+\t\tc = *(char*)fill_and_use(1);\n+\t\tval = (val << 7) + (c & 127);\n+\t}\n+\treturn val;\n+}\n+\n static const char *open_pack_file(const char *pack_name)\n {\n \tif (from_stdin) {\n@@ -315,9 +340,6 @@ static void parse_pack_header(void)\n \tuse(sizeof(struct pack_header));\n }\n \n-static NORETURN void bad_object(unsigned long offset, const char *format,\n-\t\t       ...) __attribute__((format (printf, 2, 3)));\n-\n static NORETURN void bad_object(unsigned long offset, const char *format, ...)\n {\n \tva_list params;\n@@ -455,55 +477,41 @@ static void *unpack_entry_data(unsigned long offset, unsigned long size,\n \treturn buf == fixed_buf ? NULL : buf;\n }\n \n+static void read_typesize_v2(struct object_entry *obj)\n+{\n+\tunsigned char c = *(char*)fill_and_use(1);\n+\tunsigned shift;\n+\n+\tobj->type = (c >> 4) & 7;\n+\tobj->size = (c & 15);\n+\tshift = 4;\n+\twhile (c & 128) {\n+\t\tc = *(char*)fill_and_use(1);\n+\t\tobj->size += (c & 0x7f) << shift;\n+\t\tshift += 7;\n+\t}\n+}\n+\n static void *unpack_raw_entry(struct object_entry *obj,\n \t\t\t      union delta_base *delta_base,\n \t\t\t      unsigned char *sha1)\n {\n-\tunsigned char *p;\n-\tunsigned long size, c;\n-\toff_t base_offset;\n-\tunsigned shift;\n \tvoid *data;\n+\tuintmax_t val;\n \n \tobj->idx.offset = consumed_bytes;\n \tinput_crc32 = crc32(0, NULL, 0);\n \n-\tp = fill(1);\n-\tc = *p;\n-\tuse(1);\n-\tobj->type = (c >> 4) & 7;\n-\tsize = (c & 15);\n-\tshift = 4;\n-\twhile (c & 0x80) {\n-\t\tp = fill(1);\n-\t\tc = *p;\n-\t\tuse(1);\n-\t\tsize += (c & 0x7f) << shift;\n-\t\tshift += 7;\n-\t}\n-\tobj->size = size;\n+\tread_typesize_v2(obj);\n \n \tswitch (obj->type) {\n \tcase OBJ_REF_DELTA:\n-\t\thashcpy(delta_base->sha1, fill(20));\n-\t\tuse(20);\n+\t\thashcpy(delta_base->sha1, fill_and_use(20));\n \t\tbreak;\n \tcase OBJ_OFS_DELTA:\n \t\tmemset(delta_base, 0, sizeof(*delta_base));\n-\t\tp = fill(1);\n-\t\tc = *p;\n-\t\tuse(1);\n-\t\tbase_offset = c & 127;\n-\t\twhile (c & 128) {\n-\t\t\tbase_offset += 1;\n-\t\t\tif (!base_offset || MSB(base_offset, 7))\n-\t\t\t\tbad_object(obj->idx.offset, _(\"offset value overflow for delta base object\"));\n-\t\t\tp = fill(1);\n-\t\t\tc = *p;\n-\t\t\tuse(1);\n-\t\t\tbase_offset = (base_offset << 7) + (c & 127);\n-\t\t}\n-\t\tdelta_base->offset = obj->idx.offset - base_offset;\n+\t\tval = read_varint();\n+\t\tdelta_base->offset = obj->idx.offset - val;\n \t\tif (delta_base->offset <= 0 || delta_base->offset >= obj->idx.offset)\n \t\t\tbad_object(obj->idx.offset, _(\"delta base offset is out of bound\"));\n \t\tbreak;\n-- \n1.8.2.83.gc99314b\n"},{"id":"227009","messageId":"1378550599-25365-4-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 03/12] index-pack: do not allocate buffer for unpacking deltas in the first pass","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:10Z","receivedAt":"2013-09-07T10:43:10Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"We do need deltas until the second pass. Allocating a buffer for it\nthen freeing later is wasteful is unnecessary. Make it use fixed_buf\n(aka large blob code path).\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 3 ++-\n 1 file changed, 2 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 5b1395d..a47cc34 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -446,7 +446,8 @@ static void *unpack_entry_data(unsigned long offset, unsigned long size,\n \t\tgit_SHA1_Update(&c, hdr, hdrlen);\n \t} else\n \t\tsha1 = NULL;\n-\tif (type == OBJ_BLOB && size > big_file_threshold)\n+\tif (is_delta_type(type) ||\n+\t     (type == OBJ_BLOB && size > big_file_threshold))\n \t\tbuf = fixed_buf;\n \telse\n \t\tbuf = xmalloc(size);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227010","messageId":"1378550599-25365-5-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 04/12] index-pack: split inflate/digest code out of unpack_entry_data","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:11Z","receivedAt":"2013-09-07T10:43:11Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 62 +++++++++++++++++++++++++++++++---------------------\n 1 file changed, 37 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex a47cc34..0dd7193 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -429,33 +429,19 @@ static int is_delta_type(enum object_type type)\n \treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n }\n \n-static void *unpack_entry_data(unsigned long offset, unsigned long size,\n-\t\t\t       enum object_type type, unsigned char *sha1)\n+static void read_and_inflate(unsigned long offset,\n+\t\t\t     void *buf, unsigned long size,\n+\t\t\t     unsigned long wraparound,\n+\t\t\t     git_SHA_CTX *ctx,\n+\t\t\t     unsigned char *sha1)\n {\n-\tstatic char fixed_buf[8192];\n-\tint status;\n \tgit_zstream stream;\n-\tvoid *buf;\n-\tgit_SHA_CTX c;\n-\tchar hdr[32];\n-\tint hdrlen;\n-\n-\tif (!is_delta_type(type)) {\n-\t\thdrlen = sprintf(hdr, \"%s %lu\", typename(type), size) + 1;\n-\t\tgit_SHA1_Init(&c);\n-\t\tgit_SHA1_Update(&c, hdr, hdrlen);\n-\t} else\n-\t\tsha1 = NULL;\n-\tif (is_delta_type(type) ||\n-\t     (type == OBJ_BLOB && size > big_file_threshold))\n-\t\tbuf = fixed_buf;\n-\telse\n-\t\tbuf = xmalloc(size);\n+\tint status;\n \n \tmemset(&stream, 0, sizeof(stream));\n \tgit_inflate_init(&stream);\n \tstream.next_out = buf;\n-\tstream.avail_out = buf == fixed_buf ? sizeof(fixed_buf) : size;\n+\tstream.avail_out = wraparound ? wraparound : size;\n \n \tdo {\n \t\tunsigned char *last_out = stream.next_out;\n@@ -464,17 +450,43 @@ static void *unpack_entry_data(unsigned long offset, unsigned long size,\n \t\tstatus = git_inflate(&stream, 0);\n \t\tuse(input_len - stream.avail_in);\n \t\tif (sha1)\n-\t\t\tgit_SHA1_Update(&c, last_out, stream.next_out - last_out);\n-\t\tif (buf == fixed_buf) {\n+\t\t\tgit_SHA1_Update(ctx, last_out, stream.next_out - last_out);\n+\t\tif (wraparound) {\n \t\t\tstream.next_out = buf;\n-\t\t\tstream.avail_out = sizeof(fixed_buf);\n+\t\t\tstream.avail_out = wraparound;\n \t\t}\n \t} while (status == Z_OK);\n \tif (stream.total_out != size || status != Z_STREAM_END)\n \t\tbad_object(offset, _(\"inflate returned %d\"), status);\n \tgit_inflate_end(&stream);\n \tif (sha1)\n-\t\tgit_SHA1_Final(sha1, &c);\n+\t\tgit_SHA1_Final(sha1, ctx);\n+}\n+\n+static void *unpack_entry_data(unsigned long offset, unsigned long size,\n+\t\t\t       enum object_type type, unsigned char *sha1)\n+{\n+\tstatic char fixed_buf[8192];\n+\tvoid *buf;\n+\tgit_SHA_CTX c;\n+\tchar hdr[32];\n+\tint hdrlen;\n+\n+\tif (!is_delta_type(type)) {\n+\t\thdrlen = sprintf(hdr, \"%s %lu\", typename(type), size) + 1;\n+\t\tgit_SHA1_Init(&c);\n+\t\tgit_SHA1_Update(&c, hdr, hdrlen);\n+\t} else\n+\t\tsha1 = NULL;\n+\tif (is_delta_type(type) ||\n+\t     (type == OBJ_BLOB && size > big_file_threshold))\n+\t\tbuf = fixed_buf;\n+\telse\n+\t\tbuf = xmalloc(size);\n+\n+\tread_and_inflate(offset, buf, size,\n+\t\t\t buf == fixed_buf ? sizeof(fixed_buf) : 0,\n+\t\t\t &c, sha1);\n \treturn buf == fixed_buf ? NULL : buf;\n }\n \n-- \n1.8.2.83.gc99314b\n"},{"id":"227011","messageId":"1378550599-25365-6-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 05/12] index-pack: parse v4 header and dictionaries","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:12Z","receivedAt":"2013-09-07T10:43:12Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 61 +++++++++++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 60 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 0dd7193..59b6c56 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -11,6 +11,7 @@\n #include \"exec_cmd.h\"\n #include \"streaming.h\"\n #include \"thread-utils.h\"\n+#include \"packv4-parse.h\"\n \n static const char index_pack_usage[] =\n \"git index-pack [-v] [-o <index-file>] [--keep | --keep=<msg>] [--verify] [--strict] (<pack-file> | --stdin [--fix-thin] [<pack-file>])\";\n@@ -70,6 +71,8 @@ struct delta_entry {\n static struct object_entry *objects;\n static struct delta_entry *deltas;\n static struct thread_local nothread_data;\n+static unsigned char *sha1_table;\n+static struct packv4_dict *name_dict, *path_dict;\n static int nr_objects;\n static int nr_deltas;\n static int nr_resolved_deltas;\n@@ -81,6 +84,7 @@ static int do_fsck_object;\n static int verbose;\n static int show_stat;\n static int check_self_contained_and_connected;\n+static int packv4;\n \n static struct progress *progress;\n \n@@ -300,6 +304,21 @@ static uintmax_t read_varint(void)\n \treturn val;\n }\n \n+static void *read_data(int size)\n+{\n+\tconst int max = sizeof(input_buffer);\n+\tvoid *buf;\n+\tchar *p;\n+\tp = buf = xmalloc(size);\n+\twhile (size) {\n+\t\tint to_fill = size > max ? max : size;\n+\t\tmemcpy(p, fill_and_use(to_fill), to_fill);\n+\t\tp += to_fill;\n+\t\tsize -= to_fill;\n+\t}\n+\treturn buf;\n+}\n+\n static const char *open_pack_file(const char *pack_name)\n {\n \tif (from_stdin) {\n@@ -332,7 +351,9 @@ static void parse_pack_header(void)\n \t/* Header consistency check */\n \tif (hdr->hdr_signature != htonl(PACK_SIGNATURE))\n \t\tdie(_(\"pack signature mismatch\"));\n-\tif (!pack_version_ok(hdr->hdr_version))\n+\tif (hdr->hdr_version == htonl(4))\n+\t\tpackv4 = 1;\n+\telse if (!pack_version_ok(hdr->hdr_version))\n \t\tdie(_(\"pack version %\"PRIu32\" unsupported\"),\n \t\t\tntohl(hdr->hdr_version));\n \n@@ -1013,6 +1034,31 @@ static void *threaded_second_pass(void *data)\n }\n #endif\n \n+static struct packv4_dict *read_dict(void)\n+{\n+\tunsigned long size;\n+\tunsigned char *data;\n+\tstruct packv4_dict *dict;\n+\n+\tsize = read_varint();\n+\tdata = xmallocz(size);\n+\tread_and_inflate(consumed_bytes, data, size, 0, NULL, NULL);\n+\tdict = pv4_create_dict(data, size);\n+\tif (!dict)\n+\t\tdie(\"unable to parse dictionary\");\n+\treturn dict;\n+}\n+\n+static void parse_dictionaries(void)\n+{\n+\tif (!packv4)\n+\t\treturn;\n+\n+\tsha1_table = read_data(20 * nr_objects);\n+\tname_dict = read_dict();\n+\tpath_dict = read_dict();\n+}\n+\n /*\n  * First pass:\n  * - find locations of all objects;\n@@ -1651,6 +1697,7 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \tparse_pack_header();\n \tobjects = xcalloc(nr_objects + 1, sizeof(struct object_entry));\n \tdeltas = xcalloc(nr_objects, sizeof(struct delta_entry));\n+\tparse_dictionaries();\n \tparse_pack_objects(pack_sha1);\n \tresolve_deltas();\n \tconclude_pack(fix_thin_pack, curr_pack, pack_sha1);\n@@ -1661,6 +1708,9 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \tif (show_stat)\n \t\tshow_pack_info(stat_only);\n \n+\tif (packv4)\n+\t\tdie(\"we're not there yet\");\n+\n \tidx_objects = xmalloc((nr_objects) * sizeof(struct pack_idx_entry *));\n \tfor (i = 0; i < nr_objects; i++)\n \t\tidx_objects[i] = &objects[i].idx;\n@@ -1677,6 +1727,15 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \tfree(objects);\n \tfree(index_name_buf);\n \tfree(keep_name_buf);\n+\tfree(sha1_table);\n+\tif (name_dict) {\n+\t\tfree((void*)name_dict->data);\n+\t\tfree(name_dict);\n+\t}\n+\tif (path_dict) {\n+\t\tfree((void*)path_dict->data);\n+\t\tfree(path_dict);\n+\t}\n \tif (pack_name == NULL)\n \t\tfree((void *) curr_pack);\n \tif (index_name == NULL)\n-- \n1.8.2.83.gc99314b\n"},{"id":"227012","messageId":"1378550599-25365-7-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 06/12] index-pack: make sure all objects are registered in v4's SHA-1 table","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:13Z","receivedAt":"2013-09-07T10:43:13Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 22 ++++++++++++++++++++--\n 1 file changed, 20 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 59b6c56..db2370d 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -742,6 +742,19 @@ static int check_collison(struct object_entry *entry)\n \treturn 0;\n }\n \n+static void check_against_sha1table(struct object_entry *obj)\n+{\n+\tconst unsigned char *found;\n+\tif (!packv4)\n+\t\treturn;\n+\n+\tfound = bsearch(obj->idx.sha1, sha1_table, nr_objects, 20,\n+\t\t\t(int (*)(const void *, const void *))hashcmp);\n+\tif (!found)\n+\t\tdie(_(\"object %s not found in SHA-1 table\"),\n+\t\t    sha1_to_hex(obj->idx.sha1));\n+}\n+\n static void sha1_object(const void *data, struct object_entry *obj_entry,\n \t\t\tunsigned long size, enum object_type type,\n \t\t\tconst unsigned char *sha1)\n@@ -910,6 +923,7 @@ static void resolve_delta(struct object_entry *delta_obj,\n \t\tbad_object(delta_obj->idx.offset, _(\"failed to apply delta\"));\n \thash_sha1_file(result->data, result->size,\n \t\t       typename(delta_obj->real_type), delta_obj->idx.sha1);\n+\tcheck_against_sha1table(delta_obj);\n \tsha1_object(result->data, NULL, result->size, delta_obj->real_type,\n \t\t    delta_obj->idx.sha1);\n \tcounter_lock();\n@@ -1087,8 +1101,12 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\t/* large blobs, check later */\n \t\t\tobj->real_type = OBJ_BAD;\n \t\t\tnr_delays++;\n-\t\t} else\n-\t\t\tsha1_object(data, NULL, obj->size, obj->type, obj->idx.sha1);\n+\t\t\tcheck_against_sha1table(obj);\n+\t\t} else {\n+\t\t\tcheck_against_sha1table(obj);\n+\t\t\tsha1_object(data, NULL, obj->size, obj->type,\n+\t\t\t\t    obj->idx.sha1);\n+\t\t}\n \t\tfree(data);\n \t\tdisplay_progress(progress, i+1);\n \t}\n-- \n1.8.2.83.gc99314b\n"},{"id":"227013","messageId":"1378550599-25365-8-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 07/12] index-pack: parse v4 commit format","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:14Z","receivedAt":"2013-09-07T10:43:14Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 95 ++++++++++++++++++++++++++++++++++++++++++++++++++--\n 1 file changed, 92 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex db2370d..210b78d 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -304,6 +304,30 @@ static uintmax_t read_varint(void)\n \treturn val;\n }\n \n+static const unsigned char *read_sha1ref(void)\n+{\n+\tunsigned int index = read_varint();\n+\tif (!index) {\n+\t\tstatic unsigned char sha1[20];\n+\t\thashcpy(sha1, fill_and_use(20));\n+\t\treturn sha1;\n+\t}\n+\tindex--;\n+\tif (index >= nr_objects)\n+\t\tbad_object(consumed_bytes,\n+\t\t\t   _(\"bad index in read_sha1ref\"));\n+\treturn sha1_table + index * 20;\n+}\n+\n+static const unsigned char *read_dictref(struct packv4_dict *dict)\n+{\n+\tunsigned int index = read_varint();\n+\tif (index >= dict->nb_entries)\n+\t\tbad_object(consumed_bytes,\n+\t\t\t   _(\"bad index in read_dictref\"));\n+\treturn  dict->data + dict->offsets[index];\n+}\n+\n static void *read_data(int size)\n {\n \tconst int max = sizeof(input_buffer);\n@@ -484,6 +508,59 @@ static void read_and_inflate(unsigned long offset,\n \t\tgit_SHA1_Final(sha1, ctx);\n }\n \n+static void *unpack_commit_v4(unsigned int offset,\n+\t\t\t      unsigned long size,\n+\t\t\t      unsigned char *sha1)\n+{\n+\tunsigned int nb_parents;\n+\tconst unsigned char *committer, *author, *ident;\n+\tunsigned long author_time, committer_time;\n+\tgit_SHA_CTX ctx;\n+\tchar hdr[32];\n+\tint hdrlen;\n+\tint16_t committer_tz, author_tz;\n+\tstruct strbuf dst;\n+\n+\tstrbuf_init(&dst, size);\n+\n+\tstrbuf_addf(&dst, \"tree %s\\n\", sha1_to_hex(read_sha1ref()));\n+\tnb_parents = read_varint();\n+\twhile (nb_parents--)\n+\t\tstrbuf_addf(&dst, \"parent %s\\n\", sha1_to_hex(read_sha1ref()));\n+\n+\tcommitter_time = read_varint();\n+\tident = read_dictref(name_dict);\n+\tcommitter_tz = (ident[0] << 8) | ident[1];\n+\tcommitter = ident + 2;\n+\n+\tauthor_time = read_varint();\n+\tident = read_dictref(name_dict);\n+\tauthor_tz = (ident[0] << 8) | ident[1];\n+\tauthor = ident + 2;\n+\n+\tif (author_time & 1)\n+\t\tauthor_time = committer_time + (author_time >> 1);\n+\telse\n+\t\tauthor_time = committer_time - (author_time >> 1);\n+\n+\tstrbuf_addf(&dst,\n+\t\t    \"author %s %lu %+05d\\n\"\n+\t\t    \"committer %s %lu %+05d\\n\",\n+\t\t    author, author_time, author_tz,\n+\t\t    committer, committer_time, committer_tz);\n+\n+\tif (dst.len > size)\n+\t\tbad_object(offset, _(\"bad commit\"));\n+\n+\thdrlen = sprintf(hdr, \"commit %lu\", size) + 1;\n+\tgit_SHA1_Init(&ctx);\n+\tgit_SHA1_Update(&ctx, hdr, hdrlen);\n+\tgit_SHA1_Update(&ctx, dst.buf, dst.len);\n+\tread_and_inflate(offset, dst.buf + dst.len, size - dst.len,\n+\t\t\t 0, &ctx, sha1);\n+\treturn dst.buf;\n+}\n+\n static void *unpack_entry_data(unsigned long offset, unsigned long size,\n \t\t\t       enum object_type type, unsigned char *sha1)\n {\n@@ -493,6 +570,9 @@ static void *unpack_entry_data(unsigned long offset, unsigned long size,\n \tchar hdr[32];\n \tint hdrlen;\n \n+\tif (type == OBJ_PV4_COMMIT)\n+\t\treturn unpack_commit_v4(offset, size, sha1);\n+\n \tif (!is_delta_type(type)) {\n \t\thdrlen = sprintf(hdr, \"%s %lu\", typename(type), size) + 1;\n \t\tgit_SHA1_Init(&c);\n@@ -536,7 +616,13 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \tobj->idx.offset = consumed_bytes;\n \tinput_crc32 = crc32(0, NULL, 0);\n \n-\tread_typesize_v2(obj);\n+\tif (packv4) {\n+\t\tval = read_varint();\n+\t\tobj->type = val & 15;\n+\t\tobj->size = val >> 4;\n+\t} else\n+\t\tread_typesize_v2(obj);\n+\tobj->real_type = obj->type;\n \n \tswitch (obj->type) {\n \tcase OBJ_REF_DELTA:\n@@ -554,6 +640,10 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n \t\tbreak;\n+\n+\tcase OBJ_PV4_COMMIT:\n+\t\tobj->real_type = OBJ_COMMIT;\n+\t\tbreak;\n \tdefault:\n \t\tbad_object(obj->idx.offset, _(\"unknown object type %d\"), obj->type);\n \t}\n@@ -1092,7 +1182,6 @@ static void parse_pack_objects(unsigned char *sha1)\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n \t\tvoid *data = unpack_raw_entry(obj, &delta->base, obj->idx.sha1);\n-\t\tobj->real_type = obj->type;\n \t\tif (is_delta_type(obj->type)) {\n \t\t\tnr_deltas++;\n \t\t\tdelta->obj_no = i;\n@@ -1104,7 +1193,7 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\tcheck_against_sha1table(obj);\n \t\t} else {\n \t\t\tcheck_against_sha1table(obj);\n-\t\t\tsha1_object(data, NULL, obj->size, obj->type,\n+\t\t\tsha1_object(data, NULL, obj->size, obj->real_type,\n \t\t\t\t    obj->idx.sha1);\n \t\t}\n \t\tfree(data);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227014","messageId":"1378550599-25365-9-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 08/12] index-pack: parse v4 tree format","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:15Z","receivedAt":"2013-09-07T10:43:15Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 104 +++++++++++++++++++++++++++++++++++++++++++++++++--\n 1 file changed, 100 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 210b78d..51ca64b 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -319,6 +319,21 @@ static const unsigned char *read_sha1ref(void)\n \treturn sha1_table + index * 20;\n }\n \n+static const unsigned char *read_sha1table_ref(void)\n+{\n+\tconst unsigned char *sha1 = read_sha1ref();\n+\tif (sha1 < sha1_table || sha1 >= sha1_table + nr_objects * 20) {\n+\t\tunsigned char *found;\n+\t\tfound = bsearch(sha1, sha1_table, nr_objects, 20,\n+\t\t\t\t(int (*)(const void *, const void *))hashcmp);\n+\t\tif (!found)\n+\t\t\tbad_object(consumed_bytes,\n+\t\t\t\t   _(\"SHA-1 %s not found in SHA-1 table\"),\n+\t\t\t\t   sha1_to_hex(sha1));\n+\t}\n+\treturn sha1;\n+}\n+\n static const unsigned char *read_dictref(struct packv4_dict *dict)\n {\n \tunsigned int index = read_varint();\n@@ -561,17 +576,93 @@ static void *unpack_commit_v4(unsigned int offset,\n \treturn dst.buf;\n }\n \n-static void *unpack_entry_data(unsigned long offset, unsigned long size,\n-\t\t\t       enum object_type type, unsigned char *sha1)\n+/*\n+ * v4 trees are actually kind of deltas and we don't do delta in the\n+ * first pass. This function only walks through a tree object to find\n+ * the end offset, register object dependencies and performs limited\n+ * validation.\n+ */\n+static void *unpack_tree_v4(struct object_entry *obj,\n+\t\t\t    unsigned int offset, unsigned long size,\n+\t\t\t    unsigned char *sha1)\n+{\n+\tunsigned int nr = read_varint();\n+\tconst unsigned char *last_base = NULL;\n+\tstruct strbuf sb = STRBUF_INIT;\n+\twhile (nr) {\n+\t\tunsigned int copy_start_or_path = read_varint();\n+\t\tif (copy_start_or_path & 1) { /* copy_start */\n+\t\t\tunsigned int copy_count = read_varint();\n+\t\t\tif (copy_count & 1) { /* first delta */\n+\t\t\t\tlast_base = read_sha1table_ref();\n+\t\t\t} else if (!last_base)\n+\t\t\t\tbad_object(offset,\n+\t\t\t\t\t   _(\"bad copy count index in unpack_tree_v4\"));\n+\t\t\tcopy_count >>= 1;\n+\t\t\tif (!copy_count)\n+\t\t\t\tbad_object(offset,\n+\t\t\t\t\t   _(\"bad copy count index in unpack_tree_v4\"));\n+\t\t\tnr -= copy_count;\n+\t\t} else {\t/* path */\n+\t\t\tunsigned int path_idx = copy_start_or_path >> 1;\n+\t\t\tconst unsigned char *entry_sha1;\n+\n+\t\t\tif (path_idx >= path_dict->nb_entries)\n+\t\t\t\tbad_object(offset,\n+\t\t\t\t\t   _(\"bad path index in unpack_tree_v4\"));\n+\t\t\tentry_sha1 = read_sha1ref();\n+\t\t\tnr--;\n+\n+\t\t\tif (!last_base) {\n+\t\t\t\tconst unsigned char *path;\n+\t\t\t\tunsigned mode;\n+\n+\t\t\t\tpath = path_dict->data + path_dict->offsets[path_idx];\n+\t\t\t\tmode = (path[0] << 8) | path[1];\n+\t\t\t\tstrbuf_addf(&sb, \"%o %s%c\", mode, path+2, '\\0');\n+\t\t\t\tstrbuf_add(&sb, entry_sha1, 20);\n+\t\t\t\tif (sb.len > size)\n+\t\t\t\t\tbad_object(offset,\n+\t\t\t\t\t\t   _(\"tree larger than expected\"));\n+\t\t\t}\n+\t\t}\n+\t}\n+\n+\tif (last_base) {\n+\t\tstrbuf_release(&sb);\n+\t\treturn NULL;\n+\t} else {\n+\t\tgit_SHA_CTX ctx;\n+\t\tchar hdr[32];\n+\t\tint hdrlen;\n+\n+\t\tif (sb.len != size)\n+\t\t\tbad_object(offset, _(\"tree size mismatch\"));\n+\n+\t\thdrlen = sprintf(hdr, \"tree %lu\", size) + 1;\n+\t\tgit_SHA1_Init(&ctx);\n+\t\tgit_SHA1_Update(&ctx, hdr, hdrlen);\n+\t\tgit_SHA1_Update(&ctx, sb.buf, size);\n+\t\tgit_SHA1_Final(sha1, &ctx);\n+\t\treturn strbuf_detach(&sb, NULL);\n+\t}\n+}\n+\n+static void *unpack_entry_data(struct object_entry *obj, unsigned char *sha1)\n {\n \tstatic char fixed_buf[8192];\n \tvoid *buf;\n \tgit_SHA_CTX c;\n \tchar hdr[32];\n \tint hdrlen;\n+\tunsigned long offset = obj->idx.offset;\n+\tunsigned long size = obj->size;\n+\tenum object_type type = obj->type;\n \n \tif (type == OBJ_PV4_COMMIT)\n \t\treturn unpack_commit_v4(offset, size, sha1);\n+\tif (type == OBJ_PV4_TREE)\n+\t\treturn unpack_tree_v4(obj, offset, size, sha1);\n \n \tif (!is_delta_type(type)) {\n \t\thdrlen = sprintf(hdr, \"%s %lu\", typename(type), size) + 1;\n@@ -640,16 +731,19 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n \t\tbreak;\n-\n \tcase OBJ_PV4_COMMIT:\n \t\tobj->real_type = OBJ_COMMIT;\n \t\tbreak;\n+\tcase OBJ_PV4_TREE:\n+\t\tobj->real_type = OBJ_TREE;\n+\t\tbreak;\n+\n \tdefault:\n \t\tbad_object(obj->idx.offset, _(\"unknown object type %d\"), obj->type);\n \t}\n \tobj->hdr_size = consumed_bytes - obj->idx.offset;\n \n-\tdata = unpack_entry_data(obj->idx.offset, obj->size, obj->type, sha1);\n+\tdata = unpack_entry_data(obj, sha1);\n \tobj->idx.crc32 = input_crc32;\n \treturn data;\n }\n@@ -1186,6 +1280,8 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\tnr_deltas++;\n \t\t\tdelta->obj_no = i;\n \t\t\tdelta++;\n+\t\t} else if (!data && obj->type == OBJ_PV4_TREE) {\n+\t\t\t/* delay sha1_object() until second pass */\n \t\t} else if (!data) {\n \t\t\t/* large blobs, check later */\n \t\t\tobj->real_type = OBJ_BAD;\n-- \n1.8.2.83.gc99314b\n"},{"id":"227015","messageId":"1378550599-25365-10-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 09/12] index-pack: move delta base queuing code to unpack_raw_entry","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:16Z","receivedAt":"2013-09-07T10:43:16Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"For v2, ofs-delta and ref-delta can only have queue one delta base at\na time. A v4 tree can have more than one delta base. Move the queuing\ncode up to unpack_raw_entry() and give unpack_tree_v4() more\nflexibility to add its bases.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 46 ++++++++++++++++++++++++++++++----------------\n 1 file changed, 30 insertions(+), 16 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 51ca64b..c5a8f68 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -576,6 +576,25 @@ static void *unpack_commit_v4(unsigned int offset,\n \treturn dst.buf;\n }\n \n+static void add_sha1_delta(struct object_entry *obj,\n+\t\t\t   const unsigned char *sha1)\n+{\n+\tstruct delta_entry *delta = deltas + nr_deltas;\n+\tdelta->obj_no = obj - objects;\n+\thashcpy(delta->base.sha1, sha1);\n+\tnr_deltas++;\n+}\n+\n+static void add_ofs_delta(struct object_entry *obj,\n+\t\t\t  off_t offset)\n+{\n+\tstruct delta_entry *delta = deltas + nr_deltas;\n+\tdelta->obj_no = obj - objects;\n+\tmemset(&delta->base, 0, sizeof(delta->base));\n+\tdelta->base.offset = offset;\n+\tnr_deltas++;\n+}\n+\n /*\n  * v4 trees are actually kind of deltas and we don't do delta in the\n  * first pass. This function only walks through a tree object to find\n@@ -698,17 +717,16 @@ static void read_typesize_v2(struct object_entry *obj)\n }\n \n static void *unpack_raw_entry(struct object_entry *obj,\n-\t\t\t      union delta_base *delta_base,\n \t\t\t      unsigned char *sha1)\n {\n \tvoid *data;\n-\tuintmax_t val;\n+\toff_t offset;\n \n \tobj->idx.offset = consumed_bytes;\n \tinput_crc32 = crc32(0, NULL, 0);\n \n \tif (packv4) {\n-\t\tval = read_varint();\n+\t\tuintmax_t val = read_varint();\n \t\tobj->type = val & 15;\n \t\tobj->size = val >> 4;\n \t} else\n@@ -717,14 +735,14 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \n \tswitch (obj->type) {\n \tcase OBJ_REF_DELTA:\n-\t\thashcpy(delta_base->sha1, fill_and_use(20));\n+\t\tadd_sha1_delta(obj, fill_and_use(20));\n \t\tbreak;\n \tcase OBJ_OFS_DELTA:\n-\t\tmemset(delta_base, 0, sizeof(*delta_base));\n-\t\tval = read_varint();\n-\t\tdelta_base->offset = obj->idx.offset - val;\n-\t\tif (delta_base->offset <= 0 || delta_base->offset >= obj->idx.offset)\n-\t\t\tbad_object(obj->idx.offset, _(\"delta base offset is out of bound\"));\n+\t\toffset = obj->idx.offset - read_varint();\n+\t\tif (offset <= 0 || offset >= obj->idx.offset)\n+\t\t\tbad_object(obj->idx.offset,\n+\t\t\t\t   _(\"delta base offset is out of bound\"));\n+\t\tadd_ofs_delta(obj, offset);\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n@@ -1266,7 +1284,6 @@ static void parse_dictionaries(void)\n static void parse_pack_objects(unsigned char *sha1)\n {\n \tint i, nr_delays = 0;\n-\tstruct delta_entry *delta = deltas;\n \tstruct stat st;\n \n \tif (verbose)\n@@ -1275,12 +1292,9 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\t\tnr_objects);\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n-\t\tvoid *data = unpack_raw_entry(obj, &delta->base, obj->idx.sha1);\n-\t\tif (is_delta_type(obj->type)) {\n-\t\t\tnr_deltas++;\n-\t\t\tdelta->obj_no = i;\n-\t\t\tdelta++;\n-\t\t} else if (!data && obj->type == OBJ_PV4_TREE) {\n+\t\tvoid *data = unpack_raw_entry(obj, obj->idx.sha1);\n+\t\tif (is_delta_type(obj->type) ||\n+\t\t    (!data && obj->type == OBJ_PV4_TREE)) {\n \t\t\t/* delay sha1_object() until second pass */\n \t\t} else if (!data) {\n \t\t\t/* large blobs, check later */\n-- \n1.8.2.83.gc99314b\n"},{"id":"227016","messageId":"1378550599-25365-11-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 10/12] index-pack: record all delta bases in v4 (tree and ref-delta)","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:17Z","receivedAt":"2013-09-07T10:43:17Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 33 ++++++++++++++++++++++++++++++---\n 1 file changed, 30 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex c5a8f68..33722e1 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -24,6 +24,7 @@ struct object_entry {\n \tenum object_type real_type;\n \tunsigned delta_depth;\n \tint base_object_no;\n+\tint nr_bases;\t\t/* only valid for v4 trees */\n };\n \n union delta_base {\n@@ -489,6 +490,11 @@ static int is_delta_type(enum object_type type)\n \treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n }\n \n+static int is_delta_tree(const struct object_entry *obj)\n+{\n+\treturn obj->type == OBJ_PV4_TREE && obj->nr_bases > 0;\n+}\n+\n static void read_and_inflate(unsigned long offset,\n \t\t\t     void *buf, unsigned long size,\n \t\t\t     unsigned long wraparound,\n@@ -595,6 +601,20 @@ static void add_ofs_delta(struct object_entry *obj,\n \tnr_deltas++;\n }\n \n+static void add_tree_delta_base(struct object_entry *obj,\n+\t\t\t\tconst unsigned char *base,\n+\t\t\t\tint delta_start)\n+{\n+\tint i;\n+\n+\tfor (i = delta_start; i < nr_deltas; i++)\n+\t\tif (!hashcmp(base, deltas[i].base.sha1))\n+\t\t\treturn;\n+\n+\tadd_sha1_delta(obj, base);\n+\tobj->nr_bases++;\n+}\n+\n /*\n  * v4 trees are actually kind of deltas and we don't do delta in the\n  * first pass. This function only walks through a tree object to find\n@@ -608,12 +628,14 @@ static void *unpack_tree_v4(struct object_entry *obj,\n \tunsigned int nr = read_varint();\n \tconst unsigned char *last_base = NULL;\n \tstruct strbuf sb = STRBUF_INIT;\n+\tint delta_start = nr_deltas;\n \twhile (nr) {\n \t\tunsigned int copy_start_or_path = read_varint();\n \t\tif (copy_start_or_path & 1) { /* copy_start */\n \t\t\tunsigned int copy_count = read_varint();\n \t\t\tif (copy_count & 1) { /* first delta */\n \t\t\t\tlast_base = read_sha1table_ref();\n+\t\t\t\tadd_tree_delta_base(obj, last_base, delta_start);\n \t\t\t} else if (!last_base)\n \t\t\t\tbad_object(offset,\n \t\t\t\t\t   _(\"bad copy count index in unpack_tree_v4\"));\n@@ -735,9 +757,15 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \n \tswitch (obj->type) {\n \tcase OBJ_REF_DELTA:\n-\t\tadd_sha1_delta(obj, fill_and_use(20));\n+\t\tif (packv4)\n+\t\t\tadd_sha1_delta(obj, read_sha1table_ref());\n+\t\telse\n+\t\t\tadd_sha1_delta(obj, fill_and_use(20));\n \t\tbreak;\n \tcase OBJ_OFS_DELTA:\n+\t\tif (packv4)\n+\t\t\tdie(_(\"pack version 4 does not support ofs-delta type (offset %lu)\"),\n+\t\t\t    obj->idx.offset);\n \t\toffset = obj->idx.offset - read_varint();\n \t\tif (offset <= 0 || offset >= obj->idx.offset)\n \t\t\tbad_object(obj->idx.offset,\n@@ -1293,8 +1321,7 @@ static void parse_pack_objects(unsigned char *sha1)\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n \t\tvoid *data = unpack_raw_entry(obj, obj->idx.sha1);\n-\t\tif (is_delta_type(obj->type) ||\n-\t\t    (!data && obj->type == OBJ_PV4_TREE)) {\n+\t\tif (is_delta_type(obj->type) || is_delta_tree(obj)) {\n \t\t\t/* delay sha1_object() until second pass */\n \t\t} else if (!data) {\n \t\t\t/* large blobs, check later */\n-- \n1.8.2.83.gc99314b\n"},{"id":"227017","messageId":"1378550599-25365-12-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 11/12] index-pack: skip looking for ofs-deltas in v4 as they are not allowed","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:18Z","receivedAt":"2013-09-07T10:43:18Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 11 +++++++----\n 1 file changed, 7 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 33722e1..1fa74f4 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -1171,10 +1171,13 @@ static struct base_data *find_unresolved_deltas_1(struct base_data *base,\n \t\tfind_delta_children(&base_spec,\n \t\t\t\t    &base->ref_first, &base->ref_last, OBJ_REF_DELTA);\n \n-\t\tmemset(&base_spec, 0, sizeof(base_spec));\n-\t\tbase_spec.offset = base->obj->idx.offset;\n-\t\tfind_delta_children(&base_spec,\n-\t\t\t\t    &base->ofs_first, &base->ofs_last, OBJ_OFS_DELTA);\n+\t\tif (!packv4) {\n+\t\t\tmemset(&base_spec, 0, sizeof(base_spec));\n+\t\t\tbase_spec.offset = base->obj->idx.offset;\n+\t\t\tfind_delta_children(&base_spec,\n+\t\t\t\t\t    &base->ofs_first, &base->ofs_last,\n+\t\t\t\t\t    OBJ_OFS_DELTA);\n+\t\t}\n \n \t\tif (base->ref_last == -1 && base->ofs_last == -1) {\n \t\t\tfree(base->data);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227018","messageId":"1378550599-25365-13-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 12/12] index-pack: resolve v4 one-base trees","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-07T10:43:19Z","receivedAt":"2013-09-07T10:43:19Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This is the most common case for delta trees. In fact it's the only\nkind that's produced by packv4-create. It fits well in the way\nindex-pack resolves deltas and benefits from threading (the set of\nobjects depending on this base does not overlap with the set of\nobjects depending on another base)\n\nMulti-base trees will be probably processed differently.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 194 ++++++++++++++++++++++++++++++++++++++++++++++-----\n 1 file changed, 178 insertions(+), 16 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 1fa74f4..4a24bc3 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -12,6 +12,8 @@\n #include \"streaming.h\"\n #include \"thread-utils.h\"\n #include \"packv4-parse.h\"\n+#include \"varint.h\"\n+#include \"tree-walk.h\"\n \n static const char index_pack_usage[] =\n \"git index-pack [-v] [-o <index-file>] [--keep | --keep=<msg>] [--verify] [--strict] (<pack-file> | --stdin [--fix-thin] [<pack-file>])\";\n@@ -38,8 +40,8 @@ struct base_data {\n \tstruct object_entry *obj;\n \tvoid *data;\n \tunsigned long size;\n-\tint ref_first, ref_last;\n-\tint ofs_first, ofs_last;\n+\tint ref_first, ref_last, tree_first;\n+\tint ofs_first, ofs_last, tree_last;\n };\n \n #if !defined(NO_PTHREADS) && defined(NO_THREAD_SAFE_PREAD)\n@@ -437,6 +439,7 @@ static struct base_data *alloc_base_data(void)\n \tmemset(base, 0, sizeof(*base));\n \tbase->ref_last = -1;\n \tbase->ofs_last = -1;\n+\tbase->tree_last = -1;\n \treturn base;\n }\n \n@@ -670,6 +673,8 @@ static void *unpack_tree_v4(struct object_entry *obj,\n \t}\n \n \tif (last_base) {\n+\t\tif (nr_deltas - delta_start > 1)\n+\t\t\tdie(\"sorry guys, multi-base trees are not supported yet\");\n \t\tstrbuf_release(&sb);\n \t\treturn NULL;\n \t} else {\n@@ -794,6 +799,83 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \treturn data;\n }\n \n+static void *patch_one_base_tree(const struct object_entry *src,\n+\t\t\t\t const unsigned char *src_buf,\n+\t\t\t\t const unsigned char *delta_buf,\n+\t\t\t\t unsigned long delta_size,\n+\t\t\t\t unsigned long *dst_size)\n+{\n+\tunsigned int nr;\n+\tconst unsigned char *last_base = NULL;\n+\tstruct strbuf sb = STRBUF_INIT;\n+\tconst unsigned char *p = delta_buf;\n+\n+\tnr = decode_varint(&p);\n+\twhile (nr && p < delta_buf + delta_size) {\n+\t\tunsigned int copy_start_or_path = decode_varint(&p);\n+\t\tif (copy_start_or_path & 1) { /* copy_start */\n+\t\t\tstruct tree_desc desc;\n+\t\t\tstruct name_entry entry;\n+\t\t\tunsigned int copy_count = decode_varint(&p);\n+\t\t\tunsigned int copy_start = copy_start_or_path >> 1;\n+\t\t\tif (!src)\n+\t\t\t\tdie(\"we are not supposed to copy from another tree!\");\n+\t\t\tif (copy_count & 1) { /* first delta */\n+\t\t\t\tunsigned int id = decode_varint(&p);\n+\t\t\t\tif (!id) {\n+\t\t\t\t\tlast_base = p;\n+\t\t\t\t\tp += 20;\n+\t\t\t\t} else\n+\t\t\t\t\tlast_base = sha1_table + (id - 1) * 20;\n+\t\t\t\tif (hashcmp(last_base, src->idx.sha1))\n+\t\t\t\t\tdie(_(\"bad tree base in patch_one_base_tree\"));\n+\t\t\t} else if (!last_base)\n+\t\t\t\tdie(_(\"bad copy count index in patch_one_base_tree\"));\n+\t\t\tcopy_count >>= 1;\n+\t\t\tif (!copy_count)\n+\t\t\t\tdie(_(\"bad copy count index in patch_one_base_tree\"));\n+\t\t\tnr -= copy_count;\n+\n+\t\t\tinit_tree_desc(&desc, src_buf, src->size);\n+\t\t\twhile (tree_entry(&desc, &entry)) {\n+\t\t\t\tif (copy_start)\n+\t\t\t\t\tcopy_start--;\n+\t\t\t\telse if (copy_count) {\n+\t\t\t\t\tstrbuf_addf(&sb, \"%o %s%c\", entry.mode, entry.path, '\\0');\n+\t\t\t\t\tstrbuf_add(&sb, entry.sha1, 20);\n+\t\t\t\t\tcopy_count--;\n+\t\t\t\t} else\n+\t\t\t\t\tbreak;\n+\t\t\t}\n+\t\t} else {\t/* path */\n+\t\t\tunsigned int path_idx = copy_start_or_path >> 1;\n+\t\t\tconst unsigned char *path;\n+\t\t\tunsigned mode;\n+\t\t\tunsigned int id;\n+\t\t\tconst unsigned char *entry_sha1;\n+\n+\t\t\tif (path_idx >= path_dict->nb_entries)\n+\t\t\t\tdie(_(\"bad path index in unpack_tree_v4\"));\n+\t\t\tid = decode_varint(&p);\n+\t\t\tif (!id) {\n+\t\t\t\tentry_sha1 = p;\n+\t\t\t\tp += 20;\n+\t\t\t} else\n+\t\t\t\tentry_sha1 = sha1_table + (id - 1) * 20;\n+\t\t\tnr--;\n+\n+\t\t\tpath = path_dict->data + path_dict->offsets[path_idx];\n+\t\t\tmode = (path[0] << 8) | path[1];\n+\t\t\tstrbuf_addf(&sb, \"%o %s%c\", mode, path+2, '\\0');\n+\t\t\tstrbuf_add(&sb, entry_sha1, 20);\n+\t\t}\n+\t}\n+\tif (nr != 0 || p != delta_buf + delta_size)\n+\t\tdie(_(\"bad delta tree\"));\n+\t*dst_size = sb.len;\n+\treturn sb.buf;\n+}\n+\n static void *unpack_data(struct object_entry *obj,\n \t\t\t int (*consume)(const unsigned char *, unsigned long, void *),\n \t\t\t void *cb_data)\n@@ -855,8 +937,33 @@ static void *unpack_data(struct object_entry *obj,\n \treturn data;\n }\n \n+static void *get_tree_v4_from_pack(struct object_entry *obj,\n+\t\t\t\t   unsigned long *len_p)\n+{\n+\toff_t from = obj[0].idx.offset + obj[0].hdr_size;\n+\tunsigned long len = obj[1].idx.offset - from;\n+\tunsigned char *data;\n+\tssize_t n;\n+\n+\tdata = xmalloc(len);\n+\tn = pread(pack_fd, data, len, from);\n+\tif (n < 0)\n+\t\tdie_errno(_(\"cannot pread pack file\"));\n+\tif (!n)\n+\t\tdie(Q_(\"premature end of pack file, %lu byte missing\",\n+\t\t       \"premature end of pack file, %lu bytes missing\",\n+\t\t       len),\n+\t\t    len);\n+\tif (len_p)\n+\t\t*len_p = len;\n+\treturn data;\n+}\n+\n static void *get_data_from_pack(struct object_entry *obj)\n {\n+\tif (obj->type == OBJ_PV4_COMMIT || obj->type == OBJ_PV4_TREE)\n+\t\tdie(\"BUG: unsupported code path\");\n+\n \treturn unpack_data(obj, NULL, NULL);\n }\n \n@@ -1096,14 +1203,25 @@ static void *get_base_data(struct base_data *c)\n \t\tstruct object_entry *obj = c->obj;\n \t\tstruct base_data **delta = NULL;\n \t\tint delta_nr = 0, delta_alloc = 0;\n+\t\tunsigned long size, len;\n \n-\t\twhile (is_delta_type(c->obj->type) && !c->data) {\n+\t\twhile ((is_delta_type(c->obj->type) ||\n+\t\t\t(c->base && c->obj->type == OBJ_PV4_TREE)) &&\n+\t\t       !c->data) {\n \t\t\tALLOC_GROW(delta, delta_nr + 1, delta_alloc);\n \t\t\tdelta[delta_nr++] = c;\n \t\t\tc = c->base;\n \t\t}\n \t\tif (!delta_nr) {\n-\t\t\tc->data = get_data_from_pack(obj);\n+\t\t\tif (c->obj->type == OBJ_PV4_TREE) {\n+\t\t\t\tvoid *tree_v4 = get_tree_v4_from_pack(obj, &len);\n+\t\t\t\tc->data = patch_one_base_tree(NULL, NULL,\n+\t\t\t\t\t\t\t      tree_v4, len, &size);\n+\t\t\t\tif (size != obj->size)\n+\t\t\t\t\tdie(\"size mismatch\");\n+\t\t\t\tfree(tree_v4);\n+\t\t\t} else\n+\t\t\t\tc->data = get_data_from_pack(obj);\n \t\t\tc->size = obj->size;\n \t\t\tget_thread_data()->base_cache_used += c->size;\n \t\t\tprune_base_data(c);\n@@ -1113,11 +1231,18 @@ static void *get_base_data(struct base_data *c)\n \t\t\tc = delta[delta_nr - 1];\n \t\t\tobj = c->obj;\n \t\t\tbase = get_base_data(c->base);\n-\t\t\traw = get_data_from_pack(obj);\n-\t\t\tc->data = patch_delta(\n-\t\t\t\tbase, c->base->size,\n-\t\t\t\traw, obj->size,\n-\t\t\t\t&c->size);\n+\t\t\tif (c->obj->type == OBJ_PV4_TREE) {\n+\t\t\t\traw = get_tree_v4_from_pack(obj, &len);\n+\t\t\t\tc->data = patch_one_base_tree(c->base->obj, base,\n+\t\t\t\t\t\t\t      raw, len, &size);\n+\t\t\t\tif (size != obj->size)\n+\t\t\t\t\tdie(\"size mismatch\");\n+\t\t\t} else {\n+\t\t\t\traw = get_data_from_pack(obj);\n+\t\t\t\tc->data = patch_delta(base, c->base->size,\n+\t\t\t\t\t\t      raw, obj->size,\n+\t\t\t\t\t\t      &c->size);\n+\t\t\t}\n \t\t\tfree(raw);\n \t\t\tif (!c->data)\n \t\t\t\tbad_object(obj->idx.offset, _(\"failed to apply delta\"));\n@@ -1133,6 +1258,8 @@ static void resolve_delta(struct object_entry *delta_obj,\n \t\t\t  struct base_data *base, struct base_data *result)\n {\n \tvoid *base_data, *delta_data;\n+\tint tree_v4 = delta_obj->type == OBJ_PV4_TREE;\n+\tunsigned long tree_size;\n \n \tdelta_obj->real_type = base->obj->real_type;\n \tif (show_stat) {\n@@ -1143,10 +1270,18 @@ static void resolve_delta(struct object_entry *delta_obj,\n \t\tdeepest_delta_unlock();\n \t}\n \tdelta_obj->base_object_no = base->obj - objects;\n-\tdelta_data = get_data_from_pack(delta_obj);\n+\tif (tree_v4)\n+\t\tdelta_data = get_tree_v4_from_pack(delta_obj, &tree_size);\n+\telse\n+\t\tdelta_data = get_data_from_pack(delta_obj);\n \tbase_data = get_base_data(base);\n \tresult->obj = delta_obj;\n-\tresult->data = patch_delta(base_data, base->size,\n+\tif (tree_v4)\n+\t\tresult->data = patch_one_base_tree(base->obj, base_data,\n+\t\t\t\t\t\t   delta_data, tree_size,\n+\t\t\t\t\t\t   &result->size);\n+\telse\n+\t\tresult->data = patch_delta(base_data, base->size,\n \t\t\t\t   delta_data, delta_obj->size, &result->size);\n \tfree(delta_data);\n \tif (!result->data)\n@@ -1164,7 +1299,8 @@ static void resolve_delta(struct object_entry *delta_obj,\n static struct base_data *find_unresolved_deltas_1(struct base_data *base,\n \t\t\t\t\t\t  struct base_data *prev_base)\n {\n-\tif (base->ref_last == -1 && base->ofs_last == -1) {\n+\tif (base->ref_last == -1 && base->ofs_last == -1 &&\n+\t    base->tree_last == -1) {\n \t\tunion delta_base base_spec;\n \n \t\thashcpy(base_spec.sha1, base->obj->idx.sha1);\n@@ -1177,9 +1313,15 @@ static struct base_data *find_unresolved_deltas_1(struct base_data *base,\n \t\t\tfind_delta_children(&base_spec,\n \t\t\t\t\t    &base->ofs_first, &base->ofs_last,\n \t\t\t\t\t    OBJ_OFS_DELTA);\n+\t\t} else {\n+\t\t\thashcpy(base_spec.sha1, base->obj->idx.sha1);\n+\t\t\tfind_delta_children(&base_spec,\n+\t\t\t\t\t    &base->tree_first, &base->tree_last,\n+\t\t\t\t\t    OBJ_PV4_TREE);\n \t\t}\n \n-\t\tif (base->ref_last == -1 && base->ofs_last == -1) {\n+\t\tif (base->ref_last == -1 && base->ofs_last == -1 &&\n+\t\t    base->tree_last == -1) {\n \t\t\tfree(base->data);\n \t\t\treturn NULL;\n \t\t}\n@@ -1213,6 +1355,25 @@ static struct base_data *find_unresolved_deltas_1(struct base_data *base,\n \t\treturn result;\n \t}\n \n+\twhile (base->tree_first <= base->tree_last) {\n+\t\tstruct object_entry *child = objects + deltas[base->tree_first].obj_no;\n+\t\tstruct base_data *result;\n+\n+\t\tassert(child->type == OBJ_PV4_TREE);\n+\t\tif (child->nr_bases > 1) {\n+\t\t\t/* maybe resolved in the third pass or something */\n+\t\t\tbase->tree_first++;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tresult = alloc_base_data();\n+\t\tresolve_delta(child, base, result);\n+\t\tif (base->tree_first == base->tree_last)\n+\t\t\tfree_base_data(base);\n+\n+\t\tbase->tree_first++;\n+\t\treturn result;\n+\t}\n+\n \tunlink_base_data(base);\n \treturn NULL;\n }\n@@ -1266,7 +1427,8 @@ static void *threaded_second_pass(void *data)\n \t\tcounter_unlock();\n \t\twork_lock();\n \t\twhile (nr_dispatched < nr_objects &&\n-\t\t       is_delta_type(objects[nr_dispatched].type))\n+\t\t       (is_delta_type(objects[nr_dispatched].type) ||\n+\t\t\tis_delta_tree(objects + nr_dispatched)))\n \t\t\tnr_dispatched++;\n \t\tif (nr_dispatched >= nr_objects) {\n \t\t\twork_unlock();\n@@ -1411,7 +1573,7 @@ static void resolve_deltas(void)\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n \n-\t\tif (is_delta_type(obj->type))\n+\t\tif (is_delta_type(obj->type) || is_delta_tree(obj))\n \t\t\tcontinue;\n \t\tresolve_base(obj);\n \t\tdisplay_progress(progress, nr_resolved_deltas);\n@@ -1956,7 +2118,7 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \t\tshow_pack_info(stat_only);\n \n \tif (packv4)\n-\t\tdie(\"we're not there yet\");\n+\t\topts.version = 3;\n \n \tidx_objects = xmalloc((nr_objects) * sizeof(struct pack_idx_entry *));\n \tfor (i = 0; i < nr_objects; i++)\n-- \n1.8.2.83.gc99314b\n"},{"id":"227048","messageId":"alpine.LFD.2.03.1309072211050.20709@syhkavp.arg","threadId":"34852","inReplyTo":"1378550599-25365-6-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 05/12] index-pack: parse v4 header and dictionaries","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-08T02:14:29Z","receivedAt":"2013-09-08T02:14:29Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 7 Sep 2013, Nguyễn Thái Ngọc Duy wrote:\n\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n\n[...]\n\n> @@ -1677,6 +1727,15 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n>  \tfree(objects);\n>  \tfree(index_name_buf);\n>  \tfree(keep_name_buf);\n> +\tfree(sha1_table);\n> +\tif (name_dict) {\n> +\t\tfree((void*)name_dict->data);\n> +\t\tfree(name_dict);\n> +\t}\n> +\tif (path_dict) {\n> +\t\tfree((void*)path_dict->data);\n> +\t\tfree(path_dict);\n> +\t}\n\nThe freeing of dictionary tables should probably have its own function \nin packv4-parse.c.  and a call to it added in free_pack_by_name() as \nwell.\n\n\nNicolas\n"},{"id":"227052","messageId":"alpine.LFD.2.03.1309072240230.20709@syhkavp.arg","threadId":"34852","inReplyTo":"1378550599-25365-9-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 08/12] index-pack: parse v4 tree format","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-08T02:52:05Z","receivedAt":"2013-09-08T02:52:05Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 7 Sep 2013, Nguyễn Thái Ngọc Duy wrote:\n\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  builtin/index-pack.c | 104 +++++++++++++++++++++++++++++++++++++++++++++++++--\n>  1 file changed, 100 insertions(+), 4 deletions(-)\n> \n> diff --git a/builtin/index-pack.c b/builtin/index-pack.c\n> index 210b78d..51ca64b 100644\n> --- a/builtin/index-pack.c\n> +++ b/builtin/index-pack.c\n> @@ -319,6 +319,21 @@ static const unsigned char *read_sha1ref(void)\n>  \treturn sha1_table + index * 20;\n>  }\n>  \n> +static const unsigned char *read_sha1table_ref(void)\n> +{\n> +\tconst unsigned char *sha1 = read_sha1ref();\n> +\tif (sha1 < sha1_table || sha1 >= sha1_table + nr_objects * 20) {\n> +\t\tunsigned char *found;\n> +\t\tfound = bsearch(sha1, sha1_table, nr_objects, 20,\n> +\t\t\t\t(int (*)(const void *, const void *))hashcmp);\n> +\t\tif (!found)\n> +\t\t\tbad_object(consumed_bytes,\n> +\t\t\t\t   _(\"SHA-1 %s not found in SHA-1 table\"),\n> +\t\t\t\t   sha1_to_hex(sha1));\n> +\t}\n> +\treturn sha1;\n> +}\n> +\n>  static const unsigned char *read_dictref(struct packv4_dict *dict)\n>  {\n>  \tunsigned int index = read_varint();\n> @@ -561,17 +576,93 @@ static void *unpack_commit_v4(unsigned int offset,\n>  \treturn dst.buf;\n>  }\n>  \n> -static void *unpack_entry_data(unsigned long offset, unsigned long size,\n> -\t\t\t       enum object_type type, unsigned char *sha1)\n> +/*\n> + * v4 trees are actually kind of deltas and we don't do delta in the\n> + * first pass. This function only walks through a tree object to find\n> + * the end offset, register object dependencies and performs limited\n> + * validation.\n> + */\n> +static void *unpack_tree_v4(struct object_entry *obj,\n> +\t\t\t    unsigned int offset, unsigned long size,\n> +\t\t\t    unsigned char *sha1)\n> +{\n> +\tunsigned int nr = read_varint();\n> +\tconst unsigned char *last_base = NULL;\n> +\tstruct strbuf sb = STRBUF_INIT;\n> +\twhile (nr) {\n> +\t\tunsigned int copy_start_or_path = read_varint();\n> +\t\tif (copy_start_or_path & 1) { /* copy_start */\n> +\t\t\tunsigned int copy_count = read_varint();\n> +\t\t\tif (copy_count & 1) { /* first delta */\n> +\t\t\t\tlast_base = read_sha1table_ref();\n> +\t\t\t} else if (!last_base)\n> +\t\t\t\tbad_object(offset,\n> +\t\t\t\t\t   _(\"bad copy count index in unpack_tree_v4\"));\n\nHere the error message could be a little more explicit i.e. \"missing \ndelta base\" or the like in order to distinguish from the next error.\n\n> +\t\t\tcopy_count >>= 1;\n> +\t\t\tif (!copy_count)\n> +\t\t\t\tbad_object(offset,\n> +\t\t\t\t\t   _(\"bad copy count index in unpack_tree_v4\"));\n> +\t\t\tnr -= copy_count;\n\nAlso make sure copy_count <= nr here.\n\n> +\t\t} else {\t/* path */\n> +\t\t\tunsigned int path_idx = copy_start_or_path >> 1;\n> +\t\t\tconst unsigned char *entry_sha1;\n> +\n> +\t\t\tif (path_idx >= path_dict->nb_entries)\n> +\t\t\t\tbad_object(offset,\n> +\t\t\t\t\t   _(\"bad path index in unpack_tree_v4\"));\n> +\t\t\tentry_sha1 = read_sha1ref();\n> +\t\t\tnr--;\n> +\n> +\t\t\tif (!last_base) {\n\nI've been confused for a while here by the use of last_base in the non \ndelta path.  A comment indicating why this used here might be helpful to \nthose unfamiliar with the format.\n\n> +\t\t\t\tconst unsigned char *path;\n> +\t\t\t\tunsigned mode;\n> +\n> +\t\t\t\tpath = path_dict->data + path_dict->offsets[path_idx];\n> +\t\t\t\tmode = (path[0] << 8) | path[1];\n> +\t\t\t\tstrbuf_addf(&sb, \"%o %s%c\", mode, path+2, '\\0');\n> +\t\t\t\tstrbuf_add(&sb, entry_sha1, 20);\n> +\t\t\t\tif (sb.len > size)\n> +\t\t\t\t\tbad_object(offset,\n> +\t\t\t\t\t\t   _(\"tree larger than expected\"));\n> +\t\t\t}\n> +\t\t}\n> +\t}\n> +\n> +\tif (last_base) {\n> +\t\tstrbuf_release(&sb);\n> +\t\treturn NULL;\n> +\t} else {\n> +\t\tgit_SHA_CTX ctx;\n> +\t\tchar hdr[32];\n> +\t\tint hdrlen;\n> +\n> +\t\tif (sb.len != size)\n> +\t\t\tbad_object(offset, _(\"tree size mismatch\"));\n> +\n> +\t\thdrlen = sprintf(hdr, \"tree %lu\", size) + 1;\n> +\t\tgit_SHA1_Init(&ctx);\n> +\t\tgit_SHA1_Update(&ctx, hdr, hdrlen);\n> +\t\tgit_SHA1_Update(&ctx, sb.buf, size);\n> +\t\tgit_SHA1_Final(sha1, &ctx);\n> +\t\treturn strbuf_detach(&sb, NULL);\n> +\t}\n> +}\n> +\n> +static void *unpack_entry_data(struct object_entry *obj, unsigned char *sha1)\n>  {\n>  \tstatic char fixed_buf[8192];\n>  \tvoid *buf;\n>  \tgit_SHA_CTX c;\n>  \tchar hdr[32];\n>  \tint hdrlen;\n> +\tunsigned long offset = obj->idx.offset;\n> +\tunsigned long size = obj->size;\n> +\tenum object_type type = obj->type;\n>  \n>  \tif (type == OBJ_PV4_COMMIT)\n>  \t\treturn unpack_commit_v4(offset, size, sha1);\n> +\tif (type == OBJ_PV4_TREE)\n> +\t\treturn unpack_tree_v4(obj, offset, size, sha1);\n>  \n>  \tif (!is_delta_type(type)) {\n>  \t\thdrlen = sprintf(hdr, \"%s %lu\", typename(type), size) + 1;\n> @@ -640,16 +731,19 @@ static void *unpack_raw_entry(struct object_entry *obj,\n>  \tcase OBJ_BLOB:\n>  \tcase OBJ_TAG:\n>  \t\tbreak;\n> -\n>  \tcase OBJ_PV4_COMMIT:\n>  \t\tobj->real_type = OBJ_COMMIT;\n>  \t\tbreak;\n> +\tcase OBJ_PV4_TREE:\n> +\t\tobj->real_type = OBJ_TREE;\n> +\t\tbreak;\n> +\n>  \tdefault:\n>  \t\tbad_object(obj->idx.offset, _(\"unknown object type %d\"), obj->type);\n>  \t}\n>  \tobj->hdr_size = consumed_bytes - obj->idx.offset;\n>  \n> -\tdata = unpack_entry_data(obj->idx.offset, obj->size, obj->type, sha1);\n> +\tdata = unpack_entry_data(obj, sha1);\n>  \tobj->idx.crc32 = input_crc32;\n>  \treturn data;\n>  }\n> @@ -1186,6 +1280,8 @@ static void parse_pack_objects(unsigned char *sha1)\n>  \t\t\tnr_deltas++;\n>  \t\t\tdelta->obj_no = i;\n>  \t\t\tdelta++;\n> +\t\t} else if (!data && obj->type == OBJ_PV4_TREE) {\n> +\t\t\t/* delay sha1_object() until second pass */\n>  \t\t} else if (!data) {\n>  \t\t\t/* large blobs, check later */\n>  \t\t\tobj->real_type = OBJ_BAD;\n> -- \n> 1.8.2.83.gc99314b\n> \n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n> \n"},{"id":"227056","messageId":"alpine.LFD.2.03.1309072312090.20709@syhkavp.arg","threadId":"34852","inReplyTo":"1378550599-25365-13-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 12/12] index-pack: resolve v4 one-base trees","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-08T03:28:07Z","receivedAt":"2013-09-08T03:28:07Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 7 Sep 2013, Nguyễn Thái Ngọc Duy wrote:\n\n> This is the most common case for delta trees. In fact it's the only\n> kind that's produced by packv4-create. It fits well in the way\n> index-pack resolves deltas and benefits from threading (the set of\n> objects depending on this base does not overlap with the set of\n> objects depending on another base)\n> \n> Multi-base trees will be probably processed differently.\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  builtin/index-pack.c | 194 ++++++++++++++++++++++++++++++++++++++++++++++-----\n>  1 file changed, 178 insertions(+), 16 deletions(-)\n> \n> diff --git a/builtin/index-pack.c b/builtin/index-pack.c\n> index 1fa74f4..4a24bc3 100644\n> --- a/builtin/index-pack.c\n> +++ b/builtin/index-pack.c\n> @@ -12,6 +12,8 @@\n>  #include \"streaming.h\"\n>  #include \"thread-utils.h\"\n>  #include \"packv4-parse.h\"\n> +#include \"varint.h\"\n> +#include \"tree-walk.h\"\n>  \n>  static const char index_pack_usage[] =\n>  \"git index-pack [-v] [-o <index-file>] [--keep | --keep=<msg>] [--verify] [--strict] (<pack-file> | --stdin [--fix-thin] [<pack-file>])\";\n> @@ -38,8 +40,8 @@ struct base_data {\n>  \tstruct object_entry *obj;\n>  \tvoid *data;\n>  \tunsigned long size;\n> -\tint ref_first, ref_last;\n> -\tint ofs_first, ofs_last;\n> +\tint ref_first, ref_last, tree_first;\n> +\tint ofs_first, ofs_last, tree_last;\n>  };\n>  \n>  #if !defined(NO_PTHREADS) && defined(NO_THREAD_SAFE_PREAD)\n> @@ -437,6 +439,7 @@ static struct base_data *alloc_base_data(void)\n>  \tmemset(base, 0, sizeof(*base));\n>  \tbase->ref_last = -1;\n>  \tbase->ofs_last = -1;\n> +\tbase->tree_last = -1;\n>  \treturn base;\n>  }\n>  \n> @@ -670,6 +673,8 @@ static void *unpack_tree_v4(struct object_entry *obj,\n>  \t}\n>  \n>  \tif (last_base) {\n> +\t\tif (nr_deltas - delta_start > 1)\n> +\t\t\tdie(\"sorry guys, multi-base trees are not supported yet\");\n>  \t\tstrbuf_release(&sb);\n>  \t\treturn NULL;\n>  \t} else {\n> @@ -794,6 +799,83 @@ static void *unpack_raw_entry(struct object_entry *obj,\n>  \treturn data;\n>  }\n>  \n> +static void *patch_one_base_tree(const struct object_entry *src,\n> +\t\t\t\t const unsigned char *src_buf,\n> +\t\t\t\t const unsigned char *delta_buf,\n> +\t\t\t\t unsigned long delta_size,\n> +\t\t\t\t unsigned long *dst_size)\n> +{\n> +\tunsigned int nr;\n> +\tconst unsigned char *last_base = NULL;\n> +\tstruct strbuf sb = STRBUF_INIT;\n> +\tconst unsigned char *p = delta_buf;\n> +\n> +\tnr = decode_varint(&p);\n> +\twhile (nr && p < delta_buf + delta_size) {\n> +\t\tunsigned int copy_start_or_path = decode_varint(&p);\n> +\t\tif (copy_start_or_path & 1) { /* copy_start */\n> +\t\t\tstruct tree_desc desc;\n> +\t\t\tstruct name_entry entry;\n> +\t\t\tunsigned int copy_count = decode_varint(&p);\n> +\t\t\tunsigned int copy_start = copy_start_or_path >> 1;\n> +\t\t\tif (!src)\n> +\t\t\t\tdie(\"we are not supposed to copy from another tree!\");\n> +\t\t\tif (copy_count & 1) { /* first delta */\n> +\t\t\t\tunsigned int id = decode_varint(&p);\n> +\t\t\t\tif (!id) {\n> +\t\t\t\t\tlast_base = p;\n> +\t\t\t\t\tp += 20;\n> +\t\t\t\t} else\n> +\t\t\t\t\tlast_base = sha1_table + (id - 1) * 20;\n> +\t\t\t\tif (hashcmp(last_base, src->idx.sha1))\n> +\t\t\t\t\tdie(_(\"bad tree base in patch_one_base_tree\"));\n> +\t\t\t} else if (!last_base)\n> +\t\t\t\tdie(_(\"bad copy count index in patch_one_base_tree\"));\n> +\t\t\tcopy_count >>= 1;\n> +\t\t\tif (!copy_count)\n> +\t\t\t\tdie(_(\"bad copy count index in patch_one_base_tree\"));\n> +\t\t\tnr -= copy_count;\n> +\n> +\t\t\tinit_tree_desc(&desc, src_buf, src->size);\n> +\t\t\twhile (tree_entry(&desc, &entry)) {\n> +\t\t\t\tif (copy_start)\n> +\t\t\t\t\tcopy_start--;\n> +\t\t\t\telse if (copy_count) {\n> +\t\t\t\t\tstrbuf_addf(&sb, \"%o %s%c\", entry.mode, entry.path, '\\0');\n> +\t\t\t\t\tstrbuf_add(&sb, entry.sha1, 20);\n> +\t\t\t\t\tcopy_count--;\n> +\t\t\t\t} else\n> +\t\t\t\t\tbreak;\n> +\t\t\t}\n> +\t\t} else {\t/* path */\n> +\t\t\tunsigned int path_idx = copy_start_or_path >> 1;\n> +\t\t\tconst unsigned char *path;\n> +\t\t\tunsigned mode;\n> +\t\t\tunsigned int id;\n> +\t\t\tconst unsigned char *entry_sha1;\n> +\n> +\t\t\tif (path_idx >= path_dict->nb_entries)\n> +\t\t\t\tdie(_(\"bad path index in unpack_tree_v4\"));\n> +\t\t\tid = decode_varint(&p);\n> +\t\t\tif (!id) {\n> +\t\t\t\tentry_sha1 = p;\n> +\t\t\t\tp += 20;\n> +\t\t\t} else\n> +\t\t\t\tentry_sha1 = sha1_table + (id - 1) * 20;\n\nYou should verify that id doesn't overflow the sha1 table here.\nSimilarly in other places.\n\n\nNicolas\n"},{"id":"227057","messageId":"CACsJy8BnJ-SLz2WUdmfJQt69x5N7VSMD1iXEpFsFfx1Qh=janQ@mail.gmail.com","threadId":"34852","inReplyTo":"alpine.LFD.2.03.1309072312090.20709@syhkavp.arg","subject":"Re: [PATCH 12/12] index-pack: resolve v4 one-base trees","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T03:44:35Z","receivedAt":"2013-09-08T03:44:35Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Sep 8, 2013 at 10:28 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n>> @@ -794,6 +799,83 @@ static void *unpack_raw_entry(struct object_entry *obj,\n>>       return data;\n>>  }\n>>\n>> +static void *patch_one_base_tree(const struct object_entry *src,\n>> +                              const unsigned char *src_buf,\n>> +                              const unsigned char *delta_buf,\n>> +                              unsigned long delta_size,\n>> +                              unsigned long *dst_size)\n>> +{\n>> +     unsigned int nr;\n>> +     const unsigned char *last_base = NULL;\n>> +     struct strbuf sb = STRBUF_INIT;\n>> +     const unsigned char *p = delta_buf;\n>> +\n>> +     nr = decode_varint(&p);\n>> +     while (nr && p < delta_buf + delta_size) {\n>> +             unsigned int copy_start_or_path = decode_varint(&p);\n>> +             if (copy_start_or_path & 1) { /* copy_start */\n>> +                     struct tree_desc desc;\n>> +                     struct name_entry entry;\n>> +                     unsigned int copy_count = decode_varint(&p);\n>> +                     unsigned int copy_start = copy_start_or_path >> 1;\n>> +                     if (!src)\n>> +                             die(\"we are not supposed to copy from another tree!\");\n>> +                     if (copy_count & 1) { /* first delta */\n>> +                             unsigned int id = decode_varint(&p);\n>> +                             if (!id) {\n>> +                                     last_base = p;\n>> +                                     p += 20;\n>> +                             } else\n>> +                                     last_base = sha1_table + (id - 1) * 20;\n>> +                             if (hashcmp(last_base, src->idx.sha1))\n>> +                                     die(_(\"bad tree base in patch_one_base_tree\"));\n>> +                     } else if (!last_base)\n>> +                             die(_(\"bad copy count index in patch_one_base_tree\"));\n>> +                     copy_count >>= 1;\n>> +                     if (!copy_count)\n>> +                             die(_(\"bad copy count index in patch_one_base_tree\"));\n>> +                     nr -= copy_count;\n>> +\n>> +                     init_tree_desc(&desc, src_buf, src->size);\n>> +                     while (tree_entry(&desc, &entry)) {\n>> +                             if (copy_start)\n>> +                                     copy_start--;\n>> +                             else if (copy_count) {\n>> +                                     strbuf_addf(&sb, \"%o %s%c\", entry.mode, entry.path, '\\0');\n>> +                                     strbuf_add(&sb, entry.sha1, 20);\n>> +                                     copy_count--;\n>> +                             } else\n>> +                                     break;\n>> +                     }\n>> +             } else {        /* path */\n>> +                     unsigned int path_idx = copy_start_or_path >> 1;\n>> +                     const unsigned char *path;\n>> +                     unsigned mode;\n>> +                     unsigned int id;\n>> +                     const unsigned char *entry_sha1;\n>> +\n>> +                     if (path_idx >= path_dict->nb_entries)\n>> +                             die(_(\"bad path index in unpack_tree_v4\"));\n>> +                     id = decode_varint(&p);\n>> +                     if (!id) {\n>> +                             entry_sha1 = p;\n>> +                             p += 20;\n>> +                     } else\n>> +                             entry_sha1 = sha1_table + (id - 1) * 20;\n>\n> You should verify that id doesn't overflow the sha1 table here.\n> Similarly in other places.\n\nI think it's unnecessary. All trees must have been checked by\nunpack_tree_v4() in the first pass. Overflow should be caught there if\nfound.\n\n-- \nDuy\n"},{"id":"227084","messageId":"1378624960-8919-1-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378550599-25365-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 00/14] pack v4 support in index-pack","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:26Z","receivedAt":"2013-09-08T07:22:26Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Mostly cleanups after Nico's comments. The diff against v2 is\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 4a24bc3..88340b5 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -22,8 +22,8 @@ struct object_entry {\n \tstruct pack_idx_entry idx;\n \tunsigned long size;\n \tunsigned int hdr_size;\n-\tenum object_type type;\n-\tenum object_type real_type;\n+\tenum object_type type;\t/* type as written in pack */\n+\tenum object_type real_type; /* type after delta resolving */\n \tunsigned delta_depth;\n \tint base_object_no;\n \tint nr_bases;\t\t/* only valid for v4 trees */\n@@ -194,8 +194,10 @@ static int mark_link(struct object *obj, int type, void *data)\n \treturn 0;\n }\n \n-/* The content of each linked object must have been checked\n-   or it must be already present in the object database */\n+/*\n+ * The content of each linked object must have been checked or it must\n+ * be already present in the object database\n+ */\n static unsigned check_object(struct object *obj)\n {\n \tif (!obj)\n@@ -289,6 +291,19 @@ static inline void *fill_and_use(int bytes)\n \treturn p;\n }\n \n+static void check_against_sha1table(const unsigned char *sha1)\n+{\n+\tconst unsigned char *found;\n+\tif (!packv4)\n+\t\treturn;\n+\n+\tfound = bsearch(sha1, sha1_table, nr_objects, 20,\n+\t\t\t(int (*)(const void *, const void *))hashcmp);\n+\tif (!found)\n+\t\tdie(_(\"object %s not found in SHA-1 table\"),\n+\t\t    sha1_to_hex(sha1));\n+}\n+\n static NORETURN void bad_object(unsigned long offset, const char *format,\n \t\t       ...) __attribute__((format (printf, 2, 3)));\n \n@@ -325,15 +340,8 @@ static const unsigned char *read_sha1ref(void)\n static const unsigned char *read_sha1table_ref(void)\n {\n \tconst unsigned char *sha1 = read_sha1ref();\n-\tif (sha1 < sha1_table || sha1 >= sha1_table + nr_objects * 20) {\n-\t\tunsigned char *found;\n-\t\tfound = bsearch(sha1, sha1_table, nr_objects, 20,\n-\t\t\t\t(int (*)(const void *, const void *))hashcmp);\n-\t\tif (!found)\n-\t\t\tbad_object(consumed_bytes,\n-\t\t\t\t   _(\"SHA-1 %s not found in SHA-1 table\"),\n-\t\t\t\t   sha1_to_hex(sha1));\n-\t}\n+\tif (sha1 < sha1_table || sha1 >= sha1_table + nr_objects * 20)\n+\t\tcheck_against_sha1table(sha1);\n \treturn sha1;\n }\n \n@@ -346,21 +354,6 @@ static const unsigned char *read_dictref(struct packv4_dict *dict)\n \treturn  dict->data + dict->offsets[index];\n }\n \n-static void *read_data(int size)\n-{\n-\tconst int max = sizeof(input_buffer);\n-\tvoid *buf;\n-\tchar *p;\n-\tp = buf = xmalloc(size);\n-\twhile (size) {\n-\t\tint to_fill = size > max ? max : size;\n-\t\tmemcpy(p, fill_and_use(to_fill), to_fill);\n-\t\tp += to_fill;\n-\t\tsize -= to_fill;\n-\t}\n-\treturn buf;\n-}\n-\n static const char *open_pack_file(const char *pack_name)\n {\n \tif (from_stdin) {\n@@ -532,8 +525,7 @@ static void read_and_inflate(unsigned long offset,\n \t\tgit_SHA1_Final(sha1, ctx);\n }\n \n-static void *unpack_commit_v4(unsigned int offset,\n-\t\t\t      unsigned long size,\n+static void *unpack_commit_v4(unsigned int offset, unsigned long size,\n \t\t\t      unsigned char *sha1)\n {\n \tunsigned int nb_parents;\n@@ -622,7 +614,8 @@ static void add_tree_delta_base(struct object_entry *obj,\n  * v4 trees are actually kind of deltas and we don't do delta in the\n  * first pass. This function only walks through a tree object to find\n  * the end offset, register object dependencies and performs limited\n- * validation.\n+ * validation. For v4 trees that have no dependencies, we do\n+ * uncompress and calculate their SHA-1.\n  */\n static void *unpack_tree_v4(struct object_entry *obj,\n \t\t\t    unsigned int offset, unsigned long size,\n@@ -641,9 +634,9 @@ static void *unpack_tree_v4(struct object_entry *obj,\n \t\t\t\tadd_tree_delta_base(obj, last_base, delta_start);\n \t\t\t} else if (!last_base)\n \t\t\t\tbad_object(offset,\n-\t\t\t\t\t   _(\"bad copy count index in unpack_tree_v4\"));\n+\t\t\t\t\t   _(\"missing delta base unpack_tree_v4\"));\n \t\t\tcopy_count >>= 1;\n-\t\t\tif (!copy_count)\n+\t\t\tif (!copy_count || copy_count > nr)\n \t\t\t\tbad_object(offset,\n \t\t\t\t\t   _(\"bad copy count index in unpack_tree_v4\"));\n \t\t\tnr -= copy_count;\n@@ -657,6 +650,13 @@ static void *unpack_tree_v4(struct object_entry *obj,\n \t\t\tentry_sha1 = read_sha1ref();\n \t\t\tnr--;\n \n+\t\t\t/*\n+\t\t\t * Attempt to rebuild a canonical (base) tree.\n+\t\t\t * If last_base is set, this tree depends on\n+\t\t\t * another tree, which we have no access at this\n+\t\t\t * stage, so reconstruction must be delayed until\n+\t\t\t * the second pass.\n+\t\t\t */\n \t\t\tif (!last_base) {\n \t\t\t\tconst unsigned char *path;\n \t\t\t\tunsigned mode;\n@@ -694,6 +694,11 @@ static void *unpack_tree_v4(struct object_entry *obj,\n \t}\n }\n \n+/*\n+ * Unpack an entry data in the streamed pack, calculate the object\n+ * SHA-1 if it's not a large blob. Otherwise just try to inflate the\n+ * object to /dev/null to determine the end of the entry in the pack.\n+ */\n static void *unpack_entry_data(struct object_entry *obj, unsigned char *sha1)\n {\n \tstatic char fixed_buf[8192];\n@@ -799,19 +804,23 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \treturn data;\n }\n \n+/*\n+ * Some checks are skipped because they are already done by\n+ * unpack_tree_v4() in the first pass.\n+ */\n static void *patch_one_base_tree(const struct object_entry *src,\n \t\t\t\t const unsigned char *src_buf,\n \t\t\t\t const unsigned char *delta_buf,\n \t\t\t\t unsigned long delta_size,\n \t\t\t\t unsigned long *dst_size)\n {\n-\tunsigned int nr;\n+\tint nr;\n \tconst unsigned char *last_base = NULL;\n \tstruct strbuf sb = STRBUF_INIT;\n \tconst unsigned char *p = delta_buf;\n \n \tnr = decode_varint(&p);\n-\twhile (nr && p < delta_buf + delta_size) {\n+\twhile (nr > 0 && p < delta_buf + delta_size) {\n \t\tunsigned int copy_start_or_path = decode_varint(&p);\n \t\tif (copy_start_or_path & 1) { /* copy_start */\n \t\t\tstruct tree_desc desc;\n@@ -829,11 +838,9 @@ static void *patch_one_base_tree(const struct object_entry *src,\n \t\t\t\t\tlast_base = sha1_table + (id - 1) * 20;\n \t\t\t\tif (hashcmp(last_base, src->idx.sha1))\n \t\t\t\t\tdie(_(\"bad tree base in patch_one_base_tree\"));\n-\t\t\t} else if (!last_base)\n-\t\t\t\tdie(_(\"bad copy count index in patch_one_base_tree\"));\n+\t\t\t}\n+\n \t\t\tcopy_count >>= 1;\n-\t\t\tif (!copy_count)\n-\t\t\t\tdie(_(\"bad copy count index in patch_one_base_tree\"));\n \t\t\tnr -= copy_count;\n \n \t\t\tinit_tree_desc(&desc, src_buf, src->size);\n@@ -841,7 +848,8 @@ static void *patch_one_base_tree(const struct object_entry *src,\n \t\t\t\tif (copy_start)\n \t\t\t\t\tcopy_start--;\n \t\t\t\telse if (copy_count) {\n-\t\t\t\t\tstrbuf_addf(&sb, \"%o %s%c\", entry.mode, entry.path, '\\0');\n+\t\t\t\t\tstrbuf_addf(&sb, \"%o %s%c\",\n+\t\t\t\t\t\t    entry.mode, entry.path, '\\0');\n \t\t\t\t\tstrbuf_add(&sb, entry.sha1, 20);\n \t\t\t\t\tcopy_count--;\n \t\t\t\t} else\n@@ -854,8 +862,6 @@ static void *patch_one_base_tree(const struct object_entry *src,\n \t\t\tunsigned int id;\n \t\t\tconst unsigned char *entry_sha1;\n \n-\t\t\tif (path_idx >= path_dict->nb_entries)\n-\t\t\t\tdie(_(\"bad path index in unpack_tree_v4\"));\n \t\t\tid = decode_varint(&p);\n \t\t\tif (!id) {\n \t\t\t\tentry_sha1 = p;\n@@ -876,6 +882,11 @@ static void *patch_one_base_tree(const struct object_entry *src,\n \treturn sb.buf;\n }\n \n+/*\n+ * Unpack entry data in the second pass when the pack is already\n+ * stored on disk. consume call back is used for large-blob case. Must\n+ * be thread safe.\n+ */\n static void *unpack_data(struct object_entry *obj,\n \t\t\t int (*consume)(const unsigned char *, unsigned long, void *),\n \t\t\t void *cb_data)\n@@ -1079,19 +1090,6 @@ static int check_collison(struct object_entry *entry)\n \treturn 0;\n }\n \n-static void check_against_sha1table(struct object_entry *obj)\n-{\n-\tconst unsigned char *found;\n-\tif (!packv4)\n-\t\treturn;\n-\n-\tfound = bsearch(obj->idx.sha1, sha1_table, nr_objects, 20,\n-\t\t\t(int (*)(const void *, const void *))hashcmp);\n-\tif (!found)\n-\t\tdie(_(\"object %s not found in SHA-1 table\"),\n-\t\t    sha1_to_hex(obj->idx.sha1));\n-}\n-\n static void sha1_object(const void *data, struct object_entry *obj_entry,\n \t\t\tunsigned long size, enum object_type type,\n \t\t\tconst unsigned char *sha1)\n@@ -1288,7 +1286,7 @@ static void resolve_delta(struct object_entry *delta_obj,\n \t\tbad_object(delta_obj->idx.offset, _(\"failed to apply delta\"));\n \thash_sha1_file(result->data, result->size,\n \t\t       typename(delta_obj->real_type), delta_obj->idx.sha1);\n-\tcheck_against_sha1table(delta_obj);\n+\tcheck_against_sha1table(delta_obj->idx.sha1);\n \tsha1_object(result->data, NULL, result->size, delta_obj->real_type,\n \t\t    delta_obj->idx.sha1);\n \tcounter_lock();\n@@ -1296,6 +1294,11 @@ static void resolve_delta(struct object_entry *delta_obj,\n \tcounter_unlock();\n }\n \n+/*\n+ * Given a base object, search for all objects depending on the base,\n+ * try to unpack one of those object. The function will be called\n+ * repeatedly until all objects are unpacked.\n+ */\n static struct base_data *find_unresolved_deltas_1(struct base_data *base,\n \t\t\t\t\t\t  struct base_data *prev_base)\n {\n@@ -1408,6 +1411,10 @@ static int compare_delta_entry(const void *a, const void *b)\n \t\t\t\t   objects[delta_b->obj_no].type);\n }\n \n+/*\n+ * Unpack all objects depending directly or indirectly on the given\n+ * object\n+ */\n static void resolve_base(struct object_entry *obj)\n {\n \tstruct base_data *base_obj = alloc_base_data();\n@@ -1417,6 +1424,7 @@ static void resolve_base(struct object_entry *obj)\n }\n \n #ifndef NO_PTHREADS\n+/* Call resolve_base() in multiple threads */\n static void *threaded_second_pass(void *data)\n {\n \tset_thread_data(data);\n@@ -1460,10 +1468,19 @@ static struct packv4_dict *read_dict(void)\n \n static void parse_dictionaries(void)\n {\n+\tint i;\n \tif (!packv4)\n \t\treturn;\n \n-\tsha1_table = read_data(20 * nr_objects);\n+\tsha1_table = xmalloc(20 * nr_objects);\n+\thashcpy(sha1_table, fill_and_use(20));\n+\tfor (i = 1; i < nr_objects; i++) {\n+\t\tunsigned char *p = sha1_table + i * 20;\n+\t\thashcpy(p, fill_and_use(20));\n+\t\tif (hashcmp(p - 20, p) >= 0)\n+\t\t\tdie(_(\"wrong order in SHA-1 table at entry %d\"), i);\n+\t}\n+\n \tname_dict = read_dict();\n \tpath_dict = read_dict();\n }\n@@ -1492,9 +1509,9 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\t/* large blobs, check later */\n \t\t\tobj->real_type = OBJ_BAD;\n \t\t\tnr_delays++;\n-\t\t\tcheck_against_sha1table(obj);\n+\t\t\tcheck_against_sha1table(obj->idx.sha1);\n \t\t} else {\n-\t\t\tcheck_against_sha1table(obj);\n+\t\t\tcheck_against_sha1table(obj->idx.sha1);\n \t\t\tsha1_object(data, NULL, obj->size, obj->real_type,\n \t\t\t\t    obj->idx.sha1);\n \t\t}\n@@ -2137,14 +2154,8 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \tfree(index_name_buf);\n \tfree(keep_name_buf);\n \tfree(sha1_table);\n-\tif (name_dict) {\n-\t\tfree((void*)name_dict->data);\n-\t\tfree(name_dict);\n-\t}\n-\tif (path_dict) {\n-\t\tfree((void*)path_dict->data);\n-\t\tfree(path_dict);\n-\t}\n+\tpv4_free_dict(name_dict);\n+\tpv4_free_dict(path_dict);\n \tif (pack_name == NULL)\n \t\tfree((void *) curr_pack);\n \tif (index_name == NULL)\ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex 82661ba..d515bb9 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -63,6 +63,14 @@ struct packv4_dict *pv4_create_dict(const unsigned char *data, int dict_size)\n \treturn dict;\n }\n \n+void pv4_free_dict(struct packv4_dict *dict)\n+{\n+\tif (dict) {\n+\t\tfree((void*)dict->data);\n+\t\tfree(dict);\n+\t}\n+}\n+\n static struct packv4_dict *load_dict(struct packed_git *p, off_t *offset)\n {\n \tstruct pack_window *w_curs = NULL;\ndiff --git a/packv4-parse.h b/packv4-parse.h\nindex 0b2405a..e6719f6 100644\n--- a/packv4-parse.h\n+++ b/packv4-parse.h\n@@ -8,6 +8,7 @@ struct packv4_dict {\n };\n \n struct packv4_dict *pv4_create_dict(const unsigned char *data, int dict_size);\n+void pv4_free_dict(struct packv4_dict *dict);\n \n void *pv4_get_commit(struct packed_git *p, struct pack_window **w_curs,\n \t\t     off_t offset, unsigned long size);\ndiff --git a/sha1_file.c b/sha1_file.c\nindex c7bf677..1528e28 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -763,6 +763,8 @@ void free_pack_by_name(const char *pack_name)\n \t\t\t}\n \t\t\tclose_pack_index(p);\n \t\t\tfree(p->bad_object_sha1);\n+\t\t\tpv4_free_dict(p->ident_dict);\n+\t\t\tpv4_free_dict(p->path_dict);\n \t\t\t*pp = p->next;\n \t\t\tif (last_found_pack == p)\n \t\t\t\tlast_found_pack = NULL;\n"},{"id":"227085","messageId":"1378624960-8919-2-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 01/14] pack v4: split pv4_create_dict() out of load_dict()","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:27Z","receivedAt":"2013-09-08T07:22:27Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"---\n packv4-parse.c | 63 ++++++++++++++++++++++++++++++++--------------------------\n packv4-parse.h |  8 ++++++++\n 2 files changed, 43 insertions(+), 28 deletions(-)\n\ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex 63bba03..82661ba 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -30,11 +30,38 @@ const unsigned char *get_sha1ref(struct packed_git *p,\n \treturn sha1;\n }\n \n-struct packv4_dict {\n-\tconst unsigned char *data;\n-\tunsigned int nb_entries;\n-\tunsigned int offsets[FLEX_ARRAY];\n-};\n+struct packv4_dict *pv4_create_dict(const unsigned char *data, int dict_size)\n+{\n+\tstruct packv4_dict *dict;\n+\tint i;\n+\n+\t/* count number of entries */\n+\tint nb_entries = 0;\n+\tconst unsigned char *cp = data;\n+\twhile (cp < data + dict_size - 3) {\n+\t\tcp += 2;  /* prefix bytes */\n+\t\tcp += strlen((const char *)cp);  /* entry string */\n+\t\tcp += 1;  /* terminating NUL */\n+\t\tnb_entries++;\n+\t}\n+\tif (cp - data != dict_size) {\n+\t\terror(\"dict size mismatch\");\n+\t\treturn NULL;\n+\t}\n+\n+\tdict = xmalloc(sizeof(*dict) + nb_entries * sizeof(dict->offsets[0]));\n+\tdict->data = data;\n+\tdict->nb_entries = nb_entries;\n+\n+\tcp = data;\n+\tfor (i = 0; i < nb_entries; i++) {\n+\t\tdict->offsets[i] = cp - data;\n+\t\tcp += 2;\n+\t\tcp += strlen((const char *)cp) + 1;\n+\t}\n+\n+\treturn dict;\n+}\n \n static struct packv4_dict *load_dict(struct packed_git *p, off_t *offset)\n {\n@@ -45,7 +72,7 @@ static struct packv4_dict *load_dict(struct packed_git *p, off_t *offset)\n \tconst unsigned char *cp;\n \tgit_zstream stream;\n \tstruct packv4_dict *dict;\n-\tint nb_entries, i, st;\n+\tint st;\n \n \t/* get uncompressed dictionary data size */\n \tsrc = use_pack(p, &w_curs, curpos, &avail);\n@@ -77,32 +104,12 @@ static struct packv4_dict *load_dict(struct packed_git *p, off_t *offset)\n \t\treturn NULL;\n \t}\n \n-\t/* count number of entries */\n-\tnb_entries = 0;\n-\tcp = data;\n-\twhile (cp < data + dict_size - 3) {\n-\t\tcp += 2;  /* prefix bytes */\n-\t\tcp += strlen((const char *)cp);  /* entry string */\n-\t\tcp += 1;  /* terminating NUL */\n-\t\tnb_entries++;\n-\t}\n-\tif (cp - data != dict_size) {\n-\t\terror(\"dict size mismatch\");\n+\tdict = pv4_create_dict(data, dict_size);\n+\tif (!dict) {\n \t\tfree(data);\n \t\treturn NULL;\n \t}\n \n-\tdict = xmalloc(sizeof(*dict) + nb_entries * sizeof(dict->offsets[0]));\n-\tdict->data = data;\n-\tdict->nb_entries = nb_entries;\n-\n-\tcp = data;\n-\tfor (i = 0; i < nb_entries; i++) {\n-\t\tdict->offsets[i] = cp - data;\n-\t\tcp += 2;\n-\t\tcp += strlen((const char *)cp) + 1;\n-\t}\n-\n \t*offset = curpos;\n \treturn dict;\n }\ndiff --git a/packv4-parse.h b/packv4-parse.h\nindex 5f9d809..0b2405a 100644\n--- a/packv4-parse.h\n+++ b/packv4-parse.h\n@@ -1,6 +1,14 @@\n #ifndef PACKV4_PARSE_H\n #define PACKV4_PARSE_H\n \n+struct packv4_dict {\n+\tconst unsigned char *data;\n+\tunsigned int nb_entries;\n+\tunsigned int offsets[FLEX_ARRAY];\n+};\n+\n+struct packv4_dict *pv4_create_dict(const unsigned char *data, int dict_size);\n+\n void *pv4_get_commit(struct packed_git *p, struct pack_window **w_curs,\n \t\t     off_t offset, unsigned long size);\n void *pv4_get_tree(struct packed_git *p, struct pack_window **w_curs,\n-- \n1.8.2.83.gc99314b\n"},{"id":"227086","messageId":"1378624960-8919-3-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 02/14] pack v4: add pv4_free_dict()","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:28Z","receivedAt":"2013-09-08T07:22:28Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"---\n packv4-parse.c | 8 ++++++++\n packv4-parse.h | 1 +\n sha1_file.c    | 2 ++\n 3 files changed, 11 insertions(+)\n\ndiff --git a/packv4-parse.c b/packv4-parse.c\nindex 82661ba..d515bb9 100644\n--- a/packv4-parse.c\n+++ b/packv4-parse.c\n@@ -63,6 +63,14 @@ struct packv4_dict *pv4_create_dict(const unsigned char *data, int dict_size)\n \treturn dict;\n }\n \n+void pv4_free_dict(struct packv4_dict *dict)\n+{\n+\tif (dict) {\n+\t\tfree((void*)dict->data);\n+\t\tfree(dict);\n+\t}\n+}\n+\n static struct packv4_dict *load_dict(struct packed_git *p, off_t *offset)\n {\n \tstruct pack_window *w_curs = NULL;\ndiff --git a/packv4-parse.h b/packv4-parse.h\nindex 0b2405a..e6719f6 100644\n--- a/packv4-parse.h\n+++ b/packv4-parse.h\n@@ -8,6 +8,7 @@ struct packv4_dict {\n };\n \n struct packv4_dict *pv4_create_dict(const unsigned char *data, int dict_size);\n+void pv4_free_dict(struct packv4_dict *dict);\n \n void *pv4_get_commit(struct packed_git *p, struct pack_window **w_curs,\n \t\t     off_t offset, unsigned long size);\ndiff --git a/sha1_file.c b/sha1_file.c\nindex c7bf677..1528e28 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -763,6 +763,8 @@ void free_pack_by_name(const char *pack_name)\n \t\t\t}\n \t\t\tclose_pack_index(p);\n \t\t\tfree(p->bad_object_sha1);\n+\t\t\tpv4_free_dict(p->ident_dict);\n+\t\t\tpv4_free_dict(p->path_dict);\n \t\t\t*pp = p->next;\n \t\t\tif (last_found_pack == p)\n \t\t\t\tlast_found_pack = NULL;\n-- \n1.8.2.83.gc99314b\n"},{"id":"227087","messageId":"1378624960-8919-4-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 03/14] index-pack: add more comments on some big functions","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:29Z","receivedAt":"2013-09-08T07:22:29Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"---\n builtin/index-pack.c | 30 ++++++++++++++++++++++++++----\n 1 file changed, 26 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 9c1cfac..1dbabe0 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -19,8 +19,8 @@ struct object_entry {\n \tstruct pack_idx_entry idx;\n \tunsigned long size;\n \tunsigned int hdr_size;\n-\tenum object_type type;\n-\tenum object_type real_type;\n+\tenum object_type type;\t/* type as written in pack */\n+\tenum object_type real_type; /* type after delta resolving */\n \tunsigned delta_depth;\n \tint base_object_no;\n };\n@@ -187,8 +187,10 @@ static int mark_link(struct object *obj, int type, void *data)\n \treturn 0;\n }\n \n-/* The content of each linked object must have been checked\n-   or it must be already present in the object database */\n+/*\n+ * The content of each linked object must have been checked or it must\n+ * be already present in the object database\n+ */\n static unsigned check_object(struct object *obj)\n {\n \tif (!obj)\n@@ -407,6 +409,11 @@ static int is_delta_type(enum object_type type)\n \treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n }\n \n+/*\n+ * Unpack an entry data in the streamed pack, calculate the object\n+ * SHA-1 if it's not a large blob. Otherwise just try to inflate the\n+ * object to /dev/null to determine the end of the entry in the pack.\n+ */\n static void *unpack_entry_data(unsigned long offset, unsigned long size,\n \t\t\t       enum object_type type, unsigned char *sha1)\n {\n@@ -522,6 +529,11 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \treturn data;\n }\n \n+/*\n+ * Unpack entry data in the second pass when the pack is already\n+ * stored on disk. consume call back is used for large-blob case. Must\n+ * be thread safe.\n+ */\n static void *unpack_data(struct object_entry *obj,\n \t\t\t int (*consume)(const unsigned char *, unsigned long, void *),\n \t\t\t void *cb_data)\n@@ -875,6 +887,11 @@ static void resolve_delta(struct object_entry *delta_obj,\n \tcounter_unlock();\n }\n \n+/*\n+ * Given a base object, search for all objects depending on the base,\n+ * try to unpack one of those object. The function will be called\n+ * repeatedly until all objects are unpacked.\n+ */\n static struct base_data *find_unresolved_deltas_1(struct base_data *base,\n \t\t\t\t\t\t  struct base_data *prev_base)\n {\n@@ -958,6 +975,10 @@ static int compare_delta_entry(const void *a, const void *b)\n \t\t\t\t   objects[delta_b->obj_no].type);\n }\n \n+/*\n+ * Unpack all objects depending directly or indirectly on the given\n+ * object\n+ */\n static void resolve_base(struct object_entry *obj)\n {\n \tstruct base_data *base_obj = alloc_base_data();\n@@ -967,6 +988,7 @@ static void resolve_base(struct object_entry *obj)\n }\n \n #ifndef NO_PTHREADS\n+/* Call resolve_base() in multiple threads */\n static void *threaded_second_pass(void *data)\n {\n \tset_thread_data(data);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227088","messageId":"1378624960-8919-5-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 04/14] index-pack: split out varint decoding code","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:30Z","receivedAt":"2013-09-08T07:22:30Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"---\n builtin/index-pack.c | 82 ++++++++++++++++++++++++++++------------------------\n 1 file changed, 45 insertions(+), 37 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 1dbabe0..5fbd517 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -277,6 +277,31 @@ static void use(int bytes)\n \tconsumed_bytes += bytes;\n }\n \n+static inline void *fill_and_use(int bytes)\n+{\n+\tvoid *p = fill(bytes);\n+\tuse(bytes);\n+\treturn p;\n+}\n+\n+static NORETURN void bad_object(unsigned long offset, const char *format,\n+\t\t       ...) __attribute__((format (printf, 2, 3)));\n+\n+static uintmax_t read_varint(void)\n+{\n+\tunsigned char c = *(char*)fill_and_use(1);\n+\tuintmax_t val = c & 127;\n+\twhile (c & 128) {\n+\t\tval += 1;\n+\t\tif (!val || MSB(val, 7))\n+\t\t\tbad_object(consumed_bytes,\n+\t\t\t\t   _(\"offset overflow in read_varint\"));\n+\t\tc = *(char*)fill_and_use(1);\n+\t\tval = (val << 7) + (c & 127);\n+\t}\n+\treturn val;\n+}\n+\n static const char *open_pack_file(const char *pack_name)\n {\n \tif (from_stdin) {\n@@ -317,9 +342,6 @@ static void parse_pack_header(void)\n \tuse(sizeof(struct pack_header));\n }\n \n-static NORETURN void bad_object(unsigned long offset, const char *format,\n-\t\t       ...) __attribute__((format (printf, 2, 3)));\n-\n static NORETURN void bad_object(unsigned long offset, const char *format, ...)\n {\n \tva_list params;\n@@ -462,55 +484,41 @@ static void *unpack_entry_data(unsigned long offset, unsigned long size,\n \treturn buf == fixed_buf ? NULL : buf;\n }\n \n+static void read_typesize_v2(struct object_entry *obj)\n+{\n+\tunsigned char c = *(char*)fill_and_use(1);\n+\tunsigned shift;\n+\n+\tobj->type = (c >> 4) & 7;\n+\tobj->size = (c & 15);\n+\tshift = 4;\n+\twhile (c & 128) {\n+\t\tc = *(char*)fill_and_use(1);\n+\t\tobj->size += (c & 0x7f) << shift;\n+\t\tshift += 7;\n+\t}\n+}\n+\n static void *unpack_raw_entry(struct object_entry *obj,\n \t\t\t      union delta_base *delta_base,\n \t\t\t      unsigned char *sha1)\n {\n-\tunsigned char *p;\n-\tunsigned long size, c;\n-\toff_t base_offset;\n-\tunsigned shift;\n \tvoid *data;\n+\tuintmax_t val;\n \n \tobj->idx.offset = consumed_bytes;\n \tinput_crc32 = crc32(0, NULL, 0);\n \n-\tp = fill(1);\n-\tc = *p;\n-\tuse(1);\n-\tobj->type = (c >> 4) & 7;\n-\tsize = (c & 15);\n-\tshift = 4;\n-\twhile (c & 0x80) {\n-\t\tp = fill(1);\n-\t\tc = *p;\n-\t\tuse(1);\n-\t\tsize += (c & 0x7f) << shift;\n-\t\tshift += 7;\n-\t}\n-\tobj->size = size;\n+\tread_typesize_v2(obj);\n \n \tswitch (obj->type) {\n \tcase OBJ_REF_DELTA:\n-\t\thashcpy(delta_base->sha1, fill(20));\n-\t\tuse(20);\n+\t\thashcpy(delta_base->sha1, fill_and_use(20));\n \t\tbreak;\n \tcase OBJ_OFS_DELTA:\n \t\tmemset(delta_base, 0, sizeof(*delta_base));\n-\t\tp = fill(1);\n-\t\tc = *p;\n-\t\tuse(1);\n-\t\tbase_offset = c & 127;\n-\t\twhile (c & 128) {\n-\t\t\tbase_offset += 1;\n-\t\t\tif (!base_offset || MSB(base_offset, 7))\n-\t\t\t\tbad_object(obj->idx.offset, _(\"offset value overflow for delta base object\"));\n-\t\t\tp = fill(1);\n-\t\t\tc = *p;\n-\t\t\tuse(1);\n-\t\t\tbase_offset = (base_offset << 7) + (c & 127);\n-\t\t}\n-\t\tdelta_base->offset = obj->idx.offset - base_offset;\n+\t\tval = read_varint();\n+\t\tdelta_base->offset = obj->idx.offset - val;\n \t\tif (delta_base->offset <= 0 || delta_base->offset >= obj->idx.offset)\n \t\t\tbad_object(obj->idx.offset, _(\"delta base offset is out of bound\"));\n \t\tbreak;\n-- \n1.8.2.83.gc99314b\n"},{"id":"227089","messageId":"1378624960-8919-6-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 05/14] index-pack: do not allocate buffer for unpacking deltas in the first pass","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:31Z","receivedAt":"2013-09-08T07:22:31Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"We do need deltas until the second pass. Allocating a buffer for it\nthen freeing later is wasteful is unnecessary. Make it use fixed_buf\n(aka large blob code path).\n---\n builtin/index-pack.c | 3 ++-\n 1 file changed, 2 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 5fbd517..78554d0 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -453,7 +453,8 @@ static void *unpack_entry_data(unsigned long offset, unsigned long size,\n \t\tgit_SHA1_Update(&c, hdr, hdrlen);\n \t} else\n \t\tsha1 = NULL;\n-\tif (type == OBJ_BLOB && size > big_file_threshold)\n+\tif (is_delta_type(type) ||\n+\t     (type == OBJ_BLOB && size > big_file_threshold))\n \t\tbuf = fixed_buf;\n \telse\n \t\tbuf = xmalloc(size);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227090","messageId":"1378624960-8919-7-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 06/14] index-pack: split inflate/digest code out of unpack_entry_data","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:32Z","receivedAt":"2013-09-08T07:22:32Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"---\n builtin/index-pack.c | 62 +++++++++++++++++++++++++++++++---------------------\n 1 file changed, 37 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 78554d0..3389262 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -431,6 +431,40 @@ static int is_delta_type(enum object_type type)\n \treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n }\n \n+static void read_and_inflate(unsigned long offset,\n+\t\t\t     void *buf, unsigned long size,\n+\t\t\t     unsigned long wraparound,\n+\t\t\t     git_SHA_CTX *ctx,\n+\t\t\t     unsigned char *sha1)\n+{\n+\tgit_zstream stream;\n+\tint status;\n+\n+\tmemset(&stream, 0, sizeof(stream));\n+\tgit_inflate_init(&stream);\n+\tstream.next_out = buf;\n+\tstream.avail_out = wraparound ? wraparound : size;\n+\n+\tdo {\n+\t\tunsigned char *last_out = stream.next_out;\n+\t\tstream.next_in = fill(1);\n+\t\tstream.avail_in = input_len;\n+\t\tstatus = git_inflate(&stream, 0);\n+\t\tuse(input_len - stream.avail_in);\n+\t\tif (sha1)\n+\t\t\tgit_SHA1_Update(ctx, last_out, stream.next_out - last_out);\n+\t\tif (wraparound) {\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = wraparound;\n+\t\t}\n+\t} while (status == Z_OK);\n+\tif (stream.total_out != size || status != Z_STREAM_END)\n+\t\tbad_object(offset, _(\"inflate returned %d\"), status);\n+\tgit_inflate_end(&stream);\n+\tif (sha1)\n+\t\tgit_SHA1_Final(sha1, ctx);\n+}\n+\n /*\n  * Unpack an entry data in the streamed pack, calculate the object\n  * SHA-1 if it's not a large blob. Otherwise just try to inflate the\n@@ -440,8 +474,6 @@ static void *unpack_entry_data(unsigned long offset, unsigned long size,\n \t\t\t       enum object_type type, unsigned char *sha1)\n {\n \tstatic char fixed_buf[8192];\n-\tint status;\n-\tgit_zstream stream;\n \tvoid *buf;\n \tgit_SHA_CTX c;\n \tchar hdr[32];\n@@ -459,29 +491,9 @@ static void *unpack_entry_data(unsigned long offset, unsigned long size,\n \telse\n \t\tbuf = xmalloc(size);\n \n-\tmemset(&stream, 0, sizeof(stream));\n-\tgit_inflate_init(&stream);\n-\tstream.next_out = buf;\n-\tstream.avail_out = buf == fixed_buf ? sizeof(fixed_buf) : size;\n-\n-\tdo {\n-\t\tunsigned char *last_out = stream.next_out;\n-\t\tstream.next_in = fill(1);\n-\t\tstream.avail_in = input_len;\n-\t\tstatus = git_inflate(&stream, 0);\n-\t\tuse(input_len - stream.avail_in);\n-\t\tif (sha1)\n-\t\t\tgit_SHA1_Update(&c, last_out, stream.next_out - last_out);\n-\t\tif (buf == fixed_buf) {\n-\t\t\tstream.next_out = buf;\n-\t\t\tstream.avail_out = sizeof(fixed_buf);\n-\t\t}\n-\t} while (status == Z_OK);\n-\tif (stream.total_out != size || status != Z_STREAM_END)\n-\t\tbad_object(offset, _(\"inflate returned %d\"), status);\n-\tgit_inflate_end(&stream);\n-\tif (sha1)\n-\t\tgit_SHA1_Final(sha1, &c);\n+\tread_and_inflate(offset, buf, size,\n+\t\t\t buf == fixed_buf ? sizeof(fixed_buf) : 0,\n+\t\t\t &c, sha1);\n \treturn buf == fixed_buf ? NULL : buf;\n }\n \n-- \n1.8.2.83.gc99314b\n"},{"id":"227091","messageId":"1378624960-8919-8-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 07/14] index-pack: parse v4 header and dictionaries","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:33Z","receivedAt":"2013-09-08T07:22:33Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"---\n builtin/index-pack.c | 49 ++++++++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 48 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 3389262..83e6e79 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -11,6 +11,7 @@\n #include \"exec_cmd.h\"\n #include \"streaming.h\"\n #include \"thread-utils.h\"\n+#include \"packv4-parse.h\"\n \n static const char index_pack_usage[] =\n \"git index-pack [-v] [-o <index-file>] [--keep | --keep=<msg>] [--verify] [--strict] (<pack-file> | --stdin [--fix-thin] [<pack-file>])\";\n@@ -70,6 +71,8 @@ struct delta_entry {\n static struct object_entry *objects;\n static struct delta_entry *deltas;\n static struct thread_local nothread_data;\n+static unsigned char *sha1_table;\n+static struct packv4_dict *name_dict, *path_dict;\n static int nr_objects;\n static int nr_deltas;\n static int nr_resolved_deltas;\n@@ -81,6 +84,7 @@ static int do_fsck_object;\n static int verbose;\n static int show_stat;\n static int check_self_contained_and_connected;\n+static int packv4;\n \n static struct progress *progress;\n \n@@ -334,7 +338,9 @@ static void parse_pack_header(void)\n \t/* Header consistency check */\n \tif (hdr->hdr_signature != htonl(PACK_SIGNATURE))\n \t\tdie(_(\"pack signature mismatch\"));\n-\tif (!pack_version_ok(hdr->hdr_version))\n+\tif (hdr->hdr_version == htonl(4))\n+\t\tpackv4 = 1;\n+\telse if (!pack_version_ok(hdr->hdr_version))\n \t\tdie(_(\"pack version %\"PRIu32\" unsupported\"),\n \t\t\tntohl(hdr->hdr_version));\n \n@@ -1035,6 +1041,40 @@ static void *threaded_second_pass(void *data)\n }\n #endif\n \n+static struct packv4_dict *read_dict(void)\n+{\n+\tunsigned long size;\n+\tunsigned char *data;\n+\tstruct packv4_dict *dict;\n+\n+\tsize = read_varint();\n+\tdata = xmallocz(size);\n+\tread_and_inflate(consumed_bytes, data, size, 0, NULL, NULL);\n+\tdict = pv4_create_dict(data, size);\n+\tif (!dict)\n+\t\tdie(\"unable to parse dictionary\");\n+\treturn dict;\n+}\n+\n+static void parse_dictionaries(void)\n+{\n+\tint i;\n+\tif (!packv4)\n+\t\treturn;\n+\n+\tsha1_table = xmalloc(20 * nr_objects);\n+\thashcpy(sha1_table, fill_and_use(20));\n+\tfor (i = 1; i < nr_objects; i++) {\n+\t\tunsigned char *p = sha1_table + i * 20;\n+\t\thashcpy(p, fill_and_use(20));\n+\t\tif (hashcmp(p - 20, p) >= 0)\n+\t\t\tdie(_(\"wrong order in SHA-1 table at entry %d\"), i);\n+\t}\n+\n+\tname_dict = read_dict();\n+\tpath_dict = read_dict();\n+}\n+\n /*\n  * First pass:\n  * - find locations of all objects;\n@@ -1673,6 +1713,7 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \tparse_pack_header();\n \tobjects = xcalloc(nr_objects + 1, sizeof(struct object_entry));\n \tdeltas = xcalloc(nr_objects, sizeof(struct delta_entry));\n+\tparse_dictionaries();\n \tparse_pack_objects(pack_sha1);\n \tresolve_deltas();\n \tconclude_pack(fix_thin_pack, curr_pack, pack_sha1);\n@@ -1683,6 +1724,9 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \tif (show_stat)\n \t\tshow_pack_info(stat_only);\n \n+\tif (packv4)\n+\t\tdie(\"we're not there yet\");\n+\n \tidx_objects = xmalloc((nr_objects) * sizeof(struct pack_idx_entry *));\n \tfor (i = 0; i < nr_objects; i++)\n \t\tidx_objects[i] = &objects[i].idx;\n@@ -1699,6 +1743,9 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \tfree(objects);\n \tfree(index_name_buf);\n \tfree(keep_name_buf);\n+\tfree(sha1_table);\n+\tpv4_free_dict(name_dict);\n+\tpv4_free_dict(path_dict);\n \tif (pack_name == NULL)\n \t\tfree((void *) curr_pack);\n \tif (index_name == NULL)\n-- \n1.8.2.83.gc99314b\n"},{"id":"227092","messageId":"1378624960-8919-9-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 08/14] index-pack: make sure all objects are registered in v4's SHA-1 table","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:34Z","receivedAt":"2013-09-08T07:22:34Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"---\n builtin/index-pack.c | 22 ++++++++++++++++++++--\n 1 file changed, 20 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 83e6e79..efb969a 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -288,6 +288,19 @@ static inline void *fill_and_use(int bytes)\n \treturn p;\n }\n \n+static void check_against_sha1table(const unsigned char *sha1)\n+{\n+\tconst unsigned char *found;\n+\tif (!packv4)\n+\t\treturn;\n+\n+\tfound = bsearch(sha1, sha1_table, nr_objects, 20,\n+\t\t\t(int (*)(const void *, const void *))hashcmp);\n+\tif (!found)\n+\t\tdie(_(\"object %s not found in SHA-1 table\"),\n+\t\t    sha1_to_hex(sha1));\n+}\n+\n static NORETURN void bad_object(unsigned long offset, const char *format,\n \t\t       ...) __attribute__((format (printf, 2, 3)));\n \n@@ -907,6 +920,7 @@ static void resolve_delta(struct object_entry *delta_obj,\n \t\tbad_object(delta_obj->idx.offset, _(\"failed to apply delta\"));\n \thash_sha1_file(result->data, result->size,\n \t\t       typename(delta_obj->real_type), delta_obj->idx.sha1);\n+\tcheck_against_sha1table(delta_obj->idx.sha1);\n \tsha1_object(result->data, NULL, result->size, delta_obj->real_type,\n \t\t    delta_obj->idx.sha1);\n \tcounter_lock();\n@@ -1103,8 +1117,12 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\t/* large blobs, check later */\n \t\t\tobj->real_type = OBJ_BAD;\n \t\t\tnr_delays++;\n-\t\t} else\n-\t\t\tsha1_object(data, NULL, obj->size, obj->type, obj->idx.sha1);\n+\t\t\tcheck_against_sha1table(obj->idx.sha1);\n+\t\t} else {\n+\t\t\tcheck_against_sha1table(obj->idx.sha1);\n+\t\t\tsha1_object(data, NULL, obj->size, obj->type,\n+\t\t\t\t    obj->idx.sha1);\n+\t\t}\n \t\tfree(data);\n \t\tdisplay_progress(progress, i+1);\n \t}\n-- \n1.8.2.83.gc99314b\n"},{"id":"227093","messageId":"1378624960-8919-10-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 09/14] index-pack: parse v4 commit format","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:35Z","receivedAt":"2013-09-08T07:22:35Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"---\n builtin/index-pack.c | 94 ++++++++++++++++++++++++++++++++++++++++++++++++++--\n 1 file changed, 91 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex efb969a..473514a 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -319,6 +319,30 @@ static uintmax_t read_varint(void)\n \treturn val;\n }\n \n+static const unsigned char *read_sha1ref(void)\n+{\n+\tunsigned int index = read_varint();\n+\tif (!index) {\n+\t\tstatic unsigned char sha1[20];\n+\t\thashcpy(sha1, fill_and_use(20));\n+\t\treturn sha1;\n+\t}\n+\tindex--;\n+\tif (index >= nr_objects)\n+\t\tbad_object(consumed_bytes,\n+\t\t\t   _(\"bad index in read_sha1ref\"));\n+\treturn sha1_table + index * 20;\n+}\n+\n+static const unsigned char *read_dictref(struct packv4_dict *dict)\n+{\n+\tunsigned int index = read_varint();\n+\tif (index >= dict->nb_entries)\n+\t\tbad_object(consumed_bytes,\n+\t\t\t   _(\"bad index in read_dictref\"));\n+\treturn  dict->data + dict->offsets[index];\n+}\n+\n static const char *open_pack_file(const char *pack_name)\n {\n \tif (from_stdin) {\n@@ -484,6 +508,58 @@ static void read_and_inflate(unsigned long offset,\n \t\tgit_SHA1_Final(sha1, ctx);\n }\n \n+static void *unpack_commit_v4(unsigned int offset, unsigned long size,\n+\t\t\t      unsigned char *sha1)\n+{\n+\tunsigned int nb_parents;\n+\tconst unsigned char *committer, *author, *ident;\n+\tunsigned long author_time, committer_time;\n+\tgit_SHA_CTX ctx;\n+\tchar hdr[32];\n+\tint hdrlen;\n+\tint16_t committer_tz, author_tz;\n+\tstruct strbuf dst;\n+\n+\tstrbuf_init(&dst, size);\n+\n+\tstrbuf_addf(&dst, \"tree %s\\n\", sha1_to_hex(read_sha1ref()));\n+\tnb_parents = read_varint();\n+\twhile (nb_parents--)\n+\t\tstrbuf_addf(&dst, \"parent %s\\n\", sha1_to_hex(read_sha1ref()));\n+\n+\tcommitter_time = read_varint();\n+\tident = read_dictref(name_dict);\n+\tcommitter_tz = (ident[0] << 8) | ident[1];\n+\tcommitter = ident + 2;\n+\n+\tauthor_time = read_varint();\n+\tident = read_dictref(name_dict);\n+\tauthor_tz = (ident[0] << 8) | ident[1];\n+\tauthor = ident + 2;\n+\n+\tif (author_time & 1)\n+\t\tauthor_time = committer_time + (author_time >> 1);\n+\telse\n+\t\tauthor_time = committer_time - (author_time >> 1);\n+\n+\tstrbuf_addf(&dst,\n+\t\t    \"author %s %lu %+05d\\n\"\n+\t\t    \"committer %s %lu %+05d\\n\",\n+\t\t    author, author_time, author_tz,\n+\t\t    committer, committer_time, committer_tz);\n+\n+\tif (dst.len > size)\n+\t\tbad_object(offset, _(\"bad commit\"));\n+\n+\thdrlen = sprintf(hdr, \"commit %lu\", size) + 1;\n+\tgit_SHA1_Init(&ctx);\n+\tgit_SHA1_Update(&ctx, hdr, hdrlen);\n+\tgit_SHA1_Update(&ctx, dst.buf, dst.len);\n+\tread_and_inflate(offset, dst.buf + dst.len, size - dst.len,\n+\t\t\t 0, &ctx, sha1);\n+\treturn dst.buf;\n+}\n+\n /*\n  * Unpack an entry data in the streamed pack, calculate the object\n  * SHA-1 if it's not a large blob. Otherwise just try to inflate the\n@@ -498,6 +574,9 @@ static void *unpack_entry_data(unsigned long offset, unsigned long size,\n \tchar hdr[32];\n \tint hdrlen;\n \n+\tif (type == OBJ_PV4_COMMIT)\n+\t\treturn unpack_commit_v4(offset, size, sha1);\n+\n \tif (!is_delta_type(type)) {\n \t\thdrlen = sprintf(hdr, \"%s %lu\", typename(type), size) + 1;\n \t\tgit_SHA1_Init(&c);\n@@ -541,7 +620,13 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \tobj->idx.offset = consumed_bytes;\n \tinput_crc32 = crc32(0, NULL, 0);\n \n-\tread_typesize_v2(obj);\n+\tif (packv4) {\n+\t\tval = read_varint();\n+\t\tobj->type = val & 15;\n+\t\tobj->size = val >> 4;\n+\t} else\n+\t\tread_typesize_v2(obj);\n+\tobj->real_type = obj->type;\n \n \tswitch (obj->type) {\n \tcase OBJ_REF_DELTA:\n@@ -559,6 +644,10 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n \t\tbreak;\n+\n+\tcase OBJ_PV4_COMMIT:\n+\t\tobj->real_type = OBJ_COMMIT;\n+\t\tbreak;\n \tdefault:\n \t\tbad_object(obj->idx.offset, _(\"unknown object type %d\"), obj->type);\n \t}\n@@ -1108,7 +1197,6 @@ static void parse_pack_objects(unsigned char *sha1)\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n \t\tvoid *data = unpack_raw_entry(obj, &delta->base, obj->idx.sha1);\n-\t\tobj->real_type = obj->type;\n \t\tif (is_delta_type(obj->type)) {\n \t\t\tnr_deltas++;\n \t\t\tdelta->obj_no = i;\n@@ -1120,7 +1208,7 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\tcheck_against_sha1table(obj->idx.sha1);\n \t\t} else {\n \t\t\tcheck_against_sha1table(obj->idx.sha1);\n-\t\t\tsha1_object(data, NULL, obj->size, obj->type,\n+\t\t\tsha1_object(data, NULL, obj->size, obj->real_type,\n \t\t\t\t    obj->idx.sha1);\n \t\t}\n \t\tfree(data);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227094","messageId":"1378624960-8919-11-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 10/14] index-pack: parse v4 tree format","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:36Z","receivedAt":"2013-09-08T07:22:36Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"---\n builtin/index-pack.c | 105 +++++++++++++++++++++++++++++++++++++++++++++++++--\n 1 file changed, 101 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 473514a..dcb6409 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -334,6 +334,14 @@ static const unsigned char *read_sha1ref(void)\n \treturn sha1_table + index * 20;\n }\n \n+static const unsigned char *read_sha1table_ref(void)\n+{\n+\tconst unsigned char *sha1 = read_sha1ref();\n+\tif (sha1 < sha1_table || sha1 >= sha1_table + nr_objects * 20)\n+\t\tcheck_against_sha1table(sha1);\n+\treturn sha1;\n+}\n+\n static const unsigned char *read_dictref(struct packv4_dict *dict)\n {\n \tunsigned int index = read_varint();\n@@ -561,21 +569,105 @@ static void *unpack_commit_v4(unsigned int offset, unsigned long size,\n }\n \n /*\n+ * v4 trees are actually kind of deltas and we don't do delta in the\n+ * first pass. This function only walks through a tree object to find\n+ * the end offset, register object dependencies and performs limited\n+ * validation. For v4 trees that have no dependencies, we do\n+ * uncompress and calculate their SHA-1.\n+ */\n+static void *unpack_tree_v4(struct object_entry *obj,\n+\t\t\t    unsigned int offset, unsigned long size,\n+\t\t\t    unsigned char *sha1)\n+{\n+\tunsigned int nr = read_varint();\n+\tconst unsigned char *last_base = NULL;\n+\tstruct strbuf sb = STRBUF_INIT;\n+\twhile (nr) {\n+\t\tunsigned int copy_start_or_path = read_varint();\n+\t\tif (copy_start_or_path & 1) { /* copy_start */\n+\t\t\tunsigned int copy_count = read_varint();\n+\t\t\tif (copy_count & 1) { /* first delta */\n+\t\t\t\tlast_base = read_sha1table_ref();\n+\t\t\t} else if (!last_base)\n+\t\t\t\tbad_object(offset,\n+\t\t\t\t\t   _(\"missing delta base unpack_tree_v4\"));\n+\t\t\tcopy_count >>= 1;\n+\t\t\tif (!copy_count || copy_count > nr)\n+\t\t\t\tbad_object(offset,\n+\t\t\t\t\t   _(\"bad copy count index in unpack_tree_v4\"));\n+\t\t\tnr -= copy_count;\n+\t\t} else {\t/* path */\n+\t\t\tunsigned int path_idx = copy_start_or_path >> 1;\n+\t\t\tconst unsigned char *entry_sha1;\n+\n+\t\t\tif (path_idx >= path_dict->nb_entries)\n+\t\t\t\tbad_object(offset,\n+\t\t\t\t\t   _(\"bad path index in unpack_tree_v4\"));\n+\t\t\tentry_sha1 = read_sha1ref();\n+\t\t\tnr--;\n+\n+\t\t\t/*\n+\t\t\t * Attempt to rebuild a canonical (base) tree.\n+\t\t\t * If last_base is set, this tree depends on\n+\t\t\t * another tree, which we have no access at this\n+\t\t\t * stage, so reconstruction must be delayed until\n+\t\t\t * the second pass.\n+\t\t\t */\n+\t\t\tif (!last_base) {\n+\t\t\t\tconst unsigned char *path;\n+\t\t\t\tunsigned mode;\n+\n+\t\t\t\tpath = path_dict->data + path_dict->offsets[path_idx];\n+\t\t\t\tmode = (path[0] << 8) | path[1];\n+\t\t\t\tstrbuf_addf(&sb, \"%o %s%c\", mode, path+2, '\\0');\n+\t\t\t\tstrbuf_add(&sb, entry_sha1, 20);\n+\t\t\t\tif (sb.len > size)\n+\t\t\t\t\tbad_object(offset,\n+\t\t\t\t\t\t   _(\"tree larger than expected\"));\n+\t\t\t}\n+\t\t}\n+\t}\n+\n+\tif (last_base) {\n+\t\tstrbuf_release(&sb);\n+\t\treturn NULL;\n+\t} else {\n+\t\tgit_SHA_CTX ctx;\n+\t\tchar hdr[32];\n+\t\tint hdrlen;\n+\n+\t\tif (sb.len != size)\n+\t\t\tbad_object(offset, _(\"tree size mismatch\"));\n+\n+\t\thdrlen = sprintf(hdr, \"tree %lu\", size) + 1;\n+\t\tgit_SHA1_Init(&ctx);\n+\t\tgit_SHA1_Update(&ctx, hdr, hdrlen);\n+\t\tgit_SHA1_Update(&ctx, sb.buf, size);\n+\t\tgit_SHA1_Final(sha1, &ctx);\n+\t\treturn strbuf_detach(&sb, NULL);\n+\t}\n+}\n+\n+/*\n  * Unpack an entry data in the streamed pack, calculate the object\n  * SHA-1 if it's not a large blob. Otherwise just try to inflate the\n  * object to /dev/null to determine the end of the entry in the pack.\n  */\n-static void *unpack_entry_data(unsigned long offset, unsigned long size,\n-\t\t\t       enum object_type type, unsigned char *sha1)\n+static void *unpack_entry_data(struct object_entry *obj, unsigned char *sha1)\n {\n \tstatic char fixed_buf[8192];\n \tvoid *buf;\n \tgit_SHA_CTX c;\n \tchar hdr[32];\n \tint hdrlen;\n+\tunsigned long offset = obj->idx.offset;\n+\tunsigned long size = obj->size;\n+\tenum object_type type = obj->type;\n \n \tif (type == OBJ_PV4_COMMIT)\n \t\treturn unpack_commit_v4(offset, size, sha1);\n+\tif (type == OBJ_PV4_TREE)\n+\t\treturn unpack_tree_v4(obj, offset, size, sha1);\n \n \tif (!is_delta_type(type)) {\n \t\thdrlen = sprintf(hdr, \"%s %lu\", typename(type), size) + 1;\n@@ -644,16 +736,19 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n \t\tbreak;\n-\n \tcase OBJ_PV4_COMMIT:\n \t\tobj->real_type = OBJ_COMMIT;\n \t\tbreak;\n+\tcase OBJ_PV4_TREE:\n+\t\tobj->real_type = OBJ_TREE;\n+\t\tbreak;\n+\n \tdefault:\n \t\tbad_object(obj->idx.offset, _(\"unknown object type %d\"), obj->type);\n \t}\n \tobj->hdr_size = consumed_bytes - obj->idx.offset;\n \n-\tdata = unpack_entry_data(obj->idx.offset, obj->size, obj->type, sha1);\n+\tdata = unpack_entry_data(obj, sha1);\n \tobj->idx.crc32 = input_crc32;\n \treturn data;\n }\n@@ -1201,6 +1296,8 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\tnr_deltas++;\n \t\t\tdelta->obj_no = i;\n \t\t\tdelta++;\n+\t\t} else if (!data && obj->type == OBJ_PV4_TREE) {\n+\t\t\t/* delay sha1_object() until second pass */\n \t\t} else if (!data) {\n \t\t\t/* large blobs, check later */\n \t\t\tobj->real_type = OBJ_BAD;\n-- \n1.8.2.83.gc99314b\n"},{"id":"227095","messageId":"1378624960-8919-12-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 11/14] index-pack: move delta base queuing code to unpack_raw_entry","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:37Z","receivedAt":"2013-09-08T07:22:37Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"For v2, ofs-delta and ref-delta can only have queue one delta base at\na time. A v4 tree can have more than one delta base. Move the queuing\ncode up to unpack_raw_entry() and give unpack_tree_v4() more\nflexibility to add its bases.\n---\n builtin/index-pack.c | 46 ++++++++++++++++++++++++++++++----------------\n 1 file changed, 30 insertions(+), 16 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex dcb6409..8f2d929 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -568,6 +568,25 @@ static void *unpack_commit_v4(unsigned int offset, unsigned long size,\n \treturn dst.buf;\n }\n \n+static void add_sha1_delta(struct object_entry *obj,\n+\t\t\t   const unsigned char *sha1)\n+{\n+\tstruct delta_entry *delta = deltas + nr_deltas;\n+\tdelta->obj_no = obj - objects;\n+\thashcpy(delta->base.sha1, sha1);\n+\tnr_deltas++;\n+}\n+\n+static void add_ofs_delta(struct object_entry *obj,\n+\t\t\t  off_t offset)\n+{\n+\tstruct delta_entry *delta = deltas + nr_deltas;\n+\tdelta->obj_no = obj - objects;\n+\tmemset(&delta->base, 0, sizeof(delta->base));\n+\tdelta->base.offset = offset;\n+\tnr_deltas++;\n+}\n+\n /*\n  * v4 trees are actually kind of deltas and we don't do delta in the\n  * first pass. This function only walks through a tree object to find\n@@ -703,17 +722,16 @@ static void read_typesize_v2(struct object_entry *obj)\n }\n \n static void *unpack_raw_entry(struct object_entry *obj,\n-\t\t\t      union delta_base *delta_base,\n \t\t\t      unsigned char *sha1)\n {\n \tvoid *data;\n-\tuintmax_t val;\n+\toff_t offset;\n \n \tobj->idx.offset = consumed_bytes;\n \tinput_crc32 = crc32(0, NULL, 0);\n \n \tif (packv4) {\n-\t\tval = read_varint();\n+\t\tuintmax_t val = read_varint();\n \t\tobj->type = val & 15;\n \t\tobj->size = val >> 4;\n \t} else\n@@ -722,14 +740,14 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \n \tswitch (obj->type) {\n \tcase OBJ_REF_DELTA:\n-\t\thashcpy(delta_base->sha1, fill_and_use(20));\n+\t\tadd_sha1_delta(obj, fill_and_use(20));\n \t\tbreak;\n \tcase OBJ_OFS_DELTA:\n-\t\tmemset(delta_base, 0, sizeof(*delta_base));\n-\t\tval = read_varint();\n-\t\tdelta_base->offset = obj->idx.offset - val;\n-\t\tif (delta_base->offset <= 0 || delta_base->offset >= obj->idx.offset)\n-\t\t\tbad_object(obj->idx.offset, _(\"delta base offset is out of bound\"));\n+\t\toffset = obj->idx.offset - read_varint();\n+\t\tif (offset <= 0 || offset >= obj->idx.offset)\n+\t\t\tbad_object(obj->idx.offset,\n+\t\t\t\t   _(\"delta base offset is out of bound\"));\n+\t\tadd_ofs_delta(obj, offset);\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n@@ -1282,7 +1300,6 @@ static void parse_dictionaries(void)\n static void parse_pack_objects(unsigned char *sha1)\n {\n \tint i, nr_delays = 0;\n-\tstruct delta_entry *delta = deltas;\n \tstruct stat st;\n \n \tif (verbose)\n@@ -1291,12 +1308,9 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\t\tnr_objects);\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n-\t\tvoid *data = unpack_raw_entry(obj, &delta->base, obj->idx.sha1);\n-\t\tif (is_delta_type(obj->type)) {\n-\t\t\tnr_deltas++;\n-\t\t\tdelta->obj_no = i;\n-\t\t\tdelta++;\n-\t\t} else if (!data && obj->type == OBJ_PV4_TREE) {\n+\t\tvoid *data = unpack_raw_entry(obj, obj->idx.sha1);\n+\t\tif (is_delta_type(obj->type) ||\n+\t\t    (!data && obj->type == OBJ_PV4_TREE)) {\n \t\t\t/* delay sha1_object() until second pass */\n \t\t} else if (!data) {\n \t\t\t/* large blobs, check later */\n-- \n1.8.2.83.gc99314b\n"},{"id":"227096","messageId":"1378624960-8919-13-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 12/14] index-pack: record all delta bases in v4 (tree and ref-delta)","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:38Z","receivedAt":"2013-09-08T07:22:38Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"---\n builtin/index-pack.c | 33 ++++++++++++++++++++++++++++++---\n 1 file changed, 30 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 8f2d929..e903a49 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -24,6 +24,7 @@ struct object_entry {\n \tenum object_type real_type; /* type after delta resolving */\n \tunsigned delta_depth;\n \tint base_object_no;\n+\tint nr_bases;\t\t/* only valid for v4 trees */\n };\n \n union delta_base {\n@@ -482,6 +483,11 @@ static int is_delta_type(enum object_type type)\n \treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n }\n \n+static int is_delta_tree(const struct object_entry *obj)\n+{\n+\treturn obj->type == OBJ_PV4_TREE && obj->nr_bases > 0;\n+}\n+\n static void read_and_inflate(unsigned long offset,\n \t\t\t     void *buf, unsigned long size,\n \t\t\t     unsigned long wraparound,\n@@ -587,6 +593,20 @@ static void add_ofs_delta(struct object_entry *obj,\n \tnr_deltas++;\n }\n \n+static void add_tree_delta_base(struct object_entry *obj,\n+\t\t\t\tconst unsigned char *base,\n+\t\t\t\tint delta_start)\n+{\n+\tint i;\n+\n+\tfor (i = delta_start; i < nr_deltas; i++)\n+\t\tif (!hashcmp(base, deltas[i].base.sha1))\n+\t\t\treturn;\n+\n+\tadd_sha1_delta(obj, base);\n+\tobj->nr_bases++;\n+}\n+\n /*\n  * v4 trees are actually kind of deltas and we don't do delta in the\n  * first pass. This function only walks through a tree object to find\n@@ -601,12 +621,14 @@ static void *unpack_tree_v4(struct object_entry *obj,\n \tunsigned int nr = read_varint();\n \tconst unsigned char *last_base = NULL;\n \tstruct strbuf sb = STRBUF_INIT;\n+\tint delta_start = nr_deltas;\n \twhile (nr) {\n \t\tunsigned int copy_start_or_path = read_varint();\n \t\tif (copy_start_or_path & 1) { /* copy_start */\n \t\t\tunsigned int copy_count = read_varint();\n \t\t\tif (copy_count & 1) { /* first delta */\n \t\t\t\tlast_base = read_sha1table_ref();\n+\t\t\t\tadd_tree_delta_base(obj, last_base, delta_start);\n \t\t\t} else if (!last_base)\n \t\t\t\tbad_object(offset,\n \t\t\t\t\t   _(\"missing delta base unpack_tree_v4\"));\n@@ -740,9 +762,15 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \n \tswitch (obj->type) {\n \tcase OBJ_REF_DELTA:\n-\t\tadd_sha1_delta(obj, fill_and_use(20));\n+\t\tif (packv4)\n+\t\t\tadd_sha1_delta(obj, read_sha1table_ref());\n+\t\telse\n+\t\t\tadd_sha1_delta(obj, fill_and_use(20));\n \t\tbreak;\n \tcase OBJ_OFS_DELTA:\n+\t\tif (packv4)\n+\t\t\tdie(_(\"pack version 4 does not support ofs-delta type (offset %lu)\"),\n+\t\t\t    obj->idx.offset);\n \t\toffset = obj->idx.offset - read_varint();\n \t\tif (offset <= 0 || offset >= obj->idx.offset)\n \t\t\tbad_object(obj->idx.offset,\n@@ -1309,8 +1337,7 @@ static void parse_pack_objects(unsigned char *sha1)\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n \t\tvoid *data = unpack_raw_entry(obj, obj->idx.sha1);\n-\t\tif (is_delta_type(obj->type) ||\n-\t\t    (!data && obj->type == OBJ_PV4_TREE)) {\n+\t\tif (is_delta_type(obj->type) || is_delta_tree(obj)) {\n \t\t\t/* delay sha1_object() until second pass */\n \t\t} else if (!data) {\n \t\t\t/* large blobs, check later */\n-- \n1.8.2.83.gc99314b\n"},{"id":"227097","messageId":"1378624960-8919-14-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 13/14] index-pack: skip looking for ofs-deltas in v4 as they are not allowed","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:39Z","receivedAt":"2013-09-08T07:22:39Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"---\n builtin/index-pack.c | 11 +++++++----\n 1 file changed, 7 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex e903a49..ce06473 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -1173,10 +1173,13 @@ static struct base_data *find_unresolved_deltas_1(struct base_data *base,\n \t\tfind_delta_children(&base_spec,\n \t\t\t\t    &base->ref_first, &base->ref_last, OBJ_REF_DELTA);\n \n-\t\tmemset(&base_spec, 0, sizeof(base_spec));\n-\t\tbase_spec.offset = base->obj->idx.offset;\n-\t\tfind_delta_children(&base_spec,\n-\t\t\t\t    &base->ofs_first, &base->ofs_last, OBJ_OFS_DELTA);\n+\t\tif (!packv4) {\n+\t\t\tmemset(&base_spec, 0, sizeof(base_spec));\n+\t\t\tbase_spec.offset = base->obj->idx.offset;\n+\t\t\tfind_delta_children(&base_spec,\n+\t\t\t\t\t    &base->ofs_first, &base->ofs_last,\n+\t\t\t\t\t    OBJ_OFS_DELTA);\n+\t\t}\n \n \t\tif (base->ref_last == -1 && base->ofs_last == -1) {\n \t\t\tfree(base->data);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227098","messageId":"1378624960-8919-15-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378624960-8919-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 14/14] index-pack: resolve v4 one-base trees","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T07:22:40Z","receivedAt":"2013-09-08T07:22:40Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This is the most common case for delta trees. In fact it's the only\nkind that's produced by packv4-create. It fits well in the way\nindex-pack resolves deltas and benefits from threading (the set of\nobjects depending on this base does not overlap with the set of\nobjects depending on another base)\n\nMulti-base trees will be probably processed differently.\n---\n builtin/index-pack.c | 195 ++++++++++++++++++++++++++++++++++++++++++++++-----\n 1 file changed, 179 insertions(+), 16 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex ce06473..88340b5 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -12,6 +12,8 @@\n #include \"streaming.h\"\n #include \"thread-utils.h\"\n #include \"packv4-parse.h\"\n+#include \"varint.h\"\n+#include \"tree-walk.h\"\n \n static const char index_pack_usage[] =\n \"git index-pack [-v] [-o <index-file>] [--keep | --keep=<msg>] [--verify] [--strict] (<pack-file> | --stdin [--fix-thin] [<pack-file>])\";\n@@ -38,8 +40,8 @@ struct base_data {\n \tstruct object_entry *obj;\n \tvoid *data;\n \tunsigned long size;\n-\tint ref_first, ref_last;\n-\tint ofs_first, ofs_last;\n+\tint ref_first, ref_last, tree_first;\n+\tint ofs_first, ofs_last, tree_last;\n };\n \n #if !defined(NO_PTHREADS) && defined(NO_THREAD_SAFE_PREAD)\n@@ -430,6 +432,7 @@ static struct base_data *alloc_base_data(void)\n \tmemset(base, 0, sizeof(*base));\n \tbase->ref_last = -1;\n \tbase->ofs_last = -1;\n+\tbase->tree_last = -1;\n \treturn base;\n }\n \n@@ -670,6 +673,8 @@ static void *unpack_tree_v4(struct object_entry *obj,\n \t}\n \n \tif (last_base) {\n+\t\tif (nr_deltas - delta_start > 1)\n+\t\t\tdie(\"sorry guys, multi-base trees are not supported yet\");\n \t\tstrbuf_release(&sb);\n \t\treturn NULL;\n \t} else {\n@@ -800,6 +805,84 @@ static void *unpack_raw_entry(struct object_entry *obj,\n }\n \n /*\n+ * Some checks are skipped because they are already done by\n+ * unpack_tree_v4() in the first pass.\n+ */\n+static void *patch_one_base_tree(const struct object_entry *src,\n+\t\t\t\t const unsigned char *src_buf,\n+\t\t\t\t const unsigned char *delta_buf,\n+\t\t\t\t unsigned long delta_size,\n+\t\t\t\t unsigned long *dst_size)\n+{\n+\tint nr;\n+\tconst unsigned char *last_base = NULL;\n+\tstruct strbuf sb = STRBUF_INIT;\n+\tconst unsigned char *p = delta_buf;\n+\n+\tnr = decode_varint(&p);\n+\twhile (nr > 0 && p < delta_buf + delta_size) {\n+\t\tunsigned int copy_start_or_path = decode_varint(&p);\n+\t\tif (copy_start_or_path & 1) { /* copy_start */\n+\t\t\tstruct tree_desc desc;\n+\t\t\tstruct name_entry entry;\n+\t\t\tunsigned int copy_count = decode_varint(&p);\n+\t\t\tunsigned int copy_start = copy_start_or_path >> 1;\n+\t\t\tif (!src)\n+\t\t\t\tdie(\"we are not supposed to copy from another tree!\");\n+\t\t\tif (copy_count & 1) { /* first delta */\n+\t\t\t\tunsigned int id = decode_varint(&p);\n+\t\t\t\tif (!id) {\n+\t\t\t\t\tlast_base = p;\n+\t\t\t\t\tp += 20;\n+\t\t\t\t} else\n+\t\t\t\t\tlast_base = sha1_table + (id - 1) * 20;\n+\t\t\t\tif (hashcmp(last_base, src->idx.sha1))\n+\t\t\t\t\tdie(_(\"bad tree base in patch_one_base_tree\"));\n+\t\t\t}\n+\n+\t\t\tcopy_count >>= 1;\n+\t\t\tnr -= copy_count;\n+\n+\t\t\tinit_tree_desc(&desc, src_buf, src->size);\n+\t\t\twhile (tree_entry(&desc, &entry)) {\n+\t\t\t\tif (copy_start)\n+\t\t\t\t\tcopy_start--;\n+\t\t\t\telse if (copy_count) {\n+\t\t\t\t\tstrbuf_addf(&sb, \"%o %s%c\",\n+\t\t\t\t\t\t    entry.mode, entry.path, '\\0');\n+\t\t\t\t\tstrbuf_add(&sb, entry.sha1, 20);\n+\t\t\t\t\tcopy_count--;\n+\t\t\t\t} else\n+\t\t\t\t\tbreak;\n+\t\t\t}\n+\t\t} else {\t/* path */\n+\t\t\tunsigned int path_idx = copy_start_or_path >> 1;\n+\t\t\tconst unsigned char *path;\n+\t\t\tunsigned mode;\n+\t\t\tunsigned int id;\n+\t\t\tconst unsigned char *entry_sha1;\n+\n+\t\t\tid = decode_varint(&p);\n+\t\t\tif (!id) {\n+\t\t\t\tentry_sha1 = p;\n+\t\t\t\tp += 20;\n+\t\t\t} else\n+\t\t\t\tentry_sha1 = sha1_table + (id - 1) * 20;\n+\t\t\tnr--;\n+\n+\t\t\tpath = path_dict->data + path_dict->offsets[path_idx];\n+\t\t\tmode = (path[0] << 8) | path[1];\n+\t\t\tstrbuf_addf(&sb, \"%o %s%c\", mode, path+2, '\\0');\n+\t\t\tstrbuf_add(&sb, entry_sha1, 20);\n+\t\t}\n+\t}\n+\tif (nr != 0 || p != delta_buf + delta_size)\n+\t\tdie(_(\"bad delta tree\"));\n+\t*dst_size = sb.len;\n+\treturn sb.buf;\n+}\n+\n+/*\n  * Unpack entry data in the second pass when the pack is already\n  * stored on disk. consume call back is used for large-blob case. Must\n  * be thread safe.\n@@ -865,8 +948,33 @@ static void *unpack_data(struct object_entry *obj,\n \treturn data;\n }\n \n+static void *get_tree_v4_from_pack(struct object_entry *obj,\n+\t\t\t\t   unsigned long *len_p)\n+{\n+\toff_t from = obj[0].idx.offset + obj[0].hdr_size;\n+\tunsigned long len = obj[1].idx.offset - from;\n+\tunsigned char *data;\n+\tssize_t n;\n+\n+\tdata = xmalloc(len);\n+\tn = pread(pack_fd, data, len, from);\n+\tif (n < 0)\n+\t\tdie_errno(_(\"cannot pread pack file\"));\n+\tif (!n)\n+\t\tdie(Q_(\"premature end of pack file, %lu byte missing\",\n+\t\t       \"premature end of pack file, %lu bytes missing\",\n+\t\t       len),\n+\t\t    len);\n+\tif (len_p)\n+\t\t*len_p = len;\n+\treturn data;\n+}\n+\n static void *get_data_from_pack(struct object_entry *obj)\n {\n+\tif (obj->type == OBJ_PV4_COMMIT || obj->type == OBJ_PV4_TREE)\n+\t\tdie(\"BUG: unsupported code path\");\n+\n \treturn unpack_data(obj, NULL, NULL);\n }\n \n@@ -1093,14 +1201,25 @@ static void *get_base_data(struct base_data *c)\n \t\tstruct object_entry *obj = c->obj;\n \t\tstruct base_data **delta = NULL;\n \t\tint delta_nr = 0, delta_alloc = 0;\n+\t\tunsigned long size, len;\n \n-\t\twhile (is_delta_type(c->obj->type) && !c->data) {\n+\t\twhile ((is_delta_type(c->obj->type) ||\n+\t\t\t(c->base && c->obj->type == OBJ_PV4_TREE)) &&\n+\t\t       !c->data) {\n \t\t\tALLOC_GROW(delta, delta_nr + 1, delta_alloc);\n \t\t\tdelta[delta_nr++] = c;\n \t\t\tc = c->base;\n \t\t}\n \t\tif (!delta_nr) {\n-\t\t\tc->data = get_data_from_pack(obj);\n+\t\t\tif (c->obj->type == OBJ_PV4_TREE) {\n+\t\t\t\tvoid *tree_v4 = get_tree_v4_from_pack(obj, &len);\n+\t\t\t\tc->data = patch_one_base_tree(NULL, NULL,\n+\t\t\t\t\t\t\t      tree_v4, len, &size);\n+\t\t\t\tif (size != obj->size)\n+\t\t\t\t\tdie(\"size mismatch\");\n+\t\t\t\tfree(tree_v4);\n+\t\t\t} else\n+\t\t\t\tc->data = get_data_from_pack(obj);\n \t\t\tc->size = obj->size;\n \t\t\tget_thread_data()->base_cache_used += c->size;\n \t\t\tprune_base_data(c);\n@@ -1110,11 +1229,18 @@ static void *get_base_data(struct base_data *c)\n \t\t\tc = delta[delta_nr - 1];\n \t\t\tobj = c->obj;\n \t\t\tbase = get_base_data(c->base);\n-\t\t\traw = get_data_from_pack(obj);\n-\t\t\tc->data = patch_delta(\n-\t\t\t\tbase, c->base->size,\n-\t\t\t\traw, obj->size,\n-\t\t\t\t&c->size);\n+\t\t\tif (c->obj->type == OBJ_PV4_TREE) {\n+\t\t\t\traw = get_tree_v4_from_pack(obj, &len);\n+\t\t\t\tc->data = patch_one_base_tree(c->base->obj, base,\n+\t\t\t\t\t\t\t      raw, len, &size);\n+\t\t\t\tif (size != obj->size)\n+\t\t\t\t\tdie(\"size mismatch\");\n+\t\t\t} else {\n+\t\t\t\traw = get_data_from_pack(obj);\n+\t\t\t\tc->data = patch_delta(base, c->base->size,\n+\t\t\t\t\t\t      raw, obj->size,\n+\t\t\t\t\t\t      &c->size);\n+\t\t\t}\n \t\t\tfree(raw);\n \t\t\tif (!c->data)\n \t\t\t\tbad_object(obj->idx.offset, _(\"failed to apply delta\"));\n@@ -1130,6 +1256,8 @@ static void resolve_delta(struct object_entry *delta_obj,\n \t\t\t  struct base_data *base, struct base_data *result)\n {\n \tvoid *base_data, *delta_data;\n+\tint tree_v4 = delta_obj->type == OBJ_PV4_TREE;\n+\tunsigned long tree_size;\n \n \tdelta_obj->real_type = base->obj->real_type;\n \tif (show_stat) {\n@@ -1140,10 +1268,18 @@ static void resolve_delta(struct object_entry *delta_obj,\n \t\tdeepest_delta_unlock();\n \t}\n \tdelta_obj->base_object_no = base->obj - objects;\n-\tdelta_data = get_data_from_pack(delta_obj);\n+\tif (tree_v4)\n+\t\tdelta_data = get_tree_v4_from_pack(delta_obj, &tree_size);\n+\telse\n+\t\tdelta_data = get_data_from_pack(delta_obj);\n \tbase_data = get_base_data(base);\n \tresult->obj = delta_obj;\n-\tresult->data = patch_delta(base_data, base->size,\n+\tif (tree_v4)\n+\t\tresult->data = patch_one_base_tree(base->obj, base_data,\n+\t\t\t\t\t\t   delta_data, tree_size,\n+\t\t\t\t\t\t   &result->size);\n+\telse\n+\t\tresult->data = patch_delta(base_data, base->size,\n \t\t\t\t   delta_data, delta_obj->size, &result->size);\n \tfree(delta_data);\n \tif (!result->data)\n@@ -1166,7 +1302,8 @@ static void resolve_delta(struct object_entry *delta_obj,\n static struct base_data *find_unresolved_deltas_1(struct base_data *base,\n \t\t\t\t\t\t  struct base_data *prev_base)\n {\n-\tif (base->ref_last == -1 && base->ofs_last == -1) {\n+\tif (base->ref_last == -1 && base->ofs_last == -1 &&\n+\t    base->tree_last == -1) {\n \t\tunion delta_base base_spec;\n \n \t\thashcpy(base_spec.sha1, base->obj->idx.sha1);\n@@ -1179,9 +1316,15 @@ static struct base_data *find_unresolved_deltas_1(struct base_data *base,\n \t\t\tfind_delta_children(&base_spec,\n \t\t\t\t\t    &base->ofs_first, &base->ofs_last,\n \t\t\t\t\t    OBJ_OFS_DELTA);\n+\t\t} else {\n+\t\t\thashcpy(base_spec.sha1, base->obj->idx.sha1);\n+\t\t\tfind_delta_children(&base_spec,\n+\t\t\t\t\t    &base->tree_first, &base->tree_last,\n+\t\t\t\t\t    OBJ_PV4_TREE);\n \t\t}\n \n-\t\tif (base->ref_last == -1 && base->ofs_last == -1) {\n+\t\tif (base->ref_last == -1 && base->ofs_last == -1 &&\n+\t\t    base->tree_last == -1) {\n \t\t\tfree(base->data);\n \t\t\treturn NULL;\n \t\t}\n@@ -1215,6 +1358,25 @@ static struct base_data *find_unresolved_deltas_1(struct base_data *base,\n \t\treturn result;\n \t}\n \n+\twhile (base->tree_first <= base->tree_last) {\n+\t\tstruct object_entry *child = objects + deltas[base->tree_first].obj_no;\n+\t\tstruct base_data *result;\n+\n+\t\tassert(child->type == OBJ_PV4_TREE);\n+\t\tif (child->nr_bases > 1) {\n+\t\t\t/* maybe resolved in the third pass or something */\n+\t\t\tbase->tree_first++;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tresult = alloc_base_data();\n+\t\tresolve_delta(child, base, result);\n+\t\tif (base->tree_first == base->tree_last)\n+\t\t\tfree_base_data(base);\n+\n+\t\tbase->tree_first++;\n+\t\treturn result;\n+\t}\n+\n \tunlink_base_data(base);\n \treturn NULL;\n }\n@@ -1273,7 +1435,8 @@ static void *threaded_second_pass(void *data)\n \t\tcounter_unlock();\n \t\twork_lock();\n \t\twhile (nr_dispatched < nr_objects &&\n-\t\t       is_delta_type(objects[nr_dispatched].type))\n+\t\t       (is_delta_type(objects[nr_dispatched].type) ||\n+\t\t\tis_delta_tree(objects + nr_dispatched)))\n \t\t\tnr_dispatched++;\n \t\tif (nr_dispatched >= nr_objects) {\n \t\t\twork_unlock();\n@@ -1427,7 +1590,7 @@ static void resolve_deltas(void)\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n \n-\t\tif (is_delta_type(obj->type))\n+\t\tif (is_delta_type(obj->type) || is_delta_tree(obj))\n \t\t\tcontinue;\n \t\tresolve_base(obj);\n \t\tdisplay_progress(progress, nr_resolved_deltas);\n@@ -1972,7 +2135,7 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \t\tshow_pack_info(stat_only);\n \n \tif (packv4)\n-\t\tdie(\"we're not there yet\");\n+\t\topts.version = 3;\n \n \tidx_objects = xmalloc((nr_objects) * sizeof(struct pack_idx_entry *));\n \tfor (i = 0; i < nr_objects; i++)\n-- \n1.8.2.83.gc99314b\n"},{"id":"227126","messageId":"1378652660-6731-1-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378362001-1738-1-git-send-email-nico@fluxnic.net","subject":"[PATCH 00/11] pack v4 support in pack-objects","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:09Z","receivedAt":"2013-09-08T15:04:09Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"I can produce pack v4 on git.git with this and verify it with\nindex-pack. I'm not familiar with pack-objects code and not really\nconfident with my changes. Suggestions are welcome.\n\nAlso I chose to keep packv4-create.c in libgit.a and move test code\nout to test-packv4.c. Not sure if it's good decision. The other option\nis to copy necessary code to pack-objects.c, then delete\npackv4-create.c in the end. Either way we have the same amount of code\nmove.\n\nThin pack support is not there yet, but it should be simple on\npack-objects' end. Like the compatibility layer you added to\nsha1_file.c, this code does not take advantage of v4 as source packs\n(performance regressions entail) A lot of rooms for improvements.\n\nNguyễn Thái Ngọc Duy (11):\n  pack v4: allocate dicts from the beginning\n  pack v4: stop using static/global variables in packv4-create.c\n  pack v4: move packv4-create.c to libgit.a\n  pack v4: add version argument to write_pack_header\n  pack-write.c: add pv4_encode_in_pack_object_header\n  pack-objects: add --version to specify written pack version\n  list-objects.c: add show_tree_entry callback to traverse_commit_list\n  pack-objects: create pack v4 tables\n  pack-objects: do not cache delta for v4 trees\n  pack-objects: exclude commits out of delta objects in v4\n  pack-objects: support writing pack v4\n\n Makefile               |   4 +-\n builtin/pack-objects.c | 187 +++++++++++++++--\n builtin/rev-list.c     |   4 +-\n bulk-checkin.c         |   2 +-\n list-objects.c         |   9 +-\n list-objects.h         |   3 +-\n pack-write.c           |  36 +++-\n pack.h                 |   6 +-\n packv4-create.c        | 534 ++++---------------------------------------------\n packv4-create.h (new)  |  50 +++++\n test-packv4.c (new)    | 476 +++++++++++++++++++++++++++++++++++++++++++\n upload-pack.c          |   2 +-\n 12 files changed, 789 insertions(+), 524 deletions(-)\n create mode 100644 packv4-create.h\n create mode 100644 test-packv4.c\n\n-- \n1.8.2.83.gc99314b\n"},{"id":"227127","messageId":"1378652660-6731-2-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 01/11] pack v4: allocate dicts from the beginning","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:10Z","receivedAt":"2013-09-08T15:04:10Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"commit_ident_table and tree_path_table are local to packv4-create.c\nand test-packv4.c. Move them out of add_*_dict_entries so\nadd_*_dict_entries can be exported to pack-objects.c\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n packv4-create.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 38fa594..dbc2a03 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -181,14 +181,12 @@ static char *get_nameend_and_tz(char *from, int *tz_val)\n \treturn end;\n }\n \n-static int add_commit_dict_entries(void *buf, unsigned long size)\n+int add_commit_dict_entries(struct dict_table *commit_ident_table,\n+\t\t\t    void *buf, unsigned long size)\n {\n \tchar *name, *end = NULL;\n \tint tz_val;\n \n-\tif (!commit_ident_table)\n-\t\tcommit_ident_table = create_dict_table();\n-\n \t/* parse and add author info */\n \tname = strstr(buf, \"\\nauthor \");\n \tif (name) {\n@@ -212,14 +210,12 @@ static int add_commit_dict_entries(void *buf, unsigned long size)\n \treturn 0;\n }\n \n-static int add_tree_dict_entries(void *buf, unsigned long size)\n+static int add_tree_dict_entries(struct dict_table *tree_path_table,\n+\t\t\t\t void *buf, unsigned long size)\n {\n \tstruct tree_desc desc;\n \tstruct name_entry name_entry;\n \n-\tif (!tree_path_table)\n-\t\ttree_path_table = create_dict_table();\n-\n \tinit_tree_desc(&desc, buf, size);\n \twhile (tree_entry(&desc, &name_entry)) {\n \t\tint pathlen = tree_entry_len(&name_entry);\n@@ -659,6 +655,9 @@ static int create_pack_dictionaries(struct packed_git *p,\n \tstruct progress *progress_state;\n \tunsigned int i;\n \n+\tcommit_ident_table = create_dict_table();\n+\ttree_path_table = create_dict_table();\n+\n \tprogress_state = start_progress(\"Scanning objects\", p->num_objects);\n \tfor (i = 0; i < p->num_objects; i++) {\n \t\tstruct pack_idx_entry *obj = obj_list[i];\n@@ -666,7 +665,8 @@ static int create_pack_dictionaries(struct packed_git *p,\n \t\tenum object_type type;\n \t\tunsigned long size;\n \t\tstruct object_info oi = {};\n-\t\tint (*add_dict_entries)(void *, unsigned long);\n+\t\tint (*add_dict_entries)(struct dict_table *, void *, unsigned long);\n+\t\tstruct dict_table *dict;\n \n \t\tdisplay_progress(progress_state, i+1);\n \n@@ -679,9 +679,11 @@ static int create_pack_dictionaries(struct packed_git *p,\n \t\tswitch (type) {\n \t\tcase OBJ_COMMIT:\n \t\t\tadd_dict_entries = add_commit_dict_entries;\n+\t\t\tdict = commit_ident_table;\n \t\t\tbreak;\n \t\tcase OBJ_TREE:\n \t\t\tadd_dict_entries = add_tree_dict_entries;\n+\t\t\tdict = tree_path_table;\n \t\t\tbreak;\n \t\tdefault:\n \t\t\tcontinue;\n@@ -693,7 +695,7 @@ static int create_pack_dictionaries(struct packed_git *p,\n \t\tif (check_sha1_signature(obj->sha1, data, size, typename(type)))\n \t\t\tdie(\"packed %s from %s is corrupt\",\n \t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\t\tif (add_dict_entries(data, size) < 0)\n+\t\tif (add_dict_entries(dict, data, size) < 0)\n \t\t\tdie(\"can't process %s object %s\",\n \t\t\t\ttypename(type), sha1_to_hex(obj->sha1));\n \t\tfree(data);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227128","messageId":"1378652660-6731-3-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 02/11] pack v4: stop using static/global variables in packv4-create.c","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:11Z","receivedAt":"2013-09-08T15:04:11Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n packv4-create.c       | 103 ++++++++++++++++++++++++++++----------------------\n packv4-create.h (new) |  11 ++++++\n 2 files changed, 69 insertions(+), 45 deletions(-)\n create mode 100644 packv4-create.h\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex dbc2a03..920a0b4 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -15,6 +15,7 @@\n #include \"pack-revindex.h\"\n #include \"progress.h\"\n #include \"varint.h\"\n+#include \"packv4-create.h\"\n \n \n static int pack_compression_seen;\n@@ -145,9 +146,6 @@ static void sort_dict_entries_by_hits(struct dict_table *t)\n \trehash_entries(t);\n }\n \n-static struct dict_table *commit_ident_table;\n-static struct dict_table *tree_path_table;\n-\n /*\n  * Parse the author/committer line from a canonical commit object.\n  * The 'from' argument points right after the \"author \" or \"committer \"\n@@ -243,10 +241,10 @@ void dump_dict_table(struct dict_table *t)\n \t}\n }\n \n-static void dict_dump(void)\n+static void dict_dump(struct packv4_tables *v4)\n {\n-\tdump_dict_table(commit_ident_table);\n-\tdump_dict_table(tree_path_table);\n+\tdump_dict_table(v4->commit_ident_table);\n+\tdump_dict_table(v4->tree_path_table);\n }\n \n /*\n@@ -254,10 +252,12 @@ static void dict_dump(void)\n  * pack SHA1 table incremented by 1, or the literal SHA1 value prefixed\n  * with a zero byte if the needed SHA1 is not available in the table.\n  */\n-static struct pack_idx_entry *all_objs;\n-static unsigned all_objs_nr;\n-static int encode_sha1ref(const unsigned char *sha1, unsigned char *buf)\n+\n+int encode_sha1ref(const struct packv4_tables *v4,\n+\t\t   const unsigned char *sha1, unsigned char *buf)\n {\n+\tunsigned all_objs_nr = v4->all_objs_nr;\n+\tstruct pack_idx_entry *all_objs = v4->all_objs;\n \tunsigned lo = 0, hi = all_objs_nr;\n \n \tdo {\n@@ -284,7 +284,8 @@ static int encode_sha1ref(const unsigned char *sha1, unsigned char *buf)\n  * strict so to ensure the canonical version may always be\n  * regenerated and produce the same hash.\n  */\n-void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n+void *pv4_encode_commit(const struct packv4_tables *v4,\n+\t\t\tvoid *buffer, unsigned long *sizep)\n {\n \tunsigned long size = *sizep;\n \tchar *in, *tail, *end;\n@@ -310,7 +311,7 @@ void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n \tif (get_sha1_lowhex(in + 5, sha1) < 0)\n \t\tgoto bad_data;\n \tin += 46;\n-\tout += encode_sha1ref(sha1, out);\n+\tout += encode_sha1ref(v4, sha1, out);\n \n \t/* count how many \"parent\" lines */\n \tnb_parents = 0;\n@@ -325,7 +326,7 @@ void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n \twhile (nb_parents--) {\n \t\tif (get_sha1_lowhex(in + 7, sha1))\n \t\t\tgoto bad_data;\n-\t\tout += encode_sha1ref(sha1, out);\n+\t\tout += encode_sha1ref(v4, sha1, out);\n \t\tin += 48;\n \t}\n \n@@ -337,7 +338,7 @@ void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n \tend = get_nameend_and_tz(in, &tz_val);\n \tif (!end)\n \t\tgoto bad_data;\n-\tauthor_index = dict_add_entry(commit_ident_table, tz_val, in, end - in);\n+\tauthor_index = dict_add_entry(v4->commit_ident_table, tz_val, in, end - in);\n \tif (author_index < 0)\n \t\tgoto bad_dict;\n \tauthor_time = strtoul(end, &end, 10);\n@@ -353,7 +354,7 @@ void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n \tend = get_nameend_and_tz(in, &tz_val);\n \tif (!end)\n \t\tgoto bad_data;\n-\tcommit_index = dict_add_entry(commit_ident_table, tz_val, in, end - in);\n+\tcommit_index = dict_add_entry(v4->commit_ident_table, tz_val, in, end - in);\n \tif (commit_index < 0)\n \t\tgoto bad_dict;\n \tcommit_time = strtoul(end, &end, 10);\n@@ -436,7 +437,8 @@ static int compare_tree_entries(struct name_entry *e1, struct name_entry *e2)\n  * If a delta buffer is provided, we may encode multiple ranges of tree\n  * entries against that buffer.\n  */\n-void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n+void *pv4_encode_tree(const struct packv4_tables *v4,\n+\t\t      void *_buffer, unsigned long *sizep,\n \t\t      void *delta, unsigned long delta_size,\n \t\t      const unsigned char *delta_sha1)\n {\n@@ -551,7 +553,7 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\t\tcp += encode_varint(copy_start, cp);\n \t\t\tcp += encode_varint(copy_count, cp);\n \t\t\tif (first_delta)\n-\t\t\t\tcp += encode_sha1ref(delta_sha1, cp);\n+\t\t\t\tcp += encode_sha1ref(v4, delta_sha1, cp);\n \n \t\t\t/*\n \t\t\t * Now let's make sure this is going to take less\n@@ -577,7 +579,7 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\t}\n \n \t\tpathlen = tree_entry_len(&name_entry);\n-\t\tindex = dict_add_entry(tree_path_table, name_entry.mode,\n+\t\tindex = dict_add_entry(v4->tree_path_table, name_entry.mode,\n \t\t\t\t       name_entry.path, pathlen);\n \t\tif (index < 0) {\n \t\t\terror(\"missing tree dict entry\");\n@@ -585,7 +587,7 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\t\treturn NULL;\n \t\t}\n \t\tout += encode_varint(index << 1, out);\n-\t\tout += encode_sha1ref(name_entry.sha1, out);\n+\t\tout += encode_sha1ref(v4, name_entry.sha1, out);\n \t}\n \n \tif (copy_count) {\n@@ -596,7 +598,7 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\tcp += encode_varint(copy_start, cp);\n \t\tcp += encode_varint(copy_count, cp);\n \t\tif (first_delta)\n-\t\t\tcp += encode_sha1ref(delta_sha1, cp);\n+\t\t\tcp += encode_sha1ref(v4, delta_sha1, cp);\n \t\tif (copy_count >= min_tree_copy &&\n \t\t    cp - copy_buf < out - &buffer[copy_pos]) {\n \t\t\tout = buffer + copy_pos;\n@@ -649,14 +651,15 @@ static struct pack_idx_entry **sort_objs_by_offset(struct pack_idx_entry *list,\n \treturn sorted;\n }\n \n-static int create_pack_dictionaries(struct packed_git *p,\n+static int create_pack_dictionaries(struct packv4_tables *v4,\n+\t\t\t\t    struct packed_git *p,\n \t\t\t\t    struct pack_idx_entry **obj_list)\n {\n \tstruct progress *progress_state;\n \tunsigned int i;\n \n-\tcommit_ident_table = create_dict_table();\n-\ttree_path_table = create_dict_table();\n+\tv4->commit_ident_table = create_dict_table();\n+\tv4->tree_path_table = create_dict_table();\n \n \tprogress_state = start_progress(\"Scanning objects\", p->num_objects);\n \tfor (i = 0; i < p->num_objects; i++) {\n@@ -679,11 +682,11 @@ static int create_pack_dictionaries(struct packed_git *p,\n \t\tswitch (type) {\n \t\tcase OBJ_COMMIT:\n \t\t\tadd_dict_entries = add_commit_dict_entries;\n-\t\t\tdict = commit_ident_table;\n+\t\t\tdict = v4->commit_ident_table;\n \t\t\tbreak;\n \t\tcase OBJ_TREE:\n \t\t\tadd_dict_entries = add_tree_dict_entries;\n-\t\t\tdict = tree_path_table;\n+\t\t\tdict = v4->tree_path_table;\n \t\t\tbreak;\n \t\tdefault:\n \t\t\tcontinue;\n@@ -776,9 +779,13 @@ static unsigned int packv4_write_header(struct sha1file *f, unsigned nr_objects)\n \treturn sizeof(hdr);\n }\n \n-static unsigned long packv4_write_tables(struct sha1file *f, unsigned nr_objects,\n-\t\t\t\t\t struct pack_idx_entry *objs)\n+unsigned long packv4_write_tables(struct sha1file *f,\n+\t\t\t\t  const struct packv4_tables *v4)\n {\n+\tunsigned nr_objects = v4->all_objs_nr;\n+\tstruct pack_idx_entry *objs = v4->all_objs;\n+\tstruct dict_table *commit_ident_table = v4->commit_ident_table;\n+\tstruct dict_table *tree_path_table = v4->tree_path_table;\n \tunsigned i;\n \tunsigned long written = 0;\n \n@@ -823,7 +830,8 @@ static int write_object_header(struct sha1file *f, enum object_type type, unsign\n \treturn len;\n }\n \n-static unsigned long copy_object_data(struct sha1file *f, struct packed_git *p,\n+static unsigned long copy_object_data(struct packv4_tables *v4,\n+\t\t\t\t      struct sha1file *f, struct packed_git *p,\n \t\t\t\t      off_t offset)\n {\n \tstruct pack_window *w_curs = NULL;\n@@ -850,11 +858,13 @@ static unsigned long copy_object_data(struct sha1file *f, struct packed_git *p,\n \t\tif (base_offset <= 0 || base_offset >= offset)\n \t\t\tdie(\"delta offset out of bound\");\n \t\trevidx = find_pack_revindex(p, base_offset);\n-\t\treflen = encode_sha1ref(nth_packed_object_sha1(p, revidx->nr), buf);\n+\t\treflen = encode_sha1ref(v4,\n+\t\t\t\t\tnth_packed_object_sha1(p, revidx->nr),\n+\t\t\t\t\tbuf);\n \t\tsha1write(f, buf, reflen);\n \t\twritten += reflen;\n \t} else if (type == OBJ_REF_DELTA) {\n-\t\treflen = encode_sha1ref(src + hdrlen, buf);\n+\t\treflen = encode_sha1ref(v4, src + hdrlen, buf);\n \t\thdrlen += 20;\n \t\tsha1write(f, buf, reflen);\n \t\twritten += reflen;\n@@ -919,7 +929,8 @@ static unsigned char *get_delta_base(struct packed_git *p, off_t offset,\n \treturn sha1_buf;\n }\n \n-static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n+static off_t packv4_write_object(struct packv4_tables *v4,\n+\t\t\t\t struct sha1file *f, struct packed_git *p,\n \t\t\t\t struct pack_idx_entry *obj)\n {\n \tvoid *src, *result;\n@@ -941,7 +952,7 @@ static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \tcase OBJ_TREE:\n \t\tbreak;\n \tdefault:\n-\t\treturn copy_object_data(f, p, obj->offset);\n+\t\treturn copy_object_data(v4, f, p, obj->offset);\n \t}\n \n \t/* The rest is converted into their new format */\n@@ -955,7 +966,7 @@ static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \n \tswitch (type) {\n \tcase OBJ_COMMIT:\n-\t\tresult = pv4_encode_commit(src, &buf_size);\n+\t\tresult = pv4_encode_commit(v4, src, &buf_size);\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tif (packed_type != OBJ_TREE) {\n@@ -972,11 +983,12 @@ static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \t\t\tif (!ref || ref_type != OBJ_TREE)\n \t\t\t\tdie(\"cannot obtain delta base for %s\",\n \t\t\t\t\t\tsha1_to_hex(obj->sha1));\n-\t\t\tresult = pv4_encode_tree(src, &buf_size,\n+\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n \t\t\t\t\t\t ref, ref_size, ref_sha1);\n \t\t\tfree(ref);\n \t\t} else {\n-\t\t\tresult = pv4_encode_tree(src, &buf_size, NULL, 0, NULL);\n+\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n+\t\t\t\t\t\t NULL, 0, NULL);\n \t\t}\n \t\tbreak;\n \tdefault:\n@@ -987,7 +999,7 @@ static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \t\twarning(\"can't convert %s object %s\",\n \t\t\ttypename(type), sha1_to_hex(obj->sha1));\n \t\t/* fall back to copy the object in its original form */\n-\t\treturn copy_object_data(f, p, obj->offset);\n+\t\treturn copy_object_data(v4, f, p, obj->offset);\n \t}\n \n \t/* Use bit 3 to indicate a special type encoding */\n@@ -1041,7 +1053,7 @@ static struct packed_git *open_pack(const char *path)\n \treturn p;\n }\n \n-static void process_one_pack(char *src_pack, char *dst_pack)\n+static void process_one_pack(struct packv4_tables *v4, char *src_pack, char *dst_pack)\n {\n \tstruct packed_git *p;\n \tstruct sha1file *f;\n@@ -1061,26 +1073,26 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \tobjs = get_packed_object_list(p);\n \tp_objs = sort_objs_by_offset(objs, nr_objects);\n \n-\tcreate_pack_dictionaries(p, p_objs);\n-\tsort_dict_entries_by_hits(commit_ident_table);\n-\tsort_dict_entries_by_hits(tree_path_table);\n+\tcreate_pack_dictionaries(v4, p, p_objs);\n+\tsort_dict_entries_by_hits(v4->commit_ident_table);\n+\tsort_dict_entries_by_hits(v4->tree_path_table);\n \n \tpackname = normalize_pack_name(dst_pack);\n \tf = packv4_open(packname);\n \tif (!f)\n \t\tdie(\"unable to open destination pack\");\n \twritten += packv4_write_header(f, nr_objects);\n-\twritten += packv4_write_tables(f, nr_objects, objs);\n+\twritten += packv4_write_tables(f, v4);\n \n \t/* Let's write objects out, updating the object index list in place */\n \tprogress_state = start_progress(\"Writing objects\", nr_objects);\n-\tall_objs = objs;\n-\tall_objs_nr = nr_objects;\n+\tv4->all_objs = objs;\n+\tv4->all_objs_nr = nr_objects;\n \tfor (i = 0; i < nr_objects; i++) {\n \t\toff_t obj_pos = written;\n \t\tstruct pack_idx_entry *obj = p_objs[i];\n \t\tcrc32_begin(f);\n-\t\twritten += packv4_write_object(f, p, obj);\n+\t\twritten += packv4_write_object(v4, f, p, obj);\n \t\tobj->offset = obj_pos;\n \t\tobj->crc32 = crc32_end(f);\n \t\tdisplay_progress(progress_state, i+1);\n@@ -1114,6 +1126,7 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n \n int main(int argc, char *argv[])\n {\n+\tstruct packv4_tables v4;\n \tchar *src_pack, *dst_pack;\n \n \tif (argc == 3) {\n@@ -1131,8 +1144,8 @@ int main(int argc, char *argv[])\n \tgit_config(git_pack_config, NULL);\n \tif (!pack_compression_seen && core_compression_seen)\n \t\tpack_compression_level = core_compression_level;\n-\tprocess_one_pack(src_pack, dst_pack);\n+\tprocess_one_pack(&v4, src_pack, dst_pack);\n \tif (0)\n-\t\tdict_dump();\n+\t\tdict_dump(&v4);\n \treturn 0;\n }\ndiff --git a/packv4-create.h b/packv4-create.h\nnew file mode 100644\nindex 0000000..0c8c77b\n--- /dev/null\n+++ b/packv4-create.h\n@@ -0,0 +1,11 @@\n+#ifndef PACKV4_CREATE_H\n+#define PACKV4_CREATE_H\n+\n+struct packv4_tables {\n+\tstruct pack_idx_entry *all_objs;\n+\tunsigned all_objs_nr;\n+\tstruct dict_table *commit_ident_table;\n+\tstruct dict_table *tree_path_table;\n+};\n+\n+#endif\n-- \n1.8.2.83.gc99314b\n"},{"id":"227129","messageId":"1378652660-6731-4-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 03/11] pack v4: move packv4-create.c to libgit.a","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:12Z","receivedAt":"2013-09-08T15:04:12Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"git-packv4-create now becomes test-packv4. Code that will not be used\nby pack-objects.c is moved to test-packv4.c. It may be removed when\nthe code transition to pack-objects completes.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n Makefile            |   4 +-\n packv4-create.c     | 491 +---------------------------------------------------\n packv4-create.h     |  39 +++++\n test-packv4.c (new) | 476 ++++++++++++++++++++++++++++++++++++++++++++++++++\n 4 files changed, 525 insertions(+), 485 deletions(-)\n create mode 100644 test-packv4.c\n\ndiff --git a/Makefile b/Makefile\nindex 22fc276..af2e3e3 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -550,7 +550,6 @@ PROGRAM_OBJS += shell.o\n PROGRAM_OBJS += show-index.o\n PROGRAM_OBJS += upload-pack.o\n PROGRAM_OBJS += remote-testsvn.o\n-PROGRAM_OBJS += packv4-create.o\n \n # Binary suffix, set to .exe for Windows builds\n X =\n@@ -568,6 +567,7 @@ TEST_PROGRAMS_NEED_X += test-line-buffer\n TEST_PROGRAMS_NEED_X += test-match-trees\n TEST_PROGRAMS_NEED_X += test-mergesort\n TEST_PROGRAMS_NEED_X += test-mktemp\n+TEST_PROGRAMS_NEED_X += test-packv4\n TEST_PROGRAMS_NEED_X += test-parse-options\n TEST_PROGRAMS_NEED_X += test-path-utils\n TEST_PROGRAMS_NEED_X += test-prio-queue\n@@ -702,6 +702,7 @@ LIB_H += notes.h\n LIB_H += object.h\n LIB_H += pack-revindex.h\n LIB_H += pack.h\n+LIB_H += packv4-create.h\n LIB_H += packv4-parse.h\n LIB_H += parse-options.h\n LIB_H += patch-ids.h\n@@ -839,6 +840,7 @@ LIB_OBJS += object.o\n LIB_OBJS += pack-check.o\n LIB_OBJS += pack-revindex.o\n LIB_OBJS += pack-write.o\n+LIB_OBJS += packv4-create.o\n LIB_OBJS += packv4-parse.o\n LIB_OBJS += pager.o\n LIB_OBJS += parse-options.o\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 920a0b4..cdf82c0 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -18,9 +18,9 @@\n #include \"packv4-create.h\"\n \n \n-static int pack_compression_seen;\n-static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-static int min_tree_copy = 1;\n+int pack_compression_seen;\n+int pack_compression_level = Z_DEFAULT_COMPRESSION;\n+int min_tree_copy = 1;\n \n struct data_entry {\n \tunsigned offset;\n@@ -28,17 +28,6 @@ struct data_entry {\n \tunsigned hits;\n };\n \n-struct dict_table {\n-\tunsigned char *data;\n-\tunsigned cur_offset;\n-\tunsigned size;\n-\tstruct data_entry *entry;\n-\tunsigned nb_entries;\n-\tunsigned max_entries;\n-\tunsigned *hash;\n-\tunsigned hash_size;\n-};\n-\n struct dict_table *create_dict_table(void)\n {\n \treturn xcalloc(sizeof(struct dict_table), 1);\n@@ -139,7 +128,7 @@ static int cmp_dict_entries(const void *a_, const void *b_)\n \treturn diff;\n }\n \n-static void sort_dict_entries_by_hits(struct dict_table *t)\n+void sort_dict_entries_by_hits(struct dict_table *t)\n {\n \tqsort(t->entry, t->nb_entries, sizeof(*t->entry), cmp_dict_entries);\n \tt->hash_size = (t->nb_entries * 4 / 3) / 2;\n@@ -208,7 +197,7 @@ int add_commit_dict_entries(struct dict_table *commit_ident_table,\n \treturn 0;\n }\n \n-static int add_tree_dict_entries(struct dict_table *tree_path_table,\n+int add_tree_dict_entries(struct dict_table *tree_path_table,\n \t\t\t\t void *buf, unsigned long size)\n {\n \tstruct tree_desc desc;\n@@ -224,7 +213,7 @@ static int add_tree_dict_entries(struct dict_table *tree_path_table,\n \treturn 0;\n }\n \n-void dump_dict_table(struct dict_table *t)\n+static void dump_dict_table(struct dict_table *t)\n {\n \tint i;\n \n@@ -241,7 +230,7 @@ void dump_dict_table(struct dict_table *t)\n \t}\n }\n \n-static void dict_dump(struct packv4_tables *v4)\n+void dict_dump(struct packv4_tables *v4)\n {\n \tdump_dict_table(v4->commit_ident_table);\n \tdump_dict_table(v4->tree_path_table);\n@@ -611,103 +600,6 @@ void *pv4_encode_tree(const struct packv4_tables *v4,\n \treturn buffer;\n }\n \n-static struct pack_idx_entry *get_packed_object_list(struct packed_git *p)\n-{\n-\tunsigned i, nr_objects = p->num_objects;\n-\tstruct pack_idx_entry *objects;\n-\n-\tobjects = xmalloc((nr_objects + 1) * sizeof(*objects));\n-\tobjects[nr_objects].offset = p->pack_size - 20;\n-\tfor (i = 0; i < nr_objects; i++) {\n-\t\thashcpy(objects[i].sha1, nth_packed_object_sha1(p, i));\n-\t\tobjects[i].offset = nth_packed_object_offset(p, i);\n-\t}\n-\n-\treturn objects;\n-}\n-\n-static int sort_by_offset(const void *e1, const void *e2)\n-{\n-\tconst struct pack_idx_entry * const *entry1 = e1;\n-\tconst struct pack_idx_entry * const *entry2 = e2;\n-\tif ((*entry1)->offset < (*entry2)->offset)\n-\t\treturn -1;\n-\tif ((*entry1)->offset > (*entry2)->offset)\n-\t\treturn 1;\n-\treturn 0;\n-}\n-\n-static struct pack_idx_entry **sort_objs_by_offset(struct pack_idx_entry *list,\n-\t\t\t\t\t\t    unsigned nr_objects)\n-{\n-\tunsigned i;\n-\tstruct pack_idx_entry **sorted;\n-\n-\tsorted = xmalloc((nr_objects + 1) * sizeof(*sorted));\n-\tfor (i = 0; i < nr_objects + 1; i++)\n-\t\tsorted[i] = &list[i];\n-\tqsort(sorted, nr_objects + 1, sizeof(*sorted), sort_by_offset);\n-\n-\treturn sorted;\n-}\n-\n-static int create_pack_dictionaries(struct packv4_tables *v4,\n-\t\t\t\t    struct packed_git *p,\n-\t\t\t\t    struct pack_idx_entry **obj_list)\n-{\n-\tstruct progress *progress_state;\n-\tunsigned int i;\n-\n-\tv4->commit_ident_table = create_dict_table();\n-\tv4->tree_path_table = create_dict_table();\n-\n-\tprogress_state = start_progress(\"Scanning objects\", p->num_objects);\n-\tfor (i = 0; i < p->num_objects; i++) {\n-\t\tstruct pack_idx_entry *obj = obj_list[i];\n-\t\tvoid *data;\n-\t\tenum object_type type;\n-\t\tunsigned long size;\n-\t\tstruct object_info oi = {};\n-\t\tint (*add_dict_entries)(struct dict_table *, void *, unsigned long);\n-\t\tstruct dict_table *dict;\n-\n-\t\tdisplay_progress(progress_state, i+1);\n-\n-\t\toi.typep = &type;\n-\t\toi.sizep = &size;\n-\t\tif (packed_object_info(p, obj->offset, &oi) < 0)\n-\t\t\tdie(\"cannot get type of %s from %s\",\n-\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\n-\t\tswitch (type) {\n-\t\tcase OBJ_COMMIT:\n-\t\t\tadd_dict_entries = add_commit_dict_entries;\n-\t\t\tdict = v4->commit_ident_table;\n-\t\t\tbreak;\n-\t\tcase OBJ_TREE:\n-\t\t\tadd_dict_entries = add_tree_dict_entries;\n-\t\t\tdict = v4->tree_path_table;\n-\t\t\tbreak;\n-\t\tdefault:\n-\t\t\tcontinue;\n-\t\t}\n-\t\tdata = unpack_entry(p, obj->offset, &type, &size);\n-\t\tif (!data)\n-\t\t\tdie(\"cannot unpack %s from %s\",\n-\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\t\tif (check_sha1_signature(obj->sha1, data, size, typename(type)))\n-\t\t\tdie(\"packed %s from %s is corrupt\",\n-\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\t\tif (add_dict_entries(dict, data, size) < 0)\n-\t\t\tdie(\"can't process %s object %s\",\n-\t\t\t\ttypename(type), sha1_to_hex(obj->sha1));\n-\t\tfree(data);\n-\t}\n-\n-\tstop_progress(&progress_state);\n-\treturn 0;\n-}\n-\n static unsigned long write_dict_table(struct sha1file *f, struct dict_table *t)\n {\n \tunsigned char buffer[1024];\n@@ -757,28 +649,6 @@ static unsigned long write_dict_table(struct sha1file *f, struct dict_table *t)\n \treturn hdrlen + datalen;\n }\n \n-static struct sha1file * packv4_open(char *path)\n-{\n-\tint fd;\n-\n-\tfd = open(path, O_CREAT|O_EXCL|O_WRONLY, 0600);\n-\tif (fd < 0)\n-\t\tdie_errno(\"unable to create '%s'\", path);\n-\treturn sha1fd(fd, path);\n-}\n-\n-static unsigned int packv4_write_header(struct sha1file *f, unsigned nr_objects)\n-{\n-\tstruct pack_header hdr;\n-\n-\thdr.hdr_signature = htonl(PACK_SIGNATURE);\n-\thdr.hdr_version = htonl(4);\n-\thdr.hdr_entries = htonl(nr_objects);\n-\tsha1write(f, &hdr, sizeof(hdr));\n-\n-\treturn sizeof(hdr);\n-}\n-\n unsigned long packv4_write_tables(struct sha1file *f,\n \t\t\t\t  const struct packv4_tables *v4)\n {\n@@ -802,350 +672,3 @@ unsigned long packv4_write_tables(struct sha1file *f,\n \n \treturn written;\n }\n-\n-static int write_object_header(struct sha1file *f, enum object_type type, unsigned long size)\n-{\n-\tunsigned char buf[16];\n-\tuint64_t val;\n-\tint len;\n-\n-\t/*\n-\t * We really have only one kind of delta object.\n-\t */\n-\tif (type == OBJ_OFS_DELTA)\n-\t\ttype = OBJ_REF_DELTA;\n-\n-\t/*\n-\t * We allocate 4 bits in the LSB for the object type which should\n-\t * be good for quite a while, given that we effectively encodes\n-\t * only 5 object types: commit, tree, blob, delta, tag.\n-\t */\n-\tval = size;\n-\tif (MSB(val, 4))\n-\t\tdie(\"fixme: the code doesn't currently cope with big sizes\");\n-\tval <<= 4;\n-\tval |= type;\n-\tlen = encode_varint(val, buf);\n-\tsha1write(f, buf, len);\n-\treturn len;\n-}\n-\n-static unsigned long copy_object_data(struct packv4_tables *v4,\n-\t\t\t\t      struct sha1file *f, struct packed_git *p,\n-\t\t\t\t      off_t offset)\n-{\n-\tstruct pack_window *w_curs = NULL;\n-\tstruct revindex_entry *revidx;\n-\tenum object_type type;\n-\tunsigned long avail, size, datalen, written;\n-\tint hdrlen, reflen, idx_nr;\n-\tunsigned char *src, buf[24];\n-\n-\trevidx = find_pack_revindex(p, offset);\n-\tidx_nr = revidx->nr;\n-\tdatalen = revidx[1].offset - offset;\n-\n-\tsrc = use_pack(p, &w_curs, offset, &avail);\n-\thdrlen = unpack_object_header_buffer(src, avail, &type, &size);\n-\n-\twritten = write_object_header(f, type, size);\n-\n-\tif (type == OBJ_OFS_DELTA) {\n-\t\tconst unsigned char *cp = src + hdrlen;\n-\t\toff_t base_offset = decode_varint(&cp);\n-\t\thdrlen = cp - src;\n-\t\tbase_offset = offset - base_offset;\n-\t\tif (base_offset <= 0 || base_offset >= offset)\n-\t\t\tdie(\"delta offset out of bound\");\n-\t\trevidx = find_pack_revindex(p, base_offset);\n-\t\treflen = encode_sha1ref(v4,\n-\t\t\t\t\tnth_packed_object_sha1(p, revidx->nr),\n-\t\t\t\t\tbuf);\n-\t\tsha1write(f, buf, reflen);\n-\t\twritten += reflen;\n-\t} else if (type == OBJ_REF_DELTA) {\n-\t\treflen = encode_sha1ref(v4, src + hdrlen, buf);\n-\t\thdrlen += 20;\n-\t\tsha1write(f, buf, reflen);\n-\t\twritten += reflen;\n-\t}\n-\n-\tif (p->index_version > 1 &&\n-\t    check_pack_crc(p, &w_curs, offset, datalen, idx_nr))\n-\t\tdie(\"bad CRC for object at offset %\"PRIuMAX\" in %s\",\n-\t\t    (uintmax_t)offset, p->pack_name);\n-\n-\toffset += hdrlen;\n-\tdatalen -= hdrlen;\n-\n-\twhile (datalen) {\n-\t\tsrc = use_pack(p, &w_curs, offset, &avail);\n-\t\tif (avail > datalen)\n-\t\t\tavail = datalen;\n-\t\tsha1write(f, src, avail);\n-\t\twritten += avail;\n-\t\toffset += avail;\n-\t\tdatalen -= avail;\n-\t}\n-\tunuse_pack(&w_curs);\n-\n-\treturn written;\n-}\n-\n-static unsigned char *get_delta_base(struct packed_git *p, off_t offset,\n-\t\t\t\t     unsigned char *sha1_buf)\n-{\n-\tstruct pack_window *w_curs = NULL;\n-\tenum object_type type;\n-\tunsigned long avail, size;\n-\tint hdrlen;\n-\tunsigned char *src;\n-\tconst unsigned char *base_sha1 = NULL; ;\n-\n-\tsrc = use_pack(p, &w_curs, offset, &avail);\n-\thdrlen = unpack_object_header_buffer(src, avail, &type, &size);\n-\n-\tif (type == OBJ_OFS_DELTA) {\n-\t\tconst unsigned char *cp = src + hdrlen;\n-\t\toff_t base_offset = decode_varint(&cp);\n-\t\tbase_offset = offset - base_offset;\n-\t\tif (base_offset <= 0 || base_offset >= offset) {\n-\t\t\terror(\"delta offset out of bound\");\n-\t\t} else {\n-\t\t\tstruct revindex_entry *revidx;\n-\t\t\trevidx = find_pack_revindex(p, base_offset);\n-\t\t\tbase_sha1 = nth_packed_object_sha1(p, revidx->nr);\n-\t\t}\n-\t} else if (type == OBJ_REF_DELTA) {\n-\t\tbase_sha1 = src + hdrlen;\n-\t} else\n-\t\terror(\"expected to get a delta but got a %s\", typename(type));\n-\n-\tunuse_pack(&w_curs);\n-\n-\tif (!base_sha1)\n-\t\treturn NULL;\n-\thashcpy(sha1_buf, base_sha1);\n-\treturn sha1_buf;\n-}\n-\n-static off_t packv4_write_object(struct packv4_tables *v4,\n-\t\t\t\t struct sha1file *f, struct packed_git *p,\n-\t\t\t\t struct pack_idx_entry *obj)\n-{\n-\tvoid *src, *result;\n-\tstruct object_info oi = {};\n-\tenum object_type type, packed_type;\n-\tunsigned long obj_size, buf_size;\n-\tunsigned int hdrlen;\n-\n-\toi.typep = &type;\n-\toi.sizep = &obj_size;\n-\tpacked_type = packed_object_info(p, obj->offset, &oi);\n-\tif (packed_type < 0)\n-\t\tdie(\"cannot get type of %s from %s\",\n-\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\n-\t/* Some objects are copied without decompression */\n-\tswitch (type) {\n-\tcase OBJ_COMMIT:\n-\tcase OBJ_TREE:\n-\t\tbreak;\n-\tdefault:\n-\t\treturn copy_object_data(v4, f, p, obj->offset);\n-\t}\n-\n-\t/* The rest is converted into their new format */\n-\tsrc = unpack_entry(p, obj->offset, &type, &buf_size);\n-\tif (!src || obj_size != buf_size)\n-\t\tdie(\"cannot unpack %s from %s\",\n-\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\tif (check_sha1_signature(obj->sha1, src, buf_size, typename(type)))\n-\t\tdie(\"packed %s from %s is corrupt\",\n-\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\n-\tswitch (type) {\n-\tcase OBJ_COMMIT:\n-\t\tresult = pv4_encode_commit(v4, src, &buf_size);\n-\t\tbreak;\n-\tcase OBJ_TREE:\n-\t\tif (packed_type != OBJ_TREE) {\n-\t\t\tunsigned char sha1_buf[20], *ref_sha1;\n-\t\t\tvoid *ref;\n-\t\t\tenum object_type ref_type;\n-\t\t\tunsigned long ref_size;\n-\n-\t\t\tref_sha1 = get_delta_base(p, obj->offset, sha1_buf);\n-\t\t\tif (!ref_sha1)\n-\t\t\t\tdie(\"unable to get delta base sha1 for %s\",\n-\t\t\t\t\t\tsha1_to_hex(obj->sha1));\n-\t\t\tref = read_sha1_file(ref_sha1, &ref_type, &ref_size);\n-\t\t\tif (!ref || ref_type != OBJ_TREE)\n-\t\t\t\tdie(\"cannot obtain delta base for %s\",\n-\t\t\t\t\t\tsha1_to_hex(obj->sha1));\n-\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n-\t\t\t\t\t\t ref, ref_size, ref_sha1);\n-\t\t\tfree(ref);\n-\t\t} else {\n-\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n-\t\t\t\t\t\t NULL, 0, NULL);\n-\t\t}\n-\t\tbreak;\n-\tdefault:\n-\t\tdie(\"unexpected object type %d\", type);\n-\t}\n-\tfree(src);\n-\tif (!result) {\n-\t\twarning(\"can't convert %s object %s\",\n-\t\t\ttypename(type), sha1_to_hex(obj->sha1));\n-\t\t/* fall back to copy the object in its original form */\n-\t\treturn copy_object_data(v4, f, p, obj->offset);\n-\t}\n-\n-\t/* Use bit 3 to indicate a special type encoding */\n-\ttype += 8;\n-\thdrlen = write_object_header(f, type, obj_size);\n-\tsha1write(f, result, buf_size);\n-\tfree(result);\n-\treturn hdrlen + buf_size;\n-}\n-\n-static char *normalize_pack_name(const char *path)\n-{\n-\tchar buf[PATH_MAX];\n-\tint len;\n-\n-\tlen = strlcpy(buf, path, PATH_MAX);\n-\tif (len >= PATH_MAX - 6)\n-\t\tdie(\"name too long: %s\", path);\n-\n-\t/*\n-\t * In addition to \"foo.idx\" we accept \"foo.pack\" and \"foo\";\n-\t * normalize these forms to \"foo.pack\".\n-\t */\n-\tif (has_extension(buf, \".idx\")) {\n-\t\tstrcpy(buf + len - 4, \".pack\");\n-\t\tlen++;\n-\t} else if (!has_extension(buf, \".pack\")) {\n-\t\tstrcpy(buf + len, \".pack\");\n-\t\tlen += 5;\n-\t}\n-\n-\treturn xstrdup(buf);\n-}\n-\n-static struct packed_git *open_pack(const char *path)\n-{\n-\tchar *packname = normalize_pack_name(path);\n-\tint len = strlen(packname);\n-\tstruct packed_git *p;\n-\n-\tstrcpy(packname + len - 5, \".idx\");\n-\tp = add_packed_git(packname, len - 1, 1);\n-\tif (!p)\n-\t\tdie(\"packfile %s not found.\", packname);\n-\n-\tinstall_packed_git(p);\n-\tif (open_pack_index(p))\n-\t\tdie(\"packfile %s index not opened\", p->pack_name);\n-\n-\tfree(packname);\n-\treturn p;\n-}\n-\n-static void process_one_pack(struct packv4_tables *v4, char *src_pack, char *dst_pack)\n-{\n-\tstruct packed_git *p;\n-\tstruct sha1file *f;\n-\tstruct pack_idx_entry *objs, **p_objs;\n-\tstruct pack_idx_option idx_opts;\n-\tunsigned i, nr_objects;\n-\toff_t written = 0;\n-\tchar *packname;\n-\tunsigned char pack_sha1[20];\n-\tstruct progress *progress_state;\n-\n-\tp = open_pack(src_pack);\n-\tif (!p)\n-\t\tdie(\"unable to open source pack\");\n-\n-\tnr_objects = p->num_objects;\n-\tobjs = get_packed_object_list(p);\n-\tp_objs = sort_objs_by_offset(objs, nr_objects);\n-\n-\tcreate_pack_dictionaries(v4, p, p_objs);\n-\tsort_dict_entries_by_hits(v4->commit_ident_table);\n-\tsort_dict_entries_by_hits(v4->tree_path_table);\n-\n-\tpackname = normalize_pack_name(dst_pack);\n-\tf = packv4_open(packname);\n-\tif (!f)\n-\t\tdie(\"unable to open destination pack\");\n-\twritten += packv4_write_header(f, nr_objects);\n-\twritten += packv4_write_tables(f, v4);\n-\n-\t/* Let's write objects out, updating the object index list in place */\n-\tprogress_state = start_progress(\"Writing objects\", nr_objects);\n-\tv4->all_objs = objs;\n-\tv4->all_objs_nr = nr_objects;\n-\tfor (i = 0; i < nr_objects; i++) {\n-\t\toff_t obj_pos = written;\n-\t\tstruct pack_idx_entry *obj = p_objs[i];\n-\t\tcrc32_begin(f);\n-\t\twritten += packv4_write_object(v4, f, p, obj);\n-\t\tobj->offset = obj_pos;\n-\t\tobj->crc32 = crc32_end(f);\n-\t\tdisplay_progress(progress_state, i+1);\n-\t}\n-\tstop_progress(&progress_state);\n-\n-\tsha1close(f, pack_sha1, CSUM_CLOSE | CSUM_FSYNC);\n-\n-\treset_pack_idx_option(&idx_opts);\n-\tidx_opts.version = 3;\n-\tstrcpy(packname + strlen(packname) - 5, \".idx\");\n-\twrite_idx_file(packname, p_objs, nr_objects, &idx_opts, pack_sha1);\n-\n-\tfree(packname);\n-}\n-\n-static int git_pack_config(const char *k, const char *v, void *cb)\n-{\n-\tif (!strcmp(k, \"pack.compression\")) {\n-\t\tint level = git_config_int(k, v);\n-\t\tif (level == -1)\n-\t\t\tlevel = Z_DEFAULT_COMPRESSION;\n-\t\telse if (level < 0 || level > Z_BEST_COMPRESSION)\n-\t\t\tdie(\"bad pack compression level %d\", level);\n-\t\tpack_compression_level = level;\n-\t\tpack_compression_seen = 1;\n-\t\treturn 0;\n-\t}\n-\treturn git_default_config(k, v, cb);\n-}\n-\n-int main(int argc, char *argv[])\n-{\n-\tstruct packv4_tables v4;\n-\tchar *src_pack, *dst_pack;\n-\n-\tif (argc == 3) {\n-\t\tsrc_pack = argv[1];\n-\t\tdst_pack = argv[2];\n-\t} else if (argc == 4 && !prefixcmp(argv[1], \"--min-tree-copy=\")) {\n-\t\tmin_tree_copy = atoi(argv[1] + strlen(\"--min-tree-copy=\"));\n-\t\tsrc_pack = argv[2];\n-\t\tdst_pack = argv[3];\n-\t} else {\n-\t\tfprintf(stderr, \"Usage: %s [--min-tree-copy=<n>] <src_packfile> <dst_packfile>\\n\", argv[0]);\n-\t\texit(1);\n-\t}\n-\n-\tgit_config(git_pack_config, NULL);\n-\tif (!pack_compression_seen && core_compression_seen)\n-\t\tpack_compression_level = core_compression_level;\n-\tprocess_one_pack(&v4, src_pack, dst_pack);\n-\tif (0)\n-\t\tdict_dump(&v4);\n-\treturn 0;\n-}\ndiff --git a/packv4-create.h b/packv4-create.h\nindex 0c8c77b..c1f32fd 100644\n--- a/packv4-create.h\n+++ b/packv4-create.h\n@@ -8,4 +8,43 @@ struct packv4_tables {\n \tstruct dict_table *tree_path_table;\n };\n \n+struct dict_table {\n+\tunsigned char *data;\n+\tunsigned cur_offset;\n+\tunsigned size;\n+\tstruct data_entry *entry;\n+\tunsigned nb_entries;\n+\tunsigned max_entries;\n+\tunsigned *hash;\n+\tunsigned hash_size;\n+};\n+\n+\n+struct sha1file;\n+\n+struct dict_table *create_dict_table(void);\n+int dict_add_entry(struct dict_table *t, int val, const char *str, int str_len);\n+void destroy_dict_table(struct dict_table *t);\n+void dict_dump(struct packv4_tables *v4);\n+\n+int add_commit_dict_entries(struct dict_table *commit_ident_table,\n+\t\t\t    void *buf, unsigned long size);\n+int add_tree_dict_entries(struct dict_table *tree_path_table,\n+\t\t\t  void *buf, unsigned long size);\n+void sort_dict_entries_by_hits(struct dict_table *t);\n+\n+int encode_sha1ref(const struct packv4_tables *v4,\n+\t\t   const unsigned char *sha1, unsigned char *buf);\n+unsigned long packv4_write_tables(struct sha1file *f,\n+\t\t\t\t  const struct packv4_tables *v4);\n+void *pv4_encode_commit(const struct packv4_tables *v4,\n+\t\t\tvoid *buffer, unsigned long *sizep);\n+void *pv4_encode_tree(const struct packv4_tables *v4,\n+\t\t      void *_buffer, unsigned long *sizep,\n+\t\t      void *delta, unsigned long delta_size,\n+\t\t      const unsigned char *delta_sha1);\n+\n+void process_one_pack(struct packv4_tables *v4,\n+\t\t      char *src_pack, char *dst_pack);\n+\n #endif\ndiff --git a/test-packv4.c b/test-packv4.c\nnew file mode 100644\nindex 0000000..3b0d7a2\n--- /dev/null\n+++ b/test-packv4.c\n@@ -0,0 +1,476 @@\n+#include \"cache.h\"\n+#include \"pack.h\"\n+#include \"pack-revindex.h\"\n+#include \"progress.h\"\n+#include \"varint.h\"\n+#include \"packv4-create.h\"\n+\n+extern int pack_compression_seen;\n+extern int pack_compression_level;\n+extern int min_tree_copy;\n+\n+static struct pack_idx_entry *get_packed_object_list(struct packed_git *p)\n+{\n+\tunsigned i, nr_objects = p->num_objects;\n+\tstruct pack_idx_entry *objects;\n+\n+\tobjects = xmalloc((nr_objects + 1) * sizeof(*objects));\n+\tobjects[nr_objects].offset = p->pack_size - 20;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\thashcpy(objects[i].sha1, nth_packed_object_sha1(p, i));\n+\t\tobjects[i].offset = nth_packed_object_offset(p, i);\n+\t}\n+\n+\treturn objects;\n+}\n+\n+static int sort_by_offset(const void *e1, const void *e2)\n+{\n+\tconst struct pack_idx_entry * const *entry1 = e1;\n+\tconst struct pack_idx_entry * const *entry2 = e2;\n+\tif ((*entry1)->offset < (*entry2)->offset)\n+\t\treturn -1;\n+\tif ((*entry1)->offset > (*entry2)->offset)\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n+static struct pack_idx_entry **sort_objs_by_offset(struct pack_idx_entry *list,\n+\t\t\t\t\t\t    unsigned nr_objects)\n+{\n+\tunsigned i;\n+\tstruct pack_idx_entry **sorted;\n+\n+\tsorted = xmalloc((nr_objects + 1) * sizeof(*sorted));\n+\tfor (i = 0; i < nr_objects + 1; i++)\n+\t\tsorted[i] = &list[i];\n+\tqsort(sorted, nr_objects + 1, sizeof(*sorted), sort_by_offset);\n+\n+\treturn sorted;\n+}\n+\n+static int create_pack_dictionaries(struct packv4_tables *v4,\n+\t\t\t\t    struct packed_git *p,\n+\t\t\t\t    struct pack_idx_entry **obj_list)\n+{\n+\tstruct progress *progress_state;\n+\tunsigned int i;\n+\n+\tv4->commit_ident_table = create_dict_table();\n+\tv4->tree_path_table = create_dict_table();\n+\n+\tprogress_state = start_progress(\"Scanning objects\", p->num_objects);\n+\tfor (i = 0; i < p->num_objects; i++) {\n+\t\tstruct pack_idx_entry *obj = obj_list[i];\n+\t\tvoid *data;\n+\t\tenum object_type type;\n+\t\tunsigned long size;\n+\t\tstruct object_info oi = {};\n+\t\tint (*add_dict_entries)(struct dict_table *, void *, unsigned long);\n+\t\tstruct dict_table *dict;\n+\n+\t\tdisplay_progress(progress_state, i+1);\n+\n+\t\toi.typep = &type;\n+\t\toi.sizep = &size;\n+\t\tif (packed_object_info(p, obj->offset, &oi) < 0)\n+\t\t\tdie(\"cannot get type of %s from %s\",\n+\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\n+\t\tswitch (type) {\n+\t\tcase OBJ_COMMIT:\n+\t\t\tadd_dict_entries = add_commit_dict_entries;\n+\t\t\tdict = v4->commit_ident_table;\n+\t\t\tbreak;\n+\t\tcase OBJ_TREE:\n+\t\t\tadd_dict_entries = add_tree_dict_entries;\n+\t\t\tdict = v4->tree_path_table;\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tcontinue;\n+\t\t}\n+\t\tdata = unpack_entry(p, obj->offset, &type, &size);\n+\t\tif (!data)\n+\t\t\tdie(\"cannot unpack %s from %s\",\n+\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\t\tif (check_sha1_signature(obj->sha1, data, size, typename(type)))\n+\t\t\tdie(\"packed %s from %s is corrupt\",\n+\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\t\tif (add_dict_entries(dict, data, size) < 0)\n+\t\t\tdie(\"can't process %s object %s\",\n+\t\t\t\ttypename(type), sha1_to_hex(obj->sha1));\n+\t\tfree(data);\n+\t}\n+\n+\tstop_progress(&progress_state);\n+\treturn 0;\n+}\n+\n+static struct sha1file * packv4_open(char *path)\n+{\n+\tint fd;\n+\n+\tfd = open(path, O_CREAT|O_EXCL|O_WRONLY, 0600);\n+\tif (fd < 0)\n+\t\tdie_errno(\"unable to create '%s'\", path);\n+\treturn sha1fd(fd, path);\n+}\n+\n+static unsigned int packv4_write_header(struct sha1file *f, unsigned nr_objects)\n+{\n+\tstruct pack_header hdr;\n+\n+\thdr.hdr_signature = htonl(PACK_SIGNATURE);\n+\thdr.hdr_version = htonl(4);\n+\thdr.hdr_entries = htonl(nr_objects);\n+\tsha1write(f, &hdr, sizeof(hdr));\n+\n+\treturn sizeof(hdr);\n+}\n+\n+static int write_object_header(struct sha1file *f, enum object_type type, unsigned long size)\n+{\n+\tunsigned char buf[16];\n+\tuint64_t val;\n+\tint len;\n+\n+\t/*\n+\t * We really have only one kind of delta object.\n+\t */\n+\tif (type == OBJ_OFS_DELTA)\n+\t\ttype = OBJ_REF_DELTA;\n+\n+\t/*\n+\t * We allocate 4 bits in the LSB for the object type which should\n+\t * be good for quite a while, given that we effectively encodes\n+\t * only 5 object types: commit, tree, blob, delta, tag.\n+\t */\n+\tval = size;\n+\tif (MSB(val, 4))\n+\t\tdie(\"fixme: the code doesn't currently cope with big sizes\");\n+\tval <<= 4;\n+\tval |= type;\n+\tlen = encode_varint(val, buf);\n+\tsha1write(f, buf, len);\n+\treturn len;\n+}\n+\n+static unsigned long copy_object_data(struct packv4_tables *v4,\n+\t\t\t\t      struct sha1file *f, struct packed_git *p,\n+\t\t\t\t      off_t offset)\n+{\n+\tstruct pack_window *w_curs = NULL;\n+\tstruct revindex_entry *revidx;\n+\tenum object_type type;\n+\tunsigned long avail, size, datalen, written;\n+\tint hdrlen, reflen, idx_nr;\n+\tunsigned char *src, buf[24];\n+\n+\trevidx = find_pack_revindex(p, offset);\n+\tidx_nr = revidx->nr;\n+\tdatalen = revidx[1].offset - offset;\n+\n+\tsrc = use_pack(p, &w_curs, offset, &avail);\n+\thdrlen = unpack_object_header_buffer(src, avail, &type, &size);\n+\n+\twritten = write_object_header(f, type, size);\n+\n+\tif (type == OBJ_OFS_DELTA) {\n+\t\tconst unsigned char *cp = src + hdrlen;\n+\t\toff_t base_offset = decode_varint(&cp);\n+\t\thdrlen = cp - src;\n+\t\tbase_offset = offset - base_offset;\n+\t\tif (base_offset <= 0 || base_offset >= offset)\n+\t\t\tdie(\"delta offset out of bound\");\n+\t\trevidx = find_pack_revindex(p, base_offset);\n+\t\treflen = encode_sha1ref(v4,\n+\t\t\t\t\tnth_packed_object_sha1(p, revidx->nr),\n+\t\t\t\t\tbuf);\n+\t\tsha1write(f, buf, reflen);\n+\t\twritten += reflen;\n+\t} else if (type == OBJ_REF_DELTA) {\n+\t\treflen = encode_sha1ref(v4, src + hdrlen, buf);\n+\t\thdrlen += 20;\n+\t\tsha1write(f, buf, reflen);\n+\t\twritten += reflen;\n+\t}\n+\n+\tif (p->index_version > 1 &&\n+\t    check_pack_crc(p, &w_curs, offset, datalen, idx_nr))\n+\t\tdie(\"bad CRC for object at offset %\"PRIuMAX\" in %s\",\n+\t\t    (uintmax_t)offset, p->pack_name);\n+\n+\toffset += hdrlen;\n+\tdatalen -= hdrlen;\n+\n+\twhile (datalen) {\n+\t\tsrc = use_pack(p, &w_curs, offset, &avail);\n+\t\tif (avail > datalen)\n+\t\t\tavail = datalen;\n+\t\tsha1write(f, src, avail);\n+\t\twritten += avail;\n+\t\toffset += avail;\n+\t\tdatalen -= avail;\n+\t}\n+\tunuse_pack(&w_curs);\n+\n+\treturn written;\n+}\n+\n+static unsigned char *get_delta_base(struct packed_git *p, off_t offset,\n+\t\t\t\t     unsigned char *sha1_buf)\n+{\n+\tstruct pack_window *w_curs = NULL;\n+\tenum object_type type;\n+\tunsigned long avail, size;\n+\tint hdrlen;\n+\tunsigned char *src;\n+\tconst unsigned char *base_sha1 = NULL; ;\n+\n+\tsrc = use_pack(p, &w_curs, offset, &avail);\n+\thdrlen = unpack_object_header_buffer(src, avail, &type, &size);\n+\n+\tif (type == OBJ_OFS_DELTA) {\n+\t\tconst unsigned char *cp = src + hdrlen;\n+\t\toff_t base_offset = decode_varint(&cp);\n+\t\tbase_offset = offset - base_offset;\n+\t\tif (base_offset <= 0 || base_offset >= offset) {\n+\t\t\terror(\"delta offset out of bound\");\n+\t\t} else {\n+\t\t\tstruct revindex_entry *revidx;\n+\t\t\trevidx = find_pack_revindex(p, base_offset);\n+\t\t\tbase_sha1 = nth_packed_object_sha1(p, revidx->nr);\n+\t\t}\n+\t} else if (type == OBJ_REF_DELTA) {\n+\t\tbase_sha1 = src + hdrlen;\n+\t} else\n+\t\terror(\"expected to get a delta but got a %s\", typename(type));\n+\n+\tunuse_pack(&w_curs);\n+\n+\tif (!base_sha1)\n+\t\treturn NULL;\n+\thashcpy(sha1_buf, base_sha1);\n+\treturn sha1_buf;\n+}\n+\n+static off_t packv4_write_object(struct packv4_tables *v4,\n+\t\t\t\t struct sha1file *f, struct packed_git *p,\n+\t\t\t\t struct pack_idx_entry *obj)\n+{\n+\tvoid *src, *result;\n+\tstruct object_info oi = {};\n+\tenum object_type type, packed_type;\n+\tunsigned long obj_size, buf_size;\n+\tunsigned int hdrlen;\n+\n+\toi.typep = &type;\n+\toi.sizep = &obj_size;\n+\tpacked_type = packed_object_info(p, obj->offset, &oi);\n+\tif (packed_type < 0)\n+\t\tdie(\"cannot get type of %s from %s\",\n+\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\n+\t/* Some objects are copied without decompression */\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TREE:\n+\t\tbreak;\n+\tdefault:\n+\t\treturn copy_object_data(v4, f, p, obj->offset);\n+\t}\n+\n+\t/* The rest is converted into their new format */\n+\tsrc = unpack_entry(p, obj->offset, &type, &buf_size);\n+\tif (!src || obj_size != buf_size)\n+\t\tdie(\"cannot unpack %s from %s\",\n+\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\tif (check_sha1_signature(obj->sha1, src, buf_size, typename(type)))\n+\t\tdie(\"packed %s from %s is corrupt\",\n+\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\t\tresult = pv4_encode_commit(v4, src, &buf_size);\n+\t\tbreak;\n+\tcase OBJ_TREE:\n+\t\tif (packed_type != OBJ_TREE) {\n+\t\t\tunsigned char sha1_buf[20], *ref_sha1;\n+\t\t\tvoid *ref;\n+\t\t\tenum object_type ref_type;\n+\t\t\tunsigned long ref_size;\n+\n+\t\t\tref_sha1 = get_delta_base(p, obj->offset, sha1_buf);\n+\t\t\tif (!ref_sha1)\n+\t\t\t\tdie(\"unable to get delta base sha1 for %s\",\n+\t\t\t\t\t\tsha1_to_hex(obj->sha1));\n+\t\t\tref = read_sha1_file(ref_sha1, &ref_type, &ref_size);\n+\t\t\tif (!ref || ref_type != OBJ_TREE)\n+\t\t\t\tdie(\"cannot obtain delta base for %s\",\n+\t\t\t\t\t\tsha1_to_hex(obj->sha1));\n+\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n+\t\t\t\t\t\t ref, ref_size, ref_sha1);\n+\t\t\tfree(ref);\n+\t\t} else {\n+\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n+\t\t\t\t\t\t NULL, 0, NULL);\n+\t\t}\n+\t\tbreak;\n+\tdefault:\n+\t\tdie(\"unexpected object type %d\", type);\n+\t}\n+\tfree(src);\n+\tif (!result) {\n+\t\twarning(\"can't convert %s object %s\",\n+\t\t\ttypename(type), sha1_to_hex(obj->sha1));\n+\t\t/* fall back to copy the object in its original form */\n+\t\treturn copy_object_data(v4, f, p, obj->offset);\n+\t}\n+\n+\t/* Use bit 3 to indicate a special type encoding */\n+\ttype += 8;\n+\thdrlen = write_object_header(f, type, obj_size);\n+\tsha1write(f, result, buf_size);\n+\tfree(result);\n+\treturn hdrlen + buf_size;\n+}\n+\n+static char *normalize_pack_name(const char *path)\n+{\n+\tchar buf[PATH_MAX];\n+\tint len;\n+\n+\tlen = strlcpy(buf, path, PATH_MAX);\n+\tif (len >= PATH_MAX - 6)\n+\t\tdie(\"name too long: %s\", path);\n+\n+\t/*\n+\t * In addition to \"foo.idx\" we accept \"foo.pack\" and \"foo\";\n+\t * normalize these forms to \"foo.pack\".\n+\t */\n+\tif (has_extension(buf, \".idx\")) {\n+\t\tstrcpy(buf + len - 4, \".pack\");\n+\t\tlen++;\n+\t} else if (!has_extension(buf, \".pack\")) {\n+\t\tstrcpy(buf + len, \".pack\");\n+\t\tlen += 5;\n+\t}\n+\n+\treturn xstrdup(buf);\n+}\n+\n+static struct packed_git *open_pack(const char *path)\n+{\n+\tchar *packname = normalize_pack_name(path);\n+\tint len = strlen(packname);\n+\tstruct packed_git *p;\n+\n+\tstrcpy(packname + len - 5, \".idx\");\n+\tp = add_packed_git(packname, len - 1, 1);\n+\tif (!p)\n+\t\tdie(\"packfile %s not found.\", packname);\n+\n+\tinstall_packed_git(p);\n+\tif (open_pack_index(p))\n+\t\tdie(\"packfile %s index not opened\", p->pack_name);\n+\n+\tfree(packname);\n+\treturn p;\n+}\n+\n+void process_one_pack(struct packv4_tables *v4, char *src_pack, char *dst_pack)\n+{\n+\tstruct packed_git *p;\n+\tstruct sha1file *f;\n+\tstruct pack_idx_entry *objs, **p_objs;\n+\tstruct pack_idx_option idx_opts;\n+\tunsigned i, nr_objects;\n+\toff_t written = 0;\n+\tchar *packname;\n+\tunsigned char pack_sha1[20];\n+\tstruct progress *progress_state;\n+\n+\tp = open_pack(src_pack);\n+\tif (!p)\n+\t\tdie(\"unable to open source pack\");\n+\n+\tnr_objects = p->num_objects;\n+\tobjs = get_packed_object_list(p);\n+\tp_objs = sort_objs_by_offset(objs, nr_objects);\n+\n+\tcreate_pack_dictionaries(v4, p, p_objs);\n+\tsort_dict_entries_by_hits(v4->commit_ident_table);\n+\tsort_dict_entries_by_hits(v4->tree_path_table);\n+\n+\tpackname = normalize_pack_name(dst_pack);\n+\tf = packv4_open(packname);\n+\tif (!f)\n+\t\tdie(\"unable to open destination pack\");\n+\twritten += packv4_write_header(f, nr_objects);\n+\twritten += packv4_write_tables(f, v4);\n+\n+\t/* Let's write objects out, updating the object index list in place */\n+\tprogress_state = start_progress(\"Writing objects\", nr_objects);\n+\tv4->all_objs = objs;\n+\tv4->all_objs_nr = nr_objects;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\toff_t obj_pos = written;\n+\t\tstruct pack_idx_entry *obj = p_objs[i];\n+\t\tcrc32_begin(f);\n+\t\twritten += packv4_write_object(v4, f, p, obj);\n+\t\tobj->offset = obj_pos;\n+\t\tobj->crc32 = crc32_end(f);\n+\t\tdisplay_progress(progress_state, i+1);\n+\t}\n+\tstop_progress(&progress_state);\n+\n+\tsha1close(f, pack_sha1, CSUM_CLOSE | CSUM_FSYNC);\n+\n+\treset_pack_idx_option(&idx_opts);\n+\tidx_opts.version = 3;\n+\tstrcpy(packname + strlen(packname) - 5, \".idx\");\n+\twrite_idx_file(packname, p_objs, nr_objects, &idx_opts, pack_sha1);\n+\n+\tfree(packname);\n+}\n+\n+static int git_pack_config(const char *k, const char *v, void *cb)\n+{\n+\tif (!strcmp(k, \"pack.compression\")) {\n+\t\tint level = git_config_int(k, v);\n+\t\tif (level == -1)\n+\t\t\tlevel = Z_DEFAULT_COMPRESSION;\n+\t\telse if (level < 0 || level > Z_BEST_COMPRESSION)\n+\t\t\tdie(\"bad pack compression level %d\", level);\n+\t\tpack_compression_level = level;\n+\t\tpack_compression_seen = 1;\n+\t\treturn 0;\n+\t}\n+\treturn git_default_config(k, v, cb);\n+}\n+\n+int main(int argc, char *argv[])\n+{\n+\tstruct packv4_tables v4;\n+\tchar *src_pack, *dst_pack;\n+\n+\tif (argc == 3) {\n+\t\tsrc_pack = argv[1];\n+\t\tdst_pack = argv[2];\n+\t} else if (argc == 4 && !prefixcmp(argv[1], \"--min-tree-copy=\")) {\n+\t\tmin_tree_copy = atoi(argv[1] + strlen(\"--min-tree-copy=\"));\n+\t\tsrc_pack = argv[2];\n+\t\tdst_pack = argv[3];\n+\t} else {\n+\t\tfprintf(stderr, \"Usage: %s [--min-tree-copy=<n>] <src_packfile> <dst_packfile>\\n\", argv[0]);\n+\t\texit(1);\n+\t}\n+\n+\tgit_config(git_pack_config, NULL);\n+\tif (!pack_compression_seen && core_compression_seen)\n+\t\tpack_compression_level = core_compression_level;\n+\tprocess_one_pack(&v4, src_pack, dst_pack);\n+\tif (0)\n+\t\tdict_dump(&v4);\n+\treturn 0;\n+}\n-- \n1.8.2.83.gc99314b\n"},{"id":"227130","messageId":"1378652660-6731-5-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 04/11] pack v4: add version argument to write_pack_header","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:13Z","receivedAt":"2013-09-08T15:04:13Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 2 +-\n bulk-checkin.c         | 2 +-\n pack-write.c           | 7 +++++--\n pack.h                 | 3 +--\n 4 files changed, 8 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex f069462..33faea8 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -735,7 +735,7 @@ static void write_pack_file(void)\n \t\telse\n \t\t\tf = create_tmp_packfile(&pack_tmp_name);\n \n-\t\toffset = write_pack_header(f, nr_remaining);\n+\t\toffset = write_pack_header(f, 2, nr_remaining);\n \t\tif (!offset)\n \t\t\tdie_errno(\"unable to write pack header\");\n \t\tnr_written = 0;\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6b0b6d4..9d8f0d0 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -176,7 +176,7 @@ static void prepare_to_stream(struct bulk_checkin_state *state,\n \treset_pack_idx_option(&state->pack_idx_opts);\n \n \t/* Pretend we are going to write only one object */\n-\tstate->offset = write_pack_header(state->f, 1);\n+\tstate->offset = write_pack_header(state->f, 2, 1);\n \tif (!state->offset)\n \t\tdie_errno(\"unable to write pack header\");\n }\ndiff --git a/pack-write.c b/pack-write.c\nindex 631007e..88e4788 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -186,12 +186,15 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \treturn index_name;\n }\n \n-off_t write_pack_header(struct sha1file *f, uint32_t nr_entries)\n+off_t write_pack_header(struct sha1file *f,\n+\t\t\tint version, uint32_t nr_entries)\n {\n \tstruct pack_header hdr;\n \n \thdr.hdr_signature = htonl(PACK_SIGNATURE);\n-\thdr.hdr_version = htonl(PACK_VERSION);\n+\thdr.hdr_version = htonl(version);\n+\tif (!pack_version_ok(hdr.hdr_version))\n+\t\tdie(_(\"pack version %d is not supported\"), version);\n \thdr.hdr_entries = htonl(nr_entries);\n \tif (sha1write(f, &hdr, sizeof(hdr)))\n \t\treturn 0;\ndiff --git a/pack.h b/pack.h\nindex aa6ee7d..855f6c6 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -8,7 +8,6 @@\n  * Packed object header\n  */\n #define PACK_SIGNATURE 0x5041434b\t/* \"PACK\" */\n-#define PACK_VERSION 2\n #define pack_version_ok(v) ((v) == htonl(2) || (v) == htonl(3))\n struct pack_header {\n \tuint32_t hdr_signature;\n@@ -80,7 +79,7 @@ extern const char *write_idx_file(const char *index_name, struct pack_idx_entry\n extern int check_pack_crc(struct packed_git *p, struct pack_window **w_curs, off_t offset, off_t len, unsigned int nr);\n extern int verify_pack_index(struct packed_git *);\n extern int verify_pack(struct packed_git *, verify_fn fn, struct progress *, uint32_t);\n-extern off_t write_pack_header(struct sha1file *f, uint32_t);\n+extern off_t write_pack_header(struct sha1file *f, int, uint32_t);\n extern void fixup_pack_header_footer(int, unsigned char *, const char *, uint32_t, unsigned char *, off_t);\n extern char *index_pack_lockfile(int fd);\n extern int encode_in_pack_object_header(enum object_type, uintmax_t, unsigned char *);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227131","messageId":"1378652660-6731-6-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 05/11] pack-write.c: add pv4_encode_in_pack_object_header","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:14Z","receivedAt":"2013-09-08T15:04:14Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n pack-write.c | 29 +++++++++++++++++++++++++++++\n pack.h       |  1 +\n 2 files changed, 30 insertions(+)\n\ndiff --git a/pack-write.c b/pack-write.c\nindex 88e4788..6f11104 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -1,6 +1,7 @@\n #include \"cache.h\"\n #include \"pack.h\"\n #include \"csum-file.h\"\n+#include \"varint.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -340,6 +341,34 @@ int encode_in_pack_object_header(enum object_type type, uintmax_t size, unsigned\n \treturn n;\n }\n \n+/*\n+ * The per-object header is a pretty dense thing, which is\n+ *  - first byte: low four bits are \"size\", then three bits of \"type\",\n+ *    and the high bit is \"size continues\".\n+ *  - each byte afterwards: low seven bits are size continuation,\n+ *    with the high bit being \"size continues\"\n+ */\n+int pv4_encode_in_pack_object_header(enum object_type type,\n+\t\t\t\t     uintmax_t size, unsigned char *hdr)\n+{\n+\tuintmax_t val;\n+\tif (type < OBJ_COMMIT || type > OBJ_PV4_TREE || type == OBJ_OFS_DELTA)\n+\t\tdie(\"bad type %d\", type);\n+\n+\t/*\n+\t * We allocate 4 bits in the LSB for the object type which\n+\t * should be good for quite a while, given that we effectively\n+\t * encodes only 5 object types: commit, tree, blob, delta,\n+\t * tag.\n+\t */\n+\tval = size;\n+\tif (MSB(val, 4))\n+\t\tdie(\"fixme: the code doesn't currently cope with big sizes\");\n+\tval <<= 4;\n+\tval |= type;\n+\treturn encode_varint(val, hdr);\n+}\n+\n struct sha1file *create_tmp_packfile(char **pack_tmp_name)\n {\n \tchar tmpname[PATH_MAX];\ndiff --git a/pack.h b/pack.h\nindex 855f6c6..38f869d 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -83,6 +83,7 @@ extern off_t write_pack_header(struct sha1file *f, int, uint32_t);\n extern void fixup_pack_header_footer(int, unsigned char *, const char *, uint32_t, unsigned char *, off_t);\n extern char *index_pack_lockfile(int fd);\n extern int encode_in_pack_object_header(enum object_type, uintmax_t, unsigned char *);\n+extern int pv4_encode_in_pack_object_header(enum object_type, uintmax_t, unsigned char *);\n \n #define PH_ERROR_EOF\t\t(-1)\n #define PH_ERROR_PACK_SIGNATURE\t(-2)\n-- \n1.8.2.83.gc99314b\n"},{"id":"227132","messageId":"1378652660-6731-7-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 06/11] pack-objects: add --version to specify written pack version","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:15Z","receivedAt":"2013-09-08T15:04:15Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 6 +++++-\n 1 file changed, 5 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 33faea8..ef68fc5 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -81,6 +81,7 @@ static int num_preferred_base;\n static struct progress *progress_state;\n static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n static int pack_compression_seen;\n+static int pack_version = 2;\n \n static unsigned long delta_cache_size = 0;\n static unsigned long max_delta_cache_size = 256 * 1024 * 1024;\n@@ -735,7 +736,7 @@ static void write_pack_file(void)\n \t\telse\n \t\t\tf = create_tmp_packfile(&pack_tmp_name);\n \n-\t\toffset = write_pack_header(f, 2, nr_remaining);\n+\t\toffset = write_pack_header(f, pack_version, nr_remaining);\n \t\tif (!offset)\n \t\t\tdie_errno(\"unable to write pack header\");\n \t\tnr_written = 0;\n@@ -2455,6 +2456,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t{ OPTION_CALLBACK, 0, \"index-version\", NULL, N_(\"version[,offset]\"),\n \t\t  N_(\"write the pack index file in the specified idx format version\"),\n \t\t  0, option_parse_index_version },\n+\t\tOPT_INTEGER(0, \"version\", &pack_version, N_(\"pack version\")),\n \t\tOPT_ULONG(0, \"max-pack-size\", &pack_size_limit,\n \t\t\t  N_(\"maximum size of each output pack file\")),\n \t\tOPT_BOOL(0, \"local\", &local,\n@@ -2525,6 +2527,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t}\n \tif (pack_to_stdout != !base_name || argc)\n \t\tusage_with_options(pack_usage, pack_objects_options);\n+\tif (pack_version != 2)\n+\t\tdie(_(\"pack version %d is not supported\"), pack_version);\n \n \trp_av[rp_ac++] = \"pack-objects\";\n \tif (thin) {\n-- \n1.8.2.83.gc99314b\n"},{"id":"227133","messageId":"1378652660-6731-8-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 07/11] list-objects.c: add show_tree_entry callback to traverse_commit_list","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:16Z","receivedAt":"2013-09-08T15:04:16Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This helps construct tree dictionary in pack v4.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 2 +-\n builtin/rev-list.c     | 4 ++--\n list-objects.c         | 9 ++++++++-\n list-objects.h         | 3 ++-\n upload-pack.c          | 2 +-\n 5 files changed, 14 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex ef68fc5..b38d3dc 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -2380,7 +2380,7 @@ static void get_object_list(int ac, const char **av)\n \tif (prepare_revision_walk(&revs))\n \t\tdie(\"revision walk setup failed\");\n \tmark_edges_uninteresting(revs.commits, &revs, show_edge);\n-\ttraverse_commit_list(&revs, show_commit, show_object, NULL);\n+\ttraverse_commit_list(&revs, show_commit, NULL, show_object, NULL);\n \n \tif (keep_unreachable)\n \t\tadd_objects_in_unpacked_packs(&revs);\ndiff --git a/builtin/rev-list.c b/builtin/rev-list.c\nindex a5ec30d..b25f896 100644\n--- a/builtin/rev-list.c\n+++ b/builtin/rev-list.c\n@@ -243,7 +243,7 @@ static int show_bisect_vars(struct rev_list_info *info, int reaches, int all)\n \t\tstrcpy(hex, sha1_to_hex(revs->commits->item->object.sha1));\n \n \tif (flags & BISECT_SHOW_ALL) {\n-\t\ttraverse_commit_list(revs, show_commit, show_object, info);\n+\t\ttraverse_commit_list(revs, show_commit, NULL, show_object, info);\n \t\tprintf(\"------\\n\");\n \t}\n \n@@ -348,7 +348,7 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \t\t\treturn show_bisect_vars(&info, reaches, all);\n \t}\n \n-\ttraverse_commit_list(&revs, show_commit, show_object, &info);\n+\ttraverse_commit_list(&revs, show_commit, NULL, show_object, &info);\n \n \tif (revs.count) {\n \t\tif (revs.left_right && revs.cherry_mark)\ndiff --git a/list-objects.c b/list-objects.c\nindex 3dd4a96..6def897 100644\n--- a/list-objects.c\n+++ b/list-objects.c\n@@ -61,6 +61,7 @@ static void process_gitlink(struct rev_info *revs,\n \n static void process_tree(struct rev_info *revs,\n \t\t\t struct tree *tree,\n+\t\t\t show_tree_entry_fn show_tree_entry,\n \t\t\t show_object_fn show,\n \t\t\t struct name_path *path,\n \t\t\t struct strbuf *base,\n@@ -107,9 +108,13 @@ static void process_tree(struct rev_info *revs,\n \t\t\t\tcontinue;\n \t\t}\n \n+\t\tif (show_tree_entry)\n+\t\t\tshow_tree_entry(&entry, cb_data);\n+\n \t\tif (S_ISDIR(entry.mode))\n \t\t\tprocess_tree(revs,\n \t\t\t\t     lookup_tree(entry.sha1),\n+\t\t\t\t     show_tree_entry,\n \t\t\t\t     show, &me, base, entry.path,\n \t\t\t\t     cb_data);\n \t\telse if (S_ISGITLINK(entry.mode))\n@@ -167,6 +172,7 @@ static void add_pending_tree(struct rev_info *revs, struct tree *tree)\n \n void traverse_commit_list(struct rev_info *revs,\n \t\t\t  show_commit_fn show_commit,\n+\t\t\t  show_tree_entry_fn show_tree_entry,\n \t\t\t  show_object_fn show_object,\n \t\t\t  void *data)\n {\n@@ -196,7 +202,8 @@ void traverse_commit_list(struct rev_info *revs,\n \t\t\tcontinue;\n \t\t}\n \t\tif (obj->type == OBJ_TREE) {\n-\t\t\tprocess_tree(revs, (struct tree *)obj, show_object,\n+\t\t\tprocess_tree(revs, (struct tree *)obj,\n+\t\t\t\t     show_tree_entry, show_object,\n \t\t\t\t     NULL, &base, name, data);\n \t\t\tcontinue;\n \t\t}\ndiff --git a/list-objects.h b/list-objects.h\nindex 3db7bb6..297b2e0 100644\n--- a/list-objects.h\n+++ b/list-objects.h\n@@ -2,8 +2,9 @@\n #define LIST_OBJECTS_H\n \n typedef void (*show_commit_fn)(struct commit *, void *);\n+typedef void (*show_tree_entry_fn)(const struct name_entry *, void *);\n typedef void (*show_object_fn)(struct object *, const struct name_path *, const char *, void *);\n-void traverse_commit_list(struct rev_info *, show_commit_fn, show_object_fn, void *);\n+void traverse_commit_list(struct rev_info *, show_commit_fn, show_tree_entry_fn, show_object_fn, void *);\n \n typedef void (*show_edge_fn)(struct commit *);\n void mark_edges_uninteresting(struct commit_list *, struct rev_info *, show_edge_fn);\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 127e59a..ccf76d9 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -125,7 +125,7 @@ static int do_rev_list(int in, int out, void *user_data)\n \t\tfor (i = 0; i < extra_edge_obj.nr; i++)\n \t\t\tfprintf(pack_pipe, \"-%s\\n\", sha1_to_hex(\n \t\t\t\t\textra_edge_obj.objects[i].item->sha1));\n-\ttraverse_commit_list(&revs, show_commit, show_object, NULL);\n+\ttraverse_commit_list(&revs, show_commit, NULL, show_object, NULL);\n \tfflush(pack_pipe);\n \tfclose(pack_pipe);\n \treturn 0;\n-- \n1.8.2.83.gc99314b\n"},{"id":"227134","messageId":"1378652660-6731-9-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 08/11] pack-objects: create pack v4 tables","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:17Z","receivedAt":"2013-09-08T15:04:17Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 87 ++++++++++++++++++++++++++++++++++++++++++++++++--\n 1 file changed, 85 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex b38d3dc..69a22c1 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -18,6 +18,7 @@\n #include \"refs.h\"\n #include \"streaming.h\"\n #include \"thread-utils.h\"\n+#include \"packv4-create.h\"\n \n static const char *pack_usage[] = {\n \tN_(\"git pack-objects --stdout [options...] [< ref-list | < object-list]\"),\n@@ -61,6 +62,8 @@ static struct object_entry *objects;\n static struct pack_idx_entry **written_list;\n static uint32_t nr_objects, nr_alloc, nr_result, nr_written;\n \n+static struct packv4_tables v4;\n+\n static int non_empty;\n static int reuse_delta = 1, reuse_object = 1;\n static int keep_unreachable, unpack_unreachable, include_tag;\n@@ -2039,12 +2042,42 @@ static int add_ref_tag(const char *path, const unsigned char *sha1, int flag, vo\n \treturn 0;\n }\n \n+static int sha1_idx_sort(const void *a_, const void *b_)\n+{\n+\tconst struct pack_idx_entry *a = a_;\n+\tconst struct pack_idx_entry *b = b_;\n+\treturn hashcmp(a->sha1, b->sha1);\n+}\n+\n+static void prepare_sha1_table(void)\n+{\n+\tunsigned i;\n+\t/*\n+\t * This table includes SHA-1s that may not be present in the\n+\t * pack. One of the use of such SHA-1 is for completing thin\n+\t * packs, where index-pack does not need to add SHA-1 to the\n+\t * table at completion time.\n+\t */\n+\tv4.all_objs = xmalloc(nr_objects * sizeof(*v4.all_objs));\n+\tv4.all_objs_nr = nr_objects;\n+\tfor (i = 0; i < nr_objects; i++)\n+\t\tv4.all_objs[i] = objects[i].idx;\n+\tqsort(v4.all_objs, nr_objects, sizeof(*v4.all_objs),\n+\t      sha1_idx_sort);\n+}\n+\n static void prepare_pack(int window, int depth)\n {\n \tstruct object_entry **delta_list;\n \tuint32_t i, nr_deltas;\n \tunsigned n;\n \n+\tif (pack_version == 4) {\n+\t\tsort_dict_entries_by_hits(v4.commit_ident_table);\n+\t\tsort_dict_entries_by_hits(v4.tree_path_table);\n+\t\tprepare_sha1_table();\n+\t}\n+\n \tget_object_details();\n \n \t/*\n@@ -2191,6 +2224,34 @@ static void read_object_list_from_stdin(void)\n \n \t\tadd_preferred_base_object(line+41);\n \t\tadd_object_entry(sha1, 0, line+41, 0);\n+\n+\t\tif (pack_version == 4) {\n+\t\t\tvoid *data;\n+\t\t\tenum object_type type;\n+\t\t\tunsigned long size;\n+\t\t\tint (*add_dict_entries)(struct dict_table *, void *, unsigned long);\n+\t\t\tstruct dict_table *dict;\n+\n+\t\t\tswitch (sha1_object_info(sha1, &size)) {\n+\t\t\tcase OBJ_COMMIT:\n+\t\t\t\tadd_dict_entries = add_commit_dict_entries;\n+\t\t\t\tdict = v4.commit_ident_table;\n+\t\t\t\tbreak;\n+\t\t\tcase OBJ_TREE:\n+\t\t\t\tadd_dict_entries = add_tree_dict_entries;\n+\t\t\t\tdict = v4.tree_path_table;\n+\t\t\t\tbreak;\n+\t\t\tdefault:\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tdata = read_sha1_file(sha1, &type, &size);\n+\t\t\tif (!data)\n+\t\t\t\tdie(\"cannot unpack %s\", sha1_to_hex(sha1));\n+\t\t\tif (add_dict_entries(dict, data, size) < 0)\n+\t\t\t\tdie(\"can't process %s object %s\",\n+\t\t\t\t    typename(type), sha1_to_hex(sha1));\n+\t\t\tfree(data);\n+\t\t}\n \t}\n }\n \n@@ -2198,10 +2259,26 @@ static void read_object_list_from_stdin(void)\n \n static void show_commit(struct commit *commit, void *data)\n {\n+\tif (pack_version == 4) {\n+\t\tunsigned long size;\n+\t\tenum object_type type;\n+\t\tunsigned char *buf;\n+\n+\t\t/* commit->buffer is NULL most of the time, don't bother */\n+\t\tbuf = read_sha1_file(commit->object.sha1, &type, &size);\n+\t\tadd_commit_dict_entries(v4.commit_ident_table, buf, size);\n+\t\tfree(buf);\n+\t}\n \tadd_object_entry(commit->object.sha1, OBJ_COMMIT, NULL, 0);\n \tcommit->object.flags |= OBJECT_ADDED;\n }\n \n+static void show_tree_entry(const struct name_entry *entry, void *data)\n+{\n+\tdict_add_entry(v4.tree_path_table, entry->mode, entry->path,\n+\t\t       tree_entry_len(entry));\n+}\n+\n static void show_object(struct object *obj,\n \t\t\tconst struct name_path *path, const char *last,\n \t\t\tvoid *data)\n@@ -2380,7 +2457,9 @@ static void get_object_list(int ac, const char **av)\n \tif (prepare_revision_walk(&revs))\n \t\tdie(\"revision walk setup failed\");\n \tmark_edges_uninteresting(revs.commits, &revs, show_edge);\n-\ttraverse_commit_list(&revs, show_commit, NULL, show_object, NULL);\n+\ttraverse_commit_list(&revs, show_commit,\n+\t\t\t     pack_version == 4 ? show_tree_entry : NULL,\n+\t\t\t     show_object, NULL);\n \n \tif (keep_unreachable)\n \t\tadd_objects_in_unpacked_packs(&revs);\n@@ -2527,7 +2606,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t}\n \tif (pack_to_stdout != !base_name || argc)\n \t\tusage_with_options(pack_usage, pack_objects_options);\n-\tif (pack_version != 2)\n+\tif (pack_version != 2 && pack_version != 4)\n \t\tdie(_(\"pack version %d is not supported\"), pack_version);\n \n \trp_av[rp_ac++] = \"pack-objects\";\n@@ -2579,6 +2658,10 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tprogress = 2;\n \n \tprepare_packed_git();\n+\tif (pack_version == 4) {\n+\t\tv4.commit_ident_table = create_dict_table();\n+\t\tv4.tree_path_table = create_dict_table();\n+\t}\n \n \tif (progress)\n \t\tprogress_state = start_progress(\"Counting objects\", 0);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227135","messageId":"1378652660-6731-10-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 09/11] pack-objects: do not cache delta for v4 trees","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:18Z","receivedAt":"2013-09-08T15:04:18Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 6 +++++-\n 1 file changed, 5 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 69a22c1..665853d 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1759,8 +1759,12 @@ static void find_deltas(struct object_entry **list, unsigned *list_size,\n \t\t * and therefore it is best to go to the write phase ASAP\n \t\t * instead, as we can afford spending more time compressing\n \t\t * between writes at that moment.\n+\t\t *\n+\t\t * For v4 trees we'll need to delta differently anyway\n+\t\t * so no cache. v4 commits simply do not delta.\n \t\t */\n-\t\tif (entry->delta_data && !pack_to_stdout) {\n+\t\tif (entry->delta_data && !pack_to_stdout &&\n+\t\t    (pack_version < 4 || entry->type == OBJ_BLOB)) {\n \t\t\tentry->z_delta_size = do_compress(&entry->delta_data,\n \t\t\t\t\t\t\t  entry->delta_size);\n \t\t\tcache_lock();\n-- \n1.8.2.83.gc99314b\n"},{"id":"227136","messageId":"1378652660-6731-11-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 10/11] pack-objects: exclude commits out of delta objects in v4","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:19Z","receivedAt":"2013-09-08T15:04:19Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 5 ++++-\n 1 file changed, 4 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 665853d..daa4349 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1332,7 +1332,8 @@ static void check_object(struct object_entry *entry)\n \t\t\tbreak;\n \t\t}\n \n-\t\tif (base_ref && (base_entry = locate_object_entry(base_ref))) {\n+\t\tif (base_ref && (base_entry = locate_object_entry(base_ref)) &&\n+\t\t    (pack_version < 4 || entry->type != OBJ_COMMIT)) {\n \t\t\t/*\n \t\t\t * If base_ref was set above that means we wish to\n \t\t\t * reuse delta data, and we even found that base\n@@ -1416,6 +1417,8 @@ static void get_object_details(void)\n \t\tcheck_object(entry);\n \t\tif (big_file_threshold < entry->size)\n \t\t\tentry->no_try_delta = 1;\n+\t\tif (pack_version == 4 && entry->type == OBJ_COMMIT)\n+\t\t\tentry->no_try_delta = 1;\n \t}\n \n \tfree(sorted_by_offset);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227137","messageId":"1378652660-6731-12-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 11/11] pack-objects: support writing pack v4","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-08T15:04:20Z","receivedAt":"2013-09-08T15:04:20Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 85 ++++++++++++++++++++++++++++++++++++++++++++------\n pack.h                 |  2 +-\n 2 files changed, 76 insertions(+), 11 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex daa4349..f6586a1 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -254,8 +254,10 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \tenum object_type type;\n \tvoid *buf;\n \tstruct git_istream *st = NULL;\n+\tchar *result = \"OK\";\n \n-\tif (!usable_delta) {\n+\tif (!usable_delta ||\n+\t    (pack_version == 4 || entry->type == OBJ_TREE)) {\n \t\tif (entry->type == OBJ_BLOB &&\n \t\t    entry->size > big_file_threshold &&\n \t\t    (st = open_istream(entry->idx.sha1, &type, &size, NULL)) != NULL)\n@@ -287,7 +289,37 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \n \tif (st)\t/* large blob case, just assume we don't compress well */\n \t\tdatalen = size;\n-\telse if (entry->z_delta_size)\n+\telse if (pack_version == 4 && entry->type == OBJ_COMMIT) {\n+\t\tdatalen = size;\n+\t\tresult = pv4_encode_commit(&v4, buf, &datalen);\n+\t\tif (result) {\n+\t\t\tfree(buf);\n+\t\t\tbuf = result;\n+\t\t\ttype = OBJ_PV4_COMMIT;\n+\t\t}\n+\t} else if (pack_version == 4 && entry->type == OBJ_TREE) {\n+\t\tdatalen = size;\n+\t\tif (usable_delta) {\n+\t\t\tunsigned long base_size;\n+\t\t\tchar *base_buf;\n+\t\t\tbase_buf = read_sha1_file(entry->delta->idx.sha1, &type,\n+\t\t\t\t\t\t  &base_size);\n+\t\t\tif (!base_buf || type != OBJ_TREE)\n+\t\t\t\tdie(\"unable to read %s\",\n+\t\t\t\t    sha1_to_hex(entry->delta->idx.sha1));\n+\t\t\tresult = pv4_encode_tree(&v4, buf, &datalen,\n+\t\t\t\t\t\t base_buf, base_size,\n+\t\t\t\t\t\t entry->delta->idx.sha1);\n+\t\t\tfree(base_buf);\n+\t\t} else\n+\t\t\tresult = pv4_encode_tree(&v4, buf, &datalen,\n+\t\t\t\t\t\t NULL, 0, NULL);\n+\t\tif (result) {\n+\t\t\tfree(buf);\n+\t\t\tbuf = result;\n+\t\t\ttype = OBJ_PV4_TREE;\n+\t\t}\n+\t} else if (entry->z_delta_size)\n \t\tdatalen = entry->z_delta_size;\n \telse\n \t\tdatalen = do_compress(&buf, size);\n@@ -296,7 +328,10 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \t * The object header is a byte of 'type' followed by zero or\n \t * more bytes of length.\n \t */\n-\thdrlen = encode_in_pack_object_header(type, size, header);\n+\tif (pack_version < 4)\n+\t\thdrlen = encode_in_pack_object_header(type, size, header);\n+\telse\n+\t\thdrlen = pv4_encode_in_pack_object_header(type, size, header);\n \n \tif (type == OBJ_OFS_DELTA) {\n \t\t/*\n@@ -318,7 +353,7 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \t\tsha1write(f, header, hdrlen);\n \t\tsha1write(f, dheader + pos, sizeof(dheader) - pos);\n \t\thdrlen += sizeof(dheader) - pos;\n-\t} else if (type == OBJ_REF_DELTA) {\n+\t} else if (type == OBJ_REF_DELTA && pack_version < 4) {\n \t\t/*\n \t\t * Deltas with a base reference contain\n \t\t * an additional 20 bytes for the base sha1.\n@@ -332,6 +367,10 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \t\tsha1write(f, header, hdrlen);\n \t\tsha1write(f, entry->delta->idx.sha1, 20);\n \t\thdrlen += 20;\n+\t} else if (type == OBJ_REF_DELTA && pack_version == 4) {\n+\t\thdrlen += encode_sha1ref(&v4, entry->delta->idx.sha1,\n+\t\t\t\t\theader + hdrlen);\n+\t\tsha1write(f, header, hdrlen);\n \t} else {\n \t\tif (limit && hdrlen + datalen + 20 >= limit) {\n \t\t\tif (st)\n@@ -341,14 +380,26 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \t\t}\n \t\tsha1write(f, header, hdrlen);\n \t}\n+\n \tif (st) {\n \t\tdatalen = write_large_blob_data(st, f, entry->idx.sha1);\n \t\tclose_istream(st);\n-\t} else {\n-\t\tsha1write(f, buf, datalen);\n-\t\tfree(buf);\n+\t\treturn hdrlen + datalen;\n \t}\n \n+\tif (!result) {\n+\t\twarning(_(\"can't convert %s object %s\"),\n+\t\t\ttypename(entry->type),\n+\t\t\tsha1_to_hex(entry->idx.sha1));\n+\t\tfree(buf);\n+\t\tbuf = read_sha1_file(entry->idx.sha1, &type, &size);\n+\t\tif (!buf)\n+\t\t\tdie(_(\"unable to read %s\"),\n+\t\t\t    sha1_to_hex(entry->idx.sha1));\n+\t\tdatalen = do_compress(&buf, size);\n+\t}\n+\tsha1write(f, buf, datalen);\n+\tfree(buf);\n \treturn hdrlen + datalen;\n }\n \n@@ -368,7 +419,10 @@ static unsigned long write_reuse_object(struct sha1file *f, struct object_entry\n \tif (entry->delta)\n \t\ttype = (allow_ofs_delta && entry->delta->idx.offset) ?\n \t\t\tOBJ_OFS_DELTA : OBJ_REF_DELTA;\n-\thdrlen = encode_in_pack_object_header(type, entry->size, header);\n+\tif (pack_version < 4)\n+\t\thdrlen = encode_in_pack_object_header(type, entry->size, header);\n+\telse\n+\t\thdrlen = pv4_encode_in_pack_object_header(type, entry->size, header);\n \n \toffset = entry->in_pack_offset;\n \trevidx = find_pack_revindex(p, offset);\n@@ -404,7 +458,7 @@ static unsigned long write_reuse_object(struct sha1file *f, struct object_entry\n \t\tsha1write(f, dheader + pos, sizeof(dheader) - pos);\n \t\thdrlen += sizeof(dheader) - pos;\n \t\treused_delta++;\n-\t} else if (type == OBJ_REF_DELTA) {\n+\t} else if (type == OBJ_REF_DELTA && pack_version < 4) {\n \t\tif (limit && hdrlen + 20 + datalen + 20 >= limit) {\n \t\t\tunuse_pack(&w_curs);\n \t\t\treturn 0;\n@@ -413,6 +467,11 @@ static unsigned long write_reuse_object(struct sha1file *f, struct object_entry\n \t\tsha1write(f, entry->delta->idx.sha1, 20);\n \t\thdrlen += 20;\n \t\treused_delta++;\n+\t} else if (type == OBJ_REF_DELTA && pack_version == 4) {\n+\t\thdrlen += encode_sha1ref(&v4, entry->delta->idx.sha1,\n+\t\t\t\t\theader + hdrlen);\n+\t\tsha1write(f, header, hdrlen);\n+\t\treused_delta++;\n \t} else {\n \t\tif (limit && hdrlen + datalen + 20 >= limit) {\n \t\t\tunuse_pack(&w_curs);\n@@ -477,7 +536,9 @@ static unsigned long write_object(struct sha1file *f,\n \t\t\t\t * and we do not need to deltify it.\n \t\t\t\t */\n \n-\tif (!to_reuse)\n+\tif (!to_reuse ||\n+\t    (pack_version == 4 &&\n+\t     (entry->type == OBJ_TREE || entry->type == OBJ_COMMIT)))\n \t\tlen = write_no_reuse_object(f, entry, limit, usable_delta);\n \telse\n \t\tlen = write_reuse_object(f, entry, limit, usable_delta);\n@@ -742,6 +803,8 @@ static void write_pack_file(void)\n \t\toffset = write_pack_header(f, pack_version, nr_remaining);\n \t\tif (!offset)\n \t\t\tdie_errno(\"unable to write pack header\");\n+\t\tif (pack_version == 4)\n+\t\t\toffset += packv4_write_tables(f, &v4);\n \t\tnr_written = 0;\n \t\tfor (; i < nr_objects; i++) {\n \t\t\tstruct object_entry *e = write_order[i];\n@@ -2083,6 +2146,8 @@ static void prepare_pack(int window, int depth)\n \t\tsort_dict_entries_by_hits(v4.commit_ident_table);\n \t\tsort_dict_entries_by_hits(v4.tree_path_table);\n \t\tprepare_sha1_table();\n+\t\tpack_idx_opts.version = 3;\n+\t\tallow_ofs_delta = 0;\n \t}\n \n \tget_object_details();\ndiff --git a/pack.h b/pack.h\nindex 38f869d..fde60ec 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -8,7 +8,7 @@\n  * Packed object header\n  */\n #define PACK_SIGNATURE 0x5041434b\t/* \"PACK\" */\n-#define pack_version_ok(v) ((v) == htonl(2) || (v) == htonl(3))\n+#define pack_version_ok(v) ((v) == htonl(2) || (v) == htonl(3) || (v) == htonl(4))\n struct pack_header {\n \tuint32_t hdr_signature;\n \tuint32_t hdr_version;\n-- \n1.8.2.83.gc99314b\n"},{"id":"227149","messageId":"alpine.LFD.2.03.1309081638060.20709@syhkavp.arg","threadId":"34852","inReplyTo":"1378652660-6731-6-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 05/11] pack-write.c: add pv4_encode_in_pack_object_header","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-08T20:51:06Z","receivedAt":"2013-09-08T20:51:06Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 8 Sep 2013, Nguyễn Thái Ngọc Duy wrote:\n\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  pack-write.c | 29 +++++++++++++++++++++++++++++\n>  pack.h       |  1 +\n>  2 files changed, 30 insertions(+)\n> \n> diff --git a/pack-write.c b/pack-write.c\n> index 88e4788..6f11104 100644\n> --- a/pack-write.c\n> +++ b/pack-write.c\n> @@ -1,6 +1,7 @@\n>  #include \"cache.h\"\n>  #include \"pack.h\"\n>  #include \"csum-file.h\"\n> +#include \"varint.h\"\n>  \n>  void reset_pack_idx_option(struct pack_idx_option *opts)\n>  {\n> @@ -340,6 +341,34 @@ int encode_in_pack_object_header(enum object_type type, uintmax_t size, unsigned\n>  \treturn n;\n>  }\n>  \n> +/*\n> + * The per-object header is a pretty dense thing, which is\n> + *  - first byte: low four bits are \"size\", then three bits of \"type\",\n> + *    and the high bit is \"size continues\".\n> + *  - each byte afterwards: low seven bits are size continuation,\n> + *    with the high bit being \"size continues\"\n> + */\n\nThis comment is a bit misleading.  It looks almost like the pack v2 \nobject header encoding which is not a varint encoded value like this one \nis.\n\n> +int pv4_encode_in_pack_object_header(enum object_type type,\n> +\t\t\t\t     uintmax_t size, unsigned char *hdr)\n\nCould we have a somewhat shorter function name? \npv4_encode_object_header() should be acceptable given \"pv4\" already \nimplies a pack.\n\n> +{\n> +\tuintmax_t val;\n> +\tif (type < OBJ_COMMIT || type > OBJ_PV4_TREE || type == OBJ_OFS_DELTA)\n> +\t\tdie(\"bad type %d\", type);\n\nThis test has holes, such as types 5 and 8.\n\nI think this would be better as:\n\n\tswitch (type) {\n\tcase OBJ_COMMIT:\n\tcase OBJ_TREE:\n\tcase OBJ_BLOB:\n\tcase OBJ_TAG:\n\tcase OBJ_REF_DELTA:\n\tcase OBJ_PV4_COMMIT:\n\tcase OBJ_PV4_TREE:\n\t\tbreak;\n\tdefault:\n\t\tdie(\"bad type %d\", type);\n\t}\n\nThe compiler ought to be smart enough to optimize the contiguous case \nrange.  And that makes it explicit and obvious what we test for.\n\n> +\n> +\t/*\n> +\t * We allocate 4 bits in the LSB for the object type which\n> +\t * should be good for quite a while, given that we effectively\n> +\t * encodes only 5 object types: commit, tree, blob, delta,\n> +\t * tag.\n> +\t */\n> +\tval = size;\n> +\tif (MSB(val, 4))\n> +\t\tdie(\"fixme: the code doesn't currently cope with big sizes\");\n> +\tval <<= 4;\n> +\tval |= type;\n> +\treturn encode_varint(val, hdr);\n> +}\n> +\n>  struct sha1file *create_tmp_packfile(char **pack_tmp_name)\n>  {\n>  \tchar tmpname[PATH_MAX];\n> diff --git a/pack.h b/pack.h\n> index 855f6c6..38f869d 100644\n> --- a/pack.h\n> +++ b/pack.h\n> @@ -83,6 +83,7 @@ extern off_t write_pack_header(struct sha1file *f, int, uint32_t);\n>  extern void fixup_pack_header_footer(int, unsigned char *, const char *, uint32_t, unsigned char *, off_t);\n>  extern char *index_pack_lockfile(int fd);\n>  extern int encode_in_pack_object_header(enum object_type, uintmax_t, unsigned char *);\n> +extern int pv4_encode_in_pack_object_header(enum object_type, uintmax_t, unsigned char *);\n>  \n>  #define PH_ERROR_EOF\t\t(-1)\n>  #define PH_ERROR_PACK_SIGNATURE\t(-2)\n> -- \n> 1.8.2.83.gc99314b\n> \n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n> \n"},{"id":"227151","messageId":"alpine.LFD.2.03.1309081651490.20709@syhkavp.arg","threadId":"34852","inReplyTo":"1378652660-6731-4-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 03/11] pack v4: move packv4-create.c to libgit.a","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-08T20:56:28Z","receivedAt":"2013-09-08T20:56:28Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 8 Sep 2013, Nguyễn Thái Ngọc Duy wrote:\n\n> git-packv4-create now becomes test-packv4. Code that will not be used\n> by pack-objects.c is moved to test-packv4.c. It may be removed when\n> the code transition to pack-objects completes.\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  Makefile            |   4 +-\n>  packv4-create.c     | 491 +---------------------------------------------------\n>  packv4-create.h     |  39 +++++\n>  test-packv4.c (new) | 476 ++++++++++++++++++++++++++++++++++++++++++++++++++\n>  4 files changed, 525 insertions(+), 485 deletions(-)\n>  create mode 100644 test-packv4.c\n> \n> diff --git a/Makefile b/Makefile\n> index 22fc276..af2e3e3 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -550,7 +550,6 @@ PROGRAM_OBJS += shell.o\n>  PROGRAM_OBJS += show-index.o\n>  PROGRAM_OBJS += upload-pack.o\n>  PROGRAM_OBJS += remote-testsvn.o\n> -PROGRAM_OBJS += packv4-create.o\n>  \n>  # Binary suffix, set to .exe for Windows builds\n>  X =\n> @@ -568,6 +567,7 @@ TEST_PROGRAMS_NEED_X += test-line-buffer\n>  TEST_PROGRAMS_NEED_X += test-match-trees\n>  TEST_PROGRAMS_NEED_X += test-mergesort\n>  TEST_PROGRAMS_NEED_X += test-mktemp\n> +TEST_PROGRAMS_NEED_X += test-packv4\n>  TEST_PROGRAMS_NEED_X += test-parse-options\n>  TEST_PROGRAMS_NEED_X += test-path-utils\n>  TEST_PROGRAMS_NEED_X += test-prio-queue\n> @@ -702,6 +702,7 @@ LIB_H += notes.h\n>  LIB_H += object.h\n>  LIB_H += pack-revindex.h\n>  LIB_H += pack.h\n> +LIB_H += packv4-create.h\n>  LIB_H += packv4-parse.h\n>  LIB_H += parse-options.h\n>  LIB_H += patch-ids.h\n> @@ -839,6 +840,7 @@ LIB_OBJS += object.o\n>  LIB_OBJS += pack-check.o\n>  LIB_OBJS += pack-revindex.o\n>  LIB_OBJS += pack-write.o\n> +LIB_OBJS += packv4-create.o\n>  LIB_OBJS += packv4-parse.o\n>  LIB_OBJS += pager.o\n>  LIB_OBJS += parse-options.o\n> diff --git a/packv4-create.c b/packv4-create.c\n> index 920a0b4..cdf82c0 100644\n> --- a/packv4-create.c\n> +++ b/packv4-create.c\n> @@ -18,9 +18,9 @@\n>  #include \"packv4-create.h\"\n>  \n>  \n> -static int pack_compression_seen;\n> -static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n> -static int min_tree_copy = 1;\n> +int pack_compression_seen;\n> +int pack_compression_level = Z_DEFAULT_COMPRESSION;\n> +int min_tree_copy = 1;\n>  \n>  struct data_entry {\n>  \tunsigned offset;\n> @@ -28,17 +28,6 @@ struct data_entry {\n>  \tunsigned hits;\n>  };\n>  \n> -struct dict_table {\n> -\tunsigned char *data;\n> -\tunsigned cur_offset;\n> -\tunsigned size;\n> -\tstruct data_entry *entry;\n> -\tunsigned nb_entries;\n> -\tunsigned max_entries;\n> -\tunsigned *hash;\n> -\tunsigned hash_size;\n> -};\n\nIt doesn't seem necessary to move this structure definition to the \nheader file.  Only an opaque\n\n\tstruct dict_table;\n\nshould be needed in packv4-create.h.  That would keep the dictionary \nhandling localized.\n\n\nNicolas\n"},{"id":"227208","messageId":"CACsJy8DbMnr9Y8NyGTNd6r8hSg3zbgaLa1h-e1X7FFVHHahwpg@mail.gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-9-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 08/11] pack-objects: create pack v4 tables","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T10:40:45Z","receivedAt":"2013-09-09T10:40:45Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Sep 8, 2013 at 10:04 PM, Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n> +static void prepare_sha1_table(void)\n> +{\n> +       unsigned i;\n> +       /*\n> +        * This table includes SHA-1s that may not be present in the\n> +        * pack. One of the use of such SHA-1 is for completing thin\n> +        * packs, where index-pack does not need to add SHA-1 to the\n> +        * table at completion time.\n> +        */\n> +       v4.all_objs = xmalloc(nr_objects * sizeof(*v4.all_objs));\n> +       v4.all_objs_nr = nr_objects;\n> +       for (i = 0; i < nr_objects; i++)\n> +               v4.all_objs[i] = objects[i].idx;\n> +       qsort(v4.all_objs, nr_objects, sizeof(*v4.all_objs),\n> +             sha1_idx_sort);\n> +}\n> +\n\nfwiw this is wrong. Even in the non-thin pack case, pack-objects could\nwrite multiple packs to disk and we need different sha-1 table for\neach one. The situation is worse for thin pack because not all\npreferred_base entries end up a real dependency in the final pack. I'm\nworking on it..\n-- \nDuy\n"},{"id":"227209","messageId":"alpine.LFD.2.03.1309090900210.20709@syhkavp.arg","threadId":"34852","inReplyTo":"CACsJy8DbMnr9Y8NyGTNd6r8hSg3zbgaLa1h-e1X7FFVHHahwpg@mail.gmail.com","subject":"Re: [PATCH 08/11] pack-objects: create pack v4 tables","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-09T13:07:08Z","receivedAt":"2013-09-09T13:07:08Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 9 Sep 2013, Duy Nguyen wrote:\n\n> On Sun, Sep 8, 2013 at 10:04 PM, Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n> > +static void prepare_sha1_table(void)\n> > +{\n> > +       unsigned i;\n> > +       /*\n> > +        * This table includes SHA-1s that may not be present in the\n> > +        * pack. One of the use of such SHA-1 is for completing thin\n> > +        * packs, where index-pack does not need to add SHA-1 to the\n> > +        * table at completion time.\n> > +        */\n> > +       v4.all_objs = xmalloc(nr_objects * sizeof(*v4.all_objs));\n> > +       v4.all_objs_nr = nr_objects;\n> > +       for (i = 0; i < nr_objects; i++)\n> > +               v4.all_objs[i] = objects[i].idx;\n> > +       qsort(v4.all_objs, nr_objects, sizeof(*v4.all_objs),\n> > +             sha1_idx_sort);\n> > +}\n> > +\n> \n> fwiw this is wrong. Even in the non-thin pack case, pack-objects could\n> write multiple packs to disk and we need different sha-1 table for\n> each one. The situation is worse for thin pack because not all\n> preferred_base entries end up a real dependency in the final pack. I'm\n> working on it..\n\nIs anyone still using --max-pack-size ?\n\nI'm wondering if producing multiple packs from pack-objects is really \nuseful these days.  If I remember correctly, this was created to allow \nthe archiving of large packs onto CDROMs or the like.\n\nI'd be tempted to simply ignore this facility and get rid of its \ncomplexity if no one uses it.  Or assume that split packs will have \ninter dependencies.  Or they will be pack v2 only.\n\n\nNicolas\n"},{"id":"227214","messageId":"1378735087-4813-1-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378652660-6731-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 00/16] pack v4 support in pack-objects","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:57:51Z","receivedAt":"2013-09-09T13:57:51Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This version supports thin pack. I could clone from git.git with only\nmaster, then fetch the rest and fsck did not complain anything. I did\nnot check if I broke --max-pack-size though.\n\nInteresting patches are the ones near the end: \"prepare SHA-1 table\",\n\"support writing pack v4\" and \"support completing thin packs\"\n\nStill rough edges. If I don't find any new problems, I'll try to run\nthe test suite. My vacation days are over, so I will work at a much\nslower pace than the last couple days.\n\nNguyễn Thái Ngọc Duy (16):\n  pack v4: allocate dicts from the beginning\n  pack v4: stop using static/global variables in packv4-create.c\n  pack v4: move packv4-create.c to libgit.a\n  pack v4: add version argument to write_pack_header\n  pack_write: tighten valid object type check in\n    encode_in_pack_object_header\n  pack-write.c: add pv4_encode_object_header\n  pack-objects: add --version to specify written pack version\n  list-objects.c: add show_tree_entry callback to traverse_commit_list\n  pack-objects: do not cache delta for v4 trees\n  pack-objects: exclude commits out of delta objects in v4\n  pack-objects: create pack v4 tables\n  pack-objects: prepare SHA-1 table in v4\n  pack-objects: support writing pack v4\n  pack v4: support \"end-of-pack\" indicator in index-pack and\n    pack-objects\n  index-pack: use nr_objects_final as sha1_table size\n  index-pack: support completing thin packs v4\n\n Makefile               |   4 +-\n builtin/index-pack.c   |  95 ++++++---\n builtin/pack-objects.c | 230 ++++++++++++++++++++--\n builtin/rev-list.c     |   4 +-\n bulk-checkin.c         |   2 +-\n list-objects.c         |   9 +-\n list-objects.h         |   3 +-\n pack-write.c           |  51 ++++-\n pack.h                 |   6 +-\n packv4-create.c        | 523 ++++---------------------------------------------\n packv4-create.h (new)  |  39 ++++\n test-packv4.c (new)    | 476 ++++++++++++++++++++++++++++++++++++++++++++\n upload-pack.c          |   2 +-\n 13 files changed, 901 insertions(+), 543 deletions(-)\n create mode 100644 packv4-create.h\n create mode 100644 test-packv4.c\n\n-- \n1.8.2.83.gc99314b\n"},{"id":"227219","messageId":"1378735087-4813-2-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 01/16] pack v4: allocate dicts from the beginning","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:57:52Z","receivedAt":"2013-09-09T13:57:52Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"commit_ident_table and tree_path_table are local to packv4-create.c\nand test-packv4.c. Move them out of add_*_dict_entries so\nadd_*_dict_entries can be exported to pack-objects.c\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n packv4-create.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 38fa594..dbc2a03 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -181,14 +181,12 @@ static char *get_nameend_and_tz(char *from, int *tz_val)\n \treturn end;\n }\n \n-static int add_commit_dict_entries(void *buf, unsigned long size)\n+int add_commit_dict_entries(struct dict_table *commit_ident_table,\n+\t\t\t    void *buf, unsigned long size)\n {\n \tchar *name, *end = NULL;\n \tint tz_val;\n \n-\tif (!commit_ident_table)\n-\t\tcommit_ident_table = create_dict_table();\n-\n \t/* parse and add author info */\n \tname = strstr(buf, \"\\nauthor \");\n \tif (name) {\n@@ -212,14 +210,12 @@ static int add_commit_dict_entries(void *buf, unsigned long size)\n \treturn 0;\n }\n \n-static int add_tree_dict_entries(void *buf, unsigned long size)\n+static int add_tree_dict_entries(struct dict_table *tree_path_table,\n+\t\t\t\t void *buf, unsigned long size)\n {\n \tstruct tree_desc desc;\n \tstruct name_entry name_entry;\n \n-\tif (!tree_path_table)\n-\t\ttree_path_table = create_dict_table();\n-\n \tinit_tree_desc(&desc, buf, size);\n \twhile (tree_entry(&desc, &name_entry)) {\n \t\tint pathlen = tree_entry_len(&name_entry);\n@@ -659,6 +655,9 @@ static int create_pack_dictionaries(struct packed_git *p,\n \tstruct progress *progress_state;\n \tunsigned int i;\n \n+\tcommit_ident_table = create_dict_table();\n+\ttree_path_table = create_dict_table();\n+\n \tprogress_state = start_progress(\"Scanning objects\", p->num_objects);\n \tfor (i = 0; i < p->num_objects; i++) {\n \t\tstruct pack_idx_entry *obj = obj_list[i];\n@@ -666,7 +665,8 @@ static int create_pack_dictionaries(struct packed_git *p,\n \t\tenum object_type type;\n \t\tunsigned long size;\n \t\tstruct object_info oi = {};\n-\t\tint (*add_dict_entries)(void *, unsigned long);\n+\t\tint (*add_dict_entries)(struct dict_table *, void *, unsigned long);\n+\t\tstruct dict_table *dict;\n \n \t\tdisplay_progress(progress_state, i+1);\n \n@@ -679,9 +679,11 @@ static int create_pack_dictionaries(struct packed_git *p,\n \t\tswitch (type) {\n \t\tcase OBJ_COMMIT:\n \t\t\tadd_dict_entries = add_commit_dict_entries;\n+\t\t\tdict = commit_ident_table;\n \t\t\tbreak;\n \t\tcase OBJ_TREE:\n \t\t\tadd_dict_entries = add_tree_dict_entries;\n+\t\t\tdict = tree_path_table;\n \t\t\tbreak;\n \t\tdefault:\n \t\t\tcontinue;\n@@ -693,7 +695,7 @@ static int create_pack_dictionaries(struct packed_git *p,\n \t\tif (check_sha1_signature(obj->sha1, data, size, typename(type)))\n \t\t\tdie(\"packed %s from %s is corrupt\",\n \t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\t\tif (add_dict_entries(data, size) < 0)\n+\t\tif (add_dict_entries(dict, data, size) < 0)\n \t\t\tdie(\"can't process %s object %s\",\n \t\t\t\ttypename(type), sha1_to_hex(obj->sha1));\n \t\tfree(data);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227216","messageId":"1378735087-4813-3-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 02/16] pack v4: stop using static/global variables in packv4-create.c","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:57:53Z","receivedAt":"2013-09-09T13:57:53Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n packv4-create.c       | 103 ++++++++++++++++++++++++++++----------------------\n packv4-create.h (new) |  11 ++++++\n 2 files changed, 69 insertions(+), 45 deletions(-)\n create mode 100644 packv4-create.h\n\ndiff --git a/packv4-create.c b/packv4-create.c\nindex dbc2a03..920a0b4 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -15,6 +15,7 @@\n #include \"pack-revindex.h\"\n #include \"progress.h\"\n #include \"varint.h\"\n+#include \"packv4-create.h\"\n \n \n static int pack_compression_seen;\n@@ -145,9 +146,6 @@ static void sort_dict_entries_by_hits(struct dict_table *t)\n \trehash_entries(t);\n }\n \n-static struct dict_table *commit_ident_table;\n-static struct dict_table *tree_path_table;\n-\n /*\n  * Parse the author/committer line from a canonical commit object.\n  * The 'from' argument points right after the \"author \" or \"committer \"\n@@ -243,10 +241,10 @@ void dump_dict_table(struct dict_table *t)\n \t}\n }\n \n-static void dict_dump(void)\n+static void dict_dump(struct packv4_tables *v4)\n {\n-\tdump_dict_table(commit_ident_table);\n-\tdump_dict_table(tree_path_table);\n+\tdump_dict_table(v4->commit_ident_table);\n+\tdump_dict_table(v4->tree_path_table);\n }\n \n /*\n@@ -254,10 +252,12 @@ static void dict_dump(void)\n  * pack SHA1 table incremented by 1, or the literal SHA1 value prefixed\n  * with a zero byte if the needed SHA1 is not available in the table.\n  */\n-static struct pack_idx_entry *all_objs;\n-static unsigned all_objs_nr;\n-static int encode_sha1ref(const unsigned char *sha1, unsigned char *buf)\n+\n+int encode_sha1ref(const struct packv4_tables *v4,\n+\t\t   const unsigned char *sha1, unsigned char *buf)\n {\n+\tunsigned all_objs_nr = v4->all_objs_nr;\n+\tstruct pack_idx_entry *all_objs = v4->all_objs;\n \tunsigned lo = 0, hi = all_objs_nr;\n \n \tdo {\n@@ -284,7 +284,8 @@ static int encode_sha1ref(const unsigned char *sha1, unsigned char *buf)\n  * strict so to ensure the canonical version may always be\n  * regenerated and produce the same hash.\n  */\n-void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n+void *pv4_encode_commit(const struct packv4_tables *v4,\n+\t\t\tvoid *buffer, unsigned long *sizep)\n {\n \tunsigned long size = *sizep;\n \tchar *in, *tail, *end;\n@@ -310,7 +311,7 @@ void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n \tif (get_sha1_lowhex(in + 5, sha1) < 0)\n \t\tgoto bad_data;\n \tin += 46;\n-\tout += encode_sha1ref(sha1, out);\n+\tout += encode_sha1ref(v4, sha1, out);\n \n \t/* count how many \"parent\" lines */\n \tnb_parents = 0;\n@@ -325,7 +326,7 @@ void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n \twhile (nb_parents--) {\n \t\tif (get_sha1_lowhex(in + 7, sha1))\n \t\t\tgoto bad_data;\n-\t\tout += encode_sha1ref(sha1, out);\n+\t\tout += encode_sha1ref(v4, sha1, out);\n \t\tin += 48;\n \t}\n \n@@ -337,7 +338,7 @@ void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n \tend = get_nameend_and_tz(in, &tz_val);\n \tif (!end)\n \t\tgoto bad_data;\n-\tauthor_index = dict_add_entry(commit_ident_table, tz_val, in, end - in);\n+\tauthor_index = dict_add_entry(v4->commit_ident_table, tz_val, in, end - in);\n \tif (author_index < 0)\n \t\tgoto bad_dict;\n \tauthor_time = strtoul(end, &end, 10);\n@@ -353,7 +354,7 @@ void *pv4_encode_commit(void *buffer, unsigned long *sizep)\n \tend = get_nameend_and_tz(in, &tz_val);\n \tif (!end)\n \t\tgoto bad_data;\n-\tcommit_index = dict_add_entry(commit_ident_table, tz_val, in, end - in);\n+\tcommit_index = dict_add_entry(v4->commit_ident_table, tz_val, in, end - in);\n \tif (commit_index < 0)\n \t\tgoto bad_dict;\n \tcommit_time = strtoul(end, &end, 10);\n@@ -436,7 +437,8 @@ static int compare_tree_entries(struct name_entry *e1, struct name_entry *e2)\n  * If a delta buffer is provided, we may encode multiple ranges of tree\n  * entries against that buffer.\n  */\n-void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n+void *pv4_encode_tree(const struct packv4_tables *v4,\n+\t\t      void *_buffer, unsigned long *sizep,\n \t\t      void *delta, unsigned long delta_size,\n \t\t      const unsigned char *delta_sha1)\n {\n@@ -551,7 +553,7 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\t\tcp += encode_varint(copy_start, cp);\n \t\t\tcp += encode_varint(copy_count, cp);\n \t\t\tif (first_delta)\n-\t\t\t\tcp += encode_sha1ref(delta_sha1, cp);\n+\t\t\t\tcp += encode_sha1ref(v4, delta_sha1, cp);\n \n \t\t\t/*\n \t\t\t * Now let's make sure this is going to take less\n@@ -577,7 +579,7 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\t}\n \n \t\tpathlen = tree_entry_len(&name_entry);\n-\t\tindex = dict_add_entry(tree_path_table, name_entry.mode,\n+\t\tindex = dict_add_entry(v4->tree_path_table, name_entry.mode,\n \t\t\t\t       name_entry.path, pathlen);\n \t\tif (index < 0) {\n \t\t\terror(\"missing tree dict entry\");\n@@ -585,7 +587,7 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\t\treturn NULL;\n \t\t}\n \t\tout += encode_varint(index << 1, out);\n-\t\tout += encode_sha1ref(name_entry.sha1, out);\n+\t\tout += encode_sha1ref(v4, name_entry.sha1, out);\n \t}\n \n \tif (copy_count) {\n@@ -596,7 +598,7 @@ void *pv4_encode_tree(void *_buffer, unsigned long *sizep,\n \t\tcp += encode_varint(copy_start, cp);\n \t\tcp += encode_varint(copy_count, cp);\n \t\tif (first_delta)\n-\t\t\tcp += encode_sha1ref(delta_sha1, cp);\n+\t\t\tcp += encode_sha1ref(v4, delta_sha1, cp);\n \t\tif (copy_count >= min_tree_copy &&\n \t\t    cp - copy_buf < out - &buffer[copy_pos]) {\n \t\t\tout = buffer + copy_pos;\n@@ -649,14 +651,15 @@ static struct pack_idx_entry **sort_objs_by_offset(struct pack_idx_entry *list,\n \treturn sorted;\n }\n \n-static int create_pack_dictionaries(struct packed_git *p,\n+static int create_pack_dictionaries(struct packv4_tables *v4,\n+\t\t\t\t    struct packed_git *p,\n \t\t\t\t    struct pack_idx_entry **obj_list)\n {\n \tstruct progress *progress_state;\n \tunsigned int i;\n \n-\tcommit_ident_table = create_dict_table();\n-\ttree_path_table = create_dict_table();\n+\tv4->commit_ident_table = create_dict_table();\n+\tv4->tree_path_table = create_dict_table();\n \n \tprogress_state = start_progress(\"Scanning objects\", p->num_objects);\n \tfor (i = 0; i < p->num_objects; i++) {\n@@ -679,11 +682,11 @@ static int create_pack_dictionaries(struct packed_git *p,\n \t\tswitch (type) {\n \t\tcase OBJ_COMMIT:\n \t\t\tadd_dict_entries = add_commit_dict_entries;\n-\t\t\tdict = commit_ident_table;\n+\t\t\tdict = v4->commit_ident_table;\n \t\t\tbreak;\n \t\tcase OBJ_TREE:\n \t\t\tadd_dict_entries = add_tree_dict_entries;\n-\t\t\tdict = tree_path_table;\n+\t\t\tdict = v4->tree_path_table;\n \t\t\tbreak;\n \t\tdefault:\n \t\t\tcontinue;\n@@ -776,9 +779,13 @@ static unsigned int packv4_write_header(struct sha1file *f, unsigned nr_objects)\n \treturn sizeof(hdr);\n }\n \n-static unsigned long packv4_write_tables(struct sha1file *f, unsigned nr_objects,\n-\t\t\t\t\t struct pack_idx_entry *objs)\n+unsigned long packv4_write_tables(struct sha1file *f,\n+\t\t\t\t  const struct packv4_tables *v4)\n {\n+\tunsigned nr_objects = v4->all_objs_nr;\n+\tstruct pack_idx_entry *objs = v4->all_objs;\n+\tstruct dict_table *commit_ident_table = v4->commit_ident_table;\n+\tstruct dict_table *tree_path_table = v4->tree_path_table;\n \tunsigned i;\n \tunsigned long written = 0;\n \n@@ -823,7 +830,8 @@ static int write_object_header(struct sha1file *f, enum object_type type, unsign\n \treturn len;\n }\n \n-static unsigned long copy_object_data(struct sha1file *f, struct packed_git *p,\n+static unsigned long copy_object_data(struct packv4_tables *v4,\n+\t\t\t\t      struct sha1file *f, struct packed_git *p,\n \t\t\t\t      off_t offset)\n {\n \tstruct pack_window *w_curs = NULL;\n@@ -850,11 +858,13 @@ static unsigned long copy_object_data(struct sha1file *f, struct packed_git *p,\n \t\tif (base_offset <= 0 || base_offset >= offset)\n \t\t\tdie(\"delta offset out of bound\");\n \t\trevidx = find_pack_revindex(p, base_offset);\n-\t\treflen = encode_sha1ref(nth_packed_object_sha1(p, revidx->nr), buf);\n+\t\treflen = encode_sha1ref(v4,\n+\t\t\t\t\tnth_packed_object_sha1(p, revidx->nr),\n+\t\t\t\t\tbuf);\n \t\tsha1write(f, buf, reflen);\n \t\twritten += reflen;\n \t} else if (type == OBJ_REF_DELTA) {\n-\t\treflen = encode_sha1ref(src + hdrlen, buf);\n+\t\treflen = encode_sha1ref(v4, src + hdrlen, buf);\n \t\thdrlen += 20;\n \t\tsha1write(f, buf, reflen);\n \t\twritten += reflen;\n@@ -919,7 +929,8 @@ static unsigned char *get_delta_base(struct packed_git *p, off_t offset,\n \treturn sha1_buf;\n }\n \n-static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n+static off_t packv4_write_object(struct packv4_tables *v4,\n+\t\t\t\t struct sha1file *f, struct packed_git *p,\n \t\t\t\t struct pack_idx_entry *obj)\n {\n \tvoid *src, *result;\n@@ -941,7 +952,7 @@ static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \tcase OBJ_TREE:\n \t\tbreak;\n \tdefault:\n-\t\treturn copy_object_data(f, p, obj->offset);\n+\t\treturn copy_object_data(v4, f, p, obj->offset);\n \t}\n \n \t/* The rest is converted into their new format */\n@@ -955,7 +966,7 @@ static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \n \tswitch (type) {\n \tcase OBJ_COMMIT:\n-\t\tresult = pv4_encode_commit(src, &buf_size);\n+\t\tresult = pv4_encode_commit(v4, src, &buf_size);\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\tif (packed_type != OBJ_TREE) {\n@@ -972,11 +983,12 @@ static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \t\t\tif (!ref || ref_type != OBJ_TREE)\n \t\t\t\tdie(\"cannot obtain delta base for %s\",\n \t\t\t\t\t\tsha1_to_hex(obj->sha1));\n-\t\t\tresult = pv4_encode_tree(src, &buf_size,\n+\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n \t\t\t\t\t\t ref, ref_size, ref_sha1);\n \t\t\tfree(ref);\n \t\t} else {\n-\t\t\tresult = pv4_encode_tree(src, &buf_size, NULL, 0, NULL);\n+\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n+\t\t\t\t\t\t NULL, 0, NULL);\n \t\t}\n \t\tbreak;\n \tdefault:\n@@ -987,7 +999,7 @@ static off_t packv4_write_object(struct sha1file *f, struct packed_git *p,\n \t\twarning(\"can't convert %s object %s\",\n \t\t\ttypename(type), sha1_to_hex(obj->sha1));\n \t\t/* fall back to copy the object in its original form */\n-\t\treturn copy_object_data(f, p, obj->offset);\n+\t\treturn copy_object_data(v4, f, p, obj->offset);\n \t}\n \n \t/* Use bit 3 to indicate a special type encoding */\n@@ -1041,7 +1053,7 @@ static struct packed_git *open_pack(const char *path)\n \treturn p;\n }\n \n-static void process_one_pack(char *src_pack, char *dst_pack)\n+static void process_one_pack(struct packv4_tables *v4, char *src_pack, char *dst_pack)\n {\n \tstruct packed_git *p;\n \tstruct sha1file *f;\n@@ -1061,26 +1073,26 @@ static void process_one_pack(char *src_pack, char *dst_pack)\n \tobjs = get_packed_object_list(p);\n \tp_objs = sort_objs_by_offset(objs, nr_objects);\n \n-\tcreate_pack_dictionaries(p, p_objs);\n-\tsort_dict_entries_by_hits(commit_ident_table);\n-\tsort_dict_entries_by_hits(tree_path_table);\n+\tcreate_pack_dictionaries(v4, p, p_objs);\n+\tsort_dict_entries_by_hits(v4->commit_ident_table);\n+\tsort_dict_entries_by_hits(v4->tree_path_table);\n \n \tpackname = normalize_pack_name(dst_pack);\n \tf = packv4_open(packname);\n \tif (!f)\n \t\tdie(\"unable to open destination pack\");\n \twritten += packv4_write_header(f, nr_objects);\n-\twritten += packv4_write_tables(f, nr_objects, objs);\n+\twritten += packv4_write_tables(f, v4);\n \n \t/* Let's write objects out, updating the object index list in place */\n \tprogress_state = start_progress(\"Writing objects\", nr_objects);\n-\tall_objs = objs;\n-\tall_objs_nr = nr_objects;\n+\tv4->all_objs = objs;\n+\tv4->all_objs_nr = nr_objects;\n \tfor (i = 0; i < nr_objects; i++) {\n \t\toff_t obj_pos = written;\n \t\tstruct pack_idx_entry *obj = p_objs[i];\n \t\tcrc32_begin(f);\n-\t\twritten += packv4_write_object(f, p, obj);\n+\t\twritten += packv4_write_object(v4, f, p, obj);\n \t\tobj->offset = obj_pos;\n \t\tobj->crc32 = crc32_end(f);\n \t\tdisplay_progress(progress_state, i+1);\n@@ -1114,6 +1126,7 @@ static int git_pack_config(const char *k, const char *v, void *cb)\n \n int main(int argc, char *argv[])\n {\n+\tstruct packv4_tables v4;\n \tchar *src_pack, *dst_pack;\n \n \tif (argc == 3) {\n@@ -1131,8 +1144,8 @@ int main(int argc, char *argv[])\n \tgit_config(git_pack_config, NULL);\n \tif (!pack_compression_seen && core_compression_seen)\n \t\tpack_compression_level = core_compression_level;\n-\tprocess_one_pack(src_pack, dst_pack);\n+\tprocess_one_pack(&v4, src_pack, dst_pack);\n \tif (0)\n-\t\tdict_dump();\n+\t\tdict_dump(&v4);\n \treturn 0;\n }\ndiff --git a/packv4-create.h b/packv4-create.h\nnew file mode 100644\nindex 0000000..0c8c77b\n--- /dev/null\n+++ b/packv4-create.h\n@@ -0,0 +1,11 @@\n+#ifndef PACKV4_CREATE_H\n+#define PACKV4_CREATE_H\n+\n+struct packv4_tables {\n+\tstruct pack_idx_entry *all_objs;\n+\tunsigned all_objs_nr;\n+\tstruct dict_table *commit_ident_table;\n+\tstruct dict_table *tree_path_table;\n+};\n+\n+#endif\n-- \n1.8.2.83.gc99314b\n"},{"id":"227215","messageId":"1378735087-4813-4-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 03/16] pack v4: move packv4-create.c to libgit.a","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:57:54Z","receivedAt":"2013-09-09T13:57:54Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"git-packv4-create now becomes test-packv4. Code that will not be used\nby pack-objects.c is moved to test-packv4.c. It may be removed when\nthe code transition to pack-objects completes.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n Makefile            |   4 +-\n packv4-create.c     | 480 +---------------------------------------------------\n packv4-create.h     |  28 +++\n test-packv4.c (new) | 476 +++++++++++++++++++++++++++++++++++++++++++++++++++\n 4 files changed, 514 insertions(+), 474 deletions(-)\n create mode 100644 test-packv4.c\n\ndiff --git a/Makefile b/Makefile\nindex 22fc276..af2e3e3 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -550,7 +550,6 @@ PROGRAM_OBJS += shell.o\n PROGRAM_OBJS += show-index.o\n PROGRAM_OBJS += upload-pack.o\n PROGRAM_OBJS += remote-testsvn.o\n-PROGRAM_OBJS += packv4-create.o\n \n # Binary suffix, set to .exe for Windows builds\n X =\n@@ -568,6 +567,7 @@ TEST_PROGRAMS_NEED_X += test-line-buffer\n TEST_PROGRAMS_NEED_X += test-match-trees\n TEST_PROGRAMS_NEED_X += test-mergesort\n TEST_PROGRAMS_NEED_X += test-mktemp\n+TEST_PROGRAMS_NEED_X += test-packv4\n TEST_PROGRAMS_NEED_X += test-parse-options\n TEST_PROGRAMS_NEED_X += test-path-utils\n TEST_PROGRAMS_NEED_X += test-prio-queue\n@@ -702,6 +702,7 @@ LIB_H += notes.h\n LIB_H += object.h\n LIB_H += pack-revindex.h\n LIB_H += pack.h\n+LIB_H += packv4-create.h\n LIB_H += packv4-parse.h\n LIB_H += parse-options.h\n LIB_H += patch-ids.h\n@@ -839,6 +840,7 @@ LIB_OBJS += object.o\n LIB_OBJS += pack-check.o\n LIB_OBJS += pack-revindex.o\n LIB_OBJS += pack-write.o\n+LIB_OBJS += packv4-create.o\n LIB_OBJS += packv4-parse.o\n LIB_OBJS += pager.o\n LIB_OBJS += parse-options.o\ndiff --git a/packv4-create.c b/packv4-create.c\nindex 920a0b4..83a6336 100644\n--- a/packv4-create.c\n+++ b/packv4-create.c\n@@ -18,9 +18,9 @@\n #include \"packv4-create.h\"\n \n \n-static int pack_compression_seen;\n-static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n-static int min_tree_copy = 1;\n+int pack_compression_seen;\n+int pack_compression_level = Z_DEFAULT_COMPRESSION;\n+int min_tree_copy = 1;\n \n struct data_entry {\n \tunsigned offset;\n@@ -139,7 +139,7 @@ static int cmp_dict_entries(const void *a_, const void *b_)\n \treturn diff;\n }\n \n-static void sort_dict_entries_by_hits(struct dict_table *t)\n+void sort_dict_entries_by_hits(struct dict_table *t)\n {\n \tqsort(t->entry, t->nb_entries, sizeof(*t->entry), cmp_dict_entries);\n \tt->hash_size = (t->nb_entries * 4 / 3) / 2;\n@@ -208,7 +208,7 @@ int add_commit_dict_entries(struct dict_table *commit_ident_table,\n \treturn 0;\n }\n \n-static int add_tree_dict_entries(struct dict_table *tree_path_table,\n+int add_tree_dict_entries(struct dict_table *tree_path_table,\n \t\t\t\t void *buf, unsigned long size)\n {\n \tstruct tree_desc desc;\n@@ -224,7 +224,7 @@ static int add_tree_dict_entries(struct dict_table *tree_path_table,\n \treturn 0;\n }\n \n-void dump_dict_table(struct dict_table *t)\n+static void dump_dict_table(struct dict_table *t)\n {\n \tint i;\n \n@@ -241,7 +241,7 @@ void dump_dict_table(struct dict_table *t)\n \t}\n }\n \n-static void dict_dump(struct packv4_tables *v4)\n+void dict_dump(struct packv4_tables *v4)\n {\n \tdump_dict_table(v4->commit_ident_table);\n \tdump_dict_table(v4->tree_path_table);\n@@ -611,103 +611,6 @@ void *pv4_encode_tree(const struct packv4_tables *v4,\n \treturn buffer;\n }\n \n-static struct pack_idx_entry *get_packed_object_list(struct packed_git *p)\n-{\n-\tunsigned i, nr_objects = p->num_objects;\n-\tstruct pack_idx_entry *objects;\n-\n-\tobjects = xmalloc((nr_objects + 1) * sizeof(*objects));\n-\tobjects[nr_objects].offset = p->pack_size - 20;\n-\tfor (i = 0; i < nr_objects; i++) {\n-\t\thashcpy(objects[i].sha1, nth_packed_object_sha1(p, i));\n-\t\tobjects[i].offset = nth_packed_object_offset(p, i);\n-\t}\n-\n-\treturn objects;\n-}\n-\n-static int sort_by_offset(const void *e1, const void *e2)\n-{\n-\tconst struct pack_idx_entry * const *entry1 = e1;\n-\tconst struct pack_idx_entry * const *entry2 = e2;\n-\tif ((*entry1)->offset < (*entry2)->offset)\n-\t\treturn -1;\n-\tif ((*entry1)->offset > (*entry2)->offset)\n-\t\treturn 1;\n-\treturn 0;\n-}\n-\n-static struct pack_idx_entry **sort_objs_by_offset(struct pack_idx_entry *list,\n-\t\t\t\t\t\t    unsigned nr_objects)\n-{\n-\tunsigned i;\n-\tstruct pack_idx_entry **sorted;\n-\n-\tsorted = xmalloc((nr_objects + 1) * sizeof(*sorted));\n-\tfor (i = 0; i < nr_objects + 1; i++)\n-\t\tsorted[i] = &list[i];\n-\tqsort(sorted, nr_objects + 1, sizeof(*sorted), sort_by_offset);\n-\n-\treturn sorted;\n-}\n-\n-static int create_pack_dictionaries(struct packv4_tables *v4,\n-\t\t\t\t    struct packed_git *p,\n-\t\t\t\t    struct pack_idx_entry **obj_list)\n-{\n-\tstruct progress *progress_state;\n-\tunsigned int i;\n-\n-\tv4->commit_ident_table = create_dict_table();\n-\tv4->tree_path_table = create_dict_table();\n-\n-\tprogress_state = start_progress(\"Scanning objects\", p->num_objects);\n-\tfor (i = 0; i < p->num_objects; i++) {\n-\t\tstruct pack_idx_entry *obj = obj_list[i];\n-\t\tvoid *data;\n-\t\tenum object_type type;\n-\t\tunsigned long size;\n-\t\tstruct object_info oi = {};\n-\t\tint (*add_dict_entries)(struct dict_table *, void *, unsigned long);\n-\t\tstruct dict_table *dict;\n-\n-\t\tdisplay_progress(progress_state, i+1);\n-\n-\t\toi.typep = &type;\n-\t\toi.sizep = &size;\n-\t\tif (packed_object_info(p, obj->offset, &oi) < 0)\n-\t\t\tdie(\"cannot get type of %s from %s\",\n-\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\n-\t\tswitch (type) {\n-\t\tcase OBJ_COMMIT:\n-\t\t\tadd_dict_entries = add_commit_dict_entries;\n-\t\t\tdict = v4->commit_ident_table;\n-\t\t\tbreak;\n-\t\tcase OBJ_TREE:\n-\t\t\tadd_dict_entries = add_tree_dict_entries;\n-\t\t\tdict = v4->tree_path_table;\n-\t\t\tbreak;\n-\t\tdefault:\n-\t\t\tcontinue;\n-\t\t}\n-\t\tdata = unpack_entry(p, obj->offset, &type, &size);\n-\t\tif (!data)\n-\t\t\tdie(\"cannot unpack %s from %s\",\n-\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\t\tif (check_sha1_signature(obj->sha1, data, size, typename(type)))\n-\t\t\tdie(\"packed %s from %s is corrupt\",\n-\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\t\tif (add_dict_entries(dict, data, size) < 0)\n-\t\t\tdie(\"can't process %s object %s\",\n-\t\t\t\ttypename(type), sha1_to_hex(obj->sha1));\n-\t\tfree(data);\n-\t}\n-\n-\tstop_progress(&progress_state);\n-\treturn 0;\n-}\n-\n static unsigned long write_dict_table(struct sha1file *f, struct dict_table *t)\n {\n \tunsigned char buffer[1024];\n@@ -757,28 +660,6 @@ static unsigned long write_dict_table(struct sha1file *f, struct dict_table *t)\n \treturn hdrlen + datalen;\n }\n \n-static struct sha1file * packv4_open(char *path)\n-{\n-\tint fd;\n-\n-\tfd = open(path, O_CREAT|O_EXCL|O_WRONLY, 0600);\n-\tif (fd < 0)\n-\t\tdie_errno(\"unable to create '%s'\", path);\n-\treturn sha1fd(fd, path);\n-}\n-\n-static unsigned int packv4_write_header(struct sha1file *f, unsigned nr_objects)\n-{\n-\tstruct pack_header hdr;\n-\n-\thdr.hdr_signature = htonl(PACK_SIGNATURE);\n-\thdr.hdr_version = htonl(4);\n-\thdr.hdr_entries = htonl(nr_objects);\n-\tsha1write(f, &hdr, sizeof(hdr));\n-\n-\treturn sizeof(hdr);\n-}\n-\n unsigned long packv4_write_tables(struct sha1file *f,\n \t\t\t\t  const struct packv4_tables *v4)\n {\n@@ -802,350 +683,3 @@ unsigned long packv4_write_tables(struct sha1file *f,\n \n \treturn written;\n }\n-\n-static int write_object_header(struct sha1file *f, enum object_type type, unsigned long size)\n-{\n-\tunsigned char buf[16];\n-\tuint64_t val;\n-\tint len;\n-\n-\t/*\n-\t * We really have only one kind of delta object.\n-\t */\n-\tif (type == OBJ_OFS_DELTA)\n-\t\ttype = OBJ_REF_DELTA;\n-\n-\t/*\n-\t * We allocate 4 bits in the LSB for the object type which should\n-\t * be good for quite a while, given that we effectively encodes\n-\t * only 5 object types: commit, tree, blob, delta, tag.\n-\t */\n-\tval = size;\n-\tif (MSB(val, 4))\n-\t\tdie(\"fixme: the code doesn't currently cope with big sizes\");\n-\tval <<= 4;\n-\tval |= type;\n-\tlen = encode_varint(val, buf);\n-\tsha1write(f, buf, len);\n-\treturn len;\n-}\n-\n-static unsigned long copy_object_data(struct packv4_tables *v4,\n-\t\t\t\t      struct sha1file *f, struct packed_git *p,\n-\t\t\t\t      off_t offset)\n-{\n-\tstruct pack_window *w_curs = NULL;\n-\tstruct revindex_entry *revidx;\n-\tenum object_type type;\n-\tunsigned long avail, size, datalen, written;\n-\tint hdrlen, reflen, idx_nr;\n-\tunsigned char *src, buf[24];\n-\n-\trevidx = find_pack_revindex(p, offset);\n-\tidx_nr = revidx->nr;\n-\tdatalen = revidx[1].offset - offset;\n-\n-\tsrc = use_pack(p, &w_curs, offset, &avail);\n-\thdrlen = unpack_object_header_buffer(src, avail, &type, &size);\n-\n-\twritten = write_object_header(f, type, size);\n-\n-\tif (type == OBJ_OFS_DELTA) {\n-\t\tconst unsigned char *cp = src + hdrlen;\n-\t\toff_t base_offset = decode_varint(&cp);\n-\t\thdrlen = cp - src;\n-\t\tbase_offset = offset - base_offset;\n-\t\tif (base_offset <= 0 || base_offset >= offset)\n-\t\t\tdie(\"delta offset out of bound\");\n-\t\trevidx = find_pack_revindex(p, base_offset);\n-\t\treflen = encode_sha1ref(v4,\n-\t\t\t\t\tnth_packed_object_sha1(p, revidx->nr),\n-\t\t\t\t\tbuf);\n-\t\tsha1write(f, buf, reflen);\n-\t\twritten += reflen;\n-\t} else if (type == OBJ_REF_DELTA) {\n-\t\treflen = encode_sha1ref(v4, src + hdrlen, buf);\n-\t\thdrlen += 20;\n-\t\tsha1write(f, buf, reflen);\n-\t\twritten += reflen;\n-\t}\n-\n-\tif (p->index_version > 1 &&\n-\t    check_pack_crc(p, &w_curs, offset, datalen, idx_nr))\n-\t\tdie(\"bad CRC for object at offset %\"PRIuMAX\" in %s\",\n-\t\t    (uintmax_t)offset, p->pack_name);\n-\n-\toffset += hdrlen;\n-\tdatalen -= hdrlen;\n-\n-\twhile (datalen) {\n-\t\tsrc = use_pack(p, &w_curs, offset, &avail);\n-\t\tif (avail > datalen)\n-\t\t\tavail = datalen;\n-\t\tsha1write(f, src, avail);\n-\t\twritten += avail;\n-\t\toffset += avail;\n-\t\tdatalen -= avail;\n-\t}\n-\tunuse_pack(&w_curs);\n-\n-\treturn written;\n-}\n-\n-static unsigned char *get_delta_base(struct packed_git *p, off_t offset,\n-\t\t\t\t     unsigned char *sha1_buf)\n-{\n-\tstruct pack_window *w_curs = NULL;\n-\tenum object_type type;\n-\tunsigned long avail, size;\n-\tint hdrlen;\n-\tunsigned char *src;\n-\tconst unsigned char *base_sha1 = NULL; ;\n-\n-\tsrc = use_pack(p, &w_curs, offset, &avail);\n-\thdrlen = unpack_object_header_buffer(src, avail, &type, &size);\n-\n-\tif (type == OBJ_OFS_DELTA) {\n-\t\tconst unsigned char *cp = src + hdrlen;\n-\t\toff_t base_offset = decode_varint(&cp);\n-\t\tbase_offset = offset - base_offset;\n-\t\tif (base_offset <= 0 || base_offset >= offset) {\n-\t\t\terror(\"delta offset out of bound\");\n-\t\t} else {\n-\t\t\tstruct revindex_entry *revidx;\n-\t\t\trevidx = find_pack_revindex(p, base_offset);\n-\t\t\tbase_sha1 = nth_packed_object_sha1(p, revidx->nr);\n-\t\t}\n-\t} else if (type == OBJ_REF_DELTA) {\n-\t\tbase_sha1 = src + hdrlen;\n-\t} else\n-\t\terror(\"expected to get a delta but got a %s\", typename(type));\n-\n-\tunuse_pack(&w_curs);\n-\n-\tif (!base_sha1)\n-\t\treturn NULL;\n-\thashcpy(sha1_buf, base_sha1);\n-\treturn sha1_buf;\n-}\n-\n-static off_t packv4_write_object(struct packv4_tables *v4,\n-\t\t\t\t struct sha1file *f, struct packed_git *p,\n-\t\t\t\t struct pack_idx_entry *obj)\n-{\n-\tvoid *src, *result;\n-\tstruct object_info oi = {};\n-\tenum object_type type, packed_type;\n-\tunsigned long obj_size, buf_size;\n-\tunsigned int hdrlen;\n-\n-\toi.typep = &type;\n-\toi.sizep = &obj_size;\n-\tpacked_type = packed_object_info(p, obj->offset, &oi);\n-\tif (packed_type < 0)\n-\t\tdie(\"cannot get type of %s from %s\",\n-\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\n-\t/* Some objects are copied without decompression */\n-\tswitch (type) {\n-\tcase OBJ_COMMIT:\n-\tcase OBJ_TREE:\n-\t\tbreak;\n-\tdefault:\n-\t\treturn copy_object_data(v4, f, p, obj->offset);\n-\t}\n-\n-\t/* The rest is converted into their new format */\n-\tsrc = unpack_entry(p, obj->offset, &type, &buf_size);\n-\tif (!src || obj_size != buf_size)\n-\t\tdie(\"cannot unpack %s from %s\",\n-\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\tif (check_sha1_signature(obj->sha1, src, buf_size, typename(type)))\n-\t\tdie(\"packed %s from %s is corrupt\",\n-\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n-\n-\tswitch (type) {\n-\tcase OBJ_COMMIT:\n-\t\tresult = pv4_encode_commit(v4, src, &buf_size);\n-\t\tbreak;\n-\tcase OBJ_TREE:\n-\t\tif (packed_type != OBJ_TREE) {\n-\t\t\tunsigned char sha1_buf[20], *ref_sha1;\n-\t\t\tvoid *ref;\n-\t\t\tenum object_type ref_type;\n-\t\t\tunsigned long ref_size;\n-\n-\t\t\tref_sha1 = get_delta_base(p, obj->offset, sha1_buf);\n-\t\t\tif (!ref_sha1)\n-\t\t\t\tdie(\"unable to get delta base sha1 for %s\",\n-\t\t\t\t\t\tsha1_to_hex(obj->sha1));\n-\t\t\tref = read_sha1_file(ref_sha1, &ref_type, &ref_size);\n-\t\t\tif (!ref || ref_type != OBJ_TREE)\n-\t\t\t\tdie(\"cannot obtain delta base for %s\",\n-\t\t\t\t\t\tsha1_to_hex(obj->sha1));\n-\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n-\t\t\t\t\t\t ref, ref_size, ref_sha1);\n-\t\t\tfree(ref);\n-\t\t} else {\n-\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n-\t\t\t\t\t\t NULL, 0, NULL);\n-\t\t}\n-\t\tbreak;\n-\tdefault:\n-\t\tdie(\"unexpected object type %d\", type);\n-\t}\n-\tfree(src);\n-\tif (!result) {\n-\t\twarning(\"can't convert %s object %s\",\n-\t\t\ttypename(type), sha1_to_hex(obj->sha1));\n-\t\t/* fall back to copy the object in its original form */\n-\t\treturn copy_object_data(v4, f, p, obj->offset);\n-\t}\n-\n-\t/* Use bit 3 to indicate a special type encoding */\n-\ttype += 8;\n-\thdrlen = write_object_header(f, type, obj_size);\n-\tsha1write(f, result, buf_size);\n-\tfree(result);\n-\treturn hdrlen + buf_size;\n-}\n-\n-static char *normalize_pack_name(const char *path)\n-{\n-\tchar buf[PATH_MAX];\n-\tint len;\n-\n-\tlen = strlcpy(buf, path, PATH_MAX);\n-\tif (len >= PATH_MAX - 6)\n-\t\tdie(\"name too long: %s\", path);\n-\n-\t/*\n-\t * In addition to \"foo.idx\" we accept \"foo.pack\" and \"foo\";\n-\t * normalize these forms to \"foo.pack\".\n-\t */\n-\tif (has_extension(buf, \".idx\")) {\n-\t\tstrcpy(buf + len - 4, \".pack\");\n-\t\tlen++;\n-\t} else if (!has_extension(buf, \".pack\")) {\n-\t\tstrcpy(buf + len, \".pack\");\n-\t\tlen += 5;\n-\t}\n-\n-\treturn xstrdup(buf);\n-}\n-\n-static struct packed_git *open_pack(const char *path)\n-{\n-\tchar *packname = normalize_pack_name(path);\n-\tint len = strlen(packname);\n-\tstruct packed_git *p;\n-\n-\tstrcpy(packname + len - 5, \".idx\");\n-\tp = add_packed_git(packname, len - 1, 1);\n-\tif (!p)\n-\t\tdie(\"packfile %s not found.\", packname);\n-\n-\tinstall_packed_git(p);\n-\tif (open_pack_index(p))\n-\t\tdie(\"packfile %s index not opened\", p->pack_name);\n-\n-\tfree(packname);\n-\treturn p;\n-}\n-\n-static void process_one_pack(struct packv4_tables *v4, char *src_pack, char *dst_pack)\n-{\n-\tstruct packed_git *p;\n-\tstruct sha1file *f;\n-\tstruct pack_idx_entry *objs, **p_objs;\n-\tstruct pack_idx_option idx_opts;\n-\tunsigned i, nr_objects;\n-\toff_t written = 0;\n-\tchar *packname;\n-\tunsigned char pack_sha1[20];\n-\tstruct progress *progress_state;\n-\n-\tp = open_pack(src_pack);\n-\tif (!p)\n-\t\tdie(\"unable to open source pack\");\n-\n-\tnr_objects = p->num_objects;\n-\tobjs = get_packed_object_list(p);\n-\tp_objs = sort_objs_by_offset(objs, nr_objects);\n-\n-\tcreate_pack_dictionaries(v4, p, p_objs);\n-\tsort_dict_entries_by_hits(v4->commit_ident_table);\n-\tsort_dict_entries_by_hits(v4->tree_path_table);\n-\n-\tpackname = normalize_pack_name(dst_pack);\n-\tf = packv4_open(packname);\n-\tif (!f)\n-\t\tdie(\"unable to open destination pack\");\n-\twritten += packv4_write_header(f, nr_objects);\n-\twritten += packv4_write_tables(f, v4);\n-\n-\t/* Let's write objects out, updating the object index list in place */\n-\tprogress_state = start_progress(\"Writing objects\", nr_objects);\n-\tv4->all_objs = objs;\n-\tv4->all_objs_nr = nr_objects;\n-\tfor (i = 0; i < nr_objects; i++) {\n-\t\toff_t obj_pos = written;\n-\t\tstruct pack_idx_entry *obj = p_objs[i];\n-\t\tcrc32_begin(f);\n-\t\twritten += packv4_write_object(v4, f, p, obj);\n-\t\tobj->offset = obj_pos;\n-\t\tobj->crc32 = crc32_end(f);\n-\t\tdisplay_progress(progress_state, i+1);\n-\t}\n-\tstop_progress(&progress_state);\n-\n-\tsha1close(f, pack_sha1, CSUM_CLOSE | CSUM_FSYNC);\n-\n-\treset_pack_idx_option(&idx_opts);\n-\tidx_opts.version = 3;\n-\tstrcpy(packname + strlen(packname) - 5, \".idx\");\n-\twrite_idx_file(packname, p_objs, nr_objects, &idx_opts, pack_sha1);\n-\n-\tfree(packname);\n-}\n-\n-static int git_pack_config(const char *k, const char *v, void *cb)\n-{\n-\tif (!strcmp(k, \"pack.compression\")) {\n-\t\tint level = git_config_int(k, v);\n-\t\tif (level == -1)\n-\t\t\tlevel = Z_DEFAULT_COMPRESSION;\n-\t\telse if (level < 0 || level > Z_BEST_COMPRESSION)\n-\t\t\tdie(\"bad pack compression level %d\", level);\n-\t\tpack_compression_level = level;\n-\t\tpack_compression_seen = 1;\n-\t\treturn 0;\n-\t}\n-\treturn git_default_config(k, v, cb);\n-}\n-\n-int main(int argc, char *argv[])\n-{\n-\tstruct packv4_tables v4;\n-\tchar *src_pack, *dst_pack;\n-\n-\tif (argc == 3) {\n-\t\tsrc_pack = argv[1];\n-\t\tdst_pack = argv[2];\n-\t} else if (argc == 4 && !prefixcmp(argv[1], \"--min-tree-copy=\")) {\n-\t\tmin_tree_copy = atoi(argv[1] + strlen(\"--min-tree-copy=\"));\n-\t\tsrc_pack = argv[2];\n-\t\tdst_pack = argv[3];\n-\t} else {\n-\t\tfprintf(stderr, \"Usage: %s [--min-tree-copy=<n>] <src_packfile> <dst_packfile>\\n\", argv[0]);\n-\t\texit(1);\n-\t}\n-\n-\tgit_config(git_pack_config, NULL);\n-\tif (!pack_compression_seen && core_compression_seen)\n-\t\tpack_compression_level = core_compression_level;\n-\tprocess_one_pack(&v4, src_pack, dst_pack);\n-\tif (0)\n-\t\tdict_dump(&v4);\n-\treturn 0;\n-}\ndiff --git a/packv4-create.h b/packv4-create.h\nindex 0c8c77b..ba4929a 100644\n--- a/packv4-create.h\n+++ b/packv4-create.h\n@@ -8,4 +8,32 @@ struct packv4_tables {\n \tstruct dict_table *tree_path_table;\n };\n \n+struct dict_table;\n+struct sha1file;\n+\n+struct dict_table *create_dict_table(void);\n+int dict_add_entry(struct dict_table *t, int val, const char *str, int str_len);\n+void destroy_dict_table(struct dict_table *t);\n+void dict_dump(struct packv4_tables *v4);\n+\n+int add_commit_dict_entries(struct dict_table *commit_ident_table,\n+\t\t\t    void *buf, unsigned long size);\n+int add_tree_dict_entries(struct dict_table *tree_path_table,\n+\t\t\t  void *buf, unsigned long size);\n+void sort_dict_entries_by_hits(struct dict_table *t);\n+\n+int encode_sha1ref(const struct packv4_tables *v4,\n+\t\t   const unsigned char *sha1, unsigned char *buf);\n+unsigned long packv4_write_tables(struct sha1file *f,\n+\t\t\t\t  const struct packv4_tables *v4);\n+void *pv4_encode_commit(const struct packv4_tables *v4,\n+\t\t\tvoid *buffer, unsigned long *sizep);\n+void *pv4_encode_tree(const struct packv4_tables *v4,\n+\t\t      void *_buffer, unsigned long *sizep,\n+\t\t      void *delta, unsigned long delta_size,\n+\t\t      const unsigned char *delta_sha1);\n+\n+void process_one_pack(struct packv4_tables *v4,\n+\t\t      char *src_pack, char *dst_pack);\n+\n #endif\ndiff --git a/test-packv4.c b/test-packv4.c\nnew file mode 100644\nindex 0000000..3b0d7a2\n--- /dev/null\n+++ b/test-packv4.c\n@@ -0,0 +1,476 @@\n+#include \"cache.h\"\n+#include \"pack.h\"\n+#include \"pack-revindex.h\"\n+#include \"progress.h\"\n+#include \"varint.h\"\n+#include \"packv4-create.h\"\n+\n+extern int pack_compression_seen;\n+extern int pack_compression_level;\n+extern int min_tree_copy;\n+\n+static struct pack_idx_entry *get_packed_object_list(struct packed_git *p)\n+{\n+\tunsigned i, nr_objects = p->num_objects;\n+\tstruct pack_idx_entry *objects;\n+\n+\tobjects = xmalloc((nr_objects + 1) * sizeof(*objects));\n+\tobjects[nr_objects].offset = p->pack_size - 20;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\thashcpy(objects[i].sha1, nth_packed_object_sha1(p, i));\n+\t\tobjects[i].offset = nth_packed_object_offset(p, i);\n+\t}\n+\n+\treturn objects;\n+}\n+\n+static int sort_by_offset(const void *e1, const void *e2)\n+{\n+\tconst struct pack_idx_entry * const *entry1 = e1;\n+\tconst struct pack_idx_entry * const *entry2 = e2;\n+\tif ((*entry1)->offset < (*entry2)->offset)\n+\t\treturn -1;\n+\tif ((*entry1)->offset > (*entry2)->offset)\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n+static struct pack_idx_entry **sort_objs_by_offset(struct pack_idx_entry *list,\n+\t\t\t\t\t\t    unsigned nr_objects)\n+{\n+\tunsigned i;\n+\tstruct pack_idx_entry **sorted;\n+\n+\tsorted = xmalloc((nr_objects + 1) * sizeof(*sorted));\n+\tfor (i = 0; i < nr_objects + 1; i++)\n+\t\tsorted[i] = &list[i];\n+\tqsort(sorted, nr_objects + 1, sizeof(*sorted), sort_by_offset);\n+\n+\treturn sorted;\n+}\n+\n+static int create_pack_dictionaries(struct packv4_tables *v4,\n+\t\t\t\t    struct packed_git *p,\n+\t\t\t\t    struct pack_idx_entry **obj_list)\n+{\n+\tstruct progress *progress_state;\n+\tunsigned int i;\n+\n+\tv4->commit_ident_table = create_dict_table();\n+\tv4->tree_path_table = create_dict_table();\n+\n+\tprogress_state = start_progress(\"Scanning objects\", p->num_objects);\n+\tfor (i = 0; i < p->num_objects; i++) {\n+\t\tstruct pack_idx_entry *obj = obj_list[i];\n+\t\tvoid *data;\n+\t\tenum object_type type;\n+\t\tunsigned long size;\n+\t\tstruct object_info oi = {};\n+\t\tint (*add_dict_entries)(struct dict_table *, void *, unsigned long);\n+\t\tstruct dict_table *dict;\n+\n+\t\tdisplay_progress(progress_state, i+1);\n+\n+\t\toi.typep = &type;\n+\t\toi.sizep = &size;\n+\t\tif (packed_object_info(p, obj->offset, &oi) < 0)\n+\t\t\tdie(\"cannot get type of %s from %s\",\n+\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\n+\t\tswitch (type) {\n+\t\tcase OBJ_COMMIT:\n+\t\t\tadd_dict_entries = add_commit_dict_entries;\n+\t\t\tdict = v4->commit_ident_table;\n+\t\t\tbreak;\n+\t\tcase OBJ_TREE:\n+\t\t\tadd_dict_entries = add_tree_dict_entries;\n+\t\t\tdict = v4->tree_path_table;\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tcontinue;\n+\t\t}\n+\t\tdata = unpack_entry(p, obj->offset, &type, &size);\n+\t\tif (!data)\n+\t\t\tdie(\"cannot unpack %s from %s\",\n+\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\t\tif (check_sha1_signature(obj->sha1, data, size, typename(type)))\n+\t\t\tdie(\"packed %s from %s is corrupt\",\n+\t\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\t\tif (add_dict_entries(dict, data, size) < 0)\n+\t\t\tdie(\"can't process %s object %s\",\n+\t\t\t\ttypename(type), sha1_to_hex(obj->sha1));\n+\t\tfree(data);\n+\t}\n+\n+\tstop_progress(&progress_state);\n+\treturn 0;\n+}\n+\n+static struct sha1file * packv4_open(char *path)\n+{\n+\tint fd;\n+\n+\tfd = open(path, O_CREAT|O_EXCL|O_WRONLY, 0600);\n+\tif (fd < 0)\n+\t\tdie_errno(\"unable to create '%s'\", path);\n+\treturn sha1fd(fd, path);\n+}\n+\n+static unsigned int packv4_write_header(struct sha1file *f, unsigned nr_objects)\n+{\n+\tstruct pack_header hdr;\n+\n+\thdr.hdr_signature = htonl(PACK_SIGNATURE);\n+\thdr.hdr_version = htonl(4);\n+\thdr.hdr_entries = htonl(nr_objects);\n+\tsha1write(f, &hdr, sizeof(hdr));\n+\n+\treturn sizeof(hdr);\n+}\n+\n+static int write_object_header(struct sha1file *f, enum object_type type, unsigned long size)\n+{\n+\tunsigned char buf[16];\n+\tuint64_t val;\n+\tint len;\n+\n+\t/*\n+\t * We really have only one kind of delta object.\n+\t */\n+\tif (type == OBJ_OFS_DELTA)\n+\t\ttype = OBJ_REF_DELTA;\n+\n+\t/*\n+\t * We allocate 4 bits in the LSB for the object type which should\n+\t * be good for quite a while, given that we effectively encodes\n+\t * only 5 object types: commit, tree, blob, delta, tag.\n+\t */\n+\tval = size;\n+\tif (MSB(val, 4))\n+\t\tdie(\"fixme: the code doesn't currently cope with big sizes\");\n+\tval <<= 4;\n+\tval |= type;\n+\tlen = encode_varint(val, buf);\n+\tsha1write(f, buf, len);\n+\treturn len;\n+}\n+\n+static unsigned long copy_object_data(struct packv4_tables *v4,\n+\t\t\t\t      struct sha1file *f, struct packed_git *p,\n+\t\t\t\t      off_t offset)\n+{\n+\tstruct pack_window *w_curs = NULL;\n+\tstruct revindex_entry *revidx;\n+\tenum object_type type;\n+\tunsigned long avail, size, datalen, written;\n+\tint hdrlen, reflen, idx_nr;\n+\tunsigned char *src, buf[24];\n+\n+\trevidx = find_pack_revindex(p, offset);\n+\tidx_nr = revidx->nr;\n+\tdatalen = revidx[1].offset - offset;\n+\n+\tsrc = use_pack(p, &w_curs, offset, &avail);\n+\thdrlen = unpack_object_header_buffer(src, avail, &type, &size);\n+\n+\twritten = write_object_header(f, type, size);\n+\n+\tif (type == OBJ_OFS_DELTA) {\n+\t\tconst unsigned char *cp = src + hdrlen;\n+\t\toff_t base_offset = decode_varint(&cp);\n+\t\thdrlen = cp - src;\n+\t\tbase_offset = offset - base_offset;\n+\t\tif (base_offset <= 0 || base_offset >= offset)\n+\t\t\tdie(\"delta offset out of bound\");\n+\t\trevidx = find_pack_revindex(p, base_offset);\n+\t\treflen = encode_sha1ref(v4,\n+\t\t\t\t\tnth_packed_object_sha1(p, revidx->nr),\n+\t\t\t\t\tbuf);\n+\t\tsha1write(f, buf, reflen);\n+\t\twritten += reflen;\n+\t} else if (type == OBJ_REF_DELTA) {\n+\t\treflen = encode_sha1ref(v4, src + hdrlen, buf);\n+\t\thdrlen += 20;\n+\t\tsha1write(f, buf, reflen);\n+\t\twritten += reflen;\n+\t}\n+\n+\tif (p->index_version > 1 &&\n+\t    check_pack_crc(p, &w_curs, offset, datalen, idx_nr))\n+\t\tdie(\"bad CRC for object at offset %\"PRIuMAX\" in %s\",\n+\t\t    (uintmax_t)offset, p->pack_name);\n+\n+\toffset += hdrlen;\n+\tdatalen -= hdrlen;\n+\n+\twhile (datalen) {\n+\t\tsrc = use_pack(p, &w_curs, offset, &avail);\n+\t\tif (avail > datalen)\n+\t\t\tavail = datalen;\n+\t\tsha1write(f, src, avail);\n+\t\twritten += avail;\n+\t\toffset += avail;\n+\t\tdatalen -= avail;\n+\t}\n+\tunuse_pack(&w_curs);\n+\n+\treturn written;\n+}\n+\n+static unsigned char *get_delta_base(struct packed_git *p, off_t offset,\n+\t\t\t\t     unsigned char *sha1_buf)\n+{\n+\tstruct pack_window *w_curs = NULL;\n+\tenum object_type type;\n+\tunsigned long avail, size;\n+\tint hdrlen;\n+\tunsigned char *src;\n+\tconst unsigned char *base_sha1 = NULL; ;\n+\n+\tsrc = use_pack(p, &w_curs, offset, &avail);\n+\thdrlen = unpack_object_header_buffer(src, avail, &type, &size);\n+\n+\tif (type == OBJ_OFS_DELTA) {\n+\t\tconst unsigned char *cp = src + hdrlen;\n+\t\toff_t base_offset = decode_varint(&cp);\n+\t\tbase_offset = offset - base_offset;\n+\t\tif (base_offset <= 0 || base_offset >= offset) {\n+\t\t\terror(\"delta offset out of bound\");\n+\t\t} else {\n+\t\t\tstruct revindex_entry *revidx;\n+\t\t\trevidx = find_pack_revindex(p, base_offset);\n+\t\t\tbase_sha1 = nth_packed_object_sha1(p, revidx->nr);\n+\t\t}\n+\t} else if (type == OBJ_REF_DELTA) {\n+\t\tbase_sha1 = src + hdrlen;\n+\t} else\n+\t\terror(\"expected to get a delta but got a %s\", typename(type));\n+\n+\tunuse_pack(&w_curs);\n+\n+\tif (!base_sha1)\n+\t\treturn NULL;\n+\thashcpy(sha1_buf, base_sha1);\n+\treturn sha1_buf;\n+}\n+\n+static off_t packv4_write_object(struct packv4_tables *v4,\n+\t\t\t\t struct sha1file *f, struct packed_git *p,\n+\t\t\t\t struct pack_idx_entry *obj)\n+{\n+\tvoid *src, *result;\n+\tstruct object_info oi = {};\n+\tenum object_type type, packed_type;\n+\tunsigned long obj_size, buf_size;\n+\tunsigned int hdrlen;\n+\n+\toi.typep = &type;\n+\toi.sizep = &obj_size;\n+\tpacked_type = packed_object_info(p, obj->offset, &oi);\n+\tif (packed_type < 0)\n+\t\tdie(\"cannot get type of %s from %s\",\n+\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\n+\t/* Some objects are copied without decompression */\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TREE:\n+\t\tbreak;\n+\tdefault:\n+\t\treturn copy_object_data(v4, f, p, obj->offset);\n+\t}\n+\n+\t/* The rest is converted into their new format */\n+\tsrc = unpack_entry(p, obj->offset, &type, &buf_size);\n+\tif (!src || obj_size != buf_size)\n+\t\tdie(\"cannot unpack %s from %s\",\n+\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\tif (check_sha1_signature(obj->sha1, src, buf_size, typename(type)))\n+\t\tdie(\"packed %s from %s is corrupt\",\n+\t\t    sha1_to_hex(obj->sha1), p->pack_name);\n+\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\t\tresult = pv4_encode_commit(v4, src, &buf_size);\n+\t\tbreak;\n+\tcase OBJ_TREE:\n+\t\tif (packed_type != OBJ_TREE) {\n+\t\t\tunsigned char sha1_buf[20], *ref_sha1;\n+\t\t\tvoid *ref;\n+\t\t\tenum object_type ref_type;\n+\t\t\tunsigned long ref_size;\n+\n+\t\t\tref_sha1 = get_delta_base(p, obj->offset, sha1_buf);\n+\t\t\tif (!ref_sha1)\n+\t\t\t\tdie(\"unable to get delta base sha1 for %s\",\n+\t\t\t\t\t\tsha1_to_hex(obj->sha1));\n+\t\t\tref = read_sha1_file(ref_sha1, &ref_type, &ref_size);\n+\t\t\tif (!ref || ref_type != OBJ_TREE)\n+\t\t\t\tdie(\"cannot obtain delta base for %s\",\n+\t\t\t\t\t\tsha1_to_hex(obj->sha1));\n+\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n+\t\t\t\t\t\t ref, ref_size, ref_sha1);\n+\t\t\tfree(ref);\n+\t\t} else {\n+\t\t\tresult = pv4_encode_tree(v4, src, &buf_size,\n+\t\t\t\t\t\t NULL, 0, NULL);\n+\t\t}\n+\t\tbreak;\n+\tdefault:\n+\t\tdie(\"unexpected object type %d\", type);\n+\t}\n+\tfree(src);\n+\tif (!result) {\n+\t\twarning(\"can't convert %s object %s\",\n+\t\t\ttypename(type), sha1_to_hex(obj->sha1));\n+\t\t/* fall back to copy the object in its original form */\n+\t\treturn copy_object_data(v4, f, p, obj->offset);\n+\t}\n+\n+\t/* Use bit 3 to indicate a special type encoding */\n+\ttype += 8;\n+\thdrlen = write_object_header(f, type, obj_size);\n+\tsha1write(f, result, buf_size);\n+\tfree(result);\n+\treturn hdrlen + buf_size;\n+}\n+\n+static char *normalize_pack_name(const char *path)\n+{\n+\tchar buf[PATH_MAX];\n+\tint len;\n+\n+\tlen = strlcpy(buf, path, PATH_MAX);\n+\tif (len >= PATH_MAX - 6)\n+\t\tdie(\"name too long: %s\", path);\n+\n+\t/*\n+\t * In addition to \"foo.idx\" we accept \"foo.pack\" and \"foo\";\n+\t * normalize these forms to \"foo.pack\".\n+\t */\n+\tif (has_extension(buf, \".idx\")) {\n+\t\tstrcpy(buf + len - 4, \".pack\");\n+\t\tlen++;\n+\t} else if (!has_extension(buf, \".pack\")) {\n+\t\tstrcpy(buf + len, \".pack\");\n+\t\tlen += 5;\n+\t}\n+\n+\treturn xstrdup(buf);\n+}\n+\n+static struct packed_git *open_pack(const char *path)\n+{\n+\tchar *packname = normalize_pack_name(path);\n+\tint len = strlen(packname);\n+\tstruct packed_git *p;\n+\n+\tstrcpy(packname + len - 5, \".idx\");\n+\tp = add_packed_git(packname, len - 1, 1);\n+\tif (!p)\n+\t\tdie(\"packfile %s not found.\", packname);\n+\n+\tinstall_packed_git(p);\n+\tif (open_pack_index(p))\n+\t\tdie(\"packfile %s index not opened\", p->pack_name);\n+\n+\tfree(packname);\n+\treturn p;\n+}\n+\n+void process_one_pack(struct packv4_tables *v4, char *src_pack, char *dst_pack)\n+{\n+\tstruct packed_git *p;\n+\tstruct sha1file *f;\n+\tstruct pack_idx_entry *objs, **p_objs;\n+\tstruct pack_idx_option idx_opts;\n+\tunsigned i, nr_objects;\n+\toff_t written = 0;\n+\tchar *packname;\n+\tunsigned char pack_sha1[20];\n+\tstruct progress *progress_state;\n+\n+\tp = open_pack(src_pack);\n+\tif (!p)\n+\t\tdie(\"unable to open source pack\");\n+\n+\tnr_objects = p->num_objects;\n+\tobjs = get_packed_object_list(p);\n+\tp_objs = sort_objs_by_offset(objs, nr_objects);\n+\n+\tcreate_pack_dictionaries(v4, p, p_objs);\n+\tsort_dict_entries_by_hits(v4->commit_ident_table);\n+\tsort_dict_entries_by_hits(v4->tree_path_table);\n+\n+\tpackname = normalize_pack_name(dst_pack);\n+\tf = packv4_open(packname);\n+\tif (!f)\n+\t\tdie(\"unable to open destination pack\");\n+\twritten += packv4_write_header(f, nr_objects);\n+\twritten += packv4_write_tables(f, v4);\n+\n+\t/* Let's write objects out, updating the object index list in place */\n+\tprogress_state = start_progress(\"Writing objects\", nr_objects);\n+\tv4->all_objs = objs;\n+\tv4->all_objs_nr = nr_objects;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\toff_t obj_pos = written;\n+\t\tstruct pack_idx_entry *obj = p_objs[i];\n+\t\tcrc32_begin(f);\n+\t\twritten += packv4_write_object(v4, f, p, obj);\n+\t\tobj->offset = obj_pos;\n+\t\tobj->crc32 = crc32_end(f);\n+\t\tdisplay_progress(progress_state, i+1);\n+\t}\n+\tstop_progress(&progress_state);\n+\n+\tsha1close(f, pack_sha1, CSUM_CLOSE | CSUM_FSYNC);\n+\n+\treset_pack_idx_option(&idx_opts);\n+\tidx_opts.version = 3;\n+\tstrcpy(packname + strlen(packname) - 5, \".idx\");\n+\twrite_idx_file(packname, p_objs, nr_objects, &idx_opts, pack_sha1);\n+\n+\tfree(packname);\n+}\n+\n+static int git_pack_config(const char *k, const char *v, void *cb)\n+{\n+\tif (!strcmp(k, \"pack.compression\")) {\n+\t\tint level = git_config_int(k, v);\n+\t\tif (level == -1)\n+\t\t\tlevel = Z_DEFAULT_COMPRESSION;\n+\t\telse if (level < 0 || level > Z_BEST_COMPRESSION)\n+\t\t\tdie(\"bad pack compression level %d\", level);\n+\t\tpack_compression_level = level;\n+\t\tpack_compression_seen = 1;\n+\t\treturn 0;\n+\t}\n+\treturn git_default_config(k, v, cb);\n+}\n+\n+int main(int argc, char *argv[])\n+{\n+\tstruct packv4_tables v4;\n+\tchar *src_pack, *dst_pack;\n+\n+\tif (argc == 3) {\n+\t\tsrc_pack = argv[1];\n+\t\tdst_pack = argv[2];\n+\t} else if (argc == 4 && !prefixcmp(argv[1], \"--min-tree-copy=\")) {\n+\t\tmin_tree_copy = atoi(argv[1] + strlen(\"--min-tree-copy=\"));\n+\t\tsrc_pack = argv[2];\n+\t\tdst_pack = argv[3];\n+\t} else {\n+\t\tfprintf(stderr, \"Usage: %s [--min-tree-copy=<n>] <src_packfile> <dst_packfile>\\n\", argv[0]);\n+\t\texit(1);\n+\t}\n+\n+\tgit_config(git_pack_config, NULL);\n+\tif (!pack_compression_seen && core_compression_seen)\n+\t\tpack_compression_level = core_compression_level;\n+\tprocess_one_pack(&v4, src_pack, dst_pack);\n+\tif (0)\n+\t\tdict_dump(&v4);\n+\treturn 0;\n+}\n-- \n1.8.2.83.gc99314b\n"},{"id":"227217","messageId":"1378735087-4813-5-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 04/16] pack v4: add version argument to write_pack_header","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:57:55Z","receivedAt":"2013-09-09T13:57:55Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 2 +-\n bulk-checkin.c         | 2 +-\n pack-write.c           | 7 +++++--\n pack.h                 | 3 +--\n 4 files changed, 8 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex f069462..33faea8 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -735,7 +735,7 @@ static void write_pack_file(void)\n \t\telse\n \t\t\tf = create_tmp_packfile(&pack_tmp_name);\n \n-\t\toffset = write_pack_header(f, nr_remaining);\n+\t\toffset = write_pack_header(f, 2, nr_remaining);\n \t\tif (!offset)\n \t\t\tdie_errno(\"unable to write pack header\");\n \t\tnr_written = 0;\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6b0b6d4..9d8f0d0 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -176,7 +176,7 @@ static void prepare_to_stream(struct bulk_checkin_state *state,\n \treset_pack_idx_option(&state->pack_idx_opts);\n \n \t/* Pretend we are going to write only one object */\n-\tstate->offset = write_pack_header(state->f, 1);\n+\tstate->offset = write_pack_header(state->f, 2, 1);\n \tif (!state->offset)\n \t\tdie_errno(\"unable to write pack header\");\n }\ndiff --git a/pack-write.c b/pack-write.c\nindex 631007e..88e4788 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -186,12 +186,15 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \treturn index_name;\n }\n \n-off_t write_pack_header(struct sha1file *f, uint32_t nr_entries)\n+off_t write_pack_header(struct sha1file *f,\n+\t\t\tint version, uint32_t nr_entries)\n {\n \tstruct pack_header hdr;\n \n \thdr.hdr_signature = htonl(PACK_SIGNATURE);\n-\thdr.hdr_version = htonl(PACK_VERSION);\n+\thdr.hdr_version = htonl(version);\n+\tif (!pack_version_ok(hdr.hdr_version))\n+\t\tdie(_(\"pack version %d is not supported\"), version);\n \thdr.hdr_entries = htonl(nr_entries);\n \tif (sha1write(f, &hdr, sizeof(hdr)))\n \t\treturn 0;\ndiff --git a/pack.h b/pack.h\nindex aa6ee7d..855f6c6 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -8,7 +8,6 @@\n  * Packed object header\n  */\n #define PACK_SIGNATURE 0x5041434b\t/* \"PACK\" */\n-#define PACK_VERSION 2\n #define pack_version_ok(v) ((v) == htonl(2) || (v) == htonl(3))\n struct pack_header {\n \tuint32_t hdr_signature;\n@@ -80,7 +79,7 @@ extern const char *write_idx_file(const char *index_name, struct pack_idx_entry\n extern int check_pack_crc(struct packed_git *p, struct pack_window **w_curs, off_t offset, off_t len, unsigned int nr);\n extern int verify_pack_index(struct packed_git *);\n extern int verify_pack(struct packed_git *, verify_fn fn, struct progress *, uint32_t);\n-extern off_t write_pack_header(struct sha1file *f, uint32_t);\n+extern off_t write_pack_header(struct sha1file *f, int, uint32_t);\n extern void fixup_pack_header_footer(int, unsigned char *, const char *, uint32_t, unsigned char *, off_t);\n extern char *index_pack_lockfile(int fd);\n extern int encode_in_pack_object_header(enum object_type, uintmax_t, unsigned char *);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227218","messageId":"1378735087-4813-6-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 05/16] pack_write: tighten valid object type check in encode_in_pack_object_header","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:57:56Z","receivedAt":"2013-09-09T13:57:56Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n pack-write.c | 11 ++++++++++-\n 1 file changed, 10 insertions(+), 1 deletion(-)\n\ndiff --git a/pack-write.c b/pack-write.c\nindex 88e4788..36b88a3 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -325,8 +325,17 @@ int encode_in_pack_object_header(enum object_type type, uintmax_t size, unsigned\n \tint n = 1;\n \tunsigned char c;\n \n-\tif (type < OBJ_COMMIT || type > OBJ_REF_DELTA)\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TREE:\n+\tcase OBJ_BLOB:\n+\tcase OBJ_TAG:\n+\tcase OBJ_OFS_DELTA:\n+\tcase OBJ_REF_DELTA:\n+\t\tbreak;\n+\tdefault:\n \t\tdie(\"bad type %d\", type);\n+\t}\n \n \tc = (type << 4) | (size & 15);\n \tsize >>= 4;\n-- \n1.8.2.83.gc99314b\n"},{"id":"227220","messageId":"1378735087-4813-7-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 06/16] pack-write.c: add pv4_encode_object_header","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:57:57Z","receivedAt":"2013-09-09T13:57:57Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n pack-write.c | 33 +++++++++++++++++++++++++++++++++\n pack.h       |  1 +\n 2 files changed, 34 insertions(+)\n\ndiff --git a/pack-write.c b/pack-write.c\nindex 36b88a3..c1e9da4 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -1,6 +1,7 @@\n #include \"cache.h\"\n #include \"pack.h\"\n #include \"csum-file.h\"\n+#include \"varint.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -349,6 +350,38 @@ int encode_in_pack_object_header(enum object_type type, uintmax_t size, unsigned\n \treturn n;\n }\n \n+int pv4_encode_object_header(enum object_type type,\n+\t\t\t     uintmax_t size, unsigned char *hdr)\n+{\n+\tuintmax_t val;\n+\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TREE:\n+\tcase OBJ_BLOB:\n+\tcase OBJ_TAG:\n+\tcase OBJ_REF_DELTA:\n+\tcase OBJ_PV4_COMMIT:\n+\tcase OBJ_PV4_TREE:\n+\t\tbreak;\n+\tdefault:\n+\t\tdie(\"bad type %d\", type);\n+\t}\n+\n+\t/*\n+\t * We allocate 4 bits in the LSB for the object type which\n+\t * should be good for quite a while, given that we effectively\n+\t * encodes only 5 object types: commit, tree, blob, delta,\n+\t * tag.\n+\t */\n+\tval = size;\n+\tif (MSB(val, 4))\n+\t\tdie(\"fixme: the code doesn't currently cope with big sizes\");\n+\tval <<= 4;\n+\tval |= type;\n+\treturn encode_varint(val, hdr);\n+}\n+\n struct sha1file *create_tmp_packfile(char **pack_tmp_name)\n {\n \tchar tmpname[PATH_MAX];\ndiff --git a/pack.h b/pack.h\nindex 855f6c6..4f10fa4 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -83,6 +83,7 @@ extern off_t write_pack_header(struct sha1file *f, int, uint32_t);\n extern void fixup_pack_header_footer(int, unsigned char *, const char *, uint32_t, unsigned char *, off_t);\n extern char *index_pack_lockfile(int fd);\n extern int encode_in_pack_object_header(enum object_type, uintmax_t, unsigned char *);\n+extern int pv4_encode_object_header(enum object_type, uintmax_t, unsigned char *);\n \n #define PH_ERROR_EOF\t\t(-1)\n #define PH_ERROR_PACK_SIGNATURE\t(-2)\n-- \n1.8.2.83.gc99314b\n"},{"id":"227221","messageId":"1378735087-4813-8-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 07/16] pack-objects: add --version to specify written pack version","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:57:58Z","receivedAt":"2013-09-09T13:57:58Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 6 +++++-\n 1 file changed, 5 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 33faea8..ef68fc5 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -81,6 +81,7 @@ static int num_preferred_base;\n static struct progress *progress_state;\n static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n static int pack_compression_seen;\n+static int pack_version = 2;\n \n static unsigned long delta_cache_size = 0;\n static unsigned long max_delta_cache_size = 256 * 1024 * 1024;\n@@ -735,7 +736,7 @@ static void write_pack_file(void)\n \t\telse\n \t\t\tf = create_tmp_packfile(&pack_tmp_name);\n \n-\t\toffset = write_pack_header(f, 2, nr_remaining);\n+\t\toffset = write_pack_header(f, pack_version, nr_remaining);\n \t\tif (!offset)\n \t\t\tdie_errno(\"unable to write pack header\");\n \t\tnr_written = 0;\n@@ -2455,6 +2456,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t{ OPTION_CALLBACK, 0, \"index-version\", NULL, N_(\"version[,offset]\"),\n \t\t  N_(\"write the pack index file in the specified idx format version\"),\n \t\t  0, option_parse_index_version },\n+\t\tOPT_INTEGER(0, \"version\", &pack_version, N_(\"pack version\")),\n \t\tOPT_ULONG(0, \"max-pack-size\", &pack_size_limit,\n \t\t\t  N_(\"maximum size of each output pack file\")),\n \t\tOPT_BOOL(0, \"local\", &local,\n@@ -2525,6 +2527,8 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t}\n \tif (pack_to_stdout != !base_name || argc)\n \t\tusage_with_options(pack_usage, pack_objects_options);\n+\tif (pack_version != 2)\n+\t\tdie(_(\"pack version %d is not supported\"), pack_version);\n \n \trp_av[rp_ac++] = \"pack-objects\";\n \tif (thin) {\n-- \n1.8.2.83.gc99314b\n"},{"id":"227222","messageId":"1378735087-4813-9-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 08/16] list-objects.c: add show_tree_entry callback to traverse_commit_list","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:57:59Z","receivedAt":"2013-09-09T13:57:59Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This helps construct tree dictionary in pack v4.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 2 +-\n builtin/rev-list.c     | 4 ++--\n list-objects.c         | 9 ++++++++-\n list-objects.h         | 3 ++-\n upload-pack.c          | 2 +-\n 5 files changed, 14 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex ef68fc5..b38d3dc 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -2380,7 +2380,7 @@ static void get_object_list(int ac, const char **av)\n \tif (prepare_revision_walk(&revs))\n \t\tdie(\"revision walk setup failed\");\n \tmark_edges_uninteresting(revs.commits, &revs, show_edge);\n-\ttraverse_commit_list(&revs, show_commit, show_object, NULL);\n+\ttraverse_commit_list(&revs, show_commit, NULL, show_object, NULL);\n \n \tif (keep_unreachable)\n \t\tadd_objects_in_unpacked_packs(&revs);\ndiff --git a/builtin/rev-list.c b/builtin/rev-list.c\nindex a5ec30d..b25f896 100644\n--- a/builtin/rev-list.c\n+++ b/builtin/rev-list.c\n@@ -243,7 +243,7 @@ static int show_bisect_vars(struct rev_list_info *info, int reaches, int all)\n \t\tstrcpy(hex, sha1_to_hex(revs->commits->item->object.sha1));\n \n \tif (flags & BISECT_SHOW_ALL) {\n-\t\ttraverse_commit_list(revs, show_commit, show_object, info);\n+\t\ttraverse_commit_list(revs, show_commit, NULL, show_object, info);\n \t\tprintf(\"------\\n\");\n \t}\n \n@@ -348,7 +348,7 @@ int cmd_rev_list(int argc, const char **argv, const char *prefix)\n \t\t\treturn show_bisect_vars(&info, reaches, all);\n \t}\n \n-\ttraverse_commit_list(&revs, show_commit, show_object, &info);\n+\ttraverse_commit_list(&revs, show_commit, NULL, show_object, &info);\n \n \tif (revs.count) {\n \t\tif (revs.left_right && revs.cherry_mark)\ndiff --git a/list-objects.c b/list-objects.c\nindex 3dd4a96..6def897 100644\n--- a/list-objects.c\n+++ b/list-objects.c\n@@ -61,6 +61,7 @@ static void process_gitlink(struct rev_info *revs,\n \n static void process_tree(struct rev_info *revs,\n \t\t\t struct tree *tree,\n+\t\t\t show_tree_entry_fn show_tree_entry,\n \t\t\t show_object_fn show,\n \t\t\t struct name_path *path,\n \t\t\t struct strbuf *base,\n@@ -107,9 +108,13 @@ static void process_tree(struct rev_info *revs,\n \t\t\t\tcontinue;\n \t\t}\n \n+\t\tif (show_tree_entry)\n+\t\t\tshow_tree_entry(&entry, cb_data);\n+\n \t\tif (S_ISDIR(entry.mode))\n \t\t\tprocess_tree(revs,\n \t\t\t\t     lookup_tree(entry.sha1),\n+\t\t\t\t     show_tree_entry,\n \t\t\t\t     show, &me, base, entry.path,\n \t\t\t\t     cb_data);\n \t\telse if (S_ISGITLINK(entry.mode))\n@@ -167,6 +172,7 @@ static void add_pending_tree(struct rev_info *revs, struct tree *tree)\n \n void traverse_commit_list(struct rev_info *revs,\n \t\t\t  show_commit_fn show_commit,\n+\t\t\t  show_tree_entry_fn show_tree_entry,\n \t\t\t  show_object_fn show_object,\n \t\t\t  void *data)\n {\n@@ -196,7 +202,8 @@ void traverse_commit_list(struct rev_info *revs,\n \t\t\tcontinue;\n \t\t}\n \t\tif (obj->type == OBJ_TREE) {\n-\t\t\tprocess_tree(revs, (struct tree *)obj, show_object,\n+\t\t\tprocess_tree(revs, (struct tree *)obj,\n+\t\t\t\t     show_tree_entry, show_object,\n \t\t\t\t     NULL, &base, name, data);\n \t\t\tcontinue;\n \t\t}\ndiff --git a/list-objects.h b/list-objects.h\nindex 3db7bb6..297b2e0 100644\n--- a/list-objects.h\n+++ b/list-objects.h\n@@ -2,8 +2,9 @@\n #define LIST_OBJECTS_H\n \n typedef void (*show_commit_fn)(struct commit *, void *);\n+typedef void (*show_tree_entry_fn)(const struct name_entry *, void *);\n typedef void (*show_object_fn)(struct object *, const struct name_path *, const char *, void *);\n-void traverse_commit_list(struct rev_info *, show_commit_fn, show_object_fn, void *);\n+void traverse_commit_list(struct rev_info *, show_commit_fn, show_tree_entry_fn, show_object_fn, void *);\n \n typedef void (*show_edge_fn)(struct commit *);\n void mark_edges_uninteresting(struct commit_list *, struct rev_info *, show_edge_fn);\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 127e59a..ccf76d9 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -125,7 +125,7 @@ static int do_rev_list(int in, int out, void *user_data)\n \t\tfor (i = 0; i < extra_edge_obj.nr; i++)\n \t\t\tfprintf(pack_pipe, \"-%s\\n\", sha1_to_hex(\n \t\t\t\t\textra_edge_obj.objects[i].item->sha1));\n-\ttraverse_commit_list(&revs, show_commit, show_object, NULL);\n+\ttraverse_commit_list(&revs, show_commit, NULL, show_object, NULL);\n \tfflush(pack_pipe);\n \tfclose(pack_pipe);\n \treturn 0;\n-- \n1.8.2.83.gc99314b\n"},{"id":"227223","messageId":"1378735087-4813-10-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 09/16] pack-objects: do not cache delta for v4 trees","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:58:00Z","receivedAt":"2013-09-09T13:58:00Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 6 +++++-\n 1 file changed, 5 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex b38d3dc..9613732 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1756,8 +1756,12 @@ static void find_deltas(struct object_entry **list, unsigned *list_size,\n \t\t * and therefore it is best to go to the write phase ASAP\n \t\t * instead, as we can afford spending more time compressing\n \t\t * between writes at that moment.\n+\t\t *\n+\t\t * For v4 trees we'll need to delta differently anyway\n+\t\t * so no cache. v4 commits simply do not delta.\n \t\t */\n-\t\tif (entry->delta_data && !pack_to_stdout) {\n+\t\tif (entry->delta_data && !pack_to_stdout &&\n+\t\t    (pack_version < 4 || entry->type == OBJ_BLOB)) {\n \t\t\tentry->z_delta_size = do_compress(&entry->delta_data,\n \t\t\t\t\t\t\t  entry->delta_size);\n \t\t\tcache_lock();\n-- \n1.8.2.83.gc99314b\n"},{"id":"227224","messageId":"1378735087-4813-11-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 10/16] pack-objects: exclude commits out of delta objects in v4","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:58:01Z","receivedAt":"2013-09-09T13:58:01Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 5 ++++-\n 1 file changed, 4 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 9613732..fb2394d 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1329,7 +1329,8 @@ static void check_object(struct object_entry *entry)\n \t\t\tbreak;\n \t\t}\n \n-\t\tif (base_ref && (base_entry = locate_object_entry(base_ref))) {\n+\t\tif (base_ref && (base_entry = locate_object_entry(base_ref)) &&\n+\t\t    (pack_version < 4 || entry->type != OBJ_COMMIT)) {\n \t\t\t/*\n \t\t\t * If base_ref was set above that means we wish to\n \t\t\t * reuse delta data, and we even found that base\n@@ -1413,6 +1414,8 @@ static void get_object_details(void)\n \t\tcheck_object(entry);\n \t\tif (big_file_threshold < entry->size)\n \t\t\tentry->no_try_delta = 1;\n+\t\tif (pack_version == 4 && entry->type == OBJ_COMMIT)\n+\t\t\tentry->no_try_delta = 1;\n \t}\n \n \tfree(sorted_by_offset);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227225","messageId":"1378735087-4813-12-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 11/16] pack-objects: create pack v4 tables","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:58:02Z","receivedAt":"2013-09-09T13:58:02Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 62 ++++++++++++++++++++++++++++++++++++++++++++++++--\n 1 file changed, 60 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex fb2394d..60ea5a7 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -18,6 +18,7 @@\n #include \"refs.h\"\n #include \"streaming.h\"\n #include \"thread-utils.h\"\n+#include \"packv4-create.h\"\n \n static const char *pack_usage[] = {\n \tN_(\"git pack-objects --stdout [options...] [< ref-list | < object-list]\"),\n@@ -61,6 +62,8 @@ static struct object_entry *objects;\n static struct pack_idx_entry **written_list;\n static uint32_t nr_objects, nr_alloc, nr_result, nr_written;\n \n+static struct packv4_tables v4;\n+\n static int non_empty;\n static int reuse_delta = 1, reuse_object = 1;\n static int keep_unreachable, unpack_unreachable, include_tag;\n@@ -2052,6 +2055,11 @@ static void prepare_pack(int window, int depth)\n \tuint32_t i, nr_deltas;\n \tunsigned n;\n \n+\tif (pack_version == 4) {\n+\t\tsort_dict_entries_by_hits(v4.commit_ident_table);\n+\t\tsort_dict_entries_by_hits(v4.tree_path_table);\n+\t}\n+\n \tget_object_details();\n \n \t/*\n@@ -2198,6 +2206,34 @@ static void read_object_list_from_stdin(void)\n \n \t\tadd_preferred_base_object(line+41);\n \t\tadd_object_entry(sha1, 0, line+41, 0);\n+\n+\t\tif (pack_version == 4) {\n+\t\t\tvoid *data;\n+\t\t\tenum object_type type;\n+\t\t\tunsigned long size;\n+\t\t\tint (*add_dict_entries)(struct dict_table *, void *, unsigned long);\n+\t\t\tstruct dict_table *dict;\n+\n+\t\t\tswitch (sha1_object_info(sha1, &size)) {\n+\t\t\tcase OBJ_COMMIT:\n+\t\t\t\tadd_dict_entries = add_commit_dict_entries;\n+\t\t\t\tdict = v4.commit_ident_table;\n+\t\t\t\tbreak;\n+\t\t\tcase OBJ_TREE:\n+\t\t\t\tadd_dict_entries = add_tree_dict_entries;\n+\t\t\t\tdict = v4.tree_path_table;\n+\t\t\t\tbreak;\n+\t\t\tdefault:\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tdata = read_sha1_file(sha1, &type, &size);\n+\t\t\tif (!data)\n+\t\t\t\tdie(\"cannot unpack %s\", sha1_to_hex(sha1));\n+\t\t\tif (add_dict_entries(dict, data, size) < 0)\n+\t\t\t\tdie(\"can't process %s object %s\",\n+\t\t\t\t    typename(type), sha1_to_hex(sha1));\n+\t\t\tfree(data);\n+\t\t}\n \t}\n }\n \n@@ -2205,10 +2241,26 @@ static void read_object_list_from_stdin(void)\n \n static void show_commit(struct commit *commit, void *data)\n {\n+\tif (pack_version == 4) {\n+\t\tunsigned long size;\n+\t\tenum object_type type;\n+\t\tunsigned char *buf;\n+\n+\t\t/* commit->buffer is NULL most of the time, don't bother */\n+\t\tbuf = read_sha1_file(commit->object.sha1, &type, &size);\n+\t\tadd_commit_dict_entries(v4.commit_ident_table, buf, size);\n+\t\tfree(buf);\n+\t}\n \tadd_object_entry(commit->object.sha1, OBJ_COMMIT, NULL, 0);\n \tcommit->object.flags |= OBJECT_ADDED;\n }\n \n+static void show_tree_entry(const struct name_entry *entry, void *data)\n+{\n+\tdict_add_entry(v4.tree_path_table, entry->mode, entry->path,\n+\t\t       tree_entry_len(entry));\n+}\n+\n static void show_object(struct object *obj,\n \t\t\tconst struct name_path *path, const char *last,\n \t\t\tvoid *data)\n@@ -2387,7 +2439,9 @@ static void get_object_list(int ac, const char **av)\n \tif (prepare_revision_walk(&revs))\n \t\tdie(\"revision walk setup failed\");\n \tmark_edges_uninteresting(revs.commits, &revs, show_edge);\n-\ttraverse_commit_list(&revs, show_commit, NULL, show_object, NULL);\n+\ttraverse_commit_list(&revs, show_commit,\n+\t\t\t     pack_version == 4 ? show_tree_entry : NULL,\n+\t\t\t     show_object, NULL);\n \n \tif (keep_unreachable)\n \t\tadd_objects_in_unpacked_packs(&revs);\n@@ -2534,7 +2588,7 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t}\n \tif (pack_to_stdout != !base_name || argc)\n \t\tusage_with_options(pack_usage, pack_objects_options);\n-\tif (pack_version != 2)\n+\tif (pack_version != 2 && pack_version != 4)\n \t\tdie(_(\"pack version %d is not supported\"), pack_version);\n \n \trp_av[rp_ac++] = \"pack-objects\";\n@@ -2586,6 +2640,10 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tprogress = 2;\n \n \tprepare_packed_git();\n+\tif (pack_version == 4) {\n+\t\tv4.commit_ident_table = create_dict_table();\n+\t\tv4.tree_path_table = create_dict_table();\n+\t}\n \n \tif (progress)\n \t\tprogress_state = start_progress(\"Counting objects\", 0);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227226","messageId":"1378735087-4813-13-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 12/16] pack-objects: prepare SHA-1 table in v4","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:58:03Z","receivedAt":"2013-09-09T13:58:03Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"SHA-1 table is trickier than ident or path tables because it must\ncontains exactly the number entries in pack. In the thin pack case it\nmust also cover bases that will be appended by index-pack.\n\nThe problem is not all preferred_base entries end up becoming actually\nneeded. So we do a fake write_one() round just to get what is written\nand what is not. It also helps the case when the multiple packs are\nwritten.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 55 +++++++++++++++++++++++++++++++++++++++++++++++---\n 1 file changed, 52 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 60ea5a7..055b59d 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -434,7 +434,7 @@ static unsigned long write_object(struct sha1file *f,\n \tunsigned long limit, len;\n \tint usable_delta, to_reuse;\n \n-\tif (!pack_to_stdout)\n+\tif (f && !pack_to_stdout)\n \t\tcrc32_begin(f);\n \n \t/* apply size limit if limited packsize and not first object */\n@@ -477,6 +477,12 @@ static unsigned long write_object(struct sha1file *f,\n \t\t\t\t * and we do not need to deltify it.\n \t\t\t\t */\n \n+\tif (!f) {\n+\t\tif (usable_delta && entry->delta->idx.offset < 2)\n+\t\t\tentry->delta->idx.offset = 2;\n+\t\treturn 2;\n+\t}\n+\n \tif (!to_reuse)\n \t\tlen = write_no_reuse_object(f, entry, limit, usable_delta);\n \telse\n@@ -543,10 +549,14 @@ static enum write_one_status write_one(struct sha1file *f,\n \t\te->idx.offset = recursing;\n \t\treturn WRITE_ONE_BREAK;\n \t}\n+\tif (!f) {\n+\t\t*offset += size;\n+\t\treturn WRITE_ONE_WRITTEN;\n+\t}\n \twritten_list[nr_written++] = &e->idx;\n \n \t/* make sure off_t is sufficiently large not to wrap */\n-\tif (signed_add_overflows(*offset, size))\n+\tif (f && signed_add_overflows(*offset, size))\n \t\tdie(\"pack too large for current definition of off_t\");\n \t*offset += size;\n \treturn WRITE_ONE_WRITTEN;\n@@ -716,6 +726,39 @@ static struct object_entry **compute_write_order(void)\n \treturn wo;\n }\n \n+static int sha1_idx_sort(const void *a_, const void *b_)\n+{\n+\tconst struct pack_idx_entry *a = a_;\n+\tconst struct pack_idx_entry *b = b_;\n+\treturn hashcmp(a->sha1, b->sha1);\n+}\n+\n+/*\n+ * Do a fake writting round to detemine what's in the SHA-1 table.\n+ */\n+static void prepare_sha1_table(uint32_t start, struct object_entry **write_order)\n+{\n+\tint i = start;\n+\toff_t fake_offset = 2;\n+\tfor (; i < nr_objects; i++) {\n+\t\tstruct object_entry *e = write_order[i];\n+\t\tif (write_one(NULL, e, &fake_offset) == WRITE_ONE_BREAK)\n+\t\t\tbreak;\n+\t}\n+\n+\tv4.all_objs_nr = 0;\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\tstruct object_entry *e = write_order[i];\n+\t\tif (e->idx.offset > 0) {\n+\t\t\tv4.all_objs[v4.all_objs_nr++] = e->idx;\n+\t\t\tfprintf(stderr, \"%s in\\n\", sha1_to_hex(e->idx.sha1));\n+\t\t\te->idx.offset = 0;\n+\t\t}\n+\t}\n+\tqsort(v4.all_objs, v4.all_objs_nr, sizeof(*v4.all_objs),\n+\t      sha1_idx_sort);\n+}\n+\n static void write_pack_file(void)\n {\n \tuint32_t i = 0, j;\n@@ -739,7 +782,12 @@ static void write_pack_file(void)\n \t\telse\n \t\t\tf = create_tmp_packfile(&pack_tmp_name);\n \n-\t\toffset = write_pack_header(f, pack_version, nr_remaining);\n+\t\tif (pack_version == 4)\n+\t\t\tprepare_sha1_table(i, write_order);\n+\n+\t\toffset = write_pack_header(f, pack_version,\n+\t\t\t\t\t   pack_version < 4 ? nr_remaining : v4.all_objs_nr);\n+\n \t\tif (!offset)\n \t\t\tdie_errno(\"unable to write pack header\");\n \t\tnr_written = 0;\n@@ -2058,6 +2106,7 @@ static void prepare_pack(int window, int depth)\n \tif (pack_version == 4) {\n \t\tsort_dict_entries_by_hits(v4.commit_ident_table);\n \t\tsort_dict_entries_by_hits(v4.tree_path_table);\n+\t\tv4.all_objs = xmalloc(nr_objects * sizeof(*v4.all_objs));\n \t}\n \n \tget_object_details();\n-- \n1.8.2.83.gc99314b\n"},{"id":"227227","messageId":"1378735087-4813-14-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 13/16] pack-objects: support writing pack v4","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:58:04Z","receivedAt":"2013-09-09T13:58:04Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/pack-objects.c | 85 +++++++++++++++++++++++++++++++++++++++++++++-----\n pack.h                 |  2 +-\n 2 files changed, 78 insertions(+), 9 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 055b59d..12d9af4 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -254,6 +254,7 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \tenum object_type type;\n \tvoid *buf;\n \tstruct git_istream *st = NULL;\n+\tchar *result = \"OK\";\n \n \tif (!usable_delta) {\n \t\tif (entry->type == OBJ_BLOB &&\n@@ -287,7 +288,37 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \n \tif (st)\t/* large blob case, just assume we don't compress well */\n \t\tdatalen = size;\n-\telse if (entry->z_delta_size)\n+\telse if (pack_version == 4 && entry->type == OBJ_COMMIT) {\n+\t\tdatalen = size;\n+\t\tresult = pv4_encode_commit(&v4, buf, &datalen);\n+\t\tif (result) {\n+\t\t\tfree(buf);\n+\t\t\tbuf = result;\n+\t\t\ttype = OBJ_PV4_COMMIT;\n+\t\t}\n+\t} else if (pack_version == 4 && entry->type == OBJ_TREE) {\n+\t\tdatalen = size;\n+\t\tif (usable_delta) {\n+\t\t\tunsigned long base_size;\n+\t\t\tchar *base_buf;\n+\t\t\tbase_buf = read_sha1_file(entry->delta->idx.sha1, &type,\n+\t\t\t\t\t\t  &base_size);\n+\t\t\tif (!base_buf || type != OBJ_TREE)\n+\t\t\t\tdie(\"unable to read %s\",\n+\t\t\t\t    sha1_to_hex(entry->delta->idx.sha1));\n+\t\t\tresult = pv4_encode_tree(&v4, buf, &datalen,\n+\t\t\t\t\t\t base_buf, base_size,\n+\t\t\t\t\t\t entry->delta->idx.sha1);\n+\t\t\tfree(base_buf);\n+\t\t} else\n+\t\t\tresult = pv4_encode_tree(&v4, buf, &datalen,\n+\t\t\t\t\t\t NULL, 0, NULL);\n+\t\tif (result) {\n+\t\t\tfree(buf);\n+\t\t\tbuf = result;\n+\t\t\ttype = OBJ_PV4_TREE;\n+\t\t}\n+\t} else if (entry->z_delta_size)\n \t\tdatalen = entry->z_delta_size;\n \telse\n \t\tdatalen = do_compress(&buf, size);\n@@ -296,7 +327,10 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \t * The object header is a byte of 'type' followed by zero or\n \t * more bytes of length.\n \t */\n-\thdrlen = encode_in_pack_object_header(type, size, header);\n+\tif (pack_version < 4)\n+\t\thdrlen = encode_in_pack_object_header(type, size, header);\n+\telse\n+\t\thdrlen = pv4_encode_object_header(type, size, header);\n \n \tif (type == OBJ_OFS_DELTA) {\n \t\t/*\n@@ -318,7 +352,7 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \t\tsha1write(f, header, hdrlen);\n \t\tsha1write(f, dheader + pos, sizeof(dheader) - pos);\n \t\thdrlen += sizeof(dheader) - pos;\n-\t} else if (type == OBJ_REF_DELTA) {\n+\t} else if (type == OBJ_REF_DELTA && pack_version < 4) {\n \t\t/*\n \t\t * Deltas with a base reference contain\n \t\t * an additional 20 bytes for the base sha1.\n@@ -332,6 +366,10 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \t\tsha1write(f, header, hdrlen);\n \t\tsha1write(f, entry->delta->idx.sha1, 20);\n \t\thdrlen += 20;\n+\t} else if (type == OBJ_REF_DELTA && pack_version == 4) {\n+\t\thdrlen += encode_sha1ref(&v4, entry->delta->idx.sha1,\n+\t\t\t\t\theader + hdrlen);\n+\t\tsha1write(f, header, hdrlen);\n \t} else {\n \t\tif (limit && hdrlen + datalen + 20 >= limit) {\n \t\t\tif (st)\n@@ -341,14 +379,26 @@ static unsigned long write_no_reuse_object(struct sha1file *f, struct object_ent\n \t\t}\n \t\tsha1write(f, header, hdrlen);\n \t}\n+\n \tif (st) {\n \t\tdatalen = write_large_blob_data(st, f, entry->idx.sha1);\n \t\tclose_istream(st);\n-\t} else {\n-\t\tsha1write(f, buf, datalen);\n-\t\tfree(buf);\n+\t\treturn hdrlen + datalen;\n \t}\n \n+\tif (!result) {\n+\t\twarning(_(\"can't convert %s object %s\"),\n+\t\t\ttypename(entry->type),\n+\t\t\tsha1_to_hex(entry->idx.sha1));\n+\t\tfree(buf);\n+\t\tbuf = read_sha1_file(entry->idx.sha1, &type, &size);\n+\t\tif (!buf)\n+\t\t\tdie(_(\"unable to read %s\"),\n+\t\t\t    sha1_to_hex(entry->idx.sha1));\n+\t\tdatalen = do_compress(&buf, size);\n+\t}\n+\tsha1write(f, buf, datalen);\n+\tfree(buf);\n \treturn hdrlen + datalen;\n }\n \n@@ -368,7 +418,10 @@ static unsigned long write_reuse_object(struct sha1file *f, struct object_entry\n \tif (entry->delta)\n \t\ttype = (allow_ofs_delta && entry->delta->idx.offset) ?\n \t\t\tOBJ_OFS_DELTA : OBJ_REF_DELTA;\n-\thdrlen = encode_in_pack_object_header(type, entry->size, header);\n+\tif (pack_version < 4)\n+\t\thdrlen = encode_in_pack_object_header(type, entry->size, header);\n+\telse\n+\t\thdrlen = pv4_encode_object_header(type, entry->size, header);\n \n \toffset = entry->in_pack_offset;\n \trevidx = find_pack_revindex(p, offset);\n@@ -404,7 +457,7 @@ static unsigned long write_reuse_object(struct sha1file *f, struct object_entry\n \t\tsha1write(f, dheader + pos, sizeof(dheader) - pos);\n \t\thdrlen += sizeof(dheader) - pos;\n \t\treused_delta++;\n-\t} else if (type == OBJ_REF_DELTA) {\n+\t} else if (type == OBJ_REF_DELTA && pack_version < 4) {\n \t\tif (limit && hdrlen + 20 + datalen + 20 >= limit) {\n \t\t\tunuse_pack(&w_curs);\n \t\t\treturn 0;\n@@ -413,6 +466,11 @@ static unsigned long write_reuse_object(struct sha1file *f, struct object_entry\n \t\tsha1write(f, entry->delta->idx.sha1, 20);\n \t\thdrlen += 20;\n \t\treused_delta++;\n+\t} else if (type == OBJ_REF_DELTA && pack_version == 4) {\n+\t\thdrlen += encode_sha1ref(&v4, entry->delta->idx.sha1,\n+\t\t\t\t\theader + hdrlen);\n+\t\tsha1write(f, header, hdrlen);\n+\t\treused_delta++;\n \t} else {\n \t\tif (limit && hdrlen + datalen + 20 >= limit) {\n \t\t\tunuse_pack(&w_curs);\n@@ -460,6 +518,9 @@ static unsigned long write_object(struct sha1file *f,\n \telse\n \t\tusable_delta = 0;\t/* base could end up in another pack */\n \n+\tif (pack_version == 4 && entry->type == OBJ_TREE)\n+\t\tusable_delta = 0;\n+\n \tif (!reuse_object)\n \t\tto_reuse = 0;\t/* explicit */\n \telse if (!entry->in_pack)\n@@ -477,6 +538,10 @@ static unsigned long write_object(struct sha1file *f,\n \t\t\t\t * and we do not need to deltify it.\n \t\t\t\t */\n \n+\tif (pack_version == 4 &&\n+\t     (entry->type == OBJ_TREE || entry->type == OBJ_COMMIT))\n+\t\tto_reuse = 0;\n+\n \tif (!f) {\n \t\tif (usable_delta && entry->delta->idx.offset < 2)\n \t\t\tentry->delta->idx.offset = 2;\n@@ -790,6 +855,8 @@ static void write_pack_file(void)\n \n \t\tif (!offset)\n \t\t\tdie_errno(\"unable to write pack header\");\n+\t\tif (pack_version == 4)\n+\t\t\toffset += packv4_write_tables(f, &v4);\n \t\tnr_written = 0;\n \t\tfor (; i < nr_objects; i++) {\n \t\t\tstruct object_entry *e = write_order[i];\n@@ -2107,6 +2174,8 @@ static void prepare_pack(int window, int depth)\n \t\tsort_dict_entries_by_hits(v4.commit_ident_table);\n \t\tsort_dict_entries_by_hits(v4.tree_path_table);\n \t\tv4.all_objs = xmalloc(nr_objects * sizeof(*v4.all_objs));\n+\t\tpack_idx_opts.version = 3;\n+\t\tallow_ofs_delta = 0;\n \t}\n \n \tget_object_details();\ndiff --git a/pack.h b/pack.h\nindex 4f10fa4..ccefdbe 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -8,7 +8,7 @@\n  * Packed object header\n  */\n #define PACK_SIGNATURE 0x5041434b\t/* \"PACK\" */\n-#define pack_version_ok(v) ((v) == htonl(2) || (v) == htonl(3))\n+#define pack_version_ok(v) ((v) == htonl(2) || (v) == htonl(3) || (v) == htonl(4))\n struct pack_header {\n \tuint32_t hdr_signature;\n \tuint32_t hdr_version;\n-- \n1.8.2.83.gc99314b\n"},{"id":"227228","messageId":"1378735087-4813-15-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 14/16] pack v4: support \"end-of-pack\" indicator in index-pack and pack-objects","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:58:05Z","receivedAt":"2013-09-09T13:58:05Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"In v2, the number of objects in the pack header indicates how many\nobjects are sent. In v4 this is no longer true, that number includes\nthe base objects ommitted by pack-objects. An \"end-of-pack\" is\ninserted just before the final SHA-1 to let index-pack knows when to\nstop. The EOP is zero (in variable length encoding it means type zero,\nOBJ_NONE, and size zero)\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c   | 29 +++++++++++++++++++++++++----\n builtin/pack-objects.c | 15 +++++++++++----\n 2 files changed, 36 insertions(+), 8 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 88340b5..9036f3e 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -1493,7 +1493,7 @@ static void parse_dictionaries(void)\n  */\n static void parse_pack_objects(unsigned char *sha1)\n {\n-\tint i, nr_delays = 0;\n+\tint i, nr_delays = 0, eop = 0;\n \tstruct stat st;\n \n \tif (verbose)\n@@ -1502,7 +1502,28 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\t\tnr_objects);\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n-\t\tvoid *data = unpack_raw_entry(obj, obj->idx.sha1);\n+\t\tvoid *data;\n+\n+\t\tif (packv4) {\n+\t\t\tunsigned char *eop_byte;\n+\t\t\tflush();\n+\t\t\t/* Got End-of-Pack signal? */\n+\t\t\teop_byte = fill(1);\n+\t\t\tif (*eop_byte == 0) {\n+\t\t\t\tgit_SHA1_Update(&input_ctx, eop_byte, 1);\n+\t\t\t\tuse(1);\n+\t\t\t\t/*\n+\t\t\t\t * consumed by is used to mark the end\n+\t\t\t\t * of the object right after this\n+\t\t\t\t * loop. Undo use() effect.\n+\t\t\t\t */\n+\t\t\t\tconsumed_bytes--;\n+\t\t\t\teop = 1; /* so we don't flush() again */\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t}\n+\n+\t\tdata = unpack_raw_entry(obj, obj->idx.sha1);\n \t\tif (is_delta_type(obj->type) || is_delta_tree(obj)) {\n \t\t\t/* delay sha1_object() until second pass */\n \t\t} else if (!data) {\n@@ -1521,8 +1542,8 @@ static void parse_pack_objects(unsigned char *sha1)\n \tobjects[i].idx.offset = consumed_bytes;\n \tstop_progress(&progress);\n \n-\t/* Check pack integrity */\n-\tflush();\n+\tif (!eop)\n+\t\tflush();\t/* Check pack integrity */\n \tgit_SHA1_Final(sha1, &input_ctx);\n \tif (hashcmp(fill(20), sha1))\n \t\tdie(_(\"pack is corrupted (SHA1 mismatch)\"));\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex 12d9af4..1efb728 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -865,15 +865,22 @@ static void write_pack_file(void)\n \t\t\tdisplay_progress(progress_state, written);\n \t\t}\n \n-\t\t/*\n-\t\t * Did we write the wrong # entries in the header?\n-\t\t * If so, rewrite it like in fast-import\n-\t\t */\n \t\tif (pack_to_stdout) {\n+\t\t\tunsigned char type_zero = 0;\n+\t\t\t/*\n+\t\t\t * Pack v4 thin pack is terminated by a \"type\n+\t\t\t * 0, size 0\" in variable length encoding\n+\t\t\t */\n+\t\t\tif (pack_version == 4 && nr_written < nr_objects)\n+\t\t\t\tsha1write(f, &type_zero, 1);\n \t\t\tsha1close(f, sha1, CSUM_CLOSE);\n \t\t} else if (nr_written == nr_remaining) {\n \t\t\tsha1close(f, sha1, CSUM_FSYNC);\n \t\t} else {\n+\t\t\t/*\n+\t\t\t * Did we write the wrong # entries in the header?\n+\t\t\t * If so, rewrite it like in fast-import\n+\t\t\t */\n \t\t\tint fd = sha1close(f, sha1, 0);\n \t\t\tfixup_pack_header_footer(fd, sha1, pack_tmp_name,\n \t\t\t\t\t\t nr_written, sha1, offset);\n-- \n1.8.2.83.gc99314b\n"},{"id":"227229","messageId":"1378735087-4813-16-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 15/16] index-pack: use nr_objects_final as sha1_table size","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:58:06Z","receivedAt":"2013-09-09T13:58:06Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"nr_objects in the next patch is used to reflect the number of actual\nobjects in the stream, which may be smaller than the number recorded\nin pack header.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 13 +++++++------\n 1 file changed, 7 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 9036f3e..dc9961b 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -80,6 +80,7 @@ static int nr_objects;\n static int nr_deltas;\n static int nr_resolved_deltas;\n static int nr_threads;\n+static int nr_objects_final;\n \n static int from_stdin;\n static int strict;\n@@ -297,7 +298,7 @@ static void check_against_sha1table(const unsigned char *sha1)\n \tif (!packv4)\n \t\treturn;\n \n-\tfound = bsearch(sha1, sha1_table, nr_objects, 20,\n+\tfound = bsearch(sha1, sha1_table, nr_objects_final, 20,\n \t\t\t(int (*)(const void *, const void *))hashcmp);\n \tif (!found)\n \t\tdie(_(\"object %s not found in SHA-1 table\"),\n@@ -331,7 +332,7 @@ static const unsigned char *read_sha1ref(void)\n \t\treturn sha1;\n \t}\n \tindex--;\n-\tif (index >= nr_objects)\n+\tif (index >= nr_objects_final)\n \t\tbad_object(consumed_bytes,\n \t\t\t   _(\"bad index in read_sha1ref\"));\n \treturn sha1_table + index * 20;\n@@ -340,7 +341,7 @@ static const unsigned char *read_sha1ref(void)\n static const unsigned char *read_sha1table_ref(void)\n {\n \tconst unsigned char *sha1 = read_sha1ref();\n-\tif (sha1 < sha1_table || sha1 >= sha1_table + nr_objects * 20)\n+\tif (sha1 < sha1_table || sha1 >= sha1_table + nr_objects_final * 20)\n \t\tcheck_against_sha1table(sha1);\n \treturn sha1;\n }\n@@ -392,7 +393,7 @@ static void parse_pack_header(void)\n \t\tdie(_(\"pack version %\"PRIu32\" unsupported\"),\n \t\t\tntohl(hdr->hdr_version));\n \n-\tnr_objects = ntohl(hdr->hdr_entries);\n+\tnr_objects_final = nr_objects = ntohl(hdr->hdr_entries);\n \tuse(sizeof(struct pack_header));\n }\n \n@@ -1472,9 +1473,9 @@ static void parse_dictionaries(void)\n \tif (!packv4)\n \t\treturn;\n \n-\tsha1_table = xmalloc(20 * nr_objects);\n+\tsha1_table = xmalloc(20 * nr_objects_final);\n \thashcpy(sha1_table, fill_and_use(20));\n-\tfor (i = 1; i < nr_objects; i++) {\n+\tfor (i = 1; i < nr_objects_final; i++) {\n \t\tunsigned char *p = sha1_table + i * 20;\n \t\thashcpy(p, fill_and_use(20));\n \t\tif (hashcmp(p - 20, p) >= 0)\n-- \n1.8.2.83.gc99314b\n"},{"id":"227230","messageId":"1378735087-4813-17-git-send-email-pclouds@gmail.com","threadId":"34852","inReplyTo":"1378735087-4813-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 16/16] index-pack: support completing thin packs v4","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-09T13:58:07Z","receivedAt":"2013-09-09T13:58:07Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c | 53 +++++++++++++++++++++++++++++++++++++---------------\n 1 file changed, 38 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex dc9961b..8a6e2a3 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -97,7 +97,7 @@ static unsigned char input_buffer[4096];\n static unsigned int input_offset, input_len;\n static off_t consumed_bytes;\n static unsigned deepest_delta;\n-static git_SHA_CTX input_ctx;\n+static git_SHA_CTX input_ctx, output_ctx;\n static uint32_t input_crc32;\n static int input_fd, output_fd, pack_fd;\n \n@@ -1511,6 +1511,7 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\t/* Got End-of-Pack signal? */\n \t\t\teop_byte = fill(1);\n \t\t\tif (*eop_byte == 0) {\n+\t\t\t\toutput_ctx = input_ctx;\n \t\t\t\tgit_SHA1_Update(&input_ctx, eop_byte, 1);\n \t\t\t\tuse(1);\n \t\t\t\t/*\n@@ -1540,7 +1541,8 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\tfree(data);\n \t\tdisplay_progress(progress, i+1);\n \t}\n-\tobjects[i].idx.offset = consumed_bytes;\n+\tnr_objects = i;\n+\tobjects[nr_objects].idx.offset = consumed_bytes;\n \tstop_progress(&progress);\n \n \tif (!eop)\n@@ -1634,7 +1636,7 @@ static void conclude_pack(int fix_thin_pack, const char *curr_pack, unsigned cha\n \t\treturn;\n \t}\n \n-\tif (fix_thin_pack) {\n+\tif (fix_thin_pack && !packv4) {\n \t\tstruct sha1file *f;\n \t\tunsigned char read_sha1[20], tail_sha1[20];\n \t\tstruct strbuf msg = STRBUF_INIT;\n@@ -1661,6 +1663,26 @@ static void conclude_pack(int fix_thin_pack, const char *curr_pack, unsigned cha\n \t\tif (hashcmp(read_sha1, tail_sha1) != 0)\n \t\t\tdie(_(\"Unexpected tail checksum for %s \"\n \t\t\t      \"(disk corruption?)\"), curr_pack);\n+\t} else\tif (fix_thin_pack && packv4) {\n+\t\tstruct sha1file *f;\n+\t\tstruct strbuf msg = STRBUF_INIT;\n+\t\tint nr_unresolved = nr_deltas - nr_resolved_deltas;\n+\t\tint nr_objects_initial = nr_objects;\n+\t\tif (nr_unresolved <= 0)\n+\t\t\tdie(_(\"confusion beyond insanity\"));\n+\t\tf = sha1fd(output_fd, curr_pack);\n+\t\tf->ctx = output_ctx; /* resume sha-1 from right before EOP */\n+\t\tfix_unresolved_deltas(f, nr_unresolved);\n+\t\tif (nr_objects != nr_objects_final)\n+\t\t\tdie(_(\"pack number inconsistency, expected %u got %u\"),\n+\t\t\t    nr_objects, nr_objects_final);\n+\t\tstrbuf_addf(&msg, _(\"completed with %d local objects\"),\n+\t\t\t    nr_objects_final - nr_objects_initial);\n+\t\tstop_progress_msg(&progress, msg.buf);\n+\t\tstrbuf_release(&msg);\n+\t\tsha1close(f, pack_sha1, 0);\n+\t\twrite_or_die(output_fd, pack_sha1, 20);\n+\t\tfsync_or_die(output_fd, f->name);\n \t}\n \tif (nr_deltas != nr_resolved_deltas)\n \t\tdie(Q_(\"pack has %d unresolved delta\",\n@@ -1700,16 +1722,15 @@ static struct object_entry *append_obj_to_pack(struct sha1file *f,\n {\n \tstruct object_entry *obj = &objects[nr_objects++];\n \tunsigned char header[10];\n-\tunsigned long s = size;\n-\tint n = 0;\n-\tunsigned char c = (type << 4) | (s & 15);\n-\ts >>= 4;\n-\twhile (s) {\n-\t\theader[n++] = c | 0x80;\n-\t\tc = s & 0x7f;\n-\t\ts >>= 7;\n-\t}\n-\theader[n++] = c;\n+\tint n;\n+\n+\tif (packv4) {\n+\t\tif (nr_objects > nr_objects_final)\n+\t\t\tdie(_(\"too many objects\"));\n+\t\t/* TODO: convert OBJ_TREE to OBJ_PV4_TREE using pv4_encode_tree */\n+\t\tn = pv4_encode_object_header(type, size, header);\n+\t} else\n+\t\tn = encode_in_pack_object_header(type, size, header);\n \tcrc32_begin(f);\n \tsha1write(f, header, n);\n \tobj[0].size = size;\n@@ -1748,7 +1769,8 @@ static void fix_unresolved_deltas(struct sha1file *f, int nr_unresolved)\n \t */\n \tsorted_by_pos = xmalloc(nr_unresolved * sizeof(*sorted_by_pos));\n \tfor (i = 0; i < nr_deltas; i++) {\n-\t\tif (objects[deltas[i].obj_no].real_type != OBJ_REF_DELTA)\n+\t\tstruct object_entry *obj = objects + deltas[i].obj_no;\n+\t\tif (obj->real_type != OBJ_REF_DELTA && !is_delta_tree(obj))\n \t\t\tcontinue;\n \t\tsorted_by_pos[n++] = &deltas[i];\n \t}\n@@ -1756,10 +1778,11 @@ static void fix_unresolved_deltas(struct sha1file *f, int nr_unresolved)\n \n \tfor (i = 0; i < n; i++) {\n \t\tstruct delta_entry *d = sorted_by_pos[i];\n+\t\tstruct object_entry *obj = objects + d->obj_no;\n \t\tenum object_type type;\n \t\tstruct base_data *base_obj = alloc_base_data();\n \n-\t\tif (objects[d->obj_no].real_type != OBJ_REF_DELTA)\n+\t\tif (obj->real_type != OBJ_REF_DELTA && !is_delta_tree(obj))\n \t\t\tcontinue;\n \t\tbase_obj->data = read_sha1_file(d->base.sha1, &type, &base_obj->size);\n \t\tif (!base_obj->data)\n-- \n1.8.2.83.gc99314b\n"},{"id":"227232","messageId":"alpine.LFD.2.03.1309091047510.20709@syhkavp.arg","threadId":"34852","inReplyTo":"1378735087-4813-16-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH v2 15/16] index-pack: use nr_objects_final as sha1_table size","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-09T15:01:10Z","receivedAt":"2013-09-09T15:01:10Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 9 Sep 2013, Nguyễn Thái Ngọc Duy wrote:\n\n> nr_objects in the next patch is used to reflect the number of actual\n> objects in the stream, which may be smaller than the number recorded\n> in pack header.\n\nThis highlights an issue that has been nagging me for a while.\n\nWe decided to send the final number of objects in the thin pack header \nfor two reasons:\n\n1) it allows to properly size the SHA1 table upfront which already \n   contains entries for the omitted objects;\n\n2) the whole pack doesn't have to be re-summed again after being \n   completed on the receiving end since we don't alter the header.\n\nHowever this means that the progress meter will now be wrong and that's \nterrible !  Users *will* complain that the meter doesn't reach 100% and \nthey'll protest for being denied the remaining objects during the \ntransfer !\n\nJoking aside, we should think about doing something about it.  I was \nwondering if some kind of prefix to the pack stream could be inserted \nonto the wire when sending a pack v4.  Something like:\n\n'T', 'H', 'I', 'N', <actual_number_of_sent_objects_in_network_order>\n\nThis 8-byte prefix would simply be discarded by index-pack after being \nparsed.\n\nWhat do you think?\n\n\nNicolas\n"},{"id":"227233","messageId":"xmqqioya58cl.fsf@gitster.dls.corp.google.com","threadId":"34852","inReplyTo":"alpine.LFD.2.03.1309090900210.20709@syhkavp.arg","subject":"Re: [PATCH 08/11] pack-objects: create pack v4 tables","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-09-09T15:21:30Z","receivedAt":"2013-09-09T15:21:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@fluxnic.net> writes:\n\n> Is anyone still using --max-pack-size ?\n>\n> I'm wondering if producing multiple packs from pack-objects is really \n> useful these days.  If I remember correctly, this was created to allow \n> the archiving of large packs onto CDROMs or the like.\n\nI thought this was more about using a packfile on smaller\n(e.g. 32-bit) systems, but I may be mistaken.  2b84b5a8 (Introduce\nthe config variable pack.packSizeLimit, 2008-02-05) mentions\n\"filesystem constraints\":\n\n    Introduce the config variable pack.packSizeLimit\n    \n    \"git pack-objects\" has the option --max-pack-size to limit the\n    file size of the packs to a certain amount of bytes.  On\n    platforms where the pack file size is limited by filesystem\n    constraints, it is easy to forget this option, and this option\n    does not exist for \"git gc\" to begin with.\n"},{"id":"227240","messageId":"xmqq61u94zew.fsf@gitster.dls.corp.google.com","threadId":"34852","inReplyTo":"alpine.LFD.2.03.1309091047510.20709@syhkavp.arg","subject":"Re: [PATCH v2 15/16] index-pack: use nr_objects_final as sha1_table size","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-09-09T18:34:31Z","receivedAt":"2013-09-09T18:34:31Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@fluxnic.net> writes:\n\n> On Mon, 9 Sep 2013, Nguyễn Thái Ngọc Duy wrote:\n>\n>> nr_objects in the next patch is used to reflect the number of actual\n>> objects in the stream, which may be smaller than the number recorded\n>> in pack header.\n>\n> This highlights an issue that has been nagging me for a while.\n>\n> We decided to send the final number of objects in the thin pack header \n> for two reasons:\n>\n> 1) it allows to properly size the SHA1 table upfront which already \n>    contains entries for the omitted objects;\n>\n> 2) the whole pack doesn't have to be re-summed again after being \n>    completed on the receiving end since we don't alter the header.\n>\n> However this means that the progress meter will now be wrong and that's \n> terrible !  Users *will* complain that the meter doesn't reach 100% and \n> they'll protest for being denied the remaining objects during the \n> transfer !\n>\n> Joking aside, we should think about doing something about it.  I was \n> wondering if some kind of prefix to the pack stream could be inserted \n> onto the wire when sending a pack v4.  Something like:\n>\n> 'T', 'H', 'I', 'N', <actual_number_of_sent_objects_in_network_order>\n>\n> This 8-byte prefix would simply be discarded by index-pack after being \n> parsed.\n>\n> What do you think?\n\nI do not think it is _too_ bad if the meter jumped from 92% to 100%\nwhen we finish reading from the other end ;-), as long as we can\nreliably tell that we read the right thing.\n\nWhich brings me to a tangent.  Do we have a means to make sure that\nthe data received over the wire is bit-for-bit correct as a whole\nwhen it is a thin pack stream?  When it is a non-thin pack stream,\nwe have the checksum at the end added by sha1close() which\nindex-pack.c::parse_pack_objects() can (and does) verify.\n"},{"id":"227242","messageId":"alpine.LFD.2.03.1309091441540.20709@syhkavp.arg","threadId":"34852","inReplyTo":"xmqq61u94zew.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 15/16] index-pack: use nr_objects_final as sha1_table size","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-09T18:46:54Z","receivedAt":"2013-09-09T18:46:54Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 9 Sep 2013, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@fluxnic.net> writes:\n> \n> > On Mon, 9 Sep 2013, Nguyễn Thái Ngọc Duy wrote:\n> >\n> >> nr_objects in the next patch is used to reflect the number of actual\n> >> objects in the stream, which may be smaller than the number recorded\n> >> in pack header.\n> >\n> > This highlights an issue that has been nagging me for a while.\n> >\n> > We decided to send the final number of objects in the thin pack header \n> > for two reasons:\n> >\n> > 1) it allows to properly size the SHA1 table upfront which already \n> >    contains entries for the omitted objects;\n> >\n> > 2) the whole pack doesn't have to be re-summed again after being \n> >    completed on the receiving end since we don't alter the header.\n> >\n> > However this means that the progress meter will now be wrong and that's \n> > terrible !  Users *will* complain that the meter doesn't reach 100% and \n> > they'll protest for being denied the remaining objects during the \n> > transfer !\n> >\n> > Joking aside, we should think about doing something about it.  I was \n> > wondering if some kind of prefix to the pack stream could be inserted \n> > onto the wire when sending a pack v4.  Something like:\n> >\n> > 'T', 'H', 'I', 'N', <actual_number_of_sent_objects_in_network_order>\n> >\n> > This 8-byte prefix would simply be discarded by index-pack after being \n> > parsed.\n> >\n> > What do you think?\n> \n> I do not think it is _too_ bad if the meter jumped from 92% to 100%\n> when we finish reading from the other end ;-), as long as we can\n> reliably tell that we read the right thing.\n\nSure.  but eventually people will complain about this.  So while we're \nabout to introduce a new pack format anyway, better think of this little \ncosmetic detail now when it can be included in the pack v4 capability \nnegociation.\n\n> Which brings me to a tangent.  Do we have a means to make sure that\n> the data received over the wire is bit-for-bit correct as a whole\n> when it is a thin pack stream?  When it is a non-thin pack stream,\n> we have the checksum at the end added by sha1close() which\n> index-pack.c::parse_pack_objects() can (and does) verify.\n\nThe trailing checksum is still there.  Nothing has changed in that \nregard.\n\n\nNicolas\n"},{"id":"227244","messageId":"xmqqwqmp3jtj.fsf@gitster.dls.corp.google.com","threadId":"34852","inReplyTo":"alpine.LFD.2.03.1309091441540.20709@syhkavp.arg","subject":"Re: [PATCH v2 15/16] index-pack: use nr_objects_final as sha1_table size","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-09-09T18:56:40Z","receivedAt":"2013-09-09T18:56:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@fluxnic.net> writes:\n\n>> > ...  I was \n>> > wondering if some kind of prefix to the pack stream could be inserted \n>> > onto the wire when sending a pack v4.  Something like:\n>> >\n>> > 'T', 'H', 'I', 'N', <actual_number_of_sent_objects_in_network_order>\n>> >\n>> > This 8-byte prefix would simply be discarded by index-pack after being \n>> > parsed.\n>> >\n>> > What do you think?\n>> \n>> I do not think it is _too_ bad if the meter jumped from 92% to 100%\n>> when we finish reading from the other end ;-), as long as we can\n>> reliably tell that we read the right thing.\n>\n> Sure.  but eventually people will complain about this.  So while we're \n> about to introduce a new pack format anyway, better think of this little \n> cosmetic detail now when it can be included in the pack v4 capability \n> negociation.\n\nOh, I completely agree on that part.  When we send a self-contained\npack, would we send nothing?  That is, should the receiving end\nexpect and rely on that the sending end will send a thin pack and\nnever a fat pack when asked to send a thin pack (and vice versa)?\n\nAlso should we make the \"even though we have negotiated the protocol\nparameters, after enumerating the objects and deciding what the pack\nstream would look like, we have a bit more information to tell you\"\nthe sending side gives the receiver extensible?  I am wondering if\nthat prefix needs something like \"end of prefix\" marker (or \"here\ncomes N-bytes worth of prefix information\" upfront); we probably do\nnot need it, as the capability exchange will determine what kind of\ninformation will be sent (e.g. \"actual objects in the thin pack data\nstream\").\n"},{"id":"227246","messageId":"alpine.LFD.2.03.1309091507290.20709@syhkavp.arg","threadId":"34852","inReplyTo":"xmqqwqmp3jtj.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 15/16] index-pack: use nr_objects_final as sha1_table size","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-09T19:11:38Z","receivedAt":"2013-09-09T19:11:38Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 9 Sep 2013, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@fluxnic.net> writes:\n> \n> >> > ...  I was \n> >> > wondering if some kind of prefix to the pack stream could be inserted \n> >> > onto the wire when sending a pack v4.  Something like:\n> >> >\n> >> > 'T', 'H', 'I', 'N', <actual_number_of_sent_objects_in_network_order>\n> >> >\n> >> > This 8-byte prefix would simply be discarded by index-pack after being \n> >> > parsed.\n> >> >\n> >> > What do you think?\n> >> \n> >> I do not think it is _too_ bad if the meter jumped from 92% to 100%\n> >> when we finish reading from the other end ;-), as long as we can\n> >> reliably tell that we read the right thing.\n> >\n> > Sure.  but eventually people will complain about this.  So while we're \n> > about to introduce a new pack format anyway, better think of this little \n> > cosmetic detail now when it can be included in the pack v4 capability \n> > negociation.\n> \n> Oh, I completely agree on that part.  When we send a self-contained\n> pack, would we send nothing?  That is, should the receiving end\n> expect and rely on that the sending end will send a thin pack and\n> never a fat pack when asked to send a thin pack (and vice versa)?\n> \n> Also should we make the \"even though we have negotiated the protocol\n> parameters, after enumerating the objects and deciding what the pack\n> stream would look like, we have a bit more information to tell you\"\n> the sending side gives the receiver extensible?  I am wondering if\n> that prefix needs something like \"end of prefix\" marker (or \"here\n> comes N-bytes worth of prefix information\" upfront); we probably do\n> not need it, as the capability exchange will determine what kind of\n> information will be sent (e.g. \"actual objects in the thin pack data\n> stream\").\n\nDo we know the actual number of objects to send during the capability \nnegociation?  I don't think so as this is known only after the \n\"compressing objects\" phase, and that already depends on the capability \nnegociation before it can start.\n\n\nNicolas\n"},{"id":"227249","messageId":"xmqqob813i9c.fsf@gitster.dls.corp.google.com","threadId":"34852","inReplyTo":"alpine.LFD.2.03.1309091507290.20709@syhkavp.arg","subject":"Re: [PATCH v2 15/16] index-pack: use nr_objects_final as sha1_table size","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-09-09T19:30:23Z","receivedAt":"2013-09-09T19:30:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@fluxnic.net> writes:\n\n> Do we know the actual number of objects to send during the capability \n> negociation?\n\nNo, and that is not what I meant.  We know upfront after capability\nnegotiation (by seeing a request to give them a thin-pack) that we\nwill send, in addition to the usual packfile, the prefix that\ncarries that information and that is the important part.  That lets\nthe receiver decide whether to _expect_ to see the prefix or no\nprefix.  Without such, there needs some clue in the prefix part\nitself if there are prefixes that carry information computed after\ncapability negotiation finished (i.e. after \"object enumeration\").\n\nSorry if I was unclear.\n"},{"id":"227253","messageId":"alpine.LFD.2.03.1309091553540.20709@syhkavp.arg","threadId":"34852","inReplyTo":"xmqqob813i9c.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 15/16] index-pack: use nr_objects_final as sha1_table size","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-09T19:56:20Z","receivedAt":"2013-09-09T19:56:20Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 9 Sep 2013, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@fluxnic.net> writes:\n> \n> > Do we know the actual number of objects to send during the capability \n> > negociation?\n> \n> No, and that is not what I meant.  We know upfront after capability\n> negotiation (by seeing a request to give them a thin-pack) that we\n> will send, in addition to the usual packfile, the prefix that\n> carries that information and that is the important part.  That lets\n> the receiver decide whether to _expect_ to see the prefix or no\n> prefix.  Without such, there needs some clue in the prefix part\n> itself if there are prefixes that carry information computed after\n> capability negotiation finished (i.e. after \"object enumeration\").\n\nIn this case, if negociation concludes on \"thin\" and \"pack-version=4\" \nthen that could mean there is a prefix to be expected.\n\n\nNicolas\n"},{"id":"227289","messageId":"CACsJy8D=72RsnNKs-6EdUnhu6kurvQ5S2z5PkV3KcY7SUDHJKg@mail.gmail.com","threadId":"34852","inReplyTo":"alpine.LFD.2.03.1309091047510.20709@syhkavp.arg","subject":"Re: [PATCH v2 15/16] index-pack: use nr_objects_final as sha1_table size","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-09-10T00:45:45Z","receivedAt":"2013-09-10T00:45:45Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Sep 9, 2013 at 10:01 PM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> However this means that the progress meter will now be wrong and that's\n> terrible !  Users *will* complain that the meter doesn't reach 100% and\n> they'll protest for being denied the remaining objects during the\n> transfer !\n>\n> Joking aside, we should think about doing something about it.  I was\n> wondering if some kind of prefix to the pack stream could be inserted\n> onto the wire when sending a pack v4.  Something like:\n>\n> 'T', 'H', 'I', 'N', <actual_number_of_sent_objects_in_network_order>\n>\n> This 8-byte prefix would simply be discarded by index-pack after being\n> parsed.\n>\n> What do you think?\n\nI have no problem with this. Although I rather we generalize the case\nto support multiple packs in the same stream (in some case the server\ncan just stream away one big existing pack, followed by a smaller pack\nof recent updates), where \"thin\" is just a special pack that is not\nsaved on disk. So except for the signature difference, it should at\nleast follow the pack header (sig, version, nr_objects)\n-- \nDuy\n"},{"id":"227536","messageId":"alpine.LFD.2.03.1309121132050.20709@syhkavp.arg","threadId":"34852","inReplyTo":"CACsJy8D=72RsnNKs-6EdUnhu6kurvQ5S2z5PkV3KcY7SUDHJKg@mail.gmail.com","subject":"Re: [PATCH v2 15/16] index-pack: use nr_objects_final as sha1_table size","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2013-09-12T15:34:14Z","receivedAt":"2013-09-12T15:34:14Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 10 Sep 2013, Duy Nguyen wrote:\n\n> On Mon, Sep 9, 2013 at 10:01 PM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> > However this means that the progress meter will now be wrong and that's\n> > terrible !  Users *will* complain that the meter doesn't reach 100% and\n> > they'll protest for being denied the remaining objects during the\n> > transfer !\n> >\n> > Joking aside, we should think about doing something about it.  I was\n> > wondering if some kind of prefix to the pack stream could be inserted\n> > onto the wire when sending a pack v4.  Something like:\n> >\n> > 'T', 'H', 'I', 'N', <actual_number_of_sent_objects_in_network_order>\n> >\n> > This 8-byte prefix would simply be discarded by index-pack after being\n> > parsed.\n> >\n> > What do you think?\n> \n> I have no problem with this. Although I rather we generalize the case\n> to support multiple packs in the same stream (in some case the server\n> can just stream away one big existing pack, followed by a smaller pack\n> of recent updates), where \"thin\" is just a special pack that is not\n> saved on disk. So except for the signature difference, it should at\n> least follow the pack header (sig, version, nr_objects)\n\nExcept in this case this is not a separate pack.  This prefix is there \nto provide information that is valid only for the pack to follow and \ntherefore cannot be considered as some independent data.\n\n\nNicolas\n"}]}