{"thread":{"id":"29558","subject":"[PATCH 1/6] read-cache: use sha1file for sha1 calculation","startedAt":"2012-02-06T05:48:34Z","lastAt":"2012-02-07T17:25:43Z","messageCount":20,"participants":["Nguyễn Thái Ngọc Duy","Junio C Hamano","Nguyen Thai Ngoc Duy","Dave Zarzycki","Shawn Pearce","Thomas Rast"],"isPatch":true,"patchVersion":1,"patchTotal":6},"messages":[{"id":"183971","messageId":"1328507319-24687-1-git-send-email-pclouds@gmail.com","threadId":"29558","inReplyTo":null,"subject":"[PATCH 1/6] read-cache: use sha1file for sha1 calculation","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-06T05:48:34Z","receivedAt":"2012-02-06T05:48:34Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c |   90 +++++++++++++++-------------------------------------------\n 1 files changed, 23 insertions(+), 67 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex a51bba1..e9a20b6 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -12,6 +12,7 @@\n #include \"commit.h\"\n #include \"blob.h\"\n #include \"resolve-undo.h\"\n+#include \"csum-file.h\"\n \n static struct cache_entry *refresh_cache_entry(struct cache_entry *ce, int really);\n \n@@ -1395,73 +1396,28 @@ int unmerged_index(const struct index_state *istate)\n \treturn 0;\n }\n \n-#define WRITE_BUFFER_SIZE 8192\n-static unsigned char write_buffer[WRITE_BUFFER_SIZE];\n-static unsigned long write_buffer_len;\n-\n-static int ce_write_flush(git_SHA_CTX *context, int fd)\n-{\n-\tunsigned int buffered = write_buffer_len;\n-\tif (buffered) {\n-\t\tgit_SHA1_Update(context, write_buffer, buffered);\n-\t\tif (write_in_full(fd, write_buffer, buffered) != buffered)\n-\t\t\treturn -1;\n-\t\twrite_buffer_len = 0;\n-\t}\n-\treturn 0;\n-}\n-\n-static int ce_write(git_SHA_CTX *context, int fd, void *data, unsigned int len)\n+static int ce_write(struct sha1file *f, void *data, unsigned int len)\n {\n-\twhile (len) {\n-\t\tunsigned int buffered = write_buffer_len;\n-\t\tunsigned int partial = WRITE_BUFFER_SIZE - buffered;\n-\t\tif (partial > len)\n-\t\t\tpartial = len;\n-\t\tmemcpy(write_buffer + buffered, data, partial);\n-\t\tbuffered += partial;\n-\t\tif (buffered == WRITE_BUFFER_SIZE) {\n-\t\t\twrite_buffer_len = buffered;\n-\t\t\tif (ce_write_flush(context, fd))\n-\t\t\t\treturn -1;\n-\t\t\tbuffered = 0;\n-\t\t}\n-\t\twrite_buffer_len = buffered;\n-\t\tlen -= partial;\n-\t\tdata = (char *) data + partial;\n-\t}\n-\treturn 0;\n+\treturn sha1write(f, data, len);\n }\n \n-static int write_index_ext_header(git_SHA_CTX *context, int fd,\n+static int write_index_ext_header(struct sha1file *f,\n \t\t\t\t  unsigned int ext, unsigned int sz)\n {\n \text = htonl(ext);\n \tsz = htonl(sz);\n-\treturn ((ce_write(context, fd, &ext, 4) < 0) ||\n-\t\t(ce_write(context, fd, &sz, 4) < 0)) ? -1 : 0;\n+\treturn ((ce_write(f, &ext, 4) < 0) ||\n+\t\t(ce_write(f, &sz, 4) < 0)) ? -1 : 0;\n }\n \n-static int ce_flush(git_SHA_CTX *context, int fd)\n+static int ce_flush(struct sha1file *f)\n {\n-\tunsigned int left = write_buffer_len;\n-\n-\tif (left) {\n-\t\twrite_buffer_len = 0;\n-\t\tgit_SHA1_Update(context, write_buffer, left);\n-\t}\n-\n-\t/* Flush first if not enough space for SHA1 signature */\n-\tif (left + 20 > WRITE_BUFFER_SIZE) {\n-\t\tif (write_in_full(fd, write_buffer, left) != left)\n-\t\t\treturn -1;\n-\t\tleft = 0;\n-\t}\n+\tunsigned char sha1[20];\n+\tint fd = sha1close(f, sha1, 0);\n \n-\t/* Append the SHA1 signature at the end */\n-\tgit_SHA1_Final(write_buffer + left, context);\n-\tleft += 20;\n-\treturn (write_in_full(fd, write_buffer, left) != left) ? -1 : 0;\n+\tif (fd < 0)\n+\t\treturn -1;\n+\treturn (write_in_full(fd, sha1, 20) != 20) ? -1 : 0;\n }\n \n static void ce_smudge_racily_clean_entry(struct cache_entry *ce)\n@@ -1513,7 +1469,7 @@ static void ce_smudge_racily_clean_entry(struct cache_entry *ce)\n \t}\n }\n \n-static int ce_write_entry(git_SHA_CTX *c, int fd, struct cache_entry *ce)\n+static int ce_write_entry(struct sha1file *f, struct cache_entry *ce)\n {\n \tint size = ondisk_ce_size(ce);\n \tstruct ondisk_cache_entry *ondisk = xcalloc(1, size);\n@@ -1542,7 +1498,7 @@ static int ce_write_entry(git_SHA_CTX *c, int fd, struct cache_entry *ce)\n \t\tname = ondisk->name;\n \tmemcpy(name, ce->name, ce_namelen(ce));\n \n-\tresult = ce_write(c, fd, ondisk, size);\n+\tresult = ce_write(f, ondisk, size);\n \tfree(ondisk);\n \treturn result;\n }\n@@ -1574,7 +1530,7 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n \n int write_index(struct index_state *istate, int newfd)\n {\n-\tgit_SHA_CTX c;\n+\tstruct sha1file *f;\n \tstruct cache_header hdr;\n \tint i, err, removed, extended;\n \tstruct cache_entry **cache = istate->cache;\n@@ -1598,8 +1554,8 @@ int write_index(struct index_state *istate, int newfd)\n \thdr.hdr_version = htonl(extended ? 3 : 2);\n \thdr.hdr_entries = htonl(entries - removed);\n \n-\tgit_SHA1_Init(&c);\n-\tif (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n+\tf = sha1fd(newfd, NULL);\n+\tif (ce_write(f, &hdr, sizeof(hdr)) < 0)\n \t\treturn -1;\n \n \tfor (i = 0; i < entries; i++) {\n@@ -1608,7 +1564,7 @@ int write_index(struct index_state *istate, int newfd)\n \t\t\tcontinue;\n \t\tif (!ce_uptodate(ce) && is_racy_timestamp(istate, ce))\n \t\t\tce_smudge_racily_clean_entry(ce);\n-\t\tif (ce_write_entry(&c, newfd, ce) < 0)\n+\t\tif (ce_write_entry(f, ce) < 0)\n \t\t\treturn -1;\n \t}\n \n@@ -1617,8 +1573,8 @@ int write_index(struct index_state *istate, int newfd)\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tcache_tree_write(&sb, istate->cache_tree);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_TREE, sb.len) < 0\n-\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\terr = write_index_ext_header(f, CACHE_EXT_TREE, sb.len) < 0\n+\t\t\t|| ce_write(f, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n \t\t\treturn -1;\n@@ -1627,15 +1583,15 @@ int write_index(struct index_state *istate, int newfd)\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tresolve_undo_write(&sb, istate->resolve_undo);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_RESOLVE_UNDO,\n+\t\terr = write_index_ext_header(f, CACHE_EXT_RESOLVE_UNDO,\n \t\t\t\t\t     sb.len) < 0\n-\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\t\t|| ce_write(f, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n \t\t\treturn -1;\n \t}\n \n-\tif (ce_flush(&c, newfd) || fstat(newfd, &st))\n+\tif (ce_flush(f) || fstat(newfd, &st))\n \t\treturn -1;\n \tistate->timestamp.sec = (unsigned int)st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n-- \n1.7.8.36.g69ee2\n"},{"id":"183972","messageId":"1328507319-24687-2-git-send-email-pclouds@gmail.com","threadId":"29558","inReplyTo":"1328507319-24687-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 2/6] csum-file: make sha1 calculation optional","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-06T05:48:35Z","receivedAt":"2012-02-06T05:48:35Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n csum-file.c |   16 +++++++++++++---\n csum-file.h |    2 +-\n 2 files changed, 14 insertions(+), 4 deletions(-)\n\ndiff --git a/csum-file.c b/csum-file.c\nindex 53f5375..4c517d1 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -11,6 +11,12 @@\n #include \"progress.h\"\n #include \"csum-file.h\"\n \n+static void sha1update(struct sha1file *f, const void *data, unsigned offset)\n+{\n+\tif (f->do_sha1)\n+\t\tgit_SHA1_Update(&f->ctx, data, offset);\n+}\n+\n static void flush(struct sha1file *f, void *buf, unsigned int count)\n {\n \tif (0 <= f->check_fd && count)  {\n@@ -47,7 +53,7 @@ void sha1flush(struct sha1file *f)\n \tunsigned offset = f->offset;\n \n \tif (offset) {\n-\t\tgit_SHA1_Update(&f->ctx, f->buffer, offset);\n+\t\tsha1update(f, f->buffer, offset);\n \t\tflush(f, f->buffer, offset);\n \t\tf->offset = 0;\n \t}\n@@ -58,7 +64,10 @@ int sha1close(struct sha1file *f, unsigned char *result, unsigned int flags)\n \tint fd;\n \n \tsha1flush(f);\n-\tgit_SHA1_Final(f->buffer, &f->ctx);\n+\tif (f->do_sha1)\n+\t\tgit_SHA1_Final(f->buffer, &f->ctx);\n+\telse\n+\t\thashclr(f->buffer);\n \tif (result)\n \t\thashcpy(result, f->buffer);\n \tif (flags & (CSUM_CLOSE | CSUM_FSYNC)) {\n@@ -110,7 +119,7 @@ int sha1write(struct sha1file *f, void *buf, unsigned int count)\n \t\tbuf = (char *) buf + nr;\n \t\tleft -= nr;\n \t\tif (!left) {\n-\t\t\tgit_SHA1_Update(&f->ctx, data, offset);\n+\t\t\tsha1update(f, data, offset);\n \t\t\tflush(f, data, offset);\n \t\t\toffset = 0;\n \t\t}\n@@ -154,6 +163,7 @@ struct sha1file *sha1fd_throughput(int fd, const char *name, struct progress *tp\n \tf->tp = tp;\n \tf->name = name;\n \tf->do_crc = 0;\n+\tf->do_sha1 = 1;\n \tgit_SHA1_Init(&f->ctx);\n \treturn f;\n }\ndiff --git a/csum-file.h b/csum-file.h\nindex 3b540bd..c23ea62 100644\n--- a/csum-file.h\n+++ b/csum-file.h\n@@ -12,7 +12,7 @@ struct sha1file {\n \toff_t total;\n \tstruct progress *tp;\n \tconst char *name;\n-\tint do_crc;\n+\tint do_crc, do_sha1;\n \tuint32_t crc32;\n \tunsigned char buffer[8192];\n };\n-- \n1.7.8.36.g69ee2\n"},{"id":"183973","messageId":"1328507319-24687-3-git-send-email-pclouds@gmail.com","threadId":"29558","inReplyTo":"1328507319-24687-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 3/6] Stop producing index version 2","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-06T05:48:36Z","receivedAt":"2012-02-06T05:48:36Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"read-cache.c learned to produce version 2 or 3 depending on whether\nextended cache entries exist in 06aaaa0 (Extend index to save more flags\n- 2008-10-01), first released in 1.6.1. The purpose is to keep\ncompatibility with older git. It's been more than three years since\nthen and git has reached 1.7.9. Drop support for older git.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c                          |    8 +++-----\n t/t2104-update-index-skip-worktree.sh |   12 ------------\n 2 files changed, 3 insertions(+), 17 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex e9a20b6..fe6b0e0 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1532,26 +1532,24 @@ int write_index(struct index_state *istate, int newfd)\n {\n \tstruct sha1file *f;\n \tstruct cache_header hdr;\n-\tint i, err, removed, extended;\n+\tint i, err, removed;\n \tstruct cache_entry **cache = istate->cache;\n \tint entries = istate->cache_nr;\n \tstruct stat st;\n \n-\tfor (i = removed = extended = 0; i < entries; i++) {\n+\tfor (i = removed = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n \t\t\tremoved++;\n \n \t\t/* reduce extended entries if possible */\n \t\tcache[i]->ce_flags &= ~CE_EXTENDED;\n \t\tif (cache[i]->ce_flags & CE_EXTENDED_FLAGS) {\n-\t\t\textended++;\n \t\t\tcache[i]->ce_flags |= CE_EXTENDED;\n \t\t}\n \t}\n \n \thdr.hdr_signature = htonl(CACHE_SIGNATURE);\n-\t/* for extended format, increase version so older git won't try to read it */\n-\thdr.hdr_version = htonl(extended ? 3 : 2);\n+\thdr.hdr_version = htonl(3);\n \thdr.hdr_entries = htonl(entries - removed);\n \n \tf = sha1fd(newfd, NULL);\ndiff --git a/t/t2104-update-index-skip-worktree.sh b/t/t2104-update-index-skip-worktree.sh\nindex 1d0879b..8221ffa 100755\n--- a/t/t2104-update-index-skip-worktree.sh\n+++ b/t/t2104-update-index-skip-worktree.sh\n@@ -28,19 +28,11 @@ test_expect_success 'setup' '\n \tgit ls-files -t | test_cmp expect.full -\n '\n \n-test_expect_success 'index is at version 2' '\n-\ttest \"$(test-index-version < .git/index)\" = 2\n-'\n-\n test_expect_success 'update-index --skip-worktree' '\n \tgit update-index --skip-worktree 1 sub/1 &&\n \tgit ls-files -t | test_cmp expect.skip -\n '\n \n-test_expect_success 'index is at version 3 after having some skip-worktree entries' '\n-\ttest \"$(test-index-version < .git/index)\" = 3\n-'\n-\n test_expect_success 'ls-files -t' '\n \tgit ls-files -t | test_cmp expect.skip -\n '\n@@ -50,8 +42,4 @@ test_expect_success 'update-index --no-skip-worktree' '\n \tgit ls-files -t | test_cmp expect.full -\n '\n \n-test_expect_success 'index version is back to 2 when there is no skip-worktree entry' '\n-\ttest \"$(test-index-version < .git/index)\" = 2\n-'\n-\n test_done\n-- \n1.7.8.36.g69ee2\n"},{"id":"183974","messageId":"1328507319-24687-4-git-send-email-pclouds@gmail.com","threadId":"29558","inReplyTo":"1328507319-24687-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 4/6] Introduce index version 4 with global flags","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-06T05:48:37Z","receivedAt":"2012-02-06T05:48:37Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"v4 adds 32-bit field to cache header after 32-bit number of entries.\nIf this field is zero, fall back to v3.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n Documentation/technical/index-format.txt |    4 ++-\n cache.h                                  |    6 +++++\n read-cache.c                             |   31 ++++++++++++++++++++++-------\n 3 files changed, 32 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/technical/index-format.txt b/Documentation/technical/index-format.txt\nindex 8930b3f..2b6a38e 100644\n--- a/Documentation/technical/index-format.txt\n+++ b/Documentation/technical/index-format.txt\n@@ -12,10 +12,12 @@ GIT index format\n        The signature is { 'D', 'I', 'R', 'C' } (stands for \"dircache\")\n \n      4-byte version number:\n-       The current supported versions are 2 and 3.\n+       The current supported versions are 2, 3 and 4.\n \n      32-bit number of index entries.\n \n+     32-bit flags (version 4 only).\n+\n    - A number of sorted index entries (see below).\n \n    - Extensions\ndiff --git a/cache.h b/cache.h\nindex 9bd8c2d..c2e884a 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -105,6 +105,11 @@ struct cache_header {\n \tunsigned int hdr_entries;\n };\n \n+struct ext_cache_header {\n+\tstruct cache_header h;\n+\tunsigned int hdr_flags;\n+};\n+\n /*\n  * The \"cache_time\" is just the low 32 bits of the\n  * time. It doesn't matter if it overflows - we only\n@@ -314,6 +319,7 @@ static inline unsigned int canon_mode(unsigned int mode)\n struct index_state {\n \tstruct cache_entry **cache;\n \tunsigned int cache_nr, cache_alloc, cache_changed;\n+\tunsigned int hdr_flags;\n \tstruct string_list *resolve_undo;\n \tstruct cache_tree *cache_tree;\n \tstruct cache_time timestamp;\ndiff --git a/read-cache.c b/read-cache.c\nindex fe6b0e0..fd21af6 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1190,7 +1190,9 @@ static int verify_hdr(struct cache_header *hdr, unsigned long size)\n \n \tif (hdr->hdr_signature != htonl(CACHE_SIGNATURE))\n \t\treturn error(\"bad signature\");\n-\tif (hdr->hdr_version != htonl(2) && hdr->hdr_version != htonl(3))\n+\tif (hdr->hdr_version != htonl(2) &&\n+\t    hdr->hdr_version != htonl(3) &&\n+\t    hdr->hdr_version != htonl(4))\n \t\treturn error(\"bad index version\");\n \tgit_SHA1_Init(&c);\n \tgit_SHA1_Update(&c, hdr, size - 20);\n@@ -1320,7 +1322,12 @@ int read_index_from(struct index_state *istate, const char *path)\n \tistate->cache = xcalloc(istate->cache_alloc, sizeof(struct cache_entry *));\n \tistate->initialized = 1;\n \n-\tsrc_offset = sizeof(*hdr);\n+\tif (ntohl(hdr->hdr_version) >= 4) {\n+\t\tstruct ext_cache_header *ehdr = mmap;\n+\t\tistate->hdr_flags = ntohl(ehdr->hdr_flags);\n+\t\tsrc_offset = sizeof(*ehdr);\n+\t} else\n+\t\tsrc_offset = sizeof(*hdr);\n \tfor (i = 0; i < istate->cache_nr; i++) {\n \t\tstruct ondisk_cache_entry *disk_ce;\n \t\tstruct cache_entry *ce;\n@@ -1375,6 +1382,7 @@ int discard_index(struct index_state *istate)\n \tresolve_undo_clear_index(istate);\n \tistate->cache_nr = 0;\n \tistate->cache_changed = 0;\n+\tistate->hdr_flags = 0;\n \tistate->timestamp.sec = 0;\n \tistate->timestamp.nsec = 0;\n \tistate->name_hash_initialized = 0;\n@@ -1531,8 +1539,8 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n int write_index(struct index_state *istate, int newfd)\n {\n \tstruct sha1file *f;\n-\tstruct cache_header hdr;\n-\tint i, err, removed;\n+\tstruct ext_cache_header hdr;\n+\tint i, err, removed, hdr_size;\n \tstruct cache_entry **cache = istate->cache;\n \tint entries = istate->cache_nr;\n \tstruct stat st;\n@@ -1548,12 +1556,19 @@ int write_index(struct index_state *istate, int newfd)\n \t\t}\n \t}\n \n-\thdr.hdr_signature = htonl(CACHE_SIGNATURE);\n-\thdr.hdr_version = htonl(3);\n-\thdr.hdr_entries = htonl(entries - removed);\n+\thdr.h.hdr_signature = htonl(CACHE_SIGNATURE);\n+\tif (istate->hdr_flags) {\n+\t\thdr.h.hdr_version = htonl(4);\n+\t\thdr.hdr_flags = htonl(istate->hdr_flags);\n+\t\thdr_size = sizeof(hdr);\n+\t} else {\n+\t\thdr.h.hdr_version = htonl(3);\n+\t\thdr_size = sizeof(hdr.h);\n+\t}\n+\thdr.h.hdr_entries = htonl(entries - removed);\n \n \tf = sha1fd(newfd, NULL);\n-\tif (ce_write(f, &hdr, sizeof(hdr)) < 0)\n+\tif (ce_write(f, &hdr, hdr_size) < 0)\n \t\treturn -1;\n \n \tfor (i = 0; i < entries; i++) {\n-- \n1.7.8.36.g69ee2\n"},{"id":"183975","messageId":"1328507319-24687-5-git-send-email-pclouds@gmail.com","threadId":"29558","inReplyTo":"1328507319-24687-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 5/6] Allow to use crc32 as a lighter checksum on index","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-06T05:48:38Z","receivedAt":"2012-02-06T05:48:38Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n Documentation/git-update-index.txt |   12 +++++++-\n builtin/update-index.c             |   11 +++++++\n cache.h                            |    2 +\n read-cache.c                       |   54 ++++++++++++++++++++++++++++--------\n 4 files changed, 66 insertions(+), 13 deletions(-)\n\ndiff --git a/Documentation/git-update-index.txt b/Documentation/git-update-index.txt\nindex a3081f4..2574a4e 100644\n--- a/Documentation/git-update-index.txt\n+++ b/Documentation/git-update-index.txt\n@@ -13,7 +13,7 @@ SYNOPSIS\n \t     [--add] [--remove | --force-remove] [--replace]\n \t     [--refresh] [-q] [--unmerged] [--ignore-missing]\n \t     [(--cacheinfo <mode> <object> <file>)...]\n-\t     [--chmod=(+|-)x]\n+\t     [--chmod=(+|-)x] [--[no-]crc32]\n \t     [--assume-unchanged | --no-assume-unchanged]\n \t     [--skip-worktree | --no-skip-worktree]\n \t     [--ignore-submodules]\n@@ -109,6 +109,16 @@ you will need to handle the situation manually.\n \tset and unset the \"skip-worktree\" bit for the paths. See\n \tsection \"Skip-worktree bit\" below for more information.\n \n+--crc32::\n+--no-crc32::\n+\tNormally SHA-1 is used to check for index integrity. When the\n+\tindex is large, SHA-1 computation cost can be significant.\n+\t--crc32 will convert current index to use (cheaper) crc32\n+\tinstead. Note that later writes to index by other commands can\n+\tconvert the index back to SHA-1. Older git versions may not\n+\tunderstand crc32 index, --no-crc32 can be used to convert it\n+\tback to SHA-1.\n+\n -g::\n --again::\n \tRuns 'git update-index' itself on the paths whose index\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex a6a23fa..6913226 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -707,6 +707,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n {\n \tint newfd, entries, has_errors = 0, line_termination = '\\n';\n \tint read_from_stdin = 0;\n+\tint do_crc = -1;\n \tint prefix_length = prefix ? strlen(prefix) : 0;\n \tchar set_executable_bit = 0;\n \tstruct refresh_params refresh_args = {0, &has_errors};\n@@ -791,6 +792,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\t\t\"(for porcelains) forget saved unresolved conflicts\",\n \t\t\tPARSE_OPT_NOARG | PARSE_OPT_NONEG,\n \t\t\tresolve_undo_clear_callback},\n+\t\tOPT_BOOL(0, \"crc32\", &do_crc,\n+\t\t\t \"use crc32 as checksum instead of sha1\"),\n \t\tOPT_END()\n \t};\n \n@@ -852,6 +855,14 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t}\n \targc = parse_options_end(&ctx);\n \n+\tif (do_crc != -1) {\n+\t\tif (do_crc)\n+\t\t\tthe_index.hdr_flags |= CACHE_F_CRC;\n+\t\telse\n+\t\t\tthe_index.hdr_flags &= ~CACHE_F_CRC;\n+\t\tactive_cache_changed = 1;\n+\t}\n+\n \tif (read_from_stdin) {\n \t\tstruct strbuf buf = STRBUF_INIT, nbuf = STRBUF_INIT;\n \ndiff --git a/cache.h b/cache.h\nindex c2e884a..7352402 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -105,6 +105,8 @@ struct cache_header {\n \tunsigned int hdr_entries;\n };\n \n+#define CACHE_F_CRC\t1\t/* use crc32 instead of sha1 for index checksum */\n+\n struct ext_cache_header {\n \tstruct cache_header h;\n \tunsigned int hdr_flags;\ndiff --git a/read-cache.c b/read-cache.c\nindex fd21af6..a34878e 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1185,20 +1185,33 @@ static struct cache_entry *refresh_cache_entry(struct cache_entry *ce, int reall\n \n static int verify_hdr(struct cache_header *hdr, unsigned long size)\n {\n+\tint do_crc;\n \tgit_SHA_CTX c;\n \tunsigned char sha1[20];\n \n \tif (hdr->hdr_signature != htonl(CACHE_SIGNATURE))\n \t\treturn error(\"bad signature\");\n-\tif (hdr->hdr_version != htonl(2) &&\n-\t    hdr->hdr_version != htonl(3) &&\n-\t    hdr->hdr_version != htonl(4))\n+\tif (hdr->hdr_version == htonl(2) ||\n+\t    hdr->hdr_version == htonl(3))\n+\t\tdo_crc = 0;\n+\telse if (hdr->hdr_version == htonl(4)) {\n+\t\tstruct ext_cache_header *ehdr = (struct ext_cache_header *)hdr;\n+\t\tdo_crc = ntohl(ehdr->hdr_flags) & CACHE_F_CRC;\n+\t}\n+\telse\n \t\treturn error(\"bad index version\");\n-\tgit_SHA1_Init(&c);\n-\tgit_SHA1_Update(&c, hdr, size - 20);\n-\tgit_SHA1_Final(sha1, &c);\n-\tif (hashcmp(sha1, (unsigned char *)hdr + size - 20))\n-\t\treturn error(\"bad index file sha1 signature\");\n+\tif (do_crc) {\n+\t\tuint32_t crc = crc32(0, NULL, 0);\n+\t\tcrc = crc32(crc,(void *) hdr, size - sizeof(uint32_t));\n+\t\tif (crc != *(uint32_t*)((unsigned char *)hdr + size - sizeof(uint32_t)))\n+\t\t\treturn error(\"bad index file crc32 signature\");\n+\t} else {\n+\t\tgit_SHA1_Init(&c);\n+\t\tgit_SHA1_Update(&c, hdr, size - 20);\n+\t\tgit_SHA1_Final(sha1, &c);\n+\t\tif (hashcmp(sha1, (unsigned char *)hdr + size - 20))\n+\t\t\treturn error(\"bad index file sha1 signature\");\n+\t}\n \treturn 0;\n }\n \n@@ -1421,11 +1434,24 @@ static int write_index_ext_header(struct sha1file *f,\n static int ce_flush(struct sha1file *f)\n {\n \tunsigned char sha1[20];\n-\tint fd = sha1close(f, sha1, 0);\n+\tint fd;\n \n-\tif (fd < 0)\n-\t\treturn -1;\n-\treturn (write_in_full(fd, sha1, 20) != 20) ? -1 : 0;\n+\tif (f->do_crc) {\n+\t\tuint32_t crc;\n+\n+\t\tassert(f->do_sha1 == 0);\n+\t\tsha1flush(f);\n+\t\tcrc = crc32_end(f);\n+\t\tfd = sha1close(f, sha1, 0);\n+\t\tif (fd < 0)\n+\t\t\treturn -1;\n+\t\treturn (write_in_full(fd, &crc, sizeof(crc)) != sizeof(crc)) ? -1 : 0;\n+\t} else {\n+\t\tfd = sha1close(f, sha1, 0);\n+\t\tif (fd < 0)\n+\t\t\treturn -1;\n+\t\treturn (write_in_full(fd, sha1, 20) != 20) ? -1 : 0;\n+\t}\n }\n \n static void ce_smudge_racily_clean_entry(struct cache_entry *ce)\n@@ -1568,6 +1594,10 @@ int write_index(struct index_state *istate, int newfd)\n \thdr.h.hdr_entries = htonl(entries - removed);\n \n \tf = sha1fd(newfd, NULL);\n+\tif (istate->hdr_flags & CACHE_F_CRC) {\n+\t\tcrc32_begin(f);\n+\t\tf->do_sha1 = 0;\n+\t}\n \tif (ce_write(f, &hdr, hdr_size) < 0)\n \t\treturn -1;\n \n-- \n1.7.8.36.g69ee2\n"},{"id":"183976","messageId":"1328507319-24687-6-git-send-email-pclouds@gmail.com","threadId":"29558","inReplyTo":"1328507319-24687-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 6/6] Automatically switch to crc32 checksum for index when it's too large","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-06T05:48:39Z","receivedAt":"2012-02-06T05:48:39Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"An experiment with -O3 is done on Intel D510@1.66GHz. At around 250k\nentries, index reading time exceeds 0.5s. Switching to crc32 brings it\nback lower than 0.2s.\n\nOn 4M files index, reading time with SHA-1 takes ~8.4, crc32 2.8s.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n I know no real repositories this size though. gentoo-x86 is \"only\"\n 120k. Haven't checked libreoffice repo yet.\n\n On 2M files index, allocating one big block (i.e. reverting debed2a\n (read-cache.c: allocate index entries individually - 2011-10-24)\n saves about 0.3s. Maybe we can allocate one big block, then malloc\n separately when the block is fully used.\n\n Writing time is still high. \"git update-index --crc32\" on crc32 250k index\n takes 0.9s (so writing time is about 0.5s)\n\n A better solution may be narrow clone (or just the narrow checkout\n part), where index only contains entries from checked out\n subdirectories.\n\n Documentation/config.txt |    7 +++++++\n builtin/update-index.c   |    1 +\n cache.h                  |    1 +\n config.c                 |    5 +++++\n environment.c            |    1 +\n read-cache.c             |    8 ++++++++\n 6 files changed, 23 insertions(+), 0 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex abeb82b..55b7596 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -540,6 +540,13 @@ relatively high IO latencies.  With this set to 'true', git will do the\n index comparison to the filesystem data in parallel, allowing\n overlapping IO's.\n \n+core.crc32IndexThreshold::\n+\tUsually SHA-1 is used to check for index integerity. When the\n+\tnumber of entries in index exceeds this threshold, crc32 will\n+\tbe used instead. Zero means SHA-1 always be used. Negative\n+\tvalue disables this threshold (i.e. crc32 or SHA-1 is decided\n+\tby other means).\n+\n core.createObject::\n \tYou can set this to 'link', in which case a hardlink followed by\n \ta delete of the source are used to make sure that object creation\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 6913226..5cb51c7 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -856,6 +856,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \targc = parse_options_end(&ctx);\n \n \tif (do_crc != -1) {\n+\t\tcore_crc32_index_threshold = -1;\n \t\tif (do_crc)\n \t\t\tthe_index.hdr_flags |= CACHE_F_CRC;\n \t\telse\ndiff --git a/cache.h b/cache.h\nindex 7352402..d05856b 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -610,6 +610,7 @@ extern unsigned long pack_size_limit_cfg;\n extern int read_replace_refs;\n extern int fsync_object_files;\n extern int core_preload_index;\n+extern int core_crc32_index_threshold;\n extern int core_apply_sparse_checkout;\n \n enum branch_track {\ndiff --git a/config.c b/config.c\nindex 40f9c6d..905e071 100644\n--- a/config.c\n+++ b/config.c\n@@ -671,6 +671,11 @@ static int git_default_core_config(const char *var, const char *value)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.crc32indexthreshold\")) {\n+\t\tcore_crc32_index_threshold = git_config_int(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.createobject\")) {\n \t\tif (!strcmp(value, \"rename\"))\n \t\t\tobject_creation_mode = OBJECT_CREATION_USES_RENAMES;\ndiff --git a/environment.c b/environment.c\nindex c93b8f4..9d9dfc2 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -66,6 +66,7 @@ unsigned long pack_size_limit_cfg;\n \n /* Parallel index stat data preload? */\n int core_preload_index = 0;\n+int core_crc32_index_threshold = 250000;\n \n /* This is set by setup_git_dir_gently() and/or git_default_config() */\n char *git_work_tree_cfg;\ndiff --git a/read-cache.c b/read-cache.c\nindex a34878e..fd032d8 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1582,6 +1582,14 @@ int write_index(struct index_state *istate, int newfd)\n \t\t}\n \t}\n \n+\tif (core_crc32_index_threshold >= 0) {\n+\t\tif (core_crc32_index_threshold > 0 &&\n+\t\t    istate->cache_nr >= core_crc32_index_threshold)\n+\t\t\tistate->hdr_flags |= CACHE_F_CRC;\n+\t\telse\n+\t\t\tistate->hdr_flags &= ~CACHE_F_CRC;\n+\t}\n+\n \thdr.h.hdr_signature = htonl(CACHE_SIGNATURE);\n \tif (istate->hdr_flags) {\n \t\thdr.h.hdr_version = htonl(4);\n-- \n1.7.8.36.g69ee2\n"},{"id":"183989","messageId":"7v4nv4a131.fsf@alter.siamese.dyndns.org","threadId":"29558","inReplyTo":"1328507319-24687-3-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 3/6] Stop producing index version 2","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-02-06T07:10:58Z","receivedAt":"2012-02-06T07:10:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> read-cache.c learned to produce version 2 or 3 depending on whether\n> extended cache entries exist in 06aaaa0 (Extend index to save more flags\n> - 2008-10-01), first released in 1.6.1. The purpose is to keep\n> compatibility with older git. It's been more than three years since\n> then and git has reached 1.7.9. Drop support for older git.\n\nCc'ing this, as I suspect this would surely raise eyebrows of some people\nwho wanted to get rid of the version 3 format.\n"},{"id":"183992","messageId":"7vsjio8leo.fsf@alter.siamese.dyndns.org","threadId":"29558","inReplyTo":"1328507319-24687-1-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 1/6] read-cache: use sha1file for sha1 calculation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-02-06T07:34:55Z","receivedAt":"2012-02-06T07:34:55Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n\nHaving no explanation on any of the patch in the series without cover\nletter makes it hard to comment on anything, and not having any numbers\nmakes it even harder after guessing that this is about some performance\ntweaks for 2M entry index cases.\n\nThis is open source, and I wouldn't stop you from spending time on anything\nthat interests you.\n\nBut having said that, if you have extra Git time, I would still rather see\nyou spend it first on tying up loose ends of your topics in flight and\non helping others that touch parts that are related to areas that you have\nalready thought about, namely:\n\n (1) nd/commit-ignore-i-t-a, which I think should be marketted as fixing\n     an earlier UI mistake and presented with a clean migration path to\n     make the updated behaviour the default in the future; and\n\n (2) the negative pathspec thing that resurfaced in disguise as Albert\n     Yale's \"grep --exclude\" series.\n\nthan playing with the approach of this series.  The two reasons I suspect\nthat spending your time on this series will give us much less value than\nthe above two topics out of you are:\n\n (1) While I think 2M-entry index is an interesting issue, it does not\n     affect most of the people; and more importantly\n\n (2) I think the proper way to handle 2M-entry index case is to avoid\n     having to write and read the whole 2M-entry as a flat table in the\n     first place, not by weakening how its integrity is assured in order\n     to micro-tweak the read/write efficiency without re-examining the\n     flatness of the current in-core index [*1*].\n\nThe first patch that reuses the existing csum-file API to older code that\nwas written before csum-file was invented is probably a good thing to do,\nthough, independent of the 2M-entry issue.\n\nThanks.\n\n\n[Footnote]\n\n*1* A possible approach might be to stuff unmodified trees in the index\nwithout exploding them into its components, and as entries are modified,\nlazily expand these \"tree\" entries, while ensuring the \"unmodified\" parts\nremain unmodified by turning the files in the working tree read-only and\nrequiring the user to say \"git edit\" or \"git open\" or something before\nstarting to edit.  But as I said, I consider this not an ultra-urgent\nissue, so I haven't thought things through yet.\n"},{"id":"184003","messageId":"CACsJy8DR2rPtC_2PDp=ZEm-3B-mzh+AxaDmXxxzJ_VY5M-0oVw@mail.gmail.com","threadId":"29558","inReplyTo":"7vsjio8leo.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH 1/6] read-cache: use sha1file for sha1 calculation","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-06T08:36:27Z","receivedAt":"2012-02-06T08:36:27Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"2012/2/6 Junio C Hamano <gitster@pobox.com>:\n> This is open source, and I wouldn't stop you from spending time on anything\n> that interests you.\n\nWell fun stuff is more interesting to do.. but I guess that's about\nit. More time cut down requires bigger changes that does not fit in a\nfew hours of work.\n\n> But having said that, if you have extra Git time, I would still rather see\n> you spend it first on tying up loose ends of your topics in flight and\n> on helping others that touch parts that are related to areas that you have\n> already thought about, namely:\n>\n>  (1) nd/commit-ignore-i-t-a, which I think should be marketted as fixing\n>     an earlier UI mistake and presented with a clean migration path to\n>     make the updated behaviour the default in the future; and\n\nYeah, I was avoiding the deprecation procedure (plus providing a\nconvincing argument to push it forward). Need to look up old emails..\n\n>  (2) the negative pathspec thing that resurfaced in disguise as Albert\n>     Yale's \"grep --exclude\" series.\n\nThis is pure headache. Can't avoid it forever, I guess.\n\n> *1* A possible approach might be to stuff unmodified trees in the index\n> without exploding them into its components, and as entries are modified,\n> lazily expand these \"tree\" entries, while ensuring the \"unmodified\" parts\n> remain unmodified by turning the files in the working tree read-only and\n> requiring the user to say \"git edit\" or \"git open\" or something before\n> starting to edit.  But as I said, I consider this not an ultra-urgent\n> issue, so I haven't thought things through yet.\n\nA sparse index is something that may be achieved with narrow clone (or\nnarrow checkout in full clone) because by nature we can't have full\nindex in narrow clone. That may be the right way to go.\n-- \nDuy\n"},{"id":"184010","messageId":"E799595D-61B3-4978-BCE1-BA6A33034B55@apple.com","threadId":"29558","inReplyTo":"1328507319-24687-6-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 6/6] Automatically switch to crc32 checksum for index when it's too large","fromName":"Dave Zarzycki","fromEmail":"zarzycki@apple.com","sentAt":"2012-02-06T08:50:27Z","receivedAt":"2012-02-06T08:50:27Z","isPatch":true,"sender":{"key":"zarzycki@apple.com","avatar":"https://avatars.githubusercontent.com/u/1071982?v=4"},"body":"Which crc32 polynomial is being used? crc32c (a.k.a. Castagnoli)? It would be great if this were the same polynomial that Intel implements in hardware via SSE4.2.\n\n\nOn Feb 5, 2012, at 9:48 PM, Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n\n> An experiment with -O3 is done on Intel D510@1.66GHz. At around 250k\n> entries, index reading time exceeds 0.5s. Switching to crc32 brings it\n> back lower than 0.2s.\n> \n> On 4M files index, reading time with SHA-1 takes ~8.4, crc32 2.8s.\n"},{"id":"184006","messageId":"CACsJy8Dv9fUzL3COZKVw_KR6aF20kHaw8M4CdBXJDA9H3fbxLw@mail.gmail.com","threadId":"29558","inReplyTo":"E799595D-61B3-4978-BCE1-BA6A33034B55@apple.com","subject":"Re: [PATCH 6/6] Automatically switch to crc32 checksum for index when it's too large","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-06T08:54:37Z","receivedAt":"2012-02-06T08:54:37Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"2012/2/6 Dave Zarzycki <zarzycki@apple.com>:\n> Which crc32 polynomial is being used? crc32c (a.k.a. Castagnoli)? It would be great if this were the same polynomial that Intel implements in hardware via SSE4.2.\n\nIt's zlib's crc32.\n\n> On Feb 5, 2012, at 9:48 PM, Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n>\n>> An experiment with -O3 is done on Intel D510@1.66GHz. At around 250k\n>> entries, index reading time exceeds 0.5s. Switching to crc32 brings it\n>> back lower than 0.2s.\n>>\n>> On 4M files index, reading time with SHA-1 takes ~8.4, crc32 2.8s.\n-- \nDuy\n"},{"id":"184009","messageId":"8A80E303-311E-4B87-A4B4-4F39AA932956@apple.com","threadId":"29558","inReplyTo":"CACsJy8Dv9fUzL3COZKVw_KR6aF20kHaw8M4CdBXJDA9H3fbxLw@mail.gmail.com","subject":"Re: [PATCH 6/6] Automatically switch to crc32 checksum for index when it's too large","fromName":"Dave Zarzycki","fromEmail":"zarzycki@apple.com","sentAt":"2012-02-06T09:07:26Z","receivedAt":"2012-02-06T09:07:26Z","isPatch":true,"sender":{"key":"zarzycki@apple.com","avatar":"https://avatars.githubusercontent.com/u/1071982?v=4"},"body":"\nOn Feb 6, 2012, at 12:54 AM, Nguyen Thai Ngoc Duy <pclouds@gmail.com> wrote:\n\n> 2012/2/6 Dave Zarzycki <zarzycki@apple.com>:\n>> Which crc32 polynomial is being used? crc32c (a.k.a. Castagnoli)? It would be great if this were the same polynomial that Intel implements in hardware via SSE4.2.\n> \n> It's zlib's crc32.\n\nThat's too bad. Zlib uses crc32, not crc32c. The Intel instruction is crc32c and is 2-3 times faster than the best software based implementation.\n\nhttp://www.strchr.com/crc32_popcnt\n\n\n\n> \n>> On Feb 5, 2012, at 9:48 PM, Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n>> \n>>> An experiment with -O3 is done on Intel D510@1.66GHz. At around 250k\n>>> entries, index reading time exceeds 0.5s. Switching to crc32 brings it\n>>> back lower than 0.2s.\n>>> \n>>> On 4M files index, reading time with SHA-1 takes ~8.4, crc32 2.8s.\n> -- \n> Duy\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"184087","messageId":"CAJo=hJvtRnmvALcn3vKpYTr3j6ada8iboPjWN3cQnwwKzRvrDA@mail.gmail.com","threadId":"29558","inReplyTo":"7v4nv4a131.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH 3/6] Stop producing index version 2","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2012-02-07T03:09:15Z","receivedAt":"2012-02-07T03:09:15Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"2012/2/5 Junio C Hamano <gitster@pobox.com>:\n> Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n>\n>> read-cache.c learned to produce version 2 or 3 depending on whether\n>> extended cache entries exist in 06aaaa0 (Extend index to save more flags\n>> - 2008-10-01), first released in 1.6.1. The purpose is to keep\n>> compatibility with older git. It's been more than three years since\n>> then and git has reached 1.7.9. Drop support for older git.\n>\n> Cc'ing this, as I suspect this would surely raise eyebrows of some people\n> who wanted to get rid of the version 3 format.\n\nVersion 3 was a mistake because of the variable length record sizes.\nSaving 2 bytes on some records that don't use the extended flags makes\nthe index file *MUCH* harder to parse. So much so that we should take\nversion 3 and kill it, not encourage it as the default!\n\nIMHO, when these extended flags were added to make version 3 the\nfollowing should have happened:\n\n- All records use the larger structure format with 4 bytes for the\nflags, not 2 bytes.\n\n- Change the trailing padding after the name to be a *SINGLE* \\0 byte,\nand do not pad out to an 8 byte boundary.\n\nBoth make it really hard to process the file, and the latter happens\nonly for direct mmap usage, which we don't do anymore.\n\n\nWe also have to consider the EGit and JGit user base as part of the\necosystem. We can't just kill a file format because git-core has been\ncapable of reading its alternative since some arbitrary YYYY-MM-DD\nrelease date. We need to also consider when did some other major tools\ncatch up and also support this format?\n\nFWIW JGit released index version 3 support in version 0.9.1, which\nshipped Sep 15, 2010. JGit/EGit were more than 2 years behind here.\n\n\n<thinking type=\"wishful\" probability=\"never-happen\"\nprobably-inflating-flame-from=\"linus\">\n\nI have long wanted to scrap the current index format. I unfortunately\ndon't have the time to do it myself. But I suspect there may be a lot\nof gains by making the index format match the canonical tree format\nbetter by keeping the tree structure within a single file stream,\nnesting entries below their parent directory, and keeping tree SHA-1\ndata along with the directory entry. For one thing the index would be\nable to register an empty subdirectory, rather than ignoring them. It\nwould also better line up with the filesystem's readdir() handling,\ngiving us more sane logic to compare what readdir() tells us exists\nagainst what the index thinks should be in the same file. And the\noverall index should be smaller, because we don't have to repeat the\nsame path/to/a/file/for/every/file/in/that/same/directory/tree.\nReconstructing the path strings at read time into a flat list should\nbe pretty trivial, and still keep the parallel lstat calls running off\na flat list working well for fast status operations.\n\n</thinking>\n"},{"id":"184088","messageId":"CAJo=hJvSyhv8EUh=6ROotc3Q=zQo7vbww_ShQJP3tf1T7s889g@mail.gmail.com","threadId":"29558","inReplyTo":"1328507319-24687-5-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 5/6] Allow to use crc32 as a lighter checksum on index","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2012-02-07T03:17:04Z","receivedAt":"2012-02-07T03:17:04Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"2012/2/5 Nguyễn Thái Ngọc Duy <pclouds@gmail.com>:\n>        if (hdr->hdr_signature != htonl(CACHE_SIGNATURE))\n>                return error(\"bad signature\");\n> -       if (hdr->hdr_version != htonl(2) &&\n> -           hdr->hdr_version != htonl(3) &&\n> -           hdr->hdr_version != htonl(4))\n> +       if (hdr->hdr_version == htonl(2) ||\n> +           hdr->hdr_version == htonl(3))\n> +               do_crc = 0;\n> +       else if (hdr->hdr_version == htonl(4)) {\n> +               struct ext_cache_header *ehdr = (struct ext_cache_header *)hdr;\n> +               do_crc = ntohl(ehdr->hdr_flags) & CACHE_F_CRC;\n> +       }\n> +       else\n\nIck. Ick. Ick. Please $DEITY no.\n\nWhen it comes to data integrity codes in Git... PICK ONE AND STICK WITH IT.\n\nIf CRC-32 is good enough to protect the index content such that disk\ncorruption is probably detectable with it, lets just switch to CRC-32\nin index version 4. Don't make it optional with a new header field\nthat wasn't there in version 3 and is now only able to accept 32 bits\nof flags before we have to go and create index version 5. We already\nhave a cache extension system available with extension codes in the\nfooter of the index file. We don't need YET ANOTHER EXTENSION SYSTEM.\n\nIf CRC-32 is not good enough, and we don't want to trust it (or\nreally, YOU don't want to trust it) please do not then go and propose\nthat a less knowledgeable user should switch to CRC-32 \"because it is\nfaster\". If we don't want to rely on the error detection of CRC-32,\nthen we should be using SHA-1. Or SHA-256.\n\n\nI haven't really put a lot of thought into this. But I suspect CRC-32\nis sufficient on the index file, until it gets so big that the\nprobability of a bit flip going undetected is too high due to the size\nof the file, but then we are into the \"huge\" index size range that has\nyou trying to swap out SHA-1 for CRC-32 because SHA-1 is too slow. Uhm\nno.\n\nCRC-32 may be good enough, we use it inside of the pack-objects when\ndoing repacking locally and don't want to inflate objects to check\nSHA-1, but do want to try and detect a random bit flip caused by a\nbroken file copier. Thus far its held up well there. Given the very\ntransient nature of the index file (and how it can be mostly rebuilt\nfrom a tree object and the working directory), CRC-32 might be good\nenough. But please pick one.\n"},{"id":"184089","messageId":"159A2D07-0B02-4E85-B7AA-C668FDA9F382@apple.com","threadId":"29558","inReplyTo":"CAJo=hJvSyhv8EUh=6ROotc3Q=zQo7vbww_ShQJP3tf1T7s889g@mail.gmail.com","subject":"Re: [PATCH 5/6] Allow to use crc32 as a lighter checksum on index","fromName":"Dave Zarzycki","fromEmail":"zarzycki@apple.com","sentAt":"2012-02-07T04:04:59Z","receivedAt":"2012-02-07T04:04:59Z","isPatch":true,"sender":{"key":"zarzycki@apple.com","avatar":"https://avatars.githubusercontent.com/u/1071982?v=4"},"body":"On Feb 6, 2012, at 7:17 PM, Shawn Pearce <spearce@spearce.org> wrote:\n\n> I haven't really put a lot of thought into this. But I suspect CRC-32\n> is sufficient on the index file, until it gets so big that the\n> probability of a bit flip going undetected is too high due to the size\n> of the file, but then we are into the \"huge\" index size range that has\n> you trying to swap out SHA-1 for CRC-32 because SHA-1 is too slow. Uhm\n> no.\n\nCRCs are designed to be implemented in hardware and provide basic single-bit error checking for networking packets of disk blocks. With a good polynomial, they're reasonably effective at detecting a single-bit error within 8 or 16 kilobytes:\n\n    http://www.ece.cmu.edu/~koopman/networks/dsn02/dsn02_koopman.pdf\n"},{"id":"184091","messageId":"947C7419-52B3-46B3-A849-C4498DF9A099@apple.com","threadId":"29558","inReplyTo":"159A2D07-0B02-4E85-B7AA-C668FDA9F382@apple.com","subject":"Re: [PATCH 5/6] Allow to use crc32 as a lighter checksum on index","fromName":"Dave Zarzycki","fromEmail":"zarzycki@apple.com","sentAt":"2012-02-07T04:29:53Z","receivedAt":"2012-02-07T04:29:53Z","isPatch":true,"sender":{"key":"zarzycki@apple.com","avatar":"https://avatars.githubusercontent.com/u/1071982?v=4"},"body":"\nOn Feb 6, 2012, at 8:04 PM, Dave Zarzycki <zarzycki@apple.com> wrote:\n\n> On Feb 6, 2012, at 7:17 PM, Shawn Pearce <spearce@spearce.org> wrote:\n> \n>> I haven't really put a lot of thought into this. But I suspect CRC-32\n>> is sufficient on the index file, until it gets so big that the\n>> probability of a bit flip going undetected is too high due to the size\n>> of the file, but then we are into the \"huge\" index size range that has\n>> you trying to swap out SHA-1 for CRC-32 because SHA-1 is too slow. Uhm\n>> no.\n> \n> CRCs are designed to be implemented in hardware and provide basic single-bit error checking for networking packets of disk blocks. With a good polynomial, they're reasonably effective at detecting a single-bit error within 8 or 16 kilobytes:\n> \n>    http://www.ece.cmu.edu/~koopman/networks/dsn02/dsn02_koopman.pdf\n\ns/packets of disk blocks/packets or disk blocks/g\n"},{"id":"184093","messageId":"CACsJy8Ayqea75xeFKJNm6iT7GSUGDEfvZD17uEv7ihr4SS2LMg@mail.gmail.com","threadId":"29558","inReplyTo":"CAJo=hJvtRnmvALcn3vKpYTr3j6ada8iboPjWN3cQnwwKzRvrDA@mail.gmail.com","subject":"Re: [PATCH 3/6] Stop producing index version 2","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-07T04:50:54Z","receivedAt":"2012-02-07T04:50:54Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Feb 7, 2012 at 10:09 AM, Shawn Pearce <spearce@spearce.org> wrote:\n> 2012/2/5 Junio C Hamano <gitster@pobox.com>:\n>> Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n>>\n>>> read-cache.c learned to produce version 2 or 3 depending on whether\n>>> extended cache entries exist in 06aaaa0 (Extend index to save more flags\n>>> - 2008-10-01), first released in 1.6.1. The purpose is to keep\n>>> compatibility with older git. It's been more than three years since\n>>> then and git has reached 1.7.9. Drop support for older git.\n>>\n>> Cc'ing this, as I suspect this would surely raise eyebrows of some people\n>> who wanted to get rid of the version 3 format.\n>\n> Version 3 was a mistake because of the variable length record sizes.\n> Saving 2 bytes on some records that don't use the extended flags makes\n> the index file *MUCH* harder to parse. So much so that we should take\n> version 3 and kill it, not encourage it as the default!\n\nProbably too late for that, but it's good to know there are strong\nuser base for v2.\n\n> <thinking type=\"wishful\" probability=\"never-happen\"\n> probably-inflating-flame-from=\"linus\">\n>\n> I have long wanted to scrap the current index format. I unfortunately\n> don't have the time to do it myself. But I suspect there may be a lot\n> of gains by making the index format match the canonical tree format\n> better by keeping the tree structure within a single file stream,\n> nesting entries below their parent directory, and keeping tree SHA-1\n> data along with the directory entry. For one thing the index would be\n> able to register an empty subdirectory, rather than ignoring them. It\n> would also better line up with the filesystem's readdir() handling,\n> giving us more sane logic to compare what readdir() tells us exists\n> against what the index thinks should be in the same file. And the\n> overall index should be smaller, because we don't have to repeat the\n> same path/to/a/file/for/every/file/in/that/same/directory/tree.\n> Reconstructing the path strings at read time into a flat list should\n> be pretty trivial, and still keep the parallel lstat calls running off\n> a flat list working well for fast status operations.\n>\n> </thinking>\n\nHaven't really thought through, but I suppose we could create extended\ntree object format (there is info in cache entry that's not in tree\nentry), store index in this format, then pack together and store the\npack as part of index file. Append-only access to index would be\npossible by appending a new pack of new trees to index) I think with\ntree-based index, we could kill a big chunk of code (merging trees and\nindex together) in unpack_trees(). With further efforts to remove\nlist-based index usage, we could even kill match_pathspec_depth(),\nmaking tree_entry_interesting() the only function to match patchspec.\nBut dreams probably never come true.\n-- \nDuy\n"},{"id":"184096","messageId":"7vehu743rp.fsf@alter.siamese.dyndns.org","threadId":"29558","inReplyTo":"CAJo=hJvtRnmvALcn3vKpYTr3j6ada8iboPjWN3cQnwwKzRvrDA@mail.gmail.com","subject":"Re: [PATCH 3/6] Stop producing index version 2","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-02-07T05:21:46Z","receivedAt":"2012-02-07T05:21:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n> <thinking type=\"wishful\" probability=\"never-happen\"\n> probably-inflating-flame-from=\"linus\">\n>\n> I have long wanted to scrap the current index format. I unfortunately\n> don't have the time to do it myself. But I suspect there may be a lot\n> of gains by making the index format match the canonical tree format\n> better by keeping the tree structure within a single file stream,\n> nesting entries below their parent directory, and keeping tree SHA-1\n> data along with the directory entry.\n\nI suspect that is not so \"never-happen wishful thinking\".\n\nIn an earlier message, I alluded to a data structure that starts with a\nsingle top-level tree entry that is lazily expanded as the index entries\nare updated. The above shows that at least two of us share the same (day)\ndream, and I suspect there are others that share the same \"gut feeling\"\nthat such a tree-based structure would be the way to do large index right.\n\nIt would be a large and possibly painful change, but the good thing is\nthat the index is a local matter and we won't have to worry too much about\na flag day event.\n\n</thinking>\n"},{"id":"184110","messageId":"CACsJy8DBenQGrF4-X-rsUmJQUhx6MMQg+8Yrspjuxx6W6L48AQ@mail.gmail.com","threadId":"29558","inReplyTo":"CACsJy8Ayqea75xeFKJNm6iT7GSUGDEfvZD17uEv7ihr4SS2LMg@mail.gmail.com","subject":"Re: [PATCH 3/6] Stop producing index version 2","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-07T08:51:54Z","receivedAt":"2012-02-07T08:51:54Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Feb 7, 2012 at 11:50 AM, Nguyen Thai Ngoc Duy <pclouds@gmail.com> wrote:\n>> Version 3 was a mistake because of the variable length record sizes.\n>> Saving 2 bytes on some records that don't use the extended flags makes\n>> the index file *MUCH* harder to parse. So much so that we should take\n>> version 3 and kill it, not encourage it as the default!\n>\n> Probably too late for that, but it's good to know there are strong\n> user base for v2.\n\nOK probably not too late. We cannot kill it, but we could deprecate\nit. We can introduce a mandatory extension to store extra flags. The\nextension is basically an array of\n\nstruct ce_extended_flags {\n    int ce_index; /* points to istate->cache[ce_index] */\n    unsigned long flags;\n};\n\nOn reading the extension, extra flags is applied back in mem, the\nextension is created again when new index is written. There are only\ntwo users of index v3: skip-worktree and intent-to-add bits, which are\nnot used often, I think. Still want to kill it?\n\nSwitching from sha-1 to crc32 could be done the same way (i.e. new\nmandatory extension _at the end_ that contains crc32 checksum and skip\nsha-1 check on reading if it's all zero) if we agree to move to crc32.\n-- \nDuy\n"},{"id":"184141","messageId":"874nv2o8rs.fsf@thomas.inf.ethz.ch","threadId":"29558","inReplyTo":"CAJo=hJvtRnmvALcn3vKpYTr3j6ada8iboPjWN3cQnwwKzRvrDA@mail.gmail.com","subject":"Re: [PATCH 3/6] Stop producing index version 2","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2012-02-07T17:25:43Z","receivedAt":"2012-02-07T17:25:43Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n> I have long wanted to scrap the current index format. I unfortunately\n> don't have the time to do it myself. But I suspect there may be a lot\n> of gains by making the index format match the canonical tree format\n> better by keeping the tree structure within a single file stream,\n> nesting entries below their parent directory, and keeping tree SHA-1\n> data along with the directory entry.\n\nIf I may add to this: the one thing that I would like to see fixed about\nthe index is that it's flat out impossible to change a single thing in\nit without re\"writing\" it from scratch.\n\nI'm saying \"writing\" because it is possible to change a few things\naround, but recomputing the trailing SHA1 swamps that by a large margin\nunless you are writing to a floppy disk, so it doesn't matter.  I'm sure\nusing a CRC32 helps here, but if we're going to make an incompatible\nchange, why not go all the way?\n\nA tree layout can fix that if it is properly arranged so that if you\n'git add path/to/file', it only updates the SHA1s for path/to/file,\npath/to and path.  For this to work, the checks would have to correspond\nto the trees, perhaps even directly use the actual tree SHA1.  This\nwould at least be natural in some sense; getting to actual log(n)\ncomplexity for hilariously large directories would require dynamically\nsplitting directories where appropriate.\n\nAlong the same lines the format should allow for changing the extension\ndata for a single extension while only rehashing the new data.\n\n\nWhen I worked on cache-tree, I considered making a change to the latter\neffect, but thought the impact too great for a little gain.  Now from\nthis thread, I'm getting the impression that such a change would be ok,\neven if users would have to scrap the index if they downgrade.  Is that\nright?\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"}]}