{"thread":{"id":"35410","subject":"[PATCH v4 00/24] Index-v5","startedAt":"2013-11-27T12:00:35Z","lastAt":"2013-12-09T10:14:43Z","messageCount":41,"participants":["Thomas Gummerer","Eric Sunshine","Junio C Hamano","Duy Nguyen","Antoine Pelisse"],"isPatch":true,"patchVersion":4,"patchTotal":24},"messages":[{"id":"231164","messageId":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":null,"subject":"[PATCH v4 00/24] Index-v5","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:35Z","receivedAt":"2013-11-27T12:00:35Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Hi,\n\nprevious rounds (without api) are at $gmane/202752, $gmane/202923,\n$gmane/203088 and $gmane/203517, the previous rounds with api were at\n$gmane/229732, $gmane/230210 and $gmane/232488.  Thanks to Duy for\nreviewing the the last round and Junio, Ramsay and Eric for additional\ncomments.\n\nSince the last round I've added a POC for partial writing, resulting\nin the following performance improvements for update-index:\n\nTest                                        1063432           HEAD\n------------------------------------------------------------------------------------\n0003.2: v[23]: update-index                 0.60(0.38+0.20)   0.76(0.36+0.17) +26.7%\n0003.3: v[23]: grep nonexistent -- subdir   0.28(0.17+0.11)   0.28(0.18+0.09) +0.0%\n0003.4: v[23]: ls-files -- subdir           0.26(0.15+0.10)   0.24(0.14+0.09) -7.7%\n0003.7: v[23] update-index                  0.59(0.36+0.22)   0.58(0.36+0.20) -1.7%\n0003.9: v4: update-index                    0.46(0.28+0.17)   0.45(0.30+0.11) -2.2%\n0003.10: v4: grep nonexistent -- subdir     0.26(0.14+0.11)   0.21(0.14+0.07) -19.2%\n0003.11: v4: ls-files -- subdir             0.24(0.14+0.10)   0.20(0.12+0.08) -16.7%\n0003.14: v4 update-index                    0.49(0.31+0.18)   0.65(0.34+0.17) +32.7%\n0003.16: v5: update-index                   0.53(0.30+0.22)   0.50(0.28+0.20) -5.7%\n0003.17: v5: ls-files                       0.27(0.15+0.12)   0.27(0.17+0.10) +0.0%\n0003.18: v5: grep nonexistent -- subdir     0.02(0.01+0.01)   0.03(0.01+0.01) +50.0%\n0003.19: v5: ls-files -- subdir             0.02(0.00+0.02)   0.02(0.01+0.01) +0.0%\n0003.22: v5 update-index                    0.53(0.29+0.23)   0.02(0.01+0.01) -96.2%\n\nGiven this, I don't think a complete change of the in-core format for\nthe cache-entries is necessary to take full advantage of the new index\nfile format.  Instead some changes to the current in-core format would\nwork well with the new on-disk format.\n\nThe current in-memory format fits the internal needs of git fairly well,\nso I don't think changing it to fit a better index file format would\nmake a lot of sense, given that we can take advantage of the new format\nwith the existing in-memory format.\n\nThis series doesn't use kb/fast-hashmap yet, but that should be fairly\nsimple to change if the series is deemed a good change.  The\nperformance tests for update-index test require\ntg/perf-lib-test-perf-cleanup. \n\nOther changes, made following the review comments are:\n\ndocumentation: add documentation of the index-v5 file format\n  - Update documentation that directory flags are now 32-bits.  That\n    makes aligned access simpler\n  - offset_to_offset is no longer included in the checksum for files.\n    It's unnecessary.\n\nread-cache: read index-v5\n  - Add fix for reading with different level pathspecs given\n  - Use init_directory_entry to initialize all fields in a new\n    directory entry\n  - use memset to simplify the create_new_conflict function\n  - Add comments to explain -5 when reading directories and files\n  - Add comments for the more complex functions\n  - Add name flex_array to the end of ondisk_directory_entry for\n    simplified reading\n  - Add name flex_array to the end of ondisk_cache_entry for\n    simplified reading\n  - Move conflict reading functions to next patch\n  - mark functions as static when they are\n\nread-cache: read resolve-undo data\n  - Add comments for the more complex function\n  - Read conflicts + resolve undo data as extension\n\nread-cache: read cache-tree in index-v5\n  - Add comments for the more complex function\n  - Instead of sorting the directory entries, sort the cache-tree\n    directly.  This also required changing the algorithms with which\n    the cache entries are extracted from the directory tree.\n\nread-cache: write index-v5\n  - Free pointers allocated by super_directory\n  - Rewrite condition as suggested by Duy\n  - Don't check for CE_REMOVE'd entries in the writing code, they are\n    already checked in the compile_directory_data code\n  - Remove overly complicated directory size calculation since flags\n    are now 32-bits\n\nread-cache: write resolve-undo data for index-v5\n  - Free pointers allocated by super_directory\n  - Write conflicts + resolve undo data as extension\n\nintroduce GIT_INDEX_VERSION environment variable\n  - Add documentation for GIT_INDEX_VERSION\n\ntest-lib: allow setting the index format version\n\nRemoved commits:\n  - read-cache: don't check uid, gid, ino\n  - read-cache: use fixed width integer types (independently in pu)\n  - read-cache: clear version in discard_index()\n\nTypos fixed as suggested by Eric Sunshine\n\nThomas Gummerer (22):\n  read-cache: split index file version specific functionality\n  read-cache: move index v2 specific functions to their own file\n  read-cache: Re-read index if index file changed\n  add documentation for the index api\n  read-cache: add index reading api\n  make sure partially read index is not changed\n  grep.c: use index api\n  ls-files.c: use index api\n  documentation: add documentation of the index-v5 file format\n  read-cache: make in-memory format aware of stat_crc\n  read-cache: read index-v5\n  read-cache: read resolve-undo data\n  read-cache: read cache-tree in index-v5\n  read-cache: write index-v5\n  read-cache: write index-v5 cache-tree data\n  read-cache: write resolve-undo data for index-v5\n  update-index.c: rewrite index when index-version is given\n  introduce GIT_INDEX_VERSION environment variable\n  test-lib: allow setting the index format version\n  t1600: add index v5 specific tests\n  POC for partial writing\n  perf: add partial writing test\n\nThomas Rast (1):\n  p0003-index.sh: add perf test for the index formats\n\n Documentation/git.txt                            |    5 +\n Documentation/technical/api-in-core-index.txt    |   56 +-\n Documentation/technical/index-file-format-v5.txt |  294 +++++\n Makefile                                         |   10 +\n builtin/apply.c                                  |    2 +\n builtin/grep.c                                   |   69 +-\n builtin/ls-files.c                               |   36 +-\n builtin/update-index.c                           |   50 +-\n cache-tree.c                                     |   15 +-\n cache-tree.h                                     |    2 +\n cache.h                                          |  115 +-\n lockfile.c                                       |    2 +-\n read-cache-v2.c                                  |  561 +++++++++\n read-cache-v5.c                                  | 1406 ++++++++++++++++++++++\n read-cache.c                                     |  691 +++--------\n read-cache.h                                     |   67 ++\n resolve-undo.c                                   |    1 +\n t/perf/p0003-index.sh                            |   74 ++\n t/t1600-index-v5.sh                              |   25 +\n t/t2101-update-index-reupdate.sh                 |   12 +-\n t/test-lib-functions.sh                          |    5 +\n t/test-lib.sh                                    |    3 +\n test-index-version.c                             |    6 +\n unpack-trees.c                                   |    3 +-\n 24 files changed, 2921 insertions(+), 589 deletions(-)\n create mode 100644 Documentation/technical/index-file-format-v5.txt\n create mode 100644 read-cache-v2.c\n create mode 100644 read-cache-v5.c\n create mode 100644 read-cache.h\n create mode 100755 t/perf/p0003-index.sh\n create mode 100755 t/t1600-index-v5.sh\n\n-- \n1.8.4.2\n"},{"id":"231165","messageId":"1385553659-9928-2-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 01/24] t2104: Don't fail for index versions other than [23]","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:36Z","receivedAt":"2013-11-27T12:00:36Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"t2104 currently checks for the exact index version 2 or 3,\ndepending if there is a skip-worktree flag or not. Other\nindex versions do not use extended flags and thus cannot\nbe tested for version changes.\n\nMake this test update the index to version 2 at the beginning\nof the test. Testing the skip-worktree flags for the default\nindex format is still covered by t7011 and t7012.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n t/t2104-update-index-skip-worktree.sh | 1 +\n 1 file changed, 1 insertion(+)\n\ndiff --git a/t/t2104-update-index-skip-worktree.sh b/t/t2104-update-index-skip-worktree.sh\nindex 1d0879b..bd9644f 100755\n--- a/t/t2104-update-index-skip-worktree.sh\n+++ b/t/t2104-update-index-skip-worktree.sh\n@@ -22,6 +22,7 @@ H sub/2\n EOF\n \n test_expect_success 'setup' '\n+\tgit update-index --index-version=2 &&\n \tmkdir sub &&\n \ttouch ./1 ./2 sub/1 sub/2 &&\n \tgit add 1 2 sub/1 sub/2 &&\n-- \n1.8.4.2\n"},{"id":"231166","messageId":"1385553659-9928-3-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 02/24] read-cache: split index file version specific functionality","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:37Z","receivedAt":"2013-11-27T12:00:37Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Split index file version specific functionality to their own functions,\nto prepare for moving the index file version specific parts to their own\nfile.  This makes it easier to add a new index file format later.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n read-cache.c | 114 ++++++++++++++++++++++++++++++++++++++---------------------\n 1 file changed, 74 insertions(+), 40 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 33dd676..5a8f405 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1269,10 +1269,8 @@ struct ondisk_cache_entry_extended {\n \t\t\t    ondisk_cache_entry_extended_size(ce_namelen(ce)) : \\\n \t\t\t    ondisk_cache_entry_size(ce_namelen(ce)))\n \n-static int verify_hdr(struct cache_header *hdr, unsigned long size)\n+static int verify_hdr_version(struct cache_header *hdr, unsigned long size)\n {\n-\tgit_SHA_CTX c;\n-\tunsigned char sha1[20];\n \tint hdr_version;\n \n \tif (hdr->hdr_signature != htonl(CACHE_SIGNATURE))\n@@ -1280,10 +1278,21 @@ static int verify_hdr(struct cache_header *hdr, unsigned long size)\n \thdr_version = ntohl(hdr->hdr_version);\n \tif (hdr_version < INDEX_FORMAT_LB || INDEX_FORMAT_UB < hdr_version)\n \t\treturn error(\"bad index version %d\", hdr_version);\n+\treturn 0;\n+}\n+\n+static int verify_hdr(void *mmap, unsigned long size)\n+{\n+\tgit_SHA_CTX c;\n+\tunsigned char sha1[20];\n+\n+\tif (size < sizeof(struct cache_header) + 20)\n+\t\tdie(\"index file smaller than expected\");\n+\n \tgit_SHA1_Init(&c);\n-\tgit_SHA1_Update(&c, hdr, size - 20);\n+\tgit_SHA1_Update(&c, mmap, size - 20);\n \tgit_SHA1_Final(sha1, &c);\n-\tif (hashcmp(sha1, (unsigned char *)hdr + size - 20))\n+\tif (hashcmp(sha1, (unsigned char *)mmap + size - 20))\n \t\treturn error(\"bad index file sha1 signature\");\n \treturn 0;\n }\n@@ -1425,44 +1434,14 @@ static struct cache_entry *create_from_disk(struct ondisk_cache_entry *ondisk,\n \treturn ce;\n }\n \n-/* remember to discard_cache() before reading a different cache! */\n-int read_index_from(struct index_state *istate, const char *path)\n+static int read_index_v2(struct index_state *istate, void *mmap, unsigned long mmap_size)\n {\n-\tint fd, i;\n-\tstruct stat st;\n+\tint i;\n \tunsigned long src_offset;\n \tstruct cache_header *hdr;\n-\tvoid *mmap;\n-\tsize_t mmap_size;\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \n-\tif (istate->initialized)\n-\t\treturn istate->cache_nr;\n-\n-\tistate->timestamp.sec = 0;\n-\tistate->timestamp.nsec = 0;\n-\tfd = open(path, O_RDONLY);\n-\tif (fd < 0) {\n-\t\tif (errno == ENOENT)\n-\t\t\treturn 0;\n-\t\tdie_errno(\"index file open failed\");\n-\t}\n-\n-\tif (fstat(fd, &st))\n-\t\tdie_errno(\"cannot stat the open index\");\n-\n-\tmmap_size = xsize_t(st.st_size);\n-\tif (mmap_size < sizeof(struct cache_header) + 20)\n-\t\tdie(\"index file smaller than expected\");\n-\n-\tmmap = xmmap(NULL, mmap_size, PROT_READ | PROT_WRITE, MAP_PRIVATE, fd, 0);\n-\tif (mmap == MAP_FAILED)\n-\t\tdie_errno(\"unable to map index file\");\n-\tclose(fd);\n-\n \thdr = mmap;\n-\tif (verify_hdr(hdr, mmap_size) < 0)\n-\t\tgoto unmap;\n \n \tistate->version = ntohl(hdr->hdr_version);\n \tistate->cache_nr = ntohl(hdr->hdr_entries);\n@@ -1488,8 +1467,6 @@ int read_index_from(struct index_state *istate, const char *path)\n \t\tsrc_offset += consumed;\n \t}\n \tstrbuf_release(&previous_name_buf);\n-\tistate->timestamp.sec = st.st_mtime;\n-\tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n \twhile (src_offset <= mmap_size - 20 - 8) {\n \t\t/* After an array of active_nr index entries,\n@@ -1509,6 +1486,58 @@ int read_index_from(struct index_state *istate, const char *path)\n \t\tsrc_offset += 8;\n \t\tsrc_offset += extsize;\n \t}\n+\treturn 0;\n+unmap:\n+\tmunmap(mmap, mmap_size);\n+\tdie(\"index file corrupt\");\n+}\n+\n+/* remember to discard_cache() before reading a different cache! */\n+int read_index_from(struct index_state *istate, const char *path)\n+{\n+\tint fd;\n+\tstruct stat st;\n+\tstruct cache_header *hdr;\n+\tvoid *mmap;\n+\tsize_t mmap_size;\n+\n+\terrno = EBUSY;\n+\tif (istate->initialized)\n+\t\treturn istate->cache_nr;\n+\n+\terrno = ENOENT;\n+\tistate->timestamp.sec = 0;\n+\tistate->timestamp.nsec = 0;\n+\tfd = open(path, O_RDONLY);\n+\tif (fd < 0) {\n+\t\tif (errno == ENOENT)\n+\t\t\treturn 0;\n+\t\tdie_errno(\"index file open failed\");\n+\t}\n+\n+\tif (fstat(fd, &st))\n+\t\tdie_errno(\"cannot stat the open index\");\n+\n+\terrno = EINVAL;\n+\tmmap_size = xsize_t(st.st_size);\n+\tif (mmap_size < sizeof(struct cache_header) + 20)\n+\t\tdie(\"index file smaller than expected\");\n+\n+\tmmap = xmmap(NULL, mmap_size, PROT_READ | PROT_WRITE, MAP_PRIVATE, fd, 0);\n+\tclose(fd);\n+\tif (mmap == MAP_FAILED)\n+\t\tdie_errno(\"unable to map index file\");\n+\n+\thdr = mmap;\n+\tif (verify_hdr_version(hdr, mmap_size) < 0)\n+\t\tgoto unmap;\n+\n+\tif (verify_hdr(mmap, mmap_size) < 0)\n+\t\tgoto unmap;\n+\n+\tread_index_v2(istate, mmap, mmap_size);\n+\tistate->timestamp.sec = st.st_mtime;\n+\tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \tmunmap(mmap, mmap_size);\n \treturn istate->cache_nr;\n \n@@ -1772,7 +1801,7 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n \t\trollback_lock_file(lockfile);\n }\n \n-int write_index(struct index_state *istate, int newfd)\n+static int write_index_v2(struct index_state *istate, int newfd)\n {\n \tgit_SHA_CTX c;\n \tstruct cache_header hdr;\n@@ -1864,6 +1893,11 @@ int write_index(struct index_state *istate, int newfd)\n \treturn 0;\n }\n \n+int write_index(struct index_state *istate, int newfd)\n+{\n+\treturn write_index_v2(istate, newfd);\n+}\n+\n /*\n  * Read the index file that is potentially unmerged into given\n  * index_state, dropping any unmerged entries.  Returns true if\n-- \n1.8.4.2\n"},{"id":"231169","messageId":"1385553659-9928-4-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 03/24] read-cache: move index v2 specific functions to their own file","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:38Z","receivedAt":"2013-11-27T12:00:38Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Move index version 2 specific functions to their own file. The non-index\nspecific functions will be in read-cache.c, while the index version 2\nspecific functions will be in read-cache-v2.c.\n\nHelped-by: Nguyen Thai Ngoc Duy <pclouds@gmail.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n Makefile               |   2 +\n builtin/apply.c        |   2 +\n builtin/update-index.c |   2 +-\n cache.h                |  13 +-\n read-cache-v2.c        | 553 ++++++++++++++++++++++++++++++++++++++++++++++\n read-cache.c           | 585 +++++--------------------------------------------\n read-cache.h           |  63 ++++++\n test-index-version.c   |   6 +\n unpack-trees.c         |   3 +-\n 9 files changed, 683 insertions(+), 546 deletions(-)\n create mode 100644 read-cache-v2.c\n create mode 100644 read-cache.h\n\ndiff --git a/Makefile b/Makefile\nindex af847f8..5c28777 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -705,6 +705,7 @@ LIB_H += progress.h\n LIB_H += prompt.h\n LIB_H += quote.h\n LIB_H += reachable.h\n+LIB_H += read-cache.h\n LIB_H += reflog-walk.h\n LIB_H += refs.h\n LIB_H += remote.h\n@@ -849,6 +850,7 @@ LIB_OBJS += prompt.o\n LIB_OBJS += quote.o\n LIB_OBJS += reachable.o\n LIB_OBJS += read-cache.o\n+LIB_OBJS += read-cache-v2.o\n LIB_OBJS += reflog-walk.o\n LIB_OBJS += refs.o\n LIB_OBJS += remote.o\ndiff --git a/builtin/apply.c b/builtin/apply.c\nindex ef32e4f..a954147 100644\n--- a/builtin/apply.c\n+++ b/builtin/apply.c\n@@ -3682,6 +3682,8 @@ static void build_fake_ancestor(struct patch *list, const char *filename)\n \t\t\tdie (\"Could not add %s to temporary index\", name);\n \t}\n \n+\tif (!result.initialized)\n+\t\tinitialize_index(&result, 0);\n \tfd = open(filename, O_WRONLY | O_CREAT, 0666);\n \tif (fd < 0 || write_index(&result, fd) || close(fd))\n \t\tdie (\"Could not write temporary index to %s\", filename);\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex e3a10d7..c5bb889 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -863,7 +863,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \n \t\tif (the_index.version != preferred_index_format)\n \t\t\tactive_cache_changed = 1;\n-\t\tthe_index.version = preferred_index_format;\n+\t\tchange_cache_version(preferred_index_format);\n \t}\n \n \tif (read_from_stdin) {\ndiff --git a/cache.h b/cache.h\nindex ce377e1..290e26d 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -95,16 +95,8 @@ unsigned long git_deflate_bound(git_zstream *, unsigned long);\n  */\n #define DEFAULT_GIT_PORT 9418\n \n-/*\n- * Basic data structures for the directory cache\n- */\n \n #define CACHE_SIGNATURE 0x44495243\t/* \"DIRC\" */\n-struct cache_header {\n-\tuint32_t hdr_signature;\n-\tuint32_t hdr_version;\n-\tuint32_t hdr_entries;\n-};\n \n #define INDEX_FORMAT_LB 2\n #define INDEX_FORMAT_UB 4\n@@ -279,6 +271,7 @@ struct index_state {\n \t\t initialized : 1;\n \tstruct hash_table name_hash;\n \tstruct hash_table dir_hash;\n+\tstruct index_ops *ops;\n };\n \n extern struct index_state the_index;\n@@ -296,6 +289,8 @@ extern void free_name_hash(struct index_state *istate);\n #define active_cache_changed (the_index.cache_changed)\n #define active_cache_tree (the_index.cache_tree)\n \n+#define initialize_cache() initialize_index(&the_index, 0)\n+#define change_cache_version(version) change_index_version(&the_index, (version))\n #define read_cache() read_index(&the_index)\n #define read_cache_from(path) read_index_from(&the_index, (path))\n #define read_cache_preload(pathspec) read_index_preload(&the_index, (pathspec))\n@@ -455,6 +450,8 @@ extern void sanitize_stdfds(void);\n \t} while (0)\n \n /* Initialize and use the cache information */\n+extern void initialize_index(struct index_state *istate, int version);\n+extern void change_index_version(struct index_state *istate, int version);\n extern int read_index(struct index_state *);\n extern int read_index_preload(struct index_state *, const struct pathspec *pathspec);\n extern int read_index_from(struct index_state *, const char *path);\ndiff --git a/read-cache-v2.c b/read-cache-v2.c\nnew file mode 100644\nindex 0000000..a7d076c\n--- /dev/null\n+++ b/read-cache-v2.c\n@@ -0,0 +1,553 @@\n+#include \"cache.h\"\n+#include \"read-cache.h\"\n+#include \"resolve-undo.h\"\n+#include \"cache-tree.h\"\n+#include \"varint.h\"\n+\n+/* Mask for the name length in ce_flags in the on-disk index */\n+#define CE_NAMEMASK  (0x0fff)\n+\n+/*****************************************************************\n+ * Index File I/O\n+ *****************************************************************/\n+\n+/*\n+ * dev/ino/uid/gid/size are also just tracked to the low 32 bits\n+ * Again - this is just a (very strong in practice) heuristic that\n+ * the inode hasn't changed.\n+ *\n+ * We save the fields in big-endian order to allow using the\n+ * index file over NFS transparently.\n+ */\n+struct ondisk_cache_entry {\n+\tstruct cache_time ctime;\n+\tstruct cache_time mtime;\n+\tuint32_t dev;\n+\tuint32_t ino;\n+\tuint32_t mode;\n+\tuint32_t uid;\n+\tuint32_t gid;\n+\tuint32_t size;\n+\tunsigned char sha1[20];\n+\tuint16_t flags;\n+\tchar name[FLEX_ARRAY]; /* more */\n+};\n+\n+/*\n+ * This struct is used when CE_EXTENDED bit is 1\n+ * The struct must match ondisk_cache_entry exactly from\n+ * ctime till flags\n+ */\n+struct ondisk_cache_entry_extended {\n+\tstruct cache_time ctime;\n+\tstruct cache_time mtime;\n+\tuint32_t dev;\n+\tuint32_t ino;\n+\tuint32_t mode;\n+\tuint32_t uid;\n+\tuint32_t gid;\n+\tuint32_t size;\n+\tunsigned char sha1[20];\n+\tuint16_t flags;\n+\tuint16_t flags2;\n+\tchar name[FLEX_ARRAY]; /* more */\n+};\n+\n+/* These are only used for v3 or lower */\n+#define align_flex_name(STRUCT,len) ((offsetof(struct STRUCT,name) + (len) + 8) & ~7)\n+#define ondisk_cache_entry_size(len) align_flex_name(ondisk_cache_entry,len)\n+#define ondisk_cache_entry_extended_size(len) align_flex_name(ondisk_cache_entry_extended,len)\n+#define ondisk_ce_size(ce) (((ce)->ce_flags & CE_EXTENDED) ? \\\n+\t\t\t    ondisk_cache_entry_extended_size(ce_namelen(ce)) : \\\n+\t\t\t    ondisk_cache_entry_size(ce_namelen(ce)))\n+\n+static int verify_hdr(void *mmap, unsigned long size)\n+{\n+\tgit_SHA_CTX c;\n+\tunsigned char sha1[20];\n+\n+\tif (size < + sizeof(struct cache_header) + 20)\n+\t\tdie(\"index file smaller than expected\");\n+\n+\tgit_SHA1_Init(&c);\n+\tgit_SHA1_Update(&c, mmap, size - 20);\n+\tgit_SHA1_Final(sha1, &c);\n+\tif (hashcmp(sha1, (unsigned char *)mmap + size - 20))\n+\t\treturn error(\"bad index file sha1 signature\");\n+\treturn 0;\n+}\n+\n+static int match_stat_basic(const struct cache_entry *ce,\n+\t\t\t    struct stat *st, int changed)\n+{\n+\tchanged |= match_stat_data(&ce->ce_stat_data, st);\n+\n+\t/* Racily smudged entry? */\n+\tif (!ce->ce_stat_data.sd_size) {\n+\t\tif (!is_empty_blob_sha1(ce->sha1))\n+\t\t\tchanged |= DATA_CHANGED;\n+\t}\n+\treturn changed;\n+}\n+\n+static struct cache_entry *cache_entry_from_ondisk(struct ondisk_cache_entry *ondisk,\n+\t\t\t\t\t\t   unsigned int flags,\n+\t\t\t\t\t\t   const char *name,\n+\t\t\t\t\t\t   size_t len)\n+{\n+\tstruct cache_entry *ce = xmalloc(cache_entry_size(len));\n+\n+\tce->ce_stat_data.sd_ctime.sec = ntoh_l(ondisk->ctime.sec);\n+\tce->ce_stat_data.sd_mtime.sec = ntoh_l(ondisk->mtime.sec);\n+\tce->ce_stat_data.sd_ctime.nsec = ntoh_l(ondisk->ctime.nsec);\n+\tce->ce_stat_data.sd_mtime.nsec = ntoh_l(ondisk->mtime.nsec);\n+\tce->ce_stat_data.sd_dev   = ntoh_l(ondisk->dev);\n+\tce->ce_stat_data.sd_ino   = ntoh_l(ondisk->ino);\n+\tce->ce_mode  = ntoh_l(ondisk->mode);\n+\tce->ce_stat_data.sd_uid   = ntoh_l(ondisk->uid);\n+\tce->ce_stat_data.sd_gid   = ntoh_l(ondisk->gid);\n+\tce->ce_stat_data.sd_size  = ntoh_l(ondisk->size);\n+\tce->ce_flags = flags & ~CE_NAMEMASK;\n+\tce->ce_namelen = len;\n+\thashcpy(ce->sha1, ondisk->sha1);\n+\tmemcpy(ce->name, name, len);\n+\tce->name[len] = '\\0';\n+\treturn ce;\n+}\n+\n+/*\n+ * Adjacent cache entries tend to share the leading paths, so it makes\n+ * sense to only store the differences in later entries.  In the v4\n+ * on-disk format of the index, each on-disk cache entry stores the\n+ * number of bytes to be stripped from the end of the previous name,\n+ * and the bytes to append to the result, to come up with its name.\n+ */\n+static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n+{\n+\tconst unsigned char *ep, *cp = (const unsigned char *)cp_;\n+\tsize_t len = decode_varint(&cp);\n+\n+\tif (name->len < len)\n+\t\tdie(\"malformed name field in the index\");\n+\tstrbuf_remove(name, name->len - len, len);\n+\tfor (ep = cp; *ep; ep++)\n+\t\t; /* find the end */\n+\tstrbuf_add(name, cp, ep - cp);\n+\treturn (const char *)ep + 1 - cp_;\n+}\n+\n+static struct cache_entry *create_from_disk(struct ondisk_cache_entry *ondisk,\n+\t\t\t\t\t    unsigned long *ent_size,\n+\t\t\t\t\t    struct strbuf *previous_name)\n+{\n+\tstruct cache_entry *ce;\n+\tsize_t len;\n+\tconst char *name;\n+\tunsigned int flags;\n+\n+\t/* On-disk flags are just 16 bits */\n+\tflags = ntoh_s(ondisk->flags);\n+\tlen = flags & CE_NAMEMASK;\n+\n+\tif (flags & CE_EXTENDED) {\n+\t\tstruct ondisk_cache_entry_extended *ondisk2;\n+\t\tint extended_flags;\n+\t\tondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n+\t\textended_flags = ntoh_s(ondisk2->flags2) << 16;\n+\t\t/* We do not yet understand any bit out of CE_EXTENDED_FLAGS */\n+\t\tif (extended_flags & ~CE_EXTENDED_FLAGS)\n+\t\t\tdie(\"Unknown index entry format %08x\", extended_flags);\n+\t\tflags |= extended_flags;\n+\t\tname = ondisk2->name;\n+\t}\n+\telse\n+\t\tname = ondisk->name;\n+\n+\tif (!previous_name) {\n+\t\t/* v3 and earlier */\n+\t\tif (len == CE_NAMEMASK)\n+\t\t\tlen = strlen(name);\n+\t\tce = cache_entry_from_ondisk(ondisk, flags, name, len);\n+\n+\t\t*ent_size = ondisk_ce_size(ce);\n+\t} else {\n+\t\tunsigned long consumed;\n+\t\tconsumed = expand_name_field(previous_name, name);\n+\t\tce = cache_entry_from_ondisk(ondisk, flags,\n+\t\t\t\t\t     previous_name->buf,\n+\t\t\t\t\t     previous_name->len);\n+\n+\t\t*ent_size = (name - ((char *)ondisk)) + consumed;\n+\t}\n+\treturn ce;\n+}\n+\n+static int read_index_extension(struct index_state *istate,\n+\t\t\t\tconst char *ext, void *data, unsigned long sz)\n+{\n+\tswitch (CACHE_EXT(ext)) {\n+\tcase CACHE_EXT_TREE:\n+\t\tistate->cache_tree = cache_tree_read(data, sz);\n+\t\tbreak;\n+\tcase CACHE_EXT_RESOLVE_UNDO:\n+\t\tistate->resolve_undo = resolve_undo_read(data, sz);\n+\t\tbreak;\n+\tdefault:\n+\t\tif (*ext < 'A' || 'Z' < *ext)\n+\t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n+\t\t\t\t     ext);\n+\t\tfprintf(stderr, \"ignoring %.4s extension\\n\", ext);\n+\t\tbreak;\n+\t}\n+\treturn 0;\n+}\n+\n+static int read_index_v2(struct index_state *istate, void *mmap,\n+\t\t\t unsigned long mmap_size)\n+{\n+\tint i;\n+\tunsigned long src_offset;\n+\tstruct cache_header *hdr;\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\n+\thdr = mmap;\n+\tistate->cache_nr = ntohl(hdr->hdr_entries);\n+\tistate->cache_alloc = alloc_nr(istate->cache_nr);\n+\tistate->cache = xcalloc(istate->cache_alloc, sizeof(struct cache_entry *));\n+\n+\tif (istate->version == 4)\n+\t\tprevious_name = &previous_name_buf;\n+\telse\n+\t\tprevious_name = NULL;\n+\n+\tsrc_offset = sizeof(*hdr);\n+\tfor (i = 0; i < istate->cache_nr; i++) {\n+\t\tstruct ondisk_cache_entry *disk_ce;\n+\t\tstruct cache_entry *ce;\n+\t\tunsigned long consumed;\n+\n+\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tce = create_from_disk(disk_ce, &consumed, previous_name);\n+\t\tset_index_entry(istate, i, ce);\n+\n+\t\tsrc_offset += consumed;\n+\t}\n+\tstrbuf_release(&previous_name_buf);\n+\n+\twhile (src_offset <= mmap_size - 20 - 8) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\t\tif (read_index_extension(istate,\n+\t\t\t\t\t(const char *) mmap + src_offset,\n+\t\t\t\t\t(char *) mmap + src_offset + 8,\n+\t\t\t\t\textsize) < 0)\n+\t\t\tgoto unmap;\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\treturn 0;\n+unmap:\n+\tmunmap(mmap, mmap_size);\n+\tdie(\"index file corrupt\");\n+}\n+\n+#define WRITE_BUFFER_SIZE 8192\n+static unsigned char write_buffer[WRITE_BUFFER_SIZE];\n+static unsigned long write_buffer_len;\n+\n+static int ce_write_flush(git_SHA_CTX *context, int fd)\n+{\n+\tunsigned int buffered = write_buffer_len;\n+\tif (buffered) {\n+\t\tgit_SHA1_Update(context, write_buffer, buffered);\n+\t\tif (write_in_full(fd, write_buffer, buffered) != buffered)\n+\t\t\treturn -1;\n+\t\twrite_buffer_len = 0;\n+\t}\n+\treturn 0;\n+}\n+\n+static int ce_write(git_SHA_CTX *context, int fd, void *data, unsigned int len)\n+{\n+\twhile (len) {\n+\t\tunsigned int buffered = write_buffer_len;\n+\t\tunsigned int partial = WRITE_BUFFER_SIZE - buffered;\n+\t\tif (partial > len)\n+\t\t\tpartial = len;\n+\t\tmemcpy(write_buffer + buffered, data, partial);\n+\t\tbuffered += partial;\n+\t\tif (buffered == WRITE_BUFFER_SIZE) {\n+\t\t\twrite_buffer_len = buffered;\n+\t\t\tif (ce_write_flush(context, fd))\n+\t\t\t\treturn -1;\n+\t\t\tbuffered = 0;\n+\t\t}\n+\t\twrite_buffer_len = buffered;\n+\t\tlen -= partial;\n+\t\tdata = (char *) data + partial;\n+\t}\n+\treturn 0;\n+}\n+\n+static int write_index_ext_header(git_SHA_CTX *context, int fd,\n+\t\t\t\t  unsigned int ext, unsigned int sz)\n+{\n+\text = htonl(ext);\n+\tsz = htonl(sz);\n+\treturn ((ce_write(context, fd, &ext, 4) < 0) ||\n+\t\t(ce_write(context, fd, &sz, 4) < 0)) ? -1 : 0;\n+}\n+\n+static int ce_flush(git_SHA_CTX *context, int fd)\n+{\n+\tunsigned int left = write_buffer_len;\n+\n+\tif (left) {\n+\t\twrite_buffer_len = 0;\n+\t\tgit_SHA1_Update(context, write_buffer, left);\n+\t}\n+\n+\t/* Flush first if not enough space for SHA1 signature */\n+\tif (left + 20 > WRITE_BUFFER_SIZE) {\n+\t\tif (write_in_full(fd, write_buffer, left) != left)\n+\t\t\treturn -1;\n+\t\tleft = 0;\n+\t}\n+\n+\t/* Append the SHA1 signature at the end */\n+\tgit_SHA1_Final(write_buffer + left, context);\n+\tleft += 20;\n+\treturn (write_in_full(fd, write_buffer, left) != left) ? -1 : 0;\n+}\n+\n+static void ce_smudge_racily_clean_entry(struct index_state *istate, struct cache_entry *ce)\n+{\n+\t/*\n+\t * The only thing we care about in this function is to smudge the\n+\t * falsely clean entry due to touch-update-touch race, so we leave\n+\t * everything else as they are.  We are called for entries whose\n+\t * ce_stat_data.sd_mtime match the index file mtime.\n+\t *\n+\t * Note that this actually does not do much for gitlinks, for\n+\t * which ce_match_stat_basic() always goes to the actual\n+\t * contents.  The caller checks with is_racy_timestamp() which\n+\t * always says \"no\" for gitlinks, so we are not called for them ;-)\n+\t */\n+\tstruct stat st;\n+\n+\tif (lstat(ce->name, &st) < 0)\n+\t\treturn;\n+\tif (ce_match_stat_basic(istate, ce, &st))\n+\t\treturn;\n+\tif (ce_modified_check_fs(ce, &st)) {\n+\t\t/* This is \"racily clean\"; smudge it.  Note that this\n+\t\t * is a tricky code.  At first glance, it may appear\n+\t\t * that it can break with this sequence:\n+\t\t *\n+\t\t * $ echo xyzzy >frotz\n+\t\t * $ git-update-index --add frotz\n+\t\t * $ : >frotz\n+\t\t * $ sleep 3\n+\t\t * $ echo filfre >nitfol\n+\t\t * $ git-update-index --add nitfol\n+\t\t *\n+\t\t * but it does not.  When the second update-index runs,\n+\t\t * it notices that the entry \"frotz\" has the same timestamp\n+\t\t * as index, and if we were to smudge it by resetting its\n+\t\t * size to zero here, then the object name recorded\n+\t\t * in index is the 6-byte file but the cached stat information\n+\t\t * becomes zero --- which would then match what we would\n+\t\t * obtain from the filesystem next time we stat(\"frotz\").\n+\t\t *\n+\t\t * However, the second update-index, before calling\n+\t\t * this function, notices that the cached size is 6\n+\t\t * bytes and what is on the filesystem is an empty\n+\t\t * file, and never calls us, so the cached size information\n+\t\t * for \"frotz\" stays 6 which does not match the filesystem.\n+\t\t */\n+\t\tce->ce_stat_data.sd_size = 0;\n+\t}\n+}\n+\n+/* Copy miscellaneous fields but not the name */\n+static char *copy_cache_entry_to_ondisk(struct ondisk_cache_entry *ondisk,\n+\t\t\t\t       struct cache_entry *ce)\n+{\n+\tshort flags;\n+\n+\tondisk->ctime.sec = htonl(ce->ce_stat_data.sd_ctime.sec);\n+\tondisk->mtime.sec = htonl(ce->ce_stat_data.sd_mtime.sec);\n+\tondisk->ctime.nsec = htonl(ce->ce_stat_data.sd_ctime.nsec);\n+\tondisk->mtime.nsec = htonl(ce->ce_stat_data.sd_mtime.nsec);\n+\tondisk->dev  = htonl(ce->ce_stat_data.sd_dev);\n+\tondisk->ino  = htonl(ce->ce_stat_data.sd_ino);\n+\tondisk->mode = htonl(ce->ce_mode);\n+\tondisk->uid  = htonl(ce->ce_stat_data.sd_uid);\n+\tondisk->gid  = htonl(ce->ce_stat_data.sd_gid);\n+\tondisk->size = htonl(ce->ce_stat_data.sd_size);\n+\thashcpy(ondisk->sha1, ce->sha1);\n+\n+\tflags = ce->ce_flags;\n+\tflags |= (ce_namelen(ce) >= CE_NAMEMASK ? CE_NAMEMASK : ce_namelen(ce));\n+\tondisk->flags = htons(flags);\n+\tif (ce->ce_flags & CE_EXTENDED) {\n+\t\tstruct ondisk_cache_entry_extended *ondisk2;\n+\t\tondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n+\t\tondisk2->flags2 = htons((ce->ce_flags & CE_EXTENDED_FLAGS) >> 16);\n+\t\treturn ondisk2->name;\n+\t}\n+\telse {\n+\t\treturn ondisk->name;\n+\t}\n+}\n+\n+static int ce_write_entry(git_SHA_CTX *c, int fd, struct cache_entry *ce,\n+\t\t\t  struct strbuf *previous_name)\n+{\n+\tint size;\n+\tstruct ondisk_cache_entry *ondisk;\n+\tchar *name;\n+\tint result;\n+\n+\tif (!previous_name) {\n+\t\tsize = ondisk_ce_size(ce);\n+\t\tondisk = xcalloc(1, size);\n+\t\tname = copy_cache_entry_to_ondisk(ondisk, ce);\n+\t\tmemcpy(name, ce->name, ce_namelen(ce));\n+\t} else {\n+\t\tint common, to_remove, prefix_size;\n+\t\tunsigned char to_remove_vi[16];\n+\t\tfor (common = 0;\n+\t\t     (ce->name[common] &&\n+\t\t      common < previous_name->len &&\n+\t\t      ce->name[common] == previous_name->buf[common]);\n+\t\t     common++)\n+\t\t\t; /* still matching */\n+\t\tto_remove = previous_name->len - common;\n+\t\tprefix_size = encode_varint(to_remove, to_remove_vi);\n+\n+\t\tif (ce->ce_flags & CE_EXTENDED)\n+\t\t\tsize = offsetof(struct ondisk_cache_entry_extended, name);\n+\t\telse\n+\t\t\tsize = offsetof(struct ondisk_cache_entry, name);\n+\t\tsize += prefix_size + (ce_namelen(ce) - common + 1);\n+\n+\t\tondisk = xcalloc(1, size);\n+\t\tname = copy_cache_entry_to_ondisk(ondisk, ce);\n+\t\tmemcpy(name, to_remove_vi, prefix_size);\n+\t\tmemcpy(name + prefix_size, ce->name + common, ce_namelen(ce) - common);\n+\n+\t\tstrbuf_splice(previous_name, common, to_remove,\n+\t\t\t      ce->name + common, ce_namelen(ce) - common);\n+\t}\n+\n+\tresult = ce_write(c, fd, ondisk, size);\n+\tfree(ondisk);\n+\treturn result;\n+}\n+\n+static int write_index_v2(struct index_state *istate, int newfd)\n+{\n+\tgit_SHA_CTX c;\n+\tstruct cache_header hdr;\n+\tint i, err, removed, extended, hdr_version;\n+\tstruct cache_entry **cache = istate->cache;\n+\tint entries = istate->cache_nr;\n+\tstruct stat st;\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\n+\tfor (i = removed = extended = 0; i < entries; i++) {\n+\t\tif (cache[i]->ce_flags & CE_REMOVE)\n+\t\t\tremoved++;\n+\n+\t\t/* reduce extended entries if possible */\n+\t\tcache[i]->ce_flags &= ~CE_EXTENDED;\n+\t\tif (cache[i]->ce_flags & CE_EXTENDED_FLAGS) {\n+\t\t\textended++;\n+\t\t\tcache[i]->ce_flags |= CE_EXTENDED;\n+\t\t}\n+\t}\n+\n+\tif (!istate->version)\n+\t\tistate->version = INDEX_FORMAT_DEFAULT;\n+\n+\t/* demote version 3 to version 2 when the latter suffices */\n+\tif (istate->version == 3 || istate->version == 2)\n+\t\tistate->version = extended ? 3 : 2;\n+\n+\thdr_version = istate->version;\n+\n+\thdr.hdr_signature = htonl(CACHE_SIGNATURE);\n+\thdr.hdr_version = htonl(hdr_version);\n+\thdr.hdr_entries = htonl(entries - removed);\n+\n+\tgit_SHA1_Init(&c);\n+\tif (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n+\t\treturn -1;\n+\n+\tprevious_name = (hdr_version == 4) ? &previous_name_buf : NULL;\n+\tfor (i = 0; i < entries; i++) {\n+\t\tstruct cache_entry *ce = cache[i];\n+\t\tif (ce->ce_flags & CE_REMOVE)\n+\t\t\tcontinue;\n+\t\tif (!ce_uptodate(ce) && is_racy_timestamp(istate, ce))\n+\t\t\tce_smudge_racily_clean_entry(istate, ce);\n+\t\tif (is_null_sha1(ce->sha1)) {\n+\t\t\tstatic const char msg[] = \"cache entry has null sha1: %s\";\n+\t\t\tstatic int allow = -1;\n+\n+\t\t\tif (allow < 0)\n+\t\t\t\tallow = git_env_bool(\"GIT_ALLOW_NULL_SHA1\", 0);\n+\t\t\tif (allow)\n+\t\t\t\twarning(msg, ce->name);\n+\t\t\telse\n+\t\t\t\treturn error(msg, ce->name);\n+\t\t}\n+\t\tif (ce_write_entry(&c, newfd, ce, previous_name) < 0)\n+\t\t\treturn -1;\n+\t}\n+\tstrbuf_release(&previous_name_buf);\n+\n+\t/* Write extension data here */\n+\tif (istate->cache_tree) {\n+\t\tstruct strbuf sb = STRBUF_INIT;\n+\n+\t\tcache_tree_write(&sb, istate->cache_tree);\n+\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_TREE, sb.len) < 0\n+\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\tstrbuf_release(&sb);\n+\t\tif (err)\n+\t\t\treturn -1;\n+\t}\n+\tif (istate->resolve_undo) {\n+\t\tstruct strbuf sb = STRBUF_INIT;\n+\n+\t\tresolve_undo_write(&sb, istate->resolve_undo);\n+\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_RESOLVE_UNDO,\n+\t\t\t\t\t     sb.len) < 0\n+\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\tstrbuf_release(&sb);\n+\t\tif (err)\n+\t\t\treturn -1;\n+\t}\n+\n+\tif (ce_flush(&c, newfd) || fstat(newfd, &st))\n+\t\treturn -1;\n+\tistate->timestamp.sec = (unsigned int)st.st_mtime;\n+\tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n+\treturn 0;\n+}\n+\n+struct index_ops v2_ops = {\n+\tmatch_stat_basic,\n+\tverify_hdr,\n+\tread_index_v2,\n+\twrite_index_v2\n+};\ndiff --git a/read-cache.c b/read-cache.c\nindex 5a8f405..e081084 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -5,6 +5,7 @@\n  */\n #define NO_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"read-cache.h\"\n #include \"cache-tree.h\"\n #include \"refs.h\"\n #include \"dir.h\"\n@@ -17,26 +18,9 @@\n \n static struct cache_entry *refresh_cache_entry(struct cache_entry *ce, int really);\n \n-/* Mask for the name length in ce_flags in the on-disk index */\n-\n-#define CE_NAMEMASK  (0x0fff)\n-\n-/* Index extensions.\n- *\n- * The first letter should be 'A'..'Z' for extensions that are not\n- * necessary for a correct operation (i.e. optimization data).\n- * When new extensions are added that _needs_ to be understood in\n- * order to correctly interpret the index file, pick character that\n- * is outside the range, to cause the reader to abort.\n- */\n-\n-#define CACHE_EXT(s) ( (s[0]<<24)|(s[1]<<16)|(s[2]<<8)|(s[3]) )\n-#define CACHE_EXT_TREE 0x54524545\t/* \"TREE\" */\n-#define CACHE_EXT_RESOLVE_UNDO 0x52455543 /* \"REUC\" */\n-\n struct index_state the_index;\n \n-static void set_index_entry(struct index_state *istate, int nr, struct cache_entry *ce)\n+void set_index_entry(struct index_state *istate, int nr, struct cache_entry *ce)\n {\n \tistate->cache[nr] = ce;\n \tadd_name_hash(istate, ce);\n@@ -190,7 +174,7 @@ static int ce_compare_gitlink(const struct cache_entry *ce)\n \treturn hashcmp(sha1, ce->sha1);\n }\n \n-static int ce_modified_check_fs(const struct cache_entry *ce, struct stat *st)\n+int ce_modified_check_fs(const struct cache_entry *ce, struct stat *st)\n {\n \tswitch (st->st_mode & S_IFMT) {\n \tcase S_IFREG:\n@@ -210,7 +194,18 @@ static int ce_modified_check_fs(const struct cache_entry *ce, struct stat *st)\n \treturn 0;\n }\n \n-static int ce_match_stat_basic(const struct cache_entry *ce, struct stat *st)\n+/*\n+ * Check if the reading/writing operations are set and set them\n+ * to the correct version\n+ */\n+static void set_istate_ops(struct index_state *istate)\n+{\n+\tif (istate->version >= 2 && istate->version <= 4)\n+\t\tistate->ops = &v2_ops;\n+}\n+\n+int ce_match_stat_basic(const struct index_state *istate,\n+\t\t\tconst struct cache_entry *ce, struct stat *st)\n {\n \tunsigned int changed = 0;\n \n@@ -243,19 +238,13 @@ static int ce_match_stat_basic(const struct cache_entry *ce, struct stat *st)\n \t\tdie(\"internal error: ce_mode is %o\", ce->ce_mode);\n \t}\n \n-\tchanged |= match_stat_data(&ce->ce_stat_data, st);\n-\n-\t/* Racily smudged entry? */\n-\tif (!ce->ce_stat_data.sd_size) {\n-\t\tif (!is_empty_blob_sha1(ce->sha1))\n-\t\t\tchanged |= DATA_CHANGED;\n-\t}\n-\n+\tchanged = istate->ops->match_stat_basic(ce, st, changed);\n \treturn changed;\n }\n \n-static int is_racy_timestamp(const struct index_state *istate,\n-\t\t\t     const struct cache_entry *ce)\n+\n+int is_racy_timestamp(const struct index_state *istate,\n+\t\t      const struct cache_entry *ce)\n {\n \treturn (!S_ISGITLINK(ce->ce_mode) &&\n \t\tistate->timestamp.sec &&\n@@ -298,7 +287,7 @@ int ie_match_stat(const struct index_state *istate,\n \tif (ce->ce_flags & CE_INTENT_TO_ADD)\n \t\treturn DATA_CHANGED | TYPE_CHANGED | MODE_CHANGED;\n \n-\tchanged = ce_match_stat_basic(ce, st);\n+\tchanged = ce_match_stat_basic(istate, ce, st);\n \n \t/*\n \t * Within 1 second of this sequence:\n@@ -982,6 +971,8 @@ int add_index_entry(struct index_state *istate, struct cache_entry *ce, int opti\n {\n \tint pos;\n \n+\tif (!istate->initialized)\n+\t\tinitialize_index(istate, INDEX_FORMAT_DEFAULT);\n \tif (option & ADD_CACHE_JUST_APPEND)\n \t\tpos = istate->cache_nr;\n \telse {\n@@ -1212,13 +1203,25 @@ static struct cache_entry *refresh_cache_entry(struct cache_entry *ce, int reall\n \treturn refresh_cache_ent(&the_index, ce, really, NULL, NULL);\n }\n \n+void initialize_index(struct index_state *istate, int version)\n+{\n+\tistate->initialized = 1;\n+\tif (!version)\n+\t\tversion = INDEX_FORMAT_DEFAULT;\n+\tistate->version = version;\n+\tset_istate_ops(istate);\n+}\n+\n+void change_index_version(struct index_state *istate, int version)\n+{\n+\tistate->version = version;\n+\tset_istate_ops(istate);\n+}\n \n /*****************************************************************\n  * Index File I/O\n  *****************************************************************/\n \n-#define INDEX_FORMAT_DEFAULT 3\n-\n /*\n  * dev/ino/uid/gid/size are also just tracked to the low 32 bits\n  * Again - this is just a (very strong in practice) heuristic that\n@@ -1269,7 +1272,8 @@ struct ondisk_cache_entry_extended {\n \t\t\t    ondisk_cache_entry_extended_size(ce_namelen(ce)) : \\\n \t\t\t    ondisk_cache_entry_size(ce_namelen(ce)))\n \n-static int verify_hdr_version(struct cache_header *hdr, unsigned long size)\n+static int verify_hdr_version(struct index_state *istate,\n+\t\t\t      struct cache_header *hdr, unsigned long size)\n {\n \tint hdr_version;\n \n@@ -1278,42 +1282,7 @@ static int verify_hdr_version(struct cache_header *hdr, unsigned long size)\n \thdr_version = ntohl(hdr->hdr_version);\n \tif (hdr_version < INDEX_FORMAT_LB || INDEX_FORMAT_UB < hdr_version)\n \t\treturn error(\"bad index version %d\", hdr_version);\n-\treturn 0;\n-}\n-\n-static int verify_hdr(void *mmap, unsigned long size)\n-{\n-\tgit_SHA_CTX c;\n-\tunsigned char sha1[20];\n-\n-\tif (size < sizeof(struct cache_header) + 20)\n-\t\tdie(\"index file smaller than expected\");\n-\n-\tgit_SHA1_Init(&c);\n-\tgit_SHA1_Update(&c, mmap, size - 20);\n-\tgit_SHA1_Final(sha1, &c);\n-\tif (hashcmp(sha1, (unsigned char *)mmap + size - 20))\n-\t\treturn error(\"bad index file sha1 signature\");\n-\treturn 0;\n-}\n-\n-static int read_index_extension(struct index_state *istate,\n-\t\t\t\tconst char *ext, void *data, unsigned long sz)\n-{\n-\tswitch (CACHE_EXT(ext)) {\n-\tcase CACHE_EXT_TREE:\n-\t\tistate->cache_tree = cache_tree_read(data, sz);\n-\t\tbreak;\n-\tcase CACHE_EXT_RESOLVE_UNDO:\n-\t\tistate->resolve_undo = resolve_undo_read(data, sz);\n-\t\tbreak;\n-\tdefault:\n-\t\tif (*ext < 'A' || 'Z' < *ext)\n-\t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n-\t\t\t\t     ext);\n-\t\tfprintf(stderr, \"ignoring %.4s extension\\n\", ext);\n-\t\tbreak;\n-\t}\n+\tinitialize_index(istate, hdr_version);\n \treturn 0;\n }\n \n@@ -1322,176 +1291,6 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file());\n }\n \n-#ifndef NEEDS_ALIGNED_ACCESS\n-#define ntoh_s(var) ntohs(var)\n-#define ntoh_l(var) ntohl(var)\n-#else\n-static inline uint16_t ntoh_s_force_align(void *p)\n-{\n-\tuint16_t x;\n-\tmemcpy(&x, p, sizeof(x));\n-\treturn ntohs(x);\n-}\n-static inline uint32_t ntoh_l_force_align(void *p)\n-{\n-\tuint32_t x;\n-\tmemcpy(&x, p, sizeof(x));\n-\treturn ntohl(x);\n-}\n-#define ntoh_s(var) ntoh_s_force_align(&(var))\n-#define ntoh_l(var) ntoh_l_force_align(&(var))\n-#endif\n-\n-static struct cache_entry *cache_entry_from_ondisk(struct ondisk_cache_entry *ondisk,\n-\t\t\t\t\t\t   unsigned int flags,\n-\t\t\t\t\t\t   const char *name,\n-\t\t\t\t\t\t   size_t len)\n-{\n-\tstruct cache_entry *ce = xmalloc(cache_entry_size(len));\n-\n-\tce->ce_stat_data.sd_ctime.sec = ntoh_l(ondisk->ctime.sec);\n-\tce->ce_stat_data.sd_mtime.sec = ntoh_l(ondisk->mtime.sec);\n-\tce->ce_stat_data.sd_ctime.nsec = ntoh_l(ondisk->ctime.nsec);\n-\tce->ce_stat_data.sd_mtime.nsec = ntoh_l(ondisk->mtime.nsec);\n-\tce->ce_stat_data.sd_dev   = ntoh_l(ondisk->dev);\n-\tce->ce_stat_data.sd_ino   = ntoh_l(ondisk->ino);\n-\tce->ce_mode  = ntoh_l(ondisk->mode);\n-\tce->ce_stat_data.sd_uid   = ntoh_l(ondisk->uid);\n-\tce->ce_stat_data.sd_gid   = ntoh_l(ondisk->gid);\n-\tce->ce_stat_data.sd_size  = ntoh_l(ondisk->size);\n-\tce->ce_flags = flags & ~CE_NAMEMASK;\n-\tce->ce_namelen = len;\n-\thashcpy(ce->sha1, ondisk->sha1);\n-\tmemcpy(ce->name, name, len);\n-\tce->name[len] = '\\0';\n-\treturn ce;\n-}\n-\n-/*\n- * Adjacent cache entries tend to share the leading paths, so it makes\n- * sense to only store the differences in later entries.  In the v4\n- * on-disk format of the index, each on-disk cache entry stores the\n- * number of bytes to be stripped from the end of the previous name,\n- * and the bytes to append to the result, to come up with its name.\n- */\n-static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n-{\n-\tconst unsigned char *ep, *cp = (const unsigned char *)cp_;\n-\tsize_t len = decode_varint(&cp);\n-\n-\tif (name->len < len)\n-\t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n-\tstrbuf_add(name, cp, ep - cp);\n-\treturn (const char *)ep + 1 - cp_;\n-}\n-\n-static struct cache_entry *create_from_disk(struct ondisk_cache_entry *ondisk,\n-\t\t\t\t\t    unsigned long *ent_size,\n-\t\t\t\t\t    struct strbuf *previous_name)\n-{\n-\tstruct cache_entry *ce;\n-\tsize_t len;\n-\tconst char *name;\n-\tunsigned int flags;\n-\n-\t/* On-disk flags are just 16 bits */\n-\tflags = ntoh_s(ondisk->flags);\n-\tlen = flags & CE_NAMEMASK;\n-\n-\tif (flags & CE_EXTENDED) {\n-\t\tstruct ondisk_cache_entry_extended *ondisk2;\n-\t\tint extended_flags;\n-\t\tondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n-\t\textended_flags = ntoh_s(ondisk2->flags2) << 16;\n-\t\t/* We do not yet understand any bit out of CE_EXTENDED_FLAGS */\n-\t\tif (extended_flags & ~CE_EXTENDED_FLAGS)\n-\t\t\tdie(\"Unknown index entry format %08x\", extended_flags);\n-\t\tflags |= extended_flags;\n-\t\tname = ondisk2->name;\n-\t}\n-\telse\n-\t\tname = ondisk->name;\n-\n-\tif (!previous_name) {\n-\t\t/* v3 and earlier */\n-\t\tif (len == CE_NAMEMASK)\n-\t\t\tlen = strlen(name);\n-\t\tce = cache_entry_from_ondisk(ondisk, flags, name, len);\n-\n-\t\t*ent_size = ondisk_ce_size(ce);\n-\t} else {\n-\t\tunsigned long consumed;\n-\t\tconsumed = expand_name_field(previous_name, name);\n-\t\tce = cache_entry_from_ondisk(ondisk, flags,\n-\t\t\t\t\t     previous_name->buf,\n-\t\t\t\t\t     previous_name->len);\n-\n-\t\t*ent_size = (name - ((char *)ondisk)) + consumed;\n-\t}\n-\treturn ce;\n-}\n-\n-static int read_index_v2(struct index_state *istate, void *mmap, unsigned long mmap_size)\n-{\n-\tint i;\n-\tunsigned long src_offset;\n-\tstruct cache_header *hdr;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n-\n-\thdr = mmap;\n-\n-\tistate->version = ntohl(hdr->hdr_version);\n-\tistate->cache_nr = ntohl(hdr->hdr_entries);\n-\tistate->cache_alloc = alloc_nr(istate->cache_nr);\n-\tistate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n-\tistate->initialized = 1;\n-\n-\tif (istate->version == 4)\n-\t\tprevious_name = &previous_name_buf;\n-\telse\n-\t\tprevious_name = NULL;\n-\n-\tsrc_offset = sizeof(*hdr);\n-\tfor (i = 0; i < istate->cache_nr; i++) {\n-\t\tstruct ondisk_cache_entry *disk_ce;\n-\t\tstruct cache_entry *ce;\n-\t\tunsigned long consumed;\n-\n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(disk_ce, &consumed, previous_name);\n-\t\tset_index_entry(istate, i, ce);\n-\n-\t\tsrc_offset += consumed;\n-\t}\n-\tstrbuf_release(&previous_name_buf);\n-\n-\twhile (src_offset <= mmap_size - 20 - 8) {\n-\t\t/* After an array of active_nr index entries,\n-\t\t * there can be arbitrary number of extended\n-\t\t * sections, each of which is prefixed with\n-\t\t * extension name (4-byte) and section length\n-\t\t * in 4-byte network byte order.\n-\t\t */\n-\t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n-\t\tif (read_index_extension(istate,\n-\t\t\t\t\t (const char *) mmap + src_offset,\n-\t\t\t\t\t (char *) mmap + src_offset + 8,\n-\t\t\t\t\t extsize) < 0)\n-\t\t\tgoto unmap;\n-\t\tsrc_offset += 8;\n-\t\tsrc_offset += extsize;\n-\t}\n-\treturn 0;\n-unmap:\n-\tmunmap(mmap, mmap_size);\n-\tdie(\"index file corrupt\");\n-}\n-\n /* remember to discard_cache() before reading a different cache! */\n int read_index_from(struct index_state *istate, const char *path)\n {\n@@ -1508,10 +1307,13 @@ int read_index_from(struct index_state *istate, const char *path)\n \terrno = ENOENT;\n \tistate->timestamp.sec = 0;\n \tistate->timestamp.nsec = 0;\n+\n \tfd = open(path, O_RDONLY);\n \tif (fd < 0) {\n-\t\tif (errno == ENOENT)\n+\t\tif (errno == ENOENT) {\n+\t\t\tinitialize_index(istate, 0);\n \t\t\treturn 0;\n+\t\t}\n \t\tdie_errno(\"index file open failed\");\n \t}\n \n@@ -1520,24 +1322,23 @@ int read_index_from(struct index_state *istate, const char *path)\n \n \terrno = EINVAL;\n \tmmap_size = xsize_t(st.st_size);\n-\tif (mmap_size < sizeof(struct cache_header) + 20)\n-\t\tdie(\"index file smaller than expected\");\n-\n \tmmap = xmmap(NULL, mmap_size, PROT_READ | PROT_WRITE, MAP_PRIVATE, fd, 0);\n \tclose(fd);\n \tif (mmap == MAP_FAILED)\n \t\tdie_errno(\"unable to map index file\");\n \n \thdr = mmap;\n-\tif (verify_hdr_version(hdr, mmap_size) < 0)\n+\tif (verify_hdr_version(istate, hdr, mmap_size) < 0)\n \t\tgoto unmap;\n \n-\tif (verify_hdr(mmap, mmap_size) < 0)\n+\tif (istate->ops->verify_hdr(mmap, mmap_size) < 0)\n \t\tgoto unmap;\n \n-\tread_index_v2(istate, mmap, mmap_size);\n+\tif (istate->ops->read_index(istate, mmap, mmap_size) < 0)\n+\t\tgoto unmap;\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n+\n \tmunmap(mmap, mmap_size);\n \treturn istate->cache_nr;\n \n@@ -1568,6 +1369,7 @@ int discard_index(struct index_state *istate)\n \tfree(istate->cache);\n \tistate->cache = NULL;\n \tistate->cache_alloc = 0;\n+\tistate->ops = NULL;\n \treturn 0;\n }\n \n@@ -1581,201 +1383,6 @@ int unmerged_index(const struct index_state *istate)\n \treturn 0;\n }\n \n-#define WRITE_BUFFER_SIZE 8192\n-static unsigned char write_buffer[WRITE_BUFFER_SIZE];\n-static unsigned long write_buffer_len;\n-\n-static int ce_write_flush(git_SHA_CTX *context, int fd)\n-{\n-\tunsigned int buffered = write_buffer_len;\n-\tif (buffered) {\n-\t\tgit_SHA1_Update(context, write_buffer, buffered);\n-\t\tif (write_in_full(fd, write_buffer, buffered) != buffered)\n-\t\t\treturn -1;\n-\t\twrite_buffer_len = 0;\n-\t}\n-\treturn 0;\n-}\n-\n-static int ce_write(git_SHA_CTX *context, int fd, void *data, unsigned int len)\n-{\n-\twhile (len) {\n-\t\tunsigned int buffered = write_buffer_len;\n-\t\tunsigned int partial = WRITE_BUFFER_SIZE - buffered;\n-\t\tif (partial > len)\n-\t\t\tpartial = len;\n-\t\tmemcpy(write_buffer + buffered, data, partial);\n-\t\tbuffered += partial;\n-\t\tif (buffered == WRITE_BUFFER_SIZE) {\n-\t\t\twrite_buffer_len = buffered;\n-\t\t\tif (ce_write_flush(context, fd))\n-\t\t\t\treturn -1;\n-\t\t\tbuffered = 0;\n-\t\t}\n-\t\twrite_buffer_len = buffered;\n-\t\tlen -= partial;\n-\t\tdata = (char *) data + partial;\n-\t}\n-\treturn 0;\n-}\n-\n-static int write_index_ext_header(git_SHA_CTX *context, int fd,\n-\t\t\t\t  unsigned int ext, unsigned int sz)\n-{\n-\text = htonl(ext);\n-\tsz = htonl(sz);\n-\treturn ((ce_write(context, fd, &ext, 4) < 0) ||\n-\t\t(ce_write(context, fd, &sz, 4) < 0)) ? -1 : 0;\n-}\n-\n-static int ce_flush(git_SHA_CTX *context, int fd)\n-{\n-\tunsigned int left = write_buffer_len;\n-\n-\tif (left) {\n-\t\twrite_buffer_len = 0;\n-\t\tgit_SHA1_Update(context, write_buffer, left);\n-\t}\n-\n-\t/* Flush first if not enough space for SHA1 signature */\n-\tif (left + 20 > WRITE_BUFFER_SIZE) {\n-\t\tif (write_in_full(fd, write_buffer, left) != left)\n-\t\t\treturn -1;\n-\t\tleft = 0;\n-\t}\n-\n-\t/* Append the SHA1 signature at the end */\n-\tgit_SHA1_Final(write_buffer + left, context);\n-\tleft += 20;\n-\treturn (write_in_full(fd, write_buffer, left) != left) ? -1 : 0;\n-}\n-\n-static void ce_smudge_racily_clean_entry(struct cache_entry *ce)\n-{\n-\t/*\n-\t * The only thing we care about in this function is to smudge the\n-\t * falsely clean entry due to touch-update-touch race, so we leave\n-\t * everything else as they are.  We are called for entries whose\n-\t * ce_stat_data.sd_mtime match the index file mtime.\n-\t *\n-\t * Note that this actually does not do much for gitlinks, for\n-\t * which ce_match_stat_basic() always goes to the actual\n-\t * contents.  The caller checks with is_racy_timestamp() which\n-\t * always says \"no\" for gitlinks, so we are not called for them ;-)\n-\t */\n-\tstruct stat st;\n-\n-\tif (lstat(ce->name, &st) < 0)\n-\t\treturn;\n-\tif (ce_match_stat_basic(ce, &st))\n-\t\treturn;\n-\tif (ce_modified_check_fs(ce, &st)) {\n-\t\t/* This is \"racily clean\"; smudge it.  Note that this\n-\t\t * is a tricky code.  At first glance, it may appear\n-\t\t * that it can break with this sequence:\n-\t\t *\n-\t\t * $ echo xyzzy >frotz\n-\t\t * $ git-update-index --add frotz\n-\t\t * $ : >frotz\n-\t\t * $ sleep 3\n-\t\t * $ echo filfre >nitfol\n-\t\t * $ git-update-index --add nitfol\n-\t\t *\n-\t\t * but it does not.  When the second update-index runs,\n-\t\t * it notices that the entry \"frotz\" has the same timestamp\n-\t\t * as index, and if we were to smudge it by resetting its\n-\t\t * size to zero here, then the object name recorded\n-\t\t * in index is the 6-byte file but the cached stat information\n-\t\t * becomes zero --- which would then match what we would\n-\t\t * obtain from the filesystem next time we stat(\"frotz\").\n-\t\t *\n-\t\t * However, the second update-index, before calling\n-\t\t * this function, notices that the cached size is 6\n-\t\t * bytes and what is on the filesystem is an empty\n-\t\t * file, and never calls us, so the cached size information\n-\t\t * for \"frotz\" stays 6 which does not match the filesystem.\n-\t\t */\n-\t\tce->ce_stat_data.sd_size = 0;\n-\t}\n-}\n-\n-/* Copy miscellaneous fields but not the name */\n-static char *copy_cache_entry_to_ondisk(struct ondisk_cache_entry *ondisk,\n-\t\t\t\t       struct cache_entry *ce)\n-{\n-\tshort flags;\n-\n-\tondisk->ctime.sec = htonl(ce->ce_stat_data.sd_ctime.sec);\n-\tondisk->mtime.sec = htonl(ce->ce_stat_data.sd_mtime.sec);\n-\tondisk->ctime.nsec = htonl(ce->ce_stat_data.sd_ctime.nsec);\n-\tondisk->mtime.nsec = htonl(ce->ce_stat_data.sd_mtime.nsec);\n-\tondisk->dev  = htonl(ce->ce_stat_data.sd_dev);\n-\tondisk->ino  = htonl(ce->ce_stat_data.sd_ino);\n-\tondisk->mode = htonl(ce->ce_mode);\n-\tondisk->uid  = htonl(ce->ce_stat_data.sd_uid);\n-\tondisk->gid  = htonl(ce->ce_stat_data.sd_gid);\n-\tondisk->size = htonl(ce->ce_stat_data.sd_size);\n-\thashcpy(ondisk->sha1, ce->sha1);\n-\n-\tflags = ce->ce_flags;\n-\tflags |= (ce_namelen(ce) >= CE_NAMEMASK ? CE_NAMEMASK : ce_namelen(ce));\n-\tondisk->flags = htons(flags);\n-\tif (ce->ce_flags & CE_EXTENDED) {\n-\t\tstruct ondisk_cache_entry_extended *ondisk2;\n-\t\tondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n-\t\tondisk2->flags2 = htons((ce->ce_flags & CE_EXTENDED_FLAGS) >> 16);\n-\t\treturn ondisk2->name;\n-\t}\n-\telse {\n-\t\treturn ondisk->name;\n-\t}\n-}\n-\n-static int ce_write_entry(git_SHA_CTX *c, int fd, struct cache_entry *ce,\n-\t\t\t  struct strbuf *previous_name)\n-{\n-\tint size;\n-\tstruct ondisk_cache_entry *ondisk;\n-\tchar *name;\n-\tint result;\n-\n-\tif (!previous_name) {\n-\t\tsize = ondisk_ce_size(ce);\n-\t\tondisk = xcalloc(1, size);\n-\t\tname = copy_cache_entry_to_ondisk(ondisk, ce);\n-\t\tmemcpy(name, ce->name, ce_namelen(ce));\n-\t} else {\n-\t\tint common, to_remove, prefix_size;\n-\t\tunsigned char to_remove_vi[16];\n-\t\tfor (common = 0;\n-\t\t     (ce->name[common] &&\n-\t\t      common < previous_name->len &&\n-\t\t      ce->name[common] == previous_name->buf[common]);\n-\t\t     common++)\n-\t\t\t; /* still matching */\n-\t\tto_remove = previous_name->len - common;\n-\t\tprefix_size = encode_varint(to_remove, to_remove_vi);\n-\n-\t\tif (ce->ce_flags & CE_EXTENDED)\n-\t\t\tsize = offsetof(struct ondisk_cache_entry_extended, name);\n-\t\telse\n-\t\t\tsize = offsetof(struct ondisk_cache_entry, name);\n-\t\tsize += prefix_size + (ce_namelen(ce) - common + 1);\n-\n-\t\tondisk = xcalloc(1, size);\n-\t\tname = copy_cache_entry_to_ondisk(ondisk, ce);\n-\t\tmemcpy(name, to_remove_vi, prefix_size);\n-\t\tmemcpy(name + prefix_size, ce->name + common, ce_namelen(ce) - common);\n-\n-\t\tstrbuf_splice(previous_name, common, to_remove,\n-\t\t\t      ce->name + common, ce_namelen(ce) - common);\n-\t}\n-\n-\tresult = ce_write(c, fd, ondisk, size);\n-\tfree(ondisk);\n-\treturn result;\n-}\n-\n static int has_racy_timestamp(struct index_state *istate)\n {\n \tint entries = istate->cache_nr;\n@@ -1801,101 +1408,9 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n \t\trollback_lock_file(lockfile);\n }\n \n-static int write_index_v2(struct index_state *istate, int newfd)\n-{\n-\tgit_SHA_CTX c;\n-\tstruct cache_header hdr;\n-\tint i, err, removed, extended, hdr_version;\n-\tstruct cache_entry **cache = istate->cache;\n-\tint entries = istate->cache_nr;\n-\tstruct stat st;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n-\n-\tfor (i = removed = extended = 0; i < entries; i++) {\n-\t\tif (cache[i]->ce_flags & CE_REMOVE)\n-\t\t\tremoved++;\n-\n-\t\t/* reduce extended entries if possible */\n-\t\tcache[i]->ce_flags &= ~CE_EXTENDED;\n-\t\tif (cache[i]->ce_flags & CE_EXTENDED_FLAGS) {\n-\t\t\textended++;\n-\t\t\tcache[i]->ce_flags |= CE_EXTENDED;\n-\t\t}\n-\t}\n-\n-\tif (!istate->version)\n-\t\tistate->version = INDEX_FORMAT_DEFAULT;\n-\n-\t/* demote version 3 to version 2 when the latter suffices */\n-\tif (istate->version == 3 || istate->version == 2)\n-\t\tistate->version = extended ? 3 : 2;\n-\n-\thdr_version = istate->version;\n-\n-\thdr.hdr_signature = htonl(CACHE_SIGNATURE);\n-\thdr.hdr_version = htonl(hdr_version);\n-\thdr.hdr_entries = htonl(entries - removed);\n-\n-\tgit_SHA1_Init(&c);\n-\tif (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n-\t\treturn -1;\n-\n-\tprevious_name = (hdr_version == 4) ? &previous_name_buf : NULL;\n-\tfor (i = 0; i < entries; i++) {\n-\t\tstruct cache_entry *ce = cache[i];\n-\t\tif (ce->ce_flags & CE_REMOVE)\n-\t\t\tcontinue;\n-\t\tif (!ce_uptodate(ce) && is_racy_timestamp(istate, ce))\n-\t\t\tce_smudge_racily_clean_entry(ce);\n-\t\tif (is_null_sha1(ce->sha1)) {\n-\t\t\tstatic const char msg[] = \"cache entry has null sha1: %s\";\n-\t\t\tstatic int allow = -1;\n-\n-\t\t\tif (allow < 0)\n-\t\t\t\tallow = git_env_bool(\"GIT_ALLOW_NULL_SHA1\", 0);\n-\t\t\tif (allow)\n-\t\t\t\twarning(msg, ce->name);\n-\t\t\telse\n-\t\t\t\treturn error(msg, ce->name);\n-\t\t}\n-\t\tif (ce_write_entry(&c, newfd, ce, previous_name) < 0)\n-\t\t\treturn -1;\n-\t}\n-\tstrbuf_release(&previous_name_buf);\n-\n-\t/* Write extension data here */\n-\tif (istate->cache_tree) {\n-\t\tstruct strbuf sb = STRBUF_INIT;\n-\n-\t\tcache_tree_write(&sb, istate->cache_tree);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_TREE, sb.len) < 0\n-\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n-\t\tstrbuf_release(&sb);\n-\t\tif (err)\n-\t\t\treturn -1;\n-\t}\n-\tif (istate->resolve_undo) {\n-\t\tstruct strbuf sb = STRBUF_INIT;\n-\n-\t\tresolve_undo_write(&sb, istate->resolve_undo);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_RESOLVE_UNDO,\n-\t\t\t\t\t     sb.len) < 0\n-\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n-\t\tstrbuf_release(&sb);\n-\t\tif (err)\n-\t\t\treturn -1;\n-\t}\n-\n-\tif (ce_flush(&c, newfd) || fstat(newfd, &st))\n-\t\treturn -1;\n-\tistate->timestamp.sec = (unsigned int)st.st_mtime;\n-\tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n-\treturn 0;\n-}\n-\n int write_index(struct index_state *istate, int newfd)\n {\n-\treturn write_index_v2(istate, newfd);\n+\treturn istate->ops->write_index(istate, newfd);\n }\n \n /*\ndiff --git a/read-cache.h b/read-cache.h\nnew file mode 100644\nindex 0000000..ceedcae\n--- /dev/null\n+++ b/read-cache.h\n@@ -0,0 +1,63 @@\n+#ifndef READ_CACHE_H\n+#define READ_CACHE_H\n+\n+/* Index extensions.\n+ *\n+ * The first letter should be 'A'..'Z' for extensions that are not\n+ * necessary for a correct operation (i.e. optimization data).\n+ * When new extensions are added that _needs_ to be understood in\n+ * order to correctly interpret the index file, pick character that\n+ * is outside the range, to cause the reader to abort.\n+ */\n+\n+#define CACHE_EXT(s) ( (s[0]<<24)|(s[1]<<16)|(s[2]<<8)|(s[3]) )\n+#define CACHE_EXT_TREE 0x54524545\t/* \"TREE\" */\n+#define CACHE_EXT_RESOLVE_UNDO 0x52455543 /* \"REUC\" */\n+\n+#define INDEX_FORMAT_DEFAULT 3\n+\n+/*\n+ * Basic data structures for the directory cache\n+ */\n+struct cache_header {\n+\tuint32_t hdr_signature;\n+\tuint32_t hdr_version;\n+\tuint32_t hdr_entries;\n+};\n+\n+struct index_ops {\n+\tint (*match_stat_basic)(const struct cache_entry *ce, struct stat *st, int changed);\n+\tint (*verify_hdr)(void *mmap, unsigned long size);\n+\tint (*read_index)(struct index_state *istate, void *mmap, unsigned long mmap_size);\n+\tint (*write_index)(struct index_state *istate, int newfd);\n+};\n+\n+extern struct index_ops v2_ops;\n+\n+#ifndef NEEDS_ALIGNED_ACCESS\n+#define ntoh_s(var) ntohs(var)\n+#define ntoh_l(var) ntohl(var)\n+#else\n+static inline uint16_t ntoh_s_force_align(void *p)\n+{\n+\tuint16_t x;\n+\tmemcpy(&x, p, sizeof(x));\n+\treturn ntohs(x);\n+}\n+static inline uint32_t ntoh_l_force_align(void *p)\n+{\n+\tuint32_t x;\n+\tmemcpy(&x, p, sizeof(x));\n+\treturn ntohl(x);\n+}\n+#define ntoh_s(var) ntoh_s_force_align(&(var))\n+#define ntoh_l(var) ntoh_l_force_align(&(var))\n+#endif\n+\n+extern int ce_modified_check_fs(const struct cache_entry *ce, struct stat *st);\n+extern int ce_match_stat_basic(const struct index_state *istate,\n+\t\t\t       const struct cache_entry *ce, struct stat *st);\n+extern int is_racy_timestamp(const struct index_state *istate, const struct cache_entry *ce);\n+extern void set_index_entry(struct index_state *istate, int nr, struct cache_entry *ce);\n+\n+#endif\ndiff --git a/test-index-version.c b/test-index-version.c\nindex 05d4699..d3c0ebd 100644\n--- a/test-index-version.c\n+++ b/test-index-version.c\n@@ -1,5 +1,11 @@\n #include \"cache.h\"\n \n+struct cache_header {\n+\tuint32_t hdr_signature;\n+\tuint32_t hdr_version;\n+\tuint32_t hdr_entries;\n+};\n+\n int main(int argc, char **argv)\n {\n \tstruct cache_header hdr;\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 35cb05e..8d07c2a 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -1035,10 +1035,9 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t}\n \n \tmemset(&o->result, 0, sizeof(o->result));\n-\to->result.initialized = 1;\n+\tinitialize_index(&o->result, o->src_index->version);\n \to->result.timestamp.sec = o->src_index->timestamp.sec;\n \to->result.timestamp.nsec = o->src_index->timestamp.nsec;\n-\to->result.version = o->src_index->version;\n \to->merge_size = len;\n \tmark_all_ce_unused(o->src_index);\n \n-- \n1.8.4.2\n"},{"id":"231168","messageId":"1385553659-9928-5-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 04/24] read-cache: Re-read index if index file changed","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:39Z","receivedAt":"2013-11-27T12:00:39Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Add the possibility of re-reading the index file, if it changed\nwhile reading.\n\nThe index file might change during the read, causing outdated\ninformation to be displayed. We check if the index file changed\nby using its stat data as heuristic.\n\nHelped-by: Ramsay Jones <ramsay@ramsay1.demon.co.uk>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n read-cache.c | 65 +++++++++++++++++++++++++++++++-----------------------------\n 1 file changed, 34 insertions(+), 31 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex e081084..51be1bb 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1294,8 +1294,8 @@ int read_index(struct index_state *istate)\n /* remember to discard_cache() before reading a different cache! */\n int read_index_from(struct index_state *istate, const char *path)\n {\n-\tint fd;\n-\tstruct stat st;\n+\tint fd, err, i;\n+\tstruct stat_validity sv;\n \tstruct cache_header *hdr;\n \tvoid *mmap;\n \tsize_t mmap_size;\n@@ -1307,43 +1307,46 @@ int read_index_from(struct index_state *istate, const char *path)\n \terrno = ENOENT;\n \tistate->timestamp.sec = 0;\n \tistate->timestamp.nsec = 0;\n-\n-\tfd = open(path, O_RDONLY);\n-\tif (fd < 0) {\n-\t\tif (errno == ENOENT) {\n-\t\t\tinitialize_index(istate, 0);\n-\t\t\treturn 0;\n+\tsv.sd = NULL;\n+\tfor (i = 0; i < 50; i++) {\n+\t\terr = 0;\n+\t\tfd = open(path, O_RDONLY);\n+\t\tif (fd < 0) {\n+\t\t\tif (errno == ENOENT) {\n+\t\t\t\tinitialize_index(istate, 0);\n+\t\t\t\treturn 0;\n+\t\t\t}\n+\t\t\tdie_errno(\"index file open failed\");\n \t\t}\n-\t\tdie_errno(\"index file open failed\");\n-\t}\n \n-\tif (fstat(fd, &st))\n-\t\tdie_errno(\"cannot stat the open index\");\n+\t\tstat_validity_update(&sv, fd);\n+\t\tif (!sv.sd)\n+\t\t\tdie_errno(\"cannot stat the open index\");\n \n-\terrno = EINVAL;\n-\tmmap_size = xsize_t(st.st_size);\n-\tmmap = xmmap(NULL, mmap_size, PROT_READ | PROT_WRITE, MAP_PRIVATE, fd, 0);\n-\tclose(fd);\n-\tif (mmap == MAP_FAILED)\n-\t\tdie_errno(\"unable to map index file\");\n+\t\terrno = EINVAL;\n+\t\tmmap_size = xsize_t(sv.sd->sd_size);\n+\t\tmmap = xmmap(NULL, mmap_size, PROT_READ | PROT_WRITE, MAP_PRIVATE, fd, 0);\n+\t\tclose(fd);\n+\t\tif (mmap == MAP_FAILED)\n+\t\t\tdie_errno(\"unable to map index file\");\n \n-\thdr = mmap;\n-\tif (verify_hdr_version(istate, hdr, mmap_size) < 0)\n-\t\tgoto unmap;\n+\t\thdr = mmap;\n+\t\tif (verify_hdr_version(istate, hdr, mmap_size) < 0)\n+\t\t\terr = 1;\n \n-\tif (istate->ops->verify_hdr(mmap, mmap_size) < 0)\n-\t\tgoto unmap;\n+\t\tif (!err && istate->ops->verify_hdr(mmap, mmap_size) < 0)\n+\t\t\terr = 1;\n \n-\tif (istate->ops->read_index(istate, mmap, mmap_size) < 0)\n-\t\tgoto unmap;\n-\tistate->timestamp.sec = st.st_mtime;\n-\tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n+\t\tif (!err && istate->ops->read_index(istate, mmap, mmap_size) < 0)\n+\t\t\terr = 1;\n+\t\tistate->timestamp.sec = sv.sd->sd_mtime.sec;\n+\t\tistate->timestamp.nsec = sv.sd->sd_mtime.nsec;\n \n-\tmunmap(mmap, mmap_size);\n-\treturn istate->cache_nr;\n+\t\tmunmap(mmap, mmap_size);\n+\t\tif (stat_validity_check(&sv, path) && !err)\n+\t\t\treturn istate->cache_nr;\n+\t}\n \n-unmap:\n-\tmunmap(mmap, mmap_size);\n \tdie(\"index file corrupt\");\n }\n \n-- \n1.8.4.2\n"},{"id":"231167","messageId":"1385553659-9928-6-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 05/24] add documentation for the index api","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:40Z","receivedAt":"2013-11-27T12:00:40Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Add documentation for the index reading api.  This also includes\ndocumentation for the new api functions introduced in the next patch.\n\nHelped-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n Documentation/technical/api-in-core-index.txt | 56 +++++++++++++++++++++++++--\n 1 file changed, 52 insertions(+), 4 deletions(-)\n\ndiff --git a/Documentation/technical/api-in-core-index.txt b/Documentation/technical/api-in-core-index.txt\nindex adbdbf5..2cf7a71 100644\n--- a/Documentation/technical/api-in-core-index.txt\n+++ b/Documentation/technical/api-in-core-index.txt\n@@ -1,14 +1,62 @@\n in-core index API\n =================\n \n+Reading API\n+-----------\n+\n+`cache`::\n+\n+\tAn array of cache entries.  This is used to access the cache\n+\tentries directly.  Use `index_name_pos` to search for the\n+\tindex of a specific cache entry.\n+\n+`read_index_filtered`::\n+\n+\tRead a part of the index, filtered by the pathspec given in\n+\tthe opts.  The function may load more than necessary, so the\n+\tcaller is still responsible for applying filters appropriately.  The\n+\tfiltering is only done for performance reasons, as it's\n+\tpossible to only read part of the index when the on-disk\n+\tformat is index-v5.\n++\n+To iterate only over the entries that match the pathspec, use\n+the for_each_index_entry function.\n+\n+`read_index`::\n+\n+\tRead the whole index file from disk.\n+\n+`index_name_pos`::\n+\n+\tFind a cache_entry with name in the index.  Returns pos if an\n+\tentry is matched exactly and -1-pos if an entry is matched\n+\tpartially. e.g.\n++\n+....\n+index:\n+\tfile1\n+\tfile2\n+\tpath/file1\n+\tzzz\n+....\n++\n+`index_name_pos(\"path/file1\", 10)` returns 2, while\n+`index_name_pos(\"path\", 4)` returns -3\n+\n+`for_each_index_entry`::\n+\n+\tIterates over all cache_entries in the index filtered by\n+\tfilter_opts in the index_state.  For each cache entry fn is\n+\texecuted with cb_data as callback data.  From within the loop\n+\tdo `return 0` to continue, or `return 1` to break the loop.\n+\n+TODO\n+----\n Talk about <read-cache.c> and <cache-tree.c>, things like:\n \n-* cache -> the_index macros\n-* read_index()\n * write_index()\n * ie_match_stat() and ie_modified(); how they are different and when to\n   use which.\n-* index_name_pos()\n * remove_index_entry_at()\n * remove_file_from_index()\n * add_file_to_index()\n@@ -18,4 +66,4 @@ Talk about <read-cache.c> and <cache-tree.c>, things like:\n * cache_tree_invalidate_path()\n * cache_tree_update()\n \n-(JC, Linus)\n+(JC, Linus, Thomas Gummerer)\n-- \n1.8.4.2\n"},{"id":"231170","messageId":"1385553659-9928-7-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 06/24] read-cache: add index reading api","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:41Z","receivedAt":"2013-11-27T12:00:41Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Add an api for access to the index file.  Currently there is only a very\nbasic api for accessing the index file, which only allows a full read of\nthe index, and lets the users of the data filter it.  The new index api\ngives the users the possibility to use only part of the index and\nprovides functions for iterating over and accessing cache entries.\n\nThis simplifies future improvements to the in-memory format, as changes\nwill be concentrated on one file, instead of the whole git source code.\n\nHelped-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n cache.h         | 41 ++++++++++++++++++++++++++++++++++++++++-\n read-cache-v2.c | 10 ++++++++--\n read-cache.c    | 47 +++++++++++++++++++++++++++++++++++++++++++----\n read-cache.h    |  3 ++-\n 4 files changed, 93 insertions(+), 8 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 290e26d..38d57e7 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -127,7 +127,7 @@ struct cache_entry {\n \tunsigned int ce_flags;\n \tunsigned int ce_namelen;\n \tunsigned char sha1[20];\n-\tstruct cache_entry *next;\n+\tstruct cache_entry *next; /* used by name_hash */\n \tchar name[FLEX_ARRAY]; /* more */\n };\n \n@@ -260,6 +260,29 @@ static inline unsigned int canon_mode(unsigned int mode)\n \n #define cache_entry_size(len) (offsetof(struct cache_entry,name) + (len) + 1)\n \n+/*\n+ * Options by which the index should be filtered when read partially.\n+ *\n+ * pathspec: The pathspec which the index entries have to match\n+ * seen: Used to return the seen parameter from match_pathspec()\n+ * max_prefix_len: The common prefix length of the pathspecs\n+ *\n+ * read_staged: used to indicate if the conflicted entries (entries\n+ *     with a stage) should be included\n+ * read_cache_tree: used to indicate if the cache-tree should be read\n+ * read_resolve_undo: used to indicate if the resolve undo data should\n+ *     be read\n+ */\n+struct filter_opts {\n+\tconst struct pathspec *pathspec;\n+\tchar *seen;\n+\tint max_prefix_len;\n+\n+\tint read_staged;\n+\tint read_cache_tree;\n+\tint read_resolve_undo;\n+};\n+\n struct index_state {\n \tstruct cache_entry **cache;\n \tunsigned int version;\n@@ -272,6 +295,7 @@ struct index_state {\n \tstruct hash_table name_hash;\n \tstruct hash_table dir_hash;\n \tstruct index_ops *ops;\n+\tstruct filter_opts *filter_opts;\n };\n \n extern struct index_state the_index;\n@@ -317,6 +341,12 @@ extern void free_name_hash(struct index_state *istate);\n #define unmerge_cache_entry_at(at) unmerge_index_entry_at(&the_index, at)\n #define unmerge_cache(pathspec) unmerge_index(&the_index, pathspec)\n #define read_blob_data_from_cache(path, sz) read_blob_data_from_index(&the_index, (path), (sz))\n+\n+/* index api */\n+#define read_cache_filtered(opts) read_index_filtered(&the_index, (opts))\n+#define read_cache_filtered_from(path, opts) read_index_filtered_from(&the_index, (path), (opts))\n+#define for_each_cache_entry(fn, cb_data) \\\n+\tfor_each_index_entry(&the_index, (fn), (cb_data))\n #endif\n \n enum object_type {\n@@ -449,6 +479,15 @@ extern void sanitize_stdfds(void);\n \t\t} \\\n \t} while (0)\n \n+/* index api */\n+extern int read_index_filtered(struct index_state *, struct filter_opts *opts);\n+extern int read_index_filtered_from(struct index_state *, const char *path, struct filter_opts *opts);\n+\n+typedef int each_cache_entry_fn(struct cache_entry *ce, void *);\n+extern int for_each_index_entry(struct index_state *istate,\n+\t\t\t\teach_cache_entry_fn, void *);\n+\n+\n /* Initialize and use the cache information */\n extern void initialize_index(struct index_state *istate, int version);\n extern void change_index_version(struct index_state *istate, int version);\ndiff --git a/read-cache-v2.c b/read-cache-v2.c\nindex a7d076c..f884c10 100644\n--- a/read-cache-v2.c\n+++ b/read-cache-v2.c\n@@ -3,6 +3,7 @@\n #include \"resolve-undo.h\"\n #include \"cache-tree.h\"\n #include \"varint.h\"\n+#include \"dir.h\"\n \n /* Mask for the name length in ce_flags in the on-disk index */\n #define CE_NAMEMASK  (0x0fff)\n@@ -202,8 +203,14 @@ static int read_index_extension(struct index_state *istate,\n \treturn 0;\n }\n \n+/*\n+ * The performance is the same if we read the whole index or only\n+ * part of it, therefore we always read the whole index to avoid\n+ * having to re-read it later.  The filter_opts will determine\n+ * what part of the index is used when retrieving the cache-entries.\n+ */\n static int read_index_v2(struct index_state *istate, void *mmap,\n-\t\t\t unsigned long mmap_size)\n+\t\t\t unsigned long mmap_size, struct filter_opts *opts)\n {\n \tint i;\n \tunsigned long src_offset;\n@@ -229,7 +236,6 @@ static int read_index_v2(struct index_state *istate, void *mmap,\n \t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n \t\tce = create_from_disk(disk_ce, &consumed, previous_name);\n \t\tset_index_entry(istate, i, ce);\n-\n \t\tsrc_offset += consumed;\n \t}\n \tstrbuf_release(&previous_name_buf);\ndiff --git a/read-cache.c b/read-cache.c\nindex 51be1bb..01f5397 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1288,11 +1288,41 @@ static int verify_hdr_version(struct index_state *istate,\n \n int read_index(struct index_state *istate)\n {\n-\treturn read_index_from(istate, get_index_file());\n+\treturn read_index_filtered_from(istate, get_index_file(), NULL);\n }\n \n-/* remember to discard_cache() before reading a different cache! */\n-int read_index_from(struct index_state *istate, const char *path)\n+int read_index_filtered(struct index_state *istate, struct filter_opts *opts)\n+{\n+\treturn read_index_filtered_from(istate, get_index_file(), opts);\n+}\n+\n+/*\n+ * Execute fn for each index entry which is currently in istate.  Data\n+ * can be given to the function using the cb_data parameter.\n+ */\n+int for_each_index_entry(struct index_state *istate, each_cache_entry_fn fn, void *cb_data)\n+{\n+\tint i, ret = 0;\n+\tstruct filter_opts *opts = istate->filter_opts;\n+\n+\tfor (i = 0; i < istate->cache_nr; i++) {\n+\t\tstruct cache_entry *ce = istate->cache[i];\n+\n+\t\tif (opts && !opts->read_staged && ce_stage(ce))\n+\t\t\tcontinue;\n+\n+\t\tif (opts && !match_pathspec_depth(opts->pathspec, ce->name, ce_namelen(ce),\n+\t\t\t\t\t\t  opts->max_prefix_len, opts->seen))\n+\t\t\tcontinue;\n+\n+\t\tif ((ret = fn(istate->cache[i], cb_data)))\n+\t\t\tbreak;\n+\t}\n+\treturn ret;\n+}\n+\n+int read_index_filtered_from(struct index_state *istate, const char *path,\n+\t\t\t     struct filter_opts *opts)\n {\n \tint fd, err, i;\n \tstruct stat_validity sv;\n@@ -1307,6 +1337,7 @@ int read_index_from(struct index_state *istate, const char *path)\n \terrno = ENOENT;\n \tistate->timestamp.sec = 0;\n \tistate->timestamp.nsec = 0;\n+\tistate->filter_opts = opts;\n \tsv.sd = NULL;\n \tfor (i = 0; i < 50; i++) {\n \t\terr = 0;\n@@ -1337,7 +1368,7 @@ int read_index_from(struct index_state *istate, const char *path)\n \t\tif (!err && istate->ops->verify_hdr(mmap, mmap_size) < 0)\n \t\t\terr = 1;\n \n-\t\tif (!err && istate->ops->read_index(istate, mmap, mmap_size) < 0)\n+\t\tif (!err && istate->ops->read_index(istate, mmap, mmap_size, opts) < 0)\n \t\t\terr = 1;\n \t\tistate->timestamp.sec = sv.sd->sd_mtime.sec;\n \t\tistate->timestamp.nsec = sv.sd->sd_mtime.nsec;\n@@ -1350,6 +1381,13 @@ int read_index_from(struct index_state *istate, const char *path)\n \tdie(\"index file corrupt\");\n }\n \n+\n+/* remember to discard_cache() before reading a different cache! */\n+int read_index_from(struct index_state *istate, const char *path)\n+{\n+\treturn read_index_filtered_from(istate, path, NULL);\n+}\n+\n int is_index_unborn(struct index_state *istate)\n {\n \treturn (!istate->cache_nr && !istate->timestamp.sec);\n@@ -1373,6 +1411,7 @@ int discard_index(struct index_state *istate)\n \tistate->cache = NULL;\n \tistate->cache_alloc = 0;\n \tistate->ops = NULL;\n+\tistate->filter_opts = NULL;\n \treturn 0;\n }\n \ndiff --git a/read-cache.h b/read-cache.h\nindex ceedcae..f920546 100644\n--- a/read-cache.h\n+++ b/read-cache.h\n@@ -28,7 +28,8 @@ struct cache_header {\n struct index_ops {\n \tint (*match_stat_basic)(const struct cache_entry *ce, struct stat *st, int changed);\n \tint (*verify_hdr)(void *mmap, unsigned long size);\n-\tint (*read_index)(struct index_state *istate, void *mmap, unsigned long mmap_size);\n+\tint (*read_index)(struct index_state *istate, void *mmap, unsigned long mmap_size,\n+\t\t\t  struct filter_opts *opts);\n \tint (*write_index)(struct index_state *istate, int newfd);\n };\n \n-- \n1.8.4.2\n"},{"id":"231171","messageId":"1385553659-9928-8-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 07/24] make sure partially read index is not changed","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:42Z","receivedAt":"2013-11-27T12:00:42Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"A partially read index file currently cannot be written to disk.  Make\nsure that never happens by erroring out when a caller tries to write a\npartially read index.  Do the same when trying to re-read a partially\nread index without having discarded it first to avoid losing any\ninformation.\n\nForcing the caller to load the right part of the index file, instead of\nre-reading it when changing it, gives a bit of a performance advantage\nby avoiding reading parts of the index twice.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n read-cache.c | 4 ++++\n 1 file changed, 4 insertions(+)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 01f5397..7020f26 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1330,6 +1330,8 @@ int read_index_filtered_from(struct index_state *istate, const char *path,\n \tvoid *mmap;\n \tsize_t mmap_size;\n \n+\tif (istate->filter_opts)\n+\t\tdie(\"BUG: cannot re-read partially read index\");\n \terrno = EBUSY;\n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1452,6 +1454,8 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n \n int write_index(struct index_state *istate, int newfd)\n {\n+\tif (istate->filter_opts)\n+\t\tdie(\"BUG: cannot write a partially read index\");\n \treturn istate->ops->write_index(istate, newfd);\n }\n \n-- \n1.8.4.2\n"},{"id":"231172","messageId":"1385553659-9928-9-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 08/24] grep.c: use index api","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:43Z","receivedAt":"2013-11-27T12:00:43Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n builtin/grep.c | 69 +++++++++++++++++++++++++++++-----------------------------\n 1 file changed, 35 insertions(+), 34 deletions(-)\n\ndiff --git a/builtin/grep.c b/builtin/grep.c\nindex 63f8603..36c2bf0 100644\n--- a/builtin/grep.c\n+++ b/builtin/grep.c\n@@ -369,41 +369,31 @@ static void run_pager(struct grep_opt *opt, const char *prefix)\n \tfree(argv);\n }\n \n-static int grep_cache(struct grep_opt *opt, const struct pathspec *pathspec, int cached)\n+struct grep_opts {\n+\tstruct grep_opt *opt;\n+\tconst struct pathspec *pathspec;\n+\tint cached;\n+\tint hit;\n+};\n+\n+static int grep_cache(struct cache_entry *ce, void *cb_data)\n {\n-\tint hit = 0;\n-\tint nr;\n-\tread_cache();\n+\tstruct grep_opts *opts = cb_data;\n \n-\tfor (nr = 0; nr < active_nr; nr++) {\n-\t\tconst struct cache_entry *ce = active_cache[nr];\n-\t\tif (!S_ISREG(ce->ce_mode))\n-\t\t\tcontinue;\n-\t\tif (!match_pathspec_depth(pathspec, ce->name, ce_namelen(ce), 0, NULL))\n-\t\t\tcontinue;\n-\t\t/*\n-\t\t * If CE_VALID is on, we assume worktree file and its cache entry\n-\t\t * are identical, even if worktree file has been modified, so use\n-\t\t * cache version instead\n-\t\t */\n-\t\tif (cached || (ce->ce_flags & CE_VALID) || ce_skip_worktree(ce)) {\n-\t\t\tif (ce_stage(ce))\n-\t\t\t\tcontinue;\n-\t\t\thit |= grep_sha1(opt, ce->sha1, ce->name, 0, ce->name);\n-\t\t}\n-\t\telse\n-\t\t\thit |= grep_file(opt, ce->name);\n-\t\tif (ce_stage(ce)) {\n-\t\t\tdo {\n-\t\t\t\tnr++;\n-\t\t\t} while (nr < active_nr &&\n-\t\t\t\t !strcmp(ce->name, active_cache[nr]->name));\n-\t\t\tnr--; /* compensate for loop control */\n-\t\t}\n-\t\tif (hit && opt->status_only)\n-\t\t\tbreak;\n-\t}\n-\treturn hit;\n+\tif (!S_ISREG(ce->ce_mode))\n+\t\treturn 0;\n+\t/*\n+\t * If CE_VALID is on, we assume worktree file and its cache entry\n+\t * are identical, even if worktree file has been modified, so use\n+\t * cache version instead\n+\t */\n+\tif (opts->cached || (ce->ce_flags & CE_VALID) || ce_skip_worktree(ce))\n+\t\topts->hit |= grep_sha1(opts->opt, ce->sha1, ce->name, 0, ce->name);\n+\telse\n+\t\topts->hit |= grep_file(opts->opt, ce->name);\n+\tif (opts->hit && opts->opt->status_only)\n+\t\treturn 1;\n+\treturn 0;\n }\n \n static int grep_tree(struct grep_opt *opt, const struct pathspec *pathspec,\n@@ -900,10 +890,21 @@ int cmd_grep(int argc, const char **argv, const char *prefix)\n \t} else if (0 <= opt_exclude) {\n \t\tdie(_(\"--[no-]exclude-standard cannot be used for tracked contents.\"));\n \t} else if (!list.nr) {\n+\t\tstruct grep_opts opts;\n+\t\tstruct filter_opts *filter_opts = xmalloc(sizeof(*filter_opts));\n+\n \t\tif (!cached)\n \t\t\tsetup_work_tree();\n \n-\t\thit = grep_cache(&opt, &pathspec, cached);\n+\t\tmemset(filter_opts, 0, sizeof(*filter_opts));\n+\t\tfilter_opts->pathspec = &pathspec;\n+\t\topts.opt = &opt;\n+\t\topts.pathspec = &pathspec;\n+\t\topts.cached = cached;\n+\t\topts.hit = 0;\n+\t\tread_cache_filtered(filter_opts);\n+\t\tfor_each_cache_entry(grep_cache, &opts);\n+\t\thit = opts.hit;\n \t} else {\n \t\tif (cached)\n \t\t\tdie(_(\"both --cached and trees are given.\"));\n-- \n1.8.4.2\n"},{"id":"231173","messageId":"1385553659-9928-10-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 09/24] ls-files.c: use index api","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:44Z","receivedAt":"2013-11-27T12:00:44Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n builtin/ls-files.c | 38 +++++++++++++++++++++++++++++++++++---\n 1 file changed, 35 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/ls-files.c b/builtin/ls-files.c\nindex e1cf6d8..22fb012 100644\n--- a/builtin/ls-files.c\n+++ b/builtin/ls-files.c\n@@ -290,6 +290,22 @@ static void prune_cache(const char *prefix)\n \tactive_nr = last;\n }\n \n+static int needs_trailing_slash_stripped(void)\n+{\n+\tint i;\n+\n+\tif (!pathspec.nr)\n+\t\treturn 0;\n+\n+\tfor (i = 0; i < pathspec.nr; i++) {\n+\t\tint len = strlen(pathspec.items[i].original);\n+\n+\t\tif (len > 1 && (pathspec.items[i].original)[len - 1] == '/')\n+\t\t\treturn 1;\n+\t}\n+\treturn 0;\n+}\n+\n /*\n  * Read the tree specified with --with-tree option\n  * (typically, HEAD) into stage #1 and then\n@@ -447,6 +463,7 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n \tstruct dir_struct dir;\n \tstruct exclude_list *el;\n \tstruct string_list exclude_list = STRING_LIST_INIT_NODUP;\n+\tstruct filter_opts *opts = xmalloc(sizeof(*opts));\n \tstruct option builtin_ls_files_options[] = {\n \t\t{ OPTION_CALLBACK, 'z', NULL, NULL, NULL,\n \t\t\tN_(\"paths are separated with NUL character\"),\n@@ -512,9 +529,6 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n \t\tprefix_len = strlen(prefix);\n \tgit_config(git_default_config, NULL);\n \n-\tif (read_cache() < 0)\n-\t\tdie(\"index file corrupt\");\n-\n \targc = parse_options(argc, argv, prefix, builtin_ls_files_options,\n \t\t\tls_files_usage, 0);\n \tel = add_exclude_list(&dir, EXC_CMDL, \"--exclude option\");\n@@ -550,6 +564,24 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n \t\t       PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n \t\t       prefix, argv);\n \n+\tif (!with_tree && !needs_trailing_slash_stripped()) {\n+\t\tmemset(opts, 0, sizeof(*opts));\n+\t\topts->pathspec = &pathspec;\n+\t\topts->read_staged = 1;\n+\t\tif (show_resolve_undo)\n+\t\t\topts->read_resolve_undo = 1;\n+\t\tif (read_cache_filtered(opts) < 0)\n+\t\t\tdie(\"index file corrupt\");\n+\t} else {\n+\t\tif (read_cache() < 0)\n+\t\t\tdie(\"index file corrupt\");\n+\t\tparse_pathspec(&pathspec, 0,\n+\t\t\t       PATHSPEC_PREFER_CWD |\n+\t\t\t       PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n+\t\t\t       prefix, argv);\n+\n+\t}\n+\n \t/* Find common prefix for all pathspec's */\n \tmax_prefix = common_prefix(&pathspec);\n \tmax_prefix_len = max_prefix ? strlen(max_prefix) : 0;\n-- \n1.8.4.2\n"},{"id":"231174","messageId":"1385553659-9928-11-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 10/24] documentation: add documentation of the index-v5 file format","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:45Z","receivedAt":"2013-11-27T12:00:45Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Add a documentation of the index file format version 5 to\nDocumentation/technical.\n\nHelped-by: Michael Haggerty <mhagger@alum.mit.edu>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nHelped-by: Thomas Rast <trast@student.ethz.ch>\nHelped-by: Nguyen Thai Ngoc Duy <pclouds@gmail.com>\nHelped-by: Robin Rosenberg <robin.rosenberg@dewire.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n Documentation/technical/index-file-format-v5.txt | 294 +++++++++++++++++++++++\n 1 file changed, 294 insertions(+)\n create mode 100644 Documentation/technical/index-file-format-v5.txt\n\ndiff --git a/Documentation/technical/index-file-format-v5.txt b/Documentation/technical/index-file-format-v5.txt\nnew file mode 100644\nindex 0000000..1bd8a23\n--- /dev/null\n+++ b/Documentation/technical/index-file-format-v5.txt\n@@ -0,0 +1,294 @@\n+GIT index format\n+================\n+\n+== The git index\n+\n+   The git index file (.git/index) documents the status of the files\n+     in the git staging area.\n+\n+   The staging area is used for preparing commits, merging, etc.\n+\n+== The git index file format\n+\n+   All binary numbers are in network byte order. Version 5 is described\n+     here. The index file consists of various sections. They appear in\n+     the following order in the file.\n+\n+   - header: the description of the index format, including it's signature,\n+     version and various other fields that are used internally.\n+\n+   - diroffsets (ndir entries of \"direcotry offset\"): A 4-byte offset\n+       relative to the beginning of the \"direntries block\" (see below)\n+       for each of the ndir directories in the index, sorted by pathname\n+       (of the directory it's pointing to). [1]\n+\n+   - direntries (ndir entries of \"directory offset\"): A directory entry\n+       for each of the ndir directories in the index, sorted by pathname\n+       (see below). [2]\n+\n+   - fileoffsets (nfile entries of \"file offset\"): A 4-byte offset\n+       relative to the beginning of the fileentries block (see below)\n+       for each of the nfile files in the index. [1]\n+\n+   - fileentries (nfile entries of \"file entry\"): A file entry for\n+       each of the nfile files in the index (see below).\n+\n+   - Extensions (Currently REUC, see below for details)\n+\n+     Extensions are identified by signature. Optional extensions can\n+     be ignored if GIT does not understand them.\n+\n+     GIT supports an arbitrary number of extension, but currently none\n+     is implemented. [3]\n+\n+     extsig (32-bits): extension signature. If the first byte is 'A'..'Z'\n+     the extension is optional and can be ignored.\n+\n+     extsize (32-bits): number of entries in the extension\n+\n+     extchecksum (32-bits): crc32 checksum of the extension signature\n+       and size.\n+\n+    - Extension data.\n+\n+== Header\n+   sig (32-bits): Signature:\n+     The signature is { 'D', 'I', 'R', 'C' } (stands for \"dircache\")\n+\n+   vnr (32-bits): Version number:\n+     The current supported versions are 2, 3, 4 and 5.\n+\n+   nfile (32-bits): number of file entries in the index.\n+\n+   ndir (32-bits): number of directories in the index.\n+\n+   fblockoffset (32-bits): offset to the file block, relative to the\n+     beginning of the file.\n+\n+   - Offset to the extensions.\n+\n+     nextensions (32-bits): number of extensions.\n+\n+     extoffset (32-bits): offset to the extension. (Possibly none, as\n+       many as indicated in the 4-byte number of extensions)\n+\n+   headercrc (32-bits): crc checksum including the header and the\n+     offsets to the extensions.\n+\n+\n+== Directory offsets (diroffsets)\n+\n+  diroffset (32-bits): offset to the directory relative to the\n+    beginning of the index file. There are ndir + 1 offsets in the\n+    diroffset table, the last is pointing to the end of the last\n+    direntry. With this last entry, we are able to replace the strlen\n+    of the directory name when reading the directory name, by\n+    calculating it from diroffset[n+1]-diroffset[n]-61.  61 is the\n+    size of the directory data, which follows each each directory +\n+    the crc sum + the NUL byte.\n+\n+  This part is needed for making the directory entries bisectable and\n+    thus allowing a binary search.\n+\n+== Directory entry (direntries)\n+\n+  Directory entries are sorted in lexicographic order by the name\n+    of their path starting with the root.\n+\n+  foffset (32-bits): offset to the lexicographically first file in\n+    the file offsets (fileoffsets), relative to the beginning of\n+    the fileoffset block.\n+\n+  cr (32-bits): offset to conflicted/resolved data at the end of the\n+    index. 0 if there is no such data. [4]\n+\n+  ncr (32-bits): number of conflicted/resolved data entries at the\n+    end of the index if the offset is non 0. If cr is 0, ncr is\n+    also 0.\n+\n+  nsubtrees (32-bits): number of subtrees this tree has in the index.\n+\n+  nfiles (32-bits): number of files in the directory, that are in\n+    the index.\n+\n+  nentries (32-bits): number of entries in the index that is covered\n+    by the tree this entry represents. (-1 if the entry is invalid).\n+    This number includes all the files in this tree, recursively.\n+\n+  objname (160-bits): object name for the object that would result\n+    from writing this span of index as a tree. This is only valid\n+    if nentries is valid, meaning the cache-tree is valid.\n+\n+  flags (32-bits): 'flags' field split into (high to low bits) (For\n+    D/F conflicts)\n+\n+    stage (2-bits): stage of the directory during merge\n+\n+    30-bit unused\n+\n+  pathname (variable length, nul terminated): relative to top level\n+    directory (without the leading slash). '/' is used as path\n+    separator. A string of length 0 ('') indicates the root directory.\n+    The special path components \".\", and \"..\" (without quotes) are\n+    disallowed. The path also includes a trailing slash. [9]\n+\n+  dircrc (32-bits): crc32 checksum for each directory entry.\n+\n+  The last 4-byte number of entries and the 160-bit object name are\n+    for the cache tree. An entry can be in an invalidated state which is\n+    represented by having -1 in the entry_count field.\n+\n+  The entries are written out in the top-down, depth-first order. The\n+    first entry represents the root level of the repository, followed by\n+    the first subtree - let's call it A - of the root level, followed by\n+    the first subtree of A, ... There is no prefix compression for\n+    directories.\n+\n+== File offsets (fileoffsets)\n+\n+  fileoffset (32-bits): offset to the file relative to the beginning of\n+    the fileentries block.\n+\n+  This part is needed for making the file entries bisectable and\n+    thus allowing a binary search. There are nfile + 1 offsets in the\n+    fileoffset table, the last is pointing to the end of the last\n+    fileentry. With this last entry, we can replace the strlen when\n+    reading each filename, by calculating its length with the offsets.\n+\n+== File entry (fileentries)\n+\n+  File entries are sorted in ascending order on the name field, after the\n+  respective offset given by the directory entries. All file names are\n+  prefix compressed, meaning the file name is relative to the directory.\n+\n+  flags (16-bits): 'flags' field split into (high to low bits)\n+\n+    assumevalid (1-bit): assume-valid flag\n+\n+    intenttoadd (1-bit): intent-to-add flag, used by \"git add -N\".\n+      Extended flag in index v3.\n+\n+    stage (2-bit): stage of the file during merge\n+\n+    skipworktree (1-bit): skip-worktree flag, used by sparse checkout.\n+      Extended flag in index v3.\n+\n+    smudged (1-bit): indicates if the file is racily smudged.\n+\n+    invalid (1-bit): This bit can be set to indicate that a file was\n+      deleted, but not yet removed from the index, because the index\n+      was only partially rewritten.  Entries with this flags should be\n+      ignored when reading the index file.\n+\n+    9-bit unused, must be zero [6]\n+\n+  mode (16-bits): file mode, split into (high to low bits)\n+\n+    objtype (4-bits): object type\n+      valid values in binary are 1000 (regular file), 1010 (symbolic\n+      link) and 1110 (gitlink)\n+\n+    3-bit unused\n+\n+    permission (9-bits): unix permission. Only 0755 and 0644 are valid\n+      for regular files. Symbolic links and gitlinks have value 0 in\n+      this field.\n+\n+  mtimes (32-bits): mtime seconds, the last time a file's data changed\n+    this is stat(2) data\n+\n+  mtimens (32-bits): mtime nanosecond fractions\n+    this is stat(2) data\n+\n+  file size (32-bits): The on-disk size, trucated to 32-bit.\n+    this is stat(2) data\n+\n+  statcrc (32-bits): crc32 checksum over ctime seconds, ctime\n+    nanoseconds, ino, dev, uid, gid (All stat(2) data\n+    except mtime and file size). If the statcrc is 0 it will\n+    be ignored. [7]\n+\n+  objhash (160-bits): SHA-1 for the represented object\n+\n+  filename (variable length, nul terminated). The exact encoding is\n+    undefined, but the filename cannot contain a NUL byte (iow, the same\n+    encoding as a UNIX pathname).\n+\n+  entrycrc (32-bits): crc32 checksum for the file entry.\n+\n+== Resolve undo extension\n+\n+  Stores resolved entries and conflicts in the index.  When a conflict\n+  is resolved (e.g. with \"git add path), a bit is flipped to indicate\n+  the resolution.  This way conflicts can be recreated (e.g. with \"git\n+  checkout -m\", in case users want to redo a conflict resolution from\n+  scratch.\n+\n+  The conflicts will also be stored in the fileentries part of the index,\n+  to simplify reading and writing of the index.\n+\n+  filename (variable length, nul terminated): filename of the entry,\n+    relative to its containing directory).\n+\n+  nfileconflicts (32-bits): number of conflicts for the file [8]\n+\n+  flags (nfileconflicts entries of \"flags\") (16-bits): 'flags' field\n+    split into:\n+\n+    conflicted (1-bit): conflicted state (conflicted/resolved) (1 if\n+      conflicted)\n+\n+    stage (2-bits): stage during merge.\n+\n+    13-bit unused\n+\n+  entry_mode (nfileconflicts entries of \"entry mode\") (16-bits):\n+    octal numbers, entry mode of eache entry in the different stages.\n+    (How many is defined by the 4-byte number before)\n+\n+  objectnames (nfileconflicts entries of \"object name\") (160-bits):\n+    object names  of the different stages.\n+\n+  conflictcrc (32-bits): crc32 checksum over conflict data.\n+\n+== Design explanations\n+\n+[1] The directory and file offsets are included in the index format\n+    to enable bisectability of the index, for binary searches.Updating\n+    a single entry and partial reading will benefit from this.\n+\n+[2] The directories are saved in their own block, to be able to\n+    quickly search for a directory in the index. They include a\n+    offset to the (lexically) first file in the directory.\n+\n+[3] The data of the cache-tree extension and the resolve undo\n+    extension is now part of the index itself, but if other extensions\n+    come up in the future, there is no need to change the index, they\n+    can simply be added at the end.\n+\n+[4] To avoid rewrites of the whole index when there are conflicts or\n+    conflicts are being resolved, conflicted data will be stored at\n+    the end of the index. To mark the conflict resolved, just a bit\n+    has to be flipped. The data will still be there, if a user wants\n+    to redo the conflict resolution.\n+\n+[5] Since only 4 modes are effectively allowed in git but 32-bit are\n+    used to store them, having a two bit flag for the mode is enough\n+    and saves 4 byte per entry.\n+\n+[6] The length of the file name was dropped, since each file name is\n+    nul terminated anyway.\n+\n+[7] Since all stat data (except mtime and ctime) is just used for\n+    checking if a file has changed a checksum of the data is enough.\n+    In addition to that Thomas Rast suggested ctime could be ditched\n+    completely (core.trustctime=false) and thus included in the\n+    checksum. This would save 24 bytes per index entry, which would\n+    be about 4 MB on the Webkit index.\n+    (Thanks for the suggestion to Michael Haggerty)\n+\n+[8] Since there can be more stage #1 entries, it is necessary to know\n+    the number of conflict data entries there are.\n+\n+[9] As Michael Haggerty pointed out on the mailing list, storing the\n+    trailing slash will simplify a few operations.\n-- \n1.8.4.2\n"},{"id":"231175","messageId":"1385553659-9928-12-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 11/24] read-cache: make in-memory format aware of stat_crc","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:46Z","receivedAt":"2013-11-27T12:00:46Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Make the in-memory format aware of the stat_crc used by index-v5.\nIt is simply ignored by index version prior to v5.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n cache.h      |  1 +\n read-cache.c | 25 +++++++++++++++++++++++++\n 2 files changed, 26 insertions(+)\n\ndiff --git a/cache.h b/cache.h\nindex 38d57e7..8c2ccc4 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -127,6 +127,7 @@ struct cache_entry {\n \tunsigned int ce_flags;\n \tunsigned int ce_namelen;\n \tunsigned char sha1[20];\n+\tuint32_t ce_stat_crc;\n \tstruct cache_entry *next; /* used by name_hash */\n \tchar name[FLEX_ARRAY]; /* more */\n };\ndiff --git a/read-cache.c b/read-cache.c\nindex 7020f26..baa052c 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -106,6 +106,29 @@ int match_stat_data(const struct stat_data *sd, struct stat *st)\n \treturn changed;\n }\n \n+static uint32_t calculate_stat_crc(struct cache_entry *ce)\n+{\n+\tunsigned int ctimens = 0;\n+\tuint32_t stat, stat_crc;\n+\n+\tstat = htonl(ce->ce_stat_data.sd_ctime.sec);\n+\tstat_crc = crc32(0, (Bytef*)&stat, 4);\n+#ifdef USE_NSEC\n+\tctimens = ce->ce_stat_data.sd_ctime.nsec;\n+#endif\n+\tstat = htonl(ctimens);\n+\tstat_crc = crc32(stat_crc, (Bytef*)&stat, 4);\n+\tstat = htonl(ce->ce_stat_data.sd_ino);\n+\tstat_crc = crc32(stat_crc, (Bytef*)&stat, 4);\n+\tstat = htonl(ce->ce_stat_data.sd_dev);\n+\tstat_crc = crc32(stat_crc, (Bytef*)&stat, 4);\n+\tstat = htonl(ce->ce_stat_data.sd_uid);\n+\tstat_crc = crc32(stat_crc, (Bytef*)&stat, 4);\n+\tstat = htonl(ce->ce_stat_data.sd_gid);\n+\tstat_crc = crc32(stat_crc, (Bytef*)&stat, 4);\n+\treturn stat_crc;\n+}\n+\n /*\n  * This only updates the \"non-critical\" parts of the directory\n  * cache, ie the parts that aren't tracked by GIT, and only used\n@@ -120,6 +143,8 @@ void fill_stat_cache_info(struct cache_entry *ce, struct stat *st)\n \n \tif (S_ISREG(st->st_mode))\n \t\tce_mark_uptodate(ce);\n+\n+\tce->ce_stat_crc = calculate_stat_crc(ce);\n }\n \n static int ce_compare_data(const struct cache_entry *ce, struct stat *st)\n-- \n1.8.4.2\n"},{"id":"231176","messageId":"1385553659-9928-13-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 12/24] read-cache: read index-v5","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:47Z","receivedAt":"2013-11-27T12:00:47Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Make git read the index file version 5 without complaining.\n\nThis version of the reader reads neither the cache-tree\nnor the resolve undo data, however, it won't choke on an\nindex that includes such data.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nHelped-by: Nguyen Thai Ngoc Duy <pclouds@gmail.com>\nHelped-by: Thomas Rast <trast@student.ethz.ch>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n Makefile        |   1 +\n cache.h         |  32 ++++-\n read-cache-v5.c | 417 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n read-cache.h    |   1 +\n 4 files changed, 450 insertions(+), 1 deletion(-)\n create mode 100644 read-cache-v5.c\n\ndiff --git a/Makefile b/Makefile\nindex 5c28777..6a1b054 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -851,6 +851,7 @@ LIB_OBJS += quote.o\n LIB_OBJS += reachable.o\n LIB_OBJS += read-cache.o\n LIB_OBJS += read-cache-v2.o\n+LIB_OBJS += read-cache-v5.o\n LIB_OBJS += reflog-walk.o\n LIB_OBJS += refs.o\n LIB_OBJS += remote.o\ndiff --git a/cache.h b/cache.h\nindex 8c2ccc4..65171e4 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -99,7 +99,7 @@ unsigned long git_deflate_bound(git_zstream *, unsigned long);\n #define CACHE_SIGNATURE 0x44495243\t/* \"DIRC\" */\n \n #define INDEX_FORMAT_LB 2\n-#define INDEX_FORMAT_UB 4\n+#define INDEX_FORMAT_UB 5\n \n /*\n  * The \"cache_time\" is just the low 32 bits of the\n@@ -121,6 +121,15 @@ struct stat_data {\n \tunsigned int sd_size;\n };\n \n+/*\n+ * The *next_ce pointer is used in read_entries_v5 for holding\n+ * all the elements of a directory, and points to the next\n+ * cache_entry in a directory.\n+ *\n+ * It is reset by the add_name_hash call in set_index_entry\n+ * to set it to point to the next cache_entry in the\n+ * correct in-memory format ordering.\n+ */\n struct cache_entry {\n \tstruct stat_data ce_stat_data;\n \tunsigned int ce_mode;\n@@ -132,11 +141,17 @@ struct cache_entry {\n \tchar name[FLEX_ARRAY]; /* more */\n };\n \n+#define CE_NAMEMASK  (0x0fff)\n #define CE_STAGEMASK (0x3000)\n #define CE_EXTENDED  (0x4000)\n #define CE_VALID     (0x8000)\n+#define CE_SMUDGED   (0x0400) /* index v5 only flag */\n #define CE_STAGESHIFT 12\n \n+#define CONFLICT_CONFLICTED (0x8000)\n+#define CONFLICT_STAGESHIFT 13\n+#define CONFLICT_STAGEMASK (0x6000)\n+\n /*\n  * Range 0xFFFF0000 in ce_flags is divided into\n  * two parts: in-memory flags and on-disk ones.\n@@ -173,6 +188,19 @@ struct cache_entry {\n #define CE_EXTENDED_FLAGS (CE_INTENT_TO_ADD | CE_SKIP_WORKTREE)\n \n /*\n+ * Representation of the extended on-disk flags in the v5 format.\n+ * They must not collide with the ordinary on-disk flags, and need to\n+ * fit in 16 bits.  Note however that v5 does not save the name\n+ * length.\n+ */\n+#define CE_INTENT_TO_ADD_V5  (0x4000)\n+#define CE_SKIP_WORKTREE_V5  (0x0800)\n+#define CE_INVALID_V5        (0x0200)\n+#if (CE_VALID|CE_STAGEMASK) & (CE_INTENTTOADD_V5|CE_SKIPWORKTREE_V5|CE_INVALID_V5)\n+#error \"v5 on-disk flags collide with ordinary on-disk flags\"\n+#endif\n+\n+/*\n  * Safeguard to avoid saving wrong flags:\n  *  - CE_EXTENDED2 won't get saved until its semantic is known\n  *  - Bits in 0x0000FFFF have been saved in ce_flags already\n@@ -213,6 +241,8 @@ static inline unsigned create_ce_flags(unsigned stage)\n #define ce_skip_worktree(ce) ((ce)->ce_flags & CE_SKIP_WORKTREE)\n #define ce_mark_uptodate(ce) ((ce)->ce_flags |= CE_UPTODATE)\n \n+#define conflict_stage(c) ((CONFLICT_STAGEMASK & (c)->flags) >> CONFLICT_STAGESHIFT)\n+\n #define ce_permissions(mode) (((mode) & 0100) ? 0755 : 0644)\n static inline unsigned int create_ce_mode(unsigned int mode)\n {\ndiff --git a/read-cache-v5.c b/read-cache-v5.c\nnew file mode 100644\nindex 0000000..9d8c8f0\n--- /dev/null\n+++ b/read-cache-v5.c\n@@ -0,0 +1,417 @@\n+#include \"cache.h\"\n+#include \"read-cache.h\"\n+#include \"resolve-undo.h\"\n+#include \"cache-tree.h\"\n+#include \"dir.h\"\n+#include \"pathspec.h\"\n+\n+#define ptr_add(x,y) ((void *)(((char *)(x)) + (y)))\n+\n+struct cache_header_v5 {\n+\tuint32_t hdr_ndir;\n+\tuint32_t hdr_fblockoffset;\n+\tuint32_t hdr_nextension;\n+};\n+\n+struct directory_entry {\n+\tstruct directory_entry **sub;\n+\tstruct directory_entry *next;\n+\tstruct directory_entry *next_hash;\n+\tstruct cache_entry *ce;\n+\tstruct cache_entry *ce_last;\n+\tuint32_t conflict_size;\n+\tuint32_t de_foffset;\n+\tuint32_t de_nsubtrees;\n+\tuint32_t de_nfiles;\n+\tuint32_t de_nentries;\n+\tunsigned char sha1[20];\n+\tuint16_t de_flags;\n+\tuint32_t de_pathlen;\n+\tchar pathname[FLEX_ARRAY];\n+};\n+\n+struct conflict_part {\n+\tstruct conflict_part *next;\n+\tuint16_t flags;\n+\tuint16_t entry_mode;\n+\tunsigned char sha1[20];\n+};\n+\n+struct conflict_entry {\n+\tstruct conflict_entry *next;\n+\tuint32_t nfileconflicts;\n+\tstruct conflict_part *entries;\n+\tuint32_t namelen;\n+\tuint32_t pathlen;\n+\tchar name[FLEX_ARRAY];\n+};\n+\n+/*****************************************************************\n+ * Index File I/O\n+ *****************************************************************/\n+\n+struct ondisk_cache_entry {\n+\tuint16_t flags;\n+\tuint16_t mode;\n+\tstruct cache_time mtime;\n+\tuint32_t size;\n+\tuint32_t stat_crc;\n+\tunsigned char sha1[20];\n+\tchar name[FLEX_ARRAY];\n+};\n+\n+struct ondisk_directory_entry {\n+\tuint32_t foffset;\n+\tuint32_t nsubtrees;\n+\tuint32_t nfiles;\n+\tuint32_t nentries;\n+\tunsigned char sha1[20];\n+\tuint32_t flags;\n+\tchar name[FLEX_ARRAY];\n+};\n+#define directory_entry_size(len) (offsetof(struct directory_entry,pathname) + (len) + 1)\n+#define conflict_entry_size(len) (offsetof(struct conflict_entry,name) + (len) + 1)\n+\n+static int check_crc32(int initialcrc, void *data,\n+\t\t       size_t len, unsigned int expected_crc)\n+{\n+\tint crc;\n+\n+\tcrc = crc32(initialcrc, (Bytef*)data, len);\n+\treturn crc == expected_crc;\n+}\n+\n+static int match_stat_crc(struct stat *st, uint32_t expected_crc)\n+{\n+\tuint32_t data, stat_crc = 0;\n+\tunsigned int ctimens = 0;\n+\n+\tdata = htonl(st->st_ctime);\n+\tstat_crc = crc32(0, (Bytef*)&data, 4);\n+#ifdef USE_NSEC\n+\tctimens = ST_CTIME_NSEC(*st);\n+#endif\n+\tdata = htonl(ctimens);\n+\tstat_crc = crc32(stat_crc, (Bytef*)&data, 4);\n+\tdata = htonl(st->st_ino);\n+\tstat_crc = crc32(stat_crc, (Bytef*)&data, 4);\n+\tdata = htonl(st->st_dev);\n+\tstat_crc = crc32(stat_crc, (Bytef*)&data, 4);\n+\tdata = htonl(st->st_uid);\n+\tstat_crc = crc32(stat_crc, (Bytef*)&data, 4);\n+\tdata = htonl(st->st_gid);\n+\tstat_crc = crc32(stat_crc, (Bytef*)&data, 4);\n+\n+\treturn stat_crc == expected_crc;\n+}\n+\n+static int match_stat_basic(const struct cache_entry *ce,\n+\t\t\t    struct stat *st,\n+\t\t\t    int changed)\n+{\n+\n+\tif (ce->ce_stat_data.sd_mtime.sec != (unsigned int)st->st_mtime)\n+\t\tchanged |= MTIME_CHANGED;\n+#ifdef USE_NSEC\n+\tif (ce->ce_stat_data.sd_mtime.nsec != ST_MTIME_NSEC(*st))\n+\t\tchanged |= MTIME_CHANGED;\n+#endif\n+\tif (ce->ce_stat_data.sd_size != (unsigned int)st->st_size)\n+\t\tchanged |= DATA_CHANGED;\n+\n+\tif (trust_ctime && ce->ce_stat_crc != 0 && !match_stat_crc(st, ce->ce_stat_crc)) {\n+\t\tchanged |= OWNER_CHANGED;\n+\t\tchanged |= INODE_CHANGED;\n+\t}\n+\t/* Racily smudged entry? */\n+\tif (ce->ce_flags & CE_SMUDGED) {\n+\t\tif (!changed && !is_empty_blob_sha1(ce->sha1) && ce_modified_check_fs(ce, st))\n+\t\t\tchanged |= DATA_CHANGED;\n+\t}\n+\treturn changed;\n+}\n+\n+static int verify_hdr(void *mmap, unsigned long size)\n+{\n+\tuint32_t *filecrc;\n+\tunsigned int header_size;\n+\tstruct cache_header *hdr;\n+\tstruct cache_header_v5 *hdr_v5;\n+\n+\tif (size < sizeof(struct cache_header)\n+\t    + sizeof (struct cache_header_v5) + 4)\n+\t\tdie(\"index file smaller than expected\");\n+\n+\thdr = mmap;\n+\thdr_v5 = ptr_add(mmap, sizeof(*hdr));\n+\t/* Size of the header + the size of the extensionoffsets */\n+\theader_size = sizeof(*hdr) + sizeof(*hdr_v5) + hdr_v5->hdr_nextension * 4;\n+\t/* Initialize crc */\n+\tfilecrc = ptr_add(mmap, header_size);\n+\tif (!check_crc32(0, hdr, header_size, ntohl(*filecrc)))\n+\t\treturn error(\"bad index file header crc signature\");\n+\treturn 0;\n+}\n+\n+static struct cache_entry *cache_entry_from_ondisk(struct ondisk_cache_entry *ondisk,\n+\t\t\t\t\t\t   char *pathname, size_t len,\n+\t\t\t\t\t\t   size_t pathlen)\n+{\n+\tstruct cache_entry *ce = xmalloc(cache_entry_size(len + pathlen));\n+\tint flags;\n+\n+\tmemset(ce, 0, cache_entry_size(len + pathlen));\n+\tflags = ntoh_s(ondisk->flags);\n+\t/*\n+\t * This entry was invalidated in the index file,\n+\t * we don't need any data from it\n+\t */\n+\tif (flags & CE_INVALID_V5)\n+\t\treturn NULL;\n+\tce->ce_stat_data.sd_mtime.sec  = ntoh_l(ondisk->mtime.sec);\n+\tce->ce_stat_data.sd_mtime.nsec = ntoh_l(ondisk->mtime.nsec);\n+\tce->ce_stat_data.sd_size       = ntoh_l(ondisk->size);\n+\tce->ce_mode       = ntoh_s(ondisk->mode);\n+\tce->ce_flags      = flags & CE_STAGEMASK;\n+\tce->ce_flags     |= flags & CE_VALID;\n+\tce->ce_flags     |= flags & CE_SMUDGED;\n+\tif (flags & CE_INTENT_TO_ADD_V5)\n+\t\tce->ce_flags |= CE_INTENT_TO_ADD;\n+\tif (flags & CE_SKIP_WORKTREE_V5)\n+\t\tce->ce_flags |= CE_SKIP_WORKTREE;\n+\tce->ce_stat_crc   = ntoh_l(ondisk->stat_crc);\n+\tce->ce_namelen    = len + pathlen;\n+\thashcpy(ce->sha1, ondisk->sha1);\n+\tmemcpy(ce->name, pathname, pathlen);\n+\tmemcpy(ce->name + pathlen, ondisk->name, len);\n+\tce->name[len + pathlen] = '\\0';\n+\treturn ce;\n+}\n+\n+static struct directory_entry *init_directory_entry(const char *pathname, int len)\n+{\n+\tstruct directory_entry *de = xmalloc(directory_entry_size(len));\n+\n+\tmemset(de, 0, directory_entry_size(len));\n+\tmemcpy(de->pathname, pathname, len);\n+\tde->de_pathlen = len;\n+\treturn de;\n+}\n+\n+static struct directory_entry *directory_entry_from_ondisk(struct ondisk_directory_entry *ondisk,\n+\t\t\t\t\t\t\t   size_t len)\n+{\n+\tstruct directory_entry *de = init_directory_entry(ondisk->name, len);\n+\n+\tde->de_flags      = ntoh_s(ondisk->flags);\n+\tde->de_foffset    = ntoh_l(ondisk->foffset);\n+\tde->de_nsubtrees  = ntoh_l(ondisk->nsubtrees);\n+\tde->de_nfiles     = ntoh_l(ondisk->nfiles);\n+\tde->de_nentries   = ntoh_l(ondisk->nentries);\n+\tde->de_pathlen    = len;\n+\thashcpy(de->sha1, ondisk->sha1);\n+\treturn de;\n+}\n+\n+/*\n+ * Read the directories recursively into a directory tree.  dir_offset\n+ * is the current offset to the directory to be read in the direntries\n+ * block, while dir_table_offset is the current offset for the directory\n+ * in the diroffsets block.\n+ */\n+static struct directory_entry *read_directories(unsigned int *dir_offset,\n+\t\t\t\t\t\tunsigned int *dir_table_offset,\n+\t\t\t\t\t\tvoid *mmap, int mmap_size)\n+{\n+\tuint32_t *filecrc, *beginning, *end;\n+\tstruct ondisk_directory_entry *disk_de;\n+\tstruct directory_entry *de;\n+\tunsigned int data_len, len, i;\n+\n+\tbeginning = ptr_add(mmap, *dir_table_offset);\n+\tend = ptr_add(mmap, *dir_table_offset + 4);\n+\t/* Calculate the namelen from the offsets (-5 = NUL byte + crc checksum) */\n+\tlen = ntoh_l(*end) - ntoh_l(*beginning) -\n+\t\toffsetof(struct ondisk_directory_entry, name) - 5;\n+\tdisk_de = ptr_add(mmap, *dir_offset);\n+\tde = directory_entry_from_ondisk(disk_de, len);\n+\n+\tdata_len = len + 1 + offsetof(struct ondisk_directory_entry, name);\n+\tfilecrc = ptr_add(mmap, *dir_offset + data_len);\n+\tif (!check_crc32(0, ptr_add(mmap, *dir_offset), data_len, ntoh_l(*filecrc)))\n+\t\tdie(\"directory crc doesn't match for '%s'\", de->pathname);\n+\n+\t*dir_table_offset += 4;\n+\t*dir_offset += data_len + 4; /* crc code */\n+\n+\tde->sub = xcalloc(de->de_nsubtrees, sizeof(struct directory_entry *));\n+\tfor (i = 0; i < de->de_nsubtrees; i++) {\n+\t\tde->sub[i] = read_directories(dir_offset, dir_table_offset,\n+\t\t\t\t\t\t   mmap, mmap_size);\n+\t}\n+\n+\treturn de;\n+}\n+\n+static int read_entry(struct cache_entry **ce, char *pathname, size_t pathlen,\n+\t\t      void *mmap, unsigned long mmap_size,\n+\t\t      unsigned int first_entry_offset,\n+\t\t      unsigned int foffsetblock)\n+{\n+\tint len;\n+\tuint32_t *filecrc, *beginning, *end, entry_offset;\n+\tstruct ondisk_cache_entry *disk_ce;\n+\n+\tbeginning = ptr_add(mmap, foffsetblock);\n+\tend = ptr_add(mmap, foffsetblock + 4);\n+\t/* Calculate the namelen from the offsets (-5 = NUL byte + crc checksum) */\n+\tlen = ntoh_l(*end) - ntoh_l(*beginning) -\n+\t\toffsetof(struct ondisk_cache_entry, name) - 5;\n+\tentry_offset = first_entry_offset + ntoh_l(*beginning);\n+\tdisk_ce = ptr_add(mmap, entry_offset);\n+\t*ce = cache_entry_from_ondisk(disk_ce, pathname, len, pathlen);\n+\tfilecrc = ptr_add(mmap, entry_offset + len + 1 + sizeof(*disk_ce));\n+\tif (!check_crc32(0,\n+\t\tptr_add(mmap, entry_offset), len + 1 + sizeof(*disk_ce),\n+\t\tntoh_l(*filecrc)))\n+\t\treturn -1;\n+\n+\treturn 0;\n+}\n+\n+/*\n+ * Read all file entries from the index.  This function is recursive to get\n+ * the ordering right. In the index file the entries are sorted def, abc/def,\n+ * abc/xyz, while in-core they are sorted abc/def, abc/xyz, def.\n+ */\n+static int read_entries(struct index_state *istate, struct directory_entry *de,\n+\t\t\tunsigned int first_entry_offset, void *mmap,\n+\t\t\tunsigned long mmap_size, unsigned int *nr,\n+\t\t\tunsigned int foffsetblock)\n+{\n+\tstruct cache_entry *ce;\n+\tint i, subdir = 0;\n+\n+\tfor (i = 0; i < de->de_nfiles; i++) {\n+\t\tunsigned int subdir_foffsetblock = de->de_foffset + foffsetblock + (i * 4);\n+\t\tif (read_entry(&ce, de->pathname, de->de_pathlen, mmap, mmap_size,\n+\t\t\t       first_entry_offset, subdir_foffsetblock) < 0)\n+\t\t\treturn -1;\n+\t\twhile (subdir < de->de_nsubtrees &&\n+\t\t       cache_name_compare(ce->name + de->de_pathlen,\n+\t\t\t\t\t  ce_namelen(ce) - de->de_pathlen,\n+\t\t\t\t\t  de->sub[subdir]->pathname + de->de_pathlen,\n+\t\t\t\t\t  de->sub[subdir]->de_pathlen - de->de_pathlen) > 0) {\n+\t\t\tread_entries(istate, de->sub[subdir], first_entry_offset, mmap,\n+\t\t\t\t     mmap_size, nr, foffsetblock);\n+\t\t\tsubdir++;\n+\t\t}\n+\t\tif (!ce)\n+\t\t\tcontinue;\n+\t\tset_index_entry(istate, (*nr)++, ce);\n+\t}\n+\tfor (i = subdir; i < de->de_nsubtrees; i++) {\n+\t\tread_entries(istate, de->sub[i], first_entry_offset, mmap,\n+\t\t\t     mmap_size, nr, foffsetblock);\n+\t}\n+\treturn 0;\n+}\n+\n+static void free_directory_tree(struct directory_entry *de) {\n+\tint i;\n+\n+\tfor (i = 0; i < de->de_pathlen; i++)\n+\t\tfree_directory_tree(de->sub[i]);\n+\tfree(de);\n+}\n+\n+/*\n+ * Read an index-v5 file filtered by the filter_opts.   If opts is NULL,\n+ * everything will be read.\n+ */\n+static int read_index_v5(struct index_state *istate, void *mmap,\n+\t\t\t unsigned long mmap_size, struct filter_opts *opts)\n+{\n+\tunsigned int entry_offset, foffsetblock, nr = 0, *extoffsets;\n+\tunsigned int dir_offset, dir_table_offset;\n+\tint need_root = 0, i;\n+\tuint32_t *offset;\n+\tstruct directory_entry *root_directory, *de, *last_de;\n+\tconst char **paths = NULL;\n+\tstruct pathspec adjusted_pathspec;\n+\tstruct cache_header *hdr;\n+\tstruct cache_header_v5 *hdr_v5;\n+\n+\thdr = mmap;\n+\thdr_v5 = ptr_add(mmap, sizeof(*hdr));\n+\tistate->cache_alloc = alloc_nr(ntohl(hdr->hdr_entries));\n+\tistate->cache = xcalloc(istate->cache_alloc, sizeof(struct cache_entry *));\n+\textoffsets = xcalloc(ntohl(hdr_v5->hdr_nextension), sizeof(int));\n+\tfor (i = 0; i < ntohl(hdr_v5->hdr_nextension); i++) {\n+\t\toffset = ptr_add(mmap, sizeof(*hdr) + sizeof(*hdr_v5));\n+\t\textoffsets[i] = htonl(*offset);\n+\t}\n+\n+\t/* Skip size of the header + crc sum + size of offsets to extensions + size of offsets */\n+\tdir_offset = sizeof(*hdr) + sizeof(*hdr_v5) + ntohl(hdr_v5->hdr_nextension) * 4 + 4\n+\t\t+ (ntohl(hdr_v5->hdr_ndir) + 1) * 4;\n+\tdir_table_offset = sizeof(*hdr) + sizeof(*hdr_v5) + ntohl(hdr_v5->hdr_nextension) * 4 + 4;\n+\troot_directory = read_directories(&dir_offset, &dir_table_offset,\n+\t\t\t\t\t  mmap, mmap_size);\n+\n+\tentry_offset = ntohl(hdr_v5->hdr_fblockoffset);\n+\tfoffsetblock = dir_offset;\n+\n+\tif (opts && opts->pathspec && opts->pathspec->nr) {\n+\t\tpaths = xmalloc((opts->pathspec->nr + 1)*sizeof(char *));\n+\t\tpaths[opts->pathspec->nr] = NULL;\n+\t\tfor (i = 0; i < opts->pathspec->nr; i++) {\n+\t\t\tchar *super = strdup(opts->pathspec->items[i].match);\n+\t\t\tint len = strlen(super);\n+\t\t\twhile (len && super[len - 1] == '/' && super[len - 2] == '/')\n+\t\t\t\tsuper[--len] = '\\0'; /* strip all but one trailing slash */\n+\t\t\twhile (len && super[--len] != '/')\n+\t\t\t\t; /* scan backwards to next / */\n+\t\t\tif (len >= 0)\n+\t\t\t\tsuper[len--] = '\\0';\n+\t\t\tif (len <= 0) {\n+\t\t\t\tneed_root = 1;\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t\tpaths[i] = super;\n+\t\t}\n+\t}\n+\n+\tif (!need_root)\n+\t\tparse_pathspec(&adjusted_pathspec, PATHSPEC_ALL_MAGIC, PATHSPEC_PREFER_CWD, NULL, paths);\n+\n+\tde = root_directory;\n+\tlast_de = de;\n+\twhile (de) {\n+\t\tif (need_root ||\n+\t\t    match_pathspec_depth(&adjusted_pathspec, de->pathname, de->de_pathlen, 0, NULL)) {\n+\t\t\tif (read_entries(istate, de, entry_offset,\n+\t\t\t\t\t mmap, mmap_size, &nr,\n+\t\t\t\t\t foffsetblock) < 0)\n+\t\t\t\treturn -1;\n+\t\t} else {\n+\t\t\tlast_de = de;\n+\t\t\tfor (i = 0; i < de->de_nsubtrees; i++) {\n+\t\t\t\tde->sub[i]->next = last_de->next;\n+\t\t\t\tlast_de->next = de->sub[i];\n+\t\t\t\tlast_de = last_de->next;\n+\t\t\t}\n+\t\t}\n+\t\tde = de->next;\n+\t}\n+\tfree_directory_tree(root_directory);\n+\tistate->cache_nr = nr;\n+\treturn 0;\n+}\n+\n+struct index_ops v5_ops = {\n+\tmatch_stat_basic,\n+\tverify_hdr,\n+\tread_index_v5,\n+\tNULL\n+};\ndiff --git a/read-cache.h b/read-cache.h\nindex f920546..7823fbb 100644\n--- a/read-cache.h\n+++ b/read-cache.h\n@@ -34,6 +34,7 @@ struct index_ops {\n };\n \n extern struct index_ops v2_ops;\n+extern struct index_ops v5_ops;\n \n #ifndef NEEDS_ALIGNED_ACCESS\n #define ntoh_s(var) ntohs(var)\n-- \n1.8.4.2\n"},{"id":"231177","messageId":"1385553659-9928-14-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 13/24] read-cache: read resolve-undo data","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:48Z","receivedAt":"2013-11-27T12:00:48Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Make git read the resolve-undo data from the index.\n\nSince the resolve-undo data is joined with the conflicts in\nthe ondisk format of the index file version 5, conflicts and\nresolved data is read at the same time, and the resolve-undo\ndata are then converted to the in-memory format.\n\nHelped-by: Thomas Rast <trast@student.ethz.ch>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n read-cache-v5.c | 160 ++++++++++++++++++++++++++++++++++++++++++++++++++++++--\n 1 file changed, 157 insertions(+), 3 deletions(-)\n\ndiff --git a/read-cache-v5.c b/read-cache-v5.c\nindex 9d8c8f0..a9c687f 100644\n--- a/read-cache-v5.c\n+++ b/read-cache-v5.c\n@@ -1,5 +1,6 @@\n #include \"cache.h\"\n #include \"read-cache.h\"\n+#include \"string-list.h\"\n #include \"resolve-undo.h\"\n #include \"cache-tree.h\"\n #include \"dir.h\"\n@@ -13,13 +14,18 @@ struct cache_header_v5 {\n \tuint32_t hdr_nextension;\n };\n \n+struct extension_header {\n+\tchar signature[4];\n+\tuint32_t size;\n+\tuint32_t crc;\n+};\n+\n struct directory_entry {\n \tstruct directory_entry **sub;\n \tstruct directory_entry *next;\n \tstruct directory_entry *next_hash;\n \tstruct cache_entry *ce;\n \tstruct cache_entry *ce_last;\n-\tuint32_t conflict_size;\n \tuint32_t de_foffset;\n \tuint32_t de_nsubtrees;\n \tuint32_t de_nfiles;\n@@ -42,7 +48,6 @@ struct conflict_entry {\n \tuint32_t nfileconflicts;\n \tstruct conflict_part *entries;\n \tuint32_t namelen;\n-\tuint32_t pathlen;\n \tchar name[FLEX_ARRAY];\n };\n \n@@ -50,6 +55,12 @@ struct conflict_entry {\n  * Index File I/O\n  *****************************************************************/\n \n+struct ondisk_conflict_part {\n+\tuint16_t flags;\n+\tuint16_t entry_mode;\n+\tunsigned char sha1[20];\n+};\n+\n struct ondisk_cache_entry {\n \tuint16_t flags;\n \tuint16_t mode;\n@@ -145,7 +156,7 @@ static int verify_hdr(void *mmap, unsigned long size)\n \thdr = mmap;\n \thdr_v5 = ptr_add(mmap, sizeof(*hdr));\n \t/* Size of the header + the size of the extensionoffsets */\n-\theader_size = sizeof(*hdr) + sizeof(*hdr_v5) + hdr_v5->hdr_nextension * 4;\n+\theader_size = sizeof(*hdr) + sizeof(*hdr_v5) + ntohl(hdr_v5->hdr_nextension) * 4;\n \t/* Initialize crc */\n \tfilecrc = ptr_add(mmap, header_size);\n \tif (!check_crc32(0, hdr, header_size, ntohl(*filecrc)))\n@@ -279,6 +290,134 @@ static int read_entry(struct cache_entry **ce, char *pathname, size_t pathlen,\n \treturn 0;\n }\n \n+static struct conflict_part *conflict_part_from_ondisk(struct ondisk_conflict_part *ondisk)\n+{\n+\tstruct conflict_part *cp = xmalloc(sizeof(struct conflict_part));\n+\n+\tcp->flags      = ntoh_s(ondisk->flags);\n+\tcp->entry_mode = ntoh_s(ondisk->entry_mode);\n+\thashcpy(cp->sha1, ondisk->sha1);\n+\treturn cp;\n+}\n+\n+struct conflict_entry *create_new_conflict(char *name, int len)\n+{\n+\tstruct conflict_entry *conflict_entry;\n+\n+\tconflict_entry = xmalloc(conflict_entry_size(len));\n+\tmemset(conflict_entry, 0, conflict_entry_size(len));\n+\tconflict_entry->namelen = len;\n+\tmemcpy(conflict_entry->name, name, len);\n+\n+\treturn conflict_entry;\n+}\n+\n+static void add_part_to_conflict_entry(struct conflict_entry *entry,\n+\t\t\t\t       struct conflict_part *conflict_part)\n+{\n+\n+\tstruct conflict_part *conflict_search;\n+\n+\tentry->nfileconflicts++;\n+\tif (!entry->entries)\n+\t\tentry->entries = conflict_part;\n+\telse {\n+\t\tconflict_search = entry->entries;\n+\t\twhile (conflict_search->next)\n+\t\t\tconflict_search = conflict_search->next;\n+\t\tconflict_search->next = conflict_part;\n+\t}\n+}\n+\n+/*\n+ * Read the resolve undo data on disk and convert it to the internal\n+ * resolve undo format.\n+ */\n+static int read_resolve_undo(struct index_state *istate,\n+\t\t\t     unsigned int offset, void *mmap,\n+\t\t\t     unsigned int entries)\n+{\n+\tint i, k;\n+\n+\tfor (i = 0; i < entries; i++) {\n+\t\tchar *name;\n+\t\tunsigned int len, *nfileconflicts, nc;\n+\t\tuint32_t *crc;\n+\t\tstruct ondisk_conflict_part *ondisk;\n+\t\tstruct conflict_part *cp;\n+\t\tstruct string_list_item *lost;\n+\t\tstruct resolve_undo_info *ui;\n+\n+\t\tname = ptr_add(mmap, offset);\n+\t\tlen = strlen(name);\n+\t\toffset += len + 1;\n+\t\tnfileconflicts = ptr_add(mmap, offset);\n+\t\tnc = ntoh_l(*nfileconflicts);\n+\t\toffset += 4;\n+\n+\t\tcrc = ptr_add(mmap, offset +\n+\t\t\t      nc * sizeof(struct ondisk_conflict_part));\n+\t\tif (!check_crc32(0, name, len + 1 + 4 +\n+\t\t\t\t nc * sizeof(struct ondisk_conflict_part),\n+\t\t\t\t ntoh_l(*crc)))\n+\t\t\treturn -1;\n+\n+\t\tondisk = ptr_add(mmap, offset);\n+\t\tcp = conflict_part_from_ondisk(ondisk);\n+\t\tif (cp->flags & CONFLICT_CONFLICTED) {\n+\t\t\toffset += nc * sizeof(struct ondisk_conflict_part) + 4;\n+\t\t\tcontinue;\n+\t\t}\n+\t\toffset += sizeof(struct ondisk_conflict_part);\n+\t\tif (!istate->resolve_undo) {\n+\t\t\tistate->resolve_undo = xcalloc(1, sizeof(struct string_list));\n+\t\t\tistate->resolve_undo->strdup_strings = 1;\n+\t\t}\n+\n+\t\tlost = string_list_insert(istate->resolve_undo, name);\n+\t\tif (!lost->util)\n+\t\t\tlost->util = xcalloc(1, sizeof(*ui));\n+\t\tui = lost->util;\n+\t\tfor (k = 0; k < 3; k++)\n+\t\t\tui->mode[k] = 0;\n+\n+\t\tui->mode[conflict_stage(cp) - 1] = cp->entry_mode;\n+\t\thashcpy(ui->sha1[conflict_stage(cp) - 1], cp->sha1);\n+\t\tfor (k = 1; k < nc; k++) {\n+\t\t\tstruct conflict_part *cp;\n+\n+\t\t\tondisk = ptr_add(mmap, offset);\n+\t\t\tcp = conflict_part_from_ondisk(ondisk);\n+\t\t\tui->mode[conflict_stage(cp) - 1] = cp->entry_mode;\n+\t\t\thashcpy(ui->sha1[conflict_stage(cp) - 1], cp->sha1);\n+\t\t\toffset += sizeof(struct ondisk_conflict_part);\n+\t\t}\n+\t\toffset += 4; /* crc */\n+\t}\n+\treturn 0;\n+}\n+\n+static int read_index_extension(struct index_state *istate,\n+\t\t\t\tvoid *mmap, unsigned int extoffset)\n+{\n+\tstruct extension_header *ehdr;\n+\n+\tehdr = ptr_add(mmap, extoffset);\n+\t/* -4 for the crc that's included in the struct */\n+\tif (!check_crc32(0, ptr_add(mmap, extoffset),\n+\t\t\t sizeof(*ehdr) - 4, ntoh_l(ehdr->crc)))\n+\t\treturn -1;\n+\n+\tswitch (CACHE_EXT(ehdr->signature)) {\n+\tcase CACHE_EXT_RESOLVE_UNDO:\n+\t\tif (read_resolve_undo(istate, extoffset + sizeof(*ehdr),\n+\t\t\t\t      mmap, ntoh_l(ehdr->size)) < 0)\n+\t\t\treturn -1;\n+\t\tbreak;\n+\t}\n+\treturn 0;\n+}\n+\n /*\n  * Read all file entries from the index.  This function is recursive to get\n  * the ordering right. In the index file the entries are sorted def, abc/def,\n@@ -404,6 +543,21 @@ static int read_index_v5(struct index_state *istate, void *mmap,\n \t\t}\n \t\tde = de->next;\n \t}\n+\n+\tif (!opts || opts->read_resolve_undo) {\n+\t\tfor (i = 0; i < ntohl(hdr_v5->hdr_nextension); i++) {\n+\t\t\t/*\n+\t\t\t * After the index entry there is a number of\n+\t\t\t * extensions, which is written in the header.\n+\t\t\t * The extensions are prefixed by extension name\n+\t\t\t * (4-byte) and length of the extension (4-byte,\n+\t\t\t * usually the number of entries in that section)\n+\t\t\t * in network byte order\n+\t\t\t */\n+\t\t\tif (read_index_extension(istate, mmap, extoffsets[i]) < 0)\n+\t\t\t\treturn -1;\n+\t\t}\n+\t}\n \tfree_directory_tree(root_directory);\n \tistate->cache_nr = nr;\n \treturn 0;\n-- \n1.8.4.2\n"},{"id":"231178","messageId":"1385553659-9928-15-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 14/24] read-cache: read cache-tree in index-v5","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:49Z","receivedAt":"2013-11-27T12:00:49Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Since the cache-tree data is saved as part of the directory data,\nwe already read it at the beginning of the index. The cache-tree\nis only converted from this directory data.\n\nThe cache-tree data is arranged in a tree, with the children sorted by\npathlen at each node, while the ondisk format is sorted lexically.\nSo we have to rebuild this format from the on-disk directory list.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n cache-tree.c    |  2 +-\n cache-tree.h    |  1 +\n read-cache-v5.c | 68 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 70 insertions(+), 1 deletion(-)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 0bbec43..1209732 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -31,7 +31,7 @@ void cache_tree_free(struct cache_tree **it_p)\n \t*it_p = NULL;\n }\n \n-static int subtree_name_cmp(const char *one, int onelen,\n+int subtree_name_cmp(const char *one, int onelen,\n \t\t\t    const char *two, int twolen)\n {\n \tif (onelen < twolen)\ndiff --git a/cache-tree.h b/cache-tree.h\nindex f1923ad..9818926 100644\n--- a/cache-tree.h\n+++ b/cache-tree.h\n@@ -25,6 +25,7 @@ struct cache_tree *cache_tree(void);\n void cache_tree_free(struct cache_tree **);\n void cache_tree_invalidate_path(struct cache_tree *, const char *);\n struct cache_tree_sub *cache_tree_sub(struct cache_tree *, const char *);\n+int subtree_name_cmp(const char *, int, const char *, int);\n \n void cache_tree_write(struct strbuf *, struct cache_tree *root);\n struct cache_tree *cache_tree_read(const char *buffer, unsigned long size);\ndiff --git a/read-cache-v5.c b/read-cache-v5.c\nindex a9c687f..01f1c88 100644\n--- a/read-cache-v5.c\n+++ b/read-cache-v5.c\n@@ -418,6 +418,73 @@ static int read_index_extension(struct index_state *istate,\n \treturn 0;\n }\n \n+static int compare_cache_tree(const void *a, const void *b)\n+{\n+\tconst struct cache_tree_sub *it1, *it2;\n+\n+\tit1 = *(const struct cache_tree_sub **) a;\n+\tit2 = *(const struct cache_tree_sub **) b;\n+\treturn subtree_name_cmp(it1->name, it1->namelen,\n+\t\t\t\tit2->name, it2->namelen);\n+}\n+\n+/*\n+ * Convert the directory entries to cache-tree entries\n+ * recursively.\n+ */\n+static struct cache_tree *convert_one(struct directory_entry *de)\n+{\n+\tint i;\n+\tstruct cache_tree *it;\n+\n+\tit = cache_tree();\n+\tit->entry_count = de->de_nentries;\n+\tif (0 <= it->entry_count)\n+\t\thashcpy(it->sha1, de->sha1);\n+\n+\t/*\n+\t * Just a heuristic -- we do not add directories that often but\n+\t * we do not want to have to extend it immediately when we do,\n+\t * hence +2.\n+\t */\n+\tit->subtree_alloc = de->de_nsubtrees + 2;\n+\tit->down = xcalloc(it->subtree_alloc, sizeof(struct cache_tree_sub *));\n+\tfor (i = 0; i < de->de_nsubtrees; i++) {\n+\t\tstruct cache_tree *sub = convert_one(de->sub[i]);\n+\t\tstruct cache_tree_sub *subtree;\n+\t\t/* -1 for removing the / at the end of the pathname */\n+\t\tint namelen = de->sub[i]->de_pathlen - de->de_pathlen - 1;\n+\n+\t\tif (!sub)\n+\t\t\tgoto free_return;\n+\n+\t\tsubtree = xmalloc(sizeof(*subtree) + namelen + 1);\n+\t\tsubtree->cache_tree = sub;\n+\t\tsubtree->namelen = namelen;\n+\t\tmemcpy(subtree->name, de->sub[i]->pathname + de->de_pathlen, namelen);\n+\t\tsubtree->name[namelen] = '\\0';\n+\t\tit->down[i] = subtree;\n+\t\tit->subtree_nr++;\n+\t}\n+\tqsort(it->down, it->subtree_nr, sizeof(struct cache_tree_sub *),\n+\t      compare_cache_tree);\n+\treturn it;\n+free_return:\n+\tcache_tree_free(&it);\n+\treturn NULL;\n+}\n+\n+/*\n+ * This function modifies the directory argument that is given to it.\n+ * Don't use it if the directory entries are still needed after.\n+ */\n+static struct cache_tree *cache_tree_convert_v5(struct directory_entry *de)\n+{\n+\tif (!de->de_nentries)\n+\t\treturn NULL;\n+\treturn convert_one(de);\n+}\n+\n /*\n  * Read all file entries from the index.  This function is recursive to get\n  * the ordering right. In the index file the entries are sorted def, abc/def,\n@@ -558,6 +625,7 @@ static int read_index_v5(struct index_state *istate, void *mmap,\n \t\t\t\treturn -1;\n \t\t}\n \t}\n+\tistate->cache_tree = cache_tree_convert_v5(root_directory);\n \tfree_directory_tree(root_directory);\n \tistate->cache_nr = nr;\n \treturn 0;\n-- \n1.8.4.2\n"},{"id":"231188","messageId":"1385553659-9928-16-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 15/24] read-cache: write index-v5","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:50Z","receivedAt":"2013-11-27T12:00:50Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Write the index version 5 file format to disk. This version doesn't\nwrite the cache-tree data and resolve-undo data to the file.\n\nThe main work is done when filtering out the directories from the\ncurrent in-memory format, where in the same turn also the conflicts\nand the file data is calculated.\n\nHelped-by: Nguyen Thai Ngoc Duy <pclouds@gmail.com>\nHelped-by: Thomas Rast <trast@student.ethz.ch>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n cache.h         |   1 +\n read-cache-v5.c | 431 +++++++++++++++++++++++++++++++++++++++++++++++++++++++-\n read-cache.c    |   4 +-\n read-cache.h    |   1 +\n 4 files changed, 435 insertions(+), 2 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 65171e4..71b98cf 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -138,6 +138,7 @@ struct cache_entry {\n \tunsigned char sha1[20];\n \tuint32_t ce_stat_crc;\n \tstruct cache_entry *next; /* used by name_hash */\n+\tstruct cache_entry *next_ce;\n \tchar name[FLEX_ARRAY]; /* more */\n };\n \ndiff --git a/read-cache-v5.c b/read-cache-v5.c\nindex 01f1c88..797022f 100644\n--- a/read-cache-v5.c\n+++ b/read-cache-v5.c\n@@ -631,9 +631,438 @@ static int read_index_v5(struct index_state *istate, void *mmap,\n \treturn 0;\n }\n \n+#define WRITE_BUFFER_SIZE 8192\n+static unsigned char write_buffer[WRITE_BUFFER_SIZE];\n+static unsigned long write_buffer_len;\n+\n+static int ce_write_flush(int fd)\n+{\n+\tunsigned int buffered = write_buffer_len;\n+\tif (buffered) {\n+\t\tif (write_in_full(fd, write_buffer, buffered) != buffered)\n+\t\t\treturn -1;\n+\t\twrite_buffer_len = 0;\n+\t}\n+\treturn 0;\n+}\n+\n+static int ce_write(uint32_t *crc, int fd, void *data, unsigned int len)\n+{\n+\tif (crc)\n+\t\t*crc = crc32(*crc, (Bytef*)data, len);\n+\twhile (len) {\n+\t\tunsigned int buffered = write_buffer_len;\n+\t\tunsigned int partial = WRITE_BUFFER_SIZE - buffered;\n+\t\tif (partial > len)\n+\t\t\tpartial = len;\n+\t\tmemcpy(write_buffer + buffered, data, partial);\n+\t\tbuffered += partial;\n+\t\tif (buffered == WRITE_BUFFER_SIZE) {\n+\t\t\twrite_buffer_len = buffered;\n+\t\t\tif (ce_write_flush(fd))\n+\t\t\t\treturn -1;\n+\t\t\tbuffered = 0;\n+\t\t}\n+\t\twrite_buffer_len = buffered;\n+\t\tlen -= partial;\n+\t\tdata = (char *) data + partial;\n+\t}\n+\treturn 0;\n+}\n+\n+static int ce_flush(int fd)\n+{\n+\tunsigned int left = write_buffer_len;\n+\n+\tif (left)\n+\t\twrite_buffer_len = 0;\n+\n+\tif (write_in_full(fd, write_buffer, left) != left)\n+\t\treturn -1;\n+\n+\treturn 0;\n+}\n+\n+static void ce_smudge_racily_clean_entry(struct cache_entry *ce)\n+{\n+\t/*\n+\t * This method shall only be called if the timestamp of ce\n+\t * is racy (check with is_racy_timestamp). If the timestamp\n+\t * is racy, the writer will set the CE_SMUDGED flag.\n+\t *\n+\t * The reader (match_stat_basic) will then take care\n+\t * of checking if the entry is really changed or not, by\n+\t * taking into account the size and the stat_crc and if\n+\t * that hasn't changed checking the sha1.\n+\t */\n+\tce->ce_flags |= CE_SMUDGED;\n+}\n+\n+static char *super_directory(char *filename)\n+{\n+\tchar *super = dirname(filename);\n+\tif (!strcmp(super, \".\"))\n+\t\treturn NULL;\n+\treturn super;\n+}\n+\n+static void ondisk_from_directory_entry(struct directory_entry *de,\n+\t\t\t\t\tstruct ondisk_directory_entry *ondisk)\n+{\n+\tondisk->foffset   = htonl(de->de_foffset);\n+\tondisk->nsubtrees = htonl(de->de_nsubtrees);\n+\tondisk->nfiles    = htonl(de->de_nfiles);\n+\tondisk->nentries  = htonl(de->de_nentries);\n+\thashcpy(ondisk->sha1, de->sha1);\n+\tondisk->flags     = htons(de->de_flags);\n+\tif (de->de_pathlen == 0) {\n+\t\tmemcpy(ondisk->name, \"\\0\", 1);\n+\t} else {\n+\t\tmemcpy(ondisk->name, de->pathname, de->de_pathlen);\n+\t\tmemcpy(ondisk->name + de->de_pathlen, \"/\\0\", 2);\n+\t}\n+}\n+\n+static void insert_directory_entry(struct directory_entry *de,\n+\t\t\t\t   struct hash_table *table,\n+\t\t\t\t   unsigned int *total_dir_len,\n+\t\t\t\t   unsigned int *ndir,\n+\t\t\t\t   uint32_t crc)\n+{\n+\tstruct directory_entry *insert;\n+\n+\tinsert = (struct directory_entry *)insert_hash(crc, de, table);\n+\tif (insert) {\n+\t\tde->next_hash = insert->next_hash;\n+\t\tinsert->next_hash = de;\n+\t}\n+\t(*ndir)++;\n+\tif (de->de_pathlen == 0)\n+\t\t(*total_dir_len)++;\n+\telse\n+\t\t*total_dir_len += de->de_pathlen + 2;\n+}\n+\n+static struct directory_entry *find_directory(char *dir, int dir_len, uint32_t *crc,\n+\t\t\t\t\t      struct hash_table *table)\n+{\n+\tstruct directory_entry *search;\n+\n+\t*crc = crc32(0, (Bytef*)dir, dir_len);\n+\tsearch = lookup_hash(*crc, table);\n+\twhile (search &&\n+\t       cache_name_compare(dir, dir_len, search->pathname, search->de_pathlen))\n+\t\tsearch = search->next_hash;\n+\treturn search;\n+}\n+\n+static struct directory_entry *get_directory(char *dir, unsigned int dir_len,\n+\t\t\t\t\t     struct hash_table *table,\n+\t\t\t\t\t     unsigned int *total_dir_len,\n+\t\t\t\t\t     unsigned int *ndir,\n+\t\t\t\t\t     struct directory_entry **current)\n+{\n+\tstruct directory_entry *tmp = NULL, *search, *new, *ret;\n+\tuint32_t crc;\n+\n+\tsearch = find_directory(dir, dir_len, &crc, table);\n+\tif (search)\n+\t\treturn search;\n+\twhile (!search) {\n+\t\tnew = init_directory_entry(dir, dir_len);\n+\t\tinsert_directory_entry(new, table, total_dir_len, ndir, crc);\n+\t\tif (!tmp)\n+\t\t\tret = new;\n+\t\telse\n+\t\t\tnew->de_nsubtrees = 1;\n+\t\tnew->next = tmp;\n+\t\ttmp = new;\n+\t\tdir = super_directory(dir);\n+\t\tdir_len = dir ? strlen(dir) : 0;\n+\t\tsearch = find_directory(dir, dir_len, &crc, table);\n+\t}\n+\tsearch->de_nsubtrees++;\n+\t(*current)->next = tmp;\n+\twhile ((*current)->next)\n+\t\t*current = (*current)->next;\n+\n+\treturn ret;\n+}\n+\n+static void ce_queue_push(struct cache_entry **head,\n+\t\t\t  struct cache_entry **tail,\n+\t\t\t  struct cache_entry *ce)\n+{\n+\tif (!*head) {\n+\t\t*head = *tail = ce;\n+\t\t(*tail)->next_ce = NULL;\n+\t\treturn;\n+\t}\n+\n+\t(*tail)->next_ce = ce;\n+\tce->next_ce = NULL;\n+\t*tail = (*tail)->next_ce;\n+}\n+\n+static struct directory_entry *compile_directory_data(struct index_state *istate,\n+\t\t\t\t\t\t      int nfile, unsigned int *ndir,\n+\t\t\t\t\t\t      unsigned int *total_dir_len,\n+\t\t\t\t\t\t      unsigned int *total_file_len)\n+{\n+\tint i, dir_len = -1;\n+\tchar *dir;\n+\tstruct directory_entry *de, *current, *search;\n+\tstruct cache_entry **cache = istate->cache;\n+\tstruct hash_table table;\n+\tuint32_t crc;\n+\n+\tinit_hash(&table);\n+\tde = init_directory_entry(\"\", 0);\n+\tcurrent = de;\n+\t*ndir = 1;\n+\t*total_dir_len = 1;\n+\tcrc = crc32(0, (Bytef*)de->pathname, de->de_pathlen);\n+\tinsert_hash(crc, de, &table);\n+\tfor (i = 0; i < nfile; i++) {\n+\t\tif (cache[i]->ce_flags & CE_REMOVE)\n+\t\t\tcontinue;\n+\n+\t\tif (dir_len < 0\n+\t\t    || !(!(dir_len < ce_namelen(cache[i]) && cache[i]->name[dir_len] != '/')\n+\t\t\t && !strchr(cache[i]->name + dir_len + 1, '/')\n+\t\t\t && !cache_name_compare(cache[i]->name, ce_namelen(cache[i]),\n+\t\t\t\t\t\tdir, dir_len))) {\n+\t\t\tdir = super_directory(strdup(cache[i]->name));\n+\t\t\tdir_len = dir ? strlen(dir) : 0;\n+\t\t\tsearch = get_directory(dir, dir_len, &table,\n+\t\t\t\t\t       total_dir_len, ndir,\n+\t\t\t\t\t       &current);\n+\t\t}\n+\t\tsearch->de_nfiles++;\n+\t\t*total_file_len += ce_namelen(cache[i]) + 1;\n+\t\tif (search->de_pathlen)\n+\t\t\t*total_file_len -= search->de_pathlen + 1;\n+\t\tce_queue_push(&(search->ce), &(search->ce_last), cache[i]);\n+\t}\n+\treturn de;\n+}\n+\n+static void ondisk_from_cache_entry(struct cache_entry *ce,\n+\t\t\t\t    struct ondisk_cache_entry *ondisk,\n+\t\t\t\t    int pathlen)\n+{\n+\tunsigned int flags;\n+\n+\tflags  = ce->ce_flags & CE_STAGEMASK;\n+\tflags |= ce->ce_flags & CE_VALID;\n+\tflags |= ce->ce_flags & CE_SMUDGED;\n+\tif (ce->ce_flags & CE_INTENT_TO_ADD)\n+\t\tflags |= CE_INTENT_TO_ADD_V5;\n+\tif (ce->ce_flags & CE_SKIP_WORKTREE)\n+\t\tflags |= CE_SKIP_WORKTREE_V5;\n+\tondisk->flags      = htons(flags);\n+\tondisk->mode       = htons(ce->ce_mode);\n+\tondisk->mtime.sec  = htonl(ce->ce_stat_data.sd_mtime.sec);\n+#ifdef USE_NSEC\n+\tondisk->mtime.nsec = htonl(ce->ce_stat_data.sd_mtime.nsec);\n+#else\n+\tondisk->mtime.nsec = 0;\n+#endif\n+\tondisk->size       = htonl(ce->ce_stat_data.sd_size);\n+\tif (!ce->ce_stat_crc)\n+\t\tce->ce_stat_crc = calculate_stat_crc(ce);\n+\tondisk->stat_crc   = htonl(ce->ce_stat_crc);\n+\thashcpy(ondisk->sha1, ce->sha1);\n+\tmemcpy(ondisk->name, ce->name + pathlen, ce_namelen(ce) - pathlen);\n+\tondisk->name[ce_namelen(ce) - pathlen] = '\\0';\n+}\n+\n+static int write_directories(struct directory_entry *de, int fd)\n+{\n+\tstruct directory_entry *current;\n+\tstruct ondisk_directory_entry *ondisk;\n+\tint current_offset, offset_write, ondisk_size, foffset;\n+\tuint32_t crc;\n+\n+\tondisk_size = offsetof(struct ondisk_directory_entry, name);\n+\tcurrent = de;\n+\tcurrent_offset = 0;\n+\tfoffset = 0;\n+\t/* Write directory offsets */\n+\twhile (current) {\n+\t\tint pathlen;\n+\n+\t\toffset_write = htonl(current_offset);\n+\t\tif (ce_write(NULL, fd, &offset_write, 4) < 0)\n+\t\t\treturn -1;\n+\t\tif (current->de_pathlen == 0)\n+\t\t\tpathlen = 0;\n+\t\telse\n+\t\t\tpathlen = current->de_pathlen + 1;\n+\t\tcurrent_offset += pathlen + 1 + ondisk_size + 4;\n+\t\tcurrent = current->next;\n+\t}\n+\t/*\n+\t * Write one more offset, which points to the end of the entries,\n+\t * because we use it for calculating the dir length, instead of\n+\t * using strlen.\n+\t */\n+\toffset_write = htonl(current_offset);\n+\tif (ce_write(NULL, fd, &offset_write, 4) < 0)\n+\t\treturn -1;\n+\tcurrent = de;\n+\t/* Write directory entries */\n+\twhile (current) {\n+\t\tint size = ondisk_size + current->de_pathlen + 1;\n+\n+\t\tcrc = 0;\n+\t\tcurrent->de_foffset = foffset;\n+\t\tif (current->de_pathlen != 0)\n+\t\t\tsize++;\n+\t\tondisk = xmalloc(size);\n+\t\tondisk_from_directory_entry(current, ondisk);\n+\t\tif (ce_write(&crc, fd, ondisk, size) < 0)\n+\t\t\treturn -1;\n+\t\tcrc = htonl(crc);\n+\t\tif (ce_write(NULL, fd, &crc, 4) < 0)\n+\t\t\treturn -1;\n+\t\tfoffset += current->de_nfiles * 4;\n+\t\tfree(ondisk);\n+\t\tcurrent = current->next;\n+\t}\n+\treturn 0;\n+}\n+\n+static int write_entries(struct index_state *istate,\n+\t\t\t struct directory_entry *de,\n+\t\t\t int entries,\n+\t\t\t int fd)\n+{\n+\tint offset, offset_write, ondisk_size;\n+\tstruct directory_entry *current;\n+\n+\toffset = 0;\n+\tondisk_size = offsetof(struct ondisk_cache_entry, name);\n+\tcurrent = de;\n+\t/* Write cache entry offsets */\n+\twhile (current) {\n+\t\tint pathlen;\n+\t\tstruct cache_entry *ce = current->ce;\n+\n+\t\tpathlen = current->de_pathlen ? current->de_pathlen + 1 : 0;\n+\t\twhile (ce) {\n+\t\t\tif (!ce_uptodate(ce) && is_racy_timestamp(istate, ce))\n+\t\t\t\tce_smudge_racily_clean_entry(ce);\n+\t\t\tif (is_null_sha1(ce->sha1)) {\n+\t\t\t\tstatic const char msg[] = \"cache entry has null sha1: %s\";\n+\t\t\t\tstatic int allow = -1;\n+\n+\t\t\t\tif (allow < 0)\n+\t\t\t\t\tallow = git_env_bool(\"GIT_ALLOW_NULL_SHA1\", 0);\n+\t\t\t\tif (allow)\n+\t\t\t\t\twarning(msg, ce->name);\n+\t\t\t\telse\n+\t\t\t\t\treturn error(msg, ce->name);\n+\t\t\t}\n+\t\t\toffset_write = htonl(offset);\n+\t\t\tif (ce_write(NULL, fd, &offset_write, 4) < 0)\n+\t\t\t\treturn -1;\n+\t\t\toffset += ce_namelen(ce) - pathlen + 1 + ondisk_size + 4;\n+\t\t\tce = ce->next_ce;\n+\t\t}\n+\t\tcurrent = current->next;\n+\t}\n+\t/*\n+\t * Write one more offset, which points to the end of the entries,\n+\t * because we use it for calculating the file length, instead of\n+\t * using strlen.\n+\t */\n+\toffset_write = htonl(offset);\n+\tif (ce_write(NULL, fd, &offset_write, 4) < 0)\n+\t\treturn -1;\n+\n+\tcurrent = de;\n+\t/* Write cache entries */\n+\twhile (current) {\n+\t\tint pathlen;\n+\t\tstruct cache_entry *ce = current->ce;\n+\n+\t\tpathlen = current->de_pathlen ? current->de_pathlen + 1 : 0;\n+\t\twhile (ce) {\n+\t\t\tint size = offsetof(struct ondisk_cache_entry, name) +\n+\t\t\t\tce_namelen(ce) - pathlen + 1;\n+\t\t\tstruct ondisk_cache_entry *ondisk = xmalloc(size);\n+\t\t\tuint32_t crc;\n+\n+\t\t\tcrc = 0;\n+\t\t\tondisk_from_cache_entry(ce, ondisk, pathlen);\n+\t\t\tif (ce_write(&crc, fd, ondisk, size) < 0)\n+\t\t\t\treturn -1;\n+\t\t\tcrc = htonl(crc);\n+\t\t\tif (ce_write(NULL, fd, &crc, 4) < 0)\n+\t\t\t\treturn -1;\n+\t\t\toffset += 4;\n+\t\t\tce = ce->next_ce;\n+\t\t}\n+\t\tcurrent = current->next;\n+\t}\n+\treturn 0;\n+}\n+\n+static int write_index_v5(struct index_state *istate, int newfd)\n+{\n+\tstruct cache_header hdr;\n+\tstruct cache_header_v5 hdr_v5;\n+\tstruct cache_entry **cache = istate->cache;\n+\tstruct directory_entry *de;\n+\tunsigned int entries = istate->cache_nr;\n+\tunsigned int i, removed, total_dir_len;\n+\tunsigned int total_file_len, foffsetblock;\n+\tunsigned int ndir;\n+\tuint32_t crc;\n+\n+\tif (istate->filter_opts)\n+\t\tdie(\"BUG: index: cannot write a partially read index\");\n+\n+\tfor (i = removed = 0; i < entries; i++) {\n+\t\tif (cache[i]->ce_flags & CE_REMOVE)\n+\t\t\tremoved++;\n+\t}\n+\thdr.hdr_signature = htonl(CACHE_SIGNATURE);\n+\thdr.hdr_version = htonl(istate->version);\n+\thdr.hdr_entries = htonl(entries - removed);\n+\thdr_v5.hdr_nextension = htonl(0); /* Currently no extensions are supported */\n+\n+\ttotal_dir_len = 0;\n+\ttotal_file_len = 0;\n+\tde = compile_directory_data(istate, entries, &ndir,\n+\t\t\t\t    &total_dir_len, &total_file_len);\n+\thdr_v5.hdr_ndir = htonl(ndir);\n+\n+\tfoffsetblock = sizeof(hdr) + sizeof(hdr_v5) + 4\n+\t\t+ (ndir + 1) * 4\n+\t\t+ total_dir_len\n+\t\t+ ndir * (offsetof(struct ondisk_directory_entry, name) + 4);\n+\thdr_v5.hdr_fblockoffset = htonl(foffsetblock + (entries - removed + 1) * 4);\n+\tcrc = 0;\n+\tif (ce_write(&crc, newfd, &hdr, sizeof(hdr)) < 0)\n+\t\treturn -1;\n+\tif (ce_write(&crc, newfd, &hdr_v5, sizeof(hdr_v5)) < 0)\n+\t\treturn -1;\n+\tcrc = htonl(crc);\n+\tif (ce_write(NULL, newfd, &crc, 4) < 0)\n+\t\treturn -1;\n+\n+\tif (write_directories(de, newfd) < 0)\n+\t\treturn -1;\n+\tif (write_entries(istate, de, entries, newfd) < 0)\n+\t\treturn -1;\n+\treturn ce_flush(newfd);\n+}\n+\n struct index_ops v5_ops = {\n \tmatch_stat_basic,\n \tverify_hdr,\n \tread_index_v5,\n-\tNULL\n+\twrite_index_v5\n };\ndiff --git a/read-cache.c b/read-cache.c\nindex baa052c..46551af 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -106,7 +106,7 @@ int match_stat_data(const struct stat_data *sd, struct stat *st)\n \treturn changed;\n }\n \n-static uint32_t calculate_stat_crc(struct cache_entry *ce)\n+uint32_t calculate_stat_crc(struct cache_entry *ce)\n {\n \tunsigned int ctimens = 0;\n \tuint32_t stat, stat_crc;\n@@ -227,6 +227,8 @@ static void set_istate_ops(struct index_state *istate)\n {\n \tif (istate->version >= 2 && istate->version <= 4)\n \t\tistate->ops = &v2_ops;\n+\tif (istate->version == 5)\n+\t\tistate->ops = &v5_ops;\n }\n \n int ce_match_stat_basic(const struct index_state *istate,\ndiff --git a/read-cache.h b/read-cache.h\nindex 7823fbb..9d66df6 100644\n--- a/read-cache.h\n+++ b/read-cache.h\n@@ -61,5 +61,6 @@ extern int ce_match_stat_basic(const struct index_state *istate,\n \t\t\t       const struct cache_entry *ce, struct stat *st);\n extern int is_racy_timestamp(const struct index_state *istate, const struct cache_entry *ce);\n extern void set_index_entry(struct index_state *istate, int nr, struct cache_entry *ce);\n+extern uint32_t calculate_stat_crc(struct cache_entry *ce);\n \n #endif\n-- \n1.8.4.2\n"},{"id":"231179","messageId":"1385553659-9928-17-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 16/24] read-cache: write index-v5 cache-tree data","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:51Z","receivedAt":"2013-11-27T12:00:51Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Write the cache-tree data for the index version 5 file format. The\nin-memory cache-tree data is converted to the ondisk format, by adding\nit to the directory entries, that were compiled from the cache-entries\nin the step before.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n read-cache-v5.c | 53 +++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 53 insertions(+)\n\ndiff --git a/read-cache-v5.c b/read-cache-v5.c\nindex 797022f..0d06cfe 100644\n--- a/read-cache-v5.c\n+++ b/read-cache-v5.c\n@@ -789,6 +789,57 @@ static struct directory_entry *get_directory(char *dir, unsigned int dir_len,\n \treturn ret;\n }\n \n+static void convert_one_to_ondisk(struct hash_table *table, struct cache_tree *it,\n+\t\t\t\t  const char *path, int pathlen, uint32_t crc)\n+{\n+\tint i, path_len = strlen(path);\n+\tstruct directory_entry *search;\n+\n+\tcrc = crc32(crc, (Bytef*)path, pathlen);\n+\tsearch = lookup_hash(crc, table);\n+\twhile (search && (path_len > search->de_pathlen\n+\t\t\t  || strcmp(path, search->pathname + search->de_pathlen - path_len)))\n+\t\tsearch = search->next_hash;\n+\tif (!search)\n+\t\treturn;\n+\t/*\n+\t * The number of subtrees is already calculated by\n+\t * compile_directory_data, therefore we only need to\n+\t * add the entry_count\n+\t */\n+\tsearch->de_nentries = it->entry_count;\n+\tif (0 <= it->entry_count)\n+\t\thashcpy(search->sha1, it->sha1);\n+\n+#if DEBUG\n+\tif (0 <= it->entry_count)\n+\t\tfprintf(stderr, \"cache-tree <%.*s> (%d ent, %d subtree) %s\\n\",\n+\t\t\tpathlen, path, it->entry_count, it->subtree_nr,\n+\t\t\tsha1_to_hex(it->sha1));\n+\telse\n+\t\tfprintf(stderr, \"cache-tree <%.*s> (%d subtree) invalid\\n\",\n+\t\t\tpathlen, path, it->subtree_nr);\n+#endif\n+\n+\tif (strcmp(path, \"\"))\n+\t\tcrc = crc32(crc, (Bytef*)\"/\", 1);\n+\tfor (i = 0; i < it->subtree_nr; i++) {\n+\t\tstruct cache_tree_sub *down = it->down[i];\n+\t\tif (i) {\n+\t\t\tstruct cache_tree_sub *prev = it->down[i-1];\n+\t\t\tif (subtree_name_cmp(down->name, down->namelen,\n+\t\t\t\t\t     prev->name, prev->namelen) <= 0)\n+\t\t\t\tdie(\"fatal - unsorted cache subtree\");\n+\t\t}\n+\t\tconvert_one_to_ondisk(table, down->cache_tree, down->name, down->namelen, crc);\n+\t}\n+}\n+\n+static void cache_tree_to_ondisk(struct hash_table *table, struct cache_tree *root)\n+{\n+\tconvert_one_to_ondisk(table, root, \"\", 0, 0);\n+}\n+\n static void ce_queue_push(struct cache_entry **head,\n \t\t\t  struct cache_entry **tail,\n \t\t\t  struct cache_entry *ce)\n@@ -844,6 +895,8 @@ static struct directory_entry *compile_directory_data(struct index_state *istate\n \t\t\t*total_file_len -= search->de_pathlen + 1;\n \t\tce_queue_push(&(search->ce), &(search->ce_last), cache[i]);\n \t}\n+\tif (istate->cache_tree)\n+\t\tcache_tree_to_ondisk(&table, istate->cache_tree);\n \treturn de;\n }\n \n-- \n1.8.4.2\n"},{"id":"231180","messageId":"1385553659-9928-18-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 17/24] read-cache: write resolve-undo data for index-v5","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:52Z","receivedAt":"2013-11-27T12:00:52Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Make git read the resolve-undo data from the index.\n\nSince the resolve-undo data is joined with the conflicts in the ondisk\nformat of the index file version 5, conflicts and resolved data is read\nat the same time, and the resolve-undo data is then converted to the\nin-memory format.\n\nHelped-by: Thomas Rast <trast@student.ethz.ch>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n read-cache-v5.c | 199 +++++++++++++++++++++++++++++++++++++++++++++++++++++---\n 1 file changed, 191 insertions(+), 8 deletions(-)\n\ndiff --git a/read-cache-v5.c b/read-cache-v5.c\nindex 0d06cfe..a5e9b5a 100644\n--- a/read-cache-v5.c\n+++ b/read-cache-v5.c\n@@ -723,6 +723,29 @@ static void ondisk_from_directory_entry(struct directory_entry *de,\n \t}\n }\n \n+static struct conflict_part *conflict_part_from_inmemory(struct cache_entry *ce)\n+{\n+\tstruct conflict_part *conflict;\n+\tint flags;\n+\n+\tconflict = xmalloc(sizeof(struct conflict_part));\n+\tflags                = CONFLICT_CONFLICTED;\n+\tflags               |= ce_stage(ce) << CONFLICT_STAGESHIFT;\n+\tconflict->flags      = flags;\n+\tconflict->entry_mode = ce->ce_mode;\n+\tconflict->next       = NULL;\n+\thashcpy(conflict->sha1, ce->sha1);\n+\treturn conflict;\n+}\n+\n+static void conflict_to_ondisk(struct conflict_part *cp,\n+\t\t\t       struct ondisk_conflict_part *ondisk)\n+{\n+\tondisk->flags      = htons(cp->flags);\n+\tondisk->entry_mode = htons(cp->entry_mode);\n+\thashcpy(ondisk->sha1, cp->sha1);\n+}\n+\n static void insert_directory_entry(struct directory_entry *de,\n \t\t\t\t   struct hash_table *table,\n \t\t\t\t   unsigned int *total_dir_len,\n@@ -789,6 +812,11 @@ static struct directory_entry *get_directory(char *dir, unsigned int dir_len,\n \treturn ret;\n }\n \n+static struct conflict_entry *create_conflict_entry_from_ce(struct cache_entry *ce)\n+{\n+\treturn create_new_conflict(ce->name, ce_namelen(ce));\n+}\n+\n static void convert_one_to_ondisk(struct hash_table *table, struct cache_tree *it,\n \t\t\t\t  const char *path, int pathlen, uint32_t crc)\n {\n@@ -840,6 +868,52 @@ static void cache_tree_to_ondisk(struct hash_table *table, struct cache_tree *ro\n \tconvert_one_to_ondisk(table, root, \"\", 0, 0);\n }\n \n+static void resolve_undo_to_ondisk(struct string_list *resolve_undo,\n+\t\t\t\t   struct conflict_entry **conflict_queue)\n+{\n+\tstruct string_list_item *item;\n+\tstruct conflict_entry *current = *conflict_queue;\n+\n+\tif (!resolve_undo)\n+\t\treturn;\n+\tfor_each_string_list_item(item, resolve_undo) {\n+\t\tstruct conflict_entry *conflict_entry;\n+\t\tstruct resolve_undo_info *ui = item->util;\n+\t\tint i, len;\n+\n+\t\tif (!ui)\n+\t\t\tcontinue;\n+\n+\t\tlen = strlen(item->string);\n+\t\twhile (current && current->next &&\n+\t\t       cache_name_compare(current->name, current->namelen,\n+\t\t\t\t\t  item->string, len))\n+\t\t\tcurrent = current->next;\n+\n+\t\tconflict_entry = create_new_conflict(item->string, len);\n+\t\tfor (i = 0; i < 3; i++) {\n+\t\t\tif (ui->mode[i]) {\n+\t\t\t\tstruct conflict_part *cp;\n+\n+\t\t\t\tcp = xmalloc(sizeof(struct conflict_part));\n+\t\t\t\tcp->flags = (i + 1) << CONFLICT_STAGESHIFT;\n+\t\t\t\tcp->entry_mode = ui->mode[i];\n+\t\t\t\tcp->next = NULL;\n+\t\t\t\thashcpy(cp->sha1, ui->sha1[i]);\n+\t\t\t\tadd_part_to_conflict_entry(conflict_entry, cp);\n+\t\t\t}\n+\t\t}\n+\t\tif (!*conflict_queue) {\n+\t\t\t*conflict_queue = conflict_entry;\n+\t\t\tconflict_entry->next = NULL;\n+\t\t\tcurrent = conflict_entry;\n+\t\t} else {\n+\t\t\tconflict_entry->next = current->next;\n+\t\t\tcurrent->next = conflict_entry;\n+\t\t}\n+\t}\n+}\n+\n static void ce_queue_push(struct cache_entry **head,\n \t\t\t  struct cache_entry **tail,\n \t\t\t  struct cache_entry *ce)\n@@ -855,15 +929,32 @@ static void ce_queue_push(struct cache_entry **head,\n \t*tail = (*tail)->next_ce;\n }\n \n+static void conflict_queue_push(struct conflict_entry **head,\n+\t\t\t\tstruct conflict_entry **tail,\n+\t\t\t\tstruct conflict_entry *conflict)\n+{\n+\tif (!*head) {\n+\t\t*head = *tail = conflict;\n+\t\t(*tail)->next = NULL;\n+\t\treturn;\n+\t}\n+\n+\t(*tail)->next = conflict;\n+\tconflict->next = NULL;\n+\t*tail = (*tail)->next;\n+}\n+\n static struct directory_entry *compile_directory_data(struct index_state *istate,\n \t\t\t\t\t\t      int nfile, unsigned int *ndir,\n \t\t\t\t\t\t      unsigned int *total_dir_len,\n-\t\t\t\t\t\t      unsigned int *total_file_len)\n+\t\t\t\t\t\t      unsigned int *total_file_len,\n+\t\t\t\t\t\t      struct conflict_entry **conflict_queue)\n {\n \tint i, dir_len = -1;\n \tchar *dir;\n \tstruct directory_entry *de, *current, *search;\n \tstruct cache_entry **cache = istate->cache;\n+\tstruct conflict_entry *conflict_entry = NULL, *tail;\n \tstruct hash_table table;\n \tuint32_t crc;\n \n@@ -894,9 +985,22 @@ static struct directory_entry *compile_directory_data(struct index_state *istate\n \t\tif (search->de_pathlen)\n \t\t\t*total_file_len -= search->de_pathlen + 1;\n \t\tce_queue_push(&(search->ce), &(search->ce_last), cache[i]);\n+\n+\t\tif (ce_stage(cache[i]) > 0) {\n+\t\t\tstruct conflict_part *conflict_part;\n+\t\t\tif (!conflict_entry ||\n+\t\t\t    cache_name_compare(conflict_entry->name, conflict_entry->namelen,\n+\t\t\t\t\t       cache[i]->name, ce_namelen(cache[i]))) {\n+\t\t\t\tconflict_entry = create_conflict_entry_from_ce(cache[i]);\n+\t\t\t\tconflict_queue_push(conflict_queue, &tail, conflict_entry);\n+\t\t\t}\n+\t\t\tconflict_part = conflict_part_from_inmemory(cache[i]);\n+\t\t\tadd_part_to_conflict_entry(conflict_entry, conflict_part);\n+\t\t}\n \t}\n \tif (istate->cache_tree)\n \t\tcache_tree_to_ondisk(&table, istate->cache_tree);\n+\tresolve_undo_to_ondisk(istate->resolve_undo, conflict_queue);\n \treturn de;\n }\n \n@@ -1062,16 +1166,82 @@ static int write_entries(struct index_state *istate,\n \treturn 0;\n }\n \n+static int write_conflict(struct conflict_entry *conflict, int fd)\n+{\n+\tstruct conflict_entry *current;\n+\tstruct conflict_part *current_part;\n+\tuint32_t crc;\n+\n+\tcurrent = conflict;\n+\twhile (current) {\n+\t\tunsigned int to_write, i;\n+\n+\t\tcrc = 0;\n+\t\tif (ce_write(&crc, fd, current->name, current->namelen) < 0)\n+\t\t\treturn -1;\n+\t\tif (ce_write(&crc, fd, \"\\0\", 1) < 0)\n+\t\t\treturn -1;\n+\t\tto_write = htonl(current->nfileconflicts);\n+\t\tif (ce_write(&crc, fd, (Bytef*)&to_write, 4) < 0)\n+\t\t\treturn -1;\n+\t\tcurrent_part = current->entries;\n+\t\tfor (i = 0; i < current->nfileconflicts; i++) {\n+\t\t\tstruct ondisk_conflict_part ondisk;\n+\n+\t\t\tconflict_to_ondisk(current_part, &ondisk);\n+\t\t\tif (ce_write(&crc, fd, (Bytef*)&ondisk, sizeof(struct ondisk_conflict_part)) < 0)\n+\t\t\t\treturn 0;\n+\t\t\tcurrent_part = current_part->next;\n+\t\t}\n+\t\tcrc = htonl(crc);\n+\t\tif (ce_write(NULL, fd, &crc, 4) < 0)\n+\t\t\treturn -1;\n+\t\tcurrent = current->next;\n+\t}\n+\treturn 0;\n+}\n+\n+static int write_resolve_undo(struct index_state *istate,\n+\t\t\t      struct conflict_entry *conflict_queue,\n+\t\t\t      int fd)\n+{\n+\tstruct conflict_entry *current;\n+\tint nr = 0;\n+\tuint32_t crc = 0, to_write;\n+\n+\t/* Just count */\n+\tfor (current = conflict_queue; current; current = current->next)\n+\t\tnr++;\n+\n+\tif (ce_write(&crc, fd, \"REUC\", 4) < 0)\n+\t\treturn -1;\n+\tto_write = htonl(nr);\n+\tif (ce_write(&crc, fd, &to_write, 4) < 0)\n+\t\treturn -1;\n+\tto_write = htonl(crc);\n+\tif (ce_write(NULL, fd, &to_write, 4) < 0)\n+\t\treturn -1;\n+\n+\tcurrent = conflict_queue;\n+\twhile (current) {\n+\t\tif (write_conflict(current, fd) < 0)\n+\t\t\treturn -1;\n+\t\tcurrent = current->next;\n+\t}\n+\treturn 0;\n+}\n+\n static int write_index_v5(struct index_state *istate, int newfd)\n {\n \tstruct cache_header hdr;\n \tstruct cache_header_v5 hdr_v5;\n \tstruct cache_entry **cache = istate->cache;\n \tstruct directory_entry *de;\n+\tstruct conflict_entry *conflict_queue = NULL;\n \tunsigned int entries = istate->cache_nr;\n \tunsigned int i, removed, total_dir_len;\n \tunsigned int total_file_len, foffsetblock;\n-\tunsigned int ndir;\n+\tunsigned int ndir, extoffset, nextension;\n \tuint32_t crc;\n \n \tif (istate->filter_opts)\n@@ -1084,24 +1254,34 @@ static int write_index_v5(struct index_state *istate, int newfd)\n \thdr.hdr_signature = htonl(CACHE_SIGNATURE);\n \thdr.hdr_version = htonl(istate->version);\n \thdr.hdr_entries = htonl(entries - removed);\n-\thdr_v5.hdr_nextension = htonl(0); /* Currently no extensions are supported */\n \n \ttotal_dir_len = 0;\n \ttotal_file_len = 0;\n \tde = compile_directory_data(istate, entries, &ndir,\n-\t\t\t\t    &total_dir_len, &total_file_len);\n+\t\t\t\t    &total_dir_len, &total_file_len,\n+\t\t\t\t    &conflict_queue);\n \thdr_v5.hdr_ndir = htonl(ndir);\n \n-\tfoffsetblock = sizeof(hdr) + sizeof(hdr_v5) + 4\n-\t\t+ (ndir + 1) * 4\n-\t\t+ total_dir_len\n-\t\t+ ndir * (offsetof(struct ondisk_directory_entry, name) + 4);\n+\tnextension = (istate->resolve_undo || conflict_queue) ? 1 : 0;\n+\tfoffsetblock = sizeof(hdr) + sizeof(hdr_v5) + (nextension * 4) + 4 +\n+\t\t(ndir + 1) * 4 + total_dir_len +\n+\t\tndir * (offsetof(struct ondisk_directory_entry, name) + 4);\n \thdr_v5.hdr_fblockoffset = htonl(foffsetblock + (entries - removed + 1) * 4);\n+\thdr_v5.hdr_nextension = htonl(nextension);\n+\n \tcrc = 0;\n \tif (ce_write(&crc, newfd, &hdr, sizeof(hdr)) < 0)\n \t\treturn -1;\n \tif (ce_write(&crc, newfd, &hdr_v5, sizeof(hdr_v5)) < 0)\n \t\treturn -1;\n+\n+\tif (nextension) {\n+\t\textoffset = foffsetblock + (entries - removed + 1) * 4 + total_file_len +\n+\t\t\t(entries - removed) * (offsetof(struct ondisk_cache_entry, name) + 4);\n+\t\textoffset = htonl(extoffset);\n+\t\tif (ce_write(&crc, newfd, &extoffset, 4) < 0)\n+\t\t\treturn -1;\n+\t}\n \tcrc = htonl(crc);\n \tif (ce_write(NULL, newfd, &crc, 4) < 0)\n \t\treturn -1;\n@@ -1110,6 +1290,9 @@ static int write_index_v5(struct index_state *istate, int newfd)\n \t\treturn -1;\n \tif (write_entries(istate, de, entries, newfd) < 0)\n \t\treturn -1;\n+\tif (nextension)\n+\t\tif (write_resolve_undo(istate, conflict_queue, newfd) < 0)\n+\t\t\treturn -1;\n \treturn ce_flush(newfd);\n }\n \n-- \n1.8.4.2\n"},{"id":"231181","messageId":"1385553659-9928-19-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 18/24] update-index.c: rewrite index when index-version is given","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:53Z","receivedAt":"2013-11-27T12:00:53Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Make update-index always rewrite the index when a index-version\nis given, even if the index already has the right version.\nThis option is used for performance testing the writer and\nreader.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n builtin/update-index.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex c5bb889..8b3f7a0 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -6,6 +6,7 @@\n #include \"cache.h\"\n #include \"quote.h\"\n #include \"cache-tree.h\"\n+#include \"read-cache.h\"\n #include \"tree-walk.h\"\n #include \"builtin.h\"\n #include \"refs.h\"\n@@ -861,8 +862,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\t\t    preferred_index_format,\n \t\t\t    INDEX_FORMAT_LB, INDEX_FORMAT_UB);\n \n-\t\tif (the_index.version != preferred_index_format)\n-\t\t\tactive_cache_changed = 1;\n+\t\tactive_cache_changed = 1;\n \t\tchange_cache_version(preferred_index_format);\n \t}\n \n-- \n1.8.4.2\n"},{"id":"231182","messageId":"1385553659-9928-20-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 19/24] p0003-index.sh: add perf test for the index formats","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:54Z","receivedAt":"2013-11-27T12:00:54Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"From: Thomas Rast <trast@inf.ethz.ch>\n\nAdd a performance test for index version [23]/4/5 by using\ngit update-index --index-version=x, thus testing both the reader\nand the writer speed of all index formats.\n\nSigned-off-by: Thomas Rast <trast@inf.ethz.ch>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n t/perf/p0003-index.sh | 63 +++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 63 insertions(+)\n create mode 100755 t/perf/p0003-index.sh\n\ndiff --git a/t/perf/p0003-index.sh b/t/perf/p0003-index.sh\nnew file mode 100755\nindex 0000000..5360175\n--- /dev/null\n+++ b/t/perf/p0003-index.sh\n@@ -0,0 +1,63 @@\n+#!/bin/sh\n+\n+test_description=\"Tests index versions [23]/4/5\"\n+\n+. ./perf-lib.sh\n+\n+test_perf_large_repo\n+\n+test_expect_success \"convert to v3\" \"\n+\tgit update-index --index-version=2\n+\"\n+\n+test_perf \"v[23]: update-index\" \"\n+\tgit update-index --index-version=2 >/dev/null\n+\"\n+\n+subdir=$(git ls-files | sed 's#/[^/]*$##' | grep -v '^$' | uniq | tail -n 30 | head -1)\n+\n+test_perf \"v[23]: grep nonexistent -- subdir\" \"\n+\ttest_must_fail git grep nonexistent -- $subdir >/dev/null\n+\"\n+\n+test_perf \"v[23]: ls-files -- subdir\" \"\n+\tgit ls-files $subdir >/dev/null\n+\"\n+\n+test_expect_success \"convert to v4\" \"\n+\tgit update-index --index-version=4\n+\"\n+\n+test_perf \"v4: update-index\" \"\n+\tgit update-index --index-version=4 >/dev/null\n+\"\n+\n+test_perf \"v4: grep nonexistent -- subdir\" \"\n+\ttest_must_fail git grep nonexistent -- $subdir >/dev/null\n+\"\n+\n+test_perf \"v4: ls-files -- subdir\" \"\n+\tgit ls-files $subdir >/dev/null\n+\"\n+\n+test_expect_success \"convert to v5\" \"\n+\tgit update-index --index-version=5\n+\"\n+\n+test_perf \"v5: update-index\" \"\n+\tgit update-index --index-version=5 >/dev/null\n+\"\n+\n+test_perf \"v5: ls-files\" \"\n+\tgit ls-files >/dev/null\n+\"\n+\n+test_perf \"v5: grep nonexistent -- subdir\" \"\n+\ttest_must_fail git grep nonexistent -- $subdir >/dev/null\n+\"\n+\n+test_perf \"v5: ls-files -- subdir\" \"\n+\tgit ls-files $subdir >/dev/null\n+\"\n+\n+test_done\n-- \n1.8.4.2\n"},{"id":"231183","messageId":"1385553659-9928-21-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 20/24] introduce GIT_INDEX_VERSION environment variable","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:55Z","receivedAt":"2013-11-27T12:00:55Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Respect a GIT_INDEX_VERSION environment variable, when a new index is\ninitialized.  Setting the environment variable will not cause existing\nindex files to be converted to another format for additional safety.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n Documentation/git.txt | 5 +++++\n read-cache.c          | 9 +++++++--\n 2 files changed, 12 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/git.txt b/Documentation/git.txt\nindex 10cddb5..2b2aad5 100644\n--- a/Documentation/git.txt\n+++ b/Documentation/git.txt\n@@ -703,6 +703,11 @@ Git so take care if using Cogito etc.\n \tindex file. If not specified, the default of `$GIT_DIR/index`\n \tis used.\n \n+'GIT_INDEX_VERSION'::\n+\tThis environment variable allows the specification of an index\n+\tversion for new repositories.  It won't affect existing index\n+\tfiles.  By default index file version 3 is used.\n+\n 'GIT_OBJECT_DIRECTORY'::\n \tIf the object storage directory is specified via this\n \tenvironment variable then the sha1 directories are created\ndiff --git a/read-cache.c b/read-cache.c\nindex 46551af..04430e5 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1233,8 +1233,13 @@ static struct cache_entry *refresh_cache_entry(struct cache_entry *ce, int reall\n void initialize_index(struct index_state *istate, int version)\n {\n \tistate->initialized = 1;\n-\tif (!version)\n-\t\tversion = INDEX_FORMAT_DEFAULT;\n+\tif (!version) {\n+\t\tchar *envversion = getenv(\"GIT_INDEX_VERSION\");\n+\t\tif (!envversion)\n+\t\t\tversion = INDEX_FORMAT_DEFAULT;\n+\t\telse\n+\t\t\tversion = atoi(envversion);\n+\t}\n \tistate->version = version;\n \tset_istate_ops(istate);\n }\n-- \n1.8.4.2\n"},{"id":"231184","messageId":"1385553659-9928-22-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 21/24] test-lib: allow setting the index format version","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:56Z","receivedAt":"2013-11-27T12:00:56Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"When running the test suite, it should be possible to set the default\nindex format for the tests.  Do that by allowing the user to add a\nTEST_GIT_INDEX_VERSION variable in config.mak setting the index version.\n\nIf it isn't set, the default version given in the source code is\nused (currently version 3).\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n Makefile                | 7 +++++++\n t/test-lib-functions.sh | 5 +++++\n t/test-lib.sh           | 3 +++\n 3 files changed, 15 insertions(+)\n\ndiff --git a/Makefile b/Makefile\nindex 6a1b054..8539548 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -342,6 +342,10 @@ all::\n # Define DEFAULT_HELP_FORMAT to \"man\", \"info\" or \"html\"\n # (defaults to \"man\") if you want to have a different default when\n # \"git help\" is called without a parameter specifying the format.\n+#\n+# Define TESTGIT_INDEX_FORMAT to 2, 3, 4 or 5 to run the test suite\n+# with a different indexfile format.  If it isn't set the index file\n+# format used is index-v[23].\n \n GIT-VERSION-FILE: FORCE\n \t@$(SHELL_PATH) ./GIT-VERSION-GEN\n@@ -2218,6 +2222,9 @@ endif\n ifdef GIT_PERF_MAKE_OPTS\n \t@echo GIT_PERF_MAKE_OPTS=\\''$(subst ','\\'',$(subst ','\\'',$(GIT_PERF_MAKE_OPTS)))'\\' >>$@\n endif\n+ifdef TEST_GIT_INDEX_VERSION\n+\t@echo TEST_GIT_INDEX_VERSION='$(subst ','\\'',$(subst ','\\'',$(TEST_GIT_INDEX_VERSION)))' >>$@\n+endif\n \n ### Detect Python interpreter path changes\n ifndef NO_PYTHON\ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex 2f79146..4034262 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -31,6 +31,11 @@ test_set_editor () {\n \texport EDITOR\n }\n \n+test_set_index_version () {\n+    GIT_INDEX_VERSION=\"$1\"\n+    export GIT_INDEX_VERSION\n+}\n+\n test_decode_color () {\n \tawk '\n \t\tfunction name(n) {\ndiff --git a/t/test-lib.sh b/t/test-lib.sh\nindex b25249e..d9e810c 100644\n--- a/t/test-lib.sh\n+++ b/t/test-lib.sh\n@@ -104,6 +104,9 @@ export GIT_AUTHOR_EMAIL GIT_AUTHOR_NAME\n export GIT_COMMITTER_EMAIL GIT_COMMITTER_NAME\n export EDITOR\n \n+GIT_INDEX_VERSION=\"$TEST_GIT_INDEX_VERSION\"\n+export GIT_INDEX_VERSION\n+\n # Add libc MALLOC and MALLOC_PERTURB test\n # only if we are not executing the test with valgrind\n if expr \" $GIT_TEST_OPTS \" : \".* --valgrind \" >/dev/null ||\n-- \n1.8.4.2\n"},{"id":"231185","messageId":"1385553659-9928-23-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 22/24] t1600: add index v5 specific tests","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:57Z","receivedAt":"2013-11-27T12:00:57Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Add a test that tests only index v5 specific corner cases, to protect\nagainst breaking them in the future.\n\nCurrently there is only one known case where the sorting is broken if\nthe index is read filtered with two different length pathspecs.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n t/t1600-index-v5.sh | 25 +++++++++++++++++++++++++\n 1 file changed, 25 insertions(+)\n create mode 100755 t/t1600-index-v5.sh\n\ndiff --git a/t/t1600-index-v5.sh b/t/t1600-index-v5.sh\nnew file mode 100755\nindex 0000000..fe68976\n--- /dev/null\n+++ b/t/t1600-index-v5.sh\n@@ -0,0 +1,25 @@\n+#!/bin/sh\n+\n+test_description=\"Test index-v5 specific corner cases\"\n+\n+. ./test-lib.sh\n+\n+test_set_index_version 5\n+\n+test_expect_success 'setup' '\n+\tmkdir -p abc/def def &&\n+\ttouch abc/def/xyz def/xyz &&\n+\tgit add . &&\n+\tgit commit -m \"test commit\"\n+'\n+\n+test_expect_success 'ls-files ordering correct' '\n+\tcat <<-\\EOF >expected &&\n+\tabc/def/xyz\n+\tdef/xyz\n+\tEOF\n+\tgit ls-files abc/def/xyz def/xyz >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_done\n-- \n1.8.4.2\n"},{"id":"231187","messageId":"1385553659-9928-24-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 23/24] POC for partial writing","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:58Z","receivedAt":"2013-11-27T12:00:58Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"This makes update-index use both partial reading and partial writing.\nPartial reading is only used no option other than the paths is passed to\nthe command.\n\nThis passes the test suite, but doesn't behave correctly when a write\nfails.  A log should be written to the lock file, in order to be able to\nrecover if a write fails.\n---\n builtin/update-index.c |  43 +++++++++++---\n cache-tree.c           |  13 +++++\n cache-tree.h           |   1 +\n cache.h                |  27 ++++++++-\n lockfile.c             |   2 +-\n read-cache-v2.c        |   2 +\n read-cache-v5.c        | 154 ++++++++++++++++++++++++++++++++++++++++---------\n read-cache.c           |  30 ++++++++++\n read-cache.h           |   1 +\n resolve-undo.c         |   1 +\n 10 files changed, 237 insertions(+), 37 deletions(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 8b3f7a0..69f0949 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -56,6 +56,7 @@ static int mark_ce_flags(const char *path, int flag, int mark)\n \t\telse\n \t\t\tactive_cache[pos]->ce_flags &= ~flag;\n \t\tcache_tree_invalidate_path(active_cache_tree, path);\n+\t\tthe_index.needs_rewrite = 1;\n \t\tactive_cache_changed = 1;\n \t\treturn 0;\n \t}\n@@ -99,6 +100,8 @@ static int add_one_path(const struct cache_entry *old, const char *path, int len\n \tmemcpy(ce->name, path, len);\n \tce->ce_flags = create_ce_flags(0);\n \tce->ce_namelen = len;\n+\tif (old)\n+\t\tce->entry_pos = old->entry_pos;\n \tfill_stat_cache_info(ce, st);\n \tce->ce_mode = ce_mode_from_stat(old, st->st_mode);\n \n@@ -268,6 +271,7 @@ static void chmod_path(int flip, const char *path)\n \t\tgoto fail;\n \t}\n \tcache_tree_invalidate_path(active_cache_tree, path);\n+\tthe_index.needs_rewrite = 1;\n \tactive_cache_changed = 1;\n \treport(\"chmod %cx '%s'\", flip, path);\n \treturn;\n@@ -706,15 +710,18 @@ static int reupdate_callback(struct parse_opt_ctx_t *ctx,\n \n int cmd_update_index(int argc, const char **argv, const char *prefix)\n {\n-\tint newfd, entries, has_errors = 0, line_termination = '\\n';\n+\tint newfd, has_errors = 0, line_termination = '\\n';\n \tint read_from_stdin = 0;\n \tint prefix_length = prefix ? strlen(prefix) : 0;\n \tint preferred_index_format = 0;\n \tchar set_executable_bit = 0;\n \tstruct refresh_params refresh_args = {0, &has_errors};\n \tint lock_error = 0;\n+\tstruct filter_opts opts;\n+\tstruct pathspec pathspec;\n \tstruct lock_file *lock_file;\n \tstruct parse_opt_ctx_t ctx;\n+\tint i, needs_full_read = 0;\n \tint parseopt_state = PARSE_OPT_UNKNOWN;\n \tstruct option options[] = {\n \t\tOPT_BIT('q', NULL, &refresh_args.flags,\n@@ -810,9 +817,23 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \tif (newfd < 0)\n \t\tlock_error = errno;\n \n-\tentries = read_cache();\n-\tif (entries < 0)\n-\t\tdie(\"cache corrupted\");\n+\tfor (i = 0; i < argc; i++) {\n+\t\tif (!prefixcmp(argv[i], \"--\"))\n+\t\t\tneeds_full_read = 1;\n+\t}\n+\tif (!needs_full_read) {\n+\t\tmemset(&opts, 0, sizeof(struct filter_opts));\n+\t\tparse_pathspec(&pathspec, 0,\n+\t\t\t       PATHSPEC_PREFER_CWD |\n+\t\t\t       PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n+\t\t\t       prefix, argv + 1);\n+\t\topts.pathspec = &pathspec;\n+\t\tif (read_cache_filtered(&opts) < 0)\n+\t\t\tdie(\"cache corrupted\");\n+\t} else {\n+\t\tif (read_cache() < 0)\n+\t\t\tdie(\"cache corrupted\");\n+\t}\n \n \t/*\n \t * Custom copy of parse_options() because we want to handle\n@@ -862,6 +883,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\t\t    preferred_index_format,\n \t\t\t    INDEX_FORMAT_LB, INDEX_FORMAT_UB);\n \n+\t\tthe_index.needs_rewrite = 1;\n \t\tactive_cache_changed = 1;\n \t\tchange_cache_version(preferred_index_format);\n \t}\n@@ -890,17 +912,22 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t}\n \n \tif (active_cache_changed) {\n+\t\tint r;\n \t\tif (newfd < 0) {\n \t\t\tif (refresh_args.flags & REFRESH_QUIET)\n \t\t\t\texit(128);\n \t\t\tunable_to_lock_index_die(get_index_file(), lock_error);\n \t\t}\n-\t\tif (write_cache(newfd, active_cache, active_nr) ||\n-\t\t    commit_locked_index(lock_file))\n+\t\tr = write_cache_partial(newfd);\n+\t\tif (r < 0)\n \t\t\tdie(\"Unable to write new index file\");\n+\t\telse if (r == 0)\n+\t\t\tcommit_lock_file(lock_file);\n+\t\telse\n+\t\t\tremove_lock_file();\n+\t} else {\n+\t\trollback_lock_file(lock_file);\n \t}\n \n-\trollback_lock_file(lock_file);\n-\n \treturn has_errors ? 1 : 0;\n }\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 1209732..a3d18bb 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -123,6 +123,15 @@ void cache_tree_invalidate_path(struct cache_tree *it, const char *path)\n \t\treturn;\n \tslash = strchr(path, '/');\n \tit->entry_count = -1;\n+\t/*\n+\t * Mark the cache_tree directory entry as invalid too. The\n+\t * entry_count defines if the tree is valid, so we don't need\n+\t * to reset any other field.\n+\t */\n+\tif (it->de_ref) {\n+\t\tit->de_ref->de_nentries = -1;\n+\t\tit->de_ref->changed = 1;\n+\t}\n \tif (!slash) {\n \t\tint pos;\n \t\tnamelen = strlen(path);\n@@ -140,6 +149,10 @@ void cache_tree_invalidate_path(struct cache_tree *it, const char *path)\n \t\t\t\tsizeof(struct cache_tree_sub *) *\n \t\t\t\t(it->subtree_nr - pos - 1));\n \t\t\tit->subtree_nr--;\n+\t\t\tif (it->de_ref) {\n+\t\t\t\tit->de_ref->de_nsubtrees--;\n+\t\t\t\tit->de_ref->changed = 1;\n+\t\t\t}\n \t\t}\n \t\treturn;\n \t}\ndiff --git a/cache-tree.h b/cache-tree.h\nindex 9818926..eaf14a9 100644\n--- a/cache-tree.h\n+++ b/cache-tree.h\n@@ -18,6 +18,7 @@ struct cache_tree {\n \tunsigned char sha1[20];\n \tint subtree_nr;\n \tint subtree_alloc;\n+\tstruct directory_entry *de_ref;\n \tstruct cache_tree_sub **down;\n };\n \ndiff --git a/cache.h b/cache.h\nindex 71b98cf..1a634dc 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -137,11 +137,31 @@ struct cache_entry {\n \tunsigned int ce_namelen;\n \tunsigned char sha1[20];\n \tuint32_t ce_stat_crc;\n+\tunsigned int entry_pos;\n+\tunsigned int changed;\n \tstruct cache_entry *next; /* used by name_hash */\n \tstruct cache_entry *next_ce;\n \tchar name[FLEX_ARRAY]; /* more */\n };\n \n+struct directory_entry {\n+\tstruct directory_entry **sub;\n+\tstruct directory_entry *next;\n+\tstruct directory_entry *next_hash;\n+\tstruct cache_entry *ce;\n+\tstruct cache_entry *ce_last;\n+\tuint32_t de_foffset;\n+\tuint32_t de_nsubtrees;\n+\tuint32_t de_nfiles;\n+\tuint32_t de_nentries;\n+\tunsigned char sha1[20];\n+\tuint16_t de_flags;\n+\tuint32_t de_pathlen;\n+\tuint32_t entry_pos;\n+\tunsigned int changed;\n+\tchar pathname[FLEX_ARRAY];\n+};\n+\n #define CE_NAMEMASK  (0x0fff)\n #define CE_STAGEMASK (0x3000)\n #define CE_EXTENDED  (0x4000)\n@@ -317,13 +337,15 @@ struct filter_opts {\n \n struct index_state {\n \tstruct cache_entry **cache;\n+\tstruct directory_entry *root_directory;\n \tunsigned int version;\n \tunsigned int cache_nr, cache_alloc, cache_changed;\n \tstruct string_list *resolve_undo;\n \tstruct cache_tree *cache_tree;\n \tstruct cache_time timestamp;\n \tunsigned name_hash_initialized : 1,\n-\t\t initialized : 1;\n+\t\t initialized : 1,\n+\t\t needs_rewrite : 1;\n \tstruct hash_table name_hash;\n \tstruct hash_table dir_hash;\n \tstruct index_ops *ops;\n@@ -353,6 +375,7 @@ extern void free_name_hash(struct index_state *istate);\n #define is_cache_unborn() is_index_unborn(&the_index)\n #define read_cache_unmerged() read_index_unmerged(&the_index)\n #define write_cache(newfd, cache, entries) write_index(&the_index, (newfd))\n+#define write_cache_partial(newfd) write_index_partial(&the_index, (newfd))\n #define discard_cache() discard_index(&the_index)\n #define unmerged_cache() unmerged_index(&the_index)\n #define cache_name_pos(name, namelen) index_name_pos(&the_index,(name),(namelen))\n@@ -529,6 +552,7 @@ extern int read_index_from(struct index_state *, const char *path);\n extern int is_index_unborn(struct index_state *);\n extern int read_index_unmerged(struct index_state *);\n extern int write_index(struct index_state *, int newfd);\n+extern int write_index_partial(struct index_state *, int newfd);\n extern int discard_index(struct index_state *);\n extern int unmerged_index(const struct index_state *);\n extern int verify_path(const char *path);\n@@ -613,6 +637,7 @@ extern NORETURN void unable_to_lock_index_die(const char *path, int err);\n extern int hold_lock_file_for_update(struct lock_file *, const char *path, int);\n extern int hold_lock_file_for_append(struct lock_file *, const char *path, int);\n extern int commit_lock_file(struct lock_file *);\n+extern void remove_lock_file(void);\n extern void update_index_if_able(struct index_state *, struct lock_file *);\n \n extern int hold_locked_index(struct lock_file *, int);\ndiff --git a/lockfile.c b/lockfile.c\nindex 8fbcb6a..c150e5c 100644\n--- a/lockfile.c\n+++ b/lockfile.c\n@@ -7,7 +7,7 @@\n static struct lock_file *lock_file_list;\n static const char *alternate_index_output;\n \n-static void remove_lock_file(void)\n+void remove_lock_file(void)\n {\n \tpid_t me = getpid();\n \ndiff --git a/read-cache-v2.c b/read-cache-v2.c\nindex f884c10..1fec892 100644\n--- a/read-cache-v2.c\n+++ b/read-cache-v2.c\n@@ -555,5 +555,7 @@ struct index_ops v2_ops = {\n \tmatch_stat_basic,\n \tverify_hdr,\n \tread_index_v2,\n+\twrite_index_v2,\n+\t/* Partial writing is the same as writing the full index for v2 */\n \twrite_index_v2\n };\ndiff --git a/read-cache-v5.c b/read-cache-v5.c\nindex a5e9b5a..13436a3 100644\n--- a/read-cache-v5.c\n+++ b/read-cache-v5.c\n@@ -20,22 +20,6 @@ struct extension_header {\n \tuint32_t crc;\n };\n \n-struct directory_entry {\n-\tstruct directory_entry **sub;\n-\tstruct directory_entry *next;\n-\tstruct directory_entry *next_hash;\n-\tstruct cache_entry *ce;\n-\tstruct cache_entry *ce_last;\n-\tuint32_t de_foffset;\n-\tuint32_t de_nsubtrees;\n-\tuint32_t de_nfiles;\n-\tuint32_t de_nentries;\n-\tunsigned char sha1[20];\n-\tuint16_t de_flags;\n-\tuint32_t de_pathlen;\n-\tchar pathname[FLEX_ARRAY];\n-};\n-\n struct conflict_part {\n \tstruct conflict_part *next;\n \tuint16_t flags;\n@@ -246,7 +230,7 @@ static struct directory_entry *read_directories(unsigned int *dir_offset,\n \t\toffsetof(struct ondisk_directory_entry, name) - 5;\n \tdisk_de = ptr_add(mmap, *dir_offset);\n \tde = directory_entry_from_ondisk(disk_de, len);\n-\n+\tde->entry_pos = *dir_offset;\n \tdata_len = len + 1 + offsetof(struct ondisk_directory_entry, name);\n \tfilecrc = ptr_add(mmap, *dir_offset + data_len);\n \tif (!check_crc32(0, ptr_add(mmap, *dir_offset), data_len, ntoh_l(*filecrc)))\n@@ -281,6 +265,7 @@ static int read_entry(struct cache_entry **ce, char *pathname, size_t pathlen,\n \tentry_offset = first_entry_offset + ntoh_l(*beginning);\n \tdisk_ce = ptr_add(mmap, entry_offset);\n \t*ce = cache_entry_from_ondisk(disk_ce, pathname, len, pathlen);\n+\t(*ce)->entry_pos = entry_offset;\n \tfilecrc = ptr_add(mmap, entry_offset + len + 1 + sizeof(*disk_ce));\n \tif (!check_crc32(0,\n \t\tptr_add(mmap, entry_offset), len + 1 + sizeof(*disk_ce),\n@@ -439,6 +424,7 @@ static struct cache_tree *convert_one(struct directory_entry *de)\n \n \tit = cache_tree();\n \tit->entry_count = de->de_nentries;\n+\tit->de_ref = de;\n \tif (0 <= it->entry_count)\n \t\thashcpy(it->sha1, de->sha1);\n \n@@ -523,14 +509,6 @@ static int read_entries(struct index_state *istate, struct directory_entry *de,\n \treturn 0;\n }\n \n-static void free_directory_tree(struct directory_entry *de) {\n-\tint i;\n-\n-\tfor (i = 0; i < de->de_pathlen; i++)\n-\t\tfree_directory_tree(de->sub[i]);\n-\tfree(de);\n-}\n-\n /*\n  * Read an index-v5 file filtered by the filter_opts.   If opts is NULL,\n  * everything will be read.\n@@ -626,7 +604,7 @@ static int read_index_v5(struct index_state *istate, void *mmap,\n \t\t}\n \t}\n \tistate->cache_tree = cache_tree_convert_v5(root_directory);\n-\tfree_directory_tree(root_directory);\n+\tistate->root_directory = root_directory;\n \tistate->cache_nr = nr;\n \treturn 0;\n }\n@@ -696,6 +674,7 @@ static void ce_smudge_racily_clean_entry(struct cache_entry *ce)\n \t * that hasn't changed checking the sha1.\n \t */\n \tce->ce_flags |= CE_SMUDGED;\n+\tce->changed = 1;\n }\n \n static char *super_directory(char *filename)\n@@ -1231,6 +1210,103 @@ static int write_resolve_undo(struct index_state *istate,\n \treturn 0;\n }\n \n+static int write_ce_if_necessary(struct cache_entry *ce, void *cb_data)\n+{\n+\tint *fdx = cb_data, pathlen, size;\n+\tint fd = *fdx;\n+\tchar *dir;\n+\tstruct ondisk_cache_entry *ondisk;\n+\tuint32_t crc;\n+\n+\tassert(ce->entry_pos != 0);\n+\t/* TODO I'm just using the_index out of lazyness here */\n+\tif (!ce_uptodate(ce) && is_racy_timestamp(&the_index, ce))\n+\t\tce_smudge_racily_clean_entry(ce);\n+\tif (!ce->changed)\n+\t\treturn 0;\n+\tif (is_null_sha1(ce->sha1)) {\n+\t\tstatic const char msg[] = \"cache entry has null sha1: %s\";\n+\t\tstatic int allow = -1;\n+\n+\t\tif (allow < 0)\n+\t\t\tallow = git_env_bool(\"GIT_ALLOW_NULL_SHA1\", 0);\n+\t\tif (allow)\n+\t\t\twarning(msg, ce->name);\n+\t\telse\n+\t\t\treturn error(msg, ce->name);\n+\t}\n+\tdir = super_directory(ce->name);\n+\tpathlen = dir ? strlen(dir) + 1 : 0;\n+\tsize = offsetof(struct ondisk_cache_entry, name) +\n+\t\tce_namelen(ce) - pathlen + 1;\n+\tondisk = xmalloc(size);\n+\t\n+\tcrc = 0;\n+\tondisk_from_cache_entry(ce, ondisk, pathlen);\n+\tif (lseek(fd, ce->entry_pos, SEEK_SET) < ce->entry_pos)\n+\t\tdie(\"eror ce seeking\");\n+\tif (ce_write(&crc, fd, ondisk, size) < 0)\n+\t\treturn -1;\n+\tcrc = htonl(crc);\n+\tif (ce_write(NULL, fd, &crc, 4) < 0)\n+\t\treturn -1;\n+\treturn ce_flush(fd);\n+}\n+\n+static void ondisk_from_directory_entry_partial(struct directory_entry *de,\n+\t\t\t\t\t\tstruct ondisk_directory_entry *ondisk)\n+{\n+\tondisk->foffset   = htonl(de->de_foffset);\n+\tondisk->nsubtrees = htonl(de->de_nsubtrees);\n+\tondisk->nfiles    = htonl(de->de_nfiles);\n+\tondisk->nentries  = htonl(de->de_nentries);\n+\thashcpy(ondisk->sha1, de->sha1);\n+\tondisk->flags     = htons(de->de_flags);\n+\tif (de->de_pathlen == 0) {\n+\t\tmemcpy(ondisk->name, \"\\0\", 1);\n+\t} else {\n+\t\tmemcpy(ondisk->name, de->pathname, de->de_pathlen);\n+\t\tmemcpy(ondisk->name + de->de_pathlen - 1, \"/\\0\", 2);\n+\t}\n+}\n+\n+static int write_directories_partial(struct directory_entry *de, int fd)\n+{\n+\tint ondisk_size = offsetof(struct ondisk_directory_entry, name);\n+\tint size = ondisk_size + de->de_pathlen + 1;\n+\tint i;\n+\tuint32_t crc;\n+\tstruct ondisk_directory_entry *ondisk;\n+\n+\tif (de->changed) {\n+\t\tcrc = 0;\n+\t\tondisk = xmalloc(size);\n+\t\tondisk_from_directory_entry_partial(de, ondisk);\n+\t\tif (lseek(fd, de->entry_pos, SEEK_SET) < de->entry_pos)\n+\t\t\tdie(\"error directory seeking\");;\n+\t\tif (ce_write(&crc, fd, ondisk, size) < 0)\n+\t\t\treturn -1;\n+\t\tcrc = htonl(crc);\n+\t\tif (ce_write(NULL, fd, &crc, 4) < 0)\n+\t\t\treturn -1;\n+\t\tfree(ondisk);\n+\t\tif (ce_flush(fd) < 0)\n+\t\t\treturn -1;\n+\t}\n+\tfor (i = 0; i < de->de_nsubtrees; i++) {\n+\t\tif (write_directories_partial(de->sub[i], fd) < 0)\n+\t\t\treturn -1;\n+\t}\n+\treturn 0;\n+}\n+\n+static int write_partial(struct index_state *istate, int fd)\n+{\n+\twrite_directories_partial(istate->root_directory, fd);\n+\n+\treturn for_each_index_entry(istate, write_ce_if_necessary, &fd);\n+}\n+\n static int write_index_v5(struct index_state *istate, int newfd)\n {\n \tstruct cache_header hdr;\n@@ -1296,9 +1372,33 @@ static int write_index_v5(struct index_state *istate, int newfd)\n \treturn ce_flush(newfd);\n }\n \n+static int write_index_partial_v5(struct index_state *istate, int newfd)\n+{\n+\tint fd;\n+\tchar *path = get_index_file();\n+\n+\tif (istate->needs_rewrite || istate->cache_nr == 0)\n+\t\treturn write_index_v5(istate, newfd);\n+\tif (istate->filter_opts && istate->needs_rewrite)\n+\t\tdie(\"BUG: cannot write a partially read index\");\n+\tfd = open(path, O_RDWR, 0666);\n+\tif (fd < 0) {\n+\t\tif (errno == ENOENT)\n+\t\t\tdie(\"no index file exists cannot do a partial write\");\n+\t\tdie_errno(\"index file opening for writing failed\");\n+\t}\n+\n+\tif (write_partial(istate, fd) < 0)\n+\t\treturn -1;\n+\tif (ce_flush(fd) < 0)\n+\t\treturn -1;\n+\treturn 1;\n+}\n+\n struct index_ops v5_ops = {\n \tmatch_stat_basic,\n \tverify_hdr,\n \tread_index_v5,\n-\twrite_index_v5\n+\twrite_index_v5,\n+\twrite_index_partial_v5\n };\ndiff --git a/read-cache.c b/read-cache.c\nindex 04430e5..1cad0e2 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -32,6 +32,9 @@ static void replace_index_entry(struct index_state *istate, int nr, struct cache\n \n \tremove_name_hash(istate, old);\n \tset_index_entry(istate, nr, ce);\n+\tce->changed = 1;\n+\tif (ce->entry_pos == 0)\n+\t\tistate->needs_rewrite = 1;\n \tistate->cache_changed = 1;\n }\n \n@@ -494,6 +497,7 @@ int remove_index_entry_at(struct index_state *istate, int pos)\n \n \trecord_resolve_undo(istate, ce);\n \tremove_name_hash(istate, ce);\n+\tistate->needs_rewrite = 1;\n \tistate->cache_changed = 1;\n \tistate->cache_nr--;\n \tif (pos >= istate->cache_nr)\n@@ -520,6 +524,7 @@ void remove_marked_cache_entries(struct index_state *istate)\n \t\telse\n \t\t\tce_array[j++] = ce_array[i];\n \t}\n+\tistate->needs_rewrite = 1;\n \tistate->cache_changed = 1;\n \tistate->cache_nr = j;\n }\n@@ -1024,6 +1029,7 @@ int add_index_entry(struct index_state *istate, struct cache_entry *ce, int opti\n \t\t\tistate->cache + pos,\n \t\t\t(istate->cache_nr - pos - 1) * sizeof(ce));\n \tset_index_entry(istate, pos, ce);\n+\tistate->needs_rewrite = 1;\n \tistate->cache_changed = 1;\n \treturn 0;\n }\n@@ -1108,6 +1114,8 @@ static struct cache_entry *refresh_cache_ent(struct index_state *istate,\n \tsize = ce_size(ce);\n \tupdated = xmalloc(size);\n \tmemcpy(updated, ce, size);\n+\tupdated->changed = 1;\n+\tupdated->entry_pos = ce->entry_pos;\n \tfill_stat_cache_info(updated, &st);\n \t/*\n \t * If ignore_valid is not set, we should leave CE_VALID bit\n@@ -1201,6 +1209,8 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n \t\t\t\t * means the index is not valid anymore.\n \t\t\t\t */\n \t\t\t\tce->ce_flags &= ~CE_VALID;\n+\t\t\t\t/* TODO: remove this maybe? */\n+\t\t\t\tistate->needs_rewrite = 1;\n \t\t\t\tistate->cache_changed = 1;\n \t\t\t}\n \t\t\tif (quiet)\n@@ -1241,6 +1251,8 @@ void initialize_index(struct index_state *istate, int version)\n \t\t\tversion = atoi(envversion);\n \t}\n \tistate->version = version;\n+\tistate->needs_rewrite = 0;\n+\tistate->root_directory = NULL;\n \tset_istate_ops(istate);\n }\n \n@@ -1427,6 +1439,16 @@ int is_index_unborn(struct index_state *istate)\n \treturn (!istate->cache_nr && !istate->timestamp.sec);\n }\n \n+static void free_directory_tree(struct directory_entry *de) {\n+\tint i;\n+\n+\tif (!de)\n+\t\treturn;\n+\tfor (i = 0; i < de->de_pathlen; i++)\n+\t\tfree_directory_tree(de->sub[i]);\n+\tfree(de);\n+}\n+\n int discard_index(struct index_state *istate)\n {\n \tint i;\n@@ -1435,6 +1457,7 @@ int discard_index(struct index_state *istate)\n \t\tfree(istate->cache[i]);\n \tresolve_undo_clear_index(istate);\n \tistate->cache_nr = 0;\n+\tistate->needs_rewrite = 0;\n \tistate->cache_changed = 0;\n \tistate->timestamp.sec = 0;\n \tistate->timestamp.nsec = 0;\n@@ -1446,6 +1469,8 @@ int discard_index(struct index_state *istate)\n \tistate->cache_alloc = 0;\n \tistate->ops = NULL;\n \tistate->filter_opts = NULL;\n+\tfree_directory_tree(istate->root_directory);\n+\tistate->root_directory = NULL;\n \treturn 0;\n }\n \n@@ -1491,6 +1516,11 @@ int write_index(struct index_state *istate, int newfd)\n \treturn istate->ops->write_index(istate, newfd);\n }\n \n+int write_index_partial(struct index_state *istate, int newfd)\n+{\n+\treturn istate->ops->write_index_partial(istate, newfd);\n+}\n+\n /*\n  * Read the index file that is potentially unmerged into given\n  * index_state, dropping any unmerged entries.  Returns true if\ndiff --git a/read-cache.h b/read-cache.h\nindex 9d66df6..e7f36ae 100644\n--- a/read-cache.h\n+++ b/read-cache.h\n@@ -31,6 +31,7 @@ struct index_ops {\n \tint (*read_index)(struct index_state *istate, void *mmap, unsigned long mmap_size,\n \t\t\t  struct filter_opts *opts);\n \tint (*write_index)(struct index_state *istate, int newfd);\n+\tint (*write_index_partial)(struct index_state *istate, int newfd);\n };\n \n extern struct index_ops v2_ops;\ndiff --git a/resolve-undo.c b/resolve-undo.c\nindex c09b006..c496c20 100644\n--- a/resolve-undo.c\n+++ b/resolve-undo.c\n@@ -110,6 +110,7 @@ void resolve_undo_clear_index(struct index_state *istate)\n \tstring_list_clear(resolve_undo, 1);\n \tfree(resolve_undo);\n \tistate->resolve_undo = NULL;\n+\tistate->needs_rewrite = 1;\n \tistate->cache_changed = 1;\n }\n \n-- \n1.8.4.2\n"},{"id":"231186","messageId":"1385553659-9928-25-git-send-email-t.gummerer@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"[PATCH v4 24/24] perf: add partial writing test","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-27T12:00:59Z","receivedAt":"2013-11-27T12:00:59Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Add a test that uses update-index and exercises the partial writing code\npath.\n---\n t/perf/p0003-index.sh | 33 +++++++++++++++++++++++++++++++++\n 1 file changed, 33 insertions(+)\n\ndiff --git a/t/perf/p0003-index.sh b/t/perf/p0003-index.sh\nindex 5360175..d1f590b 100755\n--- a/t/perf/p0003-index.sh\n+++ b/t/perf/p0003-index.sh\n@@ -5,6 +5,7 @@ test_description=\"Tests index versions [23]/4/5\"\n . ./perf-lib.sh\n \n test_perf_large_repo\n+test_checkout_worktree\n \n test_expect_success \"convert to v3\" \"\n \tgit update-index --index-version=2\n@@ -15,6 +16,7 @@ test_perf \"v[23]: update-index\" \"\n \"\n \n subdir=$(git ls-files | sed 's#/[^/]*$##' | grep -v '^$' | uniq | tail -n 30 | head -1)\n+file=$(git ls-files | tail -n 30 | head -1)\n \n test_perf \"v[23]: grep nonexistent -- subdir\" \"\n \ttest_must_fail git grep nonexistent -- $subdir >/dev/null\n@@ -24,6 +26,16 @@ test_perf \"v[23]: ls-files -- subdir\" \"\n \tgit ls-files $subdir >/dev/null\n \"\n \n+test_expect_success \"v[23] update-index prepare\" \"\n+\techo x >$file\n+\"\n+\n+test_perf_cleanup \"v[23] update-index\" \"\n+\tgit update-index $file\n+\" \"\n+\tgit reset\n+\"\n+\n test_expect_success \"convert to v4\" \"\n \tgit update-index --index-version=4\n \"\n@@ -40,6 +52,17 @@ test_perf \"v4: ls-files -- subdir\" \"\n \tgit ls-files $subdir >/dev/null\n \"\n \n+test_expect_success \"v4 update-index prepare\" \"\n+\techo x >$file\n+\"\n+\n+test_perf_cleanup \"v4 update-index\" \"\n+\tgit update-index $file\n+\" \"\n+\tgit reset\n+\"\n+\n+\n test_expect_success \"convert to v5\" \"\n \tgit update-index --index-version=5\n \"\n@@ -60,4 +83,14 @@ test_perf \"v5: ls-files -- subdir\" \"\n \tgit ls-files $subdir >/dev/null\n \"\n \n+test_expect_success \"v5 update-index prepare\" \"\n+\techo x >$file\n+\"\n+\n+test_perf_cleanup \"v5 update-index\" \"\n+\tgit update-index $file\n+\" \"\n+\tgit reset\n+\"\n+\n test_done\n-- \n1.8.4.2\n"},{"id":"231216","messageId":"CAPig+cSHcL62EW5z5n68jQcS4BWW9cZ=GqRwZaoyYM69NE55+w@mail.gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-21-git-send-email-t.gummerer@gmail.com","subject":"Re: [PATCH v4 20/24] introduce GIT_INDEX_VERSION environment variable","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2013-11-27T21:57:38Z","receivedAt":"2013-11-27T21:57:38Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Wed, Nov 27, 2013 at 7:00 AM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n> Respect a GIT_INDEX_VERSION environment variable, when a new index is\n> initialized.  Setting the environment variable will not cause existing\n> index files to be converted to another format for additional safety.\n>\n> Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n> ---\n> diff --git a/read-cache.c b/read-cache.c\n> index 46551af..04430e5 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1233,8 +1233,13 @@ static struct cache_entry *refresh_cache_entry(struct cache_entry *ce, int reall\n>  void initialize_index(struct index_state *istate, int version)\n>  {\n>         istate->initialized = 1;\n> -       if (!version)\n> -               version = INDEX_FORMAT_DEFAULT;\n> +       if (!version) {\n> +               char *envversion = getenv(\"GIT_INDEX_VERSION\");\n> +               if (!envversion)\n> +                       version = INDEX_FORMAT_DEFAULT;\n> +               else\n> +                       version = atoi(envversion);\n\nDo you want to check that atoi() returned a valid value and emit a\ndiagnostic if it did not?\n\n> +       }\n>         istate->version = version;\n>         set_istate_ops(istate);\n>  }\n> --\n> 1.8.4.2\n"},{"id":"231217","messageId":"xmqqk3ftsdk4.fsf@gitster.dls.corp.google.com","threadId":"35410","inReplyTo":"CAPig+cSHcL62EW5z5n68jQcS4BWW9cZ=GqRwZaoyYM69NE55+w@mail.gmail.com","subject":"Re: [PATCH v4 20/24] introduce GIT_INDEX_VERSION environment variable","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-11-27T22:08:27Z","receivedAt":"2013-11-27T22:08:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Eric Sunshine <sunshine@sunshineco.com> writes:\n\n> On Wed, Nov 27, 2013 at 7:00 AM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n>> Respect a GIT_INDEX_VERSION environment variable, when a new index is\n>> initialized.  Setting the environment variable will not cause existing\n>> index files to be converted to another format for additional safety.\n>>\n>> Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n>> ---\n>> diff --git a/read-cache.c b/read-cache.c\n>> index 46551af..04430e5 100644\n>> --- a/read-cache.c\n>> +++ b/read-cache.c\n>> @@ -1233,8 +1233,13 @@ static struct cache_entry *refresh_cache_entry(struct cache_entry *ce, int reall\n>>  void initialize_index(struct index_state *istate, int version)\n>>  {\n>>         istate->initialized = 1;\n>> -       if (!version)\n>> -               version = INDEX_FORMAT_DEFAULT;\n>> +       if (!version) {\n>> +               char *envversion = getenv(\"GIT_INDEX_VERSION\");\n>> +               if (!envversion)\n>> +                       version = INDEX_FORMAT_DEFAULT;\n>> +               else\n>> +                       version = atoi(envversion);\n>\n> Do you want to check that atoi() returned a valid value and emit a\n> diagnostic if it did not?\n\n\nGood eyes.\n\nWe use strtoul() for this kind of thing instead of atoi() for format\nchecking.  The code also needs to make sure that the value obtained\nthusly are among the versions that are supported.\n\nThanks.\n"},{"id":"231237","messageId":"20131128095615.GA25478@goose.lan","threadId":"35410","inReplyTo":"xmqqk3ftsdk4.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v4 20/24] introduce GIT_INDEX_VERSION environment variable","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-28T09:57:02Z","receivedAt":"2013-11-28T09:57:02Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 11/27, Junio C Hamano wrote:\n> Eric Sunshine <sunshine@sunshineco.com> writes:\n> \n> > On Wed, Nov 27, 2013 at 7:00 AM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n> >> Respect a GIT_INDEX_VERSION environment variable, when a new index is\n> >> initialized.  Setting the environment variable will not cause existing\n> >> index files to be converted to another format for additional safety.\n> >>\n> >> Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n> >> ---\n> >> diff --git a/read-cache.c b/read-cache.c\n> >> index 46551af..04430e5 100644\n> >> --- a/read-cache.c\n> >> +++ b/read-cache.c\n> >> @@ -1233,8 +1233,13 @@ static struct cache_entry *refresh_cache_entry(struct cache_entry *ce, int reall\n> >>  void initialize_index(struct index_state *istate, int version)\n> >>  {\n> >>         istate->initialized = 1;\n> >> -       if (!version)\n> >> -               version = INDEX_FORMAT_DEFAULT;\n> >> +       if (!version) {\n> >> +               char *envversion = getenv(\"GIT_INDEX_VERSION\");\n> >> +               if (!envversion)\n> >> +                       version = INDEX_FORMAT_DEFAULT;\n> >> +               else\n> >> +                       version = atoi(envversion);\n> >\n> > Do you want to check that atoi() returned a valid value and emit a\n> > diagnostic if it did not?\n> \n> \n> Good eyes.\n> \n> We use strtoul() for this kind of thing instead of atoi() for format\n> checking.  The code also needs to make sure that the value obtained\n> thusly are among the versions that are supported.\n> \n> Thanks.\n\nThanks both.  Will use strtoul and check the value in the re-roll.\n\n-- \nThomas\n"},{"id":"231292","messageId":"CACsJy8C5Px6d5dKOG8mbKYvLPHVFOBhJ8i5_6oRB33zB1Rmvhg@mail.gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-13-git-send-email-t.gummerer@gmail.com","subject":"Re: [PATCH v4 12/24] read-cache: read index-v5","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-11-30T09:17:35Z","receivedAt":"2013-11-30T09:17:35Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Nov 27, 2013 at 7:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -132,11 +141,17 @@ struct cache_entry {\n>         char name[FLEX_ARRAY]; /* more */\n>  };\n>\n> +#define CE_NAMEMASK  (0x0fff)\n\nCE_NAMEMASK is redefined in read-cache-v2.c in \"read-cache: move index\nv2 specific functions to their own file\". My gcc is smart enough to\nsee the two defines are about the same value and does not warn me. But\nwe should remove one (likely this one as I see no use of this macro\noutside read-cache-v2.c)\n\n>  #define CE_STAGEMASK (0x3000)\n>  #define CE_EXTENDED  (0x4000)\n>  #define CE_VALID     (0x8000)\n> +#define CE_SMUDGED   (0x0400) /* index v5 only flag */\n>  #define CE_STAGESHIFT 12\n>\n> +#define CONFLICT_CONFLICTED (0x8000)\n> +#define CONFLICT_STAGESHIFT 13\n> +#define CONFLICT_STAGEMASK (0x6000)\n> +\n>  /*\n>   * Range 0xFFFF0000 in ce_flags is divided into\n>   * two parts: in-memory flags and on-disk ones.\n\n> diff --git a/read-cache-v5.c b/read-cache-v5.c\n> new file mode 100644\n> index 0000000..9d8c8f0\n> --- /dev/null\n> +++ b/read-cache-v5.c\n> +static int read_index_v5(struct index_state *istate, void *mmap,\n> +                        unsigned long mmap_size, struct filter_opts *opts)\n> +{\n> +       unsigned int entry_offset, foffsetblock, nr = 0, *extoffsets;\n> +       unsigned int dir_offset, dir_table_offset;\n> +       int need_root = 0, i;\n> +       uint32_t *offset;\n> +       struct directory_entry *root_directory, *de, *last_de;\n> +       const char **paths = NULL;\n> +       struct pathspec adjusted_pathspec;\n> +       struct cache_header *hdr;\n> +       struct cache_header_v5 *hdr_v5;\n> +\n> +       hdr = mmap;\n> +       hdr_v5 = ptr_add(mmap, sizeof(*hdr));\n> +       istate->cache_alloc = alloc_nr(ntohl(hdr->hdr_entries));\n> +       istate->cache = xcalloc(istate->cache_alloc, sizeof(struct cache_entry *));\n> +       extoffsets = xcalloc(ntohl(hdr_v5->hdr_nextension), sizeof(int));\n> +       for (i = 0; i < ntohl(hdr_v5->hdr_nextension); i++) {\n> +               offset = ptr_add(mmap, sizeof(*hdr) + sizeof(*hdr_v5));\n> +               extoffsets[i] = htonl(*offset);\n> +       }\n> +\n> +       /* Skip size of the header + crc sum + size of offsets to extensions + size of offsets */\n> +       dir_offset = sizeof(*hdr) + sizeof(*hdr_v5) + ntohl(hdr_v5->hdr_nextension) * 4 + 4\n> +               + (ntohl(hdr_v5->hdr_ndir) + 1) * 4;\n> +       dir_table_offset = sizeof(*hdr) + sizeof(*hdr_v5) + ntohl(hdr_v5->hdr_nextension) * 4 + 4;\n> +       root_directory = read_directories(&dir_offset, &dir_table_offset,\n> +                                         mmap, mmap_size);\n> +\n> +       entry_offset = ntohl(hdr_v5->hdr_fblockoffset);\n> +       foffsetblock = dir_offset;\n> +\n> +       if (opts && opts->pathspec && opts->pathspec->nr) {\n> +               paths = xmalloc((opts->pathspec->nr + 1)*sizeof(char *));\n> +               paths[opts->pathspec->nr] = NULL;\n\nPut this statement here\n\nGUARD_PATHSPEC(opts->pathspec,\n      PATHSPEC_FROMTOP |\n      PATHSPEC_MAXDEPTH |\n      PATHSPEC_LITERAL |\n      PATHSPEC_GLOB |\n      PATHSPEC_ICASE);\n\nThis says the mentioned magic is safe in this code. New magic may or\nmay not be and needs to be checked (soonest by me, I'm going to add\nnegative pathspec and I'll need to look into how it should be handled\nin this code block).\n\n> +               for (i = 0; i < opts->pathspec->nr; i++) {\n> +                       char *super = strdup(opts->pathspec->items[i].match);\n> +                       int len = strlen(super);\n\nYou should only check as far as items[i].nowildcard_len, not strlen().\nThe rest could be wildcards and stuff and not so reliable.\n\n> +                       while (len && super[len - 1] == '/' && super[len - 2] == '/')\n> +                               super[--len] = '\\0'; /* strip all but one trailing slash */\n> +                       while (len && super[--len] != '/')\n> +                               ; /* scan backwards to next / */\n> +                       if (len >= 0)\n> +                               super[len--] = '\\0';\n> +                       if (len <= 0) {\n> +                               need_root = 1;\n> +                               break;\n> +                       }\n> +                       paths[i] = super;\n> +               }\n\nAnd maybe put the comment \"FIXME: consider merging this code with\ncreate_simplify() in dir.c\" somewhere. It's for me to look for things\nto do when I'm bored ;-)\n\n> +       }\n> +\n> +       if (!need_root)\n> +               parse_pathspec(&adjusted_pathspec, PATHSPEC_ALL_MAGIC, PATHSPEC_PREFER_CWD, NULL, paths);\n\nI would go with PATHSPEC_PREFER_FULL instead of _CWD as it's safer.\nLooking only at this function without caller context, it's hard to say\nif _CWD is the right choice.\n\n> +\n> +       de = root_directory;\n> +       last_de = de;\n\nThis statement is redundant. last_de is only used in one code block\nbelow and it's always re-initialized before entering the loop to skip\nsubdirs.\n\n> +       while (de) {\n> +               if (need_root ||\n> +                   match_pathspec_depth(&adjusted_pathspec, de->pathname, de->de_pathlen, 0, NULL)) {\n> +                       if (read_entries(istate, de, entry_offset,\n> +                                        mmap, mmap_size, &nr,\n> +                                        foffsetblock) < 0)\n> +                               return -1;\n> +               } else {\n> +                       last_de = de;\n> +                       for (i = 0; i < de->de_nsubtrees; i++) {\n> +                               de->sub[i]->next = last_de->next;\n> +                               last_de->next = de->sub[i];\n> +                               last_de = last_de->next;\n> +                       }\n> +               }\n> +               de = de->next;\n> +       }\n> +       free_directory_tree(root_directory);\n> +       istate->cache_nr = nr;\n> +       return 0;\n> +}\n-- \nDuy\n"},{"id":"231293","messageId":"CACsJy8D6q5Y4uLTcxD+9pY7U9_2qkO6o-_JeciWWAwJFrsqrzQ@mail.gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-10-git-send-email-t.gummerer@gmail.com","subject":"Re: [PATCH v4 09/24] ls-files.c: use index api","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-11-30T09:17:56Z","receivedAt":"2013-11-30T09:17:56Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Nov 27, 2013 at 7:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n> @@ -447,6 +463,7 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>         struct dir_struct dir;\n>         struct exclude_list *el;\n>         struct string_list exclude_list = STRING_LIST_INIT_NODUP;\n> +       struct filter_opts *opts = xmalloc(sizeof(*opts));\n>         struct option builtin_ls_files_options[] = {\n>                 { OPTION_CALLBACK, 'z', NULL, NULL, NULL,\n>                         N_(\"paths are separated with NUL character\"),\n> @@ -512,9 +529,6 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>                 prefix_len = strlen(prefix);\n>         git_config(git_default_config, NULL);\n>\n> -       if (read_cache() < 0)\n> -               die(\"index file corrupt\");\n> -\n>         argc = parse_options(argc, argv, prefix, builtin_ls_files_options,\n>                         ls_files_usage, 0);\n>         el = add_exclude_list(&dir, EXC_CMDL, \"--exclude option\");\n> @@ -550,6 +564,24 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>                        PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n>                        prefix, argv);\n>\n> +       if (!with_tree && !needs_trailing_slash_stripped()) {\n> +               memset(opts, 0, sizeof(*opts));\n> +               opts->pathspec = &pathspec;\n> +               opts->read_staged = 1;\n> +               if (show_resolve_undo)\n> +                       opts->read_resolve_undo = 1;\n> +               if (read_cache_filtered(opts) < 0)\n> +                       die(\"index file corrupt\");\n> +       } else {\n> +               if (read_cache() < 0)\n> +                       die(\"index file corrupt\");\n> +               parse_pathspec(&pathspec, 0,\n> +                              PATHSPEC_PREFER_CWD |\n> +                              PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n> +                              prefix, argv);\n\nSo we ran parse_pathspec() once (not shown in the context), if found\ntrailing slashes, we read full cache and rerun parse_pathspec()\nbecause the index is not loaded on the first run.\n\nThis is fine. Just a note for future improvement: as _SLASH_CHEAP only\nneeds to look at a few entries with cache_name_pos(), we could take\nadvantage of v5 to peek individual entries (or in v2, load full cache\nfirst). Nothing needs to be done now, I think we have not decided\nwhether to combine _SLASH_CHEAP and _SLASH_EXPENSIVE into one.\n\n> +       }\n> +\n>         /* Find common prefix for all pathspec's */\n>         max_prefix = common_prefix(&pathspec);\n>         max_prefix_len = max_prefix ? strlen(max_prefix) : 0;\n-- \nDuy\n"},{"id":"231295","messageId":"CACsJy8Bti465Lv9U+m=VnqvGQAFk=WxCsmcGmdB0=NLesF5Mtw@mail.gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-24-git-send-email-t.gummerer@gmail.com","subject":"Re: [PATCH v4 23/24] POC for partial writing","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-11-30T09:58:55Z","receivedAt":"2013-11-30T09:58:55Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Nov 27, 2013 at 7:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n> This makes update-index use both partial reading and partial writing.\n> Partial reading is only used no option other than the paths is passed to\n> the command.\n>\n> This passes the test suite,\n\nJust checking, the test suite was run with TEST_GIT_INDEX_VERSION=5, right?\n\n> but doesn't behave correctly when a write\n> fails.  A log should be written to the lock file, in order to be able to\n> recover if a write fails.\n\n>From the API point of view this looks nice (you should have hidden\nneeds_write = 1 in cache_invalidate_path and change_cache_version\nthough)  We could support partial file removal too by marking removed\nfiles \"removed\", but that impacts the reading code and may have bad\ninteraction with cache_invalidate_path/needs_rewrite. Probably not\nworth the effort until someone shows us they remove stuff often.\n\n> ---\n>  builtin/update-index.c |  43 +++++++++++---\n>  cache-tree.c           |  13 +++++\n>  cache-tree.h           |   1 +\n>  cache.h                |  27 ++++++++-\n>  lockfile.c             |   2 +-\n>  read-cache-v2.c        |   2 +\n>  read-cache-v5.c        | 154 ++++++++++++++++++++++++++++++++++++++++---------\n>  read-cache.c           |  30 ++++++++++\n>  read-cache.h           |   1 +\n>  resolve-undo.c         |   1 +\n>  10 files changed, 237 insertions(+), 37 deletions(-)\n>\n> diff --git a/builtin/update-index.c b/builtin/update-index.c\n> index 8b3f7a0..69f0949 100644\n> --- a/builtin/update-index.c\n> +++ b/builtin/update-index.c\n> @@ -56,6 +56,7 @@ static int mark_ce_flags(const char *path, int flag, int mark)\n>                 else\n>                         active_cache[pos]->ce_flags &= ~flag;\n>                 cache_tree_invalidate_path(active_cache_tree, path);\n> +               the_index.needs_rewrite = 1;\n>                 active_cache_changed = 1;\n>                 return 0;\n>         }\n> @@ -99,6 +100,8 @@ static int add_one_path(const struct cache_entry *old, const char *path, int len\n>         memcpy(ce->name, path, len);\n>         ce->ce_flags = create_ce_flags(0);\n>         ce->ce_namelen = len;\n> +       if (old)\n> +               ce->entry_pos = old->entry_pos;\n>         fill_stat_cache_info(ce, st);\n>         ce->ce_mode = ce_mode_from_stat(old, st->st_mode);\n>\n> @@ -268,6 +271,7 @@ static void chmod_path(int flip, const char *path)\n>                 goto fail;\n>         }\n>         cache_tree_invalidate_path(active_cache_tree, path);\n> +       the_index.needs_rewrite = 1;\n>         active_cache_changed = 1;\n>         report(\"chmod %cx '%s'\", flip, path);\n>         return;\n> @@ -706,15 +710,18 @@ static int reupdate_callback(struct parse_opt_ctx_t *ctx,\n>\n>  int cmd_update_index(int argc, const char **argv, const char *prefix)\n>  {\n> -       int newfd, entries, has_errors = 0, line_termination = '\\n';\n> +       int newfd, has_errors = 0, line_termination = '\\n';\n>         int read_from_stdin = 0;\n>         int prefix_length = prefix ? strlen(prefix) : 0;\n>         int preferred_index_format = 0;\n>         char set_executable_bit = 0;\n>         struct refresh_params refresh_args = {0, &has_errors};\n>         int lock_error = 0;\n> +       struct filter_opts opts;\n> +       struct pathspec pathspec;\n>         struct lock_file *lock_file;\n>         struct parse_opt_ctx_t ctx;\n> +       int i, needs_full_read = 0;\n>         int parseopt_state = PARSE_OPT_UNKNOWN;\n>         struct option options[] = {\n>                 OPT_BIT('q', NULL, &refresh_args.flags,\n> @@ -810,9 +817,23 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>         if (newfd < 0)\n>                 lock_error = errno;\n>\n> -       entries = read_cache();\n> -       if (entries < 0)\n> -               die(\"cache corrupted\");\n> +       for (i = 0; i < argc; i++) {\n> +               if (!prefixcmp(argv[i], \"--\"))\n> +                       needs_full_read = 1;\n> +       }\n> +       if (!needs_full_read) {\n> +               memset(&opts, 0, sizeof(struct filter_opts));\n> +               parse_pathspec(&pathspec, 0,\n> +                              PATHSPEC_PREFER_CWD |\n> +                              PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n> +                              prefix, argv + 1);\n> +               opts.pathspec = &pathspec;\n> +               if (read_cache_filtered(&opts) < 0)\n> +                       die(\"cache corrupted\");\n> +       } else {\n> +               if (read_cache() < 0)\n> +                       die(\"cache corrupted\");\n> +       }\n>\n>         /*\n>          * Custom copy of parse_options() because we want to handle\n> @@ -862,6 +883,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>                             preferred_index_format,\n>                             INDEX_FORMAT_LB, INDEX_FORMAT_UB);\n>\n> +               the_index.needs_rewrite = 1;\n>                 active_cache_changed = 1;\n>                 change_cache_version(preferred_index_format);\n>         }\n> @@ -890,17 +912,22 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>         }\n>\n>         if (active_cache_changed) {\n> +               int r;\n>                 if (newfd < 0) {\n>                         if (refresh_args.flags & REFRESH_QUIET)\n>                                 exit(128);\n>                         unable_to_lock_index_die(get_index_file(), lock_error);\n>                 }\n> -               if (write_cache(newfd, active_cache, active_nr) ||\n> -                   commit_locked_index(lock_file))\n> +               r = write_cache_partial(newfd);\n> +               if (r < 0)\n>                         die(\"Unable to write new index file\");\n> +               else if (r == 0)\n> +                       commit_lock_file(lock_file);\n> +               else\n> +                       remove_lock_file();\n> +       } else {\n> +               rollback_lock_file(lock_file);\n>         }\n>\n> -       rollback_lock_file(lock_file);\n> -\n>         return has_errors ? 1 : 0;\n>  }\n> diff --git a/cache-tree.c b/cache-tree.c\n> index 1209732..a3d18bb 100644\n> --- a/cache-tree.c\n> +++ b/cache-tree.c\n> @@ -123,6 +123,15 @@ void cache_tree_invalidate_path(struct cache_tree *it, const char *path)\n>                 return;\n>         slash = strchr(path, '/');\n>         it->entry_count = -1;\n> +       /*\n> +        * Mark the cache_tree directory entry as invalid too. The\n> +        * entry_count defines if the tree is valid, so we don't need\n> +        * to reset any other field.\n> +        */\n> +       if (it->de_ref) {\n> +               it->de_ref->de_nentries = -1;\n> +               it->de_ref->changed = 1;\n> +       }\n>         if (!slash) {\n>                 int pos;\n>                 namelen = strlen(path);\n> @@ -140,6 +149,10 @@ void cache_tree_invalidate_path(struct cache_tree *it, const char *path)\n>                                 sizeof(struct cache_tree_sub *) *\n>                                 (it->subtree_nr - pos - 1));\n>                         it->subtree_nr--;\n> +                       if (it->de_ref) {\n> +                               it->de_ref->de_nsubtrees--;\n> +                               it->de_ref->changed = 1;\n> +                       }\n>                 }\n>                 return;\n>         }\n> diff --git a/cache-tree.h b/cache-tree.h\n> index 9818926..eaf14a9 100644\n> --- a/cache-tree.h\n> +++ b/cache-tree.h\n> @@ -18,6 +18,7 @@ struct cache_tree {\n>         unsigned char sha1[20];\n>         int subtree_nr;\n>         int subtree_alloc;\n> +       struct directory_entry *de_ref;\n>         struct cache_tree_sub **down;\n>  };\n>\n> diff --git a/cache.h b/cache.h\n> index 71b98cf..1a634dc 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -137,11 +137,31 @@ struct cache_entry {\n>         unsigned int ce_namelen;\n>         unsigned char sha1[20];\n>         uint32_t ce_stat_crc;\n> +       unsigned int entry_pos;\n> +       unsigned int changed;\n>         struct cache_entry *next; /* used by name_hash */\n>         struct cache_entry *next_ce;\n>         char name[FLEX_ARRAY]; /* more */\n>  };\n>\n> +struct directory_entry {\n> +       struct directory_entry **sub;\n> +       struct directory_entry *next;\n> +       struct directory_entry *next_hash;\n> +       struct cache_entry *ce;\n> +       struct cache_entry *ce_last;\n> +       uint32_t de_foffset;\n> +       uint32_t de_nsubtrees;\n> +       uint32_t de_nfiles;\n> +       uint32_t de_nentries;\n> +       unsigned char sha1[20];\n> +       uint16_t de_flags;\n> +       uint32_t de_pathlen;\n> +       uint32_t entry_pos;\n> +       unsigned int changed;\n> +       char pathname[FLEX_ARRAY];\n> +};\n> +\n>  #define CE_NAMEMASK  (0x0fff)\n>  #define CE_STAGEMASK (0x3000)\n>  #define CE_EXTENDED  (0x4000)\n> @@ -317,13 +337,15 @@ struct filter_opts {\n>\n>  struct index_state {\n>         struct cache_entry **cache;\n> +       struct directory_entry *root_directory;\n>         unsigned int version;\n>         unsigned int cache_nr, cache_alloc, cache_changed;\n>         struct string_list *resolve_undo;\n>         struct cache_tree *cache_tree;\n>         struct cache_time timestamp;\n>         unsigned name_hash_initialized : 1,\n> -                initialized : 1;\n> +                initialized : 1,\n> +                needs_rewrite : 1;\n>         struct hash_table name_hash;\n>         struct hash_table dir_hash;\n>         struct index_ops *ops;\n> @@ -353,6 +375,7 @@ extern void free_name_hash(struct index_state *istate);\n>  #define is_cache_unborn() is_index_unborn(&the_index)\n>  #define read_cache_unmerged() read_index_unmerged(&the_index)\n>  #define write_cache(newfd, cache, entries) write_index(&the_index, (newfd))\n> +#define write_cache_partial(newfd) write_index_partial(&the_index, (newfd))\n>  #define discard_cache() discard_index(&the_index)\n>  #define unmerged_cache() unmerged_index(&the_index)\n>  #define cache_name_pos(name, namelen) index_name_pos(&the_index,(name),(namelen))\n> @@ -529,6 +552,7 @@ extern int read_index_from(struct index_state *, const char *path);\n>  extern int is_index_unborn(struct index_state *);\n>  extern int read_index_unmerged(struct index_state *);\n>  extern int write_index(struct index_state *, int newfd);\n> +extern int write_index_partial(struct index_state *, int newfd);\n>  extern int discard_index(struct index_state *);\n>  extern int unmerged_index(const struct index_state *);\n>  extern int verify_path(const char *path);\n> @@ -613,6 +637,7 @@ extern NORETURN void unable_to_lock_index_die(const char *path, int err);\n>  extern int hold_lock_file_for_update(struct lock_file *, const char *path, int);\n>  extern int hold_lock_file_for_append(struct lock_file *, const char *path, int);\n>  extern int commit_lock_file(struct lock_file *);\n> +extern void remove_lock_file(void);\n>  extern void update_index_if_able(struct index_state *, struct lock_file *);\n>\n>  extern int hold_locked_index(struct lock_file *, int);\n> diff --git a/lockfile.c b/lockfile.c\n> index 8fbcb6a..c150e5c 100644\n> --- a/lockfile.c\n> +++ b/lockfile.c\n> @@ -7,7 +7,7 @@\n>  static struct lock_file *lock_file_list;\n>  static const char *alternate_index_output;\n>\n> -static void remove_lock_file(void)\n> +void remove_lock_file(void)\n>  {\n>         pid_t me = getpid();\n>\n> diff --git a/read-cache-v2.c b/read-cache-v2.c\n> index f884c10..1fec892 100644\n> --- a/read-cache-v2.c\n> +++ b/read-cache-v2.c\n> @@ -555,5 +555,7 @@ struct index_ops v2_ops = {\n>         match_stat_basic,\n>         verify_hdr,\n>         read_index_v2,\n> +       write_index_v2,\n> +       /* Partial writing is the same as writing the full index for v2 */\n>         write_index_v2\n>  };\n> diff --git a/read-cache-v5.c b/read-cache-v5.c\n> index a5e9b5a..13436a3 100644\n> --- a/read-cache-v5.c\n> +++ b/read-cache-v5.c\n> @@ -20,22 +20,6 @@ struct extension_header {\n>         uint32_t crc;\n>  };\n>\n> -struct directory_entry {\n> -       struct directory_entry **sub;\n> -       struct directory_entry *next;\n> -       struct directory_entry *next_hash;\n> -       struct cache_entry *ce;\n> -       struct cache_entry *ce_last;\n> -       uint32_t de_foffset;\n> -       uint32_t de_nsubtrees;\n> -       uint32_t de_nfiles;\n> -       uint32_t de_nentries;\n> -       unsigned char sha1[20];\n> -       uint16_t de_flags;\n> -       uint32_t de_pathlen;\n> -       char pathname[FLEX_ARRAY];\n> -};\n> -\n>  struct conflict_part {\n>         struct conflict_part *next;\n>         uint16_t flags;\n> @@ -246,7 +230,7 @@ static struct directory_entry *read_directories(unsigned int *dir_offset,\n>                 offsetof(struct ondisk_directory_entry, name) - 5;\n>         disk_de = ptr_add(mmap, *dir_offset);\n>         de = directory_entry_from_ondisk(disk_de, len);\n> -\n> +       de->entry_pos = *dir_offset;\n>         data_len = len + 1 + offsetof(struct ondisk_directory_entry, name);\n>         filecrc = ptr_add(mmap, *dir_offset + data_len);\n>         if (!check_crc32(0, ptr_add(mmap, *dir_offset), data_len, ntoh_l(*filecrc)))\n> @@ -281,6 +265,7 @@ static int read_entry(struct cache_entry **ce, char *pathname, size_t pathlen,\n>         entry_offset = first_entry_offset + ntoh_l(*beginning);\n>         disk_ce = ptr_add(mmap, entry_offset);\n>         *ce = cache_entry_from_ondisk(disk_ce, pathname, len, pathlen);\n> +       (*ce)->entry_pos = entry_offset;\n>         filecrc = ptr_add(mmap, entry_offset + len + 1 + sizeof(*disk_ce));\n>         if (!check_crc32(0,\n>                 ptr_add(mmap, entry_offset), len + 1 + sizeof(*disk_ce),\n> @@ -439,6 +424,7 @@ static struct cache_tree *convert_one(struct directory_entry *de)\n>\n>         it = cache_tree();\n>         it->entry_count = de->de_nentries;\n> +       it->de_ref = de;\n>         if (0 <= it->entry_count)\n>                 hashcpy(it->sha1, de->sha1);\n>\n> @@ -523,14 +509,6 @@ static int read_entries(struct index_state *istate, struct directory_entry *de,\n>         return 0;\n>  }\n>\n> -static void free_directory_tree(struct directory_entry *de) {\n> -       int i;\n> -\n> -       for (i = 0; i < de->de_pathlen; i++)\n> -               free_directory_tree(de->sub[i]);\n> -       free(de);\n> -}\n> -\n>  /*\n>   * Read an index-v5 file filtered by the filter_opts.   If opts is NULL,\n>   * everything will be read.\n> @@ -626,7 +604,7 @@ static int read_index_v5(struct index_state *istate, void *mmap,\n>                 }\n>         }\n>         istate->cache_tree = cache_tree_convert_v5(root_directory);\n> -       free_directory_tree(root_directory);\n> +       istate->root_directory = root_directory;\n>         istate->cache_nr = nr;\n>         return 0;\n>  }\n> @@ -696,6 +674,7 @@ static void ce_smudge_racily_clean_entry(struct cache_entry *ce)\n>          * that hasn't changed checking the sha1.\n>          */\n>         ce->ce_flags |= CE_SMUDGED;\n> +       ce->changed = 1;\n>  }\n>\n>  static char *super_directory(char *filename)\n> @@ -1231,6 +1210,103 @@ static int write_resolve_undo(struct index_state *istate,\n>         return 0;\n>  }\n>\n> +static int write_ce_if_necessary(struct cache_entry *ce, void *cb_data)\n> +{\n> +       int *fdx = cb_data, pathlen, size;\n> +       int fd = *fdx;\n> +       char *dir;\n> +       struct ondisk_cache_entry *ondisk;\n> +       uint32_t crc;\n> +\n> +       assert(ce->entry_pos != 0);\n> +       /* TODO I'm just using the_index out of lazyness here */\n> +       if (!ce_uptodate(ce) && is_racy_timestamp(&the_index, ce))\n> +               ce_smudge_racily_clean_entry(ce);\n> +       if (!ce->changed)\n> +               return 0;\n> +       if (is_null_sha1(ce->sha1)) {\n> +               static const char msg[] = \"cache entry has null sha1: %s\";\n> +               static int allow = -1;\n> +\n> +               if (allow < 0)\n> +                       allow = git_env_bool(\"GIT_ALLOW_NULL_SHA1\", 0);\n> +               if (allow)\n> +                       warning(msg, ce->name);\n> +               else\n> +                       return error(msg, ce->name);\n> +       }\n> +       dir = super_directory(ce->name);\n> +       pathlen = dir ? strlen(dir) + 1 : 0;\n> +       size = offsetof(struct ondisk_cache_entry, name) +\n> +               ce_namelen(ce) - pathlen + 1;\n> +       ondisk = xmalloc(size);\n> +\n> +       crc = 0;\n> +       ondisk_from_cache_entry(ce, ondisk, pathlen);\n> +       if (lseek(fd, ce->entry_pos, SEEK_SET) < ce->entry_pos)\n> +               die(\"eror ce seeking\");\n> +       if (ce_write(&crc, fd, ondisk, size) < 0)\n> +               return -1;\n> +       crc = htonl(crc);\n> +       if (ce_write(NULL, fd, &crc, 4) < 0)\n> +               return -1;\n> +       return ce_flush(fd);\n> +}\n> +\n> +static void ondisk_from_directory_entry_partial(struct directory_entry *de,\n> +                                               struct ondisk_directory_entry *ondisk)\n> +{\n> +       ondisk->foffset   = htonl(de->de_foffset);\n> +       ondisk->nsubtrees = htonl(de->de_nsubtrees);\n> +       ondisk->nfiles    = htonl(de->de_nfiles);\n> +       ondisk->nentries  = htonl(de->de_nentries);\n> +       hashcpy(ondisk->sha1, de->sha1);\n> +       ondisk->flags     = htons(de->de_flags);\n> +       if (de->de_pathlen == 0) {\n> +               memcpy(ondisk->name, \"\\0\", 1);\n> +       } else {\n> +               memcpy(ondisk->name, de->pathname, de->de_pathlen);\n> +               memcpy(ondisk->name + de->de_pathlen - 1, \"/\\0\", 2);\n> +       }\n> +}\n> +\n> +static int write_directories_partial(struct directory_entry *de, int fd)\n> +{\n> +       int ondisk_size = offsetof(struct ondisk_directory_entry, name);\n> +       int size = ondisk_size + de->de_pathlen + 1;\n> +       int i;\n> +       uint32_t crc;\n> +       struct ondisk_directory_entry *ondisk;\n> +\n> +       if (de->changed) {\n> +               crc = 0;\n> +               ondisk = xmalloc(size);\n> +               ondisk_from_directory_entry_partial(de, ondisk);\n> +               if (lseek(fd, de->entry_pos, SEEK_SET) < de->entry_pos)\n> +                       die(\"error directory seeking\");;\n> +               if (ce_write(&crc, fd, ondisk, size) < 0)\n> +                       return -1;\n> +               crc = htonl(crc);\n> +               if (ce_write(NULL, fd, &crc, 4) < 0)\n> +                       return -1;\n> +               free(ondisk);\n> +               if (ce_flush(fd) < 0)\n> +                       return -1;\n> +       }\n> +       for (i = 0; i < de->de_nsubtrees; i++) {\n> +               if (write_directories_partial(de->sub[i], fd) < 0)\n> +                       return -1;\n> +       }\n> +       return 0;\n> +}\n> +\n> +static int write_partial(struct index_state *istate, int fd)\n> +{\n> +       write_directories_partial(istate->root_directory, fd);\n> +\n> +       return for_each_index_entry(istate, write_ce_if_necessary, &fd);\n> +}\n> +\n>  static int write_index_v5(struct index_state *istate, int newfd)\n>  {\n>         struct cache_header hdr;\n> @@ -1296,9 +1372,33 @@ static int write_index_v5(struct index_state *istate, int newfd)\n>         return ce_flush(newfd);\n>  }\n>\n> +static int write_index_partial_v5(struct index_state *istate, int newfd)\n> +{\n> +       int fd;\n> +       char *path = get_index_file();\n> +\n> +       if (istate->needs_rewrite || istate->cache_nr == 0)\n> +               return write_index_v5(istate, newfd);\n> +       if (istate->filter_opts && istate->needs_rewrite)\n> +               die(\"BUG: cannot write a partially read index\");\n> +       fd = open(path, O_RDWR, 0666);\n> +       if (fd < 0) {\n> +               if (errno == ENOENT)\n> +                       die(\"no index file exists cannot do a partial write\");\n> +               die_errno(\"index file opening for writing failed\");\n> +       }\n> +\n> +       if (write_partial(istate, fd) < 0)\n> +               return -1;\n> +       if (ce_flush(fd) < 0)\n> +               return -1;\n> +       return 1;\n> +}\n> +\n>  struct index_ops v5_ops = {\n>         match_stat_basic,\n>         verify_hdr,\n>         read_index_v5,\n> -       write_index_v5\n> +       write_index_v5,\n> +       write_index_partial_v5\n>  };\n> diff --git a/read-cache.c b/read-cache.c\n> index 04430e5..1cad0e2 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -32,6 +32,9 @@ static void replace_index_entry(struct index_state *istate, int nr, struct cache\n>\n>         remove_name_hash(istate, old);\n>         set_index_entry(istate, nr, ce);\n> +       ce->changed = 1;\n> +       if (ce->entry_pos == 0)\n> +               istate->needs_rewrite = 1;\n>         istate->cache_changed = 1;\n>  }\n>\n> @@ -494,6 +497,7 @@ int remove_index_entry_at(struct index_state *istate, int pos)\n>\n>         record_resolve_undo(istate, ce);\n>         remove_name_hash(istate, ce);\n> +       istate->needs_rewrite = 1;\n>         istate->cache_changed = 1;\n>         istate->cache_nr--;\n>         if (pos >= istate->cache_nr)\n> @@ -520,6 +524,7 @@ void remove_marked_cache_entries(struct index_state *istate)\n>                 else\n>                         ce_array[j++] = ce_array[i];\n>         }\n> +       istate->needs_rewrite = 1;\n>         istate->cache_changed = 1;\n>         istate->cache_nr = j;\n>  }\n> @@ -1024,6 +1029,7 @@ int add_index_entry(struct index_state *istate, struct cache_entry *ce, int opti\n>                         istate->cache + pos,\n>                         (istate->cache_nr - pos - 1) * sizeof(ce));\n>         set_index_entry(istate, pos, ce);\n> +       istate->needs_rewrite = 1;\n>         istate->cache_changed = 1;\n>         return 0;\n>  }\n> @@ -1108,6 +1114,8 @@ static struct cache_entry *refresh_cache_ent(struct index_state *istate,\n>         size = ce_size(ce);\n>         updated = xmalloc(size);\n>         memcpy(updated, ce, size);\n> +       updated->changed = 1;\n> +       updated->entry_pos = ce->entry_pos;\n>         fill_stat_cache_info(updated, &st);\n>         /*\n>          * If ignore_valid is not set, we should leave CE_VALID bit\n> @@ -1201,6 +1209,8 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n>                                  * means the index is not valid anymore.\n>                                  */\n>                                 ce->ce_flags &= ~CE_VALID;\n> +                               /* TODO: remove this maybe? */\n> +                               istate->needs_rewrite = 1;\n>                                 istate->cache_changed = 1;\n>                         }\n>                         if (quiet)\n> @@ -1241,6 +1251,8 @@ void initialize_index(struct index_state *istate, int version)\n>                         version = atoi(envversion);\n>         }\n>         istate->version = version;\n> +       istate->needs_rewrite = 0;\n> +       istate->root_directory = NULL;\n>         set_istate_ops(istate);\n>  }\n>\n> @@ -1427,6 +1439,16 @@ int is_index_unborn(struct index_state *istate)\n>         return (!istate->cache_nr && !istate->timestamp.sec);\n>  }\n>\n> +static void free_directory_tree(struct directory_entry *de) {\n> +       int i;\n> +\n> +       if (!de)\n> +               return;\n> +       for (i = 0; i < de->de_pathlen; i++)\n> +               free_directory_tree(de->sub[i]);\n> +       free(de);\n> +}\n> +\n>  int discard_index(struct index_state *istate)\n>  {\n>         int i;\n> @@ -1435,6 +1457,7 @@ int discard_index(struct index_state *istate)\n>                 free(istate->cache[i]);\n>         resolve_undo_clear_index(istate);\n>         istate->cache_nr = 0;\n> +       istate->needs_rewrite = 0;\n>         istate->cache_changed = 0;\n>         istate->timestamp.sec = 0;\n>         istate->timestamp.nsec = 0;\n> @@ -1446,6 +1469,8 @@ int discard_index(struct index_state *istate)\n>         istate->cache_alloc = 0;\n>         istate->ops = NULL;\n>         istate->filter_opts = NULL;\n> +       free_directory_tree(istate->root_directory);\n> +       istate->root_directory = NULL;\n>         return 0;\n>  }\n>\n> @@ -1491,6 +1516,11 @@ int write_index(struct index_state *istate, int newfd)\n>         return istate->ops->write_index(istate, newfd);\n>  }\n>\n> +int write_index_partial(struct index_state *istate, int newfd)\n> +{\n> +       return istate->ops->write_index_partial(istate, newfd);\n> +}\n> +\n>  /*\n>   * Read the index file that is potentially unmerged into given\n>   * index_state, dropping any unmerged entries.  Returns true if\n> diff --git a/read-cache.h b/read-cache.h\n> index 9d66df6..e7f36ae 100644\n> --- a/read-cache.h\n> +++ b/read-cache.h\n> @@ -31,6 +31,7 @@ struct index_ops {\n>         int (*read_index)(struct index_state *istate, void *mmap, unsigned long mmap_size,\n>                           struct filter_opts *opts);\n>         int (*write_index)(struct index_state *istate, int newfd);\n> +       int (*write_index_partial)(struct index_state *istate, int newfd);\n>  };\n>\n>  extern struct index_ops v2_ops;\n> diff --git a/resolve-undo.c b/resolve-undo.c\n> index c09b006..c496c20 100644\n> --- a/resolve-undo.c\n> +++ b/resolve-undo.c\n> @@ -110,6 +110,7 @@ void resolve_undo_clear_index(struct index_state *istate)\n>         string_list_clear(resolve_undo, 1);\n>         free(resolve_undo);\n>         istate->resolve_undo = NULL;\n> +       istate->needs_rewrite = 1;\n>         istate->cache_changed = 1;\n>  }\n>\n> --\n> 1.8.4.2\n>\n\n\n\n-- \nDuy\n"},{"id":"231296","messageId":"87zjomxjul.fsf@gmail.com","threadId":"35410","inReplyTo":"CACsJy8D6q5Y4uLTcxD+9pY7U9_2qkO6o-_JeciWWAwJFrsqrzQ@mail.gmail.com","subject":"Re: [PATCH v4 09/24] ls-files.c: use index api","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-30T10:30:26Z","receivedAt":"2013-11-30T10:30:26Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> On Wed, Nov 27, 2013 at 7:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n>> @@ -447,6 +463,7 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>>         struct dir_struct dir;\n>>         struct exclude_list *el;\n>>         struct string_list exclude_list = STRING_LIST_INIT_NODUP;\n>> +       struct filter_opts *opts = xmalloc(sizeof(*opts));\n>>         struct option builtin_ls_files_options[] = {\n>>                 { OPTION_CALLBACK, 'z', NULL, NULL, NULL,\n>>                         N_(\"paths are separated with NUL character\"),\n>> @@ -512,9 +529,6 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>>                 prefix_len = strlen(prefix);\n>>         git_config(git_default_config, NULL);\n>>\n>> -       if (read_cache() < 0)\n>> -               die(\"index file corrupt\");\n>> -\n>>         argc = parse_options(argc, argv, prefix, builtin_ls_files_options,\n>>                         ls_files_usage, 0);\n>>         el = add_exclude_list(&dir, EXC_CMDL, \"--exclude option\");\n>> @@ -550,6 +564,24 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>>                        PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n>>                        prefix, argv);\n>>\n>> +       if (!with_tree && !needs_trailing_slash_stripped()) {\n>> +               memset(opts, 0, sizeof(*opts));\n>> +               opts->pathspec = &pathspec;\n>> +               opts->read_staged = 1;\n>> +               if (show_resolve_undo)\n>> +                       opts->read_resolve_undo = 1;\n>> +               if (read_cache_filtered(opts) < 0)\n>> +                       die(\"index file corrupt\");\n>> +       } else {\n>> +               if (read_cache() < 0)\n>> +                       die(\"index file corrupt\");\n>> +               parse_pathspec(&pathspec, 0,\n>> +                              PATHSPEC_PREFER_CWD |\n>> +                              PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n>> +                              prefix, argv);\n>\n> So we ran parse_pathspec() once (not shown in the context), if found\n> trailing slashes, we read full cache and rerun parse_pathspec()\n> because the index is not loaded on the first run.\n>\n> This is fine. Just a note for future improvement: as _SLASH_CHEAP only\n> needs to look at a few entries with cache_name_pos(), we could take\n> advantage of v5 to peek individual entries (or in v2, load full cache\n> first). Nothing needs to be done now, I think we have not decided\n> whether to combine _SLASH_CHEAP and _SLASH_EXPENSIVE into one.\n\nYes that makes sense.  Adding the ability to search for path entries\nwithout reading the whole or part of the index was something I was\nthinking about, but didn't have time to do so yet.  I'll add this to my\nlist of possible future improvements.\n\n>> +       }\n>> +\n>>         /* Find common prefix for all pathspec's */\n>>         max_prefix = common_prefix(&pathspec);\n>>         max_prefix_len = max_prefix ? strlen(max_prefix) : 0;\n> -- \n> Duy\n"},{"id":"231297","messageId":"87wqjqxjdr.fsf@gmail.com","threadId":"35410","inReplyTo":"CACsJy8C5Px6d5dKOG8mbKYvLPHVFOBhJ8i5_6oRB33zB1Rmvhg@mail.gmail.com","subject":"Re: [PATCH v4 12/24] read-cache: read index-v5","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-30T10:40:32Z","receivedAt":"2013-11-30T10:40:32Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> On Wed, Nov 27, 2013 at 7:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n>> --- a/cache.h\n>> +++ b/cache.h\n>> @@ -132,11 +141,17 @@ struct cache_entry {\n>>         char name[FLEX_ARRAY]; /* more */\n>>  };\n>>\n>> +#define CE_NAMEMASK  (0x0fff)\n>\n> CE_NAMEMASK is redefined in read-cache-v2.c in \"read-cache: move index\n> v2 specific functions to their own file\". My gcc is smart enough to\n> see the two defines are about the same value and does not warn me. But\n> we should remove one (likely this one as I see no use of this macro\n> outside read-cache-v2.c)\n\nThanks for catching that, there's no need to have it here.  I'll remove\nit in the re-roll.\n\n>>  #define CE_STAGEMASK (0x3000)\n>>  #define CE_EXTENDED  (0x4000)\n>>  #define CE_VALID     (0x8000)\n>> +#define CE_SMUDGED   (0x0400) /* index v5 only flag */\n>>  #define CE_STAGESHIFT 12\n>>\n>> +#define CONFLICT_CONFLICTED (0x8000)\n>> +#define CONFLICT_STAGESHIFT 13\n>> +#define CONFLICT_STAGEMASK (0x6000)\n>> +\n>>  /*\n>>   * Range 0xFFFF0000 in ce_flags is divided into\n>>   * two parts: in-memory flags and on-disk ones.\n>\n>> diff --git a/read-cache-v5.c b/read-cache-v5.c\n>> new file mode 100644\n>> index 0000000..9d8c8f0\n>> --- /dev/null\n>> +++ b/read-cache-v5.c\n>> +static int read_index_v5(struct index_state *istate, void *mmap,\n>> +                        unsigned long mmap_size, struct filter_opts *opts)\n>> +{\n>> +       unsigned int entry_offset, foffsetblock, nr = 0, *extoffsets;\n>> +       unsigned int dir_offset, dir_table_offset;\n>> +       int need_root = 0, i;\n>> +       uint32_t *offset;\n>> +       struct directory_entry *root_directory, *de, *last_de;\n>> +       const char **paths = NULL;\n>> +       struct pathspec adjusted_pathspec;\n>> +       struct cache_header *hdr;\n>> +       struct cache_header_v5 *hdr_v5;\n>> +\n>> +       hdr = mmap;\n>> +       hdr_v5 = ptr_add(mmap, sizeof(*hdr));\n>> +       istate->cache_alloc = alloc_nr(ntohl(hdr->hdr_entries));\n>> +       istate->cache = xcalloc(istate->cache_alloc, sizeof(struct cache_entry *));\n>> +       extoffsets = xcalloc(ntohl(hdr_v5->hdr_nextension), sizeof(int));\n>> +       for (i = 0; i < ntohl(hdr_v5->hdr_nextension); i++) {\n>> +               offset = ptr_add(mmap, sizeof(*hdr) + sizeof(*hdr_v5));\n>> +               extoffsets[i] = htonl(*offset);\n>> +       }\n>> +\n>> +       /* Skip size of the header + crc sum + size of offsets to extensions + size of offsets */\n>> +       dir_offset = sizeof(*hdr) + sizeof(*hdr_v5) + ntohl(hdr_v5->hdr_nextension) * 4 + 4\n>> +               + (ntohl(hdr_v5->hdr_ndir) + 1) * 4;\n>> +       dir_table_offset = sizeof(*hdr) + sizeof(*hdr_v5) + ntohl(hdr_v5->hdr_nextension) * 4 + 4;\n>> +       root_directory = read_directories(&dir_offset, &dir_table_offset,\n>> +                                         mmap, mmap_size);\n>> +\n>> +       entry_offset = ntohl(hdr_v5->hdr_fblockoffset);\n>> +       foffsetblock = dir_offset;\n>> +\n>> +       if (opts && opts->pathspec && opts->pathspec->nr) {\n>> +               paths = xmalloc((opts->pathspec->nr + 1)*sizeof(char *));\n>> +               paths[opts->pathspec->nr] = NULL;\n>\n> Put this statement here\n>\n> GUARD_PATHSPEC(opts->pathspec,\n>       PATHSPEC_FROMTOP |\n>       PATHSPEC_MAXDEPTH |\n>       PATHSPEC_LITERAL |\n>       PATHSPEC_GLOB |\n>       PATHSPEC_ICASE);\n>\n> This says the mentioned magic is safe in this code. New magic may or\n> may not be and needs to be checked (soonest by me, I'm going to add\n> negative pathspec and I'll need to look into how it should be handled\n> in this code block).\n\nThanks, I'll add the statement in the re-roll.\n\n>> +               for (i = 0; i < opts->pathspec->nr; i++) {\n>> +                       char *super = strdup(opts->pathspec->items[i].match);\n>> +                       int len = strlen(super);\n>\n> You should only check as far as items[i].nowildcard_len, not strlen().\n> The rest could be wildcards and stuff and not so reliable.\n\nOk, will change it in the re-roll.\n\n>> +                       while (len && super[len - 1] == '/' && super[len - 2] == '/')\n>> +                               super[--len] = '\\0'; /* strip all but one trailing slash */\n>> +                       while (len && super[--len] != '/')\n>> +                               ; /* scan backwards to next / */\n>> +                       if (len >= 0)\n>> +                               super[len--] = '\\0';\n>> +                       if (len <= 0) {\n>> +                               need_root = 1;\n>> +                               break;\n>> +                       }\n>> +                       paths[i] = super;\n>> +               }\n>\n> And maybe put the comment \"FIXME: consider merging this code with\n> create_simplify() in dir.c\" somewhere. It's for me to look for things\n> to do when I'm bored ;-)\n\nHeh, thanks, will do.\n\n>> +       }\n>> +\n>> +       if (!need_root)\n>> +               parse_pathspec(&adjusted_pathspec, PATHSPEC_ALL_MAGIC, PATHSPEC_PREFER_CWD, NULL, paths);\n>\n> I would go with PATHSPEC_PREFER_FULL instead of _CWD as it's safer.\n> Looking only at this function without caller context, it's hard to say\n> if _CWD is the right choice.\n\nOk, thanks, will change.\n\n>> +\n>> +       de = root_directory;\n>> +       last_de = de;\n>\n> This statement is redundant. last_de is only used in one code block\n> below and it's always re-initialized before entering the loop to skip\n> subdirs.\n\nRight, good catch!  Will remove it.\n\n>> +       while (de) {\n>> +               if (need_root ||\n>> +                   match_pathspec_depth(&adjusted_pathspec, de->pathname, de->de_pathlen, 0, NULL)) {\n>> +                       if (read_entries(istate, de, entry_offset,\n>> +                                        mmap, mmap_size, &nr,\n>> +                                        foffsetblock) < 0)\n>> +                               return -1;\n>> +               } else {\n>> +                       last_de = de;\n>> +                       for (i = 0; i < de->de_nsubtrees; i++) {\n>> +                               de->sub[i]->next = last_de->next;\n>> +                               last_de->next = de->sub[i];\n>> +                               last_de = last_de->next;\n>> +                       }\n>> +               }\n>> +               de = de->next;\n>> +       }\n>> +       free_directory_tree(root_directory);\n>> +       istate->cache_nr = nr;\n>> +       return 0;\n>> +}\n> -- \n> Duy\n"},{"id":"231298","messageId":"87txeuxiwu.fsf@gmail.com","threadId":"35410","inReplyTo":"CACsJy8Bti465Lv9U+m=VnqvGQAFk=WxCsmcGmdB0=NLesF5Mtw@mail.gmail.com","subject":"Re: [PATCH v4 23/24] POC for partial writing","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-30T10:50:41Z","receivedAt":"2013-11-30T10:50:41Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> On Wed, Nov 27, 2013 at 7:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n>> This makes update-index use both partial reading and partial writing.\n>> Partial reading is only used no option other than the paths is passed to\n>> the command.\n>>\n>> This passes the test suite,\n>\n> Just checking, the test suite was run with TEST_GIT_INDEX_VERSION=5, right?\n\nYes, sure, I've been using that for finding problems in my approach.\n\n>> but doesn't behave correctly when a write\n>> fails.  A log should be written to the lock file, in order to be able to\n>> recover if a write fails.\n>\n> From the API point of view this looks nice (you should have hidden\n> needs_write = 1 in cache_invalidate_path and change_cache_version\n> though).\n\nOk, thanks, I'll look into that, once I have time to make more than a\nPOC for this.\n\n> We could support partial file removal too by marking removed\n> files \"removed\", but that impacts the reading code and may have bad\n> interaction with cache_invalidate_path/needs_rewrite. Probably not\n> worth the effort until someone shows us they remove stuff often.\n\nYes, right, I think it could be done, but I didn't take the time to look\ninto that for now.  The reading code should actually be fine as it is,\nas it already takes care of entries with the removed flag, but I'm not\nsure about cache_tree_invalidate_path/needs_rewrite.\n\n>> ---\n>>  builtin/update-index.c |  43 +++++++++++---\n>>  cache-tree.c           |  13 +++++\n>>  cache-tree.h           |   1 +\n>>  cache.h                |  27 ++++++++-\n>>  lockfile.c             |   2 +-\n>>  read-cache-v2.c        |   2 +\n>>  read-cache-v5.c        | 154 ++++++++++++++++++++++++++++++++++++++++---------\n>>  read-cache.c           |  30 ++++++++++\n>>  read-cache.h           |   1 +\n>>  resolve-undo.c         |   1 +\n>>  10 files changed, 237 insertions(+), 37 deletions(-)\n>>\n>> diff --git a/builtin/update-index.c b/builtin/update-index.c\n>> index 8b3f7a0..69f0949 100644\n>> --- a/builtin/update-index.c\n>> +++ b/builtin/update-index.c\n>> @@ -56,6 +56,7 @@ static int mark_ce_flags(const char *path, int flag, int mark)\n>>                 else\n>>                         active_cache[pos]->ce_flags &= ~flag;\n>>                 cache_tree_invalidate_path(active_cache_tree, path);\n>> +               the_index.needs_rewrite = 1;\n>>                 active_cache_changed = 1;\n>>                 return 0;\n>>         }\n>> @@ -99,6 +100,8 @@ static int add_one_path(const struct cache_entry *old, const char *path, int len\n>>         memcpy(ce->name, path, len);\n>>         ce->ce_flags = create_ce_flags(0);\n>>         ce->ce_namelen = len;\n>> +       if (old)\n>> +               ce->entry_pos = old->entry_pos;\n>>         fill_stat_cache_info(ce, st);\n>>         ce->ce_mode = ce_mode_from_stat(old, st->st_mode);\n>>\n>> @@ -268,6 +271,7 @@ static void chmod_path(int flip, const char *path)\n>>                 goto fail;\n>>         }\n>>         cache_tree_invalidate_path(active_cache_tree, path);\n>> +       the_index.needs_rewrite = 1;\n>>         active_cache_changed = 1;\n>>         report(\"chmod %cx '%s'\", flip, path);\n>>         return;\n>> @@ -706,15 +710,18 @@ static int reupdate_callback(struct parse_opt_ctx_t *ctx,\n>>\n>>  int cmd_update_index(int argc, const char **argv, const char *prefix)\n>>  {\n>> -       int newfd, entries, has_errors = 0, line_termination = '\\n';\n>> +       int newfd, has_errors = 0, line_termination = '\\n';\n>>         int read_from_stdin = 0;\n>>         int prefix_length = prefix ? strlen(prefix) : 0;\n>>         int preferred_index_format = 0;\n>>         char set_executable_bit = 0;\n>>         struct refresh_params refresh_args = {0, &has_errors};\n>>         int lock_error = 0;\n>> +       struct filter_opts opts;\n>> +       struct pathspec pathspec;\n>>         struct lock_file *lock_file;\n>>         struct parse_opt_ctx_t ctx;\n>> +       int i, needs_full_read = 0;\n>>         int parseopt_state = PARSE_OPT_UNKNOWN;\n>>         struct option options[] = {\n>>                 OPT_BIT('q', NULL, &refresh_args.flags,\n>> @@ -810,9 +817,23 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>>         if (newfd < 0)\n>>                 lock_error = errno;\n>>\n>> -       entries = read_cache();\n>> -       if (entries < 0)\n>> -               die(\"cache corrupted\");\n>> +       for (i = 0; i < argc; i++) {\n>> +               if (!prefixcmp(argv[i], \"--\"))\n>> +                       needs_full_read = 1;\n>> +       }\n>> +       if (!needs_full_read) {\n>> +               memset(&opts, 0, sizeof(struct filter_opts));\n>> +               parse_pathspec(&pathspec, 0,\n>> +                              PATHSPEC_PREFER_CWD |\n>> +                              PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n>> +                              prefix, argv + 1);\n>> +               opts.pathspec = &pathspec;\n>> +               if (read_cache_filtered(&opts) < 0)\n>> +                       die(\"cache corrupted\");\n>> +       } else {\n>> +               if (read_cache() < 0)\n>> +                       die(\"cache corrupted\");\n>> +       }\n>>\n>>         /*\n>>          * Custom copy of parse_options() because we want to handle\n>> @@ -862,6 +883,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>>                             preferred_index_format,\n>>                             INDEX_FORMAT_LB, INDEX_FORMAT_UB);\n>>\n>> +               the_index.needs_rewrite = 1;\n>>                 active_cache_changed = 1;\n>>                 change_cache_version(preferred_index_format);\n>>         }\n>> @@ -890,17 +912,22 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>>         }\n>>\n>>         if (active_cache_changed) {\n>> +               int r;\n>>                 if (newfd < 0) {\n>>                         if (refresh_args.flags & REFRESH_QUIET)\n>>                                 exit(128);\n>>                         unable_to_lock_index_die(get_index_file(), lock_error);\n>>                 }\n>> -               if (write_cache(newfd, active_cache, active_nr) ||\n>> -                   commit_locked_index(lock_file))\n>> +               r = write_cache_partial(newfd);\n>> +               if (r < 0)\n>>                         die(\"Unable to write new index file\");\n>> +               else if (r == 0)\n>> +                       commit_lock_file(lock_file);\n>> +               else\n>> +                       remove_lock_file();\n>> +       } else {\n>> +               rollback_lock_file(lock_file);\n>>         }\n>>\n>> -       rollback_lock_file(lock_file);\n>> -\n>>         return has_errors ? 1 : 0;\n>>  }\n>> diff --git a/cache-tree.c b/cache-tree.c\n>> index 1209732..a3d18bb 100644\n>> --- a/cache-tree.c\n>> +++ b/cache-tree.c\n>> @@ -123,6 +123,15 @@ void cache_tree_invalidate_path(struct cache_tree *it, const char *path)\n>>                 return;\n>>         slash = strchr(path, '/');\n>>         it->entry_count = -1;\n>> +       /*\n>> +        * Mark the cache_tree directory entry as invalid too. The\n>> +        * entry_count defines if the tree is valid, so we don't need\n>> +        * to reset any other field.\n>> +        */\n>> +       if (it->de_ref) {\n>> +               it->de_ref->de_nentries = -1;\n>> +               it->de_ref->changed = 1;\n>> +       }\n>>         if (!slash) {\n>>                 int pos;\n>>                 namelen = strlen(path);\n>> @@ -140,6 +149,10 @@ void cache_tree_invalidate_path(struct cache_tree *it, const char *path)\n>>                                 sizeof(struct cache_tree_sub *) *\n>>                                 (it->subtree_nr - pos - 1));\n>>                         it->subtree_nr--;\n>> +                       if (it->de_ref) {\n>> +                               it->de_ref->de_nsubtrees--;\n>> +                               it->de_ref->changed = 1;\n>> +                       }\n>>                 }\n>>                 return;\n>>         }\n>> diff --git a/cache-tree.h b/cache-tree.h\n>> index 9818926..eaf14a9 100644\n>> --- a/cache-tree.h\n>> +++ b/cache-tree.h\n>> @@ -18,6 +18,7 @@ struct cache_tree {\n>>         unsigned char sha1[20];\n>>         int subtree_nr;\n>>         int subtree_alloc;\n>> +       struct directory_entry *de_ref;\n>>         struct cache_tree_sub **down;\n>>  };\n>>\n>> diff --git a/cache.h b/cache.h\n>> index 71b98cf..1a634dc 100644\n>> --- a/cache.h\n>> +++ b/cache.h\n>> @@ -137,11 +137,31 @@ struct cache_entry {\n>>         unsigned int ce_namelen;\n>>         unsigned char sha1[20];\n>>         uint32_t ce_stat_crc;\n>> +       unsigned int entry_pos;\n>> +       unsigned int changed;\n>>         struct cache_entry *next; /* used by name_hash */\n>>         struct cache_entry *next_ce;\n>>         char name[FLEX_ARRAY]; /* more */\n>>  };\n>>\n>> +struct directory_entry {\n>> +       struct directory_entry **sub;\n>> +       struct directory_entry *next;\n>> +       struct directory_entry *next_hash;\n>> +       struct cache_entry *ce;\n>> +       struct cache_entry *ce_last;\n>> +       uint32_t de_foffset;\n>> +       uint32_t de_nsubtrees;\n>> +       uint32_t de_nfiles;\n>> +       uint32_t de_nentries;\n>> +       unsigned char sha1[20];\n>> +       uint16_t de_flags;\n>> +       uint32_t de_pathlen;\n>> +       uint32_t entry_pos;\n>> +       unsigned int changed;\n>> +       char pathname[FLEX_ARRAY];\n>> +};\n>> +\n>>  #define CE_NAMEMASK  (0x0fff)\n>>  #define CE_STAGEMASK (0x3000)\n>>  #define CE_EXTENDED  (0x4000)\n>> @@ -317,13 +337,15 @@ struct filter_opts {\n>>\n>>  struct index_state {\n>>         struct cache_entry **cache;\n>> +       struct directory_entry *root_directory;\n>>         unsigned int version;\n>>         unsigned int cache_nr, cache_alloc, cache_changed;\n>>         struct string_list *resolve_undo;\n>>         struct cache_tree *cache_tree;\n>>         struct cache_time timestamp;\n>>         unsigned name_hash_initialized : 1,\n>> -                initialized : 1;\n>> +                initialized : 1,\n>> +                needs_rewrite : 1;\n>>         struct hash_table name_hash;\n>>         struct hash_table dir_hash;\n>>         struct index_ops *ops;\n>> @@ -353,6 +375,7 @@ extern void free_name_hash(struct index_state *istate);\n>>  #define is_cache_unborn() is_index_unborn(&the_index)\n>>  #define read_cache_unmerged() read_index_unmerged(&the_index)\n>>  #define write_cache(newfd, cache, entries) write_index(&the_index, (newfd))\n>> +#define write_cache_partial(newfd) write_index_partial(&the_index, (newfd))\n>>  #define discard_cache() discard_index(&the_index)\n>>  #define unmerged_cache() unmerged_index(&the_index)\n>>  #define cache_name_pos(name, namelen) index_name_pos(&the_index,(name),(namelen))\n>> @@ -529,6 +552,7 @@ extern int read_index_from(struct index_state *, const char *path);\n>>  extern int is_index_unborn(struct index_state *);\n>>  extern int read_index_unmerged(struct index_state *);\n>>  extern int write_index(struct index_state *, int newfd);\n>> +extern int write_index_partial(struct index_state *, int newfd);\n>>  extern int discard_index(struct index_state *);\n>>  extern int unmerged_index(const struct index_state *);\n>>  extern int verify_path(const char *path);\n>> @@ -613,6 +637,7 @@ extern NORETURN void unable_to_lock_index_die(const char *path, int err);\n>>  extern int hold_lock_file_for_update(struct lock_file *, const char *path, int);\n>>  extern int hold_lock_file_for_append(struct lock_file *, const char *path, int);\n>>  extern int commit_lock_file(struct lock_file *);\n>> +extern void remove_lock_file(void);\n>>  extern void update_index_if_able(struct index_state *, struct lock_file *);\n>>\n>>  extern int hold_locked_index(struct lock_file *, int);\n>> diff --git a/lockfile.c b/lockfile.c\n>> index 8fbcb6a..c150e5c 100644\n>> --- a/lockfile.c\n>> +++ b/lockfile.c\n>> @@ -7,7 +7,7 @@\n>>  static struct lock_file *lock_file_list;\n>>  static const char *alternate_index_output;\n>>\n>> -static void remove_lock_file(void)\n>> +void remove_lock_file(void)\n>>  {\n>>         pid_t me = getpid();\n>>\n>> diff --git a/read-cache-v2.c b/read-cache-v2.c\n>> index f884c10..1fec892 100644\n>> --- a/read-cache-v2.c\n>> +++ b/read-cache-v2.c\n>> @@ -555,5 +555,7 @@ struct index_ops v2_ops = {\n>>         match_stat_basic,\n>>         verify_hdr,\n>>         read_index_v2,\n>> +       write_index_v2,\n>> +       /* Partial writing is the same as writing the full index for v2 */\n>>         write_index_v2\n>>  };\n>> diff --git a/read-cache-v5.c b/read-cache-v5.c\n>> index a5e9b5a..13436a3 100644\n>> --- a/read-cache-v5.c\n>> +++ b/read-cache-v5.c\n>> @@ -20,22 +20,6 @@ struct extension_header {\n>>         uint32_t crc;\n>>  };\n>>\n>> -struct directory_entry {\n>> -       struct directory_entry **sub;\n>> -       struct directory_entry *next;\n>> -       struct directory_entry *next_hash;\n>> -       struct cache_entry *ce;\n>> -       struct cache_entry *ce_last;\n>> -       uint32_t de_foffset;\n>> -       uint32_t de_nsubtrees;\n>> -       uint32_t de_nfiles;\n>> -       uint32_t de_nentries;\n>> -       unsigned char sha1[20];\n>> -       uint16_t de_flags;\n>> -       uint32_t de_pathlen;\n>> -       char pathname[FLEX_ARRAY];\n>> -};\n>> -\n>>  struct conflict_part {\n>>         struct conflict_part *next;\n>>         uint16_t flags;\n>> @@ -246,7 +230,7 @@ static struct directory_entry *read_directories(unsigned int *dir_offset,\n>>                 offsetof(struct ondisk_directory_entry, name) - 5;\n>>         disk_de = ptr_add(mmap, *dir_offset);\n>>         de = directory_entry_from_ondisk(disk_de, len);\n>> -\n>> +       de->entry_pos = *dir_offset;\n>>         data_len = len + 1 + offsetof(struct ondisk_directory_entry, name);\n>>         filecrc = ptr_add(mmap, *dir_offset + data_len);\n>>         if (!check_crc32(0, ptr_add(mmap, *dir_offset), data_len, ntoh_l(*filecrc)))\n>> @@ -281,6 +265,7 @@ static int read_entry(struct cache_entry **ce, char *pathname, size_t pathlen,\n>>         entry_offset = first_entry_offset + ntoh_l(*beginning);\n>>         disk_ce = ptr_add(mmap, entry_offset);\n>>         *ce = cache_entry_from_ondisk(disk_ce, pathname, len, pathlen);\n>> +       (*ce)->entry_pos = entry_offset;\n>>         filecrc = ptr_add(mmap, entry_offset + len + 1 + sizeof(*disk_ce));\n>>         if (!check_crc32(0,\n>>                 ptr_add(mmap, entry_offset), len + 1 + sizeof(*disk_ce),\n>> @@ -439,6 +424,7 @@ static struct cache_tree *convert_one(struct directory_entry *de)\n>>\n>>         it = cache_tree();\n>>         it->entry_count = de->de_nentries;\n>> +       it->de_ref = de;\n>>         if (0 <= it->entry_count)\n>>                 hashcpy(it->sha1, de->sha1);\n>>\n>> @@ -523,14 +509,6 @@ static int read_entries(struct index_state *istate, struct directory_entry *de,\n>>         return 0;\n>>  }\n>>\n>> -static void free_directory_tree(struct directory_entry *de) {\n>> -       int i;\n>> -\n>> -       for (i = 0; i < de->de_pathlen; i++)\n>> -               free_directory_tree(de->sub[i]);\n>> -       free(de);\n>> -}\n>> -\n>>  /*\n>>   * Read an index-v5 file filtered by the filter_opts.   If opts is NULL,\n>>   * everything will be read.\n>> @@ -626,7 +604,7 @@ static int read_index_v5(struct index_state *istate, void *mmap,\n>>                 }\n>>         }\n>>         istate->cache_tree = cache_tree_convert_v5(root_directory);\n>> -       free_directory_tree(root_directory);\n>> +       istate->root_directory = root_directory;\n>>         istate->cache_nr = nr;\n>>         return 0;\n>>  }\n>> @@ -696,6 +674,7 @@ static void ce_smudge_racily_clean_entry(struct cache_entry *ce)\n>>          * that hasn't changed checking the sha1.\n>>          */\n>>         ce->ce_flags |= CE_SMUDGED;\n>> +       ce->changed = 1;\n>>  }\n>>\n>>  static char *super_directory(char *filename)\n>> @@ -1231,6 +1210,103 @@ static int write_resolve_undo(struct index_state *istate,\n>>         return 0;\n>>  }\n>>\n>> +static int write_ce_if_necessary(struct cache_entry *ce, void *cb_data)\n>> +{\n>> +       int *fdx = cb_data, pathlen, size;\n>> +       int fd = *fdx;\n>> +       char *dir;\n>> +       struct ondisk_cache_entry *ondisk;\n>> +       uint32_t crc;\n>> +\n>> +       assert(ce->entry_pos != 0);\n>> +       /* TODO I'm just using the_index out of lazyness here */\n>> +       if (!ce_uptodate(ce) && is_racy_timestamp(&the_index, ce))\n>> +               ce_smudge_racily_clean_entry(ce);\n>> +       if (!ce->changed)\n>> +               return 0;\n>> +       if (is_null_sha1(ce->sha1)) {\n>> +               static const char msg[] = \"cache entry has null sha1: %s\";\n>> +               static int allow = -1;\n>> +\n>> +               if (allow < 0)\n>> +                       allow = git_env_bool(\"GIT_ALLOW_NULL_SHA1\", 0);\n>> +               if (allow)\n>> +                       warning(msg, ce->name);\n>> +               else\n>> +                       return error(msg, ce->name);\n>> +       }\n>> +       dir = super_directory(ce->name);\n>> +       pathlen = dir ? strlen(dir) + 1 : 0;\n>> +       size = offsetof(struct ondisk_cache_entry, name) +\n>> +               ce_namelen(ce) - pathlen + 1;\n>> +       ondisk = xmalloc(size);\n>> +\n>> +       crc = 0;\n>> +       ondisk_from_cache_entry(ce, ondisk, pathlen);\n>> +       if (lseek(fd, ce->entry_pos, SEEK_SET) < ce->entry_pos)\n>> +               die(\"eror ce seeking\");\n>> +       if (ce_write(&crc, fd, ondisk, size) < 0)\n>> +               return -1;\n>> +       crc = htonl(crc);\n>> +       if (ce_write(NULL, fd, &crc, 4) < 0)\n>> +               return -1;\n>> +       return ce_flush(fd);\n>> +}\n>> +\n>> +static void ondisk_from_directory_entry_partial(struct directory_entry *de,\n>> +                                               struct ondisk_directory_entry *ondisk)\n>> +{\n>> +       ondisk->foffset   = htonl(de->de_foffset);\n>> +       ondisk->nsubtrees = htonl(de->de_nsubtrees);\n>> +       ondisk->nfiles    = htonl(de->de_nfiles);\n>> +       ondisk->nentries  = htonl(de->de_nentries);\n>> +       hashcpy(ondisk->sha1, de->sha1);\n>> +       ondisk->flags     = htons(de->de_flags);\n>> +       if (de->de_pathlen == 0) {\n>> +               memcpy(ondisk->name, \"\\0\", 1);\n>> +       } else {\n>> +               memcpy(ondisk->name, de->pathname, de->de_pathlen);\n>> +               memcpy(ondisk->name + de->de_pathlen - 1, \"/\\0\", 2);\n>> +       }\n>> +}\n>> +\n>> +static int write_directories_partial(struct directory_entry *de, int fd)\n>> +{\n>> +       int ondisk_size = offsetof(struct ondisk_directory_entry, name);\n>> +       int size = ondisk_size + de->de_pathlen + 1;\n>> +       int i;\n>> +       uint32_t crc;\n>> +       struct ondisk_directory_entry *ondisk;\n>> +\n>> +       if (de->changed) {\n>> +               crc = 0;\n>> +               ondisk = xmalloc(size);\n>> +               ondisk_from_directory_entry_partial(de, ondisk);\n>> +               if (lseek(fd, de->entry_pos, SEEK_SET) < de->entry_pos)\n>> +                       die(\"error directory seeking\");;\n>> +               if (ce_write(&crc, fd, ondisk, size) < 0)\n>> +                       return -1;\n>> +               crc = htonl(crc);\n>> +               if (ce_write(NULL, fd, &crc, 4) < 0)\n>> +                       return -1;\n>> +               free(ondisk);\n>> +               if (ce_flush(fd) < 0)\n>> +                       return -1;\n>> +       }\n>> +       for (i = 0; i < de->de_nsubtrees; i++) {\n>> +               if (write_directories_partial(de->sub[i], fd) < 0)\n>> +                       return -1;\n>> +       }\n>> +       return 0;\n>> +}\n>> +\n>> +static int write_partial(struct index_state *istate, int fd)\n>> +{\n>> +       write_directories_partial(istate->root_directory, fd);\n>> +\n>> +       return for_each_index_entry(istate, write_ce_if_necessary, &fd);\n>> +}\n>> +\n>>  static int write_index_v5(struct index_state *istate, int newfd)\n>>  {\n>>         struct cache_header hdr;\n>> @@ -1296,9 +1372,33 @@ static int write_index_v5(struct index_state *istate, int newfd)\n>>         return ce_flush(newfd);\n>>  }\n>>\n>> +static int write_index_partial_v5(struct index_state *istate, int newfd)\n>> +{\n>> +       int fd;\n>> +       char *path = get_index_file();\n>> +\n>> +       if (istate->needs_rewrite || istate->cache_nr == 0)\n>> +               return write_index_v5(istate, newfd);\n>> +       if (istate->filter_opts && istate->needs_rewrite)\n>> +               die(\"BUG: cannot write a partially read index\");\n>> +       fd = open(path, O_RDWR, 0666);\n>> +       if (fd < 0) {\n>> +               if (errno == ENOENT)\n>> +                       die(\"no index file exists cannot do a partial write\");\n>> +               die_errno(\"index file opening for writing failed\");\n>> +       }\n>> +\n>> +       if (write_partial(istate, fd) < 0)\n>> +               return -1;\n>> +       if (ce_flush(fd) < 0)\n>> +               return -1;\n>> +       return 1;\n>> +}\n>> +\n>>  struct index_ops v5_ops = {\n>>         match_stat_basic,\n>>         verify_hdr,\n>>         read_index_v5,\n>> -       write_index_v5\n>> +       write_index_v5,\n>> +       write_index_partial_v5\n>>  };\n>> diff --git a/read-cache.c b/read-cache.c\n>> index 04430e5..1cad0e2 100644\n>> --- a/read-cache.c\n>> +++ b/read-cache.c\n>> @@ -32,6 +32,9 @@ static void replace_index_entry(struct index_state *istate, int nr, struct cache\n>>\n>>         remove_name_hash(istate, old);\n>>         set_index_entry(istate, nr, ce);\n>> +       ce->changed = 1;\n>> +       if (ce->entry_pos == 0)\n>> +               istate->needs_rewrite = 1;\n>>         istate->cache_changed = 1;\n>>  }\n>>\n>> @@ -494,6 +497,7 @@ int remove_index_entry_at(struct index_state *istate, int pos)\n>>\n>>         record_resolve_undo(istate, ce);\n>>         remove_name_hash(istate, ce);\n>> +       istate->needs_rewrite = 1;\n>>         istate->cache_changed = 1;\n>>         istate->cache_nr--;\n>>         if (pos >= istate->cache_nr)\n>> @@ -520,6 +524,7 @@ void remove_marked_cache_entries(struct index_state *istate)\n>>                 else\n>>                         ce_array[j++] = ce_array[i];\n>>         }\n>> +       istate->needs_rewrite = 1;\n>>         istate->cache_changed = 1;\n>>         istate->cache_nr = j;\n>>  }\n>> @@ -1024,6 +1029,7 @@ int add_index_entry(struct index_state *istate, struct cache_entry *ce, int opti\n>>                         istate->cache + pos,\n>>                         (istate->cache_nr - pos - 1) * sizeof(ce));\n>>         set_index_entry(istate, pos, ce);\n>> +       istate->needs_rewrite = 1;\n>>         istate->cache_changed = 1;\n>>         return 0;\n>>  }\n>> @@ -1108,6 +1114,8 @@ static struct cache_entry *refresh_cache_ent(struct index_state *istate,\n>>         size = ce_size(ce);\n>>         updated = xmalloc(size);\n>>         memcpy(updated, ce, size);\n>> +       updated->changed = 1;\n>> +       updated->entry_pos = ce->entry_pos;\n>>         fill_stat_cache_info(updated, &st);\n>>         /*\n>>          * If ignore_valid is not set, we should leave CE_VALID bit\n>> @@ -1201,6 +1209,8 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n>>                                  * means the index is not valid anymore.\n>>                                  */\n>>                                 ce->ce_flags &= ~CE_VALID;\n>> +                               /* TODO: remove this maybe? */\n>> +                               istate->needs_rewrite = 1;\n>>                                 istate->cache_changed = 1;\n>>                         }\n>>                         if (quiet)\n>> @@ -1241,6 +1251,8 @@ void initialize_index(struct index_state *istate, int version)\n>>                         version = atoi(envversion);\n>>         }\n>>         istate->version = version;\n>> +       istate->needs_rewrite = 0;\n>> +       istate->root_directory = NULL;\n>>         set_istate_ops(istate);\n>>  }\n>>\n>> @@ -1427,6 +1439,16 @@ int is_index_unborn(struct index_state *istate)\n>>         return (!istate->cache_nr && !istate->timestamp.sec);\n>>  }\n>>\n>> +static void free_directory_tree(struct directory_entry *de) {\n>> +       int i;\n>> +\n>> +       if (!de)\n>> +               return;\n>> +       for (i = 0; i < de->de_pathlen; i++)\n>> +               free_directory_tree(de->sub[i]);\n>> +       free(de);\n>> +}\n>> +\n>>  int discard_index(struct index_state *istate)\n>>  {\n>>         int i;\n>> @@ -1435,6 +1457,7 @@ int discard_index(struct index_state *istate)\n>>                 free(istate->cache[i]);\n>>         resolve_undo_clear_index(istate);\n>>         istate->cache_nr = 0;\n>> +       istate->needs_rewrite = 0;\n>>         istate->cache_changed = 0;\n>>         istate->timestamp.sec = 0;\n>>         istate->timestamp.nsec = 0;\n>> @@ -1446,6 +1469,8 @@ int discard_index(struct index_state *istate)\n>>         istate->cache_alloc = 0;\n>>         istate->ops = NULL;\n>>         istate->filter_opts = NULL;\n>> +       free_directory_tree(istate->root_directory);\n>> +       istate->root_directory = NULL;\n>>         return 0;\n>>  }\n>>\n>> @@ -1491,6 +1516,11 @@ int write_index(struct index_state *istate, int newfd)\n>>         return istate->ops->write_index(istate, newfd);\n>>  }\n>>\n>> +int write_index_partial(struct index_state *istate, int newfd)\n>> +{\n>> +       return istate->ops->write_index_partial(istate, newfd);\n>> +}\n>> +\n>>  /*\n>>   * Read the index file that is potentially unmerged into given\n>>   * index_state, dropping any unmerged entries.  Returns true if\n>> diff --git a/read-cache.h b/read-cache.h\n>> index 9d66df6..e7f36ae 100644\n>> --- a/read-cache.h\n>> +++ b/read-cache.h\n>> @@ -31,6 +31,7 @@ struct index_ops {\n>>         int (*read_index)(struct index_state *istate, void *mmap, unsigned long mmap_size,\n>>                           struct filter_opts *opts);\n>>         int (*write_index)(struct index_state *istate, int newfd);\n>> +       int (*write_index_partial)(struct index_state *istate, int newfd);\n>>  };\n>>\n>>  extern struct index_ops v2_ops;\n>> diff --git a/resolve-undo.c b/resolve-undo.c\n>> index c09b006..c496c20 100644\n>> --- a/resolve-undo.c\n>> +++ b/resolve-undo.c\n>> @@ -110,6 +110,7 @@ void resolve_undo_clear_index(struct index_state *istate)\n>>         string_list_clear(resolve_undo, 1);\n>>         free(resolve_undo);\n>>         istate->resolve_undo = NULL;\n>> +       istate->needs_rewrite = 1;\n>>         istate->cache_changed = 1;\n>>  }\n>>\n>> --\n>> 1.8.4.2\n>>\n>\n>\n>\n> -- \n> Duy\n"},{"id":"231300","messageId":"CALWbr2ybgpdDJLRKA4zwPmRE4LQv4VWLJM6Jv0dO-GyHREGW7g@mail.gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-13-git-send-email-t.gummerer@gmail.com","subject":"Re: [PATCH v4 12/24] read-cache: read index-v5","fromName":"Antoine Pelisse","fromEmail":"apelisse@gmail.com","sentAt":"2013-11-30T12:19:37Z","receivedAt":"2013-11-30T12:19:37Z","isPatch":true,"sender":{"key":"apelisse@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1929644?v=4"},"body":"On Wed, Nov 27, 2013 at 1:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n> Make git read the index file version 5 without complaining.\n>\n> This version of the reader reads neither the cache-tree\n> nor the resolve undo data, however, it won't choke on an\n> index that includes such data.\n>\n> Helped-by: Junio C Hamano <gitster@pobox.com>\n> Helped-by: Nguyen Thai Ngoc Duy <pclouds@gmail.com>\n> Helped-by: Thomas Rast <trast@student.ethz.ch>\n> Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n> ---\n> [...]\n> +static struct directory_entry *read_directories(unsigned int *dir_offset,\n> +                                               unsigned int *dir_table_offset,\n> +                                               void *mmap, int mmap_size)\n\nMinor nit: why is this mmap_size \"int\" while all others are \"unsigned long\" ?\n\n> [...]\n> +static int read_entry(struct cache_entry **ce, char *pathname, size_t pathlen,\n> +                     void *mmap, unsigned long mmap_size,\n> +                     unsigned int first_entry_offset,\n> +                     unsigned int foffsetblock)\n> [...]\n> +static int read_entries(struct index_state *istate, struct directory_entry *de,\n> +                       unsigned int first_entry_offset, void *mmap,\n> +                       unsigned long mmap_size, unsigned int *nr,\n> +                       unsigned int foffsetblock)\n> [...]\n> +static int read_index_v5(struct index_state *istate, void *mmap,\n> +                        unsigned long mmap_size, struct filter_opts *opts)\n"},{"id":"231308","messageId":"CALWbr2xUMHSU0MV-6nVbN4_eSMoj3Eyc_Ta_CxTwZ_Y8tLfbdQ@mail.gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-13-git-send-email-t.gummerer@gmail.com","subject":"Re: [PATCH v4 12/24] read-cache: read index-v5","fromName":"Antoine Pelisse","fromEmail":"apelisse@gmail.com","sentAt":"2013-11-30T15:26:46Z","receivedAt":"2013-11-30T15:26:46Z","isPatch":true,"sender":{"key":"apelisse@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1929644?v=4"},"body":"On Wed, Nov 27, 2013 at 1:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n> +static int verify_hdr(void *mmap, unsigned long size)\n> +{\n> +       uint32_t *filecrc;\n> +       unsigned int header_size;\n> +       struct cache_header *hdr;\n> +       struct cache_header_v5 *hdr_v5;\n> +\n> +       if (size < sizeof(struct cache_header)\n> +           + sizeof (struct cache_header_v5) + 4)\n> +               die(\"index file smaller than expected\");\n> +\n> +       hdr = mmap;\n> +       hdr_v5 = ptr_add(mmap, sizeof(*hdr));\n> +       /* Size of the header + the size of the extensionoffsets */\n> +       header_size = sizeof(*hdr) + sizeof(*hdr_v5) + hdr_v5->hdr_nextension * 4;\n> +       /* Initialize crc */\n> +       filecrc = ptr_add(mmap, header_size);\n> +       if (!check_crc32(0, hdr, header_size, ntohl(*filecrc)))\n> +               return error(\"bad index file header crc signature\");\n> +       return 0;\n> +}\n\nI find it curious that we actually need a value from the header (and\nuse it for pointer arithmetic) to check that the header is valid. The\napplication will crash before the crc is checked if\nhdr_v5->hdr_nextensions is corrupted. Or am I missing something ?\n"},{"id":"231310","messageId":"CALWbr2yaD9Z98ysEzVHiQQR_W_zEj7bp0uEgZ3Z=Tp=Yc1NnoQ@mail.gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-10-git-send-email-t.gummerer@gmail.com","subject":"Re: [PATCH v4 09/24] ls-files.c: use index api","fromName":"Antoine Pelisse","fromEmail":"apelisse@gmail.com","sentAt":"2013-11-30T15:39:52Z","receivedAt":"2013-11-30T15:39:52Z","isPatch":true,"sender":{"key":"apelisse@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1929644?v=4"},"body":"On Wed, Nov 27, 2013 at 1:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n> @@ -447,6 +463,7 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>         struct dir_struct dir;\n>         struct exclude_list *el;\n>         struct string_list exclude_list = STRING_LIST_INIT_NODUP;\n> +       struct filter_opts *opts = xmalloc(sizeof(*opts));\n>         struct option builtin_ls_files_options[] = {\n>                 { OPTION_CALLBACK, 'z', NULL, NULL, NULL,\n>                         N_(\"paths are separated with NUL character\"),\n> @@ -512,9 +529,6 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>                 prefix_len = strlen(prefix);\n>         git_config(git_default_config, NULL);\n>\n> -       if (read_cache() < 0)\n> -               die(\"index file corrupt\");\n> -\n>         argc = parse_options(argc, argv, prefix, builtin_ls_files_options,\n>                         ls_files_usage, 0);\n>         el = add_exclude_list(&dir, EXC_CMDL, \"--exclude option\");\n> @@ -550,6 +564,24 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>                        PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n>                        prefix, argv);\n>\n> +       if (!with_tree && !needs_trailing_slash_stripped()) {\n> +               memset(opts, 0, sizeof(*opts));\n> +               opts->pathspec = &pathspec;\n> +               opts->read_staged = 1;\n> +               if (show_resolve_undo)\n> +                       opts->read_resolve_undo = 1;\n> +               if (read_cache_filtered(opts) < 0)\n> +                       die(\"index file corrupt\");\n> +       } else {\n> +               if (read_cache() < 0)\n> +                       die(\"index file corrupt\");\n> +               parse_pathspec(&pathspec, 0,\n> +                              PATHSPEC_PREFER_CWD |\n> +                              PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n> +                              prefix, argv);\n> +\n> +       }\n> +\n\nWould it make sense to move the declaration of \"opts\" as a non-pointer\nto the block where it's used ?\n"},{"id":"231312","messageId":"87r49xy7ns.fsf@gmail.com","threadId":"35410","inReplyTo":"CALWbr2yaD9Z98ysEzVHiQQR_W_zEj7bp0uEgZ3Z=Tp=Yc1NnoQ@mail.gmail.com","subject":"Re: [PATCH v4 09/24] ls-files.c: use index api","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-30T20:08:23Z","receivedAt":"2013-11-30T20:08:23Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Antoine Pelisse <apelisse@gmail.com> writes:\n\n> On Wed, Nov 27, 2013 at 1:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n>> @@ -447,6 +463,7 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>>         struct dir_struct dir;\n>>         struct exclude_list *el;\n>>         struct string_list exclude_list = STRING_LIST_INIT_NODUP;\n>> +       struct filter_opts *opts = xmalloc(sizeof(*opts));\n>>         struct option builtin_ls_files_options[] = {\n>>                 { OPTION_CALLBACK, 'z', NULL, NULL, NULL,\n>>                         N_(\"paths are separated with NUL character\"),\n>> @@ -512,9 +529,6 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>>                 prefix_len = strlen(prefix);\n>>         git_config(git_default_config, NULL);\n>>\n>> -       if (read_cache() < 0)\n>> -               die(\"index file corrupt\");\n>> -\n>>         argc = parse_options(argc, argv, prefix, builtin_ls_files_options,\n>>                         ls_files_usage, 0);\n>>         el = add_exclude_list(&dir, EXC_CMDL, \"--exclude option\");\n>> @@ -550,6 +564,24 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n>>                        PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n>>                        prefix, argv);\n>>\n>> +       if (!with_tree && !needs_trailing_slash_stripped()) {\n>> +               memset(opts, 0, sizeof(*opts));\n>> +               opts->pathspec = &pathspec;\n>> +               opts->read_staged = 1;\n>> +               if (show_resolve_undo)\n>> +                       opts->read_resolve_undo = 1;\n>> +               if (read_cache_filtered(opts) < 0)\n>> +                       die(\"index file corrupt\");\n>> +       } else {\n>> +               if (read_cache() < 0)\n>> +                       die(\"index file corrupt\");\n>> +               parse_pathspec(&pathspec, 0,\n>> +                              PATHSPEC_PREFER_CWD |\n>> +                              PATHSPEC_STRIP_SUBMODULE_SLASH_CHEAP,\n>> +                              prefix, argv);\n>> +\n>> +       }\n>> +\n>\n> Would it make sense to move the declaration of \"opts\" as a non-pointer\n> to the block where it's used ?\n\nYes, I think that would make sense, will do so in the re-roll. Thanks!\n"},{"id":"231313","messageId":"87ob51y7l0.fsf@gmail.com","threadId":"35410","inReplyTo":"CALWbr2ybgpdDJLRKA4zwPmRE4LQv4VWLJM6Jv0dO-GyHREGW7g@mail.gmail.com","subject":"Re: [PATCH v4 12/24] read-cache: read index-v5","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-30T20:10:03Z","receivedAt":"2013-11-30T20:10:03Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Antoine Pelisse <apelisse@gmail.com> writes:\n\n> On Wed, Nov 27, 2013 at 1:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n>> Make git read the index file version 5 without complaining.\n>>\n>> This version of the reader reads neither the cache-tree\n>> nor the resolve undo data, however, it won't choke on an\n>> index that includes such data.\n>>\n>> Helped-by: Junio C Hamano <gitster@pobox.com>\n>> Helped-by: Nguyen Thai Ngoc Duy <pclouds@gmail.com>\n>> Helped-by: Thomas Rast <trast@student.ethz.ch>\n>> Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n>> ---\n>> [...]\n>> +static struct directory_entry *read_directories(unsigned int *dir_offset,\n>> +                                               unsigned int *dir_table_offset,\n>> +                                               void *mmap, int mmap_size)\n>\n> Minor nit: why is this mmap_size \"int\" while all others are \"unsigned long\" ?\n\nThanks for catching that.  It should be \"unsigned long\" here to.  Will\nfix in the re-roll.\n\n>> [...]\n>> +static int read_entry(struct cache_entry **ce, char *pathname, size_t pathlen,\n>> +                     void *mmap, unsigned long mmap_size,\n>> +                     unsigned int first_entry_offset,\n>> +                     unsigned int foffsetblock)\n>> [...]\n>> +static int read_entries(struct index_state *istate, struct directory_entry *de,\n>> +                       unsigned int first_entry_offset, void *mmap,\n>> +                       unsigned long mmap_size, unsigned int *nr,\n>> +                       unsigned int foffsetblock)\n>> [...]\n>> +static int read_index_v5(struct index_state *istate, void *mmap,\n>> +                        unsigned long mmap_size, struct filter_opts *opts)\n"},{"id":"231314","messageId":"87li05y6sq.fsf@gmail.com","threadId":"35410","inReplyTo":"CALWbr2xUMHSU0MV-6nVbN4_eSMoj3Eyc_Ta_CxTwZ_Y8tLfbdQ@mail.gmail.com","subject":"Re: [PATCH v4 12/24] read-cache: read index-v5","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-11-30T20:27:01Z","receivedAt":"2013-11-30T20:27:01Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Antoine Pelisse <apelisse@gmail.com> writes:\n\n> On Wed, Nov 27, 2013 at 1:00 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n>> +static int verify_hdr(void *mmap, unsigned long size)\n>> +{\n>> +       uint32_t *filecrc;\n>> +       unsigned int header_size;\n>> +       struct cache_header *hdr;\n>> +       struct cache_header_v5 *hdr_v5;\n>> +\n>> +       if (size < sizeof(struct cache_header)\n>> +           + sizeof (struct cache_header_v5) + 4)\n>> +               die(\"index file smaller than expected\");\n>> +\n>> +       hdr = mmap;\n>> +       hdr_v5 = ptr_add(mmap, sizeof(*hdr));\n>> +       /* Size of the header + the size of the extensionoffsets */\n>> +       header_size = sizeof(*hdr) + sizeof(*hdr_v5) + hdr_v5->hdr_nextension * 4;\n>> +       /* Initialize crc */\n>> +       filecrc = ptr_add(mmap, header_size);\n>> +       if (!check_crc32(0, hdr, header_size, ntohl(*filecrc)))\n>> +               return error(\"bad index file header crc signature\");\n>> +       return 0;\n>> +}\n>\n> I find it curious that we actually need a value from the header (and\n> use it for pointer arithmetic) to check that the header is valid. The\n> application will crash before the crc is checked if\n> hdr_v5->hdr_nextensions is corrupted. Or am I missing something ?\n\nGood catch, I'm the one that was missing something here.  We still need\nto use the value from the header before calculating the crc, but should\ncheck if header_size - 4 is less than the total size of the index file.\nThen even if the header is corrupted we won't read anything that is not\nmmap'ed and thus won't crash.\n\nThis guard should also be included for everything else that checks the\ncrc checksum, as that has the same problems and the calculated place in\nthe file for the crc might be after the end of the file.\n\nThanks, will fix in the re-roll.\n"},{"id":"231784","messageId":"87vbyyfi0c.fsf@gmail.com","threadId":"35410","inReplyTo":"1385553659-9928-1-git-send-email-t.gummerer@gmail.com","subject":"Re: [PATCH v4 00/24] Index-v5","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2013-12-09T10:14:43Z","receivedAt":"2013-12-09T10:14:43Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> Hi,\n>\n> previous rounds (without api) are at $gmane/202752, $gmane/202923,\n> $gmane/203088 and $gmane/203517, the previous rounds with api were at\n> $gmane/229732, $gmane/230210 and $gmane/232488.  Thanks to Duy for\n> reviewing the the last round and Junio, Ramsay and Eric for additional\n> comments.\n>\n> Since the last round I've added a POC for partial writing, resulting\n> in the following performance improvements for update-index:\n>\n> Test                                        1063432           HEAD\n> ------------------------------------------------------------------------------------\n> 0003.2: v[23]: update-index                 0.60(0.38+0.20)   0.76(0.36+0.17) +26.7%\n> 0003.3: v[23]: grep nonexistent -- subdir   0.28(0.17+0.11)   0.28(0.18+0.09) +0.0%\n> 0003.4: v[23]: ls-files -- subdir           0.26(0.15+0.10)   0.24(0.14+0.09) -7.7%\n> 0003.7: v[23] update-index                  0.59(0.36+0.22)   0.58(0.36+0.20) -1.7%\n> 0003.9: v4: update-index                    0.46(0.28+0.17)   0.45(0.30+0.11) -2.2%\n> 0003.10: v4: grep nonexistent -- subdir     0.26(0.14+0.11)   0.21(0.14+0.07) -19.2%\n> 0003.11: v4: ls-files -- subdir             0.24(0.14+0.10)   0.20(0.12+0.08) -16.7%\n> 0003.14: v4 update-index                    0.49(0.31+0.18)   0.65(0.34+0.17) +32.7%\n> 0003.16: v5: update-index                   0.53(0.30+0.22)   0.50(0.28+0.20) -5.7%\n> 0003.17: v5: ls-files                       0.27(0.15+0.12)   0.27(0.17+0.10) +0.0%\n> 0003.18: v5: grep nonexistent -- subdir     0.02(0.01+0.01)   0.03(0.01+0.01) +50.0%\n> 0003.19: v5: ls-files -- subdir             0.02(0.00+0.02)   0.02(0.01+0.01) +0.0%\n> 0003.22: v5 update-index                    0.53(0.29+0.23)   0.02(0.01+0.01) -96.2%\n>\n> Given this, I don't think a complete change of the in-core format for\n> the cache-entries is necessary to take full advantage of the new index\n> file format.  Instead some changes to the current in-core format would\n> work well with the new on-disk format.\n>\n> The current in-memory format fits the internal needs of git fairly well,\n> so I don't think changing it to fit a better index file format would\n> make a lot of sense, given that we can take advantage of the new format\n> with the existing in-memory format.\n\nAny more opinions on this series?  I've applied the changes suggested by\nDuy, Antoine and Eric locally, but I wouldn't want to spam the list with\nthe whole series without a chance of this being applied.  How do you\nwant me to proceed?\n\n> This series doesn't use kb/fast-hashmap yet, but that should be fairly\n> simple to change if the series is deemed a good change.  The\n> performance tests for update-index test require\n> tg/perf-lib-test-perf-cleanup.\n>\n> Other changes, made following the review comments are:\n>\n> documentation: add documentation of the index-v5 file format\n>   - Update documentation that directory flags are now 32-bits.  That\n>     makes aligned access simpler\n>   - offset_to_offset is no longer included in the checksum for files.\n>     It's unnecessary.\n>\n> read-cache: read index-v5\n>   - Add fix for reading with different level pathspecs given\n>   - Use init_directory_entry to initialize all fields in a new\n>     directory entry\n>   - use memset to simplify the create_new_conflict function\n>   - Add comments to explain -5 when reading directories and files\n>   - Add comments for the more complex functions\n>   - Add name flex_array to the end of ondisk_directory_entry for\n>     simplified reading\n>   - Add name flex_array to the end of ondisk_cache_entry for\n>     simplified reading\n>   - Move conflict reading functions to next patch\n>   - mark functions as static when they are\n>\n> read-cache: read resolve-undo data\n>   - Add comments for the more complex function\n>   - Read conflicts + resolve undo data as extension\n>\n> read-cache: read cache-tree in index-v5\n>   - Add comments for the more complex function\n>   - Instead of sorting the directory entries, sort the cache-tree\n>     directly.  This also required changing the algorithms with which\n>     the cache entries are extracted from the directory tree.\n>\n> read-cache: write index-v5\n>   - Free pointers allocated by super_directory\n>   - Rewrite condition as suggested by Duy\n>   - Don't check for CE_REMOVE'd entries in the writing code, they are\n>     already checked in the compile_directory_data code\n>   - Remove overly complicated directory size calculation since flags\n>     are now 32-bits\n>\n> read-cache: write resolve-undo data for index-v5\n>   - Free pointers allocated by super_directory\n>   - Write conflicts + resolve undo data as extension\n>\n> introduce GIT_INDEX_VERSION environment variable\n>   - Add documentation for GIT_INDEX_VERSION\n>\n> test-lib: allow setting the index format version\n>\n> Removed commits:\n>   - read-cache: don't check uid, gid, ino\n>   - read-cache: use fixed width integer types (independently in pu)\n>   - read-cache: clear version in discard_index()\n>\n> Typos fixed as suggested by Eric Sunshine\n>\n> Thomas Gummerer (22):\n>   read-cache: split index file version specific functionality\n>   read-cache: move index v2 specific functions to their own file\n>   read-cache: Re-read index if index file changed\n>   add documentation for the index api\n>   read-cache: add index reading api\n>   make sure partially read index is not changed\n>   grep.c: use index api\n>   ls-files.c: use index api\n>   documentation: add documentation of the index-v5 file format\n>   read-cache: make in-memory format aware of stat_crc\n>   read-cache: read index-v5\n>   read-cache: read resolve-undo data\n>   read-cache: read cache-tree in index-v5\n>   read-cache: write index-v5\n>   read-cache: write index-v5 cache-tree data\n>   read-cache: write resolve-undo data for index-v5\n>   update-index.c: rewrite index when index-version is given\n>   introduce GIT_INDEX_VERSION environment variable\n>   test-lib: allow setting the index format version\n>   t1600: add index v5 specific tests\n>   POC for partial writing\n>   perf: add partial writing test\n>\n> Thomas Rast (1):\n>   p0003-index.sh: add perf test for the index formats\n>\n>  Documentation/git.txt                            |    5 +\n>  Documentation/technical/api-in-core-index.txt    |   56 +-\n>  Documentation/technical/index-file-format-v5.txt |  294 +++++\n>  Makefile                                         |   10 +\n>  builtin/apply.c                                  |    2 +\n>  builtin/grep.c                                   |   69 +-\n>  builtin/ls-files.c                               |   36 +-\n>  builtin/update-index.c                           |   50 +-\n>  cache-tree.c                                     |   15 +-\n>  cache-tree.h                                     |    2 +\n>  cache.h                                          |  115 +-\n>  lockfile.c                                       |    2 +-\n>  read-cache-v2.c                                  |  561 +++++++++\n>  read-cache-v5.c                                  | 1406 ++++++++++++++++++++++\n>  read-cache.c                                     |  691 +++--------\n>  read-cache.h                                     |   67 ++\n>  resolve-undo.c                                   |    1 +\n>  t/perf/p0003-index.sh                            |   74 ++\n>  t/t1600-index-v5.sh                              |   25 +\n>  t/t2101-update-index-reupdate.sh                 |   12 +-\n>  t/test-lib-functions.sh                          |    5 +\n>  t/test-lib.sh                                    |    3 +\n>  test-index-version.c                             |    6 +\n>  unpack-trees.c                                   |    3 +-\n>  24 files changed, 2921 insertions(+), 589 deletions(-)\n>  create mode 100644 Documentation/technical/index-file-format-v5.txt\n>  create mode 100644 read-cache-v2.c\n>  create mode 100644 read-cache-v5.c\n>  create mode 100644 read-cache.h\n>  create mode 100755 t/perf/p0003-index.sh\n>  create mode 100755 t/t1600-index-v5.sh\n>\n> --\n> 1.8.4.2\n>\n\n--\nThomas\n"}]}