{"thread":{"id":"20539","subject":"[RFC PATCH v3 0/8] Sparse checkout","startedAt":"2009-08-11T15:43:58Z","lastAt":"2009-08-18T16:00:38Z","messageCount":53,"participants":["Nguyễn Thái Ngọc Duy","skillzero@gmail.com","Jakub Narebski","Nguyen Thai Ngoc Duy","Junio C Hamano","Johannes Sixt","Raja R Harinath","Johannes Schindelin"],"isPatch":true,"patchVersion":3,"patchTotal":8},"messages":[{"id":"120282","messageId":"1250005446-12047-1-git-send-email-pclouds@gmail.com","threadId":"20539","inReplyTo":null,"subject":"[RFC PATCH v3 0/8] Sparse checkout","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-11T15:43:58Z","receivedAt":"2009-08-11T15:43:58Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Continuing the endless RFCs of sparse checkout, this series drops the sparse hook\nin favor of .git/info/sparse. Changes from the last version\n\n\n  Prevent diff machinery from examining assume-unchanged entries on worktree\n\n    \"if (ce_uptodate(ce) || CE_VALID)\" is updated, as well as the corresponding test\n\n\n  Avoid writing to buffer in add_excludes_from_file_1()\n\n    Splitted out from the old second patch, as suggested by Johannes\n\n\n  Read .gitignore from index if it is assume-unchanged\n\n    read_index_data() is renamed. Commit message mentions add_excludes_from_file()\n\n\n  excluded_1(): support exclude \"directories\" in index\n\n    This one is new because index does not have \"directory\", more comments in the patch\n\n\n  dir.c: export excluded_1() and add_excludes_from_file_1()\n\n    New too, exported for use in unpack-trees.c\n\n\n  unpack-trees.c: generalize verify_* functions\n\n    Splitted out of the old third patch for easier review\n\n\n  Support sparse checkout in unpack_trees() and read-tree\n\n    Read .git/info/sparse instead of .git/hooks/sparse\n\n    \n  --sparse for porcelains\n    RFC patch\n\n Documentation/technical/api-directory-listing.txt |    3 +\n builtin-checkout.c                                |    4 +\n builtin-clean.c                                   |    5 +-\n builtin-ls-files.c                                |    4 +-\n builtin-merge.c                                   |    5 +-\n builtin-read-tree.c                               |    4 +-\n cache.h                                           |    3 +\n diff-lib.c                                        |    6 +-\n dir.c                                             |  101 +++++++++++------\n dir.h                                             |    4 +\n git-pull.sh                                       |    6 +-\n t/t1009-read-tree-sparse.sh                       |   47 ++++++++\n t/t3001-ls-files-others-exclude.sh                |   22 ++++\n t/t4039-diff-assume-unchanged.sh                  |   31 ++++++\n t/t7300-clean.sh                                  |   19 ++++\n unpack-trees.c                                    |  121 ++++++++++++++++++++-\n unpack-trees.h                                    |    3 +\n 17 files changed, 340 insertions(+), 48 deletions(-)\n create mode 100755 t/t1009-read-tree-sparse.sh\n create mode 100755 t/t4039-diff-assume-unchanged.sh\n"},{"id":"120283","messageId":"1250005446-12047-2-git-send-email-pclouds@gmail.com","threadId":"20539","inReplyTo":"1250005446-12047-1-git-send-email-pclouds@gmail.com","subject":"[RFC PATCH v3 1/8] Prevent diff machinery from examining assume-unchanged entries on worktree","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-11T15:43:59Z","receivedAt":"2009-08-11T15:43:59Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n diff-lib.c                       |    6 ++++--\n t/t4039-diff-assume-unchanged.sh |   31 +++++++++++++++++++++++++++++++\n 2 files changed, 35 insertions(+), 2 deletions(-)\n create mode 100755 t/t4039-diff-assume-unchanged.sh\n\ndiff --git a/diff-lib.c b/diff-lib.c\nindex b7813af..e5b9fe0 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -162,7 +162,8 @@ int run_diff_files(struct rev_info *revs, unsigned int option)\n \t\tif (ce_uptodate(ce))\n \t\t\tcontinue;\n \n-\t\tchanged = check_removed(ce, &st);\n+\t\t/* If CE_VALID is set, don't look at workdir for file removal */\n+\t\tchanged = (ce->ce_flags & CE_VALID) ? 0 : check_removed(ce, &st);\n \t\tif (changed) {\n \t\t\tif (changed < 0) {\n \t\t\t\tperror(ce->name);\n@@ -337,6 +338,8 @@ static void do_oneway_diff(struct unpack_trees_options *o,\n \tstruct rev_info *revs = o->unpack_data;\n \tint match_missing, cached;\n \n+\t/* if the entry is not checked out, don't examine work tree */\n+\tcached = o->index_only || (idx && (idx->ce_flags & CE_VALID));\n \t/*\n \t * Backward compatibility wart - \"diff-index -m\" does\n \t * not mean \"do not ignore merges\", but \"match_missing\".\n@@ -344,7 +347,6 @@ static void do_oneway_diff(struct unpack_trees_options *o,\n \t * But with the revision flag parsing, that's found in\n \t * \"!revs->ignore_merges\".\n \t */\n-\tcached = o->index_only;\n \tmatch_missing = !revs->ignore_merges;\n \n \tif (cached && idx && ce_stage(idx)) {\ndiff --git a/t/t4039-diff-assume-unchanged.sh b/t/t4039-diff-assume-unchanged.sh\nnew file mode 100755\nindex 0000000..9d9498b\n--- /dev/null\n+++ b/t/t4039-diff-assume-unchanged.sh\n@@ -0,0 +1,31 @@\n+#!/bin/sh\n+\n+test_description='diff with assume-unchanged entries'\n+\n+. ./test-lib.sh\n+\n+# external diff has been tested in t4020-diff-external.sh\n+\n+test_expect_success 'setup' '\n+\techo zero > zero &&\n+\tgit add zero &&\n+\tgit commit -m zero &&\n+\techo one > one &&\n+\techo two > two &&\n+\tgit add one two &&\n+\tgit commit -m onetwo &&\n+\tgit update-index --assume-unchanged one &&\n+\techo borked >> one &&\n+\ttest \"$(git ls-files -v one)\" = \"h one\"\n+'\n+\n+test_expect_success 'diff-index does not examine assume-unchanged entries' '\n+\tgit diff-index HEAD^ -- one | grep -q 5626abf0f72e58d7a153368ba57db4c673c0e171\n+'\n+\n+test_expect_success 'diff-files does not examine assume-unchanged entries' '\n+\trm one &&\n+\ttest -z \"$(git diff-files -- one)\"\n+'\n+\n+test_done\n-- \n1.6.3.GIT\n"},{"id":"120284","messageId":"1250005446-12047-3-git-send-email-pclouds@gmail.com","threadId":"20539","inReplyTo":"1250005446-12047-2-git-send-email-pclouds@gmail.com","subject":"[RFC PATCH v3 2/8] Avoid writing to buffer in add_excludes_from_file_1()","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-11T15:44:00Z","receivedAt":"2009-08-11T15:44:00Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"In the next patch, the buffer that is being used within\nadd_excludes_from_file_1() comes from another function and does not\nhave extra space to put \\n at the end.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n dir.c |    5 ++---\n 1 files changed, 2 insertions(+), 3 deletions(-)\n\ndiff --git a/dir.c b/dir.c\nindex e05b850..1170d64 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -229,10 +229,9 @@ static int add_excludes_from_file_1(const char *fname,\n \n \tif (buf_p)\n \t\t*buf_p = buf;\n-\tbuf[size++] = '\\n';\n \tentry = buf;\n-\tfor (i = 0; i < size; i++) {\n-\t\tif (buf[i] == '\\n') {\n+\tfor (i = 0; i <= size; i++) {\n+\t\tif (i == size || buf[i] == '\\n') {\n \t\t\tif (entry != buf + i && entry[0] != '#') {\n \t\t\t\tbuf[i - (i && buf[i-1] == '\\r')] = 0;\n \t\t\t\tadd_exclude(entry, base, baselen, which);\n-- \n1.6.3.GIT\n"},{"id":"120286","messageId":"1250005446-12047-4-git-send-email-pclouds@gmail.com","threadId":"20539","inReplyTo":"1250005446-12047-3-git-send-email-pclouds@gmail.com","subject":"[RFC PATCH v3 3/8] Read .gitignore from index if it is assume-unchanged","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-11T15:44:01Z","receivedAt":"2009-08-11T15:44:01Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"In sparse checkout mode (aka CE_VALID or assume-unchanged) some files\nmay be missing from working directory. If some of those files are\n.gitignore, it will affect how git excludes files.\n\nBecause those files are by definition \"assume unchanged\" we can\ninstead read them from index. This adds index as a prerequisite for\ndirectory listing. At the moment directory listing is used by \"git\nclean\", \"git add\", \"git ls-files\" and \"git status\"/\"git commit\" and\nunpack_trees()-related commands.  These commands have been\nchecked/modified to populate index before doing directory listing.\n\nadd_excludes_from_file() does not enable this feature, because it\nis used to read .git/info/exclude and some explicit files specified\nby \"git ls-files\".\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n Documentation/technical/api-directory-listing.txt |    3 +\n builtin-clean.c                                   |    5 +-\n builtin-ls-files.c                                |    4 +-\n dir.c                                             |   66 ++++++++++++++------\n t/t3001-ls-files-others-exclude.sh                |   22 +++++++\n t/t7300-clean.sh                                  |   19 ++++++\n 6 files changed, 97 insertions(+), 22 deletions(-)\n\ndiff --git a/Documentation/technical/api-directory-listing.txt b/Documentation/technical/api-directory-listing.txt\nindex 5bbd18f..7d0e282 100644\n--- a/Documentation/technical/api-directory-listing.txt\n+++ b/Documentation/technical/api-directory-listing.txt\n@@ -58,6 +58,9 @@ The result of the enumeration is left in these fields::\n Calling sequence\n ----------------\n \n+* Ensure the_index is populated as it may have CE_VALID entries that\n+  affect directory listing.\n+\n * Prepare `struct dir_struct dir` and clear it with `memset(&dir, 0,\n   sizeof(dir))`.\n \ndiff --git a/builtin-clean.c b/builtin-clean.c\nindex 2d8c735..d917472 100644\n--- a/builtin-clean.c\n+++ b/builtin-clean.c\n@@ -71,8 +71,11 @@ int cmd_clean(int argc, const char **argv, const char *prefix)\n \n \tdir.flags |= DIR_SHOW_OTHER_DIRECTORIES;\n \n-\tif (!ignored)\n+\tif (!ignored) {\n+\t\tif (read_cache() < 0)\n+\t\t\tdie(\"index file corrupt\");\n \t\tsetup_standard_excludes(&dir);\n+\t}\n \n \tpathspec = get_pathspec(prefix, argv);\n \tread_cache();\ndiff --git a/builtin-ls-files.c b/builtin-ls-files.c\nindex f473220..d1a23c4 100644\n--- a/builtin-ls-files.c\n+++ b/builtin-ls-files.c\n@@ -481,6 +481,9 @@ int cmd_ls_files(int argc, const char **argv, const char *prefix)\n \t\tprefix_offset = strlen(prefix);\n \tgit_config(git_default_config, NULL);\n \n+\tif (read_cache() < 0)\n+\t\tdie(\"index file corrupt\");\n+\n \targc = parse_options(argc, argv, prefix, builtin_ls_files_options,\n \t\t\tls_files_usage, 0);\n \tif (show_tag || show_valid_bit) {\n@@ -508,7 +511,6 @@ int cmd_ls_files(int argc, const char **argv, const char *prefix)\n \tpathspec = get_pathspec(prefix, argv);\n \n \t/* be nice with submodule paths ending in a slash */\n-\tread_cache();\n \tif (pathspec)\n \t\tstrip_trailing_slash_from_submodules();\n \ndiff --git a/dir.c b/dir.c\nindex 1170d64..66b485c 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -200,11 +200,36 @@ void add_exclude(const char *string, const char *base,\n \twhich->excludes[which->nr++] = x;\n }\n \n+static void *read_assume_unchanged_from_index(const char *path, size_t *size)\n+{\n+\tint pos, len;\n+\tunsigned long sz;\n+\tenum object_type type;\n+\tvoid *data;\n+\tstruct index_state *istate = &the_index;\n+\n+\tlen = strlen(path);\n+\tpos = index_name_pos(istate, path, len);\n+\tif (pos < 0)\n+\t\treturn NULL;\n+\t/* only applies to CE_VALID entries */\n+\tif (!(istate->cache[pos]->ce_flags & CE_VALID))\n+\t\treturn NULL;\n+\tdata = read_sha1_file(istate->cache[pos]->sha1, &type, &sz);\n+\tif (!data || type != OBJ_BLOB) {\n+\t\tfree(data);\n+\t\treturn NULL;\n+\t}\n+\t*size = xsize_t(sz);\n+\treturn data;\n+}\n+\n static int add_excludes_from_file_1(const char *fname,\n \t\t\t\t    const char *base,\n \t\t\t\t    int baselen,\n \t\t\t\t    char **buf_p,\n-\t\t\t\t    struct exclude_list *which)\n+\t\t\t\t    struct exclude_list *which,\n+\t\t\t\t    int check_index)\n {\n \tstruct stat st;\n \tint fd, i;\n@@ -212,20 +237,26 @@ static int add_excludes_from_file_1(const char *fname,\n \tchar *buf, *entry;\n \n \tfd = open(fname, O_RDONLY);\n-\tif (fd < 0 || fstat(fd, &st) < 0)\n-\t\tgoto err;\n-\tsize = xsize_t(st.st_size);\n-\tif (size == 0) {\n-\t\tclose(fd);\n-\t\treturn 0;\n+\tif (fd < 0 || fstat(fd, &st) < 0) {\n+\t\tif (0 <= fd)\n+\t\t\tclose(fd);\n+\t\tif (!check_index ||\n+\t\t    (buf = read_assume_unchanged_from_index(fname, &size)) == NULL)\n+\t\t\treturn -1;\n \t}\n-\tbuf = xmalloc(size+1);\n-\tif (read_in_full(fd, buf, size) != size)\n-\t{\n-\t\tfree(buf);\n-\t\tgoto err;\n+\telse {\n+\t\tsize = xsize_t(st.st_size);\n+\t\tif (size == 0) {\n+\t\t\tclose(fd);\n+\t\t\treturn 0;\n+\t\t}\n+\t\tbuf = xmalloc(size);\n+\t\tif (read_in_full(fd, buf, size) != size) {\n+\t\t\tclose(fd);\n+\t\t\treturn -1;\n+\t\t}\n+\t\tclose(fd);\n \t}\n-\tclose(fd);\n \n \tif (buf_p)\n \t\t*buf_p = buf;\n@@ -240,17 +271,12 @@ static int add_excludes_from_file_1(const char *fname,\n \t\t}\n \t}\n \treturn 0;\n-\n- err:\n-\tif (0 <= fd)\n-\t\tclose(fd);\n-\treturn -1;\n }\n \n void add_excludes_from_file(struct dir_struct *dir, const char *fname)\n {\n \tif (add_excludes_from_file_1(fname, \"\", 0, NULL,\n-\t\t\t\t     &dir->exclude_list[EXC_FILE]) < 0)\n+\t\t\t\t     &dir->exclude_list[EXC_FILE], 0) < 0)\n \t\tdie(\"cannot use %s as an exclude file\", fname);\n }\n \n@@ -301,7 +327,7 @@ static void prep_exclude(struct dir_struct *dir, const char *base, int baselen)\n \t\tstrcpy(dir->basebuf + stk->baselen, dir->exclude_per_dir);\n \t\tadd_excludes_from_file_1(dir->basebuf,\n \t\t\t\t\t dir->basebuf, stk->baselen,\n-\t\t\t\t\t &stk->filebuf, el);\n+\t\t\t\t\t &stk->filebuf, el, 1);\n \t\tdir->exclude_stack = stk;\n \t\tcurrent = stk->baselen;\n \t}\ndiff --git a/t/t3001-ls-files-others-exclude.sh b/t/t3001-ls-files-others-exclude.sh\nindex c65bca8..fdd5dd8 100755\n--- a/t/t3001-ls-files-others-exclude.sh\n+++ b/t/t3001-ls-files-others-exclude.sh\n@@ -64,6 +64,8 @@ two/*.4\n echo '!*.2\n !*.8' >one/two/.gitignore\n \n+allignores='.gitignore one/.gitignore one/two/.gitignore'\n+\n test_expect_success \\\n     'git ls-files --others with various exclude options.' \\\n     'git ls-files --others \\\n@@ -85,6 +87,26 @@ test_expect_success \\\n        >output &&\n      test_cmp expect output'\n \n+test_expect_success 'setup sparse gitignore' '\n+\tgit add $allignores &&\n+\tgit update-index --assume-unchanged $allignores &&\n+\trm $allignores\n+'\n+\n+test_expect_success \\\n+    'git ls-files --others with various exclude options.' \\\n+    'git ls-files --others \\\n+       --exclude=\\*.6 \\\n+       --exclude-per-directory=.gitignore \\\n+       --exclude-from=.git/ignore \\\n+       >output &&\n+     test_cmp expect output'\n+\n+test_expect_success 'restore gitignore' '\n+\tgit checkout $allignores &&\n+\trm .git/index\n+'\n+\n cat > excludes-file <<\\EOF\n *.[1-8]\n e*\ndiff --git a/t/t7300-clean.sh b/t/t7300-clean.sh\nindex 929d5d4..4886d5f 100755\n--- a/t/t7300-clean.sh\n+++ b/t/t7300-clean.sh\n@@ -22,6 +22,25 @@ test_expect_success 'setup' '\n \n '\n \n+test_expect_success 'git clean with assume-unchanged .gitignore' '\n+\tgit update-index --assume-unchanged .gitignore &&\n+\trm .gitignore &&\n+\tmkdir -p build docs &&\n+\ttouch a.out src/part3.c docs/manual.txt obj.o build/lib.so &&\n+\tgit clean &&\n+\ttest -f Makefile &&\n+\ttest -f README &&\n+\ttest -f src/part1.c &&\n+\ttest -f src/part2.c &&\n+\ttest ! -f a.out &&\n+\ttest ! -f src/part3.c &&\n+\ttest -f docs/manual.txt &&\n+\ttest -f obj.o &&\n+\ttest -f build/lib.so &&\n+\tgit update-index --no-assume-unchanged .gitignore &&\n+\tgit checkout .gitignore\n+'\n+\n test_expect_success 'git clean' '\n \n \tmkdir -p build docs &&\n-- \n1.6.3.GIT\n"},{"id":"120285","messageId":"1250005446-12047-5-git-send-email-pclouds@gmail.com","threadId":"20539","inReplyTo":"1250005446-12047-4-git-send-email-pclouds@gmail.com","subject":"[RFC PATCH v3 4/8] excluded_1(): support exclude \"directories\" in index","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-11T15:44:02Z","receivedAt":"2009-08-11T15:44:02Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Index does not really have \"directories\", attempts to match \"foo/\"\nagainst index will fail unless someone tries to reconstruct directories\nfrom a list of file.\n\nObserving that dtype in this function can never be NULL (otherwise\nit would segfault), dtype NULL will be used to say \"hey.. you are\nmatching against index\" and behave properly.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n  Having dtype to segfault when dtype is NULL is nice, but I found\n  no way else to sneak the new code in. Defining DT_INDEX may clash\n  existing system definitions..\n\n\n dir.c |    6 ++++++\n 1 files changed, 6 insertions(+), 0 deletions(-)\n\ndiff --git a/dir.c b/dir.c\nindex 66b485c..c990938 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -350,6 +350,12 @@ static int excluded_1(const char *pathname,\n \t\t\tint to_exclude = x->to_exclude;\n \n \t\t\tif (x->flags & EXC_FLAG_MUSTBEDIR) {\n+\t\t\t\tif (!dtype) {\n+\t\t\t\t\tif (!prefixcmp(pathname, exclude))\n+\t\t\t\t\t\treturn to_exclude;\n+\t\t\t\t\telse\n+\t\t\t\t\t\tcontinue;\n+\t\t\t\t}\n \t\t\t\tif (*dtype == DT_UNKNOWN)\n \t\t\t\t\t*dtype = get_dtype(NULL, pathname, pathlen);\n \t\t\t\tif (*dtype != DT_DIR)\n-- \n1.6.3.GIT\n"},{"id":"120287","messageId":"1250005446-12047-6-git-send-email-pclouds@gmail.com","threadId":"20539","inReplyTo":"1250005446-12047-5-git-send-email-pclouds@gmail.com","subject":"[RFC PATCH v3 5/8] dir.c: export excluded_1() and add_excludes_from_file_1()","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-11T15:44:03Z","receivedAt":"2009-08-11T15:44:03Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"These functions are used to handle .gitignore. They are now exported\nso that sparse checkout can reuse.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n dir.c |   32 ++++++++++++++++----------------\n dir.h |    4 ++++\n 2 files changed, 20 insertions(+), 16 deletions(-)\n\ndiff --git a/dir.c b/dir.c\nindex c990938..bc35586 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -224,12 +224,12 @@ static void *read_assume_unchanged_from_index(const char *path, size_t *size)\n \treturn data;\n }\n \n-static int add_excludes_from_file_1(const char *fname,\n-\t\t\t\t    const char *base,\n-\t\t\t\t    int baselen,\n-\t\t\t\t    char **buf_p,\n-\t\t\t\t    struct exclude_list *which,\n-\t\t\t\t    int check_index)\n+int add_excludes_from_file_to_list(const char *fname,\n+\t\t\t\t   const char *base,\n+\t\t\t\t   int baselen,\n+\t\t\t\t   char **buf_p,\n+\t\t\t\t   struct exclude_list *which,\n+\t\t\t\t   int check_index)\n {\n \tstruct stat st;\n \tint fd, i;\n@@ -275,8 +275,8 @@ static int add_excludes_from_file_1(const char *fname,\n \n void add_excludes_from_file(struct dir_struct *dir, const char *fname)\n {\n-\tif (add_excludes_from_file_1(fname, \"\", 0, NULL,\n-\t\t\t\t     &dir->exclude_list[EXC_FILE], 0) < 0)\n+\tif (add_excludes_from_file_to_list(fname, \"\", 0, NULL,\n+\t\t\t\t\t   &dir->exclude_list[EXC_FILE], 0) < 0)\n \t\tdie(\"cannot use %s as an exclude file\", fname);\n }\n \n@@ -325,9 +325,9 @@ static void prep_exclude(struct dir_struct *dir, const char *base, int baselen)\n \t\tmemcpy(dir->basebuf + current, base + current,\n \t\t       stk->baselen - current);\n \t\tstrcpy(dir->basebuf + stk->baselen, dir->exclude_per_dir);\n-\t\tadd_excludes_from_file_1(dir->basebuf,\n-\t\t\t\t\t dir->basebuf, stk->baselen,\n-\t\t\t\t\t &stk->filebuf, el, 1);\n+\t\tadd_excludes_from_file_to_list(dir->basebuf,\n+\t\t\t\t\t       dir->basebuf, stk->baselen,\n+\t\t\t\t\t       &stk->filebuf, el, 1);\n \t\tdir->exclude_stack = stk;\n \t\tcurrent = stk->baselen;\n \t}\n@@ -337,9 +337,9 @@ static void prep_exclude(struct dir_struct *dir, const char *base, int baselen)\n /* Scan the list and let the last match determine the fate.\n  * Return 1 for exclude, 0 for include and -1 for undecided.\n  */\n-static int excluded_1(const char *pathname,\n-\t\t      int pathlen, const char *basename, int *dtype,\n-\t\t      struct exclude_list *el)\n+int excluded_from_list(const char *pathname,\n+\t\t       int pathlen, const char *basename, int *dtype,\n+\t\t       struct exclude_list *el)\n {\n \tint i;\n \n@@ -413,8 +413,8 @@ int excluded(struct dir_struct *dir, const char *pathname, int *dtype_p)\n \n \tprep_exclude(dir, pathname, basename-pathname);\n \tfor (st = EXC_CMDL; st <= EXC_FILE; st++) {\n-\t\tswitch (excluded_1(pathname, pathlen, basename,\n-\t\t\t\t   dtype_p, &dir->exclude_list[st])) {\n+\t\tswitch (excluded_from_list(pathname, pathlen, basename,\n+\t\t\t\t\t   dtype_p, &dir->exclude_list[st])) {\n \t\tcase 0:\n \t\t\treturn 0;\n \t\tcase 1:\ndiff --git a/dir.h b/dir.h\nindex a631446..472e11e 100644\n--- a/dir.h\n+++ b/dir.h\n@@ -69,7 +69,11 @@ extern int match_pathspec(const char **pathspec, const char *name, int namelen,\n extern int fill_directory(struct dir_struct *dir, const char **pathspec);\n extern int read_directory(struct dir_struct *, const char *path, int len, const char **pathspec);\n \n+extern int excluded_from_list(const char *pathname, int pathlen, const char *basename,\n+\t\t\t      int *dtype, struct exclude_list *el);\n extern int excluded(struct dir_struct *, const char *, int *);\n+extern int add_excludes_from_file_to_list(const char *fname, const char *base, int baselen,\n+\t\t\t\t\t  char **buf_p, struct exclude_list *which, int check_index);\n extern void add_excludes_from_file(struct dir_struct *, const char *fname);\n extern void add_exclude(const char *string, const char *base,\n \t\t\tint baselen, struct exclude_list *which);\n-- \n1.6.3.GIT\n"},{"id":"120290","messageId":"1250005446-12047-7-git-send-email-pclouds@gmail.com","threadId":"20539","inReplyTo":"1250005446-12047-6-git-send-email-pclouds@gmail.com","subject":"[RFC PATCH v3 6/8] unpack-trees.c: generalize verify_* functions","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-11T15:44:04Z","receivedAt":"2009-08-11T15:44:04Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n unpack-trees.c |   23 ++++++++++++++++++-----\n 1 files changed, 18 insertions(+), 5 deletions(-)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 720f7a1..02ea236 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -445,8 +445,9 @@ static int same(struct cache_entry *a, struct cache_entry *b)\n  * When a CE gets turned into an unmerged entry, we\n  * want it to be up-to-date\n  */\n-static int verify_uptodate(struct cache_entry *ce,\n-\t\tstruct unpack_trees_options *o)\n+static int verify_uptodate_1(struct cache_entry *ce,\n+\t\t\t\t   struct unpack_trees_options *o,\n+\t\t\t\t   const char *error_msg)\n {\n \tstruct stat st;\n \n@@ -471,7 +472,13 @@ static int verify_uptodate(struct cache_entry *ce,\n \tif (errno == ENOENT)\n \t\treturn 0;\n \treturn o->gently ? -1 :\n-\t\terror(ERRORMSG(o, not_uptodate_file), ce->name);\n+\t\terror(error_msg, ce->name);\n+}\n+\n+static int verify_uptodate(struct cache_entry *ce,\n+\t\t\t   struct unpack_trees_options *o)\n+{\n+\treturn verify_uptodate_1(ce, o, ERRORMSG(o, not_uptodate_file));\n }\n \n static void invalidate_ce_path(struct cache_entry *ce, struct unpack_trees_options *o)\n@@ -579,8 +586,9 @@ static int icase_exists(struct unpack_trees_options *o, struct cache_entry *dst,\n  * We do not want to remove or overwrite a working tree file that\n  * is not tracked, unless it is ignored.\n  */\n-static int verify_absent(struct cache_entry *ce, const char *action,\n-\t\t\t struct unpack_trees_options *o)\n+static int verify_absent_1(struct cache_entry *ce, const char *action,\n+\t\t\t\t struct unpack_trees_options *o,\n+\t\t\t\t const char *error_msg)\n {\n \tstruct stat st;\n \n@@ -660,6 +668,11 @@ static int verify_absent(struct cache_entry *ce, const char *action,\n \t}\n \treturn 0;\n }\n+static int verify_absent(struct cache_entry *ce, const char *action,\n+\t\t\t struct unpack_trees_options *o)\n+{\n+\treturn verify_absent_1(ce, action, o, ERRORMSG(o, would_lose_untracked));\n+}\n \n static int merged_entry(struct cache_entry *merge, struct cache_entry *old,\n \t\tstruct unpack_trees_options *o)\n-- \n1.6.3.GIT\n"},{"id":"120288","messageId":"1250005446-12047-8-git-send-email-pclouds@gmail.com","threadId":"20539","inReplyTo":"1250005446-12047-7-git-send-email-pclouds@gmail.com","subject":"[RFC PATCH v3 7/8] Support sparse checkout in unpack_trees() and read-tree","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-11T15:44:05Z","receivedAt":"2009-08-11T15:44:05Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This patch makes unpack_trees() look at .git/info/sparse [1] to\ndetermine which files should stay in working directory, after\nmerging, by:\n\n - setting CE_VALID properly so that other operations correctly ignore\n   missing files\n - driving check_updates() to add/remove files in accordance to\n   CE_VALID\n\nThe feature is disabled by default. Use \"read-tree --sparse\" to enable it.\n\n[1] .git/info/sparse has the same syntax as .git/info/exclude. Files\nthat match the patterns will be set as CE_VALID.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin-read-tree.c         |    4 +-\n cache.h                     |    3 +\n t/t1009-read-tree-sparse.sh |   47 ++++++++++++++++++++\n unpack-trees.c              |   98 ++++++++++++++++++++++++++++++++++++++++++-\n unpack-trees.h              |    3 +\n 5 files changed, 153 insertions(+), 2 deletions(-)\n create mode 100755 t/t1009-read-tree-sparse.sh\n\ndiff --git a/builtin-read-tree.c b/builtin-read-tree.c\nindex 9c2d634..888f136 100644\n--- a/builtin-read-tree.c\n+++ b/builtin-read-tree.c\n@@ -31,7 +31,7 @@ static int list_tree(unsigned char *sha1)\n }\n \n static const char * const read_tree_usage[] = {\n-\t\"git read-tree [[-m [--trivial] [--aggressive] | --reset | --prefix=<prefix>] [-u [--exclude-per-directory=<gitignore>] | -i]]  [--index-output=<file>] <tree-ish1> [<tree-ish2> [<tree-ish3>]]\",\n+\t\"git read-tree [[-m [--trivial] [--aggressive] | --reset | --prefix=<prefix>] [-u [--exclude-per-directory=<gitignore>] | -i]] [--sparse] [--index-output=<file>] <tree-ish1> [<tree-ish2> [<tree-ish3>]]\",\n \tNULL\n };\n \n@@ -98,6 +98,8 @@ int cmd_read_tree(int argc, const char **argv, const char *unused_prefix)\n \t\t  PARSE_OPT_NONEG, exclude_per_directory_cb },\n \t\tOPT_SET_INT('i', NULL, &opts.index_only,\n \t\t\t    \"don't check the working tree after merging\", 1),\n+\t\tOPT_SET_INT(0, \"sparse\", &opts.apply_sparse,\n+\t\t\t    \"apply sparse checkout filter\", 1),\n \t\tOPT_END()\n \t};\n \ndiff --git a/cache.h b/cache.h\nindex 1a2a3c9..dfad54a 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -177,6 +177,9 @@ struct cache_entry {\n #define CE_HASHED    (0x100000)\n #define CE_UNHASHED  (0x200000)\n \n+/* Only remove in work directory, not index */\n+#define CE_WT_REMOVE (0x400000)\n+\n /*\n  * Extended on-disk flags\n  */\ndiff --git a/t/t1009-read-tree-sparse.sh b/t/t1009-read-tree-sparse.sh\nnew file mode 100755\nindex 0000000..f70852c\n--- /dev/null\n+++ b/t/t1009-read-tree-sparse.sh\n@@ -0,0 +1,47 @@\n+#!/bin/sh\n+\n+test_description='sparse checkout tests'\n+\n+. ./test-lib.sh\n+\n+test_expect_success 'setup' '\n+\ttest_commit one &&\n+\tmkdir two &&\n+\ttest_commit two two/two.t two.t\n+'\n+\n+test_expect_success 'read-tree without .git/info/sparse' '\n+\tgit read-tree --sparse -m -u HEAD &&\n+\ttest -f one.t &&\n+\ttest -f two/two.t\n+'\n+\n+test_expect_success 'read-tree with empty .git/info/sparse' '\n+\techo > .git/info/sparse &&\n+\tgit read-tree --sparse -m -u HEAD &&\n+\ttest -f one.t &&\n+\ttest -f two/two.t\n+'\n+\n+test_expect_success 'read-tree --sparse' '\n+\techo \"one.t\" > .git/info/sparse &&\n+\tgit read-tree --sparse -m -u HEAD &&\n+\ttest ! -f one.t &&\n+\ttest -f two/two.t\n+'\n+\n+test_expect_success 'read-tree --sparse foo where foo is \"directory\"' '\n+\techo \"two\" > .git/info/sparse &&\n+\tgit read-tree --sparse -m -u HEAD &&\n+\ttest -f one.t &&\n+\ttest -f two/two.t\n+'\n+\n+test_expect_success 'read-tree --sparse foo/' '\n+\techo \"two/\" > .git/info/sparse &&\n+\tgit read-tree --sparse -m -u HEAD &&\n+\ttest -f one.t &&\n+\ttest ! -f two/two.t\n+'\n+\n+test_done\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 02ea236..d18d333 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -32,6 +32,12 @@ static struct unpack_trees_error_msgs unpack_plumbing_errors = {\n \n \t/* bind_overlap */\n \t\"Entry '%s' overlaps with '%s'.  Cannot bind.\",\n+\n+\t/* sparse_not_uptodate_file */\n+\t\"Entry '%s' not uptodate. Cannot update sparse checkout.\",\n+\n+\t/* would_lose_orphaned */\n+\t\"Working tree file '%s' would be %s by sparse checkout update.\",\n };\n \n #define ERRORMSG(o,fld) \\\n@@ -78,7 +84,7 @@ static int check_updates(struct unpack_trees_options *o)\n \tif (o->update && o->verbose_update) {\n \t\tfor (total = cnt = 0; cnt < index->cache_nr; cnt++) {\n \t\t\tstruct cache_entry *ce = index->cache[cnt];\n-\t\t\tif (ce->ce_flags & (CE_UPDATE | CE_REMOVE))\n+\t\t\tif (ce->ce_flags & (CE_UPDATE | CE_REMOVE | CE_WT_REMOVE))\n \t\t\t\ttotal++;\n \t\t}\n \n@@ -92,6 +98,13 @@ static int check_updates(struct unpack_trees_options *o)\n \tfor (i = 0; i < index->cache_nr; i++) {\n \t\tstruct cache_entry *ce = index->cache[i];\n \n+\t\tif (ce->ce_flags & CE_WT_REMOVE) {\n+\t\t\tdisplay_progress(progress, ++cnt);\n+\t\t\tif (o->update)\n+\t\t\t\tunlink_entry(ce);\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tif (ce->ce_flags & CE_REMOVE) {\n \t\t\tdisplay_progress(progress, ++cnt);\n \t\t\tif (o->update)\n@@ -118,6 +131,74 @@ static int check_updates(struct unpack_trees_options *o)\n \treturn errs != 0;\n }\n \n+static int verify_uptodate_sparse(struct cache_entry *ce, struct unpack_trees_options *o);\n+static int verify_absent_sparse(struct cache_entry *ce, const char *action, struct unpack_trees_options *o);\n+static int apply_sparse_checkout(struct unpack_trees_options *o)\n+{\n+\tstruct index_state *index = &o->result;\n+\tstruct exclude_list el;\n+\tint i, ret = 0;\n+\n+\tmemset(&el, 0, sizeof(el));\n+\tif (add_excludes_from_file_to_list(git_path(\"info/sparse\"), \"\", 0, NULL, &el, 0) < 0)\n+\t\treturn 0;\n+\n+\tfor (i = 0; i < index->cache_nr; i++) {\n+\t\tstruct cache_entry *ce = index->cache[i];\n+\t\tconst char *basename;\n+\t\tint was_valid = ce->ce_flags & CE_VALID;\n+\n+\t\tif (ce_stage(ce))\n+\t\t\tcontinue;\n+\n+\t\tbasename = strrchr(ce->name, '/');\n+\t\tbasename = basename ? basename+1 : ce->name;\n+\t\tif (excluded_from_list(ce->name, ce_namelen(ce), basename, NULL, &el) > 0)\n+\t\t\tce->ce_flags |= CE_VALID;\n+\t\telse\n+\t\t\tce->ce_flags &= ~CE_VALID;\n+\n+\t\t/*\n+\t\t * We only care about files getting into the checkout area\n+\t\t * If merge strategies want to remove some, go ahead\n+\t\t */\n+\t\tif (ce->ce_flags & CE_REMOVE)\n+\t\t\tcontinue;\n+\n+\t\tif (!was_valid && (ce->ce_flags & CE_VALID)) {\n+\t\t\t/*\n+\t\t\t * If CE_UPDATE is set, verify_uptodate() must be called already\n+\t\t\t * also stat info may have lost after merged_entry() so calling\n+\t\t\t * verify_uptodate() again may fail\n+\t\t\t */\n+\t\t\tif (!(ce->ce_flags & CE_UPDATE) && verify_uptodate_sparse(ce, o)) {\n+\t\t\t\tret = -1;\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t\tce->ce_flags |= CE_WT_REMOVE;\n+\t\t}\n+\t\tif (was_valid && !(ce->ce_flags & CE_VALID)) {\n+\t\t\tif (verify_absent_sparse(ce, \"overwritten\", o)) {\n+\t\t\t\tret = -1;\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t\tce->ce_flags |= CE_UPDATE;\n+\t\t}\n+\n+\t\t/* merge strategies may set CE_UPDATE outside checkout area */\n+\t\tif (ce->ce_flags & CE_VALID)\n+\t\t\tce->ce_flags &= ~CE_UPDATE;\n+\n+\t}\n+\n+\tfor (i = 0;i < el.nr;i++)\n+\t\tfree(el.excludes[i]);\n+\tif (el.excludes)\n+\t\tfree(el.excludes);\n+\n+\treturn ret;\n+}\n+\n static inline int call_unpack_fn(struct cache_entry **src, struct unpack_trees_options *o)\n {\n \tint ret = o->fn(src, o);\n@@ -416,6 +497,9 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \tif (o->trivial_merges_only && o->nontrivial_merge)\n \t\treturn unpack_failed(o, \"Merge requires file-level merging\");\n \n+\tif (o->apply_sparse && apply_sparse_checkout(o))\n+\t\treturn unpack_failed(o, NULL);\n+\n \to->src_index = NULL;\n \tret = check_updates(o) ? (-2) : 0;\n \tif (o->dst_index)\n@@ -481,6 +565,12 @@ static int verify_uptodate(struct cache_entry *ce,\n \treturn verify_uptodate_1(ce, o, ERRORMSG(o, not_uptodate_file));\n }\n \n+static int verify_uptodate_sparse(struct cache_entry *ce,\n+\t\t\t\t  struct unpack_trees_options *o)\n+{\n+\treturn verify_uptodate_1(ce, o, ERRORMSG(o, sparse_not_uptodate_file));\n+}\n+\n static void invalidate_ce_path(struct cache_entry *ce, struct unpack_trees_options *o)\n {\n \tif (ce)\n@@ -674,6 +764,12 @@ static int verify_absent(struct cache_entry *ce, const char *action,\n \treturn verify_absent_1(ce, action, o, ERRORMSG(o, would_lose_untracked));\n }\n \n+static int verify_absent_sparse(struct cache_entry *ce, const char *action,\n+\t\t\t struct unpack_trees_options *o)\n+{\n+\treturn verify_absent_1(ce, action, o, ERRORMSG(o, would_lose_orphaned));\n+}\n+\n static int merged_entry(struct cache_entry *merge, struct cache_entry *old,\n \t\tstruct unpack_trees_options *o)\n {\ndiff --git a/unpack-trees.h b/unpack-trees.h\nindex d19df44..a09077b 100644\n--- a/unpack-trees.h\n+++ b/unpack-trees.h\n@@ -14,6 +14,8 @@ struct unpack_trees_error_msgs {\n \tconst char *not_uptodate_dir;\n \tconst char *would_lose_untracked;\n \tconst char *bind_overlap;\n+\tconst char *sparse_not_uptodate_file;\n+\tconst char *would_lose_orphaned;\n };\n \n struct unpack_trees_options {\n@@ -28,6 +30,7 @@ struct unpack_trees_options {\n \t\t     skip_unmerged,\n \t\t     initial_checkout,\n \t\t     diff_index_cached,\n+\t\t     apply_sparse,\n \t\t     gently;\n \tconst char *prefix;\n \tint pos;\n-- \n1.6.3.GIT\n"},{"id":"120289","messageId":"1250005446-12047-9-git-send-email-pclouds@gmail.com","threadId":"20539","inReplyTo":"1250005446-12047-8-git-send-email-pclouds@gmail.com","subject":"[RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-11T15:44:06Z","receivedAt":"2009-08-11T15:44:06Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This series is useless until now because no one would use read-tree to\ncheckout. At least with this, you can really use/test the series.\nPorcelain design was originally \"if you have .git/info/sparse,\nporcelains will use it, if you don't like that, remove\n.git/info/sparse\" while plumblings have an option to\nenable/disable this feature.\n\nAnd I still like that behavior. How about we enable sparse checkout\nby default for porcelains and make a config option to disable it?\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin-checkout.c |    4 ++++\n builtin-merge.c    |    5 ++++-\n git-pull.sh        |    6 +++++-\n 3 files changed, 13 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin-checkout.c b/builtin-checkout.c\nindex 446cac7..cec21ab 100644\n--- a/builtin-checkout.c\n+++ b/builtin-checkout.c\n@@ -30,6 +30,7 @@ struct checkout_opts {\n \tint force;\n \tint writeout_stage;\n \tint writeout_error;\n+\tint apply_sparse;\n \n \tconst char *new_branch;\n \tint new_branch_log;\n@@ -402,6 +403,7 @@ static int merge_working_tree(struct checkout_opts *opts,\n \t\ttopts.dir = xcalloc(1, sizeof(*topts.dir));\n \t\ttopts.dir->flags |= DIR_SHOW_IGNORED;\n \t\ttopts.dir->exclude_per_dir = \".gitignore\";\n+\t\ttopts.apply_sparse = opts->apply_sparse;\n \t\ttree = parse_tree_indirect(old->commit->object.sha1);\n \t\tinit_tree_desc(&trees[0], tree->buffer, tree->size);\n \t\ttree = parse_tree_indirect(new->commit->object.sha1);\n@@ -594,6 +596,8 @@ int cmd_checkout(int argc, const char **argv, const char *prefix)\n \t\tOPT_BOOLEAN('m', \"merge\", &opts.merge, \"merge\"),\n \t\tOPT_STRING(0, \"conflict\", &conflict_style, \"style\",\n \t\t\t   \"conflict style (merge or diff3)\"),\n+\t\tOPT_SET_INT(0, \"sparse\", &opts.apply_sparse,\n+\t\t\t    \"apply sparse checkout filter\", 1),\n \t\tOPT_END(),\n \t};\n \tint has_dash_dash;\ndiff --git a/builtin-merge.c b/builtin-merge.c\nindex 0b12fb3..c14b91d 100644\n--- a/builtin-merge.c\n+++ b/builtin-merge.c\n@@ -43,7 +43,7 @@ static const char * const builtin_merge_usage[] = {\n \n static int show_diffstat = 1, option_log, squash;\n static int option_commit = 1, allow_fast_forward = 1;\n-static int allow_trivial = 1, have_message;\n+static int allow_trivial = 1, have_message, apply_sparse;\n static struct strbuf merge_msg;\n static struct commit_list *remoteheads;\n static unsigned char head[20], stash[20];\n@@ -172,6 +172,7 @@ static struct option builtin_merge_options[] = {\n \tOPT_CALLBACK('m', \"message\", &merge_msg, \"message\",\n \t\t\"message to be used for the merge commit (if any)\",\n \t\toption_parse_message),\n+\tOPT_SET_INT(0, \"sparse\", &apply_sparse, \"apply sparse checkout filter\", 1),\n \tOPT__VERBOSITY(&verbosity),\n \tOPT_END()\n };\n@@ -494,6 +495,7 @@ static int read_tree_trivial(unsigned char *common, unsigned char *head,\n \topts.verbose_update = 1;\n \topts.trivial_merges_only = 1;\n \topts.merge = 1;\n+\topts.apply_sparse = apply_sparse;\n \ttrees[nr_trees] = parse_tree_indirect(common);\n \tif (!trees[nr_trees++])\n \t\treturn -1;\n@@ -646,6 +648,7 @@ static int checkout_fast_forward(unsigned char *head, unsigned char *remote)\n \topts.verbose_update = 1;\n \topts.merge = 1;\n \topts.fn = twoway_merge;\n+\topts.apply_sparse = apply_sparse;\n \n \ttrees[nr_trees] = parse_tree_indirect(head);\n \tif (!trees[nr_trees++])\ndiff --git a/git-pull.sh b/git-pull.sh\nindex 0f24182..ba583bf 100755\n--- a/git-pull.sh\n+++ b/git-pull.sh\n@@ -20,6 +20,7 @@ strategy_args= diffstat= no_commit= squash= no_ff= log_arg= verbosity=\n curr_branch=$(git symbolic-ref -q HEAD)\n curr_branch_short=$(echo \"$curr_branch\" | sed \"s|refs/heads/||\")\n rebase=$(git config --bool branch.$curr_branch_short.rebase)\n+sparse=\n while :\n do\n \tcase \"$1\" in\n@@ -65,6 +66,9 @@ do\n \t--no-r|--no-re|--no-reb|--no-reba|--no-rebas|--no-rebase)\n \t\trebase=false\n \t\t;;\n+\t--sparse)\n+\t\tsparse=--sparse\n+\t\t;;\n \t-h|--h|--he|--hel|--help)\n \t\tusage\n \t\t;;\n@@ -201,5 +205,5 @@ merge_name=$(git fmt-merge-msg $log_arg <\"$GIT_DIR/FETCH_HEAD\") || exit\n test true = \"$rebase\" &&\n \texec git-rebase $diffstat $strategy_args --onto $merge_head \\\n \t${oldremoteref:-$merge_head}\n-exec git-merge $diffstat $no_commit $squash $no_ff $log_arg $strategy_args \\\n+exec git-merge $sparse $diffstat $no_commit $squash $no_ff $log_arg $strategy_args \\\n \t\"$merge_name\" HEAD $merge_head $verbosity\n-- \n1.6.3.GIT\n"},{"id":"120307","messageId":"2729632a0908111418m57e03d8as9c122cbb52efc21a@mail.gmail.com","threadId":"20539","inReplyTo":"1250005446-12047-8-git-send-email-pclouds@gmail.com","subject":"Re: [RFC PATCH v3 7/8] Support sparse checkout in unpack_trees() and read-tree","fromName":"","fromEmail":"skillzero@gmail.com","sentAt":"2009-08-11T21:18:55Z","receivedAt":"2009-08-11T21:18:55Z","isPatch":true,"sender":{"key":"skillzero@gmail.com","avatar":null},"body":"2009/8/11 Nguyễn Thái Ngọc Duy <pclouds@gmail.com>:\n> [1] .git/info/sparse has the same syntax as .git/info/exclude. Files\n> that match the patterns will be set as CE_VALID.\n\nDoes this mean it will only support excluding paths you don't want\nrather than letting you only include paths you do want?\n\nI'm currently using your other patch series that lets you include or\nexclude paths (via config variable) and I find that I mostly use the\ninclude side of it with only a few excluded paths. This is because I\ntypically want to include only a small subset of the repository so\nusing excludes would require a pretty large list and any time somebody\nadds new files, I'd have to update the exclude list.\n\nI appreciate the flexibility of the script to control what is included\nor excluded, but like some other comments here, I like the simplicity\nof having built-in support for including/excluding paths without\nhaving to write a script to do it. Some of my projects run on Windows\nso scripting is more difficult there.\n"},{"id":"120312","messageId":"m3ab26owub.fsf@localhost.localdomain","threadId":"20539","inReplyTo":"2729632a0908111418m57e03d8as9c122cbb52efc21a@mail.gmail.com","subject":"Re: [RFC PATCH v3 7/8] Support sparse checkout in unpack_trees() and read-tree","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-08-11T21:38:04Z","receivedAt":"2009-08-11T21:38:04Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"skillzero@gmail.com writes:\n> 2009/8/11 Nguyễn Thái Ngọc Duy <pclouds@gmail.com>:\n\n> > [1] .git/info/sparse has the same syntax as .git/info/exclude. Files\n> > that match the patterns will be set as CE_VALID.\n> \n> Does this mean it will only support excluding paths you don't want\n> rather than letting you only include paths you do want?\n\nErrr... what I read is that paths set by .git/info/sparse would be\nexcluded from checkout (marked as assume-unchanged / CE_VALID).\n\nBut if it is the same mechanism as gitignore, then you can use ! \nprefix to set files (patterns) to include, e.g.\n\n  !Documentation/\n  *\n\n(I think rules are processed top-down, first matching wins).\n \n> I'm currently using your other patch series that lets you include or\n> exclude paths (via config variable) and I find that I mostly use the\n> include side of it with only a few excluded paths. This is because I\n> typically want to include only a small subset of the repository so\n> using excludes would require a pretty large list and any time somebody\n> adds new files, I'd have to update the exclude list.\n\nNot true, see above.\n\n-- \nJakub Narebski\n\nGit User's Survey 2009\nhttp://tinyurl.com/GitSurvey2009\n"},{"id":"120315","messageId":"2729632a0908111503i7f035c1aw4e84151eab821006@mail.gmail.com","threadId":"20539","inReplyTo":"m3ab26owub.fsf@localhost.localdomain","subject":"Re: [RFC PATCH v3 7/8] Support sparse checkout in unpack_trees() and read-tree","fromName":"","fromEmail":"skillzero@gmail.com","sentAt":"2009-08-11T22:03:10Z","receivedAt":"2009-08-11T22:03:10Z","isPatch":true,"sender":{"key":"skillzero@gmail.com","avatar":null},"body":"On Tue, Aug 11, 2009 at 2:38 PM, Jakub Narebski<jnareb@gmail.com> wrote:\n> skillzero@gmail.com writes:\n>> 2009/8/11 Nguyễn Thái Ngọc Duy <pclouds@gmail.com>:\n>\n>> > [1] .git/info/sparse has the same syntax as .git/info/exclude. Files\n>> > that match the patterns will be set as CE_VALID.\n>>\n>> Does this mean it will only support excluding paths you don't want\n>> rather than letting you only include paths you do want?\n>\n> Errr... what I read is that paths set by .git/info/sparse would be\n> excluded from checkout (marked as assume-unchanged / CE_VALID).\n>\n> But if it is the same mechanism as gitignore, then you can use !\n> prefix to set files (patterns) to include, e.g.\n>\n>  !Documentation/\n>  *\n>\n> (I think rules are processed top-down, first matching wins).\n\nI wasn't sure because the .gitignore negation stuff mentions negating\na previously ignored pattern. But for sparse patterns, there likely\nwouldn't be a previous pattern. Include patterns are a little\ndifferent in that if there are no include patterns (but maybe some\nexclude patterns), I think the expectation is that everything will be\nincluded (minus excludes), but if you have some include patterns then\nonly those paths will be included (minus any excludes).\n\nIt's great if it already supports includes as well as excludes\n(although it's a little confusing to say !Documentation to mean\n\"include it\"), but I wasn't sure from the comment so I was just\nasking.\n"},{"id":"120348","messageId":"fcaeb9bf0908111830n50bd4733h5033c6f13a45999@mail.gmail.com","threadId":"20539","inReplyTo":"2729632a0908111503i7f035c1aw4e84151eab821006@mail.gmail.com","subject":"Re: [RFC PATCH v3 7/8] Support sparse checkout in unpack_trees() and read-tree","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-12T01:30:45Z","receivedAt":"2009-08-12T01:30:45Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Aug 12, 2009 at 5:03 AM, <skillzero@gmail.com> wrote:\n> On Tue, Aug 11, 2009 at 2:38 PM, Jakub Narebski<jnareb@gmail.com> wrote:\n>> skillzero@gmail.com writes:\n>>> 2009/8/11 Nguyễn Thái Ngọc Duy <pclouds@gmail.com>:\n>>\n>>> > [1] .git/info/sparse has the same syntax as .git/info/exclude. Files\n>>> > that match the patterns will be set as CE_VALID.\n>>>\n>>> Does this mean it will only support excluding paths you don't want\n>>> rather than letting you only include paths you do want?\n>>\n>> Errr... what I read is that paths set by .git/info/sparse would be\n>> excluded from checkout (marked as assume-unchanged / CE_VALID).\n>>\n>> But if it is the same mechanism as gitignore, then you can use !\n>> prefix to set files (patterns) to include, e.g.\n>>\n>>  !Documentation/\n>>  *\n>>\n>> (I think rules are processed top-down, first matching wins).\n>\n> I wasn't sure because the .gitignore negation stuff mentions negating\n> a previously ignored pattern. But for sparse patterns, there likely\n> wouldn't be a previous pattern.\n\nNo problem. We put pattern '*' at top (match everything). Previous\npattern issue solved.\n\n> Include patterns are a little\n> different in that if there are no include patterns (but maybe some\n> exclude patterns), I think the expectation is that everything will be\n> included (minus excludes), but if you have some include patterns then\n> only those paths will be included (minus any excludes).\n\nLet's say you want to include foo/ and bar/ only, this should work:\n\n*\n!foo/\n!bar/\n\nThe evaluating order is from bottom up. When it first matches 'bar/',\nbecause it a negate pattern, it returns \"no don't match\" and stops.\nWhen it matches neither foo/ nor bar/ then it will be caught by '*'\nand return \"yes it matches\" - that means \"ignored\" from checkout area.\nIn the end only foo/* and bar/* survive.\n\nI think it's as easy as writing exclude patterns once you figure out '*'.\n-- \nDuy\n"},{"id":"120354","messageId":"7vocqlbv7a.fsf@alter.siamese.dyndns.org","threadId":"20539","inReplyTo":"1250005446-12047-4-git-send-email-pclouds@gmail.com","subject":"Re: [RFC PATCH v3 3/8] Read .gitignore from index if it is assume-unchanged","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-08-12T02:51:53Z","receivedAt":"2009-08-12T02:51:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> diff --git a/Documentation/technical/api-directory-listing.txt b/Documentation/technical/api-directory-listing.txt\n> index 5bbd18f..7d0e282 100644\n> --- a/Documentation/technical/api-directory-listing.txt\n> +++ b/Documentation/technical/api-directory-listing.txt\n> @@ -58,6 +58,9 @@ The result of the enumeration is left in these fields::\n>  Calling sequence\n>  ----------------\n>  \n> +* Ensure the_index is populated as it may have CE_VALID entries that\n> +  affect directory listing.\n> +\n\nWhen you want to enumerate all paths in the work tree, instead of not just\nthe untracked ones, it used to be possible to first run read_directory()\nbefore calling read_cache().  You are now forbidding this.\n\nI do not think it is hard to resurrect the feature if it is necessary (add\nan option to dir_struct and teach dir_add_name() not to ignore paths the\nindex knows about), and I do not think none of the existing code relies on\nit anymore (I think \"git add\" used to), but there may be some codepath I\nforgot about, which is a concern.\n\n> diff --git a/builtin-clean.c b/builtin-clean.c\n> index 2d8c735..d917472 100644\n> --- a/builtin-clean.c\n> +++ b/builtin-clean.c\n> @@ -71,8 +71,11 @@ int cmd_clean(int argc, const char **argv, const char *prefix)\n>  \n>  \tdir.flags |= DIR_SHOW_OTHER_DIRECTORIES;\n>  \n> -\tif (!ignored)\n> +\tif (!ignored) {\n> +\t\tif (read_cache() < 0)\n> +\t\t\tdie(\"index file corrupt\");\n>  \t\tsetup_standard_excludes(&dir);\n> +\t}\n>  \n>  \tpathspec = get_pathspec(prefix, argv);\n>  \tread_cache();\n\nWouldn't it be much cleaner to move the existing read_cache() up, like you\ndid for ls-files, instead of conditionally reading the index at a random\nplace in the program sequence depending on the combinations of options?\n"},{"id":"120362","messageId":"2729632a0908112159y13a088a1w5580cf042a40bec8@mail.gmail.com","threadId":"20539","inReplyTo":"fcaeb9bf0908111830n50bd4733h5033c6f13a45999@mail.gmail.com","subject":"Re: [RFC PATCH v3 7/8] Support sparse checkout in unpack_trees() and read-tree","fromName":"","fromEmail":"skillzero@gmail.com","sentAt":"2009-08-12T04:59:17Z","receivedAt":"2009-08-12T04:59:17Z","isPatch":true,"sender":{"key":"skillzero@gmail.com","avatar":null},"body":"On Tue, Aug 11, 2009 at 6:30 PM, Nguyen Thai Ngoc Duy<pclouds@gmail.com> wrote:\n\n> I think it's as easy as writing exclude patterns once you figure out '*'.\n\nThat solves it for me. Thanks.\n"},{"id":"120378","messageId":"7v3a7xa6e5.fsf@alter.siamese.dyndns.org","threadId":"20539","inReplyTo":"1250005446-12047-9-git-send-email-pclouds@gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-08-12T06:33:06Z","receivedAt":"2009-08-12T06:33:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> @@ -594,6 +596,8 @@ int cmd_checkout(int argc, const char **argv, const char *prefix)\n>  \t\tOPT_BOOLEAN('m', \"merge\", &opts.merge, \"merge\"),\n>  \t\tOPT_STRING(0, \"conflict\", &conflict_style, \"style\",\n>  \t\t\t   \"conflict style (merge or diff3)\"),\n> +\t\tOPT_SET_INT(0, \"sparse\", &opts.apply_sparse,\n> +\t\t\t    \"apply sparse checkout filter\", 1),\n\nShouldn't this be BOOLEAN not INT, i.e. \"--[no-]sparse\"?  That way, you\ncould enable it by simply the presense of $GIT_DIR/info/sparse.\n\nIt could also require core.sparseworktree configuration set to true if we\nare really paranoid, but without the actual sparse specification file\nflipping that configuration to true would not be useful anyway, so in\npractice, giving --sparse-work-tree option to these Porcelain commands\nwould be no-op, but --no-sparse-work-tree option would be useful to\nignore $GIT_DIR/info/sparse and populate the work tree fully.\n\nOr am I missing something?\n"},{"id":"120382","messageId":"4A826FD4.5080201@viscovery.net","threadId":"20539","inReplyTo":"1250005446-12047-9-git-send-email-pclouds@gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Johannes Sixt","fromEmail":"j.sixt@viscovery.net","sentAt":"2009-08-12T07:31:32Z","receivedAt":"2009-08-12T07:31:32Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Nguyễn Thái Ngọc Duy schrieb:\n> This series is useless until now because no one would use read-tree to\n> checkout. At least with this, you can really use/test the series.\n> Porcelain design was originally \"if you have .git/info/sparse,\n> porcelains will use it, if you don't like that, remove\n> .git/info/sparse\" while plumblings have an option to\n> enable/disable this feature.\n> \n> And I still like that behavior. How about we enable sparse checkout\n> by default for porcelains and make a config option to disable it?\n\nI would enable sparse checkout by default even for plumbing. Whether the\ncheckout area is sparse should always be governed by .git/info/sparse.\nThis way, existing scripts and aliases should automatically work in sparse\nworktrees.\n\nBTW, the name .git/info/sparse is perhaps a bit too technical in the sense\nthat only git developers know that this feature runs under the name\n\"sparse checkout\". Perhaps it should be named\n\n   .git/info/indexonly\n   .git/info/nocheckout\n\nor so.\n\n-- Hannes\n"},{"id":"120396","messageId":"fcaeb9bf0908120253p192125a4mbb6a0838fc90f10e@mail.gmail.com","threadId":"20539","inReplyTo":"4A826FD4.5080201@viscovery.net","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-12T09:53:26Z","receivedAt":"2009-08-12T09:53:26Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"2009/8/12 Johannes Sixt <j.sixt@viscovery.net>:\n> BTW, the name .git/info/sparse is perhaps a bit too technical in the sense\n> that only git developers know that this feature runs under the name\n> \"sparse checkout\". Perhaps it should be named\n>\n>   .git/info/indexonly\n>   .git/info/nocheckout\n>\n> or so.\n\nI did not like the name \"sparse\" either. Another option is\n.git/info/assume-unchanged.\n-- \nDuy\n"},{"id":"120398","messageId":"fcaeb9bf0908120301q17812b5cw5e19def7887d31db@mail.gmail.com","threadId":"20539","inReplyTo":"7v3a7xa6e5.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-12T10:01:33Z","receivedAt":"2009-08-12T10:01:33Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"2009/8/12 Junio C Hamano <gitster@pobox.com>:\n> Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n>\n>> @@ -594,6 +596,8 @@ int cmd_checkout(int argc, const char **argv, const char *prefix)\n>>               OPT_BOOLEAN('m', \"merge\", &opts.merge, \"merge\"),\n>>               OPT_STRING(0, \"conflict\", &conflict_style, \"style\",\n>>                          \"conflict style (merge or diff3)\"),\n>> +             OPT_SET_INT(0, \"sparse\", &opts.apply_sparse,\n>> +                         \"apply sparse checkout filter\", 1),\n>\n> Shouldn't this be BOOLEAN not INT, i.e. \"--[no-]sparse\"?  That way, you\n> could enable it by simply the presense of $GIT_DIR/info/sparse.\n\nThis patch was written carelessly. I wanted to have something to test.\nIf you agree on option name \"--sparse\" then yes BOOLEAN is better.\n\n> It could also require core.sparseworktree configuration set to true if we\n> are really paranoid, but without the actual sparse specification file\n> flipping that configuration to true would not be useful anyway, so in\n> practice, giving --sparse-work-tree option to these Porcelain commands\n> would be no-op, but --no-sparse-work-tree option would be useful to\n> ignore $GIT_DIR/info/sparse and populate the work tree fully.\n>\n> Or am I missing something?\n\nSounds good (and --sparse-work-tree is apparently better than\n--sparse). So let's enable it by default, add --no-sparse-work-tree to\ndisable it and wait until some one complains, then we'll add\ncore.sparseworktree. I think core.sparseworktree can also be used to\nspecify what spec file to be used instead of the default\n.git/info/sparse, if users like to switch among some well-defined spec\nfiles.\n-- \nDuy\n"},{"id":"120421","messageId":"87ljlpvy4r.fsf@hariville.hurrynot.org","threadId":"20539","inReplyTo":"fcaeb9bf0908120253p192125a4mbb6a0838fc90f10e@mail.gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Raja R Harinath","fromEmail":"harinath@hurrynot.org","sentAt":"2009-08-12T15:40:36Z","receivedAt":"2009-08-12T15:40:36Z","isPatch":true,"sender":{"key":"harinath@hurrynot.org","avatar":"https://avatars.githubusercontent.com/u/4610?v=4"},"body":"Hi,\n\nNguyen Thai Ngoc Duy <pclouds@gmail.com> writes:\n\n> 2009/8/12 Johannes Sixt <j.sixt@viscovery.net>:\n>> BTW, the name .git/info/sparse is perhaps a bit too technical in the sense\n>> that only git developers know that this feature runs under the name\n>> \"sparse checkout\". Perhaps it should be named\n>>\n>>   .git/info/indexonly\n>>   .git/info/nocheckout\n>>\n>> or so.\n>\n> I did not like the name \"sparse\" either. Another option is\n> .git/info/assume-unchanged.\n\nOr .git/info/doppelgangers, or even .git/info/doppelgängers :-)\n\n- Hari\n"},{"id":"120493","messageId":"fcaeb9bf0908122337j7d783f59l8ce7a125bf0dbc7e@mail.gmail.com","threadId":"20539","inReplyTo":"7vocqlbv7a.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC PATCH v3 3/8] Read .gitignore from index if it is assume-unchanged","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-13T06:37:21Z","receivedAt":"2009-08-13T06:37:21Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"2009/8/12 Junio C Hamano <gitster@pobox.com>:\n> Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n>\n>> diff --git a/Documentation/technical/api-directory-listing.txt b/Documentation/technical/api-directory-listing.txt\n>> index 5bbd18f..7d0e282 100644\n>> --- a/Documentation/technical/api-directory-listing.txt\n>> +++ b/Documentation/technical/api-directory-listing.txt\n>> @@ -58,6 +58,9 @@ The result of the enumeration is left in these fields::\n>>  Calling sequence\n>>  ----------------\n>>\n>> +* Ensure the_index is populated as it may have CE_VALID entries that\n>> +  affect directory listing.\n>> +\n>\n> When you want to enumerate all paths in the work tree, instead of not just\n> the untracked ones, it used to be possible to first run read_directory()\n> before calling read_cache().  You are now forbidding this.\n\nEither I phrased it badly, or I don't follow you. If you don't call\nread_cache() before read_directory(), the_index should be empty and\nread_assume_unchanged_from_index() will be no-op. So read_directory()\nbehavior does not change in this case.\n\n> I do not think it is hard to resurrect the feature if it is necessary (add\n> an option to dir_struct and teach dir_add_name() not to ignore paths the\n> index knows about), and I do not think none of the existing code relies on\n> it anymore (I think \"git add\" used to), but there may be some codepath I\n> forgot about, which is a concern.\n\nHmm.. \"git add\" loaded index early since the first version of\nbuiltin-add.c. I have checked all code path that can lead to\nread_directory_recursively(). In all cases, index is loaded before\nread_dir..() is called.\n\n>> diff --git a/builtin-clean.c b/builtin-clean.c\n>> index 2d8c735..d917472 100644\n>> --- a/builtin-clean.c\n>> +++ b/builtin-clean.c\n>> @@ -71,8 +71,11 @@ int cmd_clean(int argc, const char **argv, const char *prefix)\n>>\n>>       dir.flags |= DIR_SHOW_OTHER_DIRECTORIES;\n>>\n>> -     if (!ignored)\n>> +     if (!ignored) {\n>> +             if (read_cache() < 0)\n>> +                     die(\"index file corrupt\");\n>>               setup_standard_excludes(&dir);\n>> +     }\n>>\n>>       pathspec = get_pathspec(prefix, argv);\n>>       read_cache();\n>\n> Wouldn't it be much cleaner to move the existing read_cache() up, like you\n> did for ls-files, instead of conditionally reading the index at a random\n> place in the program sequence depending on the combinations of options?\n\nAgreed. read_cache() is called right below anyway.\n-- \nDuy\n"},{"id":"120495","messageId":"fcaeb9bf0908130020meaed129j5d6a4f04a6878bd0@mail.gmail.com","threadId":"20539","inReplyTo":"7v3a7xa6e5.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-13T07:20:56Z","receivedAt":"2009-08-13T07:20:56Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"2009/8/12 Junio C Hamano <gitster@pobox.com>:\n> It could also require core.sparseworktree configuration set to true if we\n> are really paranoid, but without the actual sparse specification file\n> flipping that configuration to true would not be useful anyway, so in\n> practice, giving --sparse-work-tree option to these Porcelain commands\n> would be no-op, but --no-sparse-work-tree option would be useful to\n> ignore $GIT_DIR/info/sparse and populate the work tree fully.\n\nOnly part \"ignore $GIT_DIR/info/sparse\" is correct.\n\"--no-sparse-work-tree\" would not clear CE_VALID from all entries in\nindex (which is good, if you are using CE_VALID for another purpose).\n\nTo quit sparse checkout, you must create an empty\n$GIT_DIR/info/sparse, then do \"git checkout\" or \"git read-tree -m -u\nHEAD\" so that the tree is full populated, then you can remove\n$GIT_DIR/info/sparse. Quite unintuitive..\n-- \nDuy\n"},{"id":"120498","messageId":"4A83C2CB.7030300@viscovery.net","threadId":"20539","inReplyTo":"87ljlpvy4r.fsf@hariville.hurrynot.org","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Johannes Sixt","fromEmail":"j.sixt@viscovery.net","sentAt":"2009-08-13T07:37:47Z","receivedAt":"2009-08-13T07:37:47Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Raja R Harinath schrieb:\n> Hi,\n> \n> Nguyen Thai Ngoc Duy <pclouds@gmail.com> writes:\n> \n>> 2009/8/12 Johannes Sixt <j.sixt@viscovery.net>:\n>>> BTW, the name .git/info/sparse is perhaps a bit too technical in the sense\n>>> that only git developers know that this feature runs under the name\n>>> \"sparse checkout\". Perhaps it should be named\n>>>\n>>>   .git/info/indexonly\n>>>   .git/info/nocheckout\n>>>\n>>> or so.\n>> I did not like the name \"sparse\" either. Another option is\n>> .git/info/assume-unchanged.\n> \n> Or .git/info/doppelgangers, or even .git/info/doppelgängers :-)\n\nHeh!\n\n   .git/info/phantoms\n   git checkout --no-phantoms\n   git read-tree --phantoms\n\n-- Hannes\n"},{"id":"120499","messageId":"m3skfwnihn.fsf@localhost.localdomain","threadId":"20539","inReplyTo":"fcaeb9bf0908130020meaed129j5d6a4f04a6878bd0@mail.gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-08-13T09:58:03Z","receivedAt":"2009-08-13T09:58:03Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Nguyen Thai Ngoc Duy <pclouds@gmail.com> writes:\n> 2009/8/12 Junio C Hamano <gitster@pobox.com>:\n\n> > It could also require core.sparseworktree configuration set to true if we\n> > are really paranoid, but without the actual sparse specification file\n> > flipping that configuration to true would not be useful anyway, so in\n> > practice, giving --sparse-work-tree option to these Porcelain commands\n> > would be no-op, but --no-sparse-work-tree option would be useful to\n> > ignore $GIT_DIR/info/sparse and populate the work tree fully.\n> \n> Only part \"ignore $GIT_DIR/info/sparse\" is correct.\n> \"--no-sparse-work-tree\" would not clear CE_VALID from all entries in\n> index (which is good, if you are using CE_VALID for another purpose).\n> \n> To quit sparse checkout, you must create an empty\n> $GIT_DIR/info/sparse, then do \"git checkout\" or \"git read-tree -m -u\n> HEAD\" so that the tree is full populated, then you can remove\n> $GIT_DIR/info/sparse. Quite unintuitive..\n\nHmmm... this looks like either argument for introducing --full option\nto git-checkout (ignore CE_VALID bit, checkout everything, and clean\nCE_VALID (?))...\n\n...or for going with _separate_ bit for partial checkout, like in the\nvery first version of this series, which otherwise functions like\nCE_VALID, or is just used to mark that CE_VALID was set using sparse.\n\nFood for thought.\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"120524","messageId":"fcaeb9bf0908130538x396b1208s43d312107e3e198c@mail.gmail.com","threadId":"20539","inReplyTo":"m3skfwnihn.fsf@localhost.localdomain","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-13T12:38:09Z","receivedAt":"2009-08-13T12:38:09Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On 8/13/09, Jakub Narebski <jnareb@gmail.com> wrote:\n> Nguyen Thai Ngoc Duy <pclouds@gmail.com> writes:\n>  > 2009/8/12 Junio C Hamano <gitster@pobox.com>:\n>\n>  > > It could also require core.sparseworktree configuration set to true if we\n>  > > are really paranoid, but without the actual sparse specification file\n>  > > flipping that configuration to true would not be useful anyway, so in\n>  > > practice, giving --sparse-work-tree option to these Porcelain commands\n>  > > would be no-op, but --no-sparse-work-tree option would be useful to\n>  > > ignore $GIT_DIR/info/sparse and populate the work tree fully.\n>  >\n>  > Only part \"ignore $GIT_DIR/info/sparse\" is correct.\n>  > \"--no-sparse-work-tree\" would not clear CE_VALID from all entries in\n>  > index (which is good, if you are using CE_VALID for another purpose).\n>  >\n>  > To quit sparse checkout, you must create an empty\n>  > $GIT_DIR/info/sparse, then do \"git checkout\" or \"git read-tree -m -u\n>  > HEAD\" so that the tree is full populated, then you can remove\n>  > $GIT_DIR/info/sparse. Quite unintuitive..\n>\n>\n> Hmmm... this looks like either argument for introducing --full option\n>  to git-checkout (ignore CE_VALID bit, checkout everything, and clean\n>  CE_VALID (?))...\n>\n>  ...or for going with _separate_ bit for partial checkout, like in the\n>  very first version of this series, which otherwise functions like\n>  CE_VALID, or is just used to mark that CE_VALID was set using sparse.\n\nIn my opinion, making an empty .git/info/sparse to fully populate\nworktree is not too bad. I wanted to have plumbing-level support in\ngit so that you could try sparse checkout on your projects (possibly\nwith a few additional scripts to make your life easier). Then good\nPorcelain UI may emerge later (or in worst case, people would roll\ntheir own sparse checkout).\n-- \nDuy\n"},{"id":"120653","messageId":"200908142223.07994.jnareb@gmail.com","threadId":"20539","inReplyTo":"fcaeb9bf0908130538x396b1208s43d312107e3e198c@mail.gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-08-14T20:23:06Z","receivedAt":"2009-08-14T20:23:06Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Dnia czwartek 13. sierpnia 2009 14:38, Nguyen Thai Ngoc Duy napisał:\n> On 8/13/09, Jakub Narebski <jnareb@gmail.com> wrote:\n>> Nguyen Thai Ngoc Duy <pclouds@gmail.com> writes:\n>>> 2009/8/12 Junio C Hamano <gitster@pobox.com>:\n>>\n>>>> It could also require core.sparseworktree configuration set to true if we\n>>>> are really paranoid, but without the actual sparse specification file\n>>>> flipping that configuration to true would not be useful anyway, so in\n>>>> practice, giving --sparse-work-tree option to these Porcelain commands\n>>>> would be no-op, but --no-sparse-work-tree option would be useful to\n>>>> ignore $GIT_DIR/info/sparse and populate the work tree fully.\n>>>\n>>> Only part \"ignore $GIT_DIR/info/sparse\" is correct.\n>>> \"--no-sparse-work-tree\" would not clear CE_VALID from all entries in\n>>> index (which is good, if you are using CE_VALID for another purpose).\n>>>\n>>> To quit sparse checkout, you must create an empty\n>>> $GIT_DIR/info/sparse, then do \"git checkout\" or \"git read-tree -m -u\n>>> HEAD\" so that the tree is full populated, then you can remove\n>>> $GIT_DIR/info/sparse. Quite unintuitive..\n>>\n>>\n>> Hmmm... this looks like either argument for introducing --full option\n>>  to git-checkout (ignore CE_VALID bit, checkout everything, and clean\n>>  CE_VALID (?))...\n>>\n>>  ...or for going with _separate_ bit for partial checkout, like in the\n>>  very first version of this series, which otherwise functions like\n>>  CE_VALID, or is just used to mark that CE_VALID was set using sparse.\n> \n> In my opinion, making an empty .git/info/sparse to fully populate\n> worktree is not too bad. I wanted to have plumbing-level support in\n> git so that you could try sparse checkout on your projects (possibly\n> with a few additional scripts to make your life easier). Then good\n> Porcelain UI may emerge later (or in worst case, people would roll\n> their own sparse checkout).\n\nDeciding whether sparse checkout should use CE_VALID only, or should it\n(as it was in the very first version of series) use additional flag, \neither CE_NO_CHECKOUT, or CE_VALID_IS_USED_HERE_FOR_SPARSE_CHECKOUT ;-)\nis a design decision about *plumbing-level* support.\n\nNote that shallow clone, while using the same mechanism as grafts file,\nnevertheless use separate file; so perhaps sparse checkout while using\nthe same mechanism as --assume-unchanged should use additional flag.\n\n\nBTW. you might want to use GIT_SPARSE_FILE, similar to GIT_INDEX_FILE;\nsee the fact that plumbing doesn't have .gitignore not .git/info/excludes\nhardcoded... well, except for --standard-excludes.  This way full\ncheckout would be as simple as using\n\n  $ GIT_SPARSE_FILE= git checkout -- .\n\n-- \nJakub Narebski\nPoland\n"},{"id":"120684","messageId":"7veird4yyi.fsf@alter.siamese.dyndns.org","threadId":"20539","inReplyTo":"200908142223.07994.jnareb@gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-08-15T02:01:41Z","receivedAt":"2009-08-15T02:01:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n>>> Hmmm... this looks like either argument for introducing --full option\n>>>  to git-checkout (ignore CE_VALID bit, checkout everything, and clean\n>>>  CE_VALID (?))...\n>>>\n>>>  ...or for going with _separate_ bit for partial checkout, like in the\n>>>  very first version of this series, which otherwise functions like\n>>>  CE_VALID, or is just used to mark that CE_VALID was set using sparse.\n\nHow would a separate bit help?  Just like you need to clear CE_VALID bit\nto revert the index into a normal (or \"non sparse\") state somehow, you\nwould need to have a way to clear that separate bit anyway.\n\nA separate bit would help only if you want to handle assume-unchanged and\nsparse checkout independently. But my impression was that the recent lstat\nreduction effort addressed the issue assume-unchanged were invented to\nwork around in the first place.\n\nCf. http://thread.gmane.org/gmane.comp.version-control.git/123218/focus=123252\n\nThere is no reason to use assume-unchanged to tell git not to lstat to see\nif a path is up-to-date by promising that you are not going to touch it\nafter you checked it out.\n\nSo I do not understand why you would want a separate bit, nor why you\nthink a separate bit would help when changing the index state from sparse\nto non-sparse (or vice versa).\n"},{"id":"120751","messageId":"200908160137.30384.jnareb@gmail.com","threadId":"20539","inReplyTo":"7veird4yyi.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-08-15T23:37:24Z","receivedAt":"2009-08-15T23:37:24Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sat, 15 Aug 2009, Junio C Hamano wrote:\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n>>>> Hmmm... this looks like either argument for introducing --full option\n>>>>  to git-checkout (ignore CE_VALID bit, checkout everything, and clean\n>>>>  CE_VALID (?))...\n>>>>\n>>>>  ...or for going with _separate_ bit for partial checkout, like in the\n>>>>  very first version of this series, which otherwise functions like\n>>>>  CE_VALID, or is just used to mark that CE_VALID was set using sparse.\n> \n> How would a separate bit help?  Just like you need to clear CE_VALID bit\n> to revert the index into a normal (or \"non sparse\") state somehow, you\n> would need to have a way to clear that separate bit anyway.\n> \n> A separate bit would help only if you want to handle assume-unchanged and\n> sparse checkout independently. But my impression was that the recent lstat\n> reduction effort addressed the issue assume-unchanged were invented to\n> work around in the first place.\n\nWell, if we assume that we don't need (don't want) to handle\nassume-unchanged and sparse checkout independently, then of course the\nidea of having separate or additional bit for sparse doesn't make sense.\n \n-- \nJakub Narebski\nPoland\n"},{"id":"120758","messageId":"alpine.DEB.1.00.0908161002460.8306@pacific.mpi-cbg.de","threadId":"20539","inReplyTo":"200908160137.30384.jnareb@gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-08-16T08:14:49Z","receivedAt":"2009-08-16T08:14:49Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 16 Aug 2009, Jakub Narebski wrote:\n\n> On Sat, 15 Aug 2009, Junio C Hamano wrote:\n> > Jakub Narebski <jnareb@gmail.com> writes:\n> > \n> >>>> Hmmm... this looks like either argument for introducing --full \n> >>>> option to git-checkout (ignore CE_VALID bit, checkout everything, \n> >>>> and clean CE_VALID (?))...\n> >>>>\n> >>>>  ...or for going with _separate_ bit for partial checkout, like in \n> >>>>  the very first version of this series, which otherwise functions \n> >>>>  like CE_VALID, or is just used to mark that CE_VALID was set using \n> >>>>  sparse.\n> > \n> > How would a separate bit help?  Just like you need to clear CE_VALID \n> > bit to revert the index into a normal (or \"non sparse\") state somehow, \n> > you would need to have a way to clear that separate bit anyway.\n> > \n> > A separate bit would help only if you want to handle assume-unchanged \n> > and sparse checkout independently. But my impression was that the \n> > recent lstat reduction effort addressed the issue assume-unchanged \n> > were invented to work around in the first place.\n> \n> Well, if we assume that we don't need (don't want) to handle \n> assume-unchanged and sparse checkout independently, then of course the \n> idea of having separate or additional bit for sparse doesn't make sense.\n\nFor the shallow/graft issue, we had a similar discussion.  Back then, I \nwas convinced that shallow commits and grafted commits were something \nfundamentally different, and my recent patch to pack-objects shows that: \nshallow commits do not have the real parents in the current repository, \nand that makes them different from other grafted commits.\n\nNow, if you want to say that assume-unchanged and sparse are two \nfundamentally different things, I would be interested in some equally \nconvincing argument as for the shallow/graft issue.\n\nThere is a fundamental difference, I grant you that: the working directory \ndoes not contain the \"sparse'd away\" files while the same is not true for \nassume-unchanged files.\n\nBut does that matter?  The corresponding files are still in the index and \nthe repository.\n\nIOW under what circumstances would you want to be able to discern between \nassume-unchanged and \"sparse'd away\" files in the working directory?\n\nI could _imagine_ that you'd want a tool that allows you to change the \nfocus of the sparse checkout together with the working directory.  \nExample: you have a sparse checkout of Documentation/ and now you want to \nhave t/, too.  Just changing .git/info/sparse will not be enough.\n\nThe question is if the tool to change the \"sparseness\" [*1*] should not \nchange .git/info/sparse itself; if it does not, it would be good to be \nable to discern between the \"assume-unchanged\" and \"sparse'd away\" files.\n\nAlthough it might be enough to traverse the index and check the presence \nof the assume-unchanged files in the working directory to determine which \nfiles are sparse, and which ones are merely assume-unchanged.\n\nCiao,\nDscho\n\nFootnote [*1*]: I think we need some nice and clear nomenclature here.  \nAny English wizards with a good taste of naming things?\n"},{"id":"120851","messageId":"alpine.DEB.1.00.0908171101090.4991@intel-tinevez-2-302","threadId":"20539","inReplyTo":"alpine.DEB.1.00.0908161002460.8306@pacific.mpi-cbg.de","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-08-17T09:08:27Z","receivedAt":"2009-08-17T09:08:27Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 16 Aug 2009, Johannes Schindelin wrote:\n\n> [...] if you want to say that assume-unchanged and sparse are two \n> fundamentally different things, I would be interested in some equally \n> convincing argument as for the shallow/graft issue.\n> \n> There is a fundamental difference, I grant you that: the working \n> directory does not contain the \"sparse'd away\" files while the same is \n> not true for assume-unchanged files.\n> \n> But does that matter?  The corresponding files are still in the index \n> and the repository.\n> \n> IOW under what circumstances would you want to be able to discern \n> between assume-unchanged and \"sparse'd away\" files in the working \n> directory?\n> \n> I could _imagine_ that you'd want a tool that allows you to change the \n> focus of the sparse checkout together with the working directory.  \n> Example: you have a sparse checkout of Documentation/ and now you want \n> to have t/, too.  Just changing .git/info/sparse will not be enough.\n> \n> The question is if the tool to change the \"sparseness\" [*1*] should not \n> change .git/info/sparse itself; if it does not, it would be good to be \n> able to discern between the \"assume-unchanged\" and \"sparse'd away\" \n> files.\n> \n> Although it might be enough to traverse the index and check the presence \n> of the assume-unchanged files in the working directory to determine \n> which files are sparse, and which ones are merely assume-unchanged.\n> \n> Ciao,\n> Dscho\n> \n> Footnote [*1*]: I think we need some nice and clear nomenclature here.  \n> Any English wizards with a good taste of naming things?\n\nTurns out that somebody on IRC had a problem that requires to have \nsparse'd out files which _do_ have working directory copies.\n\nSo just having the assume-changed bit may not be enough.\n\nThe scenario is this: the repository contains a file that users are \nsupposed to change, but not commit to (only the super-intelligent inventor \nof this scenario is allowed to).  As this repository is originally a \nsubversion one, there is no problem: people just do not switch branches.\n\nBut this guy uses git-svn, so he does switch branches, and to avoid \ncommitting the file by mistake, he marked it assume-unchanged.  Only that \na branch switch overwrites the local changes.\n\nI suggested the use of the sparse feature, and mark this file (and this \nfile alone) as sparse'd-out.\n\nIs this an intended usage scenario?  Then we cannot reuse the \nassume-changed bit [*1*].\n\nCiao,\nDscho\n\nFootnote [*1*]: in this particular scenario, we could still discern \nbetween sparse'd-out and regular assume-unchanged file, because \n.git/info/sparse knows about the file.  But the design is now brittle, and \nit is not hard at all to come up with a situation where it breaks.\n"},{"id":"120880","messageId":"fcaeb9bf0908170549w26b008bdhe67f113a58ecb4eb@mail.gmail.com","threadId":"20539","inReplyTo":"alpine.DEB.1.00.0908171101090.4991@intel-tinevez-2-302","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-17T12:49:22Z","receivedAt":"2009-08-17T12:49:22Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Aug 17, 2009 at 4:08 PM, Johannes\nSchindelin<Johannes.Schindelin@gmx.de> wrote:\n> Turns out that somebody on IRC had a problem that requires to have\n> sparse'd out files which _do_ have working directory copies.\n>\n> So just having the assume-changed bit may not be enough.\n>\n> The scenario is this: the repository contains a file that users are\n> supposed to change, but not commit to (only the super-intelligent inventor\n> of this scenario is allowed to).  As this repository is originally a\n> subversion one, there is no problem: people just do not switch branches.\n>\n> But this guy uses git-svn, so he does switch branches, and to avoid\n> committing the file by mistake, he marked it assume-unchanged.\n\nHmm.. never thought of this use before. If he does not want to commit\nby mistake, should he add to-be-committed changes to index and do \"git\ncommit\" without \"-a\" (even better, do \"git diff --cached\" first)?\n\n> Only that a branch switch overwrites the local changes.\n\nI don't think branch switch overwrites changes in this case. Whenever\nGit is to touch worktree files, it ignores assumed-unchanged bit and\ndoes lstat() to make sure worktree files are up to date.\n\n> I suggested the use of the sparse feature, and mark this file (and this\n> file alone) as sparse'd-out.\n\nSparse checkout only removes a file if its assume-unchanged bit\nchanges from 0 to 1. If it's already 1, it does not care whether there\nis a corresponding file in worktree. So something like this should\nwork:\n\ngit checkout my-branch\ngit update-index --assume-unchanged that-special-file\necho that-special-file > .git/info/sparse\n# edit that-special-file\ngit commit -a\n# do whatever you want, git pull/checkout/read-tree... won't touch\nthat-special-file because it's assume-unchanged already\n\nToo subtle?\n\nAnyway I would not recommend this. the versions of that-special-file\nin worktree and and in index will diverse. When you unmark\nassume-unchanged (be it sparse checkout or plain assume-unchanged),\nyou may have already forgot what changes you made to this file and\n\"git diff\" would not help.\n\n> Is this an intended usage scenario?  Then we cannot reuse the\n> assume-changed bit [*1*].\n\nIt'd be great if people tell us all the scenarios they have. My use\ncould be too limited.\n-- \nDuy\n"},{"id":"120886","messageId":"alpine.DEB.1.00.0908171524150.4991@intel-tinevez-2-302","threadId":"20539","inReplyTo":"fcaeb9bf0908170549w26b008bdhe67f113a58ecb4eb@mail.gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-08-17T13:35:40Z","receivedAt":"2009-08-17T13:35:40Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 17 Aug 2009, Nguyen Thai Ngoc Duy wrote:\n\n> On Mon, Aug 17, 2009 at 4:08 PM, Johannes\n> Schindelin<Johannes.Schindelin@gmx.de> wrote:\n> > Turns out that somebody on IRC had a problem that requires to have \n> > sparse'd out files which _do_ have working directory copies.\n> >\n> > So just having the assume-changed bit may not be enough.\n> >\n> > The scenario is this: the repository contains a file that users are \n> > supposed to change, but not commit to (only the super-intelligent \n> > inventor of this scenario is allowed to).  As this repository is \n> > originally a subversion one, there is no problem: people just do not \n> > switch branches.\n> >\n> > But this guy uses git-svn, so he does switch branches, and to avoid \n> > committing the file by mistake, he marked it assume-unchanged.\n> \n> Hmm.. never thought of this use before. If he does not want to commit by \n> mistake, should he add to-be-committed changes to index and do \"git \n> commit\" without \"-a\" (even better, do \"git diff --cached\" first)?\n\nYou probably agree that this would be a _very_ fragile setup.  Very easy \nto make mistakes.\n\nBut we try to get away from that, don't we?  Git had a reputation to be \neasy fsck up for long enough.\n\n> > Only that a branch switch overwrites the local changes.\n> \n> I don't think branch switch overwrites changes in this case. Whenever\n> Git is to touch worktree files, it ignores assumed-unchanged bit and\n> does lstat() to make sure worktree files are up to date.\n\nWell, it does there, thankyouverymuch.\n\nThe problem of course is that the other branch has an ancient version of \nthat file (which should _not_ overwrite the current, modified version!), \ni.e. \"git diff HEAD..other -- file\" does not come empty.\n\nAs 'file' is assume-unchanged, zinnnng, the file gets \"updated\".\n\n> > I suggested the use of the sparse feature, and mark this file (and \n> > this file alone) as sparse'd-out.\n> \n> Sparse checkout only removes a file if its assume-unchanged bit\n> changes from 0 to 1.\n\nThe problem is not removing, but overwriting.\n\nAnd in this respect, 'assume-unchanged' is a very different beast from \n'sparse'.  I am growing more and more convinced that you cannot just reuse \nthe assume-unchanged bit.\n\n> If it's already 1, it does not care whether there is a corresponding \n> file in worktree. So something like this should work:\n> \n> git checkout my-branch\n> git update-index --assume-unchanged that-special-file\n> echo that-special-file > .git/info/sparse\n> # edit that-special-file\n> git commit -a\n> # do whatever you want, git pull/checkout/read-tree... won't touch\n> that-special-file because it's assume-unchanged already\n\n... except if you changed .git/info/sparse and a formerly sparse'd-out \nfile is overwritten by \"pull\".  Not good.\n\n> Anyway I would not recommend this. the versions of that-special-file in \n> worktree and and in index will diverse. When you unmark assume-unchanged \n> (be it sparse checkout or plain assume-unchanged), you may have already \n> forgot what changes you made to this file and \"git diff\" would not help.\n\nMy point is that we should take the current implementation as Dictated By \nThe Dear Lord, but change it if the limitation is too severe.\n\nAnd I do contend that 'assume-unchanged' is dissimilar enough from \n'sparse' to merit a change.\n\n> > Is this an intended usage scenario?  Then we cannot reuse the \n> > assume-changed bit [*1*].\n> \n> It'd be great if people tell us all the scenarios they have. My use \n> could be too limited.\n\nThe use case I would have is where a collaborator wants to work only on \none subdirectory and the top-level directory.  All other subdirectories \nare of no interest to him.\n\nAnother use case: documentation.  I do not have that use case yet, but I \nknow about people who do.  Specifying what you _want_ to have checked out \nis much more straight-forward here than the opposite.\n\nCiao,\nDscho\n\n"},{"id":"120895","messageId":"fcaeb9bf0908170741v210e7f4et9f1c68bc9a81ca65@mail.gmail.com","threadId":"20539","inReplyTo":"alpine.DEB.1.00.0908171524150.4991@intel-tinevez-2-302","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-17T14:41:13Z","receivedAt":"2009-08-17T14:41:13Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Aug 17, 2009 at 8:35 PM, Johannes\nSchindelin<Johannes.Schindelin@gmx.de> wrote:\n> Hi,\n>\n> On Mon, 17 Aug 2009, Nguyen Thai Ngoc Duy wrote:\n>\n>> On Mon, Aug 17, 2009 at 4:08 PM, Johannes\n>> Schindelin<Johannes.Schindelin@gmx.de> wrote:\n>> > Turns out that somebody on IRC had a problem that requires to have\n>> > sparse'd out files which _do_ have working directory copies.\n>> >\n>> > So just having the assume-changed bit may not be enough.\n>> >\n>> > The scenario is this: the repository contains a file that users are\n>> > supposed to change, but not commit to (only the super-intelligent\n>> > inventor of this scenario is allowed to).  As this repository is\n>> > originally a subversion one, there is no problem: people just do not\n>> > switch branches.\n>> >\n>> > But this guy uses git-svn, so he does switch branches, and to avoid\n>> > committing the file by mistake, he marked it assume-unchanged.\n>>\n>> Hmm.. never thought of this use before. If he does not want to commit by\n>> mistake, should he add to-be-committed changes to index and do \"git\n>> commit\" without \"-a\" (even better, do \"git diff --cached\" first)?\n>\n> You probably agree that this would be a _very_ fragile setup.  Very easy\n> to make mistakes.\n>\n> But we try to get away from that, don't we?  Git had a reputation to be\n> easy fsck up for long enough.\n\nWell.. of course I don't want Git to keep that reputation :-)\n\n>> > Only that a branch switch overwrites the local changes.\n>>\n>> I don't think branch switch overwrites changes in this case. Whenever\n>> Git is to touch worktree files, it ignores assumed-unchanged bit and\n>> does lstat() to make sure worktree files are up to date.\n>\n> Well, it does there, thankyouverymuch.\n>\n> The problem of course is that the other branch has an ancient version of\n> that file (which should _not_ overwrite the current, modified version!),\n> i.e. \"git diff HEAD..other -- file\" does not come empty.\n>\n> As 'file' is assume-unchanged, zinnnng, the file gets \"updated\".\n\nThen it is a bug. Assume-unchanged as in reading is good.\nAssume-unchanged in writing sounds scary. Something like this should\nfix it (not well tested though). It's on top of my series, but you can\nadapt it to 'next' or 'master' easily.\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex eb47676..7b9ddf6 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -538,7 +538,9 @@ static int verify_uptodate_1(struct cache_entry *ce,\n {\n \tstruct stat st;\n\n-\tif (o->index_only || o->reset || ce_uptodate(ce))\n+\tif (o->index_only || o->reset ||\n+\t    /* we are going to update worktree, don't trust ce_uptodate if\nit is CE_VALID'd */\n+\t    (!(ce->ce_flags & CE_VALID) && ce_uptodate(ce)))\n \t\treturn 0;\n\n \tif (!lstat(ce->name, &st)) {\n\n\n>> > I suggested the use of the sparse feature, and mark this file (and\n>> > this file alone) as sparse'd-out.\n>>\n>> Sparse checkout only removes a file if its assume-unchanged bit\n>> changes from 0 to 1.\n>\n> The problem is not removing, but overwriting.\n>\n> And in this respect, 'assume-unchanged' is a very different beast from\n> 'sparse'.  I am growing more and more convinced that you cannot just reuse\n> the assume-unchanged bit.\n\nAnd assume-unchanged bit could get lost during index merging, which\nmay cause unexpected effect if sparse checkout bases off\nassume-unchanged. Let me think more of it tonight.\n\n>> If it's already 1, it does not care whether there is a corresponding\n>> file in worktree. So something like this should work:\n>>\n>> git checkout my-branch\n>> git update-index --assume-unchanged that-special-file\n>> echo that-special-file > .git/info/sparse\n>> # edit that-special-file\n>> git commit -a\n>> # do whatever you want, git pull/checkout/read-tree... won't touch\n>> that-special-file because it's assume-unchanged already\n>\n> ... except if you changed .git/info/sparse and a formerly sparse'd-out\n> file is overwritten by \"pull\".  Not good.\n\nAgain, I think it's a bug.\n\n>> > Is this an intended usage scenario?  Then we cannot reuse the\n>> > assume-changed bit [*1*].\n>>\n>> It'd be great if people tell us all the scenarios they have. My use\n>> could be too limited.\n>\n> The use case I would have is where a collaborator wants to work only on\n> one subdirectory and the top-level directory.  All other subdirectories\n> are of no interest to him.\n>\n> Another use case: documentation.  I do not have that use case yet, but I\n> know about people who do.\n\nTranslators usually checkout one or two files (I am Vietnamese\nTranslation Coordinator of GNOME, but well... I check them all out. I\nsuppose \"normal\" translators would not want to do like I do.)\n\n>  Specifying what you _want_ to have checked out\n> is much more straight-forward here than the opposite.\n\nI think it depends on type of projects. For documentation projects,\nyou may want a few files. For software projects, usually you need\neverything _except_ a few big directories. For WebKit, it's a bunch of\ntest data that I don't care about. Firmware in hardware-related\nprojects or media files in game projects fall in the same category. I\ndon't have strong opinion on this. Either include or exclude is fine\nto me.\n-- \nDuy\n"},{"id":"120898","messageId":"alpine.DEB.1.00.0908171712220.4991@intel-tinevez-2-302","threadId":"20539","inReplyTo":"fcaeb9bf0908170741v210e7f4et9f1c68bc9a81ca65@mail.gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-08-17T15:19:31Z","receivedAt":"2009-08-17T15:19:31Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 17 Aug 2009, Nguyen Thai Ngoc Duy wrote:\n\n> On Mon, Aug 17, 2009 at 8:35 PM, Johannes\n> Schindelin<Johannes.Schindelin@gmx.de> wrote:\n>\n> > The problem of course is that the other branch has an ancient version \n> > of that file (which should _not_ overwrite the current, modified \n> > version!), i.e. \"git diff HEAD..other -- file\" does not come empty.\n> >\n> > As 'file' is assume-unchanged, zinnnng, the file gets \"updated\".\n> \n> Then it is a bug. Assume-unchanged as in reading is good.\n> Assume-unchanged in writing sounds scary. Something like this should\n> fix it (not well tested though). It's on top of my series, but you can\n> adapt it to 'next' or 'master' easily.\n\nNo.\n\nThe purpose of 'assume-unchanged' is to tell Git that it has no business \nchecking that the file is unchanged.  It should _assume_ that it is \nunchanged.  That's what this flag says.\n\nSo do you agree that assume-changed is not quite similar enough to sparse \nto use the same bit?\n\n> > Another use case: documentation.  I do not have that use case yet, but \n> > I know about people who do.\n> \n> Translators usually checkout one or two files (I am Vietnamese \n> Translation Coordinator of GNOME, but well... I check them all out. I \n> suppose \"normal\" translators would not want to do like I do.)\n\nExactly.\n\necho /Documentation/ > .git/info/sparse\n\nRemember: the documentation contributors are the least programming-savvy \ncontributors of any project.\n\n> >  Specifying what you _want_ to have checked out is much more \n> > straight-forward here than the opposite.\n> \n> I think it depends on type of projects. For documentation projects, you \n> may want a few files. For software projects, usually you need everything \n> _except_ a few big directories. For WebKit, it's a bunch of test data \n> that I don't care about. Firmware in hardware-related projects or media \n> files in game projects fall in the same category. I don't have strong \n> opinion on this. Either include or exclude is fine to me.\n\nOkay, let me just ask: if you have a sparse checkout, what would you think \nI mean when I talk about the \"sparse files\"?\n\nCiao,\nDscho\n"},{"id":"120901","messageId":"7vtz06xxao.fsf@alter.siamese.dyndns.org","threadId":"20539","inReplyTo":"alpine.DEB.1.00.0908171101090.4991@intel-tinevez-2-302","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-08-17T15:41:35Z","receivedAt":"2009-08-17T15:41:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> The scenario is this: the repository contains a file that users are \n> supposed to change, but not commit to (only the super-intelligent inventor \n> of this scenario is allowed to).  As this repository is originally a \n> subversion one, there is no problem: people just do not switch branches.\n>\n> But this guy uses git-svn, so he does switch branches, and to avoid \n> committing the file by mistake, he marked it assume-unchanged.  Only that \n> a branch switch overwrites the local changes.\n\nIf it is a problem that a branch switch overwrites the local changes in\nassume-unchanged file, perhaps that is what this person needs to change?\n\nLet's step back a bit and think.\n\nLocal changes in git do not belong to any particular branch.  They belong\nto the work tree and the index.  Hence you (1) can switch from branch A to\nbranch B iff the branches do not have difference in the path with local\nchanges, and (2) have to stash save, switch branches and then stash pop if\nyou have local changes to paths that are different between branches you\nare switching between.\n\nHow should assume-unchanged play with this philosophy?\n\nI'd say that assume-unchanged is a promise you make git that you won't\nchange these paths, and in return to the promise git will give you faster\nresponse by not running lstat on them.  Having changes in such paths is\nyour problem and you deserve these chanegs to be lost.  At least, that is\nthe interpretation according to the original assume-unchanged semantics.\n\nIf some paths should not be committed, I'd say it should be handled by a\npre commit hook, and not assume-unchanged.\n\nIs checking with \"diff --cached\" on the paths and either erroring out (or\nbetter yet resetting the problematic paths in the index) an option?\n"},{"id":"120903","messageId":"200908171802.00588.jnareb@gmail.com","threadId":"20539","inReplyTo":"alpine.DEB.1.00.0908161002460.8306@pacific.mpi-cbg.de","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-08-17T16:01:58Z","receivedAt":"2009-08-17T16:01:58Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sun, 16 Aug 2009, Johannes Schindelin wrote:\n> On Sun, 16 Aug 2009, Jakub Narebski wrote:\n>> On Sat, 15 Aug 2009, Junio C Hamano wrote:\n>>> Jakub Narebski <jnareb@gmail.com> writes:\n>>> \n>>>>>> Hmmm... this looks like either argument for introducing --full \n>>>>>> option to git-checkout (ignore CE_VALID bit, checkout everything, \n>>>>>> and clean CE_VALID (?))...\n>>>>>>\n>>>>>>  ...or for going with _separate_ bit for partial checkout, like in \n>>>>>>  the very first version of this series, which otherwise functions \n>>>>>>  like CE_VALID, or is just used to mark that CE_VALID was set using \n>>>>>>  sparse.\n>>> \n>>> How would a separate bit help?  Just like you need to clear CE_VALID \n>>> bit to revert the index into a normal (or \"non sparse\") state somehow, \n>>> you would need to have a way to clear that separate bit anyway.\n>>> \n>>> A separate bit would help only if you want to handle assume-unchanged \n>>> and sparse checkout independently. But my impression was that the \n>>> recent lstat reduction effort addressed the issue assume-unchanged \n>>> were invented to work around in the first place.\n>> \n>> Well, if we assume that we don't need (don't want) to handle \n>> assume-unchanged and sparse checkout independently, then of course the \n>> idea of having separate or additional bit for sparse doesn't make sense.\n> \n> For the shallow/graft issue, we had a similar discussion.  Back then, I \n> was convinced that shallow commits and grafted commits were something \n> fundamentally different, and my recent patch to pack-objects shows that: \n> shallow commits do not have the real parents in the current repository, \n> and that makes them different from other grafted commits.\n> \n> Now, if you want to say that assume-unchanged and sparse are two \n> fundamentally different things, I would be interested in some equally \n> convincing argument as for the shallow/graft issue.\n> \n> There is a fundamental difference, I grant you that: the working directory \n> does not contain the \"sparse'd away\" files while the same is not true for \n> assume-unchanged files.\n> \n> But does that matter?  The corresponding files are still in the index and \n> the repository.\n> \n> IOW under what circumstances would you want to be able to discern between \n> assume-unchanged and \"sparse'd away\" files in the working directory?\n\n>From what I understand it, assume-unchanged is performance optimization.\nSparse checkout is about files which are (assumed to) not be in working\ndirectory, which means that they have to be assume-unchanged for git to\nnot try to access working area version of files which aren't there.\n\n$GIT_DIR/info/sparse (or how it would be named; the name 'sparse' \ndoesn't tell us whether patterns are about the files that are checked\nout, or are about files which are not present in working directory)\nis about specifying which files to checkout with \"git checkout --sparse\"\n(or core.sparse / checkout.sparse = true).\n\n> I could _imagine_ that you'd want a tool that allows you to change the \n> focus of the sparse checkout together with the working directory.  \n> Example: you have a sparse checkout of Documentation/ and now you want to \n> have t/, too.  Just changing .git/info/sparse will not be enough.\n> \n> The question is if the tool to change the \"sparseness\" [*1*] should not \n> change .git/info/sparse itself; if it does not, it would be good to be \n> able to discern between the \"assume-unchanged\" and \"sparse'd away\" files.\n> \n> Although it might be enough to traverse the index and check the presence \n> of the assume-unchanged files in the working directory to determine which \n> files are sparse, and which ones are merely assume-unchanged.\n\nThere are quite a few possibilities: file can be marked \"sparse\" in\nindex (which also implies also marking it \"assume-unchanged\", if \n\"assume-unchanged\" doesn't work alone as \"sparse\" index bit) or not,\nfile can match 'no-checkout' pattern in $GIT_DIR/info/sparse or not,\nfile can be present in working directory or not:\n\n * match no-checkout\n   - assume-unchanged\n     + present in working directory\n     + absent from working directory\n   - no assume-unchanged\n     + present\n     + absent\n * doesn't match no-checkout\n   - assume-unchanged\n     + present\n     + absent\n   - no assume-unchanged\n     + present\n     + absent\n\n> Footnote [*1*]: I think we need some nice and clear nomenclature here.  \n> Any English wizards with a good taste of naming things?\n \nEnglish is not my native language, but what about:\n\n $GIT_DIR/info/\n    no-checkout\n    exclude-checkout\n    workdir-exclude\n    ignore-change\n    assume-unchanged\n    ghosts\n  \n-- \nJakub Narebski\nPoland\n"},{"id":"120916","messageId":"fcaeb9bf0908170906m739be4abs80ee3a6fc7135fbd@mail.gmail.com","threadId":"20539","inReplyTo":"7vtz06xxao.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-17T16:06:00Z","receivedAt":"2009-08-17T16:06:00Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Aug 17, 2009 at 10:41 PM, Junio C Hamano<gitster@pobox.com> wrote:\n> How should assume-unchanged play with this philosophy?\n>\n> I'd say that assume-unchanged is a promise you make git that you won't\n> change these paths, and in return to the promise git will give you faster\n> response by not running lstat on them.  Having changes in such paths is\n> your problem and you deserve these chanegs to be lost.  At least, that is\n> the interpretation according to the original assume-unchanged semantics.\n\nBut commit 5f73076 (\"Assume unchanged\" git) says [1] it favors safety\nover performance? Otherwise I'd need to resurrect no-checkout bit.\n\n[1] excerpt from the mentioned commit:\n--<--\n    Index entries marked with CE_VALID bit are assumed to be\n    unchanged most of the time.  However, there are cases that\n    CE_VALID bit is ignored for the sake of safety and usability:\n\n     - while \"git-read-tree -m\" or git-apply need to make sure\n       that the paths involved in the merge do not have local\n       modifications.  This sacrifices performance for safety.\n\n     - when git-checkout-index -f -q -u -a tries to see if it needs\n       to checkout the paths.  Otherwise you can never check\n       anything out ;-).\n\n     - when git-update-index --really-refresh (a new flag) tries to\n       see if the index entry is up to date.  You can start with\n       everything marked as CE_VALID and run this once to drop\n       CE_VALID bit for paths that are modified.\n--<--\n-- \nDuy\n"},{"id":"120921","messageId":"fcaeb9bf0908170913l20d3cc0ma81052589a8a685f@mail.gmail.com","threadId":"20539","inReplyTo":"alpine.DEB.1.00.0908171712220.4991@intel-tinevez-2-302","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-17T16:13:19Z","receivedAt":"2009-08-17T16:13:19Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Aug 17, 2009 at 10:19 PM, Johannes\nSchindelin<Johannes.Schindelin@gmx.de> wrote:\n> Hi,\n>\n> On Mon, 17 Aug 2009, Nguyen Thai Ngoc Duy wrote:\n>\n>> On Mon, Aug 17, 2009 at 8:35 PM, Johannes\n>> Schindelin<Johannes.Schindelin@gmx.de> wrote:\n>>\n>> > The problem of course is that the other branch has an ancient version\n>> > of that file (which should _not_ overwrite the current, modified\n>> > version!), i.e. \"git diff HEAD..other -- file\" does not come empty.\n>> >\n>> > As 'file' is assume-unchanged, zinnnng, the file gets \"updated\".\n>>\n>> Then it is a bug. Assume-unchanged as in reading is good.\n>> Assume-unchanged in writing sounds scary. Something like this should\n>> fix it (not well tested though). It's on top of my series, but you can\n>> adapt it to 'next' or 'master' easily.\n>\n> No.\n>\n> The purpose of 'assume-unchanged' is to tell Git that it has no business\n> checking that the file is unchanged.  It should _assume_ that it is\n> unchanged.  That's what this flag says.\n>\n> So do you agree that assume-changed is not quite similar enough to sparse\n> to use the same bit?\n\nIf you define it that way, yes I agree.\n\n>> > Another use case: documentation.  I do not have that use case yet, but\n>> > I know about people who do.\n>>\n>> Translators usually checkout one or two files (I am Vietnamese\n>> Translation Coordinator of GNOME, but well... I check them all out. I\n>> suppose \"normal\" translators would not want to do like I do.)\n>\n> Exactly.\n>\n> echo /Documentation/ > .git/info/sparse\n>\n> Remember: the documentation contributors are the least programming-savvy\n> contributors of any project.\n\n[wanted to make a joke here, but it seemed destructive, snipped]\n\n>> >  Specifying what you _want_ to have checked out is much more\n>> > straight-forward here than the opposite.\n>>\n>> I think it depends on type of projects. For documentation projects, you\n>> may want a few files. For software projects, usually you need everything\n>> _except_ a few big directories. For WebKit, it's a bunch of test data\n>> that I don't care about. Firmware in hardware-related projects or media\n>> files in game projects fall in the same category. I don't have strong\n>> opinion on this. Either include or exclude is fine to me.\n>\n> Okay, let me just ask: if you have a sparse checkout, what would you think\n> I mean when I talk about the \"sparse files\"?\n\nIf I have to answer in 2 seconds, \"sparse files\" are files in working\ndirectory. If I have more time, I tend to think that in \"sparse\n<something>\", something should be a container, an area, therefore\n\"sparse files\" do not make sense to me while \"sparse\ncheckout/worktree\" does. So, .git/info/sparse-checkout (with \"in\"\npatterns)?\n-- \nDuy\n"},{"id":"120922","messageId":"alpine.DEB.1.00.0908171817570.4991@intel-tinevez-2-302","threadId":"20539","inReplyTo":"7vtz06xxao.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-08-17T16:19:54Z","receivedAt":"2009-08-17T16:19:54Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 17 Aug 2009, Junio C Hamano wrote:\n\n> I'd say that assume-unchanged is a promise you make git that you won't \n> change these paths, and in return to the promise git will give you \n> faster response by not running lstat on them.  Having changes in such \n> paths is your problem and you deserve these chanegs to be lost.  At \n> least, that is the interpretation according to the original \n> assume-unchanged semantics.\n\nThat's why I did not suggest using assume-unchanged (which the guy did \npreviously, and was burnt, deservedly, as you say).\n\nHowever, my illustration of the scenario was only to one end, namely to \nconvince all of you that assume-changed != sparse.\n\nAnd maybe to the end to explain that sparse checkout could help this guy.\n\nCiao,\nDscho\n"},{"id":"120931","messageId":"7vvdkmwfqs.fsf@alter.siamese.dyndns.org","threadId":"20539","inReplyTo":"7vtz06xxao.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-08-17T16:46:03Z","receivedAt":"2009-08-17T16:46:03Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Local changes in git do not belong to any particular branch.  They belong\n> to the work tree and the index.  Hence you (1) can switch from branch A to\n> branch B iff the branches do not have difference in the path with local\n> changes, and (2) have to stash save, switch branches and then stash pop if\n> you have local changes to paths that are different between branches you\n> are switching between.\n>\n> How should assume-unchanged play with this philosophy?\n>\n> I'd say that assume-unchanged is a promise you make git that you won't\n> change these paths, and in return to the promise git will give you faster\n> response by not running lstat on them.  Having changes in such paths is\n> your problem and you deserve these chanegs to be lost.  At least, that is\n> the interpretation according to the original assume-unchanged semantics.\n\nHaving said that, we could (re)define assume-unchanged to mean \"I may or\nmay not have changes to these paths, but I do not mean to commit them, so\ndo not show them as modified when I ask you for diff.  But the changes are\nprecious nevertheless\".\n\nI think the writeout codepath pays attention to assume-unchanged bit\nalready for that reason (CE_MATCH_IGNORE_VALID is all about this issue).\n\nSo with that, how should assume-unchanged play with the \"local changes\nbelong to the index and the work tree\"?\n\n - When adding to the index, the changes should be ignored;\n\n - When checking out of the index?  I.e. the user tells \"git checkout\n   path\" when path is marked as assume-unchanged.  Such an explicit\n   request should probably lose the local changes in the work tree.\n\n - When checking out of a commit?  The same deal.\n\n - When switching branches?\n\n   - If the branches do not touch assume-unchanged paths, we should keep\n     changes _and_ assume-unchanged bit.  I do not know if that is what\n     the current code does.\n\n   - If the branches do touch assume-unchanged paths, what should happen?\n     We shouldn't blindly overwrite the local changes, so at least we\n     should change the code to error out if we do not already do so.  But\n     then what?  How does the user deal with this?  Perhaps...\n\n     - Drop assume-unchanged temporarily;\n     - Stash save;\n     - Switch;\n     - Stash pop;\n     - Add assume-unchanged again.\n\n     ???\n\nIs such an updated (or \"corrected\") assume-unchanged any different from a\nsparse checkout?  After all, paths that are not to be checked out in a\nsparse checkout are \"pretend that the lack of these paths are illusion--they\nare logically there.  I do not intend to commit their removal, and I do not\nwant to lose the sparseness across branch switch\".\n\nThere is one nit about this.  If a path is outside the checkout area,\nshould it unconditionally stay outside the checkout area when you switch\nbranches?  I may be interested in not checking out Documentation/\nsubdirectory and that may hold true for all _my_ branches, and it is a\nsane thing not to complain \"Oops, you actually removed Makefile in\nDocumentation/ in your work tree in reality, and you are switching to\nanother branch that has a different Makefile --- it is a delete-modify\nconflict you need to resolve, and we won't let you switch branches\" in\nsuch a case.\n\nBut is that generally true in all \"sparse checkout\" settings?\n\nIt is unfortunate that this message raises more questions than it answers,\nbut I think a sparse checkout will have to answer them, whether it uses a\nbit separate from assume-unchanged or it reuses the assume-unchanged bit.\n"},{"id":"120957","messageId":"7vws52uvxq.fsf@alter.siamese.dyndns.org","threadId":"20539","inReplyTo":"alpine.DEB.1.00.0908171817570.4991@intel-tinevez-2-302","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-08-17T18:39:13Z","receivedAt":"2009-08-17T18:39:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> However, my illustration of the scenario was only to one end, namely to \n> convince all of you that assume-changed != sparse.\n>\n> And maybe to the end to explain that sparse checkout could help this guy.\n\nHow?  If sparse is _not to check it out_, then that is not what the person\nis doing either.  It feels to me that you are suggesting an inappropriate\nhack to replace another inappropriate hack, suggesting to use a hacksaw\nbecause an earlier attempt to use a hammer did not quite work to drive the\nscrew in.\n\nI never said assume-unchanged _is_ sparse.  You cannot mark an index entry\nthat does not exist, obviously you need more (either the earlier \"hook\nthat tells what should/shouldn't exist\", or \"the pattern\").\n\nBut I think the work-tree semantics you need to _implement_ sparse matches\nwhat you would want from assume-unchanged.  Not the original, draconian\none that updates the work tree by saying \"you promised me you wouldn't\nchange them\", but the updated one that tells git to pretend that the local\nchange is not there but still keep the local modification, including\ndeletion.  The work-tree \"local changes\" sparse makes is a small subset of\npossible local changes assume-unchanged would need to support.  It only\ndeletes work tree files.\n"},{"id":"121002","messageId":"alpine.DEB.1.00.0908172339040.8306@pacific.mpi-cbg.de","threadId":"20539","inReplyTo":"7vvdkmwfqs.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-08-17T21:45:30Z","receivedAt":"2009-08-17T21:45:30Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 17 Aug 2009, Junio C Hamano wrote:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n> \n> > Local changes in git do not belong to any particular branch.  They \n> > belong to the work tree and the index.  Hence you (1) can switch from \n> > branch A to branch B iff the branches do not have difference in the \n> > path with local changes, and (2) have to stash save, switch branches \n> > and then stash pop if you have local changes to paths that are \n> > different between branches you are switching between.\n> >\n> > How should assume-unchanged play with this philosophy?\n> >\n> > I'd say that assume-unchanged is a promise you make git that you won't \n> > change these paths, and in return to the promise git will give you \n> > faster response by not running lstat on them.  Having changes in such \n> > paths is your problem and you deserve these chanegs to be lost.  At \n> > least, that is the interpretation according to the original \n> > assume-unchanged semantics.\n> \n> Having said that, we could (re)define assume-unchanged to mean \"I may or\n> may not have changes to these paths, but I do not mean to commit them, so\n> do not show them as modified when I ask you for diff.  But the changes are\n> precious nevertheless\"\n\nI am hesitant.  The feature was introduced because of some report that Git \nwas too slow.  While the speed has increased dramatically (the report was \nfor Windows), we cannot work miracles: the file system layer is just not \ncooperative.\n\nSo I could imagine that redefining the meaning of assume-unchanged results \nin a substantially longer runtime again, which some people (yours truly) \nmight interpret as regression.\n\n> I think the writeout codepath pays attention to assume-unchanged bit \n> already for that reason (CE_MATCH_IGNORE_VALID is all about this issue).\n\nIf I were that reporter, I would not be happy, I guess.  Basically, when \nnew files come in and I marked all the files as \"assume unchanged; I know \nwhat I'm doing!\" Git would tell me \"no you're an idiot, dummy, I know \nbetter, and I will check all over again!\".\n\n> So with that, how should assume-unchanged play with the \"local changes \n> belong to the index and the work tree\"?\n> \n>  - When adding to the index, the changes should be ignored;\n> \n>  - When checking out of the index?  I.e. the user tells \"git checkout\n>    path\" when path is marked as assume-unchanged.  Such an explicit\n>    request should probably lose the local changes in the work tree.\n> \n>  - When checking out of a commit?  The same deal.\n> \n>  - When switching branches?\n> \n>    - If the branches do not touch assume-unchanged paths, we should keep\n>      changes _and_ assume-unchanged bit.  I do not know if that is what\n>      the current code does.\n> \n>    - If the branches do touch assume-unchanged paths, what should happen?\n>      We shouldn't blindly overwrite the local changes, so at least we\n>      should change the code to error out if we do not already do so.  But\n>      then what?  How does the user deal with this?  Perhaps...\n> \n>      - Drop assume-unchanged temporarily;\n>      - Stash save;\n>      - Switch;\n>      - Stash pop;\n>      - Add assume-unchanged again.\n> \n>      ???\n\nIn my book all this is overly complicated.  If I tell Git to assume a file \nis unchanged, it is not Git's business to question me.\n\n> Is such an updated (or \"corrected\") assume-unchanged any different from a\n> sparse checkout?  After all, paths that are not to be checked out in a\n> sparse checkout are \"pretend that the lack of these paths are illusion--they\n> are logically there.  I do not intend to commit their removal, and I do not\n> want to lose the sparseness across branch switch\".\n> \n> There is one nit about this.  If a path is outside the checkout area,\n> should it unconditionally stay outside the checkout area when you switch\n> branches?  I may be interested in not checking out Documentation/\n> subdirectory and that may hold true for all _my_ branches, and it is a\n> sane thing not to complain \"Oops, you actually removed Makefile in\n> Documentation/ in your work tree in reality, and you are switching to\n> another branch that has a different Makefile --- it is a delete-modify\n> conflict you need to resolve, and we won't let you switch branches\" in\n> such a case.\n> \n> But is that generally true in all \"sparse checkout\" settings?\n> \n> It is unfortunate that this message raises more questions than it answers,\n> but I think a sparse checkout will have to answer them, whether it uses a\n> bit separate from assume-unchanged or it reuses the assume-unchanged bit.\n\nI think you will come around and agree that the original, very simple \ntherefore powerful, concept of \"assume-unchanged\" should be, well, \nunchanged, and not be bent to half-fit the original intention and half-fit \nthe sparse intention.\n\nRather, I agree with Nguy�n that the no-checkout bit (which is definitely \nfree to behave differently from assume-unchanged) is needed.\n\nMaybe I contradict myself here with what I said after the third iteration \nof the sparse checkout series, but that only proves that I am able to \nlearn.\n\nCiao,\nDscho\n"},{"id":"121003","messageId":"alpine.DEB.1.00.0908172347220.8306@pacific.mpi-cbg.de","threadId":"20539","inReplyTo":"7vws52uvxq.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-08-17T22:02:55Z","receivedAt":"2009-08-17T22:02:55Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 17 Aug 2009, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > However, my illustration of the scenario was only to one end, namely \n> > to convince all of you that assume-changed != sparse.\n> >\n> > And maybe to the end to explain that sparse checkout could help this \n> > guy.\n> \n> How?  If sparse is _not to check it out_, then that is not what the \n> person is doing either.  It feels to me that you are suggesting an \n> inappropriate hack to replace another inappropriate hack, suggesting to \n> use a hacksaw because an earlier attempt to use a hammer did not quite \n> work to drive the screw in.\n\nNot exactly.\n\nWhat does \"sparse checkout\" mean, really?  It means that Git should only \ncheck out a part of the tracked files, and not even so much as look \noutside.  It means to me that everything outside of that focus is clearly \nto be handled as all the other untracked data.\n\nAnd here comes the problem: if something is treated untracked because it \nwas outside of the sparse checkout, then I want it to be treated as \nuntracked _even if_ I happened to broaden the checkout by editing \n.git/info/sparse.  The file did not just magically become subject to \noverwriting just because I edited .git/info/sparse (which could be a \nsimple mistake).  So the index _needs_ to know that the sparse'd-out \nattribute is something completely different from the assume-unchanged \nattribute, even if Git should _handle_ the files with those attributes \npretty similar _most_ of the time.\n\n> I never said assume-unchanged _is_ sparse.  You cannot mark an index \n> entry that does not exist, obviously you need more (either the earlier \n> \"hook that tells what should/shouldn't exist\", or \"the pattern\").\n\nRight.\n\n> But I think the work-tree semantics you need to _implement_ sparse \n> matches what you would want from assume-unchanged.  Not the original, \n> draconian one that updates the work tree by saying \"you promised me you \n> wouldn't change them\", but the updated one that tells git to pretend \n> that the local change is not there but still keep the local \n> modification, including deletion.  The work-tree \"local changes\" sparse \n> makes is a small subset of possible local changes assume-unchanged would \n> need to support.  It only deletes work tree files.\n\nAs I tried to convince you already, it is not wise to mix up the two \nmeanings.  They _are_ different: in one case, we _have_ a file, and we \neven _expect_ the file to actually have the same contents as what is \nrecorded in the index.  In the other case, we do _not_ have a file, so we \ndo _not_ even expect the file to have the same contents.\n\nIn fact, in the latter case (the sparse case) we do not want to look for \nthe file; not for the reason that we expect the contents to be the same \nanyway, but because we expect it not even to be there!\n\nSo while the _technical_ side is pretty much the same (most of the time, I \nillustrated a corner case, it it is very easy to think of other corner \ncases that might even be inadvertent, all the more reason to protect the \nuser) -- don't look for the file -- the _semantics_ are _very_ different.\n\nAnd you see that they are different when all of a sudden you cannot take \nthe _absence_ of the file as the indicator for \"assume-unchanged\" and \n\"sparse\".\n\nIn fact, with the semantics implied by the label 'assume-unchanged', it \ncould well be argued that making the file _absent_ (for the sparse \ncheckout) is a dirty trick.  This is not what \"assume that the file is \nunchanged\" implies at all.\n\nSo let's just keep the semantics utterly simple and stupid, and have an\n\n- assumed-unchanged bit, which assumes that a file is there, but that the \n  contents need not to be checked for performance reasons, and\n\n- a no-checkout bit, which assumes that the user never checked out that \n  file (if it exists, it comes from somewhere else, and needs to be \n  protected like untracked files that would be overwritten by a branch \n  switch).\n\nI hope this explanation was clear.\n\nCiao,\nDscho\n"},{"id":"121019","messageId":"2729632a0908171602m3c05c97bx9ce31e8960df9198@mail.gmail.com","threadId":"20539","inReplyTo":"alpine.DEB.1.00.0908172347220.8306@pacific.mpi-cbg.de","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"","fromEmail":"skillzero@gmail.com","sentAt":"2009-08-17T23:02:13Z","receivedAt":"2009-08-17T23:02:13Z","isPatch":true,"sender":{"key":"skillzero@gmail.com","avatar":null},"body":"On Mon, Aug 17, 2009 at 3:02 PM, Johannes\nSchindelin<Johannes.Schindelin@gmx.de> wrote:\n\n> And here comes the problem: if something is treated untracked because it\n> was outside of the sparse checkout, then I want it to be treated as\n> untracked _even if_ I happened to broaden the checkout by editing\n> .git/info/sparse.  The file did not just magically become subject to\n> overwriting just because I edited .git/info/sparse (which could be a\n> simple mistake).\n\nMaybe I'm misunderstanding what you're saying, but why would you want\na file that's become part of the checkout by editing .git/info/sparse\nto still be treated as untracked?\n\nIf I have a file on that's excluded via .git/info/sparse then I edit\n.git/info/sparse to include it and switch to a branch that doesn't\nhave that file, I'd expect that file to be deleted from the working\ncopy if the content matches what's in the repository. If it's modified\nthen I'd expect the branch switch to fail (like it would without a\nsparse checkout).\n"},{"id":"121023","messageId":"alpine.DEB.1.00.0908180111340.8306@pacific.mpi-cbg.de","threadId":"20539","inReplyTo":"2729632a0908171602m3c05c97bx9ce31e8960df9198@mail.gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-08-17T23:16:39Z","receivedAt":"2009-08-17T23:16:39Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 17 Aug 2009, skillzero@gmail.com wrote:\n\n> On Mon, Aug 17, 2009 at 3:02 PM, Johannes\n> Schindelin<Johannes.Schindelin@gmx.de> wrote:\n> \n> > And here comes the problem: if something is treated untracked because \n> > it was outside of the sparse checkout, then I want it to be treated as \n> > untracked _even if_ I happened to broaden the checkout by editing \n> > .git/info/sparse.  The file did not just magically become subject to \n> > overwriting just because I edited .git/info/sparse (which could be a \n> > simple mistake).\n> \n> Maybe I'm misunderstanding what you're saying, but why would you want a \n> file that's become part of the checkout by editing .git/info/sparse to \n> still be treated as untracked?\n> \n> If I have a file on that's excluded via .git/info/sparse then I edit \n> .git/info/sparse to include it and switch to a branch that doesn't have \n> that file, I'd expect that file to be deleted from the working copy if \n> the content matches what's in the repository. If it's modified then I'd \n> expect the branch switch to fail (like it would without a sparse \n> checkout).\n\nFirst things first: with sparse checkout, you should not check out \n_anything_ outside of the focus of the sparse checkout.\n\nSo I contend that you would only end up with a sparse'd-out file \nthat was formerly tracked if you did something wrong.  That should not \nhappen.\n\nEven if: all the more reason to have a flag that indicated that this file \nis not sparsed'd-out -- contradicting .git/info/sparse.\n\nThe thing is: we need a way to determine quickly and without any \nambiguity whether a file is tracked, assumed unchanged, or sparse'd-out \n(which Nguyễn calls no-checkout).\n\nAnd if we change .git/info/sparse, that state _must not_ change.  We did \nnot touch the file by editing .git/info/sparse, so the state must be \nunchanged.\n\nWhether \"git checkout\" should realize that a checked out file (which has \nno changes, mind you!) needs to be deleted and marked no-checkout is a \ndifferent question.\n\nCiao,\nDscho\n\n"},{"id":"121031","messageId":"200908180217.35963.jnareb@gmail.com","threadId":"20539","inReplyTo":"alpine.DEB.1.00.0908180111340.8306@pacific.mpi-cbg.de","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-08-18T00:17:33Z","receivedAt":"2009-08-18T00:17:33Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Johannes Schindelin wrote:\n\n> The thing is: we need a way to determine quickly and without any \n> ambiguity whether a file is tracked, assumed unchanged, or sparse'd-out \n> (which Nguyễn calls no-checkout).\n\nLet's reiterate: \"assume-unchanged\" is about telling git that it should\nassume for performance reasons that state of file in working directory\nis the same as state of file in the index.  But, from what was said in\nthis thread, there are situations where git for correctness reasons\nignores performance hack.\n\n\"no-checkout\" bit is about telling git that the file is not present\nin working directory, and it has to use version from the index.  Then\nthere is a question if there is file in working area (e.g. from applying\npatch) which corresponds to a \"no-checkout\" file in index (corresponds\nbecause of rename detection).\n\n> And if we change .git/info/sparse, that state _must not_ change.  We did \n> not touch the file by editing .git/info/sparse, so the state must be \n> unchanged.\n\nI think this situation (and the issue of correctness vs \"assume-unchanged\"\nmentioned above) hints that \"no-checkout\" and \"assume-unchanged\" should\nbe separate bits, even if both tell git to use version from index.\n\nThere is e.g. question if \"git grep\" should search \"no-checkout\" files;\nin the \"assume-unchanged\" case it should, I think, search index version.\n\n\nP.S. I wonder if it would be worth resurrecting series adding support\nfor directories in index (which can help performance and 'empty \ndirectories' issue)...  It would help, I think, with sparse checkout.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"121034","messageId":"2729632a0908171723n1c70798bp3169c813d5734c70@mail.gmail.com","threadId":"20539","inReplyTo":"alpine.DEB.1.00.0908180111340.8306@pacific.mpi-cbg.de","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"","fromEmail":"skillzero@gmail.com","sentAt":"2009-08-18T00:23:39Z","receivedAt":"2009-08-18T00:23:39Z","isPatch":true,"sender":{"key":"skillzero@gmail.com","avatar":null},"body":"On Mon, Aug 17, 2009 at 4:16 PM, Johannes\nSchindelin<Johannes.Schindelin@gmx.de> wrote:\n> Hi,\n>\n> On Mon, 17 Aug 2009, skillzero@gmail.com wrote:\n>\n>> On Mon, Aug 17, 2009 at 3:02 PM, Johannes\n>> Schindelin<Johannes.Schindelin@gmx.de> wrote:\n>>\n>> > And here comes the problem: if something is treated untracked because\n>> > it was outside of the sparse checkout, then I want it to be treated as\n>> > untracked _even if_ I happened to broaden the checkout by editing\n>> > .git/info/sparse.  The file did not just magically become subject to\n>> > overwriting just because I edited .git/info/sparse (which could be a\n>> > simple mistake).\n>>\n>> Maybe I'm misunderstanding what you're saying, but why would you want a\n>> file that's become part of the checkout by editing .git/info/sparse to\n>> still be treated as untracked?\n>>\n>> If I have a file on that's excluded via .git/info/sparse then I edit\n>> .git/info/sparse to include it and switch to a branch that doesn't have\n>> that file, I'd expect that file to be deleted from the working copy if\n>> the content matches what's in the repository. If it's modified then I'd\n>> expect the branch switch to fail (like it would without a sparse\n>> checkout).\n>\n> First things first: with sparse checkout, you should not check out\n> _anything_ outside of the focus of the sparse checkout.\n>\n> So I contend that you would only end up with a sparse'd-out file\n> that was formerly tracked if you did something wrong.  That should not\n> happen.\n\nI was thinking if you copied the file there manually and changed\n.git/info/sparse to include it. I would expect git checkout, git\nstatus, etc. to act as just as if I had never excluded it via\n.git/info/sparse, similar to .gitignore and .git/info/exclude.\n\n> The thing is: we need a way to determine quickly and without any\n> ambiguity whether a file is tracked, assumed unchanged, or sparse'd-out\n> (which Nguyễn calls no-checkout).\n>\n> And if we change .git/info/sparse, that state _must not_ change.  We did\n> not touch the file by editing .git/info/sparse, so the state must be\n> unchanged.\n\nI don't know enough to have an opinion on assume-unchanged vs\nno-checkout, but if you edit .git/info/sparse it seems like it should\naffect whether git cares about a file or not. If a file previously had\nthe no-checkout bit and you change .git/info/sparse to include the\nfile, the next time you do something with git, I would expect it to\nstart caring about that path. For example, I can edit .gitignore and\n.git/info/exclude and it notices the next time I use git without\nhaving to do anything special.\n"},{"id":"121035","messageId":"2729632a0908171734p16d6ee7dm5f62848f7625ffbc@mail.gmail.com","threadId":"20539","inReplyTo":"200908180217.35963.jnareb@gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"","fromEmail":"skillzero@gmail.com","sentAt":"2009-08-18T00:34:06Z","receivedAt":"2009-08-18T00:34:06Z","isPatch":true,"sender":{"key":"skillzero@gmail.com","avatar":null},"body":"On Mon, Aug 17, 2009 at 5:17 PM, Jakub Narebski<jnareb@gmail.com> wrote:\n\n> There is e.g. question if \"git grep\" should search \"no-checkout\" files;\n> in the \"assume-unchanged\" case it should, I think, search index version.\n\nI would like it to git grep to not search paths outside the sparse\narea (although --no-sparse would be nice for git grep in case you did\nwant to search everything). The main reason I want sparse checkouts is\nfor performance reasons. For example, git grep can take 10 minutes on\nmy full repository so excluding paths outside the sparse area would\nreduce that to a few seconds.\n"},{"id":"121038","messageId":"200908180249.48404.jnareb@gmail.com","threadId":"20539","inReplyTo":"200908180217.35963.jnareb@gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-08-18T00:49:46Z","receivedAt":"2009-08-18T00:49:46Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jakub Narebski wrote:\n> Johannes Schindelin wrote:\n> \n> > The thing is: we need a way to determine quickly and without any \n> > ambiguity whether a file is tracked, assumed unchanged, or sparse'd-out \n> > (which Nguyễn calls no-checkout).\n> \n> Let's reiterate: \"assume-unchanged\" is about telling git that it should\n> assume for performance reasons that state of file in working directory\n> is the same as state of file in the index.  But, from what was said in\n> this thread, there are situations where git for correctness reasons\n> ignores performance hack.\n> \n> \"no-checkout\" bit is about telling git that the file is not present\n> in working directory, and it has to use version from the index.  Then\n> there is a question if there is file in working area (e.g. from applying\n> patch) which corresponds to a \"no-checkout\" file in index (corresponds\n> because of rename detection).\n\nAlso there is a question if one might want to use them together.  I think\nit is not inconceivable ;-)  One might want for example to limit checkout\nto some subdirectory, but within that directory one might want to use \nassume-unchanged bit, because filesystem performance sucks (FAT, NFS).\nNow couple that with changing in sparse patterns...\n\n-- \nJakub Narebski\nPoland\n"},{"id":"121043","messageId":"fcaeb9bf0908171843x6ab0763dqff7e8aea0443c374@mail.gmail.com","threadId":"20539","inReplyTo":"2729632a0908171734p16d6ee7dm5f62848f7625ffbc@mail.gmail.com","subject":"Re: [RFC PATCH v3 8/8] --sparse for porcelains","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-18T01:43:37Z","receivedAt":"2009-08-18T01:43:37Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Aug 18, 2009 at 7:34 AM, <skillzero@gmail.com> wrote:\n> On Mon, Aug 17, 2009 at 5:17 PM, Jakub Narebski<jnareb@gmail.com> wrote:\n>\n>> There is e.g. question if \"git grep\" should search \"no-checkout\" files;\n>> in the \"assume-unchanged\" case it should, I think, search index version.\n>\n> I would like it to git grep to not search paths outside the sparse\n> area (although --no-sparse would be nice for git grep in case you did\n> want to search everything). The main reason I want sparse checkouts is\n> for performance reasons. For example, git grep can take 10 minutes on\n> my full repository so excluding paths outside the sparse area would\n> reduce that to a few seconds.\n\nThat's a porcelain question that I'd leave it for now. FWIW you can do\nsomething like this:\n\ngit ls-files -v|grep '^H'|cut -c 2-|xargs git grep\n\n/me misses \"cleartool find\"\n-- \nDuy\n"},{"id":"121058","messageId":"200908180825.55289.jnareb@gmail.com","threadId":"20539","inReplyTo":"fcaeb9bf0908171843x6ab0763dqff7e8aea0443c374@mail.gmail.com","subject":"git find (was: [RFC PATCH v3 8/8] --sparse for porcelains)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-08-18T06:25:53Z","receivedAt":"2009-08-18T06:25:53Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, Aug 18, 2009, Nguyen Thai Ngoc Duy wrote:\n> On Tue, Aug 18, 2009 at 7:34 AM, <skillzero@gmail.com> wrote:\n\n> > I would like it to git grep to not search paths outside the sparse\n> > area (although --no-sparse would be nice for git grep in case you did\n> > want to search everything). The main reason I want sparse checkouts is\n> > for performance reasons. For example, git grep can take 10 minutes on\n> > my full repository so excluding paths outside the sparse area would\n> > reduce that to a few seconds.\n> \n> That's a porcelain question that I'd leave it for now. FWIW you can do\n> something like this:\n> \n> git ls-files -v|grep '^H'|cut -c 2-|xargs git grep\n> \n> /me misses \"cleartool find\"\n\nWell, I also think that it would be nice and useful to have \"git find\"\nin addition to current \"git grep\".\n\n-- \nJakub Narebski\nPoland\n"},{"id":"121106","messageId":"fcaeb9bf0908180735s583bfdcajc354723c9faa48@mail.gmail.com","threadId":"20539","inReplyTo":"200908180825.55289.jnareb@gmail.com","subject":"Re: git find (was: [RFC PATCH v3 8/8] --sparse for porcelains)","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-08-18T14:35:05Z","receivedAt":"2009-08-18T14:35:05Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Aug 18, 2009 at 1:25 PM, Jakub Narebski<jnareb@gmail.com> wrote:\n> Well, I also think that it would be nice and useful to have \"git find\"\n> in addition to current \"git grep\".\n\nCan you make a draft on how you want \"git find\" to be? Except the\n\"-exec\" part, Git allows us to search using various commands\n(ls-files, rev-list, log). I don't think a single \"git find\" can cover\nthem all. I was thinking about putting more find-options to search\ncommands we already have. ls-files would support -exec, for example.\n\nA few things that I'd love to have supported:\n - --depth for ls-files (probably all pathspec-as-argument commands)\n - logical combination of search criteria\n - unified blob locator. git-show understands SHA-1:/path/to/blob\nsyntax. What if git-log can output using similar syntax, then feed\nthem to git-grep in order to grep through (across commits)?\n-- \nDuy\n"},{"id":"121110","messageId":"200908181800.42136.jnareb@gmail.com","threadId":"20539","inReplyTo":"fcaeb9bf0908180735s583bfdcajc354723c9faa48@mail.gmail.com","subject":"Re: git find (was: [RFC PATCH v3 8/8] --sparse for porcelains)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-08-18T16:00:38Z","receivedAt":"2009-08-18T16:00:38Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, Aug 18, 2009, Nguyen Thai Ngoc Duy wrote:\n> On Tue, Aug 18, 2009 at 1:25 PM, Jakub Narebski<jnareb@gmail.com> wrote:\n> >\n> > Well, I also think that it would be nice and useful to have \"git find\"\n> > in addition to current \"git grep\".\n> \n> Can you make a draft on how you want \"git find\" to be? Except the\n> \"-exec\" part, Git allows us to search using various commands\n> (ls-files, rev-list, log). I don't think a single \"git find\" can cover\n> them all. I was thinking about putting more find-options to search\n> commands we already have. ls-files would support -exec, for example.\n\nBoth git-rev-list and git-ls-files are plumbing, not porcelain.  Among\ntools / commands you have mentioned only git-log is porcelain.\n\nYou need to process output of git-ls-files if you want to use more\ncomplicated search criteria. \n\n> \n> A few things that I'd love to have supported:\n>  - --depth for ls-files (probably all pathspec-as-argument commands)\n>  - logical combination of search criteria\n>  - unified blob locator. git-show understands SHA-1:/path/to/blob\n> syntax. What if git-log can output using similar syntax, then feed\n> them to git-grep in order to grep through (across commits)?\n\nDraft specification for git-find.  git-find, like git-grep, searches\nthe filesystem dimension, and not time dimension like git-log.\n\ngit-find(1)\n===========\n\nNAME\n----\ngit-find - Search for files in a repository\n\nSYNOPSIS\n--------\n'git find' [--cached] [-z|--null] [(<tree> | <path>)...] [<expression>]\n\nOPTIONS\n-------\n--cached::\n        Instead of searching in the working tree files, check\n        the blobs registered in the index file.\n\nEXPRESSIONS\n-----------\nThe expression is made up of options (which affect overall operation rather\nthan the processing of a specific file, and always return true), tests\n(which return a true or false value), and actions (which have side effects\nand return a true or false value), all separated by operators. `--and`  is\nassumed where the operator is omitted.  If the expression contains no\nactions other than `--prune`, `--print` is performed on all files for which\nthe expression is true.\n\nOPTIONS\n~~~~~~~\n--max-depth <levels>::\n        Descend  at  most levels (a non-negative integer) levels of \n        directories below the command line arguments.   `--max-depth 0`\n        means only apply the tests and actions to the command line \n        arguments.\n\n--min-depth <levels>::\n        Do not apply any tests or actions at levels less than levels \n        (a non-negative integer).  `--min-depth 1` means process all\n        files except the  command line arguments.\n\nTESTS\n~~~~~\n--false::\n        Always false.\n\n--true::\n        Always true.\n\n--name <pattern>::\n--iname <pattern>::\n--path <pattern>::\n--ipath <pattern>::\n        [Entire] Filename matches glob.\n\n--regex <expr>::\n--iregex <expr>::\n        Entire file name matches regular expression.\n\n--lname <pattern>::\n--ilname <pattern>::\n        True if the file is a symbolic link whose contents match glob.\n\n--size <n>[<unit>]::\n        True if the file uses N units of space, rounding up.\n\n--empty::\n        File is empty and is either a regular file or a directory.\n\n--type <C>::\n        True if file is of type C: 'd' for directory, 'f' for regular\n        file, 'l' for symbolic link, 's' for submodule, 'x' for \n        executable regular file (replaces `-perm` from 'find').\n\nACTIONS\n~~~~~~~\n(--exec | --ok) <command> ;\n        Execute command; true if 0 status is returned.\n\n(--execdir | --okdir) <command> ;\n        Like `--exec`, but the specified command is run from the \n        subdirectory containing the matched file.\n\n--print::\n--print0::\n--printf <format>::\n--fprint <file>::\n--fprint0 <file>::\n--fprintf <file> <format>::\n        True; print the full file name.\n\n--prune::\n        True; if the file is a directory, do not descend into it.\n\n--quit::\n        Exit immediately.\n\n\nOPERATORS\n~~~~~~~~~\n--and::\n--or::\n--not::\n( ... )::\n        Specify how multiple expressions are combined using Boolean\n        expressions.  `--and` is the default operator.\n\n-- \nJakub Narebski\nPoland\n"}]}