{"thread":{"id":"7591","subject":"[PATCH 0/6] Initial subproject support (RFC?)","startedAt":"2007-04-10T04:12:50Z","lastAt":"2007-04-15T23:25:00Z","messageCount":98,"participants":["Linus Torvalds","Alex Riesen","Frank Lichtenheld","Andy Parkins","Nicolas Pitre","Josef Weidendorfer","Junio C Hamano","Sam Ravnborg","David Lang","Martin Waitz","David Kågedal","Sam Vilain","Dana How","Torgil Svensson","Brian Gernhardt","Rogan Dawes","J. Bruce Fields"],"isPatch":true,"patchVersion":1,"patchTotal":6},"messages":[{"id":"38984","messageId":"Pine.LNX.4.64.0704092100110.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":null,"subject":"[PATCH 0/6] Initial subproject support (RFC?)","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T04:12:50Z","receivedAt":"2007-04-10T04:12:50Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nOk, the following is a series of six patches that implement some very \nlow-level plumbing for what I consider sane subproject support.\n\nNOTE! I want to make it very clear that this series of patches does not \nmake subprojects \"usable\". They are very core plumbing that allows people \nto think about the issues, and shows how the low-level code could (and in \nmy opinion, should) be done.\n\nSome of the early patches are just cleanups and very basic stuff required \nto actually get to the meat of it all. I actually think that they are all \nin a state where they could be applied, if only because they don't \nactually really *do* anything unless you start generating index files \nentries (and trees) that have the \"gitlink\" entries in them.\n\nI've actually done some testing with a repository that has these kinds of \nsubproject pointers in them, and no, it's really not fully fleshed out \nyet, but yes, I can actually do a commit in one of the subprojects, and \nwhen I do that, the \"raw\" diff literally looks like this:\n\n\t[torvalds@woody superproject]$ git diff --raw\n\t:160000 160000 5813084832d3c680a3436b0253639c94ed55445d 0000000... M    sub-B\n\nand I can do a \"git commit -a\" in the superproject to commit the new \nstate.\n\nNOTE! This series of six patches does not actually contain everything you \nneed to do that - in particular, this series will not actually connect up \nthe magic to make \"git add\" (and thus \"git commit\") actually create the \ngitlink entries for subprojects. That's another (quite small) patch, but I \nhaven't cleaned it up enough to be submittable yet.\n\nI split my original larger patch up into more manageable pieces, so that \nyou should be able to actually just read the patches themselves and get a \nreasonable idea about what it's doing, even *without* actually testing it. \nAnd obviously, \"make test\" still completes happily, if only because none \nof the tests actually trigger any of the new code.\n\nThe patches are all fairly small, and the two first ones are really just \ntotally independent cleanups/fixes:\n\n - diff-lib: use ce_mode_from_stat() rather than messing with modes manually:\n\n\t diff-lib.c |   15 +++------------\n\t 1 files changed, 3 insertions(+), 12 deletions(-)\n\n - Avoid overflowing name buffer in deep directory structures:\n\n\t dir.c |    3 +++\n\t 1 files changed, 3 insertions(+), 0 deletions(-)\n\n - Add 'resolve_gitlink_ref()' helper function:\n\n\t refs.c |   79 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n\t refs.h |    3 ++\n\t 2 files changed, 82 insertions(+), 0 deletions(-)\n\n - Add \"S_IFDIRLNK\" file mode infrastructure for git links:\n\n\t cache.h |   20 +++++++++++++++++++-\n\t 1 files changed, 19 insertions(+), 1 deletions(-)\n\n - Teach \"fsck\" not to follow subproject links:\n\n\t builtin-fsck.c |    9 ++++++++-\n\t tree.c         |   15 ++++++++++++++-\n\t 2 files changed, 22 insertions(+), 2 deletions(-)\n\n - Teach core object handling functions about gitlinks:\n\n\t builtin-ls-tree.c |   20 +++++++++++++++++++-\n\t cache-tree.c      |    2 +-\n\t read-cache.c      |   35 +++++++++++++++++++++++++++++++----\n\t sha1_file.c       |    3 +++\n\t 4 files changed, 54 insertions(+), 6 deletions(-)\n\nand will follow in the next few emails..\n\n\t\t\tLinus\n"},{"id":"38985","messageId":"Pine.LNX.4.64.0704092112540.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092100110.6730@woody.linux-foundation.org","subject":"[PATCH 1/6] diff-lib: use ce_mode_from_stat() rather than messing with modes manually","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T04:13:29Z","receivedAt":"2007-04-10T04:13:29Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nThe diff helpers used to do the magic mode canonicalization and all the\nother special mode handling by hand (\"trust executable bit\" and \"has\nsymlink support\" handling).\n\nThat's bogus. Use \"ce_mode_from_stat()\" that does this all for us.\n\nThis is also going to be required when we add support for links to other\ngit repositories.\n\nSigned-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n---\n diff-lib.c |   15 +++------------\n 1 files changed, 3 insertions(+), 12 deletions(-)\n\ndiff --git a/diff-lib.c b/diff-lib.c\nindex 5c5b05b..c6d1273 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -357,7 +357,7 @@ int run_diff_files(struct rev_info *revs, int silent_on_removed)\n \t\t\t\t\tcontinue;\n \t\t\t}\n \t\t\telse\n-\t\t\t\tdpath->mode = canon_mode(st.st_mode);\n+\t\t\t\tdpath->mode = ntohl(ce_mode_from_stat(ce, st.st_mode));\n \n \t\t\twhile (i < entries) {\n \t\t\t\tstruct cache_entry *nce = active_cache[i];\n@@ -374,8 +374,7 @@ int run_diff_files(struct rev_info *revs, int silent_on_removed)\n \t\t\t\t\tint mode = ntohl(nce->ce_mode);\n \t\t\t\t\tnum_compare_stages++;\n \t\t\t\t\thashcpy(dpath->parent[stage-2].sha1, nce->sha1);\n-\t\t\t\t\tdpath->parent[stage-2].mode =\n-\t\t\t\t\t\tcanon_mode(mode);\n+\t\t\t\t\tdpath->parent[stage-2].mode = ntohl(ce_mode_from_stat(nce, mode));\n \t\t\t\t\tdpath->parent[stage-2].status =\n \t\t\t\t\t\tDIFF_STATUS_MODIFIED;\n \t\t\t\t}\n@@ -424,15 +423,7 @@ int run_diff_files(struct rev_info *revs, int silent_on_removed)\n \t\tif (!changed && !revs->diffopt.find_copies_harder)\n \t\t\tcontinue;\n \t\toldmode = ntohl(ce->ce_mode);\n-\n-\t\tnewmode = canon_mode(st.st_mode);\n-\t\tif (!trust_executable_bit &&\n-\t\t    S_ISREG(newmode) && S_ISREG(oldmode) &&\n-\t\t    ((newmode ^ oldmode) == 0111))\n-\t\t\tnewmode = oldmode;\n-\t\telse if (!has_symlinks &&\n-\t\t    S_ISREG(newmode) && S_ISLNK(oldmode))\n-\t\t\tnewmode = oldmode;\n+\t\tnewmode = ntohl(ce_mode_from_stat(ce, st.st_mode));\n \t\tdiff_change(&revs->diffopt, oldmode, newmode,\n \t\t\t    ce->sha1, (changed ? null_sha1 : ce->sha1),\n \t\t\t    ce->name, NULL);\n-- \n1.5.1.110.g1e4c\n"},{"id":"38986","messageId":"Pine.LNX.4.64.0704092113320.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092100110.6730@woody.linux-foundation.org","subject":"[PATCH 2/6] Avoid overflowing name buffer in deep directory structures","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T04:13:58Z","receivedAt":"2007-04-10T04:13:58Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nThis just makes sure that when we do a read_directory(), we check\nthat the filename fits in the buffer we allocated (with a bit of\nslop)\n\nSigned-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n---\n dir.c |    3 +++\n 1 files changed, 3 insertions(+), 0 deletions(-)\n\ndiff --git a/dir.c b/dir.c\nindex 7426fde..4f5a224 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -353,6 +353,9 @@ static int read_directory_recursive(struct dir_struct *dir, const char *path, co\n \t\t\t     !strcmp(de->d_name + 1, \"git\")))\n \t\t\t\tcontinue;\n \t\t\tlen = strlen(de->d_name);\n+\t\t\t/* Ignore overly long pathnames! */\n+\t\t\tif (len + baselen + 8 > sizeof(fullname))\n+\t\t\t\tcontinue;\n \t\t\tmemcpy(fullname + baselen, de->d_name, len+1);\n \t\t\tif (simplify_away(fullname, baselen + len, simplify))\n \t\t\t\tcontinue;\n-- \n1.5.1.110.g1e4c\n"},{"id":"38991","messageId":"Pine.LNX.4.64.0704092114010.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092100110.6730@woody.linux-foundation.org","subject":"[PATCH 3/6] Add 'resolve_gitlink_ref()' helper function","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T04:14:26Z","receivedAt":"2007-04-10T04:14:26Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nThis new function resolves a ref in *another* git repository.  It's\nnamed for its intended use: to look up the git link to a subproject.\n\nIt's not actually wired up to anything yet, but we're getting closer to\nhaving fundamental plumbing support for \"links\" from one git directory\nto another, which is the basis of subproject support.\n\nSigned-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n---\n refs.c |   79 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n refs.h |    3 ++\n 2 files changed, 82 insertions(+), 0 deletions(-)\n\ndiff --git a/refs.c b/refs.c\nindex d2b7b7f..229da74 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -215,6 +215,85 @@ static struct ref_list *get_loose_refs(void)\n \n /* We allow \"recursive\" symbolic refs. Only within reason, though */\n #define MAXDEPTH 5\n+#define MAXREFLEN (1024)\n+\n+static int resolve_gitlink_packed_ref(char *name, int pathlen, const char *refname, unsigned char *result)\n+{\n+\tFILE *f;\n+\tstruct cached_refs refs;\n+\tstruct ref_list *ref;\n+\tint retval;\n+\n+\tstrcpy(name + pathlen, \"packed-refs\");\n+\tf = fopen(name, \"r\");\n+\tif (!f)\n+\t\treturn -1;\n+\tread_packed_refs(f, &refs);\n+\tref = refs.packed;\n+\tretval = -1;\n+\twhile (ref) {\n+\t\tif (!strcmp(ref->name, refname)) {\n+\t\t\tretval = 0;\n+\t\t\tmemcpy(result, ref->sha1, 20);\n+\t\t\tbreak;\n+\t\t}\n+\t\tref = ref->next;\n+\t}\n+\tfree_ref_list(refs.packed);\n+\treturn retval;\n+}\n+\n+static int resolve_gitlink_ref_recursive(char *name, int pathlen, const char *refname, unsigned char *result, int recursion)\n+{\n+\tint fd, len = strlen(refname);\n+\tchar buffer[128], *p;\n+\n+\tif (recursion > MAXDEPTH || len > MAXREFLEN)\n+\t\treturn -1;\n+\tmemcpy(name + pathlen, refname, len+1);\n+\tfd = open(name, O_RDONLY);\n+\tif (fd < 0)\n+\t\treturn resolve_gitlink_packed_ref(name, pathlen, refname, result);\n+\n+\tlen = read(fd, buffer, sizeof(buffer)-1);\n+\tclose(fd);\n+\tif (len < 0)\n+\t\treturn -1;\n+\twhile (len && isspace(buffer[len-1]))\n+\t\tlen--;\n+\tbuffer[len] = 0;\n+\n+\t/* Was it a detached head or an old-fashioned symlink? */\n+\tif (!get_sha1_hex(buffer, result))\n+\t\treturn 0;\n+\n+\t/* Symref? */\n+\tif (strncmp(buffer, \"ref:\", 4))\n+\t\treturn -1;\n+\tp = buffer + 4;\n+\twhile (isspace(*p))\n+\t\tp++;\n+\n+\treturn resolve_gitlink_ref_recursive(name, pathlen, p, result, recursion+1);\n+}\n+\n+int resolve_gitlink_ref(const char *path, const char *refname, unsigned char *result)\n+{\n+\tint len = strlen(path), retval;\n+\tchar *gitdir;\n+\n+\twhile (len && path[len-1] == '/')\n+\t\tlen--;\n+\tif (!len)\n+\t\treturn -1;\n+\tgitdir = xmalloc(len + MAXREFLEN + 8);\n+\tmemcpy(gitdir, path, len);\n+\tmemcpy(gitdir + len, \"/.git/\", 7);\n+\n+\tretval = resolve_gitlink_ref_recursive(gitdir, len+6, refname, result, 0);\n+\tfree(gitdir);\n+\treturn retval;\n+}\n \n const char *resolve_ref(const char *ref, unsigned char *sha1, int reading, int *flag)\n {\ndiff --git a/refs.h b/refs.h\nindex acedffc..f61f6d9 100644\n--- a/refs.h\n+++ b/refs.h\n@@ -60,4 +60,7 @@ extern int check_ref_format(const char *target);\n /** rename ref, return 0 on success **/\n extern int rename_ref(const char *oldref, const char *newref, const char *logmsg);\n \n+/** resolve ref in nested \"gitlink\" repository */\n+extern int resolve_gitlink_ref(const char *name, const char *refname, unsigned char *result);\n+\n #endif /* REFS_H */\n-- \n1.5.1.110.g1e4c\n"},{"id":"38990","messageId":"Pine.LNX.4.64.0704092114300.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092100110.6730@woody.linux-foundation.org","subject":"[PATCH 4/6] Add \"S_IFDIRLNK\" file mode infrastructure for git links","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T04:14:58Z","receivedAt":"2007-04-10T04:14:58Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nThis just adds the basic helper functions to recognize and work with git\ntree entries that are links to other git repositories (\"subprojects\").\nThey still aren't actually connected up to any of the code-paths, but\nnow all the infrastructure is in place.\n\nThe next commit will start actually adding actual subproject support.\n\nSigned-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n---\n cache.h |   20 +++++++++++++++++++-\n 1 files changed, 19 insertions(+), 1 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex eb57507..1b3d00e 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -25,6 +25,22 @@\n #endif\n \n /*\n+ * A \"directory link\" is a link to another git directory.\n+ *\n+ * The value 0160000 is not normally a valid mode, and\n+ * also just happens to be S_IFDIR + S_IFLNK\n+ *\n+ * NOTE! We *really* shouldn't depend on the S_IFxxx macros\n+ * always having the same values everywhere. We should use\n+ * our internal git values for these things, and then we can\n+ * translate that to the OS-specific value. It just so\n+ * happens that everybody shares the same bit representation\n+ * in the UNIX world (and apparently wider too..)\n+ */\n+#define S_IFDIRLNK\t0160000\n+#define S_ISDIRLNK(m)\t(((m) & S_IFMT) == S_IFDIRLNK)\n+\n+/*\n  * Intensive research over the course of many years has shown that\n  * port 9418 is totally unused by anything else. Or\n  *\n@@ -104,6 +120,8 @@ static inline unsigned int create_ce_mode(unsigned int mode)\n {\n \tif (S_ISLNK(mode))\n \t\treturn htonl(S_IFLNK);\n+\tif (S_ISDIR(mode) || S_ISDIRLNK(mode))\n+\t\treturn htonl(S_IFDIRLNK);\n \treturn htonl(S_IFREG | ce_permissions(mode));\n }\n static inline unsigned int ce_mode_from_stat(struct cache_entry *ce, unsigned int mode)\n@@ -121,7 +139,7 @@ static inline unsigned int ce_mode_from_stat(struct cache_entry *ce, unsigned in\n }\n #define canon_mode(mode) \\\n \t(S_ISREG(mode) ? (S_IFREG | ce_permissions(mode)) : \\\n-\tS_ISLNK(mode) ? S_IFLNK : S_IFDIR)\n+\tS_ISLNK(mode) ? S_IFLNK : S_ISDIR(mode) ? S_IFDIR : S_IFDIRLNK)\n \n #define cache_entry_size(len) ((offsetof(struct cache_entry,name) + (len) + 8) & ~7)\n \n-- \n1.5.1.110.g1e4c\n"},{"id":"38988","messageId":"Pine.LNX.4.64.0704092115020.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092100110.6730@woody.linux-foundation.org","subject":"[PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T04:15:29Z","receivedAt":"2007-04-10T04:15:29Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nSince the subprojects don't necessarily even exist in the current tree,\nmuch less in the current git repository (they are totally independent\nrepositories), we do not want to try to follow the chain from one git\nrepository to another through a gitlink.\n\nThis involves teaching fsck to ignore references to gitlink objects from\na tree and from the current index.\n\nSigned-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n---\n builtin-fsck.c |    9 ++++++++-\n tree.c         |   15 ++++++++++++++-\n 2 files changed, 22 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin-fsck.c b/builtin-fsck.c\nindex 4d8b66c..f22de8d 100644\n--- a/builtin-fsck.c\n+++ b/builtin-fsck.c\n@@ -253,6 +253,7 @@ static int fsck_tree(struct tree *item)\n \t\tcase S_IFREG | 0644:\n \t\tcase S_IFLNK:\n \t\tcase S_IFDIR:\n+\t\tcase S_IFDIRLNK:\n \t\t\tbreak;\n \t\t/*\n \t\t * This is nonstandard, but we had a few of these\n@@ -695,8 +696,14 @@ int cmd_fsck(int argc, char **argv, const char *prefix)\n \t\tint i;\n \t\tread_cache();\n \t\tfor (i = 0; i < active_nr; i++) {\n-\t\t\tstruct blob *blob = lookup_blob(active_cache[i]->sha1);\n+\t\t\tunsigned int mode;\n+\t\t\tstruct blob *blob;\n \t\t\tstruct object *obj;\n+\n+\t\t\tmode = ntohl(active_cache[i]->ce_mode);\n+\t\t\tif (S_ISDIRLNK(mode))\n+\t\t\t\tcontinue;\n+\t\t\tblob = lookup_blob(active_cache[i]->sha1);\n \t\t\tif (!blob)\n \t\t\t\tcontinue;\n \t\t\tobj = &blob->object;\ndiff --git a/tree.c b/tree.c\nindex d188c0f..dbb63fc 100644\n--- a/tree.c\n+++ b/tree.c\n@@ -143,6 +143,14 @@ struct tree *lookup_tree(const unsigned char *sha1)\n \treturn (struct tree *) obj;\n }\n \n+/*\n+ * NOTE! Tree refs to external git repositories\n+ * (ie gitlinks) do not count as real references.\n+ *\n+ * You don't have to have those repositories\n+ * available at all, much less have the objects\n+ * accessible from the current repository.\n+ */\n static void track_tree_refs(struct tree *item)\n {\n \tint n_refs = 0, i;\n@@ -152,8 +160,11 @@ static void track_tree_refs(struct tree *item)\n \n \t/* Count how many entries there are.. */\n \tinit_tree_desc(&desc, item->buffer, item->size);\n-\twhile (tree_entry(&desc, &entry))\n+\twhile (tree_entry(&desc, &entry)) {\n+\t\tif (S_ISDIRLNK(entry.mode))\n+\t\t\tcontinue;\n \t\tn_refs++;\n+\t}\n \n \t/* Allocate object refs and walk it again.. */\n \ti = 0;\n@@ -162,6 +173,8 @@ static void track_tree_refs(struct tree *item)\n \twhile (tree_entry(&desc, &entry)) {\n \t\tstruct object *obj;\n \n+\t\tif (S_ISDIRLNK(entry.mode))\n+\t\t\tcontinue;\n \t\tif (S_ISDIR(entry.mode))\n \t\t\tobj = &lookup_tree(entry.sha1)->object;\n \t\telse\n-- \n1.5.1.110.g1e4c\n"},{"id":"38989","messageId":"Pine.LNX.4.64.0704092115350.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092100110.6730@woody.linux-foundation.org","subject":"[PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T04:20:29Z","receivedAt":"2007-04-10T04:20:29Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nThis teaches the really fundamental core SHA1 object handling routines\nabout gitlinks.  We can compare trees with gitlinks in them (although we\ncan not actually generate patches for them yet - just raw git diffs),\nand they show up as commits in \"git ls-tree\".\n\nWe also know to compare gitlinks as if they were directories (ie the\nnormal \"sort as trees\" rules apply).\n\nSigned-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n---\n\nOk, that's it for now.\n\nNOTE NOTE NOTE! I'd like to note once more that this doesn't actually get \nyou working subproject support. Not only do I need to connect up a few \nmore low-level helper functions (things like \"git diff\" don't know how to \ngenerate even rudimentary \"subproject X changed\" patches, nor can you \nactually yet *add* subprojects), but quite apart from that low-level \nstuff, anything more high-level (like \"git fetch\" and friends) will need \nto know about subprojects.\n\nIn general, think of this like the early git plumbing: it's the early\n\"content-addressable filesystem\" part. The actual SCM parts going on top \nof it are yet to be done.\n\nI'm hoping/expecting that there are more people who have the ability and \nthe interest to work on the higher-level interfaces once the core plumbing \nsupport is there. There's still some plumbing to be done, but after that, \nmaybe more people (and maybe the SoC people) can start filling out the \nhigher-level details..\n\nComments on the patches/approach so far?\n\n builtin-ls-tree.c |   20 +++++++++++++++++++-\n cache-tree.c      |    2 +-\n read-cache.c      |   35 +++++++++++++++++++++++++++++++----\n sha1_file.c       |    3 +++\n 4 files changed, 54 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin-ls-tree.c b/builtin-ls-tree.c\nindex 6472610..1cb4dca 100644\n--- a/builtin-ls-tree.c\n+++ b/builtin-ls-tree.c\n@@ -6,6 +6,7 @@\n #include \"cache.h\"\n #include \"blob.h\"\n #include \"tree.h\"\n+#include \"commit.h\"\n #include \"quote.h\"\n #include \"builtin.h\"\n \n@@ -59,7 +60,24 @@ static int show_tree(const unsigned char *sha1, const char *base, int baselen,\n \tint retval = 0;\n \tconst char *type = blob_type;\n \n-\tif (S_ISDIR(mode)) {\n+\tif (S_ISDIRLNK(mode)) {\n+\t\t/*\n+\t\t * Maybe we want to have some recursive version here?\n+\t\t *\n+\t\t * Something like:\n+\t\t *\n+\t\tif (show_subprojects(base, baselen, pathname)) {\n+\t\t\tif (fork()) {\n+\t\t\t\tchdir(base);\n+\t\t\t\texec ls-tree;\n+\t\t\t}\n+\t\t\twaitpid();\n+\t\t}\n+\t\t *\n+\t\t * ..or similar..\n+\t\t */\n+\t\ttype = commit_type;\n+\t} else if (S_ISDIR(mode)) {\n \t\tif (show_recursive(base, baselen, pathname)) {\n \t\t\tretval = READ_TREE_RECURSIVE;\n \t\t\tif (!(ls_options & LS_SHOW_TREES))\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 9b73c86..6369cc7 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -326,7 +326,7 @@ static int update_one(struct cache_tree *it,\n \t\t\tmode = ntohl(ce->ce_mode);\n \t\t\tentlen = pathlen - baselen;\n \t\t}\n-\t\tif (!missing_ok && !has_sha1_file(sha1))\n+\t\tif (mode != S_IFDIRLNK && !missing_ok && !has_sha1_file(sha1))\n \t\t\treturn error(\"invalid object %s\", sha1_to_hex(sha1));\n \n \t\tif (!ce->ce_mode)\ndiff --git a/read-cache.c b/read-cache.c\nindex 54573ce..8fe94cd 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -5,6 +5,7 @@\n  */\n #include \"cache.h\"\n #include \"cache-tree.h\"\n+#include \"refs.h\"\n \n /* Index extensions.\n  *\n@@ -91,6 +92,23 @@ static int ce_compare_link(struct cache_entry *ce, size_t expected_size)\n \treturn match;\n }\n \n+static int ce_compare_gitlink(struct cache_entry *ce)\n+{\n+\tunsigned char sha1[20];\n+\n+\t/*\n+\t * We don't actually require that the .git directory\n+\t * under DIRLNK directory be a valid git directory. It\n+\t * might even be missing (in case nobody populated that\n+\t * sub-project).\n+\t *\n+\t * If so, we consider it always to match.\n+\t */\n+\tif (resolve_gitlink_ref(ce->name, \"HEAD\", sha1) < 0)\n+\t\treturn 0;\n+\treturn hashcmp(sha1, ce->sha1);\n+}\n+\n static int ce_modified_check_fs(struct cache_entry *ce, struct stat *st)\n {\n \tswitch (st->st_mode & S_IFMT) {\n@@ -102,6 +120,9 @@ static int ce_modified_check_fs(struct cache_entry *ce, struct stat *st)\n \t\tif (ce_compare_link(ce, xsize_t(st->st_size)))\n \t\t\treturn DATA_CHANGED;\n \t\tbreak;\n+\tcase S_IFDIRLNK:\n+\t\t/* No need to do anything, we did the exact compare in \"match_stat_basic\" */\n+\t\tbreak;\n \tdefault:\n \t\treturn TYPE_CHANGED;\n \t}\n@@ -127,6 +148,12 @@ static int ce_match_stat_basic(struct cache_entry *ce, struct stat *st)\n \t\t    (has_symlinks || !S_ISREG(st->st_mode)))\n \t\t\tchanged |= TYPE_CHANGED;\n \t\tbreak;\n+\tcase S_IFDIRLNK:\n+\t\tif (!S_ISDIR(st->st_mode))\n+\t\t\tchanged |= TYPE_CHANGED;\n+\t\telse if (ce_compare_gitlink(ce))\n+\t\t\tchanged |= DATA_CHANGED;\n+\t\tbreak;\n \tdefault:\n \t\tdie(\"internal error: ce_mode is %o\", ntohl(ce->ce_mode));\n \t}\n@@ -250,9 +277,9 @@ int base_name_compare(const char *name1, int len1, int mode1,\n \t\treturn cmp;\n \tc1 = name1[len];\n \tc2 = name2[len];\n-\tif (!c1 && S_ISDIR(mode1))\n+\tif (!c1 && (S_ISDIR(mode1) || S_ISDIRLNK(mode1)))\n \t\tc1 = '/';\n-\tif (!c2 && S_ISDIR(mode2))\n+\tif (!c2 && (S_ISDIR(mode2) || S_ISDIRLNK(mode1)))\n \t\tc2 = '/';\n \treturn (c1 < c2) ? -1 : (c1 > c2) ? 1 : 0;\n }\n@@ -334,8 +361,8 @@ int add_file_to_cache(const char *path, int verbose)\n \tif (lstat(path, &st))\n \t\tdie(\"%s: unable to stat (%s)\", path, strerror(errno));\n \n-\tif (!S_ISREG(st.st_mode) && !S_ISLNK(st.st_mode))\n-\t\tdie(\"%s: can only add regular files or symbolic links\", path);\n+\tif (!S_ISREG(st.st_mode) && !S_ISLNK(st.st_mode) && !S_ISDIR(st.st_mode))\n+\t\tdie(\"%s: can only add regular files, symbolic links or git-directories\", path);\n \n \tnamelen = strlen(path);\n \tsize = cache_entry_size(namelen);\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 4304fe9..ab915fa 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -13,6 +13,7 @@\n #include \"commit.h\"\n #include \"tag.h\"\n #include \"tree.h\"\n+#include \"refs.h\"\n \n #ifndef O_NOATIME\n #if defined(__linux__) && (defined(__i386__) || defined(__PPC__))\n@@ -2332,6 +2333,8 @@ int index_path(unsigned char *sha1, const char *path, struct stat *st, int write\n \t\t\t\t     path);\n \t\tfree(target);\n \t\tbreak;\n+\tcase S_IFDIR:\n+\t\treturn resolve_gitlink_ref(path, \"HEAD\", sha1);\n \tdefault:\n \t\treturn error(\"%s: unsupported file type\", path);\n \t}\n-- \n1.5.1.110.g1e4c\n"},{"id":"38992","messageId":"Pine.LNX.4.64.0704092133550.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092100110.6730@woody.linux-foundation.org","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T04:46:12Z","receivedAt":"2007-04-10T04:46:12Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 9 Apr 2007, Linus Torvalds wrote:\n> \n> NOTE! This series of six patches does not actually contain everything you \n> need to do that - in particular, this series will not actually connect up \n> the magic to make \"git add\" (and thus \"git commit\") actually create the \n> gitlink entries for subprojects. That's another (quite small) patch, but I \n> haven't cleaned it up enough to be submittable yet.\n\nHere is, for your enjoyment, the last patch I used to actually test this \nall. I do *not* submit it as a patch for actual inclusion - the other \npatches in the series are, I think, ready to actually be merged. This one \nis not.\n\nIt's broken for a few reasons:\n\n - it allows you to do \"git add subproject\" to add the subproject to the \n   index (and then use \"git commit\" to commit it), but even something as \n   simple as \"git commit -a\" doesn't work right, because the sequence that \n   \"git commit -a\" uses to update the index doesn't work with the current \n   state of the plumbing (ie the\n\n\tgit-diff-files --name-only -z |\n\t\tgit-update-index --remove -z --stdin\n\n   thing doesn't work right.\n\n - even for \"git add\", the logic isn't really right. It should take the \n   old index state into account to decide if it wants to add it as a \n   subproject. \n\nso this patch really isn't very good, but it allows people who are \ninterested to perhaps actually test something. For example, my test repo \nwas actually created with this:\n\n\t[torvalds@woody superproject]$ git log --raw\n\tcommit 649ad968bdd79cb3b0f50feb819b7e9b134d3a1a\n\tAuthor: Linus Torvalds <torvalds@woody.linux-foundation.org>\n\tDate:   Mon Apr 9 21:36:53 2007 -0700\n\t\n\t    This commits the modification to sub-project B\n\t\n\t:160000 160000 5813084832d3c680a3436b0253639c94ed55445d 17d246a35f27a46762328281eb6e9d4558f91e9d M      sub-B\n\n\tcommit f3c55ffcc000a8c0fecc6801e8909d084e3d419e\n\tAuthor: Linus Torvalds <torvalds@woody.linux-foundation.org>\n\tDate:   Mon Apr 9 16:12:29 2007 -0700\n\t\n\t    Superproject with two subprojects\n\t\n\t:000000 160000 0000000... c0daf4c85d48879ab450a6a887bbb241eb0de00a A    sub-A\n\t:000000 160000 0000000... 5813084832d3c680a3436b0253639c94ed55445d A    sub-B\n\n\tcommit 45eb14edb43b10e3d3ac7a495a1ec861e85dc36f\n\tAuthor: Linus Torvalds <torvalds@woody.linux-foundation.org>\n\tDate:   Mon Apr 9 15:36:24 2007 -0700\n\t\n\t    Add top-level Makefile for super-project\n\t\n\t:000000 100644 0000000... 57e8394... A  Makefile\n\nso you can see how things look at a low level (ie a \"gitlink\" is just a \ntree entry with mode 0160000, and the SHA1 is just the SHA1 of the HEAD \ncommit in the subproject)\n\n\t\tLinus\n\n---\ndiff --git a/dir.c b/dir.c\nindex 4f5a224..ef284a2 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -378,6 +378,14 @@ static int read_directory_recursive(struct dir_struct *dir, const char *path, co\n \t\t\t\t\tcontinue;\n \t\t\t\t/* fallthrough */\n \t\t\tcase DT_DIR:\n+\t\t\t\t/* Does it have a git directory? If so, it's a DIRLNK */\n+\t\t\t\tif (!dir->no_dirlinks) {\n+\t\t\t\t\tmemcpy(fullname + baselen + len, \"/.git/\", 7);\n+\t\t\t\t\tif (!stat(fullname, &st)) {\n+\t\t\t\t\t\tif (S_ISDIR(st.st_mode))\n+\t\t\t\t\t\t\tbreak;\n+\t\t\t\t\t}\n+\t\t\t\t}\n \t\t\t\tmemcpy(fullname + baselen + len, \"/\", 2);\n \t\t\t\tlen++;\n \t\t\t\tif (dir->show_other_directories &&\ndiff --git a/dir.h b/dir.h\nindex 33c31f2..1931609 100644\n--- a/dir.h\n+++ b/dir.h\n@@ -33,7 +33,8 @@ struct dir_struct {\n \tint nr, alloc;\n \tunsigned int show_ignored:1,\n \t\t     show_other_directories:1,\n-\t\t     hide_empty_directories:1;\n+\t\t     hide_empty_directories:1,\n+\t\t     no_dirlinks;\n \tstruct dir_entry **entries;\n \n \t/* Exclude info */\n"},{"id":"38996","messageId":"20070410084022.GB2813@planck.djpig.de","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092115350.6730@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Frank Lichtenheld","fromEmail":"frank@lichtenheld.de","sentAt":"2007-04-10T08:40:23Z","receivedAt":"2007-04-10T08:40:23Z","isPatch":true,"sender":{"key":"frank@lichtenheld.de","avatar":"https://gravatar.com/avatar/b9f1d4b120e138f157c9e480d0818197c474628923786adb98f30017cdb99c3c?d=mp&s=160"},"body":"On Mon, Apr 09, 2007 at 09:20:29PM -0700, Linus Torvalds wrote:\n> diff --git a/sha1_file.c b/sha1_file.c\n> index 4304fe9..ab915fa 100644\n> --- a/sha1_file.c\n> +++ b/sha1_file.c\n> @@ -13,6 +13,7 @@\n>  #include \"commit.h\"\n>  #include \"tag.h\"\n>  #include \"tree.h\"\n> +#include \"refs.h\"\n>  \n>  #ifndef O_NOATIME\n>  #if defined(__linux__) && (defined(__i386__) || defined(__PPC__))\n> @@ -2332,6 +2333,8 @@ int index_path(unsigned char *sha1, const char *path, struct stat *st, int write\n>  \t\t\t\t     path);\n>  \t\tfree(target);\n>  \t\tbreak;\n> +\tcase S_IFDIR:\n> +\t\treturn resolve_gitlink_ref(path, \"HEAD\", sha1);\n>  \tdefault:\n>  \t\treturn error(\"%s: unsupported file type\", path);\n>  \t}\n\nNot that I have time right now to look up the exact context (only read\nthe patch), but I would've expected a \"case S_IFDIRLNK:\" here?\n\nGruesse,\n-- \nFrank Lichtenheld <frank@lichtenheld.de>\nwww: http://www.djpig.de/\n"},{"id":"38995","messageId":"81b0412b0704100238l38ad3765w6c06878e2db654a7@mail.gmail.com","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092114010.6730@woody.linux-foundation.org","subject":"Re: [PATCH 3/6] Add 'resolve_gitlink_ref()' helper function","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-04-10T09:38:29Z","receivedAt":"2007-04-10T09:38:29Z","isPatch":true,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/10/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> +int resolve_gitlink_ref(const char *path, const char *refname, unsigned char *result)\n> +{\n> +       int len = strlen(path), retval;\n> +       char *gitdir;\n> +\n> +       while (len && path[len-1] == '/')\n> +               len--;\n> +       if (!len)\n> +               return -1;\n> +       gitdir = xmalloc(len + MAXREFLEN + 8);\n> +       memcpy(gitdir, path, len);\n> +       memcpy(gitdir + len, \"/.git/\", 7);\n\nCan't a subproject be bare?\n"},{"id":"39007","messageId":"81b0412b0704100431i1e58c74tbe19ea490e470b8b@mail.gmail.com","threadId":"7591","inReplyTo":"20070410084022.GB2813@planck.djpig.de","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-04-10T11:31:40Z","receivedAt":"2007-04-10T11:31:40Z","isPatch":true,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/10/07, Frank Lichtenheld <frank@lichtenheld.de> wrote:\n> On Mon, Apr 09, 2007 at 09:20:29PM -0700, Linus Torvalds wrote:\n> > +     case S_IFDIR:\n> > +             return resolve_gitlink_ref(path, \"HEAD\", sha1);\n> >       default:\n> >               return error(\"%s: unsupported file type\", path);\n> >       }\n>\n> Not that I have time right now to look up the exact context (only read\n> the patch), but I would've expected a \"case S_IFDIRLNK:\" here?\n>\n\nNo, the st_mode comes directly from file system. It knows nothing about\ndirlinks.\n"},{"id":"39008","messageId":"81b0412b0704100604x2841d96aq194d3dedd303c588@mail.gmail.com","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092133550.6730@woody.linux-foundation.org","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-04-10T13:04:33Z","receivedAt":"2007-04-10T13:04:33Z","isPatch":true,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/10/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> Here is, for your enjoyment, the last patch I used to actually test this\n> all. I do *not* submit it as a patch for actual inclusion - the other\n> patches in the series are, I think, ready to actually be merged. This one\n> is not.\n>\n> It's broken for a few reasons:\n>\n>  - it allows you to do \"git add subproject\" to add the subproject to the\n>    index (and then use \"git commit\" to commit it), but even something as\n>    simple as \"git commit -a\" doesn't work right, because the sequence that\n>    \"git commit -a\" uses to update the index doesn't work with the current\n>    state of the plumbing (ie the\n>\n>         git-diff-files --name-only -z |\n>                 git-update-index --remove -z --stdin\n>\n>    thing doesn't work right.\n>\n>  - even for \"git add\", the logic isn't really right. It should take the\n>    old index state into account to decide if it wants to add it as a\n>    subproject.\n>\n\nThe other thing which will be missed a lot (I miss it that much)\nis a subproject-recursive git-commit and git-status.\nIt is very possible that the default should be different for\nthe git-commit and git-status: git-commit is likely to have it\noff whereas git-status will very much depend on how fast\nthe usual response is (or wished for). An integrator on very fast\nmachine may like it on for both, a subproject developer can have\nit off for both (to avoid accidental commits and generally being\nnot interested in anything besides his code), an occasional person\ncan have the status defaulting to on and commit to off - to avoid\naccidental commits in subprojects which are just tracked.\n\nA separate config option and a command-line switch, probably.\n"},{"id":"39024","messageId":"Pine.LNX.4.64.0704100750290.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"20070410084022.GB2813@planck.djpig.de","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T14:55:58Z","receivedAt":"2007-04-10T14:55:58Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Apr 2007, Frank Lichtenheld wrote:\n\n> On Mon, Apr 09, 2007 at 09:20:29PM -0700, Linus Torvalds wrote:\n> > @@ -2332,6 +2333,8 @@ int index_path(unsigned char *sha1, const char *path, struct stat *st, int write\n> >  \t\t\t\t     path);\n> >  \t\tfree(target);\n> >  \t\tbreak;\n> > +\tcase S_IFDIR:\n> > +\t\treturn resolve_gitlink_ref(path, \"HEAD\", sha1);\n> >  \tdefault:\n> >  \t\treturn error(\"%s: unsupported file type\", path);\n> >  \t}\n> \n> Not that I have time right now to look up the exact context (only read\n> the patch), but I would've expected a \"case S_IFDIRLNK:\" here?\n\nSo we have this strange (and worrying) dualism inside git: we use the same \nmacros *both* for \"stat data\" *and* for \"git-internal file modes\".\n\nSo sometimes a mode is the result of a [l]stat() call like above, and then \na gitlink is just a directory and we use S_IFDIR. And if it comes from the \nindex, then it uses the internal git representation, and is S_IFDIRLNK.\n\nI'm not very happy about it, but I'm actually most unhappy about it since \nI could imagine that the constants themselves are different on different \nOS's (eg VMS - a Unix-related OS will use the same constants for \nhistorical reasons).\n\nIn this particular place (index-path), we obviously not only have a stat() \nresult, but more importantly, we never come here for a \"normal\" directory, \nsince a normal directory would have been expanded into its component paths \nby the \"read_directory()\" logic.\n\nSo that interaction with directory expansion is somewhat non-obvious: \nnormal directories are expanded recursively into the files they contain, \nwhile git directories end up being visible to internals as real \ndirectories, and are turned into gitlinks by code like the above.\n\n\t\tLinus\n"},{"id":"39022","messageId":"Pine.LNX.4.64.0704100756060.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"81b0412b0704100238l38ad3765w6c06878e2db654a7@mail.gmail.com","subject":"Re: [PATCH 3/6] Add 'resolve_gitlink_ref()' helper function","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T14:58:27Z","receivedAt":"2007-04-10T14:58:27Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Apr 2007, Alex Riesen wrote:\n>\n> On 4/10/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> > +int resolve_gitlink_ref(const char *path, const char *refname, unsigned\n> > char *result)\n> > +{\n> > +       int len = strlen(path), retval;\n> > +       char *gitdir;\n> > +\n> > +       while (len && path[len-1] == '/')\n> > +               len--;\n> > +       if (!len)\n> > +               return -1;\n> > +       gitdir = xmalloc(len + MAXREFLEN + 8);\n> > +       memcpy(gitdir, path, len);\n> > +       memcpy(gitdir + len, \"/.git/\", 7);\n> \n> Can't a subproject be bare?\n\nNot when it is checked out, no. That's what \"checked out\" means ;)\n\nIf a subproject is bare, it never gets resolved, because it's never \nchecked out in a superproject.\n\nSo a subproject *can* be bare, but when it's bare it is just a totally \nregular independent git project, simply by *definition* of not being \nchecked out inside a superproject.\n\nBut hey, that was just a design decision of mine, and if people can argue \nfor it being wrong, I don't think I'm married to it ;)\n\n\t\tLinus\n"},{"id":"39031","messageId":"Pine.LNX.4.64.0704100758430.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"81b0412b0704100604x2841d96aq194d3dedd303c588@mail.gmail.com","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T15:13:58Z","receivedAt":"2007-04-10T15:13:58Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Apr 2007, Alex Riesen wrote:\n> \n> The other thing which will be missed a lot (I miss it that much)\n> is a subproject-recursive git-commit and git-status.\n\nNote that I was definitely planning on adding them too, but they are at a \nhigher level. \n\nSo the long-term plan is/was to add a flag to \"git diff\" (and \"git \nls-tree\" etc) to say \"recurse into subprojects\".\n\nYou cound perhaps even make that flag the default with some .git/config \noption, if your superproject is small enough.\n\nBut this series of 6 (and the seventh ugly hack) is literally meant for \njust the really core object-handling stuff, and even there it's not really \ncomplete.\n\nFor example, you cannot even clone a superproject yet, simply because \ngit-upload-pack doesn't know that it's not supposed to follow the gitlink \nthings etc. So there's a lot of details left even for the really *core* \nstuff, but I wanted to post the series of six patches because those six \npatches are actually enough to reach the point where you can start looking \nat individual problems (like \"git upload-pack\") and fix them \nincrementally.\n\nSo I'd like this to be merged somewhere, not because \"it works\" or \"it's \ncomplete\", but because it's in a shape where I think a lot of people can \nstart fixing small details. \n\nFor example, with just two smallish updates:\n - teach \"git upload-pack\" not to try to follow gitlinks\n - teach \"git read-tree\" to check out a git-link as just an empty \n   subdirectory\nyou should already be pretty close to being able to clone a superproject. \nYou'd still have to clone the subprojects one-by-one manually, and that \nwould be more of a porcelain'ish issue to teach git clone to fetch \nsubmodules too (with some \".gitmodules\" file that contains the rules for \nthat!)\n\nBut no, I didn't do any of that. I literally did just the \"tree object \nformat change\" to support the *notion* of gitlinks - not all the pieces to \nthen actually *implement* the notion are done by a long shot.\n\nI think everybody agrees that we need some kind of subproject support, and \nthe KDE repository certainly shows that subprojects need to be truly \nindependent (because if they aren't, you end up with all the scaling \nissues that we see now - including something as simple as just \"fsck\" \ntaking way way too long unless you have 4GB of RAM or more), and this sets \nthe basic rules for that.\n\nBut they really are pretty low-level rules. For example, to go back to the \nKDE thing: we'd also need to teach *importers* to import certain \nsubdirectories as submodules (or have a git->git translator that turns a \nsubdirectory into a separate submodule).\n\nSo those are examples of things that obviously need to be done, and that \nmy patches do not address in *any* way. They are really low-level plumbing \nsupport, kind of like the old original days when you had to run \n\n\tgit-update-index ...\n\ttree=$(git-write-tree)\n\tcommit=$(git-commit-tree -p $parent $tree <$msgfile)\n\nby hand. A few monts later it was \"git commit -a\", but it started out \nwith just fairly low-level plumbing..\n\n\t\tLinus\n"},{"id":"39019","messageId":"81b0412b0704100835vbbfe8e7o2df2f121ce088589@mail.gmail.com","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704100756060.6730@woody.linux-foundation.org","subject":"Re: [PATCH 3/6] Add 'resolve_gitlink_ref()' helper function","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-04-10T15:35:09Z","receivedAt":"2007-04-10T15:35:09Z","isPatch":true,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/10/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> >\n> > Can't a subproject be bare?\n>\n> Not when it is checked out, no. That's what \"checked out\" means ;)\n>\n> If a subproject is bare, it never gets resolved, because it's never\n> checked out in a superproject.\n>\n> So a subproject *can* be bare, but when it's bare it is just a totally\n> regular independent git project, simply by *definition* of not being\n> checked out inside a superproject.\n>\n> But hey, that was just a design decision of mine, and if people can argue\n> for it being wrong, I don't think I'm married to it ;)\n\nI didn't actually had a use case in mind as I asked it.\nAfter a bit of thinking I could imagine a repo which is\nused for integration exclusively (no compilation or looking\nat the files at all).\n"},{"id":"39026","messageId":"81b0412b0704100848n69c99f55xa7cc96087cad7e31@mail.gmail.com","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704100758430.6730@woody.linux-foundation.org","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-04-10T15:48:13Z","receivedAt":"2007-04-10T15:48:13Z","isPatch":true,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/10/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> So I'd like this to be merged somewhere, not because \"it works\" or \"it's\n> complete\", but because it's in a shape where I think a lot of people can\n> start fixing small details.\n\nIt is already \"merged somewhere\": as soon as the patches left landed\non vger, it is not possible to loose (and even destroy) them.\nThe feature is just too much sought after.\n\n> For example, with just two smallish updates:\n>  - teach \"git upload-pack\" not to try to follow gitlinks\n>  - teach \"git read-tree\" to check out a git-link as just an empty\n>    subdirectory\n\nwhich also should fix switching between the branches with subprojects.\n"},{"id":"39027","messageId":"Pine.LNX.4.64.0704100849170.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"81b0412b0704100835vbbfe8e7o2df2f121ce088589@mail.gmail.com","subject":"Re: [PATCH 3/6] Add 'resolve_gitlink_ref()' helper function","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T15:52:02Z","receivedAt":"2007-04-10T15:52:02Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Apr 2007, Alex Riesen wrote:\n> \n> After a bit of thinking I could imagine a repo which is\n> used for integration exclusively (no compilation or looking\n> at the files at all).\n\nWell, you also cannot *commit* to a bare repository, so it's a bit \npointless for integration reasons. You'd still have to commit all changes \nsomewhere else.\n\nThat said, it's definitely designed so that if you want to automate \ntracking other peoples bare repositories, you can do so: you'd just have \nto *really* script it with something like\n\n\tgit update-index --cacheinfo 0160000 <sha1> <dirname>\n\n(which is how you could create those commits to a bare repo too, so it's \nnot like this is really even any different)\n\n\t\tLinus\n"},{"id":"39021","messageId":"200704101754.17480.Josef.Weidendorfer@gmx.de","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704100756060.6730@woody.linux-foundation.org","subject":"Re: [PATCH 3/6] Add 'resolve_gitlink_ref()' helper function","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2007-04-10T15:54:17Z","receivedAt":"2007-04-10T15:54:17Z","isPatch":true,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Tuesday 10 April 2007, Linus Torvalds wrote:\n> \n> On Tue, 10 Apr 2007, Alex Riesen wrote:\n> >\n> > On 4/10/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> > > +int resolve_gitlink_ref(const char *path, const char *refname, unsigned\n> > > char *result)\n> > > +{\n> > > +       int len = strlen(path), retval;\n> > > +       char *gitdir;\n> > > +\n> > > +       while (len && path[len-1] == '/')\n> > > +               len--;\n> > > +       if (!len)\n> > > +               return -1;\n> > > +       gitdir = xmalloc(len + MAXREFLEN + 8);\n> > > +       memcpy(gitdir, path, len);\n> > > +       memcpy(gitdir + len, \"/.git/\", 7);\n> > \n> > Can't a subproject be bare?\n> \n> Not when it is checked out, no. That's what \"checked out\" means ;)\n> \n> If a subproject is bare, it never gets resolved, because it's never \n> checked out in a superproject.\n> \n> So a subproject *can* be bare, but when it's bare it is just a totally \n> regular independent git project, simply by *definition* of not being \n> checked out inside a superproject.\n> \n> But hey, that was just a design decision of mine, and if people can argue \n> for it being wrong, I don't think I'm married to it ;)\n\nIt would be nice if a redirection via a \"gitdir = ...\" line\nin .git/link of the subproject (when existing) would be possible.\nThis was part of the light-weight checkout proposal.\n\nIn contrast to contrib/workdir/git-new-workdir, this would allow\nfor (to be implemented) magic symlinks to stay intact when\nmoving the submodule directory around.\n\nHowever, this can be added later.\n\nJosef\n\nPS: I wonder how long it takes to move the official KDE repository over to git ;-)\n"},{"id":"39017","messageId":"81b0412b0704100857h7550b3f9r1772dc5789c80426@mail.gmail.com","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704100849170.6730@woody.linux-foundation.org","subject":"Re: [PATCH 3/6] Add 'resolve_gitlink_ref()' helper function","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-04-10T15:57:03Z","receivedAt":"2007-04-10T15:57:03Z","isPatch":true,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/10/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> >\n> > After a bit of thinking I could imagine a repo which is\n> > used for integration exclusively (no compilation or looking\n> > at the files at all).\n>\n> Well, you also cannot *commit* to a bare repository, so it's a bit\n> pointless for integration reasons. You'd still have to commit all changes\n> somewhere else.\n\nYes. Subprojects are push-only for storing and reference purposes.\nSuperproject can have integrated data checks in Makefiles.\n\n> That said, it's definitely designed so that if you want to automate\n> tracking other peoples bare repositories, you can do so: you'd just have\n> to *really* script it with something like\n>\n>         git update-index --cacheinfo 0160000 <sha1> <dirname>\n>\n> (which is how you could create those commits to a bare repo too, so it's\n> not like this is really even any different)\n\nNice :)\n"},{"id":"39015","messageId":"Pine.LNX.4.64.0704100852550.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"81b0412b0704100848n69c99f55xa7cc96087cad7e31@mail.gmail.com","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T16:07:49Z","receivedAt":"2007-04-10T16:07:49Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Apr 2007, Alex Riesen wrote:\n> \n> It is already \"merged somewhere\": as soon as the patches left landed\n> on vger, it is not possible to loose (and even destroy) them.\n> The feature is just too much sought after.\n\nWell, unless it hits something like Junios 'pu' (or 'next') branch, or \nsomebody (like you?) ends up maintaining a repo with this, it's just \nunnecessarily hard to have lots of people working together on it..\n\nI'm obviously interested in working on it, but at the same time, I don't \nexpect to be a primary *user* of it, so I'm hoping others will come in and \nstart looking at it.\n\nIt looks promising that you're getting involved, but I suspect you may be \na bit too optimistic when you say \"just too much sought after\". We've been \n*talking* about subprojects for a long long time, and we've had other \npatches fail. So...\n\n> > For example, with just two smallish updates:\n> >  - teach \"git upload-pack\" not to try to follow gitlinks\n> >  - teach \"git read-tree\" to check out a git-link as just an empty\n> >    subdirectory\n> \n> which also should fix switching between the branches with subprojects.\n\nYes. It would require either git-read-tree or the git-checkout script \naround it knowing to then also check out the subproject branches.\n\nIt's actually not *entirely* obvious what you should do when you switch \nbranches (or even just do a \"git reset --hard\") in the superproject. The \nbranches in the subprojects are likely to be totally different from the \nsuperproject, so as far as I can see, you end up having two choices when \nyou reset a subproject:\n\n - either basically create a \"disconnected HEAD\" in the subproject(s) when \n   you switch them around as a consequence of resetting/switching the \n   branch in the superproject.\n\n - or you'd stay on the same branch in the subproject, and just reset that \n   branch..\n\n - or you describe the branch name in the \".gitmodules\" file in the\n   superproject, and use whatever branch in the submodule that is \n   described in the supermodule that you reset/check-out.\n\n - or possibly other policies.\n\nSo there is bound to be various \"policy\" issues like this worth sorting \nout. I don't think they matter that deeply.\n\nI would _personally_ tend to like the notion of using \".gitmodules\" in the \nsupermodule to describe things like this, exactly because it's a policy \ndecision - not something that git itself should really decide about, but \nthat the supermodule maintainers can just decide to agree on.\n\nBut I haven't really even thought about all the things I'd want to have in \nthe .gitmodules. We'd obviously need to list the default URL's for the \nsubmodules some way etc, but I haven't really sat down and thought about \nwhat all the higher-level porcelain really would need to know.\n\nI suspect that somebody who has used and set up CVS \"modules\" setups \nshould be thinking about that. I've been a \"stupid user\" for CVS modules \nsetups, but I've never actually needed to really know how they *work*.\n\n\t\tLinus\n"},{"id":"39029","messageId":"Pine.LNX.4.64.0704100909160.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"81b0412b0704100857h7550b3f9r1772dc5789c80426@mail.gmail.com","subject":"Re: [PATCH 3/6] Add 'resolve_gitlink_ref()' helper function","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T16:16:07Z","receivedAt":"2007-04-10T16:16:07Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Apr 2007, Alex Riesen wrote:\n>\n> On 4/10/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> > That said, it's definitely designed so that if you want to automate\n> > tracking other peoples bare repositories, you can do so: you'd just have\n> > to *really* script it with something like\n> > \n> >         git update-index --cacheinfo 0160000 <sha1> <dirname>\n> > \n> > (which is how you could create those commits to a bare repo too, so it's\n> > not like this is really even any different)\n> \n> Nice :)\n\nWell, the *really* nice thing about doing it like this is that you can \nactually update subprojects without even having them even be *local* to \nwhere you do the superproject.\n\nIOW, you could literally build up the superproject by saying that you want \nto track \"all git projects I care about\" somewhere else, and do a series \nof automated\n\n\tgit ls-remote sub-project-xyzzy tracking-branch-xyzzy | ...\n\nand basically create the \"superproject\" without ever actually downloading \nor populating the subprojects at all.\n\nThen, if everything is set up correctly, you can basically use the \nsuperproject as an \"auto-mirror\" - whenever you want to get all the \nprojects you care about, you just clone that superproject, and (once \nyou've taught \"git clone\" to fetch the subprojects, of course ;^) you'd \nbasically fetch them all from their appropriate locations - without ever \nhaving the actual superproject have to even *really* care about it.\n\nSo basically, a superproject could be used as just a \"gathering point\", \nwithout having to actually *contain* any of the subprojects. The actual \nsources for subprojects may be on totally different servers. That's what \nreal distribution is all about.\n\n\t\tLinus\n"},{"id":"39030","messageId":"200704101828.37453.Josef.Weidendorfer@gmx.de","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092115350.6730@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2007-04-10T16:28:37Z","receivedAt":"2007-04-10T16:28:37Z","isPatch":true,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Tuesday 10 April 2007, Linus Torvalds wrote:\n> ...\n> +\tif (resolve_gitlink_ref(ce->name, \"HEAD\", sha1) < 0)\n> +\t\treturn 0;\n> +\treturn hashcmp(sha1, ce->sha1);\n\nSo this does mean that the SHA1 of a gitlink entry corresponds\nto the commit in the subproject?\n\nI wonder if it is not useful to be able to add some attribute(s)\nto a gitlink, i.e. first reference a gitlink object in the superproject,\nwhich then references the submodule commit, and also holds some\nfurther attributes. These attributes can not be put into the subproject,\nas it should be independent.\n\nAn example for such an attribute would be a subproject name/ID.\nAn argument for this: The user should be able to specify some policies\nfor submodules, like \"do not clone/checkout this submodule\". But the\npath where the submodule resides in a given commit is not useful here,\nas a submodule can reside at different paths in the history of the\nsupermodule.\n\nJosef\n"},{"id":"39032","messageId":"81b0412b0704100943k44a0e0b3vf39000bf1eb20e8f@mail.gmail.com","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704100852550.6730@woody.linux-foundation.org","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-04-10T16:43:59Z","receivedAt":"2007-04-10T16:43:59Z","isPatch":true,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/10/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> >\n> > It is already \"merged somewhere\": as soon as the patches left landed\n> > on vger, it is not possible to loose (and even destroy) them.\n> > The feature is just too much sought after.\n>\n> Well, unless it hits something like Junios 'pu' (or 'next') branch, or\n> somebody (like you?) ends up maintaining a repo with this, it's just\n> unnecessarily hard to have lots of people working together on it..\n>\n> I'm obviously interested in working on it, but at the same time, I don't\n> expect to be a primary *user* of it, so I'm hoping others will come in and\n> start looking at it.\n>\n> It looks promising that you're getting involved, but I suspect you may be\n> a bit too optimistic when you say \"just too much sought after\". We've been\n> *talking* about subprojects for a long long time, and we've had other\n> patches fail. So...\n\nThe people who need the feature are still using other VCS.\nSome do not even know about git, the others are more interested\nin their own projects than in hacking on git (like KDE or Ubuntu\npeople). And then there are commercial projects with thirdparty\nlibraries, components or data. The other VCS' provide the feature,\neven if they do it wrong and badly (I never could go back in time in my\nday-work project, always asked myself what was the point of using\nPerforce at all).\nSo, I suspect it is the people who are unable or unwilling\nto contribute to git (to anything, really) who need the feature most.\n"},{"id":"39034","messageId":"81b0412b0704100950s32645423r439d04197ee8cd78@mail.gmail.com","threadId":"7591","inReplyTo":"200704101828.37453.Josef.Weidendorfer@gmx.de","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-04-10T16:50:00Z","receivedAt":"2007-04-10T16:50:00Z","isPatch":true,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/10/07, Josef Weidendorfer <Josef.Weidendorfer@gmx.de> wrote:\n> On Tuesday 10 April 2007, Linus Torvalds wrote:\n> > ...\n> > +     if (resolve_gitlink_ref(ce->name, \"HEAD\", sha1) < 0)\n> > +             return 0;\n> > +     return hashcmp(sha1, ce->sha1);\n>\n> So this does mean that the SHA1 of a gitlink entry corresponds\n> to the commit in the subproject?\n\nRight.\n\n> I wonder if it is not useful to be able to add some attribute(s)\n> to a gitlink, i.e. first reference a gitlink object in the superproject,\n> which then references the submodule commit, and also holds some\n> further attributes. These attributes can not be put into the subproject,\n> as it should be independent.\n\nThese attributes can be put into a file in superproject tree and\nchecked in at the same as the gitlink. No real need for introducing\nanother object type (right now there is no gitlink object type, just\nan entry in tree with special mode).\n"},{"id":"39028","messageId":"200704101923.31404.Josef.Weidendorfer@gmx.de","threadId":"7591","inReplyTo":"81b0412b0704100950s32645423r439d04197ee8cd78@mail.gmail.com","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2007-04-10T17:23:31Z","receivedAt":"2007-04-10T17:23:31Z","isPatch":true,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Tuesday 10 April 2007, Alex Riesen wrote:\n> On 4/10/07, Josef Weidendorfer <Josef.Weidendorfer@gmx.de> wrote:\n> > On Tuesday 10 April 2007, Linus Torvalds wrote:\n> > > ...\n> > > +     if (resolve_gitlink_ref(ce->name, \"HEAD\", sha1) < 0)\n> > > +             return 0;\n> > > +     return hashcmp(sha1, ce->sha1);\n> >\n> > So this does mean that the SHA1 of a gitlink entry corresponds\n> > to the commit in the subproject?\n> \n> Right.\n> \n> > I wonder if it is not useful to be able to add some attribute(s)\n> > to a gitlink, i.e. first reference a gitlink object in the superproject,\n> > which then references the submodule commit, and also holds some\n> > further attributes. These attributes can not be put into the subproject,\n> > as it should be independent.\n> \n> These attributes can be put into a file in superproject tree and\n> checked in at the same as the gitlink. No real need for introducing\n> another object type (right now there is no gitlink object type, just\n> an entry in tree with special mode).\n\nLike... .gitattributes ? ;-)\nOk, this could work; however, there of course is the possibility of\ninconsistencies when e.g. manually moving subprojects around.\n\nHow is consistency ensured for .gitattributes ?\nI see that for .gitignore consistency, the user is responsible.\n\nJosef\n"},{"id":"39040","messageId":"Pine.LNX.4.64.0704101122510.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"200704101828.37453.Josef.Weidendorfer@gmx.de","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T18:45:22Z","receivedAt":"2007-04-10T18:45:22Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Apr 2007, Josef Weidendorfer wrote:\n\n> On Tuesday 10 April 2007, Linus Torvalds wrote:\n> > ...\n> > +\tif (resolve_gitlink_ref(ce->name, \"HEAD\", sha1) < 0)\n> > +\t\treturn 0;\n> > +\treturn hashcmp(sha1, ce->sha1);\n> \n> So this does mean that the SHA1 of a gitlink entry corresponds\n> to the commit in the subproject?\n\nYes.\n\n> I wonder if it is not useful to be able to add some attribute(s)\n> to a gitlink, i.e. first reference a gitlink object in the superproject,\n> which then references the submodule commit, and also holds some\n> further attributes. These attributes can not be put into the subproject,\n> as it should be independent.\n\nThe special \"link\" object has come up before, and I actually thought I'd \ndo it that way first, but there were a few reasons why I didn't:\n\n - I tend to like \"minimal\", and the patches I sent out really are pretty \n   minimal, in the sense that they introduce just _one_ new concept, in \n   one place (it's basically a \"tree entry\" - so it shows up in tree \n   reading and writing, and nowhere else. The index, of course, is the \n   staging area for trees, so the index was also affected, but that was \n   really a very direct result of that \"it's a new tree entry\" thing).\n\n - in a \"link\" object, the only thing that would normally *change* is \n   really just the commit SHA1. Everything else is really pretty static. \n   As such, I decided that it's just a waste of a perfectly fine object to \n   have several thousands of the \"link\" objects that really only differ in \n   the pointer to the commit.\n\n - the \"static\" part, which you might as well have somewhere else, tends \n   to be stuff that you would need to be able to override locally, and as \n   such it does *not* really have a global meaning that is useful \n   historically.\n\n   For example, the things that you'd want to associate with the gitlink \n   are things like \"where would I find the repository that the commit is \n   part of\" and \"what is a description of that submodule\" and \"what are \n   the relationships between the submodules\". These are things that aren't \n   necessarily even totally independent: in CVS, for example, you have \n   module names that are really not submodules themselves, but are really \n   just aliases for *collections* of submodules.\n\n   So a 1:1 link object simply wouldn't make much sense anyway, and you'd \n   want to override those defaults with site-specific ones (maybe there is \n   a \"canonical\" address for the submodule repository, but if you have a \n   copy of it locally on-site, when you clone, you'd rather use the \n   *local* copy over the standard site, for example).\n\nSo all of this just made me say:\n - the tree entry just contains the commit ID of the subproject, and \n   *nothing* else.\n - any incidental data probably isn't 1:1 with tree entries anyway (both \n   over time: you have tree entries being updated with new commit ID's, \n   but the incidental data does *not* change, and over \"space\": different \n   repositories might want to use their local preferences for incidental \n   rules)\n - which all implies that the extra information should go in a separate \n   file that actually describes the modules.\n\nIn fact, it shouldn't be _one_ separate file: it should be at least two, \nsince you'd want to have the *defaults* (which get cloned along with the \nsuperproject) in a revision-controlled file, and then have local *extra* \ninformation that is local. \n\nThis is exactly the same as the situation with the \".gitignore\" file \n(which is revision-controlled and cloned with the respository) and the \n\".git/ignore\" file (which is repository-local).\n\nI've been thinking either \".gitmodules\" (and \".git/modules\") or to just \nextend the \".git/config\" file parser to *also* parse a version-controlled \n\".gitconfig\" file, and just describe the modules there. The config file \nreally has pretty nice syntax, and I think module descriptions in many \nways end up similar to remote branch descriptions, so it would fit in \nthere, I think.\n\n(But there's nothing that says that the \".gitmodules\" file couldn't just \nuse the same parser as the git config file, so I don't really strongly \ncare either way. I just think it would be nice to be able to say\n\n\t[module \"kdelibs\"]\n\t\tdir = kdelibs\n\t\turl = git://git.kde.org/kdelibs\n\t\tdescription = \"Basic KDE libraries module\"\n\n\t[module \"base\"]\n\t\talias = \"kdelibs\", \"kdebase\", \"kdenetwork\"\n\nor whatever. You get the idea..)\n\n\t\tLinus\n"},{"id":"39016","messageId":"200704102004.08329.andyparkins@gmail.com","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704101122510.6730@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-04-10T19:04:04Z","receivedAt":"2007-04-10T19:04:04Z","isPatch":true,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Tuesday 2007, April 10, Linus Torvalds wrote:\n\n> (But there's nothing that says that the \".gitmodules\" file couldn't\n> just use the same parser as the git config file, so I don't really\n> strongly care either way. I just think it would be nice to be able to\n> say\n>\n> \t[module \"kdelibs\"]\n> \t\tdir = kdelibs\n> \t\turl = git://git.kde.org/kdelibs\n> \t\tdescription = \"Basic KDE libraries module\"\n>\n> \t[module \"base\"]\n> \t\talias = \"kdelibs\", \"kdebase\", \"kdenetwork\"\n>\n> or whatever. You get the idea..)\n\nWould it be nicer if .gitmodules were line-based to aid in merging?\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"39048","messageId":"Pine.LNX.4.64.0704101219280.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"200704102004.08329.andyparkins@gmail.com","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T19:20:24Z","receivedAt":"2007-04-10T19:20:24Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Apr 2007, Andy Parkins wrote:\n> \n> Would it be nicer if .gitmodules were line-based to aid in merging?\n\nI seriously doubt you'll ever be merging or changing this a lot. So I \ndon't think it's a huge concern.\n\n\t\tLinus\n"},{"id":"39023","messageId":"200704102129.04548.Josef.Weidendorfer@gmx.de","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704101122510.6730@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2007-04-10T19:29:04Z","receivedAt":"2007-04-10T19:29:04Z","isPatch":true,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Tuesday 10 April 2007, Linus Torvalds wrote:\n> \t[module \"kdelibs\"]\n> \t\tdir = kdelibs\n> \t\turl = git://git.kde.org/kdelibs\n> \t\tdescription = \"Basic KDE libraries module\"\n> \n> \t[module \"base\"]\n> \t\talias = \"kdelibs\", \"kdebase\", \"kdenetwork\"\n\nSo when moving the kdelibs submodule around, you would\nhave to update the .gitmodules file.\n\nI like it.\n\nJosef\n"},{"id":"39051","messageId":"7v6484vxd5.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704100852550.6730@woody.linux-foundation.org","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-10T19:32:38Z","receivedAt":"2007-04-10T19:32:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Tue, 10 Apr 2007, Alex Riesen wrote:\n>> \n>> It is already \"merged somewhere\": as soon as the patches left landed\n>> on vger, it is not possible to loose (and even destroy) them.\n>> The feature is just too much sought after.\n>\n> Well, unless it hits something like Junios 'pu' (or 'next') branch, or \n> somebody (like you?) ends up maintaining a repo with this, it's just \n> unnecessarily hard to have lots of people working together on it..\n\nWell, I was planning to apply this directly on 'master' after\ngiving them another pass.\n"},{"id":"39054","messageId":"Pine.LNX.4.63.0704101239170.27318@qynat.qvtvafvgr.pbz","threadId":"7591","inReplyTo":"200704102004.08329.andyparkins@gmail.com","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-04-10T19:41:56Z","receivedAt":"2007-04-10T19:41:56Z","isPatch":true,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Tue, 10 Apr 2007, Andy Parkins wrote:\n\n> On Tuesday 2007, April 10, Linus Torvalds wrote:\n>\n>> (But there's nothing that says that the \".gitmodules\" file couldn't\n>> just use the same parser as the git config file, so I don't really\n>> strongly care either way. I just think it would be nice to be able to\n>> say\n>>\n>> \t[module \"kdelibs\"]\n>> \t\tdir = kdelibs\n>> \t\turl = git://git.kde.org/kdelibs\n>> \t\tdescription = \"Basic KDE libraries module\"\n>>\n>> \t[module \"base\"]\n>> \t\talias = \"kdelibs\", \"kdebase\", \"kdenetwork\"\n>>\n>> or whatever. You get the idea..)\n>\n> Would it be nicer if .gitmodules were line-based to aid in merging?\n\nthis is very similar to the problem I asked about with merging config files a \ncouple weeks ago. the answer then was that when we get .gitattributes we should \nbe able to specify content specific merge programs that could deal with this \nsort of thing on a per-file basis. That sounds like the answer to your concern \nas well, rather then makeing things order dependant and otherwise harder to read \nto make it able to be merged with the current tools (which assume line-based \norder-dependant content)\n\nDavid Lang\n"},{"id":"39052","messageId":"Pine.LNX.4.64.0704101235160.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"200704102129.04548.Josef.Weidendorfer@gmx.de","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T19:45:57Z","receivedAt":"2007-04-10T19:45:57Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Apr 2007, Josef Weidendorfer wrote:\n> \n> So when moving the kdelibs submodule around, you would\n> have to update the .gitmodules file.\n\nRight. The assumption here is:\n - submodules almost never actually change. You might add a new one \n   occasionally, and once a decade you might do some bigger \n   re-organization, but in general it's pretty much static.\n - when you do move submodules around, it's probably a big flag-day anyway \n   (ie I would expect that it's a big reorg, and that you'd quite likely \n   expect developers to have to re-check out their tree if you did major \n   surgery).\n\nThat's certainly how it works under CVS. I bet we can make it much nicer \nthan CVS, but the point is, people really don't expect submodules to be \nsomething that you move around very dynamically. You want to be *able* to \nmove them around, but it's not a normal operation.\n\n> I like it.\n\nThe advantage with splitting things out like this is that it allows you \nmuch more flexibility than something automatic and deeply integrated does. \n\nYou can still edit the modules setup even if you yourself might not even \nhave that particular module checked out! That may sound insane, but it's \nactually *required* for things like \"oh, the standard server for that \nmodule went away, I need to edit the module settings to get it from xyz \ninstead\".\n\n\t\tLinus\n"},{"id":"39053","messageId":"7v1wisvvtc.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"200704102004.08329.andyparkins@gmail.com","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-10T20:06:07Z","receivedAt":"2007-04-10T20:06:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Andy Parkins <andyparkins@gmail.com> writes:\n\n> On Tuesday 2007, April 10, Linus Torvalds wrote:\n>\n>> (But there's nothing that says that the \".gitmodules\" file couldn't\n>> just use the same parser as the git config file, so I don't really\n>> strongly care either way. I just think it would be nice to be able to\n>> say\n>>\n>> \t[module \"kdelibs\"]\n>> \t\tdir = kdelibs\n>> \t\turl = git://git.kde.org/kdelibs\n>> \t\tdescription = \"Basic KDE libraries module\"\n>>\n>> \t[module \"base\"]\n>> \t\talias = \"kdelibs\", \"kdebase\", \"kdenetwork\"\n>>\n>> or whatever. You get the idea..)\n>\n> Would it be nicer if .gitmodules were line-based to aid in merging?\n\nI personally feel that if there are cases that merge conflict is\nhard to resolve, there is something wrong in the communication\nbetween project members.  In other words, merging this *should*\nbe hard.\n\nReally, if somebody wants to have project X at directory sub/X/\nand somebody else wants the same at directory X/, merging the\nmodules file would be the least of your concern -- resulting\ntoplevel would not build correctly until you decide which tree\nhierarchy should be picked, and later exchange of results among\nproject members would not be usable easily to half the people\nwho picked the hierarchy differently from you did.\n"},{"id":"39058","messageId":"Pine.LNX.4.64.0704101302480.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"7v6484vxd5.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T20:11:55Z","receivedAt":"2007-04-10T20:11:55Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Apr 2007, Junio C Hamano wrote:\n> \n> Well, I was planning to apply this directly on 'master' after\n> giving them another pass.\n\nGoodie. I gave them another pass myself, and noticed a small leak and a \nstupid copy-paste problem, fixed thus..\n\n\t\tLinus\n\n---\ndiff --git a/read-cache.c b/read-cache.c\nindex 8fe94cd..f458f50 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -279,7 +279,7 @@ int base_name_compare(const char *name1, int len1, int mode1,\n \tc2 = name2[len];\n \tif (!c1 && (S_ISDIR(mode1) || S_ISDIRLNK(mode1)))\n \t\tc1 = '/';\n-\tif (!c2 && (S_ISDIR(mode2) || S_ISDIRLNK(mode1)))\n+\tif (!c2 && (S_ISDIR(mode2) || S_ISDIRLNK(mode2)))\n \t\tc2 = '/';\n \treturn (c1 < c2) ? -1 : (c1 > c2) ? 1 : 0;\n }\ndiff --git a/refs.c b/refs.c\nindex 229da74..11a67a8 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -229,6 +229,7 @@ static int resolve_gitlink_packed_ref(char *name, int pathlen, const char *refna\n \tif (!f)\n \t\treturn -1;\n \tread_packed_refs(f, &refs);\n+\tfclose(f);\n \tref = refs.packed;\n \tretval = -1;\n \twhile (ref) {\n"},{"id":"39060","messageId":"7vwt0kugmy.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704101219280.6730@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-10T20:19:17Z","receivedAt":"2007-04-10T20:19:17Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Tue, 10 Apr 2007, Andy Parkins wrote:\n>> \n>> Would it be nicer if .gitmodules were line-based to aid in merging?\n>\n> I seriously doubt you'll ever be merging or changing this a lot. So I \n> don't think it's a huge concern.\n\nI think Andy's comment comes from our earlier discussion on the\nother in-tree configuration, .gitattributes file.\n\nWe were talking about using in-tree .gitattributes for deciding\nif we apply crlf to each paths and other things like which 3-way\nfile-level merge backend to apply, and need to make the system\ngracefully degrade even when in-tree .gitattributes have\nconflict markers during a merge.  And for that purpose, it is\ncertainly easier to arrange \"pick each line, while ignoring <<<\nor === or >>>, and if there are conflicting duplicates do\nsomething sensible about them\", if the file is line oriented.\n\nBut I do not think the .gitmodules thing needs that.  If we have\nconflicting (or non-conflicting for that matter) submodule\nmoves, that's a _MAJOR_ project re-organization, and I do not\nthink we would even want to automatically descend into\nsubmodules for merging or checking-out when we have such a\nsituation in the higher level project.\n"},{"id":"39056","messageId":"Pine.LNX.4.64.0704101325580.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"7vwt0kugmy.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-10T20:33:59Z","receivedAt":"2007-04-10T20:33:59Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Apr 2007, Junio C Hamano wrote:\n> \n> But I do not think the .gitmodules thing needs that.  If we have\n> conflicting (or non-conflicting for that matter) submodule\n> moves, that's a _MAJOR_ project re-organization, and I do not\n> think we would even want to automatically descend into\n> submodules for merging or checking-out when we have such a\n> situation in the higher level project.\n\n100% agreed. \n\nAlso, note that while the \".gitmodules\" (or whatever) file will be \nrequired to do things like \"git pull\", the basic tree-level logic that I \nsent out obviously doesn't need/use .gitmodules at all.\n\nSo there's a very real issue where a repository with submodules still \n\"works\", even with a .gitmodules file that is totally scrogged and doesn't \nhave the right information (yet), it's just that it may simply not be able \nto do all the operations because it cannot figure out where to pull \nmissing subproject data from etc..\n\nSo there is no reason to believe that we need to magically and \nautomatically resolve conflicts - if conflicts happen, functionality is \nreduced, but it's not reduced so much that you cannot use the tree and try \nto resolve them (which is important, btw, since often before you commit \nyour fix for the conflicts you'd want to *test* that fix, so we definitely \ndon't want these kinds of files to be so central that it gets hard to get \nnormal work done without them).\n\nIt really boils down to the same design issue: the way I think submodules \nshould work is that they are very loosely coupled with the supermodule. \nThe fact that the \".gitmodules\" file isn't *that* critical comes largely \nfrom that loose coupling.\n\n\t\tLinus\n"},{"id":"39046","messageId":"7vk5wkuf35.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704101302480.6730@woody.linux-foundation.org","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-10T20:52:46Z","receivedAt":"2007-04-10T20:52:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Tue, 10 Apr 2007, Junio C Hamano wrote:\n>> \n>> Well, I was planning to apply this directly on 'master' after\n>> giving them another pass.\n>\n> Goodie. I gave them another pass myself, and noticed a small leak and a \n> stupid copy-paste problem, fixed thus..\n\nYeah, I noticed the first one but not the second.  Thanks.\n\n> diff --git a/read-cache.c b/read-cache.c\n> index 8fe94cd..f458f50 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -279,7 +279,7 @@ int base_name_compare(const char *name1, int len1, int mode1,\n>  \tc2 = name2[len];\n>  \tif (!c1 && (S_ISDIR(mode1) || S_ISDIRLNK(mode1)))\n>  \t\tc1 = '/';\n> -\tif (!c2 && (S_ISDIR(mode2) || S_ISDIRLNK(mode1)))\n> +\tif (!c2 && (S_ISDIR(mode2) || S_ISDIRLNK(mode2)))\n>  \t\tc2 = '/';\n>  \treturn (c1 < c2) ? -1 : (c1 > c2) ? 1 : 0;\n>  }\n> diff --git a/refs.c b/refs.c\n> index 229da74..11a67a8 100644\n> --- a/refs.c\n> +++ b/refs.c\n> @@ -229,6 +229,7 @@ static int resolve_gitlink_packed_ref(char *name, int pathlen, const char *refna\n>  \tif (!f)\n>  \t\treturn -1;\n>  \tread_packed_refs(f, &refs);\n> +\tfclose(f);\n>  \tref = refs.packed;\n>  \tretval = -1;\n>  \twhile (ref) {\n\nBy the way,...\n\nPeople occasionally ask \"how would I make a small fix to a\ncommit that is buried in the history\", so let me take a moment\nto give them a recipe.\n\nLet's say while reviewing the code after applying all of the\n6-series, you noticed the above thinko.  First find out which\ncommit caused it:\n\n$ git checkout lt/gitlink\n$ git blame -L229,+7 master.. -- refs.c\nb60108a1 (Linus Torvalds 2007-04-09 21:14:26 -0700 229) \tif (!f)\nb60108a1 (Linus Torvalds 2007-04-09 21:14:26 -0700 230) \t\tre..\nb60108a1 (Linus Torvalds 2007-04-09 21:14:26 -0700 231) \tread_packe..\nb60108a1 (Linus Torvalds 2007-04-09 21:14:26 -0700 232) \tref = refs..\nb60108a1 (Linus Torvalds 2007-04-09 21:14:26 -0700 233) \tretval = -..\nb60108a1 (Linus Torvalds 2007-04-09 21:14:26 -0700 234) \twhile (ref..\nb60108a1 (Linus Torvalds 2007-04-09 21:14:26 -0700 235) \t\tif..\n\nThe commit to fix is b60108a1 (this is what I have in my private\nrepo, and I'll be rebuilding the series with this example, so\nyou will never see this commit object name in the end result\nI'll be pushing out).  So I detach the HEAD at that commit and\nmake a fix:\n\n$ git checkout b60108a1\n$ edit refs.c\n$ git diff; # just to make sure\n$ git commit -a --amend\n\nAt this point, the detached HEAD and the original branch look\nlike this:\n\n$ git show-branch lt/gitlink HEAD\n! [lt/gitlink] Teach core object handling functions about gitlinks\n * [HEAD] Add 'resolve_gitlink_ref()' helper function\n--\n * [HEAD] Add 'resolve_gitlink_ref()' helper function\n+  [lt/gitlink] Teach core object handling functions about gitlinks\n+  [lt/gitlink^] Teach \"fsck\" not to follow subproject links\n+  [lt/gitlink~2] Add \"S_IFDIRLNK\" file mode infrastructure for git links\n+  [lt/gitlink~3] Add 'resolve_gitlink_ref()' helper function\n+* [HEAD^] Avoid overflowing name buffer in deep directory structures\n\nWe fixed lt/gitlink~3 and the fixed-up commit is at HEAD.  We\nwant to rebase the rest of lt/gitlink on top of HEAD, like this:\n\n$ git rebase HEAD lt/gitlink\n\nThis will take us back on lt/gitlink branch, set the tip of the\nbranch to the commit we just made with the fix-up, and the first\nround will try to apply the change lt/gitlink~3 brings in on top\nof our HEAD.  This _will_ fail, but that is to be expected, as\nwe intend to replace that with what we just amended.  Just reset\nit away and keep going.\n\n$ git reset --hard\n$ git rebase --skip\n\nDealing with the other one in read-cache.c can be handled\nsimilarly after this. Luckily blame finds out that it is the\nlast in the series (i.e. at the tip of lt/gitlink branch), so\nusual \"fix the topmost commit\" procedure applies.\n\n$ edit read-cache.c\n$ git diff ;# checking...\n$ git commit -a --amend\n"},{"id":"39042","messageId":"20070410210207.GA20955@uranus.ravnborg.org","threadId":"7591","inReplyTo":"7vk5wkuf35.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Sam Ravnborg","fromEmail":"sam@ravnborg.org","sentAt":"2007-04-10T21:02:07Z","receivedAt":"2007-04-10T21:02:07Z","isPatch":true,"sender":{"key":"sam@ravnborg.org","avatar":"https://gravatar.com/avatar/168a912606ed0742d840bb365e3cc21db390c36531a58341dc7a069cc1f15f62?d=mp&s=160"},"body":"On Tue, Apr 10, 2007 at 01:52:46PM -0700, Junio C Hamano wrote:\n> \n> People occasionally ask \"how would I make a small fix to a\n> commit that is buried in the history\", so let me take a moment\n> to give them a recipe.\n\nThat recipe looks ummm complicated...\nWhat I usually do is:\n\ngit format-patch HEAD~4..HEAD\ngit reset --hard HAED~4\npatch -p1 < 0004*\n...edit...\ndelete diff from 0004*\ngit diff >> 0004*\ngit reset --hard\ngit am 000*\n\n\nMaybe this is as complicated as your example but this\nis very simple to deal with.\nAnd I do not destroy history or anything.\n\nBut that said I do not use topic brances but simply\nclone my local repository as needed.\nAnd I always deal with a linear history.\n\n\n[I post this mostly to check if this is insane\nand I need to understand the way you propose to do stuff]\n\n\tSam\n"},{"id":"39020","messageId":"alpine.LFD.0.98.0704101701030.28181@xanadu.home","threadId":"7591","inReplyTo":"7vk5wkuf35.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-04-10T21:03:13Z","receivedAt":"2007-04-10T21:03:13Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 10 Apr 2007, Junio C Hamano wrote:\n\n> By the way,...\n> \n> People occasionally ask \"how would I make a small fix to a\n> commit that is buried in the history\", so let me take a moment\n> to give them a recipe.\n\nThis is definitively good Documentation/howto/ material.\n\n\nNicolas\n"},{"id":"39041","messageId":"7vfy77vs1j.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"20070410210207.GA20955@uranus.ravnborg.org","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-10T21:27:36Z","receivedAt":"2007-04-10T21:27:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Sam Ravnborg <sam@ravnborg.org> writes:\n\n> On Tue, Apr 10, 2007 at 01:52:46PM -0700, Junio C Hamano wrote:\n>> \n>> People occasionally ask \"how would I make a small fix to a\n>> commit that is buried in the history\", so let me take a moment\n>> to give them a recipe.\n>\n> That recipe looks ummm complicated...\n> What I usually do is:\n>\n> git format-patch HEAD~4..HEAD\n> git reset --hard HAED~4\n> patch -p1 < 0004*\n> ...edit...\n> delete diff from 0004*\n> git diff >> 0004*\n> git reset --hard\n> git am 000*\n>\n>\n> Maybe this is as complicated as your example but this\n> is very simple to deal with.\n> And I do not destroy history or anything.\n>\n> But that said I do not use topic brances but simply\n> clone my local repository as needed.\n> And I always deal with a linear history.\n>\n>\n> [I post this mostly to check if this is insane\n> and I need to understand the way you propose to do stuff]\n\nIt's really the same.  You keep 000* file, I keep them in the\noriginal branch and have \"git rebase\" take care of the details.\n"},{"id":"39096","messageId":"20070411080641.GF21701@admingilde.org","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092115350.6730@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2007-04-11T08:06:41Z","receivedAt":"2007-04-11T08:06:41Z","isPatch":true,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nthanks Linus for your nice implementation.  Your core code is so much\nnicer than my hacked-up prototype :-).\n\nI only had little time to actually have a look at it but the core is\nvery similiar to my approach and I'll try to rebase some of my code on\ntop of yours in the following days.\n\nThe only thing I disagree with you is in using HEAD of the submodule:\n\nOn Mon, Apr 09, 2007 at 09:20:29PM -0700, Linus Torvalds wrote:\n> +static int ce_compare_gitlink(struct cache_entry *ce)\n> +{\n> +\tunsigned char sha1[20];\n> +\n> +\t/*\n> +\t * We don't actually require that the .git directory\n> +\t * under DIRLNK directory be a valid git directory. It\n> +\t * might even be missing (in case nobody populated that\n> +\t * sub-project).\n> +\t *\n> +\t * If so, we consider it always to match.\n> +\t */\n> +\tif (resolve_gitlink_ref(ce->name, \"HEAD\", sha1) < 0)\n> +\t\treturn 0;\n> +\treturn hashcmp(sha1, ce->sha1);\n> +}\n\n\n> @@ -2332,6 +2333,8 @@ int index_path(unsigned char *sha1, const char *path, struct stat *st, int write\n>  \t\t\t\t     path);\n>  \t\tfree(target);\n>  \t\tbreak;\n> +\tcase S_IFDIR:\n> +\t\treturn resolve_gitlink_ref(path, \"HEAD\", sha1);\n>  \tdefault:\n>  \t\treturn error(\"%s: unsupported file type\", path);\n>  \t}\n\nAlways using HEAD of the submodule makes branches in the submodule\nuseless.\n\nWhenever you do a checkout in the supermodule you also have to update\nthe submodule and this update has to change the same thing which is read\nabove.\nUpdating the branch which HEAD points to is dangerous.  You could\noverwrite some unrelated branch just because the user forgot to switch\nback to his supermodule-tracking-branch.  The user would always have to\nmake sure that all the submodules are in the correct state for an update\nof the supermodule.\nUpdating HEAD directly is possible now and may make some sense, but you\nstill get problems when you want to switch to some temporary branch in\nthe submodule.  You have no chance to get back to the original supermodule\nversion and now your temporary submodule branch gets shown as the new\nsubmodule version which should be part of the supermodule.\nThe submodule version which is stored in the supermodules tree is kind\nof a hidden/remote reference/branch.  When working on a remote branch\nwe first create a local working branch and then sync it with the remote\none.  I think that it makes sense to use the same model for submodules:\nhave one local branch in the submodule which is used for all work that\nis done in the supermodule context.\n\nSo my advice is:\nAlways read and write one dedicated branch (hardcoded \"master\" or\nconfigurable) when the supermodule wants to access a submodule.\n\nThen you have two type of branches:\nYou can branch the supermodule and have you own branch of the entire\nproject with all submodules.  Use this if you want to commit your\nwork on the submodule into the supermodule.\nYou can also branch the submodule to effectively disconnect the\nsubmodule from the supermodule temporarily.  You can use this to\ndo some experimental/debugging stuff which should not yet go into\nthe supermodule.  Once you want this branch to show up in the\nsupermodule, just merge it to \"master\" and commit it to the supermodule\n(and now its in the supermodule branch, the submodule branch is not\nneeded any more).\n\n\nSee also the discussion about it in the messages around\nhttp://marc.info/?l=git&m=116636334226668&w=2\n\n-- \nMartin Waitz\n"},{"id":"39097","messageId":"87d52bib9e.fsf@morpheus.local","threadId":"7591","inReplyTo":"7vk5wkuf35.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"David Kågedal","fromEmail":"davidk@lysator.liu.se","sentAt":"2007-04-11T08:08:29Z","receivedAt":"2007-04-11T08:08:29Z","isPatch":true,"sender":{"key":"davidk@lysator.liu.se","avatar":"https://avatars.githubusercontent.com/u/60530?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> $ git rebase HEAD lt/gitlink\n> \n> This will take us back on lt/gitlink branch, set the tip of the\n> branch to the commit we just made with the fix-up, and the first\n> round will try to apply the change lt/gitlink~3 brings in on top\n> of our HEAD.  This _will_ fail, but that is to be expected, as\n> we intend to replace that with what we just amended.  Just reset\n> it away and keep going.\n> \n> $ git reset --hard\n> $ git rebase --skip\n\nWouldn't\n\n$ git rebase --onto HEAD lt/gitlink~3 lt/gitlink\n\ndo the trick in one step?\n\n-- \nDavid Kågedal\n"},{"id":"39099","messageId":"81b0412b0704110129q56ee0628jafe8fca808ef9ef8@mail.gmail.com","threadId":"7591","inReplyTo":"20070411080641.GF21701@admingilde.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-04-11T08:29:24Z","receivedAt":"2007-04-11T08:29:24Z","isPatch":true,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/11/07, Martin Waitz <tali@admingilde.org> wrote:\n> So my advice is:\n> Always read and write one dedicated branch (hardcoded \"master\" or\n> configurable) when the supermodule wants to access a submodule.\n\nIn this case it does not correspond to the working tree anymore.\nHEAD is the \"closest\" to working tree of submodule.\n"},{"id":"39100","messageId":"20070411083236.GG21701@admingilde.org","threadId":"7591","inReplyTo":"81b0412b0704100604x2841d96aq194d3dedd303c588@mail.gmail.com","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2007-04-11T08:32:36Z","receivedAt":"2007-04-11T08:32:36Z","isPatch":true,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Tue, Apr 10, 2007 at 03:04:33PM +0200, Alex Riesen wrote:\n> The other thing which will be missed a lot (I miss it that much)\n> is a subproject-recursive git-commit and git-status.\n\ngit-status should really point out if a subproject has any changes,\nas it does for files.  Only that a submodule may have more types of\npossible changes: has new commits which are not yet in the supermodule\nindex, has an dirty index of its own, dirty working directory.\n\nBut for commit it really does not make any sense.  The commit in the\nsubmodule is totally independent to the commit in the supermodule.\nYou'd want the the submodule commit message to not refer to any\nsupermodule stuff (as you likely want to reuse the submodule in other\nsupermodules), while the supermodule commit is much more high-level and\nonly records that the submodule got changed.\n\nWhen viewed from the supermodule, a submodule is just part of its tree,\njust as normal files.  So a submodule commit is conceptually similiar to\nchanging a file, and you don't change files while you commit, also ;-).\n\n-- \nMartin Waitz\n"},{"id":"39101","messageId":"20070411083642.GH21701@admingilde.org","threadId":"7591","inReplyTo":"81b0412b0704110129q56ee0628jafe8fca808ef9ef8@mail.gmail.com","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2007-04-11T08:36:42Z","receivedAt":"2007-04-11T08:36:42Z","isPatch":true,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Wed, Apr 11, 2007 at 10:29:24AM +0200, Alex Riesen wrote:\n> On 4/11/07, Martin Waitz <tali@admingilde.org> wrote:\n> >So my advice is:\n> >Always read and write one dedicated branch (hardcoded \"master\" or\n> >configurable) when the supermodule wants to access a submodule.\n> \n> In this case it does not correspond to the working tree anymore.\n> HEAD is the \"closest\" to working tree of submodule.\n\nyes.\n\nThis has been discussed in length already.\nPlease have a look at the archives.\n\nYour working tree now contains a complete git repository which has\nfeatures which are not available for normal files.  Notable, you\nhave the possibility to create branches in the submodule.\nIf you insist in using HEAD you throw away those submodule capabilities.\n\n-- \nMartin Waitz\n"},{"id":"39102","messageId":"81b0412b0704110142l377231d7j85285a87ef73ce41@mail.gmail.com","threadId":"7591","inReplyTo":"20070411083236.GG21701@admingilde.org","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-04-11T08:42:57Z","receivedAt":"2007-04-11T08:42:57Z","isPatch":true,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/11/07, Martin Waitz <tali@admingilde.org> wrote:\n> > The other thing which will be missed a lot (I miss it that much)\n> > is a subproject-recursive git-commit and git-status.\n>\n> git-status should really point out if a subproject has any changes,\n\nOnly if I want it to. HEAD change check (which is cheap enough\nto be done unconditionally) can be done always.\n\n> But for commit it really does not make any sense.  The commit in the\n> submodule is totally independent to the commit in the supermodule.\n\nRight. Perhaps not a commit in submodule but a recursive check\nfor working directory changes in submodules. So that you can\nmake that you don't make a superproject commit which cannot\nbe resolved to what you had in all the working directories:\n\n  git commit -a --check-clean-subprojects\n"},{"id":"39103","messageId":"81b0412b0704110149g50426a5fh149fe8607f9c163a@mail.gmail.com","threadId":"7591","inReplyTo":"20070411083642.GH21701@admingilde.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-04-11T08:49:18Z","receivedAt":"2007-04-11T08:49:18Z","isPatch":true,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/11/07, Martin Waitz <tali@admingilde.org> wrote:\n> > >Always read and write one dedicated branch (hardcoded \"master\" or\n> > >configurable) when the supermodule wants to access a submodule.\n> >\n> > In this case it does not correspond to the working tree anymore.\n> > HEAD is the \"closest\" to working tree of submodule.\n>\n> yes.\n\n\"Yes\" what? It should _not_ correspond to HEAD?\n\n> This has been discussed in length already.\n> Please have a look at the archives.\n\nI should. But at least a short summary of the reasons\nwould be nice.\n\n> Your working tree now contains a complete git repository which has\n> features which are not available for normal files.  Notable, you\n> have the possibility to create branches in the submodule.\n> If you insist in using HEAD you throw away those submodule capabilities.\n>\n\nIn this (a very special, I believe) case, why not use git update-index\n--cacheinfo?\n"},{"id":"39086","messageId":"20070411085755.GI21701@admingilde.org","threadId":"7591","inReplyTo":"81b0412b0704110142l377231d7j85285a87ef73ce41@mail.gmail.com","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2007-04-11T08:57:55Z","receivedAt":"2007-04-11T08:57:55Z","isPatch":true,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Wed, Apr 11, 2007 at 10:42:57AM +0200, Alex Riesen wrote:\n> On 4/11/07, Martin Waitz <tali@admingilde.org> wrote:\n> >> The other thing which will be missed a lot (I miss it that much)\n> >> is a subproject-recursive git-commit and git-status.\n> >\n> >git-status should really point out if a subproject has any changes,\n> \n> Only if I want it to. HEAD change check (which is cheap enough\n> to be done unconditionally) can be done always.\n\nYes, that's the equivalent of checking normal files.\nThe recursive check for dirty files/index should be configurable.\n\n> >But for commit it really does not make any sense.  The commit in the\n> >submodule is totally independent to the commit in the supermodule.\n> \n> Right. Perhaps not a commit in submodule but a recursive check\n> for working directory changes in submodules. So that you can\n> make that you don't make a superproject commit which cannot\n> be resolved to what you had in all the working directories:\n> \n>  git commit -a --check-clean-subprojects\n\nFor -a such a check may even make sense unconditionally.\nAnd without -a I don't see any value in such a check.\nSo we can just add that check to -a if we see that dirty submodules\nare a problem for users.\n\n-- \nMartin Waitz\n"},{"id":"39090","messageId":"7virc3p8zr.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"20070411083642.GH21701@admingilde.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-11T09:15:36Z","receivedAt":"2007-04-11T09:15:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Waitz <tali@admingilde.org> writes:\n\n> Your working tree now contains a complete git repository which has\n> features which are not available for normal files.  Notable, you\n> have the possibility to create branches in the submodule.\n> If you insist in using HEAD you throw away those submodule capabilities.\n\nWhy?  If you are working in the parent module (e.g integration)\nand notice breakage due to a bug in a submodule, it is very\nplausible that you would want to cd into the directory you have\nthe submodule checked out, which has its own .git/ as its\nrepository, and perform a fix-up there, with the goal of coming\nup with a commit usable by the parent project pointed at by the\nHEAD of the submodule repository.  And while working toward that\ngoal, you will use branches, rebase, rewind or use StGIT there\nin that submodule repository.  It does not forbid you from using\nany of these things -- as long as you end up with a good commit\nat HEAD that the supermodule can use.\n\nOnce you come up with a suitable commit sitting at HEAD of the\nsubmodule repository, you cd up to the parent module.  Top-level\ngit-diff would notice that the commit recorded at the submodule\npath has been updated (because you now have a good commit at\nHEAD of the submodule repository, while earlier the one in your\nindex was a dud).\n\nSo it is not clear to me what your argument about throwing away\ncapabilities is.\n"},{"id":"39104","messageId":"20070411092047.GJ21701@admingilde.org","threadId":"7591","inReplyTo":"81b0412b0704110149g50426a5fh149fe8607f9c163a@mail.gmail.com","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2007-04-11T09:20:47Z","receivedAt":"2007-04-11T09:20:47Z","isPatch":true,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Wed, Apr 11, 2007 at 10:49:18AM +0200, Alex Riesen wrote:\n> On 4/11/07, Martin Waitz <tali@admingilde.org> wrote:\n> >> >Always read and write one dedicated branch (hardcoded \"master\" or\n> >> >configurable) when the supermodule wants to access a submodule.\n> >>\n> >> In this case it does not correspond to the working tree anymore.\n> >> HEAD is the \"closest\" to working tree of submodule.\n> >\n> >yes.\n> \n> \"Yes\" what? It should _not_ correspond to HEAD?\n\nNot neccessarily, yes.\n\nBranches in the submodule make no sense unless they are independent\nfrom supermodule branches.  And then changing to another branch in\nthe submodule automatically means that your current submodule working\ndirectory should be independent to the supermodule.\n\ngit-status in the supermodule should of course warn when a submodule\nis on a different branch, so that you don't accidently loose submodule\ncommits which did not get committed to the supermodule.\n\n> >Your working tree now contains a complete git repository which has\n> >features which are not available for normal files.  Notable, you\n> >have the possibility to create branches in the submodule.\n> >If you insist in using HEAD you throw away those submodule capabilities.\n> >\n> \n> In this (a very special, I believe) case, why not use git update-index\n> --cacheinfo?\n\nI think misunderstood each other.\nFor me branching is not special case.\n\n-- \nMartin Waitz\n"},{"id":"39106","messageId":"7vabxfp873.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"87d52bib9e.fsf@morpheus.local","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-11T09:32:48Z","receivedAt":"2007-04-11T09:32:48Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"David Kågedal <davidk@lysator.liu.se> writes:\n\n> Junio C Hamano <junkio@cox.net> writes:\n>\n>> ...  This _will_ fail, but that is to be expected, as\n>> we intend to replace that with what we just amended.  Just reset\n>> it away and keep going.\n>> \n>> $ git reset --hard\n>> $ git rebase --skip\n>\n> Wouldn't\n>\n> $ git rebase --onto HEAD lt/gitlink~3 lt/gitlink\n>\n> do the trick in one step?\n\nIt is probably more Kosher, and I used to always do that, but it\nis much longer to type, and I use both perhaps 50%/50% depending\non the mood.\n\nWhen the fix-up only adds stuff, 3-way merge would say that the\ncommit before fixing up (lt/gitlink~3 in our example, which you\nare explicitly excluding, while I am letting rebase to see it)\nhas already been applied, in which case the procedure would not\neven stop.  The case illustrated in my message which only adds a\nforgotten line \"fclose(f)\" falls into that category.\n"},{"id":"39108","messageId":"200704111047.01271.andyparkins@gmail.com","threadId":"7591","inReplyTo":"20070411080641.GF21701@admingilde.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-04-11T09:47:00Z","receivedAt":"2007-04-11T09:47:00Z","isPatch":true,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Wednesday 2007 April 11 09:06, Martin Waitz wrote:\n\n> The only thing I disagree with you is in using HEAD of the submodule:\n\nI know we've had this discussion before, but I'm going to bring it up again - \nmainly because Linus's implementation exactly matches what I envisaged when \nwe originally spoke of this.  I think in your \"Updating the branch which HEAD \npoints to is dangerous\" section, the main thing you're not taking into \naccount is that git can make detached checkouts.  Updating HEAD is not \ndangerous - updating refs is; and I don't think anyone is proposing that a \nsubmodule ref should ever be updated by a supermodule.\n\nI think you're also too strongly focussed on the idea that the supermodule \ntracks submodule branches - it cannot branches are not part of \"the\" \nrepository they point at \"a\" repository.  References are outside the \nrepository pointing in, and hence the supermodule cannot refer to them at its \ncore.\n\nNow, if you check out a revision in the supermodule, that's going to look up \nthe submodule revision stored in the DIRLINK tree entry which will recurse \ninto the submodule and checkout that revision - almost certainly as a \ndetached HEAD.  There are three possibilities then:\n - The submodule revision is in the past and no submodule branch points at it\n - The submodule revision is current and a submodule branch points at it\n - The submodule revision is current and multiple submodule branches point at \n   it\nThe supermodule checkout will have to make a decision whether to update the \nsubmodule HEAD (in one case it's obvious: a revision in the past has to be \ndetached HEAD as there is no suitable branch).  It's also possible that the \nsingle submodule branch case is easy - undetach HEAD; however I don't think \nthat is universally correct.\n\nI know you're very much in favour of making branches in the submodule \ncorrespond to branches in the supermodule, but I just don't see a way of \nmaking it work - the supermodule cannot know about submodule branches, \nbranches are not part of the repository, they just point at the repository.  \nMy branches could be different from your branches.\n\nIt may be that some handy configuration settings and some clever porcelain \ncould keep them in sync for your working repository - but it's never going to \nbe the case that checking out \"master\" in the supermodule can be universally \nresolved to mean \"checkout master in the submodule\".\n\nThe way submodules should be treated is that the whole submodule is analogous \nto a single repository-tracked file - that's essentially what a submodule is \nin the end but the content of the \"file\" is the submodule revision.\n\nThere is one difference from ordinary files, a submodule has two \"modified\" \nstates, not one:\n 1. HEAD of submodule is different from DIRLINK revision\n 2. Submodule is dirty\n\nIn state (1) the submodule has to have git-add run on it in the supermodule, \njust as you would with a modified file, to get it into the index (or not if \nyou don't want to commit that change).  In state (2) it should be impossible \nto git-add, because the state of the submodule doesn't represent something \nthat could be restored - there is nothing reasonable that could be written to \nthe DIRLINK tree object.  This is certainly a porcelain issue, because it's \nonly really a warning that \"git-add\" isn't doing what you think it's doing \nwhen the submodule is dirty.\n\nNow, if you change branch in the submodule, the supermodule will see that as a \nchange in the submodule (as it should).  If you changed back, it will be \nrestored and the supermodule will again see it as unchanged.  If you commit \non the submodule, the supermodule will see that as a change and you'll have \nto git-add the submodule and commit in the supermodule.  The submodule is on \nwhatever branch it is on - at all times.\n\nThe only time I can see this causing difficulties is when you want to checkout \nthe tip of a submodule branch - how is the supermodule to know when it is \ncorrect to change HEAD from being detached to being attached?  I suppose it's \ngot to be config-based; and out-of-tree config at that.\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"39110","messageId":"20070411100328.GK21701@admingilde.org","threadId":"7591","inReplyTo":"7virc3p8zr.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2007-04-11T10:03:29Z","receivedAt":"2007-04-11T10:03:29Z","isPatch":true,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Wed, Apr 11, 2007 at 02:15:36AM -0700, Junio C Hamano wrote:\n> Martin Waitz <tali@admingilde.org> writes:\n> \n> > Your working tree now contains a complete git repository which has\n> > features which are not available for normal files.  Notable, you\n> > have the possibility to create branches in the submodule.\n> > If you insist in using HEAD you throw away those submodule capabilities.\n> \n> Why?  If you are working in the parent module (e.g integration)\n> and notice breakage due to a bug in a submodule, it is very\n> plausible that you would want to cd into the directory you have\n> the submodule checked out, which has its own .git/ as its\n> repository, and perform a fix-up there, with the goal of coming\n> up with a commit usable by the parent project pointed at by the\n> HEAD of the submodule repository.  And while working toward that\n> goal, you will use branches, rebase, rewind or use StGIT there\n> in that submodule repository.  It does not forbid you from using\n> any of these things -- as long as you end up with a good commit\n> at HEAD that the supermodule can use.\n\nthat's perfectly fine.\nI only require one more thing: make sure that your commit is on\none dedicated branch (simply by merging your working/rebased/whatever\nbranch into the dedicated one) and not on some random one.\n\nAgain: for your above example this is not neccessary and using HEAD\nwould indeed be perfectly fine.\n\nBut you also have to update the submodule when you do a checkout in\nthe supermodule.  So what do you update?  Updating 'HEAD' is not\nvery concrete, please have a look at my initial mail to Linus.\n\nWhat is stored in the supermodule?  It stores a reference to a specific\npoint in the history of the submodule.  As such I am convinced that\nthe right counterpart inside the submodule is a refs/heads/whatever,\nand not the branch selector HEAD.\nYou can have other branches next to the one which is tracked by the\nsupermodule.  If you always update HEAD you don't have a clear\ndistinction between the branch which is tracked and other branches.\n\n> Once you come up with a suitable commit sitting at HEAD of the\n> submodule repository, you cd up to the parent module.  Top-level\n> git-diff would notice that the commit recorded at the submodule\n> path has been updated (because you now have a good commit at\n> HEAD of the submodule repository, while earlier the one in your\n> index was a dud).\n> \n> So it is not clear to me what your argument about throwing away\n> capabilities is.\n\nIf the supermodule just updates some random submodule branch I happen to\nuse at the time of a supermodule pull then submodule branches are\nof much lower value.\nSuddenly you have to make sure for yourself that the correct branch\ngets updated.\nFor me, different branches should be independent and I want git to\nalways update the correct one.\n\n-- \nMartin Waitz\n"},{"id":"39111","messageId":"20070411113150.GL21701@admingilde.org","threadId":"7591","inReplyTo":"200704111047.01271.andyparkins@gmail.com","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2007-04-11T11:31:50Z","receivedAt":"2007-04-11T11:31:50Z","isPatch":true,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Wed, Apr 11, 2007 at 10:47:00AM +0100, Andy Parkins wrote:\n> On Wednesday 2007 April 11 09:06, Martin Waitz wrote:\n> \n> > The only thing I disagree with you is in using HEAD of the submodule:\n> \n> I know we've had this discussion before, but I'm going to bring it up again - \n> mainly because Linus's implementation exactly matches what I envisaged when \n> we originally spoke of this.  I think in your \"Updating the branch which HEAD \n> points to is dangerous\" section, the main thing you're not taking into \n> account is that git can make detached checkouts.  Updating HEAD is not \n> dangerous - updating refs is; and I don't think anyone is proposing that a \n> submodule ref should ever be updated by a supermodule.\n\nThen we already agree on the most important part.\nMy argument is mostly against updating the ref which is behind HEAD, not\nHEAD per se.  And I haven't thought about using detached HEADs until I\nwrote the mail.\n\n> I think you're also too strongly focussed on the idea that the supermodule \n> tracks submodule branches - it cannot branches are not part of \"the\" \n> repository they point at \"a\" repository.  References are outside the \n> repository pointing in, and hence the supermodule cannot refer to them at its \n> core.\n\nNo, that may be an misunderstanding because my very first prototype\nreally did track branches.  In the meantime I changed my mind, my\ncurrent prototypes all track submodule commits directly.\nBut in doing so we create a branch of its own: remember, a branch in\ngit is just a moving reference into the history.  Such a reference\ncan be stored in .git/refs/heads or it can be stored in the index/tree of\nthe supermodule.  The difference is not really big.\nSo we do not track a branch, but we create a branch by tracking.\n\n> Now, if you check out a revision in the supermodule, that's going to look up \n> the submodule revision stored in the DIRLINK tree entry which will recurse \n> into the submodule and checkout that revision - almost certainly as a \n> detached HEAD.  There are three possibilities then:\n>  - The submodule revision is in the past and no submodule branch points at it\n>  - The submodule revision is current and a submodule branch points at it\n>  - The submodule revision is current and multiple submodule branches point at \n>    it\n> The supermodule checkout will have to make a decision whether to update the \n> submodule HEAD (in one case it's obvious: a revision in the past has to be \n> detached HEAD as there is no suitable branch).  It's also possible that the \n> single submodule branch case is easy - undetach HEAD; however I don't think \n> that is universally correct.\n\nI don't like to guess which branches to update.\nI'd prefer to just unconditionally update one specific one.\n\n> I know you're very much in favour of making branches in the submodule \n> correspond to branches in the supermodule, but I just don't see a way of \n> making it work - the supermodule cannot know about submodule branches, \n> branches are not part of the repository, they just point at the repository.  \n> My branches could be different from your branches.\n\nThat would not work, you are right.\nPlease see my above comment about tracking & branches.\n\n> The way submodules should be treated is that the whole submodule is analogous \n> to a single repository-tracked file - that's essentially what a submodule is \n> in the end but the content of the \"file\" is the submodule revision.\n\nWholeheartedly agreed.\n\n> Now, if you change branch in the submodule, the supermodule will see\n> that as a change in the submodule (as it should).  If you changed\n> back, it will be restored and the supermodule will again see it as\n> unchanged.  If you commit on the submodule, the supermodule will see\n> that as a change and you'll have to git-add the submodule and commit\n> in the supermodule.  The submodule is on whatever branch it is on - at\n> all times.\n\n> The only time I can see this causing difficulties is when you want to\n> checkout the tip of a submodule branch - how is the supermodule to\n> know when it is correct to change HEAD from being detached to being\n> attached?  I suppose it's got to be config-based; and out-of-tree\n> config at that.\n\nAgain, doing things conditionally here just adds to confusion.\nJust have one dedicated branch and be done with it.\n\n-- \nMartin Waitz\n"},{"id":"39119","messageId":"Pine.LNX.4.64.0704110753360.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"20070411080641.GF21701@admingilde.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-11T15:16:10Z","receivedAt":"2007-04-11T15:16:10Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 11 Apr 2007, Martin Waitz wrote:\n> \n> I only had little time to actually have a look at it but the core is\n> very similiar to my approach and I'll try to rebase some of my code on\n> top of yours in the following days.\n> \n> The only thing I disagree with you is in using HEAD of the submodule:\n\nWell, I don't actually see much choice. HEAD is just shorthand for \n\"whatever is checked out\".\n\n> Always using HEAD of the submodule makes branches in the submodule\n> useless.\n\nNo. \n\nBranches in submodules actually in many ways are *more* important than \nbranches in supermodules - it's just that with the CVS mentality, you \nwould never actually see that, because CVS obviously doesn't really \nsupport such a notion.\n\nSo I'd argue that branches in submodules give you:\n\n - you can develop the submodule *independently* of the supermodule, but \n   still be able to easily merge back and forth.\n\n   Quite often, the submodule would be developed entirely _outside_ of the \n   supermodule, and the \"branch\" that gets the most development would thus\n   actually be the \"vendor branch\", entirely outside the supermodule. Call \n   that the \"main\" branch or whatever, inside the supermodule it would \n   often be something like the remote \"remotes/origin/master\" branch.\n\n   So inside the supermodule, the HEAD would generally point to something \n   that is *not* necessarily the \"main development\" branch, because the \n   supermodule maintainer would quite logically and often have his own \n   modifications to the original project on that branch. It migth be a \n   detached branch, or just a local branch inside the submodule.\n\n - branches inside submodules are *also* very useful even inside the \n   supermodule, ie they again allow topic work to be fetched into the\n   submodule *without* having to actually be part of the supermodule,\n   or as a way to track a certain experimental branch of the supermodule.\n\n   I suspect that most supermodule usage is as an \"integrator\" branch, \n   which means that the supermodule tends to follow the \"main \n   development\", and the whole point of the supermodule is largely to have \n   a collection of \"stable things that work together\". \n\n   In contrast, branches within submodules are useful for doing all the \n   development that is *not* yet ready to be committed to the supermodule, \n   exactly because it's not yet been tested in the full \"make World\" kind \n   of situation.\n\n> Whenever you do a checkout in the supermodule you also have to update\n> the submodule and this update has to change the same thing which is read\n> above.\n\nI suspect (but will not guarantee) that the right approach is that a \nsupermodule checkout usually just uses a \"detached HEAD\" setup. Within the \ncontext of the supermodule, only the actual commit SHA1 matters, not what \nbranch it was developed on (side note: I haven't decided if we should \nallow the SHA1 to be a signed tag object too - the current patches \nobviously don't care since they never follow the SHA1 anyway, and it might \nbe a good idea).\n\nSo I strongly suspect (and that is what the patch series embodies) that as \nfar as the supermodule is concerned, it should *not* matter at all what \nbranch the subproject was on. The subproject can use branches for \ndevelopment, and the supermodule really doesn't care what the local \nbranchname was when a commit was made - because branch-names are *local* \nthings, and a branch that is called \"experimental\" in one environment \nmight be called \"master\" in another.\n\nSo once the commit hits the superproject, the branch identities just go \naway (only as far as the superproject is concerned, of course - the \nsubproject still stays with whatever branches it has), and the only thing \nthat matters is the commit SHA1.\n\n> Updating the branch which HEAD points to is dangerous.\n\nI would strongly suggest that the *superproject* never really change the \nstatus of the subproject HEAD, except it updates it for \"pull/reset\", and \nthen it just would use whatever the subproject decided to use.\n\nThe subproject HEAD policy would be entirely under the control of the \nsubproject. If the subproject wants to use a branch to track the \nsuperproject, go wild: have a real branch that is called \"my-integration\" \nand make HEAD a symref to that (and thus any work in the superproject will \nupdate that branch - something that is visible when you pull directly from \nthat subproject!)\n\nBut quite often, I suspect that a subproject would just use a detached \nHEAD. The subproject may have branches of its own, of course, but you can \nthink of HEAD as not being connected to any of it's \"own\" branches, but \nsimply being the \"superproject branch\". That's a fairly accurate picture \nof reality, and using \"detached HEAD\" sounds like a very natural thing to \ndo in that situation.\n\nSo I really think you can do both, and I think using HEAD inside the \nsuperproject gives you exactly that flexibility - you can decide on a \nper-subproject basis whether HEAD should track a real local branch in a \nsubproject, or whether it should be detached.\n\n(Side note: if you do *not* use detatched HEAD, I suspect the .gitmodules \nfile could also contain the branchname to be used for the subproject \ntracking, but I think that's a detail, and quite debatable)\n\n> So my advice is:\n> Always read and write one dedicated branch (hardcoded \"master\" or\n> configurable) when the supermodule wants to access a submodule.\n\nSo the main reasons I don't think that is a good idea are:\n\n - it's less flexible: see above on why you might want to use a dedicated \n   branch *or* just detached HEAD, and why you might want to choose your \n   own name for the dedicated branch.\n\n - it's also going to be quite confusing when the superproject sees \n   something *else* than what is actually checked out. This is an equally \n   strong argument for just using HEAD - when we actually implement a\n\n\t git diff --subproject\n\n   flag that recurses into the subproject, if you don't use HEAD inside \n   the subproject, that suddenly becomes a *very* confusing thing.\n\nIn other words, I really think HEAD is absolutely the right thing to use, \nbut that said, I obviously wrote \"resolve_gitlink_ref()\" so that it can \ntake any ref-name, and we *can* change that later, or make it a per-module \nconfig option or whatever.\n\n\t\tLinus\n"},{"id":"39130","messageId":"7vmz1eof3m.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"20070411100328.GK21701@admingilde.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-11T20:01:17Z","receivedAt":"2007-04-11T20:01:17Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Waitz <tali@admingilde.org> writes:\n\n> What is stored in the supermodule?  It stores a reference to a specific\n> point in the history of the submodule.  As such I am convinced that\n> the right counterpart inside the submodule is a refs/heads/whatever,\n> and not the branch selector HEAD.\n\nBecause 'submodule' is a project on its own, it can make\nprogress while the parent project is still using the stable\ncommit.  Think of this:\n\n - Your application uses product of another project as a\n   library (e.g. you are doing video application and embedding\n   ffmpeg).\n   \n - Your 'master' commit records a commit in the library\n   subproject.  Maybe library subproject declared stable 1.0 and\n   that is what you used to integrate.\n\n - But being an independent project on its own, the library\n   project can make progress, outside the context of this\n   aggregated work (i.e. your application).  Next time you do:\n\n\t$ cd ffmpeg ; git fetch\n\n   there may not be any branch that points at the exact \"stable 1.0\"\n   commit.\n\nWhen you do a \"checkout -f --recurse-into-subprojects\" from the\ntoplevel, I suspect that you would need to detach HEAD in the\nsubproject repository grafted in your application tree to move\nit to the exact commit the toplevel project (i.e. your\napplication) wants, and match the working tree to that commit.\nThe toplevel simply should _not_ have to care what branch that\ncommit comes from.\n"},{"id":"39145","messageId":"20070411221930.GN21701@admingilde.org","threadId":"7591","inReplyTo":"7vmz1eof3m.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2007-04-11T22:19:30Z","receivedAt":"2007-04-11T22:19:30Z","isPatch":true,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Wed, Apr 11, 2007 at 01:01:17PM -0700, Junio C Hamano wrote:\n> When you do a \"checkout -f --recurse-into-subprojects\" from the\n> toplevel, I suspect that you would need to detach HEAD in the\n> subproject repository grafted in your application tree to move\n> it to the exact commit the toplevel project (i.e. your\n> application) wants, and match the working tree to that commit.\n> The toplevel simply should _not_ have to care what branch that\n> commit comes from.\n\nyes.\n\nBut why does everybody want to detach the submodule HEAD, instead\nof creating one 'special' branch which holds the commit which is\nused by the supermodule?\n\nIf you then want to switch to another submodule branch you loose\nthe reference that comes from the supermodule.\n\nI want to create the extra branch exactly _because_ there is\nindependent work going on in the submodule (or the project it is\nbased on).  As you can switch between detached HEAD and an\nindependent branch you can also switch between the 'supermodule branch'\nand independent branches -- only that you can easily switch back\nif you have an branch of your own.\n\nBTW: I also think that your --recurse-into-subprojects should\nbe implied.\nIf you check out one index entry, you should be able to read it\nback afterwards.  That is a nice property everyone expects from\nnormal files and we should try to keep that for submodules.\nWhen checkout_entry wants to touch a submodule we can simply rewrite\nthe 'supermodule branch' in the submodule.  If HEAD happens to point\nto it we also read-tree the submodule.\nThis is easy to understand and implement and I have some good experience\nwith this model.\n\n-- \nMartin Waitz\n"},{"id":"39146","messageId":"Pine.LNX.4.64.0704111532540.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"20070411221930.GN21701@admingilde.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-11T22:36:31Z","receivedAt":"2007-04-11T22:36:31Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 12 Apr 2007, Martin Waitz wrote:\n> \n> But why does everybody want to detach the submodule HEAD, instead\n> of creating one 'special' branch which holds the commit which is\n> used by the supermodule?\n\nI don't think \"everybody\" wants it.\n\nBut the point is, *regardless* of whether you want a \"detached HEAD\" or \nyou want a \"'special' branch\", you should always use HEAD to look up the \ncommit, and using HEAD *allows* both (ie just make HEAD a symref to the \n'special' branch if you want that behaviour).\n\nAnd if you *do* use a special branch, HEAD *must* match that special \nbranch anyway, since when you commit in the supermodule, the only \nbehaviour that makes sense is to commit the currently checked out state!\n\n> I want to create the extra branch exactly _because_ there is\n> independent work going on in the submodule (or the project it is\n> based on).\n\nAnd that is entirely appropriate.\n\nBut that still means that HEAD must point to that branch (when in the \nsubmodule), since that branch must be the one that is checked out. If it \nisn't the branch that is checked out, normal operations like \"git diff\" \netc wouldn't make sense from the supermodule.\n\nAnd that is why *regardless* of whether you use a special branch or not, \nHEAD is the right thing to look up.\n\n\t\tLinus\n"},{"id":"39147","messageId":"461D6432.90205@vilain.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704092115020.6730@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2007-04-11T22:41:54Z","receivedAt":"2007-04-11T22:41:54Z","isPatch":true,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> Since the subprojects don't necessarily even exist in the current tree,\n> much less in the current git repository (they are totally independent\n> repositories), we do not want to try to follow the chain from one git\n> repository to another through a gitlink.\n>   \n\nDoes this consider the case where the intent of the subprojects are to\ncollate multiple, small projects into one bigger project?\n\nIn that case, you might want to keep all of the subprojects in the same\ngit repository.\n\nSam.\n"},{"id":"39148","messageId":"Pine.LNX.4.64.0704111545040.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"461D6432.90205@vilain.net","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-11T22:48:05Z","receivedAt":"2007-04-11T22:48:05Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 12 Apr 2007, Sam Vilain wrote:\n>\n> Linus Torvalds wrote:\n> > Since the subprojects don't necessarily even exist in the current tree,\n> > much less in the current git repository (they are totally independent\n> > repositories), we do not want to try to follow the chain from one git\n> > repository to another through a gitlink.\n> >   \n> \n> Does this consider the case where the intent of the subprojects are to\n> collate multiple, small projects into one bigger project?\n> \n> In that case, you might want to keep all of the subprojects in the same\n> git repository.\n\nI assume you mean \"you might want to keep all of the subprojects' objects \nin the same git object directory\".\n\nAnd yes, that's absolutely true, but it's technically no different from \njust using GIT_OBJECT_DIRECTORY to share objects between totally unrelated \nprojects, or using git/alternates to share objects between (probably \n*less* unrelated repositories, but still clearly individual repos).\n\nSo the main point of superproject/subprojects is to allow independence \n(because independence is what allows it to scale), but there is nothing to \nsay that things *have* to kept totally isolated. \n\n\t\t\tLinus\n"},{"id":"39149","messageId":"461D65E4.6060106@vilain.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704110753360.6730@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2007-04-11T22:49:08Z","receivedAt":"2007-04-11T22:49:08Z","isPatch":true,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> (Side note: if you do *not* use detatched HEAD, I suspect the .gitmodules \n> file could also contain the branchname to be used for the subproject \n> tracking, but I think that's a detail, and quite debatable)\n>   \n\nTo discuss this detail, what about keeping refs, such as\nrefs/submodules/branch/path/* (or some other convention) which are\nupdated on commit? Then you can also easily clone just the submodule.\n\nSam.\n"},{"id":"39151","messageId":"461D6858.4090007@vilain.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704111545040.6730@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2007-04-11T22:59:36Z","receivedAt":"2007-04-11T22:59:36Z","isPatch":true,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> > Does this consider the case where the intent of the subprojects are to\n> > collate multiple, small projects into one bigger project?\n> > \n> > In that case, you might want to keep all of the subprojects in the same\n> > git repository.\n>\n> I assume you mean \"you might want to keep all of the subprojects' objects \n> in the same git object directory\".\n>\n> And yes, that's absolutely true, but it's technically no different from \n> just using GIT_OBJECT_DIRECTORY to share objects between totally unrelated \n> projects, or using git/alternates to share objects between (probably \n> *less* unrelated repositories, but still clearly individual repos).\n>   \n\nWould that be the only distinction?\n\nWould submodules be descended into for object reachability questions?\n\n> So the main point of superproject/subprojects is to allow independence \n> (because independence is what allows it to scale), but there is nothing to \n> say that things *have* to kept totally isolated. \n>   \n\nI'm particularly interested in repositories with, say, thousands of\nsubmodules but only a few hundred meg. I really want to avoid the\nsituation where each of those submodules gets checked or descended into\nseparately for updates etc.\n\nSam.\n"},{"id":"39155","messageId":"Pine.LNX.4.63.0704111600390.28394@qynat.qvtvafvgr.pbz","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704111605210.6730@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-04-11T23:05:56Z","receivedAt":"2007-04-11T23:05:56Z","isPatch":true,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Wed, 11 Apr 2007, Linus Torvalds wrote:\n> On Thu, 12 Apr 2007, Sam Vilain wrote:\n>\n> The reason a *full* global fsck is so expensive is that it would have an\n> absolutely humungous working set, and effectively keep everything in\n> memory through it all. Doing it in stages (\"fsck smaller individiual trees\n> separately\") is actually the same amount of absolute work, but the working\n> set never grows, so it scales much better.\n>\n> (fsck'ing projects individually also happens to allow you to do the\n> sub-project fsck's in parallel across multiple CPU's or multiple machines,\n> so it actually scales much better that way too - but the big problem\n> tends to be excessive memory use, so the \"SMP parallel version\" only\n> makes sense if you have tons of memory and can afford to do these things\n> at the same time!)\n\nwould it make sense to have a --multiple-project option for fsck that would let \nyou specify multiple 'projects' that share a object set and have the default \nchecking not do the reachability checks that cause problems in this case?\n\nThen people can share the objects if they want to and still do a full check, but \nwould get warned that the full check would take a lot of time. which is not a \nbig problem for a housekeeping thing that's run infrequently to find unreachable \nobjects (which is something that should seldom happen in a well managed project)\n\nDavid Lang\n"},{"id":"39152","messageId":"Pine.LNX.4.64.0704111605210.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"461D6858.4090007@vilain.net","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-11T23:16:53Z","receivedAt":"2007-04-11T23:16:53Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 12 Apr 2007, Sam Vilain wrote:\n> >\n> > And yes, that's absolutely true, but it's technically no different from \n> > just using GIT_OBJECT_DIRECTORY to share objects between totally unrelated \n> > projects, or using git/alternates to share objects between (probably \n> > *less* unrelated repositories, but still clearly individual repos).\n> \n> Would that be the only distinction?\n> \n> Would submodules be descended into for object reachability questions?\n\nI think we'll eventually want that *regardless* of how the object handling \nis done (a kind of \"cross-submodule boundary check\"), but I think that's \nactually outside of the scope of the current fsck.\n\nThe current fsck goes to great lengths to make sure that the internal \nconsistency of a repository is good. That's also why it takes so long, and \nwhy it is such an expensive operation to do (notably when you do a \n\"--full\" check).\n\nIn contrast, the \"cross-submodule boundary check\" is a much cheaper \noperation, *if* you have already verified that the projects are internally \nconsistent. It literally boils down to doing a very simplified commit \nchain walker that only parses tree objects and simply spits out the \nSHA1's of the sub-tree commits (and their location in the tree), and then \na separate phase that just verifies those against the submodules.\n\nAnd that separate phase - once you've done the fsck for all the \n*individual* repositories - is truly trivial. It's literally just a matter \nof \"is that SHA1 a valid commit object\". That's *cheap*.\n\nSee?\n\n> I'm particularly interested in repositories with, say, thousands of\n> submodules but only a few hundred meg. I really want to avoid the\n> situation where each of those submodules gets checked or descended into\n> separately for updates etc.\n\nSo I think that the way to verify a superproject is:\n\n - fsck each and every project totally independently. This is something \n   you have to do *anyway*.\n\n - either as you fsck, or as a separate phase after the fsck, just \n   traverse the trees and spit out \"these are the SHA1's of subprojects\"\n\n - finally, just go through the list of SHA1's (after every project has \n   been fsck'd) and verify that they exist (since if they exist, they will \n   have everything that is reachable from them, as that's one of the \n   things that the *local* fsck verifies)\n\nNotice? At no point do you actually need to do a \"global fsck\". You can do \ntotally independent local fsck's, and then a really cheap test of \nconnectedness once those fsck's have completed.\n\nThe reason a *full* global fsck is so expensive is that it would have an \nabsolutely humungous working set, and effectively keep everything in \nmemory through it all. Doing it in stages (\"fsck smaller individiual trees \nseparately\") is actually the same amount of absolute work, but the working \nset never grows, so it scales much better.\n\n(fsck'ing projects individually also happens to allow you to do the \nsub-project fsck's in parallel across multiple CPU's or multiple machines, \nso it actually scales much better that way too - but the big problem \ntends to be excessive memory use, so the \"SMP parallel version\" only \nmakes sense if you have tons of memory and can afford to do these things \nat the same time!)\n\n\t\t\tLinus\n"},{"id":"39153","messageId":"56b7f5510704111630p5eee3c31hbbdf2b90eac7723d@mail.gmail.com","threadId":"7591","inReplyTo":"461D6858.4090007@vilain.net","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Dana How","fromEmail":"danahow@gmail.com","sentAt":"2007-04-11T23:30:08Z","receivedAt":"2007-04-11T23:30:08Z","isPatch":true,"sender":{"key":"danahow@gmail.com","avatar":null},"body":"On 4/11/07, Sam Vilain <sam@vilain.net> wrote:\n> Linus Torvalds wrote:\n> > > Does this consider the case where the intent of the subprojects are to\n> > > collate multiple, small projects into one bigger project?\n> > >\n> > > In that case, you might want to keep all of the subprojects in the same\n> > > git repository.\n> >\n> > I assume you mean \"you might want to keep all of the subprojects' objects\n> > in the same git object directory\".\n> >\n> > And yes, that's absolutely true, but it's technically no different from\n> > just using GIT_OBJECT_DIRECTORY to share objects between totally unrelated\n> > projects, or using git/alternates to share objects between (probably\n> > *less* unrelated repositories, but still clearly individual repos).\n> >\n>\n> Would that be the only distinction?\n>\n> Would submodules be descended into for object reachability questions?\n>\n> > So the main point of superproject/subprojects is to allow independence\n> > (because independence is what allows it to scale), but there is nothing to\n> > say that things *have* to kept totally isolated.\n> >\n>\n> I'm particularly interested in repositories with, say, thousands of\n> submodules but only a few hundred meg. I really want to avoid the\n> situation where each of those submodules gets checked or descended into\n> separately for updates etc.\n\nThis seems slightly related to the hazy picture I'm forming of how\nI'd like to use git at our site.  Essentially, everyone would have their\nown working tree with .git directory, but .git/objects is a symlink\nto a shared object repository.  How do you fully run git-fsck on this\nshared object repository?  The actual heads (roots) are distributed amongst\nmany .git/refs directories (I suppose you could do something akin\nto git-fsck $(cat /somepaths*/.git/refs/*), but that means you know\nwhere all the repositories are).  So in this setup, maybe I'd want to run\nfsck twice: the first time checking everything but not complaining about\ndangling commit objects [but listing them?], and maybe a 2nd finding\nall these in the users' repos [still need to know where these are].\nPlease note this is just a thought experiment at this point.\n\nAnyway,  git started out with a 1:1 relationship between working tree,\nindex, and object repository. Various things could weaken that --\nalternates, subprojects with different relationships to their object\nrepositories, etc. -- so special commands like git fsck which\nfocus mostly on the object repository may need a little tweaking eventually.\n\n-- \nDana L. How  danahow@gmail.com  +1 650 804 5991 cell\n"},{"id":"39164","messageId":"Pine.LNX.4.63.0704111628240.28394@qynat.qvtvafvgr.pbz","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704111646000.6730@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-04-11T23:30:19Z","receivedAt":"2007-04-11T23:30:19Z","isPatch":true,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Wed, 11 Apr 2007, Linus Torvalds wrote:\n\n> On Wed, 11 Apr 2007, David Lang wrote:\n>>\n>> would it make sense to have a --multiple-project option for fsck that would\n>> let you specify multiple 'projects' that share a object set and have the\n>> default checking not do the reachability checks that cause problems in this\n>> case?\n>\n> Well, the thing is, sharing object directories actually makes things\n> *harder* to check, rather than easier.\n>\n> It can be a nice space optimization, and yes, if there really is a lot of\n> shared state, it can make it much cheaper to do some of the checks, but\n> right now we have absolutely *no* way for fsck to then do the reachability\n> check, because there is no way to tell fsck where all the refs are (since\n> now the refs come in from multiple repositories!)\n\nthis is why I was suggesting a --multiple-project option to let you tell fsck \nabout all of the repositories that it needs to look for refs in.\n\n> So the individual objects get cheaper to fsck (no need to fsck shared\n> objects over and over again), but the reachability gets much harder to\n> fsck.\n\nagreed.\n\n> It's not an insurmountable problem, or even necessarily a very large one,\n> but it boils down to one very basic issue:\n>\n> - nobody seems to actually *use* the shared object directory model!\n>\n> The thing is, with pack-files and alternates directories, a lot of the\n> original reasons for shared object directories simply don't exist..\n\nI suspect that if it coudl be checked it would be used more, especially with the \nsubproject support.\n\nDavid Lang\n"},{"id":"39154","messageId":"461D70E1.10901@vilain.net","threadId":"7591","inReplyTo":"200704101828.37453.Josef.Weidendorfer@gmx.de","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2007-04-11T23:36:01Z","receivedAt":"2007-04-11T23:36:01Z","isPatch":true,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Josef Weidendorfer wrote:\n> An example for such an attribute would be a subproject name/ID.\n> An argument for this: The user should be able to specify some policies\n> for submodules, like \"do not clone/checkout this submodule\". But the\n> path where the submodule resides in a given commit is not useful here,\n> as a submodule can reside at different paths in the history of the\n> supermodule.\n>   \n\nI mentioned this briefly on another strand of this thread, but I think\nthat the simplest way to do this would be to just make refs/subproject/*\npopulate itself sensibly when you commit in the superproject.\n\nI mentioned refs/subprojects/path/branch before, but I think it would\nprobably be the sort of thing that should be in the .git/config\n\nSam.\n"},{"id":"39156","messageId":"461D73AD.9000205@vilain.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704101235160.6730@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2007-04-11T23:47:57Z","receivedAt":"2007-04-11T23:47:57Z","isPatch":true,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> On Tue, 10 Apr 2007, Josef Weidendorfer wrote:\n>   \n>> So when moving the kdelibs submodule around, you would\n>> have to update the .gitmodules file.\n>>     \n>\n> Right. The assumption here is:\n>  - submodules almost never actually change. You might add a new one \n>    occasionally, and once a decade you might do some bigger \n>    re-organization, but in general it's pretty much static.\n>  - when you do move submodules around, it's probably a big flag-day anyway \n>    (ie I would expect that it's a big reorg, and that you'd quite likely \n>    expect developers to have to re-check out their tree if you did major \n>    surgery).\n>   \n\nAlso, in the Perl 5 Perforce conversion there are a number of\n\"submodules\" (ie, bundled modules with their own history) that move\naround a lot. In some tree representations used during the conversion\nprocess they might even appear twice in a given tree with differing\nversions.\n\nSam.\n"},{"id":"39157","messageId":"Pine.LNX.4.64.0704111646000.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"Pine.LNX.4.63.0704111600390.28394@qynat.qvtvafvgr.pbz","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-11T23:53:15Z","receivedAt":"2007-04-11T23:53:15Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 11 Apr 2007, David Lang wrote:\n> \n> would it make sense to have a --multiple-project option for fsck that would\n> let you specify multiple 'projects' that share a object set and have the\n> default checking not do the reachability checks that cause problems in this\n> case?\n\nWell, the thing is, sharing object directories actually makes things \n*harder* to check, rather than easier.\n\nIt can be a nice space optimization, and yes, if there really is a lot of \nshared state, it can make it much cheaper to do some of the checks, but \nright now we have absolutely *no* way for fsck to then do the reachability \ncheck, because there is no way to tell fsck where all the refs are (since \nnow the refs come in from multiple repositories!)\n\nSo the individual objects get cheaper to fsck (no need to fsck shared \nobjects over and over again), but the reachability gets much harder to \nfsck.\n\nIt's not an insurmountable problem, or even necessarily a very large one, \nbut it boils down to one very basic issue:\n\n - nobody seems to actually *use* the shared object directory model!\n\nThe thing is, with pack-files and alternates directories, a lot of the \noriginal reasons for shared object directories simply don't exist..\n\n\t\tLinus\n"},{"id":"39158","messageId":"20070411235447.GO21701@admingilde.org","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704110753360.6730@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2007-04-11T23:54:50Z","receivedAt":"2007-04-11T23:54:50Z","isPatch":true,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Wed, Apr 11, 2007 at 08:16:10AM -0700, Linus Torvalds wrote:\n> Branches in submodules actually in many ways are *more* important than \n> branches in supermodules - it's just that with the CVS mentality, you \n> would never actually see that, because CVS obviously doesn't really \n> support such a notion.\n\nI fully agree with you about the importance of submodule branches.\nIn fact, I want to make them even more important and useable!\n\nAnd by the way, I long forgot about CVS ;-)\n\n\n> So I'd argue that branches in submodules give you:\n> \n>  - you can develop the submodule *independently* of the supermodule, but \n>    still be able to easily merge back and forth.\n> \n>    Quite often, the submodule would be developed entirely _outside_ of the \n>    supermodule, and the \"branch\" that gets the most development would thus\n>    actually be the \"vendor branch\", entirely outside the supermodule. Call \n>    that the \"main\" branch or whatever, inside the supermodule it would \n>    often be something like the remote \"remotes/origin/master\" branch.\n> \n>    So inside the supermodule, the HEAD would generally point to something \n>    that is *not* necessarily the \"main development\" branch, because the \n>    supermodule maintainer would quite logically and often have his own \n>    modifications to the original project on that branch. It migth be a \n>    detached branch, or just a local branch inside the submodule.\n\nI fully agree.\n\n>  - branches inside submodules are *also* very useful even inside the \n>    supermodule, ie they again allow topic work to be fetched into the\n>    submodule *without* having to actually be part of the supermodule,\n>    or as a way to track a certain experimental branch of the supermodule.\n> \n>    I suspect that most supermodule usage is as an \"integrator\" branch, \n>    which means that the supermodule tends to follow the \"main \n>    development\", and the whole point of the supermodule is largely to have \n>    a collection of \"stable things that work together\". \n> \n>    In contrast, branches within submodules are useful for doing all the \n>    development that is *not* yet ready to be committed to the supermodule, \n>    exactly because it's not yet been tested in the full \"make World\" kind \n>    of situation.\n\nI fully agree.\nYou are just so much better in describing things than I am...\n\n> > Whenever you do a checkout in the supermodule you also have to update\n> > the submodule and this update has to change the same thing which is read\n> > above.\n> \n> I suspect (but will not guarantee) that the right approach is that a \n> supermodule checkout usually just uses a \"detached HEAD\" setup. Within the \n> context of the supermodule, only the actual commit SHA1 matters, not what \n> branch it was developed on (side note: I haven't decided if we should \n> allow the SHA1 to be a signed tag object too - the current patches \n> obviously don't care since they never follow the SHA1 anyway, and it might \n> be a good idea).\n\nIf you use a detached HEAD then you can no longer switch back to it\nonce you used some other (independent) branch (for testing or whatever).\nThis is my main argument: If you just update some 'special'\nrefs/heads/from-supermodule (or whatever, maybe get it from\n.gitmodules/config) you can still switch between branches, making them\nmore useful IMHO.\n\nIf we create some other way to easily get to the commit referenced by\nthe index of the supermodule then a detached HEAD is ok for me, too.\nBut why create two things (this not-yet-existing way to get the\nsupermodule index entry, plus submodules HEAD) for the same thing?\nWhy not simply create a new refs/heads/whatever?\nThis is easy and everybody knows how to work with it.\n\n> So I strongly suspect (and that is what the patch series embodies) that as \n> far as the supermodule is concerned, it should *not* matter at all what \n> branch the subproject was on. The subproject can use branches for \n> development, and the supermodule really doesn't care what the local \n> branchname was when a commit was made - because branch-names are *local* \n> things, and a branch that is called \"experimental\" in one environment \n> might be called \"master\" in another.\n\nFully agree.\n\nPlease don't confuse my \"I always want to use one dedicated branch\" with\n\"I always want to use one special branch from the submodule project\".\nThis refs/heads/whatever I am talking about is _purely_ for ease of\nuse of the submodule inside the supermodule.  It is in no way linked\nto the branchnames that are used by the submodule project.\nWell, besides that you can merge back and forth between them, of course.\n\n> So once the commit hits the superproject, the branch identities just go \n> away (only as far as the superproject is concerned, of course - the \n> subproject still stays with whatever branches it has), and the only thing \n> that matters is the commit SHA1.\n\nFully agree.\n\n> > Updating the branch which HEAD points to is dangerous.\n> \n> I would strongly suggest that the *superproject* never really change the \n> status of the subproject HEAD, except it updates it for \"pull/reset\", and \n> then it just would use whatever the subproject decided to use.\n> \n> The subproject HEAD policy would be entirely under the control of the \n> subproject. If the subproject wants to use a branch to track the \n> superproject, go wild: have a real branch that is called \"my-integration\" \n> and make HEAD a symref to that (and thus any work in the superproject will \n> update that branch - something that is visible when you pull directly from \n> that subproject!)\n\nSo you now have this nice \"my-integration\" branch lying next to other\nindependent (not-supermodule-related) branches.\nIf you want to _switch_ to one of these unrelated branches you obviously\nhave to change HEAD, and suddenly your unrelated branches are\nconsidered to be part of the supermodule (ok, not yet part of its\nindex of course, but now all supermodule operations would work on\nthis unrelated branch).\n\nI want to preserve these unrelated branches and see them as a strong\nfeature.  Branches in submodules should be independent from the\nsupermodule _because_ the supermodule has no notion of which branch\nis used.\n\n> But quite often, I suspect that a subproject would just use a detached \n> HEAD. The subproject may have branches of its own, of course, but you can \n> think of HEAD as not being connected to any of it's \"own\" branches, but \n> simply being the \"superproject branch\". That's a fairly accurate picture \n> of reality, and using \"detached HEAD\" sounds like a very natural thing to \n> do in that situation.\n\nOnly that you loose your nice detached HEAD view once you start using\nthose nice branches inside your submodule.\n\n> So I really think you can do both, and I think using HEAD inside the \n> superproject gives you exactly that flexibility - you can decide on a \n> per-subproject basis whether HEAD should track a real local branch in a \n> subproject, or whether it should be detached.\n> \n> (Side note: if you do *not* use detatched HEAD, I suspect the .gitmodules \n> file could also contain the branchname to be used for the subproject \n> tracking, but I think that's a detail, and quite debatable)\n> \n> > So my advice is:\n> > Always read and write one dedicated branch (hardcoded \"master\" or\n> > configurable) when the supermodule wants to access a submodule.\n> \n> So the main reasons I don't think that is a good idea are:\n> \n>  - it's less flexible: see above on why you might want to use a dedicated \n>    branch *or* just detached HEAD, and why you might want to choose your \n>    own name for the dedicated branch.\n\nIn terms of flexibility it is important what you can do with the\nsubmodule.  Being able to use branches just like in a normal\nrepository (\"switch the branch to go to an other, unrelated branch\")\nis a plus for me.\n\nA detached HEAD does not give the same level of flexibility as a real\nhead.\n\n>  - it's also going to be quite confusing when the superproject sees \n>    something *else* than what is actually checked out.\n\nWell, the user explicitly expressed his intent to switch to another\nbranch!  In a normal repository you are not confused about the working\ndirectory not being in sync with \"master\", and we always prominently state\nwhich branch you are on.  Of course this has to be clear for submodules,\ntoo.  So if you do git-status in the supermodule it should print some\n\"submodule is on different branch\"-dirty marker.\n\nAt least I had some situations where I wanted to use something like\nthis: use some experimental brach which should not be directly touched\nby the supermodule.  Instead provide a method (\"git merge\nfrom-supermodule\") to sync your working branch with new stuff from\nthe supermodule.\n\n>    This is an equally strong argument for just using HEAD - when we\n>    actually implement a\n> \n> \t git diff --subproject\n> \n>    flag that recurses into the subproject, if you don't use HEAD inside \n>    the subproject, that suddenly becomes a *very* confusing thing.\n\nThis is right.  Suddenly we have one more player in the field which\nyou can diff against.\n\nBefore submodules:\ntree <-> index <-> working file\n\nsubmodules always using HEAD:\ntree <-> index <-> submodule HEAD <-> submodule working dir\n\nsubmodules using some dedicated branch:\ntree <-> index <-> subm. \"from-supermodule\" <-> subm. HEAD <-> subm. wd\n\nI haven't thought about which diff really makes sense in which\nsituation.\n\n\n-- \nMartin Waitz\n"},{"id":"39163","messageId":"56b7f5510704111700r1cb6923ehbfb742512014aebc@mail.gmail.com","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704111646000.6730@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Dana How","fromEmail":"danahow@gmail.com","sentAt":"2007-04-12T00:00:52Z","receivedAt":"2007-04-12T00:00:52Z","isPatch":true,"sender":{"key":"danahow@gmail.com","avatar":null},"body":"On 4/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> It's not an insurmountable problem, or even necessarily a very large one,\n> but it boils down to one very basic issue:\n>\n>  - nobody seems to actually *use* the shared object directory model!\n\nCool -- my previous email makes me either a git idiot or a git pioneer!\n\nSo I'll think through my usage model some more and\nlook over the fsck source.\n\nUntil then,\n-- \nDana L. How  danahow@gmail.com  +1 650 804 5991 cell\n"},{"id":"39165","messageId":"461D7741.50501@vilain.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704111646000.6730@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2007-04-12T00:03:13Z","receivedAt":"2007-04-12T00:03:13Z","isPatch":true,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> It can be a nice space optimization, and yes, if there really is a lot of \n> shared state, it can make it much cheaper to do some of the checks, but \n> right now we have absolutely *no* way for fsck to then do the reachability \n> check, because there is no way to tell fsck where all the refs are (since \n> now the refs come in from multiple repositories!)\n>   \n\nWell, not if the refs are only gitlinks because there is no checkout.\n\n> So the individual objects get cheaper to fsck (no need to fsck shared \n> objects over and over again), but the reachability gets much harder to \n> fsck.\n>\n> It's not an insurmountable problem, or even necessarily a very large one, \n> but it boils down to one very basic issue:\n>\n>  - nobody seems to actually *use* the shared object directory model!\n>\n> The thing is, with pack-files and alternates directories, a lot of the \n> original reasons for shared object directories simply don't exist..\n\nI think that's just the chicken-and-egg problem. Once this happens I\nthink we'll see people aggregating all sorts of related repositories\nwith this feature, and possibly making much richer histories by tracking\nportions of their trees as subprojects rather than just a subdirectory.\n\nSam.\n"},{"id":"39166","messageId":"461D798B.3040008@vilain.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704101325580.6730@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2007-04-12T00:12:59Z","receivedAt":"2007-04-12T00:12:59Z","isPatch":true,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> So there's a very real issue where a repository with submodules still \n> \"works\", even with a .gitmodules file that is totally scrogged and doesn't \n> have the right information (yet), it's just that it may simply not be able \n> to do all the operations because it cannot figure out where to pull \n> missing subproject data from etc..\n>   \n\nWhoa... \"missing\" subproject data?\n\nSurely, unless you're doing lightweight/shallow clones, if you have a\ngitlink you've also got the dependent repository? Otherwise the\nreachability rule will be broken.\n\nSam.\n"},{"id":"39167","messageId":"Pine.LNX.4.64.0704111659240.6730@woody.linux-foundation.org","threadId":"7591","inReplyTo":"461D73AD.9000205@vilain.net","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-12T00:13:55Z","receivedAt":"2007-04-12T00:13:55Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 12 Apr 2007, Sam Vilain wrote:\n> \n> Also, in the Perl 5 Perforce conversion there are a number of\n> \"submodules\" (ie, bundled modules with their own history) that move\n> around a lot. In some tree representations used during the conversion\n> process they might even appear twice in a given tree with differing\n> versions.\n\nThat should actually be something that is fairly natural to handle with \nthe current git submodule design - there's absolutely no problem with \nhaving the same subproject showing up in multiple different places in the \ntree (and each place obviously will have its own commit).\n\nHowever, it causes some questions at two points:\n\n - What do you do in the \".gitmodules\" file, where you describe the \n   submodule setup?\n\n   This is not so much a _problem_ as a \"how do you want to handle it\" \n   issue.\n\n   Would people want such a module to show up as \"one module\" that is just \n   visible in the tree in multiple places? Or do people prefer to think of \n   of it as completely separate modules that just happen to have the same \n   base repository?\n\n   I don't think it's clear that one or the other is the \"right way\" to \n   see things, and I don't think git really should care. I suspect it's \n   more likely to be a detail that some importer script just has to \n   resolve one way or the other.\n\n   The core git infrastructure needs to be able to have one module show up \n   in multiple places over time anyway, so I don't think there is any real \n   reason not to allow the same module to show up in multiple places even \n   within one single commit.. (Ie it's really mostly about the .gitmodules \n   file *syntax* - but if we use the config file syntax, it's actually \n   very natural to allow multiple entries for the module directory name)\n\n   At the same time, there are reasons why you might want to consider them \n   separate modules too - maybe you want to *descibe* them separately, and \n   maybe one of the copies is used for \"legacy support\", and you might be \n   in a situation where you want to check out only one of the copies and \n   not the other (and thus describing them as two *different* modules \n   rather than two versions of the *same* module actually makes sense!).\n\n   So I think this is something where we are technically neutral, but \n   where we may have non-technical issues to choose one representation \n   over another (and those issues may have more to do with the *importer* \n   than with any git issues - if importing from Perforce, it probably \n   makes most sense to make the import behave as much as possible the way \n   Perforce did in that case, and I have *no* idea what that is ;)\n\n - After a conversion is done, and you're no longer talking about a \n   historical archive, but a \"going forward\" concern, exactly how \n   automatic is subproject movement going to be, and what are downstream \n   developers that pull these things supposed to do when a subproject that \n   they have checked out is moved?\n\n   This is mostly a UI issue. I suspect that the initial answer is: \"you \n   may have to un-check-out a subproject, then pull the superproject, and \n   then re-check-it-out to get it in the new location\". Simply because \n   it's going to be a lot easier to do than actually having \"git pull\" \n   notice when subprojects move.\n\n   IOW, that is more of a \"just how nice do we want to be to people\", and \n   I _think_ the answer is: \"as nice as possible, but some things are more \n   important than others, and some things might take longer before they \n   are really pleasant to do\" ;)\n\nHmm?\n\n\t\tLinus\n"},{"id":"39170","messageId":"7vslb6mnva.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704111605210.6730@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-12T00:34:49Z","receivedAt":"2007-04-12T00:34:49Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> So I think that the way to verify a superproject is:\n>\n>  - fsck each and every project totally independently. This is something \n>    you have to do *anyway*.\n>\n>  - either as you fsck, or as a separate phase after the fsck, just \n>    traverse the trees and spit out \"these are the SHA1's of subprojects\"\n>\n>  - finally, just go through the list of SHA1's (after every project has \n>    been fsck'd) and verify that they exist (since if they exist, they will \n>    have everything that is reachable from them, as that's one of the \n>    things that the *local* fsck verifies)\n\nThe small detail in the last step is wrong, though.  Even if\nthey EXIST, they may be isolated commits that are note connected\nto refs, and fsck in the repository would not have warned about\nunreachable trees from such unconnected commits.  So you would\nneed to do a reachability from these commits to the refs in the\nsubproject.\n\nThis would be similar to the quick-fetch topic I sent out a\ncouple of patches for, that implements logic to skip fetching\nobjects from your alternate.  You would have rev-list --objects\ntraverse from them with \"--not --all\" in the subproject\nrepository and make sure it does not trigger \"I could not list\nall objects reachable from the commits you wanted because such\nand such tree/blob are missing\".\n\n    That reminds me of one thing I haven't verified.  I am not\n    absolutely sure that rev-list --objects makes sure that\n    blobs it lists exist (trees are checked as it needs to read\n    them, and if they are missing or corrupt it would notice and\n    barf).  When it is used for the purpose of this \"subproject\n    boundary fsck\" and the quick-fetch, it should.  Perhaps a\n    specialized option to check deeper than usual is needed.  I\n    dunno.\n\n> Notice? At no point do you actually need to do a \"global fsck\". You can do \n> totally independent local fsck's, and then a really cheap test of \n> connectedness once those fsck's have completed.\n\nThis is still true.\n"},{"id":"39172","messageId":"20070412003539.GP21701@admingilde.org","threadId":"7591","inReplyTo":"461D798B.3040008@vilain.net","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2007-04-12T00:35:41Z","receivedAt":"2007-04-12T00:35:41Z","isPatch":true,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Thu, Apr 12, 2007 at 12:12:59PM +1200, Sam Vilain wrote:\n> Linus Torvalds wrote:\n> > So there's a very real issue where a repository with submodules still \n> > \"works\", even with a .gitmodules file that is totally scrogged and doesn't \n> > have the right information (yet), it's just that it may simply not be able \n> > to do all the operations because it cannot figure out where to pull \n> > missing subproject data from etc..\n> >   \n> \n> Whoa... \"missing\" subproject data?\n> \n> Surely, unless you're doing lightweight/shallow clones, if you have a\n> gitlink you've also got the dependent repository? Otherwise the\n> reachability rule will be broken.\n\nWith submodules you actually have a natural cutting point where\nyou can say: no, I don't want to get that.\nSo for submodules the reachability rule is a little bit more relaxed.\n\nAnd when you fetch the superproject you now need some way to fetch\nthe new submodule objects.  They may be in the same upstream repository\nbut it may make sense to have this configurable.\n\n-- \nMartin Waitz\n"},{"id":"39174","messageId":"e7bda7770704111742i2ac12cbas50fd7a3ba5c21cd8@mail.gmail.com","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704101122510.6730@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2007-04-12T00:42:43Z","receivedAt":"2007-04-12T00:42:43Z","isPatch":true,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 4/10/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n\n>  - I tend to like \"minimal\",\n\n>  - in a \"link\" object, the only thing that would normally *change* is\n>   really just the commit SHA1. Everything else is really pretty static.\n\nI like this concept.\n\n\n>       [module \"kdelibs\"]\n>               dir = kdelibs\n>               url = git://git.kde.org/kdelibs\n>               description = \"Basic KDE libraries module\"\n>\n>       [module \"base\"]\n>               alias = \"kdelibs\", \"kdebase\", \"kdenetwork\"\n\nI guess this file could also cover the case where the superproject is\nonly interested in a small subset of the subproject. For example if I\nonly uses some header-files in a library and want\n\"/lib1/src/interface\" in the subproject end up as \"/includes/lib1\" in\nthe superproject. Could single files be handled in a similar way?\n\nAlthough this is just an example, external links shouldn't be\nspecified in the same configuration file as project internal things\n(which should be version-controlled). If the url configuration gets\noverwritten with checkouts there will be problems bisecting if the url\nchanges over time.\n"},{"id":"39178","messageId":"20070412005654.GQ21701@admingilde.org","threadId":"7591","inReplyTo":"e7bda7770704111742i2ac12cbas50fd7a3ba5c21cd8@mail.gmail.com","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2007-04-12T00:56:54Z","receivedAt":"2007-04-12T00:56:54Z","isPatch":true,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Thu, Apr 12, 2007 at 02:42:43AM +0200, Torgil Svensson wrote:\n> I guess this file could also cover the case where the superproject is\n> only interested in a small subset of the subproject. For example if I\n> only uses some header-files in a library and want\n> \"/lib1/src/interface\" in the subproject end up as \"/includes/lib1\" in\n> the superproject. Could single files be handled in a similar way?\n\nConceptionally this information would have to be part of the\nsupermodule tree (after all it changes how your tree is set up).\n\nI think it makes more sense to make users think about which part\nof their tree can be reused and make them choose submodule boundaries\nwisely so that the above partial-checkout is not needed.\n\n> Although this is just an example, external links shouldn't be\n> specified in the same configuration file as project internal things\n> (which should be version-controlled). If the url configuration gets\n> overwritten with checkouts there will be problems bisecting if the url\n> changes over time.\n\nMost of the time we may not need to add any per-submodule URL\ninformation anyway.  If you fetch a new supermodule version, you\ncan get the new submodule from the same source (or from a per-submodule\nsource which can be determined by looking at and munching the supermodule URL).\n\n-- \nMartin Waitz\n"},{"id":"39182","messageId":"Pine.LNX.4.64.0704111850240.4061@woody.linux-foundation.org","threadId":"7591","inReplyTo":"7vslb6mnva.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-12T01:52:46Z","receivedAt":"2007-04-12T01:52:46Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 11 Apr 2007, Junio C Hamano wrote:\n>\n> The small detail in the last step is wrong, though.  Even if\n> they EXIST, they may be isolated commits that are note connected\n> to refs, and fsck in the repository would not have warned about\n> unreachable trees from such unconnected commits.\n\nThe superproject *is* a ref.\n\nYou cannot prune the subprojects on their own. That's the *only* real \nspecial rule about subprojects. Exactly because pruning them on their own \nis not a valid op to do.\n\nIt's the same way with an source of \"alternate\" objects (or a shared \nobject directory) - you'd better not prune them, because other projects \nmay have refs to them that you don't know about locally. So this isn't \nsomethign new to subprojects.\n\n\t\tLinus\n"},{"id":"39184","messageId":"5D00E27D-5DB4-4F66-A4BB-752300F9D05E@silverinsanity.com","threadId":"7591","inReplyTo":"20070411235447.GO21701@admingilde.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Brian Gernhardt","fromEmail":"benji@silverinsanity.com","sentAt":"2007-04-12T01:57:23Z","receivedAt":"2007-04-12T01:57:23Z","isPatch":true,"sender":{"key":"benji@silverinsanity.com","avatar":"https://gravatar.com/avatar/e06c101dbc25c68114d859b4a9ec7cf8a2c52fd2b0270ef0eac0e2e63ff22311?d=mp&s=160"},"body":"\nOn Apr 11, 2007, at 7:54 PM, Martin Waitz wrote:\n\n> Before submodules:\n> tree <-> index <-> working file\n>\n> submodules always using HEAD:\n> tree <-> index <-> submodule HEAD <-> submodule working dir\n>\n> submodules using some dedicated branch:\n> tree <-> index <-> subm. \"from-supermodule\" <-> subm. HEAD <->  \n> subm. wd\n\nWhy can't can't we extend checkout with an option to look for an  \nenclosing git project, find the gitlink in the index, and check out  \nthat commit?  That allows you to return to the original state without  \nneeding to bother with new special branches.\n\nAnd instead of recording the path in a .gitmodules file, why not a  \nlist of git directories we search for the commit?  Allows moving of  \nsubprojects without suddenly breaking configuration files.  When we  \nfind the appropriate git dir, we can use a .gitlink file or symlinks  \nto attach the directory to it's repository.\n\nI dislike moving git in the direction of enforcing more policy  \ninstead of less, and of making it less capable of handling content  \nmovement instead of more.\n\n~~ Brian\n"},{"id":"39185","messageId":"7vy7kyl5br.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704111850240.4061@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-12T02:00:40Z","receivedAt":"2007-04-12T02:00:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Wed, 11 Apr 2007, Junio C Hamano wrote:\n>>\n>> The small detail in the last step is wrong, though.  Even if\n>> they EXIST, they may be isolated commits that are note connected\n>> to refs, and fsck in the repository would not have warned about\n>> unreachable trees from such unconnected commits.\n>\n> The superproject *is* a ref.\n\nBut when you fsck the subproject repository in isolation in the\nearlier step in your procedure, that is not taken into account,\nis it?\n\nThe situation I had in mind was not about pruning, but an\nearlier fetch, either the native one that unpacks the objects\ninto loose form or a http walker, fetched a commit near the tip\nbut was interrupted/killed before finishing the fetch nor\nupdating the ref.  The tip of such an incomplete commit chain\nwould be reported dangling.  They are ahead of your refs but\nthey may lack commits and trees to complete the chain back to\nyour refs yet.  When the higher-level project points at such a\ncommit, the existence of the commit is not a proof that\neverything needed to complete the commit is available.\n\nWe need to prove that separately, and that was my suggestion to\nrun a \"rev-list --objects $those-commits --not --all\" in the\nsubproject repository, simlar to what the quick-fetch topic\ndoes.\n"},{"id":"39186","messageId":"Pine.LNX.4.64.0704111854160.4061@woody.linux-foundation.org","threadId":"7591","inReplyTo":"461D798B.3040008@vilain.net","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-12T02:01:48Z","receivedAt":"2007-04-12T02:01:48Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\n[ Dang. Power failure in the middle of writing emails. Can't remember \n  which one was lost. Am rewriting some of this reply in abbreviated form.  ]\n\nOn Thu, 12 Apr 2007, Sam Vilain wrote:\n>\n> Linus Torvalds wrote:\n> > So there's a very real issue where a repository with submodules still \n> > \"works\", even with a .gitmodules file that is totally scrogged and doesn't \n> > have the right information (yet), it's just that it may simply not be able \n> > to do all the operations because it cannot figure out where to pull \n> > missing subproject data from etc..\n> >   \n> \n> Whoa... \"missing\" subproject data?\n\nAbsolutely. Not just subproject data. The whole subproject is often \nmissing.\n\nIf I fetch the KDE superproject, I generally do *not* want every single \nsubproject. In fact, I'd likely just want one or two subprojects.\n\nThe notion that all subprojects are populated is a *bug*. I would \npersonally refuse to use such a setup. Even CVS can handle that just fine, \nwe certainly don't want to be worse than CVS here.\n\nIf you just track a project, it's quite common to only check out the \"src\" \nmodule, and *not* fetch things like the \"validation\" or \"test\" module if \nyou're just following along. \n\nOr you might fetch the \"kdebase\" module, but that sure doesn't mean that \nyou want all the other ones (kdevelop source code? full kdelibs sources? \nIf I'm only interested in kwin and some other random app? No thanks!).\n\n> Surely, unless you're doing lightweight/shallow clones, if you have a\n> gitlink you've also got the dependent repository? Otherwise the\n> reachability rule will be broken.\n\nThe reachability rule *must* be breakable. That's why fsck currently \ndoesn't care AT ALL.\n\nIt's much better to break that rule than to even check it! I'd rather \nleave fsck like it is now, than to *ever* fix it, if the \"fix\" involves \n\"you have to always fetch all submodules to shut fsck up\".\n\n\t\tLinus\n"},{"id":"39188","messageId":"7vmz1el51o.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"7vy7kyl5br.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-12T02:06:43Z","receivedAt":"2007-04-12T02:06:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>\n>> On Wed, 11 Apr 2007, Junio C Hamano wrote:\n>>>\n>>> The small detail in the last step is wrong, though.  Even if\n>>> they EXIST, they may be isolated commits that are note connected\n>>> to refs, and fsck in the repository would not have warned about\n>>> unreachable trees from such unconnected commits.\n>>\n>> The superproject *is* a ref.\n>\n> But when you fsck the subproject repository in isolation in the\n> earlier step in your procedure, that is not taken into account,\n> is it?\n\nAh, forget about this.  The HEAD, which is in the tree of the\nhigher-level project, is a ref.  Silly me.\n"},{"id":"39190","messageId":"Pine.LNX.4.64.0704111903060.4061@woody.linux-foundation.org","threadId":"7591","inReplyTo":"Pine.LNX.4.63.0704111628240.28394@qynat.qvtvafvgr.pbz","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-12T02:14:31Z","receivedAt":"2007-04-12T02:14:31Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 11 Apr 2007, David Lang wrote:\n>\n> On Wed, 11 Apr 2007, Linus Torvalds wrote:\n> > \n> > It can be a nice space optimization, and yes, if there really is a lot of\n> > shared state, it can make it much cheaper to do some of the checks, but\n> > right now we have absolutely *no* way for fsck to then do the reachability\n> > check, because there is no way to tell fsck where all the refs are (since\n> > now the refs come in from multiple repositories!)\n> \n> this is why I was suggesting a --multiple-project option to let you tell fsck\n> about all of the repositories that it needs to look for refs in.\n\nWell, just from a personal observation:\n - I would *personally* actually refuse to share objects with anybody \n   else.\n\nI just find the idea too scary. Somebody doing something bad to their \nobject store by mistake (running \"git prune\" without realizing that there \nare *my* objects there too, or just deciding that they want to play with \nthe object directory by hand, or running a new fancy experimental importer \nthat has a subtle bug wrt object handling or anything like that).\n\nI'll endorse use \"alternates\" files, but partly because I know the main \nproject is safe (any alternates usage is in the \"satellite\" clones anyway, \nand they will never write to the alternate object directory), and partly \nbecause at least for the kernel, we don't have branches that get reset in \nthe main project, so there's no reason to fear that a \"git repack -a -d\" \nwill ever screw up any of the satellite repositories even by mistake.\n\nBut for git projects, even alternates isn't safe, in case somebody bases \ntheir own work on a version of \"pu\" that eventually goes away (even with \nreflogs, pruning *eventually* takes place).\n\nSo I tend to think that alternates and shared object directories are \nreally for \"temporary\" stuff, or for *managed* repositories that are at \ngit *hosting* sites (eg repo.or.cz), and where there is some other safety \ninvolved, ie users don't actually access the object directories directly \nin any way.\n\nSo I've at least personally come to the conclusion that for a *developer* \n(as opposed to a hosting site!), shared object directories just never make \nsense. The downsides are just too big. Even alternates is something where \nyou just need to be fairly careful!\n\n\t\tLinus\n"},{"id":"39191","messageId":"Pine.LNX.4.64.0704111917130.4061@woody.linux-foundation.org","threadId":"7591","inReplyTo":"7vmz1el51o.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-12T02:28:00Z","receivedAt":"2007-04-12T02:28:00Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 11 Apr 2007, Junio C Hamano wrote:\n> \n> Ah, forget about this.  The HEAD, which is in the tree of the\n> higher-level project, is a ref.  Silly me.\n\nWell, not entirely \"silly you\".\n\nIf you do a \"git reset\" in the superproject, that will obviously have to \nrewrite the heads in the subproject.\n\nI do suspect that we should always enable reflogs for the subprojects, so \nthat pruning is safe even for these kinds of situations, but that \ndoesn't resolve all issues.\n\nFor example: to manage *cloning* of the extra stuff, you might actually \nwant to have externally visible refs, and while I suspect the main \nsolution will always be to just do good maintenance (ie \"don't do 'git \nbisect' and _never_ rewrite history in the main superproject!!\"), I don't \nthink it's out of the question to add other safety nets too..\n\nSo for example, while I'm not sure it's necessary, I don't think it would \nbe *wrong* if we might eventually end up having *other* safety features \nlike adding a totally separate \"refs/superprojects/xyzzy\" ref structure. \n\nOr something like that.. Just to make the refs more visible both \nexternally and internally, and to make it much harder to make stupid \nmistakes without realizing it.\n\nI suspect a lot of this will depend on just how many mistakes people make. \nI don't think we've so far had a single problem with alternates files, \nre-basing, and people then pruning away objects used by other repositories \nby mistake, so maybe people really don't make those kinds of mistakes.\n\nSo maybe we don't need any extra safety nets at all. But who knows..\n\n\t\tLinus\n"},{"id":"39192","messageId":"7vbqhul3yk.fsf@assigned-by-dhcp.cox.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704111903060.4061@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-12T02:30:11Z","receivedAt":"2007-04-12T02:30:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> But for git projects, even alternates isn't safe, in case somebody bases \n> their own work on a version of \"pu\" that eventually goes away (even with \n> reflogs, pruning *eventually* takes place).\n>\n> So I tend to think that alternates and shared object directories are \n> really for \"temporary\" stuff, or for *managed* repositories that are at \n> git *hosting* sites (eg repo.or.cz), and where there is some other safety \n> involved, ie users don't actually access the object directories directly \n> in any way.\n\nActually that is not even true for repo.or.cz -- the site lets\npeople to create *forks* of the main project, and I recall it is\nimplemented in terms of alternates.\n\nThat's one of the reasons I never asked to take over git.git\nrepository there.  I have alt-git.git instead, which does not\nallow forks.\n"},{"id":"39194","messageId":"461DADF5.2010902@vilain.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704111854160.4061@woody.linux-foundation.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2007-04-12T03:56:37Z","receivedAt":"2007-04-12T03:56:37Z","isPatch":true,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Linus Torvalds wrote:\n>> Whoa... \"missing\" subproject data?\n>>     \n> Absolutely. Not just subproject data. The whole subproject is often \n> missing.\n>\n> If I fetch the KDE superproject, I generally do *not* want every single \n> subproject. In fact, I'd likely just want one or two subprojects.\n>   \n\nOk, but couldn't this be considered a variation of a lightweight checkout?\n\nThe only reason I'm worried about this is the case where the\nsuperproject contains *thousands* of subprojects. Eg, a superproject for\nall repo.or.cz projects. Say in a day 200 projects get updated with a\nfew commits - do you have to do 200 pulls or just one? But maybe that\nproblem can be solved in another way, or maybe it won't really hurt so\nmuch in practice and still be faster/more efficient than rsync mirroring.\n\nThis is especially the case in concert with gittorrent, which will need\nmodifications to support sharing multiple repositories (not that that's\na huge issue, given there's no implementation yet).\n\n>> Surely, unless you're doing lightweight/shallow clones, if you have a\n>> gitlink you've also got the dependent repository? Otherwise the\n>> reachability rule will be broken.\n>>     \n>\n> The reachability rule *must* be breakable. That's why fsck currently \n> doesn't care AT ALL.\n>\n> It's much better to break that rule than to even check it! I'd rather \n> leave fsck like it is now, than to *ever* fix it, if the \"fix\" involves \n> \"you have to always fetch all submodules to shut fsck up\".\n>   \n\nWell fsck can be fixed easily enough to not descend, like lightweight\ncheckouts.\n\nWhat I really want to avoid is the situation where you can't checkout,\neven though you didn't indicate a shallow/lightweight clone.\n\nWhat else might this decision impact? Obviously with a smaller base you\nhave fewer delta targets, though that's probably not a real issue.\n\nSam.\n"},{"id":"39219","messageId":"200704121712.08706.Josef.Weidendorfer@gmx.de","threadId":"7591","inReplyTo":"20070411235447.GO21701@admingilde.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2007-04-12T15:12:08Z","receivedAt":"2007-04-12T15:12:08Z","isPatch":true,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Thursday 12 April 2007, you wrote:\n> If you use a detached HEAD then you can no longer switch back to it\n> once you used some other (independent) branch (for testing or whatever).\n> This is my main argument: If you just update some 'special'\n> refs/heads/from-supermodule (or whatever, maybe get it from\n> .gitmodules/config) you can still switch between branches, making them\n> more useful IMHO.\n\nThe supermodule checkout could create a .git/SUPER_HEAD for this.\nOK, that is a special kind of reference.\n\nOr introduce \"git --super ...\" with works with the superproject.\nForm a submodule directory, a \"git --super checkout .\" could reset the\nsubmodule checkout. \n\nJosef\n"},{"id":"39223","messageId":"Pine.LNX.4.63.0704121016240.29621@qynat.qvtvafvgr.pbz","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704111903060.4061@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-04-12T17:18:40Z","receivedAt":"2007-04-12T17:18:40Z","isPatch":true,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Wed, 11 Apr 2007, Linus Torvalds wrote:\n\n> So I tend to think that alternates and shared object directories are\n> really for \"temporary\" stuff, or for *managed* repositories that are at\n> git *hosting* sites (eg repo.or.cz), and where there is some other safety\n> involved, ie users don't actually access the object directories directly\n> in any way.\n>\n> So I've at least personally come to the conclusion that for a *developer*\n> (as opposed to a hosting site!), shared object directories just never make\n> sense. The downsides are just too big. Even alternates is something where\n> you just need to be fairly careful!\n\nI was actually thinking that hosting sites (and things like gitorrent) would be \nthe ones that would get the most benifit from shareing objects. the amount saved \nfor any individual developer is probably fairly minor (and the individual \ndeveloper could run a script to look across their objects and hard-link them \ntogeather if they care about the space)\n\nDavid Lang\n"},{"id":"39224","messageId":"56b7f5510704121132g3961060amb394978bb49093e6@mail.gmail.com","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704111903060.4061@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Dana How","fromEmail":"danahow@gmail.com","sentAt":"2007-04-12T18:32:57Z","receivedAt":"2007-04-12T18:32:57Z","isPatch":true,"sender":{"key":"danahow@gmail.com","avatar":null},"body":"On 4/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> On Wed, 11 Apr 2007, David Lang wrote:\n> > this is why I was suggesting a --multiple-project option to let you tell fsck\n> > about all of the repositories that it needs to look for refs in.\n>\n> Well, just from a personal observation:\n>  - I would *personally* actually refuse to share objects with anybody\n>    else.\n>\n> I just find the idea too scary. Somebody doing something bad to their\n> object store by mistake (running \"git prune\" without realizing that there\n> are *my* objects there too, or just deciding that they want to play with\n> the object directory by hand, or running a new fancy experimental importer\n> that has a subtle bug wrt object handling or anything like that).\n>\n> I'll endorse use \"alternates\" files, but partly because I know the main\n> project is safe (any alternates usage is in the \"satellite\" clones anyway,\n> and they will never write to the alternate object directory), and partly\n> because at least for the kernel, we don't have branches that get reset in\n> the main project, so there's no reason to fear that a \"git repack -a -d\"\n> will ever screw up any of the satellite repositories even by mistake.\n>\n> But for git projects, even alternates isn't safe, in case somebody bases\n> their own work on a version of \"pu\" that eventually goes away (even with\n> reflogs, pruning *eventually* takes place).\n>\n> So I tend to think that alternates and shared object directories are\n> really for \"temporary\" stuff, or for *managed* repositories that are at\n> git *hosting* sites (eg repo.or.cz), and where there is some other safety\n> involved, ie users don't actually access the object directories directly\n> in any way.\n>\n> So I've at least personally come to the conclusion that for a *developer*\n> (as opposed to a hosting site!), shared object directories just never make\n> sense. The downsides are just too big. Even alternates is something where\n> you just need to be fairly careful!\n\nThese arguments all seem pretty convincing to me --\nmaybe the problem is that I'm not a \"*developer*\" right now.\nInstead I'm part of a multi-developer *site*.\nBelow I talk about a possible way we could use git\nwithout changing it (since I recognize this would be a minority usage pattern).\n\nWe use perforce to manage a mixed hardware/software project\n(I'm the 55GB check-out guy, remember?).  We have at least 3 different\nkinds of data with different usage patterns, and using perforce for\neverything in one centralized server was not the best solution.\n\nEach user (\"client\") has their own worktree and the perforce\nrepository is on a shared central server.  You can consider perforce\nto have the equivalent of git's index, but it is stored on the server,\nin one file (\"db.have\") covering all clients.  Obviously that becomes a\nbottleneck -- and recently db.have got larger than the total cache RAM on\nthe server, which really slowed things down until we moved to a larger\nserver.  But repository architecture aside,  the real problem has been\nperforce's usability.  Frequently one contributor,  having gotten ahead\nof the team,  needs to share this more recent work with only a few\npeople.  This could be done with p4 branching,  but this is really clunky.\nSo instead the work is pushed out (submitted) to everyone, causing\ninstability; this is partially remedied by doing it in smaller chunks.\nAnother perforce problem is that tagging consumes a lot of server\nspace (and may slow things down as well).\n\nSome of this data will stay in perforce, some will move into revision\ncontrol built-in to some of our other tools, and I'd like to try to move some\nof it into git.  The main attraction for the last group is the lightweight\nbranching that would allow early/tentative work to be easily shared.\nI think the subproject work currently being discussed is going to\nbe very helpful as well -- the perforce equivalent is chaotic.\n\nWe could give each user a work tree and an object repository,\nand then have a \"release\" repository.  Unfortunately,  this would be\nslower to use than the current perforce \"solution\": users would check\nin to their local repository, at the speed of gzip, anyone checking\nit out would do so at the speed of gzip, and all work would need\nto be resubmitted (using perforce jargon here) to the central repo,\nagain at the speed of gzip.  Currently, people either submit or\ncheck out from the central repo, and it's all done at the speed of\na network copy.  This speed issue is important because of\nthe size of a commit we'd like to share (but not yet release):\nabout 40 files, half of them control files of several KB each, 1/4 of them\ndesign files of several MB each, and the last 1/4 detailed design\nfiles 100X larger.  These 40 files will reference (include) 50 others\nof several KB each sprinkled through-out the hierarchy, a few of which\nmight have changed.  And yes, almost all of these are generated files,\nbut the generation time, and the instability of the tool and script environment,\npreclude forcing the other users to regenerate them, like you would\nwith a .o file.\n\nSo, there are 2 alternative set-ups. In one, everyone uses a shared\nobject repository (everyone's .git/objects is a symlink to it). In this\nrepository, objects/. , objects/?? , objects/pack , and objects/info all\nhave \"sticky\" set, and we do the appropriate machinations to make\nall files read-only. There would be an additional phantom user \"git\"\nwho owns the shared object repository (the only user whose .git/objects\nis not a symlink).  Users would commit to their own repositories,\nwhich would write data to the shared object repository and\nupdate their refs (e.g. HEAD). To \"release\", push to the ~git repository.\nThis push would be like a current push -- fast-forward only, figure out the list\nof objects that need to be transmitted -- but instead of transmitting the\nobjects, change their ownership to ~git and then update ~git's refs.\nSince users can share local commits, maybe the ~git ownership\nchange should happen at commit time.  This all seems do-able\nwithout change in git; instead I'd add a few bash wrapper scripts\n(and see below for fsck and pack/prune).\n\nAnother setup is like the previous, but make the central repo have\nits own hidden object repository. You would push to it using the\nstandard git command.\n\nFinally, users could run git-fsck [with misleading output];\nthey could run git-prune{,-packed}, but these commands wouldn't\nbe able to delete anything.  If we don't want users to pack,\nthen ~git/.git/objects/pack would be writable only by ~git.\nSo basically, normal people wouldn't do the things in this paragraph.\n\nTo do meaningful and safe fsck/prune on the shared repository\nas ~git,  I'd add some scripting.  If you require all users'\nGIT_DIR's to look like /home/USER/*/.git , then you can get all\ntheir refs and do a meaningful fsck.  If not, you could do a fsck\n--unreachable as ~git and filter the result by date and/or type.\n(This sort of corresponds to abandoned changesets in perforce.)\nOnce you have an fsck method you like, its filtered output (i.e.,\n--unreachable objects you want to keep) can be fed to git-prune.\n\nCare would also be required with git-repack/git-prune-packed,\nbut it seems mostly addressable with scheduling.\n\nIf I proceed down this path,  I'd like to implement this procedure\nwithout any change in git's .c or .sh files.  It's clear this is a\nminority use and should not depend on anything being maintained\nfor it inside git.  I would write a few bash scripts and a README/HOWTO\nfor possible inclusion in contrib.\n\nBTW,\nhas anyone ever thought of writing an \"Administrator's Manual\" for git?\n\nThanks,\n-- \nDana L. How  danahow@gmail.com  +1 650 804 5991 cell\n"},{"id":"39226","messageId":"Pine.LNX.4.64.0704121203290.4061@woody.linux-foundation.org","threadId":"7591","inReplyTo":"56b7f5510704121132g3961060amb394978bb49093e6@mail.gmail.com","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-12T19:17:34Z","receivedAt":"2007-04-12T19:17:34Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 12 Apr 2007, Dana How wrote:\n> \n> These arguments all seem pretty convincing to me --\n> maybe the problem is that I'm not a \"*developer*\" right now.\n> Instead I'm part of a multi-developer *site*.\n\nYes.\n\nThe issues for hosting sites are very different from the issues of \nindividual developers having their own git repositories, and I agree 100% \nthat both alternates and shared object directories make tons of sense for \nhosting.\n\n> Below I talk about a possible way we could use git\n> without changing it (since I recognize this would be a minority usage\n> pattern).\n\nI hope it wouldn't even be a minority usage pattern. I am a firm believer \nthat distributed SCM's and git in particular makes a lot more sense for \nsource control hosting than CVS or SVN do. I'm really disappointed with \nthings like sourceforge, and part of the problem is literally that a \ncentralized SCM is really *fundamentally* wrong for a hosting entity. \n\nUsing a distributed SCM just makes _so_ much more sense for hosting \nprojects, and I've actually very much wanted to try to make sure that git \ncan help people who host things. \n\nIt's not my *own* primary use, but I think it's a very important usage \npattern, even though it's very different froma \"normal developer\" private \nsandbox case.\n\nSo I think your case is really very interesting. I'd love to help figure \nout how to help you guys with git, but because it's not how I personally \nwork, I can really just try to help when you actually hit a problem - \nyou'll have to figure out what your usage patterns actually are on your \nown ;)\n\nAnd btw, I think the shared object model really works very well, but I \nthink it has to be paired with some stricter rules than people who use \ntheir own repos tend to have. For example, end-point developers have \nbecome very used to rebasing and generally rewriting history (or just \nresetting to an older state), and that's something that works find in a \n\"local repository\" setup, but it's also the kinds of patterns that can \nreally screw you in a hosted and shared-object environment.\n\nAs to your two setups: I would suggest you go with the \"hidden\" shared \nversion (ie people use the remote access pull/push to a server, and the \n*server* uses a shared object repository for multiple repositories), \nrather than having a user-visible globally shared object directory. Even \nwith sticky bits and controlled group access etc, I think it's just safer \nto have that extra level of indirection.\n\n(Partly because a globally visible shared object directory also implies \nthat you'd use a networked filesystem, and I suspect a lot of developers \nwould actually be a lot happier having their own development repositories \non their own local disks, or at least some \"group disk\", rather than have \none big and performance-critical network share. Even if you use some \ncompetent NetApp box and a modern network filesystem, it's just one less \ncritical infrastructure piece that needs to be really beefy).\n\n\t\tLinus\n"},{"id":"39231","messageId":"e7bda7770704121423i3a984c65g2d21436833b5f0a8@mail.gmail.com","threadId":"7591","inReplyTo":"20070412005654.GQ21701@admingilde.org","subject":"Re: [PATCH 6/6] Teach core object handling functions about gitlinks","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2007-04-12T21:23:16Z","receivedAt":"2007-04-12T21:23:16Z","isPatch":true,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 4/12/07, Martin Waitz <tali@admingilde.org> wrote:\n\n> On Thu, Apr 12, 2007 at 02:42:43AM +0200, Torgil Svensson wrote:\n> > I guess this file could also cover the case where the superproject is\n> > only interested in a small subset of the subproject. For example if I\n> > only uses some header-files in a library and want\n> > \"/lib1/src/interface\" in the subproject end up as \"/includes/lib1\" in\n> > the superproject. Could single files be handled in a similar way?\n>\n> Conceptionally this information would have to be part of the\n> supermodule tree (after all it changes how your tree is set up).\n\nI agree. This could be included in the module config file which in\nturn is version-controlled.\n\n\n> I think it makes more sense to make users think about which part\n> of their tree can be reused and make them choose submodule boundaries\n> wisely so that the above partial-checkout is not needed.\n\nSometimes you can't control upstream projects the way you want it.\nAlso, splitting up projects for the potential need of future\nsuperprojects has several obvious disadvantages (multiple changelogs,\nversions etc). I don't see the subfolder checkout thing as a problem\nsince the core plumbing in Linus's implementation doesn't care what's\nbeneath the commit link. The subfolder checkout can \"easily\" be done\nin a porcelain.\n\nIt's more problematic if you want to cherry-pick individual files in a\nsubproject. Here, I think the tight connection between links and\ndirectories to be too restrictive. Why does a subproject commit-link\nhave to be represented as a folder?\n\n//Torgil\n"},{"id":"39273","messageId":"461F46B5.2020007@dawes.za.net","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704121203290.4061@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2007-04-13T09:00:37Z","receivedAt":"2007-04-13T09:00:37Z","isPatch":true,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Linus Torvalds wrote:\n> \n> Yes.\n> \n> The issues for hosting sites are very different from the issues of \n> individual developers having their own git repositories, and I agree 100% \n> that both alternates and shared object directories make tons of sense for \n> hosting.\n> \n>> Below I talk about a possible way we could use git\n>> without changing it (since I recognize this would be a minority usage\n>> pattern).\n> \n> I hope it wouldn't even be a minority usage pattern. I am a firm believer \n> that distributed SCM's and git in particular makes a lot more sense for \n> source control hosting than CVS or SVN do. I'm really disappointed with \n> things like sourceforge, and part of the problem is literally that a \n> centralized SCM is really *fundamentally* wrong for a hosting entity. \n> \n> Using a distributed SCM just makes _so_ much more sense for hosting \n> projects, and I've actually very much wanted to try to make sure that git \n> can help people who host things. \n\n\n> And btw, I think the shared object model really works very well, but I \n> think it has to be paired with some stricter rules than people who use \n> their own repos tend to have. For example, end-point developers have \n> become very used to rebasing and generally rewriting history (or just \n> resetting to an older state), and that's something that works find in a \n> \"local repository\" setup, but it's also the kinds of patterns that can \n> really screw you in a hosted and shared-object environment.\n> \n\nWould it not make sense for a hosting environment to say, if you are \nusing alternates, or shared object directories, then you need to include \n*all* the refs in *all* the projects if you ever do an fsck?\n\nI'm not sure how well git will scale in this case, although it just \nshould be a matter of how well git scales to dealing with a single \nproject with tens of thousands of refs/tags/etc. The only problem might \nbe in passing all those refs/tags to fsck in one go. STDIN, I guess?\n\nRogan\n"},{"id":"39290","messageId":"Pine.LNX.4.64.0704130816310.28042@woody.linux-foundation.org","threadId":"7591","inReplyTo":"461F46B5.2020007@dawes.za.net","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-13T15:23:57Z","receivedAt":"2007-04-13T15:23:57Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 13 Apr 2007, Rogan Dawes wrote:\n> \n> Would it not make sense for a hosting environment to say, if you are using\n> alternates, or shared object directories, then you need to include *all* the\n> refs in *all* the projects if you ever do an fsck?\n\nYes. And it shouldn't be hard to add support to do it. It's just not been \ndone.\n\nA lot of git programs already take refs on stdin, but fsck just doesn't do \nit (it can do it from the command line, but you'd run out of command line \nspace very quickly).\n\nMore natural would be to just list all the git repos by git repo pathname \n(and there, usually the command line probably *is* long enough), but \nsomebody would just have to do it. It's probably not very much code: just \niterate over each repo both when adding refs and when actually doing the \nfsck itself.\n\n> I'm not sure how well git will scale in this case, although it just should be\n> a matter of how well git scales to dealing with a single project with tens of\n> thousands of refs/tags/etc. The only problem might be in passing all those\n> refs/tags to fsck in one go. STDIN, I guess?\n\nFor a real shared object directory, passing the refs to stdin (and \nteaching fsck about a \"--stdin\" flag) would be consistent with what we do \nfor many other commands, so yes, that would work.\n\nHowever, fsck actually tends to want not just the refs, but actually \nthings like the index files and reflog files too, because those add other \nreachability info, which is why it's probably more natural to just give \nfsck the list of related repositories and let it figure them out.\n\nThat's also what you'd want to do for \"alternates\", since now there is no \nlonger a single object directory either, but multiple separate (but \nrelated) ones.\n\nSomebody would just have to write the code.. The basic rules are really \nall in \"git/builtin-fsck.c\": cmd_fsck(). Hint hint.\n\n\t\t\tLinus\n"},{"id":"39367","messageId":"56b7f5510704142350w66576704la62648c6f5d990f0@mail.gmail.com","threadId":"7591","inReplyTo":"Pine.LNX.4.64.0704121203290.4061@woody.linux-foundation.org","subject":"Re: [PATCH 5/6] Teach \"fsck\" not to follow subproject links","fromName":"Dana How","fromEmail":"danahow@gmail.com","sentAt":"2007-04-15T06:50:04Z","receivedAt":"2007-04-15T06:50:04Z","isPatch":true,"sender":{"key":"danahow@gmail.com","avatar":null},"body":"On 4/12/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> On Thu, 12 Apr 2007, Dana How wrote:\n> > These arguments all seem pretty convincing to me --\n> > maybe the problem is that I'm not a \"*developer*\" right now.\n> > Instead I'm part of a multi-developer *site*.\n> The issues for hosting sites are very different from the issues of\n> individual developers having their own git repositories, and I agree 100%\n> that both alternates and shared object directories make tons of sense for\n> hosting.\nFor clarity I should have written *office* instead of *site* to\ndescribe my situation,\nbut the mention of a NetApp below indicates no actual confusion occurred.\n\n> As to your two setups: I would suggest you go with the \"hidden\" shared\n> version (ie people use the remote access pull/push to a server, and the\n> *server* uses a shared object repository for multiple repositories),\n> rather than having a user-visible globally shared object directory. Even\n> with sticky bits and controlled group access etc, I think it's just safer\n> to have that extra level of indirection.\n>\n> (Partly because a globally visible shared object directory also implies\n> that you'd use a networked filesystem, and I suspect a lot of developers\n> would actually be a lot happier having their own development repositories\n> on their own local disks, or at least some \"group disk\", rather than have\n> one big and performance-critical network share. Even if you use some\n> competent NetApp box and a modern network filesystem, it's just one less\n> critical infrastructure piece that needs to be really beefy).\nWe did go down the local disk route, but after two significant losses of\nindividuals' work,  it was decreed that (perforce) work trees must be\non the NetApp.  So we already made the investment in beefiness --\nfor different reasons -- and I need to conform to these decisions for\nthe moment.\n\nAfter reliability, the other big criterion (especially with our\npenchant for large files)\nwill be speed. With perforce,  users now see submit={1 copy to server},\nsync={1 copy from server}.  In the short term I can't get away with changing\nthis to submit={copy working to indiv repo, copy indiv repo to shared repo}\nand sync={copy shared repo to indiv repo, copy indiv repo to working},\nbecause at first everyone will be trying to emulate what they did in perforce.\n\nSo probably I'll start out with either a very small testgroup,\nor one shared object repository with sticky/group tricks on the NetApp.\nOnce git's collaboration advantages are apparent,\nI'll switch to the hidden repository model which I prefer as well.\nAnd hopefully these collaboration advantages will also mean people\nwill commit more often and local disks can come back into favor --\nand then the \"extra\" local repo file copy operations will be less noticeable.\n\nIn any event, I have some scripting to do to learn more about our usage\npatterns and pushing our datasets throught git.  I also need to finish\nthe pack-splitting patch (after 64b index goes in). Finally,  before all that,\nI'll be out of the country for the next ~10 days...\n\nThanks,\n-- \nDana L. How  danahow@gmail.com  +1 650 804 5991 cell\n"},{"id":"39422","messageId":"20070415232123.GA13515@fieldses.org","threadId":"7591","inReplyTo":"alpine.LFD.0.98.0704101701030.28181@xanadu.home","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2007-04-15T23:21:24Z","receivedAt":"2007-04-15T23:21:24Z","isPatch":true,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Tue, Apr 10, 2007 at 05:03:13PM -0400, Nicolas Pitre wrote:\n> This is definitively good Documentation/howto/ material.\n\nThere's actually something similar already in \"modifying a single\ncommit\" in the \"user manual\":\n\nhttp://www.kernel.org/pub/software/scm/git/docs/user-manual.html#id276844\n\nBut it uses a throw-away branch instead of the detached head, and uses\nrebase --onto instead of rebasing and then --skip'ing.\n\n--b.\n"},{"id":"39423","messageId":"20070415232459.GB13515@fieldses.org","threadId":"7591","inReplyTo":"7vabxfp873.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 0/6] Initial subproject support (RFC?)","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2007-04-15T23:25:00Z","receivedAt":"2007-04-15T23:25:00Z","isPatch":true,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Wed, Apr 11, 2007 at 02:32:48AM -0700, Junio C Hamano wrote:\n> David Kågedal <davidk@lysator.liu.se> writes:\n> \n> > Junio C Hamano <junkio@cox.net> writes:\n> >\n> >> ...  This _will_ fail, but that is to be expected, as\n> >> we intend to replace that with what we just amended.  Just reset\n> >> it away and keep going.\n> >> \n> >> $ git reset --hard\n> >> $ git rebase --skip\n> >\n> > Wouldn't\n> >\n> > $ git rebase --onto HEAD lt/gitlink~3 lt/gitlink\n> >\n> > do the trick in one step?\n> \n> It is probably more Kosher, and I used to always do that, but it\n> is much longer to type,\n\nAlso remembering which commit you amended is a pain sometimes.  So I\nusually do\n\n\tgit tag base lt/gitlink~3\n\tgit checkout base\n\t... edit and amend ...\n\tgit rebase --onto HEAD base lt/gitlink\n\tgit tag -d base\n\nBut the trick of letting the rebase fail and skipping looks less\ncumbersome.\n\n--b.\n"}]}