{"thread":{"id":"3867","subject":"Recent unresolved issues","startedAt":"2006-04-14T09:31:36Z","lastAt":"2006-05-09T13:09:13Z","messageCount":81,"participants":["Junio C Hamano","Petr Baudis","sean","Carl Worth","Linus Torvalds","Johannes Schindelin","Jakub Narebski","Pavel Roskin","Daniel Barkalow","Martin Langhoff","Jeff King","Sergey Vlasov","Theodore Tso","David Woodhouse","Bertrand Jacquin","Nicolas Pitre"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"18638","messageId":"7v64lcqz9j.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":null,"subject":"Recent unresolved issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-14T09:31:36Z","receivedAt":"2006-04-14T09:31:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Here is a list of topics in the recent git traffic that I feel\ninadequately addressed.  I've commented on some of them to give\npeople a feel for what my priorities are.  Somebody might want\nto rehash the ones low on my priority list to conclusion with a\nconcrete proposal if they cared about them enough.  The list is\n*not* ordered in any way.\n\nAlso please add whatever I missed (or dismissed).  I am hoping\nthis will be a good basis for 1.4 to-do list.\n\n* Message-ID: <Pine.LNX.4.64.0604121828370.14565@g5.osdl.org>\n  Common option parsing (Linus Torvalds)\n\n* Message-ID: <Pine.LNX.4.64.0604050855080.2550@localhost.localdomain>\n  Binary diff output? (Nicolas Pitre)\n\n  I do not think this is needed for our primary audience (the\n  kernel project), but I am sure it would be helpful for some\n  other projects if we allowed them to exchange patches that\n  describe binary file changes via e-mail, so I am not\n  dismissing this.  Needs to wait \"option parsing\".\n\n* Message-ID: <Pine.LNX.4.64.0604111725590.14565@g5.osdl.org>\n  Colored diff? (Linus Torvalds)\n\n  I am not opposed to it, but I'd like to do that internally if\n  we go this route.  Needs to wait \"option parsing\".  Also\n  Message-ID: <3536.10.10.10.24.1114117965.squirrel@linux1> is\n  slightly related to this.\n\n* Message-ID: <7vek02ynif.fsf@assigned-by-dhcp.cox.net>\n  diff --with-raw, --with-stat? (me)\n\n  I think \"git diff\" can be internalized next, after \"option\n  parsing\" unification.  When that is done, --with-stat would\n  help internalize format-patch's process_one(), and it would be\n  trivial to do \"git log --pretty=format-patch master..next\".\n\n* #irc 2006-04-10\n  Shallow clones (Carl Worth).\n\n  The experiment last round did not work out very well, but as\n  existing repositories get bigger, and more projects being\n  migrated from foreign SCM systems, this would become a\n  must-have from would-be-nice-to-have.\n\n  I am beginning to think using \"graft\" to cauterize history\n  for this, while it technically would work, would not be so\n  helpful to users, so the design needs to be worked out again.\n\n* Message-ID: <E1FMH3o-0001B5-Dw@jdl.com>\n  git status does not distinguish contents changes and mode\n  changes; it just says \"modified\" (Jon Loeliger).\n\n  Unconditionally changing the status letter would break\n  Porcelains so we would need an extra option to do this.\n  An outline patch has been already prepared -- this perhaps has\n  to wait until we sort out the \"option parsing\" one.\n\n* Message-ID: <tnxmzf9sh7k.fsf@arm.com>\n  git could use diff3 instead of merge which is a wrapper around\n  diff3. (Catalin Marinas)\n\n  If having \"diff3\" is a lot more common than having \"merge\", I\n  do not have problem with this; \"merge\" being a wrapper to\n  \"diff3\", people who have been happy with the current code\n  would certainly have \"diff3\" installed so changing to \"diff3\"\n  would not break them.\n\n* Message-ID: <81b0412b0603020649u99a2035i3b8adde8ddce9410@mail.gmail.com>\n  Windows problems summary (Alex Riesen)\n\n  A good list to keep in mind.\n\n* Message-ID: <Pine.LNX.4.64.0604030730040.3781@g5.osdl.org>\n  Huge packfiles (Linus Torvalds)\n\n  Because I do not think asking users to break up packs to\n  manageable and mmap()able size is too much to ask, I would not\n  be advocating for updating the pack idx to 64-bit offset and\n  mmap()ing parts of a packfile, at least too strongly.\n\n  However, we currently lack tool support or recepe for users\n  with such a repository to easily break up packs.\n\n* Message-ID: <1143856098.3555.48.camel@dv>\n  Per branch property, esp. where to merge from (Pavel Roskin)\n\n  This involves user-level \"world model\" design, which is more\n  Porcelainish than Plumbing, and as people know I do not do\n  Porcelain well; interested parties need to come up with what\n  they want and how they want to use it.\n"},{"id":"18640","messageId":"20060414160246.GZ27689@pasky.or.cz","threadId":"3867","inReplyTo":"7v64lcqz9j.fsf@assigned-by-dhcp.cox.net","subject":"Re: Recent unresolved issues","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-04-14T16:02:46Z","receivedAt":"2006-04-14T16:02:46Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Apr 14, 2006 at 11:31:36AM CEST, I got a letter\nwhere Junio C Hamano <junkio@cox.net> said that...\n> Here is a list of topics in the recent git traffic that I feel\n> inadequately addressed.  I've commented on some of them to give\n> people a feel for what my priorities are.  Somebody might want\n> to rehash the ones low on my priority list to conclusion with a\n> concrete proposal if they cared about them enough.  The list is\n> *not* ordered in any way.\n\nNice summary!\n\n> * Message-ID: <tnxmzf9sh7k.fsf@arm.com>\n>   git could use diff3 instead of merge which is a wrapper around\n>   diff3. (Catalin Marinas)\n> \n>   If having \"diff3\" is a lot more common than having \"merge\", I\n>   do not have problem with this; \"merge\" being a wrapper to\n>   \"diff3\", people who have been happy with the current code\n>   would certainly have \"diff3\" installed so changing to \"diff3\"\n>   would not break them.\n\nI've decided to bite the bullet and made Cogito use diff3 instead of\nmerge as of now. Let's see if anybody complains...\n\n> * Message-ID: <1143856098.3555.48.camel@dv>\n>   Per branch property, esp. where to merge from (Pavel Roskin)\n> \n>   This involves user-level \"world model\" design, which is more\n>   Porcelainish than Plumbing, and as people know I do not do\n>   Porcelain well; interested parties need to come up with what\n>   they want and how they want to use it.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"18648","messageId":"BAYC1-PASMTP051B8E0C784B7630DEBE35AEC00@CEZ.ICE","threadId":"3867","inReplyTo":"7v64lcqz9j.fsf@assigned-by-dhcp.cox.net","subject":"Re: Recent unresolved issues","fromName":"sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2006-04-14T19:10:30Z","receivedAt":"2006-04-14T19:10:30Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Fri, 14 Apr 2006 02:31:36 -0700\nJunio C Hamano <junkio@cox.net> wrote:\n\n> * Message-ID: <Pine.LNX.4.64.0604111725590.14565@g5.osdl.org>\n>   Colored diff? (Linus Torvalds)\n> \n>   I am not opposed to it, but I'd like to do that internally if\n>   we go this route.  Needs to wait \"option parsing\".  Also\n>   Message-ID: <3536.10.10.10.24.1114117965.squirrel@linux1> is\n>   slightly related to this.\n\nMoving it internal sounds like a good idea.  Would you be open to\nincluding the GIT_DIFF_PAGER option now anyway?   It has utility\nbeyond just color diffs.\n\nSean\n"},{"id":"18649","messageId":"20060414192448.GB27689@pasky.or.cz","threadId":"3867","inReplyTo":"7v64lcqz9j.fsf@assigned-by-dhcp.cox.net","subject":"Re: Recent unresolved issues","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-04-14T19:24:48Z","receivedAt":"2006-04-14T19:24:48Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Apr 14, 2006 at 11:31:36AM CEST, I got a letter\nwhere Junio C Hamano <junkio@cox.net> said that...\n> * Message-ID: <Pine.LNX.4.64.0604111725590.14565@g5.osdl.org>\n>   Colored diff? (Linus Torvalds)\n> \n>   I am not opposed to it, but I'd like to do that internally if\n>   we go this route.  Needs to wait \"option parsing\".  Also\n>   Message-ID: <3536.10.10.10.24.1114117965.squirrel@linux1> is\n>   slightly related to this.\n\nIt might be worthwhile to make Git and Cogito compatible if you offer\ncolors customization. Cogito lets the user customize the colors through\nthe $CG_COLORS variable (see cg-diff(1)).\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"18660","messageId":"87irpb7oma.wl%cworth@cworth.org","threadId":"3867","inReplyTo":"7v64lcqz9j.fsf@assigned-by-dhcp.cox.net","subject":"Re: Recent unresolved issues: shallow clone","fromName":"Carl Worth","fromEmail":"cworth@cworth.org","sentAt":"2006-04-14T22:56:29Z","receivedAt":"2006-04-14T22:56:29Z","isPatch":false,"sender":{"key":"cworth@cworth.org","avatar":"https://gravatar.com/avatar/3746dc28cde609bdbd7f939058356e7e2bbd16d21e32274df0725eb3d998bc5b?d=mp&s=160"},"body":"On Fri, 14 Apr 2006 02:31:36 -0700, Junio C Hamano wrote:\n>   Shallow clones (Carl Worth).\n> \n>   The experiment last round did not work out very well, but as\n>   existing repositories get bigger, and more projects being\n>   migrated from foreign SCM systems, this would become a\n>   must-have from would-be-nice-to-have.\n> \n>   I am beginning to think using \"graft\" to cauterize history\n>   for this, while it technically would work, would not be so\n>   helpful to users, so the design needs to be worked out again.\n\nAs context, here is some of what you mentioned in IRC:\n\n>>\tSuppose you have this:\n>>\n>>\tA---B---C\n>>\t \\       \\ \n>>\t  D---E---F---G\n>>\t \n>>\tand you made a shallow clone of C (because that is where the\n>>\tupstream master was when you made that clone).  Then the\n>>\tupstream updated the master branch tip to G.\n>>\n>>\tThe next update from upstream to your shallow clone would break.\n>>\tThe upstream says: I have G at master.\n>>\tYou say: I want G then.  By the way, I have C.\n>>\n>>\tWhat it means to tell the other end \"I have X\" is to promise\n>>\tthat you have X and _everything_ behind it.  So the upstream\n>>\twould send objects necessary to complete D, E, F and G for\n>>\t\"somebody who already have A and B\".  As a consequence, you\n>>\twould not see A nor B.\n>>\n>>\tEven if the only thing you are interested in is to be in sync\n>>\twith the tip of the upstream, you can end up with an\n>>\tincomplete tree for G, if some of the blobs or trees contained\n>>\tin G already exist in A or B.  They are not sent -- because\n>>\tyou told the upstream that you have everything necessary to\n>>\tget to C.\n\nSo that's an argument against using a cauterizing graft for the\nshallow clone of C. It definitely confuses the existing protocol to\nsay \"I have C\" if I have only a cauterized C, (its tree only, but none\nof the commits that should be backing C).\n\nI also read over some of your discussion of extending the protocol\nwith a new \"shallow\" extension.\n\nI'm wondering if the shallow clone support couldn't be achieved\nthrough a simpler tweak to the protocol semantics, (and no change to\nprotocol syntax), that would avoid the problem above. Specifically,\nfor shallow stuff, could we just do the same \"want\" and \"have\"\nconversation with tree objects rather than commit objects?\n\nSo, in the scenario above, the original shallow clone of C would be:\n\n\tWant C->tree, have nothing.\n\nand the later shallow update to G would be:\n\n\tWant G->tree, have C->tree\n\nA final step of a shallow clone would then require creating a new\nparent-less commit object so that there's something to point refs/head\nat, (or maybe rather than being parentless, they could be chained\ntogether with each update?).\n\nI admit that this would result in a rather atypical kind of\nrepository, but it would contain plenty of valid trees and blobs, so\nit should conceptually be fairly easy to promote such a thing to a\nfull repository.\n\nBut, even without any tool support for promotion, the ability to do\nshallow clone and shallow updates would still provide a useful\ncapability [*].\n\n-Carl\n\n[*] For reference, what I'm looking for here is a way to justify\nproviding git support for jhbuild, which is a tool used by testers of\nGNOME and other software to efficiently track the latest development\nof an arbitrarily large number of packages. It's currently primarily a\nCVS-based thing. Switching to git would be a huge win for the\nincremental updates, but would currently cause quite a hit for the\nfirst clone.\n"},{"id":"18661","messageId":"Pine.LNX.4.64.0604141637230.3701@g5.osdl.org","threadId":"3867","inReplyTo":"7v64lcqz9j.fsf@assigned-by-dhcp.cox.net","subject":"Re: Recent unresolved issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-04-14T23:52:13Z","receivedAt":"2006-04-14T23:52:13Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 14 Apr 2006, Junio C Hamano wrote:\n> \n> * Message-ID: <Pine.LNX.4.64.0604121828370.14565@g5.osdl.org>\n>   Common option parsing (Linus Torvalds)\n\nOk, here's a first cut at starting this.\n\nThis basically does a few things that are sadly somewhat interdependent, \nand nontrivial to split out\n\n - get rid of \"struct log_tree_opt\"\n\n   The fields in \"log_tree_opt\" are moved into \"struct rev_info\", and all \n   users of log_tree_opt are changed to use the rev_info struct instead.\n\n - add the parsing for the log_tree_opt arguments to \"setup_revision()\"\n\n - make setup_revision set a flag (revs->diff) if the diff-related \n   arguments were used. This allows \"git log\" to decide whether it wants \n   to show diffs or not.\n\n - make setup_revision() also initialize the diffopt part of rev_info \n   (which we had from before, but we just didn't initialize it)\n\n - make setup_revision() do all the \"finishing touches\" on it all (it will \n   do the proper flag combination logic, and call \"diff_setup_done()\")\n\nNow, that was the easy and straightforward part.\n\nThe slightly more involved part is that some of the programs that want to \nuse the new-and-improved rev_info parsing don't actually want _commits_, \nthey may want tree'ish arguments instead. That meant that I had to change \nsetup_revision() to parse the arguments not into the \"revs->commits\" list, \nbut into the \"revs->pending_objects\" list.\n\nThen, when we do \"prepare_revision_walk()\", we walk that list, and create \nthe sorted commit list from there. \n\nThis actually cleaned some stuff up, but it's the less obvious part of the \npatch, and re-organized the \"revision.c\" logic somewhat. It actually paves \nthe way for splitting argument parsing _entirely_ out of \"revision.c\", \nsince now the argument parsing really is totally independent of the commit \nwalking: that didn't use to be true, since there was lots of overlap with \nget_commit_reference() handling etc, now the _only_ overlap is the shared \n(and trivial) \"add_pending_object()\" thing.\n\nHowever, I didn't do that file split, just because I wanted the diff \nitself to be smaller, and show the actual changes more clearly. If this \ngets accepted, I'll do further cleanups then - that includes the file \nsplit, but also using the new infrastructure to do a nicer \"git diff\" etc.\n\nEven in this form, it actually ends up removing more lines than it adds.\n\nIt's nice to note how simple and straightforward this makes the built-in \n\"git log\" command, even though it continues to support all the diff flags \ntoo. It doesn't get much simpler that this.\n\nI think this is worth merging soonish, because it does allow for future \ncleanup and even more sharing of code. However, it obviously touches \n\"revision.c\", which is subtle. I've tested that it passes all the tests we \nhave, and it passes my \"looks sane\" detector, but somebody else should \nalso give it a good look-over.\n\nSigned-off-by: Linus Torvalds <torvalds@osdl.org>\n---\n\n diff-tree.c |   91 ++++++++++++++++-------------------\n git.c       |   68 ++------------------------\n log-tree.c  |   60 ++---------------------\n log-tree.h  |   22 ++------\n revision.c  |  155 ++++++++++++++++++++++++++++++++++++++++++++++++-----------\n revision.h  |   18 +++++++\n 6 files changed, 202 insertions(+), 212 deletions(-)\n\ndiff --git a/diff-tree.c b/diff-tree.c\nindex 2b79dd0..54157e4 100644\n--- a/diff-tree.c\n+++ b/diff-tree.c\n@@ -3,7 +3,7 @@ #include \"diff.h\"\n #include \"commit.h\"\n #include \"log-tree.h\"\n \n-static struct log_tree_opt log_tree_opt;\n+static struct rev_info log_tree_opt;\n \n static int diff_tree_commit_sha1(const unsigned char *sha1)\n {\n@@ -62,66 +62,55 @@ int main(int argc, const char **argv)\n {\n \tint nr_sha1;\n \tchar line[1000];\n-\tunsigned char sha1[2][20];\n-\tconst char *prefix = setup_git_directory();\n-\tstatic struct log_tree_opt *opt = &log_tree_opt;\n+\tstruct object *tree1, *tree2;\n+\tstatic struct rev_info *opt = &log_tree_opt;\n+\tstruct object_list *list;\n \tint read_stdin = 0;\n \n \tgit_config(git_diff_config);\n \tnr_sha1 = 0;\n-\tinit_log_tree_opt(opt);\n+\targc = setup_revisions(argc, argv, opt, NULL);\n \n-\tfor (;;) {\n-\t\tint opt_cnt;\n-\t\tconst char *arg;\n+\twhile (--argc > 0) {\n+\t\tconst char *arg = *++argv;\n \n-\t\targv++;\n-\t\targc--;\n-\t\targ = *argv;\n-\t\tif (!arg)\n-\t\t\tbreak;\n-\n-\t\tif (*arg != '-') {\n-\t\t\tif (nr_sha1 < 2 && !get_sha1(arg, sha1[nr_sha1])) {\n-\t\t\t\tnr_sha1++;\n-\t\t\t\tcontinue;\n-\t\t\t}\n-\t\t\tbreak;\n-\t\t}\n-\n-\t\topt_cnt = log_tree_opt_parse(opt, argv, argc);\n-\t\tif (opt_cnt < 0)\n-\t\t\tusage(diff_tree_usage);\n-\t\telse if (opt_cnt) {\n-\t\t\targv += opt_cnt - 1;\n-\t\t\targc -= opt_cnt - 1;\n-\t\t\tcontinue;\n-\t\t}\n-\n-\t\tif (!strcmp(arg, \"--\")) {\n-\t\t\targv++;\n-\t\t\targc--;\n-\t\t\tbreak;\n-\t\t}\n \t\tif (!strcmp(arg, \"--stdin\")) {\n \t\t\tread_stdin = 1;\n \t\t\tcontinue;\n \t\t}\n \t\tusage(diff_tree_usage);\n \t}\n-\n-\tif (opt->combine_merges)\n-\t\topt->ignore_merges = 0;\n-\n-\t/* We can only do dense combined merges with diff output */\n-\tif (opt->dense_combined_merges)\n-\t\topt->diffopt.output_format = DIFF_FORMAT_PATCH;\n-\n-\tif (opt->diffopt.output_format == DIFF_FORMAT_PATCH)\n-\t\topt->diffopt.recursive = 1;\n \n-\tdiff_tree_setup_paths(get_pathspec(prefix, argv), opt);\n-\tdiff_setup_done(&opt->diffopt);\n+\t/*\n+\t * NOTE! \"setup_revisions()\" will have inserted the revisions\n+\t * it parsed in reverse order. So if you do\n+\t *\n+\t *\tgit-diff-tree a b\n+\t *\n+\t * the commit list will be \"b\" -> \"a\" -> NULL, so we reverse\n+\t * the order of the objects if the first one is not marked\n+\t * UNINTERESTING.\n+\t */\n+\tnr_sha1 = 0;\n+\tlist = opt->pending_objects;\n+\tif (list) {\n+\t\tnr_sha1++;\n+\t\ttree1 = list->item;\n+\t\tlist = list->next;\n+\t\tif (list) {\n+\t\t\tnr_sha1++;\n+\t\t\ttree2 = tree1;\n+\t\t\ttree1 = list->item;\n+\t\t\tif (list->next)\n+\t\t\t\tusage(diff_tree_usage);\n+\t\t\t/* Switch them around if the second one was uninteresting.. */\n+\t\t\tif (tree2->flags & UNINTERESTING) {\n+\t\t\t\tstruct object *tmp = tree2;\n+\t\t\t\ttree2 = tree1;\n+\t\t\t\ttree1 = tmp;\n+\t\t\t}\n+\t\t}\n+\t}\n \n \tswitch (nr_sha1) {\n \tcase 0:\n@@ -129,10 +118,12 @@ int main(int argc, const char **argv)\n \t\t\tusage(diff_tree_usage);\n \t\tbreak;\n \tcase 1:\n-\t\tdiff_tree_commit_sha1(sha1[0]);\n+\t\tdiff_tree_commit_sha1(tree1->sha1);\n \t\tbreak;\n \tcase 2:\n-\t\tdiff_tree_sha1(sha1[0], sha1[1], \"\", &opt->diffopt);\n+\t\tdiff_tree_sha1(tree1->sha1,\n+\t\t\t       tree2->sha1,\n+\t\t\t       \"\", &opt->diffopt);\n \t\tlog_tree_diff_flush(opt);\n \t\tbreak;\n \t}\ndiff --git a/git.c b/git.c\nindex 78ed403..e8d1fcc 100644\n--- a/git.c\n+++ b/git.c\n@@ -287,74 +287,18 @@ static int cmd_log(int argc, const char \n \tint abbrev = DEFAULT_ABBREV;\n \tint abbrev_commit = 0;\n \tconst char *commit_prefix = \"commit \";\n-\tstruct log_tree_opt opt;\n \tint shown = 0;\n-\tint do_diff = 0;\n-\tint full_diff = 0;\n \n-\tinit_log_tree_opt(&opt);\n \targc = setup_revisions(argc, argv, &rev, \"HEAD\");\n-\twhile (1 < argc) {\n-\t\tconst char *arg = argv[1];\n-\t\tif (!strncmp(arg, \"--pretty\", 8)) {\n-\t\t\tcommit_format = get_commit_format(arg + 8);\n-\t\t\tif (commit_format == CMIT_FMT_ONELINE)\n-\t\t\t\tcommit_prefix = \"\";\n-\t\t}\n-\t\telse if (!strcmp(arg, \"--no-abbrev\")) {\n-\t\t\tabbrev = 0;\n-\t\t}\n-\t\telse if (!strcmp(arg, \"--abbrev\")) {\n-\t\t\tabbrev = DEFAULT_ABBREV;\n-\t\t}\n-\t\telse if (!strcmp(arg, \"--abbrev-commit\")) {\n-\t\t\tabbrev_commit = 1;\n-\t\t}\n-\t\telse if (!strncmp(arg, \"--abbrev=\", 9)) {\n-\t\t\tabbrev = strtoul(arg + 9, NULL, 10);\n-\t\t\tif (abbrev && abbrev < MINIMUM_ABBREV)\n-\t\t\t\tabbrev = MINIMUM_ABBREV;\n-\t\t\telse if (40 < abbrev)\n-\t\t\t\tabbrev = 40;\n-\t\t}\n-\t\telse if (!strcmp(arg, \"--full-diff\")) {\n-\t\t\tdo_diff = 1;\n-\t\t\tfull_diff = 1;\n-\t\t}\n-\t\telse {\n-\t\t\tint cnt = log_tree_opt_parse(&opt, argv+1, argc-1);\n-\t\t\tif (0 < cnt) {\n-\t\t\t\tdo_diff = 1;\n-\t\t\t\targv += cnt;\n-\t\t\t\targc -= cnt;\n-\t\t\t\tcontinue;\n-\t\t\t}\n-\t\t\tdie(\"unrecognized argument: %s\", arg);\n-\t\t}\n+\tif (argc > 1)\n+\t\tdie(\"unrecognized argument: %s\", argv[1]);\n \n-\t\targc--; argv++;\n-\t}\n-\n-\tif (do_diff) {\n-\t\topt.diffopt.abbrev = abbrev;\n-\t\topt.verbose_header = 0;\n-\t\topt.always_show_header = 0;\n-\t\topt.no_commit_id = 1;\n-\t\tif (opt.combine_merges)\n-\t\t\topt.ignore_merges = 0;\n-\t\tif (opt.dense_combined_merges)\n-\t\t\topt.diffopt.output_format = DIFF_FORMAT_PATCH;\n-\t\tif (opt.diffopt.output_format == DIFF_FORMAT_PATCH)\n-\t\t\topt.diffopt.recursive = 1;\n-\t\tif (!full_diff && rev.prune_data)\n-\t\t\tdiff_tree_setup_paths(rev.prune_data, &opt.diffopt);\n-\t\tdiff_setup_done(&opt.diffopt);\n-\t}\n+\trev.no_commit_id = 1;\n \n \tprepare_revision_walk(&rev);\n \tsetup_pager();\n \twhile ((commit = get_revision(&rev)) != NULL) {\n-\t\tif (shown && do_diff && commit_format != CMIT_FMT_ONELINE)\n+\t\tif (shown && rev.diff && commit_format != CMIT_FMT_ONELINE)\n \t\t\tputchar('\\n');\n \t\tfputs(commit_prefix, stdout);\n \t\tif (abbrev_commit && abbrev)\n@@ -388,8 +332,8 @@ static int cmd_log(int argc, const char \n \t\tpretty_print_commit(commit_format, commit, ~0, buf,\n \t\t\t\t    LOGSIZE, abbrev);\n \t\tprintf(\"%s\\n\", buf);\n-\t\tif (do_diff)\n-\t\t\tlog_tree_commit(&opt, commit);\n+\t\tif (rev.diff)\n+\t\t\tlog_tree_commit(&rev, commit);\n \t\tshown = 1;\n \t\tfree(commit->buffer);\n \t\tcommit->buffer = NULL;\ndiff --git a/log-tree.c b/log-tree.c\nindex 3d40482..04a68e0 100644\n--- a/log-tree.c\n+++ b/log-tree.c\n@@ -3,58 +3,8 @@ #include \"diff.h\"\n #include \"commit.h\"\n #include \"log-tree.h\"\n \n-void init_log_tree_opt(struct log_tree_opt *opt)\n+int log_tree_diff_flush(struct rev_info *opt)\n {\n-\tmemset(opt, 0, sizeof *opt);\n-\topt->ignore_merges = 1;\n-\topt->header_prefix = \"\";\n-\topt->commit_format = CMIT_FMT_RAW;\n-\tdiff_setup(&opt->diffopt);\n-}\n-\n-int log_tree_opt_parse(struct log_tree_opt *opt, const char **av, int ac)\n-{\n-\tconst char *arg;\n-\tint cnt = diff_opt_parse(&opt->diffopt, av, ac);\n-\tif (0 < cnt)\n-\t\treturn cnt;\n-\targ = *av;\n-\tif (!strcmp(arg, \"-r\"))\n-\t\topt->diffopt.recursive = 1;\n-\telse if (!strcmp(arg, \"-t\")) {\n-\t\topt->diffopt.recursive = 1;\n-\t\topt->diffopt.tree_in_recursive = 1;\n-\t}\n-\telse if (!strcmp(arg, \"-m\"))\n-\t\topt->ignore_merges = 0;\n-\telse if (!strcmp(arg, \"-c\"))\n-\t\topt->combine_merges = 1;\n-\telse if (!strcmp(arg, \"--cc\")) {\n-\t\topt->dense_combined_merges = 1;\n-\t\topt->combine_merges = 1;\n-\t}\n-\telse if (!strcmp(arg, \"-v\")) {\n-\t\topt->verbose_header = 1;\n-\t\topt->header_prefix = \"diff-tree \";\n-\t}\n-\telse if (!strncmp(arg, \"--pretty\", 8)) {\n-\t\topt->verbose_header = 1;\n-\t\topt->header_prefix = \"diff-tree \";\n-\t\topt->commit_format = get_commit_format(arg+8);\n-\t}\n-\telse if (!strcmp(arg, \"--root\"))\n-\t\topt->show_root_diff = 1;\n-\telse if (!strcmp(arg, \"--no-commit-id\"))\n-\t\topt->no_commit_id = 1;\n-\telse if (!strcmp(arg, \"--always\"))\n-\t\topt->always_show_header = 1;\n-\telse\n-\t\treturn 0;\n-\treturn 1;\n-}\n-\n-int log_tree_diff_flush(struct log_tree_opt *opt)\n-{\n \tdiffcore_std(&opt->diffopt);\n \tif (diff_queue_is_empty()) {\n \t\tint saved_fmt = opt->diffopt.output_format;\n@@ -73,7 +23,7 @@ int log_tree_diff_flush(struct log_tree_\n \treturn 1;\n }\n \n-static int diff_root_tree(struct log_tree_opt *opt,\n+static int diff_root_tree(struct rev_info *opt,\n \t\t\t  const unsigned char *new, const char *base)\n {\n \tint retval;\n@@ -93,7 +43,7 @@ static int diff_root_tree(struct log_tre\n \treturn retval;\n }\n \n-static const char *generate_header(struct log_tree_opt *opt,\n+static const char *generate_header(struct rev_info *opt,\n \t\t\t\t   const unsigned char *commit_sha1,\n \t\t\t\t   const unsigned char *parent_sha1,\n \t\t\t\t   const struct commit *commit)\n@@ -129,7 +79,7 @@ static const char *generate_header(struc\n \treturn this_header;\n }\n \n-static int do_diff_combined(struct log_tree_opt *opt, struct commit *commit)\n+static int do_diff_combined(struct rev_info *opt, struct commit *commit)\n {\n \tunsigned const char *sha1 = commit->object.sha1;\n \n@@ -142,7 +92,7 @@ static int do_diff_combined(struct log_t\n \treturn 0;\n }\n \n-int log_tree_commit(struct log_tree_opt *opt, struct commit *commit)\n+int log_tree_commit(struct rev_info *opt, struct commit *commit)\n {\n \tstruct commit_list *parents;\n \tunsigned const char *sha1 = commit->object.sha1;\ndiff --git a/log-tree.h b/log-tree.h\nindex da166c6..91a909b 100644\n--- a/log-tree.h\n+++ b/log-tree.h\n@@ -1,23 +1,11 @@\n #ifndef LOG_TREE_H\n #define LOG_TREE_H\n \n-struct log_tree_opt {\n-\tstruct diff_options diffopt;\n-\tint show_root_diff;\n-\tint no_commit_id;\n-\tint verbose_header;\n-\tint ignore_merges;\n-\tint combine_merges;\n-\tint dense_combined_merges;\n-\tint always_show_header;\n-\tconst char *header_prefix;\n-\tconst char *header;\n-\tenum cmit_fmt commit_format;\n-};\n+#include \"revision.h\"\n \n-void init_log_tree_opt(struct log_tree_opt *);\n-int log_tree_diff_flush(struct log_tree_opt *);\n-int log_tree_commit(struct log_tree_opt *, struct commit *);\n-int log_tree_opt_parse(struct log_tree_opt *, const char **, int);\n+void init_log_tree_opt(struct rev_info *);\n+int log_tree_diff_flush(struct rev_info *);\n+int log_tree_commit(struct rev_info *, struct commit *);\n+int log_tree_opt_parse(struct rev_info *, const char **, int);\n \n #endif\ndiff --git a/revision.c b/revision.c\nindex 0505f3f..99077af 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -116,21 +116,27 @@ static void add_pending_object(struct re\n \tadd_object(obj, &revs->pending_objects, NULL, name);\n }\n \n-static struct commit *get_commit_reference(struct rev_info *revs, const char *name, const unsigned char *sha1, unsigned int flags)\n+static struct object *get_reference(struct rev_info *revs, const char *name, const unsigned char *sha1, unsigned int flags)\n {\n \tstruct object *object;\n \n \tobject = parse_object(sha1);\n \tif (!object)\n \t\tdie(\"bad object %s\", name);\n+\tobject->flags |= flags;\n+\treturn object;\n+}\n+\n+static struct commit *handle_commit(struct rev_info *revs, struct object *object, const char *name)\n+{\n+\tunsigned long flags = object->flags;\n \n \t/*\n \t * Tag object? Look what it points to..\n \t */\n \twhile (object->type == tag_type) {\n \t\tstruct tag *tag = (struct tag *) object;\n-\t\tobject->flags |= flags;\n-\t\tif (revs->tag_objects && !(object->flags & UNINTERESTING))\n+\t\tif (revs->tag_objects && !(flags & UNINTERESTING))\n \t\t\tadd_pending_object(revs, object, tag->tag);\n \t\tobject = parse_object(tag->tagged->sha1);\n \t\tif (!object)\n@@ -143,7 +149,6 @@ static struct commit *get_commit_referen\n \t */\n \tif (object->type == commit_type) {\n \t\tstruct commit *commit = (struct commit *)object;\n-\t\tobject->flags |= flags;\n \t\tif (parse_commit(commit) < 0)\n \t\t\tdie(\"unable to parse commit %s\", name);\n \t\tif (flags & UNINTERESTING) {\n@@ -449,14 +454,6 @@ static void limit_list(struct rev_info *\n \t\t}\n \t}\n \trevs->commits = newlist;\n-}\n-\n-static void add_one_commit(struct commit *commit, struct rev_info *revs)\n-{\n-\tif (!commit || (commit->object.flags & SEEN))\n-\t\treturn;\n-\tcommit->object.flags |= SEEN;\n-\tcommit_list_insert(commit, &revs->commits);\n }\n \n static int all_flags;\n@@ -464,8 +461,8 @@ static struct rev_info *all_revs;\n \n static int handle_one_ref(const char *path, const unsigned char *sha1)\n {\n-\tstruct commit *commit = get_commit_reference(all_revs, path, sha1, all_flags);\n-\tadd_one_commit(commit, all_revs);\n+\tstruct object *object = get_reference(all_revs, path, sha1, all_flags);\n+\tadd_pending_object(all_revs, object, \"\");\n \treturn 0;\n }\n \n@@ -494,6 +491,11 @@ void init_revisions(struct rev_info *rev\n \n \trevs->topo_setter = topo_sort_default_setter;\n \trevs->topo_getter = topo_sort_default_getter;\n+\n+\trevs->header_prefix = \"\";\n+\trevs->commit_format = CMIT_FMT_RAW;\n+\n+\tdiff_setup(&revs->diffopt);\n }\n \n /*\n@@ -526,13 +528,14 @@ int setup_revisions(int argc, const char\n \n \tflags = 0;\n \tfor (i = 1; i < argc; i++) {\n-\t\tstruct commit *commit;\n+\t\tstruct object *object;\n \t\tconst char *arg = argv[i];\n \t\tunsigned char sha1[20];\n \t\tchar *dotdot;\n \t\tint local_flags;\n \n \t\tif (*arg == '-') {\n+\t\t\tint opts;\n \t\t\tif (!strncmp(arg, \"--max-count=\", 12)) {\n \t\t\t\trevs->max_count = atoi(arg + 12);\n \t\t\t\tcontinue;\n@@ -638,6 +641,78 @@ int setup_revisions(int argc, const char\n \t\t\t}\n \t\t\tif (!strcmp(arg, \"--unpacked\")) {\n \t\t\t\trevs->unpacked = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"-r\")) {\n+\t\t\t\trevs->diff = 1;\n+\t\t\t\trevs->diffopt.recursive = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"-t\")) {\n+\t\t\t\trevs->diff = 1;\n+\t\t\t\trevs->diffopt.recursive = 1;\n+\t\t\t\trevs->diffopt.tree_in_recursive = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"-m\")) {\n+\t\t\t\trevs->ignore_merges = 0;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"-c\")) {\n+\t\t\t\trevs->diff = 1;\n+\t\t\t\trevs->combine_merges = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"--cc\")) {\n+\t\t\t\trevs->diff = 1;\n+\t\t\t\trevs->dense_combined_merges = 1;\n+\t\t\t\trevs->combine_merges = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"-v\")) {\n+\t\t\t\trevs->verbose_header = 1;\n+\t\t\t\trevs->header_prefix = \"diff-tree \";\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strncmp(arg, \"--pretty\", 8)) {\n+\t\t\t\trevs->verbose_header = 1;\n+\t\t\t\trevs->header_prefix = \"diff-tree \";\n+\t\t\t\trevs->commit_format = get_commit_format(arg+8);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"--root\")) {\n+\t\t\t\trevs->show_root_diff = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"--no-commit-id\")) {\n+\t\t\t\trevs->no_commit_id = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"--always\")) {\n+\t\t\t\trevs->always_show_header = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"--no-abbrev\")) {\n+\t\t\t\trevs->abbrev = 0;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"--abbrev\")) {\n+\t\t\t\trevs->abbrev = DEFAULT_ABBREV;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"--abbrev-commit\")) {\n+\t\t\t\trevs->abbrev_commit = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"--full-diff\")) {\n+\t\t\t\trevs->diff = 1;\n+\t\t\t\trevs->full_diff = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\topts = diff_opt_parse(&revs->diffopt, argv+i, argc-i);\n+\t\t\tif (opts > 0) {\n+\t\t\t\trevs->diff = 1;\n+\t\t\t\ti += opts - 1;\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\t*unrecognized++ = arg;\n@@ -656,15 +731,15 @@ int setup_revisions(int argc, const char\n \t\t\t\tthis = \"HEAD\";\n \t\t\tif (!get_sha1(this, from_sha1) &&\n \t\t\t    !get_sha1(next, sha1)) {\n-\t\t\t\tstruct commit *exclude;\n-\t\t\t\tstruct commit *include;\n+\t\t\t\tstruct object *exclude;\n+\t\t\t\tstruct object *include;\n \n-\t\t\t\texclude = get_commit_reference(revs, this, from_sha1, flags ^ UNINTERESTING);\n-\t\t\t\tinclude = get_commit_reference(revs, next, sha1, flags);\n+\t\t\t\texclude = get_reference(revs, this, from_sha1, flags ^ UNINTERESTING);\n+\t\t\t\tinclude = get_reference(revs, next, sha1, flags);\n \t\t\t\tif (!exclude || !include)\n \t\t\t\t\tdie(\"Invalid revision range %s..%s\", arg, next);\n-\t\t\t\tadd_one_commit(exclude, revs);\n-\t\t\t\tadd_one_commit(include, revs);\n+\t\t\t\tadd_pending_object(revs, exclude, this);\n+\t\t\t\tadd_pending_object(revs, include, next);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\t*dotdot = '.';\n@@ -689,16 +764,16 @@ int setup_revisions(int argc, const char\n \t\t\trevs->prune_data = get_pathspec(revs->prefix, argv + i);\n \t\t\tbreak;\n \t\t}\n-\t\tcommit = get_commit_reference(revs, arg, sha1, flags ^ local_flags);\n-\t\tadd_one_commit(commit, revs);\n+\t\tobject = get_reference(revs, arg, sha1, flags ^ local_flags);\n+\t\tadd_pending_object(revs, object, arg);\n \t}\n-\tif (def && !revs->commits) {\n+\tif (def && !revs->pending_objects) {\n \t\tunsigned char sha1[20];\n-\t\tstruct commit *commit;\n+\t\tstruct object *object;\n \t\tif (get_sha1(def, sha1) < 0)\n \t\t\tdie(\"bad default revision '%s'\", def);\n-\t\tcommit = get_commit_reference(revs, def, sha1, 0);\n-\t\tadd_one_commit(commit, revs);\n+\t\tobject = get_reference(revs, def, sha1, 0);\n+\t\tadd_pending_object(revs, object, def);\n \t}\n \n \tif (revs->topo_order || revs->unpacked)\n@@ -708,13 +783,37 @@ int setup_revisions(int argc, const char\n \t\tdiff_tree_setup_paths(revs->prune_data, &revs->diffopt);\n \t\trevs->prune_fn = try_to_simplify_commit;\n \t}\n+\tif (revs->combine_merges) {\n+\t\trevs->ignore_merges = 0;\n+\t\tif (revs->dense_combined_merges)\n+\t\t\trevs->diffopt.output_format = DIFF_FORMAT_PATCH;\n+\t}\n+\tif (revs->diffopt.output_format == DIFF_FORMAT_PATCH)\n+\t\trevs->diffopt.recursive = 1;\n+\tif (!revs->full_diff && revs->prune_data)\n+\t\tdiff_tree_setup_paths(revs->prune_data, &revs->diffopt);\n+\tdiff_setup_done(&revs->diffopt);\n \n \treturn left;\n }\n \n void prepare_revision_walk(struct rev_info *revs)\n {\n-\tsort_by_date(&revs->commits);\n+\tstruct object_list *list;\n+\n+\tlist = revs->pending_objects;\n+\trevs->pending_objects = NULL;\n+\twhile (list) {\n+\t\tstruct commit *commit = handle_commit(revs, list->item, list->name);\n+\t\tif (commit) {\n+\t\t\tif (!(commit->object.flags & SEEN)) {\n+\t\t\t\tcommit->object.flags |= SEEN;\n+\t\t\t\tinsert_by_date(commit, &revs->commits);\n+\t\t\t}\n+\t\t}\n+\t\tlist = list->next;\n+\t}\n+\n \tif (revs->limited)\n \t\tlimit_list(revs);\n \tif (revs->topo_order)\ndiff --git a/revision.h b/revision.h\nindex 8970b57..9a45986 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -38,6 +38,24 @@ struct rev_info {\n \t\t\tboundary:1,\n \t\t\tparents:1;\n \n+\t/* Diff flags */\n+\tunsigned int\tdiff:1,\n+\t\t\tfull_diff:1,\n+\t\t\tshow_root_diff:1,\n+\t\t\tno_commit_id:1,\n+\t\t\tverbose_header:1,\n+\t\t\tignore_merges:1,\n+\t\t\tcombine_merges:1,\n+\t\t\tdense_combined_merges:1,\n+\t\t\talways_show_header:1;\n+\n+\t/* Format info */\n+\tunsigned int\tabbrev_commit:1;\n+\tunsigned int\tabbrev;\n+\tenum cmit_fmt\tcommit_format;\n+\tconst char\t*header_prefix;\n+\tconst char\t*header;\n+\n \t/* special limits */\n \tint max_count;\n \tunsigned long max_age;\n"},{"id":"18662","messageId":"Pine.LNX.4.63.0604150211490.28913@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3867","inReplyTo":"87irpb7oma.wl%cworth@cworth.org","subject":"Re: Recent unresolved issues: shallow clone","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-04-15T00:17:20Z","receivedAt":"2006-04-15T00:17:20Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 14 Apr 2006, Carl Worth wrote:\n\n> I also read over some of your discussion of extending the protocol\n> with a new \"shallow\" extension.\n> \n> I'm wondering if the shallow clone support couldn't be achieved\n> through a simpler tweak to the protocol semantics, (and no change to\n> protocol syntax), that would avoid the problem above. Specifically,\n> for shallow stuff, could we just do the same \"want\" and \"have\"\n> conversation with tree objects rather than commit objects?\n\nIt would not help your problem at all. \"have commit\" really means that you \nhave the commit and all its ancestors and their combined tree objects and \nthe combined tree objects' blob objects.\n\nIf you have a cauterized history, you know that you are lacking some of \nthem. But you don't know which ones.\n\nNow, issuing a pull could mean to get an object which was present in an \nold revision, which you unfortunately do not have (because you have a cut \noff history). Boom.\n\nI know, this is probably unlikely, but not at all *impossible*, so you \nhave to take care of that case. And you need a protocol extension for \nthat.\n\nHth,\nDscho\n"},{"id":"18663","messageId":"Pine.LNX.4.64.0604141717280.3701@g5.osdl.org","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0604141637230.3701@g5.osdl.org","subject":"Re: Recent unresolved issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-04-15T00:19:24Z","receivedAt":"2006-04-15T00:19:24Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 14 Apr 2006, Linus Torvalds wrote:\n> \n> It's nice to note how simple and straightforward this makes the built-in \n> \"git log\" command, even though it continues to support all the diff flags \n> too. It doesn't get much simpler that this.\n\nGaah. Missed this important part, which causes the thing to ignore the \n\"--pretty=xyzzy\" argument, since it would always use its own default \nformat that is no longer ever changed.\n\nI even tested that it works for git-diff-tree, just not for git log. Duh.\n\n(Found the hard way - after I had already used the broken git version for \ndoing several merges, and the \"--pretty=oneline\" format didn't work and \nscrewed up the merge message ;)\n\n\t\tLinus\n\n----\ndiff --git a/git.c b/git.c\nindex e8d1fcc..d5a4a24 100644\n--- a/git.c\n+++ b/git.c\n@@ -283,9 +283,6 @@ static int cmd_log(int argc, const char \n \tstruct rev_info rev;\n \tstruct commit *commit;\n \tchar *buf = xmalloc(LOGSIZE);\n-\tstatic enum cmit_fmt commit_format = CMIT_FMT_DEFAULT;\n-\tint abbrev = DEFAULT_ABBREV;\n-\tint abbrev_commit = 0;\n \tconst char *commit_prefix = \"commit \";\n \tint shown = 0;\n \n@@ -298,11 +295,11 @@ static int cmd_log(int argc, const char \n \tprepare_revision_walk(&rev);\n \tsetup_pager();\n \twhile ((commit = get_revision(&rev)) != NULL) {\n-\t\tif (shown && rev.diff && commit_format != CMIT_FMT_ONELINE)\n+\t\tif (shown && rev.diff && rev.commit_format != CMIT_FMT_ONELINE)\n \t\t\tputchar('\\n');\n \t\tfputs(commit_prefix, stdout);\n-\t\tif (abbrev_commit && abbrev)\n-\t\t\tfputs(find_unique_abbrev(commit->object.sha1, abbrev),\n+\t\tif (rev.abbrev_commit && rev.abbrev)\n+\t\t\tfputs(find_unique_abbrev(commit->object.sha1, rev.abbrev),\n \t\t\t      stdout);\n \t\telse\n \t\t\tfputs(sha1_to_hex(commit->object.sha1), stdout);\n@@ -325,12 +322,12 @@ static int cmd_log(int argc, const char \n \t\t\t     parents = parents->next)\n \t\t\t\tparents->item->object.flags &= ~TMP_MARK;\n \t\t}\n-\t\tif (commit_format == CMIT_FMT_ONELINE)\n+\t\tif (rev.commit_format == CMIT_FMT_ONELINE)\n \t\t\tputchar(' ');\n \t\telse\n \t\t\tputchar('\\n');\n-\t\tpretty_print_commit(commit_format, commit, ~0, buf,\n-\t\t\t\t    LOGSIZE, abbrev);\n+\t\tpretty_print_commit(rev.commit_format, commit, ~0, buf,\n+\t\t\t\t    LOGSIZE, rev.abbrev);\n \t\tprintf(\"%s\\n\", buf);\n \t\tif (rev.diff)\n \t\t\tlog_tree_commit(&rev, commit);\n"},{"id":"18664","messageId":"7vr73zn0rb.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"87irpb7oma.wl%cworth@cworth.org","subject":"Re: Recent unresolved issues: shallow clone","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-15T00:25:12Z","receivedAt":"2006-04-15T00:25:12Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Carl Worth <cworth@cworth.org> writes:\n\n> On Fri, 14 Apr 2006 02:31:36 -0700, Junio C Hamano wrote:\n>>   I am beginning to think using \"graft\" to cauterize history\n>>   for this, while it technically would work, would not be so\n>>   helpful to users, so the design needs to be worked out again.\n>\n> As context, here is some of what you mentioned in IRC:\n>\n>>>\tSuppose you have this:\n>>>\n>>>\tA---B---C\n>>>\t \\       \\ \n>>>\t  D---E---F---G\n>>>\t \n>>>\tand you made a shallow clone of C (because that is where the\n>>>\tupstream master was when you made that clone).  Then the\n>>>\tupstream updated the master branch tip to G.\n>>>\n>>>\tThe next update from upstream to your shallow clone would break.\n>>>\tThe upstream says: I have G at master.\n>>>\tYou say: I want G then.  By the way, I have C.\n>>>\n>>>\tWhat it means to tell the other end \"I have X\" is to promise\n>>>\tthat you have X and _everything_ behind it.  So the upstream\n>>>\twould send objects necessary to complete D, E, F and G for\n>>>\t\"somebody who already have A and B\".  As a consequence, you\n>>>\twould not see A nor B.\n>>>\n>>>\tEven if the only thing you are interested in is to be in sync\n>>>\twith the tip of the upstream, you can end up with an\n>>>\tincomplete tree for G, if some of the blobs or trees contained\n>>>\tin G already exist in A or B.  They are not sent -- because\n>>>\tyou told the upstream that you have everything necessary to\n>>>\tget to C.\n>\n> So that's an argument against using a cauterizing graft for the\n> shallow clone of C. It definitely confuses the existing protocol to\n> say \"I have C\" if I have only a cauterized C, (its tree only, but none\n> of the commits that should be backing C).\n\nThat's what I meant by \"graft technically works but is\ninconvenient\". \n\nMaybe after the update to G happens (which means you now have C,\nF, G but not A B D E commits), the client side could enumerate\ncommits on \"rev-list ^C G\" and cauterize the ones with missing\nparents (in this case, F does not have one of its parents).\nWhile doing this would help keeping the resulting commit\nancestry sane, it does not solve the problem of missing blobs\nand trees.  See below.\n\n> So, in the scenario above, the original shallow clone of C would be:\n>\n> \tWant C->tree, have nothing.\n>\n> and the later shallow update to G would be:\n>\n> \tWant G->tree, have C->tree\n\nWhen you ask for G, you do not know what G^{tree} is, so that is\nfantasy without a protocol extention.  To solve the missing\nblobs/trees problem we would probably need a protocol extention\nthat says it wants to receive enough data to complete trees and\nblobs associated with the commits being sent _without_ assuming\nthe recipient has any trees or blobs other than what are\ncontained in \"have\" commits.  Then after such a successful\ntransfer, missing parents of commits listed in \"rev-list ^C G\"\nare the ones from the side branch, so the client can cauterize\nthem (F in the above example) appropriately without bothering\nthe server.\n\nHowever, I think this \"do not assume I have any trees behind the\ncommits I explicitly say I have\" must be an option, because it\nmakes the resulting transfer unnecessarily more expensive for\nnormal uses.  A fetch of the Linux kernel once a day would\nupdate about a couple of hundered commits, each of which touches\nonly 3 paths on average (so that would be 600 files out of\n18,000 file tree.  When side-branch merges are involved, usually\nmany things in G (and F) are unchanged since either A or C, but\nthe extention we are discussing forbids reusing what are found\nin A (it still allows reusing what are found in C).\n\n> A final step of a shallow clone would then require creating a new\n> parent-less commit object so that there's something to point refs/head\n> at, (or maybe rather than being parentless, they could be chained\n> together with each update?).\n\nRewriting commit objects transferred to the cloner is something\nyou would _not_ want to do (e.g. rewriting F commits to say it\nhas only one parent C).  The history based on that would diverge\nfrom parents and would become unmergeable.  It is cleaner to\njust make a new graft entry to say \"As far as this repository is\nconcerned, F has one parent C\".  Shallowness of the repository\nand its slightly different view of history is a local matter.\n"},{"id":"18665","messageId":"7vlku7n05x.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0604141637230.3701@g5.osdl.org","subject":"Re: Recent unresolved issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-15T00:38:02Z","receivedAt":"2006-04-15T00:38:02Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> On Fri, 14 Apr 2006, Junio C Hamano wrote:\n>> \n>> * Message-ID: <Pine.LNX.4.64.0604121828370.14565@g5.osdl.org>\n>>   Common option parsing (Linus Torvalds)\n>\n> Ok, here's a first cut at starting this.\n\nHmph.  You did it while I was still thinking about it ;-).\n\nI was thinking long because I had an impression that anything\nbased on revision.c interface, if it wants to do a tree-diff on\nthe commit stream, would need two different diff options.  One\nis used by revision.c internally so that it can use its own\nadd_remove/change for parent pruning, and another to control the\nway diff is run by the user of revision.c.\n\nI'll take a look at it tonight.  Thanks.\n"},{"id":"18666","messageId":"Pine.LNX.4.64.0604141737580.3701@g5.osdl.org","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0604141717280.3701@g5.osdl.org","subject":"Re: Recent unresolved issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-04-15T00:39:29Z","receivedAt":"2006-04-15T00:39:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 14 Apr 2006, Linus Torvalds wrote:\n> \n> Gaah. Missed this important part, which causes the thing to ignore the \n> \"--pretty=xyzzy\" argument, since it would always use its own default \n> format that is no longer ever changed.\n\nAnd here's one more fixup: get the default format right, and don't prefix \nthe \"oneline\" format.\n\n\t\tLinus\n\n----\ndiff --git a/git.c b/git.c\nindex d5a4a24..437e9b5 100644\n--- a/git.c\n+++ b/git.c\n@@ -291,6 +291,8 @@ static int cmd_log(int argc, const char \n \t\tdie(\"unrecognized argument: %s\", argv[1]);\n \n \trev.no_commit_id = 1;\n+\tif (rev.commit_format == CMIT_FMT_ONELINE)\n+\t\tcommit_prefix = \"\";\n \n \tprepare_revision_walk(&rev);\n \tsetup_pager();\ndiff --git a/revision.c b/revision.c\nindex 99077af..0f98960 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -493,7 +493,7 @@ void init_revisions(struct rev_info *rev\n \trevs->topo_getter = topo_sort_default_getter;\n \n \trevs->header_prefix = \"\";\n-\trevs->commit_format = CMIT_FMT_RAW;\n+\trevs->commit_format = CMIT_FMT_DEFAULT;\n \n \tdiff_setup(&revs->diffopt);\n }\n"},{"id":"18667","messageId":"Pine.LNX.4.64.0604141748070.3701@g5.osdl.org","threadId":"3867","inReplyTo":"7vlku7n05x.fsf@assigned-by-dhcp.cox.net","subject":"Re: Recent unresolved issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-04-15T00:49:56Z","receivedAt":"2006-04-15T00:49:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 14 Apr 2006, Junio C Hamano wrote:\n> \n> I was thinking long because I had an impression that anything\n> based on revision.c interface, if it wants to do a tree-diff on\n> the commit stream, would need two different diff options.  One\n> is used by revision.c internally so that it can use its own\n> add_remove/change for parent pruning, and another to control the\n> way diff is run by the user of revision.c.\n\nI think you're right, and I've probably broken \"--full-diff\" (causing the \nrevparse to also use the empty set of paths). Gaah.\n\n\t\tLinus\n"},{"id":"18668","messageId":"Pine.LNX.4.64.0604141751270.3701@g5.osdl.org","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0604141748070.3701@g5.osdl.org","subject":"Re: Recent unresolved issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-04-15T00:56:53Z","receivedAt":"2006-04-15T00:56:53Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 14 Apr 2006, Linus Torvalds wrote:\n> \n> I think you're right, and I've probably broken \"--full-diff\" (causing the \n> revparse to also use the empty set of paths). Gaah.\n\nIn fact, it's broken path-limited revisions entirely. Duh. We should have \na test for that, so I would have noticed.\n\nI think we need two diffopt structures there - one for the actual diff, \nand one for the pruning.\n\n\t\tLinus\n"},{"id":"18670","messageId":"Pine.LNX.4.64.0604141808480.3701@g5.osdl.org","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0604141751270.3701@g5.osdl.org","subject":"Re: Recent unresolved issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-04-15T01:09:20Z","receivedAt":"2006-04-15T01:09:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOk, fourth time lucky?\n\n\t\tLinus\n\n----\ndiff --git a/revision.c b/revision.c\nindex 0f98960..2061ca8 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -246,7 +246,7 @@ int rev_compare_tree(struct rev_info *re\n \t\treturn REV_TREE_DIFFERENT;\n \ttree_difference = REV_TREE_SAME;\n \tif (diff_tree_sha1(t1->object.sha1, t2->object.sha1, \"\",\n-\t\t\t   &revs->diffopt) < 0)\n+\t\t\t   &revs->pruning) < 0)\n \t\treturn REV_TREE_DIFFERENT;\n \treturn tree_difference;\n }\n@@ -269,7 +269,7 @@ int rev_same_tree_as_empty(struct rev_in\n \tempty.size = 0;\n \n \ttree_difference = 0;\n-\tretval = diff_tree(&empty, &real, \"\", &revs->diffopt);\n+\tretval = diff_tree(&empty, &real, \"\", &revs->pruning);\n \tfree(tree);\n \n \treturn retval >= 0 && !tree_difference;\n@@ -476,9 +476,9 @@ static void handle_all(struct rev_info *\n void init_revisions(struct rev_info *revs)\n {\n \tmemset(revs, 0, sizeof(*revs));\n-\trevs->diffopt.recursive = 1;\n-\trevs->diffopt.add_remove = file_add_remove;\n-\trevs->diffopt.change = file_change;\n+\trevs->pruning.recursive = 1;\n+\trevs->pruning.add_remove = file_add_remove;\n+\trevs->pruning.change = file_change;\n \trevs->lifo = 1;\n \trevs->dense = 1;\n \trevs->prefix = setup_git_directory();\n@@ -780,8 +780,10 @@ int setup_revisions(int argc, const char\n \t\trevs->limited = 1;\n \n \tif (revs->prune_data) {\n-\t\tdiff_tree_setup_paths(revs->prune_data, &revs->diffopt);\n+\t\tdiff_tree_setup_paths(revs->prune_data, &revs->pruning);\n \t\trevs->prune_fn = try_to_simplify_commit;\n+\t\tif (!revs->full_diff)\n+\t\t\tdiff_tree_setup_paths(revs->prune_data, &revs->diffopt);\n \t}\n \tif (revs->combine_merges) {\n \t\trevs->ignore_merges = 0;\n@@ -790,8 +792,6 @@ int setup_revisions(int argc, const char\n \t}\n \tif (revs->diffopt.output_format == DIFF_FORMAT_PATCH)\n \t\trevs->diffopt.recursive = 1;\n-\tif (!revs->full_diff && revs->prune_data)\n-\t\tdiff_tree_setup_paths(revs->prune_data, &revs->diffopt);\n \tdiff_setup_done(&revs->diffopt);\n \n \treturn left;\ndiff --git a/revision.h b/revision.h\nindex 9a45986..6eaa904 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -61,8 +61,9 @@ struct rev_info {\n \tunsigned long max_age;\n \tunsigned long min_age;\n \n-\t/* paths limiting */\n+\t/* diff info for patches and for paths limiting */\n \tstruct diff_options diffopt;\n+\tstruct diff_options pruning;\n \n \ttopo_sort_set_fn_t topo_setter;\n \ttopo_sort_get_fn_t topo_getter;\n"},{"id":"18671","messageId":"7vacanmxhe.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0604141748070.3701@g5.osdl.org","subject":"Re: Recent unresolved issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-15T01:35:57Z","receivedAt":"2006-04-15T01:35:57Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> On Fri, 14 Apr 2006, Junio C Hamano wrote:\n>> \n>> I was thinking long because I had an impression that anything\n>> based on revision.c interface, if it wants to do a tree-diff on\n>> the commit stream, would need two different diff options.  One\n>> is used by revision.c internally so that it can use its own\n>> add_remove/change for parent pruning, and another to control the\n>> way diff is run by the user of revision.c.\n>\n> I think you're right, and I've probably broken \"--full-diff\" (causing the \n> revparse to also use the empty set of paths). Gaah.\n\nAnother thing is that some revision.c users are not interested\nin taking diff options at all.\n\nI was going to suggest a new structure that captures struct\nrev_info, struct log_tree_opt and miscellaneous bits cmd_log\nuses such as do_diff, full_diff, etc., and move the option\nparser out of cmd_log() to a separate function, and have that\nshared across cmd_log(), cmd_show(), cmd_whatchanged(), and\ncmd_diff() without affecting any of the existing revision.c\nusers.  That way, \"rev-list --cc HEAD\" will remain nonsense.\n\nOne nice property your approach has is that it makes\n\"git diff-tree a..b\" magically starts working, unlike what\nI suggested above.\n"},{"id":"18672","messageId":"7vwtdrlh8w.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"7vr73zn0rb.fsf@assigned-by-dhcp.cox.net","subject":"Re: Recent unresolved issues: shallow clone","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-15T02:11:59Z","receivedAt":"2006-04-15T02:11:59Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> Maybe after the update to G happens (which means you now have C,\n> F, G but not A B D E commits), the client side could enumerate\n> commits on \"rev-list ^C G\" and cauterize the ones with missing\n> parents (in this case, F does not have one of its parents).\n> While doing this would help keeping the resulting commit\n> ancestry sane, it does not solve the problem of missing blobs\n> and trees.  See below.\n\nActually, it is more involved than the above.\n\nThe sender would give you A B D E as well, so we would not be\nable to cauterise at F; instead you would do so at A, making\nyour shallow clone a bit deeper.  When you look at the objects\nthe parent gave you by running \"rev-list ^C G\", you would notice\nthat you do not have any of real parents of A, and add a new\ngraft.  While you are at it, you would hopefully notice that the\nreal parent of commit C is something you now have -- so you\nremove the graft entry for C.\n\nHowever, depending on what you care about more, having the\nsender not to send the side branch and cauterising the result at\nF _might_ be a better thing to do.  This is probably quite\ninvolved and I offhand would not know how to efficiently\ncompute.\n"},{"id":"18673","messageId":"7vsloflgr9.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0604141808480.3701@g5.osdl.org","subject":"Re: Recent unresolved issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-15T02:22:34Z","receivedAt":"2006-04-15T02:22:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Ok, fourth time lucky?\n>\n> \t\tLinus\n\nWith this, perhaps.\n\n---\ndiff --git a/revision.c b/revision.c\nindex 2061ca8..1d26e0d 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -792,6 +792,7 @@ int setup_revisions(int argc, const char\n \t}\n \tif (revs->diffopt.output_format == DIFF_FORMAT_PATCH)\n \t\trevs->diffopt.recursive = 1;\n+\trevs->diffopt.abbrev = revs->abbrev;\n \tdiff_setup_done(&revs->diffopt);\n \n \treturn left;\n"},{"id":"18675","messageId":"Pine.LNX.4.64.0604142104140.3701@g5.osdl.org","threadId":"3867","inReplyTo":"7vacanmxhe.fsf@assigned-by-dhcp.cox.net","subject":"Re: Recent unresolved issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-04-15T04:09:31Z","receivedAt":"2006-04-15T04:09:31Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 14 Apr 2006, Junio C Hamano wrote:\n> \n> Another thing is that some revision.c users are not interested\n> in taking diff options at all.\n\nWell, it's easy enough to do something like\n\n\tif (rev->diff)\n\t\tusage(no_diff_cmd_usage);\n\nfor something like that.\n\n> I was going to suggest a new structure that captures struct\n> rev_info, struct log_tree_opt and miscellaneous bits cmd_log\n> uses such as do_diff, full_diff, etc., and move the option\n> parser out of cmd_log() to a separate function, and have that\n> shared across cmd_log(), cmd_show(), cmd_whatchanged(), and\n> cmd_diff() without affecting any of the existing revision.c\n> users.  That way, \"rev-list --cc HEAD\" will remain nonsense.\n\nWell, I actually was going to make git-rev-list just take the diff \noptions, and it ends up doing the same thing as \"git log\" with them. \nThere's no real downside.\n\n> One nice property your approach has is that it makes\n> \"git diff-tree a..b\" magically starts working, unlike what\n> I suggested above.\n\nYeah. It just fell out automatically from using the rev-list parsing.\n\nAlthough, the thing is, once we have a built-in \"git diff\", there's \nactually little enough reason to ever use the old \"git-diff-tree\" vs \n\"git-diff-index\" vs \"git-diff-files\" at all. \n\nIt might actually be nice to prune some of the tons of git commands. At \nsome point, the fact that\n\n\techo bin/git-* | wc -w\n\nreturns 122 just makes you go \"Hmm..\".\n\n\t\tLinus\n"},{"id":"18677","messageId":"7vodz3l964.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0604142104140.3701@g5.osdl.org","subject":"Re: Recent unresolved issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-15T05:06:27Z","receivedAt":"2006-04-15T05:06:27Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Well, it's easy enough to do something like\n>\n> \tif (rev->diff)\n> \t\tusage(no_diff_cmd_usage);\n>\n> for something like that.\n\nYou're right.  I've swallowed all four patches with a fixlet on\ntop; thanks.\n\n> Although, the thing is, once we have a built-in \"git diff\", there's \n> actually little enough reason to ever use the old \"git-diff-tree\" vs \n> \"git-diff-index\" vs \"git-diff-files\" at all. \n\nTrue, unless you are writing a Porcelain, that is.\n\n> It might actually be nice to prune some of the tons of git commands. At \n> some point, the fact that\n>\n> \techo bin/git-* | wc -w\n>\n> returns 122 just makes you go \"Hmm..\".\n\nYes, but I thought the plan to deal with that was to set gitexecdir\nsomewhere other than $(prefix)/bin; removing git-diff-* siblings\nwould be unfriendly to Porcelains.\n"},{"id":"18678","messageId":"7vu08vjra5.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0604141751270.3701@g5.osdl.org","subject":"Re: Recent unresolved issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-15T06:18:10Z","receivedAt":"2006-04-15T06:18:10Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> On Fri, 14 Apr 2006, Linus Torvalds wrote:\n>> \n>> I think you're right, and I've probably broken \"--full-diff\" (causing the \n>> revparse to also use the empty set of paths). Gaah.\n>\n> In fact, it's broken path-limited revisions entirely. Duh. We should have \n> a test for that, so I would have noticed.\n>\n> I think we need two diffopt structures there - one for the actual diff, \n> and one for the pruning.\n\nAlthough I've already decided to merge it up, there are small\nfallout from this.  I've fixed the ones I noticed, but there\nprobably remain some backward compatibility issues in commands\nthat I do not usually use.  We'll see.\n\nAlso I merged the commit prettyprinter change, but I was hoping\nwe could instead pass down the commit object and commit format\nto places where log message needs to be output and write it out\nto the standard output instead of formatting in core.\n"},{"id":"18682","messageId":"7vk69ri5cp.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"7vu08vjra5.fsf@assigned-by-dhcp.cox.net","subject":"Re: Recent unresolved issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-15T08:57:10Z","receivedAt":"2006-04-15T08:57:10Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> Although I've already decided to merge it up, there are small\n> fallout from this.  I've fixed the ones I noticed, but there\n> probably remain some backward compatibility issues in commands\n> that I do not usually use.  We'll see.\n\nI am very to sorry to say this, but...\n\n\n\t\t\t\tPain\n\n\n\"git log\" wants default abbrev (to show Merge: lines and\n\"whatchanged -r\" output compactly) while \"git diff-tree -r\" by\ndefault wants to show full SHA1 unless asked, which means\n\"memset(revs, 0, sizeof(*revs))\" in revision.c::init_revisions()\nneeds to be defeated by the caller.\n\n\"git rev-list\" wants to know if any --pretty was specified to\nset verbose_header, but there is no way to tell if the user did\nnot say anything or said --pretty because revs->commit_format\nwill be CMIT_FMT_DEFAULT either way.  This is the worst breakage\nI found so far -- \"git rev-list --pretty\" no longer works,\nalthough \"git rev-list --header\" works so you probably did not\nnotice the breakage with gitk.\n\nHonestly, the longer I look at it, the more I feel that this way\nmight break more things than it fixes.  I haven't even looked at\nblame.c or http-push.c to see what's broken yet.\n"},{"id":"18688","messageId":"Pine.LNX.4.63.0604151343010.25269@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3867","inReplyTo":"7vk69ri5cp.fsf@assigned-by-dhcp.cox.net","subject":"Re: Recent unresolved issues","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-04-15T11:46:02Z","receivedAt":"2006-04-15T11:46:02Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sat, 15 Apr 2006, Junio C Hamano wrote:\n\n> Honestly, the longer I look at it, the more I feel that this way\n> might break more things than it fixes.  I haven't even looked at\n> blame.c or http-push.c to see what's broken yet.\n\nI do not have time to look at this closely, but it sounds to me like you \nneed a two-stage approach:\n\nsetup_diff_options(&options);\n[... set defaults ...]\nhandle_cmdline_arguments(&options);\n[... possibly check if the user overrode some defaults ...]\n\nI think that the unified option parsing is the right approach.\n\nCiao,\nDscho\n"},{"id":"18702","messageId":"Pine.LNX.4.64.0604150958140.3701@g5.osdl.org","threadId":"3867","inReplyTo":"7vk69ri5cp.fsf@assigned-by-dhcp.cox.net","subject":"Re: Recent unresolved issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-04-15T16:59:50Z","receivedAt":"2006-04-15T16:59:50Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 15 Apr 2006, Junio C Hamano wrote:\n> \n> \t\t\t\tPain\n> \n> \"git log\" wants default abbrev (to show Merge: lines and\n> \"whatchanged -r\" output compactly) while \"git diff-tree -r\" by\n> default wants to show full SHA1 unless asked, which means\n> \"memset(revs, 0, sizeof(*revs))\" in revision.c::init_revisions()\n> needs to be defeated by the caller.\n\nI'd suggest just moving the call to \"init_revisions()\" out from \n\"setup_revisions()\" entirely.\n\nSo the calling sequence would be something like this:\n\n\tinit_revisions(&rev);\n\t.. any localized setup ..\n\tsetup_revisions(&rev);\n\nwhich isn't really all that painful, and allows us maximal flexibility for \ndifferent defaults etc.\n\n\t\tLinus\n"},{"id":"18703","messageId":"Pine.LNX.4.64.0604151016220.3701@g5.osdl.org","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0604150958140.3701@g5.osdl.org","subject":"Re: Recent unresolved issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-04-15T17:17:05Z","receivedAt":"2006-04-15T17:17:05Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 15 Apr 2006, Linus Torvalds wrote:\n> \n> So the calling sequence would be something like this:\n> \n> \tinit_revisions(&rev);\n> \t.. any localized setup ..\n> \tsetup_revisions(&rev);\n\nBtw, I can certainly understand if you don't want to do this before 1.3.x. \nSince there's no actual user-visible advantage to it, it's probably worth \ndropping for now.\n\n\t\tLinus\n"},{"id":"18725","messageId":"7vlku6c4yu.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0604151016220.3701@g5.osdl.org","subject":"Re: Recent unresolved issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-16T08:14:17Z","receivedAt":"2006-04-16T08:14:17Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Btw, I can certainly understand if you don't want to do this before 1.3.x. \n> Since there's no actual user-visible advantage to it, it's probably worth \n> dropping for now.\n\nWell, I bit the bullet, fixed-up the remaining issues I found in\nrev-list in your grand unified version, reverted the revert, and\nported the changes for log/whatchanged/show.\n\nFor now this lives in the \"next\" branch and _will_ not graduate\nbefore 1.3.0, of course.\n"},{"id":"19503","messageId":"7v4q065hq0.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"7v64lcqz9j.fsf@assigned-by-dhcp.cox.net","subject":"Unresolved issues #2","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-04T08:15:03Z","receivedAt":"2006-05-04T08:15:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Here is a list of topics in the recent git traffic that I feel\ninadequately addressed.  I've commented on some of them to give\npeople a feel for what my priorities are.  Somebody might want\nto rehash the ones low on my priority list to conclusion with a\nconcrete proposal if they cared about them enough.\n\nThe list is *not* ordered in any way, except that the entries\nkept from the previous issue of this message have been pushed\ndown to the bottom.  I will probably start dropping some entries\nthat did not get any reaction from the list in future issues of\nthis message, but for now I kept all of them from the first one.\n\n* Message-ID: <4fb292fa0604290630r19edd7ejf88642e33b350d1d@mail.gmail.com>\n  Content-type charset for send-email (Bertrand Jacquin)\n\n  The output from format-patch by default is unmarked, which\n  means the commit message part is UTF-8 (by strong convention),\n  and the contents of the diff is whatever the contents of the\n  file is encoded in.\n\n  David Woodhouse did a patch to allow specifying charset on the\n  command line (and default to UTF-8) which is a move in the\n  right direction, but Bertrand's system seems to have trouble\n  with it.\n\n  I think if we were to do this we probably need to teach\n  format-patch to optionally do multi-part.  We may not\n  necessarily want to mark the payload to be in the same\n  encoding as the commit message (not that git-apply cares -- to\n  it, the payload is just 8-bit unencoded text, but we would\n  want to protect it from getting mangled by e-mail transport).\n\n* Message-ID: <Pine.LNX.4.64.0604291006270.3701@g5.osdl.org>\n  Perhaps \"note\" field in commit objects are useful?\n\n* Message-ID: <Pine.LNX.4.63.0604301524080.2646@wbgn013.biozentrum.uni-wuerzburg.de>\n\n  An optional \"git fetch --store newname URL refspecs...\" to\n  create an equivalent of remotes file so newname can then be\n  used as a short-hand.  I still have somewhat negative reaction\n  to it, but I am willing to apply it if there are enough people\n  who want this.\n\n-- carried over from the first issue of this list.\n\n* Message-ID: <Pine.LNX.4.64.0604050855080.2550@localhost.localdomain>\n  Binary diff output? (Nicolas Pitre)\n\n  I do not think this is needed for our primary audience (the\n  kernel project), but I am sure it would be helpful for some\n  other projects if we allowed them to exchange patches that\n  describe binary file changes via e-mail, so I am not\n  dismissing this.\n\n* #irc 2006-04-10\n  Shallow clones (Carl Worth).\n\n  The experiment last round did not work out very well, but as\n  existing repositories get bigger, and more projects being\n  migrated from foreign SCM systems, this would become a\n  must-have from would-be-nice-to-have.\n\n  I am beginning to think using \"graft\" to cauterize history\n  for this, while it technically would work, would not be so\n  helpful to users, so the design needs to be worked out again.\n\n* Message-ID: <E1FMH3o-0001B5-Dw@jdl.com>\n  git status does not distinguish contents changes and mode\n  changes; it just says \"modified\" (Jon Loeliger).\n\n  Unconditionally changing the status letter would break\n  Porcelains so we would need an extra option to do this.\n  An outline patch has been already prepared -- this perhaps has\n  to wait until we sort out the \"option parsing\" one.\n\n* Message-ID: <tnxmzf9sh7k.fsf@arm.com>\n  git could use diff3 instead of merge which is a wrapper around\n  diff3. (Catalin Marinas)\n\n  If having \"diff3\" is a lot more common than having \"merge\", I\n  do not have problem with this; \"merge\" being a wrapper to\n  \"diff3\", people who have been happy with the current code\n  would certainly have \"diff3\" installed so changing to \"diff3\"\n  would not break them.\n\n* Message-ID: <81b0412b0603020649u99a2035i3b8adde8ddce9410@mail.gmail.com>\n  Windows problems summary (Alex Riesen)\n\n  A good list to keep in mind.\n\n* Message-ID: <Pine.LNX.4.64.0604030730040.3781@g5.osdl.org>\n  Huge packfiles (Linus Torvalds)\n\n  Because I do not think asking users to break up packs to\n  manageable and mmap()able size is too much to ask, I would not\n  be advocating for updating the pack idx to 64-bit offset and\n  mmap()ing parts of a packfile, at least too strongly.\n\n  However, we currently lack tool support or recepe for users\n  with such a repository to easily break up packs.\n\n* Message-ID: <1143856098.3555.48.camel@dv>\n  Per branch property, esp. where to merge from (Pavel Roskin)\n\n  This involves user-level \"world model\" design, which is more\n  Porcelainish than Plumbing, and as people know I do not do\n  Porcelain well; interested parties need to come up with what\n  they want and how they want to use it.\n"},{"id":"19504","messageId":"e3ce4j$chl$1@sea.gmane.org","threadId":"3867","inReplyTo":"7v4q065hq0.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-04T08:32:22Z","receivedAt":"2006-05-04T08:32:22Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n> * #irc 2006-04-10\n>   Shallow clones (Carl Worth).\n> \n>   The experiment last round did not work out very well, but as\n>   existing repositories get bigger, and more projects being\n>   migrated from foreign SCM systems, this would become a\n>   must-have from would-be-nice-to-have.\n> \n>   I am beginning to think using \"graft\" to cauterize history\n>   for this, while it technically would work, would not be so\n>   helpful to users, so the design needs to be worked out again.\n\nPerhaps use comment for marking graft as cauterizing history?\n\nThere was also talk about proposed git-splithist, which would move some of\nthe history to other (historical, archive) repository.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19507","messageId":"7vwtd240e0.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"e3ce4j$chl$1@sea.gmane.org","subject":"Re: Unresolved issues #2","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-04T09:14:47Z","receivedAt":"2006-05-04T09:14:47Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n>>   I am beginning to think using \"graft\" to cauterize history\n>>   for this, while it technically would work, would not be so\n>>   helpful to users, so the design needs to be worked out again.\n>\n> Perhaps use comment for marking graft as cauterizing history?\n\n?\n\n> There was also talk about proposed git-splithist, which would move some of\n> the history to other (historical, archive) repository.\n\nI stayed out from that discussion, but my impression was that\nyou could essentially do the same thing as what Linus did when\nhe started the recent kernel history since v2.6.12-rc2 without\nany tool support.\n\nThe older kernel history from BKCVS was resurrected later by\nindependent parties and Linus's history can be grafted onto it,\nbut if you have an existing history stored in git, you could do:\n(1) take a snapshot of the tip of your development with \"git\ntar-tree HEAD\"; (2) extract it into an empty repository and\nstart a new history; (3) build on top of the truncated history;\nand (4) graft that onto the history that stopped at (1), which\nyou tentatively abandoned, as needed.\n\n\n\n\t\n"},{"id":"19508","messageId":"e3ch94$o7e$1@sea.gmane.org","threadId":"3867","inReplyTo":"7vwtd240e0.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-04T09:26:00Z","receivedAt":"2006-05-04T09:26:00Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n>>>   I am beginning to think using \"graft\" to cauterize history\n>>>   for this, while it technically would work, would not be so\n>>>   helpful to users, so the design needs to be worked out again.\n>>\n>> Perhaps use comment for marking graft as cauterizing history?\n> \n> ?\n\nFor example:\n\n# begin shallow clone\n<sha1 of commit 1> # no parents... - cut-off commit\n<sha1 of commit 2>\n...\n<sha1 of commmit n>\n# end shallow clone\n\nI don't think it is very good idea, though...\n\n>> There was also talk about proposed git-splithist, which would move some\n>> of the history to other (historical, archive) repository.\n> \n> I stayed out from that discussion, but my impression was that\n> you could essentially do the same thing as what Linus did when\n> he started the recent kernel history since v2.6.12-rc2 without\n> any tool support.\n> \n> The older kernel history from BKCVS was resurrected later by\n> independent parties and Linus's history can be grafted onto it,\n> but if you have an existing history stored in git, you could do:\n> (1) take a snapshot of the tip of your development with \"git\n> tar-tree HEAD\"; (2) extract it into an empty repository and\n> start a new history; (3) build on top of the truncated history;\n> and (4) graft that onto the history that stopped at (1), which\n> you tentatively abandoned, as needed.\n\nI have thought about splitting not at current tip(s), but for example at 1\nyear ago. Current repository would have history cautherized using grafts\n(although it would be nice to have option to omit grafts and reach to\nhistoric repository), and archive/history repository ending with commits up\nto (but not including) the cut-off (cauterization) points.\n\nIIRC the problem with 'shallow clone' was telling which commits the clone\nhas, and how to join commits and recauterize history.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19509","messageId":"20060504095827.GW27689@pasky.or.cz","threadId":"3867","inReplyTo":"7v4q065hq0.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-05-04T09:58:27Z","receivedAt":"2006-05-04T09:58:27Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, May 04, 2006 at 10:15:03AM CEST, I got a letter\nwhere Junio C Hamano <junkio@cox.net> said that...\n> * Message-ID: <1143856098.3555.48.camel@dv>\n>   Per branch property, esp. where to merge from (Pavel Roskin)\n> \n>   This involves user-level \"world model\" design, which is more\n>   Porcelainish than Plumbing, and as people know I do not do\n>   Porcelain well; interested parties need to come up with what\n>   they want and how they want to use it.\n\nOh, my holey memory. In Cogito, I have just implemented a solution\nsuggested by Martin Mares, which is pretty simple, non-obtrusive\nand will work equally fine with remotes as well as remote branches:\n\n\tif [ $branch != master ] && [ -s .git/branches/$branch-origin ]\n\t\torigin=.git/branches/$branch-origin\n\telse\n\t\torigin=.git/branches/origin\n\tfi\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"19518","messageId":"1146757531.5294.26.camel@dv","threadId":"3867","inReplyTo":"20060504095827.GW27689@pasky.or.cz","subject":"Re: Unresolved issues #2","fromName":"Pavel Roskin","fromEmail":"proski@gnu.org","sentAt":"2006-05-04T15:45:31Z","receivedAt":"2006-05-04T15:45:31Z","isPatch":false,"sender":{"key":"proski@gnu.org","avatar":null},"body":"Hello, Petr!\n\nOn Thu, 2006-05-04 at 11:58 +0200, Petr Baudis wrote:\n\n> Oh, my holey memory. In Cogito, I have just implemented a solution\n> suggested by Martin Mares, which is pretty simple, non-obtrusive\n> and will work equally fine with remotes as well as remote branches:\n> \n> \tif [ $branch != master ] && [ -s .git/branches/$branch-origin ]\n> \t\torigin=.git/branches/$branch-origin\n> \telse\n> \t\torigin=.git/branches/origin\n> \tfi\n\nIsn't \".git/branches\" obsolete, at least in git?  I'm surprised it's\nstill referenced in git sources.\n\nWhat is the future of \".git/branches\"?  Is it becoming a Cogito specific\nbranch database?  Or is it now a database or branch dependencies?\n\nI would prefer to have one single standard for branch origins that could\nbe used by git, StGIT and Cogito.  Using a location that is obsolete\noutside Cogito is probably the worst possible approach.  I'd rather have\na separate directory, e.g. .git/origins or something.\n\nAlternatively, we could reuse .git/refs by having files with \"Pull:\" but\nwithout \"URI:\", e.g.\n\n$ cat .git/refs/branch\nPull: branch-origin:branch\n\n-- \nRegards,\nPavel Roskin\n"},{"id":"19520","messageId":"87mzdx7mh9.wl%cworth@cworth.org","threadId":"3867","inReplyTo":"7v4q065hq0.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Carl Worth","fromEmail":"cworth@cworth.org","sentAt":"2006-05-04T17:01:38Z","receivedAt":"2006-05-04T17:01:38Z","isPatch":false,"sender":{"key":"cworth@cworth.org","avatar":"https://gravatar.com/avatar/3746dc28cde609bdbd7f939058356e7e2bbd16d21e32274df0725eb3d998bc5b?d=mp&s=160"},"body":"On Thu, 04 May 2006 01:15:03 -0700, Junio C Hamano wrote:\n> * #irc 2006-04-10\n>   Shallow clones (Carl Worth).\n> \n>   The experiment last round did not work out very well, but as\n>   existing repositories get bigger, and more projects being\n>   migrated from foreign SCM systems, this would become a\n>   must-have from would-be-nice-to-have.\n> \n>   I am beginning to think using \"graft\" to cauterize history\n>   for this, while it technically would work, would not be so\n>   helpful to users, so the design needs to be worked out again.\n\nI've been meaning to follow up with some thoughts on this topic, so\nthanks for the tickler.\n\nFor the one use case I had, (track latest tree), I had thrown out the\nidea of using \"faked\", parent-less commit objects to point to the tree\nof interest. Junio pointed out that there's no protocol to learn the\nname of a remote commit's tree from the name of the commit. I worked\naround that by simply making the parent-less commit object on the\nserver side, (branch name of \"master-shallow\", say).\n\nThat seemed to work just fine, and if someone really wanted to do\nthis, they could use a hook to maintain the master-shallow branch,\nand no change to git itself would be needed. But there's a very\nminimal amount of interesting functionality in this, and it's not\nclear that it's much better than git-tar-tree. So I'm considering that\nidea dead.\n\nMeanwhile, a more general ability to use shallow clones would still be\nvery useful. I think what I'd like to be able to do is to pass\nrev-list limiting options (--max-count, --max-age via --since,\netc.). That would limit the expansion of the WANT commits, and then\nthe existing logic to compute the necessary objects needed to satisfy\nthe list of desired commits should do the right thing.\n\nThen, in order for this to actually be useful, when returning objects\nfrom a limited fetch like this, the server should provide a list of\ncommits that should be noted as cauterized, (whether through the\nexisting grafts mechanism or otherwise).\n\nAdditionally, when doing a fetch into a tree that has any such\ncauterized commits, the client must also provide its list of\ncauterized commits. So the conversation changes from \"I WANT\n<fetch-heads> and I HAVE <heads>\" to one of \"I WANT <fetch-heads>, and\nI HAVE <heads>, except that I'm MISSING <cauterized-commits>\".\n\nFinally, whenever a fetch receives an commit object that is in its\nlist of cauterized commits, it should remove that commit from the\nlist. This allows a shallow clone to be naturally migrated to\nsomething unshallow. And the user can do this as incrementally as\ndesired based on the need to see more history:\n\nget a bit:\n\tgit fetch somewhere --since=2.weeks.ago\n\nthen a bit more:\n\tgit fetch somewhere --since=1.year.ago\n\nthen get it all:\n\tgit fetch somewhere\n\nMaybe that's no different from Junio's original proposal. If not, what\ndo you see in the above that wouldn't work?\n\n-Carl\n"},{"id":"19523","messageId":"Pine.LNX.4.64.0605041627310.6713@iabervon.org","threadId":"3867","inReplyTo":"7v4q065hq0.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2006-05-04T20:41:58Z","receivedAt":"2006-05-04T20:41:58Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Thu, 4 May 2006, Junio C Hamano wrote:\n\n> * Message-ID: <Pine.LNX.4.63.0604301524080.2646@wbgn013.biozentrum.uni-wuerzburg.de>\n> \n>   An optional \"git fetch --store newname URL refspecs...\" to\n>   create an equivalent of remotes file so newname can then be\n>   used as a short-hand.  I still have somewhat negative reaction\n>   to it, but I am willing to apply it if there are enough people\n>   who want this.\n\nI was just about to suggest something for this general use. It's currently \nkind of a pain to deal with the situation where you've got stuff on your \nworkstation that you want to version control in a shared repository on a \nserver.\n\nI think it shouldn't be on fetch, though; I think a \"git remote\" command \nfor describing, creating, and modifying remotes would be better, since you \nalso sometimes want to add a \"Push:\" line.\n\nMaybe:\n\n git remote <name>: Print info about <name>\n git remote add <name> <URL> [<direction> ...]: create a remote\n git remote <name> <direction> ...: modify a remote\n\n where <direction> is either:\n  pull <remote> <local> or\n  push <local> <remote>\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"19525","messageId":"Pine.LNX.4.64.0605041715500.3611@g5.osdl.org","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0605041627310.6713@iabervon.org","subject":"Re: Unresolved issues #2","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-04T21:33:20Z","receivedAt":"2006-05-04T21:33:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 4 May 2006, Daniel Barkalow wrote:\n> \n> I think it shouldn't be on fetch, though; I think a \"git remote\" command \n> for describing, creating, and modifying remotes would be better, since you \n> also sometimes want to add a \"Push:\" line.\n\nI don't think this is wrong, but I think it's more important to try to \ndecide on how we want to represent this information first, and stabilize \nthat.\n\nI realize that git has gotten a lot more porcelainish over time, but at \nthe same time, now you're really starting to argue about syntax that \nreally ends up being often a feature of the development environment. If \nyou did development using an IDE that knows about git, I think the \n\"remote\" information ends up being not necessarily a git command at all, \nbut really an interface in the IDE.\n\nI'm actually growing pretty fond of the config file interfaces that Dscho \nis pushing. I really like the idea of \"git pull\" doing different things \ndepending on which branch is active at the time, because different \nbranches really can have different sources they come from.\n\nAlways pulling from the same default source seems wrong, and having to \nremember whose source some branch is associated with is just not all that \nuser-friendly, but perhaps more importantly, it's also going to result in \npeople making mistakes, pulling from the wrong branch (because they didn't \nthink about where they were), and then having strange merges that they \nmight not notice were wrong until it's too late and they pushed the result \nout.\n\nSo Johannes' patches seem to move into that direction, and having it all \nin the config file actually seems to be quite readable.\n\nAnd that, in turn, may mean that a lot of porcelains really only care \nabout that syntax, and then they may update the config file any way they \nplease (whether by hand, or by using \"git repo-config\" or by using \"git \nremote\").\n\nSo I'd argue that (a) yes, we do want to have the \"proto porcelain\" that \nsets remote branch information without the user having to know the magic \n\"git repo-config\" incantation, or know which file in .git/remotes/ to \nedit, but that (b) it's even more important to try to decide on what the \nremote description format _is_.\n\nI personally have just two preferences:\n\n - I'd like each branch I'm on to have a \"default source\" for pulling (and \n   _maybe_ for pushing too). I'd like to just say \"git pull\", and it would \n   automatically select the appropriate thing to pull from.\n\n - maybe the same per-branch thing for \"push\", but more importantly for \n   me, I like to push to multiple destinations, and I'd like the \n   description format to be sane. I think it may already be sane in the \n   form it is in now (supporting both config file _and_ .git/remotes/ \n   formats), I'd just like us to decide on exactly what the meaning is, \n   and hopefully get to the point where we can tell porcelain how to use \n   that meaning to their advantage (and not change it)\n\nOthers may disagree, or (equally importantly), may have additional \npreferences. We should try to find something that works for everybody, and \nthat is easy to work with.\n\n\t\tLinus\n"},{"id":"19536","messageId":"7v1wv92u7o.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"87mzdx7mh9.wl%cworth@cworth.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-05T00:25:47Z","receivedAt":"2006-05-05T00:25:47Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Carl Worth <cworth@cworth.org> writes:\n\n> ... So the conversation changes from \"I WANT\n> <fetch-heads> and I HAVE <heads>\" to one of \"I WANT <fetch-heads>, and\n> I HAVE <heads>, except that I'm MISSING <cauterized-commits>\".\n>\n> Finally, whenever a fetch receives an commit object that is in its\n> list of cauterized commits, it should remove that commit from the\n> list. This allows a shallow clone to be naturally migrated to\n> something unshallow. And the user can do this as incrementally as\n> desired based on the need to see more history:\n>\n> get a bit:\n> \tgit fetch somewhere --since=2.weeks.ago\n>\n> then a bit more:\n> \tgit fetch somewhere --since=1.year.ago\n>\n> then get it all:\n> \tgit fetch somewhere\n>\n> Maybe that's no different from Junio's original proposal. If not, what\n> do you see in the above that wouldn't work?\n\nLack of actual code to do all that ;-)\n\nJokes aside, I think listing the updated conversation elements\nlike you did above is a good step forward.\n\nThe vocabulary we would want from the requestor side is probably\n(at least):\n\n\tI WANT to have these\n        I HAVE these\n        I'm MISSING these\n        Don't bother with these this time around (--since, ^v2.6.16, ...)\n\nI am not sure how we would want to encode the last one and have\nit used by rev-list on the upload-pack end safely and sanely.\n\nAnd the responder side needs to be able to say, \"Now you are\nMISSING these, remember it and tell me you are missing them next\ntime you make a request\".  That would be, in the simplest case,\na list of commit IDs to cauterize, but I am not sure what is the\nright way to come up with that list.  Especially I do not know\nif --boundary would/should work with --objects.\n"},{"id":"19544","messageId":"46a038f90605042217n3261b14cxd63f35a31223848e@mail.gmail.com","threadId":"3867","inReplyTo":"7v1wv92u7o.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-05T05:17:10Z","receivedAt":"2006-05-05T05:17:10Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/5/06, Junio C Hamano <junkio@cox.net> wrote:\n> The vocabulary we would want from the requestor side is probably\n> (at least):\n>\n>         I WANT to have these\n>         I HAVE these\n>         I'm MISSING these\n>         Don't bother with these this time around (--since, ^v2.6.16, ...)\n\nThinking... does the MISSING part matter at all? It seems that what\nreally matters are the \"ignore rules\". The pull may bring in a new\nmerge of a long-running branch, whose mergebase falls out of the\nignore rules.\n\nIn that case, the server should apply the ignore rules. Except that\nlater merges in the local repo would perhaps have to deal with missing\npart of the history. I suspect it should refuse to merge something we\ndon't have all the merging parts for.\n\ncheers,\n\n\nmartin\n"},{"id":"19545","messageId":"87bqud6o4p.wl%cworth@cworth.org","threadId":"3867","inReplyTo":"46a038f90605042217n3261b14cxd63f35a31223848e@mail.gmail.com","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Carl Worth","fromEmail":"cworth@cworth.org","sentAt":"2006-05-05T05:23:34Z","receivedAt":"2006-05-05T05:23:34Z","isPatch":false,"sender":{"key":"cworth@cworth.org","avatar":"https://gravatar.com/avatar/3746dc28cde609bdbd7f939058356e7e2bbd16d21e32274df0725eb3d998bc5b?d=mp&s=160"},"body":"On Fri, 5 May 2006 17:17:10 +1200, \"Martin Langhoff\" wrote:\n> >         I WANT to have these\n> >         I HAVE these\n> >         I'm MISSING these\n> >         Don't bother with these this time around (--since, ^v2.6.16, ...)\n> \n> Thinking... does the MISSING part matter at all?\n\nYes.\n\nImagine doing a shallow clone and then fetching a tree that includes a\nblob that existed before MISSING. If we say HAVE without MISSING then\nthe server will not send that blob and we'll be left with a broken\ntree.\n\n> In that case, the server should apply the ignore rules. Except that\n> later merges in the local repo would perhaps have to deal with missing\n> part of the history. I suspect it should refuse to merge something we\n> don't have all the merging parts for.\n\nYeah, shallow clones can shake up the conventions a bit. It's\ndefinitely common for a repository to only have a single parent-less\ncommit, such that there is always an identifiable merge base for any\npair of revisions. Shallow clones would make (effectively) parent-less\ncommits much more common.\n\nShould be fun to see what things fall over with this...\n\n-Carl\n"},{"id":"19546","messageId":"e3eouo$1fm$1@sea.gmane.org","threadId":"3867","inReplyTo":"87bqud6o4p.wl%cworth@cworth.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-05T05:48:11Z","receivedAt":"2006-05-05T05:48:11Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Carl Worth wrote:\n\n> On Fri, 5 May 2006 17:17:10 +1200, \"Martin Langhoff\" wrote:\n\n>> In that case, the server should apply the ignore rules. Except that\n>> later merges in the local repo would perhaps have to deal with missing\n>> part of the history. I suspect it should refuse to merge something we\n>> don't have all the merging parts for.\n> \n> Yeah, shallow clones can shake up the conventions a bit. It's\n> definitely common for a repository to only have a single parent-less\n> commit, such that there is always an identifiable merge base for any\n> pair of revisions. Shallow clones would make (effectively) parent-less\n> commits much more common.\n\nI wonder if it would be possible for git to:\na) as for a fetch which would bring all the commits up to the merge base\n   (and merge base has to be calculated on the server side I think),\n   i.e. give command to use (for fetch or for force baseless merge)\nb) fetch the commits\nc) do merge\nd) optionally re-cauterize history again\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19558","messageId":"Pine.LNX.4.64.0605050806370.3622@g5.osdl.org","threadId":"3867","inReplyTo":"7v1wv92u7o.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-05T15:10:50Z","receivedAt":"2006-05-05T15:10:50Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 4 May 2006, Junio C Hamano wrote:\n> \n> Jokes aside, I think listing the updated conversation elements\n> like you did above is a good step forward.\n> \n> The vocabulary we would want from the requestor side is probably\n> (at least):\n> \n> \tI WANT to have these\n>         I HAVE these\n>         I'm MISSING these\n>         Don't bother with these this time around (--since, ^v2.6.16, ...)\n\nActually, I think we can do something simpler that _most_ people might be \nhappy with.\n\nNamely just have a mode to \"git-send-pack\" that uses the \"--no-walk\" flag \nto generate the object list to send.\n\nWhat that does is to never walk the object history: so it will just use \nthe \"I HAVE THESE\" and \"I WANT THESE\" commit references to directly \ngenerate the list of commits, and then walks the trees to generate the \nlist of trees/blobs that differ between the particular end-points.\n\nWe already have the \"no_walk\" flag internally, we just don't expose it.\n\nSo what you'd get is a _really_ cut down history that doesn't contain any \ncommit history at all (just distinct \"points in commit history time\"), but \nthat _does_ contain all the objects that the commits point to.\n\n\t\tLinus\n"},{"id":"19559","messageId":"e3fqb9$hed$1@sea.gmane.org","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0605050806370.3622@g5.osdl.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-05T15:18:06Z","receivedAt":"2006-05-05T15:18:06Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds wrote:\n\n> So what you'd get is a _really_ cut down history that doesn't contain any\n> commit history at all (just distinct \"points in commit history time\"), but\n> that _does_ contain all the objects that the commits point to.\n\nSo we would get 'skin-deep clone' rather than 'shallow' one?\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19560","messageId":"8764kk7akm.wl%cworth@cworth.org","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0605050806370.3622@g5.osdl.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Carl Worth","fromEmail":"cworth@cworth.org","sentAt":"2006-05-05T15:31:05Z","receivedAt":"2006-05-05T15:31:05Z","isPatch":false,"sender":{"key":"cworth@cworth.org","avatar":"https://gravatar.com/avatar/3746dc28cde609bdbd7f939058356e7e2bbd16d21e32274df0725eb3d998bc5b?d=mp&s=160"},"body":"On Fri, 5 May 2006 08:10:50 -0700 (PDT), Linus Torvalds wrote:\n> \n> Namely just have a mode to \"git-send-pack\" that uses the \"--no-walk\" flag \n> to generate the object list to send.\n> \n> What that does is to never walk the object history: so it will just use \n> the \"I HAVE THESE\" and \"I WANT THESE\" commit references to directly \n> generate the list of commits, and then walks the trees to generate the \n> list of trees/blobs that differ between the particular end-points.\n\nOh, I think that's a great idea. I had proposed cutting the WANT list\ndown to a single commit, but I wasn't clever enough to think to also\ncut down the HAVE walking to solve the problems that would have been\nintroduced by shallow clones.\n\nAnd I think the resulting behavior is quite reasonable. I think one\ncan argue a sort of zero-one-infinity rule here. With history, either\nyou don't care about any of it, or else you really should care about\nall of it.\n\n-Carl\n"},{"id":"19562","messageId":"Pine.LNX.4.64.0605050848230.3622@g5.osdl.org","threadId":"3867","inReplyTo":"e3fqb9$hed$1@sea.gmane.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-05T15:59:44Z","receivedAt":"2006-05-05T15:59:44Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 May 2006, Jakub Narebski wrote:\n\n> Linus Torvalds wrote:\n> \n> > So what you'd get is a _really_ cut down history that doesn't contain any\n> > commit history at all (just distinct \"points in commit history time\"), but\n> > that _does_ contain all the objects that the commits point to.\n> \n> So we would get 'skin-deep clone' rather than 'shallow' one?\n\nWell, it's really shallow, but perhaps more importantly, I think it should \nbe really easy, and have totally unambiguous semantics. Never any question \nof how far back to go, and I think we already really do have all the \nsupport logic for doing it.\n\nNow, we don't actually expose the internal \"no_walk\" flag with a \n\"--no-walk\" command line argument parsing, but that's a one-liner.\n\nThere's another approach that might be a bit friendlier, which is again to \nwalk only the objects of the WANT/HAVE things, but then _do_ walk the \nhistory for just commit objects. Something close to what I think the http \nfetch thing does if you pass it \"-c -t\". That too shouldn't require too \nmuch extra complexity, and it would mean that \"git log\" at least works.\n\nOf course, that would require another slight difference to \"rev-list.c\", \nwhere we'd only recurse into trees of selected commit objects (ie we'd \nhave to mark the HAVE/WANT commits specially, but it's not exactly \ncomplex either).\n\nOf course, the complexity of _both_ of these approaches is really in the \nfsck stage, and all the crud you need to then do other things with these \npared-down repos. For example, do you allow cloning? And do you just \nautomatically notice that you're cloning a shallow repo, and only do a \nshallow clone. Etc etc..\n\n\t\tLinus\n"},{"id":"19594","messageId":"7vhd43vgnm.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0605041715500.3611@g5.osdl.org","subject":"Re: Unresolved issues #2","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-06T05:58:05Z","receivedAt":"2006-05-06T05:58:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> I'm actually growing pretty fond of the config file interfaces that Dscho \n> is pushing. I really like the idea of \"git pull\" doing different things \n> depending on which branch is active at the time, because different \n> branches really can have different sources they come from.\n\n> Always pulling from the same default source seems wrong,...\n\n> So Johannes' patches seem to move into that direction, and having it all \n> in the config file actually seems to be quite readable.\n\nI share the same reasoning and that is why I am carrying the\nseries in \"next\".  I think per branch attributes are wonderful\nthings.\n\n> So I'd argue that (a) yes, we do want to have the \"proto porcelain\" that \n> sets remote branch information without the user having to know the magic \n> \"git repo-config\" incantation, or know which file in .git/remotes/ to \n> edit, but that (b) it's even more important to try to decide on what the \n> remote description format _is_.\n\nIs it format you care about or the semantics?\n\n> I personally have just two preferences:\n>\n>  - I'd like each branch I'm on to have a \"default source\" for pulling (and \n>    _maybe_ for pushing too). I'd like to just say \"git pull\", and it would \n>    automatically select the appropriate thing to pull from.\n>\n>  - maybe the same per-branch thing for \"push\", but more importantly for \n>    me, I like to push to multiple destinations, and I'd like the \n>    description format to be sane. I think it may already be sane in the \n>    form it is in now (supporting both config file _and_ .git/remotes/ \n>    formats), I'd just like us to decide on exactly what the meaning is, \n>    and hopefully get to the point where we can tell porcelain how to use \n>    that meaning to their advantage (and not change it)\n>\n> Others may disagree, or (equally importantly), may have additional \n> preferences. We should try to find something that works for everybody, and \n> that is easy to work with.\n\nIn my day job, I maintain a base code for a generic application\nin \"master\", various topics, mostly branched from \"master\" but\nsometimes from another topic branch, and one branch each per\ncustomer installation, which pulls from the master, topics and\ncontains specific customizations.  While on master or any one of\ngeneric topic branch, I need to remember not to pull from\ninstallation branches.  For that matter, the installation\nbranches should not be pulled into anything else.  So not just\n\"this branch usually merges from there\", but \"this branch should\nnot be merged into others\" (mark \"installation branches\" as\nsuch), and \"this branch should never merge from that one\" (mark\n\"master\" with \"installation branches\") would prevent mistakes.\n\nOne thing I noticed in \"What's in libata.git\" Jeff did by\nmimicking my \"What's in git.git\" was that the description for\neach topic branch included where it branched from (iow, what\nother branch it builds on).  This is sometimes derivable, but\nhaving it as a property for a branch is very handy.\n"},{"id":"19595","messageId":"46a038f90605052323o29f8bfadr7426f97d8dfc2319@mail.gmail.com","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0605050848230.3622@g5.osdl.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-06T06:23:03Z","receivedAt":"2006-05-06T06:23:03Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/6/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> Of course, that would require another slight difference to \"rev-list.c\",\n> where we'd only recurse into trees of selected commit objects (ie we'd\n> have to mark the HAVE/WANT commits specially, but it's not exactly\n> complex either).\n\nWould it make sense to make all the shallow clone clone machinery walk\neverything and trim only blob objects? In that case, all the machinery\nthat walks commits/trees would remain intact -- we only have to deal\nwith the case of not having blob objects, which affects less\ncodepaths.\n\nIt means that for a merge or checkout involving stuff we \"don't have\",\nit's trivial to know we are missing, and so we can  attempt a fetch of\nthe missing objects or tell the user how to request them them before\nretrying.\n\nAnd in any case commits and trees are lightweight and compress well...\n\n> Of course, the complexity of _both_ of these approaches is really in the\n> fsck stage, and all the crud you need to then do other things with these\n> pared-down repos. For example, do you allow cloning? And do you just\n> automatically notice that you're cloning a shallow repo, and only do a\n> shallow clone. Etc etc..\n\nDefinitely.\n\ncheers,\n\n\nmartin\n"},{"id":"19597","messageId":"7vbqubvdbr.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"46a038f90605052323o29f8bfadr7426f97d8dfc2319@mail.gmail.com","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-06T07:10:00Z","receivedAt":"2006-05-06T07:10:00Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n\n> On 5/6/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>> Of course, that would require another slight difference to \"rev-list.c\",\n>> where we'd only recurse into trees of selected commit objects (ie we'd\n>> have to mark the HAVE/WANT commits specially, but it's not exactly\n>> complex either).\n>\n> Would it make sense to make all the shallow clone clone machinery walk\n> everything and trim only blob objects? In that case, all the machinery\n> that walks commits/trees would remain intact -- we only have to deal\n> with the case of not having blob objects, which affects less\n> codepaths.\n>\n> It means that for a merge or checkout involving stuff we \"don't have\",\n> it's trivial to know we are missing, and so we can  attempt a fetch of\n> the missing objects or tell the user how to request them them before\n> retrying.\n>\n> And in any case commits and trees are lightweight and compress well...\n\nCommit maybe, but is this based on a hard fact?  \n\nEarlier Linus said something about \"git log\" working on\ncommit-only copy, but obviously you would want at least trees\nfor the path limiting part to work, so having commits and trees\nwould be handy, but my impression was that at least for deep\nproject like the kernel trees tend to be nonnegligible (a commit\nconsists of 18K paths and 1200 trees or something like that).\n"},{"id":"19605","messageId":"Pine.LNX.4.64.0605060821430.16343@g5.osdl.org","threadId":"3867","inReplyTo":"7vhd43vgnm.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-06T15:26:36Z","receivedAt":"2006-05-06T15:26:36Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 May 2006, Junio C Hamano wrote:\n>\n> > So I'd argue that (a) yes, we do want to have the \"proto porcelain\" that \n> > sets remote branch information without the user having to know the magic \n> > \"git repo-config\" incantation, or know which file in .git/remotes/ to \n> > edit, but that (b) it's even more important to try to decide on what the \n> > remote description format _is_.\n> \n> Is it format you care about or the semantics?\n\nI _personally_ care about the semantics, but not very deeply - since I \ntend to actually have just one main branch, and a couple of throw-away \nones if I ended up working on something.\n\nBut I think that for this thing to become useful, we want to care about \nthe format - or at least the interface to the different users (with the \nacknowledgement that \"users\" should often be porcelain above us).\n\nRight now we've basically had people hand-editing the remotes files, and I \nthink cogito still uses the older branches format that came from cogito in \nthe first place. I think we should just try to decide on a config file \nformat, and make it easy for cogito etc to use it.\n\n\t\tLinus\n"},{"id":"19606","messageId":"BAYC1-PASMTP10F63ADF30C26A29D070C5AEAA0@CEZ.ICE","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0605060821430.16343@g5.osdl.org","subject":"Re: Unresolved issues #2","fromName":"sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2006-05-06T15:35:49Z","receivedAt":"2006-05-06T15:35:49Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Sat, 6 May 2006 08:26:36 -0700 (PDT)\nLinus Torvalds <torvalds@osdl.org> wrote:\n\n> I _personally_ care about the semantics, but not very deeply - since I \n> tend to actually have just one main branch, and a couple of throw-away \n> ones if I ended up working on something.\n> \n> But I think that for this thing to become useful, we want to care about \n> the format - or at least the interface to the different users (with the \n> acknowledgement that \"users\" should often be porcelain above us).\n> \n> Right now we've basically had people hand-editing the remotes files, and I \n> think cogito still uses the older branches format that came from cogito in \n> the first place. I think we should just try to decide on a config file \n> format, and make it easy for cogito etc to use it.\n\nLinus,\n\nWondering why you feel so strongly that most \"users\" shouldn't be real people.\nWhat is wrong with continuing to make git easier for developers to use without\nneeding any extra software?\n\nSean\n"},{"id":"19607","messageId":"Pine.LNX.4.64.0605060923050.16343@g5.osdl.org","threadId":"3867","inReplyTo":"BAYC1-PASMTP10F63ADF30C26A29D070C5AEAA0@CEZ.ICE","subject":"Re: Unresolved issues #2","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-06T16:30:48Z","receivedAt":"2006-05-06T16:30:48Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 6 May 2006, sean wrote:\n> \n> Wondering why you feel so strongly that most \"users\" shouldn't be real people.\n> What is wrong with continuing to make git easier for developers to use without\n> needing any extra software?\n\nBasically, it boils down to the end result.\n\nIf you design things for \"people\", then things tend to become hard to \nautomate, and it's hard to make wrappers around it. Maybe you've even made \nthe interfaces interactive, and thus any wrappers around it are simply \nscrewed, or need to do insane things.\n\nOn the other hand, if you design things for automation, doing a \"people \nwrapper\" that uses the automation should be trivial if the design is even \nremotely any good at all.\n\nIn other words: you should always design things for automation, and \nconsider the \"people interface\" to be be just _one_ wrapper layer among \nmany.\n\nThis has worked really well in git. The whole system was designed from the \nstart to be all about scripting and automation, and the \"people wrappers\" \ntend to be trivial scripts around it.\n\nThis was even more obvious when we had a number of basically one-liner \nscripts like \"git log\", which just did some trivial wrapping around\n\n\tgit-rev-list | git-diff-tree --stdin | $PAGER\n\n(Now we still have that trivial wrapper, but you just need to look into C \ncode to see it, so it's not _as_ obviously trivial).\n\nContrast this with going the other way: if you talk about the interfaces \nthat _people_ want first, you immediately start doing pretty-printing, \nnice parsing, maybe interactive stuff that asks questions. Nice GUIs. And \nthe end result is CRAP. Exactly because it lost its ability to be generic.\n\nTo some degree, this is the fundamental difference between the Windows and \nthe UNIX mindset. At least it used to be.\n\n\t\tLinus\n"},{"id":"19608","messageId":"BAYC1-PASMTP0824AA77198F95FE28B79DAEAA0@CEZ.ICE","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0605060923050.16343@g5.osdl.org","subject":"Re: Unresolved issues #2","fromName":"sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2006-05-06T16:53:23Z","receivedAt":"2006-05-06T16:53:23Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Sat, 6 May 2006 09:30:48 -0700 (PDT)\nLinus Torvalds <torvalds@osdl.org> wrote:\n\n> Basically, it boils down to the end result.\n> \n> If you design things for \"people\", then things tend to become hard to \n> automate, and it's hard to make wrappers around it. Maybe you've even made \n> the interfaces interactive, and thus any wrappers around it are simply \n> screwed, or need to do insane things.\n\nOkay, I mistook the scope of you comments to apply to all of git rather than\nas a reminder that we can't forget about the toolkit design.  So I take it\nyou're not at all against git including higher level user commands; just so\nlong as they're built on top of lower level toolkit commands that other\nporcelain can use as well.\n\nIn this particular case I see \"git repo-config\" as the low level command that\nany porcelain can use to access the remotes information and the proposed\n\"git remotes\" as a simple convenience wrapper on top of this.  Of course,\neveryone has to agree on the config file format; but that is true whether\nthe human-friendly wrapper exists or not.\n\nSean\n"},{"id":"19609","messageId":"Pine.LNX.4.64.0605061008340.16343@g5.osdl.org","threadId":"3867","inReplyTo":"BAYC1-PASMTP0824AA77198F95FE28B79DAEAA0@CEZ.ICE","subject":"Re: Unresolved issues #2","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-06T17:20:15Z","receivedAt":"2006-05-06T17:20:15Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 6 May 2006, sean wrote:\n>\n> Okay, I mistook the scope of you comments to apply to all of git rather than\n> as a reminder that we can't forget about the toolkit design.  So I take it\n> you're not at all against git including higher level user commands; just so\n> long as they're built on top of lower level toolkit commands that other\n> porcelain can use as well.\n\nCorrect. I think we've been able to handle that balance particularly well \nso far. Or maybe the porcelains don't complain enough.\n\n> In this particular case I see \"git repo-config\" as the low level command that\n> any porcelain can use to access the remotes information and the proposed\n> \"git remotes\" as a simple convenience wrapper on top of this.  Of course,\n> everyone has to agree on the config file format; but that is true whether\n> the human-friendly wrapper exists or not.\n\nI agree, but my point is that in order for a porcelain to _use_ \n\"repo-config\", the config file format needs to be defined somewhere, and \nwe need to tell people that it's not changing. Are we there yet?\n\nThat was my argument for why we should concentrate not on what the user \nwrapper should be named, but why we should look at what the low-level \nmeaning of these things are.\n\nFinally, I think \"git repo-config\" is buggy. Try with this .config file:\n\n\t[user]\n\t\tname = Bozo the Clown\n\t\temail = bozo@circus.com\n\n\t[core]\n\t\tfilemode = true\n\n\t[merge]\n\t\tsummary = true\n\nand then do\n\n\tgit repo-config core.gitproxy 'dummy example'\n\nand look where it ends up. For me, it ends up at the end, in the \"[merge]\" \nsection, which is obviously bogus.\n\nSo we'd really be screwing with porcelain if we made them use this ;)\n\n\t\tLinus\n"},{"id":"19616","messageId":"7vvesirh0q.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0605061008340.16343@g5.osdl.org","subject":"Re: Unresolved issues #2","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-06T21:16:05Z","receivedAt":"2006-05-06T21:16:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Finally, I think \"git repo-config\" is buggy. Try with this .config file:\n> ...\n> So we'd really be screwing with porcelain if we made them use this ;)\n\nThanks Linus and Sean for bringing this up and fixing it.\n\nI have a vague feeling that this may not be the last breakage of\nthe repo-config command.  My first reaction to the repo-config\ncode was \"eek\".  It tries to reuse as much the existing material\nas possible -- I understand it was done that way in order to\npreserve the comments and blank lines from the original config\nfile intact, but it just felt very error prone (demonstrated by\ncases like this and the other one Sean brought up) and generally\nwrong.\n\nIt might make sense to rewrite it to parse and read the existing\nconfiguration as a whole, do necessary manupulations on the\nparsed internal representation in-core, and write the result out\nfrom scratch.  That would fix another of my pet peeve: after an\ninvocation of repo-config to remove the last variable in a\nsection, it leaves an empty section header in.\n\n        $ git repo-config foo.bar true\n        $ cat .git/config \n        [core]\n                repositoryformatversion = 0\n                filemode = true\n        [foo]\n                bar = true\n        $ git repo-config foo1.baz false\n        $ git repo-config --unset foo.bar\n        $ cat .git/config \n        [core]\n                repositoryformatversion = 0\n                filemode = true\n        [foo]\n        [foo1]\n                baz = false\n"},{"id":"19617","messageId":"Pine.LNX.4.63.0605062332420.6423@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3867","inReplyTo":"7vvesirh0q.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-05-06T21:33:26Z","receivedAt":"2006-05-06T21:33:26Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sat, 6 May 2006, Junio C Hamano wrote:\n\n> Linus Torvalds <torvalds@osdl.org> writes:\n> \n> > Finally, I think \"git repo-config\" is buggy. Try with this .config file:\n> > ...\n> > So we'd really be screwing with porcelain if we made them use this ;)\n> \n> Thanks Linus and Sean for bringing this up and fixing it.\n> \n> I have a vague feeling that this may not be the last breakage of\n> the repo-config command.  My first reaction to the repo-config\n> code was \"eek\".  It tries to reuse as much the existing material\n> as possible -- I understand it was done that way in order to\n> preserve the comments and blank lines from the original config\n> file intact, but it just felt very error prone (demonstrated by\n> cases like this and the other one Sean brought up) and generally\n> wrong.\n\nIt was done because the very syntax of the config suggests it be a \nuser-editable file. I do not want to mess with the comments more than \nnecessary.\n\n> It might make sense to rewrite it to parse and read the existing\n> configuration as a whole, do necessary manupulations on the\n> parsed internal representation in-core, and write the result out\n> from scratch.  That would fix another of my pet peeve: after an\n> invocation of repo-config to remove the last variable in a\n> section, it leaves an empty section header in.\n\nDoes it really hurt? I think not.\n\nAnyway, I'll look into this.\n\nCiao,\nDscho\n"},{"id":"19619","messageId":"Pine.LNX.4.64.0605061442490.16343@g5.osdl.org","threadId":"3867","inReplyTo":"Pine.LNX.4.63.0605062332420.6423@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Unresolved issues #2","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-06T21:51:05Z","receivedAt":"2006-05-06T21:51:05Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 6 May 2006, Johannes Schindelin wrote:\n> \n> It was done because the very syntax of the config suggests it be a \n> user-editable file.\n\nYeah, I personally much prefer user-friendly config files. Any format that \nthinks that \"easy parsing\" is more important than \"visually obvious\" is \nbad. So I obviously think that XML is a horrid piece of cr*p (has anybody \never noticed I have strong opinions?) and totally unreadable.\n\nI think \"git repo-config\" is doing a reasonable job of editing a file that \nis really designed to be user-friendly. That said, the code _is_ a bit \nscary.\n\nIt might be worthwhile to re-write config.c to read the config file into \nmemory and work on it in-memory instead of doing the funky mixed usage \n(using fgetc/ftell to read it, but then switching over to mmap when \nrewriting it).\n\nIOW, maybe that \"static FILE *config_file\" should be changed to something \nmore like \"static const char *config_buffer; unsigned int len;\" instead, \nand at least make both the reading and writing use the same buffer rather \nthan mixing stdio and mmap..\n\n\t\tLinus\n"},{"id":"19631","messageId":"e3jfmh$1vc$1@sea.gmane.org","threadId":"3867","inReplyTo":"7vvesirh0q.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-07T00:41:00Z","receivedAt":"2006-05-07T00:41:00Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n> It might make sense to rewrite it to parse and read the existing\n> configuration as a whole, do necessary manupulations on the\n> parsed internal representation in-core, and write the result out\n> from scratch.\n\nOr perhaps do git repo-config read and change config file in two passes:\nread and build some kind of index (beginning of section, end of\nsection/last variable in section, number of elements in section), then in\nsecond pass add some information if needed.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19638","messageId":"46a038f90605062308x53995076k7bf45f0aebcae0c6@mail.gmail.com","threadId":"3867","inReplyTo":"7vbqubvdbr.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-07T06:08:03Z","receivedAt":"2006-05-07T06:08:03Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/6/06, Junio C Hamano <junkio@cox.net> wrote:\n> \"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n> >\n> > It means that for a merge or checkout involving stuff we \"don't have\",\n> > it's trivial to know we are missing, and so we can  attempt a fetch of\n> > the missing objects or tell the user how to request them them before\n> > retrying.\n> >\n> > And in any case commits and trees are lightweight and compress well...\n>\n> Commit maybe, but is this based on a hard fact?\n\nNo hard facts here :( but I think it's reasonable to assume that the\ntrees delta/compress reasonably well, as a given commit will change\njust a few entries in each tree.\n\nI might try and hack a shallow local clone of the kernel and pack it\ntightly to see what it yields.\n\ncheers,\n\n\n\nmartin\n"},{"id":"19640","messageId":"20060507075631.GA24423@coredump.intra.peff.net","threadId":"3867","inReplyTo":"46a038f90605062308x53995076k7bf45f0aebcae0c6@mail.gmail.com","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-05-07T07:56:31Z","receivedAt":"2006-05-07T07:56:31Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, May 07, 2006 at 06:08:03PM +1200, Martin Langhoff wrote:\n\n> >> And in any case commits and trees are lightweight and compress well...\n> >Commit maybe, but is this based on a hard fact?\n> No hard facts here :( but I think it's reasonable to assume that the\n> trees delta/compress reasonably well, as a given commit will change\n> just a few entries in each tree.\n\nA few hard facts (using Linus' linux-2.6 tree):\n  - original packsize: 120996 kilobytes\n  - unpacked: 233338 objects, 1417476 kilobytes\n    This is an 11.7:1 compression ratio (of course, much of this is\n    wasted space from the 4k block size in the filesystem)\n  - There were 87915 total blob objects, of which 19321 were in the\n    current tree. I removed all non-current blobs to produce a \"shallow\"\n    tree.\n  - The shallow tree unpacked: 164744 objects, 761960 kilobytes\n    IOW, about half of the unpacked disk usage was old blobs.\n  - Shallow commit/tree/tag objects packed (using 1.3.1\n    git-pack-objects):\n      Total 164744, written 164744 (delta 92322), reused 0 (delta 0)\n      size: 108088\n    The compression ratio here is only 7.0:1\n  - Total savings by going shallow: 10.7%\n\nSo basically, trees and commits DON'T compress as well as historical\nblobs (potentially because git-pack-objects isn't currently optimized\nfor this -- I haven't checked). As a result, we're saving only 10% by\ngoing shallow instead of a potential 50%.\n\n-Peff\n"},{"id":"19641","messageId":"20060507120149.40e9f749.vsu@altlinux.ru","threadId":"3867","inReplyTo":"46a038f90605062308x53995076k7bf45f0aebcae0c6@mail.gmail.com","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Sergey Vlasov","fromEmail":"vsu@altlinux.ru","sentAt":"2006-05-07T08:01:49Z","receivedAt":"2006-05-07T08:01:49Z","isPatch":false,"sender":{"key":"vsu@altlinux.ru","avatar":"https://avatars.githubusercontent.com/u/616082?v=4"},"body":"On Sun, 7 May 2006 18:08:03 +1200 Martin Langhoff wrote:\n\n> On 5/6/06, Junio C Hamano <junkio@cox.net> wrote:\n> > \"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n> > >\n> > > It means that for a merge or checkout involving stuff we \"don't have\",\n> > > it's trivial to know we are missing, and so we can  attempt a fetch of\n> > > the missing objects or tell the user how to request them them before\n> > > retrying.\n> > >\n> > > And in any case commits and trees are lightweight and compress well...\n> >\n> > Commit maybe, but is this based on a hard fact?\n> \n> No hard facts here :( but I think it's reasonable to assume that the\n> trees delta/compress reasonably well, as a given commit will change\n> just a few entries in each tree.\n> \n> I might try and hack a shallow local clone of the kernel and pack it\n> tightly to see what it yields.\n\nFor linux v2.6.16:\n\n7,3M commits-b41b04a36afebdba3b70b74f419fc7d97249bd7f.pack\n 24M commits_trees-8397f1c2a885527acd07e2caa8c95df626451493.pack\n 97M full-c7b2747a674ff55cb4a59dabebe419f191e360df.pack\n\nFor comparizon, a single version in packed form:\n\n 51M v2.6.12-rc2-4f3526b6815eb63da6c43ed85be1494bb776e2c5.pack\n\nMade with\n\ngit-rev-list v2.6.16 | git-pack-objects commits\ngit-rev-list --objects --no-blobs v2.6.16 | git-pack-objects commits_trees\ngit-rev-list --objects v2.6.16 | git-pack-objects full\n\nand this hack to git-rev-list:\n\ndiff --git a/revision.c b/revision.c\nindex f2a9f25..b5a929e 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -636,6 +636,10 @@ int setup_revisions(int argc, const char\n \t\t\t\trevs->blob_objects = 1;\n \t\t\t\tcontinue;\n \t\t\t}\n+\t\t\tif (!strcmp(arg, \"--no-blobs\")) {\n+\t\t\t\trevs->blob_objects = 0;\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tif (!strcmp(arg, \"--objects-edge\")) {\n \t\t\t\trevs->tag_objects = 1;\n \t\t\t\trevs->tree_objects = 1;\n\nSo trees are definitely not lightweight, and commits are rather large\ntoo.\n"},{"id":"19643","messageId":"7vy7xekwbs.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"Pine.LNX.4.63.0605062332420.6423@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Unresolved issues #2","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-07T09:39:35Z","receivedAt":"2006-05-07T09:39:35Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> It was done because the very syntax of the config suggests it be a \n> user-editable file. I do not want to mess with the comments more than \n> necessary.\n\nI personally feel that is a lost cause _unless_ you come up with\na reasonable convention for where to put comments, stress that\nrule to the user in the documentation, _and_ make repo-config to\nfollow that rule as well.\n\nWe _do_ want to treat config file as hand editable and cat\nreviewable file, not an unreadable gunk like xml, so trying to\npreserve user comments is important and I am not opposed to that\nyou did (at least some of) it.  But as the code currently\nstands, what it does is at best half baked, at worst somewhat\nconfusing.\n\nA demonstration.  What is wrong with this picture?\n\n        $ cat .git/config\n        [core]\n                repositoryformatversion = 0\n                ; are the mode bits trustworthy?\n                filemode = true ; yes, on ext3 \n                ; We want symlinked HEAD because we will bisect\n                ; recent kernel history.\n                prefersymlinkrefs = true\n        $ git repo-config core.prefersymlinkrefs false\n        $ git repo-config core.filemode false\n        $ cat .git/config\n        [core]\n                repositoryformatversion = 0\n                ; are the mode bits trustworthy?\n                filemode = false\n                ; We want symlinked HEAD because we will bisect\n                ; recent kernel history.\n                prefersymlinkrefs = false\n\t$ exit\n\nThe comment given to \"filemode\" is \"reasonable\" in that it\ndescribes what the value that is set to the variable does, and\nlosing the original comment given to its \"true\" when we set it\nto false is better than keeping it, so that part happens to be\ndoing the right thing -- only because I knew what repo-config\nwould do to the comments and arranged original comments in the\nfile that way.\n\nBut what about prefersymlinkrefs one?  When setting the variable\nto such a non-standard value, it is unreasonable for people to\nwant to justify why with a comment like the above.  But after\nresetting the value the comment becomes stale.\n\nIt gets worse:\n\n        $ git repo-config --unset core.filemode\n        $ cat .git/config\n        [core]\n                repositoryformatversion = 0\n                ; please please use symlinks please\n                prefersymlinkrefs = false\n                ; are the mode bits trustworthy?\n\t$ exit\n\nThere now is a confusing trailing comment left that does not\ncomment anything.  Removing core.filemode is not so common, but\nthis can happen whenever you remove any variable, so we can use\nany other variable as an example.\n\nNow, enough being negative and pointing out problems.  Time to\nbecome constructive.  Probably a reasonable convention would be\nto define the config file format to be something like this:\n\n        <comment that applies to the section>\n        [section]\n                <comment that applies to the variable stands on\n\t\t its own before the variable>\n                variable [= value] <comment that applies to the\n        \t\t\t    fact the variable is set to\n                                    this particular value starts\n\t\t\t\t    on the same line as the\n                                    \"variable = value\" thing>\n\n - when a variable is reset to another value, remove the\n   \"value comment\";\n - when a variable disappears, remove \"variable comment\";\n - when a section disappears, remove \"section comment\";\n - otherwise leave the comment intact.\n\nThen we could tell the user the rule is like above, and tell\nthem to structure the file with comments that way, if they ever\nwant to edit the file by hand.\n\nNow if we wanted to do something like the above, I suspect that\nit would be easier and less error prone to first scan the config\nfile, note what appears where, and do the processing in-core,\nand then write the results out, perhaps using data structures\nlike these:\n\n        struct config_section {\n            char *pre_comment;\n            char *name; /* e.g. \"core\" */\n            struct config_section *next; /* next section */\n            struct config_var *vars; /* pointer to the first one */\n        };\n        struct config_var {\n            char *pre_comment;\n            char *name;\n            char *value; /* \"existence\" bool may have NULL,\n                          * otherwise probably a string \"= value\"\n                          */\n            char *value_comment;\n            struct config_var *next; /* pointer to the next one\n                                      * in the section\n                                      */\n        };\n\nObviously, data structures like these would make it even easier\nif we decide we do _not_ care about comments (we would just lose\nx_comment fields, parse the thing and write the resulting list\nout).\n"},{"id":"19644","messageId":"7vu082kw6o.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"7vy7xekwbs.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-07T09:42:39Z","receivedAt":"2006-05-07T09:42:39Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> But what about prefersymlinkrefs one?  When setting the variable\n> to such a non-standard value, it is unreasonable for people to\n> want to justify why with a comment like the above.\n\nObviously I was not reading what I was typing.  It is very\nreasonable for people to want to do that.  Sorry.\n"},{"id":"19645","messageId":"Pine.LNX.4.63.0605071330210.22231@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3867","inReplyTo":"7vy7xekwbs.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-05-07T11:31:53Z","receivedAt":"2006-05-07T11:31:53Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 7 May 2006, Junio C Hamano wrote:\n\n> [...] Probably a reasonable convention would be to define the config \n> file format to be something like this:\n> \n>         <comment that applies to the section>\n>         [section]\n>                 <comment that applies to the variable stands on\n> \t\t its own before the variable>\n>                 variable [= value] <comment that applies to the\n>         \t\t\t    fact the variable is set to\n>                                     this particular value starts\n> \t\t\t\t    on the same line as the\n>                                     \"variable = value\" thing>\n> \n>  - when a variable is reset to another value, remove the\n>    \"value comment\";\n>  - when a variable disappears, remove \"variable comment\";\n>  - when a section disappears, remove \"section comment\";\n>  - otherwise leave the comment intact.\n> \n> Then we could tell the user the rule is like above, and tell\n> them to structure the file with comments that way, if they ever\n> want to edit the file by hand.\n> \n> Now if we wanted to do something like the above, I suspect that\n> it would be easier and less error prone to first scan the config\n> file, note what appears where, and do the processing in-core,\n> and then write the results out, perhaps using data structures\n> like these:\n> \n>         struct config_section {\n>             char *pre_comment;\n>             char *name; /* e.g. \"core\" */\n>             struct config_section *next; /* next section */\n>             struct config_var *vars; /* pointer to the first one */\n>         };\n>         struct config_var {\n>             char *pre_comment;\n>             char *name;\n>             char *value; /* \"existence\" bool may have NULL,\n>                           * otherwise probably a string \"= value\"\n>                           */\n>             char *value_comment;\n>             struct config_var *next; /* pointer to the next one\n>                                       * in the section\n>                                       */\n>         };\n> \n> Obviously, data structures like these would make it even easier\n> if we decide we do _not_ care about comments (we would just lose\n> x_comment fields, parse the thing and write the resulting list\n> out).\n\nSounds very reasonable.\n\nCiao,\nDscho\n"},{"id":"19646","messageId":"e3km6q$f7p$1@sea.gmane.org","threadId":"3867","inReplyTo":"7vy7xekwbs.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-07T11:38:16Z","receivedAt":"2006-05-07T11:38:16Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n>             char *value; /* \"existence\" bool may have NULL,\n>                           * otherwise probably a string \"= value\"\n>                           */\n\nProbably \" = value\" to preserve whitespace (e.g. justify on equal sign in\nhand crafted config file).\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19648","messageId":"e3ksoq$is$1@sea.gmane.org","threadId":"3867","inReplyTo":"7v1wv92u7o.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-07T13:30:15Z","receivedAt":"2006-05-07T13:30:15Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n> The vocabulary we would want from the requestor side is probably\n> (at least):\n> \n>         I WANT to have these\n>         I HAVE these\n>         I'm MISSING these\n>         Don't bother with these this time around (--since, ^v2.6.16, ...)\n\nWouldn't it be easier (sorry, no code yet) to have the following:\n\n        I WANT to have these\n        I HAVE these\n        These are GRAFT PARENTLESS        \n\nwith the target side sending list of all parentless commits in the\ninfo/grafts file. The source side will then do the grafting 'in memory' and\nsend the packs like normal, only with those cauterizing grafts in place.\n\nNow I'm waiting for someone to say that it is too simple and cannot be done,\nor that shallow clone/shallow fetch uses this method...\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19650","messageId":"Pine.LNX.4.64.0605070802590.16343@g5.osdl.org","threadId":"3867","inReplyTo":"20060507075631.GA24423@coredump.intra.peff.net","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-07T15:27:02Z","receivedAt":"2006-05-07T15:27:02Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 7 May 2006, Jeff King wrote:\n>\n>   - Total savings by going shallow: 10.7%\n> \n> So basically, trees and commits DON'T compress as well as historical\n> blobs (potentially because git-pack-objects isn't currently optimized\n> for this -- I haven't checked). As a result, we're saving only 10% by\n> going shallow instead of a potential 50%.\n\nThe biggest size savers from packing is (in rough order of relevance, if \nI recall the rough statistics I did):\n\n - avoiding block boundaries. \n - delta packing of blobs\n - delta packing of trees\n - regular compression\n\nThe block boundaries are huge, we have tons of small objects, and that was \none of the primary reasons for packing. I'd suspect that this is a 3:1 \nfactor for a lot of things for many \"common\" filesystem setups. You \nprobably didn't even account for the size of inodes in your \"du\" setup.\n\nAnd blobs with history generally delta very well (_much_ better than \nregular compression).\n\nTrees should _delta_ very well, but they basically don't compress, \nespecially after deltaing. The SHA1's are totally incompressible (in a \ntree they aren't even ASCII), and as a deta, the names won't compress much \neither because they are short.\n\nCommits are fairly small, shouldn't delta all that much, and they don't \neven compress _that_ well either (they're normal text and often have some \nredundancy with the committer and author being the same, but they are \nshort and have some fairly incompressible elements, so..)\n\nThe thing with trees in particular is that they are very common for the \nkernel (and probably not so much for many other projects). A single commit \nends up quite commonly being just one commit object, one blob (that deltas \nreally well), and three or four trees. Merges often have no new blobs at \nall, just several new trees and the commit object.\n\nSo a huge amount of the wins from packing come from the file _history_, \nthe part that a shallow clone (on purpose) leaves behind. \n\nThe regular compression will pick up a fair amount of slack with the \nblobs, but it's a much smaller factor than the delta compression for \nsomething that has a long history.\n\nIt's somewhat interesting to note that over the year that we've used git, \nthe kernel pack-size hasn't even increased all that much. I forget exactly \nwhat it was when we started packing, but it was on the order of ~75M. It \nis now 115M for me. And the old linux-history thing (full BK history over \nthree years) is 177M - not much more than twice the size of just a few \nkernel versions - with some higher packing ratios..\n\nExactly because blobs delta so incredibly well.\n\n\t\tLinus\n"},{"id":"19668","messageId":"46a038f90605071627i6a335f61lf5e35291bfbe340c@mail.gmail.com","threadId":"3867","inReplyTo":"20060507120149.40e9f749.vsu@altlinux.ru","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-07T23:27:24Z","receivedAt":"2006-05-07T23:27:24Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/7/06, Sergey Vlasov <vsu@altlinux.ru> wrote:\n> For linux v2.6.16:\n>\n> 7,3M commits-b41b04a36afebdba3b70b74f419fc7d97249bd7f.pack\n>  24M commits_trees-8397f1c2a885527acd07e2caa8c95df626451493.pack\n>  97M full-c7b2747a674ff55cb4a59dabebe419f191e360df.pack\n\nWith this pack arrangement, do you get any noticeable difference in\nwalking commits? How about walking commits+trees with git-log <path> ?\n\nI wonder whether segregating packs by object type would make things better...\n\ncheers,\n\n\n\nmartin\n"},{"id":"19669","messageId":"7vveshif1i.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"46a038f90605071627i6a335f61lf5e35291bfbe340c@mail.gmail.com","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-07T23:35:53Z","receivedAt":"2006-05-07T23:35:53Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n\n> On 5/7/06, Sergey Vlasov <vsu@altlinux.ru> wrote:\n>> For linux v2.6.16:\n>>\n>> 7,3M commits-b41b04a36afebdba3b70b74f419fc7d97249bd7f.pack\n>>  24M commits_trees-8397f1c2a885527acd07e2caa8c95df626451493.pack\n>>  97M full-c7b2747a674ff55cb4a59dabebe419f191e360df.pack\n>\n> With this pack arrangement, do you get any noticeable difference in\n> walking commits? How about walking commits+trees with git-log <path> ?\n>\n> I wonder whether segregating packs by object type would make things better...\n\nIt shouldn't.  The existing packfile is designed to make \"git\nlog\" very efficient, by making it cheap to look only at the\ncommit message and ancestry information.\n\nThe objects are sorted first by type in the pack with the\nexisting code already, and commits come first.  Try this.\n\n        git repack -a -d\n        git show-index <.git/objects/pack/pack-*.idx |\n        sort -n |\n        while read offset objectname\n        do\n                git cat-file -t \"$objectname\"\n        done\n"},{"id":"19671","messageId":"46a038f90605071644g7c7628celc097f69b310c2db5@mail.gmail.com","threadId":"3867","inReplyTo":"7vveshif1i.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-07T23:44:37Z","receivedAt":"2006-05-07T23:44:37Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/8/06, Junio C Hamano <junkio@cox.net> wrote:\n> It shouldn't.  The existing packfile is designed to make \"git\n> log\" very efficient, by making it cheap to look only at the\n> commit message and ancestry information.\n>\n> The objects are sorted first by type in the pack with the\n> existing code already, and commits come first.  Try this.\n\nThanks for the explanation! My ignorance in the pre-coffee stage of\nMonday is... exemplary. I should know better than assume that there\nare trivial optimizations for git that I can come up with ;-)\n\nI guess I should get more into the pack stuff and educate myself --\neveryone's having fun with it... grumble, will have to brush up my C\nto get to play.\n\ncheers,\n\n\nmartin\n"},{"id":"19676","messageId":"20060508003338.GB17138@thunk.org","threadId":"3867","inReplyTo":"20060507075631.GA24423@coredump.intra.peff.net","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2006-05-08T00:33:38Z","receivedAt":"2006-05-08T00:33:38Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Sun, May 07, 2006 at 03:56:31AM -0400, Jeff King wrote:\n> On Sun, May 07, 2006 at 06:08:03PM +1200, Martin Langhoff wrote:\n> \n> > >> And in any case commits and trees are lightweight and compress well...\n> > >Commit maybe, but is this based on a hard fact?\n> > No hard facts here :( but I think it's reasonable to assume that the\n> > trees delta/compress reasonably well, as a given commit will change\n> > just a few entries in each tree.\n> \n> A few hard facts (using Linus' linux-2.6 tree):\n>   - original packsize: 120996 kilobytes\n>   - unpacked: 233338 objects, 1417476 kilobytes\n>     This is an 11.7:1 compression ratio (of course, much of this is\n>     wasted space from the 4k block size in the filesystem)\n\nIf there are 233338 objects, then the average wasted space due to\ninternal fragmentation is 233338 * 2k, or 466676 kilobytes, or only\n36% of the wasted space.  Most of the savings is probably coming from\nthe compression and delta packing.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"19681","messageId":"Pine.LNX.4.64.0605071744210.3718@g5.osdl.org","threadId":"3867","inReplyTo":"20060508003338.GB17138@thunk.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-08T00:50:42Z","receivedAt":"2006-05-08T00:50:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 7 May 2006, Theodore Tso wrote:\n>> \n> If there are 233338 objects, then the average wasted space due to\n> internal fragmentation is 233338 * 2k, or 466676 kilobytes, or only\n> 36% of the wasted space.\n\nThat's not necessarily true.\n\nThat assumes a randomly distributed filesize. File sizes are _not_ random, \nand in particular if you have the distribution leaning towards <2kB being \ncommon, you can actually get >50% fragmentation.\n\nBtw, I hit this when some people argued that the page size should be made \n64kB. The above (incorrect) logic implies that you waste 32kB on average \nper file. That's not true, if a large fraction of your files are small, in \nwhich case you may actually be wastign closer to 60kB on average from \nusing a big page-size, because about half of the kernel files are actually \nsmaller than 4kB (or something. I forget the exact statistics, I did them \nwith a script at some point).\n\nAnyway, with inode overhead and a lot of objects being just a couple of \nhundred bytes, I think I estimated at some point that you actually lost \ncloser to 3kB per object.\n\nMany of the objects actually end up being smaller than the inode they end \nup allocating ;(\n\n\t\t\tLinus\n"},{"id":"19688","messageId":"20060508012632.GD17138@thunk.org","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0605071744210.3718@g5.osdl.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2006-05-08T01:26:32Z","receivedAt":"2006-05-08T01:26:32Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Sun, May 07, 2006 at 05:50:42PM -0700, Linus Torvalds wrote:\n> \n> \n> On Sun, 7 May 2006, Theodore Tso wrote:\n> >> \n> > If there are 233338 objects, then the average wasted space due to\n> > internal fragmentation is 233338 * 2k, or 466676 kilobytes, or only\n> > 36% of the wasted space.\n> \n> That's not necessarily true.\n> \n> That assumes a randomly distributed filesize. File sizes are _not_ random, \n> and in particular if you have the distribution leaning towards <2kB being \n> common, you can actually get >50% fragmentation.\n> \n> Btw, I hit this when some people argued that the page size should be made \n> 64kB. The above (incorrect) logic implies that you waste 32kB on average \n> per file. That's not true, if a large fraction of your files are small, in \n> which case you may actually be wastign closer to 60kB on average from \n> using a big page-size, because about half of the kernel files are actually \n> smaller than 4kB (or something. I forget the exact statistics, I did them \n> with a script at some point).\n> \n> Anyway, with inode overhead and a lot of objects being just a couple of \n> hundred bytes, I think I estimated at some point that you actually lost \n> closer to 3kB per object.\n\nI just ran the numbers on filesizes of a kernel tree I had handy,\nwhich happened to be 2.6.16.11.  With no object files, git files,\netc. the average loss was 2351 bytes --- not that far away from the\naverage of 2048 bytes.  Granted, it may be there is more different\nversions of small objects causing a skewing of the distributions of\ngit objects in the 2.6 tree, but I'm not familiar enough with the git\nporcelain to be able to make it disgorge the sizes of the repository\nto do the math.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"19695","messageId":"Pine.LNX.4.64.0605071853290.3718@g5.osdl.org","threadId":"3867","inReplyTo":"20060508012632.GD17138@thunk.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-08T02:04:48Z","receivedAt":"2006-05-08T02:04:48Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 7 May 2006, Theodore Tso wrote:\n> \n> I just ran the numbers on filesizes of a kernel tree I had handy,\n> which happened to be 2.6.16.11.  With no object files, git files,\n> etc. the average loss was 2351 bytes --- not that far away from the\n> average of 2048 bytes.\n\nIs that without compression?\n\ngit objects are compressed, and common types (trees) tend to be smaller \nthan your normal C file.\n\nSo git objects tend to be _smaller_ than the regular files. By about 30%. \nIn addition, the non-blob git objects themselves tend to be smaller still.\n\nSo for example, right now I have just 58 unpacked objects (I repack pretty \noften). But of those 58 objects, exactly _fifty_ are smaller than 2kB, and \n38 are smaller than 1kB. The median size is 771 bytes.\n\nOn master.kernel.org, I've not repacked as recently, so I've got 2268 \nunpacked objects. But the median size there is 773 bytes, so it looks like \nthe numbers are statistically pretty stable.\n\n\t\t\tLinus\n"},{"id":"19696","messageId":"20060508022432.GA26076@thunk.org","threadId":"3867","inReplyTo":"Pine.LNX.4.64.0605071853290.3718@g5.osdl.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2006-05-08T02:24:32Z","receivedAt":"2006-05-08T02:24:32Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Sun, May 07, 2006 at 07:04:48PM -0700, Linus Torvalds wrote:\n> Is that without compression?\n\nYes, without compression.  So yes, that probably explains the\ndifference between your numbers and mine. \n\nThat brings up an interesting question though --- why not skip\ncompressing files that are under 4k (or whatever the filesystem\nblocksize happens to be) if they are unpacked?  It burns CPU time;\nmaybe not enough to be human-noticeable, but it's still not buying you\nanything.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"19699","messageId":"Pine.LNX.4.64.0605071939291.3718@g5.osdl.org","threadId":"3867","inReplyTo":"20060508022432.GA26076@thunk.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-08T02:42:44Z","receivedAt":"2006-05-08T02:42:44Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 7 May 2006, Theodore Tso wrote:\n> \n> That brings up an interesting question though --- why not skip\n> compressing files that are under 4k (or whatever the filesystem\n> blocksize happens to be) if they are unpacked?  It burns CPU time;\n> maybe not enough to be human-noticeable, but it's still not buying you\n> anything.\n\nWell, other filesystems don't have 4kB issues. Reiser can do smaller \nthings iirc, and you might obviously have a ext3 filesystem with a 1kB \nblocksize too. And with tails on FFS, you might have a filesystem with a \n8kB blocksize, but despite that it might lay out <1kB files well.\n\nAnyway, packing makes all this basically a non-issue. There are no block \nboundaries in a pack-file, and you only use a single inode. And you'd \nobviously want to pack for other reasons anyway (ie the delta compression \nwill makea huge difference over time).\n\n\t\tLinus\n"},{"id":"19701","messageId":"7v3bfli602.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"e3km6q$f7p$1@sea.gmane.org","subject":"Re: Unresolved issues #2","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-08T02:51:09Z","receivedAt":"2006-05-08T02:51:09Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> Junio C Hamano wrote:\n>\n>>             char *value; /* \"existence\" bool may have NULL,\n>>                           * otherwise probably a string \"= value\"\n>>                           */\n>\n> Probably \" = value\" to preserve whitespace (e.g. justify on equal sign in\n> hand crafted config file).\n\nProbably even better is to remove the separate *value_comment,\nand make this thing point at the entire \" = value ; this is the\ncomment for the value\\n\" thing.\n"},{"id":"19702","messageId":"7vy7xdgram.fsf@assigned-by-dhcp.cox.net","threadId":"3867","inReplyTo":"e3ksoq$is$1@sea.gmane.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-08T02:54:09Z","receivedAt":"2006-05-08T02:54:09Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> Wouldn't it be easier (sorry, no code yet) to have the following:\n>\n>         I WANT to have these\n>         I HAVE these\n>         These are GRAFT PARENTLESS        \n>\n> with the target side sending list of all parentless commits in the\n> ... The source side will then do the grafting 'in memory' and\n> send the packs like normal, only with those cauterizing grafts in place.\n\nI think that is essentially the outline of shallow clone\nproposal, except that you have to be careful and take not just\n\"parentless\" but other grafts (e.g. one that removes one parent\nfrom a merge commit to pretend that a side branch did not exist)\ninto account as well.  I do not remember if I already coded it\nor not -- I might have.\n"},{"id":"19703","messageId":"e3mfss$rnd$1@sea.gmane.org","threadId":"3867","inReplyTo":"7vy7xdgram.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-08T04:02:53Z","receivedAt":"2006-05-08T04:02:53Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n>> Wouldn't it be easier (sorry, no code yet) to have the following:\n>>\n>>         I WANT to have these\n>>         I HAVE these\n>>         These are GRAFT PARENTLESS\n>>\n>> with the target side sending list of all parentless commits in the\n>> ... The source side will then do the grafting 'in memory' and\n>> send the packs like normal, only with those cauterizing grafts in place.\n> \n> I think that is essentially the outline of shallow clone\n> proposal, except that you have to be careful and take not just\n> \"parentless\" but other grafts (e.g. one that removes one parent\n> from a merge commit to pretend that a side branch did not exist)\n> into account as well.  I do not remember if I already coded it\n> or not -- I might have.\n\nHaving grafts file being used for both joining history and cauterizing\nhistory makes re-cauterizing (e.g. changing depth of clone) difficult at\nbest...\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19704","messageId":"e3mh51$1eq$1@sea.gmane.org","threadId":"3867","inReplyTo":"7vy7xdgram.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-08T04:24:18Z","receivedAt":"2006-05-08T04:24:18Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n>> Wouldn't it be easier (sorry, no code yet) to have the following:\n>>\n>>         I WANT to have these\n>>         I HAVE these\n>>         These are GRAFT PARENTLESS\n>>\n>> with the target side sending list of all parentless commits in the\n>> ... The source side will then do the grafting 'in memory' and\n>> send the packs like normal, only with those cauterizing grafts in place.\n> \n> I think that is essentially the outline of shallow clone\n> proposal, except that you have to be careful and take not just\n> \"parentless\" but other grafts (e.g. one that removes one parent\n> from a merge commit to pretend that a side branch did not exist)\n> into account as well.  I do not remember if I already coded it\n> or not -- I might have.\n\nSo, let it be all grafts removing some or all parents from commit.\n\nAnd that proposal would work I think also for the fetch, not only clone.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19705","messageId":"20060508042429.GA20249@coredump.intra.peff.net","threadId":"3867","inReplyTo":"20060508003338.GB17138@thunk.org","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-05-08T04:24:29Z","receivedAt":"2006-05-08T04:24:29Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, May 07, 2006 at 08:27:02AM -0700, Linus Torvalds wrote:\n\n> factor for a lot of things for many \"common\" filesystem setups. You \n> probably didn't even account for the size of inodes in your \"du\" setup.\n\nMy numbers came from git-count-objects, which uses the st_blocks sum for\nall objects. The actual du numbers showing space wasted by block\nboundaries are:\n  du -c ??: 1429216\n  du -c --apparent-size ??: 792277\nSo it's about 45% wasted space.\n\nOn Sun, May 07, 2006 at 08:33:38PM -0400, Theodore Tso wrote:\n\n> If there are 233338 objects, then the average wasted space due to\n> internal fragmentation is 233338 * 2k, or 466676 kilobytes, or only\n> 36% of the wasted space.  Most of the savings is probably coming from\n> the compression and delta packing.\n\nAs Linus indicated, that assumes a uniform distribution of file sizes\n(and my numbers above show that it is, in fact, somewhat higher). FYI,\nthe mean and median of usage of the final 4K block in the linux-2.6\nrepository are 1309 and 912 bytes, respectively.\n\n-Peff\n"},{"id":"19716","messageId":"Pine.LNX.4.64.0605080813590.3718@g5.osdl.org","threadId":"3867","inReplyTo":"20060508042429.GA20249@coredump.intra.peff.net","subject":"Re: Unresolved issues #2 (shallow clone again)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-08T15:32:36Z","receivedAt":"2006-05-08T15:32:36Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 8 May 2006, Jeff King wrote:\n>\n> On Sun, May 07, 2006 at 08:27:02AM -0700, Linus Torvalds wrote:\n> \n> > factor for a lot of things for many \"common\" filesystem setups. You \n> > probably didn't even account for the size of inodes in your \"du\" setup.\n> \n> My numbers came from git-count-objects, which uses the st_blocks sum for\n> all objects. The actual du numbers showing space wasted by block\n> boundaries are:\n>   du -c ??: 1429216\n>   du -c --apparent-size ??: 792277\n> So it's about 45% wasted space.\n\nAnd that's actually ignoring inode sizes and directory sizes (well, it \ndoesn't \"ignore\" directory sizes - it counts them - but if you compare it \nto a straight packed format, it's still overhead).\n\nAnyway, looks like it's about 2:1, not 3:1 like I claimed, but the point \nbeing that blocking factors tend to be at least on the same order of \nmagnitude as just plain compression (which also tends to be in the 2:1 \narea for normal, fairly easily compressible, stuff).\n\nThe delta-packing obviously is much bigger for any project with real \nhistory. In traditional setups (where you always delta-pack within one \nthing, ie at the level of individual SCCS/RCS files), the delta packing \nobviously _also_ avoids blocking issues, since it means that a thousand \nrevisions of the same file will all share the same inode.\n\nSo because git uses a whole-file model, it obviously makes the blocking \nissues with its unpacked format _much_ higher than for any traditional \nmedium - no conglomeration of different versions of the file in the same \nfilesystem object. On the other hand, the packed format also tends to be \neven _more_ efficient than a traditional one, so the end result of it all \nis apparently a pretty big net win even in space consumption).\n\nSide note: I realize that some people think the packs are ugly and \nstrange. They aren't linear versions of a file, and instead appear as a \nfairly random \"jumble\". And they can't be incrementally re-packed, and you \nhave to generate a whole new pack-file (which can be incremental in \n_content_, of course). So people think they are ugly.\n\nI'd argue that they are beautiful. They are beautiful because they _don't_ \ncontain history in themselves (the objects they contain encode the history \nof course, but the pack-file itself does not).\n\nAnd they are beautiful because we can use the exact same format for \nstreaming data over the network as for the database itself (that, of \ncourse, was just about _the_ design consideration). Show me another system \nthat has exactly the same (not \"similar\", not \"same concepts\": _same_) \nnetwork protocol as it internal database.\n\nAnd they are beautiful exactly because their lack of any internal \nstructure allows you to pack things by criteria _you_ care about, ie the \nwhole \"sort things by recency\" thing, so that commonly accessed data can \nbe packed at the head of the pack-file - exactly because the pack-file \ndoesn't have any internal structure of its own that you need to worry \nabout and that constrains your sorting.\n\nThe same thing is what allows you to delta any blob against any other \nblob - without worrying about history or other random pack-file rules. You \ncan do packign purely by how well you want to pack, not by any secondary \nconstraints.\n\nAnd the \"no incremental updates\" may sound like a huge downside, but it's \nall the same basic git logic: objects and filesystem contents are \nimmutable, and that allows us to avoid a lot of locking overhead. Locking \nis _hard_. Locking is _inefficient_. And locking really really screws you \nwhen you miss it.\n\nSo I'll happily say that pack-files are strange, and that you have to get \na bit used to the notion that they should be repacked \"asynchronously\". \nBut it's really a matter of \"getting used to it\", because once you do, \nyou'll see that it's actually an absolutely huge deal, and you'll learn to \nlove the bomb^H^H^H^Hpack-file.\n\n\t\t\tLinus \"pack-files rule\" Torvalds\n"},{"id":"19780","messageId":"1147174809.2794.12.camel@pmac.infradead.org","threadId":"3867","inReplyTo":"7v4q065hq0.fsf@assigned-by-dhcp.cox.net","subject":"Re: Unresolved issues #2","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2006-05-09T11:40:09Z","receivedAt":"2006-05-09T11:40:09Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Thu, 2006-05-04 at 01:15 -0700, Junio C Hamano wrote:\n> \n> * Message-ID:\n> <4fb292fa0604290630r19edd7ejf88642e33b350d1d@mail.gmail.com>\n>   Content-type charset for send-email (Bertrand Jacquin)\n> \n>   The output from format-patch by default is unmarked, which\n>   means the commit message part is UTF-8 (by strong convention),\n>   and the contents of the diff is whatever the contents of the\n>   file is encoded in.\n\nEmail without a Content-Type: header is supposed to be ASCII. If it\ncontains 8-bit characters, it's invalid. It'll be interpreted by\ndifferent systems in different ways -- not necessarily as UTF-8. Some\nmay even just reject it, on grounds of RFC non-compliance.\n\n>   David Woodhouse did a patch to allow specifying charset on the\n>   command line (and default to UTF-8) which is a move in the\n>   right direction, but Bertrand's system seems to have trouble\n>   with it.\n\nI thought Bertrand then confirmed that he was having trouble _before_\napplying my patch, too? His response when I asked it it appears without\nmy patch was \"[it] appear without in 1.3.1 and I can't seed mail with\ntoo. Also, 1.2.4 work fine here (without patch).\"\n\n>   I think if we were to do this we probably need to teach\n>   format-patch to optionally do multi-part.  We may not\n>   necessarily want to mark the payload to be in the same\n>   encoding as the commit message (not that git-apply cares -- to\n>   it, the payload is just 8-bit unencoded text, but we would\n>   want to protect it from getting mangled by e-mail transport). \n\nI'm not sure about that. The payload is patches, isn't it? That's just\ntext, too -- we aren't going to deal with diffs of binary content very\nwell _anyway_, are we?\n\nObviously, there's nothing to stop people from storing binary blobs in\nGIT, but unless you want to start sending actual _blobs_ as attachments\ninstead of sending patches, I think there's no need to play with MIME\nmultipart stuff.\n\nI've no particular objection to it, but it's a separate issue to\nBertanrd's. That's a bug-fix, while multipart is an RFE without much\npoint, IMO.\n\n-- \ndwmw2\n"},{"id":"19781","messageId":"4fb292fa0605090453s390b58f3oc6c7607ea9f2f728@mail.gmail.com","threadId":"3867","inReplyTo":"1147174809.2794.12.camel@pmac.infradead.org","subject":"Re: Unresolved issues #2","fromName":"Bertrand Jacquin","fromEmail":"beber.mailing@gmail.com","sentAt":"2006-05-09T11:53:48Z","receivedAt":"2006-05-09T11:53:48Z","isPatch":false,"sender":{"key":"beber.mailing@gmail.com","avatar":null},"body":"On 5/9/06, David Woodhouse <dwmw2@infradead.org> wrote:\n> On Thu, 2006-05-04 at 01:15 -0700, Junio C Hamano wrote:\n> >\n> > * Message-ID:\n> > <4fb292fa0604290630r19edd7ejf88642e33b350d1d@mail.gmail.com>\n>\n> >   David Woodhouse did a patch to allow specifying charset on the\n> >   command line (and default to UTF-8) which is a move in the\n> >   right direction, but Bertrand's system seems to have trouble\n> >   with it.\n>\n> I thought Bertrand then confirmed that he was having trouble _before_\n> applying my patch, too? His response when I asked it it appears without\n> my patch was \"[it] appear without in 1.3.1 and I can't seed mail with\n> too. Also, 1.2.4 work fine here (without patch).\"\n\nOk, to make short :\ngit-send-email 1.2.4 :\nNo EOF error on my smtp server.\ngit-send-email 1.3.1 :\nEOF error on my smtp server.\n\nI upgraded to 1.3.1 when I received patch from you, don't test 1.3.1\nand then applied your patch, you test. And so test failed.\n\n--\nBeber\n#e.fr@freenode\n"},{"id":"19783","messageId":"Pine.LNX.4.64.0605090907260.24505@localhost.localdomain","threadId":"3867","inReplyTo":"1147174809.2794.12.camel@pmac.infradead.org","subject":"Re: Unresolved issues #2","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-05-09T13:09:13Z","receivedAt":"2006-05-09T13:09:13Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 9 May 2006, David Woodhouse wrote:\n\n> I'm not sure about that. The payload is patches, isn't it? That's just\n> text, too -- we aren't going to deal with diffs of binary content very\n> well _anyway_, are we?\n\nYes we do.  GIT now has its own email friendly binary patch format.\n\n\nNicolas\n"}]}