{"thread":{"id":"7619","subject":"Weird shallow-tree conversion state, and branches of shallow trees","startedAt":"2007-04-12T00:53:36Z","lastAt":"2007-04-17T21:51:18Z","messageCount":34,"participants":["Robin H. Johnson","Johannes Schindelin","David Lang","Shawn O. Pearce","Nguyen Thai Ngoc Duy","Jakub Narebski","Linus Torvalds","Andy Parkins","Bill Lear","Theodore Tso","Julian Phillips","Sven Verdoolaege","Junio C Hamano","Daniel Barkalow"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"39176","messageId":"20070412005336.GA18378@curie-int.orbis-terrarum.net","threadId":"7619","inReplyTo":null,"subject":"Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2007-04-12T00:53:36Z","receivedAt":"2007-04-12T00:53:36Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"I was doing some random tests with shallow trees, and ran into two\nissues - the first is a shallow tree that doesn't extend anymore when it\nshould, and the second is some branched shallow tree trouble.\n\n1.\n(I was using the kernel.org Git repo for my testing here)\n> git clone --depth 1 git://GIT-REMOTE-URL\n> # do some local commits\n> git pull --depth 1000000 # some very large number, to try and add all the history\n\nAt this point, I noticed that my tree still seemed to be shallow, and no\nmatter what I tried, I couldn't un-shallow it.\n\n.git/shallow contained a single line:\n> 9c405082d96ed7a7ed830f9861dbad9a32e4d268\n\nAnd moving the shallow file out the way, fsck --full gets me:\n> broken link from  commit 9c405082d96ed7a7ed830f9861dbad9a32e4d268\n>               to  commit bb3e781d7f6259eb414cbecd8bad74cd4a188b41\n> broken link from  commit 9c405082d96ed7a7ed830f9861dbad9a32e4d268\n>               to  commit 9bfbe261923f4e9d89f65e6755fa6501aa6531b0\n> missing commit bb3e781d7f6259eb414cbecd8bad74cd4a188b41\n> missing commit 9bfbe261923f4e9d89f65e6755fa6501aa6531b0\n\nAny ideas on why it's not going to full depth?\nI don't have a reliable test case for this yet, sometimes it does go deep\nproperly, sometimes it doesn't.\n\n2.\nAgain about shallow repos, a development problem I ran into.\n> git clone --depth 1 git://GIT-REMOTE-URL\n> git checkout -b working-branch\n> # do various work, and git-commit the changes\n> git checkout master \n> git pull\n> # some time goes by, and you want the latest upstream changes\n> git checkout working-branch\n> git pull . master\n\nThe last pull from the local master fails. This seems weird, because if\nworking-branch development is done on the master instead, the earlier pull\nnever complains. So in this case, the working-branch should be able to pull\nfrom the local master branch fine.\n\nThis bug basically stops people from being able to take a shallow clone of a\nrepository with a lot of history, and have multiple working branches on it.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Council Member\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"39329","messageId":"Pine.LNX.4.64.0704141019290.18655@racer.site","threadId":"7619","inReplyTo":"20070412005336.GA18378@curie-int.orbis-terrarum.net","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-04-14T08:56:10Z","receivedAt":"2007-04-14T08:56:10Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 11 Apr 2007, Robin H. Johnson wrote:\n\n> I was doing some random tests with shallow trees, and ran into two \n> issues - the first is a shallow tree that doesn't extend anymore when it \n> should, and the second is some branched shallow tree trouble.\n\nAh! Seems we finally have a user for shallow clones! ;-)\n\nSeriously again: I am at fault for putting the shallow support into Git, \nfailing to provide sensible test cases. This was partly due to my \nlaziness, and partly due to the overwhelming lack of demand.\n\nI am in the middle of moving (haven't reached my destination yet), so I \nwill take a couple more days until I can look into your problems. If you \nfind out in the meantime what is happening, please share the information \nwith us.\n\nCiao,\nDscho\n"},{"id":"39353","messageId":"Pine.LNX.4.63.0704141655390.31807@qynat.qvtvafvgr.pbz","threadId":"7619","inReplyTo":"20070415000330.GG3778@curie-int.orbis-terrarum.net","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-04-15T00:02:47Z","receivedAt":"2007-04-15T00:02:47Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Sat, 14 Apr 2007, Robin H. Johnson wrote:\n\n> On Sat, Apr 14, 2007 at 10:56:10AM +0200, Johannes Schindelin wrote:\n>> Ah! Seems we finally have a user for shallow clones! ;-)\n> Heh. I'm specifically looking at git, trying to resolve the deficiencies\n> that were identified during by one of our (Gentoo) SoC2006 projects, on\n> the potential migration of the Gentoo CVS. Git has matured tremendously\n> since then.\n>\n> The primary Gentoo CVS module (gentoo-x86), has 234672 files tracked,\n> and 1309603 CVS revisions. Between 350k and 500k changesets, depending\n> on how you merge those revisions.\n>\n> Couple of the things that were identified either in the SoC project, or\n> since then.\n> - Shallow history checkouts are important to our low-bandwidth\n>  ebuild-tree developers (people in places with 33.6k modems, because\n>  the phone lines don't work well enough for 56k), or other high latency\n>  setups.\n\nnote that for people on low-bandwideth lines, makeing too shallow a checkout can \nactually end up costing more over time (they will have to pull full revisions \nsince they don't have the earlier versions to just pull a diff against)\n\n> - Shallow tree (subtree) checkouts, for the developers that focus on\n>  specific portions of large modules and have no interest in the rest of\n>  the that tree. Eg. Releng does their work in gentoo/src/releng.\n\nthis could either be shallow tree or subproject, depending on how you end up \norginizing things.\n\n> - ACLs specific to subtree commits. Something similar to the cvs_acls.pl\n>  that FreeBSD uses would be great. Eg gentoo-x86/sec-policy/ is\n>  restricted to members of the security team (SELinux policies).\n\nsince git isn't designed with a single repository, it also doesn't need to worry \nabout acl's (in fact, i don't think it has the concept of permissions at all). \nthis is up to the people maintaining the 'master' repository to pull from the \nright people\n\n> - CVS Keyword-like behavior, to specifically place the path and revision\n>  of certain files into the file directly, for ease of tracking when the\n>  file is removed from it's original surrounding. I know this one is\n>  going to draw some flack, but it's a very common practice for a user\n>  to copy a file out of the CVS tree, make some modifications, and then\n>  post the entire changed version up, esp. when the size of the changes\n>  exceeds the size of diff.\n\nI'm not understanding why you need this. git tracks the file content, not the \ndiffs betwen files. a developer does their work and git figures out when you do \na pull if it's better to send the file or a diff (and if you are sending a diff, \nwhat you are doing the diff against, it may not be the file that had that name \nbefore)\n\nthere's no need to place the path and revision in the file itself.\n\nDavid Lang\n"},{"id":"39351","messageId":"20070415000330.GG3778@curie-int.orbis-terrarum.net","threadId":"7619","inReplyTo":"Pine.LNX.4.64.0704141019290.18655@racer.site","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2007-04-15T00:03:30Z","receivedAt":"2007-04-15T00:03:30Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Sat, Apr 14, 2007 at 10:56:10AM +0200, Johannes Schindelin wrote:\n> Ah! Seems we finally have a user for shallow clones! ;-)\nHeh. I'm specifically looking at git, trying to resolve the deficiencies\nthat were identified during by one of our (Gentoo) SoC2006 projects, on\nthe potential migration of the Gentoo CVS. Git has matured tremendously\nsince then.\n\nThe primary Gentoo CVS module (gentoo-x86), has 234672 files tracked,\nand 1309603 CVS revisions. Between 350k and 500k changesets, depending\non how you merge those revisions.\n\nCouple of the things that were identified either in the SoC project, or\nsince then.\n- Shallow history checkouts are important to our low-bandwidth\n  ebuild-tree developers (people in places with 33.6k modems, because\n  the phone lines don't work well enough for 56k), or other high latency\n  setups.\n- Shallow tree (subtree) checkouts, for the developers that focus on\n  specific portions of large modules and have no interest in the rest of\n  the that tree. Eg. Releng does their work in gentoo/src/releng.\n- ACLs specific to subtree commits. Something similar to the cvs_acls.pl\n  that FreeBSD uses would be great. Eg gentoo-x86/sec-policy/ is\n  restricted to members of the security team (SELinux policies).\n- CVS Keyword-like behavior, to specifically place the path and revision\n  of certain files into the file directly, for ease of tracking when the\n  file is removed from it's original surrounding. I know this one is\n  going to draw some flack, but it's a very common practice for a user\n  to copy a file out of the CVS tree, make some modifications, and then\n  post the entire changed version up, esp. when the size of the changes\n  exceeds the size of diff.\n\n> Seriously again: I am at fault for putting the shallow support into Git, \n> failing to provide sensible test cases. This was partly due to my \n> laziness, and partly due to the overwhelming lack of demand.\nI still haven't figured out a decent testcase for this, I need to dig\nharder.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Council Member\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"39360","messageId":"20070415020139.GB2689@curie-int.orbis-terrarum.net","threadId":"7619","inReplyTo":"Pine.LNX.4.63.0704141655390.31807@qynat.qvtvafvgr.pbz","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2007-04-15T02:01:39Z","receivedAt":"2007-04-15T02:01:39Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Sat, Apr 14, 2007 at 05:02:47PM -0700, David Lang wrote:\n> > - Shallow history checkouts are important to our low-bandwidth\n> >  ebuild-tree developers (people in places with 33.6k modems, because\n> >  the phone lines don't work well enough for 56k), or other high latency\n> >  setups.\n>  note that for people on low-bandwideth lines, makeing too shallow a checkout \n>  can actually end up costing more over time (they will have to pull full \n>  revisions since they don't have the earlier versions to just pull a diff \n>  against)\nYes, I'm aware that it may be more efficient over the long term for them\nto pull given blocks, and I'm going to recommend that developers have a\nfull history anyway, but I suspect that they will still make heavy use\nof shallow trees, esp. as some do throwaway trees often.\n(This one is a moot point anyway, the shallow history support in Git is\npretty much done baring the bugs I posted about previously).\n\n> > - Shallow tree (subtree) checkouts, for the developers that focus on\n> >  specific portions of large modules and have no interest in the rest of\n> >  the that tree. Eg. Releng does their work in gentoo/src/releng.\n>  this could either be shallow tree or subproject, depending on how you end up \n>  orginizing things.\nshallow tree, because we really do have people that check out arbitrary\nsub-divisions (the web translation teams come to mind, they just have\ncheckouts of English and their own language), and going sub-project\nwould be insane for that.\n\n> > - ACLs specific to subtree commits. Something similar to the cvs_acls.pl\n> >  that FreeBSD uses would be great. Eg gentoo-x86/sec-policy/ is\n> >  restricted to members of the security team (SELinux policies).\n>  since git isn't designed with a single repository, it also doesn't need to \n>  worry about acl's (in fact, i don't think it has the concept of permissions \n>  at all). this is up to the people maintaining the 'master' repository to \n>  pull from the right people\nI should have mentioned that we aren't following the kernel model here.\nAll of the developers will have git+ssh access to the central tree, to\npush their own changes to it. On a similar tangent, in some subtrees\n(our documentation mainly) we have server-side validation tests before\nthe commit is accepted. The 'update' hook documentation suggests that\nACLs should be possible and implemented via that.\n\n> > - CVS Keyword-like behavior, to specifically place the path and revision\n> >  of certain files into the file directly, for ease of tracking when the\n> >  file is removed from it's original surrounding. I know this one is\n> >  going to draw some flack, but it's a very common practice for a user\n> >  to copy a file out of the CVS tree, make some modifications, and then\n> >  post the entire changed version up, esp. when the size of the changes\n> >  exceeds the size of diff.\n>  I'm not understanding why you need this. git tracks the file content, not \n>  the diffs betwen files. a developer does their work and git figures out when \n>  you do a pull if it's better to send the file or a diff (and if you are \n>  sending a diff, what you are doing the diff against, it may not be the file \n>  that had that name before)\nThe tree that goes out to users is NOT git or CVS. What you point to\nhere is impossible unless we forced all of the users to migrate to git\n(a truly herculean task if there was ever one).\nIt's a tarball or an rsync of an automatically managed CVS checkout.\n(Tarballs go onto the release media, and are also widely used by those\nthat sneaker-net their trees to machines for security reasons).\nAlternatively, the users browse the viewcvs, and pull something from the\nAttic. Regardless of where they get the file from, the problem is that\nthe file doesn't contain any markers to help the developers merge it\nback again.\n\nA frequent occurrence of this is where the user takes rev X of a file\n(because it was the latest one at the time), makes a local (non\nversion-controlled) copy, and submits it back our Bugzilla some months\ndown the line. Thanks to the $Header$ in the file he submits, we can\nproduce a diff against the original revision, and figure out how best to\nmerge it with the latest revision.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Council Member\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"39364","messageId":"20070415043146.GB2229@spearce.org","threadId":"7619","inReplyTo":"20070415020139.GB2689@curie-int.orbis-terrarum.net","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-04-15T04:31:46Z","receivedAt":"2007-04-15T04:31:46Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Robin H. Johnson\" <robbat2@gentoo.org> wrote:\n> On Sat, Apr 14, 2007 at 05:02:47PM -0700, David Lang wrote:\n> > > - Shallow history checkouts are important to our low-bandwidth\n> > >  ebuild-tree developers (people in places with 33.6k modems, because\n> > >  the phone lines don't work well enough for 56k), or other high latency\n> > >  setups.\n> >  note that for people on low-bandwideth lines, makeing too shallow a checkout \n> >  can actually end up costing more over time (they will have to pull full \n> >  revisions since they don't have the earlier versions to just pull a diff \n> >  against)\n\nMail them a DVD of the Git import, have them load it locally,\nand use --reference for all future clones.  With Git its possible\nto build fast throwaway trees from any random URL, so long as you\nkeep at least one repository available locally to act as a reference.\n\nThe speed at which a DVD (or small box of CDs) travels through the\nvarious postal systems might very well be faster than 33.6k modem.\n:-)\n \n> I should have mentioned that we aren't following the kernel model here.\n> All of the developers will have git+ssh access to the central tree, to\n> push their own changes to it. On a similar tangent, in some subtrees\n> (our documentation mainly) we have server-side validation tests before\n> the commit is accepted. The 'update' hook documentation suggests that\n> ACLs should be possible and implemented via that.\n\nYes.  I run probably the most paranoid update hook in existance.\nIf you want a copy let me know, I'll send it to you.  Its a Perl\nscript that verifies the 'committer ' line matches the UNIX uid (by\ndoing a table lookup) for every new commit or tag being introduced\nto the repository.  It also verifies that the user can update that\nbranch, create it, delete it, or rewind it.\n\nIt sounds like you would need to add some additional rules about\nspecific paths being modified only by certain people in certain\nbranches (for the SELinux stuff), and running other validations in\nthe documentation (whatever that is).\n\n> The tree that goes out to users is NOT git or CVS. What you point to\n> here is impossible unless we forced all of the users to migrate to git\n> (a truly herculean task if there was ever one).\n> It's a tarball or an rsync of an automatically managed CVS checkout.\n> (Tarballs go onto the release media, and are also widely used by those\n> that sneaker-net their trees to machines for security reasons).\n> Alternatively, the users browse the viewcvs, and pull something from the\n> Attic. Regardless of where they get the file from, the problem is that\n> the file doesn't contain any markers to help the developers merge it\n> back again.\n\nGit won't do this for you.  We specifically don't mangle source[*1*].\n\nWhat you could do is create a program that mangles the files before\ndelivery.  You would probably want to do something like:\n\n  $Id: 7fbf239:path/to/file$\n\nwhere 7fbf239 is the earliest commit that introduced that particular\nversion of path/to/file, even if that is months old.  That would\nbe most like what CVS would do.  8 char abbreviated commits should\nbe reasonably stable, and not too long to read or copy and paste.\nA format like the above would also be easy to grab and copy into\na Git command line.\n\nIf we had a Git library that could access the repository, this would\na pretty easy program to write.  You are basically blaming each path\nin the current HEAD commit on the parent, until you cannot blame\nanyone else for that path.  You do this blame on the entire tree,\nand then output the munged structure (or only the files you want\nmunged).\n\nIts good we have a GSoC project working on libification!  ;-)\n\n[*1*] Yes, I'm ignoring the nutso crlf support that's now in...  Even\n      though I work on Windows, the only true line ending is LF.  ;-)\n\n-- \nShawn.\n"},{"id":"39366","messageId":"fcaeb9bf0704142257x3761ef2cie3996420b3bcd24a@mail.gmail.com","threadId":"7619","inReplyTo":"20070415043146.GB2229@spearce.org","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2007-04-15T05:57:13Z","receivedAt":"2007-04-15T05:57:13Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On 4/15/07, Shawn O. Pearce <spearce@spearce.org> wrote:\n> > The tree that goes out to users is NOT git or CVS. What you point to\n> > here is impossible unless we forced all of the users to migrate to git\n> > (a truly herculean task if there was ever one).\n> > It's a tarball or an rsync of an automatically managed CVS checkout.\n> > (Tarballs go onto the release media, and are also widely used by those\n> > that sneaker-net their trees to machines for security reasons).\n> > Alternatively, the users browse the viewcvs, and pull something from the\n> > Attic. Regardless of where they get the file from, the problem is that\n> > the file doesn't contain any markers to help the developers merge it\n> > back again.\n>\n> Git won't do this for you.  We specifically don't mangle source[*1*].\n>\n> What you could do is create a program that mangles the files before\n> delivery.  You would probably want to do something like:\n>\n>   $Id: 7fbf239:path/to/file$\n>\n> where 7fbf239 is the earliest commit that introduced that particular\n> version of path/to/file, even if that is months old.  That would\n> be most like what CVS would do.  8 char abbreviated commits should\n> be reasonably stable, and not too long to read or copy and paste.\n> A format like the above would also be easy to grab and copy into\n> a Git command line.\n>\n> If we had a Git library that could access the repository, this would\n> a pretty easy program to write.  You are basically blaming each path\n> in the current HEAD commit on the parent, until you cannot blame\n> anyone else for that path.  You do this blame on the entire tree,\n> and then output the munged structure (or only the files you want\n> munged).\n>\n> Its good we have a GSoC project working on libification!  ;-)\n>\n> [*1*] Yes, I'm ignoring the nutso crlf support that's now in...  Even\n>       though I work on Windows, the only true line ending is LF.  ;-)\n\nCan we add an attribute like Subversion's svn:keywords? If the\nattribute is set, we expand keywords when checkout and remove\nexpansion in memory before doing any git operations. It's some kind of\nI/O filter for working directory access.\n\n-- \nDuy\n"},{"id":"39368","messageId":"evsp0k$qoo$1@sea.gmane.org","threadId":"7619","inReplyTo":"fcaeb9bf0704142257x3761ef2cie3996420b3bcd24a@mail.gmail.com","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-04-15T08:54:11Z","receivedAt":"2007-04-15T08:54:11Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Nguyen Thai Ngoc Duy wrote:\n\n> On 4/15/07, Shawn O. Pearce <spearce@spearce.org> wrote:\n>> > The tree that goes out to users is NOT git or CVS. What you point to\n>> > here is impossible unless we forced all of the users to migrate to git\n>> > (a truly herculean task if there was ever one).\n>> > It's a tarball or an rsync of an automatically managed CVS checkout.\n>> > (Tarballs go onto the release media, and are also widely used by those\n>> > that sneaker-net their trees to machines for security reasons).\n>> > Alternatively, the users browse the viewcvs, and pull something from\nthe\n>> > Attic. Regardless of where they get the file from, the problem is that\n>> > the file doesn't contain any markers to help the developers merge it\n>> > back again.\n>>\n>> Git won't do this for you.  We specifically don't mangle source[*1*].\n>>\n>> What you could do is create a program that mangles the files before\n>> delivery.  You would probably want to do something like:\n>>\n>>   $Id: 7fbf239:path/to/file$\n>>\n>> where 7fbf239 is the earliest commit that introduced that particular\n>> version of path/to/file, even if that is months old.  That would\n>> be most like what CVS would do.  8 char abbreviated commits should\n>> be reasonably stable, and not too long to read or copy and paste.\n>> A format like the above would also be easy to grab and copy into\n>> a Git command line.\n>>\n>> If we had a Git library that could access the repository, this would\n>> a pretty easy program to write.  You are basically blaming each path\n>> in the current HEAD commit on the parent, until you cannot blame\n>> anyone else for that path.  You do this blame on the entire tree,\n>> and then output the munged structure (or only the files you want\n>> munged).\n>>\n>> Its good we have a GSoC project working on libification!  ;-)\n>>\n>> [*1*] Yes, I'm ignoring the nutso crlf support that's now in...  Even\n>>       though I work on Windows, the only true line ending is LF.  ;-)\n> \n> Can we add an attribute like Subversion's svn:keywords? If the\n> attribute is set, we expand keywords when checkout and remove\n> expansion in memory before doing any git operations. It's some kind of\n> I/O filter for working directory access.\n\nThere was some talk about keyword expansion, and it is doable IIRC.\nCheck out threads containing:\n  Message-ID: <20070301175200.GA21433@informatik.uni-freiburg.de>\n  http://permalink.gmane.org/gmane.comp.version-control.git/41108\n(with some inane totally irrelevant subject)  \n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"39370","messageId":"20070415094438.GD2689@curie-int.orbis-terrarum.net","threadId":"7619","inReplyTo":"20070415043146.GB2229@spearce.org","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2007-04-15T09:44:38Z","receivedAt":"2007-04-15T09:44:38Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Sun, Apr 15, 2007 at 12:31:46AM -0400, Shawn O. Pearce wrote:\n> Mail them a DVD of the Git import, have them load it locally,\n> and use --reference for all future clones.  With Git its possible\n> to build fast throwaway trees from any random URL, so long as you\n> keep at least one repository available locally to act as a reference.\nOk, that makes it even more worthwhile for them to keep one tree\nlocally, I didn't think of that :-).\n\n> > the commit is accepted. The 'update' hook documentation suggests that\n> > ACLs should be possible and implemented via that.\n> Yes.  I run probably the most paranoid update hook in existance.\n> If you want a copy let me know, I'll send it to you.  Its a Perl\n> script that verifies the 'committer ' line matches the UNIX uid (by\n> doing a table lookup) for every new commit or tag being introduced\n> to the repository.  It also verifies that the user can update that\n> branch, create it, delete it, or rewind it.\n> \n> It sounds like you would need to add some additional rules about\n> specific paths being modified only by certain people in certain\n> branches (for the SELinux stuff), and running other validations in\n> the documentation (whatever that is).\nYes please, it would be greatly appreciated. I'll hack path ACLs into\nit, and feed it back to contrib/? (CVS and SVN ship ACL stuff in their\ncontrib/, so we could probably follow suite safely).\n\n> What you could do is create a program that mangles the files before\n> delivery.  You would probably want to do something like:\n> \n>   $Id: 7fbf239:path/to/file$\nThere's one core problem with mangled after the fact there:\nIt's going to break checksum/gpg verification later.\nHere's the existing CVS process as a comparison.\n1. Developer creates/changes foo-1.2.ebuild. (cvs add, but not cvs ci).\n2. Runs the local verify+commit tool (repoman).\n(these steps are done by repoman now)\n3. Generates the initial Manifest (contains SHA256/MD5/RIPEMD160 etc.).\n4. Commits the initial Manifest AND the files from the developer.\n5. Gegenerated Manifest because of any keywords in the files.\n6. Manifest is clear-signed with gpg.\n7. Signed Manifest is committed.\n\nWe can't require the re-processing of the files before they can be\nverified, as that removes the ability for users to easily verify them\nwith standard tools (md5sum,sha256sum).\n\nThe direct conversion of such a process to insert the $Id$ and then\nre-commit that $Id$ runs into chicken-and-egg problems as well, so\neither git needs to insert the keyword, or the file can't be changed.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Council Member\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"39378","messageId":"Pine.LNX.4.64.0704151115270.5473@woody.linux-foundation.org","threadId":"7619","inReplyTo":"fcaeb9bf0704142257x3761ef2cie3996420b3bcd24a@mail.gmail.com","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-15T18:18:28Z","receivedAt":"2007-04-15T18:18:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 15 Apr 2007, Nguyen Thai Ngoc Duy wrote:\n> \n> Can we add an attribute like Subversion's svn:keywords? If the\n> attribute is set, we expand keywords when checkout and remove\n> expansion in memory before doing any git operations. It's some kind of\n> I/O filter for working directory access.\n\nNNOOo-oooo...\n\nKeyword substitution is just *stupid*. It's an inexcusable braindamage. \nDon't do it. It leads to all kinds of idiotic problems downstream, and it \nreally doesn't help *anything* except for \"but I'm used to it\". There are \nabsolutely no valid uses for it.\n\nIf you want to tag your files somehow, do it in \"git archive\" when \nexporting it, but not in the working tree. And realize that once you \nexport it with the stupid keyword expansion, diffs etc will all be \ncorrupted, and will not - AND MUST NOT - apply to the uncorrupted working \ntree.\n\n\t\t\tLinus\n"},{"id":"39385","messageId":"200704152051.35639.andyparkins@gmail.com","threadId":"7619","inReplyTo":"Pine.LNX.4.64.0704151115270.5473@woody.linux-foundation.org","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-04-15T19:51:31Z","receivedAt":"2007-04-15T19:51:31Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Sunday 2007, April 15, Linus Torvalds wrote:\n\n> Keyword substitution is just *stupid*. It's an inexcusable\n> braindamage. Don't do it. It leads to all kinds of idiotic problems\n> downstream, and it really doesn't help *anything* except for \"but I'm\n> used to it\". There are absolutely no valid uses for it.\n\nYou're right that it can cause problems, but it is certainly not the \ncase that there are no valid uses for it.  I've mentioned it before but \nI'll say it again, because it is the only feature I miss from \nsubversion and I can't see why it is invalid.\n\nI keep diagrams for a project in SVG format in the repository, this \nworks very well because SVG is so nicely ASCII.  In the title block of \nthe diagram I put \"$Id$\", then in subversion, after checking in and \nupdating it got expanded to\n\n $Id: diagram.svg 148 2002-07-28 21:30:43Z andyp $\n\nNow, I print out that diagram and pin it to my wall - sometimes copies \nof it are given to others.  I do this on a regular basis.  The diagram \nis big and complicated and all versions of it look very similar.  In \nshort it is very convenient to have the version of the file actually \nprinted on the piece of paper.  This is a piece of paper remember, \nthere is no way to hash the daigram, or even look at the underlying \nsource.  When someone comes to me with a random version of the diagram, \nI can use that ID to checkout exactly the revision that that diagram \nrefers to.\n\nPlease explain to me why that is not a valid use.\n\n> If you want to tag your files somehow, do it in \"git archive\" when\n> exporting it, but not in the working tree. And realize that once you\n> export it with the stupid keyword expansion, diffs etc will all be\n> corrupted, and will not - AND MUST NOT - apply to the uncorrupted\n> working tree.\n\nAll of the problems you describe apply equally to CRLF conversion, and \nyet there seems to be no problem with implementing that.  In fact the \nproblem there is significantly worse, as it changes every line of the \nfile.\n\nNow, solving the keyword problem is not simple, obviously, but it's \ncertainly not impossible.  On git-add the expanded tags get unexpanded \nso $Tag: blah blah blah$ becomes $Tag$; on checkout they get expanded. \nSimilarly while calculating diffs - the diff engine unexpands as it \ngoes so the lines with the keywords in them are not seen as different \nregardless of the expanded part.\n\nApplying diffs from some external source doesn't corrupt anything - \nbecause the diff engine is, by definition, going to unexpand the \nkeywords when it compares.\n\nSo, someone sends you a diff that has this:\n\n- /* $Id: diagram.svg 148 2002-07-28 21:30:43Z andyp $ */\n+ /* $Id: diagram.svg 149 2002-07-29 20:32:47Z andyp $ */\n\nAnd you apply it to the working tree - well, that line will be seen as \nthis by the diff engine:\n\n- /* $Id$ */\n+ /* $Id$ */\n\nNo change.  Obviously this is entirely optional and would be activated \non a per-file basis.  For git it would be even more useful because of \nall the information actually available.  I'd love to have git-keywords \nlike these:\n\n $Commit: 2bfe3cec92be4f5e3bfc0e71ed560df4a726c07b$\n $Object: b1bd9e46c2bd64e00b671ff5ed512d9c12b53309$\n $Describe: v1.5.1.1-83-g2bfe3ce$\n $Id: cache.h v1.5.1.1-83-g2bfe3ce $\n\nFeelings seem very strong about this; I've seen comments again and again \nabout how braindamaged it is and I just can't see it - please, help me \nsee - what is it that is so utterly broken about it?  I can see that it \nadds a complication to many parts, but I can't see why it is seen as so \nevil.\n\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"39402","messageId":"Pine.LNX.4.64.0704151317180.5473@woody.linux-foundation.org","threadId":"7619","inReplyTo":"200704152051.35639.andyparkins@gmail.com","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-15T20:51:42Z","receivedAt":"2007-04-15T20:51:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 15 Apr 2007, Andy Parkins wrote:\n> \n> You're right that it can cause problems, but it is certainly not the \n> case that there are no valid uses for it.\n\nI'm sorry, but you're just wrong.\n\nThere are no valid uses for it in the working tree. Full stop.\n\nThere are valid uses to tag sources with some revision information WHEN IT \nLEAVES THE REVISION CONTROLLED ENVIRONMENT, but not one second before \nthat.\n\n> I keep diagrams for a project in SVG format in the repository, this \n> works very well because SVG is so nicely ASCII.  In the title block of \n> the diagram I put \"$Id$\", then in subversion, after checking in and \n> updating it got expanded to\n> \n>  $Id: diagram.svg 148 2002-07-28 21:30:43Z andyp $\n> \n> Now, I print out that diagram and pin it to my wall - sometimes copies \n> of it are given to others.  I do this on a regular basis.\n\nAnd is there *any* reason why you don't just do that as an \"export\" \noption, when it's very clear that people won't send diffs that include it \nand that will cause all the endless problems that keyword expansion \ncauses?\n\nWhy would you ever have the pain and suffering of using it within the \nsource control issue? Especially since you would be a *lot* better off \nusing just an export script that can do a lot better than CVS/SVN keyword \nexpansion could ever do (ie you can add all sorts of more relevant \ninformation than just a date and user name!)\n\n> Please explain to me why that is not a valid use.\n\nIt's not a valid use because there are many SO MUCH BETTER WAYS to get the \nsame thing, that have none of the downsides of keyword expansion?\n\nYour argument is akin to saying that \"Why isn't it a valid use to replace \nthe steering wheel in my car with a mouth-operated joystick under the \npassenger side seat?\"\n\nSure, you *can* steer a car by mouthing at it while having your head under \nthe passenger side seat, and your butt sticking out through the moonroof \n(\"We could add a periscope so that I can see where I'm going!\")\n\nBut that's not an argument *for* doing it, when there are ways that are \nobviously much better, and don't _need_ the periscope!\n\nSee?\n\nThe fact that you *can* do something is not a valid argument for it being \na valid use. You *can* do stupid things, but if you can get to the same \nend result by not doing stupid things, wouldn't you prefer that instead?\n\nHere's a small makefile snippet for you:\n\n\n\t%.prt: %.svg\n\t\tsed 's/\\$$Id\\$$/\\$$ $(shell git log --pretty=format:\"%h: %s (%an)\" --abbrev-commit -1 file.svg) \\$$/g' < $< > $@\n\nwhich would need some work (it doesn't quote things right - in reality \nyou'd write a simple script to do this properly).\n\nSee? No need for a periscope, and your butt can be toasty warm too if you \njust add a seat heater option...\n\n\t\tLinus\n"},{"id":"39434","messageId":"17954.48933.484379.593657@lisa.zopyra.com","threadId":"7619","inReplyTo":"Pine.LNX.4.64.0704151317180.5473@woody.linux-foundation.org","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Bill Lear","fromEmail":"rael@zopyra.com","sentAt":"2007-04-16T00:11:17Z","receivedAt":"2007-04-16T00:11:17Z","isPatch":false,"sender":{"key":"rael@zopyra.com","avatar":"https://gravatar.com/avatar/c4f2d2790ca3828d3b4e7dfebabf61d2fe94fd82fa49cdac2a5295dd2d46a874?d=mp&s=160"},"body":"On Sunday, April 15, 2007 at 13:51:42 (-0700) Linus Torvalds writes:\n>On Sun, 15 Apr 2007, Andy Parkins wrote:\n>> \n>> You're right that it can cause problems, but it is certainly not the \n>> case that there are no valid uses for it.\n>\n>I'm sorry, but you're just wrong.\n>\n>There are no valid uses for it in the working tree. Full stop.\n>\n>There are valid uses to tag sources with some revision information WHEN IT \n>LEAVES THE REVISION CONTROLLED ENVIRONMENT, but not one second before \n>that. ...\n\nNot that Linus needs any back-up from me, but I second this, very\nstrongly.  Decorating source code with release information is a proper\nfunction of release management tools, not the SCM system.  We had a\nsimilar argument in our company about this, sparked by a criticism of\ngit for not having keyword (version number) substitution, and I argued\nthat having such substitution functions in the SCM was out-of-place\nand a crutch for weak release procedures.  It's easy with a proper\nmake system to put whatever information you want from the SCM into the\nrelease product.\n\nThis would probably be as crazy as asking for saving and restoring\ntimestamps in the working tree on checkout of branches, and we know\nhow insane that is...\n\n\nBill\n"},{"id":"39440","messageId":"20070416021729.GH2689@curie-int.orbis-terrarum.net","threadId":"7619","inReplyTo":"Pine.LNX.4.64.0704151317180.5473@woody.linux-foundation.org","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2007-04-16T02:17:29Z","receivedAt":"2007-04-16T02:17:29Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Sun, Apr 15, 2007 at 01:51:42PM -0700, Linus Torvalds wrote:\n> There are valid uses to tag sources with some revision information WHEN IT \n> LEAVES THE REVISION CONTROLLED ENVIRONMENT, but not one second before \n> that.\nNobody has addressed the single problem that I have with adding it when\nit's leaving the environment, and that's still of paramount concern to\nme. Simply put, there is a conflict between being able to add revision\ninformation of stuff leaving the environment, and those additions\nbreaking previous checksums (which may be digitally signed, and thus\nbreaking the signatures).\n\nI'll reduce it further from my previous example.\n\n1. Developer commits some change to file A.\n2. The checksum file is updated because A changed (the checksum file\n   explicitly does not contain keywords).\n3. Developer signs the checksum file, and commits it.\n\nIf during the export process (which is undertaken elsewhere, by a\ndifferent person or script), file A now has an expansion applied to it,\nyou break the checksum file, which you CANNOT redo, because you lose the\ndeveloper's digital signature on the checksum file!\n\nUsing the existing git-verify-tag mechanisms are not suitable, because\nit is the exported information that must be verifiable.\n\nThere's FOUR possible solutions here:\n1. The commit to file A does the keywords - Which Linus is against.\n2. An ADDITIONAL commit to file A, after the initial commit, as a\n   scripted addition of the keywords, but before the checksum is\n   updated. I think this is messy myself, as you'd have to insert the\n   data from the N-1 commit always.\n3. Lose the ability to tag the files leaving the environment.\n4. Stop digitally signing the checksum file (which then leaves the\n   possibility for other attacks).\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Council Member\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"39441","messageId":"20070416030103.GB27533@thunk.org","threadId":"7619","inReplyTo":"20070416021729.GH2689@curie-int.orbis-terrarum.net","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-04-16T03:01:03Z","receivedAt":"2007-04-16T03:01:03Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Sun, Apr 15, 2007 at 07:17:29PM -0700, Robin H. Johnson wrote:\n> Nobody has addressed the single problem that I have with adding it when\n> it's leaving the environment, and that's still of paramount concern to\n> me. Simply put, there is a conflict between being able to add revision\n> information of stuff leaving the environment, and those additions\n> breaking previous checksums (which may be digitally signed, and thus\n> breaking the signatures).\n> \n> I'll reduce it further from my previous example.\n> \n> 1. Developer commits some change to file A.\n> 2. The checksum file is updated because A changed (the checksum file\n>    explicitly does not contain keywords).\n> 3. Developer signs the checksum file, and commits it.\n> \n> If during the export process (which is undertaken elsewhere, by a\n> different person or script), file A now has an expansion applied to it,\n> you break the checksum file, which you CANNOT redo, because you lose the\n> developer's digital signature on the checksum file!\n\nSimple, the release engineer runs a script which exports the tree,\nexpanding any keywords and updating the checksum file as necessary,\nand then the release engineer signs the checksum file!  As has already\nbeen stated, if this doesn't work, you probably don't have a well\ndefined and formal release process. \n\nJust because a developer has signed a checksum doesn't mean that the\ntree is suitable for release; that's the job of the release engineer\nto confirm, probably after running a set of regression test suites.\nAnd in fact, with git, it's pointless for the developer to sign a\nchecksum file and then commit it, since git is already maintaining\nchecksums as an integral part of how revisions are named.  \n\n\t\t\t\t\t- Ted\n"},{"id":"39442","messageId":"fcaeb9bf0704152023xaa119a4s8590452ff03befcf@mail.gmail.com","threadId":"7619","inReplyTo":"20070416030103.GB27533@thunk.org","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2007-04-16T03:23:26Z","receivedAt":"2007-04-16T03:23:26Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On 4/16/07, Theodore Tso <tytso@mit.edu> wrote:\n> Simple, the release engineer runs a script which exports the tree,\n> expanding any keywords and updating the checksum file as necessary,\n> and then the release engineer signs the checksum file!  As has already\n> been stated, if this doesn't work, you probably don't have a well\n> defined and formal release process.\n>\n> Just because a developer has signed a checksum doesn't mean that the\n> tree is suitable for release; that's the job of the release engineer\n> to confirm, probably after running a set of regression test suites.\n> And in fact, with git, it's pointless for the developer to sign a\n> checksum file and then commit it, since git is already maintaining\n> checksums as an integral part of how revisions are named.\n\nChanging Gentoo release process won't make Git the best choice while\nother SCM candidates can provide the same functionalities that Gentoo\nneeds without changing the process.\n-- \nDuy\n"},{"id":"39443","messageId":"20070416033209.GI2689@curie-int.orbis-terrarum.net","threadId":"7619","inReplyTo":"20070416030103.GB27533@thunk.org","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2007-04-16T03:32:09Z","receivedAt":"2007-04-16T03:32:09Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Sun, Apr 15, 2007 at 11:01:03PM -0400, Theodore Tso wrote:\n> On Sun, Apr 15, 2007 at 07:17:29PM -0700, Robin H. Johnson wrote:\n> > Nobody has addressed the single problem that I have with adding it when\n> > it's leaving the environment, and that's still of paramount concern to\n> > me. Simply put, there is a conflict between being able to add revision\n> > information of stuff leaving the environment, and those additions\n> > breaking previous checksums (which may be digitally signed, and thus\n> > breaking the signatures).\n> > \n> > I'll reduce it further from my previous example.\n> > \n> > 1. Developer commits some change to file A.\n> > 2. The checksum file is updated because A changed (the checksum file\n> >    explicitly does not contain keywords).\n> > 3. Developer signs the checksum file, and commits it.\n> > \n> > If during the export process (which is undertaken elsewhere, by a\n> > different person or script), file A now has an expansion applied to it,\n> > you break the checksum file, which you CANNOT redo, because you lose the\n> > developer's digital signature on the checksum file!\n> \n> Simple, the release engineer runs a script which exports the tree,\n> expanding any keywords and updating the checksum file as necessary,\n> and then the release engineer signs the checksum file!  As has already\n> been stated, if this doesn't work, you probably don't have a well\n> defined and formal release process. \nThe checksum file (named Manifest) we are talking about is for a single\nsubdirectory, and is signed as proof that it was not modified between\nthe developer and submission to the tree. \n\nAs I wrote originally, this is the Gentoo distribution tree, it's NOT\ndelineated by well-defined releases in the conventional sense.\n\nThere are presently 11571 Manifest files in the tree. Our tools will\nnot allow commits to each package of things that radically break the\npackage (semantic correctness and some automatic validation, but thinkos\ncan still get through the checks).\n\nThe 'release' process for the tree runs automatically every 30 minutes,\nand consists of more validation checks, updating a cache directory,\nproducing a signed master Manifest [1] and publishing everything to the\nrsync servers.\n\n> Just because a developer has signed a checksum doesn't mean that the\n> tree is suitable for release; that's the job of the release engineer\n> to confirm, probably after running a set of regression test suites.\n> And in fact, with git, it's pointless for the developer to sign a\n> checksum file and then commit it, since git is already maintaining\n> checksums as an integral part of how revisions are named.  \nThe entire point of the checksums is to allow end users to validate\ncontent that has been exported, with only minimal tools.\n\n[1] The master Manifest stage is only in production for the tree\ntarballs, and NOT in the rsync production at the moment, but will be\nwithin the next month. It exists solely to allow the detection of\ncompromised mirrors.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Council Member\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"39512","messageId":"200704161003.07679.andyparkins@gmail.com","threadId":"7619","inReplyTo":"Pine.LNX.4.64.0704151317180.5473@woody.linux-foundation.org","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-04-16T09:03:05Z","receivedAt":"2007-04-16T09:03:05Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Sunday 2007 April 15 21:51, Linus Torvalds wrote:\n\n> > Now, I print out that diagram and pin it to my wall - sometimes copies\n> > of it are given to others.  I do this on a regular basis.\n>\n> And is there *any* reason why you don't just do that as an \"export\"\n> option, when it's very clear that people won't send diffs that include it\n\nOf course there is a reason - the file I edit is the SVG itself, in inkscape \nwhile editing that file I press \"print\" to get a print out.  Why on earth \nwould I want to jump through hoops by closing the file I'm editing, running \nsome export script to a temporary file that I don't want, then open up \nInkscape again, check the export looks okay and then print - on what planet \nis /that/ simpler?  Worse, there is more chance that I'll lose changes once \nthere are two copies of the same file floating around.  Which one am I \nediting and which one am I printing?  Have I run the script yet?  When I \naccidentally make changes to the wrong one, I've now got to merge those \nchanges by hand back to the file they should have been in in the first place.\n\n> It's not a valid use because there are many SO MUCH BETTER WAYS to get the\n> same thing, that have none of the downsides of keyword expansion?\n\nI'm sorry, but we have different definitions of SO MUCH BETTER; it is _more_ \ntrouble for me the user to have to run scripts just to print the file that is \nalready on my screen, than not.\n\n> Your argument is akin to saying that \"Why isn't it a valid use to replace\n> the steering wheel in my car with a mouth-operated joystick under the\n> passenger side seat?\"\n\nI'd actually say that that is your argument - you want me to add steps to a \nprocess to get the same result.  I just want the steering wheel, you want the \nsteering wheel plus script that I run first to install the steering wheel and \ncorrectly adapt it for the current car.  In my version the process is \"I \npress print\"; the fact that is hard for the version control system is \nirrelevant - the whole point of tools like git is to do work for me, not the \nother way around.\n\n> The fact that you *can* do something is not a valid argument for it being\n> a valid use. You *can* do stupid things, but if you can get to the same\n> end result by not doing stupid things, wouldn't you prefer that instead?\n\nIt's not an accurate analogy at all.  Your conclusion is your supposition - \nit's stupid because it's stupid.  I don't understand what the huge problems \nare - all you've done is say again that it's a problem to have keyword \nexpansion.  Why?  What problem does it actually cause?\n\nI'm not just being argumentative - I still have not understood what terrible \nevil it is that keyword expansion causes but crlf conversion does not.\n\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"39514","messageId":"200704161010.47059.andyparkins@gmail.com","threadId":"7619","inReplyTo":"17954.48933.484379.593657@lisa.zopyra.com","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-04-16T09:10:45Z","receivedAt":"2007-04-16T09:10:45Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Monday 2007 April 16 01:11, Bill Lear wrote:\n\n> Not that Linus needs any back-up from me, but I second this, very\n> strongly.  Decorating source code with release information is a proper\n> function of release management tools, not the SCM system.  We had a\n> similar argument in our company about this, sparked by a criticism of\n> git for not having keyword (version number) substitution, and I argued\n> that having such substitution functions in the SCM was out-of-place\n> and a crutch for weak release procedures.  It's easy with a proper\n> make system to put whatever information you want from the SCM into the\n> release product.\n\nI'm not disagreeing with any of this - there are certainly cases when \nexpansion is completely the wrong tool.  That doesn't mean there are no cases \nwhere it would be useful.\n\nThe case I keep banging on about is that where nothing is made and this is not \na release.  I don't want to make a release, I just want to print out the \ncurrent version of a file and have something that appears on the printout \nthat would allow me to identify what version of the file that printout was \nfrom.  Are you seriously suggesting I should run release scripts just for \nthat?\n\nIt's not something you want - fine - not a problem for me that you wouldn't \nuse it.  The thing that is bothering me is that everyone keeps waving their \nhands while chanting \"keyword expansion evil\", while not giving an example of \nwhat problem it causes.  By this I mean \"problem for the end user\", \nnot \"problem in writing the support\" - if it's impractical to implement then \nthat's fine, say that.\n\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"39525","messageId":"Pine.LNX.4.64.0704160745040.5473@woody.linux-foundation.org","threadId":"7619","inReplyTo":"20070416021729.GH2689@curie-int.orbis-terrarum.net","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-16T14:59:32Z","receivedAt":"2007-04-16T14:59:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 15 Apr 2007, Robin H. Johnson wrote:\n>\n> Nobody has addressed the single problem that I have with adding it when\n> it's leaving the environment, and that's still of paramount concern to\n> me. Simply put, there is a conflict between being able to add revision\n> information of stuff leaving the environment, and those additions\n> breaking previous checksums (which may be digitally signed, and thus\n> breaking the signatures).\n\nDon't be silly. \n\nYou can just checksum without the ID. Which you have to do with git \nanyway, since any expanded ID *itself* would be part of any ID, which \nmeans that under git, you *physically*cannot* make an ID string be part of \nthe source control environment anyway, unless you did the SHA1 while \nignoring the $Id$ expansion.\n\nIn other words, the problem you talk about exists *regardless*. You \nsuggest pushing that problem into the SCM layer, and de-stabilizing the \nSCM and causing EVERYBODY ELSE provlems.\n\nAnd I'm telling you that if you want the idiocy of keyword expansion, you \ncan have it, BUT YOU CANNOT HAVE IT IN THE SCM.\n\nBecause *every* *single* problem you have with keyword expansion (whether \nit be checksums or anything else) will be MUCH MUCH worse if you do it at \nthe SCM level!\n\nReally. \n\nWhen you talk about your \"single problem\", why the HELL do you think that \nproblem goes away just because you try to deal with it inside the SCM? \nTrust me, the problem does *not* go away, it gets *bigger*.\n\nYou're trying to push it into the SCM, because _you_ don't want to deal \nwith the inevitable problems that keywords cause. But face it, the SCM \nwants to deal with them *even*less*, because they are much worse there, \nand more importantly, you'd be trying to push them into a level where most \nusers have gotten over the braindamage and no longer want it!\n\nSo you're trying to make *everybody* suffer, just because you cannot do it \nright. \n\nAnd suffer people do. There's a reason people are so negative about \nkeyword expansion: we've _seen_ those problems first-hand. \n\nSo the proper solution is:\n - don't do keyword expansion on the \"originals\".\n - add release information when you do a release. \n - if you want to sign releases, do so *after* the release. That's what a \n   release process is all about.\n - if you're so damn lazy that you can't be bothered to do the signing of \n   the release, don't ask others to do stupid things because *you* do \n   something stupid - just make sure that whatever release information you \n   add can be *removed*, so that you can verify an exact match.\n\nFor example, look at how \"git archive\" does this. It actually adds release \ninformation to the tar-file. It's hidden as a magic header, but that also \nmeans that since it's *separate* from the source code, it avoids all the \nproblems with keyword expansion, and now you can (for example) diff the \ntar-ball source tree with the git tree, and you will not get spurious AND \nINCORRECT differences! And any checksums would still be valid!\n\nAnd the same kind of thing can be done even if you absolutely have to \nembed the information on a file-by-file basis. Just make sure that you do \nit in some reversible manner. But preferably you generate a separate file \n(eg my hypothetical Makefile example that actually generates a \"prt\" file \nfrom a \"svg\" file) so that you have the original and can do any diff or \nvalidation efforts on *that*.\n\n\t\t\tLinus\n"},{"id":"39527","messageId":"Pine.LNX.4.64.0704160805280.5473@woody.linux-foundation.org","threadId":"7619","inReplyTo":"fcaeb9bf0704152023xaa119a4s8590452ff03befcf@mail.gmail.com","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-16T15:08:32Z","receivedAt":"2007-04-16T15:08:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 16 Apr 2007, Nguyen Thai Ngoc Duy wrote:\n> \n> Changing Gentoo release process won't make Git the best choice while\n> other SCM candidates can provide the same functionalities that Gentoo\n> needs without changing the process.\n\nAhh, the old \"argument by blackmail\" approach.\n\nYou know what? Nobody really cares. Arguing by blackmail (\"we'll use \nsomething else then\") just means that you should go somewhere else. If you \ncannot respond intelligently to intelligent arguments, you really *are* \nbetter off using SVN. \n\nA billion flies aren't exactly wrong: crap really *is* good. If you're a \nfly or a maggot.\n\nBut if you ever actually want to be something *more* than a crap eater, \ncome back then.\n\n\t\t\tLinus\n"},{"id":"39528","messageId":"Pine.LNX.4.64.0704161612120.5400@reaper.quantumfyre.co.uk","threadId":"7619","inReplyTo":"200704161010.47059.andyparkins@gmail.com","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2007-04-16T15:17:10Z","receivedAt":"2007-04-16T15:17:10Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Mon, 16 Apr 2007, Andy Parkins wrote:\n\n> On Monday 2007 April 16 01:11, Bill Lear wrote:\n>\n>> Not that Linus needs any back-up from me, but I second this, very\n>> strongly.  Decorating source code with release information is a proper\n>> function of release management tools, not the SCM system.  We had a\n>> similar argument in our company about this, sparked by a criticism of\n>> git for not having keyword (version number) substitution, and I argued\n>> that having such substitution functions in the SCM was out-of-place\n>> and a crutch for weak release procedures.  It's easy with a proper\n>> make system to put whatever information you want from the SCM into the\n>> release product.\n>\n> I'm not disagreeing with any of this - there are certainly cases when\n> expansion is completely the wrong tool.  That doesn't mean there are no cases\n> where it would be useful.\n>\n> The case I keep banging on about is that where nothing is made and this is not\n> a release.  I don't want to make a release, I just want to print out the\n> current version of a file and have something that appears on the printout\n> that would allow me to identify what version of the file that printout was\n> from.  Are you seriously suggesting I should run release scripts just for\n> that?\n>\n> It's not something you want - fine - not a problem for me that you wouldn't\n> use it.  The thing that is bothering me is that everyone keeps waving their\n> hands while chanting \"keyword expansion evil\", while not giving an example of\n> what problem it causes.  By this I mean \"problem for the end user\",\n> not \"problem in writing the support\" - if it's impractical to implement then\n> that's fine, say that.\n>\n\nWhat I don't understand is why the people who want keyword expansion don't \nsimply write a little wrapper script, a keyworded git as it were (you \ncould even call it gitk for maximum confusion :P).\n\nIn the script you simply:\n\n1) collapse all keywords\n2) call appropriate git function\n3) expand keywords again\n\nwouldn't that do what people want without having to change the git code at \nall?  You could probably even get it into contrib ..\n\n(In the case of gentoo, you could even change the ebuild so that the real \ngit is installed as raw_git or something, and the wrapper is installed as \ngit - though personally I wouldn't want to do that)\n\n-- \nJulian\n\n  ---\nYou may get an opportunity for advancement today.  Watch it!\n"},{"id":"39531","messageId":"20070416155456.GS955MdfPADPa@greensroom.kotnet.org","threadId":"7619","inReplyTo":"200704161003.07679.andyparkins@gmail.com","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Sven Verdoolaege","fromEmail":"skimo@kotnet.org","sentAt":"2007-04-16T15:54:56Z","receivedAt":"2007-04-16T15:54:56Z","isPatch":false,"sender":{"key":"skimo@kotnet.org","avatar":null},"body":"On Mon, Apr 16, 2007 at 10:03:05AM +0100, Andy Parkins wrote:\n> there are two copies of the same file floating around.  Which one am I \n> editing and which one am I printing?\n\nTurn off write permissions on the generated file.\n\n> Have I run the script yet?  When I \n\nUse a post-commit hook.\n\n> I'm not just being argumentative - I still have not understood what terrible \n> evil it is that keyword expansion causes but crlf conversion does not.\n\nFor one thing, this keyword expansion thing requires the SCM to modify\nthe file during commit.  (Hey, my editor says something changed the file.\nDo I have the file opened in another session?  Oh, it's the stupid\nkeyword expansion!)  AFAIU, crlf conversion will not change the working\ntree copy of your file on commit.\n\nskimo\n"},{"id":"39532","messageId":"Pine.LNX.4.64.0704160814300.5473@woody.linux-foundation.org","threadId":"7619","inReplyTo":"200704161003.07679.andyparkins@gmail.com","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-16T15:58:17Z","receivedAt":"2007-04-16T15:58:17Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 16 Apr 2007, Andy Parkins wrote:\n> \n> It's not an accurate analogy at all.  Your conclusion is your supposition - \n> it's stupid because it's stupid.  I don't understand what the huge problems \n> are - all you've done is say again that it's a problem to have keyword \n> expansion.  Why?  What problem does it actually cause?\n\nThe easiest way to explain it is that keyword expansion is like crlf, just \na million times worse (but if you were to do it in git, you'd literally \ndo it in the same path that does crlf expansion).\n\nLike crlf:\n - it requires you to be careful about binary vs non-binary, and corrupts \n   binary files subtly.\n - it never appears to be a problem as long as you stay inside the \"same \n   system\", because everybody just agrees.\n\nBut why did I actually implement auto-CRLF, if I'm so against it? Because \nkeyword expansion has a lot of problems that CRLF does *not* have:\n\n - pretty much every single tool out there actually handles CRLF \n   automatically. When you send emails from a CRLF system to a non-CRLF \n   system, the CRLF will just be removed. Why? Because tools *outside* the \n   SCM already know about \"text vs binary\", and while you can certainly \n   screw it up (use a CRLF system to generate a kernel patch and send it \n   as a binary attachment, and it won't apply for me, for example), you \n   actually have to work at it a bit.\n\n - A transformation like LF<->CRLF is \"stateless\". Anybody can translate a \n   file between CRLF and LF without having to know anything at all, so \n   even *if* somebody sends me a patch with CRLF (and it actually happens: \n   the amount of whitespace damage that people can do with email is just \n   surprisingly high, and people occasionally use Windows machines to send \n   me kernel patches, probably because they send email from some other \n   machine than the one they did development on).\n\n - Related to the statelessness: CRLF is a \"global\" operation, and doesn't \n   depend on file history or placement. Keyword expansion explicitly does\n   *not* work that way, since the whole *point* of keywords is to make it \n   depend on its place in history!\n\nAn example of real-world problems with that lack of statelessness of \nkeywords is something as simple as \"git rebase\". Think about what it does: \nit moves a commit around in history. But then think about *how* it does \nthat.\n\n[ Ok, take a break here, and think about why \"keyword expansion\" might be \n  a problem for \"git rebase\" in a way that CRLF is not, before you read on ]\n\nHint: the reason statefulness is broken for things like \"git rebase\" is \nthat the natural operation for something like that is to generate a patch, \nand carry it forward. Now, what is in the patch? Keywords. Will the patch \napply to the target? Yes? No?\n\nSee? Keywords means that you suddenly have merge problems with something \nas simple as patches. Does this matter in CVS? Not often. CVS is so \nlimited that you cannot much do those operations anyway, but if you've \never done a merge in CVS, keyword expansion tends to be one of the things \nthat just make it more complicated. So now you have to remember flags like \nlike \"-kk\" that disable keywords.\n\n(Not a lot of people actually do merges in CVS - branches are hard to use \nto begin with, so the only people who do branches tend to be pretty \nhardcode CVS people, and once you've learnt enough to do a branch, keyword \nexpansion is the least of your problems. But it's *one* reason - however \nsmall compared to the other reasons - that doing things like merging in \nCVS is just more painful than it should be)\n\nOr what about generating a diff between two branches? Keywords are a total \n*nightmare*. Do you realize just how *fast* git is in diff generation. \nHave you ever done \"cvs diff\"? Have you ever *thought* about how git can \nbe so fast? Hint: we don't even *look* at the contents for most files. But \nif the content is \"generated\" depending on history, you just screwed that \nup too.\n\nOr what about something as seemingly unrelated as \"git grep\". You may not \neven *realize* how nasty a problem it is when you have two different \nrepresentations of the same data: one that has keywords in it and is \nchecked out, and one that does not. Which one should you choose? Which one \nis the right one? What about the git optimization of using the checked-out \ndata because it doesn't need any unpacking?\n\nAgain, none of these things are problems with CRLF: CRLF is an issue that \nis pretty much *defined* to not matter for text-files. If you do a \"grep\", \nit doesn't matter if lines end in LF or CRLF. If you do a diff, line \nending differences (a) shouldn't exist in the first place because they are \nstateless and (b) even if they were to exist, they shouldn't change the \ndiff, because LF and CRLF are the same in text.\n\nAnd the whole keyword issue gets *worse* when you move between \nrepositories. If you stay \"inside\" the SCM, you can generally teach it to \nignore them. For example, going back to the \"git rebase\" example (or the \n\"git grep\" one, for that matter), you can just define that it's done \nwithout keyword expansion.\n\nBut when you move the data between people? That's exactly where keyword \nexpansion is enabled, and now you not only make things like \"git diff\" \nfundamentally broken and much much slower (in fact, it *cannot*work* in \nthe git model, because we don't even *have* tree history, so you cannot \nadd keywords to a tree!), you also guarantee that the end result is much \nless useful, because now when you send the patch to others, they'll have \nall the same issues that you had to work around locally.\n\nI don't know if I can convince you, but take it from me, keyword expansion \nis fundamentally broken in the first place, but it's *more* so with git \nthan with CVS, for example.\n\nIn CVS, the reason you can do keyword expansion in the first place is:\n\n - it's file-based to begin with. A file actually *has* history in CVS, in \n   a way it fundmanentally does *not* have in git. So when you generate a \n   diff on a file, the revision information is \"just there\". That's simply \n   not true in git. There *is* no per-file revision information. You \n   cannot know who touched the file last, for example, without starting \n   from a commit, and doing very expensive things.\n\n - it's slow to begin with. This is related to the above thing: exactly \n   because CVS is file-based and not content-based, when you do things \n   like \"cvs diff\" you will walk files individually anyway. People \n   *accept* (and I cannot imagine why) that an empty \"cvs diff\" on some \n   big project will take minutes. And the problems aren't even about \n   keyword expansion - keyword expansion is just a small detail.\n\n - it's centralized in more ways than one. You are simply not expected to \n   work by applying patches between two unrelated CVS trees. It's not \n   done. It cannot work. The closest you get is \n\t(a) merging. Which is *hell*. Again, keyword expansion is just a\n\t    small detail in why it's hell, and people don't generally pick \n\t    it up exactly because the merge problems are so much bigger.\n\t(b) applying patches from the outside from people who do *not* use \n\t    CVS, and thus don't generally touch things around the \n\t    keywords (but even here, you actually end up having problems\n\t    occasionally).\n\n - CVS really fundamentally has so many other problems that keyword\n   expansion just isn't on peoples radar. Yeah, it can corrupt data, but \n   you're more likely to corrupt data with binary files other ways, so \n   it's just not an issue.\n\nSo basically, other (more fundamental) design mistakes in CVS make \nkeywords seem like a better idea there, but all the keyword problems are \njust magnified ten-fold by the fact that git doesn't make those _other_ \nmistakes that CVS does.\n\nAnd don't get me wrong: I think RCS was a great step forward, and CVS was \ntoo. A few decades ago. But in git, we sometimes have to teach people to \n*not* make the mistakes they did with CVS. Keyword expansion is a small \ndetail, and happily few enough people used it in CVS that it's so far not \nbeen a huge problem to teach people not to do it.\n\nWe had to teach people that there's a difference between doing a local \nrepository commit, and pushing that commit to a shared central point. \nThat's a much more fundamental difference, and it's a lot harder to get \nyour brain to accept that kind of change. In contrast, keywords look \n\"trivial\", but they really aren't. It's a fundamentally broken notion, \neven if it *sounds* like a small detail.\n\nI'll finish off trying to explain the problem in fundamental git terms: \nsay you have a repository with two branches, A and B, and different \nhistory  on a file \"xyzzy\" in those two branches, but because they both \nended up applying the same patches, the actual file contents do end up \nbeing 100% identical. So they have the same SHA1.\n\nWhat is\n\n\tgit diff A..B -- xyzzy\n\nsupposed to print?\n\nAnd *I* claim that if you don't get an immediate and empty diff, your \nsystem is TOTALLY BROKEN.\n\nAnd now think about what keywords do. And realize that keywords are \nTOTALLY BROKEN!\n\n\t\t\tLinus\n"},{"id":"39533","messageId":"fcaeb9bf0704160906m3a2ffe60gf28570be1014403e@mail.gmail.com","threadId":"7619","inReplyTo":"Pine.LNX.4.64.0704160805280.5473@woody.linux-foundation.org","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2007-04-16T16:06:45Z","receivedAt":"2007-04-16T16:06:45Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On 4/16/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Mon, 16 Apr 2007, Nguyen Thai Ngoc Duy wrote:\n> >\n> > Changing Gentoo release process won't make Git the best choice while\n> > other SCM candidates can provide the same functionalities that Gentoo\n> > needs without changing the process.\n>\n> Ahh, the old \"argument by blackmail\" approach.\n>\n> You know what? Nobody really cares. Arguing by blackmail (\"we'll use\n> something else then\") just means that you should go somewhere else. If you\n> cannot respond intelligently to intelligent arguments, you really *are*\n> better off using SVN.\n\nAll right. I didn't mean to blackmail you or any Git developer. What I\nwanted to say is that Gentoo is currently using an old, brain-damaged\nSCM called CVS. I would like it to use Git but Git in its current\nstate can not fully replace CVS regarding to Gentoo usage. To do that\nGentoo needs some changes itself but Gentoo repositories are big ones\nand it's just hard to change such beast s from bottom up. So I would\nlike to see a compromise from Git (which, I think, does not harm other\nprojects from using Git) to ease the migration.\n\n>\n> A billion flies aren't exactly wrong: crap really *is* good. If you're a\n> fly or a maggot.\n>\n> But if you ever actually want to be something *more* than a crap eater,\n> come back then.\n>\n\nI would want to _slowly_ evolve from a crap eater to something better\nbecause I couldn't become a non-crap eater in a flash :)\n\n>                         Linus\n>\n\n\n-- \nDuy\n"},{"id":"39555","messageId":"Pine.LNX.4.64.0704160934050.5473@woody.linux-foundation.org","threadId":"7619","inReplyTo":"20070416033209.GI2689@curie-int.orbis-terrarum.net","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-16T17:00:04Z","receivedAt":"2007-04-16T17:00:04Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 15 Apr 2007, Robin H. Johnson wrote:\n>\n> The checksum file (named Manifest) we are talking about is for a single\n> subdirectory, and is signed as proof that it was not modified between\n> the developer and submission to the tree. \n\nWell, in git, you can actyally just take the tree entry for that \nsubdirectory, and it already is cryptographic proof that two \nsubdirectories match.\n\n(It's not signed, but if you actually want to sign it, you can do so, \neither inside git - by using a tag object that points to that \nsubdirectory - or outside git by just creating a Manifest that contains a \nlist of subdirectories and their tree SHA1's, and signing that).\n\nIn fact, in git, there's an explicit command to generate that \"Manifest of \ndirectories in the top level\", and it's called\n\n\tgit ls-tree HEAD\n\nand it will give you cryptographic hashes of each file/directory in the \ntop level of a repository. So just sign that, ie do\n\n\tgit ls-tree HEAD > Manifest\n\tgpg -sa -u \"$username\" Manifest \n\nor something like that. And you're done. Add the \"-r\" flag to get the \nrecursive manifest containing *all* files, rather than just the SHA1's of \nthe directories themselves.\n\nOf course, you could just sign and tag the HEAD itself, which is what the \nkernel does, since one signature will guarantee everything under it.\n\n> As I wrote originally, this is the Gentoo distribution tree, it's NOT\n> delineated by well-defined releases in the conventional sense.\n\nWe do that for the daily (or rather, nightly) snapshots for the kernel. \nThere's no \"Manifest\", but look at\n\n\thttp://www.kernel.org/pub/linux/kernel/v2.6/snapshots/\n\nand you'll see files like\n\n\tpatch-2.6.21-rc6-git8.bz2       15-Apr-2007 07:01   38K \t \n\tpatch-2.6.21-rc6-git8.bz2.sign  15-Apr-2007 07:01  248   \n\tpatch-2.6.21-rc6-git8.gz        15-Apr-2007 07:01   42K  \n\tpatch-2.6.21-rc6-git8.gz.sign   15-Apr-2007 07:01  248   \n\tpatch-2.6.21-rc6-git8.id        15-Apr-2007 07:01   41   \n\tpatch-2.6.21-rc6-git8.log       15-Apr-2007 07:01   63K  \n\tpatch-2.6.21-rc6-git8.sign      15-Apr-2007 07:01  248  \n\nwhere only the patches are signed, but the system *could* have signed the \nID file too (the 41-byte \"patch-2.6.21-rc6-git8.id\" contains the 40-byte \nHEX representation of the SHA of the HEAD of the snapshot, and a newline).\n\nThat 41-byte ID file really is sufficient to describe the whole thing, \nafter all (although you then need to have the git tree in question to \nactually get the list of files, aka the \"Manifest\", so if you want that \nlist, you'd have to do the \"git ls-tree\" thing.\n\n> There are presently 11571 Manifest files in the tree. Our tools will\n> not allow commits to each package of things that radically break the\n> package (semantic correctness and some automatic validation, but thinkos\n> can still get through the checks).\n\nSure. And every single Manifest file is pointless *inside* git, since git \nmaintains its own cryptographically secure manifest file anyway. But it's \ntrivial to generate them for external use, if you want to.\n\n> The 'release' process for the tree runs automatically every 30 minutes,\n> and consists of more validation checks, updating a cache directory,\n> producing a signed master Manifest [1] and publishing everything to the\n> rsync servers.\n\nThat sounds like the nightly snapshots the kernel does, except we only do \nthem nightly, and we don't actually validate anythign at all, we just sign \nthings as being from the \"master.kernel.org\" site (so the signature does \nmean something, but only that *that* site thinks it is valid).\n\n> The entire point of the checksums is to allow end users to validate\n> content that has been exported, with only minimal tools.\n\nIf you do a single 41-byte thing, you could use git itself to validate the \nwhole tree. But if you want to have people able to validate any random \nsingle file in a tar-file without having git installed, you'd have to:\n\n - have the \"full manifest\" (aka \"git ls-tree -r HEAD\")\n\n - have a trivial script that generates \"git ID's\" of files, which looks \n   something like this:\n\n\t#!/bin/sh\n\t# generate a \"git ID\" for one or more files\n\twhile test -n \"$1\"\n\tdo\n\t\tfile=\"$1\"\n\t\tlen=$(stat --format \"%s\" \"$file\")\n\t\techo -n \" $file (blob $len): \"\n\t\t# Generate the \"git ID\" for a blob:\n\t\t( echo -e -n \"blob $len\\0\" ; cat \"$file\") | sha1sum\n\t\tshift\n\tdone\n\nand now you can check each file in the Manifest even without having git \ninstalled.\n\n\t\t\tLinus\n"},{"id":"39568","messageId":"7vwt0crts2.fsf@assigned-by-dhcp.cox.net","threadId":"7619","inReplyTo":"200704161003.07679.andyparkins@gmail.com","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-16T19:41:49Z","receivedAt":"2007-04-16T19:41:49Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Andy Parkins <andyparkins@gmail.com> writes:\n\n> On Sunday 2007 April 15 21:51, Linus Torvalds wrote:\n>\n>> > Now, I print out that diagram and pin it to my wall - sometimes copies\n>> > of it are given to others.  I do this on a regular basis.\n>>\n>> And is there *any* reason why you don't just do that as an \"export\"\n>> option, when it's very clear that people won't send diffs that include it\n>\n> Of course there is a reason - the file I edit is the SVG itself, in inkscape \n> while editing that file I press \"print\" to get a print out.  Why on earth \n> would I want to jump through hoops by closing the file I'm editing, running \n> some export script to a temporary file that I don't want, then open up \n> Inkscape again, check the export looks okay and then print - on what planet \n> is /that/ simpler?\n\nI have one question.\n\nIn your workflow, when do you \"print\"?\n\nIf you did this:\n\n\t$ cvs update draw.svg\n        $ inkscape draw.svg\n        ... do more editing\n        ... press \"PRINT\"\n\t$ cvs diff draw.svg\n\nthe final \"cvs diff\" would say you have such and such changes to\nthe drawing file you just printed since the checked-in version.\nHowever, doesn't \"$Id: ... $\" embedded in the printed copy say\nit is from the last checked-in version?\n\nIs inkscape aware of the \"$Id: ... $\" keyword and modifies such\nstring by munging it to \"$Id: ..., modified $\", once you make a\nlocal modification to the document?  Otherwise you cannot tell\nif the printed copy is pristine and match what the $Id$ keyword\nclaims it is.\n\nOr maybe in your workflow, such a local modification may not\nactually matter because you made a habit of not making a drastic\nedit before printing.\n\nOr perhaps maybe you never print a locally modified copy.\n\nDoes Inkscape have a batch mode operation?  It might be an\noption to have something like this in the Makefile if it does (I\ndo not know if it does, and if so what the syntax is, so this is\ntotally made up):\n\n        print:: draw.svg\n                describe=$(git describe HEAD) && \\\n                git cat-file -p HEAD:draw.svg | \\\n                sed -e 's/$$Id$$/$$Id: '\"$$described\"'/g' | \\\n                inkscape --print --stdin\n\t.PHONY: print\n"},{"id":"39674","messageId":"200704162155.25114.andyparkins@gmail.com","threadId":"7619","inReplyTo":"7vwt0crts2.fsf@assigned-by-dhcp.cox.net","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-04-16T20:55:23Z","receivedAt":"2007-04-16T20:55:23Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Monday 2007, April 16, Junio C Hamano wrote:\n\n> In your workflow, when do you \"print\"?\n\nAfter a save and commit.  Otherwise - as you point out, the id is wrong.\n\n> the final \"cvs diff\" would say you have such and such changes to\n> the drawing file you just printed since the checked-in version.\n> However, doesn't \"$Id: ... $\" embedded in the printed copy say\n> it is from the last checked-in version?\n\nYep.  You will get no argument from me that keywords are by no means \ndefinitive.\n\n> Is inkscape aware of the \"$Id: ... $\" keyword and modifies such\n> string by munging it to \"$Id: ..., modified $\", once you make a\n\nNope.  Inkscape knows nothing about the expansion.  However, even if I \nwasn't careful to only print out checked in files, it would still \nnarrow down the possible versions to one of two.\n\n> local modification to the document?  Otherwise you cannot tell\n> if the printed copy is pristine and match what the $Id$ keyword\n> claims it is.\n\nCorrect.  Every user of keywords is aware that the keyword doesn't \nupdate all the time - in fact there's nothing to stop you changing the \nkeyword yourself to an utter lie.  I think the assumption is that you \naren't fighting your own tools though.\n\n> Or maybe in your workflow, such a local modification may not\n> actually matter because you made a habit of not making a drastic\n> edit before printing.\n\nYep.\n\n> Or perhaps maybe you never print a locally modified copy.\n\nYep.  In fact, for me, most of the time I'm printing a diagram that was \nchecked in a number of revisions ago.  It's not the case that I \nmodify-print.  However, that's just me.\n\n> Does Inkscape have a batch mode operation?  It might be an\n> option to have something like this in the Makefile if it does (I\n> do not know if it does, and if so what the syntax is, so this is\n> totally made up):\n\nI think it does as it happens; and your little script is just the sort \nof thing I will use when I get around to fixing this hole.\n\nHowever, it's missing the point to take my example as an unsolved \nproblem - there are plenty of ways I can get what I want; I brought it \nup merely as a counter to the statement that there were no valid \nsituations for wanting keyword expansion.\n\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"39580","messageId":"Pine.LNX.4.63.0704161528130.30610@qynat.qvtvafvgr.pbz","threadId":"7619","inReplyTo":"Pine.LNX.4.64.0704160814300.5473@woody.linux-foundation.org","subject":"Re: Weird shallow-tree conversion state, and branches of shallowtrees","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-04-16T23:25:37Z","receivedAt":"2007-04-16T23:25:37Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"I have a different situation where I'm interested in keyword expansions, and am \nwaiting for the appropriate hooks to be added to git to allow be to use it.\n\nI have a bunch of config files on different servers that are logicly equivalent, \neven though they have different values in some fields there is a translation \ntable in my software that tells it what to do.\n\nI'd really like to have a version control repository that I can \nshare/more/replicate across the machines. to do this on checkin the software \nwould need to run my helper to create a 'generic' version and check that in. on \ncheckout it would need to run my helper to take the generic version and make the \nhost specific version.\n\na lot of the problems taht you refer to in your message apply to most of the \nthings that have been discussed related to gitattributes.\n\nif improperly used it can corrupt the data (either by the checkin/checkout \nmunging or inappropriately merging things)\n\nit breaks the 1-1 coorespondance between the packed version and the checked out \nversion.\n\nOn Mon, 16 Apr 2007, Linus Torvalds wrote:\n\n> On Mon, 16 Apr 2007, Andy Parkins wrote:\n>\n> [ Ok, take a break here, and think about why \"keyword expansion\" might be\n>  a problem for \"git rebase\" in a way that CRLF is not, before you read on ]\n>\n> Hint: the reason statefulness is broken for things like \"git rebase\" is\n> that the natural operation for something like that is to generate a patch,\n> and carry it forward. Now, what is in the patch? Keywords. Will the patch\n> apply to the target? Yes? No?\n\nif you send a patch, that patch needs to be relative to the connonical version, \nnamely what's checked into the SCM. if your patch includes keywords it won't \napply cleanly to a checked-out of the file. any mergeing and merge resolution \nneeds to be based on the connonical version (i.e. one that doesn't go through \nthe checkin/out conversion)\n\n> See? Keywords means that you suddenly have merge problems with something\n> as simple as patches. Does this matter in CVS? Not often. CVS is so\n> limited that you cannot much do those operations anyway, but if you've\n> ever done a merge in CVS, keyword expansion tends to be one of the things\n> that just make it more complicated. So now you have to remember flags like\n> like \"-kk\" that disable keywords.\n\nI don't think the problems with patches are insurmountable. if everyone in the \nproject is useing git then you don't have to worry about anything, things will \njust work (except for manually fixing failed merges)\n\nI would definantly agree that sprinkling a little of this into a large project \nis going to massivly confuse people\n\n> Or what about generating a diff between two branches? Keywords are a total\n> *nightmare*. Do you realize just how *fast* git is in diff generation.\n> Have you ever done \"cvs diff\"? Have you ever *thought* about how git can\n> be so fast? Hint: we don't even *look* at the contents for most files. But\n> if the content is \"generated\" depending on history, you just screwed that\n> up too.\n\nyou do a diff of the connonical files in the repository, the same way you do \ntoday.\n\n> Or what about something as seemingly unrelated as \"git grep\". You may not\n> even *realize* how nasty a problem it is when you have two different\n> representations of the same data: one that has keywords in it and is\n> checked out, and one that does not. Which one should you choose? Which one\n> is the right one? What about the git optimization of using the checked-out\n> data because it doesn't need any unpacking?\n\nthe one with the keywords is the one to choose. and you suffer a performance hit \nbecouse you can't use the checked-out version (without running it through the \nconversion, which is a performance hit itslef)\n\n> And the whole keyword issue gets *worse* when you move between\n> repositories. If you stay \"inside\" the SCM, you can generally teach it to\n> ignore them. For example, going back to the \"git rebase\" example (or the\n> \"git grep\" one, for that matter), you can just define that it's done\n> without keyword expansion.\n\nright, this would avoid most of the problems\n\n> But when you move the data between people? That's exactly where keyword\n> expansion is enabled, and now you not only make things like \"git diff\"\n> fundamentally broken and much much slower (in fact, it *cannot*work* in\n> the git model, because we don't even *have* tree history, so you cannot\n> add keywords to a tree!), you also guarantee that the end result is much\n> less useful, because now when you send the patch to others, they'll have\n> all the same issues that you had to work around locally.\n\nwhy would you do keyword expansion when moving the files between different \npeople's repositories? or is that still considered 'inside the SCM'?\n\n> I don't know if I can convince you, but take it from me, keyword expansion\n> is fundamentally broken in the first place, but it's *more* so with git\n> than with CVS, for example.\n>\n> In CVS, the reason you can do keyword expansion in the first place is:\n>\n> - it's file-based to begin with. A file actually *has* history in CVS, in\n>   a way it fundmanentally does *not* have in git. So when you generate a\n>   diff on a file, the revision information is \"just there\". That's simply\n>   not true in git. There *is* no per-file revision information. You\n>   cannot know who touched the file last, for example, without starting\n>   from a commit, and doing very expensive things.\n\nthis is a valid argument against the keyword being a version string. it's not \nnessasarily relavent to other uses.\n\n> - it's slow to begin with. This is related to the above thing: exactly\n>   because CVS is file-based and not content-based, when you do things\n>   like \"cvs diff\" you will walk files individually anyway. People\n>   *accept* (and I cannot imagine why) that an empty \"cvs diff\" on some\n>   big project will take minutes. And the problems aren't even about\n>   keyword expansion - keyword expansion is just a small detail.\n\nif you define the keyword to be equivalent there is no need to look at the \ncontent of all the files.\n\n> - it's centralized in more ways than one. You are simply not expected to\n>   work by applying patches between two unrelated CVS trees. It's not\n>   done. It cannot work. The closest you get is\n> \t(a) merging. Which is *hell*. Again, keyword expansion is just a\n> \t    small detail in why it's hell, and people don't generally pick\n> \t    it up exactly because the merge problems are so much bigger.\n> \t(b) applying patches from the outside from people who do *not* use\n> \t    CVS, and thus don't generally touch things around the\n> \t    keywords (but even here, you actually end up having problems\n> \t    occasionally).\n\nexternal patches could be a problem, but there are two ways to deal with them.\n\n1. have the patch be against the version of the file with the keywords expanded, \nand have the result checked in (collapsing the keywords)\n\n2. have the patch be against the version of the file with the keywords \ncollapsed. this _would_ require the ability to bypass the expansion of the \nkeywords and is not something you would want to do very much.\n\nof these two, I suspect that #1 would make sense in most cases, and should be \nthe default.\n\n> I'll finish off trying to explain the problem in fundamental git terms:\n> say you have a repository with two branches, A and B, and different\n> history  on a file \"xyzzy\" in those two branches, but because they both\n> ended up applying the same patches, the actual file contents do end up\n> being 100% identical. So they have the same SHA1.\n>\n> What is\n>\n> \tgit diff A..B -- xyzzy\n>\n> supposed to print?\n>\n> And *I* claim that if you don't get an immediate and empty diff, your\n> system is TOTALLY BROKEN.\n\nI agree, and what I've been talking about above would produce exactly this.\n\n> And now think about what keywords do. And realize that keywords are\n> TOTALLY BROKEN!\n\nit may be that we are thinking of different things when we use the term \n'keywords', and that may be why we are seeing different levels of problems.\n\nDavid Lang\n"},{"id":"39597","messageId":"Pine.LNX.4.64.0704162346550.27922@iabervon.org","threadId":"7619","inReplyTo":"20070416033209.GI2689@curie-int.orbis-terrarum.net","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2007-04-17T04:16:42Z","receivedAt":"2007-04-17T04:16:42Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sun, 15 Apr 2007, Robin H. Johnson wrote:\n\n> The checksum file (named Manifest) we are talking about is for a single\n> subdirectory, and is signed as proof that it was not modified between\n> the developer and submission to the tree. \n\nSo the process has to be:\n\n1. Developer commits changes to files.\n2. Checksum utility finds the checksums of the files with IDs added where \n   the master site updater will add them.\n3. Developer signs checksums.\n4. Developer commits checksums.\n5. Developer pushes changes to master site.\n6. Master site checks out files, adds IDs, and updates live tree.\n7. End user fetches tree.\n8. End user checks checksums, which match, because the master site and the \n   developer checksum scripts agree on what the end user will see.\n\nThe only difference is that developers working out of the version control \nhave to generate the checksums with a tool that knows how the IDs will be \nadded, and check the checksums with this tool as well, because working \ndirectories don't have IDs in them.\n\nReally, it's approximately the same as having the version control system \ndo it, except that it's in the project-specific development tools instead \nof the version control system.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"39629","messageId":"200704171045.20897.andyparkins@gmail.com","threadId":"7619","inReplyTo":"Pine.LNX.4.64.0704160814300.5473@woody.linux-foundation.org","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-04-17T09:45:19Z","receivedAt":"2007-04-17T09:45:19Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Monday 2007 April 16 16:58, Linus Torvalds wrote:\n\nThank you for the detailed response.  My apologies for the delay in replying, \nI did write and send a response, but it's gone missing in the world of \ngoogle's SMTP server.  I'll try and resend when I return home.\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"39694","messageId":"Pine.LNX.4.63.0704171250120.1696@qynat.qvtvafvgr.pbz","threadId":"7619","inReplyTo":"Pine.LNX.4.63.0704161528130.30610@qynat.qvtvafvgr.pbz","subject":"Re: Weird shallow-tree conversion state, and branches of shallowtrees","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-04-17T19:50:46Z","receivedAt":"2007-04-17T19:50:46Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"sorry for the re-send, I didn't see this go through the list and it's relavent \nto the current discussion\n\nOn Mon, 16 Apr 2007, David Lang wrote:\n\n> Date: Mon, 16 Apr 2007 16:25:37 -0700 (PDT)\n> \n> I have a different situation where I'm interested in keyword expansions, and \n> am waiting for the appropriate hooks to be added to git to allow be to use \n> it.\n>\n> I have a bunch of config files on different servers that are logicly \n> equivalent, even though they have different values in some fields there is a \n> translation table in my software that tells it what to do.\n>\n> I'd really like to have a version control repository that I can \n> share/more/replicate across the machines. to do this on checkin the software \n> would need to run my helper to create a 'generic' version and check that in. \n> on checkout it would need to run my helper to take the generic version and \n> make the host specific version.\n>\n> a lot of the problems taht you refer to in your message apply to most of the \n> things that have been discussed related to gitattributes.\n>\n> if improperly used it can corrupt the data (either by the checkin/checkout \n> munging or inappropriately merging things)\n>\n> it breaks the 1-1 coorespondance between the packed version and the checked \n> out version.\n>\n> On Mon, 16 Apr 2007, Linus Torvalds wrote:\n>\n>> On Mon, 16 Apr 2007, Andy Parkins wrote:\n>> \n>> [ Ok, take a break here, and think about why \"keyword expansion\" might be\n>>  a problem for \"git rebase\" in a way that CRLF is not, before you read on ]\n>> \n>> Hint: the reason statefulness is broken for things like \"git rebase\" is\n>> that the natural operation for something like that is to generate a patch,\n>> and carry it forward. Now, what is in the patch? Keywords. Will the patch\n>> apply to the target? Yes? No?\n>\n> if you send a patch, that patch needs to be relative to the connonical \n> version, namely what's checked into the SCM. if your patch includes keywords \n> it won't apply cleanly to a checked-out of the file. any mergeing and merge \n> resolution needs to be based on the connonical version (i.e. one that doesn't \n> go through the checkin/out conversion)\n>\n>> See? Keywords means that you suddenly have merge problems with something\n>> as simple as patches. Does this matter in CVS? Not often. CVS is so\n>> limited that you cannot much do those operations anyway, but if you've\n>> ever done a merge in CVS, keyword expansion tends to be one of the things\n>> that just make it more complicated. So now you have to remember flags like\n>> like \"-kk\" that disable keywords.\n>\n> I don't think the problems with patches are insurmountable. if everyone in \n> the project is useing git then you don't have to worry about anything, things \n> will just work (except for manually fixing failed merges)\n>\n> I would definantly agree that sprinkling a little of this into a large \n> project is going to massivly confuse people\n>\n>> Or what about generating a diff between two branches? Keywords are a total\n>> *nightmare*. Do you realize just how *fast* git is in diff generation.\n>> Have you ever done \"cvs diff\"? Have you ever *thought* about how git can\n>> be so fast? Hint: we don't even *look* at the contents for most files. But\n>> if the content is \"generated\" depending on history, you just screwed that\n>> up too.\n>\n> you do a diff of the connonical files in the repository, the same way you do \n> today.\n>\n>> Or what about something as seemingly unrelated as \"git grep\". You may not\n>> even *realize* how nasty a problem it is when you have two different\n>> representations of the same data: one that has keywords in it and is\n>> checked out, and one that does not. Which one should you choose? Which one\n>> is the right one? What about the git optimization of using the checked-out\n>> data because it doesn't need any unpacking?\n>\n> the one with the keywords is the one to choose. and you suffer a performance \n> hit becouse you can't use the checked-out version (without running it through \n> the conversion, which is a performance hit itslef)\n>\n>> And the whole keyword issue gets *worse* when you move between\n>> repositories. If you stay \"inside\" the SCM, you can generally teach it to\n>> ignore them. For example, going back to the \"git rebase\" example (or the\n>> \"git grep\" one, for that matter), you can just define that it's done\n>> without keyword expansion.\n>\n> right, this would avoid most of the problems\n>\n>> But when you move the data between people? That's exactly where keyword\n>> expansion is enabled, and now you not only make things like \"git diff\"\n>> fundamentally broken and much much slower (in fact, it *cannot*work* in\n>> the git model, because we don't even *have* tree history, so you cannot\n>> add keywords to a tree!), you also guarantee that the end result is much\n>> less useful, because now when you send the patch to others, they'll have\n>> all the same issues that you had to work around locally.\n>\n> why would you do keyword expansion when moving the files between different \n> people's repositories? or is that still considered 'inside the SCM'?\n>\n>> I don't know if I can convince you, but take it from me, keyword expansion\n>> is fundamentally broken in the first place, but it's *more* so with git\n>> than with CVS, for example.\n>> \n>> In CVS, the reason you can do keyword expansion in the first place is:\n>> \n>> - it's file-based to begin with. A file actually *has* history in CVS, in\n>>   a way it fundmanentally does *not* have in git. So when you generate a\n>>   diff on a file, the revision information is \"just there\". That's simply\n>>   not true in git. There *is* no per-file revision information. You\n>>   cannot know who touched the file last, for example, without starting\n>>   from a commit, and doing very expensive things.\n>\n> this is a valid argument against the keyword being a version string. it's not \n> nessasarily relavent to other uses.\n>\n>> - it's slow to begin with. This is related to the above thing: exactly\n>>   because CVS is file-based and not content-based, when you do things\n>>   like \"cvs diff\" you will walk files individually anyway. People\n>>   *accept* (and I cannot imagine why) that an empty \"cvs diff\" on some\n>>   big project will take minutes. And the problems aren't even about\n>>   keyword expansion - keyword expansion is just a small detail.\n>\n> if you define the keyword to be equivalent there is no need to look at the \n> content of all the files.\n>\n>> - it's centralized in more ways than one. You are simply not expected to\n>>   work by applying patches between two unrelated CVS trees. It's not\n>>   done. It cannot work. The closest you get is\n>> \t(a) merging. Which is *hell*. Again, keyword expansion is just a\n>> \t    small detail in why it's hell, and people don't generally pick\n>> \t    it up exactly because the merge problems are so much bigger.\n>> \t(b) applying patches from the outside from people who do *not* use\n>> \t    CVS, and thus don't generally touch things around the\n>> \t    keywords (but even here, you actually end up having problems\n>> \t    occasionally).\n>\n> external patches could be a problem, but there are two ways to deal with \n> them.\n>\n> 1. have the patch be against the version of the file with the keywords \n> expanded, and have the result checked in (collapsing the keywords)\n>\n> 2. have the patch be against the version of the file with the keywords \n> collapsed. this _would_ require the ability to bypass the expansion of the \n> keywords and is not something you would want to do very much.\n>\n> of these two, I suspect that #1 would make sense in most cases, and should be \n> the default.\n>\n>> I'll finish off trying to explain the problem in fundamental git terms:\n>> say you have a repository with two branches, A and B, and different\n>> history  on a file \"xyzzy\" in those two branches, but because they both\n>> ended up applying the same patches, the actual file contents do end up\n>> being 100% identical. So they have the same SHA1.\n>> \n>> What is\n>>\n>> \tgit diff A..B -- xyzzy\n>> \n>> supposed to print?\n>> \n>> And *I* claim that if you don't get an immediate and empty diff, your\n>> system is TOTALLY BROKEN.\n>\n> I agree, and what I've been talking about above would produce exactly this.\n>\n>> And now think about what keywords do. And realize that keywords are\n>> TOTALLY BROKEN!\n>\n> it may be that we are thinking of different things when we use the term \n> 'keywords', and that may be why we are seeing different levels of problems.\n>\n> David Lang\n>\n"},{"id":"39709","messageId":"7v8xcqofru.fsf@assigned-by-dhcp.cox.net","threadId":"7619","inReplyTo":"200704162155.25114.andyparkins@gmail.com","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-17T21:24:53Z","receivedAt":"2007-04-17T21:24:53Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Andy Parkins <andyparkins@gmail.com> writes:\n\n> However, it's missing the point to take my example as an unsolved \n> problem - there are plenty of ways I can get what I want; I brought it \n> up merely as a counter to the statement that there were no valid \n> situations for wanting keyword expansion.\n\nThat's actually quite different from what you said.\n\nAndy Parkins <andyparkins@gmail.com> writes:\n\n> Of course there is a reason - the file I edit is the SVG\n> itself, in inkscape while editing that file I press \"print\" to\n> get a print out.  Why on earth would I want to jump through\n> hoops by closing the file I'm editing, running some export\n> script to a temporary file that I don't want, then open up\n> Inkscape again, check the export looks okay and then print -\n> on what planet is /that/ simpler?\n\nYou were claiming that with built-in keyword expansion what you\nwant becomes /simpler/.  I questioned that.\n\nMaybe it's just me, who is not a GUI person [*1*], but to me,\nhaving to start inkscape, mouse around to find the \"Print\"\nbutton and print feels much more cumbersome than simply typing\n\"make print\".\n\n[Footnote]\n\n*1* Not in the sense I do not program GUIy applications, but in\nthe sense that I do not usually _use_ GUI applications.\n"},{"id":"39711","messageId":"200704172251.20831.andyparkins@gmail.com","threadId":"7619","inReplyTo":"7v8xcqofru.fsf@assigned-by-dhcp.cox.net","subject":"Re: Weird shallow-tree conversion state, and branches of shallow trees","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-04-17T21:51:18Z","receivedAt":"2007-04-17T21:51:18Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Tuesday 2007, April 17, Junio C Hamano wrote:\n> Andy Parkins <andyparkins@gmail.com> writes:\n> > However, it's missing the point to take my example as an unsolved\n> > problem - there are plenty of ways I can get what I want; I brought\n> > it up merely as a counter to the statement that there were no valid\n> > situations for wanting keyword expansion.\n>\n> That's actually quite different from what you said.\n\nSorry; I didn't express it very well - the thing that started all this \nwas the statement that there was no valid use case for keywords.  I \njust gave an example.  I felt that the thread was moving away from \nkeywords and towards solving my particular problem - which is all \nappreciated, but wasn't the point.  Running makefile recipes or extra \nscripts are all valid methods and pragmatic \nworking-with-what-git-does-now solutions.  I wanted to distinguish \nbetween what I could do now and what I could do with keyword support.\n\n> You were claiming that with built-in keyword expansion what you\n> want becomes /simpler/.  I questioned that.\n\nWell it does from the point of view of pressing \"print\".\n\n> Maybe it's just me, who is not a GUI person [*1*], but to me,\n> having to start inkscape, mouse around to find the \"Print\"\n> button and print feels much more cumbersome than simply typing\n> \"make print\".\n\nAgain, that was addressing my particular problem - good stuff.  However, \nit's just luck that inkscape has a batch mode - there's no guarantee \nfor that.\n\nI could just swap the example around a bit, what about if it was an \nOpenOffice document that I want to have transparent \ncompression/decompression and I've set the properties tag to \ncontain \"$Id$\".  There is no amount of scripting that will enable batch \nprinting of that.\n\nAnyway - I've wasted enough of your time with this foolishness now.  \nIt's dropped, consider me silenced on this subject ;-)\n\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"}]}