{"thread":{"id":"348","subject":"A shortcoming of the git repo format","startedAt":"2005-04-27T05:43:13Z","lastAt":"2005-04-28T15:08:26Z","messageCount":29,"participants":["H. Peter Anvin","C. Scott Ananian","Linus Torvalds","Dave Jones","Petr Baudis","Brian O'Mahoney","Tom Lord","Gerhard Schrenk","Jon Seymour","Daniel Barkalow","David A. Wheeler","David Lang","Paul Jackson","Ryan Anderson","Morgan Schweers","David Woodhouse","Barry Silverman"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"1847","messageId":"426F2671.1080105@zytor.com","threadId":"348","inReplyTo":null,"subject":"A shortcoming of the git repo format","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-27T05:43:13Z","receivedAt":"2005-04-27T05:43:13Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Most of git's files are starting to converge toward an RFC822-like \nheader with (tag, data) and a free-form section.  This is a good thing. \n  However, there is one problem with this, and that is that without \nknowing every possible tag, a program reading the git repository cannot \nsafely tell what is a link to another git object and what is not.  When \nI did my repository conversion tools, I simply assumed any string of 20 \nhexadecimal digits was a pointer, but this is probably a bad idea in the \nlong run.\n\nAdditionally, there is the question of the handling of strings that may \ncontain \\n or even \\0 (which may be necessary for some applications).\n\nOne solution to all of this would be to define a quoting standard for \nstrings, and simply require that all free-format strings (like the \nauthor fields) or at least strings that match [0-9a-f]{20}, are always \nquoted.\n\nI propose the following:\n\n- Any string containing control characters or \\ must be quoted;\n- \\xXX produces control characters; other characters following \\ are \nverbatim.\n\nThus,\n\nlink 0123456789abcdef0123\n\n... is a link to an object, whereas ...\n\nstring \\0123456789abcdef0123\n\n... is a string.\n\nstring1  This string begins with a space\nstring2 This string has an embedded newline (\"\\x0a\")\n\n... are both valid strings; the first contains a leading space and the \nsecond an embedded newline.\n\nI'll implement this and integrate it tomorrow.\n\n\t-hpa\n\n\n"},{"id":"1862","messageId":"Pine.LNX.4.61.0504271058120.5008@cag.csail.mit.edu","threadId":"348","inReplyTo":"426F2671.1080105@zytor.com","subject":"Re: A shortcoming of the git repo format","fromName":"C. Scott Ananian","fromEmail":"cscott@cscott.net","sentAt":"2005-04-27T15:00:46Z","receivedAt":"2005-04-27T15:00:46Z","isPatch":false,"sender":{"key":"cscott@cscott.net","avatar":"https://gravatar.com/avatar/3551c2aefb299a0c45807f7677f5b26d8a5be4a4af359b4bf4fabbdd1f2b990e?d=mp&s=160"},"body":"On Tue, 26 Apr 2005, H. Peter Anvin wrote:\n\n> Additionally, there is the question of the handling of strings that may \n> contain \\n or even \\0 (which may be necessary for some applications).\n\nWhile we're at it, I'll just mention that '\\0' is a rather bad delimiter \nfor zlib-compressed files; it usually ends up enlarging the file by three \nor more bytes compared to using any whitespace character.  The reason is \nobvious: \\0 isn't actually used anywhere else in the compressed contents, \nso it tends to pollute zlib's dictionary.\n\nIt's probably too late to do anything about this, but hey.\n  --scott\n\nSoviet  STANDEL Yakima JMTRAX Hussein Ft. Meade algorithm JMBLUG CIA \nSEQUIN Bejing Morwenstow Boston nuclear Sigint Ft. Bragg ZRBRIEF Peking\n                          ( http://cscott.net/ )\n"},{"id":"1865","messageId":"Pine.LNX.4.58.0504270820370.18901@ppc970.osdl.org","threadId":"348","inReplyTo":"426F2671.1080105@zytor.com","subject":"Re: A shortcoming of the git repo format","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-27T15:22:07Z","receivedAt":"2005-04-27T15:22:07Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 26 Apr 2005, H. Peter Anvin wrote:\n> \n> One solution to all of this would be to define a quoting standard for \n> strings, and simply require that all free-format strings (like the \n> author fields) or at least strings that match [0-9a-f]{20}, are always \n> quoted.\n\ngit uses more of the \".newsrc\" format, in that it just knows which \ncharacters are legal or not.\n\nTo find the email address, look for the first '<'. To find the date, look \nfor the first '>'. Those characters are not allowed in the name or the \nemail, so they act as well-defined delimeters.\n\n\t\tLinus\n"},{"id":"1882","messageId":"426FD3EE.5000404@zytor.com","threadId":"348","inReplyTo":"Pine.LNX.4.58.0504270820370.18901@ppc970.osdl.org","subject":"Re: A shortcoming of the git repo format","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-27T18:03:26Z","receivedAt":"2005-04-27T18:03:26Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Tue, 26 Apr 2005, H. Peter Anvin wrote:\n> \n>>One solution to all of this would be to define a quoting standard for \n>>strings, and simply require that all free-format strings (like the \n>>author fields) or at least strings that match [0-9a-f]{20}, are always \n>>quoted.\n> \n> \n> git uses more of the \".newsrc\" format, in that it just knows which \n> characters are legal or not.\n> \n> To find the email address, look for the first '<'. To find the date, look \n> for the first '>'. Those characters are not allowed in the name or the \n> email, so they act as well-defined delimeters.\n> \n\nThat's true for email addresses, but the point was to distinguish links \nto other git objects from any other kind of text.  Currently there is no \nsuch delimiter for that.  Another solution than the one I posted would \nbe to define such a delimiter, for example '<' + 20 hex character + '>' \n(which would be distinguished from email addresses by the lack of an @ \nsign.)  That would be a repo change, though.\n\nGiven no prior constraints, I would probably argue for a format which \nmakes the data type known as a matter of syntax, using \"...\" quoted \nstrings for *ALL* arbitrary strings, a different syntax for numbers and \nlinks, and leaving the door open for new data types like lists in the \nfuture.\n\n\t-hpa\n\n"},{"id":"1885","messageId":"20050427183239.GE19011@redhat.com","threadId":"348","inReplyTo":"426FD3EE.5000404@zytor.com","subject":"Re: A shortcoming of the git repo format","fromName":"Dave Jones","fromEmail":"davej@redhat.com","sentAt":"2005-04-27T18:32:40Z","receivedAt":"2005-04-27T18:32:40Z","isPatch":false,"sender":{"key":"davej@redhat.com","avatar":null},"body":"On Wed, Apr 27, 2005 at 11:03:26AM -0700, H. Peter Anvin wrote:\n > Linus Torvalds wrote:\n > >\n > >On Tue, 26 Apr 2005, H. Peter Anvin wrote:\n > >\n > >>One solution to all of this would be to define a quoting standard for \n > >>strings, and simply require that all free-format strings (like the \n > >>author fields) or at least strings that match [0-9a-f]{20}, are always \n > >>quoted.\n > >\n > >\n > >git uses more of the \".newsrc\" format, in that it just knows which \n > >characters are legal or not.\n > >\n > >To find the email address, look for the first '<'. To find the date, look \n > >for the first '>'. Those characters are not allowed in the name or the \n > >email, so they act as well-defined delimeters.\n > >\n > \n > That's true for email addresses, but the point was to distinguish links \n > to other git objects from any other kind of text.  Currently there is no \n > such delimiter for that.\n\nThat actually broke one of my first git scripts when one of the\nchangelog texts started a line with 'tree '.  I hacked around it\nby making my script only grep in the 'head -n4' lines, but this\nseems somewhat fragile having to make assumptions that the field\nI want to see is in the first 4 lines.\n\n\t\tDave\n\n"},{"id":"1886","messageId":"426FDE48.1050700@zytor.com","threadId":"348","inReplyTo":"20050427183239.GE19011@redhat.com","subject":"Re: A shortcoming of the git repo format","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-27T18:47:36Z","receivedAt":"2005-04-27T18:47:36Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Dave Jones wrote:\n> \n> That actually broke one of my first git scripts when one of the\n> changelog texts started a line with 'tree '.  I hacked around it\n> by making my script only grep in the 'head -n4' lines, but this\n> seems somewhat fragile having to make assumptions that the field\n> I want to see is in the first 4 lines.\n> \n\nYou have the delimiter for that; there is an empty line between the \nheader and the free-form body, similar as for RFC822.\n\n\t-hpa\n"},{"id":"1889","messageId":"Pine.LNX.4.58.0504271154470.18901@ppc970.osdl.org","threadId":"348","inReplyTo":"426FD3EE.5000404@zytor.com","subject":"Re: A shortcoming of the git repo format","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-27T19:11:04Z","receivedAt":"2005-04-27T19:11:04Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 27 Apr 2005, H. Peter Anvin wrote:\n> \n> That's true for email addresses, but the point was to distinguish links \n> to other git objects from any other kind of text.\n\nNo, that's definitely _not_ the point.\n\nI repeat: git does not do any free-form parsin AT ALL. The links are in \nwell-defined places, and you do not ever search for them. And that's \nreally very very important.\n\n> Currently there is no  such delimiter for that.\n\nThere absolutely is.\n\nFor a \"commit\", the format is\n\n - first line is exactly 46 bytes: five bytes of \"tree \", 40 bytes of hex \n   sha1, and one byte of \"\\n\".\n\n   NOTHING ELSE. Not extra spaces at the end, not extra spaces at the \n   beginning or the middle. It's ASCII, but it's not free-format ASCII.\n\n - the next <n> (where 'n' can be 0 or more) lines are _exactly_ 48 bytes\n   each:  seven bytes of \"parent \", 40 bytes of hex sha1, and one byte of \n   \"\\n\".\n\n   NOTHING ELSE.\n\n - the next lines are \"author \" and \"committer \". They have well-defined \n   delimters for their fields, and no sha1's. The fields cannot contain \n   '<', '>' or newlines, since those are the field/line delimeters.\n\nThere is no free-format text _anywhere_ that git parses. No room for \nguesses, no room for mistakes, no room for anything half-way questionable.\n\nAnd fsck actually enforces this. We do _not_ just use \"gets()\" to read one \nline at a time. We literally verify that the lines are 46/48 bytes long, \nand have the delimeters in the expected places.\n\nSame goes for \"tree\" and \"tag\" objects. They all have fixed-format stuff. \nA \"tree\" entry is always\n\n\t\"%o <space> %s\" \\0 [ 20 bytes of sha1 ]\n\nwith \"%o\" being \"mode\", and \"%s\" being \"path\". We don't guess. \n\nAnd this really is _important_. Exactly because we name things by the SHA1\nhash of the contents, we MUST NOT have flexible formats. Having a format\nwhich allows non-canonical representations (extra spaces etc) would mean\nthat two trees that were identical would depend on how you happened to\nformat them.\n\nSo there's really two issues:\n - we don't guess or parse contents. We have strict rules, and that makes \n   git more reliable. There are no gray areas. There's \"right\" and there \n   is \"wrong\", and the right one works, and the wrong one gets flagged as \n   being wrong and the tools refuse to touch it.\n - there is only _one_ right way to do things, and that means that the \n   the content is well-defined, and thus the SHA1 of the content is \n   well-defined.\n\nFor example, another rule is that a \"tree\" object is always sorted by \nthe bytes in the filename (not by entry, btw: a directory called \"foo\" \nwill sort as \"foo/\", even though the _entry_ only shows \"foo\"). That rule \nnot only makes a lot of operations faster, but again, it means that there \nis only _one_ way to represent a tree validly.\n\nIOW, you _cannot_ represent a tree any other way (and I've been too lazy\nto check this in fsck, but it's alway sbeen my plan), and that is exactly \nwhy we can just compare the hashes of the results - because there is no \nrandom component of \"layout\" in the contents.\n\nThis really is important. It means that if you get to the same two tree\ncontents in totally unrelated ways (you unpack a tar-file and encode it in\ngit, or you have 5 years of git history and check it out), the \"tree\" will\nmatch _exactly_. There's no history. There's no \"optional\" stuff. Since\nthe contents of the trees are the same, the SHA1 of the two trees will be\nthe same. Exactly because git refuses to touch any free-format stuff.\n\n\t\tLinus\n"},{"id":"1890","messageId":"Pine.LNX.4.58.0504271212460.18901@ppc970.osdl.org","threadId":"348","inReplyTo":"20050427183239.GE19011@redhat.com","subject":"Re: A shortcoming of the git repo format","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-27T19:15:10Z","receivedAt":"2005-04-27T19:15:10Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 27 Apr 2005, Dave Jones wrote:\n> \n> That actually broke one of my first git scripts when one of the\n> changelog texts started a line with 'tree '.  I hacked around it\n> by making my script only grep in the 'head -n4' lines, but this\n> seems somewhat fragile having to make assumptions that the field\n> I want to see is in the first 4 lines.\n\nIt's not an assumption.\n\nIT'S THE LAW.\n\nThe speed of light is not \"an assumption\". It is.\n\nThe tree is in the first line of a commit. You don't even need to parse \nit, you do\n\n\ttree=$(cat-file commit $head | sed 's/tree //;q')\n\nand that's it. No parsing.\n\nGit doesn't guess. Git knows.\n\n\t\tLinus\n"},{"id":"1892","messageId":"20050427193903.GG22956@pasky.ji.cz","threadId":"348","inReplyTo":"20050427183239.GE19011@redhat.com","subject":"Re: A shortcoming of the git repo format","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-27T19:39:03Z","receivedAt":"2005-04-27T19:39:03Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Wed, Apr 27, 2005 at 08:32:40PM CEST, I got a letter\nwhere Dave Jones <davej@redhat.com> told me that...\n> That actually broke one of my first git scripts when one of the\n> changelog texts started a line with 'tree '.  I hacked around it\n> by making my script only grep in the 'head -n4' lines, but this\n> seems somewhat fragile having to make assumptions that the field\n> I want to see is in the first 4 lines.\n\nThe tree field is now always at the first line, but generally the header\npart is variable-sized; you have multiple parent lines in case of\nmerges.\n\nJust stop reading at the first newline.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"1893","messageId":"426FEC38.1060507@khandalf.com","threadId":"348","inReplyTo":"Pine.LNX.4.58.0504271154470.18901@ppc970.osdl.org","subject":"The git repo format","fromName":"Brian O'Mahoney","fromEmail":"omb@khandalf.com","sentAt":"2005-04-27T19:47:04Z","receivedAt":"2005-04-27T19:47:04Z","isPatch":false,"sender":{"key":"omb@khandalf.com","avatar":null},"body":"In understanding how to work with 'git' I had a number of initial\ndifficulties which are mostly covered by the e-mail from Linus below.\n\nMost of these are already covered in the README:\n\nfor objects, ie blob, commit, tag, tree: inflate, then\n<type>\\s<size>\\0<data>\n\nwhere <data> is in the form, described by Linus below\n\nwhen you look at them closely, all the formats are simple,\nun-ambiguous, and very easy to parse.\n\nThe index is also easy to parse, but there is a detail,\nafter the 3-int header the records are padded to a multiple\nof 8 bytes. The detail is in cache.h.\n\nMaybe the README needs to re-inforce this.\n\nBrian\n\n> I repeat: git does not do any free-form parsin AT ALL.\n========================================================\n\n The links are in well-defined places, and you do not ever search for them.\n\nAnd that's really very very important.\n\n\n> For a \"commit\", the format is\n> \n>  - first line is exactly 46 bytes: five bytes of \"tree \", 40 bytes of hex \n>    sha1, and one byte of \"\\n\".\n> \n>    NOTHING ELSE. Not extra spaces at the end, not extra spaces at the \n>    beginning or the middle. It's ASCII, but it's not free-format ASCII.\n> \n>  - the next <n> (where 'n' can be 0 or more) lines are _exactly_ 48 bytes\n>    each:  seven bytes of \"parent \", 40 bytes of hex sha1, and one byte of \n>    \"\\n\".\n> \n>    NOTHING ELSE.\n> \n>  - the next lines are \"author \" and \"committer \". They have well-defined \n>    delimters for their fields, and no sha1's. The fields cannot contain \n>    '<', '>' or newlines, since those are the field/line delimeters.\n> \n> There is no free-format text _anywhere_ that git parses. No room for \n> guesses, no room for mistakes, no room for anything half-way questionable.\n> \n> And fsck actually enforces this. We do _not_ just use \"gets()\" to read one \n> line at a time. We literally verify that the lines are 46/48 bytes long, \n> and have the delimeters in the expected places.\n> \n> Same goes for \"tree\" and \"tag\" objects. They all have fixed-format stuff. \n> A \"tree\" entry is always\n> \n> \t\"%o <space> %s\" \\0 [ 20 bytes of sha1 ]\n> \n> with \"%o\" being \"mode\", and \"%s\" being \"path\". We don't guess. \n> \n> And this really is _important_. Exactly because we name things by the SHA1\n> hash of the contents, we MUST NOT have flexible formats. Having a format\n> which allows non-canonical representations (extra spaces etc) would mean\n> that two trees that were identical would depend on how you happened to\n> format them.\n> \n> So there's really two issues:\n>  - we don't guess or parse contents. We have strict rules, and that makes \n>    git more reliable. There are no gray areas. There's \"right\" and there \n>    is \"wrong\", and the right one works, and the wrong one gets flagged as \n>    being wrong and the tools refuse to touch it.\n>  - there is only _one_ right way to do things, and that means that the \n>    the content is well-defined, and thus the SHA1 of the content is \n>    well-defined.\n> \n> For example, another rule is that a \"tree\" object is always sorted by \n> the bytes in the filename (not by entry, btw: a directory called \"foo\" \n> will sort as \"foo/\", even though the _entry_ only shows \"foo\"). That rule \n> not only makes a lot of operations faster, but again, it means that there \n> is only _one_ way to represent a tree validly.\n> \n> IOW, you _cannot_ represent a tree any other way (and I've been too lazy\n> to check this in fsck, but it's alway sbeen my plan), and that is exactly \n> why we can just compare the hashes of the results - because there is no \n> random component of \"layout\" in the contents.\n> \n> This really is important. It means that if you get to the same two tree\n> contents in totally unrelated ways (you unpack a tar-file and encode it in\n> git, or you have 5 years of git history and check it out), the \"tree\" will\n> match _exactly_. There's no history. There's no \"optional\" stuff. Since\n> the contents of the trees are the same, the SHA1 of the two trees will be\n> the same. Exactly because git refuses to touch any free-format stuff.\n\n"},{"id":"1900","messageId":"426FF8C4.8080809@zytor.com","threadId":"348","inReplyTo":"Pine.LNX.4.58.0504271154470.18901@ppc970.osdl.org","subject":"Re: A shortcoming of the git repo format","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-27T20:40:36Z","receivedAt":"2005-04-27T20:40:36Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> No, that's definitely _not_ the point.\n> \n> I repeat: git does not do any free-form parsin AT ALL. The links are in \n> well-defined places, and you do not ever search for them. And that's \n> really very very important.\n> \n\nI know that.  However, is that going to be true for all versions of the \nrepository format over all time?  If so, the repository format is brittle.\n\n > > Currently there is no  such delimiter for that.\n >\n > There absolutely is.\n >\n > For a \"commit\", the format is...\n\nMy point was that with a syntactic delimiter, one can write a tool that \ndoesn't necessarily know everything about every tag, including future \ntags which may not have been invented when the tool was written.\n\nOne can simply say \"we don't do that\"; finding an unknown tag is always \na fatal error.  That means the format is more brittle, but brittle does \nmean it breaks as opposed to getting deformed in some, potentially \nundesirable way.\n\n\t-hpa\n"},{"id":"1903","messageId":"200504272049.NAA14598@emf.net","threadId":"348","inReplyTo":"426FF8C4.8080809@zytor.com","subject":"Re: A shortcoming of the git repo format","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-27T20:49:52Z","receivedAt":"2005-04-27T20:49:52Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n   From: \"H. Peter Anvin\" <hpa@zytor.com>\n\n   Linus Torvalds wrote:\n   > \n   > No, that's definitely _not_ the point.\n   > \n   > I repeat: git does not do any free-form parsin AT ALL. The links are in \n   > well-defined places, and you do not ever search for them. And that's \n   > really very very important.\n   > \n\n   I know that.  However, is that going to be true for all versions of the \n   repository format over all time?  If so, the repository format is brittle.\n\nI think one has to understand Linus' posts as coming from the\n\"head-down, steaming ahead for *MY* project cause you all suck\"\nperspective and impose corresponding filters on his declarations of\n\"LAW\".  At least that's the only way *I* can make sense of his latest\ncontributions.\n\nIf you get git, just do the right thing -- Linus be damned.\n\n-t\n\n"},{"id":"1904","messageId":"Pine.LNX.4.58.0504271352110.18901@ppc970.osdl.org","threadId":"348","inReplyTo":"426FF8C4.8080809@zytor.com","subject":"Re: A shortcoming of the git repo format","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-27T20:56:49Z","receivedAt":"2005-04-27T20:56:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 27 Apr 2005, H. Peter Anvin wrote:\n> \n> I know that.  However, is that going to be true for all versions of the \n> repository format over all time?  If so, the repository format is brittle.\n\nI agree, it's brittle by design, exactly because I think it's very \nimportant not to allow any variations.\n\nHOWEVER, that's where \"convert-cache\" comes in. Any one particular format \nmay be brittle, but if we accept that, and just say \"we can upgrade by \nconverting the cache\", then we should be ok. IOW, we can change from one \nbrittle format with 160-bit SHA1 names to _another_ brittle format with \n256-bit SHA1 (or other) names.\n\n> My point was that with a syntactic delimiter, one can write a tool that \n> doesn't necessarily know everything about every tag, including future \n> tags which may not have been invented when the tool was written.\n\nNow, I kind of agree with that, but not on a \"object level\".\n\nBut exactly because the object level is \"brittle by design\", and because I \nthe way to fix that is convert-cache (which may do _big_ changes to the \nformat), I really don't think that the objects should ever be looked at \nexcept with very precise tools.\n\nBut when it comes to \"higher-level information\", I agree with you 100%.\n\nFor example, this _is_ actually why I wanted pasky to change the format of \n\"git log\" (now cg-log). Exactly so that the output of that isn't brittle, \nit now prepends spaces to the free-form part.\n\n\t\tLinus\n"},{"id":"1906","messageId":"20050427205812.GA4412@frodo","threadId":"348","inReplyTo":"426F2671.1080105@zytor.com","subject":"Re: A shortcoming of the git repo format","fromName":"Gerhard Schrenk","fromEmail":"gps@mittelerde.physik.uni-konstanz.de","sentAt":"2005-04-27T20:58:12Z","receivedAt":"2005-04-27T20:58:12Z","isPatch":false,"sender":{"key":"gps@mittelerde.physik.uni-konstanz.de","avatar":null},"body":"* H. Peter Anvin <hpa@zytor.com> [2005-04-27 07:43]:\n> Most of git's files are starting to converge toward an RFC822-like \n> header with (tag, data) and a free-form section.  This is a good\n> thing.\n\nI really hate RFC822-like data structures. Why? Lazy straightforward\npeople (who have written to much mails) tend to break the relational\ndata\nmodell and don't realize what they loose. Usually they introduce\nnon-atomar tags like\n\nTag: value1, value2\n\nand game over. You have just broken the first normal form (1NF). In the \nend the relational normalization process is just not to break the\nfunctional dependencies of your data. It's worth it.\n\nI'm reacting like pawlov's dog and really don't know what I'm talking\nabout (namely git). But please don't do the same error and just\nassociate\nrelational = sql = crap. The shell's operator stream paradigma fits very\ngood to the relational modell. It's certainly closer to the relational\nalgebra than sql...\n\nTake care\nGerhard\n\n"},{"id":"1907","messageId":"426FFD27.4030604@zytor.com","threadId":"348","inReplyTo":"200504272049.NAA14598@emf.net","subject":"Re: A shortcoming of the git repo format","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-27T20:59:19Z","receivedAt":"2005-04-27T20:59:19Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Tom Lord wrote:\n> \n> I think one has to understand Linus' posts as coming from the\n> \"head-down, steaming ahead for *MY* project cause you all suck\"\n> perspective and impose corresponding filters on his declarations of\n> \"LAW\".  At least that's the only way *I* can make sense of his latest\n> contributions.\n> \n> If you get git, just do the right thing -- Linus be damned.\n> \n\nIt's fair for Linus to want to make things behave a certain way in a \nproject.  There are design decisions which have tradeoffs both ways -- \nrobust (but subject to partial information issues) versus brittle (but \nsafe.)\n\nThat's part of why I prefer to ask first.\n\n\t-hpa\n"},{"id":"1925","messageId":"2cfc403205042715513d8123f3@mail.gmail.com","threadId":"348","inReplyTo":"426FDE48.1050700@zytor.com","subject":"Re: A shortcoming of the git repo format","fromName":"Jon Seymour","fromEmail":"jon.seymour@gmail.com","sentAt":"2005-04-27T22:51:15Z","receivedAt":"2005-04-27T22:51:15Z","isPatch":false,"sender":{"key":"jon.seymour@gmail.com","avatar":"https://avatars.githubusercontent.com/u/207131?v=4"},"body":"On 4/28/05, H. Peter Anvin <hpa@zytor.com> wrote:\n> Dave Jones wrote:\n> >\n> > That actually broke one of my first git scripts when one of the\n> > changelog texts started a line with 'tree '.  I hacked around it\n> > by making my script only grep in the 'head -n4' lines, but this\n> > seems somewhat fragile having to make assumptions that the field\n> > I want to see is in the first 4 lines.\n> >\n> \n> You have the delimiter for that; there is an empty line between the\n> header and the free-form body, similar as for RFC822.\n> \n\n...and a relatively simple way to use that rule to extract just the\nheader lines:\n\n      sed -n \"1,/^\\$/p\"                     # with the separator line\n\nor, either one of these to remove the separator line as well:\n\n      sed -n \"1,/^\\$/s/^\\(..*\\)/\\1/p\"  \n      sed -n \"1,/^\\$/p\" | tr -s \\\\012\n\njon\n-- \nhomepage: http://www.zeta.org.au/~jon/\nblog: http://orwelliantremors.blogspot.com/\n"},{"id":"1937","messageId":"Pine.LNX.4.21.0504271939100.30848-100000@iabervon.org","threadId":"348","inReplyTo":"426FF8C4.8080809@zytor.com","subject":"Re: A shortcoming of the git repo format","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-04-27T23:50:13Z","receivedAt":"2005-04-27T23:50:13Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Wed, 27 Apr 2005, H. Peter Anvin wrote:\n\n> One can simply say \"we don't do that\"; finding an unknown tag is always \n> a fatal error.  That means the format is more brittle, but brittle does \n> mean it breaks as opposed to getting deformed in some, potentially \n> undesirable way.\n\nIf you find an object with an unknown tag, you can't do much with it\nanyway, even if it has a format that matches generic rules. Sure, you\ncould trace reachability through it, but that's only helpful for a couple\nof generic programs (fsck and pull), and those programs ought to\nadditionally have some clue about what's going on if they're going to act\nappropriately.\n\nOn the other hand, it is probably true that programs should be able to\ndeal abstractly with new tags if built with a libgit that supports them,\nbut that's something that we can arrange a bit later.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"1940","messageId":"427026AB.4070809@zytor.com","threadId":"348","inReplyTo":"Pine.LNX.4.21.0504271939100.30848-100000@iabervon.org","subject":"Re: A shortcoming of the git repo format","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-27T23:56:27Z","receivedAt":"2005-04-27T23:56:27Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Daniel Barkalow wrote:\n> \n> If you find an object with an unknown tag, you can't do much with it\n> anyway, even if it has a format that matches generic rules. Sure, you\n> could trace reachability through it, but that's only helpful for a couple\n> of generic programs (fsck and pull), and those programs ought to\n> additionally have some clue about what's going on if they're going to act\n> appropriately.\n> \n> On the other hand, it is probably true that programs should be able to\n> deal abstractly with new tags if built with a libgit that supports them,\n> but that's something that we can arrange a bit later.\n> \n> \t-Daniel\n\nThere are a fair number of tools one may want that deal with reachability.\n\n\t-hpa\n\n"},{"id":"1949","messageId":"4270320D.5090708@dwheeler.com","threadId":"348","inReplyTo":"Pine.LNX.4.58.0504271352110.18901@ppc970.osdl.org","subject":"Re: A shortcoming of the git repo format","fromName":"David A. Wheeler","fromEmail":"dwheeler@dwheeler.com","sentAt":"2005-04-28T00:45:01Z","receivedAt":"2005-04-28T00:45:01Z","isPatch":false,"sender":{"key":"dwheeler@dwheeler.com","avatar":"https://avatars.githubusercontent.com/u/813150?v=4"},"body":"Linus Torvalds wrote:\n> \n> On Wed, 27 Apr 2005, H. Peter Anvin wrote:\n> \n>>I know that.  However, is that going to be true for all versions of the \n>>repository format over all time?  If so, the repository format is brittle.\n> \n> I agree, it's brittle by design, exactly because I think it's very \n> important not to allow any variations.\n\nIn the short term, not allowing any variations is probably a\ngood thing, it'll winnow out mistakes.  Creating a format that\nCOULD change in the future is, however, a very good way of avoiding\ngetting boxed into a corner if it turns out a mistake has been made.\n\n> HOWEVER, that's where \"convert-cache\" comes in. Any one particular format \n> may be brittle, but if we accept that, and just say \"we can upgrade by \n> converting the cache\", then we should be ok. IOW, we can change from one \n> brittle format with 160-bit SHA1 names to _another_ brittle format with \n> 256-bit SHA1 (or other) names.\n\nThere's a disadvantage to that, unfortunately: invalidating signatures.\nYes, you can get people to re-sign their stuff... assuming you can\nfind them & convince them to do it (ha!).  More than likely,\nyou'll lose signatures that way.  Probably not your TOP priority,\nbut there are advantages to being able to go back & years later\nSHOW that someone really did sign something.\n\nIn the long run, I'd really like to see (at least) signed commits,\nand that those signatures would \"stick around\" cleanly into the future.\n\"Breaks\" can be handled other ways, but it is DEFINITELY a pain,\nand an avoidable one.\n\n--- David A. Wheeler\n"},{"id":"1952","messageId":"Pine.LNX.4.62.0504271743580.4990@qynat.qvtvafvgr.pbz","threadId":"348","inReplyTo":"4270320D.5090708@dwheeler.com","subject":"Re: A shortcoming of the git repo format","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2005-04-28T00:46:54Z","receivedAt":"2005-04-28T00:46:54Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Wed, 27 Apr 2005, David A. Wheeler wrote:\n\n> Linus Torvalds wrote:\n>> \n>> On Wed, 27 Apr 2005, H. Peter Anvin wrote:\n>> \n>>> I know that.  However, is that going to be true for all versions of the \n>>> repository format over all time?  If so, the repository format is brittle.\n<<SNIP>> \n>> HOWEVER, that's where \"convert-cache\" comes in. Any one particular format \n>> may be brittle, but if we accept that, and just say \"we can upgrade by \n>> converting the cache\", then we should be ok. IOW, we can change from one \n>> brittle format with 160-bit SHA1 names to _another_ brittle format with \n>> 256-bit SHA1 (or other) names.\n>\n> There's a disadvantage to that, unfortunately: invalidating signatures.\n> Yes, you can get people to re-sign their stuff... assuming you can\n> find them & convince them to do it (ha!).  More than likely,\n> you'll lose signatures that way.  Probably not your TOP priority,\n> but there are advantages to being able to go back & years later\n> SHOW that someone really did sign something.\n\nall you have to do is to make sure that convert-cache doesn't loose any \ndata and you can always convert back (through as many steps as needed) to \ncheck signatures.\n\nno matter what you do, if you change the thing that's being signed the \nsignature is worthless, it doesn't matter if you change it in a flexible \nor a brittle way, it's different. the brittle approach actually makes it \neasier to go backwards as you KNOW exactly what it needs to be, there's no \npossiblity that a later tag was there, but being ignored (except for the \nsignature)\n\n> In the long run, I'd really like to see (at least) signed commits,\n> and that those signatures would \"stick around\" cleanly into the future.\n> \"Breaks\" can be handled other ways, but it is DEFINITELY a pain,\n> and an avoidable one.\n>\n> --- David A. Wheeler\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n-- \nThere are two ways of constructing a software design. One way is to make it so simple that there are obviously no deficiencies. And the other way is to make it so complicated that there are no obvious deficiencies.\n  -- C.A.R. Hoare\n"},{"id":"1956","messageId":"Pine.LNX.4.58.0504271722260.18901@ppc970.osdl.org","threadId":"348","inReplyTo":"200504272049.NAA14598@emf.net","subject":"Re: A shortcoming of the git repo format","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-28T00:57:07Z","receivedAt":"2005-04-28T00:57:07Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 27 Apr 2005, Tom Lord wrote:\n> \n> I think one has to understand Linus' posts as coming from the\n> \"head-down, steaming ahead for *MY* project cause you all suck\"\n> perspective and impose corresponding filters on his declarations of\n> \"LAW\".\n\nI'm really being very head-strong on these things, and much more so than I\nnormally am, because quite frankly, I see \"git\" as a very different\nproject from Linux.\n\n(Which is not to say that I'm not opinionated even normally, but I'm \nnormally a bit more open to listen to other people ;)\n\nThere's two huge differences between git and Linux, and I'm really sorry\nif they make me act as an asshole, but they are important to me:\n\n - with Linux, lots of people know what the \"right thing\" is, because the \n   UNIX mindset has really been a kind of \"social background\" that has \n   been around for long enough that it has institutionalized knowledge\n   about what an OS is supposed to do.\n\n   This means that in 99% of all technical discussions about the kernel, \n   people are already coming at the problem roughly from the same \n   stand-point. It's not _universally_ true, but I really think that the \n   institutionalized (but not always conscious) philosophy of UNIX is what \n   has made it a lot easier to talk about almost all kernel issues,\n   because people have generally the same expectations of what is \"good\".\n\n   Doing development is a lot about communication. Writing code in many \n   ways is secondary - it's much more important to try to make sure that \n   everybody knows what the goals are, because the _real_ pain in \n   development ends up being not the coding, but the much more fundamental \n   disagreements that happen when people really have totally different \n   expectations of what the end result is going to be.\n\n   SCM's don't have this. Quite the reverse. I see 30 years of \"CVS\" being \n   the common language for a lot of people, and the fact is, most of the\n   people on this mailing list probably never _really_ used BK, and do not\n   really understand very deeply about how the distributed model actually \n   ends up workign in _practice_.\n\n   I think a lot of people understand it intellectually, but I really do \n   think that we're lackign the kind of \"institutionalized\" knowledge\n   where people understand things at a much more visceral level.\n\n - With Linux, I never had something I needed to get _done_. Even when I \n   started, it was just for fun, and by the time others joined in, the \n   system already did much more than I initially envisioned, so everything \n   was really \"gravy\".\n\n   With git, this isn't the case. The _only_ reason I started git in the \n   first place is that I knew better than pretty much anybody else what my\n   needs were, and I was forced to act on them because nothing out there \n   really solved the problem for me.\n\nIn other words: I _know_ that I've been unpleasant. I'm sorry about that, \nbut I am trying to explain _why_ I'm being an asshole about things, more \nso than I usually am.\n\nI'm not actually all that interested in SCM's. I'd have been much happier\nif I never had to start doing git in the first place. But circumstances\nnot only forced me to do my own, it also so happens that I don't believe\nthat there are many people around that have ever really _seen_ what my\nkind of development requirements are.\n\nWhat does that boil down to? It means, for example, that to me it doesn't\nmatter one _whit_ if you've been doing SCM's for the last thirty years,\nand you can do xdelta algorithms in your sleep.\n\nQuite the reverse: such a person \"knows\" a lot of things, but I'm pretty\ndamn sure that such a person has _never_ actually worked on a system that\nworks the way the kernel development does, which means that most of the\nthings that person \"knows\" are things that may need to be un-learnt.\n\nAnd because I don't actually _care_ about SCM's, and only care about\ngetting to the point where I (once more) don't have to even think about\nthe SCM that I use for the kernel, I also don't have much incentive to\nworry about CM models that may well be very valid outside of kernel work.\n\nSee? When it comes to my Linux work, I'm very inclusive. Linux already\ndoes everything _I_ need it to do, so in many ways, all that really\nmotivates me to improve it are really about other peoples needs, and as\nsuch, I'm really really interested in what _other_ people want. I still\nsay \"no, that's now how we do things\", but that's much less contentious.\n\nIn contrast, with git, I'm totally uninterested in anything that doesn't\nmake my kernel work go faster or more smoothly, and does so _today_. Which \nmakes me a cantancerous old bastard, and bit the heads off anybody who \nisn't focused on that one thing.\n\nAnd I really _am_ sorry. I don't actually _like_ being nasty about these \nthings. But when it comes to git, I have one motivation, and one \nmotivation only, and being nice about it isn't going to help. \n\nThe good news? I actually think my needs are very basic. Once gits gets to \nthe point where it does what I need it to do, I don't really have any \nmotivation to say \"this is how we do it\" any more. And I think we're \nactually getting to that point fairly soon. That's not saying git is \n\"done\", any less than Linux was \"done\" in 1992. It's just that at that \npoint I don't have any reason to be a nasty control freak any more.\n\nIn fact, I don't see myself even maintaining the project, especially since\nthere seem to be others that are more motivated to do so than I am. Then\nI'll just go back into my dark kernel cave, and hopefully I don't have to\ncome out again for a while.\n\nBut for now, the _only_ point of git is as a kernel maintenance tool. \nThere are tons of other SCM systems that are probably better for other \nprojects, so if git is \"just another SCM project\", then git is totally \npointless. So for now, the absolutely _only_ thing that matters for git \ndesign (as far as I'm concerned) is \"how well does it suit Linus\".\n\n\t\t\tLinus\n"},{"id":"1967","messageId":"20050427183421.172b4d48.pj@sgi.com","threadId":"348","inReplyTo":"Pine.LNX.4.58.0504271722260.18901@ppc970.osdl.org","subject":"Re: A shortcoming of the git repo format","fromName":"Paul Jackson","fromEmail":"pj@sgi.com","sentAt":"2005-04-28T01:34:21Z","receivedAt":"2005-04-28T01:34:21Z","isPatch":false,"sender":{"key":"pj@sgi.com","avatar":null},"body":"Dang ... don't apologize too much ... it's fun watching Linus be a\ncranky git.\n\nThis is turning into something neat, something different and special,\nand no way we'd have gotten here using the usual ways or means.\n\nAnd we're all pretty damn confident that you won't be playing SCM\ndictator for long - tools are obviously not your first love.\n\nEvery China Shop needs a good Bull now and then.\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"1969","messageId":"Pine.LNX.4.21.0504272143260.30848-100000@iabervon.org","threadId":"348","inReplyTo":"427026AB.4070809@zytor.com","subject":"Re: A shortcoming of the git repo format","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-04-28T01:51:06Z","receivedAt":"2005-04-28T01:51:06Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Wed, 27 Apr 2005, H. Peter Anvin wrote:\n\n> There are a fair number of tools one may want that deal with reachability.\n\nDo you agree that installing a new libgit.so when you want to apply such a\ntool to a new tag is sufficient? If the library is shared, and everything\nfor parsing the objects (to the point of getting struct object filled\nout) is in the library, and you want to have some tool able to validate or\nuse any new tag that you want reachability-only tools to process, not\nhaving a standard header proto-format for future tags isn't a problem,\nsince you'll get upgrades to the parser portion of all of your tools\ntogether.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n\n"},{"id":"1971","messageId":"427042EB.8030501@zytor.com","threadId":"348","inReplyTo":"Pine.LNX.4.21.0504272143260.30848-100000@iabervon.org","subject":"Re: A shortcoming of the git repo format","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-28T01:56:59Z","receivedAt":"2005-04-28T01:56:59Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Daniel Barkalow wrote:\n> On Wed, 27 Apr 2005, H. Peter Anvin wrote:\n>  \n>>There are a fair number of tools one may want that deal with reachability.\n>  \n> Do you agree that installing a new libgit.so when you want to apply such a\n> tool to a new tag is sufficient? If the library is shared, and everything\n> for parsing the objects (to the point of getting struct object filled\n> out) is in the library, and you want to have some tool able to validate or\n> use any new tag that you want reachability-only tools to process, not\n> having a standard header proto-format for future tags isn't a problem,\n> since you'll get upgrades to the parser portion of all of your tools\n> together.\n> \n\nOnly if language bindings are created for this library.\n\n\t-hpa\n"},{"id":"1976","messageId":"200504280214.TAA19991@emf.net","threadId":"348","inReplyTo":"Pine.LNX.4.58.0504271722260.18901@ppc970.osdl.org","subject":"Re: A shortcoming of the git repo format","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-28T02:14:23Z","receivedAt":"2005-04-28T02:14:23Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n   > I think a lot of people understand it intellectually, but I really do \n   > think that we're lackign the kind of \"institutionalized\" knowledge\n   > where people understand things at a much more visceral level.\n\nI know that Arch and its progeny, as they stand, don't seduce you\nbut you should be made aware that the Arch community is one where\ngood SCM sense that you would agree with (although you might not\nrecognize it at once) is well on the path to being institutionalized.\nIt's gratifying/amazing/inspiring to see a bunch of folk catch up \non the topic.\n\nOne thing there's still a shortage of in my world is folks steeped\nin both perspectives: \"unix\" /and/ SCM.  Thus, I get folks who have\npretty decent SCM ideas in the abstract -- plus utterly terrible \nideas about how to make them real.\n\nThere is a higher-level bug I think you'll eventually viscerally \nfeel yourself, related to:\n\n   > I think a lot of people understand it intellectually, but I really do \n   > think that we're lackign the kind of \"institutionalized\" knowledge\n   > where people understand things at a much more visceral level.\n\nOnce you get to the BK or Arch level of SCM, beyond that there are\nmany possible paths.  Many of those are false paths -- imaginary\n(unrealizable) ideals about how things like merging can work and\nbe good.   Some people seem to get stuck on those paths.\n\n   > With git, this isn't the case. The _only_ reason I started git in the \n   > first place is that I knew better than pretty much anybody else what my\n   > needs were, and I was forced to act on them because nothing out there \n   > really solved the problem for me.\n\nThat's debatable but neither here nor there.  Supposing that Arch\nwere /perfect/ for your needs today (which I don't claim) -- `git'\nwould still have been the better route to take (though my reasons\nprobably aren't the same as yours).\n\n   > I'm not actually all that interested in SCM's.\n\nIn a certain way: same here, oddly enough.  Go figure.\n\n   > Quite the reverse: such a person \"knows\" a lot of things, but I'm pretty\n   > damn sure that such a person has _never_ actually worked on a system that\n   > works the way the kernel development does\n\nI've been avoiding the topic of how kernel development works ever since\ni realized, that with each additional detail you reveal, i have little\nbut yellow and red cards to raise.   Doesn't seem productive to have that\nfight when the option of simply improving the situation is open.\n\n   > And I really _am_ sorry. I don't actually _like_ being nasty about these \n   > things.\n\nIt's healthy enough that you are, for your sanity and others.  Just \nbe tolerant of others pointing that out.\n\n   > The good news? I actually think my needs are very basic.\n\nSo it would seem.  This is partly because the process you advertise\nyourself as doing is, sorry, garbage.  It's understandable why it\nhappens to work for now, but it's garbage nonetheless.  Not your fault --\nyou haven't been afforded the degrees of freedom to do better, afaict.\n\n   > But for now, the _only_ point of git is as a kernel maintenance tool. \n\nMath is math.  You don't get to say what it means.\n\n-t\n\n\n"},{"id":"1983","messageId":"20050428033737.GA30308@mythryan2.michonline.com","threadId":"348","inReplyTo":"Pine.LNX.4.58.0504271722260.18901@ppc970.osdl.org","subject":"Re: A shortcoming of the git repo format","fromName":"Ryan Anderson","fromEmail":"ryan@michonline.com","sentAt":"2005-04-28T03:37:37Z","receivedAt":"2005-04-28T03:37:37Z","isPatch":false,"sender":{"key":"ryan@michonline.com","avatar":null},"body":"On Wed, Apr 27, 2005 at 05:57:07PM -0700, Linus Torvalds wrote:\n> On Wed, 27 Apr 2005, Tom Lord wrote:\n> \n> I'm not actually all that interested in SCM's. I'd have been much happier\n> if I never had to start doing git in the first place. But circumstances\n> not only forced me to do my own, it also so happens that I don't believe\n> that there are many people around that have ever really _seen_ what my\n> kind of development requirements are.\n\nOddly, I was trying to answer \"Why distributed?\" in a discussion the\n\"Joel On Software\" forum.\n\nThe particular thread I posted on, well, was kind of stupid, but in case\nanyone is curious: http://discuss.joelonsoftware.com/default.asp?joel.3.115346.51\n\nWhat I said might help give an overview of how Linux development works,\nfrom my point of view.  I only occassionally poke at interesting things\non the periphery on whims, but I poke at the SCMy aspects of it, so\nmaybe it's relevant. \n\n Here's an overview of how the distributed world of Linux works:\n\n 1. Linus has his personal tree.  He pushes it out on a regular basis to\n rsync.kernel.org (well, kinda - that's where it ends up at).\n\n 2. \"Trusted lieutenants\" have their own trees.  Some keep these on\n *.kernel.org, some don't.\n\n 3. Lots of other people have personal trees.  These can be pretty much\n anywhere.\n\n These trees are in a variety of formats today, some are in \"git\", some\n are still in BitKeeper, some are from a tarball, some are tarball +\n patches, some are git + patches.\n\n There are a variety of merging methods:\n\n a.  Provide a publicly accessible repository.  (Formerly BK, now \"git\")\n that Linus, or a maintainer (i.e, \"trusted lieutenant\") can grab it\n from.  In the email where this location is given, the patch is usually\n included, at least in a summary format.\n\n b.  Provide a series of emails, with a description per email followed,\n inline, with a patch.\n\n These merging methods can be done directly with Linus, or with anyone\n else who is interested.  (Generally, merging with Linus is for arch and\n subsystem maintainers, or random small things that are either obviously\n correct, useful, or just don't fit elsewhere.)\n\n So, that's the merge process, for the most part.\n\n Now, most patches these days are going through Andrew Morton - even if\n he's not actually submitting them personally, he's probably putting\n them into his tree for testing purposes.  (Networking changes go direct\n to Linus, but Andrew keeps an up to date version of them in his -mm\n series of kernels.)\n\n If code isn't accepted, well, one of a couple things happens:\n 1. The patch is silently ignored.  (This is less of a problem these\n days.)\n\n 2. The patch is commented on and someone says, \"No\".  (Generally, this\n happens a few times for \"new\" code, as people try to get the concept to\n fit into the kernel in the cleanest way.  There are a lot of style nits\n at this point, but also discussions of \"Is this the right way to do\n this?\" and \"Do we need a more general method to do this instead of this\n hack?\")\n\n Verifying that testing has occurred is less important than you might\n think.  This is basically because small patches either come with a\n description of the bug they fix and an expert in that area will ACK the\n patch, they touch an area that few people use and so the submitter is\n probably the best qualified person to provide a patch and they'll only\n hurt themselves if they haven't tested it, or, via the history of your\n submissions to the kernel, you are known to not submit bad code, so\n there's an expectation of quality.\n\n Furthermore, an incredible amount of testing occurs in the major public\n trees (Linus/-mm) between a release, so most absolutely major bugs are\n spotted fairly quickly, and if the problem is systemic in a change,\n that change can be reverted until the code improves.\n\n On the topic of checking into private branches - it's not so much a\n matter of \"the parent never sees the changes\" as \"the parent doesn't\n see them right now\".\n\n FWIW, at my place of employment, we switched from CVS to BitKeeper last\n summer, and it is significantly more pleasant to work with, in all\n aspects.\n\n Currently our entire development staff is working from home.  This\n still works well, as we can all check in locally, and submit changes to\n the master repository when changes are ready.  Between having a partner\n company in Japan working on our code, and our development staff working\n from home offices, we would have a horrific time getting any\n centralized SCM product to perform well.  With purely local\n repositories, local branching, and submissions via email or ssh, the\n process still works well and is *fast*.  CVS over slow network links is\n certainly not *fast*, and I'd be very surprised if Perforce is\n significantly better in that regard.\n\n I'll just say this, in closing - working with a decentralized SCM tool\n changes the way you work.  There is a Linux Kernel developer that I am\n aware of that keeps 27 or so seperate branches on his machine, so he\n can keep all the logically unrelated changes seperate from each other.\n He builds kernels off an additional branch that merges all the others\n together, and submits changes to Linus via 2 or 3 \"rollup\" trees he\n maintains.\n\n You just don't work like that in a centralized SCM, because branching\n isn't painless, in the same way.\n\n-- \n\nRyan Anderson\n  sometimes Pug Majere\n"},{"id":"2018","messageId":"b8464fde050428013144672950@mail.gmail.com","threadId":"348","inReplyTo":"Pine.LNX.4.58.0504271722260.18901@ppc970.osdl.org","subject":"Re: A shortcoming of the git repo format","fromName":"Morgan Schweers","fromEmail":"mschweers@gmail.com","sentAt":"2005-04-28T08:31:12Z","receivedAt":"2005-04-28T08:31:12Z","isPatch":false,"sender":{"key":"mschweers@gmail.com","avatar":null},"body":"Greetings,\n\nThis is off topic, but this is a great paragraph, and an incredibly\nconcise and valuable lesson for pre-architect software developers.\n\nOn 4/27/05, Linus Torvalds <torvalds@osdl.org> wrote:\n\n[...deletia...]\n\n>    Doing development is a lot about communication. Writing code in many\n>    ways is secondary - it's much more important to try to make sure that\n>    everybody knows what the goals are, because the _real_ pain in\n>    development ends up being not the coding, but the much more fundamental\n>    disagreements that happen when people really have totally different\n>    expectations of what the end result is going to be.\n\n[...deletia...]\n\n>                         Linus\n\n--  Morgan Schweers\n"},{"id":"2035","messageId":"1114695600.27227.123.camel@hades.cambridge.redhat.com","threadId":"348","inReplyTo":"426FD3EE.5000404@zytor.com","subject":"Re: A shortcoming of the git repo format","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-28T13:39:59Z","receivedAt":"2005-04-28T13:39:59Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Wed, 2005-04-27 at 11:03 -0700, H. Peter Anvin wrote:\n> > To find the email address, look for the first '<'. To find the date, look \n> > for the first '>'. Those characters are not allowed in the name or the \n> > email, so they act as well-defined delimeters.\n> > \n> \n> That's true for email addresses,\n\nNot in general. You can have just about any character, including @, <\nand >, in either a display-name or a local-part.\n\nFor git we actually _remove_ any instances of '<' and '>' from both\n'AUTHOR_NAME' and 'AUTHOR_EMAIL', so what you say becomes true.\n\nI still say these shouldn't be considered email addresses, any more than\nthe 'user@host.domain' you see when you connect to an IRC server is\nconsidered an IP address.\n\n-- \ndwmw2\n\n"},{"id":"2040","messageId":"IGEMLBGAECDFPIKMIMLCEEGECHAA.barry@disus.com","threadId":"348","inReplyTo":"Pine.LNX.4.58.0504271722260.18901@ppc970.osdl.org","subject":"RE: A shortcoming of the git repo format","fromName":"Barry Silverman","fromEmail":"barry@disus.com","sentAt":"2005-04-28T15:08:26Z","receivedAt":"2005-04-28T15:08:26Z","isPatch":false,"sender":{"key":"barry@disus.com","avatar":null},"body":">>In contrast, with git, I'm totally uninterested in anything that doesn't\n>>make my kernel work go faster or more smoothly, and does so _today_. Which\n>>makes me a cantancerous old bastard, and bit the heads off anybody who\n>>isn't focused on that one thing.\n\nFocus is the totally operative word here!\n\nIf you really want to feel good about the world, re-read the initial set of\ngit postings that Linus made on April 7th:\nhttp://kerneltrap.org/node/4982\n\nContrast the picture today with the fact that three weeks ago:\nApril 7:\n1) the kernel workflow was at a standstill\n2) git was just a totally unproven concept in Linus' head, that could have\nended up as a band-aid while a REAL SCM (...sound of choking from the\nwings...) was chosen\n3) the performance issues in dealing with both the size of the kernel\nproject, and the velocity of the changes were completely up in the air\n\nToday:\n1) the kernel workflow has restarted, and has already made its first\nmilestone\n2) git is solid in architecture, is maintained and updated by a proven set\nof developers, and has been demonstrated to have all the performance\nnecessary going forward\n3) the primary traffic on the mailing list is related to tactical issues -\nnot architecture, or strategy, or big-ticket item stuff - with the\noccasional flame about \"renames\" ;-)\n\nAre there any large strategic issues left to be resolved for git?, or is it\njust a matter of getting all the kernel developers over the learning curve,\nand iterating the details of the workflow to make everyone maximally\nproductive?\n\nHow long do you think it will take for the kernel workflow to get back to\nits height during the BK days?\n\nThe achievement of going from a complete standstill, to full velocity kernel\nworkflow production in a couple of months has got to be something everyone\ninvolved should be intensely proud of.\nThanks, Linus, for being such a \"cantancerous old bastard\". I don't think it\ncould have happened if you were anything but....\n\nBarry Silverman\n\n"}]}