{"thread":{"id":"19010","subject":"[doc] User Manual Suggestion","startedAt":"2009-04-22T19:38:52Z","lastAt":"2009-05-03T01:48:26Z","messageCount":90,"participants":["David Abrahams","J. Bruce Fields","Michael Witten","Jeff King","Johan Herland","Daniel Barkalow","Björn Steinbrink","Felipe Contreras","Mark Lodato"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"111988","messageId":"m24owgqy0j.fsf@boostpro.com","threadId":"19010","inReplyTo":null,"subject":"[doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-22T19:38:52Z","receivedAt":"2009-04-22T19:38:52Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nhttp://www.kernel.org/pub/software/scm/git/docs/user-manual.html#how-to-check-out\ncovers \"git reset\" way too early, IMO, before one has the conceptual\nfoundation necessary to understand what it means to \"modify the current\nbranch to point at v2.6.17\".  If this operation must be covered this\nearly in the manual, it should probably not be until\nhttp://www.kernel.org/pub/software/scm/git/docs/user-manual.html#manipulating-branches\n\nHTH,\n\n-- \nDave Abrahams\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"112100","messageId":"20090423175717.GA30198@fieldses.org","threadId":"19010","inReplyTo":"m24owgqy0j.fsf@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2009-04-23T17:57:17Z","receivedAt":"2009-04-23T17:57:17Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Wed, Apr 22, 2009 at 03:38:52PM -0400, David Abrahams wrote:\n> \n> http://www.kernel.org/pub/software/scm/git/docs/user-manual.html#how-to-check-out\n> covers \"git reset\" way too early, IMO, before one has the conceptual\n> foundation necessary to understand what it means to \"modify the current\n> branch to point at v2.6.17\".  If this operation must be covered this\n> early in the manual, it should probably not be until\n> http://www.kernel.org/pub/software/scm/git/docs/user-manual.html#manipulating-branches\n\nI agree; we should suggest just a git-checkout (to a detached HEAD)\ninstead, though that needs a little explanation so people aren't scared\nby the warning message it gives.\n\nI also have a longstanding todo to experiment with rewriting the\nbeginning to use detached heads more and defer branch management till\nlater.\n\n--b.\n"},{"id":"112109","messageId":"b4087cc50904231137g67b4b84eu3b61bf174ba37d7f@mail.gmail.com","threadId":"19010","inReplyTo":"20090423175717.GA30198@fieldses.org","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-23T18:37:05Z","receivedAt":"2009-04-23T18:37:05Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Thu, Apr 23, 2009 at 12:57, J. Bruce Fields <bfields@fieldses.org> wrote:\n> On Wed, Apr 22, 2009 at 03:38:52PM -0400, David Abrahams wrote:\n>>\n>> http://www.kernel.org/pub/software/scm/git/docs/user-manual.html#how-to-check-out\n>> covers \"git reset\" way too early, IMO, before one has the conceptual\n>> foundation necessary to understand what it means to \"modify the current\n>> branch to point at v2.6.17\".  If this operation must be covered this\n>> early in the manual, it should probably not be until\n>> http://www.kernel.org/pub/software/scm/git/docs/user-manual.html#manipulating-branches\n>\n> I agree; we should suggest just a git-checkout (to a detached HEAD)\n> instead, though that needs a little explanation so people aren't scared\n> by the warning message it gives.\n\nEveryone talks about \"before one has the conceptual foundation\nnecessary to understand\". Well, here's an idea: The git documentation\nshould start with the concepts!\n\nWhy don't the docs start out defining blobs and trees and the object\ndatabase and references into that database? The reason everything is\nso confusing is that the understanding is brushed under the tutorial\nrug. People need to learn how to think before they can effectively\nlearn to start doing.\n"},{"id":"112127","messageId":"20090423201636.GD3056@coredump.intra.peff.net","threadId":"19010","inReplyTo":"b4087cc50904231137g67b4b84eu3b61bf174ba37d7f@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-23T20:16:36Z","receivedAt":"2009-04-23T20:16:36Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Apr 23, 2009 at 01:37:05PM -0500, Michael Witten wrote:\n\n> Everyone talks about \"before one has the conceptual foundation\n> necessary to understand\". Well, here's an idea: The git documentation\n> should start with the concepts!\n> \n> Why don't the docs start out defining blobs and trees and the object\n> database and references into that database? The reason everything is\n> so confusing is that the understanding is brushed under the tutorial\n> rug. People need to learn how to think before they can effectively\n> learn to start doing.\n\nI agree with you, but not everyone does (and you can find prior debates\nin the list archives). The user-manual is pretty \"top down\". There are\nsome \"bottom-up\" resources available, but I haven't seen one pointed to\nas \"definitive\". I think it might actually be nice for there to be a\nparallel to the user manual that follows the bottom-up approach, and\npeople could read the one that appeals most to them (or if they have a\nlot of time on their hands, read both and hopefully it makes sense in\nthe middle ;) ).\n\nBut we would need somebody to volunteer to write it. I would be happy to\nhelp out, but I'm too short on time at the moment to be the driving\nforce.\n\n-Peff\n"},{"id":"112131","messageId":"b4087cc50904231345x2613308eh640e50f4a2680890@mail.gmail.com","threadId":"19010","inReplyTo":"20090423201636.GD3056@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-23T20:45:46Z","receivedAt":"2009-04-23T20:45:46Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Thu, Apr 23, 2009 at 15:16, Jeff King <peff@peff.net> wrote:\n> On Thu, Apr 23, 2009 at 01:37:05PM -0500, Michael Witten wrote:\n>\n>> Everyone talks about \"before one has the conceptual foundation\n>> necessary to understand\". Well, here's an idea: The git documentation\n>> should start with the concepts!\n>>\n>> Why don't the docs start out defining blobs and trees and the object\n>> database and references into that database? The reason everything is\n>> so confusing is that the understanding is brushed under the tutorial\n>> rug. People need to learn how to think before they can effectively\n>> learn to start doing.\n>\n> I agree with you, but not everyone does (and you can find prior debates\n> in the list archives). The user-manual is pretty \"top down\". There are\n> some \"bottom-up\" resources available, but I haven't seen one pointed to\n> as \"definitive\".I think it might actually be nice for there to be a\n> parallel to the user manual that follows the bottom-up approach, and\n> people could read the one that appeals most to them (or if they have a\n> lot of time on their hands, read both and hopefully it makes sense in\n> the middle ;) ).\n\nI think the main problem, then, is that the tools have a UI that is\nsomewhere in the middle.\n\nHowever, a discussion of blobs, trees, commits, objects, and\nreferences isn't necessarily low-level. It seems to me that it is a\nhigh-level understanding of the git world. Without those\n*definitions*, people are left to their own wrong, inconsistent\nthoughts.\n\nThe low-level stuff is HOW those concepts have been used in the\nimplementation of git: Where certain files are stored, how certain\nbytes are organized in memory, what are the underlying porcelain\ntools, etc. That what's low-level.\n\n> But we would need somebody to volunteer to write it. I would be happy to\n> help out, but I'm too short on time at the moment to be the driving\n> force.\n\nMaybe I'll try to write something, but it won't take place quickly,\neither. I'd want to read ALL of the existing documentation first.\n"},{"id":"112136","messageId":"D912CAB9-E437-4733-8A87-97EE47E3FBBB@boostpro.com","threadId":"19010","inReplyTo":"20090423201636.GD3056@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-23T21:26:34Z","receivedAt":"2009-04-23T21:26:34Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 23, 2009, at 4:16 PM, Jeff King wrote:\n\n> On Thu, Apr 23, 2009 at 01:37:05PM -0500, Michael Witten wrote:\n>\n>> Everyone talks about \"before one has the conceptual foundation\n>> necessary to understand\". Well, here's an idea: The git documentation\n>> should start with the concepts!\n>>\n>> Why don't the docs start out defining blobs and trees and the object\n>> database and references into that database? The reason everything is\n>> so confusing is that the understanding is brushed under the tutorial\n>> rug. People need to learn how to think before they can effectively\n>> learn to start doing.\n>\n> I agree with you, but not everyone does (and you can find prior  \n> debates\n> in the list archives). The user-manual is pretty \"top down\".\n\nAnd that's a problem because so many things are badly named.  It also  \nleaves out lots of top\n\n> There are\n> some \"bottom-up\" resources available, but I haven't seen one pointed  \n> to\n> as \"definitive\".\n\nI've been pointed at:\n\n1. http://eagain.net/articles/git-for-computer-scientists\n2. http://www.newartisans.com/2008/04/git-from-the-bottom-up.html\n\nwhich, IMO, should be read in that order.  I've just sent John Wiegley  \na huge pile of editorial commentary on #2, which I think could improve  \nthings.\n\nBut that said, \"laying conceptual foundation\" doesn't imply bottom- \nup!  In fact, I don't think the first one is particularly bottom-up\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112137","messageId":"B873CD38-2CFE-4138-8A77-8957FA3DB81C@boostpro.com","threadId":"19010","inReplyTo":"b4087cc50904231345x2613308eh640e50f4a2680890@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-23T21:31:13Z","receivedAt":"2009-04-23T21:31:13Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 23, 2009, at 4:45 PM, Michael Witten wrote:\n\n> On Thu, Apr 23, 2009 at 15:16, Jeff King <peff@peff.net> wrote:\n>> On Thu, Apr 23, 2009 at 01:37:05PM -0500, Michael Witten wrote:\n>>\n>>> Everyone talks about \"before one has the conceptual foundation\n>>> necessary to understand\". Well, here's an idea: The git  \n>>> documentation\n>>> should start with the concepts!\n>>>\n>>> Why don't the docs start out defining blobs and trees and the object\n>>> database and references into that database? The reason everything is\n>>> so confusing is that the understanding is brushed under the tutorial\n>>> rug. People need to learn how to think before they can effectively\n>>> learn to start doing.\n>>\n>> I agree with you, but not everyone does (and you can find prior  \n>> debates\n>> in the list archives). The user-manual is pretty \"top down\". There  \n>> are\n>> some \"bottom-up\" resources available, but I haven't seen one  \n>> pointed to\n>> as \"definitive\".I think it might actually be nice for there to be a\n>> parallel to the user manual that follows the bottom-up approach, and\n>> people could read the one that appeals most to them (or if they  \n>> have a\n>> lot of time on their hands, read both and hopefully it makes sense in\n>> the middle ;) ).\n>\n> I think the main problem, then, is that the tools have a UI that is\n> somewhere in the middle.\n\nWell, \"the UI\" (how many do we really have for Git?) is spread across  \nthe spectrum.  The git command-line alone lets you do incredibly low- \nlevel things that \"nobody should ever do\" and some really high-level  \nthings that are everyone's bread-and-butter.  There's no obvious  \ndistinction.\n\n> However, a discussion of blobs, trees, commits, objects, and\n> references isn't necessarily low-level. It seems to me that it is a\n> high-level understanding of the git world. Without those\n> *definitions*, people are left to their own wrong, inconsistent\n> thoughts.\n\n1000% agreed.\n\n> The low-level stuff is HOW those concepts have been used in the\n> implementation of git: Where certain files are stored, how certain\n> bytes are organized in memory, what are the underlying porcelain\n> tools, etc. That what's low-level.\n\nYep\n\n>> But we would need somebody to volunteer to write it. I would be  \n>> happy to\n>> help out, but I'm too short on time at the moment to be the driving\n>> force.\n>\n> Maybe I'll try to write something, but it won't take place quickly,\n> either. I'd want to read ALL of the existing documentation first.\n\nSee you in a couple years ;-)\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112140","messageId":"200904240051.46233.johan@herland.net","threadId":"19010","inReplyTo":"D912CAB9-E437-4733-8A87-97EE47E3FBBB@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2009-04-23T22:51:46Z","receivedAt":"2009-04-23T22:51:46Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Thursday 23 April 2009, David Abrahams wrote:\n> On Apr 23, 2009, at 4:16 PM, Jeff King wrote:\n> > There are some \"bottom-up\" resources available, but I haven't seen one\n> > pointed to as \"definitive\".\n> I've been pointed at:\n>\n> 1. http://eagain.net/articles/git-for-computer-scientists\n> 2. http://www.newartisans.com/2008/04/git-from-the-bottom-up.html\n\nThere's also http://www.eecs.harvard.edu/~cduan/technical/git/ which I think \nis a great bottom-up introduction:\n- not too heavy on the concepts\n- shows how the concepts relates to common git commands\n- short enough to be covered in just 1-2 sessions.\n\nIn fact, I'm loosely planning a presentation on Git (for $dayjob), and I'm \nprobably going to base it on this introduction.\n\n\nHave fun! :)\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"112147","messageId":"b4087cc50904231730i1e8a005cpaf1921e23df11da6@mail.gmail.com","threadId":"19010","inReplyTo":"200904240051.46233.johan@herland.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T00:30:56Z","receivedAt":"2009-04-24T00:30:56Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Thu, Apr 23, 2009 at 17:51, Johan Herland <johan@herland.net> wrote:\n> On Thursday 23 April 2009, David Abrahams wrote:\n>> On Apr 23, 2009, at 4:16 PM, Jeff King wrote:\n>> > There are some \"bottom-up\" resources available, but I haven't seen one\n>> > pointed to as \"definitive\".\n>> I've been pointed at:\n>>\n>> 1. http://eagain.net/articles/git-for-computer-scientists\n>> 2. http://www.newartisans.com/2008/04/git-from-the-bottom-up.html\n>\n> There's also http://www.eecs.harvard.edu/~cduan/technical/git/ which I think\n> is a great bottom-up introduction:\n> - not too heavy on the concepts\n\nI really don't understand this mentality. Concepts are the only things\nthat are important. From concepts falls all else.\n"},{"id":"112148","messageId":"b4087cc50904231731p7a3cb652g18dbb2cf5744111f@mail.gmail.com","threadId":"19010","inReplyTo":"B873CD38-2CFE-4138-8A77-8957FA3DB81C@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T00:31:30Z","receivedAt":"2009-04-24T00:31:30Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Thu, Apr 23, 2009 at 16:31, David Abrahams <dave@boostpro.com> wrote:\n>\n>\n>> However, a discussion of blobs, trees, commits, objects, and\n>> references isn't necessarily low-level. It seems to me that it is a\n>> high-level understanding of the git world. Without those\n>> *definitions*, people are left to their own wrong, inconsistent\n>> thoughts.\n>\n> 1000% agreed.\n\nI think this is a case in point:\n\n    http://marc.info/?l=git&m=124052299832318&w=2\n"},{"id":"112150","messageId":"20090424022900.GB6321@fieldses.org","threadId":"19010","inReplyTo":"b4087cc50904231137g67b4b84eu3b61bf174ba37d7f@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2009-04-24T02:29:00Z","receivedAt":"2009-04-24T02:29:00Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Thu, Apr 23, 2009 at 01:37:05PM -0500, Michael Witten wrote:\n> On Thu, Apr 23, 2009 at 12:57, J. Bruce Fields <bfields@fieldses.org> wrote:\n> > On Wed, Apr 22, 2009 at 03:38:52PM -0400, David Abrahams wrote:\n> >>\n> >> http://www.kernel.org/pub/software/scm/git/docs/user-manual.html#how-to-check-out\n> >> covers \"git reset\" way too early, IMO, before one has the conceptual\n> >> foundation necessary to understand what it means to \"modify the current\n> >> branch to point at v2.6.17\".  If this operation must be covered this\n> >> early in the manual, it should probably not be until\n> >> http://www.kernel.org/pub/software/scm/git/docs/user-manual.html#manipulating-branches\n> >\n> > I agree; we should suggest just a git-checkout (to a detached HEAD)\n> > instead, though that needs a little explanation so people aren't scared\n> > by the warning message it gives.\n> \n> Everyone talks about \"before one has the conceptual foundation\n> necessary to understand\". Well, here's an idea: The git documentation\n> should start with the concepts!\n> \n> Why don't the docs start out defining blobs and trees and the object\n> database and references into that database? The reason everything is\n> so confusing is that the understanding is brushed under the tutorial\n> rug. People need to learn how to think before they can effectively\n> learn to start doing.\n\nOK, but let's not over-generalize: the person that just wants to figure\nout whether the driver for their network card was fixed in today's\nnetwork devel tree shouldn't have to sit through a discussion of the\nobject database.  And even among readers that are in it for the long\nhaul, I think many people will react better to something that gives them\nat least a little concrete how-to information up front.\n\nSo the goal was always to find a tutorial route through the material\nthat would allow us to introduce the concepts as we go along.\n\nAnd I agree that I haven't succeeded at that--patches welcomed,\nincluding patches that, say, move more of the current chapter 7 to an\nearlier place.  (But this has to be done carefully, and I'd still rather\nit not be the *very* first thing.)\n\nI've unfortunately had a lot less time to work on this, but am happy to\nat least help review patches.\n\n--b.\n"},{"id":"112151","messageId":"b4087cc50904231934h17a090d4ie2a091843457eced@mail.gmail.com","threadId":"19010","inReplyTo":"20090424022900.GB6321@fieldses.org","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T02:34:56Z","receivedAt":"2009-04-24T02:34:56Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Thu, Apr 23, 2009 at 21:29, J. Bruce Fields <bfields@fieldses.org> wrote:\n> OK, but let's not over-generalize: the person that just wants to figure\n> out whether the driver for their network card was fixed in today's\n> network devel tree shouldn't have to sit through a discussion of the\n> object database.  And even among readers that are in it for the long\n> haul, I think many people will react better to something that gives them\n> at least a little concrete how-to information up front.\n\nA quick shell synopsis is probably what you want then. Beyond that,\ncasual users should be ignored; quick instructions are usually\nprovided by each project anyway.\n"},{"id":"112153","messageId":"C30426B2-CF6F-48D0-A7D3-F96D4D153057@boostpro.com","threadId":"19010","inReplyTo":"20090424022900.GB6321@fieldses.org","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-24T04:06:12Z","receivedAt":"2009-04-24T04:06:12Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 23, 2009, at 10:29 PM, J. Bruce Fields wrote:\n\n> On Thu, Apr 23, 2009 at 01:37:05PM -0500, Michael Witten wrote:\n>> On Thu, Apr 23, 2009 at 12:57, J. Bruce Fields  \n>> <bfields@fieldses.org> wrote:\n>>> On Wed, Apr 22, 2009 at 03:38:52PM -0400, David Abrahams wrote:\n>>>>\n>>>> http://www.kernel.org/pub/software/scm/git/docs/user-manual.html#how-to-check-out\n>>>> covers \"git reset\" way too early, IMO, before one has the  \n>>>> conceptual\n>>>> foundation necessary to understand what it means to \"modify the  \n>>>> current\n>>>> branch to point at v2.6.17\".  If this operation must be covered  \n>>>> this\n>>>> early in the manual, it should probably not be until\n>>>> http://www.kernel.org/pub/software/scm/git/docs/user-manual.html#manipulating-branches\n>>>\n>>> I agree; we should suggest just a git-checkout (to a detached HEAD)\n>>> instead, though that needs a little explanation so people aren't  \n>>> scared\n>>> by the warning message it gives.\n>>\n>> Everyone talks about \"before one has the conceptual foundation\n>> necessary to understand\". Well, here's an idea: The git documentation\n>> should start with the concepts!\n>>\n>> Why don't the docs start out defining blobs and trees and the object\n>> database and references into that database? The reason everything is\n>> so confusing is that the understanding is brushed under the tutorial\n>> rug. People need to learn how to think before they can effectively\n>> learn to start doing.\n>\n> OK, but let's not over-generalize: the person that just wants to  \n> figure\n> out whether the driver for their network card was fixed in today's\n> network devel tree shouldn't have to sit through a discussion of the\n> object database.\n\nThose people don't need a VCS.  They should download a snapshot or use  \na web interface.  Seriously.  There's no way you can make even the  \nbest-designed VCS simple enough to justify the time it takes to learn  \nenough just to use it for that.\n\n> And even among readers that are in it for the long\n> haul, I think many people will react better to something that gives  \n> them\n> at least a little concrete how-to information up front.\n\nPeople (well, people like me) should get a brief \"hello, world\" demo  \nup front, to give them a feel for the flavor of the system, but  \n[important:] it shouldn't attempt to be instructive.  Fundamental  \nconcepts are next.  How-to information can come after that, or after  \nthe reference information.\n\n> So the goal was always to find a tutorial route through the material\n> that would allow us to introduce the concepts as we go along.\n\nMaybe that will work for some people, but it *really* won't work for  \nme.  You can't start throwing around terms of art without defining  \nthem unless you want to raise more questions than you're answering.  I  \nwould be surprised if it wasn't the same for many tech people.\n\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112187","messageId":"20090424141037.GD15038@fieldses.org","threadId":"19010","inReplyTo":"C30426B2-CF6F-48D0-A7D3-F96D4D153057@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2009-04-24T14:10:37Z","receivedAt":"2009-04-24T14:10:37Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Fri, Apr 24, 2009 at 12:06:12AM -0400, David Abrahams wrote:\n>\n> On Apr 23, 2009, at 10:29 PM, J. Bruce Fields wrote:\n>\n>> On Thu, Apr 23, 2009 at 01:37:05PM -0500, Michael Witten wrote:\n>>> On Thu, Apr 23, 2009 at 12:57, J. Bruce Fields  \n>>> <bfields@fieldses.org> wrote:\n>>> Why don't the docs start out defining blobs and trees and the object\n>>> database and references into that database? The reason everything is\n>>> so confusing is that the understanding is brushed under the tutorial\n>>> rug. People need to learn how to think before they can effectively\n>>> learn to start doing.\n>>\n>> OK, but let's not over-generalize: the person that just wants to  \n>> figure\n>> out whether the driver for their network card was fixed in today's\n>> network devel tree shouldn't have to sit through a discussion of the\n>> object database.\n>\n> Those people don't need a VCS.  They should download a snapshot or use a \n> web interface.  Seriously.  There's no way you can make even the  \n> best-designed VCS simple enough to justify the time it takes to learn  \n> enough just to use it for that.\n>\n>> And even among readers that are in it for the long\n>> haul, I think many people will react better to something that gives  \n>> them\n>> at least a little concrete how-to information up front.\n>\n> People (well, people like me) should get a brief \"hello, world\" demo up \n> front, to give them a feel for the flavor of the system, but  \n> [important:] it shouldn't attempt to be instructive.  Fundamental  \n> concepts are next.  How-to information can come after that, or after the \n> reference information.\n>\n>> So the goal was always to find a tutorial route through the material\n>> that would allow us to introduce the concepts as we go along.\n>\n> Maybe that will work for some people, but it *really* won't work for me.  \n> You can't start throwing around terms of art without defining them unless \n> you want to raise more questions than you're answering.  I would be \n> surprised if it wasn't the same for many tech people.\n\nI agree that (with rare exceptions) terms shouldn't be used before\nthey're defined.  I don't agree with all of the above, but I think we\ncould come to a satisfactory compromise.  I'll see if I can find a few\nhours this weekend to at least sketch a new organization.  But, as I've\nsaid, I'm short on time and could really use some help.\n\n--b.\n"},{"id":"112186","messageId":"20090424141139.GC10761@coredump.intra.peff.net","threadId":"19010","inReplyTo":"b4087cc50904231345x2613308eh640e50f4a2680890@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-24T14:11:39Z","receivedAt":"2009-04-24T14:11:39Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Apr 23, 2009 at 03:45:46PM -0500, Michael Witten wrote:\n\n> However, a discussion of blobs, trees, commits, objects, and\n> references isn't necessarily low-level. It seems to me that it is a\n> high-level understanding of the git world. Without those\n> *definitions*, people are left to their own wrong, inconsistent\n> thoughts.\n> \n> The low-level stuff is HOW those concepts have been used in the\n> implementation of git: Where certain files are stored, how certain\n> bytes are organized in memory, what are the underlying porcelain\n> tools, etc. That what's low-level.\n\nI think I wasn't clear in my original message. I didn't mean teaching\nlow-level stuff like plumbing or file layouts. By \"bottom-up\" I really\nmeant teaching concepts (like objects, their types, and references),\nfrom which user operations and workflows can be explained (or often\ndeduced by the user). Whereas a top-down approach would _start_ with\nworkflows and say \"To accomplish X, do Y\".\n\nSo I think we are in agreement about the right \"level\" to start at.\n\n-Peff\n"},{"id":"112188","messageId":"20090424141847.GD10761@coredump.intra.peff.net","threadId":"19010","inReplyTo":"B873CD38-2CFE-4138-8A77-8957FA3DB81C@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-24T14:18:47Z","receivedAt":"2009-04-24T14:18:47Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Apr 23, 2009 at 05:31:13PM -0400, David Abrahams wrote:\n\n>> I think the main problem, then, is that the tools have a UI that is\n>> somewhere in the middle.\n>\n> Well, \"the UI\" (how many do we really have for Git?) is spread across the \n> spectrum.  The git command-line alone lets you do incredibly low-level \n> things that \"nobody should ever do\" and some really high-level things that \n> are everyone's bread-and-butter.  There's no obvious distinction.\n\nI think this is a bit better than it used to be. Plumbing commands are\nmostly hidden outside of the user's PATH. Unfortunately there are still\nsome warts, like the fact that users may be referred to \"git help\nrev-parse\" to learn about how revisions are specified. But they have to\nwade through the information on the \"rev-parse\" command, which is\nsomething that most users will never need to know or care about.\n\nA lot of that is historical baggage. The original git was not a VCS but\nrather a _toolkit_ for building a VCS. So the natural place for talking\nabout parsing revisions was rev-parse, because that was the only way to\naccess the revision parsing code. :)\n\nI think a lot of documentation like the \"specifying revisions\" section\nof rev-parse might benefit from being split into its own \"concept\"\nsection, like gitrevisions(7). And commands which allow specifying\nrevisions (at least the major ones, like log, diff, etc) should\nreference it (but not include it directly, as we do with some\ndocumentation snippets -- the point is to make the user aware that they\nare learning a separate concept that can be applied in multiple places,\nand to give that concept a name).\n\n-Peff\n"},{"id":"112189","messageId":"20090424142058.GF15038@fieldses.org","threadId":"19010","inReplyTo":"20090424141847.GD10761@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2009-04-24T14:20:58Z","receivedAt":"2009-04-24T14:20:58Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Fri, Apr 24, 2009 at 10:18:47AM -0400, Jeff King wrote:\n> On Thu, Apr 23, 2009 at 05:31:13PM -0400, David Abrahams wrote:\n> \n> >> I think the main problem, then, is that the tools have a UI that is\n> >> somewhere in the middle.\n> >\n> > Well, \"the UI\" (how many do we really have for Git?) is spread across the \n> > spectrum.  The git command-line alone lets you do incredibly low-level \n> > things that \"nobody should ever do\" and some really high-level things that \n> > are everyone's bread-and-butter.  There's no obvious distinction.\n> \n> I think this is a bit better than it used to be. Plumbing commands are\n> mostly hidden outside of the user's PATH. Unfortunately there are still\n> some warts, like the fact that users may be referred to \"git help\n> rev-parse\" to learn about how revisions are specified. But they have to\n> wade through the information on the \"rev-parse\" command, which is\n> something that most users will never need to know or care about.\n> \n> A lot of that is historical baggage. The original git was not a VCS but\n> rather a _toolkit_ for building a VCS. So the natural place for talking\n> about parsing revisions was rev-parse, because that was the only way to\n> access the revision parsing code. :)\n> \n> I think a lot of documentation like the \"specifying revisions\" section\n> of rev-parse might benefit from being split into its own \"concept\"\n> section, like gitrevisions(7). And commands which allow specifying\n> revisions (at least the major ones, like log, diff, etc) should\n> reference it (but not include it directly, as we do with some\n> documentation snippets -- the point is to make the user aware that they\n> are learning a separate concept that can be applied in multiple places,\n> and to give that concept a name).\n\nI'd be in favor of that.\n\n--b.\n"},{"id":"112190","messageId":"b4087cc50904240730n42e605e1od37d88d43e00f142@mail.gmail.com","threadId":"19010","inReplyTo":"20090424141139.GC10761@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T14:30:20Z","receivedAt":"2009-04-24T14:30:20Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 09:11, Jeff King <peff@peff.net> wrote:\n> I think I wasn't clear in my original message. I didn't mean teaching\n> low-level stuff like plumbing or file layouts. By \"bottom-up\" I really\n> meant teaching concepts (like objects, their types, and references),\n> from which user operations and workflows can be explained (or often\n> deduced by the user). Whereas a top-down approach would _start_ with\n> workflows and say \"To accomplish X, do Y\".\n\nI knew you would make exactly this rebuttle ;-D\n\nHowever, notice that you can't reasonably be expected to understand\n\"accomplish X\" without having concepts like objects and references.\nThe reason most people get by is that git's operation can be\ncompatible with a number of other theories people might have already\npicked up from using computers. The trouble starts when their existing\ntheories don't mesh well with the underlying git theory, leading the\nuser to develop the equivalent of epicycles in order to explain to\nhimself whats going on.\n\nBasically, the problem is that the documentation is currently catering\nfor people, who just want to download source files (as Bruce basically\nsaid); a quick shell synopsis for this is fine, but there needs to be\ndocumentation solely devoted to understanding git fully and precisely.\n"},{"id":"112191","messageId":"b4087cc50904240733u53a2c9a0o3d0943dc7de38324@mail.gmail.com","threadId":"19010","inReplyTo":"b4087cc50904240730n42e605e1od37d88d43e00f142@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T14:33:49Z","receivedAt":"2009-04-24T14:33:49Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 09:30, Michael Witten <mfwitten@gmail.com> wrote:\n> there needs to be\n> documentation solely devoted to understanding git fully and precisely.\n\nA user should be able to read from top-to bottom in one-pass----no\njumping around or later clarifications.\n"},{"id":"112195","messageId":"20090424150442.GA11245@coredump.intra.peff.net","threadId":"19010","inReplyTo":"b4087cc50904240730n42e605e1od37d88d43e00f142@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-24T15:04:42Z","receivedAt":"2009-04-24T15:04:42Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 24, 2009 at 09:30:20AM -0500, Michael Witten wrote:\n\n> On Fri, Apr 24, 2009 at 09:11, Jeff King <peff@peff.net> wrote:\n> > I think I wasn't clear in my original message. I didn't mean teaching\n> > low-level stuff like plumbing or file layouts. By \"bottom-up\" I really\n> > meant teaching concepts (like objects, their types, and references),\n> > from which user operations and workflows can be explained (or often\n> > deduced by the user). Whereas a top-down approach would _start_ with\n> > workflows and say \"To accomplish X, do Y\".\n> \n> I knew you would make exactly this rebuttle ;-D\n> \n> However, notice that you can't reasonably be expected to understand\n> \"accomplish X\" without having concepts like objects and references.\n\nHeh. I don't think you also predicted the paragraph that I ended up\ndeleting, which made it more clear that I was not trying to rebut, but\nrather agree.\n\nLike you, I think that not teaching concepts first leads to confusion\nlater.  Version control (or at least git) is just complex enough that\nyou are much better off understanding what is happening than simply\nfollowing a recipe. So when your recipe doesn't go as planned, or you\ndon't know which recipe to use, or you need some variant of a recipe,\nyou have some basis for understanding what to do.\n\nBut users in the past have really seemed to want to start with recipes,\nso that they can be productive as soon as possible (and I think some\npeople have said that the top-down ordering just makes more sense to\nthem, so it may just be a matter of learning style). And I think the\nuser manual is somewhat of a response to that request, since the\ncommand manpages are very bottom-up (but are also quite confusing, just\nbecause of their size, and because concept information is scattered\nthroughout).\n\nSo I am advocating for more bottom-up documentation (which I think you\nare), but I don't necessarily think it should _replace_ the top-down\ndocumentation (which I'm not sure is your position or not).\n\n> The reason most people get by is that git's operation can be\n> compatible with a number of other theories people might have already\n> picked up from using computers. The trouble starts when their existing\n> theories don't mesh well with the underlying git theory, leading the\n> user to develop the equivalent of epicycles in order to explain to\n> himself whats going on.\n\nEpicycles? I thought commit orbits were defined by the ether through\nthey flowed.\n\n-Peff\n"},{"id":"112196","messageId":"b4087cc50904240818w45bd1cfaq8bbc83e10a6e3781@mail.gmail.com","threadId":"19010","inReplyTo":"20090424150442.GA11245@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T15:18:15Z","receivedAt":"2009-04-24T15:18:15Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 10:04, Jeff King <peff@peff.net> wrote:\n> On Fri, Apr 24, 2009 at 09:30:20AM -0500, Michael Witten wrote:\n>\n>> On Fri, Apr 24, 2009 at 09:11, Jeff King <peff@peff.net> wrote:\n>> > I think I wasn't clear in my original message. I didn't mean teaching\n>> > low-level stuff like plumbing or file layouts. By \"bottom-up\" I really\n>> > meant teaching concepts (like objects, their types, and references),\n>> > from which user operations and workflows can be explained (or often\n>> > deduced by the user). Whereas a top-down approach would _start_ with\n>> > workflows and say \"To accomplish X, do Y\".\n>>\n>> I knew you would make exactly this rebuttle ;-D\n>>\n>> However, notice that you can't reasonably be expected to understand\n>> \"accomplish X\" without having concepts like objects and references.\n>\n> Heh. I don't think you also predicted the paragraph that I ended up\n> deleting, which made it more clear that I was not trying to rebut, but\n> rather agree.\n\nIndeed. I saw that last sentence of yours, but I consciously ignored\nit, because I like to argue ;-)\n\n> Like you, I think that not teaching concepts first leads to confusion\n> later.  Version control (or at least git) is just complex enough that\n> you are much better off understanding what is happening than simply\n> following a recipe. So when your recipe doesn't go as planned, or you\n> don't know which recipe to use, or you need some variant of a recipe,\n> you have some basis for understanding what to do.\n\nThat, my friend, is the most important lesson of learning.\n\n> But users in the past have really seemed to want to start with recipes,\n> so that they can be productive as soon as possible (and I think some\n> people have said that the top-down ordering just makes more sense to\n> them, so it may just be a matter of learning style). And I think the\n> user manual is somewhat of a response to that request, since the\n> command manpages are very bottom-up (but are also quite confusing, just\n> because of their size, and because concept information is scattered\n> throughout).\n>\n> So I am advocating for more bottom-up documentation (which I think you\n> are), but I don't necessarily think it should _replace_ the top-down\n> documentation (which I'm not sure is your position or not).\n\nI think that we've already got that tutorial-esque style covered (I\nhaven't read it in a while):\n\n    http://www.kernel.org/pub/software/scm/git/docs/gittutorial.html\n\nHowever, the User Manual should make a Mathematician happy.\n\n>> The reason most people get by is that git's operation can be\n>> compatible with a number of other theories people might have already\n>> picked up from using computers. The trouble starts when their existing\n>> theories don't mesh well with the underlying git theory, leading the\n>> user to develop the equivalent of epicycles in order to explain to\n>> himself whats going on.\n>\n> Epicycles? I thought commit orbits were defined by the ether through\n> they flowed.\n\nActually, those commit orbits are defined by the giant glass sphere to\nwhich they are attached.\n"},{"id":"112218","messageId":"91225E09-505A-4CD6-AC8E-FBB500A95984@boostpro.com","threadId":"19010","inReplyTo":"20090424141847.GD10761@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-24T17:28:35Z","receivedAt":"2009-04-24T17:28:35Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 24, 2009, at 10:18 AM, Jeff King wrote:\n\n> On Thu, Apr 23, 2009 at 05:31:13PM -0400, David Abrahams wrote:\n>\n>>> I think the main problem, then, is that the tools have a UI that is\n>>> somewhere in the middle.\n>>\n>> Well, \"the UI\" (how many do we really have for Git?) is spread  \n>> across the\n>> spectrum.  The git command-line alone lets you do incredibly low- \n>> level\n>> things that \"nobody should ever do\" and some really high-level  \n>> things that\n>> are everyone's bread-and-butter.  There's no obvious distinction.\n>\n> I think this is a bit better than it used to be. Plumbing commands are\n> mostly hidden outside of the user's PATH.\n\nHuh?\n\ngit hash-object\ngit cat-file -t ...\ngit ls-tree\ngit rev-parse\ngit write-tree\ngit commit-tree\n\n   ...\n\nThese are just some of the ones I learned about by reading John  \nWiegley's \"Git From the Bottom Up.\"\n\nMaybe I'm wrong about rev-parse, but for the most part, having all  \nthese low-level commands available through the same executable that's  \nused for \"git add,\" \"git merge,\" \"git commit,\" et. al. makes the whole  \nshebang hard to approach.  It would be better for users if the low- \nlevel stuff was accessed some other way.\n\n> A lot of that is historical baggage. The original git was not a VCS  \n> but\n> rather a _toolkit_ for building a VCS. So the natural place for  \n> talking\n> about parsing revisions was rev-parse, because that was the only way  \n> to\n> access the revision parsing code. :)\n\nI understand that, but it doesn't change the present reality.\n\n> I think a lot of documentation like the \"specifying revisions\" section\n> of rev-parse might benefit from being split into its own \"concept\"\n> section, like gitrevisions(7).\n\nYes, please.\n\n\n[excuse me, but what the #@&*! is \"porcelainish\" supposed to mean? (http://www.kernel.org/pub/software/scm/git/docs/git-rev-parse.html \n)]\n\n> And commands which allow specifying\n> revisions (at least the major ones, like log, diff, etc) should\n> reference it (but not include it directly, as we do with some\n> documentation snippets -- the point is to make the user aware that  \n> they\n> are learning a separate concept that can be applied in multiple  \n> places,\n> and to give that concept a name).\n\n\nVery nice.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112220","messageId":"20090424173852.GF17365@fieldses.org","threadId":"19010","inReplyTo":"b4087cc50904240818w45bd1cfaq8bbc83e10a6e3781@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2009-04-24T17:38:52Z","receivedAt":"2009-04-24T17:38:52Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Fri, Apr 24, 2009 at 10:18:15AM -0500, Michael Witten wrote:\n> I think that we've already got that tutorial-esque style covered (I\n> haven't read it in a while):\n> \n>     http://www.kernel.org/pub/software/scm/git/docs/gittutorial.html\n> \n> However, the User Manual should make a Mathematician happy.\n\nI'm all for making mathematicians happy.  But, again, help?:\n\n\t- Specific examples?\n\t- Patches?  Please, patches?\n\t- Suggested text?\n\t- Suggested outline?\n\nThere's no shortage of high-level ideas.  What there's always a need for\nmore of is people willing to submit patches, respond to review, etc.\n\n--b.\n"},{"id":"112224","messageId":"20090424181539.GB11360@coredump.intra.peff.net","threadId":"19010","inReplyTo":"91225E09-505A-4CD6-AC8E-FBB500A95984@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-24T18:15:39Z","receivedAt":"2009-04-24T18:15:39Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 24, 2009 at 01:28:35PM -0400, David Abrahams wrote:\n\n>> I think this is a bit better than it used to be. Plumbing commands are\n>> mostly hidden outside of the user's PATH.\n>\n> Huh?\n>\n> git hash-object\n> git cat-file -t ...\n> git ls-tree\n> git rev-parse\n> git write-tree\n> git commit-tree\n\nHow did you find out about them? They are not in your PATH, so shell\ncompletion doesn't find them. They are not in the programmable bash\ncompletion. They are not in the short command list git gives you when\nyou type \"git help\" or \"git\" without arguments.\n\nSo you must have read about them somewhere...\n\n> These are just some of the ones I learned about by reading John Wiegley's \n> \"Git From the Bottom Up.\"\n\n...like here. So if that document gave you the impression that those are\npart of an everyday git workflow, then I think the document is at fault,\nnot git itself.\n\nI admit I haven't read \"Git From the Bottom Up\" carefully, but I think\nwhat Michael is proposing would probably start a little higher from the\nbottom than that document. You can give the concepts of the object\ntypes, show them in pretty-printed form with \"git show\", and not worry\nabout telling the user \"this is how 'git commit' could be implemented in\nterms of primitive operations\". And then you can avoid most of the\nlow-level commands entirely.\n\n> Maybe I'm wrong about rev-parse, but for the most part, having all these \n> low-level commands available through the same executable that's used for \n> \"git add,\" \"git merge,\" \"git commit,\" et. al. makes the whole shebang hard \n> to approach.  It would be better for users if the low-level stuff was \n> accessed some other way.\n\nPerhaps. The general approach is to make those commands accessible as\n\"git foo\", but not to _advertise_ them in the same way as the porcelain\ncommands. The idea was to give a uniform calling convention without\nunnecessarily confusing users by presenting a large number of\ninfrequently-used commands.\n\nAt any rate, it is too late to change the calling convention for\nplumbing. The whole point of them is to be a stable interface for\nscripting. Changing them to \"git low-level rev-parse\" (if it was even\nsomething that we wanted to do, which I don't think it is) would break\neveryone's scripts.\n\n>> A lot of that is historical baggage. The original git was not a VCS but\n>> rather a _toolkit_ for building a VCS. So the natural place for talking\n>> about parsing revisions was rev-parse, because that was the only way to\n>> access the revision parsing code. :)\n>\n> I understand that, but it doesn't change the present reality.\n\nRight. I'm just trying to say how we got here, which I think is relevant\nbecause it gives a hint of what directions we can go in. In other words,\nnobody _designed_ what we have now. It evolved into this state, which\nobviously has some drawbacks. So I think you won't find much resistance\nin trying to evolve the documentation to present git more as a coherent\ntool, and less as a set of unrelated commands.\n\n> [excuse me, but what the #@&*! is \"porcelainish\" supposed to mean? \n> (http://www.kernel.org/pub/software/scm/git/docs/git-rev-parse.html)]\n\nHeh. That one is particularly egregious, because it rests on several\nlayers of git jargon. The low-level tools are plumbing, like pipes and\nvalves. The high-level commands intended for end users are porcelain,\nlike sinks and toilets. The -ish suffix is often used in git to refer to\na type, or something we can convert into a type (like a \"tree-ish\" could\nbe a tree object, or a commit object which points to a tree, or a tag\nobject which points to a commit which points to a tree). So I think by\nsaying \"porcelain-ish\" here, the author meant \"not just porcelain, but\nother things which take revisions and behave sort of like porcelain\".\n\nWhich is a truly horrible thing to throw at a new user who just wants to\nsee how to specify a revision.\n\nSo yeah, if you are saying that could be worded better, I absolutely\nagree. There are a lot of spots like that. They are getting fixed slowly\nover time. I'm not sure if that is enough, or if somebody knowledgeable\nreally needs to take a sledge hammer to the existing documentation and\njust reorganize and rewrite a lot of it.\n\n-Peff\n"},{"id":"112225","messageId":"20090424182752.GC11360@coredump.intra.peff.net","threadId":"19010","inReplyTo":"20090424173852.GF17365@fieldses.org","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-24T18:27:52Z","receivedAt":"2009-04-24T18:27:52Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 24, 2009 at 01:38:52PM -0400, J. Bruce Fields wrote:\n\n> On Fri, Apr 24, 2009 at 10:18:15AM -0500, Michael Witten wrote:\n> > I think that we've already got that tutorial-esque style covered (I\n> > haven't read it in a while):\n> > \n> >     http://www.kernel.org/pub/software/scm/git/docs/gittutorial.html\n> > \n> > However, the User Manual should make a Mathematician happy.\n> \n> I'm all for making mathematicians happy.  But, again, help?:\n> \n> \t- Specific examples?\n> \t- Patches?  Please, patches?\n> \t- Suggested text?\n> \t- Suggested outline?\n> \n> There's no shortage of high-level ideas.  What there's always a need for\n> more of is people willing to submit patches, respond to review, etc.\n\nI usually hate to \"me too\", but I really want to second this notion. We\nhave been getting minor documentation fixups trickling in, and I think\nthose really help, and maybe they eventually would make the\ndocumentation perfect. But I have the feeling we would benefit from\nsomebody taking ownership and considering the big picture of how the\ndocumentation fits together, and then really pushing it forward with\nsomething concrete.\n\n-Peff\n"},{"id":"112226","messageId":"20090424183534.GJ17365@fieldses.org","threadId":"19010","inReplyTo":"20090424182752.GC11360@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2009-04-24T18:35:34Z","receivedAt":"2009-04-24T18:35:34Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Fri, Apr 24, 2009 at 02:27:52PM -0400, Jeff King wrote:\n> On Fri, Apr 24, 2009 at 01:38:52PM -0400, J. Bruce Fields wrote:\n> \n> > On Fri, Apr 24, 2009 at 10:18:15AM -0500, Michael Witten wrote:\n> > > I think that we've already got that tutorial-esque style covered (I\n> > > haven't read it in a while):\n> > > \n> > >     http://www.kernel.org/pub/software/scm/git/docs/gittutorial.html\n> > > \n> > > However, the User Manual should make a Mathematician happy.\n> > \n> > I'm all for making mathematicians happy.  But, again, help?:\n> > \n> > \t- Specific examples?\n> > \t- Patches?  Please, patches?\n> > \t- Suggested text?\n> > \t- Suggested outline?\n> > \n> > There's no shortage of high-level ideas.  What there's always a need for\n> > more of is people willing to submit patches, respond to review, etc.\n> \n> I usually hate to \"me too\", but I really want to second this notion. We\n> have been getting minor documentation fixups trickling in, and I think\n> those really help, and maybe they eventually would make the\n> documentation perfect. But I have the feeling we would benefit from\n> somebody taking ownership and considering the big picture of how the\n> documentation fits together, and then really pushing it forward with\n> something concrete.\n\nYup, and dealing seriously with objections, getting concensus for the\nresulting solutions, etc--in other words, being a maintainer.  I thought\nI'd be able to do that at some point, but just haven't consistently had\nthe time.\n\nThat said, several smaller suggestions have been made which could be\nhandled now:\n\n\t- I don't think I've seen objections to the idea of a\n\t  git-revision-specifying manpage, whatever you want to call\n\t  it--so probably that just needs someone to write the patch.\n\t- There've been complaints about terms being used before they're\n\t  defined sufficiently well.  I can believe it, but: specific\n\t  examples would help!\n\n--b.\n"},{"id":"112227","messageId":"20090424185208.GM17365@fieldses.org","threadId":"19010","inReplyTo":"34BD51FF-0908-48A8-BBBC-E27B0EFB32E5@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2009-04-24T18:52:08Z","receivedAt":"2009-04-24T18:52:08Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Fri, Apr 24, 2009 at 02:32:36PM -0400, David Abrahams wrote:\n>\n> On Apr 24, 2009, at 1:38 PM, J. Bruce Fields wrote:\n>\n>> On Fri, Apr 24, 2009 at 10:18:15AM -0500, Michael Witten wrote:\n>>> I think that we've already got that tutorial-esque style covered (I\n>>> haven't read it in a while):\n>>>\n>>>    http://www.kernel.org/pub/software/scm/git/docs/gittutorial.html\n>>>\n>>> However, the User Manual should make a Mathematician happy.\n>>\n>> I'm all for making mathematicians happy.  But, again, help?:\n>>\n>> \t- Specific examples?\n>> \t- Patches?  Please, patches?\n>> \t- Suggested text?\n>> \t- Suggested outline?\n>>\n>> There's no shortage of high-level ideas.  What there's always a need  \n>> for\n>> more of is people willing to submit patches, respond to review, etc.\n>\n>\n> I'll probably try to write something myself once I figure this stuff  \n> out.\n\nThat would be great, thanks.  Several people have gone off and posted\ntheir own tutorials someplace, and that's fine, but it would be\nespecially helpful if you could contribute to the actual Documentation/\ndirectory.  That may mean arguing with people and making compromises.\nBut it also means the results will be distributed with git, will be\nintegrated with other git documentation, and will get first-class\ntechnical review.\n\nI'd also encourage incrementally improving existing documentation where\npossible instead of starting over from scratch.  But having broken that\nrule myself a couple times I'm hardly in a position to insist.  If you\nmust start over, at least think about how to replace or fit it in with\nexisting documentation.\n\n--b.\n"},{"id":"112228","messageId":"0FC64949-689C-43A4-B656-9618E808962B@boostpro.com","threadId":"19010","inReplyTo":"20090424181539.GB11360@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-24T19:00:19Z","receivedAt":"2009-04-24T19:00:19Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 24, 2009, at 2:15 PM, Jeff King wrote:\n\n> On Fri, Apr 24, 2009 at 01:28:35PM -0400, David Abrahams wrote:\n>\n>>> I think this is a bit better than it used to be. Plumbing commands  \n>>> are\n>>> mostly hidden outside of the user's PATH.\n>>\n>> Huh?\n>>\n>> git hash-object\n>> git cat-file -t ...\n>> git ls-tree\n>> git rev-parse\n>> git write-tree\n>> git commit-tree\n>\n> How did you find out about them?\n\nThe first time?\n\n  $ man git\n\n> They are not in your PATH, so shell\n> completion doesn't find them.\n\nHuh?  `which git` works.  ls-tree is an argument to git as far as I  \nknow.\n\nYes, I know there are aliases like git-ls-tree somewhere, but that  \nonly adds to the sense that all commands are equal.\n\n> They are not in the programmable bash\n> completion. They are not in the short command list git gives you when\n> you type \"git help\" or \"git\" without arguments.\n>\n> So you must have read about them somewhere..\n\n   $ man git\n\nwhich makes no distinction.\n\n   $ xxx [--]help\n\nis usually OK if I already know xxx pretty well and just want a  \nrefresher.  If know I'll need a little more than that, I use man  \nstraight away.\n\n>> These are just some of the ones I learned about by reading John  \n>> Wiegley's\n>> \"Git From the Bottom Up.\"\n>\n> ...like here.\n\nThat's where I learned *what they do*.\n\n> So if that document gave you the impression that those are\n> part of an everyday git workflow, then I think the document is at  \n> fault,\n> not git itself.\n\nIt didn't.\n\n> I admit I haven't read \"Git From the Bottom Up\" carefully, but I think\n> what Michael is proposing would probably start a little higher from  \n> the\n> bottom than that document.\n\nYes, please.  \"Git for Computer Scientists\" is a great foundation.   \n From there add more information about naming things so I know what  \nthings like remotes/origin/master mean when I see them in gitk, and  \nI'm off to the races.\n\n> You can give the concepts of the object\n> types, show them in pretty-printed form with \"git show\", and not worry\n> about telling the user \"this is how 'git commit' could be  \n> implemented in\n> terms of primitive operations\". And then you can avoid most of the\n> low-level commands entirely.\n\nYes, that's fine.  Although I think there may be some things in GFTBU  \nthat are good fundamental concepts.  There's a nice list of terms with  \ndefinitions early in the document.\n\n>> Maybe I'm wrong about rev-parse, but for the most part, having all  \n>> these\n>> low-level commands available through the same executable that's  \n>> used for\n>> \"git add,\" \"git merge,\" \"git commit,\" et. al. makes the whole  \n>> shebang hard\n>> to approach.  It would be better for users if the low-level stuff was\n>> accessed some other way.\n>\n> Perhaps. The general approach is to make those commands accessible as\n> \"git foo\", but not to _advertise_ them in the same way as the  \n> porcelain\n> commands.\n\nWhat is \"porcelain,\" please?  This is one among many examples of  \njargon used only (or encountered by me for the first time) in the Git  \ncommunity.\n\n> The idea was to give a uniform calling convention without\n> unnecessarily confusing users by presenting a large number of\n> infrequently-used commands.\n\nIt's not working, I'm sorry to say.\n\n> At any rate, it is too late to change the calling convention for\n> plumbing.\n\nI disagree.  You can leave the old functionality there in a  \n\"deprecated\" state and change the way you advertise it.  It would even  \nhelp a lot if the plumbing were all spelled \"git-xxx\" and the high  \nlevel stuff were \"git xxx.\"\n\n> The whole point of them is to be a stable interface for\n> scripting. Changing them to \"git low-level rev-parse\" (if it was even\n> something that we wanted to do, which I don't think it is) would break\n> everyone's scripts.\n\nSee above.\n\n>> [excuse me, but what the #@&*! is \"porcelainish\" supposed to mean?\n>> (http://www.kernel.org/pub/software/scm/git/docs/git-rev-parse.html)]\n>\n> Heh. That one is particularly egregious, because it rests on several\n> layers of git jargon. The low-level tools are plumbing, like pipes and\n> valves.\n\n? I use the valves on my kitchen sink all the time.\n\n> The high-level commands intended for end users are porcelain,\n> like sinks and toilets. The -ish suffix is often used in git to  \n> refer to\n> a type, or something we can convert into a type (like a \"tree-ish\"  \n> could\n> be a tree object, or a commit object which points to a tree, or a tag\n> object which points to a commit which points to a tree). So I think by\n> saying \"porcelain-ish\" here, the author meant \"not just porcelain, but\n> other things which take revisions and behave sort of like porcelain\".\n\nbah. humbug.\n\n> Which is a truly horrible thing to throw at a new user who just  \n> wants to\n> see how to specify a revision.\n\nyeeeeah.\n\n> So yeah, if you are saying that could be worded better, I absolutely\n> agree. There are a lot of spots like that. They are getting fixed  \n> slowly\n> over time. I'm not sure if that is enough, or if somebody  \n> knowledgeable\n> really needs to take a sledge hammer to the existing documentation and\n> just reorganize and rewrite a lot of it.\n\n\nI'm thinking the latter.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112229","messageId":"b4087cc50904241212x97b91e6t894f14d7f1146f89@mail.gmail.com","threadId":"19010","inReplyTo":"20090424173852.GF17365@fieldses.org","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T19:12:27Z","receivedAt":"2009-04-24T19:12:27Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 12:38, J. Bruce Fields <bfields@fieldses.org> wrote:\n> I'm all for making mathematicians happy.  But, again, help?:\n\nI intend to help, but I have a terrible tendency to shave the GNU;\nright now, I'm waist deep in shavings.\n"},{"id":"112232","messageId":"20090424202403.GB13561@coredump.intra.peff.net","threadId":"19010","inReplyTo":"0FC64949-689C-43A4-B656-9618E808962B@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-24T20:24:04Z","receivedAt":"2009-04-24T20:24:04Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 24, 2009 at 03:00:19PM -0400, David Abrahams wrote:\n\n>> How did you find out about them?\n>\n> The first time?\n>\n>  $ man git\n>\n> [...]\n>\n> which makes no distinction [between porcelain and plumbing].\n\nReally? The command list in my version is divided into \"HIGH-LEVEL\nCOMMANDS (PORCELAIN)\" and \"LOW-LEVEL COMMANDS (PLUMBING)\", with the\ncommands you mentioned falling into the latter. And skimming \"git log\nDocumentation/git.txt\", it looks like it has been that way for some\ntime.\n\nThere is a little discussion under the plumbing section of what plumbing\nis. It could perhaps be more emphatic in warning regular users away.\n\n>> They are not in your PATH, so shell\n>> completion doesn't find them.\n>\n> Huh?  `which git` works.  ls-tree is an argument to git as far as I know.\n\nYes, but shell completion will never present you with the text\n\"ls-tree\". You have to have found out about it somewhere else (and\ncompletion used to show, because git-ls-tree was in the PATH).\n\n>   $ xxx [--]help\n>\n> is usually OK if I already know xxx pretty well and just want a  \n> refresher.  If know I'll need a little more than that, I use man straight \n> away.\n\ngit --help shows a list of common commands, but otherwise \"git help\nfoo\" and \"git foo --help\" _do_ show the manpage. It may be that \"man\ngit\" could use some cleanup; specific suggestions are welcome.\n\n> What is \"porcelain,\" please?  This is one among many examples of jargon \n> used only (or encountered by me for the first time) in the Git community.\n\nI think I ended up explaining it later in my email, but let me know if\nyou are still confused.\n\n>> The idea was to give a uniform calling convention without\n>> unnecessarily confusing users by presenting a large number of\n>> infrequently-used commands.\n>\n> It's not working, I'm sorry to say.\n\nRight, that's why I'm trying to figure out why you are hung up on the\nlow-level commands. The idea was that you wouldn't need to be exposed to\nthem at all, but obviously you were (or if you were exposed, it would be\nin a list that was clearly marked as \"this is low-level stuff that you\ndon't really need to worry about\". So I'm trying to figure out where it\nwent wrong.\n\n>> At any rate, it is too late to change the calling convention for\n>> plumbing.\n>\n> I disagree.  You can leave the old functionality there in a \"deprecated\" \n> state and change the way you advertise it.\n\nBut does that really help? It means that \"git hash-object\" is still\nthere, which I thought was the problem you had. You can argue that it\nwouldn't be advertised to users, and so wouldn't be a problem, but that\nis _already_ the strategy we are using. So either that strategy is fine,\nin which case we are on the right track but may still have some work to\ndo in properly implementing it. Or it's not, in which case your proposal\nis no better.\n\n> It would even help a lot if the plumbing were all spelled \"git-xxx\"\n> and the high level stuff were \"git xxx.\"\n\nDifferentating calling conventions like that was proposed when dashed\nforms were deprecated and removed from the PATH. But if we had dashed\nforms for plumbing (i.e., not forwarding them via the \"git\" wrapper),\nthen you have to do one of:\n\n  - put them in the user's PATH. Now tab completion or looking in your\n    PATH means you see _just_ the plumbing commands, and none of the\n    high level ones. Which is one of the reasons they were removed from\n    the PATH in the first place (due to numerous user complaints).\n\n  - put them elsewhere, and force plumbing users to add $GIT_EXEC_PATH\n    to their PATH. That becomes very annoying for casual plumbing users.\n    If you come to the mailing list with a problem, I would have to jump\n    through extra hoops to ask you to show me the output of \"git\n    ls-files\".\n\nNot to mention that the git wrapper does other useful things besides\nsimply exec'ing. For example, it supports --git-dir, --bare, etc.\nSo the problem is that the low-level commands _are_ still useful, and\nmany people still want to call them, just like regular git commands.\nIt's just that they are numerous and low-level, which makes them\ndaunting for new users.\n\nAnd it has become obvious over several years of the git mailing list\nthat users, once they see mention of a command, must start investigating \nit to find out if and how it is useful. And I am not saying that is a\nfailing of users; on the contrary, I think it is quite a healthy\nbehavior on a unix-ish system. But it means that if we want not to\nadvertise low-level commands, we have to be very careful about the ways\nin which we mention them.\n\nPerhaps it would make sense for each plumbing command's man page to\nstart with something like \"this is a low-level command used for\nscripting git or investigating its internals. For high-level use, you\nmay be more interested in $X\", where $X may be \"git commit\" for\nwrite-tree, commit-tree, etc. And that would at least help intercept\nusers before they get too confused.\n\n>>> [excuse me, but what the #@&*! is \"porcelainish\" supposed to mean?\n>>> (http://www.kernel.org/pub/software/scm/git/docs/git-rev-parse.html)]\n>>\n>> Heh. That one is particularly egregious, because it rests on several\n>> layers of git jargon. The low-level tools are plumbing, like pipes and\n>> valves.\n>\n> ? I use the valves on my kitchen sink all the time.\n\nSorry, I meant the ones under the sink, that you would use if you were\nreplacing the faucet. I would call the ones above \"taps\". But hopefully\nyou get a sense of the distinction between plumbing and porcelain.\n\n-Peff\n"},{"id":"112233","messageId":"200904242230.13239.johan@herland.net","threadId":"19010","inReplyTo":"b4087cc50904231730i1e8a005cpaf1921e23df11da6@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2009-04-24T20:30:12Z","receivedAt":"2009-04-24T20:30:12Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Friday 24 April 2009, Michael Witten wrote:\n> On Thu, Apr 23, 2009 at 17:51, Johan Herland <johan@herland.net> wrote:\n> > There's also http://www.eecs.harvard.edu/~cduan/technical/git/ which I\n> > think is a great bottom-up introduction:\n> > - not too heavy on the concepts\n>\n> I really don't understand this mentality. Concepts are the only things\n> that are important. From concepts falls all else.\n\nSorry for not being clear: Concepts are indeed (and should be) important. \nWhat I mean is that the concepts introduced are short and simple enough for \nnovice users to understand (without much VCS experience, if any at all). If \nwe start off _too_ detailed, we risk loosing the audience, and no one is \nbetter off.\n\nLike Jeff King said elsewhere in this thread: We want to start a little \nhigher from the bottom. The above introduction does not focus on blobs or \ntrees, but manages to introduce Git in a useful manner by starting off with \nonly two concepts: commits and refs. With only these two concepts, and \nshowing how high-level commands (remember: no plumbing) work with these \nconcepts, I believe it is possible to teach anyone to use Git well. Of \ncourse, as users progress towards becoming power-users, more concepts are \nneeded, but I don't think these are needed from the start.\n\nAs Einstein might have said: As simple as possible, but no simpler.\n\n\nHave fun!\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"112235","messageId":"2B5084A3-9BDB-4463-8530-3C8AB2E09A1F@boostpro.com","threadId":"19010","inReplyTo":"20090424202403.GB13561@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-24T21:06:27Z","receivedAt":"2009-04-24T21:06:27Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 24, 2009, at 4:24 PM, Jeff King wrote:\n\n> On Fri, Apr 24, 2009 at 03:00:19PM -0400, David Abrahams wrote:\n>\n>>> How did you find out about them?\n>>\n>> The first time?\n>>\n>> $ man git\n>>\n>> [...]\n>>\n>> which makes no distinction [between porcelain and plumbing].\n>\n> Really? The command list in my version is divided into \"HIGH-LEVEL\n> COMMANDS (PORCELAIN)\" and \"LOW-LEVEL COMMANDS (PLUMBING)\", with the\n> commands you mentioned falling into the latter. And skimming \"git log\n> Documentation/git.txt\", it looks like it has been that way for some\n> time.\n\nSorry, you are totally right.\n\nThe list is just so crazy-long; I may have skimmed it.\n\n>> Huh?  `which git` works.  ls-tree is an argument to git as far as I  \n>> know.\n>\n> Yes, but shell completion will never present you with the text\n> \"ls-tree\". You have to have found out about it somewhere else (and\n> completion used to show, because git-ls-tree was in the PATH).\n>\n>>  $ xxx [--]help\n>>\n>> is usually OK if I already know xxx pretty well and just want a\n>> refresher.  If know I'll need a little more than that, I use man  \n>> straight\n>> away.\n>\n> git --help shows a list of common commands, but otherwise \"git help\n> foo\" and \"git foo --help\" _do_ show the manpage. It may be that \"man\n> git\" could use some cleanup; specific suggestions are welcome.\n>\n>> What is \"porcelain,\" please?  This is one among many examples of  \n>> jargon\n>> used only (or encountered by me for the first time) in the Git  \n>> community.\n>\n> I think I ended up explaining it later in my email, but let me know if\n> you are still confused.\n\nNope; I'm fine now.  It's not a great analogy, because everyone who  \nuses a sink ends up dealing with spigots and valves, but I get it.\n\n>>> The idea was to give a uniform calling convention without\n>>> unnecessarily confusing users by presenting a large number of\n>>> infrequently-used commands.\n>>\n>> It's not working, I'm sorry to say.\n>\n> Right, that's why I'm trying to figure out why you are hung up on the\n> low-level commands. The idea was that you wouldn't need to be  \n> exposed to\n> them at all, but obviously you were (or if you were exposed, it  \n> would be\n> in a list that was clearly marked as \"this is low-level stuff that you\n> don't really need to worry about\". So I'm trying to figure out where  \n> it\n> went wrong.\n\nI'm sorry that I can't be much help in that department.  If I really  \nknew how I ended up with that wrong impression, I probably would have  \ncorrected it already.  It's weird; git is composed of ideas that are  \nall very familiar to me (reference-counted management of immutable  \ndata, hashing, etc.) yet for me, getting to know it has been really  \ntough.  By contrast, for example, subversion was instantly  \nunderstandable when I pawed through the SVN book.\n\n>>> At any rate, it is too late to change the calling convention for\n>>> plumbing.\n>>\n>> I disagree.  You can leave the old functionality there in a  \n>> \"deprecated\"\n>> state and change the way you advertise it.\n>\n> But does that really help? It means that \"git hash-object\" is still\n> there, which I thought was the problem you had. You can argue that it\n> wouldn't be advertised to users, and so wouldn't be a problem, but  \n> that\n> is _already_ the strategy we are using. So either that strategy is  \n> fine,\n> in which case we are on the right track but may still have some work  \n> to\n> do in properly implementing it. Or it's not, in which case your  \n> proposal\n> is no better.\n\nYou've got me stumped there, I have to admit.\n\n>> It would even help a lot if the plumbing were all spelled \"git-xxx\"\n>> and the high level stuff were \"git xxx.\"\n>\n> Differentating calling conventions like that was proposed when dashed\n> forms were deprecated and removed from the PATH. But if we had dashed\n> forms for plumbing (i.e., not forwarding them via the \"git\" wrapper),\n> then you have to do one of:\n>\n>  - put them in the user's PATH. Now tab completion or looking in your\n>    PATH means you see _just_ the plumbing commands, and none of the\n>    high level ones. Which is one of the reasons they were removed from\n>    the PATH in the first place (due to numerous user complaints).\n>\n>  - put them elsewhere, and force plumbing users to add $GIT_EXEC_PATH\n>    to their PATH. That becomes very annoying for casual plumbing  \n> users.\n>    If you come to the mailing list with a problem, I would have to  \n> jump\n>    through extra hoops to ask you to show me the output of \"git\n>    ls-files\".\n\nI see your point.\n\n   llgit xxx\n\n?\n\n> Not to mention that the git wrapper does other useful things besides\n> simply exec'ing. For example, it supports --git-dir, --bare, etc.\n> So the problem is that the low-level commands _are_ still useful, and\n> many people still want to call them, just like regular git commands.\n> It's just that they are numerous and low-level, which makes them\n> daunting for new users.\n>\n> And it has become obvious over several years of the git mailing list\n> that users, once they see mention of a command, must start  \n> investigating\n> it to find out if and how it is useful. And I am not saying that is a\n> failing of users; on the contrary, I think it is quite a healthy\n> behavior on a unix-ish system. But it means that if we want not to\n> advertise low-level commands, we have to be very careful about the  \n> ways\n> in which we mention them.\n>\n> Perhaps it would make sense for each plumbing command's man page to\n> start with something like \"this is a low-level command used for\n> scripting git or investigating its internals. For high-level use, you\n> may be more interested in $X\", where $X may be \"git commit\" for\n> write-tree, commit-tree, etc. And that would at least help intercept\n> users before they get too confused.\n\nSounds like a great idea to me.\n\n>>>> [excuse me, but what the #@&*! is \"porcelainish\" supposed to mean?\n>>>> (http://www.kernel.org/pub/software/scm/git/docs/git-rev- \n>>>> parse.html)]\n>>>\n>>> Heh. That one is particularly egregious, because it rests on several\n>>> layers of git jargon. The low-level tools are plumbing, like pipes  \n>>> and\n>>> valves.\n>>\n>> ? I use the valves on my kitchen sink all the time.\n>\n> Sorry, I meant the ones under the sink, that you would use if you were\n> replacing the faucet. I would call the ones above \"taps\". But  \n> hopefully\n> you get a sense of the distinction between plumbing and porcelain.\n\n\nI know, but the point is, they're not porcelain.  They're \"plumbing  \nfixtures.\"\n\nI think UI/API works way better than porcelain/plumbing.  We are,  \nafter all, programmers.  It would also be good to link to a definition  \nany time you use a term of art in the docs.  I would even do that in  \nthe case of UI/API since the distinction could appear to be subtle.\n\nI should also say, most of the docs and interfaces I see in Git (and  \nits wrappers, web interfaces, etc.) give the SHA1 hashes way too much  \nexposure.  The times when it's actually more convenient to use a hash  \ninstead of one of the other notations are rare, and if hashes weren't  \nso exposed I bet most interfaces would make those other names more  \navailable.  One reason I think hashes retain their prominent exposure  \nis that you have no other reasonably stable way of referring to  \ncommits, since branch~NN counts backward from HEAD.  Adding such a  \nthing would help.\n\nOh, one other specific issue: the rev-parse manpage uses $GIT_DIR  \nwithout saying what it is.  I *think* that means the root of the  \nworking copy and has nothing to do with environment variables, but  \nit's hard to be sure, and if I'm right about that, it's misleading  \nnotation.\n\nSomeone needs to get gitiseasy.org/gitiseasy.net and then provide  \ncontent that lives up to the name :^)\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112240","messageId":"alpine.LNX.2.00.0904241655090.2147@iabervon.org","threadId":"19010","inReplyTo":"200904242230.13239.johan@herland.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2009-04-24T21:34:00Z","receivedAt":"2009-04-24T21:34:00Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 24 Apr 2009, Johan Herland wrote:\n\n> On Friday 24 April 2009, Michael Witten wrote:\n> > On Thu, Apr 23, 2009 at 17:51, Johan Herland <johan@herland.net> wrote:\n> > > There's also http://www.eecs.harvard.edu/~cduan/technical/git/ which I\n> > > think is a great bottom-up introduction:\n> > > - not too heavy on the concepts\n> >\n> > I really don't understand this mentality. Concepts are the only things\n> > that are important. From concepts falls all else.\n> \n> Sorry for not being clear: Concepts are indeed (and should be) important. \n> What I mean is that the concepts introduced are short and simple enough for \n> novice users to understand (without much VCS experience, if any at all). If \n> we start off _too_ detailed, we risk loosing the audience, and no one is \n> better off.\n> \n> Like Jeff King said elsewhere in this thread: We want to start a little \n> higher from the bottom. The above introduction does not focus on blobs or \n> trees, but manages to introduce Git in a useful manner by starting off with \n> only two concepts: commits and refs.\n\nI'd say that blobs and trees are an implementation detail of \"the full \ncontent of a version of the project\", not something conceptually \nimportant. Likewise, the date representation used in commits isn't \nimportant. It might be worth saying that git purposefully discards any \ninformation in your filesystem that is just incidental and not project \ncontent, like whether other users on the system where the working \ndirectory is can access your files; but a full enumeration of what the \n\"content\" and \"incidental\" categories contain can go in an appendix or \nsomething.\n\n(FWIW, git originally didn't use tree objects for subdirectories or mask\nout the g+w bit from tree entries. These weren't conceptual changes, but \nimplementation details.)\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"112241","messageId":"20090424213848.GA14493@coredump.intra.peff.net","threadId":"19010","inReplyTo":"alpine.LNX.2.00.0904241655090.2147@iabervon.org","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-24T21:38:48Z","receivedAt":"2009-04-24T21:38:48Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 24, 2009 at 05:34:00PM -0400, Daniel Barkalow wrote:\n\n> I'd say that blobs and trees are an implementation detail of \"the full \n> content of a version of the project\", not something conceptually \n> important. Likewise, the date representation used in commits isn't \n\nI disagree. I think it's important to note that trees and blobs have a\nname, and you can refer to them. Once you know that, the fact that you\ncan do:\n\n  git show master\n  git show master:Documentation\n  git show master:Makefile\n\njust makes sense. You are always just specifying an object, but the type\nis different for each (and show \"does the right thing\" based on object\ntype).\n\nNo, that isn't critical for understanding how _commit_ operations work,\nbut I think that is exactly the sort of conceptual knowledge that let\npeople use git more fully.\n\n-Peff\n"},{"id":"112247","messageId":"b4087cc50904241518w625a9890vecdd36bb937e76d5@mail.gmail.com","threadId":"19010","inReplyTo":"20090424213848.GA14493@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T22:18:44Z","receivedAt":"2009-04-24T22:18:44Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 16:38, Jeff King <peff@peff.net> wrote:\n> On Fri, Apr 24, 2009 at 05:34:00PM -0400, Daniel Barkalow wrote:\n>\n>> I'd say that blobs and trees are an implementation detail of \"the full\n>> content of a version of the project\", not something conceptually\n>> important. Likewise, the date representation used in commits isn't\n> ...\n> No, that isn't critical for understanding how _commit_ operations work,\n> but I think that is exactly the sort of conceptual knowledge that let\n> people use git more fully.\n\nI think the key conlusion here is that the main concepts are *objects*\nand references to those objects. One type of object is not necessarily\nmore low-level or high-level than another type of object; each type of\nobject is the most important type of object for a particular task in\nor view of the git world.\n\n> I disagree. I think it's important to note that trees and blobs have a\n> name, and you can refer to them. Once you know that, the fact that you\n> can do:\n>\n>  git show master\n>  git show master:Documentation\n>  git show master:Makefile\n>\n> just makes sense. You are always just specifying an object, but the type\n> is different for each (and show \"does the right thing\" based on object\n> type).\n\nIn fact, I think it's important to note that the notation:\n\n    git show master:Makefile\n\nactually involves a translation from a Unix filesystem address to a\ngit object address that is then used to find the relevant data.\n\nIn fact, I think masking this kind of thing with a catch-all word\n'reference' is a bad idea. Rather than being hidden, it should be\nexposed: I think it would be beneficial to use the word 'address'\nrather than 'reference' when talking about the SHA-1 names. Then HEAD\ncould be called a pointer variable, etc.\n\nSo, a pointer variable's value is an object address that is the\nlocation of an object in git 'memory'. I think using this approach\nwould make things significantly more transparent.\n"},{"id":"112249","messageId":"b4087cc50904241525w7de597bfq7be89796947a14cc@mail.gmail.com","threadId":"19010","inReplyTo":"b4087cc50904241518w625a9890vecdd36bb937e76d5@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T22:25:52Z","receivedAt":"2009-04-24T22:25:52Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 17:18, Michael Witten <mfwitten@gmail.com> wrote:\n> In fact, I think masking this kind of thing with a catch-all word\n> 'reference' is a bad idea. Rather than being hidden, it should be\n> exposed: I think it would be beneficial to use the word 'address'\n> rather than 'reference' when talking about the SHA-1 names. Then HEAD\n> could be called a pointer variable, etc.\n>\n> So, a pointer variable's value is an object address that is the\n> location of an object in git 'memory'. I think using this approach\n> would make things significantly more transparent.\n\nIn fact, it's not particularly important that SHA-1 is used to compute\nthe address into git memory. The only thing that's important is that\nthe address is determined by content alone (I'm not even sure that\nspecifying that the address is a cryptographically sound hash of the\ncontent is important; shouldn't that follow from the declaration that\nit must be uniquely based on content alone?); the fact that's a SHA-1\nis purely an implementation detail, and so it shouldn't appear\nprominently in the documentation.\n\nSo, what do you say?\n\nLet's start a reformation of the git terminology to use analogies that\nhave been around since the dawn of computing: 'memory', 'address', and\n'pointer'.\n"},{"id":"112253","messageId":"20090424224517.GA10155@atjola.homenet","threadId":"19010","inReplyTo":"2B5084A3-9BDB-4463-8530-3C8AB2E09A1F@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-24T22:45:17Z","receivedAt":"2009-04-24T22:45:17Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.24 17:06:27 -0400, David Abrahams wrote:\n> On Apr 24, 2009, at 4:24 PM, Jeff King wrote:\n>> On Fri, Apr 24, 2009 at 03:00:19PM -0400, David Abrahams wrote:\n>>> It would even help a lot if the plumbing were all spelled \"git-xxx\"\n>>> and the high level stuff were \"git xxx.\"\n>>\n>> Differentating calling conventions like that was proposed when dashed\n>> forms were deprecated and removed from the PATH. But if we had dashed\n>> forms for plumbing (i.e., not forwarding them via the \"git\" wrapper),\n>> then you have to do one of:\n>>\n>>  - put them in the user's PATH. Now tab completion or looking in your\n>>    PATH means you see _just_ the plumbing commands, and none of the\n>>    high level ones. Which is one of the reasons they were removed\n>>    from the PATH in the first place (due to numerous user\n>>    complaints).\n>>\n>>  - put them elsewhere, and force plumbing users to add $GIT_EXEC_PATH\n>>    to their PATH. That becomes very annoying for casual plumbing\n>>    users. If you come to the mailing list with a problem, I would\n>>    have to jump through extra hoops to ask you to show me the output\n>>    of \"git ls-files\".\n>\n> I see your point.\n>\n>   llgit xxx\n>\n> ?\n\nIf that was the exclusive way of calling the low-level commands, that\nwould still break existing scripts. And if you keep e.g. \"git\nwrite-tree\" and just add \"llgit write-tree\" as an alias, that will IMHO\njust cause more confusion once old and new git users meet. And I agree\nwith Peff, it's not important whether it's \"git foo\", \"llgit foo\", \"git\nlowlevel foo\" or something else. It's just about how much your users\nreally _need_ to know and how you tell them to use the stuff.\n\n> I think UI/API works way better than porcelain/plumbing. We are, after\n> all, programmers.\n\nWe are programmers, but not all git users are programmers.\n\n> It would also be good to link to a definition any time you use a term\n> of art in the docs. I would even do that in the case of UI/API since\n> the distinction could appear to be subtle.\n>\n> I should also say, most of the docs and interfaces I see in Git (and\n> its wrappers, web interfaces, etc.) give the SHA1 hashes way too much\n> exposure. The times when it's actually more convenient to use a hash\n> instead of one of the other notations are rare,\n\nHow often do you need a name for a commit shown by a command and can\naccept that it is not stable? I usually need a name because I\nwant to reference that commit later on, either because I need to talk to\nother users, or because I'm working on something and might need to look\nat that commit now and then, regardless on my current state of things.\nOne big exception in my workflow is when I use \"git blame\", then I\nusually just need the name once to look at the full commit. But then I\nprefer a 7-8 characters long sha-1 prefix to something like\nimprove_foo_speed~132^12~1^3. And \"pseudo-stable\" numbers have been\ndiscussed to death.\n\n> and if hashes weren't so exposed I bet most interfaces would make\n> those other names more available. One reason I think hashes retain\n> their prominent exposure is that you have no other reasonably stable\n> way of referring to commits, since branch~NN counts backward from\n> HEAD. Adding such a thing would help.\n\nIt counts backwards from \"branch\".\n\n> Oh, one other specific issue: the rev-parse manpage uses $GIT_DIR\n> without saying what it is. I *think* that means the root of the\n> working copy and has nothing to do with environment variables, but\n> it's hard to be sure, and if I'm right about that, it's misleading\n> notation.\n\n$GIT_DIR means the .git directory of a non-bare repo.\n\nBjörn\n"},{"id":"112257","messageId":"alpine.LNX.2.00.0904241852500.2147@iabervon.org","threadId":"19010","inReplyTo":"b4087cc50904241525w7de597bfq7be89796947a14cc@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2009-04-24T23:11:40Z","receivedAt":"2009-04-24T23:11:40Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 24 Apr 2009, Michael Witten wrote:\n\n> On Fri, Apr 24, 2009 at 17:18, Michael Witten <mfwitten@gmail.com> wrote:\n> > In fact, I think masking this kind of thing with a catch-all word\n> > 'reference' is a bad idea. Rather than being hidden, it should be\n> > exposed: I think it would be beneficial to use the word 'address'\n> > rather than 'reference' when talking about the SHA-1 names. Then HEAD\n> > could be called a pointer variable, etc.\n> >\n> > So, a pointer variable's value is an object address that is the\n> > location of an object in git 'memory'. I think using this approach\n> > would make things significantly more transparent.\n> \n> In fact, it's not particularly important that SHA-1 is used to compute\n> the address into git memory. The only thing that's important is that\n> the address is determined by content alone (I'm not even sure that\n> specifying that the address is a cryptographically sound hash of the\n> content is important; shouldn't that follow from the declaration that\n> it must be uniquely based on content alone?); the fact that's a SHA-1\n> is purely an implementation detail, and so it shouldn't appear\n> prominently in the documentation.\n> \n> So, what do you say?\n> \n> Let's start a reformation of the git terminology to use analogies that\n> have been around since the dawn of computing: 'memory', 'address', and\n> 'pointer'.\n\nI actually think calling them \"sha1s\" is better, simply because this bit \nof jargon doesn't mean anything else (git deals with email, so \"address\" \nis overloaded). And the term is already in use for this particular case, \nand it doesn't mean anything else at all (since, of course, the crypto \nthing is \"SHA-1\", not \"sha1\"), and it's short (which is important for \nmaking it easy to look at usage help).\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"112258","messageId":"20090424231436.GA15058@coredump.intra.peff.net","threadId":"19010","inReplyTo":"alpine.LNX.2.00.0904241852500.2147@iabervon.org","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-24T23:14:36Z","receivedAt":"2009-04-24T23:14:36Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 24, 2009 at 07:11:40PM -0400, Daniel Barkalow wrote:\n\n> > Let's start a reformation of the git terminology to use analogies that\n> > have been around since the dawn of computing: 'memory', 'address', and\n> > 'pointer'.\n> \n> I actually think calling them \"sha1s\" is better, simply because this bit \n> of jargon doesn't mean anything else (git deals with email, so \"address\" \n> is overloaded). And the term is already in use for this particular case, \n> and it doesn't mean anything else at all (since, of course, the crypto \n> thing is \"SHA-1\", not \"sha1\"), and it's short (which is important for \n> making it easy to look at usage help).\n\nJunio suggested \"object name\" in another thread, which I think is nicely\ndescriptive.\n\nFWIW, I think the pointer nomenclature has terrible connotations. I\nthink everyone who works on git groks pointers just fine, but aren't\nthey generally reviled among the progrmaming populace as the most\ncomplex and error-prone part of learning to program? Do we really need\nto increase git's reputation as complex and error-prone? ;)\n\n-Peff\n"},{"id":"112259","messageId":"20090424231632.GB10155@atjola.homenet","threadId":"19010","inReplyTo":"b4087cc50904241518w625a9890vecdd36bb937e76d5@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-24T23:16:32Z","receivedAt":"2009-04-24T23:16:32Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.24 17:18:44 -0500, Michael Witten wrote:\n> On Fri, Apr 24, 2009 at 16:38, Jeff King <peff@peff.net> wrote:\n> > On Fri, Apr 24, 2009 at 05:34:00PM -0400, Daniel Barkalow wrote:\n> >\n> >> I'd say that blobs and trees are an implementation detail of \"the full\n> >> content of a version of the project\", not something conceptually\n> >> important. Likewise, the date representation used in commits isn't\n> > ...\n> > No, that isn't critical for understanding how _commit_ operations work,\n> > but I think that is exactly the sort of conceptual knowledge that let\n> > people use git more fully.\n> \n> I think the key conlusion here is that the main concepts are *objects*\n> and references to those objects. One type of object is not necessarily\n> more low-level or high-level than another type of object; each type of\n> object is the most important type of object for a particular task in\n> or view of the git world.\n> \n> > I disagree. I think it's important to note that trees and blobs have a\n> > name, and you can refer to them. Once you know that, the fact that you\n> > can do:\n> >\n> >  git show master\n> >  git show master:Documentation\n> >  git show master:Makefile\n> >\n> > just makes sense. You are always just specifying an object, but the type\n> > is different for each (and show \"does the right thing\" based on object\n> > type).\n> \n> In fact, I think it's important to note that the notation:\n> \n>     git show master:Makefile\n> \n> actually involves a translation from a Unix filesystem address to a\n> git object address that is then used to find the relevant data.\n\nHm? Resolving master:Makefile means to first find what master is, most\nlikely the shortname for refs/heads/master. That usually references a\ncommit object (by its name). The \"<tree-ish>:<path>\" syntax then causes\ngit to lookup the tree referenced by that commit (again, by its name).\nAnd then the tree entry for \"Makefile\" is looked up, leading to the name\nfor the object identified by \"master:Makefile\".\n\n> In fact, I think masking this kind of thing with a catch-all word\n> 'reference' is a bad idea.\n\n\"master:Makefile\" is not a reference. Just \"master\" is a shortname for a\nreference, the full name might be refs/heads/master.\n\ngit has:\n - object names (which happen to be SHA-1 hashes).\n - references (which reference objects by their names)\n - symbolic references (which reference other references by their names)\n\nThe \"<tree-ish>:<path>\" syntax is not called \"reference\".\n\n> Rather than being hidden, it should be exposed: I think it would be\n> beneficial to use the word 'address' rather than 'reference' when\n> talking about the SHA-1 names. Then HEAD could be called a pointer\n> variable, etc.\n\nWhat's wrong with just calling the object name \"object name\"? References\nare something different, and the above \"master:Makefile\" is yet a\ndifferent thing, using the \"extended SHA1\" syntax to identify an object.\n\n> So, a pointer variable's value is an object address that is the\n> location of an object in git 'memory'. I think using this approach\n> would make things significantly more transparent.\n\nBut then HEAD would be a pointer pointer variable (symbolic ref), unless\nyou have a detached HEAD.\n\nBjörn\n"},{"id":"112260","messageId":"b4087cc50904241618g5f1bed21m47f0e4bd1d0ca7ff@mail.gmail.com","threadId":"19010","inReplyTo":"20090424231436.GA15058@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T23:18:50Z","receivedAt":"2009-04-24T23:18:50Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 18:14, Jeff King <peff@peff.net> wrote:\n> but aren't\n> they generally reviled among the progrmaming populace as the most\n> complex and error-prone part of learning to program?\n\nAnd now you know why people struggle with git; as I said in a previous email:\n\n    http://marc.info/?l=git&m=124022418313288&w=2\n\n    I think that the human brain struggles with indirection.\n    Consider that so many programmers have a hard time\n    understanding pointers; no wonderso many people\n    find git's underlying concepts boggling.\n\nOf course, the difference here is that we're not asking people to do\nmemory management; we have garbage collection.\n"},{"id":"112263","messageId":"alpine.LNX.2.00.0904241911590.2147@iabervon.org","threadId":"19010","inReplyTo":"20090424213848.GA14493@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2009-04-24T23:21:26Z","receivedAt":"2009-04-24T23:21:26Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 24 Apr 2009, Jeff King wrote:\n\n> On Fri, Apr 24, 2009 at 05:34:00PM -0400, Daniel Barkalow wrote:\n> \n> > I'd say that blobs and trees are an implementation detail of \"the full \n> > content of a version of the project\", not something conceptually \n> > important. Likewise, the date representation used in commits isn't \n> \n> I disagree. I think it's important to note that trees and blobs have a\n> name, and you can refer to them. Once you know that, the fact that you\n> can do:\n> \n>   git show master\n>   git show master:Documentation\n>   git show master:Makefile\n> \n> just makes sense. You are always just specifying an object, but the type\n> is different for each (and show \"does the right thing\" based on object\n> type).\n> \n> No, that isn't critical for understanding how _commit_ operations work,\n> but I think that is exactly the sort of conceptual knowledge that let\n> people use git more fully.\n\nYeah, I'll agree with that. They're good to explain as \"these are things \ngit can tell you about\", but they're not relevant to the discussion of \n\"what is history\".\n\n(And, actually, I think git has a few usability warts due to relying too \nmuch on command line arguments being objects; it would be quite nice if \n\"git blame 1a2b3c:Makefile\" worked despite this technically being \nincoherent.)\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"112264","messageId":"20090424232531.GA15136@coredump.intra.peff.net","threadId":"19010","inReplyTo":"alpine.LNX.2.00.0904241911590.2147@iabervon.org","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-24T23:25:31Z","receivedAt":"2009-04-24T23:25:31Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 24, 2009 at 07:21:26PM -0400, Daniel Barkalow wrote:\n\n> (And, actually, I think git has a few usability warts due to relying too \n> much on command line arguments being objects; it would be quite nice if \n> \"git blame 1a2b3c:Makefile\" worked despite this technically being \n> incoherent.)\n\nYeah, I think another is that \"git show master:file\" will not do CRLF or\nother filters, and \"git diff master:file other:file\" will not respect\ndiff settings. I think all of those could be solved by path lookup\nattaching a \"here is a pathname I used to get to this object\" string,\nwhich can then be accessed as appropriate.\n\nIt is not all that different conceptually than what \"git rev-list\n--objects\" does.\n\n-Peff\n"},{"id":"112265","messageId":"b4087cc50904241626h166c6b3bqa4ec714d4cb5662a@mail.gmail.com","threadId":"19010","inReplyTo":"alpine.LNX.2.00.0904241852500.2147@iabervon.org","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T23:26:06Z","receivedAt":"2009-04-24T23:26:06Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 18:11, Daniel Barkalow <barkalow@iabervon.org> wrote:\n> On Fri, 24 Apr 2009, Michael Witten wrote:\n>\n>> On Fri, Apr 24, 2009 at 17:18, Michael Witten <mfwitten@gmail.com> wrote:\n>> > In fact, I think masking this kind of thing with a catch-all word\n>> > 'reference' is a bad idea. Rather than being hidden, it should be\n>> > exposed: I think it would be beneficial to use the word 'address'\n>> > rather than 'reference' when talking about the SHA-1 names. Then HEAD\n>> > could be called a pointer variable, etc.\n>> >\n>> > So, a pointer variable's value is an object address that is the\n>> > location of an object in git 'memory'. I think using this approach\n>> > would make things significantly more transparent.\n>>\n>> In fact, it's not particularly important that SHA-1 is used to compute\n>> the address into git memory. The only thing that's important is that\n>> the address is determined by content alone (I'm not even sure that\n>> specifying that the address is a cryptographically sound hash of the\n>> content is important; shouldn't that follow from the declaration that\n>> it must be uniquely based on content alone?); the fact that's a SHA-1\n>> is purely an implementation detail, and so it shouldn't appear\n>> prominently in the documentation.\n>>\n>> So, what do you say?\n>>\n>> Let's start a reformation of the git terminology to use analogies that\n>> have been around since the dawn of computing: 'memory', 'address', and\n>> 'pointer'.\n>\n> I actually think calling them \"sha1s\" is better, simply because this bit\n> of jargon doesn't mean anything else (git deals with email, so \"address\"\n> is overloaded).\n\nI don't know if I buy that reason; the human brain is pretty good with context.\n\nI would at least like 'location' better.\n\n> And the term is already in use for this particular case,\n> and it doesn't mean anything else at all (since, of course, the crypto\n> thing is \"SHA-1\", not \"sha1\"), and it's short (which is important for\n> making it easy to look at usage help).\n\nWhat happens when SHA-1 is shown to be broken or there is a better\nalternative? Then we'll see \"sha1 for historical reasons\"... bleh!\n"},{"id":"112267","messageId":"b4087cc50904241629u76454b1chc6e84e95066a9100@mail.gmail.com","threadId":"19010","inReplyTo":"alpine.LNX.2.00.0904241911590.2147@iabervon.org","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T23:29:22Z","receivedAt":"2009-04-24T23:29:22Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 18:21, Daniel Barkalow <barkalow@iabervon.org> wrote:\n> \"git blame 1a2b3c:Makefile\" worked despite this technically being\n> incoherent.\n\nIt seems to work on my end, and it's perfectly coherent if you\nconsider git-blame to be overloaded to handle both pointers and\naddresses (or references and object names, if you prefer).\n"},{"id":"112268","messageId":"b4087cc50904241631t8913c47ke3b2027b466ee1e9@mail.gmail.com","threadId":"19010","inReplyTo":"20090424231436.GA15058@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-24T23:31:26Z","receivedAt":"2009-04-24T23:31:26Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 18:14, Jeff King <peff@peff.net> wrote:\n> Junio suggested \"object name\" in another thread, which I think is nicely\n> descriptive.\n\nThe reason I don't like \"object name\" is that \"name\" has connotations\nthat don't go well with the idea of referencing. Isn't \"address\" (or\n\"location\") better in this sense?\n"},{"id":"112269","messageId":"20090424233509.GA15341@coredump.intra.peff.net","threadId":"19010","inReplyTo":"b4087cc50904241631t8913c47ke3b2027b466ee1e9@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-24T23:35:09Z","receivedAt":"2009-04-24T23:35:09Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 24, 2009 at 06:31:26PM -0500, Michael Witten wrote:\n\n> On Fri, Apr 24, 2009 at 18:14, Jeff King <peff@peff.net> wrote:\n> > Junio suggested \"object name\" in another thread, which I think is nicely\n> > descriptive.\n> \n> The reason I don't like \"object name\" is that \"name\" has connotations\n> that don't go well with the idea of referencing. Isn't \"address\" (or\n> \"location\") better in this sense?\n\nI'm not sure I agree, but if you are concerned with \"name\", then I think\nsomething like \"object id\" or \"object identifier\" would probably be\nbetter. \"address\" and \"location\" imply to me that they are part of a\ncontiguous set. And while technically they may be considered addresses\nof a sparse 2^160 array, I'm not sure that explanation is really helping\nnew users understand what is going on.\n\nWhat the user really cares about is that it is persistent and\nunambiguous.\n\n-Peff\n"},{"id":"112271","messageId":"b4087cc50904241701jb78ce50m122fef475b0f1de7@mail.gmail.com","threadId":"19010","inReplyTo":"20090424231632.GB10155@atjola.homenet","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-25T00:01:48Z","receivedAt":"2009-04-25T00:01:48Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"2009/4/24 Björn Steinbrink <B.Steinbrink@gmx.de>:\n>> In fact, I think it's important to note that the notation:\n>>\n>>     git show master:Makefile\n>>\n>> actually involves a translation from a Unix filesystem address to a\n>> git object address that is then used to find the relevant data.\n>\n> Hm? Resolving master:Makefile means to first find what master is, most\n> likely the shortname for refs/heads/master. That usually references a\n> commit object (by its name). The \"<tree-ish>:<path>\" syntax then causes\n> git to lookup the tree referenced by that commit (again, by its name).\n> And then the tree entry for \"Makefile\" is looked up, leading to the name\n> for the object identified by \"master:Makefile\".\n\nFirstly, your head is too bound to low-level implementation.\n\nSecondly, you've basically just expounded upon what I said. The\nMakefile part is for humans to write using a filesystem path (address)\nthat is mapped into what I call a git address. The point is that the\nuser is interfacing between two theories of content storage.\n\n>> In fact, I think masking this kind of thing with a catch-all word\n>> 'reference' is a bad idea.\n>\n> \"master:Makefile\" is not a reference. Just \"master\" is a shortname for a\n> reference, the full name might be refs/heads/master.\n>\n> git has:\n>  - object names (which happen to be SHA-1 hashes).\n>  - references (which reference objects by their names)\n>  - symbolic references (which reference other references by their names)\n>\n> The \"<tree-ish>:<path>\" syntax is not called \"reference\".\n\nI will admit that I used this term wrongly then, and that git has a\nset of terminologies much closer to what I proposed:\n\n    * object addresses: object names\n    * pointers: references\n    * handle: symbolic reference (I don't know, I just now made that one up)\n\nI was under the impression that object names were in fact called\nreferences and that things like '[refs/heads/]master' were just\nconsidered conveniences. I'm glad to have been disabused; though I\nlike my terms better ;-D\n\n>> Rather than being hidden, it should be exposed: I think it would be\n>> beneficial to use the word 'address' rather than 'reference' when\n>> talking about the SHA-1 names. Then HEAD could be called a pointer\n>> variable, etc.\n>\n> What's wrong with just calling the object name \"object name\"?\n\nWhat's wrong with calling the object address \"object address\"?\n\nAs I've stated: \"address\", \"pointer\", and \"handle\" are an analogy to\nterminology that has been around for ages. In fact, another name for\n\"pointer\" is \"reference\".\n\n> are something different, and the above \"master:Makefile\" is yet a\n> different thing, using the \"extended SHA1\" syntax to identify an object.\n\nIt is certainly something different. It's an interface between\ntheories of content storage.\n\n>> So, a pointer variable's value is an object address that is the\n>> location of an object in git 'memory'. I think using this approach\n>> would make things significantly more transparent.\n>\n> But then HEAD would be a pointer pointer variable (symbolic ref), unless\n> you have a detached HEAD.\n\nWe call those handles.\n"},{"id":"112273","messageId":"4E155CC5-B20A-4B79-8CBF-9D1E0E36920F@boostpro.com","threadId":"19010","inReplyTo":"20090424213848.GA14493@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-25T00:19:18Z","receivedAt":"2009-04-25T00:19:18Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 24, 2009, at 5:38 PM, Jeff King wrote:\n\n> On Fri, Apr 24, 2009 at 05:34:00PM -0400, Daniel Barkalow wrote:\n>\n>> I'd say that blobs and trees are an implementation detail of \"the  \n>> full\n>> content of a version of the project\", not something conceptually\n>> important. Likewise, the date representation used in commits isn't\n>\n> I disagree. I think it's important to note that trees and blobs have a\n> name, and you can refer to them. Once you know that, the fact that you\n> can do:\n>\n>  git show master\n>  git show master:Documentation\n>  git show master:Makefile\n>\n> just makes sense. You are always just specifying an object, but the  \n> type\n> is different for each (and show \"does the right thing\" based on object\n> type).\n\n\nI don't believe you need to know about trees and blobs to make sense  \nof that.  Those are just directories and files.  The whole idea that  \ntrees are a more-general thing that could be used to represent  \nsomething other than directory structure and blobs could be used to  \nrepresent something other than file contents is way below most  \npeoples' need-to-know threshold.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112274","messageId":"b4087cc50904241719w64bc6074xe5b8d341ef9f51ed@mail.gmail.com","threadId":"19010","inReplyTo":"20090424233509.GA15341@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-25T00:19:50Z","receivedAt":"2009-04-25T00:19:50Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 18:35, Jeff King <peff@peff.net> wrote:\n> On Fri, Apr 24, 2009 at 06:31:26PM -0500, Michael Witten wrote:\n>\n>> On Fri, Apr 24, 2009 at 18:14, Jeff King <peff@peff.net> wrote:\n>> > Junio suggested \"object name\" in another thread, which I think is nicely\n>> > descriptive.\n>>\n>> The reason I don't like \"object name\" is that \"name\" has connotations\n>> that don't go well with the idea of referencing. Isn't \"address\" (or\n>> \"location\") better in this sense?\n>\n> I'm not sure I agree, but if you are concerned with \"name\", then I think\n> something like \"object id\" or \"object identifier\" would probably be\n> better. \"address\" and \"location\" imply to me that they are part of a\n> contiguous set. And while technically they may be considered addresses\n> of a sparse 2^160 array, I'm not sure that explanation is really helping\n> new users understand what is going on.\n\nYou make an interesting point about implied contiguousness, but I\ndon't think any git operation is in danger of evoking that thought. I\nmainly like the idea of \"address\" and \"location\", because they go\nextremely well with \"pointer\", \"handle\" and the idea of a \"git store\n(memory)\". Most importantly, this is an analogy that has been around a\nlong time.\n"},{"id":"112275","messageId":"b4087cc50904241726l1b38495apf3df4d8e10254902@mail.gmail.com","threadId":"19010","inReplyTo":"4E155CC5-B20A-4B79-8CBF-9D1E0E36920F@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-25T00:26:32Z","receivedAt":"2009-04-25T00:26:32Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Fri, Apr 24, 2009 at 19:19, David Abrahams <dave@boostpro.com> wrote:\n>>  git show master\n>>  git show master:Documentation\n>>  git show master:Makefile\n>>\n>> just makes sense. You are always just specifying an object, but the type\n>> is different for each (and show \"does the right thing\" based on object\n>> type).\n>\n> I don't believe you need to know about trees and blobs to make sense of\n> that.  Those are just directories and files.\n\nI still think the key is that commits and blobs and trees are all\nobjects, and the important things are the concepts of objects, object\naddresses, object pointers, and handles (or, what everyone else calls\nobjects, object names, references, and symbolic references).\n\nAlso, you've mixed in the theory of file system addressing in with the\ntheory of git addressing. I think it's important to realize that the\ntool 'git show' is actually providing a translation between the two\nworlds. There's not really any need for paths to be considered a\nfundamental git concept; simply, git tools know how to translate\nbetween both worlds.\n"},{"id":"112276","messageId":"20090425003531.GA18125@coredump.intra.peff.net","threadId":"19010","inReplyTo":"4E155CC5-B20A-4B79-8CBF-9D1E0E36920F@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-25T00:35:31Z","receivedAt":"2009-04-25T00:35:31Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 24, 2009 at 08:19:18PM -0400, David Abrahams wrote:\n\n>>  git show master\n>>  git show master:Documentation\n>>  git show master:Makefile\n>>\n> I don't believe you need to know about trees and blobs to make sense of \n> that.  Those are just directories and files.  The whole idea that trees \n> are a more-general thing that could be used to represent something other \n> than directory structure and blobs could be used to represent something \n> other than file contents is way below most peoples' need-to-know \n> threshold.\n\nActually, it is not the generally of trees that I think is interesting\nthere, but the generality of _objects_. That is, each of those things is\na first-class object, and has a unique name by which it can be referred.\nThe examples above are just _one_ of the ways you can refer to the same\nobjects.\n\n-Peff\n"},{"id":"112277","messageId":"78D97574-74AB-4A4D-AEB2-874BFBB4345E@boostpro.com","threadId":"19010","inReplyTo":"20090424224517.GA10155@atjola.homenet","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-25T00:39:12Z","receivedAt":"2009-04-25T00:39:12Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 24, 2009, at 6:45 PM, Björn Steinbrink wrote:\n\n>> I think UI/API works way better than porcelain/plumbing. We are,  \n>> after\n>> all, programmers.\n>\n> We are programmers, but not all git users are programmers.\n\nI'm sure you will admit that the vast majority are programmers.  This  \nis about speaking effectively to your primary audience.\n\n>> It would also be good to link to a definition any time you use a term\n>> of art in the docs. I would even do that in the case of UI/API since\n>> the distinction could appear to be subtle.\n>>\n>> I should also say, most of the docs and interfaces I see in Git (and\n>> its wrappers, web interfaces, etc.) give the SHA1 hashes way too much\n>> exposure. The times when it's actually more convenient to use a hash\n>> instead of one of the other notations are rare,\n>\n> How often do you need a name for a commit shown by a command and can\n> accept that it is not stable?\n\nI can accept it as long as it's stable inside my own repo.  Maybe I  \nneed the SHA1 to talk about it wherever it may roam.  I think you  \ncould count in the other direction (i.e. from the roots instead of the  \nleaves) to get fairly stable symbolic names.\n\nAlso, I don't think I need to see the hashes for trees and blobs most  \nof the time.\n\n> I usually need a name because I\n> want to reference that commit later on, either because I need to  \n> talk to\n> other users, or because I'm working on something and might need to  \n> look\n> at that commit now and then, regardless on my current state of things.\n> One big exception in my workflow is when I use \"git blame\", then I\n> usually just need the name once to look at the full commit. But then I\n> prefer a 7-8 characters long sha-1 prefix to something like\n> improve_foo_speed~132^12~1^3. And \"pseudo-stable\" numbers have been\n> discussed to death.\n\nOkay, I \"say uncle.\"\n\n>> and if hashes weren't so exposed I bet most interfaces would make\n>> those other names more available. One reason I think hashes retain\n>> their prominent exposure is that you have no other reasonably stable\n>> way of referring to commits, since branch~NN counts backward from\n>> HEAD. Adding such a thing would help.\n>\n> It counts backwards from \"branch\".\n\nRight, thanks.\n\n>> Oh, one other specific issue: the rev-parse manpage uses $GIT_DIR\n>> without saying what it is. I *think* that means the root of the\n>> working copy and has nothing to do with environment variables, but\n>> it's hard to be sure, and if I'm right about that, it's misleading\n>> notation.\n>\n> $GIT_DIR means the .git directory of a non-bare repo.\n\n\nThanks for clarifying.  But don't neglect to fix the docs so the next  \nguy doesn't have to ask ;-)\n\nBTW, \"[non-]bare repo\" is yet another Git-specific jargon.  I know  \nwhat it means... again, only because I asked someone.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112278","messageId":"BD5D0ECC-5540-42BF-811D-040925225881@boostpro.com","threadId":"19010","inReplyTo":"alpine.LNX.2.00.0904241852500.2147@iabervon.org","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-25T00:41:12Z","receivedAt":"2009-04-25T00:41:12Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 24, 2009, at 7:11 PM, Daniel Barkalow wrote:\n\n> I actually think calling them \"sha1s\" is better, simply because this  \n> bit\n> of jargon doesn't mean anything else (git deals with email, so  \n> \"address\"\n> is overloaded). And the term is already in use for this particular  \n> case,\n> and it doesn't mean anything else at all (since, of course, the crypto\n> thing is \"SHA-1\", not \"sha1\"), and it's short (which is important for\n> making it easy to look at usage help).\n\n\nThe word \"hash\" would be an improvement.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112279","messageId":"B8EE4B7B-CBCE-47D6-97B2-2F505A218B58@boostpro.com","threadId":"19010","inReplyTo":"b4087cc50904241701jb78ce50m122fef475b0f1de7@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-25T00:48:57Z","receivedAt":"2009-04-25T00:48:57Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 24, 2009, at 8:01 PM, Michael Witten wrote:\n\n>> What's wrong with just calling the object name \"object name\"?\n>\n> What's wrong with calling the object address \"object address\"?\n\n\nNeither captures the connection to the object's contents.  I think  \n\"value ID\" would be closer, but it's probably too horrible.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112280","messageId":"1A9F6DB0-983F-4A5B-B3B7-33227C11F36A@boostpro.com","threadId":"19010","inReplyTo":"20090425003531.GA18125@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-25T00:53:37Z","receivedAt":"2009-04-25T00:53:37Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 24, 2009, at 8:35 PM, Jeff King wrote:\n\n> On Fri, Apr 24, 2009 at 08:19:18PM -0400, David Abrahams wrote:\n>\n>>> git show master\n>>> git show master:Documentation\n>>> git show master:Makefile\n>>>\n>> I don't believe you need to know about trees and blobs to make  \n>> sense of\n>> that.  Those are just directories and files.  The whole idea that  \n>> trees\n>> are a more-general thing that could be used to represent something  \n>> other\n>> than directory structure and blobs could be used to represent  \n>> something\n>> other than file contents is way below most peoples' need-to-know\n>> threshold.\n>\n> Actually, it is not the generally of trees that I think is interesting\n> there, but the generality of _objects_. That is, each of those  \n> things is\n> a first-class object, and has a unique name by which it can be  \n> referred.\n\n\nI'm sorry, but I think most people would find that so unremarkable  \nthat making a big deal about it would lead to \"what am I missing here\"  \nconfusion.  Maybe a person who's exclusively used CVS (or older)  \ntechnologies before coming to Git would be happy to know that, but  \nit's sort of obvious.  In CVS the lack of first-class directories  \nsticks out like a sore thumb.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112289","messageId":"94a0d4530904250318w7f368ea6hb5d59558aedc5c4f@mail.gmail.com","threadId":"19010","inReplyTo":"20090424231436.GA15058@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2009-04-25T10:18:01Z","receivedAt":"2009-04-25T10:18:01Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sat, Apr 25, 2009 at 2:14 AM, Jeff King <peff@peff.net> wrote:\n> On Fri, Apr 24, 2009 at 07:11:40PM -0400, Daniel Barkalow wrote:\n>\n>> > Let's start a reformation of the git terminology to use analogies that\n>> > have been around since the dawn of computing: 'memory', 'address', and\n>> > 'pointer'.\n>>\n>> I actually think calling them \"sha1s\" is better, simply because this bit\n>> of jargon doesn't mean anything else (git deals with email, so \"address\"\n>> is overloaded). And the term is already in use for this particular case,\n>> and it doesn't mean anything else at all (since, of course, the crypto\n>> thing is \"SHA-1\", not \"sha1\"), and it's short (which is important for\n>> making it easy to look at usage help).\n>\n> Junio suggested \"object name\" in another thread, which I think is nicely\n> descriptive.\n\nIt's not a name, it's an identification, so how about \"id\"? You have\ntree ids, commit ids, blob ids, and so on.\n\n-- \nFelipe Contreras\n"},{"id":"112290","messageId":"94a0d4530904250335w7a88c4d6lafd12a8ce6f98127@mail.gmail.com","threadId":"19010","inReplyTo":"20090424185208.GM17365@fieldses.org","subject":"Re: [doc] User Manual Suggestion","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2009-04-25T10:35:38Z","receivedAt":"2009-04-25T10:35:38Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Fri, Apr 24, 2009 at 9:52 PM, J. Bruce Fields <bfields@fieldses.org> wrote:\n> That would be great, thanks.  Several people have gone off and posted\n> their own tutorials someplace, and that's fine, but it would be\n> especially helpful if you could contribute to the actual Documentation/\n> directory.  That may mean arguing with people and making compromises.\n> But it also means the results will be distributed with git, will be\n> integrated with other git documentation, and will get first-class\n> technical review.\n>\n> I'd also encourage incrementally improving existing documentation where\n> possible instead of starting over from scratch.  But having broken that\n> rule myself a couple times I'm hardly in a position to insist.  If you\n> must start over, at least think about how to replace or fit it in with\n> existing documentation.\n\nPeople will continue to write git documentation from scratch because\nthere is a huge gap from the top-bottom approach to a point where you\nactually \"get git\", and people are trying to find short-cuts so that\nother people can really get it too.\n\nI spent years using git simply repeating the templates I had seen in\nmultiple places until I stumbled upon \"git from the bottom up\" and\nthen I finally understood the beauty and simplicity of git's design.\nFrom that point I understood why many command didn't do what I\nexpected.\n\nNote that \"bottom\" doesn't mean plumbing, the \"plumbing\" is usually\nreferred to the git.git tools, but you can work with git low-level\nobjects through your own implementation as people like Scott Chacon\nhave indeed done (git-ruby). \"bottom\" then means git basic building\nblocks: blobs, trees, commits, refs.\n\nIdeally the UI should expose the basic concepts of git, but instead\nits is hiding them, so no wonder people *need* special documentation\nto 'understand git conceptually', or learn 'git from the bottom up',\netc.\n\n-- \nFelipe Contreras\n"},{"id":"112318","messageId":"alpine.LNX.2.00.0904251445030.2147@iabervon.org","threadId":"19010","inReplyTo":"b4087cc50904241626h166c6b3bqa4ec714d4cb5662a@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2009-04-25T18:55:36Z","receivedAt":"2009-04-25T18:55:36Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 24 Apr 2009, Michael Witten wrote:\n\n> > And the term is already in use for this particular case,\n> > and it doesn't mean anything else at all (since, of course, the crypto\n> > thing is \"SHA-1\", not \"sha1\"), and it's short (which is important for\n> > making it easy to look at usage help).\n> \n> What happens when SHA-1 is shown to be broken or there is a better\n> alternative? Then we'll see \"sha1 for historical reasons\"... bleh!\n\nWhy do you think SHA-1 has anything to do with it? Git's sha1s could just \nas easily be 160 bits of a SHA-256 hash and there wouldn't be any \nuser-visible difference. The term doesn't imply any particular significant \nconnection to a particular algorithm. It could be like \"pencil lead\", \nwhich has never been made of lead, but is called that for no particularly \nimportant reason.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"112319","messageId":"b4087cc50904251216p617e347bmdc70e109298fa9b2@mail.gmail.com","threadId":"19010","inReplyTo":"alpine.LNX.2.00.0904251445030.2147@iabervon.org","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-25T19:16:50Z","receivedAt":"2009-04-25T19:16:50Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Sat, Apr 25, 2009 at 13:55, Daniel Barkalow <barkalow@iabervon.org> wrote:\n> On Fri, 24 Apr 2009, Michael Witten wrote:\n>\n>> > And the term is already in use for this particular case,\n>> > and it doesn't mean anything else at all (since, of course, the crypto\n>> > thing is \"SHA-1\", not \"sha1\"), and it's short (which is important for\n>> > making it easy to look at usage help).\n>>\n>> What happens when SHA-1 is shown to be broken or there is a better\n>> alternative? Then we'll see \"sha1 for historical reasons\"... bleh!\n>\n> Why do you think SHA-1 has anything to do with it?\n\nWell, it's named sha1.\n\n> Git's sha1s could just\n> as easily be 160 bits of a SHA-256 hash and there wouldn't be any\n> user-visible difference. The term doesn't imply any particular significant\n> connection to a particular algorithm.\n\nThen give it a generic name like 'hash'.\n\n> It could be like \"pencil lead\", which has never been made of lead,\n> but is called that for no particularly important reason.\n\nHence the perennial:\n\n    \"Hey! Did you know that pencil lead isn't lead at all?\"\n\nto which someone might respond:\n\n    \"Why do you think lead has anything to do with it?\"\n\nLook familiar?\n"},{"id":"112320","messageId":"94a0d4530904251224g6b228448q276436f17f7e5cc3@mail.gmail.com","threadId":"19010","inReplyTo":"b4087cc50904251216p617e347bmdc70e109298fa9b2@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2009-04-25T19:24:22Z","receivedAt":"2009-04-25T19:24:22Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sat, Apr 25, 2009 at 10:16 PM, Michael Witten <mfwitten@gmail.com> wrote:\n> On Sat, Apr 25, 2009 at 13:55, Daniel Barkalow <barkalow@iabervon.org> wrote:\n>> On Fri, 24 Apr 2009, Michael Witten wrote:\n>>\n>>> > And the term is already in use for this particular case,\n>>> > and it doesn't mean anything else at all (since, of course, the crypto\n>>> > thing is \"SHA-1\", not \"sha1\"), and it's short (which is important for\n>>> > making it easy to look at usage help).\n>>>\n>>> What happens when SHA-1 is shown to be broken or there is a better\n>>> alternative? Then we'll see \"sha1 for historical reasons\"... bleh!\n>>\n>> Why do you think SHA-1 has anything to do with it?\n>\n> Well, it's named sha1.\n>\n>> Git's sha1s could just\n>> as easily be 160 bits of a SHA-256 hash and there wouldn't be any\n>> user-visible difference. The term doesn't imply any particular significant\n>> connection to a particular algorithm.\n>\n> Then give it a generic name like 'hash'.\n\nFor most purposes in the documentation sha1's are used as ids, so why\ndon't use \"id\" instead? Like 'commit id'. The fact that the id is also\na hash sum is hardly relevant for the user.\n\n-- \nFelipe Contreras\n"},{"id":"112321","messageId":"E85677CA-FA7E-4777-97DF-9B295E89B83A@boostpro.com","threadId":"19010","inReplyTo":"94a0d4530904251224g6b228448q276436f17f7e5cc3@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-25T19:36:24Z","receivedAt":"2009-04-25T19:36:24Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 25, 2009, at 3:24 PM, Felipe Contreras wrote:\n\n> On Sat, Apr 25, 2009 at 10:16 PM, Michael Witten  \n> <mfwitten@gmail.com> wrote:\n>> On Sat, Apr 25, 2009 at 13:55, Daniel Barkalow  \n>> <barkalow@iabervon.org> wrote:\n>>> On Fri, 24 Apr 2009, Michael Witten wrote:\n>>>\n>>>>> And the term is already in use for this particular case,\n>>>>> and it doesn't mean anything else at all (since, of course, the  \n>>>>> crypto\n>>>>> thing is \"SHA-1\", not \"sha1\"), and it's short (which is  \n>>>>> important for\n>>>>> making it easy to look at usage help).\n>>>>\n>>>> What happens when SHA-1 is shown to be broken or there is a better\n>>>> alternative? Then we'll see \"sha1 for historical reasons\"... bleh!\n>>>\n>>> Why do you think SHA-1 has anything to do with it?\n>>\n>> Well, it's named sha1.\n>>\n>>> Git's sha1s could just\n>>> as easily be 160 bits of a SHA-256 hash and there wouldn't be any\n>>> user-visible difference. The term doesn't imply any particular  \n>>> significant\n>>> connection to a particular algorithm.\n>>\n>> Then give it a generic name like 'hash'.\n>\n> For most purposes in the documentation sha1's are used as ids, so why\n> don't use \"id\" instead? Like 'commit id'. The fact that the id is also\n> a hash sum is hardly relevant for the user.\n\n\nWhere it's relevant when the user notices that two distinct files have  \nthe same id (because they happen to have the same contents) and  \nwonders what's up.\n\nIt's not a foregone conclusion that objects with the same value have  \nidentical ids, but it's immediately apparent if the id is known to be  \na hash.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112322","messageId":"94a0d4530904251353p1e00e015o67825c8726817894@mail.gmail.com","threadId":"19010","inReplyTo":"E85677CA-FA7E-4777-97DF-9B295E89B83A@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2009-04-25T20:53:54Z","receivedAt":"2009-04-25T20:53:54Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sat, Apr 25, 2009 at 10:36 PM, David Abrahams <dave@boostpro.com> wrote:\n>\n> On Apr 25, 2009, at 3:24 PM, Felipe Contreras wrote:\n> Where it's relevant when the user notices that two distinct files have the\n> same id (because they happen to have the same contents) and wonders what's\n> up.\n>\n> It's not a foregone conclusion that objects with the same value have\n> identical ids, but it's immediately apparent if the id is known to be a\n> hash.\n\nThat's true.\n\nhash +1\n\n-- \nFelipe Contreras\n"},{"id":"112394","messageId":"20090426112802.GC10155@atjola.homenet","threadId":"19010","inReplyTo":"E85677CA-FA7E-4777-97DF-9B295E89B83A@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-26T11:28:02Z","receivedAt":"2009-04-26T11:28:02Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.25 15:36:24 -0400, David Abrahams wrote:\n> Where it's relevant when the user notices that two distinct files have  \n> the same id (because they happen to have the same contents) and wonders \n> what's up.\n\nWhy would the user have to care about the object files in the repo? And\nwhy would your implementation save the same object twice, in two\ndistinct files? The SHA-1 hash is created from the object, that means\nthe its type, size and data. It's not an id of a file in the working\ntree, but of an object.\n\n> It's not a foregone conclusion that objects with the same value have  \n> identical ids, but it's immediately apparent if the id is known to be a \n> hash.\n\nYou can't have two objects with the same contents to begin with, same\ncontent => same object.  You can just have that one object stored\nmultiple times in different places (for sane implementations this likely\nmeans that you have more than one repo to look at, and each has its own\ncopy of that object, but that's nothing you as an user should have to\ncare about).\n\nIt's an identity relation: same name/id => same object. Unlike e.g. a\nhash-table where you are expected to deal with collisions, and having\nthe same hash doesn't mean that you have identical data. But that's not\ntrue of git, it expects an identity relation, which is IMHO better\nexpressed through \"object name\" or \"object id\". You can still say that\nthe name/id is generated by using a hash function, but the important\npart is that the name/id is used to _uniquely_ identify an object, which\nisn't apparent when you call it a hash.\n\nBjörn\n"},{"id":"112388","messageId":"FA5C0FFA-2DCD-4DF1-9A94-C2A26A9DCAE9@boostpro.com","threadId":"19010","inReplyTo":"20090426112802.GC10155@atjola.homenet","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-26T13:55:34Z","receivedAt":"2009-04-26T13:55:34Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 26, 2009, at 7:28 AM, Björn Steinbrink wrote:\n\n> On 2009.04.25 15:36:24 -0400, David Abrahams wrote:\n>> Where it's relevant when the user notices that two distinct files  \n>> have\n>> the same id (because they happen to have the same contents) and  \n>> wonders\n>> what's up.\n>\n> Why would the user have to care about the object files in the repo?\n\nWhat a strange question.  I have no idea how to answer.  It seems self- \nevident to me that users of a VCS care that their files are stored in  \nit.\n\n> And\n> why would your implementation save the same object twice, in two\n> distinct files?\n\nOne could easily have the expectation that contents can be duplicated  \nbecause there are numerous precedents in everyone's experience of  \ncomputing, for example in filesystems and in any programming language  \nthat is not pure-functional.\n\n> The SHA-1 hash is created from the object, that means\n> the its type, size and data. It's not an id of a file in the working\n> tree, but of an object\n\nAll true.  All somewhat subtle distinctions that are not nearly as  \napparent unless you actually use the word \"hash\" as I have been  \nadvocating.\n\n>> It's not a foregone conclusion that objects with the same value have\n>> identical ids, but it's immediately apparent if the id is known to  \n>> be a\n>> hash.\n>\n> You can't have two objects with the same contents to begin with, same\n> content => same object.\n\nIn the Git world, I agree.  In general, I disagree.  The fact that is  \nso in the Git world is reinforced by the notion that the id of an  \nobject is a hash of its contents.\n\n> You can just have that one object stored\n> multiple times in different places (for sane implementations this  \n> likely\n> means that you have more than one repo to look at, and each has its  \n> own\n> copy of that object, but that's nothing you as an user should have to\n> care about).\n\n> It's an identity relation: same name/id => same object. Unlike e.g. a\n> hash-table where you are expected to deal with collisions, and having\n> the same hash doesn't mean that you have identical data.  But that's  \n> not\n> true of git, it expects an identity relation, which is IMHO better\n> expressed through \"object name\" or \"object id\".\n\nYes, that's true in the Git world (though not necessarily elsewhere),  \nor at least you hope it is.  In fact, there's no guarantee that SHA1  \ncollisions won't occur; it's just exremely unlikely.  In fact, if you  \ngoogle it you can find some interesting papers about SHA1 collision.\n\nAnother way to express what you wrote above:\n\n    same same id => same hash ?=> same contents => same object\n\nwhere ?=> means \"almost certainly implies.\"  What you left out was the  \nimplication in the other direction, which is a true guarantee at all  \nsteps, and \"hash\" is well-understood to mean\n\n    same contents => same hash\n\n> You can still say that\n> the name/id is generated by using a hash function, but the important\n> part is that the name/id is used to _uniquely_ identify an object,  \n> which\n> isn't apparent when you call it a hash.\n\n\nI think the implication is important in both directions.  Neither one  \nis self-evident to a new user.  Maybe the right answer is 'hash id'.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112398","messageId":"b4087cc50904260936t4aa05c92ldab3b8ce8c9bd33a@mail.gmail.com","threadId":"19010","inReplyTo":"20090426112802.GC10155@atjola.homenet","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-26T16:36:04Z","receivedAt":"2009-04-26T16:36:04Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"2009/4/26 Björn Steinbrink <B.Steinbrink@gmx.de>:\n> On 2009.04.25 15:36:24 -0400, David Abrahams wrote:\n>> Where it's relevant when the user notices that two distinct files have\n>> the same id (because they happen to have the same contents) and wonders\n>> what's up.\n>...\n> And why would your implementation save the same object twice, in two\n> distinct files?\n\nThis question makes me think that you don't understand the parent's\npoint. He's not talking about implementation details; in fact, there's\nno reason to mix the git world and the file system world at all in\nthis discussion.\n\nDavid is pointing out that a user might notice that two different\ntrees list the same blob. This can be startling if you have incomplete\npicture about what's going on.\n\nFrom a practical point of view, you might argue that not too many\npeople are looking at trees and blobs; however, it seems to me that\nmost people are afraid to use any of git's most useful features\nprecisely because they don't understand the git model and they don't\nunderstand that nothing is ever lost unless you explicitly clean up\nunreferenced objects---they don't see how easy it is manipulate their\nrepos. I argue that if they are given the full knowledge of git's\nconcepts, then they will be able to reason about their repo actions\nwith confidence, even if they only work with commits.\n\nI think the key is to stress in the documentation the idea that there\nare 2 separate worlds (the git object world and the working\ndirectory's file system world) and that the git tools provide an\ninterface between them; this seems like a small and unnecessarily\nacademic point, but I believe that it's important to working with\nconfidence.\n\n> ...\n> You can't have two objects with the same contents to begin with, same\n> content => same object.  You can just have that one object stored\n> multiple times in different places (for sane implementations this likely\n> means that you have more than one repo to look at, and each has its own\n> copy of that object, but that's nothing you as an user should have to\n> care about).\n\nIndeed it's nothing you should care about. It's an implementation\ndetail again; theoretically, every repo is in the same git world where\nall git objects are stored---in a sense, a particular repo state is\nitself an object of this world.\n\n> It's an identity relation: same name/id => same object. Unlike e.g. a\n> hash-table where you are expected to deal with collisions, and having\n> the same hash doesn't mean that you have identical data.\n\nHowever, having the same *cryptographic* hash does mean that you have\nidentical data.\n\nThe overall point is this: The documentation should force people to\nlearn the right ideas, so that they can have confidence to think\nbeyond blog-post templates for using git.\n"},{"id":"112375","messageId":"20090426175613.GA4942@atjola.homenet","threadId":"19010","inReplyTo":"FA5C0FFA-2DCD-4DF1-9A94-C2A26A9DCAE9@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-26T17:56:13Z","receivedAt":"2009-04-26T17:56:13Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.26 09:55:34 -0400, David Abrahams wrote:\n>\n> On Apr 26, 2009, at 7:28 AM, Björn Steinbrink wrote:\n>\n>> On 2009.04.25 15:36:24 -0400, David Abrahams wrote:\n>>> Where it's relevant when the user notices that two distinct files\n>>> have the same id (because they happen to have the same contents) and\n>>> wonders what's up.\n>>\n>> Why would the user have to care about the object files in the repo?\n>\n> What a strange question. I have no idea how to answer. It seems\n> self- evident to me that users of a VCS care that their files are\n> stored in it.\n\n_Their_ files. The files that come from/end up in the working tree. I\ncared about those when I used SVN, too. But I never went to the SVN repo\nto find out if there are two equal files in it. We're talking about\nobject names, and those belong to objects, not files in the working\ntree.\n\n>> And why would your implementation save the same object twice, in two\n>> distinct files?\n>\n> One could easily have the expectation that contents can be duplicated  \n> because there are numerous precedents in everyone's experience of  \n> computing, for example in filesystems and in any programming language  \n> that is not pure-functional.\n\nThat's not answering my question. I asked why you come up with an\nimplementation that is \"broken\" enough to save the same object twice\nwith different file names.  If the implementation does not do that, your\n\"when the user notices that two distinct files has the same id\" is\nimmediately invalid. The user cannot come into that situation then. And\nanyway, when the user notices something, that's a discovery, not an\nexpectation.\n\n>> The SHA-1 hash is created from the object, that means\n>> the its type, size and data. It's not an id of a file in the working\n>> tree, but of an object\n>\n> All true.  All somewhat subtle distinctions that are not nearly as  \n> apparent unless you actually use the word \"hash\" as I have been  \n> advocating.\n\nHu? How does saying \"object hash\" instead of \"object id\" make it any\nmore apparent that a file in the working tree is something else than a\ngit object?\n\n>>> It's not a foregone conclusion that objects with the same value have\n>>> identical ids, but it's immediately apparent if the id is known to  \n>>> be a\n>>> hash.\n>>\n>> You can't have two objects with the same contents to begin with, same\n>> content => same object.\n>\n> In the Git world, I agree.  In general, I disagree.\n\nI don't think were discussing a term to describe something that\nidentifies an object in general. So, \"in general\" you can disagree as\nmuch as you want, but for git that doesn't matter at all.\n\n> The fact that is so in the Git world is reinforced by the notion that\n> the id of an object is a hash of its contents.\n>\n>> You can just have that one object stored multiple times in different\n>> places (for sane implementations this  likely means that you have\n>> more than one repo to look at, and each has its  own copy of that\n>> object, but that's nothing you as an user should have to care about).\n>\n>> It's an identity relation: same name/id => same object. Unlike e.g. a\n>> hash-table where you are expected to deal with collisions, and having\n>> the same hash doesn't mean that you have identical data.  But that's\n>> not true of git, it expects an identity relation, which is IMHO\n>> better expressed through \"object name\" or \"object id\".\n>\n> Yes, that's true in the Git world (though not necessarily elsewhere), or \n> at least you hope it is.  In fact, there's no guarantee that SHA1  \n> collisions won't occur; it's just exremely unlikely.  In fact, if you  \n> google it you can find some interesting papers about SHA1 collision.\n\nSure, it's an assumption that has been made and is required to hold true\nfor git to work.\n\n> Another way to express what you wrote above:\n>\n>    same same id => same hash ?=> same contents => same object\n>\n> where ?=> means \"almost certainly implies.\"\n\nNo, that chain shows how git could be \"unreliable\" when you get hash\ncollisions. You could put that into a chapter that explains the\nimplications of the way git generates its object ids. But it's not very\ninteresting when you use git and (implicitly) trust the assumption that\nno collisions happen.\n\nFor that case, you need a different chain:\n\nsame name/id ==> same object ==> same content\n\nThat's interesting when you e.g. want to \"access\" some object or when\nyou look at a tree that references the same object twice. For example\nwhen both references are for file entries, you know that those files\nhave the same content. That it is a hash doesn't matter, the id could be\nanything that uniquely identifies an object. The \"same object ==> same\ncontent\" part should be pretty obvious, so you only need to know that\nthe \"same name/id ==> same object\" part is true, i.e. that the object\nname/id uniquely identifies the object. And that _is_ true, simply\nbecause you cannot have two objects in the same repo that have the same\nhash and thus the same id. Even if you get a collision, you'll still\nhave just one object.  And that's not something that a term that\ncontains the word \"hash\" is telling me, it would instead tell me that it\nis not something that really uniquely identifies an object, although git\nuses it as such.\n\n\nOnly when you want to explain how git manages to avoid duplicated\nstorage of fully identical contents, then you need to mention that the\nobject names are the hashes of the full object contents. But that's not\nwhat you actually use the object names for.\n\nsame content ==> same content hash ==> object name/id ==> same object\n\n(Actually, you need an additional detail: \"same\nfile/symlink/directory/... contents ==> same object contents\", which\ncan't be made explicit by just saying that you use a hash).\n\nYour chain was in the wrong order and explains neither the \"a tree that\nhas the same object name/id for two entries\" case (because of the\nuncertainity of the \"same hash ?=> same content\" part), nor, when read\nin the other direction, where all implications are true, why same\ncontent leads to the same object (as it already starts at the object\nlevel).\n\n>> You can still say that the name/id is generated by using a hash\n>> function, but the important part is that the name/id is used to\n>> _uniquely_ identify an object,  which isn't apparent when you call it\n>> a hash.\n>\n> I think the implication is important in both directions.  Neither one is \n> self-evident to a new user.  Maybe the right answer is 'hash id'.\n\ngit could work different. Just moving the storage of the filenames from\nthe tree objects to the blobs would mean that you'd get different\nobjects for files that have the same content but different names. You'd\nstill have a hash of the object contents as the object name, but\nsuddenly you get more objects. Just saying \"hash\" or \"hash id\" doesn't\nmagically explain all the other things.\n\n\nBjörn\n"},{"id":"112373","messageId":"20090426181244.GB4942@atjola.homenet","threadId":"19010","inReplyTo":"b4087cc50904260936t4aa05c92ldab3b8ce8c9bd33a@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-26T18:12:44Z","receivedAt":"2009-04-26T18:12:44Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.26 11:36:04 -0500, Michael Witten wrote:\n> 2009/4/26 Björn Steinbrink <B.Steinbrink@gmx.de>:\n> > On 2009.04.25 15:36:24 -0400, David Abrahams wrote:\n> >> Where it's relevant when the user notices that two distinct files have\n> >> the same id (because they happen to have the same contents) and wonders\n> >> what's up.\n> >...\n> > And why would your implementation save the same object twice, in two\n> > distinct files?\n> \n> This question makes me think that you don't understand the parent's\n> point. He's not talking about implementation details; in fact, there's\n> no reason to mix the git world and the file system world at all in\n> this discussion.\n> \n> David is pointing out that a user might notice that two different\n> trees list the same blob. This can be startling if you have incomplete\n> picture about what's going on.\n\nDavid said that the user encounters two distinct files with the same id.\nThe ids are properties of the objects. So he must have meant object\nfiles, or he attributed the id to the wrong thing. I assumed that he\ndidn't mix those things up and really meant the object files, thus my\nreply.\n\n> >From a practical point of view, you might argue that not too many\n> people are looking at trees and blobs;\n\nHeh, I'd rather argue that too _few_ people have looked at commits and\ntrees at least once, whether it's an actual object or a graph like in\ngit for computer scientists.\n\n> however, it seems to me that most people are afraid to use any of\n> git's most useful features precisely because they don't understand the\n> git model and they don't understand that nothing is ever lost unless\n> you explicitly clean up unreferenced objects---they don't see how easy\n> it is manipulate their repos. I argue that if they are given the full\n> knowledge of git's concepts, then they will be able to reason about\n> their repo actions with confidence, even if they only work with\n> commits.\n\nAgreed.\n\n> I think the key is to stress in the documentation the idea that there\n> are 2 separate worlds (the git object world and the working\n> directory's file system world) and that the git tools provide an\n> interface between them; this seems like a small and unnecessarily\n> academic point, but I believe that it's important to working with\n> confidence.\n\nAgreed. That's also why I asked David why the user would look at the\nobject files in the repo (the .git dir). To some degree those are also\nan implementation detail. The user works with the working tree and uses\nthe git tools to modify the repo.\n\n> > It's an identity relation: same name/id => same object. Unlike e.g. a\n> > hash-table where you are expected to deal with collisions, and having\n> > the same hash doesn't mean that you have identical data.\n> \n> However, having the same *cryptographic* hash does mean that you have\n> identical data.\n\nThat's the _assumption_ that git makes. Hash collisions are always\npossible, just hard to create intentionally when the hash function has\nnot yet been broken.\n\nBjörn\n"},{"id":"112363","messageId":"F2B2D447-57B4-459C-8A0D-A94C12AE791C@boostpro.com","threadId":"19010","inReplyTo":"20090426175613.GA4942@atjola.homenet","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-26T20:17:43Z","receivedAt":"2009-04-26T20:17:43Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 26, 2009, at 1:56 PM, Björn Steinbrink wrote:\n\n> On 2009.04.26 09:55:34 -0400, David Abrahams wrote:\n>>\n>> On Apr 26, 2009, at 7:28 AM, Björn Steinbrink wrote:\n>>\n>>> On 2009.04.25 15:36:24 -0400, David Abrahams wrote:\n>>>> Where it's relevant when the user notices that two distinct files\n>>>> have the same id (because they happen to have the same contents)  \n>>>> and\n>>>> wonders what's up.\n>>>\n>>> Why would the user have to care about the object files in the repo?\n>>\n>> What a strange question. I have no idea how to answer. It seems\n>> self- evident to me that users of a VCS care that their files are\n>> stored in it.\n>\n> _Their_ files. The files that come from/end up in the working tree. I\n> cared about those when I used SVN, too. But I never went to the SVN  \n> repo\n> to find out if there are two equal files in it. We're talking about\n> object names, and those belong to objects, not files in the working\n> tree.\n\nI'm telling you, many new users who aren't already versed in Git will  \nnaturally associate the SHA1 codes exposed by the interface with the  \nfiles they've checked in without understand that they actually  \nidentify object files (another poorly chosen Git name, if I've manage  \nto deduce what it means) rather than directly corresponding to states  \nof their files. And anyway, if you want to get into implementation  \ndetails, SHA1s don't always identify object files because blobs get  \ndelta-compressed.\n\n>>> And why would your implementation save the same object twice, in two\n>>> distinct files?\n>>\n>> One could easily have the expectation that contents can be duplicated\n>> because there are numerous precedents in everyone's experience of\n>> computing, for example in filesystems and in any programming language\n>> that is not pure-functional.\n>\n> That's not answering my question. I asked why you come up with an\n> implementation that is \"broken\" enough to save the same object twice\n> with different file names.\n\nI don't know what you mean by \"come up with an implementation.\"  I'm  \nnot inventing an implementation.  I'm saying, new users inevitably and  \ninexorably develop a mental model of the system they're learning  \nabout, and they don't always develop the right mental model, and I'm  \nsaying that it's easy to see how they can fall into incorrect  \nassumptions.  The word \"hash\" helps a bit with avoiding one of those  \nassumptions.\n\n> If the implementation does not do that, your\n> \"when the user notices that two distinct files has the same id\" is\n> immediately invalid. The user cannot come into that situation then.\n\nI think this is why Git remains more opaque than it should be.  You  \ncan't assume that people will naturally develop the smartest possible  \nmental model of a VCS, even with faced with some hints in the form of  \na partial understanding of Git.\n\n> And\n> anyway, when the user notices something, that's a discovery, not an\n> expectation.\n\nIt's better to give people something to connect their discoveries to  \n(e.g. \"oh, I see, they call those things hashes, so it makes sense  \nthat these two identical things are stored once\")\n\n>>> The SHA-1 hash is created from the object, that means\n>>> the its type, size and data. It's not an id of a file in the working\n>>> tree, but of an object\n>>\n>> All true.  All somewhat subtle distinctions that are not nearly as\n>> apparent unless you actually use the word \"hash\" as I have been\n>> advocating.\n>\n> Hu? How does saying \"object hash\" instead of \"object id\" make it any\n> more apparent that a file in the working tree is something else than a\n> git object?\n\nIt makes it apparent that two identical things can only have one ID,  \nand thus must correspond to one object.\n\n>>> You can't have two objects with the same contents to begin with,  \n>>> same\n>>> content => same object.\n>>\n>> In the Git world, I agree.  In general, I disagree.\n>\n> I don't think were discussing a term to describe something that\n> identifies an object in general. So, \"in general\" you can disagree as\n> much as you want, but for git that doesn't matter at all.\n\nYou don't think the general rules of the computing world and existing  \nmeanings of terms have an impact on a new user's ability to grok Git?   \nIf not, we don't have much to discuss.\n\n>> The fact that is so in the Git world is reinforced by the notion that\n>> the id of an object is a hash of its contents.\n>>\n>>> You can just have that one object stored multiple times in different\n>>> places (for sane implementations this  likely means that you have\n>>> more than one repo to look at, and each has its  own copy of that\n>>> object, but that's nothing you as an user should have to care  \n>>> about).\n>>\n>>> It's an identity relation: same name/id => same object. Unlike  \n>>> e.g. a\n>>> hash-table where you are expected to deal with collisions, and  \n>>> having\n>>> the same hash doesn't mean that you have identical data.  But that's\n>>> not true of git, it expects an identity relation, which is IMHO\n>>> better expressed through \"object name\" or \"object id\".\n>>\n>> Yes, that's true in the Git world (though not necessarily  \n>> elsewhere), or\n>> at least you hope it is.  In fact, there's no guarantee that SHA1\n>> collisions won't occur; it's just exremely unlikely.  In fact, if you\n>> google it you can find some interesting papers about SHA1 collision.\n>\n> Sure, it's an assumption that has been made and is required to hold  \n> true\n> for git to work.\n>\n>> Another way to express what you wrote above:\n>>\n>>   same same id => same hash ?=> same contents => same object\n>>\n>> where ?=> means \"almost certainly implies.\"\n>\n> No, that chain shows how git could be \"unreliable\" when you get hash\n> collisions. You could put that into a chapter that explains the\n> implications of the way git generates its object ids. But it's not  \n> very\n> interesting when you use git and (implicitly) trust the assumption  \n> that\n> no collisions happen.\n\nMy point in mentioning that it's not certain was to point out that you  \nleft out the implication that actually /is/ certain, even across repos.\n\n> Only when you want to explain how git manages to avoid duplicated\n> storage of fully identical contents, then you need to mention that the\n> object names are the hashes of the full object contents. But that's  \n> not\n> what you actually use the object names for.\n>\n> same content ==> same content hash ==> object name/id ==> same object\n>\n> (Actually, you need an additional detail: \"same\n> file/symlink/directory/... contents ==> same object contents\", which\n> can't be made explicit by just saying that you use a hash).\n>\n> Your chain was in the wrong order\n\nIf you think there's a right order, you haven't understood that all  \nthe arrows are bidirectional.\n\n> and explains neither the \"a tree that\n> has the same object name/id for two entries\" case (because of the\n> uncertainity of the \"same hash ?=> same content\" part), nor, when read\n> in the other direction, where all implications are true, why same\n> content leads to the same object (as it already starts at the object\n> level).\n\n>> I think the implication is important in both directions.  Neither  \n>> one is\n>> self-evident to a new user.  Maybe the right answer is 'hash id'.\n>\n> git could work different. Just moving the storage of the filenames  \n> from\n> the tree objects to the blobs would mean that you'd get different\n> objects for files that have the same content but different names.  \n> You'd\n> still have a hash of the object contents as the object name, but\n> suddenly you get more objects. Just saying \"hash\" or \"hash id\" doesn't\n> magically explain all the other things.\n\n\nBut that's a strawman.  I'm not claiming that it magically explains  \nall the other things.  I'm just claiming that it helps in avoiding  \nsome possible misunderstandings.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112362","messageId":"79E58E5E-8A59-420B-A8B0-624C2F59A5FB@boostpro.com","threadId":"19010","inReplyTo":"20090426181244.GB4942@atjola.homenet","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-26T20:20:43Z","receivedAt":"2009-04-26T20:20:43Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 26, 2009, at 2:12 PM, Björn Steinbrink wrote:\n\n> That's also why I asked David why the user would look at the\n> object files in the repo (the .git dir).\n\n\nFor what it's worth, I didn't understand what you meant by \"object  \nfiles\" until now.  I never claimed they would look at those files, at  \nleast not intentionally.  But just look at any web interface to a Git  \nrepo and you'll see why they might encounter the object file names  \neven before they've installed git on their own machine.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112403","messageId":"20090426222532.GA12338@atjola.homenet","threadId":"19010","inReplyTo":"F2B2D447-57B4-459C-8A0D-A94C12AE791C@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-26T22:25:32Z","receivedAt":"2009-04-26T22:25:32Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.26 16:17:43 -0400, David Abrahams wrote:\n>\n> On Apr 26, 2009, at 1:56 PM, Björn Steinbrink wrote:\n>\n>> On 2009.04.26 09:55:34 -0400, David Abrahams wrote:\n>>>\n>>> On Apr 26, 2009, at 7:28 AM, Björn Steinbrink wrote:\n>>>\n>>>> On 2009.04.25 15:36:24 -0400, David Abrahams wrote:\n>>>>> Where it's relevant when the user notices that two distinct files\n>>>>> have the same id (because they happen to have the same contents)\n>>>>> and wonders what's up.\n>>>>\n>>>> Why would the user have to care about the object files in the repo?\n>>>\n>>> What a strange question. I have no idea how to answer. It seems\n>>> self- evident to me that users of a VCS care that their files are\n>>> stored in it.\n>>\n>> _Their_ files. The files that come from/end up in the working tree. I\n>> cared about those when I used SVN, too. But I never went to the SVN\n>> repo to find out if there are two equal files in it. We're talking\n>> about object names, and those belong to objects, not files in the\n>> working tree.\n>\n> I'm telling you, many new users who aren't already versed in Git will\n> naturally associate the SHA1 codes exposed by the interface with the\n> files they've checked in without understand that they actually\n> identify object files (another poorly chosen Git name, if I've manage\n> to deduce what it means)\n\nHm, not sure if that name is really important. The way objects are\nstored is an implementation detail. Usually, we're just talking about\n\"objects\" not the files the loose objects are stored in (loose object =\nan object stored in its own file, not in a pack file). But as you\ncomplained about it, how would you call a file in which an object is\nstored?\n\n> rather than directly corresponding to states\n> of their files. And anyway, if you want to get into implementation\n> details, SHA1s don't always identify object files because blobs get\n> delta-compressed.\n\nTrue, they identify the object, it's not even necessesary to mention\ndelta compression, just having the object in a pack file causes the\nobject name to no longer identify the file in which the object can be\nfound. Heck, the object might be in a different repo when you use\nalternates ;-). And I think I never explicitly said that they\nidentify a file storing an object, but implied that by \"accepting\" your\nexample and assuming that you meant two object files having the same id.\nI should have said that your \"two distinct files have the same id\" makes\nno sense and should have asked what you mean.\n\n>>>> And why would your implementation save the same object twice, in\n>>>> two distinct files?\n>>>\n>>> One could easily have the expectation that contents can be\n>>> duplicated because there are numerous precedents in everyone's\n>>> experience of computing, for example in filesystems and in any\n>>> programming language that is not pure-functional.\n>>\n>> That's not answering my question. I asked why you come up with an\n>> implementation that is \"broken\" enough to save the same object twice\n>> with different file names.\n>\n> I don't know what you mean by \"come up with an implementation.\"  I'm\n> not inventing an implementation.\n\nSorry, \"come up with\" is clearly wrong. \"Assume\" or \"expect\" or so might\nhave been more correct. But I think we could agree that you misused the\n\"id\" term by using it for files, and what ensued confused both of us? If\nyou didn't mean the stored objects by \"files\", then that part of the\ndiscussion was just based on a misunderstanding and can be ignored.\n\n> I'm saying, new users inevitably and inexorably develop a mental model\n> of the system they're learning about, and they don't always develop\n> the right mental model, and I'm saying that it's easy to see how they\n> can fall into incorrect assumptions.  The word \"hash\" helps a bit with\n> avoiding one of those assumptions.\n\nI've not met a lot of people that were actually confused about the fact\nthat the same object might be \"reused\" for tree entries with different\nnames. But most (all?) of those that were confused knew that the objects\nare identified by hashes, but expected the filenames to be part of the\nobject and didn't know about tree objects.\n\n>> If the implementation does not do that, your \"when the user notices\n>> that two distinct files has the same id\" is immediately invalid. The\n>> user cannot come into that situation then.\n>\n> I think this is why Git remains more opaque than it should be.  You\n> can't assume that people will naturally develop the smartest possible\n> mental model of a VCS, even with faced with some hints in the form of\n> a partial understanding of Git.\n\nI don't think I understand what you mean here. If users don't understand\nthe data model, that's caused by missing/bad documentation or because\nthe user doesn't want to read the existing documentation. (I'll make no\nassumptions here, it's been some time since I had a close look at the\ndocs). But I've been talking about how the given implementation stores\ndata in the repository. Could you explain?\n\n>> And anyway, when the user notices something, that's a discovery, not\n>> an expectation.\n>\n> It's better to give people something to connect their discoveries to\n> (e.g. \"oh, I see, they call those things hashes, so it makes sense\n> that these two identical things are stored once\")\n\nWe're talking about seeing, for example,  the same object name more than\nonce, for different \"files\", in e.g. gitweb, right? Then the \"Hu? Isn't\nthe filename part of the object?\" thing might still apply. The user can\nstill very easily make a wrong guess.\n\nAs Michael said in another mail, the important point is probably rather\nto teach people to make a distinction between files and directories in\nthe working tree and the contents stored in the git objects. And that's\nnot accomplished by saying that the id is a hash, when the user doesn't\nknow what the hash is based upon.\n\nSomewhat related: I'm trying to remember if I ever had problems\nexplaining the concept of hardlinks to someone, but I don't remember any\nsuch situation anymore. There are no hashes involved there, and I feel\nlike that was quite easy to grasp for most people I talked to. It's\npretty similar, separating content from names.\n\n>>>> The SHA-1 hash is created from the object, that means the its type,\n>>>> size and data. It's not an id of a file in the working tree, but of\n>>>> an object\n>>>\n>>> All true.  All somewhat subtle distinctions that are not nearly as\n>>> apparent unless you actually use the word \"hash\" as I have been\n>>> advocating.\n>>\n>> Hu? How does saying \"object hash\" instead of \"object id\" make it any\n>> more apparent that a file in the working tree is something else than\n>> a git object?\n>\n> It makes it apparent that two identical things can only have one ID,\n> and thus must correspond to one object.\n\nSee above, the user needs to know \"what\" is identical in the first\nplace.\n\n>>>> You can't have two objects with the same contents to begin with,\n>>>> same content => same object.\n>>>\n>>> In the Git world, I agree.  In general, I disagree.\n>>\n>> I don't think were discussing a term to describe something that\n>> identifies an object in general. So, \"in general\" you can disagree as\n>> much as you want, but for git that doesn't matter at all.\n>\n> You don't think the general rules of the computing world and existing\n> meanings of terms have an impact on a new user's ability to grok Git?\n> If not, we don't have much to discuss.\n\nThis was probably also based on the files+id misunderstanding combined\nwith the fact that you used the term \"object\" where I thought that you\nmeant a \"git object\" (you probably didn't, right?). Because when talking\nabout \"git objects\" you actually can't have two different ones with the\nsame \"value\" (I guess you mean type, size and content when you say\n\"value\", right?)\n\nAnd admittedly, for this one, the \"hash\" term _would_ help to get the\nuser to understand that in git you cannot have two different objects\nwith the same contents and that this makes git different and efficient.\nBut I still don't buy that this is important for understanding the basic\ndata model. It's a nice hint why git can always quickly tell that two\nthings are equal and why the repository size doesn't explode. But the\nimportant part is the separation of names and content, that trees give\nnames to the contents stored in blobs. The \"hash\" name would only help\nto understand its efficiency once you already understood the data model.\nSee below.\n\n>>> The fact that is so in the Git world is reinforced by the notion\n>>> that the id of an object is a hash of its contents.\n>>>\n>>>> You can just have that one object stored multiple times in\n>>>> different places (for sane implementations this  likely means that\n>>>> you have more than one repo to look at, and each has its  own copy\n>>>> of that object, but that's nothing you as an user should have to\n>>>> care  about).\n>>>\n>>>> It's an identity relation: same name/id => same object. Unlike\n>>>> e.g. a hash-table where you are expected to deal with collisions,\n>>>> and  having the same hash doesn't mean that you have identical\n>>>> data.  But that's not true of git, it expects an identity relation,\n>>>> which is IMHO better expressed through \"object name\" or \"object\n>>>> id\".\n>>>\n>>> Yes, that's true in the Git world (though not necessarily\n>>> elsewhere), or at least you hope it is.  In fact, there's no\n>>> guarantee that SHA1 collisions won't occur; it's just exremely\n>>> unlikely.  In fact, if you google it you can find some interesting\n>>> papers about SHA1 collision.\n>>\n>> Sure, it's an assumption that has been made and is required to hold\n>> true for git to work.\n>>\n>>> Another way to express what you wrote above:\n>>>\n>>>   same same id => same hash ?=> same contents => same object\n>>>\n>>> where ?=> means \"almost certainly implies.\"\n>>\n>> No, that chain shows how git could be \"unreliable\" when you get hash\n>> collisions. You could put that into a chapter that explains the\n>> implications of the way git generates its object ids. But it's not\n>> very interesting when you use git and (implicitly) trust the\n>> assumption  that no collisions happen.\n>\n> My point in mentioning that it's not certain was to point out that you\n> left out the implication that actually /is/ certain, even across\n> repos.\n\nAnd my point is that this is not important for understanding the basic\ndata model, but only how git efficiently implements it, and which\nassumptions it has to make.\n\n>> Only when you want to explain how git manages to avoid duplicated\n>> storage of fully identical contents, then you need to mention that\n>> the object names are the hashes of the full object contents. But\n>> that's  not what you actually use the object names for.\n>>\n>> same content ==> same content hash ==> object name/id ==> same object\n>>\n>> (Actually, you need an additional detail: \"same\n>> file/symlink/directory/... contents ==> same object contents\", which\n>> can't be made explicit by just saying that you use a hash).\n>>\n>> Your chain was in the wrong order\n>\n> If you think there's a right order, you haven't understood that all\n> the arrows are bidirectional.\n\nThere's one that is not truly bidirectional.\n\nid <=> hash <?=> contents <=> object\n\nI can't go from id/hash to contents/object without hitting the \"hash =>\ncontent\" assumption. I had two chains for a reason.\n\n\tid => object => content => hash\nand\n\tcontent => hash => id => object\n\nare guaranteed, at least within a single repo.\n\nWhile:\n\tcontent => hash => id => object ?=> content\n\nhas a non-guaranteed part again, just an assumption, at least when the\nfirst and last \"content\" mean the same content. If you get a collision,\nyou rather have a guarantee that one version of the content is \"not in\nthe repo\".\n\nAnd as I said, that fact, that the identifier is not globally unique,\nalong with the fact that git cannot have two different objects with the\nsame contents or name is not required to understand how commits, tree\nand blobs go together to store the history of a project. It's IM(NS?)HO\nfar more important to understand the separation of names and content.\nThen you can understand that multiple names can be associated with the\nsame object holding some content (which can be done with other kinds of\nids as well, even with more than one object having the same contents,\njust not necessarily as efficiently). And that objects have a name that\nis used to identify the object. And only then can you understand and\nappreciate how hashes help to efficiently implement that model, knowing\nwhich data is used to calculate the hash.\n\n>> and explains neither the \"a tree that has the same object name/id for\n>> two entries\" case (because of the uncertainity of the \"same hash ?=>\n>> same content\" part), nor, when read in the other direction, where all\n>> implications are true, why same content leads to the same object (as\n>> it already starts at the object level).\n>\n>>> I think the implication is important in both directions.  Neither\n>>> one is self-evident to a new user.  Maybe the right answer is 'hash\n>>> id'.\n>>\n>> git could work different. Just moving the storage of the filenames\n>> from the tree objects to the blobs would mean that you'd get\n>> different objects for files that have the same content but different\n>> names.  You'd still have a hash of the object contents as the object\n>> name, but suddenly you get more objects. Just saying \"hash\" or \"hash\n>> id\" doesn't magically explain all the other things.\n>\n> But that's a strawman.  I'm not claiming that it magically explains\n> all the other things.  I'm just claiming that it helps in avoiding\n> some possible misunderstandings.\n\nAnd I think that it doesn't help much at all and might confuse users,\nbecause they expect the hash to be based on the wrong stuff. It's just\nimportant that the \"thing\" is used to identify an object.\n\nBjörn\n"},{"id":"112356","messageId":"20090426224230.GB12338@atjola.homenet","threadId":"19010","inReplyTo":"B8EE4B7B-CBCE-47D6-97B2-2F505A218B58@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-26T22:42:30Z","receivedAt":"2009-04-26T22:42:30Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.24 20:48:57 -0400, David Abrahams wrote:\n>\n> On Apr 24, 2009, at 8:01 PM, Michael Witten wrote:\n>\n>>> What's wrong with just calling the object name \"object name\"?\n>>\n>> What's wrong with calling the object address \"object address\"?\n>\n> Neither captures the connection to the object's contents.  I think  \n> \"value ID\" would be closer, but it's probably too horrible.\n\nI think I asked this in another mail, but I'm quite tired, so just to\nmake sure: What do you mean by \"value\"? I might be weird (I'm not a\nnative speaker, so I probably make funny and wrong connotations from\ntime to time), but while I can accept \"content\" to include the type and\nsize of the object, the term \"value\" makes me want to exclude those\npieces of meta data. So \"value\" somehow feels wrong to me, as the hash\ncovers those two fields.\n\nBjörn\n"},{"id":"112354","messageId":"20090426233539.GC12338@atjola.homenet","threadId":"19010","inReplyTo":"78D97574-74AB-4A4D-AEB2-874BFBB4345E@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-26T23:35:39Z","receivedAt":"2009-04-26T23:35:39Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.24 20:39:12 -0400, David Abrahams wrote:\n>\n> On Apr 24, 2009, at 6:45 PM, Björn Steinbrink wrote:\n>\n>>> I think UI/API works way better than porcelain/plumbing. We are,\n>>> after all, programmers.\n>>\n>> We are programmers, but not all git users are programmers.\n>\n> I'm sure you will admit that the vast majority are programmers.  This is \n> about speaking effectively to your primary audience.\n\nMy experience says to at least drop the \"vast\", but that might be\nbiased, due to the fact that the non-programmers probably need more time\nwhen you explain things to them.\n\nBut thinking about it again, I don't think I like UI/API regardless of\nthat. High-Level/Low-Level yes, but API? No. The plumbing is meant to be\nstable so it can serve as an API, and it also has options that only make\nsense when you use it that way (e.g. the parse-opt support in rev-parse)\nbut I also happen to just use those programs as a UI. For example\nls-files, ls-remote, or apply.\n\nAnd git(1) also has the sections titled \"HIGH-LEVEL COMMANDS\n(PORCELAIN)\" and \"LOW_LEVEL COMMANDS (PLUMBING)\". So if we were to get\nrid of the porcelain and plumbing terms, then _I_ would go for just\n\"high-level commands\" and \"low-level commands\".\n\n>>> I should also say, most of the docs and interfaces I see in Git (and\n>>> its wrappers, web interfaces, etc.) give the SHA1 hashes way too much\n>>> exposure. The times when it's actually more convenient to use a hash\n>>> instead of one of the other notations are rare,\n>>\n>> How often do you need a name for a commit shown by a command and can\n>> accept that it is not stable?\n>\n> I can accept it as long as it's stable inside my own repo.  Maybe I\n> need the SHA1 to talk about it wherever it may roam.  I think you\n> could count in the other direction (i.e. from the roots instead of the\n> leaves) to get fairly stable symbolic names.\n\nI'm sure this has been discussed in the earlier \"stable revision\nnumbers\" threads as well, so you can find more information there, but I\njust want to mention that one drawback of this is that those numbers\nstill have no notion of \"commit age\". You could have 5000 commits in\nyour repo, and then you fetch someone elses stuff that might have some\nvery old stuff that you don't have yet. And that gets high numbers now.\nSo 5051 might be older than 200. Doesn't exactly help to make those\nnumbers \"useful\". Just like the \"gaps\" you get by using e.g. rebase -i\nor other means that cause commits to be garbage collected.\n\n> Also, I don't think I need to see the hashes for trees and blobs most of \n> the time.\n\nOK, I think finally see what you might mean there. I'm almost\nexclusively using the CLI and gitk and seldomly see tree/blob object\nnames in a prominent way unless I ask for them. But I just noticed that\ngitweb is at least showing a \"daunting\" number of object names without\nfurther details when you ask for a \"commit\", while the \"commitdiff\" is\ncloser to what \"git show <commit>\" would show. And yeah, I think that\ncould be improved, moving the object name more into the \"background\" (I\ndon't think it should be completely removed, just be less prominent).\nAny other \"high-level\" tool that you noticed being noisy about tree/blob\nhashes?\n\n>>> Oh, one other specific issue: the rev-parse manpage uses $GIT_DIR\n>>> without saying what it is. I *think* that means the root of the\n>>> working copy and has nothing to do with environment variables, but\n>>> it's hard to be sure, and if I'm right about that, it's misleading\n>>> notation.\n>>\n>> $GIT_DIR means the .git directory of a non-bare repo.\n>\n>\n> Thanks for clarifying.  But don't neglect to fix the docs so the next\n> guy doesn't have to ask ;-)\n\nHm, I provide the information, you provide the patch? ;-) Hm, maybe I'll\nfind some time to provide one myself. But my git and general todo lists\nalready grew beyond all limits...\n\n> BTW, \"[non-]bare repo\" is yet another Git-specific jargon.  I know what \n> it means... again, only because I asked someone.\n\nAt least \"bare repository\" appears as an entry in the glossary\n(gitglossary(7), also reachable via \"git help glossary\").\n\nBjörn\n"},{"id":"112401","messageId":"20090426234159.GD12338@atjola.homenet","threadId":"19010","inReplyTo":"20090424232531.GA15136@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-26T23:41:59Z","receivedAt":"2009-04-26T23:41:59Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.24 19:25:31 -0400, Jeff King wrote:\n> On Fri, Apr 24, 2009 at 07:21:26PM -0400, Daniel Barkalow wrote:\n> \n> > (And, actually, I think git has a few usability warts due to relying too \n> > much on command line arguments being objects; it would be quite nice if \n> > \"git blame 1a2b3c:Makefile\" worked despite this technically being \n> > incoherent.)\n> \n> Yeah, I think another is that \"git show master:file\" will not do CRLF or\n> other filters, and \"git diff master:file other:file\" will not respect\n> diff settings. I think all of those could be solved by path lookup\n> attaching a \"here is a pathname I used to get to this object\" string,\n> which can then be accessed as appropriate.\n> \n> It is not all that different conceptually than what \"git rev-list\n> --objects\" does.\n\nIt's also something that hash-object already does in some way. To apply\ne.g. attributes to content that you supply via stdin.\n\nBjörn\n"},{"id":"112353","messageId":"20090427000007.GE12338@atjola.homenet","threadId":"19010","inReplyTo":"b4087cc50904241629u76454b1chc6e84e95066a9100@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-27T00:00:07Z","receivedAt":"2009-04-27T00:00:07Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.24 18:29:22 -0500, Michael Witten wrote:\n> On Fri, Apr 24, 2009 at 18:21, Daniel Barkalow <barkalow@iabervon.org> wrote:\n> > \"git blame 1a2b3c:Makefile\" worked despite this technically being\n> > incoherent.\n> \n> It seems to work on my end, and it's perfectly coherent if you\n> consider git-blame to be overloaded to handle both pointers and\n> addresses (or references and object names, if you prefer).\n\nFails for me. And it's technically incoherent in that it makes no sense\nto use blame with a blob object. 1a2b3c:Makefile identifies \"just\" a\nblob object. And that has no parents and no history, just contents. Only\nthe commit objects have the references that connect them to form a\nhistory.\n\nFor example, you could have a history like this:\n\nA---B---C---D---E\n\nAnd a file \"foo\" that has the same contents for A and E. Then \"A:foo\"\nand \"E:foo\" lead to the same blob object, and you can't uniquely go from\nthat blob object to any commit object. So technically, you can't tell if\n\"git blame E:foo\" means \"git blame E foo\" or \"git blame A foo\" (and you\ncan add a bunch of complexity by having, for example, a second file\nwith a different name that had the same content at some point).\n\nTo make that coherent, you must change the definition of the\n<tree-ish>:<path> syntax so that the context in which the path is\nresolved is kept, it must no longer just identify an object, but\nsomething more complex.\n\nBjörn\n"},{"id":"112350","messageId":"m2ab62g9fg.fsf@boostpro.com","threadId":"19010","inReplyTo":"20090426222532.GA12338@atjola.homenet","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-27T01:41:23Z","receivedAt":"2009-04-27T01:41:23Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\non Sun Apr 26 2009, Björn Steinbrink <B.Steinbrink-AT-gmx.de> wrote:\n\n>> I think this is why Git remains more opaque than it should be.  You\n>> can't assume that people will naturally develop the smartest possible\n>> mental model of a VCS, even with faced with some hints in the form of\n>> a partial understanding of Git.\n>\n> I don't think I understand what you mean here. If users don't understand\n> the data model, that's caused by missing/bad documentation or because\n> the user doesn't want to read the existing documentation. (I'll make no\n> assumptions here, it's been some time since I had a close look at the\n> docs). But I've been talking about how the given implementation stores\n> data in the repository. Could you explain?\n\nYou don't have to \"not want to read the documentation\" to have an\nincomplete mental model.  The mental model development doesn't happen\nupon finishing the documentation; it happens while the person is\nlearning.  Halfway through the docs, I have an incomplete mental model.\nIf you make it hard enough for me, maybe I never finish and I retain\nthat incomplete model forever.  The more you can help people avoid\nincorrect assumptions as they read along, the easier it will be for them\nto grok the next bit they are reading, and the less likely they are to\nbecome discouraged.\n\n-- \nDave Abrahams\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"112434","messageId":"33614929-B5B2-4B54-BF18-81ADBBCC4925@boostpro.com","threadId":"19010","inReplyTo":"20090426222532.GA12338@atjola.homenet","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-27T16:30:58Z","receivedAt":"2009-04-27T16:30:58Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 26, 2009, at 6:25 PM, Björn Steinbrink wrote:\n\n> On 2009.04.26 16:17:43 -0400, David Abrahams wrote:\n>>\n\n>> I'm telling you, many new users who aren't already versed in Git will\n>> naturally associate the SHA1 codes exposed by the interface with the\n>> files they've checked in without understand that they actually\n>> identify object files (another poorly chosen Git name, if I've manage\n>> to deduce what it means)\n>\n> Hm, not sure if that name is really important. The way objects are\n> stored is an implementation detail. Usually, we're just talking about\n> \"objects\" not the files the loose objects are stored in (loose  \n> object =\n> an object stored in its own file, not in a pack file). But as you\n> complained about it, how would you call a file in which an object is\n> stored?\n\n\"Object\" is OK. \"Object file\" is overloaded and confusing.  I'd just  \nsay there are \"Git data files\" or \"files in Git's object store\", some  \nof which store single objects whose id is the same as the filename,  \nand some of which store multiple objects.\n\n>> rather than directly corresponding to states\n>> of their files. And anyway, if you want to get into implementation\n>> details, SHA1s don't always identify object files because blobs get\n>> delta-compressed.\n>\n> True, they identify the object, it's not even necessesary to mention\n> delta compression, just having the object in a pack file causes the\n> object name to no longer identify the file in which the object can be\n> found.\n\nRight.\n\n> Heck, the object might be in a different repo when you use\n> alternates ;-). And I think I never explicitly said that they\n> identify a file storing an object, but implied that by \"accepting\"  \n> your\n> example and assuming that you meant two object files having the same  \n> id.\n\nYes, that assumption was wrong, and then when you responded using the  \nterm \"object file\" I didn't know what it meant.\n\n> I should have said that your \"two distinct files have the same id\"  \n> makes\n> no sense and should have asked what you mean.\n>\n>>>>> And why would your implementation save the same object twice, in\n>>>>> two distinct files?\n>>>>\n>>>> One could easily have the expectation that contents can be\n>>>> duplicated because there are numerous precedents in everyone's\n>>>> experience of computing, for example in filesystems and in any\n>>>> programming language that is not pure-functional.\n>>>\n>>> That's not answering my question. I asked why you come up with an\n>>> implementation that is \"broken\" enough to save the same object twice\n>>> with different file names.\n>>\n>> I don't know what you mean by \"come up with an implementation.\"  I'm\n>> not inventing an implementation.\n>\n> Sorry, \"come up with\" is clearly wrong. \"Assume\" or \"expect\" or so  \n> might\n> have been more correct.\n\nI think I explained why one might make that assumption.\n\n> But I think we could agree that you misused the\n> \"id\" term by using it for files, and what ensued confused both of  \n> us? If\n> you didn't mean the stored objects by \"files\", then that part of the\n> discussion was just based on a misunderstanding and can be ignored.\n\nI meant what the user thinks of as files stored in the repository.\n\n>> I'm saying, new users inevitably and inexorably develop a mental  \n>> model\n>> of the system they're learning about, and they don't always develop\n>> the right mental model, and I'm saying that it's easy to see how they\n>> can fall into incorrect assumptions.  The word \"hash\" helps a bit  \n>> with\n>> avoiding one of those assumptions.\n>\n> I've not met a lot of people that were actually confused about the  \n> fact\n> that the same object might be \"reused\" for tree entries with different\n> names. But most (all?) of those that were confused knew that the  \n> objects\n> are identified by hashes, but expected the filenames to be part of the\n> object and didn't know about tree objects.\n\nWell, there's certainly precedent for the idea that the filenames are  \ndistinct from file contents.\n\n>>> And anyway, when the user notices something, that's a discovery, not\n>>> an expectation.\n>>\n>> It's better to give people something to connect their discoveries to\n>> (e.g. \"oh, I see, they call those things hashes, so it makes sense\n>> that these two identical things are stored once\")\n>\n> We're talking about seeing, for example,  the same object name more  \n> than\n> once, for different \"files\", in e.g. gitweb, right? Then the \"Hu?  \n> Isn't\n> the filename part of the object?\" thing might still apply. The user  \n> can\n> still very easily make a wrong guess.\n>\n> As Michael said in another mail, the important point is probably  \n> rather\n> to teach people to make a distinction between files and directories in\n> the working tree and the contents stored in the git objects. And  \n> that's\n> not accomplished by saying that the id is a hash, when the user  \n> doesn't\n> know what the hash is based upon.\n>\n> Somewhat related: I'm trying to remember if I ever had problems\n> explaining the concept of hardlinks to someone, but I don't remember  \n> any\n> such situation anymore. There are no hashes involved there, and I feel\n> like that was quite easy to grasp for most people I talked to. It's\n> pretty similar, separating content from names.\n\nThe difference is that hardlinks are only generated explicitly.  You'd  \nneed something like a hash to generate them automatically and  \nimplicitly.\n\n>>>>> You can't have two objects with the same contents to begin with,\n>>>>> same content => same object.\n>>>>\n>>>> In the Git world, I agree.  In general, I disagree.\n>>>\n>>> I don't think were discussing a term to describe something that\n>>> identifies an object in general. So, \"in general\" you can disagree  \n>>> as\n>>> much as you want, but for git that doesn't matter at all.\n>>\n>> You don't think the general rules of the computing world and existing\n>> meanings of terms have an impact on a new user's ability to grok Git?\n>> If not, we don't have much to discuss.\n>\n> This was probably also based on the files+id misunderstanding combined\n> with the fact that you used the term \"object\" where I thought that you\n> meant a \"git object\" (you probably didn't, right?).\n\nI didn't.  I meant the general notion of \"object\" in computing.  I'm  \ntrying to talk about how the language used by Git's docs can bias  \npeople toward correct or incorrect understandings of Git as they're  \nlearning.\n\n> Because when talking\n> about \"git objects\" you actually can't have two different ones with  \n> the\n> same \"value\" (I guess you mean type, size and content when you say\n> \"value\", right?)\n\nYes.  Size is a function of content, so that adds nothing, and whether  \nit even makes sense to say that two things of different type have  \nidentical content is debatable.\n\n> And admittedly, for this one, the \"hash\" term _would_ help to get the\n> user to understand that in git you cannot have two different objects\n> with the same contents and that this makes git different and  \n> efficient.\n> But I still don't buy that this is important for understanding the  \n> basic\n> data model. It's a nice hint why git can always quickly tell that two\n> things are equal and why the repository size doesn't explode. But the\n> important part is the separation of names and content, that trees give\n> names to the contents stored in blobs.\n\nBut there's nothing unique about that; it's not distinct from what  \nfilesystems do.\n\n> The \"hash\" name would only help\n> to understand its efficiency once you already understood the data  \n> model.\n\nIt would help to reinforce that an object's id is a function of its  \ncontents.  It would help to make clear why the same object can be  \nidentified in the same way across all repos.\n\n>>>> Another way to express what you wrote above:\n>>>>\n>>>>  same same id => same hash ?=> same contents => same object\n>>>>\n>>>> where ?=> means \"almost certainly implies.\"\n>>>\n>>> No, that chain shows how git could be \"unreliable\" when you get hash\n>>> collisions. You could put that into a chapter that explains the\n>>> implications of the way git generates its object ids. But it's not\n>>> very interesting when you use git and (implicitly) trust the\n>>> assumption  that no collisions happen.\n>>\n>> My point in mentioning that it's not certain was to point out that  \n>> you\n>> left out the implication that actually /is/ certain, even across\n>> repos.\n>\n> And my point is that this is not important for understanding the basic\n> data model, but only how git efficiently implements it, and which\n> assumptions it has to make.\n\nLook, you're talking to someone who has just had to go through the  \nprocess of learning all this stuff.  What I'm telling you is based on  \nmy experiences.  Just one datapoint, to be sure, but knowing that it's  \na hash was crucial for me.\n\n>> If you think there's a right order, you haven't understood that all\n>> the arrows are bidirectional.\n>\n> There's one that is not truly bidirectional.\n>\n> id <=> hash <?=> contents <=> object\n>\n> I can't go from id/hash to contents/object without hitting the \"hash  \n> =>\n> content\" assumption.\n\nQuite right.  You can't derive contents from the hash.\n\n>> But that's a strawman.  I'm not claiming that it magically explains\n>> all the other things.  I'm just claiming that it helps in avoiding\n>> some possible misunderstandings.\n>\n> And I think that it doesn't help much at all and might confuse users,\n> because they expect the hash to be based on the wrong stuff. It's just\n> important that the \"thing\" is used to identify an object.\n\n\nOK, I give up.  *I* now understand the system, and it's starting to  \nlook like too much of a struggle to improve things for others, so they  \ncan fend for themselves I guess.\n\nThanks for the lively discussion, anyway.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112437","messageId":"b4087cc50904270952k5a7383e6hd3b29104a7aac1fd@mail.gmail.com","threadId":"19010","inReplyTo":"33614929-B5B2-4B54-BF18-81ADBBCC4925@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-04-27T16:52:03Z","receivedAt":"2009-04-27T16:52:03Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"2009/4/27 David Abrahams <dave@boostpro.com>:\n>\n> I didn't.  I meant the general notion of \"object\" in computing.  I'm trying\n> to talk about how the language used by Git's docs can bias people toward\n> correct or incorrect understandings of Git as they're learning.\n\nActually, I believe object was first used to describe anything stored\nin memory. Given that, I still think my usage of C pointer terminology\nis superior to everything; it's just the case that objects are content\naddressable in the git world and location addressable in the C world.\n"},{"id":"112600","messageId":"20090429063448.GA22448@coredump.intra.peff.net","threadId":"19010","inReplyTo":"1A9F6DB0-983F-4A5B-B3B7-33227C11F36A@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-29T06:34:48Z","receivedAt":"2009-04-29T06:34:48Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 24, 2009 at 08:53:37PM -0400, David Abrahams wrote:\n\n>> Actually, it is not the generally of trees that I think is interesting\n>> there, but the generality of _objects_. That is, each of those things is\n>> a first-class object, and has a unique name by which it can be  \n>> referred.\n>\n> I'm sorry, but I think most people would find that so unremarkable that \n> making a big deal about it would lead to \"what am I missing here\"  \n> confusion.  Maybe a person who's exclusively used CVS (or older)  \n> technologies before coming to Git would be happy to know that, but it's \n> sort of obvious.  In CVS the lack of first-class directories sticks out \n> like a sore thumb.\n\nSadly, I was away from email all weekend and so missed the ensuing storm\nin this thread. :) However, I did want to respond to this one point.\n\nTo me (and I am talking from personal experience, so it really may be\n_just_ me), an important part of understanding git was understanding the\nobject storage. That is, half of the idea of git is a big database of\ncontent-addressable objects. The _other_ half is the actual VCS built on\ntop of it. ;)\n\nAnd by understanding that, and the places where objects refer to each\nother (commits point to other commits and to trees, trees point to\nblobs, blobs are always leaves), I find it easier to understand what\neach operation is doing. And that if I'm unsure of something, I can\nalways inspect it at many levels.\n\nI don't know. Maybe that is too low-level for most people. I did end up\nworking on git, so perhaps I am inordinately interested.\n\n-Peff\n"},{"id":"112634","messageId":"D1C7AEA3-1565-4E86-AFF0-9EC2F0D79FCC@boostpro.com","threadId":"19010","inReplyTo":"20090429063448.GA22448@coredump.intra.peff.net","subject":"Re: [doc] User Manual Suggestion","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-04-29T13:27:11Z","receivedAt":"2009-04-29T13:27:11Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Apr 29, 2009, at 2:34 AM, Jeff King wrote:\n\n> On Fri, Apr 24, 2009 at 08:53:37PM -0400, David Abrahams wrote:\n>\n>>> Actually, it is not the generally of trees that I think is  \n>>> interesting\n>>> there, but the generality of _objects_. That is, each of those  \n>>> things is\n>>> a first-class object, and has a unique name by which it can be\n>>> referred.\n>>\n>> I'm sorry, but I think most people would find that so unremarkable  \n>> that\n>> making a big deal about it would lead to \"what am I missing here\"\n>> confusion.  Maybe a person who's exclusively used CVS (or older)\n>> technologies before coming to Git would be happy to know that, but  \n>> it's\n>> sort of obvious.  In CVS the lack of first-class directories sticks  \n>> out\n>> like a sore thumb.\n>\n> Sadly, I was away from email all weekend and so missed the ensuing  \n> storm\n> in this thread. :) However, I did want to respond to this one point.\n>\n> To me (and I am talking from personal experience, so it really may be\n> _just_ me), an important part of understanding git was understanding  \n> the\n> object storage. That is, half of the idea of git is a big database of\n> content-addressable objects.\n\nAbsolutely, it's important to know that everything is content- \naddressable (which essentially communicates the same important  \ninformation as \"the object's id is a hash of its contents\").  I was  \ntrying to say that the fact that each one is a \"first-class\" object   \nand has a unique name is not particularly remarkable.\n\n--\nDavid Abrahams\nBoostPro Computing\nhttp://boostpro.com\n"},{"id":"112636","messageId":"20090429140520.GA31343@coredump.intra.peff.net","threadId":"19010","inReplyTo":"D1C7AEA3-1565-4E86-AFF0-9EC2F0D79FCC@boostpro.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-29T14:05:20Z","receivedAt":"2009-04-29T14:05:20Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Apr 29, 2009 at 09:27:11AM -0400, David Abrahams wrote:\n\n>> object storage. That is, half of the idea of git is a big database of\n>> content-addressable objects.\n>\n> Absolutely, it's important to know that everything is content-addressable \n> (which essentially communicates the same important information as \"the \n> object's id is a hash of its contents\").  I was trying to say that the \n> fact that each one is a \"first-class\" object  and has a unique name is not \n> particularly remarkable.\n\nI see. I consider those concepts inextricably linked. But I suppose you\ncould explain one without the other.\n\nAnyway, thanks for the perspective.\n\n-Peff\n"},{"id":"112882","messageId":"20090502155348.GB6135@atjola.homenet","threadId":"19010","inReplyTo":"b4087cc50904241701jb78ce50m122fef475b0f1de7@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-05-02T15:53:48Z","receivedAt":"2009-05-02T15:53:48Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.24 19:01:48 -0500, Michael Witten wrote:\n> 2009/4/24 Björn Steinbrink <B.Steinbrink@gmx.de>:\n> >> In fact, I think it's important to note that the notation:\n> >>\n> >>     git show master:Makefile\n> >>\n> >> actually involves a translation from a Unix filesystem address to a\n> >> git object address that is then used to find the relevant data.\n> >\n> > Hm? Resolving master:Makefile means to first find what master is, most\n> > likely the shortname for refs/heads/master. That usually references a\n> > commit object (by its name). The \"<tree-ish>:<path>\" syntax then causes\n> > git to lookup the tree referenced by that commit (again, by its name).\n> > And then the tree entry for \"Makefile\" is looked up, leading to the name\n> > for the object identified by \"master:Makefile\".\n> \n> Firstly, your head is too bound to low-level implementation.\n> \n> Secondly, you've basically just expounded upon what I said. The\n> Makefile part is for humans to write using a filesystem path (address)\n> that is mapped into what I call a git address. The point is that the\n> user is interfacing between two theories of content storage.\n\nSorry, that part missed a few sentences I thought I had written. It was\nmeant to show where the term \"reference\" is used. I just walked along\nyour example, as that was right there, and I didn't have to come up with\nsomething else ;-)\n\nOf course there are two \"parts\", just like scp uses <host>:<path>.\n\n> >> Rather than being hidden, it should be exposed: I think it would be\n> >> beneficial to use the word 'address' rather than 'reference' when\n> >> talking about the SHA-1 names. Then HEAD could be called a pointer\n> >> variable, etc.\n> >\n> > What's wrong with just calling the object name \"object name\"?\n> \n> What's wrong with calling the object address \"object address\"?\n\nThe term \"object name\" is already used in the docs, so you'll have to\nprove that it's bad and needs to be replaced.\n\n> As I've stated: \"address\", \"pointer\", and \"handle\" are an analogy to\n> terminology that has been around for ages. In fact, another name for\n> \"pointer\" is \"reference\".\n\nAFAIK a pointer is just one kind of reference. C++ references are\nanother kind, file descriptors are yet another. A reference is one piece\nof data that lets me access a different piece of data.\n\nAnd there are probably plenty of examples where you could apply that\nanalogy, yet nobody (I know) does. Arrays, database tables, ...\n\nAnd \"memory\" usually means \"RAM\" to me, not \"WORM\"-memory (well,\nactually, you can also delete and then rewrite, but not modify). So the\nanalogy would even hurt my mental model (just like the \"commit --amend\"\ncommand might be consider harmful, because it actually creates a new\ncommit, but some users actually think the original commit is modified).\n\n> >> So, a pointer variable's value is an object address that is the\n> >> location of an object in git 'memory'. I think using this approach\n> >> would make things significantly more transparent.\n> >\n> > But then HEAD would be a pointer pointer variable (symbolic ref), unless\n> > you have a detached HEAD.\n> \n> We call those handles.\n\nIsn't a handle basically an opaque/abstract reference, at least in\n\"modern\" usage? Symvolic references aren't. The user is free to create\nand manipulate them, and gets full access to the things referenced by\nthem. And saying that HEAD is a reference, that might be symbolic is\nIMHO by far easier to understand than saying that HEAD might be a\npointer or a handle.\n\nBjörn\n"},{"id":"112884","messageId":"b4087cc50905021136l5209777bs2209bab385deeef6@mail.gmail.com","threadId":"19010","inReplyTo":"20090502155348.GB6135@atjola.homenet","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-05-02T18:36:35Z","receivedAt":"2009-05-02T18:36:35Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"2009/5/2 Björn Steinbrink <B.Steinbrink@gmx.de>:\n>> As I've stated: \"address\", \"pointer\", and \"handle\" are an analogy to\n>> terminology that has been around for ages. In fact, another name for\n>> \"pointer\" is \"reference\".\n>\n> AFAIK a pointer is just one kind of reference. C++ references are\n> another kind...\n\nActually, a C++ reference is a pointer with restrictions (AFAIK).\n\n> A reference is one piece of data that lets me access a different\n> piece of data.\n\nThe key word there is 'access', which implies some kind of storage (or memory).\n\n>\n> And there are probably plenty of examples where you could apply that\n> analogy, yet nobody (I know) does. Arrays, database tables, ...\n\nWell, this terminology is certainly used with arrays in C, because\narray elements can be accessed with pointers.\n\nAlso, databases use a much different scheme for addressing information\nthan does memory.\n\nHowever, you're probably correct that pointer terminology doesn't\nexist much outside of C/C++ and older languages (Ada?).\n\n>\n> And \"memory\" usually means \"RAM\" to me, not \"WORM\"-memory (well,\n> actually, you can also delete and then rewrite, but not modify).\n\nWell, I don't see how Random Access Memory really conflicts. One\ncertainly can access objects in the object memory/store randomly. The\nmain difference is that the computer store is addressed by location,\nwheras the git store is addressed by content.\n\nAlso, I would say that conceptually deletion is an implementation\ndetail. Because git's object store is content addressable, one could\nthink of it as already containing all possible objects (of course, I'm\nassuming that the 160-bit hash is also an implementation detail; an\ninfinite number of objects implies infinitely large addresses, though\nthe nonsignificant zeros could be disregarded as with real numbers or\nsomething. I don't know, I'm making this up as I go :-D). That the git\ntools ever complain no such object exists is an implementation detail\nresulting from our finite storage in reality.\n\n> So the\n> analogy would even hurt my mental model (just like the \"commit --amend\"\n> command might be consider harmful, because it actually creates a new\n> commit, but some users actually think the original commit is modified).\n\nActually, this is why it's so important to have the underlying\nconcepts at hand. Understanding that objects are simply addressed by\ncontent (that is, objects are immutable) completely extirpates this\nkind of confusion.\n\n>> >> So, a pointer variable's value is an object address that is the\n>> >> location of an object in git 'memory'. I think using this approach\n>> >> would make things significantly more transparent.\n>> >\n>> > But then HEAD would be a pointer pointer variable (symbolic ref), unless\n>> > you have a detached HEAD.\n>>\n>> We call those handles.\n>\n> Isn't a handle basically an opaque/abstract reference, at least in\n> \"modern\" usage? Symvolic references aren't. The user is free to create\n> and manipulate them, and gets full access to the things referenced by\n> them. And saying that HEAD is a reference, that might be symbolic is\n> IMHO by far easier to understand than saying that HEAD might be a\n> pointer or a handle.\n\nFair enough. Call them symbolic pointers; however, I don't really see\nthe problem with pointer pointers.\n\nIn any case, I *think* my point is that it's important to understand\nthat git uses content addressing; at first I was emphatic about the\nidea of 'addressing', so I went with pointer terminology (which works\nquite well, in my opinion). However, I think the 'content' part is\nmore important, which is why 'object hash' is loads better than\n'object name' or 'object id'. Also, at least the documentation could\nsay that 'objects are addressed by their hashes', which says a whole\nlot in one quick sentence about how git works.\n"},{"id":"112888","messageId":"20090502211110.GC6135@atjola.homenet","threadId":"19010","inReplyTo":"b4087cc50905021136l5209777bs2209bab385deeef6@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-05-02T21:11:10Z","receivedAt":"2009-05-02T21:11:10Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.05.02 13:36:35 -0500, Michael Witten wrote:\n> 2009/5/2 Björn Steinbrink <B.Steinbrink@gmx.de>:\n> >> As I've stated: \"address\", \"pointer\", and \"handle\" are an analogy to\n> >> terminology that has been around for ages. In fact, another name for\n> >> \"pointer\" is \"reference\".\n> >\n> > AFAIK a pointer is just one kind of reference. C++ references are\n> > another kind...\n> \n> Actually, a C++ reference is a pointer with restrictions (AFAIK).\n\nI'm not really aware of what the C++ standard says about it, but from a\nusage point of view, they're IMHO different enough to consider them as\ntruly different types of references.\n\n> > And there are probably plenty of examples where you could apply that\n> > analogy, yet nobody (I know) does. Arrays, database tables, ...\n> \n> Well, this terminology is certainly used with arrays in C, because\n> array elements can be accessed with pointers.\n\nBut when you apply the analogy, then the array is the memory, and an\ninteger is an address and an index variable is a pointer.\n\n> Also, databases use a much different scheme for addressing information\n> than does memory.\n\nI don't see any inherent problem in saying that the primary key\ndetermines the address of a row. (It just gets funny when you have a\ntable schema without a primary key *g*)\n\n> > And \"memory\" usually means \"RAM\" to me, not \"WORM\"-memory (well,\n> > actually, you can also delete and then rewrite, but not modify).\n> \n> Well, I don't see how Random Access Memory really conflicts. One\n> certainly can access objects in the object memory/store randomly. The\n> main difference is that the computer store is addressed by location,\n> wheras the git store is addressed by content.\n\nWhen I have a (non const) pointer in C I can write to the memory\nlocation it references. With git, I can't do that. (\"RWRAM\" would have\nbeen more correct, I'm damaged by the common usage of RAM as meaning\nRWRAM).\n\n> Also, I would say that conceptually deletion is an implementation\n> detail.\n\nYeah, thus I put it in parentheses, just to show that, in practise, we\ndon't even have WORM-memory (but still taking the hash collision problem\ninto account, so we need to write once).\n\n> Because git's object store is content addressable, one could\n> think of it as already containing all possible objects (of course, I'm\n> assuming that the 160-bit hash is also an implementation detail; an\n> infinite number of objects implies infinitely large addresses, though\n> the nonsignificant zeros could be disregarded as with real numbers or\n> something. I don't know, I'm making this up as I go :-D). That the git\n> tools ever complain no such object exists is an implementation detail\n> resulting from our finite storage in reality.\n\nI prefer to take the hash collision into account when looking at things\nlike that, but yeah, one could look at it like that.\n\n> > So the analogy would even hurt my mental model (just like the\n> > \"commit --amend\" command might be consider harmful, because it\n> > actually creates a new commit, but some users actually think the\n> > original commit is modified).\n> \n> Actually, this is why it's so important to have the underlying\n> concepts at hand. Understanding that objects are simply addressed by\n> content (that is, objects are immutable) completely extirpates this\n> kind of confusion.\n\nI never disagreed with that, though I put more emphasis on the plain\nobject relationships and their immutability than on the fact that hashes\nare used. Having that part right (how objects work together to form\nhistory) is a large part of what you need to understand all the rest.\n\n> >> >> So, a pointer variable's value is an object address that is the\n> >> >> location of an object in git 'memory'. I think using this approach\n> >> >> would make things significantly more transparent.\n> >> >\n> >> > But then HEAD would be a pointer pointer variable (symbolic ref), unless\n> >> > you have a detached HEAD.\n> >>\n> >> We call those handles.\n> >\n> > Isn't a handle basically an opaque/abstract reference, at least in\n> > \"modern\" usage? Symvolic references aren't. The user is free to create\n> > and manipulate them, and gets full access to the things referenced by\n> > them. And saying that HEAD is a reference, that might be symbolic is\n> > IMHO by far easier to understand than saying that HEAD might be a\n> > pointer or a handle.\n> \n> Fair enough. Call them symbolic pointers; however, I don't really see\n> the problem with pointer pointers.\n\nYou called them handles anyway ;-) But seriously, it's that \"pointer\"\ntriggers C for me. And having an entity that can switch between being a\npointer and a pointer pointer needs casting or a union (or a struct if\nyou want to), not something I'd like to have to think about in my mental\nmodel of git.\n\n> In any case, I *think* my point is that it's important to understand\n> that git uses content addressing; at first I was emphatic about the\n> idea of 'addressing', so I went with pointer terminology (which works\n> quite well, in my opinion). However, I think the 'content' part is\n> more important, which is why 'object hash' is loads better than\n> 'object name' or 'object id'. Also, at least the documentation could\n> say that 'objects are addressed by their hashes', which says a whole\n> lot in one quick sentence about how git works.\n\nHm, like chapter 7 \"Git concepts\"?\n\n>>>>>>\nThe Object Database\n\nWe already saw in the section called “Understanding History: Commits”\nthat all commits are stored under a 40-digit \"object name\". In fact, all\nthe information needed to represent the history of a project is stored\nin objects with such names. In each case the name is calculated by\ntaking the SHA-1 hash of the contents of the object. The SHA-1 hash is a\ncryptographic hash function. What that means to us is that it is\nimpossible to find two different objects with the same name. This has a\nnumber of advantages; among others:\n\n    * Git can quickly determine whether two objects are identical or\n      not, just by comparing names.\n    * Since object names are computed the same way in every repository,\n      the same content stored in two repositories will always be stored\n      under the same name.\n    * Git can detect errors when it reads an object, by checking that\n      the object's name is still the SHA-1 hash of its contents.\n<<<<<<\n\nBjörn\n"},{"id":"112890","messageId":"b4087cc50905021613j1d269757g8c599b484208c188@mail.gmail.com","threadId":"19010","inReplyTo":"20090502211110.GC6135@atjola.homenet","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-05-02T23:13:24Z","receivedAt":"2009-05-02T23:13:24Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"2009/5/2 Björn Steinbrink <B.Steinbrink@gmx.de>:\n>> In any case, I *think* my point is that it's important to understand\n>> that git uses content addressing; at first I was emphatic about the\n>> idea of 'addressing', so I went with pointer terminology (which works\n>> quite well, in my opinion). However, I think the 'content' part is\n>> more important, which is why 'object hash' is loads better than\n>> 'object name' or 'object id'. Also, at least the documentation could\n>> say that 'objects are addressed by their hashes', which says a whole\n>> lot in one quick sentence about how git works.\n>\n> Hm, like chapter 7 \"Git concepts\"?\n\nThat's exactly the problem. It should be in chapter 0.\n\nI also dislike the use of 'name' rather than 'hash'; a name is\nsomething provided by the user, but a hash is something computed. The\nuse of sha[-]1 is even more egregious.\n"},{"id":"112891","messageId":"20090502233232.GD6135@atjola.homenet","threadId":"19010","inReplyTo":"b4087cc50905021613j1d269757g8c599b484208c188@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-05-02T23:32:32Z","receivedAt":"2009-05-02T23:32:32Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.05.02 18:13:24 -0500, Michael Witten wrote:\n> 2009/5/2 Björn Steinbrink <B.Steinbrink@gmx.de>:\n> >> In any case, I *think* my point is that it's important to understand\n> >> that git uses content addressing; at first I was emphatic about the\n> >> idea of 'addressing', so I went with pointer terminology (which works\n> >> quite well, in my opinion). However, I think the 'content' part is\n> >> more important, which is why 'object hash' is loads better than\n> >> 'object name' or 'object id'. Also, at least the documentation could\n> >> say that 'objects are addressed by their hashes', which says a whole\n> >> lot in one quick sentence about how git works.\n> >\n> > Hm, like chapter 7 \"Git concepts\"?\n> \n> That's exactly the problem. It should be in chapter 0.\n\nI'm not opposed to re-ordering stuff. Though I often think that having\ncommands and concepts \"together\" is better.  Maybe we just need that\ntwice? Once the plain data model, and once a \"hands on\" version where\nthe effects of the commands are described in terms of the data model.\n\nThe former \"sucks\" for those that want to just \"dive in\" (but might\nstill be happy to get told what their actions do), the latter sucks when\nyou just want to look something up.\n\nHm?\n\nBjörn\n"},{"id":"112900","messageId":"b4087cc50905021810y28ab2019ob99857670383ba46@mail.gmail.com","threadId":"19010","inReplyTo":"20090502233232.GD6135@atjola.homenet","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-05-03T01:10:14Z","receivedAt":"2009-05-03T01:10:14Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"2009/5/2 Björn Steinbrink <B.Steinbrink@gmx.de>:\n>> > Hm, like chapter 7 \"Git concepts\"?\n>>\n>> That's exactly the problem. It should be in chapter 0.\n>\n> I'm not opposed to re-ordering stuff. Though I often think that having\n> commands and concepts \"together\" is better.  Maybe we just need that\n> twice? Once the plain data model, and once a \"hands on\" version where\n> the effects of the commands are described in terms of the data model.\n>\n> The former \"sucks\" for those that want to just \"dive in\" (but might\n> still be happy to get told what their actions do), the latter sucks when\n> you just want to look something up.\n\nIndeed. I think the key is to split up the documentation for these 2 paths.\n\n    http://marc.info/?l=git&m=124058631814726&w=2\n\nThe mixing of the 2 is what makes everyone unhappy.\n"},{"id":"112902","messageId":"ca433830905021818v5c99cce4q57902eba8ff3d7fc@mail.gmail.com","threadId":"19010","inReplyTo":"b4087cc50905021613j1d269757g8c599b484208c188@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Mark Lodato","fromEmail":"lodatom@gmail.com","sentAt":"2009-05-03T01:18:29Z","receivedAt":"2009-05-03T01:18:29Z","isPatch":false,"sender":{"key":"lodatom@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58860?v=4"},"body":"2009/5/2 Michael Witten <mfwitten@gmail.com>:\n> I also dislike the use of 'name' rather than 'hash'; a name is\n> something provided by the user, but a hash is something computed. The\n> use of sha[-]1 is even more egregious.\n\nWhat about \"identifier\" as a compromise between \"hash\" and \"name\"?\nThis is really what we're talking about - a way of identifying\nobjects.\n"},{"id":"112903","messageId":"b4087cc50905021826o2b259009pda15dcaaa135ca38@mail.gmail.com","threadId":"19010","inReplyTo":"ca433830905021818v5c99cce4q57902eba8ff3d7fc@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2009-05-03T01:26:03Z","receivedAt":"2009-05-03T01:26:03Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Sat, May 2, 2009 at 20:18, Mark Lodato <lodatom@gmail.com> wrote:\n> 2009/5/2 Michael Witten <mfwitten@gmail.com>:\n>> I also dislike the use of 'name' rather than 'hash'; a name is\n>> something provided by the user, but a hash is something computed. The\n>> use of sha[-]1 is even more egregious.\n>\n> What about \"identifier\" as a compromise between \"hash\" and \"name\"?\n> This is really what we're talking about - a way of identifying\n> objects.\n\nIt's the same problem, in my opinion. '[Cryptographic] hash' says so\nmuch more and still remains quite generic.\n\nAlso, continuing with 'sha1' doesn't seem satisfactory:\n\n    http://marc.info/?l=git&m=124068702303042&w=2\n"},{"id":"112904","messageId":"20090503014826.GE6135@atjola.homenet","threadId":"19010","inReplyTo":"b4087cc50905021810y28ab2019ob99857670383ba46@mail.gmail.com","subject":"Re: [doc] User Manual Suggestion","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-05-03T01:48:26Z","receivedAt":"2009-05-03T01:48:26Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.05.02 20:10:14 -0500, Michael Witten wrote:\n> 2009/5/2 Björn Steinbrink <B.Steinbrink@gmx.de>:\n> >> > Hm, like chapter 7 \"Git concepts\"?\n> >>\n> >> That's exactly the problem. It should be in chapter 0.\n> >\n> > I'm not opposed to re-ordering stuff. Though I often think that having\n> > commands and concepts \"together\" is better.  Maybe we just need that\n> > twice? Once the plain data model, and once a \"hands on\" version where\n> > the effects of the commands are described in terms of the data model.\n> >\n> > The former \"sucks\" for those that want to just \"dive in\" (but might\n> > still be happy to get told what their actions do), the latter sucks when\n> > you just want to look something up.\n> \n> Indeed. I think the key is to split up the documentation for these 2 paths.\n> \n>     http://marc.info/?l=git&m=124058631814726&w=2\n> \n> The mixing of the 2 is what makes everyone unhappy.\n\nI'm not sure which part of that email you're referring to (and I'm\ngetting tired, 3:20am...). I'm just seeing the paragraph where Jeff has\nsaid that we have a split, between the tutorial and the manual. And what\nI tried to said, is that we might need the tutorial to be less of a\n\"recipe collection\", but more of a hands-on introduction that actively\nexplains the data model and how data is manipulated by using the\ncommands. And the user manual might become less example oriented,\nfocussing more on concepts, giving examples in addition. So that we have\nboth approaches, hands-on and theoretical, but both keeping the data\nmodel in mind, at least to some extend.\n\nFor example the \"hands on\" version might rather create a \"toy\"\nrepository than importing an existing project right away, to get a\nsmaller scope of things to describe at once, and to be able to show e.g.\nfull \"graphs\" of the early repo as it evolves. Users that simply don't\nwant to care can still skip over the explanations and suffer^Wjust pick\nup the commands. You could e.g. say \"To create a lightweight tag you use\n..., which adds a new reference, while ... adds an annotated tags, which\nis a real tag object, with a message and a tagger and which can possibly\nbe signed using your GPG key.\" And maybe explain the tag object a bit\nfurther.\n\nWhile the manual might, for example, have a section \"Tags\" instead of\nthe current \"Creating tags\"(*), where the different types of tags are\ndescribed, how they fit into the data model, what the different types of\ntags mean, and only then give examples how to create them.\n\nLots of possible work...\n\nBjörn\n\n(*) Why's that in the \"Exploring git history chapter\"? Let's see if I\ncan sort out my local asciidoc problems and find some time to provide\nsome basic patches for that. Though I still haven't managed to get the\none for the git-push man page done... *sigh*\n"}]}