{"thread":{"id":"21659","subject":"Hey - A Conceptual Simplication....","startedAt":"2009-11-18T12:55:45Z","lastAt":"2009-11-20T15:07:47Z","messageCount":25,"participants":["George Dennie","Jan Krüger","Thomas Rast","Jason Sewall","Jonathan del Strother","Jakub Narebski","Linus Torvalds","Björn Steinbrink","Junio C Hamano","Dmitry Potapov","david@lang.hm"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"127839","messageId":"005a01ca684e$71a1d710$54e58530$@com","threadId":"21659","inReplyTo":null,"subject":"Hey - A Conceptual Simplication....","fromName":"George Dennie","fromEmail":"gdennie@pospeople.com","sentAt":"2009-11-18T12:55:45Z","receivedAt":"2009-11-18T12:55:45Z","isPatch":false,"sender":{"key":"gdennie@pospeople.com","avatar":null},"body":"A Clean checkout command might be...\n\nThe Git model does not seem to go far enough conceptually, for some\nunexplainable reason...\n\nIn particular, why is Git not treating the entire working tree as the\nversioned document (qualified of course by the .gitignore file). \n\nInstead, Git is treating a manually maintained list of files within the\nworking tree as the versioned document, this list being initialized and\nmanually amended by the \"Git add/rm/mv\" commands, etc. \n\nThe result is conceptual complexity and rather counter-intuitive behavior.\nFor example, adding and renaming files outside of Git is not considered\nediting the version until you subsequently do a \"Git Add .\" Contrast that\nwith editing or deleting files outside of Git. Yet adding and renaming files\nand folders is a significant part of substantive projects, especially in the\nearly stages and experimental branches.\n\nGranted, this is not a big deal functionally, but what is being lost is\nconceptual simplicity (and consistency, in my book) and conceptual\nsimplicity is a key value point, if not THE key.\n\nAlso can we augment checkout to totally CLEAN the working directory prior to\na restore. If necessary we can augment .gitignore to stipulate those files\nor folders that should be excluded from the cleaning. This suggestion is in\nrecognition of the fact that if you  are not versioning the file, it is\ntypically trash; which becomes the case when the entire working treat is\ntreated as the versioned document.\n\nConsequently, I recommend the following new commands:\n\t\"Git commit -x\"   -- performs a \"Git add .\" then a \"Git commit\"\n\t\"Git checkout -x\" -- that clean the working tree prior to perform a\ncheckout\n\nP.S.\nGreat your work.\n\nGeorge Dennie, BMath\nThe Point Of Sale People\nwww.pospeople.com\nBUS: 416-496-2921\nFAX: 416-496-9496\n"},{"id":"127849","messageId":"57518fd10911180518y4dbb65e2kf28d8ccc88bfb13b@mail.gmail.com","threadId":"21659","inReplyTo":"005a01ca684e$71a1d710$54e58530$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Jonathan del Strother","fromEmail":"maillist@steelskies.com","sentAt":"2009-11-18T13:18:29Z","receivedAt":"2009-11-18T13:18:29Z","isPatch":false,"sender":{"key":"jon.delstrother@bestbefore.tv","avatar":"https://gravatar.com/avatar/754e21ab701c00e2d21fc261187254c34b2a1c0b959d9ee5be1a295990be3081?d=mp&s=160"},"body":"2009/11/18 George Dennie <gdennie@pospeople.com>:\n> A Clean checkout command might be...\n>\n> The Git model does not seem to go far enough conceptually, for some\n> unexplainable reason...\n>\n> In particular, why is Git not treating the entire working tree as the\n> versioned document (qualified of course by the .gitignore file).\n>\n> Instead, Git is treating a manually maintained list of files within the\n> working tree as the versioned document, this list being initialized and\n> manually amended by the \"Git add/rm/mv\" commands, etc.\n>\n> The result is conceptual complexity and rather counter-intuitive behavior.\n> For example, adding and renaming files outside of Git is not considered\n> editing the version until you subsequently do a \"Git Add .\" Contrast that\n> with editing or deleting files outside of Git. Yet adding and renaming files\n> and folders is a significant part of substantive projects, especially in the\n> early stages and experimental branches.\n>\n> Granted, this is not a big deal functionally, but what is being lost is\n> conceptual simplicity (and consistency, in my book) and conceptual\n> simplicity is a key value point, if not THE key.\n>\n> Also can we augment checkout to totally CLEAN the working directory prior to\n> a restore. If necessary we can augment .gitignore to stipulate those files\n> or folders that should be excluded from the cleaning. This suggestion is in\n> recognition of the fact that if you  are not versioning the file, it is\n> typically trash; which becomes the case when the entire working treat is\n> treated as the versioned document.\n>\n> Consequently, I recommend the following new commands:\n>        \"Git commit -x\"   -- performs a \"Git add .\" then a \"Git commit\"\n>        \"Git checkout -x\" -- that clean the working tree prior to perform a\n> checkout\n>\n\n\nPerhaps try 'git commit -a' and 'git checkout -f' ?\n"},{"id":"127840","messageId":"20091118142512.1313744e@perceptron","threadId":"21659","inReplyTo":"005a01ca684e$71a1d710$54e58530$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Jan Krüger","fromEmail":"jk@jk.gs","sentAt":"2009-11-18T13:25:12Z","receivedAt":"2009-11-18T13:25:12Z","isPatch":false,"sender":{"key":"jk@jk.gs","avatar":"https://avatars.githubusercontent.com/u/1774?v=4"},"body":"Hi,\n\n> The result is conceptual complexity and rather counter-intuitive\n> behavior. For example, adding and renaming files outside of Git is\n> not considered editing the version until you subsequently do a \"Git\n> Add .\" Contrast that with editing or deleting files outside of Git.\n> Yet adding and renaming files and folders is a significant part of\n> substantive projects, especially in the early stages and experimental\n> branches.\n\nyet even now, people routinely add huge amounts of files they didn't\nactually want to add, and then have to expend a huge amount of effort\nto get them out of the history again (particularly if that history has\nalready been published).\n\nWhat you are describing is a workflow that is even fuller of potential\nfor wrong turns than the current standard workflow is. If simplicity\nleads to a greater potential for errors, how is it a good thing?\n\nThis kind of workflow actually involves more work for the user. She now\nhas to meticulously maintain an accurate list of ignore patterns,\nparticularly because of this:\n\n> Also can we augment checkout to totally CLEAN the working directory\n> prior to a restore. If necessary we can augment .gitignore to\n> stipulate those files or folders that should be excluded from the\n> cleaning.\n\nSo if I forget to add a certain pattern, my file is lost forever? Uhh...\n\n> This suggestion is in recognition of the fact that if you\n> are not versioning the file, it is typically trash\n\nJust how typical is that, though? I wouldn't want to be the one to\njudge that.\n\nIn light of my concerns, I oppose adding your suggestions to the\nofficial CLI of git and I suggest that you create your own commands to\nenable this kind of workflow. For example:\n\ngit config --global alias.commitx '!git add . && git commit'\ngit config --global alias.checkoutx '!git clean && git checkout'\n\nJan\n"},{"id":"127843","messageId":"200911181430.13537.trast@student.ethz.ch","threadId":"21659","inReplyTo":"005a01ca684e$71a1d710$54e58530$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-11-18T13:30:11Z","receivedAt":"2009-11-18T13:30:11Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"George Dennie wrote:\n>\n> Instead, Git is treating a manually maintained list of files within the\n> working tree as the versioned document, this list being initialized and\n> manually amended by the \"Git add/rm/mv\" commands, etc. \n\nThis feature is called the \"index\", and is not merely a list of the\nfiles, but also their content.  Please read\n\n  http://tomayko.com/writings/the-thing-about-git\n\nfor a nice explanation why this is a good and useful thing.\n\n> \t\"Git commit -x\"   -- performs a \"Git add .\" then a \"Git commit\"\n> \t\"Git checkout -x\" -- that clean the working tree prior to perform a checkout\n\nThat would require supernaturally good maintenance of your .gitignore\nto avoid adding or (worse) nuking files by accident.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"127844","messageId":"31e9dd080911180531r1f693d7bi3d9408ef8219cce0@mail.gmail.com","threadId":"21659","inReplyTo":"005a01ca684e$71a1d710$54e58530$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Jason Sewall","fromEmail":"jasonsewall@gmail.com","sentAt":"2009-11-18T13:31:42Z","receivedAt":"2009-11-18T13:31:42Z","isPatch":false,"sender":{"key":"jasonsewall@gmail.com","avatar":null},"body":"On Wed, Nov 18, 2009 at 7:55 AM, George Dennie <gdennie@pospeople.com> wrote:\n>\n> In particular, why is Git not treating the entire working tree as the\n> versioned document (qualified of course by the .gitignore file).\n>\n> Instead, Git is treating a manually maintained list of files within the\n> working tree as the versioned document, this list being initialized and\n> manually amended by the \"Git add/rm/mv\" commands, etc.\n\nIsn't fastidiously maintaining a .gitignore file to contain everything\nyou *don't* want in the project more confusing than explicitly\nspecifying things you *do* want in the project?\n\n> The result is conceptual complexity and rather counter-intuitive behavior.\n> For example, adding and renaming files outside of Git is not considered\n> editing the version until you subsequently do a \"Git Add .\" Contrast that\n> with editing or deleting files outside of Git. Yet adding and renaming files\n> and folders is a significant part of substantive projects, especially in the\n> early stages and experimental branches.\n>\n> Granted, this is not a big deal functionally, but what is being lost is\n> conceptual simplicity (and consistency, in my book) and conceptual\n> simplicity is a key value point, if not THE key.\n\nIn fact, it's a big deal in functionality, but the utility is in being\nable to to specify exactly what I want to be part of each commit. One\nof git's great features is the ability to specify *exactly* what you\nwant to be part of each commit, down to the line. This means that each\ncommit can be extremely fine grained and represent specific bug fixes\nand or features.\n\nIf you have a bunch of debugging code sitting around in your working\ntree after you've tracked down a problem, you don't want to commit all\nof those printfs, etc. - you want to commit the fix. This has\nramifications from making diffs of history cleaner to making git\nbisect actually useful.\n\n> Also can we augment checkout to totally CLEAN the working directory prior to\n> a restore. If necessary we can augment .gitignore to stipulate those files\n> or folders that should be excluded from the cleaning. This suggestion is in\n> recognition of the fact that if you  are not versioning the file, it is\n> typically trash; which becomes the case when the entire working treat is\n> treated as the versioned document.\n\nThis is even worse. It's already pretty easy to trash your working\ndirectory by reflexively typing git checkout -f, and you want to\n\n> Consequently, I recommend the following new commands:\n>        \"Git commit -x\"   -- performs a \"Git add .\" then a \"Git commit\"\n>        \"Git checkout -x\" -- that clean the working tree prior to perform a\n> checkout\n\nI see that Jan has replied with some loaded guns, *ahem* aliases. Go\nahead and use them, but I recommend you look at the diffs in git.git\nor some other repository that takes advantage of making commits as\ncompact as possible, and learn how to use git add -p.\n\nJason\n"},{"id":"127868","messageId":"008401ca6880$33d7e550$9b87aff0$@com","threadId":"21659","inReplyTo":"20091118142512.1313744e@perceptron","subject":"RE: Hey - A Conceptual Simplication....","fromName":"George Dennie","fromEmail":"gdennie@pospeople.com","sentAt":"2009-11-18T18:51:56Z","receivedAt":"2009-11-18T18:51:56Z","isPatch":false,"sender":{"key":"gdennie@pospeople.com","avatar":null},"body":"Thanks Jan, Jason, Jonathan, and Thomas for your response, your thoughts and\nconcerns are enlightening....\n\nJan Kruger wrote...\n> git config --global alias.commitx '!git add . && git commit'\n> git config --global alias.checkoutx '!git clean && git checkout'\n\nThank you. Being new to git, I did not know that such aliasing was available\nwithin it.\n\nJason Sewell wrote...\n> If you have a bunch of debugging code sitting around in your working tree\nafter you've tracked down a \n> problem, you don't want to commit all of those printfs, etc. - you want to\ncommit the fix. This has \n> ramifications from making diffs of history cleaner to making git bisect\nactually useful.\n\nOne of the concerns I have with the manual pick-n-commit is that you can\nforget a file or two. Consequently, unless you do a clean checkout and test\nof the commit, you don't know that your publishable version even compiles.\nIt seems safer to commit the entirety of your work in its working state and\nthen do a clean checkout from a dedicated publishable branch and manually\nmerge the changes in that, test, and commit.\n\nIt seems the intuitive model is to treat version control as applying to the\nwhole document, not parts of it. In this respect the document is defined by\nthe IDE, namely the entire solution, warts and all. When you start\nselectively saving parts of the document then you are doing two things,\nversioning and publishing; and at the same time. This was a critical flaw in\nolder version control approaches because the software solution document is a\nfile system sub-tree.\n\nWhat you termed the debugging/printf's I would treat as a distinctions\nbetween a debug vs. a release version that may be suitably delineated by\n#define's or preferably separate unit tests assemblies. If I must prune\nprior to committing; however, then it seems reverting spurious printf's may\noffer a more reliable and automatable technique than ensuring that I have\nadded all the new class files, resource files, text files, sub projects,\netc; that may constitute the \"fix.\" Once so selectively reverted I can test\nand commit such a publishable version.\n\nJason Sewell wrote...\n>  Isn't fastidiously maintaining a .gitignore file to contain everything\nyou *don't* want in the project more confusing \n> than explicitly specifying things you *do* want in the project?\n\nThis is git ignore for \"cleaning prior to a check\" and git ignore for\n\"adding to index\" and is not an either or. You would specify what you don't\nwant to version tracked as normal but you can also stipulate what you don't\nwant to be deleted during a clean restore (which should otherwise completely\nwipe the folder prior to restoring a specific commit). This would permit\nembedding non-version elements within the version tree for whatever reason\nyou find necessary.\n\nThomas Rast wrote...\n> That would require supernaturally good maintenance of your .gitignore to\navoid adding or (worse) nuking files by accident.\n\nOn the contrary, the approach would all but eliminate the possibility of\nloss of data since you would not manually (and therefore error prone-ingly)\npruning until after a commit. In fact, one might default automatic commits\n(if required) prior to checkouts or at least an alert system when\nuncommitted changes exists.\n\nThanks again for your input.\n"},{"id":"127869","messageId":"m37htnd3kb.fsf@localhost.localdomain","threadId":"21659","inReplyTo":"008401ca6880$33d7e550$9b87aff0$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-11-18T19:40:45Z","receivedAt":"2009-11-18T19:40:45Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"George Dennie\" <gdennie@pospeople.com> writes:\n\n> Thanks Jan, Jason, Jonathan, and Thomas for your response, your thoughts and\n> concerns are enlightening....\n \n> Jason Sewell wrote...\n>\n> > If you have a bunch of debugging code sitting around in your working tree\n> > after you've tracked down a problem, you don't want to commit all\n> > of those printfs, etc. - you want to commit the fix. This has\n> > ramifications from making diffs of history cleaner to making git\n> > bisect actually useful.\n> \n> One of the concerns I have with the manual pick-n-commit is that you can\n> forget a file or two.\n\nI don't think that this concern is valid.  \n\nThe files which make project are those defined in Makefile or\nequivalent project file, _not_ all files (or even all files of\nspecific type / extension) that do happen to reside in given\ndirectory.  And those files whould be known to git, either added when\nimporting project into git, or added when they were created.  And if\nthey are known it is enough to use \"git commit -a\" to pick all\nchanges.\n\nSo I don't see how you can 'forget a file or two'.\n\nAre those *theoretical* concerns, or is it something that happened to\nyou doring using git?\n\n> Consequently, unless you do a clean checkout and test\n> of the commit, you don't know that your publishable version even compiles.\n> It seems safer to commit the entirety of your work in its working state and\n> then do a clean checkout from a dedicated publishable branch and manually\n> merge the changes in that, test, and commit.\n\nThat's what\n\n  git stash --keep-index\n\nis for.  \n\nThat, and continuous integration repository, with it's hooks.\n\n> \n> It seems the intuitive model is to treat version control as applying to the\n> whole document, not parts of it. In this respect the document is defined by\n> the IDE, namely the entire solution, warts and all.\n\nYes, and IDE has project file which defines which files are in\nproject, just like version control system has it's tracked files.\n\n> When you start\n> selectively saving parts of the document then you are doing two things,\n> versioning and publishing; and at the same time. This was a critical flaw in\n> older version control approaches because the software solution document is a\n> file system sub-tree.\n\nAtomic commits are important, but the distinction between tracked\nfiles, (untracked) ignored files, and files in \"limbo\" state (neither\ntracked nor ignored) is orthogonal to having atomic commits.\n\n> Jason Sewell wrote...\n>\n> >  Isn't fastidiously maintaining a .gitignore file to contain\n> > everything you *don't* want in the project more confusing than\n> > explicitly specifying things you *do* want in the project?  \n> \n> This is git ignore for \"cleaning prior to a check\" and git ignore for\n> \"adding to index\" and is not an either or. You would specify what you don't\n> want to version tracked as normal but you can also stipulate what you don't\n> want to be deleted during a clean restore (which should otherwise completely\n> wipe the folder prior to restoring a specific commit). This would permit\n> embedding non-version elements within the version tree for whatever reason\n> you find necessary.\n\nAnd this is supposedly easier to use?  I don't think so.\n\n> Thomas Rast wrote...\n>\n> > That would require supernaturally good maintenance of your\n> > .gitignore to avoid adding or (worse) nuking files by accident.\n> \n> On the contrary, the approach would all but eliminate the possibility of\n> loss of data since you would not manually (and therefore error prone-ingly)\n> pruning until after a commit. In fact, one might default automatic commits\n> (if required) prior to checkouts or at least an alert system when\n> uncommitted changes exists.\n\nWhat?  I cannot understand you here.\n\nI think that automatic pruning of non-versioned files is _more_ error\nprone than manual deleting of files.  And much more error prone that\njust keeping non-ignored and non-tracked files.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"127870","messageId":"31e9dd080911181152h665d5d9dr5c0736c0ca3234c1@mail.gmail.com","threadId":"21659","inReplyTo":"m37htnd3kb.fsf@localhost.localdomain","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Jason Sewall","fromEmail":"jasonsewall@gmail.com","sentAt":"2009-11-18T19:52:06Z","receivedAt":"2009-11-18T19:52:06Z","isPatch":false,"sender":{"key":"jasonsewall@gmail.com","avatar":null},"body":"Sorry for the 2x post, George; forgot to include the list in my reply....\n\nOn Wed, Nov 18, 2009 at 1:51 PM, George Dennie <gdennie@pospeople.com> wrote:\n[some cleanup of quote line wrapping]\n> Jason Sewall wrote...\n>> If you have a bunch of debugging code sitting around in your\n>> working tree after you've tracked down a problem, you don't want to\n>> commit all of those printfs, etc. - you want to commit the\n>> fix. This has ramifications from making diffs of history cleaner to\n>> making git bisect actually useful.\n\n> One of the concerns I have with the manual pick-n-commit is that you\n> can forget a file or two. Consequently, unless you do a clean\n> checkout and test of the commit, you don't know that your\n> publishable version even compiles.  It seems safer to commit the\n> entirety of your work in its working state and then do a clean\n> checkout from a dedicated publishable branch and manually merge the\n> changes in that, test, and commit.\n\nI find git status very useful in preparing a commit; untracked (and\n'un-ignored') files are listed right there and I can if there are new\nsource files that are not present but not tracked.  You could even add\na 'pre-commit hook' to make sure that you don't have any untracked *.c\n(or whatever) files before you actually make the commit.\n\nAs to 'publishable' version, it's probably a good idea to run 'make\ndistcheck' or the equivalent before making a release anyway.\n\n> It seems the intuitive model is to treat version control as applying\n> to the whole document, not parts of it. In this respect the document\n> is defined by the IDE, namely the entire solution, warts and\n> all. When you start selectively saving parts of the document then\n> you are doing two things, versioning and publishing; and at the same\n> time. This was a critical flaw in older version control approaches\n> because the software solution document is a file system sub-tree.\n\nI find this leads to big, shapeless commits and, as I mentioned\nbefore, it seriously limits the utility of 'git bisect'.  I also fail\nto see how 'selectively saving parts of the document' is versioning\nand publishing - what is the publishing part?  The act of committing\nis one thing (and 'saving parts of the document' is one conceivable\nname for it) and publishing another.  Your workflow may vary, but\nbefore actually 'publishing' (perhaps pushing out to a public repo, or\nmerging into a public branch), it's probably a good idea to test the\ncode with whatever system you use anyway.\n\n> What you termed the debugging/printf's I would treat as a\n> distinctions between a debug vs. a release version that may be\n> suitably delineated by #define's or preferably separate unit tests\n> assemblies. If I must prune prior to committing; however, then it\n> seems reverting spurious printf's may offer a more reliable and\n> automatable technique than ensuring that I have added all the new\n> class files, resource files, text files, sub projects, etc; that may\n> constitute the \"fix.\" Once so selectively reverted I can test and\n> commit such a publishable version.\n\nWhat if you are hacking away and make changes to several parts of the\ncode at once?  Making the commits as fine-grained as possible makes it\neasier to cherry-pick, bisect, and understand the history.\n\nAs to debugging code, I admit I sometimes will use git gui or git add\n-p to stage just what I want and then put whatever is 'left over' in a\nbranch that I might use again later if another bug comes up.  Then I\ncan reset --hard my 'working' branch and the debugging code is gone.\n\n> Jason Sewell wrote...\n>>  Isn't fastidiously maintaining a .gitignore file to contain\n>> everything you *don't* want in the project more confusing than\n>> explicitly specifying things you *do* want in the project?\n>\n> This is git ignore for \"cleaning prior to a check\" and git ignore\n> for \"adding to index\" and is not an either or. You would specify\n> what you don't want to version tracked as normal but you can also\n> stipulate what you don't want to be deleted during a clean restore\n> (which should otherwise completely wipe the folder prior to\n> restoring a specific commit). This would permit embedding\n> non-version elements within the version tree for whatever reason you\n> find necessary.\n\nPerhaps I don't understand your scheme, but it sounds like you're\nadvocating 2 .gitignores:\n\n* .gitignore_track; with everything you don't automatically staged but\n which can be trashed by your cleaning checkout\n* .gitignore_keep; with things you don't want staged but which\n  shouldn't be deleted by git during cleaning\n\nThat seems even more confusing.  I'm actually having trouble seeing\nwhy you want this untracked-file nuking checkout at all.  Care to give\nan example?\n\n> Thomas Rast wrote...\n>> That would require supernaturally good maintenance of your\n>> .gitignore to\n> avoid adding or (worse) nuking files by accident.\n>\n> On the contrary, the approach would all but eliminate the\n> possibility of loss of data since you would not manually (and\n> therefore error prone-ingly) pruning until after a commit. In fact,\n> one might default automatic commits (if required) prior to checkouts\n> or at least an alert system when uncommitted changes exists.\n\nWho is pruning after a commit?  Once nice thing about checkout is that\nit will refuse to move to a different commit if there are files that\nwill get trashed.  Then you can say 'oops, I should stash/commit/nuke\nthat stuff before I change HEAD.\n\nJason\n"},{"id":"127872","messageId":"alpine.LFD.2.00.0911181224110.2793@localhost.localdomain","threadId":"21659","inReplyTo":"005a01ca684e$71a1d710$54e58530$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-11-18T20:36:17Z","receivedAt":"2009-11-18T20:36:17Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 18 Nov 2009, George Dennie wrote:\n> \n> The Git model does not seem to go far enough conceptually, for some\n> unexplainable reason...\n\nOthers already mentioned this, but the concept you missed is the git \n'index', which is actually very central (it is actually the first part of \ngit written, before even the object database) but is something that most \npeople who get started with git can (and do) ignore.\n\nNow, admittedly, for casual use it's not always clear _why_ the index is \nso central, so the fact that you overlooked it is certainly easy to \nunderstand. Just take my word for it: to truly understand git, you do need \nto understand the index.\n\nYou can ignore it for a long time, because one of the primary reasons for \nit existing is about performance. That happens to be a primary goal of \ngit, of course, but some people always think it's \"just performance\". It's \nway more fundamental than that.\n\nSo the way you can start getting used to the index is to think of it as a \nway to avoid having to do a full 'readdir()' on the whole tree to figure \nout what is in there, and avoiding having to read all the files to check \nthat their contents still match.\n\nOf course, if that was _all_ the index did, it could be seen purely as a \ncache, and have no semantic visibility at all. And that's not the case: \nthe index does have real semantic visibility.\n\nThe first time you'll see it is when you decide to stage your changes in \nparts. The index is what allows you to _not_ always commit all your \nchanges exactly because git keeps track of something more than _just_ your \nwhole current working tree.\n\nA special case (but a really useful one) of the \"staging your changes in \nparts\" is when you do merges. Now, most people don't do merges like I do \n(what, average of 5 merges per day, day in and day out), so most people \ndon't care quite as deeply as I do, but if you ever do a merge where 99% \nmerged cleanly, and 1% did not (which is the common case for conflicts), \nyou'll really understand why having a system that keeps track of the parts \nthat merged cleanly is _critical_. \n\nSo for merges, the index keeps track of what merged cleanly, and what \ndidn't, and what the original state for the not-clean stuff was. And as \nsomebody who probably does more merges than likely any other human in the \nhistory of the world, I can state with some authority that any source \ncontrol model that doesn't have this is fundamentally broken.\n\nSo the index is really _really_ important. Even if you can ignore it most \nof the time. And the index is why you don't have a model of \"always just \ntrack the exact tree state\".\n\n\t\t\tLinus\n"},{"id":"127881","messageId":"009401ca68bc$7e4b12b0$7ae13810$@com","threadId":"21659","inReplyTo":"31e9dd080911181152h665d5d9dr5c0736c0ca3234c1@mail.gmail.com","subject":"RE: Hey - A Conceptual Simplication....","fromName":"George Dennie","fromEmail":"gdennie@pospeople.com","sentAt":"2009-11-19T02:03:31Z","receivedAt":"2009-11-19T02:03:31Z","isPatch":false,"sender":{"key":"gdennie@pospeople.com","avatar":null},"body":"Thanks Linus, Jason, and Jakub...\n\nLinus Torvalds wrote....\n>On Wed, 18 Nov 2009, George Dennie wrote:\n>> \n>> The Git model does not seem to go far enough conceptually, for some \n>> unexplainable reason...\n>\n> Others already mentioned this, but the concept you missed is the git 'index', which is actually very \n> central (it is actually the first part of git written, before even the object database) but is something \n> that most people who get started with git can (and do) ignore.\n\nUhmmm, subtle. I hear you. Thanks for the heads up. But before that, I just put these two cents down...\n\nOne of the persistent problems with software documentation is that it often fails to define the \"functional or usage\" model, apart from a dry list of commands. I am sure there are many good reasons for this. For one thing, explaining stuff is hard. Now, I have not had occasions to do merges, as such. So I am finding the justification for the index vague. I am wondering whether this might be a great space to describe the functional model of git in a way that more clearly justifies the index...\n\nSpecifically, can there be a succinct description of the usage or functional model of Git that necessarily incorporates the index. \n\nFor example, the functional notion of the repository seems well defined: a growing web of immutable commits each created as either an isolated commit or more typically an update and/or merger of one or more pre-existing commits. \n\nWith such a description the rest of the structure becomes almost implicit: Commits may be annotated such as with release number labels. Commits that have not been linked to such as by an update or merger remain dangling like loose threads in the web and are called branches. Branches may be given special labels that the repository will then automatically update so as to refer to the latest commit to that branch.\n\nI don't yet have such a clear model for the index. Yes it is a staging platform, but so is the IDE....I'll do more reading.\n\nJason Sewell wrote....\n> I find this leads to big, shapeless commits and, as I mentioned before, it seriously limits the utility \n> of 'git bisect'.  I also fail to see how 'selectively saving parts of the document' is versioning and \n> publishing - what is the publishing part?  The act of committing is one thing (and 'saving...\n\nThe notion of a shapeless commit is curious. Intuitively, I consider a commit as capturing the state of my work at a transactional boundary (i.e. a successful unit test...or even lunch break). However, your characterization of \"shape\" suggest that you are constructing something other than the immediate functionality of the software. Consequently, your software document is not really the solution files alone but also this commit history that you meticulously craft. \n\nFurther, the participating of the IDE is not to compose within itself the committable document but rather to contribute to such a document in pieces. In fact, the closest metaphor to this process/workflow seems to be submitting articles to a magazine; except you are both the writer and editor/graphic artist; and each edition of the magazine becoming the committable version. \n\nWith this metaphor the index does play a clear role as a layout board of sorts for the complete magazine. And also clearly, the IDE does not \"functionally\" edit the entire committable document but rather parts of it. Even though it may effectively have the entirety of the index in its working tree; Git requires that it be submitted to the index which is the true committable document. \n\nIt begs the question, why is the working tree (the IDE document) so closely tied to the repository since it really amounts to a scratch pad. In fact, while the index may be attach to the working tree, the repository can be anywhere and have more than one index attached...yeah, I know, having a personal dedicated repository is cheap. (A great example of how expediency, the proximity of the repository, might obscure the functional model by making what is arbitrary and due to convention appear a functional necessity...; if, in fact, my above conclusion is correct of course :)\n\n> What if you are hacking away and make changes to several parts of the code at once?  Making the commits \n> as fine-grained as possible makes it easier to cherry-pick, bisect, and understand the history.\n\nYou know Jason, it is often hard to isolate my changes to specific files. I have come to appreciate unit tests as a means of delineating changes. However, clearly the historically record of your solution tree is of substantially value to you. It is something I will have to pay closer attention in my case.\n\n> Perhaps I don't understand your scheme, but it sounds like you're advocating 2 .gitignores:\n>\n> * .gitignore_track; with everything you don't automatically staged but  which can be trashed by your cleaning checkout\n> * .gitignore_keep; with things you don't want staged but which shouldn't be deleted by git during cleaning\n\nYep, that may be one implementation...but essentially the current .gitignores list exclusionary filters for the \"git add .\" command. The suggestion was to augment it to also include exclusionary filters for the proposed \"git checkout -clean\" command.  By perhaps prefixing \"+\" and \"-\" symbols to the listed elements you can designate each filter's participation in the \"do not add\" and \"do not delete\" activities, respectively. However, this suggest was with the presumption that the work tree was the committable document, but clearly it is not.\n\n> Who is pruning after a commit?  Once nice thing about checkout is that it will refuse to move to a \n> different commit if there are files that will get trashed.  Then you can say 'oops, I should \n> stash/commit/nuke that stuff before I change HEAD.\n\nNot trashing files is a nice thing by checkout. However, are you referring to changes added to the index or changes made in the working tree but not yet added to the index. Base on my current understanding of the functional model, you would be referring to the index since the working tree is little more than a scratch pad. The pruning comment was in recognition that the working tree was not expected to be committable in its entirety.\n\nGeorge.\n\nThanks again for your input and if you have the time I welcome your response.\n"},{"id":"127888","messageId":"20091119074226.GA23304@atjola.homenet","threadId":"21659","inReplyTo":"009401ca68bc$7e4b12b0$7ae13810$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-11-19T07:42:34Z","receivedAt":"2009-11-19T07:42:34Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.11.18 21:03:31 -0500, George Dennie wrote:\n> Jason Sewell wrote....\n> > I find this leads to big, shapeless commits and, as I mentioned\n> > before, it seriously limits the utility of 'git bisect'.  I also\n> > fail to see how 'selectively saving parts of the document' is\n> > versioning and publishing - what is the publishing part?  The act of\n> > committing is one thing (and 'saving...\n> \n> The notion of a shapeless commit is curious. Intuitively, I consider a\n> commit as capturing the state of my work at a transactional boundary\n> (i.e. a successful unit test...or even lunch break). However, your\n> characterization of \"shape\" suggest that you are constructing\n> something other than the immediate functionality of the software.\n> Consequently, your software document is not really the solution files\n> alone but also this commit history that you meticulously craft. \n\nYour \"lunch break\" as a transaction boundary is a great example of\nsomething that probably most people on this list would consider to\ncreate commits that need rewriting before publishing them. Let's take an\nextreme example:\n\nYou work on adding a feature to some webmail site that adds colors to\nthe mail being displayed, using different colors for the headers, quoted\nsections and the text from the sender. The colors should be configurable\nby the user.\n\n*work*\ngit commit -m \"Go for a coffee\"\n*work*\ngit commit -m \"Lunch break\"\n*work*\ngit commit -m \"Meeting\"\n*work*\ngit commit -m \"Time to go home\"\n\n*come back to work*\n*work*\ngit commit -m \"Finished the mail coloring support\"\n\nThis gives you:\n\n* Finished the mail coloring support\n|\n* Time to go home\n|\n* Meeting\n|\n* Lunch break\n|\n* Go for a coffee\n\nSuch a history is basically completely useless. It's (ab)using the VCS\nas a plain code dump. In a week, you'll be able to see that you had a\nmeeting that day, but it doesn't tell you anything about what you did to\nthe project. And even with less \"insane\" commit messages, the\n\"transactional boundaries\" are totally arbitrary. They're aligned to\nthings you did that have absolutely nothing to do with the stuff you're\ntracking in your VCS.\n\nA far more useful history might look like this:\n\n* Colorize quoted text in a mail, depending on its quoting depth\n|\n* Parse mails into a tree structure to represent sections of quoted text\n|\n* Colorize mail headers\n|\n* Add support for the user to change the colors used for mails\n|\n* Add configuration variable for the colors used for mails\n\n\nAt each step, something functionally changed about the software. The\ncommit messages tell you something about how the software evolved. And\nif you get bogus values for the colors in the configuration, you can be\n90% sure, by only looking at the commit messages, that you have a bug in\nthe \"Add support for the user to change the colors ...\" commit, and not\nin one of the others. So you can run \"git show $that_commit\" to see the\ndiff of the changes you made in that commit and quickly check them for\nyour bug.\n\nAnd while that's not sooo useful for commits that added new\nfunctionality, it's extremely useful for commits that just made small\nchanges to existing functionality. Finding a bug in a large piece of\ncode (say 2000 lines) isn't trivial. But if you know that a commit that\nchanged 5 lines in that code is responsible for the breakage, all you\nhave to do is to identify the faulty change, which is a lot easier.\n\nAnd with a large history, where it's not obvious in which commit\nsomething got broken, \"git bisect\" can help to quickly find the bad\ncommit. Now consider \"git bisect\" finding your \"Lunch break\" commit.\nLooking at the commit message tells nothing. The diff is pretty much\narbitrary, might be huge. Not much help. Finding the \"Add support for\nthe user to change the colors ...\" commit already tells you something\njust because of the commit message. And the diff is about just one\nspecific change. It's all nicely separated, and that's a huge value.\n\nUsing git and producing nice commits is about _documenting_ the history\nof your code. And having small, self-contained and well separated\ncommits is key to that.\n\n\nAnd the index can be a great help with that. Given the above example,\nyou might already have some code to use the configured colors, just for\ntesting, so things aren't so boring. Maybe even some hack-up of the\ncode you'll be using later. If that part of the code would be committed\nright away, you'd mess up your commit, because it wouldn't be about a\nsingle change anymore, but would also have your testing code in there.\nBad.\n\nBut you don't want to throw the testing code away either, because it's\nuseful right now, and you might need it later, because it might evolve\ninto the final code used for the actual coloring. So, what now? You have\ncode that you want to commit, and some code you don't want to commit,\nand which needs to go away temporarily, so you can test without it. No\nproblem, here comes the index.\n\nSay you have:\nconfig.c     # Has changes for the colors\nshow_mail.c  # Has changes to use the colors\nwhatever.c   # Has some changes for both\n\nYou do:\ngit add config.c       # Add to the index\ngit add -p whatever.c  # Only add some hunks to the index\n\nSo now the index has what you want to commit, and the working tree still\nhas everything.\n\ngit stash save --keep-index\n\nNow your working tree and index only have the things you want to commit.\nYou run your unit tests, everythings fine. You commit and get a nice\nclean commit, for which you write a useful commit message.\n\ngit stash pop\n\nYou've got your changes back that you didn't want to commit just yet,\nand you can continue working.\n\n\nAnother use-case I have found for myself is to use the index to separate\nreviewed and not-yet-reviewed changes. Before I commit, I always review\nthe diff of the things I'm going to commit. So I start out with \"git\ndiff\" and start reading. When I finished reviewing a file, I can do \"git\nadd $that_file\", so the diff for that file will no longer be shown by\n\"git diff\". That nicely cuts down the size of the \"git diff\" output to\nthings I'm still interested in. Quite useful when you are forced to do a\nlarge commit, because you did some refactoring. If I find a bug during\nthe review, I can fix that and re-run \"git diff\", which will only show\nchanges to me that I didn't declare as \"good\" already by adding them to\nthe index.\n\n\nSure, it takes some pratice and discipline to generate a nice, useful\nhistory. But that's not much different from writing code. Others will\nhate you for writing unreadable spaghetti code, and so will they hate\nyou for producing a useless history that tells them that you had lunch,\ninstead of telling them what you did to the code ;-)\n\nBjörn\n"},{"id":"127896","messageId":"200911191127.28768.jnareb@gmail.com","threadId":"21659","inReplyTo":"009401ca68bc$7e4b12b0$7ae13810$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-11-19T10:27:27Z","receivedAt":"2009-11-19T10:27:27Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Side-note: you are employing very strange line wrapping... you should\nword wrap your lines so they do not exceed 70-76 characters, and you\nshould not (except when required for readability) rewrap quoted text.\n\nOn Thu, 19 Nov 2009, George Dennie wrote:\n> Thanks Linus, Jason, and Jakub...\n> \n> Linus Torvalds wrote....\n>>On Wed, 18 Nov 2009, George Dennie wrote:\n>>> \n>>> The Git model does not seem to go far enough conceptually, for some \n>>> unexplainable reason...\n>>\n>> Others already mentioned this, but the concept you missed is the git\n>> 'index', which is actually very central (it is actually the first\n>> part of git written, before even the object database) but is\n>> something that most people who get started with git can (and do)\n>> ignore. \n> \n> Uhmmm, subtle. I hear you. Thanks for the heads up. But before that,\n> I just put these two cents down... \n\n> [...] Now, I have not had occasions to do merges, as such. So I am\n> finding the justification for the index vague. [...]\n\nErrr... you didn't do any merges?  What is then your experience with\nusing version control, then?\n\n\nAs for using index during merge: merge is joining two (or more) lines\nof history (lines of development), bringing contents of another branch\ninto current branch.  Some of changes are independent, for example\nif one branch changes one file, and other branch changed other file.\nThis is so called trivial merge, example of tree-level merge.  Even\nif branches merged touch the same file, if changes were made in separate\nsections of file git can merge changes (using three-way merge / diff3\nalgorithm).\n\nThe problem starts if there are changes which touch the same sections\nof a file.  This generates so called merge conflict (contents conflict),\nand you have to resolve such conflict manually.\n\nDuring merge index helps to manage information about yet unmerged parts.\nLet's assume for example that you made a mistake in merge resolution in\nsome file, and you want to scratch your attempt and try it anew. \nWithout index it would be very hard to do without trashing resolutions \nof other conflicts.\n\n> For example, the functional notion of the repository seems well\n> defined: a growing web of immutable commits each created as either\n> an isolated commit or more typically an update and/or merger of\n> one or more pre-existing commits.\n\nIf by \"web\" you mean DAG (Directed Acyclic Graph) of commits, then\nyes, it is _part_ of repository.\n\nThere are also refs (branches, tags, remote-tracking branches), which \nare also part of repository, very important part.  Those are named\nreferences into DAG of commits.\n\n\nAs to commits being created as update of existing commit or from \nscratch: that would depend on the way of development.  Merge commits\nare much, much more rare than ordinary commits (especially that git \nfavors fast-forwards by default when there is no need for merge).\n\n> \n> With such a description the rest of the structure becomes almost\n> implicit: Commits may be annotated such as with release number labels.\n> Commits that have not been linked to such as by an update or merger\n> remain dangling like loose threads in the web and are called branches.\n> Branches may be given special labels that the repository will then\n> automatically update so as to refer to the latest commit to that\n> branch.      \n\nAlmost right.\n \n> I don't yet have such a clear model for the index. Yes it is a staging\n> platform, but so is the IDE....I'll do more reading. \n\nThe index is area where you prepare commits, if needed.  But you\ndon't need to care that there is something like the index, and prepare\nyour commits in working area.  But when you need it, it is there.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"127926","messageId":"00d401ca6954$a29fa020$e7dee060$@com","threadId":"21659","inReplyTo":"20091119074226.GA23304@atjola.homenet","subject":"RE: Hey - A Conceptual Simplication....","fromName":"George Dennie","fromEmail":"gdennie@pospeople.com","sentAt":"2009-11-19T20:12:35Z","receivedAt":"2009-11-19T20:12:35Z","isPatch":false,"sender":{"key":"gdennie@pospeople.com","avatar":null},"body":"Thanks Jakub Narebski and Björn Steinbrink...Nice description Björn.\n\nI think an important piece of conceptual information missing from the docs\nis a concise list of the conceptual properties defining the context of the\nworking tree, index, and repository during normal use. This itemization\nwould go far in explaining the synergies between the various commands. \n\nFunctionally, all the commands merely manipulate these properties. If these\nproperties were summarize in context one would expect that would represent a\nvery complete functional model of Git. A user could review the description\nfigure what they wanted to do and then find the command(s) to accomplish it.\n\n\nPresently this knowledge is accreted over time as oppose to merely being\nread and in the space of a few minutes \"groked\" (of course it could be that\nI am particularly limited :).\n\nFor example, towards a functional model, is this close? (note: all\nproperties can be blank/empty)...\n\nREPOSITORIES\n\tCollection of Commits\n\tCollection of Branches\n\t\t-- collection of commits without children\n\t\t-- as a result each commits either augments\n\t\t-- and existing branch or creates a new one\n\tMaster Branch\n\t\t-- typically the publishable development history\n\nINDEX\n\tCollections of Parent/Merge Commits\n\t\t-- the commit will use all these as its parent\n\n\tStaged Commit \n\t\t-- these changes are shown relative to the working tree\n\n\tDefault Branch\n\t\t-- the history the staged commit is suppose to augment\n\n\tCollection of Stashes\n\t\t-- these are not copies of the working tree since they\n\t\t-- only contain \"versioned\" files/folders and so is not\n\t\t-- a backup\n\nWORKING_TREE\n\tCollection of Files and Folders\n\t\n\nAs far as I can tell, the working tree is not suppose to be stateful, but it\nseems the commands treat it as such.\n\nWhat is interesting is that branches serve to encourage a serialized view of\ncommits. More than structure, they are like books in a library narrating a\ndevelopment story. Consequently, and interestingly, they are as much the\npurpose of the repository as the commits they organize...which is\ninteresting.\n\n\nAgain, thanks for your patients.\n\nGeorge.\n"},{"id":"127931","messageId":"7vocmy9pdm.fsf@alter.siamese.dyndns.org","threadId":"21659","inReplyTo":"00d401ca6954$a29fa020$e7dee060$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-11-19T21:27:33Z","receivedAt":"2009-11-19T21:27:33Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"George Dennie\" <gdennie@pospeople.com> writes:\n\n> REPOSITORIES\n> \tCollection of Commits\n\nOk.\n\n> \tCollection of Branches\n> \t\t-- collection of commits without children\n\nWrong.\n\n> \t\t-- as a result each commits either augments\n> \t\t-- and existing branch or creates a new one\n\nOk.\n\n> \tMaster Branch\n> \t\t-- typically the publishable development history\n\nNot necessarily.\n\n> INDEX\n> \tCollections of Parent/Merge Commits\n> \t\t-- the commit will use all these as its parent\n\nWrong.\n\n> \tStaged Commit \n> \t\t-- these changes are shown relative to the working tree\n\nA new word for me.  I doubt we need to have such a concept.\n\n> \tDefault Branch\n> \t\t-- the history the staged commit is suppose to augment\n\nWe typically call it \"the current branch\".  It is \"the branch whose tip\nwill advance by one commit when you make a new commit\" and determined by\nHEAD.\n\n> \tCollection of Stashes\n> \t\t-- these are not copies of the working tree since they\n> \t\t-- only contain \"versioned\" files/folders and so is not\n> \t\t-- a backup\n\nI think it is better to say what these _are_, instead of saying what they\nare not.  These are not yoghurt cups, these are nor bicycles, these are\nnot knitting needles.  Listing what they are not does not give you more\ninformation.\n\n> WORKING_TREE\n> \tCollection of Files and Folders\n\nOk.\n\n> As far as I can tell, the working tree is not suppose to be stateful, but it\n> seems the commands treat it as such.\n\nI am not sure what you are trying to say by \"stateful\" here.  A work tree\nhas files and directories, and if you edit one of the files of course it\nchanges its state.\n\n----------------------------------------------------------------\n\nA branch is just a pointer to one commit (or nothingness, if it is unborn,\nbut that is such a special case you do not have to worry about yet until\nyou understand git more).\n\nThe commit can have many children, but you do not care about them when\nlooking at the branch, as there is no \"parent-to-children\" pointer.\n\nThe pointer that represents a branch moves to another commit by\ndifferent operations.\n\n - If you make a new commit while on the branch, it points to the new\n   commit.  This is the most typical, and is done by many every-day\n   commands, such as \"commit\", \"am\", \"merge\", \"cherry-pick\", \"revert\".\n\n   Typically the new commit B is a direct child of the commit the branch\n   used to point at A, and B has A as its first parent.\n\n - There are commands that let you violate the above, i.e. you can change\n   what commit the branch pointer points at, and the new commit A does not\n   have to be a direct child of the commit currently pointed by the\n   branch.  \"reset\" and \"rebase\" are examples of such commands and are to\n   rewrite the history.\n\nThere is the \"current branch\" that you are on.  It is recorded in HEAD\n(cat .git/HEAD to see it).  When you create a new commit, the tip of the\nbranch HEAD points at is updated to point at the new commit.  Since the\nnew commit is made a direct child of the current commit, this will appear\nto the users as \"advancing the branch\".\n\nThe state (contents of files and symlinks together with where they are in\nthe tree) to be commited next is recorded in the index.  \"git add\" and\nfriends are used to update this state in the index, and \"git diff\" with\nvarious options allow you to view the difference between this state and\nwork tree or arbitrary commit.\n"},{"id":"127944","messageId":"200911200149.19528.jnareb@gmail.com","threadId":"21659","inReplyTo":"00d401ca6954$a29fa020$e7dee060$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-11-20T00:49:18Z","receivedAt":"2009-11-20T00:49:18Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Thu, 19 Nov 2009, George Dennie wrote:\n\n> Thanks Jakub Narebski and Björn Steinbrink...Nice description Björn.\n> \n> I think an important piece of conceptual information missing from the docs\n> is a concise list of the conceptual properties defining the context of the\n> working tree, index, and repository during normal use. This itemization\n> would go far in explaining the synergies between the various commands.\n\nIf you didn't find sufficient description of underlying concepts behind\ngit in \"Git User's Manual\" (distributed with Git), \"Git Community Book\"\nor \"Pro Git\", take a look at the following documents:\n\n * \"Git for Computer Scientists\"\n * \"Git From Bottom's Up\"\n * \"The Git Parable\"\n\n> Functionally, all the commands merely manipulate these properties. If these\n> properties were summarize in context one would expect that would represent a\n> very complete functional model of Git. A user could review the description\n> figure what they wanted to do and then find the command(s) to accomplish it.\n\nI disagree.  While understanding underlying concepts of Git helps with\nfinding a way to get what one wants to achieve, I don't think that the way\npresented here would work in practice.\n\n> Presently this knowledge is accreted over time as oppose to merely being\n> read and in the space of a few minutes \"groked\" (of course it could be that\n> I am particularly limited :).\n\nIt is documented, see referenced mentioned above.\n\n> For example, towards a functional model, is this close? (note: all\n> properties can be blank/empty)...\n> \n> REPOSITORIES\n> \tCollection of Commits\n\nDirect Acyclic Graph of Commits, where edges in graph point from commit\nto zero or more its parents.\n\n> \tCollection of Branches\n> \t\t-- collection of commits without children\n\nErrr... what?  Commit doesn't *have* [pointer to] children.  Also branch\ncan point to commit for which there exists other commit which has given\ncommit as parent (up-to-date or fast-forward situation, e.g.)\n\n\n    a---b---c            <--- branch_a\n             \\\n              \\-d---e    <--- branch_b\n\nBranches (or branch heads / branch tips) are named references into DAG\nof commits, points where DAG of commits grow.\n\n> \t\t-- as a result each commits either augments\n> \t\t-- and existing branch or creates a new one\n\nCommits do not create a new branch.  New commits must be crated on\nexisting branch (or on unnamed branch aka detached HEAD, but that is\nadvanced usage).\n\n> \tMaster Branch\n> \t\t-- typically the publishable development history\n\nTANSTAAMB. There ain't such thing as a master branch. ;-)))))\n\nWell, at least not in a sense of there being a branch that is a trunk\nbranch distinguished by _technical_ means.\n\n> \n> INDEX\n> \tCollections of Parent/Merge Commits\n> \t\t-- the commit will use all these as its parent\n\nNo.  The index is set of versions of files (blobs) that would go as\na contents (tree) of a next commit (if you use \"git commit', not \n\"git commit -a\").\n\n> \n> \tStaged Commit \n> \t\t-- these changes are shown relative to the working tree\n\nErrr.... what?\n\n> \n> \tDefault Branch\n> \t\t-- the history the staged commit is suppose to augment\n\nErrr... what?\n\nIf by \"default branch\" you mean \"current branch\", it is currently checked\nout branch, where new commit would go, pointed by HEAD symbolic reference.\n\n\n> WORKING_TREE\n> \tCollection of Files and Folders\n> \t\n> \n> As far as I can tell, the working tree is not suppose to be stateful, but it\n> seems the commands treat it as such.\n\nStateful?\n\nWorking tree / working area is a working area.  It can be disconnected from\nrepository via core.worktree, --work-tree option and GIT_WORK_TREE \nenvironment, see also contrib/workdir/git-new-workdir\n\n\n> Again, thanks for your patients.\n\npatience.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"127952","messageId":"20091120013545.GA22556@dpotapov.dyndns.org","threadId":"21659","inReplyTo":"008401ca6880$33d7e550$9b87aff0$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2009-11-20T01:35:46Z","receivedAt":"2009-11-20T01:35:46Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Wed, Nov 18, 2009 at 01:51:56PM -0500, George Dennie wrote:\n> \n> One of the concerns I have with the manual pick-n-commit is that you can\n> forget a file or two.\n\nIt is more difficult to make this mistake with Git than many others\nVCSes, because Git shows the list of files that are changed but not\ncommitted as well as the list of untracked files when you try to commit\nsomething. So, it has never been a real issue for me in practice...\n\n> Consequently, unless you do a clean checkout and test\n> of the commit, you don't know that your publishable version even compiles.\n\nIf you want to be sure that clean checkout will be compiled, the only\nway to guarantee that is to do a clean checkout. Even if you commit all\nfiles except those that are specified in .gitignore, it is not enough to\nbe sure that a clean checkout will be compiled... But in most cases, you\ndo not need to do that to be *reasonable* sure that a clean checkout\nwill be compiled later, and if you have any doubts, you can do a clean\ncheckout and testing _after_ committing your changes. There is no reason\nto be afraid to commit something that may not work if you can amend that\nlater (until you publish your changes).\n\n> It seems safer to commit the entirety of your work in its working state and\n> then do a clean checkout from a dedicated publishable branch and manually\n> merge the changes in that, test, and commit.\n\nMaybe I did not understand your words, but I am not sure what is gained\nin this way... Clearly there is no reason to publish a work that you\nhave not tested yet. And no one cares about crap that you keep in your\nworking tree either... So, a better approach is to commit your changes\nas a series of patches that can be reviewed easily, then do all testing\nand then publish them for integration with the main development branch.\n\n> \n> It seems the intuitive model is to treat version control as applying to the\n> whole document, not parts of it. In this respect the document is defined by\n> the IDE, namely the entire solution, warts and all.\n\nThis is a very bogus idea. If you want to preserve all warts etc, you\njust do backup of the whole disk and now you have a state that can be\ncompiled any time later (provided that your hardware do not change too\nmuch). In my experience, in most cases when I was not able to compile\nan old version were caused not by forgetting to commit something, but\nchanging in the environment (like new compiler, new libraries, etc).\n\nBut when your commits are fine-grained, you can always cherry-pick the\ncorresponding fix-up and compile this old version if it is necessary.\n\nIn my experience, the value of VCS history is the ability to look at it\n(sometimes many years later) and understand who wrote this line and why.\nAlso, nearly all cases when I had to compile some old version were due\nto bisecting some tricky bug. In both cases, having fine-grained commits\nwas crucial to success.\n\n> When you start\n> selectively saving parts of the document then you are doing two things,\n> versioning and publishing; and at the same time.\n\nNo, you don't. Committing some changes and publishing them are two\nseparated operations in Git, and that it is pretty much fundamental.\nNormally, you commit changes in a few separated patches, review them to\nmake sure that changes match commit messages, do all testing, and only\nthen you publish them.\n\n\nDmitry\n"},{"id":"127965","messageId":"20091120014843.GB22556@dpotapov.dyndns.org","threadId":"21659","inReplyTo":"009401ca68bc$7e4b12b0$7ae13810$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2009-11-20T01:48:44Z","receivedAt":"2009-11-20T01:48:44Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Wed, Nov 18, 2009 at 09:03:31PM -0500, George Dennie wrote:\n> \n> For example, the functional notion of the repository seems well\n> defined: a growing web of immutable commits each created as either an\n> isolated commit or more typically an update and/or merger of one or\n> more pre-existing commits. \n\nIn Git, commits are not immutable. One thing that many Git users do\nis git-rebase, which in essense is re-writing or re-ordering exising\ncommits. So, you can change history in Git, but you should never change\nthe published history. (Of course, that leads to the question what is\nconsidered as published history. For instance, commits merged on the\nproposed-updates branch are usually not considered to be \"published\",\nso they can be re-written or discarded later).\n\nSo, the correct way to use Git is to find the right balance between\nthe need to clean up after mistakes (using git-rebase) and not doing\ntoo much, so you will not lose important history or create problems\nfor other peoples.\n\n> \n> The notion of a shapeless commit is curious. Intuitively, I consider a\n> commit as capturing the state of my work at a transactional boundary\n> (i.e. a successful unit test...or even lunch break).\n\nNo, it is not what Git commits were intended for. In Git, a commit is\na change intended to achieve some goal. Basically, you send a patch\nto maintainer, and you should explain what this patch does and why it\nis useful... If your explanation is \"I have a lunch break now\", it is\nvery bad explanation, thus a bad patch.\n\n\nDmitry\n"},{"id":"127966","messageId":"alpine.DEB.2.00.0911191754540.10307@asgard.lang.hm","threadId":"21659","inReplyTo":"20091120014843.GB22556@dpotapov.dyndns.org","subject":"Re: Hey - A Conceptual Simplication....","fromName":"","fromEmail":"david@lang.hm","sentAt":"2009-11-20T01:55:21Z","receivedAt":"2009-11-20T01:55:21Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 20 Nov 2009, Dmitry Potapov wrote:\n\n> On Wed, Nov 18, 2009 at 09:03:31PM -0500, George Dennie wrote:\n>>\n>> For example, the functional notion of the repository seems well\n>> defined: a growing web of immutable commits each created as either an\n>> isolated commit or more typically an update and/or merger of one or\n>> more pre-existing commits.\n>\n> In Git, commits are not immutable. One thing that many Git users do\n> is git-rebase, which in essense is re-writing or re-ordering exising\n> commits. So, you can change history in Git, but you should never change\n> the published history. (Of course, that leads to the question what is\n> considered as published history. For instance, commits merged on the\n> proposed-updates branch are usually not considered to be \"published\",\n> so they can be re-written or discarded later).\n>\n> So, the correct way to use Git is to find the right balance between\n> the need to clean up after mistakes (using git-rebase) and not doing\n> too much, so you will not lose important history or create problems\n> for other peoples.\n\nthe typical advice is to clean up before you make changes public, but not \nafterwords.\n\nDavid Lang\n\n>>\n>> The notion of a shapeless commit is curious. Intuitively, I consider a\n>> commit as capturing the state of my work at a transactional boundary\n>> (i.e. a successful unit test...or even lunch break).\n>\n> No, it is not what Git commits were intended for. In Git, a commit is\n> a change intended to achieve some goal. Basically, you send a patch\n> to maintainer, and you should explain what this patch does and why it\n> is useful... If your explanation is \"I have a lunch break now\", it is\n> very bad explanation, thus a bad patch.\n>\n>\n> Dmitry\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n"},{"id":"127967","messageId":"20091120023103.GC22556@dpotapov.dyndns.org","threadId":"21659","inReplyTo":"00d401ca6954$a29fa020$e7dee060$@com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2009-11-20T02:31:04Z","receivedAt":"2009-11-20T02:31:04Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Thu, Nov 19, 2009 at 03:12:35PM -0500, George Dennie wrote:\n> \n> I think an important piece of conceptual information missing from the docs\n> is a concise list of the conceptual properties defining the context of the\n> working tree, index, and repository during normal use. This itemization\n> would go far in explaining the synergies between the various commands. \n\nSpeaking about \"normal use\"... I suggest you read about Git workflows:\n\n$ git help gitworkflows\n\n> \n> Functionally, all the commands merely manipulate these properties. If these\n> properties were summarize in context one would expect that would represent a\n> very complete functional model of Git. A user could review the description\n> figure what they wanted to do and then find the command(s) to accomplish it.\n\nIt is like to say that driving a car merely means to manipulate its\ncomponents, so if these components were summarized, it would be all\nthat one needs to know to drive a car...\n\nWhile I don't dispute that basic understanding of key Git concepts is\nimportant, understanding of a typical Git workflow cannot be deduced\nfrom knowledge of separate parts. Now if I were to describe Git just in\na few words, I would say that Git repository is just a DAG of objects,\nthe working tree is the place where you work, and the index is what\nhelps you to create fine-grained commits and do merges. But it says\nvery little (if anything) about how to use it.\n\n\nDmitry\n"},{"id":"127968","messageId":"20091120023540.GA17796@atjola.homenet","threadId":"21659","inReplyTo":"20091120014843.GB22556@dpotapov.dyndns.org","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-11-20T02:35:40Z","receivedAt":"2009-11-20T02:35:40Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.11.20 04:48:44 +0300, Dmitry Potapov wrote:\n> On Wed, Nov 18, 2009 at 09:03:31PM -0500, George Dennie wrote:\n> > \n> > For example, the functional notion of the repository seems well\n> > defined: a growing web of immutable commits each created as either an\n> > isolated commit or more typically an update and/or merger of one or\n> > more pre-existing commits. \n> \n> In Git, commits are not immutable.\n\nCommit _are_ immutable. Like all git objects (blob, tree, commits, tag).\n\"Rewriting\" history actually means creating a new history (adding\nobjects), and then changing a ref (most often a branch head) to\nreference the new instead of the old history.\n\nBjörn\n"},{"id":"127969","messageId":"20091120025617.GD22556@dpotapov.dyndns.org","threadId":"21659","inReplyTo":"alpine.DEB.2.00.0911191754540.10307@asgard.lang.hm","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2009-11-20T02:56:18Z","receivedAt":"2009-11-20T02:56:18Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Thu, Nov 19, 2009 at 05:55:21PM -0800, david@lang.hm wrote:\n> On Fri, 20 Nov 2009, Dmitry Potapov wrote:\n>\n>> So, the correct way to use Git is to find the right balance between\n>> the need to clean up after mistakes (using git-rebase) and not doing\n>> too much, so you will not lose important history or create problems\n>> for other peoples.\n>\n> the typical advice is to clean up before you make changes public, but not \n> afterwords.\n\nTrue, except patches may get additional clean up or improvements based\non review feedback, or even get some small fix-ups while they live on\n'pu'. But re-writing something that other people may base their work on\nis clearly wrong. On the other hand, rebasing a large series of patches\neven if it has never been published may be a wrong way to go, because\nyou replace well tested states with some others, which were not tested.\nSo if it is a long and complex series of patches, chances are high that\nyou can break something in it. So, it requires some judgement when to\nuse git-rebase and when git-merge.\n\n\nDmitry\n"},{"id":"127970","messageId":"20091120030801.GE22556@dpotapov.dyndns.org","threadId":"21659","inReplyTo":"20091120023540.GA17796@atjola.homenet","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2009-11-20T03:08:02Z","receivedAt":"2009-11-20T03:08:02Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Fri, Nov 20, 2009 at 03:35:40AM +0100, Björn Steinbrink wrote:\n> On 2009.11.20 04:48:44 +0300, Dmitry Potapov wrote:\n> > On Wed, Nov 18, 2009 at 09:03:31PM -0500, George Dennie wrote:\n> > > \n> > > For example, the functional notion of the repository seems well\n> > > defined: a growing web of immutable commits each created as either an\n> > > isolated commit or more typically an update and/or merger of one or\n> > > more pre-existing commits. \n> > \n> > In Git, commits are not immutable.\n> \n> Commit _are_ immutable. Like all git objects (blob, tree, commits, tag).\n> \"Rewriting\" history actually means creating a new history (adding\n> objects), and then changing a ref (most often a branch head) to\n> reference the new instead of the old history.\n\nI stand corrected. All objects in Git repository are actually immutable,\nbut because references can be changed (and tools like git-rebase change\nit automatically), it _appears_ like editing existing commits, but in\nfact old commits do not disappear immediately. Even if there is no other\nbranches or tags that refer to old commits, git-reflog stores references\nto them for 30 days after that the garbage collector can remove them.\n\n\nDmitry\n"},{"id":"127976","messageId":"7vr5rt90d3.fsf@alter.siamese.dyndns.org","threadId":"21659","inReplyTo":"200911200149.19528.jnareb@gmail.com","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-11-20T06:27:52Z","receivedAt":"2009-11-20T06:27:52Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> If you didn't find sufficient description of underlying concepts behind\n> git in \"Git User's Manual\" (distributed with Git), \"Git Community Book\"\n> or \"Pro Git\", take a look at the following documents:\n>\n>  * \"Git for Computer Scientists\"\n>  * \"Git From Bottom's Up\"\n>  * \"The Git Parable\"\n> ...\n> It is documented, see referenced mentioned above.\n\nI actually would want ourselves step back a bit and make sure that anybody\nwho is completely new to git won't get confused with the concepts after\ns/he reads our \"Git User's Manual\" and nothing else.  Listing five or six\ndocuments and \"you'll find information somewhere among these\" *might* be\nthe best thing we could do at this very second, but we should strive to do\nbetter than that.\n"},{"id":"127977","messageId":"7vmy2h904e.fsf@alter.siamese.dyndns.org","threadId":"21659","inReplyTo":"20091120013545.GA22556@dpotapov.dyndns.org","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-11-20T06:33:05Z","receivedAt":"2009-11-20T06:33:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Dmitry Potapov <dpotapov@gmail.com> writes:\n\n> It is more difficult to make this mistake with Git than many others\n> VCSes, because Git shows the list of files that are changed but not\n> committed as well as the list of untracked files when you try to commit\n> something.\n\nNot really in practice.  Too many people carry their existing practice of\nusing -m to write a useless single liner commit log message that they\nacquired while using their previous SCM.  Arguably, useless log messages\nare less of a problem on systems like CVS/SVN because they do not do\nuseful log summarization such as \"log -- paths...\" or \"shortlog\", so they\ncan be excused for learning the practice in the first place, though.\n\nThat incidentally is exactly why earlier we (mostly me and Linus)\nrecommended people not to teach \"commit -m\" to new people, but of course\nnobody listened ;-).\n"},{"id":"128021","messageId":"20091120150747.GF22556@dpotapov.dyndns.org","threadId":"21659","inReplyTo":"7vmy2h904e.fsf@alter.siamese.dyndns.org","subject":"Re: Hey - A Conceptual Simplication....","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2009-11-20T15:07:47Z","receivedAt":"2009-11-20T15:07:47Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Thu, Nov 19, 2009 at 10:33:05PM -0800, Junio C Hamano wrote:\n> Dmitry Potapov <dpotapov@gmail.com> writes:\n> \n> > It is more difficult to make this mistake with Git than many others\n> > VCSes, because Git shows the list of files that are changed but not\n> > committed as well as the list of untracked files when you try to commit\n> > something.\n> \n> Not really in practice.  Too many people carry their existing practice of\n> using -m to write a useless single liner commit log message that they\n> acquired while using their previous SCM.\n\nWell, at least, Git allows to avoid this mistake and produce good commit\nmessages, but you are right it is difficult to break old bad habits...\n\n> Arguably, useless log messages\n> are less of a problem on systems like CVS/SVN because they do not do\n> useful log summarization such as \"log -- paths...\" or \"shortlog\", so they\n> can be excused for learning the practice in the first place, though.\n\nI think quite often commits in CVS/SVN cannot be summarized, because a\nsingle commit often contains what would be a short series of patches in\nGit plus a few separated fix-ups that are completely unrelated to the\nwhole series. It is trivial to split your changes in a few separate\ncommits in Git, but it is difficult to do that with CVS/SVN.\n\n> That incidentally is exactly why earlier we (mostly me and Linus)\n> recommended people not to teach \"commit -m\" to new people, but of course\n> nobody listened ;-).\n\nThose who got used to '-m' in another VCS will quickly find it on their\nown... BTW, Git User's Manual uses \"git commit -m\" 8 times in different\nexamples, largely to explain what is committed here, and I think it is\nsimilar with other introductions to Git. Though, clearly '-m' is rarely\nuseful in practice...\n\n\nDmitry\n"}]}