{"thread":{"id":"12266","subject":"Question about your git habits","startedAt":"2008-02-23T00:37:14Z","lastAt":"2008-02-23T19:28:58Z","messageCount":29,"participants":["Chase Venters","Tommy Thorn","Steven Walter","Jan Engelhardt","Junio C Hamano","Al Viro","J.C. Pizarro","Daniel Barkalow","Rene Herman","Jeff Garzik","Willy Tarreau","Sam Ravnborg","Mike Hommey","Samuel Tardieu","Charles Bailey","Jakub Narebski"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"69632","messageId":"200802221837.37680.chase.venters@clientec.com","threadId":"12266","inReplyTo":null,"subject":"Question about your git habits","fromName":"Chase Venters","fromEmail":"chase.venters@clientec.com","sentAt":"2008-02-23T00:37:14Z","receivedAt":"2008-02-23T00:37:14Z","isPatch":false,"sender":{"key":"chase.venters@clientec.com","avatar":null},"body":"I've been making myself more familiar with git lately and I'm curious what \nhabits others have adopted. (I know there are a few documents in circulation \nthat deal with using git to work on the kernel but I don't think this has \nbeen specifically covered).\n\nMy question is: If you're working on multiple things at once, do you tend to \nclone the entire repository repeatedly into a series of separate working \ndirectories and do your work there, then pull that work (possibly comprising \na series of \"temporary\" commits) back into a separate local master \nrespository with --squash, either into \"master\" or into a branch containing \nthe new feature?\n\nOr perhaps you create a temporary topical branch for each thing you are \nworking on, and commit arbitrary changes then checkout another branch when \nyou need to change gears, finally --squashing the intermediate commits when a \nparticular piece of work is done?\n\nI'm using git to manage my project and I'm trying to determine the most \noptimal workflow I can. I figure that I'm going to have an \"official\" master \nrepository for the project, and I want to keep the revision history clean in \nthat repository (ie, no messy intermediate commits that don't compile or only \nimplement a feature half way).\n\nOn older projects I was using a certalized revision control system like \n*cough* Subversion *cough* and I'd create separate branches which I'd check \nout into their own working trees.\n\nIt seems to me that having multiple working trees (effectively, cloning \nthe \"master\" repository every time I need to make anything but a trivial \nchange) would be most effective under git as well as it doesn't require \ncreating messy, intermediate commits in the first place (but allows for them \nif they are used). But I wonder how that approach would scale with a project \nwhose git repo weighed hundreds of megs or more. (With a centralized rcs, of \ncourse, you don't have to lug around a copy of the whole project history in \neach working tree.)\n\nInsight appreciated, and I apologize if I've failed to RTFM somewhere.\n\nThanks,\nChase\n"},{"id":"69634","messageId":"47BF765F.9010706@thorn.ws","threadId":"12266","inReplyTo":"200802221837.37680.chase.venters@clientec.com","subject":"Re: Question about your git habits","fromName":"Tommy Thorn","fromEmail":"tommy-git@thorn.ws","sentAt":"2008-02-23T01:26:55Z","receivedAt":"2008-02-23T01:26:55Z","isPatch":false,"sender":{"key":"tommy-git@thorn.ws","avatar":null},"body":"Chase Venters wrote:\n> My question is: If you're working on multiple things at once, do you tend to \n> clone the entire repository repeatedly into a series of separate working \n> directories and do your work there, then pull that work (possibly comprising \n> a series of \"temporary\" commits) back into a separate local master \n> respository with --squash, either into \"master\" or into a branch containing \n> the new feature?\n>   \n\nIMO, that approach scales poorly and involves a lot of overhead.\n\n> Or perhaps you create a temporary topical branch for each thing you are \n> working on, and commit arbitrary changes then checkout another branch when \n> you need to change gears, finally --squashing the intermediate commits when a \n> particular piece of work is done?\n>   \n\nSpot on.\n\n\nDistribution prune for relevance.\n\nTommy\n"},{"id":"69635","messageId":"20080223012818.GA27745@dervierte","threadId":"12266","inReplyTo":"200802221837.37680.chase.venters@clientec.com","subject":"Re: Question about your git habits","fromName":"Steven Walter","fromEmail":"stevenrwalter@gmail.com","sentAt":"2008-02-23T01:28:18Z","receivedAt":"2008-02-23T01:28:18Z","isPatch":false,"sender":{"key":"stevenrwalter@gmail.com","avatar":"https://avatars.githubusercontent.com/u/79127?v=4"},"body":"On Fri, Feb 22, 2008 at 06:37:14PM -0600, Chase Venters wrote:\n> My question is: If you're working on multiple things at once, do you tend to \n> clone the entire repository repeatedly into a series of separate working \n> directories and do your work there, then pull that work (possibly comprising \n> a series of \"temporary\" commits) back into a separate local master \n> respository with --squash, either into \"master\" or into a branch containing \n> the new feature?\n> \n> Or perhaps you create a temporary topical branch for each thing you are \n> working on, and commit arbitrary changes then checkout another branch when \n> you need to change gears, finally --squashing the intermediate commits when a \n> particular piece of work is done?\n\nI favor the second approach: single working copy, multiple branches.  My\nfeeling is that wanting multiple workspaces is a holdover from using\nsubversion.  For me, it is much faster to \"git commit -a -m wip\"\nand then switch branches, than it would be to clone a whole new\nrepository and manage the inter-repository relationships.\n\nDon't get so down on the \"intermediate commits,\" either.  For one,\nwhenever I switch back to a branch with a \"wip\" commit, I usually do a\n\"git reset HEAD^\" to remove it and get my working tree back where it\nwas.  There are also nifty tools like interactive rebase that assist\nyou in rewriting history to produce a set of clean, atomic commits.\nIt's not imperative to make your first draft perfection in git.\n\n[...]\n\n> Insight appreciated, and I apologize if I've failed to RTFM somewhere.\n\nNo worries, I remember being in your situation once.  git opens up\na host of opportunities with its flexibility, and getting started I\nwas consistently stumped by which of the many paths I should choose.\n-- \n-Steven Walter <stevenrwalter@gmail.com>\nFreedom is the freedom to say that 2 + 2 = 4\nB2F1 0ECC E605 7321 E818  7A65 FC81 9777 DC28 9E8F \n"},{"id":"69636","messageId":"Pine.LNX.4.64.0802230221140.21077@fbirervta.pbzchgretzou.qr","threadId":"12266","inReplyTo":"200802221837.37680.chase.venters@clientec.com","subject":"Re: Question about your git habits","fromName":"Jan Engelhardt","fromEmail":"jengelh@computergmbh.de","sentAt":"2008-02-23T01:37:00Z","receivedAt":"2008-02-23T01:37:00Z","isPatch":false,"sender":{"key":"jengelh@computergmbh.de","avatar":null},"body":"\nOn Feb 22 2008 18:37, Chase Venters wrote:\n>\n>I've been making myself more familiar with git lately and I'm curious what \n>habits others have adopted. (I know there are a few documents in circulation \n>that deal with using git to work on the kernel but I don't think this has \n>been specifically covered).\n>\n>My question is: If you're working on multiple things at once,\n\nImpossible; Humans only have one core with only seven registers --\naccording to CodingStyle chapter 6 paragraph 4.\n\n>do you tend to clone the entire repository repeatedly into a series\n>of separate working directories\n\nToo time consuming on consumer drives with projects the size of Linux.\n\n>and do your work there, then pull\n>that work (possibly comprising a series of \"temporary\" commits) back\n>into a separate local master respository with --squash, either into\n>\"master\" or into a branch containing the new feature?\n\nNo, just commit the current unfinished work to a new branch and deal\nwith it later (cherry-pick, rebase, reset --soft, commit --amend -i,\nyou name it). Or if all else fails, use git-stash.\n\nYou do not have to push these temporary branches at all, so it is\nmuch nicer than svn. (Once all the work is done and cleanly in\nmaster, you can kill off all branches without having a record\nof their previous existence.)\n\n>Or perhaps you create a temporary topical branch for each thing you\n>are working on, and commit arbitrary changes then checkout another\n>branch when you need to change gears, finally --squashing the\n>intermediate commits when a particular piece of work is done?\n\nif I don't collect arbitrary changes, I don't need squashing\n(see reset --soft/amend above)\n"},{"id":"69638","messageId":"7vk5kw4fep.fsf@gitster.siamese.dyndns.org","threadId":"12266","inReplyTo":"200802221837.37680.chase.venters@clientec.com","subject":"Re: Question about your git habits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-02-23T01:42:22Z","receivedAt":"2008-02-23T01:42:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Chase Venters <chase.venters@clientec.com> writes:\n\n[jc: kernel-list removed from CC: as this does not have anything\nto do with them]\n\n> My question is: If you're working on multiple things at once,\n> do you tend to clone the entire repository repeatedly into a\n> series of separate working directories and do your work there,\n> then pull that work (possibly comprising a series of\n> \"temporary\" commits) back into a separate local master\n> respository with --squash, either into \"master\" or into a\n> branch containing the new feature?\n>\n> Or perhaps you create a temporary topical branch for each\n> thing you are working on, and commit arbitrary changes then\n> checkout another branch when you need to change gears, finally\n> --squashing the intermediate commits when a particular piece\n> of work is done?\n\nIt is a matter of taste, but in any case, you should not have to\nsquash that often.  If you find you are always squashing because\nyou work on one thing and then switch to another thing before\nyou are done with the former, something is wrong.\n\n\tClarification: I am not saying squashing is wrong.  I am\n\tjust saying you should not have to.\n\nIf you want to park what you were working on before switching to\ndo something else, you can (and probably should) commit and it\nis a very valid thing to do (an alternative is \"git stash\").\n\nWhen resuming, if that parked commit was half-baked and\nsomething you do not want to go back to later, then the next\ncommit (be it another commit that merely \"parks\" before getting\ndistracted to do something else, or a commit that finally gets\neverything \"finito\") can be made with \"commit --amend\".  That\nway, your sequences of commits will consist of only logically\nseparate units, without half-baked ones you had to create only\nbecause you switched branches.\n\nSome people prefer to use multiple simultanous work trees.  You\ncertainly can use \"clone\" to achieve this.  And local clone is\nvery cheap as it shares the object database from the origin by\ndefault.\n\nMany people prefer to use topic branches, and working in a\nsingle repository with multiple branches and switching branches\nwithout ever cd'ing around is certainly a possible and very\nvalid way to work.  As long as your build infrastructure is sane\n(e.g. your project does not have a central header file that any\nlittle subsystem change needs to modify and included by\neverybody, which tends to screw up make quite badly), switching\nbranches would not incur too much recompilation either and it\nobviously will save disk space not having to leave multiple\ncheckout around.\n\nYou can also work with a single repository, multiple branches\nand have multiple simultaneous work trees attached to that\nsingle repository, by using contrib/workdir/git-new-workdir\nscript.\n"},{"id":"69639","messageId":"20080223014445.GK27894@ZenIV.linux.org.uk","threadId":"12266","inReplyTo":"Pine.LNX.4.64.0802230221140.21077@fbirervta.pbzchgretzou.qr","subject":"Re: Question about your git habits","fromName":"Al Viro","fromEmail":"viro@zeniv.linux.org.uk","sentAt":"2008-02-23T01:44:45Z","receivedAt":"2008-02-23T01:44:45Z","isPatch":false,"sender":{"key":"viro@zeniv.linux.org.uk","avatar":null},"body":"On Sat, Feb 23, 2008 at 02:37:00AM +0100, Jan Engelhardt wrote:\n\n> >do you tend to clone the entire repository repeatedly into a series\n> >of separate working directories\n> \n> Too time consuming on consumer drives with projects the size of Linux.\n\ngit clone -l -s\n\nis not particulary slow...\n"},{"id":"69640","messageId":"7vfxvk4f07.fsf@gitster.siamese.dyndns.org","threadId":"12266","inReplyTo":"20080223014445.GK27894@ZenIV.linux.org.uk","subject":"Re: Question about your git habits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-02-23T01:51:04Z","receivedAt":"2008-02-23T01:51:04Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Al Viro <viro@ZenIV.linux.org.uk> writes:\n\n> On Sat, Feb 23, 2008 at 02:37:00AM +0100, Jan Engelhardt wrote:\n>\n>> >do you tend to clone the entire repository repeatedly into a series\n>> >of separate working directories\n>> \n>> Too time consuming on consumer drives with projects the size of Linux.\n>\n> git clone -l -s\n>\n> is not particulary slow...\n\nHow big is a checkout of a single revision of kernel these days,\ncompared to a well-packed history since v2.6.12-rc2?\n\nThe cost of writing out the work tree files isn't ignorable and\nprobably more than writing out the repository data (which -s\nsaves for you).\n"},{"id":"69641","messageId":"20080223020913.GL27894@ZenIV.linux.org.uk","threadId":"12266","inReplyTo":"7vfxvk4f07.fsf@gitster.siamese.dyndns.org","subject":"Re: Question about your git habits","fromName":"Al Viro","fromEmail":"viro@zeniv.linux.org.uk","sentAt":"2008-02-23T02:09:13Z","receivedAt":"2008-02-23T02:09:13Z","isPatch":false,"sender":{"key":"viro@zeniv.linux.org.uk","avatar":null},"body":"On Fri, Feb 22, 2008 at 05:51:04PM -0800, Junio C Hamano wrote:\n> Al Viro <viro@ZenIV.linux.org.uk> writes:\n> \n> > On Sat, Feb 23, 2008 at 02:37:00AM +0100, Jan Engelhardt wrote:\n> >\n> >> >do you tend to clone the entire repository repeatedly into a series\n> >> >of separate working directories\n> >> \n> >> Too time consuming on consumer drives with projects the size of Linux.\n> >\n> > git clone -l -s\n> >\n> > is not particulary slow...\n> \n> How big is a checkout of a single revision of kernel these days,\n> compared to a well-packed history since v2.6.12-rc2?\n> \n> The cost of writing out the work tree files isn't ignorable and\n> probably more than writing out the repository data (which -s\n> saves for you).\n\nDepends...  I'm using ext2 for that and noatime everywhere, so that might\nchange the picture, but IME it's fast enough...  As for the size, it gets\nto ~320Mb on disk, which is comparable to the pack size (~240-odd Mb).\n"},{"id":"69647","messageId":"998d0e4a0802221846u795160b2r3acb0839ced74d29@mail.gmail.com","threadId":"12266","inReplyTo":"998d0e4a0802221736q4e4c3a28l101522912f7d3caf@mail.gmail.com","subject":"Re: Question about your git habits","fromName":"J.C. Pizarro","fromEmail":"jcpiza@gmail.com","sentAt":"2008-02-23T02:46:19Z","receivedAt":"2008-02-23T02:46:19Z","isPatch":false,"sender":{"key":"jcpiza@gmail.com","avatar":null},"body":"2008/2/23, Chase Venters <chase.venters@clientec.com> wrote:\n >\n > ... blablabla\n >\n >  My question is: If you're working on multiple things at once, do you tend to\n >  clone the entire repository repeatedly into a series of separate working\n >  directories and do your work there, then pull that work (possibly comprising\n >  a series of \"temporary\" commits) back into a separate local master\n >  respository with --squash, either into \"master\" or into a branch containing\n >  the new feature?\n >\n > ... blablabla\n >\n >  I'm using git to manage my project and I'm trying to determine the most\n >  optimal workflow I can. I figure that I'm going to have an \"official\" master\n >  repository for the project, and I want to keep the revision history clean in\n >  that repository (ie, no messy intermediate commits that don't\ncompile or only\n >  implement a feature half way).\n\n\nI recomend you to use these complementary tools\n\n   1. google: gitk screenshots  ( e.g. http://lwn.net/Articles/140350/ )\n\n   2. google: \"git-gui\" screenshots\n         ( e.g. http://www.spearce.org/2007/01/git-gui-screenshots.html )\n\n   3. google: gitweb color meld\n\n   ;)\n"},{"id":"69648","messageId":"998d0e4a0802221847m431aa136xa217333b0517b962@mail.gmail.com","threadId":"12266","inReplyTo":"998d0e4a0802221823h3ba53097gf64fcc2ea826302b@mail.gmail.com","subject":"Re: Question about your git habits","fromName":"J.C. Pizarro","fromEmail":"jcpiza@gmail.com","sentAt":"2008-02-23T02:47:07Z","receivedAt":"2008-02-23T02:47:07Z","isPatch":false,"sender":{"key":"jcpiza@gmail.com","avatar":null},"body":"On 2008/2/23, Al Viro <viro@zeniv.linux.org.uk> wrote:\n > On Fri, Feb 22, 2008 at 05:51:04PM -0800, Junio C Hamano wrote:\n >  > Al Viro <viro@ZenIV.linux.org.uk> writes:\n >  >\n >  > > On Sat, Feb 23, 2008 at 02:37:00AM +0100, Jan Engelhardt wrote:\n >  > >\n >  > >> >do you tend to clone the entire repository repeatedly into a series\n >  > >> >of separate working directories\n >  > >>\n >  > >> Too time consuming on consumer drives with projects the size of Linux.\n >  > >\n >  > > git clone -l -s\n >  > >\n >  > > is not particulary slow...\n >  >\n >  > How big is a checkout of a single revision of kernel these days,\n >  > compared to a well-packed history since v2.6.12-rc2?\n >  >\n >  > The cost of writing out the work tree files isn't ignorable and\n >  > probably more than writing out the repository data (which -s\n >  > saves for you).\n >\n >\n > Depends...  I'm using ext2 for that and noatime everywhere, so that might\n >  change the picture, but IME it's fast enough...  As for the size, it gets\n >  to ~320Mb on disk, which is comparable to the pack size (~240-odd Mb).\n\n\nYesterday, i had git cloned git://foo.com/bar.git   ( 777 MiB )\n Today, i've git cloned git://foo.com/bar.git   ( 779 MiB )\n\n Both repos are different binaries , and i used 777 MiB + 779 MiB = 1556 MiB\n of bandwidth in two days. It's much!\n\n Why don't we implement \"binary delta between old git repo and recent git repo\"\n with \"SHA1 built git repo verifier\"?\n\n Suppose the size cost of this binary delta is e.g. around 52 MiB instead of\n 2 MiB due to numerous mismatching of binary parts, then the bandwidth\n in two days will be 777 MiB + 52 MiB = 829 MiB instead of 1556 MiB.\n\n Unfortunately, this \"binary delta of repos\" is not implemented yet :|\n"},{"id":"69650","messageId":"alpine.LNX.1.00.0802222249480.19024@iabervon.org","threadId":"12266","inReplyTo":"200802221837.37680.chase.venters@clientec.com","subject":"Re: Question about your git habits","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2008-02-23T04:10:48Z","receivedAt":"2008-02-23T04:10:48Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 22 Feb 2008, Chase Venters wrote:\n\n> I've been making myself more familiar with git lately and I'm curious what \n> habits others have adopted. (I know there are a few documents in circulation \n> that deal with using git to work on the kernel but I don't think this has \n> been specifically covered).\n> \n> My question is: If you're working on multiple things at once, do you tend to \n> clone the entire repository repeatedly into a series of separate working \n> directories and do your work there, then pull that work (possibly comprising \n> a series of \"temporary\" commits) back into a separate local master \n> respository with --squash, either into \"master\" or into a branch containing \n> the new feature?\n> \n> Or perhaps you create a temporary topical branch for each thing you are \n> working on, and commit arbitrary changes then checkout another branch when \n> you need to change gears, finally --squashing the intermediate commits when a \n> particular piece of work is done?\n\nI find that the sequence of changes I make is pretty much unrelated to the \nsequence of changes that end up in the project's history, because my \nchanges as I make them involve writing a lot of stubs (so I can build) and \nthen filling them out. It's beneficial to have version control on this so \nthat, if I screw up filling out a stub, I can get back to where I was.\n\nHaving made a complete series, I then generate a new series of commits, \neach of which does one thing, without any bugs that I've resolved, such \nthat the net result is the end of the messy history, except with any \ndebugging or useless stuff skipped. It's this series that gets merged into \nthe project history, and I discard the other history.\n\nThe real trick is that the early patches in a lot of series often refactor \nexisting code in ways that are generally good and necessary for your \neventual outcome, but which you'd never think of until you've written more \nof the series. Generating a new commit sequence is necessary to end up \nwith a history where it looks from the start like you know where you're \ngoing and have everything done that needs to be done when you get to the \npoint of needing it. Furthermore, you want to be able to test these \ncommits in isolation, without the distraction of the changes that actually \nprompted them, which means that you want to have your working tree is a \nstate that you never actually had it in as you were developing the end \nresult.\n\nThis means that you'll usually want to rewrite commits for any series that \nisn't a single obvious patch, so it's not a big deal to commit any time \nyou want to work on some different branch.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"69652","messageId":"47BFA37F.10806@keyaccess.nl","threadId":"12266","inReplyTo":"200802221837.37680.chase.venters@clientec.com","subject":"Re: Question about your git habits","fromName":"Rene Herman","fromEmail":"rene.herman@keyaccess.nl","sentAt":"2008-02-23T04:39:27Z","receivedAt":"2008-02-23T04:39:27Z","isPatch":false,"sender":{"key":"rene.herman@keyaccess.nl","avatar":null},"body":"On 23-02-08 01:37, Chase Venters wrote:\n\n> Or perhaps you create a temporary topical branch for each thing you are \n> working on, and commit arbitrary changes then checkout another branch\n> when you need to change gears, finally --squashing the intermediate\n> commits when a particular piece of work is done?\n\nNo very specific advice to give but this is what I do and then pull all \n(compilable) topic branches into a \"local\" branch for complation. Just \nwanted to remark that a definite downside is that switching branches a lot \nalso touches the tree a lot and hence tends to trigger quite unwelcome \namounts of recompiles. Using ccache would proably be effective in this \nsituation but I keep neglecting to check it out...\n\nRene\n"},{"id":"69657","messageId":"47BFA938.3050504@garzik.org","threadId":"12266","inReplyTo":"alpine.LNX.1.00.0802222249480.19024@iabervon.org","subject":"Re: Question about your git habits","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2008-02-23T05:03:52Z","receivedAt":"2008-02-23T05:03:52Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Daniel Barkalow wrote:\n> I find that the sequence of changes I make is pretty much unrelated to the \n> sequence of changes that end up in the project's history, because my \n> changes as I make them involve writing a lot of stubs (so I can build) and \n> then filling them out. It's beneficial to have version control on this so \n> that, if I screw up filling out a stub, I can get back to where I was.\n> \n> Having made a complete series, I then generate a new series of commits, \n> each of which does one thing, without any bugs that I've resolved, such \n> that the net result is the end of the messy history, except with any \n> debugging or useless stuff skipped. It's this series that gets merged into \n> the project history, and I discard the other history.\n> \n> The real trick is that the early patches in a lot of series often refactor \n> existing code in ways that are generally good and necessary for your \n> eventual outcome, but which you'd never think of until you've written more \n> of the series.\n\nThat summarizes well how I do original development, too.  Whether its a \nbranch of an existing repo, or a newly cloned repo, when working on new \ncode I will do a first pass, committing as I go to provide useful \ncheckpoints.\n\nOnce I reach a satisfactory state, I'll refactor the patches so that \nthey make sense for upstream submission.\n\n\tJeff\n"},{"id":"69663","messageId":"20080223085634.GW8953@1wt.eu","threadId":"12266","inReplyTo":"200802221837.37680.chase.venters@clientec.com","subject":"Re: Question about your git habits","fromName":"Willy Tarreau","fromEmail":"w@1wt.eu","sentAt":"2008-02-23T08:56:34Z","receivedAt":"2008-02-23T08:56:34Z","isPatch":false,"sender":{"key":"w@1wt.eu","avatar":"https://avatars.githubusercontent.com/u/8141789?v=4"},"body":"On Fri, Feb 22, 2008 at 06:37:14PM -0600, Chase Venters wrote:\n> It seems to me that having multiple working trees (effectively, cloning \n> the \"master\" repository every time I need to make anything but a trivial \n> change) would be most effective under git as well as it doesn't require \n> creating messy, intermediate commits in the first place (but allows for them \n> if they are used). But I wonder how that approach would scale with a project \n> whose git repo weighed hundreds of megs or more. (With a centralized rcs, of \n> course, you don't have to lug around a copy of the whole project history in \n> each working tree.)\n\nTake a look at git-new-workdir in git's contrib directory. I'm using it a\nlot now. It makes it possible to set up as many workdirs as you want, sharing\nthe same repo. It's very dangerous if you're not rigorous, but it saves a lot\nof time when you work on several branches at a time, which is even more true\nfor a project's documentation. The real thing to care about is not to have\nthe same branch checked out at several places.\n\nRegards,\nWilly\n"},{"id":"69665","messageId":"20080223091013.GB12161@uranus.ravnborg.org","threadId":"12266","inReplyTo":"200802221837.37680.chase.venters@clientec.com","subject":"Re: Question about your git habits","fromName":"Sam Ravnborg","fromEmail":"sam@ravnborg.org","sentAt":"2008-02-23T09:10:13Z","receivedAt":"2008-02-23T09:10:13Z","isPatch":false,"sender":{"key":"sam@ravnborg.org","avatar":"https://gravatar.com/avatar/168a912606ed0742d840bb365e3cc21db390c36531a58341dc7a069cc1f15f62?d=mp&s=160"},"body":"On Fri, Feb 22, 2008 at 06:37:14PM -0600, Chase Venters wrote:\n> I've been making myself more familiar with git lately and I'm curious what \n> habits others have adopted. (I know there are a few documents in circulation \n> that deal with using git to work on the kernel but I don't think this has \n> been specifically covered).\n> \n> My question is: If you're working on multiple things at once, do you tend to \n> clone the entire repository repeatedly into a series of separate working \n> directories and do your work there, then pull that work (possibly comprising \n> a series of \"temporary\" commits) back into a separate local master \n> respository with --squash, either into \"master\" or into a branch containing \n> the new feature?\n\nThe simple (for me) workflow I use is to create a clone of the\nkernel for each 'topic' I work on.\nSo at the same time I may have one or maybe up to five clones of the\nkernel.\n\nWhen I want to combine thing I use git format-patch and git am.\nOften there is some amount of editing done before combining stuff\nespecially for larger changes where the first in the serie is often\npreparational work that were identified in random order when I did\nthe inital work.\n\n\tSam\n"},{"id":"69666","messageId":"20080223091855.GA18942@glandium.org","threadId":"12266","inReplyTo":"alpine.LNX.1.00.0802222249480.19024@iabervon.org","subject":"Re: Question about your git habits","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-02-23T09:18:55Z","receivedAt":"2008-02-23T09:18:55Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Fri, Feb 22, 2008 at 11:10:48PM -0500, Daniel Barkalow wrote:\n> I find that the sequence of changes I make is pretty much unrelated to the \n> sequence of changes that end up in the project's history, because my \n> changes as I make them involve writing a lot of stubs (so I can build) and \n> then filling them out. It's beneficial to have version control on this so \n> that, if I screw up filling out a stub, I can get back to where I was.\n> \n> Having made a complete series, I then generate a new series of commits, \n> each of which does one thing, without any bugs that I've resolved, such \n> that the net result is the end of the messy history, except with any \n> debugging or useless stuff skipped. It's this series that gets merged into \n> the project history, and I discard the other history.\n> \n> The real trick is that the early patches in a lot of series often refactor \n> existing code in ways that are generally good and necessary for your \n> eventual outcome, but which you'd never think of until you've written more \n> of the series. Generating a new commit sequence is necessary to end up \n> with a history where it looks from the start like you know where you're \n> going and have everything done that needs to be done when you get to the \n> point of needing it. Furthermore, you want to be able to test these \n> commits in isolation, without the distraction of the changes that actually \n> prompted them, which means that you want to have your working tree is a \n> state that you never actually had it in as you were developing the end \n> result.\n> \n> This means that you'll usually want to rewrite commits for any series that \n> isn't a single obvious patch, so it's not a big deal to commit any time \n> you want to work on some different branch.\n\nI do that so much that I have this alias:\n        reorder = !sh -c 'git rebase -i --onto $0 $0 $1'\n\n... and actually pass it only one argument most of the time.\n\nMike\n"},{"id":"69671","messageId":"2008-02-23-11-39-23+trackit+sam@rfc1149.net","threadId":"12266","inReplyTo":"7vk5kw4fep.fsf@gitster.siamese.dyndns.org","subject":"Re: Question about your git habits","fromName":"Samuel Tardieu","fromEmail":"sam@rfc1149.net","sentAt":"2008-02-23T10:39:23Z","receivedAt":"2008-02-23T10:39:23Z","isPatch":false,"sender":{"key":"sam@rfc1149.net","avatar":"https://avatars.githubusercontent.com/u/44656?v=4"},"body":">>>>> \"Junio\" == Junio C Hamano <gitster@pobox.com> writes:\n\nJunio> Many people prefer to use topic branches, and working in a\nJunio> single repository with multiple branches and switching branches\nJunio> without ever cd'ing around is certainly a possible and very\nJunio> valid way to work.  As long as your build infrastructure is\nJunio> sane (e.g. your project does not have a central header file\nJunio> that any little subsystem change needs to modify and included\nJunio> by everybody, which tends to screw up make quite badly),\nJunio> switching branches would not incur too much recompilation\nJunio> either and it obviously will save disk space not having to\nJunio> leave multiple checkout around.\n\nAnd even in this case (central header file), ccache will greatly\ndecrease compilation time in the case of a C/C++ project.\n\n  Sam\n-- \nSamuel Tardieu -- sam@rfc1149.net -- http://www.rfc1149.net/\n"},{"id":"69673","messageId":"20080223113952.GA4936@hashpling.org","threadId":"12266","inReplyTo":"998d0e4a0802221847m431aa136xa217333b0517b962@mail.gmail.com","subject":"Re: Question about your git habits","fromName":"Charles Bailey","fromEmail":"charles@hashpling.org","sentAt":"2008-02-23T11:39:52Z","receivedAt":"2008-02-23T11:39:52Z","isPatch":false,"sender":{"key":"charles@hashpling.org","avatar":"https://avatars.githubusercontent.com/u/1668475?v=4"},"body":"On Sat, Feb 23, 2008 at 03:47:07AM +0100, J.C. Pizarro wrote:\n> \n> Yesterday, i had git cloned git://foo.com/bar.git   ( 777 MiB )\n>  Today, i've git cloned git://foo.com/bar.git   ( 779 MiB )\n> \n>  Both repos are different binaries , and i used 777 MiB + 779 MiB = 1556 MiB\n>  of bandwidth in two days. It's much!\n> \n>  Why don't we implement \"binary delta between old git repo and recent git repo\"\n>  with \"SHA1 built git repo verifier\"?\n> \n>  Suppose the size cost of this binary delta is e.g. around 52 MiB instead of\n>  2 MiB due to numerous mismatching of binary parts, then the bandwidth\n>  in two days will be 777 MiB + 52 MiB = 829 MiB instead of 1556 MiB.\n> \n>  Unfortunately, this \"binary delta of repos\" is not implemented yet :|\n\nIt sounds like what concerns you is the bandwith to git://foo.bar. If\nyou are cloning the first repository to somewhere were the first\nclone is accessible and bandwidth between the clones is not an issue,\nthen you should be able to use the --reference parameter to git clone\nto just fetch the missing ~2 MiB from foo.bar.\n\nA \"binary delta of repos\" should just be an 'incremental' pack file\nand the git protocol should support generating an appropriate one. I'm\nnot quite sure what \"not implemented yet\" feature you are looking for.\n"},{"id":"69676","messageId":"m3ablrddna.fsf@localhost.localdomain","threadId":"12266","inReplyTo":"200802221837.37680.chase.venters@clientec.com","subject":"Re: Question about your git habits","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-23T13:07:58Z","receivedAt":"2008-02-23T13:07:58Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[removed linux-kernel list from Cc]\n\nChase Venters <chase.venters@clientec.com> writes:\n\n> My question is: If you're working on multiple things at once, do you\n> tend to clone the entire repository repeatedly into a series of\n> separate working directories and do your work there, then pull that\n> work (possibly comprising a series of \"temporary\" commits) back into\n> a separate local master respository with --squash, either into\n> \"master\" or into a branch containing the new feature?\n\nAlternate solution is to use multiple working trees (multiple working\ndirectories) with single repository, although it is still a bit\nfragile; you should take care to not checkout same branch multiple\ntimes. IIRC when discussing \".git\" as a file representing symlink,\nthere were some discussion on how to improve multiple-workspaces\nworkflow.\n \n> Or perhaps you create a temporary topical branch for each thing you\n> are working on, and commit arbitrary changes then checkout another\n> branch when you need to change gears, finally --squashing the\n> intermediate commits when a particular piece of work is done?\n\nI personally prefer this workflow, but I do not work as a main\ncontributor nor maintainer of large project.\n\nAs to intermediate commits: if you feel the need to interrupt your\nwork which is not quite ready for final commit, you can either use\n\"git stash\" command, or commit it as WIP commit, then when going back\njust \"git commit --amend\" it.\n\nMoreover, when working on some larger topic, which needs to be split\ninto individual commits for beter history clarity, and for better\nbisectability, you usually rewrite history before submitting\n(publishing) your changes. You usually have to reorder commits (for\nexample moving improvements to infrastructure before commits\nintroducing new feature), split commits (separating just noticed\nbugfix from a feature commit), squash commits (joining feature commit\nand its bugfix) etc. You can use \"git rebase --interactive\" for that,\nor one of Quilt-like patch management interfaces for git: StGit (which\nI personally use) or Guilt (idea based on mq: Mercurial queues\nextension).\n\n[...]\n\n> It seems to me that having multiple working trees (effectively, cloning \n> the \"master\" repository every time I need to make anything but a trivial \n> change) would be most effective under git as well as it doesn't require \n> creating messy, intermediate commits in the first place (but allows for them \n> if they are used). But I wonder how that approach would scale with a project \n> whose git repo weighed hundreds of megs or more. (With a centralized rcs, of \n> course, you don't have to lug around a copy of the whole project history in \n> each working tree.)\n\nYou can always clone using --shared option to set-up alternates; this\nway only new objects (new commits) would be stored in the clone. This\nof course need for clone and source to be on the same filesystem.\n\nBy default git-clone on local filesystem uses hardlinks, so it also\nshould not be so hard on disk space.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"69677","messageId":"998d0e4a0802230508w12f236baiaf2d9ab5f364670a@mail.gmail.com","threadId":"12266","inReplyTo":"20080223113952.GA4936@hashpling.org","subject":"Re: Question about your git habits","fromName":"J.C. Pizarro","fromEmail":"jcpiza@gmail.com","sentAt":"2008-02-23T13:08:35Z","receivedAt":"2008-02-23T13:08:35Z","isPatch":false,"sender":{"key":"jcpiza@gmail.com","avatar":null},"body":"On 2008/2/23, Charles Bailey <charles@hashpling.org> wrote:\n> On Sat, Feb 23, 2008 at 03:47:07AM +0100, J.C. Pizarro wrote:\n>  >\n>  > Yesterday, i had git cloned git://foo.com/bar.git   ( 777 MiB )\n>  >  Today, i've git cloned git://foo.com/bar.git   ( 779 MiB )\n>  >\n>  >  Both repos are different binaries , and i used 777 MiB + 779 MiB = 1556 MiB\n>  >  of bandwidth in two days. It's much!\n>  >\n>  >  Why don't we implement \"binary delta between old git repo and recent git repo\"\n>  >  with \"SHA1 built git repo verifier\"?\n>  >\n>  >  Suppose the size cost of this binary delta is e.g. around 52 MiB instead of\n>  >  2 MiB due to numerous mismatching of binary parts, then the bandwidth\n>  >  in two days will be 777 MiB + 52 MiB = 829 MiB instead of 1556 MiB.\n>  >\n>  >  Unfortunately, this \"binary delta of repos\" is not implemented yet :|\n>\n>\n> It sounds like what concerns you is the bandwith to git://foo.bar. If\n>  you are cloning the first repository to somewhere were the first\n>  clone is accessible and bandwidth between the clones is not an issue,\n>  then you should be able to use the --reference parameter to git clone\n>  to just fetch the missing ~2 MiB from foo.bar.\n>\n>  A \"binary delta of repos\" should just be an 'incremental' pack file\n>  and the git protocol should support generating an appropriate one. I'm\n>  not quite sure what \"not implemented yet\" feature you are looking for.\n\nBut if the repos are aggressively repacked then the bit to bit differences\nare not ~2 MiB.\n"},{"id":"69679","messageId":"20080223131749.GA5811@hashpling.org","threadId":"12266","inReplyTo":"998d0e4a0802230508w12f236baiaf2d9ab5f364670a@mail.gmail.com","subject":"Re: Question about your git habits","fromName":"Charles Bailey","fromEmail":"charles@hashpling.org","sentAt":"2008-02-23T13:17:49Z","receivedAt":"2008-02-23T13:17:49Z","isPatch":false,"sender":{"key":"charles@hashpling.org","avatar":"https://avatars.githubusercontent.com/u/1668475?v=4"},"body":"On Sat, Feb 23, 2008 at 02:08:35PM +0100, J.C. Pizarro wrote:\n> \n> But if the repos are aggressively repacked then the bit to bit differences\n> are not ~2 MiB.\n\nIt shouldn't matter how aggressively the repositories are packed or what\nthe binary differences are between the pack files are. git clone\nshould (with the --reference option) generate a new pack for you with\nonly the missing objects. If these objects are ~52 MiB then a lot has\nbeen committed to the repository, but you're not going to be able to\nget around a big download any other way.\n"},{"id":"69681","messageId":"998d0e4a0802230536w74e93ec3s40c77d52b183a419@mail.gmail.com","threadId":"12266","inReplyTo":"20080223131749.GA5811@hashpling.org","subject":"Re: Question about your git habits","fromName":"J.C. Pizarro","fromEmail":"jcpiza@gmail.com","sentAt":"2008-02-23T13:36:59Z","receivedAt":"2008-02-23T13:36:59Z","isPatch":false,"sender":{"key":"jcpiza@gmail.com","avatar":null},"body":"On 2008/2/23, Charles Bailey <charles@hashpling.org> wrote:\n> On Sat, Feb 23, 2008 at 02:08:35PM +0100, J.C. Pizarro wrote:\n>  >\n>  > But if the repos are aggressively repacked then the bit to bit differences\n>  > are not ~2 MiB.\n>\n>\n> It shouldn't matter how aggressively the repositories are packed or what\n>  the binary differences are between the pack files are. git clone\n>  should (with the --reference option) generate a new pack for you with\n>  only the missing objects. If these objects are ~52 MiB then a lot has\n>  been committed to the repository, but you're not going to be able to\n>  get around a big download any other way.\n\nYou're wrong, nothing has to be commited ~52 MiB to the repository.\n\nI'm not saying \"commit\", i'm saying\n\n\"Assume A & B binary git repos and delta_B-A another binary file, i\nrequest built\nB' = A + delta_B-A where is verified SHA1(B') = SHA1(B) for avoiding\ncorrupting\".\n\nAssume B is the higher repacked version of \"A + minor commits of the day\"\nas if B was optimizing 24 hours more the minimum spanning tree. Wow!!!\n"},{"id":"69682","messageId":"20080223140153.GB5811@hashpling.org","threadId":"12266","inReplyTo":"998d0e4a0802230536w74e93ec3s40c77d52b183a419@mail.gmail.com","subject":"Re: Question about your git habits","fromName":"Charles Bailey","fromEmail":"charles@hashpling.org","sentAt":"2008-02-23T14:01:53Z","receivedAt":"2008-02-23T14:01:53Z","isPatch":false,"sender":{"key":"charles@hashpling.org","avatar":"https://avatars.githubusercontent.com/u/1668475?v=4"},"body":"On Sat, Feb 23, 2008 at 02:36:59PM +0100, J.C. Pizarro wrote:\n> On 2008/2/23, Charles Bailey <charles@hashpling.org> wrote:\n> >\n> > It shouldn't matter how aggressively the repositories are packed or what\n> >  the binary differences are between the pack files are. git clone\n> >  should (with the --reference option) generate a new pack for you with\n> >  only the missing objects. If these objects are ~52 MiB then a lot has\n> >  been committed to the repository, but you're not going to be able to\n> >  get around a big download any other way.\n> \n> You're wrong, nothing has to be commited ~52 MiB to the repository.\n> \n> I'm not saying \"commit\", i'm saying\n> \n> \"Assume A & B binary git repos and delta_B-A another binary file, i\n> request built\n> B' = A + delta_B-A where is verified SHA1(B') = SHA1(B) for avoiding\n> corrupting\".\n> \n> Assume B is the higher repacked version of \"A + minor commits of the day\"\n> as if B was optimizing 24 hours more the minimum spanning tree. Wow!!!\n> \n\nI'm not sure that I understand where you are going with this.\nOriginally, you stated that if you clone a 775 MiB repository on day\none, and then you clone it again on day two when it was 777 MiB, then\nyou currently have to download 775 + 777 MiB of data, whereas you\ncould download a 52 MiB binary diff. I have no idea where that value\nof 52 MiB comes from, and I've no idea how many objects were committed\nbetween day one and day two. If we're going to talk about details,\nthen you need to provide more details about your scenario.\n\nHaving said that, here is my original point in some more detail. git\nrepositories are not binary blobs, they are object databases. Better\nthan this, they are databases of immutable objects. This means that to\nget the difference between one database and another, you only need to\nadd the objects that are missing from the other database. If the two\ndatabases are actually a database and the same database at short time\ninterval later, then almost all the objects are going to be common and\nthe difference will be a small set of objects. Using git:// this set\nof objects can be efficiently transfered as a pack file. You may have\na corner case scenario where the following isn't true, but in my\nexperience an incremental pack file will be a more compact\nrepresentation of this difference than a binary difference of two\naggressively repacked git repositories as generated by a generic\nbinary difference engine.\n\nI'm sorry if I've misunderstood your last point. Perhaps you could\nexpand in the exact issue that are having if I have, as I'm not sure\nthat I've really answered your last message.\n"},{"id":"69683","messageId":"20080223140820.GA4303@glandium.org","threadId":"12266","inReplyTo":"998d0e4a0802221847m431aa136xa217333b0517b962@mail.gmail.com","subject":"Re: Question about your git habits","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-02-23T14:08:20Z","receivedAt":"2008-02-23T14:08:20Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Sat, Feb 23, 2008 at 03:47:07AM +0100, J.C. Pizarro wrote:\n> On 2008/2/23, Al Viro <viro@zeniv.linux.org.uk> wrote:\n>  > On Fri, Feb 22, 2008 at 05:51:04PM -0800, Junio C Hamano wrote:\n>  >  > Al Viro <viro@ZenIV.linux.org.uk> writes:\n>  >  >\n>  >  > > On Sat, Feb 23, 2008 at 02:37:00AM +0100, Jan Engelhardt wrote:\n>  >  > >\n>  >  > >> >do you tend to clone the entire repository repeatedly into a series\n>  >  > >> >of separate working directories\n>  >  > >>\n>  >  > >> Too time consuming on consumer drives with projects the size of Linux.\n>  >  > >\n>  >  > > git clone -l -s\n>  >  > >\n>  >  > > is not particulary slow...\n>  >  >\n>  >  > How big is a checkout of a single revision of kernel these days,\n>  >  > compared to a well-packed history since v2.6.12-rc2?\n>  >  >\n>  >  > The cost of writing out the work tree files isn't ignorable and\n>  >  > probably more than writing out the repository data (which -s\n>  >  > saves for you).\n>  >\n>  >\n>  > Depends...  I'm using ext2 for that and noatime everywhere, so that might\n>  >  change the picture, but IME it's fast enough...  As for the size, it gets\n>  >  to ~320Mb on disk, which is comparable to the pack size (~240-odd Mb).\n> \n> \n> Yesterday, i had git cloned git://foo.com/bar.git   ( 777 MiB )\n>  Today, i've git cloned git://foo.com/bar.git   ( 779 MiB )\n\nWhy do you need to clone it again ? Just git fetch from it.\n\nMike\n"},{"id":"69685","messageId":"998d0e4a0802230910o1cd087f1y6b2398cfde4cfe08@mail.gmail.com","threadId":"12266","inReplyTo":"20080223140153.GB5811@hashpling.org","subject":"Re: Question about your git habits","fromName":"J.C. Pizarro","fromEmail":"jcpiza@gmail.com","sentAt":"2008-02-23T17:10:58Z","receivedAt":"2008-02-23T17:10:58Z","isPatch":false,"sender":{"key":"jcpiza@gmail.com","avatar":null},"body":"On 2008/2/23, Charles Bailey <charles@hashpling.org> wrote:\n> On Sat, Feb 23, 2008 at 02:36:59PM +0100, J.C. Pizarro wrote:\n>  > On 2008/2/23, Charles Bailey <charles@hashpling.org> wrote:\n>  > >\n>\n> > > It shouldn't matter how aggressively the repositories are packed or what\n>  > >  the binary differences are between the pack files are. git clone\n>  > >  should (with the --reference option) generate a new pack for you with\n>  > >  only the missing objects. If these objects are ~52 MiB then a lot has\n>  > >  been committed to the repository, but you're not going to be able to\n>  > >  get around a big download any other way.\n>  >\n>  > You're wrong, nothing has to be commited ~52 MiB to the repository.\n>  >\n>  > I'm not saying \"commit\", i'm saying\n>  >\n>  > \"Assume A & B binary git repos and delta_B-A another binary file, i\n>  > request built\n>  > B' = A + delta_B-A where is verified SHA1(B') = SHA1(B) for avoiding\n>  > corrupting\".\n>  >\n>  > Assume B is the higher repacked version of \"A + minor commits of the day\"\n>  > as if B was optimizing 24 hours more the minimum spanning tree. Wow!!!\n>  >\n>\n>\n> I'm not sure that I understand where you are going with this.\n>  Originally, you stated that if you clone a 775 MiB repository on day\n>  one, and then you clone it again on day two when it was 777 MiB, then\n>  you currently have to download 775 + 777 MiB of data, whereas you\n>  could download a 52 MiB binary diff. I have no idea where that value\n>  of 52 MiB comes from, and I've no idea how many objects were committed\n>  between day one and day two. If we're going to talk about details,\n>  then you need to provide more details about your scenario.\n\nI don't said that \"A & B binary git repos\" are binary files, but i said that\ndelta_B-A is a binary file.\n\nI said ago ~15 hours \"Suppose the size cost of this binary delta is e.g. around\n52 MiB instead of 2 MiB due to numerous mismatching of binary parts ...\"\n\nThe binary delta is different to the textual delta (between lines of texts)\n used in the git scheme (the commits or changesets use textual deltas).\nThe textual delta can be compressed resulting a smaller binary object.\nCollecting binary objects and some more is the git repository.\nYou can't apply textual delta of git repository, only binary delta.\nYou can apply binary delta of both git-repacked repositories if there\nis a program\n that generates binary delta of both directories but it's not implement yet.\nThe SHA1 verifier is useful for avoid the corrupting of the generated repository\n (if it's corrupted then it has to be cloned again delta or whole\nuntil non-corrupted).\nAn example of same SHA1 of both directories can be implemented as same SHA1\n of sorted SHA1s of contents, filenames and properties. Anything\nalterated, added\n or eliminated from them implies different SHA1.\n\nDon't you understand i'm saying? I will give you a practical example.\n1. zip -r -8  foo1.zip foo1  # in foo1 there are tons of information\nas from git repo\n2. mv foo1 foo2 ; cp bar.txt foo2/\n3. zip -r -9 foo2.zip foo2   # still little bit more optimized (=\nhigher repacked)\n4. Apply binary delta between foo1.zip & foo2.zip with a supposed program\n     deltaier and you get delta_foo1_foo2.bin. The size(delta_foo1_foo2.bin) is\n     not nearly ~( size(foo2.zip) - size(foo1.zip) )\n5. Apply hexadecimal diff and you will understand why it gives the exemplar\n     ~52 MiB instead of ~2 MiB that i said it.\n6. You will know some identical parts in both foo1.zip and foo2.zip.\n     Identical parts are good for smaller binary deltas. It's possible to get\n     still smaller binary deltas when their identical parts are in\nrandom offsets\n     or random locations depending of how deltaier program is advanced.\n7. Same above but instead of both files, apply binary delta of both directories.\n\n>  Having said that, here is my original point in some more detail. git\n>  repositories are not binary blobs, they are object databases. Better\n>  than this, they are databases of immutable objects. This means that to\n>  get the difference between one database and another, you only need to\n>  add the objects that are missing from the other database.\n\nDatabases of immutable objects <--- You're wrong because you confuse.\nThere are mutable objects as the better deltas of min. spanning tree.\n\nThe missing objects are not only the missing sources that you're thinking,\nthey can be any thing (blob, tree, commit, tag, etc.). The deltas of the\nminimum spanning tree too are objects of the database that can be erased\nor added when the spanning tree is alterated (because the alterated spanning\ntree is smaller than previous) for better repack. Best repack is still\nNP-problem\nand to solve this bigger NP-problem of each day is 24/365 (eternal computing).\n\nThe git database is the top-level \".git/\" directory but it has repacked binary\ninformation and has always some size measured normally in MiBs that i was\nsaying above.\n\n>                                                                        If the two\n>  databases are actually a database and the same database at short time\n>  interval later, then almost all the objects are going to be common and\n>  the difference will be a small set of objects. Using git:// this set\n>  of objects can be efficiently transfered as a pack file.\n\nYou're saying    repacked(A) + new objects   with the bandwith cost of\nnew objects\nbut i'm saying  rerepacked(A+new objects)   with the bandwith cost of\nbinary delta\n                                   where delta is repacked(A) -\nrerepacked(A+new objects)\n                                         and rerepacked(X) is more\ntime repacking again X.\n\n>                                                                                     You may have\n>  a corner case scenario where the following isn't true, but in my\n>  experience an incremental pack file will be a more compact\n>  representation of this difference than a binary difference of two\n>  aggressively repacked git repositories as generated by a generic\n>  binary difference engine.\n\nYes, it's more simple and compact, but the eternal repacking 24/365 can do it\n e.g. 30% smaller after few weeks when the incremental pack has made nothing.\n\nIt's good idea that the weekly user picks the binary delta and the\ndaily developer\n picks the incremental pack. Put both modes working in the git server.\n\n>  I'm sorry if I've misunderstood your last point. Perhaps you could\n>  expand in the exact issue that are having if I have, as I'm not sure\n>  that I've really answered your last message.\n\n   Misunderstood can be dissappeared ;)\n"},{"id":"69688","messageId":"20080223181631.GA9405@hashpling.org","threadId":"12266","inReplyTo":"998d0e4a0802230910o1cd087f1y6b2398cfde4cfe08@mail.gmail.com","subject":"Re: Question about your git habits","fromName":"Charles Bailey","fromEmail":"charles@hashpling.org","sentAt":"2008-02-23T18:16:31Z","receivedAt":"2008-02-23T18:16:31Z","isPatch":false,"sender":{"key":"charles@hashpling.org","avatar":"https://avatars.githubusercontent.com/u/1668475?v=4"},"body":"I've cut the cc'list down to just the git mailing list as this isn't a\nlinux kernel issue.\n\nOn Sat, Feb 23, 2008 at 06:10:58PM +0100, J.C. Pizarro wrote:\n> Don't you understand i'm saying? I will give you a practical example.\n> 1. zip -r -8  foo1.zip foo1  # in foo1 there are tons of information\n> as from git repo\n> 2. mv foo1 foo2 ; cp bar.txt foo2/\n> 3. zip -r -9 foo2.zip foo2   # still little bit more optimized (=\n> higher repacked)\n> 4. Apply binary delta between foo1.zip & foo2.zip with a supposed program\n>      deltaier and you get delta_foo1_foo2.bin. The size(delta_foo1_foo2.bin) is\n>      not nearly ~( size(foo2.zip) - size(foo1.zip) )\n> 5. Apply hexadecimal diff and you will understand why it gives the exemplar\n>      ~52 MiB instead of ~2 MiB that i said it.\n> 6. You will know some identical parts in both foo1.zip and foo2.zip.\n>      Identical parts are good for smaller binary deltas. It's possible to get\n>      still smaller binary deltas when their identical parts are in\n> random offsets\n>      or random locations depending of how deltaier program is advanced.\n> 7. Same above but instead of both files, apply binary delta of both directories.\n\nI totally understand what you are saying here with your zip example.\nIn fact this supports my original interpretation of what you were\nsaying. There size of the difference between the 775 MiB repository\nand the 777 MiB repository is 52 MiB, not because there is 52 MiB of\nnew data in the latter repoistory but because of the difficulty in\ngenerating a minimal binary delta between the two.\n\nThis is why I suggest that an incremental pack file will probably make\na better method of supplying a 'diff' between the two.\n\n> Databases of immutable objects <--- You're wrong because you confuse.\n> There are mutable objects as the better deltas of min. spanning tree.\n> \n> The missing objects are not only the missing sources that you're thinking,\n> they can be any thing (blob, tree, commit, tag, etc.). The deltas of the\n> minimum spanning tree too are objects of the database that can be erased\n> or added when the spanning tree is alterated (because the alterated spanning\n> tree is smaller than previous) for better repack. Best repack is still\n> NP-problem\n> and to solve this bigger NP-problem of each day is 24/365 (eternal computing).\n> \n> The git database is the top-level \".git/\" directory but it has repacked binary\n> information and has always some size measured normally in MiBs that i was\n> saying above.\n\nYou're confusing two things together here. Conceptually, the git\ndatabase is a database of immutable objects. How it is stored is a\nlower level implementation detail (albeit a very important one in\npractice). The delta chains in the pack files are nothing to do with\ngit objects.\n\n> >                                                                        If the two\n> >  databases are actually a database and the same database at short time\n> >  interval later, then almost all the objects are going to be common and\n> >  the difference will be a small set of objects. Using git:// this set\n> >  of objects can be efficiently transfered as a pack file.\n> \n> You're saying    repacked(A) + new objects   with the bandwith cost of\n> new objects\n> but i'm saying  rerepacked(A+new objects)   with the bandwith cost of\n> binary delta\n>                                    where delta is repacked(A) -\n> rerepacked(A+new objects)\n>                                          and rerepacked(X) is more\n> time repacking again X.\n\nYou seem to be comparing something that I've said with something that\nyou said. Originally I thought that you were making a bandwidth\nargument, now you seem to be making a repacking time argument. Is X\nsupposed to represent to second cloned repository?\n\nIf you git clone with --reference or git fetch from a non-dumb source\nrepository then the remote end will generate a packfile of just the\nobjects that you need to update the local repository. If the remote\nside is fully packed then A can reuse the delta information it already\nhas to generate this pack efficiently. On the local side, there is no\nneed to unpack these objects at all. The pack can just be placed in\nthe repository and used as as.\n\n> >                                                                                     You may have\n> >  a corner case scenario where the following isn't true, but in my\n> >  experience an incremental pack file will be a more compact\n> >  representation of this difference than a binary difference of two\n> >  aggressively repacked git repositories as generated by a generic\n> >  binary difference engine.\n> \n> Yes, it's more simple and compact, but the eternal repacking 24/365 can do it\n>  e.g. 30% smaller after few weeks when the incremental pack has made nothing.\n\nWhat do you mean by 'the eternal repacking 24/365'? What is it trying\nto achieve?\n\n> It's good idea that the weekly user picks the binary delta and the\n> daily developer\n>  picks the incremental pack. Put both modes working in the git server.\n\nWhat is the weekly user? Why would the 'binary delta' be better than\nan incremental pack in this case?\n"},{"id":"69690","messageId":"998d0e4a0802231019v2bd14a9nc5b2bbad3c2923d6@mail.gmail.com","threadId":"12266","inReplyTo":"998d0e4a0802230910o1cd087f1y6b2398cfde4cfe08@mail.gmail.com","subject":"Re: Question about your git habits","fromName":"J.C. Pizarro","fromEmail":"jcpiza@gmail.com","sentAt":"2008-02-23T18:19:06Z","receivedAt":"2008-02-23T18:19:06Z","isPatch":false,"sender":{"key":"jcpiza@gmail.com","avatar":null},"body":"The google's gmail made a crap my last message that it did wrap\nmy message of X lines to the crap of (X+o) lines misconfiguring\nmy original lines of the message.\n\n    I don't see the motives of Google crapping my original lines\n    of the messages that i had sended.\n"},{"id":"69695","messageId":"998d0e4a0802231047t1338439cj1a1c98f046e6ebaf@mail.gmail.com","threadId":"12266","inReplyTo":"20080223181631.GA9405@hashpling.org","subject":"Re: Question about your git habits","fromName":"J.C. Pizarro","fromEmail":"jcpiza@gmail.com","sentAt":"2008-02-23T18:47:13Z","receivedAt":"2008-02-23T18:47:13Z","isPatch":false,"sender":{"key":"jcpiza@gmail.com","avatar":null},"body":"On 2008/2/23, Charles Bailey <charles@hashpling.org> wrote:\n> You're confusing two things together here. Conceptually, the git\n>  database is a database of immutable objects. How it is stored is a\n>  lower level implementation detail (albeit a very important one in\n>  practice). The delta chains in the pack files are nothing to do with\n>  git objects.\n\nIn Documentation/git-repack.txt says:\n\ngit-repack is used to combine all objects that do not currently\nreside in a \"pack\", into a pack. It can also be used to re-organize\nexisting packs into a single, more efficient pack.\n\nA pack is a collection of objects, individually compressed, with\ndelta compression applied, stored in a single file, with an\nassociated index file.\n\n### Can you explain me that delta chains in the pack files are\n nothing to do with git objects? ###\n\nPacks are used to reduce the load on mirror systems, backup engines,\ndisk storage, etc.\n\n> You seem to be comparing something that I've said with something that\n>  you said. Originally I thought that you were making a bandwidth\n>  argument, now you seem to be making a repacking time argument. Is X\n>  supposed to represent to second cloned repository?\n\nYes, X is as the 2nd cloned repository but highly repacked, same size is not.\n\n>\n>  If you git clone with --reference or git fetch from a non-dumb source\n>  repository then the remote end will generate a packfile of just the\n>  objects that you need to update the local repository. If the remote\n>  side is fully packed then A can reuse the delta information it already\n>  has to generate this pack efficiently. On the local side, there is no\n>  need to unpack these objects at all. The pack can just be placed in\n>  the repository and used as as.\n\nIs not it redundant to place git objects and pack files in the same repo?\n1. Or erase the unnecesary pack files because there are git objects.\n2. Or erase some git objects because there are delta chains in pack files\n     that can generate the same git objects erased previously.\n\n> What do you mean by 'the eternal repacking 24/365'? What is it trying\n>  to achieve?\n\nIt's an uninterrumpted computing that is generating a sequence of\nspanning trees in convergence to smaller packs.\n   Each smaller spanning tree is found, the pack file is updated.\n\n> What is the weekly user? Why would the 'binary delta' be better than\n>  an incremental pack in this case?\n\nBecause the user wants to clone weekly 240 MiB in 1st week, 220 MiB in\n2nd week, 205 MiB in 3rd week, .... 100 MiB repo! in Nth week instead of\n240+1+1+1+1 MiB of incremental packs.\n\nWhat is better for the user in the Nth week, 100 MiB repo or 244 MiB repo?\n\n   ;)\n"},{"id":"69698","messageId":"20080223192858.GA10655@hashpling.org","threadId":"12266","inReplyTo":"998d0e4a0802231047t1338439cj1a1c98f046e6ebaf@mail.gmail.com","subject":"Re: Question about your git habits","fromName":"Charles Bailey","fromEmail":"charles@hashpling.org","sentAt":"2008-02-23T19:28:58Z","receivedAt":"2008-02-23T19:28:58Z","isPatch":false,"sender":{"key":"charles@hashpling.org","avatar":"https://avatars.githubusercontent.com/u/1668475?v=4"},"body":"On Sat, Feb 23, 2008 at 07:47:13PM +0100, J.C. Pizarro wrote:\n> On 2008/2/23, Charles Bailey <charles@hashpling.org> wrote:\n> > You're confusing two things together here. Conceptually, the git\n> >  database is a database of immutable objects. How it is stored is a\n> >  lower level implementation detail (albeit a very important one in\n> >  practice). The delta chains in the pack files are nothing to do with\n> >  git objects.\n> \n> In Documentation/git-repack.txt says:\n> \n> git-repack is used to combine all objects that do not currently\n> reside in a \"pack\", into a pack. It can also be used to re-organize\n> existing packs into a single, more efficient pack.\n> \n> A pack is a collection of objects, individually compressed, with\n> delta compression applied, stored in a single file, with an\n> associated index file.\n> \n> ### Can you explain me that delta chains in the pack files are\n>  nothing to do with git objects? ###\n\nIt's an abstraction thing. Perhaps I should have said that git objects\nhave nothing to do with pack files to indicate the direction of the\ndependency.\n\n> Is not it redundant to place git objects and pack files in the same repo?\n> 1. Or erase the unnecesary pack files because there are git objects.\n> 2. Or erase some git objects because there are delta chains in pack files\n>      that can generate the same git objects erased previously.\n\nOnly if they overlap, but usually they don't.\n\n> > What is the weekly user? Why would the 'binary delta' be better than\n> >  an incremental pack in this case?\n> \n> Because the user wants to clone weekly 240 MiB in 1st week, 220 MiB in\n> 2nd week, 205 MiB in 3rd week, .... 100 MiB repo! in Nth week instead of\n> 240+1+1+1+1 MiB of incremental packs.\n> \n> What is better for the user in the Nth week, 100 MiB repo or 244 MiB repo?\n> \n\nThat depends, doesn't it. If the everyday workflow is quicker and\neasier a 244 MiB clone could well be acceptable, but if it's not there\nis always the option of a repack. I don't buy the premise that people\nwant to be continually repacking to find the ultimate pack file, I\ndon't think that the gain over a one-shot repack is ever going to be\nworth it.\n"}]}