{"thread":{"id":"3265","subject":"Handling large files with GIT","startedAt":"2006-02-08T09:14:18Z","lastAt":"2006-02-16T20:32:11Z","messageCount":39,"participants":["Martin Langhoff","Johannes Schindelin","Linus Torvalds","Junio C Hamano","Florian Weimer","Greg KH","Ben Clifford","Jeff Garzik","Ian Molton","Keith Packard","Sam Vilain","Fredrik Kuivinen"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"15711","messageId":"46a038f90602080114r2205d72cmc2b5c93f6fffe03d@mail.gmail.com","threadId":"3265","inReplyTo":null,"subject":"Handling large files with GIT","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-02-08T09:14:18Z","receivedAt":"2006-02-08T09:14:18Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"Roland Stigge recently pointed out a use case using very large files\nwhere GIT has some serious limitations. He is one of several Debian\ndevelopers keeping their homedir under version control with SVN (blame\nJoey Hess for this - http://www.kitenet.net/~joey/svnhome.html ).\n\nSVN does reasonably well tracking his >1GB mbox file. Now, I don't\nknow if I like the idea of putting my own mbox file under version\ncontrol, but it looks like projects with large and slow-changing files\nwould be in trouble with GIT. Not literally trouble, but gross\ninefficiencies.\n\nThe problems are two. At commit time, a full copy is stored in the\nobject database until git-repack && git-prune-packed are called. And\nduring the transfer over the git protocol we send the full object,\neven if both ends have objects that are good candidates for a small\ndelta.\n\nI'm not strong on either aspect of git (packfile format or git\nprotocol), and I don't personally deal with large files. So feel free\nto ignore me for the time being. If it ever itches, you might get a\npatch...\n\ncheers,\n\n\nmartin\n"},{"id":"15712","messageId":"Pine.LNX.4.63.0602081248270.31700@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3265","inReplyTo":"46a038f90602080114r2205d72cmc2b5c93f6fffe03d@mail.gmail.com","subject":"Re: Handling large files with GIT","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-02-08T11:54:24Z","receivedAt":"2006-02-08T11:54:24Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 8 Feb 2006, Martin Langhoff wrote:\n\n> Roland Stigge recently pointed out a use case using very large files\n> where GIT has some serious limitations.\n\nThat is intentional: git handles source code very well, where you tend to \nhave small files, and it handles branches very well, where you tend to \nhave mostly the same files in different branches.\n\nI am uncertain if it is possible to extend git to handle large files \ngracefully, without slowing it down for its main use case.\n\n[thinking] A potentially silly idea just hit me: We could virtually cut \nevery file into 256kB chunks. That would not affect source code at all: \nanybody producing a 256kB C file should be shot anyway.\n\nIf the files just keep growing, this should help enormously. If the files \nchange subtly, the diff algorithm should work quite well on 'em.\n\nComments?\n\nCiao,\nDscho\n"},{"id":"15718","messageId":"Pine.LNX.4.64.0602080815180.2458@g5.osdl.org","threadId":"3265","inReplyTo":"Pine.LNX.4.63.0602081248270.31700@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-08T16:34:22Z","receivedAt":"2006-02-08T16:34:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 8 Feb 2006, Johannes Schindelin wrote:\n> \n> I am uncertain if it is possible to extend git to handle large files \n> gracefully, without slowing it down for its main use case.\n\nIndeed. The git architecture simply sucks for big objects. It was \ndiscussed somewhat durign the early stages, but a lot of it really is \npretty fundamental. The fact that all the operations work on a full \nobject, and the delta's are (on purpose) just a very specific and limited \nkind of size compression is just very ingrained.\n\n> [thinking] A potentially silly idea just hit me: We could virtually cut \n> every file into 256kB chunks. That would not affect source code at all: \n> anybody producing a 256kB C file should be shot anyway.\n\nIt probably wouldn't help that much, really. And it would probably impact \nsource code users too: I bet we'd have bugs. It would be a very strange \nspecial case.\n\nIt also would only help for things that purely grow at the end. Which \nisn't even true for a mailbox: it may or may not be true for your INBOX, \nbut anybody who _uses_ a mailbox format to read his email will be adding \nstatus flags to the mbox format (or deleting mbox entries etc). \n\nSo every time a small change happened that changed the offset, you'd have \nan explosion of these 256kB chunk objects, and while the delta would work \n(probably slowly - remember how the git deltification algorithm tries to \ncompare against the ten \"nearest\" neighbors), at _commit_ time you'd have \nto write that 1GB (compressed) out anyway.\n\nRealistically, I think the answer is that git just doesn't work for his \nusage case. There's two alternatives:\n\n - convince him to not have big mailboxes (an answer I don't particularly \n   like: it's a tool limitation, and you shouldn't change your behaviour \n   just because the tool doesn't work for it - you should just try to find \n   the right tool).\n\n   That said: git should actually work beautifully for email if you \n   _don't_ keep it as one big mbox. You could probably very reasonably use \n   git as a database backend, where each email is its own object, and you \n   can have many different ways of indexing them into trees (by content, \n   by date, by author, by thread).\n\n   But that's very different from the suggested \"home directory\" setup \n   would be.\n\n - try to work around some of the worst git issues. While I don't think \n   the 256kB blockign thing would help (the git protocol would still \n   always send the base versions), there _are_ probably things that could \n   be done. They'd be very invasive, though, and somebody would seriously \n   have to look at the architectural issues.\n\n   For example, right now the decision to send only \"self-contained\" packs \n   in the git protocol was a very conscious one: it's much safer, and it \n   makes the unpacking a lot easier (the unpacking doesn't ever have to \n   even read any other objects than the stream it gets). It's also (for \n   packs that we use on-disk) the only sane way to avoid nasty inter-pack \n   dependencies.\n\n   But for the git protocol, the inter-pack dependencies don't matter, \n   if we'd always unpack the thing on reception if it is not a \n   self-contained pack. So we _could_ allow delta's that depend on the \n   receiver already having the objects we delta against.\n\n   However, the deltification itself is likely very slow, exactly because \n   git (again, very much by design) generates the deltas dynamically \n   rather than depending on things already being in delta format.\n\nPersonally, I think the answer is \"git is good for lots of small files\". \nIt's very much what git was designed for, and the fact that it doesn't \nwork for everything is a trade-off for the things it _does_ work well for.\n\n\t\t\tLinus\n"},{"id":"15720","messageId":"Pine.LNX.4.64.0602080853480.2458@g5.osdl.org","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602080815180.2458@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-08T17:01:13Z","receivedAt":"2006-02-08T17:01:13Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 8 Feb 2006, Linus Torvalds wrote:\n>\n> The fact that all the operations work on a full object, and the delta's \n> are (on purpose) just a very specific and limited kind of size \n> compression is just very ingrained.\n\nSide note: the original explicit git \"delta\" objects by Nicolas Pitre \nwould have handled this large-file-case much more gracefully. \n\nThe pack-files had absolutely huge advantages, though, so I think we (I) \ndid the right thing there in making the delta code only a very specific \nspecial case..\n\nIt is possible that we could re-introduce the \"explicit delta\" object, \nthough (it's not incompatible with also doing pack-files, it's just that \npack-files made 99% of all the arguments for an explicit delta go away).\n\n\t\tLinus\n"},{"id":"15726","messageId":"7v4q39pq4t.fsf@assigned-by-dhcp.cox.net","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602080853480.2458@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-08T20:11:30Z","receivedAt":"2006-02-08T20:11:30Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Side note: the original explicit git \"delta\" objects by Nicolas Pitre \n> would have handled this large-file-case much more gracefully. \n\nTrue.\n\n> The pack-files had absolutely huge advantages, though, so I think we (I) \n> did the right thing there in making the delta code only a very specific \n> special case..\n\nWell the blame for ripping that out falls on me, actually...\n\n> It is possible that we could re-introduce the \"explicit delta\" object, \n> though (it's not incompatible with also doing pack-files, it's just that \n> pack-files made 99% of all the arguments for an explicit delta go away).\n\nI do not remember we had 'rev-list --objects' support for Nico's\nexplicit delta object chains.  If we didn't that would be a new\ndevelopment that needs to be done to resurrect it.  I know\npack-objects never had support for it so obviously that needs to\nbe added as well.  Probably explicit delta objects should always\nbe packed in full without spending cost to find delta candidates.\n\nPersonally I feel that post-1.2.0 would be a good time to start\nlooking at enhancing the pack generation chain, rev-list piped\nto pack-objects.  This \"large files\" use case is helped by\nless self-contained packs while \"shallow clone\" use case\nwe discussed earlier is helped by more self-contained packs (we\nhad a discussion long time ago on this and I think we have the\ncode to do so [*1*]). \n\nAn addition to pack-objects is needed to make it capable to read\na list of objects that we do not want to include in the\nresulting pack but can be used as base objects for delitified.\n\nBTW, as to the \"shallow clone\", I changed my mind and am\ninclined to agree with Johannes that handling cut-offs\ndifferently from grafts is easier for dealing with later \"give\nme more history\" operation, so I am planning to chuck my jc/clone\ntopic branch that I have included in the proposed updates so\nfar.\n\n[Footnote]\n\n*1* http://article.gmane.org/gmane.comp.version-control.git/5779\n"},{"id":"15729","messageId":"87slqty2c8.fsf@mid.deneb.enyo.de","threadId":"3265","inReplyTo":"46a038f90602080114r2205d72cmc2b5c93f6fffe03d@mail.gmail.com","subject":"Re: Handling large files with GIT","fromName":"Florian Weimer","fromEmail":"fw@deneb.enyo.de","sentAt":"2006-02-08T21:20:39Z","receivedAt":"2006-02-08T21:20:39Z","isPatch":false,"sender":{"key":"fw@deneb.enyo.de","avatar":null},"body":"* Martin Langhoff:\n\n> SVN does reasonably well tracking his >1GB mbox file. Now, I don't\n> know if I like the idea of putting my own mbox file under version\n> control, but it looks like projects with large and slow-changing files\n> would be in trouble with GIT.\n\nTo my surprise, it's not that bad.  The Debian testing-security team\nuses a single 1.8 MB file (400 KB compressed) to keep vulnerability\ndata.  Most changes to that file involve just a few lines.  But even\nin this extreme case, git doesn't compare too badly against Subversion\nif you pack regularly (but not too often).  Disk usage is actually\n*below* Subversion FSFS even with --depth=10 (the default,\nunfortunately a bit hard to override).\n\nI plan to do another experiment for GCC, which contains marvels such\nas:\n\n  35905  126056 1379093 gcc/ChangeLog-2005\n  12610   61215  417584 gcc/combine.c\n\nBut the outcome will likely be quite similar to the secure-testing\ncase: comparable disk space usage, not a difference in the order of\none or more magnitudes.\n\nBut Subversion still has got a significant adventage: I can get a\nworking copy without downloading full history (several gigabytes in\nGCC's case).  There's also the slight drawback that you shouldn't pack\ntoo often, otherwise you'll reduce its effectiveness.  You can always\nrun \"git-repack -a -d\", but it's rather expensive.  This means that\nyou need to keep compressed fulltexts from a few dozen revisions, but\nI don't think this is a huge burden.  All in all, the compressed\nfulltexts/packs model is a pretty good trade-off between disk usage,\nend user usability nad code complexity.\n\nIn your mbox case, you should simply try Maildir.  The tree object\n(which lists all files in the Maildir folder) will still be rather\nlarge (about 40 to 50 bytes per message stored), though.\n"},{"id":"15734","messageId":"46a038f90602081435x49e53a1cgdc56040a19768adb@mail.gmail.com","threadId":"3265","inReplyTo":"87slqty2c8.fsf@mid.deneb.enyo.de","subject":"Re: Handling large files with GIT","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-02-08T22:35:36Z","receivedAt":"2006-02-08T22:35:36Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 2/9/06, Florian Weimer <fw@deneb.enyo.de> wrote:\n> In your mbox case, you should simply try Maildir.  The tree object\n> (which lists all files in the Maildir folder) will still be rather\n> large (about 40 to 50 bytes per message stored), though.\n\nI did suggest maildir,  where GIT is bound to do well as the content\nof the emails doesn't change but they just move around a lot. Though\nyes, trees are going to be nasty.\n\nBut the interesting case I gues is the general one of large files\nchanging slowly. My guess is that supporting delta transfers in the\ngit protocol would make it a lot more manageable. For local storage\ngit isn't so bad, and the problem is perhaps harder to resolve.\n\ncheers,\n\n\nmartin\n"},{"id":"15764","messageId":"20060209045420.GB15924@kroah.com","threadId":"3265","inReplyTo":"87slqty2c8.fsf@mid.deneb.enyo.de","subject":"Re: Handling large files with GIT","fromName":"Greg KH","fromEmail":"greg@kroah.com","sentAt":"2006-02-09T04:54:20Z","receivedAt":"2006-02-09T04:54:20Z","isPatch":false,"sender":{"key":"greg@kroah.com","avatar":"https://gravatar.com/avatar/5bb5aa0cc2e01c00ec899d11130c07796bc186e465bae57bc34873b13b72c7c8?d=mp&s=160"},"body":"On Wed, Feb 08, 2006 at 10:20:39PM +0100, Florian Weimer wrote:\n> * Martin Langhoff:\n> \n> > SVN does reasonably well tracking his >1GB mbox file. Now, I don't\n> > know if I like the idea of putting my own mbox file under version\n> > control, but it looks like projects with large and slow-changing files\n> > would be in trouble with GIT.\n> \n> To my surprise, it's not that bad.  The Debian testing-security team\n> uses a single 1.8 MB file (400 KB compressed) to keep vulnerability\n> data.  Most changes to that file involve just a few lines.  But even\n> in this extreme case, git doesn't compare too badly against Subversion\n> if you pack regularly (but not too often).  Disk usage is actually\n> *below* Subversion FSFS even with --depth=10 (the default,\n> unfortunately a bit hard to override).\n\nI have a project that has 2.5Mb files, and git handles them just fine,\neven on my old slow laptop.\n\nBut when I tried to use it to backup my old email archive a few months\nago, running about 2Gb in about 300 files, it took forever.  Luckily I\nonly archive stuff off every other month or so, otherwise it would be\nunusable.\n\nHowever, I did notice one problem.  When cloning from one machine to\nanother, for a project that is already fully packed, it seems that the\nwhole project is packed again before sending it accross the wire.  With\nan archive this big, that takes over an hour for my slow old fileserver.\nI ended up just rsyncing over the files and pointing the parent back to\nthe original.  Is there anyway to not repack everything if it's not\nneeded?\n\nthanks,\n\ngreg k-h\n"},{"id":"15766","messageId":"46a038f90602082138x5ce072e9kff69e3d677bd0162@mail.gmail.com","threadId":"3265","inReplyTo":"20060209045420.GB15924@kroah.com","subject":"Re: Handling large files with GIT","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-02-09T05:38:26Z","receivedAt":"2006-02-09T05:38:26Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 2/9/06, Greg KH <greg@kroah.com> wrote:\n> Is there anyway to not repack everything if it's not\n> need?\n\nI often cg-clone large projects over stupid protocols (http/rsync) and\nthen hand-edit the branches/origin or remotes/origin file to say\n\"git://\" instead.\n\nIt's called cheating but it works great ;-)\n\n\n\nm\n"},{"id":"16008","messageId":"Pine.OSX.4.64.0602131416530.25089@piva.hawaga.org.uk","threadId":"3265","inReplyTo":"46a038f90602081435x49e53a1cgdc56040a19768adb@mail.gmail.com","subject":"Re: Handling large files with GIT","fromName":"Ben Clifford","fromEmail":"benc@hawaga.org.uk","sentAt":"2006-02-13T01:26:06Z","receivedAt":"2006-02-13T01:26:06Z","isPatch":false,"sender":{"key":"benc@hawaga.org.uk","avatar":"https://gravatar.com/avatar/c7ce083471287f8e77b69dd147f757799d1efdd740727fd7e3003f33e88be898?d=mp&s=160"},"body":"\nOn Thu, 9 Feb 2006, Martin Langhoff wrote:\n\n> I did suggest maildir,  where GIT is bound to do well as the content\n> of the emails doesn't change but they just move around a lot. Though\n> yes, trees are going to be nasty.\n\nI've been keeping maildir in git for a few months, with mail being \ndelivered into a git repo on one (permanently connected) host and me \nmerging that branch into a repo on my laptop for reading (the intention \nbeing that I should be able to sync it back to the permanently connected \nhost as I sometimes read mail there.\n\nAlas, the merge part of this absolutely sucks -- as time goes by, its \ngetting slower and slower (its taking an hour or so to do the merge, which \nhas got to the point of being barely usable -- if it wasn't for the neat \nhack-value, I'd have given up on this by now).\n\nI haven't really probed whats happening when I'm doing the merges in any \ndepth, but I see a lot of index manipulation happening (git update-index, \nI think) to add and remove files where each invocation of that seems to be \ntaking almost a whole second.\n\nI wonder if the present merge algorithms perform especially badly in the \ncase of a large number of files with lots of renames (and so lots of \nadds/removes) but no content changes? The merge should be able to happen \nentirely in the index, I think.\n\nPerhaps one way to proceed would be for me to write a move-optimised merge \nstrategy where I flip the index round and instead of saying \"how has the \ncontent inside this filename changed?\" I instead say \"how has the filename \nassociated with this content <hash> changed?\"\n\nA special-case on top of a move-optimised merge might be some \nmaildir-aware filename handling that knows how to resolve conflicts when a \nparticular content-hash has been renamed to two different names (eg. when \none flag is added to a message in one repo and a different flag is added \nto a message in another repo).\n\nAny advice/thoughts/suggestions-that-this-is-a-stupid-thing-to-do would be \ngreatly appreciated.\n\n-- \n"},{"id":"16016","messageId":"Pine.LNX.4.64.0602121939070.3691@g5.osdl.org","threadId":"3265","inReplyTo":"Pine.OSX.4.64.0602131416530.25089@piva.hawaga.org.uk","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-13T03:42:41Z","receivedAt":"2006-02-13T03:42:41Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 13 Feb 2006, Ben Clifford wrote:\n> \n> I've been keeping maildir in git for a few months, with mail being delivered\n> into a git repo on one (permanently connected) host and me merging that branch\n> into a repo on my laptop for reading (the intention being that I should be\n> able to sync it back to the permanently connected host as I sometimes read\n> mail there.\n> \n> Alas, the merge part of this absolutely sucks -- as time goes by, its getting\n> slower and slower (its taking an hour or so to do the merge, which has got to\n> the point of being barely usable -- if it wasn't for the neat hack-value, I'd\n> have given up on this by now).\n\nIf it takes an hour per merge, it _is_ unusable. I consider 15 _seconds_ \nto be pretty unusable.\n\nCan you do a\n\n\tgit-ls-tree -r -t HEAD\n\tgit-ls-tree -r -t HEAD^1\n\tgit-ls-tree -r -t HEAD^2\n\nafter a merge, and put the three resulting files up somewhere public (I \nassume the filenames aren't going to be anything private, I don't know how \nmaildir organizes stuff) so that people can get an idea of what ends up \nbeing involved there..\n\n\t\tLinus\n"},{"id":"16017","messageId":"46a038f90602122040r55ba8504o9f4b783ea9571f85@mail.gmail.com","threadId":"3265","inReplyTo":"Pine.OSX.4.64.0602131416530.25089@piva.hawaga.org.uk","subject":"Re: Handling large files with GIT","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-02-13T04:40:20Z","receivedAt":"2006-02-13T04:40:20Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 2/13/06, Ben Clifford <benc@hawaga.org.uk> wrote:\n> Any advice/thoughts/suggestions-that-this-is-a-stupid-thing-to-do would be\n> greatly appreciated.\n\nHow many files are you talking about?\n\n\nm\n"},{"id":"16019","messageId":"Pine.LNX.4.64.0602122049010.3691@g5.osdl.org","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602121939070.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-13T04:57:25Z","receivedAt":"2006-02-13T04:57:25Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 12 Feb 2006, Linus Torvalds wrote:\n> \n> If it takes an hour per merge, it _is_ unusable. I consider 15 _seconds_ \n> to be pretty unusable.\n\nBtw, one thing to realize is that git is inherently a lot better at \nhandling lots of files in _subdirectories_, especially if those \nsubdirectories don't change.\n\nI've never used maildir layout, but if it is a couple of large _flat_ \nsubdirectories, git will potentially handle that a lot worse than if you \nhave a hierarchy of directories.\n\nI say \"potentially\", because if the directories are all mutable and \nchange, then the flat approach is better. But if they tend to have some \nkind of stability, a lot of git operations (diffing and merging in \nparticular) are able to see that two subdirectories are 100% equal, and \nwill entirely skip them.\n\nThis is a large part of why git performs well on the kernel. Most merges \ndon't actually touch all - or even a very big percentage - of the over \nthousand subdirectories in the kernel. Git can quickly see and ignore the \nwhole subdirectory when that happens - the SHA1 is exactly the same, so \ngit knows that every file under that subdirectory (and every recursive \ndirectory) is the same.\n\nIn contrast, if you have a million files in one directory, and 10 of them \nchanged, git will still have to check the SHA1's for matches for the other \n999,990 files. Which is going to be slow.\n\nThat said, I suspect there's space for optimization. \n\n\t\tLinus\n"},{"id":"16020","messageId":"Pine.LNX.4.64.0602122058260.3691@g5.osdl.org","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602122049010.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-13T05:05:55Z","receivedAt":"2006-02-13T05:05:55Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 12 Feb 2006, Linus Torvalds wrote:\n> \n> This is a large part of why git performs well on the kernel. Most merges \n> don't actually touch all - or even a very big percentage - of the over \n> thousand subdirectories in the kernel. Git can quickly see and ignore the \n> whole subdirectory when that happens - the SHA1 is exactly the same, so \n> git knows that every file under that subdirectory (and every recursive \n> directory) is the same.\n\nFinal note: this means, for example, that git is relatively bad at \ntracking a \"hashed\" nested file directory (like the one git itself uses), \nbecause new files will end up randomly appearing in every directory, and \nno directory is ever \"stable\".\n\nIn contrast, if the directory structure is - for example - something where \nyou index files by date, and subdirectories with older dates are thus much \nmore naturally likely to be quiescent, the \"this tree is the same\" \noptimizations work very well.\n\nBasically, a lot of the git speed optimizations depend on \"on average, \nthings stay the same\". We may have 18,000+ files in the kernel, but most \npatches will change maybe five of them. There's a lot of fairly static \ncontent and the changes have a certain level of \"locality\". It's normally \na hundred-line patch to one file, not a hundred files that had one-liners. \nAnd when 20 files are changed, most of them tend to be in the same \nsubdirectory, etc etc.\n\nTaking advantage of those kinds of things is what makes git good at \nhandling software projects. But it wouldn't necessarily be how you lay out \na mail directory, for example. An automated file store might want to \nspread out the changes on purpose.\n\n\t\tLinus\n"},{"id":"16022","messageId":"43F01F5A.5020808@pobox.com","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602122049010.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Jeff Garzik","fromEmail":"jgarzik@pobox.com","sentAt":"2006-02-13T05:55:38Z","receivedAt":"2006-02-13T05:55:38Z","isPatch":false,"sender":{"key":"jgarzik@pobox.com","avatar":null},"body":"Linus Torvalds wrote:\n> I've never used maildir layout, but if it is a couple of large _flat_ \n> subdirectories,\n\nThat's what it is :/   One directory per mail folder, with each email an \nindividual file in that dir.\n\nOn the side I run a mini-ISP that offers (among other things) \nSMTP/IMAP/POP via exim/dovecot.  I use maildir for my customers, and am \ndesperately searching for a better solution.  David Woodhouse and I talk \noff and on about using git for distributed mail storage.\n\nOTOH, I wonder if a P2P-ish, sha1-indexed network service wouldn't be \nbetter for both a git fallback and email storage.\n\n\tJeff\n"},{"id":"16063","messageId":"1139810847.4183.85.camel@evo.keithp.com","threadId":"3265","inReplyTo":"43F01F5A.5020808@pobox.com","subject":"Re: Handling large files with GIT","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-02-13T06:07:27Z","receivedAt":"2006-02-13T06:07:27Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Mon, 2006-02-13 at 00:55 -0500, Jeff Garzik wrote:\n> Linus Torvalds wrote:\n> > I've never used maildir layout, but if it is a couple of large _flat_ \n> > subdirectories,\n> \n> That's what it is :/   One directory per mail folder, with each email an \n> individual file in that dir.\n\nand, named to include a hash of the contents so that they are always\nunique within a multi-folder mail store (makes refile easier).\n\n-- \nkeith.packard@intel.com\n"},{"id":"16048","messageId":"Pine.LNX.4.64.0602130806070.3691@g5.osdl.org","threadId":"3265","inReplyTo":"43F01F5A.5020808@pobox.com","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-13T16:19:10Z","receivedAt":"2006-02-13T16:19:10Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 13 Feb 2006, Jeff Garzik wrote:\n>\n> Linus Torvalds wrote:\n> > I've never used maildir layout, but if it is a couple of large _flat_\n> > subdirectories,\n> \n> That's what it is :/   One directory per mail folder, with each email an\n> individual file in that dir.\n\nOk.\n\nAnyway, I double-checked, and I'm wrong anyway. While the \"static \ndirectories\" thing is a huge performance optimization for doing many \nthings (diffing trees, file history in git-rev-list, etc etc), for merging \nit doesn't help. We always end up expanding the whole tree.\n\nWhich is kind of sad.\n\nIt's inevitable in one sense: we do the merge in the index, after all, and \nthe index - unlike the tree structures - is a flat file (like the \n\"manifest\" in mercurial or monotone). It's also represented that way in \nmemory. \n\nHowever, it is a total and complete waste in other cases.\n\nThinking more about it, this is also why merging causes all the horrible \nindex performance: not only do we (unnecessarily) read the same trees over \nand over again only to collapse them back to stage0 later when they are \nthe same, but because we keep the index in a linear format, when we read \nthe other trees, we'll have to move things around with memmove() (just the \npointers, but still).\n\nWe'd actually be a _lot_ better off if we split \"git-read-tree\" up into \ntwo phases: one that did the recursive tree operation (which can optimize \nthe \"same tree everywhere\" case), and the second stage that actually \npopulated the index.\n\nI'll have to think about this. It would be an absolutely _huge_ \noptimization for merging in certain patterns, it just doesn't matter for \nsomething like the kernel with \"just\" 18,000 files and not a lot of \nstrange merging going on.\n\nIn contrast, I can see a mail archive easily having hundreds of thousands \nof individual emails. At which time it's horribly stupid to read them all \nin three times (for a merge - base, origin, new) and do so in a pretty \ninefficient manner.\n\nHo humm. It doesn't look _hard_ per se, and I think the two-stage \ngit-read-tree is actually also what the recursive merge strategy wants \nanyway (it can't use the index - it really just wants to get a list of \nconflict information). So this definitely sounds like the RightThing(tm) \nto do anyway, and it fits the git data structures really well.\n\nSo no downsides. Except that this is some rather core code, and you can't \nafford to get it wrong. And the fact that I'm a lazy bastard, of course.\n\n\t\t\tLinus\n"},{"id":"16062","messageId":"43F113A5.2080506@f2s.com","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602122058260.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Ian Molton","fromEmail":"spyro@f2s.com","sentAt":"2006-02-13T23:17:57Z","receivedAt":"2006-02-13T23:17:57Z","isPatch":false,"sender":{"key":"spyro@f2s.com","avatar":null},"body":"Linus Torvalds wrote:\n\n> Taking advantage of those kinds of things is what makes git good at \n> handling software projects. But it wouldn't necessarily be how you lay out \n> a mail directory, for example. An automated file store might want to \n> spread out the changes on purpose.\n\nIndeed...\n\nIm curious as to why anyone would want to use a SCM tool on a mail dir \nanyway - surely no-one edits their pasnt mails and needs to keep logs?\n\nsurely incremental backups would be a better way to manage something \nlike this ?\n"},{"id":"16061","messageId":"46a038f90602131519l1d281a32o80dbc5621fbf00af@mail.gmail.com","threadId":"3265","inReplyTo":"43F113A5.2080506@f2s.com","subject":"Re: Handling large files with GIT","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-02-13T23:19:23Z","receivedAt":"2006-02-13T23:19:23Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 2/14/06, Ian Molton <spyro@f2s.com> wrote:\n> Im curious as to why anyone would want to use a SCM tool on a mail dir\n> anyway - surely no-one edits their pasnt mails and needs to keep logs?\n>\n> surely incremental backups would be a better way to manage something\n> like this ?\n\nWell, the maildir arrangement is friendlier to something content-smart\nlike git than to something content-stupid like a backup tool. Files\ninside a maildir change name/location to reflect changes in status,\nbut their content tends to remain the same.\n\ngit does great in this scenario, except for the \"dealing with a\nbazillion files\" part of it.\n\ncheers,\n\n\nmartin\n"},{"id":"16065","messageId":"46a038f90602131607k63aa32abpa273b2325fc4b1b8@mail.gmail.com","threadId":"3265","inReplyTo":"1139810847.4183.85.camel@evo.keithp.com","subject":"Re: Handling large files with GIT","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-02-14T00:07:34Z","receivedAt":"2006-02-14T00:07:34Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 2/13/06, Keith Packard <keithp@keithp.com> wrote:\n> and, named to include a hash of the contents\n\nreally? I thought it was something like md5(hostname+datestamp+random).\n\ncheers,\n\nm\n"},{"id":"16131","messageId":"Pine.LNX.4.63.0602141953000.22451@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3265","inReplyTo":"43F113A5.2080506@f2s.com","subject":"Re: Handling large files with GIT","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-02-14T18:56:48Z","receivedAt":"2006-02-14T18:56:48Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 13 Feb 2006, Ian Molton wrote:\n\n> Im curious as to why anyone would want to use a SCM tool on a mail dir \n> anyway - surely no-one edits their pasnt mails and needs to keep logs?\n> \n> surely incremental backups would be a better way to manage something \n> like this?\n\nPoint is, if you want to read your email on different computers (like one \ndesktop and one laptop), you are quite well off managing them with git. Of \ncourse, you could rsync them from/to the other computer. But rsync is slow \nonce you accumulated enough files, since it has to compare the hashes of \ntons of files (or file chunks). Git knows if they have changed.\n\nHth,\nDscho\n"},{"id":"16138","messageId":"Pine.LNX.4.64.0602141108050.3691@g5.osdl.org","threadId":"3265","inReplyTo":"Pine.LNX.4.63.0602141953000.22451@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-14T19:52:26Z","receivedAt":"2006-02-14T19:52:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 14 Feb 2006, Johannes Schindelin wrote:\n> \n> Point is, if you want to read your email on different computers (like one \n> desktop and one laptop), you are quite well off managing them with git. Of \n> course, you could rsync them from/to the other computer. But rsync is slow \n> once you accumulated enough files, since it has to compare the hashes of \n> tons of files (or file chunks). Git knows if they have changed.\n\nYes. I actually think that git would be a _wonderful_ email tracking tool, \nbut that may not mean that it's a wonderful tool for tracking all \nparticular email layout possibilities. It clearly is _not_ a wonderful \ntool for tracking mbox-style email setups, for example ;)\n\nI suspect we actually could make the \"one linear directory\" setup perform \npretty well. It wouldn't be the best possible layout (by far), but I think \nour problems there are just because of some decisions we've (me, mostly) \nmade that didn't take that layout into account. I don't think the problems \nare in any way fundamental.\n\nThat said, I think git could do much better if the layout was optimized \nfor git. For example, in the maildir thing, there's two issues: the flat \ndirectory structure is sub-optimal, but the other thing seems to be that \nmaildir apparently saves metadata in the filename.\n\nSaving meta-data in the filename should actually work wonderfully well \nwith git, but both merging and git-diff-tree consider the filename to be \nthe \"index\", so they optimize for that. You could do indexing the other \nway around, and consider the contents to be the index (and the filename is \nthe \"status\"), but that's obviously not sane for a sw project, even if it \nmight be exactly what you want to do for mail handling.\n\n\t\tLinus\n"},{"id":"16151","messageId":"43F249F7.5060008@vilain.net","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602141108050.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2006-02-14T21:21:59Z","receivedAt":"2006-02-14T21:21:59Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> That said, I think git could do much better if the layout was optimized \n> for git. For example, in the maildir thing, there's two issues: the flat \n> directory structure is sub-optimal, but the other thing seems to be that \n> maildir apparently saves metadata in the filename.\n> \n> Saving meta-data in the filename should actually work wonderfully well \n> with git, but both merging and git-diff-tree consider the filename to be \n> the \"index\", so they optimize for that. You could do indexing the other \n> way around, and consider the contents to be the index (and the filename is \n> the \"status\"), but that's obviously not sane for a sw project, even if it \n> might be exactly what you want to do for mail handling.\n\nThis seems to me to be another use case where git could gain orders of\nmagnitude speed improvement by either explicit (\"forensic\") history\nobjects, or a history analysis cache.\n\nSam.\n"},{"id":"16158","messageId":"Pine.LNX.4.64.0602141357300.3691@g5.osdl.org","threadId":"3265","inReplyTo":"43F249F7.5060008@vilain.net","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-14T22:01:33Z","receivedAt":"2006-02-14T22:01:33Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 15 Feb 2006, Sam Vilain wrote:\n> \n> This seems to me to be another use case where git could gain orders of\n> magnitude speed improvement by either explicit (\"forensic\") history\n> objects, or a history analysis cache.\n\nWell, the thing is, it could get that _without_ any cache too.\n\nThe problem really isn't that we couldn't make things faster, the problem \nis that at least for _me_ the thing is fast enough.\n\nIf somebody is interested in making the \"lots of filename changes\" case go \nfast, I'd be more than happy to walk them through what they'd need to \nchange. I'm just not horribly motivated to do it myself. Hint, hint.\n\n\t\t\tLinus\n"},{"id":"16163","messageId":"7vy80dpo9g.fsf@assigned-by-dhcp.cox.net","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602141357300.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-14T22:30:03Z","receivedAt":"2006-02-14T22:30:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> If somebody is interested in making the \"lots of filename changes\" case go \n> fast, I'd be more than happy to walk them through what they'd need to \n> change. I'm just not horribly motivated to do it myself. Hint, hint.\n\nIn case anybody is wondering, I share the same feeling.  I\ncannot say I'd be \"more than happy to\" clean up potential\nbreakages during the development of such changes, but if the\nchange eventually would help certain use cases, I can be\npersuaded to help debugging such a mess ;-).\n"},{"id":"16173","messageId":"43F27878.50701@vilain.net","threadId":"3265","inReplyTo":"7vy80dpo9g.fsf@assigned-by-dhcp.cox.net","subject":"Re: Handling large files with GIT","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2006-02-15T00:40:24Z","receivedAt":"2006-02-15T00:40:24Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Junio C Hamano wrote:\n> Linus Torvalds <torvalds@osdl.org> writes:\n> \n>>If somebody is interested in making the \"lots of filename changes\" case go \n>>fast, I'd be more than happy to walk them through what they'd need to \n>>change. I'm just not horribly motivated to do it myself. Hint, hint.\n> \n> In case anybody is wondering, I share the same feeling.  I\n> cannot say I'd be \"more than happy to\" clean up potential\n> breakages during the development of such changes, but if the\n> change eventually would help certain use cases, I can be\n> persuaded to help debugging such a mess ;-).\n\nExcellent.  Any speculations on where they might fit?  Clearly, it needs\nto be out of the \"tree\".\n\nDealing with the three cases I mentioned before in my Warnocked post;\n\n   1. caching - I'll consider this an \"under the hood\" thing, it really\n                doesn't matter, so long as the tools all know.\n\n   2. forensic - extra stuff at the end of the commit object?\n\n      eg\n         Copied: /new/path from /old/path:commit:c0bb171d..\n           (for SVN case where history matters)\n         Copied: /new/path from blob:b10b1d..\n           (for general pre-caching case)\n         Merged: /new/path from /old/path:commit:C0bb171d..\n           (for an SVK clone, so we know that subsequent merges on\n            /new/path need only merge from /old/path starting at commit\n            C0bb171d..)\n\n   3. retrospective - as above, but allow to specify old versions.\n\n      eg\n         Copied: /new/path:C0bb171d1 from /old/path:commit:c0bb171d2...\n           (for SVN case where history matters)\n\nMartin, is that enough for your CVS case?\n\nSam.\n"},{"id":"16177","messageId":"7vslqlo0wo.fsf@assigned-by-dhcp.cox.net","threadId":"3265","inReplyTo":"43F27878.50701@vilain.net","subject":"Re: Handling large files with GIT","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-15T01:39:51Z","receivedAt":"2006-02-15T01:39:51Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Sam Vilain <sam@vilain.net> writes:\n\n> ...  Clearly, it needs to be out of the \"tree\".\n\nOK.\n\n>   2. forensic - extra stuff at the end of the commit object?\n\n(except \"extra at the end of commit\", which does not make it out\nof the tree).\n\n>      eg\n>         Copied: /new/path from /old/path:commit:c0bb171d..\n>           (for SVN case where history matters)\n>         Copied: /new/path from blob:b10b1d..\n>           (for general pre-caching case)\n>         Merged: /new/path from /old/path:commit:C0bb171d..\n>           (for an SVK clone, so we know that subsequent merges on\n>            /new/path need only merge from /old/path starting at commit\n>            C0bb171d..)\n\nI am not sure if recording the bare SVN ``copied'' is very\nuseful.  You would need to infer things from what SVN did to\ntell if the copy is a tree copy inside a project (e.g. cp -r\ni386 x86_64), tagging (e.g. svn-cp rHEAD trunk tags/v1.2), or\nbranching, wouldn't you?  SVK merge ticket is a bit more useful\nin that sense.\n\nSo far, git philosophy is to record things you _know_ about and\ndefer such guesswork to the future, so limiting what you record\nto what you can actually see from the foreign SCM would be more\nin line with it.  For the same reason, if you are talking about\nmaildir managed under git, you should not have record anything\nother than what git already records: \"we used to have these\nfiles, now we have these instead\".\n\nBut I thought you were talking about caching what earlier\ninference declared what happened, so that you do not have to do\nthe same inference every time.  If that is the case, SVN level\n\"Copied:\" is probably not what you would want to record, I\nsuspect.  You would do some inference with the given information\n(\"SVN says it copied this tree to that tree, what was it that it\nreally wanted to do?  Was it a copy, or was it to create a\nbranch which was implemented as a copy?\"), and record that,\nhoping that information would help your other operations this\ntime and later.\n\nSo I think the order of questions you should be asking is:\n\n   - what operations are you trying to help?\n\n   - what information you would need to achieve those operations\n     better?\n\n   - among the second one, what will be necessary to be set in\n     stone (IOW, cannot be computed later), and what are\n     computable but expensive to recompute every time?\n\nAn example from an ancient thread.\n\nWith criss-cross merge between renamed trees, it was conjectured\nthat recording renames detected earlier would help later merges.\nI think you should arrive at the list of \"what we should record\"\nby thinking things in this order:\n\n (1) currently criss-cross merge between renamed trees does not\n     work well (realization of the status quo);\n\n (2) if we had this kind of information it would work better,\n     here are the things we need to record when a new commit is\n     made, and here is how to compute other information that can\n     be inferred, and here is how to use that information to\n     make the merge work better (solution without caching);\n\n (3) but it is expensive to recompute information we said\n     computable in (2) if we were to do so every time.  Let's\n     cache it.\n\nI am getting an impression that you are doing only the first\nhalf of (2) without other parts, which somewhat bothers me.\n"},{"id":"16178","messageId":"Pine.LNX.4.64.0602141741210.3691@g5.osdl.org","threadId":"3265","inReplyTo":"7vy80dpo9g.fsf@assigned-by-dhcp.cox.net","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-15T02:05:30Z","receivedAt":"2006-02-15T02:05:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 14 Feb 2006, Junio C Hamano wrote:\n\n> Linus Torvalds <torvalds@osdl.org> writes:\n> \n> > If somebody is interested in making the \"lots of filename changes\" case go \n> > fast, I'd be more than happy to walk them through what they'd need to \n> > change. I'm just not horribly motivated to do it myself. Hint, hint.\n> \n> In case anybody is wondering, I share the same feeling.  I\n> cannot say I'd be \"more than happy to\" clean up potential\n> breakages during the development of such changes, but if the\n> change eventually would help certain use cases, I can be\n> persuaded to help debugging such a mess ;-).\n\nActually, I got interested in seeing how hard this is, and wrote a simple \nfirst cut at doing a tree-optimized merger.\n\nLet me shout a bit first:\n\n  THIS IS WORKING CODE, BUT BE CAREFUL: IT'S A TECHNOLOGY DEMONSTRATION \n  RATHER THAN THE FINAL PRODUCT!\n\nWith that out of the way, let me descibe what this does (and then describe \nthe missing parts).\n\nThis is basically a three-way merge that works entirely on the \"tree\" \nlevel, rather than on the index. A lot of the _concepts_ are the same, \nthough, and if you're familiar with the results of an index merge, some of \nthe output will make more sense.\n\nYou give it three trees: the base tree (tree 0), and the two branches to \nbe merged (tree 1 and tree 2 respectively). It will then walk these three \ntrees, and resolve them as it goes along.\n\nThe interesting part is:\n - it can resolve whole sub-directories in one go, without actually even \n   looking recursively at them. A whole subdirectory will resolve the same \n   way as any individual files will (although that may need some \n   modification, see later).\n - if it has a \"content conflict\", for subdirectories that means \"try to \n   do a recursive tree merge\", while for non-subdirectories it's just a \n   content conflict and we'll output the stage 1/2/3 information.\n - a successful merge will output a single stage 0 (\"merged\") entry, \n   potentially for a whole subdirectory.\n - it outputs all the resolve information on stdout, so something like the \n   recursive resolver can pretty easily parse it all.\n\nNow, the caveats:\n - we probably need to be more careful about subdirectory resolves. The \n   trivial case (both branches have the exact same subdirectory) is a \n   trivial resolve, but the other cases (\"branch1 matches base, branch2 is \n   different\" probably can't be silently just resolved to the \"branch2\" \n   subdirectory state, since it might involve renames into - or out of - \n   that subdirectory)\n - we do not track the current index file at all, so this does not do the \n   \"check that index matches branch1\" logic that the three-way merge in \n   git-read-tree does. The theory is that we'd do a full three-way merge \n   (ignoring the index and working directory), and then to update the \n   working tree, we'd do a two-way \"git-read-tree branch1->result\"\n - I didn't actually make it do all the trivial resolve cases that \n   git-read-tree does. It's a technology demonstration.\n\nFinally (a more serious caveat):\n - doing things through stdout may end up being so expensive that we'd \n   need to do something else. In particular, it's likely that I should \n   not actually output the \"merge results\", but instead output a \"merge \n   results as they _differ_ from branch1\"\n\nHowever, I think this patch is already interesting enough that people who \nare interested in merging trees might want to look at it. Please keep in \nmind that tech _demo_ part, and in particular, keep in mind the final \n\"serious caveat\" part.\n\nIn many ways, the really _interesting_ part of a merge is not the result, \nbut how it _changes_ the branch we're merging into. That's particularly \nimportant as it should hopefully also mean that the output size for any \nreasonable case is minimal (and tracks what we actually need to do to the \ncurrent state to create the final result).\n\nThe code very much is organized so that doing the result as a \"diff \nagainst branch1\" should be quite easy/possible. I was actually going to do \nit, but I decided that it probably makes the output harder to read. I \ndunno.\n\nAnyway, let's think about this kind of approach.. Note how the code itself \nis actually quite small and short, although it's prbably pretty \"dense\".\n\nAs an interesting test-case, I'd suggest this merge in the kernel:\n\n\tgit-merge-tree $(git-merge-base 4cbf876 7d2babc) 4cbf876 7d2babc\n\nwhich resolves beautifully (there are no actual file-level conflicts), and \nyou can look at the output of that command to start thinking about what \nit does.\n\nThe interesting part (perhaps) is that timing that command for me shows \nthat it takes all of 0.004 seconds.. (the git-merge-base thing takes \nconsiderably more ;)\n\nThe point is, we _can_ do the actual merge part really really quickly. \n\n\t\tLinus\n\nPS. Final note: when I say that it is \"WORKING CODE\", that is obviously by \nmy standards. IOW, I tested it once and it gave reasonable results - so it \nmust be perfect.\n\nWhether it works for anybody else, or indeed for any other test-case, is \nnot my problem ;)\n\n---\ndiff-tree f0e6b454ff873429237322c846603d2e1fffc867 (from 6a9b87972f27edfe53da4ce016adf4c0cd42f5e6)\nAuthor: Linus Torvalds <torvalds@osdl.org>\nDate:   Tue Feb 14 17:39:15 2006 -0800\n\n    Add \"git-merge-tree\" functionality\n    \n    This is basically a tree-optimized merge.  Or rather, it is the first\n    stages _towards_ such a merge.\n    \n    Given a base tree and two branches to merge, it will do a trivial merge,\n    optimizing away the case of identical subdirectories, and resolving\n    trivial merges.  It outputs the list of file/directory resolves.\n    \n    Signed-off-by: Linus Torvalds <torvalds@osdl.org>\n\ndiff --git a/Makefile b/Makefile\nindex d40aa6a..4d04f49 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -151,7 +151,7 @@ PROGRAMS = \\\n \tgit-upload-pack$X git-verify-pack$X git-write-tree$X \\\n \tgit-update-ref$X git-symbolic-ref$X git-check-ref-format$X \\\n \tgit-name-rev$X git-pack-redundant$X git-repo-config$X git-var$X \\\n-\tgit-describe$X\n+\tgit-describe$X git-merge-tree$X\n \n # what 'all' will build and 'install' will install, in gitexecdir\n ALL_PROGRAMS = $(PROGRAMS) $(SIMPLE_PROGRAMS) $(SCRIPTS)\ndiff --git a/merge-tree.c b/merge-tree.c\nnew file mode 100644\nindex 0000000..0d6d434\n--- /dev/null\n+++ b/merge-tree.c\n@@ -0,0 +1,238 @@\n+#include \"cache.h\"\n+#include \"diff.h\"\n+\n+static const char merge_tree_usage[] = \"git-merge-tree <base-tree> <branch1> <branch2>\";\n+static int resolve_directories = 1;\n+\n+static void merge_trees(struct tree_desc t[3], const char *base);\n+\n+static void *fill_tree_descriptor(struct tree_desc *desc, const unsigned char *sha1)\n+{\n+\tunsigned long size = 0;\n+\tvoid *buf = NULL;\n+\n+\tif (sha1) {\n+\t\tbuf = read_object_with_reference(sha1, \"tree\", &size, NULL);\n+\t\tif (!buf)\n+\t\t\tdie(\"unable to read tree %s\", sha1_to_hex(sha1));\n+\t}\n+\tdesc->size = size;\n+\tdesc->buf = buf;\n+\treturn buf;\n+}\n+\n+struct name_entry {\n+\tconst unsigned char *sha1;\n+\tconst char *path;\n+\tunsigned int mode;\n+\tint pathlen;\n+};\n+\n+static void entry_clear(struct name_entry *a)\n+{\n+\tmemset(a, 0, sizeof(*a));\n+}\n+\n+static int entry_compare(struct name_entry *a, struct name_entry *b)\n+{\n+\treturn base_name_compare(\n+\t\t\ta->path, a->pathlen, a->mode,\n+\t\t\tb->path, b->pathlen, b->mode);\n+}\n+\n+static void entry_extract(struct tree_desc *t, struct name_entry *a)\n+{\n+\ta->sha1 = tree_entry_extract(t, &a->path, &a->mode);\n+\ta->pathlen = strlen(a->path);\n+}\n+\n+/* An empty entry never compares same, not even to another empty entry */\n+static int same_entry(struct name_entry *a, struct name_entry *b)\n+{\n+\treturn\ta->sha1 &&\n+\t\tb->sha1 &&\n+\t\t!memcmp(a->sha1, b->sha1, 20) &&\n+\t\ta->mode == b->mode;\n+}\n+\n+static void resolve(const char *base, struct name_entry *result)\n+{\n+\tprintf(\"0 %06o %s %s%s\\n\", result->mode, sha1_to_hex(result->sha1), base, result->path);\n+}\n+\n+static int unresolved_directory(const char *base, struct name_entry n[3])\n+{\n+\tint baselen;\n+\tchar *newbase;\n+\tstruct name_entry *p;\n+\tstruct tree_desc t[3];\n+\tvoid *buf0, *buf1, *buf2;\n+\n+\tif (!resolve_directories)\n+\t\treturn 0;\n+\tp = n;\n+\tif (!p->mode) {\n+\t\tp++;\n+\t\tif (!p->mode)\n+\t\t\tp++;\n+\t}\n+\tif (!S_ISDIR(p->mode))\n+\t\treturn 0;\n+\tbaselen = strlen(base);\n+\tnewbase = xmalloc(baselen + p->pathlen + 2);\n+\tmemcpy(newbase, base, baselen);\n+\tmemcpy(newbase + baselen, p->path, p->pathlen);\n+\tmemcpy(newbase + baselen + p->pathlen, \"/\", 2);\n+\n+\tbuf0 = fill_tree_descriptor(t+0, n[0].sha1);\n+\tbuf1 = fill_tree_descriptor(t+1, n[1].sha1);\n+\tbuf2 = fill_tree_descriptor(t+2, n[2].sha1);\n+\tmerge_trees(t, newbase);\n+\n+\tfree(buf0);\n+\tfree(buf1);\n+\tfree(buf2);\n+\tfree(newbase);\n+\treturn 1;\n+}\n+\n+static void unresolved(const char *base, struct name_entry n[3])\n+{\n+\tif (unresolved_directory(base, n))\n+\t\treturn;\n+\tprintf(\"1 %06o %s %s%s\\n\", n[0].mode, sha1_to_hex(n[0].sha1), base, n[0].path);\n+\tprintf(\"2 %06o %s %s%s\\n\", n[1].mode, sha1_to_hex(n[1].sha1), base, n[1].path);\n+\tprintf(\"3 %06o %s %s%s\\n\", n[2].mode, sha1_to_hex(n[2].sha1), base, n[2].path);\n+}\n+\n+/*\n+ * Merge two trees together (t[1] and t[2]), using a common base (t[0])\n+ * as the origin.\n+ *\n+ * This walks the (sorted) trees in lock-step, checking every possible\n+ * name. Note that directories automatically sort differently from other\n+ * files (see \"base_name_compare\"), so you'll never see file/directory\n+ * conflicts, because they won't ever compare the same.\n+ *\n+ * IOW, if a directory changes to a filename, it will automatically be\n+ * seen as the directory going away, and the filename being created.\n+ *\n+ * Think of this as a three-way diff.\n+ *\n+ * The output will be either:\n+ *  - successful merge\n+ *\t \"0 mode sha1 filename\"\n+ *    NOTE NOTE NOTE! FIXME! We really really need to walk the index\n+ *    in parallel with this too!\n+ * \n+ *  - conflict:\n+ *\t\"1 mode sha1 filename\"\n+ *\t\"2 mode sha1 filename\"\n+ *\t\"3 mode sha1 filename\"\n+ *    where not all of the 1/2/3 lines may exist, of course.\n+ *\n+ * The successful merge rules are the same as for the three-way merge\n+ * in git-read-tree.\n+ */\n+static void merge_trees(struct tree_desc t[3], const char *base)\n+{\n+\tfor (;;) {\n+\t\tstruct name_entry entry[3];\n+\t\tunsigned int mask = 0;\n+\t\tint i, last;\n+\n+\t\tlast = -1;\n+\t\tfor (i = 0; i < 3; i++) {\n+\t\t\tif (!t[i].size)\n+\t\t\t\tcontinue;\n+\t\t\tentry_extract(t+i, entry+i);\n+\t\t\tif (last >= 0) {\n+\t\t\t\tint cmp = entry_compare(entry+i, entry+last);\n+\n+\t\t\t\t/*\n+\t\t\t\t * Is the new name bigger than the old one?\n+\t\t\t\t * Ignore it\n+\t\t\t\t */\n+\t\t\t\tif (cmp > 0)\n+\t\t\t\t\tcontinue;\n+\t\t\t\t/*\n+\t\t\t\t * Is the new name smaller than the old one?\n+\t\t\t\t * Ignore all old ones\n+\t\t\t\t */\n+\t\t\t\tif (cmp < 0)\n+\t\t\t\t\tmask = 0;\n+\t\t\t}\n+\t\t\tmask |= 1u << i;\n+\t\t\tlast = i;\n+\t\t}\n+\t\tif (!mask)\n+\t\t\tbreak;\n+\n+\t\t/*\n+\t\t * Update the tree entries we've walked, and clear\n+\t\t * all the unused name-entries.\n+\t\t */\n+\t\tfor (i = 0; i < 3; i++) {\n+\t\t\tif (mask & (1u << i)) {\n+\t\t\t\tupdate_tree_entry(t+i);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tentry_clear(entry + i);\n+\t\t}\n+\n+\t\t/* Same in both? */\n+\t\tif (same_entry(entry+1, entry+2)) {\n+\t\t\tif (entry[0].sha1) {\n+\t\t\t\tresolve(base, entry+1);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\n+\t\tif (same_entry(entry+0, entry+1)) {\n+\t\t\tif (entry[2].sha1) {\n+\t\t\t\tresolve(base, entry+2);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\n+\t\tif (same_entry(entry+0, entry+2)) {\n+\t\t\tif (entry[1].sha1) {\n+\t\t\t\tresolve(base, entry+1);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\n+\t\tunresolved(base, entry);\n+\t}\n+}\n+\n+static void *get_tree_descriptor(struct tree_desc *desc, const char *rev)\n+{\n+\tunsigned char sha1[20];\n+\tvoid *buf;\n+\n+\tif (get_sha1(rev, sha1) < 0)\n+\t\tdie(\"unknown rev %s\", rev);\n+\tbuf = fill_tree_descriptor(desc, sha1);\n+\tif (!buf)\n+\t\tdie(\"%s is not a tree\", rev);\n+\treturn buf;\n+}\n+\n+int main(int argc, char **argv)\n+{\n+\tstruct tree_desc t[3];\n+\tvoid *buf1, *buf2, *buf3;\n+\n+\tif (argc < 4)\n+\t\tusage(merge_tree_usage);\n+\n+\tbuf1 = get_tree_descriptor(t+0, argv[1]);\n+\tbuf2 = get_tree_descriptor(t+1, argv[2]);\n+\tbuf3 = get_tree_descriptor(t+2, argv[3]);\n+\tmerge_trees(t, \"\");\n+\tfree(buf1);\n+\tfree(buf2);\n+\tfree(buf3);\n+\treturn 0;\n+}\n"},{"id":"16179","messageId":"46a038f90602141807s468c421dm5c4b68cfcf87e7@mail.gmail.com","threadId":"3265","inReplyTo":"43F27878.50701@vilain.net","subject":"Re: Handling large files with GIT","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-02-15T02:07:52Z","receivedAt":"2006-02-15T02:07:52Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 2/15/06, Sam Vilain <sam@vilain.net> wrote:\n> Excellent.  Any speculations on where they might fit?  Clearly, it needs\n> to be out of the \"tree\".\n\nI think Junio & Linus are talking about alternative mergers, something\nthat can be called instead of git-read-tree -m (which is the way\nmerges seem to kick off). Or perhaps an additional flag to\ngit-read-tree to be used in conjunction with -m, something like\n--optimize-for-identity that lets git-read-tree know to do a first\npass keying things on file identity rather than file path.\n\nSo we are _not_ touching the object database, at all. Only optimising\nmerges for very large trees there mv is a popular operation. All the\ncases you discuss can be tackled very efficiently without making *any*\nchange to the object database.\n\n> Martin, is that enough for your CVS case?\n\nOh, I don't need it at all. It's just that there's been some lazy talk\nof tracking mboxes and maildirs with git, and look where it's led.\nBlame Roland Stigge who got me started down this track.\n\nI'm sure it's because the other optimisations are a lot harder to\ntackle ;-) though Linus mentions that it'd be trivial for\ngit-read-tree -m to detect unchanged directories and perhaps do things\na bit faster. Not as revolutionary as an --optimize-for-identity but\nnot as risky either.\n\nIn any case, don't count in me for any of this git-checkout hacking. I\nknow better than start learning C posting patches to *this* list.\n\ncheers,\n\n\nmartin\n"},{"id":"16180","messageId":"Pine.LNX.4.64.0602141811050.3691@g5.osdl.org","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602141741210.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-15T02:18:21Z","receivedAt":"2006-02-15T02:18:21Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 14 Feb 2006, Linus Torvalds wrote:\n> \n> Finally (a more serious caveat):\n>  - doing things through stdout may end up being so expensive that we'd \n>    need to do something else. In particular, it's likely that I should \n>    not actually output the \"merge results\", but instead output a \"merge \n>    results as they _differ_ from branch1\"\n> \n> In many ways, the really _interesting_ part of a merge is not the result, \n> but how it _changes_ the branch we're merging into. That's particularly \n> important as it should hopefully also mean that the output size for any \n> reasonable case is minimal (and tracks what we actually need to do to the \n> current state to create the final result).\n\nHere, btw, is the trivial diff to turn my previous \"tree-resolve\" into a \n\"resolve tree relative to the current branch\".\n\nIn particular, it makes the example merge perhaps even more interesting, \nand makes the \"merging directories and merging files should use different \nheuristics more obvious\". It's quite instructive, I think.\n\nSo if you want to test this, the merge I have been testing with is the \nlast infiniband merge in the kernel:\n\n\tgit-merge-tree 3c3b809 4cbf876 7d2babc\n\nand you'll need to spend a few moments on thinking about what the \n\"directory merge\" thing there means: in particular, we should probably \nmake the\n\n\tif (entry[2].sha1) {\n\ntest be\n\n\tif (entry[2].sha && !S_ISDIR(entry[2].mode)) {\n\n(and same for \"resolve to entry[1]\" case for that matter) so that we never \ncreate a \"resolve()\" that picks a whole subdirectory from one of the \nbranches.\n\nThe current logic is \"logical\", just probably not what we want.\n\n\t\tLinus\n\n----\ndiff --git a/merge-tree.c b/merge-tree.c\nindex 0d6d434..0bf871c 100644\n--- a/merge-tree.c\n+++ b/merge-tree.c\n@@ -55,9 +55,19 @@ static int same_entry(struct name_entry \n \t\ta->mode == b->mode;\n }\n \n-static void resolve(const char *base, struct name_entry *result)\n+static void resolve(const char *base, struct name_entry *branch1, struct name_entry *result)\n {\n-\tprintf(\"0 %06o %s %s%s\\n\", result->mode, sha1_to_hex(result->sha1), base, result->path);\n+\tchar branch1_sha1[50];\n+\n+\t/* If it's already branch1, don't bother showing it */\n+\tif (!branch1)\n+\t\treturn;\n+\tmemcpy(branch1_sha1, sha1_to_hex(branch1->sha1), 41);\n+\n+\tprintf(\"0 %06o->%06o %s->%s %s%s\\n\",\n+\t\tbranch1->mode, result->mode,\n+\t\tbranch1_sha1, sha1_to_hex(result->sha1),\n+\t\tbase, result->path);\n }\n \n static int unresolved_directory(const char *base, struct name_entry n[3])\n@@ -183,21 +193,21 @@ static void merge_trees(struct tree_desc\n \t\t/* Same in both? */\n \t\tif (same_entry(entry+1, entry+2)) {\n \t\t\tif (entry[0].sha1) {\n-\t\t\t\tresolve(base, entry+1);\n+\t\t\t\tresolve(base, NULL, entry+1);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t}\n \n \t\tif (same_entry(entry+0, entry+1)) {\n \t\t\tif (entry[2].sha1) {\n-\t\t\t\tresolve(base, entry+2);\n+\t\t\t\tresolve(base, entry+1, entry+2);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t}\n \n \t\tif (same_entry(entry+0, entry+2)) {\n \t\t\tif (entry[1].sha1) {\n-\t\t\t\tresolve(base, entry+1);\n+\t\t\t\tresolve(base, NULL, entry+1);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t}\n"},{"id":"16181","messageId":"Pine.LNX.4.64.0602141829080.3691@g5.osdl.org","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602141811050.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-15T02:33:02Z","receivedAt":"2006-02-15T02:33:02Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 14 Feb 2006, Linus Torvalds wrote:\n> \n> Here, btw, is the trivial diff to turn my previous \"tree-resolve\" into a \n> \"resolve tree relative to the current branch\".\n\nGaah. It was trivial, and it happened to work fine for my test-case, but \nwhen I started looking at not doing that extremely aggressive subdirectory \nmerging, that showed a few other issues...\n\nSo in case people want to try, here's a third patch. Oh, and it's against \nmy _original_ path, not incremental to the middle one (ie both patches two \nand three are against patch #1, it's not a nice series).\n\nNow I'm really done, and won't be sending out any more patches today. \nSorry for the noise.\n\n\t\tLinus\n----\ndiff --git a/merge-tree.c b/merge-tree.c\nindex 0d6d434..6381118 100644\n--- a/merge-tree.c\n+++ b/merge-tree.c\n@@ -55,9 +55,26 @@ static int same_entry(struct name_entry \n \t\ta->mode == b->mode;\n }\n \n-static void resolve(const char *base, struct name_entry *result)\n+static const char *sha1_to_hex_zero(const unsigned char *sha1)\n {\n-\tprintf(\"0 %06o %s %s%s\\n\", result->mode, sha1_to_hex(result->sha1), base, result->path);\n+\tif (sha1)\n+\t\treturn sha1_to_hex(sha1);\n+\treturn \"0000000000000000000000000000000000000000\";\n+}\n+\n+static void resolve(const char *base, struct name_entry *branch1, struct name_entry *result)\n+{\n+\tchar branch1_sha1[50];\n+\n+\t/* If it's already branch1, don't bother showing it */\n+\tif (!branch1)\n+\t\treturn;\n+\tmemcpy(branch1_sha1, sha1_to_hex_zero(branch1->sha1), 41);\n+\n+\tprintf(\"0 %06o->%06o %s->%s %s%s\\n\",\n+\t\tbranch1->mode, result->mode,\n+\t\tbranch1_sha1, sha1_to_hex_zero(result->sha1),\n+\t\tbase, result->path);\n }\n \n static int unresolved_directory(const char *base, struct name_entry n[3])\n@@ -100,9 +117,12 @@ static void unresolved(const char *base,\n {\n \tif (unresolved_directory(base, n))\n \t\treturn;\n-\tprintf(\"1 %06o %s %s%s\\n\", n[0].mode, sha1_to_hex(n[0].sha1), base, n[0].path);\n-\tprintf(\"2 %06o %s %s%s\\n\", n[1].mode, sha1_to_hex(n[1].sha1), base, n[1].path);\n-\tprintf(\"3 %06o %s %s%s\\n\", n[2].mode, sha1_to_hex(n[2].sha1), base, n[2].path);\n+\tif (n[0].sha1)\n+\t\tprintf(\"1 %06o %s %s%s\\n\", n[0].mode, sha1_to_hex(n[0].sha1), base, n[0].path);\n+\tif (n[1].sha1)\n+\t\tprintf(\"2 %06o %s %s%s\\n\", n[1].mode, sha1_to_hex(n[1].sha1), base, n[1].path);\n+\tif (n[2].sha1)\n+\t\tprintf(\"3 %06o %s %s%s\\n\", n[2].mode, sha1_to_hex(n[2].sha1), base, n[2].path);\n }\n \n /*\n@@ -183,21 +203,21 @@ static void merge_trees(struct tree_desc\n \t\t/* Same in both? */\n \t\tif (same_entry(entry+1, entry+2)) {\n \t\t\tif (entry[0].sha1) {\n-\t\t\t\tresolve(base, entry+1);\n+\t\t\t\tresolve(base, NULL, entry+1);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t}\n \n \t\tif (same_entry(entry+0, entry+1)) {\n-\t\t\tif (entry[2].sha1) {\n-\t\t\t\tresolve(base, entry+2);\n+\t\t\tif (entry[2].sha1 && !S_ISDIR(entry[2].mode)) {\n+\t\t\t\tresolve(base, entry+1, entry+2);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t}\n \n \t\tif (same_entry(entry+0, entry+2)) {\n-\t\t\tif (entry[1].sha1) {\n-\t\t\t\tresolve(base, entry+1);\n+\t\t\tif (entry[1].sha1 && !S_ISDIR(entry[1].mode)) {\n+\t\t\t\tresolve(base, NULL, entry+1);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t}\n"},{"id":"16182","messageId":"Pine.LNX.4.64.0602141953081.3691@g5.osdl.org","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602141829080.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-15T03:58:03Z","receivedAt":"2006-02-15T03:58:03Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 14 Feb 2006, Linus Torvalds wrote:\n> \n> So in case people want to try, here's a third patch. Oh, and it's against \n> my _original_ path, not incremental to the middle one (ie both patches two \n> and three are against patch #1, it's not a nice series).\n> \n> Now I'm really done, and won't be sending out any more patches today. \n\nStill true. I've just been thinking about the last state.\n\nAs far as I can tell, the output from git-merge-tree with that fix to only \nsimplify subdirectories that match exactly in all of base/branch1/branch2 \nis precisely the output that git-merge-recursive actually wants.\n\nRather than doing a three-way merge with \"git-read-tree\", and then doing \n\"git-ls-files --unmerged\", I think this gives the same result much more \nefficiently.\n\nThat said, I can't follow the python code, so maybe I'm missing something. \nFredrik cc'd, in case he can put me right.\n\n\t\tLinus\n"},{"id":"16183","messageId":"43F2A828.2050102@vilain.net","threadId":"3265","inReplyTo":"7vslqlo0wo.fsf@assigned-by-dhcp.cox.net","subject":"Re: Handling large files with GIT","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2006-02-15T04:03:52Z","receivedAt":"2006-02-15T04:03:52Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Junio C Hamano wrote:\n> So I think the order of questions you should be asking is:\n> \n>   1 - what operations are you trying to help?\n\nPrimarily, tracing history when dealing with history/changeset based\nrevision systems like SVN or darcs, and doing this in a manner that we\ncan make guarantees about behaving in the same way as those systems\nwould.\n\n>   2 - what information you would need to achieve those operations\n>       better?\n\nMinimally, this tuple:\n\n   ( merge|copy, source_path, source_tree|source_commit,\n     destination_path, destination_commit )\n\nIt makes sense to record this with commits, as conceptually it is a part\nof the intended commit history along with the change comment.\n\n>   3 - among the second one, what will be necessary to be set in\n>       stone (IOW, cannot be computed later), and what are\n>       computable but expensive to recompute every time?\n\nThe only operation you cannot automatically and with certainty detect a \nrename and change content without inserting a dummy commit between the \nname change and the content change.  But in a sense this is the same as\nmy suggestion - using the commit object history to record information\nthat normally doesn't matter when you are doing content-keyed\noperations.\n\n> I am getting an impression that you are doing only the first\n> half of (2) without other parts, which somewhat bothers me.\n\nWell, thank you for spending so much time to reply to me given that was\nyour assessment.  I think the best direction from here would be to start\nmolding some porcelain, then I can cross this bridge when I come to it\nrather than simply speculating and hand-waving.\n\nBesides, I can always prototype it for discussion using the commit\ndescription as a surrogate container for the information.\n\nSam.\n\nps I also responded to the rest of your e-mail, but decided that the \nanswers to the above questions were more important.\n\n >>  2. forensic - extra stuff at the end of the commit object?\n > (except \"extra at the end of commit\", which does not make it out\n > of the tree).\n\nIt is a part of the repository, but more a property of the commit itself\n- like the commit description.  Like somebody writing \"I renamed this\nfile to that file and changed its contents\", but in a parsable form\nthat can _optionally_ be used to prevent the relevant git-core tools\nfrom having to do content comparison, or perhaps something subtler like\nincreasing the score of the recorded history branch when scoring\nalternatives looking for history.\n\n >>     eg\n >>        Copied: /new/path from /old/path:commit:c0bb171d..\n >>          (for SVN case where history matters)\n >>        Copied: /new/path from blob:b10b1d..\n >>          (for general pre-caching case)\n >>        Merged: /new/path from /old/path:commit:C0bb171d..\n >>          (for an SVK clone, so we know that subsequent merges on\n >>           /new/path need only merge from /old/path starting at commit\n >>           C0bb171d..)\n > I am not sure if recording the bare SVN ``copied'' is very\n > useful.  You would need to infer things from what SVN did to\n > tell if the copy is a tree copy inside a project (e.g. cp -r\n > i386 x86_64), tagging (e.g. svn-cp rHEAD trunk tags/v1.2), or\n > branching, wouldn't you?  SVK merge ticket is a bit more useful\n > in that sense.\n\nIn the SVN model there really is no difference between these cases.  Of\ncourse the actual representation of these in the object does not matter;\nthe above is the what, not the how.  But in general, SVN only records\ncopying; it has no repository concept of merge, branch, tag, rename.\nSVK adds merging to the picture.\n\nRepresenting an SVN tree copy as a new sub-tree in a git repository\nshould still be a \"cheap copy\", it's just that all the tools will not\n(and probably should not) see it as a branch but a copy.\n\n > So far, git philosophy is to record things you _know_ about and\n > defer such guesswork to the future, so limiting what you record\n > to what you can actually see from the foreign SCM would be more\n > in line with it.\n\nYes, and if I am mirroring an SVN repository, then I only know that in\nthat repository, the history /was recorded/ as such.  Not the history\n/is/ as such, that's a different question, and is the guesswork worth\nbeing defered to the future.\n\n > For the same reason, if you are talking about\n > maildir managed under git, you should not have record anything\n > other than what git already records: \"we used to have these\n > files, now we have these instead\".\n\nOk.  As Martin pointed out, the Maildir situation is actually a simple\ncase.  In a sense, I hijacked a vaguely related thread to resolve my\nWarnock dilemma :)\n\n > But I thought you were talking about caching what earlier\n > inference declared what happened, so that you do not have to do\n > the same inference every time.  If that is the case, SVN level\n > \"Copied:\" is probably not what you would want to record, I\n > suspect.  You would do some inference with the given information\n > (\"SVN says it copied this tree to that tree, what was it that it\n > really wanted to do?  Was it a copy, or was it to create a\n > branch which was implemented as a copy?\"), and record that,\n > hoping that information would help your other operations this\n > time and later.\n\nWell, this is already guesswork defered to the future that the\nSubversion authors inflict on the users of Subversion repositories.  If\nyou read the Subversion manual you will find recommendations to\nstudiously record this information and to use a standard repository\nlayout so that other people will understand what your copies were\nintended to be.\n"},{"id":"16204","messageId":"7vd5hpj6ab.fsf@assigned-by-dhcp.cox.net","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602141953081.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-15T09:54:52Z","receivedAt":"2006-02-15T09:54:52Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> As far as I can tell, the output from git-merge-tree with that fix to only \n> simplify subdirectories that match exactly in all of base/branch1/branch2 \n> is precisely the output that git-merge-recursive actually wants.\n\nThe matches the recollection I had last time I mucked with the\ncode.  Currently it is set up to do one path at a time in both\nindex and working tree, so it would not be a trivial rewrite,\nbut merge-tree based approach would speed things up quite a\nbit.\n\nI was thinking about implementing mergers as a pipeline:\n\n\tgit-merge-tree O A B |\n        git-merge-renaming A |\n        git-merge-aggressive A |\n        git-merge-filemerge\n\ngit-merge-tree (yours) does not do trivial collapsing, and\nproduce raw-diff from A.  git-merge-renaming reads it, finds\ncopied/renamed entries (maybe reusing parts of diffcore), and\nwrites out the results in the same format as merge-tree output\n(that's why I am giving A on the command line -- so it can also\nread A if it wanted to. it may need to talk about what a path in\nA was even when merge-tree did not say anything about that\npath).  Then git-merge-aggressive (bad naming, I know, it only\ncorresponds to the flag of the same name in read-tree) will\ncollapse git-merge-one-file equivalent stage collapsing.  The\nremainder is fed to file-level merger for postprocessing.\nEverything except the last step would work on a data format that\nmerge-tree outputs.\n"},{"id":"16219","messageId":"Pine.LNX.4.64.0602150715470.3691@g5.osdl.org","threadId":"3265","inReplyTo":"7vd5hpj6ab.fsf@assigned-by-dhcp.cox.net","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-15T15:44:08Z","receivedAt":"2006-02-15T15:44:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 15 Feb 2006, Junio C Hamano wrote:\n> \n> I was thinking about implementing mergers as a pipeline:\n> \n> \tgit-merge-tree O A B |\n>         git-merge-renaming A |\n>         git-merge-aggressive A |\n>         git-merge-filemerge\n\nGreat minds think alike.\n\n> git-merge-tree (yours) does not do trivial collapsing, and\n> produce raw-diff from A.\n\n(It does _truly_ trivial collapsing, but I think we both agree: it doesn't \ndo anything that we used to go git-merge-one-file on)\n\n> git-merge-renaming reads it, finds\n> copied/renamed entries (maybe reusing parts of diffcore), and\n> writes out the results in the same format as merge-tree output\n\nI was considering perhaps doing a first cut at that in git-merge-tree \nalready. Not sure.\n\nOne issue is that I think I may have to change the output format if I do \nthat. I should anyway. \n\nWhy?\n\nIt's hard to see where \"one event\" stops, and another starts. I stupidly \ninitially thought that you can do it entirely based on looking at the \nnumbers, but you can't. Right now you have to look at the pathname too, \nwhich is kind of sad, and doesn't work after rename detection (since then \nthe pathnames won't be sorted any more, and one \"event\" can have different \npathnames in different stages).\n\n[ Side note: it doesn't even work for file/directory conflicts, which can \n  have the same name, but are two different \"events\". So you'd actually \n  have to look at both mode _and_ filename to sort out if two lines that \n  start with \"1\" and \"3\" respectively are one event (removal in first \n  branch) or two events (\"1\" on one file: removal in both branches + \"3\" \n  on another file: add in second branch) ]\n\nSo to do the rename output, you can't use the same format as merge-tree \nuses right _now_. We'd have to add a marker to mark what the event \nboundaries are.\n\nThe \"mark\" could be a running \"event number\", or even as easy as an \nalternating character (\"#\" vs \" \" as the second character in the line or \nsimilar)\n\nSo instead of\n\n\t2 100644 ff280e2e1613e808e4d7844376134dfa2bb1fc21 Documentation/cputopology.txt\n\t2 100644 28c5b7d1eb90f0ccd8e0307c170f89bd7954dc9c Documentation/hwmon/f71805f\n\t1 100644 b88953dfd58022aef1680c266c7438605b146fc8 Documentation/i2c/busses/i2c-sis69x\n\t3 100644 b88953dfd58022aef1680c266c7438605b146fc8 Documentation/i2c/busses/i2c-sis69x\n\t2 100644 00a009b977e92b1a942d1138afdccf1b725df956 Documentation/i2c/busses/i2c-sis96x\n\t2 100644 90a5e9e5bef1daa9d0f0621e209827f0d180f384 Documentation/unshare.txt\n\t2 100644 5127f39fa9bf9a384a6529c6d5deb1002e945de5 arch/arm/mach-s3c2410/s3c2400-gpio.c\n\t2 100644 8b2394e1ed4088c3b8d38e87e58bde2f38152bf7 arch/arm/mach-s3c2410/s3c2400.h\n\t ...\n\nit migth be\n\n\t2#100644 ff280e2e1613e808e4d7844376134dfa2bb1fc21 Documentation/cputopology.txt\n\t2 100644 28c5b7d1eb90f0ccd8e0307c170f89bd7954dc9c Documentation/hwmon/f71805f\n\t1#100644 b88953dfd58022aef1680c266c7438605b146fc8 Documentation/i2c/busses/i2c-sis69x\n\t3#100644 b88953dfd58022aef1680c266c7438605b146fc8 Documentation/i2c/busses/i2c-sis69x\n\t2 100644 00a009b977e92b1a942d1138afdccf1b725df956 Documentation/i2c/busses/i2c-sis96x\n\t2#100644 90a5e9e5bef1daa9d0f0621e209827f0d180f384 Documentation/unshare.txt\n\t2 100644 5127f39fa9bf9a384a6529c6d5deb1002e945de5 arch/arm/mach-s3c2410/s3c2400-gpio.c\n\t2#100644 8b2394e1ed4088c3b8d38e87e58bde2f38152bf7 arch/arm/mach-s3c2410/s3c2400.h\n\t ...\n\nwhere you can clearly see the \"grouping\" without having to even look at \nthe filename.\n\n(The example I show actually has a rename-with-modifications that was made \non the first branch: notice that i2c-sis69x vs i2c-sis96x thing?)\n\nI don't know exactly what the \"after rename detection\" output format would \nbe, but it _might_ turn that\n\n\t...\n\t1#b889... i2c-sis69x\n\t3#b889... i2c-sis69x\n\t2 00a0... i2c-sis96x\n\t...\n\ninto one event:\n\n\t...\n\t1#b889... i2c-sis69x\n\t2#00a0... i2c-sis96x\n\t3#b889... i2c-sis69x\n\t...\n\nand then the actual file-merge logic would have to merge the names as well \nas the file contents (and in this case, the final name would thus be \n\"i2c-sis96x\", since one branch hadn't changed it).\n\nHmm?\n\n\t\tLinus\n"},{"id":"16222","messageId":"Pine.LNX.4.64.0602150904310.3691@g5.osdl.org","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602150715470.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-15T17:16:21Z","receivedAt":"2006-02-15T17:16:21Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nBtw, some actual numbers: I did the recent kernel networking merge (which \nis a trivial in-index merge) with the standard three-way\n\n\tgit-read-tree -m <base> <branch> <branch>\n\nand with the new git-merge-tree to compare performance.\n\nDoing git-read-tree takes ~0.35s, while git-merge-tree took 0.015s.\n\nNow, that's not a really fair comparison, because the end result is very \ndifferent: the git-read-tree has populated the index, ready for a \ngit-writet-ree, while the git-merge-tree has not. \n\nHowever, the interesting part is that especially for a trivial merge, we \ndon't actually _want_ to necessarily populate the index, because doing a \n\"git-write-tree\" is actually a pretty expensive operation (on the kernel, \nit will try to write 1000+ directory trees, most of which already exist. \nAdmittedly we don't actually have to write the objects, since we figure \nout that they already exist, but we have to do the SHA1 calculations to \ndo so).\n\nSo if we made the git-merge-tree based merge work entirely on trees all \nthe way, and never even necessarily populate the index at all (unless it \nhas to, due to actual data conflicts that want to be fixed up), that would \nactually be another performance advantage. The only downside there is that \nwe would literally have to write the resulting tree objects by hand (ie \nwe'd need a new helper for doing that, and another thing to validate).\n\nAnyway, that should almost certainly make it possible to scale up git \nmerges to hundreds of thousands of files without huge performance problems \n(still, that depends a bit on layout - again, flat directory structures \nwon't scale as well, so it might not be enough for maildir handling).\n\nBut just at a guess, I think there's at least an order of magnitude to be \nhad there. So if a maildir merge currently takes an hour, at least we \nshould be able to get it down to a few minutes.\n\nBen, are you interested in trying this out in your maildir experiments?\n\n\t\tLinus\n"},{"id":"16238","messageId":"Pine.LNX.4.64.0602151915010.916@g5.osdl.org","threadId":"3265","inReplyTo":"7vd5hpj6ab.fsf@assigned-by-dhcp.cox.net","subject":"Re: Handling large files with GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-16T03:25:32Z","receivedAt":"2006-02-16T03:25:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nBtw, here's one last gasp on this thread: it generalizes the notion of \ntraversing several trees in sync, which could be used to do the n-way diff \nfor the \"-c\" and \"--cc\" style merge diffs a lot more efficiently.\n\nI didn't check, but I'm pretty sure that this would bring the cost of \ndoing the 12-way diff down to way under a second. Right now:\n\n\t[torvalds@g5 linux]$ time git-diff-tree -c 9fdb62a > /dev/null \n\n\treal    0m1.279s\n\tuser    0m1.272s\n\tsys     0m0.008s\n\nand that's a bit too much. We I'd really have expected us to be able to do \nbetter.\n\nIt should be possible to do this as a \n\n\ttraverse_trees(12, &trees, \"\", combined_diff_callback);\n\nfairly cheaply (and quickly throw away anything where any of the parents \nwas the same as the result).\n\nJunio, that \"traverse_trees()\" logic is totally independent of whether we \nactually do \"git-merge-tree\" or not, so if you want to, I could split up \nthe patches the other way (and merge \"traverse_trees()\" first as a new \ninterface, independently).\n\n\t\tLinus\n\n----\ngit-merge-tree: generalize the \"traverse <n> trees in sync\" functionality\n\nIt's actually very useful for other things too. Notably, we could do the\ncombined diff a lot more efficiently with this.\n\nSigned-off-by: Linus Torvalds <torvalds@osdl.org>\n\ndiff --git a/merge-tree.c b/merge-tree.c\nindex 6381118..2a9a013 100644\n--- a/merge-tree.c\n+++ b/merge-tree.c\n@@ -125,44 +125,19 @@ static void unresolved(const char *base,\n \t\tprintf(\"3 %06o %s %s%s\\n\", n[2].mode, sha1_to_hex(n[2].sha1), base, n[2].path);\n }\n \n-/*\n- * Merge two trees together (t[1] and t[2]), using a common base (t[0])\n- * as the origin.\n- *\n- * This walks the (sorted) trees in lock-step, checking every possible\n- * name. Note that directories automatically sort differently from other\n- * files (see \"base_name_compare\"), so you'll never see file/directory\n- * conflicts, because they won't ever compare the same.\n- *\n- * IOW, if a directory changes to a filename, it will automatically be\n- * seen as the directory going away, and the filename being created.\n- *\n- * Think of this as a three-way diff.\n- *\n- * The output will be either:\n- *  - successful merge\n- *\t \"0 mode sha1 filename\"\n- *    NOTE NOTE NOTE! FIXME! We really really need to walk the index\n- *    in parallel with this too!\n- * \n- *  - conflict:\n- *\t\"1 mode sha1 filename\"\n- *\t\"2 mode sha1 filename\"\n- *\t\"3 mode sha1 filename\"\n- *    where not all of the 1/2/3 lines may exist, of course.\n- *\n- * The successful merge rules are the same as for the three-way merge\n- * in git-read-tree.\n- */\n-static void merge_trees(struct tree_desc t[3], const char *base)\n+typedef void (*traverse_callback_t)(int n, unsigned long mask, struct name_entry *entry, const char *base);\n+\n+static void traverse_trees(int n, struct tree_desc *t, const char *base, traverse_callback_t callback)\n {\n+\tstruct name_entry *entry = xmalloc(n*sizeof(*entry));\n+\n \tfor (;;) {\n \t\tstruct name_entry entry[3];\n-\t\tunsigned int mask = 0;\n+\t\tunsigned long mask = 0;\n \t\tint i, last;\n \n \t\tlast = -1;\n-\t\tfor (i = 0; i < 3; i++) {\n+\t\tfor (i = 0; i < n; i++) {\n \t\t\tif (!t[i].size)\n \t\t\t\tcontinue;\n \t\t\tentry_extract(t+i, entry+i);\n@@ -182,7 +157,7 @@ static void merge_trees(struct tree_desc\n \t\t\t\tif (cmp < 0)\n \t\t\t\t\tmask = 0;\n \t\t\t}\n-\t\t\tmask |= 1u << i;\n+\t\t\tmask |= 1ul << i;\n \t\t\tlast = i;\n \t\t}\n \t\tif (!mask)\n@@ -192,38 +167,77 @@ static void merge_trees(struct tree_desc\n \t\t * Update the tree entries we've walked, and clear\n \t\t * all the unused name-entries.\n \t\t */\n-\t\tfor (i = 0; i < 3; i++) {\n-\t\t\tif (mask & (1u << i)) {\n+\t\tfor (i = 0; i < n; i++) {\n+\t\t\tif (mask & (1ul << i)) {\n \t\t\t\tupdate_tree_entry(t+i);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tentry_clear(entry + i);\n \t\t}\n+\t\tcallback(n, mask, entry, base);\n+\t}\n+\tfree(entry);\n+}\n \n-\t\t/* Same in both? */\n-\t\tif (same_entry(entry+1, entry+2)) {\n-\t\t\tif (entry[0].sha1) {\n-\t\t\t\tresolve(base, NULL, entry+1);\n-\t\t\t\tcontinue;\n-\t\t\t}\n+/*\n+ * Merge two trees together (t[1] and t[2]), using a common base (t[0])\n+ * as the origin.\n+ *\n+ * This walks the (sorted) trees in lock-step, checking every possible\n+ * name. Note that directories automatically sort differently from other\n+ * files (see \"base_name_compare\"), so you'll never see file/directory\n+ * conflicts, because they won't ever compare the same.\n+ *\n+ * IOW, if a directory changes to a filename, it will automatically be\n+ * seen as the directory going away, and the filename being created.\n+ *\n+ * Think of this as a three-way diff.\n+ *\n+ * The output will be either:\n+ *  - successful merge\n+ *\t \"0 mode sha1 filename\"\n+ *    NOTE NOTE NOTE! FIXME! We really really need to walk the index\n+ *    in parallel with this too!\n+ * \n+ *  - conflict:\n+ *\t\"1 mode sha1 filename\"\n+ *\t\"2 mode sha1 filename\"\n+ *\t\"3 mode sha1 filename\"\n+ *    where not all of the 1/2/3 lines may exist, of course.\n+ *\n+ * The successful merge rules are the same as for the three-way merge\n+ * in git-read-tree.\n+ */\n+static void threeway_callback(int n, unsigned long mask, struct name_entry *entry, const char *base)\n+{\n+\t/* Same in both? */\n+\tif (same_entry(entry+1, entry+2)) {\n+\t\tif (entry[0].sha1) {\n+\t\t\tresolve(base, NULL, entry+1);\n+\t\t\treturn;\n \t\t}\n+\t}\n \n-\t\tif (same_entry(entry+0, entry+1)) {\n-\t\t\tif (entry[2].sha1 && !S_ISDIR(entry[2].mode)) {\n-\t\t\t\tresolve(base, entry+1, entry+2);\n-\t\t\t\tcontinue;\n-\t\t\t}\n+\tif (same_entry(entry+0, entry+1)) {\n+\t\tif (entry[2].sha1 && !S_ISDIR(entry[2].mode)) {\n+\t\t\tresolve(base, entry+1, entry+2);\n+\t\t\treturn;\n \t\t}\n+\t}\n \n-\t\tif (same_entry(entry+0, entry+2)) {\n-\t\t\tif (entry[1].sha1 && !S_ISDIR(entry[1].mode)) {\n-\t\t\t\tresolve(base, NULL, entry+1);\n-\t\t\t\tcontinue;\n-\t\t\t}\n+\tif (same_entry(entry+0, entry+2)) {\n+\t\tif (entry[1].sha1 && !S_ISDIR(entry[1].mode)) {\n+\t\t\tresolve(base, NULL, entry+1);\n+\t\t\treturn;\n \t\t}\n-\n-\t\tunresolved(base, entry);\n \t}\n+\n+\tunresolved(base, entry);\n+}\n+\n+static void merge_trees(struct tree_desc t[3], const char *base)\n+{\n+\ttraverse_trees(3, t, base, threeway_callback);\n }\n \n static void *get_tree_descriptor(struct tree_desc *desc, const char *rev)\n"},{"id":"16239","messageId":"7vlkwckml7.fsf@assigned-by-dhcp.cox.net","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602151915010.916@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-16T03:29:40Z","receivedAt":"2006-02-16T03:29:40Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Junio, that \"traverse_trees()\" logic is totally independent of whether we \n> actually do \"git-merge-tree\" or not, so if you want to, I could split up \n> the patches the other way (and merge \"traverse_trees()\" first as a new \n> interface, independently).\n\nI won't have time to look at the actual patch tonight but I am\ninterested.  I think the general idea should work nice with both\nmulti-base and octopus merge cases as well ;-).\n"},{"id":"16286","messageId":"20060216203211.GA14408@c165.ib.student.liu.se","threadId":"3265","inReplyTo":"Pine.LNX.4.64.0602141953081.3691@g5.osdl.org","subject":"Re: Handling large files with GIT","fromName":"Fredrik Kuivinen","fromEmail":"freku045@student.liu.se","sentAt":"2006-02-16T20:32:11Z","receivedAt":"2006-02-16T20:32:11Z","isPatch":false,"sender":{"key":"frekui@gmail.com","avatar":"https://avatars.githubusercontent.com/u/13770967?v=4"},"body":"On Tue, Feb 14, 2006 at 07:58:03PM -0800, Linus Torvalds wrote:\n> \n> \n> On Tue, 14 Feb 2006, Linus Torvalds wrote:\n> > \n> > So in case people want to try, here's a third patch. Oh, and it's against \n> > my _original_ path, not incremental to the middle one (ie both patches two \n> > and three are against patch #1, it's not a nice series).\n> > \n> > Now I'm really done, and won't be sending out any more patches today. \n> \n> Still true. I've just been thinking about the last state.\n> \n> As far as I can tell, the output from git-merge-tree with that fix to only \n> simplify subdirectories that match exactly in all of base/branch1/branch2 \n> is precisely the output that git-merge-recursive actually wants.\n> \n> Rather than doing a three-way merge with \"git-read-tree\", and then doing \n> \"git-ls-files --unmerged\", I think this gives the same result much more \n> efficiently.\n> \n> That said, I can't follow the python code, so maybe I'm missing something. \n> Fredrik cc'd, in case he can put me right.\n> \n\nI don't think you miss anything. I _think_ (I haven't looked at this\ntoo close yet) that it shouldn't be too much work to make\ngit-merge-recursive make use of the git-merge-tree thing.\n\n- Fredrik\n"}]}