{"thread":{"id":"36710","subject":"GIT and large files","startedAt":"2014-05-20T15:37:41Z","lastAt":"2014-05-20T19:01:34Z","messageCount":10,"participants":["Stewart, Louis (IS)","Jason Pyeron","Marius Storm-Olsen","Junio C Hamano","Thomas Braun","Konstantin Khomoutov"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"242276","messageId":"C755E6FBF6DC4447BEF161CE48BDE0BD2F0CD53E@XMBVAG73.northgrum.com","threadId":"36710","inReplyTo":null,"subject":"GIT and large files","fromName":"Stewart, Louis (IS)","fromEmail":"louis.stewart@ngc.com","sentAt":"2014-05-20T15:37:41Z","receivedAt":"2014-05-20T15:37:41Z","isPatch":false,"sender":{"key":"louis.stewart@ngc.com","avatar":null},"body":"Can GIT handle versioning of large 20+ GB files in a directory?\n\nLou Stewart\nAOCWS Software Configuration Management\n757-269-2388\n"},{"id":"242283","messageId":"8D5146DC3E984386B017EC4722C56D93@black","threadId":"36710","inReplyTo":"C755E6FBF6DC4447BEF161CE48BDE0BD2F0CD53E@XMBVAG73.northgrum.com","subject":"RE: GIT and large files","fromName":"Jason Pyeron","fromEmail":"jpyeron@pdinc.us","sentAt":"2014-05-20T16:03:18Z","receivedAt":"2014-05-20T16:03:18Z","isPatch":false,"sender":{"key":"jpyeron@pdinc.us","avatar":"https://gravatar.com/avatar/c2e53452caa53d940768a1ffc9cf76196d851b9b534b7a39cd39852a70a0508f?d=mp&s=160"},"body":"> -----Original Message-----\n> From: Stewart, Louis (IS)\n> Sent: Tuesday, May 20, 2014 11:38\n> \n> Can GIT handle versioning of large 20+ GB files in a directory?\n\nAre you asking 20 files of a GB each or files 20GB each?\n\nA what and why may help with the underlying questions.\n\nv/r,\n\nJason Pyeron\n\n--\n-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-\n-                                                               -\n- Jason Pyeron                      PD Inc. http://www.pdinc.us -\n- Principal Consultant              10 West 24th Street #100    -\n- +1 (443) 269-1555 x333            Baltimore, Maryland 21218   -\n-                                                               -\n-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-\nThis message is copyright PD Inc, subject to license 20080407P00.\n\n \n"},{"id":"242280","messageId":"C755E6FBF6DC4447BEF161CE48BDE0BD2F0CD5D7@XMBVAG73.northgrum.com","threadId":"36710","inReplyTo":"CALygMcCifDd4LAddZJ4tNcqqwBSvb6BGzTODHBzshBOjCwSrHQ@mail.gmail.com","subject":"RE: EXT :Re: GIT and large files","fromName":"Stewart, Louis (IS)","fromEmail":"louis.stewart@ngc.com","sentAt":"2014-05-20T16:53:15Z","receivedAt":"2014-05-20T16:53:15Z","isPatch":false,"sender":{"key":"louis.stewart@ngc.com","avatar":null},"body":"The files in question would be in directory containing many files some small other huge (example: text files, docs,and jpgs are Mbs, but executables and ova images are GBs, etc).\n\nLou\n\nFrom: Gary Fixler [mailto:gfixler@gmail.com] \nSent: Tuesday, May 20, 2014 12:09 PM\nTo: Stewart, Louis (IS)\nCc: git@vger.kernel.org\nSubject: EXT :Re: GIT and large files\n\nTechnically yes, but from a practical standpoint, not really. Facebook recently revealed that they have a 54GB git repo[1], but I doubt it has 20+GB files in it. I've put 18GB of photos into a git repo, but everything about the process was fairly painful, and I don't plan to do it again.\nAre your files non-mergeable binaries (e.g. videos)? The biggest problem here is with branching and merging. Conflict resolution with non-mergeable assets ends up an us-vs-them fight, and I don't understand all of the particulars of that. From git's standpoint it's simple - you just have to choose one or the other. From a workflow standpoint, you end up causing trouble if two people have changed an asset, and both people consider their change important. Centralized systems get around this problem with locks.\nGit could do this, and I've thought about it quite a bit. I work in games - we have code, but also a lot of binaries, that I'd like to keep in sync with the code. For awhile I considered suggesting some ideas to this group, but I'm pretty sure the locking issue makes it a non-starter. The basic idea - skipping locking for the moment - would be to allow setting git attributes by file type, file size threshold, folder, etc., to allow git to know that some files are considered \"bigfiles.\" These could be placed into the objects folder, but I'd actually prefer they go into a .git/bigfile folder. They'd still be saved as contents under their hash, but a normal git transfer wouldn't send them. They'd be in the tree as 'big' or 'bigfile' (instead of 'blob', 'tree', or 'commit' (for submodules)).\n\nGit would warn you on push that there were bigfiles to send, and you could add, say, --with-big to also send them, or send them later with, say, `git push --big`. They'd simply be zipped up and sent over, without any packfile fanciness. When you clone, you wouldn't get the bigfiles, unless you specified --with-big, and it would warn you that there are also bigfiles, and tell you what command to run to get also get them (`git fetch --big`, perhaps). Git status would always let you know if you were missing bigfiles. I think hopping around between commits would follow the same strategy, you'd always have to, e.g. `git checkout foo --with-big`, or `git checkout foo` and then `git update big` (or whatever - I'm not married to any of these names).\n\nResolving conflicts on merge would simply have to be up to you. It would be documented clearly that you're entering weird territory, and that your team has to deal with bigfiles somehow, perhaps with some suggested strategies (\"Pass the conch?\"). I could imagine some strategies for this. Maybe bigfiles require connecting to a blessed repo to grab the right to make a commit on it. That has many problems, of course, and now I can feel everyone reading this shifting uneasily in their seats :)\n-g\n\n[1] https://twitter.com/feross/status/459259593630433280\n\nOn Tue, May 20, 2014 at 8:37 AM, Stewart, Louis (IS) <louis.stewart@ngc.com> wrote:\nCan GIT handle versioning of large 20+ GB files in a directory?\n\nLou Stewart\nAOCWS Software Configuration Management\n757-269-2388\n\n--\nTo unsubscribe from this list: send the line \"unsubscribe git\" in\nthe body of a message to majordomo@vger.kernel.org\nMore majordomo info at  http://vger.kernel.org/majordomo-info.html\n\n"},{"id":"242285","messageId":"537B8C12.9020002@gmail.com","threadId":"36710","inReplyTo":"C755E6FBF6DC4447BEF161CE48BDE0BD2F0CD53E@XMBVAG73.northgrum.com","subject":"Re: GIT and large files","fromName":"Marius Storm-Olsen","fromEmail":"mstormo@gmail.com","sentAt":"2014-05-20T17:08:34Z","receivedAt":"2014-05-20T17:08:34Z","isPatch":false,"sender":{"key":"mstormo@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1500?v=4"},"body":"On 5/20/2014 10:37 AM, Stewart, Louis (IS) wrote:\n> Can GIT handle versioning of large 20+ GB files in a directory?\n\nMaybe you're looking for git-annex?\n\nhttps://git-annex.branchable.com/\n\n-- \n.marius\n"},{"id":"242287","messageId":"xmqqmwec1i9f.fsf@gitster.dls.corp.google.com","threadId":"36710","inReplyTo":"C755E6FBF6DC4447BEF161CE48BDE0BD2F0CD53E@XMBVAG73.northgrum.com","subject":"Re: GIT and large files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-05-20T17:18:04Z","receivedAt":"2014-05-20T17:18:04Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Stewart, Louis (IS)\" <louis.stewart@ngc.com> writes:\n\n> Can GIT handle versioning of large 20+ GB files in a directory?\n\nI think you can \"git add\" such files, push/fetch histories that\ncontains such files over the wire, and \"git checkout\" such files,\nbut naturally reading, processing and writing 20+GB would take some\ntime.  In order to run operations that need to see the changes,\ne.g. \"git log -p\", a real content-level merge, etc., you would also\nneed sufficient memory because we do things in-core.\n"},{"id":"242288","messageId":"C755E6FBF6DC4447BEF161CE48BDE0BD2F0CD631@XMBVAG73.northgrum.com","threadId":"36710","inReplyTo":"xmqqmwec1i9f.fsf@gitster.dls.corp.google.com","subject":"RE: EXT :Re: GIT and large files","fromName":"Stewart, Louis (IS)","fromEmail":"louis.stewart@ngc.com","sentAt":"2014-05-20T17:24:12Z","receivedAt":"2014-05-20T17:24:12Z","isPatch":false,"sender":{"key":"louis.stewart@ngc.com","avatar":null},"body":"Thanks for the reply.  I just read the intro to GIT and I am concerned about the part that it will copy the whole repository to the developers work area.  They really just need the one directory and files under that one directory. The history has TBs of data.\n\nLou\n\n-----Original Message-----\nFrom: Junio C Hamano [mailto:gitster@pobox.com] \nSent: Tuesday, May 20, 2014 1:18 PM\nTo: Stewart, Louis (IS)\nCc: git@vger.kernel.org\nSubject: EXT :Re: GIT and large files\n\n\"Stewart, Louis (IS)\" <louis.stewart@ngc.com> writes:\n\n> Can GIT handle versioning of large 20+ GB files in a directory?\n\nI think you can \"git add\" such files, push/fetch histories that contains such files over the wire, and \"git checkout\" such files, but naturally reading, processing and writing 20+GB would take some time.  In order to run operations that need to see the changes, e.g. \"git log -p\", a real content-level merge, etc., you would also need sufficient memory because we do things in-core.\n"},{"id":"242295","messageId":"xmqqoaysz59s.fsf@gitster.dls.corp.google.com","threadId":"36710","inReplyTo":"C755E6FBF6DC4447BEF161CE48BDE0BD2F0CD631@XMBVAG73.northgrum.com","subject":"Re: EXT :Re: GIT and large files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-05-20T18:14:39Z","receivedAt":"2014-05-20T18:14:39Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Stewart, Louis (IS)\" <louis.stewart@ngc.com> writes:\n\n> Thanks for the reply.  I just read the intro to GIT and I am\n> concerned about the part that it will copy the whole repository to\n> the developers work area.  They really just need the one directory\n> and files under that one directory. The history has TBs of data.\n\nThen you will spend time reading, processing and writing TBs of data\nwhen you clone, unless your developers do something to limit the\nhistory they fetch, e.g. by shallowly cloning.\n\n>\n> Lou\n>\n> -----Original Message-----\n> From: Junio C Hamano [mailto:gitster@pobox.com] \n> Sent: Tuesday, May 20, 2014 1:18 PM\n> To: Stewart, Louis (IS)\n> Cc: git@vger.kernel.org\n> Subject: EXT :Re: GIT and large files\n>\n> \"Stewart, Louis (IS)\" <louis.stewart@ngc.com> writes:\n>\n>> Can GIT handle versioning of large 20+ GB files in a directory?\n>\n> I think you can \"git add\" such files, push/fetch histories that contains such files over the wire, and \"git checkout\" such files, but naturally reading, processing and writing 20+GB would take some time.  In order to run operations that need to see the changes, e.g. \"git log -p\", a real content-level merge, etc., you would also need sufficient memory because we do things in-core.\n"},{"id":"242296","messageId":"C755E6FBF6DC4447BEF161CE48BDE0BD2F0CD670@XMBVAG73.northgrum.com","threadId":"36710","inReplyTo":"xmqqoaysz59s.fsf@gitster.dls.corp.google.com","subject":"RE: EXT :Re: GIT and large files","fromName":"Stewart, Louis (IS)","fromEmail":"louis.stewart@ngc.com","sentAt":"2014-05-20T18:18:08Z","receivedAt":"2014-05-20T18:18:08Z","isPatch":false,"sender":{"key":"louis.stewart@ngc.com","avatar":null},"body":">From you response then there is a method to only obtain the Project, Directory and Files (which could hold 80 GBs of data) and not the rest of the Repository that contained the full overall Projects?\n\n-----Original Message-----\nFrom: Junio C Hamano [mailto:gitster@pobox.com] \nSent: Tuesday, May 20, 2014 2:15 PM\nTo: Stewart, Louis (IS)\nCc: git@vger.kernel.org\nSubject: Re: EXT :Re: GIT and large files\n\n\"Stewart, Louis (IS)\" <louis.stewart@ngc.com> writes:\n\n> Thanks for the reply.  I just read the intro to GIT and I am concerned \n> about the part that it will copy the whole repository to the \n> developers work area.  They really just need the one directory and \n> files under that one directory. The history has TBs of data.\n\nThen you will spend time reading, processing and writing TBs of data when you clone, unless your developers do something to limit the history they fetch, e.g. by shallowly cloning.\n\n>\n> Lou\n>\n> -----Original Message-----\n> From: Junio C Hamano [mailto:gitster@pobox.com]\n> Sent: Tuesday, May 20, 2014 1:18 PM\n> To: Stewart, Louis (IS)\n> Cc: git@vger.kernel.org\n> Subject: EXT :Re: GIT and large files\n>\n> \"Stewart, Louis (IS)\" <louis.stewart@ngc.com> writes:\n>\n>> Can GIT handle versioning of large 20+ GB files in a directory?\n>\n> I think you can \"git add\" such files, push/fetch histories that contains such files over the wire, and \"git checkout\" such files, but naturally reading, processing and writing 20+GB would take some time.  In order to run operations that need to see the changes, e.g. \"git log -p\", a real content-level merge, etc., you would also need sufficient memory because we do things in-core.\n"},{"id":"242301","messageId":"1400610440.14137.18.camel@thomas-debian-x64","threadId":"36710","inReplyTo":"C755E6FBF6DC4447BEF161CE48BDE0BD2F0CD631@XMBVAG73.northgrum.com","subject":"Re: EXT :Re: GIT and large files","fromName":"Thomas Braun","fromEmail":"thomas.braun@virtuell-zuhause.de","sentAt":"2014-05-20T18:27:20Z","receivedAt":"2014-05-20T18:27:20Z","isPatch":false,"sender":{"key":"thomas.braun@virtuell-zuhause.de","avatar":"https://avatars.githubusercontent.com/u/1185677?v=4"},"body":"Am Dienstag, den 20.05.2014, 17:24 +0000 schrieb Stewart, Louis (IS):\n> Thanks for the reply.  I just read the intro to GIT and I am concerned\n> about the part that it will copy the whole repository to the developers\n> work area.  They really just need the one directory and files under\n> that one directory. The history has TBs of data.\n> \n> Lou\n> \n> -----Original Message-----\n> From: Junio C Hamano [mailto:gitster@pobox.com] \n> Sent: Tuesday, May 20, 2014 1:18 PM\n> To: Stewart, Louis (IS)\n> Cc: git@vger.kernel.org\n> Subject: EXT :Re: GIT and large files\n> \n> \"Stewart, Louis (IS)\" <louis.stewart@ngc.com> writes:\n> \n> > Can GIT handle versioning of large 20+ GB files in a directory?\n> \n> I think you can \"git add\" such files, push/fetch histories that\n> contains such files over the wire, and \"git checkout\" such files, but\n> naturally reading, processing and writing 20+GB would take some time. \n> In order to run operations that need to see the changes, e.g. \"git log\n> -p\", a real content-level merge, etc., you would also need sufficient\n> memory because we do things in-core.\n\nPreventing that a clone fetches the whole history can be done with the\n--depth option of git clone.\n\nThe question is what do you want to do with these 20G files?\nJust store them in the repo and *very* occasionally change them?\nFor that you need a 64bit compiled version of git with enough ram. 32G\ndoes the trick here. Everything with git 1.9.1.\n\nDoing some tests on my machine with a normal harddisc gives (sorry for\nLC_ALL != C):\n$time git add file.dat; time git commit -m \"add file\"; time git status\n\nreal    16m17.913s\nuser    13m3.965s\nsys     0m22.461s\n[master 15fa953] add file\n 1 file changed, 0 insertions(+), 0 deletions(-)\n create mode 100644 file.dat\n\nreal    15m36.666s\nuser    13m26.962s\nsys     0m16.185s\n# Auf Branch master\nnichts zu committen, Arbeitsverzeichnis unverändert\n\nreal    11m58.936s\nuser    11m50.300s\nsys     0m5.468s\n\n$ls -lh\n-rw-r--r-- 1 thomas thomas 20G Mai 20 19:01 file.dat\n\nSo this works but aint fast.\n\nPlaying some tricks with --assume-unchanged helps here:\n$git update-index --assume-unchanged file.dat\n$time git status\n# Auf Branch master\nnichts zu committen, Arbeitsverzeichnis unverändert\n\nreal    0m0.003s\nuser    0m0.000s\nsys     0m0.000s\n\nThis trick is only save if you *know* that file.dat does not change.\n\nAnd btw I also set \n$cat .gitattributes \n*.dat -delta\nas delta compresssion should be skipped in any case.\n\nPushing and pulling these files to and from a server needs some tweaking\non the server side, otherwise the occasional git gc might kill the box.\n \nBtw. I happily have files with 1.5GB in my git repositories and also\nchange them. And also work with git for windows. So in this region of\nfile sizes things work quite well.\n"},{"id":"242310","messageId":"20140520230134.25cbbffe1b0ce95de60024a5@domain007.com","threadId":"36710","inReplyTo":"C755E6FBF6DC4447BEF161CE48BDE0BD2F0CD670@XMBVAG73.northgrum.com","subject":"Re: EXT :Re: GIT and large files","fromName":"Konstantin Khomoutov","fromEmail":"flatworm@users.sourceforge.net","sentAt":"2014-05-20T19:01:34Z","receivedAt":"2014-05-20T19:01:34Z","isPatch":false,"sender":{"key":"flatworm@users.sourceforge.net","avatar":null},"body":"On Tue, 20 May 2014 18:18:08 +0000\n\"Stewart, Louis (IS)\" <louis.stewart@ngc.com> wrote:\n\n> From you response then there is a method to only obtain the Project,\n> Directory and Files (which could hold 80 GBs of data) and not the\n> rest of the Repository that contained the full overall Projects?\n\nPlease google the phrase \"Git shallow cloning\".\n\nI would also recommend to read up on git-annex [1].\n\nYou might also consider using Subversion as it seems you do not need\nmost benefits Git has over it and want certain benefits Subversion has\nover Git:\n* You don't need a distributed VCS (as you don't want each developer to\n  have a full clone).\n* You only need a single slice of the repository history at any given\n  revision on a developer's machine, and this is *almost* what\n  Subversion does: it will keep the so-called \"base\" (or \"pristine\")\n  versions of files comprising the revision you will check out, plus\n  the checked out files theirselves.  So, twice the space of the files\n  comprising a revision.\n* Subversion allows you to check out only a single folder out of the\n  entire revision.\n* IIRC, subversion supports locks, when a developer might tell the\n  server they're editing a file, and this will prevent other devs from\n  locking the same file.  This might be used to serialize editions of\n  huge and/or unmergeable files.  Git can't do that (without\n  non-standard tools deployed on the side or a centralized \"meeting\n  point\" repository).\n\nMy point is that while Git is fantastic for managing source code\nprojects and project of similar types with regard to their contents,\nit seems your requirements are mainly not suitable for the use case\nGit is best tailored for.  Your apparent lack of familiarity with Git\nmight as well bite you later should you pick it right now.  At least\nplease consider reading a book or some other introduction-level\nmaterial on Git to get the feeling of typical workflows used with it.\n\n\n1. https://git-annex.branchable.com/\n"}]}