{"thread":{"id":"18494","subject":"large(25G) repository in git","startedAt":"2009-03-23T21:10:11Z","lastAt":"2009-03-26T16:35:17Z","messageCount":16,"participants":["Adam Heath","Nicolas Pitre","Andreas Ericsson","david@lang.hm","Sam Hocevar","Marcel M. Cary"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"109105","messageId":"49C7FAB3.7080301@brainfood.com","threadId":"18494","inReplyTo":null,"subject":"large(25G) repository in git","fromName":"Adam Heath","fromEmail":"doogie@brainfood.com","sentAt":"2009-03-23T21:10:11Z","receivedAt":"2009-03-23T21:10:11Z","isPatch":false,"sender":{"key":"doogie@brainfood.com","avatar":"https://gravatar.com/avatar/c9e42ee14e1998b527796b1d8a35d75842d6a6aa885411ea1d5c97cb2d6923e6?d=mp&s=160"},"body":"We maintain a website in git.  This website has a bunch of backend\nserver code, and a bunch of data files.  Alot of these files are full\nvideos.\n\nWe use git, so that the distributed nature of website development can\nbe supported.  Quite often, you'll have a production server, with\nonline changes occurring(we support in-browser editting of content), a\npreview server, where large-scale code changes can be previewed, then\na development server, one per programmer(or more).\n\nLast friday, I was doing a checkin on the production server, and found\n1.6G of new files.  git was quite able at committing that.  However,\npushing was problematic.  I was pushing over ssh; so, a new ssh\nconnection was open to the preview server.  After doing so, git tried\nto create a new pack file.  This took *ages*, and the ssh connection\ndied.  So did git, when it finally got done with the new pack, and\ndiscovered the ssh connection was gone.\n\nSo, to work around that, I ran git gc.  When done, I discovered that\ngit repacked the *entire* repository.  While not something I care for,\nI can understand that, and live with it.  It just took *hours* to do so.\n\nThen, what really annoys me, is that when I finally did the push, it\ntried sending the single 27G pack file, when the remote already had\n25G of the repository in several different packs(the site was an\nhg->git conversion).  This part is just unacceptable.\n\nSo, here are my questions/observations:\n\n1: Handle the case of the ssh connection dying during git push(seems\nsimple).\n\n2: Is there an option to tell git to *not* be so thorough when trying\nto find similiar files.  videos/doc/pdf/etc aren't always very\ndeltafiable, so I'd be happy to just do full content compares.\n\n3: delta packs seem to be poorly done.  it seems that if one repo gets\nrepacked completely, that the entire new pack gets sent, when the\ntarget has most of the objects already.\n\n4: Are there any config options I can set to help in this?  There are\ntons of options, and some documentation as to what each one does, but\nno recommended practices type doc, that describes what should be done\nfor different kinds of workflows.\n\nps: Thank you for your time.  I hope that someone has answers for me.\n\npps: I'm not subscribed, please cc me.  If I need to be subscribed,\nI'll do so, if told.\n"},{"id":"109138","messageId":"alpine.LFD.2.00.0903232056520.26337@xanadu.home","threadId":"18494","inReplyTo":"49C7FAB3.7080301@brainfood.com","subject":"Re: large(25G) repository in git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-03-24T01:19:57Z","receivedAt":"2009-03-24T01:19:57Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 23 Mar 2009, Adam Heath wrote:\n\n> Last friday, I was doing a checkin on the production server, and found\n> 1.6G of new files.  git was quite able at committing that.  However,\n> pushing was problematic.  I was pushing over ssh; so, a new ssh\n> connection was open to the preview server.  After doing so, git tried\n> to create a new pack file.  This took *ages*, and the ssh connection\n> died.  So did git, when it finally got done with the new pack, and\n> discovered the ssh connection was gone.\n\nStrange.  You could instruct ssh to keep the connection up with the \nServerAliveInterval option (see the ssh_config man page).\n\n> So, to work around that, I ran git gc.  When done, I discovered that\n> git repacked the *entire* repository.  While not something I care for,\n> I can understand that, and live with it.  It just took *hours* to do so.\n> \n> Then, what really annoys me, is that when I finally did the push, it\n> tried sending the single 27G pack file, when the remote already had\n> 25G of the repository in several different packs(the site was an\n> hg->git conversion).  This part is just unacceptable.\n\nThis shouldn't happen either.  When pushing, git reconstruct a pack with \nonly the necessary objects to transmit.  Are you sure it was really \ntrying to send a 27G pack?\n\n> So, here are my questions/observations:\n> \n> 1: Handle the case of the ssh connection dying during git push(seems\n> simple).\n\nSee above.\n\n> 2: Is there an option to tell git to *not* be so thorough when trying\n> to find similiar files.  videos/doc/pdf/etc aren't always very\n> deltafiable, so I'd be happy to just do full content compares.\n\nLook at the gitattribute documentation.  One thing that the doc appears \nto be missing is information about the \"delta\" attribute.  You can \ndisable delta compression on a file pattern that way.\n\n> 3: delta packs seem to be poorly done.  it seems that if one repo gets\n> repacked completely, that the entire new pack gets sent, when the\n> target has most of the objects already.\n\nThis is not supposed to happen.  Please provide more details if you can.\n\n\nNicolas\n"},{"id":"109186","messageId":"49C8A102.6090408@op5.se","threadId":"18494","inReplyTo":"49C7FAB3.7080301@brainfood.com","subject":"Re: large(25G) repository in git","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2009-03-24T08:59:46Z","receivedAt":"2009-03-24T08:59:46Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Adam Heath wrote:\n> We maintain a website in git.  This website has a bunch of backend\n> server code, and a bunch of data files.  Alot of these files are full\n> videos.\n> \n\nFirst of all, I'm going to hint that you would be far better off\nkeeping the media files in a separate repository, linked in as a\nsubmodule in git and with tweaked configuration settings with the\nspecific aim of handling huge files.\n\nThe basis of such a repository is probably the following config\nsettings, since media files very rarely compress enough to be\nworth the effort, and their own compressed formats make them\nvery unsuitable delta candidates:\n[pack]\n   # disable delta-based packing\n   depth = 1\n   # disable compression\n   compression = 0\n\n[gc]\n   # don't auto-pack, ever\n   auto = 0\n   # never automatically consolidate un-.keep'd packs\n   autopacklimit = 0\n\nYou will have to manually repack this repository from time to\ntime, and it's almost certainly a good idea to mark the\nresulting packs with .keep to avoid copying tons of data.\nWhen packs are being created, objects can be copied from\nexisting packs, and send-pack will make use of that so that what\ngoes over the wire will simply be copied from the existing packs.\n\nYMMV. If you do come up with settings that work fine for huge\nrepos made up of mostly media files, please share your findings.\n\n> We use git, so that the distributed nature of website development can\n> be supported.  Quite often, you'll have a production server, with\n> online changes occurring(we support in-browser editting of content), a\n> preview server, where large-scale code changes can be previewed, then\n> a development server, one per programmer(or more).\n> \n> Last friday, I was doing a checkin on the production server, and found\n> 1.6G of new files.  git was quite able at committing that.  However,\n> pushing was problematic.  I was pushing over ssh; so, a new ssh\n> connection was open to the preview server.  After doing so, git tried\n> to create a new pack file.  This took *ages*, and the ssh connection\n> died.  So did git, when it finally got done with the new pack, and\n> discovered the ssh connection was gone.\n> \n> So, to work around that, I ran git gc.  When done, I discovered that\n> git repacked the *entire* repository.  While not something I care for,\n> I can understand that, and live with it.  It just took *hours* to do so.\n> \n\nI'm not sure what, if any, magic \"git gc\" applies before spawning\n\"git repack\", but running \"git repack\" directly would almost certainly\nhave produced an incremental pack. Perhaps we need to make gc less\nmagic.\n\n> Then, what really annoys me, is that when I finally did the push, it\n> tried sending the single 27G pack file, when the remote already had\n> 25G of the repository in several different packs(the site was an\n> hg->git conversion).  This part is just unacceptable.\n> \n\nAgreed. I've never run across that problem, so I can only assume it\nhas something to do with many huge files being in the pack.\n\n> So, here are my questions/observations:\n> \n> 1: Handle the case of the ssh connection dying during git push(seems\n> simple).\n> \n\nNot necessarily all that simple (we do not want to touch the ssh\npassword if we can possibly avoid it, but the user shouldn't have\nto type it more than once), but certainly doable. Easier would\nprobably be to recommend adding the proper SSH config variables,\nas has been stated elsewhere.\n\n> 2: Is there an option to tell git to *not* be so thorough when trying\n> to find similiar files.  videos/doc/pdf/etc aren't always very\n> deltafiable, so I'd be happy to just do full content compares.\n> \n\nSee above. I *think* you can also do this with git-attributes, but\nI'm not sure. However, keeping the large media files in a sub-module\nwould nicely solve that problem anyway, and is probably a good idea\neven with git-attributes support for pack delta- and compression\nsettings.\n\n> 3: delta packs seem to be poorly done.  it seems that if one repo gets\n> repacked completely, that the entire new pack gets sent, when the\n> target has most of the objects already.\n> \n\nThis is certainly not the case for most repositories. I believe there's\nsomething being triggered from repositories with many huge files though.\n\n> 4: Are there any config options I can set to help in this?  There are\n> tons of options, and some documentation as to what each one does, but\n> no recommended practices type doc, that describes what should be done\n> for different kinds of workflows.\n> \n\nhttp://www.thousandparsec.net/~tim/media+git.pdf probably holds all the\nrelevant information when it comes to storing large media files with\ngit. I have not checked and have no inclination to do so.\n\n> ps: Thank you for your time.  I hope that someone has answers for me.\n> \n\nAnswers aplenty, I hope. I have neither time nor interest in developing\nthis though, so the task of creating patches and/or documentation will\nhave to fall to someone else.\n\n> pps: I'm not subscribed, please cc me.  If I need to be subscribed,\n> I'll do so, if told.\n\nSubscribing won't be necessary. The custom on git@vger is to always Cc\nall who participate in the discussion, and only cull those who state\nthey're no longer interested in the topic.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n\nConsidering the successes of the wars on alcohol, poverty, drugs and\nterror, I think we should give some serious thought to declaring war\non peace.\n"},{"id":"109228","messageId":"49C91F87.3050105@brainfood.com","threadId":"18494","inReplyTo":"alpine.LFD.2.00.0903232056520.26337@xanadu.home","subject":"Re: large(25G) repository in git","fromName":"Adam Heath","fromEmail":"doogie@brainfood.com","sentAt":"2009-03-24T17:59:35Z","receivedAt":"2009-03-24T17:59:35Z","isPatch":false,"sender":{"key":"doogie@brainfood.com","avatar":"https://gravatar.com/avatar/c9e42ee14e1998b527796b1d8a35d75842d6a6aa885411ea1d5c97cb2d6923e6?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n\n> Strange.  You could instruct ssh to keep the connection up with the \n> ServerAliveInterval option (see the ssh_config man page).\n\nSure, could do that.  Already have a separate ssh config entry for\nthis host.  But why should a connection be kept open for that long?\nWhy not close and re-open?\n\nConsider the case of other protocol access.  http/git/ssh.  Should\nthey *all* be changed to allow for this?  Wouldn't it be simpler to\njust make git smarter?\n\n>> So, to work around that, I ran git gc.  When done, I discovered that\n>> git repacked the *entire* repository.  While not something I care for,\n>> I can understand that, and live with it.  It just took *hours* to do so.\n>>\n>> Then, what really annoys me, is that when I finally did the push, it\n>> tried sending the single 27G pack file, when the remote already had\n>> 25G of the repository in several different packs(the site was an\n>> hg->git conversion).  This part is just unacceptable.\n> \n> This shouldn't happen either.  When pushing, git reconstruct a pack with \n> only the necessary objects to transmit.  Are you sure it was really \n> trying to send a 27G pack?\n\nOf course I'm sure.  I wouldn't have sent the email if it didn't\nhappen.  And, I have the bandwidthd graph and lost time to prove it.\n\nAfter I ran git push, ssh timed out, the temp pack that was created\nwas then removed, as git complained about the connection being gone.\n\nI then decided to do a 'git gc', which collapsed all the separate\npacks into one.  This allowed git push to proceed quickly, but at that\npoint, it started sending the entire pack.\n\nIt's entirely possible that the temp pack created by git push was\nincremental; it just took too long to create it, so it got aborted.\n\nBut, doing git gc shouldn't cause things to be resent.\n\nThe machines in question have done push before.  Even small amounts;\njust the set of objects that are newer.  It's just this time, when the\n1.6G of new data was added, git ended up creating a new pack file,\nthat contained the entire repo, and then tried sending that.\n\nI forgot to mention previously, that the source machine was running\ngit 1.5.6.5, and was pushing to 1.5.6.3.\n\nI've tried duplicating this problem on a machine with 1.6.1.3, but\neither I don't fully understand the issue enough to replicate it, or\nthe newer git doesn't have the problem.\n\n>> 2: Is there an option to tell git to *not* be so thorough when trying\n>> to find similiar files.  videos/doc/pdf/etc aren't always very\n>> deltafiable, so I'd be happy to just do full content compares.\n> \n> Look at the gitattribute documentation.  One thing that the doc appears \n> to be missing is information about the \"delta\" attribute.  You can \n> disable delta compression on a file pattern that way.\n\nUm, if it's missing documentation, then how am I supposed to know\nabout it?  google does give me info, tho.  Thanks for the pointer.\n\n> \n>> 3: delta packs seem to be poorly done.  it seems that if one repo gets\n>> repacked completely, that the entire new pack gets sent, when the\n>> target has most of the objects already.\n> \n> This is not supposed to happen.  Please provide more details if you can.\n\nWell, I haven't been able to replicate it with a script.  I might have\nto actually clone this huge repo, do history removal, and reapply the\nchanges, just to see if I can get it to fail.  But that will take time.\n"},{"id":"109229","messageId":"alpine.LFD.2.00.0903241404080.26337@xanadu.home","threadId":"18494","inReplyTo":"49C91F87.3050105@brainfood.com","subject":"Re: large(25G) repository in git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-03-24T18:31:31Z","receivedAt":"2009-03-24T18:31:31Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 24 Mar 2009, Adam Heath wrote:\n\n> Nicolas Pitre wrote:\n> \n> > Strange.  You could instruct ssh to keep the connection up with the \n> > ServerAliveInterval option (see the ssh_config man page).\n> \n> Sure, could do that.  Already have a separate ssh config entry for\n> this host.  But why should a connection be kept open for that long?\n> Why not close and re-open?\n\nBecause it is way more complex for git to do that than for ssh to keep \nthe connection alive.  And normally there is no need as git is supposed \nto be faster than that.\n\n> Consider the case of other protocol access.  http/git/ssh.  Should\n> they *all* be changed to allow for this?  Wouldn't it be simpler to\n> just make git smarter?\n\nMaking git faster is the solution, not working around the issue.\n\n> >> So, to work around that, I ran git gc.  When done, I discovered that\n> >> git repacked the *entire* repository.  While not something I care for,\n> >> I can understand that, and live with it.  It just took *hours* to do so.\n> >>\n> >> Then, what really annoys me, is that when I finally did the push, it\n> >> tried sending the single 27G pack file, when the remote already had\n> >> 25G of the repository in several different packs(the site was an\n> >> hg->git conversion).  This part is just unacceptable.\n> > \n> > This shouldn't happen either.  When pushing, git reconstruct a pack with \n> > only the necessary objects to transmit.  Are you sure it was really \n> > trying to send a 27G pack?\n> \n> Of course I'm sure.  I wouldn't have sent the email if it didn't\n> happen.  And, I have the bandwidthd graph and lost time to prove it.\n\nAs much as I would like to believe you, this doesn't help fixing the \nproblem if you don't provide more information about this.  For example, \nthe output from git during the whole operation might give us the \nbeginning of a clue.  Otherwise, all I can tell you is that such thing \nis not supposed to happen.\n\n> After I ran git push, ssh timed out, the temp pack that was created\n> was then removed, as git complained about the connection being gone.\n\nOn a push, there is no creation of a temp pack.  It is always produced \non the fly and pushed straight via the ssh connection.\n\n> I then decided to do a 'git gc', which collapsed all the separate\n> packs into one.  This allowed git push to proceed quickly, but at that\n> point, it started sending the entire pack.\n\nIf this was really the case, then this is definitely a bug.  Please take \na snapshot of your screen with git messages if this ever happens again.\n\n> It's entirely possible that the temp pack created by git push was\n> incremental; it just took too long to create it, so it got aborted.\n\nThe push operation has multiple phases.  You should see \"counting \nobjects\", \"compressing objects\" and \"writing objects\".  Could you give \nus an approximation of how long each of those phases took?\n\n> But, doing git gc shouldn't cause things to be resent.\n\nIndeed.\n\n> The machines in question have done push before.  Even small amounts;\n> just the set of objects that are newer.  It's just this time, when the\n> 1.6G of new data was added, git ended up creating a new pack file,\n> that contained the entire repo, and then tried sending that.\n\nAnd this is wrong.\n\n> I forgot to mention previously, that the source machine was running\n> git 1.5.6.5, and was pushing to 1.5.6.3.\n> \n> I've tried duplicating this problem on a machine with 1.6.1.3, but\n> either I don't fully understand the issue enough to replicate it, or\n> the newer git doesn't have the problem.\n\nThat's possible.  Maybe others on the list might recall possible issues \nrelated to this that might have been fixed during that time.\n\n> >> 2: Is there an option to tell git to *not* be so thorough when trying\n> >> to find similiar files.  videos/doc/pdf/etc aren't always very\n> >> deltafiable, so I'd be happy to just do full content compares.\n> > \n> > Look at the gitattribute documentation.  One thing that the doc appears \n> > to be missing is information about the \"delta\" attribute.  You can \n> > disable delta compression on a file pattern that way.\n> \n> Um, if it's missing documentation, then how am I supposed to know\n> about it?\n\nAsking on the list, like you did.  However this attribute should be \ndocumented as well of course.  I even think that someone posted a patch \nfor it a while ago which might have been dropped.\n\n\nNicolas\n"},{"id":"109230","messageId":"alpine.DEB.1.10.0903241130240.16753@asgard.lang.hm","threadId":"18494","inReplyTo":"49C91F87.3050105@brainfood.com","subject":"Re: large(25G) repository in git","fromName":"","fromEmail":"david@lang.hm","sentAt":"2009-03-24T18:33:07Z","receivedAt":"2009-03-24T18:33:07Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Tue, 24 Mar 2009, Adam Heath wrote:\n\n> Nicolas Pitre wrote:\n>\n>> Strange.  You could instruct ssh to keep the connection up with the\n>> ServerAliveInterval option (see the ssh_config man page).\n>\n> Sure, could do that.  Already have a separate ssh config entry for\n> this host.  But why should a connection be kept open for that long?\n> Why not close and re-open?\n\nwhat if the server you are connecting to is behind a load balancer? how do \nyou know that your new connection will go to the same server? if the \nclient never reconnects, how long should the server keep it's resources \ntied up 'just in case'. if something connects to the server, how does it \nknow if it's something reconnecting or connecting for the first time? (or \nsomeone connecting with the intent of messing up someone else's fetch)\n\nhaving the client reconnect to finish a single transaction starts getting \n_really_ ugly.\n\nDavid Lang\n"},{"id":"109242","messageId":"49C948C1.2070404@brainfood.com","threadId":"18494","inReplyTo":"alpine.LFD.2.00.0903241404080.26337@xanadu.home","subject":"Re: large(25G) repository in git","fromName":"Adam Heath","fromEmail":"doogie@brainfood.com","sentAt":"2009-03-24T20:55:29Z","receivedAt":"2009-03-24T20:55:29Z","isPatch":false,"sender":{"key":"doogie@brainfood.com","avatar":"https://gravatar.com/avatar/c9e42ee14e1998b527796b1d8a35d75842d6a6aa885411ea1d5c97cb2d6923e6?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n> Because it is way more complex for git to do that than for ssh to keep \n> the connection alive.  And normally there is no need as git is supposed \n> to be faster than that.\n\nSure, I'll buy that.\n\n>>>> So, to work around that, I ran git gc.  When done, I discovered that\n>>>> git repacked the *entire* repository.  While not something I care for,\n>>>> I can understand that, and live with it.  It just took *hours* to do so.\n>>>>\n>>>> Then, what really annoys me, is that when I finally did the push, it\n>>>> tried sending the single 27G pack file, when the remote already had\n>>>> 25G of the repository in several different packs(the site was an\n>>>> hg->git conversion).  This part is just unacceptable.\n>>> This shouldn't happen either.  When pushing, git reconstruct a pack with \n>>> only the necessary objects to transmit.  Are you sure it was really \n>>> trying to send a 27G pack?\n>> Of course I'm sure.  I wouldn't have sent the email if it didn't\n>> happen.  And, I have the bandwidthd graph and lost time to prove it.\n> \n> As much as I would like to believe you, this doesn't help fixing the \n> problem if you don't provide more information about this.  For example, \n> the output from git during the whole operation might give us the \n> beginning of a clue.  Otherwise, all I can tell you is that such thing \n> is not supposed to happen.\n\nFirst off, you've put a bad tone on this.  It appears that you are\nsaying I'm mistaken, and it didn't send all that data.  \"It can't\nhappen, so it didn't happen.\"  Believe me, if it hadn't resent all\nthis data, I wouldn't have even sent the email.\n\nIn any event, we got lucky.  I *do* have a log of the push side of\nthis problem.  I doubt it's enough to figure out the actual cause tho.\n\n==\nofbiz@lnxwww10:/job/@anon-site@> git push bf-yum\nCounting objects: 96637, done.\n\nCompressing objects:   6% (2413/34478)   478)\nRead from remote host @anon-site-dev@.brainfood.com: Connection reset\nby peer\nCompressing objects:  27% (9458/34478)\n\nCompressing objects: 100% (34478/34478), done.\nerror: pack-objects died with strange error\nerror: failed to push some refs to 'ssh://bf-yum/@anon-site@'\nofbiz@lnxwww10:/job/@anon-site@>\nofbiz@lnxwww10:/job/@anon-site@>\nofbiz@lnxwww10:/job/@anon-site@>\nofbiz@lnxwww10:/job/@anon-site@> git push bf-yum\nCounting objects: 96637, done.\nKilled by signal 2.:   5% (1866/34478)\n\nofbiz@lnxwww10:/job/@anon-site@> git gc\nCounting objects: 96637, done.\nCompressing objects:  27% (9453/34478)\n\nCompressing objects: 100% (34478/34478), done.\nWriting objects: 100% (96637/96637), done.\nTotal 96637 (delta 48713), reused 88929 (delta 43905)\nRemoving duplicate objects: 100% (256/256), done.\nofbiz@lnxwww10:/job/@anon-site@>\nofbiz@lnxwww10:/job/@anon-site@>\nofbiz@lnxwww10:/job/@anon-site@> du .git -sc\n26797788        .git\n26797788        total\nofbiz@lnxwww10:/job/@anon-site@> git push bf-yum\nCounting objects: 96637, done.\nCompressing objects: 100% (29670/29670), done.\nWriting objects: 100% (96637/96637), 25.49 GiB | 226 KiB/s, done.\nTotal 96637 (delta 48713), reused 96637 (delta 48713)\nTo ssh://bf-yum/@anon-site@\n * [new branch]      master -> lnxwww10\n==\nofbiz@lnxwww10:/job/@anon-site@> ls .git/objects/pack/ -l\ntotal 26762436\n-r--r--r-- 1 ofbiz users     3452052 2009-03-21 23:11\npack-0d7b399006ae0a57ff3df07fdcaedbaeb7e63d0a.idx\n-r--r--r-- 1 ofbiz users 27374508409 2009-03-21 23:11\npack-0d7b399006ae0a57ff3df07fdcaedbaeb7e63d0a.pack\n==\n\nI have a bf-yum remote defined, that pushes to the remote branch; once\nit gets there, I then do a merge on the target machine.\n\nThe 'killed by signal 2' is when I ctrl-c.\n\nThe second group was done from another window.  There's only a single\npack file now.\n\nThe @anon-site@ stuff is me removing client identifiers.  It's the\nonly editting I did to the screen log.\n\n> \n>> After I ran git push, ssh timed out, the temp pack that was created\n>> was then removed, as git complained about the connection being gone.\n> \n> On a push, there is no creation of a temp pack.  It is always produced \n> on the fly and pushed straight via the ssh connection.\n\nNo.  I saw a temp file in strace.  It *was* created on the local disk,\nand *not* sent on the fly.\n\n>> I then decided to do a 'git gc', which collapsed all the separate\n>> packs into one.  This allowed git push to proceed quickly, but at that\n>> point, it started sending the entire pack.\n> \n> If this was really the case, then this is definitely a bug.  Please take \n> a snapshot of your screen with git messages if this ever happens again.\n\nSee above.\n\n> \n>> It's entirely possible that the temp pack created by git push was\n>> incremental; it just took too long to create it, so it got aborted.\n> \n> The push operation has multiple phases.  You should see \"counting \n> objects\", \"compressing objects\" and \"writing objects\".  Could you give \n> us an approximation of how long each of those phases took?\n\nWell, counting was quick enough.  compression took at *least* 2 hours,\nmight have been 4 or more.  This all started friday evening.  I was\nwatching it a bit at the beginning, but then went out, and it died\nafter I got back to it.\n\n>> I forgot to mention previously, that the source machine was running\n>> git 1.5.6.5, and was pushing to 1.5.6.3.\n>>\n>> I've tried duplicating this problem on a machine with 1.6.1.3, but\n>> either I don't fully understand the issue enough to replicate it, or\n>> the newer git doesn't have the problem.\n> \n> That's possible.  Maybe others on the list might recall possible issues \n> related to this that might have been fixed during that time.\n\nWell, I looked at the release notes between all these versions.\nNothing stands out, but I'm aware that the changelog/release note\nentry for some change doesn't always describe the actual bug that\ncaused the change to occur.\n\n>> Um, if it's missing documentation, then how am I supposed to know\n>> about it?\n> \n> Asking on the list, like you did.  However this attribute should be \n> documented as well of course.  I even think that someone posted a patch \n> for it a while ago which might have been dropped.\n\nWhat I'd like, is a way to say a certain pattern of files should only\nbe deduped, and not deltafied.  This would handle the case of exact\ncopies, or renames, which would still be a win for us, but generally\nwhen a new video(or doc or pdf) is uploaded, it's alot of work to try\nand deltafy, for very little benefit.\n"},{"id":"109244","messageId":"20090324210427.GC30959@zoy.org","threadId":"18494","inReplyTo":"49C7FAB3.7080301@brainfood.com","subject":"Re: large(25G) repository in git","fromName":"Sam Hocevar","fromEmail":"sam@zoy.org","sentAt":"2009-03-24T21:04:28Z","receivedAt":"2009-03-24T21:04:28Z","isPatch":false,"sender":{"key":"sam@zoy.org","avatar":"https://gravatar.com/avatar/1fc1e5d8c3a8d737f14572135671adfbc0e61ffba5c3bc1d1b8a6f4aac764470?d=mp&s=160"},"body":"On Mon, Mar 23, 2009, Adam Heath wrote:\n> We maintain a website in git.  This website has a bunch of backend\n> server code, and a bunch of data files.  Alot of these files are full\n> videos.\n> \n> [...]\n> \n> Last friday, I was doing a checkin on the production server, and found\n> 1.6G of new files.  git was quite able at committing that.  However,\n> pushing was problematic.  I was pushing over ssh; so, a new ssh\n> connection was open to the preview server.  After doing so, git tried\n> to create a new pack file.  This took *ages*, and the ssh connection\n> died.  So did git, when it finally got done with the new pack, and\n> discovered the ssh connection was gone.\n\n   As stated several times by Linus and others, Git was not designed\nto handle large files. My stance on the issue is that before trying\nto optimise operations so that they perform well on large files, too,\nGit should usually avoid such operations, especially deltification.\nOne notable exception would be someone storing their mailbox in Git,\nwhere deltification is a major space saver. But usually, these large\nfiles are binary blobs that do not benefit from delta search (or even\ncompression).\n\n   Since I also need to handle large files (80 GiB repository), I am\ncleaning up some fixes I did, which can be seen in the git-bigfiles\nproject (http://caca.zoy.org/wiki/git-bigfiles). I have not yet tried\nto change git-push (because I submit through git-p4), but I hope to\naddress it, too. As time goes I believe some of them could make it into\nmainstream Git.\n\n   In your particular case, I would suggest setting pack.packSizeLimit\nto something lower. This would reduce the time spent generating a new\npack file if the problem were to happen again.\n\nRegards,\n-- \nSam.\n"},{"id":"109247","messageId":"49C95453.9080503@brainfood.com","threadId":"18494","inReplyTo":"20090324210427.GC30959@zoy.org","subject":"Re: large(25G) repository in git","fromName":"Adam Heath","fromEmail":"doogie@brainfood.com","sentAt":"2009-03-24T21:44:51Z","receivedAt":"2009-03-24T21:44:51Z","isPatch":false,"sender":{"key":"doogie@brainfood.com","avatar":"https://gravatar.com/avatar/c9e42ee14e1998b527796b1d8a35d75842d6a6aa885411ea1d5c97cb2d6923e6?d=mp&s=160"},"body":"Sam Hocevar wrote:\n>    As stated several times by Linus and others, Git was not designed\n> to handle large files. My stance on the issue is that before trying\n> to optimise operations so that they perform well on large files, too,\n> Git should usually avoid such operations, especially deltification.\n> One notable exception would be someone storing their mailbox in Git,\n> where deltification is a major space saver. But usually, these large\n> files are binary blobs that do not benefit from delta search (or even\n> compression).\n\nYeah, in this case, I *know* that my binary blobs are completely\ndifferent, and it's just a waste of time for git to come to the same\nconclusion.  I'd be perfectly willing to have some knob I could turn\nthat would tell git this.\n\n>    Since I also need to handle large files (80 GiB repository), I am\n> cleaning up some fixes I did, which can be seen in the git-bigfiles\n> project (http://caca.zoy.org/wiki/git-bigfiles). I have not yet tried\n> to change git-push (because I submit through git-p4), but I hope to\n> address it, too. As time goes I believe some of them could make it into\n> mainstream Git.\n\nI'd almost be willing to help.  I know the basic premise to how git\nworks, but the devil is in the details, and I don't have time right\nnow to learn the internals.\n\nYet another thing to add to my todo list.\n\n>    In your particular case, I would suggest setting pack.packSizeLimit\n> to something lower. This would reduce the time spent generating a new\n> pack file if the problem were to happen again.\n\nYeah, saw that one, but *after* I had this problem.  The default, if\nnot set, is unlimited, which in this case, is definately *not* what we\nwant.\n"},{"id":"109257","messageId":"49C9602F.2000306@brainfood.com","threadId":"18494","inReplyTo":"49C8A102.6090408@op5.se","subject":"Re: large(25G) repository in git","fromName":"Adam Heath","fromEmail":"doogie@brainfood.com","sentAt":"2009-03-24T22:35:27Z","receivedAt":"2009-03-24T22:35:27Z","isPatch":false,"sender":{"key":"doogie@brainfood.com","avatar":"https://gravatar.com/avatar/c9e42ee14e1998b527796b1d8a35d75842d6a6aa885411ea1d5c97cb2d6923e6?d=mp&s=160"},"body":"Andreas Ericsson wrote:\n> First of all, I'm going to hint that you would be far better off\n> keeping the media files in a separate repository, linked in as a\n> submodule in git and with tweaked configuration settings with the\n> specific aim of handling huge files.\n\nAlready do that.  We have a custom overlay/union-type filesystem, that\nmakes use of a small base directory, where code resides, then each\nsub-website is where the content is.\n\nIt's just finding documentation thru google that describes the\nworkflow we are doing is difficult.\n\n> The basis of such a repository is probably the following config\n> settings, since media files very rarely compress enough to be\n> worth the effort, and their own compressed formats make them\n> very unsuitable delta candidates:\n> [pack]\n>   # disable delta-based packing\n>   depth = 1\n>   # disable compression\n>   compression = 0\n> \n> [gc]\n>   # don't auto-pack, ever\n>   auto = 0\n>   # never automatically consolidate un-.keep'd packs\n>   autopacklimit = 0\n\nThanks for the pointers!\n\n> You will have to manually repack this repository from time to\n> time, and it's almost certainly a good idea to mark the\n> resulting packs with .keep to avoid copying tons of data.\n> When packs are being created, objects can be copied from\n> existing packs, and send-pack will make use of that so that what\n> goes over the wire will simply be copied from the existing packs.\n> \n> YMMV. If you do come up with settings that work fine for huge\n> repos made up of mostly media files, please share your findings.\n\nI'll use these as a basis.\n\n>> So, to work around that, I ran git gc.  When done, I discovered that\n>> git repacked the *entire* repository.  While not something I care for,\n>> I can understand that, and live with it.  It just took *hours* to do so.\n>>\n> \n> I'm not sure what, if any, magic \"git gc\" applies before spawning\n> \"git repack\", but running \"git repack\" directly would almost certainly\n> have produced an incremental pack. Perhaps we need to make gc less\n> magic.\n\nThe repo should only be converted into a single .pack, if the user\nexplicitily wants it.  Any automatic gc call, or called without args,\nshould just take any loose objects and pack them up.  But that's my\nopinion.\n\n> Not necessarily all that simple (we do not want to touch the ssh\n> password if we can possibly avoid it, but the user shouldn't have\n> to type it more than once), but certainly doable. Easier would\n> probably be to recommend adding the proper SSH config variables,\n> as has been stated elsewhere.\n\nssh-agent, or password-less anonymous ssh(I've got a custom login\nscript inside authorized_keys on the remote).\n\n> See above. I *think* you can also do this with git-attributes, but\n> I'm not sure. However, keeping the large media files in a sub-module\n> would nicely solve that problem anyway, and is probably a good idea\n> even with git-attributes support for pack delta- and compression\n> settings.\n\nThe site would *still* be > 25G in size, at the least, and constantly\ngetting bigger.  This site contains copies of ad videos from their\ncompetitors, plus their own, and is used to market their international\ncompany.\n\n> http://www.thousandparsec.net/~tim/media+git.pdf probably holds all the\n> relevant information when it comes to storing large media files with\n> git. I have not checked and have no inclination to do so.\n\nhttp://caca.zoy.org/wiki/git-bigfiles is another one.\n"},{"id":"109279","messageId":"alpine.LFD.2.00.0903242025110.26337@xanadu.home","threadId":"18494","inReplyTo":"49C95453.9080503@brainfood.com","subject":"Re: large(25G) repository in git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-03-25T00:28:33Z","receivedAt":"2009-03-25T00:28:33Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 24 Mar 2009, Adam Heath wrote:\n\n> Sam Hocevar wrote:\n> >    In your particular case, I would suggest setting pack.packSizeLimit\n> > to something lower. This would reduce the time spent generating a new\n> > pack file if the problem were to happen again.\n> \n> Yeah, saw that one, but *after* I had this problem.  The default, if\n> not set, is unlimited, which in this case, is definately *not* what we\n> want.\n\nIn your particular case, if the problem is actually what I think it is, \nthe pack.packSizeLimit wouldn't have made any difference.  This setting \naffects local repacking only and has no effect what so ever on the push \noperation.\n\n\nNicolas\n"},{"id":"109285","messageId":"49C9818A.9040507@brainfood.com","threadId":"18494","inReplyTo":"alpine.LFD.2.00.0903242025110.26337@xanadu.home","subject":"Re: large(25G) repository in git","fromName":"Adam Heath","fromEmail":"doogie@brainfood.com","sentAt":"2009-03-25T00:57:46Z","receivedAt":"2009-03-25T00:57:46Z","isPatch":false,"sender":{"key":"doogie@brainfood.com","avatar":"https://gravatar.com/avatar/c9e42ee14e1998b527796b1d8a35d75842d6a6aa885411ea1d5c97cb2d6923e6?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n> On Tue, 24 Mar 2009, Adam Heath wrote:\n> \n>> Sam Hocevar wrote:\n>>>    In your particular case, I would suggest setting pack.packSizeLimit\n>>> to something lower. This would reduce the time spent generating a new\n>>> pack file if the problem were to happen again.\n>> Yeah, saw that one, but *after* I had this problem.  The default, if\n>> not set, is unlimited, which in this case, is definately *not* what we\n>> want.\n> \n> In your particular case, if the problem is actually what I think it is, \n> the pack.packSizeLimit wouldn't have made any difference.  This setting \n> affects local repacking only and has no effect what so ever on the push \n> operation.\n\nOoh.  Care to enlighten those of us not blessed with git internal\nknowledge?\n\nOn another note, anyone have a goat I can buy, for the sacrifice?\n"},{"id":"109286","messageId":"alpine.LFD.2.00.0903241709090.26337@xanadu.home","threadId":"18494","inReplyTo":"49C948C1.2070404@brainfood.com","subject":"Re: large(25G) repository in git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-03-25T01:21:19Z","receivedAt":"2009-03-25T01:21:19Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 24 Mar 2009, Adam Heath wrote:\n\n> Nicolas Pitre wrote:\n> > As much as I would like to believe you, this doesn't help fixing the \n> > problem if you don't provide more information about this.  For example, \n> > the output from git during the whole operation might give us the \n> > beginning of a clue.  Otherwise, all I can tell you is that such thing \n> > is not supposed to happen.\n> \n> First off, you've put a bad tone on this.  It appears that you are\n> saying I'm mistaken, and it didn't send all that data.  \"It can't\n> happen, so it didn't happen.\"  Believe me, if it hadn't resent all\n> this data, I wouldn't have even sent the email.\n\nI don't know you.  All I had is the information you provided which was \nrather incomplete.  So don't be offended if I ask for more.  I'm trying \nto help you after all.\n\nAnd especially in this case, the problem seems not to be about \npacking...\n\n> In any event, we got lucky.  I *do* have a log of the push side of\n> this problem.  I doubt it's enough to figure out the actual cause tho.\n\nWell, I think it might.\n\n> ==\n> Counting objects: 96637, done.\n> Compressing objects: 100% (29670/29670), done.\n> Writing objects: 100% (96637/96637), 25.49 GiB | 226 KiB/s, done.\n> Total 96637 (delta 48713), reused 96637 (delta 48713)\n> To ssh://bf-yum/@anon-site@\n>  * [new branch]      master -> lnxwww10\n\nWas that branch really new on the remote side?  If no, then this is \nhighly suspicious.  If somehow the previously aborted push attempt \nscrewed the remote refs, then the local client would think that the \nremote is empty and conclude that all commits have to be pushed.\n\n> >> After I ran git push, ssh timed out, the temp pack that was created\n> >> was then removed, as git complained about the connection being gone.\n> > \n> > On a push, there is no creation of a temp pack.  It is always produced \n> > on the fly and pushed straight via the ssh connection.\n> \n> No.  I saw a temp file in strace.  It *was* created on the local disk,\n> and *not* sent on the fly.\n\nA temp pack is created on the receiving side, not the sending side \nthough.  The sending side is piping the pack data on its standard output \nwhich is connected to ssh's standard input.\n\n> >> Um, if it's missing documentation, then how am I supposed to know\n> >> about it?\n> > \n> > Asking on the list, like you did.  However this attribute should be \n> > documented as well of course.  I even think that someone posted a patch \n> > for it a while ago which might have been dropped.\n> \n> What I'd like, is a way to say a certain pattern of files should only\n> be deduped, and not deltafied.  This would handle the case of exact\n> copies, or renames, which would still be a win for us, but generally\n> when a new video(or doc or pdf) is uploaded, it's alot of work to try\n> and deltafy, for very little benefit.\n\nRenamed/duplicated files are always stored uniquely by design.  Git \nstore file data into objects which are named after the SHA1 of their \ncontent.\n\nIn order to not attempt any delta on PDF files for example, you need to \nadd a negative delta attribute line such as:\n\n*.pdf\t-delta\n\neither in a file called .gitattributes which gets versionned \nand distributed, or in .git/info/attributes in which case it'll remain \nlocal.  Any file matching *.pdf won't be delta compressed.\n\n\nNicolas\n"},{"id":"109288","messageId":"alpine.LFD.2.00.0903242124470.26337@xanadu.home","threadId":"18494","inReplyTo":"49C9818A.9040507@brainfood.com","subject":"Re: large(25G) repository in git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-03-25T01:47:34Z","receivedAt":"2009-03-25T01:47:34Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 24 Mar 2009, Adam Heath wrote:\n\n> Nicolas Pitre wrote:\n> > On Tue, 24 Mar 2009, Adam Heath wrote:\n> > \n> >> Sam Hocevar wrote:\n> >>>    In your particular case, I would suggest setting pack.packSizeLimit\n> >>> to something lower. This would reduce the time spent generating a new\n> >>> pack file if the problem were to happen again.\n> >> Yeah, saw that one, but *after* I had this problem.  The default, if\n> >> not set, is unlimited, which in this case, is definately *not* what we\n> >> want.\n> > \n> > In your particular case, if the problem is actually what I think it is, \n> > the pack.packSizeLimit wouldn't have made any difference.  This setting \n> > affects local repacking only and has no effect what so ever on the push \n> > operation.\n> \n> Ooh.  Care to enlighten those of us not blessed with git internal\n> knowledge?\n\nSee my previous email for a likely explanation about your issue.\n\nAs to the pack.packSizeLimit setting: it is used when repacking only in \norder to avoid big packs on systems that might have issues dealing with \nlarge files.  During a repack, if the currently produced pack is about \nto get over that limit, then the pack is closed and a new one is \nstarted.  You therefore end up with many packs.\n\nThe transfer protocol used during a fetch or a push uses the pack format \nstreamed over the network, but only one pack can be transferred that \nway.  Maybe the reception of a pack during a network transfer should be \nsplit according to pack.packSizeLimit as well, but this is currently not \nimplemented at all. No one complained about that either, so I'm guessing \nthat \nsplitting a large pack, if needed, by using 'git repack' after a \nclone/fetch is good enough.\n\nPersonally, I don't think actively splitting packs into smaller ones is \nthat useful, unless you wish to archive them on a file system which \ncannot handle files larger than 2GB or the like.\n\n> On another note, anyone have a goat I can buy, for the sacrifice?\n\nBeware the wrath of Git...\n\n\nNicolas\n"},{"id":"109546","messageId":"49CBA2AB.30304@oak.homeunix.org","threadId":"18494","inReplyTo":"49C7FAB3.7080301@brainfood.com","subject":"Re: large(25G) repository in git","fromName":"Marcel M. Cary","fromEmail":"marcel@oak.homeunix.org","sentAt":"2009-03-26T15:43:39Z","receivedAt":"2009-03-26T15:43:39Z","isPatch":false,"sender":{"key":"marcel@oak.homeunix.org","avatar":"https://gravatar.com/avatar/2bb524e4f383167b7e256bb93256c88353748d9873c34cde0fd461f1165baa0f?d=mp&s=160"},"body":"Adam Heath wrote:\n> We maintain a website in git.  This website has a bunch of backend\n> server code, and a bunch of data files.  Alot of these files are full\n> videos.\n>\n> We use git, so that the distributed nature of website development can\n> be supported.  Quite often, you'll have a production server, with\n> online changes occurring(we support in-browser editting of content), a\n> preview server, where large-scale code changes can be previewed, then\n> a development server, one per programmer(or more).\n\nMy company manages code in a similar way, except we avoid this kind of\nissue (with 100 gigabytes of user-uploaded images and other data) by not\nchecking in the data.  We even went so far is as to halve the size of\nour repository by removing 2GB of non-user-supplied images -- rounded\ncorners, background gradients, logos, etc, etc.  This made Git\nnoticeably faster.\n\nWhile I'd love to be able to handle your kind of use case and data size\nwith Git in that way, it's a little beyond the intended usage to handle\nhundreds of gigabytes of binary data, I think.\n\nI imagine as your web site grows, which I'm assuming is your goal, your\nproblems with scaling Git will continue to be a challenge.\n\nMaybe you can find a way to:\n\n* Get along with less data in your non-production environments; we're\nhoping to be able to do this eventually\n\n* Find other ways to copy it; we use rsync even though it does take\nforever to crawl over the file system\n\n* Put your data files in a separate Git repository, at least, assuming\nyour checkin, update, and release code more often than your video files.\n That way you'll experience pain less often, and maybe even be able to\ntune your repository differently.\n\nMarcel\n"},{"id":"109549","messageId":"49CBAEC5.6070606@brainfood.com","threadId":"18494","inReplyTo":"49CBA2AB.30304@oak.homeunix.org","subject":"Re: large(25G) repository in git","fromName":"Adam Heath","fromEmail":"doogie@brainfood.com","sentAt":"2009-03-26T16:35:17Z","receivedAt":"2009-03-26T16:35:17Z","isPatch":false,"sender":{"key":"doogie@brainfood.com","avatar":"https://gravatar.com/avatar/c9e42ee14e1998b527796b1d8a35d75842d6a6aa885411ea1d5c97cb2d6923e6?d=mp&s=160"},"body":"Marcel M. Cary wrote:\n> My company manages code in a similar way, except we avoid this kind of\n> issue (with 100 gigabytes of user-uploaded images and other data) by not\n> checking in the data.  We even went so far is as to halve the size of\n> our repository by removing 2GB of non-user-supplied images -- rounded\n> corners, background gradients, logos, etc, etc.  This made Git\n> noticeably faster.\n\nDisk space is cheap.\n\n> While I'd love to be able to handle your kind of use case and data size\n> with Git in that way, it's a little beyond the intended usage to handle\n> hundreds of gigabytes of binary data, I think.\n> \n> I imagine as your web site grows, which I'm assuming is your goal, your\n> problems with scaling Git will continue to be a challenge.\n> \n> Maybe you can find a way to:\n> \n> * Get along with less data in your non-production environments; we're\n> hoping to be able to do this eventually\n\nWe do that by only cloning/checking out certain modules.\n\nHowever, as is always the case, sometimes a bug occurs with production\ndata, and you need to use the real data to track it down.\n\n> * Find other ways to copy it; we use rsync even though it does take\n> forever to crawl over the file system\n> \n> * Put your data files in a separate Git repository, at least, assuming\n> your checkin, update, and release code more often than your video files.\n>  That way you'll experience pain less often, and maybe even be able to\n> tune your repository differently.\n\nAs already mentioned, our sub-sites *are* in separate repos.  There's\na base repository, that has just the event/backend code.  Then 32\n*other* repositories, where the actual websites are.\n\nWe want to use *some* kind of versioning system.  Being able to have\nhistory of *all* changes is extremely useful.  Not to mention being\nable to track what each separate user does as they modify their files\nthru their browser.\n\nsubversion is just right out.  It's centralized.  It leaves poop all\nover the place.\n\nmercurial is just right out.  If you do several *separate* commits of\n*separate* files, but don't push for some time period, then eventually\ndo a push/pull, where the sum total of the changes is larger than some\nvalue, mercurial will fail when it tries to then update the local\ndirectory.  This limit is based on 2G, a hard-coded python limit(even\non a 64-bit host), because mercurial reads the entire set of changes\ninto a python string.\n\ngit mmaps files, does window scanning of the pack files.  It *might*\nread a single file all into memory, for compression purposes; I'm not\ncertain on this.  We certainly haven't hit any limits that cause it to\nfail outright.\n\nI haven't tried any others.\n"}]}