{"thread":{"id":"28201","subject":"git for game development?","startedAt":"2011-08-23T23:06:47Z","lastAt":"2011-08-27T15:32:29Z","messageCount":9,"participants":["Lawrence Brett","Junio C Hamano","Jeff King","Marat Radchenko","J.H.","Michael Witten"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"174129","messageId":"416D1A48-9916-4E44-A200-3A13C39C4D70@gmail.com","threadId":"28201","inReplyTo":null,"subject":"git for game development?","fromName":"Lawrence Brett","fromEmail":"lcbrett@gmail.com","sentAt":"2011-08-23T23:06:47Z","receivedAt":"2011-08-23T23:06:47Z","isPatch":false,"sender":{"key":"lcbrett@gmail.com","avatar":null},"body":"Hello,\n\nI am very interested in using git for game development.  I will be working\nwith a lot of binaries (textures, 3d assets, etc.) in addition to source\nfiles.  I'd like to be able to version these files, but I understand that\nbig binaries aren't git's forte.  I've found several possible workarounds\n(git submodules, git-media, git-annex), but the one that seems most\npromising is bup.  I started a thread on the bup mailing list to ask about\nthe best way to use bup with git for my purposes.  One of the respondents\nsuggested forking git itself to include bup functionality, thereby extending\ngit to handle binaries efficiently.\n\nMy question for this group is:  would there be interest in incorporating\nthis sort of functionality into git core?  I would certainly find it\ncompelling as a user, but have no idea how it would fit into the bigger\npicture.\n\nThanks in advance!\n\nCliff\n\nP.S.  I also heartily welcome any advice/insight on my use case.  :-)\n"},{"id":"174131","messageId":"7vaaazr9fu.fsf@alter.siamese.dyndns.org","threadId":"28201","inReplyTo":"416D1A48-9916-4E44-A200-3A13C39C4D70@gmail.com","subject":"Re: git for game development?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-08-23T23:32:05Z","receivedAt":"2011-08-23T23:32:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lawrence Brett <lcbrett@gmail.com> writes:\n\n> My question for this group is:  would there be interest in incorporating\n> this sort of functionality into git core?  I would certainly find it\n> compelling as a user, but have no idea how it would fit into the bigger\n> picture.\n\nI personally think it is too early for you to ask that question; until you\nset up a workable workflow around bup or a combination of bup and git, get\nused to its use, and find out what the real pain points are if you used\nonly git without bup, that is.\n\nEfforts to tweak tools by people who are not yet familiar with the tools\nthey are trying to use unfortunately often tend to go in wrong directions\nand become wasted effort.\n"},{"id":"174135","messageId":"20110824012418.GA19091@sigill.intra.peff.net","threadId":"28201","inReplyTo":"416D1A48-9916-4E44-A200-3A13C39C4D70@gmail.com","subject":"Re: git for game development?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-08-24T01:24:18Z","receivedAt":"2011-08-24T01:24:18Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Aug 23, 2011 at 04:06:47PM -0700, Lawrence Brett wrote:\n\n> I am very interested in using git for game development.  I will be working\n> with a lot of binaries (textures, 3d assets, etc.) in addition to source\n> files.  I'd like to be able to version these files, but I understand that\n> big binaries aren't git's forte.  I've found several possible workarounds\n> (git submodules, git-media, git-annex), but the one that seems most\n> promising is bup.  I started a thread on the bup mailing list to ask about\n> the best way to use bup with git for my purposes.  One of the respondents\n> suggested forking git itself to include bup functionality, thereby extending\n> git to handle binaries efficiently.\n> \n> My question for this group is:  would there be interest in incorporating\n> this sort of functionality into git core?  I would certainly find it\n> compelling as a user, but have no idea how it would fit into the bigger\n> picture.\n\nSomething bup-like in git-core might eventually be good. But IIRC, bup\nintroduces new object types, which mixes the abstract view of the data\nformat (i.e., commits, trees, and blobs indexed by sha1) with the\nimplementation details (e.g., now we have both loose objects in their\nown files as well as delta-compressed objects in packfiles).\n\nThat means that bup-git clients and non-bup git clients don't interact\nvery well. Where non-bup is either a client that doesn't understand the\nbup objects, or one that chooses not to use bup-like encoding for\nparticular blobs.\n\nI don't remember all of the details of bup, but if it's possible to\nimplement something similar at a lower level (i.e., at the layer of\npackfiles or object storage), then it can be a purely local thing, and\nthe compatibility issues can go away.\n\n-Peff\n\nPS I also agree with Junio's comment that we are not at the \"planning a\nsolution\" stage with big files, but rather at the \"trying it and getting\nexperience on what works and what doesn't\" stage.\n"},{"id":"174167","messageId":"7vwre2pw3m.fsf@alter.siamese.dyndns.org","threadId":"28201","inReplyTo":"20110824012418.GA19091@sigill.intra.peff.net","subject":"Re: git for game development?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-08-24T17:17:49Z","receivedAt":"2011-08-24T17:17:49Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> I don't remember all of the details of bup, but if it's possible to\n> implement something similar at a lower level (i.e., at the layer of\n> packfiles or object storage), then it can be a purely local thing, and\n> the compatibility issues can go away.\n\nI tend to agree, and we might be closer than we realize.\n\nI suspect that people with large binary assets were scared away by rumors\nthey heard second-hand, based on bad experiences other people had before\nany of the recent efforts made in various \"large Git\" topics, and they\nthemselves haven't tried recent versions of Git enough to be able to tell\nwhat the remaining pain points are. I wouldn't be surprised if none of the\ncore Git people tried shoving huge binary assets in test repositories with\nrecent versions of Git---I certainly haven't.\n\nWe used to always map the blob data as a whole for anything we do, but\nthese days, with changes like your abb371a (diff: don't retrieve binary\nblobs for diffstat, 2011-02-19) and my recent \"send large blob straight to\na new pack\" and \"stream large data out to the working tree without holding\neverything in core while checking out\" topics, I suspect that the support\nfor local usage of large blobs might be sufficiently better than the old\ndays. Git might even be usable locally without anything else, which I find\nimplausible, but I wouldn't be surprised if there remained only a handful\nminor things remaining that we need to add to make it usable.\n\nPeople toyed around with ideas to have a separate object store\nrepresentation for large and possibly incompressible blobs (a possible\ncomplaint being that it is pointless to send them even to its own\npackfile). One possible implementation would be to add a new huge\nhierarchy under $GIT_DIR/objects/, compute the object name exactly the\nsame way for huge blobs as we normally would (i.e. hash concatenation of\nobject header and then contents) to decide which subdirectory under the\n\"huge\" hierarchy to store the data (huge/[0-9a-f]{2}/[0-9a-f]{38}/ like we\ndo for loose objects, or perhaps huge/[0-9a-f]{40}/ expecting that there\nwon't be very many). The data can be stored unmodified as a file in that\ndirectory, with type stored in a separate file---that way, we won't have\nto compress, but we just copy. You still need to hash it at least once to\ncome up with the object name, but that is what gives us integrity checks,\nis unavoidable and is not going to change.\n\nThe sha1_object_info() layer can learn to return the type and size from\nsuch a representation, and you can further tweak the same places as the\n\"streaming checkout\" and the \"checkin to a pack\" topics touched to support\nsuch a representation.\n\nI would suspect that the local object representation is _not_ the largest\npain point; such a separate object store representation is not buying us\nvery much over a simpler \"single large blob in a separate packfile\", and\nif the counter-argument is \"no, decompressing still costs a lot\", then the\nreal issue might be we decompress and look at the data when we do not have\nto (i.e. issues similar to what abb371a addressed), not \"decompress vs\nstraight copy make a bit difference\".\n\nI would further suspect that we _might_ need a better support for local\nrepacking and object transfer, with or without such a third object\nrepresentation.\n"},{"id":"174171","messageId":"20110824182632.GA22659@sigill.intra.peff.net","threadId":"28201","inReplyTo":"7vwre2pw3m.fsf@alter.siamese.dyndns.org","subject":"Re: git for game development?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-08-24T18:26:32Z","receivedAt":"2011-08-24T18:26:32Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Aug 24, 2011 at 10:17:49AM -0700, Junio C Hamano wrote:\n\n> I suspect that people with large binary assets were scared away by rumors\n> they heard second-hand, based on bad experiences other people had before\n> any of the recent efforts made in various \"large Git\" topics, and they\n> themselves haven't tried recent versions of Git enough to be able to tell\n> what the remaining pain points are. I wouldn't be surprised if none of the\n> core Git people tried shoving huge binary assets in test repositories with\n> recent versions of Git---I certainly haven't.\n\nI haven't tried anything really big in a while. My personal interest in\nbig file support has been:\n\n  1. Mid-sized photos and videos (objects top out around 50M, total repo\n     size is 4G packed). Most commits are additions or tweaks of exif\n     tags (so they delta well). Using gitattributes (and especially\n     textconv caching), it's really quite pleasant to use. Doing a full\n     repack is my only complaint; the delta-compression isn't bad, but\n     just the I/O on rewriting the whole thing is a killer.\n\n  2. Storing an entire audio collection in flac. Median file size is\n     only around 20M, but the whole repo is 120G.  Obviously compression\n     doesn't buy much, so a git repo plus checkout is 240G, which is\n     pretty hefty for most laptops. I played with this early on, but\n     gave up; the data storage model just doesn't make sense.\n\nThe two common use cases that aren't represented here are:\n\n  3. Big files, not just big repos. I.e., files that are 1G or more.\n\n  4. Medium-big files that don't delta well (e.g., metadata tweaks do\n     delta well; rewriting media assets for a game don't delta well).\n\nI think recent changes (like putting big files straight to packs) make\n(3) and (4) reasonably pleasant.\n\nI'm not sure of the right answer for (1). The repack is the only\nannoying thing. But not repacking is not satisfying, either.  You don't\nget deltas where they are applicable, and the server is always\nre-examining the pack for possible deltas on fetch and push. Some sort\nof hybrid loose-pack storage would be nice: store delta chains for big\nfiles in their own individual packs, but otherwise keep everything in a\nseparate pack. We would want some kind of meta-index over all of these\nlittle pack-files, not just individual pack-file indices.\n\nBut (2) is the hardest one. It would be nice if we had some kind of\nlocal-remote hybrid storage, where objects were fetched on demand from\nsomewhere else. For example, developers on workstations with a fast\nlocal network to a storage server wouldn't have to replicate all of the\nobjects locally. And for a true distributed setup, when the fast network\nisn't there, it would be nice to fail gracefully (which maybe just means\nsaying \"sorry, we can't do 'log -p' right now; try 'log --raw'\").\n\nI wonder how close one can get on (2) using alternates and a\nnetwork-mounted filesystem.\n\n> People toyed around with ideas to have a separate object store\n> representation for large and possibly incompressible blobs (a possible\n> complaint being that it is pointless to send them even to its own\n> packfile). One possible implementation would be to add a new huge\n> hierarchy under $GIT_DIR/objects/, compute the object name exactly the\n> same way for huge blobs as we normally would (i.e. hash concatenation of\n> object header and then contents) to decide which subdirectory under the\n> \"huge\" hierarchy to store the data (huge/[0-9a-f]{2}/[0-9a-f]{38}/ like we\n> do for loose objects, or perhaps huge/[0-9a-f]{40}/ expecting that there\n> won't be very many). The data can be stored unmodified as a file in that\n> directory, with type stored in a separate file---that way, we won't have\n> to compress, but we just copy. You still need to hash it at least once to\n> come up with the object name, but that is what gives us integrity checks,\n> is unavoidable and is not going to change.\n\nYeah. I think one of the bonuses there is that some filesystems are\ncapable of referencing the same inodes in a copy-on-write way, so \"add\"\nand \"checkout\" cease to be a copy operation, but rather an inode-linking\noperation. Which is a big win, both for speed and storage.\n\nI've had dreams of using hard-linking to do something similar, but it's\njust not safe enough without some filesystem-level copy-on-write\nprotection.\n\n-Peff\n"},{"id":"174209","messageId":"loom.20110825T081519-218@post.gmane.org","threadId":"28201","inReplyTo":"416D1A48-9916-4E44-A200-3A13C39C4D70@gmail.com","subject":"One MMORPG git facts","fromName":"Marat Radchenko","fromEmail":"marat@slonopotamus.org","sentAt":"2011-08-25T06:53:57Z","receivedAt":"2011-08-25T06:53:57Z","isPatch":false,"sender":{"key":"marat@slonopotamus.org","avatar":"https://avatars.githubusercontent.com/u/92637?v=4"},"body":"Lawrence Brett <lcbrett <at> gmail.com> writes:\n\n> \n> Hello,\n> \n> I am very interested in using git for game development.  I will be working\n> with a lot of binaries (textures, 3d assets, etc.) in addition to source\n> files.  I'd like to be able to version these files, but I understand that\n> big binaries aren't git's forte.\n\nDefine \"big\".\n\nI have one MMORPG here under Git. 250k revisions, 500k files in working dir\n(7Gb), 200 commits daily, 250Gb Git repo, SVN upstream repo of ~1Tb.\n\nSome facts:\n1. It is unusable on 32bit machine (here and there hits memory limit for a\nsingle process\n2. It is unusable on Windows (because there's no 64bit msysgit)\n3. git status is 3s with hot disk caches (7mins with cold)\n4. History traversal means really massive I/O.\n5. Current setup: 120Gb 10k rpm disk for everything but .git/objects/pack,\nseparate 500Gb (will be upgraded to 1Tb soon) disk for packs\n6. git gc is PAIN. I do it on weekends because it takes more than a day to run.\nAlso, limits for git pack-objects should be configured VERY carefully, it can\neither run out of ram or take weeks to run if configured improperly.\n7. With default gc settings, git wants to gc daily (but gc takes more than a\nday, so if you follow its desire, you're in gc loop). I set objects limit to a\nvery high value and invoke gc manually.\n8. svn users cannot sensibly do status on whole working copy (more than 10 mins)\n9. svn users only update witha nightly script (40 mins)\n10. git commit is several seconds because it writes 70Mb commit file.\n11. It is a good idea to run git status often so that working copy info isn't\nevicted from OS disk caches (remember, 3s vs 7min)\n12. Cloning git repo is one more pain. 100mbps network here, so fetching 250Gb\ntakes some time. But worse, if cloning via git:// protocol, after fetching git\nsits for several hours in \"Resolving deltas\" stage. So, for initial cloning\nrsync is used.\n13. Here and there i hit scalability issues in various git commands (which i\nreport to maillist and most [well, all, except the one i reported this week] of\nwhich get fixed)\n\nHope this helps to get the idea of how git behaves on a large scale. Overall,\ni'm happy with it and won't return to svn.\n"},{"id":"174214","messageId":"4E560053.1080005@eaglescrag.net","threadId":"28201","inReplyTo":"loom.20110825T081519-218@post.gmane.org","subject":"Re: One MMORPG git facts","fromName":"J.H.","fromEmail":"warthog9@eaglescrag.net","sentAt":"2011-08-25T07:57:07Z","receivedAt":"2011-08-25T07:57:07Z","isPatch":false,"sender":{"key":"warthog9@kernel.org","avatar":"https://avatars.githubusercontent.com/u/2334704?v=4"},"body":"On 08/24/2011 11:53 PM, Marat Radchenko wrote:\n> Lawrence Brett <lcbrett <at> gmail.com> writes:\n> \n>>\n>> Hello,\n>>\n>> I am very interested in using git for game development.  I will be working\n>> with a lot of binaries (textures, 3d assets, etc.) in addition to source\n>> files.  I'd like to be able to version these files, but I understand that\n>> big binaries aren't git's forte.\n> \n> Define \"big\".\n> \n> I have one MMORPG here under Git. 250k revisions, 500k files in working dir\n> (7Gb), 200 commits daily, 250Gb Git repo, SVN upstream repo of ~1Tb.\n\nGiven the differences, I'm morbidly curious, which actually ends up\nbeing the more usable version control system of a project of this scale?\n It sounds like (from what you've said) git is generally faster,\nassuming it can get enough resources (which can obviously be hard at the\nscales your talking).\n\n- John 'Warthog9' Hawley\n"},{"id":"174234","messageId":"1314288121.8665.2.camel@n900.home.ru","threadId":"28201","inReplyTo":"4E560053.1080005@eaglescrag.net","subject":"Re: One MMORPG git facts","fromName":"Marat Radchenko","fromEmail":"marat@slonopotamus.org","sentAt":"2011-08-25T16:02:01Z","receivedAt":"2011-08-25T16:02:01Z","isPatch":false,"sender":{"key":"marat@slonopotamus.org","avatar":"https://avatars.githubusercontent.com/u/92637?v=4"},"body":"On 08/25/2011 11:57:07 MSD, J.H. <warthog9@eaglescrag.net> wrote:\n> Given the differences, I'm morbidly curious, which actually ends up\n> being the more usable version control system of a project of this scale?\n>   It sounds like (from what you've said) git is generally faster,\n> assuming it can get enough resources (which can obviously be hard at the\n> scales your talking).\n\nHard to compare (especially because I don't have pure git environment but git-svn clone). I have give you some \n\nFirst, there are lots of non-geek people working on MMORPG (quest designers, modellers, text writers, map designers). Many of them find it hard to understand DVCS concepts and prefer living with linear history in a single branch (svn trunk). Their work is highly isolated from each other (for ex, maps are split in \"regions\" and only one person is allowed to edit one region simultanuously, only one modeller works on a particular model, each quest has a person responsible for it) so they don't hit conflicts as often as programmers do. And since svn up of whole tree takes 40 mins, they don't update during work day but have nightly script for that so the only thing they regularly use is svn commit.\n\nSecond, there's TortoiseSVN that allows easy (for non-geeks) GUI history inspection.\n\nThird, we have 200 commits per day (8 work hours), that's one commit each 2.4 mins (actually, much less during lunch and much more in the morning caused by the fact that programmers are not allowed to commit after 16:00), so you copy is outdated all the time. If upstream repo was git, one would have to pull + push in those 2.4 mins, otherwise she would hit non-ff push. This could be fixed by using separate repos, though that would complicate git setup even more.\n\nOn the other hand, git is really great for programmers. Heck, svn still doesn't have anything like \"git log -u\" (well, afaik, they finally added it in 1.7)! Stash/bisect/local commits/history rewriting/cheap branching(almost no branching happens in svn repo because that either involves either fetching of 7Gb or [if svn switch is used] 10-30mins [don't remember exactly] of massive I/O. [sorry for two-level nesting]) are very handy on daily basis. Also, git allows easy sharing of experimental changes between programmers without touching shared server. It also allows atomic commits for the whole working copy (which is important when programmer changes server, client or data at the same time).\n\nThere are some decisions that were made and without which repo could be smaller (for example, client/server binaries are commited daily so that designers can use them on next day), however these decisions were made long before i joined the project and are very likely to stay this way.\n\nTo sum this up: git is a wonderful (and very powerful) tool for programmers but too complex for non-tech users.\n"},{"id":"174380","messageId":"CAMOZ1BsiFdSwhi2xMx7_-hsKYccUTf09W-4UpK8CwQjqY4cpig@mail.gmail.com","threadId":"28201","inReplyTo":"7vwre2pw3m.fsf@alter.siamese.dyndns.org","subject":"Re: git for game development?","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2011-08-27T15:32:29Z","receivedAt":"2011-08-27T15:32:29Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Wed, Aug 24, 2011 at 17:17, Junio C Hamano <gitster@pobox.com> wrote:\n> Jeff King <peff@peff.net> writes:\n>\n>> I don't remember all of the details of bup, but if it's possible to\n>> implement something similar at a lower level (i.e., at the layer of\n>> packfiles or object storage), then it can be a purely local thing, and\n>> the compatibility issues can go away.\n>\n> I tend to agree, and we might be closer than we realize.\n>\n> I suspect that people with large binary assets were scared away by rumors\n> they heard second-hand, based on bad experiences other people had before\n> any of the recent efforts made in various \"large Git\" topics, and they\n> themselves haven't tried recent versions of Git enough to be able to tell\n> what the remaining pain points are. I wouldn't be surprised if none of the\n> core Git people tried shoving huge binary assets in test repositories with\n> recent versions of Git---I certainly haven't.\n>\n> We used to always map the blob data as a whole for anything we do, but\n> these days, with changes like your abb371a (diff: don't retrieve binary\n> blobs for diffstat, 2011-02-19) and my recent \"send large blob straight to\n> a new pack\" and \"stream large data out to the working tree without holding\n> everything in core while checking out\" topics, I suspect that the support\n> for local usage of large blobs might be sufficiently better than the old\n> days. Git might even be usable locally without anything else, which I find\n> implausible, but I wouldn't be surprised if there remained only a handful\n> minor things remaining that we need to add to make it usable.\n>\n> People toyed around with ideas to have a separate object store\n> representation for large and possibly incompressible blobs (a possible\n> complaint being that it is pointless to send them even to its own\n> packfile). One possible implementation would be to add a new huge\n> hierarchy under $GIT_DIR/objects/, compute the object name exactly the\n> same way for huge blobs as we normally would (i.e. hash concatenation of\n> object header and then contents) to decide which subdirectory under the\n> \"huge\" hierarchy to store the data (huge/[0-9a-f]{2}/[0-9a-f]{38}/ like we\n> do for loose objects, or perhaps huge/[0-9a-f]{40}/ expecting that there\n> won't be very many). The data can be stored unmodified as a file in that\n> directory, with type stored in a separate file---that way, we won't have\n> to compress, but we just copy. You still need to hash it at least once to\n> come up with the object name, but that is what gives us integrity checks,\n> is unavoidable and is not going to change.\n>\n> The sha1_object_info() layer can learn to return the type and size from\n> such a representation, and you can further tweak the same places as the\n> \"streaming checkout\" and the \"checkin to a pack\" topics touched to support\n> such a representation.\n>\n> I would suspect that the local object representation is _not_ the largest\n> pain point; such a separate object store representation is not buying us\n> very much over a simpler \"single large blob in a separate packfile\", and\n> if the counter-argument is \"no, decompressing still costs a lot\", then the\n> real issue might be we decompress and look at the data when we do not have\n> to (i.e. issues similar to what abb371a addressed), not \"decompress vs\n> straight copy make a bit difference\".\n\nI've added Avery to the Cc list, because he really needs to chime in here.\n\nI am completely unqualified to make a comment about this, but I think\nthat it would be silly to ignore the insights that Avery has about\nstoring large objects; `bup' uses rolling checksums and a `bloom\nfilter' implementation and who knows what else.\n"}]}