{"thread":{"id":"16136","subject":"Git and Media repositories....","startedAt":"2008-11-02T19:50:28Z","lastAt":"2008-11-09T04:58:41Z","messageCount":6,"participants":["Tim Ansell","Johannes Schindelin","Jakub Narebski","Santi Béjar","Nguyen Thai Ngoc Duy"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"94689","messageId":"1225655428.11693.10.camel@vaio","threadId":"16136","inReplyTo":null,"subject":"Git and Media repositories....","fromName":"Tim Ansell","fromEmail":"mithro@mithis.com","sentAt":"2008-11-02T19:50:28Z","receivedAt":"2008-11-02T19:50:28Z","isPatch":false,"sender":{"key":"mithro@mithis.com","avatar":"https://gravatar.com/avatar/1df49ebfaff79712e08212ff66228e688973c823b6452c23268f9a8c0e4d341b?d=mp&s=160"},"body":"Hey guys,\n\nLast week at the gittogether I lead some discussions about how we could\nmake Git better support large media repositories (which is one area\nwhere Subversion still make sense). It was suggested that I post to this\nlist to get a discussion going. \n\nThe general idea is that we always clone the complete meta-data (tags,\ncommits and trees) and then only clone blobs when they are needed (using\nsomething like alternates). This allows us to support shallow, narrow\nand sparse checkouts while still being able to perform operations such\nas committing and merging.\n\nYou can find a copy of the summary presentation at\n http://www.thousandparsec.net/~tim/media+git.pdf\n\nI have started working on adapting git to check a remote http alternate\nto provide a proof of concept.\n\nI appreciate any help or suggestions.\n\nTim 'mithro' Ansell\n"},{"id":"94721","messageId":"alpine.DEB.1.00.0811030755270.22125@pacific.mpi-cbg.de.mpi-cbg.de","threadId":"16136","inReplyTo":"1225655428.11693.10.camel@vaio","subject":"Re: Git and Media repositories....","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-11-03T06:56:33Z","receivedAt":"2008-11-03T06:56:33Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 2 Nov 2008, Tim Ansell wrote:\n\n> Last week at the gittogether I lead some discussions about how we could \n> make Git better support large media repositories (which is one area \n> where Subversion still make sense). It was suggested that I post to this \n> list to get a discussion going.\n> \n> The general idea is that we always clone the complete meta-data (tags, \n> commits and trees) and then only clone blobs when they are needed (using \n> something like alternates). This allows us to support shallow, narrow \n> and sparse checkouts while still being able to perform operations such \n> as committing and merging.\n> \n> You can find a copy of the summary presentation at\n>  http://www.thousandparsec.net/~tim/media+git.pdf\n> \n> I have started working on adapting git to check a remote http alternate \n> to provide a proof of concept.\n> \n> I appreciate any help or suggestions.\n\nYou might find this message (and others from the same time frame and \nauthor) pretty interesting:\n\nhttp://article.gmane.org/gmane.comp.version-control.git/48485\n\nCiao,\nDscho\n"},{"id":"94737","messageId":"m3ljw1f8qv.fsf@localhost.localdomain","threadId":"16136","inReplyTo":"1225655428.11693.10.camel@vaio","subject":"Re: Git and Media repositories....","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-11-03T09:40:55Z","receivedAt":"2008-11-03T09:40:55Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Tim Ansell <mithro@mithis.com> writes:\n\n> Last week at the GitTogether I lead some discussions about how we could\n> make Git better support large media repositories (which is one area\n> where Subversion still make sense). It was suggested that I post to this\n> list to get a discussion going. \n> \n> The general idea is that we always clone the complete meta-data (tags,\n> commits and trees) and then only clone blobs when they are needed (using\n> something like alternates). This allows us to support shallow, narrow\n> and sparse checkouts while still being able to perform operations such\n> as committing and merging.\n> \n> You can find a copy of the summary presentation at\n>  http://www.thousandparsec.net/~tim/media+git.pdf\n> \n> I have started working on adapting git to check a remote http alternate\n> to provide a proof of concept.\n> \n> I appreciate any help or suggestions.\n\nDana How (CC-ed) worked on better support for large files, but in\ncorporate setting.  The solution that was the result of all discussion\nand all patches (not all accpeted) was to create kept packfile for\nthose large files, and share those packfiles (perhaps via alternates)\nusing network filesystem, instead of keeping separate copies and\ntrasferring them on fetch / push.\n\n\n>From what I remember there was one serious attempt (by serious I mean\nhere with patches) to add 'lazy clone' / 'sparse clone' / 'remote\nalternates', using some kind of \"stub\" objects and trasferring objects\nlazily.  This patch was fairly intrusive, and didn't get accepted.\nI think you can find it in archives.  Unfortunately I haven't bookmarked\nthis thread...\n\nThe problem with lazy clone is that git assumes in many places that if\nit has some object, it has all its dependencies.  Lazy clone\n(on-demand object loading) breaks this assumption... although in your\ncase (only blobs of large size can be asked to be loaded lazily) it is\nmigitated somehow.\n\n\nI also think that you would have to have 'sparse checkout' support.\nIf you don't have blob in object repository (and don't want to have it\nthere), you can not check it out.  Fortunately this feature is quite\nalive, and worked on by Duy (pclouds), see \"What's cooking...\"\n(nd/narrow branch in 'pu').\n\nHTH\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"95124","messageId":"m3r65nelo1.fsf@localhost.localdomain","threadId":"16136","inReplyTo":"1225655428.11693.10.camel@vaio","subject":"Re: Git and Media repositories....","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-11-07T13:00:43Z","receivedAt":"2008-11-07T13:00:43Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Tim Ansell <mithro@mithis.com> writes:\n\n> Last week at the gittogether I lead some discussions about how we could\n> make Git better support large media repositories (which is one area\n> where Subversion still make sense). It was suggested that I post to this\n> list to get a discussion going. \n> \n> The general idea is that we always clone the complete meta-data (tags,\n> commits and trees) and then only clone blobs when they are needed (using\n> something like alternates). This allows us to support shallow, narrow\n> and sparse checkouts while still being able to perform operations such\n> as committing and merging.\n[...]\n\nWell, the *workaround* you could currently use is to put large media\nfiles in separate subdirectory, and make this subdirectory into\nsubmodule.  This uses the fact that you can selectively clone\nsubmodules, or leave them as a stubs...\n\n...and this is also the code you might want to look at when\nimplementings stubs for 'remote' blob objects\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"95126","messageId":"adf1fd3d0811070519qc54369vf0da3bc28182460a@mail.gmail.com","threadId":"16136","inReplyTo":"1225655428.11693.10.camel@vaio","subject":"Re: Git and Media repositories....","fromName":"Santi Béjar","fromEmail":"santi@agolina.net","sentAt":"2008-11-07T13:19:05Z","receivedAt":"2008-11-07T13:19:05Z","isPatch":false,"sender":{"key":"santi@agolina.net","avatar":null},"body":"On Sun, Nov 2, 2008 at 8:50 PM, Tim Ansell <mithro@mithis.com> wrote:\n> Hey guys,\n>\n\n[...]\n\n>\n> The general idea is that we always clone the complete meta-data (tags,\n> commits and trees) and then only clone blobs when they are needed (using\n> something like alternates). This allows us to support shallow, narrow\n> and sparse checkouts while still being able to perform operations such\n> as committing and merging.\n>\n\nA related use case could be to remove a blob from a repo but being\nable to work normally with it, similar to:\n\nhttp://wiki.freebsd.org/VCSFeatureObliterate\n\nSanti\n"},{"id":"95248","messageId":"fcaeb9bf0811082058i2d0dab9et3de216f7bdb8036d@mail.gmail.com","threadId":"16136","inReplyTo":"adf1fd3d0811070519qc54369vf0da3bc28182460a@mail.gmail.com","subject":"Re: Git and Media repositories....","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2008-11-09T04:58:41Z","receivedAt":"2008-11-09T04:58:41Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On 11/7/08, Santi Béjar <santi@agolina.net> wrote:\n> On Sun, Nov 2, 2008 at 8:50 PM, Tim Ansell <mithro@mithis.com> wrote:\n>  > Hey guys,\n>  >\n>\n>  [...]\n>\n>\n>  >\n>  > The general idea is that we always clone the complete meta-data (tags,\n>  > commits and trees) and then only clone blobs when they are needed (using\n>  > something like alternates). This allows us to support shallow, narrow\n>  > and sparse checkouts while still being able to perform operations such\n>  > as committing and merging.\n>  >\n>\n>\n> A related use case could be to remove a blob from a repo but being\n>  able to work normally with it, similar to:\n>\n>  http://wiki.freebsd.org/VCSFeatureObliterate\n\nMaybe another use case: encrypted blobs (those are generally\nunavailable until corrected password is given, so they are \"holes\" in\ncheckout/clone). It could be used to store sensitive content (in $HOME\nfor example)\n-- \nDuy\n"}]}