git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Git and Media repositories....

From
Jakub Narebski <jnareb@gmail.com>
Date
Nov 3, 2008, 09:40 UTC
Message-ID
<m3ljw1f8qv.fsf@localhost.localdomain>
In-Reply-To
<1225655428.11693.10.camel@vaio>
Tim Ansell <mithro@mithis.com> writes:
Show 18 quoted lines
> Last week at the GitTogether I lead some discussions about how we could
> make Git better support large media repositories (which is one area
> where Subversion still make sense). It was suggested that I post to this
> list to get a discussion going. 
> 
> The general idea is that we always clone the complete meta-data (tags,
> commits and trees) and then only clone blobs when they are needed (using
> something like alternates). This allows us to support shallow, narrow
> and sparse checkouts while still being able to perform operations such
> as committing and merging.
> 
> You can find a copy of the summary presentation at
>  http://www.thousandparsec.net/~tim/media+git.pdf
> 
> I have started working on adapting git to check a remote http alternate
> to provide a proof of concept.
> 
> I appreciate any help or suggestions.

Dana How (CC-ed) worked on better support for large files, but in corporate setting. The solution that was the result of all discussion and all patches (not all accpeted) was to create kept packfile for those large files, and share those packfiles (perhaps via alternates) using network filesystem, instead of keeping separate copies and trasferring them on fetch / push.

>From what I remember there was one serious attempt (by serious I mean

here with patches) to add 'lazy clone' / 'sparse clone' / 'remote alternates', using some kind of "stub" objects and trasferring objects lazily. This patch was fairly intrusive, and didn't get accepted. I think you can find it in archives. Unfortunately I haven't bookmarked this thread...

The problem with lazy clone is that git assumes in many places that if it has some object, it has all its dependencies. Lazy clone (on-demand object loading) breaks this assumption... although in your case (only blobs of large size can be asked to be loaded lazily) it is migitated somehow.

I also think that you would have to have 'sparse checkout' support. If you don't have blob in object repository (and don't want to have it there), you can not check it out. Fortunately this feature is quite alive, and worked on by Duy (pclouds), see "What's cooking..." (nd/narrow branch in 'pu').

HTH
-- 
Jakub Narebski
Poland
ShadeHawk on #git
Previous: Johannes SchindelinNext: Jakub Narebski
Message 3 of 6 in “Git and Media repositories....”
  1. Tim AnsellNov 2, 2008
  2. Johannes SchindelinNov 3, 2008
  3. Jakub NarebskiNov 3, 2008
  4. Jakub NarebskiNov 7, 2008
  5. Santi BéjarNov 7, 2008
  6. Nguyen Thai Ngoc DuyNov 9, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.