git/list[1] front-page[2] threads[3] people[4] search[5] about
 

big files in git was: Re: Features from GitSurvey 2010

From
Ddavid@lang.hm <david@lang.hm>
Date
Feb 1, 2011, 22:50 UTC
Message-ID
<alpine.DEB.2.00.1102011443380.10088@asgard.lang.hm>
In-Reply-To
<201102011451.17456.jnareb@gmail.com>
On Tue, 1 Feb 2011, Jakub Narebski wrote:
Show 26 quoted lines
> On Sun, 30 Jan 2011, Jonathan Nieder wrote:
>> Hi Dmitry,
>>
>> Dmitry S. Kravtsov wrote:
>>
>>> I want to dedicate my coursework at University to implementation of
>>> some useful git feature. So I'm interesting in some kind of list of
>>> development status of these features
>> [...]
>>> Or I'll be glad to know what features are now 'free' and what are
>>> currently in active development.
>>
>> Interesting question.  The short answer is that they are all "free".
>> Generally people seem to be happy to learn of an alternative approach
>> to what they have been working on.
>>
>> [For the following pointers, the easiest way to follow up is probably
>> to search the mailing list archives.]
>>
>>> better support for big files (large media)
>>
>> For a conservative approach, you might want to get in touch with Sam
>> Hocevar, Nicolas Pitre, and Miklos Vajna.  The idea is to stream big
>> files directly to pack and not waste time trying to compress them.
>
> There is also, supposedly stalled, git-bigfiles project.

why is the clean/smudge approach that came through the list a week or two ago not acceptable?

While people talked about how it would be nice to store the large files on $remote_destination, just create a .git/bigfiles and store them in there.

with the ability to pass the filename to the clean/smudge scripts, you can even avoid the copy (replacing it with a mv) and have a working, if bare-bones system.

Then people can create/submit enhanced versions of these scripts that store the large files elsewhere if they want, but we would be past the "git can't handle large files" into "git handles large files less efficiently", which is a much better place to be.

If nobody else has time to take those e-mails and create a set of clean/smudge scripts, I'll do so later this week (unless there is some reason why they wouldn't be acceptable)

I guess the only question is how to tell what files need to be handled this way, but can't we have something in .gitattributes about the file size? (and if that's a problem for checking files out, have the stored file be a sparse file, that way it's large, but doesn't take much space on sane filesystems)

David Lang
Previous: Nicolas PitreNext: Nicolas Pitre
Message 32 of 36 in “Features from GitSurvey 2010”
  1. Dmitry S. KravtsovJan 29, 2011
  2. Jonathan NiederJan 29, 2011
  3. Jakub NarebskiFeb 1, 2011
  4. Nguyen Thai Ngoc DuyFeb 1, 2011
  5. Shawn PearceFeb 1, 2011
  6. Shawn PearceFeb 1, 2011
  7. Nguyen Thai Ngoc DuyFeb 1, 2011
  8. Junio C HamanoFeb 1, 2011
  9. Nicolas PitreFeb 1, 2011
  10. Nguyen Thai Ngoc DuyFeb 1, 2011
  11. Shawn PearceFeb 1, 2011
  12. Nicolas PitreFeb 1, 2011
  13. Shawn PearceFeb 2, 2011
  14. Nicolas PitreFeb 2, 2011
  15. david@lang.hmFeb 2, 2011
  16. Geert BoschFeb 3, 2011
  17. Narrow clone (Re: features from GitSurvey 2010)Jonathan Nieder, Feb 3, 2011
  18. Geert BoschFeb 3, 2011
  19. Jonathan NiederFeb 3, 2011
  20. Jonathan NiederFeb 3, 2011
  21. Nicolas PitreFeb 3, 2011
  22. Tracking empty directoriesJonathan Nieder, Feb 1, 2011
  23. Nguyen Thai Ngoc DuyFeb 1, 2011
  24. Ilari LiusvaaraFeb 1, 2011
  25. Jakub NarebskiFeb 1, 2011
  26. Ilari LiusvaaraFeb 1, 2011
  27. Jonathan NiederFeb 1, 2011
  28. Jakub NarebskiFeb 1, 2011
  29. Nguyen Thai Ngoc DuyFeb 2, 2011
  30. Kevin P. FlemingFeb 2, 2011
  31. Nicolas PitreFeb 1, 2011
  32. big files in git was: Re: Features from GitSurvey 2010david@lang.hm, Feb 1, 2011
  33. Nicolas PitreFeb 3, 2011
  34. Matthieu MoyFeb 1, 2011
  35. Jonathan NiederFeb 1, 2011
  36. Matthieu MoyFeb 1, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.