git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Handling large files with GIT

From
FWFlorian Weimer <fw@deneb.enyo.de>
Date
Feb 8, 2006, 21:20 UTC
Message-ID
<87slqty2c8.fsf@mid.deneb.enyo.de>
In-Reply-To
<46a038f90602080114r2205d72cmc2b5c93f6fffe03d@mail.gmail.com>
* Martin Langhoff:
> SVN does reasonably well tracking his >1GB mbox file. Now, I don't
> know if I like the idea of putting my own mbox file under version
> control, but it looks like projects with large and slow-changing files
> would be in trouble with GIT.

To my surprise, it's not that bad. The Debian testing-security team uses a single 1.8 MB file (400 KB compressed) to keep vulnerability data. Most changes to that file involve just a few lines. But even in this extreme case, git doesn't compare too badly against Subversion if you pack regularly (but not too often). Disk usage is actually *below* Subversion FSFS even with --depth=10 (the default, unfortunately a bit hard to override).

I plan to do another experiment for GCC, which contains marvels such as:

  35905  126056 1379093 gcc/ChangeLog-2005
  12610   61215  417584 gcc/combine.c

But the outcome will likely be quite similar to the secure-testing case: comparable disk space usage, not a difference in the order of one or more magnitudes.

But Subversion still has got a significant adventage: I can get a working copy without downloading full history (several gigabytes in GCC's case). There's also the slight drawback that you shouldn't pack too often, otherwise you'll reduce its effectiveness. You can always run "git-repack -a -d", but it's rather expensive. This means that you need to keep compressed fulltexts from a few dozen revisions, but I don't think this is a huge burden. All in all, the compressed fulltexts/packs model is a pretty good trade-off between disk usage, end user usability nad code complexity.

In your mbox case, you should simply try Maildir. The tree object (which lists all files in the Maildir folder) will still be rather large (about 40 to 50 bytes per message stored), though.

Previous: Junio C HamanoNext: Martin Langhoff
Message 6 of 39 in “Handling large files with GIT”
  1. Martin LanghoffFeb 8, 2006
  2. Johannes SchindelinFeb 8, 2006
  3. Linus TorvaldsFeb 8, 2006
  4. Linus TorvaldsFeb 8, 2006
  5. Junio C HamanoFeb 8, 2006
  6. Florian WeimerFeb 8, 2006
  7. Martin LanghoffFeb 8, 2006
  8. Ben CliffordFeb 13, 2006
  9. Linus TorvaldsFeb 13, 2006
  10. Linus TorvaldsFeb 13, 2006
  11. Linus TorvaldsFeb 13, 2006
  12. Ian MoltonFeb 13, 2006
  13. Martin LanghoffFeb 13, 2006
  14. Johannes SchindelinFeb 14, 2006
  15. Linus TorvaldsFeb 14, 2006
  16. Sam VilainFeb 14, 2006
  17. Linus TorvaldsFeb 14, 2006
  18. Junio C HamanoFeb 14, 2006
  19. Sam VilainFeb 15, 2006
  20. Junio C HamanoFeb 15, 2006
  21. Sam VilainFeb 15, 2006
  22. Martin LanghoffFeb 15, 2006
  23. Linus TorvaldsFeb 15, 2006
  24. Linus TorvaldsFeb 15, 2006
  25. Linus TorvaldsFeb 15, 2006
  26. Linus TorvaldsFeb 15, 2006
  27. Junio C HamanoFeb 15, 2006
  28. Linus TorvaldsFeb 15, 2006
  29. Linus TorvaldsFeb 15, 2006
  30. Linus TorvaldsFeb 16, 2006
  31. Junio C HamanoFeb 16, 2006
  32. Fredrik KuivinenFeb 16, 2006
  33. Jeff GarzikFeb 13, 2006
  34. Keith PackardFeb 13, 2006
  35. Martin LanghoffFeb 14, 2006
  36. Linus TorvaldsFeb 13, 2006
  37. Martin LanghoffFeb 13, 2006
  38. Greg KHFeb 9, 2006
  39. Martin LanghoffFeb 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.