git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Avery Pennarun's git-subtree?

From
Avery Pennarun <apenwarr@gmail.com>
Date
Jul 27, 2010, 19:15 UTC
Message-ID
<AANLkTi=6SDQ2A0Zxf8DiSSNzSfUS43M7wmCkKKraOd8w@mail.gmail.com>
In-Reply-To
<201007261051.41663.jnareb@gmail.com>
On Mon, Jul 26, 2010 at 4:51 AM, Jakub Narebski <jnareb@gmail.com> wrote:
Show 10 quoted lines
> On Sat, 24 Jul 2010 00:50, Avery Pennarun wrote:
>> My bup project (http://github.com/apenwarr/bup) is all about huge
>> repositories.  It handles repositories with hundreds of gigabytes, and
>> trees containing millions of files (entire filesystems), quite nicely.
>>  Of course, it's not a version control system, so it won't solve your
>> problems.  It's just evidence that large repositories are actually
>> quite manageable without changing the fundamentals of git.
>
> There is also git-bigfiles project, although it is more about large
> [binary] files than large repositories per se (many files, long history).

Right. git-bigfiles is valuable, but it's valuable with or without submodules. (If you have large blobs, submodules won't save you.)

bup happens to have its own way of dealing with large files too, but it may not be applicable to git. It does result in lots and lots of smaller objects, though, which is why I know git repositories are fundamentally capable of handling lots and lots of smaller objects :)

> Note that with 'bup' you might not see problems with large repositories
> because it does not examine code paths that are slow in large repositories
> (gc, log, path-delimited log).

gc is a huge problem. bup avoids it entirely (it foregoes delta compression); git gc fails completely on such large repositories (100+ GB). There's no reason this has to be true forever, but yes, to support really big repos, git gc would need to be improved somewhat. For most reasonably sane repos (a few GB) you can get reasonable performance by just making your biggest packfiles .keep so they don't keep getting repacked all the time.

Compared to that, log feels like not a problem at all :) At least performance-wise. The thing that sucks about log using git-subtree, of course, is that you get all these log messages from multiple projects jammed together into a single repo, which is rarely what you want, even if it's fast. I think the "best" solution is a single repo with all your objects, but still keeping the histories of each submodule separate.

Show 13 quoted lines
>> IMHO, the correct answer here is to have an inotify-based daemon prod
>> at the .git/index automatically when files get updated, so that git
>> itself doesn't have to stat/readdir through the entire tree in order
>> to do any of its operations.  (Windows also has something like inotify
>> that would work.)  If you had this, then git
>> status/diff/checkout/commit would be just as fast with zillions of
>> files as with 10 files.  Sooner or later, if nobody implements this, I
>> promise I'll get around to it since inotify is actually easy to code
>> for :)
>
> IIUC the problem is that inotify is not automatically recursive, so
> daemon would have to take care of adding inotify trigger to each newly
> created subdirectory.

Yeah, the inotify API is kind of gross that way. But it can be done, and people do. (eg. the beagle project)

Show 13 quoted lines
>> Also note that the only reason submodules are faster here is that
>> they're ignoring possibly important changes.  Notably, when you do
>> 'git status' from the top level, it won't warn you if you have any
>> not-yet-committed files in any of your submodules.  Personally, I
>> consider that to be really important information, but to obtain it
>> would make 'git status' take just as long as without submodules, so
>> you wouldn't get any benefit.  (I think nowadays there's a way to get
>> this recursive status information if you want it, but it'll be slow of
>> course.)
>
> Errr... didn't it got improved in recent git?  I think git-status now
> includes information about submodules if configured so / unless configured
> otherwise.  Isn't it?

Yes, but you're still left with the choice between slow (checks all files in all submodules) and not slow (might miss stuff). This isn't a submodule question, really, it's an overall performance question with huge checkouts with or without submodules.

Show 7 quoted lines
>>> We chose git-submodule over git-subtree mainly because git-submodule lets us
>>> selectively checkout different parts of our code.  (AFAIK sparse checkouts
>>> aren't yet an option.)
>
> Sparse checkouts are here, IIRC, but they do not solve problem of disk
> space (they are still in repository, even if not checked out), and speed
> (they still need to be fetched, even if not checked out).

Hmm, don't mix bandwidth usage (and thus the slowness of fetch) with slowness during everyday usage. I don't mind a slow fetch now and then, but 'git status' should be fast. AFAIK, sparse checkouts *should* make git status faster. If they don't, it's probably just a bug.

Have fun,
Avery
Previous: Jakub NarebskiNext: Marc Branchaud
Message 57 of 58 in “Avery Pennarun's git-subtree?”
  1. Bryan LarsenJul 21, 2010
  2. Ævar Arnfjörð BjarmasonJul 21, 2010
  3. Avery PennarunJul 21, 2010
  4. Ævar Arnfjörð BjarmasonJul 21, 2010
  5. Avery PennarunJul 21, 2010
  6. Avery PennarunJul 21, 2010
  7. Jens LehmannJul 21, 2010
  8. Avery PennarunJul 22, 2010
  9. Ævar Arnfjörð BjarmasonJul 21, 2010
  10. Bryan LarsenJul 22, 2010
  11. Jakub NarebskiJul 24, 2010
  12. Avery PennarunJul 22, 2010
  13. Jonathan NiederJul 22, 2010
  14. Avery PennarunJul 22, 2010
  15. Ævar Arnfjörð BjarmasonJul 22, 2010
  16. Avery PennarunJul 22, 2010
  17. Jens LehmannJul 23, 2010
  18. Eugene SajineJul 26, 2010
  19. Elijah NewrenJul 22, 2010
  20. Avery PennarunJul 22, 2010
  21. Chris WebbJul 23, 2010
  22. Avery PennarunJul 23, 2010
  23. Jens LehmannJul 23, 2010
  24. Avery PennarunJul 23, 2010
  25. Jens LehmannJul 23, 2010
  26. Jens LehmannJul 23, 2010
  27. Bryan LarsenJul 23, 2010
  28. Jens LehmannJul 23, 2010
  29. Bryan LarsenJul 23, 2010
  30. Avery PennarunJul 23, 2010
  31. Jens LehmannJul 25, 2010
  32. Avery PennarunJul 27, 2010
  33. Jens LehmannJul 27, 2010
  34. Marc BranchaudJul 23, 2010
  35. Avery PennarunJul 23, 2010
  36. skillzero@gmail.comJul 24, 2010
  37. Avery PennarunJul 24, 2010
  38. skillzero@gmail.comJul 24, 2010
  39. Nguyen Thai Ngoc DuyJul 25, 2010
  40. Jakub NarebskiJul 28, 2010
  41. Jakub NarebskiJul 26, 2010
  42. Marc BranchaudJul 26, 2010
  43. Linus TorvaldsJul 26, 2010
  44. Bryan LarsenJul 26, 2010
  45. Linus TorvaldsJul 26, 2010
  46. Avery PennarunJul 27, 2010
  47. Junio C HamanoJul 27, 2010
  48. Avery PennarunJul 27, 2010
  49. Junio C HamanoJul 27, 2010
  50. Jens LehmannJul 27, 2010
  51. Jakub NarebskiJul 26, 2010
  52. Avery PennarunJul 27, 2010
  53. Marc BranchaudJul 28, 2010
  54. Jakub NarebskiJul 28, 2010
  55. Sverre RabbelierJul 24, 2010
  56. Jakub NarebskiJul 26, 2010
  57. Avery PennarunJul 27, 2010
  58. Marc BranchaudJul 26, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.