git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Avery Pennarun's git-subtree?

From
Marc Branchaud <marcnarc@xiplink.com>
Date
Jul 26, 2010, 15:15 UTC
Message-ID
<4C4DA683.9020102@xiplink.com>
In-Reply-To
<AANLkTi=LHYDhY=424YZpO3yGqGGsxpY2Sj8=ULNKvAQX@mail.gmail.com>
On 10-07-23 06:50 PM, Avery Pennarun wrote:
Show 13 quoted lines
> On Fri, Jul 23, 2010 at 11:19 AM, Marc Branchaud <marcnarc@xiplink.com> wrote:
>> On 10-07-22 03:41 PM, Avery Pennarun wrote:
>>> 1) Sometimes I want to clone only some subdirs of a project
>>> 2) Sometimes I don't want the entire history because it's too big.
>>> 3) Super huge git repositories start to degrade in performance.
>>
>> The reason we turned to submodules is precisely to deal with repository size.
> 
> I believe that's very common.
> 
> However, I wonder whether that's actually a good reason for git to
> develop better submodules, or actually just a good reason for git to
> get better support for handling huge repositories.

I think that's a fundamental question, but part of the problem in coming up with an answer is that there's no agreed-upon definition of how to handle huge repos. People have provided tools that answer the question in ways they like, but I think the fact that these issues keep coming up is proof that git isn't there yet.

Show 9 quoted lines
>>  Our code base encompasses the entire FreeBSD tree plus different versions of
>> the Linux kernel, along with various third-party libraries & apps.  You don't
>> need everything to build a given product (a FreeBSD product doesn't use any
>> Linux kernels, for example) but because all the products share common code we
>> need to be able to branch and tag the common code along with the uncommon code.
> 
> Honest question: do you care about the wasted disk space and download
> time for these extra files?  Or just the fact that git gets slow when
> you have them?

It's not the disk space or the extra download time. It's how long takes to checkout all those files, and how long it takes to "git status" in a unified repo.

Show 8 quoted lines
>> So a straight "git clone" that would need to fetch all of FreeBSD plus 4
>> different Linux kernels and check all that out is a major problem, especially
>> for our automated build system (which could definitely be implemented better,
>> but still).
> 
> To be absolutely pedantic, the four linux kernels likely share most of
> their objects and so you're only paying the cost (at least during
> fetch) of including it once :)

That is true, but like I said the problem is the checkout. Our different products use different kernels (or FreeBSD):

	Product 1 -- Linux vX
	Product 2 -- Linux vY
	Product 3 -- FreeBSD
(Luckily we're currently only using one version of FreeBSD...)
All the products use common code.  When we release, we need to tag the common
code and the particular Linux kernel (or FreeBSD) we built the product with.
 We can't stuff all the Linux kernels into a single submodule, because then
the repo will be "dirty" if we checkout a different Linux kernel to build a
different product.  Even in a unified repo we'd need the kernels to live in
their own trees.

So we've ended up with individual submodules for each Linux kernel, and we've taught our automated build to only clone/checkout the kernel it needs to build the target product. Otherwise the checkout I/O overshadows the actual build time, especially when we try to run several builds in parallel on one slave machine.

Show 6 quoted lines
> (If you're actually using git-submodule and each copy of the kernel is
> its own module, then it might be cloning the kernel four times
> separately, in which case the objects *don't* get shared, so this ends
> up being much more expensive than it should be.  That could be fixed
> by slightly improving git-submodule to share some objects rather than
> rearchitecting it though.)
Even with the --reference parameter, it's still a problem.
Show 27 quoted lines
>>  In truth it's the checkout that takes the most time by far,
>> though commands like git-status also take inconveniently long.
> 
> Yeah, git could stand to be optimized a bit here.  And since Windows
> stats files about 10x slower than Linux, this problem occurs about 10x
> sooner on Windows, which makes using git on Windows (which sadly I
> have to do sometimes) extremely painful compared to Linux.
> 
> IMHO, the correct answer here is to have an inotify-based daemon prod
> at the .git/index automatically when files get updated, so that git
> itself doesn't have to stat/readdir through the entire tree in order
> to do any of its operations.  (Windows also has something like inotify
> that would work.)  If you had this, then git
> status/diff/checkout/commit would be just as fast with zillions of
> files as with 10 files.  Sooner or later, if nobody implements this, I
> promise I'll get around to it since inotify is actually easy to code
> for :)
> 
> Also note that the only reason submodules are faster here is that
> they're ignoring possibly important changes.  Notably, when you do
> 'git status' from the top level, it won't warn you if you have any
> not-yet-committed files in any of your submodules.  Personally, I
> consider that to be really important information, but to obtain it
> would make 'git status' take just as long as without submodules, so
> you wouldn't get any benefit.  (I think nowadays there's a way to get
> this recursive status information if you want it, but it'll be slow of
> course.)

I'm happy with a "git status" that can ignore uninitialized submodules and still probe into initialized/cloned ones. I agree that it's important for "git status" to be correct.

Show 7 quoted lines
>> We chose git-submodule over git-subtree mainly because git-submodule lets us
>> selectively checkout different parts of our code.  (AFAIK sparse checkouts
>> aren't yet an option.)
> 
> Fair enough.  If you could confirm or deny my theory that this is
> *entirely* a performance related concern (as opposed to disk space /
> download time), that would be helpful.

Consider it confirmed. Honestly, disk space is a complete non-issue. It's always nice to have faster download times, but it hasn't been an issue for us and there are already several ways to work around it anyway.

Show 9 quoted lines
>>  We didn't really consider git-subtree because it's
>> not an official part of git, and we didn't want to have to teach (and nag)
>> all our developers to install and maintain it in addition to keeping up with
>> git itself.
> 
> Arguably, this is a vote for including git-subtree into the core
> (which was Bryan's point when he started this thread); it obviously is
> being rejected sometimes by git users simply because it's not in the
> core, even though it could help them.

Yes, I have no objection to seeing git-subtree becoming an official part of git. My only complaint would be that it doesn't really help git deal with huge repos.

		M.
Previous: Avery Pennarun
Message 58 of 58 in “Avery Pennarun's git-subtree?”
  1. Bryan LarsenJul 21, 2010
  2. Ævar Arnfjörð BjarmasonJul 21, 2010
  3. Avery PennarunJul 21, 2010
  4. Ævar Arnfjörð BjarmasonJul 21, 2010
  5. Avery PennarunJul 21, 2010
  6. Avery PennarunJul 21, 2010
  7. Jens LehmannJul 21, 2010
  8. Avery PennarunJul 22, 2010
  9. Ævar Arnfjörð BjarmasonJul 21, 2010
  10. Bryan LarsenJul 22, 2010
  11. Jakub NarebskiJul 24, 2010
  12. Avery PennarunJul 22, 2010
  13. Jonathan NiederJul 22, 2010
  14. Avery PennarunJul 22, 2010
  15. Ævar Arnfjörð BjarmasonJul 22, 2010
  16. Avery PennarunJul 22, 2010
  17. Jens LehmannJul 23, 2010
  18. Eugene SajineJul 26, 2010
  19. Elijah NewrenJul 22, 2010
  20. Avery PennarunJul 22, 2010
  21. Chris WebbJul 23, 2010
  22. Avery PennarunJul 23, 2010
  23. Jens LehmannJul 23, 2010
  24. Avery PennarunJul 23, 2010
  25. Jens LehmannJul 23, 2010
  26. Jens LehmannJul 23, 2010
  27. Bryan LarsenJul 23, 2010
  28. Jens LehmannJul 23, 2010
  29. Bryan LarsenJul 23, 2010
  30. Avery PennarunJul 23, 2010
  31. Jens LehmannJul 25, 2010
  32. Avery PennarunJul 27, 2010
  33. Jens LehmannJul 27, 2010
  34. Marc BranchaudJul 23, 2010
  35. Avery PennarunJul 23, 2010
  36. skillzero@gmail.comJul 24, 2010
  37. Avery PennarunJul 24, 2010
  38. skillzero@gmail.comJul 24, 2010
  39. Nguyen Thai Ngoc DuyJul 25, 2010
  40. Jakub NarebskiJul 28, 2010
  41. Jakub NarebskiJul 26, 2010
  42. Marc BranchaudJul 26, 2010
  43. Linus TorvaldsJul 26, 2010
  44. Bryan LarsenJul 26, 2010
  45. Linus TorvaldsJul 26, 2010
  46. Avery PennarunJul 27, 2010
  47. Junio C HamanoJul 27, 2010
  48. Avery PennarunJul 27, 2010
  49. Junio C HamanoJul 27, 2010
  50. Jens LehmannJul 27, 2010
  51. Jakub NarebskiJul 26, 2010
  52. Avery PennarunJul 27, 2010
  53. Marc BranchaudJul 28, 2010
  54. Jakub NarebskiJul 28, 2010
  55. Sverre RabbelierJul 24, 2010
  56. Jakub NarebskiJul 26, 2010
  57. Avery PennarunJul 27, 2010
  58. Marc BranchaudJul 26, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.