git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Avery Pennarun's git-subtree?

From
Sskillzero@gmail.com <skillzero@gmail.com>
Date
Jul 24, 2010, 19:40 UTC
Message-ID
<AANLkTikx5EtQ0yvdkqN1Q1QAudFZfbd+_jpoa9ztLrz1@mail.gmail.com>
In-Reply-To
<AANLkTimLayG_HFxGdq+Tt8hU_MApBpSdHHiYPxcakpRJ@mail.gmail.com>
On Fri, Jul 23, 2010 at 6:20 PM, Avery Pennarun <apenwarr@gmail.com> wrote:
Show 19 quoted lines
> On Fri, Jul 23, 2010 at 8:58 PM,  <skillzero@gmail.com> wrote:
>> On Fri, Jul 23, 2010 at 3:50 PM, Avery Pennarun <apenwarr@gmail.com> wrote:
>>> Honest question: do you care about the wasted disk space and download
>>> time for these extra files?  Or just the fact that git gets slow when
>>> you have them?
>>
>> I have the similar situation to the original poster (huge trees) and
>> for me it's all three: disk space, download time, and performance. My
>> tree has a few relatively small (< 20 MB) shared directories of common
>> code, a few large (2-6 GB) directories of code for OS's, and then
>> several medium size (< 500 MB) directories for application code. The
>> application developers only care about the app+shared directories (and
>> are very annoyed by the massive space and performance impact of the OS
>> directories).
>
> Given how cheap disk space is nowadays, I'm curious about this.  Are
> they really just annoyed by the performance problem, and they complain
> about the extra size because they blame the performance on the extra
> files?  Or are they honestly short of disk space?

I think it's both space and performance. When you're using SSD drives, storage still pretty expensive. A 128 GB or less SSD is pretty common in a laptop so you can run out pretty quick, especially when you're working concurrently on a few different branches at the same time. It's useful to keep multiple working copies (e.g. git-new-workdir) because rebuild time can be significant when switching branches.

> Similarly, are all your developers located at the same office?  If so,
> then bandwidth ought not be an issue.

Bandwidth isn't a big problem because you don't need to re-download the repo very often. However, people work at home a lot where bandwidth is more limited. The biggest complaint I hear about bandwidth is that people tend to re-download when something goes wrong (i.e. inexperience with git resulting in a repository they can't recover due to git resets, etc).

Show 5 quoted lines
> I'm pushing extra hard on this because I believe there are lots of
> opportunities to just improve git performance on huge repositories.
> And if the only *real* reason people need to split repositories is
> that performance goes down, then that's fixable, and you may need
> neither git-submodule nor git-subtree.

Performance degradation is my biggest complaint with large repositories. Your inotify/FSEvents/etc daemon idea sounds interesting to deal with the stat issue.

Show 7 quoted lines
> This is indeed a problem with large repositories.  Of course,
> splitting them with git-submodule is kind of cheating, because it just
> makes git-status *not look* to see if those files are dirty or not.
> If they are dirty and you forget to commit them, you'll never know
> until someone tells you later.  It would be functionally equivalent to
> just have git-status not look inside certain subdirs of a single
> repository.

I think it's only cheating if you're using all of the submodules. The main purpose of submodules for me (although I don't currently use submodules) would be so I don't need to keep modules on disk that I don't care about. If a developer is working on an app, they don't need the OS directories/modules so they get much faster git status/etc and there wouldn't be other directories to have dirty files in. That said, if I was using git submodule, I'd want git status to show me all the submodules that were checked out.

Show 5 quoted lines
>> (although just having all those objects in
>> the .git directory still slows it down quite a bit).
>
> You're the second person who has mentioned this today (the first one
> was to me in a private email).  I'd like to understand this better.

What I'm basing this on is that even when I'm using a sparse checkout such that I have only a small subset of the files in my working directory, git status seems singifncantly slower for me than an equivalent git repository that only has that subset of files. That's not very scientific, but that's what made me think just having a large .git directory with lots of objects/history slows down git status even if the working copy doesn't have a lot of files.

I will try to experiment and see if I can narrow it down with some real numbers.

BTW...what's the policy on CC'ing people on git mailing list replies? Should it be trimmed or not? I've received complaints in the past, but I was never really clear what the recommended policy is.

Previous: Avery PennarunNext: Nguyen Thai Ngoc Duy
Message 38 of 58 in “Avery Pennarun's git-subtree?”
  1. Bryan LarsenJul 21, 2010
  2. Ævar Arnfjörð BjarmasonJul 21, 2010
  3. Avery PennarunJul 21, 2010
  4. Ævar Arnfjörð BjarmasonJul 21, 2010
  5. Avery PennarunJul 21, 2010
  6. Avery PennarunJul 21, 2010
  7. Jens LehmannJul 21, 2010
  8. Avery PennarunJul 22, 2010
  9. Ævar Arnfjörð BjarmasonJul 21, 2010
  10. Bryan LarsenJul 22, 2010
  11. Jakub NarebskiJul 24, 2010
  12. Avery PennarunJul 22, 2010
  13. Jonathan NiederJul 22, 2010
  14. Avery PennarunJul 22, 2010
  15. Ævar Arnfjörð BjarmasonJul 22, 2010
  16. Avery PennarunJul 22, 2010
  17. Jens LehmannJul 23, 2010
  18. Eugene SajineJul 26, 2010
  19. Elijah NewrenJul 22, 2010
  20. Avery PennarunJul 22, 2010
  21. Chris WebbJul 23, 2010
  22. Avery PennarunJul 23, 2010
  23. Jens LehmannJul 23, 2010
  24. Avery PennarunJul 23, 2010
  25. Jens LehmannJul 23, 2010
  26. Jens LehmannJul 23, 2010
  27. Bryan LarsenJul 23, 2010
  28. Jens LehmannJul 23, 2010
  29. Bryan LarsenJul 23, 2010
  30. Avery PennarunJul 23, 2010
  31. Jens LehmannJul 25, 2010
  32. Avery PennarunJul 27, 2010
  33. Jens LehmannJul 27, 2010
  34. Marc BranchaudJul 23, 2010
  35. Avery PennarunJul 23, 2010
  36. skillzero@gmail.comJul 24, 2010
  37. Avery PennarunJul 24, 2010
  38. skillzero@gmail.comJul 24, 2010
  39. Nguyen Thai Ngoc DuyJul 25, 2010
  40. Jakub NarebskiJul 28, 2010
  41. Jakub NarebskiJul 26, 2010
  42. Marc BranchaudJul 26, 2010
  43. Linus TorvaldsJul 26, 2010
  44. Bryan LarsenJul 26, 2010
  45. Linus TorvaldsJul 26, 2010
  46. Avery PennarunJul 27, 2010
  47. Junio C HamanoJul 27, 2010
  48. Avery PennarunJul 27, 2010
  49. Junio C HamanoJul 27, 2010
  50. Jens LehmannJul 27, 2010
  51. Jakub NarebskiJul 26, 2010
  52. Avery PennarunJul 27, 2010
  53. Marc BranchaudJul 28, 2010
  54. Jakub NarebskiJul 28, 2010
  55. Sverre RabbelierJul 24, 2010
  56. Jakub NarebskiJul 26, 2010
  57. Avery PennarunJul 27, 2010
  58. Marc BranchaudJul 26, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.