git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Performance issue: initial git clone causes massive repack

From
Robin H. Johnson <robbat2@gentoo.org>
Date
Apr 5, 2009, 07:04 UTC
Message-ID
<20090405070412.GB869@curie-int>
In-Reply-To
<20090405035453.GB12927@vidovic>

Before I answer the rest of your post, I'd like to note that the matter of which choice between single-repo, repo-per-package, repo-per-category has been flogged to death within Gentoo.

I did not come to the Git mailing list to rehash those choices. I came here to find a solution to the performance problem. While it shows up with our repo, I'm certain that we're not the only people with the problem. The GSoC 2009 ideas contain a potential project for caching the generated packs, which, while having value in itself, could be partially avoided by sending suitable pre-built packs (if they exist) without any repacking.

On Sun, Apr 05, 2009 at 05:54:53AM +0200, Nicolas Sebrecht wrote:
Show 8 quoted lines
> > That causes incredibly bloat unfortunately.
> > 
> > I'll summarize why here for the git mailing list. Most our developers
> > have the entire tree checked out, and in informal surveys, would like to
> > continue to do so. There are ~13500 packages right now 
> Each developer doesn't work on so many packages, right ? From my point
> of view, checkin'out the entire tree is the wrong way on how to do
> things.

Also, I should note that working on the tree isn't the only reason to have the tree checked out. While the great majority of Gentoo users have their trees purely from rsync, there is nothing stopping you from using a tree from CVS (anonCVS for the users, master CVS server for the developers).

A quick bit of stats run show that while some developers only touch a few packages, there are at least 200 developers that have done a major change to 100 or more packages.

> > Without tail packing, the Gentoo tree is presently around 520MiB (you
> > can fit it into ~190MiB with tail packing). This means that
> > repo-per-package would have an overhead in the range of 400%.
> Don't know about the business for Gentoo, but HDD is cheap.
There's no reason to have bloat just for the layout to change.
> Also, I'd like to know how much space you will gain with the CVS to Git >
> migration.  How bigger is a CVS repo against a Git one ?
For the CVS checkouts right now: 
- ~410MiB of content (w/ 4kb inodes)
- ~240MiB of CVS overhead (w/ 4kb inodes)
(sorry about the earlier 520MiB number, I forgot to exclude a local dir
of stats data on my box when I ran du quickly).
Our experimental Git, with only a single repo for gentoo-x86:
- ~410MiB of content (w/ 4kb inodes)
- 80MiB - 1.6GiB of Git total overhead.

80MiB of overhead is the total overhead with a shallow clone at depth 1. 1.6GiB is with the full history.

And per-package numbers, because we DID do an experimental conversion,
last year, although the packs might not have been optimal:
- ~410MiB of content (w/ 4kb inodes)
- 4.7GiB of Git total overhead, with a breakdown:
  - 1.9GiB in inode waste
  - 2.8GiB in packs
> One repo per category could be a good compromise assuming one seperate
> branch per package, then.
Other downsides to repo-per-category and repo-per-package:
- Raises difficulty in adding a new package/category. 
  You cannot just do 'mkdir && vi ... && git add && git commit' anymore.
- The name of the directory for both of the category AND the package are not
  specified in the ebuild, as such, unless they are checked out to the right
  location, you will get breakage (definitely in the package name, and
  about 10% of the time with categories).
- You cannot use git-cvsserver with them cleanly and have the correct
  behavior (we DO have developers that want to use the CVS emulation
  layer) - adding a category or a package would NOT trigger the
  addition of a new repo on the server when needed.
- Does NOT present a good base for anybody wanting to branch the entire
  tree themselves.
  
Show 7 quoted lines
> > Additionally, there's a lot of commonality between ebuilds and packages,
> > and having repo-per-package means that the compression algorithms can't
> > make use of it - dictionary algorithms are effective at compression for
> > a reason.
> Please, no. We are in the long term issues. Compression will be
> efficient. It's all about the content of the files and dictionary
> algorithms certainly will do a good job over the ebuilds revisions.

We're already on track to drop the CVS $Header$, and thereafter, some of the ebuilds are already on track to be smaller. Here's our prototype dev-perl/Sub-Name-0.04. ==== # Copyright 1999-2009 Gentoo Foundation # Distributed under the terms of the GNU General Public License v2 MODULE_AUTHOR=XMATH inherit perl-module DESCRIPTION="(re)name a sub" LICENSE="|| ( Artistic GPL-2 )" SLOT="0" KEYWORDS="~amd64 ~x86" IUSE="" SRC_TEST=do ====

We can have all the CPAN packages from CPAN author XMATH, with changing only the DESCRIPTION string. KEYWORDS then just changes over the package lifespan.

-- 
Robin Hugh Johnson
Gentoo Linux Developer & Infra Guy
E-Mail     : robbat2@gentoo.org
GnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85
Previous: Nicolas SebrechtNext: Nicolas Sebrecht
Message 6 of 97 in “Performance issue: initial git clone causes massive repack”
  1. Robin H. JohnsonApr 4, 2009
  2. Nicolas SebrechtApr 5, 2009
  3. Robin H. JohnsonApr 5, 2009
  4. Nicolas SebrechtApr 5, 2009
  5. Nicolas SebrechtApr 5, 2009
  6. Robin H. JohnsonApr 5, 2009
  7. Nicolas SebrechtApr 5, 2009
  8. Shawn O. PearceApr 5, 2009
  9. Robin H. JohnsonApr 5, 2009
  10. Robin H. JohnsonApr 5, 2009
  11. Shawn O. PearceApr 5, 2009
  12. david@lang.hmApr 5, 2009
  13. Sverre RabbelierApr 5, 2009
  14. Nicolas PitreApr 6, 2009
  15. Björn SteinbrinkApr 7, 2009
  16. Jakub NarebskiApr 7, 2009
  17. Nicolas PitreApr 7, 2009
  18. Jakub NarebskiApr 7, 2009
  19. Jon SmirlApr 7, 2009
  20. Nicolas PitreApr 7, 2009
  21. Björn SteinbrinkApr 7, 2009
  22. Nicolas PitreApr 7, 2009
  23. Björn SteinbrinkApr 7, 2009
  24. Nicolas PitreApr 7, 2009
  25. Björn SteinbrinkApr 7, 2009
  26. Nicolas PitreApr 8, 2009
  27. Robin H. JohnsonApr 10, 2009
  28. Nicolas PitreApr 11, 2009
  29. Mike HommeyApr 11, 2009
  30. Johannes SchindelinApr 14, 2009
  31. Nicolas PitreApr 14, 2009
  32. Robin H. JohnsonApr 14, 2009
  33. Nicolas PitreApr 14, 2009
  34. Nguyen Thai Ngoc DuyApr 15, 2009
  35. Robin H. JohnsonApr 15, 2009
  36. Junio C HamanoApr 15, 2009
  37. Nicolas PitreApr 15, 2009
  38. Sam VilainApr 22, 2009
  39. Mike RalphsonApr 22, 2009
  40. Pieter de BieApr 22, 2009
  41. Johannes SchindelinApr 22, 2009
  42. Shawn O. PearceApr 22, 2009
  43. Andreas EricssonApr 22, 2009
  44. Johannes SchindelinApr 22, 2009
  45. Christian CouderApr 23, 2009
  46. Nicolas PitreApr 22, 2009
  47. Sam VilainApr 22, 2009
  48. Björn SteinbrinkApr 22, 2009
  49. Nicolas PitreApr 22, 2009
  50. Johannes SchindelinApr 22, 2009
  51. Nicolas PitreApr 23, 2009
  52. Johannes SchindelinApr 14, 2009
  53. Jeff KingApr 7, 2009
  54. Björn SteinbrinkApr 7, 2009
  55. process_{tree,blob}: Remove useless xstrdup callsBjörn Steinbrink, Apr 8, 2009
  56. Linus TorvaldsApr 10, 2009
  57. Linus TorvaldsApr 11, 2009
  58. Linus TorvaldsApr 11, 2009
  59. Nicolas PitreApr 11, 2009
  60. Björn SteinbrinkApr 11, 2009
  61. Björn SteinbrinkApr 11, 2009
  62. Linus TorvaldsApr 11, 2009
  63. Linus TorvaldsApr 11, 2009
  64. Björn SteinbrinkApr 11, 2009
  65. Björn SteinbrinkApr 11, 2009
  66. Linus TorvaldsApr 11, 2009
  67. Björn SteinbrinkApr 11, 2009
  68. Linus TorvaldsApr 11, 2009
  69. Björn SteinbrinkApr 11, 2009
  70. Linus TorvaldsApr 11, 2009
  71. Nicolas SebrechtApr 5, 2009
  72. david@lang.hmApr 5, 2009
  73. Robin RosenbergApr 5, 2009
  74. Nicolas PitreApr 6, 2009
  75. Junio C HamanoApr 6, 2009
  76. Nicolas PitreApr 6, 2009
  77. Jon SmirlApr 6, 2009
  78. Nicolas PitreApr 6, 2009
  79. Jon SmirlApr 6, 2009
  80. Shawn O. PearceApr 6, 2009
  81. Nicolas PitreApr 6, 2009
  82. Jon SmirlApr 6, 2009
  83. Nicolas PitreApr 6, 2009
  84. Matthieu MoyApr 6, 2009
  85. Nicolas PitreApr 6, 2009
  86. Robin H. JohnsonApr 6, 2009
  87. Nicolas PitreApr 6, 2009
  88. Martin LanghoffApr 7, 2009
  89. Jeff KingApr 5, 2009
  90. Robin H. JohnsonApr 5, 2009
  91. Robin H. JohnsonApr 5, 2009
  92. Nguyen Thai Ngoc DuyApr 6, 2009
  93. Nicolas PitreApr 6, 2009
  94. Nicolas PitreApr 6, 2009
  95. Robin H. JohnsonApr 6, 2009
  96. Mark LevedahlApr 11, 2009
  97. Robin H. JohnsonApr 6, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.