git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] multi item packed files

From
CMChris Mason <mason@suse.com>
Date
Apr 22, 2005, 20:32 UTC
Message-ID
<200504221632.26278.mason@suse.com>
In-Reply-To
<Pine.LNX.4.58.0504221230020.2344@ppc970.osdl.org>
On Friday 22 April 2005 15:43, Linus Torvalds wrote:
Show 11 quoted lines
> On Fri, 22 Apr 2005, Chris Mason wrote:
> > The problem I see for git is that once you have enough data, it should
> > degrade over and over again somewhat quickly.
>
> I really doubt that.
>
> There's a more or less constant amount of new data added all the time: the
> number of changes does _not_ grow with history. The number of changes
> grows with the amount of changes going on in the tree, and while that
> isn't exactly constant, it definitely is not something that grows very
> fast.
>From a filesystem point of view, it's not the number of changes that matters, 

it's the distance between them. The amount of new data is constant, but the speed of accessing the new data is affected by the bulk of old data on disk.

Even with defragging you hopefully end up with a big chunk of the disk where everything is in order. Then you add a new file and it goes either somewhere behind that big chunk or in front of it. The next new file might go somewhere behind or in front etc etc. Having a big chunk just means the new files are likely to be farther apart making reads of the new data very seeky.

Show 8 quoted lines
>
> Btw, this is how git is able to be so fast in the first place. Git is fast
> because it knows that the "size of the change" is a lot smaller than the
> "size of the repository", so it fundamentally at all points tries to make
> sure that it only ever bothers with stuff that has changed.
>
> Stuff that hasn't changed, it ignores very _very_ efficiently.
>
git as a write engine is very fast, and we definitely write more then we read.
Show 5 quoted lines
> > I grabbed Ingo's tarball of 28,000 patches since 2.4.0 and applied them
> > all into git on ext3 (htree).  It only took ~2.5 hrs to apply.
>
> Ok, I'd actually wish it took even less, but that's still a pretty
> impressive average of three patches a second.

Yeah, and this was a relatively old machine with slowish drives. One run to apply into my packed tree is finished and only took 2 hours. But, I had 'tuned' it to make bigger packed files, and the end result is 2MB compressed objects. Great for compression rate, but my dumb format doesn't hold up well for reading it back.

If I pack every 64k (uncompressed), the checkout-tree time goes down to 3m14s. That's a very big difference considering how stupid my code is .git was only 20% smaller with 64k chunks. I should be able to do better...I'll do one more run.

Show 15 quoted lines
>
> > Anyway, I ended up with a 2.6GB .git directory.  Then I:
> >
> > rm .git/index
> > umount ; mount again
> > time read-tree `tree-id` (24.45s)
> > time checkout-cache --prefix=../checkout/ -a -f (4m30s)
> >
> > --prefix is neat ;)
>
> That sounds pretty acceptable. Four minutes is a long time, but I assume
> that the whole point of the exercise was to try to test worst-case
> behaviour.  We can certainly make sure that real usage gets lower numbers
> than that (in particular, my "real usage" ends up being 100% in the disk
> cache ;)

I had a tree with 28,000 patches. If we pretend that one bk changeset will equal one git changeset, we'd have 64,000 patches (57k without empty mergesets), and it probably wouldn't fit into ram anymore ;) Our bk cset rate was about 24k/year, so we'll have to trim very aggressively to have reasonable performance.

For a working tree that's fine, but we need some fast central place to pull the working .git trees from, and we're really going to feel the random io there.

-chris
Previous: Linus TorvaldsNext: Chris Mason
Message 14 of 17 in “multi item packed files”
  1. multi item packed filesChris Mason, Apr 21, 2005
  2. Linus TorvaldsApr 21, 2005
  3. Chris MasonApr 21, 2005
  4. Krzysztof HalasaApr 21, 2005
  5. Linus TorvaldsApr 21, 2005
  6. Krzysztof HalasaApr 22, 2005
  7. Martin UeckerApr 22, 2005
  8. Chris MasonApr 21, 2005
  9. Linus TorvaldsApr 21, 2005
  10. Chris MasonApr 22, 2005
  11. Linus TorvaldsApr 22, 2005
  12. Chris MasonApr 22, 2005
  13. Linus TorvaldsApr 22, 2005
  14. Chris MasonApr 22, 2005
  15. Chris MasonApr 22, 2005
  16. Chris MasonApr 25, 2005
  17. Krzysztof HalasaApr 22, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.