Re: Creating objects manually and repack
- From
Jon Smirl <jonsmirl@gmail.com>
- Date
- Aug 4, 2006, 16:11 UTC
- Message-ID
- <9e4733910608040911p443a1360k6d9d1aab00039100@mail.gmail.com>
- In-Reply-To
- <44D36F64.5040404@gmail.com>
On 8/4/06, A Large Angry SCM <gitzilla@gmail.com> wrote:
Show 19 quoted lines
> Jon Smirl wrote: > > On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote: > >> I'd suggest against it, but you can (and should) just repack often enough > >> that you shouldn't ever have gigabytes of objects "in flight". I'd have > >> expected that with a repack every few ten thousand files, and most files > >> being on the order of a few kB, you'd have been more than ok, but > >> especially if you have large files, you may want to make things "every > >> <n> > >> bytes" rather than "every <n> files". > > > > How about forking off a pack-objects and handing it one file name at a > > time over a pipe. When I hand it the next file name I delete the first > > file. Does pack-objects make multiple passes over the files? This > > model would let me hand it all 1M files. > > > > Why don't you just write the pack file directly? Pack files without > deltas have a very simple structure, and git-index-pack will create a > pack index file for the pack file you give it.
That is under consideration but the undeltafied pack is about 12GB and it takes forever (about a day) to deltafy it. I'm not convinced yet that an undeltafied pack is any faster than just having the objects in the directories.
The same data in a deltafied pack is 700MB. That is a tremendous difference in the amount of IO needed. The strategy has to be to avoid IO, nothing I am doing is ever CPU bound.
-- Jon Smirl jonsmirl@gmail.com