Re: kernel.org and GIT tree rebuilding
- From
Linus Torvalds <torvalds@osdl.org>
- Date
- Jun 26, 2005, 19:19 UTC
- Message-ID
- <Pine.LNX.4.58.0506261206170.19755@ppc970.osdl.org>
- In-Reply-To
- <7vzmtdq7wy.fsf@assigned-by-dhcp.cox.net>
On Sun, 26 Jun 2005, Junio C Hamano wrote:
> > My preference is to do things in this order: > > (0) concatenate pack and idx files;
Actually, I was originally planning to do that, but now that I have thought about what read_sha1_file() would actually do, I think it's more efficient to leave the index as a separate file.
In particular, what you'd normally do is that if you can't look up the file in the regular object directory, you start going through the pack files. You can do it by having GIT_ALTERNATE_OBJECT_DIRECTORIES point to a pack file, but I actually would prefer the notion of just adding a
.git/objects/pack
subdirectory, and having object lookup just automatically open and map all index files in that subdirectory.
And the thing is, you really just want to map the index files, the data files can be so big that you can't afford to map them (ie a really big project might have several pack-files a gig each or something like that).
And the most efficient way to map just the index file is to keep it separate, because then the "stat()" will just get the information directly, and you then just mmap that.
The alternative is to first read the index of the index (to figure out how big the index is), and then map the rest. But that just seems a lot messier than just mapping the index file directly.
And when creating these things, we do need to create the data file (which can be big enough that it doesn't fit in memory) first, so we have to have a separate file for it, we can't just stream it out to stdout.
Now, when _sending_ the pack-files, linearizing them is easy: you just send the index first, and the data file immediately afterwards. The index tells how big it is, so there's no need to even add any markers: you can do something like 'git-send-script' with something simple like
git-rev-list ... | git-pack-file tmp-pack && cat tmp-pack.idx tmp-pack.data | ssh other git-receive-script
So let's just keep the index/data files separate.
Linus