git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: kernel.org and GIT tree rebuilding

From
Linus Torvalds <torvalds@osdl.org>
Date
Jun 26, 2005, 19:19 UTC
Message-ID
<Pine.LNX.4.58.0506261206170.19755@ppc970.osdl.org>
In-Reply-To
<7vzmtdq7wy.fsf@assigned-by-dhcp.cox.net>
On Sun, 26 Jun 2005, Junio C Hamano wrote:
> 
> My preference is to do things in this order:
> 
>  (0) concatenate pack and idx files;

Actually, I was originally planning to do that, but now that I have thought about what read_sha1_file() would actually do, I think it's more efficient to leave the index as a separate file.

In particular, what you'd normally do is that if you can't look up the file in the regular object directory, you start going through the pack files. You can do it by having GIT_ALTERNATE_OBJECT_DIRECTORIES point to a pack file, but I actually would prefer the notion of just adding a

	.git/objects/pack

subdirectory, and having object lookup just automatically open and map all index files in that subdirectory.

And the thing is, you really just want to map the index files, the data files can be so big that you can't afford to map them (ie a really big project might have several pack-files a gig each or something like that).

And the most efficient way to map just the index file is to keep it separate, because then the "stat()" will just get the information directly, and you then just mmap that.

The alternative is to first read the index of the index (to figure out how big the index is), and then map the rest. But that just seems a lot messier than just mapping the index file directly.

And when creating these things, we do need to create the data file (which can be big enough that it doesn't fit in memory) first, so we have to have a separate file for it, we can't just stream it out to stdout.

Now, when _sending_ the pack-files, linearizing them is easy: you just send the index first, and the data file immediately afterwards. The index tells how big it is, so there's no need to even add any markers: you can do something like 'git-send-script' with something simple like

	git-rev-list ... | git-pack-file tmp-pack &&
	cat tmp-pack.idx tmp-pack.data | ssh other git-receive-script
So let's just keep the index/data files separate.
		Linus
Previous: Junio C HamanoNext: Junio C Hamano
Message 8 of 38 in “kernel.org and GIT tree rebuilding”
  1. David S. MillerJun 25, 2005
  2. Jeff GarzikJun 25, 2005
  3. Linus TorvaldsJun 25, 2005
  4. Jeff GarzikJun 25, 2005
  5. Linus TorvaldsJun 25, 2005
  6. Linus TorvaldsJun 26, 2005
  7. Junio C HamanoJun 26, 2005
  8. Linus TorvaldsJun 26, 2005
  9. Junio C HamanoJun 26, 2005
  10. Chris MasonJun 26, 2005
  11. Chris MasonJun 26, 2005
  12. Linus TorvaldsJun 26, 2005
  13. Linus TorvaldsJun 26, 2005
  14. Nicolas PitreJun 28, 2005
  15. Linus TorvaldsJun 28, 2005
  16. Nicolas PitreJun 28, 2005
  17. Linus TorvaldsJun 28, 2005
  18. Bugfix: initialize pack_base to NULL.Junio C Hamano, Jun 28, 2005
  19. Nicolas PitreJun 29, 2005
  20. Nicolas PitreJun 29, 2005
  21. Linus TorvaldsJun 29, 2005
  22. Linus TorvaldsJun 29, 2005
  23. Last mile for 1.0 againJunio C Hamano, Jun 29, 2005
  24. Add git-verify-pack command.Junio C Hamano, Jun 29, 2005
  25. Linus TorvaldsJun 29, 2005
  26. Daniel BarkalowJul 4, 2005
  27. Junio C HamanoJul 4, 2005
  28. Linus TorvaldsJul 4, 2005
  29. Daniel BarkalowJul 4, 2005
  30. Junio C HamanoJul 4, 2005
  31. Daniel BarkalowJul 5, 2005
  32. Junio C HamanoJul 5, 2005
  33. Marco CostalbaJul 5, 2005
  34. Junio C HamanoJun 25, 2005
  35. Obtain sha1_file_info() for deltified pack entry properly.Junio C Hamano, Jun 28, 2005
  36. Junio C HamanoJun 28, 2005
  37. 2/3 git-cat-file: use sha1_object_info() on '-t'.Junio C Hamano, Jun 28, 2005
  38. 3/3 git-cat-file: '-s' to find out object size.Junio C Hamano, Jun 28, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.