git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: "git-send-pack"

From
Linus Torvalds <torvalds@osdl.org>
Date
Jul 2, 2005, 01:56 UTC
Message-ID
<Pine.LNX.4.58.0507011831060.2977@ppc970.osdl.org>
In-Reply-To
<42C5D553.80905@timesys.com>
On Fri, 1 Jul 2005, Mike Taht wrote:
Show 5 quoted lines
> 
> You are getting closer and closer to where something like bitTorrent or 
> a multicast protocol makes sense. The problem isn't just the number of 
> outstanding commit objects but the number of machines and developers 
> that want to grab those commits at the same time.

I don't think so. First off, I don't think the decision is kernel- specific, in the sense that I at least use git for sparse and git itself too, so the solution should make sense for small projects as well.

Also, even for the kernel, the total dataset right now (after three months or whatever) is a 60MB pack. It's not like we're sending DVD's or even CD's worth of data around - we're sending the equivalent of 20MB per _month_. That's really not a lot of data. You could easily keep up with a slow modem.

Also, the number of people involved isn't _that_ big. We're talking a few thousand people who actively would update their trees for a big project, and many smaller projects have anything from a couple to maybe a hundred. A few mirrors, and you don't have any problem.

So I think that the problem is actually not that big, and we just need to find an acceptable format. Quite frankly, it might be perfectly acceptable for kernel.org to run a simple packing script once a week which packs everything into one single file, and even if that means that the mirrors will have to re-get everything once a week, that actually sounds acceptable.

It's obviously a _stupid_ way to handle the rsync problem, so there's bound to be some cleaner solution, but the point is that we can probably make mirroring acceptable even with a really really stupid approach. I'd be a bit ashamed of just how ugly it is, but it would likely _work_ fine. You'd create 52 pack-files in a year, but each pack-file is likely just ten megabytes each.

Oh, each pack-file should also be associated with the list of "refs" that were used to generate that pack-file, so make that 104 files per project year (but the list of "refs" would usually be something small, like

	refs/heads/master       4a89a04f1ee21a7c1f4413f1ad7dcfac50ff9b63
	refs/tags/v2.6.11       5dc01c595e6c6ec9ccda4f6f69c131c0dd945f8c
	refs/tags/v2.6.11-tree  5dc01c595e6c6ec9ccda4f6f69c131c0dd945f8c
	refs/tags/v2.6.12       26791a8bcf0e6d33f43aef7682bdb555236d56de
	refs/tags/v2.6.12-rc2   9e734775f7c22d2f89943ad6c745571f1930105f
	refs/tags/v2.6.12-rc3   0397236d43e48e821cce5bbe6a80a1a56bb7cc3a
	refs/tags/v2.6.12-rc4   ebb5573ea8beaf000d4833735f3e53acb9af844c
	refs/tags/v2.6.12-rc5   06f6d9e2f140466eeb41e494e14167f90210f89d
	refs/tags/v2.6.12-rc6   701d7ecec3e0c6b4ab9bb824fd2b34be4da63b7e
	refs/tags/v2.6.13-rc1   733ad933f62e82ebc92fed988c7f0795e64dea62
which was trivially generated from my current tree with
	for i in refs/*/*; do echo -ne $i"\t"; cat $i; done

so now you can use the refs associated with the previous pack-file as the list of refs you're _not_ interested in, and the current list of refs as the list you _are_ interested in, and generate the new pack-file.

Generating the pack-file would literally be something like
	obj=$(git-rev-parse $(cut -f2 new-list) --not $(cut -f2 old-list))
	git-rev-list $obj | git-pack-objects --stdin > new-pack

so a few one-liners like this, run from a cron-job once a week, should just do it.

		Linus
Previous: H. Peter AnvinNext: H. Peter Anvin
Message 20 of 86 in “"git-send-pack"”
  1. Linus TorvaldsJun 30, 2005
  2. A Large Angry SCMJun 30, 2005
  3. A Large Angry SCMJun 30, 2005
  4. Linus TorvaldsJun 30, 2005
  5. Jan HarkesJun 30, 2005
  6. Mike TahtJun 30, 2005
  7. Linus TorvaldsJun 30, 2005
  8. Matthias UrlichsJul 1, 2005
  9. Linus TorvaldsJun 30, 2005
  10. Junio C HamanoJun 30, 2005
  11. Daniel BarkalowJun 30, 2005
  12. Linus TorvaldsJun 30, 2005
  13. H. Peter AnvinJun 30, 2005
  14. Linus TorvaldsJun 30, 2005
  15. H. Peter AnvinJun 30, 2005
  16. Linus TorvaldsJul 1, 2005
  17. H. Peter AnvinJul 1, 2005
  18. Mike TahtJul 1, 2005
  19. H. Peter AnvinJul 2, 2005
  20. Linus TorvaldsJul 2, 2005
  21. H. Peter AnvinJul 2, 2005
  22. Linus TorvaldsJul 2, 2005
  23. H. Peter AnvinJul 2, 2005
  24. Linus TorvaldsJul 2, 2005
  25. H. Peter AnvinJul 2, 2005
  26. Tony LuckJul 2, 2005
  27. H. Peter AnvinJul 2, 2005
  28. A Large Angry SCMJul 2, 2005
  29. Daniel BarkalowJun 30, 2005
  30. Linus TorvaldsJun 30, 2005
  31. Daniel BarkalowJul 1, 2005
  32. Linus TorvaldsJun 30, 2005
  33. Dan HolmsandJun 30, 2005
  34. Daniel BarkalowJun 30, 2005
  35. Linus TorvaldsJun 30, 2005
  36. H. Peter AnvinJun 30, 2005
  37. Linus TorvaldsJun 30, 2005
  38. H. Peter AnvinJun 30, 2005
  39. H. Peter AnvinJun 30, 2005
  40. Linus TorvaldsJun 30, 2005
  41. H. Peter AnvinJun 30, 2005
  42. Matthias UrlichsJul 1, 2005
  43. Jan HarkesJul 1, 2005
  44. TagsEric W. Biederman, Jul 1, 2005
  45. H. Peter AnvinJul 1, 2005
  46. Eric W. BiedermanJul 1, 2005
  47. H. Peter AnvinJul 1, 2005
  48. Eric W. BiedermanJul 1, 2005
  49. Daniel BarkalowJul 1, 2005
  50. H. Peter AnvinJul 2, 2005
  51. Eric W. BiedermanJul 2, 2005
  52. H. Peter AnvinJul 2, 2005
  53. Eric W. BiedermanJul 2, 2005
  54. H. Peter AnvinJul 2, 2005
  55. Eric W. BiedermanJul 2, 2005
  56. Matthias UrlichsJul 2, 2005
  57. H. Peter AnvinJul 2, 2005
  58. Linus TorvaldsJul 2, 2005
  59. H. Peter AnvinJul 2, 2005
  60. A Large Angry SCMJul 2, 2005
  61. Linus TorvaldsJul 2, 2005
  62. A Large Angry SCMJul 2, 2005
  63. Linus TorvaldsJul 3, 2005
  64. Petr BaudisJul 2, 2005
  65. Linus TorvaldsJul 2, 2005
  66. Dan HolmsandJul 3, 2005
  67. Kevin SmithJul 3, 2005
  68. Eric W. BiedermanJul 5, 2005
  69. Daniel BarkalowJul 5, 2005
  70. Eric W. BiedermanJul 5, 2005
  71. Linus TorvaldsJul 5, 2005
  72. Junio C HamanoJul 5, 2005
  73. Matthias UrlichsJul 6, 2005
  74. Eric W. BiedermanJul 7, 2005
  75. Linus TorvaldsJul 2, 2005
  76. Jan HarkesJul 2, 2005
  77. Jan HarkesJul 2, 2005
  78. Matthias UrlichsJul 2, 2005
  79. Petr BaudisJul 1, 2005
  80. H. Peter AnvinJul 1, 2005
  81. Matthias UrlichsJul 1, 2005
  82. Petr BaudisJul 1, 2005
  83. H. Peter AnvinJul 1, 2005
  84. Daniel BarkalowJul 1, 2005
  85. Petr BaudisJul 1, 2005
  86. Daniel BarkalowJun 30, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.