git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git pack/unpack over bittorrent - works!

From
Nicolas Pitre <nico@fluxnic.net>
Date
Sep 3, 2010, 00:29 UTC
Message-ID
<alpine.LFD.2.00.1009021931340.19366@xanadu.home>
In-Reply-To
<AANLkTikSHXivniUk-1KU30Ws23ebnbDhOmjKmpmVH-Y9@mail.gmail.com>
On Thu, 2 Sep 2010, Luke Kenneth Casson Leighton wrote:
Show 15 quoted lines
> On Thu, Sep 2, 2010 at 9:45 PM, Jakub Narebski <jnareb@gmail.com> wrote:
> 
> > If I remember the discussion stalled (i.e. no working implementation),
> > and one of the latest proposals was to have some way of recovering
> > objects from partially downloaded file, and a way to request packfile
> > without objects that got already downloaded.
> 
>  oo.  ouch.  i can understand why things stalled, then.  you're
> effectively adding an extra layer in, and even if you could add a
> unique naming scheme on those objects (if one doesn't already exist?),
> those object might (or might not!) come up the second time round (for
> reasons mentioned already - threads resulting in different deltas
> being picked etc.) ... and if they weren't picked for the re-generated
> pack, you'd have to _delete_ them from the receiving end so as to
> avoid polluting the recipient's object store haaarrgh *spit*, *cough*.

Well, actually there is no need to delete anything. Git can cope with duplicated objects just fine. A subsequent gc will get rid of the duplicates automatically.

Show 14 quoted lines
>  what _might_ work however iiiiIiis... to split the pack-object into
> two parts.  or, to add an "extra part", to be more precise:
> 
> a) complete list of all objects.  _just_ the list of objects.
> b) existing pack-object format/structure.
> 
> in this way, the sender having done all the hard work already of
> determining what objects are to go into a pack-object, transfers that
> *first*.  _theeen_ you begin transferring the pack-object.  theeeen,
> if the pack-object transfer is ever interrupted, you simply send back
> that list of objects, and ask "uhh, you know that list of objects we
> were talking about?  well, here it is *splat* - are you able to
> recreate the pack-object from that, for me, and if so please gimme
> again"
Well, it isn't that simple.

First, a resumable clone is useful only when there is a big transfer in play. Otherwise it isn't worth the trouble.

So, if the clone is big, then this list of objects can be in the millions. For example my linux kernel repo with a couple branches currently has:

$ git rev-list --all --objects | wc -l 2808136

So, 2808136 objects, with 20-byte SHA1 for each of them, and you have a 54 MB object list to transfer already. This is a significant overhead that we prefer to avoid, given the actual pack transfer which is:

$ git pack-objects --all --stdout --progress < /dev/null | wc -c Counting objects: 2808136, done. Compressing objects: 100% (384219/384219), done. 645201934 Total 2808136 (delta 2422420), reused 2788225 (delta 2402700)

The output from wc is 645201934 = 615 MB for this repository. Hence the list of object alone is quite significant.

And even then, what if the transfer crashes during that object list transfer? On flaky connections this might happen within 54 MB.

> and, 10^N-1 times out of 10^N, for reasons that shawn kindly
> explained, i bet you the answer would be "yes".

For the list of objects, sure. But that isn't a big deal. It is easy enough to tell the remote about the commits we already have and ask for the rest. With a commit SHA1, the remote can figure out all the objects we have. But all is in that determination of the latest commit we have. If we get a partial pack, it is possible to somehow salvage as many objects from it, and determine what top commit(s) that correspond to. It is possible to set your local repo just as if you had requested a shallow clone and then the resume would simply be a deepening of that shallow clone.

But usually the very first commit in a pack is huge as it typically isn't delta compressed (a delta chain has to start somewhere). And this first commit will roughly represent the same size as a tarball for that commit. And if you don't get at least that first commit then you are screwed. Or if you don't get a complete second commit when deepening a clone you are still screwed.

Another issue is what to do with objects that are themselves huge.

Yet another issue: what to do with all those objects I've got in my partial pack, but that I can't connect to any commit yet. We don't want them transferred again but it isn't easy to tell the remote about them.

You could tell the remote: "I have this pack for this commit from this commit but I got only this amount of bytes from it, please resume transfer here." But as mentioned before the pack stream is not deterministic, and we really don't want to make it single-threaded on a server. Furthermore this is a lot of work for the server as even if the pack stream is deterministic, then the server still has to recreate the first part of the pack just to throw it away until the desired offset is reached. And caching pack results also has all sorts of implications we've prefered to avoid on a server for security reasons (better keep serving operations read-only).

> ... um... in fact... um... i believe i'm merely talking about the .idx
> index file, aren't i?  because... um... the index file contains the
> list of object refs in the pack, yes?

In one pack, yes. You might have multiple packs. And that doesn't mean that all the objects from a pack are all relevant to the actual branches you are willing to export.

Show 5 quoted lines
> sooo.... taking a wild guess, here: if you were to parse the .idx file
> and extract the list of object-refs, and then pass that to "git
> pack-objects --window=0 --delta=0", would you end up with the exact
> same pack file, because you'd forced git pack-objects to only return
> that specific list of object-refs?

If you do this i.e. turn off delta compression, then the 615 MB repository above will turn itself into a multi-gigabyte pack!

Nicolas
Previous: Luke Kenneth Casson LeightonNext: Nguyen Thai Ngoc Duy
Message 26 of 88 in “git pack/unpack over bittorrent - works!”
  1. Luke Kenneth Casson LeightonSep 1, 2010
  2. Nguyen Thai Ngoc DuySep 1, 2010
  3. Luke Kenneth Casson LeightonSep 2, 2010
  4. Luke Kenneth Casson LeightonSep 2, 2010
  5. Ævar Arnfjörð BjarmasonSep 2, 2010
  6. A Large Angry SCMSep 2, 2010
  7. Luke Kenneth Casson LeightonSep 2, 2010
  8. Luke Kenneth Casson LeightonSep 2, 2010
  9. A Large Angry SCMSep 2, 2010
  10. Jeff KingSep 2, 2010
  11. Nicolas PitreSep 2, 2010
  12. A Large Angry SCMSep 2, 2010
  13. Nicolas PitreSep 2, 2010
  14. Luke Kenneth Casson LeightonSep 2, 2010
  15. Shawn O. PearceSep 2, 2010
  16. Luke Kenneth Casson LeightonSep 2, 2010
  17. Luke Kenneth Casson LeightonSep 2, 2010
  18. Nicolas PitreSep 3, 2010
  19. Luke Kenneth Casson LeightonSep 3, 2010
  20. Junio C HamanoSep 3, 2010
  21. Brandon CaseySep 2, 2010
  22. Luke Kenneth Casson LeightonSep 2, 2010
  23. Jakub NarebskiSep 2, 2010
  24. Luke Kenneth Casson LeightonSep 2, 2010
  25. Luke Kenneth Casson LeightonSep 2, 2010
  26. Nicolas PitreSep 3, 2010
  27. Nguyen Thai Ngoc DuySep 3, 2010
  28. Luke Kenneth Casson LeightonSep 3, 2010
  29. Luke Kenneth Casson LeightonSep 3, 2010
  30. Luke Kenneth Casson LeightonSep 3, 2010
  31. Luke Kenneth Casson LeightonSep 2, 2010
  32. Casey DahlinSep 2, 2010
  33. A Large Angry SCMSep 2, 2010
  34. Nicolas PitreSep 2, 2010
  35. Luke Kenneth Casson LeightonSep 2, 2010
  36. A Large Angry SCMSep 2, 2010
  37. Nicolas PitreSep 2, 2010
  38. Theodore TsoSep 3, 2010
  39. Luke Kenneth Casson LeightonSep 3, 2010
  40. Junio C HamanoSep 3, 2010
  41. Ted Ts'oSep 3, 2010
  42. Nicolas PitreSep 3, 2010
  43. Luke Kenneth Casson LeightonSep 3, 2010
  44. Nguyen Thai Ngoc DuySep 4, 2010
  45. Nguyen Thai Ngoc DuySep 4, 2010
  46. Artur SkawinaSep 4, 2010
  47. Nicolas PitreSep 4, 2010
  48. Artur SkawinaSep 4, 2010
  49. Nicolas PitreSep 4, 2010
  50. Luke Kenneth Casson LeightonSep 4, 2010
  51. Luke Kenneth Casson LeightonSep 4, 2010
  52. Nicolas PitreSep 5, 2010
  53. Luke Kenneth Casson LeightonSep 5, 2010
  54. Nicolas PitreSep 5, 2010
  55. Luke Kenneth Casson LeightonSep 6, 2010
  56. Nicolas PitreSep 6, 2010
  57. Luke Kenneth Casson LeightonSep 6, 2010
  58. Junio C HamanoSep 6, 2010
  59. Nicolas PitreSep 6, 2010
  60. Luke Kenneth Casson LeightonSep 7, 2010
  61. Luke Kenneth Casson LeightonSep 7, 2010
  62. Artur SkawinaSep 4, 2010
  63. Theodore TsoSep 4, 2010
  64. Kyle MoffettSep 4, 2010
  65. Theodore TsoSep 4, 2010
  66. Luke Kenneth Casson LeightonSep 4, 2010
  67. Nicolas PitreSep 5, 2010
  68. Luke Kenneth Casson LeightonSep 5, 2010
  69. Nicolas PitreSep 4, 2010
  70. Theodore TsoSep 4, 2010
  71. Luke Kenneth Casson LeightonSep 4, 2010
  72. Luke Kenneth Casson LeightonSep 4, 2010
  73. Ted Ts'oSep 4, 2010
  74. Luke Kenneth Casson LeightonSep 4, 2010
  75. Ted Ts'oSep 4, 2010
  76. Luke Kenneth Casson LeightonSep 5, 2010
  77. Jakub NarebskiSep 4, 2010
  78. Luke Kenneth Casson LeightonSep 4, 2010
  79. Jakub NarebskiSep 4, 2010
  80. Luke Kenneth Casson LeightonSep 4, 2010
  81. Ted Ts'oSep 4, 2010
  82. Tomas CarneckySep 5, 2010
  83. Nicolas PitreSep 5, 2010
  84. Luke Kenneth Casson LeightonSep 5, 2010
  85. Nicolas PitreSep 6, 2010
  86. Luke Kenneth Casson LeightonSep 4, 2010
  87. Artur SkawinaSep 4, 2010
  88. Artur SkawinaSep 4, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.