git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git pack/unpack over bittorrent - works!

From
LLLuke Kenneth Casson Leighton <luke.leighton@gmail.com>
Date
Sep 3, 2010, 10:54 UTC
Message-ID
<AANLkTimoz6Ux18mtmm4oVr21p6RmKURgcWWK8JkWvmu5@mail.gmail.com>
In-Reply-To
<alpine.LFD.2.00.1009021931340.19366@xanadu.home>
have a bit more time.
On Fri, Sep 3, 2010 at 1:29 AM, Nicolas Pitre <nico@fluxnic.net> wrote:
Show 6 quoted lines
>> pack, you'd have to _delete_ them from the receiving end so as to
>> avoid polluting the recipient's object store haaarrgh *spit*, *cough*.
>
> Well, actually there is no need to delete anything.  Git can cope with
> duplicated objects just fine.  A subsequent gc will get rid of the
> duplicates automatically.
 excellent.  good to hear.
Show 29 quoted lines
>>  what _might_ work however iiiiIiis... to split the pack-object into
>> two parts.  or, to add an "extra part", to be more precise:
>>
>> a) complete list of all objects.  _just_ the list of objects.
>> b) existing pack-object format/structure.
>>
>> in this way, the sender having done all the hard work already of
>> determining what objects are to go into a pack-object, transfers that
>> *first*.  _theeen_ you begin transferring the pack-object.  theeeen,
>> if the pack-object transfer is ever interrupted, you simply send back
>> that list of objects, and ask "uhh, you know that list of objects we
>> were talking about?  well, here it is *splat* - are you able to
>> recreate the pack-object from that, for me, and if so please gimme
>> again"
>
> Well, it isn't that simple.
>
> First, a resumable clone is useful only when there is a big transfer in
> play.  Otherwise it isn't worth the trouble.
>
> So, if the clone is big, then this list of objects can be in the
> millions.  For example my linux kernel repo with a couple branches
> currently has:
>
> $ git rev-list --all --objects | wc -l
> 2808136
>
> So, 2808136 objects, with 20-byte SHA1 for each of them, and you have a
> 54 MB object list to transfer already.
 ok:
 a) that's fine.  first time, you have to do that, you have to do that.
 b) i have some ideas in mind, to say things like "i already have the
following objects up to here, please give me a list of everything
since then".
 > And even then, what if the transfer crashes during that object list
> transfer?  On flaky connections this might happen within 54 MB.
 that's fine: i envisage the object list being cached at the remote
end (by the first seed), and also being a "shared file", such that
there may even be complete copies of that "file" out there already,
such that resumption is a non-issue.
Show 12 quoted lines
>> and, 10^N-1 times out of 10^N, for reasons that shawn kindly
>> explained, i bet you the answer would be "yes".
>
> For the list of objects, sure.  But that isn't a big deal.  It is easy
> enough to tell the remote about the commits we already have and ask for
> the rest.  With a commit SHA1, the remote can figure out all the objects
> we have. But all is in that determination of the latest commit we have.
> If we get a partial pack, it is possible to somehow salvage as many
> objects from it, and determine what top commit(s) that correspond to.
> It is possible to set your local repo just as if you had requested a
> shallow clone and then the resume would simply be a deepening of that
> shallow clone.
 i'll need to re-read this when i have more time.  apologies.
> Another issue is what to do with objects that are themselves huge.
 that's fine, too: in fact, that's the perfect scenario where a
file-sharing protocol excels.
Show 7 quoted lines
> Yet another issue: what to do with all those objects I've got in my
> partial pack, but that I can't connect to any commit yet.  We don't want
> them transferred again but it isn't easy to tell the remote about them.
>
> You could tell the remote: "I have this pack for this commit from this
> commit but I got only this amount of bytes from it, please resume
> transfer here."
 ok i have a couple of ideas/thoughts

a) one of which was to send the commit index list back to the remote end, but that would be baaaaad as it could be 54mb as you say, so it would be necessary to say "here is the SHA1 of the index file you gave me earlier, do you still have it, if so please can we resume

b) as long as _somebody_ has a complete copy, distributed throughout the file-sharing network, of that complete pack, "resume transfer here" isn't .... the concept is moot. the bittorrent protocol covers that concept of "resume" very very easily.

> But as mentioned before the pack stream is not
> deterministic,
 one a one-off basis, it is; and even then, i believe that it could be
made to "not matter".  so you ask a server for a pack object and get a
different SHA-1?  so what, you just make that part of the
file-sharing-network unique key: {ref}-{objref}-{SHA-1} instead of
just {ref}-{objref}.  if the connection's lost, wow big deal, you just
ask again and you end up with a different SHA-1.  you're back to a
situation which is actually no different from and no less efficient
than the present http transfer system.
 ... but i'd rather avoid this scenario, if possible.
Show 7 quoted lines
> and we really don't want to make it single-threaded on a
> server.  Furthermore this is a lot of work for the server as even if the
> pack stream is deterministic, then the server still has to recreate the
> first part of the pack just to throw it away until the desired offset is
> reached.  And caching pack results also has all sorts of implications
> we've prefered to avoid on a server for security reasons (better keep
> serving operations read-only).
 i've already thrown out the idea of cacheing the pack objects
themselves, but am still exploring the concept of cacheing the .idx
file, even for short periods of time.
 so the server does a lot of work creating that .idx file, but it
contains the complete list of all objects, which you _could_ just
obtain again by just asking explicitly for each and every single one
of those objects, no more, no less, no deltas, no windows - the list,
the whole list and nothing but the list.
Show 7 quoted lines
>> ... um... in fact... um... i believe i'm merely talking about the .idx
>> index file, aren't i?  because... um... the index file contains the
>> list of object refs in the pack, yes?
>
> In one pack, yes.  You might have multiple packs.  And that doesn't mean
> that all the objects from a pack are all relevant to the actual branches
> you are willing to export.
 yes that's fine.  multiple packs are considered to be independent
files of the "VFS layer" in the file-sharing network.  that's taken
care of.  what i need to know is: can you recreate a pack object given
the list of objects in its .idx file?
Show 8 quoted lines
>> sooo.... taking a wild guess, here: if you were to parse the .idx file
>> and extract the list of object-refs, and then pass that to "git
>> pack-objects --window=0 --delta=0", would you end up with the exact
>> same pack file, because you'd forced git pack-objects to only return
>> that specific list of object-refs?
>
> If you do this i.e. turn off delta compression, then the 615 MB
> repository above will turn itself into a multi-gigabyte pack!
 ok this was covered in my previous post, hope it's clearer.  perhaps
"git pack-objects --window=0 --delta=0 <
{list-of-objects-extracted-from-the-idx-file}" isn't the way to
achieve what i envisage - if not, does anyone have any ideas on how
extracting the exact list of objects as previously given by a .idx
file can be achieved?
l.
Previous: Luke Kenneth Casson LeightonNext: Luke Kenneth Casson Leighton
Message 30 of 88 in “git pack/unpack over bittorrent - works!”
  1. Luke Kenneth Casson LeightonSep 1, 2010
  2. Nguyen Thai Ngoc DuySep 1, 2010
  3. Luke Kenneth Casson LeightonSep 2, 2010
  4. Luke Kenneth Casson LeightonSep 2, 2010
  5. Ævar Arnfjörð BjarmasonSep 2, 2010
  6. A Large Angry SCMSep 2, 2010
  7. Luke Kenneth Casson LeightonSep 2, 2010
  8. Luke Kenneth Casson LeightonSep 2, 2010
  9. A Large Angry SCMSep 2, 2010
  10. Jeff KingSep 2, 2010
  11. Nicolas PitreSep 2, 2010
  12. A Large Angry SCMSep 2, 2010
  13. Nicolas PitreSep 2, 2010
  14. Luke Kenneth Casson LeightonSep 2, 2010
  15. Shawn O. PearceSep 2, 2010
  16. Luke Kenneth Casson LeightonSep 2, 2010
  17. Luke Kenneth Casson LeightonSep 2, 2010
  18. Nicolas PitreSep 3, 2010
  19. Luke Kenneth Casson LeightonSep 3, 2010
  20. Junio C HamanoSep 3, 2010
  21. Brandon CaseySep 2, 2010
  22. Luke Kenneth Casson LeightonSep 2, 2010
  23. Jakub NarebskiSep 2, 2010
  24. Luke Kenneth Casson LeightonSep 2, 2010
  25. Luke Kenneth Casson LeightonSep 2, 2010
  26. Nicolas PitreSep 3, 2010
  27. Nguyen Thai Ngoc DuySep 3, 2010
  28. Luke Kenneth Casson LeightonSep 3, 2010
  29. Luke Kenneth Casson LeightonSep 3, 2010
  30. Luke Kenneth Casson LeightonSep 3, 2010
  31. Luke Kenneth Casson LeightonSep 2, 2010
  32. Casey DahlinSep 2, 2010
  33. A Large Angry SCMSep 2, 2010
  34. Nicolas PitreSep 2, 2010
  35. Luke Kenneth Casson LeightonSep 2, 2010
  36. A Large Angry SCMSep 2, 2010
  37. Nicolas PitreSep 2, 2010
  38. Theodore TsoSep 3, 2010
  39. Luke Kenneth Casson LeightonSep 3, 2010
  40. Junio C HamanoSep 3, 2010
  41. Ted Ts'oSep 3, 2010
  42. Nicolas PitreSep 3, 2010
  43. Luke Kenneth Casson LeightonSep 3, 2010
  44. Nguyen Thai Ngoc DuySep 4, 2010
  45. Nguyen Thai Ngoc DuySep 4, 2010
  46. Artur SkawinaSep 4, 2010
  47. Nicolas PitreSep 4, 2010
  48. Artur SkawinaSep 4, 2010
  49. Nicolas PitreSep 4, 2010
  50. Luke Kenneth Casson LeightonSep 4, 2010
  51. Luke Kenneth Casson LeightonSep 4, 2010
  52. Nicolas PitreSep 5, 2010
  53. Luke Kenneth Casson LeightonSep 5, 2010
  54. Nicolas PitreSep 5, 2010
  55. Luke Kenneth Casson LeightonSep 6, 2010
  56. Nicolas PitreSep 6, 2010
  57. Luke Kenneth Casson LeightonSep 6, 2010
  58. Junio C HamanoSep 6, 2010
  59. Nicolas PitreSep 6, 2010
  60. Luke Kenneth Casson LeightonSep 7, 2010
  61. Luke Kenneth Casson LeightonSep 7, 2010
  62. Artur SkawinaSep 4, 2010
  63. Theodore TsoSep 4, 2010
  64. Kyle MoffettSep 4, 2010
  65. Theodore TsoSep 4, 2010
  66. Luke Kenneth Casson LeightonSep 4, 2010
  67. Nicolas PitreSep 5, 2010
  68. Luke Kenneth Casson LeightonSep 5, 2010
  69. Nicolas PitreSep 4, 2010
  70. Theodore TsoSep 4, 2010
  71. Luke Kenneth Casson LeightonSep 4, 2010
  72. Luke Kenneth Casson LeightonSep 4, 2010
  73. Ted Ts'oSep 4, 2010
  74. Luke Kenneth Casson LeightonSep 4, 2010
  75. Ted Ts'oSep 4, 2010
  76. Luke Kenneth Casson LeightonSep 5, 2010
  77. Jakub NarebskiSep 4, 2010
  78. Luke Kenneth Casson LeightonSep 4, 2010
  79. Jakub NarebskiSep 4, 2010
  80. Luke Kenneth Casson LeightonSep 4, 2010
  81. Ted Ts'oSep 4, 2010
  82. Tomas CarneckySep 5, 2010
  83. Nicolas PitreSep 5, 2010
  84. Luke Kenneth Casson LeightonSep 5, 2010
  85. Nicolas PitreSep 6, 2010
  86. Luke Kenneth Casson LeightonSep 4, 2010
  87. Artur SkawinaSep 4, 2010
  88. Artur SkawinaSep 4, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.