git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git pack/unpack over bittorrent - works!

From
Nicolas Pitre <nico@fluxnic.net>
Date
Sep 2, 2010, 23:09 UTC
Message-ID
<alpine.LFD.2.00.1009021624170.19366@xanadu.home>
In-Reply-To
<AANLkTinFPxsY6frVnga8u15aovQarfWreBYJfri6ywoK@mail.gmail.com>
On Thu, 2 Sep 2010, Luke Kenneth Casson Leighton wrote:
> nicolas, thanks for responding: you'll see this some time in the
> future when you catch up, it's not a high priority, nothing new, just
> thinking out loud, for benefit of archives.
Well, I might as well pay more attention to *you* now.  :-)
Show 7 quoted lines
> >> * is it possible to _make_ the repository guaranteed to produce
> >> identical pack objects?
> >
> > Sure, but performance will suck.
> 
>  that's fiiine :)  as i've learned on the pyjamas project, it's rare
> that you have speed and interoperability at the same time...

Well, did you hear about this thing called Git? It appears that those Git developers are performance freaks. :-) Yet, Git is interoperable across almost all versions ever released because we made sure that only fundamental things are defined and relied upon. And that excludes actual delta pairing and pack object ordering. That's why a pack file may have many different byte sequences and yet still represent the same canonical data.

Show 5 quoted lines
>  if the pack-objects are going to vary, then the VFS layer idea is
> blown completely out the water, except for the absolute basic
> meta-info such as "refs/heads/*".  so i might as well just use
> "actual" bittorrent to transfer packs via
> {ref}-{commitref}-{SHA-1}.torrent.
For the archive benefit, here's what I think on the whole idea.

The BitTorrent model is simply unappropriate for Git. It doesn't fit to the Git model at all as BitTorrent works on stable and static data, and requires a lot of people wanting that same data.

When you perform a fetch, Git does actually negociate with the server to figure out what's missing locally, and the server does produce a custom pack for you that is optimized so that only what's needed for you to be up to date is transferred.

Even if you try to cache a set of packs to suit the BitTorrent static data model, you'll need so many packs to cover all the possible gaps between a server and a random number of clients each with a random repository state. Of course it is possible to have bigger packs covering larger gaps, but then you lose the biggest advantage that the smart Git protocol has. And with smaller, more fine grained packs, you'll end up with so many of them that finding a live torrent for the actual one you need is going to be difficult.

> ho hum, drawing board we come...

Yep. Instead of transferring packs, a BitTorrent-alike transfer should be based on the transfer of _objects_. Therefore you can make the correspondance between file chunks in BitTorrent with objects in a Git aware system. So, when contacting a peer, you could negociate what is the set of objects that the peer has that you don't, and vice versa. Objects in Git are stable and immutable, and they all have a unique SHA1 signature. And to optimize the negociation, the pack index content can be used, first by exchanging the content of the first level fan-out table and ignoring those entries that are equal. This for each peer.

Then, each peer make requests to connected peers for objects that those peers have but that isn't available locally, just like chunks in BitTorrent.

But here's the twist to make this scale well. Since the object sender knows what objects the receiver already has, it therefore can choose the object encoding. Meaning that the sender can simply *reuse* a delta encoding for an object it is requested to send if the requestor already has the base object for this delta.

So in most cases, the object to send will be small, especially if it is a delta object. That should fit the chunk model. But if an object is bigger than a certain treshold, then its transfer could be chunked across multiple peers just like classic BitTorrent. In this case, the chunking would need to be done on the non delta uncompressed object data as this is the only thing that is universally stable (doesn't mean that the _transfer_ of those chunks can't be compressed).

Now this design has many open questions, such as finding out what is the latest set of refs amongst all the peers, whether or not what we have locally are ancestors of the remote refs, etc.

And of course, while this will make for a speedy object transfer, the resulting mess on the receiver's end will have to be validated and repacked in the end. So overall this might not end up being faster overall for the fetcher.

Nicolas
Previous: A Large Angry SCMNext: Theodore Tso
Message 37 of 88 in “git pack/unpack over bittorrent - works!”
  1. Luke Kenneth Casson LeightonSep 1, 2010
  2. Nguyen Thai Ngoc DuySep 1, 2010
  3. Luke Kenneth Casson LeightonSep 2, 2010
  4. Luke Kenneth Casson LeightonSep 2, 2010
  5. Ævar Arnfjörð BjarmasonSep 2, 2010
  6. A Large Angry SCMSep 2, 2010
  7. Luke Kenneth Casson LeightonSep 2, 2010
  8. Luke Kenneth Casson LeightonSep 2, 2010
  9. A Large Angry SCMSep 2, 2010
  10. Jeff KingSep 2, 2010
  11. Nicolas PitreSep 2, 2010
  12. A Large Angry SCMSep 2, 2010
  13. Nicolas PitreSep 2, 2010
  14. Luke Kenneth Casson LeightonSep 2, 2010
  15. Shawn O. PearceSep 2, 2010
  16. Luke Kenneth Casson LeightonSep 2, 2010
  17. Luke Kenneth Casson LeightonSep 2, 2010
  18. Nicolas PitreSep 3, 2010
  19. Luke Kenneth Casson LeightonSep 3, 2010
  20. Junio C HamanoSep 3, 2010
  21. Brandon CaseySep 2, 2010
  22. Luke Kenneth Casson LeightonSep 2, 2010
  23. Jakub NarebskiSep 2, 2010
  24. Luke Kenneth Casson LeightonSep 2, 2010
  25. Luke Kenneth Casson LeightonSep 2, 2010
  26. Nicolas PitreSep 3, 2010
  27. Nguyen Thai Ngoc DuySep 3, 2010
  28. Luke Kenneth Casson LeightonSep 3, 2010
  29. Luke Kenneth Casson LeightonSep 3, 2010
  30. Luke Kenneth Casson LeightonSep 3, 2010
  31. Luke Kenneth Casson LeightonSep 2, 2010
  32. Casey DahlinSep 2, 2010
  33. A Large Angry SCMSep 2, 2010
  34. Nicolas PitreSep 2, 2010
  35. Luke Kenneth Casson LeightonSep 2, 2010
  36. A Large Angry SCMSep 2, 2010
  37. Nicolas PitreSep 2, 2010
  38. Theodore TsoSep 3, 2010
  39. Luke Kenneth Casson LeightonSep 3, 2010
  40. Junio C HamanoSep 3, 2010
  41. Ted Ts'oSep 3, 2010
  42. Nicolas PitreSep 3, 2010
  43. Luke Kenneth Casson LeightonSep 3, 2010
  44. Nguyen Thai Ngoc DuySep 4, 2010
  45. Nguyen Thai Ngoc DuySep 4, 2010
  46. Artur SkawinaSep 4, 2010
  47. Nicolas PitreSep 4, 2010
  48. Artur SkawinaSep 4, 2010
  49. Nicolas PitreSep 4, 2010
  50. Luke Kenneth Casson LeightonSep 4, 2010
  51. Luke Kenneth Casson LeightonSep 4, 2010
  52. Nicolas PitreSep 5, 2010
  53. Luke Kenneth Casson LeightonSep 5, 2010
  54. Nicolas PitreSep 5, 2010
  55. Luke Kenneth Casson LeightonSep 6, 2010
  56. Nicolas PitreSep 6, 2010
  57. Luke Kenneth Casson LeightonSep 6, 2010
  58. Junio C HamanoSep 6, 2010
  59. Nicolas PitreSep 6, 2010
  60. Luke Kenneth Casson LeightonSep 7, 2010
  61. Luke Kenneth Casson LeightonSep 7, 2010
  62. Artur SkawinaSep 4, 2010
  63. Theodore TsoSep 4, 2010
  64. Kyle MoffettSep 4, 2010
  65. Theodore TsoSep 4, 2010
  66. Luke Kenneth Casson LeightonSep 4, 2010
  67. Nicolas PitreSep 5, 2010
  68. Luke Kenneth Casson LeightonSep 5, 2010
  69. Nicolas PitreSep 4, 2010
  70. Theodore TsoSep 4, 2010
  71. Luke Kenneth Casson LeightonSep 4, 2010
  72. Luke Kenneth Casson LeightonSep 4, 2010
  73. Ted Ts'oSep 4, 2010
  74. Luke Kenneth Casson LeightonSep 4, 2010
  75. Ted Ts'oSep 4, 2010
  76. Luke Kenneth Casson LeightonSep 5, 2010
  77. Jakub NarebskiSep 4, 2010
  78. Luke Kenneth Casson LeightonSep 4, 2010
  79. Jakub NarebskiSep 4, 2010
  80. Luke Kenneth Casson LeightonSep 4, 2010
  81. Ted Ts'oSep 4, 2010
  82. Tomas CarneckySep 5, 2010
  83. Nicolas PitreSep 5, 2010
  84. Luke Kenneth Casson LeightonSep 5, 2010
  85. Nicolas PitreSep 6, 2010
  86. Luke Kenneth Casson LeightonSep 4, 2010
  87. Artur SkawinaSep 4, 2010
  88. Artur SkawinaSep 4, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.