git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git pack/unpack over bittorrent - works!

From
KMKyle Moffett <kyle@moffetthome.net>
Date
Sep 4, 2010, 05:23 UTC
Message-ID
<AANLkTikVf=X8cLP9s6W9VGOt0EHE4J5MYsBpgKYhrAri@mail.gmail.com>
In-Reply-To
<04755B03-EE1D-48FA-8894-33AA8E2661C0@mit.edu>
Ted,

I think your "canonical pack" idea has value, but I'd be inclined to try to optimize more for the "common case" of developing on a fast local network with many local checkouts, where you occasionally push/fetch external sources via a slow link.

Specifically, let's look at the very reasonable scenario of a developer working over a slow DSL or dialup connection. He's probably got many copies of various GIT repositories cloned all over the place (hey, disk is cheap!), but right now he just wants a fresh clean copy of somebody else's new tree with whatever its 3 feature branches are. Furthermore, he's probably even got 80% of the commit objects from that tree archived in his last clone from linux-next.

In theory he could very carefully arrange his repositories with judicious use of alternate object directories. From personal experience, though, such arrangements are *VERY* prone to accidentally purging wanted objects; unless you *never* ever delete a branch in the "reference" repository.

So I think the real problem to solve would be: Given a collection of local computers each with many local repositories, what is the best way to optimize a clone of a "new" remote repository (over a slow link) by copying most of the data from other local repositories accessible via a fast link?

The goal would be to design a P2P protocol capable of rapidly and efficiently building distributed searchable indexes of ordered commits that identify which peer(s) contain that each commit.

When you attempt to perform a "git fetch --peer" from a repository, it would quickly connect to a few of the metadata index nodes in the P2P network and use them to negotiate "have"s with the upstream server. The client would then sequentially perform the local "fetch" operations necessary to obtain all the objects it used to minimize the commit range with the server. Once all of those "fetch" operations completed, it could proceed to fetch objects from the server normally.

Some amount of design and benchmarking would need to be done in order to figure out the most efficient indexing algorithm for finding a minimal set of "have"s of potentially thousands of refs, many with independent root commits. For example if the index was grouped according to "root commit" (of which there may be more than one), you *should* be able to quickly ask the server about a small list of root commits and then only continue asking about commits whose roots are all known to the server.

The actual P2P software would probably involve 2 different daemon processes. The first would communicate with each other and with the repositories, maintaining the ref and commit indexes. These daemons would advertise themselves with Avahi, or alternatively in an enterprise environment they would be managed by your sysadmins and be automatically discovered using DNS-SD. Clients looking to perform a P2P fetch would first ask these.

The second daemon would be a modified git-daemon that connects to the advertised "index" daemons and advertises its own refs and commit lists, as well as its IP address and port.

My apologies if there are any blatant typos or thinkos, it's a bit later here than I would normally be writing about technical topics.

Cheers, Kyle Moffett

Previous: Theodore TsoNext: Theodore Tso
Message 64 of 88 in “git pack/unpack over bittorrent - works!”
  1. Luke Kenneth Casson LeightonSep 1, 2010
  2. Nguyen Thai Ngoc DuySep 1, 2010
  3. Luke Kenneth Casson LeightonSep 2, 2010
  4. Luke Kenneth Casson LeightonSep 2, 2010
  5. Ævar Arnfjörð BjarmasonSep 2, 2010
  6. A Large Angry SCMSep 2, 2010
  7. Luke Kenneth Casson LeightonSep 2, 2010
  8. Luke Kenneth Casson LeightonSep 2, 2010
  9. A Large Angry SCMSep 2, 2010
  10. Jeff KingSep 2, 2010
  11. Nicolas PitreSep 2, 2010
  12. A Large Angry SCMSep 2, 2010
  13. Nicolas PitreSep 2, 2010
  14. Luke Kenneth Casson LeightonSep 2, 2010
  15. Shawn O. PearceSep 2, 2010
  16. Luke Kenneth Casson LeightonSep 2, 2010
  17. Luke Kenneth Casson LeightonSep 2, 2010
  18. Nicolas PitreSep 3, 2010
  19. Luke Kenneth Casson LeightonSep 3, 2010
  20. Junio C HamanoSep 3, 2010
  21. Brandon CaseySep 2, 2010
  22. Luke Kenneth Casson LeightonSep 2, 2010
  23. Jakub NarebskiSep 2, 2010
  24. Luke Kenneth Casson LeightonSep 2, 2010
  25. Luke Kenneth Casson LeightonSep 2, 2010
  26. Nicolas PitreSep 3, 2010
  27. Nguyen Thai Ngoc DuySep 3, 2010
  28. Luke Kenneth Casson LeightonSep 3, 2010
  29. Luke Kenneth Casson LeightonSep 3, 2010
  30. Luke Kenneth Casson LeightonSep 3, 2010
  31. Luke Kenneth Casson LeightonSep 2, 2010
  32. Casey DahlinSep 2, 2010
  33. A Large Angry SCMSep 2, 2010
  34. Nicolas PitreSep 2, 2010
  35. Luke Kenneth Casson LeightonSep 2, 2010
  36. A Large Angry SCMSep 2, 2010
  37. Nicolas PitreSep 2, 2010
  38. Theodore TsoSep 3, 2010
  39. Luke Kenneth Casson LeightonSep 3, 2010
  40. Junio C HamanoSep 3, 2010
  41. Ted Ts'oSep 3, 2010
  42. Nicolas PitreSep 3, 2010
  43. Luke Kenneth Casson LeightonSep 3, 2010
  44. Nguyen Thai Ngoc DuySep 4, 2010
  45. Nguyen Thai Ngoc DuySep 4, 2010
  46. Artur SkawinaSep 4, 2010
  47. Nicolas PitreSep 4, 2010
  48. Artur SkawinaSep 4, 2010
  49. Nicolas PitreSep 4, 2010
  50. Luke Kenneth Casson LeightonSep 4, 2010
  51. Luke Kenneth Casson LeightonSep 4, 2010
  52. Nicolas PitreSep 5, 2010
  53. Luke Kenneth Casson LeightonSep 5, 2010
  54. Nicolas PitreSep 5, 2010
  55. Luke Kenneth Casson LeightonSep 6, 2010
  56. Nicolas PitreSep 6, 2010
  57. Luke Kenneth Casson LeightonSep 6, 2010
  58. Junio C HamanoSep 6, 2010
  59. Nicolas PitreSep 6, 2010
  60. Luke Kenneth Casson LeightonSep 7, 2010
  61. Luke Kenneth Casson LeightonSep 7, 2010
  62. Artur SkawinaSep 4, 2010
  63. Theodore TsoSep 4, 2010
  64. Kyle MoffettSep 4, 2010
  65. Theodore TsoSep 4, 2010
  66. Luke Kenneth Casson LeightonSep 4, 2010
  67. Nicolas PitreSep 5, 2010
  68. Luke Kenneth Casson LeightonSep 5, 2010
  69. Nicolas PitreSep 4, 2010
  70. Theodore TsoSep 4, 2010
  71. Luke Kenneth Casson LeightonSep 4, 2010
  72. Luke Kenneth Casson LeightonSep 4, 2010
  73. Ted Ts'oSep 4, 2010
  74. Luke Kenneth Casson LeightonSep 4, 2010
  75. Ted Ts'oSep 4, 2010
  76. Luke Kenneth Casson LeightonSep 5, 2010
  77. Jakub NarebskiSep 4, 2010
  78. Luke Kenneth Casson LeightonSep 4, 2010
  79. Jakub NarebskiSep 4, 2010
  80. Luke Kenneth Casson LeightonSep 4, 2010
  81. Ted Ts'oSep 4, 2010
  82. Tomas CarneckySep 5, 2010
  83. Nicolas PitreSep 5, 2010
  84. Luke Kenneth Casson LeightonSep 5, 2010
  85. Nicolas PitreSep 6, 2010
  86. Luke Kenneth Casson LeightonSep 4, 2010
  87. Artur SkawinaSep 4, 2010
  88. Artur SkawinaSep 4, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.