git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git pack/unpack over bittorrent - works!

From
Nicolas Pitre <nico@fluxnic.net>
Date
Sep 6, 2010, 16:51 UTC
Message-ID
<alpine.LFD.2.00.1009061025210.19366@xanadu.home>
In-Reply-To
<AANLkTi=CEOj40Sj+zegvX+ry8-y6p7UwsyqdtoHB1d-T@mail.gmail.com>
On Mon, 6 Sep 2010, Luke Kenneth Casson Leighton wrote:
Show 12 quoted lines
> On Mon, Sep 6, 2010 at 12:52 AM, Nicolas Pitre <nico@fluxnic.net> wrote:
> 
> > And object enumeration has absolutely nothing to do with packs, nor .idx
> > files for that matter.
> 
>  mmm packs not being to do with object enumeration i get.  i
> understand that .idx files contain "lists of objects" which isn't the
> same thing (and also happen to contain pointers/offsets to the objects
> of its associated .pack)
> 
>  at some point i'd really like to know what the object list is (not
> the objects themselves) that comes out of "git pack-objects --thin"

You need to feed 'git pack-objects' a list of objects in the first place for it to pack anything. So you must have that list even before pack-objects can produce any output. And that list is usually generated by 'git rev-list'. So, typically, you'd do:

	git rev-list --objects <commit_range> | git pack-objects foo

But these days the ability to enumerate objects was integrated into pack-objects directly, so you can do:

	echo "<commit_range>" | git pack-objects --revs foo
But you should get the idea.
Show 7 quoted lines
> > So... I hope you understand now that there is no relation between
> > commits and .idx files.  The only exception is when you do create a
> > custom pack with 'git pack-objects'.
> 
>  yes.  ahh... that's what i've been doing: using "git pack-objects
> --thin".  and the reason for that is because i've seen it used in the
> http implementation of "git fetch".

Well, the HTTP implementation is a rather tricky example as there are actually two implementations: one that is dumb and only slurps packs and loose objects out of a remote .git/ directory, and another that is smart enough to carry the smarter Git protocol across HTTP requests.

When using the "smart" Git protocol, the client tells the server what it already has, and then the server uses pack-objects to produce a pack with only those objects that the client doesn't have, and stream that pack directly without even storing it on disk. The server doesn't even produce a .idx file in that case. It is up to the client to store the pack on disk, and feed it through 'git index-pack' to construct a .idx file for it locally.

>  so, my questions up until now regarding .pack and .idx have all been
> targetted at that, and based on that context, _not_ the packs+idx
> files that are in .git/
Tell me if the above clears them up.
Show 5 quoted lines
> > If you want all commits then you just need --all instead of HEAD.
> 
>  no, i want commits separated and individual and "compoundable".  the plan is:
> 
> * to get the ref associated with refs/heads/master
You can do:
	git rev-parse refs/heads/master
> * to get the list of all commits associated with that master ref
Just use (without the --objects argument):
	git rev-list refs/heads/master
> * to work out how far local deviates from remote along that list of commits

That's an operation that only the peer with the most recent commits can do, unless you transfer that huge list of commits from above across the network. So, on a server (i.e. the peer sending objects) you'd do:

	git rev-list <refs_that_I_publish> --not <refs_that_the_remote_has>
> * to get the objects which will make up the missing commits (if they
> aren't already in the local store)

Again, that's a task for the peer with objects to offer. It just has to use the above rev-list invocation and add the --objects argument to it (or feed the equivalent ref specifications to pack-objects directly as shown previously).

> * to apply those commits in the correct order

Why would you care about this? There is nothing to "apply" as all you have to do is simply transfer objects.

> in other words, the plan is to follow what git http fetch and/org git
> git:// fetch does as much as possible (ok, perhaps not).

Well... I don't think it would be easy to do the same in a P2P context. Those fetch operations are totally stream oriented between 2 peers, and not many to many.

> the reason for getting the objects individually (blobs etc.) should be
> clear: prior commits _could_ have resulted in that exact object having
> been obtained already.

Sure. But objects known to exist on the remote side won't be listed by rev-list.

> so far i have implemented:
> 
> * get the master ref using git for-each-ref
> * get the list of all commits using git rev-list
So far so good.
> * enumerate the list of objects associated with an individual commit by:
>     i) creating a CUSTOM pack+idx using git pack-objects {ref}
>     ii) *parsing* the idx file using gitdb's FileIndex to get the list
> of objects

That's where you're going so much out of your way to give you trouble. A simple rev-list would give you that list:

	git rev-list --objects <this_commit> --not <this_commit''s_parents>
That's it.
>     iii) transferring that list to the local machine
> * requesting *individual* objects from the enumerated list out of the idx file
>    by using a CUSTOM "git pack-objects --thin {ref} < {ref}" command

That's where you'll have to get your hands real dirty and write actual code to serve individual objects but not through cat-file. In a P2P setup you'd want to transfer as little amount of data as possible, meaning that you'd want to serve deltas as much as possible. It's then a matter of finding if the requested object exists already in delta form, if so then whether or not its base object is something that the other end has in which case you send that as is, otherwise figuring out if that would be worth creating a delta against another object known to exist at the other end.

On the receiving end, you'd simply have to store those objects in the .git/objects/ directories as loose objects after expanding the deltas, or even stuff everything into a pack and run 'git index-pack' on it when the transfer is complete (and run 'git repack' to optimize the pack eventually).

Once the transfer is complete, you do a reachability and validity check on the refs you are supposed to have received the objects for, and if everything is OK then you update the refs and you're done.

Show 9 quoted lines
> that's as far as i've got, before you mentioned that it would be
> better to use "git rev-list --objects commit1..commit2" and to use
> "git cat-file" to obtain the actual object [what's not clear in this
> plan is how to store that cat'ed file at the local end, hence the
> continued use of git pack-objects --thin {ref} < {ref}]
> 
> the prior implementation was to treat the custom pack-object as if it
> was "the atomic leaf-node operation" instead of individual objects
> (blobs, trees).

Well, OK. But suppose that you have only 2 new commits with a big amount of objects for each. Typically the very first commit of a project corresponds to the import of that project into Git, and it is equivalent to the whole work tree. Don't you want to spread the request for those objects across as many peers as possible?

>  what i _have_ been doing however is custom-generating pack-objects
> and associated pack-indexes (just like git http fetch) _including_
> using the --thin option because that's what git http fetch does.

Well, let's get back to that HTTP fetch which has a double personality. The "smart" HTTP fetch doesn't involve any pack index at all. It ends up streaming pack-objects stdout's output over the net and the pack index is recreated on the other end.

However, what the _dumb_ HTTP fetch does (and that is the same idea for the FTP fetch, or the rsync fetch) is to dig into the remote's .git directory and grab those .idx files, look into them to see if the corresponding .pack file actually contain the wanted objects, so to only downloads the needed packs afterwards. And those dumb protocols are what they are: dumb. They usually end up transferring way more data than actually necessary.

Show 6 quoted lines
> > If an object that does get transmitted is
> > actually a delta against an object that is only part of a branch that is
> > not published, then the delta will be expanded and redone against
> > another suitable object before transmission.
> 
>  and that's handled by git pack-objects --thin (am i right?)
Right.  But the --thin flag here is unrelated to this.
What --thin does is to tell pack-objects that it can produce deltas 
against objects that will _not_ be included in the produced pack.  That 
is OK only if the consumer of that pack is 1) aware of that fact and
2) is going to "fix" the pack by appending those objects to the pack 
from a local copy.  For a pack to be "valid" in your .git directory, it 
has to be self contained with regards to deltas.  It is not allowed to 
have deltas across different packs as this makes the issue of delta 
loops extremely difficult to deal with in the context of incremental 
repacks.
Show 9 quoted lines
>  so.  we have a hierarchical plan: get the commit list, get a
> per-commit object-list, get the objects (if needed), store the
> objects.
> 
>  problem: despite looking through virtually every single builtin/*.c
> file which uses write_sha1_file (which i believe i have correctly
> identified, from examining git unpack-objects, as being the function
> which stores actual objects, including their type), i do not see a git
> command (yet) which performs the reverse operation of "git cat-file".
It is 'git hash-object'.
Nicolas
Previous: Luke Kenneth Casson LeightonNext: Luke Kenneth Casson Leighton
Message 56 of 88 in “git pack/unpack over bittorrent - works!”
  1. Luke Kenneth Casson LeightonSep 1, 2010
  2. Nguyen Thai Ngoc DuySep 1, 2010
  3. Luke Kenneth Casson LeightonSep 2, 2010
  4. Luke Kenneth Casson LeightonSep 2, 2010
  5. Ævar Arnfjörð BjarmasonSep 2, 2010
  6. A Large Angry SCMSep 2, 2010
  7. Luke Kenneth Casson LeightonSep 2, 2010
  8. Luke Kenneth Casson LeightonSep 2, 2010
  9. A Large Angry SCMSep 2, 2010
  10. Jeff KingSep 2, 2010
  11. Nicolas PitreSep 2, 2010
  12. A Large Angry SCMSep 2, 2010
  13. Nicolas PitreSep 2, 2010
  14. Luke Kenneth Casson LeightonSep 2, 2010
  15. Shawn O. PearceSep 2, 2010
  16. Luke Kenneth Casson LeightonSep 2, 2010
  17. Luke Kenneth Casson LeightonSep 2, 2010
  18. Nicolas PitreSep 3, 2010
  19. Luke Kenneth Casson LeightonSep 3, 2010
  20. Junio C HamanoSep 3, 2010
  21. Brandon CaseySep 2, 2010
  22. Luke Kenneth Casson LeightonSep 2, 2010
  23. Jakub NarebskiSep 2, 2010
  24. Luke Kenneth Casson LeightonSep 2, 2010
  25. Luke Kenneth Casson LeightonSep 2, 2010
  26. Nicolas PitreSep 3, 2010
  27. Nguyen Thai Ngoc DuySep 3, 2010
  28. Luke Kenneth Casson LeightonSep 3, 2010
  29. Luke Kenneth Casson LeightonSep 3, 2010
  30. Luke Kenneth Casson LeightonSep 3, 2010
  31. Luke Kenneth Casson LeightonSep 2, 2010
  32. Casey DahlinSep 2, 2010
  33. A Large Angry SCMSep 2, 2010
  34. Nicolas PitreSep 2, 2010
  35. Luke Kenneth Casson LeightonSep 2, 2010
  36. A Large Angry SCMSep 2, 2010
  37. Nicolas PitreSep 2, 2010
  38. Theodore TsoSep 3, 2010
  39. Luke Kenneth Casson LeightonSep 3, 2010
  40. Junio C HamanoSep 3, 2010
  41. Ted Ts'oSep 3, 2010
  42. Nicolas PitreSep 3, 2010
  43. Luke Kenneth Casson LeightonSep 3, 2010
  44. Nguyen Thai Ngoc DuySep 4, 2010
  45. Nguyen Thai Ngoc DuySep 4, 2010
  46. Artur SkawinaSep 4, 2010
  47. Nicolas PitreSep 4, 2010
  48. Artur SkawinaSep 4, 2010
  49. Nicolas PitreSep 4, 2010
  50. Luke Kenneth Casson LeightonSep 4, 2010
  51. Luke Kenneth Casson LeightonSep 4, 2010
  52. Nicolas PitreSep 5, 2010
  53. Luke Kenneth Casson LeightonSep 5, 2010
  54. Nicolas PitreSep 5, 2010
  55. Luke Kenneth Casson LeightonSep 6, 2010
  56. Nicolas PitreSep 6, 2010
  57. Luke Kenneth Casson LeightonSep 6, 2010
  58. Junio C HamanoSep 6, 2010
  59. Nicolas PitreSep 6, 2010
  60. Luke Kenneth Casson LeightonSep 7, 2010
  61. Luke Kenneth Casson LeightonSep 7, 2010
  62. Artur SkawinaSep 4, 2010
  63. Theodore TsoSep 4, 2010
  64. Kyle MoffettSep 4, 2010
  65. Theodore TsoSep 4, 2010
  66. Luke Kenneth Casson LeightonSep 4, 2010
  67. Nicolas PitreSep 5, 2010
  68. Luke Kenneth Casson LeightonSep 5, 2010
  69. Nicolas PitreSep 4, 2010
  70. Theodore TsoSep 4, 2010
  71. Luke Kenneth Casson LeightonSep 4, 2010
  72. Luke Kenneth Casson LeightonSep 4, 2010
  73. Ted Ts'oSep 4, 2010
  74. Luke Kenneth Casson LeightonSep 4, 2010
  75. Ted Ts'oSep 4, 2010
  76. Luke Kenneth Casson LeightonSep 5, 2010
  77. Jakub NarebskiSep 4, 2010
  78. Luke Kenneth Casson LeightonSep 4, 2010
  79. Jakub NarebskiSep 4, 2010
  80. Luke Kenneth Casson LeightonSep 4, 2010
  81. Ted Ts'oSep 4, 2010
  82. Tomas CarneckySep 5, 2010
  83. Nicolas PitreSep 5, 2010
  84. Luke Kenneth Casson LeightonSep 5, 2010
  85. Nicolas PitreSep 6, 2010
  86. Luke Kenneth Casson LeightonSep 4, 2010
  87. Artur SkawinaSep 4, 2010
  88. Artur SkawinaSep 4, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.