git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] rev-list: add "--full-objects" flag.

From
Eric W. Biederman <ebiederm@xmission.com>
Date
Jul 9, 2005, 21:09 UTC
Message-ID
<m1pstrr8k1.fsf@ebiederm.dsl.xmission.com>
In-Reply-To
<Pine.LNX.4.58.0507071928220.25104@g5.osdl.org>
Linus Torvalds <torvalds@osdl.org> writes:
Show 9 quoted lines
> On Thu, 7 Jul 2005, Junio C Hamano wrote:
>> 
>> However it does not automatically mean that the avenue I have
>> been pursuing would not work; the server side preparation needs
>> to be a bit more careful than what I sent, which unconditionally
>> runs "prune-packed".  It instead should leave the files that
>> "--whole-trees" would have packed as plain SHA1 files, so that
>> the bulk is obtained by statically generated packs and the rest
>> can be handled in the commit-chain walker as before.
> The "fetch one object, parse it, fetch the next one, parse that.." 
> approach is just horrible.

Agreed. That does not cover up latency at all and depending on the parsing cost can potentially even keep you from having anything on your network connection for a noticeable amount of time.

Show 5 quoted lines
> I ended up preferring the "rsync" thing even though rsync sucked badly on
> big object stores too, if only because when rsync got working, it at least
> nicely pipelined the transfers, and would transfer things ten times faster
> than git-ssh-pull did (maybe I'm exaggerating, but I don't think so, it
> really felt that way).

This feels to me like an implementation issue (no pipelining) rather than a design issue (pipelining is impossible).

Show 6 quoted lines
> And the thing is, if you purely follow one tree (which is likely the
> common case for a lot of users), then you are actually always likely
> better off with the "mirror it" model. Which is _not_ a good model for
> developers (for example, me rsync'ing from Jeff's kernel repository always
> got me hundreds of useless objects), but it's fine for somebody who
> actually just wants to track somebody else.

I assume the problem with the mirror it model was simply there were to many objects?

> And then you really can use just rsync or wget or ncftpget or anything
> else that has a "fetch recursively, optimizing existing objects" mode.

Sane. But with an intelligent fetcher and a little extra information a dumb server should still be able to not fetch branches we care nothing about. I think that extra information is simply commit object graph and which packs those commit objects are in. I assume the commit graph information will be fairly modest.

Once you have that extra information you can generate incremental packs whenever you upload to the server, and you can make the incremental packs per branch.

That should allow an dumb fetcher to look at the list of commits and just fetch those packs it cares about, and since it only has to look one place first it should be fairly sane.

The core idea is that if the dumb-server-preparation can anticipate common access patterns (mirror a branch) and give enough information so that can be done cheaply and pipelined I don't expect it to be much worse than an intelligent fetcher.

The current intelligent fetch currently has a problem that it cannot be used to bootstrap a repository. If you don't have an ancestor of what you are fetching you can't fetch it.

Eric
Previous: Linus TorvaldsNext: Linus Torvalds
Message 25 of 66 in “[ANNOUNCE] Cogito-0.12”
  1. Petr BaudisJul 3, 2005
  2. Brian GerstJul 6, 2005
  3. Petr BaudisJul 7, 2005
  4. Junio C HamanoJul 7, 2005
  5. Linus TorvaldsJul 7, 2005
  6. Junio C HamanoJul 7, 2005
  7. Linus TorvaldsJul 7, 2005
  8. Junio C HamanoJul 7, 2005
  9. Junio C HamanoJul 7, 2005
  10. Eric W. BiedermanJul 7, 2005
  11. Linus TorvaldsJul 7, 2005
  12. Eric W. BiedermanJul 8, 2005
  13. Dumb servers (was: [ANNOUNCE] Cogito-0.12)Kevin Smith, Jul 8, 2005
  14. Linus TorvaldsJul 8, 2005
  15. Petr BaudisJul 7, 2005
  16. Linus TorvaldsJul 7, 2005
  17. Pull efficiently from a dumb git store.Junio C Hamano, Jul 7, 2005
  18. rev-list: add "--objects=self-sufficient" flag.Junio C Hamano, Jul 7, 2005
  19. Linus TorvaldsJul 7, 2005
  20. rev-list: add "--full-objects" flag.Junio C Hamano, Jul 8, 2005
  21. Linus TorvaldsJul 8, 2005
  22. Linus TorvaldsJul 8, 2005
  23. Junio C HamanoJul 8, 2005
  24. Linus TorvaldsJul 8, 2005
  25. Eric W. BiedermanJul 9, 2005
  26. Linus TorvaldsJul 10, 2005
  27. Junio C HamanoJul 10, 2005
  28. Sven VerdoolaegeJul 10, 2005
  29. Linus TorvaldsJul 10, 2005
  30. Eric W. BiedermanJul 11, 2005
  31. Linus TorvaldsJul 11, 2005
  32. Eric W. BiedermanJul 12, 2005
  33. Linus TorvaldsJul 12, 2005
  34. Eric W. BiedermanJul 12, 2005
  35. Linus TorvaldsJul 12, 2005
  36. Eric W. BiedermanJul 12, 2005
  37. Linus TorvaldsJul 12, 2005
  38. Linus TorvaldsJul 11, 2005
  39. Give --full-objects flag to rev-list when preparing a dumb server.Junio C Hamano, Jul 8, 2005
  40. Use --objects=self-sufficient flag to rev-list.Junio C Hamano, Jul 7, 2005
  41. Tony LuckJul 7, 2005
  42. Junio C HamanoJul 7, 2005
  43. Linus TorvaldsJul 7, 2005
  44. Tony LuckJul 8, 2005
  45. Linus TorvaldsJul 8, 2005
  46. Russell KingJul 9, 2005
  47. Russell KingJul 9, 2005
  48. Junio C HamanoJul 9, 2005
  49. Linus TorvaldsJul 10, 2005
  50. Linus TorvaldsJul 10, 2005
  51. Russell KingJul 10, 2005
  52. Junio C HamanoJul 10, 2005
  53. Russell KingJul 10, 2005
  54. Linus TorvaldsJul 10, 2005
  55. Russell KingJul 10, 2005
  56. Linus TorvaldsJul 10, 2005
  57. Russell KingJul 10, 2005
  58. Linus TorvaldsJul 10, 2005
  59. Russell KingJul 10, 2005
  60. Petr BaudisJul 10, 2005
  61. Chris WrightJul 11, 2005
  62. Linus TorvaldsJul 8, 2005
  63. Petr BaudisJul 8, 2005
  64. Daniel BarkalowJul 8, 2005
  65. Chris WrightJul 7, 2005
  66. Check packs and then files.Junio C Hamano, Jul 11, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.