git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git-daemon on NSLU2

From
Jon Smirl <jonsmirl@gmail.com>
Date
Aug 24, 2007, 23:46 UTC
Message-ID
<9e4733910708241646x7b285574t94c3d7eb32bb60c9@mail.gmail.com>
In-Reply-To
<fanmmk$f5q$1@sea.gmane.org>
On 8/24/07, Jakub Narebski <jnareb@gmail.com> wrote:
Show 5 quoted lines
> There was idea to special case clone (just concatenate the packs, the
> receiving side as someone told there can detect pack boundaries; do not
> forget to pack loose objects, first), instead of using generic fetch --all
> for clone, bnut no code. Code speaks louder than words (although if someone
> would provide details of pack boundary detection...)

A related concept, initial clone of a repository does the equivalent of repack -a on the repo before transmitting it. Why aren't we saving those results by switching the repo onto the new pack file? Then the next clone that comes along won't have to do anything but send the file.

But this logic can be flipped around, if the remote needs any object from the pack file, just send them the whole pack file and let the remote sort it out. Using this logic you can still minimize the IO statistically.

When a remote does a fetch you have to pack all of the loose objects. When the loose object pile reaches 20MB or so, the fetch can trigger a repack of the oldest half into a pack that is kept by the tree and replaces those older loose objects. For future fetches simply apply the rule of sending the whole pack if any object is needed.

The repack of the 10MB of older objects can be kicked out to another process and copied into the tree when it is finished. At that point the loose objects can be deleted. The git db can tolerate a process copying in a new packfile and deleting the old objects while other processes may be using the database, right?

This model shouldn't statistically change the amount of data very much. If you haven't synced your tree in a month a few too many objects may get sent to you. However, it should dramatically reduce the IO load on the server cause by git protocol initial clones.

-- 
Jon Smirl
jonsmirl@gmail.com
Previous: Jakub NarebskiNext: Junio C Hamano
Message 11 of 30 in “git-daemon on NSLU2”
  1. Jon SmirlAug 24, 2007
  2. Shawn O. PearceAug 24, 2007
  3. Jon SmirlAug 24, 2007
  4. Nicolas PitreAug 24, 2007
  5. Jon SmirlAug 24, 2007
  6. Nicolas PitreAug 24, 2007
  7. Jon SmirlAug 24, 2007
  8. Jakub NarebskiAug 24, 2007
  9. Junio C HamanoAug 24, 2007
  10. Jakub NarebskiAug 24, 2007
  11. Jon SmirlAug 24, 2007
  12. Junio C HamanoAug 25, 2007
  13. David KastrupAug 25, 2007
  14. Salikh ZakirovAug 25, 2007
  15. Nicolas PitreAug 25, 2007
  16. Linus TorvaldsAug 24, 2007
  17. Jon SmirlAug 25, 2007
  18. Jeff KingAug 26, 2007
  19. Jon SmirlAug 26, 2007
  20. Linus TorvaldsAug 26, 2007
  21. Jon SmirlAug 26, 2007
  22. Linus TorvaldsAug 26, 2007
  23. Jon SmirlAug 26, 2007
  24. Linus TorvaldsAug 26, 2007
  25. Junio C HamanoAug 26, 2007
  26. Theodore TsoAug 27, 2007
  27. Linus TorvaldsAug 27, 2007
  28. Daniel HulmeAug 26, 2007
  29. Jakub NarebskiAug 27, 2007
  30. Jon SmirlAug 24, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.