Re: Git transfer protocols (was: Re: Git-commits mailing list feed)
- From
Jan Harkes <jaharkes@cs.cmu.edu>
- Date
- Apr 23, 2005, 22:22 UTC
- Message-ID
- <20050423222226.GB16751@delft.aura.cs.cmu.edu>
- In-Reply-To
- <426ABE1B.7000905@timesys.com>
On Sat, Apr 23, 2005 at 02:28:59PM -0700, Mike Taht wrote:
Show 9 quoted lines
> Jan Harkes wrote: > > >rsync works fine for now, but people are already looking at implementing > >smarter (more efficient) ways to synchronize git repositories by > >grabbing missing commits, and from there fetching any missing tree and > >file blobs. However there is no such linkage to discover missing tag > >objects, only a full rsync would be able to get them and for that it has > >to send the name of every object in the repository to the other side to > >check for any missing ones.
Actually I just realized that I personally probably wouldn't care about most of the tags that people might add to their trees. Maybe once in a while, but the tag would probably be obtained through email or the web.
> I think that one reason why rsync is inefficient for git is that it
...
> lastly, Monotone has it's own "netsync" protocol > (via http://www.venge.net/monotone/faq.html)
Interesting, probably something like any of these might end up useful to replace rsync for mirroring full git repositories.
I'm actually more selfish than that and am thinking on how I expect to use git.
See, I don't care about most of the objects in the repository. In practice I would probably pull only the latest 'head' once in a while look for missing commits to give me a quick overview of what has changed. Then if a diff between the new head and my current tree shows that anything might have changed in an area I actually do care about, such as the VFS, I'd want something that does a quick binary search to identify the commits where the changes occured. But for that I only need to look at a limited number of tree objects.
As long as I know that someone, somewhere is archiving the whole repository I can always come back later and fill in the blanks.
HTTP/1.1 with persistent connections and some request interleaving is probably the fastest and most server friendly way to grab those objects I really care about.
Jan