threads / discuss / 2669

Re: Linux 2.6.15-rc2

Subject: Re: Linux 2.6.15-rc2

## tl;dr

10 messages between Nov 24, 2005 and Nov 25, 2005.

replies: 9people: 5as markdown or json

Ed Tomlinson· Nov 24, 2005, 12:37 UTC · lore
On Saturday 19 November 2005 22:40, Linus Torvalds wrote:
> There it is (or will soon be - the tar-ball and patches are still 
> uploading, and mirroring can obviously take some time after that).

Something strange here. After a cg-update, I had no tag for rc2. Checking showed no problems so I used cg-clone to get another copy of the repository. Still no rc2.

ed@grover:/usr/src/2.6$ cg-version cogito-0.16rc2 (73874dddeec2d0a8e5cd343eec762d98314def63) ed@grover:/usr/src/2.6$ git --version git version 0.99.9.GIT

cg-clone http://www.kernel.org/pub/scm/linux/kernel/git/torvalds/linux-2.6.git 2.6
It looks to be the tag that is missing, gitk show commits after Nov 19.
Both git and cg were  updated just prior to the cg-update (~Nov 22 8pm EST).
What is happening?

TIA Ed Tomlinson

Andreas Ericsson· Nov 24, 2005, 13:07 UTC · re: Ed Tomlinson · lore
Ed Tomlinson wrote:
Show 11 quoted lines
> Something strange here.   After a cg-update, I had no tag for rc2.   Checking
> showed no problems so I used cg-clone to get another copy of the repository.
> Still no rc2.
> 
> ed@grover:/usr/src/2.6$ cg-version
> cogito-0.16rc2 (73874dddeec2d0a8e5cd343eec762d98314def63)
> ed@grover:/usr/src/2.6$ git --version
> git version 0.99.9.GIT
> 
> cg-clone http://www.kernel.org/pub/scm/linux/kernel/git/torvalds/linux-2.6.git 2.6
> 

This happened a while ago to someone else too. Apparently the http transport needs serverside help (git-update-server-info or some such must be run on the remote side).

Unless you're restricted by firewalls and other you could try

git clone git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux-2.6.git 2.6

which works flawlessly for me although it takes quite some time to transfer all the data.

Linus, HPA: Are the packs cached on kernel.org? It seems to be at least a minute before the transfers start.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Tel: +46 8-230225                  Fax: +46 8-230231
Linus Torvalds· Nov 24, 2005, 18:44 UTC · re: Andreas Ericsson · lore
On Thu, 24 Nov 2005, Andreas Ericsson wrote:
Show 5 quoted lines
> 
> git clone git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux-2.6.git 2.6
> 
> which works flawlessly for me although it takes quite some time to transfer
> all the data.

The initial clone is very expensive for the native git protocol: the protocol is designed to scale well for incremental updates (ie you have a _huge_ repository that has changed just a bit, and the protocol should work well for that), and that makes the initial clone quite expensive as it marshalls the whole damn repository into this nice packed format.

So it's often nicer (certainly on the remote server) to use "rsync" for the initial clone, and then only after that start using the git protocol.

(This is in no way really fundamental, and the server could cache the packs it generates for initial clones, but that isn't implemented yet, and probably won't be for some times).

Of course, especially if you're mostly bandwidth-constrained and the server side is not under a big load, using the native git protocol may actually be faster anyway. Because it's always going to generate the nicest packing, while rsync:// will just use whatever packing that the server happens to have at that point (but I do repack every few weeks, so rsync for the initial clone should never be horribly bad - and since I just repacked, it should get that "perfect" pack too).

		Linus
Junio C Hamano· Nov 24, 2005, 19:42 UTC · re: Linus Torvalds · lore
Linus Torvalds <torvalds@osdl.org> writes:
> (This is in no way really fundamental, and the server could cache the 
> packs it generates for initial clones, but that isn't implemented yet, and 
> probably won't be for some times).
Performance perceived by cloners is helped by
    $ mkdir -p .git/pack-cache
    $ git-rev-list --objects --all | git-pack-objects .git/pack-cache/pack

on the server side. This exact example of preparing by the repository maintainer is optimizing for a wrong case, and I do not think it is worth doing in practice, but this will give you the lower bound when server side cache is implemented to do it on demand.

Linus Torvalds· Nov 24, 2005, 19:57 UTC · re: Junio C Hamano · lore
On Thu, 24 Nov 2005, Junio C Hamano wrote:
Show 5 quoted lines
> 
> Performance perceived by cloners is helped by
> 
>     $ mkdir -p .git/pack-cache
>     $ git-rev-list --objects --all | git-pack-objects .git/pack-cache/pack

That really doesn't work very well. I push to that tree often several times a day, and you'd have to re-do the cache each time.

So it would be much better if git-pack-objects would just always cache its output in .git/pack-cache - along with some logic to just get rid of old ones regularly.

Since git-pack-objects has to generate the pack _anyway_, it might as well save it away when it does - so that if you have lots of people doing clones or pulling, you'd only need to run it once for a particular set of objects, and you'd not have to do any extra (or unnecessary) maintenance.

		Linus
Junio C Hamano· Nov 24, 2005, 21:02 UTC · re: Linus Torvalds · lore
Linus Torvalds <torvalds@osdl.org> writes:
> Since git-pack-objects has to generate the pack _anyway_, it might as well 
> save it away when it does - so that if you have lots of people doing 
> clones or pulling, you'd only need to run it once for a particular set of 
> objects, and you'd not have to do any extra (or unnecessary) maintenance.

Caching itself is relatively easy (just implement an equivalent of tee inside pack-objects ourselves). More problematic is pruning. We could do it from cron based on atime _if_ the filesystem is not mounted noatime but without arranging a reasonably way for automated pruning this would become a disk hog and extra maintenance burden, which is why I did not implement the dynamic caching part in the initial round.

Since git-daemon would be the primary user of pack-cache/, this implies a repository writable by git-daemon user on public machine (not master), which is an extra thing to note.

Linus Torvalds· Nov 24, 2005, 18:37 UTC · re: Ed Tomlinson · lore
On Thu, 24 Nov 2005, Ed Tomlinson wrote:
> 
> What is happening?

The http transport isn't very good for git, so git adds various special files to make it work at all. They need to be specially updated, and I hadn't done that.

Using the native git protocol through git://git.kernel.org/.. gets around it, as does using rsync.

I just repacked and updated it now, so how http should work too, although inefficiently (because it will get a whole new pack - just one of the disadvantages of the non-native protocols).

		Linus
Nick Hengeveld· Nov 24, 2005, 19:52 UTC · re: Linus Torvalds · lore
On Thu, Nov 24, 2005 at 10:37:15AM -0800, Linus Torvalds wrote:
> I just repacked and updated it now, so how http should work too, although 
> inefficiently (because it will get a whole new pack - just one of the 
> disadvantages of the non-native protocols).

There's room to improve on that particular inefficiency. The http commit walker could use Range: headers to fetch loose objects directly from inside a pack if it didn't make sense to fetch the entire pack. For this to work, pack fetches would need to be deferred until the entire tree had been walked, and the commit walker could decide whether to fetch the pack or loose objects based on the percentage of packed objects it needed to fetch. It would also need to fetch all tag/commit/tree objects using ranges to be able to fully walk the tree.

-- 
For a successful technology, reality must take precedence over public
relations, for nature cannot be fooled.
Ed Tomlinson· Nov 25, 2005, 02:50 UTC · re: Nick Hengeveld · lore
On Thursday 24 November 2005 14:52, Nick Hengeveld wrote:
Show 14 quoted lines
> On Thu, Nov 24, 2005 at 10:37:15AM -0800, Linus Torvalds wrote:
> 
> > I just repacked and updated it now, so how http should work too, although 
> > inefficiently (because it will get a whole new pack - just one of the 
> > disadvantages of the non-native protocols).
> 
> There's room to improve on that particular inefficiency.  The http
> commit walker could use Range: headers to fetch loose objects directly
> from inside a pack if it didn't make sense to fetch the entire pack.
> For this to work, pack fetches would need to be deferred until the
> entire tree had been walked, and the commit walker could decide whether
> to fetch the pack or loose objects based on the percentage of packed
> objects it needed to fetch.  It would also need to fetch all
> tag/commit/tree objects using ranges to be able to fully walk the tree.

Alternately, when creating a new archive the client could ask the server what protocols are active. It could then use the best one for the clone and update the .git/origin files with the optimal one for incremental pulls.

Thoughts? Ed Tomlinson

Andreas Ericsson· Nov 25, 2005, 08:42 UTC · re: Ed Tomlinson · lore
Ed Tomlinson wrote:
Show 23 quoted lines
> On Thursday 24 November 2005 14:52, Nick Hengeveld wrote:
> 
>>On Thu, Nov 24, 2005 at 10:37:15AM -0800, Linus Torvalds wrote:
>>
>>
>>>I just repacked and updated it now, so how http should work too, although 
>>>inefficiently (because it will get a whole new pack - just one of the 
>>>disadvantages of the non-native protocols).
>>
>>There's room to improve on that particular inefficiency.  The http
>>commit walker could use Range: headers to fetch loose objects directly
>>from inside a pack if it didn't make sense to fetch the entire pack.
>>For this to work, pack fetches would need to be deferred until the
>>entire tree had been walked, and the commit walker could decide whether
>>to fetch the pack or loose objects based on the percentage of packed
>>objects it needed to fetch.  It would also need to fetch all
>>tag/commit/tree objects using ranges to be able to fully walk the tree.
> 
> 
> Alternately, when creating a new archive the client could ask the server
> what protocols are active.  It could then use the best one for the clone and
> update the .git/origin files with the optimal one for incremental pulls.
> 

This would only work with the git protocol, and since that's the fastest protocol (theoretically that is, Pasky seems to have gotten other figures but I'm not sure I believe those) it should really only ever return itself which wouldn't make much sense.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Tel: +46 8-230225                  Fax: +46 8-230231

← back to recent threads