threads / discuss / 9666

repo.or.cz wishes?

Subject: repo.or.cz wishes?

## tl;dr

28 messages between Aug 26, 2007 and Sep 1, 2007.

replies: 27people: 13as markdown or json

Petr Baudis· Aug 26, 2007, 23:59 UTC · lore
  Hi,
  I've just finally killed the HTTP auth for project administration that
was destroying everyone's lives, and added support for resetting
forgotten passwords, two main things that seemed to be the popular nits
of the repo.or.cz audience.
  So now I wonder, what is the thing you miss most there? Any cool stuff
repo.or.cz could (preferrably easily) do and doesn't?
  And please don't ask for smaller roundtrip times of requests for
administrator assistance. ;-)
  Thanks,
-- 
				Petr "Pasky" Baudis
Ever try. Ever fail. No matter. // Try again. Fail again. Fail better.
		-- Samuel Beckett
Sven Verdoolaege· Aug 27, 2007, 00:16 UTC · re: Petr Baudis · lore

Re: repo.or.cz wishes?

On Mon, Aug 27, 2007 at 01:59:44AM +0200, Petr Baudis wrote:
>   So now I wonder, what is the thing you miss most there? Any cool stuff
> repo.or.cz could (preferrably easily) do and doesn't?

Just a minor nit, but how about dropping the "git+" from the Push URL?

Jakub was also talking about support in gitweb for specifying the location of submodules. It would be nice if admins could set this information, wherever it ends up getting stored.

skimo
Petr Baudis· Aug 27, 2007, 00:41 UTC · re: Sven Verdoolaege · lore

Re: repo.or.cz wishes?

On Mon, Aug 27, 2007 at 02:16:34AM CEST, Sven Verdoolaege wrote:
Show 6 quoted lines
> On Mon, Aug 27, 2007 at 01:59:44AM +0200, Petr Baudis wrote:
> >   So now I wonder, what is the thing you miss most there? Any cool stuff
> > repo.or.cz could (preferrably easily) do and doesn't?
> 
> Just a minor nit, but how about dropping the "git+" from the
> Push URL?

I'm a major proponent of the "git+" - it's just the correct thing to specify. ssh:// by itself means secure _shell_, and that's not what the URL means - ssh is literaily just a transport layer for the git protocol. This is not my invention but fairly standard thing which plenty of people use, and it makes it possible to select proper protocol handlers and so on, shall something generic crunch on the URL. I've never actually understood why do some people dislike it.

> Jakub was also talking about support in gitweb for specifying
> the location of submodules.  It would be nice if admins could
> set this information, wherever it ends up getting stored.

Hmm, this shouldn't be very hard to do if the support will get into gitweb. And adding the support to gitweb shouldn't be that hard either. :-) OTOH, it's not something that would get me terribly excited, so I guess I'll wait for the gitweb side.

-- 
				Petr "Pasky" Baudis
Early to rise and early to bed makes a male healthy and wealthy and dead.
                -- James Thurber
Linus Torvalds· Aug 27, 2007, 18:23 UTC · re: Petr Baudis · lore

Re: repo.or.cz wishes?

On Mon, 27 Aug 2007, Petr Baudis wrote:
Show 9 quoted lines
> On Mon, Aug 27, 2007 at 02:16:34AM CEST, Sven Verdoolaege wrote:
> > On Mon, Aug 27, 2007 at 01:59:44AM +0200, Petr Baudis wrote:
> > >   So now I wonder, what is the thing you miss most there? Any cool stuff
> > > repo.or.cz could (preferrably easily) do and doesn't?
> > 
> > Just a minor nit, but how about dropping the "git+" from the
> > Push URL?
> 
> I'm a major proponent of the "git+"
I'd say "only", not "major".
It makes no sense.
> - it's just the correct thing to specify. ssh:// by itself means secure 
> _shell_
No it isn't, and no it doesn't.
It makes no sense what-so-ever.

"ssh://" is the *protocol*. What is actually done over the protocol is specified by the program.

This is not at all git specific. Try running "ssh" vs "scp" some day, and you'll notice the exact same thing: they both use the ssh _protocol_, but no, your statement that "ssh://" by itself means "secure _shell_" is total and utter garbage.

It means nothing at all of the kind. 

"ssh://" means the ssh protocol. It is that unambiguous, and that simple. Saying "git+ssh://" is totally idiotic, always has been, and always will be.

It's as stupid as it would be to require people to say
	scp cp+ssh://host/filename .

and nobody sane would *ever* advocate something that stupid. It's not how it's done.

So why do you continue to advocate "git+ssh://", when nobody else does, and several people have asked you not to.

And yes, I realize that SVN does it. SVN for some unfathomable reason uses "svn+ssh://", but let's face it, the SVN developers have neither taste nor brains. They don't know any better.

			Linus
Junio C Hamano· Aug 27, 2007, 18:58 UTC · re: Linus Torvalds · lore

Re: repo.or.cz wishes?

Linus Torvalds <torvalds@linux-foundation.org> writes:
Show 11 quoted lines
> On Mon, 27 Aug 2007, Petr Baudis wrote:
> ...
> It's as stupid as it would be to require people to say
>
> 	scp cp+ssh://host/filename .
>
> and nobody sane would *ever* advocate something that stupid. It's not how 
> it's done.
>
> So why do you continue to advocate "git+ssh://", when nobody else does, 
> and several people have asked you not to.
Chuckles...

I'd rather see hostname:/path/to/file on the page as it tends to be even shorter.

Matthieu Moy· Aug 27, 2007, 19:09 UTC · re: Linus Torvalds · lore

Re: repo.or.cz wishes?

Linus Torvalds <torvalds@linux-foundation.org> writes:
Show 6 quoted lines
> It's as stupid as it would be to require people to say
>
> 	scp cp+ssh://host/filename .
>
> and nobody sane would *ever* advocate something that stupid. It's not how 
> it's done.
You could find a better example.

scp doesn't accept URL syntax at all. So, no, it doesn't know about cp+ssh://, but doesn't know ssh:// either.

-- 
Matthieu
Martin Mares· Aug 27, 2007, 20:05 UTC · re: Linus Torvalds · lore

Re: repo.or.cz wishes?

Hello, world!\n
> "ssh://" is the *protocol*. What is actually done over the protocol is 
> specified by the program.
> 
> This is not at all git specific.
Really?
What does `ssh://what.the.hell.org/some/file' per se mean?

SSH is a protocol, but rather in the sense similar to TLS, not to HTTP. If it has some addressable objects, which could be referred to by the path part of the URL, they should be the programs to execute at the remote server, i.e., in our case the path to the GIT client binary, and certainly not the name of the repository, which has nothing to do with the SSH protocol.

(Just for completeness: I do not advocate using git+ssh, but your arguments against it look somewhat illogical.)

				Have a nice fortnight
-- 
Martin `MJ' Mares                          <mj@ucw.cz>   http://mj.ucw.cz/
Faculty of Math and Physics, Charles University, Prague, Czech Rep., Earth
Only dead fish swim with the stream.
Jing Xue· Aug 27, 2007, 21:27 UTC · re: Martin Mares · lore

Re: repo.or.cz wishes?

Quoting Martin Mares <mj@ucw.cz>:
Show 8 quoted lines
> What does `ssh://what.the.hell.org/some/file' per se mean?
>
> SSH is a protocol, but rather in the sense similar to TLS, not to HTTP.
> If it has some addressable objects, which could be referred to by the
> path part of the URL, they should be the programs to execute at the
> remote server, i.e., in our case the path to the GIT client binary,
> and certainly not the name of the repository, which has nothing to do
> with the SSH protocol.

Not to advocate either way (me being completely new to git), but as far as ssh is concerned, I don't think that the addressable objects necessarily have to be executables.

Quoting RFC4251: "The Secure Shell (SSH) Protocol is a protocol for secure remote login and other secure network services over an insecure network."

That reads rather vague to me.
-- 
Jing Xue
Linus Torvalds· Aug 27, 2007, 22:27 UTC · re: Martin Mares · lore

Re: repo.or.cz wishes?

On Mon, 27 Aug 2007, Martin Mares wrote:
> 
> What does `ssh://what.the.hell.org/some/file' per se mean?
So what does "http://what.the.hell.org/some/file" mean?
Does it mean that you have to start a web browser? Should we make that be
	git+http://what.the.hell.org/some/file
to make it clear that we're doing "git work" over the "http" protocol?
Pretty obviously not.
> SSH is a protocol, but rather in the sense similar to TLS, not to HTTP.

What does *that* mean? A protocol is a protocol. Your argument that protocols are "different" is pointless. Some protocols are usable for git, others aren't. OF COURSE different protocols are different. They are different in different ways.

Git uses URL's to say how to access something, which includes a protocol, an optional host, and a location within the host. It's quite obvious what they mean, and it's *also* obvious that the meaning is git-specific.

Here's what it boils down to:
 - do you think it is sensible to write
	git clone git+file:///some/directory
	git clone git+http://host/directory
	git clone git+rsync://host/directory
   when cloning from the local filesystem, over http, or over rsync 
   respectively? The first one, btw, actually uses the "git protocol". The 
   two others do not, but since a user shouldn't care, it would be really 
   stupid to try to make some internal implementation detail show up in 
   the URL scheme.
 - if you really think that the above is sensible, then explain why.
 - if you think that is TOTALLY IDIOTIC, then explain why "ssh://" is so 
   magically special that it would somehow make sense to say "git+" for 
   it?

As to your TLS example: if we were to do "git over TLS", it would make perfect sense to use either "tls://" (although "gits://" might be more natural, not because tls is wrong, but because people have gotten used to "https://") if we were to have a "secure git" port. Or maybe we'd use the same port number that we already have assigned for git, and just add some "use TLS to authenticate/encrypt", and use "tls://" for that. It makes perfect sense.

In short: you should just ask yourself: what is the most natural thing for a *user* to type to "git clone". And no, the "git+" prefix never makes sense.

			Linus
Sam Vilain· Aug 27, 2007, 22:58 UTC · re: Linus Torvalds · lore

Re: repo.or.cz wishes?

Linus Torvalds wrote:
Show 5 quoted lines
>  - if you really think that the above is sensible, then explain why.
> 
>  - if you think that is TOTALLY IDIOTIC, then explain why "ssh://" is so 
>    magically special that it would somehow make sense to say "git+" for 
>    it?

This is also useful for foreign SCM support; the idea of supporting svn+ssh:// "directly" with git remote and the likes.

I don't usually write git+ssh://, but I do consider it to be the form which is more in the spirit of application interoperability. It says what it is, which is ssh tunnelled git protocol.

Show 7 quoted lines
> As to your TLS example: if we were to do "git over TLS", it would make 
> perfect sense to use either "tls://" (although "gits://" might be more 
> natural, not because tls is wrong, but because people have gotten used to 
> "https://") if we were to have a "secure git" port. Or maybe we'd use the 
> same port number that we already have assigned for git, and just add some 
> "use TLS to authenticate/encrypt", and use "tls://" for that. It makes 
> perfect sense.

The scheme is bad because it doesn't integrate with other appliations. Seeing the URI in a web page they have no way of knowing which application or port this tls:// URI refers to. It's not *universal*.

This is fine for URIs passed into git, but bad if you want to link to it from elsewhere.

Sam.
Linus Torvalds· Aug 27, 2007, 23:17 UTC · re: Sam Vilain · lore

Re: repo.or.cz wishes?

On Tue, 28 Aug 2007, Sam Vilain wrote:
> 
> This is fine for URIs passed into git, but bad if you want to link to it
> from elsewhere.
..and by that logic, you should add "git+" to *everything*, not just ssh.
Which simply isn't practical or sane - only damn annoying.
			Linus
Jakub Narebski· Aug 27, 2007, 23:27 UTC · re: Linus Torvalds · lore

Re: repo.or.cz wishes?

Linus Torvalds wrote:
Show 9 quoted lines
> 
> On Tue, 28 Aug 2007, Sam Vilain wrote:
> > 
> > This is fine for URIs passed into git, but bad if you want to link to it
> > from elsewhere.
> 
> ..and by that logic, you should add "git+" to *everything*, not just ssh.
> 
> Which simply isn't practical or sane - only damn annoying.

Not exactly. You can browse using http:// and file:// protocols, rsync:// is simply rsync, while ssh:// (or git_ssh://) can be limited using git-shell.

-- 
Jakub Narebski
Poland
Linus Torvalds· Aug 27, 2007, 23:38 UTC · re: Jakub Narebski · lore

Re: repo.or.cz wishes?

On Tue, 28 Aug 2007, Jakub Narebski wrote:
> 
> Not exactly. You can browse using http:// and file:// protocols,
> rsync:// is simply rsync, while ssh:// (or git_ssh://) can be limited
> using git-shell.

Bullshit. You carefully left out "git://", since that doesn't fit your "argument".

The fact is, all git URL's make sense for *git*, not necessarily for anything else. They may have incidental meanings outside of git, but certainly nothing that is really *sensible*.

And that is how it was designed to be. The URL's are for *git*, not for other uses. If you want to do cross-SCM tools, you need to let them know it's a "git" thing wheher it's browsable or not, so the argument that ssh is something "different" is bogus crapola.

Just face it, ssh is in no way different from any of the other git URL specifiers.

			Linus
Sam Vilain· Aug 27, 2007, 23:30 UTC · re: Linus Torvalds · lore

Re: repo.or.cz wishes?

Linus Torvalds wrote:
Show 7 quoted lines
> 
> On Tue, 28 Aug 2007, Sam Vilain wrote:
>> This is fine for URIs passed into git, but bad if you want to link to it
>> from elsewhere.
> 
> ..and by that logic, you should add "git+" to *everything*, not just ssh.
> Which simply isn't practical or sane - only damn annoying.

Why annoying and impractical, if you don't ever have to specify it unless you want to write a URI which is portable between applications?

Sam.
Linus Torvalds· Aug 27, 2007, 23:34 UTC · re: Sam Vilain · lore

Re: repo.or.cz wishes?

On Tue, 28 Aug 2007, Sam Vilain wrote:
> 
> Why annoying and impractical, if you don't ever have to specify it
> unless you want to write a URI which is portable between applications?

Sure. I'm perfectly happy to make connect.c just ignore any "git+" prefix, and let people do it.

What I object to is:
 - the totally *idiotic* notion that "ssh" is somehow different
 - encouraging people to actually *use* that inconvenient format

The fact is, nobody really cares. We've happily used the non-"git+" forms for over two years, and there has never *ever* been a case of actual confusion. So allowing the "git+" prefix everywhere may be _logical_, but it's still totally idiotic and user-unfriendly.

		Linus
Jakub Narebski· Aug 27, 2007, 23:16 UTC · re: Linus Torvalds · lore

Re: repo.or.cz wishes?

Linus Torvalds wrote:
Show 7 quoted lines
> As to your TLS example: if we were to do "git over TLS", it would make 
> perfect sense to use either "tls://" (although "gits://" might be more 
> natural, not because tls is wrong, but because people have gotten used to 
> "https://") if we were to have a "secure git" port. Or maybe we'd use the 
> same port number that we already have assigned for git, and just add some 
> "use TLS to authenticate/encrypt", and use "tls://" for that. It makes 
> perfect sense.

I like gits:// idea for "git over TLS", and I'm against "tls://". I wonder if it would be hard to implement "git overt TLS"? We could resurrect patch which allowed push over git protocol, onnly restricting pushing to gits protocol.

-- 
Jakub Narebski
Poland
Jakub Narebski· Aug 27, 2007, 21:58 UTC · re: Petr Baudis · lore

Re: repo.or.cz wishes?

On Monday, 27 August 2007, Petr "Pasky" Baudis wrote:
Show 5 quoted lines
> On Mon, Aug 27, 2007 at 02:16:34AM CEST, Sven Verdoolaege wrote:
>> On Mon, Aug 27, 2007 at 01:59:44AM +0200, Petr Baudis wrote:
>>>
>>>   So now I wonder, what is the thing you miss most there? Any cool stuff
>>> repo.or.cz could (preferrably easily) do and doesn't?

Is it now possible to _upload_ SSH key, instad of copy'n'paste it when creating repository/account?

Would it be reasonable to limit repository name length, or at least modify gitweb to truncate (cut) it if it is too long?

Show 10 quoted lines
>> Just a minor nit, but how about dropping the "git+" from the
>> Push URL?
> 
> I'm a major proponent of the "git+" - it's just the correct thing to
> specify. ssh:// by itself means secure _shell_, and that's not what the
> URL means - ssh is literaily just a transport layer for the git
> protocol. This is not my invention but fairly standard thing which
> plenty of people use, and it makes it possible to select proper protocol
> handlers and so on, shall something generic crunch on the URL. I've
> never actually understood why do some people dislike it.

First, it is in the context of git, so one can say that "git+" is implied. Documentation mentions only "ssh://". That said I prefer "git+ssh://" to "ssh://" alone.

Second, IIRC Linus prefers scp-like syntax for SSH protocol, namely "[user]@host:/path/to/repo", so perhaps that one should be used instead.

Show 8 quoted lines
>> Jakub was also talking about support in gitweb for specifying
>> the location of submodules.  It would be nice if admins could
>> set this information, wherever it ends up getting stored.
> 
> Hmm, this shouldn't be very hard to do if the support will get into
> gitweb. And adding the support to gitweb shouldn't be that hard either.
> :-) OTOH, it's not something that would get me terribly excited, so I
> guess I'll wait for the gitweb side.

The problem is that repo.or.cz needs support for that in gitweb, while gitweb in turn needs support for that in git. This needs git consensus on how to specify object database location (or just gitdir) for submodules, to have later submodule support in gitweb.

-- 
Jakub Narebski
Poland
Sam Vilain· Aug 27, 2007, 02:40 UTC · re: Petr Baudis · lore

Re: repo.or.cz wishes?

Petr Baudis wrote:
Show 13 quoted lines
>   Hi,
>
>   I've just finally killed the HTTP auth for project administration that
> was destroying everyone's lives, and added support for resetting
> forgotten passwords, two main things that seemed to be the popular nits
> of the repo.or.cz audience.
>
>   So now I wonder, what is the thing you miss most there? Any cool stuff
> repo.or.cz could (preferrably easily) do and doesn't?
>
>   And please don't ask for smaller roundtrip times of requests for
> administrator assistance. ;-)
>   

I'd like to see the service mirrored, including potentially a repository which contains the meta/auth information (sans password hashes, perhaps) for admins on accounts.

This of course opens a large can of worms when it comes to achieving decentralization of the service as a whole, however I think those questions will be best answered once the information is available for setting up mirrors.

I've also got hardware, bandwidth and some tuits.
Sam.
Johannes Schindelin· Aug 27, 2007, 08:35 UTC · re: Petr Baudis · lore

Re: repo.or.cz wishes?

Hi,
On Mon, 27 Aug 2007, Petr Baudis wrote:
>   So now I wonder, what is the thing you miss most there?

I wonder if this is repo.or.cz specific, but we recently had a problem where the blobs of a project went away, and the forked project still relied on them. Any ideas how to solve that issue?

Ciao, Dscho

Uwe Kleine-König· Aug 27, 2007, 19:49 UTC · re: Petr Baudis · lore

Re: repo.or.cz wishes?

Hello Petr,
Petr Baudis wrote:
>   So now I wonder, what is the thing you miss most there? Any cool stuff
> repo.or.cz could (preferrably easily) do and doesn't?
I have two wishes (even though I don't use repo.or.cz regularly).
- searching for sha1 id.
  OK, I know how to create an URL for that, but that's inconvient.
- When looking on forks of say git.git, I want the "main" project shown,
  too.  That is http://repo.or.cz/w/git.git?a=forks should include a
  link to http://repo.or.cz/w/git.git.

After rereading these may more be gitweb wishes than repo.or.cz, but anyhow ...

Best regards Uwe

-- 
Uwe Kleine-König

http://www.google.com/search?q=half+a+cup+in+teaspoons
Sven Verdoolaege· Aug 29, 2007, 07:32 UTC · lore

Re: repo.or.cz wishes?

On Tue, Aug 28, 2007 at 11:56:10PM +0200, Jakub Narebski wrote:
Show 14 quoted lines
> On Tue, 28 August 2007, Sven Verdoolaege wrote:
> > On Mon, Aug 27, 2007 at 11:58:42PM +0200, Jakub Narebski wrote:
> >>
> >> The problem is that repo.or.cz needs support for that in gitweb, while
> >> gitweb in turn needs support for that in git. This needs git consensus
> >> on how to specify object database location (or just gitdir) for
> >> submodules, to have later submodule support in gitweb.
> > 
> > What would be the use of that (outside of gitweb) ?
> 
> For the hypothetical (planned?) future '--recurse-submodules' option
> to git-diff family, git-ls-tree and git-ls-files, git-fetch and git-push
> (but I think not git-pull), perhaps git-log (besides what it supports
> by the way of git-diff-tree), maybe even git-status and git-commit.

Ah... you're talking about bare repositories, right? For non-bare repos, I'd assume you would only recurse for those submodules that you have actually checked out in your working tree.

skimo
Jakub Narebski· Aug 29, 2007, 23:12 UTC · re: Sven Verdoolaege · lore

Re: repo.or.cz wishes?

On Wed, Aug 29, 2007, Sven Verdoolaege wrote:
Show 19 quoted lines
> On Tue, Aug 28, 2007 at 11:56:10PM +0200, Jakub Narebski wrote:
>> On Tue, 28 August 2007, Sven Verdoolaege wrote:
>>> On Mon, Aug 27, 2007 at 11:58:42PM +0200, Jakub Narebski wrote:
>>>>
>>>> The problem is that repo.or.cz needs support for that in gitweb, while
>>>> gitweb in turn needs support for that in git. This needs git consensus
>>>> on how to specify object database location (or just gitdir) for
>>>> submodules, to have later submodule support in gitweb.
>>> 
>>> What would be the use of that (outside of gitweb) ?
>> 
>> For the hypothetical (planned?) future '--recurse-submodules' option
>> to git-diff family, git-ls-tree and git-ls-files, git-fetch and git-push
>> (but I think not git-pull), perhaps git-log (besides what it supports
>> by the way of git-diff-tree), maybe even git-status and git-commit.
> 
> Ah... you're talking about bare repositories, right?
> For non-bare repos, I'd assume you would only recurse for those
> submodules that you have actually checked out in your working tree.

For bare repositories, and for repositories which move around (at least once change path). Gitweb uses repository like it is bare...

-- 
Jakub Narebski
Poland
Petr Baudis· Aug 29, 2007, 09:54 UTC · lore

Re: repo.or.cz wishes?

On Wed, Aug 29, 2007 at 06:20:05AM CEST, Shawn O. Pearce wrote:
Show 27 quoted lines
> Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:
> > On Tue, 28 Aug 2007, Theodore Tso wrote:
> > > On Tue, Aug 28, 2007 at 12:10:59AM -0400, Shawn O. Pearce wrote:
> > > > At day-job I have a hard rule that you cannot even push into an A, let 
> > > > alone rewind a branch in it or delete a branch from it.
> > > 
> > > Why don't you even allow people to push into A?  That should be safe....
> > 
> > Nope:
> > 
> > for b in $(git ls-remote /that/other/repo | sed "s/^[^ ]* //")
> > do
> > 	git push /that/other/repo :$b
> > done
> 
> Well, at day-job I use contrib/hooks/update-paranoid to deny all
> push access into my A's (/that/other/repo).  But that could just
> as easily be configured to allow branch creation and branch update
> (fast-forward) but no rewind or delete.
> 
> When I symlink A's refs into B I also don't allow B to update,
> create, rewind or delete the symlinked refs via push.  This way
> you can't do something weird to A like upload new objects into B's
> ODB but then change A's refs to point to objects that A's own ODB
> doesn't have.
> 
> Hmm, I wonder of Pasky handles that correctly on repo.or.cz...

I don't handle it at all, but if you don't have permissions to modify A you simply won't be able to do anything weird to A. If you have the permissions, I'm still not sure if Git will keep symlinked refs over ref updates; if so, hey, you had the permissions for A and it's your reponsibility if you screw up.

-- 
				Petr "Pasky" Baudis
Early to rise and early to bed makes a male healthy and wealthy and dead.
                -- James Thurber
Petr Baudis· Aug 29, 2007, 09:58 UTC · lore

Re: repo.or.cz wishes?

On Wed, Aug 29, 2007 at 07:08:21AM CEST, Junio C Hamano wrote:
Show 25 quoted lines
> "Shawn O. Pearce" <spearce@spearce.org> writes:
> 
> > Theodore Tso <tytso@mit.edu> wrote:
> >> On Tue, Aug 28, 2007 at 12:10:59AM -0400, Shawn O. Pearce wrote:
> >> > Its what happens when you use `git clone --shared A B` and the
> > ...
> >> This has been discussed before, and it wouldn't be *that* hard to have
> >> "git clone --shared" create a backpointer from B to A, so that
> >> "git-prune" could also search the B's refs and not prune anything that
> >> is in A which is reachable from heads in A and B.
> >
> > Not if I already have a pointer from B to A's refs.  repo.or.cz
> > also has this same pointer:
> >
> > 	git clone --shared A B
> > 	ln -s A/refs B/refs/forkee
> 
> Two things to watch out for are (1) packed refs won't be
> protected with this trick, and (2) symrefs in refs/ hierarchy
> will point at wrong place if you did this.  The latter hopefully
> won't be a problem because the trick being discussed is only to
> add reachability and not _using_ the borrowed refs for anything
> (iow, this makes B/refs/forkee/remote/origin/HEAD incorrectly
> point at refs/remotes/origin/master, but what it really should
> point at is B/refs/forkee/remote/origin/master).

BTW gitweb actually uses refs/forkee/ to add funny ref tags to commits, which was completely unintended but is actually in the end quite handy (though the tags should be modified to look less confusing).

-- 
				Petr "Pasky" Baudis
Early to rise and early to bed makes a male healthy and wealthy and dead.
                -- James Thurber
Theodore Tso· Aug 29, 2007, 11:13 UTC · lore

Re: repo.or.cz wishes?

On Wed, Aug 29, 2007 at 12:15:23AM -0400, Shawn O. Pearce wrote:
Show 14 quoted lines
> > This is morally the same, but it makes the hardlink step easier (only
> > one pack to link from A to B), and by using git-gc mit makes it
> > conceptually easier for people to understand what's going on.
> > 
> > git --git-dir=A gc
> > ln A/.git/objects/pack/* B/.git/objects/pack
> > git --git-dir=B gc --prune
> > git --git-dir=A prune
> 
> No, it won't work.
> 
> The problem is that during the first `git --git-dir=A gc` call
> you are deleting packfiles that may contain objects that B needs.
> *poof*.  

But "git-gc" without the --prune doesn't delete any objects. So it should always be safe to use git-gc even if there are repositories that are relying on that repo's ODB. It's only if you use git-gc --prune that you could get in troudble. It might delete some packfiles containing objects needed by B, but only after consolidating all of the objects into a single packfile that contains all of the objects that had always been in A's ODB.

So I don't see why this wouldn't work.
						- Ted
Shawn O. Pearce· Aug 31, 2007, 21:09 UTC · re: Theodore Tso · lore

Re: repo.or.cz wishes?

Theodore Tso <tytso@mit.edu> wrote:
Show 14 quoted lines
> On Wed, Aug 29, 2007 at 12:15:23AM -0400, Shawn O. Pearce wrote:
> > > 
> > > git --git-dir=A gc
> > > ln A/.git/objects/pack/* B/.git/objects/pack
> > > git --git-dir=B gc --prune
> > > git --git-dir=A prune
> > 
> > No, it won't work.
> > 
> > The problem is that during the first `git --git-dir=A gc` call
> > you are deleting packfiles that may contain objects that B needs.
> > *poof*.  
> 
> But "git-gc" without the --prune doesn't delete any objects.

Yes, it does delete objects. Even without --prune. That is because git-gc is running `git-repack -a -d -l`. repack -a means repack all objects reachable from the current refs. The -d means delete the packfiles that existed when the repack started, as it is assumed that all needed (reachable) objects were copied into the new output packfile(s). The -d also means delete any loose objects that are now packed (git-prune-packed).

Yet there may be objects in A that A cannot reach anymore (deleted or rewound branch) but that B needs and B does not have a copy of. If these objects were in one of the prior packfiles of A and is not in the new packfile(s) of A then those objects are gone. *poof*.

Show 7 quoted lines
> So it
> should always be safe to use git-gc even if there are repositories
> that are relying on that repo's ODB.  It's only if you use git-gc
> --prune that you could get in troudble.  It might delete some
> packfiles containing objects needed by B, but only after consolidating
> all of the objects into a single packfile that contains all of the
> objects that had always been in A's ODB.
But when we repack we don't repack everything in A's ODB, we only
repack the things that A can reach.  If A cannot reach something
because a branch was rewound or deleted it won't survive the repack.
Then the repack is behaving like at least partially like gc --prune.
 
> So I don't see why this wouldn't work.

It only works if A cannot delete a branch or rewind a branch. In other words, once an object is stored in A's ODB it must always be reachable from A's refs.

-- 
Shawn.
Linus Torvalds· Aug 29, 2007, 17:11 UTC · lore

Re: repo.or.cz wishes?

On Wed, 29 Aug 2007, Shawn O. Pearce wrote:
Show 6 quoted lines
> 
> Not if I already have a pointer from B to A's refs.  repo.or.cz
> also has this same pointer:
> 
> 	git clone --shared A B
> 	ln -s A/refs B/refs/forkee
Now, this doesn't work well with packed refs, I'm afraid.

So I suspect that if we really want to support something like this, we'd need to do more than just avoid the recursion when you cross-link.

		Linus
Shawn O. Pearce· Sep 1, 2007, 02:58 UTC · re: Linus Torvalds · lore

Re: repo.or.cz wishes?

Linus Torvalds <torvalds@linux-foundation.org> wrote:
Show 9 quoted lines
> On Wed, 29 Aug 2007, Shawn O. Pearce wrote:
> > 
> > Not if I already have a pointer from B to A's refs.  repo.or.cz
> > also has this same pointer:
> > 
> > 	git clone --shared A B
> > 	ln -s A/refs B/refs/forkee
> 
> Now, this doesn't work well with packed refs, I'm afraid.
No, it doesn't work well.  So I actually also avoid packing A's refs.
Which is yet another reason why my A's don't allow pushing, that
way nobody goes nuts and creates a ton of refs in there.  With only
refs/heads/master and it being unpacked its not a big deal.
 
> So I suspect that if we really want to support something like this, we'd 
> need to do more than just avoid the recursion when you cross-link.

Yes. I've been thinking about trying to better share the ODB and the ref database between repositories, but it has been low priority for me.

I rely on this ref symlinking/alternate ODB trick a lot at day-job to help me cope with an ugly situation I created across a number of repositories. Most of our codebase came from one Git repository, but has been refactored and split into about 10 different Git repositories. I did that refactoring by just cloning and deleting the uninteresting content, so each repository actually has a huge block of its history in common with the other 9.

One such A is "common-crap.git" that is the shared common history. Since its strictly history nobody changes that repository, and everyone borrows objects from it. This reduces my common working set by about 900MiB, as the history lives in only one packfile and not in 10.

There are obviously other ways to deal with this:
 - start the 10 repositories over again and use info/grafts to
   reinsert the old history when/if required;
 - just hardlink the same .keep'd packfile into the 10 repositories,
   since it is held by .keep it won't be touched during repack.

So one reason it has been low priority for me to improve upon is because there's more than one way to solve the problem, and the particular solution I have settled upon may not be the best solution for anyone.

Though I think we can all agree that repo.or.cz's use of forks is increasingly more popular, and one of the more powerful social features of git. Better supporting it out of the box by making it easier to setup and manage can only be a good thing for our users.

-- 
Shawn.

← back to recent threads