threads / rfc / 43106

Re: [RFC] Submodules in GIT

Subject: Re: [RFC] Submodules in GIT

## tl;dr

82 messages between Nov 20, 2006 and Dec 2, 2006.

replies: 81people: 14as markdown or json

Jakub Narebski· Nov 20, 2006, 22:16 UTC · lore
Martin Waitz wrote:
Show 6 quoted lines
> A submodule really is part of the parent tree, so it is very natural to
> add the link to the submodule commit into the GIT tree data structure.
> In addition to links to blobs and other trees, they can now also hold
> a link to a commit, which in turn has the pointers to the submodule tree
> and its history.  In order to differenciate a submodule entry with
> normal file or directory entries, they get a special file mode.

Erm... isn't a _type_ of tree entry saved somewhere? Currently it can be only 'tree' or 'blob', what you do is adding 'commit' (then permissions are permissions of top tree of module, of course).

By the way, in todo branch, in Subpro.txt, there is talk about adding link to submodule trees in _commit object_... well link to submodule tree or commit, with the "mount point".

-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Martin Waitz· Nov 20, 2006, 22:28 UTC · re: Jakub Narebski · lore
hoi :)
On Mon, Nov 20, 2006 at 11:16:45PM +0100, Jakub Narebski wrote:
Show 11 quoted lines
> Martin Waitz wrote:
> > A submodule really is part of the parent tree, so it is very natural to
> > add the link to the submodule commit into the GIT tree data structure.
> > In addition to links to blobs and other trees, they can now also hold
> > a link to a commit, which in turn has the pointers to the submodule tree
> > and its history.  In order to differenciate a submodule entry with
> > normal file or directory entries, they get a special file mode.
>
> Erm... isn't a _type_ of tree entry saved somewhere? Currently it can
> be only 'tree' or 'blob', what you do is adding 'commit' (then permissions
> are permissions of top tree of module, of course).

It is saved inside the object which is being refered to. Right now tree objects are also identified by their file mode and not by the type of object which is referenced.

> By the way, in todo branch, in Subpro.txt, there is talk about adding
> link to submodule trees in _commit object_... well link to submodule tree
> or commit, with the "mount point".

But isn't the submodule really part of the tree? Right now the commit is used to construct the history of one project. And a submodule is not part of the history of the parent, it is part of the parent's tree.

-- 
Martin Waitz
Junio C Hamano· Nov 20, 2006, 22:43 UTC · re: Jakub Narebski · lore
Jakub Narebski <jnareb@gmail.com> writes:
> By the way, in todo branch, in Subpro.txt, there is talk about adding
> link to submodule trees in _commit object_... well link to submodule tree
> or commit, with the "mount point".

That was shot down by Linus and I agree with him. "bind" was a bad idea because binding of a particular subproject commit into a tree is a property of the tree, not one of the commits that happen to have that tree.

Jakub Narebski· Nov 20, 2006, 23:02 UTC · re: Junio C Hamano · lore
Junio C Hamano wrote:
Show 10 quoted lines
> Jakub Narebski <jnareb@gmail.com> writes:
> 
>> By the way, in todo branch, in Subpro.txt, there is talk about adding
>> link to submodule trees in _commit object_... well link to submodule tree
>> or commit, with the "mount point".
> 
> That was shot down by Linus and I agree with him.  "bind" was a
> bad idea because binding of a particular subproject commit into
> a tree is a property of the tree, not one of the commits that
> happen to have that tree.
  
"bind" was kind of "mount tree" idea; I agree that adding subproject
commits to trees is better idea than adding commits or trees to
superproject commit object.
By the way, what permissions get the subproject tree?

I wonder if it makes sense to be able to add tag objects instead of commit objects to trees (depeel to tree or blob)...

-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Martin Waitz· Nov 20, 2006, 23:52 UTC · re: Jakub Narebski · lore
hoi :)
On Tue, Nov 21, 2006 at 12:02:34AM +0100, Jakub Narebski wrote:
> By the way, what permissions get the subproject tree?

In my approach no permissions are saved in the object database, only the special bit to mark the submodule. When checking out, the directory is created 0777 modulo umask, just as other directories. Then the submodule contents are checked out with their normal permissions.

-- 
Martin Waitz
Sam Vilain· Nov 21, 2006, 01:31 UTC · re: Jakub Narebski · lore
Jakub Narebski wrote:
> I wonder if it makes sense to be able to add tag objects instead
> of commit objects to trees (depeel to tree or blob)...
>   

I'd say "as well as", and the semantics should be that to something browsing the filesystem, a tag looks like the type of object it refers to. eg, tag a tree, it's a tree, tag a commit, it's a sub-project/tree, tag a blob, it's a file.

The use case I'm thinking of is semi-transparent storing of archives; instead of storing the archive body, store a tag which contains the "extra" information - like the gzip headers for a gz and which compression options are needed to reproduce the same output stream. For a tar, the per-file information such as the filestamps, owner and permissions are recorded, and it points to a tree. A clever porcelain could detect these file types, and make sure the uncompressed streams are stored.

People who are using clients which don't understand these tag objects in between will get the contents of the node checked out instead, so instead of getting "foo.tar.gz" as a file, I got a "foo.tar.gz/" directory.

Linus Torvalds· Nov 20, 2006, 23:05 UTC · re: Junio C Hamano · lore
On Mon, 20 Nov 2006, Junio C Hamano wrote:
Show 5 quoted lines
> 
> That was shot down by Linus and I agree with him.  "bind" was a
> bad idea because binding of a particular subproject commit into
> a tree is a property of the tree, not one of the commits that
> happen to have that tree.

Yes. I think it would be a _fine_ idea to have a new tree-entry type that points to a sub-commit, but it really does need to be on a "tree level", not a commit level.

If it's on a tree level, getting things like "git diff" etc to work is not impossible, and it will also fit very well into the whole git infrastructure.

So right now a tree entry can be another tree or a blob - and the only extension would be to add a "commit" type (which would largely _act_ as a tree entry, at least for sorting, ie it would use the same "sorts as if it had a '/' at the end" logic).

Now, to get everything to work seamlessly within such a commit thing might be a fair amount of work, but I'm not sure you even _need_ to. It might be ok to just say "subproject 'xyzzy' differs" in the diff, for example, and have some rudimentary support for "git status" etc talking about subprojects that need to be committed.

J. Bruce Fields· Nov 20, 2006, 23:25 UTC · re: Linus Torvalds · lore
On Mon, Nov 20, 2006 at 03:05:47PM -0800, Linus Torvalds wrote:
Show 12 quoted lines
> 
> 
> On Mon, 20 Nov 2006, Junio C Hamano wrote:
> > 
> > That was shot down by Linus and I agree with him.  "bind" was a
> > bad idea because binding of a particular subproject commit into
> > a tree is a property of the tree, not one of the commits that
> > happen to have that tree.
> 
> Yes. I think it would be a _fine_ idea to have a new tree-entry type that 
> points to a sub-commit, but it really does need to be on a "tree level", 
> not a commit level.

Would it also be possible to allow the "Tree:" line in the commit object to refer to a commit, or does the root of the project need to be a special case?

Martin Waitz· Nov 20, 2006, 23:33 UTC · re: J. Bruce Fields · lore
hoi :)
On Mon, Nov 20, 2006 at 06:25:07PM -0500, J. Bruce Fields wrote:
Show 8 quoted lines
> On Mon, Nov 20, 2006 at 03:05:47PM -0800, Linus Torvalds wrote:
> > Yes. I think it would be a _fine_ idea to have a new tree-entry type that 
> > points to a sub-commit, but it really does need to be on a "tree level", 
> > not a commit level.
> 
> Would it also be possible to allow the "Tree:" line in the commit object
> to refer to a commit, or does the root of the project need to be a
> special case?

this would then be something like the branch-archival proposal. The user interface for such a beast would be difficult, as you have to somehow specify if you mean the inner or outer repository.

-- 
Martin Waitz
J. Bruce Fields· Nov 21, 2006, 18:01 UTC · re: Martin Waitz · lore
On Tue, Nov 21, 2006 at 12:33:34AM +0100, Martin Waitz wrote:
Show 6 quoted lines
> On Mon, Nov 20, 2006 at 06:25:07PM -0500, J. Bruce Fields wrote:
> > Would it also be possible to allow the "Tree:" line in the commit object
> > to refer to a commit, or does the root of the project need to be a
> > special case?
> 
> this would then be something like the branch-archival proposal.

Do you have any pointers to previous discussion? (A couple obvious searches don't turn up anything for me.)

Martin Waitz· Nov 21, 2006, 19:32 UTC · re: J. Bruce Fields · lore
On Tue, Nov 21, 2006 at 01:01:27PM -0500, J. Bruce Fields wrote:
Show 10 quoted lines
> On Tue, Nov 21, 2006 at 12:33:34AM +0100, Martin Waitz wrote:
> > On Mon, Nov 20, 2006 at 06:25:07PM -0500, J. Bruce Fields wrote:
> > > Would it also be possible to allow the "Tree:" line in the commit object
> > > to refer to a commit, or does the root of the project need to be a
> > > special case?
> > 
> > this would then be something like the branch-archival proposal.
> 
> Do you have any pointers to previous discussion?  (A couple obvious
> searches don't turn up anything for me.)
Aug 04 Eric W. Biederman    [RFC][PATCH] Branch history

I really think that using subprojects can be used for this workflow, too. But adding a submodule directly to the root is not really possible, we'd have to use special user interfaces for that, even when the git-core might be able to handle it. But what might be possible is to have one toplevel history-tracking repository in e.g. ~/src and then add all the repositories you work with as a submodule. Whenever you want to record the history of some project, you can simply commit it to ~/src.

-- 
Martin Waitz
Martin Waitz· Nov 20, 2006, 23:29 UTC · re: Linus Torvalds · lore
hoi :)
On Mon, Nov 20, 2006 at 03:05:47PM -0800, Linus Torvalds wrote:
Show 5 quoted lines
> Now, to get everything to work seamlessly within such a commit thing 
> might be a fair amount of work, but I'm not sure you even _need_ to. It 
> might be ok to just say "subproject 'xyzzy' differs" in the diff, for 
> example, and have some rudimentary support for "git status" etc talking 
> about subprojects that need to be committed.
this is exactly the status of my implementation at the moment ;-)

Well, it does not yet explicitly tell that a subproject diffs, but it just creates a diff of the two commit objects.

I guess we need some command line option to say if we only want to know about that the submodule changes or if the diff should recurse into it.

-- 
Martin Waitz
Junio C Hamano· Nov 21, 2006, 00:10 UTC · re: Linus Torvalds · lore
Linus Torvalds <torvalds@osdl.org> writes:
Show 5 quoted lines
> Now, to get everything to work seamlessly within such a commit thing 
> might be a fair amount of work, but I'm not sure you even _need_ to. It 
> might be ok to just say "subproject 'xyzzy' differs" in the diff, for 
> example, and have some rudimentary support for "git status" etc talking 
> about subprojects that need to be committed.

I agree with the static "diff" part, and probably "checkout" and "merge" are not all that difficult.

However, if I recall correctly, it was rather nightmarish to make this also work for reachability traversal necessary for pack generation. It was painful enough even when the bind was at the commit level (which was way simpler to handle), but to do this the right way, the bind needs to be done at the tree level, and "rev-list --objects foo..bar" would need some way to limit the commit ancestry chain of subproject at the same time, by computing the commit ancestry of the embedded commits in the trees.

Jakub Narebski· Nov 21, 2006, 00:42 UTC · re: Junio C Hamano · lore
Junio C Hamano wrote:
Show 20 quoted lines
> Linus Torvalds <torvalds@osdl.org> writes:
> 
>> Now, to get everything to work seamlessly within such a commit thing 
>> might be a fair amount of work, but I'm not sure you even _need_ to. It 
>> might be ok to just say "subproject 'xyzzy' differs" in the diff, for 
>> example, and have some rudimentary support for "git status" etc talking 
>> about subprojects that need to be committed.
> 
> I agree with the static "diff" part, and probably "checkout" and
> "merge" are not all that difficult.
> 
> However, if I recall correctly, it was rather nightmarish to
> make this also work for reachability traversal necessary for
> pack generation.  It was painful enough even when the bind was
> at the commit level (which was way simpler to handle), but to do
> this the right way, the bind needs to be done at the tree level,
> and "rev-list --objects foo..bar" would need some way to limit
> the commit ancestry chain of subproject at the same time, by
> computing the commit ancestry of the embedded commits in the
> trees.

Perhaps it would be best to join those two subproject support solutions together: "bind" tree/commit mount header in commit object, and "commit" entry in a tree. But I agree that revision walking needs to be rewamped... well, unless you always have project and subproject in the same repository, and subprojects are branches in the project too...

-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Martin Waitz· Nov 21, 2006, 06:21 UTC · re: Jakub Narebski · lore
hoi :)
On Tue, Nov 21, 2006 at 01:42:22AM +0100, Jakub Narebski wrote:
> Perhaps it would be best to join those two subproject support
> solutions together: "bind" tree/commit mount header in commit
> object, and "commit" entry in a tree.

But which is the autoritative source then? Does it give any more information?

The advantage in your proposal would be that submodules would be visible immediately when looking at the commit, without having to traverse the entire tree. This may be worthwhile when showing the combined history of parent and submodules.

But still this looks like "caching submodule information in the commit object" and I do not know if we really want to do that.

-- 
Martin Waitz
Jakub Narebski· Nov 21, 2006, 10:04 UTC · re: Martin Waitz · lore
Martin Waitz wrote:
Show 7 quoted lines
> On Tue, Nov 21, 2006 at 01:42:22AM +0100, Jakub Narebski wrote:
>> Perhaps it would be best to join those two subproject support
>> solutions together: "bind" tree/commit mount header in commit
>> object, and "commit" entry in a tree.
> 
> But which is the autoritative source then?
> Does it give any more information?

Both should contain the same information, otherwise repository is corrupt (is in inconsistent state).

"bind" header in commit objects is meant as a kind of shortcut, to ease reachability checking (you don't need to recurse into directories).

Show 5 quoted lines
> The advantage in your proposal would be that submodules would
> be visible immediately when looking at the commit,
> without having to traverse the entire tree.
> This may be worthwhile when showing the combined history of parent
> and submodules.
That was the idea.
> But still this looks like "caching submodule information in the
> commit object" and I do not know if we really want to do that.

Well, we would be repeating information, sure. But we can put additional information in "bind" header except sha1 of commit and mount point... although I cannot think what... :)

-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Martin Waitz· Nov 21, 2006, 11:49 UTC · re: Jakub Narebski · lore
hoi :)
On Tue, Nov 21, 2006 at 11:04:46AM +0100, Jakub Narebski wrote:
> "bind" header in commit objects is meant as a kind of shortcut, to ease
> reachability checking (you don't need to recurse into directories).

Well, but you already have to recurse to find all objects which are reachable by a commit, so you don't loose anything.

Show 7 quoted lines
> > The advantage in your proposal would be that submodules would
> > be visible immediately when looking at the commit,
> > without having to traverse the entire tree.
> > This may be worthwhile when showing the combined history of parent
> > and submodules.
> 
> That was the idea.

On the other hand that only has to be done once anyway. After you traversed the tree once you can create your own (in memory) cache of submodules connected to the tree. While walking the commits backwards, you only have to check those parts of the tree which have changed. So it may even be suitable for larger repositories. But clearly it is not as low as with the in-commit cache. So we have to weight complexity of the data storage with runtime complexity. Opinions?

-- 
Martin Waitz
Martin Waitz· Nov 21, 2006, 06:27 UTC · re: Junio C Hamano · lore
hoi :)
On Mon, Nov 20, 2006 at 04:10:50PM -0800, Junio C Hamano wrote:
Show 9 quoted lines
> However, if I recall correctly, it was rather nightmarish to
> make this also work for reachability traversal necessary for
> pack generation.  It was painful enough even when the bind was
> at the commit level (which was way simpler to handle), but to do
> this the right way, the bind needs to be done at the tree level,
> and "rev-list --objects foo..bar" would need some way to limit
> the commit ancestry chain of subproject at the same time, by
> computing the commit ancestry of the embedded commits in the
> trees.

This at least seems to work already. The UNINTERESTING flag is recursively set for the submodule commits while walking the object chain.

But I must admit that I only did very simple tests up to now. Do you have any special constellations in mind which were difficult to support?

-- 
Martin Waitz
Junio C Hamano· Nov 21, 2006, 07:36 UTC · re: Martin Waitz · lore
Martin Waitz <tali@admingilde.org> writes:
Show 15 quoted lines
> On Mon, Nov 20, 2006 at 04:10:50PM -0800, Junio C Hamano wrote:
>
>> However, if I recall correctly, it was rather nightmarish to
>> make this also work for reachability traversal necessary for
>> pack generation.  It was painful enough even when the bind was
>> at the commit level (which was way simpler to handle), but to do
>> this the right way, the bind needs to be done at the tree level,
>> and "rev-list --objects foo..bar" would need some way to limit
>> the commit ancestry chain of subproject at the same time, by
>> computing the commit ancestry of the embedded commits in the
>> trees.
>
> This at least seems to work already.
> The UNINTERESTING flag is recursively set for the submodule
> commits while walking the object chain.

I think that is fine as long as we somehow enforce the topology of submodule to be similar to the toplevel topology. Otherwise I suspect it leads to unintuitive behaviour.

Suppose that the ancestry chain for the toplevel are A, A~1, A~2 and you asked for "A~2..A". A submodule is bound at tree "sub/" and suppose A:sub/ == B, A~1:sub/ == C, and A~2:sub/ == D.

Now further suppose the ancestry chain for B, C and D are like this:

              o---C
             /     \
     ...o---o---D---B

A naive implementation of "--objects A~2..A" would propagate UNINTERESTING to D and mark B and C unmarked. Would it however be reasonable to include commits marked as 'o'?

I am not trying to be negative here, but just raising things that I did not think through when I tried to tackle it the last time...

Martin Waitz· Nov 21, 2006, 07:55 UTC · re: Junio C Hamano · lore
hoi :)
On Mon, Nov 20, 2006 at 11:36:55PM -0800, Junio C Hamano wrote:
Show 18 quoted lines
> I think that is fine as long as we somehow enforce the topology
> of submodule to be similar to the toplevel topology.  Otherwise
> I suspect it leads to unintuitive behaviour.
> 
> Suppose that the ancestry chain for the toplevel are A, A~1, A~2
> and you asked for "A~2..A".  A submodule is bound at tree "sub/"
> and suppose A:sub/ == B, A~1:sub/ == C, and A~2:sub/ == D.
> 
> Now further suppose the ancestry chain for B, C and D are like
> this:
> 
>               o---C
>              /     \
>      ...o---o---D---B
> 
> A naive implementation of "--objects A~2..A" would propagate
> UNINTERESTING to D and mark B and C unmarked.  Would it however
> be reasonable to include commits marked as 'o'?

I think it is reasonable to just go on as in a normal repository. That is, pretend we want to list D..B and mark all commits which are reachable.

-- 
Martin Waitz
Yann Dirson· Nov 21, 2006, 22:31 UTC · re: Linus Torvalds · lore
On Mon, Nov 20, 2006 at 03:05:47PM -0800, Linus Torvalds wrote:
Show 10 quoted lines
> On Mon, 20 Nov 2006, Junio C Hamano wrote:
> > 
> > That was shot down by Linus and I agree with him.  "bind" was a
> > bad idea because binding of a particular subproject commit into
> > a tree is a property of the tree, not one of the commits that
> > happen to have that tree.
> 
> Yes. I think it would be a _fine_ idea to have a new tree-entry type that 
> points to a sub-commit, but it really does need to be on a "tree level", 
> not a commit level.

I'm not sure I get the reason why the submodule should not be recorded on "commit level".

What I'm thinking of would be that the submodule tree would just be a standard antry of a tree in the supermodule, and we could record the submodule commit (pointing to the submodule tree) in the supermodule commit.

This idea came when thinking about implementing partial merges. That is, when different people are responsible for different parts of the tree, and thus when merging a given branch, each dev has to make only a partial merge of the full tree. Having submodule commits referenced directly from the supercommit would make it much easier to finalize the merge (ie. merging the full project while taking into account that some subtrees have been merged already).

Best regards,
Linus Torvalds· Nov 21, 2006, 22:51 UTC · re: Yann Dirson · lore
On Tue, 21 Nov 2006, Yann Dirson wrote:
> 
> I'm not sure I get the reason why the submodule should not be recorded
> on "commit level".
Because that would be STUPID.

What does the submodules have to do with the commit level? Nothing. Nada. Zero.

Submodules are _directories_. They can be anywhere in the directory tree. If you try to encode that in a commit message, you're going to totally break the whole notion of trying to "diff" two trees.

All of git is designed around the notion that a tree is the directory structure. If you put directory structure somewhere else, you totally screw all abstractions.

Now, if that weren't enough, let me enumerate _another_ reason why it's idiotic and wrong, namely the fact that a "commit" is fundamnetally the wrong place to add something like that _anyway_. Quite apart from the fact that we describe directory trees with (wait for it): "tree objects", the thing is, a commit is about a totally different _dimension_ altogether.

The only and _whole_ point of a "commit" is to describe the "time dimension". Something that doesn't always change in time should not be in a commit object, because it is by definition not what a commit is all about. A commit should describe the relationship of itself to other commits, ie it's a "how did this change".

And a sub-project simply doesn't even _do_ that. Much of the time, a subproject stays constant, and is not something that comes and goes on an individual commit basis.

I don't understand why people are so fixated with putting things in the wrong object. WHY do people want to put crap in the "commit" object? People have wanted to put "rename" information there (which is stupid for all the same reasons: renames _remain_. They aren't a one-time event. If something was renamed in commit X, it will _remain_ renamed in commit X+1, so it's clearly not really a "commit X" thing)

Think of it this way:
 - if something _only_ makes sense on an _individual commit_ level, it 
   goes into the "commit object". But if it makes sense for "git diff",
   then it MUST NOT be in a commit object, because you do "git diff" over
   a big _range_ of commit objects.

Think "git show". The "author" of a commit is only associated with a _single_ commit. It thus goes into the commit object, and nowhere else. Same goes for time, and commit message. A commit message is fundamentally a "this explains this _one_ commit".

But anything that you expect to have in a "range" of commits MUST NOT be in a "commit object". If I do "git diff v2.6.13..v2.6.14", and I expect the behaviour you want to encode to show up (and dammit, subprojects very much fall under that heading - exactly the same way renames must have meaning _outside_ of a single commit) then clearly it is NOT something that is associated with any individual commits. It's something that is associated with the _state_ of the project.

And the _state_ of the project is the "tree". Not the commit. The commit is about the _history_ of the project.

So please understand this: "commit" is about the time-dimension ("history"). "tree" is about the space-dimension ("state"). The two are _related_, but they are also very much different concepts, and "related" does not mean "you can mix them up".

Sub-projects are clearly not about "time". They are about "state".
Linus Torvalds· Nov 21, 2006, 22:59 UTC · re: Linus Torvalds · lore
On Tue, 21 Nov 2006, Linus Torvalds wrote:
> 
> Submodules are _directories_.

Side note - you can do submodules other ways, but if you do, you'll almost certainly go crazy.

You could, for example, make submodules be some kind of "union filesystem", where you allow overlapping trees. It's conceptually possible. It's also horribly horribly wrong, if only because I guarantee that you'll have so many problems with it that you will only end up with a mess that is even worse than "branches" in CVS.

Yann Dirson· Nov 21, 2006, 23:54 UTC · re: Linus Torvalds · lore
On Tue, Nov 21, 2006 at 02:51:56PM -0800, Linus Torvalds wrote:
Show 11 quoted lines
> 
> 
> On Tue, 21 Nov 2006, Yann Dirson wrote:
> > 
> > I'm not sure I get the reason why the submodule should not be recorded
> > on "commit level".
> 
> Because that would be STUPID.
> 
> What does the submodules have to do with the commit level? Nothing. Nada. 
> Zero.

Oh, I see I may have expressed something in the wrong way :) Namely, I brought an idea coming from partial merges into a discussion on submodules, because when thinking about the former, I realized we could maybe use similar mechanisms for both.

Note that the proposal I outlined did not break the tree, in that the sumodule tree is still in the same place. In the case of a partial merge, the info that a subtree has been merged in this commit is indeeed part of the commit itself.

I agree that the subtree case is somewhat different, and my idea may not apply to submodules after all :)

A question would be, do "submodules" have to be permanent objects ? I suppose it depends on what people want to use them for. Indeed, the "submodule" names strongly carries the idea of a permanent subset of the repository. My proposal partial merges could be seen as using transient submodules: they do not matter much during most of the repo life.

Put it another way, I see the proposal of allowing tree entries to be commits in addition to trees and blobs, akin to recording the submodule _history_ inside the _tree_, which I feel precisely violates the distinction you want to keep between those 2 concepts.

> And a sub-project simply doesn't even _do_ that. Much of the time, a 
> subproject stays constant, and is not something that comes and goes on an 
> individual commit basis. 

What about the case of a subproject that would evolve fast, and for which we may not want intermediate versions to be part of the supermodule ? (just exploring an idea without real connection to the one discussed above)

I mean, I have a tree in which the whole software for an embedded platform is stored, including kernel, apps, etc. While working on the kernel, I may want to do several commits to that submodule, and may not want to commit to the supermodule for each kernel commit, only when I feel the kernel is stable enough.

One may argue I just have to use a branch. Anyway, there will be a need for submodule-specific branches - eg. kernel.org ones in my case.

An alternative would be to allow committing to the submodule without creating matching supermodule commits, and let the user decide when he wants to commit at the higher level. That way, 2 successive supermodule commits could have non-successive "subcommits".

Best regards,
Shawn Pearce· Nov 22, 2006, 03:40 UTC · re: Yann Dirson · lore
Yann Dirson <ydirson@altern.org> wrote:
> Put it another way, I see the proposal of allowing tree entries to be
> commits in addition to trees and blobs, akin to recording the submodule
> _history_ inside the _tree_, which I feel precisely violates the
> distinction you want to keep between those 2 concepts.
No.  Linus is right.  Submodule commits belong in the tree.

We want to record a specific subtree within a larger tree. There are three ways we can refer to a tree: by its tree SHA1, by a commit which points at the tree SHA1, or by a tag which points at a commit which points at the tree SHA1, or by a tag which points at a tag which points at a commit which points at a tree SHA1. Which is basically a tree-ish.

The advantage of linking to the commit-ish (commit or tag) and
not the tree-ish for a submodule is that it also provides you quick
access to answer the "how did this tree arive at this state" question
as the answer cannot come solely from the top level commit chain.
The reason... keep reading...
 
> What about the case of a subproject that would evolve fast, and for
> which we may not want intermediate versions to be part of the
> supermodule ?  (just exploring an idea without real connection to the
> one discussed above)

Right. The submodule is free to be committed to an infinite number of times for any given commit in the supermodule.

It is expected that users will commit to a submodule say hundreds of times for every commit they make to the supermodule. Or thousands. This is especially true if the submodule is some very large project, e.g. the Linux kernel, and the supermodule "upgrades" the kernel it is using after 3 months of staying on the same version. Suddenly the supermodule has only 1 commit which covers maybe 10,000 commits in the submodule.

Yet we still want to be able to efficiently perform operations like "git bisect" within the scope of that submodule, to help narrow down a particular bug that is within that submodule. To do that we need the commit chain (all 10,000 of those commits) in the submodule. To get those we really need a commit-ish and not a tree-ish, as going from a tree-ish to a commit-ish is not only not unique but is also pretty infeasible to do (you need to scan *every* commit).

Yann Dirson· Nov 23, 2006, 23:23 UTC · re: Shawn Pearce · lore
On Tue, Nov 21, 2006 at 10:40:56PM -0500, Shawn Pearce wrote:
Show 7 quoted lines
> Yet we still want to be able to efficiently perform operations like
> "git bisect" within the scope of that submodule, to help narrow down
> a particular bug that is within that submodule.  To do that we need
> the commit chain (all 10,000 of those commits) in the submodule.
> To get those we really need a commit-ish and not a tree-ish, as
> going from a tree-ish to a commit-ish is not only not unique but
> is also pretty infeasible to do (you need to scan *every* commit).

We don't need to have commits in the tree for this. We'll just have submodule commits which are not attached to a supermodule commit, and we can access the whole submodule history through the submodule .git/HEAD, just like we do for a standard git project.

Or do I miss something else ?
Best regards,
Shawn Pearce· Nov 25, 2006, 06:53 UTC · re: Yann Dirson · lore
Yann Dirson <ydirson@altern.org> wrote:
Show 13 quoted lines
> On Tue, Nov 21, 2006 at 10:40:56PM -0500, Shawn Pearce wrote:
> > Yet we still want to be able to efficiently perform operations like
> > "git bisect" within the scope of that submodule, to help narrow down
> > a particular bug that is within that submodule.  To do that we need
> > the commit chain (all 10,000 of those commits) in the submodule.
> > To get those we really need a commit-ish and not a tree-ish, as
> > going from a tree-ish to a commit-ish is not only not unique but
> > is also pretty infeasible to do (you need to scan *every* commit).
> 
> We don't need to have commits in the tree for this.  We'll just have
> submodule commits which are not attached to a supermodule commit, and we
> can access the whole submodule history through the submodule .git/HEAD,
> just like we do for a standard git project.
No.  You cannot do that.

How do we setup .git/HEAD when bisecting the supermodule? Or merging it? Or doing anything else with it?

Ideally the .git/HEAD of every submodule should seek to the commit that points at the tree of the submodule which the supermodule is referencing. This lets you then perform a bisect within the submodule when you identify the supermodule commit which caused the breakage.

We need the submodule commits to do this. Doing it without is too expensive.

Yann Dirson· Nov 25, 2006, 11:12 UTC · re: Shawn Pearce · lore
On Sat, Nov 25, 2006 at 01:53:38AM -0500, Shawn Pearce wrote:
Show 10 quoted lines
> Yann Dirson <ydirson@altern.org> wrote:
> > We don't need to have commits in the tree for this.  We'll just have
> > submodule commits which are not attached to a supermodule commit, and we
> > can access the whole submodule history through the submodule .git/HEAD,
> > just like we do for a standard git project.
> 
> No.  You cannot do that.
> 
> How do we setup .git/HEAD when bisecting the supermodule?
> Or merging it?  Or doing anything else with it?

Would there be any problem assuming git-update-ref would take care of updating it ?

> Ideally the .git/HEAD of every submodule should seek to the commit
> that points at the tree of the submodule which the supermodule
> is referencing.
You mean, whenever we seek the HEAD of the supermodule, right ?
> This lets you then perform a bisect within the
> submodule when you identify the supermodule commit which caused
> the breakage.
 
That is, first bisect the supermodule (which naturally bisects the
submodule with rough granularity, assuming there are many submodule
commits for at least some supermodule commits), then bisect the submodule
between the two commits identified at supermodule level, right ?
> We need the submodule commits to do this.  Doing it without is
> too expensive.
Maybe I missed something again, but I'm still not convinced :)
Linus Torvalds· Nov 25, 2006, 18:57 UTC · re: Yann Dirson · lore
On Sat, 25 Nov 2006, Yann Dirson wrote:
Show 9 quoted lines
> 
> > This lets you then perform a bisect within the
> > submodule when you identify the supermodule commit which caused
> > the breakage.
>  
> That is, first bisect the supermodule (which naturally bisects the
> submodule with rough granularity, assuming there are many submodule
> commits for at least some supermodule commits), then bisect the submodule
> between the two commits identified at supermodule level, right ?
Right. That is how you _must_ do it.
The reason is:
 - the supermodule will not track every release of the submodule. One of 
   the biggest reasons for using submodules in the first place is that the 
   submodules have their own development _independently_ of the 
   supermodule, and usually the supermodule will import new versions of 
   submodules only occasionally (eg the supermodule might choose to track 
   only major releases of the submodule, for example)
   (And yes, I realize that this is not necessarily the only submodule 
   usage: sometimes the submodules are literally _only_ developed as 
   submodules, and you'd never develop them independently. It depends on 
   the situation)
 - As a resule of the above, you MUST NOT do bisection at the submodule 
   level at first: it's entirely possible that the supermodule never ever 
   actually used the submodule state at a finer granularity, and 
   "bisecting" into such state would be idiotic (it's really no different 
   from "bisecting" a regular commit by splitting up a commit into patches 
   against individual files - sure, it's a smaller granularity, but it's a 
   granularity that never _existed_, and was never tested or intended to 
   work!)
So yes, you should expect that
 (a) submodule changes "jump around" in the supermodule - even to the 
     point of going backwards in time as far as the submodule is concerned 
     (ie the supermodule might have tested a new release of a submodule, 
     committed that, found a problem, and decided to just go back to an 
     earlier version of the submodule again, and committed that again)
 (b) This implies very much that there can be a n:m relationship between 
     submodule and supermodule commits. A supermodule commit does _not_ 
     imply a commit in the submodule (it might commit changes to the 
     top-level makefile or to _another_ submodule), but equally, a 
     submodule commit does _not_ imply a commit in the supermodule 
     (because the submodule might be independently changed in some other 
     repository where it's the _primary_ development, not a submodule)

So you shouldn't expect submodules to be very "tightly" coupled, and I don't think you even want the workflow to _be_ that tight. I think it's ok if submodules show up as such, and that "git diff" etc don't try to make it all "seamless".

It often _shouldn't_ be seamless: you should be able to commit to a supermodule without committing the submodule state: it's really no different from committing individual files (it migth be somethign that is _discouraged_ as a workflow for some project, the same way you might discourage using "git commit one/file" over "git commit -a", and for the same reason: you're committing some state that doesn't match what your tree actually looks like).

Similarly, doing a "git commit -a" within a submodule should really just commit _that_ submodule, and not even _try_ to know about supermodules etc, because the submodule really should be a totally independent git repository.

[ Side note: you may well want to set up submodules so that they share the 
  object store with the supermodule: that may be the simplest way to make 
  operations that traverse things recursively work out, since it means 
  that you can do object lookups for everythign you traverse without 
  having to even think about it.
  On the other hand, this could equally easily be done by just making 
  every submodule an "alternates" directory in the supermodule: that keeps 
  the object databases separate, but means that anybody in the supermodule 
  will always be able to look up all the objects in the submodules. So 
  even here, we certainly _can_ keep things separated, without even 
  introducing any new concepts. ]

So I actually think that submodules should at least start out as something rather independent, where a "commit -a" in the supermodule will _only_ commit the supermodule itself - and if you haven't committed the submodule yet, you'll just get the current HEAD state of the submodule.

Add some trivial help in "git status" to _warn_ about the fact that submodules haven't been committed and are dirty, but I really think that it should be a very explicit thing where you really do see things as submodules, not as "one big module".

Steven Grimm· Nov 25, 2006, 19:19 UTC · re: Linus Torvalds · lore
Linus Torvalds wrote:
> So I actually think that submodules should at least start out as something 
> rather independent, where a "commit -a" in the supermodule will _only_ 
> commit the supermodule itself - and if you haven't committed the submodule 
> yet, you'll just get the current HEAD state of the submodule.

That would make it impossible to atomically commit a change that affects two submodules, yes? I think cross-submodule commit is highly desirable and will be a fairly common use case for submodules if it's supported. For example, if you have "client" and "server" submodules and someone makes a protocol change, you don't want some unwitting developer to pull just half of the change and end up with incompatible code in the two submodules.

I have no problem with making the "only commit the supermodule" behavior the default and requiring a command-line option for the "commit everything" case, but I think "commit everything" is useful. And honestly IMO it should be the default since it'll behave in a less surprising way; when I do a "commit -a" I expect all my changes to be committed, whether they're in submodules or not.

Linus Torvalds· Nov 25, 2006, 19:30 UTC · re: Steven Grimm · lore
On Sat, 25 Nov 2006, Steven Grimm wrote:
Show 8 quoted lines
> Linus Torvalds wrote:
> > So I actually think that submodules should at least start out as something
> > rather independent, where a "commit -a" in the supermodule will _only_
> > commit the supermodule itself - and if you haven't committed the submodule
> > yet, you'll just get the current HEAD state of the submodule.
> 
> That would make it impossible to atomically commit a change that affects two
> submodules, yes?
No. Quite the reverse. What you do is:
 (a) commit both submodules INDEPENDENTLY.
 (b) then commit the supermodule that contains the submodules.

And note how the important part here is that committing in a submodule DOES NOT AFFECT THE SUPERMODULE AT ALL!

The git trees are _independent_. That's important. You should _not_ try to mix them up and make a commit in one commit anything AT ALL in some other tree, exctly because it gets impossible to do (a) interesting things and (b) atomic commits otherwise.

Note that this is true also in the case of a submodule that itself contains a submodule. That doesn't change anything - you still need to be able to view _each_ layer as an independent thing.

Yann Dirson· Nov 25, 2006, 23:49 UTC · re: Linus Torvalds · lore
On Sat, Nov 25, 2006 at 11:30:47AM -0800, Linus Torvalds wrote:
> The git trees are _independent_. That's important.

I'm not sure how independant you mean them to be. The approach I've tried to describe so far assumes that, although you can look at each submodule independently from the supermodule or any other submodule, you can still look at the supermodule as a single tree of it own.

Eg, so that if one part of an appliance/ modules ends up promoted to a lib/ module, GIT can still show that as a move within the supermodule. If we insist that the submodules get committed independently before we make a supermodule commit tying those together, I fear it may make things like such "move/copy detection" more tricky ?

Also, I'd rather expect "git-commit -a" outside of any submodule to commit everything in the supermodule, triggering submodule commits as an intermediate step when needed - just like "git-commit -a" does not require to manually specify subdirectories to inclue in the commit. I'd rather expect a special flag to exclude submodules from a commit.

Best regards, --

Sven Verdoolaege· Nov 26, 2006, 01:14 UTC · re: Yann Dirson · lore
FWIW, here's my view on this issue.
On Sun, Nov 26, 2006 at 12:49:08AM +0100, Yann Dirson wrote:
Show 5 quoted lines
> Also, I'd rather expect "git-commit -a" outside of any submodule to
> commit everything in the supermodule, triggering submodule commits as an
> intermediate step when needed - just like "git-commit -a" does not
> require to manually specify subdirectories to inclue in the commit.  I'd
> rather expect a special flag to exclude submodules from a commit.

A commit should record the content changes that have been made, not change any content itself. Some VCSs change the contents of a file when you commit them (e.g., keyword substitution). Git, rightly, doesn't do that. Likewise, when you commit in the superproject, it should simply record the changes to the "content" of the subproject and not change it. And the content of the subproject is a commit, so a commit in the superproject should not change the content of the subproject by creating another commit in the subproject.

Yann Dirson· Nov 26, 2006, 01:32 UTC · re: Sven Verdoolaege · lore
On Sun, Nov 26, 2006 at 02:14:20AM +0100, Sven Verdoolaege wrote:
Show 5 quoted lines
> Likewise, when you commit in the superproject, it should simply record
> the changes to the "content" of the subproject and not change it.
> And the content of the subproject is a commit, so a commit in the
> superproject should not change the content of the subproject by creating
> another commit in the subproject.

I've realized after suggesting that how much that idea was inadequate - sorry for the noise.

However, I'm not yet buying the idea that "the content of the subproject is a commit" :)

Best regards,
Linus Torvalds· Nov 26, 2006, 03:39 UTC · re: Yann Dirson · lore
On Sun, 26 Nov 2006, Yann Dirson wrote:
Show 6 quoted lines
> 
> Also, I'd rather expect "git-commit -a" outside of any submodule to
> commit everything in the supermodule, triggering submodule commits as an
> intermediate step when needed - just like "git-commit -a" does not
> require to manually specify subdirectories to inclue in the commit.  I'd
> rather expect a special flag to exclude submodules from a commit.

So, how do you do commit messages? It generally doesn't make sense to share the same commit message for submodules - the sub-commits generally do different things.

I'd actually suggest that "git commit -a" with non-clean submodules error out for that reason, with something like

	submodule 'src/xyzzy' is not up-to-date, please commit changes to 
	that first.

exactly because you really generally should consider the submodule commits to be a separate phase.

Daniel Barkalow· Nov 26, 2006, 08:05 UTC · re: Linus Torvalds · lore
On Sat, 25 Nov 2006, Linus Torvalds wrote:
Show 11 quoted lines
> On Sun, 26 Nov 2006, Yann Dirson wrote:
> > 
> > Also, I'd rather expect "git-commit -a" outside of any submodule to
> > commit everything in the supermodule, triggering submodule commits as an
> > intermediate step when needed - just like "git-commit -a" does not
> > require to manually specify subdirectories to inclue in the commit.  I'd
> > rather expect a special flag to exclude submodules from a commit.
> 
> So, how do you do commit messages? It generally doesn't make sense to 
> share the same commit message for submodules - the sub-commits generally 
> do different things.

The same way you do the first commit message. Ask independantly for each commit message in sequence with enough context in the comment section that you know what you're talking about.

Show 8 quoted lines
> I'd actually suggest that "git commit -a" with non-clean submodules error 
> out for that reason, with something like
> 
> 	submodule 'src/xyzzy' is not up-to-date, please commit changes to 
> 	that first.
> 
> exactly because you really generally should consider the submodule commits 
> to be a separate phase.

I think this is getting close to the classic usability blunder of having the program tell you what you should have done instead of what you did, and then making you do it yourself, rather than just doing it.

Just have it run "git commit -a" in each dirty submodule recursively as part of preparing the index, since that's what the user wants to do anyway, and nothing already done would be affected.

"git commit -a -m <message>" should probably fail, of course.
	-Daniel
Andreas Ericsson· Nov 28, 2006, 09:36 UTC · re: Daniel Barkalow · lore
Daniel Barkalow wrote:
Show 33 quoted lines
> On Sat, 25 Nov 2006, Linus Torvalds wrote:
> 
>> On Sun, 26 Nov 2006, Yann Dirson wrote:
>>> Also, I'd rather expect "git-commit -a" outside of any submodule to
>>> commit everything in the supermodule, triggering submodule commits as an
>>> intermediate step when needed - just like "git-commit -a" does not
>>> require to manually specify subdirectories to inclue in the commit.  I'd
>>> rather expect a special flag to exclude submodules from a commit.
>> So, how do you do commit messages? It generally doesn't make sense to 
>> share the same commit message for submodules - the sub-commits generally 
>> do different things.
> 
> The same way you do the first commit message. Ask independantly for each 
> commit message in sequence with enough context in the comment section that 
> you know what you're talking about.
> 
>> I'd actually suggest that "git commit -a" with non-clean submodules error 
>> out for that reason, with something like
>>
>> 	submodule 'src/xyzzy' is not up-to-date, please commit changes to 
>> 	that first.
>>
>> exactly because you really generally should consider the submodule commits 
>> to be a separate phase.
> 
> I think this is getting close to the classic usability blunder of having 
> the program tell you what you should have done instead of what you did, 
> and then making you do it yourself, rather than just doing it.
> 
> Just have it run "git commit -a" in each dirty submodule recursively as 
> part of preparing the index, since that's what the user wants to do 
> anyway, and nothing already done would be affected.
> 

Running "commit -a" is definitely the wrong thing to do, as it prevents one from using the index at all. Erroring out if the submodules are dirty, or just accepting the fact that they are and taking whatever commit HEAD points to is *always* preferrable.

I'd actually prefer the second solution here and let git print a list of submodules with dirty state and ask for some sort of user-response before creating the actual commit. As non-interactive commits should always be clean, requiring user intervention on non-clean state should be a safe thing to do.

> "git commit -a -m <message>" should probably fail, of course.
> 

Why? There's no reason to rob this command of its power just because we're using submodules.

Show 7 quoted lines
> 	-Daniel
> *This .sig left intentionally blank*
> -
> To unsubscribe from this list: send the line "unsubscribe git" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html
> 
-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Daniel Barkalow· Nov 28, 2006, 17:28 UTC · re: Andreas Ericsson · lore
On Tue, 28 Nov 2006, Andreas Ericsson wrote:
Show 14 quoted lines
> Daniel Barkalow wrote:
> > On Sat, 25 Nov 2006, Linus Torvalds wrote:
> > > I'd actually suggest that "git commit -a" with non-clean submodules error
> > > out for that reason
> > 
> > Just have it run "git commit -a" in each dirty submodule recursively as part
> > of preparing the index, since that's what the user wants to do anyway, and
> > nothing already done would be affected.
> > 
> 
> Running "commit -a" is definitely the wrong thing to do, as it prevents one
> from using the index at all. Erroring out if the submodules are dirty, or just
> accepting the fact that they are and taking whatever commit HEAD points to is
> *always* preferrable.

I don't think anyone would actually use the index in submodules but not in the supermodule. If submodules are seen mostly as ordinary directories as far as the supermodule's working directory is concerned, it wouldn't make sense to not commit dirty state in a subdirectory with -a just because it's a submodule.

It would be wrong to do "commit -a" in submodules if the supermodule weren't being committed with -a, of course.

Show 5 quoted lines
> > "git commit -a -m <message>" should probably fail, of course.
> > 
> 
> Why? There's no reason to rob this command of its power just because we're
> using submodules.

It should fail if there are dirty submodules, because the user needs to provide a commit message for each of them, and only one commit message can be provided this way, and -m inhibits invoking an editor.

	-Daniel
Sven Verdoolaege· Nov 28, 2006, 18:08 UTC · re: Daniel Barkalow · lore
On Tue, Nov 28, 2006 at 12:28:47PM -0500, Daniel Barkalow wrote:
> It would be wrong to do "commit -a" in submodules if the supermodule 
> weren't being committed with -a, of course.

What if you say "git commit submodule" ? I sure hope you wouldn't want to do a "commit -a" in the submodule. One of the nice features of git is that you can still perform most operations if you have a dirty state and I would very much want to be able to commit only some changes in the submodule and then only commit that change in submodule commits in the supermodule without having my other changes in the submodule committed as well.

If you agree with the above, then why should "git commit -a" do any different from "git commit submodule" if submodule was the only thing that got changed ?

Daniel Barkalow· Nov 28, 2006, 18:37 UTC · re: Sven Verdoolaege · lore
On Tue, 28 Nov 2006, Sven Verdoolaege wrote:
Show 5 quoted lines
> On Tue, Nov 28, 2006 at 12:28:47PM -0500, Daniel Barkalow wrote:
> > It would be wrong to do "commit -a" in submodules if the supermodule 
> > weren't being committed with -a, of course.
> 
> What if you say "git commit submodule" ?
Obviously no -a, as I said.
> If you agree with the above, then why should "git commit -a"
> do any different from "git commit submodule" if submodule was
> the only thing that got changed ?

If submodule was the only thing that got changed, it's not dirty; if it were dirty, some of its contents would also have gotten changed. Surely:

"git commit submodule/foo bar"

should do "git commit foo" in submodule, and then commit the supermodule with the new commit for the submodule and the change to bar. And so "submodule/foo" is something you could commit changes to, so it should get picked up by -a.

Of course, if submodule *is* the *only* thing that changed (e.g., you did a fast-forward merge in it, or you've previously committed it completely), there won't be a "commit -a" in it, because that would just generate a gratuitous commit.

	-Daniel
Sven Verdoolaege· Nov 28, 2006, 19:06 UTC · re: Daniel Barkalow · lore
On Tue, Nov 28, 2006 at 01:37:54PM -0500, Daniel Barkalow wrote:
> If submodule was the only thing that got changed, it's not dirty; if it 
> were dirty, some of its contents would also have gotten changed.

For me, the commit is the only "content" of the subproject that the superproject should care about, so the submodule being dirty or not is completely irrelevant (for committing), but it seems you see the subproject more as a (working) tree than as a commit. Of course, as Linus already mentioned, a "git commit" could still warn you if the subproject was dirty.

> Surely:
> 
> "git commit submodule/foo bar"

I wouldn't dream of doing such an operation, because it doesn't make sense to me. (So as far as I'm concerned, you can make it do whatever you'd like it to do.) You can only commit the subproject as a whole.

> should do "git commit foo" in submodule, and then commit the supermodule 
> with the new commit for the submodule and the change to bar. And so
> "submodule/foo" is something you could commit changes to, so it should get 
> picked up by -a.
Daniel Barkalow· Nov 28, 2006, 20:41 UTC · re: Sven Verdoolaege · lore
On Tue, 28 Nov 2006, Sven Verdoolaege wrote:
Show 8 quoted lines
> On Tue, Nov 28, 2006 at 01:37:54PM -0500, Daniel Barkalow wrote:
> > If submodule was the only thing that got changed, it's not dirty; if it 
> > were dirty, some of its contents would also have gotten changed.
> 
> For me, the commit is the only "content" of the subproject that the
> superproject should care about, so the submodule being dirty or not
> is completely irrelevant (for committing), but it seems you see the
> subproject more as a (working) tree than as a commit.
I think we agree on the tree/commit/object database model part.

I think we disagree on how the working *directories* relate. I see the checked-out state of a submodule as being relevant to the checked-out state of the supermodule, such that dirty state in the submodule directory is dirty state in the supermodule directory.

Show 7 quoted lines
> > Surely:
> > 
> > "git commit submodule/foo bar"
> 
> I wouldn't dream of doing such an operation, because it doesn't make
> sense to me.  (So as far as I'm concerned, you can make it do whatever
> you'd like it to do.)  You can only commit the subproject as a whole.

I'm thinking that users of subprojects will often want to work on the subprojects rather than exclusively using commits prepared by other people, and it's too much trouble to have to do the work in a repository for just the subproject and pull it into the superproject's submodule to test it. So the submodule working directory needs to function as a working directory for the subproject. Then

  "cd submodule; git commit foo"
does the obvious thing, but that should be the same as
  "git commit submodule/foo" (since it normally is)

and then it makes sense to let you do multiple commits with a single command when the paths end in different modules, since that's obviously what you're requesting, and then -a must do all of them.

	-Daniel
Shawn Pearce· Nov 28, 2006, 21:10 UTC · re: Daniel Barkalow · lore
Daniel Barkalow <barkalow@iabervon.org> wrote:
Show 9 quoted lines
>   "cd submodule; git commit foo"
> 
> does the obvious thing, but that should be the same as
> 
>   "git commit submodule/foo" (since it normally is)
> 
> and then it makes sense to let you do multiple commits with a single 
> command when the paths end in different modules, since that's obviously 
> what you're requesting, and then -a must do all of them.

Except what if the submodules have different commit message standards? E.g. one requires signoff and another doesn't? Or one allows privately held information (e.g. its your coporate project) and one doesn't (e.g. its an open source project you use/contribute to)?

But slightly more practical: the change message for the superproject might simply be "resolved bug X, caused by ...". Which may make a lot of sense to the top level project, but makes no sense at all in a submodule involved in the fix as the submodule's developer community doesn't even know what "X" is, let alone how "..." could have caused it.

So you really need to think twice before you apply the same commit message to every project, as each commit message needs to make sense with that one submodule's limited scope, or within the supermodule's larger scope.

But if you really still think that the same commit message makes sense everywhere, we have 'git commit -F'. Write it out in a file and hand it off to -F in each module. This would be easier if git-ls-files grew a new option:

	vi ~/msg
	for m in $(git ls-files --submodules); do git commit -F ~/msg; done
	git commit -F ~/msg
Daniel Barkalow· Nov 28, 2006, 21:32 UTC · re: Shawn Pearce · lore
On Tue, 28 Nov 2006, Shawn Pearce wrote:
Show 10 quoted lines
> Daniel Barkalow <barkalow@iabervon.org> wrote:
> > and then it makes sense to let you do multiple commits with a single 
> > command when the paths end in different modules, since that's obviously 
> > what you're requesting, and then -a must do all of them.
> 
> Except what if the submodules have different commit message
> standards?  E.g. one requires signoff and another doesn't?  Or one
> allows privately held information (e.g. its your coporate project)
> and one doesn't (e.g. its an open source project you use/contribute
> to)?

I don't think you'd ever want the same commit message for commits in two projects. In any case where you'd commit a submodule in the process of committing a supermodule, git would do this by recursively calling git-commit, which would prompt for separate commit messages.

	-Daniel
Linus Torvalds· Nov 28, 2006, 21:53 UTC · re: Daniel Barkalow · lore
On Tue, 28 Nov 2006, Daniel Barkalow wrote:
> 
> I don't think you'd ever want the same commit message for commits in two 
> projects.

I don't know about "ever", but yes, I do think submodule commits are generally totally separate things from supermodule commits.

> In any case where you'd commit a submodule in the process of 
> committing a supermodule, git would do this by recursively calling 
> git-commit, which would prompt for separate commit messages.

That certainly works, although I'm not convinved that it's necessarily a hugely important detail.

I suspect there may well be more important things UI-wise wrt submodules than the "you may have to commit submodules separately" question.

For example, doing a "git pull" is a lot more interesting, since that actually has the potential of having to resolve conflicts in submodules before the supermodule can be committed. Getting all the "git reset" behaviour right for when you decide "oops, that was too complicated" is probably a lot more important than whether you have to have a separate "commit subproject" phase for the simple cases of doing a bog-standard "git commit -a".

Stephan Feder· Nov 30, 2006, 14:00 UTC · lore
Andy Parkins wrote:
Show 13 quoted lines
> On Thursday 2006 November 30 11:57, sf wrote:
> 
>> > Worse, if you allow that to happen, the supermodule can commit a state
>> > that cannot be retrieved from the submodule's repository.  The ONLY thing
>> > a supermodule can record about a submodule is a commit.
>>
>> So what? You have a submodule commit that only exists in the
>> supermodule. I fail to see the problem. The changes you made to the
>> submodule _in the supermodule_ can later be pulled from wherever you want.
> 
> Eh?  The files aren't stored in the supermodule, they're stored in the 
> submodule.  The ONLY thing in the supermodule is the commit hash.  The 
> objects for the submodule are still /in/ the submodule.

But you have got the submodule on your local disk anyway. So just setup alternates and the supermodule contains all of the submodule.

> It sounds like you're suggesting that the supermodule commit includes files 
> from the submodule?  How can that work?   The two aren't separate entities 
> then, it's just one big repository. 

It works as it always works in git: The supermodule commit contains the submodule commit, the submodule commit contains the submodule files, so the supermodule contains the submodule (at least the part of the submodule that is visible). It _must_ be one repository but it need not be big (once more, use alternates).

Show 7 quoted lines
> I mean, what would this supermodule commit look like?  Would it include a 
> commit message?  Which module should that commit message be about?  Should 
> the commit's parents be stored?  Which parents, the submodule HEAD or the 
> supermodule HEAD?  Which tree object should it link to?  The one in the 
> submodule doesn't exist, so it'll have to be a freshly made up one for the 
> supermodule - except now you've put submodule paths in the supermodule.  
> Nope.  That's never going to work.

Again I do not see the problem. Probably I have a much simpler picture of submodules: They are just commits in the supermodule's tree. Everything else follows naturally from how git currently behaves.

Of course it works. It is simple, it is the git way.
Am I missing the point?
Regards
Stephan
-- 
b.i.t.
beratungsgesellschaft für informations-technologie mbh
Stephan Feder
elisabethenstr. 62   fon: +49(0)6151/827575
64283 darmstadt      fax: +49(0)6151/827576
mailto:sf@b-i-t.de   www: http://www.b-i-t.de
Andy Parkins· Nov 30, 2006, 14:49 UTC · re: Stephan Feder · lore
On Thursday 2006 November 30 14:00, Stephan Feder wrote:
> Again I do not see the problem. Probably I have a much simpler picture
> of submodules: They are just commits in the supermodule's tree.
> Everything else follows naturally from how git currently behaves.

How are these commits any different from just having one big repository? If some of the development of the submodule is contained in the supermodule then it's not a submodule anymore.

Why bother with all the effort to make a separation between submodule and supermodule and then store the submodule commits in the supermodule. That's not supermodule/submodule git - that's just normal git.

Surely the whole point of having submodule's is so that you can take the submodule away. Let me give you an example. Let's say I have a project that uses the libxcb library (some random project out in the world that uses git). I've arranged it something like this:

myproject (git root)
 |----- src
 |----- doc
 `----- libxcb (git root)

This works fine; with one problem. When I make a commit in myproject, there is no link into the particular snapshot of the libxcb that I used at that moment. If libxcb moves on, and makes incompatible changes, then when I checkout an old version of myproject, it won't compile any more because I'll need to find out which commit of libxcb I used at the time.

Submodules will solve this problem. In the future I'll be able to check out any commit of myproject and it will automatically checkout the right commit from the libxcb repository. Now let's say I'm working away and find a bug in libxcb; I fix it, commit it. That change had better be stored in the libxcb repository, and had better make no reference to the myproject repository. If it doesn't, I'm going to have to pollute the libxcb upstream repository with myproject if I want to share those fixes.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Sven Verdoolaege· Nov 30, 2006, 15:20 UTC · re: Andy Parkins · lore
On Thu, Nov 30, 2006 at 02:49:53PM +0000, Andy Parkins wrote:
> How are these commits any different from just having one big repository?  If 
You can work on the submodule independently.
> some of the development of the submodule is contained in the supermodule then 
> it's not a submodule anymore.
On the contrary, that's exactly what a submodule is supposed to be.
> Why bother with all the effort to make a separation between submodule and 
> supermodule and then store the submodule commits in the supermodule.  That's 
> not supermodule/submodule git - that's just normal git.
[..]
Show 5 quoted lines
> myproject (git root)
>  |----- src
>  |----- doc
>  `----- libxcb (git root)
> 
[..]
> 
> Submodules will solve this problem.  In the future I'll be able to check out 
> any commit of myproject and it will automatically checkout the right commit 
> from the libxcb repository.

How are you going to checkout the right commit of the lixcb repo if you didn't store it in the supermodule ?

Andy Parkins· Nov 30, 2006, 15:30 UTC · re: Sven Verdoolaege · lore
On Thursday 2006 November 30 15:20, Sven Verdoolaege wrote:
> You can work on the submodule independently.
It's not independent if any part of it is in the supermodule.
> > some of the development of the submodule is contained in the supermodule
> > then it's not a submodule anymore.
>
> On the contrary, that's exactly what a submodule is supposed to be.
I don't think so.  I think it's just made some complicated normal repository.
> How are you going to checkout the right commit of the lixcb repo if
> you didn't store it in the supermodule ?

Well, I know what the commit is /that/ was all that was stored. So I (actually supermodule-git does):

cd $DIRECTORY_ASSOCIATED_WITH_SUBMODULE git checkout -f $COMMIT_FROM_SUPERMODULE

Obviously, this is grossly simplified. It also requires that HEAD be allowed to be an arbitrary commit rather than a branch, but that's already been generally agreed upon as a good thing.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Andreas Ericsson· Nov 30, 2006, 15:50 UTC · re: Andy Parkins · lore
Andy Parkins wrote:
Show 12 quoted lines
> On Thursday 2006 November 30 15:20, Sven Verdoolaege wrote:
> 
>> You can work on the submodule independently.
> 
> It's not independent if any part of it is in the supermodule.
> 
>>> some of the development of the submodule is contained in the supermodule
>>> then it's not a submodule anymore.
>> On the contrary, that's exactly what a submodule is supposed to be.
> 
> I don't think so.  I think it's just made some complicated normal repository.
> 

I believe that Andy meant "development history" in his above scentence. Naturally, using the code from the submodule while being capable of developing the submodule separately from the supermodule is what submodules are all about.

Show 13 quoted lines
>> How are you going to checkout the right commit of the lixcb repo if
>> you didn't store it in the supermodule ?
> 
> Well, I know what the commit is /that/ was all that was stored.  So I 
> (actually supermodule-git does):
> 
> cd $DIRECTORY_ASSOCIATED_WITH_SUBMODULE
> git checkout -f $COMMIT_FROM_SUPERMODULE
> 
> Obviously, this is grossly simplified.  It also requires that HEAD be allowed 
> to be an arbitrary commit rather than a branch, but that's already been 
> generally agreed upon as a good thing.
> 
It has? We're not talking supermodule specific things anymore, are we?
-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Andy Parkins· Nov 30, 2006, 16:08 UTC · re: Andreas Ericsson · lore
On Thursday 2006 November 30 15:50, Andreas Ericsson wrote:
Show 5 quoted lines
> > Obviously, this is grossly simplified.  It also requires that HEAD be
> > allowed to be an arbitrary commit rather than a branch, but that's
> > already been generally agreed upon as a good thing.
>
> It has? We're not talking supermodule specific things anymore, are we?

Not entirely, although I think it's going to be handy for submodules. It was in a thread about remotes branches. By allowing checkout of any commit rather than only those that have a ref/heads/ entry, you effectively have a read-only checkout. You obviously couldn't commit to a repository like this, because HEAD wouldn't point at anything that is changeable. It would be very easy to just git-branch from there and start work though.

I think it's going to be necessary for the submodule work, because without it the supermodule will have to create it's own temporary branches in the submodule in order to checkout an arbitrary commit.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Sven Verdoolaege· Nov 30, 2006, 16:33 UTC · re: Andy Parkins · lore
On Thu, Nov 30, 2006 at 03:30:49PM +0000, Andy Parkins wrote:
Show 5 quoted lines
> On Thursday 2006 November 30 15:20, Sven Verdoolaege wrote:
> > How are you going to checkout the right commit of the lixcb repo if
> > you didn't store it in the supermodule ?
> 
> Well, I know what the commit is /that/ was all that was stored.  So I 

Then I have no idea what you are talking about. A commit _contains_ all the history that lead up to that commit, so if you have the commit, then you also have the history.

Andy Parkins· Dec 1, 2006, 00:01 UTC · re: Sven Verdoolaege · lore
On Thursday 2006, November 30 16:33, Sven Verdoolaege wrote:
Show 5 quoted lines
> > Well, I know what the commit is /that/ was all that was stored.  So I
>
> Then I have no idea what you are talking about.
> A commit _contains_ all the history that lead up to that commit,
> so if you have the commit, then you also have the history.

It's not so much an actual commit, as a reference to a commit in another repository.

Andy
-- 
Dr Andrew Parkins, M Eng (Hons), AMIEE
Jakub Narebski· Dec 1, 2006, 00:11 UTC · re: Andy Parkins · lore
Andy Parkins wrote:
Show 10 quoted lines
> On Thursday 2006, November 30 16:33, Sven Verdoolaege wrote:
>>>
>>> Well, I know what the commit is /that/ was all that was stored.  So I
>>
>> Then I have no idea what you are talking about.
>> A commit _contains_ all the history that lead up to that commit,
>> so if you have the commit, then you also have the history.
> 
> It's not so much an actual commit, as a reference to a commit in another 
> repository.

Hmmm... I thought the idea was that submodule commit is available in the object repository, be it via alternates mechanism pointing to the submodule repository for alternate storage, or submodule being in "unrelated" branch/tracking branch.

-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Sven Verdoolaege· Dec 1, 2006, 09:32 UTC · re: Andy Parkins · lore
On Fri, Dec 01, 2006 at 12:01:54AM +0000, Andy Parkins wrote:
Show 9 quoted lines
> On Thursday 2006, November 30 16:33, Sven Verdoolaege wrote:
> > > Well, I know what the commit is /that/ was all that was stored.  So I
> >
> > Then I have no idea what you are talking about.
> > A commit _contains_ all the history that lead up to that commit,
> > so if you have the commit, then you also have the history.
> 
> It's not so much an actual commit, as a reference to a commit in another 
> repository.

This is heresy. Any object referenced in a tree should be in the repo (possibly via alternates).

Andy Parkins· Dec 1, 2006, 10:19 UTC · re: Sven Verdoolaege · lore
On Friday 2006 December 01 09:32, Sven Verdoolaege wrote:
> This is heresy.  Any object referenced in a tree should be in the repo
> (possibly via alternates).

The "submodule" object would be in the local repository. That would refer to another object, and is merely part of the submodule object. Just as the "Author" and "Commiter" fields are part of the commit object but aren't actual objects in the tree.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Martin Waitz· Nov 30, 2006, 17:19 UTC · re: Andy Parkins · lore
hoi :)
On Thu, Nov 30, 2006 at 03:30:49PM +0000, Andy Parkins wrote:
Show 9 quoted lines
> Well, I know what the commit is /that/ was all that was stored.  So I 
> (actually supermodule-git does):
> 
> cd $DIRECTORY_ASSOCIATED_WITH_SUBMODULE
> git checkout -f $COMMIT_FROM_SUPERMODULE
> 
> Obviously, this is grossly simplified.  It also requires that HEAD be allowed 
> to be an arbitrary commit rather than a branch, but that's already been 
> generally agreed upon as a good thing.
It's not that easy.

You also have to make sure that all your submodule commits that _ever_ have been part of your submodule have be stay in your repository forever. Consider that your submodule switches to an other branch and some old commits are not referenced by the current version any more. These old commits still have to survive a git-prune, if they have been part of some old supermodule version. So you really have to connect both object databases and it's not enough to just store the commit sha1 without actually parsing it by the GIT core.

-- 
Martin Waitz
sf· Nov 30, 2006, 16:05 UTC · re: Andy Parkins · lore
Andy Parkins wrote:
Show 9 quoted lines
> On Thursday 2006 November 30 14:00, Stephan Feder wrote:
> 
>> Again I do not see the problem. Probably I have a much simpler picture
>> of submodules: They are just commits in the supermodule's tree.
>> Everything else follows naturally from how git currently behaves.
> 
> How are these commits any different from just having one big repository?  If 
> some of the development of the submodule is contained in the supermodule then 
> it's not a submodule anymore.

Right now you only have commits of the top directory aka the super project. Every subdirectory is just that: a directory (which git stores as trees).

Now, if you have a subdirectory that git stores as a commit, not a tree, you have a subproject. It is a directory with history, and because the commit is part of your superprject, you have access to this history.

> Why bother with all the effort to make a separation between submodule and 
> supermodule and then store the submodule commits in the supermodule.  That's 
> not supermodule/submodule git - that's just normal git.

No, it is not. Currently, there is no way to store a commit within the contents of another commit. You can only store trees and blobs.

Show 15 quoted lines
> Surely the whole point of having submodule's is so that you can take the 
> submodule away.  Let me give you an example.  Let's say I have a project that 
> uses the libxcb library (some random project out in the world that uses git).  
> I've arranged it something like this:
> 
> myproject (git root)
>  |----- src
>  |----- doc
>  `----- libxcb (git root)
> 
> This works fine; with one problem.  When I make a commit in myproject, there 
> is no link into the particular snapshot of the libxcb that I used at that 
> moment.  If libxcb moves on, and makes incompatible changes, then when I 
> checkout an old version of myproject, it won't compile any more because I'll 
> need to find out which commit of libxcb I used at the time.
OK.
> Submodules will solve this problem.  In the future I'll be able to check out 
> any commit of myproject and it will automatically checkout the right commit 
> from the libxcb repository.
OK, I am still with you so far.
Show 5 quoted lines
> Now let's say I'm working away and find a bug in 
> libxcb; I fix it, commit it.  That change had better be stored in the libxcb 
> repository, and had better make no reference to the myproject repository.  If 
> it doesn't, I'm going to have to pollute the libxcb upstream repository with 
> myproject if I want to share those fixes.
Here comes the part where we did not meet before.

Of course you do not make any reference from your subproject to your superproject. You do exactly what you do in git today when you work with different branches:

Step 1: You fix a bug in myproject's subdirectory libxcb.

Step 2: You commit to myproject. myproject now contains a new commit object in path libxcb. (How to do that is up to the UI but at the repository level the outcome should be obvious). This commit is local to your repository.

Step 3: You propose your changes to the libxcb upstream (it might not be a repository you have write access to). I use the following made up syntax (see man git-rev-parse):

A suffix : followed by a path, _followed by a suffix //::_ names the _revision_ at the given path in the tree-ish object named by the part before the colon.

Step 3a: Generate a patch
git diff libxcb//^..libxcb//
Step 3b: Push your changes
git push <libxcb-repository> HEAD:libxcb//:<branch in libxcb-repository>
Step 3c: Let your changes be pulled

"Hello, please pull <myproject-repository> HEAD:libxcb//:<branch in libxcb-repository>"

Step 4: Pull upstream version (hopefully with your changes, otherwise you have to merge)

git pull <libxcb-repository> <branch in libxcb-repository>::HEAD:libxcb//
See, it works.
 From what I understand you want to do the commit and push steps in one 
go. How do you want to record local (to your superproject) changes to 
the subproject?
Regards
sf· Nov 30, 2006, 16:12 UTC · re: sf · lore

sf wrote: ...

> A suffix : followed by a path, _followed by a suffix //::_ names the 
> _revision_ at the given path in the tree-ish object named by the part 
> before the colon.
Sorry, that was supposed to read: followed by a suffix //
Andy Parkins· Dec 1, 2006, 09:19 UTC · re: sf · lore
On Thursday 2006 November 30 16:05, sf wrote:
> Step 2: You commit to myproject. myproject now contains a new commit
> object in path libxcb. (How to do that is up to the UI but at the
> repository level the outcome should be obvious). This commit is local to
> your repository.

Let's imagine a supermodule repository, and guess at it in more detail (I'll abbreviate some of the less interesting output):

$ git-cat-file -p HEAD tree fb02e78085ecf2f29045603df858b5362e5bf8a4 parent 4f2dba685507e4a8e07dac298c4024feaec6bd7d author Andy Parkins committer Andy Parkins $ git-cat-file -p fb02e78085ecf2f29045603df858b5362e5bf8a4 100644 blob 46bd4e284a57e2faa539e7b72d62a38867075af5 Makefile 040000 tree 49ea01373a986a3db44d66702714aa75059ffa2c doc 040000 subm d0a877464dc0198667a3e27ed3af8448ddacf947 libxcb

The "subm" type is our new ODB object that's going to store whatever we will need to access the submodule. "libxcb" has already told us where this submodule is in the supermodule tree.

$ git-cat-file -p d0a877464dc0198667a3e27ed3af8448ddacf947 submodulecommithash ccddf1d4b0cf7fd3a699d8b33cf5bc4c5c4435b7 submoduleurlhint git://anongit.freedesktop.org/git/xcb/libxcb

Here "submodulecommithash" is telling us what commit in the submodule is stored in this supermodule tree. The "submoduleurlhint" is to help when git-clone is used to clone this supermodule.

They key thing I wanted to point out here is the line:
  submodulecommithash ccddf1d4b0cf7fd3a699d8b33cf5bc4c5c4435b7
This is the ONLY link you have to the submodule.  I think this line represents 
the fundamental difference between our thinking on submodules.
I say:
 submodulecommithash points at a commit /in the submodule/
You say:
 "This commit is local to your repository".  i.e. it points at a commit in
 the supermodule, which in turn implies that the local commit object points
 at a local tree and local parents.

My question is therefore: tell me what that local commit's tree and parent's are? At the moment I am having difficulty understanding what meaningful things you could have in those fields.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Martin Waitz· Dec 1, 2006, 09:57 UTC · re: Andy Parkins · lore
hoi :)
On Fri, Dec 01, 2006 at 09:19:04AM +0000, Andy Parkins wrote:
Show 12 quoted lines
> Let's imagine a supermodule repository, and guess at it in more detail (I'll 
> abbreviate some of the less interesting output):
> 
> $ git-cat-file -p HEAD
> tree fb02e78085ecf2f29045603df858b5362e5bf8a4
> parent 4f2dba685507e4a8e07dac298c4024feaec6bd7d
> author Andy Parkins
> committer Andy Parkins 
> $ git-cat-file -p fb02e78085ecf2f29045603df858b5362e5bf8a4
> 100644 blob 46bd4e284a57e2faa539e7b72d62a38867075af5    Makefile
> 040000 tree 49ea01373a986a3db44d66702714aa75059ffa2c    doc
> 040000 subm d0a877464dc0198667a3e27ed3af8448ddacf947    libxcb
at the moment, it is:
  140000 commit ccddf1d4b0cf7fd3a699d8b33cf5bc4c5c4435b7  libxcb
Show 7 quoted lines
> The "subm" type is our new ODB object that's going to store whatever we will 
> need to access the submodule.  "libxcb" has already told us where this 
> submodule is in the supermodule tree.
> 
> $ git-cat-file -p d0a877464dc0198667a3e27ed3af8448ddacf947
> submodulecommithash ccddf1d4b0cf7fd3a699d8b33cf5bc4c5c4435b7
> submoduleurlhint git://anongit.freedesktop.org/git/xcb/libxcb
So why do you need the url hint committed to the supermodule?
We don't store remote information in the object database, too.
Remember: this is still a distributed project, there is no one URL to
any submodule.
> I say:
>  submodulecommithash points at a commit /in the submodule/

But unluckily, this does not work. You really have to be able to traverse the entire commit chain from the supermodule into all submodules.

-- 
Martin Waitz
Andy Parkins· Dec 1, 2006, 10:29 UTC · re: Martin Waitz · lore
On Friday 2006 December 01 09:57, Martin Waitz wrote:
> So why do you need the url hint committed to the supermodule?
> We don't store remote information in the object database, too.

That's why it was a hint, probably configured when you first create the submodule connection.

> Remember: this is still a distributed project, there is no one URL to
> any submodule.

That point applies equally to your "tracking a submodule branch" point, except mine is only a URL hint, to help when first cloning that supermodule. In truth, the clone will be perfectly able to get the submodule objects from the upstream supermodule, maintaining the distributed nature easily.

> > I say:
> >  submodulecommithash points at a commit /in the submodule/
>
> But unluckily, this does not work.

Eh? "Not work", we're talking about code that doesn't even exist, of course it doesn't "work". Do you mean "doesn't work if we're using my implementation of submodules"? Well that hardly seems like a fair attack.

> You really have to be able to traverse the entire commit chain
> from the supermodule into all submodules.

You can: when you hit a submodule tree object you set GIT_DIR to that submodule and continue. If you don't do it like that then you have stored submodule trees in the supermodule and it's no longer a separate repository. Why you'd want to - I have no idea. What purpose would you have for traversing the commit chain into the submodules? The commit in the submodule is just a note of where that submodule was during the supermodule commit in question.

I notice though that you avoided my question: what does YOUR submodule object contain? I really do want to know, as there is obviously a fundamental difference in what I think a submodule does and what you (and maybe everybody else) thinks a submodule does. I'm perfectly willing to accept I'm wrong, but not without understanding how your method is going to work.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Sven Verdoolaege· Dec 1, 2006, 10:42 UTC · re: Andy Parkins · lore
On Fri, Dec 01, 2006 at 10:29:26AM +0000, Andy Parkins wrote:
> I notice though that you avoided my question: what does YOUR submodule object 
> contain?

He showed it to you in the example. The "submodule object" is the COMMIT of the submodule itself.

Andy Parkins· Dec 1, 2006, 11:02 UTC · re: Sven Verdoolaege · lore
On Friday 2006 December 01 10:42, Sven Verdoolaege wrote:
> He showed it to you in the example.  The "submodule object" is the COMMIT
> of the submodule itself.
That's no different from mine.  I need more detail than that.

Is that commit in the submodule or the supermodule? If it's in the submodule then we're talking about the same thing, as that's all I want. If it's in the supermodule then I want to know what the tree object that that commit points to contains. I also want to know how we tell the difference between a commit-in-supermodule and a commit-in-supermodule-which-is-actually-in-submodule.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Sven Verdoolaege· Dec 1, 2006, 11:10 UTC · re: Andy Parkins · lore
On Fri, Dec 01, 2006 at 11:02:15AM +0000, Andy Parkins wrote:
Show 6 quoted lines
> On Friday 2006 December 01 10:42, Sven Verdoolaege wrote:
> 
> > He showed it to you in the example.  The "submodule object" is the COMMIT
> > of the submodule itself.
> 
> That's no different from mine.  I need more detail than that.

You were proposing to create an extra object containing some random value that is disconnected from the repo.

> Is that commit in the submodule or the supermodule?
It's in BOTH.  That's why it's a *sub*module.
Someone else can try to expain it you.
sf· Dec 1, 2006, 11:45 UTC · re: Sven Verdoolaege · lore
Sven Verdoolaege wrote:
Show 14 quoted lines
> On Fri, Dec 01, 2006 at 11:02:15AM +0000, Andy Parkins wrote:
>> On Friday 2006 December 01 10:42, Sven Verdoolaege wrote:
>> 
>> > He showed it to you in the example.  The "submodule object" is the COMMIT
>> > of the submodule itself.
>> 
>> That's no different from mine.  I need more detail than that.
> 
> You were proposing to create an extra object containing some random value
> that is disconnected from the repo.
> 
>> Is that commit in the submodule or the supermodule?
> 
> It's in BOTH.  That's why it's a *sub*module.

I would say it is only in the supermodule because that is the branch you are working on. If you are working on the submodule in an independent branch then you can pull from the submodule commit. But you do not want to pull the supermodule commit itself but only the commit in path libxcb (see my proposed syntax).

Regards
Stephan
Andy Parkins· Dec 1, 2006, 12:12 UTC · re: Sven Verdoolaege · lore
On Friday 2006 December 01 11:10, Sven Verdoolaege wrote:
> You were proposing to create an extra object containing some random value
> that is disconnected from the repo.

Right, I think I've finally understood what Martin (and you) are proposing. You want every commit in the submodule to be propagated up to the supermodule as well. Okay.

I don't think it's right, but at least I understand.

It seems wrong because it's making commits in the supermodule that aren't commits to do with that project. In my libxcb example; why should every project use libxcb in have to store the entire history of libxcb? When examining the supermodule history, I won't care about how libxcb got to the state its in, and it's just noise in the supermodule history. What if I use 10 submodules, the supermodule history won't show you anything useful - it's just unrelated submodule commits.

It gets worse, this is why I was asking for more detail: this commit that you're storing in the supermodule. It's the same commit as is in the submodule? What would the parent commit of that commit be? It has to be the same in both, because the commit-hash forces it to be.

The only possibility would be that it's NOT the same hash in both, because the parents in the supermodule are inapplicable to the submodule, and the parent in the submodule is independent from the supermodule. That means you have to store two commits: one for the submodule commit and one for the supermodule commit. So what are you going to write in the supermodule commit? Answer: a submodule commit hash - exactly as I said.

> > Is that commit in the submodule or the supermodule?
>
> It's in BOTH.  That's why it's a *sub*module.

If it's in BOTH then the supermodule is a normal git repository. You aren't tracking the submodule, you're just including it en masse. Using semantics to justify a position isn't a very strong argument, calling it a "sub" module is just an easy bit of naming for us to hang the discussion on, it isn't necessarily a mathematical subset and superset.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Martin Waitz· Dec 1, 2006, 12:28 UTC · re: Andy Parkins · lore
hoi :)
On Fri, Dec 01, 2006 at 12:12:34PM +0000, Andy Parkins wrote:
Show 10 quoted lines
> On Friday 2006 December 01 11:10, Sven Verdoolaege wrote:
> 
> > You were proposing to create an extra object containing some random value
> > that is disconnected from the repo.
> 
> Right, I think I've finally understood what Martin (and you) are
> proposing.  You want every commit in the submodule to be propagated up
> to the supermodule as well.  Okay.
> 
> I don't think it's right, but at least I understand.

Please note that the submodule commits are not part of the supermodule commit chain, they are part of the supermodule _tree_.

> It seems wrong because it's making commits in the supermodule that aren't 
> commits to do with that project.

Of course they are part of your project, just like all the tree and blob objects, too.

> In my libxcb example; why should every project use libxcb in have to
> store the entire history of libxcb?

Because you want to be able to use the submodule as a repository of its own, too. Be able to look at its history if you want to. Be able to merge with new versions of the submodule. This is what distiguishes a submodule from a pure file-based import of another project.

> When examining the supermodule history, I won't care about how libxcb
> got to the state its in, and it's just noise in the supermodule
> history.  What if I use 10 submodules, the supermodule history won't
> show you anything useful - it's just unrelated submodule commits.
Again: the submodules are part of your supermodule _tree_, not it's
commit chain.  So you won't see the submodule commits when you invoke
git-log in the supermodule.
> It gets worse, this is why I was asking for more detail: this commit
> that you're storing in the supermodule.  It's the same commit as is in
> the submodule?
It is _the_ commit from the submodule, yes.
> What would the parent commit of that commit be?  It has to be the same
> in both, because the commit-hash forces it to be.

It is the commit of the submodule, so its parents point to the submodule history.

Show 6 quoted lines
> > > Is that commit in the submodule or the supermodule?
> >
> > It's in BOTH.  That's why it's a *sub*module.
> 
> If it's in BOTH then the supermodule is a normal git repository.  You aren't 
> tracking the submodule, you're just including it en masse.

The submodule is part of the entire project, so yes, it is included. And the supermodule tracks submodule development by storing references to the submodule history that was used at that time.

Lets try to paint a little diagram:

belongint to: /--------- supermodule -------\ /---- submodule -------\

commit -> tree +-> blob
  |            +-> tree -> ...
  |            +-----------------> commit -> tree -> ...
  v                                  |
commit -> tree +-> ...               v
  |            +-----------------> commit -> ...
  |                                  |
  |                                  v
  |                                commit -> ...
  v                                  |
commit -> tree +-> ...               v
               +-----------------> commit

Both have their independent history, but they are linked as some submodule versions are part of the supermodule tree.

-- 
Martin Waitz
Andy Parkins· Dec 1, 2006, 14:11 UTC · re: Martin Waitz · lore
On Friday 2006 December 01 12:28, Martin Waitz wrote:
Show 5 quoted lines
> > It seems wrong because it's making commits in the supermodule that aren't
> > commits to do with that project.
>
> Of course they are part of your project, just like all the tree and blob
> objects, too.

I wouldn't go as far as that; just because I use libxcb doesn't mean I want it's history merged with mine. However, I think my worries are unfounded, your comment about being able to independently clone the libxcb tree helped me there. If I've understood; while the objects themselves are stored in the supermodule ODB, they are still independent. In fact, they're only in the supermodule tree because it's most convenient to keep them there; it sounds like it's very easy to strip them out again.

> It is the commit of the submodule, so its parents point to the submodule
> history.
Again, if I'm understanding, it's a bit like when you have an additional root 
in a normal git repository, for example:
 
 * -- * -- * -- * (project1)
       \
        * -- * -- * (project1/stable)
             
   * -- * -- * -- * (project2)

Then to make project2 a submodule of project1, one of the project1 trees simply refers to a commit in project2.

I think my original idea for how this works was correct with one minor flaw, and from that flaw all the other concerns flowed. I imagined that there were two object databases - one for the supermodule and one for the submodule. The fault was that there aren't two ODBs there are two roots. Which of course is a far easier way to blend to repositories. Apart from that, I think I'm entirely in sync, and it was merely my wanting to put each of these roots in their own repository that caused all the confusion.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Martin Waitz· Dec 1, 2006, 15:12 UTC · re: Andy Parkins · lore
hoi :)
On Fri, Dec 01, 2006 at 02:11:19PM +0000, Andy Parkins wrote:
> If I've understood; while the objects themselves are stored in the
> supermodule ODB, they are still independent.  In fact, they're only in
> the supermodule tree because it's most convenient to keep them there;
> it sounds like it's very easy to strip them out again.
Yes.
Show 11 quoted lines
> Again, if I'm understanding, it's a bit like when you have an
> additional root in a normal git repository, for example:
>
>  * -- * -- * -- * (project1)
>        \
>         * -- * -- * (project1/stable)
>
>    * -- * -- * -- * (project2)
> 
> Then to make project2 a submodule of project1, one of the project1
> trees simply refers to a commit in project2.
Exactly.
-- 
Martin Waitz
Martin Waitz· Dec 1, 2006, 11:46 UTC · re: Andy Parkins · lore
hoi :)
On Fri, Dec 01, 2006 at 11:02:15AM +0000, Andy Parkins wrote:
Show 6 quoted lines
> On Friday 2006 December 01 10:42, Sven Verdoolaege wrote:
> 
> > He showed it to you in the example.  The "submodule object" is the COMMIT
> > of the submodule itself.
> 
> That's no different from mine.
Well, there simply is no proxy object inbetween.
> Is that commit in the submodule or the supermodule?

Well, logically that commit belongs to the submodule and is referenced by the tree in the supermodule. Phyisically it is stored in the projects object database which is shared between the supermodule and all submodules (at least in my implementation).

> I also want to know how we tell the difference between a
> commit-in-supermodule and a
> commit-in-supermodule-which-is-actually-in-submodule.
There is no difference.
-- 
Martin Waitz
Andy Parkins· Dec 1, 2006, 12:16 UTC · re: Martin Waitz · lore
On Friday 2006 December 01 11:46, Martin Waitz wrote:
> > That's no different from mine.
>
> Well, there simply is no proxy object inbetween.

That's fine, I was only using the proxy object to allow additional information into the submodule object. Actually, I think it would always be better to use a proxy object otherwise you have an error in the tree object, because it will refer to an object that does not exist. The proxy object is allowed to refer to objects that don't exist because it's not a tree object.

Show 7 quoted lines
> > Is that commit in the submodule or the supermodule?
>
> Well, logically that commit belongs to the submodule and is referenced
> by the tree in the supermodule.
> Phyisically it is stored in the projects object database which is
> shared between the supermodule and all submodules (at least in my
> implementation).

Hmmm, "shared"? It must still be in the submodule physically though, and presumably the supermodule uses alternatives to get access to it? Otherwise the submodule will be impossible to separate from the supermodule.

Show 5 quoted lines
> > I also want to know how we tell the difference between a
> > commit-in-supermodule and a
> > commit-in-supermodule-which-is-actually-in-submodule.
>
> There is no difference.

Okay. I think I'm still a bit lost then. I suppose I'll wait for your patches to understand.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Martin Waitz· Dec 1, 2006, 12:34 UTC · re: Andy Parkins · lore
hoi :)
On Fri, Dec 01, 2006 at 12:16:00PM +0000, Andy Parkins wrote:
Show 6 quoted lines
> That's fine, I was only using the proxy object to allow additional
> information into the submodule object.  Actually, I think it would
> always be better to use a proxy object otherwise you have an error in
> the tree object, because it will refer to an object that does not
> exist.  The proxy object is allowed to refer to objects that don't
> exist because it's not a tree object.

It is exactly the aim of my implementation to not have any reference to something that is not accessible in the supermodule repository.

Show 12 quoted lines
> > > Is that commit in the submodule or the supermodule?
> >
> > Well, logically that commit belongs to the submodule and is referenced
> > by the tree in the supermodule.
> > Phyisically it is stored in the projects object database which is
> > shared between the supermodule and all submodules (at least in my
> > implementation).
> 
> Hmmm, "shared"?  It must still be in the submodule physically though,
> and presumably the supermodule uses alternatives to get access to it?
> Otherwise the submodule will be impossible to separate from the
> supermodule.

Yes, you can't separate it my just moving it out of the supermodule, but you can always clone the submodule alone.

> Okay.  I think I'm still a bit lost then.  I suppose I'll wait for your
> patches to understand.

have a look at http://git.admingilde.org/tali/git.git/module2. If you want to try it out, have a look at t/t7500-submodule.sh on how to create submodules.

-- 
Martin Waitz
Andy Parkins· Dec 1, 2006, 13:59 UTC · re: Martin Waitz · lore
On Friday 2006 December 01 12:34, Martin Waitz wrote:
> It is exactly the aim of my implementation to not have any reference to
> something that is not accessible in the supermodule repository.

Okay - I think you've put me right in another reply on this point - the submodule commit is in the supermodule; that was the part I hadn't got.

> Yes, you can't separate it my just moving it out of the supermodule,
> but you can always clone the submodule alone.

Ah - now that clarifies things a lot. The fact that you can't separate it by moving it implies lots of things that take away many of my earlier worries.

> have a look at http://git.admingilde.org/tali/git.git/module2.
> If you want to try it out, have a look at t/t7500-submodule.sh on how to
> create submodules.

Thanks. I will look hard at this :-) My apologies for bothering you so much with all these questions. I just got a bit interested in it all :-)

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Martin Waitz· Dec 1, 2006, 14:07 UTC · re: Andy Parkins · lore
hoi :)
On Fri, Dec 01, 2006 at 01:59:58PM +0000, Andy Parkins wrote:
> My apologies for bothering you so much with all these questions.

You were not bothering me. Those were really interesting and valid questions. In fact, it was a long way for me to come to the implementation I have now. And I really did ask many of those questions to me, too.

I should really write a nice paper about all of that, I think.
> I just got a bit interested in it all :-)
Good :-)
-- 
Martin Waitz
Martin Waitz· Dec 1, 2006, 11:31 UTC · re: Andy Parkins · lore
hoi :)
On Fri, Dec 01, 2006 at 10:29:26AM +0000, Andy Parkins wrote:
Show 11 quoted lines
> On Friday 2006 December 01 09:57, Martin Waitz wrote:
> 
> > So why do you need the url hint committed to the supermodule?
> > We don't store remote information in the object database, too.
> 
> That's why it was a hint, probably configured when you first create the 
> submodule connection.
> 
> In truth, the clone will be perfectly able to get the submodule
> objects from the upstream supermodule, maintaining the distributed
> nature easily.

that's exactly the reason why the hint is not needed. Althogh you need to have one common project object database, storing the objects of all modules.

Show 9 quoted lines
> > > I say:
> > >  submodulecommithash points at a commit /in the submodule/
> >
> > But unluckily, this does not work.
> 
> Eh?  "Not work", we're talking about code that doesn't even exist, of
> course it doesn't "work".   Do you mean "doesn't work if we're using
> my implementation of submodules"?  Well that hardly seems like a fair
> attack.

Well, at first I started exactly as you described: only store the submodule commit sha1 in the parent somewhere, but don't traverse it. So this is a fair attack: your implementation already exists in http://git.admingilde.org/tali/git.git/module ;-) (ok, yes, it really is different to what you described as I stored the sha1 differently, but I really learned that it is important to be able to traverse the entire commit chain, from the root of the project to the deepest submodule.)

Show 7 quoted lines
> > You really have to be able to traverse the entire commit chain
> > from the supermodule into all submodules.
> 
> You can: when you hit a submodule tree object you set GIT_DIR to that
> submodule and continue.  If you don't do it like that then you have
> stored submodule trees in the supermodule and it's no longer a
> separate repository.

Well, a submodule repository _is_ special in some ways: fsck and prune have to take the references from the supermodule into account. In this sense it is _not_ separate from the supermodule.

I think that is important for the submodule repository to be independent in other ways than its object database: you should be able to exchange commits with other repositories (be they stand-alone or a submodule in another supermodule). You should be able to use log/diff/blame/whatever inside the submodule.

All this does not need an object database of its own. So I chose to do it the easy way and use one object database for the entire project - and disallow git-prune in a submodule. There may be other/better ways to do this, but you have to be able to access all objects which belong the project inside the toplevel project repository.

> Why you'd want to - I have no idea.  What
> purpose would you have for traversing the commit chain into the
> submodules?  The commit in the submodule is just a note of where that
> submodule was during the supermodule commit in question.
Things get much simpler if you have one big graph of objects.

clone and especially fetch/pull naturally work at once. You can ask for all objects inside the whole project which are needed to be transferred between project version A and B, including all submodules.

You can even have one bare repository for the whole project.
> I notice though that you avoided my question: what does YOUR submodule
> object contain?  I really do want to know, as there is obviously a
> fundamental difference in what I think a submodule does and what you
> (and maybe everybody else) thinks a submodule does.

It really only stores the commit of the submodule directly. So there is no new submodule object type. The parent has a direct link to the submodule commit in his tree object and in its index. In order to separate them from normal files or normal subdirectories, they get a special mode: they are represented as socket.

-- 
Martin Waitz
Andy Parkins· Dec 1, 2006, 12:20 UTC · re: Martin Waitz · lore
On Friday 2006 December 01 11:31, Martin Waitz wrote:
Show 5 quoted lines
> It really only stores the commit of the submodule directly.
> So there is no new submodule object type.  The parent has a direct link
> to the submodule commit in his tree object and in its index.  In order
> to separate them from normal files or normal subdirectories, they get a
> special mode: they are represented as socket.

Okay. I think I've got it now. I'm not convinced that the way you've chosen is the correct way, primarily because the separation between supermodule and submodule is not strong. Regardless, as you're doing it, you get to pick :-) Is there a public repository I can look at to see what you've done? I'm interested in the sort of plumbing changes needed to make something like this work.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Martin Waitz· Dec 1, 2006, 12:37 UTC · re: Andy Parkins · lore
hoi :)
On Fri, Dec 01, 2006 at 12:20:42PM +0000, Andy Parkins wrote:
> Is there a public repository I can look at to see what you've done?
> I'm interested in the sort of plumbing changes needed to make
> something like this work.
link is in the mail that started this thread ;-).
-- 
Martin Waitz
Jakub Narebski· Dec 2, 2006, 15:16 UTC · re: Martin Waitz · lore
Martin Waitz wrote:
Show 8 quoted lines
> hoi :)
> 
> On Fri, Dec 01, 2006 at 12:20:42PM +0000, Andy Parkins wrote:
>> Is there a public repository I can look at to see what you've done?
>> I'm interested in the sort of plumbing changes needed to make
>> something like this work.
> 
> link is in the mail that started this thread ;-).
And on GitWiki as well:
  http://git.or.cz/gitwiki/SubprojectSupport
-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Andy Parkins· Dec 1, 2006, 08:49 UTC · lore
On Thursday 2006 November 30 18:57, Andreas Ericsson wrote:
(agree with everything in your mail)
> The only problem I'm seeing atm is that the supermodule somehow has to
> mark whatever commits it's using from the submodule inside the submodule
> repo so that they effectively become un-prunable, otherwise the
> supermodule may some day find itself with a history that it can't restore.

What about submodule/.git/refs/supermodule/commit12345678, where "12345678" is the hash of the supermodule commit? This gives a convenient route in the submodule to which commit contains that commit from the submodule; but doesn't write anything into the submodule repository itself. It's just a tag with a different intent.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Andreas Ericsson· Dec 1, 2006, 09:33 UTC · re: Andy Parkins · lore
Andy Parkins wrote:
Show 15 quoted lines
> On Thursday 2006 November 30 18:57, Andreas Ericsson wrote:
> 
> (agree with everything in your mail)
> 
>> The only problem I'm seeing atm is that the supermodule somehow has to
>> mark whatever commits it's using from the submodule inside the submodule
>> repo so that they effectively become un-prunable, otherwise the
>> supermodule may some day find itself with a history that it can't restore.
> 
> What about submodule/.git/refs/supermodule/commit12345678, where "12345678" is 
> the hash of the supermodule commit?  This gives a convenient route in the 
> submodule to which commit contains that commit from the submodule; but 
> doesn't write anything into the submodule repository itself.  It's just a tag 
> with a different intent.
> 

True, but this makes one repo of the submodule special. Let's say you have this layout

mozilla/.git mozilla/openssl/.git mozilla/xlat/.git

Now, we can be reasonably sure that the 'xlat' repo is something the mozilla core team can push to, or at least we can consider the core repo owners an official "vendor" of tags for the submodule repo. I'm fairly certain openssl authors won't be too happy with allowing the thousands of projects using its code to push tags to its official repo though.

Now that I think about it more, I realize this is completely irrelevant as the ui can create the tags in the submodule with info only from the the supermodule, which means the submodule repo will only be special if it's connected to the supermodule. We just need a command for creating those tags in the submodule repo so people who use the same submodule code for several projects can use the alternates mechanism effectively.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Andy Parkins· Dec 1, 2006, 10:38 UTC · re: Andreas Ericsson · lore
On Friday 2006 December 01 09:33, Andreas Ericsson wrote:
> True, but this makes one repo of the submodule special. Let's say you
> have this layout
In a way, but it's information that doesn't need to be transmitted.
Show 9 quoted lines
> mozilla/.git
> mozilla/openssl/.git
> mozilla/xlat/.git
>
> Now, we can be reasonably sure that the 'xlat' repo is something the
> mozilla core team can push to, or at least we can consider the core repo
> owners an official "vendor" of tags for the submodule repo. I'm fairly
> certain openssl authors won't be too happy with allowing the thousands
> of projects using its code to push tags to its official repo though.

No need, when cloning a supermodule, it will make those special tags automatically in the submodule repo. They are only there to prevent prune from destroying those referenced commits after all. If the submodule is cloned directly, they aren't needed anyway, and those objects won't be part of the dependency chain so wouldn't be downloaded.

Show 6 quoted lines
> Now that I think about it more, I realize this is completely irrelevant
> as the ui can create the tags in the submodule with info only from the
> the supermodule, which means the submodule repo will only be special if
> it's connected to the supermodule. We just need a command for creating
> those tags in the submodule repo so people who use the same submodule
> code for several projects can use the alternates mechanism effectively.

Is that even necessary? git-clone of a supermodule will make those tags automatically. If a submodule was alternative-cloned into a different supermodule, well then THAT supermodule would make the right tags for itself. Ah, I think I see what you mean now though, a method would be needed for creating those tags if we managed to manually get a submodule repository in to the supermodule - then supermodule-clone wouldn't have run. Perhaps they could be checked for at commit time and recreated then?

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE

← back to recent threads