threads / discuss / 31051

Feature request: fetch --prune by default

Subject: Feature request: fetch --prune by default

## tl;dr

52 messages between Jul 19, 2012 and Jun 20, 2013.

replies: 51people: 11as markdown or json

Alexey Muranov· Jul 19, 2012, 07:30 UTC · lore
Hello,
i would like
`git fetch --prune <remote>`
to be the default behavior of
`git fetch <remote>`

In fact, i think this is the only reasonable behavior. Keeping copies of deleted remote branches after `fetch` is more confusing than useful.

(Excuse me if this question has already been discussed.)
Thank you.
Alexey Muranov.
Jeff King· Jul 19, 2012, 11:55 UTC · re: Alexey Muranov · lore

Re: Feature request: fetch --prune by default

On Thu, Jul 19, 2012 at 09:30:59AM +0200, Alexey Muranov wrote:
Show 10 quoted lines
> i would like
> 
> `git fetch --prune <remote>`
> 
> to be the default behavior of
> 
> `git fetch <remote>`
> 
> In fact, i think this is the only reasonable behavior.
> Keeping copies of deleted remote branches after `fetch` is more confusing than useful.

I agree it would be much less confusing. However, one downside is that we do not keep reflogs on deleted branches (and nor did the commits in remote branches necessarily make it into the HEAD reflog). That makes "git fetch" a potentially destructive operation (you irrevocably lose the notion of which remote branches pointed where before the fetch, and you open up new commits to immediate pruning by "gc --auto".

So I think it would be a lot more palatable if we kept reflogs on deleted branches. That, in turn, has a few open issues, such as how to manage namespace conflicts (e.g., the fact that a deleted "foo" branch can conflict with a new "foo/bar" branch).

-Peff
Dan Johnson· Jul 19, 2012, 14:03 UTC · re: Jeff King · lore

Re: Feature request: fetch --prune by default

On Thu, Jul 19, 2012 at 7:55 AM, Jeff King <peff@peff.net> wrote:
Show 24 quoted lines
> On Thu, Jul 19, 2012 at 09:30:59AM +0200, Alexey Muranov wrote:
>
>> i would like
>>
>> `git fetch --prune <remote>`
>>
>> to be the default behavior of
>>
>> `git fetch <remote>`
>>
>> In fact, i think this is the only reasonable behavior.
>> Keeping copies of deleted remote branches after `fetch` is more confusing than useful.
>
> I agree it would be much less confusing. However, one downside is that
> we do not keep reflogs on deleted branches (and nor did the commits in
> remote branches necessarily make it into the HEAD reflog). That makes
> "git fetch" a potentially destructive operation (you irrevocably lose
> the notion of which remote branches pointed where before the fetch, and
> you open up new commits to immediate pruning by "gc --auto".
>
> So I think it would be a lot more palatable if we kept reflogs on
> deleted branches. That, in turn, has a few open issues, such as how to
> manage namespace conflicts (e.g., the fact that a deleted "foo" branch
> can conflict with a new "foo/bar" branch).

In the meantime, would it make sense to introduce a configuration variable to request this behavior?

If so, should it be global?
fetch.prune = always
or per-remote?
remote.<name>.prune = always

The global option seems to be more in line with what Alexey is looking for, but the per-remote one is similar to the tagopt option, which is a similar idea.

Of course, this might be just a waste of time to introduce a feature no one would use, in which case we obviously should not introduce such options.

-- 
-Dan
Stefan Haller· Jul 19, 2012, 15:11 UTC · re: Dan Johnson · lore

Re: Feature request: fetch --prune by default

Dan Johnson <computerdruid@gmail.com> wrote:
Show 8 quoted lines
> In the meantime, would it make sense to introduce a configuration
> variable to request this behavior?
> 
> fetch.prune = always
> 
> Of course, this might be just a waste of time to introduce a feature
> no one would use, in which case we obviously should not introduce such
> options.
I would use it.
-- 
Stefan Haller
Berlin, Germany
http://www.haller-berlin.de/
Junio C Hamano· Aug 16, 2012, 23:22 UTC · re: Dan Johnson · lore

Re: Feature request: fetch --prune by default

Dan Johnson <computerdruid@gmail.com> writes:
Show 25 quoted lines
> On Thu, Jul 19, 2012 at 7:55 AM, Jeff King <peff@peff.net> wrote:
> ...
>> So I think it would be a lot more palatable if we kept reflogs on
>> deleted branches. That, in turn, has a few open issues, such as how to
>> manage namespace conflicts (e.g., the fact that a deleted "foo" branch
>> can conflict with a new "foo/bar" branch).
>
> In the meantime, would it make sense to introduce a configuration
> variable to request this behavior?
>
> If so, should it be global?
>
> fetch.prune = always
>
> or per-remote?
>
> remote.<name>.prune = always
>
> The global option seems to be more in line with what Alexey is looking
> for, but the per-remote one is similar to the tagopt option, which is
> a similar idea.
>
> Of course, this might be just a waste of time to introduce a feature
> no one would use, in which case we obviously should not introduce such
> options.

I was reading through the backlog today and noticed that this topic veered into the "reflog graveyard" tangent. I wasn't involved in the main topic, but I think having both configuration variables, remote.<remote>.prune taking precedence over fetch.prune, as long as we make sure "fetch --no-prune" will override any configured default, is not a bad thing per-se.

As long as the users who elect to use this feature are aware of the pruning of the refs and logs, that is, but "branch [-r] -d" has been the way to lose both the branch and its log for a long time, so I do not see a big issue there, either.

The log graveyard is an independently interesting idea, which I may ping separately, but I consider it pretty much orthogonal to this particular topic.

Jeff King· Aug 21, 2012, 06:51 UTC · re: Junio C Hamano · lore

Re: Feature request: fetch --prune by default

On Thu, Aug 16, 2012 at 04:22:54PM -0700, Junio C Hamano wrote:
Show 34 quoted lines
> > In the meantime, would it make sense to introduce a configuration
> > variable to request this behavior?
> >
> > If so, should it be global?
> >
> > fetch.prune = always
> >
> > or per-remote?
> >
> > remote.<name>.prune = always
> >
> > The global option seems to be more in line with what Alexey is looking
> > for, but the per-remote one is similar to the tagopt option, which is
> > a similar idea.
> >
> > Of course, this might be just a waste of time to introduce a feature
> > no one would use, in which case we obviously should not introduce such
> > options.
> 
> I was reading through the backlog today and noticed that this topic
> veered into the "reflog graveyard" tangent.  I wasn't involved in
> the main topic, but I think having both configuration variables,
> remote.<remote>.prune taking precedence over fetch.prune, as long as
> we make sure "fetch --no-prune" will override any configured
> default, is not a bad thing per-se.
> 
> As long as the users who elect to use this feature are aware of the
> pruning of the refs and logs, that is, but "branch [-r] -d" has been
> the way to lose both the branch and its log for a long time, so I do
> not see a big issue there, either.
> 
> The log graveyard is an independently interesting idea, which I may
> ping separately, but I consider it pretty much orthogonal to this
> particular topic.

Yeah, I think that is sensible. As long as the _default_ is not to prune, and people are making a conscious choice to prune, I don't see a problem at all.

The log graveyard is orthogonal to the proposed option, but I think it would be a necessary step before flipping the default for that option to "true".

-Peff
Sam Roberts· Jun 20, 2013, 19:22 UTC · re: Dan Johnson · lore

Re: Feature request: fetch --prune by default

I would use the config feature to turn on --prune for fetch, and was surprised that it wasn't available - I hit this thread because I figured I somehow missed it in the config docs.

Having both global and local settings seems nice.

-- View this message in context: http://git.661346.n2.nabble.com/Feature-request-fetch-prune-by-default-tp7563241p7590048.html Sent from the git mailing list archive at Nabble.com.

Alexey Muranov· Jul 19, 2012, 16:21 UTC · re: Jeff King · lore

Re: Feature request: fetch --prune by default

On 19 Jul 2012, at 13:55, Jeff King wrote:
Show 19 quoted lines
> On Thu, Jul 19, 2012 at 09:30:59AM +0200, Alexey Muranov wrote:
> 
>> i would like
>> 
>> `git fetch --prune <remote>`
>> 
>> to be the default behavior of
>> 
>> `git fetch <remote>`
>> 
>> In fact, i think this is the only reasonable behavior.
>> Keeping copies of deleted remote branches after `fetch` is more confusing than useful.
> 
> I agree it would be much less confusing. However, one downside is that
> we do not keep reflogs on deleted branches (and nor did the commits in
> remote branches necessarily make it into the HEAD reflog). That makes
> "git fetch" a potentially destructive operation (you irrevocably lose
> the notion of which remote branches pointed where before the fetch, and
> you open up new commits to immediate pruning by "gc --auto".

I do not still understand very well some aspects of Git, like the exact purpose of "remote tracking branches" (are they for pull or for push?), so i may be wrong. However, i thought that a user was not expected to follow the moves of a remote branch of which the user is not an owner: if the user needs to follow the brach and not lose its commits, he/she should create a remote tracking branch.

> So I think it would be a lot more palatable if we kept reflogs on
> deleted branches. That, in turn, has a few open issues, such as how to
> manage namespace conflicts (e.g., the fact that a deleted "foo" branch
> can conflict with a new "foo/bar" branch).
I prefer to think of a remote branch and its local copy as the same thing, which are physically different only because of current real world/hardware/software limitations, which make it necessary to keep a local cache of remote data.  With this approach, reflogs should be deleted with the branch, and there will be no namespace conflicts.
Alexey.
Konstantin Khomoutov· Jul 19, 2012, 17:34 UTC · re: Alexey Muranov · lore

Re: Feature request: fetch --prune by default

On Thu, 19 Jul 2012 18:21:21 +0200 Alexey Muranov <alexey.muranov@gmail.com> wrote:

[...]
> I do not still understand very well some aspects of Git, like the
> exact purpose of "remote tracking branches" (are they for pull or for
> push?), so i may be wrong.

This is wery well explained in the Pro Git book, for instance. And in numerous blog posts etc.

> However, i thought that a user was not
> expected to follow the moves of a remote branch of which the user is
> not an owner: if the user needs to follow the brach and not lose its
> commits, he/she should create a remote tracking branch.

This would present another namespacing issue: how would you name the branches you're interested in so that they don't clash with your own personal local branches? You'd have to invent a scheme which would encode the remote's name in a branch name. But remote branches already do just this. So you create a remote tracking branch when you intend to actually *develop* something on that branch with the final intention to push that work back.

Show 10 quoted lines
> > So I think it would be a lot more palatable if we kept reflogs on
> > deleted branches. That, in turn, has a few open issues, such as how
> > to manage namespace conflicts (e.g., the fact that a deleted "foo"
> > branch can conflict with a new "foo/bar" branch).
> 
> I prefer to think of a remote branch and its local copy as the same
> thing, which are physically different only because of current real
> world/hardware/software limitations, which make it necessary to keep
> a local cache of remote data.  With this approach, reflogs should be
> deleted with the branch, and there will be no namespace conflicts.

It appears, the distributed nature of a DVCS did not fully sink into your mindset yet. ;-) Looks like you mentally treat a Git remote as a thing being used to access a centralized "reference" server which maintains a master copy of a repository, of which you happen to also have a local copy. Then it's quite logically to think that if someone deleted a branch in the master copy, everyone "downstream" should have the same remote branch deleted to be in sync with that master copy. But this is not the only way to organize your work. You could fetch from someone else's repository and be interested in their branch "foo", but think what happens when you fetch next time from that repo and see Git happily deleting your local branch thatremote/foo simply because someone with push access deleted that branch from the repo. This might *not* be what you really want or expect.

Alexey Muranov· Jul 19, 2012, 21:20 UTC · re: Konstantin Khomoutov · lore

Re: Feature request: fetch --prune by default

On 19 Jul 2012, at 19:34, Konstantin Khomoutov wrote:
Show 9 quoted lines
> On Thu, 19 Jul 2012 18:21:21 +0200
> Alexey Muranov <alexey.muranov@gmail.com> wrote:
> 
> [...]
>> I do not still understand very well some aspects of Git, like the
>> exact purpose of "remote tracking branches" (are they for pull or for
>> push?), so i may be wrong.
> This is wery well explained in the Pro Git book, for instance.
> And in numerous blog posts etc.
I have read the Pro Gut book and numerous blog posts, but i keep forgetting the explanation because it does not make much sense to me:
"Tracking branches are local branches that have a direct relationship to a remote branch.  If you’re on a tracking branch and type git push, Git automatically knows which server and branch to push to.  Also, running git pull while on one of these branches fetches all the remote references and then automatically merges in the corresponding remote branch." etc.
Why the same "direct relationship" for push and pull?  What happens if one of the branches was reset (yes, i know, "push -f").  Most importantly, what is the purpose of it? It is natural to expect that you might be pushing to and pulling from different remotes, i can even imagine pulling from more than one.
Show 11 quoted lines
>> However, i thought that a user was not
>> expected to follow the moves of a remote branch of which the user is
>> not an owner: if the user needs to follow the brach and not lose its
>> commits, he/she should create a remote tracking branch.
> This would present another namespacing issue: how would you name the
> branches you're interested in so that they don't clash with your own
> personal local branches?  You'd have to invent a scheme which would
> encode the remote's name in a branch name.  But remote branches already
> do just this.  So you create a remote tracking branch when you intend
> to actually *develop* something on that branch with the final intention
> to push that work back.
But i am not interested in remote branches, they are just fetched automatically when i do "git fetch".  You cannot commit to a remote branch, and i think it is not common to checkout them without a "-b" option.  If i am interested in them, i name them somehow.  I think this is the only practical way if i do not want to chase reflogs, because the owner of the branch can reset or rebase it anytime.  I do not develop on tracking branches.  In fact, i am not even using "git pull".
Show 24 quoted lines
>>> So I think it would be a lot more palatable if we kept reflogs on
>>> deleted branches. That, in turn, has a few open issues, such as how
>>> to manage namespace conflicts (e.g., the fact that a deleted "foo"
>>> branch can conflict with a new "foo/bar" branch).
>> 
>> I prefer to think of a remote branch and its local copy as the same
>> thing, which are physically different only because of current real
>> world/hardware/software limitations, which make it necessary to keep
>> a local cache of remote data.  With this approach, reflogs should be
>> deleted with the branch, and there will be no namespace conflicts.
> It appears, the distributed nature of a DVCS did not fully sink into
> your mindset yet. ;-)
> Looks like you mentally treat a Git remote as a thing being used to
> access a centralized "reference" server which maintains a master copy
> of a repository, of which you happen to also have a local copy.
> Then it's quite logically to think that if someone deleted a branch in
> the master copy, everyone "downstream" should have the same
> remote branch deleted to be in sync with that master copy.
> But this is not the only way to organize your work.
> You could fetch from someone else's repository and be interested in
> their branch "foo", but think what happens when you fetch next time from
> that repo and see Git happily deleting your local branch thatremote/foo
> simply because someone with push access deleted that branch from the
> repo.  This might *not* be what you really want or expect.
But this is true that the object store of Git can be viewed as a single centralized repository.  The fact that not everybody has access to every object in Git is a limitation and not a benefit.  These are the branches which are individual, and i do not think it is a good habit to treat every reference that was ever fetched with "git fetch" as your own, and put reflogs of all fetched remote branches under Git version control :D.
If i care about "thatremote/foo" branch, i "track" it, i do not plan to go through reflogs if it is rebased.
Alexey.
Alexey Muranov· Jul 19, 2012, 21:57 UTC · re: Alexey Muranov · lore

Re: Feature request: fetch --prune by default

I just want to correct my mistake in what i've just sent:
On 19 Jul 2012, at 23:20, Alexey Muranov wrote:
> because the owner of the branch can reset or rebase it anytime.  I do not develop on tracking branches.  In fact, i am not even using "git pull".
> I do not develop on tracking branches.

Of course i develop on "tracking" branches, i just got confused once again by pull/push thing: i develop on branches that track origin, not upstream. I think they should be called "remotely tracked branches", so there would be "remote tracking branches" for pull and "remotely tracked branches" for push.

Alexey.
Johannes Sixt· Jul 20, 2012, 07:11 UTC · re: Alexey Muranov · lore

Re: Feature request: fetch --prune by default

Am 7/19/2012 23:20, schrieb Alexey Muranov:
Show 21 quoted lines
> On 19 Jul 2012, at 19:34, Konstantin Khomoutov wrote:
> 
>> On Thu, 19 Jul 2012 18:21:21 +0200 Alexey Muranov
>> <alexey.muranov@gmail.com> wrote:
>> 
>> [...]
>>> I do not still understand very well some aspects of Git, like the 
>>> exact purpose of "remote tracking branches" (are they for pull or
>>> for push?), so i may be wrong.
>> This is wery well explained in the Pro Git book, for instance. And in
>> numerous blog posts etc.
> 
> I have read the Pro Gut book and numerous blog posts, but i keep
> forgetting the explanation because it does not make much sense to me:
> 
> "Tracking branches are local branches that have a direct relationship
> to a remote branch.  If you’re on a tracking branch and type git push,
> Git automatically knows which server and branch to push to.  Also,
> running git pull while on one of these branches fetches all the remote
> references and then automatically merges in the corresponding remote
> branch." etc.

Note the difference between "tracking branch" and "remote tracking branch"! The "remote tracking branches" are the refs in the refs/remotes/ hierarchy. The "tracking branches" are your own local branches that you have created with 'git branch topic thatremote/topic' (or perhaps 'git checkout -b'). The paragraph talks about the latter.

-- Hannes
Alexey Muranov· Jul 20, 2012, 07:28 UTC · re: Johannes Sixt · lore

Re: Feature request: fetch --prune by default

On 20 Jul 2012, at 09:11, Johannes Sixt wrote:
Show 28 quoted lines
> Am 7/19/2012 23:20, schrieb Alexey Muranov:
>> On 19 Jul 2012, at 19:34, Konstantin Khomoutov wrote:
>> 
>>> On Thu, 19 Jul 2012 18:21:21 +0200 Alexey Muranov
>>> <alexey.muranov@gmail.com> wrote:
>>> 
>>> [...]
>>>> I do not still understand very well some aspects of Git, like the 
>>>> exact purpose of "remote tracking branches" (are they for pull or
>>>> for push?), so i may be wrong.
>>> This is wery well explained in the Pro Git book, for instance. And in
>>> numerous blog posts etc.
>> 
>> I have read the Pro Gut book and numerous blog posts, but i keep
>> forgetting the explanation because it does not make much sense to me:
>> 
>> "Tracking branches are local branches that have a direct relationship
>> to a remote branch.  If you’re on a tracking branch and type git push,
>> Git automatically knows which server and branch to push to.  Also,
>> running git pull while on one of these branches fetches all the remote
>> references and then automatically merges in the corresponding remote
>> branch." etc.
> 
> Note the difference between "tracking branch" and "remote tracking
> branch"! The "remote tracking branches" are the refs in the refs/remotes/
> hierarchy. The "tracking branches" are your own local branches that you
> have created with 'git branch topic thatremote/topic' (or perhaps 'git
> checkout -b'). The paragraph talks about the latter.
Hannes, thanks for the explanation, so i was confused once again.
Various blog posts do not make the terminology clear, for example
http://gitready.com/beginner/2009/03/09/remote-tracking-branches.html
sais that there are only "two types of branches: local, and remote-tracking", while i think it depends on perspective.
There are in fact
1. remote,
2. remote-tracking (which are local!),
3. truly local:
  a) which are tracking some remote-tracking(!) branches,
  b) and which are not tracking.
I think i was also misguided by Konstantin, who wrote that "you create a remote tracking branch when you intend to actually *develop* something on that branch" :).
-Alexey.
Junio C Hamano· Aug 16, 2012, 23:27 UTC · re: Alexey Muranov · lore

Re: Feature request: fetch --prune by default

Alexey Muranov <alexey.muranov@gmail.com> writes:
Show 17 quoted lines
> On 20 Jul 2012, at 09:11, Johannes Sixt wrote:
> ...
>> Note the difference between "tracking branch" and "remote tracking
>> branch"! The "remote tracking branches" are the refs in the refs/remotes/
>> hierarchy. The "tracking branches" are your own local branches that you
>> have created with 'git branch topic thatremote/topic' (or perhaps 'git
>> checkout -b'). The paragraph talks about the latter.
>
> Hannes, thanks for the explanation, so i was confused once again.
>
> Various blog posts do not make the terminology clear, for example
> http://gitready.com/beginner/2009/03/09/remote-tracking-branches.html
> sais that there are only "two types of branches: local, and remote-tracking"...
> ...
> I think i was also misguided by Konstantin, who wrote that "you
> create a remote tracking branch when you intend to actually
> *develop* something on that branch" :).
I was re-reading the backlog today, and saw this topic fizzled out.

We obviously cannot fix third-party documentation that teach lies to people, but is there something we can do to improve our own documentation with respect to this confusion?

As I wrote it elsewhere, I try to avoid the bareword "tracking" in general, and call the local branch you build on something like "your 'next' branch that forked from origin/next remote tracking branch" myself. Perhaps we can start from checking the documentation with such a phrasing discipline?

Alexey Muranov· Jul 19, 2012, 16:40 UTC · re: Jeff King · lore

Re: Feature request: fetch --prune by default

On 19 Jul 2012, at 13:55, Jeff King wrote:
Show 6 quoted lines
> I agree it would be much less confusing. However, one downside is that
> we do not keep reflogs on deleted branches (and nor did the commits in
> remote branches necessarily make it into the HEAD reflog). That makes
> "git fetch" a potentially destructive operation (you irrevocably lose
> the notion of which remote branches pointed where before the fetch, and
> you open up new commits to immediate pruning by "gc --auto".

If i understand correctly, existence of a reflog entry will not stop "gc" from removing a commit, will it? In this case, if a remote branch was rebased or reset, commits can be lost anyway, right?

Alexey.
Dan Johnson· Jul 19, 2012, 16:48 UTC · re: Alexey Muranov · lore

Re: Feature request: fetch --prune by default

On Thu, Jul 19, 2012 at 12:40 PM, Alexey Muranov <alexey.muranov@gmail.com> wrote:

Show 11 quoted lines
> On 19 Jul 2012, at 13:55, Jeff King wrote:
>
>> I agree it would be much less confusing. However, one downside is that
>> we do not keep reflogs on deleted branches (and nor did the commits in
>> remote branches necessarily make it into the HEAD reflog). That makes
>> "git fetch" a potentially destructive operation (you irrevocably lose
>> the notion of which remote branches pointed where before the fetch, and
>> you open up new commits to immediate pruning by "gc --auto".
>
> If i understand correctly, existence of a reflog entry will not stop "gc" from removing a commit, will it?
> In this case, if a remote branch was rebased or reset, commits can be lost anyway, right?

From the git-gc man page: git gc tries very hard to be safe about the garbage it collects. In particular, it will keep not only objects referenced by your current set of branches and tags, but also objects referenced by the index, remote-tracking branches, refs saved by git filter-branch in refs/original/, or reflogs (which may reference commits in branches that were later amended or rewound).

So yes, a reflog entry does stop gc from removing objects, including commits. It will expire old reflog entries (90 days by default) though, so it's not like they will stay around forever.

-- 
-Dan
Alexey Muranov· Jul 19, 2012, 16:51 UTC · re: Dan Johnson · lore

Re: Feature request: fetch --prune by default

On 19 Jul 2012, at 18:48, Dan Johnson wrote:
Show 11 quoted lines
> From the git-gc man page:
> git gc tries very hard to be safe about the garbage it collects. In
> particular, it will keep not only objects referenced by your current
> set of branches and tags, but also objects referenced by the index,
> remote-tracking branches, refs saved by git filter-branch in
> refs/original/, or reflogs (which may reference commits in branches
> that were later amended or rewound).
> 
> So yes, a reflog entry does stop gc from removing objects, including
> commits. It will expire old reflog entries (90 days by default)
> though, so it's not like they will stay around forever.
Dan, thanks for the explanation.
Alexey.
Jeff King· Jul 19, 2012, 21:32 UTC · re: Jeff King · lore

[RFC/PATCH 0/3] reflog graveyard

On Thu, Jul 19, 2012 at 07:55:58AM -0400, Jeff King wrote:
> So I think it would be a lot more palatable if we kept reflogs on
> deleted branches. That, in turn, has a few open issues, such as how to
> manage namespace conflicts (e.g., the fact that a deleted "foo" branch
> can conflict with a new "foo/bar" branch).

Here is a patch series to address that. I think I have smoothed out most of the rough edges, but I wouldn't be surprised if there are some other corner cases. One that I notice is that "git log -g" will stop walking when it hits a null sha1 in the reflog.

  [1/3]: retain reflogs for deleted refs
  [2/3]: teach sha1_name to look in graveyard reflogs
  [3/3]: add tests for reflogs of deleted refs
-Peff
Jeff King· Jul 19, 2012, 21:33 UTC · re: Jeff King · lore

[PATCH 1/3] retain reflogs for deleted refs

When a ref is deleted, we completely delete its reflog on the spot, leaving very little help for the user to reverse the action. One can sometimes reconstruct the missing entries based on the HEAD reflog, but not always; the deleted entries may not have ever been on HEAD (for example, in the case of a refs/remotes branch that was pruned). That leaves "git fsck --lost-found", which can be quite tedious.

Instead, let's keep the reflogs for deleted refs around until their entries naturally expire according to the regular reflog expiration rules.

This cannot be done by simply leaving the reflog files in place. The ref namespace does not allow D/F conflicts, so a ref "foo" would block the creation of another ref "foo/bar", and vice versa. This limitation is acceptable for two refs to exist simultaneously, but should not have an impact if one of the refs is deleted.

This patch moves reflog entries into a special "graveyard" namespace, and appends a tilde (~) character, which is not allowed in a valid ref name. This means that the deleted reflogs of these refs:

   refs/heads/a
   refs/heads/a/b
   refs/heads/a/b/c
will be stored in:
   logs/graveyard/refs/heads/a~
   logs/graveyard/refs/heads/a/b~
   logs/graveyard/refs/heads/a/b/c~

Putting them in the graveyard namespace ensures they will not conflict with live refs, and the tilde prevents D/F conflicts within the graveyard namespace.

The implementation is fairly straightforward, but it's worth noting a few things:

  1. Updates to "logs/graveyard/refs/heads/foo~" happen
     under the ref-lock for "refs/heads/foo". So deletion
     still takes a single lock, and anyone touching the
     reflog directly needs to reverse the transformation to
     find the correct lockfile.
  2. We append entries to the graveyard reflog rather than
     simply renaming the file into place. This means that
     if you create and delete a branch repeatedly, the
     graveyard will contain the concatenation of all
     iterations.
  3. We do not resurrect dead entries when a new ref is
     created with the same name. However, it would be
     possible to build an "undelete" feature on top of this
     if one was so inclined.
  4. The for_each_reflog code has been loosened to allow
     reflogs that do not have a matching ref. In this case,
     the callback is passed the null_sha1, and callers must
     be prepared to handle this case (the only caller that
     cares is the reflog expiration code, which is updated
     here).

Only one test needed to be updated; t7701 tries to create unreachable objects by deleting branches. Of course that no longer works, which is the intent of this patch. The test now works around it by removing the graveyard logs.

Signed-off-by: Jeff King <peff@peff.net>
---
 builtin/reflog.c                     |  9 +++--
 refs.c                               | 69 +++++++++++++++++++++++++++++++++---
 refs.h                               |  3 ++
 t/t7701-repack-unpack-unreachable.sh |  5 ++-
 4 files changed, 79 insertions(+), 7 deletions(-)
diff --git a/builtin/reflog.c b/builtin/reflog.c
index b3c9e27..e79a2ca 100644
--- a/builtin/reflog.c
+++ b/builtin/reflog.c
@@ -359,6 +359,7 @@ static int expire_reflog(const char *ref, const unsigned char *sha1, int unused,
 	struct commit *tip_commit;
 	struct commit_list *tips;
 	int status = 0;
+	int updateref = cmd->updateref && !is_null_sha1(sha1);
 
 	memset(&cb, 0, sizeof(cb));
 
@@ -367,6 +368,10 @@ static int expire_reflog(const char *ref, const unsigned char *sha1, int unused,
 	 * getting updated.
 	 */
 	lock = lock_any_ref_for_update(ref, sha1, 0);
+	if (!lock && is_null_sha1(sha1))
+		lock = lock_any_ref_for_update(
+				graveyard_reflog_to_refname(ref),
+				sha1, 0);
 	if (!lock)
 		return error("cannot lock ref '%s'", ref);
 	log_file = git_pathdup("logs/%s", ref);
@@ -426,7 +431,7 @@ static int expire_reflog(const char *ref, const unsigned char *sha1, int unused,
 			status |= error("%s: %s", strerror(errno),
 					newlog_path);
 			unlink(newlog_path);
-		} else if (cmd->updateref &&
+		} else if (updateref &&
 			(write_in_full(lock->lock_fd,
 				sha1_to_hex(cb.last_kept_sha1), 40) != 40 ||
 			 write_str_in_full(lock->lock_fd, "\n") != 1 ||
@@ -438,7 +443,7 @@ static int expire_reflog(const char *ref, const unsigned char *sha1, int unused,
 			status |= error("cannot rename %s to %s",
 					newlog_path, log_file);
 			unlink(newlog_path);
-		} else if (cmd->updateref && commit_ref(lock)) {
+		} else if (updateref && commit_ref(lock)) {
 			status |= error("Couldn't set %s", lock->ref_name);
 		} else {
 			adjust_shared_perm(log_file);
diff --git a/refs.c b/refs.c
index da74a2b..553de77 100644
--- a/refs.c
+++ b/refs.c
@@ -4,6 +4,8 @@
 #include "tag.h"
 #include "dir.h"
 
+static void mark_reflog_deleted(struct ref_lock *lock);
+
 /*
  * Make sure "ref" is something reasonable to have under ".git/refs/";
  * We do not like it if:
@@ -1780,7 +1782,7 @@ int delete_ref(const char *refname, const unsigned char *sha1, int delopt)
 	 */
 	ret |= repack_without_ref(refname);
 
-	unlink_or_warn(git_path("logs/%s", lock->ref_name));
+	mark_reflog_deleted(lock);
 	invalidate_ref_cache(NULL);
 	unlock_ref(lock);
 	return ret;
@@ -2385,9 +2387,8 @@ static int do_for_each_reflog(struct strbuf *name, each_ref_fn fn, void *cb_data
 			} else {
 				unsigned char sha1[20];
 				if (read_ref_full(name->buf, sha1, 0, NULL))
-					retval = error("bad ref for %s", name->buf);
-				else
-					retval = fn(name->buf, sha1, 0, cb_data);
+					hashcpy(sha1, null_sha1);
+				retval = fn(name->buf, sha1, 0, cb_data);
 			}
 			if (retval)
 				break;
@@ -2552,3 +2553,63 @@ char *shorten_unambiguous_ref(const char *refname, int strict)
 	free(short_name);
 	return xstrdup(refname);
 }
+
+char *refname_to_graveyard_reflog(const char *ref)
+{
+	return git_path("logs/graveyard/%s~", ref);
+}
+
+char *graveyard_reflog_to_refname(const char *log)
+{
+	static struct strbuf buf = STRBUF_INIT;
+
+	if (!prefixcmp(log, "graveyard/"))
+		log += 10;
+
+	strbuf_reset(&buf);
+	strbuf_addstr(&buf, log);
+	if (buf.len > 0 && buf.buf[buf.len-1] == '~')
+		strbuf_setlen(&buf, buf.len - 1);
+
+	return buf.buf;
+}
+
+static int copy_reflog_entries(const char *dst, const char *src)
+{
+	int fdi, fdo, status;
+
+	fdi = open(src, O_RDONLY);
+	if (fdi < 0)
+		return errno == ENOENT ? 0 : -1;
+
+	fdo = open(dst, O_WRONLY | O_APPEND | O_CREAT, 0666);
+	if (fdo < 0) {
+		close(fdi);
+		return -1;
+	}
+
+	status = copy_fd(fdi, fdo);
+	if (close(fdo) < 0)
+		return -1;
+	if (status < 0 || adjust_shared_perm(dst) < 0)
+		return -1;
+	return 0;
+}
+
+static void mark_reflog_deleted(struct ref_lock *lock)
+{
+	static const char msg[] = "ref deleted";
+	const char *log = git_path("logs/%s", lock->ref_name);
+	char *grave = refname_to_graveyard_reflog(lock->ref_name);
+
+	if (log_ref_write(lock->ref_name, lock->old_sha1, null_sha1, msg) < 0)
+		warning("unable to update reflog for %s: %s",
+			lock->ref_name, strerror(errno));
+
+	if (safe_create_leading_directories(grave) < 0 ||
+	    copy_reflog_entries(grave, log) < 0)
+		warning("unable to copy reflog entries to graveyard: %s",
+			strerror(errno));
+
+	unlink_or_warn(log);
+}
diff --git a/refs.h b/refs.h
index d6c2fe2..9d14558 100644
--- a/refs.h
+++ b/refs.h
@@ -111,6 +111,9 @@ int for_each_recent_reflog_ent(const char *refname, each_reflog_ent_fn fn, long,
  */
 extern int for_each_reflog(each_ref_fn, void *);
 
+char *refname_to_graveyard_reflog(const char *ref);
+char *graveyard_reflog_to_refname(const char *log);
+
 #define REFNAME_ALLOW_ONELEVEL 1
 #define REFNAME_REFSPEC_PATTERN 2
 #define REFNAME_DOT_COMPONENT 4
diff --git a/t/t7701-repack-unpack-unreachable.sh b/t/t7701-repack-unpack-unreachable.sh
index b8d4cde..c06b715 100755
--- a/t/t7701-repack-unpack-unreachable.sh
+++ b/t/t7701-repack-unpack-unreachable.sh
@@ -38,7 +38,9 @@ test_expect_success '-A with -d option leaves unreachable objects unpacked' '
 	git show $csha1 &&
 	git show $tsha1 &&
 	# now expire the reflog, while keeping reachable ones but expiring
-	# unreachables immediately
+	# unreachables immediately; also remove any graveyard reflogs
+	# from deleted branches that would keep things reachable
+	rm -rf .git/logs/graveyard &&
 	test_tick &&
 	sometimeago=$(( $test_tick - 10000 )) &&
 	git reflog expire --expire=$sometimeago --expire-unreachable=$test_tick --all &&
@@ -76,6 +78,7 @@ test_expect_success '-A without -d option leaves unreachable objects packed' '
 	test 1 = $(ls -1 .git/objects/pack/pack-*.pack | wc -l) &&
 	packfile=$(ls .git/objects/pack/pack-*.pack) &&
 	git branch -D transient_branch &&
+	rm -rf .git/logs/graveyard &&
 	test_tick &&
 	git repack -A -l &&
 	test ! -f "$fsha1path" &&
-- 
1.7.10.5.40.g059818d
Alexey Muranov· Jul 19, 2012, 22:23 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

Jeff,
i have no idea about Git source and little idea of how it is working internally, but reading through your message i wonder: wouldn't it be a good idea to timestamp the dead reflogs ?
Alexey.
On 19 Jul 2012, at 23:33, Jeff King wrote:
Show 257 quoted lines
> When a ref is deleted, we completely delete its reflog on
> the spot, leaving very little help for the user to reverse
> the action. One can sometimes reconstruct the missing
> entries based on the HEAD reflog, but not always; the
> deleted entries may not have ever been on HEAD (for example,
> in the case of a refs/remotes branch that was pruned). That
> leaves "git fsck --lost-found", which can be quite tedious.
> 
> Instead, let's keep the reflogs for deleted refs around
> until their entries naturally expire according to the
> regular reflog expiration rules.
> 
> This cannot be done by simply leaving the reflog files in
> place. The ref namespace does not allow D/F conflicts, so a
> ref "foo" would block the creation of another ref "foo/bar",
> and vice versa. This limitation is acceptable for two refs
> to exist simultaneously, but should not have an impact if
> one of the refs is deleted.
> 
> This patch moves reflog entries into a special "graveyard"
> namespace, and appends a tilde (~) character, which is
> not allowed in a valid ref name. This means that the deleted
> reflogs of these refs:
> 
>   refs/heads/a
>   refs/heads/a/b
>   refs/heads/a/b/c
> 
> will be stored in:
> 
>   logs/graveyard/refs/heads/a~
>   logs/graveyard/refs/heads/a/b~
>   logs/graveyard/refs/heads/a/b/c~
> 
> Putting them in the graveyard namespace ensures they will
> not conflict with live refs, and the tilde prevents D/F
> conflicts within the graveyard namespace.
> 
> The implementation is fairly straightforward, but it's worth
> noting a few things:
> 
>  1. Updates to "logs/graveyard/refs/heads/foo~" happen
>     under the ref-lock for "refs/heads/foo". So deletion
>     still takes a single lock, and anyone touching the
>     reflog directly needs to reverse the transformation to
>     find the correct lockfile.
> 
>  2. We append entries to the graveyard reflog rather than
>     simply renaming the file into place. This means that
>     if you create and delete a branch repeatedly, the
>     graveyard will contain the concatenation of all
>     iterations.
> 
>  3. We do not resurrect dead entries when a new ref is
>     created with the same name. However, it would be
>     possible to build an "undelete" feature on top of this
>     if one was so inclined.
> 
>  4. The for_each_reflog code has been loosened to allow
>     reflogs that do not have a matching ref. In this case,
>     the callback is passed the null_sha1, and callers must
>     be prepared to handle this case (the only caller that
>     cares is the reflog expiration code, which is updated
>     here).
> 
> Only one test needed to be updated; t7701 tries to create
> unreachable objects by deleting branches. Of course that no
> longer works, which is the intent of this patch. The test
> now works around it by removing the graveyard logs.
> 
> Signed-off-by: Jeff King <peff@peff.net>
> ---
> builtin/reflog.c                     |  9 +++--
> refs.c                               | 69 +++++++++++++++++++++++++++++++++---
> refs.h                               |  3 ++
> t/t7701-repack-unpack-unreachable.sh |  5 ++-
> 4 files changed, 79 insertions(+), 7 deletions(-)
> 
> diff --git a/builtin/reflog.c b/builtin/reflog.c
> index b3c9e27..e79a2ca 100644
> --- a/builtin/reflog.c
> +++ b/builtin/reflog.c
> @@ -359,6 +359,7 @@ static int expire_reflog(const char *ref, const unsigned char *sha1, int unused,
> 	struct commit *tip_commit;
> 	struct commit_list *tips;
> 	int status = 0;
> +	int updateref = cmd->updateref && !is_null_sha1(sha1);
> 
> 	memset(&cb, 0, sizeof(cb));
> 
> @@ -367,6 +368,10 @@ static int expire_reflog(const char *ref, const unsigned char *sha1, int unused,
> 	 * getting updated.
> 	 */
> 	lock = lock_any_ref_for_update(ref, sha1, 0);
> +	if (!lock && is_null_sha1(sha1))
> +		lock = lock_any_ref_for_update(
> +				graveyard_reflog_to_refname(ref),
> +				sha1, 0);
> 	if (!lock)
> 		return error("cannot lock ref '%s'", ref);
> 	log_file = git_pathdup("logs/%s", ref);
> @@ -426,7 +431,7 @@ static int expire_reflog(const char *ref, const unsigned char *sha1, int unused,
> 			status |= error("%s: %s", strerror(errno),
> 					newlog_path);
> 			unlink(newlog_path);
> -		} else if (cmd->updateref &&
> +		} else if (updateref &&
> 			(write_in_full(lock->lock_fd,
> 				sha1_to_hex(cb.last_kept_sha1), 40) != 40 ||
> 			 write_str_in_full(lock->lock_fd, "\n") != 1 ||
> @@ -438,7 +443,7 @@ static int expire_reflog(const char *ref, const unsigned char *sha1, int unused,
> 			status |= error("cannot rename %s to %s",
> 					newlog_path, log_file);
> 			unlink(newlog_path);
> -		} else if (cmd->updateref && commit_ref(lock)) {
> +		} else if (updateref && commit_ref(lock)) {
> 			status |= error("Couldn't set %s", lock->ref_name);
> 		} else {
> 			adjust_shared_perm(log_file);
> diff --git a/refs.c b/refs.c
> index da74a2b..553de77 100644
> --- a/refs.c
> +++ b/refs.c
> @@ -4,6 +4,8 @@
> #include "tag.h"
> #include "dir.h"
> 
> +static void mark_reflog_deleted(struct ref_lock *lock);
> +
> /*
>  * Make sure "ref" is something reasonable to have under ".git/refs/";
>  * We do not like it if:
> @@ -1780,7 +1782,7 @@ int delete_ref(const char *refname, const unsigned char *sha1, int delopt)
> 	 */
> 	ret |= repack_without_ref(refname);
> 
> -	unlink_or_warn(git_path("logs/%s", lock->ref_name));
> +	mark_reflog_deleted(lock);
> 	invalidate_ref_cache(NULL);
> 	unlock_ref(lock);
> 	return ret;
> @@ -2385,9 +2387,8 @@ static int do_for_each_reflog(struct strbuf *name, each_ref_fn fn, void *cb_data
> 			} else {
> 				unsigned char sha1[20];
> 				if (read_ref_full(name->buf, sha1, 0, NULL))
> -					retval = error("bad ref for %s", name->buf);
> -				else
> -					retval = fn(name->buf, sha1, 0, cb_data);
> +					hashcpy(sha1, null_sha1);
> +				retval = fn(name->buf, sha1, 0, cb_data);
> 			}
> 			if (retval)
> 				break;
> @@ -2552,3 +2553,63 @@ char *shorten_unambiguous_ref(const char *refname, int strict)
> 	free(short_name);
> 	return xstrdup(refname);
> }
> +
> +char *refname_to_graveyard_reflog(const char *ref)
> +{
> +	return git_path("logs/graveyard/%s~", ref);
> +}
> +
> +char *graveyard_reflog_to_refname(const char *log)
> +{
> +	static struct strbuf buf = STRBUF_INIT;
> +
> +	if (!prefixcmp(log, "graveyard/"))
> +		log += 10;
> +
> +	strbuf_reset(&buf);
> +	strbuf_addstr(&buf, log);
> +	if (buf.len > 0 && buf.buf[buf.len-1] == '~')
> +		strbuf_setlen(&buf, buf.len - 1);
> +
> +	return buf.buf;
> +}
> +
> +static int copy_reflog_entries(const char *dst, const char *src)
> +{
> +	int fdi, fdo, status;
> +
> +	fdi = open(src, O_RDONLY);
> +	if (fdi < 0)
> +		return errno == ENOENT ? 0 : -1;
> +
> +	fdo = open(dst, O_WRONLY | O_APPEND | O_CREAT, 0666);
> +	if (fdo < 0) {
> +		close(fdi);
> +		return -1;
> +	}
> +
> +	status = copy_fd(fdi, fdo);
> +	if (close(fdo) < 0)
> +		return -1;
> +	if (status < 0 || adjust_shared_perm(dst) < 0)
> +		return -1;
> +	return 0;
> +}
> +
> +static void mark_reflog_deleted(struct ref_lock *lock)
> +{
> +	static const char msg[] = "ref deleted";
> +	const char *log = git_path("logs/%s", lock->ref_name);
> +	char *grave = refname_to_graveyard_reflog(lock->ref_name);
> +
> +	if (log_ref_write(lock->ref_name, lock->old_sha1, null_sha1, msg) < 0)
> +		warning("unable to update reflog for %s: %s",
> +			lock->ref_name, strerror(errno));
> +
> +	if (safe_create_leading_directories(grave) < 0 ||
> +	    copy_reflog_entries(grave, log) < 0)
> +		warning("unable to copy reflog entries to graveyard: %s",
> +			strerror(errno));
> +
> +	unlink_or_warn(log);
> +}
> diff --git a/refs.h b/refs.h
> index d6c2fe2..9d14558 100644
> --- a/refs.h
> +++ b/refs.h
> @@ -111,6 +111,9 @@ int for_each_recent_reflog_ent(const char *refname, each_reflog_ent_fn fn, long,
>  */
> extern int for_each_reflog(each_ref_fn, void *);
> 
> +char *refname_to_graveyard_reflog(const char *ref);
> +char *graveyard_reflog_to_refname(const char *log);
> +
> #define REFNAME_ALLOW_ONELEVEL 1
> #define REFNAME_REFSPEC_PATTERN 2
> #define REFNAME_DOT_COMPONENT 4
> diff --git a/t/t7701-repack-unpack-unreachable.sh b/t/t7701-repack-unpack-unreachable.sh
> index b8d4cde..c06b715 100755
> --- a/t/t7701-repack-unpack-unreachable.sh
> +++ b/t/t7701-repack-unpack-unreachable.sh
> @@ -38,7 +38,9 @@ test_expect_success '-A with -d option leaves unreachable objects unpacked' '
> 	git show $csha1 &&
> 	git show $tsha1 &&
> 	# now expire the reflog, while keeping reachable ones but expiring
> -	# unreachables immediately
> +	# unreachables immediately; also remove any graveyard reflogs
> +	# from deleted branches that would keep things reachable
> +	rm -rf .git/logs/graveyard &&
> 	test_tick &&
> 	sometimeago=$(( $test_tick - 10000 )) &&
> 	git reflog expire --expire=$sometimeago --expire-unreachable=$test_tick --all &&
> @@ -76,6 +78,7 @@ test_expect_success '-A without -d option leaves unreachable objects packed' '
> 	test 1 = $(ls -1 .git/objects/pack/pack-*.pack | wc -l) &&
> 	packfile=$(ls .git/objects/pack/pack-*.pack) &&
> 	git branch -D transient_branch &&
> +	rm -rf .git/logs/graveyard &&
> 	test_tick &&
> 	git repack -A -l &&
> 	test ! -f "$fsha1path" &&
> -- 
> 1.7.10.5.40.g059818d
> 
Jeff King· Jul 20, 2012, 14:26 UTC · re: Alexey Muranov · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On Fri, Jul 20, 2012 at 12:23:12AM +0200, Alexey Muranov wrote:
> i have no idea about Git source and little idea of how it is working
> internally, but reading through your message i wonder: wouldn't it be
> a good idea to timestamp the dead reflogs ?

Each individual entry in the reflog has its own timestamp, and the entries are expired individually over time as "git gc" is run. Or did you mean something else?

-Peff
Alexey Muranov· Jul 20, 2012, 14:32 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On 20 Jul 2012, at 16:26, Jeff King wrote:
Show 9 quoted lines
> On Fri, Jul 20, 2012 at 12:23:12AM +0200, Alexey Muranov wrote:
> 
>> i have no idea about Git source and little idea of how it is working
>> internally, but reading through your message i wonder: wouldn't it be
>> a good idea to timestamp the dead reflogs ?
> 
> Each individual entry in the reflog has its own timestamp, and the
> entries are expired individually over time as "git gc" is run. Or did
> you mean something else?
Yes, sorry, i was not clear, i meant to put dead reflogs into subdirectories yyyy-mm-dd, or maybe yyyy-mm-dd-hhmmss.
-Alexey.
Junio C Hamano· Jul 19, 2012, 22:36 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

Jeff King <peff@peff.net> writes:
> Only one test needed to be updated; t7701 tries to create
> unreachable objects by deleting branches. Of course that no
> longer works, which is the intent of this patch. The test
> now works around it by removing the graveyard logs.

I think the work-around indicates the need for regular users to be able to also discover, prune and delete these logs. Do we have "prune reflog for _this_ ref (or these refs), removing entries that are older than this threshold"? If so the codepath would need to know about the graveyard and the implementation detail of the tilde suffix so that the end users do not need to know about them.

I like the general direction. Perhaps a long distant future direction could be to also use the same trick in the ref namespace so that we can have 'next' branch itself, and 'next/foo', 'next/bar' forks that are based on the 'next' branch at the same time (it obviously is a totally unrelated topic)?

Jeff King· Jul 20, 2012, 14:43 UTC · re: Junio C Hamano · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On Thu, Jul 19, 2012 at 03:36:09PM -0700, Junio C Hamano wrote:
Show 13 quoted lines
> Jeff King <peff@peff.net> writes:
> 
> > Only one test needed to be updated; t7701 tries to create
> > unreachable objects by deleting branches. Of course that no
> > longer works, which is the intent of this patch. The test
> > now works around it by removing the graveyard logs.
> 
> I think the work-around indicates the need for regular users to be
> able to also discover, prune and delete these logs.  Do we have
> "prune reflog for _this_ ref (or these refs), removing entries that
> are older than this threshold"?  If so the codepath would need to
> know about the graveyard and the implementation detail of the tilde
> suffix so that the end users do not need to know about them.

We do have it: "git reflog expire --expire=now deleted-branch" is the right way to do it. Unfortunately, it does not work with my patch. The dwim_log correctly notes that a reflog exists (because it checks that the "graveyard" version of the ref exists), but then expire_reflog does not correctly fallback when opening the log (it usually has to do the _reverse_ translation, because it gets the graveyard log name from for_each_reflog, and has to find the correct lock).

I'll fix it in my re-roll, and then have t7701 use it.
Show 5 quoted lines
> I like the general direction.  Perhaps a long distant future
> direction could be to also use the same trick in the ref namespace
> so that we can have 'next' branch itself, and 'next/foo', 'next/bar'
> forks that are based on the 'next' branch at the same time (it
> obviously is a totally unrelated topic)?

I would love that, as it would mean we could simply leave the reflogs in place without having a separate graveyard namespace. Which means there wouldn't need to be any reflog-specific translation at all, and bugs like the one above wouldn't exist.

But it would mean that you cannot naively run
  echo $sha1 >.git/refs/heads/foo

anymore. I suspect that the packed-refs conversion rooted out many scripts that did not use update-ref and rev-parse to access refs, but the above does still work today. So I suspect there would be some fallout. Not to mention that older versions of git would be completely broken, which would mean we need a lengthy deprecation period while everybody upgrades to versions of git that support the reading side.

-Peff
Jeff King· Jul 20, 2012, 15:07 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On Fri, Jul 20, 2012 at 10:43:37AM -0400, Jeff King wrote:
Show 16 quoted lines
> > I think the work-around indicates the need for regular users to be
> > able to also discover, prune and delete these logs.  Do we have
> > "prune reflog for _this_ ref (or these refs), removing entries that
> > are older than this threshold"?  If so the codepath would need to
> > know about the graveyard and the implementation detail of the tilde
> > suffix so that the end users do not need to know about them.
> 
> We do have it: "git reflog expire --expire=now deleted-branch" is the
> right way to do it. Unfortunately, it does not work with my patch. The
> dwim_log correctly notes that a reflog exists (because it checks that
> the "graveyard" version of the ref exists), but then expire_reflog does
> not correctly fallback when opening the log (it usually has to do the
> _reverse_ translation, because it gets the graveyard log name from
> for_each_reflog, and has to find the correct lock).
> 
> I'll fix it in my re-roll, and then have t7701 use it.

I noticed I ignored the "discover" and "delete" parts of your paragraph. As far as deletion goes, I think we can ignore it; expiring all entries is equivalent.

Discovery is harder. Certainly these should not show up in normal ref-listing output. I'd be content to leave them slightly hidden as a first step, and people who know they are looking for the pre-deletion contents of the "foo" branch can access it by name. Probably a second step would be a fancier interface to help with listing and resurrecting dead branches, possibly including branch config.

In other words, I want to focus on getting the ref-level plumbing right, and then we can care about the porcelain later.

-Peff
Junio C Hamano· Jul 20, 2012, 15:39 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

Jeff King <peff@peff.net> writes:
Show 6 quoted lines
> I noticed I ignored the "discover" and "delete" parts of your paragraph.
> As far as deletion goes, I think we can ignore it; expiring all entries
> is equivalent.
> ...
> In other words, I want to focus on getting the ref-level plumbing right,
> and then we can care about the porcelain later.

Yeah, I agree that is a reasonable way forward. for-each-ref with a new option (--include-dead or something) can wait.

Junio C Hamano· Jul 20, 2012, 15:42 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

Jeff King <peff@peff.net> writes:
Show 10 quoted lines
> But it would mean that you cannot naively run
>
>   echo $sha1 >.git/refs/heads/foo
>
> anymore. I suspect that the packed-refs conversion rooted out many
> scripts that did not use update-ref and rev-parse to access refs, but
> the above does still work today. So I suspect there would be some
> fallout. Not to mention that older versions of git would be completely
> broken, which would mean we need a lengthy deprecation period while
> everybody upgrades to versions of git that support the reading side.

We have that "core.repositoryversion" thing, so we could treat it just like "update-index --index-version 4" to make it a "flag day event for each repository, on the day of end-user's choice".

Jeff King· Jul 20, 2012, 15:50 UTC · re: Junio C Hamano · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On Fri, Jul 20, 2012 at 08:42:57AM -0700, Junio C Hamano wrote:
Show 16 quoted lines
> Jeff King <peff@peff.net> writes:
> 
> > But it would mean that you cannot naively run
> >
> >   echo $sha1 >.git/refs/heads/foo
> >
> > anymore. I suspect that the packed-refs conversion rooted out many
> > scripts that did not use update-ref and rev-parse to access refs, but
> > the above does still work today. So I suspect there would be some
> > fallout. Not to mention that older versions of git would be completely
> > broken, which would mean we need a lengthy deprecation period while
> > everybody upgrades to versions of git that support the reading side.
> 
> We have that "core.repositoryversion" thing, so we could treat it
> just like "update-index --index-version 4" to make it a "flag day
> event for each repository, on the day of end-user's choice".

True. The code to handle both cases would be pretty nasty, though, mostly because we do not isolate the filesystem calls at all right now (i.e., there are a lot of calls to git_path("logs/%s", refname) in the code. Which is probably not too bad, but there are a lot of implicit reverse-conversions (e.g., walking the hierarchy and assuming that the path you find is a refname).

If we are seriously considering doing this for the full refs namespace anytime soon, then I'd be tempted to hold off the reflog graveyard until then. The code would be a lot simpler and less error-prone if we didn't have to convert between the namespaces (you would simply not get the reflog retention behavior in the old repositoryformatversion).

-Peff
Junio C Hamano· Aug 16, 2012, 23:29 UTC · re: Junio C Hamano · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

Junio C Hamano <gitster@pobox.com> writes:
Show 5 quoted lines
> I like the general direction.  Perhaps a long distant future
> direction could be to also use the same trick in the ref namespace
> so that we can have 'next' branch itself, and 'next/foo', 'next/bar'
> forks that are based on the 'next' branch at the same time (it
> obviously is a totally unrelated topic)?

I notice that I was responsible for making this topic veer in the wrong direction by bringing up a new feature "having 'next' and 'next/bar' at the same time" which nobody asked. Perhaps we can drop that for now to simplify the scope of the topic, to bring the log graveyard back on track?

Michael Haggerty· Jul 20, 2012, 09:49 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On 07/19/2012 11:33 PM, Jeff King wrote:
Show 7 quoted lines
> [...]
> This cannot be done by simply leaving the reflog files in
> place. The ref namespace does not allow D/F conflicts, so a
> ref "foo" would block the creation of another ref "foo/bar",
> and vice versa. This limitation is acceptable for two refs
> to exist simultaneously, but should not have an impact if
> one of the refs is deleted.
This is a great feature.
Show 18 quoted lines
> This patch moves reflog entries into a special "graveyard"
> namespace, and appends a tilde (~) character, which is
> not allowed in a valid ref name. This means that the deleted
> reflogs of these refs:
>
>     refs/heads/a
>     refs/heads/a/b
>     refs/heads/a/b/c
>
> will be stored in:
>
>     logs/graveyard/refs/heads/a~
>     logs/graveyard/refs/heads/a/b~
>     logs/graveyard/refs/heads/a/b/c~
>
> Putting them in the graveyard namespace ensures they will
> not conflict with live refs, and the tilde prevents D/F
> conflicts within the graveyard namespace.

I agree with Junio that long-term, it would be nice to allow references "foo" and "foo/bar" to exist simultaneously. To get there, we would have to redesign the mapping between reference names and the filenames used for the references and for the reflogs.

The easiest thing would be to mark files and directories differently; something like

     $GIT_DIR/{,logs/}refs/heads/a/b/c~
or
     $GIT_DIR/{,logs/}refs/heads~/a~/b~/c

i.e., munging either directory or file names to strings that are illegal in refnames such that it is unambiguous from the name whether a path is a file or directory.

And *if* we did that, then we wouldn't need a separate "graveyard" namespace, would we? The reflogs for dead references could live among those for living references.

Therefore, I think it would be good if we would choose a convention now for dead reflogs that is compatible with this hoped-for future.

The first convention, "logs/refs/heads/a/b/c~" is not usable because a reflog for a dead reference with this name would conflict with a reflog for a live reference "heads/a" or "heads/a/b" that uses the current filename convention.

But the second convention, "logs/refs/heads~/a~/b~/c, cannot conflict with current reflog files. And it would be a step towards allowing "foo" and "foo/bar" at the same time. What do you think about using a convention like this instead of the one that you proposed?

Another minor concern is the choice of trailing tilde in the file or directory names. Given that emacs creates backup files by appending a tilde to the filename, (1) it would be easy to inadvertently create such files, which git might try to interpret as reflogs and (2) there might be tools that innately "know" to skip such files in their processing. ack-grep, a replacement for grep, is an example that springs to mind. I know that I have written backup scripts that ignore files matching "*~", and a garbage-removal script that removes files matching "*~". Probably it is less precarious to name directories rather than files with trailing tildes, but either one could be a surprise for sysadmins.

Other possibilities (according to git-check-ref-format(1)):
     refs/.heads/.a/.b/c
     refs/heads./a./b./c (problematic on some Windows filesystems?)
     refs/heads../a../b../c
     refs/heads~dir/a~dir/b~dir/c (or some other suffix)
     refs/heads..a..b..c (not recommended because it flattens directory 
hierarchy)
Show 8 quoted lines
> The implementation is fairly straightforward, but it's worth
> noting a few things:
>
>    1. Updates to "logs/graveyard/refs/heads/foo~" happen
>       under the ref-lock for "refs/heads/foo". So deletion
>       still takes a single lock, and anyone touching the
>       reflog directly needs to reverse the transformation to
>       find the correct lockfile.
This should be documented in the code.
Show 5 quoted lines
>    2. We append entries to the graveyard reflog rather than
>       simply renaming the file into place. This means that
>       if you create and delete a branch repeatedly, the
>       graveyard will contain the concatenation of all
>       iterations.
Good.
>    3. We do not resurrect dead entries when a new ref is
>       created with the same name. However, it would be
>       possible to build an "undelete" feature on top of this
>       if one was so inclined.
Nice prospect.
Show 29 quoted lines
> [...]> diff --git a/refs.c b/refs.c
> index da74a2b..553de77 100644
> --- a/refs.c
> +++ b/refs.c
> [...]
> @@ -2552,3 +2553,63 @@ char *shorten_unambiguous_ref(const char *refname, int strict)
>   	free(short_name);
>   	return xstrdup(refname);
>   }
> +
> +char *refname_to_graveyard_reflog(const char *ref)
> +{
> +	return git_path("logs/graveyard/%s~", ref);
> +}
> +
> +char *graveyard_reflog_to_refname(const char *log)
> +{
> +	static struct strbuf buf = STRBUF_INIT;
> +
> +	if (!prefixcmp(log, "graveyard/"))
> +		log += 10;
> +
> +	strbuf_reset(&buf);
> +	strbuf_addstr(&buf, log);
> +	if (buf.len > 0 && buf.buf[buf.len-1] == '~')
> +		strbuf_setlen(&buf, buf.len - 1);
> +
> +	return buf.buf;
> +}

Given the names of these two functions, I was surprised that they aren't inverses of each other.

Function comments would be nice, too, especially for the latter.
Michael
-- 
Michael Haggerty
mhagger@alum.mit.edu
http://softwareswirl.blogspot.com/
Jeff King· Jul 20, 2012, 15:44 UTC · re: Michael Haggerty · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On Fri, Jul 20, 2012 at 11:49:07AM +0200, Michael Haggerty wrote:
Show 23 quoted lines
> >This patch moves reflog entries into a special "graveyard"
> >namespace, and appends a tilde (~) character, which is
> >not allowed in a valid ref name. This means that the deleted
> >reflogs of these refs:
> >
> >    refs/heads/a
> >    refs/heads/a/b
> >    refs/heads/a/b/c
> >
> >will be stored in:
> >
> >    logs/graveyard/refs/heads/a~
> >    logs/graveyard/refs/heads/a/b~
> >    logs/graveyard/refs/heads/a/b/c~
> >
> >Putting them in the graveyard namespace ensures they will
> >not conflict with live refs, and the tilde prevents D/F
> >conflicts within the graveyard namespace.
> 
> I agree with Junio that long-term, it would be nice to allow
> references "foo" and "foo/bar" to exist simultaneously.  To get
> there, we would have to redesign the mapping between reference names
> and the filenames used for the references and for the reflogs.

Yes, I would really like that, as it could make the alternate namespace go away, which is the source of about half the code in my patches (i.e., we would only need to loosen the reflog reading code to handle reflogs that do not have a matching ref).

But I fear that the fallouts from that will be much, much larger. Even with just this change, older versions of git will be slightly unhappy (e.g., you will get some extra warnings during fsck and reflog expiration about these reflogs). But changing the on-disk representation of the refs namespace will mean a totally new representation of locking. That's going to break old versions of git completely, and possibly even some user scripts.

Show 9 quoted lines
> The easiest thing would be to mark files and directories differently;
> something like
> 
>     $GIT_DIR/{,logs/}refs/heads/a/b/c~
> [...]
> The first convention, "logs/refs/heads/a/b/c~" is not usable because
> a reflog for a dead reference with this name would conflict with a
> reflog for a live reference "heads/a" or "heads/a/b" that uses the
> current filename convention.

Right. That's what I started with, then created the graveyard hierarchy to avoid conflicts between the "old" namespace (that cannot handle D/F conflicts) and the "new" one (that can, because it represents files and directories differently).

Show 7 quoted lines
> or
> 
>     $GIT_DIR/{,logs/}refs/heads~/a~/b~/c
> 
> i.e., munging either directory or file names to strings that are
> illegal in refnames such that it is unambiguous from the name whether
> a path is a file or directory.

This one can have conflicts in the opposite direction if you don't have any directories. E.g., you have $GIT_DIR/foo, a deleted ref, which has no tildes because it has no directories in the path. But you want to create foo/bar under the "old" system, which cannot happen (under the new system, it is fine, but the point of this exercise is to overlay the old and new systems).

That may be an OK tradeoff. We are restrictive in what goes into the top-level. Although I notice that you did not mark "refs" in the above example. So you could have the same problem with "refs/stash", for example. Again, though, we don't tend to have arbitrary data at the top-level (and I think refs/stash gets special cased in a couple places already). So it might be an acceptable limitation.

If we want to be pedantic, my patch causes conflicts for top-level refs called "graveyard" (although I know we have talked about restricting top-level refs to [A-Z_-], I don't recall if that has actually happened).

> And *if* we did that, then we wouldn't need a separate "graveyard"
> namespace, would we?  The reflogs for dead references could live
> among those for living references.

Right, assuming the limitation above is OK. But note that it doesn't really save us any code. We still have to convert between refnames and graveyard versions. _Eventually_ if the refnames were all converted, that code could go away.

> But the second convention, "logs/refs/heads~/a~/b~/c, cannot conflict
> with current reflog files.  And it would be a step towards allowing
> "foo" and "foo/bar" at the same time.  What do you think about using
> a convention like this instead of the one that you proposed?

I think it's reasonable. As I said, it doesn't save any code _now_, but since I am pulling a convention out of thin air, it might as well be one that has a possibility of converging in the future (all other things being equal, of course; I do find marking the directories a little uglier to read, but that is mostly because of the tilde).

Show 7 quoted lines
> Another minor concern is the choice of trailing tilde in the file or
> directory names.  Given that emacs creates backup files by appending
> a tilde to the filename, (1) it would be easy to inadvertently create
> such files, which git might try to interpret as reflogs and (2) there
> might be tools that innately "know" to skip such files in their
> processing. ack-grep, a replacement for grep, is an example that
> springs to mind.

The use of "~" for backup files was actually something that made me choose it, since these are, after all, backups of the reflog. But they are probably more precious than editor backup files, so the special treatment they're given by other programs is probably not desirable.

Show 8 quoted lines
> Other possibilities (according to git-check-ref-format(1)):
> 
>     refs/.heads/.a/.b/c
>     refs/heads./a./b./c (problematic on some Windows filesystems?)
>     refs/heads../a../b../c
>     refs/heads~dir/a~dir/b~dir/c (or some other suffix)
>     refs/heads..a..b..c (not recommended because it flattens
> directory hierarchy)

I don't like leading-dot, because those files are also often skipped by directory traversal of some programs (and certainly they are confusing to work with if you try to use "ls" to debug your $GIT_DIR/logs directory). Trailing dot is less ugly to me, but I do wonder about its special meaning as an extension separator. Double-dots just look gross.

Note that we have a few other magic characters available, too. Colon is probably the least offensive (metacharacters like *, ?, and [ just make things unnecessarily painful for shell users).

So I think a suffix like ":d" is probably the least horrible.
-Peff
Johannes Sixt· Jul 20, 2012, 16:37 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

Am 20.07.2012 17:44, schrieb Jeff King:
> So I think a suffix like ":d" is probably the least horrible.

Not so. It does not work on Windows :-( in the expected way. Trying to open a file with a colon-separated suffix either opens a resource fork on NTFS or fails with "invalid path".

-- Hannes
Jeff King· Jul 20, 2012, 17:09 UTC · re: Johannes Sixt · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On Fri, Jul 20, 2012 at 06:37:02PM +0200, Johannes Sixt wrote:
Show 6 quoted lines
> Am 20.07.2012 17:44, schrieb Jeff King:
> > So I think a suffix like ":d" is probably the least horrible.
> 
> Not so. It does not work on Windows :-( in the expected way. Trying to
> open a file with a colon-separated suffix either opens a resource fork
> on NTFS or fails with "invalid path".

Bleh. It seems that we did too good a job in coming up with a list of disallowed ref characters; they really are things you don't want in your filenames at all. :)

-Peff
Alexey Muranov· Jul 22, 2012, 11:03 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On 20 Jul 2012, at 19:09, Jeff King wrote:
Show 12 quoted lines
> On Fri, Jul 20, 2012 at 06:37:02PM +0200, Johannes Sixt wrote:
> 
>> Am 20.07.2012 17:44, schrieb Jeff King:
>>> So I think a suffix like ":d" is probably the least horrible.
>> 
>> Not so. It does not work on Windows :-( in the expected way. Trying to
>> open a file with a colon-separated suffix either opens a resource fork
>> on NTFS or fails with "invalid path".
> 
> Bleh. It seems that we did too good a job in coming up with a list of
> disallowed ref characters; they really are things you don't want in your
> filenames at all. :)
How about using '@' as an escape character ?
-Alexey.
Nguyen Thai Ngoc Duy· Jul 26, 2012, 12:47 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On Sat, Jul 21, 2012 at 12:09 AM, Jeff King <peff@peff.net> wrote:
Show 12 quoted lines
> On Fri, Jul 20, 2012 at 06:37:02PM +0200, Johannes Sixt wrote:
>
>> Am 20.07.2012 17:44, schrieb Jeff King:
>> > So I think a suffix like ":d" is probably the least horrible.
>>
>> Not so. It does not work on Windows :-( in the expected way. Trying to
>> open a file with a colon-separated suffix either opens a resource fork
>> on NTFS or fails with "invalid path".
>
> Bleh. It seems that we did too good a job in coming up with a list of
> disallowed ref characters; they really are things you don't want in your
> filenames at all. :)

So we haven't found any way to present both branches "foo" and "foo/bar" on file system at the same time. How about when we a new branch introduces such a conflict, we push the new branch directly to packed-refs? If we need either of them on a separate file, for fast update for example, then we unpack just one and repack all refs that conflict with it. Attempting to update two conflict branches in parallel may impact performance, but I don't think that happens often.

-- 
Duy
Alexey Muranov· Jul 26, 2012, 16:26 UTC · re: Nguyen Thai Ngoc Duy · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On 26 Jul 2012, at 14:47, Nguyen Thai Ngoc Duy wrote:
Show 9 quoted lines
> So we haven't found any way to present both branches "foo" and
> "foo/bar" on file system at the same time. How about when we a new
> branch introduces such a conflict, we push the new branch directly to
> packed-refs? If we need either of them on a separate file, for fast
> update for example, then we unpack just one and repack all refs that
> conflict with it. Attempting to update two conflict branches in
> parallel may impact performance, but I don't think that happens often.
> -- 
> Duy
How about simply deprecating "/" in branch name?
-Alexey.
Matthieu Moy· Jul 26, 2012, 16:41 UTC · re: Alexey Muranov · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

Alexey Muranov <alexey.muranov@gmail.com> writes:
Show 13 quoted lines
> On 26 Jul 2012, at 14:47, Nguyen Thai Ngoc Duy wrote:
>
>> So we haven't found any way to present both branches "foo" and
>> "foo/bar" on file system at the same time. How about when we a new
>> branch introduces such a conflict, we push the new branch directly to
>> packed-refs? If we need either of them on a separate file, for fast
>> update for example, then we unpack just one and repack all refs that
>> conflict with it. Attempting to update two conflict branches in
>> parallel may impact performance, but I don't think that happens often.
>> -- 
>> Duy
>
> How about simply deprecating "/" in branch name?

Err, it's not like nobody's using this feature (Junio does a heavy use of it in particular) ...

-- 
Matthieu Moy
http://www-verimag.imag.fr/~moy/
Jeff King· Jul 26, 2012, 16:59 UTC · re: Matthieu Moy · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On Thu, Jul 26, 2012 at 06:41:09PM +0200, Matthieu Moy wrote:
> > How about simply deprecating "/" in branch name?
> 
> Err, it's not like nobody's using this feature (Junio does a heavy use
> of it in particular) ...

Not to mention git itself, as it splits up the refs/remotes hierarchy into subdirectories. I think deprecating "/" is out of the question.

-Peff
Alexey Muranov· Jul 26, 2012, 17:24 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On 26 Jul 2012, at 18:59, Jeff King wrote:
> Not to mention git itself, as it splits up the refs/remotes hierarchy
> into subdirectories. I think deprecating "/" is out of the question.
> 
> -Peff
Ok, i guess you know better than me, my vision of Git is probably still too simplistic.
-Alexey.
Junio C Hamano· Jul 26, 2012, 17:46 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

Jeff King <peff@peff.net> writes:
Show 12 quoted lines
> On Fri, Jul 20, 2012 at 06:37:02PM +0200, Johannes Sixt wrote:
>
>> Am 20.07.2012 17:44, schrieb Jeff King:
>> > So I think a suffix like ":d" is probably the least horrible.
>> 
>> Not so. It does not work on Windows :-( in the expected way. Trying to
>> open a file with a colon-separated suffix either opens a resource fork
>> on NTFS or fails with "invalid path".
>
> Bleh. It seems that we did too good a job in coming up with a list of
> disallowed ref characters; they really are things you don't want in your
> filenames at all. :)

Why do no need to even worry about ~ vs : vs whatever in the first place?

With a flag-day per repository "core.repositoryformatversion = 1", you do not have to worry about mixture of old-style refs and new ones, so refs/heads/next-d/log could be a topic branch 'next/log' that is based on an integration branch 'next' branch that physically resides at refs/heads/next-f or an entry refs/heads/next in packed refs. Only the API functions in refs.c should care, no?

Jeff King· Jul 26, 2012, 17:52 UTC · re: Junio C Hamano · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On Thu, Jul 26, 2012 at 10:46:01AM -0700, Junio C Hamano wrote:
Show 13 quoted lines
> > Bleh. It seems that we did too good a job in coming up with a list of
> > disallowed ref characters; they really are things you don't want in your
> > filenames at all. :)
> 
> Why do no need to even worry about ~ vs : vs whatever in the first
> place?
> 
> With a flag-day per repository "core.repositoryformatversion = 1",
> you do not have to worry about mixture of old-style refs and new
> ones, so refs/heads/next-d/log could be a topic branch 'next/log'
> that is based on an integration branch 'next' branch that physically
> resides at refs/heads/next-f or an entry refs/heads/next in packed
> refs.  Only the API functions in refs.c should care, no?

I think the point was that Michael wanted to select a standard that could be used for graveyard reflogs _now_, but which would eventually match the format we use for active refs. And that requires a character that is not valid in a refname.

Given that the change of format for actives refs would require a flag day, keeping the graveyard scheme mixable with the current ref rules may not be worth caring about, though.

-Peff
Alexey Muranov· Jul 22, 2012, 11:10 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On 20 Jul 2012, at 17:44, Jeff King wrote:
Show 20 quoted lines
> On Fri, Jul 20, 2012 at 11:49:07AM +0200, Michael Haggerty wrote:
> 
>>> This patch moves reflog entries into a special "graveyard"
>>> namespace, and appends a tilde (~) character, which is
>>> not allowed in a valid ref name. This means that the deleted
>>> reflogs of these refs:
>>> 
>>>   refs/heads/a
>>>   refs/heads/a/b
>>>   refs/heads/a/b/c
>>> 
>>> will be stored in:
>>> 
>>>   logs/graveyard/refs/heads/a~
>>>   logs/graveyard/refs/heads/a/b~
>>>   logs/graveyard/refs/heads/a/b/c~
>>> 
>>> Putting them in the graveyard namespace ensures they will
>>> not conflict with live refs, and the tilde prevents D/F
>>> conflicts within the graveyard namespace.
Sorry if this idea is stupid or if i miss something, but how about putting deleted reflogs for

refs/heads/a refs/heads/a/b refs/heads/a/b/c

to

refs/heads/a@yyyy-mm-dd-hhmmss refs/heads/a/b@yyyy-mm-dd-hhmmss refs/heads/a/b/c@yyyy-mm-dd-hhmmss

with the time they were deleted?
-Alexey.
Alexey Muranov· Jul 22, 2012, 11:12 UTC · re: Alexey Muranov · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On 22 Jul 2012, at 13:10, Alexey Muranov wrote:
Show 15 quoted lines
> Sorry if this idea is stupid or if i miss something, but how about putting deleted reflogs for
> 
> refs/heads/a
> refs/heads/a/b
> refs/heads/a/b/c
> 
> to
> 
> refs/heads/a@yyyy-mm-dd-hhmmss
> refs/heads/a/b@yyyy-mm-dd-hhmmss
> refs/heads/a/b/c@yyyy-mm-dd-hhmmss
> 
> with the time they were deleted?
> 
> -Alexey.
Sorry, i meant to:

logs/refs/heads/a@yyyy-mm-dd-hhmmss logs/refs/heads/a/b@yyyy-mm-dd-hhmmss logs/refs/heads/a/b/c@yyyy-mm-dd-hhmmss

-Alexey.
Jeff King· Jul 22, 2012, 13:14 UTC · re: Alexey Muranov · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On Sun, Jul 22, 2012 at 01:10:55PM +0200, Alexey Muranov wrote:
Show 27 quoted lines
> >>>   refs/heads/a
> >>>   refs/heads/a/b
> >>>   refs/heads/a/b/c
> >>> 
> >>> will be stored in:
> >>> 
> >>>   logs/graveyard/refs/heads/a~
> >>>   logs/graveyard/refs/heads/a/b~
> >>>   logs/graveyard/refs/heads/a/b/c~
> >>> 
> >>> Putting them in the graveyard namespace ensures they will
> >>> not conflict with live refs, and the tilde prevents D/F
> >>> conflicts within the graveyard namespace.
> 
> Sorry if this idea is stupid or if i miss something, but how about putting deleted reflogs for
> 
> refs/heads/a
> refs/heads/a/b
> refs/heads/a/b/c
> 
> to
> 
> refs/heads/a@yyyy-mm-dd-hhmmss
> refs/heads/a/b@yyyy-mm-dd-hhmmss
> refs/heads/a/b/c@yyyy-mm-dd-hhmmss
> 
> with the time they were deleted?

I like the readability of the resulting file names, but it has three problems:

  1. "@" is allowed in ref names, so you may be conflicting with
     existing refs. You could fix that by using "@{...}", which is
     disallowed. E.g., refs/heads/a@{yyyy-mm-dd-hhmmss}.
  2. It makes lookup slightly more expensive, because to find a reflog
     for "refs/heads/a", I have to scan "logs/refs/heads" looking for
     any matching entries of the form "a@{.*}". This is probably not a
     huge deal in practice, though it does make the code more complex.
  3. Most importantly, it does not resolve D/F conflicts (it has the
     same problem as "logs/refs/heads/a~"). If you delete "foo/bar", you
     will end up with "logs/refs/heads/foo/bar@{...}". That will prevent
     D/F conflicts with a new branch "foo/bar/baz", but will still have
     a problem with just "foo".
     You need to either mark each directory to avoid the conflict
     (Michael suggested something like "refs/heads~/foo~/bar"), or
     you need to put the deleted logs into a separate hierarchy (I used
     "logs/graveyard" in my patch).
-Peff
Alexey Muranov· Jul 22, 2012, 14:40 UTC · re: Jeff King · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On 22 Jul 2012, at 15:14, Jeff King wrote:
Show 5 quoted lines
>  3. Most importantly, it does not resolve D/F conflicts (it has the
>     same problem as "logs/refs/heads/a~"). If you delete "foo/bar", you
>     will end up with "logs/refs/heads/foo/bar@{...}". That will prevent
>     D/F conflicts with a new branch "foo/bar/baz", but will still have
>     a problem with just "foo".
Unfortunately i do not really follow this, because i have not seen any directories in "logs/refs/heads/", i only saw files named after local branches there. I do not know how directories are used there.
-Alexey.
Jeff King· Jul 22, 2012, 15:50 UTC · re: Alexey Muranov · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

On Sun, Jul 22, 2012 at 04:40:14PM +0200, Alexey Muranov wrote:
Show 9 quoted lines
> >  3. Most importantly, it does not resolve D/F conflicts (it has the
> >     same problem as "logs/refs/heads/a~"). If you delete "foo/bar", you
> >     will end up with "logs/refs/heads/foo/bar@{...}". That will prevent
> >     D/F conflicts with a new branch "foo/bar/baz", but will still have
> >     a problem with just "foo".
> 
> Unfortunately i do not really follow this, because i have not seen any
> directories in "logs/refs/heads/", i only saw files named after local
> branches there. I do not know how directories are used there.

The user is free to have branch names with slashes, in which case they are represented in the filesystem as directories. Even without using slashes in your branch names, you already have subdirectories in refs/remotes.

-Peff
Johannes Sixt· Jul 20, 2012, 16:32 UTC · re: Michael Haggerty · lore

Re: [PATCH 1/3] retain reflogs for deleted refs

Am 20.07.2012 11:49, schrieb Michael Haggerty:
> Other possibilities (according to git-check-ref-format(1)):
> 
>     refs/.heads/.a/.b/c
>     refs/heads./a./b./c (problematic on some Windows filesystems?)
Yes. Probably all filesystems.
>     refs/heads../a../b../c
Same here.
>     refs/heads~dir/a~dir/b~dir/c (or some other suffix)
>     refs/heads..a..b..c (not recommended because it flattens directory
> hierarchy)
-- Hannes
Jeff King· Jul 19, 2012, 21:33 UTC · re: Jeff King · lore

[PATCH 2/3] teach sha1_name to look in graveyard reflogs

The previous commit introduced graveyard reflogs, where the reflog for a deleted branch "foo" appears in "logs/graveyard/refs/heads/foo~".

This patch teaches dwim_log to search for these logs if the ref does not exist, and teaches read_ref_at to fall back to them when the literal reflog does not exist. This allows "deleted@{1}" to refer to the final commit of a deleted branch (either to view or to re-create the branch). You can also go further back, or refer to the deleted reflog entries by time. Accessing deleted@{0} will yield the null sha1.

Similarly, for_each_reflog_ent learns to fallback to graveyard refs, which allows the reflog walker to work. However, this is slightly less friendly, as the revision parser expects the matching ref to exist before it realizes that we are interested in the reflog. Therefore you must use "git log -g deleted@{1}" insted of "git log -g deleted" to walk a deleted reflog.

In both cases, we also tighten up the mode-checking when opening the reflogs. dwim_log checks that the entry we found is a regular file (not a directory) to avoid D/F confusion (e.g., you ask for "foo" but "foo/bar" exists and we find the "foo" but it is a directory).

However, read_ref_at and for_each_reflog_ent did not do this check, and relied on earlier parts of the code to have verified the log they are about to open. This meant that even before this patch, a race condition in changing refs between dwim_log and the actual read could cause bizarre errors (e.g., read_ref_at would open and try to mmap a directory). This patch makes it even easier to trigger those conditions (because the ref namespace and the fallback graveyard namespace can have D/F ambiguity for a certain path). To solve this, we check the mode of the file we open and treat it as if it did not exist if it is not a regular file (this is the same way dwim_log handles it).

Signed-off-by: Jeff King <peff@peff.net>
---
 refs.c | 46 +++++++++++++++++++++++++++++++++++-----------
 1 file changed, 35 insertions(+), 11 deletions(-)
diff --git a/refs.c b/refs.c
index 553de77..551a0f9 100644
--- a/refs.c
+++ b/refs.c
@@ -1590,9 +1590,16 @@ int dwim_log(const char *str, int len, unsigned char *sha1, char **log)
 
 		mksnpath(path, sizeof(path), *p, len, str);
 		ref = resolve_ref_unsafe(path, hash, 1, NULL);
-		if (!ref)
-			continue;
-		if (!stat(git_path("logs/%s", path), &st) &&
+		if (!ref) {
+			if (!stat(refname_to_graveyard_reflog(path), &st) &&
+			    S_ISREG(st.st_mode)) {
+				it = path;
+				hashcpy(hash, null_sha1);
+			}
+			else
+				continue;
+		}
+		else if (!stat(git_path("logs/%s", path), &st) &&
 		    S_ISREG(st.st_mode))
 			it = path;
 		else if (strcmp(ref, path) &&
@@ -2201,9 +2208,16 @@ int read_ref_at(const char *refname, unsigned long at_time, int cnt,
 
 	logfile = git_path("logs/%s", refname);
 	logfd = open(logfile, O_RDONLY, 0);
-	if (logfd < 0)
-		die_errno("Unable to read log '%s'", logfile);
-	fstat(logfd, &st);
+	if (logfd < 0 || fstat(logfd, &st) < 0 || !S_ISREG(st.st_mode)) {
+		const char *deleted_log = refname_to_graveyard_reflog(refname);
+
+		if (logfd >= 0)
+			close(logfd);
+		logfd = open(deleted_log, O_RDONLY);
+		if (logfd < 0 || fstat(logfd, &st) < 0 || !S_ISREG(st.st_mode))
+			die_errno("Unable to read log '%s'", logfile);
+		logfile = deleted_log;
+	}
 	if (!st.st_size)
 		die("Log %s is empty.", logfile);
 	mapsz = xsize_t(st.st_size);
@@ -2296,18 +2310,28 @@ int for_each_recent_reflog_ent(const char *refname, each_reflog_ent_fn fn, long
 {
 	const char *logfile;
 	FILE *logfp;
+	struct stat st;
 	struct strbuf sb = STRBUF_INIT;
 	int ret = 0;
 
 	logfile = git_path("logs/%s", refname);
 	logfp = fopen(logfile, "r");
-	if (!logfp)
-		return -1;
+	if (!logfp || fstat(fileno(logfp), &st) < 0 || !S_ISREG(st.st_mode)) {
+		logfile = refname_to_graveyard_reflog(refname);
+
+		if (logfp)
+			fclose(logfp);
+		logfp = fopen(logfile, "r");
+		if (!logfp)
+			return -1;
+		if (fstat(fileno(logfp), &st) < 0 || !S_ISREG(st.st_mode)) {
+			fclose(logfp);
+			return -1;
+		}
+	}
 
 	if (ofs) {
-		struct stat statbuf;
-		if (fstat(fileno(logfp), &statbuf) ||
-		    statbuf.st_size < ofs ||
+		if (st.st_size < ofs ||
 		    fseek(logfp, -ofs, SEEK_END) ||
 		    strbuf_getwholeline(&sb, logfp, '\n')) {
 			fclose(logfp);
-- 
1.7.10.5.40.g059818d
Junio C Hamano· Jul 19, 2012, 22:39 UTC · re: Jeff King · lore

Re: [PATCH 2/3] teach sha1_name to look in graveyard reflogs

Jeff King <peff@peff.net> writes:
Show 40 quoted lines
> The previous commit introduced graveyard reflogs, where the
> reflog for a deleted branch "foo" appears in
> "logs/graveyard/refs/heads/foo~".
>
> This patch teaches dwim_log to search for these logs if the
> ref does not exist, and teaches read_ref_at to fall back to
> them when the literal reflog does not exist.  This allows
> "deleted@{1}" to refer to the final commit of a deleted
> branch (either to view or to re-create the branch).  You can
> also go further back, or refer to the deleted reflog entries
> by time. Accessing deleted@{0} will yield the null sha1.
>
> Similarly, for_each_reflog_ent learns to fallback to
> graveyard refs, which allows the reflog walker to work.
> However, this is slightly less friendly, as the revision
> parser expects the matching ref to exist before it realizes
> that we are interested in the reflog. Therefore you must use
> "git log -g deleted@{1}" insted of "git log -g deleted" to
> walk a deleted reflog.
>
> In both cases, we also tighten up the mode-checking when
> opening the reflogs. dwim_log checks that the entry we found
> is a regular file (not a directory) to avoid D/F confusion
> (e.g., you ask for "foo" but "foo/bar" exists and we find
> the "foo" but it is a directory).
>
> However, read_ref_at and for_each_reflog_ent did not do this
> check, and relied on earlier parts of the code to have
> verified the log they are about to open. This meant that
> even before this patch, a race condition in changing refs
> between dwim_log and the actual read could cause bizarre
> errors (e.g., read_ref_at would open and try to mmap a
> directory). This patch makes it even easier to trigger those
> conditions (because the ref namespace and the fallback
> graveyard namespace can have D/F ambiguity for a certain
> path). To solve this, we check the mode of the file we open
> and treat it as if it did not exist if it is not a regular
> file (this is the same way dwim_log handles it).
>
> Signed-off-by: Jeff King <peff@peff.net>

This may or may not be related, but I vaguely recall that "log -g" traversal hack had a corner case where the walking stops prematurely upon seeing a gap (or creation/deletion that has 0{40})? Do you recall if we have ever dealt with that?

The patch seems fine from a cursory look.  Thanks.
Jeff King· Jul 20, 2012, 15:53 UTC · re: Junio C Hamano · lore

Re: [PATCH 2/3] teach sha1_name to look in graveyard reflogs

On Thu, Jul 19, 2012 at 03:39:24PM -0700, Junio C Hamano wrote:
Show 12 quoted lines
> > Similarly, for_each_reflog_ent learns to fallback to
> > graveyard refs, which allows the reflog walker to work.
> > However, this is slightly less friendly, as the revision
> > parser expects the matching ref to exist before it realizes
> > that we are interested in the reflog. Therefore you must use
> > "git log -g deleted@{1}" insted of "git log -g deleted" to
> > walk a deleted reflog.
> 
> This may or may not be related, but I vaguely recall that "log -g"
> traversal hack had a corner case where the walking stops prematurely
> upon seeing a gap (or creation/deletion that has 0{40})?  Do you
> recall if we have ever dealt with that?
>From my tests, I think it is probably still broken (if you do a delete,

create, delete sequence on a branch and then walk the reflog, it stops prematurely at the 0{40} sha1).

But what _should_ it show for such an entry? There is no commit to show in the reflog walker, but it would still be nice to say "BTW, there was a deletion even here". Obviously just skipping it and showing the next entry would be better than the current behavior of stopping the traversal, but I feel like there must be some better behavior.

-Peff
Junio C Hamano· Jul 22, 2012, 20:53 UTC · re: Jeff King · lore

Re: [PATCH 2/3] teach sha1_name to look in graveyard reflogs

Jeff King <peff@peff.net> writes:
Show 5 quoted lines
> But what _should_ it show for such an entry? There is no commit to show
> in the reflog walker, but it would still be nice to say "BTW, there was
> a deletion even here". Obviously just skipping it and showing the next
> entry would be better than the current behavior of stopping the
> traversal, but I feel like there must be some better behavior.

Like showing an entry that says "Ref deleted here", which should be easy to do by creating a phoney commit object and inserting it to the queue the reflog walker uses, I would guess.

Jeff King· Jul 19, 2012, 21:33 UTC · re: Jeff King · lore

[PATCH 3/3] add tests for reflogs of deleted refs

These tests cover the basic functionality of retaining reflogs for deleted refs.

Signed-off-by: Jeff King <peff@peff.net>
---
 t/t1413-reflog-deletion.sh | 74 ++++++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 74 insertions(+)
 create mode 100755 t/t1413-reflog-deletion.sh
diff --git a/t/t1413-reflog-deletion.sh b/t/t1413-reflog-deletion.sh
new file mode 100755
index 0000000..e00d038
--- /dev/null
+++ b/t/t1413-reflog-deletion.sh
@@ -0,0 +1,74 @@
+#!/bin/sh
+
+test_description='test retention of reflog after ref deletion'
+. ./test-lib.sh
+
+test_expect_success 'setup deleted branch' '
+	test_tick && echo one >file && git add file && git commit -m one &&
+	test_tick && echo two >file && git add file && git commit -m two &&
+	git checkout -b foo/bar &&
+	test_tick && echo three >file && git add file && git commit -m three &&
+	git checkout master &&
+	git branch -D foo/bar &&
+	rm -f .git/logs/HEAD
+'
+
+test_expect_success 'branch is no longer accessible' '
+	test_must_fail git rev-parse --verify foo/bar
+'
+
+test_expect_success 'final reflog is null sha1' '
+	echo $_z40 >expect &&
+	git rev-parse --verify foo/bar@{0} >actual &&
+	test_cmp expect actual
+'
+
+test_expect_success 'deleted reflog entries are accessible' '
+	cat >expect <<-\EOF &&
+	three
+	two
+	EOF
+	{
+		git log -1 --format=%s foo/bar@{1}
+		git log -1 --format=%s foo/bar@{2}
+	} >actual &&
+	test_cmp expect actual
+'
+
+test_expect_success 'reflog walker can find deleted entries' '
+	cat >expect <<-\EOF &&
+	three
+	two
+	EOF
+	git log -g --format=%s foo/bar@{1} >actual &&
+	test_cmp expect actual
+'
+
+test_expect_success 'can still create/delete same ref' '
+	git branch foo/bar &&
+	git branch -D foo/bar
+'
+
+test_expect_success 'can still create/delete parent ref' '
+	git branch foo &&
+	git branch -D foo
+'
+
+test_expect_success 'can still create/delete child ref' '
+	git branch foo/bar/baz &&
+	git branch -D foo/bar/baz
+'
+
+test_expect_success 'deleted reflog entries are still reachable' '
+	>expect &&
+	git fsck --unreachable >actual &&
+	test_cmp expect actual
+'
+
+test_expect_success 'deleted reflog entries are expired normally' '
+	git reflog expire --all --expire=now &&
+	git fsck --unreachable >actual &&
+	test_line_count = 3 actual
+'
+
+test_done
-- 
1.7.10.5.40.g059818d

← back to recent threads