threads / discuss / 5438

Re: Problematic git pack

Subject: Re: Problematic git pack

## tl;dr

4 messages between Aug 31, 2006 and Aug 31, 2006.

replies: 3people: 4as markdown or json

Sergio Callegari· Aug 31, 2006, 08:45 UTC · lore

What can I say... I had never seen before such an action at such a rapid pace following the indication of a potential problem. Thanks Linus and Junio and everybody who might have contributed.

>   Junio could then generate a new pack with the one corrupted object 
>   fixed, which obviously meant that all the deltas now worked too.
>   
Excellent news...
Show 9 quoted lines
>   This is my (probably final) analysis of the resulting differences.. ]
>
> On Wed, 30 Aug 2006, Junio C Hamano wrote:
> > 
> > Ok, I was going to attach the resurrected pack that should
> > contain everything your corrupt pack had, but it is a bit too
> > large, so I'll place it here [*1*].  Drop me a note when you
> > retrieved it, so that I can remove it.
>   

Junio, can you please send me privately details about [*1*] so I can retrieve the pack also?

I also have another question... (maybe it was answered in some previous thread on this list, in this case a pointer would be enough). Now I am going to have the fixed archive and also a new archive, which I restarted from the latest working copy I had of my project. Is there any way to automatically do real "surgery" to attach one to the other and get a single archive with all the history? Obviously, if I try to change a commit object to modify its parents, its signature changes, so I need to modify its childs and so on, is this correct? Alternatively I belive that grafts should be a way to go... I had never used them before, do all git tools support them? Particularly do they get pushed and pulled correctly?

Show 5 quoted lines
> So the _real_ difference is literally just the one byte at offset 0151000 
> (decimal 53760) which in the fixed pack is 0x96, and in the corrupt pack 
> it is 0x94. That's a single-bit difference (bit #1 has been cleared).
>
>   

So, possibly, the alpha particle theory could be the plausible one in the end...

Show 5 quoted lines
> Now, that makes me feel happy on one level, because it's almost certainly 
> a hardware problem - subtle memory corruption, or disk corruption that 
> happened when either reading or writing the image. Sergio may not be that 
> happy about it, of course.
>   

The bad thing is that I don't know which of my two machines (the laptop or the desktop) caused the issue!

> Finally, this also points out that the corrupted packs _can_ be fixed, but 
> I think Sergio was a bit lucky (to offset all the bad luck). Sergio still 
> had access to the original file that had had its object corrupted. 

Actually, this could possibly be a not so rare case... In my tree I had the development of some LaTeX documents and packages (code like, the really "precious" files) and a few binary objects (images and openoffice files mainly, by far less precious). Since the binary objects were so much overwhelming in size with regard to the text ones, assuming a single error the probability of having it in a non-code object was much larger than that of having it in a precious code object. Also commit and tree objects should be much smaller than data objects. This assumption is the reason which initally pushed me to ask help to try to unpack at least all the correct objects (one of my first questions was: does git unpack-objects die on the first error or is there a way to convince it to simply skip the wrong object (or the delta against a wrong object)... If git unpack-objects can gain an option like --continue-on-errors and if checkout/reset can also get an option to do the same (i.e. in a tree with missing objects, checkout all that can be found), I believe that one is at a good point already... Finally, having a command to create an object out of a single file (contrary of git cat-file) could help re-creating the missing objects...

Show 11 quoted lines
> And it 
> took a fair amount of work, and some git hacking by somebody who really 
> understood git (Junio).
>
> Maybe we'll end up having some of that effort being useful and checked in, 
> and we'll eventually have more infrastructure for fixing these things, but 
> I suspect that in most cases, even a _single_ bit of corruption will 
> generally result in so much havoc that nobody should depend on that. It's 
> a lot better to have backups.
>
> 			Linus
Johannes Schindelin· Aug 31, 2006, 11:15 UTC · re: Sergio Callegari · lore
Hi,
On Thu, 31 Aug 2006, Sergio Callegari wrote:
> Now I am going to have the fixed archive and also a new archive, which I
> restarted from the latest working copy I had of my project.
> Is there any way to automatically do real "surgery" to attach one to the other
> and get a single archive with all the history?
You can "graft" the new onto the old branch:

If <40-hex-chars-old> is the commit id of the youngest commit of the reconstructed branch, and <40-hex-chars-new> is the commit id of the initial commit of the newly started branch, you can put this line into .git/info/grafts:

<40-hex-chars-new> <40-hex-chars-old>

This will make git believe that the initial commit is no initial commit, but has the old head as single parent. And yes, AFAICT all git tools support this. I used this technique many times to be able to merge unrelated developments.

NOTE! This is the quickest way if you want to have the history _locally_.

If you want to be able to distribute it (or synchronize it between your laptop and PC _with git!_), you can rewrite the history by either git-rebase, or by using cg-admin-rewritehist if you are using cogito.

Unfortunately, I do not use cogito nor git-rebase, so if you want to walk that path, others have to help. (And most likely, we'd put the result into Documentation/howto/.)

Ciao, Dscho

P.S.: Of course, if you do not insist on a super clean history, you can fake a merge. Just put <40-hex-chars-old> into .git/MERGE_HEAD and commit. This will pretend that your new head and your old head were merged, and the result is the new head. This _should_ even work with git-bisect, but it is slightly ugly.

Linus Torvalds· Aug 31, 2006, 21:33 UTC · re: Sergio Callegari · lore
On Thu, 31 Aug 2006, Sergio Callegari wrote:
>
> Junio, can you please send me privately details about [*1*] so I can retrieve
> the pack also?
He already did, search for "members.cox.net" in your email archive (it's 
Message-ID: <7v7j0qihwl.fsf@assigned-by-dhcp.cox.net> to be precise).
Show 6 quoted lines
> I also have another question... (maybe it was answered in some previous thread
> on this list, in this case a pointer would be enough).
> Now I am going to have the fixed archive and also a new archive, which I
> restarted from the latest working copy I had of my project.
> Is there any way to automatically do real "surgery" to attach one to the other
> and get a single archive with all the history?
Yes. This is just what a "grafts" file is for.

Put the old pack/idx files into the .git/objects/packs directory, and then you can create "fake parenthood" information in a ".git/info/grafts" file by just adding text-lines of the format "<sha1> <fakeparentsha1>" (with each SHA being the regular 40-byte hex representation).

Show 5 quoted lines
> Obviously, if I try to change a commit object to modify its parents, its
> signature changes, so I need to modify its childs and so on, is this correct?
> Alternatively I belive that grafts should be a way to go... I had never used
> them before, do all git tools support them? Particularly do they get pushed
> and pulled correctly?

Nope, they won't get pushed and pulled correctly, you need to put the grafts files in all repositories. Alternatively, you can re-create the whole history, I think cogito had some history re-writing tool.

Show 6 quoted lines
> > So the _real_ difference is literally just the one byte at offset 0151000
> > (decimal 53760) which in the fixed pack is 0x96, and in the corrupt pack it
> > is 0x94. That's a single-bit difference (bit #1 has been cleared).
> 
> So, possibly, the alpha particle theory could be the plausible one in the
> end...

Yes. It's just that Junio's original theory required it to not just hit a memory cell, it also had to hit it at _just_ the right time in between being written and the SHA1 of the buffer being computed. So the original theory was very unlikely indeed.

My theory of the corruption just causing a re-computed SHA1 when repacking (and silently copying the corruption without realizing it) meant that there was no such small and unlikely window, but that any regular memory (or disk) corruption could easily have caused it at any time, and then a subsequent re-pack "fixed" the SHA1 to match the corruption..

> The bad thing is that I don't know which of my two machines (the laptop or the
> desktop) caused the issue!

I'd suggest running memtest86 for a few days on both (not necessarily at the same time - keep one working machine to do you job on ;)

Show 8 quoted lines
> > Finally, this also points out that the corrupted packs _can_ be fixed, but I
> > think Sergio was a bit lucky (to offset all the bad luck). Sergio still had
> > access to the original file that had had its object corrupted. 
>
> Actually, this could possibly be a not so rare case... In my tree I had the
> development of some LaTeX documents and packages (code like, the really
> "precious" files) and a few binary objects (images and openoffice files
> mainly, by far less precious).

Sure. In your case you had checked in generated files too, and yes, they were the larger ones. That's not true in general - in many other projects, the _directory_ structure (ie the git "tree" objects) will be a large portion of the project, and probably more likely to be corrupt. Now, to some degree the tree objects are likely the ones easiest to "repair" (because you can try to look at the history and figure things out by hand), but at the same time, people also tend to have deeper delta-chains and it would just be _very_ painful.

So I do think you were somewhat lucky.
> Finally, having a command to create an object out of a single file (contrary
> of git cat-file) could help re-creating the missing objects...
Hmm. Like "git-hash-object"?
			Linus

← back to recent threads