threads / discuss / 5381

Re: Problem with pack

Subject: Re: Problem with pack

## tl;dr

6 messages between Aug 25, 2006 and Aug 26, 2006.

replies: 5people: 4as markdown or json

Sergio Callegari· Aug 25, 2006, 10:07 UTC · lore
Show 12 quoted lines
>
> > git verify-pack -v pack-ebcdfbbda07e5a3e4136aa1f499990b35685bab4.idx
> > fatal: failed to read delta-pack base object 2849bd2bd8a76bbca37df2a4c8e8b990811d01a7
>
> Eeeh! Not good.
>
> > 1) I am working on both a pc and a notebook, syncing the two everytime I move
> > from one to the other.
>
> So, you still have one "good" version? Please make a backup immediately. 
> (If only to reproduce the problem.)
>   

I have a good working tree, but unfortunately I realized that there was a problem with the pack only _after_ the sync: I was not expecting this kind of problem, so I silly did a repack as the last thing, I went home, I attached the laptop to the net, I run unison, I started to work and I realized that there was a problem when I attempted a new repack which failed complaining about the corrupted pack...

So actually, I do not even know where the corruption came from (an hd error, the sync tool, ...)

I only have the corrupted pack and its index and a good last working tree.

BTW, it would be nice to have some "security measure" in git reset... e.g. an option to trigger the following behavior:

- saving all current changes in a temporary commit
- checking that the current HEAD can be re-checked out before the reset
Show 5 quoted lines
> Since unpack-objects does not use the index, it cannot extract anything 
> after the first error. We _could_ enhance unpack-objects to be nice and 
> optionally take a pack-index to try to reconstruct as many objects as 
> possible.
>   

That would be very useful... Btw, even without that, if I understand correctly, git packs are collections of compressed objects, each of which has its own header stating how long is the compressed object itself. In my case, the error is in inflating one object (git unpack-objects says inflate returns -3)... so shouldn't there be a way to try to skip to the next object even in this case?

Show 5 quoted lines
> BTW I'd recommend not syncing with unison, but with the git transports: If 
> your PC and Laptop are connected, you could do something like
>
> 	git pull laptop:my_project/.git
>   

Actually, the project, including the git archive gets syncronized as a part of a syncronization process including all my Documents directory (the project is in fact a LaTeX manual with somehow complex LaTeX packages and classes). Syncronizing in this way actually worked very well so far, because at once I was getting in sync all my working trees and all my repos...

Sergio
Andreas Ericsson· Aug 25, 2006, 10:20 UTC · re: Sergio Callegari · lore
Sergio Callegari wrote:
Show 32 quoted lines
>>
>> > git verify-pack -v pack-ebcdfbbda07e5a3e4136aa1f499990b35685bab4.idx
>> > fatal: failed to read delta-pack base object 
>> 2849bd2bd8a76bbca37df2a4c8e8b990811d01a7
>>
>> Eeeh! Not good.
>>
>> > 1) I am working on both a pc and a notebook, syncing the two 
>> everytime I move
>> > from one to the other.
>>
>> So, you still have one "good" version? Please make a backup 
>> immediately. (If only to reproduce the problem.)
>>   
> I have a good working tree, but unfortunately I realized that there was 
> a problem with the pack only _after_ the sync:
> I was not expecting this kind of problem, so I silly did a repack as the 
> last thing, I went home, I attached the laptop to the net, I run unison, 
> I started to work and I realized that there was a problem when I 
> attempted a new repack which failed complaining about the corrupted pack...
> 
> So actually, I do not even know where the corruption came from (an hd 
> error, the sync tool, ...)
> 
> I only have the corrupted pack and its index and a good last working tree.
> 
> BTW, it would be nice to have some "security measure" in git reset... 
> e.g. an option to trigger the following behavior:
> 
> - saving all current changes in a temporary commit
> - checking that the current HEAD can be re-checked out before the reset
> 

The recommended way is to do a throw-away branch to commit your temporary commit to (or the 'master' so long as you remember to use reset). The current HEAD can always, barring object database errors, be checked out if 'git status' reports no changes in the working tree. Unfortunately, the object database is often enormous, so doing a full fsck-objects before each change to the branch you're on (which is basically what a reset is), would take far too long to be viable.

Show 12 quoted lines
>> Since unpack-objects does not use the index, it cannot extract 
>> anything after the first error. We _could_ enhance unpack-objects to 
>> be nice and optionally take a pack-index to try to reconstruct as many 
>> objects as possible.
>>   
> That would be very useful...
> Btw, even without that, if I understand correctly, git packs are 
> collections of compressed objects, each of which has its own header 
> stating how long is the compressed object itself. In my case, the error 
> is in inflating one object (git unpack-objects says inflate returns 
> -3)... so shouldn't there be a way to try to skip to the next object 
> even in this case?

It should be possible, assuming the pack index is still intact. The pack index is where the headers are stored, afaik.

Show 13 quoted lines
>> BTW I'd recommend not syncing with unison, but with the git 
>> transports: If your PC and Laptop are connected, you could do 
>> something like
>>
>>     git pull laptop:my_project/.git
>>   
> Actually, the project, including the git archive gets syncronized as a 
> part of a syncronization process including all my Documents directory 
> (the project is in fact a LaTeX manual with somehow complex LaTeX 
> packages and classes). Syncronizing in this way actually worked very 
> well so far, because at once I was getting in sync all my working trees 
> and all my repos...
> 

The largest benefit of using git's synchronization methods is that you immediately get a pack-file verification, and also that you never risk overwriting anything in either repo if you've forgotten to sync between the two (say you've made changes on your laptop, forgot to send them to your workstation, then made changes on your workstation and then you try to sync them). It's possible to recover from such a situation using the lost-found tool, but it can be cumbersome, and uncommitted changes, as well as changes to the working tree, are lost forever.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Tel: +46 8-230225                  Fax: +46 8-230231
Jakub Narebski· Aug 25, 2006, 10:41 UTC · re: Andreas Ericsson · lore
Andreas Ericsson wrote:
Show 22 quoted lines
>>> BTW I'd recommend not syncing with unison, but with the git 
>>> transports: If your PC and Laptop are connected, you could do 
>>> something like
>>>
>>>     git pull laptop:my_project/.git
>>>   
>> Actually, the project, including the git archive gets syncronized as a 
>> part of a syncronization process including all my Documents directory 
>> (the project is in fact a LaTeX manual with somehow complex LaTeX 
>> packages and classes). Syncronizing in this way actually worked very 
>> well so far, because at once I was getting in sync all my working trees 
>> and all my repos...
>> 
> 
> The largest benefit of using git's synchronization methods is that you 
> immediately get a pack-file verification, and also that you never risk 
> overwriting anything in either repo if you've forgotten to sync between 
> the two (say you've made changes on your laptop, forgot to send them to 
> your workstation, then made changes on your workstation and then you try 
> to sync them). It's possible to recover from such a situation using the 
> lost-found tool, but it can be cumbersome, and uncommitted changes, as 
> well as changes to the working tree, are lost forever.

Unison (which if I remember correctly uses rsync, or rsync over ssh) detect such case and ask user what to do if both sides changed a file (copy from one side, copy from second side, view diff, merge,...).

But you can always tell unison to ignore git object database
  ignore = Name .git
and perhaps also ignore working directories under git control
  ignore = Path path/to/working/dir/
-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Junio C Hamano· Aug 25, 2006, 10:58 UTC · re: Andreas Ericsson · lore
Andreas Ericsson <ae@op5.se> writes:
Show 9 quoted lines
>> Btw, even without that, if I understand correctly, git packs are
>> collections of compressed objects, each of which has its own header
>> stating how long is the compressed object itself. In my case, the
>> error is in inflating one object (git unpack-objects says inflate
>> returns -3)... so shouldn't there be a way to try to skip to the
>> next object even in this case?
>
> It should be possible, assuming the pack index is still intact. The
> pack index is where the headers are stored, afaik.

The problem Sergio seems to be having is because somehow he does not have a base object that another object that is in the pack depends on, because the latter object is stored in deltified form.

This should never happen unless .pack itself is corrupted (git-pack-objects, unless explicitly told to do so with --thin flag to git-rev-list upstream, would not make a delta against objects not in the same pack).

When a delta is written to the pack file, unless its base object has already written out, git-pack-objects writes out the base object immediately after that deltified object. So one possibility is that the pack was truncated soon after the delta that is having trouble with finding its base object. In such a case, the proposed recovery measure of skipping the corruption and keep going would not buy you that much. On the other hand, if the corruption is in the middle (e.g. a single disk block was wiped out), having .idx file might help you resync.

Does the pack pass git-verify-pack test, I wonder?
Junio C Hamano· Aug 26, 2006, 10:09 UTC · re: Sergio Callegari · lore
Sergio Callegari <scallegari@arces.unibo.it> writes:
Show 5 quoted lines
> I was not expecting this kind of problem, so I silly did a repack as
> the last thing, I went home, I attached the laptop to the net, I run
> unison, I started to work and I realized that there was a problem when
> I attempted a new repack which failed complaining about the corrupted
> pack...

Sorry about the mixed "intended audience" of this message, asking Sergio for a bit more info as the end user who had problems with git, and at the same time describing the code level analysis of possible cause to ask for help from git developers. Nico CC'ed because he seems to be the person who knows the best around this area including zlib.

- - -

Earlier you said that the mothership has 1.4.2, and the note has 1.4.0. The sequence of events as I understand are:

	- repack -a -d on the mothership with 1.4.2; no problems
          observed.
        - transfer the results to note; this was done behind git so
          no problems observed.
        - tried to repack on note with 1.4.0; got "failed to
          read delta-pack base object" error.

Can you make the pack/idx available to the public for postmortem?

Also I wonder if the pack can be read by 1.4.2.

Earlier you said "unpack-objects <$that-pack.pack" fails with "error code -3 in inflate..." What exact error do you get? I am guessing that it is get_data() which says:

	"inflate returned %d\n"
(side note: we should not say \n there).
        static void *get_data(unsigned long size)
        {
                z_stream stream;
                void *buf = xmalloc(size);
                memset(&stream, 0, sizeof(stream));
                stream.next_out = buf;
                stream.avail_out = size;
                stream.next_in = fill(1);
                stream.avail_in = len;
                inflateInit(&stream);
                for (;;) {
                        int ret = inflate(&stream, 0);
                        use(len - stream.avail_in);
                        if (stream.total_out == size && ret == Z_STREAM_END)
                                break;
                        if (ret != Z_OK)
                                die("inflate returned %d\n", ret);
                        stream.next_in = fill(1);
                        stream.avail_in = len;
                }
                inflateEnd(&stream);
                return buf;
        }

This pattern appears practically everywhere. When inflate() returns, we expect its return value to be either Z_OK or Z_STREAM_END and everything else is treated as an error. Also when we receive Z_STREAM_END the resulting length had better be the size we expect. I do not have any problem with the latter but have always felt uneasy about the former but being no zlib expert had been using the code as is. zlib.h says Z_BUF_ERROR is not fatal -- it just means this round of call with the given input and output buffer did not make any progress. I've been wondering if it is possible for inflate to eat some input but that was not enough to produce one byte of output, and what would return in such a case...

About the error message you got while attempting to repack, the exact same error message appears in two places, but the one that is emitted is the one in sha1_file.c::unpack_delta_entry(). The other one is in unpack-objects.c and is not run by repack.

This function is called after:
	- we find that we need to access an object;
	- we decided to use the .pack/.idx pair; when mmap'ing
          the .idx file, we validate that the pair is not
          corrupt by calling check_packed_git_idx().  This does
          the sha-1 checksum of both files;
	- we find the location of the deltified object in the
          .pack file by looking at the corresponding .idx;
	- we read from that location a handful bytes, find that
          it is deltified and learn its base object name.

So it is not likely that the .pack/.idx pair was corrupted after it was written (i.e. not a bit rot).

Now, in unpack_delta_entry_function():
	- we find the location of its base object in the .pack
          file by looking at the .idx again; if this fails, we
          would die with a different error message "failed to
          find delta-pack base object";
	- we call unpack_entry_gently(); we would see the error
          message you saw only when this function returns NULL.

So unpack_entry_gently() while reading the base object returned NULL. Let's see how it can:

	- we read the data for the base object in the pack we
          identified earlier.  First we learn its type and size.
	- the base object could be also a deltified object, in
          which case it would recursively call unpack_delta_entry();
	  however, that function would never return NULL.  it
          either succeeds or die()s.  So the base must not have
          been OBJ_DELTA type.
        - the object type recorded there for the base object
          could have been something bogus, in which case we
          would return NULL.  But that would mean pack-object in
          1.4.2 generated a bogus pack on the mothership.
	- if the object type is not bogus, unpack_non_delta_entry()
          is called to extract the data for the base object.
          this decompresses the data stream, and if it does
          not inflate well it would return NULL.
So there are only a few ways you can get that error message.
	- pack-objects in 1.4.2 produced an invalid pack by
          recording:
          - bogus object type for the base object, or
          - incorrectly recording the offset of the base object in
            the pack file in .idx, or
	  - incorrectly recording the size of the base object in
            the pack file; 
 
 	  and checksumed the bogus resulting pack/idx pair as if
 	  nothing was wrong.
	- pack-objects in 1.4.2 produced a deflated stream that
          made unpack_non_delta_entry() unhappy.

Other changes between 1.4.0 and 1.4.2 that I do not think are related are:

	- we slightly changed the way data is deflated with
          commit 12f6c30 to favour speed over compression, but
          this only affects loose objects and not packs.
	- we introduced a new file format for loose objects with
          commit 93821bd, but this needs to be explicitly
          enabled by .git/config option.  Even if the mothership
          1.4.2 had recorded loose objects in the new format,
          the process to create the pack by first expanding and
          then recompressing with the old-and-proven code, so
          this should not affect the resulting pack.  Even if
          such a loose object were copied to the note with 1.4.0
          together with the pack, object reading code always
          favours what's in the pack, so it should not even be
          touched.  Actually, the codepath that emits the error
          message does not read the base object from anywhere
          other than from the same pack.
	- we slightly changed the way data for the object to be
          deltified and the base object is read in pack-objects
          with commit 560b25a; I did not see anything obviously
          wrong with that change.
        - we slightly changed the way we pick the base object
          when making a delta with commits 8dbbd14 and 51d1e83;
          these should not change what happens after a pair is
          decided to be used as delta and its base.
Junio C Hamano· Aug 26, 2006, 10:31 UTC · re: Junio C Hamano · lore
Junio C Hamano <junkio@cox.net> writes:
Show 13 quoted lines
> Earlier you said "unpack-objects <$that-pack.pack" fails with
> "error code -3 in inflate..."  What exact error do you get?
> I am guessing that it is get_data() which says:
>
> 	"inflate returned %d\n"
>
> (side note: we should not say \n there).
> ...
> This pattern appears practically everywhere...
> ...  I've been
> wondering if it is possible for inflate to eat some input but
> that was not enough to produce one byte of output, and what [it]
> would return in such a case...

I do not think this fear does not apply to this particular case; return value -3 is Z_DATA_ERROR, so the deflated stream is corrupt.

> So there are only a few ways you can get that error message.
> ...
I just realized there is another not so inplausible explanation.

When the problematic pack was made on the mothership, csum-file.c::sha1write_compressed() gave the data for the base object to zlib, an alpha particle hit a memory cell that contained zlib output buffer (resulting in a corrupt deflated stream in variable "out"), and sha1write() wrote it out while computing the right checksum.

Is the memory on your mothership reliable (I do not want to make this message sound like one on the kernel list, but memtest86 might be in order)?

← back to recent threads