threads / discuss / 43863

Working with zip files

Subject: Working with zip files

## tl;dr

14 messages between Aug 16, 2016 and Aug 19, 2016.

replies: 13people: 5as markdown or json

Nikolaus Rath· Aug 16, 2016, 16:25 UTC · lore
Hello,

I would like to store Simulink models in a Git repository. Unfortunately, the file format is binary. But luckily, the binary format happens to be a zipfile containing nicely formatted XML files.

Is there a way to teach Git to take advantage of this when storing, diff-ing and merging these files?

Best, -Nikolaus

-- 
GPG encrypted emails preferred. Key id: 0xD113FCAC3C4E599F
Fingerprint: ED31 791B 2C5C 1613 AF38 8B8A D113 FCAC 3C4E 599F

             »Time flies like an arrow, fruit flies like a Banana.«
David Lang· Aug 16, 2016, 16:27 UTC · re: Nikolaus Rath · lore

Re: Working with zip files

On Tue, 16 Aug 2016, Nikolaus Rath wrote:
Show 7 quoted lines
> I would like to store Simulink models in a Git
> repository. Unfortunately, the file format is binary. But luckily, the
> binary format happens to be a zipfile containing nicely formatted XML
> files.
>
> Is there a way to teach Git to take advantage of this when storing,
> diff-ing and merging these files?

you should be able to use clean/smudge to have git store the files uncompressed, which will help a lot.

I think there's a way to tell it to do a xml aware diff/patch, but I don't remember how.

David Lang
Nikolaus Rath· Aug 16, 2016, 16:32 UTC · re: David Lang · lore

Re: Working with zip files

On Aug 16 2016, David Lang <david@lang.hm> wrote:
Show 12 quoted lines
> On Tue, 16 Aug 2016, Nikolaus Rath wrote:
>
>> I would like to store Simulink models in a Git
>> repository. Unfortunately, the file format is binary. But luckily, the
>> binary format happens to be a zipfile containing nicely formatted XML
>> files.
>>
>> Is there a way to teach Git to take advantage of this when storing,
>> diff-ing and merging these files?
>
> you should be able to use clean/smudge to have git store the files
> uncompressed, which will help a lot.
Cool, I'll look into that.
> I think there's a way to tell it to do a xml aware diff/patch, but I
> don't remember how.

Oh, I didn't even want to go that far. I'm perfectly happy if it does a text-based diff/patch of the contained XML files. Would clean/smudge provide that already?

Best, -Nikolaus

-- 
GPG encrypted emails preferred. Key id: 0xD113FCAC3C4E599F
Fingerprint: ED31 791B 2C5C 1613 AF38 8B8A D113 FCAC 3C4E 599F

             »Time flies like an arrow, fruit flies like a Banana.«
David Lang· Aug 16, 2016, 16:48 UTC · re: Nikolaus Rath · lore

Re: Working with zip files

On Tue, 16 Aug 2016, Nikolaus Rath wrote:
Show 22 quoted lines
> On Aug 16 2016, David Lang <david@lang.hm> wrote:
>> On Tue, 16 Aug 2016, Nikolaus Rath wrote:
>>
>>> I would like to store Simulink models in a Git
>>> repository. Unfortunately, the file format is binary. But luckily, the
>>> binary format happens to be a zipfile containing nicely formatted XML
>>> files.
>>>
>>> Is there a way to teach Git to take advantage of this when storing,
>>> diff-ing and merging these files?
>>
>> you should be able to use clean/smudge to have git store the files
>> uncompressed, which will help a lot.
>
> Cool, I'll look into that.
>
>> I think there's a way to tell it to do a xml aware diff/patch, but I
>> don't remember how.
>
> Oh, I didn't even want to go that far. I'm perfectly happy if it does a
> text-based diff/patch of the contained XML files. Would clean/smudge
> provide that already?
yes.
Junio C Hamano· Aug 16, 2016, 16:58 UTC · re: David Lang · lore

Re: Working with zip files

David Lang <david@lang.hm> writes:
Show 5 quoted lines
> you should be able to use clean/smudge to have git store the files
> uncompressed, which will help a lot.
>
> I think there's a way to tell it to do a xml aware diff/patch, but I
> don't remember how.

I do not know about "patch" (in the sense of "git apply"), but "git diff" (and "git log -p") can take advantage of the clean/smudge mechanism. I used to deal with a file format that is gzipped xml so my clean filter was "gzip -dc" while the smudge was "gzip -cn". Essentially, this stors the xml before compression in the repository so blobs delta well with each other and also the revisions are made textually diff-able.

Nikolaus's case has one extra layer of complexity in that the "file" is actually an archive of multiple files. The clean/smudge pair he writes need to be a filter that flattens the archive into a single human-readable text byte stream and its reverse.

Jakub Narębski· Aug 16, 2016, 19:56 UTC · re: Junio C Hamano · lore

Re: Working with zip files

W dniu 16.08.2016 o 18:58, Junio C Hamano pisze:
> David Lang <david@lang.hm> writes:
> 
>> you should be able to use clean/smudge to have git store the files
>> uncompressed, which will help a lot.

You can find rezip clean/smudge filter (originally intended for OpenDocument Format (ODF), that is OpenOffice.org etc.) that stores zip or zip-archive (like ODT, jar, etc.) uncompressed. I think you can find it on GitWiki, but I might be mistaken.

Show 15 quoted lines
>> I think there's a way to tell it to do a xml aware diff/patch, but I
>> don't remember how.
> 
> I do not know about "patch" (in the sense of "git apply"), but "git
> diff" (and "git log -p") can take advantage of the clean/smudge
> mechanism.  I used to deal with a file format that is gzipped xml so
> my clean filter was "gzip -dc" while the smudge was "gzip -cn".
> Essentially, this stores the xml before compression in the repository
> so blobs delta well with each other and also the revisions are
> made textually diff-able.
> 
> Nikolaus's case has one extra layer of complexity in that the "file"
> is actually an archive of multiple files.  The clean/smudge pair he
> writes need to be a filter that flattens the archive into a single
> human-readable text byte stream and its reverse.

There is also `textconv` filter that can be used instead; it might be 'unzip -c' (extract files to stdout, with filenames), or 'unzip -p' (same, without filenames).

-- 
Jakub Narębski
Junio C Hamano· Aug 16, 2016, 20:19 UTC · re: Jakub Narębski · lore

Re: Working with zip files

Jakub Narębski <jnareb@gmail.com> writes:
> There is also `textconv` filter that can be used instead; it might
> be 'unzip -c' (extract files to stdout, with filenames), or 'unzip -p'
> (same, without filenames).

That assumes that the in-repository data is zipped binary blob; the result won't delta well, will it?

Jakub Narębski· Aug 18, 2016, 12:16 UTC · re: Junio C Hamano · lore

Re: Working with zip files

W dniu 16.08.2016 o 22:19, Junio C Hamano pisze:
Show 8 quoted lines
> Jakub Narębski <jnareb@gmail.com> writes:
> 
>> There is also `textconv` filter that can be used instead; it might
>> be 'unzip -c' (extract files to stdout, with filenames), or 'unzip -p'
>> (same, without filenames).
> 
> That assumes that the in-repository data is zipped binary blob; the
> result won't delta well, will it?

Full solution would involve `clean` filter to rezip with no compression (which should delta well) and optional `smudge` filter to recompress; if round-trip bit-for-bit equality is needed, the original zip parameters must be saved somewhere, e.g. as ZIP archive comments. This was mentioned in the earlier part of my email (which might have been not clear):

JN>> You can find rezip clean/smudge filter (originally intended for
JN>> OpenDocument Format (ODF), that is OpenOffice.org etc.) that stores
JN>> zip or zip-archive (like ODT, jar, etc.) uncompressed.  I think
JN>> you can find it on GitWiki, but I might be mistaken.
 
Using 'unzip -c' as separate / additional `textconv` filter for diff
generation allows to separate the problem of deltifiable storage format
from textual representation for diff-ing.
Though best results could be had with `diff` and `merge` drivers...
-- 
Jakub Narębski
David Lang· Aug 18, 2016, 16:56 UTC · re: Jakub Narębski · lore

Re: Working with zip files

On Thu, 18 Aug 2016, Jakub Narębski wrote:
Show 10 quoted lines
> JN>> You can find rezip clean/smudge filter (originally intended for
> JN>> OpenDocument Format (ODF), that is OpenOffice.org etc.) that stores
> JN>> zip or zip-archive (like ODT, jar, etc.) uncompressed.  I think
> JN>> you can find it on GitWiki, but I might be mistaken.
>
> Using 'unzip -c' as separate / additional `textconv` filter for diff
> generation allows to separate the problem of deltifiable storage format
> from textual representation for diff-ing.
>
> Though best results could be had with `diff` and `merge` drivers...

can you point at an example of how to do this? when I went looking about a year ago to deal with single-line json data I wasn't able to find anything good. I ended up using clean/smudge to pretty-print the json so it was easier to handle.

David Lang
Jakub Narębski· Aug 18, 2016, 17:45 UTC · re: David Lang · lore

Re: Working with zip files

On 18 August 2016 at 18:56, David Lang <david@lang.hm> wrote:
Show 18 quoted lines
> On Thu, 18 Aug 2016, Jakub Narębski wrote:
>
>> JN>> You can find rezip clean/smudge filter (originally intended for
>> JN>> OpenDocument Format (ODF), that is OpenOffice.org etc.) that stores
>> JN>> zip or zip-archive (like ODT, jar, etc.) uncompressed.  I think
>> JN>> you can find it on GitWiki, but I might be mistaken.
>>
>> Using 'unzip -c' as separate / additional `textconv` filter for diff
>> generation allows to separate the problem of deltifiable storage format
>> from textual representation for diff-ing.
>>
>> Though best results could be had with `diff` and `merge` drivers...
>
>
> can you point at an example of how to do this? when I went looking about a
> year ago to deal with single-line json data I wasn't able to find anything
> good. I ended up using clean/smudge to pretty-print the json so it was
> easier to handle.

Pro Git has a chapter "Customizing Git - Git Attributes" about gitattributes https://git-scm.com/book/en/v2/Customizing-Git-Git-Attributes

The section "Diffing Binary Files" has two examples: docx2txt (with wrapper) for DOCX (MS Word) files, and exiftool for images. For JSON you could use some prettyprinter / formatter like pp-json.

"Performing text diffs of binary files" section of gitattributes(1) manpage covers 'textconv' vs 'diff', and uses 'exif' tool as textconv example.

HTH
-- 
Jakub Narębski




-- 
Jakub Narebski
David Lang· Aug 19, 2016, 03:00 UTC · re: Jakub Narębski · lore

Re: Working with zip files

On Thu, 18 Aug 2016, Jakub Narębski wrote:
Show 29 quoted lines
> On 18 August 2016 at 18:56, David Lang <david@lang.hm> wrote:
>> On Thu, 18 Aug 2016, Jakub Narębski wrote:
>>
>>> JN>> You can find rezip clean/smudge filter (originally intended for
>>> JN>> OpenDocument Format (ODF), that is OpenOffice.org etc.) that stores
>>> JN>> zip or zip-archive (like ODT, jar, etc.) uncompressed.  I think
>>> JN>> you can find it on GitWiki, but I might be mistaken.
>>>
>>> Using 'unzip -c' as separate / additional `textconv` filter for diff
>>> generation allows to separate the problem of deltifiable storage format
>>> from textual representation for diff-ing.
>>>
>>> Though best results could be had with `diff` and `merge` drivers...
>>
>>
>> can you point at an example of how to do this? when I went looking about a
>> year ago to deal with single-line json data I wasn't able to find anything
>> good. I ended up using clean/smudge to pretty-print the json so it was
>> easier to handle.
>
> Pro Git has a chapter "Customizing Git - Git Attributes" about gitattributes
> https://git-scm.com/book/en/v2/Customizing-Git-Git-Attributes
>
> The section "Diffing Binary Files" has two examples: docx2txt (with wrapper)
> for DOCX (MS Word) files, and exiftool for images. For JSON you could use
> some prettyprinter / formatter like pp-json.
>
> "Performing text diffs of binary files" section of gitattributes(1) manpage
> covers 'textconv' vs 'diff', and uses 'exif' tool as textconv example.

As I read that section, it only applies to the human readable output of git diff.

And the merge section only talks about the default of using patch vs accepting a specific version in a merge.

It seems to me that what I'm looking for would be something to tell git to use a different command instead of diff/patch internally when creating and using the bundles.

David Lang
Nikolaus Rath· Aug 16, 2016, 21:14 UTC · re: David Lang · lore

Re: Working with zip files

On Aug 16 2016, David Lang <david@lang.hm> wrote:
Show 12 quoted lines
> On Tue, 16 Aug 2016, Nikolaus Rath wrote:
>
>> I would like to store Simulink models in a Git
>> repository. Unfortunately, the file format is binary. But luckily, the
>> binary format happens to be a zipfile containing nicely formatted XML
>> files.
>>
>> Is there a way to teach Git to take advantage of this when storing,
>> diff-ing and merging these files?
>
> you should be able to use clean/smudge to have git store the files
> uncompressed, which will help a lot.
Having looked at that, I'm not sure if this really helps:

As I understand, the smudge command is run on checkout to convert the blob in the repository to the format that is desired in the working tree. But this is the opposite of what I need: on checkout, I need to convert the text data in the repository to a blob in the working tree.

Furthermore, I need to convert multiple text files into one blob, will smudge/clean seem to do just 1:1 conversions.

Am I missing something? Are there any other options?

Best, Nikolaus

-- 
GPG encrypted emails preferred. Key id: 0xD113FCAC3C4E599F
Fingerprint: ED31 791B 2C5C 1613 AF38 8B8A D113 FCAC 3C4E 599F

             »Time flies like an arrow, fruit flies like a Banana.«
Jacob Keller· Aug 17, 2016, 05:31 UTC · re: Nikolaus Rath · lore

Re: Working with zip files

On Tue, Aug 16, 2016 at 2:14 PM, Nikolaus Rath <Nikolaus@rath.org> wrote:
Show 25 quoted lines
> On Aug 16 2016, David Lang <david@lang.hm> wrote:
>> On Tue, 16 Aug 2016, Nikolaus Rath wrote:
>>
>>> I would like to store Simulink models in a Git
>>> repository. Unfortunately, the file format is binary. But luckily, the
>>> binary format happens to be a zipfile containing nicely formatted XML
>>> files.
>>>
>>> Is there a way to teach Git to take advantage of this when storing,
>>> diff-ing and merging these files?
>>
>> you should be able to use clean/smudge to have git store the files
>> uncompressed, which will help a lot.
>
> Having looked at that, I'm not sure if this really helps:
>
> As I understand, the smudge command is run on checkout to convert the
> blob in the repository to the format that is desired in the working
> tree. But this is the opposite of what I need: on checkout, I need to
> convert the text data in the repository to a blob in the working tree.
>
> Furthermore, I need to convert multiple text files into one blob, will
> smudge/clean seem to do just 1:1 conversions.
>
> Am I missing something? Are there any other options?

You want to store the contents of the zip file as *one* blob that is the uncompressed contents of the archive somehow concatenated together. That should still be a 1:1 relationship.

You won't store one blob per file in the zip.

Thanks, Jake

David Lang· Aug 17, 2016, 09:58 UTC · re: Nikolaus Rath · lore

Re: Working with zip files

On Tue, 16 Aug 2016, Nikolaus Rath wrote:
Show 25 quoted lines
> On Aug 16 2016, David Lang <david@lang.hm> wrote:
>> On Tue, 16 Aug 2016, Nikolaus Rath wrote:
>>
>>> I would like to store Simulink models in a Git
>>> repository. Unfortunately, the file format is binary. But luckily, the
>>> binary format happens to be a zipfile containing nicely formatted XML
>>> files.
>>>
>>> Is there a way to teach Git to take advantage of this when storing,
>>> diff-ing and merging these files?
>>
>> you should be able to use clean/smudge to have git store the files
>> uncompressed, which will help a lot.
>
> Having looked at that, I'm not sure if this really helps:
>
> As I understand, the smudge command is run on checkout to convert the
> blob in the repository to the format that is desired in the working
> tree. But this is the opposite of what I need: on checkout, I need to
> convert the text data in the repository to a blob in the working tree.
>
> Furthermore, I need to convert multiple text files into one blob, will
> smudge/clean seem to do just 1:1 conversions.
>
> Am I missing something? Are there any other options?

so the smudge command would zip the file and the clean command would unzip the file (assuming it's a single file, if the zip is multiple files, you will have to add something to combine them)

you want the working tree to have a zip file and the repository to have text.
David Lang

← back to recent threads