{"thread":{"id":"43863","subject":"Working with zip files","startedAt":"2016-08-16T16:25:17Z","lastAt":"2016-08-19T07:21:45Z","messageCount":14,"participants":["Nikolaus Rath","David Lang","Junio C Hamano","Jakub Narębski","Jacob Keller"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"299475","messageId":"87y43wwujd.fsf@thinkpad.rath.org","threadId":"43863","inReplyTo":null,"subject":"Working with zip files","fromName":"Nikolaus Rath","fromEmail":"nikolaus@rath.org","sentAt":"2016-08-16T16:25:10Z","receivedAt":"2016-08-16T16:25:17Z","isPatch":false,"sender":{"key":"nikolaus@rath.org","avatar":"https://gravatar.com/avatar/d3b6cdc023a0665882d5aac373c0115bcb36f730c2e4a4051061fefbaf7dd041?d=mp&s=160"},"body":"Hello,\n\nI would like to store Simulink models in a Git\nrepository. Unfortunately, the file format is binary. But luckily, the\nbinary format happens to be a zipfile containing nicely formatted XML\nfiles.\n\nIs there a way to teach Git to take advantage of this when storing,\ndiff-ing and merging these files?\n\nBest,\n-Nikolaus\n\n-- \nGPG encrypted emails preferred. Key id: 0xD113FCAC3C4E599F\nFingerprint: ED31 791B 2C5C 1613 AF38 8B8A D113 FCAC 3C4E 599F\n\n             »Time flies like an arrow, fruit flies like a Banana.«\n"},{"id":"299476","messageId":"alpine.DEB.2.02.1608160926330.11774@nftneq.ynat.uz","threadId":"43863","inReplyTo":"87y43wwujd.fsf@thinkpad.rath.org","subject":"Re: Working with zip files","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2016-08-16T16:27:55Z","receivedAt":"2016-08-16T16:28:03Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Tue, 16 Aug 2016, Nikolaus Rath wrote:\n\n> I would like to store Simulink models in a Git\n> repository. Unfortunately, the file format is binary. But luckily, the\n> binary format happens to be a zipfile containing nicely formatted XML\n> files.\n>\n> Is there a way to teach Git to take advantage of this when storing,\n> diff-ing and merging these files?\n\nyou should be able to use clean/smudge to have git store the files uncompressed, \nwhich will help a lot.\n\nI think there's a way to tell it to do a xml aware diff/patch, but I don't \nremember how.\n\nDavid Lang\n"},{"id":"299480","messageId":"87shu4wu6o.fsf@thinkpad.rath.org","threadId":"43863","inReplyTo":"alpine.DEB.2.02.1608160926330.11774@nftneq.ynat.uz","subject":"Re: Working with zip files","fromName":"Nikolaus Rath","fromEmail":"nikolaus@rath.org","sentAt":"2016-08-16T16:32:47Z","receivedAt":"2016-08-16T16:43:06Z","isPatch":false,"sender":{"key":"nikolaus@rath.org","avatar":"https://gravatar.com/avatar/d3b6cdc023a0665882d5aac373c0115bcb36f730c2e4a4051061fefbaf7dd041?d=mp&s=160"},"body":"On Aug 16 2016, David Lang <david@lang.hm> wrote:\n> On Tue, 16 Aug 2016, Nikolaus Rath wrote:\n>\n>> I would like to store Simulink models in a Git\n>> repository. Unfortunately, the file format is binary. But luckily, the\n>> binary format happens to be a zipfile containing nicely formatted XML\n>> files.\n>>\n>> Is there a way to teach Git to take advantage of this when storing,\n>> diff-ing and merging these files?\n>\n> you should be able to use clean/smudge to have git store the files\n> uncompressed, which will help a lot.\n\nCool, I'll look into that.\n\n> I think there's a way to tell it to do a xml aware diff/patch, but I\n> don't remember how.\n\nOh, I didn't even want to go that far. I'm perfectly happy if it does a\ntext-based diff/patch of the contained XML files. Would clean/smudge\nprovide that already?\n\n\nBest,\n-Nikolaus\n\n-- \nGPG encrypted emails preferred. Key id: 0xD113FCAC3C4E599F\nFingerprint: ED31 791B 2C5C 1613 AF38 8B8A D113 FCAC 3C4E 599F\n\n             »Time flies like an arrow, fruit flies like a Banana.«\n"},{"id":"299481","messageId":"alpine.DEB.2.02.1608160948010.11774@nftneq.ynat.uz","threadId":"43863","inReplyTo":"87shu4wu6o.fsf@thinkpad.rath.org","subject":"Re: Working with zip files","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2016-08-16T16:48:39Z","receivedAt":"2016-08-16T16:48:47Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Tue, 16 Aug 2016, Nikolaus Rath wrote:\n\n> On Aug 16 2016, David Lang <david@lang.hm> wrote:\n>> On Tue, 16 Aug 2016, Nikolaus Rath wrote:\n>>\n>>> I would like to store Simulink models in a Git\n>>> repository. Unfortunately, the file format is binary. But luckily, the\n>>> binary format happens to be a zipfile containing nicely formatted XML\n>>> files.\n>>>\n>>> Is there a way to teach Git to take advantage of this when storing,\n>>> diff-ing and merging these files?\n>>\n>> you should be able to use clean/smudge to have git store the files\n>> uncompressed, which will help a lot.\n>\n> Cool, I'll look into that.\n>\n>> I think there's a way to tell it to do a xml aware diff/patch, but I\n>> don't remember how.\n>\n> Oh, I didn't even want to go that far. I'm perfectly happy if it does a\n> text-based diff/patch of the contained XML files. Would clean/smudge\n> provide that already?\n\nyes.\n"},{"id":"299483","messageId":"xmqqeg5oejmn.fsf@gitster.mtv.corp.google.com","threadId":"43863","inReplyTo":"alpine.DEB.2.02.1608160926330.11774@nftneq.ynat.uz","subject":"Re: Working with zip files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-08-16T16:58:08Z","receivedAt":"2016-08-16T16:58:51Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"David Lang <david@lang.hm> writes:\n\n> you should be able to use clean/smudge to have git store the files\n> uncompressed, which will help a lot.\n>\n> I think there's a way to tell it to do a xml aware diff/patch, but I\n> don't remember how.\n\nI do not know about \"patch\" (in the sense of \"git apply\"), but \"git\ndiff\" (and \"git log -p\") can take advantage of the clean/smudge\nmechanism.  I used to deal with a file format that is gzipped xml so\nmy clean filter was \"gzip -dc\" while the smudge was \"gzip -cn\".\nEssentially, this stors the xml before compression in the repository\nso blobs delta well with each other and also the revisions are\nmade textually diff-able.\n\nNikolaus's case has one extra layer of complexity in that the \"file\"\nis actually an archive of multiple files.  The clean/smudge pair he\nwrites need to be a filter that flattens the archive into a single\nhuman-readable text byte stream and its reverse.\n"},{"id":"299499","messageId":"34d64f4f-3cda-385c-cdce-5f1852d545e3@gmail.com","threadId":"43863","inReplyTo":"xmqqeg5oejmn.fsf@gitster.mtv.corp.google.com","subject":"Re: Working with zip files","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-08-16T19:56:08Z","receivedAt":"2016-08-16T19:56:31Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 16.08.2016 o 18:58, Junio C Hamano pisze:\n> David Lang <david@lang.hm> writes:\n> \n>> you should be able to use clean/smudge to have git store the files\n>> uncompressed, which will help a lot.\n\nYou can find rezip clean/smudge filter (originally intended for\nOpenDocument Format (ODF), that is OpenOffice.org etc.) that stores\nzip or zip-archive (like ODT, jar, etc.) uncompressed.  I think\nyou can find it on GitWiki, but I might be mistaken.\n\n>> I think there's a way to tell it to do a xml aware diff/patch, but I\n>> don't remember how.\n> \n> I do not know about \"patch\" (in the sense of \"git apply\"), but \"git\n> diff\" (and \"git log -p\") can take advantage of the clean/smudge\n> mechanism.  I used to deal with a file format that is gzipped xml so\n> my clean filter was \"gzip -dc\" while the smudge was \"gzip -cn\".\n> Essentially, this stores the xml before compression in the repository\n> so blobs delta well with each other and also the revisions are\n> made textually diff-able.\n> \n> Nikolaus's case has one extra layer of complexity in that the \"file\"\n> is actually an archive of multiple files.  The clean/smudge pair he\n> writes need to be a filter that flattens the archive into a single\n> human-readable text byte stream and its reverse.\n\nThere is also `textconv` filter that can be used instead; it might\nbe 'unzip -c' (extract files to stdout, with filenames), or 'unzip -p'\n(same, without filenames).\n\n-- \nJakub Narębski\n"},{"id":"299500","messageId":"xmqq8tvwcvrc.fsf@gitster.mtv.corp.google.com","threadId":"43863","inReplyTo":"34d64f4f-3cda-385c-cdce-5f1852d545e3@gmail.com","subject":"Re: Working with zip files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-08-16T20:19:03Z","receivedAt":"2016-08-16T20:19:12Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narębski <jnareb@gmail.com> writes:\n\n> There is also `textconv` filter that can be used instead; it might\n> be 'unzip -c' (extract files to stdout, with filenames), or 'unzip -p'\n> (same, without filenames).\n\nThat assumes that the in-repository data is zipped binary blob; the\nresult won't delta well, will it?\n\n\n"},{"id":"299514","messageId":"87pop8wh5w.fsf@thinkpad.rath.org","threadId":"43863","inReplyTo":"alpine.DEB.2.02.1608160926330.11774@nftneq.ynat.uz","subject":"Re: Working with zip files","fromName":"Nikolaus Rath","fromEmail":"nikolaus@rath.org","sentAt":"2016-08-16T21:14:03Z","receivedAt":"2016-08-16T21:16:08Z","isPatch":false,"sender":{"key":"nikolaus@rath.org","avatar":"https://gravatar.com/avatar/d3b6cdc023a0665882d5aac373c0115bcb36f730c2e4a4051061fefbaf7dd041?d=mp&s=160"},"body":"On Aug 16 2016, David Lang <david@lang.hm> wrote:\n> On Tue, 16 Aug 2016, Nikolaus Rath wrote:\n>\n>> I would like to store Simulink models in a Git\n>> repository. Unfortunately, the file format is binary. But luckily, the\n>> binary format happens to be a zipfile containing nicely formatted XML\n>> files.\n>>\n>> Is there a way to teach Git to take advantage of this when storing,\n>> diff-ing and merging these files?\n>\n> you should be able to use clean/smudge to have git store the files\n> uncompressed, which will help a lot.\n\nHaving looked at that, I'm not sure if this really helps:\n\nAs I understand, the smudge command is run on checkout to convert the\nblob in the repository to the format that is desired in the working\ntree. But this is the opposite of what I need: on checkout, I need to\nconvert the text data in the repository to a blob in the working tree.\n\nFurthermore, I need to convert multiple text files into one blob, will\nsmudge/clean seem to do just 1:1 conversions.\n\nAm I missing something? Are there any other options?\n\nBest,\nNikolaus\n-- \nGPG encrypted emails preferred. Key id: 0xD113FCAC3C4E599F\nFingerprint: ED31 791B 2C5C 1613 AF38 8B8A D113 FCAC 3C4E 599F\n\n             »Time flies like an arrow, fruit flies like a Banana.«\n"},{"id":"299525","messageId":"CA+P7+xp2q8JWqqtqj-ATd=Ox9snTpC0DE9ND76VhV-5hOvVQdw@mail.gmail.com","threadId":"43863","inReplyTo":"87pop8wh5w.fsf@thinkpad.rath.org","subject":"Re: Working with zip files","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2016-08-17T05:31:50Z","receivedAt":"2016-08-17T05:32:18Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Tue, Aug 16, 2016 at 2:14 PM, Nikolaus Rath <Nikolaus@rath.org> wrote:\n> On Aug 16 2016, David Lang <david@lang.hm> wrote:\n>> On Tue, 16 Aug 2016, Nikolaus Rath wrote:\n>>\n>>> I would like to store Simulink models in a Git\n>>> repository. Unfortunately, the file format is binary. But luckily, the\n>>> binary format happens to be a zipfile containing nicely formatted XML\n>>> files.\n>>>\n>>> Is there a way to teach Git to take advantage of this when storing,\n>>> diff-ing and merging these files?\n>>\n>> you should be able to use clean/smudge to have git store the files\n>> uncompressed, which will help a lot.\n>\n> Having looked at that, I'm not sure if this really helps:\n>\n> As I understand, the smudge command is run on checkout to convert the\n> blob in the repository to the format that is desired in the working\n> tree. But this is the opposite of what I need: on checkout, I need to\n> convert the text data in the repository to a blob in the working tree.\n>\n> Furthermore, I need to convert multiple text files into one blob, will\n> smudge/clean seem to do just 1:1 conversions.\n>\n> Am I missing something? Are there any other options?\n\nYou want to store the contents of the zip file as *one* blob that is\nthe uncompressed contents of the archive somehow concatenated\ntogether. That should still be a 1:1 relationship.\n\nYou won't store one blob per file in the zip.\n\nThanks,\nJake\n"},{"id":"299534","messageId":"alpine.DEB.2.02.1608170256450.11774@nftneq.ynat.uz","threadId":"43863","inReplyTo":"87pop8wh5w.fsf@thinkpad.rath.org","subject":"Re: Working with zip files","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2016-08-17T09:58:09Z","receivedAt":"2016-08-17T09:58:24Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Tue, 16 Aug 2016, Nikolaus Rath wrote:\n\n> On Aug 16 2016, David Lang <david@lang.hm> wrote:\n>> On Tue, 16 Aug 2016, Nikolaus Rath wrote:\n>>\n>>> I would like to store Simulink models in a Git\n>>> repository. Unfortunately, the file format is binary. But luckily, the\n>>> binary format happens to be a zipfile containing nicely formatted XML\n>>> files.\n>>>\n>>> Is there a way to teach Git to take advantage of this when storing,\n>>> diff-ing and merging these files?\n>>\n>> you should be able to use clean/smudge to have git store the files\n>> uncompressed, which will help a lot.\n>\n> Having looked at that, I'm not sure if this really helps:\n>\n> As I understand, the smudge command is run on checkout to convert the\n> blob in the repository to the format that is desired in the working\n> tree. But this is the opposite of what I need: on checkout, I need to\n> convert the text data in the repository to a blob in the working tree.\n>\n> Furthermore, I need to convert multiple text files into one blob, will\n> smudge/clean seem to do just 1:1 conversions.\n>\n> Am I missing something? Are there any other options?\n\nso the smudge command would zip the file and the clean command would unzip the \nfile (assuming it's a single file, if the zip is multiple files, you will have \nto add something to combine them)\n\nyou want the working tree to have a zip file and the repository to have text.\n\nDavid Lang\n"},{"id":"299599","messageId":"12866c04-f910-2a83-b445-6eada3d2efc9@gmail.com","threadId":"43863","inReplyTo":"xmqq8tvwcvrc.fsf@gitster.mtv.corp.google.com","subject":"Re: Working with zip files","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-08-18T12:16:27Z","receivedAt":"2016-08-18T12:18:11Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 16.08.2016 o 22:19, Junio C Hamano pisze:\n> Jakub Narębski <jnareb@gmail.com> writes:\n> \n>> There is also `textconv` filter that can be used instead; it might\n>> be 'unzip -c' (extract files to stdout, with filenames), or 'unzip -p'\n>> (same, without filenames).\n> \n> That assumes that the in-repository data is zipped binary blob; the\n> result won't delta well, will it?\n\nFull solution would involve `clean` filter to rezip with no compression\n(which should delta well) and optional `smudge` filter to recompress;\nif round-trip bit-for-bit equality is needed, the original zip parameters\nmust be saved somewhere, e.g. as ZIP archive comments.  This was mentioned\nin the earlier part of my email (which might have been not clear):\n\nJN>> You can find rezip clean/smudge filter (originally intended for\nJN>> OpenDocument Format (ODF), that is OpenOffice.org etc.) that stores\nJN>> zip or zip-archive (like ODT, jar, etc.) uncompressed.  I think\nJN>> you can find it on GitWiki, but I might be mistaken.\n \nUsing 'unzip -c' as separate / additional `textconv` filter for diff\ngeneration allows to separate the problem of deltifiable storage format\nfrom textual representation for diff-ing.\n\nThough best results could be had with `diff` and `merge` drivers...\n\n-- \nJakub Narębski\n"},{"id":"299646","messageId":"alpine.DEB.2.02.1608180954400.11774@nftneq.ynat.uz","threadId":"43863","inReplyTo":"12866c04-f910-2a83-b445-6eada3d2efc9@gmail.com","subject":"Re: Working with zip files","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2016-08-18T16:56:16Z","receivedAt":"2016-08-19T01:15:15Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Thu, 18 Aug 2016, Jakub Narębski wrote:\n\n> JN>> You can find rezip clean/smudge filter (originally intended for\n> JN>> OpenDocument Format (ODF), that is OpenOffice.org etc.) that stores\n> JN>> zip or zip-archive (like ODT, jar, etc.) uncompressed.  I think\n> JN>> you can find it on GitWiki, but I might be mistaken.\n>\n> Using 'unzip -c' as separate / additional `textconv` filter for diff\n> generation allows to separate the problem of deltifiable storage format\n> from textual representation for diff-ing.\n>\n> Though best results could be had with `diff` and `merge` drivers...\n\ncan you point at an example of how to do this? when I went looking about a year \nago to deal with single-line json data I wasn't able to find anything good. I \nended up using clean/smudge to pretty-print the json so it was easier to handle.\n\nDavid Lang"},{"id":"299678","messageId":"alpine.DEB.2.02.1608181954560.11774@nftneq.ynat.uz","threadId":"43863","inReplyTo":"CANQwDwebocSzwrLRNoZMCkXy0D2HrSrXSyLxV=jNJ4xzCWDhHw@mail.gmail.com","subject":"Re: Working with zip files","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2016-08-19T03:00:22Z","receivedAt":"2016-08-19T03:01:50Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Thu, 18 Aug 2016, Jakub Narębski wrote:\n\n> On 18 August 2016 at 18:56, David Lang <david@lang.hm> wrote:\n>> On Thu, 18 Aug 2016, Jakub Narębski wrote:\n>>\n>>> JN>> You can find rezip clean/smudge filter (originally intended for\n>>> JN>> OpenDocument Format (ODF), that is OpenOffice.org etc.) that stores\n>>> JN>> zip or zip-archive (like ODT, jar, etc.) uncompressed.  I think\n>>> JN>> you can find it on GitWiki, but I might be mistaken.\n>>>\n>>> Using 'unzip -c' as separate / additional `textconv` filter for diff\n>>> generation allows to separate the problem of deltifiable storage format\n>>> from textual representation for diff-ing.\n>>>\n>>> Though best results could be had with `diff` and `merge` drivers...\n>>\n>>\n>> can you point at an example of how to do this? when I went looking about a\n>> year ago to deal with single-line json data I wasn't able to find anything\n>> good. I ended up using clean/smudge to pretty-print the json so it was\n>> easier to handle.\n>\n> Pro Git has a chapter \"Customizing Git - Git Attributes\" about gitattributes\n> https://git-scm.com/book/en/v2/Customizing-Git-Git-Attributes\n>\n> The section \"Diffing Binary Files\" has two examples: docx2txt (with wrapper)\n> for DOCX (MS Word) files, and exiftool for images. For JSON you could use\n> some prettyprinter / formatter like pp-json.\n>\n> \"Performing text diffs of binary files\" section of gitattributes(1) manpage\n> covers 'textconv' vs 'diff', and uses 'exif' tool as textconv example.\n\nAs I read that section, it only applies to the human readable output of git \ndiff.\n\nAnd the merge section only talks about the default of using patch vs accepting a \nspecific version in a merge.\n\nIt seems to me that what I'm looking for would be something to tell git to use a \ndifferent command instead of diff/patch internally when creating and using the \nbundles.\n\nDavid Lang"},{"id":"299685","messageId":"CANQwDwebocSzwrLRNoZMCkXy0D2HrSrXSyLxV=jNJ4xzCWDhHw@mail.gmail.com","threadId":"43863","inReplyTo":"alpine.DEB.2.02.1608180954400.11774@nftneq.ynat.uz","subject":"Re: Working with zip files","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-08-18T17:45:48Z","receivedAt":"2016-08-19T07:21:45Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On 18 August 2016 at 18:56, David Lang <david@lang.hm> wrote:\n> On Thu, 18 Aug 2016, Jakub Narębski wrote:\n>\n>> JN>> You can find rezip clean/smudge filter (originally intended for\n>> JN>> OpenDocument Format (ODF), that is OpenOffice.org etc.) that stores\n>> JN>> zip or zip-archive (like ODT, jar, etc.) uncompressed.  I think\n>> JN>> you can find it on GitWiki, but I might be mistaken.\n>>\n>> Using 'unzip -c' as separate / additional `textconv` filter for diff\n>> generation allows to separate the problem of deltifiable storage format\n>> from textual representation for diff-ing.\n>>\n>> Though best results could be had with `diff` and `merge` drivers...\n>\n>\n> can you point at an example of how to do this? when I went looking about a\n> year ago to deal with single-line json data I wasn't able to find anything\n> good. I ended up using clean/smudge to pretty-print the json so it was\n> easier to handle.\n\nPro Git has a chapter \"Customizing Git - Git Attributes\" about gitattributes\nhttps://git-scm.com/book/en/v2/Customizing-Git-Git-Attributes\n\nThe section \"Diffing Binary Files\" has two examples: docx2txt (with wrapper)\nfor DOCX (MS Word) files, and exiftool for images. For JSON you could use\nsome prettyprinter / formatter like pp-json.\n\n\"Performing text diffs of binary files\" section of gitattributes(1) manpage\ncovers 'textconv' vs 'diff', and uses 'exif' tool as textconv example.\n\nHTH\n-- \nJakub Narębski\n\n\n\n\n-- \nJakub Narebski\n"}]}