{"thread":{"id":"15539","subject":"Management of opendocument (openoffice.org) files in git","startedAt":"2008-09-15T22:40:01Z","lastAt":"2008-09-23T11:08:37Z","messageCount":10,"participants":["Sergio Callegari","Matthieu Moy","Johannes Sixt","Avery Pennarun","Stephen R. van den Berg","Robin Rosenberg","Peter Krefting"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"90782","messageId":"loom.20080915T222909-709@post.gmane.org","threadId":"15539","inReplyTo":null,"subject":"Management of opendocument (openoffice.org) files in git","fromName":"Sergio Callegari","fromEmail":"sergio.callegari@gmail.com","sentAt":"2008-09-15T22:40:01Z","receivedAt":"2008-09-15T22:40:01Z","isPatch":false,"sender":{"key":"sergio.callegari@gmail.com","avatar":"https://gravatar.com/avatar/c98f41317e0422c1e630385de0e3970227b8e5ad15f35ba8586066467cc833bc?d=mp&s=160"},"body":"Hi,\n\nManagement of opendocument files in git has been discussed a short time ago.\nHere is an helper script that may help achieving better density in git packs\ncontaing blobs from openoffice files.\n\nTo try it, save the following as \"rezip\" with execution permission:\n\n-----8<----------------------- \n\n#! /bin/bash\n#\n# (c) 2008 Sergio Callegari\n#\n# Rewrites a zip archive, possibly changing the compression level\n\nUSAGE='Usage: rezip [options] [file]\nwith options:\n  [-h | --help]            Gives help\n  [-p ?]                   Lists known profiles\n  [--unzip_opts options]   Pass options to unzip helper to read zip file\n  [--zip_opts options]     Pass options to zip helper to write zip file\n  [-p | --profile profile] Get options for helpers from profile\n\nRewrites a zip archive, possibily changing the compression level.\nIf the archive name is unspecified, then the command operates like a filter,\nreading from standard input and writing to standard output.\nOptions can be manually provided to the unzip process doing the read and to\nthe zip process doing the write. Alternatively a profile can be used to set\noptions automatically.'\n\nPROFILES=\"ODF_UNCOMPRESS ODF_COMPRESS\"\n\nPROFILE_UNZIP_ODF_UNCOMPRESS='-b -qq -X'\nPROFILE_ZIP_ODF_UNCOMPRESS='-q -r -D -0'\nPROFILE_UNZIP_ODF_COMPRESS='-b -qq -X'\nPROFILE_ZIP_ODF_COMPRESS='-q -r -D -6'\n\ndie()\n{\n    echo \"$1\" >&$2\n    exit $3\n}\n\nUNZIP_OPTS=\"\"\nZIP_OPTS=\"\"\n\nwhile true ; do\n    case \"$1\" in\n        -h | --help)\n            die \"$USAGE\" 1 0 ;;\n        -p | --profile)\n            if [ \"$2\" = \"?\" ] ; then\n                die \"Avalilable profiles: ${PROFILES}\" 1 0 ;\n            else\n                profile=$2\n                shift\n                profile_unzip=PROFILE_UNZIP_${profile}\n                profile_zip=PROFILE_ZIP_${profile}\n                UNZIP_OPTS=${!profile_unzip}\n                ZIP_OPTS=${!profile_zip}\n            fi ;;\n        --unzip_opts)\n            UNZIP_OPTS=${UNZIP_OPTS} $2\n            shift ;;\n        --zip_opts)\n            ZIP_OPTS=${ZIP_OPTS} $2\n            shift ;;\n        -*)\n            die \"$USAGE\" 2 1 ;;\n        *)\n            break ;;\n    esac\n    shift\ndone\n\nif [ $# = 0 ] ; then\n    tmpcopy=$(mktemp rezip.zip.XXXXXX)\n    cat > $tmpcopy\n    filename=\"$tmpcopy\"\nelse\n    tmpcopy=\"\"\n    filename=\"$1\"\nfi\n\nworkdir=$(mktemp -d -t rezip.workdir.XXXXXX)\ncurdir=$(pwd)\n\ncd $workdir\nunzip $UNZIP_OPTS \"$curdir/$filename\"\nzip $ZIP_OPTS \"$curdir/$filename\" .\ncd $curdir\nrm -fr $workdir\nif [ ! -z \"$tmpcopy\" ] ; then\n  cat $filename\n  rm $tmpcopy\nfi\n\n--------8<------------------------\n\nthen put in your .git/config something like\n\n[filter \"opendocument\"]\n        clean = \"rezip -p ODF_UNCOMPRESS\"\n        smudge = \"rezip -p ODF_COMPRESS\"\n\nand finally set gitattributes as\n\n*.odt filter=opendocument\n*.ods filter=opendocument\n*.odp filter=opendocument\n\nNote:\n   with this you might experience some delay on operations like\ngit status\ngit add\ngit commit -a\ngit checkout\n\ndepending on the size of the opendocument files being tracked.\n\nBefore using on anything sensitive, please test that it does what it should.\n\nThe script should probably be made more robust against unexpected situations.\n\nHope it can be useful to someone.\n\nSergio\n"},{"id":"90810","messageId":"vpqwshctwr7.fsf@bauges.imag.fr","threadId":"15539","inReplyTo":"loom.20080915T222909-709@post.gmane.org","subject":"Re: Management of opendocument (openoffice.org) files in git","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2008-09-16T06:45:00Z","receivedAt":"2008-09-16T06:45:00Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Sergio Callegari <sergio.callegari@gmail.com> writes:\n\n> Hi,\n>\n> Management of opendocument files in git has been discussed a short time ago.\n> Here is an helper script that may help achieving better density in git packs\n> containg blobs from openoffice files.\n\nIf you don't get \"oh, sh*t, I lost data with it\"-kind of feedback, can\nyou add it to the wiki:\n\nhttp://git.or.cz/gitwiki/GitTips#head-1cdd4ab777e74f12d1ffa7f0a793e46dd06e5945\n\nThanks,\n\n-- \nMatthieu\n"},{"id":"90812","messageId":"48CF5B90.5050800@viscovery.net","threadId":"15539","inReplyTo":"loom.20080915T222909-709@post.gmane.org","subject":"Re: Management of opendocument (openoffice.org) files in git","fromName":"Johannes Sixt","fromEmail":"j.sixt@viscovery.net","sentAt":"2008-09-16T07:09:04Z","receivedAt":"2008-09-16T07:09:04Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Sergio Callegari schrieb:\n> if [ $# = 0 ] ; then\n>     tmpcopy=$(mktemp rezip.zip.XXXXXX)\n>     cat > $tmpcopy\n>     filename=\"$tmpcopy\"\n> else\n>     tmpcopy=\"\"\n>     filename=\"$1\"\n> fi\n> \n> workdir=$(mktemp -d -t rezip.workdir.XXXXXX)\n> curdir=$(pwd)\n> \n> cd $workdir\n> unzip $UNZIP_OPTS \"$curdir/$filename\"\n> zip $ZIP_OPTS \"$curdir/$filename\" .\n> cd $curdir\n> rm -fr $workdir\n> if [ ! -z \"$tmpcopy\" ] ; then\n>   cat $filename\n>   rm $tmpcopy\n> fi\n\nYou don't need a temporay zip filename in filter mode:\n\n  unzip $UNZIP_OPTS /dev/stdin  # works for me, but not 100% portable\n  zip $ZIP_OPTS - .             # writes to stdout\n\n> then put in your .git/config something like\n> \n> [filter \"opendocument\"]\n>         clean = \"rezip -p ODF_UNCOMPRESS\"\n>         smudge = \"rezip -p ODF_COMPRESS\"\n\nIs the smudge filter really necessary? Can't OOo work with files at\ncompression level 0?\n\n-- Hannes\n"},{"id":"90815","messageId":"48CF630F.4090808@gmail.com","threadId":"15539","inReplyTo":"48CF5B90.5050800@viscovery.net","subject":"Re: Management of opendocument (openoffice.org) files in git","fromName":"Sergio Callegari","fromEmail":"sergio.callegari@gmail.com","sentAt":"2008-09-16T07:41:03Z","receivedAt":"2008-09-16T07:41:03Z","isPatch":false,"sender":{"key":"sergio.callegari@gmail.com","avatar":"https://gravatar.com/avatar/c98f41317e0422c1e630385de0e3970227b8e5ad15f35ba8586066467cc833bc?d=mp&s=160"},"body":"Johannes Sixt wrote:\n>\n> You don't need a temporay zip filename in filter mode:\n>\n>   unzip $UNZIP_OPTS /dev/stdin  # works for me, but not 100% portable\n>   zip $ZIP_OPTS - .             # writes to stdout\n>\n>   \nThe unzip documentation says \"Archives read from standard input are not \nyet supported\", so I was a bit worried about using the /dev/stdin \nthing.  Might it be that there are subtle cases where unzip needs to \nseek or rewind?\n>> then put in your .git/config something like\n>>\n>> [filter \"opendocument\"]\n>>         clean = \"rezip -p ODF_UNCOMPRESS\"\n>>         smudge = \"rezip -p ODF_COMPRESS\"\n>>     \n>\n> Is the smudge filter really necessary? Can't OOo work with files at\n> compression level 0?\n>   \nYes, you can live perfectly without smudge.  But at times it is not that \nnice. Just think of finding a directory with say 15 lectures as impress \nslides taking 10 times the space it needs, particularly if you need to \npass those files to someone else. As a matter of fact, ODF xml is very \nverbose and compresses particularly well having long tags.\nBut you might want to compress -1 rather than the default in smudge to \nspeed it up a little.  Can be done either adding a new profile to the \nscript (say ODF_COMPRESS_FAST) or by adding --zip_opts -1 to the smudge \ncommand line.\nAlso, we might want to add some -n suffixes to zip, to prevent it from \ntrying to compress a few things like .png or .jpeg images and that have \ntheir own compression.  That should gain us some speed in smudging.\n\nIn any case - but this is just my feeling - it is much more disturbing \nthe delay that the clean filter introduces in operations like add or \nstatus or commit, than the one introduced by the (slower) smudge filter \nin checkout.  There must be some psychological reason for that.  \nPossibly we are \"programmed\" to accept waiting when we need to get \nsomething and conversely we are impatient when someone should accept \nsomething from us.\n\nSergio\n"},{"id":"90816","messageId":"48CF6332.3020803@gmail.com","threadId":"15539","inReplyTo":"vpqwshctwr7.fsf@bauges.imag.fr","subject":"Re: Management of opendocument (openoffice.org) files in git","fromName":"Sergio Callegari","fromEmail":"sergio.callegari@gmail.com","sentAt":"2008-09-16T07:41:38Z","receivedAt":"2008-09-16T07:41:38Z","isPatch":false,"sender":{"key":"sergio.callegari@gmail.com","avatar":"https://gravatar.com/avatar/c98f41317e0422c1e630385de0e3970227b8e5ad15f35ba8586066467cc833bc?d=mp&s=160"},"body":"Matthieu Moy wrote:\n> Sergio Callegari <sergio.callegari@gmail.com> writes:\n>\n>   \n>> Hi,\n>>\n>> Management of opendocument files in git has been discussed a short time ago.\n>> Here is an helper script that may help achieving better density in git packs\n>> containg blobs from openoffice files.\n>>     \n>\n> If you don't get \"oh, sh*t, I lost data with it\"-kind of feedback, can\n> you add it to the wiki:\n>\n> http://git.or.cz/gitwiki/GitTips#head-1cdd4ab777e74f12d1ffa7f0a793e46dd06e5945\n>\n> Thanks,\n>\n>   \nSure.  I'll wait a few days for feedback (also from myself), then I'll\nadd it there.\nI've already got a couple of corrections and suggestions from Paolo.\nWould it be useful also to add a note about how to filter-branches with\na plain \"--tree-filter true\" to convert archives so that they take\nadvantage of storing ODF stuff uncompressed?\nIf proper, I can add that too.\n\n\nSergio\n"},{"id":"90817","messageId":"48CF65CC.6080509@viscovery.net","threadId":"15539","inReplyTo":"48CF630F.4090808@gmail.com","subject":"Re: Management of opendocument (openoffice.org) files in git","fromName":"Johannes Sixt","fromEmail":"j.sixt@viscovery.net","sentAt":"2008-09-16T07:52:44Z","receivedAt":"2008-09-16T07:52:44Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Sergio Callegari schrieb:\n> Johannes Sixt wrote:\n>>\n>> You don't need a temporay zip filename in filter mode:\n>>\n>>   unzip $UNZIP_OPTS /dev/stdin  # works for me, but not 100% portable\n>>   zip $ZIP_OPTS - .             # writes to stdout\n>>\n>>   \n> The unzip documentation says \"Archives read from standard input are not\n> yet supported\", so I was a bit worried about using the /dev/stdin\n> thing.  Might it be that there are subtle cases where unzip needs to\n> seek or rewind?\n\nI didn't test thoroughly nor did I read the documentation. So if the\ndocumentation says stdin is a no-go, you better do what it says. ;)\n\n> In any case - but this is just my feeling - it is much more disturbing\n> the delay that the clean filter introduces in operations like add or\n> status or commit, than the one introduced by the (slower) smudge filter\n> in checkout.\n\nMy feeling is that the temporary tree that is written slows it down. If\nrezip were a true filter it could be faster.\n\n-- Hannes\n"},{"id":"90843","messageId":"32541b130809160904v7acc73cfm4856c33d40555e94@mail.gmail.com","threadId":"15539","inReplyTo":"48CF630F.4090808@gmail.com","subject":"Re: Management of opendocument (openoffice.org) files in git","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-09-16T16:04:44Z","receivedAt":"2008-09-16T16:04:44Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Tue, Sep 16, 2008 at 3:41 AM, Sergio Callegari\n<sergio.callegari@gmail.com> wrote:\n> Johannes Sixt wrote:\n>>\n>> You don't need a temporay zip filename in filter mode:\n>>\n>>  unzip $UNZIP_OPTS /dev/stdin  # works for me, but not 100% portable\n>>  zip $ZIP_OPTS - .             # writes to stdout\n>>\n>>\n>\n> The unzip documentation says \"Archives read from standard input are not yet\n> supported\", so I was a bit worried about using the /dev/stdin thing.  Might\n> it be that there are subtle cases where unzip needs to seek or rewind?\n\nIIRC zip files keep their index at the end of the file, which means\nzipping in a pipeline is efficient (you can write all the blocks\nfirst, then drop the final index at the end) but unzipping that way is\nreally hard.\n\nunzipping from /dev/stdin seems to work if stdin is seekable, otherwise not.\n\n       unzip /dev/stdin <filename.zip    # works\n       cat filename.zip | unzip /dev/stdin    # doesn't work\n\nHave fun,\n\nAvery\n"},{"id":"90860","messageId":"20080916192830.GA16455@cuci.nl","threadId":"15539","inReplyTo":"32541b130809160904v7acc73cfm4856c33d40555e94@mail.gmail.com","subject":"Re: Management of opendocument (openoffice.org) files in git","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-16T19:28:30Z","receivedAt":"2008-09-16T19:28:30Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Avery Pennarun wrote:\n>On Tue, Sep 16, 2008 at 3:41 AM, Sergio Callegari\n><sergio.callegari@gmail.com> wrote:\n>> Johannes Sixt wrote:\n>IIRC zip files keep their index at the end of the file, which means\n>zipping in a pipeline is efficient (you can write all the blocks\n>first, then drop the final index at the end) but unzipping that way is\n>really hard.\n\nWell, the index *is* at the end, yes, however, almost all (if not all) the\ninformation in the index is present directly in front of the files as\nwell, so unzipping from stdin is possible without seeks (though the\nstandard unzip doesn't support that (yet) because it tries to verify\nintegrity and speed up lists using the index at the end).\n-- \nSincerely,\n           Stephen R. van den Berg.\n\nHuman beings were created by water to transport it uphill.\n"},{"id":"90865","messageId":"200809162313.56444.robin.rosenberg.lists@dewire.com","threadId":"15539","inReplyTo":"32541b130809160904v7acc73cfm4856c33d40555e94@mail.gmail.com","subject":"Re: Management of opendocument (openoffice.org) files in git","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2008-09-16T21:13:56Z","receivedAt":"2008-09-16T21:13:56Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"tisdagen den 16 september 2008 18.04.44 skrev Avery Pennarun:\n> unzipping from /dev/stdin seems to work if stdin is seekable, otherwise not.\n> \n>        unzip /dev/stdin <filename.zip    # works\n>        cat filename.zip | unzip /dev/stdin    # doesn't work\n\nTry a cousin of zip for extraction:\n\n\tcat filename.zip | jar x # works\n\n> Have fun,\nAlways.\n\n-- robin\n"},{"id":"91376","messageId":"Pine.LNX.4.64.0809231202560.28506@ds9.cixit.se","threadId":"15539","inReplyTo":"loom.20080915T222909-709@post.gmane.org","subject":"Re: Management of opendocument (openoffice.org) files in git","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2008-09-23T11:08:37Z","receivedAt":"2008-09-23T11:08:37Z","isPatch":false,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Sergio Callegari:\n\n> To try it, save the following as \"rezip\" with execution permission:\n\nI had some problems when I tried to implement this a Windows machine,\nit did not handle paths with spaces in them properly, and \"Documents\nand Settings\" does contain spaces.\n\nThe following patch fixes that for me:\n---\n rezip |   12 ++++++------\n 1 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/rezip b/rezip\nindex 15f83a4..845e875 100755\n--- a/rezip\n+++ b/rezip\n@@ -66,7 +66,7 @@ done\n \n if [ $# = 0 ] ; then\n     tmpcopy=$(mktemp rezip.zip.XXXXXX)\n-    cat > $tmpcopy\n+    cat > \"$tmpcopy\"\n     filename=\"$tmpcopy\"\n else\n     tmpcopy=\"\"\n@@ -76,12 +76,12 @@ fi\n workdir=$(mktemp -d -t rezip.workdir.XXXXXX)\n curdir=$(pwd)\n \n-cd $workdir\n+cd \"$workdir\"\n unzip $UNZIP_OPTS \"$curdir/$filename\"\n zip $ZIP_OPTS \"$curdir/$filename\" .\n-cd $curdir\n-rm -fr $workdir\n+cd \"$curdir\"\n+rm -fr \"$workdir\"\n if [ ! -z \"$tmpcopy\" ] ; then\n-  cat $filename\n-  rm $tmpcopy\n+  cat \"$filename\"\n+  rm \"$tmpcopy\"\n fi\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"}]}