threads / discuss / 15539

Management of opendocument (openoffice.org) files in git

Subject: Management of opendocument (openoffice.org) files in git

## tl;dr

10 messages between Sep 15, 2008 and Sep 23, 2008.

replies: 9people: 7as markdown or json

Sergio Callegari· Sep 15, 2008, 22:40 UTC · lore
Hi,

Management of opendocument files in git has been discussed a short time ago. Here is an helper script that may help achieving better density in git packs containg blobs from openoffice files.

To try it, save the following as "rezip" with execution permission:
-----8<----------------------- 

#! /bin/bash # # (c) 2008 Sergio Callegari # # Rewrites a zip archive, possibly changing the compression level

USAGE='Usage: rezip [options] [file]
with options:
  [-h | --help]            Gives help
  [-p ?]                   Lists known profiles
  [--unzip_opts options]   Pass options to unzip helper to read zip file
  [--zip_opts options]     Pass options to zip helper to write zip file
  [-p | --profile profile] Get options for helpers from profile

Rewrites a zip archive, possibily changing the compression level. If the archive name is unspecified, then the command operates like a filter, reading from standard input and writing to standard output. Options can be manually provided to the unzip process doing the read and to the zip process doing the write. Alternatively a profile can be used to set options automatically.'

PROFILES="ODF_UNCOMPRESS ODF_COMPRESS"

PROFILE_UNZIP_ODF_UNCOMPRESS='-b -qq -X' PROFILE_ZIP_ODF_UNCOMPRESS='-q -r -D -0' PROFILE_UNZIP_ODF_COMPRESS='-b -qq -X' PROFILE_ZIP_ODF_COMPRESS='-q -r -D -6'

die()
{
    echo "$1" >&$2
    exit $3
}

UNZIP_OPTS="" ZIP_OPTS=""

while true ; do
    case "$1" in
        -h | --help)
            die "$USAGE" 1 0 ;;
        -p | --profile)
            if [ "$2" = "?" ] ; then
                die "Avalilable profiles: ${PROFILES}" 1 0 ;
            else
                profile=$2
                shift
                profile_unzip=PROFILE_UNZIP_${profile}
                profile_zip=PROFILE_ZIP_${profile}
                UNZIP_OPTS=${!profile_unzip}
                ZIP_OPTS=${!profile_zip}
            fi ;;
        --unzip_opts)
            UNZIP_OPTS=${UNZIP_OPTS} $2
            shift ;;
        --zip_opts)
            ZIP_OPTS=${ZIP_OPTS} $2
            shift ;;
        -*)
            die "$USAGE" 2 1 ;;
        *)
            break ;;
    esac
    shift
done
if [ $# = 0 ] ; then
    tmpcopy=$(mktemp rezip.zip.XXXXXX)
    cat > $tmpcopy
    filename="$tmpcopy"
else
    tmpcopy=""
    filename="$1"
fi

workdir=$(mktemp -d -t rezip.workdir.XXXXXX) curdir=$(pwd)

cd $workdir
unzip $UNZIP_OPTS "$curdir/$filename"
zip $ZIP_OPTS "$curdir/$filename" .
cd $curdir
rm -fr $workdir
if [ ! -z "$tmpcopy" ] ; then
  cat $filename
  rm $tmpcopy
fi
--------8<------------------------
then put in your .git/config something like
[filter "opendocument"]
        clean = "rezip -p ODF_UNCOMPRESS"
        smudge = "rezip -p ODF_COMPRESS"
and finally set gitattributes as

*.odt filter=opendocument *.ods filter=opendocument *.odp filter=opendocument

Note:
   with this you might experience some delay on operations like
git status
git add
git commit -a
git checkout
depending on the size of the opendocument files being tracked.
Before using on anything sensitive, please test that it does what it should.
The script should probably be made more robust against unexpected situations.
Hope it can be useful to someone.
Sergio
Matthieu Moy· Sep 16, 2008, 06:45 UTC · re: Sergio Callegari · lore

Re: Management of opendocument (openoffice.org) files in git

Sergio Callegari <sergio.callegari@gmail.com> writes:
Show 5 quoted lines
> Hi,
>
> Management of opendocument files in git has been discussed a short time ago.
> Here is an helper script that may help achieving better density in git packs
> containg blobs from openoffice files.

If you don't get "oh, sh*t, I lost data with it"-kind of feedback, can you add it to the wiki:

http://git.or.cz/gitwiki/GitTips#head-1cdd4ab777e74f12d1ffa7f0a793e46dd06e5945
Thanks,
-- 
Matthieu
Sergio Callegari· Sep 16, 2008, 07:41 UTC · re: Matthieu Moy · lore

Re: Management of opendocument (openoffice.org) files in git

Matthieu Moy wrote:
Show 18 quoted lines
> Sergio Callegari <sergio.callegari@gmail.com> writes:
>
>   
>> Hi,
>>
>> Management of opendocument files in git has been discussed a short time ago.
>> Here is an helper script that may help achieving better density in git packs
>> containg blobs from openoffice files.
>>     
>
> If you don't get "oh, sh*t, I lost data with it"-kind of feedback, can
> you add it to the wiki:
>
> http://git.or.cz/gitwiki/GitTips#head-1cdd4ab777e74f12d1ffa7f0a793e46dd06e5945
>
> Thanks,
>
>   

Sure. I'll wait a few days for feedback (also from myself), then I'll add it there. I've already got a couple of corrections and suggestions from Paolo. Would it be useful also to add a note about how to filter-branches with a plain "--tree-filter true" to convert archives so that they take advantage of storing ODF stuff uncompressed? If proper, I can add that too.

Sergio
Johannes Sixt· Sep 16, 2008, 07:09 UTC · re: Sergio Callegari · lore

Re: Management of opendocument (openoffice.org) files in git

Sergio Callegari schrieb:
Show 21 quoted lines
> if [ $# = 0 ] ; then
>     tmpcopy=$(mktemp rezip.zip.XXXXXX)
>     cat > $tmpcopy
>     filename="$tmpcopy"
> else
>     tmpcopy=""
>     filename="$1"
> fi
> 
> workdir=$(mktemp -d -t rezip.workdir.XXXXXX)
> curdir=$(pwd)
> 
> cd $workdir
> unzip $UNZIP_OPTS "$curdir/$filename"
> zip $ZIP_OPTS "$curdir/$filename" .
> cd $curdir
> rm -fr $workdir
> if [ ! -z "$tmpcopy" ] ; then
>   cat $filename
>   rm $tmpcopy
> fi
You don't need a temporay zip filename in filter mode:
  unzip $UNZIP_OPTS /dev/stdin  # works for me, but not 100% portable
  zip $ZIP_OPTS - .             # writes to stdout
Show 5 quoted lines
> then put in your .git/config something like
> 
> [filter "opendocument"]
>         clean = "rezip -p ODF_UNCOMPRESS"
>         smudge = "rezip -p ODF_COMPRESS"

Is the smudge filter really necessary? Can't OOo work with files at compression level 0?

-- Hannes
Sergio Callegari· Sep 16, 2008, 07:41 UTC · re: Johannes Sixt · lore

Re: Management of opendocument (openoffice.org) files in git

Johannes Sixt wrote:
Show 7 quoted lines
>
> You don't need a temporay zip filename in filter mode:
>
>   unzip $UNZIP_OPTS /dev/stdin  # works for me, but not 100% portable
>   zip $ZIP_OPTS - .             # writes to stdout
>
>   

The unzip documentation says "Archives read from standard input are not yet supported", so I was a bit worried about using the /dev/stdin thing. Might it be that there are subtle cases where unzip needs to seek or rewind?

Show 10 quoted lines
>> then put in your .git/config something like
>>
>> [filter "opendocument"]
>>         clean = "rezip -p ODF_UNCOMPRESS"
>>         smudge = "rezip -p ODF_COMPRESS"
>>     
>
> Is the smudge filter really necessary? Can't OOo work with files at
> compression level 0?
>   

Yes, you can live perfectly without smudge. But at times it is not that nice. Just think of finding a directory with say 15 lectures as impress slides taking 10 times the space it needs, particularly if you need to pass those files to someone else. As a matter of fact, ODF xml is very verbose and compresses particularly well having long tags. But you might want to compress -1 rather than the default in smudge to speed it up a little. Can be done either adding a new profile to the script (say ODF_COMPRESS_FAST) or by adding --zip_opts -1 to the smudge command line. Also, we might want to add some -n suffixes to zip, to prevent it from trying to compress a few things like .png or .jpeg images and that have their own compression. That should gain us some speed in smudging.

In any case - but this is just my feeling - it is much more disturbing the delay that the clean filter introduces in operations like add or status or commit, than the one introduced by the (slower) smudge filter in checkout. There must be some psychological reason for that. Possibly we are "programmed" to accept waiting when we need to get something and conversely we are impatient when someone should accept something from us.

Sergio
Johannes Sixt· Sep 16, 2008, 07:52 UTC · re: Sergio Callegari · lore

Re: Management of opendocument (openoffice.org) files in git

Sergio Callegari schrieb:
Show 12 quoted lines
> Johannes Sixt wrote:
>>
>> You don't need a temporay zip filename in filter mode:
>>
>>   unzip $UNZIP_OPTS /dev/stdin  # works for me, but not 100% portable
>>   zip $ZIP_OPTS - .             # writes to stdout
>>
>>   
> The unzip documentation says "Archives read from standard input are not
> yet supported", so I was a bit worried about using the /dev/stdin
> thing.  Might it be that there are subtle cases where unzip needs to
> seek or rewind?

I didn't test thoroughly nor did I read the documentation. So if the documentation says stdin is a no-go, you better do what it says. ;)

> In any case - but this is just my feeling - it is much more disturbing
> the delay that the clean filter introduces in operations like add or
> status or commit, than the one introduced by the (slower) smudge filter
> in checkout.

My feeling is that the temporary tree that is written slows it down. If rezip were a true filter it could be faster.

-- Hannes
Avery Pennarun· Sep 16, 2008, 16:04 UTC · re: Sergio Callegari · lore

Re: Management of opendocument (openoffice.org) files in git

On Tue, Sep 16, 2008 at 3:41 AM, Sergio Callegari <sergio.callegari@gmail.com> wrote:

Show 12 quoted lines
> Johannes Sixt wrote:
>>
>> You don't need a temporay zip filename in filter mode:
>>
>>  unzip $UNZIP_OPTS /dev/stdin  # works for me, but not 100% portable
>>  zip $ZIP_OPTS - .             # writes to stdout
>>
>>
>
> The unzip documentation says "Archives read from standard input are not yet
> supported", so I was a bit worried about using the /dev/stdin thing.  Might
> it be that there are subtle cases where unzip needs to seek or rewind?

IIRC zip files keep their index at the end of the file, which means zipping in a pipeline is efficient (you can write all the blocks first, then drop the final index at the end) but unzipping that way is really hard.

unzipping from /dev/stdin seems to work if stdin is seekable, otherwise not.
       unzip /dev/stdin <filename.zip    # works
       cat filename.zip | unzip /dev/stdin    # doesn't work
Have fun,
Avery
Stephen R. van den Berg· Sep 16, 2008, 19:28 UTC · re: Avery Pennarun · lore

Re: Management of opendocument (openoffice.org) files in git

Avery Pennarun wrote:
Show 7 quoted lines
>On Tue, Sep 16, 2008 at 3:41 AM, Sergio Callegari
><sergio.callegari@gmail.com> wrote:
>> Johannes Sixt wrote:
>IIRC zip files keep their index at the end of the file, which means
>zipping in a pipeline is efficient (you can write all the blocks
>first, then drop the final index at the end) but unzipping that way is
>really hard.

Well, the index *is* at the end, yes, however, almost all (if not all) the information in the index is present directly in front of the files as well, so unzipping from stdin is possible without seeks (though the standard unzip doesn't support that (yet) because it tries to verify integrity and speed up lists using the index at the end).

-- 
Sincerely,
           Stephen R. van den Berg.

Human beings were created by water to transport it uphill.
Robin Rosenberg· Sep 16, 2008, 21:13 UTC · re: Avery Pennarun · lore

Re: Management of opendocument (openoffice.org) files in git

tisdagen den 16 september 2008 18.04.44 skrev Avery Pennarun:
> unzipping from /dev/stdin seems to work if stdin is seekable, otherwise not.
> 
>        unzip /dev/stdin <filename.zip    # works
>        cat filename.zip | unzip /dev/stdin    # doesn't work
Try a cousin of zip for extraction:
	cat filename.zip | jar x # works
> Have fun,
Always.
-- robin
Peter Krefting· Sep 23, 2008, 11:08 UTC · re: Sergio Callegari · lore

Re: Management of opendocument (openoffice.org) files in git

Sergio Callegari:
> To try it, save the following as "rezip" with execution permission:

I had some problems when I tried to implement this a Windows machine, it did not handle paths with spaces in them properly, and "Documents and Settings" does contain spaces.

The following patch fixes that for me:
---
 rezip |   12 ++++++------
 1 files changed, 6 insertions(+), 6 deletions(-)
diff --git a/rezip b/rezip
index 15f83a4..845e875 100755
--- a/rezip
+++ b/rezip
@@ -66,7 +66,7 @@ done
 
 if [ $# = 0 ] ; then
     tmpcopy=$(mktemp rezip.zip.XXXXXX)
-    cat > $tmpcopy
+    cat > "$tmpcopy"
     filename="$tmpcopy"
 else
     tmpcopy=""
@@ -76,12 +76,12 @@ fi
 workdir=$(mktemp -d -t rezip.workdir.XXXXXX)
 curdir=$(pwd)
 
-cd $workdir
+cd "$workdir"
 unzip $UNZIP_OPTS "$curdir/$filename"
 zip $ZIP_OPTS "$curdir/$filename" .
-cd $curdir
-rm -fr $workdir
+cd "$curdir"
+rm -fr "$workdir"
 if [ ! -z "$tmpcopy" ] ; then
-  cat $filename
-  rm $tmpcopy
+  cat "$filename"
+  rm "$tmpcopy"
 fi
-- 
\\// Peter - http://www.softwolves.pp.se/

← back to recent threads