{"thread":{"id":"645","subject":"[PATCH] packed delta git","startedAt":"2005-05-17T22:57:45Z","lastAt":"2005-05-19T19:30:06Z","messageCount":5,"participants":["Chris Mason","Thomas Glanzmann","Nicolas Pitre"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"3484","messageId":"200505171857.46370.mason@suse.com","threadId":"645","inReplyTo":null,"subject":"[PATCH] packed delta git","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-05-17T22:57:45Z","receivedAt":"2005-05-17T22:57:45Z","isPatch":true,"sender":{"key":"mason@suse.com","avatar":null},"body":"Hello everyone,\n\nHere's a new version of my packed git patch, diffed on top of Nicolas' \ndelta code (link below to that).  It doesn't change the core git commands\nto create packed/delta files, that is done via a new git-pack command.  The\ngit-pack usage is very simple:\n\ngit-pack [<reference_sha1>:]<target_sha1> [ <next_sha1> ... ]\n\nIf you use the ref:target notation, it will create a delta between those \ntwo objects and relink the target into the resulting packed file.  If you\ndon't provide a reference sha1, the sha1 file is relinked into the packed file.\n\nA script is provided (git-pack-changes-script) that walks the rev-list output \nand puts your whole repository into packed/delta form.  This is just\na starting point, it needs more knobs for forward/reverse deltas,\nmax depth etc.  The script does delta tree files and packs both trees and \ncommits in with the blobs.\n\ngit-diff-tree -t was added so that it will show the sha1 of subtrees that\ndiffer while it processes things.\n\nThere is no way to unpack a file yet, so please use these with caution.\n\nI did tests against the current 2.6 git tree (ext3)\n\n                                      vanilla              delta+pack\n.git size                          191M                62M\ncheckout-cache (cold)     2m13s              38.9s\ncheckout-cache (hot)      8s  (2s user)      14.9s (11.46s user)\n\n2.6.11 without any changesets:\n                                      unpacked          packed\n.git size                           91M                  55M\n\nBecause the 2.6 kernel repo has 2.6.11 floating around in there as a tree\nwith no commit, you have to do a few steps in order to pack things in. \n\n# step one, pack all the files in 2.6.11 together\ngit-ls-tree -r v2.6.11 | awk '{print $3}' | xargs git-pack\n\n# step two, make packed deltas between trees for 2.6.11 and 2.6.12-rc2\n# (git-pack-changes-script -t only works with trees, not commits/tags)\ngit-pack-changes-script -t c39ae07f393806ccf406ef966e9a15afc43cc36a 1da177e4c3f41524e886b7f1b8a0c1fc7321cac2 | xargs git-pack\n\n# step three, use git-pack-changes to walk whole rev-list and pack/delta\n# everything else\ngit-pack-changes-script | xargs git-pack\n\nAnd finally, here's the patch.  It is on top of Nicolas' code, which you\ncan grab here:\n\nhttp://marc.theaimsgroup.com/?l=git&m=111587004902021&w=2\n\nSigned-off-by: Chris Mason <mason@suse.com>\n--\n\ndiff -urN linus.delta/cache.h linus/cache.h\n--- linus.delta/cache.h\t2005-05-17 16:51:51.410686192 -0400\n+++ linus/cache.h\t2005-05-17 15:17:16.015477408 -0400\n@@ -77,6 +77,16 @@\n \tchar name[0];\n };\n \n+struct packed_item {\n+\t/* length of compressed data */\n+\tunsigned long len;\n+\tstruct packed_item *next;\n+\t/* sha1 of uncompressed data */\n+\tchar sha1[20];\n+\t/* compressed data */\n+\tchar *data;\n+};\n+\n #define CE_NAMEMASK  (0x0fff)\n #define CE_STAGEMASK (0x3000)\n #define CE_STAGESHIFT 12\n@@ -135,8 +145,10 @@\n \n /* Read and unpack a sha1 file into memory, write memory to a sha1 file */\n extern void * map_sha1_file(const unsigned char *sha1, unsigned long *size);\n-extern void * unpack_sha1_file(void *map, unsigned long mapsize, char *type, unsigned long *size);\n+extern void * unpack_sha1_file(const unsigned char *sha1, void *map, unsigned long mapsize, char *type, unsigned long *size, const unsigned char *recur_sha1, int *chain);\n+extern void * raw_unpack_sha1_file(void *map, unsigned long mapsize, char *type, unsigned long *size);\n extern void * read_sha1_file(const unsigned char *sha1, char *type, unsigned long *size);\n+extern void * read_sha1_delta_ref(const unsigned char *sha1, char *type, unsigned long *size, const unsigned char *, int *chain);\n extern int write_sha1_file(char *buf, unsigned long len, const char *type, unsigned char *return_sha1);\n \n extern int check_sha1_signature(unsigned char *sha1, void *buf, unsigned long size, const char *type);\n@@ -152,6 +164,10 @@\n extern int get_sha1(const char *str, unsigned char *sha1);\n extern int get_sha1_hex(const char *hex, unsigned char *sha1);\n extern char *sha1_to_hex(const unsigned char *sha1);\t/* static buffer result! */\n+extern int pack_sha1_buffer(void *buf, unsigned long buf_len, char *type,\n+                            unsigned char *returnsha1, unsigned char *refsha1, \n+\t\t\t    struct packed_item **, int max_depth);\n+int write_packed_list(struct packed_item *head);\n \n /* General helper functions */\n extern void usage(const char *err);\nFiles linus.delta/.cache.h.swp and linus/.cache.h.swp differ\ndiff -urN linus.delta/diff-tree.c linus/diff-tree.c\n--- linus.delta/diff-tree.c\t2005-05-17 16:51:51.413685736 -0400\n+++ linus/diff-tree.c\t2005-05-16 20:08:41.000000000 -0400\n@@ -5,6 +5,7 @@\n static int silent = 0;\n static int verbose_header = 0;\n static int ignore_merges = 1;\n+static int show_tree_diffs = 0;\n static int recursive = 0;\n static int read_stdin = 0;\n static int line_termination = '\\n';\n@@ -114,6 +115,7 @@\n \tconst unsigned char *sha1, *sha2;\n \tint cmp, pathlen1, pathlen2;\n \tchar old_sha1_hex[50];\n+\tint retval = 0;\n \n \tsha1 = extract(tree1, size1, &path1, &mode1);\n \tsha2 = extract(tree2, size2, &path2, &mode2);\n@@ -143,11 +145,11 @@\n \t}\n \n \tif (recursive && S_ISDIR(mode1)) {\n-\t\tint retval;\n \t\tchar *newbase = malloc_base(base, path1, pathlen1);\n \t\tretval = diff_tree_sha1(sha1, sha2, newbase);\n \t\tfree(newbase);\n-\t\treturn retval;\n+\t\tif (!show_tree_diffs)\n+\t\t\treturn retval;\n \t}\n \n \tif (header) {\n@@ -155,7 +157,7 @@\n \t\theader = NULL;\n \t}\n \tif (silent)\n-\t\treturn 0;\n+\t\treturn retval;\n \n \tif (generate_patch) {\n \t\tif (!S_ISDIR(mode1))\n@@ -168,7 +170,7 @@\n \t\t       old_sha1_hex, sha1_to_hex(sha2), base, path1,\n \t\t       line_termination);\n \t}\n-\treturn 0;\n+\treturn retval;\n }\n \n static int interesting(void *tree, unsigned long size, const char *base)\n@@ -387,7 +389,7 @@\n }\n \n static char *diff_tree_usage =\n-\"diff-tree [-p] [-r] [-z] [--stdin] [-m] [-s] [-v] <tree sha1> <tree sha1>\";\n+\"diff-tree [-p] [-r] [-z] [--stdin] [-m] [-s] [-v] [-t] <tree sha1> <tree sha1>\";\n \n int main(int argc, char **argv)\n {\n@@ -428,6 +430,10 @@\n \t\t\tsilent = 1;\n \t\t\tcontinue;\n \t\t}\n+\t\tif (!strcmp(arg, \"-t\")) {\n+\t\t\tshow_tree_diffs = 1;\n+\t\t\tcontinue;\n+\t\t}\n \t\tif (!strcmp(arg, \"-v\")) {\n \t\t\tverbose_header = 1;\n \t\t\theader_prefix = \"diff-tree \";\ndiff -urN linus.delta/git-pack-changes-script linus/git-pack-changes-script\n--- linus.delta/git-pack-changes-script\t1969-12-31 19:00:00.000000000 -0500\n+++ linus/git-pack-changes-script\t2005-05-17 16:57:44.768967568 -0400\n@@ -0,0 +1,161 @@\n+#!/usr/bin/perl\n+#\n+# script to search through the rev-list output and generate delta history\n+# you can specify either a start and stop commit or two trees to search.\n+# with no command line args it searches the entire revision history.\n+# output is suitable for piping to xargs git-pack\n+\n+use strict;\n+\n+my $ret;\n+my $i;\n+my @wanted = ();\n+my $argc = scalar(@ARGV);\n+my $commit;\n+my $stop;\n+my %delta = ();\n+\n+sub add_delta($$) {\n+    my ($ref, $target) = @_;\n+    if (defined($delta{$target})) {\n+        return;\n+    }\n+    if ($target eq $delta{$ref}) {\n+\tprint $ref;\n+\treturn 1;\n+    }\n+    $delta{$target} = $ref;\n+    print \"$ref:$target\\n\";\n+\n+}\n+sub print_usage() {\n+    print STDERR \"usage: pack-changes [-c commit] [-s stop commit] [-t tree1 tree2]\\n\";\n+    exit(1);\n+}\n+\n+sub find_tree($) {\n+    my ($commit) = @_;\n+    open(CM, \"git-cat-file commit $commit|\") || die \"git-cat-file failed\";\n+    while(<CM>) {\n+        chomp;\n+\tmy @words = split;\n+\tif ($words[0] eq \"tree\") {\n+\t    return $words[1];\n+\t} elsif ($words[0] ne \"parent\") {\n+\t    last;\n+\t}\n+    }\n+    close(CM);\n+    if ($? && ($ret = $? >> 8)) {\n+        die \"cat-file $commit failed with $ret\";\n+    }\n+    return undef;\n+}\n+\n+sub test_diff($$) {\n+    my ($a, $b) = @_;\n+    open(DT, \"git-diff-tree -r -t $a $b|\") || die \"diff-tree failed\";\n+    while(<DT>) {\n+        chomp;\n+\tmy @words = split;\n+\tmy $sha1 = $words[2];\n+\tmy $change = $words[0];\n+\tif ($change =~ m/^\\*/) {\n+\t    @words = split(\"->\", $sha1);\n+\t    add_delta($words[0], $words[1]);\n+\t} elsif ($change =~ m/^\\-/) {\n+\t    next;\n+\t} else {\n+\t    print \"$sha1\\n\";\n+\t}\n+    }\n+    close(DT);\n+    if ($? && ($ret = $? >> 8)) {\n+\tdie \"git-diff-tree failed with $ret\";\n+    }\n+    return 0;\n+}\n+\n+for ($i = 0 ; $i < $argc ; $i++)  {\n+    if ($ARGV[$i] eq \"-c\") {\n+    \tif ($i == $argc - 1) {\n+\t    print_usage();\n+\t}\n+\t$commit = $ARGV[++$i];\n+    } elsif ($ARGV[$i] eq \"-s\") {\n+    \tif ($i == $argc - 1) {\n+\t    print_usage();\n+\t}\n+\t$stop = $ARGV[++$i];\n+    } elsif ($ARGV[$i] eq \"-t\") {\n+        if ($argc != 3 || $i != 0) {\n+\t    print_usage();\n+\t}\n+\tif (test_diff($ARGV[1], $ARGV[2])) {\n+\t    die \"test_diff failed\\n\";\n+\t}\n+\tadd_delta($ARGV[1], $ARGV[2]);\n+\texit(0);\n+    }\n+}\n+\n+if (!defined($commit)) {\n+    $commit = `commit-id`;\n+    if ($?) {\n+    \tprint STDERR \"commit-id failed, try using -c to specify a commit\\n\";\n+\texit(1);\n+    }\n+    chomp $commit;\n+}\n+\n+open(RL, \"git-rev-list $commit|\") || die \"rev-list failed\";\n+while(<RL>) {\n+    chomp;\n+    my $cur = $_;\n+    my $cur_tree;\n+    my $parent_tree;\n+    my $parent_commit = undef;\n+    open(PARENT, \"git-cat-file commit $cur|\") || die \"cat-file failed\";\n+    while(<PARENT>) {\n+        chomp;\n+\tmy @words = split;\n+\tif ($words[0] eq \"tree\") {\n+\t    $cur_tree = $words[1];\n+\t    next;\n+\t} elsif ($words[0] ne \"parent\") {\n+\t    last;\n+\t}\n+\t$parent_commit = $words[1];\n+\tmy $next = <PARENT>;\n+\t# ignore merge sets for now\n+\tif ($next =~ m/^parent/) {\n+\t    last;\n+\t}\n+\tif (test_diff($words[1], $cur)) {\n+\t    die \"test_diff failed\\n\";\n+\t}\n+\t$parent_tree = find_tree($words[1]);\n+\tif (!defined($parent_tree)) {\n+\t    die \"failed to find tree for $words[1]\\n\";\n+\t}\n+\tadd_delta($parent_tree, $cur_tree);\n+\tprint \"$cur\\n\";\n+\tlast;\n+    }\n+    close(PARENT);\n+    if (!defined($parent_commit)) {\n+        print STDERR \"parentless commit $cur\\n\";\n+    }\n+    if ($? && ($ret = $? >> 8)) {\n+        die \"cat-file failed with $ret\";\n+    }\n+    if ($cur eq $stop) {\n+        last;\n+    }\n+}\n+close(RL);\n+\n+if ($? && ($ret = $? >> 8)) {\n+    die \"rev-list failed with $ret\";\n+}\n+\ndiff -urN linus.delta/Makefile linus/Makefile\n--- linus.delta/Makefile\t2005-05-17 16:52:56.698760896 -0400\n+++ linus/Makefile\t2005-05-17 16:50:48.335275112 -0400\n@@ -22,7 +22,7 @@\n \tgit-unpack-file git-export git-diff-cache git-convert-cache \\\n \tgit-http-pull git-rpush git-rpull git-rev-list git-mktag \\\n \tgit-diff-tree-helper git-tar-tree git-local-pull git-write-blob \\\n-\tgit-mkdelta\n+\tgit-mkdelta git-pack\n \n all: $(PROG)\n \ndiff -urN linus.delta/mkdelta.c linus/mkdelta.c\n--- linus.delta/mkdelta.c\t2005-05-17 16:52:56.700760592 -0400\n+++ linus/mkdelta.c\t2005-05-16 17:21:37.000000000 -0400\n@@ -95,7 +95,7 @@\n \tunsigned long mapsize;\n \tvoid *map = map_sha1_file(sha1, &mapsize);\n \tif (map) {\n-\t\tvoid *buffer = unpack_sha1_file(map, mapsize, type, size);\n+\t\tvoid *buffer = raw_unpack_sha1_file(map, mapsize, type, size);\n \t\tmunmap(map, mapsize);\n \t\tif (buffer)\n \t\t\treturn buffer;\ndiff -urN linus.delta/object.c linus/object.c\n--- linus.delta/object.c\t2005-05-17 16:51:51.419684824 -0400\n+++ linus/object.c\t2005-05-17 15:16:52.291084064 -0400\n@@ -107,7 +107,7 @@\n \t\tstruct object *obj;\n \t\tchar type[100];\n \t\tunsigned long size;\n-\t\tvoid *buffer = unpack_sha1_file(map, mapsize, type, &size);\n+\t\tvoid *buffer = unpack_sha1_file(sha1, map, mapsize, type, &size, NULL, NULL);\n \t\tmunmap(map, mapsize);\n \t\tif (!buffer)\n \t\t\treturn NULL;\ndiff -urN linus.delta/pack.c linus/pack.c\n--- linus.delta/pack.c\t1969-12-31 19:00:00.000000000 -0500\n+++ linus/pack.c\t2005-05-17 17:13:44.615048784 -0400\n@@ -0,0 +1,122 @@\n+/*\n+ * pack and delta files in a GIT database\n+ * (C) 2005 Chris Mason <mason@suse.com>\n+ * This code is free software; you can redistribute it and/or modify\n+ * it under the terms of the GNU General Public License version 2 as\n+ * published by the Free Software Foundation.\n+ */\n+#include \"cache.h\"\n+#include \"delta.h\"\n+\n+static char *pack_usage = \"pack [ --max-depth=N ] [<reference_sha1>:]<target_sha1> [ <next_sha1> ... ]\";\n+\n+static int pack_sha1(unsigned char *sha1, unsigned char *refsha1, \n+\t\t     struct packed_item **head, struct packed_item **tail, \n+\t\t     unsigned long *packed_size, unsigned long *packed_nr, \n+\t\t     int max_depth)\n+{\n+\tstruct packed_item *item;\n+\tchar *buffer;\n+\tunsigned long size;\n+\tint ret;\n+\tchar type[20];\n+\tunsigned char retsha1[20];\n+\n+\tbuffer = read_sha1_file(sha1, type, &size);\n+\tif (!buffer) {\n+\t\tfprintf(stderr, \"failed to read %s\\n\", sha1_to_hex(sha1));\n+\t\treturn -1;\n+\t}\n+\tret = pack_sha1_buffer(buffer, size, type, retsha1, refsha1, &item, max_depth);\n+\tfree(buffer);\n+\tif (memcmp(sha1, retsha1, 20)) {\n+\t\tfprintf(stderr, \"retsha1 %s \", sha1_to_hex(retsha1));\n+\t\tfprintf(stderr, \"sha1 %s\\n\", sha1_to_hex(sha1));\n+\t\treturn -1;\n+\t}\n+\tif (memcmp(item->sha1, sha1, 20)) {\n+\t\tfprintf(stderr, \"item sha1 %s \", sha1_to_hex(item->sha1));\n+\t\tfprintf(stderr, \"sha1 %s\\n\", sha1_to_hex(sha1));\n+\t\treturn -1;\n+\t}\n+\tif (ret)\n+\t\treturn ret;\n+\tif (item) {\n+\t\tif (*tail)\n+\t\t\t(*tail)->next = item;\n+\t\t*tail = item;\n+\t\tif (!*head)\n+\t\t\t*head = item;\n+\t\t*packed_size += item->len;\n+\t\t(*packed_nr)++;\n+\t\tif (*packed_size > (512 * 1024) || *packed_nr > 1024) {\n+\t\t\tret = write_packed_list(*head);\n+\t\t\tif (ret)\n+\t\t\t\treturn ret;\n+\t\t\t*head = NULL;\n+\t\t\t*tail = NULL;\n+\t\t\t*packed_size = 0;\n+\t\t\t*packed_nr = 0;\n+\t\t}\n+\t}\n+\treturn 0;\n+}\n+int main(int argc, char **argv)\n+{\n+\tint i;\n+\tstruct packed_item *head = NULL;\n+\tstruct packed_item *tail = NULL;\n+\tunsigned long packed_size = 0;\n+\tunsigned long packed_nr = 0;\n+\tint verbose;\n+\tint depth_max = 16;\n+\tint ret;\n+\n+\tfor (i = 1; i < argc; i++) {\n+\t\tif (!strcmp(argv[i], \"-v\")) {\n+\t\t\tverbose = 1;\n+\t\t} else if (!strcmp(argv[i], \"-d\") && i+1 < argc) {\n+\t\t\tdepth_max = atoi(argv[++i]);\n+\t\t} else if (!strncmp(argv[i], \"--max-depth=\", 12)) {\n+\t\t\tdepth_max = atoi(argv[i]+12);\n+\t\t} else\n+\t\t\tbreak;\n+\t}\n+\tif (i == argc)\n+\t\tusage(pack_usage);\n+\twhile(i < argc) {\n+\t\tunsigned char sha1[20];\n+\t\tunsigned char refsha1[20];\n+\t\tunsigned char *target;\n+\t\tunsigned char *ref = NULL;\n+\t\ttarget = strchr(argv[i], ':');\n+\t\tif (target) {\n+\t\t\t*target = '\\0';\n+\t\t\ttarget++;\n+\t\t\tref = argv[i];\n+\t\t} else {\n+\t\t\ttarget = argv[i];\n+\t\t}\n+\t\tif (get_sha1_hex(target, sha1)) {\n+\t\t\tfprintf(stderr, \"unable to parse sha1 %s\\n\", argv[i]);\n+\t\t\texit(1);\n+\t\t}\n+\t\tif (ref) {\n+\t\t\tif (get_sha1_hex(ref, refsha1)) {\n+\t\t\t\tfprintf(stderr, \"unable to parse sha1 %s\\n\", argv[i]);\n+\t\t\t\texit(1);\n+\t\t\t}\n+\t\t\tref = refsha1;\n+\t\t}\n+\t\tret = pack_sha1(sha1, ref, &head, &tail, &packed_size, &packed_nr, depth_max);\n+\t\tif (ret) {\n+\t\t\tfprintf(stderr, \"pack_sha1 failed! %d\\n\", ret);\n+\t\t\texit(1);\n+\t\t}\n+\t\ti++;\n+\t}\n+\n+\tif (head)\n+\t\twrite_packed_list(head);\n+\treturn 0;\n+}\ndiff -urN linus.delta/sha1_file.c linus/sha1_file.c\n--- linus.delta/sha1_file.c\t2005-05-17 16:52:56.697761048 -0400\n+++ linus/sha1_file.c\t2005-05-17 18:16:54.004973952 -0400\n@@ -180,39 +180,132 @@\n \treturn map;\n }\n \n-void * unpack_sha1_file(void *map, unsigned long mapsize, char *type, unsigned long *size)\n+/*\n+ * looks through buf for the header entry corresponding to sha1.  returns\n+ * 0 an entry is found and sets offset to the offset of the packed item\n+ * in the file.  The offset is relative to the start of the packed items\n+ * so you have to add in the length of the header before using it\n+ * -1 is returned if the sha1 could not be found\n+ */\n+static int find_packed_header(const unsigned char *sha1, char *buf, unsigned long buf_len, unsigned long *offset)\n+{\n+\tchar *p;\n+\tp = buf;\n+\n+\t*offset = 0;\n+\twhile(p < buf + buf_len) {\n+\t\tunsigned long item_len;\n+\t\tunsigned char item_sha[20];\n+\t\tmemcpy(item_sha, p, 20);\n+\t\tsscanf(p + 20, \"%lu\", &item_len);\n+\t\tp += 20 + strlen(p + 20) + 1;\n+\t\tif (memcmp(item_sha, sha1, 20) == 0)\n+\t\t\treturn 0;\n+\t\t*offset += item_len;\n+\t}\n+\treturn -1;\n+}\n+\n+\n+/*\n+ * uncompresses a data segment without any extra delta/packed processing\n+ */\n+static void * _unpack_sha1_file(z_stream *stream, void *map, \n+                                unsigned long mapsize, char *type, \n+\t\t\t\tunsigned long *size)\n {\n \tint ret, bytes;\n-\tz_stream stream;\n \tchar buffer[8192];\n \tchar *buf;\n \n \t/* Get the data stream */\n-\tmemset(&stream, 0, sizeof(stream));\n-\tstream.next_in = map;\n-\tstream.avail_in = mapsize;\n-\tstream.next_out = buffer;\n-\tstream.avail_out = sizeof(buffer);\n-\n-\tinflateInit(&stream);\n-\tret = inflate(&stream, 0);\n-\tif (ret < Z_OK)\n+\tmemset(stream, 0, sizeof(*stream));\n+\tstream->next_in = map;\n+\tstream->avail_in = mapsize;\n+\tstream->next_out = buffer;\n+\tstream->avail_out = sizeof(buffer);\n+\n+\tinflateInit(stream);\n+\tret = inflate(stream, 0);\n+\tif (ret < Z_OK) {\n \t\treturn NULL;\n-\tif (sscanf(buffer, \"%10s %lu\", type, size) != 2)\n+\t}\n+\tif (sscanf(buffer, \"%10s %lu\", type, size) != 2) {\n \t\treturn NULL;\n-\n+\t}\n \tbytes = strlen(buffer) + 1;\n \tbuf = xmalloc(*size);\n \n-\tmemcpy(buf, buffer + bytes, stream.total_out - bytes);\n-\tbytes = stream.total_out - bytes;\n+\tmemcpy(buf, buffer + bytes, stream->total_out - bytes);\n+\tbytes = stream->total_out - bytes;\n \tif (bytes < *size && ret == Z_OK) {\n-\t\tstream.next_out = buf + bytes;\n-\t\tstream.avail_out = *size - bytes;\n-\t\twhile (inflate(&stream, Z_FINISH) == Z_OK)\n+\t\tstream->next_out = buf + bytes;\n+\t\tstream->avail_out = *size - bytes;\n+\t\twhile (inflate(stream, Z_FINISH) == Z_OK)\n \t\t\t/* nothing */;\n \t}\n-\tinflateEnd(&stream);\n+\tinflateEnd(stream);\n+\treturn buf;\n+}\n+\n+void * raw_unpack_sha1_file(void *map, unsigned long mapsize, char *type, unsigned long *size)\n+{\n+\tz_stream stream;\n+\treturn _unpack_sha1_file(&stream, map, mapsize, type, size);\n+}\n+\n+void * unpack_sha1_file(const unsigned char *sha1, void *map, \n+\t\t\tunsigned long mapsize, char *type, unsigned long *size, \n+\t\t\tconst unsigned char *recur_sha1,\n+\t\t\tint *chain)\n+{\n+\tz_stream stream;\n+\tchar *buf;\n+\tunsigned long offset;\n+\tunsigned long header_len;\n+\tbuf = _unpack_sha1_file(&stream, map, mapsize, type, size);\n+\tif (!buf)\n+\t\treturn buf;\n+\tif (!strcmp(type, \"delta\")) {\n+\t\tchar *delta_ref;\n+\t\tunsigned long delta_size;\n+\t\tchar *newbuf;\n+\t\tunsigned long newsize;\n+\t\tif (recur_sha1 && memcmp(buf, recur_sha1, 20) == 0) {\n+\t\t\tfree(buf);\n+\t\t\treturn NULL;\n+\t\t}\n+\t\tif (chain)\n+\t\t\t*chain += 1;\n+\t\tdelta_ref = read_sha1_delta_ref(buf, type, &delta_size, recur_sha1, chain);\n+\t\tif (!delta_ref) {\n+\t\t\tfprintf(stderr, \"failed to read delta %s\\n\", sha1_to_hex(buf));\n+\t\t\tfree(buf);\n+\t\t\treturn NULL;\n+\t\t}\n+\t\tnewbuf = patch_delta(delta_ref, delta_size, buf+20, *size-20, &newsize);\n+\t\tif (!newbuf) {\n+\t\t\tfprintf(stderr, \"patch_delta failed %s %lu\\n\", sha1_to_hex(buf), delta_size);\n+\t\t}\n+\t\tfree(buf);\n+\t\tfree(delta_ref);\n+\t\t*size = newsize;\n+\t\treturn newbuf;\n+\n+\t} else if (!strcmp(type, \"packed\")) {\n+\t\tif (!sha1) {\n+\t\t\tfree(buf);\n+\t\t\treturn NULL;\n+\t\t}\n+\t\theader_len = *size;\n+\t\tif (find_packed_header(sha1, buf, header_len, &offset)) {\n+\t\t\tfree(buf);\n+\t\t\treturn NULL;\n+\t\t}\n+\t\toffset += stream.total_in;\n+\t\tfree(buf);\n+\t\tbuf = unpack_sha1_file(sha1, map+offset, mapsize-offset, type, size, recur_sha1, chain);\n+\t}\n \treturn buf;\n }\n \n@@ -223,21 +316,26 @@\n \n \tmap = map_sha1_file(sha1, &mapsize);\n \tif (map) {\n-\t\tbuf = unpack_sha1_file(map, mapsize, type, size);\n+\t\tbuf = unpack_sha1_file(sha1, map, mapsize, type, size, NULL, NULL);\n+\t\tmunmap(map, mapsize);\n+\t\treturn buf;\n+\t}\n+\treturn NULL;\n+}\n+\n+/*\n+ * the same as read_sha1_file except chain is used to count the length\n+ * of any delta chains hit while unpacking\n+ */\n+void * read_sha1_delta_ref(const unsigned char *sha1, char *type, unsigned long *size, const unsigned char *recur_sha1, int *chain)\n+{\n+\tunsigned long mapsize;\n+\tvoid *map, *buf;\n+\n+\tmap = map_sha1_file(sha1, &mapsize);\n+\tif (map) {\n+\t\tbuf = unpack_sha1_file(sha1, map, mapsize, type, size, recur_sha1, chain);\n \t\tmunmap(map, mapsize);\n-\t\tif (buf && !strcmp(type, \"delta\")) {\n-\t\t\tvoid *ref = NULL, *delta = buf;\n-\t\t\tunsigned long ref_size, delta_size = *size;\n-\t\t\tbuf = NULL;\n-\t\t\tif (delta_size > 20)\n-\t\t\t\tref = read_sha1_file(delta, type, &ref_size);\n-\t\t\tif (ref)\n-\t\t\t\tbuf = patch_delta(ref, ref_size,\n-\t\t\t\t\t\t  delta+20, delta_size-20, \n-\t\t\t\t\t\t  size);\n-\t\t\tfree(delta);\n-\t\t\tfree(ref);\n-\t\t}\n \t\treturn buf;\n \t}\n \treturn NULL;\n@@ -482,3 +580,322 @@\n \t\tmunmap(buf, size);\n \treturn ret;\n }\n+\n+static void *compress_buffer(void *buf, unsigned long buf_len, char *metadata, \n+                             int metadata_size, unsigned long *compsize)\n+{\n+\tchar *compressed;\n+\tz_stream stream;\n+\tunsigned long size;\n+\n+\t/* Set it up */\n+\tmemset(&stream, 0, sizeof(stream));\n+\tsize = deflateBound(&stream, buf_len + metadata_size);\n+\tcompressed = xmalloc(size);\n+\n+\t/*\n+\t * ASCII size + nul byte\n+\t */\t\n+\tstream.next_in = metadata;\n+\tstream.avail_in = metadata_size;\n+\tstream.next_out = compressed;\n+\tstream.avail_out = size;\n+\tdeflateInit(&stream, Z_BEST_COMPRESSION);\n+\twhile (deflate(&stream, 0) == Z_OK)\n+\t\t/* nothing */;\n+\n+\tstream.next_in = buf;\n+\tstream.avail_in = buf_len;\n+\t/* Compress it */\n+\twhile (deflate(&stream, Z_FINISH) == Z_OK)\n+\t\t/* nothing */;\n+\tdeflateEnd(&stream);\n+\tsize = stream.total_out;\n+\t*compsize = size;\n+\treturn compressed;\n+}\n+\n+/*\n+ * generates a delta for buf against refsha1 and returns a compressed buffer\n+ * with the results.  NULL is returned on error, or when the delta could\n+ * not be done.  This might happen if the delta is larger then either the\n+ * refsha1 or the buffer, or the delta chain is too long.\n+ */\n+void *delta_buffer(void *buf, unsigned long buf_len, char *metadata, \n+                   int metadata_size, unsigned long *compsize, \n+\t\t   unsigned char *sha1, unsigned char *refsha1, int max_chain)\n+{\n+\tchar *compressed;\n+\tchar *refbuffer = NULL;\n+\tchar reftype[20];\n+\tunsigned long refsize = 0;\n+\tchar *delta;\n+\tunsigned long delta_size;\n+\tchar *lmetadata = xmalloc(220);\n+\tunsigned long lmetadata_size;\n+\tint chain_length = 0;\n+\n+\tif (buf_len == 0)\n+\t\treturn NULL;\n+\trefbuffer = read_sha1_delta_ref(refsha1, reftype, &refsize, sha1, &chain_length);\n+\n+\tif (chain_length > max_chain) {\n+\t\tfree(refbuffer);\n+\t\treturn NULL;\n+\t}\n+\t/* note, we could just continue without the delta here */\n+\tif (!refbuffer) {\n+\t\tfree(refbuffer);\n+\t\treturn NULL;\n+\t}\n+\tdelta = diff_delta(refbuffer, refsize, buf, buf_len, &delta_size);\n+\tfree(refbuffer);\n+\tif (!delta)\n+\t\treturn NULL;\n+\tif (delta_size > refsize || delta_size > buf_len) {\n+\t\tfree(delta);\n+\t\treturn NULL;\n+\t}\n+\tif (delta_size < 10) {\n+\t\tfree(delta);\n+\t\treturn NULL;\n+\t}\n+\tlmetadata_size = 1 + sprintf(lmetadata, \"%s %lu\",\"delta\",delta_size+20);\n+\tmemcpy(lmetadata + lmetadata_size, refsha1, 20);\n+\tlmetadata_size += 20;\n+\tcompressed = compress_buffer(delta, delta_size, lmetadata, lmetadata_size, compsize);\n+\tfree(lmetadata);\n+\tfree(delta);\n+\treturn compressed;\n+}\n+\n+/*\n+ * returns a newly malloc'd packed item with a compressed buffer for buf.  \n+ * If refsha1 is non-null, attempts a delta against it.  The sha1 of buf \n+ * is returned via returnsha1.\n+ */\n+int pack_sha1_buffer(void *buf, unsigned long buf_len, char *type,\n+\t\t     unsigned char *returnsha1,\n+\t\t     unsigned char *refsha1,\n+\t\t     struct packed_item **packed_item, int max_depth)\n+{\n+\tunsigned char sha1[20];\n+\tSHA_CTX c;\n+\tchar *compressed = NULL;\n+\tunsigned long size;\n+\tstruct packed_item *item;\n+\tchar *metadata = xmalloc(200);\n+\tint metadata_size;\n+\n+\t*packed_item = NULL;\n+\n+\tmetadata_size = 1 + sprintf(metadata, \"%s %lu\", type, buf_len);\n+\n+\t/* Sha1.. */\n+\tSHA1_Init(&c);\n+\tSHA1_Update(&c, metadata, metadata_size);\n+\tSHA1_Update(&c, buf, buf_len);\n+\tSHA1_Final(sha1, &c);\n+\n+\tif (returnsha1)\n+\t\tmemcpy(returnsha1, sha1, 20);\n+\n+\tif (refsha1)\n+\t\tcompressed = delta_buffer(buf, buf_len, metadata, \n+\t\t                          metadata_size, &size, sha1, \n+\t\t\t\t\t  refsha1, max_depth);\n+\tif (!compressed)\n+\t\tcompressed = compress_buffer(buf, buf_len, metadata, \n+\t\t                             metadata_size, &size);\n+\tfree(metadata);\n+\tif (!compressed)\n+\t\treturn -1;\n+\n+\titem = xmalloc(sizeof(struct packed_item));\n+\tmemcpy(item->sha1, sha1, 20);\n+\titem->len = size;\n+\titem->next = NULL;\n+\titem->data = compressed;\n+\t*packed_item = item;\n+\treturn 0;\n+}\n+\n+static char *create_packed_header(struct packed_item *head, unsigned long *size)\n+{\n+\tchar *metadata = NULL;\n+\tint metadata_size = 0;\n+\t*size = 0;\n+\tint entry_size = 0;\n+\n+\twhile(head) {\n+\t\tchar *p;\n+\t\tmetadata = realloc(metadata, metadata_size + 220);\n+\t\tif (!metadata)\n+\t\t\treturn NULL;\n+\t\tp = metadata+metadata_size;\n+\t\tmemcpy(p, head->sha1, 20);\n+\t\tp += 20;\n+\t\tentry_size = 1 + sprintf(p, \"%lu\", head->len);\n+\t\tmetadata_size += entry_size + 20;\n+\t\thead = head->next;\n+\t}\n+\t*size = metadata_size;\n+\treturn metadata;\n+}\n+\n+#define WRITE_BUFFER_SIZE 8192\n+static char write_buffer[WRITE_BUFFER_SIZE];\n+static unsigned long write_buffer_len;\n+\n+static int c_write(int fd, void *data, unsigned int len)\n+{\n+\twhile (len) {\n+\t\tunsigned int buffered = write_buffer_len;\n+\t\tunsigned int partial = WRITE_BUFFER_SIZE - buffered;\n+\t\tif (partial > len)\n+\t\t\tpartial = len;\n+\t\tmemcpy(write_buffer + buffered, data, partial);\n+\t\tbuffered += partial;\n+\t\tif (buffered == WRITE_BUFFER_SIZE) {\n+\t\t\tif (write(fd, write_buffer, WRITE_BUFFER_SIZE) != WRITE_BUFFER_SIZE)\n+\t\t\t\treturn -1;\n+\t\t\tbuffered = 0;\n+\t\t}\n+\t\twrite_buffer_len = buffered;\n+\t\tlen -= partial;\n+\t\tdata += partial;\n+ \t}\n+ \treturn 0;\n+}\n+\n+static int c_flush(int fd)\n+{\n+\tif (write_buffer_len) {\n+\t\tint left = write_buffer_len;\n+\t\tif (write(fd, write_buffer, left) != left)\n+\t\t\treturn -1;\n+\t\twrite_buffer_len = 0;\n+\t}\n+\treturn 0;\n+}\n+\n+/*\n+ * creates a new packed file for all the items in head.  hard links are\n+ * made from the sha1 of all the items back to the packd file, and then\n+ * the packed file is unlinked.\n+ */\n+int write_packed_list(struct packed_item *head)\n+{\n+\tunsigned char sha1[20];\n+\tSHA_CTX c;\n+\tchar filename[PATH_MAX];\n+\tchar *metadata = xmalloc(200);\n+\tchar *header;\n+\tint metadata_size;\n+\tint fd;\n+\tint ret = 0;\n+\tunsigned long header_len;\n+\tstruct packed_item *item;\n+\tchar *compressed;\n+\tz_stream stream;\n+\tunsigned long size;\n+\n+\theader = create_packed_header(head, &header_len);\n+\tmetadata_size = 1+sprintf(metadata, \"packed %lu\", header_len);\n+\t/* \n+\t * the header contains the sha1 of each item, so we only sha1 the\n+\t * header\n+\t */ \n+\tSHA1_Init(&c);\n+\tSHA1_Update(&c, metadata, metadata_size);\n+\tSHA1_Update(&c, header, header_len);\n+\tSHA1_Final(sha1, &c);\n+\n+\tif (access(sha1_file_name(sha1), F_OK) == 0)\n+\t\tgoto out_nofile;\n+\n+\tsnprintf(filename, sizeof(filename), \"%s/obj_XXXXXX\", get_object_directory());\n+\tfd = mkstemp(filename);\n+\tif (fd < 0) {\n+\t\tret = -errno;\n+\t\tgoto out_nofile;\n+\t}\n+\n+       /* compress just the header info */\n+        memset(&stream, 0, sizeof(stream));\n+        deflateInit(&stream, Z_BEST_COMPRESSION);\n+\tsize = deflateBound(&stream, header_len + metadata_size);\n+        compressed = xmalloc(size);\n+\n+        stream.next_in = metadata;\n+        stream.avail_in = metadata_size;\n+        stream.next_out = compressed;\n+        stream.avail_out = size;\n+        while (deflate(&stream, 0) == Z_OK)\n+                /* nothing */;\n+        stream.next_in = header;\n+        stream.avail_in = header_len;\n+        while (deflate(&stream, Z_FINISH) == Z_OK)\n+                /* nothing */;\n+        deflateEnd(&stream);\n+        size = stream.total_out;\n+\n+\tc_write(fd, compressed, size);\n+\tfree(compressed);\n+\n+\titem = head;\n+\twhile(item) {\n+\t\tif (c_write(fd, item->data, item->len)) {\n+\t\t\tret = -EIO;\n+\t\t\tgoto out;\n+\t\t}\n+\t\titem = item->next;\n+\t}\n+\tif (c_flush(fd)) {\n+\t\tret = -EIO;\n+\t\tgoto out;\n+\t}\n+\titem = head;\n+\twhile(item) {\n+\t\tchar *item_file;\n+\t\tchar item_tmp[PATH_MAX];\n+\t\tstruct packed_item *next = item->next;\n+\t\tint name_iter = 0;\n+\t\titem_file = sha1_file_name(item->sha1);\n+\t\twhile(1) {\n+\t\t\t/* ugly stuff.  We want to atomically replace any old objects\n+\t\t\t * with the same sha1, making sure they don't get deleted\n+\t\t\t * if any step along the way fails\n+\t\t\t */\n+\t\t\tsnprintf(item_tmp, sizeof(item_tmp), \"%s/obj_%d\", get_object_directory(), name_iter);\n+\t\t\tif (link(filename, item_tmp)) {\n+\t\t\t\tif (errno != EEXIST) {\n+\t\t\t\t\tret = -errno;\n+\t\t\t\t\tgoto out;\n+\t\t\t\t}\n+\t\t\t} else {\n+\t\t\t\t/* link success */\n+\t\t\t\tif (rename(item_tmp, item_file)) {\n+\t\t\t\t\tret = -errno;\n+\t\t\t\t\tgoto out;\n+\t\t\t\t}\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t\tif (name_iter++ > 1000) {\n+\t\t\t\tret = -1;\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\t\t}\n+\t\tfree(item->data);\n+\t\tfree(item);\n+\t\titem = next;\n+\t}\n+out:\n+\tunlink(filename);\n+\tfchmod(fd, 0444);\n+\tclose(fd);\n+out_nofile:\n+\tfree(header);\n+\tfree(metadata);\n+\treturn ret;\n+}\n"},{"id":"3566","messageId":"200505191428.52238.mason@suse.com","threadId":"645","inReplyTo":"200505171857.46370.mason@suse.com","subject":"Re: [PATCH] packed delta git","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-05-19T18:28:50Z","receivedAt":"2005-05-19T18:28:50Z","isPatch":true,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Tuesday 17 May 2005 18:57, Chris Mason wrote:\n> Hello everyone,\n>\n> Here's a new version of my packed git patch, diffed on top of Nicolas'\n> delta code (link below to that).  It doesn't change the core git commands\n> to create packed/delta files, that is done via a new git-pack command.  The\n> git-pack usage is very simple:\n>\n> git-pack [<reference_sha1>:]<target_sha1> [ <next_sha1> ... ]\n\nMy original git-pack-changes script didn't properly limit the length of the \ndelta chains, so you could use it to create a repo that you can't later read.\n\nThe new one below fixes that, and also changes the direction of the delta.\nDeltas are now done in reverse, leaving the most recent sha1 as a whole file\nand diffing old revisions against it.\n\nThe result is the same size (62M for current linux-2.6 git tree) and faster\ncheckout times for head (9s vs 15s).  I also tested against the bk-cvs patch\nset:\n                                      vanilla               packed/delta\ncheckout-cache (hot)      (only 1.5G ram) 15s\ncheckout-cache (cold)     4m30s               1m19s  \nsize (du -sh .git)              2.5G                  227M\n\nThe steps to pack/delta the 2.6 git tree have changed:\n\n# step one, pack all of HEAD together\ngit-ls-tree -r HEAD | awk '{print $3}' | xargs git-pack\n\n# step two pack deltas for all revs back to the first commit\ngit-pack-changes-script | xargs git-pack\n\n# step three, pack a delta from 2.6.12-rc2 to 2.6.11\ngit-pack-changes-script -t 1da177e4c3f41524e886b7f1b8a0c1fc7321cac2 c39ae07f393806ccf406ef966e9a15afc43cc36a | xargs git-pack\n\nNew git-pack-changes-script:\n\n--\n#!/usr/bin/perl\n#\n# script to search through the rev-list output and generate delta history\n# you can specify either a start and stop commit or two trees to search.\n# with no command line args it searches the entire revision history.\n# output is suitable for piping to xargs git-pack\n\nuse strict;\n\nmy $ret;\nmy $i;\nmy @wanted = ();\nmy $argc = scalar(@ARGV);\nmy $commit;\nmy $stop;\nmy %delta = ();\nmy %packed = ();\n\nsub add_packed($) {\n    my ($sha1) = @_;\n    if (defined($packed{$sha1})) {\n        return 1;\n    }\n    if (defined($delta{$sha1})) {\n        return 1;\n    }\n    $packed{$sha1} = 1;\n    print \"$sha1\\n\";\n    return 0;\n}\n\nsub add_delta($$) {\n    my ($ref, $target) = @_;\n    my $chain = 0;\n    my $recur = $ref;\n    if (defined($delta{$target})) {\n        return 1;\n    }\n    if (defined($packed{$target})) {\n        return 1;\n    }\n    while(1) {\n\tlast if (!defined($delta{$recur}));\n\tif ($target eq $delta{$recur}) {\n\t    add_packed($target);\n\t    return 1;\n\t}\n\t$chain++;\n\tif ($chain > 32) {\n\t    add_packed($target);\n\t    return 1;\n\t}\n\t$recur = $delta{$recur};\n    }\n    $delta{$target} = $ref;\n    print \"$ref:$target\\n\";\n    return 0;\n}\n\nsub print_usage() {\n    print STDERR \"usage: pack-changes [-c commit] [-s stop commit] [-t tree1 tree2]\\n\";\n    exit(1);\n}\n\nsub find_tree($) {\n    my ($commit) = @_;\n    open(CM, \"git-cat-file commit $commit|\") || die \"git-cat-file failed\";\n    while(<CM>) {\n        chomp;\n\tmy @words = split;\n\tif ($words[0] eq \"tree\") {\n\t    return $words[1];\n\t} elsif ($words[0] ne \"parent\") {\n\t    last;\n\t}\n    }\n    close(CM);\n    if ($? && ($ret = $? >> 8)) {\n        die \"cat-file $commit failed with $ret\";\n    }\n    return undef;\n}\n\nsub test_diff($$) {\n    my ($a, $b) = @_;\n    open(DT, \"git-diff-tree -r -t $a $b|\") || die \"diff-tree failed\";\n    while(<DT>) {\n        chomp;\n\tmy @words = split;\n\tmy $sha1 = $words[2];\n\tmy $change = $words[0];\n\tif ($change =~ m/^\\*/) {\n\t    @words = split(\"->\", $sha1);\n\t    add_delta($words[0], $words[1]);\n\t} elsif ($change =~ m/^\\-/) {\n\t    next;\n\t} else {\n\t    add_packed($sha1);\n\t}\n    }\n    close(DT);\n    if ($? && ($ret = $? >> 8)) {\n\tdie \"git-diff-tree failed with $ret\";\n    }\n    return 0;\n}\n\nfor ($i = 0 ; $i < $argc ; $i++)  {\n    if ($ARGV[$i] eq \"-c\") {\n    \tif ($i == $argc - 1) {\n\t    print_usage();\n\t}\n\t$commit = $ARGV[++$i];\n    } elsif ($ARGV[$i] eq \"-s\") {\n    \tif ($i == $argc - 1) {\n\t    print_usage();\n\t}\n\t$stop = $ARGV[++$i];\n    } elsif ($ARGV[$i] eq \"-t\") {\n        if ($argc != 3 || $i != 0) {\n\t    print_usage();\n\t}\n\tif (test_diff($ARGV[1], $ARGV[2])) {\n\t    die \"test_diff failed\\n\";\n\t}\n\tadd_delta($ARGV[1], $ARGV[2]);\n\texit(0);\n    }\n}\n\nif (!defined($commit)) {\n    $commit = `commit-id`;\n    if ($?) {\n    \tprint STDERR \"commit-id failed, try using -c to specify a commit\\n\";\n\texit(1);\n    }\n    chomp $commit;\n}\n\nopen(RL, \"git-rev-list $commit|\") || die \"rev-list failed\";\nwhile(<RL>) {\n    chomp;\n    my $cur = $_;\n    my $cur_tree;\n    my $parent_tree;\n    my $parent_commit = undef;\n    open(PARENT, \"git-cat-file commit $cur|\") || die \"cat-file failed\";\n    while(<PARENT>) {\n        chomp;\n\tmy @words = split;\n\tif ($words[0] eq \"tree\") {\n\t    $cur_tree = $words[1];\n\t    next;\n\t} elsif ($words[0] ne \"parent\") {\n\t    last;\n\t}\n\t$parent_commit = $words[1];\n\tmy $next = <PARENT>;\n\t# ignore merge sets for now\n\tif ($next =~ m/^parent/) {\n\t    last;\n\t}\n\t# note that we run test_diff to generate a reverse\n\t# diff\n\tif (test_diff($cur, $words[1])) {\n\t    die \"test_diff failed\\n\";\n\t}\n\t$parent_tree = find_tree($words[1]);\n\tif (!defined($parent_tree)) {\n\t    die \"failed to find tree for $words[1]\\n\";\n\t}\n\tadd_delta($cur_tree, $parent_tree);\n\tadd_packed($cur);\n\tlast;\n    }\n    close(PARENT);\n    if (!defined($parent_commit)) {\n        print STDERR \"parentless commit $cur\\n\";\n    }\n    if ($? && ($ret = $? >> 8)) {\n        die \"cat-file failed with $ret\";\n    }\n    if ($cur eq $stop) {\n        last;\n    }\n}\nclose(RL);\n\nif ($? && ($ret = $? >> 8)) {\n    die \"rev-list failed with $ret\";\n}\n\n"},{"id":"3568","messageId":"20050519183810.GF8105@cip.informatik.uni-erlangen.de","threadId":"645","inReplyTo":"200505191428.52238.mason@suse.com","subject":"Re: [PATCH] packed delta git","fromName":"Thomas Glanzmann","fromEmail":"sithglan@stud.uni-erlangen.de","sentAt":"2005-05-19T18:38:10Z","receivedAt":"2005-05-19T18:38:10Z","isPatch":true,"sender":{"key":"sithglan@stud.uni-erlangen.de","avatar":null},"body":"Hello Chris,\n\n> size (du -sh .git)              2.5G                  227M\n\nwow that beats bitkeeper in size. What is missing to actual use such a\napproach in a distributed environment?\n\n\tThomas\n"},{"id":"3572","messageId":"200505191453.21566.mason@suse.com","threadId":"645","inReplyTo":"20050519183810.GF8105@cip.informatik.uni-erlangen.de","subject":"Re: [PATCH] packed delta git","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-05-19T18:53:21Z","receivedAt":"2005-05-19T18:53:21Z","isPatch":true,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Thursday 19 May 2005 14:38, Thomas Glanzmann wrote:\n> Hello Chris,\n>\n> > size (du -sh .git)              2.5G                  227M\n>\n> wow that beats bitkeeper in size. What is missing to actual use such a\n> approach in a distributed environment?\n\nIt's not quite fair to compare with bitkeeper, since my changeset comments are \nonly the name of the bk->cvs patch, and I've only got 28k changesets vs bk's \n60k or so.\n\nIn terms of actually making use of this, we need to deal with the hard linked \nfiles during push/pull.  This means using -H on rsync and teaching the \npush/pull code about packed files.\n\ngit-pack needs to be able to unpack/undelta files so that people can clean a \ntree.\n\ngit-fsck-cache needs to understand packed files and deltas.\n\n-chris\n"},{"id":"3575","messageId":"Pine.LNX.4.62.0505191529290.20274@localhost.localdomain","threadId":"645","inReplyTo":"20050519183810.GF8105@cip.informatik.uni-erlangen.de","subject":"Re: [PATCH] packed delta git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2005-05-19T19:30:06Z","receivedAt":"2005-05-19T19:30:06Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 19 May 2005, Thomas Glanzmann wrote:\n\n> Hello Chris,\n> \n> > size (du -sh .git)              2.5G                  227M\n> \n> wow that beats bitkeeper in size. What is missing to actual use such a\n> approach in a distributed environment?\n\nMe completing fsck-cache support for delta objects.\n\n\nNicolas\n"}]}