{"thread":{"id":"11972","subject":"[PATCH] RFC: git lazy clone proof-of-concept","startedAt":"2008-02-08T17:28:43Z","lastAt":"2008-02-17T18:44:01Z","messageCount":85,"participants":["Jan Holesovsky","Nicolas Pitre","Harvey Harrison","Johannes Schindelin","Mike Hommey","Jakub Narebski","Jon Smirl","Jan Hudec","Sean","Marco Costalba","Joachim B Haga","David Symonds","Andreas Ericsson","Brandon Casey","Linus Torvalds","Brian Downing","Shawn O. Pearce","Junio C Hamano"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"67974","messageId":"200802081828.43849.kendy@suse.cz","threadId":"11972","inReplyTo":null,"subject":"[PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2008-02-08T17:28:43Z","receivedAt":"2008-02-08T17:28:43Z","isPatch":true,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi,\n\nThis is my attempt to implement the 'lazy clone' I've read about a bit in the\ngit mailing list archive, but did not see implemented anywhere - the clone\nthat fetches a minimal amount of data with the possibility to download the\nrest later (transparently!) when necessary.  I am sorry to send it as a huge\npatch, not as a series of patches, but as I don't know if I chose a way that is\nacceptable for you [I'm new to the git code ;-)], I'd like to hear some\nfeedback first, and then I'll split it into smaller pieces for easier\nintegration - if OK.\n\nBackground:\n\nCurrently we are evaluating the usage of git for OpenOffice.org as one of the\ncandidates (SVN is the other one), see\n\n  http://wiki.services.openoffice.org/wiki/SCM_Migration\n\nI've provided a git import of OOo with the entire history; the problem is that\nthe pack has 2.5G, so it's not too convenient to download for casual\ndevelopers that just want to try it.  Shallow clone is not a possibility - we\ndon't get patches through mailing lists, so we need the pull/push, and also\nthanks to the OOo development cycle, we have too many living heads which\ncauses the shallow clone to download about 1.5G even with --depth 1.  Lazy\nclone sounded like the right idea to me.  With this proof-of-concept\nimplementation, just about 550M from the 2.5G is downloaded, which is still\nabout twice as much in comparison with downloading a tarball, but bearable.\n\nThe principle:\n\nDuring the initial clone, just the commit objects are downloaded.  Then, any\ntime an object is requested, it is downloaded from the remote repository if\nnot available locally.  To make this usable and performing, when a tree is\nrequested, it is downloaded together with all the subtrees and blobs at which\nit points.  Every subsequent pull (of stuff newer than what was cloned) is\nsupposed to use the normal git mechanisms.\n\nProtocol extensions:\n\nI've extended the git protocol in 2 ways:\n- added the 'commits-only' flag that is used during the clone to get a pack\n  containing just the commit objects, nothing else\n- added the 'exact-objects' flag that allows to request just few objects\n  exactly specified by the client\n\nA bit more detailed description:\n\nHere I use the term 'remote alternate' as the remote repository from which the\nobjects are downloaded when not locally available.\n\n- fetch-pack.h\n- builtin-fetch-pack.c\n  Added --commits-only, --exact-objects, and --stdin options.\n  --commits-only and --exact-objects trigger the protocol extensions described\n  above, --stdin allows fetch-pack to get the list of refs or objects on stdin\n  instead of the command line\n- transport.h\n- transport.c\n- builtin-fetch.c\n  Added --commits-only option that is passed to fetch-pack\n- builtin-unpack-objects.c\n- index-pack.c\n  Added --ignore-remote-alternates option which will avoid fetching remote\n  objects, to avoid a cycle in downloading the missing objects.\n- cache.h\n  Export the function that disables fetching remote objects.\n- git-clone.sh\n  Added handling of -s even in the case when the git:// protocol is used to\n  activate the 'remote alternates' and thus get a lazy clone.  The info where\n  to get the missing objects from is stored in the\n  objects/info/remote_alternates file.\n- sha1_file.c\n  The core of the changes.  When an object is requested, usage of 'remote\n  alternates' is on, and it is not present locally, it is downloaded.\n- upload-pack.c\n  Extended so that just the commit objects, or the exact objects are returned.\n\nLimitations/FIXMEs/TODOs:\n\nCurrently there can be just one 'remote alternate' in the\nobjects/info/remote_alternates file.  I'm not sure if it makes sense at all to\nprovide the possibility to have more of them.\n\nSome operations are too slow, like the annotate, and thus unusable [though not\ndisabled for the patient ones ;-)].\n\nNot too much tested ;-), maybe I'm leaking memory somewhere, better error\nhandling in case the pack is not available should be introduced, maybe the\nnames of the variables/functions/commands is not the best chosen, etc.\n\nEvery fetch-pack gets list of refs from the server even for the exact-objects\ncase which is unnecessary - we know what objects we want, this just wastes\nbandwidth.\n\nThe new options are not documented.\n\n\nSo - comments, ideas, questions appreciated, any help with polishing\nthis/getting this in is appreciated even more ;-)\n\nRegards,\nJan\n\nSigned-off-by: Jan Holesovsky <kendy@suse.cz>\n---\ndiff --git a/builtin-fetch-pack.c b/builtin-fetch-pack.c\nindex e68e015..69b4226 100644\n--- a/builtin-fetch-pack.c\n+++ b/builtin-fetch-pack.c\n@@ -17,7 +17,7 @@ static struct fetch_pack_args args = {\n };\n \n static const char fetch_pack_usage[] =\n-\"git-fetch-pack [--all] [--quiet|-q] [--keep|-k] [--thin] [--upload-pack=<git-upload-pack>] [--depth=<n>] [--no-progress] [-v] [<host>:]<directory> [<refs>...]\";\n+\"git-fetch-pack [--all] [--quiet|-q] [--keep|-k] [--thin] [--upload-pack=<git-upload-pack>] [--depth=<n>] [--no-progress] [--commits-only] [--exact-objects] [-v] [--stdin] [<host>:]<directory> [<refs>...|<sha1>...]\";\n \n #define COMPLETE\t(1U << 0)\n #define COMMON\t\t(1U << 1)\n@@ -141,6 +141,34 @@ static const unsigned char* get_rev(void)\n \treturn commit->object.sha1;\n }\n \n+static void send_want(int fd[2], const char *remote, int full_info)\n+{\n+\tif (full_info)\n+\t\tpacket_write(fd[1], \"want %s%s%s%s%s%s%s%s%s\\n\",\n+\t\t\t\tremote,\n+\t\t\t\t(multi_ack ? \" multi_ack\" : \"\"),\n+\t\t\t\t(use_sideband == 2 ? \" side-band-64k\" : \"\"),\n+\t\t\t\t(use_sideband == 1 ? \" side-band\" : \"\"),\n+\t\t\t\t(args.use_thin_pack ? \" thin-pack\" : \"\"),\n+\t\t\t\t(args.no_progress ? \" no-progress\" : \"\"),\n+\t\t\t\t(args.commits_only ? \" commits-only\" : \"\"),\n+\t\t\t\t(args.exact_objects ? \" exact-objects\" : \"\"),\n+\t\t\t\t\" ofs-delta\");\n+\telse\n+\t\tpacket_write(fd[1], \"want %s\\n\", remote);\n+}\n+\n+static void get_exact_objects(int fd[2], int nr_match, char **match)\n+{\n+\tint i;\n+\n+\t/* send all the objects as we got them on the command line */\n+\tfor (i = 0; i < nr_match; i++)\n+\t\tsend_want(fd, match[i], !i);\n+\n+\tpacket_flush(fd[1]);\n+}\n+\n static int find_common(int fd[2], unsigned char *result_sha1,\n \t\t       struct ref *refs)\n {\n@@ -172,17 +200,7 @@ static int find_common(int fd[2], unsigned char *result_sha1,\n \t\t\tcontinue;\n \t\t}\n \n-\t\tif (!fetching)\n-\t\t\tpacket_write(fd[1], \"want %s%s%s%s%s%s%s\\n\",\n-\t\t\t\t     sha1_to_hex(remote),\n-\t\t\t\t     (multi_ack ? \" multi_ack\" : \"\"),\n-\t\t\t\t     (use_sideband == 2 ? \" side-band-64k\" : \"\"),\n-\t\t\t\t     (use_sideband == 1 ? \" side-band\" : \"\"),\n-\t\t\t\t     (args.use_thin_pack ? \" thin-pack\" : \"\"),\n-\t\t\t\t     (args.no_progress ? \" no-progress\" : \"\"),\n-\t\t\t\t     \" ofs-delta\");\n-\t\telse\n-\t\t\tpacket_write(fd[1], \"want %s\\n\", sha1_to_hex(remote));\n+\t\tsend_want(fd, sha1_to_hex(remote), !fetching);\n \t\tfetching++;\n \t}\n \tif (is_repository_shallow())\n@@ -523,11 +541,15 @@ static int get_pack(int xd[2], char **pack_lockfile)\n \t\t\t\tstrcpy(keep_arg + s, \"localhost\");\n \t\t\t*av++ = keep_arg;\n \t\t}\n+\t\tif (args.exact_objects)\n+\t\t\t*av++ = \"--ignore-remote-alternates\";\n \t}\n \telse {\n \t\t*av++ = \"unpack-objects\";\n \t\tif (args.quiet)\n \t\t\t*av++ = \"-q\";\n+\t\tif (args.exact_objects)\n+\t\t\t*av++ = \"--ignore-remote-alternates\";\n \t}\n \tif (*hdr_arg)\n \t\t*av++ = hdr_arg;\n@@ -556,6 +578,7 @@ static struct ref *do_fetch_pack(int fd[2],\n \tunsigned char sha1[20];\n \n \tget_remote_heads(fd[0], &ref, 0, NULL, 0);\n+\n \tif (is_repository_shallow() && !server_supports(\"shallow\"))\n \t\tdie(\"Server does not support shallow clients\");\n \tif (server_supports(\"multi_ack\")) {\n@@ -573,20 +596,36 @@ static struct ref *do_fetch_pack(int fd[2],\n \t\t\tfprintf(stderr, \"Server supports side-band\\n\");\n \t\tuse_sideband = 1;\n \t}\n-\tif (!ref) {\n-\t\tpacket_flush(fd[1]);\n-\t\tdie(\"no matching remote head\");\n+\tif (!server_supports(\"remote-alternates\") &&\n+\t\t\t(args.commits_only || args.exact_objects)) {\n+\t\tif (args.verbose)\n+\t\t\tfprintf(stderr, \"Server does not support remote \"\n+\t\t\t\t\t\"alternates, ignoring %s%s\\n\",\n+\t\t\t\t\t(args.commits_only?\n+\t\t\t\t\t\t\"--commits-only \": \"\"),\n+\t\t\t\t\t(args.exact_objects? \"--exact-objects\": \"\"));\n+\t\targs.commits_only = 0;\n+\t\targs.exact_objects = 0;\n \t}\n-\tif (everything_local(&ref, nr_match, match)) {\n-\t\tpacket_flush(fd[1]);\n-\t\tgoto all_done;\n+\n+\tif (args.exact_objects)\n+\t\tget_exact_objects(fd, nr_match, match);\n+\telse {\n+\t\tif (!ref) {\n+\t\t\tpacket_flush(fd[1]);\n+\t\t\tdie(\"no matching remote head\");\n+\t\t}\n+\t\tif (everything_local(&ref, nr_match, match)) {\n+\t\t\tpacket_flush(fd[1]);\n+\t\t\tgoto all_done;\n+\t\t}\n+\t\tif (find_common(fd, sha1, ref) < 0)\n+\t\t\tif (!args.keep_pack)\n+\t\t\t\t/* When cloning, it is not unusual to have\n+\t\t\t\t * no common commit.\n+\t\t\t\t */\n+\t\t\t\tfprintf(stderr, \"warning: no common commits\\n\");\n \t}\n-\tif (find_common(fd, sha1, ref) < 0)\n-\t\tif (!args.keep_pack)\n-\t\t\t/* When cloning, it is not unusual to have\n-\t\t\t * no common commit.\n-\t\t\t */\n-\t\t\tfprintf(stderr, \"warning: no common commits\\n\");\n \n \tif (get_pack(fd, pack_lockfile))\n \t\tdie(\"git-fetch-pack: fetch failed.\");\n@@ -647,12 +686,72 @@ static void fetch_pack_setup(void)\n \tdid_setup = 1;\n }\n \n+static void read_from_stdin(int *num, char ***records)\n+{\n+\tchar buffer[4096];\n+\tsize_t records_num, leftover;\n+\tssize_t ret;\n+\n+\t*num = 0;\n+\tleftover = 0;\n+\n+\trecords_num = 4096;\n+\t(*records) = xmalloc(records_num * sizeof(char *));\n+\n+\tdo {\n+\t\tchar *p, *last;\n+\n+\t\tret = xread(0 /*stdin*/, buffer + leftover,\n+\t\t\t\tsizeof(buffer) - leftover);\n+\t\tif (ret < 0)\n+\t\t\tdie(\"read error on input: %s\", strerror(errno));\n+\n+\t\tlast = buffer;\n+\t\tfor (p = buffer; p < buffer + leftover + ret; p++)\n+\t\t\tif ((!*p || *p == '\\n') && (p != last)) {\n+\t\t\t\tif (*num >= records_num) {\n+\t\t\t\t\trecords_num *= 2;\n+\t\t\t\t\t(*records) = xrealloc(*records,\n+\t\t\t\t\t\t\t      records_num * sizeof(char*));\n+\t\t\t\t}\n+\n+\t\t\t\tif (p - last > 0) {\n+\t\t\t\t\t(*records)[*num] =\n+\t\t\t\t\t\tstrndup(last, p - last);\n+\t\t\t\t\t(*num)++;\n+\t\t\t\t}\n+\t\t\t\tlast = p + 1;\n+\t\t\t}\n+\n+\t\tleftover = p - last;\n+\t\tif (leftover >= sizeof(buffer))\n+\t\t\tdie(\"input line too long\");\n+\t\tif (leftover < 0)\n+\t\t\tleftover = 0;\n+\n+\t\tmemmove(buffer, last, leftover);\n+\t} while (ret > 0);\n+\n+\tif (leftover) {\n+\t\tif (*num >= records_num) {\n+\t\t\trecords_num *= 2;\n+\t\t\t(*records) = xrealloc(*records,\n+\t\t\t\t\t      records_num * sizeof(char*));\n+\t\t}\n+\n+\t\t(*records)[*num] = strndup(buffer, leftover);\n+\t\t(*num)++;\n+\t}\n+}\n+\n int cmd_fetch_pack(int argc, const char **argv, const char *prefix)\n {\n \tint i, ret, nr_heads;\n \tstruct ref *ref;\n \tchar *dest = NULL, **heads;\n+\tint from_stdin;\n \n+\tfrom_stdin = 0;\n \tnr_heads = 0;\n \theads = NULL;\n \tfor (i = 1; i < argc; i++) {\n@@ -696,6 +795,19 @@ int cmd_fetch_pack(int argc, const char **argv, const char *prefix)\n \t\t\t\targs.no_progress = 1;\n \t\t\t\tcontinue;\n \t\t\t}\n+\t\t\tif (!strcmp(\"--commits-only\", arg)) {\n+\t\t\t\targs.commits_only = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(\"--exact-objects\", arg)) {\n+\t\t\t\targs.exact_objects = 1;\n+\t\t\t\tdisable_remote_alternates();\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(\"--stdin\", arg)) {\n+\t\t\t\tfrom_stdin = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tusage(fetch_pack_usage);\n \t\t}\n \t\tdest = (char *)arg;\n@@ -706,14 +818,18 @@ int cmd_fetch_pack(int argc, const char **argv, const char *prefix)\n \tif (!dest)\n \t\tusage(fetch_pack_usage);\n \n+\tif (from_stdin)\n+\t\tread_from_stdin(&nr_heads, &heads);\n+\n \tref = fetch_pack(&args, dest, nr_heads, heads, NULL);\n \tret = !ref;\n \n-\twhile (ref) {\n-\t\tprintf(\"%s %s\\n\",\n-\t\t       sha1_to_hex(ref->old_sha1), ref->name);\n-\t\tref = ref->next;\n-\t}\n+\tif (!args.exact_objects)\n+\t\twhile (ref) {\n+\t\t\tprintf(\"%s %s\\n\",\n+\t\t\t\t\tsha1_to_hex(ref->old_sha1), ref->name);\n+\t\t\tref = ref->next;\n+\t\t}\n \n \treturn ret;\n }\n@@ -746,7 +862,7 @@ struct ref *fetch_pack(struct fetch_pack_args *my_args,\n \tclose(fd[1]);\n \tret = finish_connect(conn);\n \n-\tif (!ret && nr_heads) {\n+\tif (!ret && nr_heads && !args.exact_objects) {\n \t\t/* If the heads to pull were given, we should have\n \t\t * consumed all of them by matching the remote.\n \t\t * Otherwise, 'git-fetch remote no-such-ref' would\ndiff --git a/builtin-fetch.c b/builtin-fetch.c\nindex 320e235..858384a 100644\n--- a/builtin-fetch.c\n+++ b/builtin-fetch.c\n@@ -22,7 +22,7 @@ enum {\n \tTAGS_SET = 2\n };\n \n-static int append, force, keep, update_head_ok, verbose, quiet;\n+static int append, force, keep, update_head_ok, verbose, quiet, commits_only;\n static int tags = TAGS_DEFAULT;\n static const char *depth;\n static const char *upload_pack;\n@@ -45,6 +45,8 @@ static struct option builtin_fetch_options[] = {\n \t\t    \"allow updating of HEAD ref\"),\n \tOPT_STRING(0, \"depth\", &depth, \"DEPTH\",\n \t\t   \"deepen history of shallow clone\"),\n+\tOPT_BOOLEAN(0, \"commits-only\", &commits_only,\n+\t\t    \"fetch just the commit objects, leave the tree, blob, and tag objects for later\"),\n \tOPT_END()\n };\n \n@@ -602,6 +604,8 @@ int cmd_fetch(int argc, const char **argv, const char *prefix)\n \t\tset_option(TRANS_OPT_KEEP, \"yes\");\n \tif (depth)\n \t\tset_option(TRANS_OPT_DEPTH, depth);\n+\tif (commits_only)\n+\t\tset_option(TRANS_OPT_COMMITS_ONLY, \"yes\");\n \n \tif (!transport->url)\n \t\tdie(\"Where do you want to fetch from today?\");\ndiff --git a/builtin-unpack-objects.c b/builtin-unpack-objects.c\nindex 1e51865..58d8a41 100644\n--- a/builtin-unpack-objects.c\n+++ b/builtin-unpack-objects.c\n@@ -10,7 +10,7 @@\n #include \"progress.h\"\n \n static int dry_run, quiet, recover, has_errors;\n-static const char unpack_usage[] = \"git-unpack-objects [-n] [-q] [-r] < pack-file\";\n+static const char unpack_usage[] = \"git-unpack-objects [-n] [-q] [-r] [--ignore-remote-alternates] < pack-file\";\n \n /* We always read in 4kB chunks. */\n static unsigned char buffer[4096];\n@@ -359,6 +359,10 @@ int cmd_unpack_objects(int argc, const char **argv, const char *prefix)\n \t\t\t\trecover = 1;\n \t\t\t\tcontinue;\n \t\t\t}\n+\t\t\tif (!strcmp(arg, \"--ignore-remote-alternates\")) {\n+\t\t\t\tdisable_remote_alternates();\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tif (!prefixcmp(arg, \"--pack_header=\")) {\n \t\t\t\tstruct pack_header *hdr;\n \t\t\t\tchar *c;\ndiff --git a/cache.h b/cache.h\nindex 549f4bb..def7459 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -480,6 +480,8 @@ extern struct alternate_object_database {\n } *alt_odb_list;\n extern void prepare_alt_odb(void);\n \n+extern void disable_remote_alternates(void);\n+\n struct pack_window {\n \tstruct pack_window *next;\n \tunsigned char *base;\ndiff --git a/fetch-pack.h b/fetch-pack.h\nindex a7888ea..0c3b13f 100644\n--- a/fetch-pack.h\n+++ b/fetch-pack.h\n@@ -12,7 +12,9 @@ struct fetch_pack_args\n \t\tuse_thin_pack:1,\n \t\tfetch_all:1,\n \t\tverbose:1,\n-\t\tno_progress:1;\n+\t\tno_progress:1,\n+\t\tcommits_only:1,\n+\t\texact_objects:1;\n };\n \n struct ref *fetch_pack(struct fetch_pack_args *args,\ndiff --git a/git-clone.sh b/git-clone.sh\nindex b4e858c..208e9fc 100755\n--- a/git-clone.sh\n+++ b/git-clone.sh\n@@ -115,7 +115,7 @@ Perhaps git-update-server-info needs to be run there?\"\n quiet=\n local=no\n use_local_hardlink=yes\n-local_shared=no\n+shared=no\n unset template\n no_checkout=\n upload_pack=\n@@ -143,7 +143,7 @@ do\n \t--no-hardlinks)\n \t\tuse_local_hardlink=no ;;\n \t-s|--shared)\n-\t\tlocal_shared=yes ;;\n+\t\tshared=yes ;;\n \t--template)\n \t\tshift; template=\"--template=$1\" ;;\n \t-q|--quiet)\n@@ -288,7 +288,7 @@ yes)\n \t( cd \"$repo/objects\" ) ||\n \t\tdie \"cannot chdir to local '$repo/objects'.\"\n \n-\tif test \"$local_shared\" = yes\n+\tif test \"$shared\" = yes\n \tthen\n \t\tmkdir -p \"$GIT_DIR/objects/info\"\n \t\techo \"$repo/objects\" >>\"$GIT_DIR/objects/info/alternates\"\n@@ -364,11 +364,22 @@ yes)\n \t\tfi\n \t\t;;\n \t*)\n+\t\tcommits_only=\n+\t\tif test \"$shared\" = yes\n+\t\tthen\n+\t\t\tcommits_only=\"--commits-only\"\n+\t\tfi\n \t\tcase \"$upload_pack\" in\n-\t\t'') git-fetch-pack --all -k $quiet $depth $no_progress \"$repo\";;\n-\t\t*) git-fetch-pack --all -k $quiet \"$upload_pack\" $depth $no_progress \"$repo\" ;;\n+\t\t'') git-fetch-pack --all -k $quiet $depth $no_progress $commits_only \"$repo\";;\n+\t\t*) git-fetch-pack --all -k $quiet \"$upload_pack\" $depth $no_progress $commits_only \"$repo\" ;;\n \t\tesac >\"$GIT_DIR/CLONE_HEAD\" ||\n \t\t\tdie \"fetch-pack from '$repo' failed.\"\n+\t\tif test \"$shared\" = yes\n+\t\tthen\n+\t\t\t# Must be done after the fetch\n+\t\t\tmkdir -p \"$GIT_DIR/objects/info\"\n+\t\t\techo \"$repo\" >> \"$GIT_DIR/objects/info/remote_alternates\"\n+\t\tfi\n \t\t;;\n \tesac\n \t;;\ndiff --git a/index-pack.c b/index-pack.c\nindex 9fd6982..f2e6b7a 100644\n--- a/index-pack.c\n+++ b/index-pack.c\n@@ -9,7 +9,7 @@\n #include \"progress.h\"\n \n static const char index_pack_usage[] =\n-\"git-index-pack [-v] [-o <index-file>] [{ ---keep | --keep=<msg> }] { <pack-file> | --stdin [--fix-thin] [<pack-file>] }\";\n+\"git-index-pack [-v] [-o <index-file>] [{ ---keep | --keep=<msg> }] [--ignore-remote-alternates] { <pack-file> | --stdin [--fix-thin] [<pack-file>] }\";\n \n struct object_entry\n {\n@@ -746,6 +746,8 @@ int main(int argc, char **argv)\n \t\t\t\t\tpack_idx_off32_limit = strtoul(c+1, &c, 0);\n \t\t\t\tif (*c || pack_idx_off32_limit & 0x80000000)\n \t\t\t\t\tdie(\"bad %s\", arg);\n+\t\t\t} else if (!strcmp(arg, \"--ignore-remote-alternates\")) {\n+\t\t\t\tdisable_remote_alternates();\n \t\t\t} else\n \t\t\t\tusage(index_pack_usage);\n \t\t\tcontinue;\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 66a4e00..7d60be0 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -14,6 +14,7 @@\n #include \"tag.h\"\n #include \"tree.h\"\n #include \"refs.h\"\n+#include \"run-command.h\"\n \n #ifndef O_NOATIME\n #if defined(__linux__) && (defined(__i386__) || defined(__PPC__))\n@@ -411,6 +412,205 @@ static char *find_sha1_file(const unsigned char *sha1, struct stat *st)\n \treturn NULL;\n }\n \n+static char *remote_alternates = NULL;\n+static int has_remote_alt_feature = -1;\n+\n+void disable_remote_alternates(void)\n+{\n+\thas_remote_alt_feature = 0;\n+}\n+\n+static int has_remote_alternates(void)\n+{\n+\t/* FIXME: does it make sense to support more URLs inside\n+\t * remote_alternates? */\n+\tstruct stat st;\n+\tconst char remote_alt_file_name[] = \"info/remote_alternates\";\n+\tchar path[PATH_MAX + 1 + sizeof remote_alt_file_name];\n+\tint fd;\n+\tchar *map, *p;\n+\tsize_t mapsz;\n+\n+\tif (has_remote_alt_feature != -1)\n+\t\treturn has_remote_alt_feature;\n+\n+\thas_remote_alt_feature = 0;\n+\n+\tsprintf(path, \"%s/%s\", get_object_directory(),\n+\t\t\tremote_alt_file_name);\n+\tfd = open(path, O_RDONLY);\n+\tif (fd < 0)\n+\t\treturn has_remote_alt_feature;\n+\telse if (fstat(fd, &st) || (st.st_size == 0)) {\n+\t\tclose(fd);\n+\t\treturn has_remote_alt_feature;\n+\t}\n+\n+\tmapsz = xsize_t(st.st_size);\n+\tmap = xmmap(NULL, mapsz, PROT_READ, MAP_PRIVATE, fd, 0);\n+\tclose(fd);\n+\n+\t/* we support just one remote alternate for now,\n+\t * so read just the first entry */\n+\tfor (p = map; (p < map + mapsz) && (*p != '\\n'); p++)\n+\t\t;\n+\n+\tremote_alternates = strndup(map, p - map);\n+\n+\tmunmap(map, mapsz);\n+\n+\tif (remote_alternates && remote_alternates[0])\n+\t\thas_remote_alt_feature = 1;\n+\n+\treturn has_remote_alt_feature;\n+}\n+\n+struct sha1_list {\n+\tunsigned char sha1[20];\n+\tstruct sha1_list *next;\n+};\n+\n+static int has_sha1_file_locally(const unsigned char *sha1);\n+\n+static int async_dump_objects(int fd, void *data)\n+{\n+\tFILE *out = NULL;\n+\tstruct sha1_list *list;\n+\n+\tout = fdopen(fd, \"w\");\n+\n+\tlist = (struct sha1_list *)data;\n+\twhile (list) {\n+\t\tif (!has_sha1_file_locally(list->sha1))\n+\t\t\tfprintf(out, \"%s\\n\", sha1_to_hex(list->sha1));\n+\n+\t\tlist = list->next;\n+\t}\n+\n+\tfflush(out);\n+\treturn 0;\n+}\n+\n+static int fetch_remote_sha1s(struct sha1_list *objects)\n+{\n+\tstruct async dump_objects;\n+\tstruct child_process fetch_pack;\n+\tconst char *argv[20];\n+\tint argc = 0;\n+\tint err;\n+\n+\tif (!objects)\n+\t\treturn 0;\n+\n+\t/* this will fill the stdin of fetch-pack */\n+\tdump_objects.proc = async_dump_objects;\n+\tdump_objects.data = objects;\n+\n+\tif (start_async(&dump_objects))\n+\t\tdie(\"unable to send data to fetch-pack\");\n+\n+\targv[argc++] = \"fetch-pack\";\n+\targv[argc++] = \"--stdin\";\n+\targv[argc++] = \"--exact-objects\";\n+\targv[argc++] = remote_alternates;\n+\targv[argc] = NULL;\n+\n+\tmemset(&fetch_pack, 0, sizeof(fetch_pack));\n+\tfetch_pack.in = dump_objects.out;\n+\tfetch_pack.out = 1;\n+\tfetch_pack.err = 2;\n+\tfetch_pack.git_cmd = 1;\n+\tfetch_pack.argv = argv;\n+\n+\terr = run_command(&fetch_pack);\n+\n+\t/* TODO better error handling - is the object really missing, or\n+\t * was it just a temporary network error? */\n+\tif (err) {\n+\t\tfprintf(stderr, \"error %d while calling fetch-pack\\n\", err);\n+\t\treturn 0;\n+\t}\n+\n+\treturn 1;\n+}\n+\n+static struct sha1_list *remote_list = NULL;\n+\n+static int fill_remote_list(const unsigned char *sha1,\n+\t\tconst char *base, int baselen,\n+\t\tconst char *pathname, unsigned mode, int stage)\n+{\n+\tif (!has_sha1_file_locally(sha1)) {\n+\t\tstruct sha1_list *item;\n+\n+\t\titem = xmalloc(sizeof(*item));\n+\t\thashcpy(item->sha1, sha1);\n+\t\titem->next = remote_list;\n+\n+\t\tremote_list = item;\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int fetch_remote_sha1s_recursive(struct sha1_list *objects)\n+{\n+\tstruct sha1_list *list;\n+\tint ret = 0;\n+\n+\t/* first of all, fetch the missing objects */\n+\tif (!fetch_remote_sha1s(objects))\n+\t\treturn 0;\n+\n+\tremote_list = NULL;\n+\n+\tlist = objects;\n+\twhile (list) {\n+\t\tstruct tree *tree;\n+\n+\t\ttree = parse_tree_indirect(list->sha1);\n+\t\tif (tree) {\n+\t\t\tread_tree_recursive(tree, \"\", 0, 0, NULL,\n+\t\t\t\t\tfill_remote_list);\n+\t\t}\n+\n+\t\tlist = list->next;\n+\t}\n+\n+\tlist = remote_list;\n+\tif (!list)\n+\t\treturn 1; /* hooray, we have everything */\n+\n+\tret = fetch_remote_sha1s_recursive(list);\n+\n+\twhile (list) {\n+\t\tstruct sha1_list *item;\n+\n+\t\titem = list;\n+\t\tlist = list->next;\n+\n+\t\tfree(item);\n+\t}\n+\n+\treturn ret;\n+}\n+\n+static int download_remote_sha1(const unsigned char *sha1)\n+{\n+\tstruct sha1_list item;\n+\tint ret;\n+\n+\tif (!has_remote_alternates())\n+\t\treturn 0;\n+\n+\thashcpy(item.sha1, sha1);\n+\titem.next = NULL;\n+\n+\tret = fetch_remote_sha1s_recursive(&item);\n+\n+\treturn ret;\n+}\n+\n static unsigned int pack_used_ctr;\n static unsigned int pack_mmap_calls;\n static unsigned int peak_pack_open_windows;\n@@ -1880,7 +2080,7 @@ int pretend_sha1_file(void *buf, unsigned long len, enum object_type type,\n \treturn 0;\n }\n \n-void *read_sha1_file(const unsigned char *sha1, enum object_type *type,\n+static void *read_sha1_file_locally(const unsigned char *sha1, enum object_type *type,\n \t\t     unsigned long *size)\n {\n \tunsigned long mapsize;\n@@ -1897,6 +2097,7 @@ void *read_sha1_file(const unsigned char *sha1, enum object_type *type,\n \tbuf = read_packed_sha1(sha1, type, size);\n \tif (buf)\n \t\treturn buf;\n+\n \tmap = map_sha1_file(sha1, &mapsize);\n \tif (map) {\n \t\tbuf = unpack_sha1_file(map, mapsize, type, size, sha1);\n@@ -1907,6 +2108,21 @@ void *read_sha1_file(const unsigned char *sha1, enum object_type *type,\n \treturn read_packed_sha1(sha1, type, size);\n }\n \n+void *read_sha1_file(const unsigned char *sha1, enum object_type *type,\n+\t\t     unsigned long *size)\n+{\n+\tvoid *result;\n+\n+\tresult = read_sha1_file_locally(sha1, type, size);\n+\n+\t/* if it's remote, and we don't have it yet, dowload it now and try\n+\t * again */\n+\tif (!result && has_remote_alternates() && download_remote_sha1(sha1))\n+\t\tresult = read_sha1_file_locally(sha1, type, size);\n+\n+\treturn result;\n+}\n+\n void *read_object_with_reference(const unsigned char *sha1,\n \t\t\t\t const char *required_type_name,\n \t\t\t\t unsigned long *size,\n@@ -2306,7 +2522,7 @@ int has_sha1_pack(const unsigned char *sha1, const char **ignore_packed)\n \treturn find_pack_entry(sha1, &e, ignore_packed);\n }\n \n-int has_sha1_file(const unsigned char *sha1)\n+static int has_sha1_file_locally(const unsigned char *sha1)\n {\n \tstruct stat st;\n \tstruct pack_entry e;\n@@ -2316,6 +2532,18 @@ int has_sha1_file(const unsigned char *sha1)\n \treturn find_sha1_file(sha1, &st) ? 1 : 0;\n }\n \n+int has_sha1_file(const unsigned char *sha1)\n+{\n+\tif (has_sha1_file_locally(sha1))\n+\t\treturn 1;\n+\n+\t/* download it if necessary */\n+\tif (has_remote_alternates() && download_remote_sha1(sha1))\n+\t\treturn has_sha1_file_locally(sha1);\n+\n+\treturn 0;\n+}\n+\n int index_pipe(unsigned char *sha1, int fd, const char *type, int write_object)\n {\n \tstruct strbuf buf;\ndiff --git a/transport.c b/transport.c\nindex babaa21..918c390 100644\n--- a/transport.c\n+++ b/transport.c\n@@ -562,6 +562,7 @@ static int close_bundle(struct transport *transport)\n struct git_transport_data {\n \tunsigned thin : 1;\n \tunsigned keep : 1;\n+\tunsigned commits_only : 1;\n \tint depth;\n \tconst char *uploadpack;\n \tconst char *receivepack;\n@@ -589,6 +590,9 @@ static int set_git_option(struct transport *connection,\n \t\telse\n \t\t\tdata->depth = atoi(value);\n \t\treturn 0;\n+\t} else if (!strcmp(name, TRANS_OPT_COMMITS_ONLY)) {\n+\t\tdata->commits_only = !!value;\n+\t\treturn 0;\n \t}\n \treturn 1;\n }\n@@ -629,6 +633,7 @@ static int fetch_refs_via_pack(struct transport *transport,\n \targs.use_thin_pack = data->thin;\n \targs.verbose = transport->verbose > 0;\n \targs.depth = data->depth;\n+\targs.commits_only = data->commits_only;\n \n \tfor (i = 0; i < nr_heads; i++)\n \t\torigh[i] = heads[i] = xstrdup(to_fetch[i]->name);\ndiff --git a/transport.h b/transport.h\nindex 6fb4526..4076186 100644\n--- a/transport.h\n+++ b/transport.h\n@@ -53,6 +53,10 @@ struct transport *transport_get(struct remote *, const char *);\n /* Limit the depth of the fetch if not null */\n #define TRANS_OPT_DEPTH \"depth\"\n \n+/* Download only the commit objects, let the tree, blob and tag objects for\n+ * later */\n+#define TRANS_OPT_COMMITS_ONLY \"commits-only\"\n+\n /**\n  * Returns 0 if the option was used, non-zero otherwise. Prints a\n  * message to stderr if the option is not used.\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 7e04311..2d047ec 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -27,7 +27,7 @@ static const char upload_pack_usage[] = \"git-upload-pack [--strict] [--timeout=n\n static unsigned long oldest_have;\n \n static int multi_ack, nr_our_refs;\n-static int use_thin_pack, use_ofs_delta, no_progress;\n+static int use_thin_pack, use_ofs_delta, no_progress, commits_only, exact_objects;\n static struct object_array have_obj;\n static struct object_array want_obj;\n static unsigned int timeout;\n@@ -106,9 +106,15 @@ static int do_rev_list(int fd, void *create_full_pack)\n \tif (create_full_pack)\n \t\tuse_thin_pack = 0; /* no point doing it */\n \tinit_revisions(&revs, NULL);\n-\trevs.tag_objects = 1;\n-\trevs.tree_objects = 1;\n-\trevs.blob_objects = 1;\n+\tif (!commits_only) {\n+\t\trevs.tag_objects = 1;\n+\t\trevs.tree_objects = 1;\n+\t\trevs.blob_objects = 1;\n+\t} else {\n+\t\trevs.tag_objects = 0;\n+\t\trevs.tree_objects = 0;\n+\t\trevs.blob_objects = 0;\n+\t}\n \tif (use_thin_pack)\n \t\trevs.edge_hint = 1;\n \n@@ -135,6 +141,20 @@ static int do_rev_list(int fd, void *create_full_pack)\n \treturn 0;\n }\n \n+static int dump_want_objects(int fd, void *data)\n+{\n+\tint i;\n+\tpack_pipe = fdopen(fd, \"w\");\n+\n+\tfor (i = 0; i < want_obj.nr; i++) {\n+\t\tstruct object *o = want_obj.objects[i].item;\n+\t\tfprintf(pack_pipe, \"%s\\n\", sha1_to_hex(o->sha1));\n+\t}\n+\n+\tfflush(pack_pipe);\n+\treturn 0;\n+}\n+\n static void create_pack_file(void)\n {\n \tstruct async rev_list;\n@@ -148,7 +168,10 @@ static void create_pack_file(void)\n \tconst char *argv[10];\n \tint arg = 0;\n \n-\trev_list.proc = do_rev_list;\n+\tif (!exact_objects)\n+\t\trev_list.proc = do_rev_list;\n+\telse\n+\t\trev_list.proc = dump_want_objects;\n \t/* .data is just a boolean: any non-NULL value will do */\n \trev_list.data = create_full_pack ? &rev_list : NULL;\n \tif (start_async(&rev_list))\n@@ -489,6 +512,10 @@ static void receive_needs(void)\n \t\t\tuse_sideband = DEFAULT_PACKET_MAX;\n \t\tif (strstr(line+45, \"no-progress\"))\n \t\t\tno_progress = 1;\n+\t\tif (strstr(line+45, \"commits-only\"))\n+\t\t\tcommits_only = 1;\n+\t\tif (strstr(line+45, \"exact-objects\"))\n+\t\t\texact_objects = 1;\n \n \t\t/* We have sent all our refs already, and the other end\n \t\t * should have chosen out of them; otherwise they are\n@@ -498,9 +525,15 @@ static void receive_needs(void)\n \t\t * asks for something like \"master~10\" (symbolic)...\n \t\t * would it make sense?  I don't know.\n \t\t */\n-\t\to = lookup_object(sha1_buf);\n-\t\tif (!o || !(o->flags & OUR_REF))\n-\t\t\tdie(\"git-upload-pack: not our ref %s\", line+5);\n+\t\tif (!exact_objects) {\n+\t\t\to = lookup_object(sha1_buf);\n+\t\t\tif (!o || !(o->flags & OUR_REF))\n+\t\t\t\tdie(\"git-upload-pack: not our ref %s\", line+5);\n+\t\t} else {\n+\t\t\to = lookup_unknown_object(sha1_buf);\n+\t\t\tif (!o)\n+\t\t\t\tdie(\"git-upload-pack: not an object %s\", line+5);\n+\t\t}\n \t\tif (!(o->flags & WANTED)) {\n \t\t\to->flags |= WANTED;\n \t\t\tadd_object_array(o, NULL, &want_obj);\n@@ -557,7 +590,7 @@ static void receive_needs(void)\n static int send_ref(const char *refname, const unsigned char *sha1, int flag, void *cb_data)\n {\n \tstatic const char *capabilities = \"multi_ack thin-pack side-band\"\n-\t\t\" side-band-64k ofs-delta shallow no-progress\";\n+\t\t\" side-band-64k ofs-delta shallow no-progress remote-alternates\";\n \tstruct object *o = parse_object(sha1);\n \n \tif (!o)\n@@ -588,7 +621,8 @@ static void upload_pack(void)\n \tpacket_flush(1);\n \treceive_needs();\n \tif (want_obj.nr) {\n-\t\tget_common_commits();\n+\t\tif (!exact_objects)\n+\t\t\tget_common_commits();\n \t\tcreate_pack_file();\n \t}\n }\n"},{"id":"67976","messageId":"alpine.LFD.1.00.0802081250240.2732@xanadu.home","threadId":"11972","inReplyTo":"200802081828.43849.kendy@suse.cz","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-08T18:03:49Z","receivedAt":"2008-02-08T18:03:49Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 8 Feb 2008, Jan Holesovsky wrote:\n\n> Currently we are evaluating the usage of git for OpenOffice.org as one of the\n> candidates (SVN is the other one), see\n> \n>   http://wiki.services.openoffice.org/wiki/SCM_Migration\n> \n> I've provided a git import of OOo with the entire history; the problem is that\n> the pack has 2.5G, so it's not too convenient to download for casual\n> developers that just want to try it.  Shallow clone is not a possibility - we\n> don't get patches through mailing lists, so we need the pull/push, and also\n> thanks to the OOo development cycle, we have too many living heads which\n> causes the shallow clone to download about 1.5G even with --depth 1.\n\nHow did you repack your repository?\n\nWe know that current defaults are not suitable for large projects.  For \nexample, the gcc git repository shrinked from 1.5GB pack down to 230MB \nafter some tuning.\n\n\nNicolas\n"},{"id":"67977","messageId":"1202494475.31361.45.camel@brick","threadId":"11972","inReplyTo":"200802081828.43849.kendy@suse.cz","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2008-02-08T18:14:35Z","receivedAt":"2008-02-08T18:14:35Z","isPatch":true,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"On Fri, 2008-02-08 at 18:28 +0100, Jan Holesovsky wrote:\n> Hi,\n> \n> This is my attempt to implement the 'lazy clone' I've read about a bit in the\n> git mailing list archive, but did not see implemented anywhere - the clone\n> that fetches a minimal amount of data with the possibility to download the\n> rest later (transparently!) when necessary.  I am sorry to send it as a huge\n> patch, not as a series of patches, but as I don't know if I chose a way that is\n> acceptable for you [I'm new to the git code ;-)], I'd like to hear some\n> feedback first, and then I'll split it into smaller pieces for easier\n> integration - if OK.\n> \n> Background:\n> \n> Currently we are evaluating the usage of git for OpenOffice.org as one of the\n> candidates (SVN is the other one), see\n> \n>   http://wiki.services.openoffice.org/wiki/SCM_Migration\n> \n> I've provided a git import of OOo with the entire history; the problem is that\n> the pack has 2.5G, so it's not too convenient to download for casual\n> developers that just want to try it.  Shallow clone is not a possibility - we\n> don't get patches through mailing lists, so we need the pull/push, and also\n> thanks to the OOo development cycle, we have too many living heads which\n> causes the shallow clone to download about 1.5G even with --depth 1.  Lazy\n> clone sounded like the right idea to me.  With this proof-of-concept\n> implementation, just about 550M from the 2.5G is downloaded, which is still\n> about twice as much in comparison with downloading a tarball, but bearable.\n\nFor comparison, how big was the svn repo you're testing?  My experience\nhas been about 15-20 times smaller than SVN once a tuned repack has\nbeen done.\n\nCheers,\n\nHarvey\n"},{"id":"67978","messageId":"alpine.LSU.1.00.0802081809270.11591@racer.site","threadId":"11972","inReplyTo":"200802081828.43849.kendy@suse.cz","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-08T18:20:32Z","receivedAt":"2008-02-08T18:20:32Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 8 Feb 2008, Jan Holesovsky wrote:\n\n> +static void send_want(int fd[2], const char *remote, int full_info)\n> +{\n> +\tif (full_info)\n> +\t\tpacket_write(fd[1], \"want %s%s%s%s%s%s%s%s%s\\n\",\n> +\t\t\t\tremote,\n> +\t\t\t\t(multi_ack ? \" multi_ack\" : \"\"),\n> +\t\t\t\t(use_sideband == 2 ? \" side-band-64k\" : \"\"),\n> +\t\t\t\t(use_sideband == 1 ? \" side-band\" : \"\"),\n> +\t\t\t\t(args.use_thin_pack ? \" thin-pack\" : \"\"),\n> +\t\t\t\t(args.no_progress ? \" no-progress\" : \"\"),\n> +\t\t\t\t(args.commits_only ? \" commits-only\" : \"\"),\n> +\t\t\t\t(args.exact_objects ? \" exact-objects\" : \"\"),\n> +\t\t\t\t\" ofs-delta\");\n> +\telse\n> +\t\tpacket_write(fd[1], \"want %s\\n\", remote);\n> +}\n\nYou might want to make the full_info static, and only send the options the \nfirst time.\n\n> @@ -523,11 +541,15 @@ static int get_pack(int xd[2], char **pack_lockfile)\n>  \t\t\t\tstrcpy(keep_arg + s, \"localhost\");\n>  \t\t\t*av++ = keep_arg;\n>  \t\t}\n> +\t\tif (args.exact_objects)\n> +\t\t\t*av++ = \"--ignore-remote-alternates\";\n>  \t}\n>  \telse {\n>  \t\t*av++ = \"unpack-objects\";\n>  \t\tif (args.quiet)\n>  \t\t\t*av++ = \"-q\";\n> +\t\tif (args.exact_objects)\n> +\t\t\t*av++ = \"--ignore-remote-alternates\";\n>  \t}\n\nYou can move this outside of the if() instead of repeating yourself...\n\n> @@ -556,6 +578,7 @@ static struct ref *do_fetch_pack(int fd[2],\n>  \tunsigned char sha1[20];\n>  \n>  \tget_remote_heads(fd[0], &ref, 0, NULL, 0);\n> +\n>  \tif (is_repository_shallow() && !server_supports(\"shallow\"))\n>  \t\tdie(\"Server does not support shallow clients\");\n>  \tif (server_supports(\"multi_ack\")) {\n\nNot strictly necessary, right? ;-)\n\n> @@ -647,12 +686,72 @@ static void fetch_pack_setup(void)\n>  \tdid_setup = 1;\n>  }\n>  \n> +static void read_from_stdin(int *num, char ***records)\n> +{\n> +\tchar buffer[4096];\n> +\tsize_t records_num, leftover;\n> +\tssize_t ret;\n> +\n> +\t*num = 0;\n> +\tleftover = 0;\n> +\n> +\trecords_num = 4096;\n> +\t(*records) = xmalloc(records_num * sizeof(char *));\n> +\n> +\tdo {\n> +\t\tchar *p, *last;\n> +\n> +\t\tret = xread(0 /*stdin*/, buffer + leftover,\n> +\t\t\t\tsizeof(buffer) - leftover);\n> +\t\tif (ret < 0)\n> +\t\t\tdie(\"read error on input: %s\", strerror(errno));\n> +\n> +\t\tlast = buffer;\n> +\t\tfor (p = buffer; p < buffer + leftover + ret; p++)\n> +\t\t\tif ((!*p || *p == '\\n') && (p != last)) {\n> +\t\t\t\tif (*num >= records_num) {\n> +\t\t\t\t\trecords_num *= 2;\n> +\t\t\t\t\t(*records) = xrealloc(*records,\n> +\t\t\t\t\t\t\t      records_num * sizeof(char*));\n> +\t\t\t\t}\n> +\n> +\t\t\t\tif (p - last > 0) {\n> +\t\t\t\t\t(*records)[*num] =\n> +\t\t\t\t\t\tstrndup(last, p - last);\n> +\t\t\t\t\t(*num)++;\n> +\t\t\t\t}\n> +\t\t\t\tlast = p + 1;\n> +\t\t\t}\n> +\n> +\t\tleftover = p - last;\n> +\t\tif (leftover >= sizeof(buffer))\n> +\t\t\tdie(\"input line too long\");\n> +\t\tif (leftover < 0)\n> +\t\t\tleftover = 0;\n> +\n> +\t\tmemmove(buffer, last, leftover);\n> +\t} while (ret > 0);\n> +\n> +\tif (leftover) {\n> +\t\tif (*num >= records_num) {\n> +\t\t\trecords_num *= 2;\n> +\t\t\t(*records) = xrealloc(*records,\n> +\t\t\t\t\t      records_num * sizeof(char*));\n> +\t\t}\n> +\n> +\t\t(*records)[*num] = strndup(buffer, leftover);\n> +\t\t(*num)++;\n> +\t}\n> +}\n> +\n\nThis chunk could use ALLOC_GROW() quite nicely (would make it more \nreadable, and avoid errors).  Also, I'd use alloc_nr() instead of the \ndoubling.\n\n>  int cmd_fetch_pack(int argc, const char **argv, const char *prefix)\n>  {\n>  \tint i, ret, nr_heads;\n>  \tstruct ref *ref;\n>  \tchar *dest = NULL, **heads;\n> +\tint from_stdin;\n>  \n> +\tfrom_stdin = 0;\n\nYou can initialise it to 0 right away...\n\nUnfortunately, I have to go now... so I will review the rest \n(from builtin-fetch.c on) later.\n\nIt's great seeing that you work on this!\n\nThanks,\nDscho\n"},{"id":"67980","messageId":"20080208184902.GA19404@glandium.org","threadId":"11972","inReplyTo":"200802081828.43849.kendy@suse.cz","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-02-08T18:49:02Z","receivedAt":"2008-02-08T18:49:02Z","isPatch":true,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Fri, Feb 08, 2008 at 06:28:43PM +0100, Jan Holesovsky wrote:\n> Hi,\n> \n> This is my attempt to implement the 'lazy clone' I've read about a bit in the\n> git mailing list archive, but did not see implemented anywhere - the clone\n> that fetches a minimal amount of data with the possibility to download the\n> rest later (transparently!) when necessary.  I am sorry to send it as a huge\n> patch, not as a series of patches, but as I don't know if I chose a way that is\n> acceptable for you [I'm new to the git code ;-)], I'd like to hear some\n> feedback first, and then I'll split it into smaller pieces for easier\n> integration - if OK.\n> \n> Background:\n> \n> Currently we are evaluating the usage of git for OpenOffice.org as one of the\n> candidates (SVN is the other one), see\n> \n>   http://wiki.services.openoffice.org/wiki/SCM_Migration\n> \n> I've provided a git import of OOo with the entire history; the problem is that\n> the pack has 2.5G, so it's not too convenient to download for casual\n> developers that just want to try it.  Shallow clone is not a possibility - we\n> don't get patches through mailing lists, so we need the pull/push, and also\n> thanks to the OOo development cycle, we have too many living heads which\n> causes the shallow clone to download about 1.5G even with --depth 1.  Lazy\n> clone sounded like the right idea to me.  With this proof-of-concept\n> implementation, just about 550M from the 2.5G is downloaded, which is still\n> about twice as much in comparison with downloading a tarball, but bearable.\n<snip>\n\nThere are 2 things, here:\n- Probably, you can make your pack smaller with proper window sizing.\nTry taking a look at the \"Git and GCC\" that crossed borders between\nthe gcc and the git mailing lists.\n- There are tricks to do roughly what you want without modifying git.\nFor example, you can prepare several \"shared\" clones of your repo (git\nclone -s) and leave in each only a few branches. Cloning from these will\nonly pull the needed data.\n\nMike\n"},{"id":"67981","messageId":"m3ejbngtnn.fsf@localhost.localdomain","threadId":"11972","inReplyTo":"200802081828.43849.kendy@suse.cz","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-08T19:00:55Z","receivedAt":"2008-02-08T19:00:55Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jan Holesovsky <kendy@suse.cz> writes:\n\n> This is my attempt to implement the 'lazy clone' I've read about a\n> bit in the git mailing list archive, but did not see implemented\n> anywhere - the clone that fetches a minimal amount of data with the\n> possibility to download the rest later (transparently!) when\n> necessary.\n\nIt was not implemented because it was thought to be hard; git assumes\nin many places that if it has an object, it has all objects referenced\nby it.\n\nBut it is very nice of you to [try to] implement 'lazy clone'/'remote\nalternates'.\n\nCould you provide some benchmarks (time, network throughtput, latency)\nfor your implementation?\n\n> Currently we are evaluating the usage of git for OpenOffice.org as\n> one of the candidates (SVN is the other one), see\n> \n>   http://wiki.services.openoffice.org/wiki/SCM_Migration\n> \n> I've provided a git import of OOo with the entire history; the\n> problem is that the pack has 2.5G, so it's not too convenient to\n> download for casual developers that just want to try it.\n\nOne of the reasons why 'lazy clone' was not implemented was the fact\nthat by using large enough window, and larger than default delta\nlength you can repack \"archive pack\" (and keep it from trying to\nrepack using .keep files, see git-config(1)) much tighter than with\ndefault (time and CPU conserving) options, and much, much tighter than\npack which is result of fast-import driven import.\n\nBoth Mozilla import, and GCC import were packed below 0.5 GB. Warning:\nyou would need machine with large amount of memory to repack it\ntightly in sensible time!\n\n> Shallow clone is not a possibility - we don't get patches through\n> mailing lists, so we need the pull/push, and also thanks to the OOo\n> development cycle, we have too many living heads which causes the\n> shallow clone to download about 1.5G even with --depth 1.\n\nWouldn't be easier to try to fix shallow clone implementation to allow\nfor pushing from shallow to full clone (fetching from full to shallow\nis implemented), and perhaps also push/pull between two shallow\nclones?\n\nAs to many living heads: first, you don't need to fetch all\nheads. Currently git-clone has no option to select subset of heads to\nclone, but you can always use git-init + hand configuration +\ngit-remote and git-fetch for actual fetching.\n\n\nBy the way, did you try to split OpenOffice.org repository at the\ncomponents boundary into submodules (subprojects)? This would also\nlimit amount of needed download, as you don't neeed to download and\ncheckout all subprojects. \n\nThe problem of course is _how_ to split repository into\nsubmodules. Submodules should be enough self contained so the\nwhole-tree commit is alsays (or almost always) only about submodule.\n\n> Lazy clone sounded like the right idea to me.  With this\n> proof-of-concept implementation, just about 550M from the 2.5G is\n> downloaded, which is still about twice as much in comparison with\n> downloading a tarball, but bearable.\n\nDo you have any numbers for OOo repository like number of revisions,\ndepth of DAG of commits (maximum number of revisions in one line of\ncommits), number of files, size of checkout, average size of file,\netc.?\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"67982","messageId":"alpine.LSU.1.00.0802081903510.11591@racer.site","threadId":"11972","inReplyTo":"20080208184902.GA19404@glandium.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-08T19:04:46Z","receivedAt":"2008-02-08T19:04:46Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 8 Feb 2008, Mike Hommey wrote:\n\n> - There are tricks to do roughly what you want without modifying git. \n> For example, you can prepare several \"shared\" clones of your repo (git \n> clone -s) and leave in each only a few branches. Cloning from these will \n> only pull the needed data.\n\nThe problem is, of course, that the shared clones are not updated \nautomatically, whenever the big repository is updated.\n\nCiao,\nDscho\n"},{"id":"67983","messageId":"9e4733910802081126r5bf19c95rec817a1b6648ee4d@mail.gmail.com","threadId":"11972","inReplyTo":"m3ejbngtnn.fsf@localhost.localdomain","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2008-02-08T19:26:20Z","receivedAt":"2008-02-08T19:26:20Z","isPatch":true,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 2/8/08, Jakub Narebski <jnareb@gmail.com> wrote:\n> Jan Holesovsky <kendy@suse.cz> writes:\n> One of the reasons why 'lazy clone' was not implemented was the fact\n> that by using large enough window, and larger than default delta\n> length you can repack \"archive pack\" (and keep it from trying to\n> repack using .keep files, see git-config(1)) much tighter than with\n> default (time and CPU conserving) options, and much, much tighter than\n> pack which is result of fast-import driven import.\n>\n> Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n> you would need machine with large amount of memory to repack it\n> tightly in sensible time!\n\nA lot of memory is 2-4GB. Without this much memory you will trigger\nswapping and the pack process will finish in about a month. Note that\nonly one machine needs to have this kind of memory. It can be used to\nmake the optimized pack of the project history and mark it with .keep\nfiles. It doesn't take a lot of memory to use the optimized packs,\nonly to make them.\n\nThere are some patches for making repack work multi-core. Not sure if\nthey made it into the main git tree yet. These patches work almost\nlinearly. A eight hour repack will take 2.5 hours on a quad core\nmachine.\n\nThere is very good chance your 1.5GB repo will turn into 300MB if it\nis extremely packed. This is something you only need to do once, but\nyou'll probably end up doing it a dozen times trying to get it just\nright.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"67986","messageId":"alpine.LFD.1.00.0802081457170.2732@xanadu.home","threadId":"11972","inReplyTo":"9e4733910802081126r5bf19c95rec817a1b6648ee4d@mail.gmail.com","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-08T20:09:56Z","receivedAt":"2008-02-08T20:09:56Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 8 Feb 2008, Jon Smirl wrote:\n\n> There are some patches for making repack work multi-core. Not sure if\n> they made it into the main git tree yet.\n\nYes, they are.  You need to compile with\"make THREADED_DELTA_SEARCH=yes\" \nor add THREADED_DELTA_SEARCH=yes into config.mak for it to be enabled \nthough.  Then you have to set the pack.threads configuration variable \nappropriately to use it.\n\n\nNicolas\n"},{"id":"67989","messageId":"alpine.LSU.1.00.0802081905580.11591@racer.site","threadId":"11972","inReplyTo":"200802081828.43849.kendy@suse.cz","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-08T20:16:00Z","receivedAt":"2008-02-08T20:16:00Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n2nd part of my review:\n\nOn Fri, 8 Feb 2008, Jan Holesovsky wrote:\n\n> +static void read_from_stdin(int *num, char ***records)\n> +{\n> +\tchar buffer[4096];\n> +\tsize_t records_num, leftover;\n> +\tssize_t ret;\n> +\n> +\t*num = 0;\n> +\tleftover = 0;\n> +\n> +\trecords_num = 4096;\n> +\t(*records) = xmalloc(records_num * sizeof(char *));\n> +\n> +\tdo {\n> +\t\tchar *p, *last;\n> +\n> +\t\tret = xread(0 /*stdin*/, buffer + leftover,\n> +\t\t\t\tsizeof(buffer) - leftover);\n> +\t\tif (ret < 0)\n> +\t\t\tdie(\"read error on input: %s\", strerror(errno));\n> +\n> +\t\tlast = buffer;\n> +\t\tfor (p = buffer; p < buffer + leftover + ret; p++)\n> +\t\t\tif ((!*p || *p == '\\n') && (p != last)) {\n> +\t\t\t\tif (*num >= records_num) {\n> +\t\t\t\t\trecords_num *= 2;\n> +\t\t\t\t\t(*records) = xrealloc(*records,\n> +\t\t\t\t\t\t\t      records_num * sizeof(char*));\n> +\t\t\t\t}\n> +\n> +\t\t\t\tif (p - last > 0) {\n> +\t\t\t\t\t(*records)[*num] =\n> +\t\t\t\t\t\tstrndup(last, p - last);\n> +\t\t\t\t\t(*num)++;\n> +\t\t\t\t}\n> +\t\t\t\tlast = p + 1;\n> +\t\t\t}\n> +\t\tmemmove(buffer, last, leftover);\n> +\t} while (ret > 0);\n> +\n> +\tif (leftover) {\n> +\t\tif (*num >= records_num) {\n> +\t\t\trecords_num *= 2;\n> +\t\t\t(*records) = xrealloc(*records,\n> +\t\t\t\t\t      records_num * sizeof(char*));\n> +\t\t}\n> +\n> +\t\t(*records)[*num] = strndup(buffer, leftover);\n> +\t\t(*num)++;\n> +\t}\n> +}\n\nI thought about this function again.  It seems we have something similar \nin builtin-pack-objects.c, which is easier to read.  The equivalent would \nbe:\n\nstatic void read_from_stdin(int *num, char ***records)\n{\n\tchar line[4096];\n\tint alloc = 0;\n\n\t*num = 0;\n\t*records = NULL;\n\tfor (;;) {\n\t\tif (!fgets(line, sizeof(line), stdin)) {\n\t\t\tif (feof(stdin))\n\t\t\t\tbreak;\n\t\t\tif (!ferror(stdin))\n\t\t\t\tdie(\"fgets returned NULL, not EOF, nor error!\");\n\t\t\tif (errno != EINTR)\n\t\t\t\tdie(\"fgets: %s\", strerror(errno));\n\t\t\tclearerr(stdin);\n\t\t\tcontinue;\n\t\t}\n\t\tif (!line[0])\n\t\t\tcontinue;\n\t\tALLOC_GROW(*records, *num + 1, alloc);\n\t\t(*records)[(*num)++] = xstrdup(line);\n\t}\n}\t\t\n\n> diff --git a/git-clone.sh b/git-clone.sh\n> index b4e858c..208e9fc 100755\n> --- a/git-clone.sh\n> +++ b/git-clone.sh\n> @@ -115,7 +115,7 @@ Perhaps git-update-server-info needs to be run there?\"\n>  quiet=\n>  local=no\n>  use_local_hardlink=yes\n> -local_shared=no\n> +shared=no\n>  unset template\n>  no_checkout=\n>  upload_pack=\n> @@ -143,7 +143,7 @@ do\n>  \t--no-hardlinks)\n>  \t\tuse_local_hardlink=no ;;\n>  \t-s|--shared)\n> -\t\tlocal_shared=yes ;;\n> +\t\tshared=yes ;;\n>  \t--template)\n>  \t\tshift; template=\"--template=$1\" ;;\n>  \t-q|--quiet)\n> @@ -288,7 +288,7 @@ yes)\n>  \t( cd \"$repo/objects\" ) ||\n>  \t\tdie \"cannot chdir to local '$repo/objects'.\"\n>  \n> -\tif test \"$local_shared\" = yes\n> +\tif test \"$shared\" = yes\n>  \tthen\n>  \t\tmkdir -p \"$GIT_DIR/objects/info\"\n>  \t\techo \"$repo/objects\" >>\"$GIT_DIR/objects/info/alternates\"\n> @@ -364,11 +364,22 @@ yes)\n>  \t\tfi\n>  \t\t;;\n>  \t*)\n> +\t\tcommits_only=\n> +\t\tif test \"$shared\" = yes\n> +\t\tthen\n> +\t\t\tcommits_only=\"--commits-only\"\n> +\t\tfi\n>  \t\tcase \"$upload_pack\" in\n> -\t\t'') git-fetch-pack --all -k $quiet $depth $no_progress \"$repo\";;\n> -\t\t*) git-fetch-pack --all -k $quiet \"$upload_pack\" $depth $no_progress \"$repo\" ;;\n> +\t\t'') git-fetch-pack --all -k $quiet $depth $no_progress $commits_only \"$repo\";;\n> +\t\t*) git-fetch-pack --all -k $quiet \"$upload_pack\" $depth $no_progress $commits_only \"$repo\" ;;\n>  \t\tesac >\"$GIT_DIR/CLONE_HEAD\" ||\n>  \t\t\tdie \"fetch-pack from '$repo' failed.\"\n> +\t\tif test \"$shared\" = yes\n> +\t\tthen\n> +\t\t\t# Must be done after the fetch\n> +\t\t\tmkdir -p \"$GIT_DIR/objects/info\"\n> +\t\t\techo \"$repo\" >> \"$GIT_DIR/objects/info/remote_alternates\"\n> +\t\tfi\n>  \t\t;;\n>  \tesac\n>  \t;;\n\nPlease have a different option than --shared for lazy clones.  Maybe \n--lazy?  ;-)\n\nI can see why you reused --shared, though.  But let's make this more \nfool-proof: a user should explicitely ask for a lazy clone.\n\n> diff --git a/index-pack.c b/index-pack.c\n> index 9fd6982..f2e6b7a 100644\n> --- a/index-pack.c\n> +++ b/index-pack.c\n> @@ -9,7 +9,7 @@\n>  #include \"progress.h\"\n>  \n>  static const char index_pack_usage[] =\n> -\"git-index-pack [-v] [-o <index-file>] [{ ---keep | --keep=<msg> }] { <pack-file> | --stdin [--fix-thin] [<pack-file>] }\";\n> +\"git-index-pack [-v] [-o <index-file>] [{ ---keep | --keep=<msg> }] [--ignore-remote-alternates] { <pack-file> | --stdin [--fix-thin] [<pack-file>] }\";\n>  \n>  struct object_entry\n>  {\n> @@ -746,6 +746,8 @@ int main(int argc, char **argv)\n>  \t\t\t\t\tpack_idx_off32_limit = strtoul(c+1, &c, 0);\n>  \t\t\t\tif (*c || pack_idx_off32_limit & 0x80000000)\n>  \t\t\t\t\tdie(\"bad %s\", arg);\n> +\t\t\t} else if (!strcmp(arg, \"--ignore-remote-alternates\")) {\n> +\t\t\t\tdisable_remote_alternates();\n>  \t\t\t} else\n>  \t\t\t\tusage(index_pack_usage);\n>  \t\t\tcontinue;\n\nI might be missing something, but I do not believe this is necessary.  \nindex-pack only works on packs anyway.  Am I wrong?\n\n> diff --git a/sha1_file.c b/sha1_file.c\n> index 66a4e00..7d60be0 100644\n> --- a/sha1_file.c\n> +++ b/sha1_file.c\n> @@ -14,6 +14,7 @@\n>  #include \"tag.h\"\n>  #include \"tree.h\"\n>  #include \"refs.h\"\n> +#include \"run-command.h\"\n>  \n>  #ifndef O_NOATIME\n>  #if defined(__linux__) && (defined(__i386__) || defined(__PPC__))\n> @@ -411,6 +412,205 @@ static char *find_sha1_file(const unsigned char *sha1, struct stat *st)\n>  \treturn NULL;\n>  }\n>  \n> +static char *remote_alternates = NULL;\n> +static int has_remote_alt_feature = -1;\n> +\n> +void disable_remote_alternates(void)\n> +{\n> +\thas_remote_alt_feature = 0;\n> +}\n> +\n> +static int has_remote_alternates(void)\n> +{\n> +\t/* FIXME: does it make sense to support more URLs inside\n> +\t * remote_alternates? */\n\nI think it would make sense.  For example if you have a local machine \nwhich has most, but maybe not all, of the remote objects.\n\n> +\tstruct stat st;\n> +\tconst char remote_alt_file_name[] = \"info/remote_alternates\";\n\n<bikeshedding>maybe remote-alternates (note the dash instead \nof the underscore)</bikeshedding>\n\n> +\tchar path[PATH_MAX + 1 + sizeof remote_alt_file_name];\n> +\tint fd;\n> +\tchar *map, *p;\n> +\tsize_t mapsz;\n> +\n> +\tif (has_remote_alt_feature != -1)\n> +\t\treturn has_remote_alt_feature;\n> +\n> +\thas_remote_alt_feature = 0;\n> +\n> +\tsprintf(path, \"%s/%s\", get_object_directory(),\n> +\t\t\tremote_alt_file_name);\n> +\tfd = open(path, O_RDONLY);\n> +\tif (fd < 0)\n> +\t\treturn has_remote_alt_feature;\n> +\telse if (fstat(fd, &st) || (st.st_size == 0)) {\n> +\t\tclose(fd);\n> +\t\treturn has_remote_alt_feature;\n> +\t}\n> +\n> +\tmapsz = xsize_t(st.st_size);\n> +\tmap = xmmap(NULL, mapsz, PROT_READ, MAP_PRIVATE, fd, 0);\n> +\tclose(fd);\n> +\n> +\t/* we support just one remote alternate for now,\n> +\t * so read just the first entry */\n> +\tfor (p = map; (p < map + mapsz) && (*p != '\\n'); p++)\n> +\t\t;\n> +\n> +\tremote_alternates = strndup(map, p - map);\n\nSeems that you do something like the read_from_stdin() here, only from a \nfile.  It appears to me as if the function wants to be a library function \n(taking a FILE * parameter, and maybe closing it after use, or even \ntaking a filename parameter, which signifies stdin when NULL).\n\n> +struct sha1_list {\n> +\tunsigned char sha1[20];\n> +\tstruct sha1_list *next;\n> +};\n\nIt'd be probably better to make this an array which uses ALLOC_GROW() in \norder to avoid memory fragmentation/allocation overhead.\n\n> +\tmemset(&fetch_pack, 0, sizeof(fetch_pack));\n> +\tfetch_pack.in = dump_objects.out;\n> +\tfetch_pack.out = 1;\n> +\tfetch_pack.err = 2;\n> +\tfetch_pack.git_cmd = 1;\n> +\tfetch_pack.argv = argv;\n> +\n> +\terr = run_command(&fetch_pack);\n> +\n> +\t/* TODO better error handling - is the object really missing, or\n> +\t * was it just a temporary network error? */\n> +\tif (err) {\n> +\t\tfprintf(stderr, \"error %d while calling fetch-pack\\n\", err);\n> +\t\treturn 0;\n\nThat is a\n\n\t\treturn error(\"Error %d while calling fetch-pack\", err);\n\nAnd it does not really matter what type of error it is: you must report \nthe error and continue without this object.\n\n> +static int fill_remote_list(const unsigned char *sha1,\n> +\t\tconst char *base, int baselen,\n> +\t\tconst char *pathname, unsigned mode, int stage)\n> +{\n> +\tif (!has_sha1_file_locally(sha1)) {\n> +\t\tstruct sha1_list *item;\n> +\n> +\t\titem = xmalloc(sizeof(*item));\n> +\t\thashcpy(item->sha1, sha1);\n> +\t\titem->next = remote_list;\n> +\n> +\t\tremote_list = item;\n> +\t}\n> +\n> +\treturn 0;\n> +}\n> +\n> +static int fetch_remote_sha1s_recursive(struct sha1_list *objects)\n> +{\n> +\tstruct sha1_list *list;\n> +\tint ret = 0;\n> +\n> +\t/* first of all, fetch the missing objects */\n> +\tif (!fetch_remote_sha1s(objects))\n> +\t\treturn 0;\n> +\n> +\tremote_list = NULL;\n> +\n> +\tlist = objects;\n> +\twhile (list) {\n> +\t\tstruct tree *tree;\n> +\n> +\t\ttree = parse_tree_indirect(list->sha1);\n> +\t\tif (tree) {\n> +\t\t\tread_tree_recursive(tree, \"\", 0, 0, NULL,\n> +\t\t\t\t\tfill_remote_list);\n> +\t\t}\n\nThe curly brackets are not necessary.  Plus, with fill_remote_list() as \nyou defined it, it will break down with submodules (see 481f0ee6(Fix \nrev-list when showing objects involving submodules) for inspiration).\n\n> +\n> +\t\tlist = list->next;\n> +\t}\n> +\n> +\tlist = remote_list;\n> +\tif (!list)\n> +\t\treturn 1; /* hooray, we have everything */\n> +\n> +\tret = fetch_remote_sha1s_recursive(list);\n\nThis just cries out loud for a non-recursive approach: have two arrays, \nclear the second, fetch the objects in the first array, then fill the \nsecond with the objects referred to by the first array's objects.  Then \nswap the arrays.  Loop.\n\n> @@ -2316,6 +2532,18 @@ int has_sha1_file(const unsigned char *sha1)\n>  \treturn find_sha1_file(sha1, &st) ? 1 : 0;\n>  }\n>  \n> +int has_sha1_file(const unsigned char *sha1)\n> +{\n> +\tif (has_sha1_file_locally(sha1))\n> +\t\treturn 1;\n> +\n> +\t/* download it if necessary */\n> +\tif (has_remote_alternates() && download_remote_sha1(sha1))\n\nMaybe it would be nicer to have the has_remote_alternates() check only in \ndownload_remote_sha1()?  Same applies to read_sha1_file().\n\n> @@ -106,9 +106,15 @@ static int do_rev_list(int fd, void *create_full_pack)\n>  \tif (create_full_pack)\n>  \t\tuse_thin_pack = 0; /* no point doing it */\n>  \tinit_revisions(&revs, NULL);\n> -\trevs.tag_objects = 1;\n> -\trevs.tree_objects = 1;\n> -\trevs.blob_objects = 1;\n> +\tif (!commits_only) {\n> +\t\trevs.tag_objects = 1;\n> +\t\trevs.tree_objects = 1;\n> +\t\trevs.blob_objects = 1;\n> +\t} else {\n> +\t\trevs.tag_objects = 0;\n> +\t\trevs.tree_objects = 0;\n> +\t\trevs.blob_objects = 0;\n> +\t}\n\nOr\n\n\trevs.tag_objects = revs.tree_objects = revs.blob_objects\n\t\t= !commits_only;\n\n\n> @@ -498,9 +525,15 @@ static void receive_needs(void)\n>  \t\t * asks for something like \"master~10\" (symbolic)...\n>  \t\t * would it make sense?  I don't know.\n>  \t\t */\n> -\t\to = lookup_object(sha1_buf);\n> -\t\tif (!o || !(o->flags & OUR_REF))\n> -\t\t\tdie(\"git-upload-pack: not our ref %s\", line+5);\n> +\t\tif (!exact_objects) {\n> +\t\t\to = lookup_object(sha1_buf);\n> +\t\t\tif (!o || !(o->flags & OUR_REF))\n> +\t\t\t\tdie(\"git-upload-pack: not our ref %s\", line+5);\n> +\t\t} else {\n> +\t\t\to = lookup_unknown_object(sha1_buf);\n> +\t\t\tif (!o)\n> +\t\t\t\tdie(\"git-upload-pack: not an object %s\", line+5);\n> +\t\t}\n\nHmm... AFAICT lookup_unknown_object() does not return NULL.  It creates a \n\"none\" object if it did not find anything under that sha1.\n\nI think you'd rather want\n\n \t\to = lookup_object(sha1_buf);\n-\t\tif (!o || !(o->flags & OUR_REF))\n+\t\tif (!o || (!exact_objects && !(o->flags & OUR_REF)))\n \t\t\tdie(\"git-upload-pack: not our ref %s\", line+5);\n\nPuh.  What a big patch!  But as I said, it is nice to know somebody is \nworking on this.  (I do not necessarily see possibilities to break it \ndown into smaller chunks, though.)\n\nBut I think that your needs can be satisfied with partial shallow clones, \ntoo: e.g.\n\n\t$ mkdir my-new-workdir\n\t$ cd my-new-workdir\n\t$ git init\n\t$ git remote add -t master origin <url>\n\t$ git fetch --depth 1 origin\n\t$ git checkout -b master origin/master\n\nI cannot think of a proper place to make this a one-shot command.\n\nAs you probably know, I am a strong believer in semantics, so I would hate \n\"git clone\" being taught to not clone the whole repository, but only a \nsingle branch.\n\nBut hey, I have been wrong before.\n\nCiao,\nDscho\n"},{"id":"67990","messageId":"1202502007.12966.30.camel@brick","threadId":"11972","inReplyTo":"9e4733910802081126r5bf19c95rec817a1b6648ee4d@mail.gmail.com","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2008-02-08T20:19:39Z","receivedAt":"2008-02-08T20:19:39Z","isPatch":true,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"On Fri, 2008-02-08 at 14:26 -0500, Jon Smirl wrote:\n> On 2/8/08, Jakub Narebski <jnareb@gmail.com> wrote:\n> > Jan Holesovsky <kendy@suse.cz> writes:\n> > One of the reasons why 'lazy clone' was not implemented was the fact\n> > that by using large enough window, and larger than default delta\n> > length you can repack \"archive pack\" (and keep it from trying to\n> > repack using .keep files, see git-config(1)) much tighter than with\n> > default (time and CPU conserving) options, and much, much tighter than\n> > pack which is result of fast-import driven import.\n> >\n> > Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n> > you would need machine with large amount of memory to repack it\n> > tightly in sensible time!\n> \n> A lot of memory is 2-4GB. Without this much memory you will trigger\n> swapping and the pack process will finish in about a month. \n\nWell, my modest little Celeron M laptop w/ 1GB of ram did the full\nrepack overnight on the gcc repo, so a month is a bit of an\nexaggeration.\n\nCheers,\n\nHarvey\n"},{"id":"67991","messageId":"9e4733910802081224k28310b0cj171453c96802ec7f@mail.gmail.com","threadId":"11972","inReplyTo":"1202502007.12966.30.camel@brick","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2008-02-08T20:24:02Z","receivedAt":"2008-02-08T20:24:02Z","isPatch":true,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 2/8/08, Harvey Harrison <harvey.harrison@gmail.com> wrote:\n> On Fri, 2008-02-08 at 14:26 -0500, Jon Smirl wrote:\n> > On 2/8/08, Jakub Narebski <jnareb@gmail.com> wrote:\n> > > Jan Holesovsky <kendy@suse.cz> writes:\n> > > One of the reasons why 'lazy clone' was not implemented was the fact\n> > > that by using large enough window, and larger than default delta\n> > > length you can repack \"archive pack\" (and keep it from trying to\n> > > repack using .keep files, see git-config(1)) much tighter than with\n> > > default (time and CPU conserving) options, and much, much tighter than\n> > > pack which is result of fast-import driven import.\n> > >\n> > > Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n> > > you would need machine with large amount of memory to repack it\n> > > tightly in sensible time!\n> >\n> > A lot of memory is 2-4GB. Without this much memory you will trigger\n> > swapping and the pack process will finish in about a month.\n>\n> Well, my modest little Celeron M laptop w/ 1GB of ram did the full\n> repack overnight on the gcc repo, so a month is a bit of an\n> exaggeration.\n\nTry it again with window=250 and depth=250. That's how you get the\nreally small packs.\n\n>\n> Cheers,\n>\n> Harvey\n>\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"67992","messageId":"1202502302.12966.32.camel@brick","threadId":"11972","inReplyTo":"9e4733910802081224k28310b0cj171453c96802ec7f@mail.gmail.com","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2008-02-08T20:25:02Z","receivedAt":"2008-02-08T20:25:02Z","isPatch":true,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"On Fri, 2008-02-08 at 15:24 -0500, Jon Smirl wrote:\n> On 2/8/08, Harvey Harrison <harvey.harrison@gmail.com> wrote:\n> > On Fri, 2008-02-08 at 14:26 -0500, Jon Smirl wrote:\n> > > On 2/8/08, Jakub Narebski <jnareb@gmail.com> wrote:\n> > > > Jan Holesovsky <kendy@suse.cz> writes:\n> > > > One of the reasons why 'lazy clone' was not implemented was the fact\n> > > > that by using large enough window, and larger than default delta\n> > > > length you can repack \"archive pack\" (and keep it from trying to\n> > > > repack using .keep files, see git-config(1)) much tighter than with\n> > > > default (time and CPU conserving) options, and much, much tighter than\n> > > > pack which is result of fast-import driven import.\n> > > >\n> > > > Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n> > > > you would need machine with large amount of memory to repack it\n> > > > tightly in sensible time!\n> > >\n> > > A lot of memory is 2-4GB. Without this much memory you will trigger\n> > > swapping and the pack process will finish in about a month.\n> >\n> > Well, my modest little Celeron M laptop w/ 1GB of ram did the full\n> > repack overnight on the gcc repo, so a month is a bit of an\n> > exaggeration.\n> \n> Try it again with window=250 and depth=250. That's how you get the\n> really small packs.\n> \n\nYes, I know, and I did if you remember back to the gcc discussion.\n\nHarvey\n"},{"id":"67998","messageId":"9e4733910802081241g1543f873h76181c0d3165cc00@mail.gmail.com","threadId":"11972","inReplyTo":"1202502302.12966.32.camel@brick","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2008-02-08T20:41:49Z","receivedAt":"2008-02-08T20:41:49Z","isPatch":true,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 2/8/08, Harvey Harrison <harvey.harrison@gmail.com> wrote:\n> On Fri, 2008-02-08 at 15:24 -0500, Jon Smirl wrote:\n> > On 2/8/08, Harvey Harrison <harvey.harrison@gmail.com> wrote:\n> > > On Fri, 2008-02-08 at 14:26 -0500, Jon Smirl wrote:\n> > > > On 2/8/08, Jakub Narebski <jnareb@gmail.com> wrote:\n> > > > > Jan Holesovsky <kendy@suse.cz> writes:\n> > > > > One of the reasons why 'lazy clone' was not implemented was the fact\n> > > > > that by using large enough window, and larger than default delta\n> > > > > length you can repack \"archive pack\" (and keep it from trying to\n> > > > > repack using .keep files, see git-config(1)) much tighter than with\n> > > > > default (time and CPU conserving) options, and much, much tighter than\n> > > > > pack which is result of fast-import driven import.\n> > > > >\n> > > > > Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n> > > > > you would need machine with large amount of memory to repack it\n> > > > > tightly in sensible time!\n> > > >\n> > > > A lot of memory is 2-4GB. Without this much memory you will trigger\n> > > > swapping and the pack process will finish in about a month.\n> > >\n> > > Well, my modest little Celeron M laptop w/ 1GB of ram did the full\n> > > repack overnight on the gcc repo, so a month is a bit of an\n> > > exaggeration.\n> >\n> > Try it again with window=250 and depth=250. That's how you get the\n> > really small packs.\n> >\n>\n> Yes, I know, and I did if you remember back to the gcc discussion.\n\nNow that you mention it I seem to recall some changes were made to git\nduring that discussion that reduced the memory footprint and made the\noptimized gcc repack fit into 1GB. I've forgotten the exact timings\nand git is a moving target. When I was working on Mozilla it needed\n2.4GB to avoid swapping but that was with a much older git.\n\nThe rule is: if it starts swapping it is going to take way longer that\nyou are probably willing to wait. Buying more RAM is a cheap and easy\nfix.\n\nIf people are having trouble with large repositories please let the\ngit community know and your issues will probably get quickly fixed.\nWe can't fix something we don't know about.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"68010","messageId":"foihu9$110$1@ger.gmane.org","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802081905580.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-08T21:35:03Z","receivedAt":"2008-02-08T21:35:03Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Johannes Schindelin wrote:\n> On Fri, 8 Feb 2008, Jan Holesovsky wrote:\n\n>> +     struct stat st;\n>> +     const char remote_alt_file_name[] = \"info/remote_alternates\";\n> \n> <bikeshedding>maybe remote-alternates (note the dash instead \n> of the underscore)</bikeshedding>\n\nWhy not in info/alternates?\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"68014","messageId":"alpine.LSU.1.00.0802082151570.11591@racer.site","threadId":"11972","inReplyTo":"foihu9$110$1@ger.gmane.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-08T21:52:08Z","receivedAt":"2008-02-08T21:52:08Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n[I'll just Cc you, out of the goodness of my heart.]\n\nOn Fri, 8 Feb 2008, Jakub Narebski wrote:\n\n> Johannes Schindelin wrote:\n> > On Fri, 8 Feb 2008, Jan Holesovsky wrote:\n> \n> >> +     struct stat st;\n> >> +     const char remote_alt_file_name[] = \"info/remote_alternates\";\n> > \n> > <bikeshedding>maybe remote-alternates (note the dash instead \n> > of the underscore)</bikeshedding>\n> \n> Why not in info/alternates?\n\nAgain, to make the distinction clear.\n\nAlso note that info/alternates is used by the http transport (which would \nthen break semi-silently, because I expect that you usually put git:// \nurls into remote-alternates).\n\nCiao,\nDscho\n"},{"id":"68017","messageId":"20080208220356.GA22064@glandium.org","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802082151570.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-02-08T22:03:56Z","receivedAt":"2008-02-08T22:03:56Z","isPatch":true,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Fri, Feb 08, 2008 at 09:52:08PM +0000, Johannes Schindelin wrote:\n> Hi,\n> \n> [I'll just Cc you, out of the goodness of my heart.]\n> \n> On Fri, 8 Feb 2008, Jakub Narebski wrote:\n> \n> > Johannes Schindelin wrote:\n> > > On Fri, 8 Feb 2008, Jan Holesovsky wrote:\n> > \n> > >> +     struct stat st;\n> > >> +     const char remote_alt_file_name[] = \"info/remote_alternates\";\n> > > \n> > > <bikeshedding>maybe remote-alternates (note the dash instead \n> > > of the underscore)</bikeshedding>\n> > \n> > Why not in info/alternates?\n> \n> Again, to make the distinction clear.\n> \n> Also note that info/alternates is used by the http transport (which would \n> then break semi-silently, because I expect that you usually put git:// \n> urls into remote-alternates).\n\nAlso note that the http transport uses info/http-alternates for http://\nurls. By the way, it doesn't make much sense that only http-fetch uses\nit.\n\nMike\n"},{"id":"68023","messageId":"alpine.LSU.1.00.0802082234170.11591@racer.site","threadId":"11972","inReplyTo":"20080208220356.GA22064@glandium.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-08T22:34:55Z","receivedAt":"2008-02-08T22:34:55Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 8 Feb 2008, Mike Hommey wrote:\n\n> Also note that the http transport uses info/http-alternates for http:// \n> urls. By the way, it doesn't make much sense that only http-fetch uses \n> it.\n\nI think it does make sense: nobody else needs http-alternates.\n\nCiao,\nDscho\n"},{"id":"68024","messageId":"20080208225024.GA26975@glandium.org","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802082234170.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-02-08T22:50:24Z","receivedAt":"2008-02-08T22:50:24Z","isPatch":true,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Fri, Feb 08, 2008 at 10:34:55PM +0000, Johannes Schindelin wrote:\n> Hi,\n> \n> On Fri, 8 Feb 2008, Mike Hommey wrote:\n> \n> > Also note that the http transport uses info/http-alternates for http:// \n> > urls. By the way, it doesn't make much sense that only http-fetch uses \n> > it.\n> \n> I think it does make sense: nobody else needs http-alternates.\n\nIf you're setting an http-alternate, it means objects are missing in the\nrepo. If they are missing in the repo and are not in alternates, how can \nany other command needing objects out there work on the repo ?\n\nMike\n"},{"id":"68029","messageId":"alpine.LSU.1.00.0802082312520.11591@racer.site","threadId":"11972","inReplyTo":"20080208225024.GA26975@glandium.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-08T23:14:40Z","receivedAt":"2008-02-08T23:14:40Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 8 Feb 2008, Mike Hommey wrote:\n\n> On Fri, Feb 08, 2008 at 10:34:55PM +0000, Johannes Schindelin wrote:\n> \n> > On Fri, 8 Feb 2008, Mike Hommey wrote:\n> > \n> > > Also note that the http transport uses info/http-alternates for \n> > > http:// urls. By the way, it doesn't make much sense that only \n> > > http-fetch uses it.\n> > \n> > I think it does make sense: nobody else needs http-alternates.\n> \n> If you're setting an http-alternate, it means objects are missing in the \n> repo. If they are missing in the repo and are not in alternates, how can \n> any other command needing objects out there work on the repo ?\n\nThe point is: if you have a bare repository on a server that uses \nalternates, that path stored in info/alternates is usable by git-daemon.  \nBut it is not usable by git-http-fetch, since that does not have a \ngit-aware server side.  So if you want to reuse the _same_ bare repository \n_with_ alternates for both git:// transport and http:// transport, you \n_need_ to _different_ alternates: one being a path on the server, and \nanother being an http:// url for http-fetch.\n\nHth,\nDscho\n"},{"id":"68030","messageId":"20080208233856.GA31593@glandium.org","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802082312520.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-02-08T23:38:56Z","receivedAt":"2008-02-08T23:38:56Z","isPatch":true,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Fri, Feb 08, 2008 at 11:14:40PM +0000, Johannes Schindelin wrote:\n> Hi,\n> \n> On Fri, 8 Feb 2008, Mike Hommey wrote:\n> \n> > On Fri, Feb 08, 2008 at 10:34:55PM +0000, Johannes Schindelin wrote:\n> > \n> > > On Fri, 8 Feb 2008, Mike Hommey wrote:\n> > > \n> > > > Also note that the http transport uses info/http-alternates for \n> > > > http:// urls. By the way, it doesn't make much sense that only \n> > > > http-fetch uses it.\n> > > \n> > > I think it does make sense: nobody else needs http-alternates.\n> > \n> > If you're setting an http-alternate, it means objects are missing in the \n> > repo. If they are missing in the repo and are not in alternates, how can \n> > any other command needing objects out there work on the repo ?\n> \n> The point is: if you have a bare repository on a server that uses \n> alternates, that path stored in info/alternates is usable by git-daemon.  \n> But it is not usable by git-http-fetch, since that does not have a \n> git-aware server side.  So if you want to reuse the _same_ bare repository \n> _with_ alternates for both git:// transport and http:// transport, you \n> _need_ to _different_ alternates: one being a path on the server, and \n> another being an http:// url for http-fetch.\n\nBut nothing prevents you from only setting an http-alternate. Also not\nhttp-fetch can deal fine with info/alternates if it contains relative\npaths.\n\nMike\n"},{"id":"68088","messageId":"200802091525.36284.kendy@suse.cz","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802081250240.2732@xanadu.home","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2008-02-09T14:25:35Z","receivedAt":"2008-02-09T14:25:35Z","isPatch":true,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Nicolas,\n\nOn Friday 08 February 2008 19:03, Nicolas Pitre wrote:\n\n> > I've provided a git import of OOo with the entire history; the problem is\n> > that the pack has 2.5G, so it's not too convenient to download for casual\n> > developers that just want to try it.\n>\n> How did you repack your repository?\n>\n> We know that current defaults are not suitable for large projects.  For\n> example, the gcc git repository shrinked from 1.5GB pack down to 230MB\n> after some tuning.\n\nAfter the suggestions in this thread I tried to experiment with the --window \nand --depth options of git-repack, and indeed, there are still reserves.\n\nSo far I'm at 2G (saved 500M), unfortunately the aggressive values like \n--window=250 --depth=250 that someone mentioned here cause out-of-memory on a \nmachine with 8G :-(  If there's anybody brave enough here to try as well, I'd \nbe grateful.  Maybe it would be also interesting to _exactly_ locate what \ncauses the oom, and eg. exclude the object from the pack if possible.\n\nThe tree is available here:\n\ngit clone git://o3-build.services.openoffice.org/git/ooo.git\ngit clone http://o3-build.services.openoffice.org/~svn/ooo.git (the same over \nhttp://)\n\nThank you in advance!\n\nRegards,\nJan\n"},{"id":"68089","messageId":"200802091527.42696.kendy@suse.cz","threadId":"11972","inReplyTo":"1202494475.31361.45.camel@brick","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2008-02-09T14:27:42Z","receivedAt":"2008-02-09T14:27:42Z","isPatch":true,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Harvey,\n\nOn Friday 08 February 2008 19:14, Harvey Harrison wrote:\n\n> For comparison, how big was the svn repo you're testing?  My experience\n> has been about 15-20 times smaller than SVN once a tuned repack has\n> been done.\n\nAnother guy created the SVN repo, IIRC he said it had 55G.\n\nRegards,\nJan\n"},{"id":"68091","messageId":"200802091606.53001.kendy@suse.cz","threadId":"11972","inReplyTo":"20080208184902.GA19404@glandium.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2008-02-09T15:06:52Z","receivedAt":"2008-02-09T15:06:52Z","isPatch":true,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Mike,\n\nOn Friday 08 February 2008 19:49, Mike Hommey wrote:\n\n> There are 2 things, here:\n> - Probably, you can make your pack smaller with proper window sizing.\n> Try taking a look at the \"Git and GCC\" that crossed borders between\n> the gcc and the git mailing lists.\n\nJust trying this :-)\n\n> - There are tricks to do roughly what you want without modifying git.\n> For example, you can prepare several \"shared\" clones of your repo (git\n> clone -s) and leave in each only a few branches. Cloning from these will\n> only pull the needed data.\n\nGood to know about this, thank you!  The problem currently is that we are \ntrying to produce SVN and git trees containing the same data, the same number \nof branches, etc. for the sake of comparison.  If git wins, and it will be \nchosen for OOo, we'll be hopefully able to do more tuning - and I'm sure I'll \nask here for help ;-)\n\nRegards,\nJan\n"},{"id":"68093","messageId":"200802091627.25913.kendy@suse.cz","threadId":"11972","inReplyTo":"m3ejbngtnn.fsf@localhost.localdomain","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2008-02-09T15:27:25Z","receivedAt":"2008-02-09T15:27:25Z","isPatch":true,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Jakub,\n\nOn Friday 08 February 2008 20:00, Jakub Narebski wrote:\n\n> It was not implemented because it was thought to be hard; git assumes\n> in many places that if it has an object, it has all objects referenced\n> by it.\n>\n> But it is very nice of you to [try to] implement 'lazy clone'/'remote\n> alternates'.\n>\n> Could you provide some benchmarks (time, network throughtput, latency)\n> for your implementation?\n\nUnfortunately not yet :-(  The only data I have that clone done on \ngit://localhost/ooo.git took 10 minutes without the lazy clone, and 7.5 \nminutes with it - and then I sent the patch for review here ;-)  The deadline \nfor our SVN vs. git comparison for OOo is the next Friday, so I'll definitely \nhave some better data by then.\n\n> Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n> you would need machine with large amount of memory to repack it\n> tightly in sensible time!\n\nAs I answered elsewhere, unfortunately it goes out of memory even on 8G \nmachine (x86-64), so...  But still trying.\n\n> > Shallow clone is not a possibility - we don't get patches through\n> > mailing lists, so we need the pull/push, and also thanks to the OOo\n> > development cycle, we have too many living heads which causes the\n> > shallow clone to download about 1.5G even with --depth 1.\n>\n> Wouldn't be easier to try to fix shallow clone implementation to allow\n> for pushing from shallow to full clone (fetching from full to shallow\n> is implemented), and perhaps also push/pull between two shallow\n> clones?\n\nI tried to look into it a bit, but unfortunately did not see a clear way how \nto do it transparently for the user - say you pull a branch that is based off \na commit you do not have.  But of course, I could have missed something ;-)\n\n> As to many living heads: first, you don't need to fetch all\n> heads. Currently git-clone has no option to select subset of heads to\n> clone, but you can always use git-init + hand configuration +\n> git-remote and git-fetch for actual fetching.\n\nRight, might be interesting as well.  But still the missing push/pull is \nproblematic for us [or at least I see it as a problem ;-)].\n\n> By the way, did you try to split OpenOffice.org repository at the\n> components boundary into submodules (subprojects)? This would also\n> limit amount of needed download, as you don't neeed to download and\n> checkout all subprojects.\n\nYes, and got to much nicer repositories by that ;-) - by only moving some \nbinary stuff out of the CVS to a separate tree.  The problem is that the deal \nis to compare the same stuff in SVN and git - so no choice for me in fact.\n\n> The problem of course is _how_ to split repository into\n> submodules. Submodules should be enough self contained so the\n> whole-tree commit is alsays (or almost always) only about submodule.\n\nI hope it will be doable _if_ the git wins & will be chosen for OOo.\n\n> > Lazy clone sounded like the right idea to me.  With this\n> > proof-of-concept implementation, just about 550M from the 2.5G is\n> > downloaded, which is still about twice as much in comparison with\n> > downloading a tarball, but bearable.\n>\n> Do you have any numbers for OOo repository like number of revisions,\n> depth of DAG of commits (maximum number of revisions in one line of\n> commits), number of files, size of checkout, average size of file,\n> etc.?\n\nI'll try to provide the data ASAP.\n\nRegards,\nJan\n"},{"id":"68094","messageId":"200802091654.20473.kendy@suse.cz","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802082151570.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2008-02-09T15:54:20Z","receivedAt":"2008-02-09T15:54:20Z","isPatch":true,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Johannes,\n\nOne more 'thank you' for the review - now publically :-)\n\nOn Friday 08 February 2008 22:52, Johannes Schindelin wrote:\n\n> > > <bikeshedding>maybe remote-alternates (note the dash instead\n> > > of the underscore)</bikeshedding>\n> >\n> > Why not in info/alternates?\n>\n> Again, to make the distinction clear.\n\nYes; still even though the 'alteranates' and 'remote alternates' have some of \nthe ideas common, the implementation differs (and has to differ) - so I think \nyou are right even in the --lazy option for clone instead of reusing -s.\n\nFor the rest - I'll post the updated patch ASAP.\n\nRegards,\nJan\n"},{"id":"68125","messageId":"20080209212059.GB17147@efreet.light.src","threadId":"11972","inReplyTo":"20080208233856.GA31593@glandium.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2008-02-09T21:20:59Z","receivedAt":"2008-02-09T21:20:59Z","isPatch":true,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Sat, Feb 09, 2008 at 00:38:56 +0100, Mike Hommey wrote:\n> On Fri, Feb 08, 2008 at 11:14:40PM +0000, Johannes Schindelin wrote:\n> > On Fri, 8 Feb 2008, Mike Hommey wrote:\n> > > On Fri, Feb 08, 2008 at 10:34:55PM +0000, Johannes Schindelin wrote:\n> > > > On Fri, 8 Feb 2008, Mike Hommey wrote:\n> > > > > Also note that the http transport uses info/http-alternates for \n> > > > > http:// urls. By the way, it doesn't make much sense that only \n> > > > > http-fetch uses it.\n> > > > \n> > > > I think it does make sense: nobody else needs http-alternates.\n> > > \n> > > If you're setting an http-alternate, it means objects are missing in the \n> > > repo. If they are missing in the repo and are not in alternates, how can \n> > > any other command needing objects out there work on the repo ?\n> > \n> > The point is: if you have a bare repository on a server that uses \n> > alternates, that path stored in info/alternates is usable by git-daemon.  \n> > But it is not usable by git-http-fetch, since that does not have a \n> > git-aware server side.  So if you want to reuse the _same_ bare repository \n> > _with_ alternates for both git:// transport and http:// transport, you \n> > _need_ to _different_ alternates: one being a path on the server, and \n> > another being an http:// url for http-fetch.\n> \n> But nothing prevents you from only setting an http-alternate. Also not\n> http-fetch can deal fine with info/alternates if it contains relative\n> paths.\n\nThey still may not work because of whatever mapping of paths to URLs the http\nserver does. Also relative paths in info/alternates don't actually work; or\nrather, they do, but /not recursively/ (the code seems fixable, just someone\nwould have to make sure the proper base is always used).\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"68128","messageId":"20080209220551.GA30139@glandium.org","threadId":"11972","inReplyTo":"200802091525.36284.kendy@suse.cz","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-02-09T22:05:51Z","receivedAt":"2008-02-09T22:05:51Z","isPatch":true,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Sat, Feb 09, 2008 at 03:25:35PM +0100, Jan Holesovsky wrote:\n> Hi Nicolas,\n> \n> On Friday 08 February 2008 19:03, Nicolas Pitre wrote:\n> \n> > > I've provided a git import of OOo with the entire history; the problem is\n> > > that the pack has 2.5G, so it's not too convenient to download for casual\n> > > developers that just want to try it.\n> >\n> > How did you repack your repository?\n> >\n> > We know that current defaults are not suitable for large projects.  For\n> > example, the gcc git repository shrinked from 1.5GB pack down to 230MB\n> > after some tuning.\n> \n> After the suggestions in this thread I tried to experiment with the --window \n> and --depth options of git-repack, and indeed, there are still reserves.\n> \n> So far I'm at 2G (saved 500M), unfortunately the aggressive values like \n> --window=250 --depth=250 that someone mentioned here cause out-of-memory on a \n> machine with 8G :-(. If there's anybody brave enough here to try as well, I'd \n> be grateful.  Maybe it would be also interesting to _exactly_ locate what \n> causes the oom, and eg. exclude the object from the pack if possible.\n\nSpeaking of which, I haven't taken a look at builtin-pack-objects.c deep\nenough but shouldn't it be possible to do prepare_pack and\nwrite_pack_file in one pass ?\n\nMike\n"},{"id":"68132","messageId":"alpine.LFD.1.00.0802091838070.2732@xanadu.home","threadId":"11972","inReplyTo":"20080209220551.GA30139@glandium.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-09T23:38:39Z","receivedAt":"2008-02-09T23:38:39Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 9 Feb 2008, Mike Hommey wrote:\n\n> Speaking of which, I haven't taken a look at builtin-pack-objects.c deep\n> enough but shouldn't it be possible to do prepare_pack and\n> write_pack_file in one pass ?\n\nNo.\n\n\nNicolas\n"},{"id":"68159","messageId":"alpine.LFD.1.00.0802092200350.2732@xanadu.home","threadId":"11972","inReplyTo":"200802091627.25913.kendy@suse.cz","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-10T03:10:06Z","receivedAt":"2008-02-10T03:10:06Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 9 Feb 2008, Jan Holesovsky wrote:\n\n> On Friday 08 February 2008 20:00, Jakub Narebski wrote:\n> \n> > Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n> > you would need machine with large amount of memory to repack it\n> > tightly in sensible time!\n> \n> As I answered elsewhere, unfortunately it goes out of memory even on 8G \n> machine (x86-64), so...  But still trying.\n\nTry setting the following config variables as follows:\n\n\tgit config pack.deltaCacheLimit 1\n\tgit config pack.deltaCacheSize 1\n\tgit config pack.windowMemory 1g\n\nThat should help keeping memory usage somewhat bounded.\n\n\nNicolas\n"},{"id":"68171","messageId":"BAYC1-PASMTP10AF630E8A5B3D6C255317AE290@CEZ.ICE","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802092200350.2732@xanadu.home","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2008-02-10T04:59:49Z","receivedAt":"2008-02-10T04:59:49Z","isPatch":true,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Sat, 09 Feb 2008 22:10:06 -0500 (EST)\nNicolas Pitre <nico@cam.org> wrote:\n\n> On Sat, 9 Feb 2008, Jan Holesovsky wrote:\n> \n> > On Friday 08 February 2008 20:00, Jakub Narebski wrote:\n> > \n> > > Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n> > > you would need machine with large amount of memory to repack it\n> > > tightly in sensible time!\n> > \n> > As I answered elsewhere, unfortunately it goes out of memory even on 8G \n> > machine (x86-64), so...  But still trying.\n> \n> Try setting the following config variables as follows:\n> \n> \tgit config pack.deltaCacheLimit 1\n> \tgit config pack.deltaCacheSize 1\n> \tgit config pack.windowMemory 1g\n> \n> That should help keeping memory usage somewhat bounded.\n> \n\nHi Nicolas,\n\nTried that earlier today and got a 1.6G pack (on a 2G machine).  There are\nsome big objects in that repo.. over 100 are 30 to 62M in size, 400 more\nover 10M, and ~40,000 over 100K.  Would you expect a larger memory window\n(on a better machine) to help shrink the repo down any more?\n\nSean\n"},{"id":"68172","messageId":"alpine.LFD.1.00.0802100017380.2732@xanadu.home","threadId":"11972","inReplyTo":"BAYC1-PASMTP10AF630E8A5B3D6C255317AE290@CEZ.ICE","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-10T05:22:09Z","receivedAt":"2008-02-10T05:22:09Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 9 Feb 2008, Sean wrote:\n\n> On Sat, 09 Feb 2008 22:10:06 -0500 (EST)\n> Nicolas Pitre <nico@cam.org> wrote:\n> \n> > On Sat, 9 Feb 2008, Jan Holesovsky wrote:\n> > \n> > > On Friday 08 February 2008 20:00, Jakub Narebski wrote:\n> > > \n> > > > Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n> > > > you would need machine with large amount of memory to repack it\n> > > > tightly in sensible time!\n> > > \n> > > As I answered elsewhere, unfortunately it goes out of memory even on 8G \n> > > machine (x86-64), so...  But still trying.\n> > \n> > Try setting the following config variables as follows:\n> > \n> > \tgit config pack.deltaCacheLimit 1\n> > \tgit config pack.deltaCacheSize 1\n> > \tgit config pack.windowMemory 1g\n> > \n> > That should help keeping memory usage somewhat bounded.\n> > \n> \n> Hi Nicolas,\n> \n> Tried that earlier today and got a 1.6G pack (on a 2G machine).  There are\n> some big objects in that repo.. over 100 are 30 to 62M in size, 400 more\n> over 10M, and ~40,000 over 100K.  Would you expect a larger memory window\n> (on a better machine) to help shrink the repo down any more?\n\nWell, I don't think so.  Anyway, with the above pack.windowMemory \nsetting, the window probably gets shrinked if those big objects are all \nto be found in the same window.  So that would be the setting to \nincrease if you have lots of ram.\n\nFinding out what those huge objects are, and if they actually need to be \nthere, would be a good thing to do to reduce any repository size.\n\n\nNicolas\n"},{"id":"68173","messageId":"BAYC1-PASMTP059B375F7660D93F93647DAE290@CEZ.ICE","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802100017380.2732@xanadu.home","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2008-02-10T05:35:17Z","receivedAt":"2008-02-10T05:35:17Z","isPatch":true,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Sun, 10 Feb 2008 00:22:09 -0500 (EST)\nNicolas Pitre <nico@cam.org> wrote:\n\n> Well, I don't think so.  Anyway, with the above pack.windowMemory \n> setting, the window probably gets shrinked if those big objects are all \n> to be found in the same window.  So that would be the setting to \n> increase if you have lots of ram.\n\nSounds like it would be worthwhile then for Jan to try on that 8G machine\nand see what comes out the other end.\n\n> Finding out what those huge objects are, and if they actually need to be \n> there, would be a good thing to do to reduce any repository size.\n\nOkay, i've sent the sha1's of the top 500 to Jan for inspection.  It appears\nthat many of the largest objects are automatically generated i18n files that\ncould be regenerated from source files when needed rather than being checked\nin themselves; but that's for the OO folks to decide.\n\nThanks,\nSean\n"},{"id":"68179","messageId":"e5bfff550802092323u3ec3c9c8uf6e92399395efd27@mail.gmail.com","threadId":"11972","inReplyTo":"200802091525.36284.kendy@suse.cz","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2008-02-10T07:23:33Z","receivedAt":"2008-02-10T07:23:33Z","isPatch":true,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On Feb 9, 2008 3:25 PM, Jan Holesovsky <kendy@suse.cz> wrote:\n> Hi Nicolas,\n>\n> On Friday 08 February 2008 19:03, Nicolas Pitre wrote:\n>\n> > > I've provided a git import of OOo with the entire history; the problem is\n> > > that the pack has 2.5G, so it's not too convenient to download for casual\n> > > developers that just want to try it.\n> >\n\nSorry to enter so late in this thread. I just would like to ask if you\nhave evaluated a different approach for casual developers.\n\nThe approach is the one used by Linux tree.\n\nLinux git repository is not very big and can be downloaded with easy.\nOn the other end Linux history spans many more years then the repo\ndoes.\n\nThe design choice here is two have *two repositories*, one with recent\nstuff and one historical, with stuff older then version 2.6.12\n\nWe have to say that this choice come by accident due to Linus\nswitching from bitkeeper to git around 2.6.12 but today it's a more or\nless a conscious choice because there exists the git historical repo,\nconverted from bk, and this repo is still kept separated, also if\ntechnically could be grafted to the main one to create a super big\nLinux repo.\n\nAdvantage of this approach are:\n\n- Lean and fast everyday repos, where actual development occurs\n\n- Easy clone also for casual users\n\n- Possibility to have anyway the whole history when needed\n\nA variation on this theme could be to have always two repos, one with\nrecent stuff, say last 5 years of development, and one with *the\nwhole* history, not only with old stuff as in the historical Linux\ntree, in this case it's easier for people that need digging very old\nchanges to do this avoiding browsing two repos as occurs now with\nLinux.\n\nMarco\n\nP.S: Idea here is that of a kind of cache memory for git repos ;-)\n"},{"id":"68183","messageId":"854pch5f4w.fsf@lupus.strangled.net","threadId":"11972","inReplyTo":"BAYC1-PASMTP10AF630E8A5B3D6C255317AE290@CEZ.ICE","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Joachim B Haga","fromEmail":"cjhaga@fys.uio.no","sentAt":"2008-02-10T09:34:23Z","receivedAt":"2008-02-10T09:34:23Z","isPatch":true,"sender":{"key":"cjhaga@fys.uio.no","avatar":null},"body":"Sean <seanlkml@sympatico.ca> writes:\n\n>> \tgit config pack.deltaCacheLimit 1\n>> \tgit config pack.deltaCacheSize 1\n>> \tgit config pack.windowMemory 1g\n>\n> Tried that earlier today and got a 1.6G pack (on a 2G machine).  There are\n> some big objects in that repo.. over 100 are 30 to 62M in size, 400 more\n> over 10M, and ~40,000 over 100K.  Would you expect a larger memory window\n> (on a better machine) to help shrink the repo down any more?\n\nI tried without these, 1.47GiB packfile. Peak RSS ~14G.\n\n-j.\n"},{"id":"68197","messageId":"alpine.LSU.1.00.0802101207330.11591@racer.site","threadId":"11972","inReplyTo":"e5bfff550802092323u3ec3c9c8uf6e92399395efd27@mail.gmail.com","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-10T12:08:40Z","receivedAt":"2008-02-10T12:08:40Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 10 Feb 2008, Marco Costalba wrote:\n\n> Linux git repository is not very big and can be downloaded with easy. On \n> the other end Linux history spans many more years then the repo does.\n> \n> The design choice here is two have *two repositories*, one with recent \n> stuff and one historical, with stuff older then version 2.6.12\n\nI do not think that this is an option: Jan already tried a shallow clone \n(which would amount to something like what you propose), and it was still \ntoo large.\n\nCiao,\nDscho\n"},{"id":"68227","messageId":"alpine.LSU.1.00.0802101640570.11591@racer.site","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802092200350.2732@xanadu.home","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-10T16:43:04Z","receivedAt":"2008-02-10T16:43:04Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sat, 9 Feb 2008, Nicolas Pitre wrote:\n\n> On Sat, 9 Feb 2008, Jan Holesovsky wrote:\n> \n> > On Friday 08 February 2008 20:00, Jakub Narebski wrote:\n> > \n> > > Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n> > > you would need machine with large amount of memory to repack it\n> > > tightly in sensible time!\n> > \n> > As I answered elsewhere, unfortunately it goes out of memory even on 8G \n> > machine (x86-64), so...  But still trying.\n> \n> Try setting the following config variables as follows:\n> \n> \tgit config pack.deltaCacheLimit 1\n> \tgit config pack.deltaCacheSize 1\n> \tgit config pack.windowMemory 1g\n> \n> That should help keeping memory usage somewhat bounded.\n\nI tried that:\n\n$ git config pack.deltaCacheLimit 1\n$ git config pack.deltaCacheSize 1\n$ git config pack.windowMemory 2g\n$ #/usr/bin/time git repack -a -d -f --window=250 --depth=250\n$ du -s objects/\n2548137 objects/\n$ /usr/bin/time git repack -a -d -f --window=250 --depth=250\nCounting objects: 2477715, done.\nfatal: Out of memory, malloc failed411764)\nCommand exited with non-zero status 1\n9356.95user 53.33system 2:38:58elapsed 98%CPU (0avgtext+0avgdata \n0maxresident)k\n0inputs+0outputs (31929major+18088744minor)pagefaults 0swaps\n\nNote that this is on a 2.4GHz Quadcode CPU with 3.5GB RAM.\n\nI'm retrying with smaller values, but at over 2.5 hours per try, this is \ngetting tedious.\n\nCiao,\nDscho\n"},{"id":"68229","messageId":"ee77f5c20802100846g10937a49m4901f88a70a6de0@mail.gmail.com","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802101207330.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"David Symonds","fromEmail":"dsymonds@gmail.com","sentAt":"2008-02-10T16:46:30Z","receivedAt":"2008-02-10T16:46:30Z","isPatch":true,"sender":{"key":"dsymonds@gmail.com","avatar":"https://gravatar.com/avatar/b22f5051cbfc11836e36cf7a690e6cde4e225d835e13295ff98d15c7a9ee3c0f?d=mp&s=160"},"body":"On Feb 10, 2008 4:08 AM, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> Hi,\n>\n> On Sun, 10 Feb 2008, Marco Costalba wrote:\n>\n> > Linux git repository is not very big and can be downloaded with easy. On\n> > the other end Linux history spans many more years then the repo does.\n> >\n> > The design choice here is two have *two repositories*, one with recent\n> > stuff and one historical, with stuff older then version 2.6.12\n>\n> I do not think that this is an option: Jan already tried a shallow clone\n> (which would amount to something like what you propose), and it was still\n> too large.\n\nI think that was still pulling all the branches, so a shallow clone of\njust a couple of branches might be feasible.\n\n\nDave.\n"},{"id":"68233","messageId":"9e4733910802100901m729b0cdfg85ccc0ca77011249@mail.gmail.com","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802101640570.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2008-02-10T17:01:47Z","receivedAt":"2008-02-10T17:01:47Z","isPatch":true,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 2/10/08, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> Hi,\n>\n> On Sat, 9 Feb 2008, Nicolas Pitre wrote:\n>\n> > On Sat, 9 Feb 2008, Jan Holesovsky wrote:\n> >\n> > > On Friday 08 February 2008 20:00, Jakub Narebski wrote:\n> > >\n> > > > Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n> > > > you would need machine with large amount of memory to repack it\n> > > > tightly in sensible time!\n> > >\n> > > As I answered elsewhere, unfortunately it goes out of memory even on 8G\n> > > machine (x86-64), so...  But still trying.\n> >\n> > Try setting the following config variables as follows:\n> >\n> >       git config pack.deltaCacheLimit 1\n> >       git config pack.deltaCacheSize 1\n> >       git config pack.windowMemory 1g\n> >\n> > That should help keeping memory usage somewhat bounded.\n>\n> I tried that:\n>\n> $ git config pack.deltaCacheLimit 1\n> $ git config pack.deltaCacheSize 1\n> $ git config pack.windowMemory 2g\n> $ #/usr/bin/time git repack -a -d -f --window=250 --depth=250\n> $ du -s objects/\n> 2548137 objects/\n> $ /usr/bin/time git repack -a -d -f --window=250 --depth=250\n> Counting objects: 2477715, done.\n> fatal: Out of memory, malloc failed411764)\n> Command exited with non-zero status 1\n> 9356.95user 53.33system 2:38:58elapsed 98%CPU (0avgtext+0avgdata\n> 0maxresident)k\n> 0inputs+0outputs (31929major+18088744minor)pagefaults 0swaps\n>\n> Note that this is on a 2.4GHz Quadcode CPU with 3.5GB RAM.\n\nTurning on multi-core support greatly increases the memory\nconsumption; at least double the single thread case.\n\nGoing over the original repository and deleting (get all copies out of\nthe history) those giant i18n files generated by programs than Sean\nrefers to would be my first step. If you have 5,000 revisions of a\n10MB file I suspect it would take a huge amount of memory to pack.\nPlus you have to copy all of that pointless history around.\n\n>\n> I'm retrying with smaller values, but at over 2.5 hours per try, this is\n> getting tedious.\n>\n> Ciao,\n> Dscho\n>\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"68236","messageId":"alpine.LSU.1.00.0802101735530.11591@racer.site","threadId":"11972","inReplyTo":"9e4733910802100901m729b0cdfg85ccc0ca77011249@mail.gmail.com","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-10T17:36:12Z","receivedAt":"2008-02-10T17:36:12Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 10 Feb 2008, Jon Smirl wrote:\n\n> Turning on multi-core support greatly increases the memory consumption; \n> at least double the single thread case.\n\nThat's why I did not do it.\n\nCiao,\nDscho\n"},{"id":"68238","messageId":"alpine.LSU.1.00.0802101649110.11591@racer.site","threadId":"11972","inReplyTo":"ee77f5c20802100846g10937a49m4901f88a70a6de0@mail.gmail.com","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-10T17:45:00Z","receivedAt":"2008-02-10T17:45:00Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 10 Feb 2008, David Symonds wrote:\n\n> On Feb 10, 2008 4:08 AM, Johannes Schindelin \n> <Johannes.Schindelin@gmx.de> wrote:\n>\n> > On Sun, 10 Feb 2008, Marco Costalba wrote:\n> >\n> > > Linux git repository is not very big and can be downloaded with \n> > > easy. On the other end Linux history spans many more years then the \n> > > repo does.\n> > >\n> > > The design choice here is two have *two repositories*, one with \n> > > recent stuff and one historical, with stuff older then version \n> > > 2.6.12\n> >\n> > I do not think that this is an option: Jan already tried a shallow \n> > clone (which would amount to something like what you propose), and it \n> > was still too large.\n> \n> I think that was still pulling all the branches, so a shallow clone of \n> just a couple of branches might be feasible.\n\nIndeed:\n\n$ git ls-remote git://o3-build.services.openoffice.org/git/ooo.git|wc -l\n3970\n$ git ls-remote --heads git://o3-build.services.openoffice.org/git/ooo.git|\n\twc -l\n751\n\nFetching just master is a little hard on the server (it spends quite a \nlot of time deltifying -- minutes! -- especially between 80% and 95%, \nand indexing is even slower), but other than \nthat:\n\n$ /usr/bin/time git fetch --depth=1 \\\n\tgit://o3-build.services.openoffice.org/git/ooo.git \\\n\tmaster:refs/remotes/origin/master\nwarning: no common commits\nremote: Generating pack...\nremote: Done counting 79934 objects.\nremote: Deltifying 79934 objects...\nremote:  100% (79934/79934) done\nIndexing 79934 objects...\nremote: Total 79934 (delta 34549), reused 51323 (delta 20737)\n 100% (79934/79934) done\nResolving 34549 deltas...\n 100% (34549/34549) done\n* refs/remotes/origin/master: storing branch 'master' of \ngit://o3-build.services.openoffice.org/git/ooo\n  commit: 29990e4\n46.48user 4.60system 16:48.29elapsed 5%CPU (0avgtext+0avgdata \n0maxresident)k\n0inputs+0outputs (0major+941205minor)pagefaults 0swaps\n\n$ du .git/objects/pack/\n464688  .git/objects/pack/\n$ /usr/bin/time git repack -a -d -f --window=250 --depth=250\nGenerating pack...\nDone counting 79934 objects.\nDeltifying 79934 objects...\n 100% (79934/79934) done\nWriting 79934 objects...\n 100% (79934/79934) done\nTotal 79934 (delta 40013), reused 0 (delta 0)\nPack pack-350e4edca93ee75ef3d85269284a24775bf6b24f created.\nRemoving unused objects 100%...\nDone.\n1869.78user 6.66system 31:36.50elapsed 98%CPU (0avgtext+0avgdata \n0maxresident)k\n0inputs+0outputs (2031major+1753824minor)pagefaults 0swaps\n$ du .git/objects/pack/\n454636  .git/objects/pack/\n\nOf course, the clone time would be reduced dramatically if the repository \nyou clone from has only \"master\", and is fully (re-)packed.\n\nSo I was not completely correct in my assumption that a clear cut a la \nlinux-2.6 (possibly grafting historical-linux) would not help.\n\nCiao,\nDscho\n"},{"id":"68245","messageId":"alpine.LSU.1.00.0802101845320.11591@racer.site","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802101640570.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-10T18:47:32Z","receivedAt":"2008-02-10T18:47:32Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 10 Feb 2008, Johannes Schindelin wrote:\n\n> On Sat, 9 Feb 2008, Nicolas Pitre wrote:\n> \n> > On Sat, 9 Feb 2008, Jan Holesovsky wrote:\n> > \n> > > On Friday 08 February 2008 20:00, Jakub Narebski wrote:\n> > > \n> > > > Both Mozilla import, and GCC import were packed below 0.5 GB. \n> > > > Warning: you would need machine with large amount of memory to \n> > > > repack it tightly in sensible time!\n> > > \n> > > As I answered elsewhere, unfortunately it goes out of memory even on \n> > > 8G machine (x86-64), so...  But still trying.\n> > \n> > Try setting the following config variables as follows:\n> > \n> > \tgit config pack.deltaCacheLimit 1\n> > \tgit config pack.deltaCacheSize 1\n> > \tgit config pack.windowMemory 1g\n> > \n> > That should help keeping memory usage somewhat bounded.\n> \n> I tried that:\n> \n> $ git config pack.deltaCacheLimit 1\n> $ git config pack.deltaCacheSize 1\n> $ git config pack.windowMemory 2g\n> $ #/usr/bin/time git repack -a -d -f --window=250 --depth=250\n> $ du -s objects/\n> 2548137 objects/\n> $ /usr/bin/time git repack -a -d -f --window=250 --depth=250\n> Counting objects: 2477715, done.\n> fatal: Out of memory, malloc failed411764)\n> Command exited with non-zero status 1\n> 9356.95user 53.33system 2:38:58elapsed 98%CPU (0avgtext+0avgdata \n> 0maxresident)k\n> 0inputs+0outputs (31929major+18088744minor)pagefaults 0swaps\n> \n> Note that this is on a 2.4GHz Quadcode CPU with 3.5GB RAM.\n> \n> I'm retrying with smaller values, but at over 2.5 hours per try, this is \n> getting tedious.\n\nNow, _that_ is strange.  Using 150 instead of 250 brings it down even \nquicker!\n\n$ /usr/bin/time git repack -a -d -f --window=150 --depth=150\nCounting objects: 2477715, done.\nCompressing objects:  19% (481551/2411764)\nCompressing objects:  19% (482333/2411764)\nfatal: Out of memory, malloc failed411764)\nCommand exited with non-zero status 1\n7118.37user 54.15system 2:01:44elapsed 98%CPU (0avgtext+0avgdata \n0maxresident)k\n0inputs+0outputs (29834major+17122977minor)pagefaults 0swaps\n\n(I hit the Return key twice during the time I suspected it would go out of \nmemory, so it might have been really at 20%.)\n\nIdeas?\n\nCiao,\nDscho\n"},{"id":"68249","messageId":"alpine.LFD.1.00.0802101437040.2732@xanadu.home","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802101845320.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-10T19:42:44Z","receivedAt":"2008-02-10T19:42:44Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 10 Feb 2008, Johannes Schindelin wrote:\n\n> Hi,\n> \n> On Sun, 10 Feb 2008, Johannes Schindelin wrote:\n> \n> > On Sat, 9 Feb 2008, Nicolas Pitre wrote:\n> > \n> > > On Sat, 9 Feb 2008, Jan Holesovsky wrote:\n> > > \n> > > > On Friday 08 February 2008 20:00, Jakub Narebski wrote:\n> > > > \n> > > > > Both Mozilla import, and GCC import were packed below 0.5 GB. \n> > > > > Warning: you would need machine with large amount of memory to \n> > > > > repack it tightly in sensible time!\n> > > > \n> > > > As I answered elsewhere, unfortunately it goes out of memory even on \n> > > > 8G machine (x86-64), so...  But still trying.\n> > > \n> > > Try setting the following config variables as follows:\n> > > \n> > > \tgit config pack.deltaCacheLimit 1\n> > > \tgit config pack.deltaCacheSize 1\n> > > \tgit config pack.windowMemory 1g\n> > > \n> > > That should help keeping memory usage somewhat bounded.\n> > \n> > I tried that:\n> > \n> > $ git config pack.deltaCacheLimit 1\n> > $ git config pack.deltaCacheSize 1\n> > $ git config pack.windowMemory 2g\n> > $ #/usr/bin/time git repack -a -d -f --window=250 --depth=250\n> > $ du -s objects/\n> > 2548137 objects/\n> > $ /usr/bin/time git repack -a -d -f --window=250 --depth=250\n> > Counting objects: 2477715, done.\n> > fatal: Out of memory, malloc failed411764)\n> > Command exited with non-zero status 1\n> > 9356.95user 53.33system 2:38:58elapsed 98%CPU (0avgtext+0avgdata \n> > 0maxresident)k\n> > 0inputs+0outputs (31929major+18088744minor)pagefaults 0swaps\n> > \n> > Note that this is on a 2.4GHz Quadcode CPU with 3.5GB RAM.\n> > \n> > I'm retrying with smaller values, but at over 2.5 hours per try, this is \n> > getting tedious.\n> \n> Now, _that_ is strange.  Using 150 instead of 250 brings it down even \n> quicker!\n> \n> $ /usr/bin/time git repack -a -d -f --window=150 --depth=150\n> Counting objects: 2477715, done.\n> Compressing objects:  19% (481551/2411764)\n> Compressing objects:  19% (482333/2411764)\n> fatal: Out of memory, malloc failed411764)\n> Command exited with non-zero status 1\n> 7118.37user 54.15system 2:01:44elapsed 98%CPU (0avgtext+0avgdata \n> 0maxresident)k\n> 0inputs+0outputs (29834major+17122977minor)pagefaults 0swaps\n> \n> (I hit the Return key twice during the time I suspected it would go out of \n> memory, so it might have been really at 20%.)\n> \n> Ideas?\n\nYou're probably hitting the same memory allocator fragmentation issue I \nhad with the gcc repo.  On my machine with 1GB of ram, I was able to \nrepack the 1.5GB source pack just fine, but repacking the 300MB source \npack was impossible due to memory exhaustion.\n\nMy theory is that the smaller pack has many more deltas with deeper \ndelta chains, and this is stumping much harder on the memory allocator \nwhich fails to prevent fragmentation at some point.  When Jon Smirl \ntested Git using the Google memory allocator there was around 1GB less \nallocated, which might indicate that the glibc allocator has issues with \nsome of Git's workloads.\n\n\nNicolas\n"},{"id":"68250","messageId":"alpine.LFD.1.00.0802101443240.2732@xanadu.home","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802101649110.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-10T19:45:23Z","receivedAt":"2008-02-10T19:45:23Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 10 Feb 2008, Johannes Schindelin wrote:\n\n> Resolving 34549 deltas...\n>  100% (34549/34549) done\n\nWhat Git version is this?\n\nYou better try out 1.5.4 for packing comparisons.  It produces slightly \ntighter packs than 1.5.3.\n\n\nNicolas\n"},{"id":"68251","messageId":"alpine.LFD.1.00.0802101445430.2732@xanadu.home","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802101640570.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-10T19:50:49Z","receivedAt":"2008-02-10T19:50:49Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 10 Feb 2008, Johannes Schindelin wrote:\n\n> I tried that:\n> \n> $ git config pack.deltaCacheLimit 1\n> $ git config pack.deltaCacheSize 1\n> $ git config pack.windowMemory 2g\n\nThis has nothing to do with repacking memory usage, but even tighter \npacks can be obtained with:\n\n\tgit config repack.usedeltabaseoffset true\n\nThis is not the default yet.\n\n\nNicolas\n"},{"id":"68253","messageId":"9e4733910802101211v11c0e285teb36f06d9a0e5f37@mail.gmail.com","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802101437040.2732@xanadu.home","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2008-02-10T20:11:48Z","receivedAt":"2008-02-10T20:11:48Z","isPatch":true,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 2/10/08, Nicolas Pitre <nico@cam.org> wrote:\n> On Sun, 10 Feb 2008, Johannes Schindelin wrote:\n>\n> > Hi,\n> >\n> > On Sun, 10 Feb 2008, Johannes Schindelin wrote:\n> >\n> > > On Sat, 9 Feb 2008, Nicolas Pitre wrote:\n> > >\n> > > > On Sat, 9 Feb 2008, Jan Holesovsky wrote:\n> > > >\n> > > > > On Friday 08 February 2008 20:00, Jakub Narebski wrote:\n> > > > >\n> > > > > > Both Mozilla import, and GCC import were packed below 0.5 GB.\n> > > > > > Warning: you would need machine with large amount of memory to\n> > > > > > repack it tightly in sensible time!\n> > > > >\n> > > > > As I answered elsewhere, unfortunately it goes out of memory even on\n> > > > > 8G machine (x86-64), so...  But still trying.\n> > > >\n> > > > Try setting the following config variables as follows:\n> > > >\n> > > >   git config pack.deltaCacheLimit 1\n> > > >   git config pack.deltaCacheSize 1\n> > > >   git config pack.windowMemory 1g\n> > > >\n> > > > That should help keeping memory usage somewhat bounded.\n> > >\n> > > I tried that:\n> > >\n> > > $ git config pack.deltaCacheLimit 1\n> > > $ git config pack.deltaCacheSize 1\n> > > $ git config pack.windowMemory 2g\n> > > $ #/usr/bin/time git repack -a -d -f --window=250 --depth=250\n> > > $ du -s objects/\n> > > 2548137 objects/\n> > > $ /usr/bin/time git repack -a -d -f --window=250 --depth=250\n> > > Counting objects: 2477715, done.\n> > > fatal: Out of memory, malloc failed411764)\n> > > Command exited with non-zero status 1\n> > > 9356.95user 53.33system 2:38:58elapsed 98%CPU (0avgtext+0avgdata\n> > > 0maxresident)k\n> > > 0inputs+0outputs (31929major+18088744minor)pagefaults 0swaps\n> > >\n> > > Note that this is on a 2.4GHz Quadcode CPU with 3.5GB RAM.\n> > >\n> > > I'm retrying with smaller values, but at over 2.5 hours per try, this is\n> > > getting tedious.\n> >\n> > Now, _that_ is strange.  Using 150 instead of 250 brings it down even\n> > quicker!\n> >\n> > $ /usr/bin/time git repack -a -d -f --window=150 --depth=150\n> > Counting objects: 2477715, done.\n> > Compressing objects:  19% (481551/2411764)\n> > Compressing objects:  19% (482333/2411764)\n> > fatal: Out of memory, malloc failed411764)\n> > Command exited with non-zero status 1\n> > 7118.37user 54.15system 2:01:44elapsed 98%CPU (0avgtext+0avgdata\n> > 0maxresident)k\n> > 0inputs+0outputs (29834major+17122977minor)pagefaults 0swaps\n> >\n> > (I hit the Return key twice during the time I suspected it would go out of\n> > memory, so it might have been really at 20%.)\n> >\n> > Ideas?\n>\n> You're probably hitting the same memory allocator fragmentation issue I\n> had with the gcc repo.  On my machine with 1GB of ram, I was able to\n> repack the 1.5GB source pack just fine, but repacking the 300MB source\n> pack was impossible due to memory exhaustion.\n>\n> My theory is that the smaller pack has many more deltas with deeper\n> delta chains, and this is stumping much harder on the memory allocator\n> which fails to prevent fragmentation at some point.  When Jon Smirl\n> tested Git using the Google memory allocator there was around 1GB less\n> allocated, which might indicate that the glibc allocator has issues with\n> some of Git's workloads.\n\nI'm forgetting everything again, but I seem to recall that the Google\nallocator only made a significant difference with multithreading.  It\nis much better at keeping the threads from fragmenting each other.\nIt's very easy to try it, all you have to do is add another lib the\nthe link command.\n\n\n>\n>\n> Nicolas\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"68258","messageId":"alpine.LSU.1.00.0802102032211.11591@racer.site","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802101443240.2732@xanadu.home","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-10T20:32:41Z","receivedAt":"2008-02-10T20:32:41Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 10 Feb 2008, Nicolas Pitre wrote:\n\n> On Sun, 10 Feb 2008, Johannes Schindelin wrote:\n> \n> > Resolving 34549 deltas...\n> >  100% (34549/34549) done\n> \n> What Git version is this?\n> \n> You better try out 1.5.4 for packing comparisons.  It produces slightly \n> tighter packs than 1.5.3.\n\nOoops.  I thought I updated, but no: 1.5.3.6.2835.gf9ebf\n\nCiao,\nDscho\n"},{"id":"68304","messageId":"200802110220.08078.jnareb@gmail.com","threadId":"11972","inReplyTo":"200802091627.25913.kendy@suse.cz","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-11T01:20:07Z","receivedAt":"2008-02-11T01:20:07Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hi, Jan!\n\nOn Sat, 9 Feb 2008, Jan Holesovsky wrote:\n> On Friday 08 February 2008 20:00, Jakub Narebski wrote:\n> \n>> It was not implemented because it was thought to be hard; git assumes\n>> in many places that if it has an object, it has all objects referenced\n>> by it.\n>>\n>> But it is very nice of you to [try to] implement 'lazy clone'/'remote\n>> alternates'.\n>>\n>> Could you provide some benchmarks (time, network throughtput, latency)\n>> for your implementation?\n> \n> Unfortunately not yet :-(  The only data I have that clone done on \n> git://localhost/ooo.git took 10 minutes without the lazy clone, and 7.5 \n> minutes with it - and then I sent the patch for review here ;-)  The deadline \n> for our SVN vs. git comparison for OOo is the next Friday, so I'll definitely \n> have some better data by then.\n\nHere perhaps another optimization which wasn't done because git is\nfast enough on moderately-sized repositories, namely that IIRC git-clone\n(and git-fetch for sure) over native (smart) protocol recreates pack,\neven if sometimes better and simplier would be to just copy (transfer)\nexisting pack.\n\nBut this would need multi-pack \"extension\". (it should work just now\nwithout transport protocol extension, receiver must only be aware\nof the need to split resulting pack, and index them all).\n\n>> Both Mozilla import, and GCC import were packed below 0.5 GB. Warning:\n>> you would need machine with large amount of memory to repack it\n>> tightly in sensible time!\n> \n> As I answered elsewhere, unfortunately it goes out of memory even on 8G \n> machine (x86-64), so...  But still trying.\n\nI hope that would work better...\n\n>>> Shallow clone is not a possibility - we don't get patches through\n>>> mailing lists, so we need the pull/push, and also thanks to the OOo\n>>> development cycle, we have too many living heads which causes the\n>>> shallow clone to download about 1.5G even with --depth 1.\n>>\n>> Wouldn't be easier to try to fix shallow clone implementation to allow\n>> for pushing from shallow to full clone (fetching from full to shallow\n>> is implemented), and perhaps also push/pull between two shallow\n>> clones?\n> \n> I tried to look into it a bit, but unfortunately did not see a clear way how \n> to do it transparently for the user - say you pull a branch that is based off \n> a commit you do not have.  But of course, I could have missed something ;-)\n\nIf I remember correctly fetching _into_ shallow clone works correctly,\nas deepening depth of shallow clone. What is not implemented AFAIK, but\nshould be not too hard would be to allow to push from shallow clone\nto full clone. This way the network of full clones (functioning as\ncentres to publish your work) and shallow + few branches repos (working\nrepositories).\n\nI don't know if that would be enough.\n\nFor better support git would need to exchange graft-like information,\nand use union of restrictions to get correct commits.\n\n\nPerhaps it would be best to mail 'shallow clone' author...\n\n>> As to many living heads: first, you don't need to fetch all\n>> heads. Currently git-clone has no option to select subset of heads to\n>> clone, but you can always use git-init + hand configuration +\n>> git-remote and git-fetch for actual fetching.\n> \n> Right, might be interesting as well.  But still the missing push/pull is \n> problematic for us [or at least I see it as a problem ;-)].\n\nYou can configure separate 'remote's for the same repository\nwith different heads. This would work both for pull and for push.\n\n\nI think the solution proposed by Marco Costalba, namely of creating\n\"archive\" repository, and \"live\" repository, joining them if needed\nby grafts, similarly to how linux kernel has live repo, and historical\nimport repo, would be good alternative to shallow or lazy clone.\n\nThere would be \"archive\" repo (or repos), read only, with whole history,\nvery tightly packed with kept packs, with all branches and all tags,\nand \"live\" repo, with only current history (a year, or since major\nAPI change, or from today, or something like that), with only important\nbranches (or repos, each containg important for a team set of branches).\nThere would be prepared graft file to join two histories, if you have\nto examine full history. Hopefully repo would be smaller.\n\n>> By the way, did you try to split OpenOffice.org repository at the\n>> components boundary into submodules (subprojects)? This would also\n>> limit amount of needed download, as you don't neeed to download and\n>> checkout all subprojects.\n> \n> Yes, and got to much nicer repositories by that ;-) - by only moving some \n> binary stuff out of the CVS to a separate tree.  The problem is that the deal \n> is to compare the same stuff in SVN and git - so no choice for me in fact.\n\nSidenote: due to (from what I have read) heavy use of topic branches\nin OOo development, Subversion would have to be used with svnmerge\nextension, or together with SVK, to make work with it not complete\npain.\n\nIn CVS you could have ad-hoc modules, and ad-hoc partial checkouts\n(so called 'modules'), but that plays merry hell with whole tree,\natomic, recoverable state commits. In Git you have to plan carefully\nboundaries between submodules / subprojects. Additional advantage\nis that you would have boundaries more clear, and better modularity\nusually leads to better code.\n\nComparing directly Subversion and Git is a bit stupid: they promote\ndifferent workflows. From what I've read Git with its ability to very\neasily create branches, with easy _merging_ of branches, and ability\nto easily create _private_ branches (testing branches) have much\ncommon witch chosen OOo SCM workflow. Playing to strentghs of\nSubversion because that is why you used because of limits of previously\nused tools is not smart.\n\nBut if you have to, then you have to. Git would hopefully get lazy\nclone support from your effort. But perhaps it would be possible\n(if additional work) to prepare two repositories: first the same\nas Subversion (and same as now in CVS), second one \"how it should\nbe done with Git\".\n\n>> The problem of course is _how_ to split repository into\n>> submodules. Submodules should be enough self contained so the\n>> whole-tree commit is alsays (or almost always) only about submodule.\n> \n> I hope it will be doable _if_ the git wins & will be chosen for OOo.\n\nI hope that ability to work with submodules (with ability to not\nclone / checkout modules if not needed), i.e. \"svn:externals\ndone right\" to para[hrase SVN slogan, would be one of reasons to\nchose Git over Subversion.\n\n>>> Lazy clone sounded like the right idea to me.  With this\n>>> proof-of-concept implementation, just about 550M from the 2.5G is\n>>> downloaded, which is still about twice as much in comparison with\n>>> downloading a tarball, but bearable.\n>>\n>> Do you have any numbers for OOo repository like number of revisions,\n>> depth of DAG of commits (maximum number of revisions in one line of\n>> commits), number of files, size of checkout, average size of file,\n>> etc.?\n> \n> I'll try to provide the data ASAP.\n\nFor example what is the size of full checkout (all version-control\nmanaged files). Of for example it is 0.5 GB it would be hard to go\nto less that 0.5GB or so with pack size, even with compression\nof objects themselves in pack file.\n\n\nSuch large repositories, like Mozilla, GCC, or now OpenOffice.org\ntests the limits of Git. Perhaps snapshot-based distributed SCMs\ncannot deal sensibly with such large projects; I hope this is not\nthe case.\n\nI wonder if packv4 improvements, which development stalled because\n(if I understand correctly) because it didn't brough as much\nimprovements, and what is now was good enough for up-till-now\nprojects, would help with OpenOffice.org repository...\n\n\nP.S. From what I have read OOo uses CVS + some branch DB; does\nyour importer make use of this branch info database?\n\n-- \nJakub Narebski\nPoland\n"},{"id":"68307","messageId":"200802110242.27324.jnareb@gmail.com","threadId":"11972","inReplyTo":"BAYC1-PASMTP059B375F7660D93F93647DAE290@CEZ.ICE","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-11T01:42:26Z","receivedAt":"2008-02-11T01:42:26Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sun, 10 Feb 2008, Sean napisał:\n> On Sun, 10 Feb 2008 00:22:09 -0500 (EST)\n> Nicolas Pitre <nico@cam.org> wrote:\n>\n>> Finding out what those huge objects are, and if they actually need to be \n>> there, would be a good thing to do to reduce any repository size.\n> \n> Okay, i've sent the sha1's of the top 500 to Jan for inspection.  It appears\n> that many of the largest objects are automatically generated i18n files that\n> could be regenerated from source files when needed rather than being checked\n> in themselves; but that's for the OO folks to decide.\n\nGood practice is to not add generated files to version control.\nBut sometimes such files are stored if regenerating them is costly\n(./configure file in some cases, 'man' and 'html' branches in git.git).\n\nIIRC Dana How tried also to deal with repository with large binary\nfiles in repo, although in that case those had shallow history. IIRC\nthe proposed solution was to pack all such large objects undeltified\ninto separate \"large-objects\" kept pack.\n\nYou can mark large files with (undocumented except for RelNotes)\n'delta' gitattribute, but I don't know if it would help in your\ncase.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"68311","messageId":"alpine.LFD.1.00.0802102059260.2732@xanadu.home","threadId":"11972","inReplyTo":"200802110242.27324.jnareb@gmail.com","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-11T02:04:12Z","receivedAt":"2008-02-11T02:04:12Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 11 Feb 2008, Jakub Narebski wrote:\n\n> On Sun, 10 Feb 2008, Sean napisał:\n> > On Sun, 10 Feb 2008 00:22:09 -0500 (EST)\n> > Nicolas Pitre <nico@cam.org> wrote:\n> >\n> >> Finding out what those huge objects are, and if they actually need to be \n> >> there, would be a good thing to do to reduce any repository size.\n> > \n> > Okay, i've sent the sha1's of the top 500 to Jan for inspection.  It appears\n> > that many of the largest objects are automatically generated i18n files that\n> > could be regenerated from source files when needed rather than being checked\n> > in themselves; but that's for the OO folks to decide.\n> \n> Good practice is to not add generated files to version control.\n> But sometimes such files are stored if regenerating them is costly\n> (./configure file in some cases, 'man' and 'html' branches in git.git).\n> \n> IIRC Dana How tried also to deal with repository with large binary\n> files in repo, although in that case those had shallow history. IIRC\n> the proposed solution was to pack all such large objects undeltified\n> into separate \"large-objects\" kept pack.\n\nThat was to solve a completely different problem which wasn't about \nspace saving, but rather to save on 'git push' latency.\n\n\nNicolas\n"},{"id":"68356","messageId":"200802111111.54187.jnareb@gmail.com","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802102059260.2732@xanadu.home","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-11T10:11:53Z","receivedAt":"2008-02-11T10:11:53Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Mon, 11 Feb 2008, Nicolas Pitre wrote:\n> On Mon, 11 Feb 2008, Jakub Narebski wrote:\n>> On Sun, 10 Feb 2008, Sean napisał:\n>>> On Sun, 10 Feb 2008 00:22:09 -0500 (EST)\n>>> Nicolas Pitre <nico@cam.org> wrote:\n>>>\n>>>> Finding out what those huge objects are, and if they actually need to be \n>>>> there, would be a good thing to do to reduce any repository size.\n\n>> IIRC Dana How tried also to deal with repository with large binary\n>> files in repo, although in that case those had shallow history. IIRC\n>> the proposed solution was to pack all such large objects undeltified\n>> into separate \"large-objects\" kept pack.\n> \n> That was to solve a completely different problem which wasn't about \n> space saving, but rather to save on 'git push' latency.\n\nSorry, my mistake.\n\nAlthough in Dana case separating large blobs into non-packed loose\nobjects (her patches), or separate kept non-delta large blobs only\npack (proposed solution), were shared over networked filesystem.\nSo the amortized size of repository was smaller... ;-ppp\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"68358","messageId":"47B01FE7.8010207@op5.se","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802081457170.2732@xanadu.home","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2008-02-11T10:13:59Z","receivedAt":"2008-02-11T10:13:59Z","isPatch":true,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n> On Fri, 8 Feb 2008, Jon Smirl wrote:\n> \n>> There are some patches for making repack work multi-core. Not sure if\n>> they made it into the main git tree yet.\n> \n> Yes, they are.  You need to compile with\"make THREADED_DELTA_SEARCH=yes\" \n> or add THREADED_DELTA_SEARCH=yes into config.mak for it to be enabled \n> though.  Then you have to set the pack.threads configuration variable \n> appropriately to use it.\n> \n\nI sent a patch to get it to auto-detect multi-core machines, but I see\nnow that it was commented upon for finalization (by Nicolas, actually)\nand I must have missed that, thinking it had been applied because I got\nan accidental merge in my own tree.\n\nAs such, I've been using that patch the last several months without\nproblems. I'll rework them as per Nicolas' suggestions and resend.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"68477","messageId":"47B10A9D.7000702@nrlssc.navy.mil","threadId":"11972","inReplyTo":"47B01FE7.8010207@op5.se","subject":"[PATCH 1/2] pack-objects: Allow setting the #threads equal to #cpus automatically","fromName":"Brandon Casey","fromEmail":"casey@nrlssc.navy.mil","sentAt":"2008-02-12T02:55:25Z","receivedAt":"2008-02-12T02:55:25Z","isPatch":true,"sender":{"key":"drafnel@gmail.com","avatar":"https://avatars.githubusercontent.com/u/921167?v=4"},"body":"Allow pack.threads config option and --threads command line option to\naccept '0' as an argument and set the number of created threads equal\nto the number of online processors in this case.\n\nSigned-off-by: Brandon Casey <casey@nrlssc.navy.mil>\n---\n\n\nI was preparing this patch when I saw your email. I looked up your\nthe old email you were talking about. Your function is better since\nit is cross platform.\n\nWhen you redo your patch, you may want to adopt one aspect of this\none. I used a setting of zero to imply \"set number of threads to\nnumber of cpus\". This allows the user to specifically set pack.threads\nin the config file to zero with the above mentioned meaning, or to\noverride a setting in the config file from the command line with\n--threads=0. This is rather than having to delete the option from the\nconfig file.\n\n-brandon\n\n\n builtin-pack-objects.c |   22 ++++++++++++++++++----\n 1 files changed, 18 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex 692a761..5c55c11 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -1852,11 +1852,11 @@ static int git_pack_config(const char *k, const char *v)\n \t}\n \tif (!strcmp(k, \"pack.threads\")) {\n \t\tdelta_search_threads = git_config_int(k, v);\n-\t\tif (delta_search_threads < 1)\n+\t\tif (delta_search_threads < 0)\n \t\t\tdie(\"invalid number of threads specified (%d)\",\n \t\t\t    delta_search_threads);\n #ifndef THREADED_DELTA_SEARCH\n-\t\tif (delta_search_threads > 1)\n+\t\tif (delta_search_threads != 1)\n \t\t\twarning(\"no threads support, ignoring %s\", k);\n #endif\n \t\treturn 0;\n@@ -2121,10 +2121,10 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\tif (!prefixcmp(arg, \"--threads=\")) {\n \t\t\tchar *end;\n \t\t\tdelta_search_threads = strtoul(arg+10, &end, 0);\n-\t\t\tif (!arg[10] || *end || delta_search_threads < 1)\n+\t\t\tif (!arg[10] || *end || delta_search_threads < 0)\n \t\t\t\tusage(pack_usage);\n #ifndef THREADED_DELTA_SEARCH\n-\t\t\tif (delta_search_threads > 1)\n+\t\t\tif (delta_search_threads != 1)\n \t\t\t\twarning(\"no threads support, \"\n \t\t\t\t\t\"ignoring %s\", arg);\n #endif\n@@ -2234,6 +2234,20 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \tif (!pack_to_stdout && thin)\n \t\tdie(\"--thin cannot be used to build an indexable pack.\");\n \n+#ifdef THREADED_DELTA_SEARCH\n+\tif (!delta_search_threads) {\n+#if defined _SC_NPROCESSORS_ONLN\n+\t\tdelta_search_threads = sysconf(_SC_NPROCESSORS_ONLN);\n+#elif defined _SC_NPROC_ONLN\n+\t\tdelta_search_threads = sysconf(_SC_NPROC_ONLN);\n+#endif\n+\t\tif (delta_search_threads == -1)\n+\t\t\tperror(\"Could not detect number of processors\");\n+\t\tif (delta_search_threads <= 0)\n+\t\t\tdelta_search_threads = 1;\n+\t}\n+#endif\n+\n \tprepare_packed_git();\n \n \tif (progress)\n-- \n1.5.4.1.40.gdb90\n"},{"id":"68479","messageId":"47B10B7D.2030702@nrlssc.navy.mil","threadId":"11972","inReplyTo":"1202784078-23700-1-git-send-email-casey@nrlssc.navy.mil","subject":"[PATCH 2/2] pack-objects: Default to zero threads, meaning auto-assign to #cpus","fromName":"Brandon Casey","fromEmail":"casey@nrlssc.navy.mil","sentAt":"2008-02-12T02:59:09Z","receivedAt":"2008-02-12T02:59:09Z","isPatch":true,"sender":{"key":"drafnel@gmail.com","avatar":"https://avatars.githubusercontent.com/u/921167?v=4"},"body":"Additionally, update some tests for which the multi-threaded result\ndiffers from the single-threaded result and the single-threaded\nresult is expected.\n\nSigned-off-by: Brandon Casey <casey@nrlssc.navy.mil>\n---\n\n\nTwo of the tests in t5300-pack-object.sh failed when multiple\nthreads were used. My fix was to set --threads=1 for all pack-objects\ncalls. I didn't look into it any further than that. All other tests\npassed.\n\n-brandon\n\n\n builtin-pack-objects.c |    2 +-\n t/t5300-pack-object.sh |    8 ++++----\n 2 files changed, 5 insertions(+), 5 deletions(-)\n\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex 5c55c11..743de52 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -70,7 +70,7 @@ static int progress = 1;\n static int window = 10;\n static uint32_t pack_size_limit, pack_size_limit_cfg;\n static int depth = 50;\n-static int delta_search_threads = 1;\n+static int delta_search_threads = 0;\n static int pack_to_stdout;\n static int num_preferred_base;\n static struct progress *progress_state;\ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex cd3c149..16ee940 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -35,7 +35,7 @@ test_expect_success \\\n \n test_expect_success \\\n     'pack without delta' \\\n-    'packname_1=$(git pack-objects --window=0 test-1 <obj-list)'\n+    'packname_1=$(git pack-objects --threads=1 --window=0 test-1 <obj-list)'\n \n rm -fr .git2\n mkdir .git2\n@@ -66,7 +66,7 @@ cd \"$TRASH\"\n test_expect_success \\\n     'pack with REF_DELTA' \\\n     'pwd &&\n-     packname_2=$(git pack-objects test-2 <obj-list)'\n+     packname_2=$(git pack-objects --threads=1 test-2 <obj-list)'\n \n rm -fr .git2\n mkdir .git2\n@@ -96,7 +96,7 @@ cd \"$TRASH\"\n test_expect_success \\\n     'pack with OFS_DELTA' \\\n     'pwd &&\n-     packname_3=$(git pack-objects --delta-base-offset test-3 <obj-list)'\n+     packname_3=$(git pack-objects --threads=1 --delta-base-offset test-3 <obj-list)'\n \n rm -fr .git2\n mkdir .git2\n@@ -271,7 +271,7 @@ test_expect_success \\\n test_expect_success \\\n     'honor pack.packSizeLimit' \\\n     'git config pack.packSizeLimit 200 &&\n-     packname_4=$(git pack-objects test-4 <obj-list) &&\n+     packname_4=$(git pack-objects --threads=1 test-4 <obj-list) &&\n      test 3 = $(ls test-4-*.pack | wc -l)'\n \n test_done\n-- \n1.5.4.1.40.gdb90\n"},{"id":"68483","messageId":"alpine.LFD.1.00.0802112352380.2732@xanadu.home","threadId":"11972","inReplyTo":"47B10B7D.2030702@nrlssc.navy.mil","subject":"Re: [PATCH 2/2] pack-objects: Default to zero threads, meaning auto-assign to #cpus","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-12T04:57:52Z","receivedAt":"2008-02-12T04:57:52Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 11 Feb 2008, Brandon Casey wrote:\n\n> Additionally, update some tests for which the multi-threaded result\n> differs from the single-threaded result and the single-threaded\n> result is expected.\n> \n> Signed-off-by: Brandon Casey <casey@nrlssc.navy.mil>\n\nI think the first patch is OK, but having the _default_ be \nmulti-threaded is going a bit too far.  IMHO you should document the \nmeaning of the value 0, and compile with thread support whenever Posix \nthreads are available, but activating threads should be done explicitly.\n\n\nNicolas\n"},{"id":"68485","messageId":"47B1343C.8070405@op5.se","threadId":"11972","inReplyTo":"47B10A9D.7000702@nrlssc.navy.mil","subject":"Re: [PATCH 1/2] pack-objects: Allow setting the #threads equal to #cpus automatically","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2008-02-12T05:53:00Z","receivedAt":"2008-02-12T05:53:00Z","isPatch":true,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Brandon Casey wrote:\n> Allow pack.threads config option and --threads command line option to\n> accept '0' as an argument and set the number of created threads equal\n> to the number of online processors in this case.\n> \n> Signed-off-by: Brandon Casey <casey@nrlssc.navy.mil>\n> ---\n> \n> \n> I was preparing this patch when I saw your email. I looked up your\n> the old email you were talking about. Your function is better since\n> it is cross platform.\n> \n> When you redo your patch, you may want to adopt one aspect of this\n> one. I used a setting of zero to imply \"set number of threads to\n> number of cpus\". This allows the user to specifically set pack.threads\n> in the config file to zero with the above mentioned meaning, or to\n> override a setting in the config file from the command line with\n> --threads=0. This is rather than having to delete the option from the\n> config file.\n> \n\nThat make sense. Perhaps even go so far as to allow 'auto' as a\nkeyword would be nifty.\n\n>  \n> +#ifdef THREADED_DELTA_SEARCH\n> +\tif (!delta_search_threads) {\n> +#if defined _SC_NPROCESSORS_ONLN\n> +\t\tdelta_search_threads = sysconf(_SC_NPROCESSORS_ONLN);\n> +#elif defined _SC_NPROC_ONLN\n> +\t\tdelta_search_threads = sysconf(_SC_NPROC_ONLN);\n> +#endif\n> +\t\tif (delta_search_threads == -1)\n> +\t\t\tperror(\"Could not detect number of processors\");\n> +\t\tif (delta_search_threads <= 0)\n> +\t\t\tdelta_search_threads = 1;\n> +\t}\n> +#endif\n> +\n\nBut this is not so good. For one thing you've dropped windows support\nentirely. The last comment on my own patch was that get_num_active_cpus()\nshould live in a file of its own. You've taken one step back from that\nand not even kept it in its own function.\n\nI think perhaps it's time to introduce thread-compat.[ch] to deal with\nthread-related cross-platform things like this.\n\nI'll recook my patch and send it in a few minutes, using your suggestions\nand Nicolas combined.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"68545","messageId":"alpine.LSU.1.00.0802122036150.3870@racer.site","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802101845320.11591@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-12T20:37:58Z","receivedAt":"2008-02-12T20:37:58Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 10 Feb 2008, Johannes Schindelin wrote:\n\n> $ /usr/bin/time git repack -a -d -f --window=150 --depth=150\n> Counting objects: 2477715, done.\n> Compressing objects:  19% (481551/2411764)\n> Compressing objects:  19% (482333/2411764)\n> fatal: Out of memory, malloc failed411764)\n> Command exited with non-zero status 1\n> 7118.37user 54.15system 2:01:44elapsed 98%CPU (0avgtext+0avgdata \n> 0maxresident)k\n> 0inputs+0outputs (29834major+17122977minor)pagefaults 0swaps\n\nI made the window much smaller (512 megabyte), and it still runs, after 27 \nhours:\n\nCompressing objects:  20% (484132/2411764)\n\nHowever, it seems that it only worked on about 4000 objects in the last \n20(!) hours.  So, the first 19% were relatively quick.  The next percent \nnot at all.\n\nWill keep you posted,\nDscho\n"},{"id":"68547","messageId":"alpine.LFD.1.00.0802121553120.2732@xanadu.home","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802122036150.3870@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-12T21:05:42Z","receivedAt":"2008-02-12T21:05:42Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 12 Feb 2008, Johannes Schindelin wrote:\n\n> Hi,\n> \n> On Sun, 10 Feb 2008, Johannes Schindelin wrote:\n> \n> > $ /usr/bin/time git repack -a -d -f --window=150 --depth=150\n> > Counting objects: 2477715, done.\n> > Compressing objects:  19% (481551/2411764)\n> > Compressing objects:  19% (482333/2411764)\n> > fatal: Out of memory, malloc failed411764)\n> > Command exited with non-zero status 1\n> > 7118.37user 54.15system 2:01:44elapsed 98%CPU (0avgtext+0avgdata \n> > 0maxresident)k\n> > 0inputs+0outputs (29834major+17122977minor)pagefaults 0swaps\n> \n> I made the window much smaller (512 megabyte), and it still runs, after 27 \n> hours:\n> \n> Compressing objects:  20% (484132/2411764)\n> \n> However, it seems that it only worked on about 4000 objects in the last \n> 20(!) hours.  So, the first 19% were relatively quick.  The next percent \n> not at all.\n\nYeah... this repo is really a pain to repack.  I have access to a \n8-processor machine with 8GB of ram and all my repack attempts so far \nwere killed after using too much memory, despite the window memory \nlimit.  Those were threaded repack attempts, so the first 98% was really \nquick, like less than 15 minutes, but then all threads converged on this \nsmall fraction of the object space which appears to cause problems.  \nAnd then I'm presuming I ran into the same threaded memory fragmentation \nissue.  Might be worth attaching gdb to it and extract a sample of the \nobject SHA1's populating the delta window when the slowdown occurs to \nsee what they actually are...\n\nI'm attempting a single-threaded repack now.\n\n\nNicolas\n"},{"id":"68548","messageId":"alpine.LFD.1.00.0802121303450.2920@woody.linux-foundation.org","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802122036150.3870@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-02-12T21:08:30Z","receivedAt":"2008-02-12T21:08:30Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 12 Feb 2008, Johannes Schindelin wrote:\n> \n> I made the window much smaller (512 megabyte), and it still runs, after 27 \n> hours:\n\nI'd suggest making the memory window smaller yet. \n\n512MB is a *big* amount of memory, if you fill it up, and end up using an \nO(n**2) algorithm on the objects within the window (which it is: the \nrepacking algorithm is O(n) in _total_ objects, but the constant part is \nbasically O(winsize^2).\n\nI'd suggest that a reasonable window memory limit is around just a few \nmegabytes (eg 4MB to maybe 64MB). If you have \"normal\" source files, \nyou're still going to be limited by the window _count_ size (assume normal \nsource files are in the few tens of kB), and for those occasional large \nfiles, you'd better hope that the sort heursistics are good enough.\n\n\t\t\tLinus\n"},{"id":"68550","messageId":"9e4733910802121325p7ce6b58axae71f698f76dbfd2@mail.gmail.com","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802122036150.3870@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2008-02-12T21:25:58Z","receivedAt":"2008-02-12T21:25:58Z","isPatch":true,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 2/12/08, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> Hi,\n>\n> On Sun, 10 Feb 2008, Johannes Schindelin wrote:\n>\n> > $ /usr/bin/time git repack -a -d -f --window=150 --depth=150\n> > Counting objects: 2477715, done.\n> > Compressing objects:  19% (481551/2411764)\n> > Compressing objects:  19% (482333/2411764)\n> > fatal: Out of memory, malloc failed411764)\n> > Command exited with non-zero status 1\n> > 7118.37user 54.15system 2:01:44elapsed 98%CPU (0avgtext+0avgdata\n> > 0maxresident)k\n> > 0inputs+0outputs (29834major+17122977minor)pagefaults 0swaps\n>\n> I made the window much smaller (512 megabyte), and it still runs, after 27\n> hours:\n>\n> Compressing objects:  20% (484132/2411764)\n>\n> However, it seems that it only worked on about 4000 objects in the last\n> 20(!) hours.\n\nI found that out with gcc. 95% went down in no time and the last 5%\ntook two hours. The 5% that got stuck were chains with 2000+ entries.\n\nThe neat thing about the multithread code is that it will keep\nsplitting the work load. That lets all of the easy deltas finish and\nnot get stuck behind the problem objects.\n\nWith quad core on gcc one core would get stuck on the problem objects.\nThe other three would finish their list and start splitting the\nproblem list. This effectively sorts the problems to the end of the\nwork load. By printing the object hash out as they are completed you\ncan easily identify the problem objects. If I recall right on gcc the\nproblem was a configure file that had 2000 entries in its delta chain.\nThat one delta chain took over an hour to process.\n\nCould there be an N squared type problem when 2000 entry delta chains\nare encountered? Maybe something that just isn't noticeable when\ndepth/window=50. Has testing been done with really long object chains\nto make sure that only the minimal amount of work is being done? It\nseems like something is breaking down when the chain length exceeds\nthe window size.\n\nSo, the first 19% were relatively quick.  The next percent\n> not at all.\n>\n> Will keep you posted,\n> Dscho\n>\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"68552","messageId":"9e4733910802121336x42055baawf2b8f3714e2a1eb4@mail.gmail.com","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802121303450.2920@woody.linux-foundation.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2008-02-12T21:36:32Z","receivedAt":"2008-02-12T21:36:32Z","isPatch":true,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"How many diffs should it take to compress a 2000 delta chain with\nwindow/depth=250?\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"68557","messageId":"alpine.LFD.1.00.0802121356330.2920@woody.linux-foundation.org","threadId":"11972","inReplyTo":"9e4733910802121336x42055baawf2b8f3714e2a1eb4@mail.gmail.com","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-02-12T21:59:47Z","receivedAt":"2008-02-12T21:59:47Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 12 Feb 2008, Jon Smirl wrote:\n>\n> How many diffs should it take to compress a 2000 delta chain with\n> window/depth=250?\n\nThere's no fixed answer. We do various culling heurstics to avoid actually \ngenerating a diff at all if it looks unlikely to succeed etc. But in \ngeneral, the way the window works is that \n (a) we only need to generate the _unpacked_ object once\n (b) we compare each object to the \"window-1\" preceding objects, which is \n     how I got the O(windowsize^2) \n (c) but then that \"compare\" relatively seldom involves actually \n     generating a whole diff!\n\nSo the answer is: in _theory_ each object may be compared to \n(windowsize-1) other objects, but in practice it's much less than that.\n\n\t\t\tLinus\n"},{"id":"68561","messageId":"alpine.LFD.1.00.0802121412520.2920@woody.linux-foundation.org","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802121356330.2920@woody.linux-foundation.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-02-12T22:25:20Z","receivedAt":"2008-02-12T22:25:20Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 12 Feb 2008, Linus Torvalds wrote:\n>\n>  (b) we compare each object to the \"window-1\" preceding objects, which is \n>      how I got the O(windowsize^2) \n\nThat's not really true, of course. But my (broken and inexact) logic is \nthat we get one cost multiplier from the number of objects, and one from \nthe size of the objects.\n\nSo *if* we have the situation of not limiting the window size, we \nbasically have a big slowdown from raising the window in number of \nobjects: not only do we get a slowdown from comparing more objects, we \nspend relatively more time comparing the *large* ones to begin with and \nhaving more of them just makes it even more skewed - when we hit a series \nof big blocks, the window will also contain more big blocks, so it kind of \na double whammy.\n\nBut I don't think calling it O(windowsize^2) is really correct. It's still \nO(windowsize), it's just that the purely \"number-of-object\" thing doesn't \naccount for big objects being much more expensive to diff. So you really \nwant to make the *memory* limiter the big one, because that's the one that \nactually approximates how much time you end up spending.\n\nSo ignore that O(n^2) blather. It's not correct. What _is_ correct is that \nwe want to aggressively limit memory size, because CPU cost goes up \nlinearly not just with number of objects, but also super-linearly with \nsize of the object (\"super-linear\" due to bad cache behavior and in worst \ncase due to paging).\n\n\t\t\tLinus\n"},{"id":"68563","messageId":"9e4733910802121443g7b3f5977s1f7dfeb9ba6abaab@mail.gmail.com","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802121412520.2920@woody.linux-foundation.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2008-02-12T22:43:37Z","receivedAt":"2008-02-12T22:43:37Z","isPatch":true,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 2/12/08, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Tue, 12 Feb 2008, Linus Torvalds wrote:\n> >\n> >  (b) we compare each object to the \"window-1\" preceding objects, which is\n> >      how I got the O(windowsize^2)\n>\n> That's not really true, of course. But my (broken and inexact) logic is\n> that we get one cost multiplier from the number of objects, and one from\n> the size of the objects.\n>\n> So *if* we have the situation of not limiting the window size, we\n> basically have a big slowdown from raising the window in number of\n> objects: not only do we get a slowdown from comparing more objects, we\n> spend relatively more time comparing the *large* ones to begin with and\n> having more of them just makes it even more skewed - when we hit a series\n> of big blocks, the window will also contain more big blocks, so it kind of\n> a double whammy.\n>\n> But I don't think calling it O(windowsize^2) is really correct. It's still\n> O(windowsize), it's just that the purely \"number-of-object\" thing doesn't\n> account for big objects being much more expensive to diff. So you really\n> want to make the *memory* limiter the big one, because that's the one that\n> actually approximates how much time you end up spending.\n>\n> So ignore that O(n^2) blather. It's not correct. What _is_ correct is that\n> we want to aggressively limit memory size, because CPU cost goes up\n> linearly not just with number of objects, but also super-linearly with\n> size of the object (\"super-linear\" due to bad cache behavior and in worst\n> case due to paging).\n\n\nIn the gcc case I wasn't running out memory. I believe was CPU bound\nfor an hour processing a single object chain with 2000 entries. That\nsure doesn't feel like O(windowsize).\n\nMaybe someone playing the the OO repo can stick in an appropriate\nprintf and see how many diffs are really being done just to make sure\nthey match what we think the number should be.\n\n\n>\n>                         Linus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"68569","messageId":"alpine.LFD.1.00.0802121533510.2920@woody.linux-foundation.org","threadId":"11972","inReplyTo":"9e4733910802121443g7b3f5977s1f7dfeb9ba6abaab@mail.gmail.com","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-02-12T23:39:21Z","receivedAt":"2008-02-12T23:39:21Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 12 Feb 2008, Jon Smirl wrote:\n> \n> In the gcc case I wasn't running out memory. I believe was CPU bound\n> for an hour processing a single object chain with 2000 entries. That\n> sure doesn't feel like O(windowsize).\n\nWell, there's another - and totally unrelated - issue with *pre-existing* \ndelta chains that are very deep.\n\nNamely the fact that since such a deep delta chain will exhaust the \ndelta-cache, you will now have a O(n*chaindepth) behaviour when you unpack \nthe objects (in order to generate the deltas) in the first place!\n\nSo that really has nothing to do with the new window (or delta) depth at \nall, just with the _previous_ window depth.\n\nSee sha1_file.c: MAX_DELTA_CACHE.\n\nIf you have a 2000-deep delta chain, then the delta-cache should be big \nenough that you hit in it regularly without flushing it when you traverse \ndown the chain. So MAX_DELTA_CACHE should generally be at _least_ as much \nas the max delta chain length, which is obviously normally the case \n(default max delta chain length: 10).\n\nWe could probably fairly easily make that MAX_DELTA_CACHE be a config \noption, but right now you have to recompile to test that theory of mine.\n\nOr just limit your delta depth to something much smaller (ie ~100 or so)\n\n\t\tLinus\n"},{"id":"68764","messageId":"alpine.LSU.1.00.0802141917420.30505@racer.site","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802122036150.3870@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-14T19:20:39Z","receivedAt":"2008-02-14T19:20:39Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 12 Feb 2008, Johannes Schindelin wrote:\n\n> On Sun, 10 Feb 2008, Johannes Schindelin wrote:\n> \n> > $ /usr/bin/time git repack -a -d -f --window=150 --depth=150\n> > Counting objects: 2477715, done.\n> > Compressing objects:  19% (481551/2411764)\n> > Compressing objects:  19% (482333/2411764)\n> > fatal: Out of memory, malloc failed411764)\n> > Command exited with non-zero status 1\n> > 7118.37user 54.15system 2:01:44elapsed 98%CPU (0avgtext+0avgdata \n> > 0maxresident)k\n> > 0inputs+0outputs (29834major+17122977minor)pagefaults 0swaps\n> \n> I made the window much smaller (512 megabyte), and it still runs, after 27 \n> hours:\n> \n> Compressing objects:  20% (484132/2411764)\n> \n> However, it seems that it only worked on about 4000 objects in the last \n> 20(!) hours.  So, the first 19% were relatively quick.  The next percent \n> not at all.\n\nFinally!\n\nI updated to newest git+patches (git version 1.5.4.1.1353.g0d5dd), reset \nwindowMemory to 512m and restarted the process:\n\n$ /usr/bin/time git repack -a -d -f --window=250 --depth=250\nCounting objects: 2477715, done.\nCompressing objects: 100% (2411764/2411764), done.\nWriting objects: 100% (2477715/2477715), done.\nTotal 2477715 (delta 1876242), reused 0 (delta 0)\n21733.55user 175.32system 6:10:37elapsed 98%CPU (0avgtext+0avgdata \n0maxresident)k\n0inputs+0outputs (81921major+63880453minor)pagefaults 0swaps\n\nA little over 6 hours, with one core (of the four available).  Not bad, I \nsay.\n\nThe result is:\n\n$ ls -la objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack\n-rwxrwxrwx 1 root root 1638490531 2008-02-14 17:51 \nobjects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack\n\n1.6G looks much better than 2.4G, wouldn't you say?  Jan, if you want it, \nplease give me a place to upload it to.\n\nCiao,\nDscho\n"},{"id":"68765","messageId":"47B4996B.4000900@nrlssc.navy.mil","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802101445430.2732@xanadu.home","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Brandon Casey","fromEmail":"casey@nrlssc.navy.mil","sentAt":"2008-02-14T19:41:31Z","receivedAt":"2008-02-14T19:41:31Z","isPatch":true,"sender":{"key":"drafnel@gmail.com","avatar":"https://avatars.githubusercontent.com/u/921167?v=4"},"body":"Nicolas Pitre wrote:\n> On Sun, 10 Feb 2008, Johannes Schindelin wrote:\n> \n>> I tried that:\n>>\n>> $ git config pack.deltaCacheLimit 1\n>> $ git config pack.deltaCacheSize 1\n>> $ git config pack.windowMemory 2g\n> \n> This has nothing to do with repacking memory usage, but even tighter \n> packs can be obtained with:\n> \n> \tgit config repack.usedeltabaseoffset true\n> \n> This is not the default yet.\n\nI have successfully repacked this repo a few times on a 2.1GHz system with 16G.\n\nThe smallest attained pack was about 1.45G (1556569742B).\n\nThis run took about 7 hours 26 min.\n\nI ran: git repack -a -d -f --window=250 --depth=250\n\nHere are the relevent config entries:\n[pack]\n        threads = 1\n        compression = 9\n[repack]\n        usedeltabaseoffset = true\n\n\nOther runs:\n\n\n* Same as above, but with default compression:\n\n\tpack size: 1560624388\n\ttime: 7 hours 11 min\n\n\tNot much difference in time or size.\n\n\n* Multi threaded (250m window)\n[pack]\n        threads = 4\n        windowmemory = 250m\n        compression = 9\n[repack]\n        usedeltabaseoffset = true\n\n\tpack size: 1767405703\n\ttime: 3 hours\n\n\tFirst >99% took 50min. Last 10000 objects took 2hours.\n\n* Multi threaded (500m window)\n[pack]\n        threads = 4\n        windowmemory = 500m\n        compression = 9\n[repack]\n        usedeltabaseoffset = true\n\n\tpack size: 1640820903\n\ttime: forgot to time, but between 3-4 hours based on file time\n\n\tI just received Dscho's email, this is interesting to compare\n\twith his single threaded result of 1638490531. I wonder if he\n\tused deltabaseoffset? I think his machine is a little faster\n\tthan this one. So using 4 threads finished twice as fast and\n\tproduced a similar pack size. Actually, the difference could\n\tjust be the compression setting.\n\n* Deeper (git repack -a -d -f --window=250 --depth=500)\n[pack]\n        threads = 1\n        compression = 9\n[repack]\n        usedeltabaseoffset = true\n\n\tpack size: 1578263745\n\ttime: 7 hours 58 min\n\n\tLarger pack compared to --depth=250.\n\n-brandon\n"},{"id":"68766","messageId":"alpine.LSU.1.00.0802141957530.30505@racer.site","threadId":"11972","inReplyTo":"47B4996B.4000900@nrlssc.navy.mil","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-14T19:58:26Z","receivedAt":"2008-02-14T19:58:26Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 14 Feb 2008, Brandon Casey wrote:\n\n> \tI just received Dscho's email, this is interesting to compare\n> \twith his single threaded result of 1638490531. I wonder if he\n> \tused deltabaseoffset?\n\nNope.  Wanted it to be as compatible as possible.\n\nCiao,\nDscho\n"},{"id":"68767","messageId":"m3y79nb8xk.fsf@localhost.localdomain","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802141917420.30505@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-14T20:05:35Z","receivedAt":"2008-02-14T20:05:35Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Finally!\n> \n> I updated to newest git+patches (git version 1.5.4.1.1353.g0d5dd), reset \n> windowMemory to 512m and restarted the process:\n \n> A little over 6 hours, with one core (of the four available).  Not bad, I \n> say.\n> \n> The result is:\n> \n> $ ls -la objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack\n> -rwxrwxrwx 1 root root 1638490531 2008-02-14 17:51 \n> objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack\n> \n> 1.6G looks much better than 2.4G, wouldn't you say?  Jan, if you want it, \n> please give me a place to upload it to.\n\nBrandon Casey wrote:\n\n> I have successfully repacked this repo a few times on a 2.1GHz\n> system with 16G.\n> \n> The smallest attained pack was about 1.45G (1556569742B).\n\nDo you perchance know why OOo needs so large pack? Perhaps you could\ntry running contrib/stats/packinfo.pl on this pack to examine it to\nget to know what takes most space.\n\nWhat is the size of checkout, by the way?\n\nHmmm... I wonder if packv4 would help...\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"68768","messageId":"alpine.LFD.1.00.0802141501420.2732@xanadu.home","threadId":"11972","inReplyTo":"47B4996B.4000900@nrlssc.navy.mil","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-14T20:11:47Z","receivedAt":"2008-02-14T20:11:47Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 14 Feb 2008, Brandon Casey wrote:\n\n> I have successfully repacked this repo a few times on a 2.1GHz system \n> with 16G.\n> \n> The smallest attained pack was about 1.45G (1556569742B).\n> \n[...]\n> \n> * Multi threaded (250m window)\n> [pack]\n>         threads = 4\n>         windowmemory = 250m\n>         compression = 9\n> [repack]\n>         usedeltabaseoffset = true\n> \n> \tpack size: 1767405703\n> \ttime: 3 hours\n> \n> \tFirst >99% took 50min. Last 10000 objects took 2hours.\n\nRight.  That's because the algorithm to distribute the load between \nthreads ends up stealing work from other threads whenever a thread is \ndone with its own share.  So the easy objects are quickly done with by a \nfew threads until they all converge onto the hard ones.  In the non \nthreaded case, the slow down ocurs around 12%.\n\nIt looks like those hard objects are huge binary blobs.  If they could \nbe removed from the repository entirely and regenerated as needed \ninstead of being carried around then I expect the repository size would \nfall below the 500MB mark.\n\n\nNicolas\n"},{"id":"68770","messageId":"alpine.LFD.1.00.0802141512580.2732@xanadu.home","threadId":"11972","inReplyTo":"m3y79nb8xk.fsf@localhost.localdomain","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-14T20:16:35Z","receivedAt":"2008-02-14T20:16:35Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 14 Feb 2008, Jakub Narebski wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > The result is:\n> > \n> > $ ls -la objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack\n> > -rwxrwxrwx 1 root root 1638490531 2008-02-14 17:51 \n> > objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack\n> > \n> > 1.6G looks much better than 2.4G, wouldn't you say?  Jan, if you want it, \n> > please give me a place to upload it to.\n> \n> Brandon Casey wrote:\n> \n> > I have successfully repacked this repo a few times on a 2.1GHz\n> > system with 16G.\n> > \n> > The smallest attained pack was about 1.45G (1556569742B).\n> \n> Hmmm... I wonder if packv4 would help...\n\nNo.  Well, it would help a bit, maybe in the 10-20% range, but nothing \nas significant as going from 2.6G to 1.5G, or like in the GCC case, from \n1.3G to 230M.\n\n\nNicolas\n"},{"id":"68772","messageId":"alpine.LSU.1.00.0802142054080.30505@racer.site","threadId":"11972","inReplyTo":"m3y79nb8xk.fsf@localhost.localdomain","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-14T21:04:39Z","receivedAt":"2008-02-14T21:04:39Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 14 Feb 2008, Jakub Narebski wrote:\n\n> Do you perchance know why OOo needs so large pack?\n\nNo.\n\n> Perhaps you could try running contrib/stats/packinfo.pl on this pack to \n> examine it to get to know what takes most space.\n\n$ ~/git/contrib/stats/packinfo.pl < \\\nobjects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack 2>&1 | \\\ntee packinfo.txt\nIllegal division by zero at /home/imaging/git/contrib/stats/packinfo.pl \nline 141, <STDIN> line 6330855.\n\n> What is the size of checkout, by the way?\n\nI work on a bare repository, but:\n\n$ git archive origin/master | wc -c\n2010060800\n\nOr more precisely:\n\n$ echo $(($(git ls-tree -l -r origin/master | sed -n 's/^[^ ]* [^ ]* [^ ]*  \n*\\([0-9]*\\).*$/\\1/p' | tr '\\012' +)0))\n1947839459\n\nSo yes, we still have the crown of the _whole_ repository being _smaller_ \nthan a single checkout.\n\nYeah!\n\n> Hmmm... I wonder if packv4 would help...\n\nI could imagine that it does, what with it being so much better with \nstrings.  But it would come at a price of performance, I guess, as the \nstring table should be well over 64k.\n\nCiao,\nDscho\n"},{"id":"68773","messageId":"47B4ADBD.9030409@nrlssc.navy.mil","threadId":"11972","inReplyTo":"m3y79nb8xk.fsf@localhost.localdomain","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Brandon Casey","fromEmail":"casey@nrlssc.navy.mil","sentAt":"2008-02-14T21:08:13Z","receivedAt":"2008-02-14T21:08:13Z","isPatch":true,"sender":{"key":"drafnel@gmail.com","avatar":"https://avatars.githubusercontent.com/u/921167?v=4"},"body":"Jakub Narebski wrote:\n> Brandon Casey wrote:\n\n>> The smallest attained pack was about 1.45G (1556569742B).\n> \n> Do you perchance know why OOo needs so large pack? Perhaps you could\n> try running contrib/stats/packinfo.pl on this pack to examine it to\n> get to know what takes most space.\n\nEarlier in this thread Sean did some analysis and found lots of large\nobjects, and he mentioned that he sent a listing to Jan for inspection.\nI haven't heard anything more.\n\n> What is the size of checkout, by the way?\n\n2.4G\n\n-brandon\n"},{"id":"68774","messageId":"200802142300.01615.jnareb@gmail.com","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802142054080.30505@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-14T21:59:59Z","receivedAt":"2008-02-14T21:59:59Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Johannes Schindelin wrote:\n> On Thu, 14 Feb 2008, Jakub Narebski wrote:\n\n>> Perhaps you could try running contrib/stats/packinfo.pl on this pack to \n>> examine it to get to know what takes most space.\n> \n> $ ~/git/contrib/stats/packinfo.pl < \\\n> objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack 2>&1 | \\\n> tee packinfo.txt\n> Illegal division by zero at /home/imaging/git/contrib/stats/packinfo.pl \n> line 141, <STDIN> line 6330855.\n\nErrr... sorry, I should have been more explicit. What I meant here\nis the result of\n\n$ git verify-pack -v <packfile> | \\\n  ~/git/contrib/stats/packinfo.pl\n\n\n>> What is the size of checkout, by the way?\n> \n> I work on a bare repository, but:\n> \n> $ git archive origin/master | wc -c\n> 2010060800\n> \n> Or more precisely:\n> \n> $ echo $(($(git ls-tree -l -r origin/master | sed -n 's/^[^ ]* [^ ]* [^ ]*  \n> *\\([0-9]*\\).*$/\\1/p' | tr '\\012' +)0))\n> 1947839459\n> \n> So yes, we still have the crown of the _whole_ repository being _smaller_ \n> than a single checkout.\n> \n> Yeah!\n\n\nBrandon Casey wrote:\n> Jakub Narebski wrote:\n>> \n>> What is the size of checkout, by the way?\n> \n> 2.4G\n\nThat's huuuuge tree. Compared to that 1.6G (or 1.4G) packfile doesn't\nlook large.\n\nI wonder if proper subdivision into submodules (which should encourage\nbetter code by the way, see TAOUP), and perhaps partial checkouts\nwouldn't be better solution than lazy clone. But it is nice to have\nlong discussed about feature, even if at RFC stage, but with some code.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"68785","messageId":"alpine.LSU.1.00.0802142334480.30505@racer.site","threadId":"11972","inReplyTo":"200802142300.01615.jnareb@gmail.com","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-14T23:38:24Z","receivedAt":"2008-02-14T23:38:24Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 14 Feb 2008, Jakub Narebski wrote:\n\n> Johannes Schindelin wrote:\n> > On Thu, 14 Feb 2008, Jakub Narebski wrote:\n> \n> >> Perhaps you could try running contrib/stats/packinfo.pl on this pack \n> >> to examine it to get to know what takes most space.\n> > \n> > $ ~/git/contrib/stats/packinfo.pl < \\\n> > objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack 2>&1 | \\\n> > tee packinfo.txt\n> > Illegal division by zero at /home/imaging/git/contrib/stats/packinfo.pl \n> > line 141, <STDIN> line 6330855.\n> \n> Errr... sorry, I should have been more explicit. What I meant here is \n> the result of\n> \n> $ git verify-pack -v <packfile> | \\\n>   ~/git/contrib/stats/packinfo.pl\n\nHeh.  I was too lazy to look up the usage, so I just did what I thought \nwould make sense...\n\nSo here it goes:\n\n$ git verify-pack -v \nobjects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack | \n~/git/contrib/stats/packinfo.pl | tee packinfo.txt\n      all sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n4748.05 median 232 std_dev 221254.37\n all path sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n4748.05 median 232 std_dev 221254.37\n     tree sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n4748.05 median 232 std_dev 221254.37\ntree path sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n4748.05 median 232 std_dev 221254.37\n         depths: count 2477715 total 70336238 min 0 max 250 mean 28.39 \nmedian 4 std_dev 55.49\n\nSomething in my gut tells me that those four repetitive lines are not \nmeant to look like they do...\n\n> > 2.4G\n>\n> That's huuuuge tree. Compared to that 1.6G (or 1.4G) packfile doesn't \n> look large.\n> \n> I wonder if proper subdivision into submodules (which should encourage \n> better code by the way, see TAOUP), and perhaps partial checkouts \n> wouldn't be better solution than lazy clone. But it is nice to have long \n> discussed about feature, even if at RFC stage, but with some code.\n\nI think partial checkouts are wrong.  If you can work on partial \ncheckouts, chances are that what you work on should be a submodule.\n\nHaving said that, I can understand if some people do not want to have the \nhassle of test^H^H^H^Husing submodules...\n\nCiao,\nDscho\n"},{"id":"68786","messageId":"20080214235129.GU27535@lavos.net","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802142334480.30505@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Brian Downing","fromEmail":"bdowning@lavos.net","sentAt":"2008-02-14T23:51:29Z","receivedAt":"2008-02-14T23:51:29Z","isPatch":true,"sender":{"key":"bdowning@lavos.net","avatar":"https://avatars.githubusercontent.com/u/366426?v=4"},"body":"On Thu, Feb 14, 2008 at 11:38:24PM +0000, Johannes Schindelin wrote:\n> Heh.  I was too lazy to look up the usage, so I just did what I thought \n> would make sense...\n> \n> So here it goes:\n> \n> $ git verify-pack -v \n> objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack | \n> ~/git/contrib/stats/packinfo.pl | tee packinfo.txt\n>       all sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n> 4748.05 median 232 std_dev 221254.37\n>  all path sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n> 4748.05 median 232 std_dev 221254.37\n>      tree sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n> 4748.05 median 232 std_dev 221254.37\n> tree path sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n> 4748.05 median 232 std_dev 221254.37\n>          depths: count 2477715 total 70336238 min 0 max 250 mean 28.39 \n> median 4 std_dev 55.49\n> \n> Something in my gut tells me that those four repetitive lines are not \n> meant to look like they do...\n\nDo you by chance have repack.usedeltabaseoffset turned on?  That has the\nunfortunate side effect of changing the output of verify-pack -v to be\nalmost useless for my packinfo script (specifically, it no longer\nreports the parent SHA1 hash for deltas, and the script is basically all\nabout deltra tree statistics.)  I suppose that should probably be fixed,\nbut I never looked into it.\n\n-bcd\n"},{"id":"68787","messageId":"20080214235747.GV27535@lavos.net","threadId":"11972","inReplyTo":"20080214235129.GU27535@lavos.net","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Brian Downing","fromEmail":"bdowning@lavos.net","sentAt":"2008-02-14T23:57:47Z","receivedAt":"2008-02-14T23:57:47Z","isPatch":true,"sender":{"key":"bdowning@lavos.net","avatar":"https://avatars.githubusercontent.com/u/366426?v=4"},"body":"On Thu, Feb 14, 2008 at 05:51:29PM -0600, Brian Downing wrote:\n> Do you by chance have repack.usedeltabaseoffset turned on?  That has the\n> unfortunate side effect of changing the output of verify-pack -v to be\n> almost useless for my packinfo script (specifically, it no longer\n> reports the parent SHA1 hash for deltas, and the script is basically all\n> about deltra tree statistics.)  I suppose that should probably be fixed,\n> but I never looked into it.\n\nThat being said, the most useful output for figuring out where all the\nspace in the pack is going in my experience is gotten from:\n\ngit-verify-pack -v | packinfo.pl -tree -filenames\n\nThat will produce a huge amount of output, which is basically the tree\nstructure of the delta chains in the file.  If things aren't being\ndeltified together properly, it's usually pretty obvious.\n\nA delta chain in this output looks approximately like this:\n\n#   0   blob 03156f21...     1767     1767 Documentation/git-lost-found.txt @ tags/v1.2.0~142\n#   1    blob f52a9d7f...       10     1777 Documentation/git-lost-found.txt @ tags/v1.5.0-rc1~74\n#   2     blob a8cc5739...       51     1828 Documentation/git-lost+found.txt @ tags/v0.99.9h^0\n#   3      blob 660e90b1...       15     1843 Documentation/git-lost+found.txt @ master~3222^2~2\n#   4       blob 0cb8e3bb...       33     1876 Documentation/git-lost+found.txt @ master~3222^2~3\n#   2     blob e48607f0...      311     2088 Documentation/git-lost-found.txt @ tags/v1.5.2-rc3~4\n#      size: count 6 total 2187 min 10 max 1767 mean 364.50 median 51 std_dev 635.85\n# path size: count 6 total 11179 min 1767 max 2088 mean 1863.17 median 1843 std_dev 107.26\n\n# The first number after the sha1 is the object size, the second\n# number is the path size.  The statistics are across all objects in\n# the previous delta tree.  Obviously they are omitted for trees of\n# one object.\n\n# A path size is the sum of the size of the delta chain, including the\n# base object.  In other words, it's how many bytes need be read to\n# reassemble the file from deltas.\n\nThis is also quite slow, as it runs git-ls-tree -t -r on every commit in\nthe repository to assign file names to blobs.  You can leave out the\n-filenames option to not do this (if you don't care about seeing\nfilenames, that is).\n\n-bcd\n"},{"id":"68789","messageId":"alpine.LSU.1.00.0802150007480.30505@racer.site","threadId":"11972","inReplyTo":"20080214235129.GU27535@lavos.net","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-02-15T00:08:57Z","receivedAt":"2008-02-15T00:08:57Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 14 Feb 2008, Brian Downing wrote:\n\n> On Thu, Feb 14, 2008 at 11:38:24PM +0000, Johannes Schindelin wrote:\n> > Heh.  I was too lazy to look up the usage, so I just did what I \n> > thought would make sense...\n> > \n> > So here it goes:\n> > \n> > $ git verify-pack -v \n> > objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack | \n> > ~/git/contrib/stats/packinfo.pl | tee packinfo.txt\n> >       all sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n> > 4748.05 median 232 std_dev 221254.37\n> >  all path sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n> > 4748.05 median 232 std_dev 221254.37\n> >      tree sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n> > 4748.05 median 232 std_dev 221254.37 tree path sizes: count 601473 \n> > total 2855826280 min 0 max 62173032 mean 4748.05 median 232 std_dev \n> > 221254.37\n> >          depths: count 2477715 total 70336238 min 0 max 250 mean 28.39 \n> > median 4 std_dev 55.49\n> > \n> > Something in my gut tells me that those four repetitive lines are not \n> > meant to look like they do...\n> \n> Do you by chance have repack.usedeltabaseoffset turned on?\n\nOuch.  That must have been a leftover from earlier attempts.  I did not \n_mean_ to keep it, but now that I have a pretty packed repository, I think \nI'll just keep it as-is.\n\nCiao,\nDscho\n"},{"id":"68792","messageId":"200802150207.47095.jnareb@gmail.com","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802142334480.30505@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-15T01:07:46Z","receivedAt":"2008-02-15T01:07:46Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Dnia piątek 15. lutego 2008 00:38, Johannes Schindelin napisał:\n> Hi,\n> \n> On Thu, 14 Feb 2008, Jakub Narebski wrote:\n>\n>> I wonder if proper subdivision into submodules (which should\n>> encourage better code by the way, see TAOUP), and perhaps\n>> _partial checkouts_ wouldn't be better solution than _lazy clone_.\n>> But it is nice to have long discussed about feature, even if at\n>> RFC stage, but with some code. \n> \n> I think partial checkouts are wrong.  If you can work on partial \n> checkouts, chances are that what you work on should be a submodule.\n> \n> Having said that, I can understand if some people do not want to have\n> the hassle of test^H^H^H^Husing submodules...\n\nIMHO there is place for submodules, there is place for partial \ncheckouts, and perhaps there is even place for the combination of two.\n\nFor example while Documentation/ isn't a good candidate for a submodule, \nbecause as you add new feature yuou want to add to documentation, if \nyou change some feature you want to change documentation: there are \nwhole-tree commits which contain changes outside Documentation/.\nNevertheless there are some people (technical writers) which are \ninterested only in Documentation; perhaps only in few files there.\nThey would want to have partial checkout, I guess.\n\nOn the other hand cgit and msysgit use submodules, and I think it is \ngood solution. I wonder if Sourcemage Linux distro uses submodules... \nIn the case of cgit I think having git.git or its clone/fork as \nsubmodule is a good idea, but perhaps even better would be to checkout \nonly part of it: libgit or libgitthin\n\n-- \nJakub Narebski\nPoland\n"},{"id":"68794","messageId":"alpine.LFD.1.00.0802142032030.2732@xanadu.home","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802150007480.30505@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-15T01:41:28Z","receivedAt":"2008-02-15T01:41:28Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 15 Feb 2008, Johannes Schindelin wrote:\n\n> Hi,\n> \n> On Thu, 14 Feb 2008, Brian Downing wrote:\n> \n> > On Thu, Feb 14, 2008 at 11:38:24PM +0000, Johannes Schindelin wrote:\n> > > Heh.  I was too lazy to look up the usage, so I just did what I \n> > > thought would make sense...\n> > > \n> > > So here it goes:\n> > > \n> > > $ git verify-pack -v \n> > > objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack | \n> > > ~/git/contrib/stats/packinfo.pl | tee packinfo.txt\n> > >       all sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n> > > 4748.05 median 232 std_dev 221254.37\n> > >  all path sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n> > > 4748.05 median 232 std_dev 221254.37\n> > >      tree sizes: count 601473 total 2855826280 min 0 max 62173032 mean \n> > > 4748.05 median 232 std_dev 221254.37 tree path sizes: count 601473 \n> > > total 2855826280 min 0 max 62173032 mean 4748.05 median 232 std_dev \n> > > 221254.37\n> > >          depths: count 2477715 total 70336238 min 0 max 250 mean 28.39 \n> > > median 4 std_dev 55.49\n> > > \n> > > Something in my gut tells me that those four repetitive lines are not \n> > > meant to look like they do...\n> > \n> > Do you by chance have repack.usedeltabaseoffset turned on?\n> \n> Ouch.  That must have been a leftover from earlier attempts.  I did not \n> _mean_ to keep it, but now that I have a pretty packed repository, I think \n> I'll just keep it as-is.\n\nI should really come around to fixing packed_object_info_detail() for \nthe OBJ_OFS_DELTA case one day.\n\n\nNicolas\n"},{"id":"68803","messageId":"200802151034.46236.kendy@suse.cz","threadId":"11972","inReplyTo":"alpine.LSU.1.00.0802141917420.30505@racer.site","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2008-02-15T09:34:45Z","receivedAt":"2008-02-15T09:34:45Z","isPatch":true,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Johannes,\n\nOn Thursday 14 of February 2008, Johannes Schindelin wrote:\n\n> The result is:\n>\n> $ ls -la objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack\n> -rwxrwxrwx 1 root root 1638490531 2008-02-14 17:51\n> objects/pack/pack-e4dc6da0a10888ec4345490575efc587b7523b45.pack\n>\n> 1.6G looks much better than 2.4G, wouldn't you say?  Jan, if you want it,\n> please give me a place to upload it to.\n\nThank you!  In the meantime, I happened to produce something similar.  \nUnfortunately even mine was too late for another round of tests to present it \nin our git vs. svn comparison (with todays deadline) - so we just mentioned \nin the report that the tested repository still had reserves [but the numbers \nwere quite nice even with the 2.5G one ;-)].\n\n> ll minimal3.git/objects/pack/\ncelkem 1636608\n-r--r--r-- 1 kendy users   59264432 2008-02-10 15:22 \npack-909b501d3d673f10a66adfefdf8371933e7a6f3e.idx\n-r--r--r-- 1 kendy users 1614968445 2008-02-10 15:22 \npack-909b501d3d673f10a66adfefdf8371933e7a6f3e.pack\n\n> ll minimal4.git/objects/pack/\ncelkem 1644160\n-r--r--r-- 1 kendy users   59264432 2008-02-11 16:09 \npack-909b501d3d673f10a66adfefdf8371933e7a6f3e.idx\n-rw-r--r-- 1 kendy users          0 2008-02-11 16:29 \npack-909b501d3d673f10a66adfefdf8371933e7a6f3e.keep\n-r--r--r-- 1 kendy users 1622697708 2008-02-11 16:09 \npack-909b501d3d673f10a66adfefdf8371933e7a6f3e.pack\n\nThe 'minimal3' case was with '--window=250 --depth=250', 'minimal4' was \nwith '--window=250 --depth=50'\n\nI tried the --depth=50 because I read 'making it too deep affects the \nperformance on the unpacker side' in the man page.  How big the difference \ncould be in practice, please?\n\nRegards,\nJan\n"},{"id":"68806","messageId":"200802151043.21508.kendy@suse.cz","threadId":"11972","inReplyTo":"200802142300.01615.jnareb@gmail.com","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2008-02-15T09:43:20Z","receivedAt":"2008-02-15T09:43:20Z","isPatch":true,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Jakub,\n\nOn Thursday 14 of February 2008, Jakub Narebski wrote:\n\n> >> What is the size of checkout, by the way?\n> >\n> > 2.4G\n>\n> That's huuuuge tree. Compared to that 1.6G (or 1.4G) packfile doesn't\n> look large.\n>\n> I wonder if proper subdivision into submodules (which should encourage\n> better code by the way, see TAOUP), and perhaps partial checkouts\n> wouldn't be better solution than lazy clone. But it is nice to have\n> long discussed about feature, even if at RFC stage, but with some code.\n\nYes, I'd love to see the OOo tree split into several parts, I've already \nproposed a division (http://www.nabble.com/OOo-source-split-td13096065.html), \nbut it'll take some more time I'm afraid :-(\n\nRegards,\nJan\n"},{"id":"68961","messageId":"20080217081841.GS24004@spearce.org","threadId":"11972","inReplyTo":"alpine.LFD.1.00.0802142032030.2732@xanadu.home","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-02-17T08:18:41Z","receivedAt":"2008-02-17T08:18:41Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Nicolas Pitre <nico@cam.org> wrote:\n> \n> I should really come around to fixing packed_object_info_detail() for \n> the OBJ_OFS_DELTA case one day.\n\nPlease don't.\n\nObtaining the SHA-1 of your delta base would require unpacking your\ndelta base and then doing a SHA-1 hash of it.  Or alternatively\ndoing a search through the .idx for the object that starts at the\nrequested OFS.  Either way, its really expensive for a minor detail\nof output in verify-pack.  Something that any script can produce\nwith a simple reverse lookup table.\n\nIts also run after we just spent a hell of a lot of time and disk\nIO trying to verify the packfile.  We slammed through the pack\nonce to do its overall SHA-1, and then god knows how many times as\nwe iterate the objects in pack order, not delta base order, thus\ncausing the delta base cache to become overwhelmed and constantly\nfault out entries.  Pack verification is stupid and slow.  This\nwould make -v even worse.\n\n\nBut if you are going to do that, you may also want to fix the\n\"*store_size = 0 /* notyet */\" that's like 5 lines above.  :)\n\n\nBTW, why does this return const char* from typename(type) instead\nof just returning the enum object_type and letting the caller do\ntypename() if they want it?  Most of our other code that returns\ntypes returns the enum, not the string.  :-\\\n\n-- \nShawn.\n"},{"id":"68963","messageId":"7vk5l42brt.fsf@gitster.siamese.dyndns.org","threadId":"11972","inReplyTo":"20080217081841.GS24004@spearce.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-02-17T09:05:42Z","receivedAt":"2008-02-17T09:05:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> BTW, why does this return const char* from typename(type) instead\n> of just returning the enum object_type and letting the caller do\n> typename() if they want it?  Most of our other code that returns\n> types returns the enum, not the string.  :-\\\n\nIt just was not converted from the old string interface.  I\nthought you are old enough to remember ;-)\n"},{"id":"69019","messageId":"alpine.LFD.1.00.0802171335470.2732@xanadu.home","threadId":"11972","inReplyTo":"20080217081841.GS24004@spearce.org","subject":"Re: [PATCH] RFC: git lazy clone proof-of-concept","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-02-17T18:44:01Z","receivedAt":"2008-02-17T18:44:01Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 17 Feb 2008, Shawn O. Pearce wrote:\n\n> Nicolas Pitre <nico@cam.org> wrote:\n> > \n> > I should really come around to fixing packed_object_info_detail() for \n> > the OBJ_OFS_DELTA case one day.\n> \n> Please don't.\n> \n> Obtaining the SHA-1 of your delta base would require unpacking your\n> delta base and then doing a SHA-1 hash of it.  Or alternatively\n> doing a search through the .idx for the object that starts at the\n> requested OFS.\n\nI intended to use the pack index of course.  And the code already exists \nin pack-objects as find_packed_object().\n\n> Either way, its really expensive for a minor detail\n> of output in verify-pack.\n\nNot _that_ expensive actually.  Like I say, in pack-objects we do it all \nthe time.\n\n> But if you are going to do that, you may also want to fix the\n> \"*store_size = 0 /* notyet */\" that's like 5 lines above.  :)\n\nYeah, that's easy too.\n\n\nNicolas\n"}]}