{"thread":{"id":"1034","subject":"CAREFUL! No more delta object support!","startedAt":"2005-06-27T23:58:57Z","lastAt":"2005-06-29T22:24:32Z","messageCount":38,"participants":["Linus Torvalds","Junio C Hamano","Christopher Li","Daniel Barkalow","Jan Harkes","Petr Baudis","Benjamin LaHaise","Matthias Urlichs"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"5322","messageId":"20050627235857.GA21533@64m.dyndns.org","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506271755140.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-06-27T23:58:57Z","receivedAt":"2005-06-27T23:58:57Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"On Mon, Jun 27, 2005 at 06:14:40PM -0700, Linus Torvalds wrote:\n> \n> The reason? The new git understands packed files natively, which ends up \n> being a much bigger win in many many ways.\n\nInteresting. I take a look at your change, it still support delta object\ninside the pack file right? For a second I am wondering you drop the delta\nfeature completely.\n\nChris\n"},{"id":"5311","messageId":"Pine.LNX.4.58.0506271755140.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":null,"subject":"CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-28T01:14:40Z","receivedAt":"2005-06-28T01:14:40Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nSome people may have noticed already (hopefully not the hard way) that the \ncurrent git code doesn't support delta objects lying around in the object \ndirectory any more.\n\nIn other words, if you have delta objects, you need to un-deltify your \nrepository _before_ you upgrade your git binaries, or they won't be able \nto read your objects any more.\n\nThe reason? The new git understands packed files natively, which ends up \nbeing a much bigger win in many many ways.\n\nYou should be very careful about using packed files (since they are a very \nrecent addition), but what you can do to try them out is to do so in a \nseparate repository.\n\nStarting to use a packed repository is very simple indeed, and here's what \nyou need to do for git, for example:\n\nIn your regular \"git\" directory (once you have ypdated your git to a \nrecent version, in particular you need to have the \"csum-file: fix missing \nbuf pointer update\" commit), do:\n\n\tgit-rev-list --objects HEAD | git-pack-objects --window=50 --depth=50 out\n\nwhich will say something like \"Packing 3741 objects\" and result in two new \nfiles a few seconds later:\n\n\ttorvalds@ppc970:~/git> ls -lh out*\n\t-rw-r--r--  1 torvalds torvalds  89K Jun 27 17:59 out.idx\n\t-rw-r--r--  1 torvalds torvalds 1.3M Jun 27 17:59 out.pack\n\nnow, don't do anythign with those files, but instead go and create a \ndirectory somewhere else:\n\n\tcd ~\n\tmkdir packed-git-trial\n\tcd packed-git-trial\n\tgit-init-db\n\nyou have now obviously created a totally empty repository. Now, let's \npopulate that empty repository with _just_ the pack files:\n\n\tmkdir .git/objects/pack\n\tmv ~/git/out.* .git/objects/pack\n\nand then, move over your tags, in particularly the HEAD pointer, with \nsomething like\n\n\tcat ~/git/.git/HEAD > .git/HEAD\n\nand voila, you're done. Try \"gitk\", for example. Or \"git log\".\n\nNow, what's even cooler is how you can just start using this packed tree: \nfeel free to do a test-commit or something, and notice how git starts \npopulating the empty .git/objects/xx/ subdirectories with new objects. But \nit still relies on the pack-file for the old history.\n\nNow, there's still a misfeature there, which is that when you create a new\nobject, it doesn't check whether that object already exists in the\npack-file, so you'll end up with a few recent objects that you really\ndon't need (notably tree objects), and we'll fix that eventually. But\nnotice how you started with a 17MB .git/objects/ directory in your\noriginal tree, and you now have just a 1.3MB pack-file and a 90kB index\nfile that replaces all that?\n\nThere are some other issues too, like the fact that \"git-fsck-cache\"  \ndoesn't know about the pack-files yet, so it will complain about missing\nobjects etc. Also, please note that the pack-file _only_ packs the commits\nand the things reachable from them: things like tags (and your references\nin your .git/refs directory) need to be copied over separately.\n\nSo this is all very rough, still, but the basics do actually seem to work\n(ie anything that doesn't look directly at the object files - which is\npretty much all of it except for fsck and the direct-filesystem-access \nthings like \"rsync\" and \"git-local-pull\").\n\nMaybe you might not want to switch over yet, and as mentioned, rsync then\nends up not being a good way to sync (nor git-local-pull), but the\n\"git-http/ssh-pull\" family should hopefully just work.\n\nI've used a packed kernel tree too, so this has gotten _some_ testing even \non really quite big git trees. \n\n\t\t\tLinus\n"},{"id":"5315","messageId":"7vwtofi6jk.fsf@assigned-by-dhcp.cox.net","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506271755140.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-28T02:01:03Z","receivedAt":"2005-06-28T02:01:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> Now, there's still a misfeature there, which is that when you create a new\nLT> object, it doesn't check whether that object already exists in the\nLT> pack-file, so you'll end up with a few recent objects that you really\nLT> don't need (notably tree objects), and we'll fix that eventually.\n\nPatch will be sent separately.\n\nLT> ... Also, please note that the pack-file _only_ packs the commits\nLT> and the things reachable from them ...\n\nShouldn't feeding \"git-rev-list --object\" output plus\nhandcrafted list of objects in 2.6.11 tree object to\ngit-pack-objects just work???\n\nLT> Maybe you might not want to switch over yet, and as mentioned, rsync then\nLT> ends up not being a good way to sync (nor git-local-pull), but the\nLT> \"git-http/ssh-pull\" family should hopefully just work.\n\nNo.  The pull protocol Dan did expects to throw compressed\nrepresentation around on the wire (which is valid if you assume\nuncompressed transfer) and does not use read-sha1-file --\nwrite-sha1-file pair, so all three do not work.\n"},{"id":"5316","messageId":"7vr7eni6fy.fsf_-_@assigned-by-dhcp.cox.net","threadId":"1034","inReplyTo":"7vwtofi6jk.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH] Skip writing out sha1 files for objects in packed git.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-28T02:03:13Z","receivedAt":"2005-06-28T02:03:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Now, there's still a misfeature there, which is that when you\ncreate a new object, it doesn't check whether that object\nalready exists in the pack-file, so you'll end up with a few\nrecent objects that you really don't need (notably tree\nobjects), and this patch fixes it.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n apply.c          |    2 +-\n cache.h          |    2 +-\n commit-tree.c    |    2 +-\n convert-cache.c  |    6 +++---\n mktag.c          |    2 +-\n sha1_file.c      |   44 ++++++++++++++++++++++++++++++--------------\n unpack-objects.c |    4 ++--\n update-cache.c   |    2 +-\n write-tree.c     |    2 +-\n 9 files changed, 41 insertions(+), 25 deletions(-)\n\nf4f76b275cdabc038bcb4f3c7ca0d443638df88d\ndiff --git a/apply.c b/apply.c\n--- a/apply.c\n+++ b/apply.c\n@@ -1221,7 +1221,7 @@ static void add_index_file(const char *p\n \tif (lstat(path, &st) < 0)\n \t\tdie(\"unable to stat newly created file %s\", path);\n \tfill_stat_cache_info(ce, &st);\n-\tif (write_sha1_file(buf, size, \"blob\", ce->sha1) < 0)\n+\tif (write_sha1_file(buf, size, \"blob\", ce->sha1, 0) < 0)\n \t\tdie(\"unable to create backing store for newly created file %s\", path);\n \tif (add_cache_entry(ce, ADD_CACHE_OK_TO_ADD) < 0)\n \t\tdie(\"unable to add cache entry for %s\", path);\ndiff --git a/cache.h b/cache.h\n--- a/cache.h\n+++ b/cache.h\n@@ -165,7 +165,7 @@ extern int parse_sha1_header(char *hdr, \n extern int sha1_object_info(const unsigned char *, char *, unsigned long *);\n extern void * unpack_sha1_file(void *map, unsigned long mapsize, char *type, unsigned long *size);\n extern void * read_sha1_file(const unsigned char *sha1, char *type, unsigned long *size);\n-extern int write_sha1_file(void *buf, unsigned long len, const char *type, unsigned char *return_sha1);\n+extern int write_sha1_file(void *buf, unsigned long len, const char *type, unsigned char *return_sha1, int do_expand);\n \n extern int check_sha1_signature(const unsigned char *sha1, void *buf, unsigned long size, const char *type);\n \ndiff --git a/commit-tree.c b/commit-tree.c\n--- a/commit-tree.c\n+++ b/commit-tree.c\n@@ -191,7 +191,7 @@ int main(int argc, char **argv)\n \twhile (fgets(comment, sizeof(comment), stdin) != NULL)\n \t\tadd_buffer(&buffer, &size, \"%s\", comment);\n \n-\twrite_sha1_file(buffer, size, \"commit\", commit_sha1);\n+\twrite_sha1_file(buffer, size, \"commit\", commit_sha1, 0);\n \tprintf(\"%s\\n\", sha1_to_hex(commit_sha1));\n \treturn 0;\n }\ndiff --git a/convert-cache.c b/convert-cache.c\n--- a/convert-cache.c\n+++ b/convert-cache.c\n@@ -111,7 +111,7 @@ static int write_subdirectory(void *buff\n \t\tbuffer += len;\n \t}\n \n-\twrite_sha1_file(new, newlen, \"tree\", result_sha1);\n+\twrite_sha1_file(new, newlen, \"tree\", result_sha1, 0);\n \tfree(new);\n \treturn used;\n }\n@@ -251,7 +251,7 @@ static void convert_date(void *buffer, u\n \tmemcpy(new + newlen, buffer, size);\n \tnewlen += size;\n \n-\twrite_sha1_file(new, newlen, \"commit\", result_sha1);\n+\twrite_sha1_file(new, newlen, \"commit\", result_sha1, 0);\n \tfree(new);\t\n }\n \n@@ -286,7 +286,7 @@ static struct entry * convert_entry(unsi\n \tmemcpy(buffer, data, size);\n \t\n \tif (!strcmp(type, \"blob\")) {\n-\t\twrite_sha1_file(buffer, size, \"blob\", entry->new_sha1);\n+\t\twrite_sha1_file(buffer, size, \"blob\", entry->new_sha1, 0);\n \t} else if (!strcmp(type, \"tree\"))\n \t\tconvert_tree(buffer, size, entry->new_sha1);\n \telse if (!strcmp(type, \"commit\"))\ndiff --git a/mktag.c b/mktag.c\n--- a/mktag.c\n+++ b/mktag.c\n@@ -123,7 +123,7 @@ int main(int argc, char **argv)\n \tif (verify_tag(buffer, size) < 0)\n \t\tdie(\"invalid tag signature file\");\n \n-\tif (write_sha1_file(buffer, size, \"tag\", result_sha1) < 0)\n+\tif (write_sha1_file(buffer, size, \"tag\", result_sha1, 0) < 0)\n \t\tdie(\"unable to write tag file\");\n \tprintf(\"%s\\n\", sha1_to_hex(result_sha1));\n \treturn 0;\ndiff --git a/sha1_file.c b/sha1_file.c\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -891,31 +891,47 @@ void *read_object_with_reference(const u\n \t}\n }\n \n-int write_sha1_file(void *buf, unsigned long len, const char *type, unsigned char *returnsha1)\n+static char *write_sha1_file_prepare(void *buf,\n+\t\t\t\t     unsigned long len,\n+\t\t\t\t     const char *type,\n+\t\t\t\t     unsigned char *sha1,\n+\t\t\t\t     unsigned char *hdr,\n+\t\t\t\t     int *hdrlen)\n {\n-\tint size;\n-\tunsigned char *compressed;\n-\tz_stream stream;\n-\tunsigned char sha1[20];\n \tSHA_CTX c;\n-\tchar *filename;\n-\tstatic char tmpfile[PATH_MAX];\n-\tunsigned char hdr[50];\n-\tint fd, hdrlen, ret;\n \n \t/* Generate the header */\n-\thdrlen = sprintf((char *)hdr, \"%s %lu\", type, len)+1;\n+\t*hdrlen = sprintf((char *)hdr, \"%s %lu\", type, len)+1;\n \n \t/* Sha1.. */\n \tSHA1_Init(&c);\n-\tSHA1_Update(&c, hdr, hdrlen);\n+\tSHA1_Update(&c, hdr, *hdrlen);\n \tSHA1_Update(&c, buf, len);\n \tSHA1_Final(sha1, &c);\n \n+\treturn sha1_file_name(sha1);\n+}\n+\n+int write_sha1_file(void *buf, unsigned long len, const char *type,\n+\t\t    unsigned char *returnsha1, int do_expand)\n+{\n+\tint size;\n+\tunsigned char *compressed;\n+\tz_stream stream;\n+\tunsigned char sha1[20];\n+\tchar *filename;\n+\tstatic char tmpfile[PATH_MAX];\n+\tunsigned char hdr[50];\n+\tint fd, hdrlen, ret;\n+\n+\t/* Normally if we have it in the pack then we do not bother writing\n+\t * it out into .git/objects/??/?{38} file.\n+\t */\n+\tfilename = write_sha1_file_prepare(buf, len, type, sha1, hdr, &hdrlen);\n \tif (returnsha1)\n \t\tmemcpy(returnsha1, sha1, 20);\n-\n-\tfilename = sha1_file_name(sha1);\n+\tif (!do_expand && has_sha1_file(sha1))\n+\t\treturn 0;\n \tfd = open(filename, O_RDONLY);\n \tif (fd >= 0) {\n \t\t/*\n@@ -1082,7 +1098,7 @@ int index_fd(unsigned char *sha1, int fd\n \tif ((int)(long)buf == -1)\n \t\treturn -1;\n \n-\tret = write_sha1_file(buf, size, \"blob\", sha1);\n+\tret = write_sha1_file(buf, size, \"blob\", sha1, 0);\n \tif (size)\n \t\tmunmap(buf, size);\n \treturn ret;\ndiff --git a/unpack-objects.c b/unpack-objects.c\n--- a/unpack-objects.c\n+++ b/unpack-objects.c\n@@ -126,7 +126,7 @@ static int unpack_non_delta_entry(struct\n \tcase 'B': type_s = \"blob\"; break;\n \tdefault: goto err_finish;\n \t}\n-\tif (write_sha1_file(buffer, size, type_s, sha1) < 0)\n+\tif (write_sha1_file(buffer, size, type_s, sha1, 1) < 0)\n \t\tdie(\"failed to write %s (%s)\",\n \t\t    sha1_to_hex(entry->sha1), type_s);\n \tprintf(\"%s %s\\n\", sha1_to_hex(sha1), type_s);\n@@ -223,7 +223,7 @@ static int unpack_delta_entry(struct pac\n \t\tdie(\"failed to apply delta\");\n \tfree(delta_data);\n \n-\tif (write_sha1_file(result, result_size, type, sha1) < 0)\n+\tif (write_sha1_file(result, result_size, type, sha1, 1) < 0)\n \t\tdie(\"failed to write %s (%s)\",\n \t\t    sha1_to_hex(entry->sha1), type);\n \tfree(result);\ndiff --git a/update-cache.c b/update-cache.c\n--- a/update-cache.c\n+++ b/update-cache.c\n@@ -77,7 +77,7 @@ static int add_file_to_cache(char *path)\n \t\t\tfree(target);\n \t\t\treturn -1;\n \t\t}\n-\t\tif (write_sha1_file(target, st.st_size, \"blob\", ce->sha1))\n+\t\tif (write_sha1_file(target, st.st_size, \"blob\", ce->sha1, 0))\n \t\t\treturn -1;\n \t\tfree(target);\n \t\tbreak;\ndiff --git a/write-tree.c b/write-tree.c\n--- a/write-tree.c\n+++ b/write-tree.c\n@@ -76,7 +76,7 @@ static int write_tree(struct cache_entry\n \t\tnr++;\n \t}\n \n-\twrite_sha1_file(buffer, offset, \"tree\", returnsha1);\n+\twrite_sha1_file(buffer, offset, \"tree\", returnsha1, 0);\n \tfree(buffer);\n \treturn nr;\n }\n------------\n"},{"id":"5317","messageId":"Pine.LNX.4.58.0506271910390.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":"7vwtofi6jk.fsf@assigned-by-dhcp.cox.net","subject":"Re: CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-28T02:13:08Z","receivedAt":"2005-06-28T02:13:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 27 Jun 2005, Junio C Hamano wrote:\n> \n> LT> ... Also, please note that the pack-file _only_ packs the commits\n> LT> and the things reachable from them ...\n> \n> Shouldn't feeding \"git-rev-list --object\" output plus\n> handcrafted list of objects in 2.6.11 tree object to\n> git-pack-objects just work???\n\nYou could do that. And yes, we can add support for \"tag\" objects too \n(which the packing doesn't do at all right now. So this is not a \n\"fundamental\" problem, it's just a practical one right now.\n\n> > [..  git-ssh-pull hopefully working ..]\n>\n> No.  The pull protocol Dan did expects to throw compressed\n> representation around on the wire (which is valid if you assume\n> uncompressed transfer) and does not use read-sha1-file --\n> write-sha1-file pair, so all three do not work.\n\nFair enough. I'd prefer for the pull/push to push object packs around \nanyway, so there's some more work there..\n\n\t\tLinus\n"},{"id":"5318","messageId":"7vslz3yzwf.fsf@assigned-by-dhcp.cox.net","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506271910390.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-28T02:32:32Z","receivedAt":"2005-06-28T02:32:32Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> Fair enough. I'd prefer for the pull/push to push object packs around \nLT> anyway, so there's some more work there..\n\nYes, I'd prefer that too.\n\nBy the way, you broke t/t0000 with the last commit.  Now an\nempty GIT_OBJECT_DIRECTORY has 257 subdirectories.\n"},{"id":"5319","messageId":"7vmzpbyzon.fsf_-_@assigned-by-dhcp.cox.net","threadId":"1034","inReplyTo":"7vslz3yzwf.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH] Adjust to git-init-db creating $GIT_OBJECT_DIRECTORY/pack","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-28T02:37:12Z","receivedAt":"2005-06-28T02:37:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Some tests expected the directory not to exist by default.\nUpdated git-init-db prepares it properly so adjust tests to\nmatch that behaviour.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n t/t0000-basic.sh       |    6 +++---\n t/t5300-pack-object.sh |    1 -\n 2 files changed, 3 insertions(+), 4 deletions(-)\n\nde500ab0379e4db18d1511cbe91ace106eee7830\ndiff --git a/t/t0000-basic.sh b/t/t0000-basic.sh\n--- a/t/t0000-basic.sh\n+++ b/t/t0000-basic.sh\n@@ -28,11 +28,11 @@ test_expect_success \\\n     '.git/objects should be empty after git-init-db in an empty repo.' \\\n     'cmp -s /dev/null should-be-empty' \n \n-# also it should have 256 subdirectories.  257 is counting \"objects\"\n+# also it should have 257 subdirectories.  258 is counting \"objects\"\n find .git/objects -type d -print >full-of-directories\n test_expect_success \\\n-    '.git/objects should have 256 subdirectories.' \\\n-    'test $(wc -l < full-of-directories) = 257'\n+    '.git/objects should have 257 subdirectories.' \\\n+    'test $(wc -l < full-of-directories) = 258'\n \n ################################################################\n # Basics of the basics\ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -99,7 +99,6 @@ test_expect_success \\\n     'GIT_OBJECT_DIRECTORY=.git2/objects &&\n      export GIT_OBJECT_DIRECTORY &&\n      git-init-db &&\n-     mkdir .git2/objects/pack &&\n      cp test-1.pack test-1.idx .git2/objects/pack && {\n \t git-diff-tree --root -p $commit &&\n \t while read object\n------------\n"},{"id":"5320","messageId":"Pine.LNX.4.58.0506271935260.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":"7vr7eni6fy.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Skip writing out sha1 files for objects in packed git.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-28T02:43:57Z","receivedAt":"2005-06-28T02:43:57Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 27 Jun 2005, Junio C Hamano wrote:\n>\n> Now, there's still a misfeature there, which is that when you\n> create a new object, it doesn't check whether that object\n> already exists in the pack-file, so you'll end up with a few\n> recent objects that you really don't need (notably tree\n> objects), and this patch fixes it.\n> \n> Signed-off-by: Junio C Hamano <junkio@cox.net>\n\nActually, I don't think that \"do_expand\" flag should exist.\n\nIf we want to expand a packed file and really write the objects to the \n.git/objects directories, we should just not have that packed file in the \n.git/objects/pack directory.\n\nAnd if we have a pack-file in .git/objects/ that already has the object, \nthat may not be the _same_ pack-file that we're expanding at all, so if \nthat pack file already has the object, then not writing it out is actually \nthe right thing to do.\n\nThat will also simplify your patch a bit. I'll fix it up.\n\n\t\tLinus\n"},{"id":"5321","messageId":"Pine.LNX.4.58.0506271946140.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":"7vslz3yzwf.fsf@assigned-by-dhcp.cox.net","subject":"Re: CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-28T02:48:51Z","receivedAt":"2005-06-28T02:48:51Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 27 Jun 2005, Junio C Hamano wrote:\n> \n> By the way, you broke t/t0000 with the last commit.  Now an\n> empty GIT_OBJECT_DIRECTORY has 257 subdirectories.\n\nYup, I noticed that. Fix pushed out (along with another one that was \nfailing because it wanted to create the \"pack\" directory itself, and was \nunhappy when it already existed).\n\n\t\tLinus\n"},{"id":"5323","messageId":"Pine.LNX.4.58.0506272016420.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":"20050627235857.GA21533@64m.dyndns.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-28T03:30:22Z","receivedAt":"2005-06-28T03:30:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 27 Jun 2005, Christopher Li wrote:\n> On Mon, Jun 27, 2005 at 06:14:40PM -0700, Linus Torvalds wrote:\n> > \n> > The reason? The new git understands packed files natively, which ends up \n> > being a much bigger win in many many ways.\n> \n> Interesting. I take a look at your change, it still support delta object\n> inside the pack file right? For a second I am wondering you drop the delta\n> feature completely.\n\nDeltas do exist inside pack-files, yes. They just don't exist as \nindependent objects any more, so you can never get into the situation that \nyou find a delta but you don't find the delta it points to.\n\nBecause in the pack-files, there are only deltas _within_ a pack-file. You \ncan't have a delta that points to outside the pack.\n\nThis means that pack-files with few objects will inevitably be larger than\nthey could otherwise be (ie you can never have a pack file that _only_\ncontains deltas to the outside world), but it's just incredibly reassuring \nto me that a pack-file is always self-sufficient. \n\nSo when/if we start using pack-files for doing \"git pull\" etc, the \npack-file won't actually help pack things for small updates: small updates \nwill probably contain the whole changed file, unless the update has \nseveral changes to the same file (which is not unusual, of course), in \nwhich case it will only contain one version and then deltas from that.\n\nBut the savings get increasingly bigger the more history we have. That's\nalso why the packed git archive is about 1/14th of the size of the fully\nunpacked disk usage of the git project, but a packed kernel archive \"only\"  \nachieves a packing rate of 1/5th of the fully unpacked kernel archive. The\ngit archive is all history, while the kernel archive just \"appears\", and\n2/3 of the files have only one single version and thus don't delta-\ncompress at all.\n\n(Another reason is probably that the kernel has bigger files, which means\nthat it thus has relatively less loss in filesystem block padding).\n\nBut not having any outside deltas not only makes me feel safer, it also\nmeans that you can fully validate a pack archive consistency without even\nknowing what project it is from - you can check the SHA1 results of every\nfile in the pack against the index of the pack, and check that the SHA1's\nof the pack files themselves are valid. Again, this is just a data\n_consistency_ check, of course - it means that you can validate that it\ndownloaded fine, and that you don't have disk corruption, but it doesn't\nmean that the data isn't evil and nasty and buggy ;)\n\n\t\t\tLinus\n"},{"id":"5324","messageId":"7vekanyx33.fsf@assigned-by-dhcp.cox.net","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506271935260.19755@ppc970.osdl.org","subject":"Re: [PATCH] Skip writing out sha1 files for objects in packed git.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-28T03:33:20Z","receivedAt":"2005-06-28T03:33:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> If we want to expand a packed file and really write the objects to the \nLT> .git/objects directories, we should just not have that packed file in the \nLT> .git/objects/pack directory.\n\nWhat I was aiming for was this:\n\n (1) Introduce an interface to sha1_file.c that lets you say\n     \"use this file as one of the packs, although it is not\n     under .git/objects/pack\";\n\n (2) Introduce another interface to sha1_file.c that lets you\n     enumerate the index entries for a given pack file.\n\n (3) Remove the unpacking logic from unpack-object.c; instead\n     call the above interfaces to register the pack and\n     enumerate entries, and call read_sha1_file() followed by\n     write_sha1_file() with do_expand repeatedly.\n\nHowever, the infrastructure (1) and (2) may end up being a\nspecial case only to support unpack-object (and removing the\ncode duplication for unpacking), in which case what you suggest\nwould make more sense.\n\nLT> And if we have a pack-file in .git/objects/ that already has\nLT> the object, that may not be the _same_ pack-file that we're\nLT> expanding at all, so if that pack file already has the\nLT> object, then not writing it out is actually the right thing\nLT> to do.\n\nThis I have to think about a bit.\n"},{"id":"5325","messageId":"Pine.LNX.4.21.0506280049090.30848-100000@iabervon.org","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506271910390.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-28T05:09:15Z","receivedAt":"2005-06-28T05:09:15Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Mon, 27 Jun 2005, Linus Torvalds wrote:\n\n> > > [..  git-ssh-pull hopefully working ..]\n> >\n> > No.  The pull protocol Dan did expects to throw compressed\n> > representation around on the wire (which is valid if you assume\n> > uncompressed transfer) and does not use read-sha1-file --\n> > write-sha1-file pair, so all three do not work.\n> \n> Fair enough. I'd prefer for the pull/push to push object packs around \n> anyway, so there's some more work there..\n\nIt shouldn't be hard to add; the main issue is determining when\ntransfering a pack file is a good idea, because it probably doesn't make\nsense to transfer a pack file just because the source side has an object\nthat the target side wants in that pack. (If you pull from someone who\npacked up the whole history of everything, which you already have, into a\nfile with one new commit, you'd be sad to get the huge thing; you really\nwant a little custom (or just limited) pack file.)\n\nThe ideal thing is probably to pick up some tricks from Mercurial in\nfiguring out what needs to be transferred, and have the source side write\na pack file directly to the connection, which the target side would then\nsave directly. I never worked out exactly what those tricks were, though.\n\nThe next trick would be to put something in place of cleverly-chosen\nobjects to specify what pack file they're in, so that the HTTP client\ncould find things from a packed repository. (Or we could just have an\noption to unpack post-transfer.)\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5331","messageId":"7vslz2x3vg.fsf@assigned-by-dhcp.cox.net","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506271755140.19755@ppc970.osdl.org","subject":"[PATCH] Adjust fsck-cache to packed GIT and alternate object pool.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-28T08:49:39Z","receivedAt":"2005-06-28T08:49:39Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> There are some other issues too, like the fact that \"git-fsck-cache\"  \nLT> doesn't know about the pack-files yet, so it will complain about missing\nLT> objects etc.\n\nAnd here is a patch to fix it.  It is interesting to know that\nthe same problem existed for a long time in a different form and\nnobody has complained: GIT_ALTERNATE_OBJECT_DIRECTORIES.\n\nMaybe the alternate object pool mechanism is not so widely used\nand probably not very useful for everyday use.  I donno.\n\n------------\nThe fsck-cache complains if objects referred to by files in\n.git/refs/ or objects stored in files under .git/objects/??/ are\nnot found as stand-alone SHA1 files (i.e. found in alternate\nobject pools GIT_ALTERNATE_OBJECT_DIRECTORIES or packed archives\nstored under .git/objects/pack).\n\nAlthough this is a good semantics to maintain consistency of a\nsingle .git/objects directory as a self contained set of\nobjects, it sometimes is useful to consider it is OK as long as\nthese \"outside\" objects are available.\n\nThis commit introduces a new flag, --standalone, to\ngit-fsck-cache.  When it is not specified, connectivity checks\nand .git/refs pointer checks are taught that it is OK when\nexpected objects do not exist under .git/objects/?? hierarchy\nbut are available from an packed archive or in an alternate\nobject pool.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n fsck-cache.c |   20 ++++++++++++++++----\n 1 files changed, 16 insertions(+), 4 deletions(-)\n\nea4255429bb0b4b760ba2fe327f5806d8d24d8a6\ndiff --git a/fsck-cache.c b/fsck-cache.c\n--- a/fsck-cache.c\n+++ b/fsck-cache.c\n@@ -12,6 +12,7 @@\n static int show_root = 0;\n static int show_tags = 0;\n static int show_unreachable = 0;\n+static int standalone = 0;\n static int keep_cache_objects = 0; \n static unsigned char head_sha1[20];\n \n@@ -25,13 +26,17 @@ static void check_connectivity(void)\n \t\tstruct object_list *refs;\n \n \t\tif (!obj->parsed) {\n-\t\t\tprintf(\"missing %s %s\\n\",\n-\t\t\t       obj->type, sha1_to_hex(obj->sha1));\n+\t\t\tif (!standalone && has_sha1_file(obj->sha1))\n+\t\t\t\t; /* it is in pack */\n+\t\t\telse\n+\t\t\t\tprintf(\"missing %s %s\\n\",\n+\t\t\t\t       obj->type, sha1_to_hex(obj->sha1));\n \t\t\tcontinue;\n \t\t}\n \n \t\tfor (refs = obj->refs; refs; refs = refs->next) {\n-\t\t\tif (refs->item->parsed)\n+\t\t\tif (refs->item->parsed ||\n+\t\t\t    (!standalone && has_sha1_file(refs->item->sha1)))\n \t\t\t\tcontinue;\n \t\t\tprintf(\"broken link from %7s %s\\n\",\n \t\t\t       obj->type, sha1_to_hex(obj->sha1));\n@@ -315,8 +320,11 @@ static int read_sha1_reference(const cha\n \t\treturn -1;\n \n \tobj = lookup_object(sha1);\n-\tif (!obj)\n+\tif (!obj) {\n+\t\tif (!standalone && has_sha1_file(sha1))\n+\t\t\treturn 0; /* it is in pack */\n \t\treturn error(\"%s: invalid sha1 pointer %.40s\", path, hexname);\n+\t}\n \n \tobj->used = 1;\n \tmark_reachable(obj, REACHABLE);\n@@ -390,6 +398,10 @@ int main(int argc, char **argv)\n \t\t\tkeep_cache_objects = 1;\n \t\t\tcontinue;\n \t\t}\n+\t\tif (!strcmp(arg, \"--standalone\")) {\n+\t\t\tstandalone = 1;\n+\t\t\tcontinue;\n+\t\t}\n \t\tif (*arg == '-')\n \t\t\tusage(\"git-fsck-cache [--tags] [[--unreachable] [--cache] <head-sha1>*]\");\n \t}\n------------\n"},{"id":"5332","messageId":"7vekamvmxj.fsf@assigned-by-dhcp.cox.net","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506272016420.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-28T09:40:56Z","receivedAt":"2005-06-28T09:40:56Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> But the savings get increasingly bigger the more history we have. That's\nLT> also why the packed git archive is about 1/14th of the size of the fully\nLT> unpacked disk usage of the git project,...\n\nGIT archive may be an odd-ball because the project itself is so\nsmall, but a fair comparison should include the disk usage of\n256 fan-out directories.  Counting them, empty .git/objects/\nwith the 1.4MB packed archive and 90KB index file ends up being\nsomewhere around 2.4MB on my machine, compared with 17MB for the\ntraditional one.\n\nStill a good space reduction.  Good job!\n\nI am now dreaming if we someday would enhance the mechanism with\nappend-only updates to the *.pack files with complete rewrite of\nthe *.idx files, and get rid of files under .git/objects totally.\n\nThis would make things reasonably friendly to rsync.  The kernel\npack has around 60M pack with 1.1M index, so everyday use would\ninvolve incremental updates to the pack [*1*] and full download\nof the index file.\n\n[Footnote]\n\n*1* Presumably many objects are deltified against older objects\nwhich is suboptimal.  Most likely the newer objects are accessed\nfar more often and they are what we would want to keep in full\nnot as delta.  So even with this scheme we would want to have\nweekly repacking.  Interestingly enough, pack-objects gets the\nobjects via usual read_sha1_file() interface so it can produce a\nnew pack from an existing pack.\n"},{"id":"5334","messageId":"20050628103852.GB21533@64m.dyndns.org","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506272016420.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-06-28T10:38:52Z","receivedAt":"2005-06-28T10:38:52Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"That is all nice improvement to address the space usage issue.\n\nShould people just run repacking once a while or is it automaticly\nadd new object to the pack file?\n\nChris\n\n\nOn Mon, Jun 27, 2005 at 08:30:22PM -0700, Linus Torvalds wrote:\n> \n> Deltas do exist inside pack-files, yes. They just don't exist as \n> independent objects any more, so you can never get into the situation that \n> you find a delta but you don't find the delta it points to.\n> \n> Because in the pack-files, there are only deltas _within_ a pack-file. You \n> can't have a delta that points to outside the pack.\n> \n> This means that pack-files with few objects will inevitably be larger than\n> they could otherwise be (ie you can never have a pack file that _only_\n> contains deltas to the outside world), but it's just incredibly reassuring \n> to me that a pack-file is always self-sufficient. \n> \n> So when/if we start using pack-files for doing \"git pull\" etc, the \n> pack-file won't actually help pack things for small updates: small updates \n> will probably contain the whole changed file, unless the update has \n> several changes to the same file (which is not unusual, of course), in \n> which case it will only contain one version and then deltas from that.\n> \n> But the savings get increasingly bigger the more history we have. That's\n> also why the packed git archive is about 1/14th of the size of the fully\n> unpacked disk usage of the git project, but a packed kernel archive \"only\"  \n> achieves a packing rate of 1/5th of the fully unpacked kernel archive. The\n> git archive is all history, while the kernel archive just \"appears\", and\n> 2/3 of the files have only one single version and thus don't delta-\n> compress at all.\n> \n> (Another reason is probably that the kernel has bigger files, which means\n> that it thus has relatively less loss in filesystem block padding).\n> \n> But not having any outside deltas not only makes me feel safer, it also\n> means that you can fully validate a pack archive consistency without even\n> knowing what project it is from - you can check the SHA1 results of every\n> file in the pack against the index of the pack, and check that the SHA1's\n> of the pack files themselves are valid. Again, this is just a data\n> _consistency_ check, of course - it means that you can validate that it\n> downloaded fine, and that you don't have disk corruption, but it doesn't\n> mean that the data isn't evil and nasty and buggy ;)\n> \n> \t\t\tLinus\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"5335","messageId":"20050628110625.GC21533@64m.dyndns.org","threadId":"1034","inReplyTo":"7vekamvmxj.fsf@assigned-by-dhcp.cox.net","subject":"Re: CAREFUL! No more delta object support!","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-06-28T11:06:25Z","receivedAt":"2005-06-28T11:06:25Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"On Tue, Jun 28, 2005 at 02:40:56AM -0700, Junio C Hamano wrote:\n> >>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n> Still a good space reduction.  Good job!\n> \n> I am now dreaming if we someday would enhance the mechanism with\n> append-only updates to the *.pack files with complete rewrite of\n> the *.idx files, and get rid of files under .git/objects totally.\n\nNo offense my friend, this has been done. It's name is mercurial.\n\n> This would make things reasonably friendly to rsync.  The kernel\n> pack has around 60M pack with 1.1M index, so everyday use would\n> involve incremental updates to the pack [*1*] and full download\n> of the index file.\n\nIt still have other open issue. Now it would be harder to not sync\nall the heads. If I just want the clean Linus-2.6 tree, I have to\ndig it out from the pack file which mixing with other heads. \n\nYou could host different projects with it's own pack file. That\nwill lost the space saving on co-hosting projects.\n\nSo I am not convince rsync is the way to go in long run. You need\nto have your own network syncing method.\n\n> \n> [Footnote]\n> \n> *1* Presumably many objects are deltified against older objects\n> which is suboptimal.  Most likely the newer objects are accessed\n> far more often and they are what we would want to keep in full\n> not as delta.  So even with this scheme we would want to have\n> weekly repacking.  Interestingly enough, pack-objects gets the\n> objects via usual read_sha1_file() interface so it can produce a\n> new pack from an existing pack.\n\nIt sounds like you are suggesting backward delta. Keeping the\nlatest node in full and using delta to access the old one. It should\nwork but it will lose the append only property.\n\nChris\n"},{"id":"5336","messageId":"20050628144604.GA15792@delft.aura.cs.cmu.edu","threadId":"1034","inReplyTo":"7vekamvmxj.fsf@assigned-by-dhcp.cox.net","subject":"Re: CAREFUL! No more delta object support!","fromName":"Jan Harkes","fromEmail":"jaharkes@cs.cmu.edu","sentAt":"2005-06-28T14:46:04Z","receivedAt":"2005-06-28T14:46:04Z","isPatch":false,"sender":{"key":"jaharkes@cs.cmu.edu","avatar":"https://gravatar.com/avatar/cf95aecd150ca8ef33d6edc337ac4bb9e13aa4246fc3679257d578c7fddc1633?d=mp&s=160"},"body":"On Tue, Jun 28, 2005 at 02:40:56AM -0700, Junio C Hamano wrote:\n> I am now dreaming if we someday would enhance the mechanism with\n> append-only updates to the *.pack files with complete rewrite of\n> the *.idx files, and get rid of files under .git/objects totally.\n\nStop dreaming, please.\n\nThe current separate objects setup might not be space efficient, but it\nhas many other advantages.\n\n- Objects are only written only once, and from then on are only read.\n  This works well on filesystems that provide session semantics, as\n  opposed to unix semantics. And the resulting objects are perfectly\n  cacheable since they are only invalidated if someone ever decides to\n  pack the repository.\n\n- The hierarchy and the way the objects directories are updated works\n  very well in combination with AFS style directory acls. What surprised\n  me was that subdirectories in refs/heads work perfectly with all the\n  core git tools, branchnames simply become 'user/branch'.\n\n- Objects that differ in content have different naming, as a result\n  multiple developers can safely commit into a shared repository without\n  requiring locks. This is also why it is safe to pull from another\n  repository without clobbering your own history. Imagine if you\n  appended some local changes to a packed archive and the next rsync\n  wipes your local commits.\n\nI've been trying to keep an up to date document on how (and why) I use\ngit on Coda. It started pretty much the identical to jgarzik's HOWTO.\nBut it ended up a lot more complicated, to a point where I needed my own\nscripts for just about every action. Until I discovered that the\nalternate objects pool would work well in my environment.\n\n    http://www.coda.cs.cmu.edu/git.html\n\nJan\n"},{"id":"5337","messageId":"20050628145256.GA1275@pasky.ji.cz","threadId":"1034","inReplyTo":"20050628110625.GC21533@64m.dyndns.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-06-28T14:52:56Z","receivedAt":"2005-06-28T14:52:56Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Tue, Jun 28, 2005 at 01:06:25PM CEST, I got a letter\nwhere Christopher Li <git@chrisli.org> told me that...\n> On Tue, Jun 28, 2005 at 02:40:56AM -0700, Junio C Hamano wrote:\n> > >>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n> > Still a good space reduction.  Good job!\n> > \n> > I am now dreaming if we someday would enhance the mechanism with\n> > append-only updates to the *.pack files with complete rewrite of\n> > the *.idx files, and get rid of files under .git/objects totally.\n> \n> No offense my friend, this has been done. It's name is mercurial.\n> \n> > This would make things reasonably friendly to rsync.  The kernel\n> > pack has around 60M pack with 1.1M index, so everyday use would\n> > involve incremental updates to the pack [*1*] and full download\n> > of the index file.\n> \n> It still have other open issue. Now it would be harder to not sync\n> all the heads. If I just want the clean Linus-2.6 tree, I have to\n> dig it out from the pack file which mixing with other heads. \n> \n> You could host different projects with it's own pack file. That\n> will lost the space saving on co-hosting projects.\n> \n> So I am not convince rsync is the way to go in long run. You need\n> to have your own network syncing method.\n\nI think the git-*-pull tools are actually just fine. You will only need\nto have some server-side CGI gadget to frontend the file, but we need\nthat anyway to make the pull reasonably effective.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\n<Espy> be careful, some twit might quote you out of context..\n"},{"id":"5344","messageId":"Pine.LNX.4.58.0506280844190.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":"7vekanyx33.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Skip writing out sha1 files for objects in packed git.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-28T15:45:40Z","receivedAt":"2005-06-28T15:45:40Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 27 Jun 2005, Junio C Hamano wrote:\n> \n> LT> And if we have a pack-file in .git/objects/ that already has\n> LT> the object, that may not be the _same_ pack-file that we're\n> LT> expanding at all, so if that pack file already has the\n> LT> object, then not writing it out is actually the right thing\n> LT> to do.\n> \n> This I have to think about a bit.\n\nThe most trivial example is doing a \"git pull\" of a small pack-file \nupdate.\n\nWe probably don't want to leave it around as a pack-file (we'll re-pack \neverything at some later date, but we also don't want to expand the stuff \nwe already have in our _real_ pack-file).\n\n\t\tLinus\n"},{"id":"5345","messageId":"Pine.LNX.4.58.0506280846100.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":"Pine.LNX.4.21.0506280049090.30848-100000@iabervon.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-28T15:49:29Z","receivedAt":"2005-06-28T15:49:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 28 Jun 2005, Daniel Barkalow wrote:\n> \n> It shouldn't be hard to add; the main issue is determining when\n> transfering a pack file is a good idea, because it probably doesn't make\n> sense to transfer a pack file just because the source side has an object\n> that the target side wants in that pack.\n\nOh, you'd never just transfer the whole big pack-file at all: you'd just \ncreate a new one. And creatign a new one is just a matter of finding the \ncommon parent, and then doing\n\n\tgit-rev-list --objects common..HEAD | git-pack-file .git/tmp-pack\n\nand then you send the result to the other side..\n\n\t\tLinus\n"},{"id":"5346","messageId":"Pine.LNX.4.58.0506280903510.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506280846100.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-28T16:21:28Z","receivedAt":"2005-06-28T16:21:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 28 Jun 2005, Linus Torvalds wrote:\n> \n> Oh, you'd never just transfer the whole big pack-file at all: you'd just \n> create a new one. And creatign a new one is just a matter of finding the \n> common parent, and then doing\n> \n> \tgit-rev-list --objects common..HEAD | git-pack-file .git/tmp-pack\n> \n> and then you send the result to the other side..\n\nTo clarify: this also works with objects that are already in another\npack-file (now that Junio fixed the \"get size of a deltified packed\nentry\"), so you can have any number of unpacked objects in your objects\ndirectory, _and_ a pack-file (or several), and you can generate a new\ntemporary pack-file just for sending somewhere else that contains\narbistrary parts of that (ie a mix of objects that are in your \"main\"\npackfiles and objects that are unpacked).\n\nYou don't have to use \"git-rev-list\" to generate the objects, btw,\ngit-pack-file takes an arbitrary list of object ID's (plus a \"packing\nhint\" in the form of a filename that is not required, but that can help\nthe packing heuristics, and that git-rev-list does provide).\n\nI'll also fix up git-pack-file to be able to pack tag objects (and the\nunpacking to understand them), so that any valid object can be packed. \nRight now it only handles the objects that git-rev-list knows about.\n\n\t\tLinus\n"},{"id":"5347","messageId":"20050628163551.GA29410@kvack.org","threadId":"1034","inReplyTo":"20050628145256.GA1275@pasky.ji.cz","subject":"Re: CAREFUL! No more delta object support!","fromName":"Benjamin LaHaise","fromEmail":"bcrl@kvack.org","sentAt":"2005-06-28T16:35:51Z","receivedAt":"2005-06-28T16:35:51Z","isPatch":false,"sender":{"key":"bcrl@kvack.org","avatar":null},"body":"On Tue, Jun 28, 2005 at 04:52:56PM +0200, Petr Baudis wrote:\n> I think the git-*-pull tools are actually just fine. You will only need\n> to have some server-side CGI gadget to frontend the file, but we need\n> that anyway to make the pull reasonably effective.\n\nNot really -- the use of rsync for the objects fails horribly on slow \nlinks when the project scales in the number of commits.  The rsync \nprotocol has to transfer the names of each file and some information \nabout it, and that information isn't delta compressed.  This is where \nkernel.org is falling over, as well as what makes the kernel tree very \npainful to use over a dialup modem link.\n\n\t\t-ben\n-- \n\"Time is what keeps everything from happening all at once.\" -- John Wheeler\n"},{"id":"5348","messageId":"Pine.LNX.4.58.0506280921480.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":"20050628103852.GB21533@64m.dyndns.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-28T16:45:28Z","receivedAt":"2005-06-28T16:45:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 28 Jun 2005, Christopher Li wrote:\n>\n> That is all nice improvement to address the space usage issue.\n> \n> Should people just run repacking once a while or is it automaticly\n> add new object to the pack file?\n\nWhile adding a new object to a pack file is _possible_ (you add it to the\nend of the pack-file, and re-generate the index file), I would strongly\nsuggest against it for several reasons:\n\n - It's a lot more complex and expensive than just writing a new file.  \n   Much better to make the pack generation be an off-line thing, and make \n   new object creation really cheap.\n\n - it has serious locking issues, and if something goes wrong you are just \n   horribly screwed. This implies, for example, that to be safe you really \n   have to use fsync() etc at every point (and be careful about writing \n   the index), making the update even _more_ expensive. Over NFS you need \n   to be extremely careful to make sure that everybody got the right lock, \n   yadda yadda.\n\n   Packing things off-line just means that _all_ of these problems go \n   away.\n\n - There are operations that want to remove objects (I do that all the \n   time: I do something stupid, and decide to undo it, or I just do a \n   \"git-update-cache\" and notice that I need to do more work so I edit it \n   some more and actually never commit the first version)\n\n   If _adding_ to the file had some serious correctness issues, _removing_ \n   an object from a file is even worse. MUCH worse. Now you don't just \n   have to lock against other people creating new objects, now you have to \n   lock against updates (or totally re-write the whole big file and do an \n   atomic \"rename\").\n\n - it can actually generate worse packing. The current \"offline\" method \n   means that we can pack any version of a file against any other version \n   of a file, and we do. We pick the closest version we can find, and we \n   try to always pack against the bigger one (deletes are smaller deltas, \n   and the biggest one tends to be the latest version, so this not only\n   means that the delta is denser, it also means that the latest version -\n   which is likely to be the biggest and most often used - tends to be\n   non-delta).\n\n   In contrast, updating the pack file means that you always write the \n   latest version as a delta, which means that you're doing things \n   _exactly_ the wrong way around both for performance and size.\n\n - Finally: packing allows us to do optimize for locality. In particular, \n   I write out the pack file in \"recency\" order, ie the top-most objects \n   go first, and in particular, the \"commit\" objects go at the very top of \n   the file. Why? Because it means that the commit objects (which are \n   heavily used for the history generation by pretty much anything, since \n   \"git-rev-list\" will access them) are packed together, and in the right \n   order.\n\n   Again, you can't do that if you do on-line updates as opposed to \n   offline packing.\n\nSo the usage pattern I envision is to pack stuff maybe once a month\n(depending on how much changes, of course), because then you really do get\nthe best of both worlds: the simplicity of individual objects for recent\nwork and the optimal packing and ordering that you can really work on for\nthe longer range case. And your project never grows very big.\n\nBtw, I'm not claiming that my current pack format is \"optimal\" of course.  \nFor example, while I write all objects in recency order, right now that\nmeans that if a recent object has been written as a delta that depends on\nan older one, I actually write the delta first (correct) but I won't write\nthe older object until its recency ordering (wrong).\n\nThat kind of thing is trivial to fix (eventually), but it's an example of\nwhere ordering matters (ie if it's the other way around: the delta is the\nolder object, it's probably better to leave it at the end of the file,\nsince it's probably not going to be accessed much, making the effective\npacking at the head more efficicient). It's also an example of the kinds\nof things we can do exactly because we're doing the packing off-line.\n\n\t\t\tLinus\n"},{"id":"5351","messageId":"Pine.LNX.4.21.0506281251380.30848-100000@iabervon.org","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506280903510.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-28T17:04:19Z","receivedAt":"2005-06-28T17:04:19Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Tue, 28 Jun 2005, Linus Torvalds wrote:\n\n> I'll also fix up git-pack-file to be able to pack tag objects (and the\n> unpacking to understand them), so that any valid object can be packed. \n> Right now it only handles the objects that git-rev-list knows about.\n\nActually, the ideal thing would be to move the packing code into an object\nfile that git-ssh-push can include; that way it can write directly to the\nsocket instead of going through disk, and it can also go from getting the\nremote end's list of common ancestors to having a pack to send without\nneeding to exec a script.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5352","messageId":"Pine.LNX.4.58.0506281019450.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":"Pine.LNX.4.21.0506281251380.30848-100000@iabervon.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-28T17:36:15Z","receivedAt":"2005-06-28T17:36:15Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 28 Jun 2005, Daniel Barkalow wrote:\n> \n> Actually, the ideal thing would be to move the packing code into an object\n> file that git-ssh-push can include; that way it can write directly to the\n> socket instead of going through disk\n\nIt doesn't work very easily that way because the index file (which\ncontains the object list and the offsets into the pack file) cannot be\ncreated until after the pack file has been created (and we don't want to\nevaluate that one in memory, since it can be quite big).\n\nNow, what we could do is to stream out the pack file first to stdout, and\nwrite the index file afterwards. But since we don't know how big the pack\nfile will be when we start packing, and the pack-file can contain\nbasically arbitrary patterns, that requires that the receiver actually \nparse the pack-file as it comes in.\n\nThe format of the pack-file is a fairly trivial data stream of\n\n - rinse and repeat for each object:\n\n     - one character of type of file (C, T, B, G, D for \"commit\", \"tree\", \n       \"blob\", \"tag\" or \"delta\" respectively)\n\n     - four bytes of network-order unpacked data length\n\n     - [ if delta: 20 bytes of delta object ID ]\n\n     - zlib-packed data (length unknown, except we know how much we want \n       it to unpack to)\n\n - Finally at the end: 20 bytes of SHA1 of the pack-file contents (up to \n   the SHA1)\n\nso it's actually possible to pick up the objects as they come off the \nstream, since the SHA1 name is defined by the contents and you don't need \nthe index file unless you want to look things up.\n\nSo the receiver side could try this algorithm:\n\n - unpack each object in memory on the receiving side\n\n\tIf the unpack failed, it must have been the SHA1 at the end, so \n\tverify it!\n\n - if it's a delta object and you haven't seen the object it's a delta \n   against, keep it in memory.\n\n - if it's a non-delta object, just write it to the object store, and try \n   to resolve any delta objects you have pending that this new object \n   satisfies. That in turn creates other objects that may have more deltas \n   they satisfy etc.\n\nwhich looks quite doable. The delta objects are small, so keeping them in \nmemory shouldn't be a problem (especially since we _tend_ to write deltas \nafter the object they depend on).\n\nI can certainly add an option to git-pack-file that disables writing of\nthe index file, and just writes the pack-file to stdout. I'm not sure I\nwant to write the \"parse incoming pack-file\" thing, but git-unpack-objects\ncomes _reasonably_ close (but right now it seeks around using the index\nfile to resolve deltas, instead of keeping them in memory and resolving\nthem when possible). But I can make the infrastructure ready for it.\n\nSounds like a plan.\n\n\t\t\tLinus\n"},{"id":"5355","messageId":"Pine.LNX.4.58.0506281111480.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506281019450.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-28T18:17:05Z","receivedAt":"2005-06-28T18:17:05Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 28 Jun 2005, Linus Torvalds wrote:\n> \n> I can certainly add an option to git-pack-file that disables writing of\n> the index file, and just writes the pack-file to stdout.\n\nDone.\n\n>\t\t\t\t\t\t I'm not sure I\n> want to write the \"parse incoming pack-file\" thing, but git-unpack-objects\n> comes _reasonably_ close (but right now it seeks around using the index\n> file to resolve deltas, instead of keeping them in memory and resolving\n> them when possible).\n\nI'm still thinking about this one. I think I'll just do it.\n\nOne problem here is that since we don't know how big the incoming\npack-file will be, in a streaming input environment the receiver needs to\neither make the pack-file reception be the last thing it sees, or it will\nhave to live with the fact that \"git-unpack-objects\" will read some more\nthan it needs before it notices that it got it all...\n\nWe can handle the latter either by padding (make the rule be that\ngit-unpack-file will always read in chunks of 4kB max, and pad the output\nwith 4kB of zero bytes or something, and then you can execute\ngit-unpack-objects and continue reading stdin afterwards, removing any\nzeroes that git-unpack-file didn't eat), or by having git-unpack-objects \nflush anything after the final SHA1 to _its_ stdout, so that you can get \nthe following data/commands in the stream from the unpack-file thing. \nUgly, in any case.\n\n\t\tLinus\n"},{"id":"5360","messageId":"pan.2005.06.28.19.49.33.828202@smurf.noris.de","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506281111480.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Matthias Urlichs","fromEmail":"smurf@smurf.noris.de","sentAt":"2005-06-28T19:49:35Z","receivedAt":"2005-06-28T19:49:35Z","isPatch":false,"sender":{"key":"matthias@urlichs.de","avatar":"https://gravatar.com/avatar/2708905af227313eba6f2b2ae0f7d0259b5ac5d71baef58fe5a13c699ce0bbf0?d=mp&s=160"},"body":"Hi, Linus Torvalds wrote:\n\n> Ugly, in any case.\n\nWhy not chunk the thing?\n\nIn other words, the stream shouldn't be\n\n\t\"here's a big-ass packfile of unknown size\"\n\nbut an arbitrary number of\n\n\t\"here's a N-byte sized chunk of the current pack file\"\nsnippets, followed by a\n\t\"here's the SHA1 of the whole thing\"\npacket.\n\n-- \nMatthias Urlichs   |   {M:U} IT Design @ m-u-it.de   |  smurf@smurf.noris.de\nDisclaimer: The quote was selected randomly. Really. | http://smurf.noris.de\n - -\nBe like a duck -- keep calm and unruffled on the surface but paddle like the\ndevil under water.\n"},{"id":"5368","messageId":"Pine.LNX.4.21.0506281535340.30848-100000@iabervon.org","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506281111480.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-28T20:01:09Z","receivedAt":"2005-06-28T20:01:09Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Tue, 28 Jun 2005, Linus Torvalds wrote:\n\n> On Tue, 28 Jun 2005, Linus Torvalds wrote:\n> > \n> > I can certainly add an option to git-pack-file that disables writing of\n> > the index file, and just writes the pack-file to stdout.\n> \n> Done.\n\nWhat I actually meant was that it would be useful for git-ssh-push to be\nable to pack stuff as a function call rather than execing an external\nprogram, because just sticking git-ssh-push at the end of a pipeline\ndoesn't work if you don't remember what the remote side has.\n\n> >\t\t\t\t\t\t I'm not sure I\n> > want to write the \"parse incoming pack-file\" thing, but git-unpack-objects\n> > comes _reasonably_ close (but right now it seeks around using the index\n> > file to resolve deltas, instead of keeping them in memory and resolving\n> > them when possible).\n> \n> I'm still thinking about this one. I think I'll just do it.\n\nOne possibility would be to put a special type tag (like '\\0') before the\nhash, so that the format is more deterministic.\n\n> One problem here is that since we don't know how big the incoming\n> pack-file will be, in a streaming input environment the receiver needs to\n> either make the pack-file reception be the last thing it sees, or it will\n> have to live with the fact that \"git-unpack-objects\" will read some more\n> than it needs before it notices that it got it all...\n\nIn a completely streaming environment, yes; but the receiving side is the\none sending commands, so you don't run into the next thing unless you're\noverlapping requests. Failing that, we can just keep a 4k buffer of stuff\nwe've already read around; we don't have to worry about reading into\nsomething we won't want to read at all.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5362","messageId":"pan.2005.06.28.20.17.59.178402@smurf.noris.de","threadId":"1034","inReplyTo":"pan.2005.06.28.19.49.33.828202@smurf.noris.de","subject":"Re: CAREFUL! No more delta object support!","fromName":"Matthias Urlichs","fromEmail":"smurf@smurf.noris.de","sentAt":"2005-06-28T20:18:01Z","receivedAt":"2005-06-28T20:18:01Z","isPatch":false,"sender":{"key":"matthias@urlichs.de","avatar":"https://gravatar.com/avatar/2708905af227313eba6f2b2ae0f7d0259b5ac5d71baef58fe5a13c699ce0bbf0?d=mp&s=160"},"body":"I wrote:\n\n> Linus Torvalds wrote:\n> \n>> Ugly, in any case.\n> \n> Why not chunk the thing?\n\nHaving the number of files sent first would work too, I'd think.\n\nI'm wary of trying to interpret something non-decompressible as a sha1\nchunk, however -- the set of random bytes that, to zlib, look like a\nsufficiently valid zip header that it wants to read more than 20 of them\nbefore punting is certainly not zero.\n\n-- \nMatthias Urlichs   |   {M:U} IT Design @ m-u-it.de   |  smurf@smurf.noris.de\nDisclaimer: The quote was selected randomly. Really. | http://smurf.noris.de\n - -\nI was sure the old fellow would never make it\nto the other side of the curb when I struck him.\n"},{"id":"5365","messageId":"20050628203004.GH1275@pasky.ji.cz","threadId":"1034","inReplyTo":"20050628163551.GA29410@kvack.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-06-28T20:30:04Z","receivedAt":"2005-06-28T20:30:04Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Tue, Jun 28, 2005 at 06:35:51PM CEST, I got a letter\nwhere Benjamin LaHaise <bcrl@kvack.org> told me that...\n> On Tue, Jun 28, 2005 at 04:52:56PM +0200, Petr Baudis wrote:\n> > I think the git-*-pull tools are actually just fine. You will only need\n> > to have some server-side CGI gadget to frontend the file, but we need\n> > that anyway to make the pull reasonably effective.\n> \n> Not really -- the use of rsync for the objects fails horribly on slow \n> links when the project scales in the number of commits.  The rsync \n> protocol has to transfer the names of each file and some information \n> about it, and that information isn't delta compressed.  This is where \n> kernel.org is falling over, as well as what makes the kernel tree very \n> painful to use over a dialup modem link.\n\nYes. But isn't that what I'm after all saying too? git-*-pull tools\nshouldn't have that problem since they have much less overhead and only\npull stuff you need.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\n<Espy> be careful, some twit might quote you out of context..\n"},{"id":"5376","messageId":"7vmzpataae.fsf_-_@assigned-by-dhcp.cox.net","threadId":"1034","inReplyTo":"7vslz2x3vg.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH] Expose packed_git and alt_odb.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-28T21:56:57Z","receivedAt":"2005-06-28T21:56:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"The commands git-fsck-cache and probably git-*-pull needs to\nhave a way to enumerate objects contained in packed GIT archives\nand alternate object pools.  This commit exposes the data\nstructure used to keep track of them from sha1_file.c, and adds\na couple of accessor interface functions for use by the enhanced\ngit-fsck-cache command.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n cache.h     |   19 +++++++++++++++++++\n sha1_file.c |   43 ++++++++++++++++++++++++-------------------\n 2 files changed, 43 insertions(+), 19 deletions(-)\n\nda37711700d11f8c7f44fcb6819c724978c840b7\ndiff --git a/cache.h b/cache.h\n--- a/cache.h\n+++ b/cache.h\n@@ -233,4 +233,23 @@ struct checkout {\n \n extern int checkout_entry(struct cache_entry *ce, struct checkout *state);\n \n+extern struct alternate_object_database {\n+\tchar *base;\n+\tchar *name;\n+} *alt_odb;\n+extern void prepare_alt_odb(void);\n+\n+extern struct packed_git {\n+\tstruct packed_git *next;\n+\tunsigned long index_size;\n+\tunsigned long pack_size;\n+\tunsigned int *index_base;\n+\tvoid *pack_base;\n+\tunsigned int pack_last_used;\n+\tchar pack_name[0]; /* something like \".git/objects/pack/xxxxx.pack\" */\n+} *packed_git;\n+extern void prepare_packed_git(void);\n+extern int num_packed_objects(const struct packed_git *p);\n+extern int nth_packed_object_sha1(const struct packed_git *, int, unsigned char*);\n+\n #endif /* CACHE_H */\ndiff --git a/sha1_file.c b/sha1_file.c\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -184,10 +184,7 @@ char *sha1_file_name(const unsigned char\n \treturn base;\n }\n \n-static struct alternate_object_database {\n-\tchar *base;\n-\tchar *name;\n-} *alt_odb;\n+struct alternate_object_database *alt_odb;\n \n /*\n  * Prepare alternate object database registry.\n@@ -205,13 +202,15 @@ static struct alternate_object_database \n  * pointed by base fields of the array elements with one xmalloc();\n  * the string pool immediately follows the array.\n  */\n-static void prepare_alt_odb(void)\n+void prepare_alt_odb(void)\n {\n \tint pass, totlen, i;\n \tconst char *cp, *last;\n \tchar *op = NULL;\n \tconst char *alt = gitenv(ALTERNATE_DB_ENVIRONMENT) ? : \"\";\n \n+\tif (alt_odb)\n+\t\treturn;\n \t/* The first pass counts how large an area to allocate to\n \t * hold the entire alt_odb structure, including array of\n \t * structs and path buffers for them.  The second pass fills\n@@ -258,8 +257,7 @@ static char *find_sha1_file(const unsign\n \n \tif (!stat(name, st))\n \t\treturn name;\n-\tif (!alt_odb)\n-\t\tprepare_alt_odb();\n+\tprepare_alt_odb();\n \tfor (i = 0; (name = alt_odb[i].name) != NULL; i++) {\n \t\tfill_sha1_path(name, sha1);\n \t\tif (!stat(alt_odb[i].base, st))\n@@ -271,15 +269,7 @@ static char *find_sha1_file(const unsign\n #define PACK_MAX_SZ (1<<26)\n static int pack_used_ctr;\n static unsigned long pack_mapped;\n-static struct packed_git {\n-\tstruct packed_git *next;\n-\tunsigned long index_size;\n-\tunsigned long pack_size;\n-\tunsigned int *index_base;\n-\tvoid *pack_base;\n-\tunsigned int pack_last_used;\n-\tchar pack_name[0]; /* something like \".git/objects/pack/xxxxx.pack\" */\n-} *packed_git;\n+struct packed_git *packed_git;\n \n struct pack_entry {\n \tunsigned int offset;\n@@ -430,7 +420,7 @@ static void prepare_packed_git_one(char \n \t}\n }\n \n-static void prepare_packed_git(void)\n+void prepare_packed_git(void)\n {\n \tint i;\n \tstatic int run_once = 0;\n@@ -439,8 +429,7 @@ static void prepare_packed_git(void)\n \t\treturn;\n \n \tprepare_packed_git_one(get_object_directory());\n-\tif (!alt_odb)\n-\t\tprepare_alt_odb();\n+\tprepare_alt_odb();\n \tfor (i = 0; alt_odb[i].base != NULL; i++) {\n \t\talt_odb[i].name[0] = 0;\n \t\tprepare_packed_git_one(alt_odb[i].base);\n@@ -750,6 +739,22 @@ static void *unpack_entry(struct pack_en\n \treturn unpack_non_delta_entry(pack+5, size, left);\n }\n \n+int num_packed_objects(const struct packed_git *p)\n+{\n+\t/* See check_packed_git_idx and pack-objects.c */\n+\treturn (p->index_size - 20 - 20 - 4*256) / 24;\n+}\n+\n+int nth_packed_object_sha1(const struct packed_git *p, int n,\n+\t\t\t   unsigned char* sha1)\n+{\n+\tvoid *index = p->index_base + 256;\n+\tif (n < 0 || num_packed_objects(p) <= n)\n+\t\treturn -1;\n+\tmemcpy(sha1, (index + 24 * n + 4), 20);\n+\treturn 0;\n+}\n+\n static int find_pack_entry_1(const unsigned char *sha1,\n \t\t\t     struct pack_entry *e, struct packed_git *p)\n {\n------------\n"},{"id":"5375","messageId":"7vhdfita7q.fsf_-_@assigned-by-dhcp.cox.net","threadId":"1034","inReplyTo":"7vslz2x3vg.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH 3/3] Update fsck-cache (take 2)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-28T21:58:33Z","receivedAt":"2005-06-28T21:58:33Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"The fsck-cache complains if objects referred to by files in\n.git/refs/ or objects stored in files under .git/objects/??/ are\nnot found as stand-alone SHA1 files (i.e. found in alternate\nobject pools GIT_ALTERNATE_OBJECT_DIRECTORIES or packed archives\nstored under .git/objects/pack).\n\nAlthough this is a good semantics to maintain consistency of a\nsingle .git/objects directory as a self contained set of\nobjects, it sometimes is useful to consider it is OK as long as\nthese \"outside\" objects are available.\n\nThis commit introduces a new flag, --standalone, to\ngit-fsck-cache.  When it is not specified, connectivity checks\nand .git/refs pointer checks are taught that it is OK when\nexpected objects do not exist under .git/objects/?? hierarchy\nbut are available from an packed archive or in an alternate\nobject pool.\n\nAnother new flag, --full, makes git-fsck-cache to check not only\nthe current GIT_OBJECT_DIRECTORY but also objects found in\nalternate object pools and packed GIT archives.a\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n*** This completes \"the other half\" the fsck updates I did last\n*** night was missing.  Please discard that one and use this\n*** instead.\n\n Documentation/git-fsck-cache.txt |   18 +++++++++-\n fsck-cache.c                     |   71 ++++++++++++++++++++++++++++++++------\n 2 files changed, 76 insertions(+), 13 deletions(-)\n\n5cae1fa43bfeae6722d916aa764fa75d9ce1839a\ndiff --git a/Documentation/git-fsck-cache.txt b/Documentation/git-fsck-cache.txt\n--- a/Documentation/git-fsck-cache.txt\n+++ b/Documentation/git-fsck-cache.txt\n@@ -9,7 +9,7 @@ git-fsck-cache - Verifies the connectivi\n \n SYNOPSIS\n --------\n-'git-fsck-cache' [--tags] [--root] [--unreachable] [--cache] [<object>*]\n+'git-fsck-cache' [--tags] [--root] [--unreachable] [--cache] [--standalone | --full] [<object>*]\n \n DESCRIPTION\n -----------\n@@ -37,6 +37,22 @@ OPTIONS\n \tConsider any object recorded in the cache also as a head node for\n \tan unreachability trace.\n \n+--standalone::\n+\tLimit checks to the contents of GIT_OBJECT_DIRECTORY\n+\t(.git/objects), making sure that it is consistent and\n+\tcomplete without referring to objects found in alternate\n+\tobject pools listed in GIT_ALTERNATE_OBJECT_DIRECTORIES,\n+\tnor packed GIT archives found in .git/objects/pack;\n+\tcannot be used with --full.\n+\n+--full::\n+\tCheck not just objects in GIT_OBJECT_DIRECTORY\n+\t(.git/objects), but also the ones found in alternate\n+\tobject pools listed in GIT_ALTERNATE_OBJECT_DIRECTORIES,\n+\tand in packed GIT archives found in .git/objects/pack\n+\tand corresponding pack subdirectories in alternate\n+\tobject pools; cannot be used with --standalone.\n+\n It tests SHA1 and general object sanity, and it does full tracking of\n the resulting reachability and everything else. It prints out any\n corruption it finds (missing or bad objects), and if you use the\ndiff --git a/fsck-cache.c b/fsck-cache.c\n--- a/fsck-cache.c\n+++ b/fsck-cache.c\n@@ -12,6 +12,8 @@\n static int show_root = 0;\n static int show_tags = 0;\n static int show_unreachable = 0;\n+static int standalone = 0;\n+static int check_full = 0;\n static int keep_cache_objects = 0; \n static unsigned char head_sha1[20];\n \n@@ -25,13 +27,17 @@ static void check_connectivity(void)\n \t\tstruct object_list *refs;\n \n \t\tif (!obj->parsed) {\n-\t\t\tprintf(\"missing %s %s\\n\",\n-\t\t\t       obj->type, sha1_to_hex(obj->sha1));\n+\t\t\tif (!standalone && has_sha1_file(obj->sha1))\n+\t\t\t\t; /* it is in pack */\n+\t\t\telse\n+\t\t\t\tprintf(\"missing %s %s\\n\",\n+\t\t\t\t       obj->type, sha1_to_hex(obj->sha1));\n \t\t\tcontinue;\n \t\t}\n \n \t\tfor (refs = obj->refs; refs; refs = refs->next) {\n-\t\t\tif (refs->item->parsed)\n+\t\t\tif (refs->item->parsed ||\n+\t\t\t    (!standalone && has_sha1_file(refs->item->sha1)))\n \t\t\t\tcontinue;\n \t\t\tprintf(\"broken link from %7s %s\\n\",\n \t\t\t       obj->type, sha1_to_hex(obj->sha1));\n@@ -315,8 +321,11 @@ static int read_sha1_reference(const cha\n \t\treturn -1;\n \n \tobj = lookup_object(sha1);\n-\tif (!obj)\n+\tif (!obj) {\n+\t\tif (!standalone && has_sha1_file(sha1))\n+\t\t\treturn 0; /* it is in pack */\n \t\treturn error(\"%s: invalid sha1 pointer %.40s\", path, hexname);\n+\t}\n \n \tobj->used = 1;\n \tmark_reachable(obj, REACHABLE);\n@@ -366,10 +375,20 @@ static void get_default_heads(void)\n \t\tdie(\"No default references\");\n }\n \n+static void fsck_object_dir(const char *path)\n+{\n+\tint i;\n+\tfor (i = 0; i < 256; i++) {\n+\t\tstatic char dir[4096];\n+\t\tsprintf(dir, \"%s/%02x\", path, i);\n+\t\tfsck_dir(i, dir);\n+\t}\n+\tfsck_sha1_list();\n+}\n+\n int main(int argc, char **argv)\n {\n \tint i, heads;\n-\tchar *sha1_dir;\n \n \tfor (i = 1; i < argc; i++) {\n \t\tconst char *arg = argv[i];\n@@ -390,17 +409,45 @@ int main(int argc, char **argv)\n \t\t\tkeep_cache_objects = 1;\n \t\t\tcontinue;\n \t\t}\n+\t\tif (!strcmp(arg, \"--standalone\")) {\n+\t\t\tstandalone = 1;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (!strcmp(arg, \"--full\")) {\n+\t\t\tcheck_full = 1;\n+\t\t\tcontinue;\n+\t\t}\n \t\tif (*arg == '-')\n-\t\t\tusage(\"git-fsck-cache [--tags] [[--unreachable] [--cache] <head-sha1>*]\");\n+\t\t\tusage(\"git-fsck-cache [--tags] [[--unreachable] [--cache] [--standalone | --full] <head-sha1>*]\");\n \t}\n \n-\tsha1_dir = get_object_directory();\n-\tfor (i = 0; i < 256; i++) {\n-\t\tstatic char dir[4096];\n-\t\tsprintf(dir, \"%s/%02x\", sha1_dir, i);\n-\t\tfsck_dir(i, dir);\n+\tif (standalone && check_full)\n+\t\tdie(\"Only one of --standalone or --full can be used.\");\n+\tif (standalone)\n+\t\tunsetenv(\"GIT_ALTERNATE_OBJECT_DIRECTORIES\");\n+\n+\tfsck_object_dir(get_object_directory());\n+\tif (check_full) {\n+\t\tint j;\n+\t\tstruct packed_git *p;\n+\t\tprepare_alt_odb();\n+\t\tfor (j = 0; alt_odb[j].base; j++) {\n+\t\t\talt_odb[j].name[-1] = 0; /* was slash */\n+\t\t\tfsck_object_dir(alt_odb[j].base);\n+\t\t\talt_odb[j].name[-1] = '/';\n+\t\t}\n+\t\tprepare_packed_git();\n+\t\tfor (p = packed_git; p; p = p->next) {\n+\t\t\tint num = num_packed_objects(p);\n+\t\t\tfor (i = 0; i < num; i++) {\n+\t\t\t\tunsigned char sha1[20];\n+\t\t\t\tnth_packed_object_sha1(p, i, sha1);\n+\t\t\t\tif (fsck_sha1(sha1) < 0)\n+\t\t\t\t\tfprintf(stderr, \"bad sha1 entry '%s'\\n\", sha1_to_hex(sha1));\n+\n+\t\t\t}\n+\t\t}\n \t}\n-\tfsck_sha1_list();\n \n \theads = 0;\n \tfor (i = 1; i < argc; i++) {\n------------\n"},{"id":"5389","messageId":"7vll4uoulk.fsf@assigned-by-dhcp.cox.net","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506280921480.19755@ppc970.osdl.org","subject":"[PATCH] Emit base objects of a delta chain when the delta is output.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-29T00:49:27Z","receivedAt":"2005-06-29T00:49:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> While adding a new object to a pack file is _possible_ (you add it to the\nLT> end of the pack-file, and re-generate the index file), I would strongly\nLT> suggest against it for several reasons:\n\nOK, people have convinced me not to dream on ;-).\n\nLT> Btw, I'm not claiming that my current pack format is \"optimal\" of course.  \nLT> For example, while I write all objects in recency order, right now that\nLT> means that if a recent object has been written as a delta that depends on\nLT> an older one, I actually write the delta first (correct) but I won't write\nLT> the older object until its recency ordering (wrong).\n\nI agree.  \n\nHow does this one look?  Lightly tested by packing, unpacking\nwithout -n and fsck'ing, not unpacking but placing it under\n.git/objects/pack and running fsck with --full, all using the\ncurrent GIT repo.\n\n------------\nDeltas are useless by themselves and when you use them you need\nto get to their base objects.  A base object should inherit\nrecency from the most recent deltified object that is based on\nit and that is what this patch teaches git-pack-objects.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\ncd /opt/packrat/playpen/public/in-place/git/git.junio/\njit-diff\n# - master: Use enhanced diff_delta() in the similarity estimator.\n# + (working tree)\ndiff --git a/pack-objects.c b/pack-objects.c\n--- a/pack-objects.c\n+++ b/pack-objects.c\n@@ -118,6 +118,23 @@ static unsigned long write_object(struct\n \treturn hdrlen + datalen;\n }\n \n+static unsigned long write_one(struct sha1file *f,\n+\t\t\t       struct object_entry *e,\n+\t\t\t       unsigned long offset)\n+{\n+\tif (e->offset)\n+\t\t/* offset starts from header size and cannot be zero\n+\t\t * if it is written already.\n+\t\t */\n+\t\treturn offset;\n+\te->offset = offset;\n+\toffset += write_object(f, e);\n+\t/* if we are delitified, write out its base object. */\n+\tif (e->delta)\n+\t\toffset = write_one(f, e->delta, offset);\n+\treturn offset;\n+}\n+\n static void write_pack_file(void)\n {\n \tint i;\n@@ -135,11 +152,9 @@ static void write_pack_file(void)\n \thdr.hdr_entries = htonl(nr_objects);\n \tsha1write(f, &hdr, sizeof(hdr));\n \toffset = sizeof(hdr);\n-\tfor (i = 0; i < nr_objects; i++) {\n-\t\tstruct object_entry *entry = objects + i;\n-\t\tentry->offset = offset;\n-\t\toffset += write_object(f, entry);\n-\t}\n+\tfor (i = 0; i < nr_objects; i++)\n+\t\toffset = write_one(f, objects + i, offset);\n+\n \tsha1close(f, pack_file_sha1, 1);\n \tmb = offset >> 20;\n \toffset &= 0xfffff;\n\nCompilation finished at Tue Jun 28 17:43:31\n"},{"id":"5391","messageId":"Pine.LNX.4.58.0506282041360.19755@ppc970.osdl.org","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506281111480.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-29T03:53:59Z","receivedAt":"2005-06-29T03:53:59Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 28 Jun 2005, Linus Torvalds wrote:\n> \n> >\t\t\t\t\t\t I'm not sure I\n> > want to write the \"parse incoming pack-file\" thing, but git-unpack-objects\n> > comes _reasonably_ close (but right now it seeks around using the index\n> > file to resolve deltas, instead of keeping them in memory and resolving\n> > them when possible).\n> \n> I'm still thinking about this one. I think I'll just do it.\n\nOk, done. I had to basically rewrite that unpacking logic, but the end \nresult is actually slightly smaller and cleaner, and it can now unpack \nfrom a stream. That stream reading logic that uncompresses directly from \nthe stream buffer might be considered a bit too subtle (and somebody \nshould really double-check it), but hey, it works for me.\n\nIn fact, I just did this:\n\n\t#\n\t# Create empty git archive \"~/unpack\"\t\n\t#\n\tmkdir ~/unpack\n\tcd ~/unpack\n\tgit-init-db\n\n\t#\n\t# Copy the git archive there over a pipe\n\t#\n\tcd ~/git\n\tgit-rev-list --objects HEAD | git-pack-objects --depth=50 --window=50 --stdout | (cd ~/unpack ; git-unpack-objects)\n\n\t#\n\t# Go to new archive, set up the head, and fsck to verify\n\t#\n\tcd ~/unpack\n\tcat ~/git/.git/HEAD > .git/HEAD \n\tgit-fsck-cache --unreachable\n\nNow, the above is a silly example, since I _could_ just have moved the\npack file into .git/objects/pack, but that was not the point of this whole\nthing. The point was to do what a \"git-ssh-push\" would basically boil down\nto.\n\nI'd like somebody who knows zlib intimately to take a look at how I do the \nstreaming input thing (in particular, the \"use(len - stream.avail_in);\" \npart in the inflate loop in the \"get_data()\" function).\n\n\t\t\tLinus\n"},{"id":"5417","messageId":"Pine.LNX.4.58.0506291142510.14331@ppc970.osdl.org","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506271910390.19755@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-29T18:59:28Z","receivedAt":"2005-06-29T18:59:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 27 Jun 2005, Linus Torvalds wrote:\n> \n> On Mon, 27 Jun 2005, Junio C Hamano wrote:\n> > \n> > Shouldn't feeding \"git-rev-list --object\" output plus\n> > handcrafted list of objects in 2.6.11 tree object to\n> > git-pack-objects just work???\n> \n> You could do that. And yes, we can add support for \"tag\" objects too \n> (which the packing doesn't do at all right now. So this is not a \n> \"fundamental\" problem, it's just a practical one right now.\n\nOk, I've added the logic to \"git-rev-list --object\" to handle arbitrary \nobject dependencies.\n\nSo you can do things like this, if you want to:\n\n\tgit-rev-list --object HEAD ^v2.6.11-tree\n\nwhich basically generates the complete list of every object reachable from \nHEAD, but not reachable from the v2.6.11 tree. It also understands about \ntags, so if you do\n\n\tgit-rev-list --object v2.6.12 ^v2.6.11-tree\n\nthe end result will have the \"v2.6.12\" tag in it (along with all the\nobjects reachable from it, but not reachable from v2.6.11-tree).\n\nWhat does this mean? It means that you can do a \"push\" from repository \"a\" \nto repository \"b\" by doing\n\n - in \"b\", do\n\n\trefs_in_b=($(find .git/refs -type f | xargs cat))\n\n\n - in \"a\" do\n\n\trefs_in_a=($(find .git/refs -type f | xargs cat))\n\n - then, in \"a\", do\n\n\tgit-rev-list \"${refs_in_a[@]}\" --not \"${refs_in_b[@]}\" |\n\t\tgit-pack-objects --stdout > push.pack\n\n   to generate the objects pack in \"push.pack\"\n\n - then, in \"b\", do\n\n\tgit-unpack-objects < push.pack\n\nand you now have moved over _all_ the objects that were referenced in \"a\",\nbut not in \"b\". Including tags etc. So after that last stage, when you've\nunpacked the objects, the only thing left to do is to make the refs in \"b\"  \npoint to the new references from \"a\" (which basically boils down to a\n\"cp\", except it would be good to verify that the refs in \"b\" still have\nthe same values as they did before we did the object push).\n\nDaniel (or anybody else), interested? Please?\n\nOf course, you can do this one branch at a time, too, if you want to, but\nthe above was meant as an example of how you can actually do all the\nbranches in one single pack-file, which is a lot more efficient (if you do\nit one branch at a time, you'll quite possible end up transferring objects\nthat are reachable in other branches multiple times, while the \"all in one\ngo\" thing will pack each object just once).\n\nNow, have I actually _tested_ the above? Hell no. But all the heavy \nlifting should now be done for doing an efficient \"git push\" that pushes \nall branches in one go (or one at a time, it's your choice on how you end \nup using git-rev-list).\n\n\t\tLinus\n"},{"id":"5421","messageId":"Pine.LNX.4.21.0506291644120.30848-100000@iabervon.org","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506291142510.14331@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-29T21:05:32Z","receivedAt":"2005-06-29T21:05:32Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Wed, 29 Jun 2005, Linus Torvalds wrote:\n\n> and you now have moved over _all_ the objects that were referenced in \"a\",\n> but not in \"b\". Including tags etc. So after that last stage, when you've\n> unpacked the objects, the only thing left to do is to make the refs in \"b\"  \n> point to the new references from \"a\" (which basically boils down to a\n> \"cp\", except it would be good to verify that the refs in \"b\" still have\n> the same values as they did before we did the object push).\n> \n> Daniel (or anybody else), interested? Please?\n\nI'll probably get to this over the weekend.\n\n> Of course, you can do this one branch at a time, too, if you want to, but\n> the above was meant as an example of how you can actually do all the\n> branches in one single pack-file, which is a lot more efficient (if you do\n> it one branch at a time, you'll quite possible end up transferring objects\n> that are reachable in other branches multiple times, while the \"all in one\n> go\" thing will pack each object just once).\n\nIt should transfer each only once if you recalculate \"refs_in_b\" after\neach push, right? Or is the marking for \"--objects ^commit\" still not\ntight wrt object and tree files? I think branch-at-a-time is preferable\nfor the case where the source doesn't want to send quite everything, and\nthe target doesn't necessarily want everything named the same.\n\n> Now, have I actually _tested_ the above? Hell no. But all the heavy \n> lifting should now be done for doing an efficient \"git push\" that pushes \n> all branches in one go (or one at a time, it's your choice on how you end \n> up using git-rev-list).\n\nThe one thing I can think of is whether things will blow up if the target\nrepository has heads that aren't in the source, at which point the source\nhas no clue what to exclude. I.e.:\n\nparent -- new-b\n  \\\n   new-a\n\nIf I've moved the head on b forward to new-b, and a wants to push new-a\n(as a new branch, perhaps), refs_in_b has only new-b, refs_in_a has parent\nand new-a, and git-rev-list in a can't see that b has parent (and\neverything upwards of that). You probably just don't want to do this, but\nI bet that some people will (e.g. projects that synchronize through a\nshared-owner repository).\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5424","messageId":"Pine.LNX.4.58.0506291435310.14331@ppc970.osdl.org","threadId":"1034","inReplyTo":"Pine.LNX.4.21.0506291644120.30848-100000@iabervon.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-29T21:38:44Z","receivedAt":"2005-06-29T21:38:44Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 29 Jun 2005, Daniel Barkalow wrote:\n> \n> > Of course, you can do this one branch at a time, too, if you want to, but\n> > the above was meant as an example of how you can actually do all the\n> > branches in one single pack-file, which is a lot more efficient (if you do\n> > it one branch at a time, you'll quite possible end up transferring objects\n> > that are reachable in other branches multiple times, while the \"all in one\n> > go\" thing will pack each object just once).\n> \n> It should transfer each only once if you recalculate \"refs_in_b\" after\n> each push, right?\n\nYes, you can do it that way too. It will possibly not pack as well due to\ngiving you fewer opportunities for deltas, but that's likely not a huge \nissue.\n\n> The one thing I can think of is whether things will blow up if the target\n> repository has heads that aren't in the source\n\nRight. I think that's a \"feature\" of pushing: you cannot push to an \narchive that has state that you don't know about. Ie you can only push to \nsomething that is a proper subset of what you are (on a per-branch basis, \nof course - not necessarily on a \"global\" stage - so you could push just \n_one_ branch, even if another branch was ahead of where you are).\n\n\t\t\tLinus\n"},{"id":"5427","messageId":"Pine.LNX.4.21.0506291806510.30848-100000@iabervon.org","threadId":"1034","inReplyTo":"Pine.LNX.4.58.0506291435310.14331@ppc970.osdl.org","subject":"Re: CAREFUL! No more delta object support!","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-29T22:24:32Z","receivedAt":"2005-06-29T22:24:32Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Wed, 29 Jun 2005, Linus Torvalds wrote:\n\n> On Wed, 29 Jun 2005, Daniel Barkalow wrote:\n> > The one thing I can think of is whether things will blow up if the target\n> > repository has heads that aren't in the source\n> \n> Right. I think that's a \"feature\" of pushing: you cannot push to an \n> archive that has state that you don't know about. Ie you can only push to \n> something that is a proper subset of what you are (on a per-branch basis, \n> of course - not necessarily on a \"global\" stage - so you could push just \n> _one_ branch, even if another branch was ahead of where you are).\n\nThe issue is really distinguishing the \"other\" branches I don't care about\nfrom the one that I do care about. With -w, I almost certainly care about\nthe ref I'm writing, but that doesn't help for refs that are new (new\nbranches or tags), for which I care about some other thing. Also, the\nfailure is a bit hard to detect, I think, in that I could find I do\nrecognize some ancient thing that's barely useful for exclusion, and miss\nsomething that should exclude almost everything but it's been updated. In\nany case, when things go wrong we simply send stuff the recipient already\nhas, so it's not the end of the world. (And there's probably some clever\nway of dealing with it)\n\n\t-Daniel\n*This .sig left intentionally blank*\n"}]}