{"thread":{"id":"835","subject":"[PATCH] pull: gracefully recover from delta retrieval failure.","startedAt":"2005-06-05T06:11:38Z","lastAt":"2005-06-06T18:30:56Z","messageCount":8,"participants":["Junio C Hamano","Jason McMullan","Daniel Barkalow","McMullan, Jason"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"4564","messageId":"7v4qcde3j9.fsf@assigned-by-dhcp.cox.net","threadId":"835","inReplyTo":null,"subject":"[PATCH] pull: gracefully recover from delta retrieval failure.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-05T06:11:38Z","receivedAt":"2005-06-05T06:11:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This addresses a concern raised by Jason McMullan in the mailing\nlist discussion.  After retrieving and storing a potentially\ndeltified object, pull logic tries to check and fulfil its delta\ndependency.  When the pull procedure is killed at this point,\nhowever, there was no easy way to recover by re-running pull,\nsince next run would have found that we already have that\ndeltified object and happily reported success, without really\nchecking its delta dependency is satisfied.\n\nThis patch introduces --recover option to git-*-pull family\nwhich causes them to re-validate dependency of deltified objects\nwe are fetching.  A new test t5100-delta-pull.sh covers such a\nfailure mode.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n*** Linus, from now on I will go into \"calming down\" mode and\n*** refrain myself from sending you too many \"new\" stuff, until\n*** you tell me otherwise.  I will concentrate on fixes like\n*** this one and the \"diff-* -B fix\" patches I sent you earlier.\n*** Perhaps I would also work on CVS migration documents if you\n*** would like me to help you in that area as well.\n\n*** Definitely things like the idea of diff-tree switching its\n*** pathspec according rename detection results would not be\n*** something I'll be bugging you about until 1.0 happens;\n*** unless you tell me otherwise, that is.\n\n Documentation/git-http-pull.txt  |    5 ++\n Documentation/git-local-pull.txt |    5 ++\n Documentation/git-rpull.txt      |    5 ++\n pull.h                           |    4 +-\n http-pull.c                      |    4 +-\n local-pull.c                     |    4 +-\n pull.c                           |   15 +++++--\n rpull.c                          |    4 +-\n t/t5100-delta-pull.sh            |   79 ++++++++++++++++++++++++++++++++++++++\n 9 files changed, 113 insertions(+), 12 deletions(-)\n\ndiff --git a/Documentation/git-http-pull.txt b/Documentation/git-http-pull.txt\n--- a/Documentation/git-http-pull.txt\n+++ b/Documentation/git-http-pull.txt\n@@ -9,7 +9,7 @@ git-http-pull - Downloads a remote GIT r\n \n SYNOPSIS\n --------\n-'git-http-pull' [-c] [-t] [-a] [-v] [-d] commit-id url\n+'git-http-pull' [-c] [-t] [-a] [-v] [-d] [--recover] commit-id url\n \n DESCRIPTION\n -----------\n@@ -25,6 +25,9 @@ Downloads a remote GIT repository via HT\n \tDo not check for delta base objects (use this option\n \tonly when you know the remote repository is not\n \tdeltified).\n+--recover::\n+\tCheck dependency of deltified object more carefully than\n+\tusual, to recover after earlier pull that was interrupted.\n -v::\n \tReport what is downloaded.\n \ndiff --git a/Documentation/git-local-pull.txt b/Documentation/git-local-pull.txt\n--- a/Documentation/git-local-pull.txt\n+++ b/Documentation/git-local-pull.txt\n@@ -9,7 +9,7 @@ git-local-pull - Duplicates another GIT \n \n SYNOPSIS\n --------\n-'git-local-pull' [-c] [-t] [-a] [-l] [-s] [-n] [-v] [-d] commit-id path\n+'git-local-pull' [-c] [-t] [-a] [-l] [-s] [-n] [-v] [-d] [--recover] commit-id path\n \n DESCRIPTION\n -----------\n@@ -27,6 +27,9 @@ OPTIONS\n \tDo not check for delta base objects (use this option\n \tonly when you know the remote repository is not\n \tdeltified).\n+--recover::\n+\tCheck dependency of deltified object more carefully than\n+\tusual, to recover after earlier pull that was interrupted.\n -v::\n \tReport what is downloaded.\n \ndiff --git a/Documentation/git-rpull.txt b/Documentation/git-rpull.txt\n--- a/Documentation/git-rpull.txt\n+++ b/Documentation/git-rpull.txt\n@@ -10,7 +10,7 @@ git-rpull - Pulls from a remote reposito\n \n SYNOPSIS\n --------\n-'git-rpull' [-c] [-t] [-a] [-d] [-v] commit-id url\n+'git-rpull' [-c] [-t] [-a] [-d] [-v] [--recover] commit-id url\n \n DESCRIPTION\n -----------\n@@ -29,6 +29,9 @@ OPTIONS\n \tDo not check for delta base objects (use this option\n \tonly when you know the remote repository is not\n \tdeltified).\n+--recover::\n+\tCheck dependency of deltified object more carefully than\n+\tusual, to recover after earlier pull that was interrupted.\n -v::\n \tReport what is downloaded.\n \ndiff --git a/pull.h b/pull.h\n--- a/pull.h\n+++ b/pull.h\n@@ -13,7 +13,9 @@ extern int get_history;\n /** Set to fetch the trees in the commit history. **/\n extern int get_all;\n \n-/* Set to zero to skip the check for delta object base. */\n+/* Set to zero to skip the check for delta object base;\n+ * set to two to check delta dependency even for objects we already have.\n+ */\n extern int get_delta;\n \n /* Set to be verbose */\ndiff --git a/http-pull.c b/http-pull.c\n--- a/http-pull.c\n+++ b/http-pull.c\n@@ -105,6 +105,8 @@ int main(int argc, char **argv)\n \t\t\tget_history = 1;\n \t\t} else if (argv[arg][1] == 'd') {\n \t\t\tget_delta = 0;\n+\t\t} else if (!strcmp(argv[arg], \"--recover\")) {\n+\t\t\tget_delta = 2;\n \t\t} else if (argv[arg][1] == 'a') {\n \t\t\tget_all = 1;\n \t\t\tget_tree = 1;\n@@ -115,7 +117,7 @@ int main(int argc, char **argv)\n \t\targ++;\n \t}\n \tif (argc < arg + 2) {\n-\t\tusage(\"git-http-pull [-c] [-t] [-a] [-d] [-v] commit-id url\");\n+\t\tusage(\"git-http-pull [-c] [-t] [-a] [-d] [-v] [--recover] commit-id url\");\n \t\treturn 1;\n \t}\n \tcommit_id = argv[arg];\ndiff --git a/local-pull.c b/local-pull.c\n--- a/local-pull.c\n+++ b/local-pull.c\n@@ -74,7 +74,7 @@ int fetch(unsigned char *sha1)\n }\n \n static const char *local_pull_usage = \n-\"git-local-pull [-c] [-t] [-a] [-l] [-s] [-n] [-v] [-d] commit-id path\";\n+\"git-local-pull [-c] [-t] [-a] [-l] [-s] [-n] [-v] [-d] [--recover] commit-id path\";\n \n /* \n  * By default we only use file copy.\n@@ -94,6 +94,8 @@ int main(int argc, char **argv)\n \t\t\tget_history = 1;\n \t\telse if (argv[arg][1] == 'd')\n \t\t\tget_delta = 0;\n+\t\telse if (!strcmp(argv[arg], \"--recover\"))\n+\t\t\tget_delta = 2;\n \t\telse if (argv[arg][1] == 'a') {\n \t\t\tget_all = 1;\n \t\t\tget_tree = 1;\ndiff --git a/pull.c b/pull.c\n--- a/pull.c\n+++ b/pull.c\n@@ -6,6 +6,7 @@\n \n int get_tree = 0;\n int get_history = 0;\n+/* 1 means \"get delta\", 2 means \"really check delta harder */\n int get_delta = 1;\n int get_all = 0;\n int get_verbosely = 0;\n@@ -32,12 +33,16 @@ static void report_missing(const char *w\n \n static int make_sure_we_have_it(const char *what, unsigned char *sha1)\n {\n-\tint status;\n-\tif (has_sha1_file(sha1))\n+\tint status = 0;\n+\n+\tif (!has_sha1_file(sha1)) {\n+\t\tstatus = fetch(sha1);\n+\t\tif (status && what)\n+\t\t\treport_missing(what, sha1);\n+\t}\n+\telse if (get_delta < 2)\n \t\treturn 0;\n-\tstatus = fetch(sha1);\n-\tif (status && what)\n-\t\treport_missing(what, sha1);\n+\n \tif (get_delta) {\n \t\tchar delta_sha1[20];\n \t\tstatus = sha1_delta_base(sha1, delta_sha1);\ndiff --git a/rpull.c b/rpull.c\n--- a/rpull.c\n+++ b/rpull.c\n@@ -52,6 +52,8 @@ int main(int argc, char **argv)\n \t\t\tget_history = 1;\n \t\t} else if (argv[arg][1] == 'd') {\n \t\t\tget_delta = 0;\n+\t\t} else if (!strcmp(argv[arg], \"--recover\")) {\n+\t\t\tget_delta = 2;\n \t\t} else if (argv[arg][1] == 'a') {\n \t\t\tget_all = 1;\n \t\t\tget_tree = 1;\n@@ -62,7 +64,7 @@ int main(int argc, char **argv)\n \t\targ++;\n \t}\n \tif (argc < arg + 2) {\n-\t\tusage(\"git-rpull [-c] [-t] [-a] [-v] [-d] commit-id url\");\n+\t\tusage(\"git-rpull [-c] [-t] [-a] [-v] [-d] [--recover] commit-id url\");\n \t\treturn 1;\n \t}\n \tcommit_id = argv[arg];\ndiff --git a/t/t5100-delta-pull.sh b/t/t5100-delta-pull.sh\nnew file mode 100644\n--- /dev/null\n+++ b/t/t5100-delta-pull.sh\n@@ -0,0 +1,79 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2005 Junio C Hamano\n+#\n+\n+test_description='Test pulling deltified objects\n+\n+'\n+. ./test-lib.sh\n+\n+locate_obj='s|\\(..\\)|.git/objects/\\1/|'\n+\n+test_expect_success \\\n+    setup \\\n+    'cat ../README >a &&\n+    git-update-cache --add a &&\n+    a0=`git-ls-files --stage |\n+        sed -e '\\''s/^[0-7]* \\([0-9a-f]*\\) .*/\\1/'\\''` &&\n+\n+    sed -e 's/test/TEST/g' ../README >a &&\n+    git-update-cache a &&\n+    a1=`git-ls-files --stage |\n+        sed -e '\\''s/^[0-7]* \\([0-9a-f]*\\) .*/\\1/'\\''` &&\n+    tree=`git-write-tree` &&\n+    commit=`git-commit-tree $tree </dev/null` &&\n+    a0f=`echo \"$a0\" | sed -e \"$locate_obj\"` &&\n+    a1f=`echo \"$a1\" | sed -e \"$locate_obj\"` &&\n+    echo commit $commit &&\n+    echo a0 $a0 &&\n+    echo a1 $a1 &&\n+    ls -l $a0f $a1f &&\n+    echo $commit >.git/HEAD &&\n+    git-mkdelta -v $a0 $a1 &&\n+    ls -l $a0f $a1f'\n+\n+# Now commit has a tree that records delitified \"a\" whose SHA1 is a1.\n+# Create a new repo and pull this commit into it.\n+\n+test_expect_success \\\n+    'setup and cd into new repo' \\\n+    'mkdir dest && cd dest && rm -fr .git && git-init-db'\n+     \n+test_expect_success \\\n+    'pull from deltified repo into a new repo without -d' \\\n+    'rm -fr .git a && git-init-db &&\n+     git-local-pull -v -a $commit ../.git/ &&\n+     git-cat-file blob $a1 >a &&\n+     diff -u a ../a'\n+\n+test_expect_failure \\\n+    'pull from deltified repo into a new repo with -d' \\\n+    'rm -fr .git a && git-init-db &&\n+     git-local-pull -v -a -d $commit ../.git/ &&\n+     git-cat-file blob $a1 >a &&\n+     diff -u a ../a'\n+\n+test_expect_failure \\\n+    'pull from deltified repo after delta failure without --recover' \\\n+    'rm -f a &&\n+     git-local-pull -v -a $commit ../.git/ &&\n+     git-cat-file blob $a1 >a &&\n+     diff -u a ../a'\n+\n+test_expect_success \\\n+    'pull from deltified repo after delta failure with --recover' \\\n+    'rm -f a &&\n+     git-local-pull -v -a --recover $commit ../.git/ &&\n+     git-cat-file blob $a1 >a &&\n+     diff -u a ../a'\n+\n+test_expect_success \\\n+    'missing-tree or missing-blob should be re-fetched without --recover' \\\n+    'rm -f a $a0f $a1f &&\n+     git-local-pull -v -a $commit ../.git/ &&\n+     git-cat-file blob $a1 >a &&\n+     diff -u a ../a'\n+\n+test_done\n+\n------------\n\n"},{"id":"4581","messageId":"1117989532.10424.7.camel@port.evillabs.net","threadId":"835","inReplyTo":"7v4qcde3j9.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] pull: gracefully recover from delta retrieval failure.","fromName":"Jason McMullan","fromEmail":"jason.mcmullan@timesys.com","sentAt":"2005-06-05T16:38:52Z","receivedAt":"2005-06-05T16:38:52Z","isPatch":true,"sender":{"key":"jason.mcmullan@timesys.com","avatar":null},"body":"On Sat, 2005-06-04 at 23:11 -0700, Junio C Hamano wrote:\n> This addresses a concern raised by Jason McMullan in the mailing\n> list discussion.  After retrieving and storing a potentially\n> deltified object, pull logic tries to check and fulfil its delta\n> dependency.  When the pull procedure is killed at this point,\n> however, there was no easy way to recover by re-running pull,\n> since next run would have found that we already have that\n> deltified object and happily reported success, without really\n> checking its delta dependency is satisfied.\n\nI still think it would be much better if you didn't place unverified\nobjects in the database in the first place. You've taken care of delta\nobject recovery, yes, but what about unsatisfied tree objects? Or commit\nobjects? Does your algorithm require full depth scanning of the\nentire repository that is descended from the commit head?\n\nI much prefer to always leave the database in a consistent state, that\nway you only have to do O(number-of-retrieved-objects) verifications,\nnot O(number-of-commit-tree-ancestors) verifications.\n\nOr am I misunderstanding your technique here? \n\nSorry about being a pest, but this worries me. Please assuage my fears.\n\n(Or, if you'd like, I can rework pull.c to use the\n verification-before-store technique I used in my git-daemon patch, so\n all the *-pull mechanisms will be 'safe')\n\n"},{"id":"4584","messageId":"Pine.LNX.4.21.0506051256450.30848-100000@iabervon.org","threadId":"835","inReplyTo":"1117989532.10424.7.camel@port.evillabs.net","subject":"Re: [PATCH] pull: gracefully recover from delta retrieval failure.","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-05T17:24:05Z","receivedAt":"2005-06-05T17:24:05Z","isPatch":true,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sun, 5 Jun 2005, Jason McMullan wrote:\n\n> (Or, if you'd like, I can rework pull.c to use the\n>  verification-before-store technique I used in my git-daemon patch, so\n>  all the *-pull mechanisms will be 'safe')\n\nThe reasons I don't really want a verification-before-store are:\n\n - I'd like things to be resumable on error; if you've already got a bunch\n   of stuff and then the connection breaks, you should already have the\n   things you got; so, at least, the temporary locations should be\n   something predictable from the hash.\n\n - I'd like the user to be able to intentionally get partial repositories\n   of various sorts, and still have consistant information. If I don't\n   have anything written that's not in mainline, and I'm not going to be\n   applying other people's patches, and I don't have an up-to-date\n   repository, I'd like to be able to pull just mainline's head commit and\n   tree, and work from there. I don't need the history unless I want to\n   look up changes or want to merge something that's not derived entirely\n   from the head I've got.\n\nSo what I'd really like is something where you store whatever objects you\nhave, and also have extra information about what objects you know about\nbut don't have and what objects you've gotten completely. Of course, this\nneeds to be kept manageable.\n\n(Along the lines of the second one, there's a variety of partial\ninformation which is sufficient for various purposes. If I trust that\nLinus's latest tree is based on my most recent pull from him, I can\nfast-forward with just the tree. I can also merge his tree into mine with\njust the commits and his latest tree, since I must already have any common\nancestors. In all these cases, I may want to streamline my process by\ndoing \"pull tree; pull all &; checkout\" or \n\"pull tree,commits; pull all &; merge\", so that I can start on further\ndevelopment while the rest of the information fills in.)\n\nSo I'd greatly prefer to keep the metadata of what objects we have\nexplicitly, rather than implicitly in the presence or absence of files in\nthe object directory. Also, for objects which we expect to be missing, it\nwould be good to keep info on where we expect to be able to get\nthem. Then, if I'm wrong about what I actually needed, it doesn't need me\nto tell it again where to get things.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"4586","messageId":"7vekbg3de7.fsf@assigned-by-dhcp.cox.net","threadId":"835","inReplyTo":"1117989532.10424.7.camel@port.evillabs.net","subject":"Re: [PATCH] pull: gracefully recover from delta retrieval failure.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-05T17:46:24Z","receivedAt":"2005-06-05T17:46:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"JM\" == Jason McMullan <jason.mcmullan@timesys.com> writes:\n\nJM> Sorry about being a pest, but this worries me. Please assuage my fears.\n\nEarlier I said I suspected that the original code mishandled\nrecovery from a botched tree/commit dependent transfer, but that\nwas not the case.  The last test in the new test script I added\nin the patch you are responding to covers that case.\n\nJM> (Or, if you'd like, I can rework pull.c to use the\nJM>  verification-before-store technique I used in my git-daemon patch, so\nJM>  all the *-pull mechanisms will be 'safe')\n\nI would appreciate the offer.  I, however, would have to warn\nyou that the \"problem\" lies in the way the current pull\nstructure devides responsibility between the pull.c and transfer\nbackends.  The pull.c implements the dependency logic, and\ntransfer backends are to populate the database while being\noblivious of that logic.  From the purist point of view (I am\nsympathetic to your \"place only the verified objects in the\ndatabase\" principle), I am not entirely happy with that\ndivision, but at the same time I understand why it is done that\nway and even like it from practical standpoint.  Otherwise you\nneed to keep a bunch of objects somewhere outside the database\nalong with the list of \"things to rename to the final database\nname when we are done\".  You would somehow need to do clean-up\nwhen we fail in the middle _anyway_.\n\nIn other words, the current structure is optimized for non-\nfailure case, as it should be.  The original implementation\n(credit goes to Dan Barkalow) knew how to recover from failed\ntransfer (including the case that you first pull with -c or -t\nwithout using -a to miss some required objects to satisfy -a) by\nsimply running pull again, and with the --recover flag, it now\nknows how to recover from failed deltified object transfer as\nwell.\n\n"},{"id":"4600","messageId":"Pine.LNX.4.21.0506051523280.30848-100000@iabervon.org","threadId":"835","inReplyTo":"7vekbg3de7.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] pull: gracefully recover from delta retrieval failure.","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-05T20:02:57Z","receivedAt":"2005-06-05T20:02:57Z","isPatch":true,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sun, 5 Jun 2005, Junio C Hamano wrote:\n\n> >>>>> \"JM\" == Jason McMullan <jason.mcmullan@timesys.com> writes:\n> \n> JM> Sorry about being a pest, but this worries me. Please assuage my fears.\n> \n> Earlier I said I suspected that the original code mishandled\n> recovery from a botched tree/commit dependent transfer, but that\n> was not the case.  The last test in the new test script I added\n> in the patch you are responding to covers that case.\n\nIt does the O(history) method for correctness at the expense of\nefficiency; my hope is that a bit of caching can fix the efficiency issue\nas well. So the question is not really \"not safe\" as \"slow\". Of course, it\ntakes a while for this to become an issue, given the relationship of\nremote access latency to local access bandwidth. That is, you need a\nreally big history and to be getting very little new data before you'll\ncomplain.\n\n> JM> (Or, if you'd like, I can rework pull.c to use the\n> JM>  verification-before-store technique I used in my git-daemon patch, so\n> JM>  all the *-pull mechanisms will be 'safe')\n> \n> I would appreciate the offer.  I, however, would have to warn\n> you that the \"problem\" lies in the way the current pull\n> structure devides responsibility between the pull.c and transfer\n> backends.  The pull.c implements the dependency logic, and\n> transfer backends are to populate the database while being\n> oblivious of that logic.  From the purist point of view (I am\n> sympathetic to your \"place only the verified objects in the\n> database\" principle), I am not entirely happy with that\n> division, but at the same time I understand why it is done that\n> way and even like it from practical standpoint.\n\nAt one point I'd written a patch that split out the tmpfile usage of\nwrite_sha1_file(), made the filenames predictable, and used it for\neverything that writes those files. It had an \"open\" part and a\n\"close\" part (where the close also moved the file into place). This would\ngive the code better atomicity and protect against races between reading\nand validation. On the other hand, there's no reason to use an anonymous\ntemp file; just <filename>.partial or similar (with the proper open\nflags) would be sufficient and easier to clean or commit. Note that we\nwant to support /tmp and the object directory being on different\nfilesystems, also. (And all the open and place logic is nicely wrapped up\nin sha1_file.c)\n\nAside from the question of whether we want to insist that the object\ndatabase only includes objects such that everything reachable is also\npresent, we certainly want to only include objects which we have\ncompletely fetched, which are generically well-formed, and which have the\nadvertized hash, and having there never be an unvalidated file at the\nfilename would be good.\n\nBy this reasoning, a file should only be renamed after all of the delta\nrequirements are satisfied, but before tree and commit requirements are\nsatisfied. We certainly aren't going to have much use for files whose\ncontents we cannot get. This means that we'd like to have multiple\nunplaced files, but we don't need to read the contents of an unplaced\nfile.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"4642","messageId":"1118065849.8970.37.camel@jmcmullan.timesys","threadId":"835","inReplyTo":"Pine.LNX.4.21.0506051523280.30848-100000@iabervon.org","subject":"Database consistency after a successful pull","fromName":"McMullan, Jason","fromEmail":"jason.mcmullan@timesys.com","sentAt":"2005-06-06T13:50:49Z","receivedAt":"2005-06-06T13:50:49Z","isPatch":false,"sender":{"key":"jason.mcmullan@timesys.com","avatar":null},"body":"Subject Was: [PATCH] pull: gracefu[PAlly recover from delta retrieval\nfailure.]\n\n[snip lots of really good information about the thinking\n behind the design of the pull mechanisms ]\n\nOk, so would I be correct in the following assumptions\nabout the validity of a 'consistent' .git/objects database:\n\n============================================================\n\nCommits:\n\t* May have the tree they refer to in the database\n\t* Must have their parents in the database\n\nTrees:\n\t* Must have the blobs they refer to in the database\n\t* Must have the trees they refer to in the database\n\nDeltas:\n\t* Must have the referred to object in the database\n\nBlobs:\n\t* No references to check\n\n\n============================================================\n\nIn short, the database would contain:\n\n\t* The entire commit history\n\t* Selected commits would have the entire tree available\n\nCorrect, or totally mistaken? If mistaken, what are the consitency\nrules?\n\n[Oh, and does PGP signing my messages bug anybody? If so, I can stop\n doing that on this list]\n\n-- \nJason McMullan <jason.mcmullan@timesys.com>\nTimeSys Corporation\n\n"},{"id":"4649","messageId":"Pine.LNX.4.21.0506061000531.30848-100000@iabervon.org","threadId":"835","inReplyTo":"1118065849.8970.37.camel@jmcmullan.timesys","subject":"Re: Database consistency after a successful pull","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-06T16:21:19Z","receivedAt":"2005-06-06T16:21:19Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Mon, 6 Jun 2005, McMullan, Jason wrote:\n\n> Subject Was: [PATCH] pull: gracefu[PAlly recover from delta retrieval\n> failure.]\n> \n> [snip lots of really good information about the thinking\n>  behind the design of the pull mechanisms ]\n> \n> Ok, so would I be correct in the following assumptions\n> about the validity of a 'consistent' .git/objects database:\n> \n> ============================================================\n> \n> Commits:\n> \t* May have the tree they refer to in the database\n> \t* Must have their parents in the database\n\nMay have their parents in the database; we want to be able to drop ancient\nhistory from non-archival sites at some point, if nothing else.\n\n> Trees:\n> \t* Must have the blobs they refer to in the database\n> \t* Must have the trees they refer to in the database\n\nIt's probably true that there's no point to having a tree available if you\ndon't have its contents, although that's a convenient intermediate stage,\nso that you can look up the contents of the tree with the ordinary parsing\ncode. On the other hand, I could imagine an ARM developer completely\nignoring arch/i386 (and just having write-tree use the parent tree's value\nfor it).\n\n> Deltas:\n> \t* Must have the referred to object in the database\n\nYes. Can't unpack without them.\n\n> Blobs:\n> \t* No references to check\n\nRight.\n\nAlso, tags reference objects of unknown type; it's probably not vital to\nhave the object.\n\nMy bias is to call a database consistent with only deltas having the\nreferents; the rest goes towards completeness, since you have and can read\neverything that you have anything for (but may not be able to do some\nparticular operation).\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n\n"},{"id":"4650","messageId":"1118082657.8970.42.camel@jmcmullan.timesys","threadId":"835","inReplyTo":"Pine.LNX.4.21.0506061000531.30848-100000@iabervon.org","subject":"Re: Database consistency after a successful pull","fromName":"McMullan, Jason","fromEmail":"jason.mcmullan@timesys.com","sentAt":"2005-06-06T18:30:56Z","receivedAt":"2005-06-06T18:30:56Z","isPatch":false,"sender":{"key":"jason.mcmullan@timesys.com","avatar":null},"body":"On Mon, 2005-06-06 at 12:21 -0400, Daniel Barkalow wrote:\n> [snip snip]\n>\n> My bias is to call a database consistent with only deltas having the\n> referents; the rest goes towards completeness, since you have and can read\n> everything that you have anything for (but may not be able to do some\n> particular operation).\n\nNow, if we had consistent URIs for the .git/branches/* files, we could\ndo 'lazy-pull' and really have our cake and eat it too.\n\n-- \nJason McMullan <jason.mcmullan@timesys.com>\nTimeSys Corporation\n\n"}]}