{"thread":{"id":"35841","subject":"[PATCH] Introduce experimental remote object access mode","startedAt":"2014-02-11T08:54:54Z","lastAt":"2014-02-12T20:55:20Z","messageCount":3,"participants":["Shawn Pearce","Junio C Hamano"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"234608","messageId":"CAJo=hJsO=FBkiOo5fuPbToxE1SR3Lh8oim0eTAR6bH1a-TcdPA@mail.gmail.com","threadId":"35841","inReplyTo":null,"subject":"[PATCH] Introduce experimental remote object access mode","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2014-02-11T08:54:54Z","receivedAt":"2014-02-11T08:54:54Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Make it easy to experiment what remote access to objects would be\nlike if the network ran at say 1 ms round trip latency to obtain\nany object not on the local repository.\n\n  $ time git ls-tree -r HEAD\n  real 0m0.059s\n\n  $ time GIT_RTT=1 git ls-tree -r HEAD\n  real 0m27.283s\n\nYes kids, slowing down loose object access by just 1ms if all\nobjects are remote can take a simple ls-tree from 59ms to more\nthan enough time to drink tea or coffee.\n\nWhy would you do this? Perhaps you need more time in your day\nto consume tea or coffee. Set GIT_RTT and enjoy a beverage.\n\nSo-not-signed-off-by: this author or anyone else\n---\n\n  :-)\n\n sha1_file.c | 14 ++++++++++++++\n 1 file changed, 14 insertions(+)\n\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 6e8c05d..9bdcbc3 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -38,6 +38,7 @@ const unsigned char null_sha1[20];\n\n static const char *no_log_pack_access = \"no_log_pack_access\";\n static const char *log_pack_access;\n+static useconds_t rtt;\n\n /*\n  * This is meant to hold a *small* number of objects that you would\n@@ -436,9 +437,20 @@ void prepare_alt_odb(void)\n  read_info_alternates(get_object_directory(), 0);\n }\n\n+static void apply_rtt()\n+{\n+ if (!rtt) {\n+ char *rtt_str = getenv(\"GIT_RTT\");\n+ rtt = rtt_str ? strtoul(rtt_str, NULL, 10) * 1000 : 1;\n+ }\n+ if (rtt > 1)\n+ usleep(rtt);\n+}\n+\n static int has_loose_object_local(const unsigned char *sha1)\n {\n  char *name = sha1_file_name(sha1);\n+ apply_rtt();\n  return !access(name, F_OK);\n }\n\n@@ -1303,6 +1315,7 @@ void prepare_packed_git(void)\n\n  if (prepare_packed_git_run_once)\n  return;\n+\n  prepare_packed_git_one(get_object_directory(), 1);\n  prepare_alt_odb();\n  for (alt = alt_odb_list; alt; alt = alt->next) {\n@@ -1439,6 +1452,7 @@ static int open_sha1_file(const unsigned char *sha1)\n  struct alternate_object_database *alt;\n\n  fd = git_open_noatime(name);\n+ apply_rtt();\n  if (fd >= 0)\n  return fd;\n\n-- \n1.9.0.rc1.175.g0b1dcb5\n"},{"id":"234623","messageId":"xmqqppmtphx0.fsf@gitster.dls.corp.google.com","threadId":"35841","inReplyTo":"CAJo=hJsO=FBkiOo5fuPbToxE1SR3Lh8oim0eTAR6bH1a-TcdPA@mail.gmail.com","subject":"Re: [PATCH] Introduce experimental remote object access mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-02-11T19:29:15Z","receivedAt":"2014-02-11T19:29:15Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n> Why would you do this? Perhaps you need more time in your day\n> to consume tea or coffee. Set GIT_RTT and enjoy a beverage.\n\nSo the conclusion is that it is not practical to do a lazy fetch if\nit is done extremely naively at \"we want this object --- wait a bit\nand we'll give you\" level?\n\nI am wondering if we can do a bit better, like \"we want this object\n--- wait a bit, ah that's a commit, so it is likely that you may\nwant the trees and blobs associated with it, too, if not right now\nbut in a near future, let me push a pack that holds them to you\"?\n\n>\n> So-not-signed-off-by: this author or anyone else\n> ---\n>\n>   :-)\n>\n>  sha1_file.c | 14 ++++++++++++++\n>  1 file changed, 14 insertions(+)\n>\n> diff --git a/sha1_file.c b/sha1_file.c\n> index 6e8c05d..9bdcbc3 100644\n> --- a/sha1_file.c\n> +++ b/sha1_file.c\n> @@ -38,6 +38,7 @@ const unsigned char null_sha1[20];\n>\n>  static const char *no_log_pack_access = \"no_log_pack_access\";\n>  static const char *log_pack_access;\n> +static useconds_t rtt;\n>\n>  /*\n>   * This is meant to hold a *small* number of objects that you would\n> @@ -436,9 +437,20 @@ void prepare_alt_odb(void)\n>   read_info_alternates(get_object_directory(), 0);\n>  }\n>\n> +static void apply_rtt()\n> +{\n> + if (!rtt) {\n> + char *rtt_str = getenv(\"GIT_RTT\");\n> + rtt = rtt_str ? strtoul(rtt_str, NULL, 10) * 1000 : 1;\n> + }\n> + if (rtt > 1)\n> + usleep(rtt);\n> +}\n> +\n>  static int has_loose_object_local(const unsigned char *sha1)\n>  {\n>   char *name = sha1_file_name(sha1);\n> + apply_rtt();\n>   return !access(name, F_OK);\n>  }\n>\n> @@ -1303,6 +1315,7 @@ void prepare_packed_git(void)\n>\n>   if (prepare_packed_git_run_once)\n>   return;\n> +\n>   prepare_packed_git_one(get_object_directory(), 1);\n>   prepare_alt_odb();\n>   for (alt = alt_odb_list; alt; alt = alt->next) {\n> @@ -1439,6 +1452,7 @@ static int open_sha1_file(const unsigned char *sha1)\n>   struct alternate_object_database *alt;\n>\n>   fd = git_open_noatime(name);\n> + apply_rtt();\n>   if (fd >= 0)\n>   return fd;\n"},{"id":"234690","messageId":"CAJo=hJvx7vRcNk0ZtAtM99gfc-b1k9xjk_cOHco=-GgRMy55qg@mail.gmail.com","threadId":"35841","inReplyTo":"xmqqppmtphx0.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH] Introduce experimental remote object access mode","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2014-02-12T20:55:20Z","receivedAt":"2014-02-12T20:55:20Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Tue, Feb 11, 2014 at 11:29 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> Shawn Pearce <spearce@spearce.org> writes:\n>\n>> Why would you do this? Perhaps you need more time in your day\n>> to consume tea or coffee. Set GIT_RTT and enjoy a beverage.\n>\n> So the conclusion is that it is not practical to do a lazy fetch if\n> it is done extremely naively at \"we want this object --- wait a bit\n> and we'll give you\" level?\n\nYes, this is what I thought when someone proposed this hack in\nsha1_file.c to me on Monday. So I ran a quick experiment to see if my\ninstinct was right.\n\n> I am wondering if we can do a bit better, like \"we want this object\n> --- wait a bit, ah that's a commit, so it is likely that you may\n> want the trees and blobs associated with it, too, if not right now\n> but in a near future, let me push a pack that holds them to you\"?\n\nAh, smart observation. That might work. However I doubt it.\n\nI implemented a version of Git on top of Google Bigtable (and Apache\nHBase and Apache Cassandra) multiple times using JGit. tl;dr: this\napproach doesn't work in practice.\n\n\nThe naive implementation for these distributed NoSQL systems is to\nstore each object in its own row keyed by SHA-1, and lookup the object\nwhen you want it. This is very slow and is more or less what this\nstupid patch shows. Worse, none of them were able to even get close to\nthe 1ms latency I used in this example.\n\nIn another implementation (which I published into JGit as the \"DHT\"\nbackend) I stored a group of related commits together in a row. Row\ntarget sizes were in the 1-2 MiB range when using pack style\ncompression for commits, so the average row held hundreds of commits.\nReading one commit would actually slurp back a number of related\ncommits. The idea was if we need commit A now we will need B, C, D, E\n(its parents and ancestors) soon as the application walks the revision\nhistory, like rev-list or pack-objects.\n\nFor a process like pack-objects this almost seems to work. If commits\nare together we can slurp a group at a time to amortize the round trip\nlatency. Unfortunately the application can still go through hundreds\nof commits faster than the real world RTT is. So I tried to fix this\nby storing an extra metadata pointer in each row to identify the next\nrow, so next block of commits could start loading right away. Its\nstill slow, as the application can scan through data faster than the\nRTT.\n\nAt least for pack generation the traversal code does commits and\nbuilds up a list of all root trees. The root trees can be async loaded\nin batches, but the depth first traversal is still a killer. There are\nstalls while the application waits for the next subtree, even if you\ncluster the trees also into groups using depth first traversal the way\nthe packer produces pack files today.\n\n\nJunio's idea to cluster data by commit and its related trees and blobs\nis just a different data organization. It may be necessary to make two\ncopies of the data, one clustered by commits and another by\ncommit+tree+blob to satisfy different access patterns. And in the\ncommit+tree+blob case you may need multiple redundant copies of blobs\nnear commits that use them if those commits are frequently accessed.\nIts a lot of redundant disk space.\n\nWe always say disk is cheap, but disk is slow and not getting faster.\nSSDs are helping, but SSDs are expensive and have size limitations\ncompared to spinning disk. Just making many copies of data isn't\nnecessarily a solution.\n"}]}