{"thread":{"id":"25319","subject":"[RFC] New type of remote helpers","startedAt":"2010-10-03T11:33:38Z","lastAt":"2010-10-07T21:17:18Z","messageCount":21,"participants":["Tomas Carnecky","Sverre Rabbelier","Jonathan Nieder","Ramkumar Ramachandra"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"152343","messageId":"4CA86A12.6080905@dbservice.com","threadId":"25319","inReplyTo":null,"subject":"[RFC] New type of remote helpers","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2010-10-03T11:33:38Z","receivedAt":"2010-10-03T11:33:38Z","isPatch":false,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"My work has the goal of making interaction with foreign SCMs more\nnatural. The work that was done on remote helpers is the right\ndirection. But the 'import' and 'export' commands are the wrong approach\nI think. The problem I have with 'import' is that updating the refs is\nleft up to the remote helper (or git-fast-import). So you lose the nice\noutput from ls-remote/fetch: non-ff and other warnings etc. I slightly\nmodified how the remote helpers (and fast-import) work, now they behave\nexactly like 'core' git when fetching: Git tells the remote helper to\nfetch some refs, the helper does that and creates a pack and git then\nupdates the refs (or not, depending on fast-forward etc). To test this\napproach I created a simple remote helper for svn.\n\nThe work is available in the 'remote-helper-fixups' branch at\nhttp://github.com/wereHamster/git. No end-user or technical\ndocumentation is available as of yet. The svn remote helper is only an\nexample how remote helpers can use the new commands. You can see it as a\n'technology preview', use with caution.\n\nHere is an example session of how I think foreign SCMs should integrate\nwith git. For each command I explain what changes were needed. This is\nnot just some theoretical output. This is a real, working example. I did\nnot modify any of the output text.\n\n$ git ls-remote svn::/Volumes/Dump/Source/Mirror/Transmission/\nr1017 (impure)                            trunk\nr919 (impure)                             branches/nat-traversal\nr480 (impure)                             branches/0.6\n\nGit learned to understand version numbers from foreign SCMs. Git\ndisplays those as 'impure' because it knows that version exists but does\nnot know yet which git commit that version maps to.\n\n\n$ git fetch svn::/Volumes/Dump/Source/Mirror/Transmission/\n*:refs/remotes/svn/*\nFrom svn::/Volumes/Dump/Source/Mirror/Transmission\n * [new branch]      trunk      -> svn/trunk\n * [new branch]      branches/nat-traversal -> svn/branches/nat-traversal\n * [new branch]      branches/0.6 -> svn/branches/0.6\n\nGit tells the remote helper to 'fetch r1017 trunk'.  The remote helper\ndoes that, creates the pack and then tells git that it imported r1017 as\ncommit c5fed7ec. This is done with a new reply to the 'fetch' command:\n'map r1017 c5fed7ec'. The remote helper can use that to inform core git\nas which git commit the impure ref was imported. Git can then update the\nrefs. At no point does the remote helper manipulate refs directly.\n\nThe pack is created by a heavily modified git-fast-import. The existing\nfast-import not only creates the pack but also updates the refs. This is\nno longer desired as git is in charge of updating the refs. My modified\nfast-import works like this: After creating a commit, it writes it's git\nobject name to stdout. That way the remote helper can figure out as\nwhich git commits the svn revisions were imported and relay that back to\ncore git using the above described 'map' reply.\n\n\n$ git show --show-notes=svn svn/trunk\ncommit c5fed7ecc318363523d3ea2045e1c16a378bb10c\nAuthor: livings124 <livings124@localhost>\nDate:   Wed Oct 18 13:57:19 2006 +0000\n\n    more traditional toolbar icons for those afraid of change\n\nNotes (svn):\n    b4697c4a-7d4c-4a30-bd92-6745580d73b3/trunk@1017\n\nThe svn helper needs to be able to map svn revisions to git commits.\ngit-svn does this by adding the 'git-svn-id' line to each commit\nmessage. I'm using git notes for that and it seems to work just fine.\nThe note contains the repo UUID, path within the repo and revision.\n\nThere was a challenge how to update the notes ref (refs/notes/svn). As\nwith fast-import, I did not want the remote helper to do it. Neither the\nremote helper nor fast-import should be writing any refs. But core git\ncan only update refs which were discovered during transport->fetch(). I\nmodified the remote helper 'fetch' command and the transport->fetch()\nfunction to return an optional list of refs. These are the refs that the\nremote helper wants to update but which should not be presented to the\nuser (because these are internally used refs, such as my svn notes).\n\nSo the whole session between git and my svn remote helper looks like this:\n> list\n< :r1017 trunk\n> fetch :r1017 trunk\n[helper creates the pack including history up to r1017 and associated\nsvn notes]\n< map r1017 <commit corresponding to r1017>\n< silent refs/notes/svn <new commit which stores the updated svn notes>\n\n\n$ git fetch svn::/Volumes/Dump/Source/Mirror/Transmission/\n*:refs/remotes/svn/*\nFrom svn::/Volumes/Dump/Source/Mirror/Transmission\n   c5fed7e..228eaf3  trunk      -> svn/trunk\n\n$ git fetch svn::/Volumes/Dump/Source/Mirror/Transmission/\n*:refs/remotes/svn/*\nFrom svn::/Volumes/Dump/Source/Mirror/Transmission\n   228eaf3..207e5e5  trunk      -> svn/trunk\n * [new branch]      branches/scrape -> svn/branches/scrape\n * [new branch]      branches/multitracker -> svn/branches/multitracker\n * [new branch]      branches/io -> svn/branches/io\n\nUpdating the svn branches works like expected. The remote helper\nautomatically detects which branches it already imported (by going\nthrough all refs and the attached svn notes) and creates a new pack with\nthe new commits. New branches are also detected. The svn notes are\nupdated accordingly.\n\ntom\n"},{"id":"152347","messageId":"1286108511-55876-1-git-send-email-tom@dbservice.com","threadId":"25319","inReplyTo":"4CA86A12.6080905@dbservice.com","subject":"[PATCH 1/6] Remote helper: accept ':<value> <name>' as a response to 'list'","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2010-10-03T12:21:46Z","receivedAt":"2010-10-03T12:21:46Z","isPatch":true,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"Parse <value> as the remote-specific revision indicator. The ref is first stored\nas 'impure', meaning that it doesn't have any representation within git. Only\nafter the remote helper fetches that version into git, it can tell us which\ngit object (SHA1) that revision maps to. That is done with the new reply\n'map <value> <sha1>' to the 'list' command.\n\nSigned-off-by: Tomas Carnecky <tom@dbservice.com>\n---\n builtin/ls-remote.c |    7 ++++++-\n cache.h             |    2 +-\n remote.c            |    8 +++++++-\n transport-helper.c  |   41 +++++++++++++++++++++++++++++++++++++++--\n 4 files changed, 53 insertions(+), 5 deletions(-)\n\ndiff --git a/builtin/ls-remote.c b/builtin/ls-remote.c\nindex 97eed40..23dd0d2 100644\n--- a/builtin/ls-remote.c\n+++ b/builtin/ls-remote.c\n@@ -109,7 +109,12 @@ int cmd_ls_remote(int argc, const char **argv, const char *prefix)\n \t\t\tcontinue;\n \t\tif (!tail_match(pattern, ref->name))\n \t\t\tcontinue;\n-\t\tprintf(\"%s\t%s\\n\", sha1_to_hex(ref->old_sha1), ref->name);\n+\t\tif (ref->impure) {\n+\t\t\tint len = strlen(ref->impure) + strlen(\" (impure)\");\n+\t\t\tprintf(\"%s (impure)%*s  %s\\n\", ref->impure, 40 - len, \" \", ref->name);\n+\t\t} else {\n+\t\t\tprintf(\"%s      %s\\n\", sha1_to_hex(ref->old_sha1), ref->name);\n+\t\t}\n \t}\n \treturn 0;\n }\ndiff --git a/cache.h b/cache.h\nindex 2ef2fa3..23b43a6 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -905,7 +905,7 @@ struct ref {\n \tstruct ref *next;\n \tunsigned char old_sha1[20];\n \tunsigned char new_sha1[20];\n-\tchar *symref;\n+\tchar *symref, *impure;\n \tunsigned int force:1,\n \t\tmerge:1,\n \t\tnonfastforward:1,\ndiff --git a/remote.c b/remote.c\nindex 9143ec7..c355f11 100644\n--- a/remote.c\n+++ b/remote.c\n@@ -907,6 +907,7 @@ static struct ref *copy_ref(const struct ref *ref)\n \tmemcpy(cpy, ref, sizeof(struct ref) + len + 1);\n \tcpy->next = NULL;\n \tcpy->symref = ref->symref ? xstrdup(ref->symref) : NULL;\n+\tcpy->impure = ref->impure ? xstrdup(ref->impure) : NULL;\n \tcpy->remote_status = ref->remote_status ? xstrdup(ref->remote_status) : NULL;\n \tcpy->peer_ref = copy_ref(ref->peer_ref);\n \treturn cpy;\n@@ -931,6 +932,7 @@ static void free_ref(struct ref *ref)\n \tfree_ref(ref->peer_ref);\n \tfree(ref->remote_status);\n \tfree(ref->symref);\n+\tfree(ref->impure);\n \tfree(ref);\n }\n \n@@ -1453,7 +1455,11 @@ int resolve_remote_symref(struct ref *ref, struct ref *list)\n \t\treturn 0;\n \tfor (; list; list = list->next)\n \t\tif (!strcmp(ref->symref, list->name)) {\n-\t\t\thashcpy(ref->old_sha1, list->old_sha1);\n+\t\t\tif (list->impure) {\n+\t\t\t\tref->impure = xstrdup(list->impure);\n+\t\t\t} else {\n+\t\t\t\thashcpy(ref->old_sha1, list->old_sha1);\n+\t\t\t}\n \t\t\treturn 0;\n \t\t}\n \treturn 1;\ndiff --git a/transport-helper.c b/transport-helper.c\nindex acfc88e..0fe886e 100644\n--- a/transport-helper.c\n+++ b/transport-helper.c\n@@ -307,6 +307,29 @@ static int release_helper(struct transport *transport)\n \treturn 0;\n }\n \n+/* Map an impure ref to its actual value within Git. */\n+static void map_impure_ref(int nr_heads, struct ref **to_fetch, char *map)\n+{\n+\tint i;\n+\tchar *eov;\n+\n+\teov = strchr(map, ' ');\n+\tif (!eov)\n+\t\tdie(\"Malformed impure ref mapping: %s\", map);\n+\t*eov = '\\0';\n+\n+\t/* There may be multiple impure refs with the same value, be sure to\n+\t * map all of them. */\n+\tfor (i = 0; i < nr_heads; i++) {\n+\t\tstruct ref *posn = to_fetch[i];\n+\t\tif (posn->impure && !strcmp(map, posn->impure)) {\n+\t\t\tget_sha1_hex(eov + 1, posn->old_sha1);\n+\t\t\tfree(posn->impure);\n+\t\t\tposn->impure = NULL;\n+\t\t}\n+\t}\n+}\n+\n static int fetch_with_fetch(struct transport *transport,\n \t\t\t    int nr_heads, struct ref **to_fetch)\n {\n@@ -318,11 +341,17 @@ static int fetch_with_fetch(struct transport *transport,\n \n \tfor (i = 0; i < nr_heads; i++) {\n \t\tconst struct ref *posn = to_fetch[i];\n+\n \t\tif (posn->status & REF_STATUS_UPTODATE)\n \t\t\tcontinue;\n \n-\t\tstrbuf_addf(&buf, \"fetch %s %s\\n\",\n+\t\tif (posn->impure) {\n+\t\t\tstrbuf_addf(&buf, \"fetch %s %s\\n\",\n+\t\t\t\tposn->impure, posn->name);\n+\t\t} else {\n+\t\t\tstrbuf_addf(&buf, \"fetch %s %s\\n\",\n \t\t\t    sha1_to_hex(posn->old_sha1), posn->name);\n+\t\t}\n \t}\n \n \tstrbuf_addch(&buf, '\\n');\n@@ -337,12 +366,18 @@ static int fetch_with_fetch(struct transport *transport,\n \t\t\t\twarning(\"%s also locked %s\", data->name, name);\n \t\t\telse\n \t\t\t\ttransport->pack_lockfile = xstrdup(name);\n+\t\t} else if (!prefixcmp(buf.buf, \"map \")) {\n+\t\t\tmap_impure_ref(nr_heads, to_fetch, buf.buf + 4);\n \t\t}\n \t\telse if (!buf.len)\n \t\t\tbreak;\n \t\telse\n \t\t\twarning(\"%s unexpectedly said: '%s'\", data->name, buf.buf);\n \t}\n+\n+\t/* The helper may have created one or more new packs. */\n+\treprepare_packed_git();\n+\n \tstrbuf_release(&buf);\n \treturn 0;\n }\n@@ -824,7 +859,9 @@ static struct ref *get_refs_list(struct transport *transport, int for_push)\n \t\t*tail = alloc_ref(eov + 1);\n \t\tif (buf.buf[0] == '@')\n \t\t\t(*tail)->symref = xstrdup(buf.buf + 1);\n-\t\telse if (buf.buf[0] != '?')\n+\t\telse if (buf.buf[0] == ':') {\n+\t\t\t(*tail)->impure = xstrdup(buf.buf + 1);\n+\t\t} else if (buf.buf[0] != '?')\n \t\t\tget_sha1_hex(buf.buf, (*tail)->old_sha1);\n \t\tif (eon) {\n \t\t\tif (has_attribute(eon + 1, \"unchanged\")) {\n-- \n1.7.3.37.gb6088b\n"},{"id":"152349","messageId":"1286108511-55876-2-git-send-email-tom@dbservice.com","threadId":"25319","inReplyTo":"4CA86A12.6080905@dbservice.com","subject":"[PATCH 2/6] Allow more than one keepfile in the transport","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2010-10-03T12:21:47Z","receivedAt":"2010-10-03T12:21:47Z","isPatch":true,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"Git itself creates only one packfile per fetch. But other transports may\nchose to create more than one (those that use fast-import for example).\nUse an array to keep track of the pack lockfiles.\n\nSigned-off-by: Tomas Carnecky <tom@dbservice.com>\n---\n transport-helper.c |    5 +----\n transport.c        |   27 ++++++++++++++++++++++-----\n transport.h        |   11 ++++++++++-\n 3 files changed, 33 insertions(+), 10 deletions(-)\n\ndiff --git a/transport-helper.c b/transport-helper.c\nindex 0fe886e..dcaaa89 100644\n--- a/transport-helper.c\n+++ b/transport-helper.c\n@@ -362,10 +362,7 @@ static int fetch_with_fetch(struct transport *transport,\n \n \t\tif (!prefixcmp(buf.buf, \"lock \")) {\n \t\t\tconst char *name = buf.buf + 5;\n-\t\t\tif (transport->pack_lockfile)\n-\t\t\t\twarning(\"%s also locked %s\", data->name, name);\n-\t\t\telse\n-\t\t\t\ttransport->pack_lockfile = xstrdup(name);\n+\t\t\ttransport_keep(transport, name);\n \t\t} else if (!prefixcmp(buf.buf, \"map \")) {\n \t\t\tmap_impure_ref(nr_heads, to_fetch, buf.buf + 4);\n \t\t}\ndiff --git a/transport.c b/transport.c\nindex 4dba6f8..df2baa7 100644\n--- a/transport.c\n+++ b/transport.c\n@@ -518,6 +518,7 @@ static int fetch_refs_via_pack(struct transport *transport,\n \tstruct fetch_pack_args args;\n \tint i;\n \tstruct ref *refs_tmp = NULL;\n+\tchar *keepfile = NULL;\n \n \tmemset(&args, 0, sizeof(args));\n \targs.uploadpack = data->options.uploadpack;\n@@ -541,7 +542,7 @@ static int fetch_refs_via_pack(struct transport *transport,\n \n \trefs = fetch_pack(&args, data->fd, data->conn,\n \t\t\t  refs_tmp ? refs_tmp : transport->remote_refs,\n-\t\t\t  dest, nr_heads, heads, &transport->pack_lockfile);\n+\t\t\t  dest, nr_heads, heads, &keepfile);\n \tclose(data->fd[0]);\n \tclose(data->fd[1]);\n \tif (finish_connect(data->conn))\n@@ -556,6 +557,10 @@ static int fetch_refs_via_pack(struct transport *transport,\n \tfree(origh);\n \tfree(heads);\n \tfree(dest);\n+\n+\tif (keepfile)\n+\t\ttransport_keep(transport, keepfile);\n+\n \treturn (refs ? 0 : -1);\n }\n \n@@ -1116,11 +1121,15 @@ int transport_fetch_refs(struct transport *transport, struct ref *refs)\n \n void transport_unlock_pack(struct transport *transport)\n {\n-\tif (transport->pack_lockfile) {\n-\t\tunlink_or_warn(transport->pack_lockfile);\n-\t\tfree(transport->pack_lockfile);\n-\t\ttransport->pack_lockfile = NULL;\n+\tint i;\n+\n+\tfor (i = 0; i < transport->keep_nr; ++i) {\n+\t\tchar *keepfile = (char *) transport->keep[i];\n+\t\tunlink_or_warn(keepfile);\n+\t\tfree(keepfile);\n \t}\n+\n+\ttransport->keep_nr = 0;\n }\n \n int transport_connect(struct transport *transport, const char *name,\n@@ -1137,6 +1146,7 @@ int transport_disconnect(struct transport *transport)\n \tint ret = 0;\n \tif (transport->disconnect)\n \t\tret = transport->disconnect(transport);\n+\tfree(transport->keep);\n \tfree(transport);\n \treturn ret;\n }\n@@ -1188,3 +1198,10 @@ char *transport_anonymize_url(const char *url)\n literal_copy:\n \treturn xstrdup(url);\n }\n+\n+void transport_keep(struct transport *transport, const char *keepfile)\n+{\n+\tint nr = transport->keep_nr + 1;\n+\tALLOC_GROW(transport->keep, nr, transport->keep_alloc);\n+\ttransport->keep[transport->keep_nr++] = keepfile;\n+}\ndiff --git a/transport.h b/transport.h\nindex c59d973..6320d28 100644\n--- a/transport.h\n+++ b/transport.h\n@@ -78,7 +78,14 @@ struct transport {\n \t * use. disconnect() releases these resources.\n \t **/\n \tint (*disconnect)(struct transport *connection);\n-\tchar *pack_lockfile;\n+\n+\t/** The transport can create zero or more pack files which need to be\n+\t * kept until we can update the refs. This array holds the names of the\n+\t * keep files which we have to delete once the refs are updated.\n+\t **/\n+\tconst char **keep;\n+\tint keep_nr, keep_alloc;\n+\n \tsigned verbose : 3;\n \t/**\n \t * Transports should not set this directly, and should use this\n@@ -165,4 +172,6 @@ int transport_refs_pushed(struct ref *ref);\n void transport_print_push_status(const char *dest, struct ref *refs,\n \t\t  int verbose, int porcelain, int *nonfastforward);\n \n+void transport_keep(struct transport *transport, const char *keepfile);\n+\n #endif\n-- \n1.7.3.37.gb6088b\n"},{"id":"152346","messageId":"1286108511-55876-3-git-send-email-tom@dbservice.com","threadId":"25319","inReplyTo":"4CA86A12.6080905@dbservice.com","subject":"[PATCH 3/6] Allow the transport fetch command to add additional refs","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2010-10-03T12:21:48Z","receivedAt":"2010-10-03T12:21:48Z","isPatch":true,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"The fetch transport command (in particular in remote helpers) may need to create\nor update additional refs which are used internally by the helper, but which\nit doesn't want to present to the user. Those refs are refered to as 'silent'\nthroughout the code because git should be silent about their presence, but yet\nprocess those just like the other refs.\n\nExample use case:\nRemote helpers such as those for svn may chose to save the Git SHA1 -> Subversion\nrevision mapping as notes attached to the commits (as opposed to strings in the\ncommit message itself). The helper would need to update the notes on each fetch,\nbut the user should not be bothered by the presence of that ref. The remote\nhelper can update the notes tree through fast-import and then inform Git core\nthat it should silently update the notes ref.\n\nSigned-off-by: Tomas Carnecky <tom@dbservice.com>\n---\n builtin/clone.c    |   11 +++++++----\n builtin/fetch.c    |   15 ++++++++++++---\n transport-helper.c |   32 ++++++++++++++++++++++++++++----\n transport.c        |   16 ++++++++++------\n transport.h        |    6 ++++--\n 5 files changed, 61 insertions(+), 19 deletions(-)\n\ndiff --git a/builtin/clone.c b/builtin/clone.c\nindex 19ed640..78355b6 100644\n--- a/builtin/clone.c\n+++ b/builtin/clone.c\n@@ -348,13 +348,16 @@ static struct ref *wanted_peer_refs(const struct ref *refs,\n \treturn local_refs;\n }\n \n-static void write_remote_refs(const struct ref *local_refs)\n+static void write_remote_refs(const struct ref *local_refs, struct ref *silent)\n {\n \tconst struct ref *r;\n \n \tfor (r = local_refs; r; r = r->next)\n \t\tadd_extra_ref(r->peer_ref->name, r->old_sha1, 0);\n \n+\tfor (r = silent; r; r = r->next)\n+\t\tadd_extra_ref(r->name, r->old_sha1, 0);\n+\n \tpack_refs(PACK_REFS_ALL);\n \tclear_extra_refs();\n }\n@@ -369,7 +372,7 @@ int cmd_clone(int argc, const char **argv, const char *prefix)\n \tconst struct ref *refs, *remote_head;\n \tconst struct ref *remote_head_points_at;\n \tconst struct ref *our_head_points_at;\n-\tstruct ref *mapped_refs;\n+\tstruct ref *mapped_refs, *silent = NULL;\n \tstruct strbuf key = STRBUF_INIT, value = STRBUF_INIT;\n \tstruct strbuf branch_top = STRBUF_INIT, reflog_msg = STRBUF_INIT;\n \tstruct transport *transport = NULL;\n@@ -542,14 +545,14 @@ int cmd_clone(int argc, const char **argv, const char *prefix)\n \t\trefs = transport_get_remote_refs(transport);\n \t\tif (refs) {\n \t\t\tmapped_refs = wanted_peer_refs(refs, refspec);\n-\t\t\ttransport_fetch_refs(transport, mapped_refs);\n+\t\t\ttransport_fetch_refs(transport, mapped_refs, &silent);\n \t\t}\n \t}\n \n \tif (refs) {\n \t\tclear_extra_refs();\n \n-\t\twrite_remote_refs(mapped_refs);\n+\t\twrite_remote_refs(mapped_refs, silent);\n \n \t\tremote_head = find_ref_by_name(refs, \"HEAD\");\n \t\tremote_head_points_at =\ndiff --git a/builtin/fetch.c b/builtin/fetch.c\nindex 6fc5047..71db090 100644\n--- a/builtin/fetch.c\n+++ b/builtin/fetch.c\n@@ -313,7 +313,7 @@ static int update_local_ref(struct ref *ref,\n }\n \n static int store_updated_refs(const char *raw_url, const char *remote_name,\n-\t\tstruct ref *ref_map)\n+\t\tstruct ref *ref_map, struct ref *silent)\n {\n \tFILE *fp;\n \tstruct commit *commit;\n@@ -411,6 +411,13 @@ static int store_updated_refs(const char *raw_url, const char *remote_name,\n \t\t\t\tfprintf(stderr, \" %s\\n\", note);\n \t\t}\n \t}\n+\n+\t/* Also update the silent refs. */\n+\tfor (rm = silent; rm; rm = rm->next) {\n+\t\tsnprintf(note, 1024, \"note here\"); /* TODO */\n+\t\tupdate_local_ref(rm, rm->name, note);\n+\t}\n+\n \tfree(url);\n \tfclose(fp);\n \tif (rc & STORE_REF_ERROR_DF_CONFLICT)\n@@ -496,13 +503,15 @@ static int quickfetch(struct ref *ref_map)\n \n static int fetch_refs(struct transport *transport, struct ref *ref_map)\n {\n+\tstruct ref *silent = NULL;\n+\n \tint ret = quickfetch(ref_map);\n \tif (ret)\n-\t\tret = transport_fetch_refs(transport, ref_map);\n+\t\tret = transport_fetch_refs(transport, ref_map, &silent);\n \tif (!ret)\n \t\tret |= store_updated_refs(transport->url,\n \t\t\t\ttransport->remote->name,\n-\t\t\t\tref_map);\n+\t\t\t\tref_map, silent);\n \ttransport_unlock_pack(transport);\n \treturn ret;\n }\ndiff --git a/transport-helper.c b/transport-helper.c\nindex dcaaa89..c0133ca 100644\n--- a/transport-helper.c\n+++ b/transport-helper.c\n@@ -330,8 +330,29 @@ static void map_impure_ref(int nr_heads, struct ref **to_fetch, char *map)\n \t}\n }\n \n+/* `buf` points to the 'silent' response from the helper. Parse it and\n+ * add the ref to the `silent` list. */\n+static void add_silent_ref(struct strbuf *buf, struct ref **silent)\n+{\n+\tif (!silent)\n+\t\treturn;\n+\n+\tstruct ref *ref;\n+\tchar *eon = strchr(buf->buf + 7, ' ');\n+\tif (!eon)\n+\t\twarning(\"Malformed helper response: %s\", buf->buf);\n+\n+\t*eon = '\\0';\n+\tref = alloc_ref(buf->buf + 7);\n+\tget_sha1_hex(eon + 1, ref->new_sha1);\n+\t\n+\tref->next = *silent;\n+\t*silent = ref;\n+}\n+\n static int fetch_with_fetch(struct transport *transport,\n-\t\t\t    int nr_heads, struct ref **to_fetch)\n+\t\t\t    int nr_heads, struct ref **to_fetch,\n+\t\t\t    struct ref **silent)\n {\n \tstruct helper_data *data = transport->data;\n \tint i;\n@@ -365,6 +386,8 @@ static int fetch_with_fetch(struct transport *transport,\n \t\t\ttransport_keep(transport, name);\n \t\t} else if (!prefixcmp(buf.buf, \"map \")) {\n \t\t\tmap_impure_ref(nr_heads, to_fetch, buf.buf + 4);\n+\t\t} else if (!prefixcmp(buf.buf, \"silent \")) {\n+\t\t\tadd_silent_ref(&buf, silent);\n \t\t}\n \t\telse if (!buf.len)\n \t\t\tbreak;\n@@ -559,14 +582,15 @@ static int connect_helper(struct transport *transport, const char *name,\n }\n \n static int fetch(struct transport *transport,\n-\t\t int nr_heads, struct ref **to_fetch)\n+\t\t int nr_heads, struct ref **to_fetch,\n+\t\t struct ref **silent)\n {\n \tstruct helper_data *data = transport->data;\n \tint i, count;\n \n \tif (process_connect(transport, 0)) {\n \t\tdo_take_over(transport);\n-\t\treturn transport->fetch(transport, nr_heads, to_fetch);\n+\t\treturn transport->fetch(transport, nr_heads, to_fetch, silent);\n \t}\n \n \tcount = 0;\n@@ -578,7 +602,7 @@ static int fetch(struct transport *transport,\n \t\treturn 0;\n \n \tif (data->fetch)\n-\t\treturn fetch_with_fetch(transport, nr_heads, to_fetch);\n+\t\treturn fetch_with_fetch(transport, nr_heads, to_fetch, silent);\n \n \tif (data->import)\n \t\treturn fetch_with_import(transport, nr_heads, to_fetch);\ndiff --git a/transport.c b/transport.c\nindex df2baa7..eaab276 100644\n--- a/transport.c\n+++ b/transport.c\n@@ -253,7 +253,8 @@ static struct ref *get_refs_via_rsync(struct transport *transport, int for_push)\n }\n \n static int fetch_objs_via_rsync(struct transport *transport,\n-\t\t\t\tint nr_objs, struct ref **to_fetch)\n+\t\t\t\tint nr_objs, struct ref **to_fetch,\n+\t\t\t\tstruct ref **silent)\n {\n \tstruct strbuf buf = STRBUF_INIT;\n \tstruct child_process rsync;\n@@ -428,7 +429,8 @@ static struct ref *get_refs_from_bundle(struct transport *transport, int for_pus\n }\n \n static int fetch_refs_from_bundle(struct transport *transport,\n-\t\t\t       int nr_heads, struct ref **to_fetch)\n+\t\t\t       int nr_heads, struct ref **to_fetch,\n+\t\t\t       struct ref **silent)\n {\n \tstruct bundle_transport_data *data = transport->data;\n \treturn unbundle(&data->header, data->fd);\n@@ -508,7 +510,8 @@ static struct ref *get_refs_via_connect(struct transport *transport, int for_pus\n }\n \n static int fetch_refs_via_pack(struct transport *transport,\n-\t\t\t       int nr_heads, struct ref **to_fetch)\n+\t\t\t       int nr_heads, struct ref **to_fetch,\n+\t\t\t       struct ref **silent)\n {\n \tstruct git_transport_data *data = transport->data;\n \tchar **heads = xmalloc(nr_heads * sizeof(*heads));\n@@ -1083,7 +1086,8 @@ const struct ref *transport_get_remote_refs(struct transport *transport)\n \treturn transport->remote_refs;\n }\n \n-int transport_fetch_refs(struct transport *transport, struct ref *refs)\n+int transport_fetch_refs(struct transport *transport, struct ref *refs,\n+\t\t\t\tstruct ref **silent)\n {\n \tint rc;\n \tint nr_heads = 0, nr_alloc = 0, nr_refs = 0;\n@@ -1113,9 +1117,9 @@ int transport_fetch_refs(struct transport *transport, struct ref *refs)\n \t\t\theads[nr_heads++] = rm;\n \t}\n \n-\trc = transport->fetch(transport, nr_heads, heads);\n-\n+\trc = transport->fetch(transport, nr_heads, heads, silent);\n \tfree(heads);\n+\n \treturn rc;\n }\n \ndiff --git a/transport.h b/transport.h\nindex 6320d28..22daf60 100644\n--- a/transport.h\n+++ b/transport.h\n@@ -52,7 +52,7 @@ struct transport {\n \t * get_refs_list(), it should set the old_sha1 fields in the\n \t * provided refs now.\n \t **/\n-\tint (*fetch)(struct transport *transport, int refs_nr, struct ref **refs);\n+\tint (*fetch)(struct transport *transport, int refs_nr, struct ref **refs, struct ref **silent);\n \n \t/**\n \t * Push the objects and refs. Send the necessary objects, and\n@@ -149,7 +149,9 @@ int transport_push(struct transport *connection,\n \n const struct ref *transport_get_remote_refs(struct transport *transport);\n \n-int transport_fetch_refs(struct transport *transport, struct ref *refs);\n+int transport_fetch_refs(struct transport *transport, struct ref *refs,\n+\t\t\tstruct ref **silent);\n+\n void transport_unlock_pack(struct transport *transport);\n int transport_disconnect(struct transport *transport);\n char *transport_anonymize_url(const char *url);\n-- \n1.7.3.37.gb6088b\n"},{"id":"152348","messageId":"1286108511-55876-4-git-send-email-tom@dbservice.com","threadId":"25319","inReplyTo":"4CA86A12.6080905@dbservice.com","subject":"[PATCH 4/6] Rename get_mode() to decode_tree_mode() and export it","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2010-10-03T12:21:49Z","receivedAt":"2010-10-03T12:21:49Z","isPatch":true,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"Other sources (fast-import-helper.c) may want to use this function\nto parse trees.\n\nSigned-off-by: Tomas Carnecky <tom@dbservice.com>\n---\n tree-walk.c |    4 ++--\n tree-walk.h |    2 ++\n 2 files changed, 4 insertions(+), 2 deletions(-)\n\ndiff --git a/tree-walk.c b/tree-walk.c\nindex a9bbf4e..5f51f4e 100644\n--- a/tree-walk.c\n+++ b/tree-walk.c\n@@ -3,7 +3,7 @@\n #include \"unpack-trees.h\"\n #include \"tree.h\"\n \n-static const char *get_mode(const char *str, unsigned int *modep)\n+const char *decode_tree_mode(const char *str, unsigned int *modep)\n {\n \tunsigned char c;\n \tunsigned int mode = 0;\n@@ -28,7 +28,7 @@ static void decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned\n \tif (size < 24 || buf[size - 21])\n \t\tdie(\"corrupt tree file\");\n \n-\tpath = get_mode(buf, &mode);\n+\tpath = decode_tree_mode(buf, &mode);\n \tif (!path || !*path)\n \t\tdie(\"corrupt tree file\");\n \tlen = strlen(path) + 1;\ndiff --git a/tree-walk.h b/tree-walk.h\nindex 7e3e0b5..6bbde1c 100644\n--- a/tree-walk.h\n+++ b/tree-walk.h\n@@ -13,6 +13,8 @@ struct tree_desc {\n \tunsigned int size;\n };\n \n+const char *decode_tree_mode(const char *str, unsigned int *modep);\n+\n static inline const unsigned char *tree_entry_extract(struct tree_desc *desc, const char **pathp, unsigned int *modep)\n {\n \t*pathp = desc->entry.path;\n-- \n1.7.3.37.gb6088b\n"},{"id":"152351","messageId":"1286108511-55876-5-git-send-email-tom@dbservice.com","threadId":"25319","inReplyTo":"4CA86A12.6080905@dbservice.com","subject":"[PATCH 5/6] Introduce the git fast-import-helper","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2010-10-03T12:21:50Z","receivedAt":"2010-10-03T12:21:50Z","isPatch":true,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"The g-f-i-h is a heavily modified (and simplified where possible) copy\nof git-fast-import. It has a few very important changes which make it\nsuitable to be used in the new generation of remote helpers.\n\n1) It does not update refs itself. Instead, for each 'mark' it sees, it\n   writes the SHA1 of the corresponding git object to stdout. The remote\n   helper can read this data and pass it along to core git for example.\n\n2) It does not read/write mark files itself. Managing the marks is now\n   up to the application which uses g-f-i-h. To support that, a new\n   command was added: 'mark <name> <sha1>'. It can be used to feed\n   g-f-i-h with existing marks from earlier sessions. Also, marks\n   can now be arbitrary strings and not just numbers. This allows remote\n   helpers to use for example whole revision strings (r42 for svn or\n   mercurial changeset IDs).\n\n3) Memory management has been significantly simplified. No more pools\n   and custom allocators. It uses plain malloc/free. Uses `struct\n   hash_table` instead of custom data structures. This may make it\n   a bit slower than the original, but on the other hand it reduces\n   the complexity of the source code.\n\nSigned-off-by: Tomas Carnecky <tom@dbservice.com>\n---\n .gitignore           |    1 +\n Makefile             |    1 +\n fast-import-helper.c | 2201 ++++++++++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 2203 insertions(+), 0 deletions(-)\n create mode 100644 fast-import-helper.c\n\ndiff --git a/.gitignore b/.gitignore\nindex 20560b8..c8aa8c7 100644\n--- a/.gitignore\n+++ b/.gitignore\n@@ -42,6 +42,7 @@\n /git-describe\n /git-fast-export\n /git-fast-import\n+/git-fast-import-helper\n /git-fetch\n /git-fetch--tool\n /git-fetch-pack\ndiff --git a/Makefile b/Makefile\nindex 8a56b9a..f8a9c40 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -402,6 +402,7 @@ EXTRA_PROGRAMS =\n PROGRAMS += $(EXTRA_PROGRAMS)\n \n PROGRAM_OBJS += fast-import.o\n+PROGRAM_OBJS += fast-import-helper.o\n PROGRAM_OBJS += imap-send.o\n PROGRAM_OBJS += shell.o\n PROGRAM_OBJS += show-index.o\ndiff --git a/fast-import-helper.c b/fast-import-helper.c\nnew file mode 100644\nindex 0000000..b1f86a0\n--- /dev/null\n+++ b/fast-import-helper.c\n@@ -0,0 +1,2201 @@\n+\n+#include \"builtin.h\"\n+#include \"cache.h\"\n+#include \"object.h\"\n+#include \"blob.h\"\n+#include \"tree.h\"\n+#include \"commit.h\"\n+#include \"delta.h\"\n+#include \"pack.h\"\n+#include \"refs.h\"\n+#include \"csum-file.h\"\n+#include \"quote.h\"\n+#include \"exec_cmd.h\"\n+#include \"hash.h\"\n+#include \"tree-walk.h\"\n+\n+#define DEPTH_BITS 13\n+#define MAX_DEPTH ((1<<DEPTH_BITS)-1)\n+\n+struct fi_object\n+{\n+\tstruct fi_object *next;\n+\tstruct pack_idx_entry idx;\n+\tuint32_t type : TYPE_BITS,\n+\t\tdepth : DEPTH_BITS;\n+};\n+\n+struct last_object\n+{\n+\tstruct strbuf data;\n+\toff_t offset;\n+\tunsigned int depth;\n+\tunsigned no_swap : 1;\n+};\n+\n+/* An atom uniquely identifies an array of data (usually strings). */\n+struct fi_atom\n+{\n+\tstruct fi_atom *next;\n+\tunsigned short len;\n+\tchar data[FLEX_ARRAY];\n+};\n+\n+/* The stream can use marks to mark objects (blobs, commits) with a unique\n+ * string. After the object is added to the object database and we know its\n+ * SHA1, we report that to stdout. */\n+struct fi_mark\n+{\n+\tstruct fi_mark *next;\n+\tstruct fi_atom *atom;\n+\tunsigned char sha1[20];\n+};\n+\n+struct fi_tree;\n+struct fi_tree_entry\n+{\n+\tstruct fi_tree *tree;\n+\t\n+\tstruct fi_atom *name;\n+\tstruct fi_tree_entry_ms\n+\t{\n+\t\tunsigned int mode;\n+\t\tunsigned char sha1[20];\n+\t} versions[2];\n+};\n+\n+struct fi_tree\n+{\n+\tstruct fi_tree *next;\n+\tunsigned int entry_capacity;\n+\tunsigned int entry_count;\n+\tunsigned int delta_depth;\n+\tstruct fi_tree_entry *entries[FLEX_ARRAY];\n+};\n+\n+struct fi_branch\n+{\n+\tstruct fi_branch *table_next_branch;\n+\tstruct fi_branch *active_next_branch;\n+\tconst char *name;\n+\tstruct fi_tree_entry branch_tree;\n+\tuintmax_t num_notes;\n+\tunsigned active : 1;\n+\tunsigned char sha1[20];\n+};\n+\n+struct fi_tag\n+{\n+\tstruct fi_tag *next_tag;\n+\tconst char *name;\n+\tunsigned char sha1[20];\n+};\n+\n+struct hash_list\n+{\n+\tstruct hash_list *next;\n+\tunsigned char sha1[20];\n+};\n+\n+typedef enum {\n+\tWHENSPEC_RAW = 1,\n+\tWHENSPEC_RFC2822,\n+\tWHENSPEC_NOW\n+} whenspec_type;\n+\n+/* Configured limits on output */\n+static unsigned long max_depth = 10;\n+static off_t max_packsize = 32 * 1024 * 1024;\n+static uintmax_t big_file_threshold = 512 * 1024 * 1024;\n+static int force_update;\n+static int pack_compression_level = Z_DEFAULT_COMPRESSION;\n+static int pack_compression_seen;\n+\n+/* The .pack file being generated */\n+static struct sha1file *pack_file;\n+static struct packed_git *pack_data;\n+static off_t pack_size;\n+\n+/* Our last blob */\n+static struct last_object last_blob = { STRBUF_INIT, 0, 0, 0 };\n+\n+/* Branch data */\n+static struct hash_table branches;\n+static unsigned long max_active_branches = 5;\n+static unsigned long cur_active_branches;\n+static struct fi_branch *active_branches;\n+\n+/* Tag data */\n+static struct hash_table tags;\n+\n+/* Input stream parsing */\n+static whenspec_type whenspec = WHENSPEC_RAW;\n+static struct strbuf command_buf = STRBUF_INIT;\n+\n+static void end_packfile(void);\n+static struct fi_atom *to_atom(const char *s, unsigned short len);\n+static enum object_type fi_sha1_object_type(unsigned char sha1[20]);\n+\n+\n+static int git_pack_config(const char *k, const char *v, void *cb)\n+{\n+\tif (!strcmp(k, \"pack.depth\")) {\n+\t\tmax_depth = git_config_int(k, v);\n+\t\tif (max_depth > MAX_DEPTH)\n+\t\t\tmax_depth = MAX_DEPTH;\n+\t\treturn 0;\n+\t}\n+\tif (!strcmp(k, \"pack.compression\")) {\n+\t\tint level = git_config_int(k, v);\n+\t\tif (level == -1)\n+\t\t\tlevel = Z_DEFAULT_COMPRESSION;\n+\t\telse if (level < 0 || level > Z_BEST_COMPRESSION)\n+\t\t\tdie(\"bad pack compression level %d\", level);\n+\t\tpack_compression_level = level;\n+\t\tpack_compression_seen = 1;\n+\t\treturn 0;\n+\t}\n+\tif (!strcmp(k, \"pack.indexversion\")) {\n+\t\tpack_idx_default_version = git_config_int(k, v);\n+\t\tif (pack_idx_default_version > 2)\n+\t\t\tdie(\"bad pack.indexversion=%\"PRIu32,\n+\t\t\t    pack_idx_default_version);\n+\t\treturn 0;\n+\t}\n+\tif (!strcmp(k, \"pack.packsizelimit\")) {\n+\t\tmax_packsize = git_config_ulong(k, v);\n+\t\treturn 0;\n+\t}\n+\tif (!strcmp(k, \"core.bigfilethreshold\")) {\n+\t\tlong n = git_config_int(k, v);\n+\t\tbig_file_threshold = 0 < n ? n : 0;\n+\t}\n+\treturn git_default_config(k, v, cb);\n+}\n+\n+static NORETURN void die_nicely(const char *err, va_list params)\n+{\n+\tstatic int zombie;\n+\tchar message[2 * PATH_MAX];\n+\n+\tvsnprintf(message, sizeof(message), err, params);\n+\tfputs(\"fatal: \", stderr);\n+\tfputs(message, stderr);\n+\tfputc('\\n', stderr);\n+\n+\tif (!zombie) {\n+\t\tzombie = 1;\n+\t\tend_packfile();\n+\t}\n+\texit(128);\n+}\n+\n+\n+/**\n+ * Misc methods\n+ */\n+\n+static unsigned int hc_str(const char *s, size_t len)\n+{\n+\tunsigned int r = 0;\n+\twhile (len-- > 0)\n+\t\tr = r * 31 + *s++;\n+\treturn r;\n+}\n+\n+\n+/**\n+ * Command buffer\n+ */\n+\n+static int read_next_command(void)\n+{\n+\tstatic int stdin_eof = 0;\n+\n+\tif (stdin_eof) {\n+\t\treturn EOF;\n+\t}\n+\n+\tdo {\n+\t\tstrbuf_detach(&command_buf, NULL);\n+\t\tstdin_eof = strbuf_getline(&command_buf, stdin, '\\n');\n+\t\tif (stdin_eof)\n+\t\t\treturn EOF;\n+\t} while (command_buf.buf[0] == '#');\n+\t//fprintf(stderr, \"Command: %s\\n\", command_buf.buf);\n+\n+\treturn 0;\n+}\n+\n+static void skip_optional_lf(void)\n+{\n+\tif (command_buf.buf[0] == '\\n')\n+\t\tread_next_command();\n+}\n+\n+\n+/**\n+ * Object cache\n+ */\n+\n+static struct hash_table objects;\n+\n+static struct fi_object *new_object(unsigned char *sha1)\n+{\n+\tstruct fi_object *e = malloc(sizeof(struct fi_object));\n+\thashcpy(e->idx.sha1, sha1);\n+\treturn e;\n+}\n+\n+static struct fi_object *find_object(unsigned char *sha1)\n+{\n+\tunsigned int hash = *(unsigned int *) sha1;\n+\tstruct fi_object *head = lookup_hash(hash, &objects);\n+\n+\twhile (head) {\n+\t\tif (!hashcmp(sha1, head->idx.sha1))\n+\t\t\treturn head;\n+\n+\t\thead = head->next;\n+\t}\n+\n+\treturn NULL;\n+}\n+\n+static struct fi_object *insert_object(unsigned char *sha1)\n+{\n+\tunsigned int hash = *(unsigned int *) sha1;\n+\tstruct fi_object *oe, *head = lookup_hash(hash, &objects);\n+\tvoid **ptr;\n+\n+\toe = head;\n+\twhile (oe) {\n+\t\tif (!hashcmp(sha1, oe->idx.sha1))\n+\t\t\treturn oe;\n+\n+\t\toe = oe->next;\n+\t}\n+\n+\toe = new_object(sha1);\n+\toe->next = head;\n+\toe->idx.offset = 0;\n+\n+\tptr = insert_hash(hash, oe, &objects);\n+\tif (ptr)\n+\t\t*ptr = oe;\n+\n+\treturn oe;\n+}\n+\n+static int free_object(void *data)\n+{\n+\tfree(data);\n+\treturn 0;\n+}\n+\n+\n+/**\n+ * Atoms\n+ */\n+\n+static struct hash_table atoms;\n+\n+static struct fi_atom *to_atom(const char *s, unsigned short len)\n+{\n+\tunsigned int hash = hc_str(s, len);\n+\tstruct fi_atom *atom, *head = lookup_hash(hash, &atoms);\n+\n+\tatom = head;\n+\twhile (atom) {\n+\t\tif (atom->len == len && !strncmp(s, atom->data, len))\n+\t\t\treturn atom;\n+\t\tatom = atom->next;\n+\t}\n+\n+\tatom = malloc(sizeof(struct fi_atom) + len + 1);\n+\tatom->len = len;\n+\tstrncpy(atom->data, s, len);\n+\tatom->data[len] = 0;\n+\tatom->next = head;\n+\n+\tinsert_hash(hash, atom, &atoms);\n+\n+\treturn atom;\n+}\n+\n+/**\n+ * Mark cache\n+ */\n+\n+static struct hash_table marks;\n+\n+static void insert_mark(struct fi_atom *atom, unsigned char sha1[20])\n+{\n+\tunsigned int hash = hc_str(atom->data, atom->len);\n+\tstruct fi_mark *mark, *head = lookup_hash(hash, &marks);\n+\tvoid **ptr;\n+\t\n+\t/* If the mark already exists, overwrite its value. */\n+\tmark = head;\n+\twhile (mark) {\n+\t\tif (atom == mark->atom) {\n+\t\t\thashcpy(mark->sha1, sha1);\n+\t\t\tgoto out;\n+\t\t}\n+\t\t\n+\t}\n+\t\n+\tmark = malloc(sizeof(struct fi_mark));\n+\tmark->next = head;\n+\tmark->atom = atom;\n+\thashcpy(mark->sha1, sha1);\n+\n+\tptr = insert_hash(hash, mark, &marks);\n+\tif (ptr)\n+\t\t*ptr = mark;\n+\n+out:\n+\t/* Dump the mark mapping to stdout. */\n+\tfprintf(stdout, \"mark :%s %s\\n\", atom->data, sha1_to_hex(sha1));\n+\t//fprintf(stderr, \"mark :%s %s\\n\", atom->data, sha1_to_hex(sha1));\n+\tfflush(stdout); fflush(stderr);\n+}\n+\n+\n+/* Parse an atom in the string, set *atom and return the end of the atom. */\n+static const char *parse_mark_to_atom(const char *s, struct fi_atom **atom)\n+{\n+\tconst char *end = strchr(s, ' ');\n+\tunsigned int len = end ? end - s : strlen(s);\n+\t*atom = to_atom(s, len);\n+\treturn s + len;\n+}\n+\n+static const char *find_mark(const char *mark, unsigned char sha1[20])\n+{\n+\tstruct fi_atom *atom;\n+\tconst char *end = parse_mark_to_atom(mark, &atom);\n+\tunsigned int hash = hc_str(mark, end - mark);\n+\tstruct fi_mark *head = lookup_hash(hash, &marks);\n+\n+\twhile (head) {\n+\t\tif (head->atom == atom) {\n+\t\t\thashcpy(sha1, head->sha1);\n+\t\t\treturn end;\n+\t\t}\n+\t\thead = head->next;\n+\t}\n+\t\n+\tdie(\"Did not find mark '%s'\", mark);\n+}\n+\n+\n+/**\n+ * Branch cache\n+ */\n+\n+static struct fi_branch *lookup_branch(const char *name)\n+{\n+\tunsigned int hash = hc_str(name, strlen(name));\n+\tstruct fi_branch *b = lookup_hash(hash, &branches);\n+\n+\twhile (b) {\n+\t\tif (!strcmp(name, b->name))\n+\t\t\treturn b;\n+\t\tb = b->table_next_branch;\n+\t}\n+\n+\treturn NULL;\n+}\n+\n+static struct fi_branch *new_branch(const char *name)\n+{\n+\tunsigned int hash = hc_str(name, strlen(name));\n+\tstruct fi_branch *b = lookup_branch(name);\n+\tvoid **ptr;\n+\n+\tif (b)\n+\t\tdie(\"Invalid attempt to create duplicate branch: %s\", name);\n+\tswitch (check_ref_format(name)) {\n+\tcase 0: break; /* its valid */\n+\tcase CHECK_REF_FORMAT_ONELEVEL:\n+\t\tbreak; /* valid, but too few '/', allow anyway */\n+\tdefault:\n+\t\tdie(\"Branch name doesn't conform to GIT standards: %s\", name);\n+\t}\n+\n+\tb = calloc(1, sizeof(struct fi_branch));\n+\tb->name = strdup(name);\n+\tb->table_next_branch = lookup_hash(hash, &branches);\n+\tb->branch_tree.versions[0].mode = S_IFDIR;\n+\tb->branch_tree.versions[1].mode = S_IFDIR;\n+\tb->num_notes = 0;\n+\tb->active = 0;\n+\tptr = insert_hash(hash, b, &branches);\n+\tif (ptr)\n+\t\t*ptr = b;\n+\t\t\n+\treturn b;\n+}\n+\n+\n+/**\n+ * Tree cache\n+ */\n+\n+static struct fi_tree *new_tree_content(unsigned int cnt)\n+{\n+\tstruct fi_tree *t;\n+\n+\tt = malloc(sizeof(*t) + sizeof(t->entries[0]) * cnt);\n+\tt->next = NULL;\n+\tt->entry_capacity = cnt;\n+\tt->entry_count = 0;\n+\tt->delta_depth = 0;\n+\n+\treturn t;\n+}\n+\n+static void release_tree_entry(struct fi_tree_entry *e);\n+static void release_tree_content(struct fi_tree *t)\n+{\n+\tfree(t);\n+}\n+\n+static void release_tree_content_recursive(struct fi_tree *t)\n+{\n+\tunsigned int i;\n+\tfor (i = 0; i < t->entry_count; i++)\n+\t\trelease_tree_entry(t->entries[i]);\n+\trelease_tree_content(t);\n+}\n+\n+static struct fi_tree *grow_tree_content(\n+\tstruct fi_tree *t,\n+\tint amt)\n+{\n+\tstruct fi_tree *r = new_tree_content(t->entry_count + amt);\n+\tr->entry_count = t->entry_count;\n+\tr->delta_depth = t->delta_depth;\n+\tmemcpy(r->entries,t->entries,t->entry_count*sizeof(t->entries[0]));\n+\trelease_tree_content(t);\n+\treturn r;\n+}\n+\n+static struct fi_tree_entry *new_tree_entry(void)\n+{\n+\treturn xmalloc(sizeof(struct fi_tree_entry));\n+}\n+\n+static void release_tree_entry(struct fi_tree_entry *e)\n+{\n+\tfree(e);\n+}\n+\n+static struct fi_tree *dup_tree_content(struct fi_tree *s)\n+{\n+\tstruct fi_tree *d;\n+\tstruct fi_tree_entry *a, *b;\n+\tunsigned int i;\n+\n+\tif (!s)\n+\t\treturn NULL;\n+\td = new_tree_content(s->entry_count);\n+\tfor (i = 0; i < s->entry_count; i++) {\n+\t\ta = s->entries[i];\n+\t\tb = new_tree_entry();\n+\t\tmemcpy(b, a, sizeof(*a));\n+\t\tif (a->tree && is_null_sha1(b->versions[1].sha1))\n+\t\t\tb->tree = dup_tree_content(a->tree);\n+\t\telse\n+\t\t\tb->tree = NULL;\n+\t\td->entries[i] = b;\n+\t}\n+\td->entry_count = s->entry_count;\n+\td->delta_depth = s->delta_depth;\n+\n+\treturn d;\n+}\n+\n+\n+/**\n+ * Pack file handling\n+ */\n+\n+static void start_packfile(void)\n+{\n+\tstatic char tmpfile[PATH_MAX];\n+\tstruct pack_header hdr;\n+\tint pack_fd;\n+\n+\tpack_fd = odb_mkstemp(tmpfile, sizeof(tmpfile), \"pack/tmp_pack_XXXXXX\");\n+\tpack_data = xcalloc(1, sizeof(struct packed_git) + strlen(tmpfile) + 2);\n+\tstrcpy(pack_data->pack_name, tmpfile);\n+\tpack_data->pack_fd = pack_fd;\n+\tpack_file = sha1fd(pack_fd, pack_data->pack_name);\n+\n+\thdr.hdr_signature = htonl(PACK_SIGNATURE);\n+\thdr.hdr_version = htonl(2);\n+\thdr.hdr_entries = 0;\n+\tsha1write(pack_file, &hdr, sizeof(hdr));\n+\n+\tpack_size = sizeof(hdr);\n+\n+\tfor_each_hash(&objects, free_object);\n+\tfree_hash(&objects);\n+}\n+\n+static struct pack_idx_entry **create_index_iter;\n+static int fi_add_to_index(void *data)\n+{\n+\tstruct fi_object *oe = data;\n+\twhile (oe) {\n+\t\t*create_index_iter++ = &oe->idx;\n+\t\toe = oe->next;\n+\t}\n+\t\n+\treturn 0;\n+}\n+\n+static const char *create_index(void)\n+{\n+\tconst char *tmpfile;\n+\tstruct pack_idx_entry **idx, **last;\n+\n+\t/* Build the table of object IDs. */\n+\tidx = xmalloc(objects.nr * sizeof(*idx));\n+\tcreate_index_iter = idx;\n+\tfor_each_hash(&objects, fi_add_to_index);\n+\tlast = idx + objects.nr;\n+\tif (create_index_iter != last)\n+\t\tdie(\"internal consistency error creating the index\");\n+\n+\ttmpfile = write_idx_file(NULL, idx, objects.nr, pack_data->sha1);\n+\tfree(idx);\n+\treturn tmpfile;\n+}\n+\n+static char *keep_pack(const char *curr_index_name)\n+{\n+\tstatic char name[PATH_MAX];\n+\tstatic const char *keep_msg = \"fast-import\";\n+\tint keep_fd;\n+\n+\tkeep_fd = odb_pack_keep(name, sizeof(name), pack_data->sha1);\n+\tif (keep_fd < 0)\n+\t\tdie_errno(\"cannot create keep file\");\n+\twrite_or_die(keep_fd, keep_msg, strlen(keep_msg));\n+\tif (close(keep_fd))\n+\t\tdie_errno(\"failed to write keep file\");\n+\n+\tsnprintf(name, sizeof(name), \"%s/pack/pack-%s.pack\",\n+\t\t get_object_directory(), sha1_to_hex(pack_data->sha1));\n+\tif (move_temp_to_file(pack_data->pack_name, name))\n+\t\tdie(\"cannot store pack file\");\n+\n+\tsnprintf(name, sizeof(name), \"%s/pack/pack-%s.idx\",\n+\t\t get_object_directory(), sha1_to_hex(pack_data->sha1));\n+\tif (move_temp_to_file(curr_index_name, name))\n+\t\tdie(\"cannot store index file\");\n+\tfree((void *)curr_index_name);\n+\treturn name;\n+}\n+\n+static void end_packfile(void)\n+{\n+\tstruct packed_git *new_p;\n+\n+\tclear_delta_base_cache();\n+\tif (objects.nr) {\n+\t\tunsigned char cur_pack_sha1[20];\n+\t\tchar *idx_name;\n+\n+\t\tclose_pack_windows(pack_data);\n+\t\tsha1close(pack_file, cur_pack_sha1, 0);\n+\t\tfixup_pack_header_footer(pack_data->pack_fd, pack_data->sha1,\n+\t\t\t\t    pack_data->pack_name, objects.nr,\n+\t\t\t\t    cur_pack_sha1, pack_size);\n+\t\tclose(pack_data->pack_fd);\n+\t\tidx_name = keep_pack(create_index());\n+\n+\t\t/* Register the packfile with core git's machinery. */\n+\t\tnew_p = add_packed_git(idx_name, strlen(idx_name), 1);\n+\t\tif (!new_p)\n+\t\t\tdie(\"core git rejected index %s\", idx_name);\n+\t\tinstall_packed_git(new_p);\n+\t}\n+\telse {\n+\t\tclose(pack_data->pack_fd);\n+\t\tunlink_or_warn(pack_data->pack_name);\n+\t}\n+\tfree(pack_data);\n+\n+\t/* We can't carry a delta across packfiles. */\n+\tstrbuf_release(&last_blob.data);\n+\tlast_blob.offset = 0;\n+\tlast_blob.depth = 0;\n+}\n+\n+static void cycle_packfile(void)\n+{\n+\tend_packfile();\n+\tstart_packfile();\n+}\n+\n+\n+/**\n+ * Methods for storing objects\n+ */\n+\n+static int store_object(\n+\tenum object_type type,\n+\tstruct strbuf *dat,\n+\tstruct last_object *last,\n+\tunsigned char *sha1out,\n+\tstruct fi_atom *atom)\n+{\n+\tvoid *out, *delta = NULL;\n+\tstruct fi_object *e;\n+\tunsigned char hdr[96];\n+\tunsigned char sha1[20];\n+\tunsigned long hdrlen, deltalen;\n+\tgit_SHA_CTX c;\n+\tz_stream s;\n+\n+\n+\t/* Construct the header. */\n+\thdrlen = sprintf((char *)hdr,\"%s %lu\", typename(type),\n+\t\t(unsigned long)dat->len) + 1;\n+\t\n+\t/* Compute the hash of the object. */\n+\tgit_SHA1_Init(&c);\n+\tgit_SHA1_Update(&c, hdr, hdrlen);\n+\tgit_SHA1_Update(&c, dat->buf, dat->len);\n+\tgit_SHA1_Final(sha1, &c);\n+\n+\tif (sha1out)\n+\t\thashcpy(sha1out, sha1);\n+\tif (atom)\n+\t\tinsert_mark(atom, sha1);\n+\n+\t/* Determine if we should auto-checkpoint. */\n+\tif ((max_packsize && (pack_size + 60 + dat->len + hdrlen) > max_packsize)\n+\t\t|| (pack_size + 60 + dat->len + hdrlen) < pack_size) {\n+\t\tcycle_packfile();\n+\t}\n+\t\n+\t/* Insert the object into our cache, return if it already exists. */\n+\te = insert_object(sha1);\n+\n+\tif (e->idx.offset)\n+\t\treturn 1;\n+\n+\te->type = type;\n+\te->idx.offset = pack_size;\n+\n+\tmemset(&s, 0, sizeof(s));\n+\tdeflateInit(&s, pack_compression_level);\n+\t\n+\t/* Compress the data, try to create a delta against the last object. */\n+\tif (last && last->data.buf && last->depth < max_depth && dat->len > 20) {\n+\t\tdelta = diff_delta(last->data.buf, last->data.len,\n+\t\t\tdat->buf, dat->len, &deltalen, dat->len - 20);\n+\t}\n+\n+\t/* diff_delta() above can fail and return NULL! */\n+\tif (delta) {\n+\t\ts.next_in = delta;\n+\t\ts.avail_in = deltalen;\n+\t} else {\n+\t\ts.next_in = (void *)dat->buf;\n+\t\ts.avail_in = dat->len;\n+\t}\n+\n+\ts.avail_out = deflateBound(&s, s.avail_in);\n+\ts.next_out = out = xmalloc(s.avail_out);\n+\twhile (deflate(&s, Z_FINISH) == Z_OK)\n+\t\t; /* nothing */\n+\tdeflateEnd(&s);\n+\n+\t/* Write the object to the packfile. */\n+\tcrc32_begin(pack_file);\n+\tif (delta) {\n+\t\toff_t ofs = e->idx.offset - last->offset;\n+\t\tunsigned pos = sizeof(hdr) - 1;\n+\n+\t\te->depth = last->depth + 1;\n+\n+\t\thdrlen = encode_in_pack_object_header(OBJ_OFS_DELTA, deltalen, hdr);\n+\t\tsha1write(pack_file, hdr, hdrlen);\n+\t\tpack_size += hdrlen;\n+\n+\t\thdr[pos] = ofs & 127;\n+\t\twhile (ofs >>= 7)\n+\t\t\thdr[--pos] = 128 | (--ofs & 127);\n+\t\tsha1write(pack_file, hdr + pos, sizeof(hdr) - pos);\n+\t\tpack_size += sizeof(hdr) - pos;\n+\n+\t\tfree(delta);\n+\t} else {\n+\t\te->depth = 0;\n+\t\thdrlen = encode_in_pack_object_header(type, dat->len, hdr);\n+\t\tsha1write(pack_file, hdr, hdrlen);\n+\t\tpack_size += hdrlen;\n+\t}\n+\n+\tsha1write(pack_file, out, s.total_out);\n+\tpack_size += s.total_out;\n+\tfree(out);\n+\n+\t/* Update the cached object. */\n+\te->idx.crc32 = crc32_end(pack_file);\n+\n+\tif (last) {\n+\t\tif (last->no_swap) {\n+\t\t\tlast->data = *dat;\n+\t\t} else {\n+\t\t\tstrbuf_swap(&last->data, dat);\n+\t\t}\n+\t\tlast->offset = e->idx.offset;\n+\t\tlast->depth = e->depth;\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static void truncate_pack(off_t to, git_SHA_CTX *ctx)\n+{\n+\tif (ftruncate(pack_data->pack_fd, to)\n+\t || lseek(pack_data->pack_fd, to, SEEK_SET) != to)\n+\t\tdie_errno(\"cannot truncate pack to skip duplicate\");\n+\tpack_size = to;\n+\n+\t/* yes this is a layering violation */\n+\tpack_file->total = to;\n+\tpack_file->offset = 0;\n+\tpack_file->ctx = *ctx;\n+}\n+\n+static void stream_blob(uintmax_t len, unsigned char *sha1out, struct fi_atom *atom)\n+{\n+\tsize_t in_sz = 64 * 1024, out_sz = 64 * 1024;\n+\tunsigned char *in_buf = xmalloc(in_sz);\n+\tunsigned char *out_buf = xmalloc(out_sz);\n+\tstruct fi_object *e;\n+\tunsigned char sha1[20];\n+\tunsigned long hdrlen;\n+\toff_t offset;\n+\tgit_SHA_CTX c;\n+\tgit_SHA_CTX pack_file_ctx;\n+\tz_stream s;\n+\tint status = Z_OK;\n+\n+\t/* Determine if we should auto-checkpoint. */\n+\tif ((max_packsize && (pack_size + 60 + len) > max_packsize)\n+\t\t|| (pack_size + 60 + len) < pack_size)\n+\t\tcycle_packfile();\n+\n+\toffset = pack_size;\n+\n+\t/* preserve the pack_file SHA1 ctx in case we have to truncate later */\n+\tsha1flush(pack_file);\n+\tpack_file_ctx = pack_file->ctx;\n+\n+\thdrlen = snprintf((char *)out_buf, out_sz, \"blob %\" PRIuMAX, len) + 1;\n+\tif (out_sz <= hdrlen)\n+\t\tdie(\"impossibly large object header\");\n+\n+\tgit_SHA1_Init(&c);\n+\tgit_SHA1_Update(&c, out_buf, hdrlen);\n+\n+\tcrc32_begin(pack_file);\n+\n+\tmemset(&s, 0, sizeof(s));\n+\tdeflateInit(&s, pack_compression_level);\n+\n+\thdrlen = encode_in_pack_object_header(OBJ_BLOB, len, out_buf);\n+\tif (out_sz <= hdrlen)\n+\t\tdie(\"impossibly large object header\");\n+\n+\ts.next_out = out_buf + hdrlen;\n+\ts.avail_out = out_sz - hdrlen;\n+\n+\twhile (status != Z_STREAM_END) {\n+\t\tif (0 < len && !s.avail_in) {\n+\t\t\tsize_t cnt = in_sz < len ? in_sz : (size_t)len;\n+\t\t\tsize_t n = fread(in_buf, 1, cnt, stdin);\n+\t\t\tif (!n && feof(stdin))\n+\t\t\t\tdie(\"EOF in data (%\" PRIuMAX \" bytes remaining)\", len);\n+\n+\t\t\tgit_SHA1_Update(&c, in_buf, n);\n+\t\t\ts.next_in = in_buf;\n+\t\t\ts.avail_in = n;\n+\t\t\tlen -= n;\n+\t\t}\n+\n+\t\tstatus = deflate(&s, len ? 0 : Z_FINISH);\n+\n+\t\tif (!s.avail_out || status == Z_STREAM_END) {\n+\t\t\tsize_t n = s.next_out - out_buf;\n+\t\t\tsha1write(pack_file, out_buf, n);\n+\t\t\tpack_size += n;\n+\t\t\ts.next_out = out_buf;\n+\t\t\ts.avail_out = out_sz;\n+\t\t}\n+\n+\t\tswitch (status) {\n+\t\tcase Z_OK:\n+\t\tcase Z_BUF_ERROR:\n+\t\tcase Z_STREAM_END:\n+\t\t\tcontinue;\n+\t\tdefault:\n+\t\t\tdie(\"unexpected deflate failure: %d\", status);\n+\t\t}\n+\t}\n+\tdeflateEnd(&s);\n+\tgit_SHA1_Final(sha1, &c);\n+\n+\tif (sha1out)\n+\t\thashcpy(sha1out, sha1);\n+\tif (atom)\n+\t\tinsert_mark(atom, sha1);\n+\n+\te = insert_object(sha1);\n+\n+\tif (e->idx.offset) {\n+\t\ttruncate_pack(offset, &pack_file_ctx);\n+\t} else {\n+\t\te->depth = 0;\n+\t\te->type = OBJ_BLOB;\n+\t\te->idx.offset = offset;\n+\t\te->idx.crc32 = crc32_end(pack_file);\n+\t}\n+\n+\tfree(in_buf);\n+\tfree(out_buf);\n+}\n+\n+/* All calls must be guarded by find_object() or find_mark() to\n+ * ensure the 'struct fi_object' passed was written by this\n+ * process instance.  We unpack the entry by the offset, avoiding\n+ * the need for the corresponding .idx file.  This unpacking rule\n+ * works because we only use OBJ_REF_DELTA within the packfiles\n+ * created by fast-import.\n+ *\n+ * oe must not be NULL.  Such an oe usually comes from giving\n+ * an unknown SHA-1 to find_object() or an undefined mark to\n+ * find_mark().  Callers must test for this condition and use\n+ * the standard read_sha1_file() when it happens.\n+ *\n+ * oe->pack_id must not be MAX_PACK_ID.  Such an oe is usually from\n+ * find_mark(), where the mark was reloaded from an existing marks\n+ * file and is referencing an object that this fast-import process\n+ * instance did not write out to a packfile.  Callers must test for\n+ * this condition and use read_sha1_file() instead.\n+ */\n+static void *fi_unpack_entry(\n+\tstruct fi_object *oe,\n+\tenum object_type *type,\n+\tunsigned long *size)\n+{\n+\tif (pack_data->pack_size < (pack_size + 20)) {\n+\t\t/* The object is stored in the packfile we are writing to\n+\t\t * and we have modified it since the last time we scanned\n+\t\t * back to read a previously written object.  If an old\n+\t\t * window covered [p->pack_size, p->pack_size + 20) its\n+\t\t * data is stale and is not valid.  Closing all windows\n+\t\t * and updating the packfile length ensures we can read\n+\t\t * the newly written data.\n+\t\t */\n+\t\tclose_pack_windows(pack_data);\n+\t\tsha1flush(pack_file);\n+\n+\t\t/* We have to offer 20 bytes additional on the end of\n+\t\t * the packfile as the core unpacker code assumes the\n+\t\t * footer is present at the file end and must promise\n+\t\t * at least 20 bytes within any window it maps.  But\n+\t\t * we don't actually create the footer here.\n+\t\t */\n+\t\tpack_data->pack_size = pack_size + 20;\n+\t}\n+\t\n+\treturn unpack_entry(pack_data, oe->idx.offset, type, size);\n+}\n+\n+/* Same as read_sha1_file() except that it first looks in our local cache\n+ * which holds objects from our current pack. Git doesn't know anything\n+ * about those objects until we finish the pack and register it with Git. */\n+static void *fi_read_sha1_file(unsigned char sha1[20], enum object_type *type,\n+\t\t\t\t\t\t\t\t\tunsigned long *size)\n+{\n+\tstruct fi_object *obj = find_object(sha1);\n+\tif (obj)\n+\t\treturn fi_unpack_entry(obj, type, size);\n+\n+\treturn read_sha1_file(sha1, type, size);\n+}\n+\n+static enum object_type fi_sha1_object_type(unsigned char sha1[20])\n+{\n+\tstruct fi_object *obj = find_object(sha1);\n+\tif (obj) {\n+\t\treturn obj->type;\n+\t}\n+\n+\treturn sha1_object_info(sha1, NULL);\n+}\n+\n+\n+/**\n+ * Tree handling\n+ */\n+\n+static void load_tree(struct fi_tree_entry *root)\n+{\n+\tunsigned char *sha1 = root->versions[1].sha1;\n+\tstruct fi_tree *t;\n+\tunsigned long size;\n+\tenum object_type type;\n+\tchar *buf;\n+\tconst char *c;\n+\n+\troot->tree = t = new_tree_content(8);\n+\tif (is_null_sha1(sha1))\n+\t\treturn;\n+\n+\tbuf = fi_read_sha1_file(sha1, &type, &size);\n+\tif (!buf || type != OBJ_TREE)\n+\t\tdie(\"Can't load tree %s\", sha1_to_hex(sha1));\n+\n+\tc = buf;\n+\twhile (c != (buf + size)) {\n+\t\tstruct fi_tree_entry *e = new_tree_entry();\n+\n+\t\tif (t->entry_count == t->entry_capacity)\n+\t\t\troot->tree = t = grow_tree_content(t, t->entry_count);\n+\t\tt->entries[t->entry_count++] = e;\n+\n+\t\te->tree = NULL;\n+\t\tc = decode_tree_mode(c, &e->versions[1].mode);\n+\t\tif (!c)\n+\t\t\tdie(\"Corrupt mode in %s\", sha1_to_hex(sha1));\n+\t\te->versions[0].mode = e->versions[1].mode;\n+\t\te->name = to_atom(c, strlen(c));\n+\t\tc += e->name->len + 1;\n+\t\thashcpy(e->versions[0].sha1, (unsigned char *)c);\n+\t\thashcpy(e->versions[1].sha1, (unsigned char *)c);\n+\t\tc += 20;\n+\t}\n+\tfree(buf);\n+}\n+\n+static int tecmp0 (const void *_a, const void *_b)\n+{\n+\tstruct fi_tree_entry *a = *((struct fi_tree_entry**)_a);\n+\tstruct fi_tree_entry *b = *((struct fi_tree_entry**)_b);\n+\treturn base_name_compare(\n+\t\ta->name->data, a->name->len, a->versions[0].mode,\n+\t\tb->name->data, b->name->len, b->versions[0].mode);\n+}\n+\n+static int tecmp1 (const void *_a, const void *_b)\n+{\n+\tstruct fi_tree_entry *a = *((struct fi_tree_entry**)_a);\n+\tstruct fi_tree_entry *b = *((struct fi_tree_entry**)_b);\n+\treturn base_name_compare(\n+\t\ta->name->data, a->name->len, a->versions[1].mode,\n+\t\tb->name->data, b->name->len, b->versions[1].mode);\n+}\n+\n+static void mktree(struct fi_tree *t, int v, struct strbuf *b)\n+{\n+\tsize_t maxlen = 0;\n+\tunsigned int i;\n+\n+\tif (!v)\n+\t\tqsort(t->entries,t->entry_count,sizeof(t->entries[0]),tecmp0);\n+\telse\n+\t\tqsort(t->entries,t->entry_count,sizeof(t->entries[0]),tecmp1);\n+\n+\tfor (i = 0; i < t->entry_count; i++) {\n+\t\tif (t->entries[i]->versions[v].mode)\n+\t\t\tmaxlen += t->entries[i]->name->len + 34;\n+\t}\n+\n+\tstrbuf_reset(b);\n+\tstrbuf_grow(b, maxlen);\n+\tfor (i = 0; i < t->entry_count; i++) {\n+\t\tstruct fi_tree_entry *e = t->entries[i];\n+\t\tif (!e->versions[v].mode)\n+\t\t\tcontinue;\n+\t\tstrbuf_addf(b, \"%o %s%c\", (unsigned int)e->versions[v].mode,\n+\t\t\t\t\te->name->data, '\\0');\n+\t\tstrbuf_add(b, e->versions[v].sha1, 20);\n+\t}\n+}\n+\n+static void store_tree(struct fi_tree_entry *root)\n+{\n+\tstruct fi_tree *t = root->tree;\n+\tunsigned int i, j, del;\n+\tstruct last_object lo = { STRBUF_INIT, 0, 0, /* no_swap */ 1 };\n+\tstruct fi_object *le;\n+\tstruct strbuf new_tree = STRBUF_INIT;\n+\n+\tif (!is_null_sha1(root->versions[1].sha1))\n+\t\treturn;\n+\n+\tfor (i = 0; i < t->entry_count; i++) {\n+\t\tif (t->entries[i]->tree)\n+\t\t\tstore_tree(t->entries[i]);\n+\t}\n+\n+\tle = find_object(root->versions[0].sha1);\n+\tif (S_ISDIR(root->versions[0].mode) && le) {\n+\t\tmktree(t, 0, &lo.data);\n+\t\tlo.offset = le->idx.offset;\n+\t\tlo.depth = t->delta_depth;\n+\t}\n+\n+\tmktree(t, 1, &new_tree);\n+\tstore_object(OBJ_TREE, &new_tree, &lo, root->versions[1].sha1, NULL);\n+\n+\tt->delta_depth = lo.depth;\n+\tfor (i = 0, j = 0, del = 0; i < t->entry_count; i++) {\n+\t\tstruct fi_tree_entry *e = t->entries[i];\n+\t\tif (e->versions[1].mode) {\n+\t\t\te->versions[0].mode = e->versions[1].mode;\n+\t\t\thashcpy(e->versions[0].sha1, e->versions[1].sha1);\n+\t\t\tt->entries[j++] = e;\n+\t\t} else {\n+\t\t\trelease_tree_entry(e);\n+\t\t\tdel++;\n+\t\t}\n+\t}\n+\tt->entry_count -= del;\n+}\n+\n+static int tree_content_set(\n+\tstruct fi_tree_entry *root,\n+\tconst char *p,\n+\tconst unsigned char *sha1,\n+\tconst uint16_t mode,\n+\tstruct fi_tree *subtree)\n+{\n+\tstruct fi_tree *t = root->tree;\n+\tconst char *slash1;\n+\tunsigned int i, n;\n+\tstruct fi_tree_entry *e;\n+\n+\tslash1 = strchr(p, '/');\n+\tif (slash1)\n+\t\tn = slash1 - p;\n+\telse\n+\t\tn = strlen(p);\n+\tif (!n)\n+\t\tdie(\"Empty path component found in input\");\n+\tif (!slash1 && !S_ISDIR(mode) && subtree)\n+\t\tdie(\"Non-directories cannot have subtrees\");\n+\n+\tfor (i = 0; i < t->entry_count; i++) {\n+\t\te = t->entries[i];\n+\t\tif (e->name->len == n && !strncmp(p, e->name->data, n)) {\n+\t\t\tif (!slash1) {\n+\t\t\t\tif (!S_ISDIR(mode)\n+\t\t\t\t\t\t&& e->versions[1].mode == mode\n+\t\t\t\t\t\t&& !hashcmp(e->versions[1].sha1, sha1))\n+\t\t\t\t\treturn 0;\n+\t\t\t\te->versions[1].mode = mode;\n+\t\t\t\thashcpy(e->versions[1].sha1, sha1);\n+\t\t\t\tif (e->tree)\n+\t\t\t\t\trelease_tree_content_recursive(e->tree);\n+\t\t\t\te->tree = subtree;\n+\t\t\t\thashclr(root->versions[1].sha1);\n+\t\t\t\treturn 1;\n+\t\t\t}\n+\t\t\tif (!S_ISDIR(e->versions[1].mode)) {\n+\t\t\t\te->tree = new_tree_content(8);\n+\t\t\t\te->versions[1].mode = S_IFDIR;\n+\t\t\t}\n+\t\t\tif (!e->tree)\n+\t\t\t\tload_tree(e);\n+\t\t\tif (tree_content_set(e, slash1 + 1, sha1, mode, subtree)) {\n+\t\t\t\thashclr(root->versions[1].sha1);\n+\t\t\t\treturn 1;\n+\t\t\t}\n+\t\t\treturn 0;\n+\t\t}\n+\t}\n+\n+\tif (t->entry_count == t->entry_capacity)\n+\t\troot->tree = t = grow_tree_content(t, t->entry_count);\n+\te = new_tree_entry();\n+\te->name = to_atom(p, n);\n+\te->versions[0].mode = 0;\n+\thashclr(e->versions[0].sha1);\n+\tt->entries[t->entry_count++] = e;\n+\tif (slash1) {\n+\t\te->tree = new_tree_content(8);\n+\t\te->versions[1].mode = S_IFDIR;\n+\t\ttree_content_set(e, slash1 + 1, sha1, mode, subtree);\n+\t} else {\n+\t\te->tree = subtree;\n+\t\te->versions[1].mode = mode;\n+\t\thashcpy(e->versions[1].sha1, sha1);\n+\t}\n+\thashclr(root->versions[1].sha1);\n+\treturn 1;\n+}\n+\n+static int tree_content_remove(\n+\tstruct fi_tree_entry *root,\n+\tconst char *p,\n+\tstruct fi_tree_entry *backup_leaf)\n+{\n+\tstruct fi_tree *t = root->tree;\n+\tconst char *slash1;\n+\tunsigned int i, n;\n+\tstruct fi_tree_entry *e;\n+\n+\tslash1 = strchr(p, '/');\n+\tif (slash1)\n+\t\tn = slash1 - p;\n+\telse\n+\t\tn = strlen(p);\n+\n+\tfor (i = 0; i < t->entry_count; i++) {\n+\t\te = t->entries[i];\n+\t\tif (e->name->len == n && !strncmp(p, e->name->data, n)) {\n+\t\t\tif (slash1 && !S_ISDIR(e->versions[1].mode))\n+\t\t\t\t/*\n+\t\t\t\t * If p names a file in some subdirectory, and a\n+\t\t\t\t * file or symlink matching the name of the\n+\t\t\t\t * parent directory of p exists, then p cannot\n+\t\t\t\t * exist and need not be deleted.\n+\t\t\t\t */\n+\t\t\t\treturn 1;\n+\t\t\tif (!slash1 || !S_ISDIR(e->versions[1].mode))\n+\t\t\t\tgoto del_entry;\n+\t\t\tif (!e->tree)\n+\t\t\t\tload_tree(e);\n+\t\t\tif (tree_content_remove(e, slash1 + 1, backup_leaf)) {\n+\t\t\t\tfor (n = 0; n < e->tree->entry_count; n++) {\n+\t\t\t\t\tif (e->tree->entries[n]->versions[1].mode) {\n+\t\t\t\t\t\thashclr(root->versions[1].sha1);\n+\t\t\t\t\t\treturn 1;\n+\t\t\t\t\t}\n+\t\t\t\t}\n+\t\t\t\tbackup_leaf = NULL;\n+\t\t\t\tgoto del_entry;\n+\t\t\t}\n+\t\t\treturn 0;\n+\t\t}\n+\t}\n+\treturn 0;\n+\n+del_entry:\n+\tif (backup_leaf)\n+\t\tmemcpy(backup_leaf, e, sizeof(*backup_leaf));\n+\telse if (e->tree)\n+\t\trelease_tree_content_recursive(e->tree);\n+\te->tree = NULL;\n+\te->versions[1].mode = 0;\n+\thashclr(e->versions[1].sha1);\n+\thashclr(root->versions[1].sha1);\n+\treturn 1;\n+}\n+\n+static int tree_content_get(\n+\tstruct fi_tree_entry *root,\n+\tconst char *p,\n+\tstruct fi_tree_entry *leaf)\n+{\n+\tstruct fi_tree *t = root->tree;\n+\tconst char *slash1;\n+\tunsigned int i, n;\n+\tstruct fi_tree_entry *e;\n+\n+\tslash1 = strchr(p, '/');\n+\tif (slash1)\n+\t\tn = slash1 - p;\n+\telse\n+\t\tn = strlen(p);\n+\n+\tfor (i = 0; i < t->entry_count; i++) {\n+\t\te = t->entries[i];\n+\t\tif (e->name->len == n && !strncmp(p, e->name->data, n)) {\n+\t\t\tif (!slash1) {\n+\t\t\t\tmemcpy(leaf, e, sizeof(*leaf));\n+\t\t\t\tif (e->tree && is_null_sha1(e->versions[1].sha1))\n+\t\t\t\t\tleaf->tree = dup_tree_content(e->tree);\n+\t\t\t\telse\n+\t\t\t\t\tleaf->tree = NULL;\n+\t\t\t\treturn 1;\n+\t\t\t}\n+\t\t\tif (!S_ISDIR(e->versions[1].mode))\n+\t\t\t\treturn 0;\n+\t\t\tif (!e->tree)\n+\t\t\t\tload_tree(e);\n+\t\t\treturn tree_content_get(e, slash1 + 1, leaf);\n+\t\t}\n+\t}\n+\treturn 0;\n+}\n+\n+/* Parse the optional mark from the stream and return the associated atom. */\n+static struct fi_atom *parse_mark(void)\n+{\n+\tstruct fi_atom *atom = NULL;\n+\n+\tif (!prefixcmp(command_buf.buf, \"mark :\")) {\n+\t\tatom = to_atom(command_buf.buf + 6, strlen(command_buf.buf + 6));\n+\t\tread_next_command();\n+\t}\n+\n+\treturn atom;\n+}\n+\n+static int parse_data(struct strbuf *sb, uintmax_t limit, uintmax_t *len_res)\n+{\n+\tstrbuf_reset(sb);\n+\n+\tif (prefixcmp(command_buf.buf, \"data \"))\n+\t\tdie(\"Expected 'data n' command, found: %s\", command_buf.buf);\n+\n+\tif (!prefixcmp(command_buf.buf + 5, \"<<\")) {\n+\t\tchar *term = xstrdup(command_buf.buf + 5 + 2);\n+\t\tsize_t term_len = command_buf.len - 5 - 2;\n+\n+\t\tstrbuf_detach(&command_buf, NULL);\n+\t\tfor (;;) {\n+\t\t\tif (strbuf_getline(&command_buf, stdin, '\\n') == EOF)\n+\t\t\t\tdie(\"EOF in data (terminator '%s' not found)\", term);\n+\t\t\tif (term_len == command_buf.len\n+\t\t\t\t&& !strcmp(term, command_buf.buf))\n+\t\t\t\tbreak;\n+\t\t\tstrbuf_addbuf(sb, &command_buf);\n+\t\t\tstrbuf_addch(sb, '\\n');\n+\t\t}\n+\t\tfree(term);\n+\t}\n+\telse {\n+\t\tuintmax_t len = strtoumax(command_buf.buf + 5, NULL, 10);\n+\t\tsize_t n = 0, length = (size_t)len;\n+\n+\t\tif (limit && limit < len) {\n+\t\t\t*len_res = len;\n+\t\t\treturn 0;\n+\t\t}\n+\t\tif (length < len)\n+\t\t\tdie(\"data is too large to use in this context\");\n+\n+\t\twhile (n < length) {\n+\t\t\tsize_t s = strbuf_fread(sb, length - n, stdin);\n+\t\t\tif (!s && feof(stdin))\n+\t\t\t\tdie(\"EOF in data (%lu bytes remaining)\",\n+\t\t\t\t\t(unsigned long)(length - n));\n+\t\t\tn += s;\n+\t\t}\n+\t}\n+\n+\tread_next_command();\n+\tskip_optional_lf();\n+\treturn 1;\n+}\n+\n+static int validate_raw_date(const char *src, char *result, int maxlen)\n+{\n+\tconst char *orig_src = src;\n+\tchar *endp;\n+\tunsigned long num;\n+\n+\terrno = 0;\n+\n+\tnum = strtoul(src, &endp, 10);\n+\t/* NEEDSWORK: perhaps check for reasonable values? */\n+\tif (errno || endp == src || *endp != ' ')\n+\t\treturn -1;\n+\n+\tsrc = endp + 1;\n+\tif (*src != '-' && *src != '+')\n+\t\treturn -1;\n+\n+\tnum = strtoul(src + 1, &endp, 10);\n+\tif (errno || endp == src + 1 || *endp || (endp - orig_src) >= maxlen ||\n+\t    1400 < num)\n+\t\treturn -1;\n+\n+\tstrcpy(result, orig_src);\n+\treturn 0;\n+}\n+\n+static char *parse_ident(const char *buf)\n+{\n+\tconst char *gt;\n+\tsize_t name_len;\n+\tchar *ident;\n+\n+\tgt = strrchr(buf, '>');\n+\tif (!gt)\n+\t\tdie(\"Missing > in ident string: %s\", buf);\n+\tgt++;\n+\tif (*gt != ' ')\n+\t\tdie(\"Missing space after > in ident string: %s\", buf);\n+\tgt++;\n+\tname_len = gt - buf;\n+\tident = xmalloc(name_len + 24);\n+\tstrncpy(ident, buf, name_len);\n+\n+\tswitch (whenspec) {\n+\tcase WHENSPEC_RAW:\n+\t\tif (validate_raw_date(gt, ident + name_len, 24) < 0)\n+\t\t\tdie(\"Invalid raw date \\\"%s\\\" in ident: %s\", gt, buf);\n+\t\tbreak;\n+\tcase WHENSPEC_RFC2822:\n+\t\tif (parse_date(gt, ident + name_len, 24) < 0)\n+\t\t\tdie(\"Invalid rfc2822 date \\\"%s\\\" in ident: %s\", gt, buf);\n+\t\tbreak;\n+\tcase WHENSPEC_NOW:\n+\t\tif (strcmp(\"now\", gt))\n+\t\t\tdie(\"Date in ident must be 'now': %s\", buf);\n+\t\tdatestamp(ident + name_len, 24);\n+\t\tbreak;\n+\t}\n+\n+\treturn ident;\n+}\n+\n+static void parse_and_store_blob(\n+\tstruct last_object *last,\n+\tunsigned char *sha1out,\n+\tstruct fi_atom *atom)\n+{\n+\tstatic struct strbuf buf = STRBUF_INIT;\n+\tuintmax_t len;\n+\n+\tif (parse_data(&buf, big_file_threshold, &len))\n+\t\tstore_object(OBJ_BLOB, &buf, last, sha1out, atom);\n+\telse {\n+\t\tif (last) {\n+\t\t\tstrbuf_release(&last->data);\n+\t\t\tlast->offset = 0;\n+\t\t\tlast->depth = 0;\n+\t\t}\n+\t\tstream_blob(len, sha1out, atom);\n+\t\tskip_optional_lf();\n+\t}\n+}\n+\n+static void unload_one_branch(void)\n+{\n+\twhile (cur_active_branches\n+\t\t&& cur_active_branches >= max_active_branches) {\n+\t\tstruct fi_branch *e, *p = active_branches;\n+\n+\t\tif (p) {\n+\t\t\te = p->active_next_branch;\n+\t\t\tp->active_next_branch = e->active_next_branch;\n+\t\t} else {\n+\t\t\te = active_branches;\n+\t\t\tactive_branches = e->active_next_branch;\n+\t\t}\n+\t\te->active = 0;\n+\t\te->active_next_branch = NULL;\n+\t\tif (e->branch_tree.tree) {\n+\t\t\trelease_tree_content_recursive(e->branch_tree.tree);\n+\t\t\te->branch_tree.tree = NULL;\n+\t\t}\n+\t\tcur_active_branches--;\n+\t}\n+}\n+\n+static void load_branch(struct fi_branch *b)\n+{\n+\tload_tree(&b->branch_tree);\n+\tif (!b->active) {\n+\t\tb->active = 1;\n+\t\tb->active_next_branch = active_branches;\n+\t\tactive_branches = b;\n+\t\tcur_active_branches++;\n+\t}\n+}\n+\n+static unsigned char convert_num_notes_to_fanout(uintmax_t num_notes)\n+{\n+\tunsigned char fanout = 0;\n+\twhile ((num_notes >>= 8))\n+\t\tfanout++;\n+\treturn fanout;\n+}\n+\n+static void construct_path_with_fanout(const char *hex_sha1,\n+\t\tunsigned char fanout, char *path)\n+{\n+\tunsigned int i = 0, j = 0;\n+\tif (fanout >= 20)\n+\t\tdie(\"Too large fanout (%u)\", fanout);\n+\twhile (fanout) {\n+\t\tpath[i++] = hex_sha1[j++];\n+\t\tpath[i++] = hex_sha1[j++];\n+\t\tpath[i++] = '/';\n+\t\tfanout--;\n+\t}\n+\tmemcpy(path + i, hex_sha1 + j, 40 - j);\n+\tpath[i + 40 - j] = '\\0';\n+}\n+\n+static uintmax_t do_change_note_fanout(\n+\t\tstruct fi_tree_entry *orig_root, struct fi_tree_entry *root,\n+\t\tchar *hex_sha1, unsigned int hex_sha1_len,\n+\t\tchar *fullpath, unsigned int fullpath_len,\n+\t\tunsigned char fanout)\n+{\n+\tstruct fi_tree *t = root->tree;\n+\tstruct fi_tree_entry *e, leaf;\n+\tunsigned int i, tmp_hex_sha1_len, tmp_fullpath_len;\n+\tuintmax_t num_notes = 0;\n+\tunsigned char sha1[20];\n+\tchar realpath[60];\n+\n+\tfor (i = 0; t && i < t->entry_count; i++) {\n+\t\te = t->entries[i];\n+\t\ttmp_hex_sha1_len = hex_sha1_len + e->name->len;\n+\t\ttmp_fullpath_len = fullpath_len;\n+\n+\t\t/*\n+\t\t * We're interested in EITHER existing note entries (entries\n+\t\t * with exactly 40 hex chars in path, not including directory\n+\t\t * separators), OR directory entries that may contain note\n+\t\t * entries (with < 40 hex chars in path).\n+\t\t * Also, each path component in a note entry must be a multiple\n+\t\t * of 2 chars.\n+\t\t */\n+\t\tif (!e->versions[1].mode ||\n+\t\t    tmp_hex_sha1_len > 40 ||\n+\t\t    e->name->len % 2)\n+\t\t\tcontinue;\n+\n+\t\t/* This _may_ be a note entry, or a subdir containing notes */\n+\t\tmemcpy(hex_sha1 + hex_sha1_len, e->name->data,\n+\t\t       e->name->len);\n+\t\tif (tmp_fullpath_len)\n+\t\t\tfullpath[tmp_fullpath_len++] = '/';\n+\t\tmemcpy(fullpath + tmp_fullpath_len, e->name->data,\n+\t\t       e->name->len);\n+\t\ttmp_fullpath_len += e->name->len;\n+\t\tfullpath[tmp_fullpath_len] = '\\0';\n+\n+\t\tif (tmp_hex_sha1_len == 40 && !get_sha1_hex(hex_sha1, sha1)) {\n+\t\t\t/* This is a note entry */\n+\t\t\tconstruct_path_with_fanout(hex_sha1, fanout, realpath);\n+\t\t\tif (!strcmp(fullpath, realpath)) {\n+\t\t\t\t/* Note entry is in correct location */\n+\t\t\t\tnum_notes++;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\n+\t\t\t/* Rename fullpath to realpath */\n+\t\t\tif (!tree_content_remove(orig_root, fullpath, &leaf))\n+\t\t\t\tdie(\"Failed to remove path %s\", fullpath);\n+\t\t\ttree_content_set(orig_root, realpath,\n+\t\t\t\tleaf.versions[1].sha1,\n+\t\t\t\tleaf.versions[1].mode,\n+\t\t\t\tleaf.tree);\n+\t\t} else if (S_ISDIR(e->versions[1].mode)) {\n+\t\t\t/* This is a subdir that may contain note entries */\n+\t\t\tif (!e->tree)\n+\t\t\t\tload_tree(e);\n+\t\t\tnum_notes += do_change_note_fanout(orig_root, e,\n+\t\t\t\thex_sha1, tmp_hex_sha1_len,\n+\t\t\t\tfullpath, tmp_fullpath_len, fanout);\n+\t\t}\n+\n+\t\t/* The above may have reallocated the current tree_content */\n+\t\tt = root->tree;\n+\t}\n+\treturn num_notes;\n+}\n+\n+static uintmax_t change_note_fanout(struct fi_tree_entry *root,\n+\t\tunsigned char fanout)\n+{\n+\tchar hex_sha1[40], path[60];\n+\treturn do_change_note_fanout(root, root, hex_sha1, 0, path, 0, fanout);\n+}\n+\n+static void file_change_m(struct fi_branch *b)\n+{\n+\tconst char *p = command_buf.buf + 2;\n+\tstatic struct strbuf uq = STRBUF_INIT;\n+\tconst char *endp;\n+\tstruct fi_object *oe = NULL;\n+\tunsigned char sha1[20];\n+\tunsigned int mode, inline_data = 0;\n+\n+\tp = decode_tree_mode(p, &mode);\n+\tif (!p)\n+\t\tdie(\"Corrupt mode: %s\", command_buf.buf);\n+\tswitch (mode) {\n+\tcase 0644:\n+\tcase 0755:\n+\t\tmode |= S_IFREG;\n+\tcase S_IFREG | 0644:\n+\tcase S_IFREG | 0755:\n+\tcase S_IFLNK:\n+\tcase S_IFDIR:\n+\tcase S_IFGITLINK:\n+\t\t/* ok */\n+\t\tbreak;\n+\tdefault:\n+\t\tdie(\"Corrupt mode: %s\", command_buf.buf);\n+\t}\n+\n+\tif (*p == ':') {\n+\t\tp = find_mark(p + 1, sha1);\n+\t} else if (!prefixcmp(p, \"inline\")) {\n+\t\tinline_data = 1;\n+\t\tp += 6;\n+\t} else {\n+\t\tif (get_sha1_hex(p, sha1))\n+\t\t\tdie(\"Invalid SHA1: %s\", command_buf.buf);\n+\t\tp += 40;\n+\t}\n+\tif (*p++ != ' ')\n+\t\tdie(\"Missing space after SHA1: %s\", command_buf.buf);\n+\n+\tstrbuf_reset(&uq);\n+\tif (!unquote_c_style(&uq, p, &endp)) {\n+\t\tif (*endp)\n+\t\t\tdie(\"Garbage after path in: %s\", command_buf.buf);\n+\t\tp = uq.buf;\n+\t}\n+\n+\tif (S_ISGITLINK(mode)) {\n+\t\tif (inline_data)\n+\t\t\tdie(\"Git links cannot be specified 'inline': %s\",\n+\t\t\t\tcommand_buf.buf);\n+\t\telse if (oe) {\n+\t\t\tif (oe->type != OBJ_COMMIT)\n+\t\t\t\tdie(\"Not a commit (actually a %s): %s\",\n+\t\t\t\t\ttypename(oe->type), command_buf.buf);\n+\t\t}\n+\t\t/*\n+\t\t * Accept the sha1 without checking; it expected to be in\n+\t\t * another repository.\n+\t\t */\n+\t} else if (inline_data) {\n+\t\tif (S_ISDIR(mode))\n+\t\t\tdie(\"Directories cannot be specified 'inline': %s\",\n+\t\t\t\tcommand_buf.buf);\n+\t\tif (p != uq.buf) {\n+\t\t\tstrbuf_addstr(&uq, p);\n+\t\t\tp = uq.buf;\n+\t\t}\n+\t\tread_next_command();\n+\t\tparse_and_store_blob(&last_blob, sha1, 0);\n+\t} else {\n+\t\tenum object_type expected = S_ISDIR(mode) ?\n+\t\t\t\t\t\tOBJ_TREE: OBJ_BLOB;\n+\t\tenum object_type type = fi_sha1_object_type(sha1);\n+\t\tif (type < 0)\n+\t\t\tdie(\"%s not found: %s\",\n+\t\t\t\t\tS_ISDIR(mode) ?  \"Tree\" : \"Blob\",\n+\t\t\t\t\tcommand_buf.buf);\n+\t\tif (type != expected)\n+\t\t\tdie(\"Not a %s (actually a %s): %s %s\",\n+\t\t\t\ttypename(expected), typename(type),\n+\t\t\t\tcommand_buf.buf, sha1_to_hex(sha1));\n+\t}\n+\n+\ttree_content_set(&b->branch_tree, p, sha1, mode, NULL);\n+}\n+\n+static void file_change_d(struct fi_branch *b)\n+{\n+\tconst char *p = command_buf.buf + 2;\n+\tstatic struct strbuf uq = STRBUF_INIT;\n+\tconst char *endp;\n+\n+\tstrbuf_reset(&uq);\n+\tif (!unquote_c_style(&uq, p, &endp)) {\n+\t\tif (*endp)\n+\t\t\tdie(\"Garbage after path in: %s\", command_buf.buf);\n+\t\tp = uq.buf;\n+\t}\n+\ttree_content_remove(&b->branch_tree, p, NULL);\n+}\n+\n+static void file_change_cr(struct fi_branch *b, int rename)\n+{\n+\tconst char *s, *d;\n+\tstatic struct strbuf s_uq = STRBUF_INIT;\n+\tstatic struct strbuf d_uq = STRBUF_INIT;\n+\tconst char *endp;\n+\tstruct fi_tree_entry leaf;\n+\n+\ts = command_buf.buf + 2;\n+\tstrbuf_reset(&s_uq);\n+\tif (!unquote_c_style(&s_uq, s, &endp)) {\n+\t\tif (*endp != ' ')\n+\t\t\tdie(\"Missing space after source: %s\", command_buf.buf);\n+\t} else {\n+\t\tendp = strchr(s, ' ');\n+\t\tif (!endp)\n+\t\t\tdie(\"Missing space after source: %s\", command_buf.buf);\n+\t\tstrbuf_add(&s_uq, s, endp - s);\n+\t}\n+\ts = s_uq.buf;\n+\n+\tendp++;\n+\tif (!*endp)\n+\t\tdie(\"Missing dest: %s\", command_buf.buf);\n+\n+\td = endp;\n+\tstrbuf_reset(&d_uq);\n+\tif (!unquote_c_style(&d_uq, d, &endp)) {\n+\t\tif (*endp)\n+\t\t\tdie(\"Garbage after dest in: %s\", command_buf.buf);\n+\t\td = d_uq.buf;\n+\t}\n+\n+\tmemset(&leaf, 0, sizeof(leaf));\n+\tif (rename)\n+\t\ttree_content_remove(&b->branch_tree, s, &leaf);\n+\telse\n+\t\ttree_content_get(&b->branch_tree, s, &leaf);\n+\tif (!leaf.versions[1].mode)\n+\t\tdie(\"Path %s not in branch\", s);\n+\ttree_content_set(&b->branch_tree, d,\n+\t\tleaf.versions[1].sha1,\n+\t\tleaf.versions[1].mode,\n+\t\tleaf.tree);\n+}\n+\n+static void note_change_n(struct fi_branch *b, unsigned char old_fanout)\n+{\n+\tconst char *p = command_buf.buf + 2;\n+\tstatic struct strbuf uq = STRBUF_INIT;\n+\tstruct fi_object *oe = oe;\n+\tstruct fi_branch *s;\n+\tunsigned char sha1[20], commit_sha1[20];\n+\tchar path[60];\n+\tuint16_t inline_data = 0;\n+\tunsigned char new_fanout;\n+\n+\t/* <dataref> or 'inline' */\n+\tif (*p == ':') {\n+\t\tp = find_mark(p + 1, sha1);\n+\t} else if (!prefixcmp(p, \"inline\")) {\n+\t\tinline_data = 1;\n+\t\tp += 6;\n+\t} else {\n+\t\tif (get_sha1_hex(p, sha1))\n+\t\t\tdie(\"Invalid SHA1: %s\", command_buf.buf);\n+\t\toe = find_object(sha1);\n+\t\tp += 40;\n+\t}\n+\tif (*p++ != ' ')\n+\t\tdie(\"Missing space after SHA1: %s\", command_buf.buf);\n+\n+\t/* <committish> */\n+\ts = lookup_branch(p);\n+\tif (s) {\n+\t\thashcpy(commit_sha1, s->sha1);\n+\t} else if (*p == ':') {\n+\t\tfind_mark(p + 1, commit_sha1);\n+\t} else if (!get_sha1(p, commit_sha1)) {\n+\t\tenum object_type type;\n+\t\ttype = fi_sha1_object_type(commit_sha1);\n+\t\tif (type != OBJ_COMMIT)\n+\t\t\tdie(\"Can only add notes to commits\");\n+\t} else\n+\t\tdie(\"Invalid ref name or SHA1 expression: %s\", p);\n+\n+\tif (inline_data) {\n+\t\tif (p != uq.buf) {\n+\t\t\tstrbuf_addstr(&uq, p);\n+\t\t\tp = uq.buf;\n+\t\t}\n+\t\tread_next_command();\n+\t\tparse_and_store_blob(&last_blob, sha1, 0);\n+\t} else if (oe) {\n+\t\tif (oe->type != OBJ_BLOB)\n+\t\t\tdie(\"Not a blob (actually a %s): %s\",\n+\t\t\t\ttypename(oe->type), command_buf.buf);\n+\t} else if (!is_null_sha1(sha1)) {\n+\t\tenum object_type type = fi_sha1_object_type(sha1);\n+\t\tif (type < 0)\n+\t\t\tdie(\"Blob not found: %s\", command_buf.buf);\n+\t\tif (type != OBJ_BLOB)\n+\t\t\tdie(\"Not a blob (actually a %s): %s\",\n+\t\t\t    typename(type), command_buf.buf);\n+\t}\n+\n+\tconstruct_path_with_fanout(sha1_to_hex(commit_sha1), old_fanout, path);\n+\tif (tree_content_remove(&b->branch_tree, path, NULL))\n+\t\tb->num_notes--;\n+\n+\tif (is_null_sha1(sha1))\n+\t\treturn; /* nothing to insert */\n+\n+\tb->num_notes++;\n+\tnew_fanout = convert_num_notes_to_fanout(b->num_notes);\n+\tconstruct_path_with_fanout(sha1_to_hex(commit_sha1), new_fanout, path);\n+\ttree_content_set(&b->branch_tree, path, sha1, S_IFREG | 0644, NULL);\n+}\n+\n+static void file_change_deleteall(struct fi_branch *b)\n+{\n+\trelease_tree_content_recursive(b->branch_tree.tree);\n+\thashclr(b->branch_tree.versions[0].sha1);\n+\thashclr(b->branch_tree.versions[1].sha1);\n+\tload_tree(&b->branch_tree);\n+\tb->num_notes = 0;\n+}\n+\n+static void parse_from_commit(struct fi_branch *b, char *buf, unsigned long size)\n+{\n+\tif (!buf || size < 46)\n+\t\tdie(\"Not a valid commit: %s\", sha1_to_hex(b->sha1));\n+\tif (memcmp(\"tree \", buf, 5)\n+\t\t|| get_sha1_hex(buf + 5, b->branch_tree.versions[1].sha1))\n+\t\tdie(\"The commit %s is corrupt\", sha1_to_hex(b->sha1));\n+\thashcpy(b->branch_tree.versions[0].sha1,\n+\t\tb->branch_tree.versions[1].sha1);\n+}\n+\n+static void parse_from_existing(struct fi_branch *b)\n+{\n+\tif (is_null_sha1(b->sha1)) {\n+\t\thashclr(b->branch_tree.versions[0].sha1);\n+\t\thashclr(b->branch_tree.versions[1].sha1);\n+\t} else {\n+\t\tunsigned long size;\n+\t\tenum object_type type;\n+\t\tchar *buf;\n+\n+\t\tbuf = fi_read_sha1_file(b->sha1, &type, &size);\n+\t\tparse_from_commit(b, buf, size);\n+\t\tfree(buf);\n+\t}\n+}\n+\n+static int parse_from(struct fi_branch *b)\n+{\n+\tconst char *from;\n+\tstruct fi_branch *s;\n+\n+\tif (prefixcmp(command_buf.buf, \"from \"))\n+\t\treturn 0;\n+\n+\tif (b->branch_tree.tree) {\n+\t\trelease_tree_content_recursive(b->branch_tree.tree);\n+\t\tb->branch_tree.tree = NULL;\n+\t}\n+\n+\tfrom = strchr(command_buf.buf, ' ') + 1;\n+\ts = lookup_branch(from);\n+\tif (b == s)\n+\t\tdie(\"Can't create a branch from itself: %s\", b->name);\n+\telse if (s) {\n+\t\tunsigned char *t = s->branch_tree.versions[1].sha1;\n+\t\thashcpy(b->sha1, s->sha1);\n+\t\thashcpy(b->branch_tree.versions[0].sha1, t);\n+\t\thashcpy(b->branch_tree.versions[1].sha1, t);\n+\t} else if (*from == ':') {\n+\t\tfind_mark(from + 1, b->sha1);\n+\t\tif (!is_null_sha1(b->sha1)) {\n+\t\t\tunsigned long size;\n+\t\t\tenum object_type type;\n+\t\t\tchar *buf = fi_read_sha1_file(b->sha1, &type, &size);\n+\t\t\tparse_from_commit(b, buf, size);\n+\t\t\tfree(buf);\n+\t\t} else\n+\t\t\tparse_from_existing(b);\n+\t} else if (!get_sha1(from, b->sha1))\n+\t\tparse_from_existing(b);\n+\telse\n+\t\tdie(\"Invalid ref name or SHA1 expression: %s\", from);\n+\n+\tread_next_command();\n+\treturn 1;\n+}\n+\n+static struct hash_list *parse_merge(unsigned int *count)\n+{\n+\tstruct hash_list *list = NULL, *n, *e = e;\n+\tconst char *from;\n+\tstruct fi_branch *s;\n+\n+\t*count = 0;\n+\twhile (!prefixcmp(command_buf.buf, \"merge \")) {\n+\t\tfrom = strchr(command_buf.buf, ' ') + 1;\n+\t\tn = xmalloc(sizeof(*n));\n+\t\ts = lookup_branch(from);\n+\t\tif (s)\n+\t\t\thashcpy(n->sha1, s->sha1);\n+\t\telse if (*from == ':') {\n+\t\t\tfind_mark(from + 1, n->sha1);\n+\t\t} else if (!get_sha1(from, n->sha1)) {\n+\t\t\tunsigned long size;\n+\t\t\tchar *buf = read_object_with_reference(n->sha1,\n+\t\t\t\tcommit_type, &size, n->sha1);\n+\t\t\tif (!buf || size < 46)\n+\t\t\t\tdie(\"Not a valid commit: %s\", from);\n+\t\t\tfree(buf);\n+\t\t} else\n+\t\t\tdie(\"Invalid ref name or SHA1 expression: %s\", from);\n+\n+\t\tn->next = NULL;\n+\t\tif (list)\n+\t\t\te->next = n;\n+\t\telse\n+\t\t\tlist = n;\n+\t\te = n;\n+\t\t(*count)++;\n+\t\tread_next_command();\n+\t}\n+\treturn list;\n+}\n+\n+static int fi_command_commit(void)\n+{\n+\tstatic struct strbuf msg = STRBUF_INIT;\n+\tstatic struct strbuf new_data = STRBUF_INIT;\n+\tstruct fi_branch *b;\n+\tstruct fi_atom *atom;\n+\tchar *sp;\n+\tchar *author = NULL;\n+\tchar *committer = NULL;\n+\tstruct hash_list *merge_list = NULL;\n+\tunsigned int merge_count;\n+\tunsigned char prev_fanout, new_fanout;\n+\n+\t/* Obtain the branch name from the rest of our command */\n+\tsp = strchr(command_buf.buf, ' ') + 1;\n+\tb = lookup_branch(sp);\n+\tif (!b)\n+\t\tb = new_branch(sp);\n+\n+\tread_next_command();\n+\tatom = parse_mark();\n+\tif (!prefixcmp(command_buf.buf, \"author \")) {\n+\t\tauthor = parse_ident(command_buf.buf + 7);\n+\t\tread_next_command();\n+\t}\n+\tif (!prefixcmp(command_buf.buf, \"committer \")) {\n+\t\tcommitter = parse_ident(command_buf.buf + 10);\n+\t\tread_next_command();\n+\t}\n+\tif (!committer)\n+\t\tdie(\"Expected committer but didn't get one\");\n+\tparse_data(&msg, 0, NULL);\n+\t//read_next_command();\n+\tparse_from(b);\n+\tmerge_list = parse_merge(&merge_count);\n+\n+\t/* ensure the branch is active/loaded */\n+\tif (!b->branch_tree.tree || !max_active_branches) {\n+\t\tunload_one_branch();\n+\t\tload_branch(b);\n+\t}\n+\n+\tprev_fanout = convert_num_notes_to_fanout(b->num_notes);\n+\n+\t/* file_change* */\n+\twhile (command_buf.len > 0) {\n+\t\tif (!prefixcmp(command_buf.buf, \"M \"))\n+\t\t\tfile_change_m(b);\n+\t\telse if (!prefixcmp(command_buf.buf, \"D \"))\n+\t\t\tfile_change_d(b);\n+\t\telse if (!prefixcmp(command_buf.buf, \"R \"))\n+\t\t\tfile_change_cr(b, 1);\n+\t\telse if (!prefixcmp(command_buf.buf, \"C \"))\n+\t\t\tfile_change_cr(b, 0);\n+\t\telse if (!prefixcmp(command_buf.buf, \"N \"))\n+\t\t\tnote_change_n(b, prev_fanout);\n+\t\telse if (!strcmp(\"deleteall\", command_buf.buf))\n+\t\t\tfile_change_deleteall(b);\n+\t\telse\n+\t\t\tbreak;\n+\n+\t\tif (read_next_command() == EOF)\n+\t\t\tbreak;\n+\t}\n+\n+\tnew_fanout = convert_num_notes_to_fanout(b->num_notes);\n+\tif (new_fanout != prev_fanout)\n+\t\tb->num_notes = change_note_fanout(&b->branch_tree, new_fanout);\n+\n+\t/* build the tree and the commit */\n+\tstore_tree(&b->branch_tree);\n+\thashcpy(b->branch_tree.versions[0].sha1,\n+\t\tb->branch_tree.versions[1].sha1);\n+\n+\tstrbuf_reset(&new_data);\n+\tstrbuf_addf(&new_data, \"tree %s\\n\",\n+\t\tsha1_to_hex(b->branch_tree.versions[1].sha1));\n+\tif (!is_null_sha1(b->sha1))\n+\t\tstrbuf_addf(&new_data, \"parent %s\\n\", sha1_to_hex(b->sha1));\n+\twhile (merge_list) {\n+\t\tstruct hash_list *next = merge_list->next;\n+\t\tstrbuf_addf(&new_data, \"parent %s\\n\", sha1_to_hex(merge_list->sha1));\n+\t\tfree(merge_list);\n+\t\tmerge_list = next;\n+\t}\n+\tstrbuf_addf(&new_data,\n+\t\t\"author %s\\n\"\n+\t\t\"committer %s\\n\"\n+\t\t\"\\n\",\n+\t\tauthor ? author : committer, committer);\n+\tstrbuf_addbuf(&new_data, &msg);\n+\tfree(author);\n+\tfree(committer);\n+\n+\tstore_object(OBJ_COMMIT, &new_data, NULL, b->sha1, atom);\n+\n+\tread_next_command();\n+\treturn 0;\n+}\n+\n+static int fi_command_tag(void)\n+{\n+\tstatic struct strbuf msg = STRBUF_INIT;\n+\tstatic struct strbuf new_data = STRBUF_INIT;\n+\tchar *sp;\n+\tconst char *from;\n+\tchar *tagger;\n+\tstruct fi_branch *s;\n+\tstruct fi_tag *t;\n+\tint hash;\n+\tunsigned char sha1[20];\n+\tenum object_type type;\n+\n+\t/* Obtain the new tag name from the rest of our command */\n+\tsp = strchr(command_buf.buf, ' ') + 1;\n+\thash = hc_str(sp, strlen(sp));\n+\n+\tt = malloc(sizeof(struct fi_tag));\n+\tt->next_tag = lookup_hash(hash, &tags);\n+\tt->name = strdup(sp);\n+\tinsert_hash(hash, t, &tags);\n+\t\n+\tread_next_command();\n+\n+\t/* from ... */\n+\tif (prefixcmp(command_buf.buf, \"from \"))\n+\t\tdie(\"Expected from command, got %s\", command_buf.buf);\n+\tfrom = strchr(command_buf.buf, ' ') + 1;\n+\ts = lookup_branch(from);\n+\tif (s) {\n+\t\thashcpy(sha1, s->sha1);\n+\t} else if (*from == ':') {\n+\t\tfind_mark(from + 1, sha1);\n+\t} else if (get_sha1(from, sha1)) {\n+\t\tdie(\"Invalid ref name or SHA1 expression: %s\", from);\n+\t}\n+\n+\ttype = fi_sha1_object_type(sha1);\n+\tread_next_command();\n+\n+\t/* tagger ... */\n+\tif (!prefixcmp(command_buf.buf, \"tagger \")) {\n+\t\ttagger = parse_ident(command_buf.buf + 7);\n+\t\tread_next_command();\n+\t} else\n+\t\ttagger = NULL;\n+\n+\t/* tag payload/message */\n+\tparse_data(&msg, 0, NULL);\n+\n+\t/* build the tag object */\n+\tstrbuf_reset(&new_data);\n+\n+\tstrbuf_addf(&new_data,\n+\t\t    \"object %s\\n\"\n+\t\t    \"type %s\\n\"\n+\t\t    \"tag %s\\n\",\n+\t\t    sha1_to_hex(sha1), typename(type), t->name);\n+\tif (tagger)\n+\t\tstrbuf_addf(&new_data,\n+\t\t\t    \"tagger %s\\n\", tagger);\n+\tstrbuf_addch(&new_data, '\\n');\n+\tstrbuf_addbuf(&new_data, &msg);\n+\tfree(tagger);\n+\n+\tstore_object(OBJ_TAG, &new_data, NULL, t->sha1, NULL);\n+\n+\tread_next_command();\n+\treturn 0;\n+}\n+\n+static int fi_command_reset(void)\n+{\n+\tstruct fi_branch *b;\n+\tchar *sp;\n+\n+\t/* Obtain the branch name from the rest of our command */\n+\tsp = strchr(command_buf.buf, ' ') + 1;\n+\tb = lookup_branch(sp);\n+\tif (b) {\n+\t\thashclr(b->sha1);\n+\t\thashclr(b->branch_tree.versions[0].sha1);\n+\t\thashclr(b->branch_tree.versions[1].sha1);\n+\t\tif (b->branch_tree.tree) {\n+\t\t\trelease_tree_content_recursive(b->branch_tree.tree);\n+\t\t\tb->branch_tree.tree = NULL;\n+\t\t}\n+\t}\n+\telse\n+\t\tb = new_branch(sp);\n+\tread_next_command();\n+\tparse_from(b);\n+\n+\treturn 0;\n+}\n+\n+static int fi_command_blob(void)\n+{\n+\tstruct fi_atom *atom;\n+\n+\tread_next_command();\n+\tatom = parse_mark();\n+\tparse_and_store_blob(&last_blob, NULL, atom);\n+\t\n+\tread_next_command();\n+\treturn 0;\n+}\n+\n+static int fi_command_checkpoint(void)\n+{\n+\tif (objects.nr) {\n+\t\tcycle_packfile();\n+\t}\n+\tskip_optional_lf();\n+\t\n+\tread_next_command();\n+\treturn 0;\n+}\n+\n+static int fi_command_progress(void)\n+{\n+\tskip_optional_lf();\n+\tread_next_command();\n+\treturn 0;\n+}\n+\n+static int fi_command_feature(void)\n+{\n+\tchar *feature = command_buf.buf + 8;\n+\n+\tif (!prefixcmp(feature, \"date-format=\")) {\n+\t\tchar *fmt = feature + 12;\n+\t\tif (!strcmp(fmt, \"raw\")) {\n+\t\t\twhenspec = WHENSPEC_RAW;\n+\t\t} else if (!strcmp(fmt, \"rfc2822\")) {\n+\t\t\twhenspec = WHENSPEC_RFC2822;\n+\t\t} else if (!strcmp(fmt, \"now\")) {\n+\t\t\twhenspec = WHENSPEC_NOW;\n+\t\t} else {\n+\t\t\treturn 1;\n+\t\t}\n+\t} else if (!prefixcmp(feature, \"force\")) {\n+\t\tforce_update = 1;\n+\t} else {\n+\t\treturn 1;\n+\t}\n+\n+\tread_next_command();\n+\treturn 0;\n+}\n+\n+static int fi_command_option(void)\n+{\n+\tread_next_command();\n+\treturn 0;\n+}\n+\n+/* mark SP :<name> SP <value> LF \t*/\n+static int fi_command_mark(void)\n+{\n+\tstruct fi_atom *atom;\n+\tunsigned char sha1[20];\n+\tchar *end = strchr(command_buf.buf + 6, ' ');\n+\n+\tatom = to_atom(command_buf.buf + 6, end - command_buf.buf - 6);\n+\tif (get_sha1(end + 1, sha1))\n+\t\tdie(\"Invalid mark command: %s\", command_buf.buf);\n+\n+\tinsert_mark(atom, sha1);\n+\n+\tread_next_command();\n+\treturn 0;\n+}\n+\n+/* List of commands we understand. */\n+struct fi_command {\n+\tconst char *cmd;\n+\tint (*func)(void);\n+} fi_command[] = {\n+\t{ \"commit\",\t\tfi_command_commit },\n+\t{ \"tag\",\t\tfi_command_tag },\n+\t{ \"reset\",\t\tfi_command_reset },\n+\t{ \"blob\",\t\tfi_command_blob },\n+\t{ \"checkpoint\",\t\tfi_command_checkpoint },\n+\t{ \"progress\",\t\tfi_command_progress },\n+\t{ \"feature\",\t\tfi_command_feature },\n+\t{ \"option\",\t\tfi_command_option },\n+\t{ \"mark\",\t\tfi_command_mark },\n+};\n+\n+static struct fi_command *find_command(const char *cmd)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i < ARRAY_SIZE(fi_command); ++i) {\n+\t\tif (!prefixcmp(command_buf.buf, fi_command[i].cmd)) {\n+\t\t\treturn &fi_command[i];\n+\t\t}\n+\t}\n+\n+\treturn NULL;\n+}\n+\n+int main(int argc, const char **argv)\n+{\n+\tgit_extract_argv0_path(argv[0]);\n+\n+\tsetup_git_directory();\n+\tgit_config(git_pack_config, NULL);\n+\tif (!pack_compression_seen && core_compression_seen)\n+\t\tpack_compression_level = core_compression_level;\n+\n+\tset_die_routine(die_nicely);\n+\n+\t/* Initialize hash tables. */\n+\tinit_hash(&atoms);\n+\tinit_hash(&tags);\n+\tinit_hash(&branches);\n+\tinit_hash(&marks);\n+\tinit_hash(&objects);\n+\t\n+\tprepare_packed_git();\n+\tstart_packfile();\n+\t\n+\tread_next_command();\n+\twhile (command_buf.len > 0) {\n+\t\tstruct fi_command *cmd = find_command(command_buf.buf);\n+\t\tif (!cmd)\n+\t\t\tdie(\"Unsupported command: %s\", command_buf.buf);\n+\n+\t\tint err = (*cmd->func)();\n+\t\tif (err)\n+\t\t\tdie(\"Command failed\");\n+\n+\t\tfflush(stdout);\n+\t}\n+\t\n+\tend_packfile();\n+\n+\treturn 0;\n+}\n-- \n1.7.3.37.gb6088b\n"},{"id":"152350","messageId":"1286108511-55876-6-git-send-email-tom@dbservice.com","threadId":"25319","inReplyTo":"4CA86A12.6080905@dbservice.com","subject":"[PATCH 6/6] Add git-remote-svn","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2010-10-03T12:21:51Z","receivedAt":"2010-10-03T12:21:51Z","isPatch":true,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"This is an experimental git remote helper for svn repositories. It uses\nthe new git fast-import-helper. It only works with local svn repos (not\nover network). It uses notes to save the git commit -> svn revision\nmapping (refs/notes/svn).\n\nThis remote helper serves as a technology preview of what the new type\nof remote helpers can do.\n\nIt assumes that the svn repo uses the standard layout (trunk, branches).\n---\n .gitignore        |    1 +\n Makefile          |    1 +\n git-remote-svn.py |  408 +++++++++++++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 410 insertions(+), 0 deletions(-)\n create mode 100644 git-remote-svn.py\n\ndiff --git a/.gitignore b/.gitignore\nindex c8aa8c7..0a6011c 100644\n--- a/.gitignore\n+++ b/.gitignore\n@@ -114,6 +114,7 @@\n /git-remote-ftp\n /git-remote-ftps\n /git-remote-testgit\n+/git-remote-svn\n /git-repack\n /git-replace\n /git-repo-config\ndiff --git a/Makefile b/Makefile\nindex f8a9c40..eb959c9 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -387,6 +387,7 @@ SCRIPT_PERL += git-send-email.perl\n SCRIPT_PERL += git-svn.perl\n \n SCRIPT_PYTHON += git-remote-testgit.py\n+SCRIPT_PYTHON += git-remote-svn.py\n \n SCRIPTS = $(patsubst %.sh,%,$(SCRIPT_SH)) \\\n \t  $(patsubst %.perl,%,$(SCRIPT_PERL)) \\\ndiff --git a/git-remote-svn.py b/git-remote-svn.py\nnew file mode 100644\nindex 0000000..c617743\n--- /dev/null\n+++ b/git-remote-svn.py\n@@ -0,0 +1,408 @@\n+#!/usr/bin/env python\n+\n+import sys, os, re, time, subprocess\n+import svn.core, svn.repos, svn.fs\n+\n+ct_short = ['M', 'A', 'D', 'R', 'X']\n+\n+############################################################################\n+# Class which encapsulates the fast import helper. Provides methods to\n+# read/write data to it. \n+class FastImportHelper:\n+\n+\tdef start(self):\n+\t\tPIPE = subprocess.PIPE\n+\t\targs = ['git', 'fast-import-helper']\n+\t\tself.helper = subprocess.Popen(args, stdin=PIPE, stdout=PIPE)\n+\n+\tdef close(self):\n+\t\tself.helper.stdin.close()\n+\t\tself.helper.wait()\n+\t\tdel self.helper\n+\n+\tdef write(self, data):\n+\t\tself.helper.stdin.write(data)\n+\n+\t# Expect the helper to write the mark mapping to its stdout. Verify that\n+\t# the mark is the same that we've given and return the git object name\n+\tdef response(self, mark):\n+\t\tline = self.helper.stdout.readline().strip().split(' ')\n+\t\tassert(str(line[1][1:]) == str(mark))\n+\n+\t\treturn line[2]\n+\n+\t# Make a commit from the given arguments, return the git object name\n+\t# corresponding to that just created commit\n+\tdef commit(self, mark, ref, parents, author, committer, changes, message):\n+\t\tself.write(\"commit %s\\n\" % ref)\n+\t\tself.write(\"mark :%s\\n\" % mark)\n+\t\tif author:\n+\t\t\tself.write(\"author %s %s -0000\\n\" % author)\n+\t\tself.write(\"committer %s %s -0000\\n\" % committer)\n+\t\tself.write(\"data %s\\n\" % len(message))\n+\t\tself.write(message)\n+\n+\t\tparent = parents.pop(0)\n+\t\tif parent:\n+\t\t\tself.write(\"from %s\\n\" % parent)\n+\t\tfor parent in parents:\n+\t\t\tself.write(\"merge %s\\n\", parent)\n+\t\n+\t\tself.write(''.join(changes))\n+\n+\t\t# Make it happen\n+\t\tself.write(\"\\n\")\n+\t\treturn self.response(mark)\n+\n+\t# Create a blob from the given arguments. 'read' is callable object\n+\t# which returns data. Return the git object name\n+\tdef blob(self, mark, length, read):\n+\t\tself.write(\"blob\\nmark :%s\\n\" % mark)\n+\t\tself.write(\"data %s\\n\" % length)\n+\t\t\n+\t\twhile length > 0:\n+\t\t\tavail = min(length, 4096)\n+\t\t\tdata = read(avail)\n+\t\t\terr = self.write(data)\n+\t\t\tlength -= avail\n+\n+\t\t# Make it happen\n+\t\tself.write(\"\\n\")\n+\t\treturn self.response(mark)\n+\n+\n+\n+############################################################################\n+# Base class for python remote helpers. It handles the main command loop.\n+# This class also manages the fast-import-helper and marks. If you want to\n+# use the fih, call self.fih.start() first and after you're done call .close()\n+class RemoteHelper(object):\n+\n+\tdef __init__(self, kind):\n+\t\tself.kind = kind\n+\t\tself.fih = FastImportHelper()\n+\t\t\n+\t\tself.notes = []\n+\t\t\n+\t\t# nfrom is the current notes commit, we'll need that later when\n+\t\t# adding new notes. Check if that ref exists, if not set nfrom\n+\t\t# to None, if yes, get the object name and store it in nfrom\n+\t\targv = [ 'git', 'rev-parse', 'refs/notes/%s^0' % kind ]\n+\t\tPIPE = subprocess.PIPE\n+\t\tproc = subprocess.Popen(argv, stdout=PIPE, stderr=PIPE)\n+\t\tproc.wait()\n+\t\tif proc.returncode == 0:\n+\t\t\tself.nfrom = proc.stdout.readline().strip()\n+\t\telse:\n+\t\t\tself.nfrom = None\n+\n+\t# The commands we understand\n+\tCOMMANDS = ( 'capabilities', 'list', 'fetch', )\n+\n+\t# Read next command. Raise an exception if the command is invalid.\n+\t# Return a tuple (command, args,)\n+\tdef read_next_command(self):\n+\t\tline = sys.stdin.readline()\n+\t\tif not line:\n+\t\t\treturn ( None, None, )\n+\t\n+\t\tcmdline = line.strip().split()\n+\t\tif not cmdline:\n+\t\t\treturn ( None, None, )\n+\n+\t\tcmd = cmdline.pop(0)\n+\t\tif cmd not in self.COMMANDS:\n+\t\t\traise Exception(\"Invalid command '%s'\" % cmd)\n+\t\t\n+\t\treturn ( cmd, cmdline, )\n+\n+\t# Run the remote helper, process commands until the end of the world. Or\n+\t# until we're told to finish.\n+\tdef run(self):\n+\t\twhile (True):\n+\t\t\t( cmd, args, ) = self.read_next_command()\n+\t\t\tif cmd is None:\n+\t\t\t\treturn\n+\n+\t\t\tfunc = getattr(self, cmd, None)\n+\t\t\tif func is None or not callable(func):\n+\t\t\t\traise Exception(\"Command '%s' not implemented\" % cmd)\n+\n+\t\t\tresult = func(args)\n+\t\t\tsys.stdout.flush()\n+\n+\n+\t# Convenience method for writing data back to git\n+\tdef reply(self, data):\n+\t\tsys.stdout.write(data)\n+\n+\t# Return all refs and the contents of the note attached to each.\n+\t# This can be used by the remote helper to find out what the latest\n+\t# version is that we fetched into this repo.\n+\t# Returns list of tuples of (sha1, typename, refname, note,)\n+\tdef refs(self):\n+\t\trefs = []\n+\t\t\n+\t\tPIPE = subprocess.PIPE\n+\t\targs = [ 'git', 'for-each-ref' ]\n+\t\tgfer = subprocess.Popen(args, stdin=PIPE, stdout=PIPE)\n+\t\t\n+\t\t# Regular expression for matching the output from g-f-e-r\n+\t\tpattern = re.compile(r\"(.{40}) (\\w+)\t(.*)\")\n+\t\twhile (True):\n+\t\t\tline = gfer.stdout.readline()\n+\t\t\tif not line:\n+\t\t\t\tbreak \n+\n+\t\t\tmatch = pattern.match(line)\n+\t\t\t\n+\t\t\t# The sha1 and name of the ref\n+\t\t\tsha1 = match.group(1)\n+\t\t\ttypename = match.group(2)\n+\t\t\trefname = match.group(3)\n+\t\t\t\n+\t\t\t# Extract the note using `git notes show <sha>`\n+\t\t\tgit_notes_show = [ 'git', 'notes', 'show', sha1 ]\n+\t\t\t\n+\t\t\t# Set GIT_NOTES_REF to point to the notes of our kind\n+\t\t\tenv = { \"GIT_NOTES_REF\": \"refs/notes/%s\" % self.kind }\n+\t\t\t\n+\t\t\tnote = subprocess.Popen(git_notes_show, env=env, stdout=PIPE, stderr=PIPE)\n+\t\t\trefs.append(( sha1, typename, refname, note.stdout.readline() ))\n+\n+\t\tgfer.wait()\n+\n+\t\treturn refs\n+\n+\t# Attach text to an object. objects are currently limited to commits\n+\tdef note(self, obj, text):\n+\t\tself.notes.append(( obj, text, ))\n+\t\tif len(self.notes) >= 10:\n+\t\t\tself.flush()\n+\n+\t# Commit all outstanding notes. Don't forget to flush the notes before\n+\t# you close the fih\n+\tdef flush(self):\n+\t\tif len(self.notes) == 0:\n+\t\t\treturn\n+\n+\t\tnow = int(time.time())\n+\t\tmark = \"%s-notes\" % self.kind\n+\t\tref = \"refs/notes/%s\" % self.kind\n+\t\tparents = [ self.nfrom ]\n+\t\tauthor = ( 'nobody <nobody@localhost>', now, )\n+\t\tcommitter = ( 'nobody <nobody@localhost>', now, )\n+\t\tmessage = \"Update notes\"\n+\n+\t\tchanges = []\n+\t\tfor ( obj, text, ) in self.notes:\n+\t\t\tchanges.append(\"N inline %s\\ndata %s\\n\" % (obj, len(text)))\n+\t\t\tchanges.append(text)\n+\t\t\tchanges.append(\"\\n\")\n+\t\t\n+\t\tself.nfrom = self.fih.commit(mark, ref, parents, author, committer, changes, message)\n+\t\tself.notes = []\n+\n+\n+\n+############################################################################\n+# Remote helper for Subversion\n+class RemoteHelperSubversion(RemoteHelper):\n+\n+\tdef __init__(self, url):\n+\t\tsuper(RemoteHelperSubversion, self).__init__(\"svn\")\n+\n+\t\turl = svn.core.svn_path_canonicalize(url)\n+\t\tself.repo = svn.repos.svn_repos_open(url)\n+\t\tself.fs = svn.repos.svn_repos_fs(self.repo)\n+\t\tself.uuid = svn.fs.svn_fs_get_uuid(self.fs)\n+\n+\n+\t# Here follow the commands this helper implements\n+\t\n+\t# RH command 'capabilities'\n+\tdef capabilities(self, args):\n+\t\tself.reply(\"list\\nfetch\\n\\n\")\n+\n+\t# RH command 'list'\n+\tdef list(self, args):\n+\t\trev = svn.fs.svn_fs_youngest_rev(self.fs)\n+\t\troot = svn.fs.svn_fs_revision_root(self.fs, rev)\n+\n+\t\trefs = self.discover(root)\n+\t\tfor ( name, rev, ) in refs:\n+\t\t\tself.reply(\":r%s %s\\n\" % ( rev, name, ))\n+\n+\t\tif len(refs) > 0:\n+\t\t\tself.reply(\"@%s HEAD\\n\" % refs[0][0])\n+\t\tself.reply(\"\\n\")\n+\n+\t# RH command 'fetch'\n+\tdef fetch(self, args):\n+\t\t# Start the fast-import helper\n+\t\tself.fih.start()\n+\n+\t\t# Fetches are done in batches. Process fetch lines until we see a\n+\t\t# blank newline\n+\t\twhile args:\n+\t\t\t# The revision to fetch, strip the leading 'r' from 'r42'\n+\t\t\tnew = int(args[0][1:])\n+\t\t\t\n+\t\t\t# Trailing slash to ensure that it's a directory\n+\t\t\tprefix = \"/%s/\" % args[1]\n+\t\t\t\n+\t\t\t( sha1, old, ) = self.parent(args[1])\n+\t\t\tsys.stderr.write(\"Best parent: %s %s\\n\" % (old, new,))\n+\t\t\t\n+\t\t\tif old != new:\n+\t\t\t\tsha1 = self.fi(prefix, old, new, sha1)\n+\t\t\tself.reply(\"map r%s %s\\n\" % ( new, sha1 ))\n+\n+\t\t\t# Read next line, break if it's a newline (ending this fetch batch)\n+\t\t\t( cmd, args, ) = self.read_next_command()\n+\t\t\tif not cmd:\n+\t\t\t\tbreak\n+\n+\t\tself.flush()\n+\t\tself.fih.close()\n+\t\t\n+\t\t# Before finishing this command, make sure to emit the 'silent'\n+\t\t# command to register the notes\n+\t\tself.reply(\"silent refs/notes/%s %s\\n\" % (self.kind, self.nfrom, ))\n+\t\t\n+\t\tself.reply(\"\\n\")\n+\n+\t# Discover all refs (trunk, braches) in the repository\n+\tdef discover(self, root):\n+\t\trefs = []\n+\t\t\n+\t\t# First check /trunk\n+\t\tentries = svn.fs.svn_fs_dir_entries(root, \"/\")\n+\t\tnames = entries.keys()\n+\t\t\n+\t\tif 'trunk' in names:\n+\t\t\trefs.append(( 'trunk', self.rev(root, '/trunk'), ))\n+\t\t\n+\t\tif 'branches' in names:\n+\t\t\tentries = svn.fs.svn_fs_dir_entries(root, \"/branches\")\n+\t\t\tnames = entries.keys()\n+\t\t\tfor name in names:\n+\t\t\t\trefs.append(( 'branches/'+name, self.rev(root, '/branches/%s' % name), ))\n+\n+\t\treturn refs\n+\n+\t# Get the revision when `path` was last modified\n+\tdef rev(self, root, path):\n+\t\thistory = svn.fs.svn_fs_node_history(root, path)\n+\t\t\n+\t\t# Yes, this is required.\n+\t\thistory = svn.fs.svn_fs_history_prev(history, True)\n+\t\tif not history:\n+\t\t\treturn 1\n+\n+\t\t( path, rev, ) = svn.fs.svn_fs_history_location(history)\t\t\n+\t\treturn rev\n+\n+\n+\t# Find the git commit we can use as parent when importing from the\n+\t# repo with the given prefix. All commits imported from svn\n+\t# have a note attached which contains this information. But to make our\n+\t# job easier, we only scan ref heads and not the whole history.\n+\t# Go through all refs, see which one has a note that matches the given\n+\t# prefix and extract the svn revision number from the note.\n+\t# Return a tuple (sha1, rev,) which identifies the git commit and svn\n+\t# revision.\n+\tdef parent(self, prefix):\n+\t\tpattern = re.compile(r\"([0-9a-h-]+)/([^@]*)@(\\d+)\")\n+\t\tres = []\n+\t\tfor ( sha1, typename, name, note, ) in self.refs():\n+\t\t\tif typename != \"commit\":\n+\t\t\t\tcontinue\n+\n+\t\t\tmatch = pattern.match(note)\n+\t\t\tif not match:\n+\t\t\t\tcontinue\n+\n+\t\t\tif match.group(2) == prefix and match.group(1) == self.uuid:\n+\t\t\t\trev = int(match.group(3))\n+\t\t\t\tres.append(( sha1, rev ))\n+\n+\t\tif len(res) == 0:\n+\t\t\treturn ( None, 1, )\n+\n+\t\tres.sort(lambda a,b: a[1] < b[1])\n+\t\treturn res[0]\n+\n+\n+\t# Run fast import of revision `old` up to `new`, only considering files\n+\t# under the given prefix. Use `sha1` as the parent of the first commit.\n+\t# Return the git commit name that corresponds to the last revision so\n+\t# we can report it back to git.\n+\tdef fi(self, prefix, old, new, sha1):\n+\t\tfor rev in xrange(old or 1, new + 1):\n+\t\t\tsha1 = self.feed(rev, prefix, sha1)\n+\n+\t\treturn sha1\t\n+\n+\n+\t# Feed the fast-import helper with the given revision\n+\tdef feed(self, rev, prefix, sha1):\n+\t\t# Open the root at that revision and get the changes\n+\t\troot = svn.fs.svn_fs_revision_root(self.fs, rev)\n+\t\tchanges = svn.fs.svn_fs_paths_changed(root)\n+\n+\t\ti, file_changes = 1, []\n+\t\tfor path, change_type in changes.iteritems():\n+\t\t\tif svn.fs.svn_fs_is_dir(root, path):\n+\t\t\t\tcontinue\n+\t\t\tif not path.startswith(prefix):\n+\t\t\t\tcontinue\n+\t\t\n+\t\t\trealpath = path.replace(prefix, '')\n+\n+\t\t\tc_t = ct_short[change_type.change_kind]\n+\t\t\tif c_t == 'D':\n+\t\t\t\tfile_changes.append(\"D %s\\n\" % realpath)\n+\t\t\telse:\n+\t\t\t\tfile_changes.append(\"M 644 :%s %s\\n\" % (i, realpath))\n+\n+\t\t\t\tlength = int(svn.fs.svn_fs_file_length(root, path))\n+\t\t\t\tstream = svn.fs.svn_fs_file_contents(root, path)\n+\t\t\t\tread = lambda x: svn.core.svn_stream_read(stream, x)\n+\t\t\t\tself.fih.blob(i, length, read)\n+\t\t\t\tsvn.core.svn_stream_close(stream)\n+\t\t\t\ti += 1\n+\n+\t\tif len(file_changes) == 0:\n+\t\t\treturn sha1\n+\n+\t\tprops = svn.fs.svn_fs_revision_proplist(self.fs, rev)\n+\n+\t\t# Collect all the needed information to create the commit\n+\t\tmark = str(rev)\n+\t\tref = \"refs/heads/master\"\n+\t\tparents = [ sha1 ]\n+\n+\t\tsvndate = props['svn:date'][0:-8]\n+\t\tcommit_time = time.mktime(time.strptime(svndate, '%Y-%m-%dT%H:%M:%S'))\n+\t\t\n+\t\tif props.has_key('svn:author'):\n+\t\t\tauthor = \"%s <%s@localhost>\" % (props['svn:author'], props['svn:author'])\n+\t\telse:\n+\t\t\tauthor = 'nobody <nobody@localhost>'\n+\n+\t\tcommitter = ( author, int(commit_time), )\n+\t\tmessage = props['svn:log']\n+\t\t\n+\t\tsha1 = self.fih.commit(mark, ref, parents, None, committer, file_changes, message)\n+\t\t\n+\t\tnote = \"%s%s@%s\\n\" % (svn.fs.svn_fs_get_uuid(self.fs), prefix[:-1], rev)\n+\t\tself.note(sha1, note)\n+\t\t\n+\t\treturn sha1\n+\n+\n+\n+if __name__ == '__main__':\n+\thelper = RemoteHelperSubversion(sys.argv[2])\n+\thelper.run()\n-- \n1.7.3.37.gb6088b\n"},{"id":"152356","messageId":"AANLkTikQyVLyH-O-OH2yZ0B3_UKDqzcnNgtqefSCN68t@mail.gmail.com","threadId":"25319","inReplyTo":"4CA86A12.6080905@dbservice.com","subject":"Re: [RFC] New type of remote helpers","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2010-10-03T13:56:16Z","receivedAt":"2010-10-03T13:56:16Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Sun, Oct 3, 2010 at 13:33, Tomas Carnecky <tom@dbservice.com> wrote:\n> My work has the goal of making interaction with foreign SCMs more\n> natural. The work that was done on remote helpers is the right\n> direction. But the 'import' and 'export' commands are the wrong approach\n> I think.\n\nI'm not convinced that they are, but we'll see.\n\n> The problem I have with 'import' is that updating the refs is\n> left up to the remote helper (or git-fast-import). So you lose the nice\n> output from ls-remote/fetch: non-ff and other warnings etc.\n\nThis is a good point, and I like the patches addressing this.\n\n> I slightly\n> modified how the remote helpers (and fast-import) work, now they behave\n> exactly like 'core' git when fetching: Git tells the remote helper to\n> fetch some refs, the helper does that and creates a pack and git then\n> updates the refs (or not, depending on fast-forward etc).\n\nAgain, these patches I like, and would like to see them included. I'll\nprobably pick them up and send them out as part of my next\ngit-remote-hg reroll if nothing happens with them.\n\n> To test this\n> approach I created a simple remote helper for svn.\n\nI guess it suffices as a POC, but I'd have preferred to see\ncollaboration with the people working on git-remote-svn instead\n(cc-ed).\n\n> $ git ls-remote svn::/Volumes/Dump/Source/Mirror/Transmission/\n> r1017 (impure)                            trunk\n> r919 (impure)                             branches/nat-traversal\n> r480 (impure)                             branches/0.6\n>\n> Git learned to understand version numbers from foreign SCMs. Git\n> displays those as 'impure' because it knows that version exists but does\n> not know yet which git commit that version maps to.\n\nVery interesting. This is a useful feature, I approve.\n\n> $ git fetch svn::/Volumes/Dump/Source/Mirror/Transmission/\n> *:refs/remotes/svn/*\n> From svn::/Volumes/Dump/Source/Mirror/Transmission\n>  * [new branch]      trunk      -> svn/trunk\n>  * [new branch]      branches/nat-traversal -> svn/branches/nat-traversal\n>  * [new branch]      branches/0.6 -> svn/branches/0.6\n\nInteresting, if you do 'git remote add svn\nsvn::/Volumes/Dump/Source/Mirror/Transmission/' and then do 'git fetch\nsvn', do you get (more or less) the same output?\n\n> Git tells the remote helper to 'fetch r1017 trunk'.  The remote helper\n> does that, creates the pack and then tells git that it imported r1017 as\n> commit c5fed7ec. This is done with a new reply to the 'fetch' command:\n> 'map r1017 c5fed7ec'. The remote helper can use that to inform core git\n> as which git commit the impure ref was imported. Git can then update the\n> refs. At no point does the remote helper manipulate refs directly.\n\nI love this. Very elegant.\n\n> The pack is created by a heavily modified git-fast-import. The existing\n> fast-import not only creates the pack but also updates the refs. This is\n> no longer desired as git is in charge of updating the refs.\n\nNAK. I object against forking git fast-import just for this purpose.\nI'd much rather just modify git fast-import to learn to learn not\nupdate refs, which should be easy enough.\n\n> My modified\n> fast-import works like this: After creating a commit, it writes it's git\n> object name to stdout. That way the remote helper can figure out as\n> which git commits the svn revisions were imported and relay that back to\n> core git using the above described 'map' reply.\n\nWork has been underway to teach git fast-import to do just this\n(courtesy of Jonathan), no need to fork git fast-import to achieve\nthat.\n\n> $ git show --show-notes=svn svn/trunk\n> commit c5fed7ecc318363523d3ea2045e1c16a378bb10c\n> Author: livings124 <livings124@localhost>\n> Date:   Wed Oct 18 13:57:19 2006 +0000\n>\n>    more traditional toolbar icons for those afraid of change\n>\n> Notes (svn):\n>    b4697c4a-7d4c-4a30-bd92-6745580d73b3/trunk@1017\n\nVery nice! I'm definitely stealing the note-generating code for git-remote-hg.\n\n> The svn helper needs to be able to map svn revisions to git commits.\n> git-svn does this by adding the 'git-svn-id' line to each commit\n> message. I'm using git notes for that and it seems to work just fine.\n> The note contains the repo UUID, path within the repo and revision.\n\nVery elegant.\n\n> There was a challenge how to update the notes ref (refs/notes/svn). As\n> with fast-import, I did not want the remote helper to do it. Neither the\n> remote helper nor fast-import should be writing any refs. But core git\n> can only update refs which were discovered during transport->fetch(). I\n> modified the remote helper 'fetch' command and the transport->fetch()\n> function to return an optional list of refs. These are the refs that the\n> remote helper wants to update but which should not be presented to the\n> user (because these are internally used refs, such as my svn notes).\n\nNice.\n\n> So the whole session between git and my svn remote helper looks like this:\n>> list\n> < :r1017 trunk\n>> fetch :r1017 trunk\n> [helper creates the pack including history up to r1017 and associated\n> svn notes]\n> < map r1017 <commit corresponding to r1017>\n> < silent refs/notes/svn <new commit which stores the updated svn notes>\n\nLooks good.\n\n> $ git fetch svn::/Volumes/Dump/Source/Mirror/Transmission/\n> *:refs/remotes/svn/*\n> From svn::/Volumes/Dump/Source/Mirror/Transmission\n>   c5fed7e..228eaf3  trunk      -> svn/trunk\n>\n> $ git fetch svn::/Volumes/Dump/Source/Mirror/Transmission/\n> *:refs/remotes/svn/*\n> From svn::/Volumes/Dump/Source/Mirror/Transmission\n>   228eaf3..207e5e5  trunk      -> svn/trunk\n>  * [new branch]      branches/scrape -> svn/branches/scrape\n>  * [new branch]      branches/multitracker -> svn/branches/multitracker\n>  * [new branch]      branches/io -> svn/branches/io\n\nNote to other reviewers: the repository was updated in between\nsuccessive calls to 'git fetch', see below.\n\n> Updating the svn branches works like expected. The remote helper\n> automatically detects which branches it already imported (by going\n> through all refs and the attached svn notes) and creates a new pack with\n> the new commits. New branches are also detected. The svn notes are\n> updated accordingly.\n\nI assume this only works for regular svn repositories? I guess it\ndoesn't really matter to the rest of the series, since the\n'git-remote-svn' helper is more of a POC I think?\n\nMost of your extensions to the helper protocol make sense. However,\nafter re-reading your series I think we _should_ keep the 'import' and\n'export' command, so that helpers don't have to invoke 'git\nfast-import' or 'git fast-import' themselves. I suspect it will be\nmore efficient than your approach as well. Speed _is_ a very important\nconcern, to have decent support for foreign\nremotes, imports/exports should be as fast as possible. Also, this\nseries doesn't address pushing back to the foreign scm, which is very\nconvenient through the 'export' command.\n\nEither way, thank you very much for working on this!\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"152364","messageId":"20101003151304.GH17084@burratino","threadId":"25319","inReplyTo":"AANLkTikQyVLyH-O-OH2yZ0B3_UKDqzcnNgtqefSCN68t@mail.gmail.com","subject":"Re: [RFC] New type of remote helpers","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-10-03T15:13:04Z","receivedAt":"2010-10-03T15:13:04Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nSverre Rabbelier wrote:\n> On Sun, Oct 3, 2010 at 13:33, Tomas Carnecky <tom@dbservice.com> wrote:\n\n>> To test this\n>> approach I created a simple remote helper for svn.\n>\n> I guess it suffices as a POC, but I'd have preferred to see\n> collaboration with the people working on git-remote-svn instead\n> (cc-ed).\n\nJust a quick note: if this approach gets a working remote helper\nin the hands of users faster, I'm all for it.\n\nMy only concern is the name: if it is not compatible the planned\nremote helper from the summer of code project, they should probably\nget different names.  Correct me if I'm wrong, but the main\ndifferences are:\n\n - this is scripted and uses local svn working copy operations; the\n   soc project is in C and uses remote access (\"replay\")\n\n - this uses the nice ls-remote output etc.  Ram, do you think this\n   would be easy to use for remote-svn?\n\nSo, not many differences.  Maybe we can standardize the interface\nand consider them alternate implementations?\n"},{"id":"152365","messageId":"20101003153144.GA18001@burratino","threadId":"25319","inReplyTo":"1286108511-55876-5-git-send-email-tom@dbservice.com","subject":"Re: [PATCH 5/6] Introduce the git fast-import-helper","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-10-03T15:31:44Z","receivedAt":"2010-10-03T15:31:44Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Tomas Carnecky wrote:\n\n>  fast-import-helper.c | 2201 ++++++++++++++++++++++++++++++++++++++++++++++++++\n\nAaa!  Could someone send a diff of this against the usual fast-import.c?\n\nI'm sure you won't be surprised to hear I am not optimistic about the\nlong-term maintainability of two separate fast-import implementations.\nIf we get usability improvements by patching the existing fast-import,\nthat would be much better.\n"},{"id":"152366","messageId":"4CA8A504.50009@dbservice.com","threadId":"25319","inReplyTo":"20101003153144.GA18001@burratino","subject":"Re: [PATCH 5/6] Introduce the git fast-import-helper","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2010-10-03T15:45:08Z","receivedAt":"2010-10-03T15:45:08Z","isPatch":true,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"On 10/3/10 5:31 PM, Jonathan Nieder wrote:\n> Tomas Carnecky wrote:\n> \n>>  fast-import-helper.c | 2201 ++++++++++++++++++++++++++++++++++++++++++++++++++\n> \n> Aaa!  Could someone send a diff of this against the usual fast-import.c?\n> \n> I'm sure you won't be surprised to hear I am not optimistic about the\n> long-term maintainability of two separate fast-import implementations.\n> If we get usability improvements by patching the existing fast-import,\n> that would be much better.\n\nI agree. But when I started hacking on fast-import it seemed easier to\ncompletely rip out parts that I didn't need and then add the few bits\nthat I needed. I only need two new things from fast-import:\n  1) support non-numeric marks (and even this is maybe not strictly\nrequired)\n  2) dump the mark->sha1 mapping immediately after creating the object\n(I heard there is a patch somewhere that does just that)\nAll other changes are not needed. Though I think there are a few things\nwhich could be ported back to fast-import.c. I'll try to see which\nchanges make sense to be backported and will post patches.\n\ntom\n"},{"id":"152367","messageId":"AANLkTinZ6NCvKeALDBfP4z=ewkwWVwHBk=C_LmXM7OFh@mail.gmail.com","threadId":"25319","inReplyTo":"4CA8A504.50009@dbservice.com","subject":"Re: [PATCH 5/6] Introduce the git fast-import-helper","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2010-10-03T15:53:49Z","receivedAt":"2010-10-03T15:53:49Z","isPatch":true,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Sun, Oct 3, 2010 at 17:45, Tomas Carnecky <tom@dbservice.com> wrote:\n> I agree. But when I started hacking on fast-import it seemed easier to\n> completely rip out parts that I didn't need and then add the few bits\n> that I needed.\n\nSure, that would indeed be easier, but so much less maintainable :).\n\n> I only need two new things from fast-import:\n>  1) support non-numeric marks (and even this is maybe not strictly\n> required)\n\nIf this can be avoided, or worked around somehow, it would be a boon\nto performance. The current marks implementation uses a hash table\nindex by the mark number, which is O(1), very efficient.\n\n>  2) dump the mark->sha1 mapping immediately after creating the object\n> (I heard there is a patch somewhere that does just that)\n\nWhy do you need that? Wouldn't the \"write created object name to\nstdout\" not be sufficient?\n\n> All other changes are not needed. Though I think there are a few things\n> which could be ported back to fast-import.c. I'll try to see which\n> changes make sense to be backported and will post patches.\n\nI'd be very interested in those :).\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"152368","messageId":"20101003170711.GI328@kytes","threadId":"25319","inReplyTo":"20101003151304.GH17084@burratino","subject":"Re: [RFC] New type of remote helpers","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2010-10-03T17:07:12Z","receivedAt":"2010-10-03T17:07:12Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Hi Tom and Jonathan,\n\nJonathan Nieder writes:\n> Sverre Rabbelier wrote:\n> > On Sun, Oct 3, 2010 at 13:33, Tomas Carnecky <tom@dbservice.com> wrote:\n> \n> >> To test this\n> >> approach I created a simple remote helper for svn.\n> >\n> > I guess it suffices as a POC, but I'd have preferred to see\n> > collaboration with the people working on git-remote-svn instead\n> > (cc-ed).\n> \n> Just a quick note: if this approach gets a working remote helper\n> in the hands of users faster, I'm all for it.\n\nFirst off, great work on the fast-import and the remote-helper! I am\nvery impressed with the results.\n\n> My only concern is the name: if it is not compatible the planned\n> remote helper from the summer of code project, they should probably\n> get different names.  Correct me if I'm wrong, but the main\n> differences are:\n> \n>  - this is scripted and uses local svn working copy operations; the\n>    soc project is in C and uses remote access (\"replay\")\n\nYes, the name definitely needs to be changed. Maybe name it something\nalong the lines of \"local-svn\"?\n\n>  - this uses the nice ls-remote output etc.  Ram, do you think this\n>    would be easy to use for remote-svn?\n\nThis is quite awesome. Yeah, I suppose we can use it for remote-svn as\nwell.\n\n> So, not many differences.  Maybe we can standardize the interface\n> and consider them alternate implementations?\n\nThis helper can't be merged in until Tom's changes to fast-import are\nported to the current fast-import. I just hope that those changes to\nfast-import don't conflict with the changes git-remote-svn will\nneed. Frankly, I'd rather we work towards a common goal.\n\n-- Ram\n"},{"id":"152369","messageId":"4CA8BFB7.2050707@dbservice.com","threadId":"25319","inReplyTo":"AANLkTinZ6NCvKeALDBfP4z=ewkwWVwHBk=C_LmXM7OFh@mail.gmail.com","subject":"Re: [PATCH 5/6] Introduce the git fast-import-helper","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2010-10-03T17:39:03Z","receivedAt":"2010-10-03T17:39:03Z","isPatch":true,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"On 10/3/10 5:53 PM, Sverre Rabbelier wrote:\n>> I only need two new things from fast-import:\n>>  1) support non-numeric marks (and even this is maybe not strictly\n>> required)\n> \n> If this can be avoided, or worked around somehow, it would be a boon\n> to performance. The current marks implementation uses a hash table\n> index by the mark number, which is O(1), very efficient.\n\nI also use a hash table (struct hash_table from hash.h). It's indexed by\nthe atom. So it's about equally fast as the existing one but uses\nslightly more memory. I measured the speed and fih is about 5% slower\nthan fi. Also, I found out that setting max_packfile to 32MB makes the\nimport much faster (from 10 minutes down to 3m to import the sources of\ngit itself).\n\n>>  2) dump the mark->sha1 mapping immediately after creating the object\n>> (I heard there is a patch somewhere that does just that)\n> \n> Why do you need that? Wouldn't the \"write created object name to\n> stdout\" not be sufficient?\n\nI do: fprintf(stdout, \"mark :%s %s\\n\", mark, sha1_to_hex(sha1));\nOne reason why not just write the plain hash is because that's the same\nsyntax as the fih accepts in its input. This way you can do:\n  $ ( cat marks; cat fast-export-stream ) | git fast-import-helper >> marks\nand can restart at any time. Also, making the output a bit more\nstructured allows it to be easily extended in the future.\n\ntom\n"},{"id":"152427","messageId":"AANLkTi=RASZU2e+WV6kXnUH=afE=g2SoGuJFZF1QJ4=D@mail.gmail.com","threadId":"25319","inReplyTo":"4CA8BFB7.2050707@dbservice.com","subject":"Re: [PATCH 5/6] Introduce the git fast-import-helper","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2010-10-03T23:15:36Z","receivedAt":"2010-10-03T23:15:36Z","isPatch":true,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Sun, Oct 3, 2010 at 19:39, Tomas Carnecky <tom@dbservice.com> wrote:\n> I also use a hash table (struct hash_table from hash.h). It's indexed by\n> the atom. So it's about equally fast as the existing one but uses\n> slightly more memory. I measured the speed and fih is about 5% slower\n> than fi. Also, I found out that setting max_packfile to 32MB makes the\n> import much faster (from 10 minutes down to 3m to import the sources of\n> git itself).\n\nThat's curious.\n\nA 5% increase is significant. Imagine importing netbeans, (which is\nactually a use case of git-remote-hg), which takes about 4 hours. A 5%\nslowdown means the process will take more than 10 minutes longer.\n\n> I do: fprintf(stdout, \"mark :%s %s\\n\", mark, sha1_to_hex(sha1));\n> One reason why not just write the plain hash is because that's the same\n> syntax as the fih accepts in its input. This way you can do:\n>  $ ( cat marks; cat fast-export-stream ) | git fast-import-helper >> marks\n> and can restart at any time. Also, making the output a bit more\n> structured allows it to be easily extended in the future.\n\nI don't see much benefit tbh, if we want to do something like that it\ncould (relatively) easy be added to regular git fast-import with a\nfeature. So you'd start the stream with \"feature new-marks-format\" and\nonly then follow up with \"feature import-marks=...\". Ditto on the\ncommandline, `git fast-import --new-marks-format --import-marks=...\".\nIf it turns out to be very useful/popular it can be made the default\nafter warning for a full release first.\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"152634","messageId":"20101005020035.GA10818@burratino","threadId":"25319","inReplyTo":"1286108511-55876-1-git-send-email-tom@dbservice.com","subject":"Re: [PATCH 1/6] Remote helper: accept ':<value> <name>' as a response to 'list'","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-10-05T02:00:35Z","receivedAt":"2010-10-05T02:00:35Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nTomas Carnecky wrote:\n\n>                                                          The ref is first stored\n> as 'impure', meaning that it doesn't have any representation within git. Only\n> after the remote helper fetches that version into git, it can tell us which\n> git object (SHA1) that revision maps to.\n\nHmm.\n\nThe existing ls-remote output looks something like\n\n\t<object id>\tHEAD\n\t<object id>\t<refname>\n\t<object id>\t<refname>\n\t...\n\t<object id>\t<tag refname>\n\t<object id>\t<tag refname>^{}\n\t...\n\nIn particular, each line has a 40-character object id, a tab character,\nand then something like a refname.\n\nIf we want to extend that, wouldn't we need to do something like\n\n\t0000...0000\t<refname>\n\t0000...0000\t<refname> is r11\n\nto avoid breaking people's scripts?  For example, I wouldn't be surprised\nif scripts are relying on the following two properties:\n\n - each object id is 40 characters (e.g., the old fetch--tool\n   in contrib/examples relies on this)\n - each object id contains no whitespace\n\nsimply because they are convenient to script with.\n"},{"id":"152635","messageId":"20101005021133.GB10818@burratino","threadId":"25319","inReplyTo":"1286108511-55876-2-git-send-email-tom@dbservice.com","subject":"Re: [PATCH 2/6] Allow more than one keepfile in the transport","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-10-05T02:11:33Z","receivedAt":"2010-10-05T02:11:33Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Tomas Carnecky wrote:\n\n> Use an array to keep track of the pack lockfiles.\n\nShould probably use a string_list (which is the same thing :)) to\nsave some code.\n\n> --- a/transport-helper.c\n> +++ b/transport-helper.c\n> @@ -362,10 +362,7 @@ static int fetch_with_fetch(struct transport *transport,\n>  \n>  \t\tif (!prefixcmp(buf.buf, \"lock \")) {\n>  \t\t\tconst char *name = buf.buf + 5;\n> -\t\t\tif (transport->pack_lockfile)\n> -\t\t\t\twarning(\"%s also locked %s\", data->name, name);\n> -\t\t\telse\n> -\t\t\t\ttransport->pack_lockfile = xstrdup(name);\n> +\t\t\ttransport_keep(transport, name);\n\nWon't buf.buf be released before the lockfile needs to be read?\n\nSo I suspect the strdup is necessary.\n"},{"id":"152636","messageId":"20101005021833.GC10818@burratino","threadId":"25319","inReplyTo":"1286108511-55876-3-git-send-email-tom@dbservice.com","subject":"Re: [PATCH 3/6] Allow the transport fetch command to add additional refs","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-10-05T02:18:33Z","receivedAt":"2010-10-05T02:18:33Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Tomas Carnecky wrote:\n\n> Remote helpers such as those for svn may chose to save the Git SHA1 -> Subversion\n> revision mapping as notes attached to the commits (as opposed to strings in the\n> commit message itself). The helper would need to update the notes on each fetch,\n> but the user should not be bothered by the presence of that ref. The remote\n> helper can update the notes tree through fast-import and then inform Git core\n> that it should silently update the notes ref.\n\nWhat is the advantage of this approach over having fast-import update\nthe notes ref (but not the others) on its own?\n\nCurious,\nJonathan\n"},{"id":"152637","messageId":"20101005022308.GD10818@burratino","threadId":"25319","inReplyTo":"1286108511-55876-4-git-send-email-tom@dbservice.com","subject":"Re: [PATCH 4/6] Rename get_mode() to decode_tree_mode() and export it","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-10-05T02:23:08Z","receivedAt":"2010-10-05T02:23:08Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Tomas Carnecky wrote:\n\n> Other sources (fast-import-helper.c) may want to use this function\n> to parse trees.\n\nSounds like a reasonable idea, but you've left out what the function\ndoes.\n\nSo:\n\n get_mode() parses an octal \"100644 \" string as might\n appear in a tree object.  Returns a pointer to the first\n character after the mode string, or NULL on error.\n"},{"id":"152638","messageId":"20101005022623.GE10818@burratino","threadId":"25319","inReplyTo":"1286108511-55876-6-git-send-email-tom@dbservice.com","subject":"Re: [PATCH 6/6] Add git-remote-svn","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-10-05T02:26:23Z","receivedAt":"2010-10-05T02:26:23Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Tomas Carnecky wrote:\n\n> This is an experimental git remote helper for svn repositories. It uses\n> the new git fast-import-helper.\n\nNot sure how to review this without some sense of how the fast-import-helper\nworks.\n\n> It uses notes to save the git commit -> svn revision\n> mapping (refs/notes/svn).\n\nGood. :)\n\nCaching the mapping back in the other direction would be very simple.\nA simple array of commit hashes indexed by svn revision number would\nnot take up much space in memory and could easily be randomly accessed\non disk.\n"},{"id":"152933","messageId":"AANLkTi=4Vc8456TVMHPTmLg=4VyFqtje6Mnz73VsHrKs@mail.gmail.com","threadId":"25319","inReplyTo":"20101005020035.GA10818@burratino","subject":"Re: [PATCH 1/6] Remote helper: accept ':<value> <name>' as a response to 'list'","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2010-10-07T21:17:18Z","receivedAt":"2010-10-07T21:17:18Z","isPatch":true,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Tue, Oct 5, 2010 at 04:00, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> In particular, each line has a 40-character object id, a tab character,\n> and then something like a refname.\n>\n> If we want to extend that, wouldn't we need to do something like\n>\n>        0000...0000     <refname>\n>        0000...0000     <refname> is r11\n>\n> to avoid breaking people's scripts?\n\nI don't know... I'd suspect that people's scripts are going to break\nregardless, but if this is being used in scripts I agree we should\nmake it as easy as possible for scripters to parse this output.\n\n-- \nCheers,\n\nSverre Rabbelier\n"}]}