{"thread":{"id":"9051","subject":"[PATCH 0/6] Introduce commit notes","startedAt":"2007-07-15T23:19:17Z","lastAt":"2007-07-20T04:59:09Z","messageCount":42,"participants":["Johannes Schindelin","Junio C Hamano","Shawn O. Pearce","Andy Parkins","Linus Torvalds","Wincent Colaiuta","Sven Verdoolaege","Adam Hayek","Olivier Galibert"],"isPatch":true,"patchVersion":1,"patchTotal":6},"messages":[{"id":"47470","messageId":"Pine.LNX.4.64.0707152326080.14781@racer.site","threadId":"9051","inReplyTo":null,"subject":"[PATCH 0/6] Introduce commit notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-15T23:19:17Z","receivedAt":"2007-07-15T23:19:17Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nThis patch series replaces the series I sent out a while ago, which added \n\"commit annotations\".  Since \"commit notes\" was liked much better, here \nthey are.\n\nIt picks up the same idea, having a pseudo-branch whose revisions contain \na .git/objects/??/* like file structure, and whose blobs are the commit \nnotes.\n\nBy default, that pseudo-branch is \"refs/notes/commits\", but it is \noverridable by the config variable core.notesRef, which in turn can be \noverridden by the environment variable GIT_NOTES_REF.  If the given ref \ndoes not exist yet, it is interpreted as empty.\n\nThe biggest obstacle was a thinko about the scalability.  Tree objects \ntake free form name entries, and therefore a binary search by name is not \npossible.\n\nPatch 6/6 is only a WIP patch, but it shows the road ahead.  It adds code \nto generate .git/notes-index from refs/notes/commits (or any other ref you \nspecify as notes ref), which is reused until refs/notes/commits^{tree} \nchanges.  Patch 6/6 is only meant to assess which data structure yields \nbest performance, and how big the costs are.\n\nHowever, as long as there are no public, fetchable commit notes, I think \nthe first 5 patches are safe for application and testing.\n\nCiao,\nDscho\n\nJohannes Schindelin (6):\n      Rename git_one_line() to git_line_length() and export it\n      Introduce commit notes\n      Add git-notes\n      Add a test script for \"git notes\"\n      Document git-notes\n      notes: add notes-index for a substantial speedup.\n\n .gitignore                       |    1 +\n Documentation/cmd-list.perl      |    1 +\n Documentation/config.txt         |   15 ++\n Documentation/git-notes.txt      |   45 ++++\n Makefile                         |    5 +-\n cache.h                          |    1 +\n commit.c                         |   15 +-\n commit.h                         |    1 +\n config.c                         |    5 +\n environment.c                    |    1 +\n git-notes.sh                     |   61 ++++++\n notes.c                          |  416 ++++++++++++++++++++++++++++++++++++++\n notes.h                          |    9 +\n t/t3301-notes.sh                 |   63 ++++++\n t/t3302-notes-index-expensive.sh |  118 +++++++++++\n 15 files changed, 750 insertions(+), 7 deletions(-)\n create mode 100644 Documentation/git-notes.txt\n create mode 100755 git-notes.sh\n create mode 100644 notes.c\n create mode 100644 notes.h\n create mode 100755 t/t3301-notes.sh\n create mode 100755 t/t3302-notes-index-expensive.sh\n"},{"id":"47472","messageId":"Pine.LNX.4.64.0707160021390.14781@racer.site","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707152326080.14781@racer.site","subject":"[PATCH 1/6] Rename git_one_line() to git_line_length() and export it","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-15T23:22:00Z","receivedAt":"2007-07-15T23:22:00Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nThe function get_one_line() really returns the line length, not the\nwhole line, but it is really useful, so do not hide it in commit.c,\nunder the wrong name.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n commit.c |   10 +++++-----\n commit.h |    1 +\n 2 files changed, 6 insertions(+), 5 deletions(-)\n\ndiff --git a/commit.c b/commit.c\nindex 4c5dfa9..0c350bc 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -458,7 +458,7 @@ void clear_commit_marks(struct commit *commit, unsigned int mark)\n /*\n  * Generic support for pretty-printing the header\n  */\n-static int get_one_line(const char *msg, unsigned long len)\n+int get_line_length(const char *msg, unsigned long len)\n {\n \tint ret = 0;\n \n@@ -950,7 +950,7 @@ static void pp_header(enum cmit_fmt fmt,\n \tfor (;;) {\n \t\tconst char *line = *msg_p;\n \t\tchar *dst;\n-\t\tint linelen = get_one_line(*msg_p, *len_p);\n+\t\tint linelen = get_line_length(*msg_p, *len_p);\n \t\tunsigned long len;\n \n \t\tif (!linelen)\n@@ -1041,7 +1041,7 @@ static void pp_title_line(enum cmit_fmt fmt,\n \ttitle = xmalloc(title_alloc);\n \tfor (;;) {\n \t\tconst char *line = *msg_p;\n-\t\tint linelen = get_one_line(line, *len_p);\n+\t\tint linelen = get_line_length(line, *len_p);\n \t\t*msg_p += linelen;\n \t\t*len_p -= linelen;\n \n@@ -1118,7 +1118,7 @@ static void pp_remainder(enum cmit_fmt fmt,\n \tint first = 1;\n \tfor (;;) {\n \t\tconst char *line = *msg_p;\n-\t\tint linelen = get_one_line(line, *len_p);\n+\t\tint linelen = get_line_length(line, *len_p);\n \t\t*msg_p += linelen;\n \t\t*len_p -= linelen;\n \n@@ -1214,7 +1214,7 @@ unsigned long pretty_print_commit(enum cmit_fmt fmt,\n \n \t/* Skip excess blank lines at the beginning of body, if any... */\n \tfor (;;) {\n-\t\tint linelen = get_one_line(msg, len);\n+\t\tint linelen = get_line_length(msg, len);\n \t\tint ll = linelen;\n \t\tif (!linelen)\n \t\t\tbreak;\ndiff --git a/commit.h b/commit.h\nindex 467872e..fc6df23 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -60,6 +60,7 @@ enum cmit_fmt {\n \tCMIT_FMT_UNSPECIFIED,\n };\n \n+extern int get_line_length(const char *msg, unsigned long len);\n extern enum cmit_fmt get_commit_format(const char *arg);\n extern unsigned long pretty_print_commit(enum cmit_fmt fmt, const struct commit *, unsigned long len, char **buf_p, unsigned long *space_p, int abbrev, const char *subject, const char *after_subject, enum date_mode dmode);\n \n-- \n1.5.3.rc1.2718.gd2dc9-dirty\n"},{"id":"47473","messageId":"Pine.LNX.4.64.0707160022560.14781@racer.site","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707152326080.14781@racer.site","subject":"[PATCH 2/6] Introduce commit notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-15T23:23:11Z","receivedAt":"2007-07-15T23:23:11Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nCommit notes are blobs which are shown together with the commit\nmessage.  These blobs are taken from the notes ref, which you can\nconfigure by the config variable core.notesRef, which in turn can\nbe overridden by the environment variable GIT_NOTES_REF.\n\nThe notes ref is a branch which contains trees much like the\nloose object trees in .git/objects/.  In other words, to get\nat the commit notes for a given SHA-1, take the first two\nhex characters as directory name, and the remaining 38 hex\ncharacters as base name, and look that up in the notes ref.\n\nThe rationale for putting this information into a ref is this: we\nwant to be able to fetch and possibly union-merge the notes,\nmaybe even look at the date when a note was introduced, and we\nwant to store them efficiently together with the other objects.\n\nThere is one severe shortcoming, though.  Since tree objects can\ncontain file names of a variable length, it is not possible to\ndo a binary search for the correct base name in the tree object's\ncontents.  Therefore this approach does not scale well, because\nthe average lookup time will be proportional to the number of\ncommit objects, and therefore the slowdown will be quadratic in\nthat number.\n\nHowever, a remedy is near: in a later commit, a .git/notes-index\nwill be introduced, a cached mapping from commits to commit notes,\nto be written when the tree name of the notes ref changes.  In\ncase that notes-index cannot be written, the current (possibly\nslow) code will come into effect again.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/config.txt |   15 +++++++++++\n Makefile                 |    3 +-\n cache.h                  |    1 +\n commit.c                 |    5 +++\n config.c                 |    5 +++\n environment.c            |    1 +\n notes.c                  |   64 ++++++++++++++++++++++++++++++++++++++++++++++\n notes.h                  |    9 ++++++\n 8 files changed, 102 insertions(+), 1 deletions(-)\n create mode 100644 notes.c\n create mode 100644 notes.h\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex d0e9a17..5fe833d 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -285,6 +285,21 @@ core.pager::\n \tThe command that git will use to paginate output.  Can be overridden\n \twith the `GIT_PAGER` environment variable.\n \n+core.notesRef::\n+\tWhen showing commit messages, also show notes which are stored in\n+\tthe given ref.  This ref is expected to contain paths of the form\n+\t??/*, where the directory name consists of the first two\n+\tcharacters of the commit name, and the base name consists of\n+\tthe remaining 38 characters.\n++\n+If such a path exists in the given ref, the referenced blob is read, and\n+appended to the commit message, separated by a \"Notes:\" line.  If the\n+given ref itself does not exist, it is not an error, but means that no\n+notes should be print.\n++\n+This setting defaults to \"refs/notes/commits\", and can be overridden by\n+the `GIT_NOTES_REF` environment variable.\n+\n alias.*::\n \tCommand aliases for the gitlink:git[1] command wrapper - e.g.\n \tafter defining \"alias.last = cat-file commit HEAD\", the invocation\ndiff --git a/Makefile b/Makefile\nindex d7541b4..119d949 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -322,7 +322,8 @@ LIB_OBJS = \\\n \twrite_or_die.o trace.o list-objects.o grep.o match-trees.o \\\n \talloc.o merge-file.o path-list.o help.o unpack-trees.o $(DIFF_OBJS) \\\n \tcolor.o wt-status.o archive-zip.o archive-tar.o shallow.o utf8.o \\\n-\tconvert.o attr.o decorate.o progress.o mailmap.o symlinks.o remote.o\n+\tconvert.o attr.o decorate.o progress.o mailmap.o symlinks.o remote.o \\\n+\tnotes.o\n \n BUILTIN_OBJS = \\\n \tbuiltin-add.o \\\ndiff --git a/cache.h b/cache.h\nindex 328c1ad..c89cac5 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -309,6 +309,7 @@ extern size_t packed_git_window_size;\n extern size_t packed_git_limit;\n extern size_t delta_base_cache_limit;\n extern int auto_crlf;\n+extern char *notes_ref_name;\n \n #define GIT_REPO_VERSION 0\n extern int repository_format_version;\ndiff --git a/commit.c b/commit.c\nindex 0c350bc..3529b6a 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -6,6 +6,7 @@\n #include \"interpolate.h\"\n #include \"diff.h\"\n #include \"revision.h\"\n+#include \"notes.h\"\n \n int save_commit_buffer = 1;\n \n@@ -1254,6 +1255,10 @@ unsigned long pretty_print_commit(enum cmit_fmt fmt,\n \t */\n \tif (fmt == CMIT_FMT_EMAIL && offset <= beginning_of_body)\n \t\tbuf[offset++] = '\\n';\n+\n+\tif (fmt != CMIT_FMT_ONELINE)\n+\t\tget_commit_notes(commit, buf_p, &offset, space_p);\n+\n \tbuf[offset] = '\\0';\n \tfree(reencoded);\n \treturn offset;\ndiff --git a/config.c b/config.c\nindex f89a611..05d2ad6 100644\n--- a/config.c\n+++ b/config.c\n@@ -395,6 +395,11 @@ int git_default_config(const char *var, const char *value)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.notesref\")) {\n+\t\tnotes_ref_name = xstrdup(value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"user.name\")) {\n \t\tstrlcpy(git_default_name, value, sizeof(git_default_name));\n \t\treturn 0;\ndiff --git a/environment.c b/environment.c\nindex f83fb9e..2e677d3 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -34,6 +34,7 @@ char *pager_program;\n int pager_in_use;\n int pager_use_color = 1;\n int auto_crlf = 0;\t/* 1: both ways, -1: only when adding git objects */\n+char *notes_ref_name;\n \n static const char *git_dir;\n static char *git_object_dir, *git_index_file, *git_refs_dir, *git_graft_file;\ndiff --git a/notes.c b/notes.c\nnew file mode 100644\nindex 0000000..5d1bb1a\n--- /dev/null\n+++ b/notes.c\n@@ -0,0 +1,64 @@\n+#include \"cache.h\"\n+#include \"commit.h\"\n+#include \"notes.h\"\n+#include \"refs.h\"\n+\n+static int initialized;\n+\n+void get_commit_notes(const struct commit *commit,\n+\t\tchar **buf_p, unsigned long *offset_p, unsigned long *space_p)\n+{\n+\tchar name[80];\n+\tconst char *hex;\n+\tunsigned char sha1[20];\n+\tchar *msg;\n+\tunsigned long msgoffset, msglen;\n+\tenum object_type type;\n+\n+\tif (!initialized) {\n+\t\tconst char *env = getenv(GIT_NOTES_REF);\n+\t\tif (env) {\n+\t\t\tif (notes_ref_name)\n+\t\t\t\tfree(notes_ref_name);\n+\t\t\tnotes_ref_name = xstrdup(getenv(GIT_NOTES_REF));\n+\t\t} else if (!notes_ref_name)\n+\t\t\tnotes_ref_name = xstrdup(\"refs/notes/commits\");\n+\t\tif (notes_ref_name && read_ref(notes_ref_name, sha1)) {\n+\t\t\tfree(notes_ref_name);\n+\t\t\tnotes_ref_name = NULL;\n+\t\t}\n+\t\tinitialized = 1;\n+\t}\n+\tif (!notes_ref_name)\n+\t\treturn;\n+\n+\thex = sha1_to_hex(commit->object.sha1);\n+\tsnprintf(name, sizeof(name), \"%s:%.*s/%.*s\",\n+\t\t\tnotes_ref_name, 2, hex, 38, hex + 2);\n+\tif (get_sha1(name, sha1))\n+\t\treturn;\n+\n+\tif (!(msg = read_sha1_file(sha1, &type, &msglen)) || !msglen)\n+\t\treturn;\n+\t/* we will end the annotation by a newline anyway. */\n+\tif (msg[msglen - 1] == '\\n')\n+\t\tmsglen--;\n+\n+\tALLOC_GROW(*buf_p, *offset_p + 14 + msglen, *space_p);\n+\t*offset_p += sprintf(*buf_p + *offset_p, \"\\nNotes:\\n\");\n+\n+\tfor (msgoffset = 0; msgoffset < msglen;) {\n+\t\tint linelen = get_line_length(msg + msgoffset, msglen);\n+\n+\t\tALLOC_GROW(*buf_p, *offset_p + linelen + 6, *space_p);\n+\t\t*offset_p += sprintf(*buf_p + *offset_p,\n+\t\t\t\t\"    %.*s\", linelen, msg + msgoffset);\n+\t\tmsgoffset += linelen;\n+\t}\n+\tALLOC_GROW(*buf_p, *offset_p + 1, *space_p);\n+\t(*buf_p)[*offset_p] = '\\n';\n+\t(*offset_p)++;\n+\tfree(msg);\n+}\n+\n+\ndiff --git a/notes.h b/notes.h\nnew file mode 100644\nindex 0000000..aed80e7\n--- /dev/null\n+++ b/notes.h\n@@ -0,0 +1,9 @@\n+#ifndef NOTES_H\n+#define NOTES_H\n+\n+void get_commit_notes(const struct commit *commit,\n+\tchar **buf_p, unsigned long *offset_p, unsigned long *space_p);\n+\n+#define GIT_NOTES_REF \"GIT_NOTES_REF\"\n+\n+#endif\n-- \n1.5.3.rc1.2718.gd2dc9-dirty\n"},{"id":"47474","messageId":"Pine.LNX.4.64.0707160023360.14781@racer.site","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707152326080.14781@racer.site","subject":"[PATCH 3/6] Add git-notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-15T23:23:49Z","receivedAt":"2007-07-15T23:23:49Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nThis script allows you to edit and show commit notes easily.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n .gitignore   |    1 +\n Makefile     |    2 +-\n git-notes.sh |   61 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 63 insertions(+), 1 deletions(-)\n create mode 100755 git-notes.sh\n\ndiff --git a/.gitignore b/.gitignore\nindex 20ee642..125613f 100644\n--- a/.gitignore\n+++ b/.gitignore\n@@ -83,6 +83,7 @@ git-mktag\n git-mktree\n git-name-rev\n git-mv\n+git-notes\n git-pack-redundant\n git-pack-objects\n git-pack-refs\ndiff --git a/Makefile b/Makefile\nindex 119d949..10a9342 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -213,7 +213,7 @@ SCRIPT_SH = \\\n \tgit-merge-resolve.sh git-merge-ours.sh \\\n \tgit-lost-found.sh git-quiltimport.sh git-submodule.sh \\\n \tgit-filter-branch.sh \\\n-\tgit-stash.sh\n+\tgit-stash.sh git-notes.sh\n \n SCRIPT_PERL = \\\n \tgit-add--interactive.perl \\\ndiff --git a/git-notes.sh b/git-notes.sh\nnew file mode 100755\nindex 0000000..e0ad0b9\n--- /dev/null\n+++ b/git-notes.sh\n@@ -0,0 +1,61 @@\n+#!/bin/sh\n+\n+USAGE=\"(edit | show) [commit]\"\n+. git-sh-setup\n+\n+test -n \"$3\" && usage\n+\n+test -z \"$GIT_NOTES_REF\" && GIT_NOTES_REF=\"$(git config core.notesref)\"\n+test -z \"$GIT_NOTES_REF\" &&\n+\tdie \"No notes ref set.\"\n+\n+COMMIT=$(git rev-parse --verify --default HEAD \"$2\")\n+NAME=$(echo $COMMIT | sed \"s/^../&\\//\")\n+\n+case \"$1\" in\n+edit)\n+\tMESSAGE=\"$GIT_DIR\"/new-notes\n+\tGIT_NOTES_REF= git log -1 $COMMIT | sed \"s/^/#/\" > \"$MESSAGE\"\n+\n+\tGIT_INDEX_FILE=\"$MESSAGE\".idx\n+\texport GIT_INDEX_FILE\n+\n+\tCURRENT_HEAD=$(git show-ref $GIT_NOTES_REF | cut -f 1 -d ' ')\n+\tif [ -z \"$CURRENT_HEAD\" ]; then\n+\t\tPARENT=\n+\telse\n+\t\tPARENT=\"-p $OLDTIP\"\n+\t\tgit read-tree $GIT_NOTES_REF || die \"Could not read index\"\n+\t\tgit cat-file blob :$NAME >> \"$MESSAGE\" 2> /dev/null\n+\tfi\n+\n+\t${VISUAL:-${EDITOR:-vi}} \"$MESSAGE\"\n+\n+\tgrep -v ^# < \"$MESSAGE\" | git stripspace > \"$MESSAGE\".processed\n+\tmv \"$MESSAGE\".processed \"$MESSAGE\"\n+\tif [ -z \"$(cat \"$MESSAGE\")\" ]; then\n+\t\ttest -z \"$CURRENT_HEAD\" &&\n+\t\t\tdie \"Will not initialise with empty tree\"\n+\t\tgit update-index --force-remove $NAME ||\n+\t\t\tdie \"Could not update index\"\n+\telse\n+\t\tBLOB=$(git hash-object -w \"$MESSAGE\") ||\n+\t\t\tdie \"Could not write into object database\"\n+\t\tgit update-index --add --cacheinfo 0644 $BLOB $NAME ||\n+\t\t\tdie \"Could not write index\"\n+\tfi\n+\n+\tTREE=$(git write-tree) || die \"Could not write tree\"\n+\tNEW_HEAD=$(: | git commit-tree $TREE $PARENT) ||\n+\t\tdie \"Could not annotate\"\n+\tcase \"$CURRENT_HEAD\" in\n+\t'') git update-ref $GIT_NOTES_REF $NEW_HEAD ;;\n+\t*) git update-ref $GIT_NOTES_REF $NEW_HEAD $CURRENT_HEAD;;\n+\tesac\n+;;\n+show)\n+\tgit show \"$GIT_NOTES_REF\":$NAME\n+;;\n+*)\n+\tusage\n+esac\n-- \n1.5.3.rc1.2718.gd2dc9-dirty\n"},{"id":"47475","messageId":"Pine.LNX.4.64.0707160024060.14781@racer.site","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707152326080.14781@racer.site","subject":"[PATCH 4/6] Add a test script for \"git notes\"","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-15T23:24:16Z","receivedAt":"2007-07-15T23:24:16Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nIncidentally, a test for \"git notes\" implies a test for the\nwhole commit notes machinery.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/t3301-notes.sh |   63 ++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 files changed, 63 insertions(+), 0 deletions(-)\n create mode 100755 t/t3301-notes.sh\n\ndiff --git a/t/t3301-notes.sh b/t/t3301-notes.sh\nnew file mode 100755\nindex 0000000..eb50191\n--- /dev/null\n+++ b/t/t3301-notes.sh\n@@ -0,0 +1,63 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2007 Johannes E. Schindelin\n+#\n+\n+test_description='Test commit notes'\n+\n+. ./test-lib.sh\n+\n+test_expect_success setup '\n+\t: > a1 &&\n+\tgit add a1 &&\n+\ttest_tick &&\n+\tgit commit -m 1st &&\n+\t: > a2 &&\n+\tgit add a2 &&\n+\ttest_tick &&\n+\tgit commit -m 2nd\n+'\n+\n+cat > fake_editor.sh << EOF\n+echo \"\\$MSG\" > \"\\$1\"\n+echo \"\\$MSG\" >& 2\n+EOF\n+chmod a+x fake_editor.sh\n+VISUAL=\"$(pwd)\"/fake_editor.sh\n+export VISUAL\n+\n+\n+test_expect_success 'need notes ref' '\n+\t! MSG=1 git notes edit &&\n+\t! MSG=2 git notes show\n+'\n+\n+test_expect_success 'create notes' '\n+\tgit config core.notesRef refs/notes/commits &&\n+\tMSG=b1 git notes edit &&\n+cat .git/new-notes &&\n+test b1 = \"$(cat .git/new-notes)\" &&\n+\ttest 1 = $(git ls-tree refs/notes/commits | wc -l) &&\n+\ttest b1 = $(git notes show) &&\n+\tgit show HEAD^ &&\n+\t! git notes show HEAD^\n+'\n+\n+cat > expect << EOF\n+commit 268048bfb8a1fb38e703baceb8ab235421bf80c5\n+Author: A U Thor <author@example.com>\n+Date:   Thu Apr 7 15:14:13 2005 -0700\n+\n+    2nd\n+\n+Notes:\n+    b1\n+EOF\n+\n+test_expect_success 'show notes' '\n+\t! (git cat-file commit HEAD | grep b1) &&\n+\tgit log -1 > output &&\n+\tgit diff expect output\n+'\n+\n+test_done\n-- \n1.5.3.rc1.2718.gd2dc9-dirty\n"},{"id":"47476","messageId":"Pine.LNX.4.64.0707160024440.14781@racer.site","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707152326080.14781@racer.site","subject":"[PATCH 5/6] Document git-notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-15T23:24:55Z","receivedAt":"2007-07-15T23:24:55Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/cmd-list.perl |    1 +\n Documentation/git-notes.txt |   45 +++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 46 insertions(+), 0 deletions(-)\n create mode 100644 Documentation/git-notes.txt\n\ndiff --git a/Documentation/cmd-list.perl b/Documentation/cmd-list.perl\nindex 2143995..f05e291 100755\n--- a/Documentation/cmd-list.perl\n+++ b/Documentation/cmd-list.perl\n@@ -140,6 +140,7 @@ git-mergetool                           ancillarymanipulators\n git-mktag                               plumbingmanipulators\n git-mktree                              plumbingmanipulators\n git-mv                                  mainporcelain\n+git-notes                               mainporcelain\n git-name-rev                            plumbinginterrogators\n git-pack-objects                        plumbingmanipulators\n git-pack-redundant                      plumbinginterrogators\ndiff --git a/Documentation/git-notes.txt b/Documentation/git-notes.txt\nnew file mode 100644\nindex 0000000..331ed89\n--- /dev/null\n+++ b/Documentation/git-notes.txt\n@@ -0,0 +1,45 @@\n+git-notes(1)\n+============\n+\n+NAME\n+----\n+git-notes - Add commit notes\n+\n+SYNOPSIS\n+--------\n+[verse]\n+'git-notes' (edit | show) [commit\n+\n+DESCRIPTION\n+-----------\n+This command allows you to add notes to commit messages, after the\n+fact.  To discern these notes from the message stored in the commit\n+object, the notes are indented like the message, after an unindented\n+line saying \"Notes:\".\n+\n+To enable commit notes, you have to set the config variable\n+core.notesRef to something like \"refs/notes/commits\".  This setting\n+can be overridden by the environment variable \"GIT_NOTES_REF\".\n+\n+\n+SUBCOMMANDS\n+-----------\n+\n+edit::\n+\tEdit the notes for a given commit (defaults to HEAD).\n+\n+show::\n+\tShow the notes for a given commit (defaults to HEAD).\n+\n+\n+Author\n+------\n+Written by Johannes Schindelin <johannes.schindelin@gmx.de>\n+\n+Documentation\n+-------------\n+Documentation by Johannes Schindelin\n+\n+GIT\n+---\n+Part of the gitlink:git[7] suite\n-- \n1.5.3.rc1.2718.gd2dc9-dirty\n"},{"id":"47477","messageId":"Pine.LNX.4.64.0707160025480.14781@racer.site","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707152326080.14781@racer.site","subject":"[WIP PATCH 6/6] notes: add notes-index for a substantial speedup.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-15T23:26:10Z","receivedAt":"2007-07-15T23:26:10Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nActually, this commit adds two methods for a notes index:\n\n- a sorted list with a fan out to help binary search, and\n- a modified hash table.\n\nIt also adds a test which is used to determine the best algorithm.\n---\n\tNot signed off because this is not suitable to be applied as-is.\n\tIt is only meant to test the different approaches.\n\n notes.c                          |  392 ++++++++++++++++++++++++++++++++++++--\n t/t3302-notes-index-expensive.sh |  118 ++++++++++++\n 2 files changed, 490 insertions(+), 20 deletions(-)\n create mode 100755 t/t3302-notes-index-expensive.sh\n\ndiff --git a/notes.c b/notes.c\nindex 5d1bb1a..5a90abf 100644\n--- a/notes.c\n+++ b/notes.c\n@@ -1,10 +1,370 @@\n #include \"cache.h\"\n #include \"commit.h\"\n+#include \"tree-walk.h\"\n #include \"notes.h\"\n #include \"refs.h\"\n \n static int initialized;\n \n+/*\n+ * There are two choices of data structure for the notes index.\n+ *\n+ * A) Fan out enhanced sorted list.\n+ *\n+ * This is a regular sorted list with a 256 entry fan out.  In other words,\n+ * every time an entry is looked up, a binary search is performed over the\n+ * sublist defined by the first byte of the SHA-1.\n+ *\n+ * The disadvantage is an average runtime logarithmic in the number of\n+ * commit notes.  The advantages are a compact representation on disk,\n+ * and a _guaranteed_ logarithmic runtime.\n+ *\n+ * You could even squeeze out one more byte per entry, since the\n+ * first byte is known from the fan out list.  This would complicate our\n+ * algorithm, though.\n+ *\n+ * B) Hash\n+ *\n+ * This is not your classical hash.  It is _mostly_ like a hash, with\n+ * a few notable exceptions:\n+ *\n+ * - it is possibly larger than size suggests: since it is file based,\n+ *   it is easier to write at the end than to wrap around.\n+ *\n+ * - as a consequence we can make the entries _strictly_ sorted. This\n+ *   is not only nice to look at, but makes incremental updates much,\n+ *   much easier.\n+ *\n+ * The disadvantages of a hash is its loose packing.  In order to operate\n+ * reasonably well, it needs a size roughly double the number of entries.\n+ * It also has a worst runtime linear in the number of entries.\n+ *\n+ * The advantage is an expected constant lookup time.\n+ *\n+ * The performance of a hash map depends highly on a good hashing\n+ * algorithm, to avoid collisions.  Lucky us!  SHA-1 is a pretty good\n+ * hashing algorithm.\n+ *\n+ * There is another advantage to hash maps: with not much effort, the\n+ * incremental update can be performed in place, relying on O_TRUNC to\n+ * detect interruptions.  This operation has an expected constant runtime.\n+ */\n+\n+struct notes_entry {\n+\tunsigned char commit_sha1[20];\n+\tunsigned char notes_sha1[20];\n+};\n+\n+struct notes_index {\n+\tchar signature[4]; /* FANO for fan our, HASH for hash */\n+\n+\tunsigned char tree_sha1[20];\n+\tunsigned char subtree_sha1[256][20]; /* for incremental caching */\n+\toff_t offsets[256]; /* for fan out */\n+\toff_t count, size; /* for hash */\n+} notes_index;\n+\n+static int notes_index_fd;\n+static int (*get_notes)(const unsigned char *commit_sha1,\n+\t\tunsigned char *notes_sha1);\n+\n+#define GIT_NOTES_MODE \"GIT_NOTES_MODE\"\n+static int use_hash;\n+\n+static int index_uptodate_check(struct tree *tree) {\n+\tconst char *signature = use_hash ? \"HASH\" : \"FANO\";\n+\tint fd = open(git_path(\"notes-index\"), O_RDONLY);\n+\n+\tif (fd < 0)\n+\t\treturn fd;\n+\n+\tnotes_index_fd = fd;\n+\n+\treturn read_in_full(fd, &notes_index, sizeof(notes_index)) < 0||\n+\t\t\tmemcmp(notes_index.signature, signature, 4) ||\n+\t\t\tmemcmp(notes_index.tree_sha1,\n+\t\t\t\t&tree->object.sha1, 20);\n+}\n+\n+struct lock_file update_lock;\n+\n+/* this reads the remaining 38 hexchars */\n+static int get_remaining_hexchars(unsigned char *sha1, const char *path)\n+{\n+\tint i, j1, j2;\n+\tfor (i = 0; i < 38; i += 2)\n+\t\tif ((j1 = hexval(path[i])) < 0 ||\n+\t\t\t\t(j2 = hexval(path[i + 1])) < 0)\n+\t\t\treturn -1;\n+\t\telse\n+\t\t\tsha1[1 + i / 2] = (j1 << 4) | j2;\n+\treturn path[38] != '\\0';\n+}\n+\n+static int get_notes_hash_count(struct tree *tree) {\n+\tstruct tree_desc desc, desc2;\n+\tstruct name_entry entry;\n+\tvoid *buf;\n+\tunsigned long count = 0;\n+\n+\tbuf = fill_tree_descriptor(&desc, notes_index.tree_sha1);\n+\tif (!buf)\n+\t\treturn 0;\n+\twhile (tree_entry(&desc, &entry)) {\n+\t\tvoid *buf2 = fill_tree_descriptor(&desc2, entry.sha1);\n+\t\tif (!buf2)\n+\t\t\tcontinue;\n+\t\twhile (tree_entry(&desc2, &entry))\n+\t\t\tcount++;\n+\t\tfree(buf2);\n+\t}\n+\tfree(buf);\n+\n+\treturn count;\n+}\n+\n+static unsigned long get_hash_index(const unsigned char *sha1)\n+{\n+\treturn (ntohl(*(unsigned long *)sha1) % notes_index.size);\n+}\n+\n+static int write_hash_gap(int fd, unsigned char *sha1)\n+{\n+\toff_t min_offset = sizeof(notes_index) +\n+\t\tget_hash_index(sha1) * sizeof(struct notes_entry);\n+\twhile (min_offset > lseek(fd, 0, SEEK_CUR))\n+\t\tif (write_in_full(fd, null_sha1, 20) < 0 ||\n+\t\t\t\twrite_in_full(fd, null_sha1, 20) < 0)\n+\t\t\treturn error(\"Could not write gaps in notes-index\");\n+\treturn 0;\n+}\n+\n+static int update_index(struct tree *tree) {\n+\t/*\n+\t * Fan out sorted list:\n+\t *\n+\t * Write out the header, and seek back to it, in order to update it.\n+\t * Actually only seek at the end, and make sure that you write\n+\t * something big-endian.\n+\t *\n+\t * Plan for incremental: if subtree_sha1 is equal, copy out.\n+\t * Otherwise construct, and remember in the copy of the header.\n+\t *\n+\t * Hash:\n+\t *\n+\t * Always use a power of two as size.  Not the next higher one, but\n+\t * the next next higher one.\n+\t *\n+\t * Read the tree recursively, and leave as many zeros as needed\n+\t * until the next entry comes.  Or if the entry has a hash larger\n+\t * than the last free entry, write it at once.\n+\t */\n+\n+\t/* Plan for incremental: (not in-place)\n+\t * Look at tree differences.  Write null_sha1 until next, or next\n+\t * subtree.  Continue writing until original entry is null_sha1 or\n+\t * greater than current subtree entry.\n+\t */\n+\n+\tint new_fd = hold_lock_file_for_update(&update_lock,\n+\t\t\tgit_path(\"notes-index\"), 0);\n+\tstruct tree_desc desc;\n+\tstruct name_entry entry;\n+\tvoid *buf;\n+\tint i;\n+\n+\tif (new_fd < 0)\n+\t\treturn error(\"Could not construct notes-index\");\n+\n+\tmemset(&notes_index, 0, sizeof(notes_index));\n+\thashcpy(notes_index.tree_sha1, tree->object.sha1);\n+\tnotes_index.offsets[0] = sizeof(notes_index);\n+\tif (use_hash) {\n+\t\tnotes_index.count = get_notes_hash_count(tree);\n+\t\tfor (notes_index.size = 1; notes_index.size / 2\n+\t\t\t\t>= notes_index.count; notes_index.size <<= 1)\n+\t\t\t; /* do nothing */\n+\t\tmemcpy(notes_index.signature, \"HASH\", 4);\n+\t} else\n+\t\tmemcpy(notes_index.signature, \"FANO\", 4);\n+\n+\tif (write_in_full(new_fd, &notes_index, sizeof(notes_index)) < 0)\n+\t\treturn error(\"Could not write notes-index\");\n+\n+\tbuf = fill_tree_descriptor(&desc, notes_index.tree_sha1);\n+\tif (!buf)\n+\t\treturn error(\"Could not read %s for notes-index\",\n+\t\t\t\tsha1_to_hex(notes_index.tree_sha1));\n+\n+\ti = 0;\n+\twhile (tree_entry(&desc, &entry)) {\n+\t\tint j1, j2;\n+\t\tunsigned char sha1[20];\n+\t\tstruct tree_desc desc2;\n+\t\tstruct name_entry entry2;\n+\t\tvoid *buf2;\n+\n+\t\tif (!S_ISDIR(entry.mode) ||\n+\t\t\t\t(j1 = hexval(entry.path[0])) < 0 ||\n+\t\t\t\t(j2 = hexval(entry.path[1])) < 0)\n+\t\t\tcontinue;\n+\t\tsha1[0] = j1 * 16 + j2;\n+\t\twhile (++i < sha1[0])\n+\t\t\tnotes_index.offsets[i] = notes_index.offsets[i - 1];\n+\n+\t\thashcpy(notes_index.subtree_sha1[i], entry.sha1);\n+\t\tbuf2 = fill_tree_descriptor(&desc2, entry.sha1);\n+\t\tif (!buf2)\n+\t\t\tcontinue;\n+\t\twhile(tree_entry(&desc2, &entry2)) {\n+\t\t\tif (get_remaining_hexchars(sha1, entry2.path))\n+\t\t\t\tcontinue;\n+\t\t\tif (use_hash && write_hash_gap(new_fd, sha1))\n+\t\t\t\treturn -1;\n+\t\t\tif (write_in_full(new_fd, sha1, 20) < 0 ||\n+\t\t\t\t\twrite_in_full(new_fd,\n+\t\t\t\t\t\tentry2.sha1, 20) < 0)\n+\t\t\t\treturn error(\"Could not write notes-index\");\n+\t\t}\n+\t\tfree(buf2);\n+\t\tnotes_index.offsets[i] = lseek(new_fd, 0, SEEK_CUR);\n+\t}\n+\tfree(buf);\n+\n+\twhile (++i < 256)\n+\t\tnotes_index.offsets[i] = notes_index.offsets[i - 1];\n+\n+\t/* update fan_out */\n+\tlseek(new_fd, 0, SEEK_SET);\n+\twrite(new_fd, &notes_index, sizeof(notes_index));\n+\tlseek(new_fd, notes_index.offsets[255], SEEK_SET);\n+\n+\treturn close(new_fd) || commit_lock_file(&update_lock) ||\n+\t\t(notes_index_fd = open(git_path(\"notes-index\"), O_RDONLY));\n+}\n+\n+static void *notes_mmap;\n+\n+static void unmap_notes_mmap(void)\n+{\n+\tmunmap(notes_mmap, notes_index.offsets[255]);\n+}\n+\n+static int get_notes_fan_out(const unsigned char *commit_sha1,\n+\t\tunsigned char *notes_sha1)\n+{\n+\t/*\n+\t * Header is assumed to be read.\n+\t *\n+\t * mmap() the area, and bisect.\n+\t */\n+\toff_t off;\n+\tsize_t size;\n+\tint i, i2, ret = -1;\n+\tstruct notes_entry *list;\n+\n+\ti = commit_sha1[0];\n+\toff = i ? notes_index.offsets[i - 1] : sizeof(notes_index);\n+\tsize = notes_index.offsets[i] - off;\n+\tif (!size)\n+\t\treturn -1;\n+\n+\tif (!notes_mmap) {\n+\t\tnotes_mmap = xmmap(NULL, notes_index.offsets[255],\n+\t\t\t\tPROT_READ, MAP_PRIVATE, notes_index_fd, 0);\n+\t\tatexit(unmap_notes_mmap);\n+\t}\n+\n+\tlist = (void *)((char *)notes_mmap + off);\n+\n+\ti = 0;\n+\ti2 = size / sizeof(*list);\n+\twhile (i + 1 < i2) {\n+\t\tint middle = (i + i2) / 2;\n+\t\tint cmp = hashcmp(commit_sha1, list[middle].commit_sha1);\n+\t\tif (cmp < 0)\n+\t\t\ti2 = middle;\n+\t\telse if (cmp > 0)\n+\t\t\ti = middle;\n+\t\telse {\n+\t\t\thashcpy(notes_sha1, list[middle].notes_sha1);\n+\t\t\ti = middle;\n+\t\t\tret = 0;\n+\t\t\tbreak;\n+\t\t}\n+\t}\n+\tif (i == 0 && !hashcmp(commit_sha1, list[i].commit_sha1)) {\n+\t\thashcpy(notes_sha1, list[i].notes_sha1);\n+\t\tret = 0;\n+\t}\n+\n+\treturn ret;\n+}\n+\n+static int get_notes_hash(const unsigned char *commit_sha1,\n+\t\tunsigned char *notes_sha1)\n+{\n+\t/*\n+\t * Header is assumed to be read. fd is still open.\n+\t *\n+\t * Seek to hash, read until lower or equal (0000... is lower...)\n+\t */\n+\tint i = get_hash_index(commit_sha1);\n+\tstruct notes_entry entry;\n+\n+\tlseek(notes_index_fd,\n+\t\t\tsizeof(notes_index) + i * sizeof(entry), SEEK_SET);\n+\twhile (!read_in_full(notes_index_fd, &entry, sizeof(entry)) &&\n+\t\t\t!is_null_sha1(entry.commit_sha1)) {\n+\t\tint cmp = hashcmp(commit_sha1, entry.commit_sha1);\n+\t\tif (!cmp) {\n+\t\t\thashcpy(notes_sha1, entry.notes_sha1);\n+\t\t\treturn 0;\n+\t\t} else if (cmp < 0)\n+\t\t\tbreak;\n+\t}\n+\treturn -1;\n+}\n+\n+static inline void init_notes_index(void)\n+{\n+\tconst char *env;\n+\tstruct commit *notes_ref;\n+\tunsigned char sha1[20];\n+\n+\tif (initialized)\n+\t\treturn;\n+\n+\tinitialized = 1;\n+\tenv = getenv(GIT_NOTES_REF);\n+\tif (env) {\n+\t\tif (notes_ref_name)\n+\t\t\tfree(notes_ref_name);\n+\t\tnotes_ref_name = xstrdup(env);\n+\t} else if (!notes_ref_name)\n+\t\tnotes_ref_name = xstrdup(\"refs/notes/commits\");\n+\n+\tif (!notes_ref_name)\n+\t\treturn;\n+\tif (read_ref(notes_ref_name, sha1)) {\n+\t\tfree(notes_ref_name);\n+\t\tnotes_ref_name = NULL;\n+\t\treturn;\n+\t}\n+\tenv = getenv(\"GIT_NOTES_MODE\");\n+\tif (env && !strcmp(\"HASH\", env)) {\n+\t\tuse_hash = 1;\n+\t\tget_notes = get_notes_hash;\n+\t} else if (env && !strcmp(\"FANO\", env))\n+\t\tget_notes = get_notes_fan_out;\n+\tif (get_notes && !get_sha1(notes_ref_name, sha1) &&\n+\t\t\t(notes_ref = (struct commit *)parse_object(sha1)) &&\n+\t\t\tnotes_ref->object.type == OBJ_COMMIT)\n+\t\tif (index_uptodate_check(notes_ref->tree))\n+\t\t\tif (update_index(notes_ref->tree))\n+\t\t\t\tget_notes = NULL; /* disable notes-index */\n+}\n+\n void get_commit_notes(const struct commit *commit,\n \t\tchar **buf_p, unsigned long *offset_p, unsigned long *space_p)\n {\n@@ -15,28 +375,20 @@ void get_commit_notes(const struct commit *commit,\n \tunsigned long msgoffset, msglen;\n \tenum object_type type;\n \n-\tif (!initialized) {\n-\t\tconst char *env = getenv(GIT_NOTES_REF);\n-\t\tif (env) {\n-\t\t\tif (notes_ref_name)\n-\t\t\t\tfree(notes_ref_name);\n-\t\t\tnotes_ref_name = xstrdup(getenv(GIT_NOTES_REF));\n-\t\t} else if (!notes_ref_name)\n-\t\t\tnotes_ref_name = xstrdup(\"refs/notes/commits\");\n-\t\tif (notes_ref_name && read_ref(notes_ref_name, sha1)) {\n-\t\t\tfree(notes_ref_name);\n-\t\t\tnotes_ref_name = NULL;\n-\t\t}\n-\t\tinitialized = 1;\n-\t}\n-\tif (!notes_ref_name)\n-\t\treturn;\n+\tinit_notes_index();\n \n-\thex = sha1_to_hex(commit->object.sha1);\n-\tsnprintf(name, sizeof(name), \"%s:%.*s/%.*s\",\n-\t\t\tnotes_ref_name, 2, hex, 38, hex + 2);\n-\tif (get_sha1(name, sha1))\n+\tif (!notes_ref_name)\n \t\treturn;\n+\tif (get_notes) {\n+\t\tif (get_notes(commit->object.sha1, sha1))\n+\t\t\treturn;\n+\t} else {\n+\t\thex = sha1_to_hex(commit->object.sha1);\n+\t\tsnprintf(name, sizeof(name), \"%s:%.*s/%.*s\",\n+\t\t\t\tnotes_ref_name, 2, hex, 38, hex + 2);\n+\t\tif (get_sha1(name, sha1))\n+\t\t\treturn;\n+\t}\n \n \tif (!(msg = read_sha1_file(sha1, &type, &msglen)) || !msglen)\n \t\treturn;\ndiff --git a/t/t3302-notes-index-expensive.sh b/t/t3302-notes-index-expensive.sh\nnew file mode 100755\nindex 0000000..075b8e2\n--- /dev/null\n+++ b/t/t3302-notes-index-expensive.sh\n@@ -0,0 +1,118 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2007 Johannes E. Schindelin\n+#\n+\n+test_description='Test commit notes index (expensive!)'\n+\n+. ./test-lib.sh\n+\n+test -z \"$GIT_NOTES_TIMING_TESTS\" && {\n+\tsay Skipping timing tests\n+\ttest_done\n+\texit\n+}\n+\n+create_repo () {\n+\tnumber_of_commits=$1\n+\tnr=0\n+\tparent=\n+\ttest -d .git || {\n+\tgit init &&\n+\ttree=$(git write-tree) &&\n+\twhile [ $nr -lt $number_of_commits ]; do\n+\t\ttest_tick &&\n+\t\tcommit=$(echo $nr | git commit-tree $tree $parent) ||\n+\t\t\treturn\n+\t\tparent=\"-p $commit\"\n+\t\tnr=$(($nr+1))\n+\tdone &&\n+\tgit update-ref refs/heads/master $commit &&\n+\t{\n+\t\texport GIT_INDEX_FILE=.git/temp;\n+\t\tgit rev-list HEAD | cat -n | sed \"s/^[ \t][ \t]*/ /g\" |\n+\t\twhile read nr sha1; do\n+\t\t\tblob=$(echo note $nr | git hash-object -w --stdin) &&\n+\t\t\techo $sha1 | sed \"s/^../0644 $blob 0\t&\\//\"\n+\t\tdone | git update-index --index-info &&\n+\t\ttree=$(git write-tree) &&\n+\t\ttest_tick &&\n+\t\tcommit=$(echo notes | git commit-tree $tree) &&\n+\t\tgit update-ref refs/notes/commits $commit\n+\t} &&\n+\tgit config core.notesRef refs/notes/commits\n+\t}\n+}\n+\n+test_notes () {\n+\tcount=$1 &&\n+\tgit config core.notesRef refs/notes/commits &&\n+\tgit log | grep \"^    \" > output &&\n+\ti=1 &&\n+\twhile [ $i -le $count ]; do\n+\t\techo \"    $(($count-$i))\" &&\n+\t\techo \"    note $i\" &&\n+\t\ti=$(($i+1));\n+\tdone > expect &&\n+\tgit diff expect output\n+}\n+\n+cat > time_notes << EOF\n+\tmode=\\$1\n+\ti=1\n+\twhile [ \\$i -lt \\$2 ]; do\n+\t\tcase \\$1 in\n+\t\tno-notes)\n+\t\t\texport GIT_NOTES_REF=non-existing\n+\t\t;;\n+\t\tno-cash)\n+\t\t\tunset GIT_NOTES_REF\n+\t\t\texport GIT_NOTES_MODE=NONE\n+\t\t;;\n+\t\thash-cache-create)\n+\t\t\tunset GIT_NOTES_REF\n+\t\t\texport GIT_NOTES_MODE=HASH\n+\t\t\trm .git/notes-index\n+\t\t;;\n+\t\thash-cache)\n+\t\t\tunset GIT_NOTES_REF\n+\t\t\texport GIT_NOTES_MODE=HASH\n+\t\t;;\n+\t\tsorted-list-cache-create)\n+\t\t\tunset GIT_NOTES_REF\n+\t\t\texport GIT_NOTES_MODE=FANO\n+\t\t\trm .git/notes-index\n+\t\t;;\n+\t\tsorted-list-cache)\n+\t\t\tunset GIT_NOTES_REF\n+\t\t\texport GIT_NOTES_MODE=FANO\n+\t\t;;\n+\t\tesac\n+\t\tgit log >/dev/null\n+\t\ti=\\$((\\$i+1))\n+\tdone\n+EOF\n+\n+time_notes () {\n+\tfor mode in no-notes no-cash \\\n+\t\t\thash-cache-create hash-cache \\\n+\t\t\tsorted-list-cache-create sorted-list-cache; do\n+\t\techo $mode\n+\t\t/usr/bin/time sh ../trash/time_notes $mode $1\n+\tdone\n+}\n+\n+for count in 10 100 1000; do\n+\n+\ttest -d ../trash-$count || mkdir ../trash-$count\n+\t(cd ../trash-$count;\n+\n+\ttest_expect_success \"setup $count\" \"create_repo $count\"\n+\n+\ttest_expect_success 'notes work' \"test_notes $count\"\n+\n+\ttest_expect_success 'notes timing' \"time_notes 100\"\n+\t)\n+done\n+\n+test_done\n-- \n1.5.3.rc1.2718.gd2dc9-dirty\n"},{"id":"47479","messageId":"Pine.LNX.4.64.0707160031040.14781@racer.site","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707160025480.14781@racer.site","subject":"Re: [WIP PATCH 6/6] notes: add notes-index for a substantial speedup.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-15T23:33:19Z","receivedAt":"2007-07-15T23:33:19Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n[this explains what Patch 6/6 is all about:]\n\nIf GIT_NOTES_TIMING_TESTS is set, t3302 will output some timing data.\n\nIt will create three repositories, the first with 10 commits and a\ncommit note for each, the second with 100, the third with 1000.\n\nFor each repository, it times \"git log\" 100 times in several modes:\n\n- with GIT_NOTES_REF set to a non-existing ref (should be equivalent to\n  the timings without this patch series),\n\n- with no .git/notes-index,\n\n- recreating .git/notes-index as a hash map _every_ time,\n\n- creating .git/notes-index as a hash map, and using it the rest of the time,\n\n- recreating .git/notes-index as a sorted list _every_ time, and\n\n- creating .git/notes-index as a sorted list only the first time, and then\n  using it to find the notes by binary search.\n\nHere is the output:\n\n* expecting success: create_repo 10\n*   ok 1: setup 10\n\n* expecting success: test_notes 10\ndiff --git a/expect b/output\n*   ok 2: notes work\n\n* expecting success: time_notes 100\nno-notes\n0.13user 0.08system 0:00.22elapsed 98%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+45271minor)pagefaults 0swaps\nno-cash\n0.14user 0.13system 0:00.28elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+49421minor)pagefaults 0swaps\nhash-cache-create\n0.16user 0.24system 0:00.41elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+73111minor)pagefaults 0swaps\nhash-cache\n0.10user 0.08system 0:00.18elapsed 101%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+45660minor)pagefaults 0swaps\nsorted-list-cache-create\n0.23user 0.17system 0:00.40elapsed 100%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+72056minor)pagefaults 0swaps\nsorted-list-cache\n0.12user 0.08system 0:00.20elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+46854minor)pagefaults 0swaps\n*   ok 3: notes timing\n\n* expecting success: create_repo 100\n*   ok 1: setup 100\n\n* expecting success: test_notes 100\ndiff --git a/expect b/output\n*   ok 2: notes work\n\n* expecting success: time_notes 100\nno-notes\n0.38user 0.18system 0:00.56elapsed 100%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+54386minor)pagefaults 0swaps\nno-cash\n1.45user 0.66system 0:02.13elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+93980minor)pagefaults 0swaps\nhash-cache-create\n1.56user 0.98system 0:02.56elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+132604minor)pagefaults 0swaps\nhash-cache\n0.38user 0.17system 0:00.56elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+54785minor)pagefaults 0swaps\nsorted-list-cache-create\n1.56user 0.88system 0:02.47elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+124215minor)pagefaults 0swaps\nsorted-list-cache\n0.44user 0.23system 0:00.68elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+64952minor)pagefaults 0swaps\n*   ok 3: notes timing\n\n* expecting success: create_repo 1000\n*   ok 1: setup 1000\n\n* expecting success: test_notes 1000\ndiff --git a/expect b/output\n*   ok 2: notes work\n\n* expecting success: time_notes 100\nno-notes\n2.95user 1.19system 0:04.18elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+144766minor)pagefaults 0swaps\nno-cash\n23.05user 5.86system 0:33.06elapsed 87%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+639774minor)pagefaults 0swaps\nhash-cache-create\n23.86user 7.21system 0:32.67elapsed 95%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+710958minor)pagefaults 0swaps\nhash-cache\n3.16user 1.18system 0:04.35elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+145160minor)pagefaults 0swaps\nsorted-list-cache-create\n23.22user 7.32system 0:31.66elapsed 96%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+686007minor)pagefaults 0swaps\nsorted-list-cache\n3.74user 1.81system 0:05.77elapsed 96%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+241987minor)pagefaults 0swaps\n*   ok 3: notes timing\n\nResults:\n\nThese timings were taken from a desktop machine with a few background\nprocesses running, so take them with a grain of salt.\n\nAs expected, without a .git/notes-index, it scales pretty badly.  Creating\n.git/notes-index is slightly worse than that, but it typically happens\nmuch less often than looking at a commit message.  Therefore the work is\nworth it, since the lookup _with_ .git/notes-index is in the same ball park\nas no notes at all, with the hash map being better than the sorted\nlist lookup.\n\nTherefore I will go with the hash map approach, when cleaning up patch 6/6.\nBut not tonight.\n\nCiao,\nDscho\n"},{"id":"47480","messageId":"7vlkdhck8d.fsf@assigned-by-dhcp.cox.net","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707160022560.14781@racer.site","subject":"Re: [PATCH 2/6] Introduce commit notes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-15T23:36:50Z","receivedAt":"2007-07-15T23:36:50Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> The notes ref is a branch which contains trees much like the\n> loose object trees in .git/objects/.  In other words, to get\n> at the commit notes for a given SHA-1, take the first two\n> hex characters as directory name, and the remaining 38 hex\n> characters as base name, and look that up in the notes ref.\n> ...\n> However, a remedy is near: in a later commit, a .git/notes-index\n> will be introduced, a cached mapping from commits to commit notes,\n> to be written when the tree name of the notes ref changes.  In\n> case that notes-index cannot be written, the current (possibly\n> slow) code will come into effect again.\n\nI wonder if it is worth using the fan-out tree structure for the\nunderlying \"note\" trees, as the notes-index would be the primary\nway to access them.\n\nNot that I've looked at the code too deeply with an intention of\npossibly including it early.  I was hoping to see fixes to d/f\ncode in merge-recursive from either you or Alex instead ;-)\n"},{"id":"47482","messageId":"Pine.LNX.4.64.0707160049210.14781@racer.site","threadId":"9051","inReplyTo":"7vlkdhck8d.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 2/6] Introduce commit notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-15T23:52:11Z","receivedAt":"2007-07-15T23:52:11Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 15 Jul 2007, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > The notes ref is a branch which contains trees much like the\n> > loose object trees in .git/objects/.  In other words, to get\n> > at the commit notes for a given SHA-1, take the first two\n> > hex characters as directory name, and the remaining 38 hex\n> > characters as base name, and look that up in the notes ref.\n> > ...\n> > However, a remedy is near: in a later commit, a .git/notes-index\n> > will be introduced, a cached mapping from commits to commit notes,\n> > to be written when the tree name of the notes ref changes.  In\n> > case that notes-index cannot be written, the current (possibly\n> > slow) code will come into effect again.\n> \n> I wonder if it is worth using the fan-out tree structure for the\n> underlying \"note\" trees, as the notes-index would be the primary\n> way to access them.\n\nThe fan-out tree is a nice fallback solution when you cannot write the \nnotes-index.\n\n> Not that I've looked at the code too deeply with an intention of \n> possibly including it early.  I was hoping to see fixes to d/f code in \n> merge-recursive from either you or Alex instead ;-)\n\nWell, yeah.  I was kind of trying to cool off from my unpleasant \nunpack_trees() experience.\n\nBut I'll look into the issue again this week.  Promise.\n\nCiao,\nDscho\n"},{"id":"47485","messageId":"7vhco5cixe.fsf@assigned-by-dhcp.cox.net","threadId":"9051","inReplyTo":"7vlkdhck8d.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 2/6] Introduce commit notes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-16T00:05:01Z","receivedAt":"2007-07-16T00:05:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> I wonder if it is worth using the fan-out tree structure for the\n> underlying \"note\" trees, as the notes-index would be the primary\n> way to access them.\n\nActually now I think about it I think this was a stupid\nsuggestion.  Creation of a new note in a reasonably well\npopulated note tree would be made 256-fold more efficient by\nhaving the fan-out, as write-tree does not have to recompute the\nother 255 tree objects thanks to the cache-tree data being\nfresh.\n"},{"id":"47509","messageId":"7vejj96igx.fsf@assigned-by-dhcp.cox.net","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707160022560.14781@racer.site","subject":"Re: [PATCH 2/6] Introduce commit notes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-16T05:11:26Z","receivedAt":"2007-07-16T05:11:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> +core.notesRef::\n> +\tWhen showing commit messages, also show notes which are stored in\n> +\tthe given ref.  This ref is expected to contain paths of the form\n> +\t??/*, where the directory name consists of the first two\n> +\tcharacters of the commit name, and the base name consists of\n> +\tthe remaining 38 characters.\n> ++\n> +If such a path exists in the given ref, the referenced blob is read, and\n> +appended to the commit message, separated by a \"Notes:\" line.  If the\n> +given ref itself does not exist, it is not an error, but means that no\n> +notes should be print.\n> ++\n> +This setting defaults to \"refs/notes/commits\", and can be overridden by\n> +the `GIT_NOTES_REF` environment variable.\n> +\n\nThis design forces \"one blob and only one blob decorates a\ncommit\".  It certainly makes the implementation and semantics\nsimpler -- if I have this note and you have that note on the\nsame commit, comparing notes eventually should result in a merge\nof our notes.  But is it sufficient in real life usage scenarios\n(what's the use case)?  One example that was raised on the list\nis to collect \"Acked-by\", \"Tested-by\", etc., and in that case\nperhaps one set \"refs/notes/acks\" may hold the former while\n\"refs/notes/tests\" the latter.  If we wanted to show both at the\nsame time, is it the only option to put them in the same \"note\"\nblob and not use \"refs/notes/{acks,tests}\"?\n\n> diff --git a/commit.c b/commit.c\n> index 0c350bc..3529b6a 100644\n> --- a/commit.c\n> +++ b/commit.c\n> @@ -6,6 +6,7 @@\n>  #include \"interpolate.h\"\n>  #include \"diff.h\"\n>  #include \"revision.h\"\n> +#include \"notes.h\"\n>  \n>  int save_commit_buffer = 1;\n>  \n> @@ -1254,6 +1255,10 @@ unsigned long pretty_print_commit(enum cmit_fmt fmt,\n>  \t */\n>  \tif (fmt == CMIT_FMT_EMAIL && offset <= beginning_of_body)\n>  \t\tbuf[offset++] = '\\n';\n> +\n> +\tif (fmt != CMIT_FMT_ONELINE)\n> +\t\tget_commit_notes(commit, buf_p, &offset, space_p);\n> +\n>  \tbuf[offset] = '\\0';\n>  \tfree(reencoded);\n>  \treturn offset;\n\nThis makes me wonder if there are cases where \"notes\" need to be\nreencoded to honor log_output_encoding.\n\nSince more and more people live in UTF-8 only world, and this is\na _new_ feature anyway, we could declare that \"notes\" blobs MUST\nbe encoded in UTF-8 upfront, but even if we did so we would need\nreencoding to log_output_encoding, I suspect.\n\n> diff --git a/notes.c b/notes.c\n> new file mode 100644\n> index 0000000..5d1bb1a\n> --- /dev/null\n> +++ b/notes.c\n> @@ -0,0 +1,64 @@\n> +#include \"cache.h\"\n> +#include \"commit.h\"\n> +#include \"notes.h\"\n> +#include \"refs.h\"\n> +\n> +static int initialized;\n> +\n> +void get_commit_notes(const struct commit *commit,\n> +\t\tchar **buf_p, unsigned long *offset_p, unsigned long *space_p)\n> +{\n> +\tchar name[80];\n> +\tconst char *hex;\n> +\tunsigned char sha1[20];\n> +\tchar *msg;\n> +\tunsigned long msgoffset, msglen;\n> +\tenum object_type type;\n> +\n> +\tif (!initialized) {\n> +\t\tconst char *env = getenv(GIT_NOTES_REF);\n> +\t\tif (env) {\n> +\t\t\tif (notes_ref_name)\n> +\t\t\t\tfree(notes_ref_name);\n> +\t\t\tnotes_ref_name = xstrdup(getenv(GIT_NOTES_REF));\n\n\txstrdup(env)?\n\n> +\t\t} else if (!notes_ref_name)\n> +\t\t\tnotes_ref_name = xstrdup(\"refs/notes/commits\");\n\nWe would probably want to give another preprocessor constant for\nthis hardcoded string, next to GIT_NOTES_REF definition in cache.h.\n\n> +\t\tif (notes_ref_name && read_ref(notes_ref_name, sha1)) {\n> +\t\t\tfree(notes_ref_name);\n> +\t\t\tnotes_ref_name = NULL;\n> +\t\t}\n> +\t\tinitialized = 1;\n> +\t}\n> +\tif (!notes_ref_name)\n> +\t\treturn;\n> +\n> +\thex = sha1_to_hex(commit->object.sha1);\n> +\tsnprintf(name, sizeof(name), \"%s:%.*s/%.*s\",\n> +\t\t\tnotes_ref_name, 2, hex, 38, hex + 2);\n\nToo long a notes_ref_name and it won't overrun the buffer but\nthe failure is not detected, and ...\n\n> +\tif (get_sha1(name, sha1))\n> +\t\treturn;\n\n... this would fail silently, leaving the user scratching his head.\n\n> +\n> +\tif (!(msg = read_sha1_file(sha1, &type, &msglen)) || !msglen)\n> +\t\treturn;\n\nWhat's in \"type\" at this point?  Having a tree there is not an\nerror?\n\n> +\t/* we will end the annotation by a newline anyway. */\n> +\tif (msg[msglen - 1] == '\\n')\n> +\t\tmsglen--;\n> +\tALLOC_GROW(*buf_p, *offset_p + 14 + msglen, *space_p);\n> +\t*offset_p += sprintf(*buf_p + *offset_p, \"\\nNotes:\\n\");\n\nFourteen is because...\n\n> +\n> +\tfor (msgoffset = 0; msgoffset < msglen;) {\n> +\t\tint linelen = get_line_length(msg + msgoffset, msglen);\n> +\n> +\t\tALLOC_GROW(*buf_p, *offset_p + linelen + 6, *space_p);\n> +\t\t*offset_p += sprintf(*buf_p + *offset_p,\n> +\t\t\t\t\"    %.*s\", linelen, msg + msgoffset);\n\nSix is because...\n\n> +\t\tmsgoffset += linelen;\n> +\t}\n> +\tALLOC_GROW(*buf_p, *offset_p + 1, *space_p);\n> +\t(*buf_p)[*offset_p] = '\\n';\n> +\t(*offset_p)++;\n> +\tfree(msg);\n> +}\n> +\n> +\n> diff --git a/notes.h b/notes.h\n> new file mode 100644\n> index 0000000..aed80e7\n> --- /dev/null\n> +++ b/notes.h\n> @@ -0,0 +1,9 @@\n> +#ifndef NOTES_H\n> +#define NOTES_H\n> +\n> +void get_commit_notes(const struct commit *commit,\n> +\tchar **buf_p, unsigned long *offset_p, unsigned long *space_p);\n> +\n> +#define GIT_NOTES_REF \"GIT_NOTES_REF\"\n\nJudging from the existing entries in cache.h, it seems that\nGIT_NOTES_REF_ENVIRONMENT would be more appropriate preprocessor\nsymbol for this.  Also let's have this in cache.h next to\nGIT_DIR_ENVIRONMENT and friends, with another definition for\n\"refs/notes/commits\".\n"},{"id":"47508","messageId":"7v8x9h6igv.fsf@assigned-by-dhcp.cox.net","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707160023360.14781@racer.site","subject":"Re: [PATCH 3/6] Add git-notes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-16T05:11:28Z","receivedAt":"2007-07-16T05:11:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> This script allows you to edit and show commit notes easily.\n>\n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>  .gitignore   |    1 +\n>  Makefile     |    2 +-\n>  git-notes.sh |   61 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n>  3 files changed, 63 insertions(+), 1 deletions(-)\n>  create mode 100755 git-notes.sh\n>\n> diff --git a/git-notes.sh b/git-notes.sh\n> new file mode 100755\n> index 0000000..e0ad0b9\n> --- /dev/null\n> +++ b/git-notes.sh\n> @@ -0,0 +1,61 @@\n> +#!/bin/sh\n> +\n> +USAGE=\"(edit | show) [commit]\"\n> +. git-sh-setup\n> +\n> +test -n \"$3\" && usage\n> +\n> +test -z \"$GIT_NOTES_REF\" && GIT_NOTES_REF=\"$(git config core.notesref)\"\n> +test -z \"$GIT_NOTES_REF\" &&\n> +\tdie \"No notes ref set.\"\n\n\ttest -n \"${GIT_NOTES_REF=$(git config core.notesref)}\" || die\n\n> +COMMIT=$(git rev-parse --verify --default HEAD \"$2\")\n\nThis silently annotates the HEAD commit if $2 is misspelled, I\nsuspect.  Also if HEAD does not exist, COMMIT will be empty and\nthis whole command will exit with non-zero status, which you\nwould want to catch here...\n\n> +NAME=$(echo $COMMIT | sed \"s/^../&\\//\")\n\n... or here.\n\n> +case \"$1\" in\n> +edit)\n> +\tMESSAGE=\"$GIT_DIR\"/new-notes\n> +\tGIT_NOTES_REF= git log -1 $COMMIT | sed \"s/^/#/\" > \"$MESSAGE\"\n\n$MESSAGE and its associated temporary file needs to be cleaned\nup upon command exit; perhaps a trap is in order.\n\n> +\tGIT_INDEX_FILE=\"$MESSAGE\".idx\n> +\texport GIT_INDEX_FILE\n> +\n> +\tCURRENT_HEAD=$(git show-ref $GIT_NOTES_REF | cut -f 1 -d ' ')\n> +\tif [ -z \"$CURRENT_HEAD\" ]; then\n> +\t\tPARENT=\n> +\telse\n> +\t\tPARENT=\"-p $OLDTIP\"\n> +\t\tgit read-tree $GIT_NOTES_REF || die \"Could not read index\"\n> +\t\tgit cat-file blob :$NAME >> \"$MESSAGE\" 2> /dev/null\n> +\tfi\n\nI take that OLDTIP is a typo.\n\n\tif CURRENT_HEAD=$(git show-ref -s \"$GIT_NOTES_REF\")\n        then\n\t\tPARENT=\"-p $CURRENT_HEAD\"\n                ...\n\telse\n        \tPARENT=\n\tfi\n\n> +\n> +\t${VISUAL:-${EDITOR:-vi}} \"$MESSAGE\"\n> +\n> +\tgrep -v ^# < \"$MESSAGE\" | git stripspace > \"$MESSAGE\".processed\n\nMakes us wonder if we would want to teach hash-stripping to\ngit-stripspace, doesn't it?\n\n> +\tmv \"$MESSAGE\".processed \"$MESSAGE\"\n> +\tif [ -z \"$(cat \"$MESSAGE\")\" ]; then\n\nMake this 'if test -s \"$MESSAGE\"' and swap then/else clause\naround; no reason to slurp the value into your shell.\n\n> +\t\ttest -z \"$CURRENT_HEAD\" &&\n> +\t\t\tdie \"Will not initialise with empty tree\"\n> +\t\tgit update-index --force-remove $NAME ||\n> +\t\t\tdie \"Could not update index\"\n> +\telse\n> +\t\tBLOB=$(git hash-object -w \"$MESSAGE\") ||\n> +\t\t\tdie \"Could not write into object database\"\n> +\t\tgit update-index --add --cacheinfo 0644 $BLOB $NAME ||\n> +\t\t\tdie \"Could not write index\"\n> +\tfi\n> +\n> +\tTREE=$(git write-tree) || die \"Could not write tree\"\n> +\tNEW_HEAD=$(: | git commit-tree $TREE $PARENT) ||\n> +\t\tdie \"Could not annotate\"\n\nHmph.  How about \"echo Annotate $COMMIT | git commit-tree...\"?\n\n> +\tcase \"$CURRENT_HEAD\" in\n> +\t'') git update-ref $GIT_NOTES_REF $NEW_HEAD ;;\n> +\t*) git update-ref $GIT_NOTES_REF $NEW_HEAD $CURRENT_HEAD;;\n> +\tesac\n> +;;\n\nThere are some places that have \"$GIT_NOTES_REF\" in dq and some\nplaces you don't.  I think GIT_NOTES_REF begins with refs/ and\nconsists only of valid refname characters, so unless the user\nwants to shoot himself in the foot it should be Ok, but we\nprobably would want to quote it.\n\nAlso, as unquoted $CURRENT_HEAD will not even count as a\nparameter to update-ref, you do not have to do that case/esac,\nbut simply do:\n\n\tgit update-ref \"$GIT_NOTES_REF\" $NEW_HEAD $CURRENT_HEAD\n\nWould we have reflog for this ref?  What would we want to see as\nthe message if we do?\n"},{"id":"47510","messageId":"7v3azp6igt.fsf@assigned-by-dhcp.cox.net","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707160024060.14781@racer.site","subject":"Re: [PATCH 4/6] Add a test script for \"git notes\"","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-16T05:11:30Z","receivedAt":"2007-07-16T05:11:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> diff --git a/t/t3301-notes.sh b/t/t3301-notes.sh\n> new file mode 100755\n> index 0000000..eb50191\n> --- /dev/null\n> +++ b/t/t3301-notes.sh\n> @@ -0,0 +1,63 @@\n> +#!/bin/sh\n> +#\n> +# Copyright (c) 2007 Johannes E. Schindelin\n> +#\n> +\n> +test_description='Test commit notes'\n> +\n> +. ./test-lib.sh\n> +\n> +test_expect_success setup '\n> +\t: > a1 &&\n> +\tgit add a1 &&\n> +\ttest_tick &&\n> +\tgit commit -m 1st &&\n> +\t: > a2 &&\n> +\tgit add a2 &&\n> +\ttest_tick &&\n> +\tgit commit -m 2nd\n> +'\n\nDoes not test the failure mode of not having a HEAD yet.\n\n> +cat > fake_editor.sh << EOF\n> +echo \"\\$MSG\" > \"\\$1\"\n> +echo \"\\$MSG\" >& 2\n> +EOF\n\nYou can avoid all these backslashes by saying:\n\n\tcat >fake_editor.sh <<\\EOF\n        echo \"$MSG\" >\"$1\"\n        echo \"$MSG\" >&2\n\tEOF\n\n> +chmod a+x fake_editor.sh\n> +VISUAL=\"$(pwd)\"/fake_editor.sh\n> +export VISUAL\n\nNot that it hurts anybody, but do you really need that $(pwd),\ninstead of \"./fake_editor.sh\"?\n\n> +\n> +test_expect_success 'need notes ref' '\n> +\t! MSG=1 git notes edit &&\n> +\t! MSG=2 git notes show\n> +'\n> +\n> +test_expect_success 'create notes' '\n> +\tgit config core.notesRef refs/notes/commits &&\n> +\tMSG=b1 git notes edit &&\n> +cat .git/new-notes &&\n> +test b1 = \"$(cat .git/new-notes)\" &&\n> +\ttest 1 = $(git ls-tree refs/notes/commits | wc -l) &&\n> +\ttest b1 = $(git notes show) &&\n> +\tgit show HEAD^ &&\n> +\t! git notes show HEAD^\n> +'\n\nIs there particular reason for that (lack of) indentation for\nthe two lines among them?\n\nI think it is a bug to leave \".git/new-notes\" and friends\nbehind.\n\n> +\n> +cat > expect << EOF\n> +commit 268048bfb8a1fb38e703baceb8ab235421bf80c5\n> +Author: A U Thor <author@example.com>\n> +Date:   Thu Apr 7 15:14:13 2005 -0700\n> +\n> +    2nd\n> +\n> +Notes:\n> +    b1\n> +EOF\n> +\n> +test_expect_success 'show notes' '\n> +\t! (git cat-file commit HEAD | grep b1) &&\n> +\tgit log -1 > output &&\n> +\tgit diff expect output\n> +'\n> +\n> +test_done\n\nHmph.  This makes the reader wonder why this is not optional,\nperhaps linked to --decorate option somehow.\n"},{"id":"47515","messageId":"20070716060117.GF32566@spearce.org","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707160025480.14781@racer.site","subject":"Re: [WIP PATCH 6/6] notes: add notes-index for a substantial speedup.","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-07-16T06:01:17Z","receivedAt":"2007-07-16T06:01:17Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> \n> Actually, this commit adds two methods for a notes index:\n> \n> - a sorted list with a fan out to help binary search, and\n> - a modified hash table.\n> \n> It also adds a test which is used to determine the best algorithm.\n\nI know this is a nice backwards compatible way to organize notes,\nand to make them reasonably efficiently found, but I'd almost\nrather just see them crammed into the packfile alongside of the\ncommit it annotates, so that the packfile reader can quickly find\nthe annotation at the same time it finds the commit.\n\naka packv4.\n\nOk, enough dreaming for today.\n\n-- \nShawn.\n"},{"id":"47520","messageId":"200707160857.48725.andyparkins@gmail.com","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707152326080.14781@racer.site","subject":"Re: [PATCH 0/6] Introduce commit notes","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-07-16T07:57:46Z","receivedAt":"2007-07-16T07:57:46Z","isPatch":true,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Monday 2007 July 16, Johannes Schindelin wrote:\n\n> The biggest obstacle was a thinko about the scalability.  Tree objects\n> take free form name entries, and therefore a binary search by name is not\n> possible.\n\nI might be misunderstanding, but in the case of the notes tree objects isn't \nit true that the name entries aren't free form, but are guaranteed to be of a \nfixed length form:\n\n  XX/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX\n\nIn which case you can binary search?\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"47521","messageId":"7vbqec4vk2.fsf@assigned-by-dhcp.cox.net","threadId":"9051","inReplyTo":"200707160857.48725.andyparkins@gmail.com","subject":"Re: [PATCH 0/6] Introduce commit notes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-16T08:11:41Z","receivedAt":"2007-07-16T08:11:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Andy Parkins <andyparkins@gmail.com> writes:\n\n> On Monday 2007 July 16, Johannes Schindelin wrote:\n>\n>> The biggest obstacle was a thinko about the scalability.  Tree objects\n>> take free form name entries, and therefore a binary search by name is not\n>> possible.\n>\n> I might be misunderstanding, but in the case of the notes tree objects isn't \n> it true that the name entries aren't free form, but are guaranteed to be of a \n> fixed length form:\n>\n>   XX/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX\n>\n> In which case you can binary search?\n\nHmph, you are right.  In this sequence:\n\n\thex = sha1_to_hex(commit->object.sha1);\n\tsnprintf(name, sizeof(name), \"%s:%.*s/%.*s\",\n\t\t\tnotes_ref_name, 2, hex, 38, hex + 2);\n\tif (get_sha1(name, sha1))\n\t\treturn;\n\nInstead, we could read the tree object by hand in the commit\nthat is referenced by notes_ref_name, which has uniform two\nletter names for subtrees which can be binary searched, open the\ntree for that entry, again by hand, and do another binary search\nbecause that tree has uniform 38-letter names.  That certainly\ncould be done.\n\nSounds like a \"fun\" project for some definition of the word.\n"},{"id":"47555","messageId":"Pine.LNX.4.64.0707161724110.14781@racer.site","threadId":"9051","inReplyTo":"7vbqec4vk2.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 0/6] Introduce commit notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-16T16:26:31Z","receivedAt":"2007-07-16T16:26:31Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 16 Jul 2007, Junio C Hamano wrote:\n\n> Andy Parkins <andyparkins@gmail.com> writes:\n> \n> > On Monday 2007 July 16, Johannes Schindelin wrote:\n> >\n> >> The biggest obstacle was a thinko about the scalability.  Tree \n> >> objects take free form name entries, and therefore a binary search by \n> >> name is not possible.\n> >\n> > I might be misunderstanding, but in the case of the notes tree objects \n> > isn't it true that the name entries aren't free form, but are \n> > guaranteed to be of a fixed length form:\n> >\n> >   XX/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX\n> >\n> > In which case you can binary search?\n> \n> Hmph, you are right.  In this sequence:\n> \n> \thex = sha1_to_hex(commit->object.sha1);\n> \tsnprintf(name, sizeof(name), \"%s:%.*s/%.*s\",\n> \t\t\tnotes_ref_name, 2, hex, 38, hex + 2);\n> \tif (get_sha1(name, sha1))\n> \t\treturn;\n> \n> Instead, we could read the tree object by hand in the commit that is \n> referenced by notes_ref_name, which has uniform two letter names for \n> subtrees which can be binary searched, open the tree for that entry, \n> again by hand, and do another binary search because that tree has \n> uniform 38-letter names.  That certainly could be done.\n> \n> Sounds like a \"fun\" project for some definition of the word.\n\nI disagree.  One disadvantage to using tree objects is that it is much \neasier to have pilot errors.  You could even make a new working tree \nchecking out refs/notes/commits and change/add/remove files.\n\nCiao,\nDscho\n"},{"id":"47557","messageId":"Pine.LNX.4.64.0707161726430.14781@racer.site","threadId":"9051","inReplyTo":"20070716060117.GF32566@spearce.org","subject":"Re: [WIP PATCH 6/6] notes: add notes-index for a substantial speedup.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-16T16:29:58Z","receivedAt":"2007-07-16T16:29:58Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 16 Jul 2007, Shawn O. Pearce wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > \n> > Actually, this commit adds two methods for a notes index:\n> > \n> > - a sorted list with a fan out to help binary search, and\n> > - a modified hash table.\n> > \n> > It also adds a test which is used to determine the best algorithm.\n> \n> I know this is a nice backwards compatible way to organize notes,\n> and to make them reasonably efficiently found, but I'd almost\n> rather just see them crammed into the packfile alongside of the\n> commit it annotates, so that the packfile reader can quickly find\n> the annotation at the same time it finds the commit.\n> \n> aka packv4.\n> \n> Ok, enough dreaming for today.\n\nYes, I also dream of having the time to play with packv4.  If you read my \ncomments in the commit-annotation thread, you'll see that I stated that \npackv4 would solve the problem, too.\n\nThe reason I did this series was not to push commit notes, but to make \ngood for stalling Johan's efforts.  Including a proof that the commit \nnotes as I introduced them can be relatively cheap, too.\n\nCiao,\nDscho\n"},{"id":"47574","messageId":"7vzm1w2pwk.fsf@assigned-by-dhcp.cox.net","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707161724110.14781@racer.site","subject":"Re: [PATCH 0/6] Introduce commit notes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-16T17:56:43Z","receivedAt":"2007-07-16T17:56:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> Hmph, you are right.  In this sequence:\n>> \n>> \thex = sha1_to_hex(commit->object.sha1);\n>> \tsnprintf(name, sizeof(name), \"%s:%.*s/%.*s\",\n>> \t\t\tnotes_ref_name, 2, hex, 38, hex + 2);\n>> \tif (get_sha1(name, sha1))\n>> \t\treturn;\n>> \n>> Instead, we could read the tree object by hand in the commit that is \n>> referenced by notes_ref_name, which has uniform two letter names for \n>> subtrees which can be binary searched, open the tree for that entry, \n>> again by hand, and do another binary search because that tree has \n>> uniform 38-letter names.  That certainly could be done.\n>> \n>> Sounds like a \"fun\" project for some definition of the word.\n>\n> I disagree.  One disadvantage to using tree objects is that it is much \n> easier to have pilot errors.  You could even make a new working tree \n> checking out refs/notes/commits and change/add/remove files.\n\nI suspect you read me wrong.  I was saying that it is possible\nto use a specialized tree object parser in place of get_sha1()\nonly in the above code to read the tree objects that represents\na 'note'.  You obviously would want to do a sanity check such\nas:\n\n - The size of the tree object your customized tree parser is\n   fed is multiple of expected entry size (mode word + 20 SHA1 +\n   2 + NUL for fan-out, replace 2 with 38 for lower level);\n\n - mode word for the entry is sane (an entry in the fan-out tree\n   would point at a tree object, an entry in lower level would\n   point at a blob);\n\n - The name part (2 or 38) are lowercase hexadecimal strings;\n"},{"id":"47794","messageId":"Pine.LNX.4.64.0707190232570.14781@racer.site","threadId":"9051","inReplyTo":"7vzm1w2pwk.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 0/6] Introduce commit notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T01:34:28Z","receivedAt":"2007-07-19T01:34:28Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 16 Jul 2007, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> >> Hmph, you are right.  In this sequence:\n> >> \n> >> \thex = sha1_to_hex(commit->object.sha1);\n> >> \tsnprintf(name, sizeof(name), \"%s:%.*s/%.*s\",\n> >> \t\t\tnotes_ref_name, 2, hex, 38, hex + 2);\n> >> \tif (get_sha1(name, sha1))\n> >> \t\treturn;\n> >> \n> >> Instead, we could read the tree object by hand in the commit that is \n> >> referenced by notes_ref_name, which has uniform two letter names for \n> >> subtrees which can be binary searched, open the tree for that entry, \n> >> again by hand, and do another binary search because that tree has \n> >> uniform 38-letter names.  That certainly could be done.\n> >> \n> >> Sounds like a \"fun\" project for some definition of the word.\n> >\n> > I disagree.  One disadvantage to using tree objects is that it is much \n> > easier to have pilot errors.  You could even make a new working tree \n> > checking out refs/notes/commits and change/add/remove files.\n> \n> I suspect you read me wrong.  I was saying that it is possible to use a \n> specialized tree object parser in place of get_sha1() only in the above \n> code to read the tree objects that represents a 'note'.  You obviously \n> would want to do a sanity check such as:\n> \n>  - The size of the tree object your customized tree parser is\n>    fed is multiple of expected entry size (mode word + 20 SHA1 +\n>    2 + NUL for fan-out, replace 2 with 38 for lower level);\n> \n>  - mode word for the entry is sane (an entry in the fan-out tree\n>    would point at a tree object, an entry in lower level would\n>    point at a blob);\n> \n>  - The name part (2 or 38) are lowercase hexadecimal strings;\n\nIn which case it is not _that_ attractive any more, since you\n\n- have to have a fallback anyway, and\n\n- have a relatively complex thing.\n\nInstead, I want to go with the hash map approach, if only to have a O(1) \nbehaviour instead of O(log N).\n\nCiao,\nDscho\n"},{"id":"47797","messageId":"Pine.LNX.4.64.0707190258550.14781@racer.site","threadId":"9051","inReplyTo":"7vejj96igx.fsf@assigned-by-dhcp.cox.net","subject":"[REVISED PATCH 2/6] Introduce commit notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T02:30:43Z","receivedAt":"2007-07-19T02:30:43Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nCommit notes are blobs which are shown together with the commit\nmessage.  These blobs are taken from the notes ref, which you can\nconfigure by the config variable core.notesRef, which in turn can\nbe overridden by the environment variable GIT_NOTES_REF.\n\nThe notes ref is a branch which contains trees much like the\nloose object trees in .git/objects/.  In other words, to get\nat the commit notes for a given SHA-1, take the first two\nhex characters as directory name, and the remaining 38 hex\ncharacters as base name, and look that up in the notes ref.\n\nThe rationale for putting this information into a ref is this: we\nwant to be able to fetch and possibly union-merge the notes,\nmaybe even look at the date when a note was introduced, and we\nwant to store them efficiently together with the other objects.\n\nThere is one severe shortcoming, though.  Since tree objects can\ncontain file names of a variable length, it is not possible to\ndo a binary search for the correct base name in the tree object's\ncontents.  Therefore this approach does not scale well, because\nthe average lookup time will be proportional to the number of\ncommit objects, and therefore the slowdown will be quadratic in\nthat number.\n\nHowever, a remedy is near: in a later commit, a .git/notes-index\nwill be introduced, a cached mapping from commits to commit notes,\nto be written when the tree name of the notes ref changes.  In\ncase that notes-index cannot be written, the current (possibly\nslow) code will come into effect again.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n\n\tOn Sun, 15 Jul 2007, Junio C Hamano wrote:\n\n\t> This design forces \"one blob and only one blob decorates a\n\t> commit\".  It certainly makes the implementation and semantics\n\t> simpler -- if I have this note and you have that note on the\n\t> same commit, comparing notes eventually should result in a merge\n\t> of our notes.  But is it sufficient in real life usage scenarios\n\t> (what's the use case)?  One example that was raised on the list\n\t> is to collect \"Acked-by\", \"Tested-by\", etc., and in that case\n\t> perhaps one set \"refs/notes/acks\" may hold the former while\n\t> \"refs/notes/tests\" the latter.  If we wanted to show both at the\n\t> same time, is it the only option to put them in the same \"note\"\n\t> blob and not use \"refs/notes/{acks,tests}\"?\n\n\tWould that not make things even slower?  I am hesitant.\n\n\tAll other concerns should be addressed, here and in the two \n\tupcoming revised patches.\n\n Documentation/config.txt |   15 +++++++++\n Makefile                 |    3 +-\n cache.h                  |    3 ++\n commit.c                 |    5 +++\n config.c                 |    5 +++\n environment.c            |    1 +\n notes.c                  |   77 ++++++++++++++++++++++++++++++++++++++++++++++\n notes.h                  |    8 +++++\n 8 files changed, 116 insertions(+), 1 deletions(-)\n create mode 100644 notes.c\n create mode 100644 notes.h\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex d0e9a17..5fe833d 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -285,6 +285,21 @@ core.pager::\n \tThe command that git will use to paginate output.  Can be overridden\n \twith the `GIT_PAGER` environment variable.\n \n+core.notesRef::\n+\tWhen showing commit messages, also show notes which are stored in\n+\tthe given ref.  This ref is expected to contain paths of the form\n+\t??/*, where the directory name consists of the first two\n+\tcharacters of the commit name, and the base name consists of\n+\tthe remaining 38 characters.\n++\n+If such a path exists in the given ref, the referenced blob is read, and\n+appended to the commit message, separated by a \"Notes:\" line.  If the\n+given ref itself does not exist, it is not an error, but means that no\n+notes should be print.\n++\n+This setting defaults to \"refs/notes/commits\", and can be overridden by\n+the `GIT_NOTES_REF` environment variable.\n+\n alias.*::\n \tCommand aliases for the gitlink:git[1] command wrapper - e.g.\n \tafter defining \"alias.last = cat-file commit HEAD\", the invocation\ndiff --git a/Makefile b/Makefile\nindex d7541b4..119d949 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -322,7 +322,8 @@ LIB_OBJS = \\\n \twrite_or_die.o trace.o list-objects.o grep.o match-trees.o \\\n \talloc.o merge-file.o path-list.o help.o unpack-trees.o $(DIFF_OBJS) \\\n \tcolor.o wt-status.o archive-zip.o archive-tar.o shallow.o utf8.o \\\n-\tconvert.o attr.o decorate.o progress.o mailmap.o symlinks.o remote.o\n+\tconvert.o attr.o decorate.o progress.o mailmap.o symlinks.o remote.o \\\n+\tnotes.o\n \n BUILTIN_OBJS = \\\n \tbuiltin-add.o \\\ndiff --git a/cache.h b/cache.h\nindex 328c1ad..df45336 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -204,6 +204,8 @@ enum object_type {\n #define GITATTRIBUTES_FILE \".gitattributes\"\n #define INFOATTRIBUTES_FILE \"info/attributes\"\n #define ATTRIBUTE_MACRO_PREFIX \"[attr]\"\n+#define GIT_NOTES_REF_ENVIRONMENT \"GIT_NOTES_REF\"\n+#define GIT_NOTES_DEFAULT_REF \"refs/notes/commits\"\n \n extern int is_bare_repository_cfg;\n extern int is_bare_repository(void);\n@@ -309,6 +311,7 @@ extern size_t packed_git_window_size;\n extern size_t packed_git_limit;\n extern size_t delta_base_cache_limit;\n extern int auto_crlf;\n+extern char *notes_ref_name;\n \n #define GIT_REPO_VERSION 0\n extern int repository_format_version;\ndiff --git a/commit.c b/commit.c\nindex 0c350bc..8911a18 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -6,6 +6,7 @@\n #include \"interpolate.h\"\n #include \"diff.h\"\n #include \"revision.h\"\n+#include \"notes.h\"\n \n int save_commit_buffer = 1;\n \n@@ -1254,6 +1255,10 @@ unsigned long pretty_print_commit(enum cmit_fmt fmt,\n \t */\n \tif (fmt == CMIT_FMT_EMAIL && offset <= beginning_of_body)\n \t\tbuf[offset++] = '\\n';\n+\n+\tif (fmt != CMIT_FMT_ONELINE)\n+\t\tget_commit_notes(commit, buf_p, &offset, space_p, encoding);\n+\n \tbuf[offset] = '\\0';\n \tfree(reencoded);\n \treturn offset;\ndiff --git a/config.c b/config.c\nindex f89a611..05d2ad6 100644\n--- a/config.c\n+++ b/config.c\n@@ -395,6 +395,11 @@ int git_default_config(const char *var, const char *value)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.notesref\")) {\n+\t\tnotes_ref_name = xstrdup(value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"user.name\")) {\n \t\tstrlcpy(git_default_name, value, sizeof(git_default_name));\n \t\treturn 0;\ndiff --git a/environment.c b/environment.c\nindex f83fb9e..2e677d3 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -34,6 +34,7 @@ char *pager_program;\n int pager_in_use;\n int pager_use_color = 1;\n int auto_crlf = 0;\t/* 1: both ways, -1: only when adding git objects */\n+char *notes_ref_name;\n \n static const char *git_dir;\n static char *git_object_dir, *git_index_file, *git_refs_dir, *git_graft_file;\ndiff --git a/notes.c b/notes.c\nnew file mode 100644\nindex 0000000..6207f95\n--- /dev/null\n+++ b/notes.c\n@@ -0,0 +1,77 @@\n+#include \"cache.h\"\n+#include \"commit.h\"\n+#include \"notes.h\"\n+#include \"refs.h\"\n+#include \"utf8.h\"\n+\n+static int initialized;\n+\n+void get_commit_notes(const struct commit *commit,\n+\t\tchar **buf_p, unsigned long *offset_p, unsigned long *space_p,\n+\t\tconst char *output_encoding)\n+{\n+        static const char *utf8 = \"utf-8\";\n+\tchar name[80];\n+\tconst char *hex;\n+\tunsigned char sha1[20];\n+\tchar *msg;\n+\tunsigned long msgoffset, msglen;\n+\tenum object_type type;\n+\n+\tif (!initialized) {\n+\t\tconst char *env = getenv(GIT_NOTES_REF_ENVIRONMENT);\n+\t\tif (env)\n+\t\t\tnotes_ref_name = getenv(GIT_NOTES_REF_ENVIRONMENT);\n+\t\telse if (!notes_ref_name)\n+\t\t\tnotes_ref_name = GIT_NOTES_DEFAULT_REF;\n+\t\tif (notes_ref_name && read_ref(notes_ref_name, sha1))\n+\t\t\tnotes_ref_name = NULL;\n+\t\tinitialized = 1;\n+\t}\n+\tif (!notes_ref_name)\n+\t\treturn;\n+\n+\thex = sha1_to_hex(commit->object.sha1);\n+\tif (snprintf(name, sizeof(name), \"%s:%.*s/%.*s\",\n+\t\t\tnotes_ref_name, 2, hex, 38, hex + 2)\n+\t\t\t>= sizeof(name) - 1) {\n+\t\terror(\"Notes ref name too long: %.*s\", 60, notes_ref_name);\n+\t\treturn;\n+\t}\n+\tif (get_sha1(name, sha1))\n+\t\treturn;\n+\n+\tif (!(msg = read_sha1_file(sha1, &type, &msglen)) || !msglen ||\n+\t\t\ttype != OBJ_BLOB)\n+\t\treturn;\n+        if (output_encoding && *output_encoding &&\n+\t\t\tstrcmp(utf8, output_encoding)) {\n+                char *reencoded = reencode_string(msg, output_encoding, utf8);\n+\t\tif (reencoded) {\n+\t\t\tfree(msg);\n+\t\t\tmsg = reencoded;\n+\t\t\tmsglen = strlen(msg);\n+\t\t}\n+\t}\n+\t/* we will end the annotation by a newline anyway. */\n+\tif (msg[msglen - 1] == '\\n')\n+\t\tmsglen--;\n+\n+\tALLOC_GROW(*buf_p, *offset_p + 8 + msglen, *space_p);\n+\t*offset_p += sprintf(*buf_p + *offset_p, \"\\nNotes:\\n\");\n+\n+\tfor (msgoffset = 0; msgoffset < msglen;) {\n+\t\tint linelen = get_line_length(msg + msgoffset, msglen);\n+\n+\t\tALLOC_GROW(*buf_p, *offset_p + linelen + 5, *space_p);\n+\t\t*offset_p += sprintf(*buf_p + *offset_p,\n+\t\t\t\t\"    %.*s\", linelen, msg + msgoffset);\n+\t\tmsgoffset += linelen;\n+\t}\n+\tALLOC_GROW(*buf_p, *offset_p + 1, *space_p);\n+\t(*buf_p)[*offset_p] = '\\n';\n+\t(*offset_p)++;\n+\tfree(msg);\n+}\n+\n+\ndiff --git a/notes.h b/notes.h\nnew file mode 100644\nindex 0000000..fe8f209\n--- /dev/null\n+++ b/notes.h\n@@ -0,0 +1,8 @@\n+#ifndef NOTES_H\n+#define NOTES_H\n+\n+void get_commit_notes(const struct commit *commit,\n+\tchar **buf_p, unsigned long *offset_p, unsigned long *space_p,\n+\tconst char *output_encoding);\n+\n+#endif\n-- \n1.5.3.rc1.16.g9d6f-dirty\n"},{"id":"47799","messageId":"Pine.LNX.4.64.0707190331050.14781@racer.site","threadId":"9051","inReplyTo":"7v8x9h6igv.fsf@assigned-by-dhcp.cox.net","subject":"[REVISED PATCH 3/6] Add git-notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T02:31:23Z","receivedAt":"2007-07-19T02:31:23Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nThis script allows you to edit and show commit notes easily.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n .gitignore   |    1 +\n Makefile     |    2 +-\n git-notes.sh |   62 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 64 insertions(+), 1 deletions(-)\n create mode 100755 git-notes.sh\n\ndiff --git a/.gitignore b/.gitignore\nindex 20ee642..125613f 100644\n--- a/.gitignore\n+++ b/.gitignore\n@@ -83,6 +83,7 @@ git-mktag\n git-mktree\n git-name-rev\n git-mv\n+git-notes\n git-pack-redundant\n git-pack-objects\n git-pack-refs\ndiff --git a/Makefile b/Makefile\nindex 119d949..10a9342 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -213,7 +213,7 @@ SCRIPT_SH = \\\n \tgit-merge-resolve.sh git-merge-ours.sh \\\n \tgit-lost-found.sh git-quiltimport.sh git-submodule.sh \\\n \tgit-filter-branch.sh \\\n-\tgit-stash.sh\n+\tgit-stash.sh git-notes.sh\n \n SCRIPT_PERL = \\\n \tgit-add--interactive.perl \\\ndiff --git a/git-notes.sh b/git-notes.sh\nnew file mode 100755\nindex 0000000..031e911\n--- /dev/null\n+++ b/git-notes.sh\n@@ -0,0 +1,62 @@\n+#!/bin/sh\n+\n+USAGE=\"(edit | show) [commit]\"\n+. git-sh-setup\n+\n+test -n \"$3\" && usage\n+\n+test -z \"$GIT_NOTES_REF\" && GIT_NOTES_REF=\"$(git config core.notesref)\"\n+test -z \"$GIT_NOTES_REF\" && GIT_NOTES_REF=\"refs/notes/commits\"\n+\n+COMMIT=$(git rev-parse --verify --default HEAD \"$2\") || die \"Invalid ref: $2\"\n+NAME=$(echo $COMMIT | sed \"s/^../&\\//\")\n+\n+MESSAGE=\"$GIT_DIR\"/new-notes\n+trap '\n+\ttest -f \"$MESSAGE\" && rm \"$MESSAGE\"\n+' 0\n+\n+case \"$1\" in\n+edit)\n+\tGIT_NOTES_REF= git log -1 $COMMIT | sed \"s/^/#/\" > \"$MESSAGE\"\n+\n+\tGIT_INDEX_FILE=\"$MESSAGE\".idx\n+\texport GIT_INDEX_FILE\n+\n+\tCURRENT_HEAD=$(git show-ref \"$GIT_NOTES_REF\" | cut -f 1 -d ' ')\n+\tif [ -z \"$CURRENT_HEAD\" ]; then\n+\t\tPARENT=\n+\telse\n+\t\tPARENT=\"-p $CURRENT_HEAD\"\n+\t\tgit read-tree \"$GIT_NOTES_REF\" || die \"Could not read index\"\n+\t\tgit cat-file blob :$NAME >> \"$MESSAGE\" 2> /dev/null\n+\tfi\n+\n+\t${VISUAL:-${EDITOR:-vi}} \"$MESSAGE\"\n+\n+\tgrep -v ^# < \"$MESSAGE\" | git stripspace > \"$MESSAGE\".processed\n+\tmv \"$MESSAGE\".processed \"$MESSAGE\"\n+\tif [ -s \"$MESSAGE\" ]; then\n+\t\tBLOB=$(git hash-object -w \"$MESSAGE\") ||\n+\t\t\tdie \"Could not write into object database\"\n+\t\tgit update-index --add --cacheinfo 0644 $BLOB $NAME ||\n+\t\t\tdie \"Could not write index\"\n+\telse\n+\t\ttest -z \"$CURRENT_HEAD\" &&\n+\t\t\tdie \"Will not initialise with empty tree\"\n+\t\tgit update-index --force-remove $NAME ||\n+\t\t\tdie \"Could not update index\"\n+\tfi\n+\n+\tTREE=$(git write-tree) || die \"Could not write tree\"\n+\tNEW_HEAD=$(echo Annotate $COMMIT | git commit-tree $TREE $PARENT) ||\n+\t\tdie \"Could not annotate\"\n+\tgit update-ref -m \"Annotate $COMMIT\" \\\n+\t\t\"$GIT_NOTES_REF\" $NEW_HEAD $CURRENT_HEAD\n+;;\n+show)\n+\tgit show \"$GIT_NOTES_REF\":$NAME\n+;;\n+*)\n+\tusage\n+esac\n-- \n1.5.3.rc1.16.g9d6f-dirty\n"},{"id":"47798","messageId":"Pine.LNX.4.64.0707190331430.14781@racer.site","threadId":"9051","inReplyTo":"7v3azp6igt.fsf@assigned-by-dhcp.cox.net","subject":"[REVISED PATCH 4/6] Add a test script for \"git notes\"","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T02:32:05Z","receivedAt":"2007-07-19T02:32:05Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nIncidentally, a test for \"git notes\" implies a test for the\nwhole commit notes machinery.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/t3301-notes.sh |   65 ++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 files changed, 65 insertions(+), 0 deletions(-)\n create mode 100755 t/t3301-notes.sh\n\ndiff --git a/t/t3301-notes.sh b/t/t3301-notes.sh\nnew file mode 100755\nindex 0000000..ba42c45\n--- /dev/null\n+++ b/t/t3301-notes.sh\n@@ -0,0 +1,65 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2007 Johannes E. Schindelin\n+#\n+\n+test_description='Test commit notes'\n+\n+. ./test-lib.sh\n+\n+cat > fake_editor.sh << \\EOF\n+echo \"$MSG\" > \"$1\"\n+echo \"$MSG\" >& 2\n+EOF\n+chmod a+x fake_editor.sh\n+VISUAL=./fake_editor.sh\n+export VISUAL\n+\n+test_expect_success 'cannot annotate non-existing HEAD' '\n+\t! MSG=3 git notes edit\n+'\n+\n+test_expect_success setup '\n+\t: > a1 &&\n+\tgit add a1 &&\n+\ttest_tick &&\n+\tgit commit -m 1st &&\n+\t: > a2 &&\n+\tgit add a2 &&\n+\ttest_tick &&\n+\tgit commit -m 2nd\n+'\n+\n+test_expect_success 'need valid notes ref' '\n+\t! MSG=1 GIT_NOTES_REF='/' git notes edit &&\n+\t! MSG=2 GIT_NOTES_REF='/' git notes show\n+'\n+\n+test_expect_success 'create notes' '\n+\tgit config core.notesRef refs/notes/commits &&\n+\tMSG=b1 git notes edit &&\n+\ttest ! -f .git/new-notes &&\n+\ttest 1 = $(git ls-tree refs/notes/commits | wc -l) &&\n+\ttest b1 = $(git notes show) &&\n+\tgit show HEAD^ &&\n+\t! git notes show HEAD^\n+'\n+\n+cat > expect << EOF\n+commit 268048bfb8a1fb38e703baceb8ab235421bf80c5\n+Author: A U Thor <author@example.com>\n+Date:   Thu Apr 7 15:14:13 2005 -0700\n+\n+    2nd\n+\n+Notes:\n+    b1\n+EOF\n+\n+test_expect_success 'show notes' '\n+\t! (git cat-file commit HEAD | grep b1) &&\n+\tgit log -1 > output &&\n+\tgit diff expect output\n+'\n+\n+test_done\n-- \n1.5.3.rc1.16.g9d6f-dirty\n"},{"id":"47800","messageId":"Pine.LNX.4.64.0707190353570.14781@racer.site","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707190331050.14781@racer.site","subject":"Re: [REVISED PATCH 3/6] Add git-notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T02:54:51Z","receivedAt":"2007-07-19T02:54:51Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 19 Jul 2007, Johannes Schindelin wrote:\n\n> +MESSAGE=\"$GIT_DIR\"/new-notes\n> +trap '\n> +\ttest -f \"$MESSAGE\" && rm \"$MESSAGE\"\n> +' 0\n\nOh, well.  Probably this should use mktemp and should handle \nGIT_INDEX_FILE, too.\n\nWill do that tomorrow.\n\nCiao,\nDscho\n"},{"id":"47801","messageId":"alpine.LFD.0.999.0707181949490.27353@woody.linux-foundation.org","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707190258550.14781@racer.site","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-19T03:28:27Z","receivedAt":"2007-07-19T03:28:27Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 19 Jul 2007, Johannes Schindelin wrote:\n>\n> There is one severe shortcoming, though.  Since tree objects can\n> contain file names of a variable length, it is not possible to\n> do a binary search for the correct base name in the tree object's\n> contents.\n\nWell, I've been thinking about this, and that's not really entirely \ncorrect.\n\nIt *is* possible to do a binary search, it's just a bit complicated, \nbecause you have to take the \"halfway\" thing, and find the beginning of \nan entry.\n\nBut the good news is that the tree entries have a very fixed format, and \none that is actually amenable to finding where they start. It gets a bit \ncomplicated, but:\n\n - SHA1's contain random bytes, so we cannot really depend on their \n   content. Fair enough. But on the other hand, they are fixed length, \n   which means..\n\n - Each SHA1 is always preceded by a zero byte (it is what separates the \n   filename from the SHA1), and while filenames too can have arbitrary \n   content (and arbitrary length), we know that the *filename* doesn't \n   have a zero byte in it.\n\n - so finding the beginning of a tree entry should be as easy as finding \n   two zero bytes that are have at least 20 bytes in between them, and \n   then you *know* that the second zero byte is the one that starts a new \n   SHA1 (it cannot be _inside_ a SHA1: if it was, there would be less \n   than twenty bytes to the previous '\\0', and it cannot be inside the\n   filename either).\n\n - And you know that 20 bytes after that '\\0' is the next tree entry!\n\nNow, what does this mean? It means that if we actually know the filename \nwe're looking for, and we're looking at a large range, we really *could* \nstart out with binary searching. We would do something like this:\n\n\tunsigned char *start;\n\tunsigned long size;\n\n\twhile (size > 200) {\n\t\t/*\n\t\t * Look halfway, and then back up a bit, because we \n\t\t * expect it to take us about 20 characters to find\n\t\t * the zero we look for, and an additional 20\n\t\t * characters is the subsequent SHA1.\n\t\t */\n\t\tunsigned long guess = size / 2 - 40;\n\n\t\t/*\n\t\t * This is the offset past which a zero means that \n\t\t * we're good. If we don't find a zero in the first\n\t\t * twenty bytes, that means that the first zero we\n\t\t * find must be the beginning of a SHA1!\n\t\t */\n\t\tunsigned long goal_zero = guess + 20;\n\n\t\tfor (;;) {\n\t\t\tunsigned char c;\n\n\t\t\t/*\n\t\t\t * We need at least 22 characters more: the\n\t\t\t * '\\0' and the SHA1, and then the next entry.\n\t\t\t * We know the ASCII mode is 4 characters, so\n\t\t\t * we migth as well make the rule \"within 26 of\n\t\t\t * end end\".\n\t\t\t */\n\t\t\tif (guess >= size-26)\n\t\t\t\tgoto fall_back_to_linear_search;\n\t\t\tc = start[guess++];\n\t\t\tif (c)\n\t\t\t\tcontinue;\n\t\t\t/* Found it? */\n\t\t\tif (guess > goal_zero)\n\t\t\t\tbreak;\n\t\t\t/*\n\t\t\t * We found a zero that wasn't 20 bytes away, \n\t\t\t * that means we have to reset out goal..\n\t\t\t */\n\t\t\tlast_zero = guess + 20;\n\t\t}\n\t\t/*\n\t\t * \"guess\" now points to one past the '\\0': the SHA1 of\n\t\t * the previous entry. Add 20, and it points at the start\n\t\t * of a valid tree entry.\n\t\t */\n\t\tguess = guess + 20;\n\n\t\t/* Length of the entry: ascii string + '\\0' + SHA1 */\n\t\tthisentrylen = strlen(start + guess) + 1 + 20;\n\n\t\t.. compare the entry we found with\n\t\t.. the entry we are looking for!\n\t\tif (found < lookedfor) {\n\t\t\tsize = guess;\n\t\t\tcontinue;\n\t\t} else if (found == lookedfor) {\n\t\t\tYay! FOUND!\n\t\t} else {\n\t\t\tguess += thisentry;\n\t\t\tsize -= guess;\n\t\t\tstart += guess;\n\t\t\tcontinue;\n\t\t}\n\t}\n  fall_back_to_linear_search:\n\n\t.. linear search in [ start, size ] ..\n\n\nAnyway, as you can tell, the above is totally untested, but I really think \nit should work. Whether it really helps, I dunno. But if somebody is \ninterested in trying, it might be cool to see.\n\nAnd yes, the \"search for zero bytes\" is not *guaranteed* to find any \nbeginning at all, if you have lots of short names, *and* lots of zero \nbytes in the SHA1's. But while short names may be common, zero bytes in \nSHA1's are not so much (since you should expect to see a very even \ndistribution of bytes, and as such most SHA1's by far should have no zero \nbytes at all!)\n\nSo if you're really really *really* unlucky, you might end up having to \nfall back on the linear search. But it still works!\n\nCan anybody see anything wrong in my thinking above?\n\n(And the real question is whether it really helps. I suspect it does \nactually help for big directories, and that it is worth doing, but maybe \nthe magic number in \"while (size > 200)\" could be tweaked.\n\nThe logic of that was that binary searching doesn't work very well for \njust a few entries (and \"size < 200\" implies ~5-6 directory entries), but \nalso that linear search is actually perfectly good when it's just a couple \nof cache-lines, and binary searching - especially with the complication of \nhaving to find the beginning - isn't worth it unless it really means that \nwe can avoid a cache miss.\n\nOf course, it may well be that the *real* cost of the directories is just \nthe uncompression thing, and that the search is not the problem. Who \nknows? \n\n\t\t\tLinus\n"},{"id":"47806","messageId":"7vfy3l3rj0.fsf@assigned-by-dhcp.cox.net","threadId":"9051","inReplyTo":"alpine.LFD.0.999.0707181949490.27353@woody.linux-foundation.org","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-19T05:13:07Z","receivedAt":"2007-07-19T05:13:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> And yes, the \"search for zero bytes\" is not *guaranteed* to find any \n> beginning at all, if you have lots of short names, *and* lots of zero \n> bytes in the SHA1's. But while short names may be common, zero bytes in \n> SHA1's are not so much (since you should expect to see a very even \n> distribution of bytes, and as such most SHA1's by far should have no zero \n> bytes at all!)\n>\n> So if you're really really *really* unlucky, you might end up having to \n> fall back on the linear search. But it still works!\n>\n> Can anybody see anything wrong in my thinking above?\n\nAnother anchoring clue you seem not to be exploiting fully is\nthat the ASCII part must match \"^[1-7][0-7]{4,5} \" (mode bytes).\nBut the real problem of this approach of course is that this is\nnot reliable and can get a false match.  You can find your\nbeginning NUL in the SHA-1 part of one entry, and terminating\nNUL later in the SHA-1 part of next entry, and you will never\nnotice.\n\nHowever, in the case of Dscho's \"notes\" code, I do not think (1)\nyou do not have to guess like the above, and (2) the problem is\nmuch simpler.\n\nDcsho's \"note\" looks like a tree full of two-byte [0-9a-f]{2}\nnames, each of them points at another tree, with the second\nlevel tree being full of 32-byte [0-9a-f]{38} names, each of\nthem points at a blob.  So it is a much more regular, strict\nshape.  And in order to look for a note for an object whose name\nis ([0-9a-f]{2})([0-9a-f]{38}), you will find the blob that is\nat \"$1/$2\" in a \"note\".\n\nI was suggesting to have a specialized parser only to read such\ntree objects that are \"abused\" to represent notes.  You can\ncheaply validate that these trees are of expected shape.\n\n (1) Validate that size of the toplevel tree is multiple of 29 =\n     (5 + 1 + 2 + 1 + 20); the second level should be multiple\n     of 66 = (6 + 1 + 38 + 1 + 20).  These two levels of trees\n     are of fixed-entry-length that allows easy binary search.\n\n (2) While binary searching trees of either level, you can\n     validate that the entry looks like from a note (for the\n     toplevel, \"40000 [0-9a-f]{2}\\0\", for the second level,\n     \"100644 [0-9a-f]{38}\\0\").\n\nFor an added safety, a \"notes\" writer could even throw in\nsignature bytes (say, a symlink whose name is \" !\" in the\ntop-level tree, and another symlink \" !{37}\" in the second-level\ntree) to protect the reader.\n"},{"id":"47824","messageId":"E4F64312-3F86-49F3-B6BD-D148AFBAB520@wincent.com","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707190258550.14781@racer.site","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2007-07-19T09:05:32Z","receivedAt":"2007-07-19T09:05:32Z","isPatch":true,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 19/7/2007, a las 4:30, Johannes Schindelin escribió:\n\n> Commit notes are blobs which are shown together with the commit\n> message.  These blobs are taken from the notes ref, which you can\n> configure by the config variable core.notesRef, which in turn can\n> be overridden by the environment variable GIT_NOTES_REF.\n\nI was trying to look back and find out what the rationale/usage  \nscenario for these commit notes might be but Googling for 'git  \n\"commit notes\"' doesn't turn up much other than the original patch  \nyou sent a few days ago.\n\nIs this an evolution of the \"git-note: A mechanisim for providing  \nfree-form after-the-fact annotations on commits\" first introduced here?:\n\n<http://lists.zerezo.com/git/msg465441.html>\n\nCheers,\nWincent\n"},{"id":"47825","messageId":"Pine.LNX.4.64.0707191016350.14781@racer.site","threadId":"9051","inReplyTo":"E4F64312-3F86-49F3-B6BD-D148AFBAB520@wincent.com","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T09:24:57Z","receivedAt":"2007-07-19T09:24:57Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 19 Jul 2007, Wincent Colaiuta wrote:\n\n> El 19/7/2007, a las 4:30, Johannes Schindelin escribi?:\n> \n> > Commit notes are blobs which are shown together with the commit\n> > message.  These blobs are taken from the notes ref, which you can\n> > configure by the config variable core.notesRef, which in turn can\n> > be overridden by the environment variable GIT_NOTES_REF.\n> \n> I was trying to look back and find out what the rationale/usage scenario for\n> these commit notes might be but Googling for 'git \"commit notes\"' doesn't\n> turn up much other than the original patch you sent a few days ago.\n> \n> Is this an evolution of the \"git-note: A mechanisim for providing free-form\n> after-the-fact annotations on commits\" first introduced here?:\n> \n> <http://lists.zerezo.com/git/msg465441.html>\n\nAlmost.  It is an evolution of the evolution of this.\n\nhttp://thread.gmane.org/gmane.comp.version-control.git/52598/focus=52603\n\n(which started this thread you were replying to) hints at that, but you're \nright, I failed to give an explicit reference:\n\nhttp://article.gmane.org/gmane.comp.version-control.git/49588\n\nBackground: It was discussed how to go about storing notes (in the mail \nyou cited).  I was convinced that Johan's 15-strong patch series was not \noptimal, in that it tried to introduce a _second_ object store, \n_exclusively_ for commit notes, with all kinds of problems like \"how to \nfetch it?\".\n\nAfter thinking about how to avoid duplicating the object store, I posted \nmy proposal, in the second link I gave.\n\nIt was shot down, because of scalability problems.  They were not serious, \nbut hurt enough that I stalled working on it, until Alberto reminded me.\n\nSince I felt bad about shooting down Johan's patch series, and then not \ncompleting my alternative solution, I ended up working on it some more.  \nThe WIP patch 6/6 hints at what I will submit in the next days, to speed \nup in a transparent manner what would otherwise not scale well.\n\nCiao,\nDscho\n"},{"id":"47826","messageId":"7vodi83fg7.fsf@assigned-by-dhcp.cox.net","threadId":"9051","inReplyTo":"7vfy3l3rj0.fsf@assigned-by-dhcp.cox.net","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-19T09:34:00Z","receivedAt":"2007-07-19T09:34:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> ...\n>> So if you're really really *really* unlucky, you might end up having to \n>> fall back on the linear search. But it still works!\n>>\n>> Can anybody see anything wrong in my thinking above?\n> ...\n> But the real problem of this approach of course is that this is\n> not reliable and can get a false match.  You can find your\n> beginning NUL in the SHA-1 part of one entry, and terminating\n> NUL later in the SHA-1 part of next entry, and you will never\n> notice.\n\nIn other words, if you are really really *really* unlucky, not\nonly you might end up being fooled by random byte sequences in\nSHA-1 part of the tree object, you would not even notice that\nyou have to fall back on the linear search.\n\nI've long time ago concluded that if we care about reliability\n(and we do very much), a bisectable tree without breaking\nbackward compatibility is impossible.  I was hoping to find a\n\"hole\" in tree object format so that I can place an extended\nsection that is invisible to older versions of git, and place a\ntable that records offsets of each tree entries to help\nbisection and/or perhaps a hash table to help look-up, but I do\nnot think it is possible.  In the case of index file, the\noriginal file format had a hole after the cache-entry array\nwhere we can later squeeze an extension section that is\ninvisible to older versions of git.  But the tree object format\nis designed so tight that I do not see there is any place to put\nan extension section.\n\nSide note: I also think adding \"extension section\" to tree\nobject is not a good idea to begin with.  The data nor length of\nsuch a section cannot participate in hash computation to derive\nthe tree's object name so that we can still compare two tree\nobjects (with and without such extension) that have the same\ncontents by only looking at their object names.  But having\ncontents that are not counted as parts of the object's name goes\nagainst the reliability and safety of git.\n\n> ...\n> I was suggesting to have a specialized parser only to read such\n> tree objects that are \"abused\" to represent notes.  You can\n> cheaply validate that these trees are of expected shape.\n> ...\n> For an added safety, a \"notes\" writer could even throw in\n> signature bytes (say, a symlink whose name is \" !\" in the\n> top-level tree, and another symlink \" !{37}\" in the second-level\n> tree) to protect the reader.\n\nOf course, even with the above trick with relatively cheap\nvalidation based on size, entry format, and \"signature entries\",\nthe way I outlined to speed up \"notes\" access really relies on\nthe tree objects used in \"notes\" to be well formed.  If somebody\nthrows in a tree that is not really a \"note\" to refs/notes/, and\nif I am really really *really* unlucky, not only I might end up\nbeing fooled by random byte sequences in SHA-1 part of the tree\nobject, I would not even notice that I am reading garbage and\nend up giving garbage as \"note\" to the object back to the user.\n"},{"id":"47827","messageId":"Pine.LNX.4.64.0707191048030.14781@racer.site","threadId":"9051","inReplyTo":"alpine.LFD.0.999.0707181949490.27353@woody.linux-foundation.org","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T09:50:17Z","receivedAt":"2007-07-19T09:50:17Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 18 Jul 2007, Linus Torvalds wrote:\n\n> On Thu, 19 Jul 2007, Johannes Schindelin wrote:\n> >\n> > There is one severe shortcoming, though.  Since tree objects can \n> > contain file names of a variable length, it is not possible to do a \n> > binary search for the correct base name in the tree object's contents.\n> \n> Well, I've been thinking about this, and that's not really entirely \n> correct.\n> \n> It *is* possible to do a binary search, it's just a bit complicated, \n> because you have to take the \"halfway\" thing, and find the beginning of \n> an entry.\n\nI will try to work from your proposal, and do some timings.  But for the \nnotes, I really, really like the average constant running time of the hash \nmap.  As you can see from my timings in\n\nhttp://thread.gmane.org/gmane.comp.version-control.git/52598/focus=52603\n\nit does make a difference, compared to binary search.\n\nCiao,\nDscho\n"},{"id":"47828","messageId":"20070719095400.GB999MdfPADPa@greensroom.kotnet.org","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707190258550.14781@racer.site","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Sven Verdoolaege","fromEmail":"skimo@kotnet.org","sentAt":"2007-07-19T09:54:00Z","receivedAt":"2007-07-19T09:54:00Z","isPatch":true,"sender":{"key":"skimo@kotnet.org","avatar":null},"body":"On Thu, Jul 19, 2007 at 03:30:43AM +0100, Johannes Schindelin wrote:\n> +If such a path exists in the given ref, the referenced blob is read, and\n> +appended to the commit message, separated by a \"Notes:\" line.  If the\n> +given ref itself does not exist, it is not an error, but means that no\n> +notes should be print.\n\nprinted?\n\nskimo\n"},{"id":"47831","messageId":"46aec1580707190257i6e2f7e4bte61748a67549e434@mail.gmail.com","threadId":"9051","inReplyTo":"7vodi83fg7.fsf@assigned-by-dhcp.cox.net","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Adam Hayek","fromEmail":"adam.hayek@gmail.com","sentAt":"2007-07-19T09:57:06Z","receivedAt":"2007-07-19T09:57:06Z","isPatch":true,"sender":{"key":"adam.hayek@gmail.com","avatar":null},"body":"On 7/19/07, Junio C Hamano <gitster@pobox.com> wrote:\n> Side note: I also think adding \"extension section\" to tree\n> object is not a good idea to begin with.  The data nor length of\n> such a section cannot participate in hash computation to derive\n> the tree's object name so that we can still compare two tree\n> objects (with and without such extension) that have the same\n> contents by only looking at their object names.  But having\n> contents that are not counted as parts of the object's name goes\n> against the reliability and safety of git.\n\nExcuse me for being brand new to this list and to git itself, but if\nthe issue is where to put \"extra\" data to go along with a given object\nthere should be a relatively simple way to do it.  If your object's\nhash/name is X, take the string (\"%s-extra\", X), hash that, and use\nthe resulting hash as the name of the file to store whatever extra\ndata you have.  There would be the issue of when and how to move this\nnew file when the original file is moved, but old versions of git at\nleast wouldn't break, they'd just never know about the extra data in\nthe separate file.  Of course you could do endless variations of this\nto store whatever classes of extra data separately.\n"},{"id":"47833","messageId":"20070719103436.GA9143@dspnet.fr.eu.org","threadId":"9051","inReplyTo":"alpine.LFD.0.999.0707181949490.27353@woody.linux-foundation.org","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Olivier Galibert","fromEmail":"galibert@pobox.com","sentAt":"2007-07-19T10:34:36Z","receivedAt":"2007-07-19T10:34:36Z","isPatch":true,"sender":{"key":"galibert@pobox.com","avatar":null},"body":"On Wed, Jul 18, 2007 at 08:28:27PM -0700, Linus Torvalds wrote:\n> And yes, the \"search for zero bytes\" is not *guaranteed* to find any \n> beginning at all, if you have lots of short names, *and* lots of zero \n> bytes in the SHA1's. But while short names may be common, zero bytes in \n> SHA1's are not so much (since you should expect to see a very even \n> distribution of bytes, and as such most SHA1's by far should have no zero \n> bytes at all!)\n\nThe probability of a sha1 to have a zero is approximatively 0.075.\nThat's 1 in 13, more or less.\n\n  OG.\n"},{"id":"47838","messageId":"200707191158.37713.andyparkins@gmail.com","threadId":"9051","inReplyTo":"7vodi83fg7.fsf@assigned-by-dhcp.cox.net","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-07-19T10:58:36Z","receivedAt":"2007-07-19T10:58:36Z","isPatch":true,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Thursday 2007 July 19, Junio C Hamano wrote:\n\n> I've long time ago concluded that if we care about reliability\n> (and we do very much), a bisectable tree without breaking\n> backward compatibility is impossible.  I was hoping to find a\n> \"hole\" in tree object format so that I can place an extended\n\nIn the case of the notes system, is there not a big hole available because the \nlayout is under tight control?\n\n100644 blob 24631df5c6fceef7f0859903397d81f99a723197    __notes_index\n040000 tree dd3f40129c8731b1bdce1d3939de3cdc24a87783    00\n040000 tree 2b25612b5d8ee9ef469e72bbf74eab0ec00ae87f    01\n\nIn fact, this technique would work for normal tree objects too, except that \nyou'd have to be willing to pick some blob name that would always be the \nfirst entry in every tree object, and would never clash with a real file in \nthe tree.  Speaking off the top of my head, anything with \"/\" in it would be \nan invalid name so\n\n100644 blob 24631df5c6fceef7f0859903397d81f99a723197    /tree_index\n040000 tree dd3f40129c8731b1bdce1d3939de3cdc24a87783    00\n040000 tree 2b25612b5d8ee9ef469e72bbf74eab0ec00ae87f    01\n\nWould be an easy one to special case, and would be guaranteed not to clash \nwith a file in the tree.\n\nJust an idea.  I would imagine it's as daft as all my others :-)\n\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"47840","messageId":"Pine.LNX.4.64.0707191209200.14781@racer.site","threadId":"9051","inReplyTo":"200707191158.37713.andyparkins@gmail.com","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T11:10:02Z","receivedAt":"2007-07-19T11:10:02Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 19 Jul 2007, Andy Parkins wrote:\n\n> On Thursday 2007 July 19, Junio C Hamano wrote:\n> \n> > I've long time ago concluded that if we care about reliability\n> > (and we do very much), a bisectable tree without breaking\n> > backward compatibility is impossible.  I was hoping to find a\n> > \"hole\" in tree object format so that I can place an extended\n> \n> In the case of the notes system, is there not a big hole available \n> because the layout is under tight control?\n\nNo.  It is a tree object, referenced from a ref.  You can always check it \nout, modify it, and check it in.  If only by mistake.\n\nCiao,\nDscho\n"},{"id":"47865","messageId":"200707191533.48641.andyparkins@gmail.com","threadId":"9051","inReplyTo":"Pine.LNX.4.64.0707191209200.14781@racer.site","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-07-19T14:33:47Z","receivedAt":"2007-07-19T14:33:47Z","isPatch":true,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Thursday 2007 July 19, Johannes Schindelin wrote:\n\n> > In the case of the notes system, is there not a big hole available\n> > because the layout is under tight control?\n>\n> No.  It is a tree object, referenced from a ref.  You can always check it\n> out, modify it, and check it in.  If only by mistake.\n\nI was arguing for the tree-index being special cased though (ideally with an \ninvalid filename), such that it could never actually be checked out or \nchecked in, but would be maintained automatically \"git-side\".  For backwards \ncompatibility, it would be optional; and making it an invalid filename \n\nIt was only a suggestion to answer Junio's request for a \"hole\" through which \na tree-object index could be poked.\n\nIf we're only talking about the notes tree, then would it matter that it could \nbe checked out and checked in?  If someone chose to do that then it would be \ntheir own fault when the index didn't work.  If I wanted I could \nedit .git/objects/ directly - I wouldn't expect poor git to work correctly \nafterwards though.\n\n\n\nAndy\n\n\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"47879","messageId":"alpine.LFD.0.999.0707191013440.27353@woody.linux-foundation.org","threadId":"9051","inReplyTo":"7vfy3l3rj0.fsf@assigned-by-dhcp.cox.net","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-19T17:20:34Z","receivedAt":"2007-07-19T17:20:34Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 18 Jul 2007, Junio C Hamano wrote:\n> \n> Another anchoring clue you seem not to be exploiting fully is\n> that the ASCII part must match \"^[1-7][0-7]{4,5} \" (mode bytes).\n\nI did that on purpose.\n\nThe SHA1 *can* contain those characters too, so that's not really useful \nto us, and the only special character really is the NUL character (which \nis the only one cannot exists in the ASCII part - old-style trees can \ncontain '/' too, although that's going away).\n\nAlso, the mode bytes may not be visible: if we start in a long filename, \nwe'll never have looked at the mode bytes, but if we see a NUL character \nafter having seen 20 non-NUL characters (long filename), we already know \nwe got it. So I don't think we can even usefully use the other knowledge \nof the format of the ASCII part (other than to know it doesn't contain \nNUL's).\n\nOf course, we can (and should) verify that the tree entry we find is \nvalid, and *then* it makes sense to check the rules for the ASCII part. \nBut that's only after we have already found the place.\n\n> I was suggesting to have a specialized parser only to read such\n> tree objects that are \"abused\" to represent notes.  You can\n> cheaply validate that these trees are of expected shape.\n\nSure. That said, I'm less interested in the notes than I am in the cost fo \n\"git blame\", and that could be optimized by having some special code in \n\"tree_entry_interesting()\" to find the tree entries using binary search.\n\nThe special code would trigger only for:\n - large trees\n - \"opt->nr_paths == 1\"\n\nbut the latter case is the one that matters for blame in the first place, \nso..\n\n\t\tLinus\n"},{"id":"47883","messageId":"alpine.LFD.0.999.0707191032320.27353@woody.linux-foundation.org","threadId":"9051","inReplyTo":"7vodi83fg7.fsf@assigned-by-dhcp.cox.net","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-19T17:42:18Z","receivedAt":"2007-07-19T17:42:18Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 19 Jul 2007, Junio C Hamano wrote:\n> > ...\n> > But the real problem of this approach of course is that this is\n> > not reliable and can get a false match.  You can find your\n> > beginning NUL in the SHA-1 part of one entry, and terminating\n> > NUL later in the SHA-1 part of next entry, and you will never\n> > notice.\n\n[ I didn't react to this in your first email, because I thought you were \n  talking about your \"use the rules for the ASCII part\", and thought you \n  talked about how *that* was not reliable and can get a false match). But \n  it seems that you were actually talking about the NUL character test ]\n\nNope, wrong.\n\nWhy? Because there must always be a NUL *between* different SHA1's. \nThere's *always* a NUL character that precedes a SHA1. So when you have \ntwo NUL characters (with no other NUL's between them), you *know* that \nthey cannot be from two different SHA1's. If the first one was from an \nearlier SHA1, then the second one is *guaranteed* to be the one that \nhappens *before* the next SHA1.\n\nSee?\n\nYou really have two, and only two cases:\n\n - NUL's that are within 20 bytes of each other: you don't know anything \n   about them. It might be that they are both within the *same* SHA1, or \n   the first one was the one that separated the ASCII part from the SHA1, \n   or the first one was a NUL in the previous SHA1 and the second one was \n   the NUL after the ASCII part.\n\n   So two NUL's in the same 21-byte region are not reliable (ie less than \n   20 bytes in *between* them). They tell you nothing, and you must just \n   ignore them. \n\n - NUL's that are more than 20 bytes apart: the second NUL is *guaranteed* \n   to be the start of the next SHA1.\n\n   They cannot be part of the same \"NUL + sha1\", and thus the first NUL \n   *must* be from a previous SHA1 (or the NUL that preceded it). And that \n   means that the second NUL *must* be the NUL that precedes the next \n   SHA1.\n\nSo there is *no* ambiguity what-so-ever. It's not about guessing, and it's \nnot about \"luck\". If you don't find two NUL bytes separated by more than \n20 bytes, you start the linear search.\n\n> In other words, if you are really really *really* unlucky, not\n> only you might end up being fooled by random byte sequences in\n> SHA-1 part of the tree object, you would not even notice that\n> you have to fall back on the linear search.\n\nWrong. Either you find a guanteed rigth place, or you ran out of the \nbuffer and know you have to fall back on the linear search. \n\nNo fooled.\n\n> I've long time ago concluded that if we care about reliability\n> (and we do very much), a bisectable tree without breaking\n> backward compatibility is impossible.\n\nNo. You concluded incorrectly. I'm pretty damn sure the current tree \nformat is perfectly fine. It's dense, it's nice and linear, and it's \neasily bisectable.\n\n\t\tLinus\n"},{"id":"47885","messageId":"alpine.LFD.0.999.0707191043430.27353@woody.linux-foundation.org","threadId":"9051","inReplyTo":"20070719103436.GA9143@dspnet.fr.eu.org","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-19T17:50:37Z","receivedAt":"2007-07-19T17:50:37Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 19 Jul 2007, Olivier Galibert wrote:\n\n> On Wed, Jul 18, 2007 at 08:28:27PM -0700, Linus Torvalds wrote:\n> > And yes, the \"search for zero bytes\" is not *guaranteed* to find any \n> > beginning at all, if you have lots of short names, *and* lots of zero \n> > bytes in the SHA1's. But while short names may be common, zero bytes in \n> > SHA1's are not so much (since you should expect to see a very even \n> > distribution of bytes, and as such most SHA1's by far should have no zero \n> > bytes at all!)\n> \n> The probability of a sha1 to have a zero is approximatively 0.075.\n> That's 1 in 13, more or less.\n\nSure. And since we handle it fine even when it does happen, we don't care.\n\nIn fact, since we only need 20 non-zero bytes in between zeroes to know \nthat it's ok, and since the ASCII part is already 7 bytes of \"mode + \nspace\" plus <n> bytes of actual name (let's say that names average to be \nabout 8 characters - which is low: in the kernel it seems to be about 10.5 \ncharacters), we can say that the ASCII part of a tree tends to be about 15 \ncharacters.\n\nSo in order to be unlucky, it's not enough for the previous SHA1 to have a \nNUL character in it, it actually has to be in the last five bytes of the \nSHA1 - so now we're talking something like a 1:50 chance.\n\nAnd with longer names, it matters even less (to the point where it \nmatters not at all if all filenames are >= 14 characters in length).\n\nSo we can be unlucky, but it's fairly rare, and when it happens, at worst \nwe'll just need to scan to the next entry (and if we're *really* unlucky \nand it keeps happening until we scan until the end, we'll have to do the \nlinear search).\n\nThe point being that you always get the right answer, and the likelihood \nthat you have to do something slow to get that rigth answer is really \nreally low.\n\n\t\tLinus\n"},{"id":"47913","messageId":"7vbqe82afj.fsf@assigned-by-dhcp.cox.net","threadId":"9051","inReplyTo":"alpine.LFD.0.999.0707191032320.27353@woody.linux-foundation.org","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-20T00:20:00Z","receivedAt":"2007-07-20T00:20:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Thu, 19 Jul 2007, Junio C Hamano wrote:\n>> > ...\n>> > But the real problem of this approach of course is that this is\n>> > not reliable and can get a false match.  You can find your\n>> > beginning NUL in the SHA-1 part of one entry, and terminating\n>> > NUL later in the SHA-1 part of next entry, and you will never\n>> > notice.\n>\n> [ I didn't react to this in your first email, because I thought you were \n>   talking about your \"use the rules for the ASCII part\", and thought you \n>   talked about how *that* was not reliable and can get a false match). But \n>   it seems that you were actually talking about the NUL character test ]\n>\n> Nope, wrong.\n>\n> Why? Because there must always be a NUL *between* different SHA1's. \n> There's *always* a NUL character that precedes a SHA1. So when you have \n> two NUL characters (with no other NUL's between them), you *know* that \n> they cannot be from two different SHA1's. If the first one was from an \n> earlier SHA1, then the second one is *guaranteed* to be the one that \n> happens *before* the next SHA1.\n>\n> See?\n\nOk.  As usual, you are more right than I am ;-).\n"},{"id":"47923","messageId":"20070720045909.GM32566@spearce.org","threadId":"9051","inReplyTo":"7vodi83fg7.fsf@assigned-by-dhcp.cox.net","subject":"Re: [REVISED PATCH 2/6] Introduce commit notes","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-07-20T04:59:09Z","receivedAt":"2007-07-20T04:59:09Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <gitster@pobox.com> wrote:\n> I've long time ago concluded that if we care about reliability\n> (and we do very much), a bisectable tree without breaking\n> backward compatibility is impossible.  I was hoping to find a\n> \"hole\" in tree object format so that I can place an extended\n> section that is invisible to older versions of git, and place a\n> table that records offsets of each tree entries to help\n> bisection and/or perhaps a hash table to help look-up, but I do\n> not think it is possible.\n...\n> But the tree object format\n> is designed so tight that I do not see there is any place to put\n> an extension section.\n\nI came to the same conclusion the last time I thought about this\nproblem, for all the same reasons you outlined.  And came up with\npack v4.  Because the only way I could see that we could produce\na more optimal tree was to just use a different *compression* of\nthe tree, while still keeping its data the same.  Nico seemed to\nagree at the time, because he worked on the prototype with me.  :-)\n\nIts still hanging around in my fastimport repository.  But has not\nbeen merged with any recent Git, and it still needs a lot of work.\n\n-- \nShawn.\n"}]}