{"thread":{"id":"24536","subject":"inotify daemon speedup for git [POC/HACK]","startedAt":"2010-07-27T12:20:18Z","lastAt":"2010-08-13T17:58:50Z","messageCount":18,"participants":["Finn Arne Gangstad","Avery Pennarun","Joshua Juran","Sverre Rabbelier","Shawn O. Pearce","Jonathan Nieder","Ævar Arnfjörð Bjarmason","Nguyen Thai Ngoc Duy","Theodore Tso","Jakub Narebski","Enrico Weigelt"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"146471","messageId":"20100727122018.GA26780@pvv.org","threadId":"24536","inReplyTo":null,"subject":"inotify daemon speedup for git [POC/HACK]","fromName":"Finn Arne Gangstad","fromEmail":"finnag@pvv.org","sentAt":"2010-07-27T12:20:18Z","receivedAt":"2010-07-27T12:20:18Z","isPatch":false,"sender":{"key":"finnag@pvv.org","avatar":"https://gravatar.com/avatar/b421ddd58c3f0f93aa473e17b98bb8d53c221fef741746bc8cb59fae4ec6d95e?d=mp&s=160"},"body":"Reading through the thread about subtree I noticed Avery mentioning\nusing inotify to speed up git status & co.\n\nHere is a quick hack I did some time ago to test this out, to use it\ncall \"igit\" instead of \"git\" for all commmands you want to speed up.\n\nThere is one minor nit: The speedup gain is zero :) git still\ntraverses all directories to look for .gitignore files, which seems to\ntotally kill the optimisation.\n\nTo use it, put igit and git-inotify-daemon.pl in path, and do git\nconfig core.ignorestat=true in the repositories you want to test it\nwith. The igit wrapper will run git update-index --no-assume-unchanged\non all modified files before running any real git commands.\n\nTo get inotify to ignore all changes that the git commands themselves\nperform, the \"igit\" wrapper kills the currently running daemon. Then\nit reads the list of updates files, and does git-update-index\n--no-assume-unchanged on them. Then the git command is run, and\nfinally the daemon is fired up again.\n\nI had to do one tiny modification to git to make update-index ignore\nbad paths.\n\n\nigit - a git wrapper with an inotify daemon\n\nLinux only - requires inotifytools installed. This is juct a quick hack/proof\nof concept!\nupdate-index: Do not error out on bad paths, just warn\n---\n .gitignore             |    1 +\n builtin/update-index.c |    2 +-\n git-inotify-daemon.pl  |   28 ++++++++++++++++++++++++++++\n igit                   |   22 ++++++++++++++++++++++\n 4 files changed, 52 insertions(+), 1 deletions(-)\n create mode 100755 git-inotify-daemon.pl\n create mode 100755 igit\n\ndiff --git a/.gitignore b/.gitignore\nindex 14e2b6b..fa67132 100644\n--- a/.gitignore\n+++ b/.gitignore\n@@ -204,3 +204,4 @@\n *.pdb\n /Debug/\n /Release/\n+.igit-*\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 3ab214d..c905d78 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -282,7 +282,7 @@ static void update_one(const char *path, const char *prefix, int prefix_length)\n \t}\n \tif (mark_valid_only) {\n \t\tif (mark_ce_flags(p, CE_VALID, mark_valid_only == MARK_FLAG))\n-\t\t\tdie(\"Unable to mark file %s\", path);\n+\t\t\tfprintf(stderr, \"Unable to mark file %s\\n\", path);\n \t\tgoto free_return;\n \t}\n \tif (mark_skip_worktree_only) {\ndiff --git a/git-inotify-daemon.pl b/git-inotify-daemon.pl\nnew file mode 100755\nindex 0000000..a57ceef\n--- /dev/null\n+++ b/git-inotify-daemon.pl\n@@ -0,0 +1,28 @@\n+#!/usr/bin/env perl\n+# Run from igit\n+\n+use warnings;\n+use strict;\n+\n+die \"Usage: $0 <output-file>\" unless $#ARGV == 0;\n+my $output = $ARGV[0];\n+my $pid = open(INOTIFY, \"exec inotifywait -q --monitor --recursive --exclude .git -e attrib,moved_to,moved_from,move,create,delete,modify --format '%w%f' .|\") or die \"Cannot run inotifywait: $!\\n\";\n+\n+$| = 1;\n+print \"$pid\\n\";\n+\n+my %modified_files;\n+while (<INOTIFY>) {\n+    s=^./==;\n+    chomp;\n+    $modified_files{$_} = 1;\n+}\n+\n+# Output file must be opened as late as possible, it is a named pipe\n+# and the listener won't be here before inotifywait exits.\n+# open would just hang if it was done earlier.\n+open(OUT, \">$output\");\n+foreach my $key (sort keys %modified_files) {\n+    print OUT \"$key\\000\";\n+}\n+exit 0;\ndiff --git a/igit b/igit\nnew file mode 100755\nindex 0000000..60c5bb2\n--- /dev/null\n+++ b/igit\n@@ -0,0 +1,22 @@\n+#!/bin/sh\n+\n+TOPDIR=`git rev-parse --show-cdup` || exit 1\n+\n+if [ ! \"$TOPDIR\" ]; then\n+    TOPDIR=\"./\"\n+fi\n+\n+PIPE=.igit-pipe\n+PIDFILE=.igit-pid\n+\n+if [ -p ${TOPDIR}${PIPE} ] && kill -TERM `cat ${TOPDIR}${PIDFILE}`; then\n+    ( cd $TOPDIR && git update-index --verbose --no-assume-unchanged -z --stdin < $PIPE )\n+fi\n+\n+git \"$@\"\n+\n+cd $TOPDIR\n+rm -f $PIPE\n+mkfifo $PIPE\n+git-inotify-daemon.pl $PIPE > $PIDFILE 2>> .igit-errors </dev/null &\n+\n-- \n1.7.2.rc0\n\n\n- Finn Arne\n"},{"id":"146551","messageId":"AANLkTinuU6b1vmRFuBrA4Tc5H6gmC5cMP3Pa8EYz-8JE@mail.gmail.com","threadId":"24536","inReplyTo":"20100727122018.GA26780@pvv.org","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-27T23:29:41Z","receivedAt":"2010-07-27T23:29:41Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Tue, Jul 27, 2010 at 8:20 AM, Finn Arne Gangstad <finnag@pvv.org> wrote:\n> Reading through the thread about subtree I noticed Avery mentioning\n> using inotify to speed up git status & co.\n>\n> Here is a quick hack I did some time ago to test this out, to use it\n> call \"igit\" instead of \"git\" for all commmands you want to speed up.\n>\n> There is one minor nit: The speedup gain is zero :) git still\n> traverses all directories to look for .gitignore files, which seems to\n> totally kill the optimisation.\n\nHey, this is kind of cool.  Except for that last part :)\n\nActually I think the problem is a little worse than .gitignore files.\n'git status', for example (which is called by git commit), wants to\ngenerate a list of the files it *doesn't* know about.  Unfortunately,\nthose files aren't in the index at all.  So it resorts to doing\nrecursive readdir() across the entire repository.  The net result is\nabout as slow as doing that plus one stat() per file in the index.\n\nAn inotify daemon could easily keep track of which files have been\nadded that aren't in the index... but where would it put the list of\nfiles git doesn't know about?  Do they go in the index with a special\nNOT_REALLY_INDEXED flag?\n\nThis is the main question that has so far prevented me from trying to\nsolve the problem myself.\n\nThanks,\n\nAvery\n"},{"id":"146552","messageId":"9E67A084-4EDB-4CCB-A771-11B97107F4EF@gmail.com","threadId":"24536","inReplyTo":"AANLkTinuU6b1vmRFuBrA4Tc5H6gmC5cMP3Pa8EYz-8JE@mail.gmail.com","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Joshua Juran","fromEmail":"jjuran@gmail.com","sentAt":"2010-07-27T23:39:14Z","receivedAt":"2010-07-27T23:39:14Z","isPatch":false,"sender":{"key":"jjuran@gmail.com","avatar":null},"body":"On Jul 27, 2010, at 4:29 PM, Avery Pennarun wrote:\n\n> An inotify daemon could easily keep track of which files have been\n> added that aren't in the index... but where would it put the list of\n> files git doesn't know about?  Do they go in the index with a special\n> NOT_REALLY_INDEXED flag?\n\nOne option is not to write it to disk at all.  The client could  \nconsult the daemon directly.\n\nJosh\n"},{"id":"146553","messageId":"AANLkTi=oA33M4DmS5FyDx7Wn1DFrUGcmhSYkvcSYMc2r@mail.gmail.com","threadId":"24536","inReplyTo":"9E67A084-4EDB-4CCB-A771-11B97107F4EF@gmail.com","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-27T23:51:26Z","receivedAt":"2010-07-27T23:51:26Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Tue, Jul 27, 2010 at 7:39 PM, Joshua Juran <jjuran@gmail.com> wrote:\n> On Jul 27, 2010, at 4:29 PM, Avery Pennarun wrote:\n>\n>> An inotify daemon could easily keep track of which files have been\n>> added that aren't in the index... but where would it put the list of\n>> files git doesn't know about?  Do they go in the index with a special\n>> NOT_REALLY_INDEXED flag?\n>\n> One option is not to write it to disk at all.  The client could consult the\n> daemon directly.\n\nTrue.  What would the client-server protocol look like, though?  \"Give\nme the list of unknown files?\"  Does the daemon need to understand\n.gitignore or will it send back a list of all my million *.o files\nevery time?  etc.\n\nOffhandedly, I think it would be nice to have an inotify daemon just\nmaintain (something like) the git index file where it just has a list\nof *all* the files in a form that's a) random access, not just\nsequential, and b) really fast when accessed sequentially.\n\nKnowing that large numbers of files can cause slowness, I was planning\nahead for inotify when I designed bup's index file format, and it\nmeets the above criteria.  Unfortunately I screwed up other stuff\n(adding new files is too slow) and it still needs to be rewritten\nanyway.  Oh well.\n\nWhile we're here, it's probably worth mentioning that git's index file\nformat (which stores a sequential list of full paths in alphabetical\norder, instead of an actual hierarchy) does become a bottleneck when\nyou actually have a huge number of files in your repo (like literally\na million).  You can't actually binary search through the index!  The\ncurrent implementation of submodules allows you to dodge that\nscalability problem since you end up with multiple smaller index\nfiles.  Anyway, that's fixable too.\n\nHave fun,\n\nAvery\n"},{"id":"146554","messageId":"AANLkTi=6pPrQkEozTR6OXuO6C4kGk61ExTWiLD6vQ1Mp@mail.gmail.com","threadId":"24536","inReplyTo":"20100727122018.GA26780@pvv.org","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2010-07-27T23:58:31Z","receivedAt":"2010-07-27T23:58:31Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Tue, Jul 27, 2010 at 07:20, Finn Arne Gangstad <finnag@pvv.org> wrote:\n> There is one minor nit: The speedup gain is zero :) git still\n> traverses all directories to look for .gitignore files, which seems to\n> totally kill the optimisation.\n\nThis is very true. In my experience with ginormous trees even if you\n'git update-index --assume-unchanged' every file and directory it's\nstill unbearably slow due to the .gitignore files. Any solution that\naims to solve this problem should also address the .gitignore file\nproblem. Note: a safe assumption here is that to solve the problem it\nneeds to work if there are more .gitignore files than regular files\n:).\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"146555","messageId":"20100728000009.GE25268@spearce.org","threadId":"24536","inReplyTo":"AANLkTi=oA33M4DmS5FyDx7Wn1DFrUGcmhSYkvcSYMc2r@mail.gmail.com","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2010-07-28T00:00:09Z","receivedAt":"2010-07-28T00:00:09Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Avery Pennarun <apenwarr@gmail.com> wrote:\n> \n> While we're here, it's probably worth mentioning that git's index file\n> format (which stores a sequential list of full paths in alphabetical\n> order, instead of an actual hierarchy) does become a bottleneck when\n> you actually have a huge number of files in your repo (like literally\n> a million).  You can't actually binary search through the index!  The\n> current implementation of submodules allows you to dodge that\n> scalability problem since you end up with multiple smaller index\n> files.  Anyway, that's fixable too.\n\nYes.\n\nMore than once I've been tempted to rewrite the on-disk (and I guess\nin-memory) format of the index.  And then I remember how painful that\nstuff is in either C git.git or JGit, and I back away slowly.  :-)\n\nIdeally the index is organized the same way the trees are, but\nyou still can't do a really good binary search because of the\nass-backwards name sorting rule for trees.  But for performance\nreasons you still want to keep the entire index in a single file,\nan index per directory (aka SVN/CVS) is too slow for the common\ncase of <30k files.\n\n-- \nShawn.\n"},{"id":"146557","messageId":"AANLkTimkLrTwavErFkyaUTSVU-2s3me5f+cyqNFp7n+D@mail.gmail.com","threadId":"24536","inReplyTo":"20100728000009.GE25268@spearce.org","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-28T00:18:07Z","receivedAt":"2010-07-28T00:18:07Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Tue, Jul 27, 2010 at 8:00 PM, Shawn O. Pearce <spearce@spearce.org> wrote:\n> Avery Pennarun <apenwarr@gmail.com> wrote:\n>> While we're here, it's probably worth mentioning that git's index file\n>> format (which stores a sequential list of full paths in alphabetical\n>> order, instead of an actual hierarchy) does become a bottleneck when\n>> you actually have a huge number of files in your repo (like literally\n>> a million).  You can't actually binary search through the index!  The\n>> current implementation of submodules allows you to dodge that\n>> scalability problem since you end up with multiple smaller index\n>> files.  Anyway, that's fixable too.\n>\n> Yes.\n>\n> More than once I've been tempted to rewrite the on-disk (and I guess\n> in-memory) format of the index.  And then I remember how painful that\n> stuff is in either C git.git or JGit, and I back away slowly.  :-)\n>\n> Ideally the index is organized the same way the trees are, but\n> you still can't do a really good binary search because of the\n> ass-backwards name sorting rule for trees.  But for performance\n> reasons you still want to keep the entire index in a single file,\n> an index per directory (aka SVN/CVS) is too slow for the common\n> case of <30k files.\n\nReally?  What's wrong with the name sorting rule?  I kind of like it.\n\nbup's current index - after I abandoned my clone of the git one since\nit was too slow with insane numbers of files - is very fast for reads\nand in-place updates using mmap.\n\nEssentially, it's a tree, starting from the outermost leafs and\nleading toward the entry at the very end of the file, which is the\nroot.  (The idea of doing it backwards was that I could write the file\nsequentially.  In retrospect, that was probably an unnecessarily\nbrain-bending waste of time and the root should have been the first\nentry instead.)\n\nFor speed, the bup index can just mark entries as deleted using a flag\nrather than actually rewriting the whole indexfile.  Unfortunately, I\nfailed to make it sufficiently flexible to *add* new entries without\nneeding to rewrite the whole thing.  In bup, that's a big deal\n(especially since python is kind of slow and there are typically >1\nmillion files in the index).  In git, it's maybe not so bad; after\nall, the current implementation rewrites the index *every* time and\nnobody notices.\n\nAnyway, the code for it isn't too hairy, in case you want to steal some ideas:\nhttp://github.com/apenwarr/bup/blob/master/lib/bup/index.py\n\n(Disclaimer: I say this after actually spending a couple of late\nnights pulling my hair out over it.  So I'm not so hairy anymore\neither, but that doesn't prove much.)\n\nI've considered just tossing the whole thing and using sqlite instead.\n Eventually I'll do it as a benchmark to see what happens.  My past\nexperiments with sqlite have demonstrated that its performance is\nrather mind boggling (> 100k rows inserted per second as long as you\nprepare() your SQL statements).  Reading from the index would be fast,\nadding entries would be much faster than presently, but I'm not sure\nabout mass updates.  For bup sqlite would be okay, though I doubt git\nwants to take on a whole sqlite dependency.  Then again, you never\nknow.\n\nHave fun,\n\nAvery\n"},{"id":"146561","messageId":"52EDBD9A-2961-4F66-88B3-07BF873FA994@gmail.com","threadId":"24536","inReplyTo":"AANLkTimkLrTwavErFkyaUTSVU-2s3me5f+cyqNFp7n+D@mail.gmail.com","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Joshua Juran","fromEmail":"jjuran@gmail.com","sentAt":"2010-07-28T01:14:47Z","receivedAt":"2010-07-28T01:14:47Z","isPatch":false,"sender":{"key":"jjuran@gmail.com","avatar":null},"body":"On Jul 27, 2010, at 5:18 PM, Avery Pennarun wrote:\n\n> On Tue, Jul 27, 2010 at 8:00 PM, Shawn O. Pearce  \n> <spearce@spearce.org> wrote:\n>> Avery Pennarun <apenwarr@gmail.com> wrote:\n>>> While we're here, it's probably worth mentioning that git's index  \n>>> file\n>>> format (which stores a sequential list of full paths in alphabetical\n>>> order, instead of an actual hierarchy) does become a bottleneck when\n>>> you actually have a huge number of files in your repo (like  \n>>> literally\n>>> a million).  You can't actually binary search through the index!   \n>>> The\n>>> current implementation of submodules allows you to dodge that\n>>> scalability problem since you end up with multiple smaller index\n>>> files.  Anyway, that's fixable too.\n>>\n>> Yes.\n>>\n>> More than once I've been tempted to rewrite the on-disk (and I guess\n>> in-memory) format of the index.  And then I remember how painful that\n>> stuff is in either C git.git or JGit, and I back away slowly.  :-)\n>>\n>> Ideally the index is organized the same way the trees are, but\n>> you still can't do a really good binary search because of the\n>> ass-backwards name sorting rule for trees.  But for performance\n>> reasons you still want to keep the entire index in a single file,\n>> an index per directory (aka SVN/CVS) is too slow for the common\n>> case of <30k files.\n>\n> Really?  What's wrong with the name sorting rule?  I kind of like it.\n>\n> bup's current index - after I abandoned my clone of the git one since\n> it was too slow with insane numbers of files - is very fast for reads\n> and in-place updates using mmap.\n>\n> Essentially, it's a tree, starting from the outermost leafs and\n> leading toward the entry at the very end of the file, which is the\n> root.  (The idea of doing it backwards was that I could write the file\n> sequentially.  In retrospect, that was probably an unnecessarily\n> brain-bending waste of time and the root should have been the first\n> entry instead.)\n>\n> For speed, the bup index can just mark entries as deleted using a flag\n> rather than actually rewriting the whole indexfile.  Unfortunately, I\n> failed to make it sufficiently flexible to *add* new entries without\n> needing to rewrite the whole thing.  In bup, that's a big deal\n> (especially since python is kind of slow and there are typically >1\n> million files in the index).  In git, it's maybe not so bad; after\n> all, the current implementation rewrites the index *every* time and\n> nobody notices.\n\nOkay, I have an idea.  If I understand correctly, the index is a flat  \ndatabase of records including a pathname and several fixed-length  \nfields.  Since the records are not fixed-length, only sequential  \nsearch is possible, even though the records are sorted by pathname.\n\nHere's the idea:  Divide the database into blocks.  Each block  \ncontains a block header and the records belonging to a single  \ndirectory.  The block header contains the length of the block and also  \nthe offset to the next block, in bytes.  In addition to a record for  \neach indexed file in a directory, a directory's block also contains  \nrecords for subdirectories. The mode flags in a record indicate the  \nrecord type.  Directory records contain an offset in bytes to the  \nblock for that directory (in place of the SHA-1 hash).  The block list  \nis preceded by a file header, which includes the offset in bytes of  \nthe root block.  All offsets are from the beginning of the file.\n\nInstead of having to search among every file in the repository, the  \nsearch space now includes only the immediate descendants of each  \ndirectory in the target file's path.  If a directory is modified then  \nit can either be rewritten in place (if there's sufficient room) or  \nappended to the end of the file (requiring the old and new  \nsequentially preceding blocks and the parent directory's block to  \nupdate their offsets).\n\nIs this useful?\n\nJosh\n"},{"id":"146564","messageId":"AANLkTi=TQnyATgJ0LSdR3qeeCVAgu+wOFcHmHUBguPiV@mail.gmail.com","threadId":"24536","inReplyTo":"52EDBD9A-2961-4F66-88B3-07BF873FA994@gmail.com","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-28T01:31:38Z","receivedAt":"2010-07-28T01:31:38Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Tue, Jul 27, 2010 at 9:14 PM, Joshua Juran <jjuran@gmail.com> wrote:\n> Okay, I have an idea.  If I understand correctly, the index is a flat\n> database of records including a pathname and several fixed-length fields.\n>  Since the records are not fixed-length, only sequential search is possible,\n> even though the records are sorted by pathname.\n>\n> Here's the idea:  Divide the database into blocks.  Each block contains a\n> block header and the records belonging to a single directory.  The block\n> header contains the length of the block and also the offset to the next\n> block, in bytes.  In addition to a record for each indexed file in a\n> directory, a directory's block also contains records for subdirectories. The\n> mode flags in a record indicate the record type.  Directory records contain\n> an offset in bytes to the block for that directory (in place of the SHA-1\n> hash).  The block list is preceded by a file header, which includes the\n> offset in bytes of the root block.  All offsets are from the beginning of\n> the file.\n>\n> Instead of having to search among every file in the repository, the search\n> space now includes only the immediate descendants of each directory in the\n> target file's path.  If a directory is modified then it can either be\n> rewritten in place (if there's sufficient room) or appended to the end of\n> the file (requiring the old and new sequentially preceding blocks and the\n> parent directory's block to update their offsets).\n\nYeah, that's pretty much what bup's current format does, minus\nappending rewritten dirs at the end when files are added.  I've\nthought of that, but sooner or later, the file would need to be\nrewritten anyway, and then you end up with odd performance\ncharacteristics where the file expands in random ways and then shrinks\nagain when you decide it's gotten too big.  And if you do try to reuse\nempty blocks - which should mostly avoid the endless growth problem -\nyou basically just have a database, including fragmentation problems\nand multi-user concerns and all.  That's what made me think that\nsqlite might be a sensible choice, since it's already a database :)\n\nBut maybe there's some simpler way.\n\nHave fun,\n\nAvery\n"},{"id":"146574","messageId":"AANLkTinabaO3csi3TBRJKPTZ1zVGgK8-ijs6h1YUkT-n@mail.gmail.com","threadId":"24536","inReplyTo":"AANLkTi=TQnyATgJ0LSdR3qeeCVAgu+wOFcHmHUBguPiV@mail.gmail.com","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2010-07-28T06:03:14Z","receivedAt":"2010-07-28T06:03:14Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Tue, Jul 27, 2010 at 20:31, Avery Pennarun <apenwarr@gmail.com> wrote:\n> That's what made me think that\n> sqlite might be a sensible choice, since it's already a database :)\n\nSounds very sensible to me, especially the fact that (if it is indeed\nfast enough, which I can't imagine it not being) it would make\ndevelopment so much easier. At least, I think that having sqlite deal\nwith backwards comparability of your schema is easier than having to\nmanually do that? Also, sqlite is known to scale, is exactly one file\nworth of dependency, what's not to love (other than having to support\nupgrading to 'index vSqlite').\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"146575","messageId":"20100728060646.GA16400@dert.cs.uchicago.edu","threadId":"24536","inReplyTo":"AANLkTinabaO3csi3TBRJKPTZ1zVGgK8-ijs6h1YUkT-n@mail.gmail.com","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-07-28T06:06:46Z","receivedAt":"2010-07-28T06:06:46Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Sverre Rabbelier wrote:\n\n> Also, sqlite is known to scale, is exactly one file\n> worth of dependency, what's not to love (other than having to support\n> upgrading to 'index vSqlite').\n\nThe frequent fsync()-ing.  Though that seems to be a problem with\npretty much anything that does not involve rewriting the index\nwith each change.\n\nMaybe filesystems will cope better soon. :)\n"},{"id":"146590","messageId":"AANLkTinuq9Q_RADtQwvVTn-kDCw7cg7JcdkhbQnek9Tw@mail.gmail.com","threadId":"24536","inReplyTo":"20100728060646.GA16400@dert.cs.uchicago.edu","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2010-07-28T07:44:26Z","receivedAt":"2010-07-28T07:44:26Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Wed, Jul 28, 2010 at 06:06, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> Sverre Rabbelier wrote:\n>\n>> Also, sqlite is known to scale, is exactly one file\n>> worth of dependency, what's not to love (other than having to support\n>> upgrading to 'index vSqlite').\n>\n> The frequent fsync()-ing.  Though that seems to be a problem with\n> pretty much anything that does not involve rewriting the index\n> with each change.\n\nSQLite has an option to turn that off [1], but I don't know if it has\nan equivalent feature to manually call fsync when you need that.\n\nAnyway, I've been very impressed by SQLite in every way. I'd try it\nbefore designing my own fileformat, especially something involving\nbinary/sequential search. It's not a large dependency, and can easily\nbe bundled in compat/.\n\n1. http://www.sqlite.org/pragma.html#pragma_synchronous\n"},{"id":"146593","messageId":"AANLkTimqBSTRzcU++jW6izMgeA=HB00wBXQVHuSsn1oR@mail.gmail.com","threadId":"24536","inReplyTo":"AANLkTinabaO3csi3TBRJKPTZ1zVGgK8-ijs6h1YUkT-n@mail.gmail.com","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2010-07-28T08:20:55Z","receivedAt":"2010-07-28T08:20:55Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Jul 28, 2010 at 4:03 PM, Sverre Rabbelier <srabbelier@gmail.com> wrote:\n> Heya,\n>\n> On Tue, Jul 27, 2010 at 20:31, Avery Pennarun <apenwarr@gmail.com> wrote:\n>> That's what made me think that\n>> sqlite might be a sensible choice, since it's already a database :)\n>\n> Sounds very sensible to me, especially the fact that (if it is indeed\n> fast enough, which I can't imagine it not being) it would make\n> development so much easier. At least, I think that having sqlite deal\n> with backwards comparability of your schema is easier than having to\n> manually do that? Also, sqlite is known to scale, is exactly one file\n> worth of dependency, what's not to love (other than having to support\n> upgrading to 'index vSqlite').\n\nEven more sensible to replace all pack index with a single database.\nBut then we could as well drop git object store in favor of Fossil (OK\nI'm going to far).\n-- \nDuy\n"},{"id":"146613","messageId":"5C87954A-5BB2-468D-8C4E-79A97685ED0D@mit.edu","threadId":"24536","inReplyTo":"AANLkTinuq9Q_RADtQwvVTn-kDCw7cg7JcdkhbQnek9Tw@mail.gmail.com","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2010-07-28T11:08:37Z","receivedAt":"2010-07-28T11:08:37Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"\nOn Jul 28, 2010, at 3:44 AM, Ævar Arnfjörð Bjarmason wrote:\n\n> SQLite has an option to turn that off [1], but I don't know if it has\n> an equivalent feature to manually call fsync when you need that.\n\nThe right way to use SQLite is to have a memory-packed database which you check first, and where you do al of your work.  Then once you hit a stable stopping point, you commit those changes to your on-disk SQLite database, which can have proper transaction support.   That way you don't lose your database when your crappy binary-only video driver crashes on you, but you don't trash your disk performance because of the fsync() calls....\n\nIt only took a few years for firefox developers to figure this out, but the next version is supposed to finally get this right....  it'll be nice to have it NOT chewing up a third of a megabyte of SSD write endurance on every URL click....\n\n-- Ted\n"},{"id":"146618","messageId":"m3tynjkb90.fsf@localhost.localdomain","threadId":"24536","inReplyTo":"20100728000009.GE25268@spearce.org","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-07-28T13:06:22Z","receivedAt":"2010-07-28T13:06:22Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> Avery Pennarun <apenwarr@gmail.com> wrote:\n> > \n> > While we're here, it's probably worth mentioning that git's index file\n> > format (which stores a sequential list of full paths in alphabetical\n> > order, instead of an actual hierarchy) does become a bottleneck when\n> > you actually have a huge number of files in your repo (like literally\n> > a million).  You can't actually binary search through the index!  The\n> > current implementation of submodules allows you to dodge that\n> > scalability problem since you end up with multiple smaller index\n> > files.  Anyway, that's fixable too.\n> \n> Yes.\n> \n> More than once I've been tempted to rewrite the on-disk (and I guess\n> in-memory) format of the index.  And then I remember how painful that\n> stuff is in either C git.git or JGit, and I back away slowly.  :-)\n> \n> Ideally the index is organized the same way the trees are, but\n> you still can't do a really good binary search because of the\n> ass-backwards name sorting rule for trees.  But for performance\n> reasons you still want to keep the entire index in a single file,\n> an index per directory (aka SVN/CVS) is too slow for the common\n> case of <30k files.\n\nI guess that modern filesystems solve the problem of very many files\nin a single directory somehow (hash tables?).  Perhaps the index file\ncould borrow some such mechanism as an extension.\n\nIndex for index?\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"146619","messageId":"m3pqy7kb2z.fsf@localhost.localdomain","threadId":"24536","inReplyTo":"AANLkTimkLrTwavErFkyaUTSVU-2s3me5f+cyqNFp7n+D@mail.gmail.com","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-07-28T13:09:55Z","receivedAt":"2010-07-28T13:09:55Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Avery Pennarun <apenwarr@gmail.com> writes:\n\n> For speed, the bup index can just mark entries as deleted using a flag\n> rather than actually rewriting the whole indexfile.  Unfortunately, I\n> failed to make it sufficiently flexible to *add* new entries without\n> needing to rewrite the whole thing.  In bup, that's a big deal\n> (especially since python is kind of slow and there are typically >1\n> million files in the index).  In git, it's maybe not so bad; after\n> all, the current implementation rewrites the index *every* time and\n> nobody notices.\n\nSidenote: couldn't you do what e.g. Mercurial did, i.e. rewrite\ncritical for performance parts in C?\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"148012","messageId":"20100813175333.GC27540@nibiru.local","threadId":"24536","inReplyTo":"AANLkTimqBSTRzcU++jW6izMgeA=HB00wBXQVHuSsn1oR@mail.gmail.com","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Enrico Weigelt","fromEmail":"weigelt@metux.de","sentAt":"2010-08-13T17:53:33Z","receivedAt":"2010-08-13T17:53:33Z","isPatch":false,"sender":{"key":"weigelt@metux.de","avatar":null},"body":"* Nguyen Thai Ngoc Duy <pclouds@gmail.com> wrote:\n\n> But then we could as well drop git object store in favor of Fossil (OK\n> I'm going to far).\n\nYou mean venti ?\n\nActually: that's an idea I'm thinking about for quite a while :)\n\nBut venti is yet lacking delete operations and differential\ncompression. The first is unproblematic (even it would require\nrewriting the log areas in some ways to reclaim space), but\nfor differential compression, the venti store would have to\nknow a lot about the object's internal structure.\n\nI'm doing some bit reasearch in the area of distributed \ncontent-addressed objects stores , designing an superstore \ncalled \"Nebulon\" [1] with things like strong encryption and\non-demand fetching/syncing. But getting git into it seems\nto be a bit tricky, at least the hashes would change ...\n\n\ncu\n\n[1] http://www.metux.de/index.php/de/nebulon-storage-cloud.html\n-- \n----------------------------------------------------------------------\n Enrico Weigelt, metux IT service -- http://www.metux.de/\n\n phone:  +49 36207 519931  email: weigelt@metux.de\n mobile: +49 151 27565287  icq:   210169427         skype: nekrad666\n----------------------------------------------------------------------\n Embedded-Linux / Portierung / Opensource-QM / Verteilte Systeme\n----------------------------------------------------------------------\n"},{"id":"148013","messageId":"20100813175850.GD27540@nibiru.local","threadId":"24536","inReplyTo":"m3tynjkb90.fsf@localhost.localdomain","subject":"Re: inotify daemon speedup for git [POC/HACK]","fromName":"Enrico Weigelt","fromEmail":"weigelt@metux.de","sentAt":"2010-08-13T17:58:50Z","receivedAt":"2010-08-13T17:58:50Z","isPatch":false,"sender":{"key":"weigelt@metux.de","avatar":null},"body":"* Jakub Narebski <jnareb@gmail.com> wrote:\n\n> I guess that modern filesystems solve the problem of very many files\n> in a single directory somehow (hash tables?).  Perhaps the index file\n> could borrow some such mechanism as an extension.\n> \n> Index for index?\n\nhmm, if an index gets too large, it could be split into several\nones by an pathname prefix (but not necessarily one per directory)\nso when having the subdirs \"a\", \"b\", \"c\", we'll three separate \nindex files and a master index telling:\n\n    a/\tindex.001\n    b/\tindex.002\n    c/\tindex.003\n\nor even:\n\n    a/\t\tindex.001\n    b/\t\tindex.002\n    b/foo\tindex.004\n    c/\t\tindex.005\n\nthis would just add one indirection in the index-lookup by\ncomparing the key w/ index indice's prefix.\n\n\ncu\n-- \n----------------------------------------------------------------------\n Enrico Weigelt, metux IT service -- http://www.metux.de/\n\n phone:  +49 36207 519931  email: weigelt@metux.de\n mobile: +49 151 27565287  icq:   210169427         skype: nekrad666\n----------------------------------------------------------------------\n Embedded-Linux / Portierung / Opensource-QM / Verteilte Systeme\n----------------------------------------------------------------------\n"}]}