git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: inotify daemon speedup for git [POC/HACK]

From
Avery Pennarun <apenwarr@gmail.com>
Date
Jul 28, 2010, 01:31 UTC
Message-ID
<AANLkTi=TQnyATgJ0LSdR3qeeCVAgu+wOFcHmHUBguPiV@mail.gmail.com>
In-Reply-To
<52EDBD9A-2961-4F66-88B3-07BF873FA994@gmail.com>
On Tue, Jul 27, 2010 at 9:14 PM, Joshua Juran <jjuran@gmail.com> wrote:
Show 22 quoted lines
> Okay, I have an idea.  If I understand correctly, the index is a flat
> database of records including a pathname and several fixed-length fields.
>  Since the records are not fixed-length, only sequential search is possible,
> even though the records are sorted by pathname.
>
> Here's the idea:  Divide the database into blocks.  Each block contains a
> block header and the records belonging to a single directory.  The block
> header contains the length of the block and also the offset to the next
> block, in bytes.  In addition to a record for each indexed file in a
> directory, a directory's block also contains records for subdirectories. The
> mode flags in a record indicate the record type.  Directory records contain
> an offset in bytes to the block for that directory (in place of the SHA-1
> hash).  The block list is preceded by a file header, which includes the
> offset in bytes of the root block.  All offsets are from the beginning of
> the file.
>
> Instead of having to search among every file in the repository, the search
> space now includes only the immediate descendants of each directory in the
> target file's path.  If a directory is modified then it can either be
> rewritten in place (if there's sufficient room) or appended to the end of
> the file (requiring the old and new sequentially preceding blocks and the
> parent directory's block to update their offsets).

Yeah, that's pretty much what bup's current format does, minus appending rewritten dirs at the end when files are added. I've thought of that, but sooner or later, the file would need to be rewritten anyway, and then you end up with odd performance characteristics where the file expands in random ways and then shrinks again when you decide it's gotten too big. And if you do try to reuse empty blocks - which should mostly avoid the endless growth problem - you basically just have a database, including fragmentation problems and multi-user concerns and all. That's what made me think that sqlite might be a sensible choice, since it's already a database :)

But maybe there's some simpler way.
Have fun,
Avery
Previous: Joshua JuranNext: Sverre Rabbelier
Message 8 of 18 in “inotify daemon speedup for git [POC/HACK]”
  1. Finn Arne GangstadJul 27, 2010
  2. Avery PennarunJul 27, 2010
  3. Joshua JuranJul 27, 2010
  4. Avery PennarunJul 27, 2010
  5. Shawn O. PearceJul 28, 2010
  6. Avery PennarunJul 28, 2010
  7. Joshua JuranJul 28, 2010
  8. Avery PennarunJul 28, 2010
  9. Sverre RabbelierJul 28, 2010
  10. Jonathan NiederJul 28, 2010
  11. Ævar Arnfjörð BjarmasonJul 28, 2010
  12. Theodore TsoJul 28, 2010
  13. Nguyen Thai Ngoc DuyJul 28, 2010
  14. Enrico WeigeltAug 13, 2010
  15. Jakub NarebskiJul 28, 2010
  16. Jakub NarebskiJul 28, 2010
  17. Enrico WeigeltAug 13, 2010
  18. Sverre RabbelierJul 27, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.