Re: GSoC - Designing a faster index format
- From
Thomas Rast <trast@student.ethz.ch>
- Date
- Mar 26, 2012, 14:28 UTC
- Message-ID
- <87iphrjv23.fsf@thomas.inf.ethz.ch>
- In-Reply-To
- <CAKTdtZkx+7iU5T4oBNDEx-A5cgZCLU9ocdXmC9jRbD39J1zb3Q@mail.gmail.com>
elton sky <eltonsky9404@gmail.com> writes:
Show 22 quoted lines
> On Mon, Mar 26, 2012 at 12:06 PM, Nguyen Thai Ngoc Duy > <pclouds@gmail.com> wrote: >> (I think this should be on git@vger as there are many experienced devs there) >> >> On Sun, Mar 25, 2012 at 11:13 AM, elton sky <eltonsky9404@gmail.com> wrote: >>> About the new format: >>> >>> The index is a single file. Entries in the index still stored >>> sequentially as old format. The difference is they are grouped into >>> blocks. A block contains many entries and they are ordered by names. >>> Blocks are also ordered by the name of the first entry. Each block >>> contains a sha1 for entries in it. >> >> If I remove an entry in the first block, because blocks are of fixed >> size, you would need to shift all entries up by one, thus update all >> blocks? > > We need some GC here. I am not moving all blocks. Rather I would > consider merge or recycle the block. In a simple case if a block > becomes empty, I ll change the offset of new block in the header point > to this block, and make this block points to the original offset of > new block. In this way, I keep the list of empty blocks I can reuse.
[...]
Doesn't that venture into database land?
If we go that far, wouldn't it be better to use a proper database library? All other things being equal, writing such complex code from scratch is probably not a good idea.
--
Thomas Rast
trast@{inf,student}.ethz.ch