git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [Gnu-arch-users] Re: [GNU-arch-dev] [ANNOUNCEMENT] /Arch/ embraces `git'

From
TMTomas Mraz <t8m@centrum.cz>
Date
Apr 21, 2005, 10:21 UTC
Message-ID
<1114078877.5886.37.camel@perun.redhat.usu>
In-Reply-To
<86d5soa42h.fsf@speedy.lifl.fr>
On Thu, 2005-04-21 at 11:09 +0200, Denys Duchier wrote:
Show 16 quoted lines
> Tomas Mraz <t8m@centrum.cz> writes:
> 
> > If we suppose the maximum number of stored blobs in the order of milions
> > probably the optimal indexing would be 1 level [0:2] indexing or 2
> > levels [0:1] [2:3]. However it would be necessary to do some
> > benchmarking first before setting this to stone.
> 
> As I have suggested in a previous message, it is trivial to implement adaptive
> indexing: there is no need to hardwire a specific indexing scheme.  Furthermore,
> I suspect that the optimal size of subkeys may well depend on the filesystem.
> My experiments seem to indicate that subkeys of length 2 achieve an excellent
> compromise between discriminatory power and disk footprint on ext2.
> 
> Btw, if, as you indicate above, you do believe that a 1 level indexing should
> use [0:2], then it doesn't make much sense to me to also suggest that a 2 level
> indexing should use [0:1] as primary subkey :-)

Why do you think so? IMHO we should always target a similar number of files/subdirectories in a directories of the blob archive. So If I always suppose that the archive would contain at most 16 millions of files then the possible indexing schemes are either 1 level with key length 3 (each directory would contain ~4096 files) or 2 level with key length 2 (each directory would contain ~256 files). Which one is better could be of course filesystem and hardware dependent.

Of course it might be best to allow adaptive indexing but I think that first some benchmarking should be made and it's possible that some fixed scheme could be chosen as optimal.

-- 
Tomas Mraz <t8m@centrum.cz>
Previous: Denys DuchierNext: duchier@ps.uni-sb.de
Message 6 of 23 in “[ANNOUNCEMENT] /Arch/ embraces `git'”
  1. Tom LordApr 20, 2005
  2. Miles BaderApr 20, 2005
  3. duchier@ps.uni-sb.deApr 20, 2005
  4. Tomas MrazApr 20, 2005
  5. Denys DuchierApr 21, 2005
  6. Tomas MrazApr 21, 2005
  7. duchier@ps.uni-sb.deApr 21, 2005
  8. Tomas MrazApr 20, 2005
  9. Tom LordApr 21, 2005
  10. Tom LordApr 21, 2005
  11. Tom LordApr 20, 2005
  12. Denys DuchierApr 21, 2005
  13. Tom LordApr 21, 2005
  14. Tomas MrazApr 21, 2005
  15. Tom LordApr 21, 2005
  16. Tom LordApr 21, 2005
  17. Linus TorvaldsApr 22, 2005
  18. Edésio Costa e SilvaApr 22, 2005
  19. Petr BaudisApr 20, 2005
  20. C. Scott AnanianApr 20, 2005
  21. chunking (Re: [ANNOUNCEMENT] /Arch/ embraces `git')Linus Torvalds, Apr 20, 2005
  22. C. Scott AnanianApr 20, 2005
  23. blowing chunks (quick update)C. Scott Ananian, Apr 22, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.