git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: WARNING! Object DB conversion (was Re: [PATCH] write-tree performance problems)

From
C. Scott Ananian <cscott@cscott.net>
Date
Apr 20, 2005, 15:28 UTC
Message-ID
<Pine.LNX.4.61.0504201121490.2630@cag.csail.mit.edu>
In-Reply-To
<20050420151902.GA13175@macavity>
On Wed, 20 Apr 2005, Martin Uecker wrote:
Show 14 quoted lines
>> You can (and my code demonstrates/will demonstrate) still use a whole-file
>> hash to use chunking.  With content prefixes, this takes O(N ln M) time
>> (where N is the file size and M is the number of chunks) to compute all
>> hashes; if subtrees can share the same prefix, then you can do this in
>> O(N) time (ie, as fast as possible, modulo a constant factor, which is
>> '2').  You don't *need* internal hashing functions.
>
> I don't understand this paragraph. What is an internal
> hash function? Your code seems to do exactly what I want.
> The hashes are computed recusively as in a hash tree
> with O(N ln N). The only difference between your design
> and a design based on a conventional (binary) hash tree
> seems to be that data is stored in the intermediate nodes
> too.

A merkle-tree (which I think you initially pointed me at) makes the hash of the internal nodes be a hash of the chunk's hashes; ie not a straight content hash. This is roughly what my current implementation does, but I would like to identify each subtree with the hash of the *(expanded) contents of that subtree* (ie no explicit reference to subtree hashes). This makes it interoperable with non-chunked or differently-chunked representations, in that the top-level hash is *just the hash of the complete content*, not some hash-of-subtree-hashes. Does that make more sense?

The code I posted doesn't demonstrate this very well, but now that Linus has abandoned the 'hash of compressed content' stuff, my next code posting should show this more clearly.

Show 5 quoted lines
> If I don't miss anything essential, you can validate
> each treap piece at the moment you get it from the
> network with its SHA1 hash and then proceed with
> downloading the prefix and suffix tree (in parallel
> if you have more than one peer a la bittorrent).
Yes, I guess this is the detail I was going to abandon. =)
I viewed the fact that the top-level hash was dependent on the exact chunk 
makeup a 'misfeature', because it doesn't allow easy interoperability with 
existing non-chunked repos.
  --scott
WTO atomic operation Mossad Castro overthrow FSF fissionable HTAUTOMAT 
LCPANES MKDELTA Bush non-violent protest OVER THE HORIZON RADAR KUPALM
                          ( http://cscott.net/ )
Previous: Martin UeckerNext: Martin Uecker
Message 19 of 54 in “write-tree performance problems”
  1. write-tree performance problemsChris Mason, Apr 19, 2005
  2. Linus TorvaldsApr 19, 2005
  3. Chris MasonApr 19, 2005
  4. Linus TorvaldsApr 19, 2005
  5. Chris MasonApr 19, 2005
  6. Linus TorvaldsApr 19, 2005
  7. Chris MasonApr 20, 2005
  8. Linus TorvaldsApr 20, 2005
  9. Linus TorvaldsApr 20, 2005
  10. H. Peter AnvinApr 20, 2005
  11. WARNING! Object DB conversion (was Re: [PATCH] write-tree performance problems)Linus Torvalds, Apr 20, 2005
  12. Ingo MolnarApr 20, 2005
  13. Jon SeymourApr 20, 2005
  14. Martin UeckerApr 20, 2005
  15. Morten WelinderApr 20, 2005
  16. Jon SeymourApr 20, 2005
  17. C. Scott AnanianApr 20, 2005
  18. Martin UeckerApr 20, 2005
  19. C. Scott AnanianApr 20, 2005
  20. Martin UeckerApr 20, 2005
  21. Martin UeckerApr 20, 2005
  22. Blob chunking code. [First look.]C. Scott Ananian, Apr 20, 2005
  23. Blob chunking code. [Second look]C. Scott Ananian, Apr 20, 2005
  24. David WoodhouseApr 20, 2005
  25. Linus TorvaldsApr 20, 2005
  26. David WoodhouseApr 20, 2005
  27. Chris MasonApr 20, 2005
  28. C. Scott AnanianApr 20, 2005
  29. Linus TorvaldsApr 20, 2005
  30. C. Scott AnanianApr 20, 2005
  31. Linus TorvaldsApr 20, 2005
  32. Linus TorvaldsApr 20, 2005
  33. David WillmoreApr 20, 2005
  34. Linus TorvaldsApr 20, 2005
  35. Linus TorvaldsApr 20, 2005
  36. Chris MasonApr 20, 2005
  37. Linus TorvaldsApr 20, 2005
  38. Chris MasonApr 20, 2005
  39. Linus TorvaldsApr 20, 2005
  40. Chris MasonApr 20, 2005
  41. Linus TorvaldsApr 20, 2005
  42. Linus TorvaldsApr 20, 2005
  43. David S. MillerApr 20, 2005
  44. David LangApr 19, 2005
  45. Linus TorvaldsApr 19, 2005
  46. David LangApr 19, 2005
  47. Linus TorvaldsApr 19, 2005
  48. David LangApr 19, 2005
  49. Linus TorvaldsApr 19, 2005
  50. Christopher LiApr 19, 2005
  51. Olivier GalibertApr 19, 2005
  52. C. Scott AnanianApr 19, 2005
  53. Linus TorvaldsApr 20, 2005
  54. C. Scott AnanianApr 20, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.