git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: reftable [v4]: new ref storage format

From
Dave Borowitz <dborowitz@google.com>
Date
Aug 2, 2017, 12:20 UTC
Message-ID
<CAD0k6qRTa6jSgEBBX1Ux5yg4QMMWPpyOGTa471cRhtzBaS-KjQ@mail.gmail.com>
In-Reply-To
<CAJo=hJv=zJvbzfAZwspxECXrnBJR4XfJbGZegsNUCx=6uheO2Q@mail.gmail.com>
On Tue, Aug 1, 2017 at 10:38 PM, Shawn Pearce <spearce@spearce.org> wrote:
Show 23 quoted lines
>> Peff and I discussed off-list whether the lookup-by-SHA-1 feature is
>> so important in the first place. Currently, all references must be
>> scanned for the advertisement anyway,
>
> Not really. You can hide refs and allow-tip-sha1 so clients can fetch
> a ref even if it wasn't in the advertisement. We really want to use
> that wire protocol capability with Gerrit Code Review to hide the
> refs/changes/ namespace from the advertisement, but allow clients to
> fetch any of those refs if they send its current SHA-1 in a want line
> anyway.
>
> So a server could scan only the refs/{heads,tags}/ prefixes for the
> advertisement, and then leverage the lookup-by-SHA1 to verify other
> SHA-1s sent by the client.
>
>> so avoiding a second scan to vet
>> SHA-1s received from the client is at best going to reduce the effort
>> by a constant factor. Do you have numbers showing that this
>> optimization is worth it?
>
> No, but I don't think I need to do much to prove it. My 866k ref
> example advertisement right now is >62 MiB. If we do what I'm
> suggesting in the paragraphs above, the advertisement is ~51 KiB.

That being said, our bias towards minimizing the number of ref scans is rooted in our experience where scanning 866k refs takes 5 seconds to get the response from the storage backend into the git server. Cutting ref scans from 2 to 1 (or 1 to 0) is a big deal in that case. But that 5s number is based on our current, slow storage, not on reftable. If migrating to reftable turns each 5s scan into a 400ms scan, we might be able to live with that, even if we don't have fast lookup by SHA-1.

>> OTOH a mythical protocol v2 might reduce the need to scan the
>> references for advertisement, so maybe this optimization will be more
>> helpful in the future?

I haven't been following the status of the proposal, but I was assuming a client-speaks-first protocol would also imply the client asking for refnames, not SHA-1s, in which case lookup by SHA-1 is no longer relevant.

Previous: Jeff KingNext: Jeff King
Message 19 of 34 in “Re: reftable [v4]: new ref storage format”
  1. Shawn PearceJul 31, 2017
  2. Dave BorowitzJul 31, 2017
  3. Stefan BellerJul 31, 2017
  4. Shawn PearceJul 31, 2017
  5. Junio C HamanoJul 31, 2017
  6. Shawn PearceJul 31, 2017
  7. Shawn PearceAug 1, 2017
  8. Michael HaggertyAug 1, 2017
  9. Shawn PearceAug 1, 2017
  10. Michael HaggertyAug 2, 2017
  11. Shawn PearceAug 1, 2017
  12. Shawn PearceAug 1, 2017
  13. Michael HaggertyAug 2, 2017
  14. Shawn PearceAug 2, 2017
  15. Jeff KingAug 2, 2017
  16. Shawn PearceAug 2, 2017
  17. Junio C HamanoAug 2, 2017
  18. Jeff KingAug 2, 2017
  19. Dave BorowitzAug 2, 2017
  20. Jeff KingAug 2, 2017
  21. Michael HaggertyAug 3, 2017
  22. Shawn PearceAug 3, 2017
  23. Michael HaggertyAug 3, 2017
  24. Shawn PearceAug 4, 2017
  25. Shawn PearceAug 5, 2017
  26. Dave BorowitzAug 1, 2017
  27. Shawn PearceAug 1, 2017
  28. Junio C HamanoAug 2, 2017
  29. Jeff KingAug 2, 2017
  30. Shawn PearceAug 3, 2017
  31. Junio C HamanoAug 3, 2017
  32. Shawn PearceAug 3, 2017
  33. Junio C HamanoAug 3, 2017
  34. Stefan BellerAug 2, 2017

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.