git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] cherry: cache patch-ids to avoid repeating work

From
Geoffrey Irving <irving@naml.us>
Date
Jul 11, 2008, 14:58 UTC
Message-ID
<7f9d599f0807110758y6c4ea7bepd726daf4fe5f074c@mail.gmail.com>
In-Reply-To
<7v3amglxmb.fsf@gitster.siamese.dyndns.org>
On Thu, Jul 10, 2008 at 11:54 PM, Junio C Hamano <gitster@pobox.com> wrote:
Show 24 quoted lines
> "Geoffrey Irving" <irving@naml.us> writes:
>
>>>> Oops: avoiding the infinite loop only requires reading expected O(1)
>>>> entries on load, so I can fix that if you like.  It would only be all of
>>>> them if it actually did detect the infinite loop.
>>>
>>> I have to admit that you lost me there.  AFAIR the patch-id cache is a
>>> simple commit->patch_id store, right?  Then there should be no way to get
>>> an infinite loop.
>>
>> If every entry is nonnull, find_helper loops forever.
>
> Isn't it sufficient to make this part check the condition as well?
>
> +       if (cache->count >= cache->size)
> +       {
> +               warning("%s is corrupt: count %"PRIu32" >= size %"PRIu32,
> +                       filename, cache->count, cache->size);
> +               goto empty;
> +       }
>
> At runtime you keep the invariants that hashtable always has at most 3/4
> full and whoever wrote the file you are reading must have honored that as
> well, or there is something fishy going on.

Good point. There's no reason not to check the 3/4 condition. It isn't sufficient to avoid the infinite loop, though, since we don't verify that count is accurate.

Another route would to eliminate the count field entirely, and replace the count >= size/4*3 check with a statistical one based on the entries seen so far. The main advantage of that would be to make incremental writes simpler by avoiding the need to update the header. The disadvantage is that there would be a small chance that the map would grow in size despite being almost empty. Thoughts on whether we should do that?

Show 9 quoted lines
>>> Besides, this is a purely local cache, no?  Never to be transmitted...  So
>>> not much chance of a malicious attack, except if you allow write access to
>>> your local repository, in which case you are endangered no matter what.
>>
>> Yep, that's why it's only a hole in quotes, and why I didn't fix it.
>
> You might want to protect yourself against file corruption, though.
> Checksumming the whole file and checking it at opening defeats the point
> of mmap'ed access, but at least the header may want to be checksummed?
Okay.  I imagine I should use sha1 as the sum?
Geoffrey
Previous: Junio C HamanoNext: Johannes Schindelin
Message 12 of 21 in “cherry: cache patch-ids to avoid repeating work”
  1. 1/3 cherry: cache patch-ids to avoid repeating workGeoffrey Irving, Jul 9, 2008
  2. Junio C HamanoJul 9, 2008
  3. Geoffrey IrvingJul 9, 2008
  4. Junio C HamanoJul 9, 2008
  5. Johannes SchindelinJul 9, 2008
  6. cherry: cache patch-ids to avoid repeating workGeoffrey Irving, Jul 10, 2008
  7. Geoffrey IrvingJul 10, 2008
  8. Johannes SchindelinJul 10, 2008
  9. Geoffrey IrvingJul 10, 2008
  10. Johannes SchindelinJul 10, 2008
  11. Junio C HamanoJul 11, 2008
  12. Geoffrey IrvingJul 11, 2008
  13. Johannes SchindelinJul 11, 2008
  14. Geoffrey IrvingJul 11, 2008
  15. Johannes SchindelinJul 11, 2008
  16. Geoffrey IrvingJul 13, 2008
  17. cherry: cache patch-ids to avoid repeating workGeoffrey Irving, Jul 15, 2008
  18. Johannes SchindelinJul 15, 2008
  19. Junio C HamanoJul 15, 2008
  20. Karl HasselströmJul 16, 2008
  21. Johan HerlandJul 16, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.