git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v5 4/4] convert: Stream from fd to required clean filter instead of mmap

From
Jeff King <peff@peff.net>
Date
Aug 26, 2014, 18:00 UTC
Message-ID
<20140826180018.GB17546@peff.net>
In-Reply-To
<xmqq4mx0mn7i.fsf@gitster.dls.corp.google.com>
On Mon, Aug 25, 2014 at 11:35:45AM -0700, Junio C Hamano wrote:
Show 15 quoted lines
> Steffen Prohaska <prohaska@zib.de> writes:
> 
> >> Couldn't we do that with an lseek (or even an mmap with offset 0)? That
> >> obviously would not work for non-file inputs, but I think we address
> >> that already in index_fd: we push non-seekable things off to index_pipe,
> >> where we spool them to memory.
> >
> > It could be handled that way, but we would be back to the original problem
> > that 32-bit git fails for large files.
> 
> Correct, and you are making an incremental improvement so that such
> a large blob can be handled _when_ the filters can successfully
> munge it back and forth.  If we fail due to out of memory when the
> filters cannot, that would be the same as without your improvement,
> so you are still making progress.

I do not think my proposal makes anything worse than Steffen's patch. _If_ you have a non-required filter, and _if_ we can run it, then we stream the filter and hopefully end up with a small enough result to fit into memory. If we cannot run the filter, we are screwed anyway (we follow the regular code path and dump the whole thing into memory; i.e., the same as without this patch series).

I think the main argument against going further is just that it is not worth the complexity. Tell people doing reduction filters they need to use "required", and that accomplishes the same thing.

Show 11 quoted lines
> >> So it seems like the ideal strategy would be:
> >> 
> >>  1. If it's seekable, try streaming. If not, fall back to lseek/mmap.
> >> 
> >>  2. If it's not seekable and the filter is required, try streaming. We
> >>     die anyway if we fail.
> 
> Puzzled...  Is it assumed that any content the filters tell us to
> use the contents from the db as-is by exiting with non-zero status
> will always be large not to fit in-core?  For small contents, isn't
> this "ideal" strategy a regression?

I am not sure what you mean by regression here. We will try to stream more often, but I do not see that as a bad thing.

-Peff
Previous: Junio C HamanoNext: Junio C Hamano
Message 13 of 15 in “Stream fd to clean filter; GIT_MMAP_LIMIT, GIT_ALLOC_LIMIT with git_parse_ulong()”
  1. 0/4 Stream fd to clean filter; GIT_MMAP_LIMIT, GIT_ALLOC_LIMIT with git_parse_ulong()Steffen Prohaska, Aug 24, 2014
  2. 1/4 convert: Refactor would_convert_to_git() to single arg 'path'Steffen Prohaska, Aug 24, 2014
  3. Junio C HamanoAug 25, 2014
  4. 2/4 Change GIT_ALLOC_LIMIT check to use git_parse_ulong()Steffen Prohaska, Aug 24, 2014
  5. Jeff KingAug 25, 2014
  6. Steffen ProhaskaAug 25, 2014
  7. Jeff KingAug 25, 2014
  8. 3/4 Introduce GIT_MMAP_LIMIT to allow testing expected mmap sizeSteffen Prohaska, Aug 24, 2014
  9. 4/4 convert: Stream from fd to required clean filter instead of mmapSteffen Prohaska, Aug 24, 2014
  10. Jeff KingAug 25, 2014
  11. Steffen ProhaskaAug 25, 2014
  12. Junio C HamanoAug 25, 2014
  13. Jeff KingAug 26, 2014
  14. Junio C HamanoAug 26, 2014
  15. Jeff KingAug 26, 2014

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.