git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: fast-import slowness when importing large files with small differences

From
Ævar Arnfjörð Bjarmason <avarab@gmail.com>
Date
Jul 3, 2018, 16:05 UTC
Message-ID
<87muv8cnk3.fsf@evledraar.gmail.com>
In-Reply-To
<20180629233538.7zxxrvou4twqyd6d@glandium.org>
On Fri, Jun 29 2018, Mike Hommey wrote:
Show 34 quoted lines
> On Sat, Jun 30, 2018 at 12:10:24AM +0200, Ævar Arnfjörð Bjarmason wrote:
>>
>> On Fri, Jun 29 2018, Mike Hommey wrote:
>>
>> > I noticed some slowness when fast-importing data from the Firefox mercurial
>> > repository, where fast-import spends more than 5 minutes importing ~2000
>> > revisions of one particular file. I reduced a testcase while still
>> > using real data. One could synthesize data with kind of the same
>> > properties, but I figured real data could be useful.
>> >
>> > To reproduce:
>> > $ git clone https://gist.github.com/b6b8edcff2005cc482cf84972adfbba9.git foo
>> > $ git init bar
>> > $ cd bar
>> > $ python ../foo/import.py ../foo/data.gz | git fast-import --depth=2000
>> >
>> > [...]
>> > So maybe it would make sense to consolidate the diff code (after all,
>> > diff-delta.c is an old specialized fork of xdiff). With manual trimming
>> > of common head and tail, this gets down to 3:33.
>> >
>> > I'll also note that Facebook has imported xdiff from the git code base
>> > into mercurial and improved performance on it, so it might also be worth
>> > looking at what's worth taking from there.
>>
>> It would be interesting to see how does this compares with a more naïve
>> approach of committing every version of this file one-at-a-time into a
>> new repository (with & without gc.auto=0). Perhaps deltaing as we go is
>> suboptimal compared to just writing out a lot of redundant data and
>> repacking it all at once later.
>
> "Just" writing 26GB? And that's only one file. If I were to do that for
> the whole repository, it would yield a > 100GB pack. Instead of < 2GB
> currently.

To clarify on my terse response. I mean to try this on an isolated test case to see to what extent the problem you're describing is unique to fast-import, and to what extent it's encountered during "normal" git use when you commit all the revisions of that file in succession.

Perhaps the difference between the two would give some hint as to how to proceed, or not.

Previous: Mike HommeyNext: Mike Hommey
Message 16 of 17 in “fast-import slowness when importing large files with small differences”
  1. Mike HommeyJun 29, 2018
  2. Stefan BellerJun 29, 2018
  3. xdiff: reduce indent heuristic overheadStefan Beller, Jun 29, 2018
  4. Junio C HamanoJun 29, 2018
  5. xdiff: reduce indent heuristic overheadStefan Beller, Jun 29, 2018
  6. Jun WuJun 30, 2018
  7. Michael HaggertyJul 1, 2018
  8. Stefan BellerJul 2, 2018
  9. Michael HaggertyJul 3, 2018
  10. xdiff: reduce indent heuristic overheadStefan Beller, Jul 27, 2018
  11. Junio C HamanoJul 3, 2018
  12. Jeff KingJun 29, 2018
  13. Stefan BellerJun 29, 2018
  14. Ævar Arnfjörð BjarmasonJun 29, 2018
  15. Mike HommeyJun 29, 2018
  16. Ævar Arnfjörð BjarmasonJul 3, 2018
  17. Mike HommeyJul 3, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.