git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git-fetching from a big repository is slow

From
GBGeert Bosch <bosch@adacore.com>
Date
Dec 14, 2006, 23:01 UTC
Message-ID
<E30DCF6F-5D3E-4CA3-85D7-CD2847B86F86@adacore.com>
In-Reply-To
<20061214194636.GO1747@spearce.org>
On Dec 14, 2006, at 14:46, Shawn Pearce wrote:
> And yet I get good delta compression on a number of ZIP formatted
> files which don't get good additional zlib compression (<3%).
> Doing the above would cause those packfiles to explode to about
> 10x their current size.

Yes, that's because for zip files each file in the archive is compressed independently. Similar things might happen when checking in uncompressed tar files with JPG's. The question is whether you prefer bad time usage or bad space usage when handling large binary blobs. Maybe we should use a faster, less precise algorithm instead of giving up.

Still, I think doing anything based on filename is a mistake. If we want to have a heuristic to prevent spending too much time on deltifying large compressed files, the heuristic should be based on content, not filename.

Maybe we could some "magic" as used by the file(1) command that allows git to say a bit more about the content of blobs. This could be used both for ordering files during deltification and to determine wether to try deltification at all.

   -Geert
Previous: Shawn PearceNext: Johannes Schindelin
Message 2 of 11 in “Re: git-fetching from a big repository is slow”
  1. Shawn PearceDec 14, 2006
  2. Geert BoschDec 14, 2006
  3. Johannes SchindelinDec 14, 2006
  4. Shawn PearceDec 14, 2006
  5. Johannes SchindelinDec 15, 2006
  6. Shawn PearceDec 15, 2006
  7. Nicolas PitreDec 15, 2006
  8. Horst H. von BrandDec 14, 2006
  9. Shawn PearceDec 14, 2006
  10. PazuDec 15, 2006
  11. Robin RosenbergDec 16, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.