git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Fw: Curiosity

From
JBJoão Victor Bonfim <joaovictorbonfim@protonmail.com>
Date
Dec 18, 2021, 00:17 UTC
Message-ID
<qzxpLxxzy2ooNpnphGZ_IjuF0yj-39e_CR6OiXgthFAz2VR_OKA2HzyY2zznYIv4DyZZFfrBiMa9M1eR_Qwj8iJHZzMBd_QsEoPIYHMEwuo=@protonmail.com>
In-Reply-To
<xmqqtuf88bw6.fsf@gitster.g>
> That is probably too application specific to be in core-git, but it
Application specific as in that it is too much of an edge case to be used by all git users?
> is probably a good application for smudge/clean filters like brian
>
> alluded to?
Perhaps.
‐‐‐‐‐‐‐ Original Message ‐‐‐‐‐‐‐
Em quinta-feira, 16 de dezembro de 2021 às 18:42, Junio C Hamano <gitster@pobox.com> escreveu:
Show 59 quoted lines
> Martin Fick mfick@codeaurora.org writes:
>
> > On 2021-12-16 14:20, João Victor Bonfim wrote:
> >
> > > > To expand on this, if what you're storing is already compressed, like
> > > >
> > > > Ogg Vorbis files or PNGs, like are found in that repository, then
> > > >
> > > > generally they will not delta well. This is also true of things like
> > > >
> > > > Microsoft Office or OpenOffice documents, because they're essentially
> > > >
> > > > Zip files.
> > > >
> > > > The delta algorithm looks for similarities between files to
> > > >
> > > > compress
> > > >
> > > > them. If a file is already compressed using something like Deflate,
> > > >
> > > > used in PNGs and Zip files, then even very similar files will
> > > >
> > > > generally
> > > >
> > > > look very different, so deltification will generally be ineffective.
> > > >
> > > > ...
> > > >
> > > > Maybe I am thinking too outside the box, but wouldn't it be quite more
> > > >
> > > > effective for git to identify compressed files, specially on edge cases
> > > >
> > > > where the compression doesn't have a good chemistry with delta
> > > >
> > > > compression,
> > > >
> > > > decompress them for repo storage while also storing the compression
> > > >
> > > > algorithm as some metadata tag (like a text string or an ID code
> > > >
> > > > decided
> > > >
> > > > beforehand), and, when creating the work mirrors, return the
> > > >
> > > > compression
> > > >
> > > > to its default state before checkout?
> >
> > I suspect that for most algorithms and their implementations, this would
> >
> > not result in repeatable "recompressed" results. Thus the checked-out
> >
> > files might be different every time you checked them out. :(
>
> That is probably too application specific to be in core-git, but it
>
> is probably a good application for smudge/clean filters like brian
>
> alluded to?
Previous: Junio C HamanoNext: João Victor Bonfim
Message 8 of 14 in “Fw: Curiosity”
  1. João Victor BonfimDec 15, 2021
  2. Junio C HamanoDec 15, 2021
  3. João Victor BonfimDec 15, 2021
  4. brian m. carlsonDec 16, 2021
  5. João Victor BonfimDec 16, 2021
  6. Martin FickDec 16, 2021
  7. Junio C HamanoDec 16, 2021
  8. João Victor BonfimDec 18, 2021
  9. João Victor BonfimDec 18, 2021
  10. Junio C HamanoDec 18, 2021
  11. João Victor BonfimDec 18, 2021
  12. Martin FickDec 18, 2021
  13. brian m. carlsonDec 18, 2021
  14. João Victor BonfimDec 18, 2021

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.