{"thread":{"id":"57082","subject":"Fw: Curiosity","startedAt":"2021-12-15T03:52:44Z","lastAt":"2021-12-18T01:40:28Z","messageCount":14,"participants":["João Victor Bonfim","Junio C Hamano","brian m. carlson","Martin Fick"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"444140","messageId":"wVwq9WVLpVt7MNLmIYOWCFKVSf8l532MD_vu4yTA8hl1fCARnW8nOUJjxYmKSzFw1SnPp5iYRD-aW4gAT2HnyQbC5aLBOvyT6npn88lxwNQ=@protonmail.com","threadId":"57082","inReplyTo":"Wlh_w2gSCDQ2ieJnIY7TStWrzxbwP98SNRIFMTYpva7SRFipqk63HEYFVF7wFn1oSHOkQNsjWGOa5L49vyRlvSLbuZqpmvOaDOHmFkdt2zw=@protonmail.com","subject":"Fw: Curiosity","fromName":"João Victor Bonfim","fromEmail":"joaovictorbonfim@protonmail.com","sentAt":"2021-12-15T03:52:38Z","receivedAt":"2021-12-15T03:52:44Z","isPatch":false,"sender":{"key":"joaovictorbonfim@protonmail.com","avatar":null},"body":"I sent this message to Junio Hamano kinda of forever ago, since then I haven't been able to address it or do anything about it really (I am personally making a report on Git for the conclusion of my technician course so I can get my certification, yada yada yada, couldn't get to it). These days I have been reading Junio's responses on the git mailing list archive (https://marc.info/?l=git or rather https://marc.info/?a=118086005800002&r=1&w=4) from May to now to see if Junio said anything. Junio didn't, but I did read https://marc.info/?i=xmqqpmudng5x.fsf%20()%20gitster%20!%20g and kinda of felt that was targeted at me, or people like me at least...\n\n`:-)  - me sweating in exasperation.\n\nAlso since then, I may have improved on my confusing line of thought, so here is the past message and my current version so to speak:\n\n------- Second attempt --------\n\nSince Git is almost used for everything at this point, is there any intent on providing better support for non textual file types? Why do I say this? Take this game mod which I follow as example -> https://github.com/SolariusScorch/XComFiles <- whenever I clone it Git takes a significant forever amount of time to download 452 MB of files whose some part, from my perspective, isn't being delta compressed like the text files are (since, whenever reading a log of what changes were made, git creates and undoes modes for all binary files, some of which only changed by a pixel from one colour to another).\n\nFrom my perspective it would be interesting to enhance the effectiveness/performance of git for such files, since some projects are very heavy on multimedia that isn't hard coded and those will eventually come around to using git. From a personal perspective: I pretend to create an open source game and track it with git, however it concerns me whether or not it might take forever for users to clone the repo once a few versions of a singular file of, perhaps, some Gigabytes in size aren't stored and compressed efficiently and instead all the versions are stored in full, totalling some Terabytes in storage for a few of such files.\n\n‐‐‐‐‐‐‐ Original Message ‐‐‐‐‐‐‐\nWednesday, 27, May, 2021, 22:12, João Victor Bonfim <JoaoVictorBonfim@protonmail.com> wrote:\n\n> I am assuming you are the Git maintainer, therefore the message, otherwise, forgive me.\n> Considering the ubiquity of Git as a versioning system and my internal queries about the future of software development, specially game development, is there any intent on providing support for non textual file types? What do I mean is that binary files, from my perspective as a user, are tracked in full rather than partially, which I mean is that the files are discarded and replaced if they are altered when, instead, they could have the differentiation between files tracked. Of course this would require several changes to Git so it can interpret images and so on, but I think that it could be good for software development that requires extensive multimedia use and, therefore, may require that better tracking for such material is made available.\n>\n> Do you understand where I want to get to?\n>\n> Graciously yours, João Victor Bonfim.\n"},{"id":"444156","messageId":"xmqq8rwl91yf.fsf@gitster.g","threadId":"57082","inReplyTo":"wVwq9WVLpVt7MNLmIYOWCFKVSf8l532MD_vu4yTA8hl1fCARnW8nOUJjxYmKSzFw1SnPp5iYRD-aW4gAT2HnyQbC5aLBOvyT6npn88lxwNQ=@protonmail.com","subject":"Re: Fw: Curiosity","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-12-15T18:07:20Z","receivedAt":"2021-12-15T18:07:27Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"João Victor Bonfim  <JoaoVictorBonfim@protonmail.com> writes:\n\n> I sent this message to Junio Hamano kinda of forever ago, since\n> then I haven't been able to address it or do anything about it\n> really...\n\nMy spam filter has learned that anything that goes to gitster@\naddress without cc'ed to the git@vger list are to be caught, so it\nis very plausible that I didn't see it.  Sending any inquiry here on\nthe list is the right thing to do, especially because it is likely\nthat I may not be the area expert for whatever you want to learn\nabout Git, while there are others who are more familiar with various\nparts of the system and other ways the system is used.\n\nYou will also increase your chances to be read if you made your\nmessage look more like the ones typically posted here (see the\narchive), by wrapping overly long lines, etc.\n\n> Since Git is almost used for everything at this point, is there\n> any intent on providing better support for non textual file types?\n> Why do I say this? Take this game mod which I follow as example ->\n> https://github.com/SolariusScorch/XComFiles <- whenever I clone it\n> Git takes a significant forever amount of time to download 452 MB\n> of files whose some part, from my perspective, isn't being delta\n> compressed like the text files are (since, whenever reading a log\n> of what changes were made, git creates and undoes modes for all\n> binary files, some of which only changed by a pixel from one\n> colour to another).\n\nOur delta compression does not care whether the contents are text or\nbinary, so if it is not compressed well, so it can be a sign that\nthe contents are not compressible to begin with, at least with the\nxdelta binary-diff-patch engine we use.  Improvement designs,\nalgorithms and patches are always welcome ;-)\n\n\n"},{"id":"444183","messageId":"hYaujy-hcnAn3EcYnkfxn2Iorjz5gjD3_0c1DZf_AztxWhLgmKSdLxuv3rrFKZOpRRGpdc8xAAA-ym4OrYg5iJ8STRvdChftNVYRzdYz4QY=@protonmail.com","threadId":"57082","inReplyTo":"xmqq8rwl91yf.fsf@gitster.g","subject":"Re: Fw: Curiosity","fromName":"João Victor Bonfim","fromEmail":"joaovictorbonfim@protonmail.com","sentAt":"2021-12-15T23:45:35Z","receivedAt":"2021-12-15T23:45:38Z","isPatch":false,"sender":{"key":"joaovictorbonfim@protonmail.com","avatar":null},"body":"> João Victor Bonfim JoaoVictorBonfim@protonmail.com writes:\n>\n> Our delta compression does not care whether the contents are text or\n>\n> binary, so if it is not compressed well, so it can be a sign that\n>\n> the contents are not compressible to begin with, at least with the\n>\n> xdelta binary-diff-patch engine we use. Improvement designs,\n>\n> algorithms and patches are always welcome ;-)\n\nGosh, I wish I could do anything about it.\n\nI am but a mere code monkey, haven't done much writing practice either.\n\nMaybe one day, but that is yet to be seen.\n"},{"id":"444195","messageId":"YbqiQ1B9ezF/RPOn@camp.crustytoothpaste.net","threadId":"57082","inReplyTo":"xmqq8rwl91yf.fsf@gitster.g","subject":"Re: Fw: Curiosity","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2021-12-16T02:19:47Z","receivedAt":"2021-12-16T02:19:53Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2021-12-15 at 18:07:20, Junio C Hamano wrote:\n> João Victor Bonfim  <JoaoVictorBonfim@protonmail.com> writes:\n> > Since Git is almost used for everything at this point, is there\n> > any intent on providing better support for non textual file types?\n> > Why do I say this? Take this game mod which I follow as example ->\n> > https://github.com/SolariusScorch/XComFiles <- whenever I clone it\n> > Git takes a significant forever amount of time to download 452 MB\n> > of files whose some part, from my perspective, isn't being delta\n> > compressed like the text files are (since, whenever reading a log\n> > of what changes were made, git creates and undoes modes for all\n> > binary files, some of which only changed by a pixel from one\n> > colour to another).\n> \n> Our delta compression does not care whether the contents are text or\n> binary, so if it is not compressed well, so it can be a sign that\n> the contents are not compressible to begin with, at least with the\n> xdelta binary-diff-patch engine we use.  Improvement designs,\n> algorithms and patches are always welcome ;-)\n\nTo expand on this, if what you're storing is already compressed, like\nOgg Vorbis files or PNGs, like are found in that repository, then\ngenerally they will not delta well.  This is also true of things like\nMicrosoft Office or OpenOffice documents, because they're essentially\nZip files.\n\nThe delta algorithm looks for similarities between files to compress\nthem.  If a file is already compressed using something like Deflate,\nused in PNGs and Zip files, then even very similar files will generally\nlook very different, so deltification will generally be ineffective.\n\nThere are two main solutions to this.  One is to store your data\nuncompressed in the repository and compress it as part of a build step.\nThis makes your checkouts larger, but it makes your repository smaller.\n\nThe other is to store them outside of the repository proper.  Some folks\nuse Git LFS for this, but you could also just store a manifest with file\nnames and secure hashes, plus a download location for a public server.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"444257","messageId":"xndBIO9EtrXaA932eF-0YkvHCAOL1GOKQQlIigssmcwhtZWqGxhc6I_A-lXt7vMK-j1oDrQMHUIuExlpqFS4v88nWci32qx3W5Xi1_hPpUM=@protonmail.com","threadId":"57082","inReplyTo":"YbqiQ1B9ezF/RPOn@camp.crustytoothpaste.net","subject":"Re: Fw: Curiosity","fromName":"João Victor Bonfim","fromEmail":"joaovictorbonfim@protonmail.com","sentAt":"2021-12-16T21:20:45Z","receivedAt":"2021-12-16T21:20:52Z","isPatch":false,"sender":{"key":"joaovictorbonfim@protonmail.com","avatar":null},"body":"> To expand on this, if what you're storing is already compressed, like\n> Ogg Vorbis files or PNGs, like are found in that repository, then\n> generally they will not delta well. This is also true of things like\n> Microsoft Office or OpenOffice documents, because they're essentially\n> Zip files.\n>\n> The delta algorithm looks for similarities between files to compress\n> them. If a file is already compressed using something like Deflate,\n> used in PNGs and Zip files, then even very similar files will generally\n> look very different, so deltification will generally be ineffective.\n\nThis explain why, also, Git opens a new mode every time an edit is made,\nsince it cannot recognize any similarities between the files, even\nthough there are.\n\n> There are two main solutions to this. One is to store your data\n> uncompressed in the repository and compress it as part of a build step.\n> This makes your checkouts larger, but it makes your repository smaller.\n>\n> The other is to store them outside of the repository proper. Some folks\n> use Git LFS for this, but you could also just store a manifest with file\n> names and secure hashes, plus a download location for a public server.\n\nMaybe I am thinking too outside the box, but wouldn't it be quite more\neffective for git to identify compressed files, specially on edge cases\nwhere the compression doesn't have a good chemistry with delta compression,\ndecompress them for repo storage while also storing the compression\nalgorithm as some metadata tag (like a text string or an ID code decided\nbeforehand), and, when creating the work mirrors, return the compression\nto its default state before checkout?\n\nOf course you would also need reversing functions when you want to\ncheckout the info back to repo.\n\nJust throwing ideas out there.\n\n-------------------------------\n\nJoão Victor Bonfim, any pronouns are welcome.\n\n‐‐‐‐‐‐‐Original Message ‐‐‐‐‐‐‐\n\nEm quarta-feira, 15 de dezembro de 2021 às 23:19, brian m. carlson <sandals@crustytoothpaste.net> escreveu:\n\n> On 2021-12-15 at 18:07:20, Junio C Hamano wrote:\n>\n> > João Victor Bonfim JoaoVictorBonfim@protonmail.com writes:\n> >\n> > > Since Git is almost used for everything at this point, is there\n> > >\n> > > any intent on providing better support for non textual file types?\n> > >\n> > > Why do I say this? Take this game mod which I follow as example ->\n> > >\n> > > https://github.com/SolariusScorch/XComFiles <- whenever I clone it\n> > >\n> > > Git takes a significant forever amount of time to download 452 MB\n> > >\n> > > of files whose some part, from my perspective, isn't being delta\n> > >\n> > > compressed like the text files are (since, whenever reading a log\n> > >\n> > > of what changes were made, git creates and undoes modes for all\n> > >\n> > > binary files, some of which only changed by a pixel from one\n> > >\n> > > colour to another).\n> >\n> > Our delta compression does not care whether the contents are text or\n> >\n> > binary, so if it is not compressed well, so it can be a sign that\n> >\n> > the contents are not compressible to begin with, at least with the\n> >\n> > xdelta binary-diff-patch engine we use. Improvement designs,\n> >\n> > algorithms and patches are always welcome ;-)\n>\n> To expand on this, if what you're storing is already compressed, like\n>\n> Ogg Vorbis files or PNGs, like are found in that repository, then\n>\n> generally they will not delta well. This is also true of things like\n>\n> Microsoft Office or OpenOffice documents, because they're essentially\n>\n> Zip files.\n>\n> The delta algorithm looks for similarities between files to compress\n>\n> them. If a file is already compressed using something like Deflate,\n>\n> used in PNGs and Zip files, then even very similar files will generally\n>\n> look very different, so deltification will generally be ineffective.\n>\n> There are two main solutions to this. One is to store your data\n>\n> uncompressed in the repository and compress it as part of a build step.\n>\n> This makes your checkouts larger, but it makes your repository smaller.\n>\n> The other is to store them outside of the repository proper. Some folks\n>\n> use Git LFS for this, but you could also just store a manifest with file\n>\n> names and secure hashes, plus a download location for a public server.\n> -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------\n>\n> brian m. carlson (he/him or they/them)\n>\n> Toronto, Ontario, CA\n"},{"id":"444262","messageId":"54fe7ba20109f974b61a7e6c24ba8264@codeaurora.org","threadId":"57082","inReplyTo":"xndBIO9EtrXaA932eF-0YkvHCAOL1GOKQQlIigssmcwhtZWqGxhc6I_A-lXt7vMK-j1oDrQMHUIuExlpqFS4v88nWci32qx3W5Xi1_hPpUM=@protonmail.com","subject":"Re: Fw: Curiosity","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2021-12-16T21:33:28Z","receivedAt":"2021-12-16T21:33:32Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On 2021-12-16 14:20, João Victor Bonfim wrote:\n>> To expand on this, if what you're storing is already compressed, like\n>> Ogg Vorbis files or PNGs, like are found in that repository, then\n>> generally they will not delta well. This is also true of things like\n>> Microsoft Office or OpenOffice documents, because they're essentially\n>> Zip files.\n>> \n>> The delta algorithm looks for similarities between files to compress\n>> them. If a file is already compressed using something like Deflate,\n>> used in PNGs and Zip files, then even very similar files will \n>> generally\n>> look very different, so deltification will generally be ineffective.\n...\n> Maybe I am thinking too outside the box, but wouldn't it be quite more\n> effective for git to identify compressed files, specially on edge cases\n> where the compression doesn't have a good chemistry with delta \n> compression,\n> decompress them for repo storage while also storing the compression\n> algorithm as some metadata tag (like a text string or an ID code \n> decided\n> beforehand), and, when creating the work mirrors, return the \n> compression\n> to its default state before checkout?\n\nI suspect that for most algorithms and their implementations, this would\nnot result in repeatable \"recompressed\" results. Thus the checked-out\nfiles might be different every time you checked them out. :(\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code\nAurora Forum, hosted by The Linux Foundation\n"},{"id":"444264","messageId":"xmqqtuf88bw6.fsf@gitster.g","threadId":"57082","inReplyTo":"54fe7ba20109f974b61a7e6c24ba8264@codeaurora.org","subject":"Re: Fw: Curiosity","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-12-16T21:42:33Z","receivedAt":"2021-12-16T21:43:59Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Fick <mfick@codeaurora.org> writes:\n\n> On 2021-12-16 14:20, João Victor Bonfim wrote:\n>>> To expand on this, if what you're storing is already compressed, like\n>>> Ogg Vorbis files or PNGs, like are found in that repository, then\n>>> generally they will not delta well. This is also true of things like\n>>> Microsoft Office or OpenOffice documents, because they're essentially\n>>> Zip files.\n>>> The delta algorithm looks for similarities between files to\n>>> compress\n>>> them. If a file is already compressed using something like Deflate,\n>>> used in PNGs and Zip files, then even very similar files will\n>>> generally\n>>> look very different, so deltification will generally be ineffective.\n> ...\n>> Maybe I am thinking too outside the box, but wouldn't it be quite more\n>> effective for git to identify compressed files, specially on edge cases\n>> where the compression doesn't have a good chemistry with delta\n>> compression,\n>> decompress them for repo storage while also storing the compression\n>> algorithm as some metadata tag (like a text string or an ID code\n>> decided\n>> beforehand), and, when creating the work mirrors, return the\n>> compression\n>> to its default state before checkout?\n>\n> I suspect that for most algorithms and their implementations, this would\n> not result in repeatable \"recompressed\" results. Thus the checked-out\n> files might be different every time you checked them out. :(\n\nThat is probably too application specific to be in core-git, but it\nis probably a good application for smudge/clean filters like brian\nalluded to?\n"},{"id":"444406","messageId":"1X3gQ48NK5aBDHcpYMlxESRjqubcCBKJUQu2K0dBOnTyvsXCXXoGDBg2Ff4KarK6WsZnzN3HgqHGOlCKKdF-wtZQ5tHsoAcfit2CTXMWqh4=@protonmail.com","threadId":"57082","inReplyTo":"54fe7ba20109f974b61a7e6c24ba8264@codeaurora.org","subject":"Re: Fw: Curiosity","fromName":"João Victor Bonfim","fromEmail":"joaovictorbonfim@protonmail.com","sentAt":"2021-12-18T00:15:59Z","receivedAt":"2021-12-18T00:16:06Z","isPatch":false,"sender":{"key":"joaovictorbonfim@protonmail.com","avatar":null},"body":"> I suspect that for most algorithms and their implementations, this would\n>\n> not result in repeatable \"recompressed\" results. Thus the checked-out\n>\n> files might be different every time you checked them out. :(\n\nHow or why?\n\nSincere question.\n\n‐‐‐‐‐‐‐ Original Message ‐‐‐‐‐‐‐\n\nEm quinta-feira, 16 de dezembro de 2021 às 18:33, Martin Fick <mfick@codeaurora.org> escreveu:\n\n> On 2021-12-16 14:20, João Victor Bonfim wrote:\n>\n> > > To expand on this, if what you're storing is already compressed, like\n> > >\n> > > Ogg Vorbis files or PNGs, like are found in that repository, then\n> > >\n> > > generally they will not delta well. This is also true of things like\n> > >\n> > > Microsoft Office or OpenOffice documents, because they're essentially\n> > >\n> > > Zip files.\n> > >\n> > > The delta algorithm looks for similarities between files to compress\n> > >\n> > > them. If a file is already compressed using something like Deflate,\n> > >\n> > > used in PNGs and Zip files, then even very similar files will\n> > >\n> > > generally\n> > >\n> > > look very different, so deltification will generally be ineffective.\n>\n> ...\n>\n> > Maybe I am thinking too outside the box, but wouldn't it be quite more\n> >\n> > effective for git to identify compressed files, specially on edge cases\n> >\n> > where the compression doesn't have a good chemistry with delta\n> >\n> > compression,\n> >\n> > decompress them for repo storage while also storing the compression\n> >\n> > algorithm as some metadata tag (like a text string or an ID code\n> >\n> > decided\n> >\n> > beforehand), and, when creating the work mirrors, return the\n> >\n> > compression\n> >\n> > to its default state before checkout?\n>\n> I suspect that for most algorithms and their implementations, this would\n>\n> not result in repeatable \"recompressed\" results. Thus the checked-out\n>\n> files might be different every time you checked them out. :(\n>\n> -Martin\n>\n> -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------\n>\n> The Qualcomm Innovation Center, Inc. is a member of Code\n>\n> Aurora Forum, hosted by The Linux Foundation\n"},{"id":"444407","messageId":"qzxpLxxzy2ooNpnphGZ_IjuF0yj-39e_CR6OiXgthFAz2VR_OKA2HzyY2zznYIv4DyZZFfrBiMa9M1eR_Qwj8iJHZzMBd_QsEoPIYHMEwuo=@protonmail.com","threadId":"57082","inReplyTo":"xmqqtuf88bw6.fsf@gitster.g","subject":"Re: Fw: Curiosity","fromName":"João Victor Bonfim","fromEmail":"joaovictorbonfim@protonmail.com","sentAt":"2021-12-18T00:17:48Z","receivedAt":"2021-12-18T00:17:51Z","isPatch":false,"sender":{"key":"joaovictorbonfim@protonmail.com","avatar":null},"body":"> That is probably too application specific to be in core-git, but it\n\nApplication specific as in that it is too much of an edge case to be used by all git users?\n\n> is probably a good application for smudge/clean filters like brian\n>\n> alluded to?\n\nPerhaps.\n\n‐‐‐‐‐‐‐ Original Message ‐‐‐‐‐‐‐\n\nEm quinta-feira, 16 de dezembro de 2021 às 18:42, Junio C Hamano <gitster@pobox.com> escreveu:\n\n> Martin Fick mfick@codeaurora.org writes:\n>\n> > On 2021-12-16 14:20, João Victor Bonfim wrote:\n> >\n> > > > To expand on this, if what you're storing is already compressed, like\n> > > >\n> > > > Ogg Vorbis files or PNGs, like are found in that repository, then\n> > > >\n> > > > generally they will not delta well. This is also true of things like\n> > > >\n> > > > Microsoft Office or OpenOffice documents, because they're essentially\n> > > >\n> > > > Zip files.\n> > > >\n> > > > The delta algorithm looks for similarities between files to\n> > > >\n> > > > compress\n> > > >\n> > > > them. If a file is already compressed using something like Deflate,\n> > > >\n> > > > used in PNGs and Zip files, then even very similar files will\n> > > >\n> > > > generally\n> > > >\n> > > > look very different, so deltification will generally be ineffective.\n> > > >\n> > > > ...\n> > > >\n> > > > Maybe I am thinking too outside the box, but wouldn't it be quite more\n> > > >\n> > > > effective for git to identify compressed files, specially on edge cases\n> > > >\n> > > > where the compression doesn't have a good chemistry with delta\n> > > >\n> > > > compression,\n> > > >\n> > > > decompress them for repo storage while also storing the compression\n> > > >\n> > > > algorithm as some metadata tag (like a text string or an ID code\n> > > >\n> > > > decided\n> > > >\n> > > > beforehand), and, when creating the work mirrors, return the\n> > > >\n> > > > compression\n> > > >\n> > > > to its default state before checkout?\n> >\n> > I suspect that for most algorithms and their implementations, this would\n> >\n> > not result in repeatable \"recompressed\" results. Thus the checked-out\n> >\n> > files might be different every time you checked them out. :(\n>\n> That is probably too application specific to be in core-git, but it\n>\n> is probably a good application for smudge/clean filters like brian\n>\n> alluded to?\n"},{"id":"444408","messageId":"xmqq1r2a220u.fsf@gitster.g","threadId":"57082","inReplyTo":"1X3gQ48NK5aBDHcpYMlxESRjqubcCBKJUQu2K0dBOnTyvsXCXXoGDBg2Ff4KarK6WsZnzN3HgqHGOlCKKdF-wtZQ5tHsoAcfit2CTXMWqh4=@protonmail.com","subject":"Re: Fw: Curiosity","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-12-18T00:24:33Z","receivedAt":"2021-12-18T00:24:38Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"João Victor Bonfim  <JoaoVictorBonfim@protonmail.com> writes:\n\n>> I suspect that for most algorithms and their implementations, this would\n>>\n>> not result in repeatable \"recompressed\" results. Thus the checked-out\n>>\n>> files might be different every time you checked them out. :(\n>\n> How or why?\n>\n> Sincere question.\n\nTwo immediate things that come to my mind are lossy compression\nalgorithms (jpeg pictures?) and compressors that do not necessarily\nproduce bit-for-bit identical results (e.g. gzip by default embeds\ntimestamp unless explicitly told not to from a command line option).\n"},{"id":"444410","messageId":"NcaqpHwXfrPNnhzTamuF_xESMQUHMdzLNbHXKOJY59bkJdsT63Nk4PksqQXfqygUMO0EZRFBJeS90r59Tvt2I4kXH69TSe3RwQRXQThxRRA=@protonmail.com","threadId":"57082","inReplyTo":"xmqq1r2a220u.fsf@gitster.g","subject":"Re: Fw: Curiosity","fromName":"João Victor Bonfim","fromEmail":"joaovictorbonfim@protonmail.com","sentAt":"2021-12-18T00:50:19Z","receivedAt":"2021-12-18T00:50:25Z","isPatch":false,"sender":{"key":"joaovictorbonfim@protonmail.com","avatar":null},"body":"Yeah, that sounds reasonable, Junio.\n\n‐‐‐‐‐‐‐ Original Message ‐‐‐‐‐‐‐\n\nEm sexta-feira, 17 de dezembro de 2021 às 21:24, Junio C Hamano <gitster@pobox.com> escreveu:\n\n> João Victor Bonfim JoaoVictorBonfim@protonmail.com writes:\n>\n> > > I suspect that for most algorithms and their implementations, this would\n> > >\n> > > not result in repeatable \"recompressed\" results. Thus the checked-out\n> > >\n> > > files might be different every time you checked them out. :(\n> >\n> > How or why?\n> >\n> > Sincere question.\n>\n> Two immediate things that come to my mind are lossy compression\n>\n> algorithms (jpeg pictures?) and compressors that do not necessarily\n>\n> produce bit-for-bit identical results (e.g. gzip by default embeds\n>\n> timestamp unless explicitly told not to from a command line option).\n"},{"id":"444412","messageId":"df4a5ac37e8d703fa54af91269a9a736@codeaurora.org","threadId":"57082","inReplyTo":"1X3gQ48NK5aBDHcpYMlxESRjqubcCBKJUQu2K0dBOnTyvsXCXXoGDBg2Ff4KarK6WsZnzN3HgqHGOlCKKdF-wtZQ5tHsoAcfit2CTXMWqh4=@protonmail.com","subject":"Re: Fw: Curiosity","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2021-12-18T01:06:56Z","receivedAt":"2021-12-18T01:06:59Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On 2021-12-17 17:15, João Victor Bonfim wrote:\n>> I suspect that for most algorithms and their implementations, this \n>> would\n>> \n>> not result in repeatable \"recompressed\" results. Thus the checked-out\n>> \n>> files might be different every time you checked them out. :(\n> \n> How or why?\n> \n\nHere are some reasons I can think of (I am no expert):\n\n1) Most compression formats are file formats, not exact algorithms, thus \ndifferent program implementations of similar algorithms can create \nvastly different outputs.\n\n2) The same program will evolve over time, get improvements, bug fixes, \netc. so each version of the same program could vary over time even with \nthe same settings. The same program version on different platforms could \nhave different output.\n\n3) Settings, compression programs have compression levels, perhaps \nmemory utilization parameters... The way the program measures these may \nnot be deterministic and non-repeatable.\n\n4) Threading. Some compressions algorithms, such as git repack itself, \ncan use several threads to analyze the input data. And since the timing \nbetween different threads is not deterministic, when cooperating, they \ncan have different results.\n\nMuch of this has to do with the idea that there is usually no such thing \nas \"done\" when it comes to compression. You can probably search \ninfinitely to try and find more data patterns to compress the data more. \nThus compression programs have to have limits based on heuristics (how \nfar to look ahead/behind, how many patterns to remember...) programmed \ninto them to come to an end somehow. How these limits are determined can \nsometimes be non deterministic, it may even involve system resources \n(how much RAM the machine has, how long it has run...) or system config.\n\nI hope that helps,\n\n-Martin\n\n\n> ‐‐‐‐‐‐‐ Original Message ‐‐‐‐‐‐‐\n> \n> Em quinta-feira, 16 de dezembro de 2021 às 18:33, Martin Fick\n> <mfick@codeaurora.org> escreveu:\n> \n>> On 2021-12-16 14:20, João Victor Bonfim wrote:\n>> \n>> > > To expand on this, if what you're storing is already compressed, like\n>> > >\n>> > > Ogg Vorbis files or PNGs, like are found in that repository, then\n>> > >\n>> > > generally they will not delta well. This is also true of things like\n>> > >\n>> > > Microsoft Office or OpenOffice documents, because they're essentially\n>> > >\n>> > > Zip files.\n>> > >\n>> > > The delta algorithm looks for similarities between files to compress\n>> > >\n>> > > them. If a file is already compressed using something like Deflate,\n>> > >\n>> > > used in PNGs and Zip files, then even very similar files will\n>> > >\n>> > > generally\n>> > >\n>> > > look very different, so deltification will generally be ineffective.\n>> \n>> ...\n>> \n>> > Maybe I am thinking too outside the box, but wouldn't it be quite more\n>> >\n>> > effective for git to identify compressed files, specially on edge cases\n>> >\n>> > where the compression doesn't have a good chemistry with delta\n>> >\n>> > compression,\n>> >\n>> > decompress them for repo storage while also storing the compression\n>> >\n>> > algorithm as some metadata tag (like a text string or an ID code\n>> >\n>> > decided\n>> >\n>> > beforehand), and, when creating the work mirrors, return the\n>> >\n>> > compression\n>> >\n>> > to its default state before checkout?\n>> \n>> I suspect that for most algorithms and their implementations, this \n>> would\n>> \n>> not result in repeatable \"recompressed\" results. Thus the checked-out\n>> \n>> files might be different every time you checked them out. :(\n>> \n>> -Martin\n>> \n>> -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------\n>> \n>> The Qualcomm Innovation Center, Inc. is a member of Code\n>> \n>> Aurora Forum, hosted by The Linux Foundation\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code\nAurora Forum, hosted by The Linux Foundation\n"},{"id":"444414","messageId":"Yb06k5ob+bl/oE68@camp.crustytoothpaste.net","threadId":"57082","inReplyTo":"1X3gQ48NK5aBDHcpYMlxESRjqubcCBKJUQu2K0dBOnTyvsXCXXoGDBg2Ff4KarK6WsZnzN3HgqHGOlCKKdF-wtZQ5tHsoAcfit2CTXMWqh4=@protonmail.com","subject":"Re: Fw: Curiosity","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2021-12-18T01:34:11Z","receivedAt":"2021-12-18T01:34:19Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2021-12-18 at 00:15:59, João Victor Bonfim wrote:\n> > I suspect that for most algorithms and their implementations, this would\n> >\n> > not result in repeatable \"recompressed\" results. Thus the checked-out\n> >\n> > files might be different every time you checked them out. :(\n> \n> How or why?\n> \n> Sincere question.\n\nA lossless compression algorithm has to produce an encoded value that,\nwhen decoded, must produce the original input.  Ideally, it will also\nreduce the file size of the original input.  Beyond that, there's a\ngreat deal of freedom to implement that.\n\nJust taking Deflate, which is used in zlib and gzip, as an example,\nthere are different compression settings that control the size of the\nwindow to use that affect compression speed, quality of compression\n(resulting size), and memory usage.  One might prefer using gzip -1 to\nget better performance or use less memory, or gzip -9 to reduce the file\nsize as much as possible.\n\nEven when the same settings are used, the technique used can vary\nbetween versions of the software.  For example, GitHub effectively uses\ngit archive to generate archives, and one time when they upgraded their\nservers, the compression changed in the tarballs and zip files, and\neverybody who was relying on the archives being bit-for-bit identical[0]\nhad a problem.\n\nSo it would be nearly impossible to produce bit-for-bit repeatable\nresults without specifying a specific, hard-coded implementation, and\neven in that case, the behavior might need to change for security\nreasons, so it would end up being difficult to achieve.\n\n[0] Neither Git nor GitHub provides this guarantee, so please do not\nmake this mistake.  If you need a fixed bit-for-bit tarball, save it as\na release artifact.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"444415","messageId":"jwBqM0tKGX4kLwYY1KwT3FYojoznuAHTqQ2zZVw-JCyUUXHvcxPFWjBHYxUp-lxud3rpCw4huIOSIyWdL8SoNx-ETTpNrQm85t54jQLf5ZA=@protonmail.com","threadId":"57082","inReplyTo":"Yb06k5ob+bl/oE68@camp.crustytoothpaste.net","subject":"Re: Fw: Curiosity","fromName":"João Victor Bonfim","fromEmail":"joaovictorbonfim+git-mail-list@protonmail.com","sentAt":"2021-12-18T01:40:24Z","receivedAt":"2021-12-18T01:40:28Z","isPatch":false,"sender":{"key":"joaovictorbonfim+git-mail-list@protonmail.com","avatar":null},"body":"How does one make a release artifact?\no-o\n\n‐‐‐‐‐‐‐ Original Message ‐‐‐‐‐‐‐\n\nEm sexta-feira, 17 de dezembro de 2021 às 22:34, brian m. carlson <sandals@crustytoothpaste.net> escreveu:\n\n> On 2021-12-18 at 00:15:59, João Victor Bonfim wrote:\n>\n> > > I suspect that for most algorithms and their implementations, this would\n> > >\n> > > not result in repeatable \"recompressed\" results. Thus the checked-out\n> > >\n> > > files might be different every time you checked them out. :(\n> >\n> > How or why?\n> >\n> > Sincere question.\n>\n> A lossless compression algorithm has to produce an encoded value that,\n>\n> when decoded, must produce the original input. Ideally, it will also\n>\n> reduce the file size of the original input. Beyond that, there's a\n>\n> great deal of freedom to implement that.\n>\n> Just taking Deflate, which is used in zlib and gzip, as an example,\n>\n> there are different compression settings that control the size of the\n>\n> window to use that affect compression speed, quality of compression\n>\n> (resulting size), and memory usage. One might prefer using gzip -1 to\n>\n> get better performance or use less memory, or gzip -9 to reduce the file\n>\n> size as much as possible.\n>\n> Even when the same settings are used, the technique used can vary\n>\n> between versions of the software. For example, GitHub effectively uses\n>\n> git archive to generate archives, and one time when they upgraded their\n>\n> servers, the compression changed in the tarballs and zip files, and\n>\n> everybody who was relying on the archives being bit-for-bit identical[0]\n>\n> had a problem.\n>\n> So it would be nearly impossible to produce bit-for-bit repeatable\n>\n> results without specifying a specific, hard-coded implementation, and\n>\n> even in that case, the behavior might need to change for security\n>\n> reasons, so it would end up being difficult to achieve.\n>\n> [0] Neither Git nor GitHub provides this guarantee, so please do not\n>\n> make this mistake. If you need a fixed bit-for-bit tarball, save it as\n>\n> a release artifact.\n> --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------\n>\n> brian m. carlson (he/him or they/them)\n>\n> Toronto, Ontario, CA\n"}]}