{"thread":{"id":"59162","subject":"Stability of git-archive, breaking (?) the Github universe, and a possible solution","startedAt":"2023-01-31T00:06:57Z","lastAt":"2023-02-12T17:41:26Z","messageCount":57,"participants":["Eli Schwartz","Ævar Arnfjörð Bjarmason","brian m. carlson","Konstantin Ryabitsev","Michal Suchánek","demerphq","Raymond E. Pasco","Theodore Ts'o","Junio C Hamano","Phillip Wood","Joey Hess","rsbecker@nexbridge.com","René Scharfe"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"471158","messageId":"a812a664-67ea-c0ba-599f-cb79e2d96694@gmail.com","threadId":"59162","inReplyTo":null,"subject":"Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Eli Schwartz","fromEmail":"eschwartz93@gmail.com","sentAt":"2023-01-31T00:06:44Z","receivedAt":"2023-01-31T00:06:57Z","isPatch":false,"sender":{"key":"eschwartz93@gmail.com","avatar":"https://gravatar.com/avatar/80b459bb75c0edb5e116884705adb9095fbd01cbbf91cb29a9ed23577fff7d34?d=mp&s=160"},"body":"For those that haven't seen, github changed its checksums for all\n\"source code\" artifacts attached to any git repository with tags. This\nchange is now reverted due to widespread breakage -- and the lack of\nadvance warning. The technical details of the change appear simple: they\nupgraded git.\n\nProbably the main discussion, complete with Github employees from this\nmailing list responding:\n\nhttps://github.com/bazel-contrib/SIG-rules-authors/issues/11#issuecomment-1409438954\n\nConsequences of that discussion, attempting to mitigate issues by\nwarning people that it already happened:\n\nhttps://github.blog/changelog/2023-01-30-git-archive-checksums-may-change/\n\nAnd where I first saw it: https://github.com/mesonbuild/wrapdb/pull/884\n\nHistorically speaking, git-archive has been stable minus... a bug fix or\ntwo in rare cases, specifically relating to an inability to transcribe\nthe contents of the git repo at all, I think? And the other factor is\nthe compression algorithm used, which is generally GNU gzip, and\nhistorically whatever the system `gzip` command is.\n\nAnd gzip is a stable format. It's a worn-out, battle-weary format, even\n-- it's not the best at compressing, and it's not the best at\ndecompressing, and \"all the cool kids\" are working on cooler formats,\nsuch as zstd which does indeed regularly change its byte output between\nversions. But the advantage of gzip is that it's good *enough*, and it's\nprobably *everywhere*, and it's *reliable*.\n\nGNU gzip is reproducible. busybox gzip was fixed to agree with GNU gzip\n(this is relevant to the handful of people running software forges on,\nsay, Alpine Linux):\n\nhttps://reproducible-builds.org/reports/2019-08/#upstream-news\n\n...\n\nNevertheless, I've seen the sentiment a few times that git doesn't like\ncommitting to output stability of git-archive, because it isn't\nofficially documented (but it's not entirely clear what the benefits of\nchanging are). And yet, git endeavors to do so, in order to prevent\nunnecessary breakage of people who embody Hyrum's Law and need that\nstability.\n\nEven with the new change to the compressor, git-archive is still\nreproducible, it's the internal gzip compressor that isn't. (This may be\nfixable, possibly by embedding an implementation from busybox or from\nGNU gzip? I'm not going to discuss that right now, though I think it's\nan interesting avenue of exploration.)\n\nI've thought about this now and then over the last couple of years,\nbecause I think I have a reasonable compromise that might make everyone\n(or at least most people) happy, and now seems like a good idea to\nmention it.\n\nWhat does everyone think about offering versioned git-archive outputs?\nThis could be user-selectable as an option to `git archive`, but the\nmain goal would be to select a good versioned output format depending on\nwhat is being archived. So:\n\n- first things first, un-default the internal compressor again\n- implement a v2 archive format, where the internal compressor is the\n  default -- no other changes\n- teach git to select an archive format based on the date of the object\n  being archived\n  - when given a commit/tag ID to archive, check which support frame the\n    committer date falls inside\n  - for tree IDs, always use the latest format (it always uses the\n    current date anyway)\n- schedule a date, for the sake of argument, 6 months after the next\n  scheduled release date of git version X.Y in which this change goes\n  live; bake this into the git sources as a transition date, all commits\n  or tags generated after this date fall into the next format support\n  frame\n\n\nThe end result is that for all historic commits or tags, `git archive`\nwill always produce the same output. This can be documented in the\ngit-archive manpage: \"the produced archive is guaranteed to be\nreproducible, unless you override the `tar.<format>.command` or your\nsystem compressor is not reproducible\".\n\nFor *new* commits or tags, everyone gets the benefit of fascinating,\ncool new archive formats with useful improvements at the tar container\nlevel, which is apparently a very desirable feature. The git project no\nlonger has to worry, at all, about whether users will come to complain\nabout how their build pipelines suddenly fail with checksum issues. The\ngit project can simply, fearlessly, go implement innovative new changes\nwithout giving any thought to backwards compatibility.\n\nIt is, simply, that those new changes only apply to projects which are\nstill under active development, and which push new commits or tag new\nreleases after the transition date.\n\nOld states of existing projects (regardless of whether they are still\nactively updating) can go have their old and apparently inefficient\narchives and don't get cool new stuff. That's fine. They're also\nincreasingly rarely used, because they are, after all, old -- and most\nlikely only used for historic archival purposes. If the worst comes to\nworst, well, they managed to produce a somehow useful archive with an\nolder version of git -- nothing will *break* if they don't get the cool\nnew stuff.\n\nAnd for the vast majority of new downloads for new stuff, the in-process\ncompressor saves one fork+exec and is a bit more efficient, I guess?\n\nA note on the transition date: I suggested 6 months after the scheduled\nrelease date, because this gives everyone running a software forge time\nto update git itself, and have everything ready, in time to handle the\nfirst wave of commits and tags that naturally occur after the transition\ndate. And you don't want it to be immediate, because then people will\ntake days or weeks to deploy and the most recent archives will change\n\n\nFor the purposes of this thought experiment, we assume that people don't\nroutinely set the system time to a year in the future. This will only be\ndone in situations such as, say, testing a git upgrade deployment for a\nsoftware forge.\n\n...\n\n\n\"And then no one ever complained about archive checksums changing again.\"\n\n🤞🙏🥺\n\n-- \nEli Schwartz\n"},{"id":"471162","messageId":"230131.86357rrtsg.gmgdl@evledraar.gmail.com","threadId":"59162","inReplyTo":"a812a664-67ea-c0ba-599f-cb79e2d96694@gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-01-31T07:49:12Z","receivedAt":"2023-01-31T08:29:43Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Jan 30 2023, Eli Schwartz wrote:\n\n> For those that haven't seen, github changed its checksums for all\n> \"source code\" artifacts attached to any git repository with tags. This\n> change is now reverted due to widespread breakage -- and the lack of\n> advance warning. The technical details of the change appear simple: they\n> upgraded git.\n>\n> Probably the main discussion, complete with Github employees from this\n> mailing list responding:\n>\n> https://github.com/bazel-contrib/SIG-rules-authors/issues/11#issuecomment-1409438954\n>\n> Consequences of that discussion, attempting to mitigate issues by\n> warning people that it already happened:\n>\n> https://github.blog/changelog/2023-01-30-git-archive-checksums-may-change/\n>\n> And where I first saw it: https://github.com/mesonbuild/wrapdb/pull/884\n\nMaybe I'm the only one that missed this on a first reading, but I\ncouldn't find what specific change in Git was being discussed.\n\nBut it's linked from the now-strikethrough portion of that github.blog\nURL: 4f4be00d302 (archive-tar: use internal gzip by default,\n2022-06-15), first released with v2.38.0.\n\nThat's the change to use gzip as a library instead of gzip(1), I've\nadded the author to the CC list, as well as well as others in the\ninitial ML dicsussion.\n\nThe ML discussion about that series starts at:\nhttps://lore.kernel.org/git/pull.145.git.gitgitgadget@gmail.com/\n\nFor that change specifically I had this comment at the time:\nhttps://lore.kernel.org/git/220615.86wndhwt9a.gmgdl@evledraar.gmail.com/\n\nThe response from René\n(https://lore.kernel.org/git/3ed80afd-34b3-afd8-5ffb-0187a4475ee1@web.de/)\nfills in the \"why\" missing from the commit message itself:\n\n\t\"It's to avoid a run dependency [on gzip(1)] [...] and you can\n\tset tar.tgz.command='gzip -cn' to get the old behavior.  Saving\n\tenergy is a better default, though.\n\nWe can discuss how worthwhile that trade-off is, especially in the face\nof this behavior change GitHub encounterd, but I don't think it was the\nintent with this change to change the output (but maybe René was aware\nof that, but didn't note it).\n\nWhich brings me to...\n\n> Historically speaking, git-archive has been stable minus... a bug fix or\n> two in rare cases, specifically relating to an inability to transcribe\n> the contents of the git repo at all, I think? And the other factor is\n> the compression algorithm used, which is generally GNU gzip, and\n> historically whatever the system `gzip` command is.\n>\n> And gzip is a stable format. It's a worn-out, battle-weary format, even\n> -- it's not the best at compressing, and it's not the best at\n> decompressing, and \"all the cool kids\" are working on cooler formats,\n> such as zstd which does indeed regularly change its byte output between\n> versions. But the advantage of gzip is that it's good *enough*, and it's\n> probably *everywhere*, and it's *reliable*.\n>\n> GNU gzip is reproducible. busybox gzip was fixed to agree with GNU gzip\n> (this is relevant to the handful of people running software forges on,\n> say, Alpine Linux):\n>\n> https://reproducible-builds.org/reports/2019-08/#upstream-news\n>\n> ...\n>\n> Nevertheless, I've seen the sentiment a few times that git doesn't like\n> committing to output stability of git-archive, because it isn't\n> officially documented (but it's not entirely clear what the benefits of\n> changing are). And yet, git endeavors to do so, in order to prevent\n> unnecessary breakage of people who embody Hyrum's Law and need that\n> stability.\n\n...Yes, this has been discussed many times on-list.\n\nMy recollection of those discussions in general is that we were mostly\ntalking about the \"tar\" format itself, moreso than \"gzip\", although in\nthis case it's a change in the gzip component that changed the output.\n\nIt's not clear to me (and I'm asking instead of digging myself, as I\nassume someone at GitHub has dug already) whether our change to the\n\"internal gzip\" is necessarily going to result in a different hash, or\ndid we just forget to provide some option to the library to get the same\nresult as gzip(1).\n\nA major thing you're eliding here is that even if \"tar\" or \"gzip\" is a\n\"a worn-out, battle-weary format\" that does *not* translate to it being\na trivial matter to maintain byte-for-byte compatibility in the archives\n(or compression stream) you produce, even though the resulting output\nonce un-archived or un-compressed is guaranteed to be the same.\n\nWe ship our own \"tar\" for the purposes of this discussion (the archive.c\ncode etc.), but offload the \"gzip\" part to either an external library\n(which is new in v2.38.0, and the subject of this discussion), or to\nGNU's gzip command.\n\nI have no idea if the \"gzip\" part of this would be as easy as saying\n\"we'll default to gzip(1)\", you note \"GNU gzip is reproducible. busybox\ngzip was fixed to agree with GNU gzip\", but does the same apply to other\n\"gzip(1)\"? I know of at least the BSD gzip.\n\nEven then, has even GNU gzip promised that it will forever maintain\nbyte-for-byte compatibility in its output?\n\n> Even with the new change to the compressor, git-archive is still\n> reproducible, it's the internal gzip compressor that isn't. (This may be\n> fixable, possibly by embedding an implementation from busybox or from\n> GNU gzip? I'm not going to discuss that right now, though I think it's\n> an interesting avenue of exploration.)\n\nSo first, aside from whatever the git project does about the default,\nhave you tried running the newer git version with a\ntar.tgz.command='gzip -cn' and seeing if it's compatible with the old\nversion?\n\nIt's unclear from the blog post's \"we are reverting this change for now\"\nwhether that meant a revert of the git version (probably), or a revert\nback to using gzip(1).\n\n> I've thought about this now and then over the last couple of years,\n> because I think I have a reasonable compromise that might make everyone\n> (or at least most people) happy, and now seems like a good idea to\n> mention it.\n>\n> What does everyone think about offering versioned git-archive outputs?\n> This could be user-selectable as an option to `git archive`, but the\n> main goal would be to select a good versioned output format depending on\n> what is being archived. So:\n>\n> - first things first, un-default the internal compressor again\n> - implement a v2 archive format, where the internal compressor is the\n>   default -- no other changes\n> - teach git to select an archive format based on the date of the object\n>   being archived\n>   - when given a commit/tag ID to archive, check which support frame the\n>     committer date falls inside\n>   - for tree IDs, always use the latest format (it always uses the\n>     current date anyway)\n> - schedule a date, for the sake of argument, 6 months after the next\n>   scheduled release date of git version X.Y in which this change goes\n>   live; bake this into the git sources as a transition date, all commits\n>   or tags generated after this date fall into the next format support\n>   frame\n>\n> The end result is that for all historic commits or tags, `git archive`\n> will always produce the same output. This can be documented in the\n> git-archive manpage: \"the produced archive is guaranteed to be\n> reproducible, unless you override the `tar.<format>.command` or your\n> system compressor is not reproducible\".\n>\n> For *new* commits or tags, everyone gets the benefit of fascinating,\n> cool new archive formats with useful improvements at the tar container\n> level, which is apparently a very desirable feature. The git project no\n> longer has to worry, at all, about whether users will come to complain\n> about how their build pipelines suddenly fail with checksum issues. The\n> git project can simply, fearlessly, go implement innovative new changes\n> without giving any thought to backwards compatibility.\n>\n> It is, simply, that those new changes only apply to projects which are\n> still under active development, and which push new commits or tag new\n> releases after the transition date.\n>\n> Old states of existing projects (regardless of whether they are still\n> actively updating) can go have their old and apparently inefficient\n> archives and don't get cool new stuff. That's fine. They're also\n> increasingly rarely used, because they are, after all, old -- and most\n> likely only used for historic archival purposes. If the worst comes to\n> worst, well, they managed to produce a somehow useful archive with an\n> older version of git -- nothing will *break* if they don't get the cool\n> new stuff.\n>\n> And for the vast majority of new downloads for new stuff, the in-process\n> compressor saves one fork+exec and is a bit more efficient, I guess?\n>\n> A note on the transition date: I suggested 6 months after the scheduled\n> release date, because this gives everyone running a software forge time\n> to update git itself, and have everything ready, in time to handle the\n> first wave of commits and tags that naturally occur after the transition\n> date. And you don't want it to be immediate, because then people will\n> take days or weeks to deploy and the most recent archives will change\n>\n> For the purposes of this thought experiment, we assume that people don't\n> routinely set the system time to a year in the future. This will only be\n> done in situations such as, say, testing a git upgrade deployment for a\n> software forge.\n\nThis sounds like a workable transition plan, but it assumes that we had\na really good reason to change to the \"internal gzip\" by default, and\nthat we must move forward with that change in some way.\n\nI don't think that's the case per the linked-to on-list discussion, the\naim was just to provide output if gzip(1) wasn't available, so all we'd\nneed is the pseudocode of:\n\n\t- Prepare our tar stream\n        - Try to strem it to gzip(1)\n        - If that fails with \"command does not exist\" fall back to the\n          internal one (possibly with a warning about possibly-different\n          output)\n\nThen systems without a gzip(1) could produce output (which René was\naiming for), but those with a system gzip(1) (e.g. GitHub's production\ninstallation) could just continue to use it.\n\nThat's still a band-aid on the larger questions I raised above,\ni.e. whether we'd want to forever guarantee the output of \"git archive\"\nitself, and of the \"tar.tgz.command\".\n\nMy off-the-cuff response to that is that we should probably:\n\n - Guarantee the \"git archive\" output itself (without compression),\n   leaving the out that it *may* change in the future with notice (or\n   we'd just version it)\n\n - Switch back to using gzip(1) by default, whatever gzip(1) that\n   happens to be.\n\nBut:\n\n - Promise that the total end result will be byte-for-byte the same, as\n   that would imply a promise about the external gzip(1).\n\n - Just prominently note in our docs that if you want the\n   archive->compression to be byte-for-byte with the past it's up to you\n   to ensure that your compressor gives you that guarantee.\n"},{"id":"471164","messageId":"fd39ed5f-cee4-0d60-c0c2-fae18246fb53@gmail.com","threadId":"59162","inReplyTo":"230131.86357rrtsg.gmgdl@evledraar.gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Eli Schwartz","fromEmail":"eschwartz93@gmail.com","sentAt":"2023-01-31T09:11:51Z","receivedAt":"2023-01-31T09:16:48Z","isPatch":false,"sender":{"key":"eschwartz93@gmail.com","avatar":"https://gravatar.com/avatar/80b459bb75c0edb5e116884705adb9095fbd01cbbf91cb29a9ed23577fff7d34?d=mp&s=160"},"body":"Quick response for now...\n\nOn 1/31/23 2:49 AM, Ævar Arnfjörð Bjarmason wrote:\n> So first, aside from whatever the git project does about the default,\n> have you tried running the newer git version with a\n> tar.tgz.command='gzip -cn' and seeing if it's compatible with the old\n> version?\n> \n> It's unclear from the blog post's \"we are reverting this change for now\"\n> whether that meant a revert of the git version (probably), or a revert\n> back to using gzip(1).\n\n\nI do not know which one Github internally did, but I can confirm that\nthe gzipped tarballs which github started shipping, when gunzipped,\nproduced an uncompressed tarball that was byte-identical to uncompressed\neditions of the historic ones.\n\ni.e. you could do this:\n\n```\nwget ${important_archive_release}\n\ngzip -dc < ${important_archive_localfile} | gzip -cn >\n${important_archive_localfile}.new\n```\n\nAnd:\n- they have different checksums\n- the .new file has reverted to the same checksum as historic versions\n  from last year that are frozen into manifests\n\nThat was part of my original investigation, before I located the public\nconversations.\n\n\n-- \nEli Schwartz\n"},{"id":"471165","messageId":"Y9jlWYLzZ/yy4NqD@tapette.crustytoothpaste.net","threadId":"59162","inReplyTo":"a812a664-67ea-c0ba-599f-cb79e2d96694@gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-01-31T09:54:58Z","receivedAt":"2023-01-31T09:55:04Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-01-31 at 00:06:44, Eli Schwartz wrote:\n> Nevertheless, I've seen the sentiment a few times that git doesn't like\n> committing to output stability of git-archive, because it isn't\n> officially documented (but it's not entirely clear what the benefits of\n> changing are). And yet, git endeavors to do so, in order to prevent\n> unnecessary breakage of people who embody Hyrum's Law and need that\n> stability.\n\nI'm one of the GitHub employees who chimed in there, and I'm also a Git\ncontributor in my own time (and I am speaking here only in my personal\ncapacity, since this is a personal address).  I made a change some years\nback to the archive format to fix the permissions on pax headers when\nextracted as files, and kernel.org was relying on that and broke.  Linus\nyelled at me because of that.\n\nSince then, I've been very opposed to us guaranteeing output format\nconsistency without explicitly doing so.  I had sent some patches before\nthat I don't think ever got picked up that documented this explicitly.\nI very much don't want people to come to rely on our behaviour unless we\nexplicitly guarantee it.\n\n> What does everyone think about offering versioned git-archive outputs?\n> This could be user-selectable as an option to `git archive`, but the\n> main goal would be to select a good versioned output format depending on\n> what is being archived. So:\n> \n> - first things first, un-default the internal compressor again\n> - implement a v2 archive format, where the internal compressor is the\n>   default -- no other changes\n> - teach git to select an archive format based on the date of the object\n>   being archived\n>   - when given a commit/tag ID to archive, check which support frame the\n>     committer date falls inside\n>   - for tree IDs, always use the latest format (it always uses the\n>     current date anyway)\n> - schedule a date, for the sake of argument, 6 months after the next\n>   scheduled release date of git version X.Y in which this change goes\n>   live; bake this into the git sources as a transition date, all commits\n>   or tags generated after this date fall into the next format support\n>   frame\n\nI am actually very much in favour of providing a standard, deterministic\nversion of pax (the extended tar format) that we use and documenting it\nas a standard so that other archive tools can use that.  That is, we\ndocument some canonical tar format that is bit-for-bit identical that we\n(and hopefully GNU tar and libarchive) will agree should be used to\nserialize files for software interchange.  I don't think this should be\ndependent on the date at all, but I do believe it should be versioned\nand tested, and the version number embedded as a pax header.  I think\nthis would be valuable for simply having reproducible archives in\ngeneral, including for things like Docker containers, Debian packages,\nRust crates, and more, and I'm happy to work with others on such a\nformat, as I've said in the past on the list.  People can opt-in to\nwhatever format they want when creating an archive and continue to use\nthat forever if they like.\n\nPart of the reason I think this is valuable is that once SHA-1 and\nSHA-256 interoperability is present, git archive will change the\ncontents of the archive format, since it will embed a SHA-256 hash into\nthe file instead of a SHA-1 hash, since that's what's in the repository.\nThus, we can't produce an archive that's deterministic in the face of\nSHA-1/SHA-256 interoperability concerns, and we need to create a new\nformat that doesn't contain that data embedded in it.\n\nHaving said that, I don't think this should be based on the timestamp of\nthe file, since that means that two otherwise identical archives\ndiffering in timestamp aren't ever going to be the same, and we do see\npeople who import or vendor other projects.  Nor do I think we should\nattempt to provide consistent compression, since I believe the output of\nthings like zlib has changed in the past, and we can't continually carry\nan old, potentially insecure version of zlib just because the output\nchanged.  People should be able to implement compression using gzip,\nzlib, pigz, miniz_oxide, or whatever if they want, since people\nimplement Git in many different languages, and we won't want to force\npeople using memory-safe languages like Go and Rust to explicitly use\nzlib for archives.\n\nThat may mean that it's important for people to actually decompress the\narchive before checking hashes if they want deterministic behaviour, and\nI'm okay with that.  You already have to do that if you're verifying the\nsignature on Git tarballs, since only the uncompressed tar archive is\nsigned, so I don't think this is out of the question.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"471168","messageId":"230131.86tu06rkbp.gmgdl@evledraar.gmail.com","threadId":"59162","inReplyTo":"Y9jlWYLzZ/yy4NqD@tapette.crustytoothpaste.net","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-01-31T11:31:43Z","receivedAt":"2023-01-31T11:54:08Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Jan 31 2023, brian m. carlson wrote:\n\n> Part of the reason I think this is valuable is that once SHA-1 and\n> SHA-256 interoperability is present, git archive will change the\n> contents of the archive format, since it will embed a SHA-256 hash into\n> the file instead of a SHA-1 hash, since that's what's in the repository.\n> Thus, we can't produce an archive that's deterministic in the face of\n> SHA-1/SHA-256 interoperability concerns, and we need to create a new\n> format that doesn't contain that data embedded in it.\n\nI don't see why a format change would be required in this context.\n\nIf a repository were to switch over to SHA-256 wouldn't a better\nsolution to this be to disambiguate whether you're requesting a SHA-1 or\nSHA-256 derived archive in the URL? E.g. to never serve up an archive\nwith a SHA-256 embedded in the header at:\n\n\thttps://github.com/git/git/archive/refs/tags/v2.39.1.tar.gz\n\nBut require a URL like:\n\n\thttps://github.com/git/git/archive-sha256/refs/tags/v2.39.1.tar.gz\n\nIf you did that then existing archives would continue to have the same\nbyte-for-byte content (assuming that the result of this discussion is\nthat we support that forever), but they'd always be generated with \"-c\nextensions.objectFormat=sha1\". For always-SHA256 repos such a URL would\nfail to generate anything.\n\nBut for repos that used to be SHA-1 but are now SHA-256 either URL would\nwork, but the PAX header would be different, referring to the SHA-1 or\nSHA-256 commit, respectively.\n\nWhereas your proposal seems to be that we should omit that SHA-(1|256)\nfrom the \"comment\" entirely. That would seem to require either a one-off\nchange of all existing archives, or some cut-off date (or other marker).\n\nIf you've got a cut-off, you could also just use it to decide whether to\ngenerate a SHA-1 or SHA-256 archive, and without that you'd be back to\nthe one-off breakage.\n\nI also find it very useful that we've got the commit OID in the archive,\nas it allows for round-tripping from archives back to the relevant\nrepository commit. Losing that entirely for SHA-1<->SHA-256 interop\nwould be unfortunate, especially if it turns out we could have easily\nkept it\n\n> Having said that, I don't think this should be based on the timestamp of\n> the file, since that means that two otherwise identical archives\n> differing in timestamp aren't ever going to be the same, and we do see\n> people who import or vendor other projects.\n\nYes, I agree that doing this by that sort of heuristic would be bad.\n\n> Nor do I think we should\n> attempt to provide consistent compression, since I believe the output of\n> things like zlib has changed in the past, and we can't continually carry\n> an old, potentially insecure version of zlib just because the output\n> changed.  People should be able to implement compression using gzip,\n> zlib, pigz, miniz_oxide, or whatever if they want, since people\n> implement Git in many different languages, and we won't want to force\n> people using memory-safe languages like Go and Rust to explicitly use\n> zlib for archives.\n\nAs I noted in the side-thread I think an acceptable solution would be to\npush the problem of the consistent compressor downstream. I.e. if a site\nlike GitHub wants to maintain a potentially old version of GNU gzip that\nshould be up to them.\n\nBut I think it's a valid concern that we should guarantee the stability\nof the archive format.\n"},{"id":"471186","messageId":"20230131150555.ewiwsbczwep6ltbi@meerkat.local","threadId":"59162","inReplyTo":"Y9jlWYLzZ/yy4NqD@tapette.crustytoothpaste.net","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2023-01-31T15:05:55Z","receivedAt":"2023-01-31T15:09:29Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Tue, Jan 31, 2023 at 09:54:58AM +0000, brian m. carlson wrote:\n> I'm one of the GitHub employees who chimed in there, and I'm also a Git\n> contributor in my own time (and I am speaking here only in my personal\n> capacity, since this is a personal address).  I made a change some years\n> back to the archive format to fix the permissions on pax headers when\n> extracted as files, and kernel.org was relying on that and broke.  Linus\n> yelled at me because of that.\n> \n> Since then, I've been very opposed to us guaranteeing output format\n> consistency without explicitly doing so.  I had sent some patches before\n> that I don't think ever got picked up that documented this explicitly.\n> I very much don't want people to come to rely on our behaviour unless we\n> explicitly guarantee it.\n\nI understand your position, but I also think it's one of those things that\nhappen despite your best efforts to prevent it. :)\n\nMay I suggest adding a \"git-archive --stable\" that offers this guarantee,\nsimply as a matter of codifying the fact that the world has built\ninfrastructure around git's repeatable output. Maybe just for .tar (and\n.tar.gz).\n\nI know this complicates the code and makes it more \"expensive\" to maintain,\nbut it would be dramatically less expensive than changing the established\npractices around the world.\n\n-K\n"},{"id":"471187","messageId":"6fc8e122-a190-c291-c347-258a5a2ad9c9@gmail.com","threadId":"59162","inReplyTo":"Y9jlWYLzZ/yy4NqD@tapette.crustytoothpaste.net","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Eli Schwartz","fromEmail":"eschwartz93@gmail.com","sentAt":"2023-01-31T15:56:52Z","receivedAt":"2023-01-31T15:57:02Z","isPatch":false,"sender":{"key":"eschwartz93@gmail.com","avatar":"https://gravatar.com/avatar/80b459bb75c0edb5e116884705adb9095fbd01cbbf91cb29a9ed23577fff7d34?d=mp&s=160"},"body":"On 1/31/23 4:54 AM, brian m. carlson wrote:\n> Part of the reason I think this is valuable is that once SHA-1 and\n> SHA-256 interoperability is present, git archive will change the\n> contents of the archive format, since it will embed a SHA-256 hash into\n> the file instead of a SHA-1 hash, since that's what's in the repository.\n> Thus, we can't produce an archive that's deterministic in the face of\n> SHA-1/SHA-256 interoperability concerns, and we need to create a new\n> format that doesn't contain that data embedded in it.\n\n\nI assume that whatever the reason for originally embedding the OID into\nthe file is still an applicable reason even if a new PAX format is\nestablished for the use of git-archive.\n\nIt may not be a great reason -- I don't know. Perhaps there's an\nargument to remove it. But can't that be done irrespective of\nstandardizing the PAX format?\n\n...\n\nI'm not deeply knowledgeable about the SHA-256 transition work -- or\nknowledgeable at all about it, frankly. (Also my understanding was it\nseems to have stalled as discussed in https://lwn.net/Articles/898522/\n-- I understand that you're still enthusiastic about the work? But that\ndoesn't really answer \"is there a timeframe for that to ever happen\".)\n\nBut I sort of assumed that the transition work would already have to\nembed a fair bit of information into the repository about the whole\nprocess? Would it not be possible to determine whether a given tag\nstarted life as SHA-1 or SHA-256? Maybe even just a date when the\nrepository was converted to work with both, and embed the OID based on\nwhether the tag is tagging contents that were created after that conversion?\n\nSeems to me like the problem should be solvable if people want to solve it.\n\n...\n\ngit-archive run on a commit obviously doesn't have this problem -- it\ncan simply embed the OID for the same argument it was called with. But I\nassume it's far more common to access tag-based github endpoints. :D\n\n\n> Having said that, I don't think this should be based on the timestamp of\n> the file, since that means that two otherwise identical archives\n> differing in timestamp aren't ever going to be the same, and we do see\n> people who import or vendor other projects. \n\n\nThe timestamp of the output file? Surely not. But I only suggested the\ntimestamp of the commit/tag metadata that git-archive is asked to\nproduce output for. And we would need that in order to solve the problem\nthat reproducible github API archive endpoints poses.\n\nI'm not sure what the \"import or vendor other projects\" angle here\nmeans. Do you mean people who copy a directory of files into their\nproject? Who expects this to be the same to begin with? And doesn't\nembedding the OID kill this idea, since the entire point of git commit\nsha's is that you shouldn't (it should be prohibitively unrealistic to)\nbe able to produce the same one twice in different contexts?\n\nI have never said to myself \"ah yes, I really would like to be able to\ndownload a git auto-generated tarball for project A, and compare its\nhash to the tarball for project B, and have them compare identical even\nthough they are different projects with different commits\". IMHO this\nisn't an interesting problem to solve -- the interesting problem to\nsolve is that a single absolute URL to a downloadable file should be\nable to offer documented guarantees that it will always be the same\nfile, even though it is generated on the fly.\n\n\n> Nor do I think we should\n> attempt to provide consistent compression, since I believe the output of\n> things like zlib has changed in the past, and we can't continually carry\n> an old, potentially insecure version of zlib just because the output\n> changed.  People should be able to implement compression using gzip,\n> zlib, pigz, miniz_oxide, or whatever if they want, since people\n> implement Git in many different languages, and we won't want to force\n> people using memory-safe languages like Go and Rust to explicitly use\n> zlib for archives.\n\n\nI do not think it is realistic or reasonable for people to implement\ncompression using intentionally incompatible replacements for gzip and\nexpect interoperability of any sort.\n\nI also don't think people *have* to implement compression in rust using\nzlib, but if they are going to make a git-alike that produces archives,\nit would be worth it for them to write whatever memory-safe rust is\nnecessary to memory-safely produce the same output stream of bytes. It's\nno less feasible than making sure that busybox gzip and GNU gzip produce\nthe same output, surely.\n\nAlternatively, they could just not bother with gzip at all, and make\ntheir git-alike produce zstd-compressed tarballs, which change their\nbyte outputs every time a new zstd release is published. :D Again, why\nlimit yourself to gzip if you want to be innovative anyway.\n\n\n> That may mean that it's important for people to actually decompress the\n> archive before checking hashes if they want deterministic behaviour, and\n> I'm okay with that.  You already have to do that if you're verifying the\n> signature on Git tarballs, since only the uncompressed tar archive is\n> signed, so I don't think this is out of the question.\n\n\nThis is a very kernel.org-centric view of things, I think. I have rarely\nseen PGP signatures applied to the uncompressed tar except in that\ncontext. The vast majority of tarballs with signatures have signed a\nsingle compressed tarball and don't concern themselves with, say,\nproviding a rotating backdated changeable list of compression formats\nwith a single signature covering all of them.\n\nNevertheless, in order to handle kernel.org-style tarballs, you are\nentirely correct that one should be able to handle this.\n\n>From experience, I can say that this needs to be selected on a\nper-tarball basis. Since signature files have filenames, we can match\ntheir stems and given foo.tar.asc and foo.tar.gz, check the signature of\nthe output of gzip -dc < foo.tar.gz, but given foo.tar.gz.asc and\nfoo.tar.gz, simply check the signature of the original foo.tar.gz.\n\nThis doesn't really work for checksums, because you need to settle on\none or the other everywhere or else embed decompression information into\nyour checksum metadata field.\n\nAnd for tarballs that are generated once and uploaded to ftp storage,\nnot repeatedly generated on the fly, we know the checksum will never\nlegitimately change, so we *want* to hash the compressed file.\nDecompressing kernel.org tarballs in order to run PGP on them is *slow*.\nAlthough at least one can verify the checksums first without\ndecompression, which is virtually guaranteed to catch invalid source\ncode releases, so if you ever progress to the PGP verification stage\nit's unlikely to be wasted effort -- that tarball is definitely getting\nused to build something.\n\n\n-- \nEli Schwartz\n"},{"id":"471188","messageId":"20230131162049.mgqdxcucjesw4afr@meerkat.local","threadId":"59162","inReplyTo":"6fc8e122-a190-c291-c347-258a5a2ad9c9@gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2023-01-31T16:20:49Z","receivedAt":"2023-01-31T16:20:55Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Tue, Jan 31, 2023 at 10:56:52AM -0500, Eli Schwartz wrote:\n> And for tarballs that are generated once and uploaded to ftp storage,\n> not repeatedly generated on the fly, we know the checksum will never\n> legitimately change, so we *want* to hash the compressed file.\n> Decompressing kernel.org tarballs in order to run PGP on them is *slow*.\n\nFWIW, the most correct way is:\n\n* download sha256sums.asc and verify its signature (auto-signed by infra)\n* download the tarball you want and verify that the checksum matches\n* uncompress and verify the PGP signature (signed by developer)\n\nThis script implements this workflow:\nhttps://git.kernel.org/pub/scm/linux/kernel/git/mricon/korg-helpers.git/tree/get-verified-tarball\n\n-K\n"},{"id":"471189","messageId":"df7b0b43-efa2-ea04-dc5b-9515e7f1d86f@gmail.com","threadId":"59162","inReplyTo":"20230131162049.mgqdxcucjesw4afr@meerkat.local","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Eli Schwartz","fromEmail":"eschwartz93@gmail.com","sentAt":"2023-01-31T16:34:59Z","receivedAt":"2023-01-31T16:35:29Z","isPatch":false,"sender":{"key":"eschwartz93@gmail.com","avatar":"https://gravatar.com/avatar/80b459bb75c0edb5e116884705adb9095fbd01cbbf91cb29a9ed23577fff7d34?d=mp&s=160"},"body":"On 1/31/23 11:20 AM, Konstantin Ryabitsev wrote:\n> On Tue, Jan 31, 2023 at 10:56:52AM -0500, Eli Schwartz wrote:\n>> And for tarballs that are generated once and uploaded to ftp storage,\n>> not repeatedly generated on the fly, we know the checksum will never\n>> legitimately change, so we *want* to hash the compressed file.\n>> Decompressing kernel.org tarballs in order to run PGP on them is *slow*.\n> \n> FWIW, the most correct way is:\n> \n> * download sha256sums.asc and verify its signature (auto-signed by infra)\n> * download the tarball you want and verify that the checksum matches\n> * uncompress and verify the PGP signature (signed by developer)\n> \n> This script implements this workflow:\n> https://git.kernel.org/pub/scm/linux/kernel/git/mricon/korg-helpers.git/tree/get-verified-tarball\n\n\nThis is just what I said, but with an additional first step for when you\nare updating to a new tarball and don't have your own checksums\nintegrated into your own ecosystem tracking.\n\nIn most contexts, it's utterly unacceptable to not remember the checksum\nof the file you used last time and instead simply trust PGP identity\nverification. This permits upstream the technical means to be malicious,\nand re-upload a totally different tarball with the same name, different\ncontents, and different PGP signature, and you will never notice because\nthe PGP signature is still okay.\n\nJust because I trust you all doesn't mean I should ignore existing best\npractices to make sure that I always use the same reviewed\nbyte-identical tarball -- or find out exactly why it changed.\n\n\n-- \nEli Schwartz\n"},{"id":"471202","messageId":"20230131203425.qxy5f7aappzip5om@meerkat.local","threadId":"59162","inReplyTo":"df7b0b43-efa2-ea04-dc5b-9515e7f1d86f@gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2023-01-31T20:34:25Z","receivedAt":"2023-01-31T20:34:32Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Tue, Jan 31, 2023 at 11:34:59AM -0500, Eli Schwartz wrote:\n> In most contexts, it's utterly unacceptable to not remember the checksum\n> of the file you used last time and instead simply trust PGP identity\n> verification. This permits upstream the technical means to be malicious,\n> and re-upload a totally different tarball with the same name, different\n> contents, and different PGP signature, and you will never notice because\n> the PGP signature is still okay.\n\nYes, it's true, and it's something that Sigstore tries to address.\n\nThat said, if I wanted to trojan a download and had access to both the\ninfrastructure and the developer's credentials, I wouldn't pick a months-old\nrelease for this purpose. I would wait until I see a new release coming out\nand then swap it mid-flight. This lets me defeat even transparency-log based\nsolutions like sigstore.\n\n(I'll probably be giving a talk at the Linux Security Summit titled \"How to\ntrojan the Linux Kernel\" where I'll go into some of these considerations. :))\n\n-K\n"},{"id":"471203","messageId":"20230131204550.GI19419@kitsune.suse.cz","threadId":"59162","inReplyTo":"df7b0b43-efa2-ea04-dc5b-9515e7f1d86f@gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2023-01-31T20:45:50Z","receivedAt":"2023-01-31T20:46:23Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"On Tue, Jan 31, 2023 at 11:34:59AM -0500, Eli Schwartz wrote:\n> On 1/31/23 11:20 AM, Konstantin Ryabitsev wrote:\n> > On Tue, Jan 31, 2023 at 10:56:52AM -0500, Eli Schwartz wrote:\n> >> And for tarballs that are generated once and uploaded to ftp storage,\n> >> not repeatedly generated on the fly, we know the checksum will never\n> >> legitimately change, so we *want* to hash the compressed file.\n> >> Decompressing kernel.org tarballs in order to run PGP on them is *slow*.\n> > \n> > FWIW, the most correct way is:\n> > \n> > * download sha256sums.asc and verify its signature (auto-signed by infra)\n> > * download the tarball you want and verify that the checksum matches\n> > * uncompress and verify the PGP signature (signed by developer)\n> > \n> > This script implements this workflow:\n> > https://git.kernel.org/pub/scm/linux/kernel/git/mricon/korg-helpers.git/tree/get-verified-tarball\n> \n> \n> This is just what I said, but with an additional first step for when you\n> are updating to a new tarball and don't have your own checksums\n> integrated into your own ecosystem tracking.\n> \n> In most contexts, it's utterly unacceptable to not remember the checksum\n> of the file you used last time and instead simply trust PGP identity\n> verification. This permits upstream the technical means to be malicious,\n> and re-upload a totally different tarball with the same name, different\n> contents, and different PGP signature, and you will never notice because\n> the PGP signature is still okay.\n\nBut where is the hash remembered?\n\nThe signature is a hash+signature, it you can replace that, you can also\nrepolace a hash without a signature.\n\nYou can store hashesd of anything you want locally, and indeed such\nstored hashes in some build systemns did detect some code hosting\ncorruption but that's not for upstream to do, that's something that only\nunrelated third party can do.\n\nThanks\n\nMichal\n"},{"id":"471208","messageId":"Y9mXB1LaYSUJBlwF@tapette.crustytoothpaste.net","threadId":"59162","inReplyTo":"20230131150555.ewiwsbczwep6ltbi@meerkat.local","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-01-31T22:32:39Z","receivedAt":"2023-01-31T22:32:44Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-01-31 at 15:05:55, Konstantin Ryabitsev wrote:\n> On Tue, Jan 31, 2023 at 09:54:58AM +0000, brian m. carlson wrote:\n> > I'm one of the GitHub employees who chimed in there, and I'm also a Git\n> > contributor in my own time (and I am speaking here only in my personal\n> > capacity, since this is a personal address).  I made a change some years\n> > back to the archive format to fix the permissions on pax headers when\n> > extracted as files, and kernel.org was relying on that and broke.  Linus\n> > yelled at me because of that.\n> > \n> > Since then, I've been very opposed to us guaranteeing output format\n> > consistency without explicitly doing so.  I had sent some patches before\n> > that I don't think ever got picked up that documented this explicitly.\n> > I very much don't want people to come to rely on our behaviour unless we\n> > explicitly guarantee it.\n> \n> I understand your position, but I also think it's one of those things that\n> happen despite your best efforts to prevent it. :)\n> \n> May I suggest adding a \"git-archive --stable\" that offers this guarantee,\n> simply as a matter of codifying the fact that the world has built\n> infrastructure around git's repeatable output. Maybe just for .tar (and\n> .tar.gz).\n\nIt is my intention to implement just .tar.  That's my proposal: simply a\npax-based format that serializes in a consistent way according to a\npredefined spec.\n\nAs far as whether other people want to implement consistent compression,\nthey are welcome to also write a spec and implement it.  I personally\nfeel that's too hard to get right and am not planning on working on it.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"471211","messageId":"Y9nBdZZCRZgPzB/v@tapette.crustytoothpaste.net","threadId":"59162","inReplyTo":"6fc8e122-a190-c291-c347-258a5a2ad9c9@gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-02-01T01:33:41Z","receivedAt":"2023-02-01T01:33:50Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-01-31 at 15:56:52, Eli Schwartz wrote:\n> On 1/31/23 4:54 AM, brian m. carlson wrote:\n> > Part of the reason I think this is valuable is that once SHA-1 and\n> > SHA-256 interoperability is present, git archive will change the\n> > contents of the archive format, since it will embed a SHA-256 hash into\n> > the file instead of a SHA-1 hash, since that's what's in the repository.\n> > Thus, we can't produce an archive that's deterministic in the face of\n> > SHA-1/SHA-256 interoperability concerns, and we need to create a new\n> > format that doesn't contain that data embedded in it.\n> \n> \n> I assume that whatever the reason for originally embedding the OID into\n> the file is still an applicable reason even if a new PAX format is\n> established for the use of git-archive.\n> \n> It may not be a great reason -- I don't know. Perhaps there's an\n> argument to remove it. But can't that be done irrespective of\n> standardizing the PAX format?\n> \n> ...\n> \n> I'm not deeply knowledgeable about the SHA-256 transition work -- or\n> knowledgeable at all about it, frankly. (Also my understanding was it\n> seems to have stalled as discussed in https://lwn.net/Articles/898522/\n> -- I understand that you're still enthusiastic about the work? But that\n> doesn't really answer \"is there a timeframe for that to ever happen\".)\n\nThe timeframe is when my employer pays me to work on it.  Right now,\nI've implemented functional SHA-256 repositories but am currently a bit\non the way to burnout and am very selective about what things I'm doing\noutside of work.  My hope is that my employer will find time for me to\nwork on the interop stuff soon, but I'm not at liberty to discuss this\nmore in depth at the moment.\n\n> But I sort of assumed that the transition work would already have to\n> embed a fair bit of information into the repository about the whole\n> process? Would it not be possible to determine whether a given tag\n> started life as SHA-1 or SHA-256? Maybe even just a date when the\n> repository was converted to work with both, and embed the OID based on\n> whether the tag is tagging contents that were created after that conversion?\n\nIt's designed such that the two objects are completely interoperable and\ncan be accessed by either name, depending on how the repository is\nconfigured locally.  There may be a signature for one algorithm, both,\nor neither, so it's hard to say definitively what version it's created\nwith.  That is completely intentional since the goal is to transition\nseamlessly from one to another at any point depending on the preferences\nof the owner of the local repository.\n\n> > Having said that, I don't think this should be based on the timestamp of\n> > the file, since that means that two otherwise identical archives\n> > differing in timestamp aren't ever going to be the same, and we do see\n> > people who import or vendor other projects. \n> \n> \n> The timestamp of the output file? Surely not. But I only suggested the\n> timestamp of the commit/tag metadata that git-archive is asked to\n> produce output for. And we would need that in order to solve the problem\n> that reproducible github API archive endpoints poses.\n\nI think it would simply be easier to say, \"This is the command-line\noption that implements canonical tar version 1.\"  If you want a\nreproducible archive, you use that command-line option, and your\nuncompressed tar archive is reproducible.  Otherwise, you get the same\nguarantees on reproducibility that we've always provided, which is\nabsolutely none.\n\nUsing commit and tag metadata doesn't solve the problem of trees, which\nwould use the current timestamp.  It's better to solve the problem in a\nconsistent way, which would mean embedding a fixed timestamp (probably\nthe Epoch) into those tree tarballs.\n\nIn my view, using the commit or tag timestamp is very risky, because it\nchanges the behaviour at some point in the future without notifying\npeople.  If we produce a tar archive that isn't readable by FooZip, say,\nthen nobody will realize that until we actually start producing them,\nseveral months after the release.  And, I should point out, this still\nposes problems for GitHub and other forges, because GitHub doesn't run\nthe latest release right away; we usually trail a version or two.  So\nusing the commit or tag timestamp might mean that on an upgrade,\nsuddenly the behaviour changes because the new version has a change\n(which was scheduled to have occurred in the past) but the old version\ndoesn't.\n\nIn addition, the one guarantee we've given with archives in the past is\nthat the same version of Git with the same input (flags, repository,\netc.) will produce deterministic results (that is, the same output), and\nI think we're likely to run afoul of that with a timestamp-based\napproach.  I don't want the archive to suddenly be different because I\nhappened to do \"git commit --amend\" to update just a commit message and\nwe happened to cross that timestamp threshold.\n\n> I'm not sure what the \"import or vendor other projects\" angle here\n> means. Do you mean people who copy a directory of files into their\n> project? Who expects this to be the same to begin with? And doesn't\n> embedding the OID kill this idea, since the entire point of git commit\n> sha's is that you shouldn't (it should be prohibitively unrealistic to)\n> be able to produce the same one twice in different contexts?\n\nWe have people who import the entirety of Chromium into a project at\none time to work on a browser-based project.\n\n> I have never said to myself \"ah yes, I really would like to be able to\n> download a git auto-generated tarball for project A, and compare its\n> hash to the tarball for project B, and have them compare identical even\n> though they are different projects with different commits\". IMHO this\n> isn't an interesting problem to solve -- the interesting problem to\n> solve is that a single absolute URL to a downloadable file should be\n> able to offer documented guarantees that it will always be the same\n> file, even though it is generated on the fly.\n\nI do think having identical output for identical contents is very\nvaluable.  If our goal is reproducible output, we should endeavour to\nproduce identical output for identical input.  What we're specifically\ntrying to move away from is varying output based on the same input.\n\n> I do not think it is realistic or reasonable for people to implement\n> compression using intentionally incompatible replacements for gzip and\n> expect interoperability of any sort.\n\nI disagree completely.  The gzip and zlib formats are documented in RFCs\nand have been since 1996.  There are already at least a half-dozen\ninteroperable implementations, including zlib, gzip, pigz, Go's standard\nlibrary, miniz_oxide, and the Windows archiver.  I'm sure if I searched\nI could find at least half a dozen more.\n\n> I also don't think people *have* to implement compression in rust using\n> zlib, but if they are going to make a git-alike that produces archives,\n> it would be worth it for them to write whatever memory-safe rust is\n> necessary to memory-safely produce the same output stream of bytes. It's\n> no less feasible than making sure that busybox gzip and GNU gzip produce\n> the same output, surely.\n\nI don't agree at all.  The Go standard library couldn't achieve that,\nbecause busybox and gzip are GPL and doing that would almost certainly\nrequire looking at the code, which would require the Go standard library\nto be GPL as well.  The same thing goes for zlib, which is permissively\nlicensed, and which is clearly the obvious choice if we had to settle on\na standard, since it's a shared library.\n\nThat also ignores tools like pigz which provide parallel compression and\ncan provide an order of magnitude performance increase, but which won't\nprovide an identical byte stream.  Why should we require people to use a\nsingle core if they have a very large archive that could compress\nseveral times as fast with a parallel operation?\n\nMy goal is to produce tar archives that are interoperable based on a\nspec.  That spec would be implementable by Git, GNU tar, libarchive, or\nanyone else, by reading the spec and following it.  That's very\ndifferent from saying, \"Well, just make your program do exactly the same\nthing as this other one without sharing any code.\"  If you want to write\na spec for canonical gzip, I'm interested in reading it, but I think\nit's practically going to be difficult to achieve.\n\n> > That may mean that it's important for people to actually decompress the\n> > archive before checking hashes if they want deterministic behaviour, and\n> > I'm okay with that.  You already have to do that if you're verifying the\n> > signature on Git tarballs, since only the uncompressed tar archive is\n> > signed, so I don't think this is out of the question.\n> \n> \n> This is a very kernel.org-centric view of things, I think. I have rarely\n> seen PGP signatures applied to the uncompressed tar except in that\n> context. The vast majority of tarballs with signatures have signed a\n> single compressed tarball and don't concern themselves with, say,\n> providing a rotating backdated changeable list of compression formats\n> with a single signature covering all of them.\n\nSure, and that's a valid approach if you have a consistent, persistent\ntarball.  However, Git does not persist data forever in tarballs, and\npeople want to use different versions to get the same data, which is a\nnew guarantee that we'd be providing.  That is an easy guarantee to\nprovide with tar, but not an easy guarantee to provide with the gzip\nformat, as we've all just seen.\n\n> >From experience, I can say that this needs to be selected on a\n> per-tarball basis. Since signature files have filenames, we can match\n> their stems and given foo.tar.asc and foo.tar.gz, check the signature of\n> the output of gzip -dc < foo.tar.gz, but given foo.tar.gz.asc and\n> foo.tar.gz, simply check the signature of the original foo.tar.gz.\n> \n> This doesn't really work for checksums, because you need to settle on\n> one or the other everywhere or else embed decompression information into\n> your checksum metadata field.\n\nI don't think that's absolutely required.  You need to know how to\ndecompress the archive, and you can have a hash for the tarball before\ndecompression or after decompression, as well as possibly needing to\ndeal with multiple different hash algorithms.  I've implemented this\nmyself when I was a vendor of Git and lots of other software, and we\nwould take the hash of the compressed or decompressed archive as shipped\nby the vendor and verify it, as long as the hash was sufficiently\nstrong.\n\n> And for tarballs that are generated once and uploaded to ftp storage,\n> not repeatedly generated on the fly, we know the checksum will never\n> legitimately change, so we *want* to hash the compressed file.\n> Decompressing kernel.org tarballs in order to run PGP on them is *slow*.\n> Although at least one can verify the checksums first without\n> decompression, which is virtually guaranteed to catch invalid source\n> code releases, so if you ever progress to the PGP verification stage\n> it's unlikely to be wasted effort -- that tarball is definitely getting\n> used to build something.\n\nSure, and if you want to generate tarballs once and upload them to\nstorage, go ahead.  That's always an option.  Even GitHub provides you\nthe option to do that with release assets if you want.\n\nMy proposal is to provide deterministic archives in a functionally and\npractically achievable way with nothing more than a version of Git,\nwhich I think we can do with tar, but not gzip.  I'm happy to be proven\nwrong if you can develop a spec for canonical gzip compression.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"471221","messageId":"230201.86pmatr9mj.gmgdl@evledraar.gmail.com","threadId":"59162","inReplyTo":"Y9mXB1LaYSUJBlwF@tapette.crustytoothpaste.net","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-01T09:40:57Z","receivedAt":"2023-02-01T09:57:34Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Jan 31 2023, brian m. carlson wrote:\n\n> As far as whether other people want to implement consistent compression,\n> they are welcome to also write a spec and implement it.  I personally\n> feel that's too hard to get right and am not planning on working on it.\n\n\"A spec\" here seems like overkill to me, so far on that front we've been\nshelling out to gzip(1), and the breakage/event that triggered this\nthread is rectified by starting to do that again by default.\n\nIt means that someone writing a clean-room implementation of git would\nlikely run into the same issue, if they used e.g. the Go language and a\nnative Go implementation of deflate.\n\nBut so what? We don't need to make promises for all potential git\nimplementations, just this one. So we could add a blurb like this to the\ndocs:\n\n\tAs people have come to rely on the exact \"deflate\"\n\timplementation \"git archive\" promises to invoke the system's\n\t\"gzip\" binary by default, under the assumption that its output\n\tis stable. If that's no longer the case you'll need to complain\n\tto whoever maintains your local \"gzip\".\n\nIf we wanted to be even more helpful we could bunde and ship an old\nversion of GNU gzip with our sources, and either default to that, or\noffer it as a \"--stable\" implementation of deflate.\n\nThat would be going above & beyond what's needed IMO, but still a lot\neasier than the daunting task of writing a specification that exactly\ndescribed GNU gzip's current behavior, to the point where you could\nclean-room implement it and be guaranteed byte-for-byte compatibility.\n"},{"id":"471226","messageId":"CANgJU+V0QRFwmTh8ZzY=28kmbUw=DvSLE24LioOXp6_ozq+RdA@mail.gmail.com","threadId":"59162","inReplyTo":"230201.86pmatr9mj.gmgdl@evledraar.gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2023-02-01T11:34:06Z","receivedAt":"2023-02-01T11:34:21Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On Wed, 1 Feb 2023 at 11:26, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> That would be going above & beyond what's needed IMO, but still a lot\n> easier than the daunting task of writing a specification that exactly\n> described GNU gzip's current behavior, to the point where you could\n> clean-room implement it and be guaranteed byte-for-byte compatibility.\n\nWhy does it have to be gzip? It is not that hard to come up with a\nrelatively good compression algorithm that is stable if you aren't\nexpecting super fast performance or super good compression. If all you\nneed is good enough but stability is a hard requirement then\nalgorithms like LZW are available (it has been out of patent since\n~2003), and produce reasonable results. If people want a stable\narchive then they might have to use some tool that git provides to\ndecompress and they might not get the best compression ratios, nor\nspeed, but they would get stability. You can write a decent LZW\nimplementation in a few hundred lines of code. With a bit of care you\ncould implement it in a way that allows you to compute the true hash\ndigest of the compressed data without actually decompressing it as\nwell, which would address some of the concerns that brian raised with\nregard to security I think.\n\nWhy does this email remind me of that old canard that any sufficiently\nadvanced piece of software gains the ability to send emails? :-)\n\ncheers,\nYves\n\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"471234","messageId":"20230201122152.GJ19419@kitsune.suse.cz","threadId":"59162","inReplyTo":"CANgJU+V0QRFwmTh8ZzY=28kmbUw=DvSLE24LioOXp6_ozq+RdA@mail.gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2023-02-01T12:21:52Z","receivedAt":"2023-02-01T12:21:58Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"On Wed, Feb 01, 2023 at 12:34:06PM +0100, demerphq wrote:\n> On Wed, 1 Feb 2023 at 11:26, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> > That would be going above & beyond what's needed IMO, but still a lot\n> > easier than the daunting task of writing a specification that exactly\n> > described GNU gzip's current behavior, to the point where you could\n> > clean-room implement it and be guaranteed byte-for-byte compatibility.\n> \n> Why does it have to be gzip? It is not that hard to come up with a\nhistorical reasons?\n"},{"id":"471235","messageId":"8452eb684b212b1e364bdc4709d4b202@ameretat.dev","threadId":"59162","inReplyTo":"230201.86pmatr9mj.gmgdl@evledraar.gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Raymond E. Pasco","fromEmail":"ray@ameretat.dev","sentAt":"2023-02-01T12:17:14Z","receivedAt":"2023-02-01T12:25:26Z","isPatch":false,"sender":{"key":"ray@ameretat.dev","avatar":"https://avatars.githubusercontent.com/u/115765?v=4"},"body":"February 1, 2023 4:40 AM, \"Ævar Arnfjörð Bjarmason\" <avarab@gmail.com> wrote:\n> As people have come to rely on the exact \"deflate\"\n> implementation \"git archive\" promises to invoke the system's\n> \"gzip\" binary by default, under the assumption that its output\n> is stable. If that's no longer the case you'll need to complain\n> to whoever maintains your local \"gzip\".\n\nSurely if reproducibility of .tar.gz files is the goal,\"invoke\nwhatever arbitrary binary on $PATH happens to be called gzip\" is an\npoor solution.\n\nIt is only even possible to consider stabilizing gzip output as a\ngoal for Git (although this seems ill-advised for the reasons\nBrian already discussed) in the post-2.38 world where git is\ndoing the gzipping.\n\nIf one has the requirement to substitute one's own specific\ncompressor, there is an option for that.\n"},{"id":"471237","messageId":"230201.86lelhr1wv.gmgdl@evledraar.gmail.com","threadId":"59162","inReplyTo":"Y9jlWYLzZ/yy4NqD@tapette.crustytoothpaste.net","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-01T12:42:54Z","receivedAt":"2023-02-01T12:44:22Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Jan 31 2023, brian m. carlson wrote:\n\n> Since then, I've been very opposed to us guaranteeing output format\n> consistency without explicitly doing so.  I had sent some patches before\n> that I don't think ever got picked up that documented this explicitly.\n> I very much don't want people to come to rely on our behaviour unless we\n> explicitly guarantee it.\n\nFWIW I think the reason that didn't get picked up (I went back and read\nthe discussion) is that there was some feedback on the v1, [1] suggested\n(at least to me) that you'd re-roll it, but that re-roll never seems to\nhave made it to the list.\n\n1. https://lore.kernel.org/git/YD7aDwX%2FaiRN0GZs@camp.crustytoothpaste.net/\n"},{"id":"471238","messageId":"CANgJU+VLseURimM++38WA81uFPbnoHiToOt4F4UFL9yVbQpBEw@mail.gmail.com","threadId":"59162","inReplyTo":"20230201122152.GJ19419@kitsune.suse.cz","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2023-02-01T12:48:17Z","receivedAt":"2023-02-01T12:48:33Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On Wed, 1 Feb 2023, 20:21 Michal Suchánek, <msuchanek@suse.de> wrote:\n>\n> On Wed, Feb 01, 2023 at 12:34:06PM +0100, demerphq wrote:\n> > Why does it have to be gzip? It is not that hard to come up with a\n\n> historical reasons?\n\nCurrently git doesn't advertise that archive creation is stable\nright[1]? So I wrote that with the assumption that this new\ncompression would only be used when making a new archive with a\nhypothetical new '--stable' option. So historical reasons don't come\nup. Or was there some other form of history that you meant?\n\nI'm just trying to point out here that stable compression is doable\nand doesn't need to be as complex as specifying a stable gzip format.\nI am not even saying git should just do this, just that it /could/ if\nit decided that stability was important, and that doing so wouldn't\ninvolve the complexity that Avar was implying would be needed.  Simple\ncompression like LZ variants are pretty straightforward to implement,\nachieve pretty good compression and can run pretty fast.\n\nYves\n[1] if it did the issue kicking off this thread would not have\nhappened as there would be a test that would have noticed the change.\n"},{"id":"471243","messageId":"230201.86cz6tqyvy.gmgdl@evledraar.gmail.com","threadId":"59162","inReplyTo":"CANgJU+VLseURimM++38WA81uFPbnoHiToOt4F4UFL9yVbQpBEw@mail.gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-01T13:43:15Z","receivedAt":"2023-02-01T13:49:57Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Feb 01 2023, demerphq wrote:\n\n> On Wed, 1 Feb 2023, 20:21 Michal Suchánek, <msuchanek@suse.de> wrote:\n>>\n>> On Wed, Feb 01, 2023 at 12:34:06PM +0100, demerphq wrote:\n>> > Why does it have to be gzip? It is not that hard to come up with a\n>\n>> historical reasons?\n>\n> Currently git doesn't advertise that archive creation is stable\n> right[1]? So I wrote that with the assumption that this new\n> compression would only be used when making a new archive with a\n> hypothetical new '--stable' option. So historical reasons don't come\n> up. Or was there some other form of history that you meant?\n\nWe haven't advertised it, but people have come to rely on it, as the\nwidespread breakages reported when upgrading to v2.38.0 at the start of\nthis thread show.\n\nThat's unfortunate, and those people probably shouldn't have done that,\nbut that's water under the bridge. I think it would be irresponsible to\nchange the output willy-nilly at this point, especially when it seems\nrather easy to find some compromise everyone will be happy with.\n\n> I'm just trying to point out here that stable compression is doable\n> and doesn't need to be as complex as specifying a stable gzip format.\n> I am not even saying git should just do this, just that it /could/ if\n> it decided that stability was important, and that doing so wouldn't\n> involve the complexity that Avar was implying would be needed.  Simple\n> compression like LZ variants are pretty straightforward to implement,\n> achieve pretty good compression and can run pretty fast.\n>\n> Yves\n> [1] if it did the issue kicking off this thread would not have\n> happened as there would be a test that would have noticed the change.\n\nI have some patches I'm about to submit to address issues in this\nthread, and it does add *a* test for archive output stability.\n\nBut I'm not at all confident that it's exhaustive. I just found it by\nexperiment, by locating tests ouf ours where the \"git archive\" output at\nthe end is different with gzip and \"git archive gzip\".\n\nBut is it guaranteed to find all potential cases where repository\ncontent might trigger different output with different gzip\nimplementations? I don't know, but probably not.\n"},{"id":"471247","messageId":"CANgJU+VNY-VziRijSwyb1WF9s31hKroK+2VJ0qEGiYweiA59Ug@mail.gmail.com","threadId":"59162","inReplyTo":"230201.86cz6tqyvy.gmgdl@evledraar.gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2023-02-01T15:21:56Z","receivedAt":"2023-02-01T15:22:51Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On Wed, 1 Feb 2023 at 14:49, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Wed, Feb 01 2023, demerphq wrote:\n>\n> > On Wed, 1 Feb 2023, 20:21 Michal Suchánek, <msuchanek@suse.de> wrote:\n> >>\n> >> On Wed, Feb 01, 2023 at 12:34:06PM +0100, demerphq wrote:\n> >> > Why does it have to be gzip? It is not that hard to come up with a\n> >\n> >> historical reasons?\n> >\n> > Currently git doesn't advertise that archive creation is stable\n> > right[1]? So I wrote that with the assumption that this new\n> > compression would only be used when making a new archive with a\n> > hypothetical new '--stable' option. So historical reasons don't come\n> > up. Or was there some other form of history that you meant?\n>\n> We haven't advertised it, but people have come to rely on it, as the\n> widespread breakages reported when upgrading to v2.38.0 at the start of\n> this thread show.\n>\n> That's unfortunate, and those people probably shouldn't have done that,\n> but that's water under the bridge. I think it would be irresponsible to\n> change the output willy-nilly at this point, especially when it seems\n> rather easy to find some compromise everyone will be happy with.\n>\n> > I'm just trying to point out here that stable compression is doable\n> > and doesn't need to be as complex as specifying a stable gzip format.\n> > I am not even saying git should just do this, just that it /could/ if\n> > it decided that stability was important, and that doing so wouldn't\n> > involve the complexity that Avar was implying would be needed.  Simple\n> > compression like LZ variants are pretty straightforward to implement,\n> > achieve pretty good compression and can run pretty fast.\n> >\n> > Yves\n> > [1] if it did the issue kicking off this thread would not have\n> > happened as there would be a test that would have noticed the change.\n>\n> I have some patches I'm about to submit to address issues in this\n> thread, and it does add *a* test for archive output stability.\n>\n> But I'm not at all confident that it's exhaustive. I just found it by\n> experiment, by locating tests ouf ours where the \"git archive\" output at\n> the end is different with gzip and \"git archive gzip\".\n>\n> But is it guaranteed to find all potential cases where repository\n> content might trigger different output with different gzip\n> implementations? I don't know, but probably not.\n\nBTW, I just happened to be looking at the zstd docs (I am updating\ncode that uses it), I saw this:\n\nZstandard's format is stable and documented in\n[RFC8878](https://datatracker.ietf.org/doc/html/rfc8878). Multiple\nindependent implementations are already available.\nThis repository represents the reference implementation, provided as\nan open-source dual [BSD](LICENSE) and [GPLv2](COPYING) licensed **C**\nlibrary,\nand a command line utility producing and decoding `.zst`, `.gz`, `.xz`\nand `.lz4` files.\nShould your project require another programming language,\na list of known ports and bindings is provided on [Zstandard\nhomepage](http://www.zstd.net/#other-languages).\n\nSo it sounds like that is a spec you could use. Not sure exactly what\nthey mean by \"stable\", but given the .gz compatibility maybe it would\nbe worth considering. Its a lot faster than zlib. (The library I\nsupport includes Snappy, Zlib, and Zstd, and the latter is faster and\nbetter than the other two.)\n\nYves\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"471261","messageId":"Y9q129WbseimgeBS@mit.edu","threadId":"59162","inReplyTo":"CANgJU+VNY-VziRijSwyb1WF9s31hKroK+2VJ0qEGiYweiA59Ug@mail.gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2023-02-01T18:56:27Z","receivedAt":"2023-02-01T18:56:50Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"If the goal is stable tar.gz files, Debian has a very nice soution\ncalled pristine-tar[1].  This you to store a tar.gz image which in a\nvery efficient way, by leveraging the objects in the git repository.\n\n[1] https://manpages.debian.org/unstable/pristine-tar/pristine-tar.1.en.html\n\nThe data is stored on the pristine-tar branch, and is quite efficient:\n\n% git show --stat pristine-tar\ncommit 56dded989c9e0c852b8af9ae72ffe94270bfd34a (origin/pristine-tar, github/pristine-tar, pristine-tar)\nAuthor: Theodore Ts'o <tytso@mit.edu>\nDate:   Thu Dec 30 01:06:13 2021 -0500\n\n    pristine-tar data for e2fsprogs_1.46.5.orig.tar.gz\n\n e2fsprogs_1.46.5.orig.tar.gz.asc   |  11 +++++++++++\n e2fsprogs_1.46.5.orig.tar.gz.delta | Bin 0 -> 59034 bytes\n e2fsprogs_1.46.5.orig.tar.gz.id    |   1 +\n 3 files changed, 12 insertions(+)\n\nAnd this allows me to reproduce the original tar.gz file, along with a\nGPG signature file, which is about 9 megabytes.  The *.id file\ncontains the git commit from which the tar file was generated, and\nthis is what allows the *.delta file to be as small as it is.\n\n% pristine-tar checkout e2fsprogs_1.46.5.orig.tar.gz -s e2fsprogs_1.46.5.orig.tar.gz.asc\npristine-tar: successfully generated e2fsprogs_1.46.5.orig.tar.gz\npristine-tar: successfully generated e2fsprogs_1.46.5.orig.tar.gz.asc\n\n% ls -sh e2fsprogs_1.46.5.orig.tar.gz*\n9.1M e2fsprogs_1.46.5.orig.tar.gz  4.0K e2fsprogs_1.46.5.orig.tar.gz.asc\n\n% gpg e2fsprogs_1.46.5.orig.tar.gz.asc\ngpg: WARNING: no command supplied.  Trying to guess what you mean ...\ngpg: assuming signed data in 'e2fsprogs_1.46.5.orig.tar.gz'\ngpg: Signature made Thu 30 Dec 2021 01:02:52 AM EST\ngpg:                using RSA key 2B69B954DBFE0879288137C9F2F95956950D81A3\ngpg: Good signature from \"Theodore Ts'o <tytso@mit.edu>\" [ultimate]\ngpg:                 aka \"Theodore Ts'o <tytso@debian.org>\" [ultimate]\ngpg:                 aka \"Theodore Ts'o <tytso@google.com>\" [ultimate]\nPrimary key fingerprint: 3AB0 57B7 E78D 945C 8C55  91FB D36F 769B C118 04F0\n     Subkey fingerprint: 2B69 B954 DBFE 0879 2881  37C9 F2F9 5956 950D 81A3\n\nThis is currently a Debian special, and while its functionality was\ndesigned to work well with Debian packaging workflows, but it's a\ngeneral tool that could be used in multiple contexts, not just for\nDebian packaging.\n\nIf I recall correctly, pristine-tar is currently in maintenance mode,\nand I suspect if someone was interested in investing time into making\npristine-tar more portable to other OS's, including MacOS and Windows,\nand maybe potentially even integrating into git directly, the current\nmaintainer of pristine-tar might be quite happy to let other people\ngive the code more TLC.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"471280","messageId":"Y9ry5Wxck4s/X2B+@tapette.crustytoothpaste.net","threadId":"59162","inReplyTo":"230201.86pmatr9mj.gmgdl@evledraar.gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-02-01T23:16:53Z","receivedAt":"2023-02-01T23:18:27Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-02-01 at 09:40:57, Ævar Arnfjörð Bjarmason wrote:\n> \"A spec\" here seems like overkill to me, so far on that front we've been\n> shelling out to gzip(1), and the breakage/event that triggered this\n> thread is rectified by starting to do that again by default.\n\nSure, that will fix the immediate problem.\n\n> But so what? We don't need to make promises for all potential git\n> implementations, just this one. So we could add a blurb like this to the\n> docs:\n> \n> \tAs people have come to rely on the exact \"deflate\"\n> \timplementation \"git archive\" promises to invoke the system's\n> \t\"gzip\" binary by default, under the assumption that its output\n> \tis stable. If that's no longer the case you'll need to complain\n> \tto whoever maintains your local \"gzip\".\n\nI don't think a blurb is necessary, but you're basically underscoring\nthe problem, which is that nobody is willing to promise that compression\nis consistent, but yet people want to rely on that fact.  I'm willing to\nwrite and implement a consistent tar spec and to guarantee compatibility\nwith that, but the tension here is that people also want gzip to never\nchange its byte format ever, which frankly seems unrealistic without\nexplicit guarantees.  Maybe the authors will agree to promise that, but\nit seems unlikely.\n\n> If we wanted to be even more helpful we could bunde and ship an old\n> version of GNU gzip with our sources, and either default to that, or\n> offer it as a \"--stable\" implementation of deflate.\n\nThat would probably break things, because gzip is GPLv3, and we'd need\nto ship a much older GPLv2 gzip, which would probably differ from the\ncurrent behaviour, and might also have some security problems.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"471281","messageId":"Y9rzT2r1rjhn+/HW@tapette.crustytoothpaste.net","threadId":"59162","inReplyTo":"230201.86lelhr1wv.gmgdl@evledraar.gmail.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-02-01T23:18:39Z","receivedAt":"2023-02-01T23:18:59Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-02-01 at 12:42:54, Ævar Arnfjörð Bjarmason wrote:\n> \n> On Tue, Jan 31 2023, brian m. carlson wrote:\n> \n> > Since then, I've been very opposed to us guaranteeing output format\n> > consistency without explicitly doing so.  I had sent some patches before\n> > that I don't think ever got picked up that documented this explicitly.\n> > I very much don't want people to come to rely on our behaviour unless we\n> > explicitly guarantee it.\n> \n> FWIW I think the reason that didn't get picked up (I went back and read\n> the discussion) is that there was some feedback on the v1, [1] suggested\n> (at least to me) that you'd re-roll it, but that re-roll never seems to\n> have made it to the list.\n\nThat may very well have been the case.  As mentioned upthread, I have\nvery limited time to work on Git these days, and sometimes things just\nfall through the cracks.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"471283","messageId":"xmqqh6w5x8i8.fsf@gitster.g","threadId":"59162","inReplyTo":"Y9ry5Wxck4s/X2B+@tapette.crustytoothpaste.net","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-02-01T23:37:19Z","receivedAt":"2023-02-01T23:37:24Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> I don't think a blurb is necessary, but you're basically underscoring\n> the problem, which is that nobody is willing to promise that compression\n> is consistent, but yet people want to rely on that fact.  I'm willing to\n> write and implement a consistent tar spec and to guarantee compatibility\n> with that, but the tension here is that people also want gzip to never\n> change its byte format ever, which frankly seems unrealistic without\n> explicit guarantees.  Maybe the authors will agree to promise that, but\n> it seems unlikely.\n\nJust to step back a bit, where does the distinction between\nguaranteeing the tar format stability and gzip compressed bitstream\nstability come from?  At both levels, the same thing can be\nexpressed in multiple different ways, I think, but spelling out how\nexactly the compressor compresses is more involved than spelling out\nhow entries in a tar archive is ordered and each entry is expressed,\nor something?\n\n> That would probably break things, because gzip is GPLv3, and we'd need\n> to ship a much older GPLv2 gzip, which would probably differ from the\n> current behaviour, and might also have some security problems.\n\nYup, security issues may make bit-for-bit-stability unrealistic.\nIIRC, the last time we had discussion on this topic, we settled\non stability across the same version of Git (i.e. deterministic\nresult)?\n"},{"id":"471298","messageId":"230202.86r0v8q3oz.gmgdl@evledraar.gmail.com","threadId":"59162","inReplyTo":"Y9ry5Wxck4s/X2B+@tapette.crustytoothpaste.net","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T00:42:48Z","receivedAt":"2023-02-02T01:03:18Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Feb 01 2023, brian m. carlson wrote:\n\n> [[PGP Signed Part:Undecided]]\n> On 2023-02-01 at 09:40:57, Ævar Arnfjörð Bjarmason wrote:\n>> \"A spec\" here seems like overkill to me, so far on that front we've been\n>> shelling out to gzip(1), and the breakage/event that triggered this\n>> thread is rectified by starting to do that again by default.\n>\n> Sure, that will fix the immediate problem.\n>\n>> But so what? We don't need to make promises for all potential git\n>> implementations, just this one. So we could add a blurb like this to the\n>> docs:\n>> \n>> \tAs people have come to rely on the exact \"deflate\"\n>> \timplementation \"git archive\" promises to invoke the system's\n>> \t\"gzip\" binary by default, under the assumption that its output\n>> \tis stable. If that's no longer the case you'll need to complain\n>> \tto whoever maintains your local \"gzip\".\n>\n> I don't think a blurb is necessary, but you're basically underscoring\n> the problem, which is that nobody is willing to promise that compression\n> is consistent, but yet people want to rely on that fact.  I'm willing to\n> write and implement a consistent tar spec and to guarantee compatibility\n> with that, but the tension here is that people also want gzip to never\n> change its byte format ever, which frankly seems unrealistic without\n> explicit guarantees.  Maybe the authors will agree to promise that, but\n> it seems unlikely.\n\nMaybe they won't, the point is that an upgrade of git wouldn't break\ngithub in the way that's been observed, instead that potential breakage\nwould happen whenever the OS (or whatever's providing \"gzip\") is\nupgraded.\n\nSo, if gzip promises to never change such sites can upgrade it without\nissues, but if it does they'll presumably need to pin it forever.\n\nAnd those sites that don't care about \"git archive\" stability can use\nwhatever their local \"gzip\" is, without caring that the output might\nchange.\n\n>> If we wanted to be even more helpful we could bunde and ship an old\n>> version of GNU gzip with our sources, and either default to that, or\n>> offer it as a \"--stable\" implementation of deflate.\n>\n> That would probably break things, because gzip is GPLv3, and we'd need\n> to ship a much older GPLv2 gzip, which would probably differ from the\n> current behaviour, and might also have some security problems.\n\nWe're way off in the realm of the hypothetical, I don't think we need a\ngzip fallback, we can make it the issue of the rare downstream user who\nneeds such stability.\n\nBut if we shipped a last-good gzip my understanding of software\nlicensing is that we could ship the GPLv3 version.\n\nThe issue with combining GPLv3 and GPLv2 works is if you do something\nlike upgrade our wildmatch.c to the GPLv3 version (ours is derived from\nan older GPLv2 version). Then our combined work is derived from two\ndifferent licenses.\n\nBut if you're just invoking a different process those two sources can\nuse incompatible licenses. There's established precedence for that\nthroughout the industry, and it's the FSF's position on the matter.\n\nSo if we offered to build a gzip for you from GPLv3 sources shipped\nin-tree that wouldn't infect the rest of git's GPLv2 code, any more than\nDebian shipping both git and gzip is cross-contaminating the two.\n\nIt might cause us some hassle with distributors for whom any mention of\nGPLv3 is anathema (e.g. Apple), but I understand that that's general\nparanoia about its patent clauses impacting the distributor, not a\nlicense incompatiblity.\n"},{"id":"471307","messageId":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","threadId":"59162","inReplyTo":"230131.86357rrtsg.gmgdl@evledraar.gmail.com","subject":"[PATCH 0/9] git archive: use gzip again by default, document output stabilty","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T09:32:20Z","receivedAt":"2023-02-02T09:32:44Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"As reported in\nhttps://lore.kernel.org/git/a812a664-67ea-c0ba-599f-cb79e2d96694@gmail.com/\nchanging the default \"tgz\" output method of from \"gzip(1)\" to our\ninternal \"git archive gzip\" (using zlib ) broke things for users in\nthe wild that assume that the \"git archive\" output is stable, most\nnotably GitHub: https://github.com/orgs/community/discussions/45830\n\nLeaving aside the larger question of whether we're going to promise\noutput stability for \"git archive\" in general, the motivation for that\nchange was to have a working compression method on systems that lacked\na gzip(1).\n\nAs the disruption of changing the default isn't worth it, let's use\ngzip(1) again by default, and only fall back on the new \"git archive\ngzip\" if it isn't available.\n\nThe later parts of this series then document and test for the output\nstability of the command.\n\nWe're not promising anything new there, except that we now promise\nthat we're going to use \"gzip\" as the default compressor, but that\nit's up to that command to be stable, should the user desire output\nstability.\n\nThe documentation discusses the various caveats involved, suggests\nalternatives to checksumming compressed archives, but in the end notes\nwhat's been the policy so far: We're not promising that the \"tar\"\noutput is going to be stable.\n\nThe early parts of this series (1-2/9) are clean-up for existing\nconfig drift, as later in the series we'll otherwise need to change\nthe divergent config documentation in two places.\n\nCI & branch for this at:\nhttps://github.com/avar/git/tree/avar/archive-internal-gzip-not-the-default\n\nÆvar Arnfjörð Bjarmason (9):\n  archive & tar config docs: de-duplicate configuration section\n  git config docs: document \"tar.<format>.{command,remote}\"\n  archiver API: make the \"flags\" in \"struct archiver\" an enum\n  archive: omit the shell for built-in \"command\" filters\n  archive-tar.c: move internal gzip implementation to a function\n  archive: use \"gzip -cn\" for stability, not \"git archive gzip\"\n  test-lib.sh: add a lazy GZIP prerequisite\n  archive tests: test for \"gzip -cn\" and \"git archive gzip\" stability\n  git archive docs: document output non-stability\n\n Documentation/config/tar.txt           | 29 +++++++-\n Documentation/git-archive.txt          | 96 +++++++++++++++++++-------\n archive-tar.c                          | 78 ++++++++++++++-------\n archive.h                              | 11 +--\n t/t5000-tar-tree.sh                    |  2 -\n t/t5005-archive-stability.sh           | 70 +++++++++++++++++++\n t/t5562-http-backend-content-length.sh |  2 -\n t/test-lib.sh                          |  4 ++\n 8 files changed, 231 insertions(+), 61 deletions(-)\n create mode 100755 t/t5005-archive-stability.sh\n\n-- \n2.39.1.1392.g63e6d408230\n\n"},{"id":"471308","messageId":"patch-1.9-feb3e1bebd7-20230202T093212Z-avarab@gmail.com","threadId":"59162","inReplyTo":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","subject":"[PATCH 1/9] archive & tar config docs: de-duplicate configuration section","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T09:32:21Z","receivedAt":"2023-02-02T09:32:47Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"The \"tar.umask\" documentation was initially added in [1], and was\nduplicated from the start. Then with [2] the two started drifting\napart. Let's consolidate them with a change like the ones made in the\ncommits merged in [3].\n\n1. ce1a79b6a74 (tar-tree: add the \"tar.umask\" config option,\n   2006-07-20)\n2. 687157c736d (Documentation: update tar.umask default, 2007-08-21)\n3. 7a54d740451 (Merge branch 'ab/dedup-config-and-command-docs',\n   2022-09-14)\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/tar.txt  | 4 +++-\n Documentation/git-archive.txt | 8 +-------\n 2 files changed, 4 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/config/tar.txt b/Documentation/config/tar.txt\nindex de8ff48ea9d..c68e294bbc5 100644\n--- a/Documentation/config/tar.txt\n+++ b/Documentation/config/tar.txt\n@@ -3,4 +3,6 @@ tar.umask::\n \ttar archive entries.  The default is 0002, which turns off the\n \tworld write bit.  The special value \"user\" indicates that the\n \tarchiving user's umask will be used instead.  See umask(2) and\n-\tlinkgit:git-archive[1].\n+\tlinkgit:git-archive[1] for\n+\tdetails. If `--remote` is used then only the configuration of\n+\tthe remote repository takes effect.\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex 60c040988bb..bbb407d4975 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -131,13 +131,7 @@ tar\n CONFIGURATION\n -------------\n \n-tar.umask::\n-\tThis variable can be used to restrict the permission bits of\n-\ttar archive entries.  The default is 0002, which turns off the\n-\tworld write bit.  The special value \"user\" indicates that the\n-\tarchiving user's umask will be used instead.  See umask(2) for\n-\tdetails.  If `--remote` is used then only the configuration of\n-\tthe remote repository takes effect.\n+include::config/tar.txt[]\n \n tar.<format>.command::\n \tThis variable specifies a shell command through which the tar\n-- \n2.39.1.1392.g63e6d408230\n\n"},{"id":"471309","messageId":"patch-2.9-3cf4bf5a538-20230202T093212Z-avarab@gmail.com","threadId":"59162","inReplyTo":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","subject":"[PATCH 2/9] git config docs: document \"tar.<format>.{command,remote}\"","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T09:32:22Z","receivedAt":"2023-02-02T09:32:49Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Since the \"tar.<format>.command\" and \"tar.<format>.remote\"\nconfiguration was added in [1] and [2], we have not included it in the\n\"git-config(1)\" docs themselves.\n\nSince we're including \"Documentation/config/tar.txt\" in\n\"Documentation/config/git-archive.txt\" as of the preceding commit,\nlet's move this documentation to the former, to be included in the\nlatter.\n\nThis is a move-only change, aside from changing the mention of \"`git\narchive`\" to \"linkgit:git-archive[1]\", for consistency with other such\nmentions.\n\n1. 767cf4579f0 (archive: implement configurable tar filters,\n   2011-06-21)\n2. 7b97730b764 (upload-archive: allow user to turn off filters,\n   2011-06-21)\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/tar.txt  | 18 ++++++++++++++++++\n Documentation/git-archive.txt | 18 ------------------\n 2 files changed, 18 insertions(+), 18 deletions(-)\n\ndiff --git a/Documentation/config/tar.txt b/Documentation/config/tar.txt\nindex c68e294bbc5..894c1163bb9 100644\n--- a/Documentation/config/tar.txt\n+++ b/Documentation/config/tar.txt\n@@ -6,3 +6,21 @@ tar.umask::\n \tlinkgit:git-archive[1] for\n \tdetails. If `--remote` is used then only the configuration of\n \tthe remote repository takes effect.\n+\n+tar.<format>.command::\n+\tThis variable specifies a shell command through which the tar\n+\toutput generated by linkgit:git-archive[1] should be piped. The command\n+\tis executed using the shell with the generated tar file on its\n+\tstandard input, and should produce the final output on its\n+\tstandard output. Any compression-level options will be passed\n+\tto the command (e.g., `-9`).\n++\n+The `tar.gz` and `tgz` formats are defined automatically and use the\n+magic command `git archive gzip` by default, which invokes an internal\n+implementation of gzip.\n+\n+tar.<format>.remote::\n+\tIf true, enable the format for use by remote clients via\n+\tlinkgit:git-upload-archive[1]. Defaults to false for\n+\tuser-defined formats, but true for the `tar.gz` and `tgz`\n+\tformats.\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex bbb407d4975..268e797f03a 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -133,24 +133,6 @@ CONFIGURATION\n \n include::config/tar.txt[]\n \n-tar.<format>.command::\n-\tThis variable specifies a shell command through which the tar\n-\toutput generated by `git archive` should be piped. The command\n-\tis executed using the shell with the generated tar file on its\n-\tstandard input, and should produce the final output on its\n-\tstandard output. Any compression-level options will be passed\n-\tto the command (e.g., `-9`).\n-+\n-The `tar.gz` and `tgz` formats are defined automatically and use the\n-magic command `git archive gzip` by default, which invokes an internal\n-implementation of gzip.\n-\n-tar.<format>.remote::\n-\tIf true, enable the format for use by remote clients via\n-\tlinkgit:git-upload-archive[1]. Defaults to false for\n-\tuser-defined formats, but true for the `tar.gz` and `tgz`\n-\tformats.\n-\n [[ATTRIBUTES]]\n ATTRIBUTES\n ----------\n-- \n2.39.1.1392.g63e6d408230\n\n"},{"id":"471310","messageId":"patch-3.9-9d1a68b5282-20230202T093212Z-avarab@gmail.com","threadId":"59162","inReplyTo":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","subject":"[PATCH 3/9] archiver API: make the \"flags\" in \"struct archiver\" an enum","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T09:32:23Z","receivedAt":"2023-02-02T09:32:53Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Refactor the \"#define\" pattern in the archiver.h to use a new \"enum\narchiver_flags\". This isn't a functional change, but will make adding\nnew flags in a subsequent commit easier to reason about.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n archive.h | 10 ++++++----\n 1 file changed, 6 insertions(+), 4 deletions(-)\n\ndiff --git a/archive.h b/archive.h\nindex 08bed3ed3af..6b51288c2ed 100644\n--- a/archive.h\n+++ b/archive.h\n@@ -36,13 +36,15 @@ const char *archive_format_from_filename(const char *filename);\n \n /* archive backend stuff */\n \n-#define ARCHIVER_WANT_COMPRESSION_LEVELS 1\n-#define ARCHIVER_REMOTE 2\n-#define ARCHIVER_HIGH_COMPRESSION_LEVELS 4\n+enum archiver_flags {\n+\tARCHIVER_WANT_COMPRESSION_LEVELS = 1<<0,\n+\tARCHIVER_REMOTE = 1<<1,\n+\tARCHIVER_HIGH_COMPRESSION_LEVELS = 1<<2,\n+};\n struct archiver {\n \tconst char *name;\n \tint (*write_archive)(const struct archiver *, struct archiver_args *);\n-\tunsigned flags;\n+\tenum archiver_flags flags;\n \tchar *filter_command;\n };\n void register_archiver(struct archiver *);\n-- \n2.39.1.1392.g63e6d408230\n\n"},{"id":"471311","messageId":"patch-4.9-8bc1bfd1fe2-20230202T093212Z-avarab@gmail.com","threadId":"59162","inReplyTo":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","subject":"[PATCH 4/9] archive: omit the shell for built-in \"command\" filters","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T09:32:24Z","receivedAt":"2023-02-02T09:32:55Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Since the \"tar.<format.command\" interface was added in [1] we've\npromised to invoke the shell to run if e.g. \"gzip -cn\" is\nconfigured. That common format was then added as a default in [2].\n\nBut if we have no such configuration we can safely assume that the\nuser isn't expecting the \"gzip\" to be invoked via a shell, and we can\nskip the \"sh\" process.\n\nWe are intentionally not treating a configured\n\"tar.<format>.command=<cmd>\" where \"<cmd>\" is equivalent to our\nhardcoded \"<cmd>\" the same as when the same \"<cmd>\" is specified in\nthe config. If the user has configured e.g. \"gzip -cn\" they may be\nrelying on what the shell gives them over a direct execve() of \"gzip\".\n\nThis makes us marginally faster, but the real point is to make the\nerror handling easier to deal with. When we're using the shell we\ndon't know if e.g. the \"gzip\" we spawned fails as easily,\ni.e. \"start_command()\" won't fail, because we can find the \"sh\".\n\nA subsequent commit will tweak the default that [3] introduced to be a\nfallback instead, at which point we'll need this for correctness.\n\n1. 767cf4579f0 (archive: implement configurable tar filters, 2011-06-21)\n2. 0e804e09938 (archive: provide builtin .tar.gz filter, 2011-06-21)\n3. 4f4be00d302 (archive-tar: use internal gzip by default, 2022-06-15)\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/tar.txt |  3 +++\n archive-tar.c                | 17 +++++++++++++----\n archive.h                    |  1 +\n 3 files changed, 17 insertions(+), 4 deletions(-)\n\ndiff --git a/Documentation/config/tar.txt b/Documentation/config/tar.txt\nindex 894c1163bb9..5456fc617a2 100644\n--- a/Documentation/config/tar.txt\n+++ b/Documentation/config/tar.txt\n@@ -18,6 +18,9 @@ tar.<format>.command::\n The `tar.gz` and `tgz` formats are defined automatically and use the\n magic command `git archive gzip` by default, which invokes an internal\n implementation of gzip.\n++\n+The automatically defined commands do not invoke the shell, avoiding\n+the minor overhead of an extra sh(1) process.\n \n tar.<format>.remote::\n \tIf true, enable the format for use by remote clients via\ndiff --git a/archive-tar.c b/archive-tar.c\nindex f8fad2946ef..8c5de949c64 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -367,12 +367,13 @@ static struct archiver *find_tar_filter(const char *name, size_t len)\n }\n \n static int tar_filter_config(const char *var, const char *value,\n-\t\t\t     void *data UNUSED)\n+\t\t\t     void *data)\n {\n \tstruct archiver *ar;\n \tconst char *name;\n \tconst char *type;\n \tsize_t namelen;\n+\tint *configured = data;\n \n \tif (parse_config_key(var, \"tar\", &name, &namelen, &type) < 0 || !name)\n \t\treturn 0;\n@@ -388,6 +389,9 @@ static int tar_filter_config(const char *var, const char *value,\n \t\ttar_filters[nr_tar_filters++] = ar;\n \t}\n \n+\tif (configured && *configured)\n+\t\tar->flags |= ARCHIVER_COMMAND_FROM_CONFIG;\n+\n \tif (!strcmp(type, \"command\")) {\n \t\tif (!value)\n \t\t\treturn config_error_nonbool(var);\n@@ -495,8 +499,12 @@ static int write_tar_filter_archive(const struct archiver *ar,\n \tif (args->compression_level >= 0)\n \t\tstrbuf_addf(&cmd, \" -%d\", args->compression_level);\n \n-\tstrvec_push(&filter.args, cmd.buf);\n-\tfilter.use_shell = 1;\n+\tif (ar->flags & ARCHIVER_COMMAND_FROM_CONFIG) {\n+\t\tstrvec_push(&filter.args, cmd.buf);\n+\t\tfilter.use_shell = 1;\n+\t} else {\n+\t\tstrvec_split(&filter.args, cmd.buf);\n+\t}\n \tfilter.in = -1;\n \tfilter.silent_exec_failure = 1;\n \n@@ -526,13 +534,14 @@ static struct archiver tar_archiver = {\n void init_tar_archiver(void)\n {\n \tint i;\n+\tint configured = 1;\n \tregister_archiver(&tar_archiver);\n \n \ttar_filter_config(\"tar.tgz.command\", internal_gzip_command, NULL);\n \ttar_filter_config(\"tar.tgz.remote\", \"true\", NULL);\n \ttar_filter_config(\"tar.tar.gz.command\", internal_gzip_command, NULL);\n \ttar_filter_config(\"tar.tar.gz.remote\", \"true\", NULL);\n-\tgit_config(git_tar_config, NULL);\n+\tgit_config(git_tar_config, &configured);\n \tfor (i = 0; i < nr_tar_filters; i++) {\n \t\t/* omit any filters that never had a command configured */\n \t\tif (tar_filters[i]->filter_command)\ndiff --git a/archive.h b/archive.h\nindex 6b51288c2ed..9686b3b5cc1 100644\n--- a/archive.h\n+++ b/archive.h\n@@ -40,6 +40,7 @@ enum archiver_flags {\n \tARCHIVER_WANT_COMPRESSION_LEVELS = 1<<0,\n \tARCHIVER_REMOTE = 1<<1,\n \tARCHIVER_HIGH_COMPRESSION_LEVELS = 1<<2,\n+\tARCHIVER_COMMAND_FROM_CONFIG = 1<<3,\n };\n struct archiver {\n \tconst char *name;\n-- \n2.39.1.1392.g63e6d408230\n\n"},{"id":"471312","messageId":"patch-5.9-498037b2e65-20230202T093212Z-avarab@gmail.com","threadId":"59162","inReplyTo":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","subject":"[PATCH 5/9] archive-tar.c: move internal gzip implementation to a function","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T09:32:25Z","receivedAt":"2023-02-02T09:32:56Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Refactor the code added in 76d7602631a (archive-tar: add internal gzip\nimplementation, 2022-06-15) to call the magic \"git archive gzip\"\ncommand as a function.\n\nA subsequent commit will start using this as a fallback, but for now\nthere's no functional changes here.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n archive-tar.c | 43 +++++++++++++++++++++++++------------------\n 1 file changed, 25 insertions(+), 18 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 8c5de949c64..dfc133deac7 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -465,12 +465,33 @@ static void tgz_write_block(const void *data)\n \n static const char internal_gzip_command[] = \"git archive gzip\";\n \n-static int write_tar_filter_archive(const struct archiver *ar,\n-\t\t\t\t    struct archiver_args *args)\n+static int gzip_internally(const struct archiver *ar,\n+\t\t\t   struct archiver_args *args)\n {\n #if ZLIB_VERNUM >= 0x1221\n \tstruct gz_header_s gzhead = { .os = 3 }; /* Unix, for reproducibility */\n #endif\n+\tint r;\n+\n+\twrite_block = tgz_write_block;\n+\tgit_deflate_init_gzip(&gzstream, args->compression_level);\n+#if ZLIB_VERNUM >= 0x1221\n+\tif (deflateSetHeader(&gzstream.z, &gzhead) != Z_OK)\n+\t\tBUG(\"deflateSetHeader() called too late\");\n+#endif\n+\tgzstream.next_out = outbuf;\n+\tgzstream.avail_out = sizeof(outbuf);\n+\n+\tr = write_tar_archive(ar, args);\n+\n+\ttgz_deflate(Z_FINISH);\n+\tgit_deflate_end(&gzstream);\n+\treturn r;\n+}\n+\n+static int write_tar_filter_archive(const struct archiver *ar,\n+\t\t\t\t    struct archiver_args *args)\n+{\n \tstruct strbuf cmd = STRBUF_INIT;\n \tstruct child_process filter = CHILD_PROCESS_INIT;\n \tint r;\n@@ -478,22 +499,8 @@ static int write_tar_filter_archive(const struct archiver *ar,\n \tif (!ar->filter_command)\n \t\tBUG(\"tar-filter archiver called with no filter defined\");\n \n-\tif (!strcmp(ar->filter_command, internal_gzip_command)) {\n-\t\twrite_block = tgz_write_block;\n-\t\tgit_deflate_init_gzip(&gzstream, args->compression_level);\n-#if ZLIB_VERNUM >= 0x1221\n-\t\tif (deflateSetHeader(&gzstream.z, &gzhead) != Z_OK)\n-\t\t\tBUG(\"deflateSetHeader() called too late\");\n-#endif\n-\t\tgzstream.next_out = outbuf;\n-\t\tgzstream.avail_out = sizeof(outbuf);\n-\n-\t\tr = write_tar_archive(ar, args);\n-\n-\t\ttgz_deflate(Z_FINISH);\n-\t\tgit_deflate_end(&gzstream);\n-\t\treturn r;\n-\t}\n+\tif (!strcmp(ar->filter_command, internal_gzip_command))\n+\t\treturn gzip_internally(ar, args);\n \n \tstrbuf_addstr(&cmd, ar->filter_command);\n \tif (args->compression_level >= 0)\n-- \n2.39.1.1392.g63e6d408230\n\n"},{"id":"471313","messageId":"patch-6.9-34c7ce73099-20230202T093212Z-avarab@gmail.com","threadId":"59162","inReplyTo":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","subject":"[PATCH 6/9] archive: use \"gzip -cn\" for stability, not \"git archive gzip\"","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T09:32:26Z","receivedAt":"2023-02-02T09:32:57Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"This reverts and amends [1] so that we don't use \"git archive gzip\" by\ndefault, but only fall back on it when we cannot invoke \"gzip\".\n\nAs noted in the discussion at [2] that commit first released with\nv2.38.0 caused widespread breakage in the wild: Hosting sites like\nGitHub tend to offer a feature to download tagged releases as\narchives, which are generated by some variant of \"git archive\n--format=tgz\".\n\nDownstream distributors then tend to (re-)download those archives\nas-is, hardcoding their known hash their packaging systems. See [3],\n[4] etc. for reports of those systems breaking in conjunction with\n[1].\n\nThe reason for \"why\" is entirely missing from the commit message for\n[1], but as seen in the question about that in [5] and reply at [6] at\nthe time it was to \"avoid a run[time] dependency; the build/test\ndependency remains.\".\n\nIt's not immediately apparent what the second part of that is\nreferring to, as [1] also removed the \"GZIP\" prerequisite from some\ntests. The answer is that we still have other tests that need \"GZIP\",\nbut those are invoking \"gzip(1)\" explicitly.\n\nIn any case, whatever promises we make in the future about the\nstability and non-stability of \"git archive\" output (or the derived\ncompressed artifact), this amount of fallout isn't worth it to get to\nthe stated goal in [1].\n\nLet's instead default to \"gzip -cn\" again, but if we can't find it\nfall back on \"git archive gzip\". Note that we'll only fallback if that\n\"gzip -cn\" is ours, not if it comes from the user's own\n\"tar.<format>.command\" configuration.\n\nIf we do need the fallback we'll warn about it. No such warning will\nbe emitted if the user has explicitly asked for \"git archive gzip\".\n\n1. 4f4be00d302 (archive-tar: use internal gzip by default, 2022-06-15)\n2. https://lore.kernel.org/git/a812a664-67ea-c0ba-599f-cb79e2d96694@gmail.com/\n3. https://github.com/Homebrew/homebrew-core/issues/121877\n4. https://github.com/bazel-contrib/SIG-rules-authors/issues/11\n5. https://lore.kernel.org/git/220615.86wndhwt9a.gmgdl@evledraar.gmail.com/\n6. https://lore.kernel.org/git/3ed80afd-34b3-afd8-5ffb-0187a4475ee1@web.de/\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/tar.txt |  8 ++++++--\n archive-tar.c                | 20 +++++++++++++++-----\n 2 files changed, 21 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/config/tar.txt b/Documentation/config/tar.txt\nindex 5456fc617a2..37f24baa73a 100644\n--- a/Documentation/config/tar.txt\n+++ b/Documentation/config/tar.txt\n@@ -16,8 +16,12 @@ tar.<format>.command::\n \tto the command (e.g., `-9`).\n +\n The `tar.gz` and `tgz` formats are defined automatically and use the\n-magic command `git archive gzip` by default, which invokes an internal\n-implementation of gzip.\n+command `gzip -cn` by default. An internal gzip implementation can be\n+used by specifying the value `git archive gzip`.\n++\n+If 'gzip -cn' cannot be executed we'll fall back on `git archive gzip`\n+with a warning, if you don't have a gzip(1) and would like to use the\n+internal `git archive gzip` without warning, configure it explicitly.\n +\n The automatically defined commands do not invoke the shell, avoiding\n the minor overhead of an extra sh(1) process.\ndiff --git a/archive-tar.c b/archive-tar.c\nindex dfc133deac7..26efb911ebc 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -464,6 +464,7 @@ static void tgz_write_block(const void *data)\n }\n \n static const char internal_gzip_command[] = \"git archive gzip\";\n+static const char gzip_cn_command[] = \"gzip -cn\";\n \n static int gzip_internally(const struct archiver *ar,\n \t\t\t   struct archiver_args *args)\n@@ -494,12 +495,15 @@ static int write_tar_filter_archive(const struct archiver *ar,\n {\n \tstruct strbuf cmd = STRBUF_INIT;\n \tstruct child_process filter = CHILD_PROCESS_INIT;\n+\tint filter_is_gzip_cn = 0;\n \tint r;\n \n \tif (!ar->filter_command)\n \t\tBUG(\"tar-filter archiver called with no filter defined\");\n \n-\tif (!strcmp(ar->filter_command, internal_gzip_command))\n+\tif (!strcmp(ar->filter_command, gzip_cn_command))\n+\t\tfilter_is_gzip_cn = 1;\n+\telse if (!strcmp(ar->filter_command, internal_gzip_command))\n \t\treturn gzip_internally(ar, args);\n \n \tstrbuf_addstr(&cmd, ar->filter_command);\n@@ -515,8 +519,14 @@ static int write_tar_filter_archive(const struct archiver *ar,\n \tfilter.in = -1;\n \tfilter.silent_exec_failure = 1;\n \n-\tif (start_command(&filter) < 0)\n-\t\tdie_errno(_(\"unable to start '%s' filter\"), cmd.buf);\n+\tif (start_command(&filter) < 0) {\n+\t\tif (!filter_is_gzip_cn)\n+\t\t\tdie_errno(_(\"unable to start '%s' filter\"), cmd.buf);\n+\n+\t\twarning_errno(_(\"unable to start '%s' filter, falling back to '%s'\"),\n+\t\t\t      cmd.buf, internal_gzip_command);\n+\t\treturn gzip_internally(ar, args);\n+\t}\n \tclose(1);\n \tif (dup2(filter.in, 1) < 0)\n \t\tdie_errno(_(\"unable to redirect descriptor\"));\n@@ -544,9 +554,9 @@ void init_tar_archiver(void)\n \tint configured = 1;\n \tregister_archiver(&tar_archiver);\n \n-\ttar_filter_config(\"tar.tgz.command\", internal_gzip_command, NULL);\n+\ttar_filter_config(\"tar.tgz.command\", gzip_cn_command, NULL);\n \ttar_filter_config(\"tar.tgz.remote\", \"true\", NULL);\n-\ttar_filter_config(\"tar.tar.gz.command\", internal_gzip_command, NULL);\n+\ttar_filter_config(\"tar.tar.gz.command\", gzip_cn_command, NULL);\n \ttar_filter_config(\"tar.tar.gz.remote\", \"true\", NULL);\n \tgit_config(git_tar_config, &configured);\n \tfor (i = 0; i < nr_tar_filters; i++) {\n-- \n2.39.1.1392.g63e6d408230\n\n"},{"id":"471314","messageId":"patch-7.9-0c7a8aa59e8-20230202T093212Z-avarab@gmail.com","threadId":"59162","inReplyTo":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","subject":"[PATCH 7/9] test-lib.sh: add a lazy GZIP prerequisite","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T09:32:27Z","receivedAt":"2023-02-02T09:32:59Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Move the \"gzip --version\" lazy prerequisite added in [1] and\ncopy/pasted to another test in [2] to test-lib.sh. A subsequent commit\nwill add a third user, let's first stop duplicating it.\n\n1. 96174145fc3 (t5000: simplify gzip prerequisite checks, 2013-12-03)\n2. 6c213e863ae (http-backend: respect CONTENT_LENGTH for receive-pack,\n   2018-07-27)\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n t/t5000-tar-tree.sh                    | 2 --\n t/t5562-http-backend-content-length.sh | 2 --\n t/test-lib.sh                          | 4 ++++\n 3 files changed, 4 insertions(+), 4 deletions(-)\n\ndiff --git a/t/t5000-tar-tree.sh b/t/t5000-tar-tree.sh\nindex d4730481384..e1fa34bb828 100755\n--- a/t/t5000-tar-tree.sh\n+++ b/t/t5000-tar-tree.sh\n@@ -38,8 +38,6 @@ test_lazy_prereq TAR_NEEDS_PAX_FALLBACK '\n \t)\n '\n \n-test_lazy_prereq GZIP 'gzip --version'\n-\n get_pax_header() {\n \tfile=$1\n \theader=$2=\ndiff --git a/t/t5562-http-backend-content-length.sh b/t/t5562-http-backend-content-length.sh\nindex b68ec22d3fd..e83aa336fa8 100755\n--- a/t/t5562-http-backend-content-length.sh\n+++ b/t/t5562-http-backend-content-length.sh\n@@ -3,8 +3,6 @@\n test_description='test git-http-backend respects CONTENT_LENGTH'\n . ./test-lib.sh\n \n-test_lazy_prereq GZIP 'gzip --version'\n-\n verify_http_result() {\n \t# some fatal errors still produce status 200\n \t# so check if there is the error message\ndiff --git a/t/test-lib.sh b/t/test-lib.sh\nindex 01e88781dd2..33bb9fe991f 100644\n--- a/t/test-lib.sh\n+++ b/t/test-lib.sh\n@@ -1922,6 +1922,10 @@ test_lazy_prereq LONG_IS_64BIT '\n test_lazy_prereq TIME_IS_64BIT 'test-tool date is64bit'\n test_lazy_prereq TIME_T_IS_64BIT 'test-tool date time_t-is64bit'\n \n+test_lazy_prereq GZIP '\n+\tgzip --version\n+'\n+\n test_lazy_prereq CURL '\n \tcurl --version\n '\n-- \n2.39.1.1392.g63e6d408230\n\n"},{"id":"471315","messageId":"patch-8.9-62c796da4e2-20230202T093212Z-avarab@gmail.com","threadId":"59162","inReplyTo":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","subject":"[PATCH 8/9] archive tests: test for \"gzip -cn\" and \"git archive gzip\" stability","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T09:32:28Z","receivedAt":"2023-02-02T09:33:00Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"If our test suite is instrumented to run the first \"test_cmp_bin\" in\n\"test_done\" it'll mostly pass, but fail on a few tests, such as\n\"t5319-multi-pack-index.sh\". Those tests reveal edge cases where the\noutput of \"gzip -cn\" is different than that of \"git archive gzip\" for\nthe same input.\n\nLet's extract a minimal version of the part of\n\"t5319-multi-pack-index.sh\" which triggers it, and add a test for\narchival stability.\n\nWhatever we ultimately decide to promise when it comes to this\nstability (see [1]) it'll be better to go into any behavior difference\nknowing that's what we're about to do, rather than discover widespread\nbreakage due to already released Git versions.\n\nThe \"GZIP_TRIVIALLY_STABLE\" code here is added because on OSX even a\ntrivial *.tgz generated by the two methods will be different.\n\n1. https://lore.kernel.org/git/a812a664-67ea-c0ba-599f-cb79e2d96694@gmail.com/\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n t/t5005-archive-stability.sh | 70 ++++++++++++++++++++++++++++++++++++\n 1 file changed, 70 insertions(+)\n create mode 100755 t/t5005-archive-stability.sh\n\ndiff --git a/t/t5005-archive-stability.sh b/t/t5005-archive-stability.sh\nnew file mode 100755\nindex 00000000000..c7532886920\n--- /dev/null\n+++ b/t/t5005-archive-stability.sh\n@@ -0,0 +1,70 @@\n+#!/bin/sh\n+\n+test_description='git archive stabilty'\n+\n+TEST_PASSES_SANITIZE_LEAK=true\n+. ./test-lib.sh\n+\n+create_archive_file_with_config () {\n+\tlocal file=\"$1\" &&\n+\tlocal config=\"$2\" &&\n+\tshift 2 &&\n+\n+\ttest_when_finished \"rm -rf \\\"$file\\\"\" &&\n+\tgit -c tar.tgz.command=\"$config\" archive -o \"$file\" HEAD\n+}\n+\n+setup_gzip_vs_git_archive_gzip () {\n+\tcreate_archive_file_with_config \"expect.tgz\" \"gzip -cn\" &&\n+\tcreate_archive_file_with_config \"actual.tgz\" \"git archive gzip\"\n+}\n+\n+test_lazy_prereq GZIP_TRIVIALLY_STABLE '\n+\tgit clone \"$TRASH_DIRECTORY\" . &&\n+\ttest_commit P &&\n+\tsetup_gzip_vs_git_archive_gzip &&\n+\ttest_cmp_bin expect.tgz actual.tgz\n+'\n+\n+if ! test_have_prereq GZIP_TRIVIALLY_STABLE\n+then\n+\tskip_all='skipping gzip v.s. git archive gzip tests, even trivial content differs'\n+\ttest_done\n+fi\n+\n+# The first test_expect_success is after the \"skip_all\" so we'll get\n+# the skip summary in prove(1) output.\n+test_expect_success 'setup' '\n+\ttest_commit A\n+'\n+\n+test_expect_success GZIP '\"gzip -cn\" and v.s. \"git archive gzip\" produce the same output still' '\n+\tsetup_gzip_vs_git_archive_gzip &&\n+\ttest_cmp_bin expect.tgz actual.tgz\n+'\n+\n+generate_objects () {\n+\ti=$1\n+\tiii=$(printf '%03i' $i)\n+\t{\n+\t\techo $iii &&\n+\t\ttest-tool genrandom \"$iii\" 8192\n+\t} >file_$iii &&\n+\tgit update-index --add file_$iii\n+}\n+\n+test_expect_success 'create objects with (stable) random data' '\n+\ttest_commit initial &&\n+\tfor i in $(test_seq 1 5)\n+\tdo\n+\t\tgenerate_objects $i || return 1\n+\tdone &&\n+\tgit commit -m\"add objects\"\n+'\n+\n+test_expect_success GZIP '\"gzip -cn\" and v.s. \"git archive gzip\" have differing output' '\n+\tsetup_gzip_vs_git_archive_gzip &&\n+\t! test_cmp_bin expect.tgz actual.tgz\n+'\n+\n+test_done\n-- \n2.39.1.1392.g63e6d408230\n\n"},{"id":"471316","messageId":"patch-9.9-b40833b2168-20230202T093212Z-avarab@gmail.com","threadId":"59162","inReplyTo":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","subject":"[PATCH 9/9] git archive docs: document output non-stability","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T09:32:29Z","receivedAt":"2023-02-02T09:33:02Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"There's an ongoing discussion about the output stability of \"git\narchive\"[1] as a follow-up to the incident GitHub experienced when\nupgrading to v2.38.0[2].\n\nIn a preceding commit we reverted the immediate cause of that\nincident, which was that we'd moved away from \"gzip -cn\" as the\ndefault compression method in favor of the internal \"git archive gzip\"\nin [3].\n\nLet's follow that up by documenting the non-promises we've always\nmaintained with regards to \"git archive\"'s output stability. We may\nwant to make stronger promises in this area, but this change avoids\naddressing that question.\n\nInstead we're discussing that we've changed this in the past, aren't\nchanging it willy-nilly, but it may change again in the future. The\nonly new promise here that we haven't explicitly maintained\nhistorically is that we're promising to forever shell out to the\nsystem's \"gzip\" by default. Whether it produces stable output once\nthat happens we leave up to the \"gzip\" tool.\n\nWe're also discussing the caveats & differences in output with with\nSHA-1 and SHA-256 repositories, and trying to steer users towards more\nstable alternatives. First by using \"git verify-tag\" and the like to\nverify releases, and if they really must checksum generated output, to\nencourage them to at least checksum the \"tar\" output contained within\nthe compressed output, not the compressed output itself.\n\n1. https://lore.kernel.org/git/a812a664-67ea-c0ba-599f-cb79e2d96694@gmail.com/\n2. https://github.com/orgs/community/discussions/45830\n3. 4f4be00d302 (archive-tar: use internal gzip by default, 2022-06-15)\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/git-archive.txt | 70 ++++++++++++++++++++++++++++++++++-\n 1 file changed, 69 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex 268e797f03a..78f1b033cb7 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -14,6 +14,7 @@ SYNOPSIS\n \t      [--remote=<repo> [--exec=<git-upload-archive>]] <tree-ish>\n \t      [<path>...]\n \n+[[DESCRIPTION]]\n DESCRIPTION\n -----------\n Creates an archive of the specified format containing the tree\n@@ -28,7 +29,7 @@ case the commit time as recorded in the referenced commit object is\n used instead.  Additionally the commit ID is stored in a global\n extended pax header if the tar format is used; it can be extracted\n using 'git get-tar-commit-id'. In ZIP files it is stored as a file\n-comment.\n+comment. See the <<STABILITY,OUTPUT STABILITY>> section below.\n \n OPTIONS\n -------\n@@ -202,6 +203,73 @@ EXAMPLES\n \tYou can use it specifying `--format=tar.xz`, or by creating an\n \toutput file like `-o foo.tar.xz`.\n \n+[[STABILITY]]\n+OUTPUT STABILITY\n+----------------\n+\n+The output of 'git archive' is not guaranteed to be stable, and may\n+change between versions.\n+\n+There are many valid ways to encode the same data in the tar format\n+itself. For non-`tar` arguments to the `--format` option we rely on\n+external tools (or libraries) for compressing the output we generate.\n+\n+The `tar` format contains the commit ID in the pax header (see the\n+<<DESCRIPTION>> section above). A repository that's been migrated from\n+SHA-1 to SHA-256 will therefore have different `tar` output for the\n+\"same\" commit. See `extension.objectFormat` in linkgit:git-config[1].\n+\n+Instead of relying on the output of `git archive`, you should prefer\n+to stick to git's own transport protocols, and e.g. validate releases\n+with linkgit:git-tag[1]'s `--verify` option.\n+\n+Despite the output of `git archive` having never been promised to be\n+stable, various users in the wild have come to rely on that being the\n+case.\n+\n+Most notably, large hosting providers provide a way to download a\n+given tagged release as a `git archive`. Some downstream tools then\n+expect the content of that archive to be stable. When that's changed\n+widespread breakage has been observed, see\n+https://github.com/orgs/community/discussions/45830 for one such case.\n+\n+While we won't promise that the output won't change in the future, we\n+are aware of these users, and will try to avoid changing it\n+willy-nilly. Furthermore, we make the following promises:\n+\n+* The default gzip compression tool will continue to be gzip(1). If\n+  you rely on this being e.g. GNU gzip for the purposes of stability,\n+  it's up to you to ensure that its output is stable across\n+  versions.\n++\n+\n+We in turn promise to not e.g. make the internal \"git archive gzip\"\n+implementation the default, as it produces different ouput than\n+gzip(1) in some case.\n+\n+* We will do our best not to change the \"tar\" output itself, but won't\n+  promise that we're never going to change it.\n++\n+If you must avoid using \"git\" itself for the tree validation, you\n+should be checksumming the uncompressed \"tar\" output, not e.g. the\n+compressed \"tgz\" output.\n++\n+\n+This ensures that you're only relying on the output emitted by git\n+itself, and avoiding the additional dependency on external\n+compression.\n++\n+See\n+https://git.kernel.org/pub/scm/linux/kernel/git/mricon/korg-helpers.git/tree/get-verified-tarball\n+for an implementation of that workflow.\n+\n+* We promise that a given version of git will emit stable \"tar\" output\n+  for the same tree ID (but not commit ID, see the discussion in the\n+  <<DESCRIPTION>> section above).\n++\n+While you shouldn't assume that different versions of git will emit\n+the same output, you can assume (e.g. for the purposes of caching)\n+that a given version's output is stable.\n \n SEE ALSO\n --------\n-- \n2.39.1.1392.g63e6d408230\n\n"},{"id":"471346","messageId":"Y9uPhPnNFlCju8Fo@tapette.crustytoothpaste.net","threadId":"59162","inReplyTo":"patch-9.9-b40833b2168-20230202T093212Z-avarab@gmail.com","subject":"Re: [PATCH 9/9] git archive docs: document output non-stability","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-02-02T10:25:08Z","receivedAt":"2023-02-02T10:25:14Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-02-02 at 09:32:29, Ævar Arnfjörð Bjarmason wrote:\n> +[[STABILITY]]\n> +OUTPUT STABILITY\n> +----------------\n> +\n> +The output of 'git archive' is not guaranteed to be stable, and may\n> +change between versions.\n> +\n> +There are many valid ways to encode the same data in the tar format\n> +itself. For non-`tar` arguments to the `--format` option we rely on\n> +external tools (or libraries) for compressing the output we generate.\n> +\n> +The `tar` format contains the commit ID in the pax header (see the\n> +<<DESCRIPTION>> section above). A repository that's been migrated from\n> +SHA-1 to SHA-256 will therefore have different `tar` output for the\n> +\"same\" commit. See `extension.objectFormat` in linkgit:git-config[1].\n> +\n> +Instead of relying on the output of `git archive`, you should prefer\n> +to stick to git's own transport protocols, and e.g. validate releases\n> +with linkgit:git-tag[1]'s `--verify` option.\n> +\n> +Despite the output of `git archive` having never been promised to be\n> +stable, various users in the wild have come to rely on that being the\n> +case.\n> +\n> +Most notably, large hosting providers provide a way to download a\n> +given tagged release as a `git archive`. Some downstream tools then\n> +expect the content of that archive to be stable. When that's changed\n> +widespread breakage has been observed, see\n> +https://github.com/orgs/community/discussions/45830 for one such case.\n> +\n> +While we won't promise that the output won't change in the future, we\n> +are aware of these users, and will try to avoid changing it\n> +willy-nilly. Furthermore, we make the following promises:\n> +\n> +* The default gzip compression tool will continue to be gzip(1). If\n> +  you rely on this being e.g. GNU gzip for the purposes of stability,\n> +  it's up to you to ensure that its output is stable across\n> +  versions.\n> ++\n> +\n> +We in turn promise to not e.g. make the internal \"git archive gzip\"\n> +implementation the default, as it produces different ouput than\n> +gzip(1) in some case.\n\nI think this is fine up to here.\n\n> +* We will do our best not to change the \"tar\" output itself, but won't\n> +  promise that we're never going to change it.\n> ++\n> +If you must avoid using \"git\" itself for the tree validation, you\n> +should be checksumming the uncompressed \"tar\" output, not e.g. the\n> +compressed \"tgz\" output.\n> ++\n\nI don't think I want to state this, because it implies that the changes\nI made that broke kernel.org (making tar.umask apply to pax headers)\nwouldn't have been allowed.  We should probably just state that \"we\nwon't promise that the tar output won't change between versions\". Maybe,\n\"We won't change the tar output needlessly, but it may change from time\nto time.\"  That is, we won't be \"let's change the format just to mix it\nup for users\", but if there's a valuable patch that could be applied,\nthen we might well take it.\n\nAs I said, it's my goal to provide more concrete guarantees in a future\npatch, probably this weekend.\n\n> +* We promise that a given version of git will emit stable \"tar\" output\n> +  for the same tree ID (but not commit ID, see the discussion in the\n> +  <<DESCRIPTION>> section above).\n\nI think that section contradicts this.  The tree version uses the\ncurrent timestamp, which would make the archive change based on the time\nof day.\n\n> +While you shouldn't assume that different versions of git will emit\n> +the same output, you can assume (e.g. for the purposes of caching)\n> +that a given version's output is stable.\n\nUnfortunately, this isn't actually true if someone uses export-subst.\nThat's because adding unrelated objects can increase the length of\nabbreviations, and then the tar contents can be different.  I've\nactually seen this in the wild.\n\nModulo that, yes, I agree with this.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"471348","messageId":"230202.86ilgkpcxt.gmgdl@evledraar.gmail.com","threadId":"59162","inReplyTo":"Y9uPhPnNFlCju8Fo@tapette.crustytoothpaste.net","subject":"Re: [PATCH 9/9] git archive docs: document output non-stability","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-02T10:30:45Z","receivedAt":"2023-02-02T10:41:07Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Feb 02 2023, brian m. carlson wrote:\n\n>> +* We will do our best not to change the \"tar\" output itself, but won't\n>> +  promise that we're never going to change it.\n>> ++\n>> +If you must avoid using \"git\" itself for the tree validation, you\n>> +should be checksumming the uncompressed \"tar\" output, not e.g. the\n>> +compressed \"tgz\" output.\n>> ++\n>\n> I don't think I want to state this, because it implies that the changes\n> I made that broke kernel.org (making tar.umask apply to pax headers)\n> wouldn't have been allowed.\n\nI don't see how \"we'll do our best, but it might change\" precludes that...\n\n> We should probably just state that \"we\n> won't promise that the tar output won't change between versions\". Maybe,\n\n...but it sounds like you'd like this \"softer\" promise. I think it's\nsaying the same, but picked the \"we'll try not to\" wording because I\nthink it more accurately reflects reality, but...\n\n> \"We won't change the tar output needlessly, but it may change from time\n> to time.\"  That is, we won't be \"let's change the format just to mix it\n> up for users\", but if there's a valuable patch that could be applied,\n> then we might well take it.\n\n...here we're back (at least per my reading) to basically what my\nproposed patch said. I'm happy to improve/change the wording, but I'm\nconfused about the \"because it implies\" part you noted.\n\n> As I said, it's my goal to provide more concrete guarantees in a future\n> patch, probably this weekend.\n\nI think that would be great, but also think that if we're going to make\nnew guarantees it's probably best applied on top of a series such as\nthis, which aside from the reverting back to gzip as the default\nattempts to clarify the status quo.\n>\n>> +* We promise that a given version of git will emit stable \"tar\" output\n>> +  for the same tree ID (but not commit ID, see the discussion in the\n>> +  <<DESCRIPTION>> section above).\n>\n> I think that section contradicts this.  The tree version uses the\n> current timestamp, which would make the archive change based on the time\n> of day.\n\nThanks! It's referring back to the previous discussion, but I managed to\nsomehow get the tree & commit cases reversed.\t\n\n>> +While you shouldn't assume that different versions of git will emit\n>> +the same output, you can assume (e.g. for the purposes of caching)\n>> +that a given version's output is stable.\n>\n> Unfortunately, this isn't actually true if someone uses export-subst.\n> That's because adding unrelated objects can increase the length of\n> abbreviations, and then the tar contents can be different.  I've\n> actually seen this in the wild.\n>\n> Modulo that, yes, I agree with this.\n\nI didn't know about the export-subst case, I'll add that caveat in\nthere. Thanks!\n"},{"id":"471367","messageId":"771a98ca-9540-ad4e-dfba-9d304e1dff09@dunelm.org.uk","threadId":"59162","inReplyTo":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","subject":"Re: [PATCH 0/9] git archive: use gzip again by default, document output stabilty","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2023-02-02T16:17:09Z","receivedAt":"2023-02-02T16:17:18Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Ævar\n\nOn 02/02/2023 09:32, Ævar Arnfjörð Bjarmason wrote:\n> As reported in\n> https://lore.kernel.org/git/a812a664-67ea-c0ba-599f-cb79e2d96694@gmail.com/\n> changing the default \"tgz\" output method of from \"gzip(1)\" to our\n> internal \"git archive gzip\" (using zlib ) broke things for users in\n> the wild that assume that the \"git archive\" output is stable, most\n> notably GitHub: https://github.com/orgs/community/discussions/45830\n >\n> Leaving aside the larger question of whether we're going to promise\n> output stability for \"git archive\" in general, the motivation for that\n> change was to have a working compression method on systems that lacked\n> a gzip(1).\n\nAs I recall the reduction in cpu time used to create a compressed \narchive was a factor in making it the default.\n\n> As the disruption of changing the default isn't worth it, let's use\n> gzip(1) again by default, and only fall back on the new \"git archive\n> gzip\" if it isn't available.\n\nPlaying devil's advocate for a moment as we're not going to promise that \nthe compressed output of \"git archive\" will be stable in the future \nperhaps we should use this breakage as an opportunity to highlight that \nto users and to advertize the config setting that allows them to use \ngzip for compressing archives. Reverting the change gives the misleading \nimpression that we're making a commitment to keeping the output stable. \nThe focus of this thread seems to be the problems relating to github \nwhich they have already addressed.\n\nI think there is general agreement that it is not practical to promise \nthat the compressed output of \"git archive\" is stable so maybe it is \nbetter to make that clear now while users can work around it in the \nshort term with a config setting rather than waiting until we're faced \nwith some security or other issue that forces a change to the output \nwhich users cannot work around so easily.\n\nBest Wishes\n\nPhillip\n\n\n> The later parts of this series then document and test for the output\n> stability of the command.\n> \n> We're not promising anything new there, except that we now promise\n> that we're going to use \"gzip\" as the default compressor, but that\n> it's up to that command to be stable, should the user desire output\n> stability.\n> \n> The documentation discusses the various caveats involved, suggests\n> alternatives to checksumming compressed archives, but in the end notes\n> what's been the policy so far: We're not promising that the \"tar\"\n> output is going to be stable.\n> \n> The early parts of this series (1-2/9) are clean-up for existing\n> config drift, as later in the series we'll otherwise need to change\n> the divergent config documentation in two places.\n> \n> CI & branch for this at:\n> https://github.com/avar/git/tree/avar/archive-internal-gzip-not-the-default\n> \n> Ævar Arnfjörð Bjarmason (9):\n>    archive & tar config docs: de-duplicate configuration section\n>    git config docs: document \"tar.<format>.{command,remote}\"\n>    archiver API: make the \"flags\" in \"struct archiver\" an enum\n>    archive: omit the shell for built-in \"command\" filters\n>    archive-tar.c: move internal gzip implementation to a function\n>    archive: use \"gzip -cn\" for stability, not \"git archive gzip\"\n>    test-lib.sh: add a lazy GZIP prerequisite\n>    archive tests: test for \"gzip -cn\" and \"git archive gzip\" stability\n>    git archive docs: document output non-stability\n> \n>   Documentation/config/tar.txt           | 29 +++++++-\n>   Documentation/git-archive.txt          | 96 +++++++++++++++++++-------\n>   archive-tar.c                          | 78 ++++++++++++++-------\n>   archive.h                              | 11 +--\n>   t/t5000-tar-tree.sh                    |  2 -\n>   t/t5005-archive-stability.sh           | 70 +++++++++++++++++++\n>   t/t5562-http-backend-content-length.sh |  2 -\n>   t/test-lib.sh                          |  4 ++\n>   8 files changed, 231 insertions(+), 61 deletions(-)\n>   create mode 100755 t/t5005-archive-stability.sh\n> \n"},{"id":"471369","messageId":"xmqq5yckvxtb.fsf@gitster.g","threadId":"59162","inReplyTo":"cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com","subject":"Re: [PATCH 0/9] git archive: use gzip again by default, document output stabilty","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-02-02T16:25:52Z","receivedAt":"2023-02-02T16:26:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> As the disruption of changing the default isn't worth it, let's use\n> gzip(1) again by default, and only fall back on the new \"git archive\n> gzip\" if it isn't available.\n\nIt perhaps is OK, and lets us answer \"ugh, the compressed output of\n'git archive' is unstable again\" with \"we didn't change anything,\nperhaps you changed your gzip(1)?\" when they fix bugs or improve\ncompression or whatever.  Of course that is not an overall win for\nthe end users, but in the short term until gzip gets such a change,\nwe would presumably get the \"same\" output as before.\n"},{"id":"471370","messageId":"xmqq1qn8vxej.fsf@gitster.g","threadId":"59162","inReplyTo":"Y9uPhPnNFlCju8Fo@tapette.crustytoothpaste.net","subject":"Re: [PATCH 9/9] git archive docs: document output non-stability","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-02-02T16:34:44Z","receivedAt":"2023-02-02T16:34:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n>> +* We will do our best not to change the \"tar\" output itself, but won't\n>> +  promise that we're never going to change it.\n>> ++\n>> +If you must avoid using \"git\" itself for the tree validation, you\n>> +should be checksumming the uncompressed \"tar\" output, not e.g. the\n>> +compressed \"tgz\" output.\n>> ++\n>\n> I don't think I want to state this, because it implies that the changes\n> I made that broke kernel.org (making tar.umask apply to pax headers)\n> wouldn't have been allowed.  We should probably just state that \"we\n> won't promise that the tar output won't change between versions\". Maybe,\n> \"We won't change the tar output needlessly, but it may change from time\n> to time.\"  That is, we won't be \"let's change the format just to mix it\n> up for users\", but if there's a valuable patch that could be applied,\n> then we might well take it.\n\nI agree with you.  Giving \"will do our best not to\" is still too\nstrong for that.  We won't change the format willy-nilly but when\nthere is a good reason to do so, we should be able to fix or improve\nthe output.\n\n>> +While you shouldn't assume that different versions of git will emit\n>> +the same output, you can assume (e.g. for the purposes of caching)\n>> +that a given version's output is stable.\n>\n> Unfortunately, this isn't actually true if someone uses export-subst.\n> That's because adding unrelated objects can increase the length of\n> abbreviations, and then the tar contents can be different.  I've\n> actually seen this in the wild.\n\n\"subst\" is certainly an issue, especially when the substitution is\nunstable.\n\nThere shouldn't be cross platform differences to break bit-for-bit\nstability at least for \"tar\" format, as we do not rely on any\nexternal library.  Can we say the same for \"zip\"?  I thought we\nthrow the blob at git_deflate_*() so the exact bitstream is up to\nthe libz implementation?\n"},{"id":"471371","messageId":"xmqqwn50uil7.fsf@gitster.g","threadId":"59162","inReplyTo":"771a98ca-9540-ad4e-dfba-9d304e1dff09@dunelm.org.uk","subject":"Re: [PATCH 0/9] git archive: use gzip again by default, document output stabilty","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-02-02T16:40:04Z","receivedAt":"2023-02-02T16:40:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n> ... Reverting the change\n> gives the misleading impression that we're making a commitment to\n> keeping the output stable. The focus of this thread seems to be the\n> problems relating to github which they have already addressed.\n>\n> I think there is general agreement that it is not practical to promise\n> that the compressed output of \"git archive\" is stable so maybe it is\n> better to make that clear now while users can work around it in the\n> short term with a config setting rather than waiting until we're faced\n> with some security or other issue that forces a change to the output\n> which users cannot work around so easily.\n\nI love to see somebody else play the devil's advocate role.  Thanks\nfor all of the above.\n\n"},{"id":"471386","messageId":"de8f1e338e6ee99cd3ee06b16f1edbce@ameretat.dev","threadId":"59162","inReplyTo":"771a98ca-9540-ad4e-dfba-9d304e1dff09@dunelm.org.uk","subject":"Re: [PATCH 0/9] git archive: use gzip again by default, document output stabilty","fromName":"Raymond E. Pasco","fromEmail":"ray@ameretat.dev","sentAt":"2023-02-02T19:23:46Z","receivedAt":"2023-02-02T19:23:53Z","isPatch":true,"sender":{"key":"ray@ameretat.dev","avatar":"https://avatars.githubusercontent.com/u/115765?v=4"},"body":"February 2, 2023 11:17 AM, \"Phillip Wood\" <phillip.wood123@gmail.com> wrote:\n> Playing devil's advocate for a moment as we're not going to promise that the compressed output of\n> \"git archive\" will be stable in the future perhaps we should use this breakage as an opportunity to\n> highlight that to users and to advertize the config setting that allows them to use gzip for\n> compressing archives. Reverting the change gives the misleading impression that we're making a\n> commitment to keeping the output stable. The focus of this thread seems to be the problems relating\n> to github which they have already addressed.\n> \n> I think there is general agreement that it is not practical to promise that the compressed output\n> of \"git archive\" is stable so maybe it is better to make that clear now while users can work around\n> it in the short term with a config setting rather than waiting until we're faced with some security\n> or other issue that forces a change to the output which users cannot work around so easily.\n\nReverting to the behavior of \"use some arbitrary gzip from $PATH\" would\nbe a poor decision whether or not git were willing to make some\ncommitment to gzip stability, because Git does not control arbitrary\ngzips on the user's $PATH. If Git did want to promise gzip stability, it \ncould only start from something like the current internal implementation\nalong with a vendored zlib; if it doesn't, as appears to be the case, \nthen the internal implementation is superior for the other reasons \nalready discussed.\n\nIf the user wants to depend on a particular gzip executable they supply, \nthis configuration knob already exists for them.\n\nSince there is no guarantee of stability, but there has been a popular \nmisconception that there is some such guarantee (e.g., [1]), some kind \nof STABILITY section describing how there isn't any and suggesting ways\nthe user can attain more stability via configuration seems to be a good\nidea.\n\n[1]: https://lists.reproducible-builds.org/pipermail/rb-general/2021-October/002422.html\n"},{"id":"471397","messageId":"Y9wo4iH2crlt26+d@kitenet.net","threadId":"59162","inReplyTo":"Y9q129WbseimgeBS@mit.edu","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Joey Hess","fromEmail":"id@joeyh.name","sentAt":"2023-02-02T21:19:30Z","receivedAt":"2023-02-02T21:25:46Z","isPatch":false,"sender":{"key":"id@joeyh.name","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"In my opinion as the original developer of pristine-tar, it's too\ncomplicated to be usefully used by git. The problem it solves is of a\nlarger scope than the problem git has here. (I hope.)\n\nDeveloping pristine-tar did entail much investigation of past changes in\ncompressor outputs. I know that gzip's output has sometimes not been\ndeterministic as recently as 2012, see for example\nhttps://git.savannah.gnu.org/cgit/gzip.git/commit/?id=0a284baeaedca68017f46d2646e4\n\n-- \nsee shy jo\n"},{"id":"471402","messageId":"Y9xAv1reHJRj7iKA@tapette.crustytoothpaste.net","threadId":"59162","inReplyTo":"xmqqh6w5x8i8.fsf@gitster.g","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-02-02T23:01:31Z","receivedAt":"2023-02-02T23:01:58Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-02-01 at 23:37:19, Junio C Hamano wrote:\n> \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n> \n> > I don't think a blurb is necessary, but you're basically underscoring\n> > the problem, which is that nobody is willing to promise that compression\n> > is consistent, but yet people want to rely on that fact.  I'm willing to\n> > write and implement a consistent tar spec and to guarantee compatibility\n> > with that, but the tension here is that people also want gzip to never\n> > change its byte format ever, which frankly seems unrealistic without\n> > explicit guarantees.  Maybe the authors will agree to promise that, but\n> > it seems unlikely.\n> \n> Just to step back a bit, where does the distinction between\n> guaranteeing the tar format stability and gzip compressed bitstream\n> stability come from?  At both levels, the same thing can be\n> expressed in multiple different ways, I think, but spelling out how\n> exactly the compressor compresses is more involved than spelling out\n> how entries in a tar archive is ordered and each entry is expressed,\n> or something?\n\nYes, at least with my understanding about how gzip and compression in\ngeneral work.\n\nThe tar format (and the pax format which builds on it) can mostly be\nrestricted by explaining what data is to be included in the pax and tar\nheaders and how it is to be formatted.  If we say, we will always write\nsuch and such information in the pax header and sort the keys, and we\nwrite such and such information in the tar header, then the format is\ncompletely deterministic, and we can make nice guarantees.\n\nMy understanding about how Lempel-Ziv-based compression algorithms work\nis that there's a lot more freedom to decide how best to compress things\nand that there isn't always a logical obvious choice, but I will admit\nmy understanding is relatively limited.  If someone thinks we can\neffectively succeed in supporting compression more than just relying on\ngzip, I would be delighted to be shown to be wrong.\n\n> > That would probably break things, because gzip is GPLv3, and we'd need\n> > to ship a much older GPLv2 gzip, which would probably differ from the\n> > current behaviour, and might also have some security problems.\n> \n> Yup, security issues may make bit-for-bit-stability unrealistic.\n> IIRC, the last time we had discussion on this topic, we settled\n> on stability across the same version of Git (i.e. deterministic\n> result)?\n\nYes, I think that's what we agreed.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"471410","messageId":"01a901d93760$c690d970$53b28c50$@nexbridge.com","threadId":"59162","inReplyTo":"Y9xAv1reHJRj7iKA@tapette.crustytoothpaste.net","subject":"RE: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2023-02-02T23:47:50Z","receivedAt":"2023-02-02T23:58:32Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 2, 2023 6:02 PM, brian m. carlson wrote:\n>On 2023-02-01 at 23:37:19, Junio C Hamano wrote:\n>> \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n>>\n>> > I don't think a blurb is necessary, but you're basically\n>> > underscoring the problem, which is that nobody is willing to promise\n>> > that compression is consistent, but yet people want to rely on that\n>> > fact.  I'm willing to write and implement a consistent tar spec and\n>> > to guarantee compatibility with that, but the tension here is that\n>> > people also want gzip to never change its byte format ever, which\n>> > frankly seems unrealistic without explicit guarantees.  Maybe the\n>> > authors will agree to promise that, but it seems unlikely.\n>>\n>> Just to step back a bit, where does the distinction between\n>> guaranteeing the tar format stability and gzip compressed bitstream\n>> stability come from?  At both levels, the same thing can be expressed\n>> in multiple different ways, I think, but spelling out how exactly the\n>> compressor compresses is more involved than spelling out how entries\n>> in a tar archive is ordered and each entry is expressed, or something?\n>\n>Yes, at least with my understanding about how gzip and compression in general\n>work.\n>\n>The tar format (and the pax format which builds on it) can mostly be restricted by\n>explaining what data is to be included in the pax and tar headers and how it is to be\n>formatted.  If we say, we will always write such and such information in the pax\n>header and sort the keys, and we write such and such information in the tar header,\n>then the format is completely deterministic, and we can make nice guarantees.\n>\n>My understanding about how Lempel-Ziv-based compression algorithms work is that\n>there's a lot more freedom to decide how best to compress things and that there\n>isn't always a logical obvious choice, but I will admit my understanding is relatively\n>limited.  If someone thinks we can effectively succeed in supporting compression\n>more than just relying on gzip, I would be delighted to be shown to be wrong.\n\nThe nice part about gzip is that it is generally available on virtually all platforms (or can be easily obtained). Other compression forms, like bz2, which sometimes produces more dense compression, are not necessarily available. Availability is something I would be worried about (clone and checkout failures).\n\nTar formats are also to be used carefully. Not all platform implementations of tar support all variants. \"ustar\" is fairly common but there are others that are not. Interoperability needs to be the biggest factor in this decision, IMHO, rather than compression rates.\n\nThe alternative is having git supply its own implementation, but that is a longer term migration problem, resembling the SHA-256 migration.\n\n>\n>> > That would probably break things, because gzip is GPLv3, and we'd\n>> > need to ship a much older GPLv2 gzip, which would probably differ\n>> > from the current behaviour, and might also have some security problems.\n>>\n>> Yup, security issues may make bit-for-bit-stability unrealistic.\n>> IIRC, the last time we had discussion on this topic, we settled on\n>> stability across the same version of Git (i.e. deterministic result)?\n\nIn the old days, it was export concerns. Fortunately, git never really hit those in a post-2007 timeframe. I would not bank on this issue staying off the table.\n\n--Randall\n\n"},{"id":"471416","messageId":"Y9yHWFh0ijwrqhOX@mit.edu","threadId":"59162","inReplyTo":"Y9wo4iH2crlt26+d@kitenet.net","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2023-02-03T04:02:32Z","receivedAt":"2023-02-03T04:02:45Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Feb 02, 2023 at 05:19:30PM -0400, Joey Hess wrote:\n> In my opinion as the original developer of pristine-tar, it's too\n> complicated to be usefully used by git. The problem it solves is of a\n> larger scope than the problem git has here. (I hope.)\n\nWell, the problem which I believe folks on this thread are trying to\ndeal with is a way to reconstruct a bit-for-bit compressed tarball of\na particular release in a way that minimizes the cost of storage in\nthe git tree.  One way of doing that would be to guarantee that git\narchive would return something which is always bit-for-bit identical.\nAnother way is to use something like pristine tar.\n\nI'll grant that pristine tar does solve a bit more of the problem than\nwhat has been stated, since it allows the creator of the tarball to\nremove some files, or add some auto-generated files (e.g., after\nrunning autoreconf), and so in that way, pristine tar does solve a\nsomewhat larger problem than what was expressed in this thread.\n\nThat being said, however, pristine-tar is **extremely** useful, and\nI'm very happy, and very thankful, that you wrote it.  It has been\nsuper, super useful.\n\nCheers,\n\n\t\t\t\t\t\t- Ted\n"},{"id":"471418","messageId":"20230203080629.31492-1-ray@ameretat.dev","threadId":"59162","inReplyTo":"de8f1e338e6ee99cd3ee06b16f1edbce@ameretat.dev","subject":"[PATCH] archive: document output stability concerns","fromName":"Raymond E. Pasco","fromEmail":"ray@ameretat.dev","sentAt":"2023-02-03T08:06:29Z","receivedAt":"2023-02-03T08:16:01Z","isPatch":true,"sender":{"key":"ray@ameretat.dev","avatar":"https://avatars.githubusercontent.com/u/115765?v=4"},"body":"In 4f4be00d302 (archive-tar: use internal gzip by default), the 'git\narchive' command switched to using an internal compression filter\nimplemented with zlib rather than invoking a 'gzip' binary, for the\n'.tar.gz' / '.tgz' output formats.\n\nThis change brought to light a common misconception that the output of\n'git archive' is intended to be byte-for-byte stable. While this is not\nthe case, stable archive output is desirable for many applications; we\ndiscuss concerns related to output stability and suggest ways in which\nthe user can control the compression used with the\n\"tar.<format>.command\" configuration option.\n\nSigned-off-by: Raymond E. Pasco <ray@ameretat.dev>\n---\nI think that something along these lines should be included in the\ndocs, but that the behavior should be kept the same. If it is decided\nlater to stabilize output, e.g. by vendoring a blessed zlib version\nforever, the current state as of 2.38 is the best starting point;\nand reverting a useful change because of external breakage which\nalready has a solution, while also promising instability, seems like\na poor choice.\n\n Documentation/git-archive.txt | 35 +++++++++++++++++++++++++++++++++++\n 1 file changed, 35 insertions(+)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex 60c040988b..77acdacdf8 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -178,6 +178,41 @@ appropriate export-ignore in its `.gitattributes`), adjust the checked out\n option.  Alternatively you can keep necessary attributes that should apply\n while archiving any tree in your `$GIT_DIR/info/attributes` file.\n \n+[[STABILITY]]\n+STABILITY\n+---------\n+\n+'git archive' does not guarantee that precisely identical archive files\n+will be produced for invocations on the same commit or tree.\n+\n+'git archive' uses an internal implementation of `tar` archiving\n+for the `tar` format, which includes the commit ID in an extended\n+pax header.  For the `tgz` and `tar.gz` formats, it is augmented with\n+a compression filter applied to the output, which is implemented by\n+'git archive' by linking to the system zlib.\n+\n+If the commit ID of the \"same\" commit is different, for instance in the\n+case of an object format migration from SHA-1 to SHA-256, the `tar`\n+archive will necessarily differ due to including a different ID.\n+\n+The output of the compression filter is less deterministic than\n+the output of the `tar` implementation, because the versions\n+of zlib used may differ. The internal compression filter can be\n+replaced with a particular command specified by the user using the\n+`tar.<format>.command` configuration option; for instance, a particular\n+gzip binary provided by the user could be specified here for consistent\n+output.\n+\n+The `tar` format used by 'git archive' is unlikely to change\n+frequently, but is not guaranteed to be completely stable; its output\n+will remain identical at least within the same Git version.\n+\n+The `zip` format has similar concerns to the `tar.gz` and `tgz`\n+formats; ZIP archiving is implemented internally, but the Deflate\n+compression used relies on the linked zlib. However, because archiving\n+and compression are combined into a single operation, there is no\n+user-specifiable filter command for the `zip` format.\n+\n EXAMPLES\n --------\n `git archive --format=tar --prefix=junk/ HEAD | (cd /var/tmp/ && tar xf -)`::\n-- \n2.39.1.561.g98d13ac3e7\n\n"},{"id":"471428","messageId":"230203.86sffmc1tz.gmgdl@evledraar.gmail.com","threadId":"59162","inReplyTo":"01a901d93760$c690d970$53b28c50$@nexbridge.com","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-03T13:18:58Z","receivedAt":"2023-02-03T13:32:14Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Feb 02 2023, rsbecker@nexbridge.com wrote:\n\n> On February 2, 2023 6:02 PM, brian m. carlson wrote:\n>>On 2023-02-01 at 23:37:19, Junio C Hamano wrote:\n>>> \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n>>>\n>>> > I don't think a blurb is necessary, but you're basically\n>>> > underscoring the problem, which is that nobody is willing to promise\n>>> > that compression is consistent, but yet people want to rely on that\n>>> > fact.  I'm willing to write and implement a consistent tar spec and\n>>> > to guarantee compatibility with that, but the tension here is that\n>>> > people also want gzip to never change its byte format ever, which\n>>> > frankly seems unrealistic without explicit guarantees.  Maybe the\n>>> > authors will agree to promise that, but it seems unlikely.\n>>>\n>>> Just to step back a bit, where does the distinction between\n>>> guaranteeing the tar format stability and gzip compressed bitstream\n>>> stability come from?  At both levels, the same thing can be expressed\n>>> in multiple different ways, I think, but spelling out how exactly the\n>>> compressor compresses is more involved than spelling out how entries\n>>> in a tar archive is ordered and each entry is expressed, or something?\n>>\n>>Yes, at least with my understanding about how gzip and compression in general\n>>work.\n>>\n>>The tar format (and the pax format which builds on it) can mostly be restricted by\n>>explaining what data is to be included in the pax and tar headers and how it is to be\n>>formatted.  If we say, we will always write such and such information in the pax\n>>header and sort the keys, and we write such and such information in the tar header,\n>>then the format is completely deterministic, and we can make nice guarantees.\n>>\n>>My understanding about how Lempel-Ziv-based compression algorithms work is that\n>>there's a lot more freedom to decide how best to compress things and that there\n>>isn't always a logical obvious choice, but I will admit my understanding is relatively\n>>limited.  If someone thinks we can effectively succeed in supporting compression\n>>more than just relying on gzip, I would be delighted to be shown to be wrong.\n>\n> The nice part about gzip is that it is generally available on\n> virtually all platforms (or can be easily obtained). Other compression\n> forms, like bz2, which sometimes produces more dense compression, are\n> not necessarily available. Availability is something I would be\n> worried about...\n\nI agree with all of that, gzip is in such wide use for a reason. \n\n>... (clone and checkout failures).\n\nBut how would a hypothetical obscure format for \"git archive\" contribute\nto clone or checkout failures? Are you thinking of our use of zlib for\ne.g. loose objects? That's unrelated to this discussion (and I don't\nthink anyone relies on their compressed checksum).\n\n> Tar formats are also to be used carefully. Not all platform\n> implementations of tar support all variants. \"ustar\" is fairly common\n> but there are others that are not. Interoperability needs to be the\n> biggest factor in this decision, IMHO, rather than compression rates.\n\nFor \"git archive\" whether you care about interoperability depends on the\ntarget audience of your archive, and in any case I don't see why we need\nto worry about it, except to perhaps note that some are more portable\nthan others if we e.g. had a built-in \"tar.bz2\" helper method.\n\n> The alternative is having git supply its own implementation, but that\n> is a longer term migration problem, resembling the SHA-256 migration.\n\nI've noted elsewhere in this thread that I don't see the point of\nshipping a fallback \"gzip\" beyond the \"git archive gzip\" we have\nalready, but even if we did that the scope of that seems pretty simple,\nand *much* easier than the SHA-256 migration.\n"},{"id":"471430","messageId":"230203.86o7qac1hr.gmgdl@evledraar.gmail.com","threadId":"59162","inReplyTo":"Y9yHWFh0ijwrqhOX@mit.edu","subject":"Re: Stability of git-archive, breaking (?) the Github universe, and a possible solution","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-03T13:32:08Z","receivedAt":"2023-02-03T13:39:04Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Feb 02 2023, Theodore Ts'o wrote:\n\n> On Thu, Feb 02, 2023 at 05:19:30PM -0400, Joey Hess wrote:\n>> In my opinion as the original developer of pristine-tar, it's too\n>> complicated to be usefully used by git. The problem it solves is of a\n>> larger scope than the problem git has here. (I hope.)\n>\n> Well, the problem which I believe folks on this thread are trying to\n> deal with is a way to reconstruct a bit-for-bit compressed tarball of\n> a particular release in a way that minimizes the cost of storage in\n> the git tree.  One way of doing that would be to guarantee that git\n> archive would return something which is always bit-for-bit identical.\n> Another way is to use something like pristine tar.\n\nI think that's what this side-thread has devolved into, but I honestly\ndon't see how that's useful or more than tangentally related to the\nproblem noted at the start of the thread.\n\nIf you are writing a new system that consumes \"git archive\" output\nsomething like what I'm proposing to add in [1] should nicely sidestep\nthis issue, just checksum the uncompressed archive (assuming you're OK\nwith our soft \"tar\" guarantees), or \"git tag -v\" (if you can) etc.\n\nThat part of the docs is just a summary of what Konstantin Ryabitsev\npointed out in a side-thread.\n\nOne might also imagine any other number of trivial solutions to the\nproblem, e.g. people interested in this can unpack the archive, and then\n(needs to guarantee sorted order, which I think find(1) doesn't, but\njust as a POC):\n\n\t(cd unpacked && find . -type f -printf \"%f\\n\" -exec cat {} \\; | sha256sum)\n\nOr whatever.\n\nBut any such solution to the abstract problem isn't going to help the\nexisting users whose systems broke because they were assuming certain\nthings about the \"git archive\" output.\n\nFor those users I think (as my proposed series does) we should just do\nwhatever we can do limit the disruption, as my proposed [2] does by\nswitching back to \"gzip\".\n\nFor those users who are creating new systems that might use \"git\narchive\" today we then just need to update the documentation going\nforward. Maybe those could use \"pristine-tar\", or perhaps they can use\nsome entirely different distribution mechanism.\n\n1. https://lore.kernel.org/git/patch-9.9-b40833b2168-20230202T093212Z-avarab@gmail.com/\n2. https://lore.kernel.org/git/cover-0.9-00000000000-20230202T093212Z-avarab@gmail.com/\n"},{"id":"471433","messageId":"230203.86fsbmbzwp.gmgdl@evledraar.gmail.com","threadId":"59162","inReplyTo":"771a98ca-9540-ad4e-dfba-9d304e1dff09@dunelm.org.uk","subject":"Re: [PATCH 0/9] git archive: use gzip again by default, document output stabilty","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-03T13:49:37Z","receivedAt":"2023-02-03T14:14:15Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Feb 02 2023, Phillip Wood wrote:\n\n> On 02/02/2023 09:32, Ævar Arnfjörð Bjarmason wrote:\n>> As reported in\n>> https://lore.kernel.org/git/a812a664-67ea-c0ba-599f-cb79e2d96694@gmail.com/\n>> changing the default \"tgz\" output method of from \"gzip(1)\" to our\n>> internal \"git archive gzip\" (using zlib ) broke things for users in\n>> the wild that assume that the \"git archive\" output is stable, most\n>> notably GitHub: https://github.com/orgs/community/discussions/45830\n>>\n>> Leaving aside the larger question of whether we're going to promise\n>> output stability for \"git archive\" in general, the motivation for that\n>> change was to have a working compression method on systems that lacked\n>> a gzip(1).\n>\n> As I recall the reduction in cpu time used to create a compressed\n> archive was a factor in making it the default.\n\nI read those references in 76d7602631a (archive-tar: add internal gzip\nimplementation, 2022-06-15) more of a \"it's not [much] slower\", the flip\nto the default in 4f4be00d302 (archive-tar: use internal gzip by\ndefault, 2022-06-15) didn't discuss it.\n\nSo I didn't think it was important enough to mention (even though we're\nnow back to the faster \"gzip\" method).\n\n>> As the disruption of changing the default isn't worth it, let's use\n>> gzip(1) again by default, and only fall back on the new \"git archive\n>> gzip\" if it isn't available.\n>\n> Playing devil's advocate for a moment as we're not going to promise\n> that the compressed output of \"git archive\" will be stable in the\n> future perhaps we should use this breakage as an opportunity to\n> highlight that to users and to advertize the config setting that\n> allows them to use gzip for compressing archives.\n\nIf we were trying to intentionally break things for those users we could\ndo a lot better than \"git archive gzip\", whose output is mostly the same\nas \"gzip\", we could tweak one of the headers to make it different all\nthe time.\n\nBut I think it's better to advocate for such intentional chaos-monkeying\nas a follow-up to this more conservative \"oops, we broke stuff, it's\neasy not to break it, so let's not do it'.\n\n> Reverting the change gives the misleading impression that we're making\n> a commitment to keeping the output stable.\n\nI don't see how you can conclude that from this series. It explicitly\nstates that we make no such promises, what it does is go back to\nallowing the gzip(1) command to make its own promises.\n\n> The focus of this thread seems to be the\n> problems relating to github which they have already addressed.\n\nWhich they've addressed by reverting the change, but while they're a\nmajor user of git they're not the only one. They just happened to use\n\"git archive\".\n\nI think it would be a mistake to conclude that everyone who's run into\nthis has already done so, or is aware of it.\n\n> I think there is general agreement that it is not practical to promise\n> that the compressed output of \"git archive\" is stable so maybe it is\n> better[...]\n\n...better than what? This seems to imply that this series is making new\npromises about the output stability, which it isn't doing.\n\n> [...]to make that clear now while users can work around it in the\n> short term with a config setting rather than waiting until we're faced\n> with some security or other issue that forces a change to the output\n> which users cannot work around so easily.\n\nI think it's always been clear that you can use that setting. For ages\nwe've been saying:\n\n\tThe `tar.gz` and `tgz` formats are defined automatically and use the\n\tcommand `gzip -cn` by default.\n\nThen v2.38.0 changed it to:\n\n\t[...]\n        magic command `git archive gzip` by default\n\nWhich IMO was easily missed among other \"Performance, Internal\nImplementation, Development Support etc.\" items in the release notes,\nwhich said:\n\n   Teach \"git archive\" to (optionally and then by default) avoid\n   spawning an external \"gzip\" process when creating \".tar.gz\" (and\n   \".tgz\") archives.\n\nBut I agree that all of this is subjective. To me a 2% reduction in CPU\nuse (at the cost of ~20% increse in wallclock) & some unclear benefits\nto teaching users that they can't rely on our \"gzip\" output seems\nunclear or hypothetical.\n\nWhereas the widespread breakage reported is very real, and we should\nconsider GitHub as a canary for that, not the the stand & end of its\npotential impact.\n\nAs we didn't have a strong reason to change this in the first place (and\nas my series shows, we can have our cake & eat it too if we don't have a\n\"gzip\") I think the obvious choice is to go back to using \"gzip\".\n"},{"id":"471435","messageId":"Y90soPW6KRB7PQCY@mit.edu","threadId":"59162","inReplyTo":"771a98ca-9540-ad4e-dfba-9d304e1dff09@dunelm.org.uk","subject":"Re: [PATCH 0/9] git archive: use gzip again by default, document output stabilty","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2023-02-03T15:47:44Z","receivedAt":"2023-02-03T15:48:15Z","isPatch":true,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Feb 02, 2023 at 04:17:09PM +0000, Phillip Wood wrote:\n> Playing devil's advocate for a moment as we're not going to promise that the\n> compressed output of \"git archive\" will be stable in the future perhaps we\n> should use this breakage as an opportunity to highlight that to users and to\n> advertize the config setting that allows them to use gzip for compressing\n> archives. Reverting the change gives the misleading impression that we're\n> making a commitment to keeping the output stable. The focus of this thread\n> seems to be the problems relating to github which they have already\n> addressed.\n> \n> I think there is general agreement that it is not practical to promise that\n> the compressed output of \"git archive\" is stable so maybe it is better to\n> make that clear now while users can work around it in the short term with a\n> config setting rather than waiting until we're faced with some security or\n> other issue that forces a change to the output which users cannot work\n> around so easily.\n\nI would be in favor of adding a config option that allows using the\ninternal gzip option, although leave the default to be keep things\ncompatible.\n\nThe reason for that it should be easy for a forge provider such as\nGitHub to break things, deliberately.  Sound insane?  Hear me out.\n\nAt $WORK, we have a highly reliable system, Paxos.  It is a highly\nfault-tolerant system, so it rarely fails.  But \"rarely fails\" is not\nthe same as \"never fails\".  And hopefully, things should degrade\ngracefully if there is a Paxos outage.  But as the Google SRE's are\nfond of saying, \"Hope is not a strategy\".\n\nSo periodically, the people who run the Paxos service will\ndeliberately force downtime for a short amount of time.  The fact that\nthey will do this is well advertised, and scheduled ahead of time ---\nand teams responsible for user-facing services are supposed to make\nsure that end-users don't notice when this happens.  Maybe they won't\nbe able to update configurations as easily while Paxos is down, but it\nshouldn't cause a user-visible outage.\n\nSo what I would recommend to the GitHub product manager, is that once\na quarter, on a well-advertised date, that they flip the switch and\nbreak the git archive checksums for say, an hour.  Then next quarter,\nthey advertise that the switch will be thrown for 2 hours, doubling\neach time, until it is ramped up to 16 hours.\n\nThis will provide the necessary nudge so that all of these badly\ndesigned systems that depend on downloaded archives of arbitrary git\nhubs to be stable will rethink their position, while minimizing the\nend-user customer impact.  Otherwise, I predict that Bazel, homebrew,\netc will consider to rely on this ill-considered assumption, and at\nsome point in the future, when we *do* have a much better reason to\nwant to make a change to the tar or compression algorithm, all of\nthese end users will once again scream bloody murder.\n\nOf course, this is going to be up to each forge provider to decide\nwhether they want to do this.  But we can make it easy for them to do\nthis thing, and I'd argue it is in our interest to make it easy for\nthem to do this.  Otherwise we'll get constrained in the future by the\nfear of massive user blowback, no metter what we say in our\ndocumentation regarding \"no promises --- and next time, we really\nmean it!\"\n\n\t      \t       \t       \t    \t  - Ted\n"},{"id":"471498","messageId":"Y96Zbttj+VzsSz+w@tapette.crustytoothpaste.net","threadId":"59162","inReplyTo":"xmqq1qn8vxej.fsf@gitster.g","subject":"Re: [PATCH 9/9] git archive docs: document output non-stability","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-02-04T17:46:33Z","receivedAt":"2023-02-04T17:46:39Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-02-02 at 16:34:44, Junio C Hamano wrote:\n> There shouldn't be cross platform differences to break bit-for-bit\n> stability at least for \"tar\" format, as we do not rely on any\n> external library.  Can we say the same for \"zip\"?  I thought we\n> throw the blob at git_deflate_*() so the exact bitstream is up to\n> the libz implementation?\n\nThat's also true.  There, we can't use gzip, so we do whatever libz\ndoes.  For Zip, I believe we embed a local timestamp, so the output is\nalso dependent on the time zone.  I don't know enough about the Zip\nformat to say if there are any other things that may vary.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"471499","messageId":"c3f215ca-b4ae-79a2-c14a-3c0f1799e6f7@web.de","threadId":"59162","inReplyTo":"xmqq5yckvxtb.fsf@gitster.g","subject":"Re: [PATCH 0/9] git archive: use gzip again by default, document output stabilty","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2023-02-04T18:08:39Z","receivedAt":"2023-02-04T18:08:59Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 02.02.23 um 17:25 schrieb Junio C Hamano:\n> Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n>\n>> As the disruption of changing the default isn't worth it, let's use\n>> gzip(1) again by default, and only fall back on the new \"git archive\n>> gzip\" if it isn't available.\n>\n> It perhaps is OK, and lets us answer \"ugh, the compressed output of\n> 'git archive' is unstable again\" with \"we didn't change anything,\n> perhaps you changed your gzip(1)?\" when they fix bugs or improve\n> compression or whatever.  Of course that is not an overall win for\n> the end users, but in the short term until gzip gets such a change,\n> we would presumably get the \"same\" output as before.\n\nRestoring the old default is an understandable reflex.  In theory it\nworsens consistency and stability of the output, but in practice using\nwhatever was found in $PATH did work before -- or at least it was not\nour problem if it didn't.\n\nAre there still people left that would benefit from such a step back,\nhowever?  As far as I understand forges like GitHub relied on git\narchive producing the same tgz output across versions.  That assumption\nwas violated, trust lost.  They had to learn about the configuration\noption tar.tgz.command or find some other way to cope.  Changing the\ndefault again won't undo that.\n\nRené\n\n"},{"id":"471546","messageId":"230205.86mt5r7q2e.gmgdl@evledraar.gmail.com","threadId":"59162","inReplyTo":"c3f215ca-b4ae-79a2-c14a-3c0f1799e6f7@web.de","subject":"Re: [PATCH 0/9] git archive: use gzip again by default, document output stabilty","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2023-02-05T21:30:41Z","receivedAt":"2023-02-05T21:36:15Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sat, Feb 04 2023, René Scharfe wrote:\n\n> Am 02.02.23 um 17:25 schrieb Junio C Hamano:\n>> Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n>>\n>>> As the disruption of changing the default isn't worth it, let's use\n>>> gzip(1) again by default, and only fall back on the new \"git archive\n>>> gzip\" if it isn't available.\n>>\n>> It perhaps is OK, and lets us answer \"ugh, the compressed output of\n>> 'git archive' is unstable again\" with \"we didn't change anything,\n>> perhaps you changed your gzip(1)?\" when they fix bugs or improve\n>> compression or whatever.  Of course that is not an overall win for\n>> the end users, but in the short term until gzip gets such a change,\n>> we would presumably get the \"same\" output as before.\n>\n> Restoring the old default is an understandable reflex.  In theory it\n> worsens consistency and stability of the output, but in practice using\n> whatever was found in $PATH did work before -- or at least it was not\n> our problem if it didn't.\n\n\"In theory\" because the user might be flip-flopping between different\ngzip(1) versions?\n\n> Are there still people left that would benefit from such a step back,\n> however?  As far as I understand forges like GitHub relied on git\n> archive producing the same tgz output across versions.  That assumption\n> was violated, trust lost.  They had to learn about the configuration\n> option tar.tgz.command or find some other way to cope.  Changing the\n> default again won't undo that.\n\nI think it's safe to assume that git is used by enough users that\nanything breaking at a major hosting provider is likely to have a very\nlong tail in the wild, almost all of which we'll never see in \"this\nbroke for me\" reports to this ML.\n\nSo no, that ship has clearly sailed for GitHub, but this series aims to\naddress more than that.\n\nEven if it wasn't for that breakage, I think 4/9 and 6/9 here show the\nmain problem you were trying to solve in making \"git archive gzip\" the\ndefault didn't need to be solved by changing the default. I.e. the aim\nwas to have it work when \"gzip(1)\" wasn't available, which we can do by\nfalling back only if we can't invoke it, rather than changing the\nlong-standing default.\n"},{"id":"471558","messageId":"b24fc8ae-a9f8-868f-b281-74c256447084@dunelm.org.uk","threadId":"59162","inReplyTo":"230203.86fsbmbzwp.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH 0/9] git archive: use gzip again by default, document output stabilty","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2023-02-06T14:46:57Z","receivedAt":"2023-02-06T14:47:03Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 03/02/2023 13:49, Ævar Arnfjörð Bjarmason wrote:\n> \n> On Thu, Feb 02 2023, Phillip Wood wrote: >> Reverting the change gives the misleading impression that we're making\n>> a commitment to keeping the output stable.\n> \n> I don't see how you can conclude that from this series. It explicitly\n> states that we make no such promises, what it does is go back to\n> allowing the gzip(1) command to make its own promises.\n\nThis series would not be happening if we were not reverting a change to \nthe compressed output of 'git archive'. The documentation updates are \nvery welcome but I think we're undermining the message that the \ncompressed output can change by reverting that change.\n\n>> The focus of this thread seems to be the\n>> problems relating to github which they have already addressed.\n> \n> Which they've addressed by reverting the change, but while they're a\n> major user of git they're not the only one. They just happened to use\n> \"git archive\".\n> \n> I think it would be a mistake to conclude that everyone who's run into\n> this has already done so, or is aware of it.\n\nI've spent some time trying to find reports of problems caused by this \nchange and have not seen anything apart from the issue with GitHub. \nAlthough it takes a while for new versions of git to get into linux \ndistributions if there is a widespread problem we normally hear about it \npretty quickly. This change has been in two releases now. If anyone does \nhave a problem there is an easy fix in the form of setting \ntar.<format>.command\n\n>> I think there is general agreement that it is not practical to promise\n>> that the compressed output of \"git archive\" is stable so maybe it is\n>> better[...]\n> \n> ...better than what? This seems to imply that this series is making new\n> promises about the output stability, which it isn't doing.\n\nIt's better people realize they cannot rely on the output being stable \nnow when they can safely work around the problem while working on a \nproper fix rather than waiting until the change in output is caused by a \nsecurity issue in gzip which means the work around is no longer safe.\n\nBest Wishes\n\nPhillip\n\n>> [...]to make that clear now while users can work around it in the\n>> short term with a config setting rather than waiting until we're faced\n>> with some security or other issue that forces a change to the output\n>> which users cannot work around so easily.\n> \n> I think it's always been clear that you can use that setting. For ages\n> we've been saying:\n> \n> \tThe `tar.gz` and `tgz` formats are defined automatically and use the\n> \tcommand `gzip -cn` by default.\n> \n> Then v2.38.0 changed it to:\n> \n> \t[...]\n>          magic command `git archive gzip` by default\n> \n> Which IMO was easily missed among other \"Performance, Internal\n> Implementation, Development Support etc.\" items in the release notes,\n> which said:\n> \n>     Teach \"git archive\" to (optionally and then by default) avoid\n>     spawning an external \"gzip\" process when creating \".tar.gz\" (and\n>     \".tgz\") archives.\n> \n> But I agree that all of this is subjective. To me a 2% reduction in CPU\n> use (at the cost of ~20% increse in wallclock) & some unclear benefits\n> to teaching users that they can't rely on our \"gzip\" output seems\n> unclear or hypothetical.\n> \n> Whereas the widespread breakage reported is very real,\n\nwhere are the reports of widespread berakage outside of GitHub?\n\n> and we should\n> consider GitHub as a canary for that, not the the stand & end of its\n> potential impact.\n> \n> As we didn't have a strong reason to change this in the first place (and\n> as my series shows, we can have our cake & eat it too if we don't have a\n> \"gzip\") I think the obvious choice is to go back to using \"gzip\".\n"},{"id":"471995","messageId":"29b3cd6f-6e06-e32f-dfad-ab527488ba12@web.de","threadId":"59162","inReplyTo":"230205.86mt5r7q2e.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH 0/9] git archive: use gzip again by default, document output stabilty","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2023-02-12T17:41:01Z","receivedAt":"2023-02-12T17:41:26Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 05.02.23 um 22:30 schrieb Ævar Arnfjörð Bjarmason:\n>\n> On Sat, Feb 04 2023, René Scharfe wrote:\n>\n>> Am 02.02.23 um 17:25 schrieb Junio C Hamano:\n>>> Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n>>>\n>>>> As the disruption of changing the default isn't worth it, let's use\n>>>> gzip(1) again by default, and only fall back on the new \"git archive\n>>>> gzip\" if it isn't available.\n>>>\n>>> It perhaps is OK, and lets us answer \"ugh, the compressed output of\n>>> 'git archive' is unstable again\" with \"we didn't change anything,\n>>> perhaps you changed your gzip(1)?\" when they fix bugs or improve\n>>> compression or whatever.  Of course that is not an overall win for\n>>> the end users, but in the short term until gzip gets such a change,\n>>> we would presumably get the \"same\" output as before.\n>>\n>> Restoring the old default is an understandable reflex.  In theory it\n>> worsens consistency and stability of the output, but in practice using\n>> whatever was found in $PATH did work before -- or at least it was not\n>> our problem if it didn't.\n>\n> \"In theory\" because the user might be flip-flopping between different\n> gzip(1) versions?\n\nNo flopping needed.  We can't control what's in $PATH.  There are\nOS-specific replacements for GNU gzip in NetBSD/FreeBSD/macOS and\nOpenBSD.  People could use pigz.  Or cat, for that matter.  Different\nversions of different tools might produce different output.\n\nThere are alternative to the original libz as well, e.g. libz-ng.  We\ndon't control which one or which version is installed, either, but we\ncould do so if we wanted by importing one of them like we did with\nLibXDiff.\n\n> Even if it wasn't for that breakage, I think 4/9 and 6/9 here show the\n> main problem you were trying to solve in making \"git archive gzip\" the\n> default didn't need to be solved by changing the default. I.e. the aim\n> was to have it work when \"gzip(1)\" wasn't available, which we can do by\n> falling back only if we can't invoke it, rather than changing the\n> long-standing default.\n\nThe aim was to no longer depend on gzip.  That goal was already met by\nproviding the internal implementation, without changing the default.\nGit for Windows for example could use it in their config and drop gzip.\n\nCalling gzip if available, warning if it isn't and using the internal\nimplementation adds yet more variance.  No longer allowing gzip to be a\nshell alias might confuse someone.  The automatic fallback would only\nbenefit users that don't want to touch /etc/gitconfig, have nobody to\ndo it for them and don't care about warnings -- hopefully not a big\ncrowd.\n\nI didn't intend the change of default to be that painful, but don't see\nthe point in going back now that we're through.  The new default is\nbetter -- one less dependency to care about.  And if we need to go\nback, however, then a know-good state makes more sense than a smart\nfallback with some new twists.\n\nRené\n"}]}