{"thread":{"id":"38670","subject":"weaning distributions off tarballs: extended verification of git tags","startedAt":"2015-02-28T14:48:05Z","lastAt":"2015-07-08T04:00:49Z","messageCount":13,"participants":["Colin Walters","brian m. carlson","Morten Welinder","Joey Hess","Sam Vilain","Junio C Hamano","Duy Nguyen","Michael Haggerty"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"256778","messageId":"1425134885.3150003.233627665.2E48E28B@webmail.messagingengine.com","threadId":"38670","inReplyTo":null,"subject":"weaning distributions off tarballs: extended verification of git tags","fromName":"Colin Walters","fromEmail":"walters@verbum.org","sentAt":"2015-02-28T14:48:05Z","receivedAt":"2015-02-28T14:48:05Z","isPatch":false,"sender":{"key":"walters@verbum.org","avatar":"https://gravatar.com/avatar/793625733050359d5917f5375179e5af7d2a56dca197661f6f709c4ad69cd7f2?d=mp&s=160"},"body":"Hi, \n\nTL;DR: Let's define a standard for embedding stronger checksums in tags and commit messages:\nhttps://github.com/cgwalters/homegit/blob/master/bin/git-evtag\n\nI think tarballs should go away as a source distribution mechanism in favor of pure git.  I won't go into too many details of the \"why\" here (hopefully most of you agree!) but that's the background.\n\nNow, there are a few things that the classical tarball model provides:\n\n- Version numbers compatible with dpkg/rpm/etc\n  -> Do the same with your tag names, and use a well known scheme like \"v$VERSION\"\n- The assumption that this source has been run through some tests\n  -> Broken assumption, and regardless you want to rerun tests downstream\n- Hosting providers typically offer a strong checksum over the entire source\n  -> The topic of this post\n\nThe above strawman code allows embedding the SHA256(git archive | tar).  Now,\nin order to make this work, the byte output of \"git archive\" must never change in the\nfuture.  I'm not sure how valid an assumption this is.  Timestamps are set to the\ncommit timestamp, but I could imagine someone wanting to come along later\nand tweak the output to be compatible with some variant of tar or something.\n\nWe could define the checksum to be over the stream of raw objects, sorted by their checksum,\nand that way be independent of archiving format variations.\n\nIs there agreement that something like this makes sense in the git core?  Does the\nconcept make sense?  Does anything like this exist today?  Other thoughts/objections?\n"},{"id":"256780","messageId":"20150228191403.GD514544@vauxhall.crustytoothpaste.net","threadId":"38670","inReplyTo":"1425134885.3150003.233627665.2E48E28B@webmail.messagingengine.com","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2015-02-28T19:14:03Z","receivedAt":"2015-02-28T19:14:03Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Sat, Feb 28, 2015 at 09:48:05AM -0500, Colin Walters wrote:\n>The above strawman code allows embedding the SHA256(git archive | tar).  Now,\n>in order to make this work, the byte output of \"git archive\" must never change in the\n>future.  I'm not sure how valid an assumption this is.  Timestamps are set to the\n>commit timestamp, but I could imagine someone wanting to come along later\n>and tweak the output to be compatible with some variant of tar or something.\n\nThis is not a safe assumption.  Unfortunately, kernel.org assumed that \nit was the case, and a change broke it.  Let's please not make more code \nthat does that.\n\n>We could define the checksum to be over the stream of raw objects, sorted by their checksum,\n>and that way be independent of archiving format variations.\n\nThis would be a much better idea, assuming you mean \"raw git objects\". \nFor cryptographic purposes, it's important to make the item boundaries \nunambiguous, which is usually done using the length.  Since the raw git \nobjects include the length, this is sufficient.\n\nIf you don't make the boundaries unambiguous, you get the problem you \nhave with v3 OpenPGP keys, where somebody could move bytes from one \nvalue to another, creating a different key, but with the same \nfingerprint (hash value).\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | http://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: RSA v4 4096b: 88AC E9B2 9196 305B A994 7552 F1BA 225C 0223 B187\n"},{"id":"256781","messageId":"CANv4PNmF9sTh8od9xT5tYTOF1Cv0Mev2Muf-qxQDS_6kE7EnOw@mail.gmail.com","threadId":"38670","inReplyTo":"1425134885.3150003.233627665.2E48E28B@webmail.messagingengine.com","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"Morten Welinder","fromEmail":"mwelinder@gmail.com","sentAt":"2015-02-28T20:34:41Z","receivedAt":"2015-02-28T20:34:41Z","isPatch":false,"sender":{"key":"mwelinder@gmail.com","avatar":null},"body":"Is there a point to including a different checksum inside\na git tag?  If someone can break the SHA-1 checksum\nin the repository then the recorded SHA-256 checksum can\nbe changed.  In other words, wouldn't you be just as well\noff handing someone a SHA-1 commit id?\n\nIf you can guard the SHA-256 with a signature, you can\ndo the same thing to the SHA-1.  Or the tarball for that matter.\n\nUnrelatedly, your assumptions:\n\nTar balls have too many degrees of freedom to rely on them\nbeing created identically in the future.\n\n> - The assumption that this source has been run through some tests\n\nA perfectly valid assumption for some build systems, notably\nautotools.  \"make distcheck\" is the only way my tarballs get\nmade and they only get made when the checks succeed.\n(If your point was that many projects have too few tests,\nwell, then I agree.)\n\nM.\n\n\n\nOn Sat, Feb 28, 2015 at 9:48 AM, Colin Walters <walters@verbum.org> wrote:\n> Hi,\n>\n> TL;DR: Let's define a standard for embedding stronger checksums in tags and commit messages:\n> https://github.com/cgwalters/homegit/blob/master/bin/git-evtag\n>\n> I think tarballs should go away as a source distribution mechanism in favor of pure git.  I won't go into too many details of the \"why\" here (hopefully most of you agree!) but that's the background.\n>\n> Now, there are a few things that the classical tarball model provides:\n>\n> - Version numbers compatible with dpkg/rpm/etc\n>   -> Do the same with your tag names, and use a well known scheme like \"v$VERSION\"\n> - The assumption that this source has been run through some tests\n>   -> Broken assumption, and regardless you want to rerun tests downstream\n> - Hosting providers typically offer a strong checksum over the entire source\n>   -> The topic of this post\n>\n> The above strawman code allows embedding the SHA256(git archive | tar).  Now,\n> in order to make this work, the byte output of \"git archive\" must never change in the\n> future.  I'm not sure how valid an assumption this is.  Timestamps are set to the\n> commit timestamp, but I could imagine someone wanting to come along later\n> and tweak the output to be compatible with some variant of tar or something.\n>\n> We could define the checksum to be over the stream of raw objects, sorted by their checksum,\n> and that way be independent of archiving format variations.\n>\n> Is there agreement that something like this makes sense in the git core?  Does the\n> concept make sense?  Does anything like this exist today?  Other thoughts/objections?\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"256849","messageId":"1425316197.895196.234425829.536E6C06@webmail.messagingengine.com","threadId":"38670","inReplyTo":"CANv4PNmF9sTh8od9xT5tYTOF1Cv0Mev2Muf-qxQDS_6kE7EnOw@mail.gmail.com","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"Colin Walters","fromEmail":"walters@verbum.org","sentAt":"2015-03-02T17:09:57Z","receivedAt":"2015-03-02T17:09:57Z","isPatch":false,"sender":{"key":"walters@verbum.org","avatar":"https://gravatar.com/avatar/793625733050359d5917f5375179e5af7d2a56dca197661f6f709c4ad69cd7f2?d=mp&s=160"},"body":"On Sat, Feb 28, 2015, at 03:34 PM, Morten Welinder wrote:\n> Is there a point to including a different checksum inside\n> a git tag?  If someone can break the SHA-1 checksum\n> in the repository then the recorded SHA-256 checksum can\n> be changed.  In other words, wouldn't you be just as well\n> off handing someone a SHA-1 commit id?\n\nThe issue is more about what the checksum covers, as\nwell as its strength.  Git uses a hash tree, which means\nthat an attacker only has to find a collision for *one* of\nthe objects, and the signature is still valid.  And that collision\nis valid for *every* commit that contains that object.\n\nThis topic has been covered elsewhere pretty extensively,\nhere's a link:\nhttps://www.whonix.org/forum/index.php/topic,538.msg4278.html#msg4278\n\nNow I think rough consensus is still that git is \"secure\" or\n\"secure enough\" - but with this proposal I'm just trying\nto overcome the remaining conservatism.  (Also, while those\ndiscussions were focusing on corrupting an existing repository,\nthe attack model of MITM also exists, and there\nyou don't have to worry about deltas, particularly if the\nattacker's goal is to get a downstream to do a build\nand thus execute their hostile code inside the downstream\nnetwork).\n\nIt's really not that expensive to do once per release,\nbasically free for small repositories, and for a large one like\nthe Linux kernel:\n\n$ cd ~/src/linux\n$ git describe\nv3.19-7478-g796e1c5\n$ time /bin/sh -c 'git archive --format=tar HEAD|sha256sum'\n4a5c5826cea188abd52fa50c663d17ebe1dfe531109fed4ddbd765a856f1966e  -\n\nreal\t0m3.772s\nuser\t0m6.132s\nsys\t0m0.279s\n$\n\nWith this proposal, the checksum\ncovers an entire stream of objects for a given commit at once;\nmaking it significantly harder to find a collision.  At least as good as \nchecksummed tarballs, and arguably better since it's\npre-compression.\n\nSo to implement this, perhaps something like:\n\n$ git archive --format=raw\n\nas a base primitive, and:\n\n$ git tag --archive-raw-checksum=SHA256 -s -m \"...\"\n\n?\n\n\"git fsck\" could also learn to optionally use this.\n"},{"id":"256850","messageId":"20150302181230.GA31798@kitenet.net","threadId":"38670","inReplyTo":"1425316197.895196.234425829.536E6C06@webmail.messagingengine.com","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"Joey Hess","fromEmail":"id@joeyh.name","sentAt":"2015-03-02T18:12:30Z","receivedAt":"2015-03-02T18:12:30Z","isPatch":false,"sender":{"key":"id@joeyh.name","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"I support this proposal, as someone who no longer releases tarballs\nof my software, when I can possibly avoid it. I have worried about\nsigned tags / commits only being a SHA1 break away from useless.\n\nAs to the implementation, checksumming the collection of raw objects is\ncertainly superior to tar. Colin had suggested sorting the objects by\nchecksum, but I don't think that is necessary. Just stream the commit\nobject, then its tree object, followed by the content of each object\nlisted in the tree, recursing into subtrees as necessary. That will be a\nstable stream for a given commit, or tree.\n\n-- \nsee shy jo\n"},{"id":"256857","messageId":"54F4BC18.5060702@vilain.net","threadId":"38670","inReplyTo":"20150302181230.GA31798@kitenet.net","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2015-03-02T19:38:00Z","receivedAt":"2015-03-02T19:38:00Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"On 03/02/2015 10:12 AM, Joey Hess wrote:\n> I support this proposal, as someone who no longer releases tarballs\n> of my software, when I can possibly avoid it. I have worried about\n> signed tags / commits only being a SHA1 break away from useless.\n>\n> As to the implementation, checksumming the collection of raw objects is\n> certainly superior to tar. Colin had suggested sorting the objects by\n> checksum, but I don't think that is necessary. Just stream the commit\n> object, then its tree object, followed by the content of each object\n> listed in the tree, recursing into subtrees as necessary. That will be a\n> stable stream for a given commit, or tree.\n\nI would really just do it exactly the same way that git does: checksum \nthe objects including their headers with the new hashes.  I have a hazy \nrecollection of what it would take to replace SHA-1 in git with \nsomething else; it should be possible (though tricky) to do it lazily, \nwhere a tree entry has bits (eg, some of the currently unused file mode \nbits) to denotes which hash algorithm is in use for the entry.  However \nI don't think that got past idea stage...\n\nSam\n"},{"id":"256860","messageId":"xmqqwq2z9n7c.fsf@gitster.dls.corp.google.com","threadId":"38670","inReplyTo":"54F4BC18.5060702@vilain.net","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2015-03-02T20:08:39Z","receivedAt":"2015-03-02T20:08:39Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Sam Vilain <sam@vilain.net> writes:\n\n>> As to the implementation, checksumming the collection of raw objects is\n>> certainly superior to tar. Colin had suggested sorting the objects by\n>> checksum, but I don't think that is necessary. Just stream the commit\n>> object, then its tree object, followed by the content of each object\n>> listed in the tree, recursing into subtrees as necessary. That will be a\n>> stable stream for a given commit, or tree.\n>\n> I would really just do it exactly the same way that git does: checksum\n> the objects including their headers with the new hashes.\n\nI tend to agree that it is a good idea.  I also suspect that would\nmake the implementation simpler by allowing it to share more code,\nbut I didn't look into it too deeply.\n\n> I have a\n> hazy recollection of what it would take to replace SHA-1 in git with\n> something else; it should be possible (though tricky) to do it lazily,\n> where a tree entry has bits (eg, some of the currently unused file\n> mode bits) to denotes which hash algorithm is in use for the entry.\n> However I don't think that got past idea stage...\n\nI think one reason why it didn't was because it would not work well.\nThat \"bit that tells this is a new object or old\" would mean that a\nsingle tree can have many different object names, depending on which\nof its component entries are using that bit and which aren't.  There\ngoes the \"we know two trees with the same object name are identical\nwithout recursing into them\" optimization out the window.\n\nAlso it would make it impossible to do what you suggest to Joey to\ndo, i.e. \"exactly the same way that git does\", once you start saying\nthat a tree object can be encoded in more than one different ways,\nwouldn't it?\n"},{"id":"256868","messageId":"54F4CD79.4080209@vilain.net","threadId":"38670","inReplyTo":"xmqqwq2z9n7c.fsf@gitster.dls.corp.google.com","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2015-03-02T20:52:09Z","receivedAt":"2015-03-02T20:52:09Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"On 03/02/2015 12:08 PM, Junio C Hamano wrote:\n>> I have a\n>> hazy recollection of what it would take to replace SHA-1 in git with\n>> something else; it should be possible (though tricky) to do it lazily,\n>> where a tree entry has bits (eg, some of the currently unused file\n>> mode bits) to denotes which hash algorithm is in use for the entry.\n>> However I don't think that got past idea stage...\n> I think one reason why it didn't was because it would not work well.\n> That \"bit that tells this is a new object or old\" would mean that a\n> single tree can have many different object names, depending on which\n> of its component entries are using that bit and which aren't.  There\n> goes the \"we know two trees with the same object name are identical\n> without recursing into them\" optimization out the window.\n>\n> Also it would make it impossible to do what you suggest to Joey to\n> do, i.e. \"exactly the same way that git does\", once you start saying\n> that a tree object can be encoded in more than one different ways,\n> wouldn't it?\n\nI was reasoning that people would rather not have to rewrite their whole \nhistory in order to switch checksum algorithms, and that by allowing \ntrees to be lazily converted that this would make things more \nefficient.  However, I think I see your point here that this doesn't work.\n\nHowever, as a per-commit header, then only first commit which changes \nthe hashing algorithm would have to re-checksum each of the files: but \njust in the current tree, not all the way back to the beginning of \nhistory.  The delta logic should not have to care, and these objects \nwith the same content but different object ID should pack perfectly, so \nlong as git-pack-objects knows to re-checksum objects with the available \nhash algorithms and spot matches.\n\nOther operations like diff which span commit hashing algorithms might be \nable to get away with their existing object ranking algorithms and cache \nalternate object IDs for content as they operate to facilitate exact \nmatching across hash algorithm changes.\n\nBut actually, for the original problem - just producing a signature with \na different hashing algorithm - probably it would be sufficient to just \nre-hash the current commit and the current tree recursively, and the \nmixed hash-algorithm case does not need to exist.  But I'm just thinking \nit might not be too hard to make git nicely generic, to be well prepared \nfor when a second pre-image attack on SHA-1 becomes practical.\n\nSam\n"},{"id":"256877","messageId":"CACsJy8C3=f=esBrHE8OudSa0nUbCrLaYJtLC2in3p+tcc-d9bw@mail.gmail.com","threadId":"38670","inReplyTo":"20150302181230.GA31798@kitenet.net","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2015-03-02T23:20:26Z","receivedAt":"2015-03-02T23:20:26Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Mar 3, 2015 at 1:12 AM, Joey Hess <id@joeyh.name> wrote:\n> I support this proposal, as someone who no longer releases tarballs\n> of my software, when I can possibly avoid it. I have worried about\n> signed tags / commits only being a SHA1 break away from useless.\n>\n> As to the implementation, checksumming the collection of raw objects is\n> certainly superior to tar. Colin had suggested sorting the objects by\n> checksum, but I don't think that is necessary. Just stream the commit\n> object, then its tree object, followed by the content of each object\n> listed in the tree, recursing into subtrees as necessary. That will be a\n> stable stream for a given commit, or tree.\n\nIt could be simplified a bit by using ls-tree -r (so you basically\nhave a single big tree). Then hash commit, ls-tree -r output and all\nblobs pointed by ls-tree in listed order.\n-- \nDuy\n"},{"id":"256879","messageId":"xmqqsidn7ymg.fsf@gitster.dls.corp.google.com","threadId":"38670","inReplyTo":"CACsJy8C3=f=esBrHE8OudSa0nUbCrLaYJtLC2in3p+tcc-d9bw@mail.gmail.com","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2015-03-02T23:44:55Z","receivedAt":"2015-03-02T23:44:55Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> On Tue, Mar 3, 2015 at 1:12 AM, Joey Hess <id@joeyh.name> wrote:\n>> I support this proposal, as someone who no longer releases tarballs\n>> of my software, when I can possibly avoid it. I have worried about\n>> signed tags / commits only being a SHA1 break away from useless.\n>>\n>> As to the implementation, checksumming the collection of raw objects is\n>> certainly superior to tar. Colin had suggested sorting the objects by\n>> checksum, but I don't think that is necessary. Just stream the commit\n>> object, then its tree object, followed by the content of each object\n>> listed in the tree, recursing into subtrees as necessary. That will be a\n>> stable stream for a given commit, or tree.\n>\n> It could be simplified a bit by using ls-tree -r (so you basically\n> have a single big tree). Then hash commit, ls-tree -r output and all\n> blobs pointed by ls-tree in listed order.\n\nWhat problem are you trying to solve here, though, by deliberately\ndeviating what Git internally used to store these objects?  If it is\nOK to ignore the tree boundary, then you probably do not even need\ntrees in this secondary hash for validation in the first place.\n\nFor example, you can hash a stream:\n\n    <commit object contents> +\n    N * (<pathname> + NUL + <blob object contents>)\n\nas long as the <pathname>s are sorted in a predictable order (like\nin \"the index order\") in the output.  That would be even simpler (I\nam not saying it is necessarily better, and by inference neither is\nyour \"simplification\").\n\nI was about to suggest another alternative.\n\n    Pretend as if Git internally used SHA-512 (or whatever hash you\n    want to use) instead of SHA-1, compute the object names that\n    way.  Recompute the contents of a tree object is by replacing\n    the 20-byte SHA-1 field in it with a field with whatever\n    necessary length to hold the longer object names of elements in\n    the tree.\n\nBut then a realization hit me: what new value will be placed in the\n\"parent \" field in the commit object?  You cannot have SHA-512\nvariant of commit object name without recomputing the whole history.\n\nNow, if the final objective is to replace signature of tarballs,\ndoes it matter to cover the commit object, or is it sufficient to\ncover the tree contents?\n\nAmong the ideas raised so far, I like what Joey suggested, combined\nwith \"each should have '<type> <length>NUL' header\" from Sam Vilain\nthe best.  That is, hash the stream:\n\n    \"commit <length>\" NUL + <commit object contents> +\n    \"tree <length>\" NUL + <top level tree contents> +\n    ... list the entries in the order you would find by\n    ... some defined traversal order people can agree on.\n\nwith whatever the preferred strong hash function of the age.\n"},{"id":"256881","messageId":"CACsJy8ALQ=Hs2vnpiNxbp-n_sZvNahhtE4N2H-4_Jma4yo6rVQ@mail.gmail.com","threadId":"38670","inReplyTo":"xmqqsidn7ymg.fsf@gitster.dls.corp.google.com","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2015-03-03T00:42:09Z","receivedAt":"2015-03-03T00:42:09Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Mar 3, 2015 at 6:44 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> Duy Nguyen <pclouds@gmail.com> writes:\n>\n>> On Tue, Mar 3, 2015 at 1:12 AM, Joey Hess <id@joeyh.name> wrote:\n>>> I support this proposal, as someone who no longer releases tarballs\n>>> of my software, when I can possibly avoid it. I have worried about\n>>> signed tags / commits only being a SHA1 break away from useless.\n>>>\n>>> As to the implementation, checksumming the collection of raw objects is\n>>> certainly superior to tar. Colin had suggested sorting the objects by\n>>> checksum, but I don't think that is necessary. Just stream the commit\n>>> object, then its tree object, followed by the content of each object\n>>> listed in the tree, recursing into subtrees as necessary. That will be a\n>>> stable stream for a given commit, or tree.\n>>\n>> It could be simplified a bit by using ls-tree -r (so you basically\n>> have a single big tree). Then hash commit, ls-tree -r output and all\n>> blobs pointed by ls-tree in listed order.\n>\n> What problem are you trying to solve here, though, by deliberately\n> deviating what Git internally used to store these objects?  If it is\n> OK to ignore the tree boundary, then you probably do not even need\n> trees in this secondary hash for validation in the first place.\n>\n> For example, you can hash a stream:\n>\n>     <commit object contents> +\n>     N * (<pathname> + NUL + <blob object contents>)\n>\n> as long as the <pathname>s are sorted in a predictable order (like\n> in \"the index order\") in the output.  That would be even simpler (I\n> am not saying it is necessarily better, and by inference neither is\n> your \"simplification\").\n\nI did nearly that [1]. But this morning I realized trees carry file\npermission. We should keep that in the final checksum as well.\n\n> Now, if the final objective is to replace signature of tarballs,\n> does it matter to cover the commit object, or is it sufficient to\n> cover the tree contents?\n>\n> Among the ideas raised so far, I like what Joey suggested, combined\n> with \"each should have '<type> <length>NUL' header\" from Sam Vilain\n> the best.  That is, hash the stream:\n>\n>     \"commit <length>\" NUL + <commit object contents> +\n>     \"tree <length>\" NUL + <top level tree contents> +\n>     ... list the entries in the order you would find by\n>     ... some defined traversal order people can agree on.\n>\n> with whatever the preferred strong hash function of the age.\n\nA bit harder to script, but simpler to provide from cat-file, I think.\n\n[1] http://article.gmane.org/gmane.comp.version-control.git/260211\n-- \nDuy\n"},{"id":"257071","messageId":"54F84DCB.9000900@alum.mit.edu","threadId":"38670","inReplyTo":"xmqqsidn7ymg.fsf@gitster.dls.corp.google.com","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2015-03-05T12:36:27Z","receivedAt":"2015-03-05T12:36:27Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 03/03/2015 12:44 AM, Junio C Hamano wrote:\n> [...]\n> I was about to suggest another alternative.\n> \n>     Pretend as if Git internally used SHA-512 (or whatever hash you\n>     want to use) instead of SHA-1, compute the object names that\n>     way.  Recompute the contents of a tree object is by replacing\n>     the 20-byte SHA-1 field in it with a field with whatever\n>     necessary length to hold the longer object names of elements in\n>     the tree.\n> \n> But then a realization hit me: what new value will be placed in the\n> \"parent \" field in the commit object?  You cannot have SHA-512\n> variant of commit object name without recomputing the whole history.\n> \n> Now, if the final objective is to replace signature of tarballs,\n> does it matter to cover the commit object, or is it sufficient to\n> cover the tree contents?\n\nThe original goal was to replace a tarball signature, for which the\n\"alternative\" that you described above seems quite elegant.\n\nIf the goal were really to certify the entire history, then none of the\nproposals that I have seen so far is adequate anyway, because none of\nthem propose to include better than the original SHA-1s of the parent\ncommits.\n\nIncluding other metadata from the release commit does not seem useful to\nme; how valuable is it to know the author and commit message of the last\ncommit that happened to make it into a release? It would be more useful\nto know the SHA-1 of that commit, but that would presumably be included\nelsewhere in the packaging data used by the distribution.\n\n> [...]\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\n"},{"id":"265800","messageId":"1436328049.1937003.317969577.6CBA24A0@webmail.messagingengine.com","threadId":"38670","inReplyTo":"1425134885.3150003.233627665.2E48E28B@webmail.messagingengine.com","subject":"Re: weaning distributions off tarballs: extended verification of git tags","fromName":"Colin Walters","fromEmail":"walters@verbum.org","sentAt":"2015-07-08T04:00:49Z","receivedAt":"2015-07-08T04:00:49Z","isPatch":false,"sender":{"key":"walters@verbum.org","avatar":"https://gravatar.com/avatar/793625733050359d5917f5375179e5af7d2a56dca197661f6f709c4ad69cd7f2?d=mp&s=160"},"body":"\n\nOn Sat, Feb 28, 2015, at 10:48 AM, Colin Walters wrote:\n> Hi, \n> \n> TL;DR: Let's define a standard for embedding stronger checksums in tags and commit messages:\n> https://github.com/cgwalters/homegit/blob/master/bin/git-evtag\n\n[time passes]\n\nI finally had a bit of time to pick this back up again in:\n\nhttps://github.com/cgwalters/git-evtag\n\nIt should address the core concern here about stability of `git archive`.\n\nI prototyped it out with libgit2 because it was easier, and I'd like actually to be able to use this with older versions of git.\n\nBut I think the next steps here are:\n\n- Validate the core design\n  * Tree walking order\n  * Submodule recursion\n  * Use of SHA512\n- Standardize it\n  (Would like to see at least a stupid slow shell script implementation to cross-validate)\n- Add it as an option to `git tag`?\n\nLonger term:\n- Support adding `Git-EVTag` as a git note, so I can retroactively add stronger\n  checksums to older git repositories\n- Anything else?\n"}]}