{"thread":{"id":"5317","subject":"[RFC] adding support for md5","startedAt":"2006-08-18T06:01:44Z","lastAt":"2006-08-24T10:34:22Z","messageCount":20,"participants":["David Rientjes","Nguyễn Thái Ngọc Duy","Johannes Schindelin","Trekie","Petr Baudis","Jon Smirl","Linus Torvalds","Chris Wedgwood","Junio C Hamano","Shawn Pearce"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"25526","messageId":"Pine.LNX.4.63.0608172259280.25827@chino.corp.google.com","threadId":"5317","inReplyTo":null,"subject":"[RFC] adding support for md5","fromName":"David Rientjes","fromEmail":"rientjes@google.com","sentAt":"2006-08-18T06:01:44Z","receivedAt":"2006-08-18T06:01:44Z","isPatch":false,"sender":{"key":"rientjes@google.com","avatar":null},"body":"I'd like to solicit some comments about implementing support for md5 as a \nhash function that could be determined at runtime by the user during a \nproject init-db.  md5, which I implemented as a configurable option in my \nown tree, is a 128-bit hash that is slightly more recognized than sha1.  \nLikewise, it is also available in openssl/md5.h just as sha1 is available \nthrough a library in openssl/sha1.h.  My patch to move the hash name \ncomparison was a step in this direction in isolating many of the \nparticulars of hash-specific dependencies.\n\n\t\tDavid\n"},{"id":"25536","messageId":"fcaeb9bf0608180259o21a3d482y71b50227bd070497@mail.gmail.com","threadId":"5317","inReplyTo":"Pine.LNX.4.63.0608172259280.25827@chino.corp.google.com","subject":"Re: [RFC] adding support for md5","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2006-08-18T09:59:38Z","receivedAt":"2006-08-18T09:59:38Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On 8/18/06, David Rientjes <rientjes@google.com> wrote:\n> I'd like to solicit some comments about implementing support for md5 as a\n> hash function that could be determined at runtime by the user during a\n> project init-db.  md5, which I implemented as a configurable option in my\n> own tree, is a 128-bit hash that is slightly more recognized than sha1.\n> Likewise, it is also available in openssl/md5.h just as sha1 is available\n> through a library in openssl/sha1.h.  My patch to move the hash name\n> comparison was a step in this direction in isolating many of the\n> particulars of hash-specific dependencies.\nJust curious, but why md5? Is there any benefit using md5 over sha1?\n"},{"id":"25538","messageId":"Pine.LNX.4.63.0608181209210.28360@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"5317","inReplyTo":"Pine.LNX.4.63.0608172259280.25827@chino.corp.google.com","subject":"Re: [RFC] adding support for md5","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-08-18T10:21:11Z","receivedAt":"2006-08-18T10:21:11Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 17 Aug 2006, David Rientjes wrote:\n\n> I'd like to solicit some comments about implementing support for md5 as a \n> hash function that could be determined at runtime by the user during a \n> project init-db.\n\nMake it a config variable, too, right?\n\nIt would be interesting to have other hashes, for three reasons:\n\n1. they could be faster to calculate,\n2. they could reduce clashes, and related to that,\n3. it is possible that some day SHA1 is broken, i.e. that there is an \n   algorithm to generate a different text for a given hash.\n\nAs for 2 and 3, it seems MD5 is equivalent, since another sort of attacks \nwas already successful on both SHA1 and MD5: generating two different \ntexts with the same hash.\n\nSo, 1 could be a good reason to have another hash. IIRC SHA1 is about 25% \nslower than MD5, so it could be worth it.\n\nHowever, you should know that there is _no way_ to use both hashes on the \nsame project. Yes, you could rewrite the history, trying to convert also \nthe hashes in the commit objects, but people actually started relying on \nnaming commits with the short-SHA1.\n\nI think it would be a nice thing to play through (for example, to find \nout how much impact the hash calculation has on the overall performance \nof git), but I doubt it will ever come to real use.\n\nCiao,\nDscho\n"},{"id":"25542","messageId":"44E59BF6.2070909@sinister.cz","threadId":"5317","inReplyTo":"Pine.LNX.4.63.0608172259280.25827@chino.corp.google.com","subject":"Re: [RFC] adding support for md5","fromName":"Trekie","fromEmail":"trekie@sinister.cz","sentAt":"2006-08-18T10:52:38Z","receivedAt":"2006-08-18T10:52:38Z","isPatch":false,"sender":{"key":"trekie@sinister.cz","avatar":null},"body":"Hi,\n\nI'd like to point out that while finding collision for SHA1 according to\nthe Wikipedia needs 2^63 operations and AFAIK no collision has been\nfound yet, finding collision for MD5 can be achieved in a minute and\nless (see http://cryptography.hyperlink.cz/2006/tunnels.pdf for details).\n\nDavid Brodsky\n"},{"id":"25543","messageId":"Pine.LNX.4.63.0608181255060.28360@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"5317","inReplyTo":"44E59BF6.2070909@sinister.cz","subject":"Re: [RFC] adding support for md5","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-08-18T10:56:07Z","receivedAt":"2006-08-18T10:56:07Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 18 Aug 2006, Trekie wrote:\n\n> I'd like to point out that while finding collision for SHA1 according to\n> the Wikipedia needs 2^63 operations and AFAIK no collision has been\n> found yet, finding collision for MD5 can be achieved in a minute and\n> less (see http://cryptography.hyperlink.cz/2006/tunnels.pdf for details).\n\nSHA1 has been broken (collisions have been found):\n\nhttp://www.schneier.com/blog/archives/2005/02/sha1_broken.html\n\nCiao,\nDscho\n"},{"id":"25544","messageId":"44E5A416.9040709@sinister.cz","threadId":"5317","inReplyTo":"Pine.LNX.4.63.0608181255060.28360@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [RFC] adding support for md5","fromName":"Trekie","fromEmail":"trekie@sinister.cz","sentAt":"2006-08-18T11:27:18Z","receivedAt":"2006-08-18T11:27:18Z","isPatch":false,"sender":{"key":"trekie@sinister.cz","avatar":null},"body":"Johannes Schindelin wrote:\n> SHA1 has been broken (collisions have been found):\n> \n> http://www.schneier.com/blog/archives/2005/02/sha1_broken.html\n\nI don't think you're right. That blog just says, that Wang can find\n\n\"collisions in the the full SHA-1 in 2**69 hash operations, much less\nthan the brute-force attack of 2**80 operations based on the hash length.\"\n\nThat doesn't mean any collision has been found. In academic\ncryptography, any attack that has less computational complexity than the\nexpected time needed for brute force is considered a break.\n\nIn a document (http://www.rsasecurity.com/rsalabs/node.asp?id=2927) that\nhas been released 6 months after that blog post is said a collision can\nbe found in 2^63 operations.\n\nWell, if someone use the fastest computer today\n(http://www.top500.org/system/7747) to get a collision it would take a\nday to found one.\n\nThe point is why use MD5 if anyone can compute a collision?\n\nDavid Brodsky\n"},{"id":"25545","messageId":"Pine.LNX.4.63.0608181330210.28360@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"5317","inReplyTo":"44E5A416.9040709@sinister.cz","subject":"Re: [RFC] adding support for md5","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-08-18T11:37:58Z","receivedAt":"2006-08-18T11:37:58Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 18 Aug 2006, Trekie wrote:\n\n> Johannes Schindelin wrote:\n> > SHA1 has been broken (collisions have been found):\n> > \n> > http://www.schneier.com/blog/archives/2005/02/sha1_broken.html\n> \n> I don't think you're right. That blog just says, that Wang can find\n> \n> \"collisions in the the full SHA-1 in 2**69 hash operations, much less\n> than the brute-force attack of 2**80 operations based on the hash length.\"\n\nTrue. I have not heard of a collision either.\n\n> The point is why use MD5 if anyone can compute a collision?\n\nIt does not suffice to generate collisions to make a hash unusable for our \npurposes: you would have to find a way to produce another text for a \n_given_ hash. Plus, this text would not only have to look meaningful, but \ncompile. And preferrably introduce a back door.\n\nGranted, once people find out how to generate another text, they can try \nto \"optimize\" some block between \"/*\" and \"*/\", so that the hash stays the \nsame. But AFAICT none of the breaks of SHA1 or MD5 point into such a \ndirection. Yet.\n\nBut _even if_ somebody succeeds in all that, that somebody has to convince \n_you_ to pull. And if you already have that object (the \"good\" version), \nit will not get overwritten.\n\nCiao,\nDscho\n"},{"id":"25548","messageId":"20060818123110.GQ13776@pasky.or.cz","threadId":"5317","inReplyTo":"Pine.LNX.4.63.0608181209210.28360@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [RFC] adding support for md5","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-08-18T12:31:10Z","receivedAt":"2006-08-18T12:31:10Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Hi,\n\nDear diary, on Fri, Aug 18, 2006 at 12:21:11PM CEST, I got a letter\nwhere Johannes Schindelin <Johannes.Schindelin@gmx.de> said that...\n> However, you should know that there is _no way_ to use both hashes on the \n> same project. Yes, you could rewrite the history, trying to convert also \n> the hashes in the commit objects, but people actually started relying on \n> naming commits with the short-SHA1.\n\nI don't really like having IDs ambiguous in this sense - having the same\ntype of IDs in all git-tracked projects has some cute benefits which are\nof the kind that you don't know ahead that you will need them: joining\nhistory of two distinct projects in a merge and theoretical possibility\nof having subprojects where the main project references an exact\ntree/commit of the sub project.\n\nIf we are ever going to implement support for multiple hashes, the hash\ntype should at least be part of the object id, in textual representation\nas e.g. the first letter. This can still lead to convergence issues and\nduplicate objects, but it enables smooth transition without rewriting\nthe history and it is much less confusing than just switching to a\ndifferent function.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nSnow falling on Perl. White noise covering line noise.\nHides all the bugs too. -- J. Putnam\n"},{"id":"25588","messageId":"Pine.LNX.4.63.0608181328160.30860@chino.corp.google.com","threadId":"5317","inReplyTo":"Pine.LNX.4.63.0608181209210.28360@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [RFC] adding support for md5","fromName":"David Rientjes","fromEmail":"rientjes@google.com","sentAt":"2006-08-18T20:35:03Z","receivedAt":"2006-08-18T20:35:03Z","isPatch":false,"sender":{"key":"rientjes@google.com","avatar":null},"body":"On Fri, 18 Aug 2006, Johannes Schindelin wrote:\n\n> Make it a config variable, too, right?\n> \n\nSure.  The default hash function can be a config variable so that all \nprojects started with init-db will default to a specific hash.  Other \nprojects may still be started with something like init-db -md5.\n\n> 1. they could be faster to calculate,\n> 2. they could reduce clashes, and related to that,\n> 3. it is possible that some day SHA1 is broken, i.e. that there is an \n>    algorithm to generate a different text for a given hash.\n> \n> As for 2 and 3, it seems MD5 is equivalent, since another sort of attacks \n> was already successful on both SHA1 and MD5: generating two different \n> texts with the same hash.\n> \n\nCorrect; performance was my main motivation.  sha1 is obviously the \nsecurest algorithm among the two choices, but there are more steps \ninvolved in the hash than md5 (sha1 uses 80 and md5 uses 64) and sha1 is \n160-bit compared to the 128-bit md5.  One paper I read from the \nInformation Technology Journal stated that sha1 is 25% slower than md5 \nprecisely for these reasons.\n\n> However, you should know that there is _no way_ to use both hashes on the \n> same project. Yes, you could rewrite the history, trying to convert also \n> the hashes in the commit objects, but people actually started relying on \n> naming commits with the short-SHA1.\n> \n\nI don't foresee changing a hash on a project (and thus rewriting the \nhistory) to be something that anybody would want to do.  As I said in the \nemail that started this thread, it would be configurable at runtime on \ninit-db.\n\n> I think it would be a nice thing to play through (for example, to find \n> out how much impact the hash calculation has on the overall performance \n> of git), but I doubt it will ever come to real use.\n> \n\nAgain, when working with an enormous amount of data, this could be a \nconsiderable speedup.  A terabyte is _big_.\n\n\t\tDavid\n"},{"id":"25591","messageId":"9e4733910608181452x65ca937aqbfde55caa98ff6da@mail.gmail.com","threadId":"5317","inReplyTo":"Pine.LNX.4.63.0608172259280.25827@chino.corp.google.com","subject":"Re: [RFC] adding support for md5","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-18T21:52:46Z","receivedAt":"2006-08-18T21:52:46Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/18/06, David Rientjes <rientjes@google.com> wrote:\n> I'd like to solicit some comments about implementing support for md5 as a\n> hash function that could be determined at runtime by the user during a\n> project init-db.  md5, which I implemented as a configurable option in my\n> own tree, is a 128-bit hash that is slightly more recognized than sha1.\n> Likewise, it is also available in openssl/md5.h just as sha1 is available\n> through a library in openssl/sha1.h.  My patch to move the hash name\n> comparison was a step in this direction in isolating many of the\n> particulars of hash-specific dependencies.\n\nIf I have two repositories each with 100M objects in them and I merge\nthem, what is the probability of a object id collision with MD5 (128b)\nversus SHA1 (160b)?\n\nThis is not that far fetched. I know of at least four repositories\nwith over 1M objects in them today.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"25597","messageId":"Pine.LNX.4.63.0608190416370.28360@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"5317","inReplyTo":"9e4733910608181452x65ca937aqbfde55caa98ff6da@mail.gmail.com","subject":"Re: [RFC] adding support for md5","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-08-19T02:35:55Z","receivedAt":"2006-08-19T02:35:55Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 18 Aug 2006, Jon Smirl wrote:\n\n> If I have two repositories each with 100M objects in them and I merge \n> them, what is the probability of a object id collision with MD5 (128b) \n> versus SHA1 (160b)?\n\nAssuming a uniform distribution of the hashes over our data, this is the \nbirthday problem:\n\nhttp://mathworld.wolfram.com/BirthdayProblem.html\n\n(In short, given a number of days in the year, how many people do I need \nto pick randomly until at least two of them have the same birthday?)\n\nIn our case, we want to know how many objects we need in order to probably \nhave a clash in 2^128 (approx. 3.4e38) and 2^160 (approx. 1.5e48) hashes, \nrespectively.\n\nMathworld tells us that a good approximation of the probability is\n\np = 1 - (1-n/(2d))^(n-1)\n\nwhere n is the number of objects, and d is the total number of hashes. If \nyou have 100M = 1e5 objects, you probably want the probability of a clash \nbelow 1/1e5 = 1e-5, so let's take 1e-10. Assuming n is way lower than d, \nwe can approximate\n\np = 1 - (1 - (n - 1 over 1) * n/(2d)) = n(n-1)/2d\n\nand therefore (approximately)\n\nn = sqrt(2pd)\n\nwhich amounts to 2.6e14 in the case of a 128-bit hash, and 1.7e19 in the \ncase of a 160-bit hash, both well beyond your 100M objects. BTW the \naddressable space of a 64-bit processor is about 1.9e19.\n\nIf you want to know the probability of a clash, you can use the same \napproximation:\n\nFor 100M objects: p = 1.5e-59 for 128-bit, and p = 3.3e-69 for 160-bit. \nThis is so low as to be incomprehensible.\n\nRemember that all these approximations are really crude, so do not rely on \nthe precise numbers. But they'll give you good ballpark figures (if I did \nnot make a mistake...).\n\nHth,\nDscho\n"},{"id":"25627","messageId":"Pine.LNX.4.64.0608191339010.11811@g5.osdl.org","threadId":"5317","inReplyTo":"Pine.LNX.4.63.0608172259280.25827@chino.corp.google.com","subject":"Re: [RFC] adding support for md5","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-19T20:50:32Z","receivedAt":"2006-08-19T20:50:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 17 Aug 2006, David Rientjes wrote:\n>\n> I'd like to solicit some comments about implementing support for md5 as a \n> hash function that could be determined at runtime by the user during a \n> project init-db.\n\nI would _strongly_ suggest against this. At least not md5. \n\nI can see the point of configurable hashes, but it would be for a stronger \nhash than sha1, not for a (much) weaker one.\n\nmd5 is not only shorter, it's known to be broken, and there are attacks \nout there that generate documents with the same md5 checksum quickly and \nundetectably (ie depending on what the \"document format\" is, you might \nactually not _see_ the corruption).\n\nThere's a real-life example of this (just google for \"same md5\") with a \npostscript file, which when printed out still looks \"valid\".\n\nIn contrast, sha1 is still considered \"hard\", in that while you can \nobviously always brute-force _any_ hash, the sha1 brute-forcing attack is \nconsidered to be impractical and nobody has at least shown any realistic \nversion of the above postscript kind of hack.\n\nIn my fairly limited performance analysis, I've actually been surprised by \nthe fact that the hashing has never really shown up as a major issue in \nany of my profiles. All the _real_ performance issues have been related to \nmemory usage, and things like the hash lookup (ie \"memcmp()\" was pretty \nhigh on the list - just from comparing object names during lookup).\n\nWe've also had compression issues (initial check-in) and obviously the \ndelta selection used to be a _huge_ time-waster until the pack info reuse \ncode went in. But I don't think we've ever had a load that was really \nhashing-limited.\n\nSo considering that md5 isn't _that_ much faster to compute (let's say \nthat it's ~30% slower), the biggest advantage of md5 would likely be just \nthe fact that 16 bytes is smaller than 20 bytes, and thus commit objects \nand tree objects in particular could be smaller. But you'd be better off \njust using the first 16 bytes of the sha1 than the md5 hash, if that was \nthe main goal.\n\nSo yes, maybe we'll want to make the hash choice a setup-time option, but \nif we ever do, I don't think we should make md5 even a choice. It's just \nnot a very good hash, and no new program should start using it. \n\n\t\t\tLinus\n"},{"id":"25709","messageId":"20060821204430.GA2700@tuatara.stupidest.org","threadId":"5317","inReplyTo":"Pine.LNX.4.64.0608191339010.11811@g5.osdl.org","subject":"Re: [RFC] adding support for md5","fromName":"Chris Wedgwood","fromEmail":"cw@f00f.org","sentAt":"2006-08-21T20:44:30Z","receivedAt":"2006-08-21T20:44:30Z","isPatch":false,"sender":{"key":"cw@f00f.org","avatar":null},"body":"On Sat, Aug 19, 2006 at 01:50:32PM -0700, Linus Torvalds wrote:\n\n> I can see the point of configurable hashes, but it would be for a\n> stronger hash than sha1, not for a (much) weaker one.\n\nWhy any configuration option at all?  What in practice does it really\nbuy?\n\nIf someone (eventually) wants to do something malicious (which right\nnow requires some effort and would probably not go undetected) there\nare probably easier ways to achieve this (like posting a patch with a\nnon-obvious subtle side-effect).\n"},{"id":"25724","messageId":"7vr6z9s376.fsf@assigned-by-dhcp.cox.net","threadId":"5317","inReplyTo":"20060821204430.GA2700@tuatara.stupidest.org","subject":"Re: [RFC] adding support for md5","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-08-22T06:18:53Z","receivedAt":"2006-08-22T06:18:53Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Chris Wedgwood <cw@f00f.org> writes:\n\n> On Sat, Aug 19, 2006 at 01:50:32PM -0700, Linus Torvalds wrote:\n>\n>> I can see the point of configurable hashes, but it would be for a\n>> stronger hash than sha1, not for a (much) weaker one.\n>\n> Why any configuration option at all?  What in practice does it really\n> buy?\n\nI personally am not interested in making this configurable at\nall.  The hashcmp() change on the other hand to abstract out 20\nwas a good preparation, if we ever want to switch to longer\nhashes we would know where to look.\n"},{"id":"25782","messageId":"20060823041453.GA25796@spearce.org","threadId":"5317","inReplyTo":"7vr6z9s376.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] adding support for md5","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-08-23T04:14:53Z","receivedAt":"2006-08-23T04:14:53Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <junkio@cox.net> wrote:\n> I personally am not interested in making this configurable at\n> all.  The hashcmp() change on the other hand to abstract out 20\n> was a good preparation, if we ever want to switch to longer\n> hashes we would know where to look.\n\nWhat about all of those memcpy(a, b, 20)'s?  :-)\n\nI can see us wanting to support say SHA-128 or SHA-256 in a few\nyears.  Especially as processors get faster and better attacks are\ndeveloped against SHA-1 such that its no longer really the best\ntrade-off hash function available.\n\n-- \nShawn.\n"},{"id":"25785","messageId":"7v3bbojbzj.fsf@assigned-by-dhcp.cox.net","threadId":"5317","inReplyTo":"20060823041453.GA25796@spearce.org","subject":"Re: [RFC] adding support for md5","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-08-23T04:46:08Z","receivedAt":"2006-08-23T04:46:08Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n> Junio C Hamano <junkio@cox.net> wrote:\n>> I personally am not interested in making this configurable at\n>> all.  The hashcmp() change on the other hand to abstract out 20\n>> was a good preparation, if we ever want to switch to longer\n>> hashes we would know where to look.\n>\n> What about all of those memcpy(a, b, 20)'s?  :-)\n\nSurely.  If you are inclined to, go wild.\n\n> I can see us wanting to support say SHA-128 or SHA-256 in a few\n> years.  Especially as processors get faster and better attacks are\n> developed against SHA-1 such that its no longer really the best\n> trade-off hash function available.\n"},{"id":"25789","messageId":"20060823064900.GA26340@spearce.org","threadId":"5317","inReplyTo":"7v3bbojbzj.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] adding support for md5","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-08-23T06:49:00Z","receivedAt":"2006-08-23T06:49:00Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <junkio@cox.net> wrote:\n> Shawn Pearce <spearce@spearce.org> writes:\n> \n> > Junio C Hamano <junkio@cox.net> wrote:\n> >> I personally am not interested in making this configurable at\n> >> all.  The hashcmp() change on the other hand to abstract out 20\n> >> was a good preparation, if we ever want to switch to longer\n> >> hashes we would know where to look.\n> >\n> > What about all of those memcpy(a, b, 20)'s?  :-)\n> \n> Surely.  If you are inclined to, go wild.\n\nLike this?  :-)\n\n-->--\nConvert memcpy(a,b,20) to hashcpy(a,b).\n\nThis abstracts away the size of the hash values when copying them\nfrom memory location to memory location, much as the introduction\nof hashcmp abstracted away hash value comparsion.\n\nA few call sites were using char* rather than unsigned char* so\nI added the cast rather than open hashcpy to be void*.  This is a\nreasonable tradeoff as most call sites already use unsigned char*\nand the existing hashcmp is also declared to be unsigned char*.\n\nSigned-off-by: Shawn O. Pearce <spearce@spearce.org>\n---\n blame.c                  |    4 ++--\n builtin-diff.c           |    4 ++--\n builtin-pack-objects.c   |    6 +++---\n builtin-read-tree.c      |    2 +-\n builtin-unpack-objects.c |    4 ++--\n builtin-update-index.c   |    4 ++--\n builtin-write-tree.c     |    4 ++--\n cache-tree.c             |    6 +++---\n cache.h                  |    4 ++++\n combine-diff.c           |    6 +++---\n connect.c                |    6 +++---\n convert-objects.c        |    6 +++---\n csum-file.c              |    2 +-\n diff.c                   |    4 ++--\n fetch-pack.c             |    2 +-\n fetch.c                  |    2 +-\n fsck-objects.c           |    2 +-\n http-fetch.c             |    2 +-\n http-push.c              |    6 +++---\n index-pack.c             |    4 ++--\n merge-recursive.c        |   20 ++++++++++----------\n mktree.c                 |    4 ++--\n object.c                 |    2 +-\n patch-id.c               |    2 +-\n receive-pack.c           |    4 ++--\n revision.c               |    2 +-\n send-pack.c              |    4 ++--\n sha1_file.c              |   18 +++++++++---------\n sha1_name.c              |   22 +++++++++++-----------\n tree-walk.c              |    4 ++--\n tree.c                   |    2 +-\n unpack-trees.c           |    2 +-\n 32 files changed, 85 insertions(+), 81 deletions(-)\n\ndiff --git a/blame.c b/blame.c\nindex c253b9c..5a8af72 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -176,7 +176,7 @@ static int get_blob_sha1(struct tree *t,\n \tif (i == 20)\n \t\treturn -1;\n \n-\tmemcpy(sha1, blob_sha1, 20);\n+\thashcpy(sha1, blob_sha1);\n \treturn 0;\n }\n \n@@ -191,7 +191,7 @@ static int get_blob_sha1_internal(const \n \t    strcmp(blame_file + baselen, pathname))\n \t\treturn -1;\n \n-\tmemcpy(blob_sha1, sha1, 20);\n+\thashcpy(blob_sha1, sha1);\n \treturn -1;\n }\n \ndiff --git a/builtin-diff.c b/builtin-diff.c\nindex 874f773..a659020 100644\n--- a/builtin-diff.c\n+++ b/builtin-diff.c\n@@ -192,7 +192,7 @@ static int builtin_diff_combined(struct \n \tparent = xmalloc(ents * sizeof(*parent));\n \t/* Again, the revs are all reverse */\n \tfor (i = 0; i < ents; i++)\n-\t\tmemcpy(parent + i, ent[ents - 1 - i].item->sha1, 20);\n+\t\thashcpy((unsigned char*)parent + i, ent[ents - 1 - i].item->sha1);\n \tdiff_tree_combined(parent[0], parent + 1, ents - 1,\n \t\t\t   revs->dense_combined_merges, revs);\n \treturn 0;\n@@ -290,7 +290,7 @@ int cmd_diff(int argc, const char **argv\n \t\tif (obj->type == OBJ_BLOB) {\n \t\t\tif (2 <= blobs)\n \t\t\t\tdie(\"more than two blobs given: '%s'\", name);\n-\t\t\tmemcpy(blob[blobs].sha1, obj->sha1, 20);\n+\t\t\thashcpy(blob[blobs].sha1, obj->sha1);\n \t\t\tblob[blobs].name = name;\n \t\t\tblobs++;\n \t\t\tcontinue;\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex f19f0d6..46f524d 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -534,7 +534,7 @@ static int add_object_entry(const unsign\n \tentry = objects + idx;\n \tnr_objects = idx + 1;\n \tmemset(entry, 0, sizeof(*entry));\n-\tmemcpy(entry->sha1, sha1, 20);\n+\thashcpy(entry->sha1, sha1);\n \tentry->hash = hash;\n \n \tif (object_ix_hashsz * 3 <= nr_objects * 4)\n@@ -649,7 +649,7 @@ static struct pbase_tree_cache *pbase_tr\n \t\tfree(ent->tree_data);\n \t\tnent = ent;\n \t}\n-\tmemcpy(nent->sha1, sha1, 20);\n+\thashcpy(nent->sha1, sha1);\n \tnent->tree_data = data;\n \tnent->tree_size = size;\n \tnent->ref = 1;\n@@ -799,7 +799,7 @@ static void add_preferred_base(unsigned \n \tit->next = pbase_tree;\n \tpbase_tree = it;\n \n-\tmemcpy(it->pcache.sha1, tree_sha1, 20);\n+\thashcpy(it->pcache.sha1, tree_sha1);\n \tit->pcache.tree_data = data;\n \tit->pcache.tree_size = size;\n }\ndiff --git a/builtin-read-tree.c b/builtin-read-tree.c\nindex 53087fa..c1867d2 100644\n--- a/builtin-read-tree.c\n+++ b/builtin-read-tree.c\n@@ -53,7 +53,7 @@ static void prime_cache_tree_rec(struct \n \tstruct name_entry entry;\n \tint cnt;\n \n-\tmemcpy(it->sha1, tree->object.sha1, 20);\n+\thashcpy(it->sha1, tree->object.sha1);\n \tdesc.buf = tree->buffer;\n \tdesc.size = tree->size;\n \tcnt = 0;\ndiff --git a/builtin-unpack-objects.c b/builtin-unpack-objects.c\nindex f0ae5c9..ca0ebc2 100644\n--- a/builtin-unpack-objects.c\n+++ b/builtin-unpack-objects.c\n@@ -95,7 +95,7 @@ static void add_delta_to_list(unsigned c\n {\n \tstruct delta_info *info = xmalloc(sizeof(*info));\n \n-\tmemcpy(info->base_sha1, base_sha1, 20);\n+\thashcpy(info->base_sha1, base_sha1);\n \tinfo->size = size;\n \tinfo->delta = delta;\n \tinfo->next = delta_list;\n@@ -173,7 +173,7 @@ static int unpack_delta_entry(unsigned l\n \tunsigned char base_sha1[20];\n \tint result;\n \n-\tmemcpy(base_sha1, fill(20), 20);\n+\thashcpy(base_sha1, fill(20));\n \tuse(20);\n \n \tdelta_data = get_data(delta_size);\ndiff --git a/builtin-update-index.c b/builtin-update-index.c\nindex 5dd91af..8675126 100644\n--- a/builtin-update-index.c\n+++ b/builtin-update-index.c\n@@ -142,7 +142,7 @@ static int add_cacheinfo(unsigned int mo\n \tsize = cache_entry_size(len);\n \tce = xcalloc(1, size);\n \n-\tmemcpy(ce->sha1, sha1, 20);\n+\thashcpy(ce->sha1, sha1);\n \tmemcpy(ce->name, path, len);\n \tce->ce_flags = create_ce_flags(len, stage);\n \tce->ce_mode = create_ce_mode(mode);\n@@ -333,7 +333,7 @@ static struct cache_entry *read_one_ent(\n \tsize = cache_entry_size(namelen);\n \tce = xcalloc(1, size);\n \n-\tmemcpy(ce->sha1, sha1, 20);\n+\thashcpy(ce->sha1, sha1);\n \tmemcpy(ce->name, path, namelen);\n \tce->ce_flags = create_ce_flags(namelen, stage);\n \tce->ce_mode = create_ce_mode(mode);\ndiff --git a/builtin-write-tree.c b/builtin-write-tree.c\nindex ca06149..50670dc 100644\n--- a/builtin-write-tree.c\n+++ b/builtin-write-tree.c\n@@ -50,10 +50,10 @@ int write_tree(unsigned char *sha1, int \n \tif (prefix) {\n \t\tstruct cache_tree *subtree =\n \t\t\tcache_tree_find(active_cache_tree, prefix);\n-\t\tmemcpy(sha1, subtree->sha1, 20);\n+\t\thashcpy(sha1, subtree->sha1);\n \t}\n \telse\n-\t\tmemcpy(sha1, active_cache_tree->sha1, 20);\n+\t\thashcpy(sha1, active_cache_tree->sha1);\n \n \trollback_lock_file(lock_file);\n \ndiff --git a/cache-tree.c b/cache-tree.c\nindex d9f7e1e..323c68a 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -335,7 +335,7 @@ static int update_one(struct cache_tree \n \t\toffset += sprintf(buffer + offset,\n \t\t\t\t  \"%o %.*s\", mode, entlen, path + baselen);\n \t\tbuffer[offset++] = 0;\n-\t\tmemcpy(buffer + offset, sha1, 20);\n+\t\thashcpy((unsigned char*)buffer + offset, sha1);\n \t\toffset += 20;\n \n #if DEBUG\n@@ -412,7 +412,7 @@ #if DEBUG\n #endif\n \n \tif (0 <= it->entry_count) {\n-\t\tmemcpy(buffer + *offset, it->sha1, 20);\n+\t\thashcpy((unsigned char*)buffer + *offset, it->sha1);\n \t\t*offset += 20;\n \t}\n \tfor (i = 0; i < it->subtree_nr; i++) {\n@@ -478,7 +478,7 @@ static struct cache_tree *read_one(const\n \tif (0 <= it->entry_count) {\n \t\tif (size < 20)\n \t\t\tgoto free_return;\n-\t\tmemcpy(it->sha1, buf, 20);\n+\t\thashcpy(it->sha1, (unsigned char*)buf);\n \t\tbuf += 20;\n \t\tsize -= 20;\n \t}\ndiff --git a/cache.h b/cache.h\nindex cd2ad90..df642a1 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -222,6 +222,10 @@ static inline int hashcmp(const unsigned\n {\n \treturn memcmp(sha1, sha2, 20);\n }\n+static inline void hashcpy(unsigned char *sha_dst, const unsigned char *sha_src)\n+{\n+\tmemcpy(sha_dst, sha_src, 20);\n+}\n \n int git_mkstemp(char *path, size_t n, const char *template);\n \ndiff --git a/combine-diff.c b/combine-diff.c\nindex 0682acd..466899d 100644\n--- a/combine-diff.c\n+++ b/combine-diff.c\n@@ -31,9 +31,9 @@ static struct combine_diff_path *interse\n \t\t\tmemset(p->parent, 0,\n \t\t\t       sizeof(p->parent[0]) * num_parent);\n \n-\t\t\tmemcpy(p->sha1, q->queue[i]->two->sha1, 20);\n+\t\t\thashcpy(p->sha1, q->queue[i]->two->sha1);\n \t\t\tp->mode = q->queue[i]->two->mode;\n-\t\t\tmemcpy(p->parent[n].sha1, q->queue[i]->one->sha1, 20);\n+\t\t\thashcpy(p->parent[n].sha1, q->queue[i]->one->sha1);\n \t\t\tp->parent[n].mode = q->queue[i]->one->mode;\n \t\t\tp->parent[n].status = q->queue[i]->status;\n \t\t\t*tail = p;\n@@ -927,6 +927,6 @@ void diff_tree_combined_merge(const unsi\n \tfor (parents = commit->parents, num_parent = 0;\n \t     parents;\n \t     parents = parents->next, num_parent++)\n-\t\tmemcpy(parent + num_parent, parents->item->object.sha1, 20);\n+\t\thashcpy((unsigned char*)parent + num_parent, parents->item->object.sha1);\n \tdiff_tree_combined(sha1, parent, num_parent, dense, rev);\n }\ndiff --git a/connect.c b/connect.c\nindex 7a6a73f..e501ccc 100644\n--- a/connect.c\n+++ b/connect.c\n@@ -77,7 +77,7 @@ struct ref **get_remote_heads(int in, st\n \t\tif (nr_match && !path_match(name, nr_match, match))\n \t\t\tcontinue;\n \t\tref = xcalloc(1, sizeof(*ref) + len - 40);\n-\t\tmemcpy(ref->old_sha1, old_sha1, 20);\n+\t\thashcpy(ref->old_sha1, old_sha1);\n \t\tmemcpy(ref->name, buffer + 41, len - 40);\n \t\t*list = ref;\n \t\tlist = &ref->next;\n@@ -208,7 +208,7 @@ static struct ref *try_explicit_object_n\n \tlen = strlen(name) + 1;\n \tref = xcalloc(1, sizeof(*ref) + len);\n \tmemcpy(ref->name, name, len);\n-\tmemcpy(ref->new_sha1, sha1, 20);\n+\thashcpy(ref->new_sha1, sha1);\n \treturn ref;\n }\n \n@@ -318,7 +318,7 @@ int match_refs(struct ref *src, struct r\n \t\t\tint len = strlen(src->name) + 1;\n \t\t\tdst_peer = xcalloc(1, sizeof(*dst_peer) + len);\n \t\t\tmemcpy(dst_peer->name, src->name, len);\n-\t\t\tmemcpy(dst_peer->new_sha1, src->new_sha1, 20);\n+\t\t\thashcpy(dst_peer->new_sha1, src->new_sha1);\n \t\t\tlink_dst_tail(dst_peer, dst_tail);\n \t\t}\n \t\tdst_peer->peer_ref = src;\ndiff --git a/convert-objects.c b/convert-objects.c\nindex 4e7ff75..631678b 100644\n--- a/convert-objects.c\n+++ b/convert-objects.c\n@@ -23,7 +23,7 @@ static struct entry * convert_entry(unsi\n static struct entry *insert_new(unsigned char *sha1, int pos)\n {\n \tstruct entry *new = xcalloc(1, sizeof(struct entry));\n-\tmemcpy(new->old_sha1, sha1, 20);\n+\thashcpy(new->old_sha1, sha1);\n \tmemmove(convert + pos + 1, convert + pos, (nr_convert - pos) * sizeof(struct entry *));\n \tconvert[pos] = new;\n \tnr_convert++;\n@@ -54,7 +54,7 @@ static struct entry *lookup_entry(unsign\n static void convert_binary_sha1(void *buffer)\n {\n \tstruct entry *entry = convert_entry(buffer);\n-\tmemcpy(buffer, entry->new_sha1, 20);\n+\thashcpy(buffer, entry->new_sha1);\n }\n \n static void convert_ascii_sha1(void *buffer)\n@@ -104,7 +104,7 @@ static int write_subdirectory(void *buff\n \t\tif (!slash) {\n \t\t\tnewlen += sprintf(new + newlen, \"%o %s\", mode, path);\n \t\t\tnew[newlen++] = '\\0';\n-\t\t\tmemcpy(new + newlen, (char *) buffer + len - 20, 20);\n+\t\t\thashcpy((unsigned char*)new + newlen, (unsigned char *) buffer + len - 20);\n \t\t\tnewlen += 20;\n \n \t\t\tused += len;\ndiff --git a/csum-file.c b/csum-file.c\nindex e227889..b7174c6 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -38,7 +38,7 @@ int sha1close(struct sha1file *f, unsign\n \t}\n \tSHA1_Final(f->buffer, &f->ctx);\n \tif (result)\n-\t\tmemcpy(result, f->buffer, 20);\n+\t\thashcpy(result, f->buffer);\n \tif (update)\n \t\tsha1flush(f, 20);\n \tif (close(f->fd))\ndiff --git a/diff.c b/diff.c\nindex ddf2dea..852c175 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -1107,7 +1107,7 @@ void fill_filespec(struct diff_filespec \n {\n \tif (mode) {\n \t\tspec->mode = canon_mode(mode);\n-\t\tmemcpy(spec->sha1, sha1, 20);\n+\t\thashcpy(spec->sha1, sha1);\n \t\tspec->sha1_valid = !is_null_sha1(sha1);\n \t}\n }\n@@ -1200,7 +1200,7 @@ static struct sha1_size_cache *locate_si\n \t\t\tsizeof(*sha1_size_cache));\n \te = xmalloc(sizeof(struct sha1_size_cache));\n \tsha1_size_cache[first] = e;\n-\tmemcpy(e->sha1, sha1, 20);\n+\thashcpy(e->sha1, sha1);\n \te->size = size;\n \treturn e;\n }\ndiff --git a/fetch-pack.c b/fetch-pack.c\nindex e18c148..377fede 100644\n--- a/fetch-pack.c\n+++ b/fetch-pack.c\n@@ -404,7 +404,7 @@ static int everything_local(struct ref *\n \t\t\tcontinue;\n \t\t}\n \n-\t\tmemcpy(ref->new_sha1, local, 20);\n+\t\thashcpy(ref->new_sha1, local);\n \t\tif (!verbose)\n \t\t\tcontinue;\n \t\tfprintf(stderr,\ndiff --git a/fetch.c b/fetch.c\nindex aeb6bf2..ef60b04 100644\n--- a/fetch.c\n+++ b/fetch.c\n@@ -84,7 +84,7 @@ static int process_commit(struct commit \n \tif (commit->object.flags & COMPLETE)\n \t\treturn 0;\n \n-\tmemcpy(current_commit_sha1, commit->object.sha1, 20);\n+\thashcpy(current_commit_sha1, commit->object.sha1);\n \n \tpull_say(\"walk %s\\n\", sha1_to_hex(commit->object.sha1));\n \ndiff --git a/fsck-objects.c b/fsck-objects.c\nindex 31e00d8..ae0ec8d 100644\n--- a/fsck-objects.c\n+++ b/fsck-objects.c\n@@ -356,7 +356,7 @@ static void add_sha1_list(unsigned char \n \tint nr;\n \n \tentry->ino = ino;\n-\tmemcpy(entry->sha1, sha1, 20);\n+\thashcpy(entry->sha1, sha1);\n \tnr = sha1_list.nr;\n \tif (nr == MAX_SHA1_ENTRIES) {\n \t\tfsck_sha1_list();\ndiff --git a/http-fetch.c b/http-fetch.c\nindex d1f74b4..7619b33 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -393,7 +393,7 @@ void prefetch(unsigned char *sha1)\n \tchar *filename = sha1_file_name(sha1);\n \n \tnewreq = xmalloc(sizeof(*newreq));\n-\tmemcpy(newreq->sha1, sha1, 20);\n+\thashcpy(newreq->sha1, sha1);\n \tnewreq->repo = alt;\n \tnewreq->url = NULL;\n \tnewreq->local = -1;\ndiff --git a/http-push.c b/http-push.c\nindex 4849779..ebfcc73 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -1874,7 +1874,7 @@ static int one_local_ref(const char *ref\n \tstruct ref *ref;\n \tint len = strlen(refname) + 1;\n \tref = xcalloc(1, sizeof(*ref) + len);\n-\tmemcpy(ref->new_sha1, sha1, 20);\n+\thashcpy(ref->new_sha1, sha1);\n \tmemcpy(ref->name, refname, len);\n \t*local_tail = ref;\n \tlocal_tail = &ref->next;\n@@ -1909,7 +1909,7 @@ static void one_remote_ref(char *refname\n \t}\n \n \tref = xcalloc(1, sizeof(*ref) + len);\n-\tmemcpy(ref->old_sha1, remote_sha1, 20);\n+\thashcpy(ref->old_sha1, remote_sha1);\n \tmemcpy(ref->name, refname, len);\n \t*remote_tail = ref;\n \tremote_tail = &ref->next;\n@@ -2445,7 +2445,7 @@ int main(int argc, char **argv)\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t}\n-\t\tmemcpy(ref->new_sha1, ref->peer_ref->new_sha1, 20);\n+\t\thashcpy(ref->new_sha1, ref->peer_ref->new_sha1);\n \t\tif (is_zero_sha1(ref->new_sha1)) {\n \t\t\terror(\"cannot happen anymore\");\n \t\t\trc = -3;\ndiff --git a/index-pack.c b/index-pack.c\nindex 96ea687..80bc6cb 100644\n--- a/index-pack.c\n+++ b/index-pack.c\n@@ -161,7 +161,7 @@ static void *unpack_raw_entry(unsigned l\n \tcase OBJ_DELTA:\n \t\tif (pos + 20 >= pack_limit)\n \t\t\tbad_object(offset, \"object extends past end of pack\");\n-\t\tmemcpy(delta_base, pack_base + pos, 20);\n+\t\thashcpy(delta_base, pack_base + pos);\n \t\tpos += 20;\n \t\t/* fallthru */\n \tcase OBJ_COMMIT:\n@@ -304,7 +304,7 @@ static void parse_pack_objects(void)\n \t\tif (obj->type == OBJ_DELTA) {\n \t\t\tstruct delta_entry *delta = &deltas[nr_deltas++];\n \t\t\tdelta->obj = obj;\n-\t\t\tmemcpy(delta->base_sha1, base_sha1, 20);\n+\t\t\thashcpy(delta->base_sha1, base_sha1);\n \t\t} else\n \t\t\tsha1_object(data, data_size, obj->type, obj->sha1);\n \t\tfree(data);\ndiff --git a/merge-recursive.c b/merge-recursive.c\nindex 048cca1..8a2f697 100644\n--- a/merge-recursive.c\n+++ b/merge-recursive.c\n@@ -158,7 +158,7 @@ static struct cache_entry *make_cache_en\n \tsize = cache_entry_size(len);\n \tce = xcalloc(1, size);\n \n-\tmemcpy(ce->sha1, sha1, 20);\n+\thashcpy(ce->sha1, sha1);\n \tmemcpy(ce->name, path, len);\n \tce->ce_flags = create_ce_flags(len, stage);\n \tce->ce_mode = create_ce_mode(mode);\n@@ -355,7 +355,7 @@ static struct path_list *get_unmerged(vo\n \t\t}\n \t\te = item->util;\n \t\te->stages[ce_stage(ce)].mode = ntohl(ce->ce_mode);\n-\t\tmemcpy(e->stages[ce_stage(ce)].sha, ce->sha1, 20);\n+\t\thashcpy(e->stages[ce_stage(ce)].sha, ce->sha1);\n \t}\n \n \treturn unmerged;\n@@ -636,10 +636,10 @@ static struct merge_file_info merge_file\n \t\tresult.clean = 0;\n \t\tif (S_ISREG(a->mode)) {\n \t\t\tresult.mode = a->mode;\n-\t\t\tmemcpy(result.sha, a->sha1, 20);\n+\t\t\thashcpy(result.sha, a->sha1);\n \t\t} else {\n \t\t\tresult.mode = b->mode;\n-\t\t\tmemcpy(result.sha, b->sha1, 20);\n+\t\t\thashcpy(result.sha, b->sha1);\n \t\t}\n \t} else {\n \t\tif (!sha_eq(a->sha1, o->sha1) && !sha_eq(b->sha1, o->sha1))\n@@ -648,9 +648,9 @@ static struct merge_file_info merge_file\n \t\tresult.mode = a->mode == o->mode ? b->mode: a->mode;\n \n \t\tif (sha_eq(a->sha1, o->sha1))\n-\t\t\tmemcpy(result.sha, b->sha1, 20);\n+\t\t\thashcpy(result.sha, b->sha1);\n \t\telse if (sha_eq(b->sha1, o->sha1))\n-\t\t\tmemcpy(result.sha, a->sha1, 20);\n+\t\t\thashcpy(result.sha, a->sha1);\n \t\telse if (S_ISREG(a->mode)) {\n \t\t\tint code = 1, fd;\n \t\t\tstruct stat st;\n@@ -699,7 +699,7 @@ static struct merge_file_info merge_file\n \t\t\tif (!(S_ISLNK(a->mode) || S_ISLNK(b->mode)))\n \t\t\t\tdie(\"cannot merge modes?\");\n \n-\t\t\tmemcpy(result.sha, a->sha1, 20);\n+\t\t\thashcpy(result.sha, a->sha1);\n \n \t\t\tif (!sha_eq(a->sha1, b->sha1))\n \t\t\t\tresult.clean = 0;\n@@ -1096,11 +1096,11 @@ static int process_entry(const char *pat\n \n \t\toutput(\"Auto-merging %s\", path);\n \t\to.path = a.path = b.path = (char *)path;\n-\t\tmemcpy(o.sha1, o_sha, 20);\n+\t\thashcpy(o.sha1, o_sha);\n \t\to.mode = o_mode;\n-\t\tmemcpy(a.sha1, a_sha, 20);\n+\t\thashcpy(a.sha1, a_sha);\n \t\ta.mode = a_mode;\n-\t\tmemcpy(b.sha1, b_sha, 20);\n+\t\thashcpy(b.sha1, b_sha);\n \t\tb.mode = b_mode;\n \n \t\tmfi = merge_file(&o, &a, &b,\ndiff --git a/mktree.c b/mktree.c\nindex 9324138..56205d1 100644\n--- a/mktree.c\n+++ b/mktree.c\n@@ -30,7 +30,7 @@ static void append_to_tree(unsigned mode\n \tent = entries[used++] = xmalloc(sizeof(**entries) + len + 1);\n \tent->mode = mode;\n \tent->len = len;\n-\tmemcpy(ent->sha1, sha1, 20);\n+\thashcpy(ent->sha1, sha1);\n \tmemcpy(ent->name, path, len+1);\n }\n \n@@ -64,7 +64,7 @@ static void write_tree(unsigned char *sh\n \t\toffset += sprintf(buffer + offset, \"%o \", ent->mode);\n \t\toffset += sprintf(buffer + offset, \"%s\", ent->name);\n \t\tbuffer[offset++] = 0;\n-\t\tmemcpy(buffer + offset, ent->sha1, 20);\n+\t\thashcpy((unsigned char*)buffer + offset, ent->sha1);\n \t\toffset += 20;\n \t}\n \twrite_sha1_file(buffer, offset, tree_type, sha1);\ndiff --git a/object.c b/object.c\nindex fdcfff7..60bf16b 100644\n--- a/object.c\n+++ b/object.c\n@@ -91,7 +91,7 @@ void created_object(const unsigned char \n \tobj->used = 0;\n \tobj->type = OBJ_NONE;\n \tobj->flags = 0;\n-\tmemcpy(obj->sha1, sha1, 20);\n+\thashcpy(obj->sha1, sha1);\n \n \tif (obj_hash_size - 1 <= nr_objs * 2)\n \t\tgrow_object_hash();\ndiff --git a/patch-id.c b/patch-id.c\nindex 3b4c80f..086d2d9 100644\n--- a/patch-id.c\n+++ b/patch-id.c\n@@ -47,7 +47,7 @@ static void generate_id_list(void)\n \n \t\tif (!get_sha1_hex(p, n)) {\n \t\t\tflush_current_id(patchlen, sha1, &ctx);\n-\t\t\tmemcpy(sha1, n, 20);\n+\t\t\thashcpy(sha1, n);\n \t\t\tpatchlen = 0;\n \t\t\tcontinue;\n \t\t}\ndiff --git a/receive-pack.c b/receive-pack.c\nindex 81e9190..2015316 100644\n--- a/receive-pack.c\n+++ b/receive-pack.c\n@@ -247,8 +247,8 @@ static void read_head_info(void)\n \t\t\t\treport_status = 1;\n \t\t}\n \t\tcmd = xmalloc(sizeof(struct command) + len - 80);\n-\t\tmemcpy(cmd->old_sha1, old_sha1, 20);\n-\t\tmemcpy(cmd->new_sha1, new_sha1, 20);\n+\t\thashcpy(cmd->old_sha1, old_sha1);\n+\t\thashcpy(cmd->new_sha1, new_sha1);\n \t\tmemcpy(cmd->ref_name, line + 82, len - 81);\n \t\tcmd->error_string = \"n/a (unpacker error)\";\n \t\tcmd->next = NULL;\ndiff --git a/revision.c b/revision.c\nindex 5a91d06..1d89d72 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -496,7 +496,7 @@ static int add_parents_only(struct rev_i\n \t\tit = get_reference(revs, arg, sha1, 0);\n \t\tif (it->type != OBJ_TAG)\n \t\t\tbreak;\n-\t\tmemcpy(sha1, ((struct tag*)it)->tagged->sha1, 20);\n+\t\thashcpy(sha1, ((struct tag*)it)->tagged->sha1);\n \t}\n \tif (it->type != OBJ_COMMIT)\n \t\treturn 0;\ndiff --git a/send-pack.c b/send-pack.c\nindex f7c0cfc..fd79a61 100644\n--- a/send-pack.c\n+++ b/send-pack.c\n@@ -185,7 +185,7 @@ static int one_local_ref(const char *ref\n \tstruct ref *ref;\n \tint len = strlen(refname) + 1;\n \tref = xcalloc(1, sizeof(*ref) + len);\n-\tmemcpy(ref->new_sha1, sha1, 20);\n+\thashcpy(ref->new_sha1, sha1);\n \tmemcpy(ref->name, refname, len);\n \t*local_tail = ref;\n \tlocal_tail = &ref->next;\n@@ -310,7 +310,7 @@ static int send_pack(int in, int out, in\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t}\n-\t\tmemcpy(ref->new_sha1, ref->peer_ref->new_sha1, 20);\n+\t\thashcpy(ref->new_sha1, ref->peer_ref->new_sha1);\n \t\tif (is_zero_sha1(ref->new_sha1)) {\n \t\t\terror(\"cannot happen anymore\");\n \t\t\tret = -3;\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 4b1be72..789deb7 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -558,7 +558,7 @@ struct packed_git *add_packed_git(char *\n \tp->pack_use_cnt = 0;\n \tp->pack_local = local;\n \tif ((path_len > 44) && !get_sha1_hex(path + path_len - 44, sha1))\n-\t\tmemcpy(p->sha1, sha1, 20);\n+\t\thashcpy(p->sha1, sha1);\n \treturn p;\n }\n \n@@ -589,7 +589,7 @@ struct packed_git *parse_pack_index_file\n \tp->pack_base = NULL;\n \tp->pack_last_used = 0;\n \tp->pack_use_cnt = 0;\n-\tmemcpy(p->sha1, sha1, 20);\n+\thashcpy(p->sha1, sha1);\n \treturn p;\n }\n \n@@ -971,7 +971,7 @@ int check_reuse_pack_delta(struct packed\n \tptr = unpack_object_header(p, ptr, kindp, sizep);\n \tif (*kindp != OBJ_DELTA)\n \t\tgoto done;\n-\tmemcpy(base, (unsigned char *) p->pack_base + ptr, 20);\n+\thashcpy(base, (unsigned char *) p->pack_base + ptr);\n \tstatus = 0;\n  done:\n \tunuse_packed_git(p);\n@@ -999,7 +999,7 @@ void packed_object_info_detail(struct pa\n \t\tif (p->pack_size <= offset + 20)\n \t\t\tdie(\"pack file %s records an incomplete delta base\",\n \t\t\t    p->pack_name);\n-\t\tmemcpy(base_sha1, pack, 20);\n+\t\thashcpy(base_sha1, pack);\n \t\tdo {\n \t\t\tstruct pack_entry base_ent;\n \t\t\tunsigned long junk;\n@@ -1219,7 +1219,7 @@ int nth_packed_object_sha1(const struct \n \tvoid *index = p->index_base + 256;\n \tif (n < 0 || num_packed_objects(p) <= n)\n \t\treturn -1;\n-\tmemcpy(sha1, (char *) index + (24 * n) + 4, 20);\n+\thashcpy(sha1, (unsigned char *) index + (24 * n) + 4);\n \treturn 0;\n }\n \n@@ -1236,7 +1236,7 @@ int find_pack_entry_one(const unsigned c\n \t\tint cmp = hashcmp((unsigned char *)index + (24 * mi) + 4, sha1);\n \t\tif (!cmp) {\n \t\t\te->offset = ntohl(*((unsigned int *) ((char *) index + (24 * mi))));\n-\t\t\tmemcpy(e->sha1, sha1, 20);\n+\t\t\thashcpy(e->sha1, sha1);\n \t\t\te->p = p;\n \t\t\treturn 1;\n \t\t}\n@@ -1349,7 +1349,7 @@ void *read_object_with_reference(const u\n \tunsigned long isize;\n \tunsigned char actual_sha1[20];\n \n-\tmemcpy(actual_sha1, sha1, 20);\n+\thashcpy(actual_sha1, sha1);\n \twhile (1) {\n \t\tint ref_length = -1;\n \t\tconst char *ref_type = NULL;\n@@ -1360,7 +1360,7 @@ void *read_object_with_reference(const u\n \t\tif (!strcmp(type, required_type)) {\n \t\t\t*size = isize;\n \t\t\tif (actual_sha1_return)\n-\t\t\t\tmemcpy(actual_sha1_return, actual_sha1, 20);\n+\t\t\t\thashcpy(actual_sha1_return, actual_sha1);\n \t\t\treturn buffer;\n \t\t}\n \t\t/* Handle references */\n@@ -1555,7 +1555,7 @@ int write_sha1_file(void *buf, unsigned \n \t */\n \tfilename = write_sha1_file_prepare(buf, len, type, sha1, hdr, &hdrlen);\n \tif (returnsha1)\n-\t\tmemcpy(returnsha1, sha1, 20);\n+\t\thashcpy(returnsha1, sha1);\n \tif (has_sha1_file(sha1))\n \t\treturn 0;\n \tfd = open(filename, O_RDONLY);\ndiff --git a/sha1_name.c b/sha1_name.c\nindex b6c198f..8a5809e 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -109,7 +109,7 @@ static int find_short_packed_object(int \n \t\t\t\t    !match_sha(len, match, next)) {\n \t\t\t\t\t/* unique within this pack */\n \t\t\t\t\tif (!found) {\n-\t\t\t\t\t\tmemcpy(found_sha1, now, 20);\n+\t\t\t\t\t\thashcpy(found_sha1, now);\n \t\t\t\t\t\tfound++;\n \t\t\t\t\t}\n \t\t\t\t\telse if (hashcmp(found_sha1, now)) {\n@@ -126,7 +126,7 @@ static int find_short_packed_object(int \n \t\t}\n \t}\n \tif (found == 1)\n-\t\tmemcpy(sha1, found_sha1, 20);\n+\t\thashcpy(sha1, found_sha1);\n \treturn found;\n }\n \n@@ -146,13 +146,13 @@ static int find_unique_short_object(int \n \tif (1 < has_unpacked || 1 < has_packed)\n \t\treturn SHORT_NAME_AMBIGUOUS;\n \tif (has_unpacked != has_packed) {\n-\t\tmemcpy(sha1, (has_packed ? packed_sha1 : unpacked_sha1), 20);\n+\t\thashcpy(sha1, (has_packed ? packed_sha1 : unpacked_sha1));\n \t\treturn 0;\n \t}\n \t/* Both have unique ones -- do they match? */\n \tif (hashcmp(packed_sha1, unpacked_sha1))\n \t\treturn SHORT_NAME_AMBIGUOUS;\n-\tmemcpy(sha1, packed_sha1, 20);\n+\thashcpy(sha1, packed_sha1);\n \treturn 0;\n }\n \n@@ -326,13 +326,13 @@ static int get_parent(const char *name, \n \tif (parse_commit(commit))\n \t\treturn -1;\n \tif (!idx) {\n-\t\tmemcpy(result, commit->object.sha1, 20);\n+\t\thashcpy(result, commit->object.sha1);\n \t\treturn 0;\n \t}\n \tp = commit->parents;\n \twhile (p) {\n \t\tif (!--idx) {\n-\t\t\tmemcpy(result, p->item->object.sha1, 20);\n+\t\t\thashcpy(result, p->item->object.sha1);\n \t\t\treturn 0;\n \t\t}\n \t\tp = p->next;\n@@ -353,9 +353,9 @@ static int get_nth_ancestor(const char *\n \n \t\tif (!commit || parse_commit(commit) || !commit->parents)\n \t\t\treturn -1;\n-\t\tmemcpy(sha1, commit->parents->item->object.sha1, 20);\n+\t\thashcpy(sha1, commit->parents->item->object.sha1);\n \t}\n-\tmemcpy(result, sha1, 20);\n+\thashcpy(result, sha1);\n \treturn 0;\n }\n \n@@ -407,7 +407,7 @@ static int peel_onion(const char *name, \n \t\to = deref_tag(o, name, sp - name - 2);\n \t\tif (!o || (!o->parsed && !parse_object(o->sha1)))\n \t\t\treturn -1;\n-\t\tmemcpy(sha1, o->sha1, 20);\n+\t\thashcpy(sha1, o->sha1);\n \t}\n \telse {\n \t\t/* At this point, the syntax look correct, so\n@@ -419,7 +419,7 @@ static int peel_onion(const char *name, \n \t\t\tif (!o || (!o->parsed && !parse_object(o->sha1)))\n \t\t\t\treturn -1;\n \t\t\tif (o->type == expected_type) {\n-\t\t\t\tmemcpy(sha1, o->sha1, 20);\n+\t\t\t\thashcpy(sha1, o->sha1);\n \t\t\t\treturn 0;\n \t\t\t}\n \t\t\tif (o->type == OBJ_TAG)\n@@ -526,7 +526,7 @@ int get_sha1(const char *name, unsigned \n \t\t\t    memcmp(ce->name, cp, namelen))\n \t\t\t\tbreak;\n \t\t\tif (ce_stage(ce) == stage) {\n-\t\t\t\tmemcpy(sha1, ce->sha1, 20);\n+\t\t\t\thashcpy(sha1, ce->sha1);\n \t\t\t\treturn 0;\n \t\t\t}\n \t\t\tpos++;\ndiff --git a/tree-walk.c b/tree-walk.c\nindex 3f83e98..14cc5ae 100644\n--- a/tree-walk.c\n+++ b/tree-walk.c\n@@ -179,7 +179,7 @@ static int find_tree_entry(struct tree_d\n \t\tif (cmp < 0)\n \t\t\tbreak;\n \t\tif (entrylen == namelen) {\n-\t\t\tmemcpy(result, sha1, 20);\n+\t\t\thashcpy(result, sha1);\n \t\t\treturn 0;\n \t\t}\n \t\tif (name[entrylen] != '/')\n@@ -187,7 +187,7 @@ static int find_tree_entry(struct tree_d\n \t\tif (!S_ISDIR(*mode))\n \t\t\tbreak;\n \t\tif (++entrylen == namelen) {\n-\t\t\tmemcpy(result, sha1, 20);\n+\t\t\thashcpy(result, sha1);\n \t\t\treturn 0;\n \t\t}\n \t\treturn get_tree_entry(sha1, name + entrylen, result, mode);\ndiff --git a/tree.c b/tree.c\nindex ef456be..ea386e5 100644\n--- a/tree.c\n+++ b/tree.c\n@@ -25,7 +25,7 @@ static int read_one_entry(const unsigned\n \tce->ce_flags = create_ce_flags(baselen + len, stage);\n \tmemcpy(ce->name, base, baselen);\n \tmemcpy(ce->name + baselen, pathname, len+1);\n-\tmemcpy(ce->sha1, sha1, 20);\n+\thashcpy(ce->sha1, sha1);\n \treturn add_cache_entry(ce, ADD_CACHE_OK_TO_ADD|ADD_CACHE_SKIP_DFCHECK);\n }\n \ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 467d994..3ac0289 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -200,7 +200,7 @@ #endif\n \n \t\t\tany_files = 1;\n \n-\t\t\tmemcpy(ce->sha1, posns[i]->sha1, 20);\n+\t\t\thashcpy(ce->sha1, posns[i]->sha1);\n \t\t\tsrc[i + o->merge] = ce;\n \t\t\tsubposns[i] = df_conflict_list;\n \t\t\tposns[i] = posns[i]->next;\n-- \n1.4.2.gfec68\n"},{"id":"25834","messageId":"7vac5ubn57.fsf@assigned-by-dhcp.cox.net","threadId":"5317","inReplyTo":"20060823064900.GA26340@spearce.org","subject":"Re: [RFC] adding support for md5","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-08-24T07:36:52Z","receivedAt":"2006-08-24T07:36:52Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n>> > What about all of those memcpy(a, b, 20)'s?  :-)\n>> \n>> Surely.  If you are inclined to, go wild.\n>\n> Like this?  :-)\n\nExcept some minor nits, yes.\n\n * I would have preferred two patches, one for \"master\" and one\n   for the C merge-recursive topic (or at least \"next\").\n\n * You missed a few in \"master\".\n\n * The cast in the second hunk in combine-diff.c was wrong;\n   breakage was caught by our testsuite.\n\nI've pushed out a fixed up result in \"master\" and \"next\".\n"},{"id":"25836","messageId":"20060824080807.GG25247@spearce.org","threadId":"5317","inReplyTo":"7vac5ubn57.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] adding support for md5","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-08-24T08:08:07Z","receivedAt":"2006-08-24T08:08:07Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <junkio@cox.net> wrote:\n> Shawn Pearce <spearce@spearce.org> writes:\n> \n> >> > What about all of those memcpy(a, b, 20)'s?  :-)\n> >> \n> >> Surely.  If you are inclined to, go wild.\n> >\n> > Like this?  :-)\n> \n> Except some minor nits, yes.\n> \n>  * I would have preferred two patches, one for \"master\" and one\n>    for the C merge-recursive topic (or at least \"next\").\n\nDoh.  I didn't realize this was something you were interested in\npulling into master.  Otherwise I would have done this.  Next time\nI'll try to keep that in mind.\n \n>  * You missed a few in \"master\".\n\nNot surprising since I typically work against and use next.\n \n>  * The cast in the second hunk in combine-diff.c was wrong;\n>    breakage was caught by our testsuite.\n\nOK, that's just flat out stupid of me.  I apologize for making you\nfix my mistakes.  :-)\n\n> I've pushed out a fixed up result in \"master\" and \"next\".\n\nThanks.\n\n-- \nShawn.\n"},{"id":"25843","messageId":"7vmz9ua0cx.fsf@assigned-by-dhcp.cox.net","threadId":"5317","inReplyTo":"20060824080807.GG25247@spearce.org","subject":"Re: [RFC] adding support for md5","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-08-24T10:34:22Z","receivedAt":"2006-08-24T10:34:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n>> Except some minor nits, yes.\n>> \n>>  * I would have preferred two patches, one for \"master\" and one\n>>    for the C merge-recursive topic (or at least \"next\").\n>\n> Doh.  I didn't realize this was something you were interested in\n> pulling into master.\n\nApplying them directly to \"master\" does not have much to do with\nthis.\n\nOf course, if something is obviously the right thing to do, then\nI often apply them straight to \"master\" (or apply them to\n\"maint\" and pull the result into \"master\").  In other cases, I\nprefer to fork off a new series from the tip of the \"master\"\ninto a new topic branch.  Then merge that into either \"pu\" (if I\nhave suspicion that it is not ready for even testing yet) or\n\"next\", and cook there for a while until it is ready.\n\nWhile being cooked in \"next\", what happens is that changes to\nthat specific topic are applied to the tip of the topic branch,\nand then pulled into \"next\", over and over.  Many topic branches\nare cooked simultaneously that way.  So the development history\nof \"next\" is, eh, messy.\n\nI usually test \"next\", in other words, multiple topics cooking\ntogether.  But when some changes are applied to \"master\" that\nmight interfere with an older but still not in \"master\", I pull\n\"master\" into the topic and test that topic alone in isolation.\n\nThat way, when a topic matures, we can be reasonably sure that\nit can be pulled into \"master\" without breaking things.\n\nA single patch on top of \"next\" depends on all existing topics\nthat may or may not turn out to be useful.  That makes such a\npatch less useful than otherwise be.\n\nSo if a series affects things in \"master\" and some other things\nstill not in \"master\", a preferred way, from my workflow point\nof view, is to have at least two patches: one for \"master\" and\nanother for the rest.  Then what I would do is to fork one topic\noff from \"master\" and apply the former, pull that into \"next\"\nand cook that.  That part of the topic can graduate to \"master\"\nwithout waiting for other topics in \"next\".\n\nWhat happens to the rest is a bit more involved.  In the case of\nthe hashcpy() patch, one thing that only exists in \"next\" was\nmerge-recursive.c, but the story would be the same if \"next\" had\nmore places that used memcpy(a, b, 20) than \"master\" in a file\nthat are common in two branches.  Ideally, the remainder will be\nbroken into pieces and applied as a fixup on top of existing\ntopics (in this case, C merge-recursive topic) and then merged\ninto \"next\".  This can be either done by applying the remainder\ndirectly on top of the affected topic, or forking a subtopic off\nof the affected topic (the latter is useful if the new series\nmight turn out to be dud -- the original topic will not be\ncontaminated by bad changes and can graduate to \"master\" more\neasily).\n\nIn any case, I've done a split myself and the parts that can be\napplied to \"master\" is now sitting in gl/cleanup topic and the\nremainder is sitting in gl/cleanup-next topic which was forked\noff from C merge-recursive topic (you can tell where their tips\nare by looking at \"gitk next\" output) and both are merged into\n\"next\".\n"}]}