{"thread":{"id":"270","subject":"Hash collision count","startedAt":"2005-04-23T20:27:47Z","lastAt":"2005-04-26T00:00:17Z","messageCount":14,"participants":["Jeff Garzik","Ray Heasman","Petr Baudis","David Lang","Imre Simon","Jon Seymour","Tom Lord"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"1412","messageId":"426AAFC3.800@pobox.com","threadId":"270","inReplyTo":null,"subject":"Hash collision count","fromName":"Jeff Garzik","fromEmail":"jgarzik@pobox.com","sentAt":"2005-04-23T20:27:47Z","receivedAt":"2005-04-23T20:27:47Z","isPatch":false,"sender":{"key":"jgarzik@pobox.com","avatar":null},"body":"\nIdeally a hash + collision-count pair would make the best key, rather \nthan just hash alone.\n\nA collision -will- occur eventually, and it is trivial to avoid this \nproblem:\n\n\t$n = 0\n\tattempt to store as $hash-$n\n\tif $hash-$n exists (unlikely)\n\t\t$n++\n\t\tgoto restart\n\tkey = $hash-$n\n\nTangent-as-the-reason-I-bring-this-up:\n\nOne of my long-term projects is an archive service, somewhat like \nPlan9's venti:  a multi-server key-value database, with sha1 hash as the \nkey.\n\nHowever, as the database grows into the terabyte (possibly petabyte) \nrange, the likelihood of a collision transitions rapidly from unlikely \n-> possible -> likely.\n\nSince it is -so- simple to guarantee that you avoid collisions, I'm \nhoping git will do so before the key structure is too ingrained.\n\n\tJeff\n\n\n\n"},{"id":"1413","messageId":"426AB10F.9050300@pobox.com","threadId":"270","inReplyTo":"426AAFC3.800@pobox.com","subject":"Re: Hash collision count","fromName":"Jeff Garzik","fromEmail":"jgarzik@pobox.com","sentAt":"2005-04-23T20:33:19Z","receivedAt":"2005-04-23T20:33:19Z","isPatch":false,"sender":{"key":"jgarzik@pobox.com","avatar":null},"body":"Jeff Garzik wrote:\n> \n> Ideally a hash + collision-count pair would make the best key, rather \n> than just hash alone.\n> \n> A collision -will- occur eventually, and it is trivial to avoid this \n> problem:\n> \n>     $n = 0\n>     attempt to store as $hash-$n\n>     if $hash-$n exists (unlikely)\n>         $n++\n>         goto restart\n>     key = $hash-$n\n\n\nThis of course presumes that you know that your new data does not exist \nin the cache.\n\nOtherwise, you would need to hash the on-disk $hash-0 to determine if \nits a collision or a reference.\n\n\tJeff\n\n\n"},{"id":"1423","messageId":"1114297231.10264.12.camel@maze.mythral.org","threadId":"270","inReplyTo":"426AAFC3.800@pobox.com","subject":"Re: Hash collision count","fromName":"Ray Heasman","fromEmail":"lists@mythral.org","sentAt":"2005-04-23T23:00:31Z","receivedAt":"2005-04-23T23:00:31Z","isPatch":false,"sender":{"key":"lists@mythral.org","avatar":null},"body":"On Sat, 2005-04-23 at 16:27 -0400, Jeff Garzik wrote:\n> Ideally a hash + collision-count pair would make the best key, rather \n> than just hash alone.\n> \n> A collision -will- occur eventually, and it is trivial to avoid this \n> problem:\n> \n> \t$n = 0\n> \tattempt to store as $hash-$n\n> \tif $hash-$n exists (unlikely)\n> \t\t$n++\n> \t\tgoto restart\n> \tkey = $hash-$n\n> \n\nGreat. So what have you done here? Suppose you have 32 bits of counter\nfor n. Whoopee, you just added 32 bits to your hash, using a two stage\nalgorithm. So, you have a 192 bit hash assuming you started with the 160\nbit SHA. And, one day your 32 bit counter won't be enough. Then what?\n\n> Tangent-as-the-reason-I-bring-this-up:\n> \n> One of my long-term projects is an archive service, somewhat like \n> Plan9's venti:  a multi-server key-value database, with sha1 hash as the \n> key.\n> \n> However, as the database grows into the terabyte (possibly petabyte) \n> range, the likelihood of a collision transitions rapidly from unlikely \n> -> possible -> likely.\n> \n> Since it is -so- simple to guarantee that you avoid collisions, I'm \n> hoping git will do so before the key structure is too ingrained.\n\nYou aren't solving anything. You're just putting it off, and doing it in\na way that breaks all the wonderful semantics possible by just assuming\nthat the hash is unique. All of a sudden we are doing checks of data\nthat we never did before, and we have to do the check trillions of times\nbefore the CPU time spent pays off.\n\nIf you want to use a bigger hash then use a bigger hash, but don't fool\nyourself into thinking that isn't what you are doing.\n\nCiao,\nRay\n\n\n\n"},{"id":"1430","messageId":"426AD835.5070404@pobox.com","threadId":"270","inReplyTo":"1114297231.10264.12.camel@maze.mythral.org","subject":"Re: Hash collision count","fromName":"Jeff Garzik","fromEmail":"jgarzik@pobox.com","sentAt":"2005-04-23T23:20:21Z","receivedAt":"2005-04-23T23:20:21Z","isPatch":false,"sender":{"key":"jgarzik@pobox.com","avatar":null},"body":"Ray Heasman wrote:\n> On Sat, 2005-04-23 at 16:27 -0400, Jeff Garzik wrote:\n> \n>>Ideally a hash + collision-count pair would make the best key, rather \n>>than just hash alone.\n>>\n>>A collision -will- occur eventually, and it is trivial to avoid this \n>>problem:\n>>\n>>\t$n = 0\n>>\tattempt to store as $hash-$n\n>>\tif $hash-$n exists (unlikely)\n>>\t\t$n++\n>>\t\tgoto restart\n>>\tkey = $hash-$n\n>>\n> \n> \n> Great. So what have you done here? Suppose you have 32 bits of counter\n> for n. Whoopee, you just added 32 bits to your hash, using a two stage\n> algorithm. So, you have a 192 bit hash assuming you started with the 160\n> bit SHA. And, one day your 32 bit counter won't be enough. Then what?\n\nFirst, there is no 32-bit limit.  git stores keys (aka hashes) as \nstrings.  As it should.\n\nSecond, in your scenario, it's highly unlikely you would get 4 billion \nsha1 hash collisions, even if you had the disk space to store such a git \ndatabase.\n\n\n>>Tangent-as-the-reason-I-bring-this-up:\n>>\n>>One of my long-term projects is an archive service, somewhat like \n>>Plan9's venti:  a multi-server key-value database, with sha1 hash as the \n>>key.\n>>\n>>However, as the database grows into the terabyte (possibly petabyte) \n>>range, the likelihood of a collision transitions rapidly from unlikely \n>>-> possible -> likely.\n>>\n>>Since it is -so- simple to guarantee that you avoid collisions, I'm \n>>hoping git will do so before the key structure is too ingrained.\n> \n> \n> You aren't solving anything. You're just putting it off, and doing it in\n> a way that breaks all the wonderful semantics possible by just assuming\n> that the hash is unique. All of a sudden we are doing checks of data\n> that we never did before, and we have to do the check trillions of times\n> before the CPU time spent pays off.\n\nFirst, the hash is NOT unique.\n\nSecond, you lose data if you pretend it is unique.  I don't like losing \ndata.\n\nThird, a data check only occurs in the highly unlikely case that a hash \nalready exists -- a collision.  Rather than \"trillions of times\", more \nlike \"one in a trillion chance.\"\n\n\tJeff\n\n\n\n"},{"id":"1437","messageId":"20050423234637.GS13222@pasky.ji.cz","threadId":"270","inReplyTo":"426AD835.5070404@pobox.com","subject":"Re: Hash collision count","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-23T23:46:37Z","receivedAt":"2005-04-23T23:46:37Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sun, Apr 24, 2005 at 01:20:21AM CEST, I got a letter\nwhere Jeff Garzik <jgarzik@pobox.com> told me that...\n> Second, in your scenario, it's highly unlikely you would get 4 billion \n> sha1 hash collisions, even if you had the disk space to store such a git \n> database.\n\nIt's highly unlikely you would get a _single_ collision.\n\n> First, the hash is NOT unique.\n> \n> Second, you lose data if you pretend it is unique.  I don't like losing \n> data.\n\n*sigh*\n\nWe've been through this before, haven't we?\n\n> Third, a data check only occurs in the highly unlikely case that a hash \n> already exists -- a collision.  Rather than \"trillions of times\", more \n> like \"one in a trillion chance.\"\n\nNo, a collision is pretty common thing, actually. It's the main power of\ngit, actually - when you do read-tree, modify it and do write-tree\n(typically when doing commit), everything you didn't modify (99% of\nstuff, most likely) is basically a collision - but it's ok since it\njust stays the same.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"1449","messageId":"426AE9ED.4060005@pobox.com","threadId":"270","inReplyTo":"20050423234637.GS13222@pasky.ji.cz","subject":"Re: Hash collision count","fromName":"Jeff Garzik","fromEmail":"jgarzik@pobox.com","sentAt":"2005-04-24T00:35:57Z","receivedAt":"2005-04-24T00:35:57Z","isPatch":false,"sender":{"key":"jgarzik@pobox.com","avatar":null},"body":"Petr Baudis wrote:\n> Dear diary, on Sun, Apr 24, 2005 at 01:20:21AM CEST, I got a letter\n> where Jeff Garzik <jgarzik@pobox.com> told me that...\n> \n>>Second, in your scenario, it's highly unlikely you would get 4 billion \n>>sha1 hash collisions, even if you had the disk space to store such a git \n>>database.\n> \n> \n> It's highly unlikely you would get a _single_ collision.\n\nAgreed.\n\n\n>>First, the hash is NOT unique.\n>>\n>>Second, you lose data if you pretend it is unique.  I don't like losing \n>>data.\n> \n> \n> *sigh*\n> \n> We've been through this before, haven't we?\n\n<shrug>\n\nIn messing around with archive servers, people get nervous using \n(hash,value) based storage if there isn't even a simple test for collisions.\n\nSomeone just told me that one implementation of the Venti archive \nserver[1] simply fails the write, if a data item exists with a duplicate \nhash value.  As long as git fails or does something -predictable- in the \nface of the hash collision, I'm satisfied.\n\n\tJeff\n\n\n[1] http://www.cs.bell-labs.com/sys/doc/venti/venti.html\n"},{"id":"1451","messageId":"20050424004039.GU13222@pasky.ji.cz","threadId":"270","inReplyTo":"426AE9ED.4060005@pobox.com","subject":"Re: Hash collision count","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-24T00:40:39Z","receivedAt":"2005-04-24T00:40:39Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sun, Apr 24, 2005 at 02:35:57AM CEST, I got a letter\nwhere Jeff Garzik <jgarzik@pobox.com> told me that...\n> Someone just told me that one implementation of the Venti archive \n> server[1] simply fails the write, if a data item exists with a duplicate \n> hash value.  As long as git fails or does something -predictable- in the \n> face of the hash collision, I'm satisfied.\n\n-DCOLLISION_CHECK\n\nSee the top of Makefile.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"1452","messageId":"426AEBAE.1060402@pobox.com","threadId":"270","inReplyTo":"20050424004039.GU13222@pasky.ji.cz","subject":"Re: Hash collision count","fromName":"Jeff Garzik","fromEmail":"jgarzik@pobox.com","sentAt":"2005-04-24T00:43:26Z","receivedAt":"2005-04-24T00:43:26Z","isPatch":false,"sender":{"key":"jgarzik@pobox.com","avatar":null},"body":"Petr Baudis wrote:\n> -DCOLLISION_CHECK\n\nCool.  I am happy, then :)\n\nMake sure that's enabled by default...\n\nThanks,\n\n\tJeff\n\n\n"},{"id":"1455","messageId":"1114304468.10043.45.camel@maze.mythral.org","threadId":"270","inReplyTo":"426AD835.5070404@pobox.com","subject":"Re: Hash collision count","fromName":"Ray Heasman","fromEmail":"lists@mythral.org","sentAt":"2005-04-24T01:01:08Z","receivedAt":"2005-04-24T01:01:08Z","isPatch":false,"sender":{"key":"lists@mythral.org","avatar":null},"body":"On Sat, 2005-04-23 at 19:20 -0400, Jeff Garzik wrote:\n> Ray Heasman wrote:\n> > On Sat, 2005-04-23 at 16:27 -0400, Jeff Garzik wrote:\n> > \n> >>Ideally a hash + collision-count pair would make the best key, rather \n> >>than just hash alone.\n> >>\n> >>A collision -will- occur eventually, and it is trivial to avoid this \n> >>problem:\n> >>\n> >>\t$n = 0\n> >>\tattempt to store as $hash-$n\n> >>\tif $hash-$n exists (unlikely)\n> >>\t\t$n++\n> >>\t\tgoto restart\n> >>\tkey = $hash-$n\n> >>\n> > \n> > \n> > Great. So what have you done here? Suppose you have 32 bits of counter\n> > for n. Whoopee, you just added 32 bits to your hash, using a two stage\n> > algorithm. So, you have a 192 bit hash assuming you started with the 160\n> > bit SHA. And, one day your 32 bit counter won't be enough. Then what?\n> \n> First, there is no 32-bit limit.  git stores keys (aka hashes) as \n> strings.  As it should.\n\nOh great, now we have variable length id strings too. And we'll have to\npretend the OS can store infinite length file names.\n\n> Second, in your scenario, it's highly unlikely you would get 4 billion \n> sha1 hash collisions, even if you had the disk space to store such a git \n> database.\n\nEr. So your so-unlikely-the-sun-will-burn-out-first scenario beats my\nso-unlikely-the-sun-will-burn-out-first scenario? Why am I not worried?\n\n> > You aren't solving anything. You're just putting it off, and doing it in\n> > a way that breaks all the wonderful semantics possible by just assuming\n> > that the hash is unique. All of a sudden we are doing checks of data\n> > that we never did before, and we have to do the check trillions of times\n> > before the CPU time spent pays off.\n> \n> First, the hash is NOT unique.\n\nNooooo. Really?\n\nWhy not just use a 8192 bit hash for each 1KiB of data? We could store a\nzero length file and store all the data in the filename. Guaranteed, no\nhash collisions that way.\n\nWe make an assumption that we know is right most of the time, and we use\nit because we know our computer will crash from random quantum\nfluctuations before we have a chance of bumping into the problem. You do\nknow that metastability means that every logic gate in your computer\nhardware is guaranteed to fail every \"x\" operations, where x is defined\nby process size, voltage and temperature? Sure, any failures in git\nwould be data dependent rather than random, but that just means we don't\nget to store carefully crafted blocks invented by hypothetical\ncryptographers that have completely broken SHA.\n\n> Second, you lose data if you pretend it is unique.  I don't like losing \n> data.\n\nYou lose data either way. Just we get to burn out a few extra suns\nbefore yours dies, and I can burn out whole galaxies before mine dies by\nusing a 256 bit hash, anyway.\n\n> Third, a data check only occurs in the highly unlikely case that a hash \n> already exists -- a collision.  Rather than \"trillions of times\", more \n> like \"one in a trillion chance.\"\n\nHeh. I calculate it has a 50% probability of happening after you have\nseen 10^24 input blocks. So, you are off by a factor of a trillion or\nso.\n\nAssuming we store 1 KiB blocks with a 160-bit hash, we would be able to\nstore 1000 Trillion Terabytes before the chance of hitting a collision\ngoes to 50%. To use marketing units, that is around 10 Trillion\nLibraries of Congress. Every 2 bits we add to the hash doubles the\namount of data we can store before we hit a 50% probability of\ncollision.\n\nI'm not sure how I could convince you that we're arguing about the\nnumber of angels that could dance on a pin.\n\nCiao,\nRay\n\n"},{"id":"1507","messageId":"Pine.LNX.4.62.0504240053480.32437@qynat.qvtvafvgr.pbz","threadId":"270","inReplyTo":"426AAFC3.800@pobox.com","subject":"Re: Hash collision count","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2005-04-24T07:56:24Z","receivedAt":"2005-04-24T07:56:24Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Sat, 23 Apr 2005, Jeff Garzik wrote:\n\n> Ideally a hash + collision-count pair would make the best key, rather than \n> just hash alone.\n>\n> A collision -will- occur eventually, and it is trivial to avoid this problem:\n>\n> \t$n = 0\n> \tattempt to store as $hash-$n\n> \tif $hash-$n exists (unlikely)\n> \t\t$n++\n> \t\tgoto restart\n> \tkey = $hash-$n\n>\n> Tangent-as-the-reason-I-bring-this-up:\n>\n> One of my long-term projects is an archive service, somewhat like Plan9's \n> venti:  a multi-server key-value database, with sha1 hash as the key.\n>\n> However, as the database grows into the terabyte (possibly petabyte) range, \n> the likelihood of a collision transitions rapidly from unlikely -> possible \n> -> likely.\n>\n> Since it is -so- simple to guarantee that you avoid collisions, I'm hoping \n> git will do so before the key structure is too ingrained.\n\nJeff, this can't work becouse you don't know what objects exist on other \nservers, in fact given the number of different repositories that will \neventually exist the odds are good that when the colision occures it will \nbe when object repositories get combined,\n\nDavid Lang\n-- \nThere are two ways of constructing a software design. One way is to make it so simple that there are obviously no deficiencies. And the other way is to make it so complicated that there are no obvious deficiencies.\n  -- C.A.R. Hoare\n"},{"id":"1549","messageId":"68ff9fa6050424142416fbadcd@mail.gmail.com","threadId":"270","inReplyTo":"426AEBAE.1060402@pobox.com","subject":"Re: Hash collision count","fromName":"Imre Simon","fromEmail":"imres.g@gmail.com","sentAt":"2005-04-24T21:24:14Z","receivedAt":"2005-04-24T21:24:14Z","isPatch":false,"sender":{"key":"imres.g@gmail.com","avatar":null},"body":"On 4/23/05, Jeff Garzik <jgarzik@pobox.com> wrote:\n> Petr Baudis wrote:\n> > -DCOLLISION_CHECK\n> \n> Cool.  I am happy, then :)\n> \n> Make sure that's enabled by default...\n> \n> Thanks,\n> \n>         Jeff\n\nI would like to second the suggestion that this should be enabled by default.\n\nFirst, I do think it is highly unlikely that a collision will ever be\nfound. Actually because of this if one is found it would be an\nimportant byproduct of git and it would probably influence future\nsystem designs. Hence, I believe that it is worth loosing some\nefficiency in order to illuminate better this question of the\ncollision.\n\nI would like to add a reciepe by which one would surely produce a\ncollision with two files of the same length and quite similar to each\nother. I believe that there is a popular belief that this is pretty\nmuch impossible. Of course, we do not have time to execute this\nalgorithm, but I believe that it convinces everyone that similar files\n*do* exist (actually in abundance) with the same hash.\n\n1. Take your favorite text file, at least 160 characters long. \n2. Choose 160 positions in this file.\n3. For each position choose your favorite mispelling of that character.\n4. Produce all 2^160 text files, all of the same length, choosing for\neach position either the original or the alternate character\n5. Add an arbitrary file of the same length, different from the above\n\nTwo of these files have the same sha1 hash. Or, for that matter, for\nany 160 bit  hash the same is true.\n\nCheers, Imre Simon\n"},{"id":"1554","messageId":"2cfc40320504241525c4153c2@mail.gmail.com","threadId":"270","inReplyTo":"68ff9fa6050424142416fbadcd@mail.gmail.com","subject":"Whales falling on houses - was: Hash collision count","fromName":"Jon Seymour","fromEmail":"jon.seymour@gmail.com","sentAt":"2005-04-24T22:25:32Z","receivedAt":"2005-04-24T22:25:32Z","isPatch":false,"sender":{"key":"jon.seymour@gmail.com","avatar":"https://avatars.githubusercontent.com/u/207131?v=4"},"body":".> \n> 1. Take your favorite text file, at least 160 characters long.\n> 2. Choose 160 positions in this file.\n> 3. For each position choose your favorite mispelling of that character.\n> 4. Produce all 2^160 text files, all of the same length, choosing for\n> each position either the original or the alternate character\n> 5. Add an arbitrary file of the same length, different from the above\n> \n> Two of these files have the same sha1 hash. Or, for that matter, for\n> any 160 bit  hash the same is true.\n\nIf you were to create those files at 10^9 files per second, it would\ntake you 10^38 years before you were in position to take step 5. I am\nabout to turn 38 this week. Would that I could live to 10^38.\n\nIt's absolute rubbish to say that the best solution from an\n<double-quote>engineering</double-quote> point of view is to eliminate\nthe infinitessimal possibility of a collision. Engineering is all\nabout assessing risk and making suitable trade-offs. Every day of the\nweek, \"real\" engineers accept life-threatening risks that put\nthousands of peoples lives in danger. They do it because we live in a\nworld where risk cannot be eliminated, merely reduced to an acceptable\nlevel.\n\nI can't understand that you are a prepared to drive a car or fly in a\nBoeing or Airbus that has a demonstrated risk of killing you, yet you\nwant to insist on eliminating a risk that at most might create an\ninteresting Slashdot headline: \"Jolt-crazed programmer finds SHA1\ncollision - but later dies when whale falls on house\".\n\njon.\n-- \nhomepage: http://www.zeta.org.au/~jon/\nblog: http://orwelliantremors.blogspot.com/\n"},{"id":"1677","messageId":"200504252350.QAA02241@emf.net","threadId":"270","inReplyTo":"20050423234637.GS13222@pasky.ji.cz","subject":"Re: Hash collision count","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-25T23:50:31Z","receivedAt":"2005-04-25T23:50:31Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n  From: Petr Baudis <pasky@ucw.cz>\n\n  Pasky:\n\n  > No, a collision is pretty common thing, actually. It's the main power of\n  > git, actually - when you do read-tree, modify it and do write-tree\n  > (typically when doing commit), everything you didn't modify (99% of\n  > stuff, most likely) is basically a collision - but it's ok since it\n  > just stays the same.\n\nThat is not the way people ordinarily use the word \"collision\".\nIt's pretty much the opposite of the normal way, actually.\n\n-t\n"},{"id":"1678","messageId":"20050426000017.GN13467@pasky.ji.cz","threadId":"270","inReplyTo":"200504252350.QAA02241@emf.net","subject":"Re: Hash collision count","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-26T00:00:17Z","receivedAt":"2005-04-26T00:00:17Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Tue, Apr 26, 2005 at 01:50:31AM CEST, I got a letter\nwhere Tom Lord <lord@emf.net> told me that...\n> \n>   From: Petr Baudis <pasky@ucw.cz>\n> \n>   Pasky:\n> \n>   > No, a collision is pretty common thing, actually. It's the main power of\n>   > git, actually - when you do read-tree, modify it and do write-tree\n>   > (typically when doing commit), everything you didn't modify (99% of\n>   > stuff, most likely) is basically a collision - but it's ok since it\n>   > just stays the same.\n> \n> That is not the way people ordinarily use the word \"collision\".\n> It's pretty much the opposite of the normal way, actually.\n\nYou need to quote me in the context of Jeff Garzik's\n\n> > Third, a data check only occurs in the highly unlikely case that a hash\n> > already exists -- a collision.  Rather than \"trillions of times\", more\n> > like \"one in a trillion chance.\"\n\nI just wanted to point out that the data check would hahve to occur\neverytime you didn't modify an object.\n\nKind regards,\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"}]}