From: Jeff Garzik Date: Sun, 24 Apr 2005 00:35:57 GMT Subject: Re: Hash collision count Message-ID: <426AE9ED.4060005@pobox.com> In-Reply-To: <20050423234637.GS13222@pasky.ji.cz> Petr Baudis wrote: > Dear diary, on Sun, Apr 24, 2005 at 01:20:21AM CEST, I got a letter > where Jeff Garzik told me that... > >>Second, in your scenario, it's highly unlikely you would get 4 billion >>sha1 hash collisions, even if you had the disk space to store such a git >>database. > > > It's highly unlikely you would get a _single_ collision. Agreed. >>First, the hash is NOT unique. >> >>Second, you lose data if you pretend it is unique. I don't like losing >>data. > > > *sigh* > > We've been through this before, haven't we? In messing around with archive servers, people get nervous using (hash,value) based storage if there isn't even a simple test for collisions. Someone just told me that one implementation of the Venti archive server[1] simply fails the write, if a data item exists with a duplicate hash value. As long as git fails or does something -predictable- in the face of the hash collision, I'm satisfied. Jeff [1] http://www.cs.bell-labs.com/sys/doc/venti/venti.html