Re: Object hash (was: Re: [ANNOUNCE] git-rev-size: calculate sizes of repository)
- From
- Josef Weidendorfer <josef.weidendorfer@gmx.de>
- Date
- Aug 20, 2006, 18:41 UTC
- Message-ID
- <200608202041.19644.Josef.Weidendorfer@gmx.de>
- In-Reply-To
- <Pine.LNX.4.63.0608201846110.28360@wbgn013.biozentrum.uni-wuerzburg.de>
On Sunday 20 August 2006 18:51, Johannes Schindelin wrote:
Show 10 quoted lines
> > > +static unsigned int hash_index(struct hash_map *hash, const char *sha1)
> > > +{
> > > + unsigned int index = *(unsigned int *)sha1;
> >
> > If you have the same SHA1, stored at different addresses, you get different
> > indexes for the same SHA1. Index probably should be calculated from the
> > SHA1 string.
>
> Actually, it does! "*(unsigned int *)sha1" means that the first 4 bytes
> of the sha1 are interpreted as a number.Ah, yes. That's fine.
Show 10 quoted lines
> > > +void hash_put(struct hash_map *hash, struct object *obj)
> > > +{
> > > + if (++hash->nr > hash->alloc / 2)
> > > + grow_hash(hash);
> >
> > If you insert the same object multiple times, hash->nr will get too big.
>
> First, you cannot put the same object multiple times. That is not what a
> hash does (at least in this case): it stores unique objects (identified by
> their sha1 in this case).I put it the wrong way; I should have said "if you call hash_put() multiple times with the same object". You get the same index, and nothing should change. However, you still increment hash->nr, but this error is not really important as you correct it in grow_hash().
So... sorry for the noise ;-)
Josef