From: Jeff King Date: Thu, 24 Sep 2020 07:31:25 GMT Subject: Re: [PATCH 06/13] reftable: (de)serialization for the polymorphic record type. Message-ID: <20200924073125.GD1851751@coredump.intra.peff.net> In-Reply-To: <20200924072151.GC1851751@coredump.intra.peff.net> On Thu, Sep 24, 2020 at 03:21:51AM -0400, Jeff King wrote: > > I originally had > > > > +void put_be64(uint8_t *out, uint64_t v) > > +{ > > + int i = sizeof(uint64_t); > > + while (i--) { > > + out[i] = (uint8_t)(v & 0xff); > > + v >>= 8; > > + } > > +} > > > > in my reftable library, which is portable. Is there a reason for the > > magic with htonll and friends? > > Presumably it was thought to be faster. This comes originally from the > block-sha1 code in 660231aa97 (block-sha1: support for architectures > with memory alignment restrictions, 2009-08-12). I don't know how it > compares in practice, and especially these days. > > Our fallback routines are similar to an unrolled version of what you > wrote above. We should be able to measure it pretty easily, since block-sha1 uses a lot of get_be32/put_be32. I generated a 4GB random file, built with BLK_SHA1=Yes and -O2, and timed: t/helper/test-tool sha1