{"thread":{"id":"15033","subject":"Casting and dereferencing of pointer","startedAt":"2008-08-16T09:33:30Z","lastAt":"2008-08-16T18:24:08Z","messageCount":2,"participants":["sed","Kalle Olavi Niemitalo"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"87370","messageId":"loom.20080816T093019-717@post.gmane.org","threadId":"15033","inReplyTo":null,"subject":"Casting and dereferencing of pointer","fromName":"sed","fromEmail":"sed.nivo@gmail.com","sentAt":"2008-08-16T09:33:30Z","receivedAt":"2008-08-16T09:33:30Z","isPatch":false,"sender":{"key":"sed.nivo@gmail.com","avatar":null},"body":"Maybe I should not post in this group but anyway...\n\nPlease look at my code that do the same, except of endianness.\nstatic void _parseItems(const unsigned char *pBuffer)\n{\n  unsigned int itemSize;\n  itemSize = *((unsigned int*)pBuffer); // this give system fault\n  memcpy(&itemSize, pBuffer, sizeof(unsigned int)); // this works well\n  .......\n}\n\nI'm not very experienced with C so I use git as example of good written code.\nIn object.c I've found two functions that looks like my one.\n\nstatic unsigned int hash_obj(struct object *obj, unsigned int n)\n{\n\tunsigned int hash = *(unsigned int *)obj->sha1;\n\treturn hash % n;\n}\n\nstatic int hashtable_index(const unsigned char *sha1)\n{\n\tunsigned int i;\n\tmemcpy(&i, sha1, sizeof(unsigned int));\n\treturn (int)(i % obj_hash_size);\n}\n\nI wonder why in the second used memcpy instead of:\nunsigned int i = *(unsigned int *)sha1\nMaybe there is explanation that will help to solve my problem.\nThank you\n"},{"id":"87384","messageId":"87y72w95tj.fsf@Astalo.kon.iki.fi","threadId":"15033","inReplyTo":"loom.20080816T093019-717@post.gmane.org","subject":"Re: Casting and dereferencing of pointer","fromName":"Kalle Olavi Niemitalo","fromEmail":"kon@iki.fi","sentAt":"2008-08-16T18:24:08Z","receivedAt":"2008-08-16T18:24:08Z","isPatch":false,"sender":{"key":"kon@iki.fi","avatar":null},"body":"sed <sed.nivo@gmail.com> writes:\n\n> static void _parseItems(const unsigned char *pBuffer)\n> {\n>   unsigned int itemSize;\n>   itemSize = *((unsigned int*)pBuffer); // this give system fault\n>   memcpy(&itemSize, pBuffer, sizeof(unsigned int)); // this works well\n>   .......\n> }\n\nOn some processors, you get an exception or just wrong results\nif you try to read a value from an address that is not properly\naligned for the type.  For example, a processor can require that\nthe address of any (signed or unsigned) int is divisible by 4.\nIntel 8086 compatibles have historically allowed any alignment,\nalthough even with them the code works faster if the alignment\nis right.  GCC has an __alignof__ operator with which you can\ncheck the alignment expected for a given type.\n\nSo, the *((unsigned int*)pBuffer) expression works portably only\nif the correct alignment is somehow ensured.  The memcpy() call\nworks regardless of the alignment of the pBuffer pointer.\n\n> static unsigned int hash_obj(struct object *obj, unsigned int n)\n> {\n> \tunsigned int hash = *(unsigned int *)obj->sha1;\n> \treturn hash % n;\n> }\n\nobject.h defines struct object as:\n\n| struct object {\n| \tunsigned parsed : 1;\n| \tunsigned used : 1;\n| \tunsigned type : TYPE_BITS;\n| \tunsigned flags : FLAG_BITS;\n| \tunsigned char sha1[20];\n| };\n\nThe structure has first an unsigned int that is split into\nbitfields, and then unsigned char sha1[20].  If the compiler\nworks sensibly, then __alignof__(struct object) is the same as\n__alignof__(unsigned int), and offsetof(struct object, sha1) is\nsizeof(unsigned int); so if obj points to a correctly aligned\nstruct object, then obj->sha1 should also be correctly aligned\nto be accessed as an unsigned int.\n\nIt seems somewhat brittle though: if someone adds a member\nbetween flags and sha1, then the alignment may become wrong.\nI think it would be good to add a comment in struct object\nabout the alignment requirement, or change hash_obj to use\nmemcpy (which might be slower).\n\n> static int hashtable_index(const unsigned char *sha1)\n> {\n> \tunsigned int i;\n> \tmemcpy(&i, sha1, sizeof(unsigned int));\n> \treturn (int)(i % obj_hash_size);\n> }\n\nHere, the sha1 argument does not come from struct object and\nmight not be sufficiently aligned for unsigned int, so the\nmemcpy is necessary.  If you look with git blame, you'll find\nthe commit where the memcpy was added for just this reason.\n"}]}