git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Index format v5

From
Michael Haggerty <mhagger@alum.mit.edu>
Date
May 19, 2012, 05:40 UTC
Message-ID
<4FB73268.4020204@alum.mit.edu>
In-Reply-To
<20120516215407.GA1738@tgummerer.surfnet.iacbox>
On 05/16/2012 11:54 PM, Thomas Gummerer wrote:
Show 13 quoted lines
> On 05/16, Michael Haggerty wrote:
>> I just reviewed version 1369bd855b86 of your script, and it is MUCH
>> better.  It's easy to read and review.  The functions that it
>> defines are now self-contained and could therefore be reused for
>> other purposes.  There are fewer magic numbers (though there are
>> still a few; I wonder if there is a way to get rid of those?)
>> You've done a nice job polishing up the code.
>
> Thanks for the feedback! I could get rid of the magic numbers for
> the crc code, but I'm not sure if it makes sense to replace the
> others with constants, since they only occur once in the file. I
> added comments instead explaining where those numbers come from
> instead.

I think it is possible to remove the last magic number and also to make the CRC handling easier. I have pushed some suggested changes to github [1]:

1. With the current code, trying to read a file that is less than 24 
bytes long would result in a struct.error (because it would try to 
unpack a string that is shorter than the struct) whereas the underlying 
error in this case should almost always be reported as a signature 
error.  So it is correct to read the signature separately from the rest 
of the header, but it is even more correct to check the signature 
(including its length) before reading on.
2. I introduce a class CRC to hold checksums, so that (a) the code for 
handling checksums can be encapsulated, and (b) an instance of this 
class can be passed into functions and mutated in-place, which is less 
cumbersome than requiring functions to return (data, crc) tuples.  This, 
in turn, makes possible...
3. ...a new function read_struct(f, s, crc), which reads the data for a 
struct.Struct from f, checksums it, and returns the unpacked data.  This 
function is more convenient to use than the old read_calc_crc().
4. The checksum instance can also be made responsible for verifying that 
the next four bytes in the file agree with the expected checksum.  This 
removes some more code duplication.  (See CRC.matches().)
5. You read the extension offsets using CRC_STRUCT.size, which is 
technically correct but misleading.  In fact, the extension offsets 
should be documented using their own EXTENSION_INDEX_STRUCT.  Also, it 
makes more sense to store the unpacked integer offsets to extoffsets 
rather than the raw 4-byte strings.
6. With a couple of more minor changes it is possible to replace the 
magic number "24" in read_index_entries().  With this change the 
computation is documented very explicitly and is also (somewhat) robust 
against future changes in the format.
Look over my changes and take whatever you want.
> Thanks, I have changed those. I added the docstrings for all read
> functions, they however don't seem to make sense for the print
> functions, since you're probably faster just reading the code for
> them.

That's fine. If you document nontrivial functions you are already doing much better than the git project average.

> One since I changed in addition is to those changes, is that I
> gave the exceptions names to make them more meaningful.
Good.
Michael
[1] https://github.com/mhagger/git/tree/pythonprototype
-- 
Michael Haggerty
mhagger@alum.mit.edu
http://softwareswirl.blogspot.com/
Previous: Thomas GummererNext: Thomas Gummerer
Message 46 of 49 in “Index format v5”
  1. Thomas GummererMay 3, 2012
  2. Thomas RastMay 3, 2012
  3. Junio C HamanoMay 3, 2012
  4. Michael HaggertyMay 4, 2012
  5. Robin RosenbergMay 7, 2012
  6. Ronan KeryellMay 3, 2012
  7. Thomas GummererMay 3, 2012
  8. Junio C HamanoMay 3, 2012
  9. Thomas RastMay 3, 2012
  10. Thomas RastMay 3, 2012
  11. Thomas RastMay 3, 2012
  12. Junio C HamanoMay 3, 2012
  13. Thomas GummererMay 3, 2012
  14. Robin RosenbergMay 7, 2012
  15. solo-git@goeswhere.comMay 3, 2012
  16. Nguyen Thai Ngoc DuyMay 4, 2012
  17. Thomas GummererMay 4, 2012
  18. Philip OakleyMay 4, 2012
  19. Junio C HamanoMay 4, 2012
  20. Nguyen Thai Ngoc DuyMay 6, 2012
  21. Thomas GummererMay 7, 2012
  22. Phil HordMay 6, 2012
  23. Thomas GummererMay 7, 2012
  24. Michael HaggertyMay 7, 2012
  25. Thomas GummererMay 8, 2012
  26. Nguyen Thai Ngoc DuyMay 8, 2012
  27. Nguyen Thai Ngoc DuyMay 8, 2012
  28. Thomas GummererMay 10, 2012
  29. Nguyen Thai Ngoc DuyMay 10, 2012
  30. Michael HaggertyMay 9, 2012
  31. Thomas GummererMay 10, 2012
  32. Michael HaggertyMay 10, 2012
  33. Thomas GummererMay 11, 2012
  34. Michael HaggertyMay 13, 2012
  35. Thomas GummererMay 14, 2012
  36. Michael HaggertyMay 14, 2012
  37. Thomas RastMay 14, 2012
  38. Michael HaggertyMay 15, 2012
  39. Thomas GummererMay 15, 2012
  40. Michael HaggertyMay 15, 2012
  41. Thomas GummererMay 18, 2012
  42. Michael HaggertyMay 19, 2012
  43. Thomas GummererMay 21, 2012
  44. Michael HaggertyMay 16, 2012
  45. Thomas GummererMay 16, 2012
  46. Michael HaggertyMay 19, 2012
  47. Thomas GummererMay 21, 2012
  48. Philip OakleyMay 13, 2012
  49. Thomas GummererMay 14, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.