{"thread":{"id":"34767","subject":"Index Fileformat: stat(2) info necessary? What for?","startedAt":"2013-08-26T12:34:18Z","lastAt":"2013-08-26T22:17:18Z","messageCount":2,"participants":["Erik Bernoth","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"225897","messageId":"CAB46HOmMtgD+CtWUS3CQhr+ux1a3JP=hF2Cerd2nmDWzX5pxcw@mail.gmail.com","threadId":"34767","inReplyTo":null,"subject":"Index Fileformat: stat(2) info necessary? What for?","fromName":"Erik Bernoth","fromEmail":"erik.bernoth@gmail.com","sentAt":"2013-08-26T12:34:18Z","receivedAt":"2013-08-26T12:34:18Z","isPatch":false,"sender":{"key":"erik.bernoth@gmail.com","avatar":null},"body":"Hi,\n\nI am still working on implementing git in Python for self education\npurposes. Implementing the Index in memory was no problem after I\nunderstood how its done with help of Andreas Ericsson and Junio C\nHamano.\n\nNow I want to store an Index state to the filesystem in a\ngit-compatible file format. I looked up what the Git documentation has\nto say about that [1]. Now there's a lot of information (all the\nstat(2) stuff) that gets stored about the staged files, which I never\nneeded for file-IO in Python or Java. In my eyes if a person would be\ncloning my git repository he wouldn't need it as well, because the new\ninode on his system will probably be different from mine and applying\nthe access rights onto the cloning user id and group id would also\nmake sense, because that user introduced that file to that system.\n\nThus I am now missing concrete experience in when this stat(2)\ninformation comes in handy or if it would be completely okay in a\npython-git implementation to just store the info shown with `git\nls-files -s` to a file, maybe zlib.compressed like a git object. Of\ncourse I would then lose the compatibility with git repositories,\nwhich is a shame even if it would make sense. What is your opinion?\n\n\n[1] https://github.com/git/git/blob/master/Documentation/technical/index-format.txt\n\n\nCheers\nErik\n"},{"id":"225943","messageId":"20130826221718.GA12384@sigill.intra.peff.net","threadId":"34767","inReplyTo":"CAB46HOmMtgD+CtWUS3CQhr+ux1a3JP=hF2Cerd2nmDWzX5pxcw@mail.gmail.com","subject":"Re: Index Fileformat: stat(2) info necessary? What for?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-08-26T22:17:18Z","receivedAt":"2013-08-26T22:17:18Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Aug 26, 2013 at 02:34:18PM +0200, Erik Bernoth wrote:\n\n> Now there's a lot of information (all the stat(2) stuff) that gets\n> stored about the staged files, which I never needed for file-IO in\n> Python or Java. In my eyes if a person would be cloning my git\n> repository he wouldn't need it as well, because the new inode on his\n> system will probably be different from mine and applying the access\n> rights onto the cloning user id and group id would also make sense,\n> because that user introduced that file to that system.\n\nGit does not look at the index at all when cloning; only the actual\nobjects. So that stat information is not copied (the information in the\nclone's index comes from the checkout procedure on the receiving side).\n\n> Thus I am now missing concrete experience in when this stat(2)\n> information comes in handy or if it would be completely okay in a\n> python-git implementation to just store the info shown with `git\n> ls-files -s` to a file, maybe zlib.compressed like a git object. Of\n> course I would then lose the compatibility with git repositories,\n> which is a shame even if it would make sense. What is your opinion?\n\nThe stat information is there for performance. Think about how you would\nimplement \"git diff\" between the working tree and the index (or a tree).\n\nNaively, you would have to open each file and either compare it byte for\nbyte to what is committed, or hash it and compare the hash to what is\ncommitted. We have to open the files anyway to show a real patch, of\ncourse, but most files haven't been modified, and we would like to avoid\nopening them at all.\n\nBy keeping the stat information, we can check that the version in the\nindex is at some sha1 X, and then check that the stat information\nmatches what is on the disk, and know that what is on disk also has sha1\nX. That makes comparing to the index or a tree very cheap; we only call\nlstat() once per file, rather than opening and hashing all of the bytes.\n\n-Peff\n"}]}