{"thread":{"id":"21083","subject":"[JGIT] patch-id","startedAt":"2009-09-28T22:21:00Z","lastAt":"2009-10-08T16:28:05Z","messageCount":2,"participants":["Nasser Grainawi","Shawn O. Pearce"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"123964","messageId":"4AC136CC.8040300@codeaurora.org","threadId":"21083","inReplyTo":null,"subject":"[JGIT] patch-id","fromName":"Nasser Grainawi","fromEmail":"nasser@codeaurora.org","sentAt":"2009-09-28T22:21:00Z","receivedAt":"2009-09-28T22:21:00Z","isPatch":false,"sender":{"key":"nasser@codeaurora.org","avatar":"https://avatars.githubusercontent.com/u/757421?v=4"},"body":"Hello again,\n\nI'm trying to add a public getPatchId method to the jgit Patch class and I\ncame up with some questions. Shawn previously mentioned that Patch already\ndoes the parsing of the patch; however, I can't quite wrap my head around\nhow/where/if data from that parsing is stored.\n\nIt seems Patch does some statistical number gathering, but at no point does\nit store a 'slimmed-down' version of a patch. I had the idea to just iterate\nover the FileHeader's and get the byte buffer of each, but I don't think\nthose buffers have the parsed data.\n\nIf I've mis-read the code (quite possible), someone please let me know.\nShort of that, suggestions for how to go about acquiring/storing a parsed\nrepresentation of the data with maximal existing code re-use would be\nappreciated.\n\nThanks,\nNasser\n"},{"id":"124386","messageId":"20091008162805.GE9261@spearce.org","threadId":"21083","inReplyTo":"4AC136CC.8040300@codeaurora.org","subject":"Re: [JGIT] patch-id","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2009-10-08T16:28:05Z","receivedAt":"2009-10-08T16:28:05Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Nasser Grainawi <nasser@codeaurora.org> wrote:\n> I'm trying to add a public getPatchId method to the jgit Patch class [...]\n>\n> It seems Patch does some statistical number gathering, but at no point does\n> it store a 'slimmed-down' version of a patch.\n\nIt parses the patch to create FileHeader objects, one for each\nfile mentioned in the script.  Within each FileHeader there is a\nHunkHeader object, one for each hunk present in the patch.  Within\neach HunkHeader there is an EditList composed of Edit instances;\neach Edit instance denotes a contiguous line range within that hunk.\n\nEdit instances come in one of 3 forms:\n\n  INSERT:  a run of + lines with no - lines\n  DELETE:  a run of - lines with no + lines\n  REPLACE: a mixture of - and + lines\n\nand their type is actually determined by the line numbers attached\nto them.  A INSERT has the same starting and ending line number on\nthe A side, but on the B side the ending line number is at least\none higher than the starting number.  DELETE is the reverse, and\nREPLACE has both ending numbers higher than the starting number.\n\nIIRC Edit uses 0 based offsets, so line 3 is actually position 2.\n\nThese HunkHeader and Edit instances are only available on a text\npatch, binary patches use a different representation for the\nbinary delta.  Combined diff patches (--cc format) also lack these\nHunkHeader/Edit instances as we don't have a generic n-way patch\nparser yet.\n\n> I had the idea to just iterate\n> over the FileHeader's and get the byte buffer of each, but I don't think\n> those buffers have the parsed data.\n\nThe HunkHeader and Edit instances really don't have the actual\nline data available to them, they only have the line numbers.\nTo generate a patch ID you'd need to get the line data too.\n\nWorse, IIRC the patch ID generation in C git favors a 3 line context.\n\nIn theory you could modify FileHeader or HunkHeader to produce\na RawText that uses the underlying byte[] returned by getBuffer()\nas the backing store, but create a specialized IntList which has the\nactual file line numbers mapped to the positions in the patch script.\nTo do that you'd need to re-walk the patch, like the toEditList()\nmethod in HunkHeader does.\n\nGiven that RawText you could feed it through something like\nDiffFormatter to create a patch with 3 lines of context, and hash\nthe relevant bits.\n\nBut... that seems like a lot of work.\n\nAlso, there is a class in Gerrit Code Review called EditList (not\nto be confused with JGit's EditList class!) that really should be\nmoved back over to JGit.  It has some useful routines for walking\nthrough a patch as a series of iterations.\n\n> Short of that, suggestions for how to go about acquiring/storing a parsed\n> representation of the data with maximal existing code re-use would be\n> appreciated.\n\nI'm coming up short on suggestions right now.  I'm not seeing an\neasy path to this without writing a bit of code.  I think you really\njust need to walk the patch... :-\\\n\n-- \nShawn.\n"}]}