{"thread":{"id":"4129","subject":"[RFC] Add \"rcs format diff\" support","startedAt":"2006-05-13T21:14:15Z","lastAt":"2006-05-16T20:49:00Z","messageCount":3,"participants":["Linus Torvalds","Al Viro"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"19894","messageId":"Pine.LNX.4.64.0605131405590.3866@g5.osdl.org","threadId":"4129","inReplyTo":null,"subject":"[RFC] Add \"rcs format diff\" support","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-13T21:14:15Z","receivedAt":"2006-05-13T21:14:15Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nAl was asking for the \"diff -n\" format, which is the old RCS format, and \nwhich is really easy to parse.\n\nNow, we can't use the \"-n\" flag, because we use that for something else, \nand quite frankly, I don't know what to do about the diff _header_ (RCS \nformat doesn't have a header, afaik), but this implements the actual core \n\"xdiff\" rcs-format patch emit logic, and exposes it with the \nXDL_EMIT_RCSFORMAT flag. \n\n(In order to get valid diffs, you also have to set the context to zero \nwhen you set the RCSFORMAT flag).\n\nIt also adds a \"--rcs-format\" flag to the git diff option parsing, so you \ncan test it out, but as mentioned, we will still emit the full git header.\n\nDavide - I think the \"xdiff/\" sub-part of the patch should apply fine to \nthe standard xdiff sources, but I'm not sure you're really interested. The \nheader issue doesn't matter there, of course, since xdiff doesn't output \nany headers (ie that is an issue for the higher-level user).\n\nThe biggest issue for the xdiff library was that I needed to pass down \nthe xecfg parameter deeper into the call-chain (ie down to xdl_emit_record \n& co). The rest is pretty trivial.\n\nAl - feel free to play with this. I didn't test it heavily, but it gave \nthe right output for the one case I compared with \"diff -n\". This patch is \non top of my previous patch to parse \"-U\" and \"--unified\".\n\nJunio - this is not really meant for applying, although I don't think \nthere is any real down-side to this either. \n\n\t\tLinus\n\n---\ndiff --git a/diff.c b/diff.c\nindex be925a3..fd8f454 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -568,6 +568,10 @@ static void builtin_diff(const char *nam\n \t\t\txecfg.ctxlen = strtoul(diffopts + 2, NULL, 10);\n \t\tecb.outf = fn_out;\n \t\tecb.priv = &ecbdata;\n+\t\tif (o->rcs_format) {\n+\t\t\txecfg.flags |= XDL_EMIT_RCSFORMAT;\n+\t\t\txecfg.ctxlen = 0;\n+\t\t}\n \t\txdl_diff(&mf1, &mf2, &xpp, &xecfg, &ecb);\n \t}\n \n@@ -1277,6 +1281,8 @@ int diff_opt_parse(struct diff_options *\n \t\toptions->output_format = DIFF_FORMAT_PATCH;\n \telse if (opt_arg(arg, 'U', \"unified\", &options->context))\n \t\toptions->output_format = DIFF_FORMAT_PATCH;\n+\telse if (!strcmp(arg, \"--rcs-format\"))\n+\t\toptions->rcs_format = 1;\n \telse if (!strcmp(arg, \"--patch-with-raw\")) {\n \t\toptions->output_format = DIFF_FORMAT_PATCH;\n \t\toptions->with_raw = 1;\ndiff --git a/diff.h b/diff.h\nindex bef586d..953beb9 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -29,6 +29,7 @@ struct diff_options {\n \t\t with_stat:1,\n \t\t tree_in_recursive:1,\n \t\t binary:1,\n+\t\t rcs_format:1,\n \t\t full_index:1,\n \t\t silent_on_remove:1,\n \t\t find_copies_harder:1;\ndiff --git a/xdiff/xdiff.h b/xdiff/xdiff.h\nindex 2540e8a..a52359e 100644\n--- a/xdiff/xdiff.h\n+++ b/xdiff/xdiff.h\n@@ -36,6 +36,7 @@ #define XDL_PATCH_MODEMASK ((1 << 8) - 1\n #define XDL_PATCH_IGNOREBSPACE (1 << 8)\n \n #define XDL_EMIT_FUNCNAMES (1 << 0)\n+#define XDL_EMIT_RCSFORMAT (1 << 1)\n \n #define XDL_MMB_READONLY (1 << 0)\n \ndiff --git a/xdiff/xemit.c b/xdiff/xemit.c\nindex ad5bfb1..e127469 100644\n--- a/xdiff/xemit.c\n+++ b/xdiff/xemit.c\n@@ -26,7 +26,7 @@ #include \"xinclude.h\"\n \n \n static long xdl_get_rec(xdfile_t *xdf, long ri, char const **rec);\n-static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *ecb);\n+static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *ecb, xdemitconf_t const *xecfg);\n static xdchange_t *xdl_get_hunk(xdchange_t *xscr, xdemitconf_t const *xecfg);\n \n \n@@ -40,12 +40,13 @@ static long xdl_get_rec(xdfile_t *xdf, l\n }\n \n \n-static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre, xdemitcb_t *ecb) {\n+static int xdl_emit_record(xdfile_t *xdf, long ri, char const *pre,\n+\txdemitcb_t *ecb, xdemitconf_t const *xecfg) {\n \tlong size, psize = strlen(pre);\n \tchar const *rec;\n \n \tsize = xdl_get_rec(xdf, ri, &rec);\n-\tif (xdl_emit_diffrec(rec, size, pre, psize, ecb) < 0) {\n+\tif (xdl_emit_diffrec(rec, size, pre, psize, ecb, xecfg) < 0) {\n \n \t\treturn -1;\n \t}\n@@ -129,14 +130,14 @@ int xdl_emit_diff(xdfenv_t *xe, xdchange\n \t\t\t\t      sizeof(funcbuf), &funclen);\n \t\t}\n \t\tif (xdl_emit_hunk_hdr(s1 + 1, e1 - s1, s2 + 1, e2 - s2,\n-\t\t\t\t      funcbuf, funclen, ecb) < 0)\n+\t\t\t\t      funcbuf, funclen, ecb, xecfg) < 0)\n \t\t\treturn -1;\n \n \t\t/*\n \t\t * Emit pre-context.\n \t\t */\n \t\tfor (; s1 < xch->i1; s1++)\n-\t\t\tif (xdl_emit_record(&xe->xdf1, s1, \" \", ecb) < 0)\n+\t\t\tif (xdl_emit_record(&xe->xdf1, s1, \" \", ecb, xecfg) < 0)\n \t\t\t\treturn -1;\n \n \t\tfor (s1 = xch->i1, s2 = xch->i2;; xch = xch->next) {\n@@ -144,21 +145,21 @@ int xdl_emit_diff(xdfenv_t *xe, xdchange\n \t\t\t * Merge previous with current change atom.\n \t\t\t */\n \t\t\tfor (; s1 < xch->i1 && s2 < xch->i2; s1++, s2++)\n-\t\t\t\tif (xdl_emit_record(&xe->xdf1, s1, \" \", ecb) < 0)\n+\t\t\t\tif (xdl_emit_record(&xe->xdf1, s1, \" \", ecb, xecfg) < 0)\n \t\t\t\t\treturn -1;\n \n \t\t\t/*\n \t\t\t * Removes lines from the first file.\n \t\t\t */\n \t\t\tfor (s1 = xch->i1; s1 < xch->i1 + xch->chg1; s1++)\n-\t\t\t\tif (xdl_emit_record(&xe->xdf1, s1, \"-\", ecb) < 0)\n+\t\t\t\tif (xdl_emit_record(&xe->xdf1, s1, \"-\", ecb, xecfg) < 0)\n \t\t\t\t\treturn -1;\n \n \t\t\t/*\n \t\t\t * Adds lines from the second file.\n \t\t\t */\n \t\t\tfor (s2 = xch->i2; s2 < xch->i2 + xch->chg2; s2++)\n-\t\t\t\tif (xdl_emit_record(&xe->xdf2, s2, \"+\", ecb) < 0)\n+\t\t\t\tif (xdl_emit_record(&xe->xdf2, s2, \"+\", ecb, xecfg) < 0)\n \t\t\t\t\treturn -1;\n \n \t\t\tif (xch == xche)\n@@ -171,7 +172,7 @@ int xdl_emit_diff(xdfenv_t *xe, xdchange\n \t\t * Emit post-context.\n \t\t */\n \t\tfor (s1 = xche->i1 + xche->chg1; s1 < e1; s1++)\n-\t\t\tif (xdl_emit_record(&xe->xdf1, s1, \" \", ecb) < 0)\n+\t\t\tif (xdl_emit_record(&xe->xdf1, s1, \" \", ecb, xecfg) < 0)\n \t\t\t\treturn -1;\n \t}\n \ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\nindex 21ab8e7..b0d075a 100644\n--- a/xdiff/xutils.c\n+++ b/xdiff/xutils.c\n@@ -43,10 +43,20 @@ long xdl_bogosqrt(long n) {\n \n \n int xdl_emit_diffrec(char const *rec, long size, char const *pre, long psize,\n-\t\t     xdemitcb_t *ecb) {\n+\t\t     xdemitcb_t *ecb, xdemitconf_t const *xecfg) {\n \tmmbuffer_t mb[3];\n \tint i;\n \n+\tif (xecfg->flags & XDL_EMIT_RCSFORMAT) {\n+\t\tif (*pre != '+')\n+\t\t\treturn 0;\n+\t\tmb[0].ptr = (char *) rec;\n+\t\tmb[0].size = size;\n+\t\tif (ecb->outf(ecb->priv, mb, 1) < 0)\n+\t\t\treturn -1;\n+\t\treturn 0;\n+\t}\n+\n \tmb[0].ptr = (char *) pre;\n \tmb[0].size = psize;\n \tmb[1].ptr = (char *) rec;\n@@ -249,11 +259,34 @@ long xdl_atol(char const *str, char cons\n \n \n int xdl_emit_hunk_hdr(long s1, long c1, long s2, long c2,\n-\t\t      const char *func, long funclen, xdemitcb_t *ecb) {\n+\t\t      const char *func, long funclen,\n+\t\t      xdemitcb_t *ecb, xdemitconf_t const *xecfg) {\n \tint nb = 0;\n \tmmbuffer_t mb;\n \tchar buf[128];\n \n+\tif (xecfg->flags & XDL_EMIT_RCSFORMAT) {\n+\t\tif (c1) {\n+\t\t\tbuf[nb++] = 'd';\n+\t\t\tnb += xdl_num_out(buf + nb, s1);\n+\t\t\tbuf[nb++] = ' ';\n+\t\t\tnb += xdl_num_out(buf + nb, c1);\n+\t\t\tbuf[nb++] = '\\n';\n+\t\t}\n+\t\tif (c2) {\n+\t\t\tbuf[nb++] = 'a';\n+\t\t\tnb += xdl_num_out(buf + nb, s2);\n+\t\t\tbuf[nb++] = ' ';\n+\t\t\tnb += xdl_num_out(buf + nb, c2);\n+\t\t\tbuf[nb++] = '\\n';\n+\t\t}\n+\t\tmb.ptr = buf;\n+\t\tmb.size = nb;\n+\t\tif (ecb->outf(ecb->priv, &mb, 1) < 0)\n+\t\t\treturn -1;\n+\t\treturn 0;\n+\t}\n+\n \tmemcpy(buf, \"@@ -\", 4);\n \tnb += 4;\n \ndiff --git a/xdiff/xutils.h b/xdiff/xutils.h\nindex ea38ee9..e5c6ed0 100644\n--- a/xdiff/xutils.h\n+++ b/xdiff/xutils.h\n@@ -26,7 +26,7 @@ #define XUTILS_H\n \n long xdl_bogosqrt(long n);\n int xdl_emit_diffrec(char const *rec, long size, char const *pre, long psize,\n-\t\t     xdemitcb_t *ecb);\n+\t\t     xdemitcb_t *ecb, xdemitconf_t const *xecfg);\n int xdl_cha_init(chastore_t *cha, long isize, long icount);\n void xdl_cha_free(chastore_t *cha);\n void *xdl_cha_alloc(chastore_t *cha);\n@@ -38,7 +38,8 @@ unsigned int xdl_hashbits(unsigned int s\n int xdl_num_out(char *out, long val);\n long xdl_atol(char const *str, char const **next);\n int xdl_emit_hunk_hdr(long s1, long c1, long s2, long c2,\n-\t\t      const char *func, long funclen, xdemitcb_t *ecb);\n+\t\t      const char *func, long funclen,\n+\t\t      xdemitcb_t *ecb, xdemitconf_t const *xecfg);\n \n \n \n"},{"id":"19898","messageId":"20060514001214.GB27946@ftp.linux.org.uk","threadId":"4129","inReplyTo":"Pine.LNX.4.64.0605131405590.3866@g5.osdl.org","subject":"Re: [RFC] Add \"rcs format diff\" support","fromName":"Al Viro","fromEmail":"viro@ftp.linux.org.uk","sentAt":"2006-05-14T00:12:14Z","receivedAt":"2006-05-14T00:12:14Z","isPatch":false,"sender":{"key":"viro@ftp.linux.org.uk","avatar":null},"body":"On Sat, May 13, 2006 at 02:14:15PM -0700, Linus Torvalds wrote:\n> \n> Al was asking for the \"diff -n\" format, which is the old RCS format, and \n> which is really easy to parse.\n\nHeh... And I've just managed to get around that stuff on plain git.  Have fun:\n\nUse:\n\tgit-remap-data [git-diff arguments] > map\n\tgit-remap map <old-log >remapped-old\n\tgit-remap /dev/null <new-log >remapped-new\n\tdiff -u remapped-old remapped-new\nwith old-log and new-log being build/sparse/whatever logs produced on\ntrees in question (for values of \"whatever logs\" including e.g. grep -n\nresults, etc.)\n\ngit-remap-data builds the description of how lines of old tree are mapped\nto the new one; git-remap is a filter using that data to massage log\nfrom the old tree to new one; lines of form\n<file>:<line>:<text>\nare turned into\nN:<new-file>:<new-line>:<text>\nif they survive in new tree and\nO:<file>:<line>:<text>\notherwise.\n\nHere they are; enjoy.  BTW, that puppy can be used on unified diffs with\nzero context; won't catch renames, obviously...\n\ngit-remap-data.sh:\n#!/bin/sh\nGIT_DIFF_OPTS=\"-u 0\" git-diff -M \"$@\" | git-remap\n\ngit-remap.c:\n\n/*\n * Copyright (c) 2006, Al Viro.  All rights reserved.\n * \n * Redistribution and use in source and binary forms, with or without\n * modification, are permitted provided that the following conditions\n * are met:\n * 1. Redistributions of source code must retain the above copyright\n *    notice, this list of conditions and the following disclaimer.\n * 2. Redistributions in binary form must reproduce the above copyright\n *    notice, this list of conditions and the following disclaimer in the\n *    documentation and/or other materials provided with the distribution.\n *\n * THIS SOFTWARE IS PROVIDED BY AUTHOR AND CONTRIBUTORS ``AS IS'' AND\n * ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE\n * IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE\n * ARE DISCLAIMED.  IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE\n * FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL\n * DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS\n * OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION)\n * HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT\n * LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY\n * OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF\n * SUCH DAMAGE.\n */\n\n#include <stdio.h>\n#include <stdint.h>\n#include <stdlib.h>\n#include <string.h>\n#include <limits.h>\n\nchar *prefix1 = \"a/\", *prefix2 = \"b/\";\nsize_t len1, len2;\n\nchar *line;\nsize_t size;\n\nvoid die(char *s)\n{\n\tfprintf(stderr, \"remap: %s\\n\", s);\n\texit(1);\n}\n\nvoid Enomem(void)\n{\n\tdie(\"out of memory\");\n}\n\nvoid Eio(void)\n{\n\tdie(\"IO error\");\n}\n\nint getline(FILE *f)\n{\n\tchar *s;\n\tif (!fgets(line, size, f)) {\n\t\tif (!feof(f))\n\t\t\tEio();\n\t\treturn 0;\n\t}\n\tfor (s = line + strlen(line); s[-1] != '\\n'; s = s + strlen(s)) {\n\t\tif (s == line + size - 1) {\n\t\t\tline = realloc(line, 2 * size);\n\t\t\tif (!line)\n\t\t\t\tEnomem();\n\t\t\ts = line + size - 1;\n\t\t\tsize *= 2;\n\t\t}\n\t\tif (!fgets(s, size - (s - line), f)) {\n\t\t\tif (!feof(f))\n\t\t\t\tEio();\n\t\t\treturn 1;\n\t\t}\n\t}\n\ts[-1] = '\\0';\n\treturn 1;\n}\n\n/* to == 0 -> deletion */\nstruct range_map {\n\tint from, to;\n};\n\nstruct file_map {\n\tchar *name;\n\tstruct file_map *next;\n\tchar *new_name;\n\tint count;\n\tint allocated;\n\tint last;\n\tstruct range_map ranges[];\n};\n\nstruct file_map *alloc_map(char *name)\n{\n\tstruct file_map *map;\n\n\tmap = malloc(sizeof(struct file_map) + 16 * sizeof(struct range_map));\n\tif (!map)\n\t\tEnomem();\n\tmap->name = map->new_name = strdup(name);\n\tif (!map->name)\n\t\tEnomem();\n\tmap->count = 0;\n\tmap->allocated = 16;\n\tmap->next = NULL;\n\tmap->last = 0;\n\treturn map;\n}\n\n/* this is 32bit FNV1 */\nuint32_t FNV_hash(char *name)\n{\n\tuint32_t n = 0x811c9dc5;\n\twhile (*name) {\n\t\tunsigned char c = *name++;\n\t\tn *= 0x01000193;\n\t\tn ^= c;\n\t}\n\treturn n;\n}\n\nstruct file_map *hash[1024];\n\nint hash_map(struct file_map *map)\n{\n\tint n = FNV_hash(map->name) % 1024;\n\tstruct file_map **p = &hash[n];\n\n\twhile (*p) {\n\t\tif (!strcmp((*p)->name, map->name))\n\t\t\treturn 0;\n\t\tp = &(*p)->next;\n\t}\n\t*p = map;\n\tif (map->new_name && !map->count)\n\t\treturn 0;\n\tif (map->new_name && map->ranges[0].from != 1)\n\t\treturn 0;\n\treturn 1;\n}\n\nstruct file_map *find_map(char *name)\n{\n\tstatic struct file_map *last = NULL;\n\tint n = FNV_hash(name) % 1024;\n\tstruct file_map *p;\n\n\tif (last && !strcmp(last->name, name))\n\t\treturn last;\n\n\tfor (p = hash[n]; p && strcmp(p->name, name); p = p->next)\n\t\t;\n\tif (p)\n\t\tlast = p;\n\treturn p;\n}\n\nvoid parse_map(char *name)\n{\n\tstruct file_map *map = NULL;\n\tstruct range_map *range;\n\tchar *s;\n\tFILE *f;\n\n\tf = fopen(name, \"r\");\n\tif (!f)\n\t\tdie(\"can't open map\");\n\twhile (getline(f)) {\n\t\tif (line[0] == 'D') {\n\t\t\tif (map && !hash_map(map))\n\t\t\t\tgoto Ebadmap;\n\t\t\tif (line[1] != ' ')\n\t\t\t\tgoto Ebadmap;\n\t\t\tif (strchr(line + 2, ' '))\n\t\t\t\tgoto Ebadmap;\n\t\t\tmap = alloc_map(line + 2);\n\t\t\tmap->new_name = NULL;\n\t\t\tcontinue;\n\t\t}\n\t\tif (line[0] == 'M') {\n\t\t\tif (map && !hash_map(map))\n\t\t\t\tgoto Ebadmap;\n\t\t\tif (line[1] != ' ')\n\t\t\t\tgoto Ebadmap;\n\t\t\ts = strchr(line + 2, ' ');\n\t\t\tif (!s)\n\t\t\t\tgoto Ebadmap;\n\t\t\t*s++ = '\\0';\n\t\t\tif (strchr(s, ' '))\n\t\t\t\tgoto Ebadmap;\n\t\t\tmap = alloc_map(line + 2);\n\t\t\tif (strcmp(line + 2, s)) {\n\t\t\t\tmap->new_name = strdup(s);\n\t\t\t\tif (!map->new_name)\n\t\t\t\t\tEnomem();\n\t\t\t}\n\t\t\tcontinue;\n\t\t}\n\t\tif (!map || !map->new_name)\n\t\t\tgoto Ebadmap;\n\t\tif (map->count == map->allocated) {\n\t\t\tint n = 2 * map->allocated;\n\t\t\tmap = realloc(map, sizeof(struct file_map) +\n\t\t\t\t\t   n * sizeof(struct range_map));\n\t\t\tif (!map)\n\t\t\t\tEnomem();\n\t\t\tmap->allocated = n;\n\t\t}\n\t\trange = &map->ranges[map->count++];\n\t\tif (sscanf(line, \"%d %d%*c\", &range->from, &range->to) != 2)\n\t\t\tgoto Ebadmap;\n\t\tif (range > map->ranges && range->from <= range[-1].from)\n\t\t\tgoto Ebadmap;\n\t}\n\tif (map && !hash_map(map))\n\t\tgoto Ebadmap;\n\tfclose(f);\n\treturn;\nEbadmap:\n\tdie(\"bad map\");\n}\n\nstruct range_map *find_range(struct file_map *map, int l)\n{\n\tstruct range_map *range = &map->ranges[map->last];\n\tstruct range_map *p;\n\n\tif (range->from <= l) {\n\t\tp = &map->ranges[map->count - 1];\n\t\tif (p->from > l) {\n\t\t\tfor (p = range; p->from <= l; p++)\n\t\t\t\t;\n\t\t\tp--;\n\t\t}\n\t} else {\n\t\tfor (p = map->ranges; p->from <= l; p++)\n\t\t\t;\n\t\tp--;\n\t}\n\tmap->last = p - map->ranges;\n\treturn p;\n}\n\nvoid mapline(void)\n{\n\tstruct file_map *map;\n\tstruct range_map *range;\n\tunsigned long l;\n\tchar *s1, *s2;\n\tchar *name;\n\n\ts1 = strchr(line, ':');\n\tif (!s1)\n\t\tgoto noise;\n\ts2 = strchr(line, ' ');\n\tif (s2 && s2 < s1)\n\t\tgoto noise;\n\tl = strtoul(s1 + 1, &s2, 10);\n\tif (s2 == s1 + 1 || *s2 != ':' || !l || l > INT_MAX)\n\t\tgoto noise;\n\t*s1++ = *s2++ = '\\0';\n\tname = line;\n\tmap = find_map(line);\n\tif (!map)\n\t\tgoto new;\n\tif (!map->new_name)\n\t\tgoto old;\n\tname = map->new_name;\n\trange = find_range(map, l);\n\tif (!range->to)\n\t\tgoto old;\n\tl += range->to - range->from;\nnew:\n\tprintf(\"N:%s:%lu:%s\\n\", name, l, s2);\n\treturn;\nold:\n\ts1[-1] = s2[-1] = ':';\n\tprintf(\"O:%s\\n\", line);\n\treturn;\nnoise:\n\tprintf(\"%s\\n\", line);\n}\n\nint parse_hunk(int *l1, int *l2, int *n1, int *n2)\n{\n\tunsigned long n;\n\tchar *s, *p;\n\tif (line[3] != '-')\n\t\treturn 0;\n\tn = strtoul(line + 4, &s, 10);\n\tif (s == line + 4 || n > INT_MAX)\n\t\treturn 0;\n\t*l1 = n;\n\tif (*s == ',') {\n\t\tn = strtoul(s + 1, &p, 10);\n\t\tif (p == s + 1 || n > INT_MAX)\n\t\t\treturn 0;\n\t\t*n1 = n;\n\t\tif (!n)\n\t\t\t(*l1)++;\n\t} else {\n\t\tp = s;\n\t\t*n1 = 1;\n\t}\n\tif (*p != ' ' || p[1] != '+')\n\t\treturn 0;\n\tn = strtoul(p + 2, &s, 10);\n\tif (s == p + 2 || n > INT_MAX)\n\t\treturn 0;\n\t*l2 = n;\n\tif (*s == ',') {\n\t\tn = strtoul(s + 1, &p, 10);\n\t\tif (p == s + 1 || n > INT_MAX)\n\t\t\treturn 0;\n\t\t*n2 = n;\n\t\tif (!n)\n\t\t\t(*l2)++;\n\t} else {\n\t\tp = s;\n\t\t*n2 = 1;\n\t}\n\treturn 1;\n}\n\nvoid parse_diff(void)\n{\n\tint skipping = -1, suppress = 1;\n\tchar *name1 = NULL, *name2 = NULL;\n\tint from = 1, to = 1;\n\tint l1, l2, n1, n2;\n\tenum cmd {\n\t\tDiff, Hunk, New, Del, Copy, Rename, Junk\n\t} cmd;\n\tstatic struct { const char *s; size_t len; } pref[] = {\n\t\t[Hunk] = {\"@@ \", 3},\n\t\t[Diff] = {\"diff \", 5},\n\t\t[New] = {\"new file \", 9},\n\t\t[Del] = {\"deleted file \", 12},\n\t\t[Copy] = {\"copy from \", 10},\n\t\t[Rename] = {\"rename from \", 11},\n\t\t[Junk] = {\"\", 0},\n\t};\n\tsize_t len1 = strlen(prefix1), len2 = strlen(prefix2);\n\n\twhile (getline(stdin)) {\n\t\tif (skipping > 0) {\n\t\t\tswitch (line[0]) {\n\t\t\tcase '+':\n\t\t\tcase '-':\n\t\t\tcase '\\\\':\n\t\t\t\tcontinue;\n\t\t\t}\n\t\t}\n\t\tfor (cmd = 0; strncmp(line, pref[cmd].s, pref[cmd].len); cmd++)\n\t\t\t;\n\t\tswitch (cmd) {\n\t\tcase Hunk:\n\t\t\tif (skipping < 0)\n\t\t\t\tgoto Ediff;\n\t\t\tif (!suppress) {\n\t\t\t\tif (!skipping)\n\t\t\t\t\tprintf(\"M %s %s\\n\", name1, name2);\n\t\t\t\tif (!parse_hunk(&l1, &l2, &n1, &n2))\n\t\t\t\t\tgoto Ediff;\n\t\t\t\tif (l1 > from)\n\t\t\t\t\tprintf(\"%d %d\\n\", from, to);\n\t\t\t\tif (n1)\n\t\t\t\t\tprintf(\"%d 0\\n\", l1);\n\t\t\t\tfrom = l1 + n1;\n\t\t\t\tto = l2 + n2;\n\t\t\t}\n\t\t\tskipping = 1;\n\t\t\tbreak;\n\t\tcase Diff:\n\t\t\tif (!suppress) {\n\t\t\t\tif (!skipping)\n\t\t\t\t\tprintf(\"M %s %s\\n\", name1, name2);\n\t\t\t\tprintf(\"%d %d\\n\", from, to);\n\t\t\t}\n\t\t\tfree(name1);\n\t\t\tfree(name2);\n\t\t\tname2 = strrchr(line, ' ');\n\t\t\tif (!name2)\n\t\t\t\tgoto Ediff;\n\t\t\t*name2 = '\\0';\n\t\t\tname1 = strrchr(line, ' ');\n\t\t\tif (!name1)\n\t\t\t\tgoto Ediff;\n\t\t\tif (strncmp(name1 + 1, prefix1, len1))\n\t\t\t\tgoto Ediff;\n\t\t\tif (strncmp(name2 + 1, prefix2, len2))\n\t\t\t\tgoto Ediff;\n\t\t\tname1 = strdup(name1 + len1 + 1);\n\t\t\tname2 = strdup(name2 + len2 + 1);\n\t\t\tif (!name1 || !name2)\n\t\t\t\tgoto Ediff;\n\t\t\tskipping = 0;\n\t\t\tsuppress = 0;\n\t\t\tfrom = to = 1;\n\t\t\tbreak;\n\t\tcase New:\n\t\t\tif (skipping)\n\t\t\t\tgoto Ediff;\n\t\t\tsuppress = 1;\n\t\t\tbreak;\n\t\tcase Del:\n\t\tcase Copy:\n\t\t\tif (skipping)\n\t\t\t\tgoto Ediff;\n\t\t\tprintf(\"D %s\\n\", name2);\n\t\t\tsuppress = 1;\n\t\t\tbreak;\n\t\tcase Rename:\n\t\t\tif (skipping)\n\t\t\t\tgoto Ediff;\n\t\t\tprintf(\"D %s\\n\", name2);\n\t\t\tbreak;\n\t\tdefault:\n\t\t\tbreak;\n\t\t}\n\t}\n\treturn;\nEdiff:\n\tdie(\"odd diff\");\n}\n\nint main(int argc, char **argv)\n{\n\tint skipping = 0;\n\tsize = 256;\n\tline = malloc(size);\n\tif (!line)\n\t\tEnomem();\n\tif (argc < 2) {\n\t\tparse_diff();\n\t} else {\n\t\tparse_map(argv[1]);\n\t\twhile (getline(stdin))\n\t\t\tmapline();\n\t}\n\treturn 0;\n}\n"},{"id":"20080","messageId":"20060516204900.GA9051@ftp.linux.org.uk","threadId":"4129","inReplyTo":"20060514001214.GB27946@ftp.linux.org.uk","subject":"Re: [RFC] Add \"rcs format diff\" support","fromName":"Al Viro","fromEmail":"viro@ftp.linux.org.uk","sentAt":"2006-05-16T20:49:00Z","receivedAt":"2006-05-16T20:49:00Z","isPatch":false,"sender":{"key":"viro@ftp.linux.org.uk","avatar":null},"body":"Use:\n\tdiff-remap-data <dir1> <dir2> >map\nor\n\tgit-remap-data <git-diff arguments> >map\nwill build information for remapper,\n\tgit-remap <map> <options>\nwill do line numbers remapping.\n\ngit-remap is a filter.  It takes map as argument and, in the simplest form,\nwill look at the lines in stdin that have form\n<filename>:<number>:<text>\nIf the indicated line from old tree had survived into the new one, we will\nget\nN:<new-filename>:<new-number>:<text>\non the output.  If it hadn't, we get\nO:<filename>:<number>:<text>\nLines that do not have such form are passed unchanged.\n\nEven that is already very useful for log comparison.  E.g. if old-log is\nfrom the old tree and new-log is from the new one, we can do\n\tgit-remap map <old-log >foo\n\tgit-remap /dev/null <new-log >bar\n\tdiff -u foo bar\nand have the noise due to line number changes excluded (empty map means\nidentity mapping, so the second line will simply slap N: on all lines of\nform <filename>:<number>:<text> in new-log).\n\nNote that it's not just for build logs; the thing is useful for sparse logs,\ngrep -n output, etc., etc. \n\nBehaviour described above is the default; what _really_ happens is\nthat we take lines of form\n<original_prefix><filename>:<number>:<text>\nand replace them with\n<prefix_for_new><new-filename>:<new-number>:<text>\nor\n<prefix_for_old><filename>:<number>:<text>\nDefaults are :\", \"N:\" and \"O:\" resp.; what it gives us is the ability to\ndo multiple remappings.  IOW, we can say\n\ndiff-remap-data old-tree newer-tree > map1\ndiff-remap-data newer-tree current-tree > map2\ngit-remap -o old: map1 <old-log | git-remap -p N: -o newer: -n current: map2>foo\n\nand get lines that didn't make it into the newer tree marked with old: and\notherwise be unchanged, ones that made it to newer, but not the current to\nbe marked with newer: and have the filenames/line numbers remapped and ones\nthat made it all the way be marked with current: and remapped all the way\nto current tree.\n\nThat's quite useful when you want to carry logs for a while, basically using\nthem as annotated TODO (\"logs\" here can very well be results of grep -n with\nannotations added to them).  You can have all still relevant bits stay with\nthe locations in text and see what had fallen out.\n\nNote on relation to git:\n\t* git-remap, despite the name, doesn't need git to work\n\t* diff-remap-data doesn't need git to work\n\t* git-remap-data _does_ need it.  Aside of working on revisions in\ngit repository instead of a couple of directory trees, it generates slightly\nbetter map than diff-remap-data does.  I.e. it manages to remap more lines -\nit does notice renames.\n\nThis stuff lives on ftp.linux.org.uk/pub/people/viro/remapper/; I'm not\nsure what to do with it wrt distributing - submit for inclusion into\ngit, or leave that sucker standalone.  It can be used without git, but\nOTOH having it in git would make my life easier - I wouldn't have to\nthink about packaging it myself ;-)\n\nSeriously,\n\ta) feel free to play with it; hopefully it will be useful.\n\tb) review and comments are welcome.\n\tc) so would any thoughts regarding the right way to distribute it.\n"}]}