{"thread":{"id":"770","subject":"I want to release a \"git-1.0\"","startedAt":"2005-05-30T20:00:42Z","lastAt":"2005-06-03T15:09:07Z","messageCount":64,"participants":["Linus Torvalds","jeff millar","Nicolas Pitre","Junio C Hamano","David Greaves","Dave Jones","Ryan Anderson","Chris Wedgwood","Dmitry Torokhov","Petr Baudis","Eric W. Biederman","David Lang","C. Scott Ananian","Daniel Barkalow","Brian O'Mahoney","Kay Sievers","Alexey Nezhdanov","Vincent Hanquez","Adam Kropelin"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"4265","messageId":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","threadId":"770","inReplyTo":null,"subject":"I want to release a \"git-1.0\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-30T20:00:42Z","receivedAt":"2005-05-30T20:00:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nOk, I'm at the point where I really think it's getting close to a 1.0, and\nmake another tar-ball etc. I obviously feel that it's already way superior\nto CVS, but I also realize that somebody who is used to CVS may not \nactually realize that very easily.\n\nSo before I do a 1.0 release, I want to write some stupid git tutorial for\na complete beginner that has only used CVS before, with a real example of\nhow to use raw git, and along those lines I actually want the thing to\nshow how to do something useful.\n\nSo before I do that, is there something people think is just too hard for\nsomebody coming from the CVS world to understand? I already realized that\nthe \"git-write-tree\" + \"git-commit-tree\" interfaces were just _too_ hard\nto put into a sane tutorial.\n\nI was showing off raw git to Steve Chamberlain yesterday, and showing it\nto him made some things pretty obvious - one of them being that\n\"git-init-db\" really needed to set up the initial refs etc). So I wrote\nthis silly \"git-commit-script\" to make it at least half-way palatable, but\nwhat else do people feel is \"too hard\"?\n\nI think I'll move the \"cvs2git\" script thing to git proper before the 1.0 \nrelease (again, in order to have the tutorial able to show what to do if \nyou already have an existing CVS tree), what else?\n\n\t\tLinus\n"},{"id":"4266","messageId":"429B78AD.3080109@adelphia.net","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"jeff millar","fromEmail":"wa1hco@adelphia.net","sentAt":"2005-05-30T20:33:49Z","receivedAt":"2005-05-30T20:33:49Z","isPatch":false,"sender":{"key":"wa1hco@adelphia.net","avatar":null},"body":"Linus Torvalds wrote:\n\n>So before I do that, is there something people think is just too hard for\n>somebody coming from the CVS world to understand? \n>\nI'm a fairly clueless cvs user, trying to use cg/git as a way to track a \nsingle\nuser project...using cogito, because that's easier, right?\n\nThe usage pattern that causing me problems right now.\n\ncg-init a whole directory tree (trying with /etc and a software project \ndirectory)\nnote that too many files got included (*.cache, *.backup, *.o, binaries, \netc)\nwant to stop tracking them, cg-rm also removes the file, don't want that.\n\nWhat's the best way to stop tracking files?\n\njeff\n"},{"id":"4268","messageId":"Pine.LNX.4.62.0505301644430.5330@localhost.localdomain","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2005-05-30T20:49:33Z","receivedAt":"2005-05-30T20:49:33Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 30 May 2005, Linus Torvalds wrote:\n\n> \n> Ok, I'm at the point where I really think it's getting close to a 1.0, and\n> make another tar-ball etc.\n\nAny chance you could merge my latest mkdelta patch _please_ ???\n\nI just posted it twice in the last 4 days and it still didn't appear in \nyour repository.\n\nAgain, the current version of mkdelta in your tree has a bug that can \nscrew things up, and it is fixed in the latest patch of course.\n\n\nNicolas\n"},{"id":"4269","messageId":"7vpsv81ld4.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-05-30T20:59:03Z","receivedAt":"2005-05-30T20:59:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"I'd really appreciate if you reconsider diff-* -O for inclusion\nbefore 1.0 happens.  It is probably the lowest impact among the\ndiffcore family.\n\nDon't I deserve it ;-)?\n\n"},{"id":"4270","messageId":"7vbr6s1kyv.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-05-30T21:07:36Z","receivedAt":"2005-05-30T21:07:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> I was showing off raw git to Steve Chamberlain yesterday, and showing it\nLT> to him made some things pretty obvious - one of them being that\nLT> \"git-init-db\" really needed to set up the initial refs etc). So I wrote\nLT> this silly \"git-commit-script\" to make it at least half-way palatable, but\nLT> what else do people feel is \"too hard\"?\n\nI think you need to clarify your intended audience first before\nsoliciting \"list of things that would help CVS user to convert\nto GIT\".  Specifically, which variant of GIT you are talking\nabout.\n\nI think you are talking about using the bare Plumbing.  I\nsuspect that some of the things you said \"too hard\" may be\ncoming from the fact that you did not use Cogito in the \"showing\noff\" you did.  I imagine Cogito users do not experience the\ntrouble you felt with git-init-db, since I presume they would\nrather use cg-init which IIUIC sets up the .git/refs structure\nfor its taste.\n\nHaving said that, I am in the same camp as you are in, in that\nthe (secondary) goal of my involvement in this project so far\nhas been to make the bare Plumbing confortable enough to use, to\nmake the choice of Porcelain more or less irrelevant.  As such,\nI am all for such a tutorial to convert CVS people to Plumbing\nGIT.\n\nNot that I'd volunteer writing big part of such a document.  I\nsuck at documentation, not just math ;-).\n\n"},{"id":"4271","messageId":"429B8FA0.1080903@dgreaves.com","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"David Greaves","fromEmail":"david@dgreaves.com","sentAt":"2005-05-30T22:11:44Z","receivedAt":"2005-05-30T22:11:44Z","isPatch":false,"sender":{"key":"david@dgreaves.com","avatar":"https://gravatar.com/avatar/ca67bad50999edcdd137c9a65da2381557d175bea99ae956afdabc5785e42b79?d=mp&s=160"},"body":"Linus Torvalds wrote:\n\n>So before I do a 1.0 release, I want to write some stupid git tutorial for\n>a complete beginner that has only used CVS before, with a real example of\n>how to use raw git, and along those lines I actually want the thing to\n>show how to do something useful.\n>\n>So before I do that, is there something people think is just too hard for\n>somebody coming from the CVS world to understand? I already realized that\n>the \"git-write-tree\" + \"git-commit-tree\" interfaces were just _too_ hard\n>to put into a sane tutorial.\n>\n>I was showing off raw git to Steve Chamberlain yesterday, and showing it\n>to him made some things pretty obvious - one of them being that\n>\"git-init-db\" really needed to set up the initial refs etc). So I wrote\n>this silly \"git-commit-script\" to make it at least half-way palatable, but\n>what else do people feel is \"too hard\"?\n>\n>I think I'll move the \"cvs2git\" script thing to git proper before the 1.0 \n>release (again, in order to have the tutorial able to show what to do if \n>you already have an existing CVS tree), what else?\n>  \n>\n\nIt seems to me that a tutorial for end users is inappropriate.\nYou should be writing a tutorial for porcelain implementors :)\n\nAnyway, a while back I split the commands into manipulation and\ninterrogation and then into ancillary commands and scripts. Do you\nactually agree with this grouping?\nhttp://www.kernel.org/pub/software/scm/git/docs/git.html\nIt may help to position who should be doing what.\n\nAlso, if you're writing a git-init-script, it may be that you're simply\nscripting common processes and could helpfully maintain consistency by\neither pulling some of the really trivial Cogito scripts (cg-init,\ncg-add, cg-rm) into the core 'ancillary' area or suggesting\nmodifications to Cogito as the current 'best of breed' implementation of\nthe low-level git usage process. Cogito also 'fixes' some useability\nissues such as using \"git-update-cache --add\" == \"cg-add\"\nI know you _can_ use git as an end user - but it seems that it's\ndesigned to be used by plumbers.\n\nOh, I'd also like to see something along the lines of my cg-Xignore\nbefore git hits 1.0\n\nOn the tutorial side - yesterday I started pulling together stuff from\nthe list about merging to complete the README where it says [ fixme:\ntalk about resolving merges here ]\n\nI haven't done much other than collect some discussion from the list and\nthe text from git-read-tree.txt.\nI do think this area needs more explanation as the whole 'stage' thing\nis pretty alien to CVS.\nI also noted a few people asking \"so I did this merge - what do I do now?\"\n\nThe working directory/cache/repository is also confusing sometimes -\nespecially when the cache and working-dir unexpectedly don't match.\n\nI also see in my notes: \"improve the docs around update-cache.\"\n\nDavid\n"},{"id":"4272","messageId":"20050530221214.GA29556@redhat.com","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Dave Jones","fromEmail":"davej@redhat.com","sentAt":"2005-05-30T22:12:14Z","receivedAt":"2005-05-30T22:12:14Z","isPatch":false,"sender":{"key":"davej@redhat.com","avatar":null},"body":"On Mon, May 30, 2005 at 01:00:42PM -0700, Linus Torvalds wrote:\n > \n > Ok, I'm at the point where I really think it's getting close to a 1.0, and\n > make another tar-ball etc. I obviously feel that it's already way superior\n > to CVS, but I also realize that somebody who is used to CVS may not \n > actually realize that very easily.\n > \n > So before I do a 1.0 release, I want to write some stupid git tutorial for\n > a complete beginner that has only used CVS before, with a real example of\n > how to use raw git, and along those lines I actually want the thing to\n > show how to do something useful.\n > \n > So before I do that, is there something people think is just too hard for\n > somebody coming from the CVS world to understand? I already realized that\n > the \"git-write-tree\" + \"git-commit-tree\" interfaces were just _too_ hard\n > to put into a sane tutorial.\n > \n > I was showing off raw git to Steve Chamberlain yesterday, and showing it\n > to him made some things pretty obvious - one of them being that\n > \"git-init-db\" really needed to set up the initial refs etc). So I wrote\n > this silly \"git-commit-script\" to make it at least half-way palatable, but\n > what else do people feel is \"too hard\"?\n\nI finally got around to actually trying to use git to maintain the\ncpufreq repository the last few days after reading Jeff Garzik's mini-howto[1]\n\nIt's not particularly complicated, but the number one thing that's bugged me is this..\n\n# commit changes\nGIT_AUTHOR_NAME=\"John Doe\"\t\t\\\n    GIT_AUTHOR_EMAIL=\"jdoe@foo.com\"\t\\\n    GIT_COMMITTER_NAME=\"Jeff Garzik\"\t\\\n    GIT_COMMITTER_EMAIL=\"jgarzik@pobox.com\"\t\\\n    git-commit-tree `git-write-tree`\t\\\n    -p $(cat .git/HEAD )\t\t\t\\\n    < changelog.txt\t\t\t\\\n    > .git/HEAD\n\nFor merging a lot of csets, thats a lot of typing per cset. So my .bashrc\nnow sets up GIT_COMMITTER_NAME & GIT_COMMITTER_EMAIL, because I don't\nforesee myself changing either of those anytime soon, which takes it down\nto\n    GIT_AUTHOR_NAME=\"John Doe\"      \\\n    GIT_AUTHOR_EMAIL=\"jdoe@foo.com\" \\\n    git-commit-tree `git-write-tree`    \\\n    -p $(cat .git/HEAD )            \\\n    < changelog.txt         \\\n    > .git/HEAD\n\nper-cset.  Maybe I have early on-set dementia, but the number of times\nI've typoed those two remaining environment variables is bizarre.\nI must've hit every known combination possible in my merge of ~30 patches.\n\nI could make the latter 4 lines of the above a shell alias to save some\ntyping, but those shell vars still bug me. Hmm, maybe I could create a\nwrapper that splits a \"Dave Jones <davej@redhat.com\" style string into two vars.\n\nI realise you've got a nifty bunch of tools to apply a whole mbox of\npatches, but that's not ideal if all of my patches aren't in mboxes\n(some I create myself and toss in my spool, some I pull from bugzilla etc..)\n\nTypos aside, the other thing that seems non-intuitive is the splitting up\nof the patch & changelog comment into seperate files during the patch-apply\nstage.\n\nMaybe your new git-commit-script wonder-tool fixes up all these problems\nalready, I'll take a look after food.\n\nIts pretty nifty stuff, but for merging a lot of patches in non-mbox format,\neither I'm doing something wrong, or its, well.. painful.\n\n\t\tDave\n\n[1] http://lkml.org/lkml/2005/5/26/11/index.html\n\n"},{"id":"4273","messageId":"20050530221922.GC21076@mythryan2.michonline.com","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Ryan Anderson","fromEmail":"ryan@michonline.com","sentAt":"2005-05-30T22:19:22Z","receivedAt":"2005-05-30T22:19:22Z","isPatch":false,"sender":{"key":"ryan@michonline.com","avatar":null},"body":"On Mon, May 30, 2005 at 01:00:42PM -0700, Linus Torvalds wrote:\n> \n> I think I'll move the \"cvs2git\" script thing to git proper before the 1.0 \n> release (again, in order to have the tutorial able to show what to do if \n> you already have an existing CVS tree), what else?\n\nUmm, why do you maintain two seperate \"git\" related trees?\n\nWhy not merge all of git-tools in, in a tools/ subdirectory?\n\nI've been meaning to ask the same question about \"gitweb\" for that\nmatter.  The distributions that want seperate packages for dependency\nreasons can handle that easily inside one tree, anyway, I believe.\n\nI'd guess part of this is a holdover from the fact that you needed an\nindependent tree for BitKeeper, but does it still make sense?\n\n-- \n\nRyan Anderson\n  sometimes Pug Majere\n"},{"id":"4274","messageId":"972477.0a6782ba1d3b9f05216ed520ef720fcf.ANY@taniwha.stupidest.org","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Chris Wedgwood","fromEmail":"cw@f00f.org","sentAt":"2005-05-30T22:32:42Z","receivedAt":"2005-05-30T22:32:42Z","isPatch":false,"sender":{"key":"cw@f00f.org","avatar":null},"body":"On Mon, May 30, 2005 at 01:00:42PM -0700, Linus Torvalds wrote:\n\n> So before I do that, is there something people think is just too\n> hard for somebody coming from the CVS world to understand? I already\n> realized that the \"git-write-tree\" + \"git-commit-tree\" interfaces\n> were just _too_ hard to put into a sane tutorial.\n\nI'm still at a loss how to do the equivalent of annotate.  I know a\ncouple of front ends can do this but I have no idea what command line\nmagic would be equivalent.\n"},{"id":"4276","messageId":"200505301755.15371.dtor_core@ameritech.net","threadId":"770","inReplyTo":"20050530221214.GA29556@redhat.com","subject":"Re: I want to release a \"git-1.0\"","fromName":"Dmitry Torokhov","fromEmail":"dtor_core@ameritech.net","sentAt":"2005-05-30T22:55:14Z","receivedAt":"2005-05-30T22:55:14Z","isPatch":false,"sender":{"key":"dtor_core@ameritech.net","avatar":null},"body":"On Monday 30 May 2005 17:12, Dave Jones wrote:\n> I realise you've got a nifty bunch of tools to apply a whole mbox of\n> patches, but that's not ideal if all of my patches aren't in mboxes\n> (some I create myself and toss in my spool, some I pull from bugzilla etc..)\n\nI mercilessly hacked Linus's scripts from git-tools repo to work with\nnon-mailbox patches, maybe you can make use of them too. Note that\nstripspace.c is not changed in any way whatsoever and mailsplit.c was\nchanged to handle my personal preference of having patch description\nin the form of:\n\nInput: make blah blah change\n---\n \nAnd Linus's script would eat that line.\n\n-- \nDmitry\n\n\n#include <stdio.h>\n#include <string.h>\n#include <ctype.h>\n\n/*\n * Remove empty lines from the beginning and end.\n *\n * Turn multiple consecutive empty lines into just one\n * empty line.\n */\nstatic void cleanup(char *line)\n{\n\tint len = strlen(line);\n\n\tif (len > 1 && line[len-1] == '\\n') {\n\t\tdo {\n\t\t\tunsigned char c = line[len-2];\n\t\t\tif (!isspace(c))\n\t\t\t\tbreak;\n\t\t\tline[len-2] = '\\n';\n\t\t\tlen--;\n\t\t\tline[len] = 0;\n\t\t} while (len > 1);\n\t}\n}\n\nint main(int argc, char **argv)\n{\n\tint empties = -1;\n\tchar line[1024];\n\n\twhile (fgets(line, sizeof(line), stdin)) {\n\t\tcleanup(line);\n\n\t\t/* Not just an empty line? */\n\t\tif (line[0] != '\\n') {\n\t\t\tif (empties > 0)\n\t\t\t\tputchar('\\n');\n\t\t\tempties = 0;\n\t\t\tfputs(line, stdout);\n\t\t\tcontinue;\n\t\t}\n\t\tif (empties < 0)\n\t\t\tcontinue;\n\t\tempties++;\n\t}\n\treturn 0;\n}\n\n\n/*\n * Totally braindamaged mbox splitter program.\n *\n * It just splits a mbox into a list of files: \"0001\" \"0002\" ..\n * so you can process them further from there.\n */\n#include <unistd.h>\n#include <stdlib.h>\n#include <fcntl.h>\n#include <sys/types.h>\n#include <sys/stat.h>\n#include <sys/mman.h>\n#include <string.h>\n#include <stdio.h>\n#include <ctype.h>\n#include <assert.h>\n\nstatic int usage(void)\n{\n\tfprintf(stderr, \"mailsplit <mbox> <directory>\\n\");\n\texit(1);\n}\n\nstatic int linelen(const char *map, unsigned long size)\n{\n\tint len = 0, c;\n\n\tdo {\n\t\tc = *map;\n\t\tmap++;\n\t\tsize--;\n\t\tlen++;\n\t} while (size && c != '\\n');\n\treturn len;\n}\n\nstatic int is_from_line(const char *line, int len)\n{\n\tconst char *colon;\n\n\tif (len < 20 || memcmp(\"From \", line, 5))\n\t\treturn 0;\n\n\tcolon = line + len - 2;\n\tline += 5;\n\tfor (;;) {\n\t\tif (colon < line)\n\t\t\treturn 0;\n\t\tif (*--colon == ':')\n\t\t\tbreak;\n\t}\n\n\tif (!isdigit(colon[-4]) ||\n\t    !isdigit(colon[-2]) ||\n\t    !isdigit(colon[-1]) ||\n\t    !isdigit(colon[ 1]) ||\n\t    !isdigit(colon[ 2]))\n\t\treturn 0;\n\n\t/* year */\n\tif (strtol(colon+3, NULL, 10) <= 90)\n\t\treturn 0;\n\n\t/* Ok, close enough */\n\treturn 1;\n}\n\nstatic int parse_email(const void *map, unsigned long size)\n{\n\tunsigned long offset;\n\n\tif (size < 6 || memcmp(\"From \", map, 5))\n\t\tgoto corrupt;\n\n\t/* Make sure we don't trigger on this first line */\n\tmap++; size--; offset=1;\n\n\t/*\n\t * Search for a line beginning with \"From \", and \n\t * having smething that looks like a date format.\n\t */\n\tdo {\n\t\tint len = linelen(map, size);\n\t\tif (is_from_line(map, len))\n\t\t\treturn offset;\n\t\tmap += len;\n\t\tsize -= len;\n\t\toffset += len;\n\t} while (size);\n\treturn offset;\n\ncorrupt:\n\tfprintf(stderr, \"corrupt mailbox\\n\");\n\texit(1);\n}\n\nint main(int argc, char **argv)\n{\n\tint fd, nr;\n\tstruct stat st;\n\tunsigned long size;\n\tvoid *map;\n\n\tif (argc != 3)\n\t\tusage();\n\tfd = open(argv[1], O_RDONLY);\n\tif (fd < 0) {\n\t\tperror(argv[1]);\n\t\texit(1);\n\t}\n\tif (chdir(argv[2]) < 0)\n\t\tusage();\n\tif (fstat(fd, &st) < 0) {\n\t\tperror(\"stat\");\n\t\texit(1);\n\t}\n\tsize = st.st_size;\n\tmap = mmap(NULL, size, PROT_READ, MAP_PRIVATE, fd, 0);\n\tif (-1 == (int)(long)map) {\n\t\tperror(\"mmap\");\n\t\texit(1);\n\t}\n\tclose(fd);\n\tnr = 0;\n\tdo {\n\t\tchar name[10];\n\t\tunsigned long len = parse_email(map, size);\n\t\tassert(len <= size);\n\t\tsprintf(name, \"%04d\", ++nr);\n\t\tfd = open(name, O_WRONLY | O_CREAT | O_EXCL, 0600);\n\t\tif (fd < 0) {\n\t\t\tperror(name);\n\t\t\texit(1);\n\t\t}\n\t\tif (write(fd, map, len) != len) {\n\t\t\tperror(\"write\");\n\t\t\texit(1);\n\t\t}\n\t\tclose(fd);\n\t\tmap += len;\n\t\tsize -= len;\n\t} while (size > 0);\n\treturn 0;\n}\n"},{"id":"4278","messageId":"7v7jhgz4oq.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"200505301755.15371.dtor_core@ameritech.net","subject":"Re: I want to release a \"git-1.0\"","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-05-30T23:15:17Z","receivedAt":"2005-05-30T23:15:17Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"On a related topic of making bare Plumbing a bit easier to use,\nhere is what I use to prepare patches, one patch per file, to be\nsent to Linus via e-mail.\n\nUsage:\n\n  $ git-format-patch-script HEAD linus\n\nAssuming that \"linus\" is the tip of the tree from Linus\n(typically stored in .git/branches/linus if you use Cogito), and\nHEAD is your additions on top of it, the above command will\nproduce patches in the format you have been seeing on this list\nfrom me, one file per commit, in .patches/XXXX-patch-title.txt\nfile.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\nsed -e 's/^X//' >git-format-patch-script <<\\EOF\nX#!/bin/sh\nX#\nX# Copyright (c) 2005 Junio C Hamano\nX#\nXjunio=\"$1\"\nXlinus=\"$2\"\nX\nXtmp=.tmp-series$$\nXtrap 'rm -f $tmp-*' 0 1 2 3 15\nX\nXseries=$tmp-series\nX\nXtitleScript='\nX\t1,/^$/d\nX\t: loop\nX\t/^$/b loop\nX\ts/[^-a-z.A-Z_0-9]/-/g\nX\ts/^--*//g\nX\ts/--*$//g\nX\ts/---*/-/g\nX\ts/$/.txt/\nX        s/\\.\\.\\.*/\\./g\nX\tq\nX'\nXO=\nXif test -f .git/patch-order\nXthen\nX\tO=-O.git/patch-order\nXfi\nXgit-rev-list \"$junio\" \"$linus\" >$series\nXtotal=`wc -l <$series`\nXi=$total\nXwhile read commit\nXdo\nX    title=`git-cat-file commit \"$commit\" | sed -e \"$titleScript\"`\nX    num=`printf \"%d/%d\" $i $total`\nX    file=`printf '%04d-%s' $i \"$title\"`\nX    i=`expr \"$i\" - 1`\nX    echo \"$file\"\nX    {\nX\tmailScript='\nX\t1,/^$/d\nX\t: loop\nX\t/^$/b loop\nX\ts|^|[PATCH '\"$num\"'] |\nX\t: body\nX\tp\nX\tn\nX\tb body'\nX\nX\tgit-cat-file commit \"$commit\" | sed -ne \"$mailScript\"\nX\techo '---'\nX\tgit-diff-tree -p $O \"$commit\" | diffstat -p1\nX\techo\nX\tgit-diff-tree -p $O \"$commit\"\nX    } >\".patches/$file\"\nXdone <$series\nEOF\n\n"},{"id":"4279","messageId":"200505301823.04338.dtor_core@ameritech.net","threadId":"770","inReplyTo":"200505301755.15371.dtor_core@ameritech.net","subject":"Re: I want to release a \"git-1.0\"","fromName":"Dmitry Torokhov","fromEmail":"dtor_core@ameritech.net","sentAt":"2005-05-30T23:23:02Z","receivedAt":"2005-05-30T23:23:02Z","isPatch":false,"sender":{"key":"dtor_core@ameritech.net","avatar":null},"body":"On Monday 30 May 2005 17:55, Dmitry Torokhov wrote:\n> On Monday 30 May 2005 17:12, Dave Jones wrote:\n> > I realise you've got a nifty bunch of tools to apply a whole mbox of\n> > patches, but that's not ideal if all of my patches aren't in mboxes\n> > (some I create myself and toss in my spool, some I pull from bugzilla etc..)\n> \n> I mercilessly hacked Linus's scripts from git-tools repo to work with\n> non-mailbox patches, maybe you can make use of them too. Note that\n> stripspace.c is not changed in any way whatsoever and mailsplit.c was\n> changed to handle my personal preference of having patch description\n> in the form of:\n> \n> Input: make blah blah change\n> ---\n>  \n> And Linus's script would eat that line.\n> \n\nOops, make it mailinfo.c, not mailsplit.c\n\n-- \nDmitry\n\n\n/*\n * Another stupid program, this one parsing the headers of an\n * email to figure out authorship and subject\n */\n#include <stdio.h>\n#include <stdlib.h>\n#include <string.h>\n#include <ctype.h>\n\nstatic FILE *cmitmsg, *patchfile, *filelist;\n\nstatic char line[1000];\nstatic char date[1000];\nstatic char name[1000];\nstatic char email[1000];\nstatic char subject[1000];\n\nstatic char *sanity_check(char *name, char *email)\n{\n\tint len = strlen(name);\n\tif (len < 3 || len > 60)\n\t\treturn email;\n\tif (strchr(name, '@') || strchr(name, '<') || strchr(name, '>'))\n\t\treturn email;\n\treturn name;\n}\n\nstatic int handle_from(char *line)\n{\n\tchar *at = strchr(line, '@');\n\tchar *dst;\n\n\tif (!at)\n\t\treturn 0;\n\n/*\n* If we already have one email, don't take any confusing lines\n*/\n\tif (*email && strchr(at + 1, '@'))\n\t\treturn 0;\n\n\twhile (at > line) {\n\t\tchar c = at[-1];\n\t\tif (isspace(c) || c == '<')\n\t\t\tbreak;\n\t\tat--;\n\t}\n\tdst = email;\n\tfor (;;) {\n\t\tunsigned char c = *at;\n\t\tif (!c || c == '>' || isspace(c))\n\t\t\tbreak;\n\t\t*at++ = ' ';\n\t\t*dst++ = c;\n\t}\n\t*dst++ = 0;\n\n\tat = line + strlen(line);\n\twhile (at > line) {\n\t\tunsigned char c = *--at;\n\t\tif (isalnum(c))\n\t\t\tbreak;\n\t\t*at = 0;\n\t}\n\n\tat = line;\n\tfor (;;) {\n\t\tunsigned char c = *at;\n\t\tif (!c)\n\t\t\tbreak;\n\t\tif (isalnum(c))\n\t\t\tbreak;\n\t\tat++;\n\t}\n\n\tat = sanity_check(at, email);\n\n\tstrcpy(name, at);\n\treturn 1;\n}\n\nstatic void handle_date(char *line)\n{\n\tstrcpy(date, line);\n}\n\nstatic void handle_subject(char *line)\n{\n\tstrcpy(subject, line);\n}\n\nstatic void add_subject_line(char *line)\n{\n\twhile (isspace(*line))\n\t\tline++;\n\t*--line = ' ';\n\tstrcat(subject, line);\n}\n\nstatic int check_special_line(char *line, int len)\n{\n\tstatic int cont = -1;\n\tif (!memcmp(line, \"From:\", 5) && isspace(line[5])) {\n\t\thandle_from(line + 6);\n\t\tcont = 0;\n\t\treturn 1;\n\t}\n\tif (!memcmp(line, \"Date:\", 5) && isspace(line[5])) {\n\t\thandle_date(line + 6);\n\t\tcont = 0;\n\t\treturn 1;\n\t}\n\tif (!memcmp(line, \"Subject:\", 8) && isspace(line[8])) {\n\t\thandle_subject(line + 9);\n\t\tcont = 1;\n\t\treturn 1;\n\t}\n\tif (isspace(*line)) {\n\t\tswitch (cont) {\n\t\tcase 0:\n\t\t\tfprintf(stderr,\n\t\t\t\t\"I don't do 'Date:' or 'From:' line continuations\\n\");\n\t\t\tbreak;\n\t\tcase 1:\n\t\t\tadd_subject_line(line);\n\t\t\treturn 1;\n\t\tdefault:\n\t\t\tbreak;\n\t\t}\n\t}\n\tcont = -1;\n\treturn 0;\n}\n\nstatic char *cleanup_subject(char *subject)\n{\n\tfor (;;) {\n\t\tchar *p;\n\t\tint len, remove;\n\t\tswitch (*subject) {\n\t\tcase 'r':\n\t\tcase 'R':\n\t\t\tif (!memcmp(\"e:\", subject + 1, 2)) {\n\t\t\t\tsubject += 3;\n\t\t\t\tcontinue;\n\t\t\t}\n\t\t\tbreak;\n\t\tcase ' ':\n\t\tcase '\\t':\n\t\tcase ':':\n\t\t\tsubject++;\n\t\t\tcontinue;\n\n\t\tcase '[':\n\t\t\tp = strchr(subject, ']');\n\t\t\tif (!p) {\n\t\t\t\tsubject++;\n\t\t\t\tcontinue;\n\t\t\t}\n\t\t\tlen = strlen(p);\n\t\t\tremove = p - subject;\n\t\t\tif (remove <= len * 2) {\n\t\t\t\tsubject = p + 1;\n\t\t\t\tcontinue;\n\t\t\t}\n\t\t\tbreak;\n\t\t}\n\t\treturn subject;\n\t}\n}\n\nstatic void cleanup_space(char *buf)\n{\n\tunsigned char c;\n\twhile ((c = *buf) != 0) {\n\t\tbuf++;\n\t\tif (isspace(c)) {\n\t\t\tbuf[-1] = ' ';\n\t\t\tc = *buf;\n\t\t\twhile (isspace(c)) {\n\t\t\t\tint len = strlen(buf);\n\t\t\t\tmemmove(buf, buf + 1, len);\n\t\t\t\tc = *buf;\n\t\t\t}\n\t\t}\n\t}\n}\n\n/*\n* Hacky hacky. This depends not only on -p1, but on\n* filenames not having some special characters in them,\n* like tilde.\n*/\nstatic void show_filename(char *line)\n{\n\tint len;\n\tchar *name = strchr(line, '/');\n\n\tif (!name || !isspace(*line))\n\t\treturn;\n\tname++;\n\tlen = 0;\n\tfor (;;) {\n\t\tunsigned char c = name[len];\n\t\tswitch (c) {\n\t\tdefault:\n\t\t\tlen++;\n\t\t\tcontinue;\n\n\t\tcase 0:\n\t\tcase ' ':\n\t\tcase '\\t':\n\t\tcase '\\n':\n\t\t\tbreak;\n\n/* patch tends to special-case these things.. */\n\t\tcase '~':\n\t\t\tbreak;\n\t\t}\n\t\tbreak;\n\t}\n/* remove \".orig\" from the end - common patch behaviour */\n\tif (len > 5 && !memcmp(name + len - 5, \".orig\", 5))\n\t\tlen -= 5;\n\tif (!len)\n\t\treturn;\n\tfprintf(filelist, \"%.*s\\n\", len, name);\n}\n\nstatic void handle_rest(void)\n{\n\tchar *sub = cleanup_subject(subject);\n\tcleanup_space(name);\n\tcleanup_space(date);\n\tcleanup_space(email);\n\tcleanup_space(sub);\n\tprintf(\"Author: %s\\nEmail: %s\\nSubject: %s\\nDate: %s\\n\\n\", name, email,\n\t       sub, date);\n\tFILE *out = cmitmsg;\n\n\tdo {\n/* Track filename information from the patch.. */\n\t\tif (!memcmp(\"---\", line, 3)) {\n\t\t\tout = patchfile;\n\t\t\tshow_filename(line + 3);\n\t\t}\n\n\t\tif (!memcmp(\"+++\", line, 3))\n\t\t\tshow_filename(line + 3);\n\n\t\tfputs(line, out);\n\t} while (fgets(line, sizeof(line), stdin) != NULL);\n\n\tif (out == cmitmsg) {\n\t\tfprintf(stderr, \"No patch found\\n\");\n\t\texit(1);\n\t}\n\n\tfclose(cmitmsg);\n\tfclose(patchfile);\n}\n\nstatic int eatspace(char *line)\n{\n\tint len = strlen(line);\n\twhile (len > 0 && isspace(line[len - 1]))\n\t\tline[--len] = 0;\n\treturn len;\n}\n\nstatic void handle_body(void)\n{\n\tint has_from = 0;\n\n/* First line of body can be a From: */\n\twhile (fgets(line, sizeof(line), stdin) != NULL) {\n\t\tint len = eatspace(line);\n\t\tif (!len)\n\t\t\tcontinue;\n\t\tif (!memcmp(\"From:\", line, 5) && isspace(line[5])) {\n\t\t\tif (!has_from && handle_from(line + 6)) {\n\t\t\t\thas_from = 1;\n\t\t\t\tcontinue;\n\t\t\t}\n\t\t}\n\t\tline[len] = '\\n';\n\t\thandle_rest();\n\t\tbreak;\n\t}\n}\n\nstatic void usage(void)\n{\n\tfprintf(stderr, \"mailinfo msg-file path-file filelist-file < email\\n\");\n\texit(1);\n}\n\nint main(int argc, char **argv)\n{\n\tint mail_patch = 0;\n\n\tif (argc != 4)\n\t\tusage();\n\tcmitmsg = fopen(argv[1], \"w\");\n\tif (!cmitmsg) {\n\t\tperror(argv[1]);\n\t\texit(1);\n\t}\n\tpatchfile = fopen(argv[2], \"w\");\n\tif (!patchfile) {\n\t\tperror(argv[2]);\n\t\texit(1);\n\t}\n\tfilelist = fopen(argv[3], \"w\");\n\tif (!filelist) {\n\t\tperror(argv[3]);\n\t\texit(1);\n\t}\n\twhile (fgets(line, sizeof(line), stdin) != NULL) {\n\t\tint len = eatspace(line);\n\t\tif (!len) {\n\t\t\tif (!mail_patch)\n\t\t\t\tfputs(\"\\n\", cmitmsg);\n\n\t\t\thandle_body();\n\t\t\tbreak;\n\t\t}\n\t\tif (check_special_line(line, len)) {\n\t\t\tmail_patch = 1;\n\t\t\trewind(cmitmsg);\n\t\t}\n\t\tif (!mail_patch) {\n\t\t\tline[len] = '\\n';\n\t\t\tfputs(line, cmitmsg);\n\t\t}\n\t}\n\treturn 0;\n}\n"},{"id":"4281","messageId":"526551.b8e690d40dd297a911bca6b35fc0dd2e.ANY@taniwha.stupidest.org","threadId":"770","inReplyTo":"972477.0a6782ba1d3b9f05216ed520ef720fcf.ANY@taniwha.stupidest.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Chris Wedgwood","fromEmail":"cw@f00f.org","sentAt":"2005-05-30T23:56:33Z","receivedAt":"2005-05-30T23:56:33Z","isPatch":false,"sender":{"key":"cw@f00f.org","avatar":null},"body":"On Mon, May 30, 2005 at 03:32:42PM -0700, Chris Wedgwood wrote:\n\n> I'm still at a loss how to do the equivalent of annotate.  I know a\n> couple of front ends can do this but I have no idea what command line\n> magic would be equivalent.\n\nA few people asked what does this now.  Git Tracker does, a (random)\nexample of which might be:\n\n  http://www.tglx.de/cgi-bin/gittracker/annotate/tracker-linux/torvalds/linux-2.6.git/mm/mmap.c?blob=de54acd9942f9929004921042721df5cdfe2b6c7\n"},{"id":"4282","messageId":"20050531001916.GD10439@pasky.ji.cz","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-05-31T00:19:16Z","receivedAt":"2005-05-31T00:19:16Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Mon, May 30, 2005 at 10:00:42PM CEST, I got a letter\nwhere Linus Torvalds <torvalds@osdl.org> told me that...\n> Ok, I'm at the point where I really think it's getting close to a 1.0, and\n> make another tar-ball etc. I obviously feel that it's already way superior\n> to CVS, but I also realize that somebody who is used to CVS may not \n> actually realize that very easily.\n\nCan we (well, me) count on the output format of the git commands being\nstabilized now and not change in a backwards-incompatible way from now\non? I would like to finally remove the git itself from Cogito, but for\nthat I have to be able to rely on the fact that as long as the user has\ngit version >=N, it will work (assuming that Cogito is bugless ;-).\n\n> So before I do a 1.0 release, I want to write some stupid git tutorial for\n> a complete beginner that has only used CVS before, with a real example of\n> how to use raw git, and along those lines I actually want the thing to\n> show how to do something useful.\n\nIs there actually much point in using raw git directly? You don't\nusually invoke the syscalls directly from the user programs either (and\nyou usually actually use stdio for the casual stuff). I guess the\nraw git usage can get quite long and tiresome sometimes.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"4283","messageId":"Pine.LNX.4.58.0505301748130.1876@ppc970.osdl.org","threadId":"770","inReplyTo":"20050530221214.GA29556@redhat.com","subject":"Re: I want to release a \"git-1.0\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-31T00:52:11Z","receivedAt":"2005-05-31T00:52:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 30 May 2005, Dave Jones wrote:\n>\n>     GIT_AUTHOR_NAME=\"John Doe\"      \\\n>     GIT_AUTHOR_EMAIL=\"jdoe@foo.com\" \\\n>     git-commit-tree `git-write-tree`    \\\n>     -p $(cat .git/HEAD )            \\\n>     < changelog.txt         \\\n>     > .git/HEAD\n\nYou _really_ want to script this.\n\nAlso, I'd seriously suggest you avoid using \".git/HEAD\" _and_ writing to \n.git/HEAD in the same command. Maybe it works, maybe it doesn't.\n\nSo script it with something like\n\n\t#!/bin/sh\n\texport GIT_AUTHOR_NAME=\"$1\"\n\texport GIT_AUTHOR_EMAIL=\"$2\"\n\ttree=$(git-write-tree) || exit 1\n\tcommit=$(git-commit-tree $tree -p HEAD) || exit 1\n\techo $commit > .git/HEAD\n\nand now you can just do\n\n\tcommit-as \"John Doe\" \"jdoe@foo.com\" < changelog.txt\n\nor something like that.\n\nThe git commands really are designed to be scripted.\n\n\t\tLinus\n"},{"id":"4284","messageId":"Pine.LNX.4.58.0505301752320.1876@ppc970.osdl.org","threadId":"770","inReplyTo":"20050530221922.GC21076@mythryan2.michonline.com","subject":"Re: I want to release a \"git-1.0\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-31T00:58:39Z","receivedAt":"2005-05-31T00:58:39Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 30 May 2005, Ryan Anderson wrote:\n> \n> Umm, why do you maintain two seperate \"git\" related trees?\n\nWell, my \"tools\" thing really isn't git proper, and may not make much \nsense in the git distribution.\n\nThat said, I'm actually moving things into git as they turn useful. For \nexample, I move the \"stripspace\" program into git (which means it got \nrenamed into \"git-stripspace\", since it ended up being useful for the \nstand-alone git-commit-scripts too. \n\nBut how many non-Linux projects really apply mailboxes of patches? It \ndoesn't seem to be very \"core\".\n\n> Why not merge all of git-tools in, in a tools/ subdirectory?\n\nI'll think about it. It does look like at least about half of the git\ntools end up being pretty core.\n\n> I've been meaning to ask the same question about \"gitweb\" for that\n> matter.\n\nWell, there the issue definitely boils down to \"different maintainers\". I \ndon't want to connect things that don't need to be connected. \n\n> I'd guess part of this is a holdover from the fact that you needed an\n> independent tree for BitKeeper, but does it still make sense?\n\nWell, I see the \"tools\" thing really as my private tools that may or may\nnot make sense for anybody else. Even the cvs2git thing is just so\n_stupid_, since I bet you can do it quite cleanly in perl without having\nthat strange \"convert cvsps output into a shellscript\" stage (admittedly,\nit was _really_ convenient for debugging to have that separate stage, so\nwhile it looks a bit hacky, it ended up being very powerful).\n\n\t\tLinus\n"},{"id":"4285","messageId":"Pine.LNX.4.58.0505301801520.1876@ppc970.osdl.org","threadId":"770","inReplyTo":"972477.0a6782ba1d3b9f05216ed520ef720fcf.ANY@taniwha.stupidest.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-31T01:06:25Z","receivedAt":"2005-05-31T01:06:25Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 30 May 2005, Chris Wedgwood wrote:\n> \n> I'm still at a loss how to do the equivalent of annotate.  I know a\n> couple of front ends can do this but I have no idea what command line\n> magic would be equivalent.\n\nThere isn't any. It's actually pretty nasty to do, following history \nbackwards and keeping track of lines as they are added. I know how, I'm \njust really lazy and hoping somebody else will do it, since I really end \nup not caring that much myself.\n\nI notice that Thomas Gleixner seems to have one, but that one is based on \na database, and doesn't look usable as a standalone command..\n\n\t\tLinus\n"},{"id":"4299","messageId":"m1psv7bjb6.fsf@ebiederm.dsl.xmission.com","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2005-05-31T13:45:33Z","receivedAt":"2005-05-31T13:45:33Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Ok, I'm at the point where I really think it's getting close to a 1.0, and\n> make another tar-ball etc. I obviously feel that it's already way superior\n> to CVS, but I also realize that somebody who is used to CVS may not \n> actually realize that very easily.\n\nI way behind the power curve on learning git at this point but\none piece of the puzzle that CVS has that I don't believe git does\nare multiple people committing to the same repository, especially\nremotely.  I don't see that as a down side of git but it is a common\nway people CVS so it is worth documenting.\n\nEric\n"},{"id":"4336","messageId":"7vu0kiu8pm.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301801520.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-01T02:11:49Z","receivedAt":"2005-06-01T02:11:49Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> On Mon, 30 May 2005, Chris Wedgwood wrote:\n>> \n>> I'm still at a loss how to do the equivalent of annotate.  I know a\n>> couple of front ends can do this but I have no idea what command line\n>> magic would be equivalent.\n\nLT> There isn't any. It's actually pretty nasty to do, following history \nLT> backwards and keeping track of lines as they are added. I know how, I'm \nLT> just really lazy and hoping somebody else will do it, since I really end \nLT> up not caring that much myself.\n\nLT> I notice that Thomas Gleixner seems to have one, but that one is based on \nLT> a database, and doesn't look usable as a standalone command..\n\nHere is my quick-and-dirty one done in Perl.  This is dog-slow\nand not suited for interactive use, but its algorithm should\nhandle the merges, renames and complete rewrites correctly.\n\nIts sample output for:\n\n    $ blame.perl HEAD git-commit-script\n\nlook like this (I've edited the SHA1 and names to make it a bit\nshorter, but it still does not fit on my 80-column terminal X-<).\n\nFor each line in the version in the HEAD, it outputs (TAB\nseparated) the SHA1 of the commit that is responsible for the\nline to be there, author, commiter, line number in the version\nof the guity commit and the filename in the guilty commit (this\nfile could have been renamed in which case this may not match\nthe name of the file the script was originally asked to\nannotate).  It shows that 9th line was what was in\na3e870f2... commit as 11th line done by Linus, for example.\n\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t1\tgit-commit-script\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t2\tgit-commit-script\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t3\tgit-commit-script\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t4\tgit-commit-script\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t5\tgit-commit-script\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t6\tgit-commit-script\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t7\tgit-commit-script\n:2036d841...\tJunio C....@cox.net>\tLinus T....osdl.org>\t8\tgit-commit-script\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t11\tgit-commit-script\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t12\tgit-commit-script\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t13\tgit-commit-script\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t14\tgit-commit-script\n:a3e870f2...\tLinus T....osdl.org>\tLinus T....osdl.org>\t15\tgit-commit-script\n\n------------\nA blame script for use by higher-level annotate tools.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\ndiff -u a/blame.perl b/blame.perl\n--- /dev/null\n+++ b/blame.perl\n@@ -0,0 +1,400 @@\n+#!/usr/bin/perl -w\n+\n+use strict;\n+\n+package main;\n+$::debug = 0;\n+\n+sub read_blob {\n+    my $sha1 = shift;\n+    my $fh = undef;\n+    my $result;\n+    local ($/) = undef;\n+    open $fh, '-|', 'git-cat-file', 'blob', $sha1\n+\tor die \"cannot read blob $sha1\";\n+    $result = join('', <$fh>);\n+    close $fh\n+\tor die \"failure while closing pipe to git-cat-file\";\n+    return $result;\n+}\n+\n+sub read_diff_raw {\n+    my ($parent, $filename) = @_;\n+    my $fh = undef;\n+    local ($/) = \"\\0\";\n+    my @result = (); \n+    my ($meta, $status, $sha1_1, $sha1_2, $file1, $file2);\n+    print STDERR \"* diff-cache $parent\\n\" if $::debug;\n+    open $fh, '-|', 'git-diff-cache', '-B', '-C', '--cached', '-z', $parent\n+\tor die \"cannot read git-diff-cache with $parent\";\n+    while (defined ($meta = <$fh>)) {\n+\tchomp($meta);\n+\t(undef, undef, $sha1_1, $sha1_2, $status) = split(/ /, $meta);\n+\t$file1 = <$fh>;\n+\tchomp($file1);\n+\tif ($status =~ /^[CR]/) {\n+\t    $file2 = <$fh>;\n+\t    chomp($file2);\n+\t} elsif ($status =~ /^D/) {\n+\t    next;\n+\t} else {\n+\t    $file2 = $file1;\n+\t}\n+\tif ($file2 eq $filename) {\n+\t    push @result, [$status, $sha1_1, $sha1_2, $file1, $file2];\n+\t}\n+    }\n+    close $fh\n+\tor die \"failure while closing pipe to git-diff-cache\";\n+    return @result;\n+}\n+\n+sub write_temp_blob {\n+    my ($sha1, $temp) = @_;\n+    my $fh = undef;\n+    my $blob = read_blob($sha1);\n+    open $fh, '>', $temp\n+\tor die \"cannot open temporary file $temp\";\n+    print $fh $blob;\n+    close($fh);\n+}\n+\n+package Git::Patch;\n+sub new {\n+    my ($class, $sha1_1, $sha1_2) = @_;\n+    my $self = bless [], $class;\n+    my $fh = undef;\n+    ::write_temp_blob($sha1_1, \"/tmp/blame-$$-1\");\n+    ::write_temp_blob($sha1_2, \"/tmp/blame-$$-2\");\n+    open $fh, '-|', 'diff', '-u0', \"/tmp/blame-$$-1\", \"/tmp/blame-$$-2\"\n+\tor die \"cannot read diff\";\n+    while (<$fh>) {\n+\tif (/^\\@\\@ -(\\d+)(?:,(\\d+))? \\+(\\d+)(?:,(\\d+))? \\@\\@/) {\n+\t    push @$self, [$1, (defined $2 ? $2 : 1),\n+\t\t\t  $3, (defined $4 ? $4 : 1)];\n+\t}\n+    }\n+    close $fh;\n+    unlink \"/tmp/blame-$$-1\", \"/tmp/blame-$$-2\";\n+    return $self;\n+}\n+\n+sub find_parent_line {\n+    my ($self, $commit_lineno) = @_;\n+    my $ofs = 0;\n+    for (@$self) {\n+\tmy ($line_1, $len_1, $line_2, $len_2) = @$_;\n+\tif ($commit_lineno < $line_2) {\n+\t    return $commit_lineno - $ofs;\n+\t}\n+\tif ($line_2 <= $commit_lineno && $commit_lineno < $line_2 + $len_2) {\n+\t    return -1; # changed by commit.\n+\t}\n+\t$ofs += ($len_1 - $len_2);\n+    }\n+    return $commit_lineno + $ofs;\n+}\n+\n+package Git::Commit;\n+sub new {\n+    my $class = shift;\n+    my $self = bless {\n+\tPARENT => [],\n+\tTREE => undef,\n+\tAUTHOR => undef,\n+\tCOMMITTER => undef,\n+    }, $class;\n+    my $commit_sha1 = shift;\n+    $self->{SHA1} = $commit_sha1;\n+    my $fh = undef;\n+    open $fh, '-|', 'git-cat-file', 'commit', $commit_sha1\n+\tor die \"cannot read commit object $commit_sha1\";\n+    while (<$fh>) {\n+\tchomp;\n+\tif (/^tree ([0-9a-f]{40})$/) { $self->{TREE} = $1; }\n+\telsif (/^parent ([0-9a-f]{40})$/) { push @{$self->{PARENT}}, $1; }\n+\telsif (/^author ([^>]+>)/) { $self->{AUTHOR} = $1; }\n+\telsif (/^committer ([^>]+>)/) { $self->{COMMITTER} = $1; }\n+    }\n+    close $fh\n+\tor die \"failure while closing pipe to git-cat-file\";\n+    return $self;\n+}\n+\n+sub find_file {\n+    my ($commit, $path) = @_;\n+    my $result = undef;\n+    my $fh = undef;\n+    local ($/) = \"\\0\";\n+    open $fh, '-|', 'git-ls-tree', '-z', '-r', '-d', $commit->{TREE}, $path\n+\tor die \"cannot read git-ls-tree $commit->{TREE}\";\n+    while (<$fh>) {\n+\tchomp;\n+\tif (/^[0-7]{6} blob ([0-9a-f]{40})\t(.*)$/) {\n+\t    if ($2 ne $path) {\n+\t\tdie \"$2 ne $path???\";\n+\t    }\n+\t    $result = $1;\n+\t    last;\n+\t}\n+    }\n+    close $fh\n+\tor die \"failure while closing pipe to git-ls-tree\";\n+    return $result;\n+}\n+\n+package Git::Blame;\n+sub new {\n+    my $class = shift;\n+    my $self = bless {\n+\tLINE => [],\n+\tUNKNOWN => undef,\n+\tWORK => [],\n+    }, $class;\n+    my $commit = shift;\n+    my $filename = shift;\n+    my $sha1 = $commit->find_file($filename);\n+    my $blob = ::read_blob($sha1);\n+    my @blob = (split(/\\n/, $blob));\n+    for (my $i = 0; $i < @blob; $i++) {\n+\t$self->{LINE}[$i] = +{\n+\t    COMMIT => $commit,\n+\t    FOUND => undef,\n+\t    FILENAME => $filename,\n+\t    LINENO => ($i + 1),\n+\t};\n+    }\n+    $self->{UNKNOWN} = scalar @blob;\n+    push @{$self->{WORK}}, [$commit, $filename];\n+    return $self;\n+}\n+\n+sub print {\n+    my $self = shift;\n+    my $line_termination = shift;\n+    for (my $i = 0; $i < @{$self->{LINE}}; $i++) {\n+\tmy $l = $self->{LINE}[$i];\n+\tprint ($l->{FOUND} ? ':' : '?');;\n+\tprint \"$l->{COMMIT}->{SHA1}\t\";\n+\tprint \"$l->{COMMIT}->{AUTHOR}\t\";\n+\tprint \"$l->{COMMIT}->{COMMITTER}\t\";\n+\tprint \"$l->{LINENO}\t$l->{FILENAME}\";\n+\tprint $line_termination;\n+    }\n+}\n+\n+sub take_responsibility {\n+    my ($self, $commit) = @_;\n+    for (my $i = 0; $i < @{$self->{LINE}}; $i++) {\n+\tmy $l = $self->{LINE}[$i];\n+\tif (! $l->{FOUND} && ($l->{COMMIT}->{SHA1} eq $commit->{SHA1})) {\n+\t    $l->{FOUND} = 1;\n+\t    $self->{UNKNOWN}--;\n+\t}\n+    }\n+}\n+\n+sub blame_parent {\n+    my ($self, $commit, $parent, $filename) = @_;\n+    my @diff = ::read_diff_raw($parent->{SHA1}, $filename);\n+    my $filename_in_parent;\n+    my $passed_blame_to_parent = undef;\n+    if (@diff == 0) {\n+\t# We have not touched anything.  Blame parent for everything\n+\t# that we are suspected for.\n+\tfor (my $i = 0; $i < @{$self->{LINE}}; $i++) {\n+\t    my $l = $self->{LINE}[$i];\n+\t    if (! $l->{FOUND} && ($l->{COMMIT}->{SHA1} eq $commit->{SHA1})) {\n+\t\t$l->{COMMIT} = $parent;\n+\t\t$passed_blame_to_parent = 1;\n+\t    }\n+\t}\n+\t$filename_in_parent = $filename;\n+    }\n+    elsif (@diff != 1) {\n+\t# This should not happen.\n+\tfor (@diff) {\n+\t    print \"** @$_\\n\";\n+\t}\n+\tdie \"Oops\";\n+    }\n+    else {\n+\tmy ($status, $sha1_1, $sha1_2, $file1, $file2) = @{$diff[0]};\n+\tprint STDERR \"** $status $file1 $file2\\n\" if $::debug;\n+\tif ($status =~ /N/) {\n+\t    # Either some of other parents created it, or we did.\n+\t    # At this point the only thing we know is that this\n+\t    # parent is not responsible for it.\n+\t    ;\n+\t}\n+\telse {\n+\t    my $patch = Git::Patch->new($sha1_1, $sha1_2);\n+\t    $filename_in_parent = $file1;\n+\t    for (my $i = 0; $i < @{$self->{LINE}}; $i++) {\n+\t\tmy $l = $self->{LINE}[$i];\n+\t\tif (! $l->{FOUND} && $l->{COMMIT}->{SHA1} eq $commit->{SHA1}) {\n+\t\t    # We are suspected to have introduced this line.\n+\t\t    # Does it exist in the parent?\n+\t\t    my $lineno = $l->{LINENO};\n+\t\t    my $parent_line = $patch->find_parent_line($lineno);\n+\t\t    if ($parent_line < 0) {\n+\t\t\t# No, we may be the guilty ones, or some other\n+\t\t\t# parent might be.  We do not assign blame to\n+\t\t\t# ourselves here yet.\n+\t\t\t;\n+\t\t    }\n+\t\t    else {\n+\t\t\t# This line is coming from the parent, so pass\n+\t\t\t# blame to it.\n+\t\t\t$l->{COMMIT} = $parent;\n+\t\t\t$l->{FILENAME} = $file1;\n+\t\t\t$l->{LINENO} = $parent_line;\n+\t\t\t$passed_blame_to_parent = 1;\n+\t\t    }\n+\t\t}\n+\t    }\n+\t}\n+    }\n+    if ($passed_blame_to_parent && $self->{UNKNOWN}) {\n+\tunshift @{$self->{WORK}},\n+\t[$parent, $filename_in_parent];\n+    }\n+}\n+\n+sub assign {\n+    my ($self, $commit, $filename) = @_;\n+    # We do read-tree of the current commit and diff-cache\n+    # with each parents, instead of running diff-tree.  This\n+    # is because diff-tree does not look for copies hard enough.\n+    #\n+    print STDERR \"* read-tree  $commit->{SHA1}\\n\" if $::debug;\n+    system('git-read-tree', '-m', $commit->{SHA1});\n+    for my $parent (@{$commit->{PARENT}}) {\n+\t$self->blame_parent($commit, Git::Commit->new($parent), $filename);\n+    }\n+    $self->take_responsibility($commit);\n+}\n+\n+sub assign_blame {\n+    my ($self) = @_;\n+    while ($self->{UNKNOWN} && @{$self->{WORK}}) {\n+\tmy $wk = shift @{$self->{WORK}};\n+\tmy ($commit, $filename) = @$wk;\n+\t$self->assign($commit, $filename);\n+    }\n+}\n+\n+\n+\n+################################################################\n+package main;\n+my $usage = \"blame [-z] <commit> filename\";\n+my $line_termination = \"\\n\";\n+\n+$::ENV{GIT_INDEX_FILE} = \"/tmp/blame-$$-index\";\n+unlink($::ENV{GIT_INDEX_FILE});\n+\n+if ($ARGV[0] eq '-z') {\n+    $line_termination = \"\\0\";\n+    shift;\n+}\n+\n+if (@ARGV != 2) {\n+    die $usage;\n+}\n+\n+my $head_commit = Git::Commit->new($ARGV[0]);\n+my $filename = $ARGV[1];\n+my $blame = Git::Blame->new($head_commit, $filename);\n+\n+$blame->assign_blame();\n+$blame->print($line_termination);\n+\n+unlink($::ENV{GIT_INDEX_FILE});\n+\n+__END__\n+\n+How does this work, and what do we do about merges?\n+\n+The algorithm considers that the first parent is our main line of\n+development and treats it somewhat special than other parents.  So we\n+pass on the blame to the first parent if a line has not changed from\n+it.  For lines that have changed from the first parent, we must have\n+either inherited that change from some other parent, or it could have\n+been merge conflict resolution edit we did on our own.\n+\n+The following picture illustrates how we pass on and assign blames.\n+\n+In the sample, the original O was forked into A and B and then merged\n+into M.  Line 1, 2, and 4 did not change.  Line 3 and 5 are changed in\n+A, and Line 5 and 6 are changed in B.  M made its own decision to\n+resolve merge conflicts at Line 5 to something different from A and B:\n+\n+                A: 1 2 T 4 T 6\n+               /               \\ \n+O: 1 2 3 4 5 6                  M: 1 2 T 4 M S\n+               \\               / \n+                B: 1 2 3 4 S S\n+\n+In the following picture, each line is annotated with a blame letter.\n+A lowercase blame (e.g. \"a\" for \"1\") means that commit or its ancestor\n+is the guilty party but we do not know which particular ancestor is\n+responsible for the change yet.  An uppercase blame means that we know\n+that commit is the guilty party.\n+\n+First we look at M (the HEAD) and initialize Git::Blame->{LINE} like\n+this:\n+\n+             M: 1 2 T 4 M S\n+                m m m m m m\n+\n+That is, we know all lines are results of modification made by some\n+ancestor of M, so we assign lowercase 'm' to all of them.\n+\n+Then we examine our first parent A.  Throughout the algorithm, we are\n+always only interested in the lines we are the suspect, but this being\n+the initial round, we are the suspect for all of them.  We notice that\n+1 2 T 4 are the same as the parent A, so we pass the blame for these\n+four lines to A.  M and S are different from A, so we leave them as\n+they are (note that we do not immediately take the blame for them):\n+\n+             M: 1 2 T 4 M S\n+                a a a a m m\n+\n+Next we go on to examine parent B.  Again, we are only interested in\n+the lines we are still the suspect (i.e. M and S).  We notice S is\n+something we inherited from B, so we pass the blame on to it, like\n+this:\n+\n+             M: 1 2 T 4 M S\n+                a a a a m b\n+\n+Once we exhausted the parents, we look at the results and take\n+responsibility for the remaining ones that we are still the suspect:\n+\n+             M: 1 2 T 4 M S\n+                a a a a M b\n+\n+We are done with M.  And we know commits A and B need to be examined\n+further, so we do them recursively.  When we look at A, we again only\n+look at the lines that A is the suspect:\n+\n+             A: 1 2 T 4 T 6\n+                a a a a M b\n+\n+Among 1 2 T 4, comparing against its parent O, we notice 1 2 4 are\n+the same so pass the blame for those lines to O:\n+\n+             A: 1 2 T 4 T 6\n+                o o a o M b\n+\n+A is a non-merge commit; we have already exhausted the parents and\n+take responsibility for the remaining ones that A is the suspect:\n+\n+             A: 1 2 T 4 T 6\n+                o o A o M b\n+\n+We go on like this and the final result would become:\n+\n+             O: 1 2 3 4 5 6\n+                O O A O M B\n\n"},{"id":"4337","messageId":"Pine.LNX.4.62.0505311923240.19864@qynat.qvtvafvgr.pbz","threadId":"770","inReplyTo":"7vu0kiu8pm.fsf@assigned-by-dhcp.cox.net","subject":"Re: I want to release a \"git-1.0\"","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2005-06-01T02:25:15Z","receivedAt":"2005-06-01T02:25:15Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Tue, 31 May 2005, Junio C Hamano wrote:\n\n>>>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n>\n> LT> On Mon, 30 May 2005, Chris Wedgwood wrote:\n>>>\n>>> I'm still at a loss how to do the equivalent of annotate.  I know a\n>>> couple of front ends can do this but I have no idea what command line\n>>> magic would be equivalent.\n>\n> LT> There isn't any. It's actually pretty nasty to do, following history\n> LT> backwards and keeping track of lines as they are added. I know how, I'm\n> LT> just really lazy and hoping somebody else will do it, since I really end\n> LT> up not caring that much myself.\n>\n> LT> I notice that Thomas Gleixner seems to have one, but that one is based on\n> LT> a database, and doesn't look usable as a standalone command..\n>\n> Here is my quick-and-dirty one done in Perl.  This is dog-slow\n> and not suited for interactive use, but its algorithm should\n> handle the merges, renames and complete rewrites correctly.\n\nHmm, thinking out loud. would it help to look at the deltify scripts and \nlet them find the major chunks and then look in detail only when that \nfails?\n\nDavid Lang\n"},{"id":"4338","messageId":"Pine.LNX.4.58.0505312002160.1876@ppc970.osdl.org","threadId":"770","inReplyTo":"m1psv7bjb6.fsf@ebiederm.dsl.xmission.com","subject":"Re: I want to release a \"git-1.0\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-01T03:04:11Z","receivedAt":"2005-06-01T03:04:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 31 May 2005, Eric W. Biederman wrote:\n> \n> I way behind the power curve on learning git at this point but\n> one piece of the puzzle that CVS has that I don't believe git does\n> are multiple people committing to the same repository, especially\n> remotely.  I don't see that as a down side of git but it is a common\n> way people CVS so it is worth documenting.\n\nIt's actually one thing git doesn't do per se.\n\nYou have to do a \"git-pull-script\" from the common repository side, \nthere's no \"git-push-script\". Ugly.\n\nAnyway, I wrote just a _very_ introductory thing in\nDocumentation/tutorial.txt, I'll try to update and expand on it later. It\nbasically has a really stupid example of \"how to set up a new project\".\n\n\t\tLinus\n"},{"id":"4341","messageId":"7vmzqau3es.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505312002160.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-01T04:06:19Z","receivedAt":"2005-06-01T04:06:19Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> Anyway, I wrote just a _very_ introductory thing in\nLT> Documentation/tutorial.txt, I'll try to update and expand on\nLT> it later. It basically has a really stupid example of \"how\nLT> to set up a new project\".\n\nI've spotted a couple of typos which I will leave others to fix,\nbut there is one thing I am to blame.\n\n    (Btw, current versions of git will consider the change in question to be\n    so big that it's considered a whole new file, since the diff is actually\n    bigger than the file.  So the helpful comments that git-commit-script\n    tells you for this example will say that you deleted and re-created the\n    file \"a\".  For a less contrieved example, these things are usually more\n    obvious). \n\nDo you want me to do something about this with -B (and possibly\n-C/-M), like skipping the comparison altogether if the file size\nis smaller than, say, 1k bytes or something silly like that?  Or\nnot having special case for this kind of \"contrived example\"\npreferrable?\n\n"},{"id":"4344","messageId":"7vhdgismo0.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.62.0505311923240.19864@qynat.qvtvafvgr.pbz","subject":"Re: I want to release a \"git-1.0\"","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-01T04:53:19Z","receivedAt":"2005-06-01T04:53:19Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"DL\" == David Lang <david.lang@digitalinsight.com> writes:\n\nDL> Hmm, thinking out loud. would it help to look at the deltify scripts\nDL> and let them find the major chunks and then look in detail only when\nDL> that fails?\n\nIt's unclear to me which part you are trying to help with\ndeltify algorithm [*1*].\n\nInternally, git-diff-cache -B -C is used which does use the\ndeltify to locate complete rewrites, renames and copies (that's\nwhy the script is so slow).  For passing on and assigning blames\nline by line, parsing \"diff --unified=0\" output was a lot easier\nfor this script and that was what I did in this quick-and-dirty\nversion.\n\n\n[Footnotes]\n\n*1* David says \"deltify\" and Nico calls it \"deltafy\".  I am not\na native speaker so I cannot tell, but which one is correct?\n\n"},{"id":"4352","messageId":"7vsm02pp3s.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505312002160.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-01T06:28:55Z","receivedAt":"2005-06-01T06:28:55Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> Anyway, I wrote just a _very_ introductory thing in\nLT> Documentation/tutorial.txt, I'll try to update and expand on it later. It\nLT> basically has a really stupid example of \"how to set up a new project\".\n\nLinus,\n\n        I was following your \"tutorial\" and saw the last step\n(git-whatchanged) showing the HEAD commit and diff _twice_.\n\nYou got me _WORRIED_!!!\n\nI knew it uses your faviorite diff-tree command and I was the\nmost likely suspect who broke it.  And I remember you were\nunderstandably unhappy last time I broke it (the \"diff-tree -s\"\nproblem).\n\nIt turns out that the example in the tutorial was bad.  Here is\na fix.  It is so obvious that I do not think it deserves a\nsign-off nor credit.  Please just fold it into your edit next\ntime you update the tutorial.\n\n---\ncd /opt/packrat/playpen/public/in-place/git/git.junio/\njit-diff : Documentation\n# - linus: git-apply --stat: limit lines to 79 characters\n# + (working tree)\ndiff --git a/Documentation/tutorial.txt b/Documentation/tutorial.txt\n--- a/Documentation/tutorial.txt\n+++ b/Documentation/tutorial.txt\n@@ -401,7 +401,7 @@ activity.\n To see the whole history of our pitiful little git-tutorial project, we\n can do\n \n-\tgit-whatchanged -p --root HEAD\n+\tgit-whatchanged -p --root\n \n (the \"--root\" flag is a flag to git-diff-tree to tell it to show the\n initial aka \"root\" commit as a diff too), and you will see exactly what\n\nCompilation finished at Tue May 31 23:12:32\n\n\n"},{"id":"4354","messageId":"7vacmapo18.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.62.0505301644430.5330@localhost.localdomain","subject":"Re: I want to release a \"git-1.0\"","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-01T06:52:03Z","receivedAt":"2005-06-01T06:52:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"I just remembered that I mentioned potential problems with non\nrsync pulls with delta objects, especially when the git-*-pull\ncommands are used in \"things only close to the tip\" mode,\ni.e. without \"-a\" option.  Do you think we should do something\nabout it before GIT 1.0 happens?  \n\nIt may be enough if we just tell people not to deltify their\npublic non-rsync repositories in the documentation.\n\n"},{"id":"4356","messageId":"7vu0kilc21.fsf_-_@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"7vacmapo18.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH] Add -d flag to git-pull-* family.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-01T08:24:22Z","receivedAt":"2005-06-01T08:24:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"When a remote repository is deltified, we need to get the\nobjects that a deltified object we want to obtain is based upon.\nSince checking representation type of all objects we retreive\nfrom remote side may be costly, this is made into a separate\noption -d; -a implies it for convenience and safety.\n\nRsync transport does not have this problem since it fetches\neverything the remote side has.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\nDocumentation/git-http-pull.txt  |    4 +++-\nDocumentation/git-local-pull.txt |    4 +++-\nDocumentation/git-rpull.txt      |    4 +++-\nhttp-pull.c                      |    5 ++++-\nlocal-pull.c                     |    5 ++++-\npull.c                           |   15 +++++++++++++++\npull.h                           |    3 +++\nrpull.c                          |    5 ++++-\n8 files changed, 39 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/git-http-pull.txt b/Documentation/git-http-pull.txt\n--- a/Documentation/git-http-pull.txt\n+++ b/Documentation/git-http-pull.txt\n@@ -9,7 +9,7 @@ git-http-pull - Downloads a remote GIT r\n \n SYNOPSIS\n --------\n-'git-http-pull' [-c] [-t] [-a] [-v] commit-id url\n+'git-http-pull' [-c] [-t] [-a] [-v] [-d] commit-id url\n \n DESCRIPTION\n -----------\n@@ -17,6 +17,8 @@ Downloads a remote GIT repository via HT\n \n -c::\n \tGet the commit objects.\n+-d::\n+\tGet objects that deltified objects are based upon.\n -t::\n \tGet trees associated with the commit objects.\n -a::\ndiff --git a/Documentation/git-local-pull.txt b/Documentation/git-local-pull.txt\n--- a/Documentation/git-local-pull.txt\n+++ b/Documentation/git-local-pull.txt\n@@ -9,7 +9,7 @@ git-local-pull - Duplicates another GIT \n \n SYNOPSIS\n --------\n-'git-local-pull' [-c] [-t] [-a] [-l] [-s] [-n] [-v] commit-id path\n+'git-local-pull' [-c] [-t] [-a] [-l] [-s] [-n] [-v] [-d] commit-id path\n \n DESCRIPTION\n -----------\n@@ -19,6 +19,8 @@ OPTIONS\n -------\n -c::\n \tGet the commit objects.\n+-d::\n+\tGet objects that deltified objects are based upon.\n -t::\n \tGet trees associated with the commit objects.\n -a::\ndiff --git a/Documentation/git-rpull.txt b/Documentation/git-rpull.txt\n--- a/Documentation/git-rpull.txt\n+++ b/Documentation/git-rpull.txt\n@@ -10,7 +10,7 @@ git-rpull - Pulls from a remote reposito\n \n SYNOPSIS\n --------\n-'git-rpull' [-c] [-t] [-a] [-v] commit-id url\n+'git-rpull' [-c] [-t] [-a] [-v] [-d] commit-id url\n \n DESCRIPTION\n -----------\n@@ -21,6 +21,8 @@ OPTIONS\n -------\n -c::\n \tGet the commit objects.\n+-d::\n+\tGet objects that deltified objects are based upon.\n -t::\n \tGet trees associated with the commit objects.\n -a::\ndiff --git a/http-pull.c b/http-pull.c\n--- a/http-pull.c\n+++ b/http-pull.c\n@@ -103,17 +103,20 @@ int main(int argc, char **argv)\n \t\t\tget_tree = 1;\n \t\t} else if (argv[arg][1] == 'c') {\n \t\t\tget_history = 1;\n+\t\t} else if (argv[arg][1] == 'd') {\n+\t\t\tget_delta = 1;\n \t\t} else if (argv[arg][1] == 'a') {\n \t\t\tget_all = 1;\n \t\t\tget_tree = 1;\n \t\t\tget_history = 1;\n+\t\t\tget_delta = 1;\n \t\t} else if (argv[arg][1] == 'v') {\n \t\t\tget_verbosely = 1;\n \t\t}\n \t\targ++;\n \t}\n \tif (argc < arg + 2) {\n-\t\tusage(\"git-http-pull [-c] [-t] [-a] [-v] commit-id url\");\n+\t\tusage(\"git-http-pull [-c] [-t] [-a] [-d] [-v] commit-id url\");\n \t\treturn 1;\n \t}\n \tcommit_id = argv[arg];\ndiff --git a/local-pull.c b/local-pull.c\n--- a/local-pull.c\n+++ b/local-pull.c\n@@ -74,7 +74,7 @@ int fetch(unsigned char *sha1)\n }\n \n static const char *local_pull_usage = \n-\"git-local-pull [-c] [-t] [-a] [-l] [-s] [-n] [-v] commit-id path\";\n+\"git-local-pull [-c] [-t] [-a] [-l] [-s] [-n] [-v] [-d] commit-id path\";\n \n /* \n  * By default we only use file copy.\n@@ -92,10 +92,13 @@ int main(int argc, char **argv)\n \t\t\tget_tree = 1;\n \t\telse if (argv[arg][1] == 'c')\n \t\t\tget_history = 1;\n+\t\telse if (argv[arg][1] == 'd')\n+\t\t\tget_delta = 1;\n \t\telse if (argv[arg][1] == 'a') {\n \t\t\tget_all = 1;\n \t\t\tget_tree = 1;\n \t\t\tget_history = 1;\n+\t\t\tget_delta = 1;\n \t\t}\n \t\telse if (argv[arg][1] == 'l')\n \t\t\tuse_link = 1;\ndiff --git a/pull.c b/pull.c\n--- a/pull.c\n+++ b/pull.c\n@@ -6,6 +6,7 @@\n \n int get_tree = 0;\n int get_history = 0;\n+int get_delta = 0;\n int get_all = 0;\n int get_verbosely = 0;\n static unsigned char current_commit_sha1[20];\n@@ -37,6 +38,20 @@ static int make_sure_we_have_it(const ch\n \tstatus = fetch(sha1);\n \tif (status && what)\n \t\treport_missing(what, sha1);\n+\tif (get_delta) {\n+\t\tunsigned long mapsize, size;\n+\t\tvoid *map, *buf;\n+\t\tchar type[20];\n+\n+\t\tmap = map_sha1_file(sha1, &mapsize);\n+\t\tif (map) {\n+\t\t\tbuf = unpack_sha1_file(map, mapsize, type, &size);\n+\t\t\tmunmap(map, mapsize);\n+\t\t\tif (buf && !strcmp(type, \"delta\"))\n+\t\t\t\tstatus = make_sure_we_have_it(what, buf);\n+\t\t\tfree(buf);\n+\t\t}\n+\t}\n \treturn status;\n }\n \ndiff --git a/pull.h b/pull.h\n--- a/pull.h\n+++ b/pull.h\n@@ -13,6 +13,9 @@ extern int get_history;\n /** Set to fetch the trees in the commit history. **/\n extern int get_all;\n \n+/* Set to fetch the base of delta objects.*/\n+extern int get_delta;\n+\n /* Set to be verbose */\n extern int get_verbosely;\n \ndiff --git a/rpull.c b/rpull.c\n--- a/rpull.c\n+++ b/rpull.c\n@@ -27,17 +27,20 @@ int main(int argc, char **argv)\n \t\t\tget_tree = 1;\n \t\t} else if (argv[arg][1] == 'c') {\n \t\t\tget_history = 1;\n+\t\t} else if (argv[arg][1] == 'd') {\n+\t\t\tget_delta = 1;\n \t\t} else if (argv[arg][1] == 'a') {\n \t\t\tget_all = 1;\n \t\t\tget_tree = 1;\n \t\t\tget_history = 1;\n+\t\t\tget_delta = 1;\n \t\t} else if (argv[arg][1] == 'v') {\n \t\t\tget_verbosely = 1;\n \t\t}\n \t\targ++;\n \t}\n \tif (argc < arg + 2) {\n-\t\tusage(\"git-rpull [-c] [-t] [-a] [-v] commit-id url\");\n+\t\tusage(\"git-rpull [-c] [-t] [-a] [-v] [-d] commit-id url\");\n \t\treturn 1;\n \t}\n \tcommit_id = argv[arg];\n------------------------------------------------\n\n"},{"id":"4367","messageId":"Pine.LNX.4.63.0506011032570.7439@localhost.localdomain","threadId":"770","inReplyTo":"7vu0kilc21.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Add -d flag to git-pull-* family.","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2005-06-01T14:39:25Z","receivedAt":"2005-06-01T14:39:25Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 1 Jun 2005, Junio C Hamano wrote:\n\n> When a remote repository is deltified, we need to get the\n> objects that a deltified object we want to obtain is based upon.\n> Since checking representation type of all objects we retreive\n> from remote side may be costly, this is made into a separate\n> option -d; -a implies it for convenience and safety.\n\nI wonder if making this optional makes sense.  In fact, if you believe \nhaving the option is useful then it should probably be the other \nway around i.e. to _not_ look at deltas when it is specified.  Otherwise \nyou'll end up with an incoherent repository.\n\nTo minimize the cost a lot it could be possible to uncompress just the \nfirst 40 bytes or so which is enough to determine if the object is a \ndelta and if so what object it is against.\n\nWhat do you think?\n\n\nNicolas\n"},{"id":"4369","messageId":"7vpsv6kqx0.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.63.0506011032570.7439@localhost.localdomain","subject":"Re: [PATCH] Add -d flag to git-pull-* family.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-01T16:00:59Z","receivedAt":"2005-06-01T16:00:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"NP\" == Nicolas Pitre <nico@cam.org> writes:\n\nNP> What do you think?\n\nWhat you say makes a lot more sense than my quick hack on both\ncounts.\n\n"},{"id":"4384","messageId":"Pine.LNX.4.62.0506011304001.21267@qynat.qvtvafvgr.pbz","threadId":"770","inReplyTo":"7vhdgismo0.fsf@assigned-by-dhcp.cox.net","subject":"Re: I want to release a \"git-1.0\"","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2005-06-01T20:06:31Z","receivedAt":"2005-06-01T20:06:31Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Tue, 31 May 2005, Junio C Hamano wrote:\n\n>>>>>> \"DL\" == David Lang <david.lang@digitalinsight.com> writes:\n>\n> DL> Hmm, thinking out loud. would it help to look at the deltify scripts\n> DL> and let them find the major chunks and then look in detail only when\n> DL> that fails?\n>\n> It's unclear to me which part you are trying to help with\n> deltify algorithm [*1*].\n\nI was thinking that the speedups (only look for similar sized files, etc) \nwould help narrow the search. Also each chunk that's different should be \nable to be able to be annotated as a chunk, instead of by individual line\n\n> Internally, git-diff-cache -B -C is used which does use the\n> deltify to locate complete rewrites, renames and copies (that's\n> why the script is so slow).  For passing on and assigning blames\n> line by line, parsing \"diff --unified=0\" output was a lot easier\n> for this script and that was what I did in this quick-and-dirty\n> version.\n\nI was under the impressin that the deltafy stuff was significantly faster \nthen you are suggeting that it is here\n\n> [Footnotes]\n>\n> *1* David says \"deltify\" and Nico calls it \"deltafy\".  I am not\n> a native speaker so I cannot tell, but which one is correct?\n\nNico is correct\n\nDavid Lang\n\n-- \nThere are two ways of constructing a software design. One way is to make it so simple that there are obviously no deficiencies. And the other way is to make it so complicated that there are no obvious deficiencies.\n  -- C.A.R. Hoare\n"},{"id":"4385","messageId":"Pine.LNX.4.61.0506011607480.11264@cag.csail.mit.edu","threadId":"770","inReplyTo":"Pine.LNX.4.62.0506011304001.21267@qynat.qvtvafvgr.pbz","subject":"Re: I want to release a \"git-1.0\"","fromName":"C. Scott Ananian","fromEmail":"cscott@cscott.net","sentAt":"2005-06-01T20:16:05Z","receivedAt":"2005-06-01T20:16:05Z","isPatch":false,"sender":{"key":"cscott@cscott.net","avatar":"https://gravatar.com/avatar/3551c2aefb299a0c45807f7677f5b26d8a5be4a4af359b4bf4fabbdd1f2b990e?d=mp&s=160"},"body":"On Wed, 1 Jun 2005, David Lang wrote:\n\n>> *1* David says \"deltify\" and Nico calls it \"deltafy\".  I am not\n>> a native speaker so I cannot tell, but which one is correct?\n>\n> Nico is correct\n\nAu contraire.  The common *pronunciation* may be 'delta-fy', but the \ncorrect spelling should be 'deltify'.  The google oracle agrees (1,440 vs \n54) as does the spelling of the svnadmin command.  (Of course, what google \nis really measuring is relative frequency of 'git' vs 'svn'.)\n\n$ grep '[^if]fy$' /usr/dict/american-english-large\n\nshows that the only vowels other than 'i' which preced the '-fy' morpheme \nare 'e's, and they only appear in words like 'liquefy' where the root has \nbeen substantially altered.  Most sources (eg\nhttp://www.southampton.liunet.edu/academic/pau/course/websuf.htm#IFYVERB\n) list the morpheme as '-ify'.  See\n     http://m-w.com/cgi-bin/dictionary?book=Dictionary&va=ify\nand compare\n     http://m-w.com/cgi-bin/dictionary?book=Dictionary&va=fy\n\nContrary to David's assertion, David is right.\n  --scott\n\nUnited Nations KMPLEBE AMTHUG AVBRANDY UNIFRUIT chemical agent tonight \nZPSEMANTIC ODYOKE struggle PBCABOOSE FJDEFLECT CLOWER MKSEARCH ZRBRIEF\n                          ( http://cscott.net/ )\n"},{"id":"4386","messageId":"Pine.LNX.4.21.0506011742560.30848-100000@iabervon.org","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505312002160.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-01T22:00:55Z","receivedAt":"2005-06-01T22:00:55Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Tue, 31 May 2005, Linus Torvalds wrote:\n\n> On Tue, 31 May 2005, Eric W. Biederman wrote:\n> > \n> > I way behind the power curve on learning git at this point but\n> > one piece of the puzzle that CVS has that I don't believe git does\n> > are multiple people committing to the same repository, especially\n> > remotely.  I don't see that as a down side of git but it is a common\n> > way people CVS so it is worth documenting.\n> \n> It's actually one thing git doesn't do per se.\n> \n> You have to do a \"git-pull-script\" from the common repository side, \n> there's no \"git-push-script\". Ugly.\n\nIt shouldn't be hard to do one, except that locking with rsync is going to\nbe a pain. I had a patch to make it work with the rpush/rpull pair, but I\ndidn't get its dependancies in at the time. I can dust those patches off\nagain if you want that functionality included.\n\nThe patches are essentially:\n\n - make the transport protocol handle things other than objects\n - library procedure for locking atomic update of refs files\n - fetching refs in general\n - rpull/rpush that updates a specified ref file atomically\n\nAt least the first would be very nice to get in before 1.0, since it is an\nincompatible change to the protocol.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"4387","messageId":"7vfyw1isrl.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.62.0506011304001.21267@qynat.qvtvafvgr.pbz","subject":"Re: I want to release a \"git-1.0\"","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-01T23:03:58Z","receivedAt":"2005-06-01T23:03:58Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"DL\" == David Lang <david.lang@digitalinsight.com> writes:\n\n>> Internally, git-diff-cache -B -C is used which does use the\n>> deltify to locate complete rewrites, renames and copies (that's\n>> why the script is so slow).  For passing on and assigning blames\n>> line by line, parsing \"diff --unified=0\" output was a lot easier\n>> for this script and that was what I did in this quick-and-dirty\n>> version.\n\nDL> I was under the impressin that the deltafy stuff was significantly\nDL> faster then you are suggeting that it is here\n\nI perhaps phrased it poorly.\n\nThe slow part is not a single delta operation, but having to run\nmany delta operations between all combinations of rename/copy\ncandidates, which is O(n * m) where n is the number of newly\ncreated files (counting \"broken\" ones created by -B flag) and m\nis the number of (deleted, modified and unmodified) files in the\noriginal tree.\n\n\n"},{"id":"4388","messageId":"7v7jhdisoc.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.21.0506011742560.30848-100000@iabervon.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-01T23:05:55Z","receivedAt":"2005-06-01T23:05:55Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"DB\" == Daniel Barkalow <barkalow@iabervon.org> writes:\n\nDB> It shouldn't be hard to do one, except that locking with\nDB> rsync is going to be a pain. I had a patch to make it work\nDB> with the rpush/rpull pair, but I didn't get its dependancies\nDB> in at the time. I can dust those patches off again if you\nDB> want that functionality included.\n\nTalking about pulls, wouldn't it be nicer to (re)name it to\ngit-ssh-pull, for consistency with others, especially before we\nhit 1.0?\n\n\n"},{"id":"4389","messageId":"Pine.LNX.4.63.0506012040310.17354@localhost.localdomain","threadId":"770","inReplyTo":"Pine.LNX.4.61.0506011607480.11264@cag.csail.mit.edu","subject":"Re: I want to release a \"git-1.0\"","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2005-06-02T00:43:08Z","receivedAt":"2005-06-02T00:43:08Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 1 Jun 2005, C. Scott Ananian wrote:\n\n> On Wed, 1 Jun 2005, David Lang wrote:\n> \n> > > *1* David says \"deltify\" and Nico calls it \"deltafy\".  I am not\n> > > a native speaker so I cannot tell, but which one is correct?\n> > \n> > Nico is correct\n> \n> Au contraire.  The common *pronunciation* may be 'delta-fy', but the correct\n> spelling should be 'deltify'.\n\nAinsi soit-il alors.\n\nI'm not a native english speaker either so I defer to anyone with better \nenglish knowledge.\n\n\nNicolas\n"},{"id":"4390","messageId":"Pine.LNX.4.63.0506012045190.17354@localhost.localdomain","threadId":"770","inReplyTo":"7v1x7lk8fl.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Handle deltified object correctly in git-*-pull family.","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2005-06-02T00:47:44Z","receivedAt":"2005-06-02T00:47:44Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 1 Jun 2005, Junio C Hamano wrote:\n\n> *** Dan and Nico, could you check this for correctness?  I've\n> *** tested it with a deltified core GIT repository and pulling\n> *** with local-pull from there.  I have verified that a pull\n> *** that fails with -d flag retrieves the right base-object to\n> *** complete a deltified ones.\n\nThe delta part looks fine to me.\n\n\nNicolas\n"},{"id":"4391","messageId":"Pine.LNX.4.63.0506012048330.17354@localhost.localdomain","threadId":"770","inReplyTo":"7vpsv5hbm5.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Stop inflating the whole SHA1 file only to check size.","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2005-06-02T00:51:28Z","receivedAt":"2005-06-02T00:51:28Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 1 Jun 2005, Junio C Hamano wrote:\n\n> Using the new unpack_sha1_file_partial() function, stop\n> inflating the whole SHA1 file when rename detector wants to know\n> only the filesize.\n\nBeware.  If you have delta objects you'll get the size of the delta \nitself and not the final object size, unless you recurse until a non \ndelta object is found.\n\n\nNicolas\n"},{"id":"4392","messageId":"Pine.LNX.4.58.0506011738080.1876@ppc970.osdl.org","threadId":"770","inReplyTo":"7v1x7lk8fl.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Handle deltified object correctly in git-*-pull family.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-02T00:58:33Z","receivedAt":"2005-06-02T00:58:33Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 1 Jun 2005, Junio C Hamano wrote:\n> \n> *** Linus, I have a hook in sha1_file.c to let me figure out the\n> *** size of the SHA1 file without fully expanding it.  This\n> *** patch does not use it, but you already know where I am\n> *** heading, so please leave it there ;-). \n\nArgh. This is just adding conceptual complexity without any real \nadvantage.\n\nWhy not just split out the current \"unpack_sha1_file()\" into two stages: \n\"unpack_sha1_header()\" and the rest.\n\nThen you can just decide to call \"unpack_sha1_header()\" when you want the \nheader information.\n\nHmm. I just committed something like that. If you want to just see the \ntype of an object, you can map the object in memory, and just do\n\n\tz_stream stream;\n\tchar buffer[100];\n\n\tif (unpack_sha1_header(&stream, map, mapsize, buffer, sizeof(buffer) < 0)\n\t\treturn NULL;\n\tif (sscanf(buffer, %10s %lu\", type, size) != 0)\n\t\treturn NULL;\n\t.. there you have it ..\n\nwhich is a lot simpler than worrying about callbacks etc.\n\n\t\tLinus\n"},{"id":"4393","messageId":"429E5D6B.8030406@khandalf.com","threadId":"770","inReplyTo":"Pine.LNX.4.63.0506012040310.17354@localhost.localdomain","subject":"Re: I want to release a \"git-1.0\"","fromName":"Brian O'Mahoney","fromEmail":"omb@khandalf.com","sentAt":"2005-06-02T01:14:19Z","receivedAt":"2005-06-02T01:14:19Z","isPatch":false,"sender":{"key":"omb@khandalf.com","avatar":null},"body":"Neither are _correct_, both are slang and a new word, but\nEnglish is good at that,\n\nrepresent via a delta, would be traditional,\n\nbut 'deltify' sounds nicer.\n\nNicolas Pitre wrote:\n> On Wed, 1 Jun 2005, C. Scott Ananian wrote:\n> \n> \n>>On Wed, 1 Jun 2005, David Lang wrote:\n>>\n>>\n>>>>*1* David says \"deltify\" and Nico calls it \"deltafy\".  I am not\n>>>>a native speaker so I cannot tell, but which one is correct?\n>>>\n>>>Nico is correct\n>>\n>>Au contraire.  The common *pronunciation* may be 'delta-fy', but the correct\n>>spelling should be 'deltify'.\n> \n> \n> Ainsi soit-il alors.\n> \n> I'm not a native english speaker either so I defer to anyone with better \n> english knowledge.\n> \n> \n> Nicolas\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n> \n> \n\n-- \nmit freundlichen Grüßen, Brian.\n\nDr. Brian O'Mahoney\nMobile +41 (0)79 334 8035 Email: omb@bluewin.ch\nBleicherstrasse 25, CH-8953 Dietikon, Switzerland\nPGP Key fingerprint = 33 41 A2 DE 35 7C CE 5D  F5 14 39 C9 6D 38 56 D5\n"},{"id":"4394","messageId":"7vll5th7cb.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.63.0506012048330.17354@localhost.localdomain","subject":"Re: [PATCH] Stop inflating the whole SHA1 file only to check size.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-02T01:32:04Z","receivedAt":"2005-06-02T01:32:04Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"NP\" == Nicolas Pitre <nico@cam.org> writes:\n\nNP> On Wed, 1 Jun 2005, Junio C Hamano wrote:\n>> Using the new unpack_sha1_file_partial() function, stop\n>> inflating the whole SHA1 file when rename detector wants to know\n>> only the filesize.\n\nNP> Beware.\n\nYou are right.  I cannot believe how stupid I am, falling into\nthis trap _just_ _after_ looking at the delta stuff X-<.\n\nLinus please drop that one.\n\n\n\n"},{"id":"4395","messageId":"7vhdghh6sh.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.58.0506011738080.1876@ppc970.osdl.org","subject":"Re: [PATCH] Handle deltified object correctly in git-*-pull family.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-02T01:43:58Z","receivedAt":"2005-06-02T01:43:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> Why not just split out the current \"unpack_sha1_file()\" into two stages: \nLT> \"unpack_sha1_header()\" and the rest.\nLT> which is a lot simpler than worrying about callbacks etc.\n\nAlright.\n\n\n"},{"id":"4401","messageId":"m18y1tb55c.fsf@ebiederm.dsl.xmission.com","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505312002160.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2005-06-02T07:15:59Z","receivedAt":"2005-06-02T07:15:59Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Anyway, I wrote just a _very_ introductory thing in\n> Documentation/tutorial.txt, I'll try to update and expand on it later. It\n> basically has a really stupid example of \"how to set up a new project\".\n\nSo I need to do a git checkout of the latest version of git to\nread the tutorial?  So I can figure out how to use git?\n\nCatch 22? :)\n\nEric\n"},{"id":"4404","messageId":"20050602083239.GA10775@vrfy.org","threadId":"770","inReplyTo":"m18y1tb55c.fsf@ebiederm.dsl.xmission.com","subject":"Re: I want to release a \"git-1.0\"","fromName":"Kay Sievers","fromEmail":"kay.sievers@vrfy.org","sentAt":"2005-06-02T08:32:39Z","receivedAt":"2005-06-02T08:32:39Z","isPatch":false,"sender":{"key":"kay.sievers@vrfy.org","avatar":null},"body":"On Thu, Jun 02, 2005 at 01:15:59AM -0600, Eric W. Biederman wrote:\n> Linus Torvalds <torvalds@osdl.org> writes:\n> \n> > Anyway, I wrote just a _very_ introductory thing in\n> > Documentation/tutorial.txt, I'll try to update and expand on it later. It\n> > basically has a really stupid example of \"how to set up a new project\".\n> \n> So I need to do a git checkout of the latest version of git to\n> read the tutorial?  So I can figure out how to use git?\n\nNo problem: :)\n  http://www.kernel.org/git/?p=git/git.git;a=blob;f=Documentation/tutorial.txt\n\nKay\n"},{"id":"4408","messageId":"200506021602.07258.snake@penza-gsm.ru","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505312002160.1876@ppc970.osdl.org","subject":"[PATCH] several typos in tutorial","fromName":"Alexey Nezhdanov","fromEmail":"snake@penza-gsm.ru","sentAt":"2005-06-02T12:02:07Z","receivedAt":"2005-06-02T12:02:07Z","isPatch":true,"sender":{"key":"snake@penza-gsm.ru","avatar":null},"body":"Signed-off-by: Alexey Nezhdanov <snake@penza-gsm.ru>\n---\ndiff --git a/Documentation/tutorial.txt b/Documentation/tutorial.txt\n--- a/Documentation/tutorial.txt\n+++ b/Documentation/tutorial.txt\n@@ -298,7 +298,7 @@ have committed something, we can also le\n \n Unlike \"git-diff-files\", which showed the difference between the index\n file and the working directory, \"git-diff-cache\" shows the differences\n-between a committed _tree_ and the index file.  In other words,\n+between a committed _tree_ and the working directory.  In other words,\n git-diff-cache wants a tree to be diffed against, and before we did the\n commit, we couldn't do that, because we didn't have anything to diff\n against. \n@@ -423,8 +423,8 @@ With that, you should now be having some\n can explore on your own.\n \n \n-\tCopoying archives\n-\t-----------------\n+\tCopying archives\n+\t----------------\n \n Git arhives are normally totally self-sufficient, and it's worth noting\n that unlike CVS, for example, there is no separate notion of\n\n"},{"id":"4409","messageId":"20050602124143.GA31483@snarc.org","threadId":"770","inReplyTo":"200506021602.07258.snake@penza-gsm.ru","subject":"Re: [PATCH] several typos in tutorial","fromName":"Vincent Hanquez","fromEmail":"tab@snarc.org","sentAt":"2005-06-02T12:41:43Z","receivedAt":"2005-06-02T12:41:43Z","isPatch":true,"sender":{"key":"tab@snarc.org","avatar":"https://gravatar.com/avatar/f639df78e7af804fd54c86e32a2b7b1853d3765ba8e66f878e24c933efc30a2a?d=mp&s=160"},"body":"On Thu, Jun 02, 2005 at 04:02:07PM +0400, Alexey Nezhdanov wrote:\n>  Git arhives are normally totally self-sufficient, and it's worth noting\n       ^^^^^^^\nand one more here\n\n-- \nVincent Hanquez\n"},{"id":"4410","messageId":"200506021645.15247.snake@penza-gsm.ru","threadId":"770","inReplyTo":"20050602124143.GA31483@snarc.org","subject":"Re: [PATCH] several typos in tutorial","fromName":"Alexey Nezhdanov","fromEmail":"snake@penza-gsm.ru","sentAt":"2005-06-02T12:45:15Z","receivedAt":"2005-06-02T12:45:15Z","isPatch":true,"sender":{"key":"snake@penza-gsm.ru","avatar":null},"body":"On thursday, 02 June 2005 16:41 Vincent Hanquez wrote:\n> On Thu, Jun 02, 2005 at 04:02:07PM +0400, Alexey Nezhdanov wrote:\n> >  Git arhives are normally totally self-sufficient, and it's worth noting\n>\n>        ^^^^^^^\n> and one more here\nWhy? It's ok to speak about many [existing] archives here.\n\n-- \nRespectfully\nAlexey Nezhdanov\n\n"},{"id":"4411","messageId":"20050602125124.GA31680@snarc.org","threadId":"770","inReplyTo":"200506021645.15247.snake@penza-gsm.ru","subject":"Re: [PATCH] several typos in tutorial","fromName":"Vincent Hanquez","fromEmail":"tab@snarc.org","sentAt":"2005-06-02T12:51:24Z","receivedAt":"2005-06-02T12:51:24Z","isPatch":true,"sender":{"key":"tab@snarc.org","avatar":"https://gravatar.com/avatar/f639df78e7af804fd54c86e32a2b7b1853d3765ba8e66f878e24c933efc30a2a?d=mp&s=160"},"body":"On Thu, Jun 02, 2005 at 04:45:15PM +0400, Alexey Nezhdanov wrote:\n> On thursday, 02 June 2005 16:41 Vincent Hanquez wrote:\n> > On Thu, Jun 02, 2005 at 04:02:07PM +0400, Alexey Nezhdanov wrote:\n> > >  Git arhives are normally totally self-sufficient, and it's worth noting\n> >\n> >        ^^^^^^^\n> > and one more here\n> Why? It's ok to speak about many [existing] archives here.\n\nit's missing a 'c'\n\n-- \nVincent Hanquez\n"},{"id":"4413","messageId":"200506021656.11180.snake@penza-gsm.ru","threadId":"770","inReplyTo":"20050602125124.GA31680@snarc.org","subject":"Re: [PATCH] several typos in tutorial","fromName":"Alexey Nezhdanov","fromEmail":"snake@penza-gsm.ru","sentAt":"2005-06-02T12:56:11Z","receivedAt":"2005-06-02T12:56:11Z","isPatch":true,"sender":{"key":"snake@penza-gsm.ru","avatar":null},"body":"Signed-off-by: Alexey Nezhdanov <snake@penza-gsm.ru>\n---\ndiff --git a/Documentation/tutorial.txt b/Documentation/tutorial.txt\n--- a/Documentation/tutorial.txt\n+++ b/Documentation/tutorial.txt\n@@ -298,7 +298,7 @@ have committed something, we can also le\n \n Unlike \"git-diff-files\", which showed the difference between the index\n file and the working directory, \"git-diff-cache\" shows the differences\n-between a committed _tree_ and the index file.  In other words,\n+between a committed _tree_ and the working directory.  In other words,\n git-diff-cache wants a tree to be diffed against, and before we did the\n commit, we couldn't do that, because we didn't have anything to diff\n against. \n@@ -423,10 +423,10 @@ With that, you should now be having some\n can explore on your own.\n \n \n-\tCopoying archives\n-\t-----------------\n+\tCopying archives\n+\t----------------\n \n-Git arhives are normally totally self-sufficient, and it's worth noting\n+Git archives are normally totally self-sufficient, and it's worth noting\n that unlike CVS, for example, there is no separate notion of\n \"repository\" and \"working tree\".  A git repository normally _is_ the\n working tree, with the local git information hidden in the \".git\"\n\n"},{"id":"4414","messageId":"200506021700.21824.snake@penza-gsm.ru","threadId":"770","inReplyTo":"20050602125124.GA31680@snarc.org","subject":"Re: [PATCH] several typos in tutorial","fromName":"Alexey Nezhdanov","fromEmail":"snake@penza-gsm.ru","sentAt":"2005-06-02T13:00:21Z","receivedAt":"2005-06-02T13:00:21Z","isPatch":true,"sender":{"key":"snake@penza-gsm.ru","avatar":null},"body":"On thursday, 02 June 2005 16:51 Vincent Hanquez wrote:\n> On Thu, Jun 02, 2005 at 04:45:15PM +0400, Alexey Nezhdanov wrote:\n> > On thursday, 02 June 2005 16:41 Vincent Hanquez wrote:\n> > > On Thu, Jun 02, 2005 at 04:02:07PM +0400, Alexey Nezhdanov wrote:\n> > > >  Git arhives are normally totally self-sufficient, and it's worth\n> > > > noting\n> > >\n> > >        ^^^^^^^\n> > > and one more here\n> >\n> > Why? It's ok to speak about many [existing] archives here.\n>\n> it's missing a 'c'\nok :)\n\n-- \nRespectfully\nAlexey Nezhdanov\n\n"},{"id":"4419","messageId":"Pine.LNX.4.58.0506020752130.1876@ppc970.osdl.org","threadId":"770","inReplyTo":"m18y1tb55c.fsf@ebiederm.dsl.xmission.com","subject":"Re: I want to release a \"git-1.0\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-02T14:52:30Z","receivedAt":"2005-06-02T14:52:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 2 Jun 2005, Eric W. Biederman wrote:\n> \n> So I need to do a git checkout of the latest version of git to\n> read the tutorial?  So I can figure out how to use git?\n\nJust use the gitweb thing, it's easy to read off there..\n\n\t\tLinus\n"},{"id":"4440","messageId":"7vll5s35pd.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505301253070.1876@ppc970.osdl.org","subject":"CVS migration section to the tutorial.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-02T19:43:26Z","receivedAt":"2005-06-02T19:43:26Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"I think a section to discuss \"I am used to doing 'cvs xxx' to solve\nthis problem, how do I do that in GIT\" would be a good idea.  Here is\nan example to talk about \"cvs annotate\".\n\n------------\n\nCVS annotate.\n\nThe core GIT itself does not do \"cvs annotate\" equivalent, but it has\nsomething much nicer.\n\nLet's step back a bit and think about the reason why you would want to\ndo \"cvs annotate a-file.c\" to begin with.\n\n\t- Are you really interested in _all_ the lines in that file?\n\n\t- Are you interested in lines _only_ in that file and do not\n          care if the file was created by renaming from a different\n          file?\n\nYou would use \"cvs annotate\" on a file when you have trouble with a\nfunction (or even a single \"if\" statement in that function) that\nhappens to be defined in the file, which does not do what you want it\nto do.  And you would want to find out why it was written in that way,\nbecause you are about to modify it to suit your needs, and at the same\ntime you do not want to break its current callers.  For that, you want\nto find out why the original author did things that way in the\noriginal context.  That's why you want \"cvs annotate\".  So your answer\nto the first question _should_ be \"no\".  You do not care about the\nwhole file, only a segment of it.\n\nAlso, in the original context, the same statement might have appeared\nat first in a different file and later the file was renamed to\n\"a-file.c\".  Or the entire program may have constructs similar to the\n\"if\" statement you are having trouble with in different places, that\nyou are still not aware of.  So your answer to the second question\n_should_ be \"no\" as well.\n\nAs an example, assuming that you have this piece code that you are\ninterested in in the HEAD version:\n\n\tif (frotz) {\n\t\tnitfol();\n\t}\n\nyou would use git-rev-list and git-diff-tree like this:\n\n\t$ git-rev-list HEAD |\n\t  git-diff-tree --stdin -v -p -S'if (frotz) {\n\t\tnitfol();\n\t}'\n\nWe have already talked about the \"--stdin\" form of git-diff-tree\ncommand that reads the list of commits and compares each commit with\nits parents.  What the -S flag and its argument does is called\n\"pickaxe\", a tool for software archaeologists.  When \"pickaxe\" is\nused, git-diff-tree command outputs differences between two commits\nonly if one tree has the specified string in a file and the\ncorresponding file in the other tree does not.  The above example\nlooks for a commit that has the \"if\" statement in it in a file, but\nits parent commit does not have it in the same shape in the\ncorresponding file (or the other way around, where the parent has it\nand the commit does not), and the differences between them are shown,\nalong with the commit message (thanks to the -v flag).  It does not\nshow anything for commits that do not touch this \"if\" statement.\n\nTo make things more interesting, you can give the -C flag to\ngit-diff-tree, like this:\n\n\t$ git-rev-list HEAD |\n\t  git-diff-tree --stdin -v -p -C -S'if (frotz) {\n\t\tnitfol();\n\t}'\n\nWhen the -C flag is used, file renames and copies are followed.  So if\nthe \"if\" statement in question happens to be in \"a-file.c\" in the\ncurrent HEAD commit, even if the file was originally called \"o-file.c\"\nand then renamed in an earlier commit, or if the file was created by\ncopying an existing \"o-file.c\" in an earlier commit, you will not lose\ntrack.  If the \"if\" statement did not change across such rename or\ncopy, then the commit that does rename or copy would not show in the\noutput, and if the \"if\" statement was modified while the file was\nstill called \"o-file.c\", it would find the commit that changed the\nstatement when it was in \"o-file.c\".\n\n[ BTW, the current versions of \"git-diff-tree -C\" is not eager enough\n  to find copies, and it will miss the fact that a-file.c was created\n  by copying o-file.c unless o-file.c was somehow changed in the same\n  commit.]\n\nTo make things even more interesting, you can use the --pickaxe-all\nflag in addition to the -S flag.  This causes the differences from all\nthe files contained in those two commits, not just the differences\nbetween the files that contain this changed \"if\" statement:\n\n\t$ git-rev-list HEAD |\n\t  git-diff-tree --stdin -v -p -C -S'if (frotz) {\n\t\tnitfol();\n\t}' --pickaxe-all\n\n\n"},{"id":"4464","messageId":"00e101c567cc$80c0de80$03c8a8c0@kroptech.com","threadId":"770","inReplyTo":"Pine.LNX.4.58.0505312002160.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Adam Kropelin","fromEmail":"akropel1@rochester.rr.com","sentAt":"2005-06-02T23:40:48Z","receivedAt":"2005-06-02T23:40:48Z","isPatch":false,"sender":{"key":"akropel1@rochester.rr.com","avatar":null},"body":"Linus Torvalds wrote:\n> Anyway, I wrote just a _very_ introductory thing in\n> Documentation/tutorial.txt, I'll try to update and expand on it later. \n> It\n> basically has a really stupid example of \"how to set up a new \n> project\".\n\nI've been working my way thru the tutorial, trying to up my git clue \nlevel a bit. One part where things start to go a bit pear-shaped for me \nis in the description of git-diff-files vs. git-diff-cache. The tutorial \ntakes pains to emphasize the difference between \"working directory \ncontents\", \"index file\", and \"committed tree\", and I'm on board with \nthat. What confuses me is the following:\n\n> Unlike \"git-diff-files\", which showed the difference between the index\n> file and the working directory, \"git-diff-cache\" shows the differences\n> between a committed _tree_ and the index file.\n> ...\n> [example where git-diff-cache shows difference between working\n> directory and committed tree]\n> ...\n> \"git-diff-cache\" also has a specific flag \"--cached\", which is used to\n> tell it to show the differences purely with the index file, and ignore\n> the current working directory state entirely\n\nThe example and the description of --cached seem to contradict the first \nsentence's description the tool's purpose in life. If it shows you \ndifferences between a committed tree and the index file, why is it \nlooking in my working directory at all? In order to get the behavior the \nfirst sentence describes you actually have to use --cached.\n\nAm I on right track?\n\n--Adam\n\n"},{"id":"4465","messageId":"7vll5sz54z.fsf_-_@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"7vmzqau3es.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH] Fix -B \"very-different\" logic.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-02T23:54:36Z","receivedAt":"2005-06-02T23:54:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"JCH\" == Junio C Hamano <junkio@cox.net> writes:\n\n>     (Btw, current versions of git will consider the change\n>     in question to be so big that it's considered a whole\n>     new file, since the diff is actually bigger than the\n>     file. \n\nJCH> Do you want me to do something about this with -B (and possibly\nJCH> -C/-M), like skipping the comparison altogether if the file size\nJCH> is smaller than, say, 1k bytes or something silly like that?  Or\nJCH> not having special case for this kind of \"contrived example\"\nJCH> preferrable?\n\nI was looking at the -B code.  The reason it thinks change is\ntoo big is because xdelta tells us to reconstruct the\ndestination by all new literal bytes in this small string case.\nThere is not much I can do about it.\n\nHowever I think the diffcore-break algorithm itself was basing\nits \"very_different\" computation on numbers somewhat bogus.  It\nwas counting newly inserted bytes into account, but amount of\nthose bytes should not make any difference when determining if\nthe change is a complete rewrite.\n\nI suspect that -M/-C heuristics has similar (if not the same)\nissues, but I would like to address that separately.\n\nHere is a proposed fix for -B.  It also tells diffcore-break not\nto break a file smaller than 400 bytes.  I did not make this\nnumber configurable, since that would be too many knobs to\ntweak.  If somebody feels strong enough about it, it can be made\ninto an option later, but for now that size \"feels\" reasonable.\n\n  -- >8 -- cut here -- >8 --\n\n------------\nWhat we are interested in here is how much the original source\nmaterial remains in the final result, and it does not really\nmatter how much new contents are added as part of the edit.  If\nyou remove 97 lines from an original 100-line document, it does\nnot matter if you add 47 lines of your own to make a 50-line\ndocument, or if you add 997 lines to make a 1000-line document.\nEither way, you did a complete rewrite.\n\nEarlier code counted both new material and deletions to detect\ncomplete rewrites.  This patch fixes it.  With its default\nsetting, it detects three such complete rewrites in the core-GIT\nrepository.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n count-delta.h    |    1 +\n diffcore.h       |    4 ++-\n count-delta.c    |   70 ++++++++++++++++++++++++++++++++++++++++++++++++++++++\n diffcore-break.c |   55 +++++++++++++++---------------------------\n 4 files changed, 92 insertions(+), 38 deletions(-)\n\ndiff --git a/count-delta.h b/count-delta.h\n--- a/count-delta.h\n+++ b/count-delta.h\n@@ -5,5 +5,6 @@\n #define COUNT_DELTA_H\n \n unsigned long count_delta(void *, unsigned long);\n+unsigned long count_excluded_source_material(void *, unsigned long);\n \n #endif\ndiff --git a/diffcore.h b/diffcore.h\n--- a/diffcore.h\n+++ b/diffcore.h\n@@ -10,9 +10,9 @@\n  */\n #define MAX_SCORE 60000\n #define DEFAULT_RENAME_SCORE 30000 /* rename/copy similarity minimum (50%) */\n-#define DEFAULT_BREAK_SCORE  59400 /* minimum for break to happen (99%)*/\n+#define DEFAULT_BREAK_SCORE  48000 /* minimum for break to happen    (80%) */\n \n-#define RENAME_DST_MATCHED 01\n+#define DIFF_MINIMUM_BREAK   400 /* minimum size of source that -B breaks */\n \n struct diff_filespec {\n \tunsigned char sha1[20];\ndiff --git a/count-delta.c b/count-delta.c\n--- a/count-delta.c\n+++ b/count-delta.c\n@@ -93,3 +93,73 @@ unsigned long count_delta(void *delta_bu\n \t\treturn 0;\n \treturn (src_size - copied_from_source) + added_literal;\n }\n+\n+\n+/*\n+ * What we are interested in here is how much the original source\n+ * material remains in the final result, and it does not really matter\n+ * how much new contents are added as part of the edit.  If you remove\n+ * 97 lines from an original 100-line document, it does not matter if\n+ * you add 47 lines of your own to make a 50-line document, or if you\n+ * add 997 lines to make a 1000-line document.  Either way, you did a\n+ * complete rewrite.\n+ *\n+ * Note.  We do not interprete delta fully.  Instead, we look at xdelta\n+ * instructions that copy bytes from the source, and count those copied\n+ * bytes.  Subtracting this number from the original source size yields\n+ * the number of bytes not used from the source material.  In the above\n+ * example, this number corresponds to 97-line (but we count in bytes).\n+ */\n+unsigned long count_excluded_source_material(void *delta_buf,\n+\t\t\t\t\t     unsigned long delta_size)\n+{\n+\tunsigned long copied_from_source;\n+\tconst unsigned char *data, *top;\n+\tunsigned char cmd;\n+\tunsigned long src_size, dst_size, out;\n+\n+\t/* the smallest delta size possible is 6 bytes */\n+\tif (delta_size < 6)\n+\t\treturn UINT_MAX;\n+\n+\tdata = delta_buf;\n+\ttop = delta_buf + delta_size;\n+\n+\tsrc_size = get_hdr_size(&data);\n+\tdst_size = get_hdr_size(&data);\n+\n+\tcopied_from_source = out = 0;\n+\twhile (data < top) {\n+\t\tcmd = *data++;\n+\t\tif (cmd & 0x80) {\n+\t\t\tunsigned long cp_off = 0, cp_size = 0;\n+\t\t\tif (cmd & 0x01) cp_off = *data++;\n+\t\t\tif (cmd & 0x02) cp_off |= (*data++ << 8);\n+\t\t\tif (cmd & 0x04) cp_off |= (*data++ << 16);\n+\t\t\tif (cmd & 0x08) cp_off |= (*data++ << 24);\n+\t\t\tif (cmd & 0x10) cp_size = *data++;\n+\t\t\tif (cmd & 0x20) cp_size |= (*data++ << 8);\n+\t\t\tif (cp_size == 0) cp_size = 0x10000;\n+\n+\t\t\tif (cmd & 0x40)\n+\t\t\t\t/* copy from dst */\n+\t\t\t\t;\n+\t\t\telse\n+\t\t\t\tcopied_from_source += cp_size;\n+\t\t\tout += cp_size;\n+\t\t} else {\n+\t\t\t/* write literal into dst */\n+\t\t\tout += cmd;\n+\t\t\tdata += cmd;\n+\t\t}\n+\t}\n+\n+\t/* sanity check */\n+\tif (data != top || out != dst_size)\n+\t\treturn UINT_MAX;\n+\n+\tif (src_size < copied_from_source)\n+\t\t/* we ended up overcounting and underflowed; I dunno why */\n+\t\treturn 0;\n+\treturn src_size - copied_from_source;\n+}\ndiff --git a/diffcore-break.c b/diffcore-break.c\n--- a/diffcore-break.c\n+++ b/diffcore-break.c\n@@ -13,63 +13,46 @@ static int very_different(struct diff_fi\n {\n \t/* dst is recorded as a modification of src.  Are they so\n \t * different that we are better off recording this as a pair\n-\t * of delete and create?  min_score is the minimum amount of\n-\t * new material that must exist in the dst and not in src for\n-\t * the pair to be considered a complete rewrite, and recommended\n-\t * to be set to a very high value, 99% or so.\n+\t * of delete and create?\n \t *\n-\t * The value we return represents the amount of new material\n-\t * that is in dst and not in src.  We return 0 when we do not\n-\t * want to get the filepair broken.\n+\t * We base the score on the amount of material originally from\n+\t * src that still remains in the dst.  If src was 100-line\n+\t * file among which only 3-line remains in the dst, then it is\n+\t * a complete rewrite with 97% \"change\", and it does not\n+\t * matter if the resulting file is a 15-line file or a\n+\t * 2000-line file.  On the other hand, if 40-line remains\n+\t * among those 100-lines, even if the resulting file is a\n+\t * 2000-lines file, it still is an edit with 60% \"change\",\n+\t * which may sound counter-intuitive at first but that is the\n+\t * right number to use.\n \t */\n+\n \tvoid *delta;\n-\tunsigned long delta_size, base_size;\n+\tunsigned long delta_size;\n \n \tif (!S_ISREG(src->mode) || !S_ISREG(dst->mode))\n \t\treturn 0; /* leave symlink rename alone */\n \n-\tif (diff_populate_filespec(src, 1) || diff_populate_filespec(dst, 1))\n-\t\treturn 0; /* error but caught downstream */\n-\n-\tdelta_size = ((src->size < dst->size) ?\n-\t\t      (dst->size - src->size) : (src->size - dst->size));\n-\n-\t/* Notice that we use max of src and dst as the base size,\n-\t * unlike rename similarity detection.  This is so that we do\n-\t * not mistake a large addition as a complete rewrite.\n-\t */\n-\tbase_size = ((src->size < dst->size) ? dst->size : src->size);\n-\n-\t/*\n-\t * If file size difference is too big compared to the\n-\t * base_size, we declare this a complete rewrite.\n-\t */\n-\tif (base_size * min_score < delta_size * MAX_SCORE)\n-\t\treturn MAX_SCORE;\n-\n \tif (diff_populate_filespec(src, 0) || diff_populate_filespec(dst, 0))\n \t\treturn 0; /* error but caught downstream */\n \n+\tif (src->size < DIFF_MINIMUM_BREAK)\n+\t\treturn 0; /* Too small to consider breaking */\n+\n \tdelta = diff_delta(src->data, src->size,\n \t\t\t   dst->data, dst->size,\n \t\t\t   &delta_size);\n \n-\t/* A delta that has a lot of literal additions would have\n-\t * big delta_size no matter what else it does.\n-\t */\n-\tif (base_size * min_score < delta_size * MAX_SCORE)\n-\t\treturn MAX_SCORE;\n-\n \t/* Estimate the edit size by interpreting delta. */\n-\tdelta_size = count_delta(delta, delta_size);\n+\tdelta_size = count_excluded_source_material(delta, delta_size);\n \tfree(delta);\n \tif (delta_size == UINT_MAX)\n \t\treturn 0; /* error in delta computation */\n \n-\tif (base_size < delta_size)\n+\tif (src->size < delta_size)\n \t\treturn MAX_SCORE;\n \n-\treturn delta_size * MAX_SCORE / base_size; \n+\treturn delta_size * MAX_SCORE / src->size;\n }\n \n void diffcore_break(int min_score)\n------------\n\n"},{"id":"4466","messageId":"Pine.LNX.4.58.0506021705520.1876@ppc970.osdl.org","threadId":"770","inReplyTo":"00e101c567cc$80c0de80$03c8a8c0@kroptech.com","subject":"Re: I want to release a \"git-1.0\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-03T00:06:30Z","receivedAt":"2005-06-03T00:06:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 2 Jun 2005, Adam Kropelin wrote:\n> What confuses me is the following:\n\nYeah, I'll try to clarify.\n\ngit-diff-cache can show the difference between a tree and either the index \n_or_ the working directory. Will fix up.\n\n\t\tLinus\n"},{"id":"4467","messageId":"Pine.LNX.4.58.0506021716140.1876@ppc970.osdl.org","threadId":"770","inReplyTo":"7vll5sz54z.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Fix -B \"very-different\" logic.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-03T00:21:43Z","receivedAt":"2005-06-03T00:21:43Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 2 Jun 2005, Junio C Hamano wrote:\n> \n> However I think the diffcore-break algorithm itself was basing\n> its \"very_different\" computation on numbers somewhat bogus.  It\n> was counting newly inserted bytes into account, but amount of\n> those bytes should not make any difference when determining if\n> the change is a complete rewrite.\n\nCareful. \n\nI think the amount of new code _should_ matter. Otherwise, an old empty\nfile would always be considered the source of a new file, since the diff \ndoesn't remove anything. Similarly, just because we have a boilerplate \nfile shouldn't make that always be considered a \"wonderful source\", when \npeople add the real meat to it.\n\nSo I think you're on the right track, but I don't think you should\nentirely dismiss \"lots of stuff added\" as a reason for a \"break\". I think\nthat if the new stuff is _much_ larger than the old stuff, it might as\nwell be considered a rewrite.\n\nIn particular, let's say that I used to have two files:\n\n\ta.c - small helper functions\n\tb.c - the \"meat\" of the thing\n\nand I end up deciding that I might as well collapse it all into one file, \na.c. What happens? There's almost no deletes from a.c, but there's a lot \nof new code in it. \n\nWouldn't it be _better_ if you considered the new \"a.c\" a new file, so \nthat you might notice that it's actually _closer_ to the old removed \"b.c\" \nthan the old \"a.c\"?\n\nSee what I'm saying?\n\n\t\tLinus\n"},{"id":"4468","messageId":"Pine.LNX.4.58.0506021745310.1876@ppc970.osdl.org","threadId":"770","inReplyTo":"Pine.LNX.4.58.0506021705520.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-03T00:47:45Z","receivedAt":"2005-06-03T00:47:45Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 2 Jun 2005, Linus Torvalds wrote:\n>\n> Yeah, I'll try to clarify.\n\nAdam, do you find the current version a bit more clear on this?\n\n\t\tLinus\n"},{"id":"4470","messageId":"7vis0wusv5.fsf@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"Pine.LNX.4.58.0506021716140.1876@ppc970.osdl.org","subject":"Re: [PATCH] Fix -B \"very-different\" logic.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-03T01:33:18Z","receivedAt":"2005-06-03T01:33:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> Careful. \n\nLT> I think the amount of new code _should_ matter. Otherwise, an old empty\nLT> file would always be considered the source of a new file, since the diff \nLT> doesn't remove anything. Similarly, just because we have a boilerplate \nLT> file shouldn't make that always be considered a \"wonderful source\", when \nLT> people add the real meat to it.\n\nYes, I agree that rename/copy logic should use different\nheuristics from the one I proposed for breaking.\n\nIt is my assumption that people in practice tend to make only\nsmall edits after a rename/copy just to adjust things like:\n\n - filenames mentioned in the comment of the file itself,\n\n - include paths that refer other files if the file was\n   moved/copied from a different directory,\n\n - names of functions and variables.\n\nand making sure there would not be too much new stuff is quite\nuseful to detect rename/copy source correctly as the current\nsimilarity estimator in diffcore-rename does.  I do not intend\nto touch that.\n\nThe boilderplate example you mention is a very good reason not\nto dismiss the amount of new material when doing rename/copy\ndetection.\n\nLT> In particular, let's say that I used to have two files:\n\nLT> \ta.c - small helper functions\nLT> \tb.c - the \"meat\" of the thing\n\nLT> and I end up deciding that I might as well collapse it all into one file, \nLT> a.c. What happens? There's almost no deletes from a.c, but there's a lot \nLT> of new code in it. \n\nLT> See what I'm saying?\n\nYes.  I think I do.\n\nWhen git-diff-tree -B -C runs your example, it feeds diffcore\nwith these:\n\n  :100644 100644 sha1-a-helper-only sha1-a-and-meat M   a.c\n  :100644 000000 sha1-b-stale-meat  0{40}           D   b.c\n\nThe ideal diffcore-break breaks a.c because it looks at\ninsertions as well:\n\n  :100644 000000 sha1-a-helper-only 0{40}           D   a.c\n  :000000 100644 0{40}              sha1-a-and-meat N   a.c\n  :100644 000000 sha1-b-stale-meat  0{40}           D   b.c\n\nThen diffcore-rename notices that sha1-b-stale-meat is better\nmatch than sha1-a-helper-only to produce sha1-a-and-meat, and\nresolves the above to:\n\n  :100644 100644 sha1-b-stale-meat  sha1-a-and-meat R   b.c\ta.c\n\nUp to this point is just a demonstration that I see your point.\n\nBut I still want to keep the example I gave in the original\ncommit message.  Suppose you did not have b.c file under version\ncontrol, and did the same operation.  I.e. a.c acquired a lot of\ngood stuff.  git-diff-tree -B -C feeds:\n\n  :100644 100644 sha1-a-helper-only sha1-a-and-meat M   a.c\n\nwhich is broken into:\n\n  :100644 000000 sha1-a-helper-only 0{40}           D   a.c\n  :000000 100644 0{40}              sha1-a-and-meat N   a.c\n\nUnfortunately, in this case nobody absorbs these pairs.  I want\nto allow you to add 1000 lines of new stuff to a file (which was\noriginally 100 lines long) as long as you do not remove too many\nlines from the original 100 lines without triggering \"this is a\nrewrite\" logic in this case.  So after rename/copy runs, we need\nto match these up and merge them back into the original.\n\n  :100644 100644 sha1-a-helper-only sha1-a-and-meat M   a.c\n\nWe should carry a bit more information about broken entries than\nwe currently do.  We would break a pair based on both deletion\nand insertion, just like the current code (i.e. without the\npatch you are responding to) does.  But when we do break a pair,\nwe need to mark them if the \"new\" side have enough original\nsource material remaining.  If we have such mark to tell us that\n\"these were broken but there are a good chunk of source material\nremaining\", the clean-up phase, to run after diffcore-rename\nfinishes, should be able to notice surviving broken pairs and\nmerge them back accordingly.\n\n"},{"id":"4471","messageId":"012a01c567dc$542787b0$03c8a8c0@kroptech.com","threadId":"770","inReplyTo":"Pine.LNX.4.58.0506021745310.1876@ppc970.osdl.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Adam Kropelin","fromEmail":"akropel1@rochester.rr.com","sentAt":"2005-06-03T01:34:05Z","receivedAt":"2005-06-03T01:34:05Z","isPatch":false,"sender":{"key":"akropel1@rochester.rr.com","avatar":null},"body":"Linus Torvalds wrote:\n> On Thu, 2 Jun 2005, Linus Torvalds wrote:\n>>\n>> Yeah, I'll try to clarify.\n>\n> Adam, do you find the current version a bit more clear on this?\n\nAbsolutely. I especially like the new digression explaining that \nthe --cached flag controls where file _content_ is fetched from and \nreinforcing that the index file always governs which files are involved \nin the diff.\n\nThanks!\n\n--Adam\n\n"},{"id":"4477","messageId":"7vis0vq1rz.fsf_-_@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"7vis0wusv5.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH 0/4] Fix -B \"very-different\" logic.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-03T08:32:00Z","receivedAt":"2005-06-03T08:32:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"I am sending the following four patch series:\n\n        [PATCH 1/4] Tweak count-delta interface\n        [PATCH 2/4] diff: Fix docs and add -O to diff-helper.\n        [PATCH 3/4] diff: Clean up diff_scoreopt_parse().\n        [PATCH 4/4] diff: Update -B heuristics.\n\nThe first three are preparations and cleanups I found necessary\nwhile I was working on the last one, which is the gem of this\nseries.  It addresses the concerns you raised in your message\n\"Careful.\" while keeping the semantics I wanted to have \"if you\nkeep 97 lines out of original 100-line document, it does not\nmatter if the end result is a 110-line or 1000-line document.\nYou did not do a rewrite.\"\n\nYou may have to remove the warning about git-status with this\nchange, though.\n\n"},{"id":"4478","messageId":"7vekbjq1l8.fsf_-_@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"7vis0vq1rz.fsf_-_@assigned-by-dhcp.cox.net","subject":"[PATCH 1/4] Tweak count-delta interface","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-03T08:36:03Z","receivedAt":"2005-06-03T08:36:03Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Make it return copied source and insertion separately, so that\nlater implementation of heuristics can use them more flexibly.\n\nThis does not change the heuristics implemented in\ndiffcore-rename nor diffcore-break in any way.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n count-delta.h     |    3 ++-\n diffcore.h        |    2 --\n count-delta.c     |   30 ++++++++++++++++--------------\n diffcore-break.c  |   15 +++++++++++----\n diffcore-rename.c |   15 +++++++++++----\n 5 files changed, 40 insertions(+), 25 deletions(-)\n\ndiff --git a/count-delta.h b/count-delta.h\n--- a/count-delta.h\n+++ b/count-delta.h\n@@ -4,6 +4,7 @@\n #ifndef COUNT_DELTA_H\n #define COUNT_DELTA_H\n \n-unsigned long count_delta(void *, unsigned long);\n+int count_delta(void *, unsigned long,\n+\t\tunsigned long *src_copied, unsigned long *literal_added);\n \n #endif\ndiff --git a/diffcore.h b/diffcore.h\n--- a/diffcore.h\n+++ b/diffcore.h\n@@ -12,8 +12,6 @@\n #define DEFAULT_RENAME_SCORE 30000 /* rename/copy similarity minimum (50%) */\n #define DEFAULT_BREAK_SCORE  59400 /* minimum for break to happen (99%)*/\n \n-#define RENAME_DST_MATCHED 01\n-\n struct diff_filespec {\n \tunsigned char sha1[20];\n \tchar *path;\ndiff --git a/count-delta.c b/count-delta.c\n--- a/count-delta.c\n+++ b/count-delta.c\n@@ -29,15 +29,18 @@ static unsigned long get_hdr_size(const \n /*\n  * NOTE.  We do not _interpret_ delta fully.  As an approximation, we\n  * just count the number of bytes that are copied from the source, and\n- * the number of literal data bytes that are inserted.  Number of\n- * bytes that are _not_ copied from the source is deletion, and number\n- * of inserted literal bytes are addition, so sum of them is what we\n- * return.  xdelta can express an edit that copies data inside of the\n- * destination which originally came from the source.  We do not count\n- * that in the following routine, so we are undercounting the source\n- * material that remains in the final output that way.\n+ * the number of literal data bytes that are inserted.\n+ *\n+ * Number of bytes that are _not_ copied from the source is deletion,\n+ * and number of inserted literal bytes are addition, so sum of them\n+ * is the extent of damage.  xdelta can express an edit that copies\n+ * data inside of the destination which originally came from the\n+ * source.  We do not count that in the following routine, so we are\n+ * undercounting the source material that remains in the final output\n+ * that way.\n  */\n-unsigned long count_delta(void *delta_buf, unsigned long delta_size)\n+int count_delta(void *delta_buf, unsigned long delta_size,\n+\t\tunsigned long *src_copied, unsigned long *literal_added)\n {\n \tunsigned long copied_from_source, added_literal;\n \tconst unsigned char *data, *top;\n@@ -46,7 +49,7 @@ unsigned long count_delta(void *delta_bu\n \n \t/* the smallest delta size possible is 6 bytes */\n \tif (delta_size < 6)\n-\t\treturn UINT_MAX;\n+\t\treturn -1;\n \n \tdata = delta_buf;\n \ttop = delta_buf + delta_size;\n@@ -83,13 +86,12 @@ unsigned long count_delta(void *delta_bu\n \n \t/* sanity check */\n \tif (data != top || out != dst_size)\n-\t\treturn UINT_MAX;\n+\t\treturn -1;\n \n \t/* delete size is what was _not_ copied from source.\n \t * edit size is that and literal additions.\n \t */\n-\tif (src_size + added_literal < copied_from_source)\n-\t\t/* we ended up overcounting and underflowed */\n-\t\treturn 0;\n-\treturn (src_size - copied_from_source) + added_literal;\n+\t*src_copied = copied_from_source;\n+\t*literal_added = added_literal;\n+\treturn 0;\n }\ndiff --git a/diffcore-break.c b/diffcore-break.c\n--- a/diffcore-break.c\n+++ b/diffcore-break.c\n@@ -23,7 +23,7 @@ static int very_different(struct diff_fi\n \t * want to get the filepair broken.\n \t */\n \tvoid *delta;\n-\tunsigned long delta_size, base_size;\n+\tunsigned long delta_size, base_size, src_copied, literal_added;\n \n \tif (!S_ISREG(src->mode) || !S_ISREG(dst->mode))\n \t\treturn 0; /* leave symlink rename alone */\n@@ -61,10 +61,17 @@ static int very_different(struct diff_fi\n \t\treturn MAX_SCORE;\n \n \t/* Estimate the edit size by interpreting delta. */\n-\tdelta_size = count_delta(delta, delta_size);\n+\tif (count_delta(delta, delta_size, &src_copied, &literal_added)) {\n+\t\tfree(delta);\n+\t\treturn 0;\n+\t}\n \tfree(delta);\n-\tif (delta_size == UINT_MAX)\n-\t\treturn 0; /* error in delta computation */\n+\n+\t/* Extent of damage */\n+\tif (src->size + literal_added < src_copied)\n+\t\tdelta_size = 0;\n+\telse\n+\t\tdelta_size = (src->size - src_copied) + literal_added;\n \n \tif (base_size < delta_size)\n \t\treturn MAX_SCORE;\ndiff --git a/diffcore-rename.c b/diffcore-rename.c\n--- a/diffcore-rename.c\n+++ b/diffcore-rename.c\n@@ -135,7 +135,7 @@ static int estimate_similarity(struct di\n \t * call into this function in that case.\n \t */\n \tvoid *delta;\n-\tunsigned long delta_size, base_size;\n+\tunsigned long delta_size, base_size, src_copied, literal_added;\n \tint score;\n \n \t/* We deal only with regular files.  Symlink renames are handled\n@@ -174,10 +174,17 @@ static int estimate_similarity(struct di\n \t\treturn 0;\n \n \t/* Estimate the edit size by interpreting delta. */\n-\tdelta_size = count_delta(delta, delta_size);\n-\tfree(delta);\n-\tif (delta_size == UINT_MAX)\n+\tif (count_delta(delta, delta_size, &src_copied, &literal_added)) {\n+\t\tfree(delta);\n \t\treturn 0;\n+\t}\n+\tfree(delta);\n+\n+\t/* Extent of damage */\n+\tif (src->size + literal_added < src_copied)\n+\t\tdelta_size = 0;\n+\telse\n+\t\tdelta_size = (src->size - src_copied) + literal_added;\n \n \t/*\n \t * Now we will give some score to it.  100% edit gets 0 points\n------------\n\n"},{"id":"4479","messageId":"7v8y1rq1k4.fsf_-_@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"7vis0vq1rz.fsf_-_@assigned-by-dhcp.cox.net","subject":"[PATCH 2/4] diff: Fix docs and add -O to diff-helper.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-03T08:36:43Z","receivedAt":"2005-06-03T08:36:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This patch updates diff documentation and usage strings:\n\n - clarify the semantics of -R.  It is not \"output in reverse\";\n   rather, it is \"I will feed diff backwards\".  Semantically\n   they are different when -C is involved.\n\n - describe -O in usage strings of diff-* brothers.  It was\n   implemented, documented but not described in usage text.\n\nAlso it adds -O to diff-helper.  Like -S (and unlike -M/-C/-B),\nthis option can work on sanitized diff-raw output produced by\nthe diff-* brothers.  While we are at it, the call it makes to\ndiffcore is cleaned up to use the diffcore_std() like everybody\nelse, and the declaration for the low level diffcore routines\nare moved from diff.h (public) to diffcore.h (private between\ndiff.c and diffcore backends).\n\nSigned-off-by: Junio C Hamano <junkio@cox.net> \n---\n\n Documentation/git-diff-cache.txt  |    3 ++-\n Documentation/git-diff-files.txt  |    3 ++-\n Documentation/git-diff-helper.txt |    5 ++++-\n Documentation/git-diff-tree.txt   |    2 +-\n diff.h                            |   10 +---------\n diffcore.h                        |    6 ++++++\n diff-cache.c                      |    2 +-\n diff-files.c                      |    2 +-\n diff-helper.c                     |   25 ++++++++++++++-----------\n diff-tree.c                       |    2 +-\n 10 files changed, 33 insertions(+), 27 deletions(-)\n\ndiff --git a/Documentation/git-diff-cache.txt b/Documentation/git-diff-cache.txt\n--- a/Documentation/git-diff-cache.txt\n+++ b/Documentation/git-diff-cache.txt\n@@ -57,7 +57,8 @@ OPTIONS\n \t<orderfile>, which has one shell glob pattern per line.\n \n -R::\n-\tOutput diff in reverse.\n+\tSwap two inputs; that is, show differences from cache or\n+\ton-disk file to tree contents.\n \n --cached::\n \tdo not consider the on-disk file at all\ndiff --git a/Documentation/git-diff-files.txt b/Documentation/git-diff-files.txt\n--- a/Documentation/git-diff-files.txt\n+++ b/Documentation/git-diff-files.txt\n@@ -27,7 +27,8 @@ OPTIONS\n \tRemain silent even on nonexisting files\n \n -R::\n-\tOutput diff in reverse.\n+\tSwap two inputs; that is, show differences from on-disk files\n+\tto cache contents.\n \n -B::\n \tBreak complete rewrite changes into pairs of delete and create.\ndiff --git a/Documentation/git-diff-helper.txt b/Documentation/git-diff-helper.txt\n--- a/Documentation/git-diff-helper.txt\n+++ b/Documentation/git-diff-helper.txt\n@@ -9,7 +9,7 @@ git-diff-helper - Generates patch format\n \n SYNOPSIS\n --------\n-'git-diff-helper' [-z] [-S<string>]\n+'git-diff-helper' [-z] [-S<string>] [-O<orderfile>]\n \n DESCRIPTION\n -----------\n@@ -24,6 +24,9 @@ OPTIONS\n -S<string>::\n \tLook for differences that contains the change in <string>.\n \n+-O<orderfile>::\n+\tOutput the patch in the order specified in the\n+\t<orderfile>, which has one shell glob pattern per line.\n \n See Also\n --------\ndiff --git a/Documentation/git-diff-tree.txt b/Documentation/git-diff-tree.txt\n--- a/Documentation/git-diff-tree.txt\n+++ b/Documentation/git-diff-tree.txt\n@@ -43,7 +43,7 @@ OPTIONS\n \tDetect copies as well as renames.\n \n -R::\n-\tOutput diff in reverse.\n+\tSwap two input trees.\n \n -S<string>::\n \tLook for differences that contains the change in <string>.\ndiff --git a/diff.h b/diff.h\n--- a/diff.h\n+++ b/diff.h\n@@ -35,21 +35,13 @@ extern int diff_scoreopt_parse(const cha\n #define DIFF_SETUP_REVERSE      \t1\n #define DIFF_SETUP_USE_CACHE\t\t2\n #define DIFF_SETUP_USE_SIZE_CACHE\t4\n+\n extern void diff_setup(int flags);\n \n #define DIFF_DETECT_RENAME\t1\n #define DIFF_DETECT_COPY\t2\n \n-extern void diffcore_rename(int rename_copy, int minimum_score);\n-\n #define DIFF_PICKAXE_ALL\t1\n-extern void diffcore_pickaxe(const char *needle, int opts);\n-\n-extern void diffcore_pathspec(const char **pathspec);\n-\n-extern void diffcore_order(const char *orderfile);\n-\n-extern void diffcore_break(int max_score);\n \n extern void diffcore_std(const char **paths,\n \t\t\t int detect_rename, int rename_score,\ndiff --git a/diffcore.h b/diffcore.h\n--- a/diffcore.h\n+++ b/diffcore.h\n@@ -73,6 +73,12 @@ extern struct diff_filepair *diff_queue(\n \t\t\t\t\tstruct diff_filespec *);\n extern void diff_q(struct diff_queue_struct *, struct diff_filepair *);\n \n+extern void diffcore_pathspec(const char **pathspec);\n+extern void diffcore_break(int);\n+extern void diffcore_rename(int rename_copy, int);\n+extern void diffcore_pickaxe(const char *needle, int opts);\n+extern void diffcore_order(const char *orderfile);\n+\n #define DIFF_DEBUG 0\n #if DIFF_DEBUG\n void diff_debug_filespec(struct diff_filespec *, int, const char *);\ndiff --git a/diff-cache.c b/diff-cache.c\n--- a/diff-cache.c\n+++ b/diff-cache.c\n@@ -157,7 +157,7 @@ static void mark_merge_entries(void)\n }\n \n static char *diff_cache_usage =\n-\"git-diff-cache [-p] [-r] [-z] [-m] [-M] [-C] [-R] [-S<string>] [--cached] <tree-ish> [<path>...]\";\n+\"git-diff-cache [-p] [-r] [-z] [-m] [-M] [-C] [-R] [-S<string>] [-O<orderfile>] [--cached] <tree-ish> [<path>...]\";\n \n int main(int argc, const char **argv)\n {\ndiff --git a/diff-files.c b/diff-files.c\n--- a/diff-files.c\n+++ b/diff-files.c\n@@ -7,7 +7,7 @@\n #include \"diff.h\"\n \n static const char *diff_files_usage =\n-\"git-diff-files [-p] [-q] [-r] [-z] [-M] [-C] [-R] [-S<string>] [paths...]\";\n+\"git-diff-files [-p] [-q] [-r] [-z] [-M] [-C] [-R] [-S<string>] [-O<orderfile>] [paths...]\";\n \n static int diff_output_format = DIFF_FORMAT_HUMAN;\n static int detect_rename = 0;\ndiff --git a/diff-helper.c b/diff-helper.c\n--- a/diff-helper.c\n+++ b/diff-helper.c\n@@ -7,11 +7,22 @@\n \n static const char *pickaxe = NULL;\n static int pickaxe_opts = 0;\n+static const char *orderfile = NULL;\n static int line_termination = '\\n';\n static int inter_name_termination = '\\t';\n \n+static void flush_them(int ac, const char **av)\n+{\n+\tdiffcore_std(av + 1,\n+\t\t     0, 0, /* no renames */\n+\t\t     pickaxe, pickaxe_opts,\n+\t\t     -1, /* no breaks */\n+\t\t     orderfile);\n+\tdiff_flush(DIFF_FORMAT_PATCH, 0);\n+}\n+\n static const char *diff_helper_usage =\n-\t\"git-diff-helper [-z] [-S<string>] paths...\";\n+\t\"git-diff-helper [-z] [-S<string>] [-O<orderfile>] paths...\";\n \n int main(int ac, const char **av) {\n \tstruct strbuf sb;\n@@ -131,17 +142,9 @@ int main(int ac, const char **av) {\n \t\t\t\t\t  new_path);\n \t\t\tcontinue;\n \t\t}\n-\t\tif (1 < ac)\n-\t\t\tdiffcore_pathspec(av + 1);\n-\t\tif (pickaxe)\n-\t\t\tdiffcore_pickaxe(pickaxe, pickaxe_opts);\n-\t\tdiff_flush(DIFF_FORMAT_PATCH, 0);\n+\t\tflush_them(ac, av);\n \t\tprintf(garbage_flush_format, sb.buf);\n \t}\n-\tif (1 < ac)\n-\t\tdiffcore_pathspec(av + 1);\n-\tif (pickaxe)\n-\t\tdiffcore_pickaxe(pickaxe, pickaxe_opts);\n-\tdiff_flush(DIFF_FORMAT_PATCH, 0);\n+\tflush_them(ac, av);\n \treturn 0;\n }\ndiff --git a/diff-tree.c b/diff-tree.c\n--- a/diff-tree.c\n+++ b/diff-tree.c\n@@ -397,7 +397,7 @@ static int diff_tree_stdin(char *line)\n }\n \n static char *diff_tree_usage =\n-\"git-diff-tree [-p] [-r] [-z] [--stdin] [-M] [-C] [-R] [-S<string>] [-m] [-s] [-v] [-t] <tree-ish> <tree-ish>\";\n+\"git-diff-tree [-p] [-r] [-z] [--stdin] [-M] [-C] [-R] [-S<string>] [-O<orderfile>] [-m] [-s] [-v] [-t] <tree-ish> <tree-ish>\";\n \n int main(int argc, const char **argv)\n {\n------------\n\n"},{"id":"4480","messageId":"7v1x7jq1i5.fsf_-_@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"7vis0vq1rz.fsf_-_@assigned-by-dhcp.cox.net","subject":"[PATCH 3/4] diff: Clean up diff_scoreopt_parse().","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-03T08:37:54Z","receivedAt":"2005-06-03T08:37:54Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This cleans up diff_scoreopt_parse() function that is used to\nparse the fractional notation -B, -C and -M option takes.  The\ncallers are modified to check for errors and complain.  Earlier\nthey silently ignored malformed input and falled back on the\ndefault.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n diff-cache.c      |    9 ++++++---\n diff-files.c      |   15 +++++++++++----\n diff-tree.c       |    9 ++++++---\n diff.c            |   39 +++++++++++++++++++++++++++++++++++++++\n diffcore-rename.c |   18 ------------------\n 5 files changed, 62 insertions(+), 28 deletions(-)\n\ndiff --git a/diff-cache.c b/diff-cache.c\n--- a/diff-cache.c\n+++ b/diff-cache.c\n@@ -191,17 +191,20 @@ int main(int argc, const char **argv)\n \t\t\tcontinue;\n \t\t}\n \t\tif (!strncmp(arg, \"-B\", 2)) {\n-\t\t\tdiff_break_opt = diff_scoreopt_parse(arg);\n+\t\t\tif ((diff_break_opt = diff_scoreopt_parse(arg)) == -1)\n+\t\t\t\tusage(diff_cache_usage);\n \t\t\tcontinue;\n \t\t}\n \t\tif (!strncmp(arg, \"-M\", 2)) {\n \t\t\tdetect_rename = DIFF_DETECT_RENAME;\n-\t\t\tdiff_score_opt = diff_scoreopt_parse(arg);\n+\t\t\tif ((diff_score_opt = diff_scoreopt_parse(arg)) == -1)\n+\t\t\t\tusage(diff_cache_usage);\n \t\t\tcontinue;\n \t\t}\n \t\tif (!strncmp(arg, \"-C\", 2)) {\n \t\t\tdetect_rename = DIFF_DETECT_COPY;\n-\t\t\tdiff_score_opt = diff_scoreopt_parse(arg);\n+\t\t\tif ((diff_score_opt = diff_scoreopt_parse(arg)) == -1)\n+\t\t\t\tusage(diff_cache_usage);\n \t\t\tcontinue;\n \t\t}\n \t\tif (!strcmp(arg, \"-z\")) {\ndiff --git a/diff-files.c b/diff-files.c\n--- a/diff-files.c\n+++ b/diff-files.c\n@@ -61,14 +61,21 @@ int main(int argc, const char **argv)\n \t\t\torderfile = argv[1] + 2;\n \t\telse if (!strcmp(argv[1], \"--pickaxe-all\"))\n \t\t\tpickaxe_opts = DIFF_PICKAXE_ALL;\n-\t\telse if (!strncmp(argv[1], \"-B\", 2))\n-\t\t\tdiff_break_opt = diff_scoreopt_parse(argv[1]);\n+\t\telse if (!strncmp(argv[1], \"-B\", 2)) {\n+\t\t\tif ((diff_break_opt =\n+\t\t\t     diff_scoreopt_parse(argv[1])) == -1)\n+\t\t\t\tusage(diff_files_usage);\n+\t\t}\n \t\telse if (!strncmp(argv[1], \"-M\", 2)) {\n-\t\t\tdiff_score_opt = diff_scoreopt_parse(argv[1]);\n+\t\t\tif ((diff_score_opt =\n+\t\t\t     diff_scoreopt_parse(argv[1])) == -1)\n+\t\t\t\tusage(diff_files_usage);\n \t\t\tdetect_rename = DIFF_DETECT_RENAME;\n \t\t}\n \t\telse if (!strncmp(argv[1], \"-C\", 2)) {\n-\t\t\tdiff_score_opt = diff_scoreopt_parse(argv[1]);\n+\t\t\tif ((diff_score_opt =\n+\t\t\t     diff_scoreopt_parse(argv[1])) == -1)\n+\t\t\t\tusage(diff_files_usage);\n \t\t\tdetect_rename = DIFF_DETECT_COPY;\n \t\t}\n \t\telse\ndiff --git a/diff-tree.c b/diff-tree.c\n--- a/diff-tree.c\n+++ b/diff-tree.c\n@@ -459,16 +459,19 @@ int main(int argc, const char **argv)\n \t\t}\n \t\tif (!strncmp(arg, \"-M\", 2)) {\n \t\t\tdetect_rename = DIFF_DETECT_RENAME;\n-\t\t\tdiff_score_opt = diff_scoreopt_parse(arg);\n+\t\t\tif ((diff_score_opt = diff_scoreopt_parse(arg)) == -1)\n+\t\t\t\tusage(diff_tree_usage);\n \t\t\tcontinue;\n \t\t}\n \t\tif (!strncmp(arg, \"-C\", 2)) {\n \t\t\tdetect_rename = DIFF_DETECT_COPY;\n-\t\t\tdiff_score_opt = diff_scoreopt_parse(arg);\n+\t\t\tif ((diff_score_opt = diff_scoreopt_parse(arg)) == -1)\n+\t\t\t\tusage(diff_tree_usage);\n \t\t\tcontinue;\n \t\t}\n \t\tif (!strncmp(arg, \"-B\", 2)) {\n-\t\t\tdiff_break_opt = diff_scoreopt_parse(arg);\n+\t\t\tif ((diff_break_opt = diff_scoreopt_parse(arg)) == -1)\n+\t\t\t\tusage(diff_tree_usage);\n \t\t\tcontinue;\n \t\t}\n \t\tif (!strcmp(arg, \"-z\")) {\ndiff --git a/diff.c b/diff.c\n--- a/diff.c\n+++ b/diff.c\n@@ -589,6 +589,45 @@ void diff_setup(int flags)\n \t\n }\n \n+static int parse_num(const char **cp_p)\n+{\n+\tint num, scale, ch, cnt;\n+\tconst char *cp = *cp_p;\n+\n+\tcnt = num = 0;\n+\tscale = 1;\n+\twhile ('0' <= (ch = *cp) && ch <= '9') {\n+\t\tif (cnt++ < 5) {\n+\t\t\t/* We simply ignore more than 5 digits precision. */\n+\t\t\tscale *= 10;\n+\t\t\tnum = num * 10 + ch - '0';\n+\t\t}\n+\t\t*cp++;\n+\t}\n+\t*cp_p = cp;\n+\n+\t/* user says num divided by scale and we say internally that\n+\t * is MAX_SCORE * num / scale.\n+\t */\n+\treturn (MAX_SCORE * num / scale);\n+}\n+\n+int diff_scoreopt_parse(const char *opt)\n+{\n+\tint opt1, cmd;\n+\n+\tif (*opt++ != '-')\n+\t\treturn -1;\n+\tcmd = *opt++;\n+\tif (cmd != 'M' && cmd != 'C' && cmd != 'B')\n+\t\treturn -1; /* that is not a -M, -C nor -B option */\n+\n+\topt1 = parse_num(&opt);\n+\tif (*opt != 0)\n+\t\treturn -1;\n+\treturn opt1;\n+}\n+\n struct diff_queue_struct diff_queued_diff;\n \n void diff_q(struct diff_queue_struct *queue, struct diff_filepair *dp)\ndiff --git a/diffcore-rename.c b/diffcore-rename.c\n--- a/diffcore-rename.c\n+++ b/diffcore-rename.c\n@@ -229,24 +229,6 @@ static int score_compare(const void *a_,\n \treturn b->score - a->score;\n }\n \n-int diff_scoreopt_parse(const char *opt)\n-{\n-\tint diglen, num, scale, i;\n-\tif (opt[0] != '-' || (opt[1] != 'M' && opt[1] != 'C' && opt[1] != 'B'))\n-\t\treturn -1; /* that is not a -M, -C nor -B option */\n-\tdiglen = strspn(opt+2, \"0123456789\");\n-\tif (diglen == 0 || strlen(opt+2) != diglen)\n-\t\treturn 0; /* use default */\n-\tsscanf(opt+2, \"%d\", &num);\n-\tfor (i = 0, scale = 1; i < diglen; i++)\n-\t\tscale *= 10;\n-\n-\t/* user says num divided by scale and we say internally that\n-\t * is MAX_SCORE * num / scale.\n-\t */\n-\treturn MAX_SCORE * num / scale;\n-}\n-\n void diffcore_rename(int detect_rename, int minimum_score)\n {\n \tstruct diff_queue_struct *q = &diff_queued_diff;\n------------\n\n"},{"id":"4481","messageId":"7vpsv3omtf.fsf_-_@assigned-by-dhcp.cox.net","threadId":"770","inReplyTo":"7vis0vq1rz.fsf_-_@assigned-by-dhcp.cox.net","subject":"[PATCH 4/4] diff: Update -B heuristics.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-03T08:40:28Z","receivedAt":"2005-06-03T08:40:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"As Linus pointed out on the mailing list discussion, -B should\nbreak a files that has many inserts even if it still keeps\nenough of the original contents, so that the broken pieces can\nlater be matched with other files by -M or -C.  However, if such\na broken pair does not get picked up by -M or -C, we would want\nto apply different criteria; namely, regardless of the amount of\nnew material in the result, the determination of \"rewrite\"\nshould be done by looking at the amount of original material\nstill left in the result.  If you still have the original 97\nlines from a 100-line document, it does not matter if you add\nyour own 13 lines to make a 110-line document, or if you add 903\nlines to make a 1000-line document.  It is not a rewrite but an\nin-place edit.  On the other hand, if you did lose 97 lines from\nthe original, it does not matter if you added 27 lines to make a\n30-line document or if you added 997 lines to make a 1000-line\ndocument.  You did a complete rewrite in either case.\n\nThis patch introduces a post-processing phase that runs after\ndiffcore-rename matches up broken pairs diffcore-break creates.\nThe purpose of this post-processing is to pick up these broken\npieces and merge them back into in-place modifications.  For\nthis, the score parameter -B option takes is changed into a pair\nof numbers, and it takes \"-B99/80\" format when fully spelled\nout.  The first number is the minimum amount of \"edit\" (same\ndefinition as what diffcore-rename uses, which is \"sum of\ndeletion and insertion\") that a modification needs to have to be\nbroken, and the second number is the minimum amount of \"delete\"\na surviving broken pair must have to avoid being merged back\ntogether.  It can be abbreviated to \"-B\" to use default for\nboth, \"-B9\" or \"-B9/\" to use 90% for \"edit\" but default (80%)\nfor merge avoidance, or \"-B/75\" to use default (99%) \"edit\" and\n75% for merge avoidance.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n diffcore.h       |   11 ++\n diff.c           |   18 ++++\n diffcore-break.c |  240 +++++++++++++++++++++++++++++++++++++++++++++---------\n 3 files changed, 225 insertions(+), 44 deletions(-)\n\ndiff --git a/diffcore.h b/diffcore.h\n--- a/diffcore.h\n+++ b/diffcore.h\n@@ -8,9 +8,19 @@\n  * (e.g. diffcore-rename, diffcore-pickaxe).  Never include this header\n  * in anything else.\n  */\n+\n+/* We internally use unsigned short as the score value,\n+ * and rely on an int capable to hold 32-bits.  -B can take\n+ * -Bmerge_score/break_score format and the two scores are\n+ * passed around in one int (high 16-bit for merge and low 16-bit\n+ * for break).\n+ */\n #define MAX_SCORE 60000\n #define DEFAULT_RENAME_SCORE 30000 /* rename/copy similarity minimum (50%) */\n #define DEFAULT_BREAK_SCORE  59400 /* minimum for break to happen (99%)*/\n+#define DEFAULT_MERGE_SCORE  48000 /* maximum for break-merge to happen (80%)*/\n+\n+#define MINIMUM_BREAK_SIZE     400 /* do not break a file smaller than this */\n \n struct diff_filespec {\n \tunsigned char sha1[20];\n@@ -76,6 +86,7 @@ extern void diff_q(struct diff_queue_str\n extern void diffcore_pathspec(const char **pathspec);\n extern void diffcore_break(int);\n extern void diffcore_rename(int rename_copy, int);\n+extern void diffcore_merge_broken(void);\n extern void diffcore_pickaxe(const char *needle, int opts);\n extern void diffcore_order(const char *orderfile);\n \ndiff --git a/diff.c b/diff.c\n--- a/diff.c\n+++ b/diff.c\n@@ -614,7 +614,7 @@ static int parse_num(const char **cp_p)\n \n int diff_scoreopt_parse(const char *opt)\n {\n-\tint opt1, cmd;\n+\tint opt1, opt2, cmd;\n \n \tif (*opt++ != '-')\n \t\treturn -1;\n@@ -623,9 +623,21 @@ int diff_scoreopt_parse(const char *opt)\n \t\treturn -1; /* that is not a -M, -C nor -B option */\n \n \topt1 = parse_num(&opt);\n+\tif (cmd != 'B')\n+\t\topt2 = 0;\n+\telse {\n+\t\tif (*opt == 0)\n+\t\t\topt2 = 0;\n+\t\telse if (*opt != '/')\n+\t\t\treturn -1; /* we expect -B80/99 or -B80 */\n+\t\telse {\n+\t\t\topt++;\n+\t\t\topt2 = parse_num(&opt);\n+\t\t}\n+\t}\n \tif (*opt != 0)\n \t\treturn -1;\n-\treturn opt1;\n+\treturn opt1 | (opt2 << 16);\n }\n \n struct diff_queue_struct diff_queued_diff;\n@@ -955,6 +967,8 @@ void diffcore_std(const char **paths,\n \t\tdiffcore_break(break_opt);\n \tif (detect_rename)\n \t\tdiffcore_rename(detect_rename, rename_score);\n+\tif (0 <= break_opt)\n+\t\tdiffcore_merge_broken();\n \tif (pickaxe)\n \t\tdiffcore_pickaxe(pickaxe, pickaxe_opts);\n \tif (orderfile)\ndiff --git a/diffcore-break.c b/diffcore-break.c\n--- a/diffcore-break.c\n+++ b/diffcore-break.c\n@@ -7,28 +7,58 @@\n #include \"delta.h\"\n #include \"count-delta.h\"\n \n-static int very_different(struct diff_filespec *src,\n-\t\t\t  struct diff_filespec *dst,\n-\t\t\t  int min_score)\n+static int should_break(struct diff_filespec *src,\n+\t\t\tstruct diff_filespec *dst,\n+\t\t\tint break_score,\n+\t\t\tint *merge_score_p)\n {\n \t/* dst is recorded as a modification of src.  Are they so\n \t * different that we are better off recording this as a pair\n-\t * of delete and create?  min_score is the minimum amount of\n-\t * new material that must exist in the dst and not in src for\n-\t * the pair to be considered a complete rewrite, and recommended\n-\t * to be set to a very high value, 99% or so.\n-\t *\n-\t * The value we return represents the amount of new material\n-\t * that is in dst and not in src.  We return 0 when we do not\n-\t * want to get the filepair broken.\n+\t * of delete and create?\n+\t *\n+\t * There are two criteria used in this algorithm.  For the\n+\t * purposes of helping later rename/copy, we take both delete\n+\t * and insert into account and estimate the amount of \"edit\".\n+\t * If the edit is very large, we break this pair so that\n+\t * rename/copy can pick the pieces up to match with other\n+\t * files.\n+\t *\n+\t * On the other hand, we would want to ignore inserts for the\n+\t * pure \"complete rewrite\" detection.  As long as most of the\n+\t * existing contents were removed from the file, it is a\n+\t * complete rewrite, and if sizable chunk from the original\n+\t * still remains in the result, it is not a rewrite.  It does\n+\t * not matter how much or how little new material is added to\n+\t * the file.\n+\t *\n+\t * The score we leave for such a broken filepair uses the\n+\t * latter definition so that later clean-up stage can find the\n+\t * pieces that should not have been broken according to the\n+\t * latter definition after rename/copy runs, and merge the\n+\t * broken pair that have a score lower than given criteria\n+\t * back together.  The break operation itself happens\n+\t * according to the former definition.\n+\t *\n+\t * The minimum_edit parameter tells us when to break (the\n+\t * amount of \"edit\" required for us to consider breaking the\n+\t * pair).  We leave the amount of deletion in *merge_score_p\n+\t * when we return.\n+\t *\n+\t * The value we return is 1 if we want the pair to be broken,\n+\t * or 0 if we do not.\n \t */\n \tvoid *delta;\n \tunsigned long delta_size, base_size, src_copied, literal_added;\n+\tint to_break = 0;\n+\n+\t*merge_score_p = 0; /* assume no deletion --- \"do not break\"\n+\t\t\t     * is the default.\n+\t\t\t     */\n \n \tif (!S_ISREG(src->mode) || !S_ISREG(dst->mode))\n \t\treturn 0; /* leave symlink rename alone */\n \n-\tif (diff_populate_filespec(src, 1) || diff_populate_filespec(dst, 1))\n+\tif (diff_populate_filespec(src, 0) || diff_populate_filespec(dst, 0))\n \t\treturn 0; /* error but caught downstream */\n \n \tdelta_size = ((src->size < dst->size) ?\n@@ -40,53 +70,95 @@ static int very_different(struct diff_fi\n \t */\n \tbase_size = ((src->size < dst->size) ? dst->size : src->size);\n \n-\t/*\n-\t * If file size difference is too big compared to the\n-\t * base_size, we declare this a complete rewrite.\n-\t */\n-\tif (base_size * min_score < delta_size * MAX_SCORE)\n-\t\treturn MAX_SCORE;\n-\n-\tif (diff_populate_filespec(src, 0) || diff_populate_filespec(dst, 0))\n-\t\treturn 0; /* error but caught downstream */\n-\n \tdelta = diff_delta(src->data, src->size,\n \t\t\t   dst->data, dst->size,\n \t\t\t   &delta_size);\n \n-\t/* A delta that has a lot of literal additions would have\n-\t * big delta_size no matter what else it does.\n-\t */\n-\tif (base_size * min_score < delta_size * MAX_SCORE)\n-\t\treturn MAX_SCORE;\n-\n \t/* Estimate the edit size by interpreting delta. */\n-\tif (count_delta(delta, delta_size, &src_copied, &literal_added)) {\n+\tif (count_delta(delta, delta_size,\n+\t\t\t&src_copied, &literal_added)) {\n \t\tfree(delta);\n-\t\treturn 0;\n+\t\treturn 0; /* we cannot tell */\n \t}\n \tfree(delta);\n \n-\t/* Extent of damage */\n-\tif (src->size + literal_added < src_copied)\n-\t\tdelta_size = 0;\n+\t/* Compute merge-score, which is \"how much is removed\n+\t * from the source material\".  The clean-up stage will\n+\t * merge the surviving pair together if the score is\n+\t * less than the minimum, after rename/copy runs.\n+\t */\n+\tif (src->size <= src_copied)\n+\t\tdelta_size = 0; /* avoid wrapping around */\n+\telse\n+\t\tdelta_size = src->size - src_copied;\n+\t*merge_score_p = delta_size * MAX_SCORE / src->size;\n+\t\n+\t/* Extent of damage, which counts both inserts and\n+\t * deletes.\n+\t */\n+\tif (src->size + literal_added <= src_copied)\n+\t\tdelta_size = 0; /* avoid wrapping around */\n \telse\n \t\tdelta_size = (src->size - src_copied) + literal_added;\n+\t\n+\t/* We break if the edit exceeds the minimum.\n+\t * i.e. (break_score / MAX_SCORE < delta_size / base_size)\n+\t */\n+\tif (break_score * base_size < delta_size * MAX_SCORE)\n+\t\tto_break = 1;\n \n-\tif (base_size < delta_size)\n-\t\treturn MAX_SCORE;\n-\n-\treturn delta_size * MAX_SCORE / base_size; \n+\treturn to_break;\n }\n \n-void diffcore_break(int min_score)\n+void diffcore_break(int break_score)\n {\n \tstruct diff_queue_struct *q = &diff_queued_diff;\n \tstruct diff_queue_struct outq;\n+\n+\t/* When the filepair has this much edit (insert and delete),\n+\t * it is first considered to be a rewrite and broken into a\n+\t * create and delete filepair.  This is to help breaking a\n+\t * file that had too much new stuff added, possibly from\n+\t * moving contents from another file, so that rename/copy can\n+\t * match it with the other file.\n+\t *\n+\t * int break_score; we reuse incoming parameter for this.\n+\t */\n+\n+\t/* After a pair is broken according to break_score and\n+\t * subjected to rename/copy, both of them may survive intact,\n+\t * due to lack of suitable rename/copy peer.  Or, the caller\n+\t * may be calling us without using rename/copy.  When that\n+\t * happens, we merge the broken pieces back into one\n+\t * modification together if the pair did not have more than\n+\t * this much delete.  For this computation, we do not take\n+\t * insert into account at all.  If you start from a 100-line\n+\t * file and delete 97 lines of it, it does not matter if you\n+\t * add 27 lines to it to make a new 30-line file or if you add\n+\t * 997 lines to it to make a 1000-line file.  Either way what\n+\t * you did was a rewrite of 97%.  On the other hand, if you\n+\t * delete 3 lines, keeping 97 lines intact, it does not matter\n+\t * if you add 3 lines to it to make a new 100-line file or if\n+\t * you add 903 lines to it to make a new 1000-line file.\n+\t * Either way you did a lot of additions and not a rewrite.\n+\t * This merge happens to catch the latter case.  A merge_score\n+\t * of 80% would be a good default value (a broken pair that\n+\t * has score lower than merge_score will be merged back\n+\t * together).\n+\t */\n+\tint merge_score;\n \tint i;\n \n-\tif (!min_score)\n-\t\tmin_score = DEFAULT_BREAK_SCORE;\n+\t/* See comment on DEFAULT_BREAK_SCORE and\n+\t * DEFAULT_MERGE_SCORE in diffcore.h\n+\t */\n+\tmerge_score = (break_score >> 16) & 0xFFFF;\n+\tbreak_score = (break_score & 0xFFFF);\n+\n+\tif (!break_score)\n+\t\tbreak_score = DEFAULT_BREAK_SCORE;\n+\tif (!merge_score)\n+\t\tmerge_score = DEFAULT_MERGE_SCORE;\n \n \toutq.nr = outq.alloc = 0;\n \toutq.queue = NULL;\n@@ -101,12 +173,22 @@ void diffcore_break(int min_score)\n \t\tif (DIFF_FILE_VALID(p->one) && DIFF_FILE_VALID(p->two) &&\n \t\t    !S_ISDIR(p->one->mode) && !S_ISDIR(p->two->mode) &&\n \t\t    !strcmp(p->one->path, p->two->path)) {\n-\t\t\tscore = very_different(p->one, p->two, min_score);\n-\t\t\tif (min_score <= score) {\n+\t\t\tif (should_break(p->one, p->two,\n+\t\t\t\t\t break_score, &score)) {\n \t\t\t\t/* Split this into delete and create */\n \t\t\t\tstruct diff_filespec *null_one, *null_two;\n \t\t\t\tstruct diff_filepair *dp;\n \n+\t\t\t\t/* Set score to 0 for the pair that\n+\t\t\t\t * needs to be merged back together\n+\t\t\t\t * should they survive rename/copy.\n+\t\t\t\t * Also we do not want to break very\n+\t\t\t\t * small files.\n+\t\t\t\t */\n+\t\t\t\tif ((score < merge_score) ||\n+\t\t\t\t    (p->one->size < MINIMUM_BREAK_SIZE))\n+\t\t\t\t\tscore = 0;\n+\n \t\t\t\t/* deletion of one */\n \t\t\t\tnull_one = alloc_filespec(p->one->path);\n \t\t\t\tdp = diff_queue(&outq, p->one, null_one);\n@@ -132,3 +214,77 @@ void diffcore_break(int min_score)\n \n \treturn;\n }\n+\n+static void merge_broken(struct diff_filepair *p,\n+\t\t\t struct diff_filepair *pp,\n+\t\t\t struct diff_queue_struct *outq)\n+{\n+\t/* p and pp are broken pairs we want to merge */\n+\tstruct diff_filepair *c = p, *d = pp;\n+\tif (DIFF_FILE_VALID(p->one)) {\n+\t\t/* this must be a delete half */\n+\t\td = p; c = pp;\n+\t}\n+\t/* Sanity check */\n+\tif (!DIFF_FILE_VALID(d->one))\n+\t\tdie(\"internal error in merge #1\");\n+\tif (DIFF_FILE_VALID(d->two))\n+\t\tdie(\"internal error in merge #2\");\n+\tif (DIFF_FILE_VALID(c->one))\n+\t\tdie(\"internal error in merge #3\");\n+\tif (!DIFF_FILE_VALID(c->two))\n+\t\tdie(\"internal error in merge #4\");\n+\n+\tdiff_queue(outq, d->one, c->two);\n+\tdiff_free_filespec_data(d->two);\n+\tdiff_free_filespec_data(c->one);\n+\tfree(d);\n+\tfree(c);\n+}\n+\n+void diffcore_merge_broken(void)\n+{\n+\tstruct diff_queue_struct *q = &diff_queued_diff;\n+\tstruct diff_queue_struct outq;\n+\tint i, j;\n+\n+\toutq.nr = outq.alloc = 0;\n+\toutq.queue = NULL;\n+\n+\tfor (i = 0; i < q->nr; i++) {\n+\t\tstruct diff_filepair *p = q->queue[i];\n+\t\tif (!p)\n+\t\t\t/* we already merged this with its peer */\n+\t\t\tcontinue;\n+\t\telse if (p->broken_pair &&\n+\t\t\t p->score == 0 &&\n+\t\t\t !strcmp(p->one->path, p->two->path)) {\n+\t\t\t/* If the peer also survived rename/copy, then\n+\t\t\t * we merge them back together.\n+\t\t\t */\n+\t\t\tfor (j = i + 1; j < q->nr; j++) {\n+\t\t\t\tstruct diff_filepair *pp = q->queue[j];\n+\t\t\t\tif (pp->broken_pair &&\n+\t\t\t\t    p->score == 0 &&\n+\t\t\t\t    !strcmp(pp->one->path, pp->two->path) &&\n+\t\t\t\t    !strcmp(p->one->path, pp->two->path)) {\n+\t\t\t\t\t/* Peer survived.  Merge them */\n+\t\t\t\t\tmerge_broken(p, pp, &outq);\n+\t\t\t\t\tq->queue[j] = NULL;\n+\t\t\t\t\tbreak;\n+\t\t\t\t}\n+\t\t\t}\n+\t\t\tif (q->nr <= j)\n+\t\t\t\t/* The peer did not survive, so we keep\n+\t\t\t\t * it in the output.\n+\t\t\t\t */\n+\t\t\t\tdiff_q(&outq, p);\n+\t\t}\n+\t\telse\n+\t\t\tdiff_q(&outq, p);\n+\t}\n+\tfree(q->queue);\n+\t*q = outq;\n+\n+\treturn;\n+}\n------------\n\n"},{"id":"4482","messageId":"20050603094706.GB24873@pasky.ji.cz","threadId":"770","inReplyTo":"Pine.LNX.4.21.0506011742560.30848-100000@iabervon.org","subject":"Re: I want to release a \"git-1.0\"","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-06-03T09:47:06Z","receivedAt":"2005-06-03T09:47:06Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, Jun 02, 2005 at 12:00:55AM CEST, I got a letter\nwhere Daniel Barkalow <barkalow@iabervon.org> told me that...\n> It shouldn't be hard to do one, except that locking with rsync is going to\n> be a pain. I had a patch to make it work with the rpush/rpull pair, but I\n> didn't get its dependancies in at the time.\n\nWas that the patch I was replying to recently? It didn't seem to have\nany dependencies.\n\n> I can dust those patches off again if you want that functionality included.\n> \n> The patches are essentially:\n> \n>  - make the transport protocol handle things other than objects\n>  - library procedure for locking atomic update of refs files\n>  - fetching refs in general\n>  - rpull/rpush that updates a specified ref file atomically\n> \n> At least the first would be very nice to get in before 1.0, since it is an\n> incompatible change to the protocol.\n\nI would like to have this a lot too. Pulling tags now is a PITA, and I\ndefinitively want to go in this way. So it will land at least in git-pb.\n:-) (But that's a little troublesome if you say it's incompatible\nchange.)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"4492","messageId":"Pine.LNX.4.21.0506031102490.30848-100000@iabervon.org","threadId":"770","inReplyTo":"20050603094706.GB24873@pasky.ji.cz","subject":"Re: I want to release a \"git-1.0\"","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-03T15:09:07Z","receivedAt":"2005-06-03T15:09:07Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 3 Jun 2005, Petr Baudis wrote:\n\n> Dear diary, on Thu, Jun 02, 2005 at 12:00:55AM CEST, I got a letter\n> where Daniel Barkalow <barkalow@iabervon.org> told me that...\n> > It shouldn't be hard to do one, except that locking with rsync is going to\n> > be a pain. I had a patch to make it work with the rpush/rpull pair, but I\n> > didn't get its dependancies in at the time.\n> \n> Was that the patch I was replying to recently? It didn't seem to have\n> any dependencies.\n\nThe rpush/rpull changes were at the end of a series that you were replying\nto the beginning of.\n\n> > I can dust those patches off again if you want that functionality included.\n> > \n> > The patches are essentially:\n> > \n> >  - make the transport protocol handle things other than objects\n> >  - library procedure for locking atomic update of refs files\n> >  - fetching refs in general\n> >  - rpull/rpush that updates a specified ref file atomically\n> > \n> > At least the first would be very nice to get in before 1.0, since it is an\n> > incompatible change to the protocol.\n> \n> I would like to have this a lot too. Pulling tags now is a PITA, and I\n> definitively want to go in this way. So it will land at least in git-pb.\n> :-) (But that's a little troublesome if you say it's incompatible\n> change.)\n\nThe ssh-based protocol has to change, because the current version doesn't\nhave any way of being extended. The first patch in the new set makes the\nincompatible change without adding anything new (so as to be as\nuncontroversial as possible), and now also adds a version number so that\nfuture additions should be less of a big deal. The rest of the series will\nadd the transfer of refs to the transfer mechanism and the protocol.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"}]}