{"thread":{"id":"5126","subject":"[RFC][PATCH] Branch history","startedAt":"2006-08-04T19:24:47Z","lastAt":"2006-08-05T09:30:28Z","messageCount":3,"participants":["Eric W. Biederman","Shawn Pearce"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"24739","messageId":"m1mzakpam8.fsf@ebiederm.dsl.xmission.com","threadId":"5126","inReplyTo":null,"subject":"[RFC][PATCH] Branch history","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2006-08-04T19:24:47Z","receivedAt":"2006-08-04T19:24:47Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\nThe problem:\ngit-rebase, stgit and the like destructively edit the commit history\non a branch.  Making it a challenge to go back to a known good point.\n\nrevlog and the like sort of help this but they don't address the\nissues that they capture irrelevant points and are not git-prune safe.\n\nWith current git the best technique I have found is to always make\na new branch before I would call git-rebase.\n\n\n\n\nAfter thinking about the problem some more I believe I have found\na rather simple solution to the problem of keeping branch history.\n\nFor each branch you want to keep the history of keep 2 branches.\nA normal working branch, and a second archive branch that records\nthe history of the branch you are editing.\n\nThe history can be kept simply by placing an additional commit on the\ntop of each branch.  The new commit on top of each branch will point\nto the same tree object as the previous top commit on the branch but\nit will have 2 parent commit objects.  The first parent commit object\nis the previous top commit object of the branch.  The second parent\ncommit object is the commit object on top of the previous version of\nthis branch.\n\nThe work flow is you edit a branch to your hearts comment then when\nyou get to an interesting point you commit the branch to your archive\nbranch so you can keep track of things.\n\nTo gitk and friends the archive branch looks like a series of branch\nmerges where one input branch is always the same as the merge result.\nSo all of the git tools work normally.\n\nThe implementation is trivial.\n\nThe neat thing is that it gives an immutable history of a branch that\nis actively being edited.  So if you export your archive branch people\nwill never see time roll backward.\n\n\n\nBelow is my patch to implement this idea.  Currently I am storing\nthe archive branch in .git/refs/archive/$branchname.  And calling\nthe command to commit a branch git-archive-branch.\n\nI think my initial naming is most likely lacking so suggestions\nfor something better would be appreciated.\n\nComments?\n\n\nEric\n\ndiff --git a/Makefile b/Makefile\nindex 700c77f..411ae95 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -150,7 +150,7 @@ SCRIPT_SH = \\\n \tgit-applymbox.sh git-applypatch.sh git-am.sh \\\n \tgit-merge.sh git-merge-stupid.sh git-merge-octopus.sh \\\n \tgit-merge-resolve.sh git-merge-ours.sh \\\n-\tgit-lost-found.sh git-quiltimport.sh\n+\tgit-lost-found.sh git-quiltimport.sh git-archive-branch.sh\n \n SCRIPT_PERL = \\\n \tgit-archimport.perl git-cvsimport.perl git-relink.perl \\\ndiff --git a/git-archive-branch.sh b/git-archive-branch.sh\nnew file mode 100755\nindex 0000000..00638de\n--- /dev/null\n+++ b/git-archive-branch.sh\n@@ -0,0 +1,131 @@\n+#!/bin/sh\n+\n+USAGE='[-m <message> | -F logfile] [-e]'\n+\n+. git-sh-setup\n+\n+headref=$(git-symbolic-ref HEAD | sed -e 's|^refs/heads/||')\n+headsha1=$(git-rev-parse \"$headref\")\n+archiveref=\"refs/archive/$headref\"\n+\n+\n+logfile=\n+edit_flag=\n+no_edit=\n+log_given=\n+log_message=\n+while case \"$#\" in 0) break;; esac\n+do\n+  case \"$1\" in\n+  -F|--F|-f|--f|--fi|--fil|--file)\n+      case \"$#\" in 1) usage ;; esac\n+      shift\n+      no_edit=t\n+      log_given=t$log_given\n+      logfile=\"$1\"\n+      shift\n+      ;;\n+  -F*|-f*)\n+      no_edit=t\n+      log_given=t$log_given\n+      logfile=`expr \"z$1\" : 'z-[Ff]\\(.*\\)'`\n+      shift\n+      ;;\n+  --F=*|--f=*|--fi=*|--fil=*|--file=*)\n+      no_edit=t\n+      log_given=t$log_given\n+      logfile=`expr \"z$1\" : 'z-[^=]*=\\(.*\\)'`\n+      shift\n+      ;;\n+  -e|--e|--ed|--edi|--edit)\n+      edit_flag=t\n+      shift\n+      ;;\n+  -m|--m|--me|--mes|--mess|--messa|--messag|--message)\n+      case \"$#\" in 1) usage ;; esac\n+      shift\n+      log_given=m$log_given\n+      if test \"$log_message\" = ''\n+      then\n+          log_message=\"$1\"\n+      else\n+          log_message=\"$log_message\n+\n+$1\"\n+      fi\n+      no_edit=t\n+      shift\n+      ;;\n+  -m*)\n+      log_given=m$log_given\n+      if test \"$log_message\" = ''\n+      then\n+          log_message=`expr \"z$1\" : 'z-m\\(.*\\)'`\n+      else\n+          log_message=\"$log_message\n+\n+`expr \"z$1\" : 'z-m\\(.*\\)'`\"\n+      fi\n+      no_edit=t\n+      shift\n+      ;;\n+  --m=*|--me=*|--mes=*|--mess=*|--messa=*|--messag=*|--message=*)\n+      log_given=m$log_given\n+      if test \"$log_message\" = ''\n+      then\n+          log_message=`expr \"z$1\" : 'z-[^=]*=\\(.*\\)'`\n+      else\n+          log_message=\"$log_message\n+\n+`expr \"z$1\" : 'zq-[^=]*=\\(.*\\)'`\"\n+      fi\n+      no_edit=t\n+      shift\n+      ;;\n+  esac\n+done\n+case \"$edit_flag\" in t) no_edit= ;; esac\n+\n+if test \"$log_message\" != \"\"\n+then\n+\techo \"$log_message\"\n+elif test \"$logfile\" != \"\"\n+then\n+\tif test \"$logfile\" = -\n+\tthen\n+\t\ttest -t 0 &&\n+\t\techo >&2 \"(read log message from standard input)\"\n+\t\tcat\n+\telse\n+\t\tcat <\"$logfile\"\n+\tfi\n+fi | git-stripspace > \"$GIT_DIR\"/COMMIT_EDITMSG\n+\n+case \"$no_edit\" in\n+'')\n+\tcase \"${VISUAL:-$EDITOR},$TERM\" in\n+\t,dumb)\n+\t\techo >&2 \"Terminal is dumb but no VISUAL nor EDITOR defined.\"\n+\t\techo >&2 \"Please supply the commit log message using either\"\n+\t\techo >&2 \"-m or -F option.  A boilerplate log message has\"\n+\t\techo >&2 \"been prepared in $GIT_DIR/COMMIT_EDITMSG\"\n+\t\texit 1\n+\t\t;;\n+\tesac\n+\tgit-var GIT_AUTHOR_IDENT > /dev/null || die\n+\tgit-var GIT_COMMITTER_IDENT > /dev/null || die\n+\t${VISUAL:-${EDITOR:-vi}} \"$GIT_DIR/COMMIT_EDITMSG\"\n+\t;;\n+esac\n+\n+cat $GIT_DIR/COMMIT_EDITMSG | git-stripspace > \"$GIT_DIR\"/COMMIT_MSG\n+\n+parents=\"-p $headsha1\"\n+if git-rev-parse --verify $archiveref > /dev/null 2> /dev/null; then\n+\tparents=\"$parents -p $(git-rev-parse $archiveref)\"\n+fi\n+\n+tree=$(git-cat-file commit $headsha1 | sed -n -e 's/^tree \\(.*\\)$/\\1/p') &&\n+commit=$(cat $GIT_DIR/COMMIT_MSG | git-commit-tree $tree $parents) \n+git-update-ref \"$archiveref\" $commit \n+rm -f \"$GIT_DIR/COMMIT_MSG\" \"$GIT_DIR/COMMIT_EDITMSG\"\n"},{"id":"24789","messageId":"20060805031821.GB18223@spearce.org","threadId":"5126","inReplyTo":"m1mzakpam8.fsf@ebiederm.dsl.xmission.com","subject":"Re: [RFC][PATCH] Branch history","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-08-05T03:18:21Z","receivedAt":"2006-08-05T03:18:21Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Eric W. Biederman\" <ebiederm@xmission.com> wrote:\n> \n> The problem:\n> git-rebase, stgit and the like destructively edit the commit history\n> on a branch.  Making it a challenge to go back to a known good point.\n> \n> revlog and the like sort of help this but they don't address the\n> issues that they capture irrelevant points and are not git-prune safe.\n\nHow are the points irrelevant?  Each commit/rebase/am/update-ref\nis recorded.  That's each change to the branch head.  It appears\nas though you are mainly interested in tracking across rebases,\nwhich a reflog would do, assuming you filtered the events down to\nonly those caused by rebase and ignored the others.\n\nBut yea, a reflog is not prune-safe, but it wouldn't be hard to\nmodify git-prune to also consider the reflog associated with\na ref if its using that ref as a root that must be preserved.\nAssuming anyone really wants that as a feature...\n \n> After thinking about the problem some more I believe I have found\n> a rather simple solution to the problem of keeping branch history.\n> \n> For each branch you want to keep the history of keep 2 branches.\n> A normal working branch, and a second archive branch that records\n> the history of the branch you are editing.\n\nIt would appear as though you are really only tracking rebase events,\nas everything else done on the branch is preserved since the work\nbranch is itself parent #1 for the archive branch commit.  So the\narchive branch shows every commit ever done along the main branch,\nbut also shows itself joining back quite frequently.  Further if you\narchive away the work branch without during a rebase since the last\narchive then there's really nothing happening except saving a tag\n(but as a commit!) on the archive branch.\n\nThis creates for a rather messy history, and is more-or-less what\npg does when patches get pushed onto a stack and they can't be\npushed by a simple fast-forward operation.  Reading this history\nin gitk is \"interesting\" at best.  This is the main reason I've\nbeen trying to write `tb` (a topic branch manager, fashioned after\nJunio's TO script) but I can't seem to find enough time to get it\nfinished.\n\n> The neat thing is that it gives an immutable history of a branch that\n> is actively being edited.  So if you export your archive branch people\n> will never see time roll backward.\n\nRight.  That's an interesting way of handling it, but that branch\nis also quite messy as its full of merge commits.  Although it may\nbe useful to export its going to carry along with it all of the bad\nedits and prior rebases made on that branch.  You probably wouldn't\nwant to merge that branch into a mainline, which means that branch\nis likely to be discarded at some point in the future.  When that\nhappens then nobody can track it anymore and that immutable history\njust got mutated out of existance.\n\nI think the right way to deal with these types of branches is to\npublicly publish whether or not the branch is going to be expected\nto roll backwards in time (due to a rebase type of event) then\nlet clients always update those branches during pulls, rather\nthan needing to explicitly mark them with '+' on the client side.\n\nFurther good remege tools (git-rerere on steriods) would help\nre-resolve conflicts resulting from continous rebasing.  This would\nmake it easier to maintain such a branch and carry the thing forward;\nor to leave it on its original base but to continously remerge\nit and the current mainline into a temporary working branch for\ntesting purposes.\n\nThis is largely the policy that Junio uses for the `pu` and\nthe `next` branches, as well as for the topic branches that he\ncarries for everyone else doing GIT development.  It appears to be\nworking rather well, but it certainly could be streamlined better.\nMy git-rerere2 and tb tools are an attempt to do this, but sadly\nthey aren't in a useful state yet.  Maybe because they are both\nfar more complex then what you are doing here.  :-)\n\n\nNonethless it is an interesting contribution.  Thank you for taking\nthe time to send it.\n\n-- \nShawn.\n"},{"id":"24807","messageId":"m1slkbmswb.fsf@ebiederm.dsl.xmission.com","threadId":"5126","inReplyTo":"20060805031821.GB18223@spearce.org","subject":"Re: [RFC][PATCH] Branch history","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2006-08-05T09:30:28Z","receivedAt":"2006-08-05T09:30:28Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n> \"Eric W. Biederman\" <ebiederm@xmission.com> wrote:\n>> \n>> The problem:\n>> git-rebase, stgit and the like destructively edit the commit history\n>> on a branch.  Making it a challenge to go back to a known good point.\n>> \n>> revlog and the like sort of help this but they don't address the\n>> issues that they capture irrelevant points and are not git-prune safe.\n>\n> How are the points irrelevant?  Each commit/rebase/am/update-ref\n> is recorded.  That's each change to the branch head.  It appears\n> as though you are mainly interested in tracking across rebases,\n> which a reflog would do, assuming you filtered the events down to\n> only those caused by rebase and ignored the others.\n\nIt tracks each change, it does not track the changes that humans find\ninteresting.  That can easily be a lot of noise.\n\nI don't want to see every time a head is updated any more than I want\nsingle keystroke level version control.  Way too much uninteresting\ndetail.\n\n> But yea, a reflog is not prune-safe, but it wouldn't be hard to\n> modify git-prune to also consider the reflog associated with\n> a ref if its using that ref as a root that must be preserved.\n> Assuming anyone really wants that as a feature...\n\nI do.  I also want history I can clone between repositories.\nI have times I have had to look 9 months back to see where I accidentally\ndropped a patch.\n\n>> After thinking about the problem some more I believe I have found\n>> a rather simple solution to the problem of keeping branch history.\n>> \n>> For each branch you want to keep the history of keep 2 branches.\n>> A normal working branch, and a second archive branch that records\n>> the history of the branch you are editing.\n>\n> It would appear as though you are really only tracking rebase events,\n> as everything else done on the branch is preserved since the work\n> branch is itself parent #1 for the archive branch commit.  So the\n> archive branch shows every commit ever done along the main branch,\n> but also shows itself joining back quite frequently.  Further if you\n> archive away the work branch without during a rebase since the last\n> archive then there's really nothing happening except saving a tag\n> (but as a commit!) on the archive branch.\n\nTrue.  But that is largely the wrong way to think about it.  I am\nsaving away a branch at times it is interesting to a human being.\nThere are also other tools and other methods of editing a branch\nbesides git-rebase.\n\n> This creates for a rather messy history, and is more-or-less what\n> pg does when patches get pushed onto a stack and they can't be\n> pushed by a simple fast-forward operation.  Reading this history\n> in gitk is \"interesting\" at best.  This is the main reason I've\n> been trying to write `tb` (a topic branch manager, fashioned after\n> Junio's TO script) but I can't seem to find enough time to get it\n> finished.\n\nI just took a quick look at pg, and while the mechanism may be\nsimilar I believe the goals are fundamentally different.  I am\ntrying to record the history at points human beings care about,\npg seems to do something automatically behind the scenes, with\nthe existing model.\n\nThe points I am recording the history are points at which I want a\nhuman commit message, because these are points in time meaningful to\nme.  The ideal companion would be something that could just walk my\nbranch history and pull it out.  So when generating an overview\nmessage I could easily generate a summary of how I had been editing my\npatches.\n\n\n>> The neat thing is that it gives an immutable history of a branch that\n>> is actively being edited.  So if you export your archive branch people\n>> will never see time roll backward.\n>\n> Right.  That's an interesting way of handling it, but that branch\n> is also quite messy as its full of merge commits.  Although it may\n> be useful to export its going to carry along with it all of the bad\n> edits and prior rebases made on that branch.  You probably wouldn't\n> want to merge that branch into a mainline, which means that branch\n> is likely to be discarded at some point in the future.  When that\n> happens then nobody can track it anymore and that immutable history\n> just got mutated out of existance.\n\nYes.  But it is interesting until it gets merged into mainline, and\nkeeping around in the developers own archives.  Mistakes can be\ninteresting.  I don't expect that there will be a need for keeping\nthe mistakes after a branch is perfected and merged into mainline.\nUntil the branch is perfected though I fully expect there to be bad\nbranch history edits that need to be fixed.\n\nThe point at which the immutable history goes out of existence is\nthe point where the branch stops being interesting as an entity\nin it's own right.  So I think that is exactly the right behavior.\n\n> I think the right way to deal with these types of branches is to\n> publicly publish whether or not the branch is going to be expected\n> to roll backwards in time (due to a rebase type of event) then\n> let clients always update those branches during pulls, rather\n> than needing to explicitly mark them with '+' on the client side.\n\nNot if part of the problem is distributing the work of coming up\nwith a perfect patch set.  If you don't distribute the history\nit is hard to see what someone has really changed.  You can't help\nme undo a branch editing mistake if you don't have the previous\nversion of the branch.  It is hard to verify I actually fixed what\nyou are concerned about if you don't have the old version to compare\nagainst.\n\n> Further good remege tools (git-rerere on steriods) would help\n> re-resolve conflicts resulting from continous rebasing.  This would\n> make it easier to maintain such a branch and carry the thing forward;\n> or to leave it on its original base but to continously remerge\n> it and the current mainline into a temporary working branch for\n> testing purposes.\n\nRebase is not the primary operation.  I have one basic branch that\nI have 10 copies of against v2.6.18-rc3.  Refactoring, debugging,\nand perfecting patches is a much more interesting event than rebasing.\nAlthough rebasing does happen as well.\n\nIf you look at the -mm tree it tends to have 2-3 releases before\ngetting rebased.\n\n> This is largely the policy that Junio uses for the `pu` and\n> the `next` branches, as well as for the topic branches that he\n> carries for everyone else doing GIT development.  It appears to be\n> working rather well, but it certainly could be streamlined better.\n> My git-rerere2 and tb tools are an attempt to do this, but sadly\n> they aren't in a useful state yet.  Maybe because they are both\n> far more complex then what you are doing here.  :-)\n\nTo some extent I have a very interesting subset of kernel development.\nMost of my changes are to systems that I am not a maintainer of.\nMost of my changes are substantial, and scary because they touch\nfundamental things.  Most of my change involve many interdependent\npatches, so topic branches cannot solve my problems.\n\nFor edits stgit git-rebase certainly can help, and I clearly\nanticipate better tools in that vein, as well as better tools\nfor dealing with topic branches.\n\nBut that isn't the problem I am trying to solve here.  I am trying\nto implement version control for branch edits, (with maximum\ncapability with the existing git).\n\n> Nonethless it is an interesting contribution.  Thank you for taking\n> the time to send it.\n\nWelcome. \n\nI think by making branch edit history something fundamental, we\nachieve some fairly substantial things.\n- We don't care about how the operations to edit a branch are\n  implemented, making them simpler to write.\n- We begin to allow distributed branch editing.\n- Branches become primary objects we can work with.\n\nHopefully I have stirred up the pot enough to allow some interesting\nthings.\n\nEric\n"}]}